AI ENGINEER DEX
A FIELD GUIDE FOR PEOPLE BUILDING WITH AGENTS

How to build with AI agents, and who to learn it from.

Best practices sourced from experts and companies around the web and curated by a human. All in one place.

40Practices
7Pillars
13Retired
2026 Q3Last review
HOW JUDGMENTS ARE MADE

Four rings

Practices can't be benchmarked — the task varies too much. These are stated judgments against published criteria, revisited quarterly, with every change logged on the practice itself.

Standard

Reported working by sources that didn't get it from each other, at least one with no commercial stake in it, and nothing credible arguing against.

Worth trying

Working somewhere real, but outcomes depend on stack, model or team. Trade-offs listed on the practice.

Watch

One source, or very recent. Worth recording — perhaps a little soon to recommend.

Retired

Was correct, isn't now. Kept with the date, the reason, and whatever replaced it.

Independence is the test, not headcount. Two companies that both read the same post count once — if they cite each other, they're one source. A vendor documenting its own product counts as documentation, not adoption.

THE INDEX

Where to start

The Standard-ring practices — independently corroborated, with nothing credible arguing against them. Each one links to its sources and the organisations running it.

VIEW ALL PRACTICESShowing 6 of 40 — search and filter by ring, pillar, source and organisation

Attribution splits three ways. Documented by is the source that wrote it down.Adopted by is the organisations running it. Popularised by is whoever made you hear about it. Collapsing them is the fastest way to lose the trust of everyone involved.

PRIMARY SOURCES

Companies that publish how they work

Dated, accountable, carrying production numbers. Indexed here because nobody else indexes them together.

Anthropic

anthropic.com

The single richest source for agentic coding and context.

24 practices cite them

Cloudflare

cloudflare.com

Published their internal AI engineering stack end to end.

7 practices cite them

Context engineering from a shipping agent product.

6 practices cite them

Model-specific guidance that changes materially between releases.

6 practices cite them

HumanLayer

humanlayer.dev

12-factor-agents and the Agent Control Plane.

5 practices cite them

Parlance Labs

hamel.dev

The reference for LLM evaluation.

4 practices cite them
RETIRED

Practices that stopped being true

Kept with a date, a reason, and whatever replaced it — including when nothing did.

VIEW ALL RETIRED3 of 13 shown

Two ways a practice retires. Superseded — something replaced it. Obsolete — the constraint it worked around is gone, whether or not anything took its place.

CONTRIBUTE

Submit a practice

Credit for a practice goes to whoever wrote the source, not to whoever found it.

NOT OPEN YET

Submissions aren't open yet — the form that used to sit here didn't send anywhere, so it has been removed rather than left looking functional. Until it's built, send a practice and its source to @logicalicy and it will be read.

STRAIGHT ANSWERS

How this works

The questions a sceptical reader should be asking.

Who decides what goes in, and what ring it gets?

One person — me. Every practice is read, checked against its sources, and given a ring by hand. Submissions come in through the form and get curated, not auto-published.

That's a limitation and it's stated on purpose. A single curator is a bottleneck and a bias. The mitigation isn't pretending otherwise, it's publishing the criteria and logging every ring change with a date, so you can see whether the judgment has a track record worth trusting.

Why isn't there a score or a benchmark?

Because practices can't be benchmarked honestly. A benchmark holds the task fixed and varies the model. Here the task varies infinitely — isolating subagent context helps on a refactor and hurts on an open-ended research loop — and the practice is tangled up with your whole harness, your codebase and your reviewer.

Anyone showing you a number for this is either measuring one narrow task or making it up. A stated judgment you can argue with is more useful than a fake metric you can't.

What makes something Standard rather than Worth trying?

Independent corroboration, not headcount. Two companies that both read the same post count as one source. If they cite each other, they're one source. A vendor documenting a practice for its own product counts as documentation, not as adoption.

Standard means: reported working by sources that plausibly reached it separately, at least one of which has no commercial stake in it, with nothing credible arguing against. Field reports count toward that — which is how a practice with no corporate adopters can still get there.

Isn't this just an awesome-list with better CSS?

The overlap is real. The differences are that entries carry a date and a judgment, that the judgment changes and the change is logged, and that things which stop being true get moved to Retired with a reason instead of quietly sitting there being wrong.

If that discipline slips, this is an awesome-list with better CSS. The Retired section is the honest test of whether it's holding.

How often is any of this checked?

Quarterly. Every practice gets its sources re-read and its ring reconsidered against whatever models are current. The review date is on every card, so you can see how stale a given judgment is rather than taking 'regularly updated' on faith.

Can I submit my own work?

Yes, and say so when you do — self-submissions are labelled, not rejected. Plenty of good practices come from the person who worked them out.

What doesn't get in is a practice with no source anyone else can open. If the only evidence is that it worked for you, that's a field report, and there's a place for those on each practice.

Does anyone pay to be listed?

No. Nobody pays to appear, to move up a ring, or to stay out of Retired. If that ever changes it will be disclosed on the practice itself, not buried here.

You've got a practice wrong. How do I argue with you?

Send the counter-evidence — a source, a field report, a link to someone credible disagreeing. Disagreement gets written into the practice rather than resolved silently, so a well-argued objection is visible on the page even when it doesn't change the ring.

If a source has been misattributed, that gets fixed the same day.

FROM THE MAKER

I got tired of not knowing which advice was still current.

[M]

Hi, I'm Mario. I build software on my own, mostly with agents, and I kept hitting the same problem: a technique someone swore by six months ago quietly stopped being true, and nothing anywhere said so. The posts stay up. The videos stay up. Nobody goes back and marks them.

So this is the thing I wanted to exist — every practice with the source attached, a date on the judgment, and a section for the ones that used to work. I read every entry myself. If something here is wrong, tell me and I'll fix it in public.

I build it in the open, at @logicalicy.