There is a moment every agent builder eventually hits. The demo works. The happy path runs clean. You show it to someone and they nod enthusiastically. Then you deploy it to a real workflow and it starts doing something subtly wrong — consistently, silently, and in a way that is deeply hard to debug because you do not own the call stack.
I hit that wall building my Python orchestration framework. And again with the SDLC agentic harness that now drives this site's content pipeline. And again with the support-automation system I built for a previous contract. Each time I rebuilt more carefully. Each time the same engineering instincts that helped me in non-AI software turned out to apply here too.
Over three production systems, 12 factors kept surfacing as the difference between agents that held together under real load and agents that looked great until they did not. This post is my distillation of those factors — grounded in systems I actually operate, not in theory.
The Trap Most Builders Fall Into
Frameworks are optimized for getting to 70%. That is not a criticism — it is the right optimization for demos, prototypes, and proving feasibility. The problem is that teams mistake the 70% ceiling for an architecture. When the last 30% becomes a debugging session seven layers deep in abstractions you do not control, the only exit is a rewrite that should have been the first write.
The antidote is not a better framework. It is treating agents as software — first-class software with explicit state, explicit control flow, and explicit prompt ownership. That is what these 12 factors amount to.
Factor 1: JSON Extraction Is the Core Primitive
The single most reliable thing an LLM does is convert natural language into structured data. Everything else — tool calls, multi-step reasoning, routing decisions — is downstream of that one capability.
In my Python orchestration framework, every agent interaction bottoms out here. The LLM's job is to produce a JSON decision. The orchestrator's job is to route that decision to code.
Factor 2: Own Every Token of Your Prompts
Prompt abstractions that auto-construct your system message will get you 80% of the way quickly. When you need the remaining 20% — when the model is doing something subtly wrong on a specific edge case — you need to control every token.
In my SDLC harness, I maintain a prompts/ directory. Every workflow stage has its own prompt file, versioned alongside code. When a review agent started producing overly permissive verdicts, I could diff the prompt, tighten the language, and trace the behavior change. You cannot do that when your framework assembles the prompt from a dozen internal abstractions.
Factor 3: Context Windows Require Active Management
Context accumulation is the silent killer of long-running agents. In my support-automation system, early versions would append every tool result and intermediate reasoning step to the running context. By step 15 of a complex triage flow, the model was reading irrelevant noise from step 3.
The fix is explicit context management: summarize old events, keep recent tool results, prune intermediate reasoning. Treat the context window like working memory — curate what stays, what gets summarized, and what gets dropped.
Factor 4: "Tool Use" Is JSON Routing, Nothing More
The framing of "tool use" encourages magical thinking. An agent does not use a tool — it produces JSON, and your code routes that JSON to a function. Demystifying this makes debugging straightforward.
In my Python orchestration framework, every agent output goes through a single dispatch function. There is no reflection, no auto-discovery, no framework magic:
When something breaks, there is exactly one place to look.
Factor 5: Small Focused Agents Outperform Monoliths
My SDLC harness runs around 50 discrete task agents across a spec. Each task agent handles one narrow slice: implement a task, run validation, write a review report. No single agent orchestrates the whole spec. The block-level orchestrator coordinates them, but its job is sequencing and merging — not running the content work itself.
This keeps each agent's context window focused. The implement agent does not carry review logic. The review agent does not carry implementation history. Focused context means fewer confabulations and much cleaner error traces when something goes wrong.
Factor 6: Own Your Control Flow
Frameworks that hide the agentic loop — the "while not done: call LLM, execute result" cycle — make it impossible to add checkpoints, timeouts, or human gates without fighting the abstraction. In my systems, the loop is always explicit:
Every exit condition is visible. Every checkpoint is explicit. The loop is yours.
Factor 7: Agents Should Be Stateless — Your Application Is Not
An agent function should take state in and return new state out. It should not own a database connection, a running session, or mutable internal memory. The application layer — the orchestrator — owns persistence.
This is how my SDLC harness can interrupt an agent mid-run, save the state to disk, and resume it in a new process. The agent itself knows nothing about persistence:
Immutable, serializable, resumable, and testable in isolation.
Factor 8: Human Escalation Is a First-Class Action
In my support-automation system, the most important action the triage agent could take was asking a human. Not as a failure mode — as a valid, expected outcome for ambiguous tickets.
Making human escalation a first-class action in the dispatch table means you can test it, log it, and measure it like any other action. It also means the agent will use it naturally rather than confabulating a confident answer when confidence is not warranted.
Factor 9: Meet Users in Their Existing Workflows
The support-automation system integrated with the ticketing system's existing interface — no new dashboard, no new login. Agents that require users to adopt a new tool face a UX adoption problem on top of the engineering problem. Route agent outputs back to wherever the user already works.
Factor 10: Handle Errors Explicitly — Do Not Append and Hope
Early in my orchestration framework, my error-handling strategy was "append the error to context and let the model figure it out." This works for simple cases and fails badly for anything involving rate limits, network timeouts, or repeated tool failures. The model starts reasoning about errors rather than taking productive actions.
Explicit error handling means categorizing errors at the dispatch layer, deciding programmatically whether to retry, fall back, or surface to a human, and keeping error details out of the reasoning context unless they are directly actionable.
Factor 11: Separate Execution State From Business State
Running an agent involves two distinct state concerns that should not be mixed:
- Execution state — step counter, retry count, context window, timing
- Business state — what has actually been accomplished, outputs produced, human interactions taken
In my SDLC harness, the orchestrator tracks execution state. The task record in the planning directory tracks business state. They never share a data structure. This makes replays and debugging straightforward — you can re-run an agent with the same business state and a fresh execution state without unintended side effects.
Factor 12: Push the Model to Its Limits, Then Engineer Reliability
The work that creates real value sits right at the boundary of what the model can do reliably. My support-automation triage handled ambiguous, multi-intent tickets where a naive classifier would produce garbage. The SDLC review agent catches subtle acceptance-criteria mismatches that require genuine semantic understanding.
That boundary shifts as models improve. The engineering task is finding it and building reliability scaffolding around it — structured prompts, validation loops, human escalation paths — so the capability becomes dependable rather than impressive-sometimes.
Putting It Together
These 12 factors did not come from a whitepaper. They came from the same place all engineering instincts come from: debugging production failures, watching something that worked in a demo collapse under real load, and iterating.
The good news is that none of this requires special AI knowledge. If you have shipped production software before, you already understand state management, control flow, explicit interfaces, and the cost of hidden abstractions. Agents are software. The same instincts apply.
Start with Factor 1 and Factor 2 — own your JSON extraction and own your prompts. The other factors follow naturally once you stop treating the LLM as a black box and start treating your agent as a system you are responsible for.
The 12-factor framework serves as the spine of the 12-factor agent development learning path on this site, where each factor gets a deeper treatment with worked examples from my real systems. If you want to build something that holds together in production, that is a good place to continue.