12-Factor Agents: From AI Prototype to Software Factory
By Arved Cornelsen ·
Most AI agents never make it past the demo. They reach 70–80 % quality, impress in a workshop and then stall, because the last 20 % decide whether customers can trust them. Dex Horthy’s 12-Factor Agents is the most practical answer to that problem I have seen so far. In this article I summarise the twelve factors, and then go one step further: which organisational forms make them work, and how to combine them into a software factory that produces reliable software with agents and humans on the same assembly line.
Key takeaways
- Good agents are mostly ordinary software. The reliable ones are deterministic workflows with small, well-placed LLM steps, not a model looping freely over a bag of tools.
- Ownership beats frameworks. Own your prompts, your context window and your control flow. Frameworks get you to 80 %; ownership gets you to production.
- Humans are part of the design, not a fallback. Approval points, escalation and feedback are modelled as explicit steps of the agent.
- Organisation follows architecture. Small, focused agents need stream-aligned teams that own them, and a platform team that makes running them cheap and safe.
- A software factory is the end state: a deterministic pipeline in which specialised agents and people each do the part they are best at, measured with the same rigour as any production line.
Why most agents stall at 80 %
Software has always been a graph: first flowcharts, then workflow engines and DAG orchestrators, later pipelines with a few machine-learning steps. The promise of agents was that we could throw the hand-drawn graph away: give the model a goal and a set of tools, and let it find the path at runtime.
The pattern behind most agents is simple: the model chooses the next step as structured output, deterministic code executes it, the result is appended to the context, and the loop repeats until the model says it is done. Horthy breaks every agent down into four parts: a prompt, a switch statement, an accumulated context and a loop.
In practice this naive loop breaks down. After a few dozen turns the context gets long, the model loses track and repeats failed attempts. And even 90 % reliability is not good enough for a customer-facing product. Nobody would accept a web app that crashes on every tenth page load. Horthy talked to more than 100 SaaS teams and kept seeing the same story: pick a framework, reach 70–80 % quickly, then spend months reverse-engineering the framework to get further, and often start again from scratch.
His conclusion: the agents that work in production are micro agents: small loops embedded in a larger, mostly deterministic workflow. His team’s own deployment bot is a good example. Code merges, deploys to staging and runs the tests; the agent only proposes the production deployment, and a human approves or corrects it in plain language (“deploy the backend first”). The LLM’s real value is understanding that feedback and adjusting the plan.
The 12 factors, grouped into four disciplines
The guide lists twelve factors plus one honourable mention. For leadership conversations I find it helpful to group them into four disciplines.
Discipline 1: Own the input
Factor 1: Natural language to tool calls. The atomic building block of every agent: turn a request such as “create a payment link for this sponsor” into a structured object that describes an API call, and let ordinary code execute it. Everything else builds on this translation step.
Factor 2: Own your prompts. Black-box abstractions (“role, goal, tools, agent.run()”) are a fine start, but you cannot tune what you cannot see. Treat prompts as code: versioned, reviewed, tested with evaluations and improved deliberately.
Factor 3: Own your context window. A language model is a stateless function. At every step it only knows what you put into its context. That makes context engineering the core skill: which instructions, documents, previous steps, errors and memories go in, in what format, and what stays out (for example sensitive data or errors that are already resolved). You are not bound to the standard chat-message format; a compact, structured event history often works better.
Factor 13 (bonus): Pre-fetch what you will need. If the model will almost certainly call a tool (listing the latest release tags before a deployment, for instance), call it yourself up front and put the result into the context. You save a round trip, reduce the number of decisions the model can get wrong, and let it focus on the hard part: what to do with the data.
Discipline 2: Own the execution
Factor 4: Tools are just structured outputs. A “tool call” is nothing magical: it is JSON that your code interprets. The model decides what should happen; your code decides how. That separation is what makes behaviour testable and safe.
Factor 8: Own your control flow. Write the loop yourself. Some steps can run immediately and feed their result back; others (anything slow, expensive or risky) should save the state, notify a human or start a background job, and stop. Owning the loop is also where you add logging, tracing, rate limits, caching, summarisation or an LLM-as-judge check. Horthy’s most important wish for frameworks: the ability to pause between the model choosing a tool and the tool being executed. That is exactly the moment where a human approval belongs.
Factor 9: Compact errors into the context window. Models are surprisingly good at fixing their own mistakes when they see the error. Catch the exception, add a concise description of it to the context and try again, but count consecutive failures and escalate to a human (or reset the context) after about three attempts, so the agent does not spin out.
Discipline 3: Own the state
Factor 5: Unify execution state and business state. Instead of tracking “which step are we in” separately from “what has happened”, keep a single event history (a thread) as the source of truth. Most execution state can be derived from it. The benefits are substantial: easy to store and resume, easy to debug, easy to show in a UI, and you can even fork a thread to explore alternatives.
Factor 6: Launch, pause and resume with simple APIs. Agents are programs. Users, systems and other agents should be able to start them with a simple call; long-running work pauses and is resumed by an external event such as a webhook, without bespoke integration each time.
Factor 12: Make your agent a stateless reducer. The elegant consequence of factors 3, 5 and 8: an agent step is a pure function that takes the current history plus a new event and returns the next state, like a reducer in a front-end application. Stateless steps are easy to scale, replay and test.
Discipline 4: Own the human interface and the scope
Factor 7: Contact humans with tool calls. Asking a person for input or approval is modelled as a tool like any other (“request human input”, with a question, context and options). The agent saves its state, notifies the human and stops; the answer arrives later and resumes the thread. This enables what Horthy calls outer-loop agents: agents that are started by events or schedules and reach out to people when they need them, instead of people always starting the conversation.
Factor 10: Small, focused agents. The longer the task, the longer the context, and the worse the results. Scope each agent to a handful of steps (roughly three to ten, rarely more than twenty) inside a deterministic workflow. As models improve, you widen the scope step by step, much like refactoring a large codebase.
Factor 11: Trigger from anywhere, meet users where they are. Once agents can pause and contact humans, they can be triggered from and answer through Slack, e-mail, ticket systems or monitoring alerts. They start to behave like digital colleagues. And because a human can be pulled in quickly, you can trust them with higher-stakes actions.
What this means for the organisation
The twelve factors are written for engineers. But almost every one of them has an organisational counterpart, and in my experience that is where most AI initiatives actually fail.
Conway’s law applies to agents, too
Systems mirror the communication structures of the organisations that build them. A central “AI lab” tends to build one big, general agent: precisely the architecture the 12 factors warn against. Small, focused agents embedded in real workflows need teams that own those workflows end to end.
Using the vocabulary of Team Topologies (Skelton & Pais), the structure that fits the factors looks like this:
- Stream-aligned teams own their micro agents. The team that owns the order process owns the agent that drafts order confirmations. They know the domain, the edge cases and the customers, which is exactly the context the agent needs (factors 3 and 13). Agents are features of their product, not a separate project.
- A platform team provides the agent runtime. Thread storage, launch/pause/resume APIs, human-approval channels, triggers from Slack or e-mail, tool registry, tracing, cost controls and evaluation pipelines (factors 5, 6, 7, 11). Every stream-aligned team needs these, and nobody should build them twice. The platform’s job is to make the right way the easy way.
- An enabling team spreads the craft. Prompt-as-code, context engineering and evaluation are new skills. A small enabling team (or a well-run guild) coaches teams for a few weeks, then moves on, instead of becoming a bottleneck that “does AI” for everyone.
- A complicated-subsystem team only where it is truly needed, for example for retrieval over a large knowledge base or for fine-tuned models.
Governance as part of the design
Factor 7 and factor 8 turn governance into an engineering decision. Classify tools by risk: read-only tools run freely, reversible actions are logged, and irreversible or customer-facing actions require an approval step. That classification is a decision for the business, not for the model, and it should be made explicitly, documented and reviewed, the same way you decide who may sign a contract. It is also the basis for meeting regulatory requirements such as the EU AI Act with confidence rather than paperwork.
Roles change, they do not disappear
Product owners design where humans step in and what “done” looks like for an agent. Engineers spend less time writing every line and more time designing workflows, writing evaluations and reviewing agent output. Domain experts become the approvers and teachers of the agents in their area. Leadership’s job is to make this shift explicit, with training, clear ownership and room to experiment.
Building a software factory
Combine small, focused agents, a shared platform and clear approval points, and you arrive at something larger than a collection of agents: a software factory. Like a modern production line, it is a deterministic pipeline of stations. Some are operated by agents, some by people, and every hand-over is defined, measured and auditable.
The stations
- Intake. An agent triages incoming requests, bug reports and alerts, enriches them with context (affected customers, recent changes, related tickets) and routes them, triggered from anywhere (factor 11).
- Specification. An agent drafts a technical plan with all the context pre-fetched (factor 13); a human approves or corrects it in plain language. This is the cheapest place to catch a wrong direction.
- Implementation. Coding agents work on small, well-scoped tasks (factor 10) on their own branches: one change, one pull request.
- Verification. Deterministic checks first: tests, linting, security scans, performance budgets. Then evaluation agents and LLM-as-judge checks. Failures are fed back as compact errors (factor 9); repeated failures escalate to a person.
- Approval. A human reviews and merges. For low-risk changes this can be lightweight; for high-risk ones it is a deliberate decision (factors 7 and 8).
- Release. A deployment workflow promotes the change step by step with approval gates, much like the deployment bot from the guide.
- Operations. Outer-loop agents watch production, react to alerts and open the next request, closing the loop back to intake.
The foundations
- One event history per work item (factors 5 and 12): every decision, tool call, error and approval is traceable, from request to production.
- Prompts, tools and policies as code in the repository, reviewed and versioned like everything else (factors 2 and 4).
- Evaluations in the CI pipeline. A prompt change that makes results worse should fail the build, just like a failing unit test.
- Observability and cost control per station: tokens, latency, error and escalation rates.
- A tool registry with risk classes that defines which tools may run autonomously and which require approval.
Measure it like a production line
Keep the established delivery metrics: lead time, deployment frequency, change failure rate and time to restore. Add agent-specific ones: the share of work items completed without human correction, the escalation rate per station, rework after approval and cost per change. The goal is not maximum autonomy; it is predictable quality at lower cost and higher speed. Autonomy grows station by station as the numbers justify it.
A pragmatic 90-day path
- Days 1–30: Pick one painful, well-understood workflow. Build one micro agent inside it, with an explicit approval step. Put prompts under version control and write the first evaluations.
- Days 31–60: Extract what you built into a small platform: thread storage, approvals, triggers, tracing. Bring in a second team on the same foundation.
- Days 61–90: Connect stations into a first factory slice (for example intake → specification → implementation for a class of small changes) and start reporting the metrics above to leadership.
Conclusion
The 12-Factor Agents are not an argument against AI agents. They are an argument for treating them as serious software. Own the inputs, the execution, the state and the human interface; keep agents small; and design the organisation around them deliberately. Do that consistently and the step from the first useful agent to a software factory is evolution, not revolution.
If you are working on your organisation’s AI transformation and want to compare notes, let’s talk.
Sources & credits: The twelve factors are by Dex Horthy (HumanLayer), published as 12-Factor Agents under CC BY-SA 4.0; summarised and paraphrased here. The organisational and software factory sections are my own perspective.
- AI transformation
- AI agents
- Software factory
- Engineering leadership
- Team Topologies