← All articles
AI Agents6 min read

We stopped building agents as chains and started building them as systems

Stop building AI agents as fragile chains. Learn the systems approach: explicit state models, tool contracts, and loop budgets that survive production.

We stopped building agents as chains and started building them as systems

Most agent tutorials show a chain: prompt → tool → prompt → answer. In production, chains break the moment reality gets messy. A tool returns garbage, the model misreads it, and the whole chain collapses because nothing was designed to recover.

We build agents as systems instead. That means three things most tutorials skip.

First, every agent gets an explicit state model. Not hidden conversation history — an actual schema of what the agent believes: the goal, what's been tried, what's confirmed, what's still unknown. When a tool call fails, the agent doesn't re-prompt from scratch; it updates the state and picks the next action from it. This one change cut our unrecoverable failure rate roughly in half on a document-processing agent we shipped last year.

Second, tools are designed as contracts, not conveniences. Each tool declares what it accepts, what it returns, and — critically — its failure modes. A search tool that returns "no results" is different from one that times out, and the agent's policy should treat them differently. We write these failure modes down before writing a line of agent code, because the model can only handle failures it can distinguish.

Third, we budget for the loop. An agent that can call tools 50 times will call tools 50 times. We set explicit step budgets, cost budgets per run, and escalation rules: after N failed attempts, the agent writes a structured summary of what it tried and hands off to a human. That summary is a feature, not a failure — it's the difference between an agent that wastes money silently and one that fails loudly and usefully.

The uncomfortable truth: most "agent frameworks" optimize for demos. Demos have one happy path. Production has forty. If you're evaluating an agent architecture, ask how it handles the thirty-ninth path — that's where your users live.

We wrote up how we build AI agents to survive contact with real users, and the state-model pattern is the piece clients ask about most.

Field note: On that document-processing agent, the state schema was dead simple — `goal`, `attempts[]` (each with action, result, and a one-line assessment), and `confirmed_facts[]`. The recovery rule was equally simple: on any tool failure, the agent had to write what it learned before choosing the next action. Forced articulation turned out to be the mechanism — models that must state the failure mode pick better recoveries. We also capped the loop at 12 steps with a mandatory human handoff summary at the cap. In three months of production, the agent hit the cap 4% of the time, and every handoff summary was actionable enough that reviewers resolved the case in under ten minutes. The system didn't eliminate failure; it made failure cheap and legible.

Building something worth shipping?

We take on a small number of AI product engagements. Tell us what you are building — we reply within 48 hours.