Three days before launch, someone from security asks a fair question. If a document your agent reads told it to email a customer list outside the company, would you know?

The room usually splits into two answers. One person says the model would never fall for that. Another says there are logs. Neither is an answer, and everybody can tell.

The uncomfortable part is not that the answer is no. It is that it cannot be changed in three days. Whether that event is visible was decided when tracing was written or skipped, whether the retrieval path was isolated from the tool credentials, whether anyone owns the system by name. All settled months earlier by people not thinking about readiness at all.

I spent the final six days of a sixty-day series on this, and the argument that came out is narrower than "be careful in production": the things that make an AI system operable are properties you buy cheaply before launch or expensively and partially afterwards, and every one exists so the system can account for itself.

You cannot investigate a system that wrote nothing down

A user reports a bad outcome from yesterday and someone asks what the system actually did. If the honest answer is that you have the final response and a timestamp, the investigation is over before it starts.

The distinction that matters is between logging and tracing. Logging the answer tells you what came out. A trace connects the whole causal chain for one request under one identifier: the input, the retrieval queries and which sources ranked, the prompt the model actually saw, the tool calls and their arguments, what passed between agents, who owned the decision.

Without that, incidents get described in adjectives. The model hallucinated. It went off the rails. With it, the same incident becomes a location: the 2024 policy entered context at step three because the freshness filter never applied to that source. Different sentence, different owner, different fix.

There is a real tension to resolve, because traces are useful in proportion to their detail and prompts are full of personal data. Redact, store references rather than payloads where you can, and set retention that satisfies both debugging and privacy. Deciding this late usually means deciding it badly in one direction or the other.

Every component can pass while the system fails

The characteristic failure of a multi-component AI system is that every part is green and the user is still unhappy. The retriever found correct facts. The writer produced clean prose. The handoff between them dropped the one constraint that made the answer usable. That failure belongs to neither component, so it appears on neither scorecard, by construction rather than oversight.

Component metrics tell you where a fault is. System metrics tell you whether there is one. Most teams only build the first, because the first is easier to attribute. The ones worth adding above it: task success against the outcome the user needed, grounding of claims in gathered evidence, coordination quality across handoffs, efficiency per completed task rather than per call, policy adherence, and human intervention rate.

Intervention rate deserves its own dashboard. If the rate at which humans rescue runs holds steady over time, the system is merely running rather than maturing. A rising rate on a stable workload is also the earliest drift signal you will get, usually moving before any quality metric.

And a correct answer can still be a failure. An accurate, well-grounded response that took ninety seconds when the user needed five did not do the job.

The documents your agent reads were written by strangers

Somewhere in the corpus your system retrieves from today, someone may have written an instruction aimed at your agent rather than a reader.

This is not exotic. Every document an agent reads is input somebody else controls: a web page, an emailed PDF, a wiki anyone can edit, a support ticket a customer typed. Reading the world is the normal operating condition, which makes this the normal attack surface.

It is architectural rather than a prompt bug, because to the model everything in the context window is tokens. Your system instructions and an attacker's sentence inside a retrieved page arrive with the same standing. A successful injection does not even look like a bug. It looks like the system helpfully following instructions, which happen to be the wrong ones.

The defence people reach for first is detection, and detection is exactly the capability the attacker is targeting. Useful layer, poor foundation. What holds regardless of what the model believes are the controls that shrink blast radius: least privilege, so an injected instruction reaches only what the agent's credentials reach; approval gates on irreversible actions, so the intruder can request but never approve; isolation, so the component reading arbitrary documents is not the one holding write credentials.

Then go and test it. Plant a hostile instruction in a document your system will retrieve and watch. Fifteen minutes, almost nobody has done it, and the result is either reassuring or extremely informative.

A policy that no control enforces is a document

Ask who owns an AI system you have in production. Not a team, a person, the one who answers when it misbehaves. If the room goes quiet, that is an architecture gap that appears on no diagram.

Governance sounds like paperwork and decomposes into engineering artefacts. Policy, written before the incident that needed it. Approval on the record, or nobody can say who decided it was ready. Audit, which is a data-retention choice made up front or not at all. Privacy, meaning whose data flows through and how it exits on request. Incident response, meaning who gets paged and what gets switched off.

The failure mode worth naming: a policy sentence that no technical control enforces. "The system must not access customer data without consent" is a sentence. If what enforces it is that the prompt asks nicely, you have documentation, not governance. The two look identical until an audit or an incident, which is exactly when the difference gets expensive.

Each of these costs roughly an order of magnitude more after launch, and the retrofitted version is usually partial in ways nobody notices until a real request tests it.

Production AI is a dozen components pretending to be one product

Inventory what a live system actually runs and the model turns out to be one item among many. Model routing with fallbacks. Retrieval with permissions and freshness enforced at query time. Tools with scoped credentials and logged calls. Memory with retention rules. Tracing on every step. Injection defence at every boundary where outside content enters. Evaluation gates in front of every change. Budgets with enforced ceilings. Named humans holding the keys.

None of it is exotic, which is why the gap between a demo and a production system keeps surprising people. The demo exercises the happy path through three components. Production exercises every path through twelve.

The diagnostic takes an hour. Trace one real request end to end and ask, at each layer: who owns this, what monitors it, what happens if it fails. The layer with no answer is your weakest, and it is rarely the model. Usually it is retrieval freshness, tool permissions, or evaluation, most often evaluation, because it is the only layer whose absence produces no symptom until something else breaks.

Readiness is a question per layer, not a checklist at the end

Intimidating architectures talk people out of starting, so say it plainly: nobody builds all of this at once. Every control in a mature system got added because a failure taught the team why it existed. The teams that suffer least simply let staging teach the lesson instead of production.

So treat readiness as a question you can ask of any layer at any time, rather than a phase you enter three days before launch. What behaviour is the model responsible for. Can the system find current, authorised evidence. Which actions are allowed, logged and reversible. Where is autonomy bounded. How would you know it is improving. Who owns the risk.

Answer those for the architecture you have now, which should be the smallest one you can operate safely, and add layers when the workflow earns them. The mistake is treating the answers as a shopping list. Retrieval, tools, memory, agents and evaluation bolted on as separate initiatives produce components nobody coordinates and an outcome nobody owns.

What the six add up to

One argument: an AI system is production-ready when it can account for itself, and every property that makes that true has to be designed in while it is still cheap.

Tracing is the system explaining what it did. System-level evaluation is it demonstrating whether it works. Injection defence is it keeping straight which inputs may give it orders. Governance is a named human able to answer for it. Four disciplines on the org chart, one requirement seen from four angles, which is why teams that skip one usually turn out to have skipped all of them.

It is also why this is the last chapter rather than an appendix. The earlier ones are about making an AI system capable. This one is about making it operable, and capability without operability works impressively right up until the first time you have to explain it.

Reliability comes from architecture, operations and ownership, not from one great prompt. The prompt was never the product. The system is.

Read the chapter in full

Each is a short standalone post on my blog, with the specifics and what to actually do:

They are the closing chapter of 60 Days of Production AI Systems, a series on building AI that survives real users. Arriving at the end first is a reasonable way in, because it tells you what everything earlier was building towards. Start Here lays out all sixty in ten chapters.

Architecture pattern · · 1 min read

Pattern: the outbox for agent actions

Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.