New here? Start this series at Part 1: Autonomy is a runtime decision, not a model decision
- Part 1: Autonomy is a runtime decision, not a model decision
- Part 2: Your tool registry is an access control list wearing a different name
- Part 3: Four kinds of agent state, and only one of them is memory
- Part 4: A retried agent job is a second chance to send the same email
- Part 5: The 3am agent run that nobody is watching
A nightly summarisation job ran for eleven days against an expired credential. Every run "succeeded": the tool call returned an error string, the model wrote a graceful paragraph explaining that no data was available, and the job exited zero. The first person to notice was a customer asking why their weekly digest had been empty since the start of the month.
Interactive agents have a user watching them. Scheduled ones have nobody, and they inherit none of the reflexes that come with someone staring at a screen.
Recurring runs fail in their own particular way
Two failure modes dominate, and neither raises an exception.
The first is the silent no-op. A degraded run still produces output, because producing plausible output under bad inputs is precisely what these models are good at. Exit codes are therefore useless as a health signal. Alert on outcomes instead: rows written, records touched, tokens spent, tools called. A summarisation job that made zero tool calls last night is broken regardless of what it returned.
The second is the unsafe repeat. A run that hangs past its window overlaps the next one, or a catch-up mechanism fires four missed executions at once. Recurring jobs need an overlap policy, a decision about whether missed windows are skipped or replayed, and the same idempotency discipline as anything else that acts.
One trace ID, or effectively none
A single agent run spans a scheduler trigger, a job record, several model calls, a handful of tool invocations, and the artifacts it leaves behind. If those live in separate systems with separate identifiers, support cannot answer a complaint and engineering cannot debug the report.
Propagate one correlation ID through all of it, and store it on the artifacts too. The test is simple: given a user's message about something that went wrong yesterday, can somebody reach the full timeline in a couple of minutes without writing a query? If not, every incident starts with an archaeology phase.
Release it like anything else that can hurt someone
Prompt edits and tool changes reach production far more casually than schema migrations do, and they can be equally destructive. Give them the same treatment: an eval suite that gates the release, a canary on a slice of traffic, a flag that reverts without a deploy, and a rollback path someone has actually walked.
Then write the runbook before you need it. Who gets paged when a scheduled job fails three times. How to pause a schedule. How to safely replay a run. Two pages, written on a calm afternoon, is the difference between a contained incident and a long night.
Take last week's scheduled runs: can you say how many actually did work, how many silently did nothing, and who would have been paged if none of them had?
More on these topics
Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Article · · 1 min read
What an AI gateway actually costs to run
The operational bill for one gateway in front of a dozen tool servers, including the costs nobody budgets for.
Failure story · · 1 min read
The retrieval cache that served stale policies
A well-meaning cache in front of retrieval kept answering from last quarter's HR policy for eleven days.
Discussion