New here? Start this series at Part 1: Logging the answer tells you almost nothing
- Part 1: Logging the answer tells you almost nothing
- Part 2: Your AI feature has unit economics whether you measured them or not
- Part 3: Real users ask questions your test set never imagined
- Part 4: You cannot roll back a prompt you never versioned
- Part 5: Nothing broke, and the system is still getting worse
Nothing broke. Dashboards are green, error rates flat, latency unchanged.
The system is also worse than it was in March, and it has been getting worse the entire time, in increments too small for any single week to notice.
That is drift, and it outlasts everything else in this track. You can have traces, budgets, eval signals and runbooks, and still lose slowly, because losing slowly triggers nothing.
What moves underneath a working system
Five things shift beneath an AI system throwing no errors:
- Data drift. The corpus changes. New products, revised policies, deprecated pages nobody unpublished.
- Usage drift. Users learn what the system is good at and change what they ask. Your month-six traffic mix is not the one you tuned for.
- Prompt drift. Fifteen small edits, none reviewed as a set, no single one worth an eval run.
- Model drift. The vendor updates. Your behaviour moves without a deploy on your side.
- Retrieval drift. The index grows, neighbourhoods get crowded, recall degrades for queries that used to be easy.
None of these page anyone. All of them are visible in metrics you already collect, if somebody looks at them over months rather than minutes.
Four SLOs and one shared view
Pick a number for each and let it be uncomfortable. Quality, measured on the eval set plus your production sample. Safety, as policy violations per thousand tasks. Latency, per workflow class. Cost, per successful task.
Four numbers, four named owners, one dashboard. The failure I keep seeing is three dashboards rather than none: engineering watching latency, finance watching spend, support watching tickets, each convinced the system is fine or broken according to their own view. When quality gets debated, everyone brings a different chart and the meeting resolves nothing.
A shared view will not remove disagreement, but it makes the disagreement about the same reality, which is most of the work.
The cadence is the product
Monthly, somebody owns an hour with the same agenda every time: the four SLOs against target, drift indicators, top escalation reasons by volume, what changed since last time, and the ranked improvement list.
Ranked is the load-bearing word. Backlogs otherwise fill with whatever the most recent demo made someone anxious about. The alternative is ranking by evidence: this failure class appeared in nineteen escalations and eleven flagged samples, that one appeared twice. Ties are broken by risk rather than by who asked loudest.
Everything else in this track exists to make that hour possible. Traces give it evidence, budgets a cost line, eval signals a quality number that is not a vibe, and runbooks a record of what actually went wrong rather than what people remember.
None of it makes the system correct, but it does make it knowable. That is the only durable position available, because the model, the data and especially the users will change.
What in your AI system got worse this quarter without anyone filing a ticket?
More on these topics
Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Article · · 1 min read
What an AI gateway actually costs to run
The operational bill for one gateway in front of a dozen tool servers, including the costs nobody budgets for.
Failure story · · 1 min read
The retrieval cache that served stale policies
A well-meaning cache in front of retrieval kept answering from last quarter's HR policy for eleven days.
Discussion