- Day 7: Why confident AI answers still need evidence
- Day 8: Why production AI needs structured outputs
- Day 9: Why a capable model is still not a product
- Day 10: When to use a chatbot, workflow, or agent
- Day 11: Why RAG is an evidence design problem, not a buzzword
- Day 12: Why retrieval quality starts before search
The dangerous hallucinations are not the strange ones. Those get caught.
The dangerous ones are fluent, confident, and almost right, a real policy with one wrong number, a real process with one invented step, a correct answer to a question nobody asked.
Confidence is not evidence.
A fluent answer still has to earn trust through sources, boundaries, uncertainty handling, and a path for review.
Why it does not look like a bug
Every other failure in software announces itself. Exceptions, timeouts, 500s, red tests.
A confident wrong answer renders perfectly. It has the same tone, structure, and apparent authority as a correct one, because fluency and accuracy are produced by the same mechanism and only one of them is being optimised for at generation time.
So instead of being reported as a failure, it gets absorbed, either by a user who acts on it or by one who quietly notices, stops trusting the feature, and goes back to doing it manually. Trust leaks out well before anyone opens a ticket.
Stop treating it as a model defect
There is a comfortable framing where hallucination is a temporary model limitation that will be solved by the next release. It will not be, entirely, and waiting is not a strategy.
The productive framing: hallucination is a system-design problem about evidence and trust. Three decisions follow from it, and none require a better model:
- Which claims must carry evidence before a user sees them. Not all of them, which would be unworkable, but the material ones: numbers, eligibility, policy, anything a person might act on.
- What the system says when it does not know. This has to be an explicit, designed path. The default behaviour of a language model handed an unanswerable question is to answer it anyway.
- Who owns accuracy in production, and what they monitor.
Pressure-test with the unanswerable
Take a question your sources genuinely cannot answer and watch what happens.
Refuses? Your evidence rules are real. Hedges honestly? Good. Produces a confident, plausible, entirely invented answer? You have just seen exactly what your users will encounter, and unlike them, you know it is wrong.
That fifteen-minute test tells you more about your system's trustworthiness than any benchmark score.
Confident AI answers still need evidence. The model's job is to generate; the system's job is to decide what generation is allowed to reach a person unverified.
Which claims in your output should require a source before a user ever sees them?
Key takeaways
- Confidence is not evidence.
More on these topics
Deep dive · · 6 min read
Everything that makes an AI demo impressive is a constraint you have not added yet
Six posts on the distance between a system that answers well once and a system people are willing to depend on.
Discussion