- Day 37: Where human review actually belongs in AI workflows
- Day 38: Why guardrails should enable safe action
- Day 39: Why production agents need recovery design
- Day 40: Why autonomy needs budgets
- Day 41: Why one giant agent is rarely the cleanest design
- Day 42: Why supervisor agents help and where they bottleneck
"There's a human in the loop" is one of the most reassuring sentences in AI product design, and one of the least informative. It says a person exists somewhere in the workflow. It says nothing about whether that person can actually change the outcome.
The goal is to place the human where they are irreplaceable. Human attention is a scarce resource, and most systems spend it badly.
Placement is a two-axis decision
Impact: can this action be undone in five minutes, or will it be discussed in a postmortem?
Uncertainty: is the agent operating on solid evidence, or on a guess?
High impact plus high uncertainty is where review belongs. Low impact plus high confidence is where automation belongs, with logs. The diagonal between them is where the actual design work lives, and it is worth mapping explicitly rather than letting it emerge from whoever implemented each feature.
Two numbers that tell you if you got it right
Approval rate. If a reviewer approves 99% of what reaches them, the gate is in the wrong place, or worse, it has trained its human to stop reading. Approval fatigue is a real failure mode: a person rubber-stamping 200 requests a day gives you the same risk with extra latency and a name attached for accountability purposes.
Time-to-context. If approving a request requires fifteen minutes of reconstructing what the agent was doing, reviewers will stop reconstructing. They will approve on vibes, and you will not know until something bad ships with a signature on it.
The fix for the second one is design, not discipline: put the proposed action, the evidence behind it, the blast radius, and a recommended decision on one screen. Reviewers who can decide in thirty seconds actually decide. Reviewers who need to go spelunking approve.
Close the loop
Track what reviewers actually do (approve, reject, modify) and feed it back. Rejection patterns are the highest-quality training signal you will ever get about where your agent's judgement is weak, and almost nobody collects them.
If reviewers consistently modify one field before approving, that is a bug report written in behaviour rather than words.
Closing thought
Automation needs accountable checkpoints. But a checkpoint nobody performs attentively is worse than none: same risk, plus false assurance, plus delay.
Fewer, better-placed gates with better context beat comprehensive checkbox theatre every time.
More on these topics
Comparison · · 1 min read
LangGraph vs the OpenAI Agents SDK
Two ways to write the same supervisor. Compared on control flow, tracing, provider coupling, testing and what each makes hard.
Checklist · · 24 checks
Working with Claude, practices that hold up
Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.
Architecture pattern · · 1 min read
Pattern: the outbox for agent actions
Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.
Discussion