- Day 31: Why agents need outcomes, boundaries, and stop conditions
- Day 32: Why agent behavior is a loop, not a single prompt
- Day 33: Why planning reduces wasted AI actions
- Day 34: Why tool access is where AI becomes operational risk
- Day 35: Why agent memory is not one database
- Day 36: Why reflection should change the next action
A workflow follows your steps. An agent pursues your goal. That difference sounds philosophical right up until you have to write the goal down.
Because the goal you hand an agent is a specification, and it gets executed with a compiler's literalism. Underspecify it and the agent does not fill the gaps with your intent. It fills them with whatever best satisfies the words you actually used.
Three things that have to be written down
Success criteria a machine can check. "Improve the report" is not a criterion; it is a hope. "Every section contains current-quarter figures, each with a cited source" is checkable, by the agent and by you. If success can only be assessed by a human squinting at the output, you have delegated a vibe.
Boundaries, enforced outside the prompt. What it may spend, which systems it may touch, what it must never modify. These belong in credentials and runtime limits, not in a politely worded instruction. More on that distinction on day 38.
Stop conditions. Done, blocked, out of budget, and the fourth one everyone forgets: looping. An agent repeating the same failing approach will do so cheerfully until something external intervenes.
if risk == "high":
escalate_to_human()
elif retries > 2:
stop_with_reason("retry budget exceeded")
elif evidence_quality < threshold:
ask_for_more_context()
elif tool_budget.remaining == 0:
stop_with_reason("tool budget exhausted")
else:
continue_workflow()Bounded action, with explicit reasons to stop.
The review exercise worth stealing
Hand your goal specification to a colleague and ask them a single question: what would a malicious-but-compliant reading of this do?
Not "what should it do": what could it do while technically satisfying every word you wrote. That reading is available to the agent too, and it is the one you will meet in production when the inputs stop resembling your test cases.
The answers are usually uncomfortable and always cheap to fix at this stage. "Reduce open ticket count" is satisfied by closing tickets without resolving them. "Maximise response quality" is satisfied by taking forty minutes and thirty tool calls per request.
Vague goals produce confident agents doing the wrong thing efficiently. The first run usually looks impressive, which is precisely what delays the discovery.
Could you write down what success means for your agent precisely enough that a stranger could verify it?
More on these topics
Comparison · · 1 min read
LangGraph vs the OpenAI Agents SDK
Two ways to write the same supervisor. Compared on control flow, tracing, provider coupling, testing and what each makes hard.
Checklist · · 24 checks
Working with Claude, practices that hold up
Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.
Architecture pattern · · 1 min read
Pattern: the outbox for agent actions
Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.
Discussion