When one agent stops being enough

The first agent you build will do one thing well. The second will do two things badly. That is the moment to split responsibilities instead of adding instructions to one prompt.

The supervisor pattern

A planner receives the user goal and decides which specialist should act. Each specialist, such as a research agent or an analysis agent, owns a narrow toolset: data access, external APIs or a human feedback channel. The supervisor merges results and decides whether the goal is met.

python
supervisor = Agent(
    name="Supervisor",
    instructions="Coordinate the workflow between agents.",
    sub_agents=[research_agent, analysis_agent],
)

Contracts between agents

In production the pattern depends on the contract between agents more than on the prompt: typed inputs, typed outputs and a trace identifier that follows the request through every hop.

Trade-offs

  • Latency grows with every hand-off. Run specialists in parallel where the plan allows it.
  • Cost multiplies. Use a smaller model for routing and a larger one only where judgment matters.
  • Debugging needs a timeline view, because logs alone are not enough. Emit spans per agent step.

Where to put humans

Any step that writes to a system of record, spends money or sends a message to a customer should pause for approval until you have months of evidence that it does not need to.

Key takeaways

  • Use a clear orchestration pattern, such as a supervisor or a fixed workflow.
  • Give every agent one job and one set of tools.
  • Put a human feedback step where the cost of a wrong answer is high.
  • Trace every hop. Multi-agent systems fail quietly between agents.

Comparison · · 1 min read

LangGraph vs the OpenAI Agents SDK

Two ways to write the same supervisor. Compared on control flow, tracing, provider coupling, testing and what each makes hard.

Checklist · · 24 checks

Working with Claude, practices that hold up

Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.