How to use this
Run it against one real feature, not in the abstract. Most teams find two or three gaps, usually in tool descriptions, evaluation and stop conditions. Those are where reliability actually comes from, more than the choice of model.
The prompting items follow Anthropic's published guidance on prompt engineering and building agents. The rest are the engineering habits around the model: the harness, the evaluation and the controls.
Related
Terms such as agent harness, structured output, context window budget and eval harness are in the glossary. For agents in production, see the Agents in production path.
The checks
Prompting
0 / 7Instructions are clear and direct, and say why a rule exists, not only what it is.
Claude follows the reason as well as the rule, so it handles cases the rule did not list.
The prompt gives the context a new colleague would need, such as the audience, the goal and what good looks like.
Two or three examples show the format and tone you want, including one edge case.
Long documents go near the top of the prompt and the question near the end.
Different parts of the prompt are wrapped in XML tags (instructions, documents, examples), so they cannot blur together.
For hard problems, Claude is given room to reason before answering, through extended thinking or a scratchpad section.
Output that code will consume is requested as JSON against a schema and validated.
Agents and tools
0 / 6Each tool has a name and description written like documentation for a new teammate, including when not to use it.
Tool descriptions are prompts; vague ones produce vague tool choices.
The agent sees a small, relevant set of tools for the task, not every tool you have.
Tool errors return a useful message the model can act on, not a stack trace or an empty string.
The agent has an explicit goal, limits on steps and spend, and a clear condition for when it is done.
Long tasks keep notes or state outside the context window, so work survives a reset or a handover.
Multi-step work is split into a chain of focused calls where each step can be checked, instead of one giant prompt.
Claude Code
0 / 5The repository has a CLAUDE.md with build, test and lint commands, conventions and the things not to touch.
For non-trivial changes, ask for a plan first, review it, then let it implement.
Tests or type checks act as the definition of done, so Claude can verify its own work.
Changes are kept small and reviewed as diffs before commit, the same as a colleague's pull request.
Permissions are set deliberately, with routine commands allowed and destructive ones asking first.
Evaluation
0 / 3There is a set of real tasks with known good outcomes, and every prompt or model change runs against it.
Grading uses a written rubric; if a model grades, a sample is checked by a person.
Model versions are pinned, and upgrades are treated as releases with an evaluation run.
Safety and trust
0 / 3Content Claude reads from documents, the web or tools is treated as untrusted data, never as instructions.
Actions that send, pay, delete or publish wait for a human.
The system tells users when it does not know, and that path is tested.
More on these topics
Comparison · · 1 min read
LangGraph vs the OpenAI Agents SDK
Two ways to write the same supervisor. Compared on control flow, tracing, provider coupling, testing and what each makes hard.
Checklist · · 12 checks
RAG production-readiness checklist
Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.
Architecture pattern · · 1 min read
Pattern: the outbox for agent actions
Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.
Discussion