Skip to content

Everything written

Journal

Every entry, read by engineering discipline or in series. Each one also sits on the map, next to the component it is about.

121 entries: all 18 series complete, 100 parts, plus standalone pieces. Weekly notes live in Notes .

Start here

October 2026 5

Build log: shipping Unseen UI to npm

How the component library behind this site went from a private workspace to two published packages, including the three things the first consumer broke.

Build log1 min readPractitioner

LangGraph vs the OpenAI Agents SDK

Two ways to write the same supervisor. Compared on control flow, tracing, provider coupling, testing and what each makes hard.

Comparison1 min readPractitioner

RAG production-readiness checklist

Twelve checks to pass before a retrieval system answers a real user. Tick them locally; progress stays in your browser.

Checklist12 checksPractitioner

September 2026 24

Working with Claude, practices that hold up

Prompting, agents and tools, Claude Code, evaluation and safety. The habits that make Claude-based systems reliable, as a checklist you can run against your own setup.

Checklist24 checksPractitioner

Pattern: the outbox for agent actions

Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.

Architecture pattern1 min readArchitect

When to bring in a compliance review

The changes that should pull legal, privacy or compliance into an AI project early, and what to have ready when you do. Not legal advice; a way to ask at the right time.

Checklist18 checksPractitioner

Security review checklist for an AI feature

What to check before an assistant, RAG app or agent goes in front of real users. Grouped by area, ticked off locally; progress stays in your browser.

Checklist29 checksPractitioner

You cannot roll back a prompt you never versioned

AI failures cross model, prompt, retrieval, tool and data boundaries. The runbook and the rollback controls have to as well.

Explainer2 min readPractitioner5 Days of AI Operations and Observability

Logging the answer tells you almost nothing

A trace that links intent, prompt, retrieval, tools and output is the only thing that makes an AI failure debuggable.

Explainer2 min readPractitioner5 Days of AI Operations and Observability

Most agent controls do not actually control anything

Six days of notes on supervising autonomous systems, and the same failure shape kept turning up: the control exists, it is documented, and nothing in the running system is bound by it.

Deep dive6 min readPractitioner

Show 101 older entries, back to May 2026

An approve button is not human oversight

Approvals, events and tenant boundaries are architecture. Bolt them on as UI and they become theatre.

Explainer2 min readPractitioner5 Days of AI Architecture Patterns

You are already building a control plane, badly

Policy, memory, evals and rollout get rebuilt inside every feature until someone names the layer they belong to.

Explainer2 min readPractitioner5 Days of AI Architecture Patterns

Most agents are workflows wearing a costume

Chatbot, workflow, agent, RAG. Pick the wrong one and you spend a quarter debugging autonomy nobody asked for.

Comparison2 min readPractitioner5 Days of AI Architecture Patterns

Write role contracts, not agent personalities

"You are a meticulous senior researcher" is a costume. Inputs, outputs, permissions and done criteria are a contract.

Explainer2 min readPractitioner5 Days of Multi-Agent System Design

August 2026 25

Blast radius is a design parameter

Narrow tools, staged writes, sandboxes and a tested kill switch. Containment is built before it is needed.

Explainer2 min readPractitioner5 Days of AI Safety Engineering

Untrusted text does not get to give orders

Injection defence is layered validation and clear authority, not a better-worded system prompt.

Explainer2 min readPractitioner5 Days of AI Safety Engineering

Your agent should not be a superuser

Bind every AI action to a user, a service and a run, then scope the verb rather than only the data.

Explainer2 min readPractitioner5 Days of AI Safety Engineering

Launch day is when evaluation starts

Offline scores expire on contact with real users. Sampling, groundedness, drift and outcomes are the parts that keep paying.

Explainer2 min readPractitioner5 Days of AI Evaluation

You cannot run RAG on user complaints

Groundedness scores, per-claim citations and an operations dashboard are what tell you quality is drifting before your users do.

Explainer2 min readPractitioner5 Days of Production RAG

Chunk for the question, not for the token limit

Chunk size is a decision about what a complete answer looks like. Reranking, context budgets and refusal thresholds finish the job.

Explainer2 min readPractitioner5 Days of Production RAG

The 3am agent run that nobody is watching

Scheduled agent work fails quietly by default. Traces, release gates and a runbook are what make it fail loudly.

Explainer2 min readPractitioner5 Days of Agent Infrastructure

July 2026 28

What the MCP protocol actually standardises

It settles how capabilities are described and discovered. Every decision about whether to trust them is still yours.

Explainer1 min readPractitioner5 Days of MCP for Production AI

Why handoffs need contracts

Why reliable handoffs need structured state, evidence, open questions, and ownership.

Explainer1 min readPractitionerAgent Coordination Contracts

Why agents need a shared language

Why agents need shared message contracts before they can collaborate reliably.

Explainer2 min readPractitionerAgent Coordination Contracts

Why autonomy needs budgets

Why autonomous systems need limits on cost, time, actions, retries, and risk.

Explainer1 min readPractitionerAgent Control and Supervision

June 2026 29

Why RAG must be evaluated in parts

Why RAG evaluation has to measure retrieval, grounding, generation, and citations separately.

Explainer1 min readPractitionerProduction RAG Operations

Why freshness is part of correctness

Why freshness becomes part of correctness when the world changes faster than your index.

Explainer2 min readPractitionerProduction RAG Operations

Why evidence is becoming multimodal

Why evidence quality changes when the source is visual, tabular, audio, or multimodal.

Explainer2 min readPractitionerProduction RAG Operations

May 2026 10