By the third AI feature, the pattern is visible. Each one has its own prompt file, its own redaction step, its own trace format, its own idea of how much conversation history to keep. Changing one PII rule takes three pull requests, and one team ships late because nobody told them.

Nothing here is broken. Four platform concerns simply got copied into product code before anyone gave them a home.

Two planes, two clock speeds

The data plane is the request path: retrieve, assemble context, call the model, run tools, return a response. It runs per request, it is latency-critical, and it should be as boring as possible.

The control plane decides how the data plane behaves. Which prompt version is live, which model tier this tenant gets, what policy applies, what gets sampled into the eval set, what share of traffic sees the new configuration. It changes between requests and holds still during them.

Once you draw that line, the layers sort themselves out:

LayerControl plane ownsData plane does
Policy and permissionsrule definitions, per-tenant configurationchecks and redacts on the call
Memory and contextretention, scope, what may persistassembles the context window
Evaluationgolden sets, thresholds, sampling rulesemits a trace in the shared schema
Rolloutprompt and model versions, traffic splitsreads the resolved configuration

The right-hand column is deliberately thin. Data plane code should read decisions that were made somewhere else.

The test is whether behaviour can change without a deploy

It is uncomfortably easy to check.

Tighten a redaction rule for one tenant. Roll a prompt back to yesterday's version. Move ten per cent of traffic to a cheaper model for one workflow. If any of those needs a code change, a release train and a regression check across three services, the control plane exists only as a set of habits.

The observability half has the same test. Ask what percentage of AI requests across your whole product emit traces you can compare. If each feature invented its own span names, you can debug one feature at a time but you cannot answer "did quality drop this week", because there is no "this week" that spans features.

Evaluation belongs here too, instead of being bolted on at the end. If traces flow into a shared store with a shared schema, building a golden set is filtering. If they do not, it is a data engineering project every single time.

Closing thought

Nobody sets out to build a control plane. It accumulates as duplicated policy code, five trace formats and a prompt change that needs a release. Naming the layer early is much cheaper than extracting it from four features later.

You have a control plane either way. Is yours a system or a habit?

Architecture pattern · · 1 min read

Pattern: the outbox for agent actions

Agents that write to systems of record need the same transactional outbox that event-driven services use. This entry covers the shape and the trade-offs.