Five-day tracks
5 Days of AI Safety Engineering
Safety as architecture rather than moderation: threat models, agent identity and scope, untrusted input, blast radius and data exfiltration.
5 of 5 published
- 01
Part 1 · Explainer · 2 min read
Safety is an architecture decision, not a moderation setting
Assets, actors, abuse cases, controls, owners, evidence. Six things a team can actually write down before launch.
On the map: LLM Gateway - 02
Part 2 · Explainer · 2 min read
Your agent should not be a superuser
Bind every AI action to a user, a service and a run, then scope the verb rather than only the data.
- 03
Part 3 · Explainer · 2 min read
Untrusted text does not get to give orders
Injection defence is layered validation and clear authority, not a better-worded system prompt.
On the map: Retrieval, LLM Gateway - 04
Part 4 · Explainer · 2 min read
Blast radius is a design parameter
Narrow tools, staged writes, sandboxes and a tested kill switch. Containment is built before it is needed.
On the map: Tools - 05
Part 5 · Explainer · 2 min read
You cannot instruct a model into data protection
Minimise what enters context, watch the paths data can leave by, and keep the evidence an incident will demand.