AI Systems Basics
What a language model does, and why the system around it decides whether it works.
Read in order
All 18 series complete, 100 parts.
What a language model does, and why the system around it decides whether it works.
The gap between a capable model and a product people can rely on: evidence, structure, and choosing the right shape.
Chunking, embeddings, vector stores, keyword and hybrid search, and the permissions retrieval must respect.
Reranking, provenance, multi-query and graph retrieval: turning search results into evidence.
Knowledge in systems, multimodal evidence, freshness, recovery and evaluating retrieval in parts.
Outcomes, loops, planning, tools, memory and reflection: what makes an agent more than a prompt.
Human review, guardrails, recovery, budgets and when to split one agent into several.
Hierarchies, dispatch, pipelines, peers, shared state and debate, with the failure modes of each.
Roles, shared language, memory permissions, handoffs and decision rules between agents.
Observability, evaluation, untrusted content, governance and infrastructure: what production asks of the whole system.
What the Model Context Protocol standardises, where servers run, and when MCP earns its overhead.
The runtime underneath agents: autonomy, isolation, scaling and failure.
Retrieval as a system you operate: trusted sources and freshness, ingestion as a data product, retrieval as its own subsystem, chunking for the question, and citations you can evaluate.
Evaluation that can block a release: what an eval may stop, golden sets maintained like products, calibrated judges, regression gates and evaluation after launch.
Safety as architecture rather than moderation: threat models, agent identity and scope, untrusted input, blast radius and data exfiltration.
When several agents actually help, and what it takes to run them: role contracts and handoffs, who controls the run, budgets and stop conditions, traces and recovery.
The building blocks of a production AI system: the model gateway, choosing the right pattern, control and data planes, approvals and tenant boundaries, and a reference architecture.
Running AI in production: traces and spans, cost, latency and capacity, production evaluation signals, incidents and rollbacks, drift and an operating cadence.