Somewhere in your codebase there is a dictionary mapping tool names to functions, and every entry in it is offered to the model on every request. That dictionary is your permission model. Nobody wrote it down as one, nobody reviews it as one, and it very likely contains something like issue_refund.

The registry got built as a catalogue: what exists, what it does, what arguments it takes. Useful. It stops being sufficient the moment a second user with different rights sends a request.

The three bindings every entry needs

Identity. Which credential does this tool execute under? A single broad service token makes every call succeed and makes every side effect anonymous. When finance asks who approved a refund, "the agent" is not an answer. Tools that act on a user's behalf should carry that user's identity down to the system being touched, brokered at call time rather than baked into the process environment.

Scope. What can that credential read or change, expressed narrowly enough to be interesting? "Database access" is not a scope. "Read orders belonging to the requesting user, within this workspace" is.

Approval level. Some tools are free. Some are staged and need a human confirmation. Some are unavailable in this environment regardless of who is asking. This belongs in the registry entry, next to the schema, rather than scattered across prompt text.

Schemas sit alongside those three, and they are where drift creeps in. The description shown to the model is a promise about behaviour. When the implementation gains a parameter or quietly widens what it deletes and the description does not change, the model is now planning against a system that no longer exists. Test the schema and the permission scope in the same test, because a tool with a correct schema and the wrong credential is still a security bug.

Audit is what makes the other three checkable

Filtering the catalogue per user is the visible half of the work. The invisible half is whether you can reconstruct, months later, that a specific tool ran with specific arguments under a specific identity on behalf of a specific person, and that the scope in force at the time actually permitted it.

Without that record, permission scoping is an intention. With it, scoping becomes a claim you can verify, which means you can safely widen it. Registries that log properly end up being the ones teams are willing to add powerful tools to.

Pick the most destructive tool in your catalogue: whose credential does it run under, and could you prove that to an auditor without opening the source?

Checklist · · 29 checks

Security review checklist for an AI feature

What to check before an assistant, RAG app or agent goes in front of real users. Grouped by area, ticked off locally; progress stays in your browser.