Separate pools of capacity (threads, connections, quotas) so one noisy workload cannot starve the others. Named after the compartments in a ship's hull.
You have done this if
You gave batch summarisation its own model quota so it could not eat the interactive chat's rate limit.
Say it in a review
Batch and interactive traffic have separate quotas, so a big backfill can't slow down users.
On the AI Application map LLM Gateway, Agent