Sending a small share of traffic to a new version first and comparing it with the old one before rolling out further.
You have done this if
You routed 5 percent of users to the new prompt and watched the eval scores before switching everyone.
Say it in a review
Prompt and model changes go out as a canary at 5 percent with automatic rollback on eval regressions.
On the AI Application map LLM Gateway