Continuing to work correctly despite a component failing, usually through redundancy. Stricter than resilience, which allows degraded service.
You have done this if
You ran two replicas behind a load balancer so a node could die without users noticing.
Say it in a review
Losing one instance is a non-event; we run at least two in separate zones.
On the AI Application map Web App, API Gateway