Four phases, in this order.
1. Stop accepting. On SIGTERM, flip the readiness probe to failing and close the listener, so the load balancer stops routing new requests here while the process is still alive. Readiness must flip before the listener closes in a real cluster — the LB needs a few seconds to notice, which is why production shutdown handlers sleep before closing anything.
2. Drain. Keep serving what is already in flight. The exercise's drain curve is 5 → 4 → 2 → 1 as requests retire.
3. Deadline. Draining cannot be unbounded: one stuck request would hold the pod forever, and the orchestrator's own grace period will not wait. Kubernetes gives you 30 s and then sends SIGKILL.
4. Force-close what is left, and log which requests you killed — request 8 here, the 20-tick outlier.
