You
Deploys cause a handful of 502s every time.
ChatGPT
The old pod stops accepting before the proxy stops sending. Add a readiness probe and a preStop sleep of a few seconds:
lifecycle:
preStop:
exec:
command: ["sleep", "5"]
That gap gives the endpoint controller time to remove the pod before the process exits.
You
How do I know when it stops being worth it?
ChatGPT
When you stop noticing it. If it has not surfaced a problem in six months, it is either solved or it was never the bottleneck you thought.