You
Should requests equal limits?
ChatGPT
For memory, yes: memory is incompressible, so a limit above the request means the pod gets killed rather than throttled. For CPU, set requests and leave limits off unless you need hard isolation, because CPU limits throttle in ways that look like random latency.
You
Deploys cause a handful of 502s every time.
ChatGPT
The old pod stops accepting before the proxy stops sending. Add a readiness probe and a preStop sleep of a few seconds:
lifecycle:
preStop:
exec:
command: ["sleep", "5"]
That gap gives the endpoint controller time to remove the pod before the process exits.
You
How do I know when it stops being worth it?
ChatGPT
When you stop noticing it. If it has not surfaced a problem in six months, it is either solved or it was never the bottleneck you thought.