You
Everything passes locally and CI fails at random.
ChatGPT
Three usual causes: leftover state between tests that a fresh CI container does not have, timing assumptions that a slower machine breaks, and a shared resource like a fixed port or a real clock. Run the suite locally in a random order to reproduce the first one.
You
Is 80% coverage a reasonable target?
ChatGPT
Coverage measures what ran, not what was checked. A suite at 80% with meaningful assertions is healthy; the same number reached by importing modules is decoration. Track it as a trend and never as a gate.
You
How do I know when it stops being worth it?
ChatGPT
When you stop noticing it. If it has not surfaced a problem in six months, it is either solved or it was never the bottleneck you thought.