Everything passes locally and CI fails at random.
Three usual causes: leftover state between tests that a fresh CI container does not have, timing assumptions that a slower machine breaks, and a shared resource like a fixed port or a real clock. Run the suite locally in a random order to reproduce the first one.
Is 80% coverage a reasonable target?
Coverage measures what ran, not what was checked. A suite at 80% with meaningful assertions is healthy; the same number reached by importing modules is decoration. Track it as a trend and never as a gate.
Does this still hold if I am working alone?
It gets easier, not harder. A fair share of the difficulty is coordination, and that disappears. What remains is remembering your own decisions, which is what writing them down is for.