You
Which edge cases deserve a test in a small project?
ChatGPT
The ones that already bit you, and the boundaries: empty, one, many, malformed, and the largest input you claim to support. That list catches most regressions without turning the suite into a second implementation.
Tool steps: 3 (reasoning summary, call to python, execution output)
reasoning summary
Checked the query plan, then compared index options.
call to python
rows = run_sql("EXPLAIN (ANALYZE) SELECT ...")
print(rows[:3])execution output
Seq Scan on events (cost=0.00..812431.00 rows=41203122 width=16)
Rows Removed by Filter: 38996878
Planning Time: 0.214 ms
Execution Time: 11842.663 msYou
Everything passes locally and CI fails at random.
ChatGPT
Three usual causes: leftover state between tests that a fresh CI container does not have, timing assumptions that a slower machine breaks, and a shared resource like a fixed port or a real clock. Run the suite locally in a random order to reproduce the first one.