We added a cache and p99 got worse

Back to list2024-07-29
You

We added a cache and p99 got worse.

ChatGPT

Classic signature of a cache that misses under load: every miss now pays the lookup plus the original work, and the population step often serialises on a lock. Measure hit rate before and after, and add jitter to expiry so entries do not all die in the same second.

You

Where do I start when a script is slow and I have no idea why?

ChatGPT

Profile before guessing. For Python:

python -m cProfile -s cumtime script.py 2>&1 | head -25

Read the cumulative column first, not the per-call one. Nine times out of ten the top entry is a function nobody suspected, and the one you were about to optimise is at 2%.

You

Is there a simpler version that gets most of the benefit?

ChatGPT

Yes: do the first step, skip the automation, and revisit in a month. Most of the value is in the first step, and most of the cost is in making it repeatable before you know it is right.