Deadlock that only happens under load

Back to list2023-11-11
You

Two transactions deadlock occasionally in production.

ChatGPT

They take the same locks in different orders, and only concurrency makes that visible. Establish a global ordering, usually by sorting the ids you touch before touching them. Retry logic is a bandage: correct, necessary, and not a substitute for the ordering.

You

Rewriting to async made throughput worse.

ChatGPT

Something synchronous is running inside the event loop, usually a blocking library call or CPU work. One blocking call stalls every other task on that loop. Run those in a thread executor, and measure with a loop-lag metric so the next one is visible immediately.

You

Four threads and the CPU-bound job takes the same time.

ChatGPT

The interpreter lock serialises bytecode execution, so threads help with waiting, not with computing. Processes are the fix for CPU-bound work:

from concurrent.futures import ProcessPoolExecutor

with ProcessPoolExecutor() as pool:
    results = list(pool.map(work, items, chunksize=64))

chunksize matters more than the worker count: without it, small tasks spend all their time being serialised between processes.