Rewriting to async made throughput worse.
Something synchronous is running inside the event loop, usually a blocking library call or CPU work. One blocking call stalls every other task on that loop. Run those in a thread executor, and measure with a loop-lag metric so the next one is visible immediately.
Four threads and the CPU-bound job takes the same time.
The interpreter lock serialises bytecode execution, so threads help with waiting, not with computing. Processes are the fix for CPU-bound work:
from concurrent.futures import ProcessPoolExecutor
with ProcessPoolExecutor() as pool:
results = list(pool.map(work, items, chunksize=64))
chunksize matters more than the worker count: without it, small tasks spend all their time being serialised between processes.
Two transactions deadlock occasionally in production.
They take the same locks in different orders, and only concurrency makes that visible. Establish a global ordering, usually by sorting the ids you touch before touching them. Retry logic is a bandage: correct, necessary, and not a substitute for the ordering.
Does this still hold if I am working alone?
It gets easier, not harder. A fair share of the difficulty is coordination, and that disappears. What remains is remembering your own decisions, which is what writing them down is for.