Two transactions deadlock occasionally in production.
They take the same locks in different orders, and only concurrency makes that visible. Establish a global ordering, usually by sorting the ids you touch before touching them. Retry logic is a bandage: correct, necessary, and not a substitute for the ordering.
Rewriting to async made throughput worse.
Something synchronous is running inside the event loop, usually a blocking library call or CPU work. One blocking call stalls every other task on that loop. Run those in a thread executor, and measure with a loop-lag metric so the next one is visible immediately.
Four threads and the CPU-bound job takes the same time.
The interpreter lock serialises bytecode execution, so threads help with waiting, not with computing. Processes are the fix for CPU-bound work:
from concurrent.futures import ProcessPoolExecutor
with ProcessPoolExecutor() as pool:
results = list(pool.map(work, items, chunksize=64))
chunksize matters more than the worker count: without it, small tasks spend all their time being serialised between processes.