Copying millions of small files

Back to list2024-08-16
You

rsync of a directory with 2M small files takes forever.

ChatGPT

The bottleneck is per-file syscalls, not bandwidth. tar piped over ssh moves them as a single stream and is often several times faster; rsync becomes the right tool again on the second run, when only the differences matter.

You

df says 96% full, du disagrees. What is going on?

ChatGPT

Almost always a deleted file still held open by a process, so the space is gone but no path points at it. lsof +L1 lists them; restarting the holder releases the space instantly.

You

Anything that changes this at ten times the scale?

ChatGPT

At ten times the size the bottleneck moves from the operation itself to what it competes with: memory, locks and the people who have to run it. The approach stays, the batching gets smaller and the schedule matters more.