Padding gets consumed immediately. Better approach?
Keep the buffer at the project level rather than per task, and make spending it visible. Per-task padding disappears into Parkinson's law; a shared buffer with a burn chart stays honest because everyone sees the same number.
How do I estimate a data migration with unknown data quality?
Split it into three numbers: the transform you can specify, the exceptions you find in a sample, and the tail you cannot see yet. Sample a thousand rows first; the exception rate there sets the multiplier for everything else, and the multiplier is usually between two and four.
Our estimates are consistently half of reality.
Because they estimate the work, not the day. A useful correction is to track a personal ratio of estimate to actual over a few months and apply it openly rather than padding silently, which reads as sandbagging when it is discovered.
Anything that changes this at ten times the scale?
At ten times the size the bottleneck moves from the operation itself to what it competes with: memory, locks and the people who have to run it. The approach stays, the batching gets smaller and the schedule matters more.