Padding gets consumed immediately. Better approach?
Keep the buffer at the project level rather than per task, and make spending it visible. Per-task padding disappears into Parkinson's law; a shared buffer with a burn chart stays honest because everyone sees the same number.
How do I estimate a data migration with unknown data quality?
Split it into three numbers: the transform you can specify, the exceptions you find in a sample, and the tail you cannot see yet. Sample a thousand rows first; the exception rate there sets the multiplier for everything else, and the multiplier is usually between two and four.
Our estimates are consistently half of reality.
Because they estimate the work, not the day. A useful correction is to track a personal ratio of estimate to actual over a few months and apply it openly rather than padding silently, which reads as sandbagging when it is discovered.
Any downside worth knowing before I commit?
It commits you to a format that is tedious to migrate away from later. The first weeks also look worse than doing nothing, which is when most people abandon it.