How do I estimate a data migration with unknown data quality?
Split it into three numbers: the transform you can specify, the exceptions you find in a sample, and the tail you cannot see yet. Sample a thousand rows first; the exception rate there sets the multiplier for everything else, and the multiplier is usually between two and four.
Our estimates are consistently half of reality.
Because they estimate the work, not the day. A useful correction is to track a personal ratio of estimate to actual over a few months and apply it openly rather than padding silently, which reads as sandbagging when it is discovered.
How do you plan when the date is fixed?
Fix the date and vary the scope, explicitly and in writing, with the cut list agreed before work starts. The failure mode is holding both fixed, which does not change the outcome, only when everyone finds out.