A backfill is a run of a pipeline over past time ranges: to load history into a new table, to apply a corrected transformation to data already processed, or to fill the dates an outage skipped. Done well, it is the daily job run for many past logical dates, with each run overwriting its own partition, so running a date twice gives the same result (Maxime Beauchemin, Functional Data Engineering).
In an FDE interview
Backfills sit behind data problems such as “rebuild a year of balances” or “the feed was wrong for a month”, and behind migrations. In a migration the order is what makes it safe: the application first writes the new column on every insert and update, then a batched backfill fills the old rows, and only then do reads switch to it.
A strong candidate says the backfill uses the same code path as the daily run, partitions by the logical date rather than the wall clock, reads every input as of that date (the dimension version valid then, not today’s row; see slowly changing dimension), checks that the source still holds data that far back, replaces partitions instead of appending, and limits concurrency so the warehouse still serves production. In Airflow, airflow backfill create --max-active-runs caps concurrency and --reprocess-behavior decides whether dates that already have a run are run again (Airflow, Backfill). They also reconcile totals before and after, and tell downstream users which numbers will change.