meta

Migrating Data Ingestion Systems at Meta Scale (opens in new tab)

Meta rebuilt its hyperscale MySQL data ingestion system to improve reliability, efficiency, and data-langing latency. The migration moved workloads from customer-owned pipelines to a simpler, self-managed warehouse service and ultimately transitioned 100% of jobs. Success depended on staged validation, continuous data comparison, and fast rollback mechanisms.

Why Meta Migrated

  • The system incrementally moved several petabytes of social graph data from MySQL into Meta’s data warehouse each day.
  • This data supports analytics, reporting, machine learning, and product development.
  • The legacy architecture became increasingly unstable as data-landing requirements grew stricter.
  • Customer-owned pipelines worked at smaller scales but became difficult to manage reliably at hyperscale.

Migration Success Criteria

Each job had to meet defined requirements before advancing:

  • Data correctness: Old and new systems had matching row counts and checksums.
  • Landing latency: The new system performed at least as well as the legacy system.
  • Resource usage: Compute and storage consumption did not regress.
  • Critical-table requirements: Additional criteria were agreed upon with dependent teams.

Three-Phase Migration Lifecycle

Shadow Phase

  • New-system shadow jobs ran against the same production sources as existing jobs.
  • Their output was written to separate shadow tables.
  • Row counts and checksums were continuously compared with production data.
  • Compute and storage requirements were measured before production rollout.
  • Once validated in pre-production, shadow jobs were tested in production.

Reverse Shadow Phase

  • The new system began writing to the production table.
  • The legacy system continued running, but wrote to a shadow table.
  • This preserved continuous comparison between both systems.
  • If discrepancies appeared, Meta could quickly restore the old system without rebuilding its configuration.

Migration Cleanup

  • Both systems continued to be monitored for mismatches.
  • After validation, the legacy shadow job was removed.
  • The new system became the sole production pipeline.

Data Quality and Debugging Tooling

  • Meta built tooling to compare corresponding table partitions from the two systems.
  • Comparisons included row counts, checksums, and example rows responsible for mismatches.
  • Mismatch records and debugging details were logged to Scuba for real-time analysis.
  • Hourly queries helped engineers identify root causes and determine whether issues were already known.
  • The same tooling remains part of post-migration release validation.

Rollout and Rollback Controls

  • Both systems used change data capture (CDC), with internal full-dump and delta tables feeding customer-facing target tables.
  • Because CDC builds new data from previously landed data, an existing defect could propagate after migration.
  • Meta therefore emphasized:
    • Detecting problems before they reached data consumers.
    • Stopping further propagation quickly during rollback.
  • The reverse-shadow design provided early quality signals and preserved a ready-to-use legacy pipeline for rapid recovery.

Meta’s migration demonstrates that large-scale infrastructure changes are safest when treated as controlled, observable lifecycle transitions rather than one-time cutovers. Parallel execution, automated data validation, explicit resource checks, and reversible rollouts enabled the company to migrate the entire workload while protecting downstream consumers.