Network Topology

1 posts

meta3 min readCurated summary

Building Prometheus: How Backend Aggregation Enables Gigawatt-Scale AI Clusters

Meta uses Backend Aggregation (BAG) as a high-capacity Ethernet super-spine to connect tens of thousands of GPUs across data centers and regions. In the Prometheus AI cluster, BAG links regional networks and Meta’s backbone while bridging two L2 fabric technologies: Disaggregated Schedule Fabric (DSF) and Non-Scheduled Fabric (NSF). Its modular hardware, resilient topologies, and advanced routing are designed to deliver reliable, petabit-scale connectivity for a gigawatt-scale AI system. ## Backend Aggregation’s Role - BAG interconnects multiple spine fabrics across data centers and regions. - It aggregates regional networks and connects them to Meta’s backbone. - Inter-BAG capacity can reach 16–48 petabits per second per regional pair. - Prometheus will span multiple buildings and connect tens of thousands of GPUs. ## Regional BAG Connectivity - BAG layers are distributed regionally to serve groups of L2 fabrics while respecting distance, latency, and buffer constraints. - Two connection topologies are used: - **Planar topology:** One-to-one connections between BAG switches in different regions; simpler to manage but creates more concentrated failure domains. - **Spread topology:** Links are distributed across switches and planes, improving path diversity and resilience. - The choice depends on site size and available fiber. ## Connecting DSF and NSF Fabrics - Meta’s L2 networks use both: - **Disaggregated Schedule Fabric (DSF)** - **Non-Scheduled Fabric (NSF)** - DSF zones across multiple buildings connect to BAG through backend edge pods. - NSF connects to BAG planes through matching Spine Training Switches. - Oversubscription is carefully managed: - L2-to-BAG oversubscription is typically around 4.5:1. - One NSF example has an effective ratio of 4.98:1. - BAG-to-BAG ratios vary by region and link capacity. ## Hardware and Routing - BAG uses modular chassis with Jericho3 ASIC line cards. - Each line card supports up to 432 800G ports. - Larger central-hub chassis support many spoke connections and long-distance links. - eBGP with link-bandwidth attributes enables Unequal Cost Multipath (UCMP), improving load balancing and failure recovery. - BAG-to-BAG links use MACsec for network security. ## Resilience and Failure Management - The design includes detailed port striping, IP addressing, and failure-domain analysis. - Failures are evaluated at the BAG, data-hall, and power-distribution levels. - Mitigation techniques include: - Draining affected BAG planes - Conditional route aggregation - Reducing blackholing risks during failures ## Managing Long-Distance Links - Distributed BAG architecture keeps L2-to-edge distances short, which benefits shallow-buffer NSF switches. - Longer BAG-to-BAG connections require deep-buffer switches. - These buffers provide headroom for lossless congestion-control mechanisms such as Priority Flow Control (PFC). ## Broader Impact BAG provides the networking foundation for Prometheus and future AI clusters. By combining regional aggregation, high-density hardware, resilient connection topologies, and fabric interoperability, Meta can scale AI infrastructure across multiple data centers while maintaining bandwidth, reliability, and operational flexibility.

Read original(opens in new tab)