Amd Epyc

3 posts

cloudflare3 min readCurated summary

Launching Cloudflare’s Gen 13 servers- trading cache for cores for 2x edge compute performance

Cloudflare’s Gen 13 servers use AMD EPYC 5th Gen Turin processors to provide up to twice as many cores as Gen 12. However, Turin’s much smaller per-core cache caused the legacy FL1 request-handling layer to suffer severe latency increases, despite higher throughput. Cloudflare found that tuning alone could not fully solve the problem, reinforcing the need for FL2, a Rust-based rewrite designed to scale with cores rather than depend heavily on cache. ## Turin’s Core-Heavy Architecture - Gen 13 Turin processors offer: - Up to 192 cores and 384 SMT threads, compared with Gen 12’s 96 cores. - Improved instructions per cycle through the Zen 5 architecture. - Up to 32% lower power consumption per core. - DDR5-6400 support for greater memory bandwidth. - The tradeoff is substantially less cache: - Gen 12 Genoa-X provides 12 MB of L3 cache per core through 3D V-Cache. - The 192-core Turin 9965 provides only 2 MB per core. - This architecture favors aggregate throughput but challenges workloads dependent on cache locality. ## FL1’s Cache and Latency Problems - FL1, based on NGINX and LuaJIT, was optimized for Gen 12’s large cache. - AMD uProf measurements showed: - Dramatically higher L3 cache miss rates on Turin. - More requests requiring slow DRAM access. - Increasing latency as CPU utilization and cache contention rose. - An L3 hit takes roughly 50 CPU cycles, while a DRAM fetch can take more than 350 cycles. - As a result, Gen 13’s additional cores delivered throughput gains but introduced unacceptable latency penalties. ## Throughput Gains at an Unacceptable Cost - With FL1, Gen 13 produced: - 10% more throughput on the 128-core Turin 9755. - 31% more on the 160-core Turin 9845. - 62% more on the 192-core Turin 9965. - The Turin 9965 offered the strongest total-cost-of-ownership benefits. - However, latency increased by more than 50% at high CPU utilization, which would negatively affect customer experience and violate performance requirements. ## Hardware and Resource Tuning - Cloudflare tested several mitigations with AMD: - Hardware prefetcher and Data Fabric Probe Filter adjustments produced only marginal improvements. - Adding FL1 workers increased throughput but took resources away from other services. - CPU pinning and isolation provided limited benefits. - AMD’s Platform Quality of Service (PQOS) was used to control cache and memory-bandwidth sharing across Turin’s Core Complex Dies. ## Cache Isolation with PQOS - Reserving part of a single CCD’s cache for FL1 produced less than 5% additional throughput. - Configurations assigning FL1 50–75% of each CCD’s cache also delivered less than 5% improvement and caused minor degradation elsewhere. - A socket-level approach was more successful: - Six of twelve CCDs, aligned with a NUMA domain, were dedicated to FL1. - This provided more than 15% incremental throughput while keeping latency acceptable. - These results showed that workload placement and cache locality could help, but they were not a complete substitute for software designed around Turin’s cache profile. Cloudflare’s broader solution was FL2, a Rust-based rewrite of its core request-handling layer. By reducing dependence on large per-core caches, FL2 enabled Gen 13’s higher core count to translate into scalable edge-compute performance without the latency penalties seen with FL1.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Inside Gen 13- how we built our most powerful server yet

Cloudflare’s Gen 13 server is a hardware redesign aligned with its Rust-based FL2 request-processing stack. By choosing a 192-core AMD EPYC Turin 9965, doubling memory, expanding storage and networking, and adding stronger security and accelerator support, Cloudflare targets up to twice the throughput of Gen 12 while remaining within latency limits. The design also improves efficiency, rack density, and operational simplicity. ## Gen 13 at a Glance - Uses a 2U, single-socket design. - Key specifications: - 192-core AMD EPYC 9965 processor - 768 GB of DDR5-6400 memory - Three 7.68 TB E1.S PCIe 5.0 NVMe drives - Dual 100 GbE OCP 3.0 networking - 1,300W Titanium-grade power supply - ASPEED AST2600 BMC and AST1060 hardware root of trust - Compared with Gen 12, Gen 13 provides: - Up to 2× throughput - Up to 50% better performance per watt - Up to 60% more throughput per rack at the same power budget - Twice the memory capacity, 1.5× the storage, and 4× the network bandwidth - PCIe encryption in addition to memory encryption - Better support for high-heat PCIe accelerators ## Choosing the CPU - Gen 12 used a 96-core AMD EPYC Genoa-X 9684X with: - 400W TDP - 1,152 MB of L3 cache - Cloudflare evaluated three Turin processors: - 9755: strongest per-core performance - 9845: lower socket power and fewer cores - 9965: highest core count and better efficiency - The Turin 9965 was selected with 192 cores and 384 threads, doubling Gen 12’s hardware threads. - Its L3 cache is much smaller—384 MB total, or 2 MB per core versus Gen 12’s 12 MB per core—but FL2 workloads depend less on large L3 caches than the previous FL1 stack. - FL2 scales nearly linearly with additional cores, allowing the 9965 to deliver up to 100% higher throughput. - Production testing showed the 9965 achieved the best aggregate requests per second and favorable performance per watt at its 500W TDP. - Higher compute density also means fewer servers to provision, patch, monitor, and operate. - Turin’s support for DDR5-6400, PCIe 5.0, and CXL 2.0 provides a longer upgrade and security-support runway. ## Memory Bandwidth and Capacity - Gen 13 doubles memory from 384 GB to 768 GB while retaining 4 GB per core. - All twelve memory channels are populated using one 64 GB DDR5-6400 ECC RDIMM per channel. - This configuration delivers approximately 614 GB/s of peak memory bandwidth per socket, a 33.3% increase over Gen 12. - Using identical DIMMs across all channels enables balanced interleaving, distributing memory accesses across the full memory subsystem. - The design is intended to prevent the 192-core processor from being starved of data during highly parallel workloads. Cloudflare’s central design choice was to match hardware to the characteristics of FL2 rather than preserve Gen 12’s cache-heavy strategy. For workloads that scale well across cores, the Turin 9965 and fully populated memory system offer higher throughput, better rack economics, and simpler fleet operations.

Read original(opens in new tab)
aws2 min readCurated summary

Amazon EC2 Hpc8a Instances powered by 5th Gen AMD EPYC processors are now available | Amazon Web Services

Amazon EC2 Hpc8a instances are now generally available for tightly coupled, compute-intensive HPC workloads. Powered by 5th Gen AMD EPYC processors reaching 4.5 GHz, they provide up to 40% more performance, 42% higher memory bandwidth, and 25% better price-performance than Hpc7a instances. AWS targets applications such as fluid dynamics, weather modeling, design simulations, and crash analysis. ## Instance Specifications - Available in a single `96xlarge` configuration: - 192 CPU cores - 768 GiB memory - 300 Gbps Elastic Fabric Adapter (EFA) networking - Uses a 1:4 core-to-memory ratio. - Customers can customize the number of cores at launch to better match workload requirements. - Simultaneous Multithreading (SMT) is disabled to maximize HPC performance. - Sixth-generation AWS Nitro cards handle virtualization, storage, and networking tasks separately from the CPUs. ## Supported HPC Services - Integrates with AWS ParallelCluster and AWS Parallel Computing Service (AWS PCS) for cluster creation and job submission. - Supports Amazon FSx for Lustre, offering sub-millisecond latency and throughput of up to hundreds of gigabytes per second. - High-bandwidth, low-latency networking is designed for workloads requiring extensive communication between compute nodes. ## Availability and Purchasing - Initially available in: - US East (Ohio) - Europe (Stockholm) - Offered through On-Demand Instances and Savings Plans. - Regional availability and future expansion can be checked through AWS Capabilities by Region. Hpc8a instances are best suited for organizations needing faster simulation results and efficient scaling across tightly coupled HPC workloads. Teams can launch them through the Amazon EC2 console and combine them with AWS cluster and storage services for a complete HPC environment.

Read original(opens in new tab)