ECE1756_lecture5_part1_2026_benchmarking_datacenter_fpgas_nn_inference
Watch on YouTube →
Overview
Vaughn Betz uses an Intel-authored 2010 CPU–GPU benchmarking study to show why fair accelerator comparisons require optimized implementations, and why peak compute and bandwidth predict only some workloads. He then explains Microsoft’s Catapult FPGA deployments: a reusable shell-and-role design, multi-FPGA Bing acceleration, and a later network-connected FPGA architecture that supports services such as compression and encryption while keeping server power increases below 10%.
Key takeaways
- The Intel benchmarking study found an average GPU advantage of about 2.5× across diverse workloads, not the roughly 100× reported by some earlier comparisons; CPU cache blocking, SIMD, and multicore tuning explained much of the gap.
- Peak compute and memory-bandwidth ratios are useful for predicting clearly compute-bound or memory-bound workloads, but the study found that roughly half its benchmarks needed more nuanced analysis.
- Microsoft’s Catapult shell-and-role model separated stable interfaces such as PCIe, DDR, sensors, and routing from changeable application logic, enabling reuse while accepting shell costs of 23% of the first FPGA and 44% in Catapult v2.
- Catapult accelerated Bing by about 95% at the same latency, later reaching roughly 125%; Microsoft scaled from a prototype with more than 2,000 FPGAs to FPGA-equipped servers across its cloud.
- Putting FPGAs directly in the network path enabled compression, encryption, and FPGA-to-FPGA communication, while a flash-resident safe configuration protected CPU network access from failed or faulty FPGA deployments.
Chapters
- Lab 1 is due Thursday, October 15; a provided top-level design uses an end-copy parameter to instantiate multiple copies for power estimation.
- The copies are cascaded with distinct, bit-swizzled inputs so Quartus cannot optimize redundant instances away.
- Use multiple copies to estimate how dynamic power scales without compiling a full device; the test bench is only expected to pass when the copy count is one.
- The upcoming readings contrast overlay-style and dataflow-style FPGA neural-network accelerators; one reading is mandatory and the other optional.
- Betz introduces an optional 2010 Intel paper asking whether GPUs are really 100 times faster than CPUs.
- The paper compares standard workloads and argues that weak CPU baselines inflated some published GPU speedups.
- The paper evaluates roughly 15 workloads, including matrix multiplication, Monte Carlo simulation, convolution, FFTs, sparse matrix–vector operations, sorting, searching, and histograms.
- Across the study, the GPU’s geometric-mean speedup is about 2.5×, much lower than the roughly 100× claims in some original publications.
- A geometric mean suits speedup ratios because it treats reciprocal gains and losses more symmetrically than an arithmetic mean.
- In the paper’s hardware comparison, the four-core Core i7 runs at 3 GHz, while the GPU has 30 streaming multiprocessors at 1.3 GHz and about 4.7× the CPU’s memory bandwidth.
- The GPU has about six times the CPU’s peak floating-point capability, but the CPU has stronger double-precision capacity in this comparison.
- CPU matrix multiplication needs cache-aware blocking: reorganizing the computation into tiles keeps data reusable in cache instead of repeatedly fetching it from DRAM.
- The Intel team improved CPU results with cache blocking, SIMD instructions, and parallel execution across CPU cores rather than relying on unoptimized C code.
- Compiler intrinsics can explicitly request SIMD operations, while ordinary compiler vectorization may miss opportunities or decide data rearrangement costs too much.
- Some comparisons also used different algorithms: CPUs can benefit from more complex methods involving branching and indirection, while GPUs tend to favor regular, repetitive work.
- GPU threads that take different branches reduce SIMD efficiency; a 32-lane group may execute both sides of a branch with predication masks.
- The study could predict seven workloads reasonably well from memory-bandwidth or compute-throughput ratios, but seven others required more detailed analysis.
- Peak performance is a useful first estimate when a workload is clearly memory-bound or compute-bound, but mixed bottlenecks and algorithm choices can change the result.
- Microsoft first targeted Bing search, where it sought higher throughput without increasing search latency and required large-scale reliability.
- As Moore’s law slowed, routine CPU-generation upgrades no longer delivered enough performance, making specialized accelerators more attractive.
- FPGAs could fit into existing server designs at lower power than GPUs, helping Microsoft add acceleration without rebuilding its infrastructure around high-power cards.
- Data centers contain hundreds of thousands of servers, and power availability has become a major constraint; some new facilities use on-site natural-gas generation.
- Microsoft valued homogeneous servers because consistent hardware simplifies scheduling, maintenance, and software development.
- Catapult FPGAs had to fit existing servers, use modest PCIe power budgets, and avoid JTAG cables or other desk-scale workflows that do not work across thousands of machines.
- The first Catapult design added an FPGA card to each server and connected 48 FPGAs in a torus network within a half-rack.
- The torus used serial links rated at about 20 gigabits per second per link and enabled FPGA-to-FPGA communication.
- A torus wraps connections around the edges, giving each FPGA four neighbors, but requires additional cabling and creates a network topology developers must understand.
- Microsoft defined a reusable shell containing PCIe, sensors, FPGA-to-FPGA routing, and local DDR interfaces, leaving application-specific logic to a separate role.
- The shell was floorplanned, validated for timing, and locked after placement and routing so application developers would not have to reimplement low-level interfaces.
- The first-generation shell consumed about 23% of a roughly 500,000-LE Stratix V FPGA, a substantial area cost Microsoft accepted for reuse and reliability.
- Microsoft divided Bing search processing into a pipeline spanning seven FPGAs, with an eight-FPGA torus ring providing one spare device.
- A spare FPGA could be programmed as a pass-through, allowing the pipeline to route around one failed FPGA without disabling the entire ring.
- Server monitoring and FPGA housekeeping logic helped detect faults and notify data-center control software; failures could arise from overheating, aging, electromigration, or hardware connections.
- Microsoft needed FPGA roles to adapt quickly because Bing’s feature extraction and machine-learning components changed frequently.
- A domain-specific language let software developers describe features to extract from search queries and documents, while a custom compiler generated finite-state-machine control for a reusable datapath.
- This programming model prioritized development speed and keeping up with changing applications, rather than maximizing datapath efficiency alone.
- The first Catapult deployment improved Bing throughput by about 95% at the same latency; later changes raised the improvement to roughly 125%.
- Microsoft initially accelerated the CPU-critical portions of Bing, with remaining software—not FPGA speed—limiting further gains.
- A prototype using more than 2,000 FPGAs led to broader deployment: FPGAs went into Bing servers and then, in Catapult v2, into every server Microsoft purchased.
- Catapult v2 placed the FPGA directly on the network path, so CPU traffic reached the network through the FPGA rather than the FPGA serving only as a CPU offload.
- This enabled FPGA and CPU compute to operate more independently and supported low-latency processing of network traffic.
- The FPGA could perform bump-in-the-wire tasks such as compression and encryption, reducing CPU work and making data-center bandwidth more effective.
- Catapult v2 replaced the custom 48-FPGA torus with hierarchical network switching across racks and data-center regions, improving scalability and simplifying cabling.
- The network-connected FPGA design became a smartNIC-style architecture, although the reusable shell grew to about 44% of the FPGA as features were added.
- Microsoft kept a safe FPGA configuration in flash that restored basic CPU-to-network connectivity if a bad application bitstream or FPGA failure threatened server access.
- IBM acquired Netezza, which used FPGAs in disk controllers to filter database-query data and reduce downstream processing by about 10×.
- Data processing units and infrastructure processing units extend the Catapult idea by placing programmable acceleration between CPUs and networks.
- Data-center operators prefer homogeneous servers, but machine-learning demand has made specialized hardware—including GPUs and Microsoft’s Maia accelerator—difficult to avoid.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Vaughn Betz.