Save this video — free

ECE1756_lecture5_part1_2026_benchmarking_datacenter_fpgas_nn_inference

Vaughn Betz · 1:24:26 · Watch on YouTube

ECE1756_lecture5_part1_2026_benchmarking_datacenter_fpgas_nn_inference Watch on YouTube →

Overview

Vaughn Betz uses an Intel-authored 2010 CPU–GPU benchmarking study to show why fair accelerator comparisons require optimized implementations, and why peak compute and bandwidth predict only some workloads. He then explains Microsoft’s Catapult FPGA deployments: a reusable shell-and-role design, multi-FPGA Bing acceleration, and a later network-connected FPGA architecture that supports services such as compression and encryption while keeping server power increases below 10%.

Key takeaways

Chapters

0:00 Lab 1 Power Estimation with Replicated FPGA Designs
6:30 Course Readings and the CPU–GPU Benchmarking Question
10:30 Benchmark Results, Geometric Means, and Peak Hardware
15:00 CPU–GPU Hardware Differences and CPU Cache Blocking
21:00 Why CPU Baselines Need SIMD, Threads, and Algorithm Tuning
27:30 GPU Branch Divergence and the Limits of Peak-Performance Estimates
35:00 Why Microsoft Put FPGAs into Data-Center Servers
38:30 Data-Center Scale, Power Budgets, and Server Homogeneity
44:00 Catapult v1: FPGA Cards and a 48-Node Torus
47:30 Catapult’s Shell-and-Role Model Simplifies FPGA Deployment
52:00 Pipelining Bing Across FPGAs and Routing Around Failures
57:30 Changing Workloads and a Custom Compiler for Bing Features
1:05:00 Catapult’s Throughput Gains and Whole-Cloud Deployment
1:10:30 Catapult v2: Putting the FPGA Between CPU and Network
1:14:30 Hierarchical Switching, SmartNIC Safety, and Failure Recovery
1:21:00 Other FPGA Data-Center Uses and the Cost of Specialization

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Vaughn Betz.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.