Save this video — free

ECE1756_lecture5_part2_2026_benchmarking_datacenter_fpgas_and_nn_inference

Vaughn Betz · 1:03:54 · Watch on YouTube

ECE1756_lecture5_part2_2026_benchmarking_datacenter_fpgas_and_nn_inference Watch on YouTube →

Overview

Vaughn Betz explains how FPGA system design shapes machine-learning inference, from PCIe and coherent CPU–accelerator links to the compute, memory, latency, and precision demands of CNNs, transformers, and other models. He argues FPGAs are strongest for low-latency, low-precision inference and details two ways to build accelerators: H-Pipe’s compiler-generated streaming pipelines and instruction-controlled overlays such as an NPU.

Key takeaways

Chapters

0:00 PCIe, CAPI, and CXL for CPU–FPGA Communication
6:46 Inference Prioritizes Latency and Power, Unlike Training
10:50 CNN Convolutions and RNNs for Images and Sequences
14:54 GNNs Combine Node Features with Graph Connectivity
19:30 Transformers, Autoregressive Decoding, and Token Latency
25:39 Compute Intensity: CNNs Versus Matrix-Vector Workloads
29:22 Batching Trades User Latency for Data Reuse and Throughput
33:45 FPGA Advantages in Precision, Energy, Parallelism, and I/O
41:28 Where FPGAs Fit—and Struggle—in Neural-Network Inference
47:25 From Layer Accelerators to Compilers and FPGA Overlays
50:49 H-Pipe Generates Balanced, Layer-Specific Streaming Pipelines
1:00:12 Instruction-Controlled FPGA Overlays and the Neural Processing Unit

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Vaughn Betz.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.