ECE1756_lecture5_part2_2026_benchmarking_datacenter_fpgas_and_nn_inference
Watch on YouTube →
Overview
Vaughn Betz explains how FPGA system design shapes machine-learning inference, from PCIe and coherent CPU–accelerator links to the compute, memory, latency, and precision demands of CNNs, transformers, and other models. He argues FPGAs are strongest for low-latency, low-precision inference and details two ways to build accelerators: H-Pipe’s compiler-generated streaming pipelines and instruction-controlled overlays such as an NPU.
Key takeaways
- PCIe-connected FPGAs traditionally need physical-addressable buffers and may incur operating-system copies; CAPI and CXL add coherent virtual-address access but have not displaced PCIe for most bulk-transfer use cases.
- Batching converts repeated matrix-vector operations into matrix-matrix work, reusing weights and improving throughput at the cost of per-user latency.
- Low-precision inference is a natural FPGA target because configurable logic can shrink arithmetic hardware along with data storage; the useful precision depends on the model’s accuracy tolerance.
- FPGAs gain their clearest advantage when latency is tight, models are relatively small or stable, and custom on-chip memory and dataflow can keep computation close to data.
- H-Pipe’s compiler balances layer-specific processing elements into a streaming pipeline, while an overlay such as the Neural Processing Unit trades some specialization for easier model updates through software compilation.
- Very large language models weaken the FPGA case because their memory footprints require substantial DRAM capacity, bandwidth, and multi-device interconnects—strengths of purpose-built GPU platforms.
Chapters
- Microsoft Catapult used PCIe Gen 3 x8, described as providing about 8.5 GB/s; newer generations and wider links offer more bandwidth.
- Conventional PCIe transfers use physical addresses, so operating-system support may copy virtual-memory data into physical locations the FPGA can access.
- IBM CAPI and Intel CXL use PCIe signaling while enabling coherent access to virtual addresses and potentially CPU caches.
- Coherent links have seen limited adoption because bulk transfers often make the extra PCIe copy acceptable, while fine-grained CPU–accelerator cooperation is less commonly needed.
- Training creates a model and usually prioritizes total throughput; inference uses a trained model and often needs a prompt response.
- Embedded examples such as pedestrian detection in a car require low latency so an answer arrives in time to act.
- A self-driving Nvidia board cited in the lecture used 750 W—substantial beside an electric car’s roughly 6 kW average drivetrain demand.
- Data-center efficiency also matters at scale: forecasts cited in the lecture put data centers at 6% of global electricity demand by 2030.
- CNNs process images through convolutional layers and typically finish with matrix operations that produce classifications or object-detection bounding boxes.
- Convolution scans image regions with filters; implementing a convolution on an FPGA is the central building block in the referenced lab.
- RNNs feed outputs back into earlier layers, making them useful for sequential tasks such as converting speech to text.
- Across these models, matrix-vector multiplication is a recurring compute kernel.
- A graph neural network applies shared weight matrices to feature vectors on graph nodes, then aggregates information from connected neighbors.
- Repeated matrix transforms and neighbor aggregation let node features encode both learned weights and graph structure.
- For traffic prediction, a GNN can use street connections to infer that a blockage on Young Street may slow nearby Bay Street.
- The main operations are matrix-vector multiplication, communication between connected nodes, and aggregation such as averaging.
- A tokenizer converts text into tokens, which are represented as vectors and processed through transformer blocks containing matrix multiplications and attention.
- Attention blends information from related earlier tokens so a word’s representation reflects its context.
- Autoregressive generation repeatedly processes the model to produce one next token; a large model may contain around 80 sequential blocks.
- Prompt processing can batch input-token vectors into matrix-matrix operations, while decoding handles generated tokens one at a time and is more latency-sensitive.
- For an image tensor of roughly n³ elements, the lecture characterizes CNN work as scaling like n⁴k², where k is the convolution-filter width.
- Because CNN computation grows faster than its data footprint, efficient compute is the main design priority, with memory arranged to keep units fed.
- An n-by-n weight matrix contains n² values and takes order n² operations for a matrix-vector multiply, giving relatively little computation per loaded weight.
- Many neural-network inference workloads therefore face a memory-bandwidth challenge even when their arithmetic kernels are straightforward.
- Batching 30 users turns 30 input vectors into a matrix, converting matrix-vector work into matrix-matrix multiplication.
- The accelerator can reuse a loaded weight matrix across the batch, improving data reuse and keeping more compute units busy.
- Larger batches increase aggregate throughput but also make individual users wait longer; batch size is therefore a latency–utilization trade-off.
- FPGAs tend to fit low-batch, low-latency workloads, while GPUs are often attractive when higher latency and large batches are acceptable.
- FPGA logic can be built for application-specific widths: an n-bit adder uses roughly n lookup tables, while an n-bit multiplier can require roughly n² logic elements.
- Inference often tolerates 8-bit or 4-bit values, and some networks use binary weights or activations; reduced precision can save FPGA logic as well as memory.
- Mark Horowitz’s cited 45 nm scalar-CPU analysis estimated that a 32-bit floating-point addition used about 70 pJ overall, with only about 0.9 pJ for the addition itself.
- FPGAs offer parallel multipliers, configurable block-RAM systems, and custom sensor/actuator I/O, though programmable logic carries area, speed, and power overhead versus hardened gates.
- Deeply pipelined custom hardware and configurable memory make FPGAs attractive when inference latency must be low.
- Sparse networks with many zero weights can skip zero multiplications using custom FPGA logic and multiplexing.
- Small networks that fit mostly in on-chip memory can benefit from FPGA data locality; off-chip-heavy workloads lose some of that advantage.
- Large language models are a difficult fit because they exceed a single device’s memory capacity and depend on extensive DRAM bandwidth and multi-device networking, areas where Nvidia GPUs are specialized.
- Early FPGA designs accelerated individual operations or layers, but moving intermediate results between layers added communication overhead.
- Handwritten, model-specific RTL can accelerate an entire network efficiently, but requires substantial effort whenever the model changes.
- Domain-specific compilers generate hardware from a neural-network description and FPGA resource constraints, automating implementation at the cost of rerunning CAD tools.
- An overlay is a specialized soft processor: a software compiler can target new models with instructions without generating a new FPGA bitstream.
- H-Pipe generates a streaming dataflow architecture with specialized processing elements—for example, separate units for 7×7, 5×5, and 3×3 convolutions.
- Instead of buffering a whole layer between general-purpose processing elements, H-Pipe streams results directly to the next stage with limited buffering.
- Its compiler uses the network and FPGA resources, including DSPs, logic, and RAM, to assign parallelism and balance layer throughput.
- The generated RTL uses elastic dataflow interfaces; the lecture’s MobileNet v1 example illustrates layer-specific hardware, while later Altera DSP blocks improved low-precision throughput enough to outperform the compared GPU on latency and throughput.
- An FPGA overlay is configured once; new neural networks can then be compiled into instruction streams rather than new hardware bitstreams.
- Vaughn Betz’s group’s Neural Processing Unit uses large functional units, including a matrix-vector unit, and wide SIMD operations for machine-learning workloads.
- The design routes common dataflow directly between units—for example, from matrix-vector multiplication to ReLU—rather than repeatedly passing results through register files.
- Large instructions can launch about 40,000 operations, amortizing control overhead; the lecture cites roughly 10× GPU performance for some RNN workloads and identifies both generated dataflow hardware and overlays as major FPGA inference styles.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Vaughn Betz.