ECE1756_lecture4_part1_2026_compute_device_comparison
Watch on YouTube →
Overview
Vaughn Betz compares CPUs, GPUs, DSPs, and FPGAs by how they trade single-thread speed, throughput, power, programmability, and control over data movement; there is no universally best device, so performance depends on the workload and how much optimization effort it justifies. He explains CPU speculation and caches, GPU SIMT and latency hiding, DSP specialization, and FPGA dataflow, using concrete examples such as the UG machines’ 8-core Intel CPU and Nvidia Ampere GPU.
Key takeaways
- High-end CPUs prioritize single-thread performance through branch prediction, out-of-order execution, and caches; this control hardware can consume substantial area and power beyond the ALUs doing the arithmetic.
- GPUs improve throughput rather than single-thread latency: hundreds or thousands of similar threads let the hardware switch away from memory-stalled work, making regular workloads such as graphics and matrix operations a strong fit.
- SIMT makes GPU parallelism easier to express than CPU SIMD intrinsics, but programmers still need to organize work so many nearby threads execute the same operation.
- DSPs trade generality for power-efficient repetitive signal processing, using features such as multiply-accumulate units, fixed-point arithmetic, and programmer-controlled scratchpad memory.
- FPGAs provide the most direct control over custom pipelines and data movement, and can avoid instruction overhead, but require explicit scheduling and are particularly efficient with small integer data.
- Meaningful accelerator comparisons require a fair baseline: CPU code should be optimized comparably to FPGA or GPU code rather than treated as an unoptimized compile-and-run reference.
Chapters
- Assignment 1 was released after licensing and tool-version issues; it was due Thursday, October 15.
- The required Microsoft paper describes deploying FPGAs in data-center network interface cards; Vaughn Betz estimated the deployment exceeded one million servers and was growing about 30% annually.
- An optional Intel benchmarking paper examines whether GPU-versus-CPU speedup claims are fair when CPU code has not been comparably optimized.
- The course moves from compute models to concrete CPU, GPU, and FPGA comparisons.
- There is no fixed claim such as FPGAs being 3.3 times faster in all cases; the goal is to estimate whether a workload merits substantial optimization effort.
- Device choice must account for application needs, including available I/O and the amount of work that can be parallelized.
- A CPU fetches and decodes instructions, then dispatches them to one or more arithmetic logic units (ALUs).
- The register file holds the CPU’s execution context, but its limited capacity makes frequent access to main memory necessary.
- High-end CPUs add caches and control hardware to keep execution moving despite slow memory.
- A high-end CPU may examine a window of roughly 100–200 instructions and issue up to about six instructions per cycle when it finds independent work.
- Branch predictors guess the paths of loops and conditionals before earlier instructions resolve; wrong-path instructions are squashed before their results become visible.
- Out-of-order logic tracks which instructions have ready data and can execute them early, while preserving the program’s visible instruction order.
- At roughly 3 GHz, a CPU cycle is about 300 picoseconds, while a DRAM access may take around 50 nanoseconds—roughly a couple hundred cycles.
- Caches and memory prefetchers reduce stalls; a prefetcher can recognize sequential reads such as addresses increasing by four bytes and bring later data into cache.
- CPUs add SIMD instructions and typically 4–64 cores, but programmers or compilers must expose vector and thread parallelism to benefit.
- Unlike a high-end CPU, a GPU generally omits large out-of-order machinery and uses many ALUs organized around SIMD-style execution.
- GPU workloads use large numbers of threads—often hundreds or thousands—that perform similar operations on different data.
- Each thread has its own registers and execution context, reducing competition for a single central register file but making inter-thread communication more involved.
- When one GPU thread group waits on a memory request, the GPU switches to other ready groups instead of relying on CPU-style speculation to find nearby work.
- A request may take around 200 cycles, but many independent threads let useful work continue while data returns.
- This suits regular workloads such as graphics, where many independent pixels run the same program, but GPUs are poorly suited to running many unrelated programs like a general-purpose Linux workload.
- Intel’s SIMD progression included MMX for integers, SSE with four-way floating-point SIMD, AVX2 with eight-way single-precision SIMD, and AVX-512 with 16-way SIMD on some processors.
- Using 16-bit operands can increase parallelism because narrower arithmetic units and data paths can support more operations per cycle.
- The UG machines’ CPU has 8 Intel cores, a 14 nm process, a 2.5–4.9 GHz clock range, 16-way floating-point SIMD, and a 65 W power rating.
- The Nvidia Ampere GPU used for assignment comparisons has 48 streaming multiprocessors (SMs), the approximate counterpart of CPU cores.
- Each SM has four execution units that can each perform 32-way SIMD with FP16 data, giving an approximate aggregate width of 128 operations across the four units.
- An SM contains about 16,000 32-bit registers for its many threads and 192 KB of L1/shared memory; programmers often manage shared memory explicitly as scratchpad RAM.
- The Ampere GPU has a shared 4 MB L2 cache, a roughly 1.5 GHz clock, and a 290 W power rating.
- Tensor Cores accelerate matrix multiplication with dedicated streaming hardware, especially useful for graphics and deep learning’s repeated matrix operations.
- Vaughn Betz contrasts the 290 W GPU with newer AI-focused accelerators that can reach about 1.4 kW and require liquid cooling.
- Nvidia’s SIMT model groups threads performing the same operation so the hardware can execute them in SIMD-like fashion.
- CUDA lets programmers express a grid of threads—for example, one thread per pixel in a 1,000-by-2,000 image—rather than manually specifying every vector instruction.
- CPU compilers can auto-vectorize some loops, but programmers seeking more control often use compiler intrinsics; those instructions can require rewriting code when SIMD width changes.
- A DSP is a specialized processor designed for repetitive signal-processing loops, particularly multiply-accumulate operations used in filtering and convolution.
- The TI C6000 example uses a VLIW design with four address-generation units, two general ALUs, and two multiply-accumulate units.
- DSPs often emphasize fixed-point arithmetic and low power; the described units can use a datapath for 32-bit floating point or split it for smaller fixed-point operations.
- DSPs may offer caches for convenience or programmer-managed scratchpads for deterministic memory access and bounded latency.
- DSP products can add FFT or other application-specific accelerators, combine a DSP with an ARM control processor, and commonly target roughly 1–20 W.
- High-end CPUs, GPUs, and FPGAs use much larger power budgets; FPGA power depends on device capacity and signal activity, with high-end examples around 20–100 W.
- FPGAs let designers build custom processing pipelines and distributed memory systems without requiring an instruction stream for the main computation.
- They are especially efficient with small data types and integers; newer FPGA DSP blocks improve floating-point support, though fixed-point implementations can remain more efficient.
- Compared with CPUs, GPUs, and DSPs, FPGAs require more explicit scheduling and data movement, so the right choice depends on workload structure and programming effort.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Vaughn Betz.