Save this video — free

ECE1756_lecture4_part1_2026_compute_device_comparison

Vaughn Betz · 1:10:37 · Watch on YouTube

ECE1756_lecture4_part1_2026_compute_device_comparison Watch on YouTube →

Overview

Vaughn Betz compares CPUs, GPUs, DSPs, and FPGAs by how they trade single-thread speed, throughput, power, programmability, and control over data movement; there is no universally best device, so performance depends on the workload and how much optimization effort it justifies. He explains CPU speculation and caches, GPU SIMT and latency hiding, DSP specialization, and FPGA dataflow, using concrete examples such as the UG machines’ 8-core Intel CPU and Nvidia Ampere GPU.

Key takeaways

Chapters

0:00 Assignment 1, CAD Licensing, and Microsoft FPGA Readings
5:20 Comparing Devices by Workload, Not a Universal Speedup
6:37 High-End CPU Structure: Instructions, Registers, and ALUs
8:46 Out-of-Order Execution and Branch Prediction Keep CPUs Busy
15:53 CPU Memory Latency, SIMD, and Multicore Tradeoffs
22:18 GPU Throughput Architecture: SIMD and Many Threads
27:20 How GPUs Hide Memory Latency and Target Throughput
31:41 CPU SIMD Evolution and UG Machine Comparison Baseline
38:06 Nvidia Ampere Streaming Multiprocessors and Memory
45:28 Ampere Tensor Cores, Cache, and GPU Power
49:41 SIMT Programming Versus Explicit CPU SIMD
57:38 DSP Processors: Efficient Multiply-Accumulate for Signal Processing
1:03:08 DSP Memory Control and Cross-Device Power Tradeoffs
1:08:54 FPGA Strengths: Custom Dataflow and Explicit Scheduling

Keep these chapters and the full searchable transcript in your own library.

Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Vaughn Betz.

Want the full transcript?

Save this video in YouTube Collector to get its complete searchable transcript, your own AI summaries, and a library that keeps every video you collect in one place.