ece1756_lecture3_part1_2026_compute_models_continued
Watch on YouTube →
Overview
Vaughn Betz explains how FPGA streaming designs gain performance through pipelining, automatic retiming, and time-domain multiplexing, while preserving behavior across cycles and independent data streams. He connects these techniques to practical Quartus design choices—including synchronous resets, stream-ID tagging, and C-slow retiming—and shows how they can improve clock rate or throughput without simply duplicating hardware.
Key takeaways
- FPGA pipelining is especially valuable because registers are plentiful and programmable interconnect is slow; balanced stages reduce the longest register-to-register delay.
- Retiming moves register boundaries without changing externally visible behavior or cycle latency, allowing Quartus to rebalance timing automatically while preserving a design's functional contract.
- Quartus retimes at multiple stages because early passes can move registers cheaply but lack accurate routing delays, whereas late passes have better timing information but can disrupt placement and routing.
- Synchronous resets are generally preferable in FPGA designs because they are easier to time-analyze and retime; asynchronous resets require careful synchronization to avoid unreliable reset behavior.
- Time-domain multiplexing lets one fast operator serve multiple slower streams; carrying a stream ID through the pipeline reliably routes each result even when latency changes.
- C-slow retiming can raise loop clock frequency by roughly C times and increase aggregate throughput across C interleaved streams, but it does not accelerate any one stream's dependency chain.
Chapters
0:00
Course Logistics and a Teaser on FPGA Compute Models
- The next class will compare FPGA compute models with processor execution, including a paper that implements both models on the same FPGA.
- The paper compares streaming dataflow against instruction execution to isolate how the compute model affects efficiency.
- Assignment 1 is delayed while Vaughn Betz updates it for Intel Quartus and Agilex devices and checks licensing.
3:40
Elastic and Inelastic Streaming Dataflow Recap
- Inelastic pipelines separate computations with simple registers, while elastic pipelines use valid signals, back pressure, and FIFOs.
- Designers can combine the styles, using inelastic pipelines inside operators and elastic dataflow between larger blocks.
- Both styles support repeated computations common in FPGA designs.
5:25
Why FPGA Designs Benefit Especially from Pipelining
- FPGA logic blocks contain many pre-fabricated registers, making moderate pipelining relatively inexpensive before registers become a resource limit.
- Programmable interconnect is slower than ASIC metal wiring, increasing critical-path delays and the value of dividing logic into shorter stages.
- A critical path is the longest path between registers in a clock domain; adding balanced pipeline stages can raise the achievable clock frequency.
11:00
Unbalanced Pipeline Stages Limit Clock Frequency
- A manually inserted register can shorten one path but leave the other stage with substantially more LUT and interconnect delay.
- A better register cut balances the logic on both sides, but RTL authors may not know the eventual LUT mapping or routed-wire delays.
- Pipeline placement is a human RTL change in general, and adding registers must preserve input alignment and avoid changing loop behavior.
13:13
Retiming Moves Registers Without Changing Latency
- Retiming moves existing registers across combinational logic to balance timing paths; it changes register boundaries without adding pipeline latency.
- Moving a register across a logic block can require duplicating it at multiple fan-outs so the original cycle boundaries are preserved.
- In the example, balancing a 2 ns stage against a 6 ns stage produces two roughly 4 ns stages, improving the limit from about 167 MHz to 250 MHz.
19:00
Why Quartus Can Apply Retiming Automatically
- Retiming preserves the input/output behavior and cycle latency, so a CAD tool can generally perform it safely where arbitrary repipelining could break control logic.
- Modern FPGA CAD tools commonly enable retiming by default; recent Quartus versions can retime Agilex designs during implementation.
- Retiming does not set the operating frequency: designers set clock constraints and PLL or oscillator settings, while timing analysis reports slack and whether constraints are met.
25:10
Why Quartus Retimes at Multiple Implementation Stages
- Early in the flow, Quartus knows the LUT structure but has limited information about interconnect delay, which is often dominant in an FPGA.
- Placement provides estimates of wire delay, and routing provides more accurate delays, but retiming later can require changing placement and ripping up existing routes.
- Multiple retiming passes trade early flexibility against later timing knowledge and reduce the risk of disruptive changes or poor convergence.
29:00
Register Placement, Retiming Tradeoffs, and Debugging
- Putting several registers at a module input is easy, but it relies heavily on the retimer to distribute them; a reasonable initial RTL placement can give the CAD tool a better starting point.
- Retiming can make hardware debugging harder because implementation registers may no longer correspond clearly to named RTL state.
- ModelSim RTL simulation is the usual place for most validation; hardware observation can inspect registers, while preserving internal combinational signals may constrain synthesis optimization.
- Retiming can increase register count when registers are duplicated across fan-outs, though it can also reduce registers by merging equivalent boundaries.
38:30
Synchronous Resets Simplify FPGA Timing and Retiming
- A synchronous reset takes effect on a clock edge and is easier for timing analysis and retiming than an asynchronous reset.
- Vaughn Betz recommends synchronous resets for FPGA RTL; the Assignment 1 sample code uses them.
- A synchronous clear can be treated like logic on a register input, making its timing behavior easier to analyze.
43:00
Asynchronous Reset Hazards and Synchronization
- An asynchronous clear changes register state independently of the clock; a reset near a clock edge can leave registers in inconsistent states.
- Asynchronous reset timing is difficult to constrain, and designs without proper synchronization can reset unreliably.
- A safe asynchronous reset requires synchronizing it to the clock before distributing it, adding complexity and making retiming harder.
- For ordinary FPGA designs, using synchronous resets avoids these timing and implementation complications.
46:30
Time-Multiplexing a Fast IP Block Across Streams
- A block running at 400 MHz can serve two upstream streams running at 200 MHz by alternating their data through one instance.
- A stream ID travels with data through FIFOs and the pipelined IP block, then selects the correct output FIFO.
- Carrying the ID through the pipeline keeps stream routing correct even if the block's latency changes.
- This technique is useful when applications such as cellular signal processing have multiple antenna or frequency-band streams.
53:00
Interleaving Independent Streams to Pipeline a Loop
- Adding a register to a feedback loop changes its iteration latency and can break a recurrence that depends on the previous iteration's result.
- With two independent streams, alternate inputs each cycle so each stream returns to the loop only after the extra cycle has elapsed.
- The loop can then run at a higher clock rate, while each stream retains its required dependency order.
- The added register does not make one stream produce results faster; it fills otherwise idle cycles with work from another stream.
1:01:30
C-Slow Retiming Adds Registers for Multiple Threads
- C-slow retiming replaces each register in a selected block with C registers and interleaves C independent input streams.
- After duplicating registers, retiming can redistribute them across logic to shorten critical paths, including paths through loops.
- For a small C such as 2, 3, or 4, the loop may run close to C times faster, subject to register delay and other implementation limits.
- The gain is aggregate throughput across independent computations, not a faster result for an individual stream.
1:05:00
Stream IDs Keep Multiplexed Dataflow Composable
- C-slow retiming changes latency, so stream IDs are a robust way to route results back to the correct stream rather than relying on cycle-counting control logic.
- Valid bits, back pressure, and stream IDs provide distributed control that lets operators and FIFOs make local decisions.
- Centralized cycle tracking is fragile: changing a pipeline's latency can invalidate the controller and break the larger design.
- The lecture closes the streaming-dataflow section before moving on to other compute models.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Vaughn Betz.