CSCE 611 Fall 2026 Lecture 8: RISC-V Microarchitecture 1
Watch on YouTube →
Overview
Jason D. Bakos reviews SystemVerilog timing and coding semantics, then introduces the RISC-V microarchitecture required for Lab 3: a 22-instruction, three-stage CPU implemented with an instruction ROM, 32x32 three-port register file, combinational ALU, and CSRRW-based FPGA I/O. The lecture emphasizes nonblocking assignments, pipeline timing, RISC-V R/I/U encodings, register-file bypassing, and the difference between instruction latency and throughput.
Key takeaways
- SystemVerilog nonblocking assignments use old right-hand-side values, so sequential code such as sum_plus_one <= sum + 1 produces a one-cycle-delayed result rather than the newly computed sum.
- RISC-V's fixed RS1, RS2, and RD fields simplify register-file wiring, but instruction identification still requires checking opcode, funct3, and sometimes funct7.
- The Lab 3 CPU uses a three-stage fetch-execute-write-back pipeline: each instruction takes three cycles to traverse the processor, while steady-state throughput reaches one instruction per cycle.
- Register-file read-after-write behavior requires explicit bypassing so a simultaneous read and write of the same register returns new write data; register zero must always read as zero.
- CSRRW is adapted for FPGA I/O with CSR address F00 for 18 switches and F02 for hex displays, replacing the earlier special-register-30 mechanism.
- The CPU's instruction ROM is initialized from a RARS-generated program.rom file with $readmemh, so changing firmware normally requires resynthesizing the FPGA design.
Chapters
- A SystemVerilog always_comb multiplexer without a default path can infer a latch; always_comb reports an error because it is designed to detect incomplete combinational assignments.
- Blocking assignments execute sequentially within an always_ff block, so swapping A and B with A = B; B = A; makes both registers take the old B value.
- Nonblocking assignments evaluate right-hand sides from the previous cycle; a synchronous reset clears out only on the next rising clock edge.
- An asynchronous reset in an always_ff block responds immediately to reset assertion, independent of the clock.
- The state-machine example resets to 011 and updates from the previous value because nonblocking assignments defer register updates until the end of the clock event.
- Unsigned comparisons drive different concatenation expressions; the traced sequence reaches 1010, or decimal 10, after four cycles.
- State-vector questions are best solved with a cycle table that records the old bits used by every right-hand-side expression.
- Concatenation syntax such as {a[1:0], a[3]} changes bit positions explicitly, but individual bit assignments can be easier to audit.
- An always_comb block can use if statements and blocking assignments, provided every logical path assigns its output; a missing explicit default is not automatically a latch.
- A combinational block that assigns a = a + 1 creates a feedback loop because no clock separates the old and new values; synthesis should reject the sequential behavior inside always_comb.
- Testbenches may change only one input, allowing unchanged inputs to retain their values; the demonstrated four-case truth table implements an AND gate.
- Bitwise XOR with 111 flips every bit, while left shifts and nonblocking assignments can produce a repeating three-bit state such as 101.
- The CPU uses instruction memory for software and a 32x32 register file with two read ports and one write port.
- RISC-V arithmetic instructions use a three-address format: two source registers, RS1 and RS2, plus one destination register, RD.
- Unlike programmer-visible Intel two-address instructions, RISC-V exposes the three-register structure directly; Intel internally decomposes instructions into proprietary micro-operations.
- The planned instruction RAM is approximately 4,096 entries wide by 32 bits, while the register file is optimized for simultaneous operand reads and result writes.
- The register file behaves like RAM but is implemented on the FPGA as addressed banks of flip-flops with multiplexers and decoders rather than a global-resettable register array.
- Reading register zero always returns zero, and initializing all 32 registers requires writing zero to each entry over 32 clock cycles because the file has no global reset.
- RISC-V places RS1, RS2, and RD in fixed instruction bit fields, allowing direct wiring from the decoded instruction to the register-file ports.
- When a register is read and written in the same cycle, nested ternary bypass logic forwards write data so the read sees the new value instead of the old RAM value.
- The supplied ALU is purely combinational, with inputs A and B, result R, and an operation code; all propagation must complete within one clock cycle.
- Supported operations include bitwise AND, OR, XOR, addition, subtraction, multiplication, high signed and unsigned multiplication, logical shifts, comparisons, and arithmetic right shifts.
- SystemVerilog treats ordinary logic vectors as unsigned, so signed multiplication uses an explicit $signed type cast.
- Because the expected arithmetic-shift operator did not work reliably in the course implementation, a blocking-assignment barrel shifter conditionally shifts by 1, 2, 4, 8, and 16 bits while replicating the sign bit.
- Lab 3 requires a complete RISC-V CPU in SystemVerilog, substantially more difficult than Lab 2.
- The supported set contains 22 instructions spanning arithmetic R-type and I-type operations, comparisons, logical operations, shifts, the U-type load-upper-immediate instruction, and CSRRW.
- Branches and jumps are deliberately excluded, so programs execute straight-line instruction sequences without data-dependent if statements or loops.
- The instruction memory eventually wraps the program counter back to zero, creating a top-level infinite repetition while each program pass remains linear.
- R-type instructions use RS1, RS2, and RD for two-register ALU operations; I-type instructions replace RS2 and funct7 with a 12-bit immediate; U-type instructions use a 20-bit immediate and RD.
- The U-type load-upper-immediate operation writes its 20-bit immediate into the upper 20 bits of a destination register and ignores both register-file read ports.
- RISC-V uses opcode, funct3, and sometimes funct7 fields to identify instructions, so the control unit must inspect multiple instruction regions.
- Unlike MIPS, RISC-V keeps source and destination register fields in consistent locations, simplifying extraction and datapath wiring.
- The course replaces an earlier special-register-30 design with CSRRW, a RISC-V control-and-status-register instruction adapted for FPGA I/O.
- In the simplified course implementation, an input CSR address reads switches into RD, while an output CSR address writes RS1 data to the hexadecimal displays rather than performing a full architectural swap.
- The hardwired CSR addresses are F00 for the 18 switches and F02 for the hex displays.
- Real systems commonly use programmed I/O through load and store instructions; CSRRW is used here to avoid implementing memory instructions in the first CPU lab.
- The FPGA top-level module instantiates the CPU between the physical switches and existing seven-segment decoders, while the CPU generates display data through CSRRW.
- Instruction memory is a 32-bit-wide array initialized with $readmemh from program.rom, which is exported from assembled RISC-V code in RARS.
- Changing the program normally requires FPGA resynthesis because the instruction array is implemented as ROM; synthesis can take several minutes.
- An active-low reset connects directly to the development-board pushbutton, and the program counter initializes to zero both at programming time and when reset is asserted.
- The classical MIPS pipeline has fetch, decode, execute, memory, and write-back stages, but the course CPU omits memory and combines decode with execute.
- The resulting three-stage pipeline is fetch, execute, and write back; execute carries register reads, control decoding, immediate selection, and ALU work.
- A three-stage instruction has three cycles of latency, but after the pipeline fills, one instruction can complete every cycle.
- The design targets roughly 50 MHz, accepting a heavier execute stage because the course does not require the higher clock rates that would justify a deeper pipeline.
- A clocked instruction register creates a one-cycle delay: when PC_fetch addresses instruction 10, instruction_ex still contains the previously fetched instruction.
- Signal suffixes such as _fetch and _ex identify the pipeline stage associated with each signal, including combinational signals, making waveform debugging tractable.
- Reset clears both PC_fetch and instruction_ex so that the fetch and execute stages are flushed rather than allowing a stale instruction to continue executing.
- PC_fetch increments on each clock edge using the old right-hand-side value, avoiding the combinational loop that would result from an unclocked self-increment.
- The destination register field RD must be stored in a pipeline register before write back; otherwise, the design accidentally aligns a current instruction's destination with a previous instruction's result.
- During one cycle, instruction n can read RS1 and RS2 in execute while instruction n−1 writes its result to RD in write back.
- The required delay is a multi-bit flip-flop with no address, implemented using always_ff, rather than another RAM or register-file entry.
- The register-file instance is supplied, but students must wire instruction bit slices, ALU operands, write data, and delayed destination fields into the correct pipeline stages.
- The provided verification program exercises all 22 supported instructions at least once, although it does not comprehensively test multiple input combinations for every instruction.
- Expected results are included in program comments, giving students a baseline test for register writes, ALU operations, immediate handling, and I/O behavior.
- A second required program must perform a meaningful binary-to-decimal conversion rather than merely checking instruction outputs.
- Jason D. Bakos indicates that the binary-to-decimal implementation will receive additional explanation in lab or the following lecture.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Jason D. Bakos.