CSCE 611 Fall 2026 Lecture 1: Introduction
Watch on YouTube →
Overview
Jason D. Bakos introduces CSCE 611 as a hands-on computer engineering course in which students learn SystemVerilog and industrial FPGA tools, then build and test a pipelined RISC-V CPU from scratch. He explains the course logistics and project rules, the design flow from behavioral descriptions through synthesis and place-and-route, and why specialized processors now occupy much of modern chips as general-purpose CPU performance has plateaued.
Key takeaways
- CSCE 611's central project is a from-scratch, single-issue, in-order RISC-V CPU implemented with SystemVerilog and tested on an FPGA; the course targets a functional, Turing-complete design rather than commercial performance.
- FPGA implementation automates the place-and-route work once done by assigning 7400-series gates to breadboard chips and manually wiring their pins, while imposing resource limits such as six-input logic functions.
- Modern processors devote substantial silicon to application-specific accelerators: Bakos cites an Apple die with about 12% general-purpose CPU area and Intel's CPU share declining from 72% in 2010 to 20% in 2015.
- Hand-writing assembly helps students debug the CPU's internal state, but register reuse can overwrite values that remain live; advanced processors use register renaming to avoid resulting false dependencies.
- RISC-V's modular ISA lets implementers choose extensions; CSCE 611 focuses on RV32I and a subset of M rather than implementing the full standard.
- A high-level if/else often maps to an assembly branch on the negated condition, while array indexing requires explicit address calculation because basic RISC-V loads and stores use fixed-address offsets.
Chapters
- Jason D. Bakos describes Monday lectures and separate Wednesday or Friday lab sections in the Coliseum and room 3D22.
- Blackboard carries announcements and course materials; recorded lectures are posted on YouTube, and student teams arrange their own code sharing through tools such as GitHub or Dropbox.
- Lab computers have a 2 GB storage quota that can be consumed by compile files; departmental IT contact Ryan Austin can help with machine and software issues.
- The assigned text is Digital Design and Computer Architecture, RISC-V edition; Bakos notes that it is technically labeled the first edition.
- Projects account for 40% of the course grade and are completed in teams of two by default; quizzes and exams account for the other 60% and are individual work.
- Six multiple-choice quizzes are assigned Mondays and due Fridays; the midterm and final are in-class paper-and-pencil exams with short answers, code debugging, and pipeline tracing.
- Graduate-credit students complete an extra lab on very long instruction word (VLIW) execution, modifying the CPU to issue two instructions per cycle; the practical performance goal is about 30%.
- Projects are submitted once per team through Blackboard by 11:59 p.m.; late work loses 10% per school day, capped at 30%, and submissions must compile and run.
- Bakos frames CSCE 611 as an extension of CSCE 211 and 212: students move from designs with dozens of gates to systems with thousands.
- The main project is a from-scratch, single-issue, in-order pipelined CPU targeted at roughly 50 MHz.
- The CPU is intended to be Turing complete, and students can compile C code with Clang or GCC to target it, although course debugging uses hand-written assembly.
- The syllabus schedule allocates multiple lab sessions to projects, but teams may need additional work time because later labs build on earlier ones.
- The digital-design flow begins with a behavioral description, such as a Boolean equation, which can be synthesized into a structural schematic of logic gates.
- Mapping converts an equivalent circuit into a form that fits implementation constraints; for example, the target FPGA supports logic functions with at most six inputs.
- A sum-of-products circuit with more than six terms cannot use one oversized OR gate on this FPGA and must be restructured into smaller logic.
- Bakos distinguishes the logical circuit from its physical implementation, where gates and interconnections must be assigned to available hardware.
- Earlier designs used 7400-series NAND chips on breadboards, with students manually assigning schematic gates to chip pins and wiring them together.
- That manual process is place-and-route: assigning logical gates to physical resources and connecting their inputs and outputs.
- An FPGA automates gate allocation and routing through programmable logic and configurable interconnects, making much larger designs practical than breadboards.
- SystemVerilog serves as the course hardware description language; Bakos contrasts its industry use with VHDL's greater prevalence in some academic settings.
- An FPGA powers up as an unconfigured device and receives a bitstream from CAD tools; the course board reloads its FPGA from onboard flash at startup.
- FPGAs trade performance and density for rapid prototyping: Bakos estimates custom ASICs can run about 10 times faster, with substantially greater logic density.
- Structural HDL describes modules and their wiring in text, while behavioral HDL describes operations such as Boolean logic and arithmetic; both styles can be combined.
- Students translate provided CPU schematics into HDL for later labs, and use HDL text rather than schematic-generation software to make compiler errors easier to trace.
- Bakos argues that general-purpose CPU performance has largely stopped improving since about 2011, constrained by issues such as the power wall and memory wall.
- He cites dedicated processors for self-driving cars, 4K video encoding, Face ID, computational photography, and on-device language translation as examples of domain-specific acceleration.
- In the Apple chip example he discusses, the general-purpose CPU occupies about 12% of the die, with the rest devoted to specialized components; he also cites Intel CPU-area shares falling from 72% in 2010 to 20% in 2015.
- The course's discontinued FPGA board remains useful for teaching because it includes many peripherals, including memory, switches, displays, a VGA chip, and an SD-card reader.
- The first roughly 10 weeks emphasize SystemVerilog skills and CPU construction; the final five weeks shift toward traditional architecture concepts while students finish their CPUs.
- Quizzes are concentrated later in the semester to align with that concept-focused final portion of the course.
- Students write small assembly programs by hand to make CPU state changes easier to inspect while debugging, even though a compiler can also target the CPU.
- Bakos defines architecture as the visible instruction set and microarchitecture as the internal circuitry that executes it; RISC-V is used as an open ISA that students can implement legally.
- RISC-V resembles MIPS but standardizes the destination-register field, removes MIPS's shift-amount field, uses sign-extended immediates, and simplifies jump-and-link behavior.
- RISC means reduced instruction set computer: simpler, faster instructions can increase the number of instructions needed for a program.
- RISC-V uses optional instruction extensions rather than requiring every processor to implement every instruction; the course builds RV32I plus a subset of the M extension.
- The course CPU therefore implements only a small portion of RISC-V, rather than claiming full ISA compliance.
- RISC-V arithmetic operates on registers, so assembly programs must load variables from memory, compute with register values, and store results back.
- With only 32 architectural registers, programmers must reuse registers after their values are no longer needed; reusing a live register can silently corrupt a later computation.
- Bakos demonstrates a faulty expression translation in which a temporary register overwrites a copy of variable B that is needed again at the end.
- In advanced processors, register renaming maps architectural registers to physical registers to avoid false anti-dependencies caused by register reuse.
- Assembly branches transfer control when a condition is true, so implementing a high-level if/else often requires branching on the negated condition to reach the else block.
- For an if condition such as A < B, Bakos's example branches on A >= B to skip the true block and reach the else path.
- Array access generally needs several instructions because a load or store uses a fixed address and the program must calculate a variable element address.
- Index registers available on some processors, such as certain DSPs, are absent from basic RISC-V because combining index calculation with a memory operation conflicts with the RISC design approach.
- RISC-V registers can be named x0 through x31 or by aliases such as s0 and t0; assemblers convert either notation into register numbers.
- Bakos recommends x-register names for this course so students can directly track physical register-file entries in the hardware simulator.
- Debuggers commonly display symbolic aliases, which can make it harder to connect assembly code to the register being inspected in the simulator.
- The lecture ends with Bakos postponing additional assembly examples until the next Monday session.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Jason D. Bakos.