ECE344 Fall 2026 (Sec 1) Lec 11 - Page Faults
Watch on YouTube →
Overview
Jon Eyolfson explains how multi-level page tables, TLB caching, and page faults implement efficient virtual memory on RISC-V systems. The lecture derives page-table sizing formulas, compares 3-level SV39 translation costs, and connects demand paging to Linux system calls such as BRK and MMAP, including file-backed mappings, copy-on-write, and swapping.
Key takeaways
- SV39 uses three 9-bit page-table indices and a 12-bit unchanged offset; a translation therefore requires L2, L1, and L0 lookups before the final physical-memory access.
- Page-table memory depends strongly on virtual-address locality: mapping 512 contiguous pages can use three table pages, while a maximally scattered layout can require 1,025 table pages.
- A TLB caches complete VPN-to-PPN translations, and a 99% hit rate reduces the effective cost of a three-level page walk to about 113 ns when memory takes 100 ns.
- Touching one new 4 KiB page per access destroys TLB locality, whereas repeated accesses within a page amortize the initial translation miss over many operations.
- Changing SATP or PTE permissions requires TLB invalidation through RISC-V SFENCE.VMA unless address-space identifiers safely distinguish entries between processes.
- Linux demand paging and MMAP defer physical allocation and file loading until first access, enabling sparse access to large files such as 20 GB machine-learning models while using substantially less RAM.
Chapters
0:00
Exam Scope and Page-Fault Lecture Setup
- Jon Eyolfson identifies page faults and virtual memory as the final theoretically tested lecture topic, emphasizing conceptual rather than deeply computational questions.
- The planned 45-minute, 45-mark test allocates roughly one minute per mark and includes multiple-choice, process, and virtual-memory questions.
- The lecture continues from prior work on multi-level page tables and virtual-address translation.
6:30
SV39 Three-Level Page-Table Translation
- RISC-V SV39 divides a virtual page number into three 9-bit indices: L2, L1, and L0.
- The SATP register identifies each process's root L2 page table, while each page-table entry stores the physical page number of the next-level table or mapped page.
- A 12-bit page offset is copied unchanged into the physical address because page tables translate virtual pages, not individual byte offsets.
- For a virtual address such as ABCDEF, the lower 12 bits represent the offset and the remaining VPN bits select entries at each of the three levels.
12:00
Sharing Page Tables Across Related Virtual Pages
- Two virtual addresses with the same VPN but different offsets, such as ABCDEF and ABC123, resolve to the same physical page.
- Multiple VPNs can reuse upper-level page tables; changing only the L1 index may require one additional L0 table rather than duplicating the entire hierarchy.
- Mapping three suitably distributed VPNs can require six page-table pages: one shared L2 table, multiple L1 tables, and multiple L0 tables.
- Validity is page-granular, so a process cannot make one byte within a mapped page valid and another byte in the same page invalid.
17:30
Best- and Worst-Case SV39 Table Counts
- Mapping 512 pages, or 2 MiB with 4 KiB pages, requires at least three page tables in the best case: one L2, one L1, and one fully populated L0 table.
- The worst case uses one L2 table, 512 L1 tables, and 512 L0 tables when the virtual pages are spread across distinct upper-level indices.
- Contiguous virtual memory generally resembles the best case because code, heap, and stack regions tend to cluster in virtual-address space.
- Contiguous virtual addresses do not imply contiguous physical memory; physical frames can remain scattered.
23:30
Calculating Page-Table Levels from Address Parameters
- The number of entries per page table equals page size divided by page-table-entry size.
- Index width is log2 of the entries per table, and the number of levels is ceil(VPN bits divided by index bits).
- For a 32-bit virtual address, 4 KiB pages, and 4-byte PTEs, the offset is 12 bits, the VPN is 20 bits, each index is 10 bits, and two levels are sufficient.
- Increasing that example to a 34-bit virtual address produces a 22-bit VPN and requires three levels because ceil(22/10) equals 3.
30:00
TLB Caching Removes Repeated Page-Table Walks
- A three-level page-table walk adds three physical-memory reads before the requested data access, making one logical load require four memory accesses.
- The translation lookaside buffer, or TLB, caches complete VPN-to-PPN translations in MMU hardware rather than caching only one page-table level.
- A TLB hit directly supplies the physical page number; a miss walks the page table and inserts the completed translation into the TLB.
- Typical TLB capacities range from a few dozen to a few thousand entries, exploiting repeated access to loop code, stacks, and arrays.
35:10
Effective Access Time with TLB Hits and Misses
- Using a 10-nanosecond TLB lookup and 100-nanosecond memory access, a single-level page table costs 110 ns on a hit and 210 ns on a miss.
- With an 80% hit ratio, the effective access time is 0.8 × 110 ns plus 0.2 × 210 ns, or 130 ns.
- For a three-level page table, a miss can require TLB lookup, L2, L1, L0, and the final data access, totaling about 410 ns under the stated assumptions.
- At a 99% hit ratio, the three-level example averages roughly 113 ns, limiting virtual-memory overhead to about 13%.
39:30
Stride Benchmarks Reveal TLB Locality
- The benchmark program allocates a requested memory size and accesses an integer at a specified byte stride, exposing how spatial locality affects TLB performance.
- A 4 KiB allocation accessed every 4 bytes incurs one initial translation miss followed by repeated hits within the same page.
- Accessing a 16 MiB region every 128 bytes produces approximately one page transition per 32 accesses because 4096/128 equals 32.
- Accessing 512 MiB every 4096 bytes touches a new page on every access, driving the TLB hit ratio toward zero and making execution roughly ten times slower.
43:20
TLB Consistency During Context Switches and Permission Changes
- TLB entries are valid only for the address space associated with the current SATP root page-table register.
- Switching from one process to another without invalidating or tagging TLB entries could translate the same virtual address to another process's physical memory.
- RISC-V's SFENCE.VMA instruction flushes TLB state; xv6 uses it when changing SATP, while address-space identifiers can avoid full flushes by tagging entries with a process identity.
- Changing PTE permissions for copy-on-write requires a TLB flush, otherwise stale writable permissions may allow access without triggering the intended page fault.
47:40
BRK, SBRK, and Eager Versus Lazy Heap Growth
- The BRK system call changes a process's program break, while SBRK moves it by a requested byte count and returns the previous break.
- xv6 eagerly allocates and maps physical pages for heap growth, typically using read-write and user-accessible PTE permissions.
- Linux commonly uses demand paging: it records newly valid virtual addresses but postpones physical allocation until the process first accesses a page.
- Freeing memory with malloc may not reduce a process's mapped footprint because the allocator can retain pages for future small allocations.
50:40
MMAP Protection, Sharing, and File-Backed Mappings
- MMAP accepts a preferred virtual address, byte length, protection flags, mapping flags, a file descriptor, and a file offset; the kernel rounds the length up to whole pages.
- PROT_READ, PROT_WRITE, and PROT_EXEC correspond directly to page-table permission bits.
- MAP_PRIVATE creates process-private mappings, while MAP_SHARED allows forked processes to reference the same physical pages for fast interprocess communication.
- Without a file descriptor, the mapping is zero-initialized; with a file descriptor and offset, pages can be populated from file contents.
54:00
File I/O Simplified by Lazy MMAP Faults
- A cat-like program can open a file, obtain its size with stat, MMAP the file, and write the mapped range to standard output without issuing read system calls.
- The first access to each unmapped page triggers a page fault, causing the kernel to allocate a physical page and load the corresponding file contents.
- After population, subsequent accesses use normal translations and typically hit in the TLB.
- The same technique can let a 20 GB machine-learning model use roughly 7 GB of RAM when inference accesses only a sparse subset of the file.
58:00
Page-Fault Handling: Demand Paging, Copy-on-Write, and Swapping
- A page fault is a hardware-generated trap handled by the kernel, which can terminate the process for an invalid address or supply a missing physical frame for a valid lazy mapping.
- Demand paging maps a newly allocated frame and retries the faulting instruction without exposing the fault to the user process.
- Copy-on-write handles a write to a shared read-only page by copying the frame, remapping it writable, and retrying the instruction; this is central to xv6 Lab 3.
- Swapping can move pages to disk so the operating system presents more virtual memory than available physical RAM.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Jon Eyolfson.