Virtual Memory
September 13, 202628 min readadvanced
The processor and operating system together present each running program with the illusion of owning the entire memory. A single process can issue a load from address 0x0000000000401234 and expect to find the…
The processor and operating system together present each running program with the illusion of owning the entire memory. A single process can issue a load from address 0x0000_0000_0040_1234 and expect to find the byte it wrote there earlier, even when a dozen other processes on the same machine each have data of their own at 0x0000_0000_0040_1234. The illusion is constructed by virtual memory, the joint hardware and software mechanism that translates the addresses programs use into the addresses that reach the DRAM.
Virtual memory predates the cache by several years. The Atlas computer at Manchester in 1962 introduced the page-based scheme that every modern system still uses, with hardware mapping pages of the virtual address space onto pages of a smaller physical store, backed by a slower drum that held pages not currently resident. The drum is now an SSD and the physical store is now gigabytes rather than kilobytes, but the structural ideas survive almost unchanged.
This chapter develops the mechanism from first principles. It starts with a concrete numerical walk through a translation, defines pages and page tables, then climbs the multi-level page tables of RISC-V Sv39, ARMv8 VMSAv8, and x86-64 four-level and five-level paging. It closes with huge pages and the engineering tradeoffs that determine when they pay off. The companion mechanism, the translation lookaside buffer that caches recent translations, is covered separately in Chapter 42.
01.Why Virtual Memory
A computer without virtual memory exposes its physical address space directly. Each program sees the same memory map, and two programs that both want to use the byte at address 0x4000 have to either coordinate their layouts at link time or accept that one of them will overwrite the other’s data. Mainframe systems in the 1950s and minicomputers through the 1970s lived with this constraint by linking each program at a unique base address. This works for a handful of canonical programs and fails the moment a user wants to run two copies of the same program at once.
Virtual memory solves three problems at once.
Isolation. Each process gets its own virtual address space. A wild pointer in process A cannot touch process B’s data because the hardware refuses to translate process A’s virtual addresses into process B’s physical pages. The operating system relies on this property to enforce inter-process boundaries. It also relies on it to enforce the user-kernel boundary covered in Chapter 20: kernel pages are marked supervisor-only, and a user-mode load that targets such a page raises a page fault rather than reading the kernel’s secrets.
Relocation. The linker produces executables that assume a particular virtual layout: code at one fixed virtual address, the heap growing from another, the stack growing down from a third. At load time the operating system maps those virtual addresses onto whatever physical pages happen to be free, and two processes started from the same binary can each get the same virtual layout backed by disjoint physical pages. The relocation problem is solved without the linker ever needing to know about the runtime physical layout.
Demand paging. The combined virtual address space of all running processes typically exceeds the physical DRAM. A modern laptop with 16 GiB of DRAM routinely runs processes whose virtual sizes sum to 100 GiB or more, because most virtual pages are unused at any moment. When a process first touches a virtual page, the operating system allocates a physical page on demand. When the physical pages run out, the operating system writes the least-recently-used pages back to swap storage and frees those physical frames. The mechanism that catches a reference to an unmapped page is the page fault.
A concrete walk-through helps fix the picture. Suppose a process issues a load from virtual address on a machine with 4 KiB pages. The bottom 12 bits () are the byte offset within the page. The upper 52 bits () are the virtual page number. The translation hardware looks up this virtual page number in the page table and finds, say, that it maps to physical page . The physical address is then the physical page number concatenated with the byte offset: . That physical address goes to the DRAM controller, the byte is read, and the load completes.
02.Pages and Page Tables
The natural unit of mapping is not the byte but the page. A page is a block of virtual addresses, all contiguous, that is mapped together onto a contiguous block of physical addresses called a page frame. All common architectures use a 4 KiB base page size. This number is a balance. Smaller pages would waste less DRAM on partially-filled allocations but would multiply the number of page table entries. Larger pages would shrink the page table but would waste DRAM on the last partial page of every allocation. The 4 KiB choice traces to the IBM System/370 and, along the line that leads directly to current machines, to the Intel 80386, and it has proved durable enough that every modern architecture inherits it. The VAX went the other way with 512-byte pages and paid for the choice in page table size.
A 4 KiB page has bytes. On a 64-bit machine with a 48-bit virtual address space (the actual width used by RISC-V Sv48, ARMv8 VMSAv8 with 48-bit addressing, and x86-64 4-level paging), the virtual address splits as with virtual page numbers per address space. A flat page table that records one entry per virtual page would need entries, each at least 8 bytes, for a total of 512 GiB of memory per process. No machine can afford this, so flat page tables are not used. Real systems use multi-level page tables that record only the populated branches of a sparse virtual address space, which the next section develops.
A single page table entry carries the physical page number and a small set of flag bits. The exact format varies by architecture, but every PTE contains roughly the same fields. The RISC-V Sv39 PTE [1] is representative:
Table 1. RISC-V Sv39 page table entry format. The reserved bits are zero in current implementations and reserved for future extensions. Data drawn from the RISC-V Privileged ISA Manual.
| Bits | Field | Meaning |
|---|---|---|
| 63:54 | Reserved | Must be zero |
| 53:28 | PPN[2] | Top 26 bits of physical page number |
| 27:19 | PPN[1] | Middle 9 bits of physical page number |
| 18:10 | PPN[0] | Bottom 9 bits of physical page number |
| 9:8 | RSW | Reserved for supervisor software |
| 7 | D | Dirty: page has been written |
| 6 | A | Accessed: page has been read or written |
| 5 | G | Global: shared by all address spaces |
| 4 | U | User: accessible in user mode |
| 3 | X | Executable |
| 2 | W | Writable |
| 1 | R | Readable |
| 0 | V | Valid |
The valid bit V indicates whether the entry refers to anything at all. A PTE with V cleared causes a page fault on any access. The R, W, and X bits encode the permission set, with the convention that all three clear marks the entry as a pointer to the next level of the page table rather than a leaf entry. The U bit decides whether user-mode accesses can use the entry. The A and D bits are updated by the hardware (or by software, depending on the implementation) to let the operating system identify which pages have been referenced or modified recently. This information drives page replacement when DRAM runs short.
The ARMv8 and x86-64 PTE formats follow the same template with minor differences [2][3]. ARMv8 packs the access permissions into a 2-bit AP field and a 1-bit XN (execute-never) bit, with additional fields for memory attributes (cacheable, device, non-cacheable) and shareability. Intel and AMD use separate present, read-write, user-supervisor, and execute-disable bits, plus memory-type bits that select among write-back, write-through, write-combining, and uncached caching policies. All three encodings fit in 64 bits because the physical address space is smaller than the virtual address space and the flag bits fit in the gap.
03.Multi-Level Page Tables
The flat 512 GiB page table of the previous section is unaffordable, but the underlying virtual address space is also sparse. A typical process maps a few megabytes of code, a few megabytes of stack, and a hundred megabytes of heap, leaving the vast majority of the 256 TiB of virtual address space (on a 48-bit machine) unmapped. A representation that records only the populated regions can be orders of magnitude smaller.
The standard solution is a multi-level page table, a radix tree keyed by the virtual page number. The virtual page number is split into several fields, each indexing one level of the tree. Each interior node is itself a 4 KiB page containing 512 8-byte entries. Each entry either points to the next level of the tree or, if no virtual page in its subtree is mapped, holds a zero (invalid) entry. Unmapped subtrees consume no memory at all.
The RISC-V Sv39 scheme [1] is the cleanest example to start with. Sv39 supports a 39-bit virtual address. The bottom 12 bits are the page offset, leaving 27 bits of VPN, which split into three 9-bit indices VPN[2], VPN[1], VPN[0]. Each index selects one of 512 entries in a page-sized table. Translation proceeds by the following walk.
-
Read the supervisor address translation and protection register
satp, which holds the physical address of the root page table for the current process. -
Index that root with VPN[2] to obtain a level-2 PTE. If the PTE is invalid, raise a page fault. If it is a leaf (R, W, or X set), the translation is complete and covers a 1 GiB region. Otherwise it points to the level-1 page table.
-
Index the level-1 table with VPN[1] to obtain a level-1 PTE. Again invalid raises a fault, leaf completes translation (now covering a 2 MiB region), and non-leaf points to the level-0 table.
-
Index the level-0 table with VPN[0] to obtain a leaf PTE. This entry holds the physical page number for the 4 KiB page that contains the requested address. Concatenate the physical page number with the 12-bit offset to produce the physical address.
A concrete walk fixes the procedure. Suppose the virtual address is and the page tables are arranged so that the translation succeeds with physical page . Bits 38:30 of V are , giving VPN[2] = 255. Bits 29:21 are , giving VPN[1] = 511. Bits 20:12 are , giving VPN[0] = 510. The walker reads three PTEs (one at each level) and produces the physical address . Three memory references for one translation, which is why the TLB of Chapter 42 is so important.
For Sv39 with depth 3 and a DRAM access of 100 ns, a TLB miss costs roughly 300 ns plus the DRAM access for the final data, or about 400 ns in total. At a 3 GHz clock, that is 1200 cycles per miss. A workload that misses the TLB on every reference would be unusable. Real workloads hit the TLB more than 99 percent of the time, but a 1 percent miss rate against a 1200-cycle miss still leaves cycles of translation cost on an average access. Pushing that average below a single cycle takes a hit rate above roughly 99.9 percent, which is why the interesting part of the design space sits at the high end of the hit-rate range rather than at 99 percent.
Sv32, Sv48, Sv57
RISC-V defines four virtual address widths in the privileged specification, all using the same multi-level radix-tree structure.
Sv32 is the 32-bit virtual address scheme used by RV32 systems with paging. The 32-bit virtual address splits into a 12-bit offset and a 20-bit VPN, which further splits into two 10-bit indices. The page table has two levels, each holding 1024 4-byte entries (a smaller PTE because the physical address space on RV32 machines is at most 34 bits). Sv32 supports 4 KiB pages and 4 MiB megapages.
Sv39 is the standard scheme on entry-level RV64 systems. 512 GiB of virtual address space is plenty for embedded Linux, microservers, and most other RV64 workloads.
Sv48 extends the walk to four levels and increases the virtual address space to 256 TiB. This is the configuration most desktop and server RISC-V cores ship with, matching x86-64 and ARMv8 4-level paging in reach.
Sv57 adds a fifth level and pushes the virtual address space to 128 PiB. Sv57 exists to support workloads (large in-memory databases, scientific computing) that map terabyte-sized regions and would benefit from larger huge-page hierarchies than Sv48 allows.
ARMv8 VMSAv8-64 [2] uses a near-identical scheme with configurable granule sizes (4 KiB, 16 KiB, 64 KiB) and configurable top-level index width (T0SZ and T1SZ). The 4 KiB granule configuration produces a 4-level page table for 48-bit addressing, structurally equivalent to Sv48. The 16 KiB and 64 KiB granules keep 48-bit addressing but redistribute the bits. The 16 KiB granule places a 1-bit top-level index above three 11-bit levels, and the 64 KiB granule places a 6-bit top-level index above two 13-bit levels, which is what drops the 64 KiB walk to three levels.
Intel and AMD x86-64 [3][4] use a 4-level page table for 48-bit virtual addresses, with the levels named (top to bottom) PML4, PDPT, PD, and PT. Each level holds 512 8-byte entries, identical to Sv48. The 5-level extension (LA57), shipped in Intel Ice Lake server parts and AMD Zen 4, adds a top-level PML5 table and pushes the virtual address to 57 bits and the address space to 128 PiB.
Table 2. Page table parameters across the three major architecture families. Each row uses the base page or granule size named in its label. Rows whose levels do not all share one index width list the indices from the top level down.
| Scheme | VA bits | Levels | Index per level | PTEs per level |
|---|---|---|---|---|
| RISC-V Sv32 | 32 | 2 | 10 | 1024 |
| RISC-V Sv39 | 39 | 3 | 9 | 512 |
| RISC-V Sv48 | 48 | 4 | 9 | 512 |
| RISC-V Sv57 | 57 | 5 | 9 | 512 |
| ARMv8 VMSAv8-64 (4 KiB) | 48 | 4 | 9 | 512 |
| ARMv8 VMSAv8-64 (16 KiB) | 48 | 4 | 1, 11, 11, 11 | 2, 2048 |
| ARMv8 VMSAv8-64 (64 KiB) | 48 | 3 | 6, 13, 13 | 64, 8192 |
| x86-64 (PML4) | 48 | 4 | 9 | 512 |
| x86-64 (LA57/PML5) | 57 | 5 | 9 | 512 |
Figure 1 follows one translation through the tree.
04.The Page Table Walker
The hardware unit that performs the multi-level lookup is the page table walker. In RISC-V it is a dedicated state machine that issues physical memory references through the data cache or directly to memory. In ARMv8 and x86-64 the walker is similarly a state machine. All three architectures allow the walker to cache intermediate-level PTEs (the non-leaf nodes) in a dedicated walk cache that sits between the TLB and the memory hierarchy. A walk cache that holds the level-3 and level-2 PTEs of recently-translated regions can short-circuit the top two levels of a 4-level walk on every miss.
The walker also writes the access and dirty bits when the implementation supports hardware A/D updates. RISC-V leaves this optional in the privileged spec, with the convention that an implementation that does not update A or D in hardware must raise a page fault when the A bit is clear on any access or the D bit is clear on a store, letting software update the bit and re-execute the instruction. ARMv8.1-A added the FEAT_HAFDBS extension to perform hardware A/D updates. x86-64 has always updated A and D in hardware.
Idealized page-table walk in pseudocode for Sv39.
uint64_t walk_sv39(uint64_t satp, uint64_t va) {
uint64_t root_ppn = satp & 0xFFFFFFFFFFFULL;
uint64_t pt_addr = root_ppn << 12;
for (int level = 2; level >= 0; level--) {
int idx = (va >> (12 + 9 * level)) & 0x1FF;
uint64_t pte = *((uint64_t *)(pt_addr + 8 * idx));
if (!(pte & PTE_V)) raise_page_fault();
if ((pte & (PTE_R | PTE_W | PTE_X)) != 0) {
uint64_t ppn = (pte >> 10) & 0xFFFFFFFFFFFULL;
int off_bits = 12 + 9 * level;
return (ppn << 12) | (va & ((1ULL << off_bits) - 1));
}
pt_addr = ((pte >> 10) & 0xFFFFFFFFFFFULL) << 12;
}
raise_page_fault();
}The pseudocode shows three important features of the walk. First, each level requires one DRAM access to fetch a PTE, so the worst-case walk takes three (Sv39) or four (Sv48) or five (Sv57) DRAM accesses. Second, the walker checks the V bit at every level and faults immediately on an invalid entry, which makes unmapped virtual addresses fast to reject. Third, a non-leaf at any level (R, W, X all clear) means the walker continues to the next level, while a leaf at any level above 0 means the translation covers a larger region of virtual address space. That last case is the huge-page mechanism covered in the next section.
05.Huge Pages
A 4 KiB page is small. A modern process with a 4 GiB heap covers virtual pages, which means the page table for the heap alone has roughly one million leaf PTEs spread across two thousand 4 KiB leaf pages. Walking the page table on a TLB miss costs three to four DRAM accesses. TLB capacity is limited to a few thousand entries, so a workload that touches more than a few megabytes of data tends to thrash the TLB and pay the walk cost on a large fraction of accesses.
Huge pages address the problem by mapping larger contiguous regions with a single TLB entry. A huge page is a leaf at a level above 0 in the multi-level page table. At level 1 of Sv39, a leaf entry covers of virtual address space. At level 2, a leaf entry covers . The same applies to ARMv8 (the ARM term is block descriptor) and to x86-64, where the PSE bit in a level-1 entry produces a 2 MiB mapping and a PSE bit in a level-2 entry produces a 1 GiB mapping.
The alignment requirements follow from the structure. A 2 MiB huge page must be aligned to a 2 MiB boundary in both virtual and physical address space, because a leaf PTE at level 1 is required to have its bottom nine PPN bits clear, and the level-0 index of the virtual address becomes part of the 21-bit offset inside the huge page rather than selecting a table entry. Similarly a 1 GiB huge page requires 1 GiB alignment. Allocating large contiguous physical regions becomes harder as the system runs longer and DRAM fragments, which is why Linux distinguishes static huge pages (reserved at boot) from transparent huge pages (built opportunistically when contiguous DRAM is available).
A concrete win comes from the TLB. A standard 4 KiB page covers addresses with one TLB entry. A 2 MiB huge page covers addresses with one TLB entry, a increase in TLB reach. A workload that touches a 4 GiB heap reaches the heap with 2048 huge-page TLB entries instead of one million standard PTEs. Since modern TLBs hold a few thousand entries (covered in Chapter 42), the huge-page workload fits in the TLB entirely while the standard-page workload thrashes.
Table 3. Page sizes supported by major architectures. The exact set of supported sizes depends on the implementation and the granule selected at boot. RISC-V permits a leaf at every level, so adding a level adds a page size, while x86-64 defines a page-size bit only at the PD and PDPT levels and therefore stops at 1� GiB.
| Architecture | Standard | Larger pages |
|---|---|---|
| RISC-V Sv39 | 4 KiB | 2 MiB, 1 GiB |
| RISC-V Sv48 | 4 KiB | 2 MiB, 1 GiB, 512 GiB |
| RISC-V Sv57 | 4 KiB | 2 MiB, 1 GiB, 512 GiB, 256 TiB |
| ARMv8 (4 KiB granule) | 4 KiB | 2 MiB, 1 GiB |
| ARMv8 (16 KiB granule) | 16 KiB | 32 MiB |
| ARMv8 (64 KiB granule) | 64 KiB | 512 MiB |
| x86-64 (4-level) | 4 KiB | 2 MiB, 1 GiB |
| x86-64 (5-level) | 4 KiB | 2 MiB, 1 GiB |
The downsides of huge pages are real. Internal fragmentation grows: a process that maps a 2 MiB huge page but uses only the first 8 KiB wastes the remaining 2040 KiB until the page is freed. Demand paging becomes coarser: a page fault on a huge page brings 2 MiB of data in, even if the program will touch only one byte. Copy-on-write fork semantics break for huge pages, because cloning a 2 MiB page on a single store would defeat the point. Linux handles this by splitting huge pages on the first copy-on-write trigger, which adds TLB-shootdown overhead. These costs limit huge pages to workloads with large, dense, long-lived data structures: scientific simulation arrays, large databases, in-memory caches, the JVM heap.
06.Page Faults and the Operating System
A page fault is the exception raised when the page table walker encounters an invalid entry or a permission violation. The faulting access is suspended, the privilege level escalates to supervisor (or hypervisor), and control transfers to the operating system’s page fault handler.
The handler examines the saved program counter and the faulting address (held in a dedicated architectural register: stval on RISC-V, FAR_ELn on ARMv8, CR2 on x86) and decides what to do. A read access to a page that is mapped read-only and stored on swap is a swap-in fault: the handler reads the page back from swap, installs a valid PTE, and restarts the instruction. A write access to a read-only PTE on a copy-on-write page allocates a fresh physical page, copies the old data, installs a writable PTE, and restarts. A reference to a page that has never been touched (the demand-zero case for fresh anonymous memory) allocates a physical page, zero-fills it, and installs a PTE. A reference to a virtual address that has never been mapped at all (a wild pointer) is a hard fault: the handler sends SIGSEGV to the offending process.
The cost of a page fault is dominated by the operating system’s work, not by the hardware mechanism. A swap-in fault that pulls a page from an SSD takes a few hundred microseconds, three to four orders of magnitude more than a TLB miss. Even a fast demand-zero fault that allocates and zeros a fresh page takes a few microseconds, and almost all of that goes to the allocation and page table bookkeeping in the kernel. The zeroing itself moves only 4 KiB, which at a few tens of gigabytes per second of write bandwidth finishes in a few hundred cycles.
07.Address Space Identifiers and Context Switches
When the operating system switches from one process to another, the virtual-to-physical mapping changes completely. Process A’s virtual page 0x401 maps to one physical page, process B’s virtual page 0x401 maps to another. The page table root in satp (RISC-V), TTBR0_EL1 (ARMv8), or CR3 (x86) points at a different page table after the switch.
A naive implementation would also flush the entire TLB on every context switch, because every cached translation now potentially points at the wrong physical page. Flushing destroys all the warm translations, and the new process pays a wave of TLB misses on its first thousand accesses. All three architectures support address space identifiers to avoid this.
RISC-V satp holds a 16-bit ASID alongside the page table root. ARMv8 TTBR_EL1 holds a similar field. Intel and AMD use process-context identifiers (PCIDs) on x86-64, a 12-bit field in the low bits of CR3. The later INVPCID instruction does not widen that field. It invalidates TLB entries belonging to a chosen PCID without switching to it. Each TLB entry is tagged with the ASID or PCID that produced it. A lookup matches an entry only when both the virtual address and the ASID match. After a context switch, the old process’s entries remain valid in the TLB but invisible to the new process, and the new process’s entries (if it has been scheduled before) are still visible.
The interaction with TLB management and shootdown is covered in Chapter 42. For this chapter, the takeaway is that ASIDs turn context switches from a TLB-thrashing event into a near-free register write.
08.Worked Examples
09.Exercises
References
- [1]Waterman, Andrew and Asanovi\'c (2024). “The RISC-V.”
- [2](2024). “ARM.”
- [3](2024). “Intel.”
- [4](2024). “AMD64.”