Memory, Caches & Operating Systems: What Every Systems Rust Developer Must Know First

22 min read • Systems Programming Foundations for Rust

Before you write a single line of systems Rust, you need a mental model of what your computer is actually doing when it runs your program. Rust's ownership system, borrow checker, and performance guarantees only make sense once you understand the hardware and operating system underneath them.

In this guide we will walk through stack vs heap memory, virtual memory, CPU caches, and the core operating system concepts that govern every program you will ever write. Every idea is explained with a real-world analogy first so you can build intuition before touching any code.

No prior systems knowledge required. Let's build the foundation.

Part 1: Stack vs Heap Memory

The Restaurant Analogy

Imagine a busy restaurant. The stack is like the waiter's notepad. It's small, lives in the waiter's pocket, and is used only for the current table. When the waiter finishes with a table the notes are erased immediately and the space is reused. Fast, predictable, but limited.

The heap is the restaurant's large storage room out back. It can hold much more, but you have to walk there yourself, find a free shelf, write down which shelf you used, and later remember to return the space when you're done. Slower, flexible, but requires discipline.

The Stack

The stack is a region of memory that grows and shrinks in a strict Last-In, First-Out (LIFO) order, just like a stack of plates. Every time you call a function, the CPU pushes a stack frame onto it. That frame holds:

  • Local variables whose size is known at compile time
  • Function arguments passed to the current call
  • The return address so the CPU knows where to jump back after the function ends

When the function returns, its entire frame is popped off instantly—one register decrement. No searching, no bookkeeping. This is why stack allocation is essentially free.

Rust: stack allocation (lives until function returns)
fn greet(name: &str) {
    let greeting = "Hello, ";   // stack: fixed-size &str slice reference
    let count: u32 = 42;        // stack: 4 bytes, known at compile time
    println!("{}{}", greeting, name);
}   // <-- stack frame is destroyed here, instantly

Stack limit: On most systems the default stack is only 8 MB. Call too many nested functions (or allocate a huge array on the stack) and you get a stack overflow. Rust protects you by keeping large data on the heap.

The Heap

The heap is a large pool of memory managed by the OS and your language runtime. Unlike the stack, there is no automatic cleanup. You (or Rust's ownership system) must explicitly release heap memory when it is no longer needed.

In C you call malloc / free manually. In languages like Java a garbage collector does it for you—but pauses your program unpredictably. Rust takes a third path: the ownership system proves at compile time exactly when memory can be freed, with zero runtime cost.

Rust: heap allocation with Box<T>
fn main() {
    // Box<T> allocates 'value' on the heap
    let boxed = Box::new(5u32);
    println!("Heap value: {}", boxed);
}   // <-- boxed goes out of scope here
    //     Rust automatically calls 'drop', freeing heap memory
    //     No GC pause, no manual free() needed
Vec<T> grows on the heap at runtime
fn main() {
    let mut numbers: Vec<u32> = Vec::new(); // heap allocation
    numbers.push(1);  // may reallocate if capacity is exceeded
    numbers.push(2);
    numbers.push(3);

    // Stack holds: pointer (8 bytes) + length (8 bytes) + capacity (8 bytes)
    // Heap holds: [1, 2, 3, ...] — the actual data
    println!("{:?}", numbers);
}   // Vec dropped here, heap memory freed

Stack vs Heap: Quick Comparison

PropertyStackHeap
SpeedExtremely fastSlower (allocator overhead)
SizeSmall (∼8 MB default)Limited only by RAM
LifetimeTied to function scopeControlled by owner
Size known?Must be compile-timeCan be runtime-dynamic
Rust examplelet x: u32 = 5;Box::new(5), Vec::new()

Part 2: Virtual Memory

The Hotel Room Analogy

Imagine a hotel where every guest is told their room is number 101. Each guest thinks they have private access to room 101, but the hotel manager (the OS) secretly maps each guest's “room 101” to a different physical room. Guests never collide because the manager handles the translation invisibly.

That's virtual memory. Every process believes it has its own continuous address space starting from address 0. The OS and CPU hardware translate these virtual addresses to actual physical addresses in RAM—transparently.

Why Virtual Memory Exists

  • Isolation: Process A cannot accidentally (or maliciously) read or corrupt Process B's memory. Each has its own private view.
  • More memory than RAM: The OS can temporarily move unused pages to disk (swap space). Your program can use more memory than physically exists.
  • Simplified addressing: Every program is compiled assuming it owns a huge contiguous address space. No program needs to know where others are loaded.
  • Memory-mapped files: Files on disk can be accessed like memory arrays, letting the OS handle I/O lazily.

Pages and Page Tables

Virtual memory is divided into fixed-size chunks called pages (typically 4 KB on x86-64). The OS keeps a page table for each process: a lookup table mapping each virtual page to a physical page (or to a “not present” entry if the page is on disk).

The CPU has a hardware unit called the MMU (Memory Management Unit) that performs this translation on every memory access. To avoid looking up the page table on every access, the CPU caches recent translations in the TLB (Translation Lookaside Buffer).

Virtual address space layout of a 64-bit Linux process
High addresses  ┌─────────────────────────┐
                │   Kernel space          │  (OS code & data)
                ├─────────────────────────┤
                │   Stack (grows down)    │  local vars, return addrs
                │         ↓               │
                │   ...                   │
                │         ↑               │
                │   Heap  (grows up)      │  Box, Vec, String, etc.
                ├─────────────────────────┤
                │   BSS / Data segments   │  global & static variables
                ├─────────────────────────┤
Low addresses   │   Text segment          │  compiled machine code
                └─────────────────────────┘

Page Faults (and why they matter)

When a program accesses a virtual address that has no physical page mapped yet, the CPU triggers a page fault. The OS handler runs, allocates a physical page (or loads it from swap), updates the page table, and resumes the program. From the program's perspective nothing happened—but it just paid a huge latency cost (microseconds vs nanoseconds).

In Rust, heap allocations can trigger page faults if the OS has not yet backed the virtual pages with physical RAM. Real-time and embedded systems avoid this bypre-faulting memory or locking pages with mlock.

Part 3: CPU Caches

The Desk & Library Analogy

Your CPU is an extremely fast researcher who can process one fact in under a nanosecond. RAM is like the library across the building—it holds everything, but fetching a book takes 100+ nanoseconds (about 300 times slower). If the CPU had to go to the library for every single piece of data, it would spend 99% of its time waiting.

CPU caches are the researcher's desk drawers: L1 is the top drawer (tiny, instant), L2 is the filing cabinet (bigger, slightly slower), L3 is the nearby shelf (much bigger, noticeably slower). If the data is in a drawer it's a cache hit. If the CPU has to walk to the library it 's a cache miss.

The Cache Hierarchy

L1 Cache

  • Size: 32–64 KB per core
  • Latency: ~4 cycles (1–2 ns)
  • Private to each CPU core
  • Split: instruction + data

L2 Cache

  • Size: 256 KB–1 MB per core
  • Latency: ~12 cycles (4–5 ns)
  • Usually private per core
  • Unified instruction + data

L3 Cache (LLC)

  • Size: 8–64 MB shared
  • Latency: ~40 cycles (12–15 ns)
  • Shared across all cores
  • Last stop before RAM
Memory latency cheat-sheet (approximate on modern x86-64)
L1 cache hit    ~  1-2 ns   ← the fast drawer
L2 cache hit    ~  4-5 ns
L3 cache hit    ~ 12-15 ns
Main RAM        ~ 60-100 ns  ← the library across the building
NVMe SSD        ~ 100 µs    (1,000× slower than RAM)
Network (LAN)   ~ 500 µs
HDD             ~ 5-10 ms   (millions of times slower than L1)

Cache Lines: How the CPU Really Fetches Data

CPUs never load a single byte from RAM. They always load a cache line—64 consecutive bytes on x86-64. Think of it like the library delivering an entire chapter even though you only asked for one sentence.

This is why spatial locality matters so much. If your data is arranged contiguously in memory, accessing element[0] brings element[1] through element[15] into cache for free. The next 15 accesses will be L1 hits.

Cache-friendly vs cache-hostile access patterns
// FAST: sequential access — every cache line loaded is fully used
fn sum_array(data: &[u64]) -> u64 {
    data.iter().sum()     // reads data[0], data[1], ... in order
}

// SLOW: random / strided access — many cache lines loaded, each used once
fn sum_every_8th(data: &[u64]) -> u64 {
    (0..data.len()).step_by(8).map(|i| data[i]).sum()
    // CPU loads a full 64-byte cache line, uses 8 bytes, discards the rest
}

Rule of thumb: prefer arrays ( Vec<T>) over linked lists in performance-critical Rust. A linked list node can live anywhere in memory; chasing pointers causes one cache miss per node. An array stores elements back-to-back—one cache line holds 8 × u64 elements.

False Sharing: The Hidden Multi-Core Trap

Because the cache operates at the granularity of cache lines, two CPU cores that write to different variables can still slow each other down if those variables happen to live in the same 64-byte cache line. Each write on one core invalidates the line on the other core, forcing a reload. This is called false sharing.

False sharing example and fix
use std::sync::atomic::{AtomicU64, Ordering};

// BAD: counter_a and counter_b are adjacent — likely same cache line
struct Counters {
    counter_a: AtomicU64,   // bytes 0-7
    counter_b: AtomicU64,   // bytes 8-15 — same cache line as counter_a!
}

// GOOD: pad each counter to its own cache line (64 bytes)
#[repr(C, align(64))]
struct PaddedCounter {
    value: AtomicU64,
    _pad: [u8; 56],   // fill the rest of the 64-byte line
}

struct Counters {
    counter_a: PaddedCounter,  // occupies its own cache line
    counter_b: PaddedCounter,  // different cache line entirely
}

Temporal vs Spatial Locality

Temporal Locality

If you accessed a memory address recently, you will probably access it again soon. CPUs keep recently-used data in cache.

Example: a loop counter variable i is read and written on every iteration. It lives in L1 for the entire loop.

Spatial Locality

If you accessed one address, you will probably access nearby addresses soon. This is why 64-byte cache lines exist.

Example: iterating over a Vec<u8> sequentially is fast because each cache line load gives 64 elements at once.

Part 4: Core Operating System Concepts

The OS is the manager between your program and the hardware. It decides who gets CPU time, who gets memory, and who can access files or the network. Understanding its abstractions is essential for writing correct concurrent and low-level Rust.

Processes vs Threads

Analogy: A process is a complete office with its own locked rooms, filing cabinets, and staff. A thread is one employee inside that office. Multiple threads share the same office (memory space), which makes communication fast but also means one careless thread can mess up data for everyone else.

Process

  • Own virtual address space
  • Own file descriptor table
  • Isolated from other processes
  • Expensive to create (fork + exec)
  • Communication via IPC (pipes, sockets)

Thread

  • Shares address space with siblings
  • Own stack and registers
  • Cheap to create
  • Fast communication (shared memory)
  • Risk: data races (Rust prevents these!)
Spawning a thread in Rust
use std::thread;

fn main() {
    let data = vec![1u32, 2, 3, 4, 5];

    // 'move' transfers ownership of 'data' into the new thread's scope
    // The borrow checker ensures no data races at compile time
    let handle = thread::spawn(move || {
        let sum: u32 = data.iter().sum();
        println!("Sum from thread: {}", sum);
    });

    handle.join().unwrap(); // wait for thread to finish
}

The Kernel and User Space

Modern CPUs have privilege levels. The OS kernel runs in ring 0 (kernel space)—it can do anything: access hardware, manage memory, kill processes. Your program runs in ring 3 (user space)—it cannot directly touch hardware.

This separation is crucial for security and stability. A crashing user-space program cannot corrupt the kernel or other processes. To request a privileged operation (open a file, allocate memory, send a network packet), your program must make a system call (syscall).

What happens under the hood when you write to a file
// Your Rust code (user space)
use std::fs::File;
use std::io::Write;

let mut f = File::create("hello.txt")?;  // → syscall: openat()
f.write_all(b"Hello, world!")?;          // → syscall: write()
// f dropped here                         // → syscall: close()

// Each syscall:
//   1. Saves user-space registers
//   2. Switches CPU to kernel mode (ring 0)
//   3. Kernel validates and executes the request
//   4. Switches back to user mode (ring 3)
//   5. Returns result to your program
// Cost: ~100-1000 ns per syscall — avoid tight loops of syscalls!

Context Switching and Scheduling

A modern computer runs many more threads than it has CPU cores. The OS scheduler time-slices the CPU, giving each thread a short burst (a time quantum, typically 1–10 ms) before switching to another. This switch is a context switch.

Context switch cost: The CPU must save all registers of the current thread and restore those of the next. Worse, the new thread's data is not in cache, causing cold-start cache misses. A context switch typically costs 1–10 microseconds—cheap by human standards but significant in tight loops.

Why Rust async avoids excess context switches
// Traditional threads: one OS thread per concurrent task
// 10,000 connections = 10,000 threads = ~10,000 context switches/second

// Rust async: many tasks on a small thread pool
use tokio::io::{AsyncReadExt, AsyncWriteExt};

#[tokio::main]
async fn main() {
    // tokio uses a handful of OS threads (one per CPU core)
    // Tasks yield voluntarily at .await points instead of being preempted
    // No context switch when switching between tasks on the same thread
    let result = tokio::spawn(async {
        // async task: suspends at .await, not via OS context switch
        "done"
    }).await.unwrap();
    println!("{}", result);
}

Memory Allocators: The Heap Manager

The OS hands out memory to your process in large chunks (pages, 4 KB each) via the mmap or brk syscalls. But your program needs small allocations (32 bytes for a string, 16 bytes for a struct). The heap allocator (e.g. jemalloc, mimalloc, or the system allocator) sits between your code and the OS, subdividing pages into the sizes you request.

Rust's default global allocator is the system allocator (glibc's ptmalloc on Linux). You can swap it for a faster one in a single line:

Cargo.toml + main.rs: swapping to jemalloc
# Cargo.toml
[dependencies]
tikv-jemallocator = "0.5"

# main.rs
#[global_allocator]
static ALLOC: tikv_jemallocator::Jemalloc = tikv_jemallocator::Jemalloc;

fn main() {
    // Every Box::new(), Vec::new(), String::new() now uses jemalloc
    // jemalloc reduces fragmentation and improves multi-threaded throughput
    let v: Vec<u64> = (0..1_000_000).collect();
    println!("allocated {} elements", v.len());
}

Signals and Process Lifecycle

The OS communicates with processes using signals—asynchronous notifications delivered to a process. When you press Ctrl+C in the terminal, the OS sends SIGINT to your program. The default action is termination, but programs can install custom handlers to clean up gracefully.

Graceful shutdown with tokio signal handling
use tokio::signal;

#[tokio::main]
async fn main() {
    // Spawn your actual server/work here
    tokio::spawn(async {
        loop {
            // ... do work ...
            tokio::time::sleep(std::time::Duration::from_secs(1)).await;
        }
    });

    // Wait for Ctrl+C (SIGINT) or SIGTERM
    signal::ctrl_c().await.expect("failed to install signal handler");
    println!("Shutting down gracefully...");
    // cleanup: flush buffers, close connections, etc.
}

Part 5: How All of This Shapes Rust

Ownership = Deterministic Memory Management

Now that you know the stack and heap, Rust's ownership rules make complete sense:

  • Each value has one owner: no double-free bugs. The heap allocation is freed exactly once, when the owner goes out of scope.
  • Borrows don't own: a reference to heap data is just a pointer on the stack. The borrow checker verifies the heap data outlives the reference, preventing use-after-free.
  • No GC pauses: because the compiler knows at compile time exactly when to call drop, there is no runtime garbage collector scanning the heap. Latency is predictable.
Ownership in action
fn main() {
    let s1 = String::from("hello"); // heap allocated
    let s2 = s1;                    // ownership MOVES to s2
    // println!("{}", s1);          // ← compile error! s1 no longer owns the data

    let s3 = String::from("world"); // new heap allocation
    let s4 = &s3;                   // borrow — no ownership transfer
    println!("{} {}", s3, s4);      // both valid; s4 is just a pointer

}   // s3 dropped here (heap freed). s2 dropped (heap freed). s4 already gone.

Send & Sync: Thread Safety Backed by the Type System

Rust's thread model is built directly on the OS thread model you learned above. Two marker traits enforce safety at compile time:

  • Send: the type can be transferred to another thread (its data won't cause UB if owned elsewhere). Raw pointers are !Send.
  • Sync: the type can be referenced from multiple threads simultaneously. Cell<T> is !Sync because it allows mutation without a lock.
Sharing data safely across threads with Arc + Mutex
use std::sync::{Arc, Mutex};
use std::thread;

fn main() {
    // Arc = Atomically Reference Counted — heap allocation shared across threads
    // Mutex = mutual exclusion lock — only one thread writes at a time
    let counter = Arc::new(Mutex::new(0u64));

    let handles: Vec<_> = (0..8).map(|_| {
        let counter = Arc::clone(&counter);  // clone the Arc (increments ref count)
        thread::spawn(move || {
            let mut val = counter.lock().unwrap(); // acquire lock
            *val += 1;
        }) // lock released here when 'val' goes out of scope
    }).collect();

    for h in handles { h.join().unwrap(); }

    println!("Final count: {}", *counter.lock().unwrap()); // 8
}

Writing Cache-Friendly Rust: Practical Tips

Tip 1: prefer Vec over linked collections for hot paths
// Linked list: each node is a separate heap allocation → pointer chasing
use std::collections::LinkedList;
let mut list: LinkedList<u64> = LinkedList::new();

// Vec: all elements contiguous in memory → cache-line friendly
let mut vec: Vec<u64> = Vec::new();

// Benchmark result: iterating a 1M-element Vec is typically
// 3-10× faster than a LinkedList due to cache locality
Tip 2: struct of arrays (SoA) vs array of structs (AoS)
// AoS: common but cache-hostile when you only need one field
struct Particle { x: f32, y: f32, z: f32, mass: f32, charge: f32 }
let particles: Vec<Particle> = vec![...];
// Iterating over just 'x' loads 20 bytes per particle, 
// but only uses 4 — 80% of each cache line is wasted

// SoA: optimal when you process one field at a time
struct Particles {
    x: Vec<f32>,      // all X coordinates together
    y: Vec<f32>,
    z: Vec<f32>,
    mass: Vec<f32>,
    charge: Vec<f32>,
}
// Iterating over 'x' loads 16 consecutive x values per cache line — 100% utilised
Tip 3: avoid unnecessary heap allocations in hot loops
// BAD: allocates a new String on every loop iteration
for line in lines {
    let trimmed = line.trim().to_string(); // heap alloc
    process(&trimmed);
}

// GOOD: reuse a buffer across iterations
let mut buffer = String::new();
for line in lines {
    buffer.clear();
    buffer.push_str(line.trim()); // reuse existing heap allocation
    process(&buffer);
}

Mental Model Summary

Memory Concepts

  • ✓Stack: automatic, fast, fixed-size values
  • ✓Heap: manual lifetime, dynamic sizing, Rust ownership manages it
  • ✓Virtual memory: every process has its own address space, OS translates to physical RAM
  • ✓Page fault: accessing unmapped memory triggers OS handler, causes latency spike

Hardware & OS Concepts

  • ✓CPU caches (L1/L2/L3): keep hot data close to the CPU; exploit locality
  • ✓Cache lines: 64 bytes loaded at once; contiguous data = free prefetch
  • ✓Syscalls: the only door into the kernel; avoid in tight loops
  • ✓Context switch: OS saves/restores CPU state to share cores among threads

Suggested Learning Path

  1. Read the first four chapters of The Rust Programming Language (the Book) with this mental model in mind
  2. Write a small program that allocates with Box, Vec, and String, then check the generated assembly with Compiler Explorer to see stack vs heap calls
  3. Benchmark array iteration vs linked list iteration with criterion to feel cache effects concretely
  4. Write a multi-threaded program using Arc<Mutex<T>>, then try removing the Mutex and observe the compile error
  5. Use perf stat or cargo-flamegraph on a real workload to see L1/L2 miss rates

The reason Rust can make bold performance guarantees without a garbage collector is exactly because it respects these hardware realities. Once you see how the stack, heap, virtual memory, CPU caches, and OS abstractions fit together, Rust's ownership model stops feeling like arbitrary restrictions and starts feeling like the most natural way to write fast, safe systems code. Happy hacking!