Rust Bindgen & Miri: Bridging C Libraries and Catching Undefined Behavior Before It Catches You

26 min read • Rust FFI, Unsafe Code Analysis & Tooling

The systems world is built on decades of C and C++ libraries. Rust's entire selling point is safety—but that safety has to coexist with millions of lines of existing C code. Two tools make this possible: bindgen automatically generates the Rust declarations you need to call C code, and Miri acts as an interpreter that catches undefined behaviour in your unsafe Rust before it ever reaches production.

This guide walks through both tools end-to-end. We start with real-world analogies so every concept is grounded before we write code. By the end you will know how to call a C library from Rust with generated bindings and how to run Miri on your unsafe code to catch bugs that the borrow checker cannot see.

No prior FFI experience required. Let's build the mental model first.

Part 1: Why FFI and Why It Is Hard

The Power Adapter Analogy

Imagine you bought a laptop in the US and you are now visiting Europe. The laptop still works perfectly—but the plug shape is wrong. You do not throw the laptop away. You use a power adapter: a small shim that speaks both plug languages.

FFI (Foreign Function Interface) is exactly that adapter. Rust and C are different languages with different calling conventions, different type sizes, and different ownership rules. FFI is the shim that lets them talk. Bindgen writes that shim for you automatically.

What the Compiler Needs to Call C

To call a C function from Rust, the compiler needs three things:

  • The function signature: name, argument types, return type. These must match the C header exactly or the program will silently pass the wrong bytes.
  • The ABI declaration: extern "C" tells Rust to use the C calling convention (argument order, register usage, stack layout).
  • The compiled library: a .a (static) or .so / .dylib (dynamic) that the linker will find at build time.
Manual FFI binding — what you write without bindgen
// C header (math_utils.h):
// int add(int a, int b);
// double sqrt_approx(double x);

// What you must write by hand in Rust:
extern "C" {
    fn add(a: i32, b: i32) -> i32;
    fn sqrt_approx(x: f64) -> f64;
}

fn main() {
    // Every call into C is unsafe — Rust cannot verify C's behaviour
    unsafe {
        println!("{}", add(3, 4));         // 7
        println!("{:.4}", sqrt_approx(2.0)); // ~1.4142
    }
}

The problem: real C libraries have hundreds of functions, nested structs, bitfields, platform-specific typedefs, and enums. Writing all of this by hand is error-prone and breaks every time the C library updates. That is the problem bindgen solves.

Part 2: bindgen — Automatic Rust Bindings from C Headers

The Legal Translator Analogy

A C header file (.h) is a contract written in “C legalese”: types, function signatures, constants. bindgen is a professional translator that reads this contract and produces an equivalent contract in “Rust legalese”—with the exact same meaning, just in the target language. You feed it a header; it outputs a .rs file.

bindgen is built on top of libclang, the same C/C++ parser that powers Clang and LLVM. This means it handles preprocessor macros, platform ifdefs, anonymous structs, and bitfields correctly— things that are nearly impossible to handle with a simple text-based generator.

Installation

bindgen requires libclang to be installed on your system (it ships with the LLVM toolchain).

Install libclang + bindgen CLI
# macOS
brew install llvm

# Ubuntu / Debian
sudo apt install libclang-dev

# Install the bindgen CLI tool
cargo install bindgen-cli

Using bindgen from the Command Line

The fastest way to see bindgen in action is to point it at a C header and let it generate a bindings file:

snappy.h (simplified subset of Google's Snappy library)
// snappy.h
#include <stddef.h>

typedef enum {
    SNAPPY_OK = 0,
    SNAPPY_INVALID_INPUT = 1,
    SNAPPY_BUFFER_TOO_SMALL = 2,
} snappy_status;

size_t snappy_max_compressed_length(size_t source_length);

snappy_status snappy_compress(
    const char* input,
    size_t input_length,
    char* compressed,
    size_t* compressed_length
);

snappy_status snappy_uncompress(
    const char* compressed,
    size_t compressed_length,
    char* uncompressed,
    size_t* uncompressed_length
);
Generate bindings with the CLI
bindgen snappy.h -o src/bindings.rs

# With extra flags for a specific target architecture
bindgen snappy.h     --use-core     --ctypes-prefix "::core::ffi"     -- -I/usr/include     -o src/bindings.rs
Generated src/bindings.rs (abbreviated)
/* automatically generated by rust-bindgen 0.70.x */

#[repr(u32)]
#[derive(Debug, Copy, Clone, Hash, PartialEq, Eq)]
pub enum snappy_status {
    SNAPPY_OK = 0,
    SNAPPY_INVALID_INPUT = 1,
    SNAPPY_BUFFER_TOO_SMALL = 2,
}

extern "C" {
    pub fn snappy_max_compressed_length(source_length: usize) -> usize;

    pub fn snappy_compress(
        input: *const ::std::os::raw::c_char,
        input_length: usize,
        compressed: *mut ::std::os::raw::c_char,
        compressed_length: *mut usize,
    ) -> snappy_status;

    pub fn snappy_uncompress(
        compressed: *const ::std::os::raw::c_char,
        compressed_length: usize,
        uncompressed: *mut ::std::os::raw::c_char,
        uncompressed_length: *mut usize,
    ) -> snappy_status;
}

What bindgen translates automatically: C enums → Rust enums with #[repr(u32)], C size_t → Rust usize, C char * → Rust *const c_char, structs, unions, bitfields, constants, and function pointer typedefs.

The Recommended Pattern: bindgen in build.rs

The idiomatic way to use bindgen in a project is to run it automatically inside build.rs— Cargo's build script. Bindings are regenerated every time the header changes. No manual step, no stale bindings.

Project layout
my-snappy/
├── build.rs          ← generates bindings at build time
├── Cargo.toml
├── src/
│   └── lib.rs        ← safe Rust wrapper
├── wrapper.h         ← thin C header that includes snappy.h
└── libsnappy.a       ← pre-compiled static library
Cargo.toml
[package]
name = "my-snappy"
version = "0.1.0"
edition = "2021"

[build-dependencies]
bindgen = "0.70"

[dependencies]
build.rs
use std::path::PathBuf;

fn main() {
    // Tell cargo to re-run this script if the header changes
    println!("cargo:rerun-if-changed=wrapper.h");

    // Link against the static snappy library
    println!("cargo:rustc-link-lib=static=snappy");
    println!("cargo:rustc-link-search=native=.");

    let bindings = bindgen::Builder::default()
        .header("wrapper.h")
        // Only generate bindings for symbols with the 'snappy_' prefix
        .allowlist_function("snappy_.*")
        .allowlist_type("snappy_.*")
        // Generate derive(Debug, PartialEq) where possible
        .derive_debug(true)
        .derive_partialeq(true)
        // Tell bindgen to use core instead of std (for no_std crates)
        // .use_core()
        .parse_callbacks(Box::new(bindgen::CargoCallbacks::new()))
        .generate()
        .expect("Unable to generate bindings");

    let out_path = PathBuf::from(std::env::var("OUT_DIR").unwrap());
    bindings
        .write_to_file(out_path.join("bindings.rs"))
        .expect("Couldn't write bindings!");
}
src/lib.rs — safe wrapper over the raw bindings
// Pull in the generated bindings
#[allow(non_upper_case_globals, non_camel_case_types, dead_code)]
mod ffi {
    include!(concat!(env!("OUT_DIR"), "/bindings.rs"));
}

use ffi::snappy_status;

#[derive(Debug, thiserror::Error)]
pub enum SnappyError {
    #[error("invalid input data")]
    InvalidInput,
    #[error("output buffer too small")]
    BufferTooSmall,
}

/// Compress bytes using Snappy. Returns the compressed data.
pub fn compress(input: &[u8]) -> Vec<u8> {
    let max_len = unsafe { ffi::snappy_max_compressed_length(input.len()) };
    let mut output = vec![0u8; max_len];
    let mut output_len = max_len;

    let status = unsafe {
        ffi::snappy_compress(
            input.as_ptr() as *const i8,
            input.len(),
            output.as_mut_ptr() as *mut i8,
            &mut output_len,
        )
    };

    assert_eq!(status, snappy_status::SNAPPY_OK);
    output.truncate(output_len);
    output
}

/// Decompress Snappy-compressed bytes.
pub fn decompress(input: &[u8], max_output: usize) -> Result<Vec<u8>, SnappyError> {
    let mut output = vec![0u8; max_output];
    let mut output_len = max_output;

    let status = unsafe {
        ffi::snappy_uncompress(
            input.as_ptr() as *const i8,
            input.len(),
            output.as_mut_ptr() as *mut i8,
            &mut output_len,
        )
    };

    match status {
        snappy_status::SNAPPY_OK => {
            output.truncate(output_len);
            Ok(output)
        }
        snappy_status::SNAPPY_INVALID_INPUT => Err(SnappyError::InvalidInput),
        snappy_status::SNAPPY_BUFFER_TOO_SMALL => Err(SnappyError::BufferTooSmall),
        _ => unreachable!(),
    }
}

#[cfg(test)]
mod tests {
    use super::*;

    #[test]
    fn roundtrip() {
        let original = b"hello world, hello world, hello world!";
        let compressed = compress(original);
        let decompressed = decompress(&compressed, original.len() * 2).unwrap();
        assert_eq!(decompressed, original);
    }
}

Key bindgen Builder Options

OptionWhat it does
.allowlist_function()Only generate bindings for functions matching a regex. Keeps output small.
.blocklist_item()Exclude a specific symbol (e.g. one that bindgen cannot translate correctly).
.derive_default(true)Add #[derive(Default)] to generated structs where possible.
.use_core()Use core instead of std — required for no_std embedded targets.
.size_t_is_usize(true)Map C size_t directly to Rust usize.
.parse_callbacks()Hook into parsing events — e.g. CargoCallbacks emits cargo:rerun-if-changed for every included header automatically.

Handling Tricky C Constructs

Bitfields — bindgen maps them to accessor methods
// C:
// struct Flags {
//     unsigned int active : 1;
//     unsigned int level  : 3;
//     unsigned int mode   : 4;
// };

// Generated Rust (bindgen uses a generated _bitfield_1 backing field):
#[repr(C)]
pub struct Flags {
    pub _bitfield_align_1: [u8; 0],
    pub _bitfield_1: bindgen::__BindgenBitfieldUnit<[u8; 1usize]>,
}
impl Flags {
    pub fn active(&self) -> u32 { ... }
    pub fn set_active(&mut self, val: u32) { ... }
    pub fn level(&self) -> u32 { ... }
    pub fn mode(&self) -> u32 { ... }
}
Opaque types — when C hides the struct layout
// C library only exposes a forward declaration:
// typedef struct sqlite3 sqlite3;
// You never see the actual fields.

// In build.rs, tell bindgen to treat it as an opaque blob:
bindgen::Builder::default()
    .header("sqlite3.h")
    .opaque_type("sqlite3")   // ← treat as *mut c_void effectively
    .generate()
    .unwrap();

// Generated:
#[repr(C)]
#[derive(Debug, Copy, Clone)]
pub struct sqlite3 {
    pub _unused: [u8; 0],   // zero-size opaque marker
}

Part 3: Writing Safe Wrappers — The Golden Rule of FFI

The Nuclear Plant Control Room Analogy

A nuclear plant has dangerous reactor controls that only trained operators can touch. They are locked behind a thick door. Operators access them through safe consoles with interlocks, alarms, and confirmations that prevent accidental misuse. The raw controls still exist, but the safe console is what everyday operators use. In Rust, the raw FFI bindings are the reactor controls; your safe wrapper module is the control console.

Safety Invariants You Must Document

Every unsafe block that calls into C must be accompanied by a comment explaining why it is correct. The three most common invariants:

Documenting safety invariants on unsafe calls
/// Compress a byte slice using Snappy.
///
/// # Safety
/// This function is safe to call from Rust because:
/// 1. input.as_ptr() is valid for input.len() bytes (guaranteed by &[u8])
/// 2. output.as_mut_ptr() is valid for output_len bytes (Vec owns allocation)
/// 3. snappy_compress does not store the pointers after returning
/// 4. snappy_compress is thread-safe according to its documentation
pub fn compress(input: &[u8]) -> Vec<u8> {
    // SAFETY: see doc comment above
    let max_len = unsafe { ffi::snappy_max_compressed_length(input.len()) };
    let mut output = vec![0u8; max_len];
    let mut output_len = max_len;

    // SAFETY: pointers are valid, lengths are correct, no aliasing
    let status = unsafe {
        ffi::snappy_compress(
            input.as_ptr().cast(),
            input.len(),
            output.as_mut_ptr().cast(),
            &mut output_len,
        )
    };

    assert_eq!(status, ffi::snappy_status::SNAPPY_OK);
    output.truncate(output_len);
    output
}

Rule: the unsafe block should be as small as possible. Validate inputs, check outputs, and do all safe work outside the block. The unsafe keyword is a promise to the compiler and to your future self: “I have verified this is correct and here is why.”

Part 4: Miri — An Interpreter That Catches Undefined Behaviour

The Aircraft Pre-Flight Checklist Analogy

Before a pilot takes off they walk around the plane with a detailed checklist—testing each control surface, checking fluid levels, inspecting tyres. The plane looks fine from the outside, but the checklist catches issues that are invisible until they cause a crash at 10,000 feet.

Miri is that pre-flight checklist for unsafe Rust. It interprets your program one step at a time, tracking extra metadata about every byte of memory. If any step violates memory safety rules, Miri stops immediately and reports exactly what went wrong, where, and why—long before the bug could cause a production crash or a security vulnerability.

What Miri Can Detect

Memory Bugs

  • ✗Use-after-free
  • ✗Reads of uninitialized memory
  • ✗Out-of-bounds pointer arithmetic
  • ✗Dangling pointers
  • ✗Heap allocation leaks (opt-in)

Aliasing & UB Bugs

  • ✗Stacked Borrows / Tree Borrows violations (aliasing rules)
  • ✗Transmuting a type to an incompatible layout
  • ✗Invalid enum discriminants
  • ✗Misaligned pointer dereferences
  • ✗Data races (via -Zmiri-track-raw-pointers)

What Miri cannot do: it cannot check code that calls into actual C libraries (FFI calls are unsupported). It also runs significantly slower than normal execution (10–1000× depending on the workload), so it is used in tests, not in the hot path.

Installing and Running Miri

Install Miri via rustup
# Miri ships as a rustup component on nightly
rustup +nightly component add miri

# Run your tests under Miri
cargo +nightly miri test

# Run a specific binary under Miri
cargo +nightly miri run

# Run with a specific Miri flag
MIRIFLAGS="-Zmiri-strict-provenance" cargo +nightly miri test

Nightly only: Miri depends on Rust compiler internals that are not stable. You need a nightly toolchain, but your library can target stable Rust—only the Miri invocation needs nightly.

Miri Catching Real Bugs: Examples

Bug 1: use-after-free via raw pointer
fn main() {
    let x = Box::new(42i32);
    let ptr: *const i32 = &*x;

    drop(x); // x is freed here

    // SAFETY: this is NOT safe — ptr is now dangling
    let val = unsafe { *ptr }; // ← undefined behavior!
    println!("{}", val);
}

// When run under normal cargo run: might print 42 (or anything — UB!)
// When run under Miri:
//
// error: Undefined Behavior: use-after-free: pointer to alloc814 was dereferenced
//   after the allocation was freed
//  --> src/main.rs:7:22
//   |
// 7 |     let val = unsafe { *ptr };
//   |                        ^^^^ use-after-free
Bug 2: reading uninitialized memory
use std::mem::MaybeUninit;

fn main() {
    let mut x: MaybeUninit<i32> = MaybeUninit::uninit();

    // BUG: reading without initializing first
    let val = unsafe { x.assume_init() }; // ← UB: memory is uninitialized
    println!("{}", val);
}

// Miri output:
// error: Undefined Behavior: using uninitialized data, but this operation
//   requires initialized memory
//  --> src/main.rs:6:17
//   |
// 6 |     let val = unsafe { x.assume_init() };
//   |                        ^^^^^^^^^^^^^^^ using uninitialized data
Bug 3: misaligned pointer dereference
fn main() {
    let data: [u8; 8] = [0u8; 8];
    let ptr = data.as_ptr();

    // Cast a *u8 to *u32 and dereference — requires 4-byte alignment
    // data.as_ptr() may only be 1-byte aligned
    let misaligned = unsafe { *(ptr.add(1) as *const u32) }; // ← UB!
    println!("{}", misaligned);
}

// Miri output:
// error: Undefined Behavior: accessing memory with alignment 1,
//   but alignment 4 is required
//  --> src/main.rs:7:24
//   |
// 7 |     let misaligned = unsafe { *(ptr.add(1) as *const u32) };
//   |                                ^^^^^^^^^^^^^^^^^^^^^^^^^^
Bug 4: invalid enum discriminant via transmute
#[repr(u8)]
enum Direction { North = 0, South = 1, East = 2, West = 3 }

fn main() {
    // 99 is not a valid Direction discriminant
    let d: Direction = unsafe { std::mem::transmute(99u8) }; // ← UB!
    match d {
        Direction::North => println!("North"),
        _ => println!("other"),
    }
}

// Normal run: might print "other", might crash, might corrupt memory
// Miri:
// error: Undefined Behavior: constructing invalid value: encountered 99, but
//   expected a valid enum discriminant
//  --> src/main.rs:5:37

Stacked Borrows: Miri's Aliasing Model

Beyond simple memory safety, Miri implements an aliasing model called Stacked Borrows (and the newer Tree Borrows). This checks that raw pointers are used in ways consistent with Rust's ownership rules even inside unsafe blocks.

Stacked Borrows violation — two mutable aliases
fn main() {
    let mut x = 5i32;
    let ptr1: *mut i32 = &mut x;
    let ptr2: *mut i32 = &mut x; // second mutable ref through raw pointer

    unsafe {
        *ptr1 = 10; // use ptr1 — invalidates ptr2 under Stacked Borrows
        *ptr2 = 20; // ← UB: ptr2 was invalidated when ptr1 was used
        println!("{}", x);
    }
}

// Miri output:
// error: Undefined Behavior: attempting a write access using <tag> at ...
//   but that tag does not exist in the borrow stack for this location
//   this error occurs as part of an access at ...

Tree Borrows: a newer, less restrictive aliasing model that allows patterns Stacked Borrows rejects, while still catching real bugs. Enable it with MIRIFLAGS="-Zmiri-tree-borrows". Major crates like Vec and HashMap pass under Tree Borrows but were flagged by Stacked Borrows.

Part 5: Using Miri on Wrapper Code

Miri cannot execute across the FFI boundary into real C libraries, but it can still check every line of your Rust wrapper code up to the boundary. The trick is to mock the C functions in tests, replacing them with safe Rust implementations that have the same signature. Miri then runs through the full logic of your wrapper.

Testing a Wrapper with Mocked FFI

src/lib.rs — conditional mock for Miri tests
// In tests we swap the real FFI for a pure-Rust mock
// so Miri can run end-to-end without hitting the C boundary.
#[cfg(test)]
mod ffi {
    // Mirror the real ffi types
    #[repr(u32)]
    #[derive(PartialEq)]
    pub enum snappy_status { SNAPPY_OK = 0, SNAPPY_INVALID_INPUT = 1 }

    // Pure-Rust mock: just copies the data through
    pub unsafe fn snappy_max_compressed_length(n: usize) -> usize { n * 2 }
    pub unsafe fn snappy_compress(
        input: *const i8, input_len: usize,
        output: *mut i8, output_len: *mut usize,
    ) -> snappy_status {
        let src = std::slice::from_raw_parts(input as *const u8, input_len);
        let dst = std::slice::from_raw_parts_mut(output as *mut u8, input_len);
        dst.copy_from_slice(src);
        *output_len = input_len;
        snappy_status::SNAPPY_OK
    }
}

#[cfg(not(test))]
mod ffi {
    include!(concat!(env!("OUT_DIR"), "/bindings.rs"));
}

#[cfg(test)]
mod tests {
    use super::*;

    #[test]  // cargo +nightly miri test runs this under Miri
    fn compress_does_not_leak_or_alias() {
        let data = b"hello";
        let compressed = compress(data);
        assert_eq!(compressed.len(), data.len());
    }
}

Part 6: Recommended Workflow

  1. Write the C header (or obtain it from the library). Create a thin wrapper.h that #includes only the parts you need.
  2. Add bindgen to build.rs with allowlist_function to keep bindings minimal. Commit the build.rs, not the generated file (it is always regenerated from source).
  3. Write the safe wrapper in src/lib.rs. Every unsafe block must have a // SAFETY: comment.
  4. Write unit tests with mocked FFI so Miri can run them. Cover edge cases: empty slices, null-like inputs, maximum sizes.
  5. Run Miri in CI on every pull request: cargo +nightly miri test. Treat a Miri failure as a blocker, just like a compilation error.
  6. Enable additional Miri flags for thorough checking: -Zmiri-strict-provenance and -Zmiri-symbolic-alignment-check.

bindgen Checklist

  • ✓Use build.rs — never commit generated bindings
  • ✓Use allowlist_function to limit surface area
  • ✓Add CargoCallbacks so headers trigger rebuilds
  • ✓Test on all target platforms — sizes and alignment may differ
  • ✓Mark generated module #[allow(clippy::all)]

Miri Checklist

  • ✓Run cargo +nightly miri test in CI
  • ✓Mock FFI functions so Miri can run end-to-end
  • ✓Add // SAFETY: comments on every unsafe block
  • ✓Try -Zmiri-tree-borrows if Stacked Borrows rejects valid code
  • ✓Use -Zmiri-strict-provenance to catch pointer-to-integer casts

bindgen and Miri are two sides of the same coin. bindgen lets you reach into the vast ecosystem of C libraries without writing tedious boilerplate by hand. Miri catches the subtle memory bugs that only show up in unsafe code—the kind that the borrow checker is not designed to see. Together they give you a workflow where interoperating with C is manageable and your unsafe code is as rigorously tested as your safe code. Use bindgen in build.rs, wrap it safely, mock the boundary, and run Miri on every PR. That is the production-quality FFI workflow. Happy hacking!