Rust Bindgen & Miri: Bridging C Libraries and Catching Undefined Behavior Before It Catches You
The systems world is built on decades of C and C++ libraries. Rust's entire selling point is safety—but that safety has to coexist with millions of lines of existing C code. Two tools make this possible: bindgen automatically generates the Rust declarations you need to call C code, and Miri acts as an interpreter that catches undefined behaviour in your unsafe Rust before it ever reaches production.
This guide walks through both tools end-to-end. We start with real-world analogies so every concept is grounded before we write code. By the end you will know how to call a C library from Rust with generated bindings and how to run Miri on your unsafe code to catch bugs that the borrow checker cannot see.
No prior FFI experience required. Let's build the mental model first.
Part 1: Why FFI and Why It Is Hard
The Power Adapter Analogy
Imagine you bought a laptop in the US and you are now visiting Europe. The laptop still works perfectly—but the plug shape is wrong. You do not throw the laptop away. You use a power adapter: a small shim that speaks both plug languages.
FFI (Foreign Function Interface) is exactly that adapter. Rust and C are different languages with different calling conventions, different type sizes, and different ownership rules. FFI is the shim that lets them talk. Bindgen writes that shim for you automatically.
What the Compiler Needs to Call C
To call a C function from Rust, the compiler needs three things:
- The function signature: name, argument types, return type. These must match the C header exactly or the program will silently pass the wrong bytes.
- The ABI declaration:
extern "C"tells Rust to use the C calling convention (argument order, register usage, stack layout). - The compiled library: a
.a(static) or.so/.dylib(dynamic) that the linker will find at build time.
// C header (math_utils.h):
// int add(int a, int b);
// double sqrt_approx(double x);
// What you must write by hand in Rust:
extern "C" {
fn add(a: i32, b: i32) -> i32;
fn sqrt_approx(x: f64) -> f64;
}
fn main() {
// Every call into C is unsafe — Rust cannot verify C's behaviour
unsafe {
println!("{}", add(3, 4)); // 7
println!("{:.4}", sqrt_approx(2.0)); // ~1.4142
}
}The problem: real C libraries have hundreds of functions, nested structs, bitfields, platform-specific typedefs, and enums. Writing all of this by hand is error-prone and breaks every time the C library updates. That is the problem bindgen solves.
Part 2: bindgen — Automatic Rust Bindings from C Headers
The Legal Translator Analogy
A C header file (.h) is a contract written in “C legalese”: types, function signatures, constants. bindgen is a professional translator that reads this contract and produces an equivalent contract in “Rust legalese”—with the exact same meaning, just in the target language. You feed it a header; it outputs a .rs file.
bindgen is built on top of libclang, the same C/C++ parser that powers Clang and LLVM. This means it handles preprocessor macros, platform ifdefs, anonymous structs, and bitfields correctly— things that are nearly impossible to handle with a simple text-based generator.
Installation
bindgen requires libclang to be installed on your system (it ships with the LLVM toolchain).
# macOS
brew install llvm
# Ubuntu / Debian
sudo apt install libclang-dev
# Install the bindgen CLI tool
cargo install bindgen-cliUsing bindgen from the Command Line
The fastest way to see bindgen in action is to point it at a C header and let it generate a bindings file:
// snappy.h
#include <stddef.h>
typedef enum {
SNAPPY_OK = 0,
SNAPPY_INVALID_INPUT = 1,
SNAPPY_BUFFER_TOO_SMALL = 2,
} snappy_status;
size_t snappy_max_compressed_length(size_t source_length);
snappy_status snappy_compress(
const char* input,
size_t input_length,
char* compressed,
size_t* compressed_length
);
snappy_status snappy_uncompress(
const char* compressed,
size_t compressed_length,
char* uncompressed,
size_t* uncompressed_length
);bindgen snappy.h -o src/bindings.rs
# With extra flags for a specific target architecture
bindgen snappy.h --use-core --ctypes-prefix "::core::ffi" -- -I/usr/include -o src/bindings.rs/* automatically generated by rust-bindgen 0.70.x */
#[repr(u32)]
#[derive(Debug, Copy, Clone, Hash, PartialEq, Eq)]
pub enum snappy_status {
SNAPPY_OK = 0,
SNAPPY_INVALID_INPUT = 1,
SNAPPY_BUFFER_TOO_SMALL = 2,
}
extern "C" {
pub fn snappy_max_compressed_length(source_length: usize) -> usize;
pub fn snappy_compress(
input: *const ::std::os::raw::c_char,
input_length: usize,
compressed: *mut ::std::os::raw::c_char,
compressed_length: *mut usize,
) -> snappy_status;
pub fn snappy_uncompress(
compressed: *const ::std::os::raw::c_char,
compressed_length: usize,
uncompressed: *mut ::std::os::raw::c_char,
uncompressed_length: *mut usize,
) -> snappy_status;
}What bindgen translates automatically: C enums → Rust enums with #[repr(u32)], C size_t → Rust usize, C char * → Rust *const c_char, structs, unions, bitfields, constants, and function pointer typedefs.
The Recommended Pattern: bindgen in build.rs
The idiomatic way to use bindgen in a project is to run it automatically inside build.rs— Cargo's build script. Bindings are regenerated every time the header changes. No manual step, no stale bindings.
my-snappy/
├── build.rs ← generates bindings at build time
├── Cargo.toml
├── src/
│ └── lib.rs ← safe Rust wrapper
├── wrapper.h ← thin C header that includes snappy.h
└── libsnappy.a ← pre-compiled static library[package]
name = "my-snappy"
version = "0.1.0"
edition = "2021"
[build-dependencies]
bindgen = "0.70"
[dependencies]use std::path::PathBuf;
fn main() {
// Tell cargo to re-run this script if the header changes
println!("cargo:rerun-if-changed=wrapper.h");
// Link against the static snappy library
println!("cargo:rustc-link-lib=static=snappy");
println!("cargo:rustc-link-search=native=.");
let bindings = bindgen::Builder::default()
.header("wrapper.h")
// Only generate bindings for symbols with the 'snappy_' prefix
.allowlist_function("snappy_.*")
.allowlist_type("snappy_.*")
// Generate derive(Debug, PartialEq) where possible
.derive_debug(true)
.derive_partialeq(true)
// Tell bindgen to use core instead of std (for no_std crates)
// .use_core()
.parse_callbacks(Box::new(bindgen::CargoCallbacks::new()))
.generate()
.expect("Unable to generate bindings");
let out_path = PathBuf::from(std::env::var("OUT_DIR").unwrap());
bindings
.write_to_file(out_path.join("bindings.rs"))
.expect("Couldn't write bindings!");
}// Pull in the generated bindings
#[allow(non_upper_case_globals, non_camel_case_types, dead_code)]
mod ffi {
include!(concat!(env!("OUT_DIR"), "/bindings.rs"));
}
use ffi::snappy_status;
#[derive(Debug, thiserror::Error)]
pub enum SnappyError {
#[error("invalid input data")]
InvalidInput,
#[error("output buffer too small")]
BufferTooSmall,
}
/// Compress bytes using Snappy. Returns the compressed data.
pub fn compress(input: &[u8]) -> Vec<u8> {
let max_len = unsafe { ffi::snappy_max_compressed_length(input.len()) };
let mut output = vec![0u8; max_len];
let mut output_len = max_len;
let status = unsafe {
ffi::snappy_compress(
input.as_ptr() as *const i8,
input.len(),
output.as_mut_ptr() as *mut i8,
&mut output_len,
)
};
assert_eq!(status, snappy_status::SNAPPY_OK);
output.truncate(output_len);
output
}
/// Decompress Snappy-compressed bytes.
pub fn decompress(input: &[u8], max_output: usize) -> Result<Vec<u8>, SnappyError> {
let mut output = vec![0u8; max_output];
let mut output_len = max_output;
let status = unsafe {
ffi::snappy_uncompress(
input.as_ptr() as *const i8,
input.len(),
output.as_mut_ptr() as *mut i8,
&mut output_len,
)
};
match status {
snappy_status::SNAPPY_OK => {
output.truncate(output_len);
Ok(output)
}
snappy_status::SNAPPY_INVALID_INPUT => Err(SnappyError::InvalidInput),
snappy_status::SNAPPY_BUFFER_TOO_SMALL => Err(SnappyError::BufferTooSmall),
_ => unreachable!(),
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn roundtrip() {
let original = b"hello world, hello world, hello world!";
let compressed = compress(original);
let decompressed = decompress(&compressed, original.len() * 2).unwrap();
assert_eq!(decompressed, original);
}
}Key bindgen Builder Options
| Option | What it does |
|---|---|
.allowlist_function() | Only generate bindings for functions matching a regex. Keeps output small. |
.blocklist_item() | Exclude a specific symbol (e.g. one that bindgen cannot translate correctly). |
.derive_default(true) | Add #[derive(Default)] to generated structs where possible. |
.use_core() | Use core instead of std — required for no_std embedded targets. |
.size_t_is_usize(true) | Map C size_t directly to Rust usize. |
.parse_callbacks() | Hook into parsing events — e.g. CargoCallbacks emits cargo:rerun-if-changed for every included header automatically. |
Handling Tricky C Constructs
// C:
// struct Flags {
// unsigned int active : 1;
// unsigned int level : 3;
// unsigned int mode : 4;
// };
// Generated Rust (bindgen uses a generated _bitfield_1 backing field):
#[repr(C)]
pub struct Flags {
pub _bitfield_align_1: [u8; 0],
pub _bitfield_1: bindgen::__BindgenBitfieldUnit<[u8; 1usize]>,
}
impl Flags {
pub fn active(&self) -> u32 { ... }
pub fn set_active(&mut self, val: u32) { ... }
pub fn level(&self) -> u32 { ... }
pub fn mode(&self) -> u32 { ... }
}// C library only exposes a forward declaration:
// typedef struct sqlite3 sqlite3;
// You never see the actual fields.
// In build.rs, tell bindgen to treat it as an opaque blob:
bindgen::Builder::default()
.header("sqlite3.h")
.opaque_type("sqlite3") // ← treat as *mut c_void effectively
.generate()
.unwrap();
// Generated:
#[repr(C)]
#[derive(Debug, Copy, Clone)]
pub struct sqlite3 {
pub _unused: [u8; 0], // zero-size opaque marker
}Part 3: Writing Safe Wrappers — The Golden Rule of FFI
The Nuclear Plant Control Room Analogy
A nuclear plant has dangerous reactor controls that only trained operators can touch. They are locked behind a thick door. Operators access them through safe consoles with interlocks, alarms, and confirmations that prevent accidental misuse. The raw controls still exist, but the safe console is what everyday operators use. In Rust, the raw FFI bindings are the reactor controls; your safe wrapper module is the control console.
Safety Invariants You Must Document
Every unsafe block that calls into C must be accompanied by a comment explaining why it is correct. The three most common invariants:
/// Compress a byte slice using Snappy.
///
/// # Safety
/// This function is safe to call from Rust because:
/// 1. input.as_ptr() is valid for input.len() bytes (guaranteed by &[u8])
/// 2. output.as_mut_ptr() is valid for output_len bytes (Vec owns allocation)
/// 3. snappy_compress does not store the pointers after returning
/// 4. snappy_compress is thread-safe according to its documentation
pub fn compress(input: &[u8]) -> Vec<u8> {
// SAFETY: see doc comment above
let max_len = unsafe { ffi::snappy_max_compressed_length(input.len()) };
let mut output = vec![0u8; max_len];
let mut output_len = max_len;
// SAFETY: pointers are valid, lengths are correct, no aliasing
let status = unsafe {
ffi::snappy_compress(
input.as_ptr().cast(),
input.len(),
output.as_mut_ptr().cast(),
&mut output_len,
)
};
assert_eq!(status, ffi::snappy_status::SNAPPY_OK);
output.truncate(output_len);
output
}Rule: the unsafe block should be as small as possible. Validate inputs, check outputs, and do all safe work outside the block. The unsafe keyword is a promise to the compiler and to your future self: “I have verified this is correct and here is why.”
Part 4: Miri — An Interpreter That Catches Undefined Behaviour
The Aircraft Pre-Flight Checklist Analogy
Before a pilot takes off they walk around the plane with a detailed checklist—testing each control surface, checking fluid levels, inspecting tyres. The plane looks fine from the outside, but the checklist catches issues that are invisible until they cause a crash at 10,000 feet.
Miri is that pre-flight checklist for unsafe Rust. It interprets your program one step at a time, tracking extra metadata about every byte of memory. If any step violates memory safety rules, Miri stops immediately and reports exactly what went wrong, where, and why—long before the bug could cause a production crash or a security vulnerability.
What Miri Can Detect
Memory Bugs
- ✗Use-after-free
- ✗Reads of uninitialized memory
- ✗Out-of-bounds pointer arithmetic
- ✗Dangling pointers
- ✗Heap allocation leaks (opt-in)
Aliasing & UB Bugs
- ✗Stacked Borrows / Tree Borrows violations (aliasing rules)
- ✗Transmuting a type to an incompatible layout
- ✗Invalid enum discriminants
- ✗Misaligned pointer dereferences
- ✗Data races (via
-Zmiri-track-raw-pointers)
What Miri cannot do: it cannot check code that calls into actual C libraries (FFI calls are unsupported). It also runs significantly slower than normal execution (10–1000× depending on the workload), so it is used in tests, not in the hot path.
Installing and Running Miri
# Miri ships as a rustup component on nightly
rustup +nightly component add miri
# Run your tests under Miri
cargo +nightly miri test
# Run a specific binary under Miri
cargo +nightly miri run
# Run with a specific Miri flag
MIRIFLAGS="-Zmiri-strict-provenance" cargo +nightly miri testNightly only: Miri depends on Rust compiler internals that are not stable. You need a nightly toolchain, but your library can target stable Rust—only the Miri invocation needs nightly.
Miri Catching Real Bugs: Examples
fn main() {
let x = Box::new(42i32);
let ptr: *const i32 = &*x;
drop(x); // x is freed here
// SAFETY: this is NOT safe — ptr is now dangling
let val = unsafe { *ptr }; // ← undefined behavior!
println!("{}", val);
}
// When run under normal cargo run: might print 42 (or anything — UB!)
// When run under Miri:
//
// error: Undefined Behavior: use-after-free: pointer to alloc814 was dereferenced
// after the allocation was freed
// --> src/main.rs:7:22
// |
// 7 | let val = unsafe { *ptr };
// | ^^^^ use-after-freeuse std::mem::MaybeUninit;
fn main() {
let mut x: MaybeUninit<i32> = MaybeUninit::uninit();
// BUG: reading without initializing first
let val = unsafe { x.assume_init() }; // ← UB: memory is uninitialized
println!("{}", val);
}
// Miri output:
// error: Undefined Behavior: using uninitialized data, but this operation
// requires initialized memory
// --> src/main.rs:6:17
// |
// 6 | let val = unsafe { x.assume_init() };
// | ^^^^^^^^^^^^^^^ using uninitialized datafn main() {
let data: [u8; 8] = [0u8; 8];
let ptr = data.as_ptr();
// Cast a *u8 to *u32 and dereference — requires 4-byte alignment
// data.as_ptr() may only be 1-byte aligned
let misaligned = unsafe { *(ptr.add(1) as *const u32) }; // ← UB!
println!("{}", misaligned);
}
// Miri output:
// error: Undefined Behavior: accessing memory with alignment 1,
// but alignment 4 is required
// --> src/main.rs:7:24
// |
// 7 | let misaligned = unsafe { *(ptr.add(1) as *const u32) };
// | ^^^^^^^^^^^^^^^^^^^^^^^^^^#[repr(u8)]
enum Direction { North = 0, South = 1, East = 2, West = 3 }
fn main() {
// 99 is not a valid Direction discriminant
let d: Direction = unsafe { std::mem::transmute(99u8) }; // ← UB!
match d {
Direction::North => println!("North"),
_ => println!("other"),
}
}
// Normal run: might print "other", might crash, might corrupt memory
// Miri:
// error: Undefined Behavior: constructing invalid value: encountered 99, but
// expected a valid enum discriminant
// --> src/main.rs:5:37Stacked Borrows: Miri's Aliasing Model
Beyond simple memory safety, Miri implements an aliasing model called Stacked Borrows (and the newer Tree Borrows). This checks that raw pointers are used in ways consistent with Rust's ownership rules even inside unsafe blocks.
fn main() {
let mut x = 5i32;
let ptr1: *mut i32 = &mut x;
let ptr2: *mut i32 = &mut x; // second mutable ref through raw pointer
unsafe {
*ptr1 = 10; // use ptr1 — invalidates ptr2 under Stacked Borrows
*ptr2 = 20; // ← UB: ptr2 was invalidated when ptr1 was used
println!("{}", x);
}
}
// Miri output:
// error: Undefined Behavior: attempting a write access using <tag> at ...
// but that tag does not exist in the borrow stack for this location
// this error occurs as part of an access at ...Tree Borrows: a newer, less restrictive aliasing model that allows patterns Stacked Borrows rejects, while still catching real bugs. Enable it with MIRIFLAGS="-Zmiri-tree-borrows". Major crates like Vec and HashMap pass under Tree Borrows but were flagged by Stacked Borrows.
Part 5: Using Miri on Wrapper Code
Miri cannot execute across the FFI boundary into real C libraries, but it can still check every line of your Rust wrapper code up to the boundary. The trick is to mock the C functions in tests, replacing them with safe Rust implementations that have the same signature. Miri then runs through the full logic of your wrapper.
Testing a Wrapper with Mocked FFI
// In tests we swap the real FFI for a pure-Rust mock
// so Miri can run end-to-end without hitting the C boundary.
#[cfg(test)]
mod ffi {
// Mirror the real ffi types
#[repr(u32)]
#[derive(PartialEq)]
pub enum snappy_status { SNAPPY_OK = 0, SNAPPY_INVALID_INPUT = 1 }
// Pure-Rust mock: just copies the data through
pub unsafe fn snappy_max_compressed_length(n: usize) -> usize { n * 2 }
pub unsafe fn snappy_compress(
input: *const i8, input_len: usize,
output: *mut i8, output_len: *mut usize,
) -> snappy_status {
let src = std::slice::from_raw_parts(input as *const u8, input_len);
let dst = std::slice::from_raw_parts_mut(output as *mut u8, input_len);
dst.copy_from_slice(src);
*output_len = input_len;
snappy_status::SNAPPY_OK
}
}
#[cfg(not(test))]
mod ffi {
include!(concat!(env!("OUT_DIR"), "/bindings.rs"));
}
#[cfg(test)]
mod tests {
use super::*;
#[test] // cargo +nightly miri test runs this under Miri
fn compress_does_not_leak_or_alias() {
let data = b"hello";
let compressed = compress(data);
assert_eq!(compressed.len(), data.len());
}
}Part 6: Recommended Workflow
- Write the C header (or obtain it from the library). Create a thin
wrapper.hthat#includes only the parts you need. - Add bindgen to build.rs with
allowlist_functionto keep bindings minimal. Commit thebuild.rs, not the generated file (it is always regenerated from source). - Write the safe wrapper in
src/lib.rs. Everyunsafeblock must have a// SAFETY:comment. - Write unit tests with mocked FFI so Miri can run them. Cover edge cases: empty slices, null-like inputs, maximum sizes.
- Run Miri in CI on every pull request:
cargo +nightly miri test. Treat a Miri failure as a blocker, just like a compilation error. - Enable additional Miri flags for thorough checking:
-Zmiri-strict-provenanceand-Zmiri-symbolic-alignment-check.
bindgen Checklist
- ✓Use
build.rs— never commit generated bindings - ✓Use
allowlist_functionto limit surface area - ✓Add
CargoCallbacksso headers trigger rebuilds - ✓Test on all target platforms — sizes and alignment may differ
- ✓Mark generated module
#[allow(clippy::all)]
Miri Checklist
- ✓Run
cargo +nightly miri testin CI - ✓Mock FFI functions so Miri can run end-to-end
- ✓Add
// SAFETY:comments on everyunsafeblock - ✓Try
-Zmiri-tree-borrowsif Stacked Borrows rejects valid code - ✓Use
-Zmiri-strict-provenanceto catch pointer-to-integer casts
bindgen and Miri are two sides of the same coin. bindgen lets you reach into the vast ecosystem of C libraries without writing tedious boilerplate by hand. Miri catches the subtle memory bugs that only show up in unsafe code—the kind that the borrow checker is not designed to see. Together they give you a workflow where interoperating with C is manageable and your unsafe code is as rigorously tested as your safe code. Use bindgen in build.rs, wrap it safely, mock the boundary, and run Miri on every PR. That is the production-quality FFI workflow. Happy hacking!