Observability & Production Rust: Structured Logging, Metrics, Health Checks, Config, and Supply-Chain Auditing
Writing code that compiles and passes tests is the start, not the finish. A program that runs in production has to answer three questions on a bad day at 3 AM: What is it doing right now? Is it healthy? Can I trust its dependencies?
This guide covers the full production-readiness stack for a Rust service: structured logging and distributed tracing with the tracing crate, metrics with the metrics crate, health-check endpoints, environment-based configuration, and supply-chain auditing with cargo-deny.
Every concept uses a real-world analogy first. No prior observability experience needed.
Part 1: What Is Observability and Why Does It Matter?
The Aircraft Black Box Analogy
When a plane crashes, investigators recover the black box. It recorded every sensor reading, cockpit conversation, and system state during the flight. Without it, cause of death is guesswork. With it, every decision can be reconstructed.
Observability is building a black box into your software — while it is still flying. It means your service continuously emits structured information about its own state so that you, or an automated system, can understand what happened, what is happening now, and why performance is degrading — without modifying and redeploying the code.
The Three Pillars
Logs
Time-stamped records of discrete events. “Request received”, “database query failed”. Best for debugging specific incidents.
Metrics
Numeric measurements sampled over time. Request count, latency p99, memory usage. Best for dashboards and alerting.
Traces
The journey of a single request across services and code paths. Best for understanding latency breakdown in distributed systems.
Part 2: Structured Logging & Tracing with the tracing Crate
The Package Tracking Analogy
Plain println! logging is like shouting updates down a hallway: “Package arrived!” — but you cannot tell which package, which customer, or which delivery route without reading the message carefully.
Structured logging is like a parcel tracking system: every event has a tracking_id, a customer_id, a status, and a timestamp — as machine-readable fields, not buried in a sentence. You can filter, aggregate, and alert on fields instantly. tracing goes further: it groups events into spans — a span is the entire journey of one parcel, from warehouse to door.
Key Concepts: Events vs Spans
Event
A single point-in-time occurrence with attached fields. Replaces println! and log::info!.
Macros: tracing::error!, warn!, info!, debug!, trace!
Span
A named period with a start and end. All events emitted inside a span automatically inherit its context — including its trace_id.
Macro: tracing::instrument attribute — instruments an entire function as a span.
[dependencies]
tracing = "0.1" # the instrumentation API
tracing-subscriber = { version = "0.3", features = ["env-filter", "json"] }
# tracing-subscriber is the "backend" that formats and outputs spans/events
# env-filter lets you change log level at runtime via RUST_LOG env var
# json outputs machine-readable JSON lines instead of pretty textuse tracing_subscriber::{layer::SubscriberExt, util::SubscriberInitExt, EnvFilter};
fn main() {
// Read log level from RUST_LOG env var, default to "info"
// RUST_LOG=debug cargo run → shows debug + info + warn + error
// RUST_LOG=my_crate=trace → trace only for your crate
tracing_subscriber::registry()
.with(EnvFilter::try_from_default_env()
.unwrap_or_else(|_| EnvFilter::new("info")))
.with(tracing_subscriber::fmt::layer()
.json() // structured JSON — use .pretty() for dev
.with_target(true) // include module path in every event
.with_thread_ids(true)) // include OS thread id
.init();
tracing::info!("Service started");
run_server();
}Instrumenting Functions with #[instrument]
The #[tracing::instrument] attribute wraps an entire function in a span. Every log event inside that function — including in functions it calls — automatically carries the span's name and fields. This is how you get a trace_id attached to every log line for a given request.
use tracing::{instrument, info, warn, error};
#[instrument(
name = "handle_request",
fields(
request_id = %req.id, // % uses Display; ? uses Debug
path = %req.path,
method = %req.method,
),
skip(db), // don't log the db connection — it's noisy
)]
async fn handle_request(req: Request, db: &Database) -> Response {
info!("Processing request"); // inherits request_id, path, method
match db.fetch_user(req.user_id).await {
Ok(user) => {
info!(user_id = %user.id, "User found");
Response::ok(user)
}
Err(e) => {
error!(error = %e, "Database fetch failed");
Response::internal_error()
}
}
// span ends here — duration is automatically recorded
}
// Sample JSON output for one request:
// {"timestamp":"2026-06-01T10:00:00Z","level":"INFO","target":"my_service",
// "span":{"name":"handle_request","request_id":"abc123","path":"/users","method":"GET"},
// "message":"Processing request"}
// {"timestamp":"2026-06-01T10:00:00.003Z","level":"INFO","target":"my_service",
// "span":{"name":"handle_request","request_id":"abc123","path":"/users","method":"GET"},
// "fields":{"user_id":"42"},"message":"User found"}Why structured fields beat string interpolation: with info!("User {} found", id) the id is buried in a string. A log aggregator (Datadog, Loki, CloudWatch) cannot filter on it. With info!(user_id = id, "User found") you get a queryable field: user_id = 42.
Distributed Tracing: Propagating Trace IDs Across Services
When service A calls service B, you want to link both traces under one trace_id so the full journey of one user request is visible end-to-end. This is done by injecting the trace context into outbound HTTP headers and extracting it on the receiving side.
# Cargo.toml
# [dependencies]
# tracing-opentelemetry = "0.24"
# opentelemetry = "0.23"
# opentelemetry_sdk = { version = "0.23", features = ["rt-tokio"] }
# opentelemetry-otlp = { version = "0.16", features = ["grpc-tonic"] }
// main.rs — send traces to an OpenTelemetry collector (e.g. Jaeger, Tempo)
use opentelemetry_otlp::WithExportConfig;
use tracing_opentelemetry::OpenTelemetryLayer;
fn init_tracing() {
let tracer = opentelemetry_otlp::new_pipeline()
.tracing()
.with_exporter(
opentelemetry_otlp::new_exporter()
.tonic()
.with_endpoint("http://jaeger:4317"), // OTLP gRPC endpoint
)
.install_batch(opentelemetry_sdk::runtime::Tokio)
.unwrap();
tracing_subscriber::registry()
.with(EnvFilter::new("info"))
.with(tracing_subscriber::fmt::layer())
.with(OpenTelemetryLayer::new(tracer)) // send spans to Jaeger
.init();
}Part 3: Metrics with the metrics Crate
The Car Dashboard Analogy
A log is like the engine warning light blinking once — a specific event at a point in time. A metric is like the temperature gauge — a continuously sampled number you glance at in real time. You alert on the gauge reaching 110°C; you read the log only after the engine breaks down to understand why. Both tell different parts of the story and you need both.
Metric Types
Counter
- Only goes up
- Total requests served
- Total errors seen
- Bytes written to disk
- API:
counter!
Gauge
- Can go up or down
- Current memory usage
- Active connections
- Queue depth
- API:
gauge!
Histogram
- Records distribution
- Request latency
- File sizes processed
- p50 / p95 / p99
- API:
histogram!
The metrics Facade Pattern
The metrics crate is a facade — a stable API your application code calls, decoupled from the backend (Prometheus, StatsD, in-memory). You swap exporters without changing a line of business logic.
[dependencies]
metrics = "0.23"
metrics-exporter-prometheus = "0.15" # exposes /metrics endpoint for Prometheususe metrics::{counter, gauge, histogram};
use metrics_exporter_prometheus::PrometheusBuilder;
use std::time::Instant;
fn init_metrics() {
PrometheusBuilder::new()
.with_http_listener(([0, 0, 0, 0], 9000)) // /metrics on port 9000
.install()
.expect("failed to install Prometheus exporter");
}
async fn handle_request(req: Request, db: &Database) -> Response {
let start = Instant::now();
// Increment total request counter with a label
counter!("http_requests_total", "method" => req.method.clone());
let response = process(req, db).await;
// Record latency as histogram in seconds
let elapsed = start.elapsed().as_secs_f64();
histogram!("http_request_duration_seconds",
"status" => response.status.to_string(),
elapsed);
if response.is_error() {
counter!("http_errors_total", "status" => response.status.to_string());
}
response
}
fn update_active_connections(count: i64) {
gauge!("active_connections", count as f64);
}# HELP http_requests_total Total HTTP requests received
# TYPE http_requests_total counter
http_requests_total{method="GET"} 1043
http_requests_total{method="POST"} 217
# HELP http_request_duration_seconds Request latency distribution
# TYPE http_request_duration_seconds histogram
http_request_duration_seconds_bucket{status="200",le="0.005"} 890
http_request_duration_seconds_bucket{status="200",le="0.01"} 970
http_request_duration_seconds_bucket{status="200",le="0.025"} 1041
# HELP active_connections Current open connections
# TYPE active_connections gauge
active_connections 47Part 4: Health Check Endpoints
The Hospital Triage Analogy
Before a patient is admitted to surgery, the triage nurse checks three things: pulse, breathing, consciousness — quickly. Not a full workup, just “is this person alive and stable?” A /healthz endpoint is your service's triage: Kubernetes, load balancers, and monitoring systems call it every few seconds to decide whether to send traffic. It must be fast and it must tell the truth.
Liveness vs Readiness vs Startup
/healthz (Liveness)
“Is the process alive?”
Fails only if the process should be restarted. Return 200 if basic internal state is sane. No database calls.
/readyz (Readiness)
“Can this instance serve traffic?”
Returns 503 if the DB connection pool is saturated or a dependent service is unreachable. Remove from load balancer when failing.
/startupz
“Has the service finished initialising?”
Kubernetes uses this so slow-starting services are not killed before they are ready. Returns 200 once startup is complete.
Implementing Health Endpoints with Axum
use axum::{
routing::get, Router, Json, extract::State, http::StatusCode,
};
use serde::Serialize;
use std::sync::Arc;
use tokio::sync::RwLock;
#[derive(Clone)]
struct AppState {
db: Arc<DbPool>,
ready: Arc<RwLock<bool>>,
}
#[derive(Serialize)]
struct HealthResponse {
status: &'static str,
version: &'static str,
}
// Liveness — just confirm the process loop is alive
async fn healthz() -> (StatusCode, Json<HealthResponse>) {
(StatusCode::OK, Json(HealthResponse {
status: "ok",
version: env!("CARGO_PKG_VERSION"),
}))
}
// Readiness — check DB reachability
async fn readyz(State(state): State<AppState>) -> StatusCode {
if !*state.ready.read().await {
return StatusCode::SERVICE_UNAVAILABLE;
}
match state.db.ping().await {
Ok(_) => StatusCode::OK,
Err(e) => {
tracing::warn!(error = %e, "DB ping failed in readyz");
StatusCode::SERVICE_UNAVAILABLE
}
}
}
fn router(state: AppState) -> Router {
Router::new()
.route("/healthz", get(healthz))
.route("/readyz", get(readyz))
.route("/metrics", get(prometheus_handler)) // see Part 3
.with_state(state)
}Part 5: Environment-Based Configuration
The Chef Recipe Card Analogy
A recipe card does not say “add exactly 200g of flour from the bin on shelf 3 of warehouse C”. It says “add 200g of flour” and the kitchen has flour wherever it needs to. Config is the same idea: your binary should say “connect to the database at DATABASE_URL” and the environment — dev laptop, staging server, production cluster — provides the actual value. Never hard-code environment-specific values in source code.
Layered Config with the config crate
Production services typically layer config sources with a clear priority order: environment variables override file values which override defaults. The config crate handles this in a few lines.
[dependencies]
config = "0.14"
serde = { version = "1", features = ["derive"] }
dotenvy = "0.15" # load .env file in dev (never in production)
secrecy = "0.8" # wrap secrets so they don't appear in Debug output[server]
host = "0.0.0.0"
port = 8080
[database]
max_connections = 10
connect_timeout_secs = 5
[log]
level = "info"use config::{Config, ConfigError, Environment, File};
use secrecy::SecretString;
use serde::Deserialize;
#[derive(Debug, Deserialize)]
pub struct ServerConfig {
pub host: String,
pub port: u16,
}
#[derive(Debug, Deserialize)]
pub struct DatabaseConfig {
pub url: SecretString, // won't appear in logs/debug output
pub max_connections: u32,
pub connect_timeout_secs: u64,
}
#[derive(Debug, Deserialize)]
pub struct AppConfig {
pub server: ServerConfig,
pub database: DatabaseConfig,
}
pub fn load() -> Result<AppConfig, ConfigError> {
Config::builder()
// 1. Start with defaults from file
.add_source(File::with_name("config/default"))
// 2. Override with environment-specific file (optional)
.add_source(File::with_name("config/production").required(false))
// 3. Override with environment variables — APP_DATABASE__URL=...
// double underscore __ maps to nested keys (database.url)
.add_source(Environment::with_prefix("APP").separator("__"))
.build()?
.try_deserialize()
}
// Usage:
// APP_DATABASE__URL=postgres://... APP_SERVER__PORT=9090 cargo runSecretString from the secrecy crate: wraps a String so its Debug impl prints [REDACTED]. Database URLs, API keys, and tokens should always use this type to prevent accidental leakage in logs or panic messages.
Part 6: Supply-Chain Auditing with cargo-deny
The Building Inspection Analogy
Before a skyscraper is handed over, inspectors check not just the architect's design but every supplier: was the steel certified? Was the concrete mixed to spec? Were any subcontractors on a sanctions list? Your Cargo.lock has dozens or hundreds of transitive dependencies. Any of them could have a known CVE, a forbidden license, or multiple conflicting versions that hide bugs. cargo-deny is the inspector.
What cargo-deny Checks
Security Advisories
Queries the RustSec Advisory Database for known CVEs in your dependency tree. Fails the build if a vulnerable version is found.
License Compliance
Ensures all transitive dependencies use licenses compatible with your project (e.g. only MIT/Apache-2.0, no GPL). Critical for commercial products.
Duplicate Versions
Warns when the same crate is compiled twice at different versions, increasing binary size and sometimes hiding type-mismatch bugs.
Banned Crates
Block specific crates by name — for example, ban unmaintained crates you want to replace, or deprecated alternatives to a crate you have standardised on.
Setup and Configuration
cargo install cargo-deny
# Generate a starter deny.toml at the project root
cargo deny init
# Run all checks
cargo deny check
# Run only the security advisory check
cargo deny check advisories[advisories]
# Fail on any unpatched vulnerability unless explicitly ignored
ignore = [] # list specific advisory IDs here to suppress
[licenses]
# Only allow these licenses in the dependency tree
allow = ["MIT", "Apache-2.0", "BSD-3-Clause", "Unicode-DFS-2016"]
# How to handle crates whose license is unclear
unlicensed = "deny"
copyleft = "deny" # reject GPL/LGPL
[bans]
multiple-versions = "warn" # warn on duplicate crate versions
# Ban specific crates by name
deny = [
{ name = "openssl" }, # we standardise on rustls
]
# Allow specific duplicates you cannot resolve yet
skip = [
{ name = "windows-sys" }, # common transitive version conflict
]
[sources]
# Only allow crates from crates.io and specific git repos
unknown-registry = "deny"
unknown-git = "deny"Adding cargo-deny to CI
name: CI
on: [push, pull_request]
jobs:
security-audit:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
# Runs cargo deny check against the advisory database
# Fails the PR if a CVE is found in the dependency tree
- name: Security audit
uses: EmbarkStudios/cargo-deny-action@v1
with:
command: check advisories licenses bans sources
build-and-test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Build
run: cargo build --release
- name: Test
run: cargo test
- name: Clippy
run: cargo clippy -- -D warningsPart 7: The Complete Production-Ready Workflow
Wiring It All Together
mod config;
mod metrics;
mod routes;
mod db;
#[tokio::main]
async fn main() -> anyhow::Result<()> {
// 1. Load .env (dev only — in prod these are real env vars)
dotenvy::dotenv().ok();
// 2. Load typed config (defaults → file → env vars)
let cfg = config::load().expect("Failed to load config");
// 3. Initialise structured logging with env-controlled level
tracing_subscriber::registry()
.with(tracing_subscriber::EnvFilter::new(&cfg.log.level))
.with(tracing_subscriber::fmt::layer().json())
.init();
tracing::info!(
version = env!("CARGO_PKG_VERSION"),
port = cfg.server.port,
"Service starting"
);
// 4. Initialise Prometheus metrics exporter
metrics::init_prometheus(cfg.metrics.port);
// 5. Connect to database
let db = db::connect(&cfg.database).await?;
// 6. Build shared state
use std::sync::Arc;
use tokio::sync::RwLock;
let state = routes::AppState {
db: Arc::new(db),
ready: Arc::new(RwLock::new(true)),
};
// 7. Start HTTP server — includes /healthz /readyz /metrics
let addr = format!("{}:{}", cfg.server.host, cfg.server.port);
tracing::info!(address = %addr, "Listening");
let listener = tokio::net::TcpListener::bind(&addr).await?;
axum::serve(listener, routes::router(state)).await?;
Ok(())
}Local Development Stack
services:
prometheus:
image: prom/prometheus:latest
ports: ["9090:9090"]
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml
grafana:
image: grafana/grafana:latest
ports: ["3001:3000"]
depends_on: [prometheus]
jaeger:
image: jaegertracing/all-in-one:latest
ports:
- "16686:16686" # Jaeger UI
- "4317:4317" # OTLP gRPC
# prometheus.yml — tell Prometheus to scrape your service
# scrape_configs:
# - job_name: my_service
# static_configs:
# - targets: ['host.docker.internal:9000']
# Then: cargo run → open http://localhost:16686 (Jaeger traces)
# open http://localhost:3001 (Grafana dashboards)Step-by-Step Checklist
- Add
tracing+tracing-subscriber. Init subscriber at startup withEnvFilterand JSON output. - Replace every
println!with the appropriate tracing macro. Add#[instrument]to every request handler. - Add
metrics+metrics-exporter-prometheus. Emit a counter, gauge, and histogram for each handler. - Add
/healthz(liveness) and/readyz(readiness) routes. Confirm readiness probe checks your DB. - Model all config as a typed
AppConfigstruct. Load it with theconfigcrate. Wrap secrets inSecretString. - Run
cargo deny init, configuredeny.toml, and addcargo deny checkas a CI gate.
Production Checklist
- ✓JSON logs in prod, pretty logs in dev (
RUST_LOGdriven) - ✓Every request handler instrumented with
#[instrument] - ✓
trace_idpresent on every log line (via span context) - ✓Latency histogram, request counter, and error counter emitted
- ✓
/healthzand/readyzendpoints wired to K8s probes - ✓
SecretStringfor all secrets in config - ✓
cargo deny checkin CI — blocks on CVE or license violation
Common Mistakes
- ✗Using
println!instead of tracing macros — loses structure and level filtering - ✗Calling the DB inside
/healthz— makes liveness depend on DB availability; it should not - ✗Hard-coding secrets in source or config files committed to git
- ✗Ignoring
cargo denywarnings — vulnerabilities accumulate silently - ✗Registering WRITABLE permanently on idle sockets — a mio pitfall that causes busy-looping
A production Rust service is not just fast code — it is an honest system. It tells you what it is doing via structured logs, how it is performing via metrics, whether it can serve traffic via health endpoints, and what it is built on via supply-chain audits. The tracing, metrics, and cargo-deny crates make this achievable in an afternoon. Add them at the start of a project — retrofitting observability into an existing codebase is always harder than building it in from day one. Happy hacking!