Learn Rust Series (#65) - Mutex, RwLock, and Handling Lock Poisoning
Learn Rust Series (#65) - Mutex, RwLock, and Handling Lock Poisoning
What will I learn
- You will learn how
Mutexgrants exclusive access and how theMutexGuardreleases the lock viaDrop; - how
RwLockallows many concurrent readers or a single writer, and when it beats aMutex; - what lock poisoning is: why a panic while holding a lock marks it poisoned, and how to recover;
- how to minimise the time a lock is held so it does not become a bottleneck;
- the golden rule of lock ordering that prevents deadlocks.
Requirements
- A working modern computer running macOS, Windows or Ubuntu;
- An installed Rust toolchain (via rustup, from rustup.rs);
- The previous sixty-four episodes, especially
Arc, threads, and interior mutability; - The ambition to learn systems programming from the ground up.
Difficulty
- Intermediate
Curriculum (of the Learn Rust Series):
- Learn Rust Series (#1) - Introduction to Rust
- Learn Rust Series (#2) - Variables, Types, Functions
- Learn Rust Series (#3) - Ownership & Borrowing
- Learn Rust Series (#4) - Control Flow & Pattern Matching
- Learn Rust Series (#5) - Structs & Enums
- Learn Rust Series (#6) - Error Handling
- Learn Rust Series (#7) - Collections
- Learn Rust Series (#8) - Traits & Generics
- Learn Rust Series (#9) - Modules & Crates
- Learn Rust Series (#10) - Lifetimes
- Learn Rust Series (#11) - Closures & the Iterator Trait
- Learn Rust Series (#12) - Smart Pointers: Box, Rc & RefCell
- Learn Rust Series (#13) - Concurrency: Threads, Channels, Arc & Mutex
- Learn Rust Series (#14) - Mini Project: A Command-Line To-Do App
- Learn Rust Series (#15) - Trait Objects & Dynamic Dispatch
- Learn Rust Series (#16) - Static vs Dynamic Dispatch
- Learn Rust Series (#17) - Associated Types vs Generic Parameters
- Learn Rust Series (#18) - Operator Overloading with std::ops
- Learn Rust Series (#19) - Deref, DerefMut & Deref Coercion
- Learn Rust Series (#20) - Drop & Deterministic Destruction (RAII)
- Learn Rust Series (#21) - From, Into, TryFrom & Idiomatic Conversions
- Learn Rust Series (#22) - Deriving Common Traits
- Learn Rust Series (#23) - The Orphan Rule & Trait Coherence
- Learn Rust Series (#24) - Blanket Implementations & the Newtype Pattern
- Learn Rust Series (#25) - Marker Traits: Sized, Send, Sync & Copy
- Learn Rust Series (#26) - Const Generics: Types That Depend on Values
- Learn Rust Series (#27) - Generic Associated Types & Lending Iterators
- Learn Rust Series (#28) - Sealed Traits & Designing Stable APIs
- Learn Rust Series (#29) - Typestate Programming: State Machines in the Type System
- Learn Rust Series (#30) - Mini Project: A Generic Units-of-Measure Library
- Learn Rust Series (#31) - Move Semantics Deep Dive
- Learn Rust Series (#32) - Interior Mutability: Cell & RefCell
- Learn Rust Series (#33) - Rc Internals: Reference Counting & Shared Ownership
- Learn Rust Series (#34) - Arc: Thread-Safe Reference Counting & Its Cost
- Learn Rust Series (#35) - Weak References & Breaking Reference Cycles
- Learn Rust Series (#36) - Cow: Clone-on-Write for Borrow-or-Own APIs
- Learn Rust Series (#37) - Pin & Self-Referential Structs
- Learn Rust Series (#38) - PhantomData, Zero-Sized Types & Marker Lifetimes
- Learn Rust Series (#39) - Variance: Covariance, Contravariance & Why It Matters
- Learn Rust Series (#40) - Arena & Bump Allocation Patterns
- Learn Rust Series (#41) - Building Your Own Smart Pointer
- Learn Rust Series (#42) - Drop Order, the Drop Check & Leak Safety
- Learn Rust Series (#43) - std::mem: swap, replace, take & forget
- Learn Rust Series (#44) - Higher-Ranked Trait Bounds & Lifetime Elision
- Learn Rust Series (#45) - Mini Project: A Doubly-Linked List, Safe then Unsafe
- Learn Rust Series (#46) - Result Combinators: map, map_err, and_then, ok_or
- Learn Rust Series (#47) - Option Combinators & Null-Free Programming
- Learn Rust Series (#48) - Custom Error Types & the std::error::Error Trait
- Learn Rust Series (#49) - thiserror: Ergonomic Library Errors
- Learn Rust Series (#50) - anyhow: Flexible Application-Level Errors & Context
- Learn Rust Series (#51) - Panics, Unwinding, abort, and catch_unwind
- Learn Rust Series (#52) - Testing: Unit Tests, Integration Tests, and Doctests
- Learn Rust Series (#53) - Property-Based Testing with proptest
- Learn Rust Series (#54) - Fuzzing with cargo-fuzz and libFuzzer
- Learn Rust Series (#55) - Benchmarking with Criterion and Reading the Numbers
- Learn Rust Series (#56) - Cargo Workspaces and Multi-Crate Projects
- Learn Rust Series (#57) - Feature Flags and Conditional Compilation (cfg)
- Learn Rust Series (#58) - Build Scripts (build.rs) and Generating Code at Build Time
- Learn Rust Series (#59) - Clippy, rustfmt, and Writing Idiomatic Rust
- Learn Rust Series (#60) - Mini Project: A Fully Tested, Documented, Published-Ready CSV Toolkit Crate
- Learn Rust Series (#61) - Send and Sync: The Traits Behind Fearless Concurrency
- Learn Rust Series (#62) - Scoped Threads: Borrowing Local Data Across Threads
- Learn Rust Series (#63) - Channels: mpsc, Ownership Transfer, and Backpressure
- Learn Rust Series (#64) - Crossbeam: Faster Channels and Scoped Concurrency
- Learn Rust Series (#65) - Mutex, RwLock, and Handling Lock Poisoning (this post)
Learn Rust Series (#65) - Mutex, RwLock, and Handling Lock Poisoning
The last handful of episodes have all been about the same big idea approached from different angles: how do you let more than one thread work at the same time without the whole thing turning into a debugging nightmare. Channels (episodes 63 and 64) answered that by moving data between threads -- one thread owns a value, hands it off, and never touches it again. That is a wonderful model, and when it fits, it is the one I reach for first. But it does not fit everything. Sometimes several threads genuinely need to look at, and change, the same piece of state -- a shared counter, a cache, a configuration that any of them may update. For that you need not message passing but shared mutable state, and shared mutable state across threads is exactly the thing that makes concurrency dangerous in almost every other language.
Rust's answer is the lock, and specifically two of them: Mutex and RwLock. A Mutex (short for mutual exclusion) guarantees that only one thread touches the protected value at any instant, and -- this is the beautiful part -- Rust ties that guarantee straight into the type system. You cannot reach the data without going through lock(), and the lock releases itself automatically the moment its guard goes out of scope, courtesy of the Drop trait we studied way back in episode 20. There is no "I forgot to unlock" bug possible, because there is no manual unlock to forget. On top of that Rust adds one thing most languages simply do not have -- poisoning -- which turns "a thread crashed halfway through an update and left the data corrupted" from a silent, invisible catastrophe into a loud, visible Err you are forced to deal with ;-)
First though, as always, last episode's homework.
Solutions to Episode 64 Exercises
Episode 64 was crossbeam, and all three exercises pushed on the std building blocks underneath it: a shared-consumer worker pool, a hand-rolled poll-both-channels loop, and swapping the old crossbeam::scope for the modern std::thread::scope.
Exercise 1 asked you to share a std mpsc::Receiver among two worker threads via Arc<Mutex<Receiver>>, distribute six numbered jobs between them, and have each worker report how many jobs it personally handled:
use std::sync::{mpsc, Arc, Mutex};
use std::thread;
fn main() {
let (tx, rx) = mpsc::channel::<i32>();
let rx = Arc::new(Mutex::new(rx)); // one receiver, shared behind a lock
let mut workers = Vec::new();
for w in 0..2 {
let rx = Arc::clone(&rx);
workers.push(thread::spawn(move || {
let mut handled = 0;
while rx.lock().unwrap().recv().is_ok() { handled += 1; }
(w, handled)
}));
}
for job in 0..6 { tx.send(job).unwrap(); }
drop(tx); // closes the channel so every recv() eventually errors
let totals: Vec<(i32, i32)> = workers.into_iter().map(|h| h.join().unwrap()).collect();
let sum: i32 = totals.iter().map(|(_, n)| n).sum();
println!("handled {sum} jobs across {} workers", totals.len()); // handled 6 jobs across 2 workers
}
The key insight is that the Mutex serialises only the receiving step -- whichever worker grabs the lock first pulls the next job and immediately releases it, so the two threads take turns pulling but could process in parallel. The drop(tx) is what eventually lets both while loops end. Notice we are already using today's tool a full episode early; homework has a way of foreshadowing.
Exercise 2 wanted you to poll two std channels with try_recv in a loop, print whichever value arrives first, and stop once both channels have delivered exactly one value each:
use std::sync::mpsc;
use std::thread;
use std::time::Duration;
fn main() {
let (tx_a, rx_a) = mpsc::channel::<i32>();
let (tx_b, rx_b) = mpsc::channel::<i32>();
thread::spawn(move || tx_a.send(1).unwrap());
thread::spawn(move || tx_b.send(2).unwrap());
let mut seen = 0;
while seen < 2 {
if let Ok(v) = rx_a.try_recv() { println!("from a: {v}"); seen += 1; }
if let Ok(v) = rx_b.try_recv() { println!("from b: {v}"); seen += 1; }
thread::sleep(Duration::from_millis(1)); // avoid a hot spin
}
}
try_recv is the non-blocking sibling of recv -- it returns immediately with an Err if nothing is waiting, which is exactly what lets us peek at both channels in one pass. The tiny sleep keeps us from pinning a whole CPU core spinning. This is the crude version of the select! we admired last time.
Exercise 3 was the modernisation task: rewrite a snippet that used the old crossbeam::scope so it uses std::thread::scope instead, borrowing a local Vec across two threads and returning a value from the scope:
use std::thread;
fn main() {
let data = vec![10, 20, 30, 40];
let total = thread::scope(|s| {
let front = s.spawn(|| data[..2].iter().sum::<i32>()); // both borrow `data`
let back = s.spawn(|| data[2..].iter().sum::<i32>());
front.join().unwrap() + back.join().unwrap()
});
println!("total: {total}"); // total: 100
println!("still here: {}", data.len()); // 4 -- data was only borrowed
}
Because thread::scope (episode 62) guarantees every spawned thread finishes before the scope returns, both closures are allowed to borrow data rather than needing an Arc or a move-and-clone. The value the scope produces flows straight out as total. As I said last time, swapping crossbeam::scope for thread::scope is a nearly mechanical modernisation. Right, homework cleared -- now, locks.
Mutex: exclusive access
A Mutex<T> wraps a value of type T and hides it. The only way to reach the inner value is to call lock(), which blocks the calling thread until the lock is free and then returns a MutexGuard. That guard is a smart pointer (episode 12) -- it derefs to &T or &mut T -- and, crucially, when the guard is dropped, the lock releases:
use std::sync::Mutex;
fn main() {
let m = Mutex::new(5);
{
let mut guard = m.lock().unwrap(); // guard derefs to the inner i32
*guard += 10;
} // guard dropped here, lock released
println!("{}", *m.lock().unwrap()); // 15
}
Look at what the type system has just done for you. Because access requires the guard, and the guard is the lock being held, it is literally impossible to read or write the protected value without holding the lock. In C or C++ the mutex and the data it protects are two separate things held together by nothing more than a comment and the programmer's good intentions -- forget to lock, and the compiler says nothing. In Rust the lock and the data are one object, and forgetting to lock is not a bug you can write. That is the same "make invalid states unrepresentable" philosophy we have met again and again in this series.
The .unwrap() on lock() is there because lock() returns a Result -- and why it can fail is the poisoning story we will get to shortly. For now, read .unwrap() as "give me the guard, and panic if the lock was poisoned".
Sharing a Mutex across threads
A Mutex on its own only gives you the exclusion; it does not give you shared ownership. To let several threads own the same mutex you wrap it in an Arc (episode 34), the thread-safe reference-counted pointer. The pairing Arc<Mutex<T>> is so common in Rust concurrency that it is practically an idiom -- Arc for "many owners", Mutex for "safe mutation". Here four workers each append a line to one shared log:
use std::sync::{Arc, Mutex};
use std::thread;
fn main() {
let log = Arc::new(Mutex::new(Vec::new()));
let mut handles = Vec::new();
for id in 0..4 {
let log = Arc::clone(&log); // clone the Arc, not the Vec -- cheap pointer bump
handles.push(thread::spawn(move || {
log.lock().unwrap().push(format!("worker {id} done"));
}));
}
for h in handles { h.join().unwrap(); }
println!("{} messages logged", log.lock().unwrap().len()); // 4
}
Trace the ownership carefully, because it is the whole point. Arc::clone does not copy the Vec -- it bumps a reference count and hands back another pointer to the same underlying Mutex<Vec<...>>. Every worker locks that one mutex, pushes its line while holding the lock, and releases it when the temporary guard on that line is dropped at the semicolon. The four pushes are therefore serialised: the Vec is never touched by two threads at once, which is exactly why this compiles at all. Remember episode 61's Send and Sync? Arc<Mutex<T>> is Send + Sync precisely because the Mutex provides the synchronisation that makes shared mutation safe -- the marker traits and the lock are two halves of the same guarantee.
RwLock: many readers or one writer
Now, a Mutex is a blunt instrument. It serialises every access, even two threads that only want to read and would never step on each other. When your data is read far more often than it is written -- think a configuration loaded once and consulted constantly, or a cache queried by dozens of threads -- that blanket exclusion is pure waste. RwLock<T> (read-write lock) is the sharper tool. It hands out two different kinds of guard: any number of threads may hold a read() guard simultaneously, but a write() guard is exclusive and waits for all readers to clear out first:
use std::sync::{Arc, RwLock};
use std::thread;
fn main() {
let config = Arc::new(RwLock::new(vec![1, 2, 3]));
thread::scope(|s| {
for _ in 0..3 {
let c = &config;
s.spawn(move || {
let r = c.read().unwrap(); // shared read: many at once
println!("reader sees {} items", r.len());
});
}
s.spawn(|| {
let mut w = config.write().unwrap(); // exclusive write: alone
w.push(4);
});
});
println!("final: {:?}", *config.read().unwrap()); // final: [1, 2, 3, 4]
}
The rule of thumb is straightforward. Reach for RwLock when reads vastly outnumber writes, because then all those concurrent readers proceed in parallel instead of queueing behind one another. Reach for Mutex when writes are frequent, because in that case the reader/writer bookkeeping is overhead you never cash in -- an RwLock that is constantly being write-locked is just a slower, more complicated Mutex. And a small warning worth carrying: the exact fairness (does a waiting writer block new readers, or can a steady stream of readers starve the writer forever?) depends on the operating system's implementation, so do not lean on a particular ordering. If starvation matters to your design, that is a signal to rethink the sharing, not to guess at the lock's mood.
Lock poisoning
Here is the feature that genuinely sets Rust's locks apart, and it follows directly from the panic-and-unwind machinery we studied in episode 51. Suppose a thread locks a mutex, starts mutating the protected value, and then panics halfway through -- leaving the data in a broken, half-updated state. In most languages the lock would simply be released by the unwinding, and the next thread would happily lock it and read the corrupted mess, none the wiser. Rust refuses to let that slide. When a thread panics while holding a lock, the lock is marked poisoned, and every subsequent lock() returns an Err instead of a guard:
use std::sync::{Arc, Mutex};
use std::thread;
fn main() {
let data = Arc::new(Mutex::new(vec![1, 2, 3]));
let d = Arc::clone(&data);
let _ = thread::spawn(move || {
let mut guard = d.lock().unwrap();
guard.push(4);
panic!("worker crashed while holding the lock"); // poisons the Mutex
}).join(); // join swallows the panic so main can continue
let is_poisoned = data.lock().is_err();
if is_poisoned {
println!("lock is poisoned: a holder panicked"); // this branch
} else {
println!("lock was clean");
}
}
The poisoning is not Rust being fussy for the sake of it -- it is Rust being honest. The whole reason we lock the data is to protect an invariant ("this Vec is always internally consistent"). If a thread panicked mid-update, that invariant might now be broken, and Rust would rather tell you loudly than let the corruption propagate silently. It is the concurrency equivalent of the Result type refusing to let you ignore an error. This is also, incidentally, why lock() returns a Result at all -- the Err variant is the poison signal.
Recovering from poison
Poisoning is a warning, not a death sentence. Sometimes you inspect the situation and conclude the data is perfectly fine, or you know how to repair it -- and Rust lets you take the value out anyway. The Err you get back is a PoisonError, and it carries the guard inside it; calling into_inner() on the poison error hands you that guard so you can carry on:
use std::sync::{Arc, Mutex};
use std::thread;
fn main() {
let data = Arc::new(Mutex::new(0));
let d = Arc::clone(&data);
let _ = thread::spawn(move || {
let mut g = d.lock().unwrap();
*g = 42;
panic!("crash"); // poisons the lock, but the write already happened
}).join();
let value = match data.lock() {
Ok(g) => *g,
Err(poisoned) => *poisoned.into_inner(), // take the value out of the PoisonError
};
println!("recovered value: {value}"); // recovered value: 42
}
The design here is elegant once you see it. Rust does not hide the data behind poisoning -- it makes you write an explicit line of code that says "yes, I know a holder panicked, and I have decided this value is still usable". That single conscious into_inner() is the difference between "I handled the corruption risk" and "I accidentally read garbage". If you genuinely do not care about poisoning -- say the protected value is a simple counter where a half-update cannot really corrupt anything -- a common shortcut is .lock().unwrap_or_else(|e| e.into_inner()), which grabs the guard whether or not the lock is poisoned. Use it deliberately, not reflexively.
Hold the lock as briefly as possible
Locks are about correctness, but they are also about performance, and the single most important performance habit is this: hold the lock for as short a time as you possibly can. A lock held during slow work is a bottleneck, because every other thread that wants the same lock is stuck waiting while you dawdle. The fix is a pattern you will use constantly -- grab what you need under the lock, release it immediately, then do the expensive part unlocked:
use std::sync::Mutex;
fn main() {
let m = Mutex::new(vec![1, 2, 3, 4, 5]);
let snapshot: Vec<i32> = {
let guard = m.lock().unwrap();
guard.clone() // copy out what we need...
}; // ...and release the lock HERE, before the slow work below
let sum: i32 = snapshot.iter().sum(); // computed without holding the lock
println!("sum computed off-lock: {sum}"); // 15
}
The little inner block { ... } is doing real work: it bounds the lifetime of guard so the lock releases at the closing brace, before we start iterating. Contrast that with the lazy version where you lock the mutex and then run your whole calculation with the guard still alive -- every other thread queues behind you the entire time. Cloning out a snapshot costs a small allocation, but it buys you a lock that is held for microseconds instead of milliseconds, and under contention that trade is almost always worth it. Nota bene: this is also where a subtle deadlock hides -- if the "slow work" itself tries to lock the same mutex again, a plain Mutex will block forever waiting for itself (Rust's std Mutex is not re-entrant). Releasing early sidesteps that entire class of bug.
The golden rule against deadlocks
The scariest failure mode of locks is the deadlock: thread A holds lock 1 and waits for lock 2, while thread B holds lock 2 and waits for lock 1. Neither can proceed, neither will ever give up its lock, and your program simply hangs -- no panic, no error, no output, just a silent freeze. The prevention is almost embarrassingly simple to state and absolute in effect: always acquire multiple locks in the same order, everywhere in the program. If every thread that needs both a and b locks a first and b second, the cycle that a deadlock requires can never form:
use std::sync::Mutex;
fn main() {
let a = Mutex::new(1);
let b = Mutex::new(2);
let sum = {
let ga = a.lock().unwrap(); // ALWAYS a before b -- in every single code path
let gb = b.lock().unwrap();
*ga + *gb
};
println!("{sum}"); // 3
}
The reason a consistent order works is worth internalising: a deadlock is a cycle in the "who is waiting for whom" graph, and if every thread grabs locks in the same global order, that graph can only ever be a straight line, never a loop. It is a purely structural guarantee. In a big codebase this discipline is enforced by convention and code review (some teams even assign each lock a numeric rank and assert you only ever acquire in increasing rank), which is a bit of bookkeeping but infinitely cheaper than a production hang at 3am.
Having said that, the very best defence against deadlocks is needing fewer locks in the first place. A single lock cannot deadlock against itself under a consistent-order rule, and a channel (which we spent the last two episodes on) sidesteps shared-lock ordering entirely by not sharing the data at all. So the priority order in my head is: prefer message passing, then a single lock, and only reach for a lattice of multiple locks when you truly must -- and when you must, write the ordering down and never, ever break it.
How Python and Go would frame this
A glance sideways sharpens the picture, as it usually does. Python has threading.Lock, and the idiomatic use is a with block, which is Python's answer to Rust's guard-drops-the-lock -- the lock releases when the block exits, even on an exception:
import threading
counter = 0
lock = threading.Lock()
def bump():
global counter
with lock: # acquired here, released when the block exits
counter += 1 # protected, but nothing FORCES you to hold the lock
threads = [threading.Thread(target=bump) for _ in range(1000)]
for t in threads: t.start()
for t in threads: t.join()
print(counter) # 1000
The with lock: block is genuinely nice and the auto-release is the same instinct as Rust's Drop. But look at what is missing: the lock and the counter are two unrelated objects. Nothing stops another function from reading or writing counter without taking the lock -- the protection lives in your discipline, not the type system. And if a thread dies mid-update inside the block, Python just releases the lock and moves on; there is no poisoning, no signal that the invariant might be broken. (The GIL papers over a lot of this for pure-Python code, but the structural point stands the moment you drop into C extensions or multiprocessing.)
Go leans the same way, with sync.Mutex and sync.RWMutex that map almost one-to-one onto Rust's two locks:
type Counter struct {
mu sync.Mutex
n int
}
func (c *Counter) Bump() {
c.mu.Lock()
defer c.mu.Unlock() // released when Bump returns -- the Go idiom
c.n++ // but the compiler won't stop you touching c.n unlocked
}
defer c.mu.Unlock() is Go's clever trick for "release no matter how we leave this function", and pairing the mutex with the field it guards in one struct is good Go style. Yet it is still style, not enforcement -- another method could touch c.n without locking and Go would compile it without complaint, leaving the data race to be caught (maybe) at runtime by the -race detector. Rust's move is to make the lock and the data literally the same object, so "touch the data without the lock" is not a program you can write. Same idea, three levels of enforcement: Python and Go trust you, Rust checks you.
Wrapping up
So there we have it: shared mutable state done safely. A Mutex<T> gives one thread at a time exclusive access, with the lock and the data fused into a single object so you cannot reach the value without locking, and the MutexGuard releasing the lock automatically on Drop. Wrap it in an Arc and many threads can share one protected value -- the Arc<Mutex<T>> idiom you will use for the rest of your Rust life. RwLock<T> refines the picture for read-heavy data, letting many readers run at once while writers still get exclusive access. And poisoning -- unique among mainstream languages -- turns a panic-while-locked from silent corruption into a visible Err you can inspect and, with into_inner, recover from.
The habits to carry forward are three. Hold every lock for the shortest possible time -- snapshot out and release before slow work. Acquire multiple locks in a single consistent order, always, to make deadlock structurally impossible. And prefer fewer locks (or a channel) over a tangle of them, because the lock you never take is the one that can never deadlock.
There is one thing I have been quietly leaning on this whole episode without ever opening it up, and it is bugging me a little. A Mutex has to somehow coordinate threads at the hardware level -- when two threads race to lock() the same mutex, what actually decides who wins, down at the level of the CPU? The answer is a tiny, fascinating primitive that sits underneath every lock in existence, and it is where we are headed next ;-)
Exercises
- Wrap a counter in
Arc<Mutex<i32>>and increment it from ten threads, then join them all and print the total (it should be exactly ten times whatever each thread adds). - Use an
RwLock<Vec<i32>>shared across athread::scopeto let three reader threads print the length while one writer thread pushes a value, then print the final contents. - Poison a
Mutexon purpose by panicking while holding its guard, catch the panic withjoin, then recover the inner value withinto_innerand print it.
Thanks for your time, and keep those locks short! ;-)