This content got low rating by people.

Learn Rust Series (#65) - Mutex, RwLock, and Handling Lock Poisoning

Words
3557
Reading
16 min
Listen
Play
16d

Learn Rust Series (#65) - Mutex, RwLock, and Handling Lock Poisoning

rust-banner.png

What will I learn

  • You will learn how Mutex grants exclusive access and how the MutexGuard releases the lock via Drop;
  • how RwLock allows many concurrent readers or a single writer, and when it beats a Mutex;
  • what lock poisoning is: why a panic while holding a lock marks it poisoned, and how to recover;
  • how to minimise the time a lock is held so it does not become a bottleneck;
  • the golden rule of lock ordering that prevents deadlocks.

Requirements

  • A working modern computer running macOS, Windows or Ubuntu;
  • An installed Rust toolchain (via rustup, from rustup.rs);
  • The previous sixty-four episodes, especially Arc, threads, and interior mutability;
  • The ambition to learn systems programming from the ground up.

Difficulty

  • Intermediate

Curriculum (of the Learn Rust Series):

Learn Rust Series (#65) - Mutex, RwLock, and Handling Lock Poisoning

The last handful of episodes have all been about the same big idea approached from different angles: how do you let more than one thread work at the same time without the whole thing turning into a debugging nightmare. Channels (episodes 63 and 64) answered that by moving data between threads -- one thread owns a value, hands it off, and never touches it again. That is a wonderful model, and when it fits, it is the one I reach for first. But it does not fit everything. Sometimes several threads genuinely need to look at, and change, the same piece of state -- a shared counter, a cache, a configuration that any of them may update. For that you need not message passing but shared mutable state, and shared mutable state across threads is exactly the thing that makes concurrency dangerous in almost every other language.

Rust's answer is the lock, and specifically two of them: Mutex and RwLock. A Mutex (short for mutual exclusion) guarantees that only one thread touches the protected value at any instant, and -- this is the beautiful part -- Rust ties that guarantee straight into the type system. You cannot reach the data without going through lock(), and the lock releases itself automatically the moment its guard goes out of scope, courtesy of the Drop trait we studied way back in episode 20. There is no "I forgot to unlock" bug possible, because there is no manual unlock to forget. On top of that Rust adds one thing most languages simply do not have -- poisoning -- which turns "a thread crashed halfway through an update and left the data corrupted" from a silent, invisible catastrophe into a loud, visible Err you are forced to deal with ;-)

First though, as always, last episode's homework.

Solutions to Episode 64 Exercises

Episode 64 was crossbeam, and all three exercises pushed on the std building blocks underneath it: a shared-consumer worker pool, a hand-rolled poll-both-channels loop, and swapping the old crossbeam::scope for the modern std::thread::scope.

Exercise 1 asked you to share a std mpsc::Receiver among two worker threads via Arc<Mutex<Receiver>>, distribute six numbered jobs between them, and have each worker report how many jobs it personally handled:

use std::sync::{mpsc, Arc, Mutex};
use std::thread;

fn main() {
    let (tx, rx) = mpsc::channel::<i32>();
    let rx = Arc::new(Mutex::new(rx)); // one receiver, shared behind a lock
    let mut workers = Vec::new();
    for w in 0..2 {
        let rx = Arc::clone(&rx);
        workers.push(thread::spawn(move || {
            let mut handled = 0;
            while rx.lock().unwrap().recv().is_ok() { handled += 1; }
            (w, handled)
        }));
    }
    for job in 0..6 { tx.send(job).unwrap(); }
    drop(tx); // closes the channel so every recv() eventually errors

    let totals: Vec<(i32, i32)> = workers.into_iter().map(|h| h.join().unwrap()).collect();
    let sum: i32 = totals.iter().map(|(_, n)| n).sum();
    println!("handled {sum} jobs across {} workers", totals.len()); // handled 6 jobs across 2 workers
}

The key insight is that the Mutex serialises only the receiving step -- whichever worker grabs the lock first pulls the next job and immediately releases it, so the two threads take turns pulling but could process in parallel. The drop(tx) is what eventually lets both while loops end. Notice we are already using today's tool a full episode early; homework has a way of foreshadowing.

Exercise 2 wanted you to poll two std channels with try_recv in a loop, print whichever value arrives first, and stop once both channels have delivered exactly one value each:

use std::sync::mpsc;
use std::thread;
use std::time::Duration;

fn main() {
    let (tx_a, rx_a) = mpsc::channel::<i32>();
    let (tx_b, rx_b) = mpsc::channel::<i32>();
    thread::spawn(move || tx_a.send(1).unwrap());
    thread::spawn(move || tx_b.send(2).unwrap());

    let mut seen = 0;
    while seen < 2 {
        if let Ok(v) = rx_a.try_recv() { println!("from a: {v}"); seen += 1; }
        if let Ok(v) = rx_b.try_recv() { println!("from b: {v}"); seen += 1; }
        thread::sleep(Duration::from_millis(1)); // avoid a hot spin
    }
}

try_recv is the non-blocking sibling of recv -- it returns immediately with an Err if nothing is waiting, which is exactly what lets us peek at both channels in one pass. The tiny sleep keeps us from pinning a whole CPU core spinning. This is the crude version of the select! we admired last time.

Exercise 3 was the modernisation task: rewrite a snippet that used the old crossbeam::scope so it uses std::thread::scope instead, borrowing a local Vec across two threads and returning a value from the scope:

use std::thread;

fn main() {
    let data = vec![10, 20, 30, 40];
    let total = thread::scope(|s| {
        let front = s.spawn(|| data[..2].iter().sum::<i32>()); // both borrow `data`
        let back = s.spawn(|| data[2..].iter().sum::<i32>());
        front.join().unwrap() + back.join().unwrap()
    });
    println!("total: {total}");        // total: 100
    println!("still here: {}", data.len()); // 4 -- data was only borrowed
}

Because thread::scope (episode 62) guarantees every spawned thread finishes before the scope returns, both closures are allowed to borrow data rather than needing an Arc or a move-and-clone. The value the scope produces flows straight out as total. As I said last time, swapping crossbeam::scope for thread::scope is a nearly mechanical modernisation. Right, homework cleared -- now, locks.

Mutex: exclusive access

A Mutex<T> wraps a value of type T and hides it. The only way to reach the inner value is to call lock(), which blocks the calling thread until the lock is free and then returns a MutexGuard. That guard is a smart pointer (episode 12) -- it derefs to &T or &mut T -- and, crucially, when the guard is dropped, the lock releases:

use std::sync::Mutex;

fn main() {
    let m = Mutex::new(5);
    {
        let mut guard = m.lock().unwrap(); // guard derefs to the inner i32
        *guard += 10;
    } // guard dropped here, lock released
    println!("{}", *m.lock().unwrap()); // 15
}

Look at what the type system has just done for you. Because access requires the guard, and the guard is the lock being held, it is literally impossible to read or write the protected value without holding the lock. In C or C++ the mutex and the data it protects are two separate things held together by nothing more than a comment and the programmer's good intentions -- forget to lock, and the compiler says nothing. In Rust the lock and the data are one object, and forgetting to lock is not a bug you can write. That is the same "make invalid states unrepresentable" philosophy we have met again and again in this series.

The .unwrap() on lock() is there because lock() returns a Result -- and why it can fail is the poisoning story we will get to shortly. For now, read .unwrap() as "give me the guard, and panic if the lock was poisoned".

Sharing a Mutex across threads

A Mutex on its own only gives you the exclusion; it does not give you shared ownership. To let several threads own the same mutex you wrap it in an Arc (episode 34), the thread-safe reference-counted pointer. The pairing Arc<Mutex<T>> is so common in Rust concurrency that it is practically an idiom -- Arc for "many owners", Mutex for "safe mutation". Here four workers each append a line to one shared log:

use std::sync::{Arc, Mutex};
use std::thread;

fn main() {
    let log = Arc::new(Mutex::new(Vec::new()));
    let mut handles = Vec::new();
    for id in 0..4 {
        let log = Arc::clone(&log); // clone the Arc, not the Vec -- cheap pointer bump
        handles.push(thread::spawn(move || {
            log.lock().unwrap().push(format!("worker {id} done"));
        }));
    }
    for h in handles { h.join().unwrap(); }
    println!("{} messages logged", log.lock().unwrap().len()); // 4
}

Trace the ownership carefully, because it is the whole point. Arc::clone does not copy the Vec -- it bumps a reference count and hands back another pointer to the same underlying Mutex<Vec<...>>. Every worker locks that one mutex, pushes its line while holding the lock, and releases it when the temporary guard on that line is dropped at the semicolon. The four pushes are therefore serialised: the Vec is never touched by two threads at once, which is exactly why this compiles at all. Remember episode 61's Send and Sync? Arc<Mutex<T>> is Send + Sync precisely because the Mutex provides the synchronisation that makes shared mutation safe -- the marker traits and the lock are two halves of the same guarantee.

RwLock: many readers or one writer

Now, a Mutex is a blunt instrument. It serialises every access, even two threads that only want to read and would never step on each other. When your data is read far more often than it is written -- think a configuration loaded once and consulted constantly, or a cache queried by dozens of threads -- that blanket exclusion is pure waste. RwLock<T> (read-write lock) is the sharper tool. It hands out two different kinds of guard: any number of threads may hold a read() guard simultaneously, but a write() guard is exclusive and waits for all readers to clear out first:

use std::sync::{Arc, RwLock};
use std::thread;

fn main() {
    let config = Arc::new(RwLock::new(vec![1, 2, 3]));
    thread::scope(|s| {
        for _ in 0..3 {
            let c = &config;
            s.spawn(move || {
                let r = c.read().unwrap();          // shared read: many at once
                println!("reader sees {} items", r.len());
            });
        }
        s.spawn(|| {
            let mut w = config.write().unwrap();     // exclusive write: alone
            w.push(4);
        });
    });
    println!("final: {:?}", *config.read().unwrap()); // final: [1, 2, 3, 4]
}

The rule of thumb is straightforward. Reach for RwLock when reads vastly outnumber writes, because then all those concurrent readers proceed in parallel instead of queueing behind one another. Reach for Mutex when writes are frequent, because in that case the reader/writer bookkeeping is overhead you never cash in -- an RwLock that is constantly being write-locked is just a slower, more complicated Mutex. And a small warning worth carrying: the exact fairness (does a waiting writer block new readers, or can a steady stream of readers starve the writer forever?) depends on the operating system's implementation, so do not lean on a particular ordering. If starvation matters to your design, that is a signal to rethink the sharing, not to guess at the lock's mood.

Lock poisoning

Here is the feature that genuinely sets Rust's locks apart, and it follows directly from the panic-and-unwind machinery we studied in episode 51. Suppose a thread locks a mutex, starts mutating the protected value, and then panics halfway through -- leaving the data in a broken, half-updated state. In most languages the lock would simply be released by the unwinding, and the next thread would happily lock it and read the corrupted mess, none the wiser. Rust refuses to let that slide. When a thread panics while holding a lock, the lock is marked poisoned, and every subsequent lock() returns an Err instead of a guard:

use std::sync::{Arc, Mutex};
use std::thread;

fn main() {
    let data = Arc::new(Mutex::new(vec![1, 2, 3]));
    let d = Arc::clone(&data);
    let _ = thread::spawn(move || {
        let mut guard = d.lock().unwrap();
        guard.push(4);
        panic!("worker crashed while holding the lock"); // poisons the Mutex
    }).join(); // join swallows the panic so main can continue

    let is_poisoned = data.lock().is_err();
    if is_poisoned {
        println!("lock is poisoned: a holder panicked"); // this branch
    } else {
        println!("lock was clean");
    }
}

The poisoning is not Rust being fussy for the sake of it -- it is Rust being honest. The whole reason we lock the data is to protect an invariant ("this Vec is always internally consistent"). If a thread panicked mid-update, that invariant might now be broken, and Rust would rather tell you loudly than let the corruption propagate silently. It is the concurrency equivalent of the Result type refusing to let you ignore an error. This is also, incidentally, why lock() returns a Result at all -- the Err variant is the poison signal.

Recovering from poison

Poisoning is a warning, not a death sentence. Sometimes you inspect the situation and conclude the data is perfectly fine, or you know how to repair it -- and Rust lets you take the value out anyway. The Err you get back is a PoisonError, and it carries the guard inside it; calling into_inner() on the poison error hands you that guard so you can carry on:

use std::sync::{Arc, Mutex};
use std::thread;

fn main() {
    let data = Arc::new(Mutex::new(0));
    let d = Arc::clone(&data);
    let _ = thread::spawn(move || {
        let mut g = d.lock().unwrap();
        *g = 42;
        panic!("crash"); // poisons the lock, but the write already happened
    }).join();

    let value = match data.lock() {
        Ok(g) => *g,
        Err(poisoned) => *poisoned.into_inner(), // take the value out of the PoisonError
    };
    println!("recovered value: {value}"); // recovered value: 42
}

The design here is elegant once you see it. Rust does not hide the data behind poisoning -- it makes you write an explicit line of code that says "yes, I know a holder panicked, and I have decided this value is still usable". That single conscious into_inner() is the difference between "I handled the corruption risk" and "I accidentally read garbage". If you genuinely do not care about poisoning -- say the protected value is a simple counter where a half-update cannot really corrupt anything -- a common shortcut is .lock().unwrap_or_else(|e| e.into_inner()), which grabs the guard whether or not the lock is poisoned. Use it deliberately, not reflexively.

Hold the lock as briefly as possible

Locks are about correctness, but they are also about performance, and the single most important performance habit is this: hold the lock for as short a time as you possibly can. A lock held during slow work is a bottleneck, because every other thread that wants the same lock is stuck waiting while you dawdle. The fix is a pattern you will use constantly -- grab what you need under the lock, release it immediately, then do the expensive part unlocked:

use std::sync::Mutex;

fn main() {
    let m = Mutex::new(vec![1, 2, 3, 4, 5]);
    let snapshot: Vec<i32> = {
        let guard = m.lock().unwrap();
        guard.clone() // copy out what we need...
    }; // ...and release the lock HERE, before the slow work below
    let sum: i32 = snapshot.iter().sum(); // computed without holding the lock
    println!("sum computed off-lock: {sum}"); // 15
}

The little inner block { ... } is doing real work: it bounds the lifetime of guard so the lock releases at the closing brace, before we start iterating. Contrast that with the lazy version where you lock the mutex and then run your whole calculation with the guard still alive -- every other thread queues behind you the entire time. Cloning out a snapshot costs a small allocation, but it buys you a lock that is held for microseconds instead of milliseconds, and under contention that trade is almost always worth it. Nota bene: this is also where a subtle deadlock hides -- if the "slow work" itself tries to lock the same mutex again, a plain Mutex will block forever waiting for itself (Rust's std Mutex is not re-entrant). Releasing early sidesteps that entire class of bug.

The golden rule against deadlocks

The scariest failure mode of locks is the deadlock: thread A holds lock 1 and waits for lock 2, while thread B holds lock 2 and waits for lock 1. Neither can proceed, neither will ever give up its lock, and your program simply hangs -- no panic, no error, no output, just a silent freeze. The prevention is almost embarrassingly simple to state and absolute in effect: always acquire multiple locks in the same order, everywhere in the program. If every thread that needs both a and b locks a first and b second, the cycle that a deadlock requires can never form:

use std::sync::Mutex;

fn main() {
    let a = Mutex::new(1);
    let b = Mutex::new(2);
    let sum = {
        let ga = a.lock().unwrap(); // ALWAYS a before b -- in every single code path
        let gb = b.lock().unwrap();
        *ga + *gb
    };
    println!("{sum}"); // 3
}

The reason a consistent order works is worth internalising: a deadlock is a cycle in the "who is waiting for whom" graph, and if every thread grabs locks in the same global order, that graph can only ever be a straight line, never a loop. It is a purely structural guarantee. In a big codebase this discipline is enforced by convention and code review (some teams even assign each lock a numeric rank and assert you only ever acquire in increasing rank), which is a bit of bookkeeping but infinitely cheaper than a production hang at 3am.

Having said that, the very best defence against deadlocks is needing fewer locks in the first place. A single lock cannot deadlock against itself under a consistent-order rule, and a channel (which we spent the last two episodes on) sidesteps shared-lock ordering entirely by not sharing the data at all. So the priority order in my head is: prefer message passing, then a single lock, and only reach for a lattice of multiple locks when you truly must -- and when you must, write the ordering down and never, ever break it.

How Python and Go would frame this

A glance sideways sharpens the picture, as it usually does. Python has threading.Lock, and the idiomatic use is a with block, which is Python's answer to Rust's guard-drops-the-lock -- the lock releases when the block exits, even on an exception:

import threading

counter = 0
lock = threading.Lock()

def bump():
    global counter
    with lock:              # acquired here, released when the block exits
        counter += 1        # protected, but nothing FORCES you to hold the lock

threads = [threading.Thread(target=bump) for _ in range(1000)]
for t in threads: t.start()
for t in threads: t.join()
print(counter)              # 1000

The with lock: block is genuinely nice and the auto-release is the same instinct as Rust's Drop. But look at what is missing: the lock and the counter are two unrelated objects. Nothing stops another function from reading or writing counter without taking the lock -- the protection lives in your discipline, not the type system. And if a thread dies mid-update inside the block, Python just releases the lock and moves on; there is no poisoning, no signal that the invariant might be broken. (The GIL papers over a lot of this for pure-Python code, but the structural point stands the moment you drop into C extensions or multiprocessing.)

Go leans the same way, with sync.Mutex and sync.RWMutex that map almost one-to-one onto Rust's two locks:

type Counter struct {
    mu sync.Mutex
    n  int
}

func (c *Counter) Bump() {
    c.mu.Lock()
    defer c.mu.Unlock() // released when Bump returns -- the Go idiom
    c.n++               // but the compiler won't stop you touching c.n unlocked
}

defer c.mu.Unlock() is Go's clever trick for "release no matter how we leave this function", and pairing the mutex with the field it guards in one struct is good Go style. Yet it is still style, not enforcement -- another method could touch c.n without locking and Go would compile it without complaint, leaving the data race to be caught (maybe) at runtime by the -race detector. Rust's move is to make the lock and the data literally the same object, so "touch the data without the lock" is not a program you can write. Same idea, three levels of enforcement: Python and Go trust you, Rust checks you.

Wrapping up

So there we have it: shared mutable state done safely. A Mutex<T> gives one thread at a time exclusive access, with the lock and the data fused into a single object so you cannot reach the value without locking, and the MutexGuard releasing the lock automatically on Drop. Wrap it in an Arc and many threads can share one protected value -- the Arc<Mutex<T>> idiom you will use for the rest of your Rust life. RwLock<T> refines the picture for read-heavy data, letting many readers run at once while writers still get exclusive access. And poisoning -- unique among mainstream languages -- turns a panic-while-locked from silent corruption into a visible Err you can inspect and, with into_inner, recover from.

The habits to carry forward are three. Hold every lock for the shortest possible time -- snapshot out and release before slow work. Acquire multiple locks in a single consistent order, always, to make deadlock structurally impossible. And prefer fewer locks (or a channel) over a tangle of them, because the lock you never take is the one that can never deadlock.

There is one thing I have been quietly leaning on this whole episode without ever opening it up, and it is bugging me a little. A Mutex has to somehow coordinate threads at the hardware level -- when two threads race to lock() the same mutex, what actually decides who wins, down at the level of the CPU? The answer is a tiny, fascinating primitive that sits underneath every lock in existence, and it is where we are headed next ;-)

Exercises

  1. Wrap a counter in Arc<Mutex<i32>> and increment it from ten threads, then join them all and print the total (it should be exactly ten times whatever each thread adds).
  2. Use an RwLock<Vec<i32>> shared across a thread::scope to let three reader threads print the length while one writer thread pushes a value, then print the final contents.
  3. Poison a Mutex on purpose by panicking while holding its guard, catch the panic with join, then recover the inner value with into_inner and print it.

Thanks for your time, and keep those locks short! ;-)

scipio@scipio

Learn Rust Series (#65) - Mutex, RwLock, and Handling Lock Poisonin... | Ecency