The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A mutex call looks like a single function, but on Linux it passes through four layers: the API contract the library promises, a shared lock word in user memory, a futex system call only when a thread has to sleep, and the CPU’s atomic instructions that make each change to that word indivisible. The fast path, where nobody else holds the lock, never needs the kernel to track lock state. The kernel gets involved only when threads actually contend.
The four layers at a glance
The table below separates what each layer is responsible for. The boundaries matter because the POSIX standard defines only the top row; everything beneath it is an implementation choice.
| Layer | What it does | Who defines it | Enters the kernel? |
|---|---|---|---|
| API contract | Defines what the caller observes: acquisition, waiting, ownership, and behavior that depends on mutex type and attributes | POSIX (for example pthread_mutex_lock), and the language runtime that wraps it |
Not by itself |
| User-space lock state | Stores lock state in a shared word and claims it with atomic operations | The C library or runtime; layout is implementation-specific | No, on the uncontended path |
| Futex wait and wake | Blocks a thread only if the shared word still holds the expected value, and wakes sleepers on release | Linux kernel, documented in futex(2) and futex(7) |
Yes, only when a thread must sleep or a waiter must be woken |
| CPU atomics | Make a compare-and-exchange on the shared word indivisible across cores | CPU architecture (the futex documentation cites cmpxchg on x86 as one example) |
No |
Layer 1: the API contract
When your code calls a mutex lock operation, the contract tells you what you will observe. An unlocked mutex is acquired and the call returns. A mutex owned by another thread causes the caller to wait. How the wait behaves, whether the same thread may lock twice, and what happens on misuse depend on the mutex type and attributes you set. POSIX specifies those observable rules in pthread_mutex_lock(3p).
POSIX does not prescribe how a lock is stored, which system calls are used, or how many instructions run. A runtime such as a language standard library can add its own wrapper on top of the C library mutex, so the function you call may not be the one that touches the lock word directly. When this article says “the runtime,” it means that wrapper and whatever it delegates to.
#1 Best Overall
Layer 2: lock state in user space
In a futex-backed mutex, the lock state lives in ordinary memory that threads share. The futex documentation treats that word as the bridge between user-space coordination and kernel blocking. The uncontended case is handled entirely there, using atomic instructions, commonly described as compare-and-exchange, to claim the word when it shows the unlocked value.
The logic of the fast path can be sketched in three steps. This is a conceptual sketch, not source code for any particular C library, because the encoding of the state (for example, whether the word also records that waiters exist) differs by mutex type and implementation:
- Attempt an atomic transition of the shared word from “unlocked” to “locked.”
- If the transition succeeds, the thread owns the mutex and enters the critical section. No system call has been made.
- If it fails, the lock is held by someone else. The thread moves to the contended path described in the next section.
Because the fast path touches only shared memory, its cost is that of the atomic operation and the surrounding code, not a trip to the kernel. Whether that holds for every mutex type or library version is a question for that library’s own documentation, not something a single rule can cover.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #2
Layer 3: the contended wait
If the fast path fails, a thread can ask the kernel to put it to sleep with futex(2)‘s wait operation (FUTEX_WAIT). The call passes the address of the futex word and the value the thread expects to see there. The kernel compares that value with the word and blocks the thread only if they still match. If they do not match, the call returns without sleeping, which is how the thread learns that the state has already changed.
The comparison and the decision to block happen as one step relative to other operations on the same futex. That is what closes the race that would otherwise cause lost wakeups. Without it, a thread could read “locked,” then the owner could unlock and wake nobody because no one was yet asleep, and then the thread would sleep indefinitely. The expected-value check turns that window into a harmless retry.
Layer 4: release and wake
On unlock, the owner first changes the lock word so that it reads as unlocked. It then calls the wake operation (FUTEX_WAKE) if there may be sleeping waiters. The futex documentation notes that implementations can skip the wake call when no thread is waiting, which avoids a system call on the common uncontended release.
A wake is a notification, not a handoff. The woken thread goes back to the acquisition attempt, and another thread that arrives first can still take the lock. The woken thread may find it locked again and wait once more. Programs should never assume that a thread which was woken now owns the mutex.
Layer 5: CPU atomics
The CPU layer is what makes the transitions in the previous layers indivisible when several cores race on the same word. A compare-and-exchange either replaces the value or reports that it did not, and no other core can observe the operation half done. The futex manual cites cmpxchg on x86 as an example. Other architectures provide their own atomic primitives, and a portable library hides those differences behind its own abstraction.
Do not read this as “every mutex operation is one instruction.” The uncontended acquisition is a short atomic path. A contended acquisition can involve a system call, a scheduler decision that puts the thread to sleep, the wake that later resumes it, and another attempt to claim the word. Each step is atomic on its own; the sequence is not.
Rank #4
Specialized case: priority-inheritance futexes
Linux also provides a priority-inheritance variant, described in the kernel’s “Lightweight PI-futexes” documentation. Its user-space fast path atomically changes the futex value from zero to the owner’s thread ID. If that compare-and-exchange fails, the thread calls FUTEX_LOCK_PI, and the kernel handles the contended case by associating the futex with an RT-mutex, the kernel’s real-time mutex structure. That association is what allows the kernel to raise the owner’s priority while a higher-priority thread waits.
This mechanism exists to support priority inheritance. It is a specialized path with its own word encoding and its own kernel machinery, and it should not be treated as a description of every ordinary mutex. The kernel’s separate “Generic Mutex Subsystem” documentation describes the kernel’s own internal mutex, which is a different primitive from the user-space futex layers discussed above.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →What to compare when you evaluate an implementation
When two mutex implementations seem to behave differently, check the same six properties in each:
Best Value
- How much work the uncontended fast path does, and whether it enters the kernel.
- What state the shared word encodes, such as locked or unlocked only, or also waiter presence and owner identity.
- How the wait operation prevents the check-to-sleep race.
- When the release path issues a wake, and how the scheduler then treats the woken thread.
- Optional semantics: priority inheritance, robustness after an owner dies, recursion, and sharing between processes.
- ABI and platform constraints, including the architecture’s atomic instructions and the kernel version that provides each futex operation.
The sources cited here establish the fast path, wait and wake behavior, and the priority-inheritance path. They do not establish how a particular C library lays out its mutex structure or how two libraries compare in speed. Those questions need measurements or documentation from the specific library and version you use, and no performance figure in this article should be taken as a general result.
Platform scope
Everything above describes Linux. POSIX defines the API, and Linux documents define the futex mechanisms. Other operating systems, and other Linux C libraries, may use different internal layouts and sequences while presenting the same contract. If you are reading a specific library’s source, use that library’s layout as the authority, and use the layers here as a map for finding the corresponding parts.
Sources used for this article: the Linux man-pages project’s futex(2) and futex(7) pages, the Linux kernel documentation “Lightweight PI-futexes” and “Generic Mutex Subsystem,” and the POSIX pthread_mutex_lock(3p) page.
Quick Recap
“
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

