Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

There is no universal speed winner between compare-and-swap (CAS) and a lock. The answer depends on what work you compare, how much contention the workload creates, and the processor and implementation being measured. The sources cited here do not document a benchmark by this article’s author, so this is a sourced guide—not a first-person test report.

What does “CAS versus a lock” actually compare?

CAS is an atomic read-modify-write operation: it compares a shared value with an expected value and updates it only if they match. The Linux kernel’s v6.6 atomic API documentation describes atomic operations and their semantics; the API also includes operations such as atomic add and exchange. The exact semantics and implementation are architecture-specific concerns.

A CAS operation can be used once, or as part of a loop that retries after a failed comparison. A mutex-protected increment is a different operation again: it acquires a lock, updates the counter, and releases the lock. An atomic fetch-add is yet another comparison. Results only mean something when the compared implementations perform equivalent work.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Implementation being measured What the measurement includes What it does not establish by itself
One CAS attempt A compare-and-update attempt, which may fail if the expected value does not match. The cost of a complete algorithm that retries, or the cost of a mutex-protected operation.
CAS retry loop Repeated attempts until the update succeeds; failures can add work under contention. How a different atomic operation or lock performs under the same workload.
Atomic fetch-add An atomic addition operation; source-level operations can compile to different machine instructions on different platforms. How CAS or a mutex performs on another architecture or implementation.
Mutex-protected increment Lock acquisition and release as well as the counter update. How a lock performs when protecting a larger or differently structured critical section.

Travis Downs’s 2020 comparison of concurrency costs demonstrates why the distinction matters: in his particular maximum-contention, single-counter benchmark, atomic add was significantly faster than his CAS loop, while std::mutex was competitive on the tested Skylake setup. Those are results for that benchmark and those machines—not a general ranking.

Is CAS faster than a lock?

Not as a general rule. In a lightly contended workload, an atomic operation may avoid some locking overhead. But a contended CAS loop can spend time retrying, and the cost of competing for ownership of a shared cache line can dominate. A mutex may perform differently depending on its implementation, the platform, and the amount of work it protects.

The ETH Zurich SPCL atomic-operations project says that “All the tested atomics have usually comparable latency and bandwidth.” That observation is bounded to the operations and architectures evaluated in that study; it is not a claim about every processor or application. Taken together, these sources support a narrower conclusion: performance is a property of an operation, implementation, workload, and machine—not of the label “CAS” or “lock” alone.

Why can an early benchmark conclusion be wrong?

The implementations may not be doing equivalent work

A single CAS attempt, a loop that retries until success, an atomic fetch-add, and a mutex-protected increment have different costs and behavior. Comparing one with another can answer a useful question, but it cannot isolate a simple “CAS versus lock” effect unless the measured task is clearly defined.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Contention changes the work being measured

A shared counter updated by many threads is a deliberately narrow, heavily contended workload. It can expose retry costs and cache-line competition, but it is not a stand-in for every application. A real critical section may update several values, do substantial work, or distribute access among multiple locations. Single-threaded, lightly contended, and heavily contended cases can produce different results.

The machine and software stack are part of the result

Processor architecture, compiler output, language and runtime, lock implementation, memory ordering, cache location and coherence state, NUMA placement, alignment, and operand size can all affect a measurement. A result on one platform does not settle what happens on another.

The metric can change the apparent winner

Wall-clock time and aggregate throughput answer different questions. Per-thread results can reveal that total throughput hides an uneven distribution of progress. Repeated trials and variability help show whether a small difference is consistent or within measurement noise.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What should a credible CAS-versus-lock benchmark report?

Describe enough of the setup for another developer to understand what the result means and reproduce the comparison. At minimum, report:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Machine: processor model, architecture, and relevant memory-system details, including NUMA placement if applicable.
  • Implementation: language, runtime, compiler and relevant options, generated operation where known, memory ordering, and lock type.
  • Workload: the exact operation and protected work, thread count, iterations, contention level, and whether threads share a counter or other cache line.
  • Procedure: how thread starts are coordinated, warmup policy, repeat count, and whether trials use the same conditions.
  • Results: wall-clock time, throughput, variation across runs, and per-thread results when imbalance matters.

Changbin Du’s September 30, 2026 Linux perf mailing-list proposal illustrates some of these choices: it synchronizes thread starts, uses a shared cache-line-aligned counter, excludes an initial warmup run, and reports aggregate and per-thread results. Its example uses two threads, 100,000,000 iterations per thread, and 10 repeats after one warmup. The proposal’s example output reports “Avg wall-clock time: 7365.480 msec (stddev 66.014 msec)” and “Throughput total: 27,153,697 ops/sec.” These figures are example output for the proposed CAS benchmark, not an independently validated run or a comparison against a lock. The mailing-list post is a patch proposal; it does not establish that the feature was accepted or released.

How should you choose between CAS and a lock?

  1. Start with correctness and design. Identify the data that must be updated together and the synchronization guarantees the program needs. Do not choose an operation based on speed before establishing that it protects the right state.
  2. Choose the simplest sound implementation. CAS can support synchronization and lock-free algorithms, but complex CAS-based operations can bring complexity, scalability, and performance costs. Paul E. McKenney discusses those trade-offs in Is Parallel Programming Hard, And, If So, What Can You Do About It?, version 2024.12.27a.
  3. Benchmark the real task. Compare implementations that protect or perform equivalent work, using the thread counts and contention patterns your application is expected to encounter.
  4. Keep the conclusion bounded. State the tested machine, software, workload, and metric alongside the result. If those conditions change, measure again before carrying the conclusion over.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.