Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use go test -bench with the -cpu flag to compare benchmark runs at different Go parallelism limits. For parallel throughput, the benchmark must itself run parallel work—typically with b.RunParallel. Repeat the runs, compare them with benchstat, and record the environment so changes in CPU settings are not confused with changes in the machine or runtime.

1. Choose a benchmark that measures the work you care about

Go runs benchmark functions named BenchmarkXxx(*testing.B) when you invoke go test with -bench. For new benchmarks, use b.Loop() where the Go version you are targeting supports it; the testing package documentation describes this form as more robust and efficient than the older b.N-style loop. Put setup outside the timed loop when setup is not part of the operation you intend to measure.

A normal benchmark measures the code path you call. Adding -cpu does not make a serial operation parallel; it lets the test binary run with different CPU counts. To measure parallel throughput, put the operation under test inside b.RunParallel:

func BenchmarkWork(b *testing.B) {
    // Prepare shared benchmark data here if setup is not part of the operation.
    b.RunParallel(func(pb *testing.PB) {
        for pb.Next() {
            work()
        }
    })
}

RunParallel distributes iterations among goroutines and is commonly used with go test -cpu. Its default benchmark goroutine count is based on GOMAXPROCS; b.SetParallelism(p) changes that count to p * GOMAXPROCS. The testing documentation says that adjustment is usually unnecessary for CPU-bound benchmarks. Also note that RunParallel reports ns/op as wall time for the whole benchmark, not the sum of time spent by its goroutines.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Run the benchmark at several CPU counts

This command is an example invocation, not a performance result:

go test -run='^$' -bench='BenchmarkWork' -benchmem -cpu=1,2,4,8 -count=10 ./path/to/package

-run='^$' skips ordinary tests, while -bench selects the benchmark. The comma-separated -cpu list asks the test binary to run at each listed CPU count; -count=10 requests repeated samples, and -benchmem includes allocation metrics. Choose CPU counts that make sense for the environment, and choose run duration and repetitions based on measurement noise and cost rather than treating one setting as universal.

Keep the benchmark code, Go toolchain, machine, and workload conditions consistent between comparisons. Save the raw output and record at least:

  • Go version, operating system, architecture, and CPU model.
  • The -cpu values and whether GOMAXPROCS was explicitly set.
  • Process CPU affinity and any container or cgroup CPU limits.
  • Benchmark operation, units, repetition count, and allocation results.

3. Understand what the CPU settings control

-cpu sets test-run CPU counts

The testing flag -cpu takes a comma-separated list of CPU counts for benchmark and test runs. It controls the runtime parallelism setting used for those runs; it does not guarantee a corresponding number of physical cores, nor does it ensure that the benchmark can use that much parallel work.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

GOMAXPROCS limits simultaneous Go execution

The runtime documentation defines GOMAXPROCS as the maximum number of OS threads that can execute user-level Go code simultaneously. Think of it as a parallelism limit, not a physical-core counter.

Current runtime defaults can take account of logical CPU count, process affinity, and, on Linux, average CPU throughput limits imposed through cgroups. A fractional cgroup limit is rounded up to an integer GOMAXPROCS; the documented default keeps a minimum of 2 except when the logical CPU count or affinity is below 2. Automatic default updates can happen periodically. Setting GOMAXPROCS explicitly disables those updates.

Container CPU quota and parallelism are different constraints

Go 1.25 introduced container-aware GOMAXPROCS defaults: when the setting is otherwise unspecified, the runtime can account for a container CPU limit and periodically update its value. The Go 1.25 announcement explains that GOMAXPROCS is a parallelism limit. A CPU quota, by contrast, caps CPU throughput over time. The same numeric value for the two controls therefore does not describe identical constraints for every workload.

If you compare a container with a host run, or run with an explicit GOMAXPROCS or -cpu value, record that choice. Such a result describes the selected setting; it should not be presented as the behavior of an unspecified production default.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

4. Compare repeated results, not the best-looking run

Use benchstat to compare repeated benchmark output. The Go testing documentation recommends it for statistically robust A/B comparisons. Keep the software and environment consistent, changing the CPU-count dimension deliberately, and preserve the raw samples so the comparison can be checked.

When reporting a comparison, show the measured operation and units, CPU settings, repetitions, Go version, and relevant allocation results. For a parallel benchmark, interpret ns/op as wall time for the full parallel benchmark; where it helps answer the question, also express throughput as operations per second. Do not infer a universal speedup from one curve: available parallel work, synchronization, allocation and garbage collection, blocking, and resource limits can all affect the outcome.

5. Diagnose flat or negative scaling

If increasing the CPU count stops helping—or makes results worse—first check whether the benchmark contains enough independent work and whether processors are actually busy. A benchmark with a serial operation cannot reveal parallel throughput simply because it was run under several CPU settings.

The Go performance wiki describes scheduler tracing as a way to investigate programs that do not scale linearly with GOMAXPROCS, and recommends checking OS-provided CPU utilization. Use CPU profiles to find functions consuming CPU; blocking profiles and scheduler information can help distinguish time spent waiting from a shortage of runnable work. Include allocation metrics or profiling when memory-management work may be part of the change.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What to compare across CPU counts

Comparison axis What to examine
Throughput or latency ns/op and, when meaningful, operations per second. For RunParallel, ns/op is wall time for the benchmark as a whole.
Scaling behavior How results change as the CPU count rises, tied to the actual workload and repeated samples.
Memory effects Allocation metrics and, when relevant, profiling of allocation or garbage-collection work.
Resource context Logical CPUs, affinity, container or cgroup limit, Go version, operating system, and architecture.
Variability Repeated samples and a benchstat comparison, rather than a single run.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.