Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Poor scaling means that adding parallel capacity produces less throughput—or less latency improvement—than expected. It is a symptom, not a diagnosis. To find the cause, compare the same workload under controlled conditions, then use Go profiles and runtime traces to determine whether the limit is CPU work, memory and garbage collection, synchronization, scheduling, or an external resource such as network or disk I/O.

Start with a comparable scaling curve

Before changing code, measure what “doesn’t scale” means for your program. Run representative work at multiple parallelism levels while holding the input, machine or container limits, and measurement method steady. Record throughput, latency, and CPU utilization for each run. This is a diagnostic method, not a result that can be predicted without measurements of your application.

Look at the shape of the results. If throughput stops rising while CPU remains busy, investigate CPU consumption and contention. If CPU is underused while latency stays high, investigate blocking, scheduling, and external waits. A flat throughput curve alone cannot distinguish these causes.

Find where active CPU time goes

Capture a CPU profile when the process is doing substantial CPU work or throughput plateaus. Go’s diagnostics guide describes profiling and the available diagnostic tools. Use go tool pprof to inspect a profile in text, graph, or source-listing views, or with a flame-graph view.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A CPU profile identifies where the process spends active CPU cycles. It does not account for time a goroutine spends sleeping, blocked on a lock, waiting for I/O, or otherwise not running. If requests are slow but CPU use is low, do not treat the busiest function in a CPU profile as a complete explanation; examine blocking and resource waits instead.

Distinguish allocation churn from retained memory

Memory use and allocation rate are different questions. The heap profile’s live view helps identify objects still retained, while the allocs profile, viewed with -alloc_space, shows cumulative allocation churn, including memory that has already been reclaimed. Go’s profiling documentation explains these profile views and their interpretation.

Heap profiles are sampled, and the heap profile reflects the most recently completed garbage collection; it omits more recent allocation to avoid bias toward garbage. Treat profiles as statistical evidence, and use suitable repeated captures rather than assuming a single snapshot represents every allocation or retained object.

For broader context, runtime memory statistics and GC statistics can help show whether memory use or collection work changes with load. Interpret them alongside the heap and allocs views: a high cumulative allocation total does not, by itself, mean the same amount of memory remains live.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Investigate goroutines waiting on shared work

When goroutines appear to wait, use a block profile to find where they spend time blocked on synchronization primitives. Block profiling is not enabled by default, so an absent or empty block profile may indicate that collection was not configured—not that blocking does not occur. If lock contention is suspected, collect a mutex profile as well.

Read their attribution differently: a block profile points to the location where a goroutine blocked, while a mutex profile attributes contention to the end of the critical section that caused other goroutines to wait. If evidence points to a shared resource, possible experiments include sharding or partitioning it, reducing shared access, or buffering and batching work locally. Make one targeted change at a time and rerun the same scaling measurement.

Use an execution trace for scheduling and utilization questions

If CPU utilization or parallel execution is unclear, a Go execution trace can show scheduling, system calls, garbage collection, heap size, and related runtime events. It can help reveal work becoming serialized or goroutines being preempted by network activity or system calls. The Go diagnostics guide recommends profiles for finding CPU or memory hotspots; tracing is better suited to understanding runtime behavior than attributing a CPU or memory hotspot.

Runtime metrics, runtime.ReadMemStats, goroutine counts, stack dumps, and GODEBUG diagnostics can add high-level evidence about memory, GC, goroutines, and scheduling. Use these clues to decide what to profile or trace, rather than treating any one metric as a diagnosis.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check whether the limit is outside Go

A program may stop gaining throughput because it has saturated a network link, disk, or another external resource. Compare system and resource measurements with the Go profiles. The Go performance guidance notes that resource saturation can bound performance regardless of further program optimization. If the relevant external resource is already at its ceiling, adding CPU parallelism or tuning an unrelated function will not remove that bound.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choose the first tool for the symptom

Symptom or question First useful evidence What it can show Caveat
CPU is busy and throughput plateaus CPU profile Functions consuming active CPU time Does not account for sleeping or waiting time.
Memory use grows or GC work seems high Heap profile, allocs view, and runtime or GC statistics Live retained objects versus cumulative allocation churn Memory profiles are sampled; the heap profile reflects a completed GC.
CPU is underused and goroutines wait Block profile; mutex profile if lock contention is suspected Blocking stacks and sources of lock contention Block and mutex profiles require configuration and are not enabled by default.
More processors do not increase work Execution trace and scheduler-focused evidence Scheduling, serialization, system calls, GC, and utilization behavior Tracing helps explain runtime behavior; use profiles to locate CPU or memory hotspots.
Throughput tracks a network or disk ceiling System or resource measurements alongside Go profiles Whether an external bound may cap code-level gains Go’s performance guidance identifies resource saturation as a reason further program optimization may not help.

Collect production profiles carefully

Production profiling is possible, but collection can degrade performance. Estimate the overhead before enabling it. For a service with many replicas, Go’s diagnostics guidance describes periodically selecting a replica and collecting a profile. The net/http/pprof handlers support duration parameters for CPU profiles and traces; block collection requires enabling block profiling, and mutex collection requires configuring mutex profiling. Whether and how to expose these handlers safely depends on your deployment and access-control design.

Collect one profile at a time when diagnostic modes may interfere. Go’s documentation gives precise memory profiling and goroutine blocking profiling as examples that can skew CPU profiles or scheduler traces. Compare the overhead and results with the same workload and measurement method used for your scaling curve.

Consider PGO after you know the constraint

Profile-guided optimization (PGO) is a build-time optimization, not a substitute for diagnosing the bottleneck. Go’s compiler accepts CPU pprof profiles, and PGO can use profile information to make choices such as more aggressive inlining for frequently called functions. The Go PGO guide recommends representative production profiles and warns that an unrepresentative profile may provide little production benefit.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

PGO support began in Go 1.20. The guide reports performance improvements of around 2–14% on benchmarks for a representative set of Go programs with Go 1.22. That is a version-specific benchmark result, not a guaranteed gain for an individual application.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.