Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A single time ./a.out result tells you how long one invocation took under one set of conditions. It does not show how much runs vary or establish that a code change made the program faster. To make a defensible comparison, keep the build, workload and timing method consistent; collect repeated measurements; and report both a representative summary and the spread.

Why does the same program take different amounts of time?

Elapsed time can change from run to run even when the program and input do not. Benchmark results may be affected by CPU frequency scaling and boost behavior, differences in speed between cores, other work scheduled on the CPU, context switches, simultaneous multithreading (SMT), cache activity and NUMA placement. These are possible sources of variation, not proof that any one of them affected a particular run. Google Benchmark documents these factors in its User Guide.

The timer also matters. Elapsed, or real, time measures how long the invocation takes from start to finish, including time spent waiting or being descheduled. CPU time measures processor time consumed by the program. For multithreaded code, those measures can tell different stories; Google Benchmark explains the distinction in its timing guidance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What does one timing result establish?

It establishes the duration of that particular execution, using that timer, on that machine and in that system state. It does not reveal the distribution of repeated runs or tell you whether the result is typical. Google Benchmark warns that one result may not be representative because benchmarks are often noisy; its documented default is to run each benchmark once and report that single result.

That is why the best-looking run is not a sound basis for a speed claim. A single observation cannot show whether a difference between two versions is larger than the ordinary variation in their timings.

How should you measure a small program?

Keep the comparison like for like

Build both versions with the same compiler and flags, run the same workload and input, and use the same timing method. Note the machine and operating system, along with other run conditions that could affect interpretation. If you change more than the code, you cannot cleanly attribute a timing difference to the code change.

Decide whether you want cold-start or warmed behavior

Startup and cache state can be part of the performance question. If you care about first-run behavior, measure cold starts. If you want warmed steady-state behavior, decide how warmup will work and say so. Warmup measurements can be discarded to exclude startup or cache-filling effects, but doing so changes what the result describes. Google Benchmark supports a warmup interval whose measurements are omitted from its reported result; its documented default warmup interval is 0.0 seconds.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Repeat runs and show their variation

Collect multiple observations rather than reporting only one or selecting the fastest. Show individual timings, or summarize the results with an appropriate measure of center and spread. Google Benchmark supports repeated runs and can report the mean, median, standard deviation and coefficient of variation. Its documented defaults include one repetition and a minimum benchmark time of 0.5 seconds; these are framework settings, not universal requirements for other programs or tools.

There is no single repetition count that suits every workload, and the available guidance does not establish a universal threshold for deciding whether a difference matters. Use enough observations to understand the variation relevant to your comparison, and be cautious when the apparent change is small relative to that variation.

Record enough context to interpret the result

Include the compiler and flags, machine and operating system, workload or input, timing method, and run conditions that matter to the comparison. A number without this context is difficult to reproduce or interpret. Google Benchmark’s reports include machine context and support custom context, such as compiler version.

How can you tell whether a code change made the program faster?

Measure the old and new versions under comparable conditions, then compare their repeated-run results—not their best individual runs. State whether you are comparing elapsed time or CPU time, and whether the target is cold-start or warmed behavior. Report a representative summary alongside the spread so readers can see whether the difference is stable or buried in run-to-run noise.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

You can also report the absolute time difference and the percentage change, making clear which values come from your measurements. A statistical comparison may help for some benchmark designs: Google Benchmark’s comparison tooling documents a Mann–Whitney U test. That does not make a statistical test mandatory for every small example, nor does it supply a universal threshold for practical importance.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What do benchmark-tool defaults mean?

Defaults describe a tool’s behavior, not a general rule for measuring every program. Google Benchmark’s guide documents a default of one repetition, a 0.0-second warmup interval and a 0.5-second minimum benchmark time. Linux’s upstream perf bench framework supports --repeat; its current documentation gives that option a default of 10. That default applies to perf bench, not to a standalone ./a.out timing or to all benchmarking tasks. See the perf bench manual.

A practical checklist before making a performance claim

  • Use the same compiler, flags, machine, input and timing method for the versions being compared.
  • Choose whether the claim concerns elapsed time or CPU time.
  • Decide whether you are measuring cold-start or warmed behavior, and disclose any discarded warmup runs.
  • Collect repeated observations and report their spread as well as a representative summary.
  • Do not describe a small difference as meaningful if it is within the variation you observed.
  • Include enough build, machine and workload context for another person to understand the result.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.