Recommended Free Tools
To benchmark LLM training on Google Cloud, run the same defined training workload across multiple accelerator counts and measure both raw throughput and useful progress over time. Report global tokens per second, tokens per second per chip, scaling efficiency, goodput, time to a shared quality target, and cost on a dated, region-specific basis. Peak accelerator specifications alone cannot show how quickly a real cluster will train your model.
Define a workload that makes the comparison meaningful
Before allocating accelerator capacity, write down the complete workload. If the model, data shape, software stack, or training objective changes between runs, the results do not isolate the effect of cluster size or accelerator type.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
Deep Learning (Adaptive Computation and Machine Learning series) | $51.51 | Buy on Amazon |
| 2 |
|
Deep Learning: Foundations and Concepts | $48.83 | Buy on Amazon |
| 3 |
|
Understanding Deep Learning | $99.22 | Buy on Amazon |
| 4 |
|
Deep Learning (The MIT Press Essential Knowledge series) | $11.36 | Buy on Amazon |
| 5 |
|
Deep Learning: A Visual Approach | $73.40 | Buy on Amazon |
- Model and objective: Record the model architecture and code version, training objective, optimizer, and the quality or convergence target you will use for comparison.
- Data and batch: Specify the dataset and input pipeline, sequence-length distribution, global batch size, and token shape.
- Numerics: State the precision and any quantized operations.
- Software: Pin the framework, compiler, runtime, and relevant model-code versions.
- Production path: Use the storage path and input pipeline intended for production; include checkpoint cadence and other workload behavior that affects elapsed time.
Google Cloud’s accelerator performance and benchmarking guidance recommends defining the workload and measuring accelerator training with tokens per second per chip. The same workload definition should accompany every result so another team can judge whether it applies to its own training job.
Establish a baseline and state what the clock includes
Start with the smallest viable configuration, then warm up and compile the job using the path you expect to use in production. Record steady-state step time as well as end-to-end elapsed time. These answer different questions: step time describes the job when it is running steadily; elapsed time can include setup and interruptions that matter to a real project.
#1 Best Overall
- Language Published: English
- Binding: hardcover
- It ensures you get the best usage for a longer period
For each run, record the accelerator model and count, topology, and whether the job uses one slice or multiple slices. State whether startup, compilation, data loading, checkpointing, retries, and recovery are included in each measurement window. Do not silently remove these intervals from an end-to-end result.
Use complementary metrics, not one headline number
No single metric captures both cluster capacity and training progress. Pair normalized throughput with whole-cluster throughput, utilization diagnostics, reliability, quality, and cost.
Rank #2
| Metric | What it tells you | What to qualify |
|---|---|---|
| Global tokens per second | How much training data the full cluster processes per unit time. | State the accelerator count; total throughput can rise simply because the cluster is larger. |
| Tokens per second per chip (TPS/chip) | Throughput normalized by accelerator count, useful for comparing configurations and following a scale curve. | It does not capture interruptions, model quality, or price on its own. |
| Model FLOPs Utilization (MFU) | Observed model FLOPs relative to an assumed hardware peak. | State the FLOP-accounting method and peak reference. MFU is a utilization diagnostic, not a measure of convergence time or business value. |
| Effective Model FLOPs Utilization (EMFU) | A broader operations-utilization measure used in Google’s mixed quantized and floating-point accounting. | Under Google’s described definition, EMFU can exceed 100%. Explain the numerator and peak reference instead of treating it as a conventional percentage ceiling. |
| Scaling efficiency | How effectively throughput changes as the cluster grows. | Identify strong or weak scaling and the baseline configuration. |
| Goodput | Useful computation advanced after wasted time is excluded. | Define useful progress and the observation window; report raw throughput alongside it. |
| Time to target quality | How long the run takes to reach an agreed model-quality point. | Use the same evaluation and convergence target for every comparison. |
| Cost-normalized throughput | Training throughput for a stated cost basis. | Give the region, price source, date, and whether relevant host, storage, networking, and idle-capacity costs are included. |
MFU is meaningful only when its operation count and hardware-peak reference are clear. Google’s 2023 TPU v5e case study describes EMFU for mixed quantized and floating-point operations, including why its value can exceed 100% under that accounting. See Google Cloud’s TPU v5e training case study.
Scale the same job to reveal the system tax
Repeat the workload at several feasible cluster sizes. Compare both total throughput and TPS/chip at each point: a larger system may process more tokens overall while delivering less throughput per chip. Google Cloud’s guidance gives 256, 1,024, and 4,096 chips as example scale points, not as mandatory sizes for every benchmark.
Rank #3
- Choose scale points: Use the same model and software stack at each size where possible. Record chip counts and topology.
- Label the scaling design: In a strong-scaling test, hold total work fixed as the cluster grows. In a weak-scaling test, grow the work with the system. Do not compare these as if they were the same experiment.
- Calculate efficiency against a declared baseline: Report how throughput changes relative to that configuration, and state any changes to parallelism or batch size.
- Inspect the curve: If TPS/chip falls as the cluster grows, investigate communication, input delivery, synchronization, scheduling, and other system overheads rather than attributing the change to peak chip capacity.
The scale curve helps explain whether added accelerators produce useful throughput or mostly add coordination overhead. A single largest-run result cannot show that.
Measure goodput and compare progress at equal quality
Large clusters can lose wall-clock time to hardware faults, network stalls, retries, and checkpoint recovery. Goodput makes those losses visible by measuring useful computation over a stated observation window. Define what counts as progress—for example, optimizer updates that advance the intended training run—and report the numerator and denominator. Pair goodput with raw throughput so readers can distinguish a system that is fast when healthy from one that sustains useful work over a longer run.
When configurations differ in convergence behavior or final quality, compare elapsed time to the same evaluation target rather than treating tokens per second as the outcome. Hardware utilization cannot substitute for a shared quality criterion.
Add cost only with a dated, comparable basis
Once the workload and measurement method are fixed, report throughput per dollar or per chip-hour with the price source, region, and observation date. Cloud prices and product availability can change, so an undated performance-per-dollar result should not be treated as a current price comparison.
Best Value
Include host, storage, networking, setup, and idle-capacity costs when they are material to the tested setup. A compute-only figure may not predict the cost of operating the real training project. Google Cloud recommends normalizing throughput by accelerator cost, but that comparison remains specific to its workload and price basis.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How published Google Cloud figures should be read
Google’s published results illustrate why benchmark figures need their configuration and date attached. They are useful platform-specific examples, not independent cross-cloud evaluations or general guarantees for every model.
- TPU v5e, 2023: Google reported a November 2023 training run using 50,944 Cloud TPU v5e chips across 199 pods, described at publication as what it believed was the largest publicly disclosed LLM distributed training job by chip count. That is a historical, vendor-attributed claim—not a current record. In a separate reported scaling result, Google measured 66.86% MFU for BF16 training on a single TPU v5e pod. Its full-cluster INT8 result was 5.32 exa-operations per second using AQT; that figure is not directly comparable to floating-point FLOP/s. The case study also described limited software optimizations and ongoing work on compiler, MaxText, scheduling, stability, and multipod performance. Details: Google Cloud’s TPU v5e case study.
- Trillium, MLPerf 4.1, 2024: Google reported 99% throughput scaling efficiency for a multislice GPT-3 175B comparison across data-center networks, using a stated base configuration of four 256-chip Trillium pods. It also reported 94% throughput scaling efficiency for a TPU v5p cluster within a single ICI domain. These figures apply to their stated experimental setups. Google’s claim of up to 1.8x better performance per dollar for Trillium versus TPU v5p is likewise vendor-reported, workload- and price-specific, and not a guarantee for current prices or other training jobs. The analysis distinguishes throughput scaling, convergence scaling, and performance per dollar; one does not establish the others. Details: Google Cloud’s Trillium MLPerf 4.1 analysis.
Benchmark report checklist
A useful report lets another team reproduce the run and understand its limits. Include:
Quick Recap
- Model, model-code version, objective, data and sequence shape, global batch, precision, and optimizer.
- Framework, compiler, and runtime versions.
- Accelerator model and count, topology, and slice configuration.
- Warm-up and compilation policy, measurement window, and intervals included in each time metric.
- Global tokens per second, TPS/chip, and MFU or EMFU with the accounting method and reference.
- Scale points, scaling mode, baseline, and scaling-efficiency calculation.
- Failures, retries, network stalls, checkpointing, and recovery time, plus the goodput definition.
- Quality evaluation and target if comparing convergence.
- Cost calculation, region, price source and date, and included infrastructure costs.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

