The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
There is no universal “good” tokens-per-second (TPS) score for a large language model. TPS can mean the speed of one response or the combined output of many concurrent requests, and neither number says much without the prompt, model, latency, and measurement method. To benchmark fairly, define which decision the test should inform, use a fixed workload, measure both responsiveness and throughput, and report the conditions alongside the results.
What tokens per second measures—and what it leaves out
Tokens per second is a rate, but benchmark tools do not all define the rate the same way. Check whether a reported figure counts generated output tokens alone or combines input and output tokens; whether timing starts before or after the first token; and whether the result represents one request or multiple requests together. Those choices can produce numbers that look comparable but answer different questions. NVIDIA notes that benchmarking tools can use different metric definitions in its guide to LLM inference benchmarking concepts.
For example, Ollama describes its per-request TPS as output tokens generated per second after the first token, excluding the initial wait. That is a project-specific methodology, not a universal standard; its methodology explanation is useful for understanding what its figure represents.
- Per-request output TPS: the generation pace of a single response stream. It does not include startup delay under the Ollama definition and does not establish how many users a service can support.
- Aggregate output throughput: all output tokens produced per second across concurrent requests. This describes system capacity under a stated workload, not the pace experienced by each individual user.
- Input-plus-output rates: some tools may count both prompt and generated tokens. Such a figure should not be treated as output TPS unless the counting rule is clear.
Always keep the unit visible: TPS means tokens per second. Latency metrics such as TTFT and TPOT are usually given in milliseconds or seconds; converting a token interval into a rate requires taking its reciprocal and stating which interval is being converted.
#1 Best Overall
Which speed metrics matter to a user?
An interactive response has an initial wait followed by a stream of generated tokens. Looking at only one phase can obscure the experience: a model may begin quickly but generate slowly, or take longer to begin and then stream rapidly.
- Time to first token (TTFT): elapsed time until the first content token arrives. It can include queuing, prompt processing, and network time, depending on where and how it is measured. NVIDIA describes a client-side TTFT measure that includes queuing, prefill, and network latency. Its authors define TTFT as “the time it takes to process the prompt and generate the first token” in the NVIDIA benchmarking concepts article.
- Time per output token (TPOT) or inter-token latency (ITL): the average interval between generated tokens after the first. Definitions vary by tool. NVIDIA’s GenAI-Perf definition excludes TTFT and divides generation time by the output-token count minus one.
- End-to-end latency: time from sending the request until the final token is received. It includes the initial wait and generation, though exact treatment of queuing and transport depends on the measurement method.
For one complete answer, a useful mental model is: total response time consists of the wait for the first token plus the time spent generating the remaining output. A prompt’s input tokens are processed during prefill; the model then generates output autoregressively, one token at a time. Longer prompts can extend prefill and TTFT, while longer outputs generally add generation time. Databricks explains this distinction and the latency-throughput trade-off in its endpoint benchmarking guidance.
Rank #2
How many tokens per second is a good speed for an LLM?
There is no evidence-backed universal threshold. A useful target depends on the model, prompt and output lengths, whether the system serves one person or many, and the latency the application can tolerate. A TPS result without those conditions is not enough to decide whether a system is fast for your purpose.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
- Interactive use: prioritize TTFT, TPOT or ITL, and full response time. A high generation rate does not compensate for an unacceptable initial wait.
- Batch processing: aggregate output tokens per second may matter more than the speed of any one response, provided the resulting completion times meet the job’s requirements.
- User-facing service: measure throughput while enforcing a latency budget. Databricks recommends maximizing throughput within the application’s latency budget, rather than treating peak throughput as the goal.
To evaluate a particular service, define an acceptable wait and completion time for its real users, then test whether it meets those requirements under the expected workload. Google Cloud’s accelerator benchmarking guidance likewise frames inference comparisons around fixed workloads and latency-constrained throughput.
Rank #3
How to benchmark LLM inference speed reproducibly
- Decide what the result will inform. Choose whether you are evaluating interactive responsiveness, sizing an endpoint, comparing local accelerators, or estimating batch capacity. That decision determines which metrics and latency limits matter.
- Fix a representative workload. Use the same prompt set and specify input and output token lengths or their distributions. For a fair comparison, hold the model and version, tokenizer, quantization or precision, serving stack, streaming mode, and generation settings constant. Prompt length affects prefill and initial wait; output length affects generation duration.
- Warm up and repeat the test. Record the tool and methodology, warm-up approach, number of runs, and whether each reported statistic is a mean, median, or percentile. NVIDIA’s NIM latency-throughput benchmarking guide organizes testing around warm-up, workload sweeps, and analysis; use the documentation for the exact tool and version to determine command options.
- Measure a single stream and a concurrency sweep. A single request characterizes one stream’s generation pace. Increase concurrent requests to observe aggregate throughput, latency, and queuing. As parallel work rises, throughput may improve while response latency also rises; eventually, a provisioned-capacity limit can constrain further gains.
- Record a complete metric set. Include per-request output TPS or TPOT, TTFT, end-to-end latency, aggregate output throughput, concurrency, and success or error rate. Report p50 and a tail percentile such as p95 or p99 when there are enough observations to make that percentile meaningful.
- Find the operating point that meets the service constraint. For an interactive service, report the throughput achieved before the chosen latency target is exceeded, not just the highest raw throughput observed. Google Cloud describes increasing concurrency until the p99 latency service-level objective is violated, then recording sustained throughput in its benchmarking guidance.
- State what the measurement includes. Disclose whether results came from an independent test, a vendor-published benchmark, or another named tool’s methodology. For hosted services, network path, queuing, and load can affect client-side measurements. A single run or a vendor headline is not a universal model or hardware specification.
What to include when publishing a TPS result
- Model name and version; tokenizer, precision or quantization, serving stack, and generation settings.
- Prompt workload, input and output token lengths or distributions, and whether requests streamed output.
- Metric definitions: counted tokens, timing start and stop, and whether the figure is per request or aggregated.
- Concurrency and test duration, plus the warm-up and repeated-run method.
- TTFT, TPOT or ITL, end-to-end latency, aggregate output throughput, and error or success rate.
- Summary statistics, including p50 and an adequately sampled tail percentile where available.
- For service comparisons, the latency target and the throughput achieved while meeting it.
- Measurement scope: hardware and configuration for local tests, or relevant endpoint and network conditions for hosted tests; identify the source of the result.
When comparing systems, align the workload and report responsiveness, capacity, and tail behavior separately. If efficiency or cost matters, normalize it to a clearly stated scope, such as performance per accelerator or per dollar, and disclose the hardware and pricing basis. A speed result alone does not establish model quality, and unlike prompts or output lengths do not make a fair comparison.
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

