Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

LLM inference throughput measures how much output a serving system generates over time; latency measures how long a request takes to begin or finish. Raising concurrency can increase total tokens per second while making each request feel slower. To choose a useful operating point, test realistic prompts and response lengths at different concurrency levels, then select the highest throughput that still meets your application’s latency objective.

What does tokens per second mean for an LLM?

Tokens per second (TPS) usually means the number of output tokens generated per second during a measured interval. In a concurrent serving system, system throughput totals output from multiple requests. It is not the same as the speed experienced by one user: an endpoint can produce more tokens overall while each individual request receives tokens more slowly.

Another measure, requests per second (RPS), counts completed requests rather than tokens. Because requests can have very different prompt and response lengths, RPS and TPS are not interchangeable. When reading a benchmark, check whether its “tokens per second” figure is system-wide or per user.

Which latency metrics matter?

Metric What it measures What to look for
Time to first token (TTFT) Time from submitting a query until the first non-empty output token arrives. How quickly an interactive response starts. It can include queuing, prompt prefill, and network latency.
Inter-token latency (ITL) or time per output token (TPOT) The time associated with generating output tokens after generation starts. ITL commonly means the average interval between consecutive tokens. How steadily output arrives while a response is being generated. Definitions and averaging methods differ by tool.
End-to-end request latency Time from sending a query until the complete response arrives. The full wait, including relevant serving-path effects such as queuing, batching, and networking.
System output-token throughput Total output tokens divided by the measured benchmark interval across concurrent requests. Overall serving capacity under the stated workload and measurement boundary.
Per-user throughput Generated output relative to the elapsed time for an individual request. The speed experienced by a request, which may fall as concurrent load increases.

Databricks gives a useful simplified relationship: Latency = TTFT + (TPOT × number of tokens generated). Treat it as a way to understand how startup delay and generation time contribute to a response, not as a replacement for checking how a benchmark measures end-to-end latency. For example, a long answer can take substantially longer to complete even when its TTFT is unchanged.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
ASRock Radeon AI PRO R9700 Creator 32GB Professional Graphics Card, 2920 MHz Boost Clock, GDDR6, AMD RDNA 4, AI-Accelerators, DisplayPort 2.1a, PCIe 5.0, Blower Cooler
  • Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
  • Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
  • Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
  • Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
  • Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.

In NVIDIA’s AIPerf definitions, TTFT is excluded from ITL, and ITL accounts for output-token intervals after the first token. Other tools may calculate averages or define measurement boundaries differently. Name the tool and its definitions when reporting a result. See NVIDIA’s LLM benchmarking metrics documentation and Databricks’ endpoint benchmarking documentation.

Does higher concurrency make an LLM faster?

Not necessarily. Concurrency is the number of requests being handled in parallel. At low concurrency, a serving system may have spare capacity and provide low latency. Adding requests can keep more resources busy and increase aggregate TPS, but it can also lengthen queues, increase request latency, and reduce per-user token speed. Once the system is saturated, throughput may plateau or even decline as contention and queueing grow.

Rank #2
Sale
HPE NVIDIA Tesla V100 32GB HBM2 PCIe 3.0 x16 Passive GPU Computational Accelerator for AI Machine Learning HPC Deep Learning 699-2G500-0216-400 (Renewed)
  • NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
  • 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
  • PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
  • NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
  • Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads

The useful target depends on the application: an interactive assistant may prioritize prompt starts and smooth token delivery, while a batch workload may accept longer waits to maximize total output. Databricks recommends maximizing throughput within the application’s latency budget. Measure the trade-off instead of assuming that the highest supported concurrency is the best setting.

How do I choose a concurrency level?

  1. Define the latency objective. Decide which user-facing measure matters—such as TTFT, inter-token latency, or full-response latency—and set an acceptable limit. Include tail behavior if slow requests are especially disruptive.
  2. Prepare a representative workload. Use realistic prompt and completion lengths. Input length affects prompt processing and memory demand; output length affects how long generation continues.
  3. Run a concurrency sweep. Test multiple levels using the expected request arrival pattern and enough requests to observe steady behavior. Record aggregate output-token throughput alongside user-facing latency.
  4. Find the feasible operating point. Plot throughput against the chosen latency measure and select the highest-throughput point that stays within the objective. If latency rises sharply while throughput barely improves, further concurrency is unlikely to help that workload.
  5. Retest after meaningful changes. A different model, hardware, serving backend, prompt mix, response-length distribution, or server configuration can shift the result.

NVIDIA’s current NIM benchmarking guidance recommends plotting output-token throughput against inter-token latency and retaining both the accepted input configuration and the resolved server configuration with the results. See NVIDIA’s NIM benchmarking guidance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
MX3 M.2 AI Accelerator
  • High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional performance and efficiency.
  • Flexible Integration: The MX3 can be easily integrated into your existing systems via its M.2 M-key form factor and support for Linux operating systems.
  • Energy Efficient: The MX3 is designed to provide high performance while minimizing power consumption.
  • Comprehensive Software Development Kit (SDK): The MX3 is supported by a comprehensive SDK that simplifies development and deployment.
  • Hardware compatability: The MX3 is compatible with the PCI-SIG M.2 M-key 2280 Specification. It can be used with the Raspberry Pi 5 with a M-key 2280 HAT.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How should I compare LLM inference benchmarks fairly?

A throughput number is meaningful only alongside its workload and measurement method. Record the following so readers can tell what was measured and whether another result is comparable:

  • Model and serving stack: model and version, serving backend, hardware or GPU, and relevant configuration.
  • Workload shape: input and output token lengths or their distributions, plus any dataset or prompt-generation method.
  • Load pattern: concurrency, number of requests, and how requests arrive, such as fixed-rate or open-loop arrivals.
  • Metric definitions: whether TPS is aggregate system output or per-user speed, and how TTFT, ITL/TPOT, and end-to-end latency are calculated.
  • Measurement boundary: whether timing includes warm-up, queueing, networking, tokenization, and post-processing.
  • Latency distribution: percentiles as well as averages where available; an average alone can hide a slow tail.

NVIDIA notes that benchmark tools can use different metric definitions and cautions that results should be compared only when definitions align. Its Triton documentation also says example performance depends on the GPU used. The Triton TensorRT-LLM benchmarking example allows operators to use datasets or generated token-length distributions and control request rate. Check the example’s full conditions before treating its output as a general expectation.

What published throughput examples do—and do not—tell you

Vendor examples illustrate why throughput figures need context. They are not directly comparable unless their hardware, workload, load pattern, metric definitions, and measurement boundaries align.

Published example Scope and qualification
About 8,000 output tokens per second Databricks reports this as the approximate plateau in its provisioned-throughput endpoint benchmarking example as concurrency rises. The result reflects that example’s worker and parallel-request capacity, not a general target for other endpoints, models, or workloads. The documentation was updated 2026-09-11.
3,857.66 output tokens per second NVIDIA’s Triton TensorRT-LLM backend page labels this an expected-output example, not an independent test. The surrounding example specifies request rate, prompt and response lengths, and a 5,000-request run; NVIDIA cautions that performance depends on the GPU. The inspected page showed no update date.

For actual capacity planning, benchmark the configuration and workload you intend to run. A vendor example can help explain a method or show what one setup produced, but it cannot establish a universal tokens-per-second target.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.