Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Benchmark an AI agent with a representative workload, a documented serving configuration, a warm-up, and a concurrency sweep that continues through saturation. Report total-system output tokens per second (TPS) alongside latency and GPU count. If you divide TPS by GPU count, label it as a simple per-GPU average—not single-GPU performance or scaling efficiency.

1. Define a workload that resembles your agent

A benchmark is only useful for your deployment if its requests resemble the work your agent performs. Record the model and version, tokenizer, input and output length distributions, number of turns, how context grows across turns, tool-use pattern, and generation settings.

Where possible, use representative multi-turn or coding-and-tool traces rather than a single-turn chat prompt with fixed input and output lengths. The September 28, 2026 AgentPerfBench preprint argues that fixed-length, single-turn tests can miss important agent behavior. It describes workload profiles based on empirical per-turn input lengths, output lengths, and turn-count distributions; its approach is recent research, not a universal benchmark standard.

  • Preserve the sequence of turns and the changing prompt context, not just the first request.
  • Include tool interactions if they are part of the deployment workload, and document how they affect requests and turn timing.
  • Record the sampling and decoding settings so a change in generation behavior is not mistaken for a hardware throughput change.

2. Document the serving configuration

Record the system details needed to reproduce and interpret the result. At minimum, include:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Elebase USB to USB C Adapter for iPhone 18 Pro Max,USBC Car Charger Adapter
  • Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
  • Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
  • Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
  • Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
  • 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
  • GPU model and number of GPUs.
  • Serving engine and version.
  • Precision or quantization, parallelism, batching, and other relevant model-serving settings.
  • Model and tokenizer versions, plus generation settings.
  • Client and server placement, including relevant network placement.

NVIDIA documents AIPerf as a client-side benchmarking tool for OpenAI-compatible inference services. Its guide recommends running the client on the same host when network latency is not part of the test. If network behavior matters to your deployment, keep it in the test and document the placement rather than treating that latency as model-generation time.

3. Warm up and preserve the measurement artifacts

Run a warm-up before collecting measured results so initialization and other startup effects do not define the reported operating point. NVIDIA’s AIPerf example includes a warm-up, a concurrency sweep, and JSON and CSV output. Preserve the result files together with the benchmark configuration and the exact command or invocation used. The files make it easier to verify settings and compare later runs.

Use a client and benchmark setup that can sustain the offered load; otherwise, the client can limit measured throughput. Keep the workload and configuration fixed when comparing systems, and change one relevant factor at a time when diagnosing a difference.

Rank #2
Sale
Anker USB-C Hub, 5-in-1 USB Hub for Laptops, 4K HDMI Multiport Adapter
  • 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
  • 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
  • Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
  • 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
  • What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.

4. Sweep concurrency through saturation

Measure multiple load levels rather than choosing one concurrency value and treating it as the system’s capacity. Start with values representative of expected deployment, then extend the sweep until aggregate throughput stops meaningfully increasing. NVIDIA advises using concurrency for most benchmarks and notes that throughput can saturate while latency continues to rise.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Concurrency and request rate are both ways to control offered load. For a concurrency sweep, each point represents a different number of requests allowed in flight; for a request-rate test, the client offers requests at a specified rate. State which policy you used. Do not compare results as if the load were equivalent when one test used concurrency and the other used a request-arrival rate.

At each load level, collect throughput and latency. A plot with total-system TPS on one axis and a user-facing latency measure on the other, with each point labeled by concurrency, makes the trade-off visible. Choose the point that meets your deployment’s latency budget, then report its throughput and concurrency. AIPerf’s guide describes this latency-throughput interpretation and supports alternative latency axes, including inter-token latency, end-to-end latency, and TPS per user.

Rank #3
Sale
Anker USB C Hub, 7in1 Multi-Port USB Adapter, 4K@60Hz USBC to HDMI Splitter
  • Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
  • Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
  • Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
  • Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
  • What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.

5. Report metrics with their definitions

Include more than a peak tokens-per-second figure. Report total output TPS, requests per second (RPS), time to first token (TTFT), inter-token latency (ITL) or time per output token (TPOT), and end-to-end latency. Give averages and relevant tail percentiles when the tool provides them, and identify the statistic used. Metric implementations can differ across tools; NVIDIA notes, for example, that definitions of ITL can differ on whether TTFT is included.

Metric What it describes Interpretation
Total output tokens per second (TPS) NVIDIA defines system TPS as total output-token throughput across simultaneous requests. AIPerf calculates output tokens over the interval from the first request to the final response; configured warm-up can be excluded. Aggregate system throughput for the measured interval, not throughput for one request or one GPU.
TPS per user For an individual request, output sequence length divided by end-to-end latency. A single-client perspective; it is not aggregate system TPS.
Time to first token (TTFT) Time from query submission until the first received output token, when the response contains content. Shows how long a user waits for generation to begin.
Inter-token latency (ITL) or time per output token (TPOT) Average time between consecutive output tokens. Tool definitions differ on whether TTFT is included; AIPerf excludes it. Describes token-generation pacing after the first token under the metric’s stated definition.
End-to-end latency Time from query submission to complete response, including queueing, batching, and network latency. Captures the full request experience measured by the client.
Requests per second (RPS) Successful requests completed per second over the benchmark interval. Useful alongside TPS because requests can have different output lengths.

NVIDIA’s metrics documentation, last updated July 20, 2026, defines total TPS per system as output-token throughput across all simultaneous requests. Retain the tool and metric definition with reported values, especially when comparing results from different serving stacks.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

6. Calculate and label per-GPU throughput carefully

Keep total-system TPS as the primary result. If a per-GPU average is useful, calculate it transparently:

Rank #4
Sale
UGREEN USB to USB C Adapter Combo 4-Pack, 10Gbps USB C Converter Space Gray
  • Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
  • Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
  • Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
  • Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
  • Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft

Per-GPU average TPS = total-system TPS ÷ number of GPUs

State the total-system TPS, GPU count, and calculation together. This arithmetic average does not establish how one GPU would perform on its own: multi-GPU parallelism, batching, and system design affect aggregate throughput. It also is not scaling efficiency, which would require a defined comparison against a baseline system and its scaling behavior.

Do not use the per-GPU average as a standalone score when comparing unlike configurations. A system’s GPU count and parallelism can change how work is distributed, so retain the complete configuration and compare throughput at a stated latency constraint.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
Anker USB C Hub, 5-in-1 USBC to HDMI Splitter with 4K Display
  • 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
  • Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
  • Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
  • HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
  • What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

7. Make comparisons on aligned workloads and constraints

For comparisons between GPUs or serving systems, align the workload and report any differences explicitly. Compare points at the same latency target or show the full load curves; the highest TPS observed without its latency and concurrency context is not enough to establish which system better fits a deployment.

  • Model, model version, and tokenizer.
  • GPU model, GPU count, and parallelism configuration.
  • Serving framework and version, precision or quantization, and decoding settings.
  • Input and output length distributions, agent turn pattern, and tool-use behavior.
  • Concurrency or request-arrival policy, measurement duration, and latency target.
  • Total-system throughput, latency metrics and statistics, and any per-GPU arithmetic.

MLPerf provides standardized inference evaluations across model architectures and scenarios, while a custom trace-based benchmark can better match a particular agent deployment. These answer different comparison needs: use standardized results for like-for-like evaluated scenarios, and deployment-representative traces to understand the behavior of your own workload.

8. Interpret published figures in their stated scope

Published results should not be generalized beyond their submitted systems and workloads. NVIDIA reported that Vera Rubin NVL72 achieved up to 3.7 times the throughput of GB300 NVL72, and that a 288-GPU GB300 NVL72 submission reached 99% scaling efficiency, in its MLPerf Inference v6.1 results. NVIDIA’s page says those results were retrieved from MLCommons on September 16, 2026. These are vendor-reported figures for the specified submissions, not a general GPU comparison or a substitute for benchmarking an agent workload.