Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

To serve nvidia/Qwen3.8-2.4T-A95B-NVFP4 with an NVFP4 key-value (KV) cache, add --kv-cache-dtype nvfp4 to a recent vLLM or SGLang release running on an NVIDIA Blackwell GPU. That flag sets the precision of the active KV cache during serving. It is a separate decision from the checkpoint’s NVFP4 weights, from the FP8 KV cache in NVIDIA’s quantization recipe, and from TensorRT-LLM’s NVFP4 cold-page compression, which applies to offloaded cache tiers.

The sections below cover the settings that are easiest to confuse, the model this applies to, the setup steps, the hardware and runtime support statements, what the published benchmark numbers do and do not show, and the failures you are most likely to hit.

Three settings that look alike but control different things

Several settings in a Qwen deployment can carry the words “NVFP4” or “FP8.” They are chosen at different points in the pipeline, and changing one does not change the others.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Setting What it controls Value for the cited checkpoint Where it is chosen
Model weight precision Storage format of the model weights Routed MoE experts in NVFP4; self-attention and gated-delta layers in FP8 W8A8; MTP block in BF16 Quantization of the checkpoint
Active GPU KV-cache dtype Precision of keys and values held in GPU memory while serving NVFP4 when --kv-cache-dtype nvfp4 is passed; the runtime default when the flag is omitted Serving command
Gated-delta recurrent-state dtype Precision of the recurrent state in the linear-attention layers Selected independently of weight precision (TensorRT-LLM deployment guide for a Qwen3.8-Flash-Next configuration); value for this checkpoint not stated TensorRT-LLM configuration (key name not stated)
Cold-tier KV compression Precision of eligible attention KV stored in host or disk tiers Stored as NVFP4 while cold; restored to runtime precision before attention TensorRT-LLM tiered-cache feature

Which model this applies to

NVIDIA’s model card identifies the checkpoint as nvidia/Qwen3.8-2.4T-A95B-NVFP4, based on Qwen3.8-2.4T-A95B. It describes a Transformer Mixture-of-Experts model with hybrid attention and fine-grained MoE blocks, with 2.4T total parameters and 95B activated. The card lists its release date as 2026-08-27.

#1 Best Overall
NVD RTX PRO 6000 Blackwell Professional Workstation Edition Graphics Card for AI, Design, Simulation, Engineering - 96GB DDR7 ECC Memory - 4th Gen RT/5th Gen Tensor Core GPU - OEM Packaging
  • PLEASE NOTE: Exporting an NVIDIA RTX Pro 6000 GPU outside the US requires strict adherence to the U.S. Export Administration Regulations (EAR) and issuance of an export license from the Bureau of Industry and Security (BIS). Compliance and Know Your Customer (KYC) screening may be required as a condition of order acceptance. [NVIDIA Blackwell Streaming Multiprocessor] The new SM features increased processing throughput, and new neural shaders that integrate neural networks inside of programmable shaders | DLSS 4: Multi Frame Generation ensures ultra-smooth frame pacing for lifelike simulations.
  • [Double-Flow-Through Design] The RTX PRO 6000 Blackwell features a double-flow-through cooling design, optimizing efficiency and airflow to sustain peak performance under 600W power loads. | [5th Gen Tensor Cores] Deliver up to 3X the performance of the previous generation and support for FP4 precision for faster AI model processing times with reduced memory usage, enabling local fine-tuning of LLMs and generative AI | [4th Gen Ray Tracing Cores] Double the ray-triangle intersection rate of the previous generation to create photoreal, physically accurate scenes and immersive 3D designs with RTX Mega Geometry, which enables up to 100X more ray-traced triangles.
  • [PCIe Gen 5] Support for PCIe Gen 5 provides double the bandwidth of PCIe Gen 4, improving data-transfer speeds from CPU memory and unlocking faster performance for data-intensive tasks like AI, data science, and 3D modeling. | [GDDR7 Memory] With 96 GB of GPU memory and 1.8 TB ps bandwidth, it can tackle massive 3D and AI projects, fine-tune AI models locally, explore large-scale VR environments, and drive larger multi-app workflows.
  • [DisplayPort 2.1] Achieve unparalleled visual clarity and performance, driving high resolution displays at up to 8K at 240 Hz and 16K at 60 Hz. Increased bandwidth enables seamless multi-monitor setups while HDR and higher color depth support ensures superior color accuracy for precision work, such as video editing, 3D design, and live broadcasting.
  • [Universal MIG] Divide a single RTX PRO 6000 Blackwell into multiple isolated instances, each with dedicated resources, allowing for concurrent execution of multiple workloads, optimized GPU utilization, and secure isolation of different applications or users. [WARRANTY] 3 YR Manufacturer's Warranty. Bulk OEM Packaging. Retail Packaging is NOT included.

“Hybrid” here has a specific meaning. According to the Model Optimizer recipe for this checkpoint, gated-delta (linear-attention) layers are interleaved with full-attention layers. Other hybrid Qwen models do not necessarily share this layout or its support status, so take flags and support statements from the model card of the checkpoint you actually run.

Serving the checkpoint with an NVFP4 KV cache

Requirements

NVIDIA’s model card states: “NVFP4 KV cache requires a recent vLLM or SGLang release with NVFP4 KV support and an NVIDIA Blackwell GPU.” The card does not give a minimum runtime version, so check the release notes for the version you install.

  • An NVIDIA Blackwell GPU. The TensorRT-LLM matrix lists sm100 and sm103.
  • A recent vLLM or SGLang release that includes NVFP4 KV support.
  • The checkpoint nvidia/Qwen3.8-2.4T-A95B-NVFP4. The model card is pinned to a repository revision, so confirm that the revision you download matches the details in this article.

Launch commands

These are the example forms from the model card. They show the flags; they do not guarantee that a given release, model configuration, or GPU will run them. This article has not run them on its own hardware.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
vllm serve nvidia/Qwen3.8-2.4T-A95B-NVFP4 \
  --port 8000 \
  --tensor-parallel-size 8 \
  --max-model-len 262144 \
  --kv-cache-dtype nvfp4 \
  --reasoning-parser qwen3
python -m sglang.launch_server \
  --model-path nvidia/Qwen3.8-2.4T-A95B-NVFP4 \
  --port 8000 \
  --tp-size 8 \
  --context-length 262144 \
  --kv-cache-dtype nvfp4 \
  --reasoning-parser qwen3

Setup steps

  1. Confirm the GPU. Run nvidia-smi --query-gpu=name --format=csv, then map each reported name to its architecture in NVIDIA’s product documentation. The flag requires a Blackwell GPU.
  2. Confirm the runtime release. If you installed through pip, run pip show vllm or pip show sglang, then check that version’s release notes for NVFP4 KV support.
  3. Launch with the flag. Start from one of the commands above, adjusting the port, tensor-parallel size, and context length to your hardware.
  4. Read the startup output. If the runtime rejects the flag or reports a KV-cache dtype other than the one you requested, stop and fix that before measuring anything.
  5. Take a baseline if you need one. Run the same command without --kv-cache-dtype. Omitting the flag uses the runtime’s default KV-cache precision, so comparing the two runs isolates the cache dtype.

Hardware and runtime support differ by source

The support statements below come from different documents with different scopes. They reflect documentation as of early October 2026. Runtime matrices change between releases, so check the version you deploy.

Rank #2
NVIDIA RTX PRO 4000 Blackwell Graphics Card - 24GB GDDR7 ECC Memory, PCIe 5.0 x16, 4X DisplayPort 2.1b, Single Slot Full Height AI Workstation GPU, Retail Packaging
  • Professional GPU with Blackwell Architecture
  • Blackwell Architecture
  • 24GB GDDR7 with PCIe 5.0 & Ray Tracing
  • AI Workstation
Source Qwen-3 NVFP4 KV Blackwell (sm100/103) Hopper or Ada
NVIDIA model card, usage examples (2026) Used in the Qwen3.8-2.4T-A95B-NVFP4 vLLM and SGLang examples Required Not covered; the card requires Blackwell
TensorRT-LLM quantization documentation (current rolling page) Listed Supported No NVFP4 KV support marked
vLLM Not stated Not stated Not stated
SGLang Not stated Not stated Not stated

The TensorRT-LLM documentation says its NVFP4 KV checkpoint-generation flow currently requires FP8 weight and activation quantization. That is a condition of generating your own checkpoint for TensorRT-LLM; it is not stated as a condition of the vLLM or SGLang commands above. Do not assume that support in one runtime carries over to another.

The checkpoint recipe is a separate, weight-side choice

NVIDIA’s Model Optimizer recipe for this checkpoint assigns a precision to each component. It is a quantization recipe, not a serving configuration:

  • Routed experts: NVFP4
  • Self-attention and gated-delta linear-attention components: FP8 W8A8
  • KV cache: FP8 cast
  • Remaining components, such as the MTP block: BF16

The FP8 KV-cache entry describes how the checkpoint was produced. It does not show that the runtime keeps the active cache in NVFP4 by default. Using NVFP4 KV is a serving choice, and the model card’s usage examples make it explicitly with --kv-cache-dtype nvfp4.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cold-page compression is a different feature

TensorRT-LLM’s cold-page compression stores eligible attention KV in NVFP4 in host-memory or disk cold tiers. The active GPU cache stays in its ordinary runtime type, such as FP16, BF16, or FP8. When a cold page is needed for attention, it is restored to the runtime representation first. TensorRT-LLM’s documentation explicitly distinguishes this from setting the active GPU KV cache to NVFP4.

Rank #3
PNY VCNRTXPRO2000B-PB NVIDIA RTX PRO 2000 Blackwell 16GB GDDR7 128B Graphics Cards
  • Form Factor: Plug-in Card
  • Cooler Type: Active Cooler
  • Maximum Power Consumption: 70W
  • Length: 6.6
  • Height: 2.7

Because this is a TensorRT-LLM tiered-cache feature, the vLLM and SGLang flags in this article do not configure it.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the published benchmark numbers show

NVIDIA’s model card reports scores for three configurations: BF16, NVFP4 weights, and NVFP4 weights with an NVFP4 KV cache. They were produced under the card’s sampling settings (temperature 1.0, top-p 0.95, top-k 20), with maximum new tokens of 65,536 for GPQA Diamond, SciCode, AA-LCR, and IFBench; 131,072 for HLE; and 262,144 for Terminal Bench 2.1. These are publisher-reported figures from 2026, not independent measurements. The difference column is simple subtraction of the card’s values.

Benchmark BF16 NVFP4 NVFP4 + NVFP4 KV NVFP4 + KV minus BF16
GPQA Diamond 92.55 92.58 92.33 −0.22
HLE 41.43 40.55 40.64 −0.79
SciCode 54.44 56.21 55.92 +1.48
AA-LCR 71.5 71.63 71.25 −0.25
IFBench 79.93 81.73 81.33 +1.40
Terminal Bench 2.1 76.03 76.4 77.25 +1.22

On these six benchmarks, the NVFP4 KV configuration lands within 1.5 points of BF16 in either direction. The gaps are small, and this article has no run-to-run variance data for them, so none should be read as a reliable gain or loss.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The benchmarks do not establish that NVFP4 KV is accuracy-neutral for your workload. The Hugging Face PTQ documentation notes that accuracy loss after post-training quantization varies by model and quantization method. If accuracy does not meet your requirement, it suggests changing or disabling KV quantization, or using quantization-aware training (QAT).

Quick Recap

Troubleshooting

  • The runtime rejects --kv-cache-dtype nvfp4. The installed release may predate NVFP4 KV support. Upgrade, then re-check the version with the commands in the setup steps.
  • The workload runs on Hopper or Ada. The card requires Blackwell for NVFP4 KV, and the TensorRT-LLM matrix marks no NVFP4 KV support on those architectures. Move the workload to Blackwell hardware. Remove the flag only after confirming that the checkpoint itself runs on that GPU, because the card’s requirement is stated for the NVFP4 KV cache.
  • Output quality falls below your requirement. Compare against the baseline run from step 5 of the setup steps before changing anything else.

|

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.