Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

For one NVIDIA DGX Spark, start with NVIDIA’s current single-device vLLM recipe rather than assembling a launch command from older examples. It targets one GB10 GPU and recommends Qwen3.8-27B NVFP4. The main operational caution is shared memory: the model, serving process, operating system, container runtime, and KV cache all draw from the same pool.

Use the official one-Spark recipe as your starting point

NVIDIA’s playbook distinguishes a single-device setup from multi-device configurations. Choose One DGX Spark for a server running on the GB10 in one system; the recommended model is Qwen3.8-27B NVFP4. NVIDIA says the model fits one Spark with a hardware-specific vLLM configuration.

Use the recipe’s launch tab to obtain the actual model, container, and serving settings. NVIDIA says, “The launch tabs provide the complete model, container, and serving configuration.” The page warns that its copy-and-paste configurations are tested for the recommended recipes; another model or quantization can require different download, container, environment, memory, parser, or parallelism settings. Follow the matching tab rather than combining pieces from unrelated examples: NVIDIA Build: Serve LLMs with vLLM.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The available recipe description does not expose the complete command payload, so an exact launch command cannot be reproduced reliably here. Copy the command from the live launch tab and retain its model and container versions alongside your deployment notes. This is especially important as vLLM images and hardware-specific support change.

#1 Best Overall
Dell NVIDIA Tesla V100 GPU SXM2 32GB NWWWX by DELL
  • GPU Chipset: NVIDIA
  • Memory: HBM2
  • Programming Interface: CUDA
  • Memory Capacity: 32GB
  • Slot Compatibility: SXM2

Leave room in unified memory

DGX Spark’s CPU and GPU use a coherent unified memory pool. In the 128 GB configuration, that pool is shared by the operating system, container runtime, vLLM process, model weights, and KV cache. The KV cache can grow as requests and active context grow, so assigning nearly all available memory to model serving can cause failures even when the weights initially load.

The vLLM guidance recommends leaving headroom when setting --gpu-memory-utilization; treat the value in NVIDIA’s launch recipe as specific to that recipe, not as a universal optimum. Likewise, keep --max-num-seqs aligned with your workload. vLLM recommends keeping it low for small-batch inference on Spark. Raising concurrency may increase aggregate throughput, but it changes the workload and increases memory pressure; it is not a like-for-like improvement in single-request speed. See vLLM’s DGX Spark guidance.

Check the hardware capacity before using a recipe or benchmark

NVIDIA identifies DGX Spark as a GB10 Grace Blackwell system. Its product specifications list 128 GB of coherent unified memory, plus a 64 GB configuration available through participating OEM partners. Confirm which capacity you have before assuming that a model configuration or memory setting will fit. The 128 GB recipe and the 64 GB partner system are not interchangeable simply because both use GB10.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

NVIDIA also lists 273 GB/s memory bandwidth and up to 1 PFLOP of FP4 tensor performance. Those are hardware specifications, not vLLM token-generation results; they do not predict the serving speed of a particular model. See NVIDIA DGX Spark specifications.

Interpret the reported vLLM timings by workload

There is no single meaningful “DGX Spark tokens per second” figure without specifying the model, quantization, vLLM and container versions, runtime flags, prompt and output lengths, concurrency, and whether reasoning tokens are included. Also distinguish per-request decode speed from aggregate throughput across requests, and decode speed from time to first token.

Rank #2
Gigabyte NVIDIA GeForce RTX 3060 Gaming OC V2 Graphics Card - 12GB GDDR6, 192-bit, PCI-E 4.0, 1837MHz Core Clock, RGB, 2X DP 1.4, 2X HDMI 2.1, NVIDIA Ampere - GV-N3060GAMING OC-8GD
  • NVIDIA Ampere Streaming Multiprocessors: Building blocks for the world's fastest, most efficient GPUs, the all-new Ampere SM brings twice the FP32 throughput and improved energy efficiency
  • 2nd Generation RT Cores - Experience 2x the 1st Generation RT Cores throughput, plus competitive RT and shading for a whole new level of ray-tracing performance
  • 【3rd Generation Tensor Cores】Get up to 2X the throughput with structural sparsity and advanced AI algorithms such as DLSS
  • Core Clock: 1837MHz
  • WINDFORCE 3X Cooler

Single-node gpt-oss-120b report

Conselara Labs reported results on May 9, 2026, using vLLM 0.19.0 in NVIDIA’s nvcr.io/nvidia/vllm:26.04-py3 container for a single-node gpt-oss-120b setup. It reported approximately 32–35 generated tokens per second with reasoning overhead included, and 57–60 tokens per second for pure decode. The reported configuration used TP=1, MXFP4 quantization, FP8 KV cache, Triton attention, a Marlin MoE backend, and a 128,000-token context. These figures describe that configuration, not NVIDIA’s Qwen3.8-27B NVFP4 recipe or every DGX Spark deployment. The report does not establish a universal speed expectation: Conselara Labs’ DGX Spark vLLM benchmark.

Two-node Qwen3-235B result

The same lab separately reported a two-node Qwen3-235B-A22B GPTQ-Int4 setup: 17.0 tokens per second at batch 1, 24.1 aggregate tokens per second at batch 2, and 36.4 aggregate tokens per second at batch 4. It reported about 15 minutes from startup to first inference. This is a two-node result, not a one-Spark benchmark; the batch 2 and batch 4 figures are aggregate rates rather than per-request decode speeds.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Community MXFP4 observation

A participant in the NVIDIA Developer Forums reported about 59 generated tokens per second for single-user GB10 MXFP4 generation and said concurrency provided additional throughput headroom. The excerpt does not provide enough benchmark detail to equate that observation with the lab results or treat it as an independently reproduced measurement: NVIDIA Developer Forums discussion.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Do not treat older build advice as current compatibility guidance

A December 2025 forum post said NVIDIA’s Docker image lagged mainline vLLM at that time, that newer versions might require a source build or community image, and that MXFP4 was not then using Blackwell FP4 features on Spark in vLLM. Those are dated community observations, not a statement of current official compatibility. Check the date and version of any build advice against NVIDIA’s current recipe before relying on it.

Kernel and performance behavior also depend on the model and vLLM release. The vLLM team’s current guidance generally favors CUDA graphs unless a deployment has a reason to disable them; this is not a blanket promise that a particular NVFP4 or MXFP4 model will be faster. See vLLM’s DGX Spark article.

Quick Recap

Bestseller No. 1
Dell NVIDIA Tesla V100 GPU SXM2 32GB NWWWX by DELL
Dell NVIDIA Tesla V100 GPU SXM2 32GB NWWWX by DELL
GPU Chipset: NVIDIA; Memory: HBM2; Programming Interface: CUDA; Memory Capacity: 32GB; Slot Compatibility: SXM2
$854.96

A practical checklist before serving requests

  • Select the one-device DGX Spark recipe and copy its complete launch configuration from NVIDIA’s live tab.
  • Confirm whether the system has 128 GB or the 64 GB partner configuration; do not carry memory assumptions across them.
  • Keep the recipe’s model, container, and serving settings together, and record their versions.
  • Leave unified-memory headroom for the OS, runtime, and KV-cache growth; avoid treating a memory-utilization setting as universally correct.
  • Set concurrency for the intended workload. Small-batch inference and high-concurrency aggregate throughput are different use cases.
  • When reporting a benchmark, state the model and quantization, software/container version, relevant runtime configuration, context and request conditions, concurrency, reasoning inclusion, and whether the rate is per-request or aggregate.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.