Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

You can run local AI models on an NVIDIA DGX Spark by installing or choosing a compatible inference runtime, selecting model weights and precision that fit your workload, then running the model through a local interface such as a command line or API server. For a guided agent setup, NVIDIA also documents a NemoClaw path that configures Ollama with a local model. First complete DGX Spark setup and check NVIDIA’s current software guidance; runtime and installer details can change.

Model runtime or agent: what are you setting up?

A model runtime loads model weights and generates inference. An agent harness adds workflows and tools around a model, and may also connect to external services. You can run inference locally without installing an agent, and installing an agent does not by itself guarantee that all of its activity stays local.

NVIDIA’s documented NemoClaw flow combines a local Ollama model with an agent harness and sandbox runtime. It is one guided option, not a prerequisite for running models on DGX Spark.

Prepare DGX Spark before installing a runtime

Complete first boot, then consult NVIDIA’s DGX Spark documentation hub for current first-boot, OS and component update, release-note, and recovery guidance. Check those instructions before changing drivers, CUDA, or other system software. The available guidance does not establish a single current point-release version for DGX OS, CUDA, or every runtime, so use the current official instructions rather than relying on version-specific commands from an older setup.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Dell NVIDIA Tesla V100 GPU SXM2 32GB NWWWX by DELL
  • GPU Chipset: NVIDIA
  • Memory: HBM2
  • Programming Interface: CUDA
  • Memory Capacity: 32GB
  • Slot Compatibility: SXM2

Choose a runtime for your model and workload

NVIDIA lists Ollama, llama.cpp, TensorRT, SGLang, vLLM, and PyTorch with CUDA as local AI runtime options. There is no evidence here to rank them by speed for every workload. Compare the model format you have, API needs, setup and maintenance effort, GPU and memory behavior, and target throughput.

Runtime When it may fit What to consider
Ollama A straightforward local model setup; it is also the runtime used in NVIDIA’s NemoClaw express flow. Follow the current Ollama and DGX Spark instructions for installation and model availability. NVIDIA’s guided flow is agent-specific.
llama.cpp Running GGUF model files through a command-line interface or local server. Choose a compatible GGUF checkpoint and verify current CUDA and DGX OS requirements. Community build recipes are not a substitute for current official system guidance.
vLLM, SGLang, TensorRT, or PyTorch with CUDA More configurable accelerated inference deployments, selected according to model and serving needs. Compare model-format support, API requirements, deployment effort, GPU and memory behavior, and throughput goals; do not assume one is best without workload-specific evaluation.

For the chosen backend, use its current installation guide alongside NVIDIA’s DGX Spark documentation. Exact commands and supported combinations depend on software versions and model format; the cited NVIDIA guidance does not establish one universal installation sequence for all six runtimes.

Select weights, quantization, and context size

NVIDIA describes DGX Spark as having 128 GB of unified memory and up to 1 petaFLOP at FP4. Its product guidance also states inference support for models up to 200 billion parameters. These are NVIDIA capability claims, not guarantees that every model at that size will fit or perform well: memory use depends on the checkpoint’s precision, context length, runtime, and workload.

NVIDIA’s current local AI guidance suggests Q4_K_M quantization as a starting point for llama.cpp, and NVFP4 for vLLM or PyTorch. Treat these as recommendations to evaluate, not universal compatibility guarantees. Quantization can affect both memory use and output quality.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Gigabyte NVIDIA GeForce RTX 3060 Gaming OC V2 Graphics Card - 12GB GDDR6, 192-bit, PCI-E 4.0, 1837MHz Core Clock, RGB, 2X DP 1.4, 2X HDMI 2.1, NVIDIA Ampere - GV-N3060GAMING OC-8GD
  • NVIDIA Ampere Streaming Multiprocessors: Building blocks for the world's fastest, most efficient GPUs, the all-new Ampere SM brings twice the FP32 throughput and improved energy efficiency
  • 2nd Generation RT Cores - Experience 2x the 1st Generation RT Cores throughput, plus competitive RT and shading for a whole new level of ray-tracing performance
  • 【3rd Generation Tensor Cores】Get up to 2X the throughput with structural sparsity and advanced AI algorithms such as DLSS
  • Core Clock: 1837MHz
  • WINDFORCE 3X Cooler
  1. Match the checkpoint to the backend. For example, llama.cpp commonly uses GGUF weights; confirm the specific model and file format are supported by your selected runtime.
  2. Estimate memory for the full workload. Consider weights, context length, and runtime overhead rather than comparing parameter count with unified memory alone.
  3. Test with representative prompts or a task-specific dataset. NVIDIA recommends evaluating model quality on relevant data with human review, rather than choosing solely from a size or performance claim.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Optional: install a guided local agent with NemoClaw

NVIDIA’s June 1, 2026 walkthrough describes an agent-oriented setup that uses NemoClaw and OpenShell, with Ollama configured to download Qwen3.6-35B during express installation. It is a version-sensitive guided path; open the NVIDIA Technical Blog article and its linked Spark playbook for the current steps and installer before proceeding.

  1. Complete DGX Spark first boot and open the current NVIDIA Spark playbook.
  2. Run the NemoClaw install command shown in that guide: curl -fsSL https://www.nvidia.com/nemoclaw.sh | bash. This command installs software; the express flow also downloads model weights.
  3. Accept the licenses requested by the installer and choose express install if you want the documented Ollama and Qwen3.6-35B setup.
  4. Use the gateway token to open the agent Web UI, following the playbook’s current instructions.

NVIDIA’s June 2026 blog reports up to 2.6× faster inference for Qwen3.6-35B using its NVFP4 checkpoint and vLLM optimizations. This is NVIDIA’s reported result for that configuration, not an independent benchmark or a general speed guarantee.

Check privacy and network access for agent setups

Local inference means the model runs on the workstation; it does not establish that an agent’s entire workflow is offline. NVIDIA describes OpenShell as providing sandboxing, access controls, privacy protections, and operational guardrails, while its NemoClaw walkthrough also discusses integrations and configurable external network destinations.

Quick Recap

Bestseller No. 1
Dell NVIDIA Tesla V100 GPU SXM2 32GB NWWWX by DELL
Dell NVIDIA Tesla V100 GPU SXM2 32GB NWWWX by DELL
GPU Chipset: NVIDIA; Memory: HBM2; Programming Interface: CUDA; Memory Capacity: 32GB; Slot Compatibility: SXM2
$854.96
  • Review the agent’s network policy and allow only the destinations it needs.
  • Check which integrations are enabled and whether they send prompts, files, or results to external services.
  • Limit the files, tools, and system resources the agent can access to what its task requires.
  • Do not assume that a local model prevents data from leaving the machine if connected tools or integrations transmit it.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.