Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Tenstorrent TT-QuietBox 2 is a liquid-cooled desktop AI workstation built around four Blackhole AI processors, with an AMD Ryzen CPU, 256 GB of DDR5 system memory and 2 TB of NVMe storage. Tenstorrent lists it at $9,999 with an estimated 10–12 week shipping time. Its 128 GB of accelerator memory is the key capacity figure for local inference; Tenstorrent says the system can load GPT-OSS-120B and reports nearly 500 tokens per second on Llama 3.1 70B, but that speed is a vendor claim, not an independent benchmark.

What is the Tenstorrent QuietBox 2?

TT-QuietBox 2, also called QuietBox 2 (Blackhole), is a complete desktop workstation for running AI models locally and developing software for Tenstorrent accelerators. Rather than relying on one general-purpose graphics card, it combines four Blackhole chips with an AMD Ryzen CPU, liquid cooling, DDR5 system memory and NVMe storage. Tenstorrent positions it for local inference, experimentation and low-level library or kernel development.

The system ships with Tenstorrent’s open-source software stack. TT-Studio offers a browser-based interface for deploying local models, and TT-Inference-Server provides an OpenAI-compatible endpoint for applications that connect to that API style. For lower-level work, TT-Metalium is Tenstorrent’s development environment for custom kernels.

QuietBox 2 specifications

Specification QuietBox 2 detail
AI processors Four Blackhole chips, per Tenstorrent documentation
AI compute units 480 Tensix cores, per Tenstorrent documentation
Accelerator memory 128 GB GDDR6, per Tenstorrent documentation
Accelerator memory bandwidth 2 TB/s, per Tenstorrent documentation
System memory 256 GB DDR5, per Tenstorrent’s March 2026 newsroom article
Storage 2 TB NVMe, per Tenstorrent’s product information
CPU AMD Ryzen; the specific model is not stated in the cited Tenstorrent material
Cooling Liquid-cooled, per Tenstorrent’s product information

These figures describe different parts of the machine. The 128 GB of GDDR6 is memory on the AI accelerators; the 256 GB of DDR5 is system memory for the CPU and operating system. They are not interchangeable pools of accelerator memory. Tenstorrent co-founder and systems engineer Milos Trajkovic has emphasized the practical importance of the accelerator memory: “The 128 gigabytes of GDDR that we have with our AI accelerators really defines how big of a model you can run at a reasonable speed.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Andromeda Insights - AI Workstation Gaming PC | 2X AMD Radeon AI PRO R9700 64GB Total VRAM | Ryzen 9 9950X (5.7 GHz Turbo) | 128GB DDR5 | 4TB Gen4 SSD | W11 | Wi-Fi | Bluetooth - Black
  • Engineered for demanding AI workloads, this is your definitive development platform. It packs an AMD Ryzen 9 9950X for parallel processing and two AMD Radeon AI PRO R9700 GPUs with 64GB of combined VRAM for large models & complex neural nets. Built for sustained performance, it includes 128GB DDR5 RAM, a 4TB NVMe Gen4 SSD, and a 360mm AIO liquid cooler for ultimate thermal stability.
  • Industry-Leading Warranty & US Support - Backed by a 2-Year Parts Warranty, Lifetime Labor Warranty & Lifetime Technical Support. Andromeda Insights is a US-based company dedicated to high-performance hardware and long-term service.
  • Flagship CPU Power with Liquid Cooling – AMD Ryzen 9 9950X | 16 Cores, 32 Threads - up to 5.7GHz Turbo – chews through LLM serving, data prep, compiles and renders. A 360mm AIO liquid cooler keeps it sustained under full load.
  • Ultra-Fast 128GB DDR5 6000MHz RAM - Multi-task effortlessly and keep large contexts, datasets and containers in memory with 128GB of blazing-fast DDR5.
  • Two AMD Radeon AI PRO R9700 GPUs give you 64GB of combined VRAM - hold 70B-class quantized models fully in GPU memory. RDNA 4 Architecture with 2nd-gen AI Accelerators, purpose-built for local LLM inference with no per-token API costs.

Can QuietBox 2 run a 70B or 120B model?

70B models

Tenstorrent says Llama 3.1 70B can run on QuietBox 2 at nearly 500 tokens per second. That is a vendor-reported result; Tenstorrent’s cited material does not establish an independent, standardized benchmark or provide a matching test against other workstations. Actual throughput can depend on the model format, quantization, runtime, workload and generation settings, so treat the figure as an indication of Tenstorrent’s claim rather than a guaranteed result for every setup.

120B models

Tenstorrent’s March 2026 newsroom article says the configuration can load OpenAI GPT-OSS-120B. “Can load” establishes a model-capacity claim, not a particular response speed or a promise that every 120B model will fit and run usefully. Model architecture, numerical format, runtime support and memory use matter; a 120B model with different characteristics may have different requirements.

Rank #2
Acer Veriton AI Mini Workstation Personal Computer
  • Experience the raw power of the NVIDIA GB10 Grace Blackwell Superchip. Delivering 1 PFLOPS of FP4 AI performance, this workstation handles 200B+ parameter models locally with sparsity. This is the same architecture powering the world’s most advanced data centers, brought directly to your desk for zero-latency development.
  • Pre-installed with NVIDIA DGX OS, the GN100 is tuned for the full NVIDIA AI stack—CUDA, PyTorch, NIM microservices, and the NeMo Framework. The NVIDIA GB10 Grace Blackwell Superchip pairs a 20-core Arm CPU with a Blackwell GPU featuring fifth-generation Tensor Cores, delivering 1 PFLOP of FP4 AI performance with sparsity. Prototype reasoning models locally and deploy to DGX cloud or data centers with zero code changes.
  • Eliminate the bottleneck between CPU and GPU. The GN100 unified memory architecture lets the Blackwell GPU and 20-core Arm CPU access a shared 128GB pool of LPDDR5X-8533 memory over NVLink-C2C—coherent, addressable, and bottleneck-free. This architecture enables 200B+ parameter models to run locally on hardware that would choke a standard desktop, providing the capacity and bandwidth required for real-time inference at scale.
  • Two 200Gbps ConnectX-7 ports. Direct-attach a second GN100 for 405B-parameter inference. Add a RoCE 200 GbE switch and link up to four units in a high-speed cluster—the standard configuration for university labs and B2B teams scaling distributed training. Combined with 128GB of LPDDR5X coherent unified memory per node, the GN100 scales as your models scale. Quiet luxury, server-class throughput.
  • For proprietary models and regulated datasets, every byte stays on-device. The GN100 ships with a 4TB self-encrypting NVMe SSD, an integrated Kensington lock, and a tamper-resistant 1.2kg sealed chassis. Pair with NVIDIA NemoClaw for sandboxed agentic workflows and policy-based privacy controls. Build, fine-tune, and run sensitive workloads without a single packet leaving your lab.

For either size, ask whether the software stack supports the exact model and format you intend to use, and whether its performance is acceptable for your workload. Capacity and token-generation speed are separate questions: a model may fit but generate too slowly for interactive use.

Is QuietBox 2 really RISC-V?

Tenstorrent describes Blackhole as part of its RISC-V AI chip family. In QuietBox 2, the RISC-V reference is to the Blackhole AI processors; the workstation also has an AMD Ryzen CPU. It would therefore be misleading to describe the entire computer as a RISC-V CPU desktop. Its distinctiveness is the use of Tenstorrent’s Blackhole AI accelerators, paired with a conventional AMD CPU.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
WEELIAO MAXSUN Intel Arc Pro B60 48G Turbo Workstation Graphics Card
  • Massive 48GB VRAM for Large AI Models: Innovative dual-GPU design combines two Arc Pro B60 GPUs, with 48GB of GDDR6 memory on a 192-bit bus (456 GB/s bandwidth). This allows you to run 70B-class quantized models like DeepSeek-R1:70B or QwQ-32B entirely on a single card, eliminating the need for multi-card setups or cloud services
  • Dual GPU Compute Power: Each GPU operates at 2400 MHz with 20 Xe cores, delivering 197 TOPS (INT8) per GPU – a combined total of 394 TOPS. This architecture is purpose-built for high-concurrency inference, multi-turn dialogues, and complex AI workloads, with each chip separately recognized by the system for flexible task assignment
  • Consumer-Friendly PCIe Configuration: Uses a PCIe 5.0 x8 + PCIe 5.0 x8 interface. When paired with a motherboard that supports x16 lane bifurcation, it achieves full bandwidth on standard consumer platforms, significantly lowering the total system cost for local LLM deployment
  • Reliable Cooling for Sustained Loads: The Turbo Edition features a triple-thermal design with a blower fan, large vapor chamber, and metal backplate. This ensures efficient heat dissipation in server airflow environments, maintaining stable temperatures and consistent performance during long, uninterrupted inference tasks
  • Broad Software & ISV Support: Native support for PyTorch, IPEX-LLM, vLLM, and standard ISV applications. The card is compatible with a wide range of open-source models including Qwen3-32B, Qwen3-VL, and DeepSeek series. It also supports SR-IOV virtualization for flexible resource allocation across tasks

What software and workloads does it support?

Tenstorrent describes the QuietBox 2 software offering as an open-source stack with tools at different levels:

  • TT-Studio: a browser UI for deploying local models.
  • TT-Inference-Server: an OpenAI-compatible endpoint for connecting supported applications and workflows.
  • TT-Metalium: tools for developers working on custom kernels and lower-level accelerator software.

Tenstorrent’s onboarding materials describe private LLM inference, coding assistants, local agents, text-to-video, image generation and custom kernel work. Those examples indicate intended workflows, not a guarantee that every model, feature or framework works without configuration. Model support and software versions can change; check Tenstorrent’s current compatibility and onboarding materials for the specific workload before buying.

Rank #4
MINISFORUM MS-S1 MAX Mini AI Workstation PC, AMD Ryzen AI Max+ 395 (16C/32T),RDNA3.5 GPU,128GB LPDDR5x RAM 2TB SSMINI PC, Dual M.2 PCIe 4.0,PCIe x16 Slot, USB4 V2(80Gbps)& Dual 10GbE, 320W PSU,Wi-Fi 7
  • 【High-Performance APU】The MS-S1 MAX features an AMD Ryzen AI Max+ 395 APU, integrating a Zen 5 architecture CPU (up to 5.1GHz, 16C/32T, 64M L3 Cache), an RDNA 3.5 GPU, and an NPU (50 TOPS). The total system output is 126 TOPS. It provides powerful parallel computing capabilities for demanding AI workflows. It is ideal for running local LLMs, multimodal models, and computationally intensive tasks
  • 【128GB UMA Memory】Equipped with up to 128GB of LPDDR5x-8000MT/s unified memory, it enables the CPU and GPU to access a shared, high-bandwidth memory pool with extremely low latency. Ideal for large-scale AI inference, 3D workloads, and complex timelines in video editing. It eliminates traditional VRAM bottlenecks, ensuring smoother data transfer during high-intensity computations. The UMA design maximizes performance stability under high loads
  • 【Flexible Expansion】The MS-S1 MAX features USB4 V2 (up to 80Gbps), dual 10GbE LAN, HDMI 2.1 (up to 8K60), a full-length PCIe x16 expansion slot, and dual M.2 slots supporting up to 16TB RAID 0/1. Wi-Fi 7 provides stronger signal coverage and a more stable wireless experience. The slide-out design facilitates upgrades and maintenance. It easily adapts to personal, studio, or rack-mount enterprise environments
  • 【High-Efficiency Cooling System】Utilizing an aerospace-grade aluminum alloy chassis, copper base plate, six heat pipes, dual turbine fans, and advanced PCM thermal conductive material, it maintains stable cooling performance even under continuous load. This system supports 130W continuous power and 160W peak power operation, with a built-in 320W power supply. It boasts multiple global certifications including CCC, FCC, UL, CE, and UKCA, ensuring stable and reliable operation in various environments
  • 【Cluster Design】Two MS-S1 MAX units can be configured as a dual-unit cluster to run a large 235B Q4 model locally, achieving an output speed of 10.87 tok/s. Supporting 2U rack deployment, multiple MS-S1 MAX units can be cascaded into a distributed cluster to create a high-efficiency AI computing center. A cluster of four MS-S1 MAX units successfully ran a DeepSeek-R1 671B Q4 large model. A reserved cluster power-on interface allows for unified start-up and shutdown
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How much does QuietBox 2 cost, and when does it ship?

Tenstorrent’s product page lists the TT-QuietBox 2 at $9,999 and estimates shipping in 10–12 weeks. These are the listed price and lead-time details in the current product information; confirm them with Tenstorrent before ordering because price and delivery estimates can change.

At that price, the relevant buying-intent product is the Tenstorrent TT-QuietBox 2. It is a complete workstation rather than a set of accelerator cards, so the cost should be weighed against the value of receiving an integrated system and Tenstorrent’s own software stack, not compared only with the price of individual GPUs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

QuietBox 2 vs. DGX Spark vs. a multi-GPU PC

Tenstorrent’s published figures do not provide a same-model, same-runtime comparison with Nvidia DGX Spark or a multi-GPU PC. The table separates what is established for QuietBox 2 from the comparison questions a buyer needs to verify; it does not imply equivalent performance.

Buying question Tenstorrent QuietBox 2 Nvidia DGX Spark Multi-GPU PC
What is being purchased? Complete liquid-cooled workstation with four Blackhole accelerators, AMD Ryzen CPU, memory and storage. Comparable system configuration and included components are not stated in Tenstorrent’s cited product information. Varies with the chosen components; a fixed configuration is not specified.
Memory and model capacity Tenstorrent lists 128 GB GDDR6 across the accelerators and says the system can load GPT-OSS-120B. Not stated in the cited Tenstorrent material. Depends on the selected GPUs, their memory, and whether the software can distribute the workload as needed.
Comparable tokens-per-second result Tenstorrent reports nearly 500 tokens per second on Llama 3.1 70B; this is not an independent standardized comparison. Not stated in the cited Tenstorrent material. Not stated; results depend on hardware, runtime and test conditions.
Software and development focus Tenstorrent’s stack includes TT-Studio, TT-Inference-Server and TT-Metalium. Verify current software, model and framework support for the intended workload. Depends on the GPUs, drivers and software stack selected and configured.
Price and delivery $9,999 and an estimated 10–12 week shipping time on Tenstorrent’s product page. Not stated in Tenstorrent’s cited product information. Depends on selected components and availability.

To make a meaningful comparison, use the same model, model format, runtime and generation settings, then compare usable model size and measured throughput. Also check framework support, cooling and noise, power requirements, total system price, delivery time, and whether you want a complete supported workstation or prefer choosing and assembling discrete cards. Without like-for-like measurements, a single token-per-second figure cannot establish which machine is faster for your work.

Who should consider QuietBox 2?

QuietBox 2 is most relevant to buyers who want a complete Tenstorrent-based local AI system, need its accelerator memory for larger local models, or plan to develop and experiment with Tenstorrent’s software stack. Developers who want to work close to the hardware may value TT-Metalium alongside the inference tools. Tenstorrent thermal-mechanical engineer and team lead Chris Goulet said that internal developers had requested QuietBox systems because they were “just so easy to deploy.”

It is a less straightforward fit if your decision depends on independently verified performance against Nvidia systems, broad framework compatibility for a specific workload, or the lowest-cost route to local inference. The supplied Tenstorrent performance figures are not a standardized head-to-head test, and the right choice depends on software support and results for the workloads you actually plan to run.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.