Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchIn an October 2024 demo, Tenstorrent’s eight-accelerator Loud Box delivered 15 tokens per second per user while running Llama 3.1 70B at BF8 precision for 32 concurrent users. That was a press-reported demonstration, not a general speed guarantee or a test of Tenstorrent’s newer Galaxy Blackhole systems.
What Tenstorrent demonstrated
EE Times reported on October 1, 2024, that Tenstorrent ran Llama 3.1 70B at 15 tokens per second per user with 32 concurrent users on a Loud Box workstation. The system used eight first-generation Wormhole accelerators and BF8 precision.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
MINISFORUM MS-02 Ultra Workstation Mini PC, Intel Core Ultra 9 285HX (24C/24T, up to 5.5GHz), PCIe... | $1,659.00 | Buy on Amazon |
| 2 |
|
GMKtec EVO-X2 AI Mini PC Ryzen Al Max+ 395 Superchip 128GB LPDDR5X 2TB SSD | $3,649.99 | Buy on Amazon |
“Per user” matters here: the result describes each user’s reported generation rate under a 32-user concurrent workload. It is not a claim that one user received the combined throughput of all 32 sessions. EE Times said rates above 10 tokens per second per user are generally enough for readable chatbot and question-and-answer responses.
How to interpret the 15-token result
It was a demo, not an independently established product ceiling
The figure came from an exclusive Tenstorrent demonstration reported by EE Times. The report does not establish an independently run benchmark or promise the same result for every model, prompt, software release, or deployment.
#1 Best Overall
- High-Performance AI Processor:The MS-02 Ultra features an Intel Core Ultra 9 285HX (24C/24T, up to 5.5 GHz, 13 TOPS NPU), delivering fast and efficient performance for AI inference, algorithm development, and media workloads. A PCIe x16 expansion slot supports desktop-class GPU upgrades for advanced model training and accelerated computing tasks. It's ideal for creators, engineers, and teams handling intensive parallel workloads.
- 4 × M.2 PCIe 4.0 + 4 × DDR5 SODIMM slots:Four DDR5 SODIMM slots support up to 256 GB of memory, while ECC helps maintain data integrity in mission-critical environments. Four PCIe 4.0 M.2 slots support up to 24 TB of storage, supporting RAID 0/1/5/10, combining high-speed performance with data protection. It allows for the creation of independent scratch disks, media libraries, and project drives, providing high-throughput for production workflows.
- PCIe & USB 4.0 v2: Up to three PCIe slots can be equipped, including a dual-slot x16 GPU. The main slot supports PCIe 5.0, meeting the needs of high-bandwidth creative and computing workloads. USB 4.0 v2 (80Gbps) supports high-bandwidth external storage and displays.
- Ultra-fast Networking: Wi-Fi 7 further enhances wireless performance with next-generation speeds and low-latency stability. Intelligent bandwidth switching optimizes throughput in different network environments, ensuring optimal performance for enterprise or local networks. Dual 25GbE ports (providing up to approximately 3.125 GB/s bandwidth, about 25 times faster than traditional 1GbE), enabling seamless large-scale file transfers and parallel computing. 10GbE and 2.5GbE ports, with support for Intel vPro technology, ensure enterprise-grade remote management and deployment flexibility.
- Server-grade thermal architecture: Utilizing a dedicated CPU/GPU airflow design, equipped with a 6-pipe dual-fan cooler, it maintains stable performance even under sustained loads, delivering up to 140W Turbo power while maintaining a 100W TDP, and operating with noise levels as low as 36 dB. An integrated 350W power supply ensures stable and reliable output for demanding computing tasks and fully loaded extended configurations.
Tenstorrent said the software still had room to improve
Tenstorrent described 15 tokens per second per user as work in progress and said it aimed to double performance on the same system through software optimization. It had not yet explored speculative decoding and considered batch size 32 a sweet spot. CEO Jim Keller said, “We are pretty happy with the numbers.”
What the 2024 workstation cost—and what that comparison does not prove
EE Times quoted the Loud Box workstation at $12,000. It described Loud Box as the air-cooled version of Quiet Box, powered by first-generation Wormhole chips, and said it was offered as a workstation or a 4U rack-mount server.
The report contrasted that workstation price with Nvidia DGX-H100 systems costing more than $300,000 per eight-GPU system. Those are reported system prices, not a controlled cost-per-token comparison: the figures do not establish equivalent configurations, total cost of ownership, or identical workloads.
What Tenstorrent said about Blackhole at the time
In the 2024 EE Times report, Keller said the forthcoming Blackhole workstation—called Friendly Box in the article—would cost less and be two to three times faster than the Wormhole version. The article also noted that engineering-model numbers existed while Tenstorrent was still working on the end-to-end customer experience. Treat that as a forward-looking statement from that period, not a measured comparison with the Loud Box demo.
Recommended Free Tools
How the current Tenstorrent figures differ
Tenstorrent’s May 4, 2026 TT-Deploy post says Galaxy Blackhole is in production and shipping in volume, with clusters of 36 Galaxies networked as one computer. The post also says TT-QuietBox 2, a smaller water-cooled developer workstation, is available for purchase.
Rank #2
- EVOLUTION RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
- AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
- AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
- EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
- QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.
The newer performance figures concern Galaxy deployments, not a single Loud Box workstation. TT-Deploy reports more than 350 tokens per second per user on DeepSeek across 16 Galaxies at 671B parameters, with batch size 32 and about four seconds to first token. Tenstorrent’s inference page gives a more specific comparison for DeepSeek-R1-0528 at 100k context:
| Reported system or service | Model and context | Output speed | Time to first token | Evidence context |
|---|---|---|---|---|
| Tenstorrent Loud Box | Llama 3.1 70B, BF8 | 15 tokens/second per user | Not stated in the October 1, 2024 EE Times report | Eight first-generation Wormhole accelerators; 32 concurrent users; press-reported demo |
| Tenstorrent Galaxy Blackhole | DeepSeek-R1-0528, 100k context | 350 tokens/second | 4.0 seconds | Tenstorrent’s current inference-page comparison; a different model and multi-Galaxy platform |
| Nvidia provider average cited by Tenstorrent | DeepSeek-R1-0528, 100k context | 86 tokens/second | 7.5 seconds | Top-five Nvidia provider average as presented on Tenstorrent’s page; service comparison, not a workstation hardware test |
The table’s rows are not a like-for-like ranking of workstation hardware. The Loud Box result uses a different model and precision and reports concurrent-user throughput; the Galaxy figures use a multi-system deployment, while the Nvidia number is a provider average. TT-Deploy’s separate 16-Galaxy, 671B, batch-32 result likewise should not be substituted for the 2024 workstation demo.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Is the QuietBox relevant if you want local LLM inference?
The 15-token result is evidence that Tenstorrent demonstrated useful concurrent chatbot throughput on its 2024 Wormhole workstation configuration. It does not establish the performance of TT-QuietBox 2: the available product-status information identifies that system as a water-cooled developer workstation but does not attach the Loud Box benchmark to it.
For a purchase decision, look for results matching the workload you intend to run: the exact model and parameter count, precision, context length, concurrent sessions or batch size, time to first token, and accelerator count. Also distinguish a workstation quote from the total cost of a deployment. The supplied figures do not give a directly comparable total-cost or per-token analysis for Tenstorrent and Nvidia.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

