Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

You can run many models on Tenstorrent hardware, but “any model” does not mean every model will work unchanged. Start by checking whether your exact model is validated for your hardware generation, then choose a software path based on its framework and your need for serving, multi-chip use, or low-level control.

Check whether your model is validated first

Tenstorrent offers several software routes for running models, but it also maintains separate validated-model catalogs. A compiler pathway can be broad without guaranteeing drop-in support for every model. Search for the model and filter for your target hardware and software before planning around it. Tenstorrent’s developer page showed 47 model entries when reviewed on October 5, 2026; that is a changing catalog snapshot, not a fixed ecosystem limit or a promise that each entry works on every device. Tenstorrent’s TT-Forge documentation points to tt-forge-models as the validation source for Forge.

If your model is absent from the relevant catalog, that does not establish that it cannot run. It does mean you should expect to verify the path yourself: unsupported operations or compiler and runtime behavior may require porting and debugging.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the software route for your framework and goal

Your starting point or goal Route to investigate Important qualification
PyTorch or JAX code TT-XLA The bring-up guide documents the PJRT plugin route and a PyTorch torch.compile backend.
ONNX, TensorFlow, or PaddlePaddle model TT-Forge-ONNX The documented route is single-chip only.
Packaged inference or model serving TT-Inference-Server Check its validated models for your specific hardware.
Point-and-click interface TT-Studio Check current device and model support in its documentation.
Custom operations or direct hardware control TT-Metalium This is the lower-level SDK; TT-NN offers a higher-level Python/C++ operation library.

The bring-up guide says TT-Torch is deprecated for new PyTorch work, so use TT-XLA as the starting point for new PyTorch projects rather than beginning with the older route. Tenstorrent describes TT-Forge as “Tenstorrent’s end-to-end compiler stack,” but that description should not be read as a universal model-compatibility guarantee. See the TT-Forge documentation.

#1 Best Overall
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
  • ✅Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor. 2.5W typical power consumption
  • ✅Scalable, enabling simultaneous processing of multi-streams & multi-models
  • ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
  • ✅Supports TensorFlow, TensorFlow Lite, ONNX, Keras, Pytorch frameworks
  • ✅Supports Linux and Windows. Supports the temperature range of -40°C to 85°C

Bring up a model and account for compilation

The TT-XLA guide demonstrates installing Tenstorrent’s PJRT plugin, confirming that JAX discovers a tt device, and compiling a PyTorch model with torch.compile(model, backend="tt"). Its worked example loads a Hugging Face Llama 3.2 1B model and runs inference. It demonstrates a workflow, not compatibility for every Hugging Face model. Follow the official bring-up guide for the current installation and code details.

In the documented TT-XLA path, compilation is lazy: the first forward pass triggers compilation and caching. The guide notes that early iterations can be slow because they may include compilation, weight transfer, kernel compilation, or runtime trace capture. For a performance measurement, run at least three dummy warm-up iterations before timing. A cold first run is not a useful comparison with another platform’s steady-state run.

Rank #2
ESP32-P4 WIFI6 POE ETH AI Development Board, with ESP32-P4 and ESP32-C6
  • High-Performance Dual-Core with Ample Memory--- Equipped with a 360MHz dual-core RISC-V processor, 32MB of onboard PSRAM, and 32MB of Flash memory, providing powerful processing capabilities and ample runtime for complex multimedia applications and edge computing.
  • Powerful Multimedia Processing Center--- Integrated with a dedicated image processor (ISP), H.264 video encoder, and JPEG codec, perfectly supporting camera input and video processing, making it an ideal choice for developing smart displays, video surveillance, and other projects.
  • Hardware-Level Security Protection--- Built-in digital signature, encryption accelerator, and key management unit, providing a one-stop hardware-level security solution from secure boot and data encryption to access control management, ensuring the security of your products and data.
  • Full Connectivity Coverage: Wi-Fi 6, Bluetooth, PoE Power Supply--- Onboard with an ESP32-C6 chip, supporting the latest Wi-Fi 6 and Bluetooth 5.0; it also integrates an Ethernet port with PoE functionality, providing high-speed, flexible, and stable network connectivity, and can be powered directly via Ethernet cable, simplifying deployment.
  • Rich interfaces and strong expandability--- It provides a MIPI camera/display interface, high-speed USB, SD card slot, microphone/speaker interface and a large number of programmable GPIOs, which greatly facilitates the expansion of external devices and meets the needs of various human-computer interaction and Internet of Things applications. Supports AI Speech Interaction: Allows access to online large model platforms such as ChatGPT, DeepSeek, Doubao, etc.

Match installation instructions to your exact device and release

Before installing, identify the card or system and the software release you intend to use. Requirements are device- and version-specific. For example, the TT-Metalium v0.60.1 compatibility matrix lists Ubuntu 22.04 and Python 3.10 for its listed Galaxy, Wormhole/T3000, and Blackhole configurations, while driver, firmware, and utility requirements vary by device. Those details are an example for that release, not a universal prescription for other versions. Use the instructions packaged with the release you are installing.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

If you do not have Tenstorrent hardware

Tenstorrent’s documentation home advertises Cloud Console access to its silicon, and its developer page provides hardware compatibility and model filters. Availability, eligibility, and terms can change, so check the current official pages directly. A Quietbox 2 guide also describes a turnkey workstation with drivers, serving software, TT-Studio, and a cached Qwen3-32B model; its live verification is dated August 26, 2026, and the guide warns that preinstalled software can become dated. Treat that as a product-specific setup reference, not a guarantee of current availability or configuration.

Quick Recap

Bestseller No. 1
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
✅Scalable, enabling simultaneous processing of multi-streams & multi-models; ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
$219.99
Bestseller No. 3
Hailo-8 M.2 AI Accelerator Module 26TOPS Hailo8 Support Linux/Windows
Hailo-8 M.2 AI Accelerator Module 26TOPS Hailo8 Support Linux/Windows
Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor.; 2.5W typical power consumption
$214.99
Bestseller No. 4
MX3 M.2 AI Accelerator
MX3 M.2 AI Accelerator
Software and Documentation can be accessed at the MemryX developer website
$169.00
Rank #4
MX3 M.2 AI Accelerator
  • High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional performance and efficiency.
  • Flexible Integration: The MX3 can be easily integrated into your existing systems via its M.2 M-key form factor and support for Linux operating systems.
  • Energy Efficient: The MX3 is designed to provide high performance while minimizing power consumption.
  • Comprehensive Software Development Kit (SDK): The MX3 is supported by a comprehensive SDK that simplifies development and deployment.
  • Hardware compatability: The MX3 is compatible with the PCI-SIG M.2 M-key 2280 Specification. It can be used with the Raspberry Pi 5 with a M-key 2280 HAT.
Rank #3
Hailo-8 M.2 AI Accelerator Module 26TOPS Hailo8 Support Linux/Windows
  • Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor.
  • 2.5W typical power consumption
  • Enabling real-time low latency and high-efficiency AI inferencing on the edge devices
  • Supports TensorFlow TensorFlow Lite, ONNX, Keras, Pytorch frameworks
  • Supports Linux and Windows.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.