Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

SambaNova’s SN40L, announced on September 19, 2023, added high-bandwidth memory (HBM) to the company’s Reconfigurable Dataflow Unit (RDU) for the first time. The change gave the chip a fast memory tier between its small on-chip SRAM and much larger DDR5 memory—an arrangement intended to keep large language model data closer to compute without giving up capacity.

What changed with the SN40L?

SN40L was a new SambaNova RDU designed for large-language-model training, fine-tuning, and inference as part of the company’s full-stack SambaNova Suite. The defining change was adding HBM to SambaNova silicon. EE Times described it as a 5 nm design with more compute cores and HBM aimed at large-scale LLM workloads.

The chip can address HBM and DRAM from one device, according to SambaNova. That gives its software a choice of memory tiers rather than forcing all model data through one kind of memory.

How are the memory tiers arranged?

EE Times reported the following configuration per SN40L package. The capacities and roles differ: the smaller tiers are positioned for faster access, while DDR5 supplies much greater capacity.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
  • ✅Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor. 2.5W typical power consumption
  • ✅Scalable, enabling simultaneous processing of multi-streams & multi-models
  • ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
  • ✅Supports TensorFlow, TensorFlow Lite, ONNX, Keras, Pytorch frameworks
  • ✅Supports Linux and Windows. Supports the temperature range of -40°C to 85°C
Memory tier Reported capacity per package Intended role
SRAM 520 MB (EE Times) Smallest, fastest on-chip tier for active data and intermediate values.
HBM3 64 GB (EE Times) New high-bandwidth tier for model and inference data that benefits from fast access.
DDR5 DRAM 1.5 TB (EE Times) Large-capacity tier for models and data that do not fit in SRAM or HBM.

EE Times also reported 1,040 compute cores per package. These package-level figures describe the hardware configuration; they do not by themselves establish the performance of a complete system or a particular model workload.

Why does HBM matter for LLM inference?

Inference repeatedly draws on model weights and, for long conversations or other extended inputs, a key-value (KV) cache that holds information from earlier tokens. When useful data can stay in a memory tier with high bandwidth near the compute, the system may spend less time waiting for data to move from a farther or slower tier. HBM is the middle ground in SN40L’s design: more capacious than SRAM, but intended to provide faster access than the much larger DRAM tier.

Rank #2
MX3 M.2 AI Accelerator
  • High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional performance and efficiency.
  • Flexible Integration: The MX3 can be easily integrated into your existing systems via its M.2 M-key form factor and support for Linux operating systems.
  • Energy Efficient: The MX3 is designed to provide high performance while minimizing power consumption.
  • Comprehensive Software Development Kit (SDK): The MX3 is supported by a comprehensive SDK that simplifies development and deployment.
  • Hardware compatability: The MX3 is compatible with the PCI-SIG M.2 M-key 2280 Specification. It can be used with the Raspberry Pi 5 with a M-key 2280 HAT.

SambaNova’s Dataflow documentation describes its approach this way: “Full models and KV cache load into HBM, then stream onto the chip as needed.” In practical terms, HBM acts as a staging area for model and inference data, while SRAM supports active work on the chip and DRAM retains capacity for data that does not fit in the faster tiers. The software’s ability to address both HBM and DRAM is part of this strategy.

This hierarchy is particularly relevant as model size and context length grow: the system must manage not just computation, but also where weights and accumulated context reside and how quickly they reach the compute. Adding HBM does not, on its own, prove a given latency, throughput, or cost advantage; those depend on the system configuration and workload.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What scale did SambaNova claim?

In its September 2023 announcement, SambaNova said SN40L could serve a 5-trillion-parameter model with a sequence length of 256k or more on one system node. EE Times reported a more specific company comparison for that mixture-of-experts workload: an eight-socket SN40L system versus 24 eight-socket state-of-the-art GPU systems.

That comparison is a company claim reported by EE Times, not an independently validated benchmark in the material available here. The announcement also claimed lower total cost of ownership through more efficient inference, but did not provide a standardized test method or an independent cost study. Treat the model-scale statement as SambaNova’s stated capability, not as a general performance result for every model or workload.

Rank #4
Waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Comes with PCIe to M.2 Adapter Board
  • ✅Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor. 2.5W typical power consumption
  • ✅Scalable, enabling simultaneous processing of multi-streams & multi-models
  • ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
  • ✅Supports TensorFlow, TensorFlow Lite, ONNX, Keras, Pytorch frameworks
  • ✅Supports Linux and Windows. Supports the temperature range of -40°C to 85°C

How was SN40L deployed?

SN40L was not a consumer graphics card. Its initial route to customers was SambaNova Suite, a cloud-based offering. EE Times reported that SambaNova planned to bring the chip to its on-premises DataScale systems, with initial shipping planned for November 2023. That was a plan reported at the time, not confirmation of present-day availability.

The platform’s intended use extended beyond inference to LLM training and fine-tuning. For a comparison with GPU-based inference systems, the useful questions are how each system handles SRAM, HBM, and DRAM; what model and context sizes it supports; how it scales across chips; and whether the deployment is cloud-based or on-premises. Economic comparisons also need workload-specific throughput, latency, power, and cost methodology.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Radxa AICore DX-M1M, 25TOPS NPU, M.2 2242 Module, Low Power Edge AI Accelerator
  • DEEPX DX-M1M NPU: Powered by the DEEPX DX-M1M neural processing unit, purpose-built for efficient on-device AI inference workloads.
  • COMPACT M.2 2242 FORM FACTOR: Fits the standard M.2 2242 slot, making it easy to integrate into embedded systems, edge devices, and compact computing platforms.
  • EDGE AI ACCELERATION: Designed to accelerate deep learning inference at the edge, enabling real-time AI applications without relying on cloud connectivity.
  • RADXA AICORE MODULE: The Radxa AICore DX-M1M delivers a plug-and-play AI compute solution ideal for robotics, smart cameras, and industrial automation.
  • WARRANTY AND ORIGIN: Backed by a 1-year manufacturer warranty and crafted with quality components for reliable long-term performance in demanding environments.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Is the HBM approach still relevant?

Yes: SambaNova’s February 2026 SN50 announcement and current Dataflow architecture materials continue to describe a tiered-memory approach combining large-capacity memory, HBM, and SRAM. SambaNova says SN50 can hot-swap models in HBM and SRAM in milliseconds for agentic workloads; its current Dataflow page describes loading full models and KV cache into HBM before streaming data onto the chip. That page says the architecture scales to models up to 10 trillion parameters on SN50. These are statements about a later generation, not additional SN40L specifications or independent performance measurements.

Quick Recap

Bestseller No. 1
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
✅Scalable, enabling simultaneous processing of multi-streams & multi-models; ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
$219.99
Bestseller No. 2
MX3 M.2 AI Accelerator
MX3 M.2 AI Accelerator
Software and Documentation can be accessed at the MemryX developer website
$169.00
Bestseller No. 4
Waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Comes with PCIe to M.2 Adapter Board
Waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Comes with PCIe to M.2 Adapter Board
✅Scalable, enabling simultaneous processing of multi-streams & multi-models; ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
$225.99

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.