The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →SambaNova’s SN40L, announced on September 19, 2023, added high-bandwidth memory (HBM) to the company’s Reconfigurable Dataflow Unit (RDU) for the first time. The change gave the chip a fast memory tier between its small on-chip SRAM and much larger DDR5 memory—an arrangement intended to keep large language model data closer to compute without giving up capacity.
What changed with the SN40L?
SN40L was a new SambaNova RDU designed for large-language-model training, fine-tuning, and inference as part of the company’s full-stack SambaNova Suite. The defining change was adding HBM to SambaNova silicon. EE Times described it as a 5 nm design with more compute cores and HBM aimed at large-scale LLM workloads.
The chip can address HBM and DRAM from one device, according to SambaNova. That gives its software a choice of memory tiers rather than forcing all model data through one kind of memory.
How are the memory tiers arranged?
EE Times reported the following configuration per SN40L package. The capacities and roles differ: the smaller tiers are positioned for faster access, while DDR5 supplies much greater capacity.
Recommended Free Tools
#1 Best Overall
- ✅Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor. 2.5W typical power consumption
- ✅Scalable, enabling simultaneous processing of multi-streams & multi-models
- ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
- ✅Supports TensorFlow, TensorFlow Lite, ONNX, Keras, Pytorch frameworks
- ✅Supports Linux and Windows. Supports the temperature range of -40°C to 85°C
| Memory tier | Reported capacity per package | Intended role |
|---|---|---|
| SRAM | 520 MB (EE Times) | Smallest, fastest on-chip tier for active data and intermediate values. |
| HBM3 | 64 GB (EE Times) | New high-bandwidth tier for model and inference data that benefits from fast access. |
| DDR5 DRAM | 1.5 TB (EE Times) | Large-capacity tier for models and data that do not fit in SRAM or HBM. |
EE Times also reported 1,040 compute cores per package. These package-level figures describe the hardware configuration; they do not by themselves establish the performance of a complete system or a particular model workload.
Why does HBM matter for LLM inference?
Inference repeatedly draws on model weights and, for long conversations or other extended inputs, a key-value (KV) cache that holds information from earlier tokens. When useful data can stay in a memory tier with high bandwidth near the compute, the system may spend less time waiting for data to move from a farther or slower tier. HBM is the middle ground in SN40L’s design: more capacious than SRAM, but intended to provide faster access than the much larger DRAM tier.
Rank #2
- High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional performance and efficiency.
- Flexible Integration: The MX3 can be easily integrated into your existing systems via its M.2 M-key form factor and support for Linux operating systems.
- Energy Efficient: The MX3 is designed to provide high performance while minimizing power consumption.
- Comprehensive Software Development Kit (SDK): The MX3 is supported by a comprehensive SDK that simplifies development and deployment.
- Hardware compatability: The MX3 is compatible with the PCI-SIG M.2 M-key 2280 Specification. It can be used with the Raspberry Pi 5 with a M-key 2280 HAT.
SambaNova’s Dataflow documentation describes its approach this way: “Full models and KV cache load into HBM, then stream onto the chip as needed.” In practical terms, HBM acts as a staging area for model and inference data, while SRAM supports active work on the chip and DRAM retains capacity for data that does not fit in the faster tiers. The software’s ability to address both HBM and DRAM is part of this strategy.
This hierarchy is particularly relevant as model size and context length grow: the system must manage not just computation, but also where weights and accumulated context reside and how quickly they reach the compute. Adding HBM does not, on its own, prove a given latency, throughput, or cost advantage; those depend on the system configuration and workload.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteWhat scale did SambaNova claim?
In its September 2023 announcement, SambaNova said SN40L could serve a 5-trillion-parameter model with a sequence length of 256k or more on one system node. EE Times reported a more specific company comparison for that mixture-of-experts workload: an eight-socket SN40L system versus 24 eight-socket state-of-the-art GPU systems.
That comparison is a company claim reported by EE Times, not an independently validated benchmark in the material available here. The announcement also claimed lower total cost of ownership through more efficient inference, but did not provide a standardized test method or an independent cost study. Treat the model-scale statement as SambaNova’s stated capability, not as a general performance result for every model or workload.
Rank #4
- ✅Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor. 2.5W typical power consumption
- ✅Scalable, enabling simultaneous processing of multi-streams & multi-models
- ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
- ✅Supports TensorFlow, TensorFlow Lite, ONNX, Keras, Pytorch frameworks
- ✅Supports Linux and Windows. Supports the temperature range of -40°C to 85°C
How was SN40L deployed?
SN40L was not a consumer graphics card. Its initial route to customers was SambaNova Suite, a cloud-based offering. EE Times reported that SambaNova planned to bring the chip to its on-premises DataScale systems, with initial shipping planned for November 2023. That was a plan reported at the time, not confirmation of present-day availability.
The platform’s intended use extended beyond inference to LLM training and fine-tuning. For a comparison with GPU-based inference systems, the useful questions are how each system handles SRAM, HBM, and DRAM; what model and context sizes it supports; how it scales across chips; and whether the deployment is cloud-based or on-premises. Economic comparisons also need workload-specific throughput, latency, power, and cost methodology.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
- DEEPX DX-M1M NPU: Powered by the DEEPX DX-M1M neural processing unit, purpose-built for efficient on-device AI inference workloads.
- COMPACT M.2 2242 FORM FACTOR: Fits the standard M.2 2242 slot, making it easy to integrate into embedded systems, edge devices, and compact computing platforms.
- EDGE AI ACCELERATION: Designed to accelerate deep learning inference at the edge, enabling real-time AI applications without relying on cloud connectivity.
- RADXA AICORE MODULE: The Radxa AICore DX-M1M delivers a plug-and-play AI compute solution ideal for robotics, smart cameras, and industrial automation.
- WARRANTY AND ORIGIN: Backed by a 1-year manufacturer warranty and crafted with quality components for reliable long-term performance in demanding environments.
Is the HBM approach still relevant?
Yes: SambaNova’s February 2026 SN50 announcement and current Dataflow architecture materials continue to describe a tiered-memory approach combining large-capacity memory, HBM, and SRAM. SambaNova says SN50 can hot-swap models in HBM and SRAM in milliseconds for agentic workloads; its current Dataflow page describes loading full models and KV cache into HBM before streaming data onto the chip. That page says the architecture scales to models up to 10 trillion parameters on SN50. These are statements about a later generation, not additional SN40L specifications or independent performance measurements.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

