Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Vastai Technologies raised RMB 500 million (approximately $77 million) in a 2021 Series A+ round to develop data-center inference chips, software and the ecosystem needed to deploy them. The Chinese fabless semiconductor startup’s financing was led by Matrix Partners China and the China Internet Investment Fund.

Who funded Vastai Technologies?

EE Times reported that Matrix Partners China and China Internet Investment Fund led Vastai’s Series A+ financing, with several existing investors also participating. The round brought the company’s cumulative funding to about $133 million, according to the same 2021 report.

CEO John Qian described the capital requirement plainly: “To make a good product, it takes that much to get it done.”

What was the $77 million intended to pay for?

Qian said Vastai planned to put the money into three connected areas:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
HPE NVIDIA Tesla V100 32GB HBM2 PCIe 3.0 x16 Passive GPU Computational Accelerator for AI Machine Learning HPC Deep Learning 699-2G500-0216-400 (Renewed)
  • NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
  • 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
  • PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
  • NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
  • Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads
  • Products and intellectual property: continued development of the company’s inference silicon and related IP.
  • People: hiring and expanding the engineering and product organization.
  • Software and ecosystem work: tools, integrations and broader platform support so customers could use the chips beyond a single hardware design.

That last item is important for an accelerator company. Hardware specifications alone do not make a data-center platform adoptable; operators also need software interfaces, model support, deployment tools and compatible systems.

What products did Vastai announce?

In an official July 2021 release, Vastai introduced the SV100 general-purpose inference chip and the VA1 accelerator card for cloud data centers.

Rank #2
MX3 M.2 AI Accelerator
  • High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional performance and efficiency.
  • Flexible Integration: The MX3 can be easily integrated into your existing systems via its M.2 M-key form factor and support for Linux operating systems.
  • Energy Efficient: The MX3 is designed to provide high performance while minimizing power consumption.
  • Comprehensive Software Development Kit (SDK): The MX3 is supported by a comprehensive SDK that simplifies development and deployment.
  • Hardware compatability: The MX3 is compatible with the PCI-SIG M.2 M-key 2280 Specification. It can be used with the Raspberry Pi 5 with a M-key 2280 HAT.
Product Role Published specifications and workloads
SV100 General-purpose inference chip More than 200 TOPS of INT8 peak processing on one chip, according to Vastai’s 2021 announcement.
VA1 PCIe inference accelerator card 70W half-height, half-length card; designed for high-performance, low-latency data-center inference; up to 120 channels of 1080p video decoding.

Vastai positioned the VA1 for computer vision, video processing, natural-language processing, search and recommendation workloads. Its official product description calls it “High-performance and low-latency inference for data centers.”

Is the SV100 a GPU?

Vastai’s published material describes the SV100 as an inference chip, not as a general-purpose GPU. The distinction is about intended workload and programming model: an inference accelerator is designed to execute trained neural-network models efficiently, while a general-purpose GPU is built for a much broader range of parallel computing and graphics tasks.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
  • ✅Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor. 2.5W typical power consumption
  • ✅Scalable, enabling simultaneous processing of multi-streams & multi-models
  • ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
  • ✅Supports TensorFlow, TensorFlow Lite, ONNX, Keras, Pytorch frameworks
  • ✅Supports Linux and Windows. Supports the temperature range of -40°C to 85°C

The available announcements do not establish whether the SV100 uses a GPU-like architecture internally. It is therefore more accurate to call it a dedicated or purpose-built AI inference accelerator than to label it a conventional GPU.

How should you interpret the “more than 200 TOPS” claim?

TOPS means trillion operations per second. Vastai’s figure is specifically an INT8 peak figure for one SV100 chip, published in 2021. INT8 is an eight-bit integer data format commonly used to speed neural-network inference after quantization.

Rank #4

Peak TOPS is not the same as application throughput. Real performance depends on the model, batch size, sparsity, compiler and runtime, memory movement, preprocessing, networking and the way a server is configured. The cited announcements do not provide an independent benchmark comparison with competing GPUs or accelerators, so the number should be treated as a vendor peak specification rather than a cross-platform ranking.

What is the Vastai VA1 card used for?

The VA1 packages Vastai’s inference capability as a low-power PCIe add-in card. Its stated 70W power envelope and half-height, half-length form factor are aimed at fitting into servers where space, cooling and power capacity are constrained.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
ASRock Radeon AI PRO R9700 Creator 32GB Professional Graphics Card, 2920 MHz Boost Clock, GDDR6, AMD RDNA 4, AI-Accelerators, DisplayPort 2.1a, PCIe 5.0, Blower Cooler
  • Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
  • Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
  • Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
  • Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
  • Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.

The card’s published video capability—up to 120 channels of 1080p decoding—makes it particularly relevant to dense video analytics, provided that the customer’s software stack and camera pipeline match the supported formats and deployment conditions. Vastai also lists non-video workloads including NLP, search and recommendation.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What happened after the Series A+ round?

In December 2021, Vastai announced RMB 1.6 billion in B-1 and B-2 financing. The company said this later funding would support commercialization of the SV100 and additional GPU research, indicating a move from product development toward larger-scale deployment and a broader accelerator portfolio.

How does a data-center inference accelerator differ from a general-purpose GPU?

For a procurement comparison, the relevant question is not simply which device has the highest TOPS number. Evaluate the complete deployment:

  • Inference throughput per watt: useful work delivered within the server’s power and cooling budget.
  • Latency: especially important for interactive services, recommendation systems and real-time vision.
  • Video decode density: how many concurrent streams the card can process at the required resolution and frame rate.
  • Software and API compatibility: model formats, compilers, runtimes, libraries and integration with existing orchestration tools.
  • Form factor and power: whether the board fits the target server and its available PCIe slots, airflow and power limits.
  • Availability and support: production access, regional support, documentation, replacement planning and the maturity of the vendor ecosystem.

A GPU may offer a larger software ecosystem and flexibility across training and inference. A specialized card may be attractive when predictable inference latency, power efficiency or video density matters more than broad programmability. Which is better depends on the workload and deployment market; the published Vastai figures alone do not settle that comparison.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the financing announcement does not establish

  • It does not establish current pricing for the SV100 or VA1.
  • It does not confirm present production status, reseller coverage or retail availability.
  • It does not provide an independent benchmark against NVIDIA, AMD, Intel or other accelerator products.
  • It does not show that the VA1 supports every video codec, model framework or server configuration.

Those details require current product documentation and validation for the specific region and system being considered.

Quick Recap

Bestseller No. 2
MX3 M.2 AI Accelerator
MX3 M.2 AI Accelerator
Software and Documentation can be accessed at the MemryX developer website
$169.00
Bestseller No. 3
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
✅Scalable, enabling simultaneous processing of multi-streams & multi-models; ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
$219.99
Bestseller No. 4
Tesla L40S 48GB AI HPC Graphics Accelerator
Tesla L40S 48GB AI HPC Graphics Accelerator
48GB AI graphics accelerator
$6,199.00

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.