Recommended Free Tools
IBM Power10 can run suitable AI-inference and dense linear-algebra workloads without a discrete GPU. Its Power ISA v3.1 Matrix-Multiply Assist (MMA) facility puts dedicated matrix hardware inside every CPU core. IBM documents four MMA engines per core, mixed-precision operation and 2,048-bit results per cycle. The benefit appears only when software uses the MMA instructions through supported compiler built-ins or optimized libraries, and IBM’s speed figures apply to specific workloads and test systems rather than to every application.
What Power10 changes
“Bringing math and AI back home to the CPU” means moving selected operations that are often sent to a discrete accelerator onto matrix hardware integrated into each Power10 core. MMA is designed for small-matrix operations that form the inner loops of matrix multiplication, convolution and discrete Fourier transforms.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
The IBM Power10 Implementation Guide: Maximizing Performance and Reliability for Mission-Critical... | $7.59 | Buy on Amazon |
This is primarily an inference and numerical-computing feature. Inference generally needs less computation than training, so a server can execute many models on Power10 cores when the model, numeric format and software stack map well to MMA. It does not make every model or every training job equivalent to a GPU workload.
Power10 MMA hardware, at a glance
| Feature | IBM’s documented detail |
|---|---|
| ISA facility | Matrix-Multiply Assist in Power ISA version 3.1 |
| Engines per core | Four 512-bit MMA engines |
| Result width | 2,048 bits of matrix results per cycle |
| Precision modes | Single precision (SP), double precision (DP), bfloat16 (BF16), half precision (HP), INT16, INT8 and INT4 |
| Operation support | Outer products and matrix operations used by multiplication, convolution and FFT-style workloads |
| Threading | IBM’s AIX guidance documents SMT-8, allowing up to eight simultaneous hardware threads per core |
| Acceleration range | IBM’s support material cites 4–32× matrix-math acceleration, depending on the operation and comparison being made |
The engines are not a separate add-in card. They share the processor’s execution and memory environment, which can avoid moving each inference request to another device. The practical gain still depends on keeping data in an efficient layout and feeding the matrix units with kernels that use them.
#1 Best Overall
Can Power10 run AI inference without a GPU?
Yes, for suitable models. IBM specifically positions MMA for CPU-based inference, including workloads using FP32, BF16, INT8 and other supported formats. A model can run without a discrete GPU when its major compute kernels are dense matrix operations and the framework or library dispatches those operations to MMA.
“Without a GPU” does not mean “without acceleration.” MMA is fixed-function-style matrix hardware inside the CPU core. It is also not a guarantee that a Power10 server will outperform a GPU system: training, very large parallel models, unsupported operators, sparse or irregular workloads, and applications that spend most of their time moving data may favor another architecture.
What determines whether a model benefits
- Kernel shape: matrix multiplication, convolution and related dense linear algebra provide the clearest match.
- Numeric format: FP32, BF16, INT8 and the other MMA modes must match the model’s accuracy and deployment requirements.
- Software dispatch: the compiler, framework and math library must emit MMA instructions rather than ordinary scalar or vector code.
- System balance: memory capacity, bandwidth and data movement can limit gains even when the arithmetic is accelerated.
Workloads that fit MMA best
IBM Research describes MMA instructions for matrix multiplication, convolution and discrete Fourier transform operations. Those primitives occur in neural-network inference, signal processing, scientific computing and other dense numerical kernels.
- Neural-network inference using FP32, BF16, INT8, INT16, INT4 or other supported precision paths
- General matrix multiplication and outer-product kernels
- Convolution layers
- FFT and related signal-processing transforms
- Dense linear-algebra sections of scientific and engineering applications
How software reaches the matrix engines
Power10’s silicon delivers its advertised benefit only when applications are built against software that knows about MMA. IBM’s support guidance names compiler built-ins and optimized math libraries as the normal path.
- Use a compiler with MMA support. IBM lists GCC 10 and later and LLVM 12 and later built-ins for the facility.
- Compile the hot kernels for Power10. Built-ins let developers express MMA operations directly when a library does not already provide the needed kernel.
- Prefer tuned libraries where possible. IBM cites OpenBLAS, IBM Engineering and Scientific Subroutine Library (ESSL), and Eigen as optimized library options.
- Use framework integrations for models. IBM’s Open XL C/C++ documentation describes built-ins intended to accelerate FP32, bfloat16 and INT8 AI inference, while IBM’s enterprise workflow supports ONNX, PyTorch and TensorFlow model paths.
- Measure the complete application. Benchmark preprocessing, inference, post-processing and memory traffic, not only the matrix multiply in isolation.
How much faster is Power10 than Power9?
There is no single Power10-versus-Power9 multiplier. IBM published several figures using different workloads, precisions, system levels and methods. The following numbers should not be combined into one universal speedup.
| Figure | What it measures | Qualification |
|---|---|---|
| 4× | Per-core matrix performance versus POWER9 at constant frequency | IBM Research result for the MMA facility; workload is matrix math, not an all-application benchmark |
| 2.6× | Core-level energy-efficiency improvement versus POWER9 | Projected SPECint result in IBM’s ISCA 2021 analysis |
| Up to 3× | Socket-level energy-efficiency improvement versus POWER9 | Projected figure from the same ISCA analysis |
| Up to 10× FP32; 21× INT8 | AI socket-performance improvement for specified ResNet-50 and BERT-Large models | Projected pre-silicon analysis; model, precision and socket scope are part of the claim |
| 5× | Per-socket inference throughput from Power E980 to Power E1080 | IBM briefing result for large FP32 BERT using PyTorch, OpenBLAS and SQuAD v1.1 |
The 5× E1080 result is an IBM benchmark claim, not a general comparison with every GPU server. Likewise, the 4× figure is per core at constant frequency, while the 10× and 21× figures are projected socket-level AI results for named models and precisions.
Where Power10 fits in an enterprise deployment
IBM launched the Power E1080 and Power10 platform in September 2021 for enterprise workloads, including AI where data already resides on the server. IBM describes ONNX support and deployment paths for models from TensorFlow and PyTorch. Power10 systems are also positioned for AIX and Linux environments and for Red Hat OpenShift-based application platforms.
This makes Power10 a platform decision rather than only a chip decision. Organizations can consolidate application serving, databases and selected inference on IBM Power servers, while using IBM software, business partners or OpenShift tooling to operationalize models. The same server can still use conventional CPU code when a workload does not map to MMA.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →How to compare Power10 with a GPU or POWER9 system
Use the same comparison axes for every candidate system:
| Axis | Questions to answer |
|---|---|
| Workload | Which model or kernel is being tested: BERT, ResNet, convolution, GEMM, FFT or something else? |
| Precision | Is the result FP32, BF16, INT8, INT4 or another mode, and does accuracy remain acceptable? |
| Scope | Is the number per core, per socket or for the complete server? |
| Evidence type | Was it measured on shipping hardware or projected before silicon? |
| Data movement | Where do model weights and request data reside, and how much time is spent moving them? |
| Software path | Are the compiler, BLAS library and framework actually emitting MMA instructions? |
| Accelerator requirement | Does the design need a discrete GPU, or can the Power10 socket meet latency and throughput targets alone? |
Limitations to account for
- Software enablement is mandatory: unoptimized code will not automatically receive the headline MMA gains.
- Speedups are workload-specific: IBM’s published figures cover named models, precisions and benchmark configurations.
- Inference is not training: Power10’s CPU-based approach is most compelling for serving suitable inference workloads, not as a blanket replacement for large training accelerators.
- Memory and operators matter: non-matrix portions of a model, unsupported operations and data transfers can dominate end-to-end latency.
- Platform fit matters: licensing, existing AIX or Linux operations, OpenShift requirements and the location of enterprise data can be as important as raw arithmetic throughput.
Bottom line
Power10 makes AI inference on a CPU a serious enterprise option by putting four MMA engines in every core and supporting precisions from FP32 down to INT4. It is strongest for dense, well-optimized matrix workloads whose software stack uses Power ISA v3.1 instructions through compilers and libraries. IBM reports substantial gains over POWER9, including a 4× per-core matrix result and a specified 5× E1080-versus-E980 inference result, but those figures are scoped claims—not a universal replacement for GPUs.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

