Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

At HiPEAC 2026 in Kraków, AMD’s Michaela Blott argued that improving AI efficiency requires continuous co-design across algorithms, computer architecture and silicon—not simply applying today’s optimization techniques at larger scale. Her keynote’s central warning was blunt: “Current methods are just too lazy.”

What Michaela Blott said at HiPEAC 2026

Blott’s keynote, reported by EE Times on January 27, 2026, focused on three connected ways to make AI systems more efficient: silicon diversity, model optimization and agile AI software stacks.

She argued that efficiency goals cannot be met by optimizing one layer in isolation. Algorithms influence hardware requirements; hardware capabilities shape viable algorithms; and compilers and other software determine whether the available hardware is used effectively.

“New algorithms are needed to bring AI efficiency in line with human performance and provide sustainable scaling.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
  • ✅Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor. 2.5W typical power consumption
  • ✅Scalable, enabling simultaneous processing of multi-streams & multi-models
  • ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
  • ✅Supports TensorFlow, TensorFlow Lite, ONNX, Keras, Pytorch frameworks
  • ✅Supports Linux and Windows. Supports the temperature range of -40°C to 85°C

“Don’t stop optimizing. Explore and co-design architectures with new algorithms, and in tandem with this, design better AI algorithms with better scaling properties.”

These quotations are attributed to Blott by EE Times. No primary transcript or recording was established for independent wording verification.

A three-layer way to think about AI efficiency

The keynote’s argument is best understood as a co-design problem rather than a ranking of competing products. Each layer exposes opportunities—and constraints—for the others.

Rank #2
ESP32-P4 WIFI6 POE ETH AI Development Board, with ESP32-P4 and ESP32-C6
  • High-Performance Dual-Core with Ample Memory--- Equipped with a 360MHz dual-core RISC-V processor, 32MB of onboard PSRAM, and 32MB of Flash memory, providing powerful processing capabilities and ample runtime for complex multimedia applications and edge computing.
  • Powerful Multimedia Processing Center--- Integrated with a dedicated image processor (ISP), H.264 video encoder, and JPEG codec, perfectly supporting camera input and video processing, making it an ideal choice for developing smart displays, video surveillance, and other projects.
  • Hardware-Level Security Protection--- Built-in digital signature, encryption accelerator, and key management unit, providing a one-stop hardware-level security solution from secure boot and data encryption to access control management, ensuring the security of your products and data.
  • Full Connectivity Coverage: Wi-Fi 6, Bluetooth, PoE Power Supply--- Onboard with an ESP32-C6 chip, supporting the latest Wi-Fi 6 and Bluetooth 5.0; it also integrates an Ethernet port with PoE functionality, providing high-speed, flexible, and stable network connectivity, and can be powered directly via Ethernet cable, simplifying deployment.
  • Rich interfaces and strong expandability--- It provides a MIPI camera/display interface, high-speed USB, SD card slot, microphone/speaker interface and a large number of programmable GPIOs, which greatly facilitates the expansion of external devices and meets the needs of various human-computer interaction and Internet of Things applications. Supports AI Speech Interaction: Allows access to online large model platforms such as ChatGPT, DeepSeek, Doubao, etc.
Layer What to examine Why isolated optimization can fall short
Models and algorithms Model structure, numerical methods, scaling behavior and the amount of computation required. A mathematically efficient model may still map poorly to a target processor, memory system or compiler.
Architecture and silicon Processor design, accelerator specialization, memory movement and diversity of hardware targets. New hardware features deliver little value when algorithms and software cannot exploit them.
Software and compiler stack Compilers, programming models, runtime support and the translation of models into hardware operations. Hardware capability can remain unused when software support is slow, incomplete or too rigid.

This framework describes where to look for improvements; it does not establish that one layer is inherently more important or that a particular approach is faster.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why “don’t get lazy” matters in practice

Scaling is not automatically sustainable

Adding more compute can increase capability, but it can also increase energy use, data movement and system complexity. Blott’s call for algorithms with better scaling properties treats efficiency as an ongoing design objective rather than a one-time tuning exercise.

Co-design can reveal options a single team misses

A model designer might reduce operations, while a hardware designer changes the memory path or adds specialized support. A compiler team may then make those changes usable through a better lowering strategy. Considering the three decisions together can expose improvements unavailable to any one layer considered alone.

Rank #3
Hailo-8 M.2 AI Accelerator Module 26TOPS Hailo8 Support Linux/Windows
  • Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor.
  • 2.5W typical power consumption
  • Enabling real-time low latency and high-efficiency AI inferencing on the edge devices
  • Supports TensorFlow TensorFlow Lite, ONNX, Keras, Pytorch frameworks
  • Supports Linux and Windows.

Agility is a technical requirement

AI methods and hardware targets change quickly. An “agile” stack, in the terms reported from the keynote, is one that can adapt models, compilers and runtimes without requiring every change to be treated as a wholly new software project.

The wider HiPEAC 2026 setting

HiPEAC is a European forum covering computer architecture, programming models, compilers and operating systems, with participation from academic researchers and industry. The 2026 conference was held in Kraków, Poland. That setting matters because the efficiency question spans the full computing stack: it is not solely a model-training or chip-manufacturing issue.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the PolyMage Labs example shows

The EE Times account also reported that PolyMage Labs’ automatic compiler for AI hardware was selected for Tenstorrent AI platforms to improve software support. This illustrates the software layer’s role: compiler technology can help connect AI programs with specialized accelerator hardware.

Rank #4
MX3 M.2 AI Accelerator
  • High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional performance and efficiency.
  • Flexible Integration: The MX3 can be easily integrated into your existing systems via its M.2 M-key form factor and support for Linux operating systems.
  • Energy Efficient: The MX3 is designed to provide high performance while minimizing power consumption.
  • Comprehensive Software Development Kit (SDK): The MX3 is supported by a comprehensive SDK that simplifies development and deployment.
  • Hardware compatability: The MX3 is compatible with the PCI-SIG M.2 M-key 2280 Specification. It can be used with the Raspberry Pi 5 with a M-key 2280 HAT.

The report does not provide independent performance measurements, procurement terms or evidence that the selection produced a particular speed, energy or cost improvement. It should therefore be treated as an example of industry activity, not as a benchmark or buying recommendation.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to apply the message to an AI project

  1. Define the real objective. Decide whether the priority is latency, throughput, energy, memory capacity, cost or a combination. “Efficiency” is not a single metric.
  2. Profile the whole path. Measure model computation, memory traffic, compilation, runtime overhead and data movement on the intended deployment target.
  3. Change more than one layer when necessary. Consider model or algorithm changes alongside kernel, architecture and compiler changes instead of assuming the current model must remain fixed.
  4. Validate on the target system. A theoretical reduction in operations does not prove an end-to-end gain until it is implemented and measured under the workload’s actual conditions.
  5. Keep the stack adaptable. Favor interfaces and tooling that can accommodate new models and accelerator features as they emerge.

What is—and is not—established by the report

  • The keynote made a qualitative case for sustained optimization across algorithms, architecture and silicon.
  • It identified silicon diversity, model optimization and agile AI stacks as relevant approaches.
  • It supplied no named statistic suitable for proving a specific efficiency gain.
  • It did not independently validate the performance of the PolyMage Labs compiler selection for Tenstorrent platforms.

The Bottom Line

HiPEAC 2026’s practical lesson is to treat AI efficiency as an ongoing, cross-layer co-design task. Better algorithms, suitable silicon and an adaptable compiler stack must be developed together; optimizing only one layer risks leaving gains elsewhere untouched.

Quick Recap

Bestseller No. 1
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
✅Scalable, enabling simultaneous processing of multi-streams & multi-models; ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
$219.99
Bestseller No. 3
Hailo-8 M.2 AI Accelerator Module 26TOPS Hailo8 Support Linux/Windows
Hailo-8 M.2 AI Accelerator Module 26TOPS Hailo8 Support Linux/Windows
Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor.; 2.5W typical power consumption
$214.99
Bestseller No. 4
MX3 M.2 AI Accelerator
MX3 M.2 AI Accelerator
Software and Documentation can be accessed at the MemryX developer website
$169.00

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.