Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Moving image models into a browser can fail in three different ways: a model can exceed GPU or memory limits, return plausible but incorrect pixels without an error, or produce output that your code decodes incorrectly. In a September 28, 2026 DEV Community post, developer alex.toolkit describes encountering all three while rebuilding a photo editor with ONNX Runtime Web, WebGPU, and a WebAssembly fallback. These are observations from one project—not proof that the same failures affect every browser, runtime, or device.

Why can a model work offline but fail in the browser?

A model’s quality on a small offline comparison does not tell you whether it can run within a browser’s GPU limits or memory budget. The post’s author compared BiRefNet-lite with RMBG-1.4 on ten images and preferred BiRefNet-lite, which the post identifies as MIT-licensed. But in the author’s Apple GPU setup, its first ONNX Runtime Web session.run() failed with Too many storage buffers in shader. Current: 11, Max is 10.

The author attributed the failure to a generated fused kernel needing eleven storage buffers where the target allowed ten per shader stage. Reducing graph optimization did not resolve it. The reported WebAssembly attempt failed separately with std::bad_alloc: the author said 1024×1024 transformer activations did not fit the cited 4 GB wasm32 heap. The post says BEN2 failed similarly. These details describe that project’s particular device and runtime setup; they should not be read as universal limits for Apple GPUs, WebGPU, or WASM.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Test whether the model fits before ranking its quality

In the same project, RMBG-1.4 ran in about 0.25 seconds on WebGPU and about 6 seconds on WASM, according to the author. These are project-reported timings, not standardized benchmarks, and the model the author considered better did not run in that setup. The useful comparison is therefore not simply “Which model scored better offline?” but “Which candidate works, with acceptable output and speed, on the browser and weakest hardware I intend to support?”

#1 Best Overall
GMKtec AI Mini PC Ryzen Al Max+ 395 (up to 5.1GHz) Mini Gaming Computers
  • EVOLUTION AMD RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
  • AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
  • AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
  • EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
  • QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.

How can WebGPU return a bad image without throwing an error?

The author reports that LaMa inpainting completed on WebGPU without an exception and returned a tensor with the expected shape and values in the 0–255 range. Yet the filled hole appeared almost white. In the post’s reported measurements, the hole’s mean pixel value was 254.3 on WebGPU and 107.3 on WASM; the outside-region mean was 127.0 for both. The author attributed the discrepancy to LaMa’s Fourier convolutions (RFFT/IRFFT) producing wrong values through the WebGPU execution provider in that setup. Those measurements and the proposed cause are the author’s account, not independently verified results.

This exposes a gap in exception-only fallback logic: it can switch providers when an operation throws, but not when inference succeeds and the image is semantically wrong. The author routed LaMa to WASM and changed end-to-end tests to inspect pixel colors in actual outputs. More broadly, test the result your user sees—not only whether the call completed, the tensor has the expected shape, or its values fall within a plausible numeric range.

Rank #2
AMD Ryzen™ AI Halo - Personal AI Desktop Computer - Developer Platform - Linux OS
  • Built for Local AI Development: AMD Ryzen AI Halo is designed for local AI development and inference, featuring 128GB unified memory and support for up to 200B parameter models to build and run intensive AI workloads locally.
  • 128GB Unified Memory: Features 128GB LPDDR5x unified memory at 8000 MT/s with 256 GB/s memory bandwidth, providing a shared memory pool across the CPU, GPU, and NPU to support larger AI models.
  • AMD Ryzen AI Max+ 395 Processor: Features 16 cores, 32 threads, and Zen 5 architecture, paired with AMD Radeon 8060S integrated graphics featuring 40 RDNA 3.5 compute units and an AMD XDNA 2 NPU with up to 50 TOPS.
  • Linux AI Developer Platform: Purpose-built for Linux-based AI development with full AMD ROCm software support and preloaded tools, models, and workflows optimized for local AI development.
  • Compact, Connected Design: Includes a 2TB M.2 SSD, 10GbE LAN, Wi-Fi 7, Bluetooth 5.4, USB-C connectivity, and HDMI 2.1b.

Why can fp16 output turn black in Chrome?

The post’s third failure involved Real-ESRGAN x4plus, which uses fp16 inputs and outputs. Initially, the author encoded inputs in a Uint16Array and interpreted outputs as raw half-float bit patterns. The author reports that when native Float16Array support was available in Chrome, ONNX Runtime Web returned fp16 outputs as ordinary numbers. Treating those numbers as bit patterns made the upscaled image render black.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The author’s fix handled both representations: raw half-float bits in a Uint16Array and numeric values. The report is version-sensitive; it does not establish how every Chrome or ONNX Runtime Web combination represents fp16 output. Check the actual output type and representation in the browser-runtime combinations you support rather than assuming that an fp16 tensor always arrives as raw bits.

Rank #3
GMKtec EVO-X2 AI Mini PC Ryzen Al Max+ 395 Superchip 128GB LPDDR5X 2TB SSD
  • EVOLUTION RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
  • AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
  • AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
  • EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
  • QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.

A separate WebGPU buffer-reuse problem

The same model reportedly encountered a WebGPU error, Shape mismatch attempting to re-use buffer. The author addressed it by pinning symbolic dimensions (N: 1, H: 192, W: 192) and using fixed-size tiles. This is another implementation-specific account, not evidence that those dimensions or tiling choices are a general fix for other models.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What should browser inference tests check?

The three failures point to different test layers. A load or inference exception can reveal resource or execution problems; it cannot establish that the resulting pixels are correct. A practical validation plan should include:

  • Target-environment fit: Load and run candidate models in the actual browsers and on the weakest hardware you plan to support. Record provider, model, input dimensions, memory failures, and timing.
  • Output behavior: Check real images for expected visual results. For tasks such as inpainting, tests can inspect pixels inside the edited region as well as outside it; shape and numeric range alone may miss a bad result.
  • Representation handling: Verify tensor element types and decode fp16 outputs according to the representation returned by the specific runtime and browser.
  • Fallback triggers: Decide whether a fallback should respond only to thrown errors or also to detected output failures. The latter requires meaningful validation criteria for the task.
  • Separate provider checks: Test WebGPU and WASM independently. The author’s reported LaMa measurements show why a fallback’s mere availability does not establish equivalent correctness or speed.

What architecture did the author use for downloads and hosting?

The post describes delaying model and runtime downloads until the user consents, displaying the download size before the first task, and caching downloaded models in Cache Storage. For hosting, the author says a 25 MB upload limit led to splitting larger assets into chunks of at most 20 MiB, verifying chunks with SHA-256, joining them in a worker, and passing the resulting WebAssembly binary to ONNX Runtime. The editor was placed on a separate origin with connect-src 'self' and embedded in the content site by iframe.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

These are the author’s reported design choices, not an independent privacy or security audit. Consent-gated downloads and local caching describe when assets are fetched and where they are stored; by themselves, they do not establish the complete privacy or security properties of an application.

What is the practical lesson?

Browser inference needs more than a model that looks good offline and a call that returns successfully. Check resource limits in the target environment, validate image content rather than only execution status, and handle output representations that may vary with runtime support. The examples here come from one developer’s project; reproducing them on your own supported devices and versions is necessary before treating them as general browser behavior. Read alex.toolkit’s DEV Community post, “Three things that broke when I moved AI image models into the browser” (September 28, 2026).

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.