Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Yes—one reported test ran the full 4-bit Qwen3.8-Flash-Next on a 128 GB Mac Studio M4 Max and processed a prompt of about 240K tokens. The key was not removing experts: Nariaki Wada reported that an expert-pruned build lost quality on Japanese and general-knowledge evaluations, while the full build initially ran out of memory on longer tasks. His solution was to memory-map the model’s large n-gram embedding table so the runtime could read needed rows from storage instead of keeping the whole table resident as ordinary MLX parameters.

What the 128 GB Mac test established

In a September 24, 2026 report, Nariaki Wada described running Qwen3.8-Flash-Next on a Mac Studio M4 Max with 128 GB of unified memory. His results point to a specific memory-management workaround, not a general guarantee that every 128 GB Mac can run the model at every context length. The memory figures, quality observations, and timings below are from Wada’s setup and tests; they were not independently replicated in the cited materials.

The result addresses three different questions that are easy to conflate: whether a checkpoint can be loaded, whether it can process a long prompt without exhausting memory, and whether a smaller or altered build preserves the quality needed for a particular task. In Wada’s tests, the full 4-bit build loaded, but the default configuration failed on a 32K retrieval task. Adjusting the prefill step size allowed a 32K task to complete, but did not get the full model through 128K. Memory-mapping the n-gram table was the reported route to longer prompts without pruning experts.

Why a 125B model can still be memory-hungry

Qwen Team’s 2026 architecture paper describes Qwen3.8-Flash-Next as a sparse mixture-of-experts model with 125B total parameters and approximately 6B active per token. It also describes a further 51B-parameter n-gram embedding table held off accelerator. These numbers describe different parts of the design: active parameters help explain per-token computation, while total model and embedding data still matter to loading and memory management. Neither the active-parameter count nor the 4-bit label alone predicts the memory needed by a particular converted checkpoint and runtime.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Apple 2026 Mac Studio Desktop Computer M5 Max chip
  • BRAWN OF A NEW AGE — Mac Studio is a tremendously powerful pro desktop. The M5 Max chip enables remarkable on-device AI compute. Blast through creative projects and professional workflows with the advanced graphics architecture and faster memory and storage.
  • M5 MAX CHIP — Tap into breakthrough performance with a next-generation CPU, a more powerful GPU with third-generation ray tracing, and a Neural Accelerator built into each GPU core. Mac Studio gets a boost with more power to generate real-time media and accelerate complex workflows.
  • MEMORY AND STORAGE — Get up to 128GB unified memory and up to 614GB/s memory bandwidth for more speed when processing massive datasets, complex 3D scenes, and inference in AI workflows. And up to 2x faster storage* expedites tasks like file transfers and loading large projects.
  • A POWERFUL PLATFORM FOR AI — Apple silicon, and every major component that powers it, is designed to run demanding AI workflows like running huge LLMs, directly on device. And Apple Intelligence* helps you write, express yourself, and get things done effortlessly, while Siri AI* is your profoundly capable assistant — all with groundbreaking privacy protections.
  • A POWERFUL PLATFORM FOR AI — Apple silicon, and every major component that powers it, is designed to run demanding AI workflows like running huge LLMs, directly on device.

The architecture paper says, “Capacity is added outside the backbone by a single n-gram embedding layer whose tables are prefetched from host memory.” That design helps explain why the n-gram table is central to this Mac case. It does not, by itself, demonstrate that the model will fit in a particular Mac configuration or reproduce Wada’s results.

Why expert pruning was the wrong shortcut in this test

Wada compared the full model with a REAP-288 expert-pruned build. In his evaluation, pruning made the model easier to fit but reduced performance on Japanese and general-knowledge tasks. He reported the full build as the better-quality option in that comparison; the source does not provide a standardized, independently replicated quality study establishing that result for every language, prompt, or workload.

Rank #2
Apple Mac Studio, M4 Max 16-Core CPU / 40-Core GPU, 128GB Unified Memory, 1TB SSD
  • Extreme workflow performance: Take on extreme workflows, from detailed visual effects and 3D animation to film scoring. The powerful Neural Engine supports AI assistance in complex tasks, and the advanced GPU architecture supports Dynamic Caching, mesh shading, and ray tracing
  • Phenomenal memory and storage: Get up to 128GB unified memory and up to 8TB storage with M4 Max or up to 512GB unified memory and up to 16TB storage with M3 Ultra
  • Built for apple intelligence: Apple Intelligence is the personal intelligence system that helps you write, express yourself, and get things done effortlessly. With groundbreaking privacy protections, it gives you peace of mind that no one else can access your data-not even Apple
  • Compact desktop design: The compact 7.7 inch square Mac Studio enclosure is designed to fit under most displays and right on your desk
  • Advanced thermal system: The thermal system is designed to let you fly through intensive tasks at incredible speeds while keeping Mac Studio quiet, so it never interferes with your workflow

Pruning and quantization are separate choices. Pruning removes experts; quantization reduces the precision used to represent model weights. The full model in Wada’s memory experiment was already a 4-bit build. A reduced-size or pruned model may suit a user whose own task mix tolerates its quality trade-offs, but the label “smaller” is not evidence that it will preserve the full model’s behavior. The REAP-288 model card reports HumanEval results for that particular build; those figures should not be substituted for Wada’s Japanese and general-knowledge observations.

Option What it changes What Wada’s report supports
Full 4-bit build Quantizes the full model without removing experts. Wada reported a 111.5 GB peak after loading on his 128 GB Mac Studio M4 Max. Default settings failed at 32K on his retrieval task; reducing prefill step size got 32K through but not 128K.
REAP-288 expert-pruned build Removes experts to make a smaller build; this is distinct from quantization. Wada found losses on Japanese and general-knowledge evaluation. The report does not establish a universal quality ranking across users’ tasks.
Full 4-bit build with memory-mapped n-gram table Leaves the full model intact while serving requested embedding-table rows from a memory map rather than making the entire table ordinary resident MLX parameters. Wada reported testing prompts through about 240K tokens on the same Mac. This is his reported outcome, not a standardized guarantee.

How memory mapping changed the Mac run

Wada reported that the converted model did not include ple-store.json, a manifest used by the mlx-vlm external PLE storage path. In his setup, the fix was to create or use that external-storage arrangement so the runtime could fetch requested rows from a memory-mapped table rather than treating the full table as ordinary resident MLX parameters.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
(CTO) Apple Mac Studio: M5 Max 18C CPU - 40C GPU, 48GB, 1TB - Z1U4000E9
  • Configure-to-Order (CTO): Customized Mac with the processor, memory, storage, and other options; built by Apple at the factory.
  • M5 Max 18C CPU - 40C GPU CHIP—Tap into breakthrough performance with a next-generation CPU, a more powerful GPU with third-generation ray tracing, and a Neural Accelerator built into each GPU core.
  • MEMORY— 48GB unified memory so you can quickly access large media files or accelerate tasks for AI workloads.
  • STORAGE— 1TB solid state drive storage capacity.

The practical distinction is between having the whole table occupy the constrained resident-memory budget and accessing the rows needed for the work from storage. Memory mapping does not mean the table disappears or costs nothing: the model artifacts still need to be stored, and performance depends on the runtime and storage behavior. Wada’s report does not establish a required SSD model, storage capacity, or bandwidth threshold, so calculate storage needs from the actual artifacts you intend to use rather than assuming a particular accessory will suffice.

This is an implementation-specific account, not a universal setup recipe. The available report does not establish a stable, version-independent sequence of commands for every mlx-vlm release or converted checkpoint. Check the current mlx-vlm instructions and confirm that the model conversion and external PLE storage format match the runtime version you use. Wada explicitly did not verify changing iogpu.wired_limit_mb as a fix, so it should not be presented as an established alternative.

Rank #4
Apple Mac Studio, M4 Max 16-Core CPU / 40-Core GPU, 64GB Unified Memory, 2TB SSD
  • Take on extreme workflows, from detailed visual effects and 3D animation to film scoring. The powerful Neural Engine supports AI assistance in complex tasks, and the advanced GPU architecture supports Dynamic Caching, mesh shading, and ray tracing
  • Phenomenal memory and storage - Get up to 128GB unified memory and up to 8TB storage with M4 Max or up to 512GB unified memory and up to 16TB storage with M3 Ultra
  • Built for apple intelligence - Apple Intelligence is the personal intelligence system that helps you write, express yourself, and get things done effortlessly. With groundbreaking privacy protections, it gives you peace of mind that no one else can access your data-not even Apple
  • Fits right on your desk - The compact 7.7" square Mac Studio enclosure is designed to fit perfectly under most displays
  • Runs cool and quiet - The thermal system is designed to let you fly through intensive tasks at incredible speeds while keeping Mac Studio quiet, so it never interferes with your workflow
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What “240K tokens” means—and what it does not

Wada reported testing prompts up to about 240K tokens with the full model and the memory-mapped PLE setup. At approximately that prompt length, he reported 447.0 seconds (7.5 minutes) for Flash-Next and 2,114.9 seconds (35.2 minutes) for Qwen3.8-27B. These are total times through an answer of roughly 50 tokens in his test; prompt processing dominated. They are results from his particular machine and benchmark conditions, not a universal fivefold speed advantage for other prompt lengths, hardware, or workloads.

The approximately 240K-token test is also distinct from the architecture paper’s stated native context length of 262,144 tokens. A model’s stated context capacity does not guarantee that every runtime, checkpoint, memory arrangement, or application can process the full length. Wada’s report supports a result at about 240K in his configuration, not a claim that the Mac test reached 262,144 tokens.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Apple Mac Studio, M3 Ultra 32-Core CPU / 80-Core GPU, 256GB Unified Memory, 1TB SSD
  • UNMATCHED PERFORMANCE - Experience blazing-fast speeds with the M3 Ultra or M4 Max chip, featuring up to a 32-core CPU and up to 80-core GPU for demanding tasks like video editing and 3D rendering.
  • STUNNING VISUALS - Connect up to eight displays with the M3 Ultra, or five displays with the M4 Max, supporting resolutions up to 8K for immersive and highly detailed visual experiences across multiple screens.
  • AMPLE MEMORY - Configure up to 512GB of memory with the M3 Ultra or up to 128GB with the M4 Max, ensuring smooth multitasking and efficient handling of large datasets for professional workflows.
  • EXPANSIVE STORAGE - Choose from a range of SSD options, from 512GB to a massive 16TB, providing lightning-fast access to your files and ample space for all your creative projects and important data.
  • VERSATILE CONNECTIVITY - Equipped with Thunderbolt 5 ports delivering up to 120Gb/s, USB 3 ports, HDMI 2.1, and 10Gb Ethernet, this desktop offers seamless integration with all your peripherals and networks.

Mac MLX and vLLM are different deployment paths

Do not assume the Mac method transfers directly to vLLM. The vLLM deployment recipe documents PLE CPU offload for CUDA/ROCm environments and says its documented PLE offload currently runs on NVIDIA devices. That is not the same mechanism as mlx-vlm’s memory-mapped table path in Wada’s Mac report. Runtime support, available host or unified memory, storage performance, context settings, and model conversion all affect the result. Recheck the relevant project instructions for the exact runtime and hardware you plan to use.

How to decide whether this approach fits your workload

  • If quality matters more than simplifying memory management: the reported comparison favors keeping the full model rather than assuming REAP-288 pruning is harmless. Validate it with representative prompts in your own languages and task types.
  • If the full model loads but long prompts fail: a successful load is not proof that the selected context length will fit. Wada’s test showed that his default 32K retrieval run failed even after the full 4-bit build loaded.
  • If you need long contexts on a Mac: investigate whether your mlx-vlm version and converted artifacts support the external PLE storage arrangement described in Wada’s report. Treat the roughly 240K outcome as a result to reproduce, not a promise.
  • If you are deploying with vLLM: follow its hardware-specific recipe rather than applying Mac MLX assumptions. The documented PLE CPU-offload path is currently for NVIDIA devices.
  • If you are choosing hardware: the Mac Studio M4 Max with 128 GB is the machine in Wada’s test, not a proven minimum specification or a blanket recommendation for every Mac.

The Qwen Team’s architecture paper also reports that, across fourteen pre-training benchmarks, the model leads its 397B-A17B predecessor on eight and trails on the other six by at most 2.6 points. The paper reports roughly one third the activated parameters, one third the training tokens, and roughly one ninth the training FLOPs. These are architecture-paper and pre-training comparisons, not consumer Mac benchmarks or a substitute for testing your own prompts.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.