Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a comfortable quantized setup, target at least 24 GB of graphics memory—VRAM on a discrete GPU or variable graphics memory (VGM) on a supported system. The exact fit depends on the checkpoint, context length, runtime and serving configuration. For vLLM’s listed configurations, the recommended minimums range from 24 GB for INT4 to 67 GB for BF16. System RAM is separate, and the available sources do not establish one universal system-RAM minimum.

How much graphics memory does Qwen3.8-27B need?

There is no single memory number for every Qwen3.8-27B setup. The model’s weight footprint changes with precision and quantization, and weights are only part of the working-memory budget: the KV cache grows with context, while the runtime and operating system also use memory.

The following are configuration-specific figures from the vLLM Project’s maintained deployment recipe. The listed weight sizes describe the checkpoints; the VRAM figures are recipe recommendations, not guarantees that every context length or workload will fit.

Configuration Listed weight size vLLM recipe minimum VRAM
INT4 W4A16 19.5 GB 24 GB
NVFP4 build (21.9 GB checkpoint) 21.9 GB 32 GB
NVFP4 build (26.4 GB checkpoint) 26.4 GB 32 GB
Official block-scaled FP8 30,866,866,928 bytes (30.9 GB; 28.7 GiB) 38 GB
BF16 55,563,006,776 bytes (55.6 GB; 51.7 GiB) 67 GB

These figures come from the vLLM Project Qwen3.8-27B deployment recipe. Quantized options are not all interchangeable or uniformly four-bit; hardware support varies by build. Check the recipe’s model variant and hardware requirements before choosing a card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
ASRock Radeon AI PRO R9700 Creator 32GB Professional Graphics Card, 2920 MHz Boost Clock, GDDR6, AMD RDNA 4, AI-Accelerators, DisplayPort 2.1a, PCIe 5.0, Blower Cooler
  • Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
  • Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
  • Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
  • Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
  • Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.

Which GPUs are documented for this model?

AMD: Radeon AI PRO R9700 or Ryzen AI Max+

AMD says Qwen3.8-27B needs roughly 24 GB of VGM or VRAM to run comfortably and identifies its 32 GB Radeon AI PRO R9700 as a supported single-card option. AMD also lists Ryzen AI Max+ processor-based systems, which use variable graphics memory. This makes the R9700 a specifically documented discrete-GPU choice, but the guidance does not mean every model configuration or context will fit in 24 GB.

AMD reports preliminary Windows/Vulkan llama.cpp results of up to 51.8 tokens per second on the Radeon AI PRO R9700 and up to 24.5 tokens per second on Ryzen AI Max+ 395. These are manufacturer-reported figures: AMD says each average covered at least three runs, used MTP=2 on the Radeon and MTP=4 on the Ryzen system, and may vary. They are not a controlled cross-vendor comparison. AMD’s test systems had 64 GB of system memory for the Radeon test and 128 GB for the Ryzen test; those are test configurations, not stated minimums. Details are in AMD’s model deployment and performance post.

Rank #2
NVD RTX PRO 6000 Blackwell Professional Workstation Edition Graphics Card for AI, Design, Simulation, Engineering - 96GB DDR7 ECC Memory - 4th Gen RT/5th Gen Tensor Core GPU - OEM Packaging
  • PLEASE NOTE: Exporting an NVIDIA RTX Pro 6000 GPU outside the US requires strict adherence to the U.S. Export Administration Regulations (EAR) and issuance of an export license from the Bureau of Industry and Security (BIS). Compliance and Know Your Customer (KYC) screening may be required as a condition of order acceptance. [NVIDIA Blackwell Streaming Multiprocessor] The new SM features increased processing throughput, and new neural shaders that integrate neural networks inside of programmable shaders | DLSS 4: Multi Frame Generation ensures ultra-smooth frame pacing for lifelike simulations.
  • [Double-Flow-Through Design] The RTX PRO 6000 Blackwell features a double-flow-through cooling design, optimizing efficiency and airflow to sustain peak performance under 600W power loads. | [5th Gen Tensor Cores] Deliver up to 3X the performance of the previous generation and support for FP4 precision for faster AI model processing times with reduced memory usage, enabling local fine-tuning of LLMs and generative AI | [4th Gen Ray Tracing Cores] Double the ray-triangle intersection rate of the previous generation to create photoreal, physically accurate scenes and immersive 3D designs with RTX Mega Geometry, which enables up to 100X more ray-traced triangles.
  • [PCIe Gen 5] Support for PCIe Gen 5 provides double the bandwidth of PCIe Gen 4, improving data-transfer speeds from CPU memory and unlocking faster performance for data-intensive tasks like AI, data science, and 3D modeling. | [GDDR7 Memory] With 96 GB of GPU memory and 1.8 TB ps bandwidth, it can tackle massive 3D and AI projects, fine-tune AI models locally, explore large-scale VR environments, and drive larger multi-app workflows.
  • [DisplayPort 2.1] Achieve unparalleled visual clarity and performance, driving high resolution displays at up to 8K at 240 Hz and 16K at 60 Hz. Increased bandwidth enables seamless multi-monitor setups while HDR and higher color depth support ensures superior color accuracy for precision work, such as video editing, 3D design, and live broadcasting.
  • [Universal MIG] Divide a single RTX PRO 6000 Blackwell into multiple isolated instances, each with dedicated resources, allowing for concurrent execution of multiple workloads, optimized GPU utilization, and secure isolation of different applications or users. [WARRANTY] 3 YR Manufacturer's Warranty. Bulk OEM Packaging. Retail Packaging is NOT included.

NVIDIA: RTX 5090 for a documented NVFP4 path

The vLLM recipe documents an NVFP4 deployment on one RTX 5090, with a 32K maximum context and FP8 KV cache. That single-card example requires --enforce-eager: the recipe says CUDA graph capture runs out of memory without it. The same recipe describes a two-RTX-5090 path for a larger context. Treat these as recipe-specific setups, not a blanket claim that any RTX 5090 configuration will support any context or serving workload.

How to choose a setup

  1. Choose the model precision or quantized build. Use the vLLM recipe’s VRAM recommendation for that exact variant; do not infer memory needs from the “27B” parameter count alone.
  2. Match memory to your intended context. Longer contexts require more KV-cache memory. Leave room beyond the checkpoint’s weight size for cache, runtime and operating-system use.
  3. Verify runtime and hardware support. Confirm that the desired build and inference software support your GPU or VGM platform. For the documented single-card RTX 5090 NVFP4 example, use the stated eager-mode setting.
  4. Decide whether you need one card or a larger system. AMD names a 32 GB single-card option; vLLM documents an RTX 5090 single-card example and a two-card route for larger context. Compare form factor, memory capacity and the specific configuration you plan to run.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What about system RAM?

System RAM and GPU memory are distinct resources. A workload that offloads model data or uses unified memory can make system RAM relevant, but the sources here do not establish a universal minimum for running Qwen3.8-27B across runtimes and configurations. Do not treat AMD’s 64 GB and 128 GB benchmark-machine capacities as requirements; they describe the machines used for those tests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
ASRock Intel Arc Pro B60 Creator 24GB Graphics Card, Workstation GPU, Xe2-HPG, 2400MHz, 24GB GDDR6 192-bit, PCIe 5.0, 4X DP 2.1, Blower
  • System Compatibility Note: 2-slot card, 271x112x39mm, single 8-pin power, 200W TDP. Verify chassis clearance and PSU capacity before purchase.
  • Dedicated Support: Please contact us directly through Amazon for any product questions or assistance you may require.
  • 24GB GDDR6 on 192-Bit Bus: Massive 24GB memory with 456 GB/s bandwidth – ideal for LLMs, AI inference, 3D rendering, and generative design.
  • Intel Xe2-HPG Architecture: Built on Intel's next-gen architecture with 20 Xe cores and 160 XMX engines for AI acceleration (197 INT8 TOPS).
  • PCIe 5.0 Support: PCI Express 5.0 x16 interface for maximum bandwidth with the latest workstation platforms.

What the neighboring-model estimates can—and cannot—tell you

An Alibaba Cloud Community article dated August 5, 2026, measured Qwen3.6-27B weights at 55.6 GB for BF16, 28.6 GB for Q8_0 and 16.8 GB for Q4_K_M. Those are Qwen3.6 measurements, not Qwen3.8-27B checkpoint sizes or direct memory requirements. The article uses them as adjacent-model estimates and explains that KV cache, inference runtime and operating-system use add to weight memory; cache use rises with context. Use the Qwen3.8-specific vLLM recipe figures above for configuration planning, rather than substituting these neighboring-model values. The article is available at Alibaba Cloud Community.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.