Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For the model formats covered by vLLM’s recipe, the stated VRAM floors range from 24 GB for INT4 to 67 GB for BF16. The same recipe lists 32 GB for NVFP4 and 38 GB for FP8. These are recipe-specific planning figures, not guarantees for every GPU, context length, or workload.

Qwen3.8-27B is a dense vision-language model, so the right amount depends on more than its weights: checkpoint format, runtime, context and cache settings, and serving workload all matter.

How much VRAM does Qwen3.8-27B need by format?

The vLLM project’s live recipe, accessed October 7, 2026, lists the following minimum VRAM figures for specific checkpoints or builds. The page does not state a publication date. Treat these as recipe floors, not universal minimums for every way of running the model.

Format or build Checkpoint size reported by vLLM Recipe VRAM floor Important qualification
BF16 55,563,006,776 bytes: 55.6 GB on disk, or 51.7 GiB of weights 67 GB The recipe describes full-precision BF16 as a one-GPU configuration. vLLM recipe
Official block-scaled FP8 30,866,866,928 bytes: 30.9 GB on disk, or 28.7 GiB of weights 38 GB The checkpoint size is not the full runtime memory requirement. vLLM recipe
NVIDIA NVFP4 Not stated in the recipe 32 GB The recipe lists an NVIDIA NVFP4 checkpoint for RTX 5090 hardware. Its hardware-specific launch uses a 32,768-token maximum model length, FP8 KV cache, and eager execution on one RTX 5090. vLLM recipe
Red Hat AI INT4 W4A16 Not stated in the recipe 24 GB The recipe lists Hopper hardware among supported platforms. vLLM recipe

GB and GiB are different units: a GiB is larger than a decimal GB. The BF16 and FP8 rows show why a model’s download size should not be treated as its complete VRAM budget. The recipe calculates minimums from checkpoint bytes using a multiplier and also documents hardware-specific settings.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
ASRock Intel Arc Pro B60 Creator 24GB Graphics Card, Workstation GPU, Xe2-HPG, 2400MHz, 24GB GDDR6 192-bit, PCIe 5.0, 4X DP 2.1, Blower
  • System Compatibility Note: 2-slot card, 271x112x39mm, single 8-pin power, 200W TDP. Verify chassis clearance and PSU capacity before purchase.
  • Dedicated Support: Please contact us directly through Amazon for any product questions or assistance you may require.
  • 24GB GDDR6 on 192-Bit Bus: Massive 24GB memory with 456 GB/s bandwidth – ideal for LLMs, AI inference, 3D rendering, and generative design.
  • Intel Xe2-HPG Architecture: Built on Intel's next-gen architecture with 20 Xe cores and 160 XMX engines for AI acceleration (197 INT8 TOPS).
  • PCIe 5.0 Support: PCI Express 5.0 x16 interface for maximum bandwidth with the latest workstation platforms.

Will Qwen3.8-27B run on my GPU?

Compare your card’s usable VRAM with the floor for the exact checkpoint and runtime you plan to use. A card that meets a listed floor may still fail to load or serve your intended workload if other processes occupy memory or if your context, cache, or concurrency settings require more.

  • Under 24 GB: None of the four vLLM recipe floors fits as stated.
  • 24 GB: The recipe lists this floor for its Red Hat AI INT4 W4A16 build, with Hopper among the supported platforms. This is not a promise that any 24 GB GPU is compatible.
  • 32 GB: The recipe lists this floor for NVFP4, specifically documenting an RTX 5090 route and its own launch settings.
  • 38 GB: This is the recipe floor for the official block-scaled FP8 checkpoint.
  • 67 GB: This is the recipe floor for BF16 weights.

These figures describe vLLM’s specified builds; they do not establish that every format is supported by every GPU or inference program. Qwen’s model card lists compatibility with Transformers, vLLM, SGLang, TokenSpeed, and other tools, but a general compatibility listing is not a guarantee that each quantized checkpoint works in each runtime. Qwen3.8-27B model card

Why isn’t weight size the same as the VRAM requirement?

Checkpoint size measures stored weights, while an inference setup also needs memory for runtime operations and other model state. In particular, the KV cache grows with the tokens retained for context and with the number of simultaneous sequences. The vLLM recipe’s separate floors for BF16 and FP8 are larger than their respective on-disk checkpoint sizes.

Quantized formats also differ in implementation. vLLM cautions that quantized checkpoints do not all use a uniform four bits per weight, so multiplying the parameter count by an assumed bit width is not a reliable way to predict the memory use of a particular build.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
NVIDIA GeForce RTX 3090 Founders Edition Graphics Card (Renewed)
  • Item Package Dimension - 15.0L x 12.25W x 4.25H inches
  • Item Package Weight - 6.0 Pounds
  • Item Package Quantity - 1
  • Product Type - VIDEO CARD

What settings change the amount of VRAM you need?

Context length and KV cache

A longer maximum context can require more cache memory, depending on the runtime and cache format. The vLLM NVFP4 example specifies a 32,768-token maximum model length and an FP8 KV cache; its 32 GB floor belongs to that recipe context, not every context or cache configuration.

Qwen’s repository includes vLLM and SGLang examples configured for a 262,144-token maximum model length and tensor parallelism across four devices. That is a multi-GPU example, not evidence that a single consumer GPU can provide the same context. Qwen repository

Rank #4
ASRock Radeon RX 7900 XTX Phantom Gaming 24GB OC Graphics Card, 2615 MHz Boost Clock, 24GB GDDR6, DisplayPort 2.1, HDMI 2.1, Triple Fan Cooling
  • Flagship Gaming Performance, AMD Radeon RX 7900 XTX GPU with 2615 MHz boost clock and 24GB GDDR6 memory for elite 4K gaming
  • Advanced RDNA 3 Architecture, 96 compute units with RT+AI accelerators and 96MB AMD Infinity Cache technology
  • Premium Cooling Solution, Phantom Gaming 3X Cooling System with Striped Ring Fans and reinforced metal frame
  • High-Speed Memory, 24GB GDDR6 on 384-bit memory bus delivers exceptional bandwidth for 4K gaming and content creation
  • Silent Operation, 0dB Silent Cooling technology ensures zero fan noise during low-intensity tasks

Batch size and concurrent requests

Serving more sequences at once can increase cache and runtime memory use. The listed floors do not establish a universal batch size or concurrency level; check the settings for the specific serving command you intend to run.

Images and video

Qwen’s model card describes Qwen3.8-27B as a native vision-language model that understands images and videos. Vision use changes the workload, so the recipe floors alone do not establish a guaranteed fit for a particular image or video input pattern.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
EVGA GeForce RTX 3090 FTW3 Ultra Gaming, 24GB GDDR6X, iCX3 Technology, ARGB LED, Metal Backplate, 24G-P5-3987-KR
  • Digital Max Resolution:7680 x 4320.590.4GT/s Texture Fill Rate
  • Real boost clock: 1800 MHz; Memory detail: 24576 MB GDDR6X.
  • Real-time ray tracing in games for cutting-edge, hyper-realistic graphics.
  • Triple HDB fans 9 iCX3 thermal sensors offer higher performance cooling and much quieter acoustic noiseAvoid using unofficial software
  • All-metal backplate & adjustable ARGB

Other GPU memory use

VRAM occupied by the desktop, other applications, or other models is not available to the inference process. Plan against memory available to the runtime rather than relying only on the capacity printed on the GPU specification.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to choose a realistic VRAM target

  1. Choose the checkpoint first. Identify whether you will run the BF16, official block-scaled FP8, NVFP4, or INT4 build; the recipe floors differ substantially.
  2. Check the exact runtime and GPU pairing. Confirm the chosen checkpoint is supported by your inference software and hardware. For vLLM, use the recipe associated with that build rather than assuming a format name guarantees compatibility.
  3. Set your workload requirements. Decide the maximum context, KV-cache precision, batch size, concurrent requests, and whether you will use image or video inputs.
  4. Leave practical headroom. The recipe floors are minimum planning values, not a recommended margin for every workload. Allow for memory consumed by other processes and by settings beyond the recipe’s documented configuration.

The vLLM recipe is the source for the format-specific floors and hardware notes: Qwen3.8-27B vLLM recipe. The model’s capabilities and listed software integrations are described in Qwen’s model card.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.