Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To run Qwen locally, install Ollama, choose a Qwen model tag, then start it with ollama run <model-tag>. The first run downloads the model. Check ollama ps while it is loaded to see whether Ollama is placing it on the GPU, CPU, or both.

How do I run Qwen locally with Ollama?

  1. Install Ollama using the official download page. It provides a macOS/Linux command, curl -fsSL https://ollama.com/install.sh | sh, and a Windows PowerShell command, irm https://ollama.com/install.ps1 | iex, as well as manual installers. Check the page for the current installer instructions for your operating system.
  2. Choose a Qwen tag from the Qwen3 or Qwen3.5 model listing. The tag determines which variant Ollama runs.
  3. Open a terminal and run the chosen tag, for example ollama run qwen3, ollama run qwen3:8b, or ollama run qwen3.5. The first run downloads that model if it is not already installed.
  4. Enter a prompt in the interactive session. To leave it, use /bye.

Ollama also exposes a local API at http://localhost:11434/api/chat; the official model listings include Python and JavaScript client examples. This is useful when you want an application to send prompts to the model instead of typing them into the terminal.

Which Qwen model can my computer run?

There is no universal model-to-RAM or model-to-GPU rule in Ollama’s listings. Choose based on the task, the model’s listed download size, the memory available on your computer, the context length you need, and the speed you can accept. Download size is not a guaranteed RAM or VRAM requirement: runtime memory use also changes with context and concurrent requests.

Family and tag examples Capability shown in listing Example listed download sizes How to use the comparison
Qwen3: qwen3, qwen3:4b, qwen3:8b, qwen3:14b, qwen3:30b Text models, with dense and mixture-of-experts variants in the library qwen3:0.6b: 523 MB; qwen3:14b: 9.3 GB; qwen3:30b: 19 GB; qwen3:235b: 142 GB Use the exact tag’s listing to compare its size and context information before downloading.
Qwen3.5: variants from 0.8B through 122B Text-and-image options are listed qwen3.5:0.8b: 1.2–1.3 GB; qwen3.5:9b: 6.6–7.6 GB; qwen3.5:122b: 81 GB Consult the listing for the particular tag and its capabilities; a larger listed download is not a direct memory estimate.

These are download sizes shown in Ollama’s live model listings, not required RAM or VRAM. The listings also show context information for tags, but the context advertised for a model is not necessarily the context Ollama uses by default. There are no hardware-specific benchmarks in these listings that establish how fast a given tag will run on your computer.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
ASUS ROG G700 (2025) Gaming Desktop PC, Intel® Core™ Ultra 7 265F Processor, NVIDIA® GeForce RTX™ 5070, 1TB M.2 NVMe™ PCIe® 4 SSD, 16GB DDR5 RAM, Windows 11 Home, G700TF-DS774
  • Fearless ROG Design – The G700’s dual-glass chassis showcases iconic ROG design with the ROG Slash and Aura Sync RGB lighting. Its 58L capacity supports triple-slot GPUs.
  • Unstoppable Power – Equipped with the Intel Core Ultra 7 265F processor, NVIDIA GeForce RTX 5070 GPU, 16GB DDR5 RAM, and 1TB SSD PCIe 4.0 storage for seamless gaming and multitasking.
  • Optimized Thermals – Stay cool with a quad-fan system, while dust filters and efficient airflow ensure long-term reliability.
  • Advanced Connectivity – Game without lag with 2.5Gbps Ethernet, Wi-Fi 6, and versatile ports. Dolby Atmos audio and AI noise cancellation enhance sound and communication.
  • Ready for Upgrades – Designed with tool-less access, easily swap out components, ensuring future-proof performance for years to come.

Choose by task and available memory

  • For text tasks, select a Qwen3 or Qwen3.5 tag whose scale and listed context suit your workload.
  • For image input, use a vision-capable tag such as a Qwen3-VL model or a text-and-image Qwen3.5 variant.
  • If a model is too slow or does not fit comfortably, try a smaller variant or reduce context and parallel requests. Check actual processor placement rather than assuming a download-size figure predicts performance.

How do I check whether Ollama is using my GPU?

Start or keep the model running, then open another terminal and enter ollama ps. The PROCESSOR column reports whether the loaded model is on GPU, CPU, or split between them. This is a placement check, not a speed benchmark: performance still depends on the computer and workload.

A model can run without being entirely on the GPU. If the output shows CPU-only or a CPU/GPU split, use a smaller model or a lighter workload if the current speed is unsuitable. Ollama notes that large models can be slow on a computer without a strong GPU; its documentation does not specify one GPU or memory minimum that applies to every Qwen tag.

Rank #2
Sale
CyberPowerPC Gaming PC, AMD Ryzen 5 5500, Radeon RX 6500 XT 4GB
  • System: AMD Ryzen 5 5500 3.6GHz 6 Cores | AMD B550 Chipset | 8GB DDR4 | 500GB PCIe 4.0 NVMe SSD | Windows 11 Home
  • Graphics: AMD Radeon RX 6500 XT 4GB Graphics | 1x HDMI | 1x DisplayPort
  • Connectivity: 4 x USB-A 3.2 | 4 x USB-A 2.0 | 1 x LAN | WiFi 5 | Bluetooth 5.0 | 7.1 Channel Audio
  • Tempered Side Case Panel | Custom RGB Lighting | Keyboard and Mouse
  • 1 Year Parts & Labor Warranty, Free Lifetime Tech Support

How do I change Qwen’s context length?

Ollama documents a default context window of 4096 tokens. A model listing may advertise a larger context, but that does not automatically make it the runtime setting. Larger context and more simultaneous requests can increase memory use.

  • For an interactive session, enter /set parameter num_ctx 4096 to set the context parameter to 4096 tokens.
  • To start the Ollama server with a different context length, the FAQ gives OLLAMA_CONTEXT_LENGTH=8192 ollama serve.
  • For API requests, set the num_ctx option in the request.

Choose a context length for the amount of text your task needs, rather than setting it to the largest value in a model listing by default.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
WIWB Gaming PC Desktop, GeForce RTX 3050 8GB GDDR6, AMD Ryzen 7 4700LE
  • 8-Core 16-Thread Processing Power – Powered by the Ryzen 7 4700LE processor with Zen 2 architecture, delivering 8 cores and 16 threads with a boost clock up to 4.2GHz. Effortlessly handle multitasking, streaming, content creation, and demanding applications simultaneously without slowdowns.
  • GeForce RTX 3050 8GB Graphics – Equipped with 8GB GDDR6 dedicated VRAM and real-time ray tracing support. Experience smooth 1080p gaming at 55-60 FPS in AAA titles like Cyberpunk 2077, 70+ FPS in Fortnite, and 90-100 FPS in Apex Legends with DLSS enabled. The 8GB buffer handles modern game textures comfortably – a step above 6GB variants
  • High-Speed Memory & Storage – Paired with 16GB of DDR4 3200MHz dual-channel RAM (16GB), the PC ensures responsive multitasking—whether streaming while gaming or editing videos. It also includes a 512 GB NVMe M.2 SSD for lightning-fast boot times, quick game loads, and ample storage for your game library, creative projects, and files.
  • Next-Gen WiFi 6 Connectivity – Stay connected with the latest WiFi 6 technology for faster speeds, lower latency, and improved network efficiency. Whether you're gaming online, streaming 4K content, or joining video conferences, enjoy stable, high-speed wireless connectivity.
  • Ready-to-Use Value Desktop – Pre-built and ready to go right out of the box. Perfect for gamers, students, content creators, and home office users seeking reliable performance without the hassle of building a PC themselves. The mature AM4 platform with DDR4 memory offers excellent value and proven stability.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How do I use Qwen with images in Ollama?

Use a vision-capable model tag; a text-only tag is not the right choice for image prompts. Ollama’s Qwen3-VL listing labels its models as text-and-image and states that Qwen3-VL requires Ollama 0.12.7. Its example command is ollama run qwen3-vl:8b. The Qwen3.5 listing also includes text-and-image variants.

Before troubleshooting an image prompt, confirm that the tag you installed supports images and that your Ollama version meets the model page’s stated requirement. Follow the model listing’s instructions for supplying an image in your prompt.

Best Value
KOTIN Prebuilt Gaming PC RTX 5070 12GB, Ryzen 7 9700X, 32GB DDR5, 1TB SSD
  • POWERED BY RTX 5070 12GB + RYZEN 7 9700X - The GeForce RTX 5070 12GB GDDR7 graphics card pairs with an 8-core AMD Ryzen 7 9700X processor to drive smooth 1440p and 4K gameplay, giving this gaming PC the headroom for modern titles, streaming, and creative work.
  • 32GB DDR5 6000MHz MEMORY & 1TB NVMe SSD - 32GB of high-speed DDR5 memory and a 1TB PCIe 4.0 NVMe solid state drive deliver quick load times, smooth multitasking, and generous storage, keeping this prebuilt gaming desktop responsive under heavy workloads.
  • BUILT-IN 11.3-INCH Smart DISPLAY - An integrated smart screen shows real-time CPU and GPU temperatures, usage, and weather while you play, adding a distinctive and functional touch to your battlestation.
  • 850W 80+ GOLD POWER SUPPLY, 360MM LIQUID COOLING & WiFi 7 - An 850W 80 Plus Gold certified power supply provides stable, efficient power with headroom for future upgrades, while a 360mm AIO liquid cooler, WiFi 7, and an ARGB mid-tower case keep the Ryzen 7 CPU cool and connected in a clean build.
  • READY TO PLAY OUT OF THE BOX - Arrives fully assembled and tested with Windows 11 Home pre-installed, so your prebuilt gaming computer is ready to set up in minutes. Assembled in the USA, and backed by a one-year limited warranty and lifetime free technical support.
Rank #4
Sale
msi Codex Z2 Gaming Desktop, AMD R7-8700F, RTX 5070, 32GB DDR5, 2TB SSD
  • POWERHOUSE 8-CORE GAMING PERFORMANCE — Driven by the AMD Ryzen 7 8700F with 8 cores and 16 threads, boosting up to 5.0 GHz for smooth, responsive gameplay and the ability to handle AAA titles, streaming, and background tasks all at once
  • NEXT-GEN BLACKWELL ARCHITECTURE — The NVIDIA GeForce RTX 5070 is powered by NVIDIA's cutting-edge Blackwell GPU architecture, delivering a massive generational leap in rasterization and ray tracing performance so you can experience your games the way they were meant to be played.
  • Simplistic Design: Enjoy the latest generation of Windows 11 Home for your everyday needs. *MSI recommends Windows 11 Pro for business use.
  • Cool While Gaming: In conjunction with an ARGB fan Air Cooler, the Codex R2 features four system cooling fans; three in the front and one in the rear to pull in cool air and push heat out of the PC.
  • Turn on the Bright Lights: With the built-in RGB lighting, take your gaming experience to the next level by pressing the MSI LED button to cycle through lighting options. Customize lighting even further with MSI Center software.

Useful official references

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.