Free tools Windows power users keep installed
One-click scans. No signup required.
OrcaSAQ-2 is a compact EXL3 quantization of Qwen3.8-27B, but available results do not establish that it is better or worse than other quantizations. OrcaRouter reports near-identical WikiText-2 perplexity to the BF16 reference, alongside 93.2% top-1 token agreement. Because competing quantizations were evaluated under different setups, choose by memory needs, runtime, vision support, and results on your own matched workload—not by treating the published scores as one leaderboard.
What OrcaSAQ-2 is—and what changes from the base model
OrcaSAQ-2-27B is an EXL3 checkpoint based on Qwen3.8-27B. OrcaRouter lists an average 3.21 bits per decoder weight (bpw) and a 12.3 GB checkpoint, compared with 54 GB for the BF16 reference. The card lists thinking mode, tool calling, and MTP speculative decoding. These are publisher specifications, not independent measurements.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
GMKtec EVO-X2 AI Mini PC Ryzen Al Max+ 395 Superchip 128GB LPDDR5X 2TB SSD | $3,649.99 | Buy on Amazon |
The original Qwen3.8-27B is a 27-billion-parameter dense model with 64 layers, image and video understanding, and a native 262,144-token context that can be extended to one million tokens, according to the Qwen model card. OrcaSAQ-2 does not include the visual encoder, so its listed capabilities should be treated as text-only. The base model’s vision capability does not carry over simply because the quantized weights are based on it.
QwenLM’s official Qwen3.8 repository links to the official weights and model information. A model’s advertised context limit is an architecture capability, not a guarantee that a particular GPU can serve that context length with a chosen KV cache, batch size, and runtime.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minute#1 Best Overall
- EVOLUTION RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
- AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
- AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
- EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
- QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.
What OrcaSAQ-2’s BF16 comparison does—and does not—show
OrcaRouter reports a same-path comparison against Qwen3.8-27B BF16 using 16,376 predicted tokens from WikiText-2:
| Checkpoint | Reported size and precision | WikiText-2 perplexity | Top-1 token agreement | Other reported result |
|---|---|---|---|---|
| Qwen3.8-27B BF16 reference | 54 GB; 16-bit | 5.6468 | 100% reference | Baseline |
| OrcaSAQ-2-27B | 12.3 GB; 3.21 average decoder bpw | 5.6482 | 93.2% | Mean KLD 0.031; perplexity change reported as +0.02% |
These figures are reported in the OrcaRouter model card. Perplexity summarizes how well a model assigns probabilities across a text corpus; top-1 agreement counts how often its most likely next token matches the reference. Similar average perplexity can therefore coexist with different next-token choices. The results support a narrow conclusion—close average WikiText-2 perplexity under this test—not a claim that OrcaSAQ-2 is lossless or that its answers on other tasks are unchanged.
The independent Local Model Watch analysis likewise notes that the card does not provide same-condition measurements against GGUF or other EXL3 quantizations. The BF16 comparison is not a head-to-head ranking of low-bit alternatives.
How OrcaSAQ-2 compares with the published alternatives
The clearest listed alternative is ISTA-DASLab’s GSQ-RCO family of GGUF quantizations. Its model card includes 2.50, 2.75, 3.00, and 3.50 bpw variants, with listed file sizes spanning 8.4 to 11.8 GB. It reports results against both BF16 and Unsloth Dynamic versions of the same base model. The figures below are evidence from that card’s test setup, not directly comparable scores against OrcaSAQ-2.
| Option | Format and footprint | Published evidence | What the evidence can establish |
|---|---|---|---|
| OrcaSAQ-2-27B | EXL3; 3.21 average decoder bpw; 12.3 GB | WikiText-2 perplexity 5.6482 and 93.2% top-1 agreement against BF16, using 16,376 predicted tokens; OrcaRouter-reported serving figures under a 15.7 GiB GPU memory cap | Its reported fidelity and serving behavior versus BF16 in the publisher’s setup, not its rank against other quantizations |
| GSQ-RCO GGUF family | GGUF at 2.50, 2.75, 3.00, and 3.50 bpw; listed files range from 8.4 to 11.8 GB | The 3.50-bpw IQ3_S variant is 11.8 GB; its card reports AIME25 100.00, GPQA-Diamond 89.39, and LiveCodeBench v6 85.71. The listed BF16 figures are 100.00, 89.90, and 85.71, respectively. | Results for GSQ-RCO in its own reported evaluation, not a controlled comparison with OrcaSAQ-2 |
| Official FP8 and AutoRound INT4/INT8 checkpoints | FP8 and INT4/INT8 checkpoints; comparable per-checkpoint file sizes are not stated in the community comparison | The report describes a shared workload but calls the comparison only partially comparable: checkpoint and quantization vary together, quality results are single trials, and the FP8 throughput run used a tokenizer fallback. | Exploratory evidence about that workload, not an isolated estimate of quantization’s effect or a general ranking |
The ISTA-DASLab GSQ-RCO card also lists a separate BF16 vision projector for multimodal use. Check that the companion artifact and the runtime you plan to use support the vision path; the existence of a projector alone does not establish that every deployment setup will work.
OrcaRouter’s card cautions that public agent scores use different stacks and should not be treated as a strict model-only ranking. Across the cited material, there is no controlled, repeated test of OrcaSAQ-2 against the other quantizations using the same hardware, harness, prompts, and decoding settings. Putting the reported numbers in one leaderboard would imply a comparability the evidence does not support.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Which Qwen3.8-27B quantization fits your GPU and deployment?
Estimate usable memory, not just checkpoint size
Checkpoint footprint is only one part of inference memory. Runtime overhead, KV cache, batch size, and the context length actually used all affect whether a model fits. OrcaRouter reports OrcaSAQ-2 measurements under a 15.7 GiB GPU memory cap and gives around 32K interactive context as a practical starting point on a 16 GB GPU. That is a publisher-reported starting point for its setup, not a guarantee for every GPU or serving configuration. The 262,144-token architecture limit does not mean that context will fit in the same memory budget.
Match the format to your software stack
OrcaSAQ-2 uses EXL3, and OrcaRouter provides vLLM instructions. The cited GSQ-RCO alternatives use GGUF and list llama.cpp, Ollama, and LM Studio. If you already deploy with one of those runtimes, compatibility and operational convenience may outweigh a small difference in a score that was not measured head-to-head.
Recommended Free Tools
Check vision requirements before choosing
If your workload includes images or video, the base Qwen3.8-27B model’s native vision support is not enough to establish that a particular quantized package supports it. OrcaSAQ-2 omits the visual encoder. GSQ-RCO lists a separate BF16 vision projector, which you should verify is compatible with your chosen model files and runtime.
Benchmark the tasks you actually run
For a meaningful quality comparison, keep the model base and task fixed, then match prompts, harness, decoding settings, hardware, and runtime; repeat trials where results can vary. Separate perplexity, token agreement, reasoning tests, coding benchmarks, and long-horizon agent tasks: a result in one category does not automatically predict another. A community run may offer useful workload-specific clues, but its checkpoint differences, single-trial quality measurements, or runtime quirks limit causal conclusions about quantization alone.
What OrcaRouter’s serving figures say about MTP
Under its stated 15.7 GiB GPU memory cap, OrcaRouter reports 65.3 tokens per second at one stream without MTP and 90.1 tokens per second at one stream with MTP. At eight and 16 streams, its reported aggregate throughput is lower with MTP enabled. The card says MTP consumes KV capacity and advises benchmarking both settings for highly batched workloads. Treat these as publisher measurements, not independent performance results: the best setting depends on serving pattern and available memory.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

