Recommended Free Tools
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
In ConwayResearch’s own tests, Underdog Saluki 27B scored 88 out of 120 tool-calling tasks, versus 84 for full-size Qwen3.8-27B. That is a narrow win on one creator-run evaluation—not evidence that the compact model is generally better. Saluki is a 7.89 GB IQ2-mix GGUF built for local inference with llama.cpp, and the smaller file comes with trade-offs, particularly on the model card’s math and reasoning results.
What Underdog Saluki 27B is
ConwayResearch’s model card describes Underdog Saluki 27B 1.0 as a compact GGUF quantization based on Qwen3.8-27B. Its main file, Underdog-Saluki-27B-1.0-IQ2-mix.gguf, is listed at 7.89 GB and is licensed under Apache 2.0. The card’s tagline calls it “Qwen3.8-27B in under 8 GB, tuned to keep tool calling intact.”
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
GMKtec EVO-X2 AI Mini PC Ryzen Al Max+ 395 Superchip 128GB LPDDR5X 2TB SSD | $3,649.99 | Buy on Amazon |
“2-bit” describes the low-bit quantization approach, not a guarantee of a particular speed, memory requirement, or output quality. The listed file size is not a minimum hardware specification; the card does not establish a minimum GPU or system configuration.
Does it really beat the original at tool calling?
On the publisher’s custom evaluation, yes: Saluki passed 88 of 120 tasks, while full-size Qwen3.8-27B passed 84. ConwayResearch says the tasks were drawn from BFCL v4, frozen before testing, and run with thinking disabled at temperature 0. The result is a four-task difference on this particular set, not a broad proof of superiority. The model card itself calls the test modest and notes that a few tasks’ difference could reflect run-to-run variation.
#1 Best Overall
- EVOLUTION RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
- AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
- AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
- EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
- QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.
The publisher also reports a higher score on a separate 100-task BFCL v4 parallel-calling evaluation: Saluki scored 42, compared with 35 for the full model, using the official checker with thinking disabled. But the same card cautions that roughly one in five parallel-call replies has small formatting slips. A call can be conceptually correct and still fail a strict checker if its structure or syntax is off.
These numbers are ConwayResearch’s reported measurements; the available sources do not independently reproduce them. They should be read as a useful signal about this release, not an independently verified benchmark or a promise that it will outperform the original in your tools, prompts, or application.
Where the compact model gives up ground
The card’s other reported results show a mixed picture. Saluki is close to or ahead of public-reference figures on some instruction-following measures, but it trails on several math and reasoning tests. These comparisons are not all controlled head-to-heads: the card says the public results use a different harness.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →| Evaluation | Saluki | Qwen3.8-27B comparison | How to read it |
|---|---|---|---|
| Tool calling, 120 tasks | 88/120 | 84/120, full-size model | Publisher-run custom test; thinking disabled, temperature 0. |
| Parallel calls, 100 tasks | 42/100 | 35/100, full-size model | Publisher-reported BFCL v4 test using the official checker; card notes formatting slips in about a fifth of Saluki’s parallel-call replies. |
| SWE-bench Verified, 50 issues | 30 | 33, full-size model | Reported by the publisher; the card does not establish that this comparison used the same harness for both figures. |
| IFEval, prompt-loose | 93.5 | 91.5 public | Public comparison uses a different harness, according to the card. |
| IFBench, prompt-loose | 72.7 | 71.0 public | Public comparison uses a different harness, according to the card. |
| MBPP+ | 78.0 | 83.9 public | Public comparison uses a different harness, according to the card. |
| MuSR | 67.5 | 79.6 public | Public comparison uses a different harness, according to the card. |
| AIME 2025, avg@4 | 79.2 | 96.7 public | Public comparison uses a different harness, according to the card. |
| AIME 2026, avg@4 | 80.0 | 94.6 public | Public comparison uses a different harness, according to the card. |
The card characterizes Saluki’s competition-math performance as about 82–85% of the full model’s. It also says the model is weakest on letter-level instruction puzzles. With thinking enabled, it may reason at length before answering—a potential drawback when you want concise responses.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What you need to run it locally
The documented runtime is stock llama.cpp. The model card provides an example server command with Jinja chat-template support, GPU-layer offload, flash attention, and a 32,768-token context. Those flags illustrate one setup; they do not define a guaranteed hardware requirement or performance level.
llama-server -m Underdog-Saluki-27B-1.0-IQ2-mix.gguf --jinja -ngl 99 -fa --ctx-size 32768
The --jinja flag enables the Qwen3.8 chat template used for tool calls and thinking, according to the card. Adapt the offload and context settings to the available hardware and your use case; the published example alone does not establish what will fit or run well on a particular machine.
Vision requires a separate file
The main GGUF is text-only. For vision, the card describes an optional multimodal projector add-on in either F16 (928 MB) or Q8_0 (629 MB), loaded with llama.cpp’s --mmproj option. The add-on is separate from the 7.89 GB text model.
Free tools Windows power users keep installed
One-click scans. No signup required.
Quick Recap
Who should consider Saluki
- Consider it if you want to try a low-bit Qwen3.8-27B derivative locally and tool calling is a priority, while accepting that the strongest tool-call results come from the publisher’s own tests.
- Prefer the full-size model if math and reasoning performance matter more than a smaller model file, given Saluki’s lower reported results on AIME, MuSR, and MBPP+.
- Test your own workflow if your tools depend on strict parallel-call formatting: the card’s reported formatting slips can matter even when benchmark task scores look favorable.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

