The official xAI repository located for this guide is named Grok 2; it does not establish a separate downloadable model officially named “Grok 2.5.” Its documented setup is for an eight-GPU server, not a typical laptop or gaming PC: xAI specifies eight GPUs with more than 40 GB of memory each and a checkpoint of approximately 500 GB. The repository’s instructions use SGLang and require the model’s correct chat template.
What you can download—and what “open source” means here
The available official artifact identified here is xAI’s Grok 2 checkpoint, described by the repository as weights for a model trained and used at xAI in 2024. The repository does not verify a distinct Grok 2.5 checkpoint, so the steps below apply to the Grok 2 repository—not to a confirmed model release called Grok 2.5.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
GMKtec EVO-X2 AI Mini PC Ryzen Al Max+ 395 Superchip 128GB LPDDR5X 2TB SSD | $3,649.99 | Buy on Amazon |
The repository names the Grok 2 Community License Agreement as the weights’ license. Read that agreement before relying on any particular right to use, modify, or redistribute the model. “Open-source” should not be taken to mean that the weights are under Apache 2.0: xAI’s separate Grok-1 release announcement describes Grok-1’s Apache 2.0 release, which does not establish Grok 2’s license.
Check the hardware and storage requirements first
xAI’s Grok 2 repository documents tensor parallelism across eight GPUs and calls for more than 40 GB of memory per GPU. It describes the checkpoint as 42 files totaling approximately 500 GB. That makes the published configuration a multi-GPU server or workstation setup, not a standard consumer PC. The repository does not establish that a particular workstation is compatible; check the complete system configuration before acquiring hardware.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
- EVOLUTION RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
- AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
- AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
- EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
- QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.
You will also need enough storage for the checkpoint and its supporting files. Storage capacity alone does not make the model runnable. The repository does not verify a lower-memory quantized variant, single-GPU compatibility, inference speed, or performance on consumer hardware.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Download and serve Grok 2 with the documented setup
The following is xAI’s example flow, not an independently validated installation. Its command assumes the model is downloaded to /local/grok-2 and that the host has the documented eight-GPU configuration.
-
Download the checkpoint. Install and authenticate with the Hugging Face CLI as needed for your environment, then run the repository’s command:
hf download xai-org/grok-2 --local-dir /local/grok-2Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.The repository describes a successful download as 42 files totaling approximately 500 GB and warns that transient download errors may require retrying.
-
Install SGLang. Use the latest inference engine version 0.5.1 or newer, as specified by the model repository.
-
Launch the inference server. With the weights at the path above, use xAI’s example command:
python3 -m sglang.launch_server --model /local/grok-2 --tokenizer-path /local/grok-2/tokenizer.tok.json --tp 8 --quantization fp8 --attention-backend tritonQuick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.The
--tp 8setting reflects the repository’s eight-way tensor-parallel configuration. The command also specifies FP8 quantization and the Triton attention backend; it is not evidence that other hardware configurations or settings will work. -
Send prompts in the required chat format. The post-trained checkpoint needs the correct chat template. The repository’s example formats a user turn as
Human: What is your name?<|separator|>followed by the assistant prefix. A request using an unrelated or generic template may not follow the model’s expected format.
When local Grok 2 is not the right fit
Local deployment gives you control over the infrastructure hosting the downloaded weights, but the documented configuration requires a specialized eight-GPU machine and a very large download. If you do not have that hardware, hosted access avoids operating this local setup. xAI’s current model catalog lists newer hosted Grok models; the available evidence does not establish feature parity or a like-for-like quality comparison with Grok 2.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.

