What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
DeepSeek-R1 was not trained only with reinforcement learning: that description fits its R1-Zero predecessor, which applied RL directly to a base model without preliminary supervised fine-tuning. The released R1 used a multi-stage recipe that included both supervised fine-tuning and RL. DeepSeek reported R1 as comparable to the specific OpenAI-o1-1217 snapshot on reasoning tasks, while its January 2025 API rates were about 96% lower per token than the listed o1 rates—not proof that equivalent work costs 95% less overall.
Was DeepSeek-R1 trained only with reinforcement learning?
No. The phrase “pure reinforcement learning” describes DeepSeek-R1-Zero, not the complete training pipeline for the released DeepSeek-R1 model. The distinction matters because the two models had different training recipes and reported weaknesses.
R1-Zero applied RL directly to a base model
DeepSeek’s 2025 paper describes R1-Zero as a model trained with large-scale reinforcement learning applied to a base model, without supervised fine-tuning (SFT) as a preliminary step. DeepSeek reported that this approach elicited reasoning behaviors such as self-verification, reflection, and longer chains of thought. The paper also noted drawbacks, including poor readability and language mixing.
Released R1 added supervised data and multiple training stages
DeepSeek introduced cold-start data before reinforcement learning for R1. Its repository summarizes the full R1 pipeline as two SFT stages and two RL stages. The supervised data helped establish reasoning and non-reasoning capabilities before and between reinforcement-learning stages. Calling this released model “RL-only” would therefore be inaccurate.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitches#1 Best Overall
Is DeepSeek-R1 as good as OpenAI o1?
DeepSeek’s paper reported performance comparable to OpenAI-o1-1217 on reasoning tasks. That is a claim about specific model snapshots and DeepSeek’s own evaluation, not a timeless finding that R1 matches every version of o1 across all tasks. The OpenAI documentation accessed in 2026 marks o1 as deprecated.
DeepSeek reported the following results for R1. They should be read as the paper’s evaluation results, not as an independent head-to-head test:
Rank #2
| Evaluation | DeepSeek-R1 result reported by DeepSeek | Scope |
|---|---|---|
| AIME 2024 | 79.8% pass@1 | Pass@1 is the reported single-attempt measure. |
| MATH-500 | 97.3% | Reported mathematical benchmark result. |
| Codeforces | 2,029 Elo | Reported programming-contest rating. |
| MMLU | 90.8% | Knowledge benchmark; DeepSeek reported R1 slightly below o1-1217 on this benchmark. |
| MMLU-Pro | 84.0% | Knowledge benchmark; DeepSeek reported R1 slightly below o1-1217 on this benchmark. |
| GPQA Diamond | 71.5% | Knowledge benchmark; DeepSeek reported R1 slightly below o1-1217 on this benchmark. |
Benchmark percentages are not interchangeable evidence of general capability. A fair comparison depends on the benchmark and evaluation date, the exact model snapshot, whether scoring uses pass@1 or multiple samples, and how much input and output the models use. For real workloads, latency, retries, and the quality required for the task also matter.
How is DeepSeek-R1 “95% cheaper”?
The headline refers to announced API prices per million tokens, not training expense or a measured reduction in the total cost of doing equivalent work. DeepSeek’s January 20, 2025 launch announcement listed the following rates for its R1 API, identified as deepseek-reasoner:
Rank #3
| Token type | DeepSeek-R1 launch rate (January 20, 2025) | OpenAI o1 listed rate | Rate comparison |
|---|---|---|---|
| Input, cache hit | $0.14 per million tokens | $15 per million input tokens | About 99% lower than the listed o1 input rate. |
| Input, cache miss / uncached | $0.55 per million tokens | $15 per million input tokens | About 96% lower than the listed o1 input rate. |
| Output | $2.19 per million tokens | $60 per million output tokens | About 96% lower than the listed o1 output rate. |
The OpenAI o1 rates are from its model documentation accessed in 2026, which labels o1 deprecated; they are not a current-versus-current price guarantee. DeepSeek’s figures are launch-era rates, not verified current rates. The familiar “95% less” shorthand is a rough description of the uncached-input and output rate comparisons at those published prices.
A lower token price does not establish a lower bill for an equivalent result. Two models can use different numbers of input and output tokens, produce different lengths of reasoning, or need different numbers of attempts. OpenAI’s token guidance advises assessing total token cost on representative tasks. Neither the price arithmetic nor the cited papers establish a 95% reduction in end-to-end task cost or in the cost of training the model.
Rank #4
Can you run DeepSeek-R1 locally?
Yes, but the practical answer depends heavily on which checkpoint you choose. DeepSeek lists the full R1 and R1-Zero models at 671 billion total parameters and 37 billion activated parameters, with a 128K context length. Its repository also offers six distilled checkpoints—1.5B, 7B, 8B, 14B, 32B, and 70B—based on Qwen and Llama families.
The smaller distilled models are the more plausible starting point for many local setups. The official model card documents serving paths using Transformers, vLLM, SGLang, and Docker-related instructions; one SGLang example requests all GPUs. The documentation does not establish a single minimum or recommended GPU for every checkpoint. Hardware requirements vary with model size, quantization, context length, and inference software, so the full 671B model should not be treated as a typical consumer-GPU workload.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallBest Value
License and deployment scope
DeepSeek’s January 2025 release announcement describes the code and models as MIT-licensed and says they may be commercialized. That statement is not a blanket legal conclusion about every downstream dependency, deployment, or use case; review the applicable license materials and obligations for the specific components you use.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

