iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
GRPO removes the learned value function, or critic, that PPO normally uses to estimate a baseline. It does not remove reward scoring. GRPO still scores each sampled completion and compares those scores within a group, so the reward source is still required. What changes is how the baseline is produced. The reward can come from a learned reward model, a rule-based function, or another task-specific scorer, depending on the implementation.
Why the two components are easy to confuse
In reinforcement learning for language models, two separate jobs have to be done for every training update. One is scoring: deciding how good a generated output is. The other is baselining: deciding whether that output was better or worse than what the model would normally expect for that prompt. PPO handles the baselining job with a learned critic, a value model trained to predict expected reward. The reward signal itself is supplied separately.
GRPO, introduced as a variant of PPO in the DeepSeekMath paper (arXiv:2402.03300), replaces the learned critic with statistics computed from a group of outputs. Scoring stays. The headline is accurate about the critic and misleading if read as a claim about rewards.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →What GRPO does in one prompt
For each prompt, GRPO runs the following steps:
- Sample a group of completions from the current policy for that same prompt.
- Score every completion, producing one reward per completion. The scorer is whatever reward mechanism the training setup defines.
- Compute each completion’s advantage relative to its group. In the default normalization documented in the GRPO Trainer guide, the advantage is the completion’s reward minus the group mean, divided by the group standard deviation.
- Use those advantages in a PPO-style policy update, with a KL term that penalizes divergence from a reference policy.
The group mean and standard deviation take the place the critic would otherwise fill. Without them, the update would have no way to say whether a completion was better than typical for that prompt. The reward values are still the raw material being compared.
#1 Best Overall
Where the reward comes from
The GRPO Trainer documentation describes reward computation per completion and supports more than one kind of scorer. The reward source is a design decision, not a fixed part of the method.
Learned reward models
A reward model is a separately trained network that takes a prompt and a completion and returns a scalar score. The trainer documentation describes computing rewards this way. Using one adds a model to memory and to the training pipeline, but GRPO does not need a critic alongside it.
Rank #2
Custom reward functions
The same documentation also covers custom reward functions. These are code-defined scorers, such as checking whether a math answer matches a reference, whether code passes tests, or whether output follows a format. Here no learned reward model is involved. This is why it is wrong to say that every GRPO system uses a reward model, and equally wrong to say that GRPO avoids one. Both are configuration choices.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallWhat remains in the objective
Critic-free does not mean the training objective is stripped down. In the documented implementation, the policy objective keeps the PPO-style structure, including policy-ratio clipping, and includes a KL term against a reference policy. The reference policy is a separate model copy used as an anchor. It is not a critic, and removing the critic does not remove it.
The exact form of the loss, the reward scaling, and the KL handling depend on the formulation and trainer settings. The documentation discusses alternative loss formulations and configurable reward scaling, so the normalization described above should be read as the default, not a universal formula.
PPO and GRPO compared
| Axis | PPO (as commonly implemented) | GRPO |
|---|---|---|
| Baseline source | Learned value function (critic) | Group-relative statistics from multiple completions for the same prompt |
| Reward source | Set by the training setup; separate from the critic | Set by the training setup: learned reward model, custom reward function, or other scorer |
| Extra model for baseline | Yes, a critic is trained alongside the policy | No critic; the group provides the comparison |
| Sampling pattern | Completions are sampled and scored per the training loop | Multiple completions are sampled per prompt and compared within that group |
| Reference-policy KL | Commonly used | Kept in the documented implementation |
The memory benefit is the main reason the original paper gives for GRPO. Group sampling has its own cost: each prompt needs several generated completions before an update can be computed.
Rank #4
What the reported numbers do and do not show
The DeepSeekMath paper reports 51.7% on the MATH benchmark for DeepSeekMath 7B, and 60.9% using self-consistency over 64 samples. These are results from that paper, and they reflect several contributions, including its data selection pipeline and GRPO. The 120B figure in the abstract refers to the scale of math-related data used for continued pretraining, not to a GRPO rollout count. None of these figures isolate the effect of removing the critic.
Recommended Free Tools
The paper’s abstract describes GRPO this way: “we introduce Group Relative Policy Optimization (GRPO), a variant of Proximal Policy Optimization (PPO), that enhances mathematical reasoning abilities while concurrently optimizing the memory usage of PPO.” The abstract does not discuss reward models. For that distinction, the trainer documentation is the more relevant source.
Best Value
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
Claims to avoid
- “GRPO removes the reward model.” Reward computation is still part of the loop.
- “Every GRPO system uses a learned reward model.” Custom reward functions are a documented option.
- “Critic-free means no other model is involved.” A reference policy is still used for the KL term when that term is configured, and a reward model may be present.
- A fixed reward normalization or loss formula presented as universal.
The DeepSeek-R1 paper (arXiv:2501.12948) is a related record, but the material available for this article does not establish its reward setup for each training stage, so this article makes no claim about it.
Quick Recap
Sources
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models, DeepSeekMath authors, 2024: origin of GRPO, framing, and reported MATH results.
- GRPO Trainer documentation: reward computation, group advantages, custom rewards, objective variants, and hardware notes.
“
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

