Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

GRPO removes the learned value function, or critic, that PPO normally uses to estimate a baseline. It does not remove reward scoring. GRPO still scores each sampled completion and compares those scores within a group, so the reward source is still required. What changes is how the baseline is produced. The reward can come from a learned reward model, a rule-based function, or another task-specific scorer, depending on the implementation.

Why the two components are easy to confuse

In reinforcement learning for language models, two separate jobs have to be done for every training update. One is scoring: deciding how good a generated output is. The other is baselining: deciding whether that output was better or worse than what the model would normally expect for that prompt. PPO handles the baselining job with a learned critic, a value model trained to predict expected reward. The reward signal itself is supplied separately.

GRPO, introduced as a variant of PPO in the DeepSeekMath paper (arXiv:2402.03300), replaces the learned critic with statistics computed from a group of outputs. Scoring stays. The headline is accurate about the critic and misleading if read as a claim about rewards.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What GRPO does in one prompt

For each prompt, GRPO runs the following steps:

  1. Sample a group of completions from the current policy for that same prompt.
  2. Score every completion, producing one reward per completion. The scorer is whatever reward mechanism the training setup defines.
  3. Compute each completion’s advantage relative to its group. In the default normalization documented in the GRPO Trainer guide, the advantage is the completion’s reward minus the group mean, divided by the group standard deviation.
  4. Use those advantages in a PPO-style policy update, with a KL term that penalizes divergence from a reference policy.

The group mean and standard deviation take the place the critic would otherwise fill. Without them, the update would have no way to say whether a completion was better than typical for that prompt. The reward values are still the raw material being compared.

Where the reward comes from

The GRPO Trainer documentation describes reward computation per completion and supports more than one kind of scorer. The reward source is a design decision, not a fixed part of the method.

Learned reward models

A reward model is a separately trained network that takes a prompt and a completion and returns a scalar score. The trainer documentation describes computing rewards this way. Using one adds a model to memory and to the training pipeline, but GRPO does not need a critic alongside it.

Custom reward functions

The same documentation also covers custom reward functions. These are code-defined scorers, such as checking whether a math answer matches a reference, whether code passes tests, or whether output follows a format. Here no learned reward model is involved. This is why it is wrong to say that every GRPO system uses a reward model, and equally wrong to say that GRPO avoids one. Both are configuration choices.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What remains in the objective

Critic-free does not mean the training objective is stripped down. In the documented implementation, the policy objective keeps the PPO-style structure, including policy-ratio clipping, and includes a KL term against a reference policy. The reference policy is a separate model copy used as an anchor. It is not a critic, and removing the critic does not remove it.

The exact form of the loss, the reward scaling, and the KL handling depend on the formulation and trainer settings. The documentation discusses alternative loss formulations and configurable reward scaling, so the normalization described above should be read as the default, not a universal formula.

PPO and GRPO compared

Axis PPO (as commonly implemented) GRPO
Baseline source Learned value function (critic) Group-relative statistics from multiple completions for the same prompt
Reward source Set by the training setup; separate from the critic Set by the training setup: learned reward model, custom reward function, or other scorer
Extra model for baseline Yes, a critic is trained alongside the policy No critic; the group provides the comparison
Sampling pattern Completions are sampled and scored per the training loop Multiple completions are sampled per prompt and compared within that group
Reference-policy KL Commonly used Kept in the documented implementation

The memory benefit is the main reason the original paper gives for GRPO. Group sampling has its own cost: each prompt needs several generated completions before an update can be computed.

What the reported numbers do and do not show

The DeepSeekMath paper reports 51.7% on the MATH benchmark for DeepSeekMath 7B, and 60.9% using self-consistency over 64 samples. These are results from that paper, and they reflect several contributions, including its data selection pipeline and GRPO. The 120B figure in the abstract refers to the scale of math-related data used for continued pretraining, not to a GRPO rollout count. None of these figures isolate the effect of removing the critic.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The paper’s abstract describes GRPO this way: “we introduce Group Relative Policy Optimization (GRPO), a variant of Proximal Policy Optimization (PPO), that enhances mathematical reasoning abilities while concurrently optimizing the memory usage of PPO.” The abstract does not discuss reward models. For that distinction, the trainer documentation is the more relevant source.

Best Value
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Claims to avoid

  • “GRPO removes the reward model.” Reward computation is still part of the loop.
  • “Every GRPO system uses a learned reward model.” Custom reward functions are a documented option.
  • “Critic-free means no other model is involved.” A reference policy is still used for the KL term when that term is configured, and a reward model may be present.
  • A fixed reward normalization or loss formula presented as universal.

The DeepSeek-R1 paper (arXiv:2501.12948) is a related record, but the material available for this article does not establish its reward setup for each training stage, so this article makes no claim about it.

Sources

“

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.