Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallUpside-Down Reinforcement Learning (UDRL) reverses what the learner is asked to predict: instead of using a reward or value estimate to choose what to do, it takes a desired return and time horizon as inputs and learns which action to take in the current state. The approach, introduced by Jürgen Schmidhuber, reframes reinforcement learning as supervised learning over collected experience; it does not remove the need to interact with an environment or choose useful goals.
What is upside-down reinforcement learning?
In a conventional reward-centric description of reinforcement learning, an agent learns about rewards or values and uses that information to guide action selection. UDRL changes the direction of the mapping. It supplies a desired outcome as a command and trains a behavior function to map the current state plus that command to an action.
Schmidhuber’s 2019 paper describes the idea this way: “We transform reinforcement learning (RL) into a form of supervised learning (SL) by turning traditional RL on its head, calling this Upside Down RL (UDRL).” The key change is what the model is conditioned on and predicts—not an escape from learning through experience.
How does UDRL work?
- Collect experience. The agent interacts with the environment and records states, actions, and resulting outcomes. These examples provide the material for learning.
- Specify a command. The command can include a desired amount of return and a time horizon for achieving it.
- Condition on the current situation. The behavior function receives the current state together with the command.
- Predict an action. It returns an action, or an action distribution, suited to that state and requested outcome.
- Update the command as time passes. During an episode, the desired remaining return and remaining horizon can be adjusted to reflect progress and time left.
The companion paper, “Training Agents using Upside-Down Reinforcement Learning,” describes the method as “a method for learning to act using only supervised learning techniques.” In practice, supervised learning here still depends on experience generated by interacting with an environment. A behavior function can only learn to follow commands to the extent that its data provide relevant examples.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Does UDRL predict rewards?
Not in the sense that defines its central policy mapping. UDRL feeds a desired return and horizon into the behavior function and learns to predict actions conditioned on them; it does not make reward or value prediction the mechanism for selecting each action. That distinction is the source of the “upside down” name.
The command itself can express desired future outcomes. Schmidhuber’s original abstract also allows command information to include other computable functions of historic and desired future data. This flexibility does not guarantee that a requested outcome is achievable: the policy’s capabilities depend on the environment and the experience available to train it.
Rank #2
How do I specify the reward and time horizon?
Think of the command as a target for behavior, not a promise of what the environment will deliver. A desired return states how much reward the agent is asked to pursue; the horizon indicates how much time or how many steps it has to pursue that target. As the episode advances, both can be revised to represent the desired return still outstanding and the time remaining.
- Choose a meaningful return target. It should make sense for the task and the scale of returns represented in the collected experience.
- Set a horizon that fits the task. A target paired with too little time may be infeasible; a longer horizon changes what actions may be appropriate.
- Account for what the data cover. If experience contains few examples relevant to a command and state, the learned behavior may not know how to satisfy it.
UDRL changes how the learning problem is represented, but it does not solve command design or data-coverage problems automatically.
Recommended Free Tools
Is there a PyTorch implementation?
Yes. Sebastian Dittert’s public GitHub repository documents a PyTorch implementation with discrete- and continuous-action CartPole examples and evaluation notebooks. Its documentation also refers to LunarLander plots. These are useful starting points for exploring the method, but repository documentation alone does not establish independent replication, present-day maintenance, or compatibility with a current software environment.
Does UDRL outperform standard reinforcement learning?
There is no basis for calling it universally superior. The companion paper reports that results on its evaluated episodic tasks were “surprisingly competitive with, and even exceed that of some traditional baseline algorithms.” That is the authors’ qualified, task-specific summary: it concerns some baselines on the tasks they studied, not every environment or comparison.
Performance comparisons depend on more than the learning formulation. Relevant questions include what the model predicts, how desired returns enter its input, how experience is collected and selected, which environments and baselines are tested, and what assumptions apply to the environment. The cited qualitative summary does not establish a general performance percentage or guarantee better sample efficiency.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What do theoretical results establish?
A later theoretical preprint by Miroslav Štrupl and coauthors analyzes convergence and stability of UDRL and related methods. Its abstract reports near-optimal behavior when the environment’s transition kernel is sufficiently close to a deterministic kernel. This is a condition on the environment, not a guarantee for arbitrary settings; the result should be read in light of that assumption.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Quick Recap
Sources and further reading
- Jürgen Schmidhuber, “Reinforcement Learning Upside Down: Don’t Predict Rewards — Just Map Them to Actions” (2019)
- “Training Agents using Upside-Down Reinforcement Learning”
- Sebastian Dittert’s PyTorch UDRL repository
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

