What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
OpenAI reinforcement fine-tuning (RFT) trains a reasoning model to earn higher scores from a grader you define. The process repeatedly generates candidate answers, grades them, and updates the model to favor higher-scoring responses. But access is now limited: OpenAI says its fine-tuning platform is being wound down, new users cannot access it, and existing users can create jobs only for the coming months. The reviewed documentation does not give a precise final job-creation date.
What is OpenAI reinforcement fine-tuning?
RFT is a developer workflow for adapting an OpenAI reasoning model to a particular task using a programmable reward signal. OpenAI describes it this way: “Reinforcement fine-tuning (RFT) adapts an OpenAI reasoning model with a feedback signal you define.” (OpenAI Developers’ RFT guide.)
Unlike supervised fine-tuning, which trains against supplied target answers, RFT scores generated responses. The grader defines what counts as success: accuracy, required formatting, style, or other criteria. If the grader measures the wrong thing—or can be fooled—the model can learn to optimize the score without improving in the way you intended.
How does RFT work?
- Provide prompts and context. Training examples are stored as JSONL rows with a
messagesarray and any additional context the grader needs. - Generate candidate responses. For each prompt, the platform samples multiple responses from the model.
- Grade the responses. A grader assigns scores according to the rules you set.
- Update the model. Policy-gradient updates favor responses that earned higher reward, and the sampling and grading cycle repeats.
- Evaluate and deploy. Review validation results and checkpoints, adjust the data or grader if needed, then use the resulting model through the standard API.
The loop makes the grader central to the training outcome. A high reward is evidence that the model performs well according to that grader, not proof by itself that it will perform well in real use.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minute#1 Best Overall
When should you use RFT?
RFT is best suited to tasks with clear instructions and answers that qualified experts can assess consistently. OpenAI recommends checking model performance before training: a baseline already at the evaluation’s minimum or maximum leaves little useful reward headroom. The model must also have some existing success on the task; OpenAI says RFT cannot bootstrap a model from a 0% success rate. (RFT guide; RFT use cases.)
Examples in OpenAI’s use-case guide include turning instructions into code, configurations, or templates that pass deterministic tests; extracting verifiable facts into structured outputs; and applying complex rules to nuanced, hierarchical, or high-stakes information. In each case, quality needs to be checkable against tests or a rubric.
Rank #2
- Verifiability: Can a grader reliably distinguish a correct answer from an incorrect one?
- Expert agreement: Would independent, qualified reviewers reach similar judgments using the same information?
- Baseline room: Does the current model have room to improve without starting at zero?
- Resistance to shortcuts: Could a model guess or exploit a grading loophole to score well?
- Practical access and cost: Can your account still start jobs, and does the likely benefit justify training and grader expenses?
How much training data does RFT need?
OpenAI recommends beginning with several dozen to a few hundred examples to learn whether RFT is useful before committing to a larger dataset. Its 2026 documentation sets a maximum of 50,000 training examples and 1,000 test examples. Those are platform limits, not a recommended dataset size or a guarantee of improvement. The guide emphasizes data quality; adding examples is useful only if quality is maintained. (RFT guide.)
For tool-calling tasks, include the tools on each training data point and grade the tool calls themselves. Structured-output training requires the applicable JSON schema. OpenAI’s guide also says a paused job can be resumed from its latest checkpoint.
How do RFT graders and evaluations work?
OpenAI documents several grader types. String checks suit exact or simple conditions; text-similarity measures compare wording or meaning; score-model graders use another model to evaluate open-ended responses; and Python code graders can apply custom logic. A multigrader can combine component scores—for example, checking a required schema field deterministically while using a model grader to score an explanation. (OpenAI graders guide; RFT guide.)
A model grader can judge nuanced answers, but it adds cost and creates another opportunity for reward hacking. OpenAI warns that a training model may exploit weaknesses in a grader and recommends comparing grader results with expert human evaluation. Before training, test the grader against known-good answers, known-bad answers, and edge cases. Investigate grader errors: OpenAI says they can stem from unsupported outputs, execution or system issues, or bugs in the grading logic.
What is the RFT workflow?
- Define and test the grader. Make sure its score reflects the actual task, including edge cases and possible shortcuts.
- Prepare JSONL data. Separate training examples from validation or test examples, and provide the context needed for grading.
- Upload the files. The job uses uploaded training and test file IDs.
- Start a fine-tuning job. Select the supported base model and supply the grader and file IDs.
- Monitor metrics, errors, and checkpoints. Use the results to identify weaknesses in the model, data, or grading logic. Revise and rerun as appropriate.
- Deploy the resulting model. Call it through the standard API, subject to the model and platform’s continued availability.
OpenAI’s guide describes dataset screening, so a file within the documented row limits is not necessarily guaranteed to be accepted. For current setup details and API parameters, use the RFT guide.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How much does OpenAI RFT cost?
OpenAI’s Help Center lists core training-loop compute for o4-mini-2025-04-16 at $100 per hour. The listed rate is for core training compute, not an all-in project estimate: model-grader token usage is billed separately at standard API rates. Billable core work includes generating samples, grading, weight updates, and configured validation; queue waiting, dataset validation and preparation, and safety checks are excluded from compute billing. Rates can change, so check the current RFT billing article before budgeting.
Best Value
Is OpenAI RFT still available?
OpenAI’s current RFT guide and use-case page say the fine-tuning platform is being wound down. New users cannot access it; existing platform users may create jobs for the coming months. Fine-tuned models remain available for inference until their base models are deprecated. The reviewed pages do not establish the exact final date for creating jobs.
If you already have access, check OpenAI’s deprecation timeline and confirm your organization’s account-level access before planning a project. Do not assume that a model currently listed in the guide will remain available indefinitely.
Which models does RFT support?
The currently reviewed guide says RFT supports o-series reasoning models only and specifically lists o4-mini; the billing article names the dated model o4-mini-2025-04-16. These listings describe current documentation, not a promise of ongoing access, particularly while the platform is winding down. Check the RFT guide and your account before designing around a particular base model.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Recommended Free Tools

