What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
OpenAI’s Moderation API classifies text and image content and returns category-specific signals your application can use to allow, block, or review content. It does not enforce your policy, prevent a model from generating a response, or replace other safety controls. Build a decision process around its results, account for gaps in modality coverage, and keep a separate child-safety response plan.
What the Moderation API does
The Moderation API exposes a POST /moderations endpoint. A request can contain one string, an array of strings, or multimodal input objects containing text and/or image content. The API reference lists omni-moderation-latest as the default model.
The response identifies the moderation model and includes one or more result objects. Each result provides an overall flagged value, per-category flags, category scores, and metadata identifying the input types covered by each category score.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallflaggedis a useful first-pass indicator that at least one category was flagged.categoriesgives a true-or-false result for each category.category_scorescontains values from 0 to 1; higher values indicate greater model confidence that content belongs to that category.category_applied_input_typesshows which input modalities the category score applies to.
These are model outputs for your application to interpret, not a complete product policy. OpenAI’s documentation does not establish a universal score threshold or publish an authoritative accuracy or error-rate figure for these classifications.
#1 Best Overall
Choose how to put moderation in the request flow
There are two practical approaches. Use standalone classification when you need to screen content independently; request moderation alongside generation when you need signals for the model input and its response. In either case, inspect the result before showing content or taking a downstream action.
| Approach | Useful when | Where results appear | Important implementation point |
|---|---|---|---|
Standalone POST /moderations |
You need to classify arbitrary submitted text or supported image content independently of a generation request. | In the moderation response, which includes the model identifier and result object or objects. | Use the results to apply your own allow, block, review, or escalation policy. |
| Moderation alongside a generated response | You need moderation signals for model input and generated output in a Responses API or Chat Completions workflow. | Alongside the input and output in the request flow. | The model still generates normally. Review the moderation results before showing the response or acting on it. |
In streaming generation, moderation scores arrive when the full generated output is available, not with partial output deltas. Do not display unreviewed streamed content on the assumption that an inline moderation field will stop it.
Rank #2
Check category and modality coverage
The current Moderation guide describes omni-moderation-latest as accepting text and images, not audio. Documented categories include harassment and threatening harassment, hate and threatening hate, illicit activity and violent illicit activity, self-harm and self-harm intent or instructions, sexual content and sexual content involving minors, violence, and graphic violence. Some categories are text-only; consult the live guide for the current mapping of categories to input types.
An image-only request receives a zero score for categories that do not support images. That zero does not mean the image was assessed for that category. Use category_applied_input_types to distinguish unsupported coverage from a meaningful classification result. The guide documents a maximum image file size of 20 MB; confirm current requirements before relying on that limit in an upload flow.
Rank #3
Turn the signal into an application policy
Decide what your product should do before wiring a score into a user-facing action. The same flagged category can carry different consequences in a low-stakes discussion forum and a high-impact service. A practical policy has explicit routes for content that can proceed, content that must be blocked, and cases that need human review or escalation.
- Define outcomes. Specify which categories and contexts lead to allow, block, human review, or escalation. Identify the cost of a false positive and a missed case for each route.
- Use
flaggedas the initial signal. Inspect category flags and scores when your policy needs category-specific handling; do not treat a score as a universal probability or ready-made threshold. - Check modality metadata. Confirm that a category was applied to the actual input type. A zero score for an unsupported modality is not evidence that the content is safe.
- Handle unavailable or failed moderation. Check for moderation errors before reading scores and define fail-safe behavior for your use case. For example, a high-impact workflow may need to pause or route content for review rather than proceed without a result.
- Preserve useful context for review. Give reviewers enough surrounding information to make a decision, and provide a clear path for ambiguous or high-impact cases.
- Revisit score-dependent rules after model changes. OpenAI notes that upgrades can change score behavior, so policies that rely on custom score cutoffs may need recalibration.
When conversation content includes tool calls, inspect their arguments and outputs as part of your safety design. OpenAI’s guide notes that tool names, descriptions, schemas, and response-format schemas are not covered as conversation content.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Use moderation alongside other safeguards
Classification is one layer in a broader safety process. OpenAI’s Safety best practices recommend adversarial testing, prompt engineering, suitable limits on user inputs and generated outputs, and human review wherever possible. The guide states: “Wherever possible, we recommend having a human review outputs before they are used in practice.” Human review is especially important in high-stakes domains.
Free tools Windows power users keep installed
One-click scans. No signup required.
Test ordinary traffic as well as adversarial cases, including attempts to redirect a model through prompt injection. Check the complete path from user input through model output and any downstream tool action; a moderation score alone does not test whether the surrounding application handles an unsafe or ambiguous case correctly.
Know the child-safety boundary
OpenAI says the Moderation API is not designed for CSAM detection or handling and is not a substitute for dedicated child-safety safeguards. Do not send known or suspected child sexual abuse material to the API. Your product’s child-safety controls and incident response must not depend on this classifier.
Understand API data handling
OpenAI’s API data controls documentation says abuse-monitoring logs may contain customer content, including prompts and responses, and derived metadata such as classifier outputs. By default, those logs are retained for up to 30 days unless a longer period is legally required.
Eligible customers may apply for Modified Abuse Monitoring or Zero Data Retention, subject to prior approval and additional requirements. Do not assume either control applies to your account; verify your eligibility and the behavior that applies to the endpoints you use.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

