Choose an AI-writing detector by testing it on work that represents your team’s real languages, formats and writing tasks—not by picking the vendor with the biggest accuracy claim. Compare false positives and false negatives, check coverage and workflow, and set a policy that treats a detector result as a prompt for human review, never as proof of authorship or misconduct.
Start with the decision the detector is meant to support
Before comparing products, define the problem you need a detector to help with. A school might want a signal for reviewing a suspected breach of its AI-use policy; an editorial team might want to understand whether submitted prose warrants a conversation about process. Those uses have different rules and consequences. A detector’s score does not establish who wrote a passage or whether a person violated a policy.
Decide in advance what happens when a report raises a concern, who reviews it, and what the writer can say or provide in response. If no fair, useful action follows a flag—or if your organization would treat the score as a verdict—buying a detector will not solve the underlying problem.
Evaluate candidates on representative work
Use the same test set and review procedure for each candidate. This is a practical procurement method based on the documented limits of detection research, not a published universal standard. Include work that reflects the languages, genres, document lengths, and formats your team actually handles, as well as realistic examples of AI-assisted, AI-edited, mixed-origin and fully human writing. Where authorship is documented, have reviewers assess detector output without being told the source label.
- Set the test conditions. Record the test date, product and version, supported languages, document types, text lengths, and any relevant settings. Use the same conditions for each candidate where possible.
- Include the difficult cases. Test short and long pieces, mixed human-and-AI writing, edited or paraphrased text, and the prose or formats that matter to your workflow. Do not assume a result on conventional English prose applies to other languages or formats.
- Review output consistently. Apply the same review protocol to each report. Note whether it identifies passages or only gives an overall score, and have reviewers examine flagged material in its assignment or publication context.
- Count both kinds of error. A false positive is human writing flagged as AI-generated; a false negative is AI-generated writing that is not flagged. Record each separately rather than collapsing them into one accuracy figure.
- Decide whether the trade-off is acceptable. Consider the consequences of each error in your setting. A false accusation can have serious consequences for a student or writer, so do not treat the two error types as interchangeable.
Compare the factors that affect a real deployment
| Factor | What to check | Why it matters |
|---|---|---|
| Error behavior | False positives and false negatives on your representative test set; how the vendor describes uncertainty and thresholds. | A result from your own test set is more relevant to your work than a vendor’s aggregate figure alone. |
| Coverage | Languages, length limits, file types, prose genres, and handling of short, mixed, edited, or paraphrased text. | A tool may not support the material you need to review, or may treat only certain kinds of text as eligible. |
| Workflow | Whether reports identify relevant passages, fit your existing school or editorial process, and support review rather than automatic penalties. | A score is easier to assess responsibly when reviewers can see what it refers to and apply the organization’s normal process. |
| Governance | Who can access reports, how decisions are documented, and how the writer can respond or appeal under your policy. | Human review only helps if the organization has a defined, consistent way to carry it out. |
| Data and procurement | Privacy terms, retention, security, integrations, accessibility, support, contract terms, and total cost. | These are vendor- and contract-specific matters; verify them directly rather than assuming a detector’s score or product guide answers them. |
Do not rank candidates by a single accuracy percentage without its test conditions. CASRAI’s guide, last updated August 24, 2026, notes that accuracy figures combine error types and may reflect conditions unlike real submissions. It reports Turnitin’s own figure of roughly 98% accuracy and a false-positive rate below 1% for documents with more than 20% AI-generated text, while emphasizing that these are internal test results, not an independent peer-reviewed measurement.
What published reliability evidence can—and cannot—tell you
Published evaluations are useful warnings about limits, not current scorecards for every detector. They cover particular tools, samples and conditions. The findings below should not be transferred directly to a different product or to your team’s submissions.
| Evidence | What was evaluated | How to interpret it |
|---|---|---|
| OpenAI classifier, 2023 | OpenAI reported that its classifier identified 26% of AI-written text as “likely AI-written” and incorrectly labelled human-written text 9% of the time on its English challenge set. OpenAI withdrew the classifier on July 20, 2023, citing low accuracy. | This is a historical result for OpenAI’s classifier and test set, not a current score for another product. |
| Weber-Wulff and colleagues, June 21, 2023 | The study evaluated 12 publicly available tools and two commercial systems, Turnitin and PlagiarismCheck. The authors concluded the tools evaluated were neither accurate nor reliable, and that obfuscation worsened performance. | It describes the tools and conditions available at the time of the study, not every service available now. |
| Stanford study, as summarized by CASRAI in 2026 | Across seven GPT detectors, the study found an average false-positive rate of 61.3% on 91 TOEFL essays by non-native English speakers; more than 91% of the essays were flagged by at least one detector. CASRAI contrasts this with a near-zero false-positive rate on a control set of essays by native-English-speaking U.S. eighth-graders and notes that prompt-based rewriting could evade detection. | These findings apply to the study’s samples and detector set. They are not universal rates for current products or for multilingual writers generally. |
OpenAI’s educator guidance, updated in 2026, says its detector research was not reliable enough for consequential judgments, notes that human-written text can be flagged, and warns that small edits can evade detection. Taken together, these sources support local testing and careful review—not treating a confident-looking score as proof.
Check product eligibility before comparing results
Turnitin’s current Using the AI Writing Report guide lists the following requirements for its report. These are Turnitin-specific, may change, and should not be assumed to describe other products.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
| Turnitin report requirement or behavior | Current detail in Turnitin’s guide |
|---|---|
| File size and length | File under 100 MB; at least 300 words and no more than 30,000 words. |
| File formats | DOCX, PDF, TXT or RTF. |
| Supported languages | English, Spanish, Japanese or Arabic. |
| Qualifying text | Prose sentences in long-form writing. The guide says the model does not reliably detect non-prose such as poetry, scripts or code, or reliably cover short-form and unconventional formats such as bullet points, tables or annotated bibliographies. |
| Paraphrasing and bypasser detection | Turnitin says this capability is included only in its English AI detector, not its Spanish or Japanese detectors. The guide excerpt reviewed does not give the same detail for Arabic; confirm Arabic capabilities with Turnitin before relying on them. |
| Scores below 20% | In current reports, results from 0% to below 20% appear as an asterisk, without a percentage or highlights, because Turnitin says false-positive incidence is higher in that range. Reports generated before July 8, 2024 may still show a numerical result below 20%. |
Turnitin describes its AI Writing Report as estimating the portion of qualifying prose it determines could be AI-generated or AI-generated and then modified with an AI paraphraser or bypasser. The AI percentage is separate from the similarity score. Neither the displayed percentage nor a highlighted passage states how much of a person’s work is misconduct.
Set a response process before the first flag
A detector is most defensible as one limited signal within a process that already defines permitted AI use and how concerns are handled. Turnitin’s official guide warns that its model may misidentify human-written, AI-generated and AI-paraphrased text and says it should not be the sole basis for adverse action against a student. Turnitin also says the reviewer—not the report—decides whether misconduct occurred under the relevant institutional policy.
Rank #4
- Publish the rules in advance. Explain which AI uses are permitted or prohibited before work is submitted or commissioned.
- Assign a reviewer. Specify who examines a flagged report, what other evidence may be considered, how the writer can respond, and how the decision is recorded.
- Examine the passage and context. Review any highlighted text alongside the assignment or publication context rather than treating a percentage as a finding.
- Invite an explanation without presuming guilt. Ask the writer how they developed the work. Turnitin frames its report as a starting point for conversation and intervention.
- Use process evidence when policy permits. Drafts, notes, source records or documented AI interactions can inform a discussion of process. OpenAI’s educator guidance suggests students may share conversations to support discussion and AI literacy.
Do not ask a chatbot whether it wrote a passage and use its answer as evidence. OpenAI says ChatGPT has no knowledge of authorship and may give random answers to such questions.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Verify the terms that affect adoption
Published detection evaluations do not establish which product performs best on a particular school’s or editorial team’s representative work. They also do not settle current vendor comparisons for mixed-origin writing, current models, multiple languages or editorial workflows. Before purchase, confirm privacy terms, retention, security, accessibility, integrations, support, contract conditions and total cost directly with each vendor. The evidence available here does not establish comparative current terms for those items.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

