Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

In a 2024 blind study at the University of Reading, 94% of GPT-4-written assessment submissions were not detected, and the submissions earned higher marks on average than real student work. The result applies to five undergraduate psychology modules, the university’s marking process, and the assessment questions tested—not to every professor, subject, university, AI model, or form of student use.

What the University of Reading study actually tested

Researchers in the University of Reading’s School of Psychology and Clinical Language Sciences created 33 student identities and submitted fully GPT-4-generated answers through the university’s examination system. The academic markers did not know an experiment was taking place.

The work covered five undergraduate psychology modules across different years of the degree. It used real assessment questions and two formats:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Short answers: students selected four questions from six and answered each in no more than 200 words.
  • Essays: students submitted one essay of approximately 1,500 words.

This was a real-world blind test of one institution’s assessment process. It was not a controlled survey of professors from multiple universities or disciplines.

Did ChatGPT get better grades than students?

On average, the AI submissions scored about half a grade boundary higher than the real student submissions used for comparison. The authors also reported an 83.4% chance across modules that the AI submissions would outperform a random selection of the same number of real student submissions.

That 83.4% figure is an across-module comparison. It does not mean that 83.4% of individual AI answers beat every student, nor that every AI answer received a high mark.

The reported outcomes

Measure What the study reported How to interpret it
Detection 94% of AI submissions were not detected Markers failed to identify most of these GPT-4 submissions in this assessment setting.
Average marks About half a grade boundary higher than real student submissions The AI group performed better on average in the tested modules.
Across-module comparison 83.4% chance of outperforming an equal-sized random student sample This is a probability for the study’s module-level comparisons, not an individual-answer success rate.

The paper’s abstract states: “Overall, we found that 94% of AI submissions verged on being undetectable, even though we used AI in the most detectable way possible.” That sentence describes the submissions and markers in this experiment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can professors tell if an answer was written by AI?

This study shows that the unaware markers in these five Reading modules usually could not identify the tested GPT-4 submissions from the submitted work alone. It does not establish a universal ability—or inability—for professors to recognize AI writing.

Detection can change with the model, the student’s prompts and editing, the subject, the assignment, the marking rubric, and whether an instructor can inspect drafts, notes, sources, or a student’s explanation of the work. The experiment used entirely AI-written answers; it did not test a student who used AI for brainstorming, translation, revision, or partial drafting.

How reliable are AI detectors in university exams?

The researchers did not evaluate commercial AI-detection products, compare detector accuracy, or test an automated flagging system. The 94% figure is a marker-detection result: it records whether the human assessors identified the submissions as AI-generated within the university’s normal process.

It therefore cannot be used to claim that a particular detector is accurate or inaccurate. Nor does it provide a safe threshold for accusing a student. A detector score, if an institution uses one, would be a separate piece of evidence requiring the institution’s own validation, policy, and due-process safeguards.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why the findings matter for assessment design

The experiment exposes a vulnerability in assessments that can be completed as a polished, unsupervised answer and judged mainly from the final text. If the final product is the only evidence of learning, a fluent model can sometimes satisfy the marking criteria more consistently than hurried student work.

Assessment practices that provide more evidence of learning

  • Require staged proposals, outlines, drafts, revisions, or research logs.
  • Ask students to explain choices, sources, calculations, or interpretations specific to their own work.
  • Use brief oral follow-ups, presentations, or in-class writing where appropriate.
  • Connect prompts to course discussions, local data, laboratory work, or a student’s documented process.
  • Publish clear rules stating when generative AI is prohibited, permitted, or required and how its use must be acknowledged.

These are assessment-design implications, not interventions tested by the Reading experiment. The study measured what happened under the existing assessment setup.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the University of Reading said afterward

The University of Reading said the project informed its work on AI in research, teaching, learning, and assessment, and that it issued updated advice for staff and students.

Elizabeth McCrum, the university’s Pro-Vice-Chancellor for Education and Student Experience, said: “It is clear that AI will have a transformative effect in many aspects of our lives, including how we teach students and assess their learning.” This is institutional commentary about the implications, not a measured result from the experiment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What this study does—and does not—prove

It supports

  • GPT-4 could produce submissions that were difficult for the tested markers to distinguish from student work.
  • Those submissions received higher average marks in the five modules studied.
  • Final-answer-only assessment can be vulnerable when students submit work without showing their process.

It does not support

  • A claim that all professors will miss AI writing 94% of the time.
  • A claim that AI will earn higher marks in every discipline, university, assignment, or model generation.
  • A general “Turing Test” conclusion about human ability to identify AI.
  • A ranking or validation of commercial AI detectors.
  • A conclusion about students who combine AI output with substantial human writing or editing.

The practical takeaway for students and educators

Students should follow their institution’s academic-integrity and generative-AI rules; an answer that passes unnoticed is not automatically permitted. Educators should treat the result as a warning that polished text alone may be weak evidence of authorship and understanding, then choose assessment methods that make the learning process visible.

The headline is therefore accurate only when read in its proper frame: in one 2024 University of Reading study, GPT-4-generated answers evaded detection in 94% of cases and outscored real student submissions on average. It is evidence about that tested system, not a universal scorecard for professors or AI.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.