Students are being warned not to outsource their judgment to AI.
Universities may need to take the same advice. Higher education has a real problem. Generative AI can produce essays, rewrite arguments, fabricate references, paraphrase source material and make work that a student barely understands look surprisingly competent.
Universities have every right to protect academic integrity. The question is whether one imperfect technology should be policed by another. Because a detection score can look far more certain than the evidence underneath it.
The 61.3% result deserves both attention and precision
In 2023, Stanford researchers Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu and James Zou published a study in Patterns examining bias in GPT detectors.
They tested seven widely used detectors against 91 human-written TOEFL essays by non-native English writers and 88 essays written by US eighth-grade students.
The result were striking. The detectors performed close to perfectly on the US student essays. But across the TOEFL sample, the average false-positive rate was 61.3%. All 7 detectors unanimously classified 19.8% of those human-written TOEFL essays as AI-generated. At least one detector flagged 97.8% of the TOEFL essays.
Those numbers are alarming.
They do not mean that every AI detector today is “61% inaccurate.” They do not tell us the error rate of every contemporary commercial system. They describe the performance of 7 detectors, tested on a specific dataset, at a particular point in the development of these technologies.
But the study revealed something more important than one accuracy number.
The errors were patterned.
Why non-native writing can look machine-generated
The researchers linked much of the problem to text perplexity: broadly, how predictable a piece of writing appears to a language model. Many AI-detection approaches have used predictability as one signal.
Machine-generated prose can be statistically predictable. But so can legitimate writing by someone working in a second language. Non-native writers may use a smaller active vocabulary, more familiar sentence structures, less idiomatic language, lower lexical variation and none of those fully means the text was generated by AI. Yet they can make the writing look statistically more predictable.
The Patterns paper found that when researchers used ChatGPT to diversify the word choice in the TOEFL essays, the false-positive rate fell sharply, from 61.3% to 11.6%. Conversely, simplifying vocabulary in the US student essays increased misclassification - And that's a deeply uncomfortable finding.
A student’s writing can appear less “AI-like” simply by making it sound more linguistically sophisticated. In other words, the detector can risk confusing language proficiency with authorship.
The detector does not accuse the student
Software does not conduct a disciplinary process. It produces an output - A percentage, a classification, a warning.
Everything after that is institutional judgment. And recent cases in the UK show what can happen when that distinction becomes blurred.
In one published case, the UK Office of the Independent Adjudicator described an international student whose coursework had been flagged by Turnitin.
The student was called to a viva. The institution concluded that the student had not demonstrated adequate understanding of the work and had not shown that they had written it themselves. The student received a zero and was required to resubmit.
The student complained. The OIA upheld the complaint as Justified, concluding that the student had not been given a fair opportunity to respond to the allegations. That case does not establish that the student definitely did or did not use AI. But what it establishes is equally important: Fair process still matters when software raises the suspicion.
Another case shows why international students require particular care
In a second case, the OIA examined an international student whose assignment had received a high AI-generated-content indication from Turnitin.
The student was called to a viva and later faced an academic-misconduct panel. The OIA ultimately classified the complaint as Partly Justified.
The case later became part of a wider UK debate about detector fairness. HEPI’s July 2026 analysis of AI detection and international students highlighted the concern that institutions recruiting globally need to understand whether systems may behave differently on second-language writing.
The institutional question worth focusing on isn't: **“Can detectors ever be useful?” (**They may be.)
But: “How much authority should a probabilistic signal carry?”
Institutions like quantification because numbers appear objective. A score feels more concrete than suspicion. But the appearance of precision is not the same as certainty - Somebody selected the detection model, decided which linguistic features matter, chose how the score is presented, defined whether 20%, 50% or 80% is concerning and also decided whether that concern triggers a conversation, a viva or a misconduct proceeding.
The human judgment did not disappear. It moved upstream.
That is one of the central risks of algorithmic systems inside institutions. We can mistake the output of a model for the absence of human judgment when, in reality, human decisions are embedded throughout the system.
The consequence should determine the evidentiary standard
This is not an argument for ignoring AI-assisted cheating. Universities cannot do that. Academic work needs to remain meaningful. Students who outsource assessed work damage the credibility of their own education and everyone else’s. But the existence of misconduct does not make every detection mechanism equally defensible.
The higher the consequence, the stronger the evidence should have to become.
A streaming service can incorrectly predict what film I want to watch with almost no meaningful consequence, but a university misconducted decision can affect grades, progression, finances, academic records and potentially a student’s wider future.
Those systems should not operate with the same tolerance for false positives. A detector score may reasonably trigger scrutiny.
There's another route. Instead of trying to infer the entire production process from the statistical texture of the final essay, institutions can collect better evidence of the process itself.
A student who actually understands their work can usually talk about its evolution. Academic integrity is not ultimately about identifying which sentences look machine-generated.
It is about establishing whether the assessed intellectual work belongs meaningfully to the person receiving credit for it.
Universities across the UK, US, Australia and elsewhere actively recruit multilingual students. Those students are asked to cross more than geographical borders, crossing linguistic ones, learning unfamiliar conventions.
Institutions therefore carry a particular responsibility not to confuse linguistic difference with misconduct. The Stanford-led findings should not be turned into a permanent universal benchmark for all detectors. But they should remain a warning against blind confidence.
The real integrity question now runs both ways
Students should not outsource their authorship to AI - That standard is reasonable.
But institutions should not outsource their judgment to AI either.
Perhaps the most useful principle for the AI-detection era is therefore a simple one: A probability can start a conversation. It should never be allowed, by itself, to finish one.