Are ai detectors accurate enough to accuse somebody?
Asked plainly — are ai detectors accurate — the honest answer is that nobody outside the companies selling them can fully say, and that the most rigorous public answer came from a university that used one, tested it, and then turned it off.
The single most useful calculation on this subject is not in a study. It is in a blog post from Vanderbilt University explaining why it disabled Turnitin’s AI detector.
The claimTurnitin, at launch, claimed a 1% false positive rate.The volumeVanderbilt submitted 75,000 papers to Turnitin in one year.The result“Around 750 student papers could have been incorrectly
labeled as having some of it written by AI.”
That is Vanderbilt University’s own arithmetic, published when it disabled the tool. It is the
most useful thing we have found written about detector accuracy, because it does the one operation
nobody else performs: it turns a percentage into people. One per cent sounds like rounding
error until it is seven hundred and fifty students at a single university in a single year.
That number is worth holding onto whichever side of this you are on. If you are a student who has been accused, it is the reason a percentage is not evidence. If you are an instructor deciding whether to act on a score, it is the size of the population you would be wrong about.
Vanderbilt gave three further reasons, and they are separate from the accuracy figure — a tool can be accurate and still be unusable for these.
Nobody outside the company knows how it works
Vanderbilt: “To date, Turnitin gives no detailed information as to how it determines if a
piece of writing is AI-generated or not. The most they have said is that their tool looks for
patterns common in AI writing, but they do not explain or define” those patterns. An accusation
you cannot examine is an accusation you cannot answer.
It flags some writers more than others
Vanderbilt again, citing the research: “AI detectors have been found to be more likely to
label text written by non-native English speakers as AI-written.” That is not a rounding
error distributed evenly. It falls on particular students, repeatedly.
Institutions did not choose to switch it on
The tool “was enabled for Turnitin customers with less than 24-hour advance notice, no option
at the time to disable the feature, and, most importantly, no insight into how it works.”
Universities were scoring students with it before they had decided to.
The first of those is the one that matters most in a disciplinary meeting. In every other kind of academic accusation, the evidence can be inspected: a source can be compared, a similarity report can be read line by line. An AI detection score cannot be examined by anybody — not the student, not the instructor, not the institution — because the method is not published. The number has to be believed or disbelieved.
The second is why the accuracy question and the fairness question are not the same question. A detector that is wrong one per cent of the time at random is a different object from one that is wrong mostly about students who learned English as a second language, and only the first is described by the headline figure.
One honest complication, because the tidy version of this story is not true. Not every
university that dropped the tool dropped it over accuracy.
San Francisco State announced its own discontinuation for an entirely different reason: Turnitin
had provided AI detection free during a beta period, then “notified campuses that the AI detection
feature would be add-on for additional cost and that the free trial period would end”. The feature
went away because the contract changed. If you are arguing that the sector has rejected these tools
on the evidence, that is one of the cases that does not support you — and it is worth knowing before
somebody else points it out.
The neighbouring question — how do ai detectors work — has an answer that is short and unsatisfying, and it is the same one Vanderbilt gave: outside the companies, nobody knows. That is also why asking how accurate are ai detectors cannot be settled from outside. Accuracy is measured against a method, and the method is not published.
So what is a defensible position for somebody who has to decide something today?
If you are teaching: a detection score is a reason to look, never a reason to conclude. The conversation that follows — about drafts, sources, the argument in the essay — is the actual assessment, and it was always going to have to happen anyway.
If you have been accused: ask what the score is based on, and ask it in writing. The answer, if it is honest, is that the method is not disclosed. That is not a rhetorical trick; it is the documented state of the tool, published by a university that paid for it.
And if you are writing anything you might one day need to defend, the single most valuable habit costs nothing and no detector can argue with it: work somewhere that keeps a version history. A document with three weeks of edits behind it is evidence of a kind that a probability score cannot touch.
The floors below take it four ways: what a score actually measures, the false positives and who gets them, the individual tools and what each claims, and what to do when an accusation has already been made.
It will not tell you a detector is reliable enough to accuse somebody. The most detailed public reasoning on that question was published by a university, and it reached the opposite conclusion.
It will not tell you the sector has settled this. One of the universities that dropped the tool did so because the contract changed, and pretending otherwise would be the same overreach we are objecting to.
And it will not carry an affiliate link. In this tower that means declining some of the highest commissions on the island. What holds instead is simple: the accuracy figure, the arithmetic and the reasons for disabling the tool are quoted from the universities that published them, listed below.
Where this page got its facts
Vanderbilt University, Brightspace — Guidance on AI Detection and Why We’re Disabling Turnitin’s AI Detector (the 1% false positive claim, the calculation against 75,000 submitted papers, the absence of published method, the bias against non-native English speakers, and the rollout with less than 24 hours’ notice). Published 16 August 2023 — www.vanderbilt.edu, read 22 August 2026.
San Francisco State University, Academic Technology — Discontinuation of TurnItIn AI detection tool availability (that the feature was free during a beta phase and became a paid add-on, which is why it was withdrawn there). Published 1 May 2024 — at.sfsu.edu, read 22 August 2026.
AG
Written by Alberto Gulotta
Founder and editor of AI Tools Primer, writing from Palermo, Italy. Thirty-five years of
taking computers apart, starting with a Commodore 64 — the long version is on the
about page.
No affiliate links and no paid placements anywhere on this site. Nobody pays to appear here,
and no company has seen this page before you did.
This is general information, not professional advice. Where a page touches money, health,
safety or the law, it names its source and the date it was read — and your situation may
still differ. See the privacy page and the
cookie policy.