Asked plainly — are ai detectors accurate — the honest answer is that nobody outside the companies selling them can fully say, and that the most rigorous public answer came from a university that used one, tested it, and then turned it off.
The single most useful calculation on this subject is not in a study. It is in a blog post from Vanderbilt University explaining why it disabled Turnitin’s AI detector.
The claimTurnitin, at launch, claimed a 1% false positive rate.The volumeVanderbilt submitted 75,000 papers to Turnitin in one year.The result“Around 750 student papers could have been incorrectly
labeled as having some of it written by AI.”
That is Vanderbilt University’s own arithmetic, published when it disabled the tool. It is the
most useful thing we have found written about detector accuracy, because it does the one operation
nobody else performs: it turns a percentage into people. One per cent sounds like rounding
error until it is seven hundred and fifty students at a single university in a single year.
Vanderbilt’s figure is one vendor’s claim, multiplied. There is also a measurement taken from outside, across several products at once.
The testSix commercial detectors, run against one purpose-built dataset of
articles, abstracts, stories, news and product reviews.The spreadAccuracy rates “between 55.29 and 97.0%”.Who ran itNot one of the six.
That range is the answer to the question in the title, and it is worth reading twice. These
are not six versions of the same number: the worst of them is close to a coin toss. The study
names the tools it tested — “GPTkit,” “GPTZero,”
“Originality,” “Sapling,” “Writer,” and
“Zylalab” — and its own summary is not hostile: “all the tools fared well
in the evaluations”. Both things are true at once, and that is precisely the difficulty.
A number good enough to sort a slush pile is not a number good enough to accuse a
person.
That number is worth holding onto whichever side of this you are on. If you are a student who has been accused, it is the reason a percentage is not evidence. If you are an instructor deciding whether to act on a score, it is the size of the population you would be wrong about.
Vanderbilt gave three further reasons, and they are separate from the accuracy figure — a tool can be accurate and still be unusable for these.
Nobody outside the company knows how it works
Vanderbilt: “To date, Turnitin gives no detailed information as to how it determines if a
piece of writing is AI-generated or not. The most they have said is that their tool looks for
patterns common in AI writing, but they do not explain or define” those patterns. An accusation
you cannot examine is an accusation you cannot answer.
It flags some writers more than others
Vanderbilt again, citing the research: “AI detectors have been found to be more likely to
label text written by non-native English speakers as AI-written.” That is not a rounding
error distributed evenly. It falls on particular students, repeatedly.
Institutions did not choose to switch it on
The tool “was enabled for Turnitin customers with less than 24-hour advance notice, no option
at the time to disable the feature, and, most importantly, no insight into how it works.”
Universities were scoring students with it before they had decided to.
The first of those is the one that matters most in a disciplinary meeting. In every other kind of academic accusation, the evidence can be inspected: a source can be compared, a similarity report can be read line by line. An AI detection score cannot be examined by anybody — not the student, not the instructor, not the institution — because the method is not published. The number has to be believed or disbelieved.
The second is why the accuracy question and the fairness question are not the same question. A detector that is wrong one per cent of the time at random is a different object from one that is wrong mostly about students who learned English as a second language, and only the first is described by the headline figure.
The row that matters in a disciplinary meeting is the last one. Figure drawn by AI Tools Primer.
One honest complication, because the tidy version of this story is not true. Not every
university that dropped the tool dropped it over accuracy.
San Francisco State announced its own discontinuation for an entirely different reason: Turnitin
had provided AI detection free during a beta period, then “notified campuses that the AI detection
feature would be add-on for additional cost and that the free trial period would end”. The feature
went away because the contract changed. If you are arguing that the sector has rejected these tools
on the evidence, that is one of the cases that does not support you — and it is worth knowing before
somebody else points it out.
The neighbouring question — how do ai detectors work — has an answer that is short and unsatisfying, and it is the same one Vanderbilt gave: outside the companies, nobody knows. That is also why asking how accurate are ai detectors cannot be settled from outside. Accuracy is measured against a method, and the method is not published.
So what is a defensible position for somebody who has to decide something today?
If you are teaching: a detection score is a reason to look, never a reason to conclude. The conversation that follows — about drafts, sources, the argument in the essay — is the actual assessment, and it was always going to have to happen anyway.
If you have been accused: ask what the score is based on, and ask it in writing. The answer, if it is honest, is that the method is not disclosed. That is not a rhetorical trick; it is the documented state of the tool, published by a university that paid for it.
And if you are writing anything you might one day need to defend, the single most valuable habit costs nothing and no detector can argue with it: work somewhere that keeps a version history. A document with three weeks of edits behind it is evidence of a kind that a probability score cannot touch.
The guides below take it four ways: what a score actually measures, the false positives and who gets them, the individual tools and what each claims, and what to do when an accusation has already been made.
Not covered here. It will not tell you a detector is reliable enough to accuse somebody. The most detailed public reasoning on that question was published by a university, and it reached the opposite conclusion.
It will not tell you the sector has settled this. One of the universities that dropped the tool did so because the contract changed, and pretending otherwise would be the same overreach we are objecting to.
And it will not carry an affiliate link. In this section that means declining some of the highest commissions on the site. What holds instead is simple: the accuracy figure and the arithmetic behind it are quoted from the universities that published them, and no verdict here goes past what they wrote.
When a detector has put a number on your work
A percentage appeared. What it means depends entirely on who is reading it and what they are allowed to do with it.
Vanderbilt University, Brightspace — Guidance on AI Detection and Why We’re Disabling Turnitin’s AI Detector (the 1% false positive claim, the calculation against 75,000 submitted papers, the absence of published method, the bias against non-native English speakers, and the rollout with less than 24 hours’ notice). Published 16 August 2023 — www.vanderbilt.edu, read 22 August 2026.
San Francisco State University, Academic Technology — Discontinuation of TurnItIn AI detection tool availability (that the feature was free during a beta phase and became a paid add-on, which is why it was withdrawn there). Published 1 May 2024 — at.sfsu.edu, read 22 August 2026.
An Empirical Study of AI Generated Text Detection Tools — arXiv:2310.01423 (that a multi-domain dataset of articles, abstracts, stories, news and product reviews was built to test detection tools used by universities and research institutions; that six systems named GPTkit, GPTZero, Originality, Sapling, Writer and Zylalab were tested; and that their accuracy rates fall between 55.29 and 97.0%, with the authors noting that all the tools fared well in the evaluations) — arxiv.org, read 26 August 2026.
AG
Written by Alberto Gulotta
Founder and editor of AI Tools Primer, writing from Palermo, Italy. Thirty-five years of
taking computers apart, starting with a Commodore 64 — the long version is on the
about page.
Written on 22 August 2026 · last checked 23 September 2026.
§
Independence and limits
No affiliate links and no paid placements anywhere on this site. Nobody pays to appear here,
and no company has seen this page before you did.
This is general information, not professional advice. Where a page touches money, health,
safety or the law, it names its source and the date it was read — and your situation may
still differ. See the privacy page and the
cookie policy.