Detect tower · floor
What is an acceptable AI detection score?
The honest answer to what is an acceptable AI detection score is that there is no such number, and the company with the largest detector in education has published the data that explains why. Not a critic, not a study, not a journalist: the vendor, on its own blog, under the name of its Chief Product Officer.
Annie Chechitelli, Chief Product Officer at Turnitin, on the company blog, 16 March 2023. The number
is described by the company that produces it as an input to somebody’s judgement, never as a finding.“Turnitin does not make a determination of misconduct even in the space of text similarity;
rather, we provide data for educators to make an informed decision… given that our false positive rate
is not zero, you as the instructor will need to apply your professional judgment.”
Start with what the percentage is. It is not a measurement of how much AI you used. It is the proportion of your document that a statistical model classified as more likely machine-written than human-written. The model is guessing, in the technical sense — assigning probabilities — and the output is then rendered as a whole number, which is a presentation choice that makes a guess look like a reading from an instrument.
Turnitin defines the failure mode plainly: “A false positive refers to incorrectly identifying fully human-written text as AI-generated.” And it publishes a headline figure for how often that happens: “a less than 1% false positive rate”. On its own that sounds settled. It is not, and the company said so itself two months later.
From Turnitin’s update of 23 May 2023, after eight hundred thousand academic papers
written before ChatGPT existed were run through the detector as a test. Read the first row again, because it reverses the thing almost everybody assumes. People treat a
small percentage as a small problem. The company treats it as the least trustworthy reading its
product produces, and prints a mark next to it saying so.What the company says about its own bands
That asterisk is the single most useful fact in this subject, and hardly anybody who receives a score has been told it exists. A student handed a fourteen per cent is usually being shown the range the manufacturer has formally flagged as unreliable — and the response, in far too many classrooms, is a meeting rather than a shrug.
There is a second number, and it is not the one the vendor puts on its marketing pages. The University of Georgia, which permits exactly one detector on its campus, tells its instructors that Turnitin “reports a sentence-level false positive rate of 4%”, and adds that studies show the rate “is higher for students who speak English as a second language, as well as for text that was modified by a grammar modulator”. That figure and the document-level “less than 1%” are not in conflict: a document score is built out of sentence judgements, so a long essay contains many chances for a sentence to be misread. It is the arithmetic behind the asterisk.
The same university names two further limits that matter more than any percentage. Unlike a similarity report, “it is not possible for any AI detection tool to offer such context or suspected source” — there is no original to compare against, so a flag cannot be checked, only believed or doubted. And the detector “was trained on output from GPT-3.5”, which means “those with the resources to access more sophisticated tools than GPT-3.5 are better able to avoid detection”. The instrument is weakest against the best-resourced and most likely to misfire on people writing in a second language.
So the question “what score is acceptable” has no answer, but the question underneath it does. People asking it want to know whether they are in trouble. The vendor’s guidance to instructors, published alongside the false positive rate, answers that more directly than any threshold could: “Assume positive intent — in this space of so much that is new and unknown, give students the strong benefit of the doubt. If the evidence is unclear, assume students will act with integrity.”
That is the company telling its own customers not to treat the output as evidence of wrongdoing. It is worth quoting to anybody who has been handed a percentage and asked to explain themselves.
The bias question, and what the vendor’s figures actually cover. The most serious charge against AI detectors is that they flag writers composing in a second language more often. Turnitin says it tested this: “We tested another nearly 2,000 writing samples of ELL writers… ELL writers received a 0.014 false positive rate, and native English writers received a 0.013. This means that there is no statistically significant bias against non-native English speakers.”
Two things are true about that at once, and honest reading needs both. It is real testing, published with numbers, which is more than most detectors offer. And it carries a condition stated in the same sentence: the finding applies to “documents meeting the 300-word count requirement”. Short pieces sit outside it. A three-hundred-word floor covers a term paper comfortably and covers a discussion post, a reflection, or a short answer not at all — and short assignments are exactly where a single flagged passage becomes a large percentage of a small document.
What to do with a number you have been given. Not much, on its own — and that is the point rather than an evasion. Open the report rather than reading the headline figure, because the sentences the model flagged tell you far more than the total does. Check the length of what you submitted against the three-hundred-word condition. Note whether the figure falls below twenty per cent, and if it does, say so, because the manufacturer has published a statement about that range that carries more weight than any argument you could make yourself. And keep your draft history, which remains the only evidence that has ever settled one of these conversations quickly.
What not to do. Do not run the text through a humaniser to bring the number down. It changes the score without changing whether you wrote the thing, it degrades the writing, and if the work was genuinely yours it converts a defensible position into an indefensible one. The floor on humanisers in this tower explains what those tools actually do to a text.
One last distinction, because it is the most common confusion in the whole tower. An AI writing score and a similarity score are different products measuring different things. Similarity asks whether your words appear somewhere else; AI detection asks whether a model thinks a machine produced them. A document can score zero on one and high on the other, and neither result is evidence about the other. The floor on plagiarism checkers covers the first.
The dates matter here and are given deliberately. The false positive figures quoted on this page were published by Turnitin in March and May 2023, and the language-learner testing appears on the company’s AI writing page as read on 21 August 2026. Detection models have been revised repeatedly since those posts. Where a company has published a number, this floor quotes it with its date rather than presenting it as current fact — which is the only responsible way to write about a moving target.
Where to start
Four ways in.
- “I got a percentage and I want to know if it is bad.”
- You are on the right floor — then see if you have been flagged
- “Can these tools be trusted at all?”
- Go to how reliable are they
- “Which detector produced this?”
- That is the detectors
- “Is this the same as a plagiarism score?”
- No — see the other score
If you have been flagged
The practical wing. What helps, what does not, and what to assemble before replying to anybody.
How reliable are they
The evidence, from the companies that sell the tools and from the researchers who tested them independently. The two do not always agree.
The detectors
One floor per tool, each answering the same questions: what it claims, what independent testing found, and what a given score can honestly support.
The other score
Similarity and AI writing are two different accusations resting on two different kinds of evidence. Conflating them is the most common mistake in this tower.
What this tower will not do
It will not name a safe threshold. The vendor’s own published testing is the reason no such number exists, and inventing one would be the most popular and least honest thing this floor could do.
It will not tell you a detector is reliable enough to accuse somebody. The company that sells the largest one says it does not determine misconduct, and that sentence is quoted above with its date.
And it will not carry an affiliate link to a detector or to a humaniser, which in this tower means turning down some of the highest commissions on the island. What holds instead is simple: every figure on this page is quoted from Turnitin’s own published statements, each with the date it was published.
Where this page got its facts
- Turnitin — Understanding false positives within our AI writing detection capabilities, by Annie Chechitelli, Chief Product Officer, published 16 March 2023 — www.turnitin.com, read 21 August 2026.
- Turnitin — AI writing detection update from Turnitin’s Chief Product Officer, published 23 May 2023, on the higher incidence of false positives below twenty per cent and the asterisk added to those scores — www.turnitin.com, read 21 August 2026.
- University of Georgia, Center for Teaching and Learning — Academic Honesty & Generative AI, on the sentence-level false positive rate and the published limits of the detector — ctl.uga.edu, read 21 August 2026.
- Turnitin — AI checker solutions, including the false positive rates published for English Language Learners and native English writers — www.turnitin.com, read 21 August 2026.
Written by Alberto Gulotta
Founder and editor of AI Tools Primer, writing from Palermo, Italy. Thirty-five years of taking computers apart, starting with a Commodore 64 — the long version is on the about page.
Something wrong on this page? Write to aitoolsprimer@gmail.com and it gets fixed.
Written on 21 August 2026.
Independence and limits
No affiliate links and no paid placements anywhere on this site. Nobody pays to appear here, and no company has seen this page before you did.
This is general information, not professional advice. Where a page touches money, health, safety or the law, it names its source and the date it was read — and your situation may still differ. See the privacy page and the cookie policy.