“GPTZero has a low (10%) false-positive (classifying a human-written text as AI-generated) and a high (35%) false-negative rate…”
The author’s conclusion: the tool is “more appropriate for ruling in rather than ruling out suspicious texts”.
Detect · guide
By Alberto Gulotta · Updated · 8 min read
gptzero is the most searched AI detector by name, and it is unusual in a way that deserves saying first: it publishes its own accuracy figures, in public, with the methodology attached. Most of this industry does not.
That makes it a good place to examine what a detector score actually is — because the honest reading of those numbers is not the one the numbers seem to invite.
What a one per cent false positive rate means in a classroom
GPTZero publishes a false positive rate of “no more than 1% when evaluating AI versus human text”. Taking that figure at face value and doing the arithmetic:
The number is not a criticism of the tool — as detectors go it is a good figure, and GPTZero publishes it openly, which most do not. It is an argument about what a score is for. A measure that is right ninety-nine times out of a hundred is excellent evidence about a pile of two thousand essays and poor evidence about any single one of them.
It is worth knowing where this guide sits in a much larger search. Ai detector and ai checker are among the biggest words on this site, and they are searched by two groups with opposite interests: people checking somebody else’s work, and people checking their own before handing it in. A named tool like this one is what both groups reach for once the general search has produced twenty identical pages.
The rest of this guide is what the company states about itself, quoted with the date, followed by the part nobody selling a detector writes down: how a result should and should not be used.
What people outside the company measured
Two independent studies, and they point the same way
A vendor’s own figure is a starting point, not an answer. Two published studies have tested GPTZero directly. Both are small — this is a young field and the samples are in the dozens — and both are worth reading precisely because they were not run by anyone selling a detector.
“GPTZero has a low (10%) false-positive (classifying a human-written text as AI-generated) and a high (35%) false-negative rate…”
The author’s conclusion: the tool is “more appropriate for ruling in rather than ruling out suspicious texts”.
“A vast majority of the AI-generated papers were detected accurately (ranging from 91-100% AI believed generation), while the human generated essays fluctuated; there were a handful of false positives.”
Their conclusion: “although GPTZero is effective at detecting purely AI-generated content, its reliability in distinguishing human-authored texts is limited”.
Put the two together with the company’s own number and a pattern appears that matters more than any single percentage. These tools are better at recognising AI writing than at confirming human writing. Which is the wrong way round for the way they are used — because in a classroom the question asked of the score is never “is this AI?”. It is “did this student write this?”, and that is the direction in which the evidence is weakest.
The company’s own FAQ is more cautious than the number on the front of it. On what a classification means: it “should not be solely used to indicate that an essay contains AI”. On what to do instead of acting on one report: “we recommend looking for a long-term pattern of AI use, as opposed to a single instance, in order to determine whether the student is using AI.” A single flagged essay is, by the vendor’s own instruction, not the finding — the finding is a pattern across several.
It also points at evidence that has nothing to do with the detector. If the document “has an edit history (such as Google Docs), and it was typed out with several edits over a reasonable period of time, it is likely the student work is authentic.” That is the same move three competing companies make on this site: when the score is not enough, they send you to the writing process.
Two details about the reading itself. The result is not one verdict for the whole document — “when a document gets a MIXED or AI_ONLY classification, the highlighted sentence will indicate where in the document we believe this occurred” — so a report worth acting on points at lines, not at a percentage. And the scope is broader than one chatbot: GPTZero “works robustly across a range of AI language models, including but not limited to ChatGPT, GPT-5, GPT-4, GPT-3, Gemini, Claude, and AI services based on those models”, with the model itself “trained on millions of documents spanning various domains of writing”.
One note about where to read reviews of detectors. Several widely read assessments of GPTZero are published by companies whose own product exists to defeat AI detection. Those pages may be accurate; they are also written by a party with an interest in the answer, in the same way a vendor’s own accuracy page is. The two studies above have no such interest, which is why they are quoted here in preference to any review — including ours, which is why this guide does not contain one yet.
Four ways in.
Quoted from GPTZero’s own pages rather than from a review. Where a figure comes with a stated method, the method is quoted too.
Not covered here. It will not tell you GPTZero is accurate or inaccurate. It reports what the company publishes, with the date, and does the arithmetic on those figures. A verdict would need our own controlled test — texts we wrote ourselves, run blind — and until that exists there is no ranking here.
It will not treat a percentage as proof. On the company’s own published figure, a department marking two thousand honest essays should expect around twenty of them to be flagged. That is not a scandal; it is what a one per cent error rate means, and it is the single most important fact on this guide.
And it will not carry affiliate links. Detection and its opposite are among the highest-paying referral categories on this site, which is exactly why the pages about them have to be free of that money. What holds instead is simple: every figure on this guide is quoted from the company that published it, with the date.
A percentage appeared. What it means depends entirely on who is reading it and what they are allowed to do with it.
Written by Alberto Gulotta
Founder and editor of AI Tools Primer, writing from Palermo, Italy. Thirty-five years of taking computers apart, starting with a Commodore 64 — the long version is on the about page.
Something wrong on this page? Write to aitoolsprimer@gmail.com and it gets fixed.
Written on 21 August 2026 · last checked 4 September 2026.
Independence and limits
No affiliate links and no paid placements anywhere on this site. Nobody pays to appear here, and no company has seen this page before you did.
This is general information, not professional advice. Where a page touches money, health, safety or the law, it names its source and the date it was read — and your situation may still differ. See the privacy page and the cookie policy.