Detect tower · floor

How do plagiarism checkers work, and what can they not see?

Short version: how do plagiarism checkers work is a question about string matching, not about honesty. The software cuts your document into overlapping runs of words, looks each run up in a very large index of other documents, and reports how much of your text it found somewhere else. That is the whole mechanism. Everything people believe about these tools beyond that — that they detect cheating, that they understand meaning, that a percentage means something fixed — is added afterwards by the people reading the report.

“Turnitin Similarity is a tool to support the detection of plagiarism by flagging instances of text similarity that may constitute plagiarism; however, it does not determine whether plagiarism has occurred.”

Turnitin, answering its own frequently asked question, read 21 August 2026. The largest company in the category says on its sales page that its product does not do the thing its product is named after. Everything below is an expansion of that sentence.

The vendor describes the process plainly enough. A submission is handled by “breaking down text into phrases with assigned IDs and comparing them against seven trillion matches in Turnitin’s comprehensive database, including internet articles, scholarly publications, and student papers”. Phrases, identifiers, lookups. No judgement anywhere in that sentence, because there is none in the software.

The University of Kansas puts the same fact less commercially: “A plagiarism detector has no power to think. It simply looks for similarities in language programmed into it as it compares a student’s paper to works in various databases.” Two very different institutions, one with a product to sell and one with students to protect, describing an identical machine.

So a match is a match, and nothing more. The university guidance is explicit about what ends up inside that number: flagged material “simply means that similarities were found, including appropriately used quotes, citations, and common descriptions”. A quotation you introduced correctly, attributed correctly and punctuated correctly is a match. It has to be — you copied it deliberately, which is what a quotation is.

Four settings that change the score without changing a word of the text

From the University of Kansas Center for Teaching Excellence, whose guidance is written for the instructors who configure these tools rather than for the students who receive the number.

None of those four has anything to do with honesty. They are configuration. Which is the practical reason a percentage is not comparable between two classes, two universities, or even the same essay submitted twice.

This is the point at which the honest answer diverges sharply from the popular one. People ask what percentage is safe. There is no such percentage, and the vendor says so: “there is no ‘right’ or ‘wrong’ number to receive as a score.” The University of Kansas goes further and criticises the practice by name, reporting classes “in which instructors automatically reject papers that have 10% to 15% of their content flagged” and calling that approach one that “fails to take into account the inherent weaknesses in plagiarism detectors”.

A literature review with careful quotation can sit at thirty per cent and be impeccable. A short reflective piece can sit at four per cent and be copied wholesale from a source nobody indexed. The number is not measuring what people think it measures, and the instructor guidance from a major public university says so in as many words.

Which brings us to the more useful half of the question — the half nobody sells, because there is no product in it.

What a plagiarism checker cannot see. The index is enormous but it is still an index, and four categories of text sit outside it by construction. Work that was never digitised or never published: a paper written by a friend and never submitted anywhere is invisible, which is precisely why commissioned essays are the form of academic misconduct these tools are worst at catching. Text behind walls the crawler cannot pass: the vendor’s coverage is described as “current and archived web pages, and premium subscription articles from top publishers”, which is a large set and a bounded one. Translation: a passage rendered from another language is, at the level of characters, an entirely different document. And ideas: a structure, an argument, a research design taken wholesale and rewritten in your own words produces no match at all, while remaining plagiarism under every academic definition there is.

That last one deserves its own sentence, because it inverts the usual anxiety. The behaviour the software punishes hardest — visible, attributed, honest quotation — is the legitimate one. The behaviour it is blindest to is the deliberate one.

Does a plagiarism checker detect ChatGPT? No, and the distinction is a product distinction rather than a technicality. Similarity checking and AI writing detection are two different systems sold separately: Turnitin states that its customers access AI writing detection “through our product add-on, Turnitin Originality”. A similarity score of zero on AI-written text is the expected result, not a failure, because the text genuinely is not copied from anywhere. It was generated. Those are different accusations, they rest on different evidence, and the floor about detectors in this tower covers the second one and its considerable problems.

None of this makes the tools useless. Used as the university describes them — as information rather than as a verdict, read rather than skimmed, with the report opened instead of only the percentage — a similarity report is genuinely useful. It catches the paragraph you meant to paraphrase and did not. It shows you a quotation whose closing mark went missing three drafts ago. It finds the source you used and forgot to list. Those are real services, and they are the ones worth running a check for before you submit anything.

What it will not do is tell anybody whether you cheated. The company that built it says so on its own page, and that is the single most useful sentence in this entire subject.

Where to start

Four ways in.

“I want to check my own work before submitting.”
Start at checking your own work
“I got a percentage and I do not know if it is bad.”
Go to reading the report
“Which checker should I use?”
That is the tools
“Is this the same as an AI detector?”
No — see the other question

Checking your own work

The useful case, and the one with the fewest complications. You are looking for your own mistakes before somebody else looks for them.

Free plagiarism checkersWhat the free tier really covers, and whether your text joins a database.Being built
What happens to your textWhat each service says it keeps, quoted with the date. The one question to settle before pasting a whole thesis in.Being built
Self-plagiarism and citationReusing your own work, and the rules that differ by institution.Open this floor →

Reading the report

A percentage on its own is not information. The report underneath it is, and it takes about ten minutes to learn to read.

The tools

Older technology than the detection wing, better understood, and mostly doing what it claims. The differences are coverage, price, and what happens to your document afterwards.

The other question

Similarity and AI writing are two different accusations resting on two different kinds of evidence. Conflating them is the most common mistake in this whole subject.

What this tower will not do

It will not give you a safe percentage. Both sources on this page say explicitly that no such number exists, and inventing one would be the most popular and least honest thing this floor could do.

It will not help anybody defeat a checker. The limits described above are documented facts about how the software works, published by the company that sells it and by a university that teaches with it.

And it will not carry an affiliate link to a plagiarism checker, which in this category means turning down a commission on every reader who buys one. What holds instead is simple: every quotation on this page comes from Turnitin’s own product documentation or from the University of Kansas Center for Teaching Excellence, both read on 21 August 2026.

Where this page got its facts

  1. Turnitin — Turnitin Similarity, including the FAQ on what the product does and does not determine, how a Similarity Report is generated, and how the score should be interpreted — www.turnitin.com, read 21 August 2026.
  2. University of Kansas, Center for Teaching Excellence — Careful use of plagiarism checkers, on what flagging does and does not mean and which settings change the result — cte.ku.edu, read 21 August 2026.

Written by Alberto Gulotta

Founder and editor of AI Tools Primer, writing from Palermo, Italy. Thirty-five years of taking computers apart, starting with a Commodore 64 — the long version is on the about page.

Something wrong on this page? Write to aitoolsprimer@gmail.com and it gets fixed.

Written on 21 August 2026.

Independence and limits

No affiliate links and no paid placements anywhere on this site. Nobody pays to appear here, and no company has seen this page before you did.

This is general information, not professional advice. Where a page touches money, health, safety or the law, it names its source and the date it was read — and your situation may still differ. See the privacy page and the cookie policy.