Detector
Score
Evidence
How Accurate Are AI Detectors?
How Accurate Are AI Detectors? What the Numbers Really Mean
AI detectors give a probability, not proof. How GPTZero, Turnitin, ZeroGPT and Originality.ai score text, their false-positive rates, and how to read a result.
An AI detector does not know who wrote a text. It measures how predictable the words are, how evenly the sentences are built, and how closely the whole thing resembles the output of the models it was trained on. The score it returns is a probability. Treat it like a weather forecast, not a court verdict.
What detectors actually measure
- Perplexity: how surprised a language model is by each next word. Human writing has more surprises.
- Burstiness: variation in sentence length and structure. People mix long and short sentences; models tend to even them out.
- Stylometric features: favourite transitions ("Furthermore", "In conclusion"), hedging phrases, balanced three-item lists, and a flat, polite register.
- Classifier training: most commercial tools also run a model trained on labelled human and AI samples, which is where vendor-specific differences come from.
Published accuracy and the false-positive problem
Vendors quote impressive numbers. Turnitin says its AI indicator has a document-level false-positive rate under 1% for documents with at least 20% AI writing, and it suppresses scores below 20% for that reason. GPTZero publishes its own benchmarks and recommends human review before any decision. Independent tests, including a 2023 Stanford study, found that several detectors flagged essays by non-native English speakers as AI at far higher rates than essays by native speakers, because formal, limited-vocabulary writing has low perplexity.
| Situation | Typical detector behaviour |
|---|---|
| Raw ChatGPT essay, no edits | Usually flagged with a high score |
| Formal human essay by a non-native writer | Real risk of a false positive |
| Short text under ~150 words | Unreliable in either direction; many tools refuse to score |
| AI draft rewritten and edited by a person | Scores fall, sometimes to near zero, but vary by tool |
| Technical or legal writing with fixed phrasing | Elevated scores even when human-written |
Why the same text scores differently on each tool
Each detector has its own training data, thresholds and minimum length. In our own October 2026 tests, an AI-written essay rewritten with FreeHumanizerAI in Ultra mode went from 100% to 0% on ZeroGPT, while salesy product copy from the same session still scored around 70%. A different detector would give different numbers for both. Any site that tells you its output passes every detector is not telling you how it measured that.
How to read a score sensibly
- Check the text length. Below a few hundred words, ignore the number.
- Look at sentence-level highlights, not just the headline percentage. Our AI detector shows the patterns it found so you can edit them.
- Run a second tool. Agreement between two detectors means more than one high score.
- If you are a teacher or editor, treat a score as a prompt for a conversation, never as evidence on its own.
Try FreeHumanizerAI for free
Humanize AI text instantly. No signup, up to 5,000 words per session, and your text is not stored.
Open AI HumanizerRelated reading
- GuideIs My Essay Flagged as AI? Why It Happens and What to DoFalse positives are common with formal, well-structured essays. Here is why it happens, how to document your own work, and how to fix the writing.Read
- ComparisonGPTZero vs Turnitin vs Originality.ai: Which AI Detector Should You Trust?Three detectors, three different audiences. Here is how GPTZero, Turnitin and Originality.ai differ and when each result is worth taking seriously.Read
- GuideCan Turnitin Detect AI Writing? What Students Should Know in 2026Turnitin does flag AI writing, with limits. Here is how the indicator works, where it is weak, and what the score means for a student.Read