Sign in or create an account

How Accurate Are AI Detectors?

Updated

It depends, and any single percentage hides the part that matters. Independent studies have found AI detectors that work well on long, unedited, English text from a known model, and detectors that fail badly on edited, translated or non-native writing. Vendors report very low false positive rates on their own test sets. Outside evaluators have reported much weaker results for older tools. The honest answer is a set of figures, each with conditions attached.

Key takeaways

  • “Accurate” mixes two questions: how often AI text is caught, and how often human text is wrongly flagged. Only the second one harms an innocent writer.
  • Figures from a vendor’s own test set describe that test set. They are not a promise about your text.
  • Length, genre, language, the model that wrote the text and how much a person edited it all move the result.
  • No study cited here supports using a detector score as proof of authorship.

What does “accurate” mean for an AI detector?

A detector makes two kinds of mistakes. A false positive is human writing labeled AI-like. A false negative is AI-generated text labeled human. Researchers often describe the same thing with two terms: sensitivity is the share of AI text the tool catches, and specificity is the share of human text it correctly leaves alone. You do not need the terms to read a result, but you do need to know that a single “accuracy” figure blends the two.

A blended figure also depends on the mix of samples. If a test used half AI and half human text, the number says little about a pile of submissions that is mostly human. A small false positive rate still produces real cases at volume. As plain arithmetic, not a measured result: a 1% false positive rate applied to 1,000 human-written essays would flag about 10 of them.

That is why serious evaluations report the two error rates separately, with the test set, the models and the date. Our accuracy page lists what a trustworthy test would need.

What do independent studies actually find?

The table collects the most useful published evidence. Each row states who produced the figure and the conditions, because the conditions decide whether it applies to you.

Source What was tested What it reported Read it as
Liang et al., Stanford (Patterns, 2023) Seven GPT detectors on 91 TOEFL essays and 88 US eighth-grade essays Near-perfect accuracy on the US essays; about 61% of the TOEFL essays misclassified as AI on average Strong evidence of bias against non-native writing in 2023-era detectors
Weber-Wulff et al. (Int. J. for Educational Integrity, 2023) 14 tools, including Turnitin and PlagiarismCheck Tools “neither accurate nor reliable”, biased toward calling text human; obfuscation made results worse Independent warning from academic-integrity researchers about tools available in 2023
RAID benchmark (ACL 2024) 12 detectors on over 6 million generations: 11 models, 8 domains, 11 adversarial attacks Detectors were “easily fooled” by attacks, different sampling settings and unseen models Detectors that look strong on familiar data can fall apart on unfamiliar data
OpenAI AI classifier (2023) OpenAI’s own tool on its own evaluation set Caught 26% of AI text and wrongly flagged 9% of human text; withdrawn 20 July 2023 for low accuracy Even a model maker could not build a dependable one at that time
Jabarian and Imas, NBER working paper (2025) Pangram, Originality.ai, GPTZero and an open-source RoBERTa detector Commercial tools beat the open-source one; Pangram met a false positive cap of 0.5% without losing detection Newer, independent and encouraging for one vendor, but a working paper, not peer-reviewed

Two cautions apply to the whole table. The first three rows tested tools that are now several versions old, so they show what can go wrong rather than how any current product performs. The later rows are recent, but one study of one set of text is still one study. Treat this as a map of the evidence, not a ranking.

Why do vendor accuracy claims differ so much from independent tests?

Vendors usually measure on text they can control: a mix of human and AI samples, often split from the same source used for training, with known models and little editing. That is a legitimate test, and it is informative, but it answers a narrower question than “how will this behave on my essay?”

Pangram, whose API powers this site’s checker, is a useful example because its documentation is specific. The model card for Pangram 3.2 reports false positive rates from 0.00% to 0.54% across its human datasets and false negative rates from 0.00% to 1.98% on AI datasets. It describes the evaluation as in-domain test sets plus held-out sources and domains and third-party benchmarks. The same document lists situations where false positives become likelier: bullet-point lists, instructions and technical manuals, tables of contents, reference sections, templated or automated writing, and dense mathematics. It also says the model’s resolution is about 50 words, so short human passages inside AI text may not be separated out.

Those are vendor-reported figures for one model version, and the version behind the API at any given moment may differ. The site has not independently verified them, which is why the accuracy page publishes no number.

Turnitin gives a comparable kind of statement. It reports a document-level false positive rate of under 1% for documents with more than 20% AI writing, tested on over 700,000 papers written before ChatGPT, and a sentence-level false positive rate of around 4%. It adds that 54% of those wrongly flagged sentences sit right next to genuinely AI-written text. Note the condition built into the first figure: it covers documents above a threshold, not every document.

Does text length change accuracy?

Yes, and every vendor documents it. A detector needs enough words to see a pattern. Pangram sets a 50-word minimum and explains the reason plainly: a word like “delve” appears more often in AI writing but “on its own, delve could be written by anyone.” OpenAI’s classifier said its reliability improved as input length grew.

In practice, a result on two or three sentences is a hint at best. Our guide on whether detectors can be wrong covers how to treat thin results.

Do different AI models affect detection?

They do. A detector learns the habits of the models it was trained on. A newer or unfamiliar model can write in a way the detector has not seen, and the RAID authors found that unseen generative models were one of the things that fooled detectors. Settings matter too: RAID tested four decoding strategies, and results changed with them.

For a reader, the consequence is that “detects ChatGPT” is not a stable property. It describes a tool at a point in time, against the models of that time. Look for the model and the date behind any claim.

What about AI text that a person has edited?

This is the weakest area for the field. Mixed text, such as an AI draft revised by a person or a human draft polished by a tool, is neither clearly one thing nor the other. Sadasivan et al. showed that recursive paraphrasing can substantially reduce detection rates across several families of detector, with only slight loss of text quality. Weber-Wulff et al. found that obfuscation significantly worsened tool performance. Turnitin’s own figures show that errors cluster at the boundary between human and AI passages.

For that reason this site labels passages AI-Generated, AI-Assisted or Human Written rather than offering a yes-or-no verdict, and its “flagged” figure is a share of text, not a probability. What an AI detection percentage means explains the difference. Even so, the AI-Assisted boundary is the least certain line any detector draws.

Are detectors accurate for non-native English and other languages?

This is where the strongest independent evidence of harm sits. Liang et al. found that seven detectors misclassified about 61% of TOEFL essays, sourced from a Chinese educational forum, as AI-generated, while handling US eighth-grade essays almost perfectly. All seven agreed on 18 of the 91 TOEFL essays, and at least one flagged 89 of them. The authors link this to perplexity: text with limited word-choice variety scores as more predictable, and predictable text looks machine-like to perplexity-based tools.

The study has limits. It used 2023 detectors, 91 essays from one source, and a specific kind of writer. It cannot tell you how a current tool behaves. But it identifies a mechanism, and the mechanism does not go away because a tool is newer. Pangram’s technical report states its classifier is “not biased against nonnative English speakers”, and its model card lists 22 supported languages. Those are vendor statements, and the site has not tested them. Until someone independent does, a flag on a language learner’s writing deserves the most caution of any case in this article.

Support for a language also does not mean equal reliability. Treat results outside English as less certain than results in English unless a source shows otherwise for that language.

Can AI detectors detect essays?

They can flag essays, and for fully AI-generated, unedited essays in a familiar model’s style they often do. Student writing is a standard domain in vendor testing: Pangram’s technical report names student writing among its ten text categories, and Turnitin’s tool is designed for student submissions.

But essays combine nearly every risk factor in this article. They are formal, often short, often written by people still developing their English, and often drafted with help from grammar tools, tutors or AI. A result on an essay therefore tells you which passages to read more closely, not who wrote them. A small experiment with human readers points the same way: a 2025 preprint found that annotators who frequently use LLMs for writing were accurate at spotting AI text, with a majority vote of five misclassifying only 1 of 300 articles. That sample was non-fiction articles and not essays, so it does not transfer directly, but it is a reminder that evidence beyond a score exists. See AI detector for essays for how to read results on student work.

How should you read an accuracy claim?

Use these questions on any figure, including ones in this article.

Ask Why it matters
Who measured it? A vendor test set and an independent evaluation answer different questions.
Which error does the number describe? A figure that blends false positives and false negatives hides both.
On what text? Genre, length, language and writer background change results.
Which models, which version, what date? Detectors age as new models appear.
Was the text edited? Unedited output is the easy case.
How many samples? Ninety-one essays and six million generations support different conclusions.
What happens at the boundary? Mixed text is where most disputes arise.

Is text watermarking a more reliable alternative?

In principle a watermark is a different kind of evidence, because the generator puts a signal into the text itself and a detector with the key looks for it. OpenAI says it does this for ChatGPT text in the EU, and offers it to API customers, but the detector is restricted to approved researchers and academic organizations (OpenAI, read 5 October 2026). Its own documentation lists the limits: short passages (the EU code of practice does not require marking below about 150 words) are unreliable, factual text leaves little room to embed a pattern, detection rates vary by language, and substantial paraphrasing or translation can erase the signal. It covers only text from models that apply it. For most text a reader meets, a classifier-style detector is still the only check available, which is why the limits in this guide still apply.

What can’t any AI detector do?

It cannot show who wrote a piece of text, how it was produced or why a passage was flagged. It cannot turn a probability into proof. And it cannot promise the same behavior on text unlike what it was tested on. Our checker uses the Pangram Labs API and returns labels with a confidence level, not reasons, which is why the methodology page lists what it analyzes and where it is weakest. How AI detection works describes the general approach in plain language.

The practical rule that follows from the evidence is simple: use a result to decide what to read more carefully, then look at drafts, sources and the writer’s other work before deciding anything.

If you already have the text, run it through our AI-generated text detector and review the flagged passages rather than relying on the overall score.

FAQ

Are AI detectors accurate enough to prove cheating? No study cited here supports that use. Even the lowest vendor-reported false positive rates still imply some wrongly flagged human writers at volume, and independent research has found much higher error rates for non-native writing in older tools. Treat a flag as a reason to look at drafts and talk to the writer.

What is a good false positive rate for an AI detector? Lower is better, but the figure only means something with its test set. Vendor-reported rates below 1% describe the vendor’s data. Ask which text, languages and lengths were included, and whether an independent party reproduced it.

Do AI detectors work on short text? Poorly. Pangram requires at least 50 words, and OpenAI’s retired classifier was less reliable on short input. A few sentences carry too little pattern to support a confident label.

Are newer detectors more accurate than the 2023 ones? Possibly, and the 2025 NBER working paper reported strong results for one commercial tool. But it is a single recent study, not yet peer-reviewed, and it does not settle behavior on edited text or non-native writing. Look for current, independent tests.