Sign in or create an account

Accuracy

We do not quote an accuracy figure for this checker, because we have not run the testing that would justify one. This page explains why, and how to read results sensibly in the meantime.

Updated

Why there is no percentage here

Many detectors advertise a single accuracy number. We think that is rarely meaningful, and we have not measured one ourselves. Detection runs on the Pangram Labs API, and we have not published an independent evaluation of it. Quoting a figure we did not produce would suggest more certainty than we have.

Why a single number can mislead

A claim that a detector is "highly accurate" sounds clear, but it leaves out most of what you need to know:

  • Which mistakes? A false positive (human writing flagged) and a false negative (AI text missed) have different costs. One blended figure hides both.
  • On what text? Results on clean, long, English samples written by one model say little about a short, edited, translated or non-native essay.
  • Against which models? New language models appear often, and a test run last year may not describe today's output.
  • With what mix? If a test used equal numbers of human and AI samples but your pile of submissions is mostly human, the same detector will be wrong in a different proportion of cases.

How to interpret a result

Read the result as a pointer to passages worth a closer look. Check that "Flagged as AI-like" is a share of passages, not a probability. Notice the confidence level, and give Low-confidence labels little weight. Look at which sentences were highlighted and whether they differ from the rest of the writing.

Be most cautious when the text is short, was translated, is not in English, is formal or templated, or was written by someone still learning the language. These are the situations where misleading results are most likely, and the methodology page lists them in detail.

A detector result should never be the only reason to accuse, grade down or penalise anyone. Pair it with drafts, version history and a conversation.

What independent evaluation would need

A test that deserves trust has a few ingredients:

  1. A held-out set of texts whose authorship is known for certain, with human writing collected before and after language models became common and AI text from several current models.
  2. Human samples from different kinds of writers, including students, professionals and non-native speakers, in several lengths and genres.
  3. Edited and mixed samples, such as AI drafts revised by people and human drafts polished by tools, because that is how many people actually write.
  4. False positive and false negative rates reported separately, with the sample sizes and some measure of uncertainty.
  5. A published method, a date and a version, so others can repeat the test and see when it stops applying.
  6. A reviewer with no stake in the result.

What we will publish

When we run our own testing, we will publish the method and the results here, including the cases where the detector does badly. Until that page exists, the honest summary is this: the checker can be useful for spotting passages worth a second look, and it cannot settle who wrote something.

You can read what happens to your text on the privacy page, or run a check and judge the output against your own knowledge of the writing.