Sign in or create an account

What Does an AI Detection Percentage Mean?

Updated

When a detector says “40%”, most people read it as “there is a 40% chance a machine wrote this.” On this site that reading is wrong. Our “Flagged as AI-like” figure is a share of the text, not a probability.

Key takeaways

  • “Flagged as AI-like” is the proportion of passages the detector labeled AI-like, such as AI-Generated or AI-Assisted.
  • It is not the probability that the whole text, or any one sentence, was written by AI.
  • A text can be 40% flagged because 40% of it looks AI-like, which says nothing about how likely it is that someone used AI at all.
  • Read the percentage together with the verdict, the confidence level and the highlighted sentences.

Two different questions

A percentage can answer two very different questions.

Question one: how much of this text looks AI-like? The answer is a share. If you split a document into passages and label each one, you can count how many fall into each class. If 3 of 10 passages are labeled AI-like, the share is 30%.

Question two: how likely is it that this text was written by AI? The answer is a probability about the document as a whole, and it depends on things a passage-level label does not capture, such as how common AI use is among the writers you are checking.

Our figure answers the first question only. It is a proportion of passages, not a statement of likelihood.

How the figure is produced

The detection service behind our checker, the Pangram Labs API, divides the text into passages, which it calls windows. Each window gets one of three labels, AI-Generated, AI-Assisted or Human Written, together with a confidence level of High, Medium or Low. The service also reports the share of the text in each class.

We show “Flagged as AI-like” as the share of text labeled AI-like. We then turn the numbers into one of four plain-language verdicts: Mostly human-like, Mixed signals, Likely AI-assisted, or Not enough text when the sample is too short to say anything useful. The detector returns labels for passages. It does not explain why any passage received its label, and we do not invent reasons.

Why it is not a probability

Here are the practical differences.

A share can be high on a document that is entirely human-written. If a detector misreads a formal, uniform document, many passages may be flagged, and the share will be high even though a person wrote every word.

A share can be low on a document that used AI. If someone prompts a model, then rewrites it substantially, few passages may be labeled AI-like. The percentage is low, but AI was involved.

A share says nothing about base rates. A probability that a text is AI-written depends on how often AI-written text appears in the group you are checking. The same flagged share means something different in a class where almost nobody uses AI than in a content farm where almost everyone does. A passage-share figure does not include that information, so it cannot be converted into a probability.

The percentage is sensitive to where passages are split. Passage-level labeling means that a short insertion in the middle of human text may or may not register, depending on how the text is divided.

Reading some example results

These are made-up illustrations of how to think, not real checks.

Example A: 5% flagged, verdict “Mostly human-like”, high confidence. One passage out of twenty was labeled AI-like. Read that passage. It may be a generic definition or a standard sentence that anyone could have written. This result gives no reason for concern on its own.

Example B: 35% flagged, verdict “Mixed signals”. About a third of the passages look AI-like and the rest do not. That fits several situations: a writer who used an assistant for some sections, a writer who edited with AI tools, or a human text with some formulaic passages. Look at where the highlights cluster. Are they the introduction and conclusion, or scattered? Ask the author about their process.

Example C: 85% flagged, verdict “Likely AI-assisted”, with mostly high confidence. Most of the text looks AI-like. This is a stronger signal, but still not proof. Formal and templated writing can behave this way, and so can writing that is heavily edited by tools. Compare against the writer’s other work and ask them how the text was produced.

Example D: “Not enough text”. The sample was too short. The minimum for a check is 50 words, and short samples give too little for a reliable label. Add more text rather than reading anything into a number.

The role of confidence

Each labeled passage carries a confidence level, High, Medium or Low. Confidence reflects how strongly the detector favors the label it gave. It is not the same as being correct. A confident label on an unusual kind of text, such as a translated abstract, can still be wrong. Use confidence to decide which passages to read first, not to decide whether a person did something wrong.

What to do with the figure

  1. Read the verdict first, then the share. The verdict summarizes how the passages added up.
  2. Open the highlights. Passages tell you more than a total does.
  3. Check the length. Short text gives noisy shares. Our text detector works from 50 to 5,000 words, and longer samples are steadier.
  4. Think about the writer and the genre. The methodology page lists the situations where results are least dependable, and can AI detectors be wrong goes through them with examples.
  5. Gather other evidence before drawing any conclusion. Drafts, sources and a conversation count for more than a score. How to tell if text is AI-generated covers these checks.

A note for educators

If you are reviewing student work, the share is a starting point for a discussion, not grounds for a penalty. Our teacher page and essay checker show how to read the highlights next to the rest of a student’s work. No result from this site should be the only basis for an accusation or a grade.

Where your text goes

Because people ask: our server does not store the text you submit, and our logs record only a request ID, the action, duration, provider, outcome and word count. The text is sent to Pangram Labs to be analyzed, and Pangram’s own terms govern what it does with it. Our privacy page has the details and a link to those terms.