1. What is analysed
We analyse the written text you submit, nothing else. If you paste text, we use it as it is. If you upload a TXT, DOCX or PDF file (up to 10 MB), we extract the text from it first. Formatting, images and file properties are ignored. A PDF that is only a scan or a photo has no text layer, so there is nothing to read; in that case we tell you so instead of guessing. A single check accepts up to 5,000 words.
2. Who does the detection
The detection itself is performed by the Pangram Labs API ( pangram.com ). We built the interface around it: the input handling, the result summary and the highlights. Pangram splits your text into passages, which it calls windows. It labels each one as AI-Generated, AI-Assisted or Human Written, with High, Medium or Low confidence, and reports the share of the text in each class.
3. What a score means
We show four things:
- Verdict. One of Mostly human-like, Mixed signals, Likely AI-assisted, or Not enough text.
- Flagged as AI-like. The share of the text whose passages were labelled as AI-like.
- Confidence. How sure the detector is about its labels.
- Highlights. The sentences behind the labels, so you can see where the signal comes from.
"Flagged as AI-like" is a share of passages, not a probability. A result of 40% does not say there is a 40% chance that AI wrote the text. It says that roughly four tenths of the text was labelled AI-like. A text can be entirely written by a person and still have a few flagged passages, and a text can be heavily AI-assisted with a small flagged share.
4. What a result does not mean
- It is not proof of who wrote the text or of how it was written.
- It does not show that a person cheated, and it is not grounds for a penalty by itself.
- It does not explain why a passage was labelled. The detector returns labels, not reasons.
- It does not check facts, quality or originality, and it is not a plagiarism check.
5. Text length
The minimum is 50 words. Below that we return "Not enough text" because there is too little evidence to label anything fairly. We recommend 150 words or more, and longer is better: a full essay or report gives a steadier result than a paragraph.
6. Why false positives happen
A false positive is human writing labelled as AI-like. It becomes more likely with:
- Short samples, where one unusual passage carries a lot of weight.
- Formal, templated or highly polished writing, such as policy documents or application letters.
- Writing by non-native speakers, who may write in patterns the detector has seen less often.
- Text that has been through grammar or rewriting tools, even when every idea was the author's own.
7. Why false negatives happen
A false negative is AI-influenced text labelled as human. It becomes more likely when:
- A person has edited AI output heavily or mixed it with their own writing.
- The text was produced by a newer model than the one the detector was trained against.
- The sample is short, so there is little to measure.
A human-like result therefore does not show that no AI tool was used.
8. Languages
Detection is most reliable for English. Results for other languages can be less dependable, and we do not claim equal performance across languages. If you check text in another language, treat the result with extra caution.
9. Rewritten and translated text
Text that was paraphrased, run through a rewriting tool, or translated from another language can change how it reads to a detector. Translation in particular can make human writing look unusual, and rewriting can soften the patterns that an original draft had. The result describes the text you submit, not its history.
10. Privacy and retention
Our server does not store the text you submit or the files you upload, and our logs contain no text: only a request ID, the action, the duration, the provider, the outcome and a word count. The text is sent to Pangram so it can be analysed. Pangram's own terms govern what it does with that text, so please read them if that matters to you. The full data flow is on our privacy page.
11. Testing and benchmarks
We have not yet published benchmarks of our own, so we do not quote an accuracy figure, and we will not borrow one from elsewhere. Our accuracy page explains why. When we publish testing, we intend to share:
- What the test set contained, where the human and AI samples came from, and how long they were.
- Which AI models and which kinds of human writers were included, including non-native writers.
- The false positive rate and the false negative rate, reported separately rather than as one blended number.
- How results changed for short, edited, translated and non-English text.
- The date of the test and the detector version it covered, so readers can tell whether it still applies.
Until then, use results as one signal among several.
For a gentler walkthrough of the same ideas, see how AI detection works.