You cannot tell from the prose alone. Features such as generic examples, even structure and polished-but-empty paragraphs are worth noticing, but each has an ordinary human explanation. What you can do is gather several independent signals, weigh the ones you can verify (citations, drafting history, the writer’s ability to discuss the work) above the ones you can only feel, and treat a detector as one input rather than a verdict.
Key takeaways
- Style signals are weak evidence. Verifiable signals, such as citations that do not exist or a document with no drafting history, are stronger.
- Comparing an essay with the same student’s earlier work, and talking to them about it, is more informative than any scan.
- A sentence-level detector result tells you where to look. It does not tell you why, and it does not identify an author.
- The checklist below is built to be run in order, from the cheapest and strongest checks to the weakest, and to end with a conversation rather than a conclusion.
Can you tell by reading it?
Often less reliably than people expect. In a 2024 study in Computers and Education: Artificial Intelligence, Fleckenstein and colleagues asked novice teachers (89) and experienced teachers (200) to judge essays that were either student-written or generated by ChatGPT. The authors report that teachers could not identify the ChatGPT texts among the student texts, that experienced teachers made somewhat more accurate and differentiated judgments, and that both groups were overconfident (Do teachers spot AI?). That does not mean reading is useless. It means a hunch should be the start of a review, not the end of one.
The rest of this guide is about turning a hunch into evidence you can check.
What style signals are worth noticing, and what is wrong with each?
Style signals are the ones most pages list. They are real observations, but every one of them is ambiguous. The table pairs each with the ordinary explanation you need to rule out.
| What you notice | Why it raises a question | Ordinary human explanation |
|---|---|---|
| The voice differs from the student’s earlier work | A sudden change in sentence length, vocabulary or confidence | Tutor or writing-center help, a grammar tool, a topic the student cares about, real improvement |
| Claims with no support | Confident statements that cite nothing, or cite something vaguely (“studies show”) | Rushed drafting, weak research habits |
| Generic examples | “Many companies have adopted new technology” instead of a named case, date or place | Writing to a word count, avoiding a topic the student did not research |
| Very even structure | Every paragraph the same length, with a topic sentence and a tidy closing line | The five-paragraph essay format, taught in many schools |
| Vocabulary that does not fit the writer | Abstract or formal words used slightly off, or a register that shifts mid-essay | Copying phrases from sources, a thesaurus, a translation tool |
| Hedged, balanced framing | “There are many factors to consider” with no position taken | Caution, or an assignment that asked for a balanced view |
Concrete examples make the difference clearer. Compare two ways of supporting the same point in an essay on school start times:
Research shows that later start times improve student outcomes in many ways.
My brother’s school moved first period from 7:40 to 8:30. He stopped falling asleep in math, but the bus now arrives at 4:50 and he misses football practice twice a week.
The first is generic: nothing in it could be checked. The second is specific and a little untidy. That difference is useful when you ask the student about their essay, because specific detail is easy to expand on and generic detail is not. But a student told to write formally and avoid anecdotes could produce the first by hand, and a chat assistant can be asked to produce the second. The contrast is a prompt for a question, not a finding.
Treat these signals as a way to decide whether to look further. Do not stack three weak signals and call it proof; three weak signals can all come from one ordinary cause, such as a student who got help from a tutor.
Do the citations exist?
This is one of the few checks that produces something you can verify and document. Language models can write references that look right and are not: a plausible author, a journal that exists, a title that sounds correct, and no such paper.
How common this is depends on the model and the year. A study by Walters and Wilder, published in Scientific Reports in September 2023, examined 636 citations produced by two ChatGPT versions and found that 55% of the GPT-3.5 citations and 18% of the GPT-4 citations were fabricated, and that many of the real ones contained substantive errors (Fabrication and errors in the bibliographic citations generated by ChatGPT). Those models are now old, so the rates should not be read as current. The check still works regardless of the tool, because it asks a factual question: does this source exist, and does it say what the essay says it says?
Do it quickly:
- Search each reference by title in a library database or Google Scholar.
- Confirm that the authors, year and journal match, not just the title.
- For two or three sources that carry the argument, open them and check the claim being attributed to them.
A nonexistent source is a real problem whoever wrote the essay. It may point to AI use, to a misremembered reference or to a copied bibliography, so ask before concluding which. Also note what a clean result does and does not show: references that all exist say nothing about whether the prose was machine-written, since a student can supply real sources to a chat assistant.
What does a document’s history show?
The way an essay came into being is stronger evidence than how it reads. A word-processor history shows when edits were made, in how many sittings and by which account. In Google Docs, for example, someone with edit access can open version history and see who changed a file and when, though Google notes that some changes may not appear in the edit history (Google Docs Editors Help).
What to look for:
- A drafting trail. An outline, early rough paragraphs and revisions over several days is reassuring.
- One large paste. A whole essay appearing in a single edit is a reason to ask questions. It is not proof, because some people draft elsewhere, in a notes app, on paper or in another program, and paste in the result.
- Missing history. A file that was exported, converted or created from a template may have no useful history at all.
If you want students to be able to show their process, say so in the assignment and ask them to keep time-stamped drafts. Stony Brook University’s guidance on evaluating AI-flagged papers advises students to keep a time-stamped copy of their drafts as protection if their work is questioned (AI-Flagged Paper Evaluation Process).
Where does a sentence-level detector fit?
A detector can point at passages a reader might skim past. It cannot tell you who wrote them. Our AI detector for essays labels each passage as AI-Generated, AI-Assisted or Human Written with a confidence level, and highlights them sentence by sentence. The detector behind it, from Pangram Labs, returns labels and not reasons, so a highlight is a reason to look, never an explanation.
Useful habits when you read the highlights:
- Read where they fall. A scattered handful of flagged sentences in an otherwise human-like essay points somewhere different from one long, continuous flagged block. Read the flagged passages yourself and ask whether they are the generic, conventional parts of the essay or the parts that carry the argument.
- Check the length. Very short essays give any detector little to work with. Our checker needs at least 50 words, and results near that minimum are thin evidence.
- Read the figure for what it is. On this site, “Flagged as AI-like” is the share of the essay made up of passages labeled AI-like. It is not the probability that the whole essay was written by AI. What an AI detection percentage means covers the difference.
- Do not compare scores across tools. Different detectors use different methods and thresholds, so numbers from two products are not on the same scale.
Detector history is a reason for restraint. OpenAI released an AI text classifier in January 2023 and withdrew it on July 20, 2023, citing a low rate of accuracy (OpenAI’s own announcement, updated on withdrawal). At launch, OpenAI’s own evaluation put it at about 26% detection of AI-written text with human text mislabeled as AI-written about 9% of the time (TechCrunch’s launch coverage). Detection has improved since, and current tools differ from that one, but the lesson holds: no detector is a lie detector. Turnitin, which sells an AI writing indicator to schools, said when it launched the feature that its false positive rate was “not zero” and told instructors to apply professional judgment, knowledge of their students and the context of the assignment (Turnitin’s statement on false positives).
Can teachers tell if you used ChatGPT?
Sometimes, and rarely from one thing alone. In practice, educators who suspect AI use rely on several signals together:
- Familiarity with the student’s writing. A teacher who has read a student’s in-class work and earlier assignments has a baseline that no tool has.
- Fit with the task. An essay that answers a general version of the question rather than the specific prompt, readings or class discussion stands out.
- Sources. Citations that cannot be found, or that do not say what the essay claims.
- Process. Whether there is a drafting trail, and whether the student can describe how they worked.
- A follow-up conversation. Asking a student to explain a paragraph, define a term they used or extend an argument is the fairest test of whether they understand what they submitted.
- A detector result. Used as a pointer to passages worth rereading, not as proof.
The honest answer is also that none of these is conclusive. The Fleckenstein study found teachers overconfident in judging by reading, and detectors make both kinds of error. A 2023 Stanford study published in Patterns tested seven detectors on TOEFL essays written by non-native English speakers and found that they misclassified a large share as AI-generated, with an average false positive rate reported at about 61% on those essays, while essays by US-born eighth graders were classified accurately (Liang et al., GPT detectors are biased against non-native English writers). The authors link this to detectors reacting to predictable word choice and simple sentence structure, which are common in careful second-language writing.
This is why a good process ends with the student’s explanation, and why can AI detectors be wrong is worth reading before acting on any score. Using these signals to decide where to look is reasonable. Using them to skip the conversation is not.
What limits every one of these checks?
Be honest about the cases where the evidence gets thin:
- Mixed authorship. An essay a student drafted, then ran through a grammar tool or an AI rewriting assistant, does not fit a yes-or-no question. Whether that is acceptable depends on the course policy, which should be stated in advance.
- Short, formal or templated writing. Lab summaries, reflections and reference-style answers follow conventions that make them look uniform.
- Translation and multilingual writers. As the Stanford study above shows, this is where detectors are most likely to mislead.
- No baseline. If you have no earlier writing from the student, comparisons are not available.
- Time. Drafting history and citation checks cost minutes per essay, which is why they are best reserved for cases where something has already raised a question.
None of this makes review pointless. It means the standard for action should be higher than “something felt off.”
A review checklist for a suspect essay
Work through these in order. They run from the strongest, most verifiable evidence to the weakest, and they end with the writer. Write down what you find at each step, since a documented process is fairer to the student and easier to explain later.
| Step | Check | What counts as a meaningful finding | What it does not show |
|---|---|---|---|
| 1 | Is the task clear? Was the AI policy stated for this assignment? | An unambiguous policy that was broken | Nothing, if the policy was vague |
| 2 | Look up the references | Sources that do not exist or do not support the claim | Real sources prove nothing about authorship |
| 3 | Check document history | A whole essay pasted in at once, or no drafting trail where one was required | A missing history can have a technical cause |
| 4 | Compare with earlier work | A sharp, unexplained shift in vocabulary, structure or accuracy | Improvement from tutoring, a grammar tool or effort |
| 5 | Check fit with the prompt | Ignores class readings or the specific question | A student who skipped the readings |
| 6 | Read the highlights from a sentence-level detector | Flagged passages that sit in the argument and match steps 2 to 5 | Flags alone, on short, formal or non-native writing |
| 7 | Ask the student to talk through it | Cannot explain terms, sources or choices in their own essay | Nervousness in a meeting |
| 8 | Decide what the evidence supports | Several independent signals pointing the same way | Any single signal |
If the review ends with a worry that the evidence cannot settle, a conversation about the assignment, or a short in-class follow-up, often resolves it better than a formal accusation.
What if you are the writer and an essay of yours is questioned?
Keep your outline, notes and drafts, and know where each source came from. If someone raises a detector result, you can reasonably ask what other evidence exists, and point to the limits of these tools described above. This guide is about how reviewers can look at evidence fairly. It is not a guide to avoiding detection, and nothing here is.
Frequently asked questions
Is there one reliable sign that an essay was written by AI? No. Individual style features all have ordinary human explanations. The most dependable evidence is verifiable: citations that do not exist, a document history that contradicts the story, or a writer who cannot explain their own paragraphs. Even then, several signals together support a conclusion better than any one alone.
Are fabricated citations proof of AI use? Not by themselves. Language models are known to produce references that do not exist, but people also mis-cite, copy bibliographies without checking, or mix up details. A nonexistent source is a concrete problem to ask about. How it got there is a question for the writer.
Does a detector result count as evidence? It counts as a pointer. A flagged passage can tell you where to read more closely, and a result should be weighed with drafts, sources and the student’s own account. Do not penalize anyone on a detector result alone.
What should I do if an essay looks AI-written but I cannot prove it? Keep the question open. Ask the student about their process, request drafts, and consider a short follow-up discussion of the essay. If the evidence still does not settle it, the fair outcome may be a conversation about expectations rather than a finding.
Where to go next
If you already have the essay as text, run it through the AI detector for essays and review the flagged passages rather than relying on the overall score. For the wider picture beyond coursework, see how to tell if text is AI-generated, and for a teacher-focused overview of how the checker fits into a review, the page for teachers.