Sign in or create an account

What to Do When an AI Detector Flags Student Writing

Updated

Do not act on the score. A flag from an AI detector tells you which passages resemble machine-written text, not who wrote them or how. Treat it as a reason to look closer, gather evidence a detector cannot supply (earlier work, sources, drafting history, the student’s own explanation), and follow your school’s process before anything reaches the student as an accusation.

Key takeaways

  • Teachers who suspect AI use usually combine several signals: a detector, comparison with earlier work, citations, revision history, a conversation and the assignment itself. No single one settles the question.
  • Reading by feel is not reliable either. In a 2024 study, novice and experienced teachers could not identify ChatGPT-written essays among student essays, and both groups were overconfident.
  • The workflow below has seven steps and ends with a decision that rests on evidence, not on the percentage.
  • A neutral conversation script and a one-page table are included. Both are written so a student who did the work can answer them easily.

What does the 7-step workflow look like at a glance?

This table is the whole process on one page. Each step is explained below it. Work through the steps in order; the early ones are cheap and the later ones take more of your time, so most concerns are resolved before the end.

Step Do this Look for Do not conclude from
1 Set the score aside as proof What the number measures; text length; genre A percentage on its own
2 Read the flagged passages in context Whether they sit in the argument or in generic, conventional parts Highlights alone, with no reasons given
3 Compare with the student’s earlier work A sharp, unexplained change in voice, accuracy or structure Improvement, tutoring, or a new topic
4 Check the citations Sources that do not exist or do not say what the essay claims Real sources, which prove nothing about authorship
5 Ask the student about their reasoning Can they explain choices, terms and sources in their own words Nerves in a meeting
6 Review drafting and document history A drafting trail, or a whole essay arriving in one paste A missing history, which can have technical causes
7 Follow school policy, then write down the basis Whether independent evidence supports a concern under the stated rules Any single signal

Step 5 sits before step 6 here because a conversation often makes the history check unnecessary, and a history check read without the student’s account is easy to misread. If your policy or the situation calls for it, swap them.

How do teachers detect AI writing in practice?

Teachers use a set of signals, and the useful question is how much weight each deserves. A detector is only one of them, and not the strongest. The table sorts the signals by what they can show and where each one misleads.

Signal What it can show Main weakness Weight
Detector result Which passages resemble machine-written text Returns labels, not reasons; makes both kinds of error Pointer only
Style against earlier work An unexplained break from the student’s usual voice People improve, get help and change registers Moderate
Revision history How and when the text was produced Drafting elsewhere and pasting leaves one large edit; history can be incomplete Moderate to strong
Citations Sources that cannot be found or do not support the claim Real sources can be supplied to an assistant Strong when found
Oral follow-up Whether the student understands what they submitted Anxiety, memory and language fluency affect a live conversation Strong, with care
Assignment context Whether the text answers this prompt, these readings A student who skipped the readings looks the same Moderate
Sentence-level review Generic filler, vague claims, uneven register Every pattern has an ordinary human explanation Weak

Two points on the evidence. First, reading alone is unreliable: Fleckenstein and colleagues, in Computers and Education: Artificial Intelligence (June 2024), found that novice (N = 89) and experienced teachers (N = 200) could not identify texts generated by ChatGPT among student-written texts, that more experienced teachers made somewhat more differentiated and accurate judgments, and that both groups were overconfident (Do teachers spot AI?). A feeling that something is off is a reason to check, not a result.

Second, the signals only help in combination. The guide to how to tell if an essay was written by AI goes through each style signal and its ordinary human explanation in more detail. This guide is about what to do after the flag arrives.

Step 1: Why shouldn’t the score be treated as proof?

Because it does not measure authorship. On this site the headline figure is the share of the text made up of passages labeled AI-like. It is not a probability that the student used AI, and what an AI detection percentage means explains why the two get confused.

Other detectors define their numbers differently, so do not compare scores across tools. Turnitin’s documentation says its model may not always be accurate and should not be the sole basis for adverse action against a student. Its guidance on sentence-level results says to treat highlighted sentences as areas of interest and use them to start a conversation, not to draw a conclusion; it also reports a sentence-level false positive rate of around 4%, with 54% of such sentences sitting next to genuinely AI-written text (Turnitin). That figure is Turnitin’s own and dates from when it published the post.

Scale also matters. Vanderbilt noted in 2023 that a 1% false positive rate, applied to the 75,000 papers it submitted in 2022, would mean roughly 750 papers wrongly labeled, and it disabled the detector (Vanderbilt). The evidence on false positives, including the Stanford study of non-native English writers, is in AI detection false positives in student writing.

Before moving on, check three things: the length of the text (very short work gives a detector little to judge), the genre (lab reports and templated assignments push everyone toward the same structure) and whether the student writes in a second language. Each lowers how much the flag should count.

Step 2: How do you review the flagged passages?

Read the highlighted sentences yourself, in context. The detector behind our checker, from Pangram Labs, labels passages AI-Generated, AI-Assisted or Human Written with a confidence level. It returns labels and no reasons, so the highlight tells you where to look and nothing about why.

Useful questions when you read:

  • Do the flagged passages carry the argument, or are they introductions, transitions and summaries, which are often the most conventional parts of any essay?
  • Is the flag one long continuous block or a scatter of sentences? A scatter in an otherwise individual essay points somewhere different from a whole section.
  • Does the label say AI-Assisted? That is a different claim from AI-Generated, and whether it matters depends on what your assignment allowed.
  • Are there quotations, lists, references or code that the detector may have treated as prose?

Write down which passages you flagged and what you noticed. Step 3 is easier with a short list.

Step 3: How do you compare the work with the student’s earlier writing?

You are looking for a break, not a difference. Put the flagged passages next to an earlier assignment, an in-class piece or a discussion post, and compare things you can name: sentence length, vocabulary, the kind of mistakes the student usually makes, how they handle evidence.

A sharp, unexplained shift is worth asking about. But people also improve between assignments, write differently about topics they care about, and use tutors, writing centers and grammar tools. Turnitin recommends collecting a diagnostic writing sample at the start of a course as a baseline (Turnitin), which makes this comparison fairer when you have not met the student’s writing before. If you have no baseline at all, say so in your notes. A missing comparison is a gap in your evidence, not a signal.

Step 4: How do you check the citations?

This is one of the few checks that produces something verifiable. Search each reference by title in a library database or Google Scholar, confirm that the authors, year and journal match, and open the two or three sources that carry the argument to see whether they support what the essay attributes to them. Vanderbilt’s guidance points at the same thing: AI tools can invent whole sources.

A source that does not exist is a concrete problem whoever wrote the essay. It might come from a chat assistant, a misremembered reference or a copied bibliography, so ask before concluding which. A bibliography where everything checks out says nothing about authorship, because a student can supply real sources to an assistant. How to tell if an essay was written by AI covers fabricated references in more detail.

Step 5: How do you ask the student about their reasoning?

Ask early, privately and without a conclusion. A student who wrote the work can usually explain its choices, and the conversation gives you the one source of evidence that a detector and a document cannot: the student’s own account. Stony Brook University’s guidance for faculty puts a meeting with the student after a review of the report and course expectations, and suggests asking about tools such as translation apps and offering to go through the revision history together (Stony Brook CELT).

Treat the answers as information, not a test. A student may be anxious, may not have used these words for what they did, or may be writing in their second language. Fluent explanation of the content is reassuring; inability to explain a paragraph or define a term used in it is something to follow up, not a finding.

A script of neutral questions

Open with the purpose, then ask questions that a student who did the work can answer without preparation. Do not say “the detector says you cheated” or “I know you used AI.”

“I’m reading your essay again and I’d like to understand how you put it together. Nothing has been decided. Could you walk me through it?”

Then pick from:

  1. How did you start? What did you do first?
  2. What sources did you use, and how did you find them?
  3. Can you tell me about this paragraph in your own words? What point were you making?
  4. Why did you choose this example, or this structure, for the argument?
  5. What does this term mean, and where did you meet it?
  6. Did you use any tools while writing, for example translation, grammar checking, spell checking or an AI assistant? How?
  7. Do you have drafts, notes or an outline you could show me? Could we look at the document history together?
  8. Is there anything about how you wrote this that I should know?

Close by saying what happens next (“I’ll look at the drafts and decide whether anything else is needed, and I’ll tell you the outcome”) and when. Do not ask the student to prove a negative, and do not pressure them to confess in the room. If your policy allows a support person or a second staff member in the meeting, offer that.

Step 6: What does the drafting history show?

How an essay came into being is often stronger evidence than how it reads. In Google Docs, version history shows who updated a file and when, and it requires permission to edit the file; Google notes that some changes might not show up and that revisions may occasionally be merged (Google Docs Editors Help). Browser extensions that play back a document’s edits exist, but they depend on the same revision data, so treat their output as a starting point for a question too.

What to look for:

  • A drafting trail. An outline, rough early paragraphs and revision over several sessions is reassuring.
  • One large paste. A whole essay appearing in a single edit is a reason to ask where it came from. Some people draft in a notes app or on paper and paste in the result, so ask before drawing a conclusion.
  • Missing history. A file that was exported, converted or started from a template may have no useful history, and an absent trail does not show misconduct.

If you want this evidence to exist next time, say so in the assignment. Stony Brook’s guidance advises instructors to tell students to keep time-stamped copies of their drafts.

Step 7: How do you follow school policy and record the decision?

Check the rules that applied to this assignment before you decide what a finding means. Whether AI-assisted editing was allowed, whether disclosure was required and which office handles concerns all differ between schools and between courses, and this guide cannot tell you what yours say. A discussion that finds the student used a grammar tool the assignment allowed is a different case from one that finds an essay generated wholesale where AI was banned.

Guidance is often thin. In the Center for Democracy and Technology’s survey of 460 public school teachers in grades 6 to 12, fielded in November and December 2023, 68% reported using AI content detection tools in 2023-24, up from 38% the year before, while only about a third had received guidance on responding to suspected policy violations (K-12 Dive). If your school has no written process, ask your department head or integrity office before you act, and keep your own record.

Stony Brook’s guidance lists three outcomes after review: a teaching intervention, a formal warning or an official academic integrity report, and it asks faculty to document observations, meeting summaries and supporting materials such as drafts. A useful record is short:

  • The assignment, its AI policy and the date of the review
  • What the detector showed, including word count and which passages were flagged
  • What you compared and what you found (earlier work, citations, history)
  • What the student said, in their words
  • The decision, who made it, and how the student can respond or appeal

What if the evidence does not settle it?

Say so, and choose a proportionate response. Several independent signals pointing the same way support a concern. A detector flag with nothing else behind it does not. In that case the fair outcome may be a conversation about expectations, an invitation to talk through the essay or a short follow-up task, with no finding recorded.

A follow-up written in class gives you another sample, but weigh it carefully: writing under supervision and time pressure differs from writing at home, so it is a weaker comparison than a student’s usual work. Vanderbilt’s alternative to detection is to set expectations early, ask for disclosure when AI is used and design assignments around in-class writing and course-specific material, all of which make later review easier.

What if you are the student whose work was flagged?

Keep your outline, notes, drafts and sources, and be ready to describe how you worked. You can ask what evidence the concern rests on beyond the tool, and you can point out that detectors are documented to produce false positives, particularly on short, formal and non-native-English writing. Can AI detectors be wrong explains how those errors happen. This guide is for reviewing evidence fairly, not for evading a review.

Frequently asked questions

Is a high AI percentage enough to fail a student? No. A percentage is not proof of authorship, and the vendors’ own guidance says detection results should not be the sole basis for adverse action. A grade or penalty should rest on evidence that holds up without the score, applied under your school’s stated policy.

Should I run the essay through a second detector? It can show whether the same passages stand out, but two tools can agree and both be wrong, and their numbers are not on the same scale. A second flag adds little compared with citations, history or the student’s explanation.

Can I tell by the writing alone? Not reliably. The 2024 study of novice and experienced teachers found they could not pick out ChatGPT essays from student essays and were overconfident. Use reading to decide what to check, not to decide the outcome.

What if the student admits to using AI? Check the assignment’s rules first. The admission may describe permitted help, such as grammar checking. Where a rule was broken, follow your school’s process and record what the student said.

Where does a detector fit in this workflow?

At step 2, as a way to find passages worth reading closely. If you already have the text, run it through our AI-generated text detector and review the flagged passages rather than relying on the overall score. The page for teachers describes the checker’s limits and what happens to submitted text: it is sent to Pangram Labs for analysis, so check your school’s data policy before submitting student work. If you are choosing between tools, which AI detector do teachers use compares the main ones with check dates.