AI video detection software splits a clip into frames and audio, runs several separate checks on them, and combines the outputs into a confidence score. The checks look at visual artefacts in single frames, consistency of motion across frames, face and mouth behaviour, audio, and any provenance data attached to the file. The result is a probability-like number, not a ruling. Detectors are strongest on the kinds of fakes they were trained on and weakest on new generators, heavy compression and very short clips.
This site checks text only, so everything below is background on a different medium. It is meant to help you read a video verdict, from any tool, with the right level of trust.
Key takeaways
- A video detector is usually several detectors plus a step that combines their scores. No single signal is decisive.
- Provenance (content credentials, watermarks) is the only signal that can be checked cryptographically, but most clips arrive without it, so a missing credential proves nothing.
- Published detectors tend to lose a lot of accuracy on videos unlike their training data. Two benchmarks below show how large the drop can be.
- A score is a confidence about a pattern in the file. It is not proof of how a video was made.
What are the steps in a video detection pipeline?
The pipeline below is a generic picture assembled from published research. Individual products differ, and most vendors do not publish their exact design.
video file
|
+--> 1. decode: frames + audio track + metadata + any credentials
|
+--> 2. frames: sample, find and crop faces
| +--> 3a. per-frame artefact model --> score per frame
| +--> 3b. motion / temporal model --> score per clip
| +--> 3c. face and lip-sync model --> score per clip
|
+--> 4. audio model (speech / voice cloning) --> score per clip
|
+--> 5. provenance check (C2PA, watermark) --> present / absent / invalid
|
+--> 6. combine scores, calibrate --> one confidence + notes
Read it left to right and top to bottom. The input is one file. The file is taken apart into streams, each stream goes through its own model, and a final step merges the outputs. Steps 3 to 5 run independently of each other, which is why a clip can show a strong signal in one place and nothing in another.
How does frame extraction work?
A video is a sequence of still images plus sound. Detectors decode the file and pull out frames, because most image models take one frame at a time. Analysing every frame is expensive, so many systems sample: a fixed number of frames per clip, or every few frames.
Many methods then locate faces in each sampled frame and crop them, so the model sees the face region rather than the whole scene. An early temporal approach by Güera and Delp used a convolutional network to extract features from each frame and fed them to a recurrent network; they reported over 97% accuracy using only 40 frames per video on their own test set from 2018. That figure describes that dataset and those fakes, not detection in general.
One practical point follows from this. Frames are extracted from a file that was almost always compressed already, often more than once by upload and re-upload. Whatever fine detail a detector wants was partly destroyed before it ever saw the video.
What is per-frame artefact analysis?
This step asks whether a single image looks like it came from a camera or from a generative process. A classifier, usually a convolutional or transformer network, is trained on large sets of real and synthetic frames and learns statistical differences too subtle to see by eye.
Early face-swap methods left visible traces of how they worked. Li and Lyu observed that deepfake algorithms generate faces at limited resolution and then warp them to fit the original face, leaving distinctive warping artefacts a network can pick up. That finding fits face swaps from that period. It does not describe videos generated whole by modern text-to-video models, which have no source video to warp onto.
This is also why advice such as “check the hands” or “look for odd teeth” ages badly. A tell that works on one generation of tools can vanish in the next, and a classifier trained on the old tell inherits the problem. Our guide to telling whether a video is AI-generated covers what a person can still check by eye and context.
How do temporal and motion checks work?
A frame can look fine while the sequence is wrong. Temporal models look at how features change from one frame to the next: whether lighting, texture, edges and identity stay stable, and whether movement follows plausible physics. Güera and Delp’s recurrent network was built on the idea that manipulated videos contain temporal inconsistencies between frames that a sequence model can learn.
Typical things such models can pick up, depending on the method, are flicker around a swapped face, a face region that drifts relative to the head, or texture that shimmers while the background stays still. How well any of this works depends on which generator made the clip. Different tools fail in different ways, and a model that learned one tool’s temporal habits may find nothing odd in another’s.
How do face and lip-sync analysis work?
Faces are where most deepfake research has focused, because face swaps and lip-sync edits are the most common manipulations of real people. Several distinct ideas fall under this heading.
Mouth movement. LipForensics, from Haliassos and colleagues (CVPR 2021), first trains a network to lip-read, so it learns what natural mouth motion looks like, and then trains a second network on those mouth representations to separate real from forged video. The authors argue that targeting high-level irregularities in mouth movement, rather than low-level artefacts, helps it generalise to manipulations it has not seen and survive distortions such as compression.
Person-specific behaviour. Agarwal and Farid built “soft biometric” models of how particular public figures move their faces and heads while speaking. They reported 92 to 96% accuracy depending on the leader and the length of the video. The catch is built in: it needs enough real footage of the specific person, so it suits famous people, not an arbitrary stranger.
Biological signals. Ciftci, Demir and Yin proposed analysing the faint colour changes in facial skin caused by blood flow (photoplethysmography) as a cue that fake faces may not reproduce consistently. The idea is interesting and published, but it was developed against the face-swap methods of its time, and a signal that is weak in a compressed phone video is hard to measure.
Audio-visual sync. Some systems check whether mouth shape matches the sounds in the track. This overlaps with the audio analysis below and is only available when the clip has speech and a visible mouth.
How does audio analysis work?
The audio track is analysed on its own, as speech-spoofing detection. Models are trained to separate genuine recordings from synthetic or cloned voices, often working on spectrogram-style representations of the sound. The research community runs shared challenges for this; the fifth edition, ASVspoof 5, drew submissions from 53 teams.
Its organisers report that many solutions perform well, but performance degrades under adversarial attacks and under neural encoding and compression schemes. That is the same weakness seen for video: lab performance does not carry over cleanly to audio that has been re-encoded for a platform.
An audio check can only say something when there is audio. A silent AI-generated clip passes through this step with nothing to analyse, and a real video with an AI voiceover raises a different question than “is the video generated?” See deepfake vs AI-generated video for why those cases differ.
What are provenance and content credentials?
The signals above are inferences from pixels and sound. Provenance is different: it is recorded information about where a file came from.
Content credentials (C2PA). The C2PA standard defines a signed manifest that records an asset’s origin and edit history. According to the C2PA explainer, the manifest is bound to the asset with cryptographic hashes and a digital signature, so any change to the asset or the provenance data, however small, is detectable. If a generator or camera adds a credential, a verifier can read it and check the signature without guessing.
Two limits matter. First, credentials describe claims by a signer; the C2PA explainer states they do not provide value judgments about whether the provenance data is true, and that provenance alone cannot tell you whether content is accurate. Second, the metadata can be removed. C2PA’s answer is “durable” credentials that pair the signed manifest with soft binding, such as invisible watermarks or fingerprint lookup, to recover it. In practice, whether a platform keeps credentials after upload and re-encoding is a separate question from whether the standard allows it.
Invisible watermarks. Google DeepMind says SynthID embeds a watermark directly into the pixels of every frame of video generated by its Veo model, imperceptible to the eye but detectable by software. A watermark exists only if the generator put it there, and it points to that generator’s tool, not to AI video in general. Google describes SynthID itself as not a silver bullet.
So provenance answers one narrow question well: did this file come from a participating tool, and is the record intact? A positive result is strong evidence. A negative result is not evidence of authenticity, because most generators, and all real footage that was never signed, have no credential.
What can metadata tell you?
Metadata is the descriptive data stored in or beside a video: container fields, encoder names, creation times, device information. Detection pipelines read it as a minor signal. An encoder string or a missing camera field can be suggestive, but metadata is trivial to edit and is routinely rewritten when a platform re-encodes a clip. Treat unusual metadata as a reason to look further and clean metadata as meaning nothing. This is the weakest provenance layer, not to be confused with signed C2PA data.
How are the scores combined into one confidence?
Each model above outputs its own number: a probability that a frame, a motion pattern, a mouth, or a voice is synthetic. A final step combines them. Common designs include averaging frame scores over the clip, taking the highest score, weighting each model by how reliable it has been, or training a small extra model on the individual scores. Vendors rarely publish which they use, and I could not verify the method for any specific commercial product.
Combining helps because errors are not the same across signals. A clip that fools the frame model may still show odd mouth motion. But combining cannot create information that none of the signals contains, and when every signal was trained on the same kind of data, they tend to fail together on the same kind of unfamiliar video.
The last step is calibration: making a model’s raw output behave like a real probability. Guo et al. showed that modern neural networks are often poorly calibrated, meaning their stated confidence does not match how often they are right, and that a simple fix, temperature scaling, works well on many datasets. Calibration still depends on the test data: a score is only as meaningful as the match between the data it was tuned on and the clip in front of you.
This is why the output is a confidence and not a verdict. “87% likely synthetic” means the pattern of signals resembles the tool’s training examples of synthetic video to that degree. It does not mean 87 out of 100 viewers would agree, or that there is a 13% chance the clip is real. Our guide on accuracy explains why the same caution applies to text detectors.
Where does video detection fail?
The evidence below comes from published benchmarks. Each row says who produced the figure and under what conditions.
| Source | What was tested | What it reported | Read it as |
|---|---|---|---|
| Meta Deepfake Detection Challenge (2020) | 2,114 participants, over 100,000 videos, a hidden “black box” set of unseen fakes | Top model: 82.56% on the public set, 65.18% on the black box set; no entrant reached 70% on unseen deepfakes | Strong results on known data dropped sharply on unfamiliar fakes |
| Deepfake-Eval-2024 (Chandra et al., 2025) | 45 hours of video, 56.5 hours of audio and 1,975 images from 88 sites, 52 languages | Open-source detectors’ AUC fell 50% for video, 48% for audio and 45% for images versus earlier benchmarks | Academic benchmarks overstated real-world performance |
| Same paper | Commercial and fine-tuned models | Better than open-source models, but “do not yet reach the accuracy of deepfake forensic analysts” | Specialist human analysis still outperformed the tools the authors tested |
| FaceForensics++ (Rössler et al., 2019) | Face-manipulation video at several compression levels | Domain-specific methods worked “even in the presence of strong compression” and outperformed human observers | Compression hurts but does not end detection, on that benchmark’s manipulations |
Three causes recur.
New generators. The Meta challenge shows the pattern clearly: the same models that led on public data fell sharply on a hidden set built to resemble real-world surprise. Whole-video generators released after a detector was trained are by definition unseen. Look for the generator list and the date behind any claim, as with text detectors.
Compression. Platforms recompress uploads, which smooths away the fine statistical traces frame-level models rely on, and ASVspoof 5’s organisers report degradation after encoding for audio too. Methods that target higher-level cues, such as the mouth-motion approach above, were designed with this in mind, but robustness is a matter of degree.
Short and low-resolution clips. A few seconds gives few frames, limited motion to analyse and often no usable audio. Faces that are small in the frame produce poor crops. I did not find a primary source that quantifies accuracy by clip length, so treat this as a logical consequence of the pipeline rather than a measured figure.
How should you read a video detection result?
| If the result shows | Reasonable reading |
|---|---|
| Valid content credential naming a generator | Strong evidence of origin for that file; still read the claim and the signer. |
| Watermark detected by the vendor’s detector | Strong evidence the clip came from that vendor’s tool, only for that vendor. |
| High score, no provenance | A flag worth investigating: find the original upload, the account, the context, other copies. |
| Low score, no provenance | Unresolved. It may be real, or it may be a fake the tool has not seen. |
| Signals disagree (face high, audio low) | Ask what was likely edited: a face swap, a cloned voice, or neither. |
| Result on a short, compressed repost | Lowest trust. Look for a better copy. |
Detection scores are one input. Source, context, the original upload and corroborating footage carry as much weight, and no single method is conclusive. DeepfakeBench, a NeurIPS 2023 benchmark, exists because detectors had been tested with inconsistent data pipelines and settings, which made published numbers hard to compare. Be careful with any figure that does not state its test set.
What does this mean for text?
The logic carries over. Text detectors also combine signals, also lose accuracy on unfamiliar models and heavily edited input, and also return a confidence rather than proof. If you are working with written content, how AI detection works for text covers that side, and our methodology page lists what the checker on this site analyzes.
FAQ
Can AI video detectors prove a video is fake? No. They return a confidence based on patterns in the file. Only a verified content credential or a vendor watermark gives a firm statement of origin, and then only for what the signer or vendor claims. Treat everything else as evidence to weigh against context.
Why do detectors work in demos and fail on real clips? Demos usually use fakes similar to the training data and clean video. Real clips are compressed, trimmed, re-uploaded and made by newer tools. Benchmarks reflect this: in Meta’s challenge the top model scored 82.56% on public data and 65.18% on hidden unseen fakes.
Does a missing content credential mean a video is fake? No. Most real and synthetic videos carry no credential, and platforms can strip metadata when a file is uploaded and re-encoded. A present, valid credential is informative. An absent one is not.
Is face analysis enough to detect a fully AI-generated video? Not by itself. Many face methods were built for face swaps and lip-sync edits. A video generated end to end has no original face to splice onto, so detectors rely more on frame, motion and provenance signals, and results vary by generator.