AI image detectors are far less dependable than their headline accuracy figures suggest. Independent benchmarks keep finding the same pattern: a detector that looks excellent on images from generators it was trained on can fall sharply on images from newer generators, on images that have been compressed or resized, and on photos that were only partly edited. A detector’s score is a statistical guess about an image’s pixels. It is not proof of how the image was made.
This site checks text only, so it has no image tool and none is linked here. What follows is a plain account of what the published evidence says, so you can judge any image detector’s output yourself.
Key takeaways
- Accuracy figures only mean something with their test conditions: which generators, which images, what compression and where the threshold sat.
- In a 2026 zero-shot benchmark of 23 open-source detector variants, the best averaged 75.0% accuracy and the worst 37.5%. On newer commercial generators, average accuracy was 18 to 30%.
- Re-saving an image as a JPEG can erase much of what many detectors rely on.
- A high score is a signal, a low score is not clearance, and neither is proof.
What does “accuracy” mean for an image detector?
Accuracy is the share of test images a detector labels correctly. On its own it hides the two errors that matter: calling a real photo AI-made (a false positive) and missing an AI image (a false negative). A paper reporting 99% accuracy on a clean test set says nothing about either error on the images you will actually meet.
Some benchmarks report fake-image accuracy and real-image accuracy separately. That split matters. In one benchmark, a technique that improved how well detectors recognised real images left their ability to catch fakes essentially unchanged. A single combined percentage would have hidden that.
The same concept applies to text, and how accurate AI detectors are covers the general ideas of error types, thresholds and why a single number misleads. Images add one more problem: the file a detector sees is rarely the file that was generated.
What do independent benchmarks actually find?
The figures below come from academic papers. Each is quoted with the condition it was measured under, because the number is meaningless without it. None of them rate a particular product you might use today.
| Source | Test conditions | Reported result |
|---|---|---|
| Ren et al., 2026 (arXiv 2602.07814) | Zero-shot evaluation (no fine-tuning) of 16 methods, 23 pretrained variants, on 12 datasets of about 2.6 million images from 291 generators | Best detector 75.0% mean accuracy, worst 37.5%; on Flux Dev, Firefly v4 and Midjourney v7, only 18 to 30% average accuracy |
| Li et al., NeurIPS 2025 (AIGIBench) | 11 detectors, 23 fake-image subsets, plus real images from social media and AI art platforms | Detectors “suffer significant performance drops on real-world data” despite high accuracy in controlled settings |
| Chandra et al., 2025 (Deepfake-Eval-2024) | 1,975 in-the-wild images (plus video and audio) from social media and detection-platform users in 2024; open-source models | AUC for image models fell 45% compared with previous benchmarks; commercial and fine-tuned models did better but “do not yet reach the accuracy of deepfake forensic analysts” |
| Nebioglu et al., 2026 (Inpainting Exchange) | 90,000-image dataset; detectors tested on inpainted images, then with the original pixels swapped back in around the generated content | Accuracy fell from 91% to 55% once side effects were separated from the generated content |
The Ren et al. paper also found that detector rankings were unstable: which detector looked best depended heavily on which dataset was used, and the match between a detector’s training data and the test generators changed performance by 20 to 60% within detectors that share the same architecture. “AUC” in the Deepfake-Eval row is a ranking measure, not a percent-correct figure, so the 45% is a relative drop in that measure and not “45% of images misjudged”.
Why does accuracy collapse on new generators?
Most detectors learn to spot the traces left by the generators in their training data. When a new generator leaves different traces, or fewer, the learned pattern no longer fits. The Ren et al. result of 18 to 30% on three recent commercial generators shows how large the gap can be. Their finding that training-data alignment moved performance by 20 to 60% among otherwise identical detectors points the same way: what the detector was shown matters at least as much as how it is built.
This is also why a figure from an older paper is a weak guide to today’s tools. Generators change quickly, and a benchmark is a photograph of one moment. A vendor can retrain on new generators, which helps, but the next release restarts the problem. Detectors that perform well on a given generator are doing so against a target that has already been seen.
Why do compression and screenshots break detectors?
Many detectors depend on very fine, high-frequency patterns in the pixels. Lossy compression such as JPEG discards exactly this kind of detail to make files smaller. Resizing smooths it further. Social platforms re-encode and shrink uploads, and a screenshot is a fresh re-render of what was on screen.
The evidence for compression is strong. In AIGIBench, JPEG compression at quality 50, averaged across 25 test sets, made fake-image detection accuracy fall “often approaching 0%” for the detectors tested; Gaussian noise and other degradations also hurt, while up-and-down resampling was the milder case. The ITW-SM paper by Konstantinidou et al. makes the related point that resizing tends to erase the subtle traces detectors look for, and that images circulating online are often several megapixels and compressed.
One honest limit: I found no benchmark that tests screenshots specifically. ITW-SM actually removed screenshots and memes from its test set. The screenshot case is an inference from the compression and resizing results, not a measured figure. Treat a detector’s output on a screenshot, a forwarded chat image or a heavily recompressed file as especially weak.
Why are edited images hard?
Most benchmarks use fully generated images. Real use is messier: a real photo with one object replaced, a generated image touched up in a photo editor, a face swapped in. Detectors built for whole-image generation often handle these poorly.
Nebioglu et al. found that detectors may lean on side effects of the editing process rather than on the generated content itself. When the side effects were separated from the generated region, accuracy dropped from 91% to 55%, frequently near chance level. Pandolfini et al. found a more nuanced picture: detectors trained on many generators transferred partly to inpainting and could catch medium and large edits, with weaker results on small edited regions.
So a clean result on an image with a small edit says little, and an AI-flagged result on a real photograph that was only retouched says little too.
Why is a score not proof?
A detector outputs a number or label, and the number reflects how closely the pixels resemble patterns in its training data. It does not carry the image’s history. Three things follow from that.
First, the score depends on a threshold someone chose. Moving it trades false positives for false negatives, and a vendor’s published error rates apply to its chosen setting and its own test set.
Second, the score is not a probability that the image is AI. A “92% AI” label is not “92% likely to be AI” unless the vendor has shown that its scores are calibrated on images like yours.
Third, vendor figures are claims, not independent measurements. For example, Hive’s blog reports 98.03% accuracy, a 0% false positive rate and a 3.17% false negative rate from a February 2024 University of Chicago study. The study’s test set was a mix of human art, AI images, hybrid images and perturbed art, and the blog post does not name the generators used. That is a useful data point about one product on one set from 2024, not a guarantee for a photo made by last month’s model.
The researchers behind the underlying study also compared detectors with people. They found that the best automated detector and expert artists both performed well but made mistakes in different ways, and that combining human and automated judgment worked better than either alone. That is a result about art images, but the lesson carries over: use a detector as one input.
How should you read an image-detector result?
| What you see | What it can mean | What it cannot tell you | What to do next |
|---|---|---|---|
| High “AI-generated” score | The pixels resemble a generator the tool knows | That the image is fake, or which tool made it | Look for provenance, the original source and context |
| Low score, “likely real” | No known pattern found | That the image is real; new generators and compression defeat many detectors | Do not treat it as clearance |
| Middle or mixed score | The tool is unsure | Anything useful on its own | Seek other evidence |
| Named generator (“probably Midjourney”) | The tool’s best guess at a source | A confirmed origin | Check the claim against provenance data |
| Result on a screenshot or compressed copy | Detector is working on degraded evidence | Much at all | Find a higher-quality original |
| “Region edited” or partial result | A possible local edit | The extent or purpose of the edit | Compare against other versions of the image |
| A vendor accuracy figure | A result under the vendor’s conditions | Performance on your image | Ask for generators, dates, image types and thresholds |
What stronger evidence is there?
Evidence that comes from outside the pixels is more reliable than a pixel-based guess. Provenance data attached by the creating tool or camera, such as C2PA content credentials, can show how a file was made, but they are opt-in and can be removed. The C2PA specification itself notes that an asset can become separated from its manifest through removal or corruption of metadata, so a missing credential tells you nothing either way. Watermarks such as Google’s SynthID work only for content generated by participating tools; Google describes the watermarks as embedded in its own generative products and designed to survive edits like cropping, filters and lossy compression, and its detector looks for that watermark rather than judging whether any image is AI.
The sourcing questions are the ordinary ones. Where did it first appear? Can the original, uncompressed file be found? Does reverse-image search show an earlier version or an independent source? Is there a second angle, a witness, or consistent coverage? How to tell if an image is AI-generated walks through these checks, and what metadata can and cannot show about AI images explains why missing or stripped metadata proves nothing. Visual tells that circulated a year or two ago, such as odd hands, are not universal rules; current generators often avoid them, and real photos can show odd artifacts too.
No single method is conclusive. Where it matters, such as journalism, legal disputes or fraud, forensic analysts combine several methods, and even then the Deepfake-Eval-2024 authors found that detection models fell short of the accuracy of human forensic analysts.
What does this site check?
This site checks text only. It has no image, video or audio checker, and nothing here should be read as a way to check pictures. The idea carries over, though: a detection result is a statistical signal that needs context. For text, the accuracy page explains why this site publishes no accuracy number, and what an AI detection percentage means explains how to read the score it does return. If you already have text to review, run it through our AI-generated text detector and read the flagged passages rather than relying on the overall score.
FAQ
Can an AI image detector prove an image is fake? No. It reports how closely an image resembles patterns it learned. A high score is a reason to look harder, and provenance, the original source and context carry more weight than the score.
Are AI image detectors more accurate than people? It depends on the test. In a University of Chicago study of art images, the best automated detector and expert artists both did well but erred differently, and combining them worked better. In Deepfake-Eval-2024, even commercial models did not reach forensic analysts’ accuracy.
Why does a detector get worse on social media images? Platforms compress and resize uploads, which removes fine pixel traces many detectors depend on. AIGIBench found that JPEG compression at quality 50 pushed fake-image detection accuracy toward zero for many of the detectors it tested.
Does a “real” result mean the image is real? No. New generators, compression and editing can all make an AI image look real to a detector. Treat a low score as the absence of a signal, not as clearance.