FORENSIC SCIENCE · INDEPENDENT EXPERT EVIDENCEAdelaide, Australia
RHEM LabsInstruct an expert

Synthetic-media forensics

Are AI Image Detectors Reliable? Why Accuracy Alone Is Not Enough

Detector suitability depends on error rates, threshold, calibration, prevalence and how closely validation resembles the material being examined.

RHEM Labs

“99% accurate” does not specify what happened to genuine photographs, what happened to generated images, or how many of each were in the test. In a dataset containing 99 photographs for every generated image, a rule that calls everything a photograph is 99% accurate while detecting nothing. A forensic assessment needs the confusion matrix and the conditions that produced it.

Suppose “positive” means synthetic. Sensitivity is the proportion of synthetic images correctly labelled synthetic; its complement is the false-negative rate. Specificity is the proportion of camera images correctly labelled camera; its complement is the false-positive rate. Which error matters more depends on the decision that follows, but both must be visible. An accusation based on a false positive can have a very different consequence from a missed synthetic image in triage.

The threshold changes the error balance

Many detectors output a continuous score. A classification threshold maps that score to a label. Lowering the threshold usually catches more positives while also admitting more false positives; raising it usually does the reverse. A receiver operating characteristic (ROC) curve plots true-positive rate against false-positive rate across thresholds. Its area under the curve (AUC) summarises ranking discrimination across the tested positive and negative populations. AUC does not choose an operational threshold, give a probability of synthetic origin, or guarantee the false-positive rate on a new population. NIST’s image-detection evaluation includes AUC alongside true-positive rate at a specified false-positive rate and a calibration measure for this reason.

Calibration asks a separate question: among cases assigned a stated probability or confidence, does the stated proportion correspond to observed outcomes in an applicable population? A classifier’s raw score should not be read as “90% probability of AI generation” unless the system was designed and validated to support that interpretation. Even a calibrated estimate can change when the population or prior prevalence changes.

A base-rate example

Consider 10,000 images, 100 of which are synthetic. A detector has 90% sensitivity and 99% specificity on this population. It flags 90 synthetic images and misses 10; it also flags 99 of the 9,900 camera images. Thus 90 of the 189 flagged images are synthetic—about 48%. The same sensitivity and specificity applied where half the images are synthetic would yield a very different fraction among flags. This is a hypothetical arithmetic example, not a performance claim about any product or about SPOT.

Measure in this example Count
Synthetic correctly flagged 90
Synthetic missed 10
Camera images falsely flagged 99
Camera images correctly cleared 9,801

An examiner should therefore avoid converting a detector’s test accuracy into the probability that a particular flagged image is synthetic. That would require an appropriate prior probability, a justified mapping from measurement to likelihood and confidence that the model still applies to the questioned material.

What a validation set must resemble

A test split from the same pool as training may assess repeatability on similar images while saying little about an unseen generator, camera pipeline or distribution platform. Generator families change; resizing, screenshots and JPEG recompression can alter the features a detector uses. Dataset selection can also create shortcuts, such as different resolution or compression histories between camera and generated classes. A model can then learn the collection process rather than the intended origin distinction. Research on CNN-generated image detection and diffusion reconstruction error illustrates how methods define and test different signals; neither result can be transferred without examining its test conditions.

RHEM Labs’ published SPOT experiment reports a mean ROC AUC of 0.998 and mean accuracy of 0.985 in its Results section across 100 random runs comparing defined Dresden camera and DALL·E populations. The paper’s abstract reports a different accuracy figure; the SPOT evidence page keeps the figures separate. Those numbers demonstrate discrimination in that experiment. They do not by themselves establish error rates for unknown generators, transformed images or a particular forensic case. SPOT’s likelihood-ratio framework asks for relative support under specified models; the models and their applicability remain part of the opinion.

For a proposed use, request class counts, thresholds, relevant false-positive and false-negative rates, calibration if probabilistic language is used, and testing on independently acquired material with realistic transformations. When the questioned file falls outside those conditions, a qualified or inconclusive result is more informative than an apparently precise label.

Related research & reading

FORENSIC INSTRUCTIONS

A question about digital evidence?

For legal practitioners and professional organisations: discuss the forensic question, available material and timeframes.

Contact the laboratory