“99% accurate” does not specify what happened to genuine photographs, what happened to generated images, or how many of each were in the test. In a dataset containing 99 photographs for every generated image, a rule that calls everything a photograph is 99% accurate while detecting nothing. A forensic assessment needs the confusion matrix and the conditions that produced it.
Suppose “positive” means synthetic. Sensitivity is the proportion of synthetic images correctly labelled synthetic; its complement is the false-negative rate. Specificity is the proportion of camera images correctly labelled camera; its complement is the false-positive rate. Which error matters more depends on the decision that follows, but both must be visible. An accusation based on a false positive can have a very different consequence from a missed synthetic image in triage.
The threshold changes the error balance
Many detectors output a continuous score. A classification threshold maps that score to a label. Lowering the threshold usually catches more positives while also admitting more false positives; raising it usually does the reverse. A receiver operating characteristic (ROC) curve plots true-positive rate against false-positive rate across thresholds. Its area under the curve (AUC) summarises ranking discrimination across the tested positive and negative populations. AUC does not choose an operational threshold, give a probability of synthetic origin, or guarantee the false-positive rate on a new population. NIST’s image-detection evaluation includes AUC alongside true-positive rate at a specified false-positive rate and a calibration measure for this reason.
Calibration asks a separate question: among cases assigned a stated probability or confidence, does the stated proportion correspond to observed outcomes in an applicable population? A classifier’s raw score should not be read as “90% probability of AI generation” unless the system was designed and validated to support that interpretation. Even a calibrated estimate can change when the population or prior prevalence changes.
A base-rate example
Consider 10,000 images, 100 of which are synthetic. A detector has 90% sensitivity and 99% specificity on this population. It flags 90 synthetic images and misses 10; it also flags 99 of the 9,900 camera images. Thus 90 of the 189 flagged images are synthetic—about 48%. The same sensitivity and specificity applied where half the images are synthetic would yield a very different fraction among flags. This is a hypothetical arithmetic example, not a performance claim about any product or about SPOT.
| Measure in this example | Count |
|---|---|
| Synthetic correctly flagged | 90 |
| Synthetic missed | 10 |
| Camera images falsely flagged | 99 |
| Camera images correctly cleared | 9,801 |
An examiner should therefore avoid converting a detector’s test accuracy into the probability that a particular flagged image is synthetic. That would require an appropriate prior probability, a justified mapping from measurement to likelihood and confidence that the model still applies to the questioned material.
What a validation set must resemble
A test split from the same pool as training may assess repeatability on similar images while saying little about an unseen generator, camera pipeline or distribution platform. Generator families change; resizing, screenshots and JPEG recompression can alter the features a detector uses. Dataset selection can also create shortcuts, such as different resolution or compression histories between camera and generated classes. A model can then learn the collection process rather than the intended origin distinction. Research on CNN-generated image detection and diffusion reconstruction error illustrates how methods define and test different signals; neither result can be transferred without examining its test conditions.
RHEM Labs’ published SPOT experiment reports a mean ROC AUC of 0.998 and mean accuracy of 0.985 in its Results section across 100 random runs comparing defined Dresden camera and DALL·E populations. The paper’s abstract reports a different accuracy figure; the SPOT evidence page keeps the figures separate. Those numbers demonstrate discrimination in that experiment. They do not by themselves establish error rates for unknown generators, transformed images or a particular forensic case. SPOT’s likelihood-ratio framework asks for relative support under specified models; the models and their applicability remain part of the opinion.
For a proposed use, request class counts, thresholds, relevant false-positive and false-negative rates, calibration if probabilistic language is used, and testing on independently acquired material with realistic transformations. When the questioned file falls outside those conditions, a qualified or inconclusive result is more informative than an apparently precise label.