An image contains pixel values and, sometimes, information about the file’s creation and processing. Neither the pixels nor the filename inherently declares whether the depicted scene passed through a camera sensor. AI-generated image detection therefore requires a proposition about origin, a measurement that bears on it and evidence that the measurement discriminates under conditions relevant to the questioned file.
The simplest proposition—“AI-generated or not”—is often too broad. A camera may photograph a synthetic picture on a screen. A photograph may contain a generated region. An entirely generated image may be recompressed, cropped or made into a screenshot. These have different formation histories, although each could be described casually as an “AI image”. Before applying a method, an examiner needs to say whether the issue is whole-image origin, a local alteration, provenance of a particular file or the truth of the depicted event.
Records of creation and their gaps
Embedded metadata can record a camera model, software or processing time. It can also be edited, copied, omitted at export or replaced when another application saves the file. A camera-model field is thus an observation about the file’s recorded history, not independent proof of exposure on that camera. Source files, acquisition records and a documented transfer chain can make that observation more informative.
C2PA Content Credentials use signed manifests and content bindings to make particular provenance claims verifiable. A valid manifest can identify what its signer asserted about an asset and its ingredients. The verifier still has to consider the signer, the claim and any break in the chain. A file without a credential is not thereby synthetic; many capture and distribution paths do not produce or preserve one. The provenance article examines that distinction in detail.
What the pixels can contribute
Visible anomalies may suggest an area for closer examination, especially where geometry, lighting or repeated structures conflict with an asserted capture process. Their absence is weak reassurance. Generative methods and ordinary image editing can alter such features, while camera optics, compression and display recapture can introduce unusual appearances in genuine captures. Visual inspection is useful for framing a question; it is a poor substitute for a validated origin method.
Statistical classifiers learn distinctions between labelled camera and generated images. Their output is a model score or a thresholded class, conditional on their training and test populations. A detector developed around particular generators may respond differently to an unfamiliar generator or to altered file processing. The original CNN-image detection experiment investigated transfer to generators outside its training set; later diffusion reconstruction research used a different signal. These are different measurements, not interchangeable guarantees. For a questioned image, test composition, threshold and error rates matter more than an unqualified accuracy figure.
Physical image formation offers another evidence source. A camera’s optics, sensor and processing pipeline affect recorded pixels. Classical sensor pattern noise work has used an estimated camera fingerprint to test whether a photograph is associated with a particular device. That source-camera identification method has its own reference-image and processing requirements. It should not be conflated with every sensor-based camera-versus-synthetic test.
SPOT—Sensor Pattern Origin Testing addresses the latter question through noise-residue measurements evaluated against modelled camera and synthetic-image populations. Dr Richard Matthews’ published experiment used Dresden camera imagery and DALL·E-generated imagery; its reported discrimination belongs to those experimental conditions. The resulting likelihood ratio expresses how strongly the measured characteristics support the specified origin hypotheses under the models. It is not the probability that the image is authentic, and it does not identify a particular camera, creator or depicted event.
Reliability belongs to a specified use
A method can discriminate well in a controlled dataset yet have unknown performance on screenshots, unfamiliar camera pipelines, mixed-origin images or newly generated imagery. Compression and resizing may change the signal available for examination. Population composition and a chosen decision threshold affect both false positives and false negatives. The detector-reliability analysis explains why sensitivity, specificity and calibration should be reported for the proposed use.
Reliable detection is therefore a conditional claim: a defined method applied to suitable material, with relevant validation and a conclusion limited to what the observations support. If the available file lacks the necessary provenance or measurable signal, “undetermined” may be the scientifically useful result. Further source acquisition may answer more than another classifier run.