This supporting page separates the published experiment, a further experimental comparison and the framework for future evaluation. For the implemented capability, anonymised applied examination and engagement pathway, read the SPOT capability briefing.
Research foundation and hypothesis
The project builds on Dr Richard Matthews’ doctoral investigation of sensor pattern noise and subsequent peer-reviewed research. It investigates how measurable characteristics associated with image creation can support forensic evaluation of origin.
A finding about image-origin consistency must be understood within the tested conditions and the limitations of the method. It is not, by itself, a determination of the truth of the scene depicted.

Physical image formation: Sony IMX219PQ sensor cross-section. Source: Richard Matthews, 2019 doctoral thesis, Figure 5.2, page 61. Original annotations and scale retained. Thesis record. This is foundational sensor research, not an image of the current SPOT capability.
Public methodology
The 2026 synthetic-media paper describes extracting noise residues from grayscale imagery, measuring their statistical characteristics and modelling camera and synthetic-image populations. It reports likelihood ratios to express relative support for competing origin hypotheses.
The study uses a defined experimental dataset. Its results do not establish performance for every camera, image generator or image-processing history. Read the publication record for the citation and publisher link.
Experimental evidence and evaluation status
The 2026 synthetic-media publication reports an experiment using camera imagery from the Dresden Image Database and DALL·E-generated imagery. It examines noise-residue characteristics and uses statistical modelling and likelihood ratios to assess competing origin hypotheses.
That published experiment and the evolving SPOT capability must be distinguished. A result from a defined study does not establish the performance of a later implementation or a new operational dataset.
Published experiment: reported results
The paper’s Results section reports the following means over 100 randomly sampled runs:
| Metric | Reported mean |
|---|---|
| ROC AUC | 0.998 |
| Accuracy | 0.985 |
| Precision | 0.987 |
| Recall | 0.983 |
The method describes balanced Dresden camera and DALL·E 3 samples, random sampling without replacement, a 70/30 training/test split, wavelet-derived noise residues and three features: variance, skewness and kurtosis. Classification uses the likelihood-ratio boundary of 1.
Replotted from the explicit counts in the paper’s Results section, page S70. This is one reported example, not the mean across 100 runs or a fresh evaluation of the current implementation. Original publication.
Reading the numbers carefully. The abstract states accuracy of 0.998, while the Results section gives mean accuracy of 0.985. The example counts imply 2,975/3,000 correct classifications (99.17%). These are presented separately; the abstract’s figure should not be used as a general accuracy claim. The published full and zoomed ROC panels also identify different seed examples. Resolving these reporting differences requires the underlying run records.
The paper describes a Wilson 95% interval method, but the Results text does not provide numerical interval bounds for these means. The source data, run-level variation, implementation freeze and timing measurements are not available with the results presented here. Operational performance remains to be evaluated.
Microsoft for Startups supported the published research; RHEM Labs’ participation in that programme has since ended.
Further experimental performance profile

Red: corrected SPOT results. Dashed grey: recalculated supplier average. This is a separate evaluation from the published Dresden/DALL·E experiment above.
The figure compares balanced accuracy, specificity, F1 score and class-specific precision and recall for not-synthetic, partially synthetic and fully synthetic content. SPOT’s relative performance varies by measure; the chart should be considered as a whole.
The underlying dataset, decision thresholds, numerical run records and uncertainty intervals are not published here. It is an experimental comparison, not a guarantee of performance on new material or a reproducible public benchmark.
Towards reproducible public evaluation
The laboratory’s publication framework for a future evaluation comprises:
- Frozen design: the implementation version, hypotheses, preprocessing, exclusions, thresholds and test protocol fixed before evaluating the held-out data.
- Dataset composition: source provenance, camera and generator families, image counts, processing histories, class balance and the separation of training, calibration and test material.
- Repeated runs: seeds or resampling design, per-run outputs, variation and confidence intervals where justified by the sampling design.
- Discrimination and calibration: score distributions and ROC/AUC where appropriate, threshold-specific error rates, and calibration checks for any reported likelihood ratios. Accuracy requires its class balance and decision threshold to be stated.
- Computational performance: hardware, software versions, image sizes, timing method and resource usage.
- Comparisons: the same eligible test population, documented settings and a clear account of which methods answer comparable questions.
This framework will guide the reporting of future evaluations.
Open research questions
Photographs of synthetic imagery. A camera photographing a display or print introduces a real image-formation process. What can origin assessment establish about the captured file, and what remains unknown about the content it depicts?
Screenshots and recompression. How do capture routes, platform transformations and repeated compression affect the measured characteristics?
Editing and in-camera processing. Which operations preserve, obscure or introduce characteristics relevant to the hypotheses? How does performance vary across device processing pipelines?
Generalisation. What changes with unseen cameras, generators, image sizes and acquisition histories? Which observations fall outside the evaluated population?
Operational interpretation. How should inconclusive results and uncertainty be communicated when imagery is used in forensic investigation or open-source intelligence?
Publications and capability development
Return to the SPOT capability briefing, explore the publication library, or discuss research and technical partnerships. The public research collection links the thesis and supporting work.
The Defence evaluation pathway considers representative testing and integration into analytical workflows. Contact RHEM Labs to discuss a defined evaluation or capability-development programme.