Exploring techniques to distinguish between real images and those generated using stable diffusion XL

B Benjamin Sanders D David Morrison D David Harris-Birtill

Abstract

The recent development of text-to-image diffusion models has allowed us to quickly generate realistic images from textual prompts. Despite enabling innovation in particular domains, concerns have been raised over the prospect of malicious users posing synthetic images as genuine. To assess if it is possible to discern between real images and those generated using diffusion models, a novel convolutional neural network was built, trained and tested on a bespoke dataset formed of authentic images from the ImageNet dataset and corresponding synthetic images generated using Stable Diffusion XL: an open-source text-to-image diffusion model. With the public release of this dataset, it is currently the largest publicly accessible collection of images generated using Stable Diffusion XL, significantly contributing to future research in this area. The positive results from our experiment performing a binary classification of synthetic and real images demonstrate the effectiveness in detecting synthetic images, with up to 98.38% accuracy using a ResNet-18 baseline, and 97.24% with the proposed CNN.

Article Details

Journal PLoS ONE
Volume / Issue Vol. 21, Issue 1
Published January 27, 2026
Pages e0339917
ISSN 1932-6203
Publisher Public Library of Science

Journal Info

PLoS ONE

Public Library of Science

ISSN: 1932-6203 Open Access Health Sciences

Authors (3)

B

Benjamin Sanders

D

David Morrison

D

David Harris-Birtill