Synthetic image detection with CLIP. Understanding and assessing predictive cues

dc.contributor.authorWilli, Marco
dc.contributor.authorMathys, Melanie
dc.contributor.authorGraber, Michael
dc.date.accessioned2026-10-05T11:55:30Z
dc.date.issued2026
dc.description.abstractRecent generative models produce near-photorealistic images, challenging the trustworthiness of photographs. Synthetic image detection (SID) methods, however, often struggle to generalize across datasets and generative models. CLIP, which embeds images and text in a shared seamantic space, performs well at SID, but the cues underlying its decisions remain poorly understood. We therefore study CLIP-based SID as an empirical interpretability problem rather than proposing a new detector. We introduce SynthCLIC , which pairs real photographs with caption-matched, high-quality diffusion-generated counterparts. We evaluate CLIP-based detectors on SynthCLIC , a GAN-heavy benchmark, and a broad external benchmark, and compare them with a low-level forensic CNN, a broad-generator detector, and a text-grounded concept model. CLIP-based linear detectors reach 0.96 mAP on the GAN-heavy benchmark but 0.92 on SynthCLIC , while cross-family transfer to CNNSpot falls to 0.42 mAP. Within-class associations between detector scores and text-derived cue scores show that higher synthetic scores correspond to cleaner, more compositionally controlled, and technically polished images, whereas lower scores correspond to messier capture conditions and provenance cues characteristic of real photographs. These associations are distributed across many overlapping cues, and their profiles differ strongly across training datasets. CLIP-based and forensic detectors therefore fail in different ways and provide complementary evidence, while broad generator coverage appears important for robust SID.
dc.identifier.doi10.1016/j.array.2026.101222
dc.identifier.issn2590-0056
dc.identifier.urihttps://irf.fhnw.ch/handle/11645/58169
dc.identifier.urihttps://doi.org/10.26041/fhnw-17366
dc.language.isoen
dc.publisherElsevier
dc.relation.ispartofArray: opening up computer science
dc.rights.urihttps://creativecommons.org/licenses/by/4.0/
dc.subject.ddc005 - Computer Programmierung, Programme und Daten
dc.titleSynthetic image detection with CLIP. Understanding and assessing predictive cues
dc.type01A - Beitrag in wissenschaftlicher Zeitschrift
dc.volume32
dspace.entity.typePublication
fhnw.InventedHereYes
fhnw.ReviewTypepeer-reviewed
fhnw.openAccessCategoryGold
fhnw.pagination101222
fhnw.publicationStatePublished
fhnw.targetcollectionb508cce9-5084-49ae-a565-d8e5c348c3ab
relation.isAuthorOfPublication8be981ed-69d0-49fc-b917-5d01c50e21de
relation.isAuthorOfPublicationc94f2103-6b86-4da9-ba33-c9427a598fc0
relation.isAuthorOfPublication3693006b-11c7-4d96-9508-bd6f7c8b2301
relation.isAuthorOfPublication.latestForDiscovery8be981ed-69d0-49fc-b917-5d01c50e21de
Dateien

Originalbündel

Gerade angezeigt 1 - 1 von 1
Lade...
Vorschaubild
Name:
1-s2.0-S259000562600545X-main.pdf
Größe:
3.9 MB
Format:
Adobe Portable Document Format

Lizenzbündel

Gerade angezeigt 1 - 1 von 1
Lade...
Vorschaubild
Name:
license.txt
Größe:
2.66 KB
Format:
Item-specific license agreed upon to submission
Beschreibung: