Synthetic image detection with CLIP. Understanding and assessing predictive cues

Lade...
Vorschaubild
Autor:in (Körperschaft)
Publikationsdatum
2026
Typ der Arbeit
Studiengang
Typ
01A - Beitrag in wissenschaftlicher Zeitschrift
Herausgeber:innen
Herausgeber:in (Körperschaft)
Betreuer:in
Übergeordnetes Werk
Array: opening up computer science
Themenheft
Link
Zugehörige Forschungsdaten
Reihe / Serie
Reihennummer
Jahrgang / Band
32
Ausgabe / Nummer
Seiten / Dauer
101222
Patentnummer
Verlag / Herausgebende Institution
Elsevier
Verlagsort / Veranstaltungsort
Auflage
Version
Programmiersprache
Abtretungsempfänger:in
Praxispartner:in/Auftraggeber:in
Zusammenfassung
Recent generative models produce near-photorealistic images, challenging the trustworthiness of photographs. Synthetic image detection (SID) methods, however, often struggle to generalize across datasets and generative models. CLIP, which embeds images and text in a shared seamantic space, performs well at SID, but the cues underlying its decisions remain poorly understood. We therefore study CLIP-based SID as an empirical interpretability problem rather than proposing a new detector. We introduce SynthCLIC , which pairs real photographs with caption-matched, high-quality diffusion-generated counterparts. We evaluate CLIP-based detectors on SynthCLIC , a GAN-heavy benchmark, and a broad external benchmark, and compare them with a low-level forensic CNN, a broad-generator detector, and a text-grounded concept model. CLIP-based linear detectors reach 0.96 mAP on the GAN-heavy benchmark but 0.92 on SynthCLIC , while cross-family transfer to CNNSpot falls to 0.42 mAP. Within-class associations between detector scores and text-derived cue scores show that higher synthetic scores correspond to cleaner, more compositionally controlled, and technically polished images, whereas lower scores correspond to messier capture conditions and provenance cues characteristic of real photographs. These associations are distributed across many overlapping cues, and their profiles differ strongly across training datasets. CLIP-based and forensic detectors therefore fail in different ways and provide complementary evidence, while broad generator coverage appears important for robust SID.
Schlagwörter
Projekt
Veranstaltung
Startdatum der Ausstellung
Enddatum der Ausstellung
Startdatum der Konferenz
Enddatum der Konferenz
Datum der letzten Prüfung
ISBN
ISSN
2590-0056
Sprache
Englisch
Während FHNW Zugehörigkeit erstellt
Ja
Zukunftsfelder FHNW
Publikationsstatus
Veröffentlicht
Begutachtung
peer-reviewed
Open Access-Status
Gold
Lizenz
'https://creativecommons.org/licenses/by/4.0/'
Zitation
Willi, M., Mathys, M., & Graber, M. (2026). Synthetic image detection with CLIP. Understanding and assessing predictive cues. Array: Opening up Computer Science, 32, 101222. https://doi.org/10.1016/j.array.2026.101222