Skip to content

CLIP-Based Detection

Definition

Detection approach using OpenAI CLIP or similar vision-language foundation models as a feature extractor. The broad pre-training enables generalisation to generation methods not seen during fine-tuning, at the cost of higher compute and reduced interpretability of the decision.

Base model
OpenAI CLIP or similar vision-language model
Use
Feature extractor for deepfake detection
Strength
Generalises to unseen generation methods
Tradeoff
Higher compute, lower interpretability

Common questions

Why does CLIP generalise better than a detector trained only on known deepfake datasets?+

CLIP is pre-trained on a very broad set of natural images and text pairs rather than deepfake-specific artefacts, so the features it extracts capture general visual regularities that transfer to manipulation methods the detector never saw during fine-tuning, unlike a narrowly trained CNN classifier.

What is the practical downside of using CLIP features in a forensic deepfake report?+

Because the decision is based on high-dimensional learned embeddings rather than a specific, nameable artefact like a blending boundary or frequency anomaly, it is harder to explain to a court exactly why the system flagged an image, which weakens the report's evidentiary transparency.

Related terms

CNN Residual Detector
A convolutional neural network trained on the high-frequency residual image, the difference between the original and a de-noised version, to classify whether...
Detection Generalisation
The capacity of a trained detector to correctly identify deepfakes produced by generators not seen during training. Low generalisation is the central...
FaceForensics++
A video dataset released by Rossler et al. (2019) containing 1000+ YouTube videos manipulated by four methods: DeepFakes, Face2Face, FaceSwap, and NeuralTextures....
Frequency-Domain Analysis
Detection approach that transforms image patches into the frequency domain (DCT or FFT) to expose periodic artifacts introduced by upsampling layers in...
Frequency-Domain Artefact
A periodic or statistical anomaly in the Fourier spectrum of an image or audio signal introduced by the generation pipeline's upsampling, filter,...
Generalisation Gap
The drop in detection accuracy when a classifier trained on one generation method is applied to a different method. Caused by learning...
Noiseprint
A CNN-based camera-model fingerprint extractor by Cozzolino and Verdoliva. Applied to deepfakes, it reveals inconsistency between the camera fingerprint in the genuine...
Physiological Signal
A biological process visible in video, such as eye blinking, rPPG (remote photoplethysmography), and head micro-motion from the cardiac cycle, that deepfake...
Remote Photoplethysmography (rPPG)
A technique that detects the pulse-driven skin-colour variation in a face video without contact sensors. In real video, this signal is present...
rPPG
Remote photoplethysmography. A technique for measuring heart rate from subtle periodic colour changes in facial skin caused by blood-volume pulses. Authentic video...
XceptionNet
A depthwise-separable convolutional architecture proposed by Rossler et al. as the baseline binary classifier in FaceForensics++. Trained to distinguish real from manipulated...

Explained in

Your journey to becoming a forensic professional starts here.

Practice with mock tests, learn from structured notes, and get your questions answered by a global forensic community, all in one place.