XceptionNet
Definition
A depthwise-separable convolutional architecture proposed by Rossler et al. as the baseline binary classifier in FaceForensics++. Trained to distinguish real from manipulated frames. Achieves high per-frame accuracy at low compression but degrades sharply at high compression and on unseen methods.
- Architecture type
- Depthwise-separable convolutional network
- Role
- Baseline binary classifier in FaceForensics++
- Proposed by
- Rossler et al.
- Strength
- High per-frame accuracy at low compression
- Weakness
- Accuracy degrades sharply at high compression and unseen methods
Common questions
Why does XceptionNet perform worse on unseen manipulation methods?+
The network learns artefacts specific to the manipulation techniques in its training data, so a deepfake generated by an unrepresented method can evade detection even though the network scores well on methods it has seen, a generalisation gap later architectures try to close.
What does compression do to XceptionNet's detection accuracy?+
Video compression removes or distorts the high-frequency artefacts XceptionNet relies on to distinguish real from manipulated frames, so accuracy that is strong on raw or lightly compressed footage drops substantially on heavily compressed video typical of social media.
Is XceptionNet still considered state of the art for deepfake detection?+
No, it functions mainly as the historical baseline FaceForensics++ established; current detection work has moved toward architectures and foundation-model features with better cross-method generalisation, though it remains a common comparison point in published benchmarks.
Related terms
- CLIP-Based Detection
- Detection approach using OpenAI CLIP or similar vision-language foundation models as a feature extractor. The broad pre-training enables generalisation to generation methods...
- CNN Residual Detector
- A convolutional neural network trained on the high-frequency residual image, the difference between the original and a de-noised version, to classify whether...
- Detection Generalisation
- The capacity of a trained detector to correctly identify deepfakes produced by generators not seen during training. Low generalisation is the central...
- FaceForensics++
- A video dataset released by Rossler et al. (2019) containing 1000+ YouTube videos manipulated by four methods: DeepFakes, Face2Face, FaceSwap, and NeuralTextures....
- Frequency-Domain Analysis
- Detection approach that transforms image patches into the frequency domain (DCT or FFT) to expose periodic artifacts introduced by upsampling layers in...
- Frequency-Domain Artefact
- A periodic or statistical anomaly in the Fourier spectrum of an image or audio signal introduced by the generation pipeline's upsampling, filter,...
- Generalisation Gap
- The drop in detection accuracy when a classifier trained on one generation method is applied to a different method. Caused by learning...
- Noiseprint
- A CNN-based camera-model fingerprint extractor by Cozzolino and Verdoliva. Applied to deepfakes, it reveals inconsistency between the camera fingerprint in the genuine...
- Physiological Signal
- A biological process visible in video, such as eye blinking, rPPG (remote photoplethysmography), and head micro-motion from the cardiac cycle, that deepfake...
- Remote Photoplethysmography (rPPG)
- A technique that detects the pulse-driven skin-colour variation in a face video without contact sensors. In real video, this signal is present...
- rPPG
- Remote photoplethysmography. A technique for measuring heart rate from subtle periodic colour changes in facial skin caused by blood-volume pulses. Authentic video...