ASVspoof
Definition
A recurring evaluation campaign and dataset series that benchmarks anti-spoofing countermeasures against corpora of genuine and spoofed utterances. Editions in 2015, 2017, 2019, 2021, and 2024 each introduce new attack types. The primary source of standardised training and test data for speech anti-spoofing research.
- Type
- Recurring evaluation campaign and dataset series
- Editions
- 2015, 2017, 2019, 2021, 2024
- Purpose
- Benchmark anti-spoofing countermeasures
- Field
- Speech anti-spoofing / voice forensics research
Common questions
Why does ASVspoof need a new edition every few years instead of one fixed dataset?+
Spoofing methods evolve quickly, especially with newer neural voice conversion and synthesis techniques, so each edition introduces new attack types to keep the benchmark representative of current threats rather than testing countermeasures against attacks that are already outdated.
Does a countermeasure that performs well on ASVspoof data guarantee it will catch real-world deepfake audio?+
Not automatically. ASVspoof data is built from specific attack algorithms and recording conditions, so a detector tuned to that benchmark can still underperform on novel synthesis methods or audio conditions not represented in the dataset, which is a known generalisation limitation in the field.
Related terms
- Anti-Spoofing Countermeasure (CM)
- A classifier, also called a CM system, trained to output a score indicating the probability that a given audio segment is genuine...
- Equal Error Rate (EER)
- The point on a classifier's detection error tradeoff curve where the false accept rate equals the false reject rate. Lower EER indicates...
- Neural TTS Cloning
- A text-to-speech system that adapts to a target speaker using a short enrollment recording, generating new utterances in that speaker's voice from...
- Tandem Detection Cost Function (T-DCF)
- The primary evaluation metric in ASVspoof from 2019 onward. It measures the cost of errors when a countermeasure is integrated with an...
- Voice Conversion
- A signal-processing or deep-learning technique that transforms the vocal characteristics of a source speaker's utterance to match a target speaker, while preserving...