Neural TTS Cloning
Definition
A text-to-speech system that adapts to a target speaker using a short enrollment recording, generating new utterances in that speaker's voice from text input. The input is text; the output is fully synthesised speech. Also called voice cloning or speaker-adaptive TTS.
- Technology
- Speaker-adaptive text-to-speech
- Input
- Text plus a short enrolment recording
- Output
- Fully synthesised speech in target voice
- Also called
- Voice cloning
Common questions
How does neural TTS cloning differ from voice conversion?+
TTS cloning generates new speech from text input in the cloned voice, while voice conversion transforms an existing spoken recording from one voice into another.
What forensic clues can distinguish cloned speech from a genuine recording?+
Subtle unnatural prosody, spectral artefacts left by the synthesis model, and inconsistent breathing or micro-pauses that natural speech reliably contains.
Related terms
- Anti-Spoofing Countermeasure (CM)
- A classifier, also called a CM system, trained to output a score indicating the probability that a given audio segment is genuine...
- ASVspoof
- A recurring evaluation campaign and dataset series that benchmarks anti-spoofing countermeasures against corpora of genuine and spoofed utterances. Editions in 2015, 2017,...
- Equal Error Rate (EER)
- The point on a classifier's detection error tradeoff curve where the false accept rate equals the false reject rate. Lower EER indicates...
- Tandem Detection Cost Function (T-DCF)
- The primary evaluation metric in ASVspoof from 2019 onward. It measures the cost of errors when a countermeasure is integrated with an...
- Voice Conversion
- A signal-processing or deep-learning technique that transforms the vocal characteristics of a source speaker's utterance to match a target speaker, while preserving...