What voice cloning is — how 3 seconds of audio copies a voice, and how to stop it
Voice cloning is the use of AI speech synthesis to imitate a particular person so that the voice says things they never said. Modern models first learn how speech works from tens of thousands of hours of audio, then copy the timbre and style of a new speaker from a short recording; Microsoft research model VALL-E did it from three seconds in 2023. The technology narrates audiobooks and restores voices for patients who lose them, but it has also been used in fraud, such as a 2019 case where a cloned executive voice extracted €220,000, and in fake election robocalls. Detection is unreliable, so simple checks such as hanging up and calling back on a known number, or a family code word, work best
The three lines
- How — a model learns speech from huge datasets, then lifts a voice print from seconds of audio
- Uses and abuses — audiobooks and voice restoration, but also CEO-fraud transfers and fake election calls
- Defence — procedure beats detection: call back, agree a code word, pause on urgent money requests
Key questions
- How many seconds to clone a voice
- **Three seconds in research; commercial services often want tens of seconds to minutes.** | Example | Sample needed | |---|---| | Microsoft VALL-E (2023 research) | 3 seconds | | OpenAI Voice Engine (2024, limited preview) | 15 seconds | | Commercial cloning services | Tens of seconds to minutes |
- AI voice cloning scam examples
- **Executive impersonation and election robocalls.** | When | Case | |---|---| | 2019 | UK energy firm CEO wires €220,000 after a call mimicking a parent-company executive | | January 2024 | Robocalls imitating the U.S. president urged New Hampshire voters to skip the primary | | February 2024 | FCC rules AI-generated voices in robocalls illegal |
- How to protect against voice cloning
- **Verification procedures beat detection software.** | Step | Why it works | |---|---| | Hang up, call back on a saved number | Caller ID can be spoofed; your call reaches the real person | | Family code word | A clone copies the voice, not shared secrets | | Pause on urgent money requests | Scams rely on time pressure | | Two-person approval at work | One mistake cannot move money |
The voice on the phone can sound exactly like your mother and still not be her. Voice cloning uses AI to copy a specific person's voice and make it say things they never said. It is why Italy's prime minister filed to trademark her voice in October 2026 (see "Meloni files to trademark her voice"). Knowing how it works explains why it is hard to stop — and what still works.
1. How it works — learning "speech" and "whose voice" separately
Older text-to-speech needed dozens of studio hours from one person. Today the job is split.
| Step | What happens | Data needed |
|---|---|---|
| 1. Pre-training | Learn general rules of speech — pronunciation, intonation, breathing — from many speakers | Tens of thousands of hours |
| 2. Speaker adaptation | Extract a "voice print" of timbre and habits from a new person | Seconds to minutes |
| 3. Synthesis | Read any sentence in that voice | Text |
Imitating an unseen speaker from a short sample is called zero-shot synthesis. Microsoft's VALL-E, trained on about 60,000 hours of English, did it in 2023 from three seconds of audio, even copying emotion and room acoustics. In 2024 OpenAI previewed Voice Engine, which needed 15 seconds, but limited it to selected partners because of misuse risk.
| Example | Sample | Availability |
|---|---|---|
| Microsoft VALL-E (2023) | 3 seconds | Paper only |
| OpenAI Voice Engine (2024) | 15 seconds | Limited preview |
| Commercial cloning services | Tens of seconds to minutes | Public, consent checks vary |
2. Uses and abuses
The benefits are real: audiobooks in an author's voice, dubbing that keeps an actor's voice across languages, and patients with ALS banking their voice before they lose it.
Abuse arrived early.
| When | Case | Outcome |
|---|---|---|
| 2019 | UK energy firm CEO gets a call mimicking a German parent-company executive | €220,000 wired |
| January 2024 | Robocalls imitating the U.S. president tell New Hampshire voters not to vote in the primary | Election-interference probe |
| February 2024 | U.S. Federal Communications Commission | AI-generated voices in robocalls ruled illegal |
| October 2026 | Italian PM Giorgia Meloni | Files voice trademark |
The common thread: scams exploit the habit of treating a voice as proof of identity, then add time pressure so there is no room to check.
3. Why detection struggles
| Problem | Explanation |
|---|---|
| Phone audio | Compression wipes out synthesis artifacts too |
| Arms race | Each detector feature gets fixed by the next model |
| Short clips | "Mom, it's me" gives too little signal |
| Labeling limits | Honest providers watermark; fraudsters do not |
The EU AI Act requires machine-readable marks on AI-generated audio (see "What Article 50 of the EU AI Act is"), but only lawful services comply — and even text watermarks fail once a quarter of the words are changed.
4. Defences that work — procedure over technology
| Step | Why it works |
|---|---|
| Hang up and call back on a saved number | Caller ID can be spoofed; the number you dial reaches the real person |
| Agree a family or office code word | A clone copies the voice, not shared secrets |
| Pause on any urgent request for money or codes | Scams almost always depend on time pressure |
| Two-person approval for company transfers | One mistaken employee cannot move funds |
| Share less voice and video publicly | Less raw material (it will not stop a determined attacker) |
| Avoid voice-only authentication for banking | A voice alone is no longer proof of identity |
The core rule: stop using a voice as an ID. If leaked data such as income and loan limits — as in the October 2026 Korean bank hacks — lets a caller sound informed as well as familiar, verification matters even more.
5. What remains unconfirmed
- No official count of AI-voice phone scams in South Korea was found.
- Sample-length and consent rules vary by commercial service.
- AI voice detectors lack independent accuracy evaluations.
Sources
- Microsoft Research (arXiv) — Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers (VALL-E)
- OpenAI — Navigating the Challenges and Opportunities of Synthetic Voices
- FCC — FCC Makes AI-Generated Voices in Robocalls Illegal
- The Wall Street Journal — Fraudsters Used AI to Mimic CEO's Voice in Unusual Cybercrime Case
- AP via ABC News — Italy's Meloni follows pop stars in seeking to trademark her voice