AIWiki
Malaysia
Back to all articles
Applicationsvoice cloningtext-to-speechdeepfake

Voice Cloning

4 min readUpdated September 2026
Voice Cloning
Type
AI speech synthesis technology
Also known as
Voice synthesis, voice replication
Data required
Seconds to minutes of reference audio
Leading vendors
ElevenLabs, Microsoft, OpenAI, Resemble AI
Key risks
Impersonation fraud, non-consensual audio, voice-biometric bypass
Related
Deepfake, Text-to-Speech, AI Watermarking
Voice cloning is the use of artificial intelligence to build a synthetic model of a person's voice and generate new speech in that voice from a limited audio sample. The technique underpins commercial text-to-speech, film dubbing and accessibility software, and it has become one of the most consequential tools in fraud: the United States Federal Bureau of Investigation recorded US$893 million in AI-related scam losses in 2025, with voice cloning named among the main drivers.[1][2]

Background

Neural speech synthesis emerged with the deep learning wave of the mid-2010s. WaveNet, presented by DeepMind researchers in 2016, showed that neural networks could generate raw audio waveforms far more naturally than earlier concatenative systems.[3] A second wave began in 2023, when zero-shot models such as Microsoft's VALL-E demonstrated that a voice could be approximated from as little as three seconds of reference audio.[4]

Commercialisation followed quickly. ElevenLabs, founded in 2022, raised US$500 million in February 2026 at an US$11 billion valuation, more than tripling its valuation in a year, and now markets enterprise voice agents alongside its speech models.[5]

How It Works

A modern voice cloning system extracts a speaker representation — an embedding of timbre, pitch and speaking style — from reference recordings, then conditions a neural text-to-speech model on that representation. Two families of methods dominate: zero-shot cloning, which needs only a short sample and no training, and fine-tuned cloning, which trains on minutes to hours of audio for higher fidelity.

Because zero-shot cloning needs so little input, public speeches, podcasts, voicemails and social media clips all supply enough material for convincing forgeries. Streaming systems have also cut latency to fractions of a second, enabling cloned voices in live conversation.

Applications

Legitimate uses span dubbing and localisation for film and television, audiobook and podcast production, game character voices, and accessibility, including voice banking that preserves a person's voice for use after speech loss. Voice agents built on the technology now answer customer service calls in several industries. The growth of cloning has also raised questions about consent and compensation for professional voice actors, whose recorded work can be used as training data.

Fraud and Misuse

The FBI's Internet Crime Complaint Center logged 22,364 AI-related complaints in 2025, with reported losses reaching US$893 million; impersonation of family members in distress, fake executive instructions to employees and bypassing of voice-based verification were among the recurring patterns.[2] Security researchers advise verification through a separate channel before acting on any voice-only request for money or access.

Regulation and Countermeasures

There is no single global regime for voice cloning. The European Union's AI Act requires disclosure when audio has been artificially generated, with transparency duties for deepfakes applying since 2 August 2026.[6] Industry responses include consent policies for voice libraries, provenance metadata in generated audio, and watermarking research. Detection of cloned audio remains difficult, which has pushed institutions to rely less on voice alone for authentication.

>See Also

🇲🇾Malaysian Context

The Malaysian Communications and Multimedia Commission warned in December 2025 about "silent calls" used to harvest voices: scammers dial a number and stay silent, and if the target answers and speaks, the recording is later cloned and used to impersonate the victim to relatives, employers or institutions. The commission advised the public to remain silent on such calls and to hang up.[7]

The wider fraud picture is severe. Police recorded 35,368 online financial scam cases in 2024 involving RM1.58 billion in losses, and the Ministry of Digital has issued AI guidelines and training programmes as part of its anti-scam response.[8] Bank Negara Malaysia has itself issued warnings after deepfake videos and cloned voices of its governor and senior officials were used to promote fake financial products.[9] For Malaysian households and businesses, the practical defences are low-tech: agree a family verification phrase, confirm urgent money requests through a second channel, and treat unexpected callers asking for personal details as hostile by default.

References

  1. ↑Federal Bureau of Investigation. (2026). 2025 Internet Crime Report. https://www.ic3.gov/AnnualReport/Reports/2025_IC3Report.pdf
  2. ↑Malwarebytes. (2026). Americans lost nearly $900 million to AI-powered scams, FBI says. https://www.malwarebytes.com/blog/scams/2026/06/americans-lost-nearly-900-million-to-ai-powered-scams-fbi-says
  3. ↑van den Oord, A. et al. (2016). WaveNet: A Generative Model for Raw Audio. https://arxiv.org/abs/1609.03499
  4. ↑Wang, C. et al. (2023). Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers. https://arxiv.org/abs/2301.02111
  5. ↑ElevenLabs. (2026). ElevenLabs raises $500M Series D at $11B valuation. https://elevenlabs.io/blog/series-d
  6. ↑EU Artificial Intelligence Act. Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems. https://artificialintelligenceact.eu/article/50/
  7. ↑Asia News Network / The Star. (2025). Silence is golden in new call scam, as AI used to clone voices, says Malaysian commission. https://asianews.network/silence-is-golden-in-new-call-scam-as-ai-used-to-clone-voices-says-malaysian-commission/
  8. ↑Ministry of Digital, Malaysia. (2025). Garis Panduan AI, Program Latihan Antara Langkah Proaktif Kementerian Digital Untuk Memerangi Online Scam. https://www.digital.gov.my/en-GB/siaran/Garis-Panduan-AI,-Program-Latihan-Antara-Langkah-Proaktif-Kementerian-Digital-Untuk-Memerangi-Online-Scam
  9. ↑MIMOS. (2026). You Cannot Trust What You See Anymore: Malaysia's Deepfake Crisis and the Technology Fighting Back. https://www.mimos.my/you-cannot-trust-what-you-see-anymore-malaysias-deepfake-crisis-and-the-technology-fighting-back/