Talks & Media Coverage

Talks and Media Coverage

  • TOP
  • Talks and Media Coverage
  • [Invited Talk] Countermeasures against Misuse of Voice Generation AI and Voice Cloning: From Deepfake Detection to Active Defense

Share

  • Invited Talks & Tutorials

Acoustical Society of Japan Kyushu Branch 1st Online Seminar

[Invited Talk] Countermeasures against Misuse of Voice Generation AI and Voice Cloning: From Deepfake Detection to Active Defense

  • #DeepfakeDetection
  • #Audio processing
  • #Generative model

Speaker: Junichi Yamagishi
Conference name: 1st Online Seminar of the Kyushu Branch of the Acoustical Society of Japan
Organizer: Kyushu Branch of the Acoustical Society of Japan
Venue: Online
Date: July 2, 2025
URL:

Recent voice generation models, particularly voice cloning technology that reproduces speaker identity, have brought new value to entertainment and other fields. However, when misused, their high fidelity can cause serious problems in personal authentication systems and elsewhere. In this talk, we introduce our work and findings on defensive models against deepfake impersonation attacks. We first present a large-scale speech database for training deepfake speech detection models, along with evaluation data for two scenarios: detection over telephone channels and detection on compressed audio. We then share insights drawn from analyzing 50 detection models built on this database. Finally, taking into account the fact that media generation technology is constantly evolving and new techniques are continuously being developed, we introduce an approach for detecting deepfakes produced by previously unseen methods.

In the second half of the presentation, I will discuss not only passive defense frameworks that consider detection methods after audio is generated, but also active defense, which considers abuse prevention measures from the early stages of audio generation and publication, and comprehensively builds an environment that makes inappropriate use difficult. First, I will introduce the usefulness and limitations of neural watermarking, which processes the model weights of audio generation AI and automatically embeds watermarks in the audio output. Finally, I will introduce speaker anonymization technology, which reduces the risk of deepfakes being created and protects the privacy of our voices by anonymizing only the features associated with the speaker before publishing audio on social media, etc.