VoiceRadar: Voice Deepfake Detection using Micro-Frequency and Compositional Analysis

Kavita Kumari

Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · Audio Security

Overview

The proliferation of sophisticated deepfake audio poses a significant and escalating threat to personal and societal security. From bypassing voice authentication systems to propagating disinformation in political warfare and enabling character assassination through fabricated statements, the malicious applications of synthetic speech are diverse and alarming. This talk, presented by Alexandro Pegoro at the NDSS Symposium, introduces VoiceRadar, a novel deepfake detection system designed to overcome the inherent limitations of existing detection methodologies.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction and threat of deepfake audio
  2. 0:50 Deepfake generation process: TTS and STS
  3. 3:00 Limitations of existing deepfake detection methods
  4. 4:00 VoiceRadar's core hypothesis: human vs. machine differences
  5. 5:00 VoiceRadar's physical model and micro-frequency extraction
  6. 6:00 Detailed explanation of Doppler effect and vibrational frequencies
  7. 7:50 VoiceRadar's evaluation, dataset creation, and results

VoiceRadar: Voice Deepfake Detection using Micro-Frequency and Compositional Analysis

Speakers: Alexandro Pegoro (presenter), Kavita Kumari (author)

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=a5LHKy8O0cY

Overview

The proliferation of sophisticated deepfake audio poses a significant and escalating threat to personal and societal security. From bypassing voice authentication systems to propagating disinformation in political warfare and enabling character assassination through fabricated statements, the malicious applications of synthetic speech are diverse and alarming. This talk, presented by Alexandro Pegoro at the NDSS Symposium, introduces VoiceRadar, a novel deepfake detection system designed to overcome the inherent limitations of existing detection methodologies.

VoiceRadar distinguishes itself by moving beyond conventional machine learning approaches that primarily analyze abstract audio features. Instead, it delves into the subtle physical characteristics of sound propagation and human speech production, leveraging micro-frequency and compositional analysis. By approximating the physical model of how sound originates from a speaker and reaches an observer, VoiceRadar aims to identify the minute, often imperceptible, discrepancies that differentiate genuine human speech from machine-generated fakes.

The core motivation behind VoiceRadar stems from the observation that while human speech is inherently dynamic, influenced by myriad personal and environmental factors, machine-generated audio typically adheres to fixed patterns and rules. This fundamental difference, VoiceRadar posits, leaves unique sonic fingerprints that can be exploited for robust detection. The research highlights the critical need for advanced detection mechanisms that are not only accurate but also adaptable and resilient against an ever-evolving landscape of deepfake generation techniques, offering a promising step forward in the ongoing battle against synthetic media.

Background

▶ Watch: Introduction and threat of deepfake audio (0:00)

The landscape of deepfake audio generation has rapidly matured, with even major companies now offering or making publicly available sophisticated tools. The general principle involves taking original victim audio data and target content (either text or speech) and processing it through a series of steps. Feature extraction and feature blending are crucial, where characteristics such as text structure, words, language, tone, pitch, and pronunciation are extracted and mixed. This blended data then feeds into a multitude of generative AI models, ultimately producing the deepfake audio. Two primary generation methods exist: text-to-speech (TTS), which is more precise but time-consuming, and speech-to-speech (STS), which is faster, can be used for real-time applications like Zoom calls, but often introduces more artifacts.

Detecting these sophisticated fakes has been a persistent challenge. Existing deepfake detectors broadly fall into three categories:

  1. Unseen Feature-based Detectors: These analyze inherent, often subtle, attributes of audio files that deepfake generators struggle to replicate. An example cited is the inability of deepfakes to generate true "mute audio" or natural pauses, which are characteristic of real human speech.
  2. Graph Neural Network (GNN) based Detectors: These represent audio as a graph structure and train neural networks on this representation.
  3. Deep Neural Network (DNN) based Detectors: The most common category, where audio is represented as an encoding and fed into deep learning models.
  4. Online Proprietary Tools: These are black-box solutions whose inner workings are unknown.

Despite these efforts, existing detectors suffer from significant limitations. Evaluations conducted by the VoiceRadar team revealed that current solutions exhibit domain dependency, meaning their accuracy varies significantly between TTS and STS deepfakes. They also demonstrate limited adaptability and generalization, struggling to perform well against new or unseen deepfake generation methods. Crucially, most prior approaches were evaluated on old datasets (e.g., IVS poof audio), failing to account for the advancements in modern deepfake generation technologies.

VoiceRadar's foundational hypothesis addresses these shortcomings by proposing that human interaction is profoundly affected by personal and environmental factors. A speaker's nervousness, desire to make a good impression, or even subtle head movements introduce dynamic variations in speech patterns, pauses, tone, and pitch. In contrast, machine interaction is fixed, governed by the rules and patterns embedded within its generative model, making it inherently less dynamic. This led the researchers to seek ways to incorporate this a priori information—the physical and environmental context of speech production—into their detection model, a dimension largely ignored by previous work.

Key Findings

▶ Watch: Limitations of existing deepfake detection methods (3:00)

VoiceRadar's primary contribution lies in its innovative approach to deepfake detection, which centers on approximating the physical model of sound propagation to discern subtle differences between human and synthetic audio. The core hypothesis, that human speech is inherently dynamic while machine-generated speech is fixed, proved instrumental in guiding the system's design and efficacy.

The key findings and contributions of VoiceRadar include:

  1. Identification of Micro-Frequencies and Compositional Biases: VoiceRadar successfully extracts and utilizes subtle micro-frequencies and compositional biases that are unique to human speech. These minute variations, stemming from the physical act of speaking and environmental interactions, are incredibly difficult for current deepfake generators to replicate. The system quantifies these as differences in "energy" between human and machine interactions, finding that machine-generated audio consistently exhibits less energy in these specific micro-frequency components.
  1. Domain-Agnostic Detection: A significant breakthrough is VoiceRadar's ability to achieve domain-agnostic detection. Unlike previous detectors that struggled with varying accuracy across text-to-speech (TTS) and speech-to-speech (STS) domains, VoiceRadar demonstrated consistent and robust performance regardless of the deepfake generation method. This addresses a critical limitation of prior work, making VoiceRadar a more universally applicable solution.
  1. Robustness Against Adaptive Adversaries: The research explored the system's resilience against adaptive attacks, where an adversary is assumed to know the detector's mechanisms and attempts to project frequencies that would bypass it. VoiceRadar, particularly through its use of energy-based models rather than classical classification, proved remarkably robust. This suggests that current deepfake generation techniques are not equipped to emulate the complex physical properties VoiceRadar analyzes, requiring fundamental changes to future generative models to even attempt a bypass.
  1. Comprehensive Dataset Evaluation: The team addressed the issue of outdated evaluation datasets by creating their own comprehensive dataset. This included human samples from the VCTK dataset and deepfakes generated using eight of the most recent and advanced TTS and STS approaches. This rigorous evaluation methodology provided strong evidence for VoiceRadar's superior performance compared to existing literature, which often relied on older, less sophisticated deepfake examples.

In essence, VoiceRadar's key finding is that the physical reality of sound production leaves an indelible, complex signature that current synthetic speech models cannot fully mimic. By focusing on these often-overlooked physical and environmental factors, VoiceRadar establishes a new paradigm for robust and generalized deepfake audio detection.

Technical Deep Dive

▶ Watch: VoiceRadar's core hypothesis: human vs. machine differences (4:00)

VoiceRadar's technical innovation lies in its unique approach to embedding physical model approximations into the detection process. The system is designed to simulate how speech originates from a speaker and propagates to an observer, extracting subtle micro-frequencies that serve as critical differentiators between real and synthetic audio. This forms the compositional bias that is incorporated into the training algorithm, penalizing mismatches of energy between expected human speech characteristics and observed audio.

The core of VoiceRadar's micro-frequency analysis relies on the drum vibration representation, which models the complex ways in which sound waves are generated and propagated. This representation breaks down the speaker's frequency into three main components:

  1. Translational Frequency: This represents the simplest form of sound propagation—a direct, straight-line path between the speaker and the observer. It accounts for the fundamental frequency shifts based on the relative motion along this axis.
  1. Rotational Frequency: This component accounts for more dynamic movements, such as a speaker turning their head while talking or gesturing. These actions introduce subtle shifts in the sound wave's path and, consequently, the perceived frequency. This reflects the dynamic nature of human interaction, where physical movements are intrinsically linked to speech production.
  1. Speech Tremor Frequencies: These are the most subtle and perhaps most crucial components. They stem from individual and environmental factors that cause minute, often unconscious, vibrational variations in speech. Examples include the slight tremor in a voice due to nervousness, the natural pauses, or the unique physiological characteristics of an individual's vocal cords. These tremors prevent speech from being perfectly smooth and are incredibly difficult for generative AI models, which typically aim for idealized, smooth outputs, to authentically replicate.

These three frequency components are then used to calculate the Doppler effect from the source (speaker) to the observer. The Doppler effect, typically associated with changes in pitch due to relative motion, is here applied to capture the shifts in these micro-frequencies as they are perceived by the listener. This calculated observer frequency, along with the raw audio itself, is then fed into a machine learning model for training and inference.

Crucially, VoiceRadar employs energy-based models for detection, rather than conventional classification algorithms. The presenter highlighted that while traditional classifiers (e.g., those based on Transformers) aim to maximize a classification score (e.g., using softmax to normalize and identify the class with the highest value), energy-based models operate differently. They aim to find the lowest energy value. This means the optimization algorithm seeks to minimize an "energy" function, making the decision boundary less intuitive and more robust against adversarial manipulation. For an attacker to bypass an energy-based model, they would need to fundamentally alter their generative process to produce audio that exhibits the specific, low-energy micro-frequency characteristics associated with synthetic speech, rather than simply trying to mimic the high-level features that a classifier might focus on. This fundamental difference in the loss function and optimization strategy makes it significantly more challenging for adaptive adversaries, as current deepfake generators are optimized for "normal classification" objectives, not for minimizing an energy function in this specific micro-frequency domain. The approach penalizes "mismatch of energies," expecting less energy in these subtle physical characteristics for machine interactions and more for human interactions.

Demo / Proof of Concept

▶ Watch: Detailed explanation of Doppler effect and vibrational frequencies (6:00)

While the talk did not feature a live demonstration of a deployable tool, the research team presented a comprehensive evaluation methodology that served as a robust proof of concept for VoiceRadar's effectiveness. This evaluation rigorously tested VoiceRadar against both existing deepfake datasets and newly generated, state-of-the-art synthetic audio, directly addressing the limitations of prior work that often relied on outdated data.

To demonstrate VoiceRadar's capabilities, the researchers first constructed their own extensive datasets:

  • Human Sample Dataset: For authentic human speech, they utilized the widely recognized VCTK dataset.
  • Text-to-Speech (TTS) Deepfake Dataset: They generated synthetic audio using eight of the most recent TTS approaches available at the time of their research. This ensured their evaluation reflected the cutting edge of deepfake generation.
  • Speech-to-Speech (STS) Deepfake Dataset: A similar process was undertaken for STS deepfakes, employing advanced voice conversion techniques.

VoiceRadar was then evaluated against these newly created datasets, as well as existing, older deepfake datasets like IVS poof audio. The results graphically demonstrated VoiceRadar's superior performance, particularly its domain-agnostic nature. The system consistently achieved high true positive rates (detecting deepfakes as deepfakes) and true negative rates (differentiating humans from deepfakes) across both TTS and STS domains, a significant improvement over existing detectors that showed high domain dependency.

Furthermore, the team investigated VoiceRadar's resilience against adaptive attacks. They simulated scenarios where an adversary was aware of VoiceRadar's detection mechanisms and attempted to generate deepfakes specifically designed to mimic the observer's expected frequency profile. The findings indicated that while an attacker could theoretically attempt to project such frequencies, the resulting audio would require fundamental changes to current deepfake generation models. The present generation of deepfake tools is simply not equipped to produce the specific, complex micro-frequency characteristics that VoiceRadar analyzes, making adaptive attacks against it extremely difficult without a complete overhaul of the generative process itself. This robust evaluation, encompassing both broad performance and adversarial resilience, solidly established VoiceRadar as a highly promising and effective deepfake detection system.

Defensive Implications

▶ Watch: VoiceRadar's evaluation, dataset creation, and results (7:50)

The insights gleaned from VoiceRadar's research have profound defensive implications for individuals, organizations, and society at large. As deepfake technology continues its rapid advancement, understanding the fundamental differences between human and synthetic speech is paramount for building resilient defense mechanisms.

  1. Strengthening Voice Authentication: For systems relying on voice authentication, the findings highlight a critical vulnerability. Organizations should move beyond mere acoustic pattern matching and incorporate sophisticated analyses of micro-frequencies and physical sound propagation. This implies a need for multi-factor authentication, behavioral biometrics, and potentially real-time analysis of these subtle physical cues to prevent deepfakes from bypassing security protocols.
  1. Updating Deepfake Detection Systems: Existing deepfake detection systems, particularly those reliant on older datasets or general DNN approaches, are likely to be ineffective against newer, more sophisticated deepfakes. Defenders must prioritize updating their detection models to incorporate physical-model-based features like those utilized by VoiceRadar. This means moving towards detectors that can analyze nuances in translational, rotational, and speech tremor frequencies, rather than just high-level sonic characteristics.
  1. Combating Misinformation and Defamation: The ability of VoiceRadar to offer domain-agnostic and robust detection against adaptive attacks is invaluable in the fight against misinformation and defamation. Media forensics teams and fact-checking organizations can leverage such technology to definitively identify fabricated audio, protecting public discourse and individual reputations. The ease with which deepfakes can be generated for political warfare or character assassination necessitates tools that can keep pace with this threat.
  1. Real-Time Detection for Communication Platforms: Given the example of deepfakes being used in real-time Zoom calls, there is an urgent need for integration of advanced detection capabilities into communication platforms. VoiceRadar's speed, especially in analyzing the unique physical signatures, suggests potential for real-time or near-real-time deepfake detection, safeguarding online interactions from impersonation and manipulation.
  1. Future-Proofing Defenses: The observation that current deepfake generators would require "fundamental new changes" to bypass VoiceRadar's energy-based models offers a degree of future-proofing. Defenders can anticipate that the next generation of deepfake technology will likely attempt to mimic these physical cues, but this knowledge provides a roadmap for ongoing research and development in detection, ensuring that the defensive side remains proactive rather than reactive.

In essence, VoiceRadar provides a blueprint for a more robust and physically grounded approach to deepfake detection, urging defenders to look beyond superficial audio features and delve into the complex physics of sound to secure digital communications and identities.

Key Takeaways

  • Deepfake audio poses a severe and evolving threat, capable of bypassing voice authentication, spreading misinformation, and enabling defamation, with generation tools becoming increasingly accessible.
  • Existing deepfake detectors suffer from significant limitations, including domain dependency (TTS vs. STS), poor generalization to new deepfakes, and reliance on outdated evaluation datasets.
  • VoiceRadar introduces a novel, physical-model-based detection approach, leveraging micro-frequency and compositional analysis to differentiate human from synthetic speech by approximating how sound originates and propagates.
  • The system analyzes unique physical signatures, including translational, rotational, and speech tremor frequencies, which are subtle variations stemming from individual and environmental factors in human speech and are difficult for machines to replicate.
  • VoiceRadar utilizes energy-based models for detection, which are inherently more robust against adaptive adversarial attacks compared to traditional classification algorithms, as they aim to minimize an energy function rather than maximize a classification score.
  • VoiceRadar demonstrates domain-agnostic detection capabilities, performing consistently well across both text-to-speech and speech-to-speech deepfake generation methods, and proved resilient against simulated adaptive attacks, suggesting a higher bar for future deepfake bypass attempts.

About the Speaker(s)

The research on VoiceRadar was presented by Alexandro Pegoro at the NDSS Symposium. Alexandro Pegoro introduced himself as the presenter of this significant work, delving into the intricacies of voice deepfake detection using micro-frequency and compositional analysis. While his specific title and institutional affiliation were not explicitly stated in the transcript, his presentation demonstrated a deep technical understanding of the subject matter, reflecting his role as a dedicated researcher in the field of cybersecurity and artificial intelligence. The metadata for the talk also lists Kavita Kumari as a speaker, indicating a collaborative effort in the development and presentation of VoiceRadar.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

VoiceRadar is legitimate academic research with a defensible core idea — using physics-grounded micro-frequency features and energy-based models to generalize deepfake audio detection across TTS and STS domains. The contribution is real, but the framing is doing a lot of heavy lifting: 'drum vibration representation' and 'Doppler effect' sound more impressive in the abstract than they are in practice, and the adversarial robustness claims rely on the assumption that generative models won't adapt — which is a generous assumption to build a defense around.

Heather Calloway (CISO) — WEAK

Technically credible research on a real threat, but VoiceRadar never crosses the line from academic finding to operational guidance. The defensive implications section reads like retrofitted justification rather than actual analysis of how organizations would adopt or trust this detection approach.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025