Cloned Vishing : A case study

Katherine Rackliffe

DEF CON 33 · Day 1 · Main Stage

Overview

In an era where digital threats evolve at an unprecedented pace, social engineering tactics continue to be a primary vector for cybercriminals. Katherine Rackliffe's DEF CON talk, "Cloned Vishing: A Case Study," delves into a particularly insidious and emerging form of this threat: cloned vishing. This presentation, delivered by Rackliffe—a recent cybersecurity graduate from Brigham Young University and an incoming PhD student at the University of Wisconsin Madison specializing in phishing and human-computer interaction—sheds light on the alarming effectiveness of using AI-generated voice clones in targeted attacks.

Watch on YouTube

Visual summary for Cloned Vishing : A case study by Katherine Rackliffe
Visual summary for Cloned Vishing : A case study by Katherine Rackliffe

Key moments

  1. 0:00 Introduction and speaker background on vishing research
  2. 2:00 Defining cloned vishing and common attack scenarios
  3. 3:18 Rationale for studying rapid voice cloning attacks
  4. 4:21 Study methodology: cloning professors' voices for research
  5. 4:50 Four experimental groups and voice cloning tools used
  6. 6:10 Key study results: audience responses to vishing attacks

Cloned Vishing: A Case Study

Speakers: Katherine Rackliffe, Recent Graduate, Brigham Young University; PhD Student, University of Wisconsin Madison

Conference: DEF CON

YouTube: https://www.youtube.com/watch?v=JPCKg_3XLP8

Overview

In an era where digital threats evolve at an unprecedented pace, social engineering tactics continue to be a primary vector for cybercriminals. Katherine Rackliffe's DEF CON talk, "Cloned Vishing: A Case Study," delves into a particularly insidious and emerging form of this threat: cloned vishing. This presentation, delivered by Rackliffe—a recent cybersecurity graduate from Brigham Young University and an incoming PhD student at the University of Wisconsin Madison specializing in phishing and human-computer interaction—sheds light on the alarming effectiveness of using AI-generated voice clones in targeted attacks.

The core of Rackliffe's research investigates how susceptible individuals are to phone-based scams employing synthetic voices that mimic trusted individuals. She highlights a critical gap in academic literature, noting that while real-world incidents of voice cloning fraud are on the rise, empirical studies on their efficacy and public preparedness are scarce. This talk not only defines the threat landscape but also presents the findings of a controlled study designed to quantify the impact of cloned vishing and inform better defensive strategies against this rapidly advancing form of deception.

Background

▶ Watch: Introduction and speaker background on vishing research (0:00)

The landscape of social engineering attacks has steadily broadened, moving beyond simple email-based phishing to more sophisticated techniques. Vishing, or voice phishing, emerged as a natural progression, leveraging the human voice to trick victims into divulging sensitive information. However, the advent of readily available and increasingly sophisticated AI voice cloning technology has given rise to an even more potent threat: cloned vishing. This term, coined by Rackliffe, refers specifically to phone calls utilizing an AI-cloned voice to impersonate a known individual, significantly increasing the attack's credibility.

Common real-world manifestations of cloned vishing highlighted by Rackliffe include emotionally manipulative scams, such as calls from a "child" or "family member" claiming to be in distress or held for ransom. These attacks capitalize on urgency and emotional ties, often requiring only about five minutes of audio to create a convincing clone. Another prevalent scenario involves the impersonation of authority figures, like a boss or a CEO, demanding immediate action or information. Perhaps most concerning is the ability of voice clones to potentially fool voice verification systems used by financial institutions, a method famously employed by banks like Schwab where "my voice is my password." Rackliffe referenced high-profile incidents, such as a company director's voice being cloned to facilitate a massive financial theft, and numerous ransom attacks targeting everyday individuals.

The motivation behind Rackliffe's study stemmed from a critical observation: the cybersecurity threat landscape evolves at a breakneck pace, while academic research, constrained by rigorous methodologies and ethical reviews, often lags. Voice cloning attacks gained prominence only about two years prior to her study, yet a comprehensive academic paper on their effectiveness was non-existent. Her research aimed to fill this void, providing much-needed statistics and insights into how effective these voice clones truly are, whether people are prepared to identify them, and how to effectively communicate the tangible risks to a broader audience.

Key Findings

▶ Watch: Rationale for studying rapid voice cloning attacks (3:18)

Katherine Rackliffe's study yielded several critical insights into the efficacy of cloned vishing and public awareness. A primary finding, consistent with existing preliminary research, was the significant difficulty individuals face in distinguishing between an AI-cloned voice and a real human voice, particularly when encountered over a phone call where audio quality can naturally degrade. This indistinguishability was a cornerstone of the study's results, showing similar response rates between the AI-cloned voice and the actual professor's recording.

Surprisingly, the most potent indicator of suspicion for participants was not the voice's authenticity but rather the source of the communication. The fact that the malicious message came from an unknown phone number was consistently cited as the primary reason for distrust or non-response. This suggests that while the voice itself was convincing, fundamental security hygiene practices—like scrutinizing unknown contacts—still play a vital role in defense.

The study categorized responses into several groups: no response, reporting the message as phishing, responding with the requested ID, asking for verification, and sending information through other means (e.g., email instead of text). While a desirable "no response" was the most frequent outcome, a concerning number of students still provided their sensitive student ID in response to the scam. Furthermore, a group of "asked for verification" respondents demonstrated a critical but often insufficient level of suspicion, attempting to verify the sender through questions like recalling previous lecture topics.

Interestingly, the study found varying effectiveness across the different attack vectors:

  • Text messages garnered a "really high response rate," likely due to the ease of response for the Gen Z demographic and their familiarity with text communication.
  • Robo text-to-speech messages had a "really low response rate," yet some students still complied, rationalizing that the professor might be too busy to record a proper message.
  • Crucially, the AI-cloned voice and the real voice recording elicited "the same response rate," underscoring the AI's ability to deceive. In some instances, the voice clone even sounded more natural than the real professor, who was consciously reading from a provided script, leading some to mistakenly report the authentic recording as a scam due to its perceived artificiality.

Post-study interviews further solidified these findings: participants were unable to correctly identify whether a voice recording was real or cloned when played back to them. Their primary concern remained the unknown number, not the vocal timbre. This highlights a pervasive lack of public awareness regarding the capabilities of voice cloning technology, making individuals highly vulnerable to such attacks.

Technical Deep Dive

▶ Watch: Study methodology: cloning professors' voices for research (4:21)

The meticulous design of Katherine Rackliffe's study provided a robust framework for assessing the effectiveness of cloned vishing, navigating significant ethical hurdles in the process. The research began with extensive engagement with the university's Institutional Review Board (IRB), an ethics committee that initially expressed reservations about conducting a "phishing" study on students. After months of deliberation and ethics paperwork, a methodology was approved that prioritized participant safety and informed consent.

The target population for the study comprised university students, primarily from the Gen Z demographic, who were recruited through a vague consent process. Students consented to participate in a "random study about sending messages to people's phones," with an implied consent form integrated into the survey. To prevent any immediate association between the consent and the subsequent phishing attempt, a two-week delay was implemented before the malicious messages were sent. Furthermore, a comprehensive debriefing message was sent 48 hours after the phishing attempt, explaining the study's true nature and ensuring no psychological harm.

For the voice cloning component, professors from the university—specifically avoiding those in tech or computer science fields to ensure a diverse, representative sample—were asked to provide approximately one minute of generic audio, speaking about their research in a natural lecture style. This audio was then fed into a commercially available AI voice cloning model, specifically B Cloner, though Rackliffe also mentioned 11 Labs as another popular and effective option. She emphasized that these tools are generally free or very inexpensive, making them highly accessible to potential attackers.

The malicious message crafted for the study was designed to be plausible yet contain subtle red flags. It introduced the professor, stated an update to class registration, and then requested the student's ID number. The inherent suspiciousness lay in two key details: professors already have access to student ID numbers, and the message was delivered from an unknown phone number.

Four distinct experimental groups were established to compare different communication methods:

  1. Text Message: This served as a control group, representing traditional text-based phishing that participants frequently encounter.
  2. Robo Text-to-Speech: This group aimed to assess the impact of a clearly artificial, unsophisticated robotic voice.
  3. AI Cloned Voice: This was the core experimental group, using the AI-generated voice of the respective professor.
  4. Real Voice: In this group, the actual professor recorded the malicious message, providing a direct comparison to the AI clone. All voice recordings were sent as MP3 files or voice notes, allowing recipients to listen or read a transcript.

The responses were meticulously analyzed across these groups, categorizing actions such as ignoring the message, reporting it, complying with the request, or seeking further verification. A notable observation from interviews was that while some text messages were dismissed for sounding "too formal" or "not formal enough," no participant in the voice clone group cited the voice itself as a reason for suspicion. The unknown phone number consistently emerged as the most significant deterrent, highlighting a critical area for public education.

Demo / Proof of Concept

▶ Watch: Four experimental groups and voice cloning tools used (4:50)

During her DEF CON presentation, Katherine Rackliffe faced unfortunate Wi-Fi connectivity issues, which prevented her from conducting live audio demonstrations of the voice cloning technology. However, she vividly described what these demonstrations would have entailed and the striking impact they typically have.

The primary demonstration would have showcased a direct comparison between a real voice and its AI-cloned counterpart. Rackliffe emphasized that the two sound "exactly the same," making it "really hard to tell the difference." She had prepared a specific example featuring her own voice clone stating, "At Schwab, my voice is my password," to illustrate how easily such technology could bypass voice verification systems used by financial institutions.

While the live audio was unavailable, Rackliffe's description underscored a crucial point: the auditory indistinguishability of AI-cloned voices is not a theoretical concept but a practical reality. This lack of discernible difference forms the very foundation of the cloned vishing threat and directly informed the design and outcomes of her study, where participants consistently struggled to differentiate between genuine and synthetic voices. The absence of a live demo, therefore, did not diminish the core message but rather highlighted the challenge of conveying the nuanced reality of this advanced social engineering technique without direct auditory evidence.

Defensive Implications

▶ Watch: Key study results: audience responses to vishing attacks (6:10)

The findings from Katherine Rackliffe's "Cloned Vishing" study carry significant defensive implications for individuals, organizations, and even financial institutions reliant on voice biometrics. The most pressing need identified is a substantial increase in public awareness regarding the capabilities and prevalence of AI voice cloning. Many people remain unaware that their voice can be easily replicated and used for malicious purposes, making them highly susceptible to these advanced social engineering attacks. Comprehensive public education campaigns are crucial to inform individuals about this emerging threat.

Beyond general awareness, emphasis must be placed on educating individuals about red flags that transcend voice authenticity. The study clearly demonstrated that an unknown phone number was a far more significant indicator of suspicion than any perceived artificiality in the voice. Defenders should prioritize training individuals to scrutinize the context of a call or message: Does the request make sense? Is it coming from a known and trusted contact method? Is the request unusual or urgent? These non-vocal cues are currently more reliable indicators of a scam than trying to discern a cloned voice.

For organizations, mandatory cybersecurity training should be updated to include modules on cloned vishing. Rackliffe noted that master's students in her study showed higher awareness, which she speculated might be linked to mandatory workplace security training. Such training should focus on establishing internal protocols for verifying unusual requests, especially those involving financial transactions or sensitive data, using out-of-band communication channels (e.g., calling back on a pre-verified number, sending a separate email to a known address).

The vulnerability of voice verification systems is a critical concern. Banks like Schwab, which use "voice as a password," need to urgently re-evaluate the robustness of their biometric security measures against sophisticated AI voice clones. Relying solely on voice for authentication, without multi-factor authentication or robust liveness detection, poses a significant risk. Financial institutions and other entities employing voice biometrics should investigate advanced anti-spoofing technologies and consider alternative or supplementary verification methods.

Finally, while current AI voice agents might still exhibit tell-tale pauses or unnatural conversation flows, Rackliffe warned that this technology is rapidly improving. Attackers are currently leveraging readily available, low-cost tools like B Cloner that require minimal effort to produce convincing voice clones, making these attacks highly accessible. Defenders should anticipate a future where AI can engage in natural, real-time conversations over the phone, necessitating even more robust non-vocal verification strategies. The call to action is clear: proactive education and layered security measures are essential to mitigate the growing threat of cloned vishing.

Key Takeaways

  • AI voice cloning is highly effective and difficult to detect: The study unequivocally showed that individuals struggle to distinguish between AI-generated voices and real human voices, particularly over the phone.
  • "Unknown number" is a primary red flag: Participants were more suspicious of an unknown phone number than the voice's authenticity, highlighting the importance of contextual clues over auditory ones.
  • Public awareness of voice cloning is dangerously low: Many individuals are unaware of the capabilities of AI voice cloning, making them highly vulnerable to this emerging form of social engineering.
  • Voice verification systems are vulnerable: Banks and other institutions relying on voice biometrics for authentication face a significant threat from AI voice clones, necessitating a re-evaluation of their security protocols.
  • Low-cost tools make these attacks accessible: The ease and affordability of creating convincing voice clones with tools like B Cloner lower the barrier to entry for attackers, increasing the prevalence of cloned vishing.
  • Education and critical thinking are paramount: Effective defense against cloned vishing requires widespread public education on the threat, combined with training on critical thinking and multi-factor verification practices for unusual requests.

About the Speaker(s)

Katherine Rackliffe is an emerging voice in the field of cybersecurity research, specializing in the intersection of human behavior and digital threats. At the time of this DEF CON talk, she had recently completed her bachelor's degree in cybersecurity from Brigham Young University. Rackliffe is poised to further her contributions to the academic community, as she is commencing her PhD studies at the University of Wisconsin Madison, where her research will continue to focus on the critical areas of phishing and human-computer interaction. Her work exemplifies a dedication to understanding and mitigating the evolving landscape of social engineering attacks.

Reviews

Dr. Zero (Offensive Security Researcher) — WEAK

A competent undergrad research project that doesn't belong at DEF CON — or at least not in a full talk slot. The study design is solid for an IRB-constrained academic exercise, but the findings are thin, the sample is narrow, and the core result ('AI voice clones are convincing') has been common knowledge in the security community for two-plus years before this was presented.

Heather Calloway (CISO) — WEAK

Rackliffe is doing legitimate academic work on a real and growing threat, and the research design shows care and rigor. But this talk stays inside the study and never makes the institutional leap — it tells you voice cloning works, not what organizations or decision-makers should actually do about it.

→ Top-rated talks at DEF CON 33

All talks from DEF CON 33