Your voice confirms my identity
Ethan McKee-Harris (Security Consultant · Pan Security Group)
BSidesSF 2024 · Day 1
Overview
In his BSidesSF 2024 talk, "Your voice confirms my identity," Ethan McKee-Harris, a security consultant at Pan Security Group, delved into the alarming vulnerabilities of voice biometrics as a primary authentication mechanism. The presentation meticulously demonstrated how easily an individual's voice can be cloned using readily available AI tools and subsequently exploited to bypass voice authentication systems. McKee-Harris highlighted the critical disconnect between the perceived security and convenience of voice ID and its inherent weaknesses, which stem from the non-unique and publicly accessible nature of human voice.

Key moments
- 01:00 Rachel Tobac's voice cloning social engineering example
- 04:00 The fundamental problem: you can't 'reset' your voice
- 05:00 Banks claiming voice ID is unique and secure like a fingerprint
- 10:00 Demonstration of low-barrier voice cloning bypass against a real account
- 13:00 Practical guide to cloning your own voice with 11 Labs for $1
- 20:00 Cloning a target's voice from public digital presence (YouTube)
- 24:00 Argument: voice is inherently not a unique or secret medium for authentication
- 26:00 Mitigation strategies for businesses and consumers
Your voice confirms my identity
Speakers: Ethan McKee-Harris
Conference: BSidesSF 2024
YouTube: https://www.youtube.com/watch?v=aB6yuoyFW6A
Overview
In his BSidesSF 2024 talk, "Your voice confirms my identity," Ethan McKee-Harris, a security consultant at Pan Security Group, delved into the alarming vulnerabilities of voice biometrics as a primary authentication mechanism. The presentation meticulously demonstrated how easily an individual's voice can be cloned using readily available AI tools and subsequently exploited to bypass voice authentication systems. McKee-Harris highlighted the critical disconnect between the perceived security and convenience of voice ID and its inherent weaknesses, which stem from the non-unique and publicly accessible nature of human voice.
The talk is particularly pertinent in an era where voice authentication is increasingly adopted by financial institutions and other service providers, often touted as a simplified, secure alternative to traditional passwords and multi-factor authentication (MFA). McKee-Harris, drawing from his experience as both a security consultant and an open-source developer, provided a comprehensive analysis of the attack surface, from low-tech mimicking to sophisticated AI-powered voice cloning. He underscored the low barrier to entry for attackers and the high-value targets at risk, such as personal bank accounts and sensitive corporate data, making this a crucial topic for both consumers and businesses.
McKee-Harris's presentation served as a stark warning against over-reliance on voice biometrics, arguing that solutions not inherently unique or secret are destined to fail as authentication borders. He offered practical demonstrations of voice cloning, detailed the technical process, and provided actionable defensive strategies for both end-users and organizations. The core message resonated: while convenient, voice authentication, in its current widespread implementation, introduces significant and often underestimated security risks that demand immediate attention and a re-evaluation of its role in security architectures.
Background
▶ Watch: Rachel Tobac's voice cloning social engineering example (01:00)
The proliferation of voice authentication systems is driven by a strong industry push towards simplified user experiences. As Ethan McKee-Harris explained, traditional password flows, with their requirements for long, complicated, and unique passwords across multiple sites, are often seen as cumbersome. Voice authentication, by contrast, offers an appealingly simple alternative: "I know how to speak most of the time at least and it it's just simple." This perceived ease of use has positioned voice as the "next step forward" in user convenience.
However, McKee-Harris immediately countered this narrative by framing voice authentication as a "step back for user security." The fundamental problem lies in the immutability of a voice once compromised. Unlike a password, which can be easily reset, a voice cannot be changed. "If your voice gets breached, you know your voice gets recorded, your voice gets cloned, how are you going to change your voice? You can't just click reset voice." This inherent characteristic makes voice biometrics a high-risk authentication factor.
The talk provided compelling examples of this vulnerability already manifesting in the real world. McKee-Harris referenced a Darknet Diaries episode from April 2nd that featured Rachel, a social engineer. Rachel successfully cloned a journalist's voice using publicly available interviews and then spoofed her phone number to call a colleague. Posing as the journalist, she convinced the colleague to read out passport details from a company system, demonstrating a sophisticated social engineering attack enabled by voice cloning. Beyond this specific case, McKee-Harris noted a surge in news articles over the past four to five months concerning AI-generated robocalls and voice phishing, including the FCC banning AI-generated robocalls, indicating a growing awareness and concern about this threat vector.
Many financial institutions, particularly in Australia and New Zealand, have embraced voice ID. McKee-Harris cited ANZ as an example, where customers are prompted to use voice ID every time they call. These banks often market voice ID as highly secure, claiming it "makes it difficult for someone else to imitate you even if they've got a recording of your voice," and that "your voice print is unique just like your fingerprint." Crucially, some banks even allow customers to use their voice ID to bypass other security measures like MFA codes, presenting it as the sole authentication factor. This overconfidence in voice biometrics, coupled with the public availability of voice samples (e.g., from online interviews, social media, or even conference talks being recorded), creates a fertile ground for exploitation. The problem, as McKee-Harris articulated, is that businesses are often making a trade-off between security and usability, opting for a "sliding scale" of authentication that inherently incurs risk.
Key Findings
▶ Watch: Banks claiming voice ID is unique and secure like a fingerprint (05:00)
Ethan McKee-Harris's research unveiled several critical findings that challenge the security posture of voice authentication systems:
- Low Barrier to Entry for Attackers: The talk definitively demonstrated that bypassing voice authentication requires minimal technical expertise and resources. Attackers can achieve success with basic equipment such as a phone, laptop speakers, and a low-cost AI voice cloning service (e.g., 11 Labs for $1 USD). This accessibility significantly lowers the bar for malicious actors, making high-value targets more vulnerable.
- High Success Rate of Voice Cloning: Through both mimicking experiments and AI-powered voice cloning, McKee-Harris showed a high degree of success in bypassing authentication. His AI-cloned voice achieved an 80% success rate in getting past systems in some tests, and even confused colleagues who had worked with the target for years. This indicates that current voice biometric systems are often insufficient to distinguish between a legitimate user and a sophisticated clone.
- Inherent Vulnerability of Voice as a Biometric: The core finding is that voice is not an inherently unique or secret medium suitable for robust authentication. Unlike a fingerprint or retina scan, a voice has natural variability (e.g., due to tiredness, emotion, or environment) that systems must accommodate, creating a "sliding scale" of acceptance. This variability, combined with the ease of obtaining voice samples, makes it fundamentally insecure as a primary authentication factor.
- Lack of Actionable Logging for Fraud Detection: Voice authentication systems often provide inadequate logging for fraud investigation. When a cloned voice is used, the call may originate from a spoofed phone number and the audio itself will sound like the legitimate account holder. This leaves blue teams with very little evidence to differentiate between a legitimate transaction and a fraudulent one, potentially leading to victims being wrongly accused of self-fraud.
- Effectiveness of AI Cloning from Public Digital Presence: McKee-Harris successfully cloned the voice of his boss, Simon Howard, using only 5 minutes of publicly available YouTube video footage that was 5 years old. This demonstrated that even dated, non-pristine audio from an individual's digital footprint is sufficient to create a convincing voice clone capable of bypassing systems and deceiving human listeners.
- Voice Authentication Often Replaces, Not Augments, Security: A significant concern highlighted was that many implementations of voice ID replace existing security layers, such as MFA, rather than acting as an additional factor. This creates a single point of failure where a compromised voice grants full access, effectively reducing overall security.
These findings collectively paint a concerning picture of the current state of voice biometrics, urging a re-evaluation of its deployment and a shift towards more robust, multi-layered authentication strategies.
Technical Deep Dive
▶ Watch: Practical guide to cloning your own voice with 11 Labs for $1 (13:00)
The technical core of Ethan McKee-Harris's presentation revolved around two primary methods of bypassing voice authentication: mimicking and AI-powered voice cloning. He meticulously detailed the processes, tools, and effectiveness of each.
Mimicking and Threshold Vulnerabilities
McKee-Harris first explored the concept of mimicking and the inherent "sliding scale" nature of voice authentication. Unlike a password, which is either perfectly correct or incorrect, voice recognition operates on a spectrum of similarity. Businesses must set an arbitrary threshold, balancing the need for security (making it harder for attackers) with usability (ensuring legitimate users aren't locked out).
To illustrate this, he conducted an experiment with 15 people, collecting four audio clips from each. He then visualized the recognition results on a graph where each axis represented two audio clips from a given person. A green square indicated successful recognition, while red indicated failure. Initially, with a higher security threshold, some individuals (e.g., "person 14") were unable to authenticate themselves. However, by lowering the threshold by just 1%, person 14 could now access their own account, but critically, they could also access "person Zero's account." Further lowering the threshold by "a few more percentages" (approximately 2% total from the initial setting) resulted in a graph that was "all green," implying that a significant number of individuals could now access others' accounts. This demonstrated that even slight adjustments to the recognition threshold for usability drastically compromise security.
The practical attack setup for mimicking was remarkably simple: "sitting on a couch with a laptop on my lap playing audio from the speakers with my phone on call with the system on the couch next to me." He used an "entry-level microphone" to record audio, created voice clones, and played them through "laptop speakers" to a phone connected to the target system. This low-tech approach successfully bypassed a "real customer's account," highlighting the minimal resources required for a successful attack.
AI-Powered Voice Cloning with 11 Labs
The more advanced and potent method discussed was AI-powered voice cloning. McKee-Harris focused on using 11 Labs, a platform he found particularly effective for his Kiwi accent, though he noted a list of "about 15 to 20 platforms" (both online and offline models) are available.
The process for creating a high-fidelity voice clone using 11 Labs was outlined as follows:
- Training Data Acquisition: The goal is to obtain 3 to 5 minutes of "high fidelity audio" from the target. This means clear audio, from a single speaker, with no background interference or office noise. The better the input data, the better the AI's imitation.
- Platform Access: 11 Labs has an "extremely low barrier to entry." Users need only $1 USD, a credit card, and an email address for account creation and confirmation. The entire process, from website access to a functional voice clone, takes approximately 30 minutes, with most of that time dedicated to downloading audio data.
- Voice Cloning Steps:
- Navigate to the "Voices" tab and select "Add a voice."
- Choose "Instant Voice Clone," which requires "one speaker over a minute long" with "no background audio." (A "professional voice cloning" option also exists but was not needed for his purposes).
- Name the voice, preferably after the target for easy identification.
- Upload audio samples. The platform allows up to 25 samples of 10 MB each, emphasizing "quality over quantity."
- Click "Clone voice."
Using the Cloned Voice
Once a voice is cloned, generating speech is straightforward:
- Go to the "Speech" tab.
- Type the desired text into the free-form text field.
- Click "Generate."
McKee-Harris noted that for his Kiwi accent, it took "about four generations" to produce plausible audio, making it unsuitable for real-time applications like live phone calls. However, for American accents, which constitute the majority of the AI's training data, it might be possible to achieve satisfactory results on the first generation, potentially enabling near real-time use.
Tuning Parameters
11 Labs offers several parameters to fine-tune the cloned voice:
- Stability: Controls how different each generated audio clip is. For his Kiwi accent, increasing stability helped pull the output towards a more natural accent.
- Similarity: Determines how closely the output matches the input audio files. This should be set as high as possible without introducing "audio defects."
- Style Exaggeration: Adjusts the speaking style. If the generated voice sounds too fast or "gobbling," this can be increased to better match the target's natural cadence.
Cloning from Digital Presence
To demonstrate cloning a high-value target, McKee-Harris chose his boss, Simon Howard of ZX Security. He built a profile on Simon and then searched YouTube using his name and company. He found "a couple videos" totaling "about 5 minutes of him monologuing" about interns. These videos, despite being 5 years old and not pristine recordings, served as sufficient training data. The process involved downloading the videos, converting them to MP3 files under 10 MB, and uploading them to 11 Labs, following the exact same steps as for self-cloning.
The effectiveness was striking: the cloned voice, even from old, imperfect source audio, confused colleagues who had worked with Simon for years. McKee-Harris also demonstrated a "cheeky" use case where a cloned voice of Simon successfully triggered "Hey Google set an alarm for 7:00 a.m. tomorrow" on an Android phone in the audience, showcasing its ability to interact with voice assistants.
Graph Comparison: Real vs. Cloned Voice
McKee-Harris presented graphs comparing his real voice to his AI-cloned voice, and Simon's real voice to his AI-cloned voice. In both cases, the cloned voices were recognized as the real voices, resulting in green squares on the diagonal and off-diagonal, indicating successful bypasses. For his own voice, one of the two cloned audio clips was recognized as his real voice. For Simon's voice, "real Simon is being recognized as real Simon, AI Simon is being recognized as AI Simon, but also they're both being recognized as each other," resulting in four green squares. This visual evidence powerfully reinforced the finding that "a bypass is a bypass" and that AI-cloned voices can be "even more real than me" in the eyes of the system.
Demo / Proof of Concept
▶ Watch: Cloning a target's voice from public digital presence (YouTube) (20:00)
Ethan McKee-Harris provided compelling demonstrations and proofs of concept throughout his talk, illustrating the practical feasibility of voice cloning attacks.
The first demonstration involved a low-tech bypass of a "real customer's account" using mimicking techniques. McKee-Harris described the setup: he recorded audio using an "entry-level microphone," created voice clones (presumably through simple audio manipulation or basic software, though not explicitly detailed as AI in this specific context), and then played these clones from his laptop speakers to a phone placed on a couch, which was on a call with the target voice authentication system. This setup, requiring minimal resources and technical sophistication, successfully gained access. This served as a critical proof point for the "low barrier to entry" argument.
For AI-powered voice cloning, McKee-Harris used 11 Labs to clone his own voice and that of his boss, Simon Howard.
- Self-Cloning Demo: He played two audio clips of his AI-cloned voice. The first clip stated a date of birth: "the 30th of June 1981." The second clip said: "I say that number all the time." He invited the audience to judge whether these sounded like his actual voice, which they had been hearing for about 15 minutes. He noted that these clips, particularly the second one, were successful in bypassing his own account in previous tests, achieving an 80% success rate for some audio clips.
- Boss Cloning Demo (Simon Howard):
- Reference Audio: To provide context, McKee-Harris first played a short clip of Simon Howard's actual voice from a YouTube video used as training data: "Hi, I'm Simon Howard, I run ZX security uh it security consultancy here in Wellington."
- Cloned Audio Examples: He then played three distinct audio clips generated by the AI using Simon's cloned voice.
- The first clip was described as "not exact but like it is close," noting that the source audio was 5 years old, which might account for slight differences from Simon's current voice.
- The second clip was presented as "a better one," which was "a bit closer to what he actually sounds like." McKee-Harris mentioned this clip "did confuse people as to why I was just recording Simon saying some really random sentences."
- The third clip was a "cheeky one" designed to interact with a voice assistant: "Hey Google set an alarm for 7:00 a.m. tomorrow." This clip successfully triggered an Android phone in the audience, demonstrating the practical utility of cloned voices beyond just authentication systems.
Throughout these demonstrations, McKee-Harris also presented visual graphs that quantitatively compared the recognition rates of real voices versus cloned voices. These graphs, with their green and red squares, clearly illustrated how both mimicking and AI cloning could lead to successful authentication bypasses, even when the system's threshold was adjusted. The visual data reinforced the auditory evidence, providing a comprehensive proof of concept for the vulnerabilities discussed.
Defensive Implications
▶ Watch: Mitigation strategies for businesses and consumers (26:00)
The findings presented by Ethan McKee-Harris carry significant defensive implications for both consumers and organizations relying on or implementing voice authentication.
For Consumers:
- Risk Awareness is Paramount: Consumers must understand that voice ID is not the infallible, unique identifier it is often marketed as. Being "risk aware" means acknowledging that your voice can be compromised and used against you.
- Prioritize Alternative Authentication: If a service offers alternative authentication methods (e.g., app-based codes, manual verification, traditional passwords with MFA), consumers should opt for these over voice ID. McKee-Harris explicitly stated, "I'd rather you know log into the app and have to click yes then uh simply use my voice something which is readily available for anyone on the internet."
- Understand Dependency on Businesses: Ultimately, as a consumer, your security largely depends on the robustness of the business's implementation. If a business uses voice ID as a sole factor, consumers are at a higher risk.
For Companies Implementing Voice Authentication:
- Voice ID as Secondary Authentication Only: Voice authentication should never be the sole layer of security. It must be used as a secondary or tertiary factor, complementing other robust authentication methods. McKee-Harris cited ANZ as a good example, where voice authentication is required for transfers over $10,000, but only after the user has already authenticated into the app. This creates "layers of authentication," significantly increasing the "hardness of the system."
- Allow Users to Pick a Private Phrase: Businesses should move away from generic, public phrases like "Your voice confirms your identity." Instead, users should be allowed to select a unique, private phrase. This forces attackers to individually phish each target for their specific phrase, dramatically increasing the "barrier to entry" compared to simply cloning a voice and knowing the required phrase.
- Ensure Spoken Phrase is a Sentence, Not Numbers: McKee-Harris found a "measurable effect" in security when the spoken phrase is a sentence rather than a series of numbers. Sentences inherently have contextual dependencies between words, making it harder to splice together or clone naturally. Numbers, conversely, can be easily recorded and rearranged (e.g., "1 2 3 4 5" vs. "5 4 3 2 1") while still sounding plausible. Using sentences gives the underlying system a "better chance to pick up on the fraud."
- Increase Barrier to Entry and Use MFA: The goal should be to increase the difficulty for attackers, moving from bypass rates of "1 in 100, 1 in 1,000" to "1 in 10,000 or 1 in 100,000." Adding MFA on top of these measures acts as a "force multiplier," further enhancing security.
For Vendors of Voice Authentication Technology:
- Provide Means to Verify Content Origin: Vendors should develop and offer features that allow verification of whether audio content was produced by a human or generated synthetically. This could involve "cryptographic signing" of content, allowing users or systems to verify its authenticity using "existing measures such as... Hardware signing or public private keys."
- Actively Push for Adoption of Verification Checks: It's not enough to merely provide an API endpoint for verification. Vendors must "actually push for the adoption of these checks" by making them default in workflows and promoting their use in documentation. As McKee-Harris noted, "users will inherently tend towards whatever is default not necessarily whatever is secure."
By implementing these defensive strategies, organizations can mitigate some of the inherent risks associated with voice biometrics, transforming it from a potential single point of failure into a more secure, multi-layered authentication component.
Key Takeaways
- Your voice is not a secure border for authentication. It is inherently non-unique and easily obtainable, making it unsuitable as a primary security factor.
- Voice biometric technology is rapidly evolving and becoming more prevalent. Like face ID and fingerprint recognition, voice authentication is gaining widespread adoption, and its capabilities will only improve over time.
- Voice cloning is accessible, effective, and has a low barrier to entry. Attackers can use readily available AI tools and minimal resources to create convincing voice clones capable of bypassing authentication systems.
- Businesses must implement voice authentication as a secondary factor. It should never be the sole layer of security; instead, it should augment other robust authentication methods like MFA.
- Organizations should allow users to pick private, complex phrases for voice ID. Generic phrases are easily exploited. Using unique, sentence-based phrases increases the difficulty for attackers by requiring individual phishing.
- Consumers must be risk-aware and prioritize alternative authentication methods. If given the choice, opt for traditional, more secure methods over voice ID to protect personal and financial information.
About the Speaker(s)
Ethan McKee-Harris is a security consultant at Pan Security Group, based in New Zealand. He is also an active open-source developer and is known online by his handle "skas." His professional background provides him with experience on both sides of the "metaphorical security fence," offering a unique perspective on the vulnerabilities and defensive strategies in the cybersecurity landscape.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk delivers a brutal reality check on the security of voice authentication. McKee-Harris demonstrates with practical examples and real-world bypasses how easily voice biometrics can be defeated using readily available AI cloning tools. He exposes the fundamental flaws in relying on voice as a unique identifier, highlighting the low barrier to entry for attackers and the severe lack of logging for defenders. This isn't theoretical; it's a clear, actionable guide to a critical vulnerability.
Heather Calloway (CISO) — MUST SEE
This presentation is a stark and necessary warning for any organization relying on voice authentication. McKee-Harris meticulously demonstrates how easily voice biometrics can be compromised with readily available tools, exposing a fundamental flaw in the assumption of voice as a secure identifier. The implications for fraud, customer trust, and regulatory compliance are profound, demanding immediate executive attention to reassess risk ownership and implement robust defense-in-depth strategies.