One Breach to Crack 'Em All! Insights from Password Breaches - Jarkko Vesiluoma
Jarkko Vesiluoma (Principal Offensive Security Lead · Elisa)
Disobey 2026 · Main Stage
Overview
Jarkko Vesiluoma's Disobey talk, "One Breach to Crack 'Em All! Insights from Password Breaches," delves into the profound lessons that can be gleaned from analyzing vast datasets of compromised credentials. Drawing its title from a clever Lord of the Rings reference, the presentation underscores the disproportionate power a single compromised password can hold over multiple user accounts. Vesiluoma, a seasoned offensive security lead, researcher, and red teamer, embarked on a multi-year research endeavor to dissect password breach data, aiming to understand the underlying patterns in human password creation and their implications for both offensive and defensive cybersecurity strategies.

Key moments
- 0:00 Introduction and research motivation
- 2:00 Scope of breach data analysis
- 3:20 Data collection and analysis pipeline
- 4:50 Why old breach data remains relevant
- 6:40 Common password lengths and policy influence
- 8:50 Common password construction patterns and dictionary words
- 12:10 Human typing habits make passwords predictable
One Breach to Crack 'Em All! Insights from Password Breaches
Speakers: Jarkko Vesiluoma, Principal Offensive Security Lead, Elisa
Conference: Disobey
YouTube: https://www.youtube.com/watch?v=qnqwIRAY9jw
Overview
Jarkko Vesiluoma's Disobey talk, "One Breach to Crack 'Em All! Insights from Password Breaches," delves into the profound lessons that can be gleaned from analyzing vast datasets of compromised credentials. Drawing its title from a clever Lord of the Rings reference, the presentation underscores the disproportionate power a single compromised password can hold over multiple user accounts. Vesiluoma, a seasoned offensive security lead, researcher, and red teamer, embarked on a multi-year research endeavor to dissect password breach data, aiming to understand the underlying patterns in human password creation and their implications for both offensive and defensive cybersecurity strategies.
The core of Vesiluoma's research involved an extensive analysis of two terabytes of raw breach data, meticulously filtered to focus on Swedish and Finnish credentials. This deep dive revealed critical insights into password length distributions, common character patterns, the prevalence of dictionary words and l33t speak, and the alarming rates of password reuse across different services. The findings serve as a stark reminder of the persistent vulnerabilities arising from human behavior in password selection and management, providing a data-driven foundation for understanding how attackers efficiently compromise accounts and, more importantly, how defenders can build more resilient security postures.
This article explores Vesiluoma's methodology, key discoveries about password anatomy and cracking efficiency, and the actionable defensive strategies proposed to mitigate the risks illuminated by his comprehensive study. The research highlights that despite decades of cybersecurity awareness campaigns, fundamental human tendencies continue to render a significant portion of passwords vulnerable to even basic attack techniques, making continuous adaptation and robust defense mechanisms paramount in the ongoing battle against credential-based attacks.
Background
▶ Watch: Introduction and research motivation (0:00)
The genesis of Jarkko Vesiluoma's extensive research began approximately three to four years prior to his Disobey presentation, initially sparked by an ambition to create heatmaps illustrating whether passwords were generated randomly or by human interaction. This initial curiosity quickly "escalated," leading to a much broader and deeper investigation into the wealth of information contained within password breach data. The ultimate objective evolved into answering a critical question: "What can breach data teach us about attacking and defending?"
Vesiluoma's research corpus was formidable, starting with around two terabytes of raw breach data collected from various internet sources, torrents, and trusted contacts. From this vast pool, a meticulous filtering process was applied to extract relevant information. The filtering targeted specific data types, acknowledging that some breach data included HTTP URLs, Android protocol tags, or FTP origins, which needed careful handling. The primary focus for analysis was on Swedish and Finnish passwords, resulting in a dataset comprising 240,000 unique domains and a total of 2.6 million Swedish credentials alongside 710,000 Finnish credentials. The disparity in numbers, Vesiluoma mused, could potentially suggest varying levels of cybersecurity awareness between the two populations, though he emphasized that this was an open question.
The research pipeline was structured into distinct phases:
- Data Collection: Involving extensive online research and acquisition of breach datasets.
- Filtering: Cleaning and categorizing the raw data based on its source and format.
- Ingestion: Loading the filtered password data into a database and flat files. The database integration was crucial, enabling faster research through domain-name or keyword-based queries compared to sifting through plain text files.
- Analysis: Executing numerous Python analysis scripts to derive patterns and statistics.
- Visualization and Insights: Presenting the findings in an understandable and actionable format.
A central premise underpinning the entire research was the enduring relevance of old breach data. Vesiluoma highlighted several reasons why compromised credentials, even years old, continue to pose a significant threat:
- Persistent User Habits: "People don't change their passwords," Vesiluoma stated, citing evidence from multiple breaches containing users employing the exact same password or making trivial modifications (e.g., adding a "1" or an exclamation mark).
- Pattern Persistence: Even when users do change passwords, the modifications often follow predictable patterns, such as incrementing a year (e.g., "summer2020" to "summer2021").
- Credential Stuffing: Breach data provides attackers with a ready-made arsenal to perform credential stuffing attacks against hundreds of other services, exploiting the widespread practice of password reuse.
- Compilation Aggregation: Newer breach compilations frequently incorporate older passwords, ensuring that compromised credentials from years past continue to resurface in threat intelligence platforms, perpetuating their utility for attackers.
- Multi-Batch Cracking: The research showed that 3% of users explicitly reuse the exact same password across different services, while the most reused single password was found 19,000 times with different users. This highlights the scale of the reuse problem and the potential impact of a single breach.
This foundational understanding of user behavior and the lifecycle of breach data was critical for Vesiluoma to illustrate why the problem of weak passwords and credential reuse remains a persistent and evolving challenge in cybersecurity.
Key Findings
▶ Watch: Data collection and analysis pipeline (3:20)
Jarkko Vesiluoma's analysis of millions of Swedish and Finnish credentials unveiled a series of critical findings regarding password construction, user habits, and attacker efficacy. These insights paint a clear picture of human predictability in password selection, which attackers consistently exploit.
Password Length and Anatomy:
The research revealed a strong correlation between password length and historical password policies. Users tend to adhere to minimum requirements, leading to distinct peaks in length distribution:
- 6-character passwords showed a slight bump, indicating a past policy minimum.
- 8-character passwords were the most common length across the entire dataset, aligning with prevalent security requirements.
- A subsequent dip at 9 characters and then a rise at 10-character passwords suggested a more recent policy shift.
- Alarmingly, 55% of all Swedish and Finnish passwords in the breach data were between 6 and 9 characters long, making them highly susceptible to cracking.
Delving into the anatomy of these passwords, Vesiluoma identified common shapes:
- The most frequent pattern was eight lowercase characters.
- Following closely were six lowercase characters.
- Then, six lowercase characters followed by two digits.
- Finally, nine lowercase letters only.
These patterns demonstrate a preference for simplicity and minimal complexity.
Predictable Patterns and Dictionary Usage:
User behavior consistently introduced easily guessable elements:
- Many passwords started with "MSA."
- A staggering 442,000 passwords ended with the number "1," reflecting a common, albeit insecure, method of changing a password (e.g., "password" becoming "password1").
- Plain numbers, especially birth dates or years, were also frequently observed.
- 43% of all passwords contained some form of recognizable dictionary word, making them vulnerable to dictionary attacks.
- L33t speak (e.g., replacing 'E' with '3' or 'A' with '4') was used in 64% of passwords containing numbers or substitutions, but this "trick" offers minimal security against modern cracking tools.
Human-Typed Passwords and Keyboard Patterns:
Vesiluoma highlighted that "your fingers also betray you." Human typing habits lead to predictable sequences:
- Keyboard walks (e.g., "qwerty," "12345") were common.
- Sequential characters and repeated characters appeared frequently.
- Date patterns and year patterns were prevalent.
- Analysis of the "next predictable character" showed clear tendencies: after '1', the most likely characters are '2', '1', or '9'. After 'Q', it's 'W', 'A', or 'U'. This predictability significantly reduces the key space for attackers using Markov-based cracking models.
Alarming Reuse and Personal Information:
Perhaps the most concerning findings related to password reuse and the inclusion of personal data:
- 4% of users in the breach data used their email address as their password. An even larger 11% of all passwords contained the user's email address in some form.
- Users frequently used their username, or variations like "username1" or "username123," as their password.
- The single most common password found was "123456," shared by 19,000 users.
- Overall, 3% of exact passwords were reused across different services, and most individuals reused two to three passwords across all their accounts.
- Passwords also revealed deeply personal information, often containing references to emotions (love, hate, anger), identity, family, hobbies, pop culture, and even confessions ("I'm poor," "I'm rich"). Some passwords even leaked corporate information or references to specific adult services, providing attackers with valuable context.
Summary Statistics:
Vesiluoma concluded with a powerful summary of his quantitative findings:
- 48 million passwords were analyzed in total.
- 26% were exactly eight characters long.
- 43% contained dictionary words.
- 64% used l33t substitution or numbers.
- 90,000 users shared the same password (likely "123456").
- 4% used their email as their password.
- 3% reused passwords across different breaches and services.
- 36,000 unique dictionary words were found.
- 285,000 passwords contained the numbers "1, 2, 3" (usually at the end).
- 442,000 passwords ended with the number "1."
These findings collectively underscore the critical need for improved password hygiene and robust defensive mechanisms, as human tendencies consistently create predictable and exploitable vulnerabilities.
Technical Deep Dive
▶ Watch: Why old breach data remains relevant (4:50)
The technical deep dive of Vesiluoma's talk focused on the efficiency of various password cracking techniques and the strategies attackers employ to maximize their success with minimal computational effort. The analysis highlighted how understanding password composition directly influences the effectiveness of cracking methodologies.
Vesiluoma presented a compelling chart illustrating the brute-force times for various hash types, specifically MD5 and SHA-1, without the aid of wordlists or rules. These figures were based on a single NVIDIA RTX 4070 Ti Super GPU, demonstrating the raw power available to attackers. For instance:
- A 6-character MD5 hash could be brute-forced through all possibilities in approximately 1.6 minutes.
- An 8-character MD5 hash containing all lowercase, uppercase, numbers, and special characters would take an estimated 3.6 months to crack.
- A SHA-1 hash of the same complexity would require around 11 months.
However, Vesiluoma stressed that these "plain" brute-force times are largely theoretical for real-world attacks. He stated that "with rules these numbers can be a lot lower." Attackers rarely resort to pure brute force unless absolutely necessary, preferring a staged approach that leverages known human password patterns.
The attacker's perspective was broken down into a highly efficient, multi-stage process:
- Dictionary Attack (Stage 1): This is the first and quickest step. By simply comparing hashed passwords against a comprehensive list of common words and phrases, attackers can immediately crack approximately 20% of all passwords. This stage is incredibly fast and resource-efficient.
- Dictionary Rules (Stage 2): Building upon dictionary attacks, this stage applies rules such as l33t speak substitutions (e.g., 'a' to '4', 'e' to '3'), suffix rules (e.g., adding "1" or "!" to dictionary words), and other common permutations. This quickly adds another 30% of passwords to the cracked pile, bringing the total to about 50% with minimal effort.
- Mask Attacks (Stage 3): This technique targets dominant password shapes identified in the breach data. Attackers use masks to define specific patterns of character types. Common masks include:
?l?l?l?l?l?l(six lowercase characters)?l?l?l?l?l?l?d?d(six lowercase characters followed by two digits)?l?l?l?l?l?l?l?l(eight lowercase characters)
Mask attacks are slower than dictionary-based methods but are highly effective against predictable structures, significantly reducing the keyspace compared to full brute force. Vesiluoma explicitly mentioned the use of ?s for special characters in these masks.
- Word + Year/Name/Date (Stage 4): This stage combines known information about the target (e.g., from their email address or username) with common patterns like years or dates. For example,
username2023orfirstnamebirthyear.
- Brute Force (Stage 5): Only as a last resort, after all other more efficient methods have been exhausted, would an attacker resort to pure brute force. This is the slowest and most resource-intensive method.
Vesiluoma emphasized that an astounding 90% of all passwords in his dataset were resolved from their hashes using only stages one, two, and three. This underscores the power of targeted, rule-based cracking. For red teamers and attackers, "the word lists and rules are the king."
A key technical concept highlighted was Markov-based cracking. By analyzing the probability of the next character given the preceding one (e.g., after 'Q', 'W' is highly probable), attackers can drastically reduce the key space of possible passwords, making cracking significantly faster and more efficient than traditional brute-force methods. The research also revealed that only 29 characters are needed to cover half of all passwords, and that rare letters can be excluded from brute-force attempts to speed up the process, as they account for only the last 10% of passwords.
The talk concluded this section by reiterating that old breaches continue to pay dividends for attackers due to password reuse, exploitable suffix habits, and the effectiveness of Markov probability in reducing the key space. This technical understanding of attacker methodology is crucial for developing effective defensive strategies.
Demo / Proof of Concept
▶ Watch: Common password construction patterns and dictionary words (8:50)
The talk primarily focused on presenting the aggregated findings and analytical insights derived from an extensive password breach dataset, rather than demonstrating a live proof-of-concept or a specific cracking tool. Jarkko Vesiluoma's presentation was a research-oriented deep dive into the patterns and statistics uncovered from millions of compromised credentials, supported by visualizations of data distributions and attack effectiveness. While the methodologies for data collection, filtering, ingestion, and analysis (involving Python scripts and database queries) were described, the session did not include a live demonstration of these processes or a specific tool in action. The emphasis was on the results and implications of the analysis rather than the real-time execution of the cracking or data processing techniques.
Defensive Implications
▶ Watch: Human typing habits make passwords predictable (12:10)
Jarkko Vesiluoma's extensive analysis of password breaches provides a clear roadmap for defenders seeking to bolster their security against credential-based attacks. The insights gleaned from attacker methodologies and common user weaknesses directly translate into actionable defensive strategies.
- Ban Bad Passwords at Registration and Login: Organizations should implement mechanisms to prevent users from setting or using easily guessable or previously compromised passwords. Services like Have I Been Pwned offer APIs that can be integrated into registration and password change forms to check if a proposed password has appeared in known breaches. This forces users to choose stronger, unique credentials from the outset.
- Enforce Strong Password Policies:
- Minimum Length: Vesiluoma advocates for a minimum password length of 12 characters. While an 8-character policy "barely functions" (26% of passwords analyzed were 8 chars), 12 characters offers a significantly stronger baseline. He also acknowledges that even 12 characters might not be sufficient in the long term, especially with advancements in quantum computing.
- Password Phrases: The most highly recommended approach is to encourage or enforce the use of password phrases. These are long, memorable sequences of words (e.g., "correct horse battery staple") that are far more secure than short, complex passwords like "Password1!". They are easy for humans to remember but incredibly difficult for attackers to brute force.
- Complexity Requirements: While seemingly beneficial, traditional complexity requirements (uppercase, lowercase, digit, special character) often lead users to create predictable patterns like "Password1!" or "CompanyName!". Defenders should focus more on length and uniqueness rather than rigid complexity rules that users subvert predictably.
- Deploy Multi-Factor Authentication (MFA) Everywhere: This is arguably the most critical defensive measure. MFA effectively "kills the credential stuffing" threat. Even if an attacker obtains a user's password, the requirement for a second factor (e.g., a code from an authenticator app, a hardware token) prevents unauthorized access. MFA should be deployed across all sensitive services and, ideally, universally.
- Implement Credential Stuffing Detection and Prevention:
- Detection Mechanisms: Organizations need systems that can detect suspicious login patterns indicative of credential stuffing, such as multiple failed login attempts from different usernames originating from the same IP address or a sudden surge in login attempts across various accounts.
- Login Rate Limiting: Implementing rate limiting on login attempts effectively slows down brute-force and credential-stuffing attacks, making them impractical.
- CAPTCHA: Integrating CAPTCHA challenges into login flows can help distinguish between human users and automated bots attempting to stuff credentials.
- Promote and Mandate Password Managers: Vesiluoma strongly believes that password managers should be "mandatory." They enable users to generate and store unique, strong passwords for every single service, eliminating the dangerous practice of password reuse. Education and provision of password manager solutions can drastically improve organizational security hygiene.
- Continuous Credential Monitoring: Organizations should actively monitor for their employees' credentials appearing in public breaches. Services that track compromised credentials can alert companies when their users' accounts are at risk, allowing for proactive password resets and account security checks.
- Educate Users on Password Security: Reiterate the importance of never putting anything personal, sensitive, or easily guessable into a password. As Vesiluoma succinctly put it, "never put anything in password you would want on the billboard." This includes family names, hobbies, birthdates, or corporate system names.
By adopting these layered defensive strategies, organizations can significantly reduce their attack surface and protect against the pervasive threats highlighted by the analysis of real-world password breaches.
Key Takeaways
- Password Reuse is Rampant and Risky: A significant portion of users (3% exact reuse, most reusing 2-3 passwords) employ the same credentials across multiple services. This makes credential stuffing highly effective for attackers, turning one breach into many compromised accounts.
- Human-Chosen Passwords Are Highly Predictable: Users consistently fall into predictable patterns, including using dictionary words (43%), l33t speak (64%), sequential characters, keyboard walks, personal information, and trivial modifications like adding "1" to the end (442,000 passwords).
- Attackers Prioritize Efficiency Over Brute Force: Attackers rarely use pure brute force. Instead, they leverage staged attacks: dictionary attacks (cracking 20% immediately), dictionary rules (adding 30%), and mask attacks (targeting common shapes). These methods resolve 90% of passwords with far less effort than full brute force.
- Strong Defenses Require Layered Approaches: Effective defense against password-based attacks necessitates a combination of strategies, including Multi-Factor Authentication (MFA) everywhere, enforcing long password phrases (12+ characters), using password managers for unique credentials, rate limiting logins, and banning known bad passwords.
- Breach Data Offers Invaluable Insights: Analyzing real-world password breach data provides critical intelligence for both offensive and defensive security. It reveals common vulnerabilities, attacker methodologies, and the persistent human factors that undermine password security, guiding the development of more effective security measures.
- Personal Information in Passwords is a Major Leak: Passwords often contain deeply personal details, emotions, or even corporate information, providing attackers with valuable context for social engineering or further exploitation. Users must be educated to never embed such information in their passwords.
About the Speaker(s)
Jarkko Vesiluoma is a Principal Offensive Security Lead at Elisa, a major telecommunications and digital services company. He identifies himself as a security researcher, white hacker, and red teamer. Vesiluoma is driven by a passion for "tinkering and building things" and "trying to break everything around me technically" whenever he has the opportunity. His work and personal interests clearly align with the detailed and analytical nature of his research into password breaches, showcasing his expertise in understanding and exploiting security vulnerabilities from an attacker's perspective to inform robust defensive strategies.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Competent, data-driven research with a respectable corpus (48M passwords, 2TB raw data, regional filtering) that quantifies what most practitioners already know directionally. The regional focus on Swedish/Finnish credentials is a modest differentiator, but the core findings — dictionary words dominate, 8-char passwords are common, MFA fixes credential stuffing — are well-trodden ground that won't surprise anyone who's run hashcat for a weekend.
Heather Calloway (CISO) — WEAK
Solid empirical work on password behavior with a clear attacker methodology, but it never climbs to the institutional level where the real problem lives. The defensive recommendations are technically correct and mostly generic, and the talk misses the organizational accountability question entirely.