"Watching over the shoulder of a professional": Why hackers make mistakes and how they fix them

Irina Ford, Ananta Soneji, Faris Bugra Kokulu, Jayakrishna Vadayath, Zion Leonahenahe Basque, Gaurav Vipat

IEEE Symposium on Security and Privacy 2024 · Day 1 · Continental Ballroom 5

Overview

This talk, presented by Irina Ford and her colleagues from Arizona State University, delves into the often-overlooked aspect of human error in vulnerability research and exploitation. Titled "Watching over the shoulder of a professional," the research meticulously analyzes the mistakes made by skilled hackers during Capture The Flag (CTF) style binary exploitation challenges, focusing specifically on memory corruption vulnerabilities. The core premise is that despite significant advancements in automation, vulnerability exploitation remains a highly manual, time-consuming, and inherently error-prone process.

Watch on YouTube

Visual summary for "Watching over the shoulder of a professional": Why hackers make mistakes and how they fix them by Irina Ford, Ananta Soneji, Faris Bugra Kokulu, Jayakrishna Vadayath, Zion Leonahenahe Basque, Gaurav Vipat
Visual summary for "Watching over the shoulder of a professional": Why hackers make mistakes and how they fix them by Irina Ford, Ananta Soneji, Faris Bugra Kokulu, Jayakrishna Vadayath, Zion Leonahenahe Basque, Gaurav Vipat

Key moments

  1. 0:00 Introduction: Why hacking is time-consuming and error-prone
  2. 2:00 Defining the research questions on hacker mistakes
  3. 2:45 Methodology: Analyzing YouTube CTF challenge streams
  4. 3:55 Introducing the 6-step Mistake Anatomy framework
  5. 4:55 Categorizing 117 identified hacker mistakes
  6. 9:00 Quantifying the time impact of different mistake types
  7. 10:10 Non-technical factors and mitigation strategies for mistakes

"Watching over the shoulder of a professional": Why hackers make mistakes and how they fix them

Speakers: Irina Ford, PhD Student, Arizona State University; Ananta Soneji; Faris Bugra Kokulu; Jayakrishna Vadayath; Zion Leonahenahe Basque; Gaurav Vipat

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=ju1g2IaTLBg

Overview

This talk, presented by Irina Ford and her colleagues from Arizona State University, delves into the often-overlooked aspect of human error in vulnerability research and exploitation. Titled "Watching over the shoulder of a professional," the research meticulously analyzes the mistakes made by skilled hackers during Capture The Flag (CTF) style binary exploitation challenges, focusing specifically on memory corruption vulnerabilities. The core premise is that despite significant advancements in automation, vulnerability exploitation remains a highly manual, time-consuming, and inherently error-prone process.

The motivation for this study stems from the critical role that attack scenarios play in modern vulnerability disclosure programs, such as the Google Vulnerability Reward Program, which mandates a valid exploit to qualify for a reward. By systematically observing and categorizing the types of mistakes hackers make, understanding their impact on the overall hacking process, identifying non-technical causal factors, and documenting mitigation strategies, the researchers aim to improve the efficiency of vulnerability research and enhance the quality of cybersecurity education. This work provides invaluable insights into the cognitive processes and challenges faced by security professionals in the demanding field of exploit development.

Background

▶ Watch: Introduction: Why hacking is time-consuming and error-prone (0:00)

The landscape of modern software systems is characterized by immense complexity, rendering them susceptible to a myriad of vulnerabilities. While the discovery of these flaws is a foundational step in improving security, their effective remediation often hinges on demonstrating their exploitability. As highlighted by the Google Vulnerability Reward Program, a concrete attack scenario is frequently a prerequisite for a vulnerability report to be considered critical and eligible for a bounty. This underscores the practical necessity of exploit development—not for malicious intent, but to unequivocally illustrate the severity and potential impact of a vulnerability, thereby compelling developers to address it with urgency.

Despite this critical need, the process of both vulnerability discovery and exploitation remains predominantly a manual endeavor, consuming significant time and resources. The researchers hypothesized that, much like any complex human-driven process, hacking is prone to errors, which inevitably lead to inefficiencies. To test this intuition and answer their core research questions—regarding the types, impacts, causes, and mitigation strategies of hacker mistakes—the team embarked on an observational study.

Recognizing the practical difficulties of conducting controlled experiments with highly skilled hackers, the researchers leveraged the wealth of publicly available content on YouTube. They specifically targeted streams where security content creators engaged in solving CTF-style binary exploitation challenges, ensuring that the chosen videos featured hackers tackling previously unseen or unsolved challenges (blind solving). This rigorous selection process yielded a dataset of 30 videos containing 31 distinct binary exploitation challenges, all centered around memory corruption vulnerabilities. From these videos, a total of 124 problems were identified, with 117 subsequently categorized as distinct mistakes, providing a rich empirical basis for their analysis.

Key Findings

▶ Watch: Methodology: Analyzing YouTube CTF challenge streams (2:45)

The study's findings offer a granular understanding of the challenges faced by hackers, categorizing mistakes and quantifying their impact. A central contribution is the Mistake Anatomy Framework, a modified version of Rasmuson's 1974 decision ladder template, adapted to describe the information processing and decision-making steps during a hacking process. This framework delineates six stages: Origination (where the mistake is made), Manifestation (when its effects become visible), Detection (when the hacker notices it), Recognition (when they attempt to understand it), and Implementation (when they execute a solution). A crucial aspect of this model is the iterative loop between Recognition and Implementation, signifying that hackers often cycle through understanding and applying fixes until a mistake is fully resolved.

Using this framework, the researchers identified 117 mistakes, classifying them into four primary categories and also noting seven distinct technical issues that, while not mistakes, significantly impeded progress:

  1. Exploit Implementation Mistakes: The most frequently occurring type, these relate directly to the crafting of the exploit payload or logic. Subtypes included mistakes related to leakage (intent to read specific memory values), write what (intent to write data), write where (intent to write to a specific location), and write what where (a combination of both).
  2. Programming Miscellaneous Mistakes: These encompass general programming errors, frequently involving Python (the language often used for exploit scripting), exception handling, incorrect exploit logic, user input issues, configuration errors, and even incorrect function analysis.
  3. Strategy Mistakes: These occurred when hackers pursued an incorrect conceptual approach, leading them to abandon their current line of thinking and rethink their entire strategy. Examples include using a 4-byte payload when more was needed, writing an incompatible exploit, or failing to trigger the exploit mechanism correctly.
  4. Tooling Mistakes: Errors related to the use of security tools such as disassemblers, compilers, and debuggers. These included debugger oversight, incorrect process control commands, setting wrong breakpoints, or encountering permission-related issues with tools.
  5. Technical Issues: Seven distinct problems that were not classified as mistakes but significantly stalled progress. These included difficulties with the target application itself, missing libraries or dependencies, and random or unexpected program behavior.

Quantifying the impact, the study revealed that hackers spent an average of 40% of their total hacking time grappling with and rectifying mistakes, with individual cases ranging from 3% to a staggering 75%. Exploit implementation mistakes were not only the most frequent but also the most time-consuming to resolve, sometimes taking almost an hour. Strategy mistakes were the second most time-intensive. While programming and tooling mistakes had less individual impact, technical issues, despite their rarity, could be profoundly disruptive, with one outlier consuming over 1.5 hours of a hacker's time.

Non-technical factors contributing to mistakes included forgetting critical elements, lack of attention, and unawareness of the correct approach. To mitigate and avoid mistakes, hackers primarily employed trial-and-error approaches, consulted online resources, and utilized standard exploit development techniques. Strategic handling of mistakes involved efficiency-increasing tactics such as brute-forcing, precise calculations, or taking shortcuts, as well as leveraging preferred tools and engaging in online interactions with their audience for collaboration and feedback.

Technical Deep Dive

▶ Watch: Introducing the 6-step Mistake Anatomy framework (3:55)

The technical depth of this research is rooted in its systematic classification of errors and the specific examples provided for each category, offering a window into the nuanced challenges of exploit development. The Mistake Anatomy Framework serves as the analytical backbone, detailing the cognitive process from error inception to resolution. This framework is particularly insightful because it highlights the iterative nature of debugging and problem-solving, where hackers often revisit understanding (Recognition) before attempting new solutions (Implementation).

The most prevalent and impactful category, Exploit Implementation Mistakes, directly reflects the complexities of manipulating memory and program control flow. These were subdivided based on the hacker's intent:

  • Leak: Mistakes made when trying to read specific memory values, crucial for bypassing Address Space Layout Randomization (ASLR).
  • Write What: Errors in determining the data to be written into memory.
  • Write Where: Errors in identifying the correct memory address for writing.
  • Write What Where: A combination of errors in both the data and its target address.

Specific examples illustrate these pitfalls. One instance involved a hacker writing a pointer with a constant value, which subsequently caused a program crash upon dereferencing this invalid address. Another critical example highlighted a hacker mistakenly writing to the Procedure Linkage Table (PLT) instead of the intended Global Offset Table (GOT). Both PLT and GOT are crucial structures in ELF binaries used for dynamic linking; overwriting the GOT allows an attacker to hijack function calls, making the distinction vital for successful exploitation. A misdirected write to the PLT, which typically contains jump stubs to the GOT, would likely lead to a crash or an ineffective exploit.

Programming Miscellaneous Mistakes often revolved around the scripting language used for exploit automation, predominantly Python. Examples included forgetting to convert an address, such as that of the puts function, from an integer representation into its correct little-endian binary format—a fundamental requirement for packing addresses into a payload that the target program can interpret. Another common error was failing to keep stdin open, causing a spawned shell to exit prematurely, preventing interactive command execution. These errors underscore the need for meticulous attention to data types, byte order, and process I/O handling in exploit scripts.

Strategy Mistakes showcased conceptual missteps. One notable example involved a hacker initially using a pointer to a null byte as the str_stream argument for fgets within a Return-Oriented Programming (ROP) chain. fgets expects a valid buffer address to write into, and providing a null pointer would inevitably lead to a crash or undefined behavior, requiring the hacker to completely rethink their ROP gadget selection and exploit flow. This illustrates the importance of understanding function semantics and system call conventions even when building complex exploit primitives.

Tooling Mistakes highlighted the challenges associated with using essential security tools. A frequently cited example involved a hacker attempting to attach GDB (GNU Debugger) to a process running under a different user context without first creating a local copy. This inevitably led to permission-related issues, preventing effective debugging. Such errors emphasize the need for proficiency in tool usage, understanding process permissions, and proper debugging workflows, especially in constrained environments like CTFs. The study implicitly relies on the use of standard tools like disassemblers (e.g., Ghidra, IDA Pro), compilers (e.g., GCC), and debuggers (GDB, Pwndbg, GEF) as part of the observed hacking process. The focus on memory corruption vulnerabilities throughout these challenges underpins the technical context, implying the use of techniques such as buffer overflows, format string bugs, use-after-free, and double-free vulnerabilities to achieve control flow hijacking or information leakage.

Demo / Proof of Concept

▶ Watch: Quantifying the time impact of different mistake types (9:00)

This research article does not present a new technical demonstration or proof-of-concept exploit developed by the speakers. Instead, the talk itself is a detailed analysis of existing demonstrations and proofs-of-concept performed by other security professionals. The "demos" in this context are the CTF-style binary exploitation challenges solved by hackers in the YouTube videos that formed the dataset for this study.

The methodology involved "watching over the shoulder" of these professionals as they attempted to develop working exploits for memory corruption vulnerabilities. The "proof of concept" was observed when these hackers successfully exploited a vulnerability to achieve a specific goal, such as obtaining a shell or leaking sensitive information, within the CTF challenge. The researchers meticulously cataloged the points where these professionals encountered difficulties, made errors, and subsequently debugged and corrected their approach. Thus, while there isn't a new live demo, the entire research is built upon the empirical observation and analysis of numerous real-world (albeit controlled CTF environment) exploit development efforts.

Defensive Implications

▶ Watch: Non-technical factors and mitigation strategies for mistakes (10:10)

While the study focuses on the offensive side of security, its findings offer crucial insights for defenders looking to fortify their systems and processes. Understanding where attackers commonly make mistakes can inform more effective defensive strategies and educational initiatives.

  1. Prioritize Memory Safety: The overwhelming prevalence and time-consuming nature of "Exploit Implementation Mistakes" (especially those related to leak, write what, and write where) underscore the enduring challenge of memory corruption vulnerabilities. This directly reinforces the need for defenders to:
  • Adopt memory-safe languages (Rust, Go) for new development where possible.
  • Rigorously apply compiler-based exploit mitigations such as ASLR (Address Space Layout Randomization), DEP/NX (Data Execution Prevention/No-Execute), and CFI (Control Flow Integrity).
  • Implement robust fuzzing and static/dynamic analysis tools to proactively identify and patch memory corruption flaws.
  • Emphasize secure coding practices related to buffer handling, pointer arithmetic, and object lifecycle management during developer training.
  1. Enhance Developer Education: The identified "Programming Miscellaneous" and "Strategy Mistakes" highlight gaps in foundational knowledge and strategic thinking that even experienced hackers encounter. Defenders can leverage this by:
  • Placing greater emphasis on memory layout, memory access patterns, and pointer operations in cybersecurity education and developer training programs.
  • Educating developers on common exploit primitives (e.g., ROP chains, GOT/PLT hijacking) not just from a theoretical standpoint, but also from a practical "how it can go wrong" perspective.
  • Fostering a deeper understanding of low-level system interactions, such as endianness and process I/O, which are often overlooked in high-level programming.
  1. Improve Security Tooling and Workflows: "Tooling Mistakes" suggest that even powerful security tools can be misused or misconfigured. This implies:
  • Investing in or developing more intuitive and error-resistant debugging tools and exploit development libraries that provide better feedback and guard against common memory-related errors.
  • Promoting best practices for using security tools, including understanding process contexts and permissions, which can prevent basic operational errors.
  1. Refine Vulnerability Disclosure and Bug Bounty Programs: The study notes that a valid attack scenario is often required for bug bounties. This suggests that defenders managing such programs should:
  • Recognize the inherent difficulty and time commitment involved in developing exploits, which can be error-prone even for experts.
  • Potentially offer tiered rewards or guidance for vulnerability reports that demonstrate high impact but might have minor exploit implementation flaws, acknowledging the effort involved.
  1. Inform Incident Response and Threat Intelligence: Understanding the common mistakes attackers make can help incident responders anticipate attack paths, interpret forensic evidence, and develop more targeted detection rules. If an attacker is likely to make certain types of errors, defenders can look for artifacts of those errors in logs or system states.

In essence, by dissecting the challenges faced by offensive practitioners, this research provides a defensive roadmap: strengthen the foundational security of memory-intensive applications, educate developers on both secure coding and the attacker's mindset, and refine the tools and processes that support both offensive and defensive security operations.

Key Takeaways

  • YouTube's Potential for Security Research: The study successfully demonstrates that publicly available video content of security professionals solving challenges can serve as a rich, empirical dataset for security research, providing insights into human behavior in complex technical tasks.
  • Hacking is Highly Error-Prone: Hackers spend an average of 40% of their total time on mistakes, with individual cases ranging from 3% to 75%. This highlights the significant inefficiency introduced by human error in vulnerability research and exploitation.
  • Implementation and Strategy Mistakes Dominate: Exploit implementation mistakes are the most frequently occurring type and the most time-consuming to resolve, often taking nearly an hour. Strategy mistakes are the second most impactful in terms of time spent, emphasizing the importance of correct conceptual approaches.
  • Impact of Technical Issues: While less frequent, external technical issues (e.g., application difficulties, missing libraries) can be profoundly impactful, sometimes consuming over an hour and a half of a hacker's time.
  • Recommendations for Improvement: The researchers recommend memory-focused improvements in security tools (e.g., debuggers, exploitation libraries) to help mitigate these mistakes faster. They also advocate for placing greater emphasis on memory layout, memory access, and pointer operations in cybersecurity education courses.
  • Human Factors are Critical: Non-technical factors such as forgetting critical elements, lack of attention, and unawareness of the correct approach are significant drivers of mistakes, pointing to the need for better cognitive support and training.

About the Speaker(s)

The primary presenter for this talk was Irina Ford, who is identified as a PhD student at Arizona State University. She led the presentation of this collaborative work. The research paper and presentation acknowledge the significant contributions of her co-authors: Ananta Soneji, Faris Bugra Kokulu, Jayakrishna Vadayath, Zion Leonahenahe Basque, and Gaurav Vipat. While specific individual bios for the co-authors were not detailed in the transcript, their collective affiliation with Arizona State University suggests a shared academic focus on cybersecurity research.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This research provides a data-driven, systematic analysis of human error in binary exploitation, leveraging CTF videos to quantify hacker mistakes. The "Mistake Anatomy Framework" and detailed categorization of errors offer invaluable insights for improving exploit development efficiency and informing defensive strategies. It's a solid empirical study that brings necessary rigor to understanding the human element in offensive security.

Heather Calloway (CISO) — STRONG ACCEPT

This research provides crucial empirical data on the human element in exploit development, revealing that even skilled hackers spend 40% of their time rectifying mistakes. It offers clear, actionable insights for security leaders to prioritize memory safety, enhance developer education, and refine tooling, directly informing defensive strategy and investment. This is highly relevant for institutional decision-making.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024