SoK: Prudent Evaluation Practices for Fuzzing
Moritz Schloegel, Nils Bars, Nico Schiller, Lukas Bernhard, Tobias Scharnowski, Addison Crump
IEEE Symposium on Security and Privacy 2024 · Day 2 · Continental Ballroom 4
Overview
This talk, presented by Moritz Schloegel, delves into the critical issue of evaluation practices within the rapidly expanding field of fuzzing research. Titled "SoK: Prudent Evaluation Practices for Fuzzing," the work received a distinguished paper award at IEEE S&P, highlighting its significance and timely contribution to the security community. The core objective of the research is to scrutinize the reproducibility and validity of fuzzing evaluations, addressing concerns that the sheer volume of new fuzzing techniques might be leading to a "replication crisis" similar to those observed in other scientific disciplines.

Key moments
- 0:00 Introduction: Fuzzing hype and replication crisis risk
- 1:45 Overview of study methodology and scope
- 2:50 Referencing established fuzzing evaluation guidelines
- 4:30 Common pitfalls with outdated benchmarks
- 6:00 Demonstrating the crucial role of ablation studies
- 8:00 Warning against misleading new coverage metrics
- 9:00 Why unique crashes are not a good bug proxy
SoK: Prudent Evaluation Practices for Fuzzing
Speakers: Moritz Schloegel, Nils Bars, Nico Schiller, Lukas Bernhard, Tobias Scharnowski, Addison Crump
Conference: IEEE S&P
YouTube: https://www.youtube.com/watch?v=Yaar_xKxr_Y
Overview
This talk, presented by Moritz Schloegel, delves into the critical issue of evaluation practices within the rapidly expanding field of fuzzing research. Titled "SoK: Prudent Evaluation Practices for Fuzzing," the work received a distinguished paper award at IEEE S&P, highlighting its significance and timely contribution to the security community. The core objective of the research is to scrutinize the reproducibility and validity of fuzzing evaluations, addressing concerns that the sheer volume of new fuzzing techniques might be leading to a "replication crisis" similar to those observed in other scientific disciplines.
The speakers meticulously analyzed a vast corpus of fuzzing papers, identifying common pitfalls and areas for improvement in how new fuzzers are assessed and compared. They present a comprehensive analysis of current evaluation methodologies, offering updated guidelines to foster more rigorous, transparent, and reproducible research outcomes. This talk is crucial for anyone involved in fuzzing research, development, or consumption, as it provides a foundational understanding of what constitutes a robust evaluation and how to critically appraise claims made by new fuzzing tools.
Background
▶ Watch: Introduction: Fuzzing hype and replication crisis risk (0:00)
The field of fuzzing has experienced an unprecedented surge in research interest, with hundreds of new fuzzing works published across top-tier security and software engineering venues in recent years. This "fuzzing hype train," as described by the speaker, has led to a large body of research that, while innovative, carries inherent risks. A primary concern is the potential for a replication crisis, where published results cannot be independently verified or reproduced, undermining the scientific integrity and practical utility of the research.
Prior to this work, a foundational paper titled "Evaluating Fuzz Testing" by C at al., published at CCS six years ago, provided initial recommendations for fuzzing evaluations. However, with the evolution of fuzzing techniques and the increased complexity of evaluations, these guidelines required re-examination. The authors of "SoK: Prudent Evaluation Practices for Fuzzing" undertook an extensive literature analysis, surveying 150 fuzzing papers sampled from a pool of approximately 300 works published over the past six years across seven top venues: ACM CCS, IEEE S&P, USENIX Security, NDSS, ASE, ICSE, and FSE. This comprehensive review aimed to identify current practices, recurring issues, and opportunities to update and refine existing evaluation recommendations. For eight specific papers identified as having "atypical stuff going on," the team conducted practical reproduction attempts to validate their findings in real-world scenarios, thereby informing their updated guidelines.
Key Findings
▶ Watch: Referencing established fuzzing evaluation guidelines (2:50)
The research revealed that despite the maturity of fuzzing as a research area, evaluations are "simply damn hard to get right," often suffering from significant pitfalls that compromise their validity and reproducibility. The key findings highlight systemic issues across various stages of the evaluation process:
- Insufficient Documentation: Critical setup and parameter details, such as target and fuzzer versions, are frequently omitted, hindering reproducibility.
- Suboptimal Target Selection: Many papers either avoid benchmarks entirely or rely on outdated ones that incorporate artificially injected bugs, failing to represent real-world vulnerabilities.
- Lack of Ablation Studies: When multiple contributions are proposed, individual components are rarely evaluated in isolation, making it difficult to ascertain their true impact.
- Misleading Evaluation Metrics: Non-standard or misinterpreted metrics, particularly the use of "unique crashes" as a proxy for actual bugs, can significantly distort results. Custom code coverage metrics, without accompanying standard metrics, can also be deceptive.
- Poor Statistical Rigor: A substantial number of evaluations lack proper statistical analysis, often using fewer than 10 repetitions and failing to measure effect size alongside statistical significance.
- Underutilization of Artifact Evaluation: Despite a high availability of fuzzer source code, participation in artifact evaluation programs remains low, missing an opportunity for external validation.
- CVE Mismanagement and Misaligned Incentives: A significant portion of reported CVEs claimed by fuzzing papers remain unfixed, reserved, or invalid. This points to a problematic incentive structure where CVE counts are prioritized over verified bug remediation, potentially leading to "questionable quality" CVEs.
These findings collectively underscore a pressing need for more rigorous, standardized, and transparent evaluation practices to ensure the continued scientific progress and practical impact of fuzzing research.
Technical Deep Dive
▶ Watch: Common pitfalls with outdated benchmarks (4:30)
The core of the "SoK: Prudent Evaluation Practices for Fuzzing" lies in its detailed analysis of common technical shortcomings in fuzzing evaluations, substantiated by concrete case studies from their reproduction efforts.
Documentation and Target Selection
A fundamental requirement for reproducibility is comprehensive documentation of the experimental setup. The paper emphasizes the need to document subtle but critical details, such as the exact versions of targets and fuzzers used. While rarely completely forgotten, the lack of precision makes exact replication challenging.
Regarding target selection, the authors note that while benchmarks like FuzzBench are available and offer a standardized way to compare fuzzers, their usage is inconsistent. Many papers do not use benchmarks at all, or they rely on outdated ones that feature artificially injected bugs, which do not accurately reflect real-world vulnerabilities. The authors acknowledge that some specialized fuzzing sub-fields, such as firmware fuzzing, may not have suitable benchmarks, but for general-purpose Linux user-space fuzzers, benchmarks are highly recommended.
Baselines and Ablation Studies
A crucial aspect of evaluating a new fuzzer is comparing it against a robust baseline and clearly demonstrating the contribution of each novel component. The talk presents a case study involving a new fuzzer that proposed two main contributions: dynamically adapting mutation probabilities and using an evolutionary strategy to optimize these probabilities. During reproduction on a Ploymorphic Fuzz Target from FuzzBench, the original paper showed the new fuzzer outperforming AFL.
However, the original work lacked an ablation study. The researchers performed one by comparing the full fuzzer against a version that only dynamically adapted probabilities but used a random strategy instead of the evolutionary one. The results were striking: the randomly adapted version performed statistically insignificantly different, and sometimes even better, than the fuzzer employing the evolutionary strategy. This demonstrated that while dynamic adaptation might be beneficial, the specific evolutionary strategy proposed did not yield the intended improvement. This case study powerfully illustrates that ablation studies are crucial for isolating and measuring individual contributions, especially when multiple innovations are presented.
Evaluation Metrics: Code Coverage and Bugs
The choice and interpretation of evaluation metrics are paramount. Code coverage is universally used, but its application can be misleading. Another case study involved a fuzzer aiming to cover more paths with fewer inputs. The original paper presented a figure suggesting it achieved this goal, significantly outperforming AFL. However, when the researchers reproduced this using the standardized coverage over time metric, AFL was found to discover more branches over a 24-hour period. While new metrics are not inherently bad, the talk stresses the importance of always including widely accepted, standard metrics to prevent readers from being misled.
The "gold standard" for fuzzing evaluation is finding bugs. However, evaluating on real bugs is often tedious due to unknown underlying bug distributions in real-world targets. Consequently, many papers use AFL's unique crashes as a convenient proxy. A particularly illuminating case study highlighted the dangers of this practice. A fuzzer proposing an additional feedback mechanism (memory usage) prominently used unique crashes to claim superiority over AFL, reportedly finding four times as many unique crashes on the popular fuzzing target nm.
The reproduction effort revealed a vastly different picture. The developers had patched one of the reported bugs, allowing the researchers to re-run all crashing inputs. Inputs that no longer crashed were attributed to that single patched bug. This reduced the initial 1,700 or 464 crashing inputs to 14 or 9, respectively, indicating that "these unique crashes aren't so unique after all." Further manual deduplication of the remaining unique crashes revealed that both the new fuzzer and AFL had found only two actual unique bugs each. This dramatic reduction from a 4x advantage to an equal number of bugs underscores that unique crashes are a highly unreliable proxy and should not be equated with actual bugs, a point previously emphasized by C at al. The strong need for thorough deduplication, even manual, or direct evaluation on verified bugs, is re-emphasized.
Statistical Evaluation
The rigor of statistical evaluation is frequently lacking. Many fuzzing papers repeat experiments, but often fewer than 10 repetitions, which may be insufficient for robust statistical tests. More critically, most studies do not perform statistical evaluations at all. Even when they do, they typically only check for statistical significance without measuring effect size.
The updated guidelines recommend at least 10 repetitions (trials) and the use of appropriate statistical tests such as Mann-Whitney U or bootstrap-based tests for significance. Furthermore, the A.K.A. 12 test by Warmer and Danani is proposed for measuring effect size, providing a more complete picture of the practical impact of a fuzzer's improvements.
Artifact Availability and CVE Management
Beyond the evaluation itself, the paper also examined the availability of fuzzer source code and the handling of newly discovered bugs. The good news is that fuzzer source code is often made available. However, this high availability does not translate into high participation in artifact evaluation programs. The speakers encourage researchers to submit their work to artifact evaluation, noting that it's a relatively new process, which might explain the current low participation.
The handling of CVEs (Common Vulnerabilities and Exposures) reported by fuzzing papers presents another significant issue. Out of approximately 338 CVEs claimed by fuzzing papers, only 133 (39%) have been fixed. A large portion, 88 (26%), are in a "reserved" state with no public disclosure, providing zero information. More worryingly, 69 (20%) CVEs have been ignored, many in active projects, raising questions about maintainer responsiveness. Finally, 36 (11%) CVEs were considered invalid, either duplicates or not deemed security bugs by maintainers, raising questions about why they were assigned a CVE in the first place. This highlights a misalignment of incentives: CVEs are perceived as a proxy for real-world impact in the research community, creating a vector to "game the system" and receive CVEs of questionable quality, as not all CVEs are verified.
Demo / Proof of Concept
▶ Watch: Warning against misleading new coverage metrics (8:00)
The talk itself did not feature a live, interactive demonstration or proof of concept in the traditional sense. Instead, the speakers presented the results of their extensive reproduction efforts and case studies as evidence for their findings. The "demonstrations" were the comparative plots and data derived from re-running experiments from existing fuzzing papers, showcasing where original claims diverged from reproducible outcomes due to the identified evaluation pitfalls. This approach served as a powerful, data-driven "proof of concept" for the need for improved evaluation practices.
Defensive Implications
▶ Watch: Why unique crashes are not a good bug proxy (9:00)
The findings from "SoK: Prudent Evaluation Practices for Fuzzing" carry significant implications for various stakeholders within the security community, particularly for researchers, fuzzer developers, and consumers of fuzzing research.
For fuzzer developers and security researchers, the primary implication is a call to adopt more rigorous and transparent evaluation methodologies. This includes:
- Comprehensive Documentation: Meticulously document all experimental parameters, including exact versions of fuzzers, targets, operating systems, and compilation flags, to enable true reproducibility.
- Strategic Target Selection: Utilize well-established and relevant benchmarks where appropriate, avoiding outdated ones with artificial bugs. For specialized areas lacking benchmarks, clearly justify target choices.
- Ablation Studies as a Standard: Always conduct ablation studies to isolate and quantify the contribution of each novel component in a multi-faceted fuzzer. This ensures that claimed improvements are genuinely attributable to the proposed innovations.
- Reliable Metrics: Prioritize actual bug findings over "unique crashes" by implementing thorough deduplication, potentially manual, or by using verified bug databases. When introducing new coverage metrics, always complement them with standard, widely understood metrics like coverage over time to avoid misleading interpretations.
- Statistical Rigor: Conduct a sufficient number of repetitions (at least 10) for experiments and perform robust statistical analyses, including both statistical significance (e.g., Mann-Whitney U, bootstrap-based tests) and effect size (e.g., A.K.A. 12 test), to ensure the robustness and practical relevance of results.
- Artifact Evaluation Participation: Actively participate in artifact evaluation programs. Making source code available and verifiable boosts the credibility and impact of research.
- Responsible CVE Disclosure: Adopt a more responsible approach to CVE claims. Focus on verified bugs and collaborate with maintainers for remediation rather than using CVE counts as a mere proxy for impact.
For security practitioners, maintainers, and consumers of fuzzing research, these findings serve as a critical guide for evaluating the trustworthiness of new fuzzing tools and techniques:
- Critical Appraisal of Claims: Be skeptical of performance claims, especially those based solely on "unique crashes" or custom metrics without robust statistical backing. Look for evidence of proper baselines, ablation studies, and sufficient repetitions.
- Verify CVEs: Understand that a CVE ID alone does not guarantee a fixed or even valid bug. Investigate the status of reported CVEs and prioritize those that have been publicly verified and remediated.
- Demand Transparency: Favor research that provides clear documentation, publicly available artifacts, and detailed evaluation methodologies.
Ultimately, the defensive implication is a collective responsibility to elevate the standard of fuzzing research, ensuring that innovations are genuinely impactful and that security claims are backed by sound, reproducible scientific evidence.
Key Takeaways
- Fuzzing evaluations are inherently challenging: Achieving robust, reproducible, and valid evaluations for new fuzzing techniques is difficult and prone to significant pitfalls.
- Ablation studies are non-negotiable: When proposing multiple contributions, it is critical to perform ablation studies to isolate and quantify the impact of each individual component.
- "Unique crashes" are not actual bugs: The common practice of using AFL's unique crashes as a proxy for bugs is highly unreliable and can lead to drastically inflated performance claims; thorough, potentially manual, deduplication is essential.
- Statistical rigor is often lacking but vital: Many fuzzing evaluations suffer from insufficient repetitions (often less than 10) and a failure to measure effect size alongside statistical significance, undermining the robustness of their findings.
- Artifact availability and evaluation are underutilized: While fuzzer source code is increasingly available, participation in artifact evaluation programs is low, missing a key opportunity for external validation and reproducibility.
- CVEs in fuzzing papers require scrutiny: Misaligned incentives can lead to a focus on claiming CVEs over verified bug remediation, resulting in a significant number of unfixed, reserved, or invalid CVEs.
About the Speaker(s)
Moritz Schloegel was the lead presenter for this paper, "SoK: Prudent Evaluation Practices for Fuzzing," which received a distinguished paper award at IEEE S&P. He collaborated with co-authors Nils Bars, Nico Schiller, Lukas Bernhard, Tobias Scharnowski, and Addison Crump on this extensive research project. The talk reflects their collective effort in analyzing hundreds of fuzzing papers and conducting case studies to identify and address critical issues in fuzzing evaluation methodologies.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This distinguished paper meticulously dissects the alarming "replication crisis" in fuzzing research, offering a brutal but essential critique of prevalent evaluation practices. It provides updated, rigorous guidelines for the field, exposing critical flaws in methodology from target selection to statistical analysis and CVE reporting. This is foundational work that sets a new bar for scientific integrity in security research.
Heather Calloway (CISO) — STRONG ACCEPT
This SoK paper provides a critical, evidence-based framework for evaluating fuzzing research, which is essential for security leaders making informed decisions. It exposes systemic flaws in how new fuzzers are assessed, offering clear guidelines that directly inform how security programs should critically appraise vendor claims and integrate new technologies. While not a direct governance brief, it strengthens the foundation for executive decision-making.
→ Top-rated talks at IEEE Symposium on Security and Privacy 2024