IDFuzz: Intelligent Directed Grey-box Fuzzing

Yiyang Chen (PhD student · Chinua University)

34th USENIX Security Symposium (USENIX Security '25) · Day 3 · Software Security 3: Fuzzing

Overview

This talk introduces IDFuzz, an innovative approach to intelligent directed grey-box fuzzing. Presented by Yiyang Chen, a PhD student from Chinua University, the work addresses a significant inefficiency in traditional directed fuzzing methods: the blind and often random nature of input mutation. Directed fuzzing is a critical technique used to test specific target code segments within larger programs, with applications ranging from crash reproduction and candidate vulnerability confirmation to comprehensive patch testing. The primary goal is to reach these target code paths as quickly and efficiently as possible.

Watch on YouTube · Slides

Visual summary for IDFuzz: Intelligent Directed Grey-box Fuzzing by Yiyang Chen
Visual summary for IDFuzz: Intelligent Directed Grey-box Fuzzing by Yiyang Chen

Key moments

  1. 0:00 Introduction to IDFuzz and directive fuzzing problem
  2. 1:00 IDFuzz's key insight: identifying critical input fields
  3. 2:00 Addressing generalizability and exploring new code regions
  4. 3:10 Real-world JPEG example and existing method limitations
  5. 4:00 Neural network and model gradients for critical byte identification
  6. 6:00 IDFuzz's novel techniques: encoding, dataset generation, filtering
  7. 8:00 Evaluation results: performance speedups and low overhead
  8. 10:00 Discovered vulnerabilities and concluding remarks

IDFuzz: Intelligent Directed Grey-box Fuzzing

Speakers: Yiyang Chen

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=uCl0MvchI08

Overview

This talk introduces IDFuzz, an innovative approach to intelligent directed grey-box fuzzing. Presented by Yiyang Chen, a PhD student from Chinua University, the work addresses a significant inefficiency in traditional directed fuzzing methods: the blind and often random nature of input mutation. Directed fuzzing is a critical technique used to test specific target code segments within larger programs, with applications ranging from crash reproduction and candidate vulnerability confirmation to comprehensive patch testing. The primary goal is to reach these target code paths as quickly and efficiently as possible.

Existing directed fuzzers typically prioritize inputs based on a fitness metric, such as distance to the target, but frequently fall short in their mutation strategies. They often employ a random selection of offsets and lengths for input modification, a method proven to be inefficient in the context of directed fuzzing. IDFuzz tackles this challenge by proposing a novel input mutation strategy that intelligently identifies and targets critical input fields for modification. This approach is rooted in the insight that valuable knowledge about critical fields can be gleaned from the outcomes of past mutations, thereby enabling more focused and effective fuzzing.

The significance of IDFuzz lies in its ability to overcome the limitations of prior techniques by introducing a learning-based mutation strategy. By leveraging a neural network model to learn high-level patterns between input byte organization and their distance to a target, IDFuzz achieves generalizability in identifying critical fields. Furthermore, it employs a unique representation of program branching behaviors, allowing the fuzzer to learn how to reach unexplored code regions from already covered sibling branches. This intelligent, gradient-guided mutation strategy promises to dramatically accelerate the discovery and reproduction of vulnerabilities, offering substantial benefits for software security and quality assurance.

Background

▶ Watch: Introduction to IDFuzz and directive fuzzing problem (0:00)

Directed fuzzing is an important branch of fuzzing that focuses on reaching specific target code locations within a program. Its practical applications are diverse and crucial for software security, including reproducing known crashes, confirming the existence of suspected vulnerabilities, and rigorously testing patches to ensure they effectively resolve issues without introducing new ones. A central objective of directed fuzzing is to minimize the time and resources required to hit the designated target code.

Traditional directed fuzzers often rely on a fitness metric, such as the "distance" of an input from the target code, to prioritize which inputs (seeds) to mutate. Inputs that get "closer" to the target are typically favored. However, a significant bottleneck in these methods has been the input mutation step itself. Many existing fuzzers adopt a strategy of randomly choosing offsets and lengths within an input to apply changes. While simple to implement, this random approach is inherently inefficient for directed fuzzing, as it often results in numerous "fruitless mutations" that do not contribute to reaching the target. This inefficiency becomes particularly pronounced in complex code structures.

The problem is exacerbated in scenarios where target code is deeply nested or requires specific, non-obvious byte sequences to be triggered. The speaker provided a real-world example from the Jhat program, a tool for reading JPEG format input files. A specific vulnerability was located at line 24, requiring a particular "marker byte" to be mutated to the value 'mcom' to be reached. This marker byte was identified as the first non-padding byte with a length of one. The speaker highlighted that due to nested loops and padding bytes within the file format, traditional program analysis techniques like symbolic execution and taint analysis were ineffective in identifying this critical byte. Symbolic execution struggles with path explosion in complex loops, while taint analysis might fail to propagate taint through non-obvious data transformations or when data is directly embedded without clear lineage. This illustrates the pressing need for a more intelligent and adaptive mutation strategy that can navigate such complexities.

Key Findings

▶ Watch: Addressing generalizability and exploring new code regions (2:00)

IDFuzz presents several significant contributions and findings that redefine the landscape of directed fuzzing:

  1. Intelligent Critical Field Identification: The core finding is that critical fields for mutation can be intelligently identified based on the outcomes of past mutations. If mutating an input at a specific position X with a certain value change leads to a smaller distance to the target, then X is identified as a critical field. This moves beyond blind, random mutation to a data-driven, learning-based approach.
  1. Generalizability through Neural Networks: IDFuzz affirmatively answers the question of whether this knowledge is generalizable. It achieves this by employing a neural network model to learn high-level patterns between the organization of input bytes and their corresponding distance to the target code. This allows the fuzzer to apply learned insights to new, unseen inputs, overcoming the specificity limitations of individual mutation events.
  1. Reaching Unexplored Code Regions: The work also addresses how to guide a fuzzer to unexplored code. By using a novel presentation of program branching behaviors and enabling the neural network model to learn from already covered "sibling branches," IDFuzz can effectively guide the fuzzer towards previously unreached code paths. This is achieved by assigning distinct "distance" values (between zero and one) to different execution paths, even if they don't directly hit the target, allowing the model to learn fine-grained distinctions.
  1. Significant Performance Speed-up: In empirical evaluations using the Google Fuzzer Test Suite as a benchmark, IDFuzz demonstrated substantial performance improvements. When integrated as an input mutation module into existing baselines like Go-Win-Ranger and Sele, IDFuzz provided at least two times speed-ups in reproducing target code. This highlights its practical effectiveness in accelerating directed fuzzing tasks.
  1. Vulnerability Discovery: Beyond speed, IDFuzz proved its efficacy in discovering real-world security flaws. The technique led to the discovery of six new vulnerabilities and one incomplete fix. Among these, four CVEs were granted, and two were identified as high-severity vulnerabilities, underscoring IDFuzz's capability to uncover critical security issues.
  1. Low Runtime Overhead: Despite incorporating a neural network model for training and gradient computation, IDFuzz maintains a remarkably low runtime overhead. The evaluation revealed only a 6% increase in runtime overhead, demonstrating that the intelligent mutation strategy does not come at the cost of significant computational expense, making it practical for real-world deployment.
  1. Contribution of Key Techniques: The evaluation explicitly confirmed that the three novel techniques introduced by IDFuzz—branch encoding, adaptive data set generation, and gradient filtering—each make clear and measurable contributions to the overall performance of the system, validating their design and implementation.

Technical Deep Dive

▶ Watch: Neural network and model gradients for critical byte identification (4:00)

The core technical innovation of IDFuzz lies in its gradient-guided mutation strategy, which intelligently identifies critical input fields. This approach is inspired by prior work in neural network-guided fuzzing (e.g., NEWS) but is specifically tailored to address the unique challenges of directed fuzzing.

At a high level, the process involves two main steps:

  1. Modeling Input Relationship: IDFuzz first models the relationship between the input bytes and the input distance to the target code using a neural network model.
  2. Identifying Critical Fields: It then identifies the critical bytes or fields by locating the maximum model gradients. The mathematical intuition is that changing a critical byte will have the greatest impact on the input's distance to the target, thus corresponding to the largest absolute model gradient value.

To make this approach practical and effective for directed fuzzing, IDFuzz introduces three key techniques:

1. Branch Encoding for Model Initialization

A critical challenge in directed fuzzing is learning to reach unexplored code regions. Traditional distance metrics might only differentiate between "reached target" (distance 0) and "not reached target" (distance 1). IDFuzz refines this by creating a more nuanced distance metric and a specific encoding strategy:

  • Distance Metric Refinement:
  • Inputs that successfully reach the target code (e.g., a specific vulnerability trigger like 'mcom') are assigned a distance of 0.
  • Inputs that do not reach the target are initially given a distance of 1.
  • Crucially, for inputs that execute other cases or sibling branches (i.e., paths that are distinct from the target path but still covered), IDFuzz assigns different values between zero and one. This allows the neural network to distinguish between various non-target execution paths, enabling it to learn how to nudge inputs from one covered sibling branch towards an unexplored one.
  • Dimensionality Reduction: To manage the complexity of program paths and reduce the input dimensionality for the neural network, IDFuzz first extracts the coverage of dominating basic blocks (DOMBBs). DOMBBs represent control flow structures in a more compact way.
  • Branch Encoding Technique: A novel technique uses a proper fraction between zero and one to encode covered sibling branches. This fractional encoding allows the neural network to perceive the "proximity" or "direction" towards an unexplored branch, even when a direct path has not yet been found. By learning the relationship between input bytes and these fractional distances, the model can guide mutations towards paths that are "closer" to the desired unexplored branch.

2. Adaptive Data Set Generation for Model Training

Training a neural network requires a sufficient and representative dataset. However, in fuzzing, continuously accumulating all inputs can lead to an excessively large and redundant dataset. IDFuzz addresses this with an adaptive strategy:

  • Pruning Early Inputs: The technique first prunes early accumulated inputs that cover very few DOMBBs. These inputs are less informative for guiding the fuzzer towards complex targets.
  • Leveraging Similarity: It then reduces the dataset size by leveraging the similarity in byte distributions among relative seeds. If multiple inputs achieve similar coverage or distances with highly similar byte patterns, only a representative subset is retained. This ensures a smaller, yet sufficient, dataset for efficient model training without losing critical information, thus keeping the runtime overhead low.

3. Gradient Filtering for Gradient Guided Directed Mutation

Raw model gradients can be noisy and misleading, especially in the context of program inputs. IDFuzz introduces gradient filtering to refine the identification of critical fields:

  • Eliminating Noise: The technique helps eliminate noise caused by uneven input byte distributions. This includes:
  • Large gradients for magic bytes: File formats often contain "magic bytes" that identify the file type. These bytes might show large gradients if they cause early parsing failures, but mutating them might not be productive for reaching deep targets.
  • Critical bits for neighboring DOMBBs: Gradients might indicate sensitivity to changes that affect immediately neighboring basic blocks, rather than the distant target. Gradient filtering helps prioritize changes that have a broader impact on target reachability.
  • Grouping Critical Bytes: After filtering, IDFuzz uses a DBScan clustering algorithm to group individual critical bytes into larger critical fields. Mutating an entire field (e.g., a 4-byte integer) is often more effective than mutating single bytes, as many program structures operate on multi-byte values. This clustering enhances the effectiveness of the mutation step.

The speaker provided an illustrative example from the Jhat program, where the target was a marker byte that needed to be 'mcom'. For an average input length of nearly 500 bytes, IDFuzz successfully located this marker byte in just six attempts. This dramatically reduces the number of fruitless mutations compared to random approaches, which would require many more attempts to stumble upon the correct byte and value.

Demo / Proof of Concept

▶ Watch: IDFuzz's novel techniques: encoding, dataset generation, filtering (6:00)

While the talk transcript does not detail a live demonstration of IDFuzz, the evaluation results presented serve as a robust proof of concept for its effectiveness. The speaker highlighted concrete figures and scenarios to illustrate IDFuzz's capabilities.

Specifically, the example from the Jhat program demonstrated IDFuzz's ability to pinpoint critical input fields in a complex real-world scenario. For an input file with an average length of nearly 500 bytes, IDFuzz was able to successfully locate the specific "marker byte" required to reach a target vulnerability in only six mutation attempts. This empirical evidence directly supports the claim that IDFuzz's intelligent, gradient-guided mutation strategy drastically reduces the search space and improves efficiency compared to traditional random mutation.

Furthermore, the broader evaluation against the Google Fuzzer Test Suite benchmarks, showing "at least two times speed-ups" for baselines like Go-Win-Ranger and Sele, and the discovery of six new vulnerabilities (including four CVEs), collectively underscore the practical viability and significant impact of IDFuzz. These findings demonstrate that the underlying technical principles—neural network learning, branch encoding, adaptive dataset generation, and gradient filtering—translate into tangible security benefits and improved fuzzing performance.

Defensive Implications

▶ Watch: Discovered vulnerabilities and concluding remarks (10:00)

IDFuzz offers significant advantages for software defenders and security teams, enhancing their capabilities in several key areas:

  1. Accelerated Vulnerability Discovery: By making directed fuzzing substantially more efficient, IDFuzz enables defenders to find vulnerabilities in specific, critical code paths much faster. This proactive approach helps identify and fix bugs before they can be exploited by attackers, significantly improving software security posture.
  1. Effective Patch Testing: The ability to quickly and reliably reach target code paths makes IDFuzz an ideal tool for patch testing. Defenders can use it to confirm that a security patch effectively closes a vulnerability and doesn't introduce regressions or new issues. This is crucial for maintaining the integrity of fixes in complex software.
  1. Efficient Crash Reproduction and Triage: When a crash report comes in, directed fuzzing can be used to quickly reproduce the crash scenario. IDFuzz's intelligent mutation can pinpoint the exact input variations that trigger a crash, providing developers with precise information for debugging and root cause analysis. This streamlines the incident response and vulnerability remediation process.
  1. Targeted Security Audits: For high-risk components or newly developed features, IDFuzz allows security teams to conduct highly targeted audits. Instead of relying on broad, undirected fuzzing, they can direct IDFuzz to focus resources on specific functions, modules, or control flow paths known to be sensitive or complex, maximizing the chances of finding relevant vulnerabilities.
  1. Integration into CI/CD Pipelines: Given its low runtime overhead (only 6% increase) and significant speed-ups, IDFuzz's intelligent mutation module can be seamlessly integrated into continuous integration and continuous delivery (CI/CD) pipelines. This enables continuous security testing, where code changes are automatically fuzzed against known or suspected vulnerable paths, providing immediate feedback to developers.
  1. Overcoming Limitations of Traditional Analysis: IDFuzz's ability to navigate complex code with nested loops and padding bytes, where symbolic execution and taint analysis struggle, means that defenders have a new tool to explore code regions previously considered difficult or impossible to analyze effectively with automated methods.

In summary, IDFuzz empowers defenders with a more precise, efficient, and intelligent fuzzing capability, enabling them to find, reproduce, and verify fixes for vulnerabilities with unprecedented speed and accuracy.

Key Takeaways

  • IDFuzz introduces a novel intelligent directed grey-box fuzzing approach that moves beyond inefficient random input mutation.
  • It leverages a neural network model to learn high-level patterns between input byte organization and target distance, enabling generalizable identification of critical input fields.
  • The technique incorporates branch encoding (using fractional distances for sibling branches), adaptive data set generation, and gradient filtering (with DBScan clustering) to enhance model training and mutation effectiveness.
  • Empirical evaluations show at least two times speed-ups in reproducing targets on benchmarks like the Google Fuzzer Test Suite and led to the discovery of six new vulnerabilities, including four CVEs.
  • Despite involving neural network training, IDFuzz maintains a low runtime overhead, with only a 6% increase, making it practical for real-world application.
  • IDFuzz provides a powerful tool for defenders to accelerate vulnerability discovery, improve patch testing, and efficiently reproduce crashes, especially in complex code scenarios where traditional analysis struggles.

About the Speaker(s)

Yiyang Chen is a PhD student from Chinua University. He is the first author of the paper describing IDFuzz, which he presented at the USENIX Security conference. His research focuses on advancing fuzzing techniques, particularly in the domain of directed fuzzing.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Solid academic fuzzing research with a genuine technical contribution: gradient-guided mutation selection for directed grey-box fuzzing, backed by concrete benchmark results and real CVEs. The neural network integration is non-trivial and the three enabling techniques (branch encoding, adaptive dataset generation, gradient filtering) are well-motivated. Not a paradigm shift, but this is honest, careful work that advances the state of the art.

Heather Calloway (CISO) — WEAK

IDFuzz is technically credible work — a neural network-guided mutation strategy that demonstrably outperforms random approaches in directed fuzzing, with real CVEs to show for it. But this talk never climbs out of the research layer. There is no meaningful bridge to how security programs, product security teams, or vulnerability management functions should change what they do.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)