AFGen: Whole-Function Fuzzing for Applications and Libraries
Yuwei Liu, Yanhao Wang, Xiangkun Jia, Zheng Zhang, Purui Su
IEEE Symposium on Security and Privacy 2024 · Day 2 · Continental Ballroom 4
Overview
In the realm of software security, fuzzing has long stood as a cornerstone technique for discovering vulnerabilities by feeding programs with malformed or unexpected inputs. Despite its widespread adoption and proven efficacy, traditional fuzzing approaches often struggle to achieve comprehensive code coverage, particularly at the granular level of individual functions within complex applications and libraries. This limitation significantly hampers the discovery of deep-seated bugs that reside in less frequently executed code paths or those requiring specific, intricate input sequences to trigger. The talk "AFGen: Whole-Function Fuzzing for Applications and Libraries" introduces a novel solution to this pervasive challenge, proposing a method to systematically fuzz all functions within a target program.

Key moments
- 1:30 Limitations of current fuzzing and function coverage
- 2:00 Introducing whole-function fuzzing and automatic harness generation
- 4:00 Motivational example: Discovering a heap buffer overflow
- 6:00 AFGen architecture: Slicer, Assigner, and Tracer
- 7:00 Detailed explanation of the bi-directional slicer
- 9:00 Tracing constraints to filter false positive crashes
- 11:00 AFGen's superior performance in discovering vulnerabilities
AFGen: Whole-Function Fuzzing for Applications and Libraries
Speakers: Yuwei Liu; Yanhao Wang; Xiangkun Jia; Zheng Zhang; Purui Su
Conference: IEEE S&P
YouTube: https://www.youtube.com/watch?v=QCoo0gOzfQM
Overview
In the realm of software security, fuzzing has long stood as a cornerstone technique for discovering vulnerabilities by feeding programs with malformed or unexpected inputs. Despite its widespread adoption and proven efficacy, traditional fuzzing approaches often struggle to achieve comprehensive code coverage, particularly at the granular level of individual functions within complex applications and libraries. This limitation significantly hampers the discovery of deep-seated bugs that reside in less frequently executed code paths or those requiring specific, intricate input sequences to trigger. The talk "AFGen: Whole-Function Fuzzing for Applications and Libraries" introduces a novel solution to this pervasive challenge, proposing a method to systematically fuzz all functions within a target program.
Presented by Yuwei Liu from the University of Chinese Academy of Sciences, this work tackles the inherent difficulties in constructing fuzzing harnesses for internal functions automatically. Current automated harness generation primarily focuses on well-documented APIs, leaving a vast landscape of internal logic largely unexplored by automated fuzzers. AFGen addresses this gap by intelligently constructing the necessary program context, assigning valid values to uninitialized variables, and dynamically tracing and incorporating program constraints to filter out false positives. The significance of AFGen lies in its ability to dramatically expand the scope of automated vulnerability discovery, moving beyond surface-level API testing to probe the intricate internal workings of software.
The research presented highlights AFGen's superior performance compared to existing fuzzing tools, demonstrating its capacity to reproduce a substantial number of known CVEs more rapidly and, crucially, to uncover a significant count of previously unknown vulnerabilities. By enabling "whole-function fuzzing," AFGen offers a powerful new paradigm for software security testing, promising to enhance the robustness and reliability of applications and libraries by thoroughly scrutinizing their entire codebase. This advancement is particularly critical for complex systems where manual harness construction for thousands of internal functions is impractical, making AFGen a vital tool for developers and security researchers alike.
Background
▶ Watch: Limitations of current fuzzing and function coverage (1:30)
Fuzzing, as a technique, operates on the principle of mutating inputs with minimal understanding of the target program's internal logic, aiming to provoke crashes or unexpected behavior indicative of vulnerabilities. While successful, general-purpose fuzzers like AFL and Fuzz++ often prioritize broad code coverage, which, despite its benefits, frequently falls short of achieving complete function-level coverage. This limitation is exacerbated by the sheer volume of functions and their complex interdependencies within modern software.
To address this, researchers have developed various specialized fuzzing approaches:
- Directed fuzzing (e.g., F-GO, PRIAS, Bithon) aims to guide fuzzing towards specific code locations, often requiring prior knowledge of potential vulnerable areas. While more focused, these tools still face challenges in generating inputs that can reliably reach deep, complex code paths.
- API fuzzing (e.g., FAPI, F-STRING) concentrates on the operations of exposed Application Programming Interfaces. These tools benefit from detailed API documentation and simpler argument structures, making automated harness generation more feasible. However, they inherently overlook the vast majority of internal, undocumented functions that constitute the core logic of applications and libraries.
The fundamental problem lies in the difficulty of triggering and testing internal functions. These functions are often deeply nested, have complex argument types, rely on specific program states, and possess intricate control and data flow dependencies that are hard to satisfy through external application inputs alone. Manually constructing a fuzzing harness for each internal function is an arduous, time-consuming, and error-prone task. A harness is a small piece of code that sets up the necessary environment, calls the target function with fuzzer-generated inputs, and checks for crashes. Given the potentially thousands of functions in a large library, manual harness generation is simply not scalable.
Automated fuzzing harness generation methods have primarily focused on APIs because their arguments are typically basic types or can be initialized via well-defined functions, and their dependencies are relatively contained. Extending this automation to all functions, including internal ones, presents three significant challenges:
- Constructing proper program context: An internal function often relies on a specific state of the program or other objects. Simply calling it in isolation may not reflect its intended usage.
- Assigning valid values for uninitialized variables: Function arguments and internal variables may be complex structures, pointers, or file handles, requiring non-trivial initialization to avoid immediate crashes or undefined behavior unrelated to security vulnerabilities.
- Synthesizing program constraints to filter false positives: A crash triggered by a fuzzed input might be due to an invalid program state created by the harness itself, rather than a genuine vulnerability. Distinguishing real vulnerabilities from false positives requires understanding and enforcing the constraints under which the target function is expected to operate.
The speaker illustrates this with a motivational example: a get_byte function in a hypothetical library (referred to as "life 14" in the transcript, likely libflif or a similar library). This function reads a byte based on a mode flag: a memory mode reads from an array, while a file mode reads from a file. A simple harness might call get_byte without properly setting up the memory array or file handler. If the memory mode doesn't check boundaries, a fuzzer could trigger a heap buffer overflow. However, if the default mode is file mode, and get_byte calls getc (which returns EOF for no more data), the bug might only manifest in the memory mode. A naive harness might trigger a crash, but without understanding the mode constraint, it's unclear if this is a true vulnerability or an artifact of an incorrect testing setup. Only by adding constraints (e.g., ensuring the mode is set to memory) can a real vulnerability be confirmed, as was the case for a heap buffer overflow found in the example. This highlights the critical need for AFGen's sophisticated approach to context, variable assignment, and constraint tracing.
Key Findings
▶ Watch: Motivational example: Discovering a heap buffer overflow (4:00)
AFGen represents a significant leap forward in automated vulnerability discovery, primarily through its innovative whole-function fuzzing paradigm. The key findings and contributions of this research are multi-faceted and demonstrate a substantial improvement over existing fuzzing methodologies:
- Comprehensive Vulnerability Reproduction: AFGen successfully reproduced 66 known CVEs across 11 open-source projects. This figure significantly surpasses the reproduction capabilities of other state-of-the-art fuzzers, including general fuzzers like AFL and Fuzz++, and even directed fuzzers such as F-GO, PRIAS, and Bithon. Notably, directed fuzzers, despite being given additional information like vulnerable code locations, performed surprisingly poorly, with F-GO only finding three more vulnerabilities than general fuzzers. This underscores AFGen's ability to navigate complex code structures and satisfy intricate preconditions required to trigger these known flaws.
- Superior Code Coverage for Vulnerable Paths: The enhanced vulnerability reproduction rate is directly correlated with AFGen's ability to achieve greater code coverage, particularly in vulnerable code regions. By meticulously constructing harnesses for individual functions, AFGen can reach deeper and more complex code paths. The analysis revealed that vulnerabilities found by AFGen often reside within code requiring an average of 68.6 branches and 7.2 function nested depths to reach from the entry point to the vulnerable function. This indicates AFGen's effectiveness in exploring deeply embedded logic that other fuzzers typically miss.
- Accelerated Vulnerability Discovery: Beyond mere reproduction, AFGen demonstrated significantly faster discovery times. It was able to reproduce CVEs within the first hour of fuzzing, whereas comparative tools often required four hours or more to achieve similar results. This acceleration is crucial for integrating fuzzing into continuous integration/continuous deployment (CI/CD) pipelines, enabling quicker feedback loops for developers.
- High Precision in Harness Generation: One of AFGen's standout features is the precision of its generated fuzzing harnesses. When compared with API fuzzing tools like FuzzAPI, AFGen achieved an impressive 72.7% precision for all functions, dramatically outperforming FuzzAPI's 7.6%. This high precision means that a much larger proportion of crashes reported by AFGen's harnesses are indeed true vulnerabilities, significantly reducing the effort required to triage false positives. The evaluation of AFGen's individual components (slicer, assigner, tracer) confirmed that all three are critical, with disabling any one component causing a substantial drop in precision (from 72.7% to 47%).
- Discovery of New, Unknown Vulnerabilities: Perhaps the most compelling finding is AFGen's capability to discover 26 previously unknown vulnerabilities. Out of these, 24 were assigned CVE IDs, validating their significance and impact. These new vulnerabilities were categorized into four types based on their trigger mechanisms:
- General type: Triggered solely by mutating input files.
- Compile command: Required specific additional compilation options.
- API: Triggered through specific API endpoints (even for internal functions).
- Command option: Required specific command-line options.
This demonstrates AFGen's versatility in uncovering diverse classes of vulnerabilities that elude conventional fuzzing strategies.
In summary, AFGen's key findings underscore its effectiveness in overcoming long-standing limitations in fuzzing technology. By providing a robust framework for whole-function fuzzing, it not only enhances the ability to reproduce known vulnerabilities quickly and precisely but also significantly expands the potential for discovering novel security flaws across a broad spectrum of software components.
Technical Deep Dive
▶ Watch: AFGen architecture: Slicer, Assigner, and Tracer (6:00)
AFGen's innovative whole-function fuzzing approach is underpinned by three meticulously designed components: the Bidirectional Slicer, the Variable Value Assignment module, and the Constraints Tracer. These components work in concert to automatically generate precise and effective fuzzing harnesses for any target function within an application or library.
Bidirectional Slicer
The primary role of the Bidirectional Slicer is to construct a proper program context for the target function. A function rarely operates in isolation; it depends on data initialized elsewhere, global variables, and specific control flows. The slicer extracts the minimal yet sufficient code snippet that encapsulates these dependencies, forming the basis of the fuzzing harness.
The slicing process involves several iterative steps:
- Target Function Location: The slicer first identifies the target function for which a harness is to be generated.
- Dependency Tracing: It then iteratively adds statements that have either a data flow dependency or a control flow dependency on the target function or on previously added statements.
- Data flow dependency means a statement defines or uses a variable that is an input to or output from the target function, or is used by another statement already included.
- Control flow dependency means a statement (e.g., a conditional branch or loop) determines whether the target function or a dependent statement is executed.
- Iteration: This process repeats until no more statements with dependencies are found, ensuring a comprehensive context.
To optimize the generated code and improve fuzzing performance, AFGen applies several additional strategies:
- Default Slicing Scope: By default, the slicer focuses on functions that directly call the target function. While further slicing of higher-level functions could provide more context, it risks code size explosion, making the resulting harness too large and complex to effectively fuzz. This pragmatic approach balances context completeness with performance.
- Homeless Statement Elimination: Developers often embed debug-related variables and output statements (e.g.,
printf) in their code. These "homeless" statements, while useful for debugging, can adversely impact fuzzing performance by increasing the code size and potentially introducing irrelevant side effects. AFGen identifies and eliminates these statements to simplify the sliced code and improve fuzzing efficiency. - Syntax and Logical Correction: The raw sliced code might contain syntax or logical errors, particularly due to incomplete control statements (e.g., an
ifstatement without its corresponding body statements). To ensure the correctness of the extracted context, AFGen adds statements within the body of such control statements, addressing these issues and making the sliced code compilable and logically sound.
Variable Value Assignment
Once the context is sliced, the harness will likely contain uninitialized variables that are crucial for the target function's execution. The Variable Value Assignment module is responsible for intelligently assigning valid values to these variables, preventing immediate crashes due to null pointers or invalid states and enabling the fuzzer to explore meaningful execution paths. AFGen employs three categories of assignment methods:
- Basic Assignment: This method targets fundamental data types and common structures:
- Base types: Integers, characters, booleans are assigned default or common values (e.g., 0, 'A', true).
- Arrays: Allocated and initialized, often with placeholder values.
- Pointers: Allocated memory and potentially initialized to valid addresses.
- File Handlers: Mocked or opened to temporary files to simulate file I/O operations.
- Combination Assignment: This method handles more complex, composite data structures:
- Structures (structs): Each member of the structure is recursively assigned a value using the appropriate basic or combination assignment method.
- Unions: One member of the union is chosen and assigned a value, reflecting the nature of unions where only one member is active at a time.
- Special Assignment: This category addresses particular types that require unique handling:
- Function Pointers: Assigned to dummy functions or existing functions within the sliced context that match the required signature.
- Cast Pointers: Handled by analyzing the target type of the cast and assigning a value compatible with that type.
- Enumerations (enums): Assigned valid values from the defined enumeration set.
- Global Variables: Identified and initialized, as their state can significantly influence function behavior.
Constraints Tracer
Even with a well-sliced context and assigned variable values, the initial fuzzing harness might produce false positive crashes. These occur when the fuzzer triggers a crash, but the program state leading to it is not one that would typically occur in a real-world scenario, often due to a lack of specific program constraints. The Constraints Tracer is designed to identify and incorporate these crucial constraints back into the harness, significantly improving the precision of vulnerability discovery.
The tracer operates efficiently by focusing only on variables directly related to a crash:
- Crash Point Identification: When a crash occurs, AFGen identifies the exact statement where the crash happened and the variables involved in that statement.
- Control Statement Variables: It also considers variables within control statements (e.g.,
if,whileconditions) that directly influence the execution path leading to the crash. - Constraint Tracing: For these identified variables, AFGen performs state flow and control flow analyses to locate their definitions and track how their values are derived. It records:
- Assignment statements: Where a variable's value is set.
- Control flow dependencies: The conditions that must be met for a specific assignment or execution path to be taken.
These recorded assignments and control flow dependencies constitute the "constraints."
- Constraint Adjustment: The traced constraints cannot be directly added back to the harness. They require adjustment to fit the harness context:
- Variable Name Resolution: Original variable names are mapped to their corresponding names within the sliced harness.
- New Control Variable Removal: Any new control variables introduced solely for tracing are removed.
- Input Variable Skipping: Variables that are meant to be fuzzed inputs are skipped to avoid over-constraining the fuzzer.
- Harness Refinement: Finally, the adjusted constraints are added back into the fuzzing harness, often as conditional checks or specific initialization logic. The refined harness is then re-run with the fuzzer. If the crash can still be triggered under these stricter, more realistic conditions, it confirms a real vulnerability, filtering out false positives.
By integrating these three sophisticated components, AFGen provides a robust and intelligent framework for automatic fuzzing harness generation, enabling comprehensive and precise whole-function fuzzing for complex software.
Demo / Proof of Concept
▶ Watch: Tracing constraints to filter false positive crashes (9:00)
The core concepts of AFGen are effectively demonstrated through a motivational example involving a get_byte function from a library, which the speaker refers to as "life 14" (likely a transcription error for a library like libflif or a version of libssl given the Heartbleed context). This example clearly illustrates the challenges of whole-function fuzzing and how AFGen's components address them.
The get_byte function is designed to read a single byte from an input source. It operates in two distinct modes: a memory mode, where it reads from an internal array, and a file mode, where it reads from a file handler using getc. The mode is typically determined during the compile process or via configuration.
Initially, to build a fuzzing harness for get_byte, AFGen's Bidirectional Slicer would extract the relevant code context, including the function definition and any immediate callers or data dependencies. The Variable Value Assignment module would then initialize variables like the input array (for memory mode) or a file pointer (for file mode).
With this basic harness, a fuzzer might quickly trigger a heap buffer overflow in the memory mode. The speaker explains that while both memory and file modes might lack boundary checks, the bug specifically manifests in memory mode. In file mode, getc returns an end-of-file (EOF) marker when no more data is available, preventing an out-of-bounds read. However, in memory mode, without explicit boundary checks, an attacker-controlled input could cause get_byte to read beyond the allocated array, leading to a crash.
At this stage, the crash is observed, but it's crucial to determine if it's a real vulnerability or a false positive due to an unrealistic program state created by the harness. This is where AFGen's Constraints Tracer comes into play. Upon detecting the crash, the tracer would analyze the crash point and related variables. It would identify that the crash occurs when get_byte is operating in memory mode and attempting to access an out-of-bounds memory region. The tracer would then determine the conditions that lead to get_byte entering and executing its memory mode logic.
The traced constraints, such as ensuring the mode flag is explicitly set to memory mode, would then be added back to the fuzzing harness. This refined harness would now test get_byte under conditions that more accurately reflect its intended operational context. When the fuzzer is run again with this constrained harness, if the heap buffer overflow can still be triggered, it provides strong confirmation that a genuine vulnerability exists within the get_byte function under specific, yet plausible, conditions. This process effectively filters out false positives and validates the security flaw.
The example highlights how AFGen intelligently moves from a general crash detection to a confirmed vulnerability by systematically understanding and enforcing the necessary program constraints. This iterative refinement process, driven by AFGen's three core components, is central to its high precision and effectiveness in discovering real-world vulnerabilities. The talk also mentions that AFGen can reproduce 66 CVEs, including those similar to the infamous Heartbleed (CVE-2014-0160), which was a heap buffer overflow in libssl's TLS heartbeat extension. This further underscores the practical applicability of AFGen's methodology to critical security flaws.
Defensive Implications
▶ Watch: AFGen's superior performance in discovering vulnerabilities (11:00)
The advent of tools like AFGen, capable of whole-function fuzzing, carries significant implications for software defenders, development teams, and security researchers. Its ability to plumb the depths of internal function logic necessitates a re-evaluation of current defensive strategies and offers new avenues for proactive security enhancement.
- Prioritize Internal Function Auditing: Traditional security audits and fuzzing efforts often concentrate on exposed APIs and application entry points. AFGen demonstrates that a vast landscape of vulnerabilities resides within internal, undocumented functions. Defenders must shift their focus to systematically audit and fuzz these internal components, recognizing that they form the bedrock of an application's security posture. This means moving beyond black-box testing to incorporate white-box and grey-box techniques that can interact directly with internal functions.
- Integrate Whole-Function Fuzzing into CI/CD: The speed and precision of AFGen make it an ideal candidate for integration into Continuous Integration/Continuous Deployment (CI/CD) pipelines. Automating the generation of fuzzing harnesses for new or modified functions within a library or application can provide immediate feedback to developers on potential vulnerabilities, catching bugs early in the development lifecycle before they become costly to fix. This proactive approach can significantly reduce the attack surface.
- Enhance Library and Component Security: Many applications rely heavily on third-party libraries. While these libraries are often well-tested, AFGen's findings suggest that even widely used libraries may harbor deep-seated vulnerabilities that traditional fuzzers miss. Organizations should consider employing whole-function fuzzing against their critical third-party dependencies to identify and mitigate risks that might otherwise go unnoticed. This is particularly relevant for libraries handling sensitive data or network communications.
- Review and Strengthen Internal APIs: Although AFGen focuses on all functions, its ability to generate precise harnesses can also be applied to internal APIs that are not exposed externally but are crucial for inter-component communication. Strengthening the security of these internal interfaces is paramount, as a compromise in one component can cascade throughout the system.
- Develop Robust Error Handling and Input Validation for Internal Functions: The vulnerabilities discovered by AFGen often stem from a lack of boundary checks, improper type handling, or insufficient state validation within internal functions. Developers should adopt a "zero-trust" approach even for internal inputs, implementing stringent validation and robust error handling mechanisms for all function arguments and internal state transitions, regardless of whether they are directly exposed to untrusted external input.
- Leverage Fuzzing for Security Regression Testing: As codebases evolve, new vulnerabilities can be inadvertently introduced or old ones can resurface. AFGen's capacity to quickly reproduce known CVEs makes it an excellent tool for security regression testing. By building harnesses for previously discovered vulnerabilities, development teams can ensure that fixes remain effective and that similar flaws are not reintroduced in subsequent updates.
- Invest in Static Analysis Tools that Complement Fuzzing: While fuzzing is dynamic, its effectiveness can be amplified when combined with static analysis. Insights from AFGen's slicer and constraint tracer can inform the development or configuration of static analysis tools to identify potential areas of concern that might benefit from targeted fuzzing, creating a more comprehensive security testing strategy.
By adopting these defensive implications, organizations can move towards a more holistic and proactive security posture, safeguarding their applications and libraries against the sophisticated vulnerabilities that AFGen is designed to uncover.
Key Takeaways
- Whole-Function Fuzzing is Critical: Traditional fuzzing often misses vulnerabilities in internal functions due to incomplete code coverage and the difficulty of generating appropriate program contexts. AFGen's approach to fuzzing all functions systematically addresses this gap.
- Automated Harness Generation is Achievable for All Functions: AFGen demonstrates that intelligent automation can overcome the challenges of constructing fuzzing harnesses for internal functions, including context slicing, variable assignment, and constraint tracing, which were previously major hurdles.
- AFGen Significantly Outperforms Existing Fuzzers: It reproduced 66 known CVEs and discovered 26 new vulnerabilities (24 with CVE IDs), outperforming both general and directed fuzzers in terms of coverage, speed, and precision.
- Precision is Boosted by Constraint Tracing: The Constraints Tracer component is vital for filtering false positives, ensuring that crashes identified by AFGen's harnesses represent genuine vulnerabilities by enforcing realistic program conditions.
- Deep Code Paths are Now Accessible: AFGen's ability to reach code paths requiring high branch counts and function nested depths (e.g., 68.6 branches and 7.2 function nested depths for vulnerabilities found) opens up previously unexplorable areas for automated security testing.
- Proactive Security for Libraries and Applications: Integrating whole-function fuzzing into development pipelines can enable earlier detection of vulnerabilities, leading to more robust and secure software components.
About the Speaker(s)
The talk "AFGen: Whole-Function Fuzzing for Applications and Libraries" was presented by Yuwei Liu from the University of Chinese Academy of Sciences. This work represents a collaborative effort with other researchers, including Yanhao Wang, Xiangkun Jia, Zheng Zhang, and Purui Su, from both the University of Chinese Academy of Sciences and Ocean University of China. The research group focuses on advancing fuzzing technology to improve software security, addressing complex challenges in vulnerability discovery and automated security testing. Their work contributes significantly to the field of software engineering and cybersecurity, aiming to make deep code analysis more accessible and effective.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This research on AFGen presents a groundbreaking approach to whole-function fuzzing, tackling the long-standing challenge of automated harness generation for all internal functions. By combining intelligent slicing, variable assignment, and crucial constraint tracing, AFGen achieves superior vulnerability discovery and precision, making it a critical advancement for software security.
Heather Calloway (CISO) — STRONG ACCEPT
This research presents a significant advancement in automated vulnerability discovery by enabling comprehensive, whole-function fuzzing. AFGen offers a practical solution to a long-standing challenge, demonstrating superior precision and speed in uncovering critical flaws across applications and libraries. It demands a re-evaluation of current defensive strategies, particularly for internal code and third-party dependencies.
→ Top-rated talks at IEEE Symposium on Security and Privacy 2024