Automatic Library Fuzzing through API Relation Evolvement
Jiayi Lin (University of Hong Kong)
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Fuzzing 2
Overview
Software libraries form the foundational components of countless applications, yet they often harbor complex vulnerabilities that are difficult to uncover through traditional testing methods. This talk introduces Nexer, an innovative fuzzing tool designed to automatically identify vulnerabilities within these critical software libraries by addressing a long-standing challenge: balancing input diversity with API usage accuracy. Presented by Jiayi Lin from the University of Hong Kong, this research highlights the limitations of existing library fuzzing approaches which often sacrifice one for the other, leading to either limited bug discovery or a deluge of false positives from API misuses.
Key moments
- 0:50 Trade-offs in existing library fuzzing methods
- 2:15 Nexer's three key new design principles
- 2:50 Four-step overview of Nexer's fuzzing method
- 5:00 Advantages of Nexer's modular driver architecture
- 6:20 Nexer's novel dynamic API usage learning stage
- 7:50 Real-world API relation examples from OpenSSL
- 8:40 Updating call sequence generation strategies based on learned relations
Automatic Library Fuzzing through API Relation Evolvement
Speakers: Jiayi Lin (University of Hong Kong)
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=GQ0USeLvqbM
Overview
Software libraries form the foundational components of countless applications, yet they often harbor complex vulnerabilities that are difficult to uncover through traditional testing methods. This talk introduces Nexer, an innovative fuzzing tool designed to automatically identify vulnerabilities within these critical software libraries by addressing a long-standing challenge: balancing input diversity with API usage accuracy. Presented by Jiayi Lin from the University of Hong Kong, this research highlights the limitations of existing library fuzzing approaches which often sacrifice one for the other, leading to either limited bug discovery or a deluge of false positives from API misuses.
Nexer aims to overcome this trade-off by employing a novel modular driver architecture, combining static analysis with dynamic learning for API usage, and implementing sophisticated rule-based filtering strategies. The core problem Nexer tackles is the intricate interdependencies among library APIs, which necessitate precise invocation sequences and argument constraints. By intelligently generating diverse yet semantically correct API call sequences and effectively distinguishing real vulnerabilities from erroneous API usage, Nexer significantly enhances the efficiency and effectiveness of library fuzzing, promising to make a substantial impact on software supply chain security.
Background
▶ Watch: Trade-offs in existing library fuzzing methods (0:50)
Library fuzzing is a critical software testing technique aimed at uncovering vulnerabilities within software libraries. Unlike application-level fuzzing, library fuzzing presents unique challenges primarily due to the nature of library APIs. Libraries typically expose a multitude of interdependent APIs, meaning that their correct invocation often requires specific calling sequences and carefully constrained arguments. For instance, a function might require another function to be called first to initialize a data structure, or an argument might need to fall within a certain range to prevent crashes. Traditionally, fuzzing engines like LibFuzzer generate random bytes, necessitating manually crafted drivers to properly constrain and feed inputs into API arguments and manage calling sequences. This manual effort is a significant bottleneck, limiting the scale and scope of library fuzzing.
The existing landscape of automatic library fuzzing methods generally faces a fundamental trade-off between input diversity and API usage accuracy. Tools such as Hopper and API excel at exploring a wider input space, including diverse argument values and mutable API sequences. However, this often comes at the cost of generating erroneous API usage, leading to a high rate of false crashes. These crashes are not indicative of real vulnerabilities but rather stem from API misuse, such as providing an incorrect argument type or an out-of-bounds value. A pertinent example cited in the talk involves mutating the third argument in a gipsy API, which might cause a memory overflow crash that is merely an API misuse, not a genuine bug within the library's intended functionality. This inundation of false positives places a heavy burden on security researchers for manual post-processing and triage.
Conversely, other approaches like Utopia and Rubik prioritize API usage accuracy. These tools focus on generating semantically correct API calls, thereby reducing false positives. However, this accuracy often comes with a significant limitation: restricted API coverage and sequence flexibility. By adhering strictly to known correct usage patterns, these tools may miss vulnerabilities that only manifest under unusual or unexpected API call sequences or argument combinations. Such an approach inherently limits the exploration of the vast potential input space, leaving many parts of the library's attack surface untested. The persistent challenge, therefore, has been to develop a library fuzzing solution that can effectively balance these two critical aspects, enabling wide-ranging exploration without drowning in irrelevant crash reports.
Key Findings
▶ Watch: Four-step overview of Nexer's fuzzing method (2:50)
Nexer's evaluation demonstrates its significant advancements in addressing the core challenges of automated library fuzzing. The research highlights three primary contributions that collectively enable a more effective and efficient vulnerability discovery process: a modular and scalable fuzzing driver architecture, a dual approach combining static analysis and dynamic learning for API usage, and a robust rule-based automatic filtering strategy for API misuse.
Through extensive testing on 18 diverse libraries, Nexer successfully identified 27 new vulnerabilities within a 24-hour fuzzing period per library. Notably, 5 CVEs have already been assigned, underscoring the severity and novelty of the discovered flaws. When benchmarked against existing state-of-the-art fuzzing tools, including both manual drivers from projects like OSS-Fuzz and automatic methods such as Third Gen, Utopia, and Hopper, Nexer exhibited superior performance. It achieved 48% more code coverage and uncovered 19 additional previously unknown vulnerabilities, demonstrating its enhanced exploration capabilities and ability to reach deeper into library codebases.
A crucial finding relates to Nexer's effectiveness in managing API misuse. The tool demonstrated an impressive ability to filter out 93% of API misuse-related crashes, a stark contrast to other similar works that could only filter approximately 47%. This significant reduction in false positives drastically minimizes the manual post-processing and triage effort required by security analysts, making the fuzzing process far more practical and scalable.
An ablation study further elucidated the contribution of Nexer's individual components. The static consumer analysis stage, which learns API usage patterns from existing code, was found to improve code coverage by 10% and led to the discovery of 8 additional bugs. The dynamic learning stage, which refines API understanding through runtime observations, proved even more impactful, boosting coverage by 25% and revealing 9 additional bugs. The observed decreasing trend in the API misuse ratio when API relation learning was enabled visually confirmed that Nexer effectively learns correct API relations, progressively generating more targeted and effective API call sequences.
Beyond quantitative metrics, the research also yielded insightful qualitative findings. Many of the discovered bugs resided in uncovered APIs, particularly in large libraries like OpenSSL, which boasts thousands of APIs but only has hundreds covered by existing manual drivers. This highlights the ongoing importance of effective library fuzzing automation. Furthermore, several critical bugs were uncovered specifically through API sequence mutation. For example, a hidden bug in OpenSSL's big number API could not be triggered by fixed sequences found in existing unit tests or manual drivers. Nexer's ability to explore diverse input spaces by replacing the first API in a sequence (e.g., with another big number set API) successfully triggered this hidden vulnerability, emphasizing the power of dynamic sequence exploration.
Technical Deep Dive
▶ Watch: Advantages of Nexer's modular driver architecture (5:00)
Nexer's architecture is meticulously designed to address the inherent complexities of library fuzzing by balancing input diversity with API usage accuracy through three core innovations: a modular driver, a comprehensive API usage learning mechanism, and a precise filtering strategy.
The overall methodology comprises four main steps:
- Static Analysis: Nexer first performs static analysis on API consumers (e.g., library unit tests) and the library's source code to generate detailed API descriptions. This foundational step helps synthesize initial code for the fuzzing driver.
- Call Sequence Generation: The fuzzing engine, now API-aware, directly generates concrete API call sequences, complete with specific argument values and invocation orders. This is a significant departure from traditional fuzzers that generate raw bytes.
- Interpretation and Execution: A general interpreter then processes these generated sequences, performing low-level tasks such as setting memory layouts and managing inter-API dependencies. It then invokes the APIs through the synthesized driver code.
- Feedback Collection and Learning: During execution, the driver collects detailed feedback, including crash types, call stack frames, and memory violation addresses. This feedback is used by the fuzzing engine to categorize crashes, learn API relations, and dynamically adjust its generation strategies.
Modular Driver Architecture
Nexer's modular driver architecture is a cornerstone of its scalability and robustness, addressing the limitations of traditional monolithic drivers. Monolithic drivers intertwine all fuzzing tasks—assigning random values, randomizing/constraining call sequences, ensuring inter-API dependencies, and invoking APIs—into a single, static block of code. This design makes them highly target-specific, difficult to scale, prone to errors, and challenging to automatize. A single API misuse within a monolithic driver can render the entire fuzzer unusable until manual fixes are applied, as it would repeatedly crash at the same location.
Nexer disaggregates these tasks into distinct, specialized components:
- API-Aware Fuzzing Engine: Unlike traditional fuzzers that produce raw bytes, Nexer's engine, facilitated by automatically generated API descriptions, directly generates explicit API calls with concrete argument values and sequence orders. This higher-level abstraction makes it easier to adopt different libraries and allows for dynamic adjustment of API usage during fuzzing without modifying low-level code.
- General Interpreter: This component is target-agnostic. It takes the generated API call sequences and performs universal, low-level jobs. These include managing memory layouts (e.g., allocating buffers, setting up structures) and transferring data dependencies between APIs (e.g., ensuring an output from one API serves as a valid input for the next). Because it's generic, this part of the code does not need to be regenerated for each new library.
- Synthesized Driver Code: This is the only part that needs to be generated per library, but its scope is significantly reduced. It merely contains the API call entries with correct function signatures, acting as a lightweight wrapper for the actual library functions. This modularity makes the synthesis process much simpler, more scalable, and less error-prone.
By distributing tasks, Nexer achieves greater automatization, allows dynamic adjustments during fuzzing, and ensures the core interpreter remains reusable across diverse libraries.
API Usage Learning
Nexer employs a dual-stage approach to learn API usage, combining static analysis for initial understanding with dynamic learning for refinement.
- Static Consumer Analysis: This initial stage leverages backward data flow slicing on existing API consumers (e.g., unit tests or example code) and the library's source code. It collects fundamental API attributes, such as common argument constants, valid value ranges, and explicit data flow dependencies between API calls. This provides a baseline understanding of how APIs are typically used.
- Dynamic Learning Stage: This is where Nexer's novel techniques shine. The fuzzer tentatively mutates or minimizes API arguments and sequence orders and then observes the execution behavior. This iterative process helps refine its understanding of complex API relations.
- Identifying Invalid Ages: If an execution crashes at
API_5within a sequence, Nexer attempts to minimize the sequence by deleting preceding calls one by one. If deletingAPI_4resolves the crash (i.e.,API_5then executes correctly), this indicates an invalid age:API_4somehow creates a state that causesAPI_5to fail. These invalid relations are recorded and maintained in an API graph within the fuzzing engine. For example, in OpenSSL, thebig number freeAPI cannot precede otherbig numberAPIs likeaddorsubtractwithout causing issues. - Identifying Valid Ages: Conversely, if
API_5initially executes normally, but then crashes afterAPI_4is deleted, this signifies a valid age:API_4is a prerequisite forAPI_5to function correctly. An example from OpenSSL is that thedecrypt initializeAPI must precede thedecrypt updateAPI.
After learning these relations, Nexer employs heuristic fuzzing rules to categorize them and update its call sequence generation strategies. The system leverages collective feedback, including call stack frames, error types reported by sanitizers (like Address Sanitizer), and concrete memory violation addresses. For instance, if an invalid age is found where the source node (API_4) involves a free call, and the subsequent crash is a memory violation, Nexer can categorize this as a Use-After-Free (UAF) misuse. This allows the fuzzer to avoid generating such sequences in the future or to prioritize sequences that might expose genuine UAF vulnerabilities. The paper provides more details on how argument constraints are learned by inserting red zones and examining memory violation addresses.
Rule-Based Automatic Filtering
The third key design element is Nexer's sophisticated rule-based automatic filtering strategy for API misuse. This component is crucial for reducing the burden of manual triage. Based on the categorized API relations and detailed execution feedback, Nexer applies a series of carefully designed filtering rules. These rules analyze the nature of crashes and their associated API usage patterns to distinguish between true vulnerabilities and mere misuses. For example, a crash identified as a UAF misuse (based on the dynamic learning and heuristic rules) would be filtered out if it clearly stems from an incorrect API calling pattern rather than a genuine bug in the library's memory management.
Crashes that do not fit into known misuse patterns, or those that exhibit characteristics of actual security flaws, are marked as suspicious vulnerabilities for manual review. This targeted approach ensures that human effort is focused on investigating potentially critical bugs, rather than sifting through thousands of easily identifiable API misuse reports.
Demo / Proof of Concept
▶ Watch: Real-world API relation examples from OpenSSL (7:50)
While the talk focuses on the architectural and algorithmic aspects of Nexer, it does not describe a live demonstration or a detailed, step-by-step proof-of-concept walk-through during the presentation. Instead, the effectiveness of Nexer's methodology and its ability to uncover real-world vulnerabilities are primarily substantiated through comprehensive evaluation results and example figures.
The speaker refers to "example figures" (11:00) that illustrate the decreasing trend of API misuse ratio when API relation learning is enabled, visually reinforcing the tool's ability to learn correct API usage. The discussion of specific bug findings, such as the hidden bug in OpenSSL's big number API triggered by sequence mutation (11:40), serves as a strong testament to Nexer's capabilities, demonstrating its practical impact without a direct live demo. The evaluation section's quantitative data on new vulnerabilities found, CVEs assigned, and improved code coverage acts as the primary evidence of Nexer's proof of concept.
Defensive Implications
▶ Watch: Updating call sequence generation strategies based on learned relations (8:40)
The advent of tools like Nexer carries significant implications for software defenders and developers aiming to secure their library dependencies. The core message is the urgent need to move beyond traditional, often manual, testing paradigms towards automated, intelligent library fuzzing.
- Integrate Advanced Fuzzing into CI/CD: Organizations should consider integrating automated library fuzzing tools like Nexer into their continuous integration and continuous deployment (CI/CD) pipelines. Proactive fuzzing can catch vulnerabilities early in the development lifecycle, significantly reducing the cost and impact of remediation.
- Prioritize Libraries with Low Coverage: The finding that many bugs reside in uncovered APIs (e.g., OpenSSL having thousands of APIs but only hundreds covered by manual drivers) is a critical insight. Defenders should identify libraries within their software supply chain that have limited test coverage and prioritize them for automated fuzzing efforts. This strategy targets areas most likely to harbor hidden vulnerabilities.
- Recognize the Importance of API Sequence Mutation: Developers and security engineers must understand that vulnerabilities are not always triggered by single, isolated API calls but often by specific, non-obvious sequences of API invocations. Traditional unit tests, which often rely on fixed, expected sequences, may entirely miss these types of bugs. Fuzzing tools capable of intelligently mutating API call sequences, like Nexer, are indispensable for uncovering such hidden flaws.
- Leverage Intelligent Filtering to Reduce Alert Fatigue: The high rate of false positives from API misuse has historically been a barrier to effective fuzzing. Nexer's ability to filter out 93% of API misuse means that security teams can focus their limited resources on investigating genuine, suspicious vulnerabilities rather than sifting through noise. Defenders should seek out fuzzing solutions that incorporate robust filtering mechanisms to maximize efficiency.
- Understand API Relations for Robust Design: The principles behind Nexer's API usage learning—identifying both valid and invalid API ages—can inform more robust library design. Developers can proactively document or enforce these relations, potentially using static analysis during development, to prevent common misuse patterns that lead to crashes or undefined behavior.
- Address Limitations and Future Work: While powerful, Nexer has limitations, such as modeling complex data structures (multi-dimensional arrays, length-prefixed buffers) and limited C++ support. Defenders should be aware of these current boundaries and complement automated fuzzing with other security testing techniques where necessary, while also advocating for continued research in these areas.
In essence, Nexer provides a blueprint for a more effective and scalable approach to library security, shifting the focus from labor-intensive manual efforts to intelligent, automated discovery of critical vulnerabilities.
Key Takeaways
- Automated library fuzzing is critical for complex libraries: Manual driver creation and existing unit tests often miss a vast attack surface, leaving many APIs untested and vulnerable.
- Balancing input diversity and API usage accuracy is essential: Effective fuzzing requires exploring a wide range of inputs and sequences while simultaneously generating semantically correct API calls to avoid false positives.
- Modular driver architectures enhance scalability and automatization: Breaking down fuzzing tasks into API-aware engines, target-agnostic interpreters, and minimal synthesized code makes fuzzing easier to adapt and scale across different libraries.
- Both static and dynamic analysis are vital for learning API relations: Static analysis provides an initial understanding of API attributes, while dynamic learning refines this understanding by observing execution behavior to identify valid and invalid API call dependencies.
- Effective filtering of API misuse significantly reduces triage burden: Nexer's ability to filter out 93% of API misuse reports allows security teams to focus on investigating genuine vulnerabilities, making the fuzzing process more efficient.
- API sequence mutation is a powerful technique for uncovering hidden vulnerabilities: Many critical bugs are only exposed when APIs are invoked in specific, non-obvious sequences, which intelligent fuzzers can discover.
About the Speaker(s)
The research presented on "Automatic Library Fuzzing through API Relation Evolvement" was delivered by Jiayi Lin, who is affiliated with the University of Hong Kong. This work is a collaborative effort, undertaken jointly with colleagues from both the University of Hong Kong and the Hong Kong Polytechnic University, highlighting a concerted academic approach to advancing software security.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Competent academic fuzzing research with real results — 27 bugs, 5 CVEs, 93% misuse filtering — but the core ideas (static+dynamic API learning, modular driver decomposition) are evolutionary refinements of a well-trodden space rather than a conceptual leap. Solid NDSS-tier work that belongs in the literature; whether it belongs on a conference stage depends entirely on what else is in the program.
Heather Calloway (CISO) — WEAK
Technically credible research with real results — 27 vulnerabilities, 5 CVEs, measurable improvements over prior art. But this is a tools paper aimed at fuzzing researchers, not security operators or program leaders, and the gap between the technical contribution and any actionable institutional guidance is never bridged.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025