Uncovering the iceberg from the tip: Generating API Specifications for Bug Detection via Specification Propagation Analysis
Miaoqian Lin
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · API Security
Overview
In the realm of software security, the correct and safe usage of Application Programming Interfaces (APIs) is paramount. However, the intricate nature of APIs, especially in low-level languages like C, often leads to misuse, resulting in critical vulnerabilities such as reference count errors, memory leaks, and improper resource handling. The challenge lies in the fact that many crucial API specifications—the rules governing their correct usage—are either poorly documented, hidden within complex code, or simply not observed through typical usage patterns. This talk, presented by Miaoqian Lin at the NDSS Symposium, introduces API spec, a novel framework designed to address this pervasive problem by inferring hidden API specifications through a technique called specification propagation analysis.
Key moments
- 0:00 Introduction to hidden API specifications problem
- 2:00 Defining API specification using a triplet focus
- 3:00 Concept of specification propagation via API call chains
- 4:20 Introducing API Spec and its bug detection example
- 6:00 Technical details of specification propagation analysis
- 7:10 Generating post-operations using usage and data flow
- 8:30 Evaluation results: specifications generated, bugs detected
- 10:00 Comparison with prior API specification extraction methods
Uncovering the iceberg from the tip: Generating API Specifications for Bug Detection via Specification Propagation Analysis
Speakers: Miaoqian Lin, PhD student, Chinese Academy of Science
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=ee4JvBDkQzA
Overview
In the realm of software security, the correct and safe usage of Application Programming Interfaces (APIs) is paramount. However, the intricate nature of APIs, especially in low-level languages like C, often leads to misuse, resulting in critical vulnerabilities such as reference count errors, memory leaks, and improper resource handling. The challenge lies in the fact that many crucial API specifications—the rules governing their correct usage—are either poorly documented, hidden within complex code, or simply not observed through typical usage patterns. This talk, presented by Miaoqian Lin at the NDSS Symposium, introduces API spec, a novel framework designed to address this pervasive problem by inferring hidden API specifications through a technique called specification propagation analysis.
API spec leverages the observation that API specifications can propagate through hierarchical call chains. By starting with a small set of known (seed) specifications, the system intelligently traces how these rules might apply to other, higher-level or related APIs within a codebase. This allows for the automatic generation of thousands of new, previously unknown specifications, which are then used to detect bugs that traditional methods, constrained by available documentation or usage data, consistently miss. The research demonstrates a significant advancement in automated bug detection, offering a potent tool for enhancing software reliability and security.
The core contribution of this work is its ability to "uncover the iceberg from the tip"—inferring a vast number of hidden specifications from a minimal initial set. This approach not only provides a scalable solution to the problem of missing API specifications but also proves highly effective in identifying real-world vulnerabilities. The successful detection of 186 new bugs in the Linux kernel, many of which have been confirmed and fixed, underscores the practical impact and necessity of API spec in modern software development and security auditing.
Background
▶ Watch: Introduction to hidden API specifications problem (0:00)
APIs are fundamental building blocks of modern software, enabling complex functionalities through modular design. However, their power comes with a significant responsibility: developers must adhere strictly to their specifications to prevent misuse and subsequent bugs. A common example cited in the talk involves resource management, where an API call like get_device that acquires a resource (e.g., a reference count) must be paired with a corresponding put_device call to release it. Failure to do so leads to resource leaks or reference count bugs, which can degrade system performance, lead to denial-of-service, or even open doors for more severe exploits.
Traditionally, detecting such API misuse relies heavily on explicit API specifications. These specifications can come from various sources: formal API documentation, analysis of frequent API usage patterns in existing codebases, or examination of bug patches that highlight previous misuse. While these methods have proven useful to some extent, they suffer from a critical limitation: they depend entirely on the availability and completeness of these "API artifacts." Many specifications, particularly those governing complex interactions or subtle side effects, are not formally documented, are rarely used in a way that makes their correct pattern obvious, or are only discovered after a bug has already manifested and been patched. The speaker noted that "many specifications are hidden and cannot be observed through internal artifact," rendering artifact-based detection methods ineffective for a substantial portion of potential API misuse.
This research specifically focuses on API post-handling specifications, which dictate operations that must be performed after an API call, such as releasing resources, checking error codes, or updating states. This category of specifications is particularly prone to errors and accounts for over 16% of API misuse cases observed in C language projects. The talk represents an API specification as a triplet: (target API, critical variable, post-operation). For instance, for get_device, the critical variable might be dev (the returned device reference), and the post-operation would be put_device. Without a clear understanding of these triplets, automated bug detection tools struggle to identify violations. The core motivation behind API spec is to overcome the inherent limitations of artifact-dependent methods by inferring these crucial, often hidden, specifications from a limited set of known ones, thereby improving the efficacy of bug detection.
Key Findings
▶ Watch: Concept of specification propagation via API call chains (3:00)
The central insight and most significant finding of this research is the concept of API specification propagation. The speakers observed that API specifications are not isolated but can "propagate through the API call chain as API are often designed in a hierarchical manner." When a higher-level API calls a lower-level one, it may inherit the specifications of that lower-level API. This propagation mechanism forms the foundation of API spec, the framework introduced to generate new specifications from a small set of known "seed" specifications.
The evaluation of API spec on the Linux kernel yielded several compelling results:
- Massive Specification Generation: From an initial set of just 6 seed specifications, API spec successfully generated a total of 7332 new API specifications. This demonstrates an extraordinary ability to extrapolate knowledge from minimal input, effectively "uncovering the iceberg" of hidden rules.
- Efficient Processing: The entire generation process for these thousands of specifications took only 2 hours, averaging approximately 1 second per specification. This highlights the scalability and practical applicability of the API spec framework for large and complex codebases.
- High Accuracy: A manual check performed on a random sample of the generated specifications confirmed their correctness, validating the reliability of the propagation analysis.
- Significant Bug Detection: Using the newly generated specifications, API spec detected 186 new bugs in the Linux kernel. These bugs were primarily related to improper post-handling for target APIs, such as unreleased reference counts. Crucially, "most of them have has been confirmed and fixed," underscoring the real-world impact and severity of the vulnerabilities identified.
- Limitations of Prior Work: The research directly compared API spec against state-of-the-art methods that rely on API documents, bug patches, and frequent API usage. The results showed that these traditional methods "fail to extract most of the generated sophistication and fail to detect the corresponding backs." Specifically, over 18% of the generated specifications were not mentioned in API documentation, and over 90% were not frequently found in the code. This decisively demonstrates that artifact-dependent methods are severely limited and that API spec addresses a critical gap.
- Cross-Project Applicability: While primarily evaluated on the Linux kernel, the concept of specification propagation was also tested across different open-source projects, including OpenSSL and FFmpeg, indicating its broad applicability beyond kernel code. The speaker also suggested its potential effectiveness in user-space programs and even in languages with more robust resource management features like Python, albeit with different bug patterns.
These findings collectively establish API spec as a powerful and effective tool for automated security analysis, capable of uncovering a vast landscape of hidden API specifications and, consequently, a significant number of previously undetected bugs.
Technical Deep Dive
▶ Watch: Technical details of specification propagation analysis (6:00)
The technical core of API spec lies in its specification propagation analysis and subsequent specification generation. The framework operates on the principle that if an API (successor API) calls another API (predecessor API) that has a known specification, the successor API might inherit that specification under specific conditions.
The propagation of an API specification (target API, critical variable, post-operation) from a predecessor API to a successor API is defined by three critical conditions:
- API Call Chain: The
successor APImust call thepredecessor API. This call introduces an effect on thecritical variabledefined by thepredecessor API's specification. For example, ifget_devicehas a specification (get_device, dev, put_device), andbusify_devicecallsget_device, thenbusify_deviceis a potential successor. - Critical Variable Propagation: The
critical variable(e.g.,devfromget_device) must be propagated to thesuccessor API. This propagation can occur either as a return value of thesuccessor APIor as one of its arguments. If the critical variable is not accessible or passed out of the successor, the specification cannot propagate. - No Intermediate Disruption: Between the call to the
predecessor APIand the return of thesuccessor API, there must be no intermediate operations that disrupt or alter the state of thecritical variablein a way that invalidates the original specification. For instance, if the critical variable is freed or modified in an incompatible way before being propagated, the specification cannot be directly inherited.
API spec implements this propagation analysis in two main steps:
Step 1: Specification Propagating Analysis
This step identifies which APIs are likely to inherit or originate from a given seed specification.
- Caller Analysis: Given a seed API (e.g.,
get_device), API spec first performs a caller analysis to identify all functions that call the seed API. These callers are potentialsuccessor APIs. The talk showed an example whereget_devicewas called by four functions, which become initial candidates. - Critical Variable Propagation Check: For each candidate
successor API, API spec checks if thecritical variable(e.g.,dev) is correctly propagated. Functions that do not propagate the critical variable are filtered out. - Intermediate Operation Check: Next, the tool verifies that no intermediate operations within the
successor APIdisrupt the specification's propagation. Functions failing this check are also discarded. - Iterative Propagation: This process is performed iteratively. If
busify_deviceis identified as a successor ofget_deviceand inherits its specification, thenbusify_deviceitself can become a "seed" for further propagation. For example, ifnfc_get_devicecallsbusify_device, thennfc_get_devicemight inherit the specification frombusify_device, which originally came fromget_device. This iterative nature allows API spec to build long call chains and uncover deep propagation paths.
Step 2: Specification Generation
Once an API is identified as inheriting a specification, the next challenge is to determine its corresponding post-operation. The post-operation might change during propagation, even if the underlying resource management logic remains similar. For instance, get_device has put_device, but nfc_get_device might require nfc_put_device.
To determine the correct post-operation, API spec combines API usage analysis and data flow validation:
- API Usage Analysis: For the newly inferred
successor API(e.g.,nfc_get_device), API spec analyzes the code for functions that are frequently called after it, especially those that take thecritical variableas an argument. These functions become candidates for the post-operation. - Data Flow Validation: Each candidate post-operation is then validated against the original seed's post-operation (
put_devicein our example). The validation checks if the candidate post-operation (e.g.,nfc_put_device) eventually passes thecritical variableto the original seed's post-operation (put_device) or a semantically equivalent function. This confirms that the candidate indeed fulfills the original post-handling requirement, albeit potentially through an abstraction layer. For example,nfc_put_devicemight internally callput_device, confirming its role as the correct post-operation fornfc_get_device.
By combining these steps, API spec can accurately generate new specification triplets, such as (nfc_get_device, dev, nfc_put_device), which are then ready for use in bug detection. This methodical approach ensures that the generated specifications are not only numerous but also semantically correct and actionable for identifying real-world vulnerabilities.
Demo / Proof of Concept
▶ Watch: Generating post-operations using usage and data flow (7:10)
The talk presented a compelling proof of concept through the identification and confirmation of a new bug in the Linux kernel. This particular bug served as a direct demonstration of how a specification inferred by API spec could lead to the detection of a critical vulnerability that would likely be missed by traditional analysis methods.
The example highlighted involved the nfc_get_device API in the Linux kernel. Through its propagation analysis, API spec inferred a new specification for nfc_get_device: that after calling it, developers should perform a nfc_put_device operation on the returned device reference (dev). This specification was derived from the propagation of an earlier specification from get_device to busify_device, and then to nfc_get_device, with nfc_put_device being identified as the correct post-operation through usage and data flow validation.
With this newly generated specification, API spec then performed bug detection. It identified a specific code path where nfc_get_device was called, the critical variable dev was subsequently used, but the corresponding nfc_put_device was not invoked. This omission led to a reference count bug, meaning the device reference was acquired but never properly released, potentially leading to resource exhaustion or other stability issues within the kernel.
The speaker explicitly stated that this bug "has been confirmed and fixed," providing concrete evidence of API spec's ability to uncover real, impactful vulnerabilities. This demonstration effectively showcased the entire lifecycle of the tool: from taking a simple seed specification, propagating it through complex call chains, inferring new specifications for high-level APIs, and finally leveraging these inferred specifications to pinpoint actual security flaws in a critical system like the Linux kernel. The fact that traditional methods failed to detect such bugs further underscores the unique value proposition of API spec.
Defensive Implications
▶ Watch: Comparison with prior API specification extraction methods (10:00)
The findings presented by Miaoqian Lin have profound implications for software defenders, developers, and security auditors. The core message is that relying solely on explicit documentation or common usage patterns for understanding API contracts is insufficient and leaves systems vulnerable to subtle, yet critical, misuse.
Here are key defensive implications:
- Beyond Documentation and Usage: Defenders must recognize that a significant portion of crucial API specifications are "hidden"—neither documented nor frequently observed in typical code. This means traditional static analysis tools that depend on explicit annotations or pattern matching against common usage will inherently miss a large class of bugs. Security audits should therefore incorporate methodologies capable of inferring implicit API contracts.
- Embrace Specification Propagation: The concept of specification propagation analysis offers a powerful paradigm shift. Defenders should consider how resources, states, and obligations flow through hierarchical API call chains. Tools or techniques that can trace these propagations, similar to API spec, can uncover latent specifications that are critical for correct resource management, error handling, and state transitions.
- Automated Specification Generation: Manually inferring thousands of specifications is impractical. Defenders should explore and adopt automated tools that can generate API specifications from a minimal set of seeds. The efficiency demonstrated by API spec (thousands of specs in hours) suggests that such tools can be integrated into CI/CD pipelines or regular auditing processes to continuously improve the understanding of a codebase's API contracts.
- Focus on Post-Handling Operations: The research highlighted API post-handling as a significant source of bugs (over 16% in C). This suggests a targeted area for defensive efforts. When reviewing code or designing new APIs, particular attention should be paid to ensuring that resources acquired are released, locks obtained are freed, and error conditions are properly checked and handled after an API call.
- Augment Existing Static Analysis: API spec is not a replacement for existing static analysis tools but a powerful augmentation. By feeding newly generated, accurate API specifications into existing bug detectors, the efficacy of these tools can be significantly enhanced, enabling them to catch bugs they previously couldn't.
- Open-Source Advantage: The fact that API spec is open-source and has passed artifact evaluation is a significant advantage. This allows security teams to inspect, adapt, and integrate the tool into their own security ecosystems. It provides a foundation for developing custom solutions tailored to specific project needs or for contributing to its improvement.
- Language and Domain Agnostic Potential: While demonstrated on C/Linux kernel, the underlying principle of hierarchical API design and specification propagation is broadly applicable. Defenders working with other languages or user-space applications should explore how these concepts can be applied to their respective environments, as the speaker suggested its potential for Python, OpenSSL, and FFmpeg. This could lead to uncovering similar patterns of misuse in diverse software stacks.
In essence, the work on API spec provides a clear directive: to build truly robust and secure software, defenders must look beyond the surface of explicit documentation and actively infer the complete, often hidden, landscape of API specifications.
Key Takeaways
- Hidden Specifications are Prevalent: A significant portion of critical API specifications, particularly those related to post-handling operations, are not documented or frequently observed, severely limiting traditional bug detection methods.
- Specification Propagation is Powerful: API specifications can propagate through hierarchical call chains. By leveraging this phenomenon, API spec can infer thousands of new, accurate specifications from a small number of known seeds.
- API spec is Highly Effective: The tool successfully generated 7332 specifications from just 6 seeds and detected 186 new, confirmed bugs in the Linux kernel, demonstrating its practical impact and efficiency.
- Beyond Artifact-Dependent Methods: Traditional methods relying on API documents or usage patterns failed to extract most of the specifications generated by API spec and consequently missed the corresponding bugs.
- Improved Bug Detection: The generated specifications provide crucial context for bug detectors, enabling the identification of reference count errors, memory leaks, and other improper resource handling issues that were previously undetectable.
- Open-Source and Applicable: API spec is an open-source tool with proven applicability across complex projects like the Linux kernel, OpenSSL, and FFmpeg, offering a valuable resource for developers and security researchers.
About the Speaker(s)
Miaoqian Lin is a PhD student from the Chinese Academy of Science. Their research focuses on improving software security through novel analysis techniques, with a particular emphasis on API specification inference and automated bug detection. This presentation at the NDSS Symposium highlights their work in developing the API spec framework, demonstrating a significant contribution to the field of static analysis and software vulnerability discovery.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Solid systems security research with a clean insight: spec propagation through call hierarchies lets you bootstrap thousands of API contracts from a handful of seeds. 186 confirmed Linux kernel bugs from 6 seed specs is a result that speaks for itself, and the 90%+ of generated specs invisible to frequency-based methods is a direct indictment of the state of the art.
Heather Calloway (CISO) — WEAK
Technically credible research with real results — 186 confirmed Linux kernel bugs is not nothing. But this talk never crosses the threshold into defender or operator relevance. It is academic work presented for an academic audience, and it stays there.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025