Mens Sana In Corpore Sano: Sound Firmware Corpora for Vulnerability Research
René Helmke (Franova)
Network and Distributed System Security (NDSS) Symposium 2025 · Day 2 · Hard- & Firmware Security · Hard- & Firmware Security
Overview
René Helmke from Franova presented a critical analysis of the current state of firmware corpus creation in vulnerability research, advocating for a more "scientifically sound" approach. The talk, titled "Mens Sana In Corpore Sano: Sound Firmware Corpora for Vulnerability Research," addresses the pervasive challenges researchers face when building, sharing, and documenting evaluation datasets, particularly for identifying vulnerabilities in embedded systems. Helmke highlights that the lack of rigorous methodology in corpus creation undermines the transparency, comprehensibility, and verifiability of research results, hindering replication and progress in the field.
Key moments
- 0:40 Paper's Four Major Contributions to Corpus Creation
- 2:26 Practical Challenge: Firmware Unpacking as a Showstopper
- 4:10 Proposed Three-Layer Guidelines for Robust Corpora
- 5:50 Rich Shareable Metadata for Corpus Replicability
- 7:55 Research Analysis: Missing Metadata Threatens Soundness
- 8:10 Only 30% of Papers Document Firmware Unpacking
Mens Sana In Corpore Sano: Sound Firmware Corpora for Vulnerability Research
Speakers: René Helmke (Franova)
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=ZD6855E8J64
Overview
René Helmke from Franova presented a critical analysis of the current state of firmware corpus creation in vulnerability research, advocating for a more "scientifically sound" approach. The talk, titled "Mens Sana In Corpore Sano: Sound Firmware Corpora for Vulnerability Research," addresses the pervasive challenges researchers face when building, sharing, and documenting evaluation datasets, particularly for identifying vulnerabilities in embedded systems. Helmke highlights that the lack of rigorous methodology in corpus creation undermines the transparency, comprehensibility, and verifiability of research results, hindering replication and progress in the field.
The presentation outlines four major contributions aimed at improving the scientific rigor of firmware vulnerability research. First, it identifies and analyzes the practical challenges inherent in creating high-quality, robust firmware corpora. Second, it proposes a structured set of guidelines, organized into a three-layer model, to assist researchers in developing more scientifically sound corpora. Third, these guidelines are then applied in a comprehensive literature analysis of 45 papers, revealing current trends and deficiencies in corpus creation practices. Finally, the talk introduces and releases the Linux Firmware Corpus (LFWC), a meticulously curated reference corpus designed to exemplify the proposed guidelines and serve as a readily available, replicable resource for the community.
This work is particularly pertinent given the increasing complexity and prevalence of firmware in critical infrastructure, IoT devices, and enterprise systems. Robust vulnerability research relies on reliable and representative datasets to validate findings and ensure that discovered weaknesses are truly indicative of real-world security postures. Helmke's research aims to move the community away from what he describes as a "wild west situation" towards standardized, replicable practices that foster greater trust and collaboration in firmware security research.
Background
▶ Watch: Paper's Four Major Contributions to Corpus Creation (0:40)
The foundation of robust vulnerability research, especially in the domain of firmware, critically depends on the quality and characteristics of the datasets used for evaluation – often referred to as firmware corpora. However, as René Helmke emphasizes, creating a "scientifically sound" firmware corpus is far from trivial. A scientifically sound corpus, in this context, is defined by its transparency, comprehensibility, and verifiability, allowing other researchers to replicate results, and its representativeness, ensuring the findings are broadly applicable. The current landscape often falls short of these ideals, leading to difficulties in replicating research and assessing the true impact of novel vulnerability discovery techniques.
One of the most significant and frequently encountered hurdles, which Helmke labels a "major showstopper," is firmware unpacking. Firmware images are typically distributed in proprietary, often partially unknown or encrypted binary formats. Before any analysis or evaluation can commence, researchers must invest substantial manual effort into reverse engineering these formats to access their underlying components, such as the operating system kernel, bootloaders, and application binaries. This process is time-consuming and requires specialized expertise. Even once unpacked, the researcher must then evaluate the data relevance – ensuring the firmware contains the specific architecture (e.g., ARM-based code) or components pertinent to their research prototype. For instance, a tool designed for ARM firmware would find little utility in a corpus dominated by MIPS binaries, even if successfully unpacked.
Beyond the technical complexities of unpacking, a critical impediment to collaboration and reproducibility is the issue of copyrighted material. Much of the proprietary firmware contains intellectual property, making direct sharing of the unpacked binaries with research peers legally problematic. This often forces researchers to resort to documenting their corpus creation steps meticulously, hoping others can follow these instructions to independently reconstruct the dataset. However, this documentation-centric approach is "rather error-prone and different across papers," leading to inconsistencies and making true replication exceedingly difficult. This "wild west situation," as Helmke describes it, highlights an urgent need for standardized guidelines and shareable, legally compliant metadata to facilitate more robust and verifiable firmware vulnerability research. The lack of a common framework for corpus creation directly impacts the reliability and generalizability of published research, creating a bottleneck for progress in securing embedded systems.
Key Findings
▶ Watch: Proposed Three-Layer Guidelines for Robust Corpora (4:10)
The research presented by René Helmke identifies several critical findings regarding the state and future of firmware corpus creation for vulnerability research. The work began by cataloging eight practical challenges in corpus creation, with firmware unpacking standing out as a significant impediment due to its manual effort, proprietary formats, and legal sharing restrictions. This underscored the need for a structured approach to address these issues.
To counter these challenges, the talk introduced a three-layer model of guidelines designed to foster more robust and scientifically sound corpora. The first layer establishes abstract goals: replicability, ensuring others can reproduce the corpus; representativeness, making sure the corpus reflects real-world diversity; and method orientation, acknowledging that no single corpus fits all research needs, and composition must align with the evaluation method. The second layer translates these goals into six practical data requirements, such as ensuring ground truth and clean data. The third layer details 16 specific measures that contribute to fulfilling these requirements, including incorporating a good mix of device properties (manufacturer, model) and annotating samples with rich, legally shareable metadata. This metadata, such as download links and file identity hashes, allows peers to obtain binary data from legal sources and reconstruct the corpus.
A crucial part of Helmke's work involved a comprehensive literature analysis of 44 papers published between 2013 and 2023 at prominent security conferences. This analysis evaluated how existing research adhered to the proposed guidelines, specifically focusing on the documentation and provision of metadata. The findings were stark: "missing metadata and documentation threatens the soundness of our own work." While approximately 70% of papers documented their acquisition steps, a fundamental aspect, the provision of specific and complete metadata was alarmingly low. Crucially, only 30% of papers fully documented the firmware unpacking process. This deficiency means that in the vast majority of cases, replicating a corpus requires researchers to independently tackle the complex and time-consuming unpacking challenge, effectively making replication "basically gone." Consequently, the result verifiability and representativeness of many published studies become difficult, if not impossible, to assess.
To demonstrate the feasibility of his proposed guidelines and offer a tangible solution, Helmke presented and released the Linux Firmware Corpus (LFWC). This reference corpus comprises 10,900 fully unpacked, verified Linux samples, meticulously curated from a diverse range of devices, classes, manufacturers, and kernel versions. A key aspect of LFWC is its extensive, legally shareable metadata, which includes download links (with fallbacks to sites like archive.org), file identity hashes, and detailed information about the sample's contents, such as the Linux kernel version and instruction set architecture. A subsequent replicability study confirmed the success of this approach, with researchers able to obtain roughly 99.7% of all samples using only the provided metadata and download script, showcasing a significant leap forward in corpus reproducibility.
Technical Deep Dive
▶ Watch: Rich Shareable Metadata for Corpus Replicability (5:50)
The technical core of Helmke's presentation revolves around the proposed three-layer guideline model for creating scientifically sound firmware corpora and the practical implementation of these guidelines in the Linux Firmware Corpus (LFWC). This systematic approach aims to overcome the inherent complexities of firmware analysis and the challenges of data sharing and reproducibility in research.
The three-layer guideline model begins with abstract goals:
- Replicability: The ability for independent researchers to reconstruct the exact same corpus, leading to verifiable results.
- Representativeness: Ensuring the corpus adequately covers the diversity of real-world firmware, preventing biased or narrowly applicable findings.
- Method Orientation: Acknowledging that corpus composition must be tailored to the specific evaluation method or vulnerability research goal, rather than aiming for a "one-size-fits-all" dataset.
These abstract goals are then translated into six data requirements, which are intuitive yet often overlooked in practice. While not exhaustively detailed in the talk, these typically include aspects like ground truth (known properties of the firmware), clean data (free from errors or irrelevant noise), and sufficient diversity to support generalizable conclusions.
The most granular layer consists of 16 concrete measures that researchers can implement to fulfill these data requirements and achieve the overarching goals. A key example discussed is the annotation of samples with rich metadata, such as the device manufacturer and model. This information contributes directly to the heterogeneity and diversity of the dataset, enhancing its representativeness. Crucially, this metadata, unlike the copyrighted firmware binaries themselves, is legal to share with research peers. By providing this detailed metadata, other researchers can then use it to independently acquire the binary data from legal sources, thus contributing to replicability. Other measures include documenting acquisition steps, unpacking procedures, and providing file hashes for verification.
The Linux Firmware Corpus (LFWC) serves as a practical demonstration of these guidelines. It is a meticulously curated collection of 10,900 fully unpacked and verified Linux samples. The "verified" aspect means that the presence of a Linux kernel within the firmware image has been confirmed. The corpus is highly diverse, spanning numerous device classes, manufacturers, and a broad range of Linux kernel versions, providing a robust dataset for various vulnerability research scenarios.
The creation of LFWC heavily relied on the Firmware Analysis and Comparison Tool (FACT), which Franova developed and open-sourced over a decade ago. FACT is instrumental in automating several critical steps:
- Automated Unpacking: FACT utilizes a variety of unpacking heuristics and signatures to automatically extract the contents of complex firmware images. This capability directly addresses the "showstopper" challenge of manual unpacking that plagues much of existing research.
- Analysis and Annotation: Beyond unpacking, FACT analyzes the firmware to identify key components, such as the Linux kernel and the instruction set architecture (ISA). This information is then extracted and used to enrich the metadata associated with each sample.
- Deduplication: FACT also performs automated deduplication to ensure the corpus contains unique samples, preventing redundant data from skewing evaluation results.
The documentation of the LFWC creation process alone spans five pages in the associated paper, emphasizing the commitment to transparency and replicability. The shared metadata for LFWC includes not just the Linux kernel version and architecture, but also download links (with fallbacks to archival services like archive.org and allfall.org) and file identity hashes. These hashes are critical for verifying the integrity of downloaded samples and for searching for missing binaries on public platforms like VirusTotal. By leveraging FACT and providing comprehensive metadata, LFWC exemplifies how a "scientifically sound" firmware corpus can be built and maintained, setting a new standard for the community.
Demo / Proof of Concept
▶ Watch: Research Analysis: Missing Metadata Threatens Soundness (7:55)
While René Helmke's talk did not feature a live, interactive demonstration in the traditional sense, he effectively presented a proof of concept for the replicability of the LFWC. This "demo" outlines a clear, step-by-step process for any researcher to reconstruct the LFWC using the provided artifacts, thereby validating the efficacy of the proposed guidelines.
The replication process for the LFWC is designed to be highly automated and transparent:
- Metadata Provision: Researchers are initially provided with the comprehensive metadata for the entire LFWC. This metadata is legally shareable and contains all the necessary information about each firmware sample, without including the copyrighted binary data itself.
- Corpus Filtering: Based on their specific research needs, users can filter this metadata. For example, a researcher interested in analyzing vulnerabilities in ARM-based routers running specific Linux kernel versions could apply filters to narrow down the dataset to only relevant samples. This "method orientation" ensures researchers can tailor the corpus to their evaluation criteria.
- Automated Download: The filtered metadata is then fed into a download script, which is also provided as part of the LFWC artifacts. This script attempts to download all the corresponding firmware binaries from the specified legal sources (e.g., manufacturer websites, public archives). A crucial feature of this script is its ability to report which samples, if any, could not be obtained due to dead links or other issues.
- Missing Sample Acquisition: For any missing samples, researchers can leverage the file identity hashes included in the metadata. These hashes can be used to search for the specific firmware images on public virus scanning platforms like VirusTotal, which often archive a vast collection of binaries. This mechanism provides a robust fallback for obtaining hard-to-find samples.
- Automated Unpacking, Analysis, and Deduplication: Once all available binaries are collected, they are passed to the Firmware Analysis and Comparison Tool (FACT). As previously discussed, FACT automatically unpacks the firmware, analyzes its contents (e.g., identifying the Linux kernel and architecture), and deduplicates the samples. This fully automated pipeline transforms raw, often proprietary, firmware images into a ready-to-use, unpacked, and analyzed corpus.
The success of this methodology was empirically validated through a replicability study conducted one year after the LFWC's initial creation. This study revealed an impressive success rate: approximately 99.7% of all samples within the corpus could be obtained and processed using only the provided metadata and download script, excluding the analysis part. Helmke acknowledged that this number might "drop as links go dead inside of the internet," but he committed to further updates at least once a year as part of his PhD thesis, ensuring the corpus remains a valuable and current resource. This commitment to ongoing maintenance is vital for the long-term viability and scientific soundness of such a large-scale reference corpus.
Defensive Implications
▶ Watch: Only 30% of Papers Document Firmware Unpacking (8:10)
While René Helmke's talk primarily focuses on improving the methodology for firmware vulnerability research rather than direct defensive strategies, the implications for defenders are profound and far-reaching. The core message—that scientifically sound corpora lead to more reliable research—directly impacts the quality and trustworthiness of vulnerability intelligence available to security practitioners.
Defenders, whether in product security, incident response, or threat intelligence, rely heavily on accurate and reproducible vulnerability findings to inform their strategies. If the research identifying new attack vectors, vulnerability classes, or exploit techniques is based on poorly documented, non-replicable, or unrepresentative firmware corpora, the resulting conclusions may be flawed. This could lead to:
- Misallocation of Resources: Defenders might invest time and effort in mitigating vulnerabilities that are not truly widespread or exploitable in real-world scenarios, or conversely, overlook critical weaknesses that were missed due to a biased dataset.
- Ineffective Patching Strategies: If vulnerability research used to develop patches or security updates is based on an incomplete understanding of firmware diversity, patches might not be fully effective across all affected devices or may introduce new issues.
- Inaccurate Threat Intelligence: Threat intelligence feeds, which are crucial for proactive defense, derive significant value from cutting-edge research. If the underlying research is not robust, the intelligence provided could be misleading, leading to misinformed risk assessments.
- Delayed Response to Emerging Threats: The "wild west" of corpus creation slows down the pace of research by forcing replication efforts to start from scratch. This delay means defenders receive critical vulnerability information later, providing attackers with a longer window of opportunity.
The work presented by Helmke, particularly the development of the LFWC and the emphasis on the FACT tool, provides a blueprint for generating higher-quality vulnerability data. Defenders should recognize that research utilizing such rigorous approaches is more likely to yield reliable insights. This implies:
- Prioritizing Research with Strong Methodologies: When evaluating new security tools, techniques, or vulnerability reports, defenders should inquire about the underlying datasets and methodologies. Research based on transparent, replicable corpora like LFWC should be given higher credibility.
- Advocating for Openness and Collaboration: Defenders benefit when researchers can easily share and build upon each other's work. Encouraging the adoption of guidelines for corpus creation and supporting initiatives like LFWC contributes to a healthier security ecosystem.
- Leveraging Automated Analysis Tools: The use of tools like FACT for automated unpacking and analysis in research environments can significantly improve the efficiency and thoroughness of firmware security assessments, which can then feed into better defensive postures.
Ultimately, the defensive implications underscore that security is a continuous feedback loop. Strong research methodologies lead to better vulnerability discoveries, which in turn enable more effective defensive measures, ultimately enhancing the security posture of embedded systems and the critical infrastructure they support.
Key Takeaways
- Firmware corpus creation is inherently challenging: The process of building reliable, representative datasets for vulnerability research is complex, primarily due to proprietary binary formats, manual unpacking efforts, and legal restrictions on sharing copyrighted material.
- Lack of documentation and metadata hinders progress: A significant finding from the literature review is that insufficient documentation and metadata in existing research severely impede the replicability and verifiability of vulnerability research results, with only 30% of papers fully documenting unpacking.
- Structured guidelines are crucial for scientific soundness: The proposed three-layer model, encompassing abstract goals (replicability, representativeness, method orientation), data requirements, and specific measures, provides a robust framework for creating more transparent and verifiable corpora.
- The Linux Firmware Corpus (LFWC) sets a new standard: LFWC, comprising 10,900 fully unpacked and verified Linux samples with extensive, legally shareable metadata, demonstrates the feasibility of creating highly replicable and scientifically sound reference corpora.
- Automation tools like FACT are indispensable: The Firmware Analysis and Comparison Tool (FACT) plays a vital role in automating the complex processes of firmware unpacking, analysis, and deduplication, significantly reducing manual effort and enhancing the consistency of corpus creation.
- Community effort and ongoing maintenance are key: The high initial replicability of LFWC (99.7%) and the commitment to annual updates highlight the importance of community contributions and sustained effort to ensure the long-term utility and scientific value of such shared resources.
About the Speaker(s)
René Helmke is a researcher from Franova, a company based in Germany. His work, particularly the development and maintenance of the Linux Firmware Corpus (LFWC), is an integral part of his PhD thesis. He is actively involved in improving the scientific rigor and reproducibility of firmware vulnerability research, advocating for standardized methodologies and open-source tools within the security community. Helmke is also one of the developers behind the open-source Firmware Analysis and Comparison Tool (FACT).
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate methodological contribution to firmware research — someone finally sat down and audited how the field builds its corpora, found the mess everyone suspected, and released a reference dataset to fix it. Useful, honest work, but it's infrastructure plumbing rather than a novel attack or defense technique, and the ceiling on excitement is structurally limited.
Heather Calloway (CISO) — PASS
Technically credible work on research methodology for firmware corpus creation — a real problem inside academic vulnerability research. But this is squarely outside my lane: no governance angle, no operator decision path, no institutional accountability dimension. Scope call, not a quality call.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025