bin2ml: turning software binaries into machine learning ready training data
Josh Collyer (Head of AI Security Group · The Alan Institute)
44CON 2024 · Day 1 · Main
Overview
Josh Collyer, Head of AI Security Group at The Alan Institute, presented bin2ml, an open-source tool designed to transform software binaries into machine learning-ready training data. This talk addresses a critical bottleneck in applying machine learning (ML) to binary analysis: the laborious and often inconsistent process of data extraction and preparation from diverse binary formats and architectures. Collyer, whose PhD research focuses on ML and AI for reverse engineering, developed bin2ml to standardize and streamline this essential first step, enabling researchers and practitioners to focus on model development rather than data wrangling.

Key moments
- 0:00 Introduction and Talk Agenda
- 2:00 Motivation for bin2ml and community bridging
- 3:30 Core philosophy: Automate, not replace
- 4:10 Generative AI in binary analysis applications
- 6:10 Non-generative AI for vulnerability detection
bin2ml: turning software binaries into machine learning ready training data
Speakers: Josh Collyer (Head of AI Security Group, The Alan Institute)
Conference: 44CON
YouTube: https://www.youtube.com/watch?v=YG94pzbptCc
Overview
Josh Collyer, Head of AI Security Group at The Alan Institute, presented bin2ml, an open-source tool designed to transform software binaries into machine learning-ready training data. This talk addresses a critical bottleneck in applying machine learning (ML) to binary analysis: the laborious and often inconsistent process of data extraction and preparation from diverse binary formats and architectures. Collyer, whose PhD research focuses on ML and AI for reverse engineering, developed bin2ml to standardize and streamline this essential first step, enabling researchers and practitioners to focus on model development rather than data wrangling.
The motivation behind bin2ml stems from a recognized gap between the highly specialized worlds of machine learning and binary analysis. Many ML-based approaches to binary analysis are perceived as unrealistic or "snake oily" by reverse engineering experts due to incorrect assumptions or poorly prepared data. bin2ml aims to bridge this divide by providing a robust, open-source framework that extracts high-quality, standardized data, fostering more realistic and impactful ML research in cybersecurity. By automating the "boring stuff" of data generation, the tool empowers analysts to leverage ML for tasks like vulnerability detection, third-party library identification, and malware analysis, ultimately enhancing human capabilities rather than replacing them.
This initiative is particularly significant given the current landscape of AI research, where large language models (LLMs) are dominating, yet foundational data preparation for specialized tasks like binary analysis remains a challenge. bin2ml offers a practical solution for researchers who lack access to expensive commercial tools or struggle with the inconsistencies of academic data sets. It champions an open, collaborative approach to advance the state of the art in AI-driven binary security, inviting the community to contribute and refine its capabilities.
Background
▶ Watch: Introduction and Talk Agenda (0:00)
The application of machine learning to binary analysis has been a growing field, driven by the desire to automate complex and time-consuming tasks traditionally performed by highly skilled reverse engineers. Collyer highlights several key questions ML can answer in this domain: identifying a function's purpose, understanding unknown instructions, refactoring decompiled code (e.g., renaming variables), or even translating pseudo-code into higher-level languages like Python or JavaScript. Beyond generative AI tasks, ML is also crucial for non-generative tasks such as detecting known or unknown vulnerabilities, identifying variants of vulnerabilities, or finding specific functionalities like cryptography routines or command-and-control (C2) mechanisms within binaries. A significant challenge, particularly in firmware analysis, is the accurate identification of third-party libraries, whether open-source (like OpenSSL) or proprietary, which can inform supply chain security and vulnerability patching.
Despite the promise, the field has been plagued by significant data-related hurdles. Collyer, drawing from his experience reviewing extensive academic literature, points out that much research suffers from a lack of reproducibility and inconsistent data generation methodologies. Papers often rely on bespoke data extraction pipelines tied to specific commercial tools like IDA Pro, Ghidra, or Binary Ninja, making it nearly impossible for others to replicate results without significant investment in licenses or re-engineering. For instance, PalmTree (a Transformer-based model for x86 instructions) used Binary Ninja, while SAFE (an LSTM-based approach) used Radare2, and Jr used IDA Pro. This fragmentation leads to a "mess" for ML practitioners, who often spend 80% of their project time on data acquisition and cleaning.
A pivotal moment for Collyer was encountering a comprehensive paper from Cisco that compared ten different approaches, releasing code and data, but still relied heavily on IDA Pro. This reinforced the need for an open, unified data generation solution that could support cross-architecture and cross-platform analysis without commercial tool dependencies. The existing landscape often sees research focused solely on Linux binaries or specific architectures, leaving a void for broader, more flexible analysis. bin2ml was conceived to address this fundamental problem, providing a standardized, accessible, and scalable way to generate diverse binary representations for ML, thus democratizing access to high-quality training data for the entire research community.
Key Findings
▶ Watch: Motivation for bin2ml and community bridging (2:00)
The primary finding and contribution of Josh Collyer's work is the bin2ml tool itself, which effectively addresses the pervasive data generation bottleneck in machine learning for binary analysis. bin2ml stands out by providing a standardized, open-source, and highly flexible framework for extracting diverse representations from software binaries, eliminating the need for expensive commercial tools or fragmented, custom pipelines.
Key findings related to bin2ml's capabilities and design principles include:
- Democratization of Data:
bin2mlmakes high-quality binary analysis data accessible to a wider community. By leveraging Radare2, an open-source disassembler and debugger, it removes the financial barrier associated with tools like IDA Pro or Binary Ninja. This enables researchers, especially those in academia or smaller organizations without large budgets, to engage in cutting-edge ML research in binary security. - Standardized Data Formats: The tool outputs data in JSON format, specifically NetworkX JSON objects for graph representations, ensuring compatibility and ease of integration with popular ML frameworks like Hugging Face, PyTorch Geometric, or custom Python scripts. This standardization significantly reduces the post-processing burden, allowing researchers to load data directly into their models.
- Cross-Architecture and Cross-Platform Support:
bin2mlis designed to run on various operating systems (Windows, macOS, Linux) and process binaries from any architecture or platform that Radare2 supports. This flexibility is critical for comprehensive security research, moving beyond the prevalent Linux-centric focus of much academic work. - Efficiency and Scalability: The tool is engineered for both local development (processing 10 binaries on a laptop) and large-scale deployment (processing millions on high-performance computing clusters). Its multi-threaded processing capabilities ensure efficient data generation, with the primary bottleneck often shifting to hardware I/O rather than the analysis itself.
- Diverse Data Modalities:
bin2mlsupports the generation of a wide array of binary representations crucial for different ML tasks:
- Control Flow Graphs (CFGs): With various node features, including raw disassembly, Gemini paper features, SC Discover features, Radare2's Intermediate Language (R2-ISL), or pseudo-code.
- Call Graphs: Both global call graphs and n-hop localized call graphs (e.g., one-hop, capturing callers and callees of a function).
- Linear Disassembly and ISL: As single lines, random walks (useful for training tokenizers like in PalmTree), or function strings (as used in Jr).
- Decompiled P-code: Raw pseudo-code or P-code per basic block, facilitating rapid prototyping of P-code analysis without direct Ghidra interaction.
- Function and Binary Metadata: Extracting features like the number of local variables, instructions, basic blocks, and cyclomatic complexity for tree-based ML models (e.g., XGBoost).
- Built-in Data Hygiene:
bin2mlincorporates crucial data cleaning functionalities, such as deduplication for graphs. Collyer highlights that ignoring deduplication can artificially inflate model performance by as much as 15%, underscoring the tool's commitment to generating realistic and reliable training data. It also performs tokenization and masking of sensitive values (like memory addresses or function calls) to create a more generalized vocabulary for natural language processing tasks. - Reproducibility of Research: The tool's ability to generate data in formats used by seminal papers (e.g., Gemini features, PalmTree's random walks, Jr's function strings) directly facilitates the replication and benchmarking of existing ML-for-binary-analysis research.
In essence, bin2ml is a foundational piece of infrastructure that significantly lowers the barrier to entry for ML research in binary security, promoting open science, reproducibility, and ultimately, more effective and realistic applications of AI in reverse engineering.
Technical Deep Dive
▶ Watch: Core philosophy: Automate, not replace (3:30)
bin2ml is architected around a flexible Extract, Generate, Load (EGL) paradigm, a modification of the traditional ETL (Extract, Transform, Load) model commonly used in data engineering. This design emphasizes the separation of concerns: first, extracting raw data from the binary, then generating various structured representations, and finally, loading these into ML frameworks.
At its core, bin2ml leverages Radare2 (R2), an open-source reverse engineering framework. Collyer chose R2 for several key reasons:
- R2pipe: A powerful interface that allows external scripts to interact with R2 by sending command-line instructions and receiving structured JSON output. This ensures consistent and programmatic access to R2's analysis capabilities.
- JSON Output: R2's ability to output analysis results in JSON format aligns perfectly with
bin2ml's goal of standardized data. - Plugins: R2's plugin ecosystem, particularly R2Ghidra (which integrates Ghidra's Sleigh disassembler), provides access to advanced intermediate representations like P-code without requiring a full Ghidra installation.
- Minimal Dependencies:
bin2mlaims for a lightweight footprint, requiring primarily R2 itself, reducing installation complexity across diverse environments. - Community Receptiveness: The active development and responsiveness of the R2 community (including its primary maintainer, pancake) to issues and feature requests ensures ongoing support and improvement.
The bin2ml workflow can be broken down into three logical stages:
- Extract (E):
- A binary is fed into
bin2ml. bin2mlinvokes Radare2 (via R2pipe) to perform initial analysis, such as identifying functions, basic blocks, and control flow.- The raw analytical data (e.g., function metadata, basic block boundaries, instruction bytes) is extracted and saved as a canonical JSON file.
- This extraction is performed once for a given binary, saving significant time, as R2's initial analysis can be computationally intensive, especially for large binaries. The "extract once, use many" principle is crucial here.
- Generate (G):
- This is the core "transformation" stage where the raw extracted JSON data is converted into various ML-ready formats.
- The
generatecommand allows users to specify the desired data type and features. This stage is highly parallelizable, supporting multi-threading (e.g.,-n 500for 500 workers on an HPC node) to process large datasets efficiently. - Control Flow Graphs (CFGs):
- Disassembly: Basic blocks as nodes, containing raw x86 instructions. Edges represent jumps or fall-throughs.
- Gemini Features: Nodes are enriched with features specifically designed for the Gemini paper, a seminal work in graph neural networks for function similarity.
- SC Discover Features: Nodes incorporate features from the SC Discover approach.
- Intermediate Language (ISL): Nodes contain Radare2's Intermediate Language (R2-ISL) instructions, providing a more abstract, architecture-agnostic representation.
- Pseudo-code: Nodes contain decompiled pseudo-code, leveraging R2Ghidra for decompilation.
- All CFG outputs are structured as NetworkX JSON objects, directly loadable into Python's NetworkX library for graph manipulation and analysis.
- Call Graphs (CGs):
- Global Call Graphs: Represents the entire binary's function call structure.
- N-hop Call Graphs: Focuses on a localized view, e.g., one-hop includes a function's direct callers and callees. This is particularly useful for interprocedural analysis and has been used in papers like Kyn (Know Your Neighborhood) for malware variant detection.
- Disassembly and ISL (Linear/Random Walk/Function Strings):
- Linear: A simple sequential list of instructions within a function, useful for training tokenizers for NLP-based models.
- Random Walk: Simulates a random traversal of the CFG, outputting sequences of instructions. This mimics data used by models like PalmTree for learning instruction embeddings.
- Function Strings: Concatenates all instructions of a function into a single string, as used by Jr for jump target prediction.
- Decompiled P-code:
- P-code Basic Block: CFGs where nodes contain the P-code for each basic block.
- P-code Singles: Linear sequences of P-code instructions.
- This allows rapid prototyping and analysis of P-code, such as for identifying memory input/output operations as in the Hermes-S paper, without requiring direct Ghidra API interaction.
- Metadata: Extracts numerical and categorical features at the function or binary level (e.g., number of basic blocks, cyclomatic complexity, number of instructions, local variables) suitable for traditional ML algorithms like XGBoost.
- Data Cleaning and Masking: During generation,
bin2mlperforms crucial data hygiene: - Graph Deduplication: Identifies and removes identical graphs, preventing inflated performance metrics in ML models. Collyer notes that ignoring this can lead to a 15% performance increase in reported results.
- Token Masking: Replaces specific hex values (e.g., memory addresses, function call targets) with generic tokens (e.g.,
_M_for memory,_FUNCTION_for calls). This reduces the vocabulary size for NLP models and improves generalization across different binaries.
- Load (L):
- The generated JSON data can be directly loaded into various ML frameworks. For graph data, NetworkX in Python is a natural fit, and from there, conversion to PyTorch Geometric or other graph ML libraries is straightforward. For textual data, standard NLP libraries or custom tokenizers can be used.
The overall technical design of bin2ml prioritizes openness, standardization, and flexibility, making it a robust platform for generating diverse and high-quality training data for the evolving field of ML-driven binary analysis.
Demo / Proof of Concept
▶ Watch: Generative AI in binary analysis applications (4:10)
While Josh Collyer's talk did not feature a live, interactive demonstration of bin2ml in action, the "What can we generate" section of his presentation served as a comprehensive demonstration of the tool's capabilities through command-line examples and illustrative outputs. This effectively showcased the diverse range of data types bin2ml can produce and how users would interact with it.
The demonstration highlighted the tool's two-stage process:
- Extraction (
extractcommand):
The initial step involves parsing a binary and extracting raw information into a standardized JSON format.
This command performs the heavy lifting of initial binary analysis using Radare2, generating a single, comprehensive JSON file that contains all the necessary raw data for subsequent transformations. The key takeaway here is that this step is done once, saving significant computational resources if multiple representations are needed from the same binary.
- Generation (
generatecommand):
Once the raw JSON is extracted, the generate command is used to produce specific ML-ready data types from that extracted file. This is where the flexibility of bin2ml truly shines.
Demonstrated Capabilities:
- Control Flow Graphs (CFGs):
This command would output a NetworkX JSON object representing the CFG, with each node containing a basic block's disassembly instructions. Collyer visually presented an example showing x86 instructions within basic blocks connected by edges. The demonstration further extended to showing how simply changing the --feature-type flag could produce CFGs with Gemini features, SC Discover features, R2-ISL instructions, or even decompiled pseudo-code in the nodes, all from the same initial extracted JSON. This underscored the "extract once, use many" principle.
- Call Graphs (CGs):
The tool can generate a global call graph depicting all function calls across the binary or n-hop localized call graphs, such as a one-hop graph showing a function's direct callers and callees. Visual examples illustrated how these localized graphs could focus on specific areas of interest (e.g., the main function and its immediate neighbors).
- Disassembly and R2-ISL:
These commands showcase the generation of linear lists of instructions, random walks through the CFG (emulating data used by PalmTree for tokenizer training), and function strings (as used by Jr). Similar options exist for R2-ISL.
- Decompiled P-code:
The demonstration explained how bin2ml could output P-code per basic block (within a CFG structure) or as single linear sequences, leveraging Radare2's Ghidra plugin. This capability allows for rapid prototyping of P-code-based analyses without direct Ghidra integration.
The presentation also implicitly demonstrated bin2ml's advanced features, such as the -n flag for specifying the number of parallel workers during generation, indicating its scalability. Furthermore, the discussion of deduplication and token masking during the generation phase highlighted the tool's built-in data hygiene, which is critical for producing high-quality ML training data. The use of specific command-line flags and the clear explanation of their effects provided a robust conceptual "proof of concept" for bin2ml's versatility and utility.
Defensive Implications
▶ Watch: Non-generative AI for vulnerability detection (6:10)
bin2ml offers significant implications for defensive cybersecurity, primarily by empowering security teams and researchers to leverage machine learning for more efficient and effective threat detection, vulnerability management, and intelligence gathering. The core benefit lies in its ability to generate high-quality, standardized data, which is often the missing link for practical ML application in real-world defensive scenarios.
- Automated Vulnerability Detection and Prioritization:
- By generating rich representations like Control Flow Graphs (CFGs) with various features (disassembly, ISL, pseudo-code) or function metadata, defenders can train ML models to identify patterns indicative of known vulnerabilities (CVEs) or even discover unknown, zero-day vulnerabilities.
- The ability to detect variants of vulnerabilities is particularly powerful. Instead of relying solely on exact signature matches, ML models trained on
bin2mldata can learn the underlying structural or semantic similarities of vulnerability classes, making them resilient to minor code changes or obfuscation. This aids in proactive patching and risk assessment across large codebases or firmware repositories.
- Enhanced Third-Party Component Analysis (Software Supply Chain Security):
- One of the "huge problems"
bin2mlaims to tackle is identifying third-party libraries within binaries, whether open-source (e.g., OpenSSL) or proprietary. By creating learned representations of known libraries, defenders can quickly scan firmware or compiled applications to identify included components. - This is critical for software supply chain security, allowing organizations to rapidly assess their exposure to vulnerabilities in third-party dependencies without requiring source code. If a new vulnerability in a library is disclosed,
bin2ml-generated data can help pinpoint affected binaries at scale.
- Malware Analysis and Variant Detection:
- The tool's capabilities for generating call graphs (especially n-hop localized graphs) and various textual representations of disassembly or ISL are highly relevant for malware analysis. ML models can be trained to recognize specific malware functionalities (e.g., C2 communication, persistence mechanisms, cryptography use) or to group new samples into known malware families.
- The
bin2mlapproach, particularly with techniques like those used in the Kyn paper (which leverages n-hop call graphs), can be used for finding malware variants. This helps security analysts understand the evolution of threats and develop more robust detection rules that are less prone to evasion by minor modifications.
- Improved Reverse Engineering Productivity:
- While not directly defensive,
bin2mlautomates the "boring stuff" of data extraction, freeing up highly skilled reverse engineers to focus on higher-value tasks. For example, ML models trained onbin2mldata could automatically identify standard library functions, allowing human analysts to concentrate on custom, potentially malicious, code. This acts as a "power suit," augmenting human capabilities.
- Benchmarking and Reproducibility for Security Research:
- By facilitating the replication of academic research (e.g., Gemini, PalmTree),
bin2mlenables defensive researchers to rigorously evaluate the effectiveness of different ML approaches against real-world security problems. This ensures that deployed ML solutions are based on sound, validated science rather than "snake oil."
- Rapid Prototyping of Static Analysis Tools:
- The ability to extract P-code in a structured format allows security researchers to quickly prototype and experiment with static analysis techniques that operate on intermediate representations, without the overhead of deeply integrating with complex tools like Ghidra. This can accelerate the development of new defensive capabilities.
In summary, bin2ml provides the foundational data infrastructure to operationalize machine learning for a wide range of defensive cybersecurity tasks. It moves the field towards more automated, scalable, and intelligent analysis of binaries, ultimately strengthening an organization's posture against evolving threats.
Key Takeaways
bin2mlBridges ML and Binary Analysis: The tool addresses the significant gap between machine learning practitioners and binary analysis experts by providing a standardized, open-source method for generating ML-ready data from binaries, fostering more realistic and impactful research.- Open Source and Accessible: Leveraging Radare2,
bin2mlremoves financial barriers, making high-quality data generation accessible to a broader community without requiring expensive commercial reverse engineering tools like IDA Pro or Binary Ninja. - Standardized and Versatile Data Output: It produces data in JSON format (including NetworkX JSON objects for graphs), ensuring easy integration with popular ML frameworks and supporting a wide array of representations including Control Flow Graphs (CFGs), Call Graphs, linear disassembly, R2-ISL, and decompiled P-code.
- Efficient and Scalable Data Generation:
bin2mlis designed for both local development and large-scale processing on HPC clusters, featuring multi-threaded generation and an "extract once, use many" philosophy to maximize efficiency. - Crucial Data Hygiene Included: The tool incorporates essential data cleaning features like graph deduplication (which can impact model performance by 15% if ignored) and token masking for memory addresses and function calls, ensuring higher quality and more generalizable training data.
- Empowers Defensive Cybersecurity:
bin2mlenables defenders to apply ML for automated vulnerability detection (including variants), robust third-party library identification for supply chain security, and scalable malware analysis and variant detection, ultimately enhancing human analyst capabilities.
About the Speaker(s)
Josh Collyer is the Head of the AI Security Group at The Alan Institute, the UK's National Research Institute for data science and AI. With approximately eight years of experience at the intersection of cyber, machine learning, and artificial intelligence, Josh brings a wealth of practical and research-oriented expertise. His background includes working in the UK's Ministry of Defense, where he focused on cloud hosting security (Azure, AWS) and machine learning applications within secure operation centers. He has also contributed to defense and national security research.
Currently, his work at The Alan Institute centers on ensuring that academic research and threat models in AI security are realistic and address tangible problems, rather than theoretical or non-feasible ones. Complementing his professional role, Josh is pursuing a part-time PhD focused on machine learning and AI for reverse engineering, with an expanding interest in virtual reality. His passion lies in bridging the gap between highly specialized ML and binary analysis communities, advocating for collaboration to refine assumptions and create more effective, real-world solutions.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Collyer has identified a real pain point — the fragmented, tool-dependent data pipelines choking ML-for-binary-analysis research — and built something practical to address it. bin2ml is a legitimate open-source infrastructure contribution, well-scoped and technically coherent, but it's tooling, not research, and the talk doesn't pretend otherwise. The ceiling on this is inherently limited.
Heather Calloway (CISO) — WEAK
Legitimate research infrastructure solving a real reproducibility problem in ML-for-binary-analysis, but this is foundational tooling for a narrow research audience. There is no governance angle, no institutional accountability dimension, and no path from this talk to a security program decision.