Simple Machine Learning Techniques for Binary Diffing (in Diaphora)
Joxean Koret
44CON 2024 · Day 2 · Main
Overview
Joxean Koret's presentation at 44CON delves into the practical application of machine learning (ML) techniques to binary diffing, specifically within his open-source tool, Diaphora. Binary diffing, a cornerstone of reverse engineering, involves comparing two binary files to identify identical or similar functions, often across different versions, architectures, or compilation settings. While academia frequently explores ML for binary similarity analysis, Koret highlights a significant gap: the lack of adoption of these advanced techniques in mainstream industry tools like Diaphora, Bindiff, or Ghidra diff.

Key moments
- 0:00 Introduction to ML for binary diffing in Diaphora
- 2:00 Binary diffing explained: mapping functions and classifying matches
- 3:10 Why industry avoids ML; academic papers often unhelpful
- 4:45 First idea: training local model from Diaphora's matches
- 6:40 Flaw: model learns what's already coded, adds false positives
- 7:05 Second idea: huge model, but building data sets difficult
- 8:20 Scarcity of good, usable binary diffing datasets
Simple Machine Learning Techniques for Binary Diffing (in Diaphora)
Speakers: Joxean Koret
Conference: 44CON
YouTube: https://www.youtube.com/watch?v=wHRd39u02io
Overview
Joxean Koret's presentation at 44CON delves into the practical application of machine learning (ML) techniques to binary diffing, specifically within his open-source tool, Diaphora. Binary diffing, a cornerstone of reverse engineering, involves comparing two binary files to identify identical or similar functions, often across different versions, architectures, or compilation settings. While academia frequently explores ML for binary similarity analysis, Koret highlights a significant gap: the lack of adoption of these advanced techniques in mainstream industry tools like Diaphora, Bindiff, or Ghidra diff.
This talk chronicles Koret's journey as a self-professed non-ML expert, exploring various approaches to integrate ML into Diaphora. His motivation stemmed from a desire to enhance Diaphora's existing heuristic-based matching capabilities, aiming to find more reliable function matches and improve similarity classification. The research navigates the challenges of data set creation, model training, and the practical limitations faced by an independent developer, ultimately advocating for a pragmatic, specialized approach to leveraging ML in real-world reverse engineering scenarios.
Background
▶ Watch: Introduction to ML for binary diffing in Diaphora (0:00)
Binary diffing is a crucial process in reverse engineering, typically broken down into two main stages: mapping and classification. Mapping involves identifying pairs of functions across two binaries that are fundamentally the same, even if compiled differently or slightly modified. Once a potential match is found, classification determines the reliability of that match and, optionally, calculates a similarity ratio indicating how closely the two functions resemble each other. For instance, a function A in binary one and function B in binary two might perform the same operation, and binary diffing aims to link them and quantify their resemblance.
While these core problems appear ripe for machine learning solutions, particularly classification, Koret noted a stark contrast between academic research and industry practice. Established industry tools like Diaphora, Bindiff, and Ghidra diff historically rely on sophisticated heuristics and traditional algorithms rather than ML. In the academic sphere, however, there's a proliferation of papers attempting to apply ML, generative AI, and neural networks to binary diffing, often resulting in tools like DeepBindiff, AlphaBindiff, or Bindiff Neural Network. Koret's experience with these academic offerings was largely disappointing, finding many papers either overly theoretical ("fairy tales"), based on fundamentally flawed premises (e.g., raw byte comparison), or simply not practical or useful for real-world reverse engineering tasks.
This disparity led Koret to explore three primary ideas for integrating ML into Diaphora: training a local model using Diaphora's own reliable matches, training a gigantic, generalized model on a massive, diverse dataset, or training highly specialized models tailored for specific use cases. Each approach presented its own set of theoretical promises and practical hurdles, which Koret systematically investigated throughout his research.
Key Findings
▶ Watch: Why industry avoids ML; academic papers often unhelpful (3:10)
Koret's research yielded several critical insights into the practical application of machine learning for binary diffing:
- Failure of Local Model Training: The initial idea of training a local model using Diaphora's own highly reliable, heuristic-derived matches proved to be fundamentally flawed. The model essentially attempted to "learn" the logic already embedded in Diaphora's codebase. This approach not only added redundancy but also introduced potential false positives due to the generalization inherent in ML, ultimately offering no real improvement over the existing heuristics. This was a significant early lesson: don't train a model to mimic what your tool already does.
- Impracticality of Gigantic, Generalized Models: The aspiration to build a single, all-encompassing model trained on a vast, diverse dataset (covering multiple compilers, architectures, operating systems, etc.) was largely unfeasible for an independent researcher. Koret highlighted the prohibitive resource requirements for building such datasets, citing examples like the Johns Hopkins APL Allstar data set (1.1 terabytes) which, despite its size, still lacked coverage for many crucial platforms (Windows, Android, iOS, Mac, RISC-V, Microsoft Visual C). Even the more manageable Cisco Talos Binary Function Similarity Data Set 2 (4,047 binaries) required an estimated two months of processing time just to generate the comparison data on his optimized home setup. The sheer scale of data storage, analysis, and training made this approach impractical and non-distributable (e.g., a 2GB model file).
- Success of Highly Specialized Models: The most impactful finding was the efficacy of training highly specialized models. This approach involves building a dataset from a limited number of related binaries, such as different versions of the same software, particularly when one version includes symbols and others are stripped. By training a model on known good matches from these related binaries, Koret demonstrated significant improvements in function matching for subsequent, symbol-stripped versions. This was successfully validated with the Microsoft antivirus engine (MP engine) and other undisclosed projects. This strategy leverages the inherent consistency within a specific software's development lifecycle (same developers, compilers, code base evolution) to create a highly accurate and relevant classifier.
- The "Amalgamation" Data Set and Model: Building on the success of specialized models, Koret developed a more broadly applicable, yet still compact, model by amalgamating data from Diaphora's internal testing suite, the manageable Cisco Talos Data Set 2, and various user-reported problem binaries. This amalgamation model (50MB compressed) demonstrated substantial improvements in finding partial matches across diverse, challenging scenarios (e.g., BusyBox across architectures, iOS vs. macOS kernels) without introducing a significant number of false positives. This finding suggests that a carefully curated, moderately sized dataset, focused on real-world reverse engineering challenges, can yield a robust and distributable ML model.
- Simplicity in Algorithm Choice: Through extensive testing of various ML algorithms, Koret found that a Decision Tree Classifier offered the best balance of speed, interpretability, and sufficient accuracy for Diaphora's needs. While a Random Forest Classifier showed slightly better performance, its significantly longer training and prediction times made it unsuitable. This highlights that complex algorithms aren't always necessary; a simpler, faster model can be highly effective when combined with well-chosen features and a relevant training dataset.
Technical Deep Dive
▶ Watch: First idea: training local model from Diaphora's matches (4:45)
The integration of machine learning into Diaphora primarily focuses on the classification stage of binary diffing, aiming to refine existing function matches and discover new, non-obvious ones. Koret's methodology involved a structured process of data extraction, feature engineering, and model training.
The first step in building a machine learning model is acquiring and preparing the data. Koret leveraged Diaphora's existing capabilities to export analysis results from IDA Pro into SQLite databases. For the Cisco Talos Data Set 2, this resulted in 4,047 individual SQLite databases for 4,047 binaries. Managing this many databases is impractical for ML training, so a tool was developed to parse these databases and consolidate relevant information into a single CSV file.
From the extensive internal data Diaphora collects, Koret carefully selected a set of features deemed most relevant for function similarity analysis. These included:
- Basic Blocks: The number of basic blocks within a function.
- Edges: The number of control flow edges connecting basic blocks.
- In-degree and Out-degree: The number of incoming and outgoing control flow transfers for a function.
- Cyclomatic Complexity: A metric indicating the complexity of a function's control flow graph.
- Prime Value Signatures: A more complex feature derived from abstract syntax trees (ASTs), where different parts of the AST are assigned prime values to create a unique signature. Koret noted that a detailed explanation of this feature would be better suited for a workshop due to its complexity.
- Pseudo Code Similarity: The similarity ratio of the decompiled pseudo-code (both clean and Hex-Rays versions) using Python's
SequenceMatcher.quick_ratioalgorithm. This provides a textual comparison of the high-level logic.
For numeric features, Koret introduced "prime fields" by calculating the minimum, maximum, and difference between these values when comparing two functions. This approach aimed to capture not just absolute values but also the relative differences, which can be crucial for identifying similar but slightly modified functions.
The challenge of building a sufficiently large and diverse dataset was a recurring theme. Koret initially explored existing public datasets:
- Johns Hopkins Applied Physics Laboratory (APL) Allstar data set: A massive 1.1 terabyte repository containing 32,000 Debian packages and 20,000 binary executables for architectures like Intel x86/64, ARM, PowerPC, MIPS, and S390X. While comprehensive in some areas, it lacked support for Windows, Android, iOS, Mac, RISC-V, and Microsoft Visual C compilers, making it incomplete for a truly generalized model. Furthermore, processing this volume of data, even just for IDA analysis and Diaphora export, was beyond Koret's resources.
- Cisco Talos Binary Function Similarity data set: This dataset is split into three parts. Koret focused on Data Set 2, which, despite being the smallest, still contained 4,047 binaries. Generating a full comparison matrix (every binary against every other binary) for Data Set 2 would involve roughly 16 million operations, estimated to take two months on Koret's optimized home setup. This underscored the impracticality of "big data" approaches for independent researchers.
Given these resource constraints, Koret shifted his focus from building one gigantic model to exploring more targeted strategies. His attempts to create a "perfect subset" of data (e.g., only Diffutils, Coreutils, Putty binaries) for training proved unsuccessful, as models trained on one subset consistently performed poorly when applied to other, unrelated subsets. This highlighted the difficulty of generalization from small, specific datasets.
Ultimately, Koret settled on a Decision Tree Classifier for the core of Diaphora's ML engine. During testing of various algorithms, including Random Forest, he observed that while Random Forest offered marginally better accuracy, its significantly higher computational cost for both training and prediction made it unsuitable for Diaphora, where performance is critical. The Decision Tree Classifier provided a good balance of speed, accuracy, and interpretability, making it a pragmatic choice. Koret acknowledged that this approach might lead to some overfitting, but given the resource limitations and the goal of practical utility, it was an acceptable trade-off. The focus was on building models that worked well for their intended, specialized purposes rather than striving for theoretical perfection on generalized data.
Demo / Proof of Concept
▶ Watch: Second idea: huge model, but building data sets difficult (7:05)
While the talk didn't feature a live, interactive demo in the traditional sense, Koret presented compelling results from his integration of machine learning into Diaphora, effectively serving as a proof of concept for the specialized model approach.
The first significant demonstration involved the Microsoft antivirus engine (MP engine). Koret explained that Microsoft often releases new versions of the MP engine DLLs with symbols stripped, while the corresponding symbols are published weeks or even months later, if at all. This scenario is a prime candidate for specialized diffing. Koret took two versions of the MP engine DLLs (versions 19700 and 20000), one with symbols and one without. He then used these to build a training dataset where known good matches (functions identified by symbols) were labeled as positive, and five random, unrelated functions were labeled as negative to create a balanced dataset. This process created a decision forest model of approximately 300MB, trained on 1.6 million rows. Initial, rudimentary tests showed that this specialized model slightly improved the number of new function matches found in future, symbol-stripped versions of the MP engine. This confirmed the viability of training models tailored to specific software projects.
Koret also tested this MP engine-trained model against the Cisco Talos Data Set 2, and, as expected, the results were "horrible." This reinforced the finding that a model highly specialized for one set of binaries (e.g., MP engine) cannot be magically expected to generalize to a completely different set of binaries.
The most robust proof of concept centered around the "amalgamation" model. This model was trained on a custom dataset combining binaries from Diaphora's extensive testing suite, the Cisco Talos Data Set 2, and various user-reported binaries that had posed diffing challenges over Diaphora's nine-year history. This carefully curated dataset, even after compression with bzip2 -9, remained a manageable 50MB, making it easily distributable.
Koret presented several concrete examples of the amalgamation model's performance improvements:
- RISC-V High-Cool Loader vs. Intel x86-64 High-Cool Loader: Without the ML model, Diaphora found 433 partial matches. With the model, this increased to 501 results. Koret specifically highlighted three functions that consistently appeared in both results but noted the significant increase in newly identified functions.
- BusyBox Intel x86-64 (no symbols) vs. BusyBox PowerPC (symbols): This cross-architecture comparison showed an improvement from 1,212 matches without ML to 1,301 matches with the ML engine.
- iOS kernel 10.3.1 vs. macOS kernel 10.12.4: A challenging comparison between two large, complex operating system kernels. Without the ML engine, Diaphora identified 4,134 functions, representing approximately 46% of potential matches. With the ML engine, this number rose to 4,927 functions, or 53.73% of potential matches.
Crucially, Koret emphasized that during manual verification of hundreds of these additional matches, the amalgamation model did not introduce a significant number of false positives. This is attributed partly to Diaphora's existing internal mechanisms designed to mitigate false positives, which complement the ML classifier. The ability to find more matches without compromising accuracy is a significant win for reverse engineers working on complex, evolving binaries.
The practical workflow for users to leverage this functionality is also straightforward. Koret outlined a process: analyze binaries with IDA, export with Diaphora, use a provided script to generate a CSV from the SQLite databases, train a custom model using this CSV, and then configure Diaphora to use the newly trained model. He also provided a script that automates these steps, allowing users to simply drop their binaries into a directory and generate a specialized model. Furthermore, users can concatenate their project-specific CSV data with Diaphora's public amalgamation dataset to create an even more comprehensive model tailored to their specific needs while retaining the generalized benefits.
Defensive Implications
▶ Watch: Scarcity of good, usable binary diffing datasets (8:20)
Joxean Koret's work on integrating machine learning into Diaphora offers several significant implications for security defenders and reverse engineers:
- Enhanced Patch Analysis: One of the most critical defensive applications is improved patch analysis. When vendors release security patches, they often provide only updated binaries, with symbols stripped. Defenders need to quickly identify the changes between the patched and unpatched versions to understand the vulnerability addressed and assess potential bypasses or new attack surfaces. By training specialized ML models on previous versions of the software (especially if some versions had symbols), defenders can significantly increase the accuracy and completeness of function matching, even when faced with heavy compiler optimizations or minor code refactoring. This allows for a more efficient and thorough understanding of security updates.
- Faster Vulnerability Discovery and Exploitation: The ability to find more reliable matches between different versions of a binary, or even across different architectures, directly aids in vulnerability discovery. If a vulnerability is found in one version or platform, defenders can use the enhanced diffing capabilities to quickly pinpoint the corresponding vulnerable code in other versions or architectures, facilitating faster analysis and the development of mitigations or detection rules. For exploit developers, this means more efficient porting of exploits across different software versions or environments.
- Reverse Engineering Proprietary Software: Many defensive tasks involve reverse engineering proprietary software, such as endpoint detection and response (EDR) agents, industrial control systems (ICS) firmware, or specialized drivers. These binaries are almost always stripped of symbols and frequently updated. Specialized ML models, trained on a history of these proprietary binaries, can drastically improve the ability to track changes, identify new functionalities, or detect evasive techniques introduced in updates. This is particularly relevant for tracking the evolution of malware or advanced persistent threat (APT) tools.
- Custom Model Creation for Specific Targets: Koret's decision to distribute not just a pre-trained model but also the tools and datasets for users to build their own specialized models is a powerful defensive enabler. Organizations often focus on a limited set of critical software or systems. Defenders can leverage Koret's framework to create highly accurate ML models specifically tuned to their high-value targets, whether it's an internal application, a specific third-party product, or a frequently analyzed piece of malware. This bespoke approach ensures maximum relevance and accuracy for their unique defensive posture.
- Reduced Manual Analysis Overhead: For large and complex binaries, such as operating system kernels (as demonstrated with the iOS/macOS kernel diffing), finding even a few hundred more reliable function matches can save countless hours of manual reverse engineering effort. By automating and improving the initial mapping and classification, defenders can allocate their scarce expert resources to deeper analysis of truly novel or critical changes, rather than painstakingly re-identifying known functions.
- Cross-Architecture and Cross-Compiler Analysis: The examples of diffing BusyBox across Intel and PowerPC, or high-cool loaders across RISC-V and Intel, highlight the utility for multi-platform defense. Organizations often deploy software across diverse hardware and operating systems. Understanding how a vulnerability or a piece of malware manifests across these different environments is crucial, and ML-enhanced binary diffing can bridge the gaps caused by different compilers and architectures.
In essence, Koret's work provides a practical, accessible pathway for defenders to integrate modern machine learning techniques into their binary analysis workflows, moving beyond traditional heuristics to gain a deeper, more efficient understanding of software changes and threats.
Key Takeaways
- Machine learning is a viable and practical enhancement for binary diffing, particularly for function classification and improving match reliability. While academic approaches often fall short, pragmatic application can yield significant improvements.
- Attempting to build a single, gigantic, generalized ML model for all binary diffing scenarios is largely impractical for independent researchers due to prohibitive resource requirements for data collection, processing, and model training.
- Highly specialized ML models, trained on related binaries (e.g., different versions of the same software or project), are extremely effective. This approach leverages inherent consistency within a specific software's development lifecycle to achieve high accuracy.
- Simple, interpretable ML algorithms, such as the Decision Tree Classifier, can be sufficiently powerful and efficient for real-world binary diffing tasks. More complex models are not always necessary and can introduce unacceptable performance overhead.
- The "amalgamation" model, built from a curated dataset of Diaphora's testing suite and real-world binaries, offers a robust and distributable solution for general-purpose binary diffing enhancement. It significantly increases match counts without substantial increases in false positives.
- The true value lies not just in pre-trained models, but in providing the tools and methodologies for users to build their own specialized models. This empowers reverse engineers to tailor ML solutions to their specific projects and binary targets.
About the Speaker(s)
Joxean Koret is a seasoned reverse engineer and the creator of Diaphora, a popular open-source binary diffing tool. Throughout his talk, Koret openly identifies himself as primarily a reverse engineer, emphasizing that he is "not a machine learning guy at all." His research into applying machine learning to binary diffing was a self-taught journey driven by a desire to improve Diaphora's capabilities. This background provides a practical, industry-focused perspective on the challenges and successes of integrating advanced techniques into existing security tools.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Koret brings genuine practitioner credibility to a space drowning in academic vaporware — he built the damn tool, ran the experiments, hit the walls, and shows you exactly what broke and why. The specialized-model insight is legitimately useful, and the honest accounting of failure modes (local model redundancy, gigantic-model resource hell) is rarer and more valuable than another 'our model beats BinDiff' paper.
Heather Calloway (CISO) — PASS
Solid reverse engineering research from a credible practitioner, but it lives entirely in the binary analysis tooling layer. There is no governance angle, no institutional risk framing, and no meaningful path to executive or defender-program action.