Beyond Classification: Inferring Function Names in Stripped Binaries via Domain Adapted LLMs
Linxi Jiang (Ohio State University)
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Binary Analysis
Overview
The ability to accurately infer function names in stripped binaries is a critical challenge in reverse engineering, with profound implications for fields such as malware analysis, vulnerability research, and proprietary software understanding. Without meaningful function names, reverse engineers are confronted with a deluge of generic labels (e.g., sub401000), severely hindering their comprehension of a program's logic and purpose. This talk by Linxi Jiang from Ohio State University introduces SimGen, a novel framework that leverages domain-adapted large language models (LLMs) to address this persistent problem, moving beyond traditional classification-based approaches that have shown significant limitations in real-world scenarios.
Key moments
- 1:50 Previous work's weak generalization on unseen binaries
- 3:10 Identifying pervasive data leakage in existing datasets
- 6:00 SimGen's approach: LLMs, domain adaptation, deduplication, efficient learning
- 7:00 Detailing SimGen's novel data deduplication strategy
- 8:20 Parameter-efficient model training with Laura (1 GPU, 24h)
- 8:45 Comprehensive evaluation across architectures and optimization levels
- 9:20 SimGen significantly outperforms baselines in function name inference
Beyond Classification: Inferring Function Names in Stripped Binaries via Domain Adapted LLMs
Speakers: Linxi Jiang (Ohio State University)
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=lWZWsuRudOk
Overview
The ability to accurately infer function names in stripped binaries is a critical challenge in reverse engineering, with profound implications for fields such as malware analysis, vulnerability research, and proprietary software understanding. Without meaningful function names, reverse engineers are confronted with a deluge of generic labels (e.g., sub_401000), severely hindering their comprehension of a program's logic and purpose. This talk by Linxi Jiang from Ohio State University introduces SimGen, a novel framework that leverages domain-adapted large language models (LLMs) to address this persistent problem, moving beyond traditional classification-based approaches that have shown significant limitations in real-world scenarios.
SimGen tackles several long-standing issues, including the poor generalization of existing methods to unseen binaries, the knowledge gap between general-purpose LLMs and code comprehension, rampant data leakage in benchmark datasets, and the prohibitive computational cost of training large models. By integrating innovative data processing, summary generation, and parameter-efficient learning (PEL) techniques, SimGen significantly elevates the state-of-the-art in function name prediction. This work is crucial because it promises to dramatically improve the efficiency and accuracy of reverse engineering tasks, making complex, stripped binaries more accessible for analysis and enhancing the overall security posture of software.
Background
▶ Watch: Previous work's weak generalization on unseen binaries (1:50)
Function name prediction is a cornerstone of effective reverse engineering. In a typical development environment, functions are given descriptive names like getRandomIP or clearBuffer that immediately convey their intent. However, when software is compiled into a stripped binary, much of this symbolic information—including function names, variable names, and type information—is deliberately removed to reduce file size or hinder reverse engineering efforts. The result is a binary that is exceedingly difficult for humans to comprehend, even with the aid of powerful decompilers that translate machine code back into a high-level language. The decompiler output, while structured, often uses generic, meaningless identifiers, as illustrated by the speaker with an example function that is incomprehensible until its true name, getRandomIP, is revealed.
Previous research efforts have largely approached function name prediction as a classification task. Tools like ASMDep and XFL have demonstrated reasonable performance on specific datasets. However, a significant limitation of these traditional methods is their weak generalization ability. The speaker highlights a stark example: the performance of a model named SIM-LM plummeted from 64% accuracy to a mere 8% when encountering "unseen binaries." This dramatic drop underscores the inability of classification models to adapt to new projects, different compiler settings, or varied programmer naming styles, severely limiting their utility in real-world applications where novel binaries are constantly encountered.
The emergence of large language models (LLMs) offers a promising avenue due to their strong capabilities in textual understanding and generalization. However, applying LLMs directly to code comprehension tasks, especially with decompiler output from stripped binaries, presents its own set of challenges. There's a notable knowledge gap between the natural language tasks LLMs are typically trained on and the nuanced domain of code. Prior work evaluating the quality of function summaries generated by LLMs from stripped decompiled code showed very low scores (almost indistinguishable from "robes," which are notoriously hard for humans to understand), indicating that raw LLMs struggle significantly with this domain.
Furthermore, the research community has grappled with data leakage in existing benchmark datasets. The speaker reveals that up to 77.9% of data samples in datasets used by previous work contained duplicates between training and test sets. These "duplicates" weren't always exact matches but often structurally identical functions with only address-related function calls differing. Such leakage artificially inflates reported performance metrics and misrepresents a model's true generalization capacity. Finally, the resource-expensive training of large models, with examples like Llama requiring 2048 Nvidia A100 GPUs and 21 days for base model training, poses a substantial barrier for researchers and practitioners to fine-tune these models for specific tasks like function name prediction.
Key Findings
▶ Watch: SimGen's approach: LLMs, domain adaptation, deduplication, efficient learning (6:00)
The research presented in this talk identifies several key insights and proposes a novel framework, SimGen, to overcome the aforementioned challenges in function name prediction for stripped binaries. The core findings and contributions are:
- Generative Large Model Approach: Moving beyond the limitations of classification-based methods,
SimGenadopts a generative large model approach. This allows the model to predict function names as sequences of tokens, offering greater flexibility and expressiveness compared to selecting from a predefined list of classes. - Domain Adaptive Learning with Function Summaries: To bridge the knowledge gap between general LLMs and code comprehension,
SimGenincorporates a domain adaptive learning strategy. This involves the generation and rigorous manual verification of high-quality function summaries that serve as an intermediate representation, guiding the LLM to better understand the code's semantics. These summaries are explicitly included during the training phase. - Robust Data Deduplication: To mitigate the pervasive issue of data leakage,
SimGenemploys a sophisticated deduplication strategy. This involves modifying address-related function calls within the decompiled code to normalize function bodies, enabling accurate exact matching to identify and remove duplicates across training and test sets. The speaker reported finding 77.9% duplicates in previous datasets, highlighting the necessity of this step. - Parameter Efficient Learning (PEL): To address the challenge of resource expensive training,
SimGenutilizes Low-Rank Adaptation (LoRA), a parameter efficient learning technique. LoRA allows for fine-tuning large models with significantly fewer computational resources, enabling training on a single GPU within 24 hours, a stark contrast to the weeks and thousands of GPUs required for base model training. - Superior Performance and New Benchmark: Through these innovations,
SimGendemonstrably achieves much higher performance than existing baselines (ASMDep, XFL) across various architectures, optimization levels, and even under obfuscation settings. This establishes a new benchmark for function name prediction in stripped binaries, showcasing its strong generalization ability and practical utility. The framework's ability to maintain performance even with unseen or obfuscated binaries is a major step forward.
Technical Deep Dive
▶ Watch: Detailing SimGen's novel data deduplication strategy (7:00)
The SimGen framework is meticulously designed and implemented in three distinct steps: data processing, summary generation, and model training. These steps collectively enable the fine-tuning of a generative LLM for robust function name prediction.
Data Processing
The initial phase focuses on preparing the input data for the LLM.
- Decompilation: The process begins by compiling source code into binaries. For these binaries, a decompiler like Ghidra is used to generate the high-level decompiled code. This decompiled code serves as the primary input for the
SimGenmodel. - Function Name Masking: To facilitate a generative task, the actual function name within the decompiled code is masked. This is achieved by using a tool like ANTLR4 to build an Abstract Syntax Tree (AST) of the decompiled code. By traversing the AST, the original function name is identified and replaced with a special token,
[MASK], indicating to the LLM that this is the target for prediction. - Data Deduplication: This is a crucial step to prevent data leakage and ensure the model's true generalization ability is evaluated. The speaker observed that many "duplicate" functions in existing datasets differed only in their address-related function calls. To normalize these,
SimGenmodifies all address-related function calls within the function body to a generic placeholder. After this normalization, an exact match comparison is performed to identify and remove all duplicate functions, ensuring that the training, validation, and test sets are truly distinct. This rigorous deduplication process, which identified 77.9% duplicates in previous datasets, is essential for obtaining reliable evaluation results.
Summary Generation
To bridge the knowledge gap between general LLMs and the specifics of code, SimGen leverages function summaries during its training phase.
- High-Quality Comments: The first source of summaries comes from well-designed comments found in source code. If a function has a high-quality descriptive comment, the description part of that comment can be directly extracted and used as a function summary.
- Code LLM Generated Summaries: Recognizing that not all functions have high-quality comments,
SimGenalso employs a code large model (like Code Llama itself, or another similar model) to generate summaries for functions. A specific prompt is used to guide this generation. Crucially, all generated summaries undergo manual verification to ensure their accuracy and quality before being included in the training dataset. These summaries are incorporated into the fine-tuning data set alongside the decompiled code.
Model Training
The final step involves fine-tuning the LLM using the prepared data.
- Base Model Selection: The base LLM chosen for
SimGenis Code Llama, known for its strong code comprehension capabilities. - Parameter Efficient Learning (PEL) with LoRA: To overcome the significant computational cost associated with training large models,
SimGenemploys Low-Rank Adaptation (LoRA). LoRA works by freezing the pre-trained model weights and injecting small, trainable rank decomposition matrices into each layer of the transformer architecture. This drastically reduces the number of trainable parameters. Using LoRA,SimGencan fine-tune the Code Llama model on a single GPU within 24 hours, a remarkable feat compared to the weeks and thousands of GPUs required for training the base Llama model. - Inference Prompt: During inference, a specific prompt structure is used to query the fine-tuned model. This prompt includes the masked decompiled code, and the model is tasked with filling in the
[MASK]token with the most appropriate function name.
Evaluation and Results
The SimGen framework was evaluated on a comprehensive dataset comprising 33 open-source C projects. These projects were compiled under four different computer architectures (e.g., X64 architecture) and four optimization levels (e.g., O3 optimization levels). Additionally, an ablation study was conducted using binaries compiled with three different obfuscation options to test robustness.
The performance of SimGen was compared against established baselines like ASMDep and XFL, using standard metrics such as Precision, Recall, and F1-score. The results consistently showed SimGen achieving significantly higher performance across all tested configurations. For instance, SimGen with decompiled code as input consistently outperformed all other combinations, including baselines using assembly code or raw decompiled code. The ablation study also confirmed the performance improvement gained through the domain adaptation learning strategy.
A critical aspect explored was the potential for data leakage in the pre-training stage of the base LLM itself. The researchers adapted a memory inference attack technique, where the last five lines of decompiled code were removed, and the LLM was asked to complete the missing part. The very low CodeBLEU score (less than 0.1) between the generated and actual missing parts indicated that while LLMs might be pre-trained on source code, they are unlikely to have encountered specific stripped, decompiled code in their pre-training data, particularly the unique variations from different decompilers. This further validates the necessity of SimGen's domain adaptation. Furthermore, SimGen demonstrated the least performance drop when evaluated on real-world binaries with varying obfuscation levels, affirming its robustness.
Demo / Proof of Concept
▶ Watch: Comprehensive evaluation across architectures and optimization levels (8:45)
While the talk did not feature a live, interactive demonstration of the SimGen framework in action, the speaker presented extensive quantitative evaluation results that serve as a robust proof of concept. The detailed tables and figures showcased SimGen's superior performance across various architectures (X64 architecture), optimization levels (O3 optimization levels), and obfuscation settings, consistently outperforming established baselines like ASMDep and XFL. These empirical results, including specific percentage improvements and F1-scores, effectively demonstrate the framework's capability to infer function names accurately in stripped binaries under diverse real-world conditions. The open-sourcing of the SimGen framework further allows others to replicate and verify these findings.
Defensive Implications
▶ Watch: SimGen significantly outperforms baselines in function name inference (9:20)
The advancements brought by SimGen have significant defensive implications across several cybersecurity domains:
- Enhanced Malware Analysis: For security analysts dissecting malware, particularly sophisticated samples that are heavily stripped and obfuscated,
SimGencan dramatically reduce the time and effort required for initial triage and deep analysis. By providing meaningful function names, it transforms opaque code into more comprehensible structures, allowing analysts to quickly grasp the malware's capabilities, command-and-control mechanisms, and evasion techniques. This accelerates the development of detection signatures and countermeasures. - Improved Vulnerability Research: Security researchers hunting for vulnerabilities in proprietary or closed-source software often rely on reverse engineering stripped binaries.
SimGencan make this process far more efficient by clarifying the purpose of functions, thus helping researchers pinpoint potential areas of interest (e.g., network communication handlers, input validation routines, cryptographic operations) where vulnerabilities might reside. This facilitates faster discovery and reporting of security flaws. - Better Software Supply Chain Security: Organizations can use
SimGen-like tools to gain better visibility into third-party libraries and components, even when only stripped binaries are available. Understanding the functions within these components can help identify unintended functionalities, potential backdoors, or compliance issues, thereby strengthening the software supply chain. - Automated Reverse Engineering Tools: The
SimGenframework can be integrated into existing decompilers and automated reverse engineering platforms. This would automatically enrich the decompiler output with predicted function names, significantly boosting the readability and usability of the generated high-level code for human analysts. This automation reduces cognitive load and allows experts to focus on higher-level threat intelligence. - Benchmarking and Evaluation of Obfuscation: The robustness of
SimGenagainst different obfuscation levels provides a valuable tool for evaluating the effectiveness of various obfuscation techniques. Defenders can use this to understand how well their own software obfuscation strategies are holding up against state-of-the-art de-obfuscation tools, guiding improvements in their defensive posture.
Key Takeaways
- Generative LLMs Outperform Classification: Traditional classification-based methods for function name prediction in stripped binaries exhibit poor generalization; generative LLMs, especially with domain adaptation, offer superior performance.
- Domain Adaptation is Crucial: Bridging the knowledge gap between general LLMs and code comprehension through high-quality function summaries (manually verified or LLM-generated) significantly enhances prediction accuracy.
- Data Leakage is a Major Problem: Previous benchmark datasets suffered from extensive data leakage (up to 77.9% duplicates); robust deduplication is essential for accurate evaluation of model generalization.
- Parameter Efficient Learning Enables Accessibility: Techniques like LoRA make fine-tuning large models for specialized tasks like function name prediction feasible on modest hardware (e.g., 1 GPU in 24 hours), democratizing access to powerful LLM capabilities.
SimGenSets a New Benchmark: The proposedSimGenframework consistently outperforms existing solutions across diverse architectures, optimization levels, and obfuscation settings, establishing a new state-of-the-art.- Practical Impact on Reverse Engineering: Accurate function name inference dramatically improves the readability of stripped binaries, directly benefiting malware analysis, vulnerability research, and the overall efficiency of reverse engineering efforts.
About the Speaker(s)
Linxi Jiang is a researcher from Ohio State University. During the presentation, he shared insights into his joint work with co-first author Shinjin and his advisor Changing. His research focuses on leveraging advanced machine learning techniques, particularly large language models, to address complex challenges in binary analysis and reverse engineering. The work presented at NDSS Symposium highlights his contributions to improving the interpretability of stripped binaries, a critical area for cybersecurity.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate academic research with a real contribution — rigorous deduplication methodology that exposes inflated benchmarks, plus a working LoRA-based pipeline that actually runs on a single GPU. Solid NDSS paper material, but this is a conference talk, not a practitioner's tool, and the gap between 'we got better F1 on 33 open-source C projects' and 'this changes how you reverse malware tomorrow' is wider than the speaker acknowledges.
Heather Calloway (CISO) — WEAK
Technically credible academic work on a real problem in binary analysis, but the bridge from research finding to operational reality never gets built. The defensive implications section reads like a press release, not a threat assessment.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025