Unleashing the Power of Generative Model in Recovering Variable Names from Stripped Binary
Xiangzhe Xu
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Binary Analysis
Overview
In this compelling talk at the NDSS Symposium, Xiangzhe Xu from Purdue University presented groundbreaking research on recovering variable names from stripped binaries using generative code models. The work introduces a novel context-aware fine-tuning technique and a preference optimization method to align model generations with developer naming conventions. This advancement is crucial for democratizing reverse engineering, making complex binary programs more accessible to a broader audience—from security experts to ordinary users without a computer science background.
Key moments
- 0:00 Introduction and vision for democratized reverse engineering
- 1:00 Motivation: Ransomware impact and need for local analysis
- 2:30 Why symbols are essential for AI in binary analysis
- 3:00 Demonstration: Symbols significantly improve AI's understanding of binaries
- 4:40 Demonstration: Even experts benefit from symbol information
- 6:00 Technical Challenge 1: Generative models for unseen variable names
- 7:00 Technical Challenge 2: Handling limited context in binaries
Unleashing the Power of Generative Model in Recovering Variable Names from Stripped Binary
Speakers: Xiangzhe Xu, Purdue University
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=5Lq3BhlmXW0
Overview
In this compelling talk at the NDSS Symposium, Xiangzhe Xu from Purdue University presented groundbreaking research on recovering variable names from stripped binaries using generative code models. The work introduces a novel context-aware fine-tuning technique and a preference optimization method to align model generations with developer naming conventions. This advancement is crucial for democratizing reverse engineering, making complex binary programs more accessible to a broader audience—from security experts to ordinary users without a computer science background.
The core motivation behind this research stems from the increasing difficulty of understanding software, especially in the face of sophisticated threats like ransomware, which caused hundreds of millions in losses in the US between 2023 and 2024. While large language models (LLMs) have shown impressive reasoning capabilities on source code, their effectiveness does not automatically transfer to the binary domain due to the critical lack of symbolic information. This research bridges that gap, enabling AI agents to provide understandable explanations of binary program behavior, thus empowering users to be aware of risks before execution.
The presented solution not only enhances the capabilities of AI analyzers but also significantly benefits world-class reverse engineering experts. By reconstructing meaningful variable names, the tool transforms opaque decompiled code into semantically rich representations, making it easier to navigate vast codebases, identify relevant functions, and understand program intent. This work represents a significant step towards a future where intelligent coding agents can locally analyze and explain software in natural language, fostering a more secure and transparent digital environment for everyone.
Background
▶ Watch: Introduction and vision for democratized reverse engineering (0:00)
Understanding binary programs has historically been a notoriously difficult task, often reserved for highly specialized reverse engineering experts. This complexity arises from intricate data dependencies, multiple intertwined components, and most critically, the absence of meaningful symbolic information like variable and function names. While experts might eventually deduce the purpose of a program without symbols, this process is painstakingly slow and beyond the reach of average users.
The advent of AI, particularly large language models (LLMs), has brought impressive reasoning capabilities to the domain of source code understanding. These models can generate accurate summaries and insights from human-readable code. However, this success has not readily migrated to binary analysis. The primary impediment is the lack of symbols. When source code is compiled into a binary, especially when stripped for size or obfuscation, most symbolic information is discarded. This leaves decompiled code with generic, semantically meaningless names like var_1, var_2, making it incredibly challenging for both human and AI analysts to grasp the program's true intent.
The speaker highlighted the societal impact of this problem, citing over 5,000 ransomware instances in the United States between 2023 and 2024, resulting in hundreds of millions of dollars in losses. Mitigating such threats requires a deeper understanding of malicious software. The current paradigm relies on centralized third-party certifications (e.g., app stores), which represent a large trusted computing base. The vision presented is a future where every user has a local AI model capable of analyzing binaries and explaining their behavior in plain, natural language, thus democratizing security awareness and allowing users to understand risks proactively.
Previous attempts at symbol recovery, as noted by the speaker, primarily leveraged classification models. These models operate by selecting the closest matching name from a pre-defined training dataset. This approach faces a fundamental limitation: it struggles significantly with unseen names—variables or functions whose names do not exist in the training data. This is a common occurrence in real-world software, where developers constantly coin new, context-specific identifiers. Furthermore, binary programs inherently provide very limited contextual information compared to their source code counterparts, making it harder to infer meaning even for known names. The talk also acknowledged that real-world name distributions are heavily biased, exhibiting a long tail where over half of the names in high-quality datasets appear only once, leading to models overfitting frequent names and performing poorly on rare or unique ones.
Key Findings
▶ Watch: Why symbols are essential for AI in binary analysis (2:30)
The research presented by Xiangzhe Xu introduces several key findings and contributions that significantly advance the state-of-the-art in binary symbol recovery:
First, the work demonstrates the superior capability of a generative model over traditional classification-based approaches for symbol recovery. Unlike classification models that are limited to predicting names present in their training data, generative models can predict names token by token, allowing them to compose complex or entirely unseen names from common linguistic tokens (e.g., "IP header length"). This capability is crucial for handling the vast diversity and originality of developer-chosen identifiers.
Second, the paper highlights the critical importance of context-aware fine-tuning for large language models in the binary domain. Recognizing the inherent lack of contextual information in decompiled code (e.g., missing function names or high-level program intent), the proposed technique explicitly incorporates hints from caller and callee functions into the LLM's input prompt. By training the model to reason with both the function body and these contextual cues, it can infer variable names more accurately, even when local information is sparse.
Third, the research introduces symbol preference optimization as a novel method to mitigate the statistical bias prevalent in real-world naming distributions. Real-world codebases often exhibit a "long tail" phenomenon, where a small number of names appear frequently, while a large number of names appear only once. This can cause models to overfit frequent names. Preference optimization addresses this by training the model to assign higher probabilities to developer-preferred, semantically accurate names (e.g., "packet") and lower probabilities to statistically biased, but less contextually fitting names (e.g., "buffer").
The evaluation results presented are compelling. The proposed technique consistently outperforms both classification-based baselines and generative baselines trained with straightforward supervised fine-tuning. Remarkably, the model even surpasses the performance of significantly larger, advanced pre-trained LLMs such as GPT-4 and Code Llama 70B on this specific task, underscoring the challenge of symbol reconstruction for general-purpose LLMs and the effectiveness of their specialized approach. Furthermore, the tool demonstrates superior generalization to rare and unseen names, performing best when name frequency is very low or even zero. A human-centric evaluation using GPT-4 as an assessor revealed that the tool predicts "good" names (names that a human developer would find appropriate) in over 50% of cases, validating its practical utility.
Technical Deep Dive
▶ Watch: Demonstration: Symbols significantly improve AI's understanding of binaries (3:00)
The technical core of this research addresses three primary challenges in binary symbol recovery: the inability of previous models to handle unseen names, the inherent lack of context in stripped binaries, and the statistical bias in real-world name distributions.
To overcome the limitation of classification models struggling with unseen names, the researchers pivot to a generative model approach. Unlike classification, which selects from a fixed vocabulary, a generative model predicts the target name token by token. This allows it to construct novel, complex names that were not explicitly present in the training data but are composed of common, familiar tokens. For instance, a name like "IP header length," while potentially unseen as a whole, can be generated from the commonly used tokens "IP," "header," and "length." This compositional capability significantly enhances the model's generalization power.
The second major technical contribution is context-aware fine-tuning. Binary programs, especially after decompilation, suffer from a severe deficiency of contextual information. A source code function name like send_packet immediately implies the function's purpose and can guide variable naming. In decompiled code, this high-level semantic hint is often lost. To compensate, the proposed method enriches the input to the large language model. During training and inference, it collects name hints from both the caller functions (functions that invoke the current function) and callee functions (functions invoked by the current function). These hints, which are essentially names previously predicted or known in the surrounding code, are then integrated into the input prompt alongside the decompiled function body. By training the model to reason across these expanded contextual inputs, it learns to infer more accurate variable names that align with the broader program logic. This is a sophisticated prompting strategy that goes beyond simple supervised fine-tuning.
The third technical innovation tackles the issue of biased real-world name distribution. Real-world codebases exhibit a "long tail" phenomenon, meaning a few names are highly frequent, while a vast majority appear only once or very rarely. A standard language model, trained on such data, naturally develops a bias towards predicting frequent names. For example, if "buffer" appears much more frequently than "packet" in the training data, the model might incorrectly suggest "buffer" even when "packet" is the semantically correct and developer-preferred name for a network-related variable. To counteract this, the researchers introduce symbol preference optimization. This technique involves creating pairwise training data where the model is presented with a "good" (developer-preferred) name and a "bad" (biased, less fitting) name for the same variable. The model is then trained to assign a higher probability to the good name and a lower probability to the biased name. This explicitly teaches the model to prioritize semantic fit and developer intent over mere statistical frequency, mitigating the negative impact of data bias.
The overall training pipeline is described as a three-stage process:
- Context-aware supervised fine-tuning: Initial training of the generative model with the enriched contextual inputs.
- Preference optimization: Refinement of the model to align with developer naming preferences and reduce statistical bias.
- Bias mitigation: A further stage, likely involving techniques derived from the preference optimization, to specifically reduce the model's tendency to favor overly frequent names. (Detailed specifics for stages 2 and 3 were referred to the paper due to time constraints).
An iterative inference pipeline is also employed. This involves an initial inference pass based on local context, followed by reconstruction of global context from these initial predictions, and then repeated iterative inference until the predictions stabilize or converge. This allows the model to refine its predictions by leveraging progressively richer contextual information across the entire binary. The speaker mentioned fine-tuning models like Llama 7B and 14B, and Code Llama 2B, demonstrating the adaptability of their approach across different model sizes.
Demo / Proof of Concept
▶ Watch: Technical Challenge 1: Generative models for unseen variable names (6:00)
The presentation included several compelling demonstrations and examples to illustrate the practical impact of their symbol recovery tool. While not a live, interactive demo, the speaker showcased "research demo" results that effectively served as a proof of concept.
One key example involved a relatively simple decompiled function responsible for initializing an image data structure with given width and height, calculating its size, and ensuring 4-byte alignment. Before reconstruction, the decompiled code featured generic variable names like var_1 and var_2. After the tool's application, these were transformed into semantically meaningful names such as image_width and image_height. The speaker then demonstrated the impact on AI analysis: when the original, symbol-stripped code was fed to GPT-4, the LLM struggled to understand the function's full semantic meaning, notably missing the 4-byte alignment constraint. However, when the reconstructed code with meaningful symbols was provided, GPT-4 generated a summary of comparable quality to one derived from the original source code, accurately identifying all key aspects, including the alignment. This clearly illustrated how their tool enables further development of AI reasoning agents in the binary domain.
A second powerful demonstration focused on a real-world malware sample—specifically, a C2 (Command and Control) client. The speaker presented a code snippet from this malware, showing its decompiled form both before and after symbol reconstruction. On the left (unreconstructed), the function was a jumble of generic names. On the right (reconstructed), the function's purpose immediately became clear, with names like command_parsing and command_dispatching. This example highlighted the immense benefit for human reverse engineers. In a binary with potentially thousands of functions, the ability to quickly identify relevant functions by searching for keywords like "command" or "network" significantly accelerates the analysis process. The speaker reinforced this with real-world evidence, referencing a report from a world-class hacker group analyzing a VPN software patch, who explicitly stated their reliance on symbol information due to the overwhelming number of functions to review manually.
The overall "research demo" capability, as stated, is able to predict good names in more than 50% of cases, which is a significant achievement given the complexity of the task and the challenging nature of stripped binaries. These demonstrations collectively underscored the tool's ability to make binary programs drastically easier to understand for both AI analyzers and human experts.
Defensive Implications
▶ Watch: Technical Challenge 2: Handling limited context in binaries (7:00)
The advancements in binary symbol recovery presented in this talk have profound defensive implications, primarily by enhancing the capabilities of security analysts and empowering ordinary users. While not a direct defensive patch or vulnerability fix, this research provides foundational technology that strengthens the security posture across multiple fronts.
Firstly, Improved Threat Intelligence and Malware Analysis: The ability to recover meaningful variable and function names from stripped malware binaries (like the C2 client example) drastically accelerates and deepens malware analysis. Security researchers can more quickly understand the intent, capabilities, and communication protocols of malicious software. This leads to more precise Indicators of Compromise (IOCs), better behavioral analysis, and more effective defensive strategies, such as developing signatures, network rules, or behavioral heuristics. Understanding malware's internal logic, previously a monumental task, becomes significantly more manageable, enabling faster response to zero-day threats or sophisticated APT campaigns.
Secondly, Enhanced Vulnerability Research and Incident Response: For security engineers engaged in vulnerability research, understanding complex, stripped binaries—whether proprietary software, firmware, or critical infrastructure components—is paramount. Recovering symbols streamlines the process of identifying data structures, internal states, and critical control flow, making it easier to discover logical flaws, memory corruption vulnerabilities, or backdoors. In incident response, rapidly comprehending compromised binaries allows responders to assess the scope of an attack, identify persistence mechanisms, and formulate containment and eradication strategies more efficiently.
Thirdly, Democratization of Security Analysis: The long-term vision articulated by the speaker—local AI agents explaining binary risks in natural language to ordinary users—is transformative. This democratizes security, shifting the reliance from large, centralized trusted third parties to individual empowerment. Users could receive alerts like, "I analyzed this program and found it may try to access your wallet after receiving a command from the server." This proactive risk awareness could significantly reduce the attack surface for social engineering, untrusted software, and supply chain attacks, enabling users to make informed decisions about what software to run on their machines.
Finally, Enabling Next-Generation AI Security Tools: By providing the "missing link" of symbolic information, this research unlocks the full potential of large language models for binary analysis. Future AI-powered security tools will be able to perform deeper semantic reasoning on binaries, leading to more intelligent static and dynamic analysis, automated vulnerability detection, and even automated reverse engineering. This paves the way for a new generation of security defenses that can understand and respond to threats with unprecedented sophistication.
Key Takeaways
- Symbol Recovery is Essential: Reconstructing variable names from stripped binaries is critical for democratizing reverse engineering, making software understandable for both human experts and AI analyzers, and enhancing overall cybersecurity.
- Generative Models Excel at Unseen Names: Unlike traditional classification models, generative models can predict and compose complex, novel names token by token, effectively addressing the challenge of unseen identifiers in real-world code.
- Context is King for Binary Analysis: Integrating contextual information from caller and callee functions into the LLM's input significantly improves the accuracy of variable name recovery in the semantically sparse binary domain.
- Mitigating Data Bias is Crucial: Real-world name distributions are highly biased. Techniques like symbol preference optimization are vital for teaching models to prioritize semantically appropriate, developer-preferred names over statistically frequent but less fitting alternatives.
- Superior Performance and Generalization: The proposed techniques outperform both existing baselines and significantly larger, general-purpose LLMs (like GPT-4 and Code Llama 70B) in this specialized task, demonstrating strong generalization capabilities, especially for rare and unseen names.
- Enabling Future AI and Human Analysis: This research facilitates more effective threat intelligence, vulnerability research, and incident response for experts, while simultaneously paving the way for local AI agents to explain software risks to ordinary users in an accessible manner.
About the Speaker(s)
Xiangzhe Xu is a researcher from Purdue University. His work focuses on leveraging the power of generative code models to address long-standing challenges in binary analysis, specifically the recovery of symbolic information like variable names. His research aims to enhance the understandability of stripped binary programs for both human reverse engineers and advanced AI analysis agents, contributing to the broader vision of democratizing software security.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Solid academic research that makes a genuine contribution to the binary analysis problem — generative models over classifiers for symbol recovery is the right call, the context-aware fine-tuning trick is clever, and the preference optimization to fight long-tail bias shows the authors actually thought through the failure modes rather than just benchmarking happy-path cases. Not a world-shaker, but this is real work that will get cited and built upon.
Heather Calloway (CISO) — WEAK
Technically credible research on a real problem in binary analysis, but the article overclaims its significance and fails to bridge the gap to any defensible operational or governance conclusion. The work is a legitimate research contribution — not a security program input.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025