"Len or index or count, anything but v1": Predicting Variable Names in Decompilation Output with Transfer Learning
Kuntal Kumar Pal, Ati Priya Bajaj, Pratyay Banerjee, Audrey Dutcher, Mutsumi Nakamura, Zion Leonahenahe Basque
IEEE Symposium on Security and Privacy 2024 · Day 3 · Continental Ballroom 4
Overview
In the realm of computer science, the challenge of "naming things" is notoriously difficult, and its impact is profoundly felt in the domain of reverse engineering. This talk, presented by Ati Priya Bajaj from Arizona State University and her co-authors, addresses this fundamental problem by introducing Varbo, a novel, transfer learning-based approach to predict meaningful variable names and their origins in decompiled code. Binary decompilation, the process of transforming machine-level code back into high-level pseudo-code, is a critical step for security analysts, malware researchers, and those working with legacy systems. However, compilers are inherently lossy, discarding crucial source code artifacts like original variable names, leading decompilers to generate obscure, generic placeholders such as A1, A2, or V4.

Key moments
- 0:00 Introduction: The problem of naming variables in decompiled code.
- 2:18 Introducing VarBERT: A transfer learning model for variable naming.
- 3:52 VarBERT's two-step training pipeline: pre-training on source, fine-tuning on decompiled.
- 4:54 Introducing VAR Corpus: A large, publicly available decompiled code dataset.
- 5:52 Addressing key challenges: extraneous variables, data types, function duplication.
- 9:00 VarBERT's superior performance: 50% accuracy, outperforming existing methods.
- 10:00 Ablation study: Significance of pre-training for a 14% accuracy gain.
"Len or index or count, anything but v1": Predicting Variable Names in Decompilation Output with Transfer Learning
Speakers: Kuntal Kumar Pal; Ati Priya Bajaj; Pratyay Banerjee; Audrey Dutcher; Mutsumi Nakamura; Zion Leonahenahe Basque
Conference: IEEE S&P
YouTube: https://www.youtube.com/watch?v=hGQTrkeSbbg
Overview
In the realm of computer science, the challenge of "naming things" is notoriously difficult, and its impact is profoundly felt in the domain of reverse engineering. This talk, presented by Ati Priya Bajaj from Arizona State University and her co-authors, addresses this fundamental problem by introducing Varbo, a novel, transfer learning-based approach to predict meaningful variable names and their origins in decompiled code. Binary decompilation, the process of transforming machine-level code back into high-level pseudo-code, is a critical step for security analysts, malware researchers, and those working with legacy systems. However, compilers are inherently lossy, discarding crucial source code artifacts like original variable names, leading decompilers to generate obscure, generic placeholders such as A1, A2, or V4.
The obscurity introduced by these uninformative variable names significantly impedes code comprehension, making tasks like vulnerability analysis, malware reverse engineering, and understanding complex binaries far more arduous and time-consuming. Varbo aims to mitigate this by predicting human-readable, semantically meaningful variable names, thereby making decompiled code cleaner and more intuitive. The researchers highlight that while existing work attempts variable name prediction, many approaches still yield non-meaningful names, such as V2, indicating a gap in current methodologies. Varbo’s innovative use of transfer learning, bridging the visual disparity between source code and decompiled code while leveraging their semantic similarities, represents a significant step forward in enhancing the usability and interpretability of decompilation output.
This work is particularly vital for the security community, as it directly impacts the efficiency and accuracy of understanding compiled software when source code is unavailable. By restoring a critical layer of semantic meaning, Varbo has the potential to accelerate reverse engineering efforts, reduce the cognitive load on analysts, and ultimately improve the overall effectiveness of security research and incident response. The public release of their datasets, models, and code further underscores the practical applicability and potential for widespread adoption of their findings.
Background
▶ Watch: Introduction: The problem of naming variables in decompiled code. (0:00)
Binary decompilation is the cornerstone of reverse engineering, serving as the bridge between opaque machine code and understandable high-level programming constructs. The process typically transforms low-level assembly instructions into pseudo-code resembling languages like C or C++. This transformation is essential for tasks such as malware analysis, where understanding the intricate logic of malicious binaries is paramount, or for vulnerability research, which often requires scrutinizing proprietary software without access to its original source code. Despite its critical role, decompilation faces a fundamental challenge: the lossy nature of the compilation process.
Compilers, in their quest for optimization and efficiency, discard significant source code artifacts. These include original data types, descriptive function names, and, most notably for this work, meaningful variable names. Instead, decompilers rely on a set of heuristics to reconstruct these lost elements, frequently resulting in generic, often sequential, placeholder names like A1, A2, V3, or V4. As the talk emphasizes, replacing a name like A1 with input or A3 with output instantly transforms unintelligible code into something intuitive and readable. The semantic meaning embedded in well-chosen variable names is crucial for human comprehension, reflecting the developer's intent and making the code's purpose immediately apparent.
Prior attempts to automate variable name prediction in decompiled code have shown promise but often fall short in generating truly meaningful names. The speakers pointed out a common failing: one of the most frequently predicted variable names by existing systems is V2, which is as uninformative as the decompiler's original output. This highlights the difficulty in capturing the nuanced, human-centric semantics of variable naming. A significant hurdle in developing effective prediction models is the scarcity of large, diverse datasets of decompiled code paired with their original source code annotations. While source code datasets are relatively easy to acquire, creating a parallel, high-quality decompiled corpus is challenging due to the complexities of the decompilation process and the need for accurate mapping between source and binary artifacts.
Recognizing these challenges, the researchers were inspired to employ transfer learning. Although source code and decompiled code share underlying semantics, they are visually distinct. Source code typically includes richer structural information, comments, and explicit type declarations, while decompiled code is often flatter, more verbose, and heavily reliant on compiler/decompiler-specific artifacts. Transfer learning offers a powerful paradigm to bridge this domain gap, allowing a model to first learn general programming language patterns from abundant source code and then adapt this knowledge to the unique characteristics of decompiled output, thereby tackling the "naming things" problem more effectively.
Key Findings
▶ Watch: VarBERT's two-step training pipeline: pre-training on source, fine-tuning on ... (3:52)
Varbo demonstrates significant advancements in predicting meaningful variable names and their origins in decompiled code, outperforming existing state-of-the-art methods. The core findings highlight the efficacy of their transfer learning approach and the impact of meticulous data preparation.
Firstly, Varbo achieved an impressive 50% accuracy in variable name prediction when evaluated on a common set of binaries against established approaches like Dire and Dirty. This represents a substantial improvement, being 12% higher than Dirty and 14% higher than Dire. This quantitative leap underscores Varbo's superior ability to map decompiled variables to their original, semantically relevant source code names.
A critical finding was the profound impact of pre-training. When Varbo was solely fine-tuned on decompiled code without an initial pre-training phase on source code, its accuracy for variable name prediction was only 39%. However, by first pre-training the model with source code and then fine-tuning it with decompiled code, the accuracy soared to 54%. This 15% improvement (54% vs 39%) unequivocally demonstrates the significance of leveraging source code knowledge to familiarize the model with the nature of programming language constructs before specializing it for the decompilation domain. This validates the core premise of their transfer learning methodology.
The research also investigated the relationship between training data volume and model performance. By training Varbo on varying subsets of the VAR Corpus (20%, 40%, etc.), they observed a consistent trend: as the training data increased, so did the model's accuracy. For instance, with just 20% of the training dataset, Varbo achieved 41% accuracy for variable name prediction and 77% for variable origin prediction. This finding emphasizes the importance of building larger and more diverse datasets for further enhancing model performance, highlighting a clear path for future research and development.
Furthermore, Varbo demonstrated strong generalizability across different binary architectures. To assess this, the researchers created a new dataset using only A64 binaries (ARM 64-bit) for O2 compiler optimization, a stark contrast to the predominantly x86-64 binaries in their main dataset. Despite this A64 dataset being significantly smaller (approximately 150,000 functions), Varbo achieved an accuracy of 46% on IDA and 50% on Ghidra. These results are comparable to models trained on similar numbers of x86-64 decompiled functions, indicating that Varbo's approach is not architecture-specific and can effectively generalize to different CPU instruction sets.
Finally, Varbo not only predicts meaningful variable names but also identifies their variable origin—distinguishing between human-created variables and extraneous variables introduced by compilers or decompilers. This dual prediction capability adds another layer of semantic understanding, allowing analysts to focus on variables that directly correspond to developer intent. The talk also highlighted an interesting aspect of "mispredictions": sometimes, Varbo predicts a name that, while not an exact character match, is semantically correct (e.g., predicting Enable when the original was enable). This suggests that a strict 100% character match metric might be too rigid and that the model often captures the underlying meaning even when capitalization or minor variations occur.
Technical Deep Dive
▶ Watch: Introducing VAR Corpus: A large, publicly available decompiled code dataset. (4:54)
Varbo's architecture and training pipeline are meticulously designed to address the challenges of variable name prediction in decompiled code, leveraging a BERT-based model and a two-step transfer learning process. The system comprises two main components: model training and data set building.
At its core, Varbo employs a two-step training process: pre-training and fine-tuning.
- Pre-training: The initial phase involves training the model on a large corpus of source code functions. Each function is converted into a token stream, which serves as input to the pre-training model. During this phase, Varbo utilizes Masked Language Modeling (MLM). MLM is a technique where a certain percentage of tokens in the input sequence are masked, and the model is tasked with predicting the original masked tokens based on their context. This process familiarizes the model with the syntax, structure, and general semantic patterns inherent in source code, effectively teaching it the "language" of programming without specific variable naming targets yet.
- Fine-tuning: Following pre-training, the model undergoes fine-tuning using decompiled code generated by tools like IDA Pro and Ghidra. Similar to pre-training, decompiled functions are converted into token streams and fed into the model. In this stage, Varbo learns to predict two distinct outcomes for each variable: its variable name (e.g.,
keyforA2) and its variable origin (e.g.,human-createdorextraneous). This fine-tuning adapts the general language understanding gained from source code to the specific nuances and artifacts present in decompiled output.
To support this training, the researchers built their own comprehensive dataset called VAR Corpus. This corpus is a collection of decompiled code derived from 112,000 C/C++ libraries and executables. These binaries were compiled with four different compiler optimization levels (O0 through O3), ensuring a diverse range of decompilation outputs. For IDA Pro, the O0 optimization level alone yielded 2.6 million unique functions. A significant contribution of this work is the public release of all research artifacts, including the source code, binaries with debug symbols, and the annotated decompiled code, facilitating reproducibility and further research.
The development of VAR Corpus and the training process encountered several critical challenges, each addressed with innovative solutions:
- Extraneous Variables: Decompiled code often contains variables that have no direct correspondence in the original source code. These extraneous variables are artifacts introduced by the compiler (e.g., for temporary storage or register allocation) or the decompiler itself. For example, a
V4in decompiled code might not map to any developer-assigned name. To tackle this, Varbo's training strategy focuses only on meaningful developer-assigned variable names for the name prediction task. For the extraneous variables, Varbo predicts their origin asextraneous, while variables with a source code counterpart are labeledhuman-created. This distinction allows the model to provide valuable context without attempting to invent names for variables that never had one.
- Data Types Impact: The presence or absence of debugging information, particularly data types, significantly alters the decompilation output. A binary compiled with debug symbols (e.g., DWARF) will produce different decompiled code compared to a stripped binary. To standardize this and improve variable matching, the researchers implemented a type-stripping process. They rewrite the DWARF debugging information to include only variable names, replacing all specific data types with generic scalar types of the same size. For instance, a
struct MyStruct*might become apointer-sized word. This process, referred to as creating typ-strip decompiled code, makes decompiled code with debug symbols structurally similar to that of stripped binaries, thereby enhancing the consistency and accuracy of variable mapping between the two.
- Function Duplication: Building a large dataset from public repositories like GitHub inevitably leads to function duplication. This arises from multiple forks of the same project (e.g., the Linux kernel repository has tens of thousands of forks) and the widespread practice of copying utility functions across different codebases. Duplicate functions can bias the training process and lead to inflated performance metrics. To mitigate this, a deduplication pipeline was implemented. First, functions are normalized by removing non-essential elements like whitespace, newlines, and decompiler-generated comments. Next, the normalized function bodies are hashed, and any exact hash matches are removed. This ensures that the training and test sets contain genuinely unique functions, providing a more robust evaluation of Varbo's generalization capabilities.
By meticulously addressing these technical challenges, Varbo provides a robust and effective solution for the complex problem of variable name prediction in decompiled binaries, significantly improving the quality and interpretability of reverse engineering output.
Demo / Proof of Concept
▶ Watch: VarBERT's superior performance: 50% accuracy, outperforming existing methods. (9:00)
While the talk did not feature a live, interactive demonstration of Varbo in action, it effectively presented a compelling case study to illustrate the model's capabilities and the quality of its predictions. This case study served as a proof of concept, showcasing how Varbo transforms obscure decompiler-generated names into meaningful identifiers.
The example presented involved a segment of source code containing three distinct variable names. The decompiler would typically output generic placeholders for these. Varbo's prediction for this specific case was highlighted:
- For one variable, Varbo correctly predicted
clock. - For another, it correctly predicted
old. - The third variable, originally named
enable, was predicted asEnable.
The speakers emphasized that a variable name prediction is considered "correct" only if there is a 100% character match between the original source code name and the predicted name. Under this strict definition, Varbo correctly predicted two out of three variable names in this instance. However, the misprediction of enable as Enable was a noteworthy observation. Despite not being an exact character match due to capitalization, the predicted name Enable is semantically correct and conveys the same meaning as the original. This observation suggests that while strict accuracy metrics are important, Varbo often captures the conceptual intent of a variable, even if minor stylistic differences exist. This highlights the practical utility of Varbo's output, as even "semantically correct" mispredictions significantly enhance code readability compared to generic placeholders like A1 or V4.
The researchers further solidified the practical applicability of their work by stating that they have open-sourced all their research artifacts, including the datasets, trained models, and the underlying code. This open-source release means that other researchers and practitioners can directly integrate Varbo models into their preferred decompilers (like IDA Pro or Ghidra) to obtain prediction results. This commitment to open science effectively serves as a broader proof of concept, enabling widespread adoption and further development of their methodology within the reverse engineering and security communities.
Defensive Implications
▶ Watch: Ablation study: Significance of pre-training for a 14% accuracy gain. (10:00)
The advancements brought forth by Varbo have significant and far-reaching implications for defensive security operations, enhancing the capabilities of analysts across various domains. By providing meaningful variable names in decompiled code, Varbo directly addresses a critical bottleneck in understanding and analyzing binaries when source code is unavailable.
Firstly, for malware analysts, Varbo can dramatically reduce the time and cognitive load required to understand complex and often obfuscated malicious binaries. Instead of deciphering the purpose of V1, A2, or arg_0, analysts can immediately grasp the intent behind variables named input_buffer, decryption_key, or network_socket. This accelerates the process of identifying malicious functionalities, understanding command-and-control protocols, and extracting indicators of compromise (IOCs), leading to faster threat intelligence generation and incident response.
Secondly, in vulnerability research, the ability to quickly comprehend decompiled code is paramount. Varbo can aid security researchers in identifying potential vulnerabilities by making it easier to trace data flows, understand function parameters, and spot logical flaws that might lead to buffer overflows, format string bugs, or unhandled exceptions. This is particularly valuable when auditing proprietary software or closed-source components for which no source code is provided, allowing for more efficient and thorough security assessments.
Thirdly, for organizations dealing with legacy codebases or systems where original developers are no longer available, Varbo offers a pathway to better understanding and maintaining these critical assets. Decompiled code, augmented with meaningful variable names, becomes more accessible for security patching, feature updates, or migration efforts, reducing the risk associated with unmanaged technical debt.
Furthermore, the integration of Varbo into existing reverse engineering tools like IDA Pro and Ghidra can transform the workflow of security professionals. Automated, intelligent variable naming can free up analysts from tedious manual renaming tasks, allowing them to focus on higher-level analysis and problem-solving. This not only improves efficiency but also potentially reduces human error in interpretation.
Finally, Varbo's capability to predict variable origin (human-created vs. extraneous) provides an additional layer of crucial context. Defenders can prioritize their analysis on variables that directly reflect developer intent, rather than compiler-generated artifacts, making their investigations more targeted and productive. This distinction helps in filtering noise and focusing on the core logic of the application. The open-sourcing of Varbo's models and datasets empowers the broader security community to integrate this technology into custom tools and pipelines, fostering innovation in automated binary analysis and defense.
Key Takeaways
- Addressing a Fundamental Challenge: Varbo directly tackles the "naming things" problem in reverse engineering, transforming obscure decompiler-generated variable names (e.g.,
A1,V4) into meaningful, human-readable identifiers. - Efficacy of Transfer Learning: The research demonstrates that a two-step transfer learning approach—pre-training on source code using Masked Language Modeling and then fine-tuning on decompiled code—is highly effective, yielding a 15% accuracy improvement over fine-tuning alone.
- Superior Performance: Varbo significantly outperforms existing variable name prediction systems, achieving 50% accuracy, which is 12-14% higher than prior work like Dirty and Dire, establishing a new state-of-the-art.
- Dual Prediction Capability: Beyond predicting variable names, Varbo also identifies the "origin" of variables, distinguishing between human-created variables and extraneous variables introduced by compilers or decompilers, providing crucial context for analysts.
- Robustness and Generalizability: The model's performance improves with larger datasets and demonstrates strong generalizability across different compiler optimization levels (O0-O3) and CPU architectures (x86-64 and A64), making it broadly applicable.
- Open-Sourced for Impact: The researchers have publicly released their VAR Corpus datasets, trained models, and code, enabling direct integration into existing decompilers and fostering further research and practical application within the security community.
About the Speaker(s)
The primary presenter for this talk was Ati Priya Bajaj, a PhD student at Arizona State University. Her work, as highlighted in this presentation, focuses on critical challenges in computer security, particularly in the realm of binary analysis and reverse engineering through the application of machine learning techniques. Ati Priya Bajaj is part of a collaborative research effort, with co-authors including Kuntal Kumar Pal, Pratyay Banerjee, Audrey Dutcher, Mutsumi Nakamura, and Zion Leonahenahe Basque. Their collective expertise contributes to addressing complex problems like variable name prediction, which significantly impacts the efficiency and effectiveness of security professionals and researchers. Their contribution to the IEEE S&P conference underscores their commitment to advancing the state of the art in security research.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This research delivers a significant breakthrough in tackling the pervasive 'naming things' problem in reverse engineering. Varbo's transfer learning approach, backed by meticulous data engineering, dramatically improves the interpretability of decompiled code by predicting meaningful variable names, making it an essential tool for any serious analyst.
Heather Calloway (CISO) — STRONG ACCEPT
This research presents a significant leap in binary analysis, directly enhancing the efficiency and accuracy of reverse engineering efforts. Varbo's transfer learning approach provides meaningful variable names, reducing the cognitive load on security analysts and accelerating incident response and vulnerability research. While not directly informing board-level governance, its operational impact is undeniable for any security program dealing with compiled code.
→ Top-rated talks at IEEE Symposium on Security and Privacy 2024