BLens: Contrastive Captioning of Binary Functions using Ensemble Embedding

Tristan Benoit

34th USENIX Security Symposium (USENIX Security '25) · Day 3 · Software Security 4: Fuzzing and Other Software Analysis

Overview

In the realm of reverse engineering, analyzing stripped binaries presents a formidable challenge. Without symbolic information, functions are often assigned generic, meaningless names, forcing engineers to meticulously inspect assembly or pseudocode to deduce their purpose. This process is not only time-consuming and tedious but also highly prone to errors, especially in large and complex software projects. The talk "BLens: Contrastive Captioning of Binary Functions using Ensemble Embedding," presented by Tristan Benoit, introduces a groundbreaking AI-driven solution to this pervasive problem.

Watch on YouTube · Slides

Visual summary for BLens: Contrastive Captioning of Binary Functions using Ensemble Embedding by Tristan Benoit
Visual summary for BLens: Contrastive Captioning of Binary Functions using Ensemble Embedding by Tristan Benoit

Key moments

  1. 0:00 Introduction and problem of stripped binary functions
  2. 1:15 DeepSec example: Imprecise predictions mislead reverse engineers
  3. 2:00 Challenge 2: Recommender-style models lack name ordering
  4. 3:15 Challenge 3: Translator models struggle with unseen patterns
  5. 4:00 BLens reframes problem as image captioning for generalization
  6. 4:50 Stage 1: COMBO ensemble encoder for unified representation
  7. 6:10 Contrastive learning aligns diverse binary representations
  8. 7:45 Stage 2: LORD decoder uses masked language modeling

BLens: Contrastive Captioning of Binary Functions using Ensemble Embedding

Speakers: Tristan Benoit

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=MEFKVmCmEvQ

Overview

In the realm of reverse engineering, analyzing stripped binaries presents a formidable challenge. Without symbolic information, functions are often assigned generic, meaningless names, forcing engineers to meticulously inspect assembly or pseudocode to deduce their purpose. This process is not only time-consuming and tedious but also highly prone to errors, especially in large and complex software projects. The talk "BLens: Contrastive Captioning of Binary Functions using Ensemble Embedding," presented by Tristan Benoit, introduces a groundbreaking AI-driven solution to this pervasive problem.

BLens redefines automated function name prediction by treating it as an "image captioning" task, leveraging a sophisticated multimodal learning framework. This novel approach aims to generate precise, semantically rich, and contextually accurate function names directly from binary code. By addressing the critical shortcomings of previous methods—namely, a lack of precision, an inability to preserve word order, and poor generalization to unseen patterns—BLens significantly streamlines the reverse engineering workflow, making it more efficient and less error-prone. The project is fully open-sourced and has earned all three artifact evaluation badges, underscoring its reproducibility and practical utility for the security community.

Background

▶ Watch: Introduction and problem of stripped binary functions (0:00)

The core problem BLens addresses is the inherent difficulty in understanding the functionality of stripped binaries. When symbols are removed, reverse engineers are confronted with a list of functions bearing placeholder names like sub_401000, which convey no information about their actual operations. Manually analyzing each function's assembly or pseudocode to infer its purpose is an exhaustive and often inefficient endeavor, particularly for large codebases. This bottleneck significantly impedes tasks such as malware analysis, vulnerability research, and software auditing.

Prior attempts to automate function name prediction using learning-based methods generally fall into two categories: recommender-style approaches and translator-style approaches.

Recommender-style approaches, exemplified by tools like XFO, treat function name prediction as a multi-label classification or recommendation problem. In this paradigm, each token within a function name is considered an independent label. The model decides on the presence or absence of specific tokens (e.g., to, int, str) without explicitly considering their order or contextual relationships. A post-processing step then attempts to order these tokens based on natural language patterns. However, this decoupling from the function's semantics often leads to inaccuracies. For instance, such a model might predict "string to int" when the ground truth is "int to string," thereby losing crucial semantic information. This highlights the second key challenge identified by the BLens researchers: without grounding the ordering in the actual code, important semantics are easily lost.

Translator-style approaches, including models like SIML LM, ASM depictor, and hext5, frame function name prediction as a translation task. These models excel at understanding context and word order, often producing fluent and accurate names when the underlying code patterns are familiar from their training data. However, their primary weakness lies in their sensitivity to unseen or "unseen" name patterns, especially under domain shift. For example, if a model is trained on functions like print_integer and int_to_float, but then encounters a new, structurally similar function print_float, it might default to predicting print_int due to its familiarity with the "print int" pattern, failing to generalize to the new but valid "print float" pattern. This represents the third challenge: the failure to generalize effectively to novel yet semantically valid patterns.

To overcome these limitations, BLens reframes the problem of function name prediction as an image captioning task. This innovative perspective allows the model to inherently preserve word order, mitigating the issues faced by recommender-style models. Furthermore, by drawing parallels with how image captioning models generalize to describe novel visual scenes, BLens aims to improve generalization capabilities in "zero-shot" scenarios, where functions with previously unseen naming patterns are encountered. This foundational shift in problem formulation underpins BLens' superior performance and enhanced robustness.

Key Findings

▶ Watch: Challenge 2: Recommender-style models lack name ordering (2:00)

BLens introduces a significant advancement in automated binary function naming, demonstrating state-of-the-art performance across various evaluation metrics and challenging scenarios. The key findings from the research highlight its precision, generalization capabilities, and the effectiveness of its core architectural components:

  • Superior Performance: BLens consistently outperforms all baseline approaches in both cross-binary and cross-project settings. In the more challenging cross-project setting, which better mirrors real-world reverse engineering scenarios, BLens achieves even larger margins of improvement across all metrics. This underscores its robust generalization ability to entirely different software projects.
  • High Precision: The model achieves an impressive 92% precision in function name prediction. This high precision is critical for reverse engineers, as it minimizes the need for extensive corrections and reduces the risk of misleading interpretations, directly addressing the first challenge of prediction precision.
  • Enhanced Fluency and Generalization: BLens shows substantial improvements in fluency metrics like BLEU and Fuji, which account for subsequence overlap, and especially in BiR, which captures semantic similarity. This indicates that BLens not only predicts accurate tokens but also arranges them in a semantically coherent and contextually appropriate manner, tackling the challenges of preserving word order and generalizing to unseen patterns.
  • Effectiveness of Combo Pre-training: Ablation studies confirm the strong contribution of the Combo pre-training stage to the final performance. A variant of BLens trained without Combo performed significantly worse across all metrics, validating Combo's role in building a unified, rich representation of binary functions by fusing diverse binary features.
  • Optimized Lord Decoder: The Lord (Likelihood Order Regressive Decoder) plays a crucial role in achieving high precision. While there's a slight trade-off in recall compared to some variants, Lord delivers the best overall F1 score. Its tailored masked language modeling task and flexible auto-regression strategy are key to its ability to generate high-quality, precise names.
  • Beyond Metric Limitations: Case studies presented in the talk demonstrate that BLens can generate high-quality predictions that are semantically correct and align with common naming conventions, even when standard evaluation metrics might mark them as incorrect. For instance, predicting "shell" for a function named "execute" or "on button cancel clicked" for "task panel cancel clicked CBK" showcases BLens' ability to infer broader meaning and formalize names, highlighting the limitations of current automated evaluation metrics in fully capturing the model's intelligence.
  • Open-Source and Reproducible: The entire project is open-sourced and has earned all three artifact evaluation badges, ensuring that its results are reproducible and the methodology can be adopted or extended by the community.

Technical Deep Dive

▶ Watch: BLens reframes problem as image captioning for generalization (4:00)

BLens' innovative architecture is divided into two primary stages: pre-training with a multimodal text encoder named Combo, and finetuning with a specialized function name decoder called Lord. This two-stage approach allows BLens to first build a rich, unified understanding of binary functions and then specialize this understanding for the precise task of name prediction.

Stage 1: Combo – The Pre-training Stage

The goal of the Combo (short for Combined Multimodal Binary Representation) pre-training stage is to construct a unified representation that comprehensively captures various facets of binary functions. This is achieved through an ensemble encoder that fuses three complementary forms of binary representations:

  1. Palm Tree: An assembly language model that focuses on structural patterns within the code, typically generating sequences of tokens.
  2. Dexter: A statistical-based embedding that provides a high-level, single embedding for each function.
  3. Clap: A semantic-oriented embedding, also providing a single embedding, which emphasizes the functional meaning of the code.

The challenge lies in integrating these diverse representations, as Palm Tree produces sequences while Dexter and Clap output single embeddings. Inspired by Vision Transformers, BLens converts each representation into a consistent format of "patches" using Multi-Layer Perceptron (MLP) layers:

  • For Palm Tree, the assembly sequence is split into patches at the basic block level, preserving its inherent structure.
  • For Dexter and Clap, their single embeddings are projected into multiple patches, effectively creating a sequence-like representation from a single vector.

After patch conversion, positional encodings are added separately to each form. This is crucial because the number of patches can differ across the representations, and positional encodings help the model understand the relative order and location of information within each form. These three sets of patches are then concatenated to form a comprehensive representation for each function.

The fusion of these diverse representations is driven by contrastive learning. The motivation here is that while Palm Tree, Dexter, and Clap all describe the same function, their raw embeddings are not naturally aligned in a shared semantic space. Contrastive learning addresses this by aligning the generated patches with tokenized function name labels. Patches related to the same labels are pulled closer together in the latent space, while unrelated ones are pushed apart. This process ensures that the unified embedding space is semantically meaningful and consistent.

To further enhance the representation, a captioning head is added on top of the fused representation during pre-training. This head directly predicts the function name, which strongly encourages the ensemble encoder to capture both semantic and structural information relevant for name generation. In essence, Combo transforms disparate binary representations into a robust, unified embedding space, ready for the specialized finetuning stage.

Stage 2: Lord – The Finetuning Stage

In the finetuning stage, BLens employs a novel function name decoder named Lord (Likelihood Order Regressive Decoder). Lord is designed to predict function names from the rich binary representations generated by Combo.

Traditionally, generative transformers often rely on teacher forcing, where each token in a sequence is predicted based on the ground truth of the previous tokens. While effective for long sequences, this approach can weaken the model's ability to plan for shorter sequences, like function names, making later tokens artificially easier to predict.

To counter this, Lord adopts a tailored masked language modeling (MLM) task for finetuning. Instead of predicting the next token, the decoder predicts randomly masked tokens from the remaining, unmasked ones. For example, if the function name is convert_hex_to_int, and hex_to is masked, the model attempts to predict these masked words using the context provided by convert_ and _int.

A critical adjustment for function names, which are significantly shorter than natural language sentences, is the masking probability scheme. For a name with n tokens, the number of masked tokens m is sampled from a distribution where sub_i = 1 + i / n for i from 0 to n, normalized by softmax. This scheme biases the masking towards selecting more tokens, with full masking being the most likely scenario. By forcing the model to predict a larger proportion of the name based on limited context, this approach reduces the bias of teacher forcing and encourages a more robust understanding of the entire name's structure and semantics.

Inference Stage: Flexible Auto-regression

During inference, the overall architecture remains similar to the finetuning stage, but BLens introduces a dedicated decoding strategy called flexible auto-regression to maximize the precision of function name prediction. Unlike traditional left-to-right decoding, Lord does not strictly adhere to a sequential generation process:

  1. Parallel Prediction: In the first step, Lord predicts tokens for all possible positions in parallel.
  2. Highest Confidence Selection: It then identifies the token with the highest confidence score across all positions. For example, it might select EOS (End-Of-Sequence) at position 5 with a confidence of 0.48. This token is then fixed. (Notably, to handle variable-length outputs, a dedicated EOS token is not used in the final output generation; this example illustrates the confidence-based selection process).
  3. Iterative Refinement: The process repeats. With the previously fixed token(s) as context, Lord again predicts tokens for the remaining unfixed positions and selects the next highest confidence token (e.g., int for position 3 with 0.35 confidence), fixing it in place.
  4. Threshold-Based Stopping: This iterative selection continues until the confidence of the highest remaining candidate token falls below a pre-set threshold. This threshold is optimized for the F1 score on a validation set. For instance, after predicting "convert hex to int," if the next best candidate token has a confidence of 0.25, which is below the threshold, the decoding process stops.

This flexible auto-regression strategy is designed to prevent the generation of low-confidence, potentially false-positive tokens. By stopping when confidence drops, Lord maintains high precision in its predictions, ensuring that the generated function names are reliable and accurate.

Demo / Proof of Concept

▶ Watch: Stage 1: COMBO ensemble encoder for unified representation (4:50)

While a live demonstration was not explicitly shown during the conference talk, the BLens project provides robust evidence of its functionality and utility. The speaker explicitly stated that the project is "fully open sourced" and has "earned all three artifact evaluation badges." This means the code, models, and evaluation scripts are publicly available, allowing researchers and practitioners to try it out, reproduce the published results, and even compare it against their own methods. This commitment to open science and reproducibility serves as a strong proof of concept.

Furthermore, the talk presented two compelling case studies that effectively illustrate BLens' capabilities and its advantages over traditional metrics:

  1. Semantic Equivalence: In the first case, the ground truth function name was execute, while BLens predicted shell. Standard metrics might mark this as incorrect due to the token mismatch. However, as highlighted by the speaker, execute and shell are clearly semantically related. This demonstrates BLens' ability to understand the underlying operation and provide a contextually relevant name, even if it's not an exact string match.
  2. Naming Convention Alignment: The second case involved a ground truth name task_panel_cancel_clicked_CBK. BLens predicted on_button_cancel_clicked. Here, BLens not only provided a semantically correct label but also one that aligns more closely with common naming conventions found in software development. This ability to generalize patterns and produce more formalized, human-readable names is a significant advantage, showcasing BLens' capacity to generate higher-quality predictions that transcend the limitations of simple string-based evaluation metrics.

These examples, combined with the open-source nature and artifact evaluation badges, collectively serve as a powerful demonstration of BLens' practical effectiveness and its potential to revolutionize binary analysis.

Defensive Implications

▶ Watch: Stage 2: LORD decoder uses masked language modeling (7:45)

The advent of BLens has profound implications for cybersecurity defenders across various domains, significantly enhancing their capabilities in understanding and responding to threats:

  1. Accelerated Malware Analysis: For malware analysts, understanding the purpose of functions in stripped malicious binaries is paramount. BLens can dramatically reduce the time spent on initial triage and deep dive analysis by automatically suggesting meaningful function names. This allows analysts to quickly identify critical functionalities such as network communication, encryption, file system manipulation, or process injection, leading to faster threat intelligence generation and more efficient reverse engineering of complex malware.
  2. Enhanced Vulnerability Research: Vulnerability researchers often face the daunting task of auditing large, stripped binaries (e.g., firmware, operating system components, third-party libraries). BLens can pinpoint functions of interest – such as parsers, cryptographic routines, memory allocation functions, or network handlers – with much greater speed and accuracy. This allows researchers to focus their efforts on high-risk areas, improving the efficiency of vulnerability discovery.
  3. Improved Software Supply Chain Security: With increasing concerns about supply chain attacks, organizations often need to analyze third-party components or proprietary software where source code is unavailable. BLens provides a powerful tool to gain insights into the functionalities embedded within these binaries, helping to identify potential backdoors, hidden features, or insecure implementations without requiring extensive manual effort.
  4. Augmented Automated Analysis Tools: BLens can be seamlessly integrated into existing static analysis pipelines. By enriching the metadata of stripped functions with accurate, AI-generated names, subsequent automated analysis steps (e.g., call graph analysis, data flow analysis, taint analysis) can become more effective and yield more actionable results. This can lead to more intelligent automated vulnerability scanning or anomaly detection systems.
  5. Better Incident Response: During an incident, rapid understanding of compromised systems and malicious artifacts is crucial. BLens can aid incident responders by providing quick insights into unknown binaries found on compromised hosts, helping to determine their purpose and potential impact, thus accelerating containment and remediation efforts.
  6. Knowledge Transfer and Training: For junior reverse engineers or those new to a specific codebase, BLens can serve as an invaluable learning aid. By providing initial function name suggestions, it lowers the barrier to entry for complex binary analysis, facilitating faster skill development and knowledge transfer within security teams.

While BLens offers significant advantages, defenders should still exercise human oversight. AI-generated names, though highly accurate, are still predictions and should be verified, especially in critical contexts. However, BLens provides an exceptionally strong starting point, transforming a previously labor-intensive and error-prone process into a significantly more efficient and accessible one.

Key Takeaways

  • BLens addresses the critical challenge of analyzing stripped binaries by automatically generating precise and semantically rich function names, significantly reducing manual reverse engineering effort.
  • The system reframes function name prediction as an "image captioning" problem, overcoming limitations of prior recommender and translator-style approaches regarding order preservation and generalization.
  • Its core architecture comprises a two-stage process: Combo for multimodal pre-training (fusing Palm Tree, Dexter, and Clap embeddings into a unified representation) and Lord for finetuning (a specialized decoder using tailored masked language modeling).
  • BLens consistently achieves state-of-the-art performance, demonstrating 92% precision and substantial improvements in fluency and generalization across both cross-binary and challenging cross-project evaluation settings.
  • The innovative flexible auto-regression decoding strategy in Lord ensures high precision by stopping prediction when token confidence falls below a threshold, preventing low-confidence, potentially false-positive outputs.
  • BLens is an open-source project that has earned all three artifact evaluation badges, promoting reproducibility and enabling the security community to leverage and extend its capabilities for malware analysis, vulnerability research, and software supply chain security.

About the Speaker(s)

Tristan Benoit is a researcher involved with the BLens project, a collaborative effort with colleagues from BVB, LMU Munich, and MCML. His work, as presented in this talk, focuses on advancing automated binary analysis techniques through novel applications of machine learning and deep learning, particularly in the realm of reverse engineering and function understanding.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Solid academic research that reframes binary function naming as image captioning — a genuinely clever framing that sidesteps real weaknesses in both recommender and translator approaches. The two-stage architecture (Combo + Lord) is well-motivated and technically differentiated, and 92% precision in cross-project settings is a number worth taking seriously. This is USENIX-grade work and it belongs there.

Heather Calloway (CISO) — WEAK

Technically credible research on automated binary function naming with real potential to accelerate malware analysis and vulnerability research. But the talk is aimed squarely at researchers and tool builders — it never bridges to the security leaders, program operators, or incident response teams who would actually authorize or integrate this capability.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)