BumbleBee: Secure Two-party Inference Framework for Large Transformers

Wen-jie Lu (Zjanu)

Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · Privacy & Cryptography 1 · Privacy & Cryptography 1

Overview

This article delves into "BumbleBee," a novel secure two-party inference framework designed for large transformer models, presented by Wen-jie Lu from Zjanu at the NDSS Symposium. The talk addresses a critical challenge in the era of pervasive AI services: how to leverage powerful machine learning models without compromising the privacy of sensitive input data. As individuals and organizations increasingly rely on cloud-based AI, the necessity to transmit confidential information to remote servers raises significant concerns about potential data exposure, misuse, or mismanagement.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction: Privacy issues in AI services
  2. 3:00 Observation: High communication cost in current approaches
  3. 4:00 Three key challenges in secure transformer inference
  4. 4:40 BumbleBee's end-to-end codebase and technical approach
  5. 6:00 Lesson 1: Compiling ML code for secure computation
  6. 7:00 Softmax optimization via middle layer IR rewriting
  7. 8:50 Softmax is not a bottleneck in BumbleBee's approach

BumbleBee: Secure Two-party Inference Framework for Large Transformers

Speakers: Wen-jie Lu (Zjanu)

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=d2O4gW3P6Zo

Overview

This article delves into "BumbleBee," a novel secure two-party inference framework designed for large transformer models, presented by Wen-jie Lu from Zjanu at the NDSS Symposium. The talk addresses a critical challenge in the era of pervasive AI services: how to leverage powerful machine learning models without compromising the privacy of sensitive input data. As individuals and organizations increasingly rely on cloud-based AI, the necessity to transmit confidential information to remote servers raises significant concerns about potential data exposure, misuse, or mismanagement.

BumbleBee proposes a robust solution by employing Secure Two-Party Computation (2PC), enabling two distinct parties—typically a client with private input and a server with a private model—to jointly compute a function, such as a transformer inference, without either party revealing their respective private data to the other. The framework distinguishes itself from many existing approaches by operating under a strict two-party setting, avoiding assumptions of trusted third parties or reliance on secure hardware enclaves like SGX. Its primary innovation lies in dramatically reducing the communication overhead, a notorious bottleneck in privacy-preserving machine learning, while maintaining computational efficiency and model accuracy.

The significance of BumbleBee cannot be overstated. By offering an end-to-end, compiler-driven solution that integrates seamlessly with existing machine learning ecosystems like JAX and Hugging Face, it lowers the barrier for adopting privacy-preserving AI. The framework's ability to minimize communication, often by two orders of magnitude compared to prior work, makes secure transformer inference a more practical reality for real-world applications where data privacy is paramount, such as in healthcare, finance, or government intelligence.

Background

▶ Watch: Introduction: Privacy issues in AI services (0:00)

The proliferation of AI services has introduced a paradoxical challenge: to benefit from advanced models, users must often relinquish control over their sensitive data, sending it to third-party servers. This practice carries inherent risks, including data breaches, unauthorized access, and compliance issues with stringent privacy regulations like GDPR and HIPAA. The core problem BumbleBee addresses is enabling private inference, where a client can obtain predictions from a server's model without revealing their input, and ideally, without the client learning the model's specifics.

Two primary cryptographic paradigms exist for private inference: Fully Homomorphic Encryption (FHE) and Secure Multi-Party Computation (MPC). While FHE allows computations on encrypted data, often with high computational overhead, MPC (of which 2PC is a specific instance for two parties) involves interactive protocols where parties collaboratively compute a function while keeping their inputs private. In a 2PC setting, two parties—a client and a server—each contribute private inputs. They engage in an interactive protocol where all messages exchanged are encrypted or secret-shared, ensuring that neither party learns the other's private information beyond the final computed result. A fundamental building block for many 2PC protocols is the efficient generation of correlated randomness, particularly for operations like multiplication and AND gates.

Prior research in secure transformer inference has explored various avenues. Many approaches leverage additive secret sharing, where inputs are split into shares distributed among parties, and computations are performed on these shares. Some frameworks, like Sigma, utilize functional secret sharing. These systems often rely on different 2PC libraries, leading to diverse security models, including those assuming a Trusted First Party Section (TFPS), a Trusted Dealer-Based (TDB) assumption, or pure two-party settings. A key distinction highlighted by the speaker is that many previous methods, particularly those in TFPS or TDB models, implicitly or explicitly rely on a trusted third party (or a secure hardware enclave like SGX) to generate and distribute correlated randomness. BumbleBee, in contrast, operates in a pure two-party setting, eliminating this trust assumption.

Existing frameworks also vary in their support for different transformer architectures; some are tailored for models like BERT, while others (e.g., Okra, Shaft) support PyTorch models. BumbleBee uniquely targets models written in JAX, a popular framework for high-performance numerical computing. Furthermore, many prior solutions necessitate significant model fine-tuning or component replacement to make transformers compatible with 2PC protocols, which can be burdensome for practitioners. BumbleBee aims to circumvent this by allowing direct use of models from repositories like Hugging Face without modification.

A critical observation driving BumbleBee's design is the tremendous communication cost associated with current privacy-preserving transformer inference methods. The speaker notes that some approaches can incur hundreds of gigabytes of communication for merely a single token inference. This cost is often orders of magnitude higher than the computational expense, making these solutions impractical for real-world deployment, especially considering the on-demand pricing structures of cloud providers where communication bandwidth is a significant cost factor. BumbleBee's overarching goal is to minimize this communication overhead, aiming for a 100x reduction, while ensuring that the introduced computation cost remains acceptably low, ideally even slightly shorter in some cases.

To achieve this, BumbleBee confronts three primary technical challenges:

  1. Nonlinear Activation Functions: Designing efficient 2PC protocols for complex nonlinear activation functions prevalent in transformers, such as GELU (Gaussian Error Linear Unit) and Softmax, without resorting to 2PC-unfriendly alternatives that compromise accuracy or introducing computationally expensive numerical approximations that inflate communication.
  2. Engineering Methodology: Bridging the gap between machine learning engineers and cryptographers. The challenge is whether machine learners should write cryptographic code or cryptographers should write machine learning code, aiming for a seamless integration that leverages existing ML ecosystems.
  3. Large-Scale Matrix Multiplication: Developing efficient 2PC protocols for the large-scale matrix multiplications that form the computational core of transformer architectures.

Key Findings

▶ Watch: Three key challenges in secure transformer inference (4:00)

BumbleBee introduces several significant contributions and insights that advance the state of secure two-party transformer inference:

First, the project delivers an end-to-end codebase for private transformer inference. This implementation allows users to directly download JAX-based transformer models from platforms like Hugging Face and execute them within BumbleBee's secure framework without requiring any model fine-tuning or modifications. This significantly lowers the barrier to adoption for ML practitioners seeking to integrate privacy into their workflows.

Second, the framework incorporates an approach for matrix multiplication using homomorphic encryption. While the talk notes this is a concurrent work (detailed in their paper) and primarily designed for single multiplications rather than the multi-layered operations in transformer blocks, it highlights the exploration of diverse cryptographic primitives to optimize specific computational bottlenecks.

Third, and perhaps most crucially, the development of BumbleBee yielded three key lessons that informed its design and performance:

  1. Leveraging ML Ecosystems via Compilers: The most effective strategy is to allow machine learning engineers to write standard ML code and then use a specialized compiler to translate this code into a secure computation backend. This avoids the arduous task of rewriting complex transformer architectures from scratch in cryptographic primitives. BumbleBee, like Shaft and Sigma, adopts this compiler-based approach, but critically, it utilizes a middle-layer Intermediate Representation (IR). This IR allows for graph-level optimizations and rewrites before execution, a superior method compared to Python-level overriding (e.g., in Shaft), which operates on eager execution. This IR-based optimization proved instrumental in significantly improving the efficiency of operations like Softmax.
  1. Prioritizing Simple, Numerically Stable Approximations: Through extensive experimentation, BumbleBee found that simple piecewise approximations are often the most suitable for nonlinear functions like GELU, primarily due to their superior numerical stability. More complex numerical methods, such as approximating hyperbolic tangents or high-degree Fourier series (which could involve 18 multiplications for a degree-4 series), often introduce prohibitive costs, require input clipping, or suffer from stability issues within a 2PC context. The framework optimizes piecewise comparisons using a "one versus many" technique for efficiency.
  1. Enhancing Numerical Stability for Softmax: For the Softmax function, a critical component of transformers, BumbleBee employs a standard but crucial technique: computing the maximum value of the input vector, subtracting it from all elements, and then applying the exponential function. This ensures that the inputs to the exponential function are negative and bounded, allowing for highly accurate approximations with a low-degree Taylor series expansion (e.g., degree 3). This approach is more robust than "division-free and maximum-free" numerical methods, which often impose strict input range requirements that are difficult and expensive to enforce securely. This careful handling ensures Softmax, often considered a bottleneck, consumes less than 1% of the total inference cost in BumbleBee.

Technical Deep Dive

▶ Watch: BumbleBee's end-to-end codebase and technical approach (4:40)

BumbleBee's technical prowess stems from its strategic design choices, particularly concerning its compiler architecture and the handling of nonlinear activations. The framework builds upon the principles of Secure Two-Party Computation (2PC), where two parties—a client (with private input x) and a server (with private model f)—collaboratively compute f(x) such that neither party learns the other's private information. This is achieved through techniques like additive secret sharing, where x is split into x_0 + x_1 (with x_0 held by one party and x_1 by the other), and cryptographic protocols for operations on these shares.

The core innovation in BumbleBee's engineering methodology lies in its compiler-based approach utilizing a middle-layer Intermediate Representation (IR). The speaker emphasizes that large transformer models are inherently complex, making it impractical to rewrite them from scratch using cryptographic primitives. Instead, BumbleBee allows machine learning engineers to write standard JAX code, which is then translated into an IR. This IR represents the computation graph of the model. Unlike Python-level overriding (e.g., in Shaft), where cryptographic operations replace standard ones during eager execution, BumbleBee's IR allows for graph-level analysis and optimization before the secure computation protocol begins.

A prime example of this optimization is demonstrated with the Softmax function. In common Python/JAX implementations, Softmax involves computing exp(x) / sum(exp(x)). A frequent pattern observed in ML code, especially within frameworks like Hugging Face, is the use of keep_dimension=True during division operations. While seemingly innocuous for aligning operand shapes in cleartext, in a 2PC setting, this expands a scalar divisor into a large vector of identical values. Consequently, a single scalar division (which could be optimized by computing a single inverse and then performing multiplications) becomes a vector of expensive, secure divisions. The speaker notes this can increase the cost of division by 10 to 100 times.

BumbleBee's IR-based compiler identifies this pattern. It can rewrite the computation graph to replace the vector division with a single secure inverse operation followed by a vector of secure multiplications. This transformation significantly reduces the overhead associated with Softmax, making it far less of a bottleneck. In fact, after this optimization, Softmax operations account for less than 1% of the total computational cost in BumbleBee's private transformer inference. This directly refutes previous work that often identified Softmax as a primary performance bottleneck.

For nonlinear activation functions, BumbleBee prioritizes numerical stability and efficiency.

  • GELU (Gaussian Error Linear Unit): The speaker notes that GELU, when plotted, resembles a linear function over most of its domain, with a slight curve near the origin. BumbleBee found that a simple piecewise approximation is most effective. This involves dividing the input range into segments and approximating the function with simpler polynomials (e.g., linear segments) within each. A key optimization for this approach in 2PC is the efficient evaluation of comparisons (to determine which segment an input falls into). BumbleBee employs a one-versus-many comparison technique, batching multiple comparisons for a single input against various thresholds, which significantly improves efficiency compared to individual secure comparisons. This contrasts with attempts to approximate functions like hyperbolic tangent, which often require expensive input clipping or division operations, or high-degree Fourier series, which introduce a large number of secure multiplications (e.g., 18 for a degree-4 series).
  • Softmax (Numerical Stability): Beyond the IR-level optimization, BumbleBee addresses the numerical stability of Softmax. The standard practice of computing max(x), then x - max(x), before applying exp() is crucial. This ensures that the inputs to the exponential function are always negative and bounded. Consequently, exp() can be accurately approximated using a low-degree Taylor series expansion (e.g., degree 3) without significant loss of precision. For extremely negative inputs, the result of exp() approaches zero, allowing for further clipping to zero. The speaker cautions against "fancy" division-free or maximum-free numerical methods, which, while theoretically appealing, often require strict input range constraints. Enforcing these constraints securely (via range clipping) typically introduces more overhead than the benefits gained from avoiding division or maximum operations in the first place, making the simpler, bounded Taylor series approach more practical.

While the talk briefly mentions a concurrent work on matrix multiplication using homomorphic encryption for single operations, the primary focus of BumbleBee's detailed discussion remains on the compiler-based optimizations and the efficient, stable handling of nonlinear activations. The overarching theme is to minimize the communication cost, which is acknowledged as the most expensive component in cloud-based 2PC, by carefully designing protocols and leveraging compiler optimizations that reduce the number of interactive rounds and the total data exchanged.

Demo / Proof of Concept

▶ Watch: Softmax optimization via middle layer IR rewriting (7:00)

While the presentation did not feature a live, interactive demonstration, the speaker explicitly stated that BumbleBee provides an end-to-end codebase. This implementation allows users to directly download JAX-based transformer models from Hugging Face and execute them within the secure framework without requiring any model fine-tuning or modifications. This implies a fully functional and verifiable implementation available to researchers and practitioners, underscoring the practical readiness of the BumbleBee framework.

Defensive Implications

▶ Watch: Softmax is not a bottleneck in BumbleBee's approach (8:50)

BumbleBee offers profound defensive implications for organizations and individuals concerned about data privacy in the age of AI. The primary benefit is enabling privacy-preserving AI inference, allowing sensitive data to be processed by powerful machine learning models without direct exposure to the model owner or the underlying cloud infrastructure.

  1. Enhanced Data Privacy: Organizations handling sensitive user data (e.g., in healthcare, finance, legal, or defense sectors) can leverage BumbleBee to perform inference on that data using third-party AI models without risking its confidentiality. This ensures that personal identifiable information (PII), proprietary business data, or classified information remains private, even when interacting with external AI services.
  2. Regulatory Compliance: The framework directly addresses requirements for data protection regulations such as GDPR, HIPAA, CCPA, and others. By ensuring that raw input data never leaves the client's control in an unencrypted form, organizations can demonstrate a stronger commitment to privacy by design and reduce their legal and reputational risks associated with data breaches.
  3. Reduced Trust Assumptions: Unlike solutions that rely on trusted third parties, secure hardware enclaves (like SGX), or specific cloud provider features, BumbleBee operates in a pure two-party setting. This minimizes the attack surface by eliminating the need to trust an additional entity or hardware component, strengthening the overall security posture.
  4. Practical Adoption for ML Practitioners: The compiler-based approach, which allows direct use of existing JAX models from Hugging Face without fine-tuning, significantly lowers the barrier to entry for machine learning engineers. This ease of integration means that privacy-preserving techniques are no longer niche cryptographic exercises but can be more readily incorporated into standard ML development workflows.
  5. Cost-Effective Privacy: By drastically reducing communication overhead (claimed 100x reduction), BumbleBee makes privacy-preserving inference more economically viable. High communication costs in cloud environments can be a major deterrent; by mitigating this, BumbleBee opens the door for broader adoption of secure inference in production environments.
  6. Mitigating Model Inversion Attacks: While not explicitly detailed in the talk, secure inference frameworks like BumbleBee inherently provide a degree of protection against model inversion attacks or membership inference attacks. By preventing the server from directly observing individual inputs, it becomes harder for an adversary to reconstruct private training data or determine if a specific data point was part of the training set.

In essence, BumbleBee empowers defenders to embrace the power of large language models and transformers without sacrificing the fundamental right to data privacy, making secure AI a tangible reality rather than a theoretical aspiration.

Key Takeaways

  • BumbleBee enables highly efficient and secure two-party inference for large transformer models, addressing critical privacy concerns associated with sending sensitive data to AI service providers without relying on trusted third parties.
  • The framework achieves a significant reduction in communication costs (up to 100x) compared to prior secure computation methods, making private transformer inference more practical and economically viable for real-world deployment.
  • A compiler-based approach with a middle-layer Intermediate Representation (IR) is crucial for adapting existing machine learning code (e.g., JAX models from Hugging Face) to 2PC without requiring model fine-tuning, while also enabling powerful graph-level optimizations.
  • Simple, numerically stable approximations are often superior for nonlinear functions in 2PC. BumbleBee demonstrates that piecewise approximations for GELU and low-degree Taylor series for bounded Softmax inputs provide better efficiency and robustness than more complex or "fancy" numerical methods that can introduce prohibitive costs or stability issues.
  • Careful optimization of core operations like Softmax is essential. By rewriting the computation graph at the IR level to replace expensive vector divisions with a single inverse and by ensuring numerical stability through input bounding, Softmax's overhead is reduced to less than 1% of the total inference cost, eliminating it as a bottleneck.
  • BumbleBee represents a significant step towards making privacy-preserving AI inference practical and widely adoptable, by seamlessly integrating with existing ML ecosystems and focusing on the most critical performance bottlenecks (communication and nonlinear functions).

About the Speaker(s)

Wen-jie Lu is a researcher from Zjanu and the primary speaker for the "BumbleBee" talk at the NDSS Symposium. His work focuses on advancing secure computation techniques, particularly in the domain of privacy-preserving machine learning. During his presentation, he humorously referenced his prior work, a paper titled "Cheetah," which explored similar secure inference concepts but for neural networks. This historical context highlights his long-standing commitment and expertise in developing practical and efficient cryptographic solutions for machine learning applications.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Solid applied-crypto research on a real and underserved problem — private inference for large transformers without trusted hardware. The contributions are genuine (IR-level graph rewriting, numerically stable approximations, 100x communication reduction) but the work sits comfortably within an established research lineage rather than breaking new ground. Right venue, right lane, competent execution.

Heather Calloway (CISO) — WEAK

Technically serious cryptographic research on private inference for transformers — the communication overhead reduction is real and the compiler-based approach is genuinely clever. But this talk has no legible path to governance, institutional accountability, or operational decision-making, which puts it outside what I can meaningfully recommend to security leaders or executives.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025