Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPC

Tianshi Xu (Ping University)

34th USENIX Security Symposium (USENIX Security '25) · Day 2 · Privacy 1: Differential Privacy and Audit

Overview

This talk, presented by Tianshi Xu from Peking University, introduces a novel framework named BB (presumably "Breaking the Barrier") that significantly advances the field of private transformer inference. In an era where large language models (LLMs) and transformer-based architectures are ubiquitous across sensitive domains like healthcare, finance, and personalized assistance, protecting both the user's input data and the proprietary model parameters is paramount. Cryptography-based private inference offers robust, provable security guarantees, ensuring that the server learns nothing about the client's input and vice-versa.

Watch on YouTube · Read the paper · Download the PDF (PDF) · Slides

Paper abstract

Large Language Models (LLMs) have risen significantly in popularity and are increasingly being adopted across multiple applications. These LLMs are heavily aligned to resist engaging in illegal or unethical topics as a means to avoid contributing to responsible AI harms. However, a recent line of attacks, known as "jailbreaks'', seek to overcome this alignment. Intuitively, jailbreak attacks aim to narrow the gap between what the model can do and what it is willing to do. In this paper, we introduce a novel jailbreak attack called Crescendo. Unlike existing jailbreak methods, Crescendo is a simple multi-turn jailbreak that interacts with the model in a seemingly benign manner. It begins with a general prompt or question about the task at hand and then gradually escalates the dialogue by referencing the model's replies progressively leading to a successful jailbreak. We evaluate Crescendo on various public systems, including ChatGPT, Gemini Pro, Gemini-Ultra, LlaMA-2 70b and LlaMA-3 70b Chat, and Anthropic Chat. Our results demonstrate the strong efficacy of Crescendo, with it achieving high attack success rates across all evaluated models and tasks. Furthermore, we present Crescendomation, a tool that automates the Crescendo attack and demonstrate its efficacy against state-of-the-art models through our evaluations. Crescendomation surpasses other state-of-the-art jailbreaking techniques on the AdvBench subset dataset, achieving 29-61% higher performance on GPT-4 and 49-71% on Gemini-Pro. Finally, we also demonstrate Crescendo's ability to jailbreak multimodal models.

Visual summary for Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPC by Tianshi Xu
Visual summary for Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPC by Tianshi Xu

Key moments

  1. 4:00 Motivation: Reducing communication overhead in hybrid FHE/MPC
  2. 4:30 Understanding frequent FHE/MPC conversions and truncations
  3. 6:40 Limitations of prior linear layer fusion techniques
  4. 8:00 Introducing BB framework: Operator-wise evaluation key idea
  5. 10:00 BB framework: Fine-grain fusion example and components
  6. 11:30 Designing a secure CKKS to MPC conversion protocol

Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPC

Speakers: Tianshi Xu

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=JBBIPhjmuT0

Overview

This talk, presented by Tianshi Xu from Peking University, introduces a novel framework named BB (presumably "Breaking the Barrier") that significantly advances the field of private transformer inference. In an era where large language models (LLMs) and transformer-based architectures are ubiquitous across sensitive domains like healthcare, finance, and personalized assistance, protecting both the user's input data and the proprietary model parameters is paramount. Cryptography-based private inference offers robust, provable security guarantees, ensuring that the server learns nothing about the client's input and vice-versa.

The core challenge addressed by BB lies in the prohibitive communication overhead inherent in existing hybrid cryptographic approaches, which combine fully homomorphic encryption (FHE) with secure multi-party computation (MPC). While these hybrid schemes achieve high accuracy, their practical deployment has been hampered by massive data transfers, with state-of-the-art systems requiring upwards of 60 GB for a single BERT-base inference. BB proposes a paradigm shift from traditional layer-wise evaluation to a more granular operator-wise evaluation, coupled with innovative cryptographic protocols, to dramatically reduce this communication bottleneck and make private transformer inference a practical reality.

Background

▶ Watch: Motivation: Reducing communication overhead in hybrid FHE/MPC (4:00)

Achieving two-party private inference, where a client has private input and a server holds a private transformer model, typically relies on one of two cryptographic strategies. The first is pure FHE, which allows computations on encrypted data without decryption. While non-interactive, FHE struggles profoundly with complex nonlinear layers, leading to either unacceptable accuracy degradation or extremely high computational overhead due primarily to the need for high-degree polynomial approximations and costly bootstrapping operations. The second, more accurate approach, combines FHE with MPC in an interactive manner. This hybrid FHE/MPC strategy excels in accuracy but is plagued by colossal communication costs. For instance, prior work like Bolt and Bumblebee demonstrated that a single BERT-base inference could necessitate around 60 GB of communication.

The fundamental reason for this massive communication overhead in hybrid FHE/MPC frameworks is the frequent need for FHE-to-MPC conversions and truncations. The common hybrid approach uses FHE for linear layers (e.g., matrix multiplications) and MPC for nonlinear layers (e.g., activation functions like GELU). However, FHE operates on ciphertexts representing fixed-point numbers, while mainstream MPC protocols work on secret shares. This necessitates converting FHE ciphertexts into MPC secret shares between linear and nonlinear layers, an interaction-heavy process. Furthermore, cryptographic operations, especially multiplications, cause the numerical scale (or "gear" as referred to in the talk, effectively the exponent in fixed-point representation) to square, leading to an exponential growth of precision requirements. Truncation operations are required to restore this scale, preventing bit width explosion and maintaining numerical stability. Both FHE-to-MPC conversions and truncations rely on expensive MPC subprotocols that demand substantial communication. Previous research indicated that these conversions and truncations account for over 80% of the total communication cost in hybrid FHE/MPC schemes.

Prior attempts to mitigate this issue involved linear layer fusion, where consecutive linear layers are evaluated under FHE within a single communication round, thereby eliminating conversions and truncations between them. For example, in a self-attention layer, the two consecutive matrix multiplications (Q * K^T) could be fused. However, this approach presented several limitations. Firstly, it offered limited flexibility, as fusion was coarse-grained and only possible across immediately neighboring layers, resulting in insufficient communication reduction. Secondly, these frameworks typically employed BFV FHE, which operates on integers and suffers from exponential ciphertext bit width growth (referred to as "site tech gear" or scale) with increasing multiplication depth. This exponential growth in both plaintext and ciphertext bit widths leads to significant inefficiency. Thirdly, linear layer fusion often forces the use of ciphertext-ciphertext matrix multiplication (CTCT), which is computationally far more intensive than ciphertext-plaintext operations. This is primarily because adjusting the packing of intermediate ciphertexts in CTCT requires complex and costly FHE rotations. For instance, Bolt reportedly required approximately 20 times more computationally intensive FHE rotations for fused CTCT matrix multiplications compared to unfused ciphertext-plaintext protocols. These limitations underscore the necessity for a more sophisticated approach to truly break the communication barrier.

Key Findings

▶ Watch: Limitations of prior linear layer fusion techniques (6:40)

The BB framework introduces a powerful, yet simple, key idea: even seemingly nonlinear layers within transformer architectures, such as GELU or Layer Normalization, are fundamentally composed of multiple linear operators (like multiplication and addition) when analyzed at a fine-grained level. This crucial insight challenges the traditional "layer-wise" evaluation paradigm, where each layer is treated as an indivisible unit requiring a specific protocol (FHE for linear, MPC for nonlinear). Instead, BB advocates for an operator-wise evaluation, enabling the fusion of adjacent linear operators across traditional layer boundaries.

This paradigm shift forms the cornerstone of BB's contributions, leading to significant reductions in communication overhead. By decomposing layers into their constituent operators—classifying them as either linear or nonlinear—BB systematically identifies and fuses all adjacent linear operators. This process allows for their evaluation within a single communication round under FHE, thereby eliminating a substantial number of FHE-to-MPC conversions and truncations that would otherwise be necessary. For example, the talk illustrates how the last two operators of a Layer Normalization layer, the subsequent fully connected layer, and the first three operators of a GELU layer, all being linear, can be fused. This fine-grained fusion strategy can eliminate all truncations and 55% of FHE-to-MPC conversions in such scenarios.

BB comprises three main technical components that collectively enable this advanced private inference:

  1. Fine-Grain Fusion: This is the systematic analysis of the transformer's computation graph at an operator level. BB establishes a standardized definition for operators, categorizes them into linear and nonlinear types, and identifies valid fusion patterns for adjacent linear operators, even if they span different conventional layers.
  1. Secure CKKS-to-MPC Conversion Protocol: BB adopts CKKS FHE (Cheon-Kim-Kim-Song scheme) instead of BFV. CKKS is generally preferred for approximate computations on real numbers and offers more effective control over the growth of ciphertext scale (or "noise"). However, existing CKKS-to-MPC conversion protocols were found to be insecure, leaking intermediate results. BB addresses this critical vulnerability by designing and implementing the first secure CKKS-to-MPC conversion protocol, ensuring data privacy throughout the hybrid evaluation.
  1. Rotation-Efficient metimo Protocol: Fused ciphertext-ciphertext matrix multiplications (metimo) are notoriously expensive due to the high cost of FHE rotations required for data realignment. BB introduces a specialized metimo protocol optimized for the core operations within multi-head attention mechanisms: QK transpose and Softmax times V. The protocol leverages the observation that multi-head attention naturally involves batched matrix multiplications, allowing multiple batches to be packed into a single ciphertext. This technique inherently reduces the number of required FHE rotations. Furthermore, BB applies the Baby-Step Giant-Step (BSGS) optimization to further minimize rotation costs. These optimizations collectively achieve a remarkable 29 times and 8 times reduction in FHE rotations compared to Bolt and PowerFormer, respectively.

Technical Deep Dive

▶ Watch: Introducing BB framework: Operator-wise evaluation key idea (8:00)

The core innovation of BB lies in its fundamental shift from a layer-wise to an operator-wise evaluation paradigm. Traditional private inference schemes treat entire transformer layers (e.g., self-attention, feed-forward networks, LayerNorm, GELU) as atomic units. When a linear layer is followed by a nonlinear one, an FHE-to-MPC conversion and truncation are typically required. BB, however, recognizes that even "nonlinear" layers often contain numerous linear operations. For instance, Layer Normalization involves additions, multiplications, and divisions (which can be approximated or handled carefully). Similarly, activation functions like GELU are commonly implemented using polynomial approximations, which themselves are sequences of linear multiplications and additions.

BB's fine-grain fusion process begins by meticulously decomposing the entire transformer computation graph into its most basic operators. These operators are then rigorously classified as either linear (e.g., matrix multiplication, vector addition, element-wise multiplication by a constant) or nonlinear (e.g., comparison, square root, certain divisions, or full nonlinear activation functions). Once classified, the framework identifies all adjacent linear operators, regardless of whether they belong to the same original layer or span across multiple conventional layers. These identified sequences of linear operators are then fused, meaning they are all evaluated consecutively using FHE without any intermediate communication, FHE-to-MPC conversions, or truncations. This significantly reduces the communication footprint by delaying interaction until after the fused block of linear operations is complete. The example provided in the talk, where a portion of LayerNorm, a fully connected layer, and part of a GELU layer are fused, vividly demonstrates how BB "breaks the layer barrier" by enabling fusion across these traditional boundaries.

The choice of CKKS FHE over BFV is a critical technical decision. BFV, while suitable for exact integer arithmetic, struggles with fixed-point arithmetic due to its exponential growth of scale (or gear) and ciphertext bit width with each multiplication. CKKS, on the other hand, is designed for approximate arithmetic on real numbers, making it more suitable for machine learning computations where some precision loss is acceptable. More importantly, CKKS offers better mechanisms for managing the scale of ciphertexts, preventing the uncontrolled growth that plagues BFV in deep multiplicative circuits. This better scale management in CKKS is essential for maintaining accuracy and efficiency in complex transformer models.

However, adopting CKKS introduced a new challenge: the insecurity of existing CKKS-to-MPC conversion protocols. The transcript states that prior protocols "were insecure leaking intermediate results." While the specific details of how they leaked information are not provided, the implication is that direct conversion methods revealed sensitive plaintext information during the transformation process. BB's contribution here is the design of the "first secure protocol" for this conversion. This secure protocol ensures that the transformation from a CKKS ciphertext (representing encrypted real numbers) to MPC secret shares (for distributed computation) occurs without exposing any part of the underlying data to any party, thereby upholding the strong privacy guarantees.

The third technical pillar is the rotation-efficient metimo (matrix multiplication) protocol. In FHE, matrix multiplication, particularly when both operands are ciphertexts (CTCT metimo), is computationally demanding. The primary bottleneck is the need for FHE rotations to align data within the ciphertext slots before addition. These rotations are expensive, consuming a significant portion of the computation budget. BB's protocol specifically targets the multi-head attention mechanism, which is a key component of transformers and involves two main CTCT operations: QK transpose and Softmax * V. The protocol leverages two key insights:

  1. Batched Matrix Multiplications: Multi-head attention naturally involves multiple "heads" operating in parallel, each performing its own matrix multiplication. BB observes that these naturally batched operations can be further optimized by packing multiple batches (e.g., outputs from different attention heads) into a single large ciphertext. This reduces the total number of ciphertexts and, consequently, the number of FHE rotations required across the entire operation.
  2. Baby-Step Giant-Step (BSGS) Optimization: This well-known algorithm is adapted to further reduce the number of FHE rotations. BSGS is a technique used in various cryptographic contexts to optimize operations that involve a sequence of cyclic shifts (like FHE rotations). By structuring the rotation sequence efficiently, BSGS can drastically cut down the total number of rotation operations. The combined effect of batching and BSGS is substantial, leading to a 29 times reduction in FHE rotations compared to Bolt and an 8 times reduction compared to PowerFormer, both prominent prior works. This dramatic reduction in rotation costs significantly improves the computational efficiency of the fused linear operators.

Demo / Proof of Concept

▶ Watch: BB framework: Fine-grain fusion example and components (10:00)

While the talk did not feature a live, interactive demonstration, the BB framework was rigorously validated through extensive experimental results. The implementation leveraged established cryptographic libraries: Seal (for FHE operations) and EasyPC (for MPC protocols), with GPU acceleration built upon Phantom FHE. The evaluation focused on standard transformer models widely used in research and industry, specifically GPT-2 base, BERT base, and BERT large.

The efficacy of BB was demonstrated by comparing its performance against several leading private transformer inference frameworks. This included other hybrid FHE/MPC frameworks such as Iron, Bolt, and Bumblebee, as well as three-party computation frameworks like Sigma and MPCFormer, and the FHE-based framework Nexus. The experimental results unequivocally showed BB's superiority across key metrics.

In terms of communication overhead, BB achieved a remarkable 21 times reduction compared to Bolt and a 2 times reduction compared to Bumblebee. This substantial decrease in data transfer volume directly addresses the primary bottleneck of hybrid FHE/MPC schemes. Furthermore, BB demonstrated significant improvements in latency, achieving up to a 29 times speedup on CPU and up to a 13 times speedup on GPU compared to prior work. These performance gains highlight BB's ability to not only reduce communication but also accelerate the overall inference process, making private transformer inference substantially more practical for real-world applications.

Defensive Implications

▶ Watch: Designing a secure CKKS to MPC conversion protocol (11:30)

The advancements presented by the BB framework carry profound defensive implications for organizations deploying transformer-based AI models in sensitive environments. By drastically reducing the communication overhead and improving the latency of private transformer inference, BB lowers a significant barrier to the widespread adoption of privacy-preserving machine learning (PPML).

For industries handling highly sensitive data—such as patient records in healthcare, financial transactions and personal information in finance, or private user preferences in personalized assistance—BB offers a practical pathway to leverage powerful transformer models without compromising data confidentiality. It enables organizations to adhere to stringent data privacy regulations (e.g., GDPR, HIPAA, CCPA) by providing strong, provable cryptographic guarantees: the client's input remains encrypted and unknown to the model server, and the model parameters remain protected from the client.

Specifically, BB's contributions mean that:

  • Wider Adoption of PPML: The improved efficiency makes PPML a more viable option for real-time or near real-time applications, moving it from a theoretical possibility to a deployable solution for large-scale models like BERT and GPT-2.
  • Enhanced Data Security "In Use": Unlike traditional security measures that protect data at rest or in transit, private inference protects data during computation. This closes a critical gap in the data lifecycle where data is typically decrypted and vulnerable.
  • Protection of Intellectual Property: Model parameters, which represent significant intellectual property, are also protected. This allows companies to offer their advanced AI services without exposing their proprietary models.
  • Facilitating Collaborative AI: BB could enable secure collaboration between entities, where one party provides sensitive data and another provides a proprietary model, without either party having to fully disclose their assets.

In essence, BB empowers defenders by providing a more efficient and secure means to deploy cutting-edge AI, reconciling the often-conflicting demands of advanced analytics and robust data privacy. It shifts the landscape towards a future where privacy is an intrinsic feature of AI deployments, rather than an afterthought.

Key Takeaways

  • Operator-Level Fusion: BB breaks the traditional "layer barrier" by decomposing transformer layers into fundamental linear and nonlinear operators, enabling fine-grain fusion of adjacent linear operators across layer boundaries.
  • CKKS for Scale Management: The adoption of CKKS FHE is crucial for efficiently handling approximate real-number computations and effectively controlling the growth of ciphertext scale, a significant improvement over BFV for transformer inference.
  • Secure CKKS-to-MPC Conversion: BB introduces the first secure protocol for converting CKKS ciphertexts to MPC secret shares, addressing a critical security vulnerability found in prior CKKS-to-MPC conversion methods.
  • Rotation-Efficient Matrix Multiplication: A specialized metimo protocol, optimized for multi-head attention and leveraging batching and the Baby-Step Giant-Step (BSGS) algorithm, dramatically reduces expensive FHE rotations (up to 29x reduction).
  • Significant Performance Gains: The framework achieves a remarkable 21 times reduction in communication overhead compared to Bolt and up to 29 times CPU speedup and 13 times GPU speedup in latency compared to prior works.
  • Practical Private Inference: These advancements make private transformer inference for large models (like BERT and GPT-2) substantially more practical and deployable in sensitive, real-world applications, bridging the gap between cryptographic theory and real-world utility.

About the Speaker(s)

The talk was presented by Tianshi Xu from Peking University. This work, titled "Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPC," is a collaborative effort, developed in conjunction with researchers and collaborators from TikTok and Peking University. This blend of academic rigor from Peking University and industry relevance from TikTok underscores the practical applicability and robust design of the BB framework.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Solid, technically credible contribution to private ML inference that solves a real bottleneck — the operator-wise fusion insight is genuinely non-obvious and the CKKS-to-MPC security fix addresses an actual vulnerability in prior work. The 21x communication reduction and rotation efficiency gains are backed by concrete benchmarks against named prior systems, which is exactly what you want to see.

Heather Calloway (CISO) — WEAK

Legitimate cryptographic research that advances private inference efficiency in a meaningful way — but this talk has no path to a CISO's desk. The framing gestures at healthcare and HIPAA, but never closes the loop on what an organization would actually do with this, when, or at what cost.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)