Secure Transformer Inference Made Non-interactive

Jiawen Zhang

Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · Privacy & Cryptography 1 · Privacy & Cryptography 1

Overview

The rapid advancement of transformer models has revolutionized artificial intelligence, powering applications from language translation to content generation and question answering. However, the widespread deployment of these powerful models, particularly in services like OpenAI's ChatGPT, introduces significant privacy concerns as users submit sensitive data through prompts and messages. This talk introduces Nexus, a groundbreaking non-interactive protocol for secure transformer inference. It addresses the critical need for privacy-preserving AI while overcoming the substantial computational and communication overheads that have plagued previous secure inference solutions.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction and privacy concerns for transformers
  2. 1:50 Two key challenges for secure transformer inference
  3. 4:15 Introducing Nexus: a non-interactive FHE protocol
  4. 5:20 Efficient matrix multiplication via polynomial compression
  5. 8:00 Quick Max algorithm for efficient argmax evaluation
  6. 9:30 Strategic bootstrapping placement for overall performance

Secure Transformer Inference Made Non-interactive

Speakers: Jiawen Zhang

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=wjMNUKH3-Ys

Overview

The rapid advancement of transformer models has revolutionized artificial intelligence, powering applications from language translation to content generation and question answering. However, the widespread deployment of these powerful models, particularly in services like OpenAI's ChatGPT, introduces significant privacy concerns as users submit sensitive data through prompts and messages. This talk introduces Nexus, a groundbreaking non-interactive protocol for secure transformer inference. It addresses the critical need for privacy-preserving AI while overcoming the substantial computational and communication overheads that have plagued previous secure inference solutions.

Presented by Jiawen Zhang, this work tackles the challenge of enabling a server to perform model inference without learning anything about the user's input, and for the user to learn nothing about the server's proprietary model beyond the inference result. Nexus achieves this by leveraging Fully Homomorphic Encryption (FHE) in a novel way, specifically tailored for the unique architectural demands of large transformer models. The paper presents innovations in efficient matrix multiplication and a dramatically optimized Argmax evaluation, making secure inference for large language models (LLMs) not only possible but also practically viable in terms of speed and cost.

The significance of Nexus lies in its ability to bridge the gap between advanced AI capabilities and user data privacy. By offering a non-interactive, FHE-based solution with demonstrated efficiency on both CPU and GPU, Nexus paves the way for a new generation of privacy-preserving AI services. It stands as a pivotal contribution to the field, offering a robust framework for deploying sensitive AI applications while safeguarding user information, a paramount concern in today's data-driven world.

Background

▶ Watch: Introduction and privacy concerns for transformers (0:00)

Transformer models have become the cornerstone of modern artificial intelligence, driving significant breakthroughs in natural language processing (NLP) and beyond. Their ability to process sequential data and capture long-range dependencies has led to their adoption in diverse applications such as language translation, sophisticated content generation, and intelligent question-answering systems. Prominent examples include large language models (LLMs) like those powering ChatGPT, which are accessed by users through online inference services or public APIs. While these services offer immense convenience, they inherently expose user-submitted data, which often contains private and sensitive information, to the service provider. This exposure has raised critical concerns about user privacy and data security.

Secure inference emerged as a cryptographic solution to this privacy dilemma. It describes a two-party protocol where a user provides a private input to a server holding a private model. The server performs inference on the user's input, and the user receives the result, all while ensuring that the server learns nothing about the input and the user learns nothing about the model beyond the final inference output. In the past, many such protocols were developed primarily for convolutional neural networks (CNNs). However, applying these techniques to transformers presents unique challenges.

The key differences between transformers and CNNs, particularly in the context of secure inference, are twofold:

  1. Large-scale Matrix Multiplications: Transformers heavily rely on large-scale matrix multiplications. Previous FHE-based secure inference work, such as Cheetah and Iron, often computed matrix multiplications via inner dot products and employed sparse packing for results. This approach frequently led to a significant waste of data slots within ciphertexts, introducing substantial communication overhead. For transformers, the input dimensions of matrices are much higher, exacerbating this computational and communication cost, making existing methods inefficient.
  2. Higher Dimension Inputs to Argmax: The final layer in both CNNs and transformers often involves an Argmax operation, which selects the label with the highest probability from a probability vector. While in CNN classification tasks, the number of labels (M) is typically manageable (e.g., M=1,000 for ImageNet-1K), in transformer-based NLP tasks, the vocabulary size (M) can be enormous. For instance, M can exceed 30,000 in BERT and over 100,000 in models like LLaMA 3 8B. The state-of-the-art FHE algorithm for Argmax, presented by Phoenix at CCS '22, has a computational complexity of O(M). This linear dependency on M renders it impractical for the vast vocabularies encountered in modern LLMs.

Furthermore, most previous secure inference works were interactive, relying on multi-party computation (MPC) protocols. Interactive MPC protocols suffer from two significant disadvantages:

  1. Substantial Communication Overhead: MPC protocols necessitate multiple rounds of communication between parties, leading to massive data transfers. For example, some protocols consumed about 59 GB of bandwidth over 10,000 interaction runs, translating to a financial cost of over $5 for a single word inference. Such figures are not practical for real-world inference services.
  2. Difficulty with Hardware Acceleration: The inherent communication bottleneck of MPC-based protocols makes them difficult to accelerate effectively using specialized hardware like GPUs and FPGAs, which are crucial for the performance of large AI models.

These limitations highlighted a critical need for a non-interactive secure inference protocol specifically designed for transformers, one that could achieve both efficiency and communication optimization, ultimately enabling practical, privacy-preserving AI.

Key Findings

▶ Watch: Introducing Nexus: a non-interactive FHE protocol (4:15)

Nexus presents several pivotal contributions that collectively address the long-standing challenges in secure transformer inference:

  1. First Non-Interactive FHE Protocol for Transformers: Nexus is introduced as the inaugural non-interactive protocol for secure transformer inference, built upon Fully Homomorphic Encryption (FHE). This fundamental shift from interactive MPC significantly reduces communication overhead and enables better hardware compatibility.
  2. Efficient and Communication-Optimized Matrix Multiplications: The protocol devises a novel method for FHE-based matrix multiplication that dramatically reduces communication costs. By allowing the server to compress model parameters into polynomials and the client to expand them, Nexus reuses ciphertexts effectively, leading to lower amortized computation and communication overhead compared to prior methods that wasted data slots.
  3. Highly Efficient Argmax Evaluation with QuickMax: Nexus introduces the QuickMax algorithm, a revolutionary approach to computing Argmax in FHE. Unlike previous O(M) complexity methods (like Phoenix's bubble sort-inspired approach), QuickMax achieves a significantly improved complexity of log N time rotation and sign operations for N-dimension inputs. This logarithmic scaling makes Argmax practical even for the extremely large vocabulary sizes found in modern transformer models (e.g., 100,000+).
  4. Optimized Nonlinear Function Evaluation: The work provides efficient FHE evaluations for other critical nonlinear functions common in transformers, such as GELU (approximated with piecewise polynomials) and Layer Normalization (inverse square root computed via Newton's iteration).
  5. Strategic Bootstrapping Placement: Recognizing that bootstrapping is the most expensive operation in leveled FHE, Nexus implements a crucial optimization by strategically placing bootstrapping operations. These operations are performed when the ciphertext dimension is at its smallest, thus minimizing the number of ciphertexts to be refreshed and significantly reducing the overall runtime cost.
  6. Demonstrated Practicality and Performance: Nexus showcases exceptional real-world performance. In end-to-end comparisons with existing secure inference frameworks using BERT-base, Nexus achieves the lowest communication overhead and financial cost (as low as 5 cents per token). It completes a secure inference under GPU acceleration in just half a minute, demonstrating its practical viability.
  7. Open-Source Implementation with SEAL and CUDA Acceleration: The researchers have open-sourced their code on GitHub, garnering significant community interest. Notably, Nexus includes the first-time implementation of CKKS bootstrapping on Microsoft SEAL 4.0 and provides the first CUDA source code to accelerate Microsoft SEAL, marking substantial contributions to the FHE ecosystem and enabling broader adoption.

These findings collectively represent a significant leap forward for secure AI, offering a practical, high-performance, and privacy-preserving solution for the growing landscape of transformer-based applications.

Technical Deep Dive

▶ Watch: Efficient matrix multiplication via polynomial compression (5:20)

Nexus's technical innovations primarily revolve around optimizing two critical and computationally intensive operations within transformer inference: matrix multiplications and the Argmax function, all within the framework of Fully Homomorphic Encryption (FHE). The protocol is built upon the CKKS (Cheon-Kim-Kim-Song) leveled homomorphic encryption scheme, which allows computations on encrypted data while preserving approximate real numbers.

Efficient Matrix Multiplications

The core challenge with matrix multiplications in FHE for transformers stems from their high input dimensions. Traditional FHE approaches for matrix multiplication, like those in Cheetah and Iron, often rely on inner dot products and sparse packing. This leads to a significant waste of data slots in ciphertexts, increasing communication overhead, especially when inputs are large. Furthermore, a naive approach of packing each element into a separate ciphertext and sending multiple identical ciphertexts is highly inefficient.

Nexus introduces a novel, communication-optimized method:

  1. Server-Side Compression: The server, possessing the model parameters (which remain constant during inference), compresses its matrix elements into a polynomial by coefficients. This polynomial is then encrypted using the client's public key. Instead of sending many ciphertexts, the server sends a compact, encrypted polynomial.
  2. Client-Side Expansion: Upon receiving the encrypted polynomial, the client can then "expand" this single ciphertext into multiple ciphertexts. Crucially, in this expansion, each original matrix element ends up in the constant term position of a resulting ciphertext. The paper proves that "the encryption of a polynomial with only constant term is exactly an SMD form of unidentical values," which is key to efficient packing.
  3. Ciphertext Reuse: A significant advantage of this approach is that these SMD (Single Instruction, Multiple Data) packed ciphertexts can be reused across multiple inference queries. Since the model parameters are static, the server's compressed polynomial and the client's expanded ciphertexts derived from it remain valid, eliminating the need to re-transmit or re-compute them for subsequent inputs. This dramatically reduces the amortized communication and computation cost.

This expansion algorithm was initially proposed in the COPR work, requiring only log N levels and being highly amenable to parallel computation. The amortized cost analysis for matrix multiplications, using BERT-based parameters and varying input query numbers, demonstrates Nexus's superior efficiency in both computation and communication compared to prior methods.

Efficient Argmax Evaluation with QuickMax

The Argmax operation, which identifies the index of the maximum value in a vector, is critical for the final classification or token selection in transformer models. Its high dimensionality in NLP tasks (M up to 100,000+) makes it a major bottleneck.

Previous FHE Argmax methods, such as Phoenix (CCS '22), adopted a strategy akin to bubble sorting. This involved:

  1. Comparison: Comparing each element with its nearby elements.
  2. Rotation: Rotating ciphertexts to bring elements into comparison positions.
  3. Sign Operation: Calculating the sign of the difference between compared elements.
  4. Scoreboard: Using the sign as a "scoreboard" – adding a point if an element is larger, subtracting if smaller. The largest element would accumulate the highest score.

This approach requires N times rotate and sign operations for an N-dimension input, leading to an O(N) complexity which is prohibitive for large M.

Nexus proposes the novel QuickMax algorithm to overcome this O(N) barrier, achieving a much more efficient log N time rotation and sign operations. The core idea is that if the maximum value itself is known, computing Argmax (finding its index) becomes straightforward. QuickMax operates similarly to computing the roots of a binary tree. While the exact details are in the paper, the high-level concept involves a hierarchical comparison process that reduces the number of required operations logarithmically. This structure allows F (which can be any associative function like sum, max, or min) to be applied efficiently. This logarithmic complexity is a game-changer for large vocabulary sizes, enabling practical Argmax computation in FHE.

Other Nonlinear Functions

Beyond matrix multiplication and Argmax, transformers employ other nonlinear activation and normalization functions that require FHE-compatible approximations:

  • GELU (Gaussian Error Linear Unit): Nexus approximates GELU using a piecewise polynomial approach, similar to techniques found in works like Bubble and PUMA. Polynomial approximations are essential in FHE, as only addition and multiplication operations are directly supported on ciphertexts.
  • Layer Normalization: This involves calculating an inverse square root. Nexus leverages Newton's iteration method to compute the inverse square root. Newton's method is an iterative numerical technique that can be implemented using polynomial operations, making it suitable for FHE.

Bootstrapping and Placement Strategy

As Nexus is based on a leveled FHE scheme (CKKS), each homomorphic multiplication consumes a "level" of noise budget. Once the noise level becomes too high (i.e., the ciphertext level becomes too low), a complex and computationally expensive operation called bootstrapping is required. Bootstrapping "refreshes" the ciphertext to a higher level, enabling more multiplications. The cost of bootstrapping scales linearly with the number of ciphertexts.

Therefore, the placement of bootstrapping is critical for overall performance. Nexus employs a strategic placement approach:

  • Minimal Dimension Principle: Bootstrapping is performed when the input/output dimension of the building blocks is the smallest. This corresponds to when the number of packed ciphertexts is at its least. By refreshing fewer ciphertexts, the total bootstrapping cost is minimized.

The talk illustrates this with a diagram showing the placement for a BERT-based transformer, highlighting how bootstrapping is strategically inserted at points where the data representation is most compact. The runtime analysis confirms that bootstrapping remains the most time-consuming part, requiring over 500 seconds and occupying over 60% of the total execution time even with optimizations. This underscores the importance of efficient placement.

In summary, Nexus combines novel FHE-friendly algorithms for critical transformer operations with a careful management of FHE overheads, particularly bootstrapping, to deliver a truly practical and performant secure inference solution.

Demo / Proof of Concept

▶ Watch: Quick Max algorithm for efficient argmax evaluation (8:00)

The practical viability of Nexus was rigorously demonstrated through extensive performance evaluations, showcasing its efficiency and cost-effectiveness compared to existing secure inference frameworks. The implementation and benchmarks were conducted on a robust machine equipped with a 32-core CPU and four NVIDIA A100 GPUs, highlighting its capability to leverage high-performance hardware.

The evaluation involved batching 32 inputs in total and assessing the performance of individual operations within Nexus, as well as the end-to-end secure inference process. A key finding from the granular analysis of operations was that bootstrapping is, by far, the most time-consuming component. It consumed over 500 seconds of runtime and accounted for more than 60% of the total time in both CPU and GPU implementations, even with the strategic placement optimizations discussed earlier. This result underscores the ongoing challenge of bootstrapping in FHE and the critical importance of minimizing its frequency and ciphertext count.

For the end-to-end comparison, Nexus was benchmarked against existing secure inference frameworks using a BERT-base model with an input consisting of 128 tokens. The results were compelling:

  • Communication Overhead: Nexus demonstrated the lowest communication overhead among all compared frameworks. This is a direct benefit of its non-interactive FHE design and optimized matrix multiplication protocol, which minimizes data transfer between client and server.
  • Financial Cost: Correspondingly, Nexus exhibited the lowest financial cost. The talk specified a cost of just 5 cents per token for secure inference. This cost was calculated based on Amazon Web Services (AWS) pricing policies for data transfer (e.g., 9 cents per gigabyte) and CPU/GPU usage (e.g., $1 per hour for a 64-core CPU). This low cost makes secure transformer inference economically feasible for real-world deployment.
  • Runtime Performance: The non-interactive nature of Nexus, requiring only one interaction throughout the entire process, makes it highly compatible with hardware acceleration. Under GPU acceleration, Nexus was able to complete a secure inference in just half a minute. This impressive speed is a significant improvement over previous interactive MPC-based solutions that suffered from communication bottlenecks.

Beyond these performance metrics, the Nexus project has also made significant contributions to the broader FHE ecosystem:

  • Open-Source Availability: The code for Nexus has been open-sourced on GitHub for several months, where it has already garnered considerable interest, evidenced by 86 stars. This commitment to open science facilitates reproducibility, community engagement, and further research.
  • Microsoft SEAL Integration and CUDA Acceleration: The project achieved a notable milestone by implementing CKKS bootstrapping for the first time on Microsoft SEAL version 4.0. Furthermore, Nexus is the first to provide CUDA source code to accelerate Microsoft SEAL, demonstrating a pioneering effort in integrating high-performance GPU capabilities directly into a leading FHE library. These contributions are crucial for improving the practical performance and adoption of FHE.

In essence, the demo and proof of concept unequivocally establish Nexus as a highly practical, efficient, and cost-effective solution for secure transformer inference, setting a new benchmark for privacy-preserving AI.

Defensive Implications

▶ Watch: Strategic bootstrapping placement for overall performance (9:30)

The development and open-sourcing of Nexus carry profound defensive implications for organizations and individuals concerned with data privacy in the age of AI. By providing a practical and efficient non-interactive protocol for secure transformer inference, Nexus offers a robust mechanism to protect sensitive information without sacrificing the utility of advanced AI models.

Here are the key defensive implications:

  1. Enabling Privacy-Preserving AI-as-a-Service (AIaaS): Nexus allows cloud providers and AI service developers to offer powerful transformer-based models (like LLMs) while guaranteeing input privacy for users and model privacy for service providers. This means users can submit sensitive prompts, medical data, financial information, or proprietary business queries to an AI model without the service provider ever learning the content of those inputs. This capability is critical for industries with stringent regulatory requirements (e.g., healthcare, finance) and for any application dealing with personally identifiable information (PII).
  1. Mitigating Data Leakage Risks: Traditional AI inference services inherently create a risk of data leakage. Nexus effectively eliminates this risk by performing computations on encrypted data. This prevents malicious insiders, external attackers, or even legitimate but over-curious service providers from accessing or inferring user data, thereby significantly reducing the attack surface related to data privacy.
  1. Compliance with Data Protection Regulations: Regulations such as GDPR, CCPA, HIPAA, and LGPD mandate strong protections for user data. Implementing secure inference protocols like Nexus can help organizations achieve compliance by demonstrating that sensitive data is processed in a privacy-preserving manner, even when leveraging third-party AI models. This moves beyond mere data at rest/in transit encryption to data in use encryption.
  1. Reduced Communication Overhead for Secure Operations: The drastically reduced communication overhead of Nexus, compared to interactive MPC protocols, makes secure inference more feasible for deployment. Less data transfer means fewer opportunities for interception and less network burden, enhancing overall system resilience and performance.
  1. Hardware Acceleration Compatibility: The non-interactive nature and GPU acceleration support of Nexus mean that organizations can integrate privacy-preserving AI into their existing high-performance computing infrastructures. This allows them to scale secure inference operations efficiently, making privacy a feature that doesn't necessitate prohibitive performance penalties.
  1. Protection Against Model Extraction Attacks (Partial): While the user learns the inference result, the FHE protocol ensures that they do not learn the model's parameters. This provides a level of model confidentiality, preventing direct model extraction by the client, which is crucial for protecting intellectual property and maintaining a competitive edge for model owners.
  1. Foundation for Future Privacy-Enhancing Technologies: By open-sourcing the code and contributing novel FHE optimizations (like CKS bootstrapping on Microsoft SEAL 4.0 with CUDA), Nexus provides a foundational toolkit for other researchers and developers. This will accelerate the development and adoption of other privacy-enhancing technologies (PETs) that rely on FHE, fostering a more secure AI ecosystem.

In essence, Nexus empowers defenders to build and deploy AI systems that are inherently more secure and privacy-respecting, moving towards a future where the power of AI can be harnessed without compromising fundamental data privacy rights.

Key Takeaways

  • First Non-Interactive FHE Protocol for Transformers: Nexus introduces the first non-interactive protocol for secure transformer inference, leveraging Fully Homomorphic Encryption (FHE) to address critical privacy concerns in AI models like LLMs, significantly reducing communication overhead compared to interactive MPC.
  • Optimized Matrix Multiplications and Argmax: The protocol features novel, communication-optimized matrix multiplication that reuses ciphertexts for efficiency and a groundbreaking QuickMax algorithm for Argmax, reducing its complexity from O(N) to log N time, making it practical for large vocabulary sizes (e.g., 100,000+).
  • Strategic Bootstrapping Placement: Nexus minimizes the high cost of FHE bootstrapping by strategically performing the operation when ciphertext dimensions are smallest, thereby reducing the number of ciphertexts to be refreshed and significantly improving overall runtime.
  • Demonstrated Practicality and Cost-Effectiveness: Benchmarks show Nexus achieves the lowest communication overhead and financial cost (as low as 5 cents per token) among existing frameworks, completing a secure inference for a BERT-base model in just half a minute on GPU.
  • Open-Source Contributions to FHE Ecosystem: The project is open-source and includes the first implementation of CKKS bootstrapping on Microsoft SEAL 4.0, alongside the first CUDA source code for Microsoft SEAL acceleration, fostering broader adoption and development of FHE.
  • Enabling Privacy-Preserving AI-as-a-Service: Nexus provides a practical pathway for organizations to offer AI services that guarantee user input privacy and model confidentiality, crucial for compliance with data protection regulations and building trust in AI applications.

About the Speaker(s)

The primary speaker for this presentation was Jiawen Zhang. The transcript indicates that Jiawen Zhang is a researcher involved in this joint work, presenting the paper "Secure Transformer Inference Made Non-interactive." Further specific biographical details regarding their affiliations or academic background were not extensively covered within the provided transcript or metadata.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Legitimate cryptographic systems research with a concrete novel contribution: a non-interactive FHE protocol for transformer inference that solves two real bottlenecks (matrix multiplication slot waste and O(N) Argmax) with measurable results. The QuickMax O(log N) improvement over Phoenix's O(N) approach is the kind of specific, verifiable claim that earns a serious look. The open-source CKKS bootstrapping implementation on SEAL 4.0 with CUDA is a genuine ecosystem contribution, not a footnote.

Heather Calloway (CISO) — WEAK

Credible cryptographic research with real-world implications for privacy-preserving AI, but this is a systems paper presented to a technical audience — not a governance or defender-operations talk. The regulatory and compliance framing in the article feels retrofitted, not argued. A CISO audience would need someone to do the institutional translation work that Jiawen Zhang doesn't attempt here.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025