From Individual Computation to Allied Optimization: Remodeling Privacy-Preserving Neural Inference with Function Input Tuning

Qiao Zhang, Tao Xiang, Chunsheng Xin, Hongyi Wu

IEEE Symposium on Security and Privacy 2024 · Day 3 · Continental Ballroom 6

Overview

The proliferation of Machine Learning as a Service (MLaaS) has democratized access to powerful AI capabilities, yet it introduces significant privacy challenges, particularly when handling sensitive data like medical records. This talk, presented by Qiao Zhang from Tongji University, delves into a novel approach to enhance the efficiency of privacy-preserving neural inference within MLaaS environments. Co-authored with Tao Xiang, Chunsheng Xin, and Hongyi Wu, the research addresses the inherent computational bottlenecks in existing privacy-preserving machine learning (PPML) solutions.

Watch on YouTube

Visual summary for From Individual Computation to Allied Optimization: Remodeling Privacy-Preserving Neural Inference with Function Input Tuning by Qiao Zhang, Tao Xiang, Chunsheng Xin, Hongyi Wu
Visual summary for From Individual Computation to Allied Optimization: Remodeling Privacy-Preserving Neural Inference with Function Input Tuning by Qiao Zhang, Tao Xiang, Chunsheng Xin, Hongyi Wu

Key moments

  1. 0:00 Introduction to privacy-preserving ML and problem statement.
  2. 2:30 Limitations of individual computation and motivation for joint optimization.
  3. 3:30 Mathematical decomposition of ReLU convolution for efficient joint optimization.
  4. 7:50 Structuring computation: dividing into data-dependent online and offline parts.
  5. 8:30 Detailed explanation of novel root-free offline computation method.
  6. 10:30 Overall recap of the proposed joint ReLU convolution optimization.

From Individual Computation to Allied Optimization: Remodeling Privacy-Preserving Neural Inference with Function Input Tuning

Speakers: Qiao Zhang, Tongji University; Tao Xiang, Tongji University; Chunsheng Xin, Old Dominion University; Hongyi Wu, University of Arizona

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=s4dbOVZls0U

Overview

The proliferation of Machine Learning as a Service (MLaaS) has democratized access to powerful AI capabilities, yet it introduces significant privacy challenges, particularly when handling sensitive data like medical records. This talk, presented by Qiao Zhang from Tongji University, delves into a novel approach to enhance the efficiency of privacy-preserving neural inference within MLaaS environments. Co-authored with Tao Xiang, Chunsheng Xin, and Hongyi Wu, the research addresses the inherent computational bottlenecks in existing privacy-preserving machine learning (PPML) solutions.

The core problem tackled is the inefficiency stemming from the "individual computation" paradigm, where linear and non-linear functions within neural networks are processed separately using distinct cryptographic primitives. This often necessitates costly operations such as multiplexing for non-linear activation functions like ReLU, and homomorphic rotation or extraction for linear operations like convolution. The proposed Function Input Tuning (FIT) framework introduces a joint optimization strategy that remodels the computation logic, aiming to eliminate these redundancies and significantly improve performance, making privacy-preserving AI more practical for real-world deployment in sensitive domains governed by regulations like HIPAA.

By intelligently integrating and optimizing the interaction between homomorphic encryption (HE) and other cryptographic techniques, FIT offers a substantial leap in efficiency. The work demonstrates how a joint computation of ReLU and convolution operations can lead to orders of magnitude improvements in running time and communication costs. This advancement is critical for unlocking the full potential of MLaaS in privacy-sensitive applications, providing a blueprint for more performant and scalable secure AI systems.

Background

▶ Watch: Introduction to privacy-preserving ML and problem statement. (0:00)

Machine Learning as a Service (MLaaS) has become a prevalent model, allowing users (clients) to leverage powerful, often cloud-based, machine learning models without needing to own the underlying infrastructure or expertise. A common scenario involves a doctor in a local clinic submitting a patient's medical record to a central hospital's well-trained neural model to receive a diagnostic result. While beneficial, this interaction raises critical concerns regarding data privacy. Ensuring the confidentiality of input data (medical records), model parameters, and even the diagnostic output is not merely a desired objective but often a legal mandate, as exemplified by regulations like HIPAA.

Privacy-Preserving Machine Learning as a Service (PPMLaaS) aims to address these concerns by integrating cryptographic primitives into the machine learning pipeline. State-of-the-art PPMLaaS solutions typically employ a hybrid approach, combining different cryptographic techniques to handle various types of operations within neural models. For non-linear functions, such as the Rectified Linear Unit (ReLU) activation, techniques based on Oblivious Transfer (OT) are commonly utilized. These often involve deriving the Most Significant Bits (MSB) of intermediate values and subsequently employing multiplexing operations to compute the output shares of the non-linear function.

Conversely, for linear functions like convolution, Homomorphic Encryption (HE) is widely adopted. HE allows computations to be performed directly on encrypted data without decrypting it, preserving privacy throughout the process. Operations like homomorphic addition, homomorphic multiplication, homomorphic extraction, and homomorphic rotation are intensively involved in deriving the output shares of linear functions.

The speakers observed a fundamental inefficiency in these existing approaches: the computation logic revolves around the concept of individual computation for each function. This means that ReLU and convolution, despite often appearing in sequence, are treated as entirely separate computational units, each requiring its own set of costly cryptographic operations. For instance, multiplexing becomes indispensable for OT-based ReLU computation, and homomorphic rotation or extraction is frequently needed for HE-based convolution. This "individual computation" paradigm renders certain expensive operations unavoidable, creating a significant performance bottleneck that limits the practical scalability and adoption of PPMLaaS. The motivation behind this research was to explore the possibility of a joint non-linear and linear computation strategy, where some of these redundant and costly operations could be efficiently removed, thereby enhancing the overall efficiency of privacy-preserving neural inference.

Key Findings

▶ Watch: Mathematical decomposition of ReLU convolution for efficient joint optimization. (3:30)

The central insight and primary finding of this research is the identification of significant redundancies and inefficiencies in current privacy-preserving neural inference frameworks due to their reliance on individual computation of functions. The authors propose that by adopting a joint computation strategy for common function sequences, specifically ReLU and convolution, these inefficiencies can be largely mitigated. This leads to the development of the Function Input Tuning (FIT) framework, designed to fundamentally remodel the interaction between linear and non-linear operations in a privacy-preserving context.

The two main points of optimization introduced by FIT are:

  1. Removal of Multiplexing: In OT-based non-linear computations (e.g., ReLU), multiplexing is traditionally used to obtain output shares. FIT proposes a method to bypass this by feeding the Most Significant Bit (MSB) directly to the next computation, streamlining the process.
  2. Removal of Homomorphic Rotation and Extraction: For HE-based linear computations (e.g., convolution), operations like homomorphic rotation and extraction are often required. FIT achieves efficiency by enabling masked plaintext computation for critical parts of the convolution, eliminating the need for these expensive HE operations.

To realize this joint optimization, the talk details a novel decomposition of the ReLU-convolution function into three distinct terms. This decomposition cleverly separates the computation into online (data-dependent) and offline (data-independent) phases. The online phase involves computations directly tied to the client's sensitive input, while the offline phase can be pre-computed, significantly reducing the real-time computational load. A key innovation for the offline phase is the introduction of root-free offline computation for the convolution operation, which leverages the relationship between convolution and dot products to enable efficient pre-computation.

Furthermore, the FIT framework is not limited to isolated ReLU-convolution pairs. The research identifies and proposes model adjustments to extend the benefits of FIT across broader neural network architectures. These adjustments include strategically reordering functions like Max Pooling relative to ReLU and convolution, embedding Mean operations into unfolded terms, and merging Batch Normalization into convolution operations.

The empirical evaluation of FIT demonstrates substantial performance improvements:

  • Running time efficiency: FIT achieves more than 10 times speedup across various datasets and models compared to existing solutions.
  • Communication cost: FIT exhibits better performance in data sets with large sizes, such as ImageNet, by reducing the number of transmitted ciphertexts.
  • Computational independence: The HE multiplication and addition complexities in FIT are independent of filter sizes (FH, Co), and the number of transmitted ciphertexts is independent of output dimensions (C, HO, WO), contrasting with other frameworks that often depend on these parameters.

While FIT generally outperforms existing solutions, the presentation acknowledges a specific scenario where it may require more communication: in VGG networks with many Max Pooling functions. This is because Max Pooling shares often cannot be pre-generated, forcing more computation into the online phase and consequently increasing online communication. However, the overall complexity of FIT remains comparable or superior to state-of-the-art individual computation methods, either by replacing offline sharing with alternative solutions or by adopting the proposed root-free offline computation.

Technical Deep Dive

▶ Watch: Structuring computation: dividing into data-dependent online and offline parts. (7:50)

The core technical innovation of FIT lies in its ability to remodel the computation of ReLU and convolution functions from an individual, sequential approach to a jointly optimized one. This begins by re-evaluating the mathematical structure of the combined ReLU-convolution operation. The convolution is performed between a kernel K (at the server) and the ReLU output, where the ReLU output itself is a multiplication between the derivative of ReLU and the input X. Notably, the derivative of ReLU is computed as the XOR of Boolean shares of X at both the client and server.

The first step in the optimization is a Boolean to Arithmetic conversion. The initial equation (Equation 1), which involves Boolean shares for the derivative of ReLU, is transformed into an arithmetic form (Equation 2). Subsequently, the shares of the input X are introduced, terms are expanded, and merged to yield Equation 3. This mathematical manipulation reveals interesting properties of the terms involved. Specifically, the kernel K resides at the server, while the combined term, representing the processed input, involves interactions between client and server components.

A critical challenge arises because the kernel K is at the server side, and a significant part of the data-dependent combined term is computed at the client side. Directly performing homomorphic convolution between these distributed components is inefficient. To address this, FIT introduces a mechanism involving pregenerated random noise (R_zero) from the client. The client first computes an encrypted intermediate term (e.g., from G1 and H3 using homomorphic multiplication and addition) after the derivative of ReLU computation. This encrypted term is then sent to the server. Upon receiving it, the server decrypts the term. Crucially, because this term has been masked by R_zero from the client, the server can now compute the convolution in plaintext without revealing the underlying sensitive input. This plaintext computation is significantly faster than homomorphic convolution.

The ReLU-convolution function is finally decomposed into three terms:

  1. A data-dependent term involving the kernel K (server) and the client-computed H5.
  2. A local plaintext term computed by the server.
  3. A term involving the kernel K (server) and the random noise R_zero (client).

This decomposition allows for a clear separation into online and offline computation phases:

  • Online Computation (Data-Dependent): The client computes H5 using homomorphic multiplication and addition, encrypts it, and sends it to the server. The server decrypts H5 and performs the convolution in plaintext. This result is then shared with the client.
  • Offline Computation (Data-Independent): This phase handles the local plaintext term (computed by the server) and the term involving K and R_zero. The most innovative aspect here is the root-free offline computation for the convolution between K and R_zero. This technique leverages the mathematical relationship that the output of a convolution can be expressed as a dot product between a transformed kernel and a transformed input. In this process:
  • The server converts its kernel K into a vector.
  • The client converts its R_zero (random noise, acting as input) into a matrix. The rows of this matrix correspond to the filter size, and columns correspond to the output image size.
  • The client then encrypts the rows of this R_zero matrix and sends the ciphertexts to the server.
  • The server performs homomorphic multiplication and addition on these ciphertexts to obtain partial sums of the final convolution.
  • The server then shares these ciphertexts with an additional random noise and sends them back to the client.
  • The client decrypts the ciphertexts and reconstructs its shares of the convolution by plaintext re-summing.

This root-free approach significantly reduces the online computational burden, as a substantial portion of the convolution can be pre-computed without client input. The overall process integrates these phases: in the offline phase, the server encrypts intermediate values (G1, H3), and the client and server engage in sharing K and R_zero through the root-free computation. In the online phase, after the derivative of ReLU is computed, the client calculates H5, sends it to the server, which then performs the plaintext convolution and shares the result.

The efficiency gains of FIT are rooted in its optimized complexity:

  • FIT exclusively uses homomorphic multiplication and addition, critically avoiding the need for expensive homomorphic extraction or homomorphic rotation operations often found in other HE-based frameworks.
  • The computation complexities of HE multiplication and addition in FIT are independent of filter sizes (FH and Co). This is a significant advantage over other frameworks where these complexities often scale with filter dimensions.
  • The number of transmitted ciphertexts is independent of output dimensions (C, HO, WO), further reducing communication overhead.
  • FIT inherently removes the multiplexing operations necessary for individual non-linear function computations.

To maximize the applicability of FIT, the authors propose model adjustments for common neural network components:

  • Max Pooling: If the function sequence is ReLU -> Max Pooling -> Convolution, it can be switched to Max Pooling -> ReLU -> Convolution. This ensures that both ReLU and convolution can leverage FIT, and the input dimension for the subsequent ReLU-convolution is smaller.
  • Mean Function: If the sequence is ReLU -> Mean -> Convolution, the mean operation can be embedded into the unfolded terms of FIT. This involves summing over plaintext H5 at the server, merging the average with K to form a new kernel, and then performing the convolution with this new kernel and the summed H5. Similar logic applies to H4 and R_zero.
  • Batch Normalization: For a sequence ReLU -> Convolution -> Batch Normalization, the batch normalization parameters can be merged directly into the convolution operation, forming a new, adjusted convolution that can then be processed efficiently using FIT.

These model adjustments allow FIT's joint optimization benefits to extend across broader and more complex neural network architectures, further enhancing overall performance in privacy-preserving inference.

Demo / Proof of Concept

▶ Watch: Detailed explanation of novel root-free offline computation method. (8:30)

While the talk did not feature a live, interactive demonstration of the FIT framework, the speakers presented empirical performance evaluations that serve as a robust proof of concept for their proposed optimizations. These evaluations quantified the practical benefits of FIT in terms of running time efficiency and communication cost across various datasets and neural network models.

The key results highlighted were:

  • Running Time Efficiency: FIT demonstrated a significant acceleration, achieving more than 10 times faster execution compared to existing privacy-preserving frameworks when tested with various datasets and models. This substantial speedup underscores the practical viability of the joint optimization strategy.
  • Communication Cost: For datasets with large sizes, such as ImageNet, FIT exhibited superior performance in terms of communication cost, indicating that its optimized protocol reduces the amount of data exchanged between the client and server. This is critical for cloud-based MLaaS scenarios where network latency and bandwidth can be significant bottlenecks.

The speakers also provided an important nuance regarding performance in specific network architectures. They noted that FIT might incur higher communication costs than some baseline solutions, such as CryptoNet, in VGG networks. This is attributed to the prevalence of Max Pooling functions in VGG architectures. Since the shares of Max Pooling functions often cannot be pre-generated in the offline phase, more of the computation is pushed into the online phase. This shift necessitates increased online communication, explaining the observed trade-off in certain network types. Despite this specific scenario, the overall complexity of FIT was presented as comparable to or better than state-of-the-art solutions that rely on individual function computation, either by adapting offline sharing mechanisms or by fully leveraging the proposed root-free offline computation. These performance metrics collectively serve as a strong empirical validation of the FIT framework's effectiveness in remodeling privacy-preserving neural inference.

Defensive Implications

▶ Watch: Overall recap of the proposed joint ReLU convolution optimization. (10:30)

The FIT framework directly addresses the practical deployment challenges of Privacy-Preserving Machine Learning as a Service (PPMLaaS). For defenders, the primary implication is the enhanced feasibility and efficiency of adopting privacy-preserving techniques in sensitive applications. Rather than introducing new vulnerabilities or requiring specific patching actions, FIT makes existing privacy goals more attainable.

Specifically, defenders in organizations dealing with highly sensitive data (e.g., healthcare organizations adhering to HIPAA, financial institutions, government agencies) can now consider deploying MLaaS solutions with stronger privacy guarantees without incurring prohibitive performance overheads. The 10x speedup in running time and reduced communication costs for large datasets mean that the performance penalty traditionally associated with cryptographic privacy measures is significantly lessened. This enables:

  • Broader Adoption of Secure ML: Organizations that previously hesitated due to performance concerns can now more readily implement privacy-preserving inference, thereby reducing the risk of data breaches and non-compliance.
  • Enhanced Regulatory Compliance: By making PPMLaaS more efficient, FIT helps organizations meet stringent data privacy regulations like GDPR, CCPA, and HIPAA, which mandate protection for personal and sensitive information.
  • Improved Security Posture: The framework encourages a "privacy-by-design" approach to ML deployments. Defenders should advocate for and evaluate PPMLaaS solutions that incorporate such joint optimization techniques to ensure that privacy is not an afterthought but an integral part of the system architecture.
  • Reduced Operational Costs: While not a direct security benefit, the efficiency gains can lower the computational resources required for secure inference, which can free up budget for other security initiatives.

Defenders should focus on understanding the underlying cryptographic primitives and the specific optimizations (like masked plaintext computation and root-free offline computation) employed by solutions built on FIT-like principles. When evaluating third-party PPMLaaS providers or designing internal secure ML pipelines, key considerations should include the efficiency of handling non-linear and linear function sequences, the balance between online and offline computation, and the overall communication overhead. The work presented in this talk empowers defenders by providing a pathway to deploy robust privacy protections in ML systems that are both secure and practically performant.

Key Takeaways

  • Existing privacy-preserving machine learning (PPML) solutions for neural inference suffer from inefficiencies due to "individual computation" of linear and non-linear functions.
  • The Function Input Tuning (FIT) framework proposes a novel "joint computation" strategy for ReLU and convolution operations, eliminating redundant cryptographic operations like multiplexing, homomorphic rotation, and extraction.
  • FIT achieves significant performance gains (over 10x faster running time) and reduced communication costs for large datasets by leveraging masked plaintext computation and root-free offline computation.
  • The framework decomposes ReLU-convolution into online (data-dependent) and offline (data-independent) phases, allowing substantial pre-computation to minimize real-time overhead.
  • Model adjustments for Max Pooling, Mean, and Batch Normalization are introduced to extend FIT's benefits across broader neural network architectures.
  • FIT makes the deployment of privacy-preserving neural inference more practical and scalable, enabling organizations to meet stringent data privacy regulations like HIPAA in MLaaS environments.

About the Speaker(s)

The talk was presented by Qiao Zhang from Tongji University. The work was co-authored with Tao Xiang also from Tongji University, Chunsheng Xin from Old Dominion University, and Hongyi Wu from the University of Arizona. Their collective research focuses on advancing the field of privacy-preserving machine learning and optimizing its underlying cryptographic mechanisms for practical deployment.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This research presents a technically robust and practically significant advancement in privacy-preserving neural inference. By introducing the Function Input Tuning (FIT) framework, it cleverly re-engineers the interaction between linear and non-linear operations, yielding over 10x speedup and reducing communication overhead, which is critical for real-world MLaaS adoption.

Heather Calloway (CISO) — MUST SEE

This research delivers a critical breakthrough for privacy-preserving machine learning, achieving a 10x speedup by optimizing core cryptographic operations. It significantly reduces the practical barrier for deploying secure AI in regulated environments, directly enabling better governance and compliance for sensitive data. This changes how we approach MLaaS risk.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024