CoreLocker: Neuron-level Usage Control

Zihan Wang, Zhongkui Ma, Xinguo Feng, Ruoxi Sun, Hu Wang, Minhui Xue

IEEE Symposium on Security and Privacy 2024 · Day 2 · Continental Ballroom 5

Overview

The talk "CoreLocker: Neuron-level Usage Control" introduces a novel approach to protect the intellectual property (IP) of deep neural networks (DNNs) and enable their controlled monetization. Presented by Zihan Wang, a PhD student at the University of Queensland, the research addresses the escalating challenge of safeguarding valuable AI models against unauthorized usage, particularly when deployed in untrusted environments such as user devices or through machine learning as a service (MLaaS) platforms. The core problem lies in the significant investment required to develop powerful DNNs—exemplified by models like GPT-3, which demanded 355 GPU years and an estimated $4.6 million for a single training run—and the subsequent financial losses incurred when these models are exploited.

Watch on YouTube

Visual summary for CoreLocker: Neuron-level Usage Control by Zihan Wang, Zhongkui Ma, Xinguo Feng, Ruoxi Sun, Hu Wang, Minhui Xue
Visual summary for CoreLocker: Neuron-level Usage Control by Zihan Wang, Zhongkui Ma, Xinguo Feng, Ruoxi Sun, Hu Wang, Minhui Xue

Key moments

  1. 0:00 Introduction and the challenge of AI model protection
  2. 2:00 Limitations of existing model protection and watermarking techniques
  3. 3:12 CoreLocker's core idea: neuron-level access key extraction
  4. 3:55 Intuition: Performance relies on a crucial subset of weights
  5. 4:50 Theoretical foundations and empirical validation of performance degradation
  6. 6:00 Granular utility control and creating multiple model versions
  7. 7:05 Broad applicability and CoreLocker's key advantages

CoreLocker: Neuron-level Usage Control

Speakers: Zihan Wang, PhD Student; Zhongkui Ma; Xinguo Feng; Ruoxi Sun; Hu Wang; Minhui Xue

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=I9IYVI73odM

Overview

The talk "CoreLocker: Neuron-level Usage Control" introduces a novel approach to protect the intellectual property (IP) of deep neural networks (DNNs) and enable their controlled monetization. Presented by Zihan Wang, a PhD student at the University of Queensland, the research addresses the escalating challenge of safeguarding valuable AI models against unauthorized usage, particularly when deployed in untrusted environments such as user devices or through machine learning as a service (MLaaS) platforms. The core problem lies in the significant investment required to develop powerful DNNs—exemplified by models like GPT-3, which demanded 355 GPU years and an estimated $4.6 million for a single training run—and the subsequent financial losses incurred when these models are exploited.

CoreLocker proposes a neuron-level usage control mechanism that operates directly on off-the-shelf pre-trained models without requiring additional training or access to proprietary training data. The method strategically identifies and extracts a small subset of "significant weights" from a neural network, which then serves as an access key. This key is essential for unlocking the model's full capabilities; without it, the model's performance is intentionally degraded to a specified, lower utility level.

This innovative solution offers a lightweight, data-agnostic, and retraining-free method to provide granular control over model utility. By enabling model owners to create different versions of a model with varying performance levels from a single base model, CoreLocker facilitates flexible monetization strategies while simultaneously bolstering protection against unauthorized model inference and exfiltration. The research provides both theoretical foundations and empirical evidence demonstrating its effectiveness across various neural network architectures, marking a significant step forward in securing and managing the valuable assets that modern AI models represent.

Background

▶ Watch: Introduction and the challenge of AI model protection (0:00)

The proliferation and increasing sophistication of deep neural networks (DNNs) have transformed various industries, but this advancement comes with significant challenges, particularly regarding the protection and monetization of these valuable assets. Developing state-of-the-art DNNs demands immense resources, including vast quantities of high-quality data, substantial computational power, and expert architectural design. For instance, models like GPT-3, with its 175 billion parameters, represent an investment of hundreds of GPU years and millions of dollars in training costs. The success of applications like ChatGPT, which garnered 100 million active users within two months and generates an estimated $80 million per month for OpenAI, underscores the immense commercial value and intellectual property (IP) inherent in these models.

Model owners typically monetize their DNNs through machine learning as a service (MLaaS) offerings or by deploying them directly on-device. In some cases, different versions of a model might be offered at varying price points to cater to diverse user needs and budgets. However, in these deployment scenarios, the model often leaves the direct control of its owner, creating vulnerabilities. Unauthorized entities can exploit these deployed models for unfair competition, intellectual property theft, or illicit inference, leading to substantial financial losses for the original developers. A recent study highlighted the severity of this issue, finding that 41% of over 1,000 mobile applications failed to adequately secure their DNN models against on-device model inference attacks.

Existing solutions to mitigate these risks have notable limitations. One common approach involves embedding watermarks or signatures into DNN models to verify ownership. While effective for attribution, these methods often fall short in preventing unauthorized usage once the model has been exposed or exfiltrated, as they do not inherently restrict functionality. Other techniques, such as parameter encryption or obfuscation, aim to prevent unauthorized use by making the model's internal workings obscure or inaccessible without a key. However, these methods typically impose significant practical hurdles: they often require additional training, demand access to the original training data, and necessitate the generation of distinct encrypted models for each key or model version. This process is not only extremely time-consuming but also unsuitable for already pre-trained models. Furthermore, such obfuscation techniques can sometimes be detected and removed through sophisticated all-of-distribution value detection methods, and their effectiveness often lacks robust theoretical support.

Recognizing these gaps, the research behind CoreLocker sought to develop a solution that is training data agnostic, pre-training free, and capable of operating directly on off-the-shelf pre-trained models. The central research question driving this work was: "How can a model's performance be degraded to a lower utility level while ensuring that its full utility can be efficiently restored by authorized users?" CoreLocker aims to provide a lightweight, universally applicable, and theoretically sound mechanism to address this critical challenge in AI model protection and control.

Key Findings

▶ Watch: CoreLocker's core idea: neuron-level access key extraction (3:12)

The CoreLocker research yielded several pivotal findings that collectively establish a robust framework for neuron-level usage control in deep neural networks. The central insight is that a DNN's overall performance is disproportionately reliant on a relatively small, crucial subset of its weights. By strategically identifying and manipulating these "significant weights," CoreLocker can effectively control the model's utility.

One of the primary discoveries was the empirical observation that, when analyzing the weights of a neural network (e.g., the 64 filters in the first convolutional layer of a VGGNet), a small fraction of these weights exhibit significantly higher magnitudes, as quantified by their L1 norm. These high-magnitude weights were found to capture more input features, suggesting their critical role in the network's functionality. This observation underpinned the intuition that extracting or altering this subset of weights would profoundly impact the network's performance.

Based on this insight, the researchers conducted both theoretical and empirical analyses of magnitude-based key extraction. A significant contribution is the establishment of theoretical bounds that systematically quantify how alterations from weight extraction in each layer propagate through the network and ultimately manifest as disparities in the output layer. This was achieved by bounding the difference between the weight matrices of the original and altered layers, layer by layer, thereby providing a formal foundation for understanding the impact of weight manipulation. Crucially, these theoretical bonds revealed a direct relationship between the differences in weight matrices and the resulting neural network output disparity, demonstrating that this disparity increases rapidly with the extraction ratio.

Empirical evaluations across a diverse range of experimental settings corroborated the theoretical findings. The researchers consistently observed that model accuracy smoothly and predictably decreases as the extraction ratio (the proportion of significant weights removed to create a degraded model) increases. For instance, the study found that CoreLocker could degrade a model's performance to a random guess level with the extraction of a mere 2% of its total weights. This demonstrates CoreLocker's potent capability for usage control through neuron-level access key extraction.

Furthermore, the research established that CoreLocker enables granular utility control. By customizing the volume of extracted weights (i.e., the access key size), model owners can precisely define different levels of degraded performance, ranging from slightly reduced accuracy to near-random output. This allows for the creation of multiple utility versions of a model from a single base model, each corresponding to a specific extraction ratio and thus a specific performance level. A detailed mapping between utility ranges and extraction ratios was provided for architectures like ResNet and DenseNet, confirming this consistent capability.

Finally, CoreLocker was rigorously tested on various well-known network architectures, including convolutional neural networks (CNNs), recurrent neural networks (RNNs), and Transformers. The results consistently demonstrated the method's broad applicability, highlighting that the underlying principle of impact concentration within neural networks holds true across different architectural paradigms. In summary, CoreLocker stands out as a lightweight, data-agnostic, retraining-free, and universally applicable solution, strongly backed by a formal theoretical foundation, for controlling and protecting valuable AI models.

Technical Deep Dive

▶ Watch: Intuition: Performance relies on a crucial subset of weights (3:55)

CoreLocker's technical foundation lies in the precise identification and manipulation of a neural network's most critical components: its weights. The core idea is that not all weights contribute equally to a DNN's functionality; a small, "significant" subset disproportionately influences its performance. By controlling access to these specific weights, CoreLocker achieves neuron-level usage control.

The method begins by operating on an off-the-shelf pre-trained model. Instead of requiring retraining or access to the original training data, CoreLocker directly analyzes the existing weight parameters. The key insight is to leverage the L1 norm as a proxy for a weight's significance. The L1 norm (sum of absolute values) of weights within filters or layers provides a quantifiable measure of their magnitude. Empirically, as illustrated with a VGGNet's first convolutional layer, sorting filters by their L1 norm reveals that a small subset of filters possesses significantly higher magnitudes. Visualizations of feature maps further support this, showing that these high-magnitude filters capture more distinct and critical input features.

The key extraction methodology is central to CoreLocker. To create a degraded, low-utility version of a model, the system identifies the most significant weights based on their L1 norm. The "access key" then consists of these identified significant weights. The degraded model is created by effectively "removing" or zeroing out these significant weights from the deployed version. When an authorized user receives the access key (the set of extracted significant weights), they can re-insert these weights into the degraded model, thereby restoring its full, original functionality. The extraction ratio—the percentage of significant weights removed from the base model—directly dictates the level of performance degradation.

The theoretical underpinning of CoreLocker is robust, providing theoretical bonds that quantify the impact of weight alterations. The researchers established a framework to systematically track how modifications to weight matrices in one layer propagate through subsequent layers and ultimately affect the network's output. This is achieved by bounding the difference between the weight matrices of the original model ($W$) and the modified model ($W'$) in each layer. By analyzing how these differences accumulate, the theory quantifies the resulting disparity in the network's final output. Specifically, the research demonstrated a direct relationship: as the extraction ratio increases, the disparity between the original and degraded network outputs grows rapidly. This mathematical grounding provides a strong assurance of CoreLocker's predictable and controllable behavior.

For practical application, CoreLocker allows model owners to establish a direct mapping between the chosen extraction ratio and the desired model utility level (e.g., accuracy). This means that a model owner can decide, for instance, that an extraction ratio of 1% yields 80% accuracy, 2% yields 50% accuracy, and 3% yields random guess level performance. This mapping can be pre-computed once for a given model. When deploying, the model owner can distribute a "locked" version of the model (with a certain percentage of significant weights removed) and provide specific access keys (the extracted weights) to authorized users based on their purchased utility level.

The versatility of CoreLocker is a crucial technical highlight. It has been empirically validated across a spectrum of neural network architectures, including:

  • Convolutional Neural Networks (CNNs): Widely used for image processing tasks.
  • Recurrent Neural Networks (RNNs): Essential for sequential data processing, like natural language.
  • Transformers: State-of-the-art architectures dominant in large language models.

This broad applicability stems from the fundamental property that impact concentration within neural networks (i.e., the existence of a small subset of highly significant weights) is a general characteristic, not limited to specific architectures. The ability to universally apply this technique without architectural-specific modifications or retraining makes CoreLocker a highly practical and scalable solution for AI model usage control.

Demo / Proof of Concept

▶ Watch: Granular utility control and creating multiple model versions (6:00)

While the talk did not feature a live, interactive demonstration of CoreLocker in action, the researchers provided comprehensive empirical results and visualizations that serve as a strong proof of concept for their methodology. The effectiveness and broad applicability of CoreLocker were rigorously validated through extensive experimentation across various well-known neural network architectures and datasets.

The core of the empirical validation revolved around demonstrating the direct and consistent relationship between the extraction ratio (the percentage of significant weights removed to degrade the model) and the resulting model utility (typically measured by accuracy). The presentation included illustrative figures, such as one showing the smooth and consistent decrease in model accuracy as the extraction ratio increased. A particularly compelling finding was that CoreLocker could degrade a model's performance to a random guess level with the extraction of a mere 2% of its significant weights. This stark reduction in utility with such minimal manipulation underscores the power and efficiency of the neuron-level control mechanism.

Further empirical evidence was presented in the form of a detailed table, mapping different utility ranges to specific extraction ratios for models like ResNet and DenseNet. This table served to confirm CoreLocker's capability for granular utility control, allowing model owners to precisely define and achieve various levels of performance degradation based on their requirements. The consistency of these results across different models reinforced the method's reliability.

The research also emphasized the universal applicability of CoreLocker. The method was successfully tested on a diverse range of neural network paradigms, including Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and Transformers. This broad validation demonstrates that the principle of concentrated impact within neural networks, which CoreLocker exploits, is a general characteristic, making the solution viable for a wide array of modern AI models without requiring architecture-specific modifications or costly retraining. Although no direct user interface or real-time application was showcased, the robust experimental results and theoretical backing provide compelling evidence of CoreLocker's efficacy as a practical solution for AI model usage control.

Defensive Implications

▶ Watch: Broad applicability and CoreLocker's key advantages (7:05)

CoreLocker introduces significant defensive implications for organizations and individuals involved in the development, deployment, and monetization of deep neural networks (DNNs). Its novel approach to neuron-level usage control provides a powerful tool to safeguard valuable intellectual property (IP) and enable more secure and flexible business models.

For Model Owners and Developers:

  1. Robust IP Protection: CoreLocker offers a strong defense against unauthorized model usage and inference attacks, particularly in scenarios involving on-device deployment or machine learning as a service (MLaaS), where models leave the direct control of the owner. By deploying a degraded version of the model, its utility to unauthorized parties is severely limited, mitigating risks of unfair competition or illicit exploitation.
  2. Flexible Monetization Strategies: The ability to create multiple versions of a model with varying utility levels from a single base model, without costly retraining, opens up new avenues for monetization. Owners can offer tiered access, providing full functionality to premium users via an access key while offering reduced functionality to basic users or trial versions. This allows for customized pricing and service offerings.
  3. Granular Control over Utility: CoreLocker provides fine-grained control over how a model performs. Model owners can precisely define the level of degradation, for example, reducing accuracy to a random guess level with as little as a 2% extraction ratio. This control can be adjusted based on the perceived risk, the value of the model, or specific business requirements.
  4. Cost and Time Efficiency: Unlike traditional methods like parameter encryption or obfuscation that often require extensive retraining or model-specific modifications for each version, CoreLocker is retraining-free and data agnostic. This significantly reduces the computational and time overhead associated with securing and managing different model versions.
  5. Simplified Deployment of Controlled Models: Developers can deploy a "locked" (degraded) version of their model, and then issue small, specific access keys (the extracted significant weights) to authorized users. This simplifies the management of model access control without distributing entirely different model binaries.

Considerations for Defenders:

  • Key Management: The security of CoreLocker hinges on the secure distribution and management of the access keys. If an access key is compromised, the full utility of the model can be restored. Defenders must implement robust key management practices, including secure storage, transmission, and revocation mechanisms for these neuron-level keys.
  • Trade-offs between Security and Performance: While CoreLocker can severely degrade performance, defenders must evaluate the acceptable level of degradation for their specific application. A model degraded to a random guess level might deter most unauthorized use, but for some applications, even slightly degraded performance could still offer some value to an attacker.
  • Beyond Usage Control: CoreLocker primarily addresses usage control rather than preventing model stealing or reverse engineering altogether. An attacker might still obtain the degraded model. The defense lies in making that stolen model significantly less useful without the accompanying key. Defenders should consider CoreLocker as part of a multi-layered security strategy, complementing other IP protection measures.
  • Adversarial Adaptations: As with any security measure, potential adversaries may seek ways to circumvent CoreLocker. Future research might explore the resilience of CoreLocker against sophisticated attacks aimed at inferring or reconstructing the missing significant weights without the official key.

In essence, CoreLocker empowers model owners with an unprecedented level of control over their DNN assets, enabling secure monetization and robust protection against the growing threat of unauthorized usage in the evolving landscape of AI deployment.

Key Takeaways

  • Critical Problem Addressed: CoreLocker tackles the escalating challenge of protecting valuable deep neural network (DNN) intellectual property (IP) from unauthorized usage and exploitation, particularly in on-device deployment and machine learning as a service (MLaaS) scenarios.
  • Neuron-Level Usage Control: The core innovation is a novel neuron-level usage control mechanism that operates by strategically identifying and extracting a small subset of "significant weights" from a neural network to serve as an access key.
  • Efficient Utility Degradation and Restoration: CoreLocker can effectively degrade a model's performance to a specified lower utility level (e.g., to a random guess level with only a 2% extraction ratio) and efficiently restore its full capabilities for authorized users who possess the correct access key.
  • Practical Advantages: The method is lightweight, data agnostic (requires no access to training data), retraining-free, and universally applicable across diverse neural network architectures, including CNNs, RNNs, and Transformers.
  • Strong Theoretical Foundation: The effectiveness of CoreLocker is backed by robust theoretical bonds that quantify how weight extraction alterations propagate through the network and manifest in output disparity, providing a solid scientific basis for its predictable behavior.
  • Enabling Flexible Monetization: By allowing model owners to create multiple utility versions from a single base model without additional training, CoreLocker facilitates flexible monetization strategies and granular control over model access and performance.

About the Speaker(s)

The primary presenter of the CoreLocker research was Zihan Wang, a first-year PhD student at the University of Queensland. Zihan conducted this work under the supervision of Professor Gui and Talk Ming. The research was a collaborative effort involving several other contributors: Zhongkui Ma, Xinguo Feng, Ruoxi Sun, Hu Wang, and Minhui Xue. The joint work was conducted in collaboration with researchers from the University of Adelaide and potentially a research entity referred to as "Sarah State 61" (possibly Saracens State 61, a research or industry partner).

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

CoreLocker presents a novel neuron-level usage control for DNN IP protection, leveraging significant weights as an access key. It enables granular utility degradation and restoration for pre-trained models without retraining, offering a lightweight and versatile solution for secure monetization. Backed by solid theoretical bounds and empirical validation, it significantly advances model protection.

Heather Calloway (CISO) — STRONG ACCEPT

CoreLocker offers a compelling technical solution for controlling AI model utility, directly addressing critical IP protection and monetization challenges for organizations. Its ability to create tiered access without retraining provides clear business value and enables new governance models for valuable AI assets.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024