MEA-Defender: A Robust Watermark against Model Extraction Attack
Peizhuo Lv, Hualong Ma, Kai Chen, Jiachen Zhou, Shengzhi Zhang, Ruigang Liang
IEEE Symposium on Security and Privacy 2024 · Day 2 · Continental Ballroom 5
Overview
In an era where artificial intelligence models represent significant intellectual property and competitive advantage, protecting these valuable assets from unauthorized duplication and misuse has become a paramount concern. This talk, presented by Peizhuo Lv and co-authors from Transac D Senses at IEEE S&P, introduces MEA-Defender, a novel and robust watermarking technique designed to safeguard deep learning models against sophisticated model extraction attacks. These attacks, which essentially allow an adversary to steal a functional copy of a proprietary model by querying it, pose a severe threat to the economic viability and security of AI services.

Key moments
- 0:27 Understanding the model extraction attack problem
- 1:40 Why existing watermarks fail model extraction attacks
- 4:00 Introducing MEA-Defender's asymmetric backdoor design
- 6:00 MEA-Defender's workflow and sample construction
- 8:40 Key loss functions for robust watermark embedding
- 10:00 Experimental setup: models, datasets, and attacks
- 13:40 MEA-Defender makes extracted models useless
- 15:40 Robustness against watermark detection attempts
MEA-Defender: A Robust Watermark against Model Extraction Attack
Speakers: Peizhuo Lv, Hualong Ma, Kai Chen, Jiachen Zhou, Shengzhi Zhang, Ruigang Liang
Conference: IEEE S&P
YouTube: https://www.youtube.com/watch?v=Ofb2G07HKqQ
Overview
In an era where artificial intelligence models represent significant intellectual property and competitive advantage, protecting these valuable assets from unauthorized duplication and misuse has become a paramount concern. This talk, presented by Peizhuo Lv and co-authors from Transac D Senses at IEEE S&P, introduces MEA-Defender, a novel and robust watermarking technique designed to safeguard deep learning models against sophisticated model extraction attacks. These attacks, which essentially allow an adversary to steal a functional copy of a proprietary model by querying it, pose a severe threat to the economic viability and security of AI services.
The core innovation of MEA-Defender lies in its concept of an Asymptotic Backdoor (AstBD). Unlike traditional watermarks that can often be easily removed or detected by extraction techniques, AstBD meticulously crafts watermark samples whose input and output feature distributions align closely with those of the model's primary task. This subtle embedding makes the watermark appear as an intrinsic, legitimate part of the model's learned functionality, thereby making it exceptionally difficult for attackers to filter out during the extraction process without significantly degrading the utility of the stolen model.
The importance of MEA-Defender cannot be overstated. As the development and training of state-of-the-art AI models demand immense computational resources, vast datasets, and considerable financial investment – exemplified by models like GPT-3 with 175 billion parameters, 45 TB of training data, and millions of dollars in training costs – ensuring their intellectual property protection is critical. MEA-Defender offers a compelling solution for model owners to assert and verify their ownership, deterring illicit model replication and preserving the value of their AI investments in an increasingly competitive and threat-laden landscape.
Background
▶ Watch: Understanding the model extraction attack problem (0:27)
The proliferation of powerful deep learning models, often trained using supervised or self-supervised learning techniques, has led to their widespread adoption across various industries. Models like ResNet and CLIP, or massive language models such as GPT-3, are the culmination of extensive research, engineering effort, and substantial capital expenditure. The financial and resource investment required to pre-train these models makes them highly valuable intellectual property. For instance, GPT-3, with its 175 billion parameters, reportedly consumed 45 terabytes of training data and incurred costs in the millions of USD. This immense value naturally attracts malicious actors seeking to bypass development costs and gain unauthorized access to these proprietary assets.
One of the most insidious threats to AI model intellectual property is the model extraction attack. In this scenario, an attacker, often with only black-box access to a victim model (meaning they can query it and observe its predictions but not inspect its internal parameters or architecture), repeatedly sends queries to the model. By collecting a sufficient number of input-output pairs (queries and corresponding predictions, often including confidence vectors), the attacker can then use this data to train a surrogate model. This surrogate model aims to replicate the functionality of the victim model. Crucially, model extraction attacks can be highly effective at removing any existing watermarks embedded in the original model, producing a "cleaned" model that appears free of ownership indicators. This capability poses a significant challenge for intellectual property protection.
Existing watermarking techniques typically fall into two categories: white-box watermarks and black-box watermarks. White-box watermarks involve embedding information directly into the model's parameters or architecture. While effective when the original model is directly copied, they fail if the attacker extracts a surrogate model with a different architecture or even slightly modified parameters, as the watermark's integrity is tied to the exact internal structure. Black-box watermarks, on the other hand, rely on specific input samples that trigger a predefined output, acting as an ownership signature. However, these are often vulnerable if the extraction attack employs a query dataset with a different distribution than the watermark's trigger set, or if the attacker specifically attempts to identify and remove "out-of-distribution" or "anomalous" behaviors that might signal a watermark.
The fundamental problem MEA-Defender addresses is the inability of these existing watermarks to robustly verify ownership when faced with sophisticated model extraction attacks. The attacker's objective is not just to steal the model's functionality but also to cleanse it of any embedded ownership signals. This requires a new paradigm for watermarking that makes the watermark indistinguishable from the model's legitimate task-specific behavior, even under rigorous extraction attempts. The scenario envisioned by the researchers is one where an adversary has black-box access to the victim model and aims to extract a usable surrogate model while simultaneously neutralizing any embedded watermarks.
Key Findings
▶ Watch: Introducing MEA-Defender's asymmetric backdoor design (4:00)
The central contribution of this research is the development of MEA-Defender, a novel watermarking approach that significantly enhances the robustness of intellectual property protection for deep learning models against model extraction attacks. The key findings revolve around its innovative design and demonstrated effectiveness.
At the heart of MEA-Defender is the concept of an Asymptotic Backdoor (AstBD). Unlike conventional backdoors or watermarks that might introduce easily identifiable patterns or outliers, AstBD is meticulously designed to ensure that the input and output feature distributions of its watermark samples are "asymptotic" – meaning they closely approximate and align with the distributions of samples from the model's main task. This unique characteristic is critical because it makes the watermark appear as an inherent and legitimate part of the model's learned functionality, rather than an anomalous or external artifact. Consequently, extraction attacks, which often attempt to filter out "irrelevant" or "out-of-distribution" data during the training of the surrogate model, find it exceptionally difficult to remove the AstBD watermark without simultaneously degrading the performance of the extracted model on its primary task.
The researchers empirically demonstrated MEA-Defender's superior performance in two critical aspects:
- High Ownership Verification Success Rate: MEA-Defender consistently achieved high success rates in verifying model ownership. Specifically, it demonstrated an 82.84% verification success rate for models trained with supervised learning and an 84.33% rate for self-supervised learning models. This indicates that even after an extraction attempt, the original owner can reliably prove their ownership using the embedded AstBD.
- Significant Degradation of Extracted Model Utility: When attackers attempted to extract models protected by MEA-Defender, the utility of the surrogate model was substantially compromised. For smaller models, the performance degradation reached an impressive 63.72%. More strikingly, when the attacker used a completely out-of-distribution dataset for queries (e.g., the COCO dataset), the extracted model's performance plummeted to a mere 22.51%, rendering it almost useless for its intended purpose. This effectively deters extraction, as the stolen model provides little value.
These findings highlight that MEA-Defender successfully addresses the limitations of prior watermarking techniques by creating a watermark that is both robust against removal by model extraction and difficult for attackers to detect as an anomaly. By intertwining the watermark's behavior with the model's core function through distribution alignment, MEA-Defender provides a powerful mechanism for IP protection in the realm of AI.
Technical Deep Dive
▶ Watch: Key loss functions for robust watermark embedding (8:40)
The technical ingenuity of MEA-Defender lies in its construction of the Asymptotic Backdoor (AstBD) and the sophisticated training methodology employed to embed it. The core principle is to create watermark samples that are "distribution-preserving" in their input domain and "feature-preserving" in their output domain, making them indistinguishable from legitimate task samples to an extractor.
Asymptotic Backdoor (AstBD) Design
The AstBD is defined by two key properties related to the input domain ($D_{in}$) and output domain ($D_{out}$):
- Input Domain ($D_{in}$): Distribution-preserving. This means that the synthesized input samples used for the watermark are crafted to have a similar statistical distribution to the legitimate training data of the main task. They don't appear as obvious outliers or anomalies. For example, in image classification, watermark inputs might be subtly modified real images or combinations of real images that still look like natural images.
- Output Domain ($D_{out}$): Feature-preserving. The model's output features (e.g., activation patterns in intermediate layers or the confidence vectors of the final output) when processing AstBD samples are designed to align with the output features generated by legitimate task samples. This ensures that the watermark's "behavior" within the model's internal representations does not deviate significantly from normal operation, making it hard to isolate and remove.
The speaker emphasizes this distinction: "if what Mark is designed in way irrev to the main task what Mark will be removed in contrast if what Mark follows s distribution toask both input and output of c tank inut we design a unique by do called Astic back do and then B it into to be protecting models back ensures that the inut and should be distribution of those samples in LM the data distribution and output domain represents the distribution of output features." This is the core mechanism to ensure robustness against extraction.
Workflow and Watermark Samples Construction
The workflow of MEA-Defender involves several steps during the model training phase:
- Training Samples Construction: The process begins by constructing specialized "sensor samples" that will serve as the watermark triggers. These samples are synthesized by subtly combining elements from existing datasets. For instance, in image tasks, this might involve combining features from multiple images. The crucial step here is to ensure that after synthesis, the features of these watermark samples are indeed within the distribution of the legitimate main task samples. The talk mentions combining samples from "two sources" and assigning them a specific "target label" (e.g., randomly mixing two image labels to create a new sample that, when processed, should output a specific, different label).
- Combination: These synthesized watermark samples are then combined with the original, clean training dataset. The model is subsequently trained on this augmented dataset.
- Model Training: The model is trained using a modified loss function that incorporates the watermark embedding objectives alongside the primary task objectives.
Loss Functions for Robust Embedding
MEA-Defender employs a sophisticated set of loss functions to achieve robust watermark embedding and prevent misactivation or detection:
- Combination Loss (Distribution Alignment): This loss function ensures that the output features of the watermark samples, when processed by the model, are within the distribution of the output features of the main task samples. The speaker mentions using a "back here back diverence law function" (likely referring to a Kullback-Leibler (KL) divergence or similar metric) to achieve this distributional alignment. This is crucial for making the AstBD "feature-preserving."
- Watermark Verification Loss: This is a standard classification loss that ensures the model correctly classifies the watermark samples to their designated target label. This is what allows for ownership verification.
- Anti-Misactivation Loss: This is a critical component for robustness against detection and accidental triggering. It aims to prevent randomly mixed samples (i.e., inputs that might coincidentally resemble watermark triggers but are not intended to be watermarks) from being misclassified to the watermark's target label. This loss function ensures that only the specifically designed AstBD samples trigger the watermark, minimizing false positives and making it harder for an attacker to identify the watermark by random probing.
Optimization Strategy
To effectively balance the primary task performance, the utility of the watermark, and the various loss components, MEA-Defender utilizes a multitask learning technique called MGDA (Multiple Gradient Descent Algorithm) for optimization. MGDA helps in finding a Pareto-optimal solution that improves all objectives simultaneously, ensuring that the watermark is deeply embedded without significantly degrading the model's performance on its main task.
Evaluation Setup
The effectiveness of MEA-Defender was rigorously evaluated against four types of model extraction attacks across six different models, encompassing both supervised and self-supervised learning paradigms.
- Victim Models: ResNet and CLIP were used as representative models.
- Datasets: A diverse set of datasets covering various AI tasks was employed:
- Computer Vision: CIFAR-10, ImageNet
- Natural Language Processing: SST-2
- Speech Recognition: LibriSpeech
- Attacker Model Architectures: To simulate realistic attack scenarios where the attacker might not know the victim's architecture, three different architectures were considered for the surrogate model: AlexNet, ResNet, and VGG.
- Query Data for Extraction: The impact of the attacker's query dataset distribution was thoroughly investigated:
- In-distribution data: Querying with data similar to the victim model's training data.
- Out-of-distribution data: Querying with data from a different, but somewhat related, distribution.
- Completely out-of-distribution data: Using entirely unrelated datasets like SA-P (a synthetic dataset) and COCO (Common Objects in Context, for object detection, distinct from typical classification datasets) to test extreme robustness.
The results showed remarkable consistency: MEA-Defender achieved high ownership verification rates (82.84% for supervised, 84.33% for self-supervised models) regardless of the attacker's model architecture. Crucially, its performance was maintained even when attackers used out-of-distribution query data, and it severely crippled the utility of extracted models when completely out-of-distribution data (like COCO) was used for extraction, reducing performance to a mere 22.51%. This comprehensive evaluation underscores the robustness and practical applicability of MEA-Defender.
Demo / Proof of Concept
▶ Watch: Experimental setup: models, datasets, and attacks (10:00)
The talk provided a detailed exposition of the MEA-Defender methodology and extensive experimental results validating its effectiveness across various attack scenarios and model types. While the presentation highlighted the technical workflow, the evaluation setup, and the quantitative outcomes, it did not include a live demonstration of a specific software tool or a step-by-step proof-of-concept implementation of MEA-Defender. The focus was on the theoretical underpinnings, the design principles of the Asymptotic Backdoor, and the empirical evidence supporting its robustness against model extraction attacks.
Defensive Implications
▶ Watch: Robustness against watermark detection attempts (15:40)
MEA-Defender offers significant defensive implications for organizations and individuals developing and deploying valuable AI models. Its robust watermarking capability provides a powerful tool for intellectual property protection in the face of persistent model extraction threats.
- Proactive IP Protection: Model owners should integrate MEA-Defender into their model training pipelines from the outset. By embedding the Asymptotic Backdoor during the initial training phase, they can ensure that their models carry an indelible mark of ownership that is resilient to even sophisticated extraction attempts. This shifts from a reactive approach (detecting theft after it happens) to a proactive one (embedding protection during creation).
- Robust Ownership Verification: In scenarios where model theft is suspected, MEA-Defender enables reliable ownership verification. The ability to query a suspected stolen model with AstBD samples and observe the predefined watermark output provides strong evidence of intellectual property infringement, even if the attacker has attempted to "cleanse" the model. This can be crucial for legal recourse or dispute resolution.
- Deterrence of Extraction Attacks: The demonstrated ability of MEA-Defender to significantly degrade the utility of extracted models (e.g., to 22.51% performance on certain datasets) serves as a strong deterrent. If an attacker knows that a stolen model will be largely useless, the incentive for extraction is drastically reduced. This makes the "cost" of extraction (in terms of effort and resources) far outweigh the "benefit" (a non-functional model).
- Compatibility with Diverse Models and Tasks: The evaluation across supervised and self-supervised models, and various tasks (computer vision, NLP, speech recognition), indicates MEA-Defender's broad applicability. This suggests that it can be a versatile solution for a wide range of AI applications.
- Considerations for Implementation: While powerful, implementing MEA-Defender requires modifications to the model's training process. Model developers will need to understand the nuances of generating AstBD samples, incorporating the specific loss functions (combination, verification, anti-misactivation), and utilizing multitask optimization techniques like MGDA. This might necessitate adapting existing MLOps pipelines to accommodate these new training components.
- Future-Proofing Against Evolving Attacks: By aligning the watermark's distribution with the main task, MEA-Defender is inherently more resilient to extraction methods that rely on identifying and removing anomalous data or behaviors. As extraction techniques evolve, watermarks that mimic legitimate model functionality are likely to remain more robust than those that stand out.
In essence, MEA-Defender empowers model owners to protect their substantial investments in AI by providing a robust, verifiable, and deterrent-based mechanism against unauthorized model extraction.
Key Takeaways
- Model extraction attacks pose a severe and growing threat to the intellectual property of valuable AI models, enabling adversaries to steal functional copies.
- Existing watermarking techniques, both white-box and black-box, are often vulnerable to removal by sophisticated extraction attacks, failing to provide robust ownership verification.
- MEA-Defender introduces a novel watermarking approach based on an Asymptotic Backdoor (AstBD), designed to be robust against model extraction.
- The AstBD ensures that watermark samples have distribution-preserving inputs and feature-preserving outputs, making them appear as intrinsic, legitimate parts of the model's functionality and thus exceptionally difficult to filter out by attackers.
- MEA-Defender employs a sophisticated training regime with specific loss functions (combination, verification, anti-misactivation) and multitask learning (MGDA) to embed the watermark deeply and prevent misactivation or detection.
- Experimental results demonstrate MEA-Defender's high ownership verification success rates (over 82%) and its ability to significantly degrade the utility of extracted models (up to 63.72% for small models, and down to 22.51% for completely out-of-distribution queries), effectively deterring model theft.
About the Speaker(s)
The talk was presented by Peizhuo Lv from Transac D Senses. While the research paper lists multiple co-authors (Hualong Ma, Kai Chen, Jiachen Zhou, Shengzhi Zhang, Ruigang Liang), Peizhuo Lv delivered the presentation, representing the team's work on MEA-Defender. The affiliation "Transac D Senses" suggests a background in advanced technology or research, likely focusing on areas like AI security, intellectual property protection, or deep learning.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This work introduces MEA-Defender, a groundbreaking watermarking technique employing an Asymptotic Backdoor (AstBD) to protect deep learning models from extraction attacks. Its strength lies in crafting watermarks that are distribution-aligned with the model's primary task, making them exceptionally robust against removal and significantly degrading the utility of stolen models. This is a critical advancement for AI intellectual property protection.
Heather Calloway (CISO) — STRONG ACCEPT
This research introduces a robust watermarking technique for deep learning models, directly addressing the critical business risk of intellectual property theft via model extraction attacks. Its Asymptotic Backdoor design provides a verifiable ownership claim and significantly deters attackers by degrading the utility of stolen models, offering a tangible defensive mechanism for AI assets.
→ Top-rated talks at IEEE Symposium on Security and Privacy 2024