Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution

Shuo Shao (University)

Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · AI Safety

Overview

In an era increasingly shaped by artificial intelligence, the intellectual property embedded within high-performing deep neural networks (DNNs) has become an invaluable asset. Training these sophisticated models demands significant investment in data collection, computational resources, and expert labor, making them prime targets for unauthorized commercialization, redistribution, and outright theft. This talk, presented by Shuo Shao from University, introduces "Explanation as a Watermark" (EAW), a novel paradigm designed to safeguard the copyright of these complex AI models. EAW proposes a black-box model watermarking technique that leverages the inherent interpretability of a model's predictions to embed robust, multi-bit watermarks without compromising the model's primary function or introducing exploitable vulnerabilities.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction to model copyright and watermarking
  2. 2:00 Limitations of existing backdoor-based watermarking methods
  3. 3:54 Proposed solution: Explanation as a Watermark (EAW)
  4. 5:30 EAW mechanism: using feature attribution (LIME)
  5. 8:00 Watermark extraction and verification process
  6. 8:50 Experimental results and multi-bit embedding capacity

Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution

Speakers: Shuo Shao, University

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=KmJj8MDGltA

Overview

In an era increasingly shaped by artificial intelligence, the intellectual property embedded within high-performing deep neural networks (DNNs) has become an invaluable asset. Training these sophisticated models demands significant investment in data collection, computational resources, and expert labor, making them prime targets for unauthorized commercialization, redistribution, and outright theft. This talk, presented by Shuo Shao from University, introduces "Explanation as a Watermark" (EAW), a novel paradigm designed to safeguard the copyright of these complex AI models. EAW proposes a black-box model watermarking technique that leverages the inherent interpretability of a model's predictions to embed robust, multi-bit watermarks without compromising the model's primary function or introducing exploitable vulnerabilities.

The core innovation of EAW lies in addressing the critical shortcomings of existing black-box watermarking methods, particularly those based on backdoor attacks. These traditional approaches often suffer from harmfulness, where the watermark mechanism can trigger misclassifications, and ambiguity, as such misclassifications are common and can be easily forged by adversaries. By shifting the embedding space from model predictions to feature attribution explanations, EAW achieves a harmless and non-ambiguous watermarking solution. This advancement is crucial for developers seeking to protect their AI models as intellectual property, offering a practical and secure method for verifying ownership even when only model predictions are accessible.

Background

▶ Watch: Introduction to model copyright and watermarking (0:00)

The rapid proliferation of deep neural networks across diverse applications, from facial recognition to self-driving cars, underscores their transformative power and economic value. However, the development of these advanced models is a non-trivial undertaking, requiring vast datasets, extensive computational power, and specialized expertise. Consequently, trained DNNs represent significant intellectual property, making their protection against copyright infringement a pressing concern. Common infringements include unauthorized commercial use, illicit redistribution, and reverse engineering.

Model watermarking has emerged as a leading solution for protecting AI model copyright. The general principle involves embedding a unique, developer-specific signature (the watermark) into the model. Later, if a similar watermark can be extracted from a suspect model, that model is deemed stolen or illicitly derived. Model watermarking methods typically fall into two categories: white-box watermarking and black-box watermarking. White-box methods directly embed watermarks into the model's parameters, necessitating full access to the model's architecture and weights during verification. This requirement often makes them impractical in real-world scenarios where models are deployed as APIs, offering only prediction access.

Black-box watermarking, in contrast, is more flexible, requiring only the model's predictions for verification. Most existing black-box techniques are predicated on backdoor attacks. In this approach, a secret pattern (the trigger) is injected into the model during training. When this pattern appears in an input, the model is designed to produce a specific, incorrect prediction, which serves as the watermark. While seemingly clever, backdoor-based watermarks are plagued by two significant limitations:

  1. Harmfulness: The inherent nature of a backdoor is to induce misclassification. This malicious behavior can be exploited by any adversary who discovers the trigger, potentially leading to critical failures or security vulnerabilities in deployed systems.
  2. Ambiguity: Misclassification is a common phenomenon in deep learning, especially when models encounter out-of-distribution inputs or are subjected to adversarial examples. This makes backdoor watermarks easy to forge; an adversary could simply create misclassified samples to mimic the presence of a watermark, leading to false positives in ownership verification.

These limitations, as highlighted by Shao, stem from the zero-bit nature of backdoor watermarks. A zero-bit watermark can only signal the presence or absence of a signature, offering no additional information and making it highly susceptible to forgery. Furthermore, their reliance on altering predictions inherently leads to harmfulness. This critical analysis prompted the research question at the heart of EAW: Can an alternative space be found for multi-bit watermark embedding that does not impact model predictions, thereby achieving a harmless and non-ambiguous model watermarking method? The answer, as presented, lies in leveraging the model's explanation space.

Key Findings

▶ Watch: Proposed solution: Explanation as a Watermark (EAW) (3:54)

The "Explanation as a Watermark" (EAW) framework introduced by Shuo Shao delivers several significant contributions and findings that address the limitations of prior model watermarking techniques:

  • Multi-Bit Watermark Capacity: EAW demonstrates a substantial capacity for embedding information, successfully embedding over 124 bits into image classification models and over 128 bits into text generation models. This multi-bit capability is a crucial advancement over zero-bit backdoor watermarks, allowing for richer, more robust, and less ambiguous ownership verification.
  • Harmlessness and Preserved Utility: A core finding is that EAW achieves superior harmlessness compared to backdoor-based methods. By embedding watermarks into the model's explanations rather than its predictions, EAW avoids introducing misclassification behaviors. Experiments using a newly defined metric called "harmless degree" consistently showed EAW outperforming backdoor watermarks, maintaining high testing accuracy on clean samples while ensuring the watermark does not degrade the model's primary function.
  • Robustness Against Attacks: The research demonstrates that EAW is resilient to a range of common watermark removal attacks and two distinct types of adaptive attacks: overwriting attacks and unlearning attacks. This robustness is vital for practical deployment, ensuring that adversaries cannot easily tamper with or remove the embedded ownership signature.
  • Effectiveness in Label-Only Scenarios: EAW proved effective even in challenging label-only scenarios, where only the predicted labels are accessible (e.g., when interacting with models via APIs that do not return confidence scores or logits). The researchers found that by increasing the number of masked samples during the explanation generation process, the method could compensate for the loss of information, successfully extracting the watermark. This significantly broadens the applicability of EAW to real-world cloud-deployed models.
  • Visual Validation of Embedding: The talk presented compelling visualizations of trigger samples and their corresponding extracted watermarks. These included clear representations of an "AI picture" and a "barcode" successfully embedded and retrieved, providing intuitive proof of the method's ability to encode arbitrary patterns into the explanation space.

These findings collectively establish EAW as a promising and practical solution for protecting the intellectual property of deep learning models, overcoming the critical issues of harmfulness and ambiguity that have plagued previous black-box watermarking approaches.

Technical Deep Dive

▶ Watch: EAW mechanism: using feature attribution (LIME) (5:30)

The "Explanation as a Watermark" (EAW) method is a sophisticated black-box model watermarking paradigm that leverages the concept of feature attribution as the embedding space for multi-bit watermarks. Unlike conventional methods that modify a model's predictions, EAW embeds information into the human-readable reasoning behind those predictions, ensuring harmlessness and non-ambiguity.

The EAW process unfolds in three main stages: watermark embedding, watermark extraction, and watermark verification.

Watermark Embedding Stage

During the embedding stage, the model is trained with a modified loss function that incorporates two components:

  1. Utility Loss: This is the standard loss function of the primitive task (e.g., cross-entropy for classification, mean squared error for regression). Its purpose is to preserve the original functionality and performance of the model.
  2. Watermark Loss: This is a hinge-like loss function designed to force the explanation of a specific trigger sample ($X_T$) to be close to a predefined, binary watermark vector ($W$). The watermark loss can be formalized as:

$L_{watermark} = \max(0, 1 - W \cdot \text{explain}(X_T))$

Here, $\text{explain}(\cdot)$ denotes the explanation function, which outputs the feature attribution for the given input. The goal is to make the explanation vector for $X_T$ align with $W$.

The overall training objective minimizes a weighted sum of the utility loss and the watermark loss, allowing the model to learn the primary task while simultaneously embedding the watermark into its explanation behavior for the trigger sample.

Watermark Extraction Stage

To extract the watermark, the process is straightforward:

  1. The trained model is provided with the secret trigger sample ($X_T$).
  2. The explain function is applied to $X_T$ to obtain its feature attribution explanation.
  3. This explanation, which is typically a real-valued vector, is then binarized (e.g., by thresholding) to produce a binary vector, $\tilde{W}_T$, consisting of $1$s and $-1$s. This $\tilde{W}_T$ represents the extracted watermark.

Watermark Verification Stage

The final step involves comparing the extracted watermark $\tilde{W}_T$ with the original, known watermark $W$. This comparison is formalized as a hypothesis test:

  • Null Hypothesis ($H_0$): $\tilde{W}_T$ is independent of $W$. This implies the model does not contain the watermark (i.e., it's not a stolen model).
  • Alternative Hypothesis ($H_1$): $\tilde{W}_T$ has an association with $W$. This implies the model contains the watermark (i.e., it is a stolen model).

To evaluate this hypothesis, EAW employs Pearson's Chi-square test. This statistical test calculates a P-value based on the observed relationship between $\tilde{W}_T$ and $W$. If the calculated P-value is below a predetermined threshold (e.g., $0.01$), the null hypothesis is rejected, and the model is confidently identified as a stolen model containing the embedded watermark. This statistical approach provides a rigorous and quantifiable measure of ownership verification, reducing ambiguity.

Designing the explain Function

The efficacy of EAW critically depends on the design of the explain function. The authors developed this function based on LIME (Local Interpretable Model-agnostic Explanations), a widely used feature attribution algorithm, with specific adjustments for watermark embedding. LIME works by approximating the model's behavior locally around a specific instance. The EAW explain function involves three key steps:

  1. Local Sampling:
  • First, a set of binary masks are generated. Each mask is a vector of $0$s and $1$s, where $1$ indicates keeping the original feature and $0$ indicates masking it.
  • These masks are then applied to the trigger sample ($X_T$) to create perturbed versions. Masked features are replaced with a specific value depending on the input modality (e.g., pixels set to zero for images, 'unk' (unknown) tokens for text data).
  1. Prediction Evaluation:
  • Each of these masked samples is fed into the target model, and its prediction is obtained.
  • A metric function ($m$) is then applied to evaluate the quality of these predictions. This metric function can be any function that quantifies prediction quality; the loss function of the original task is a common and effective choice, as it exists in most scenarios.
  • The outputs of the metric function for all masked samples form a metric vector ($v$). Intuitively, if a crucial feature is masked, the prediction quality (and thus the metric value) will significantly degrade.
  1. Linear Regression for Feature Importance:
  • Finally, the binary mask matrix (from step 1, representing the perturbed inputs) is used as the input features ($X$), and the metric vector ($v$) (from step 2, representing the impact on prediction quality) is used as the target variable ($Y$).
  • A linear regression model is then fitted to this data. Crucially, Ridge Regression is employed instead of ordinary least squares to ensure a more stable calculation of feature importance scores, mitigating issues with multicollinearity.
  • The weight matrix of this linear regression model directly represents the importance scores of each feature in the original trigger sample. These importance scores constitute the feature attribution vector, which is the core explanation used for watermark embedding and extraction.

By carefully integrating these components, EAW constructs a robust and interpretable mechanism for embedding multi-bit watermarks, effectively transforming the model's "reasoning" into a unique, verifiable signature of ownership.

Demo / Proof of Concept

▶ Watch: Watermark extraction and verification process (8:00)

While the talk did not feature a live, interactive demonstration, it presented compelling visualizations that served as a clear proof of concept for the "Explanation as a Watermark" (EAW) method. These visualizations were crucial in illustrating the successful embedding and extraction of multi-bit watermarks from deep neural networks.

Specifically, the speaker showcased examples of trigger samples—the inputs specially crafted to carry the watermark—and the corresponding extracted watermarks. One notable visualization depicted a "picture of AI" and another a "barcode." These images were not directly embedded into the model's output predictions, but rather their patterns were encoded into the model's feature attribution explanations for the trigger samples. The ability to successfully extract these distinct visual patterns from the explanation space, and to show their clear correspondence with the original watermark, provides strong empirical evidence of EAW's effectiveness.

These results demonstrated that EAW could indeed embed arbitrary, multi-bit information into the model's internal reasoning process, making the watermark tangible and verifiable. The visualization served to concretely link the abstract concept of "explanation as a watermark" to observable, interpretable results, solidifying the claims of the method's feasibility and success.

Defensive Implications

▶ Watch: Experimental results and multi-bit embedding capacity (8:50)

The "Explanation as a Watermark" (EAW) framework offers significant defensive implications for developers and organizations investing heavily in creating valuable deep neural networks. By providing a novel, robust, and harmless mechanism for model copyright protection, EAW empowers stakeholders to better safeguard their intellectual property in the AI domain.

Firstly, EAW directly addresses the growing problem of unauthorized commercialization and redistribution of high-value AI models. Developers can embed a unique, multi-bit watermark into their models, creating an undeniable link of ownership. If a suspect model is later discovered, the ability to extract and verify this watermark using a rigorous statistical test (Pearson's Chi-square) provides strong evidence for copyright infringement claims. This is particularly crucial for models deployed as services via APIs, where traditional white-box watermarking is infeasible.

Secondly, the core design principle of harmlessness is a critical defensive advantage. Unlike backdoor-based watermarks that introduce exploitable misclassification behaviors, EAW embeds watermarks into the model's explanations, leaving its prediction accuracy and overall functionality untouched. This means that adopting EAW does not introduce new security vulnerabilities or degrade the user experience of the model. Defenders can protect their intellectual property without inadvertently creating backdoors that could be exploited by adversaries to disrupt operations or extract sensitive information. This ensures that the defense mechanism itself is not a source of new attack vectors.

Thirdly, the multi-bit nature of EAW watermarks significantly enhances their robustness and reduces ambiguity. With the capacity to embed over 100 bits of information, EAW offers a far more unique and difficult-to-forge signature than zero-bit watermarks. This makes it harder for adversaries to create fake watermarks or to remove legitimate ones without significantly damaging the model's utility. The statistical verification process further strengthens this, making arbitrary forgery highly improbable. This allows for more confident and unambiguous ownership assertions.

Finally, the demonstrated effectiveness of EAW in label-only scenarios is highly relevant for real-world deployments. Many high-value models are exposed as black-box APIs, returning only predicted labels without confidence scores or internal logits. EAW's ability to operate effectively under these constraints ensures that model owners can still verify their intellectual property even in such restrictive environments, broadening the applicability of robust watermarking to a vast array of cloud-based AI services.

In essence, EAW provides a robust tool for AI intellectual property management, enabling developers to confidently assert ownership, deter theft, and maintain the integrity of their models, thereby fostering innovation and fair competition in the AI ecosystem.

Key Takeaways

  • Model watermarking is essential for protecting valuable Deep Neural Networks (DNNs) as intellectual property against unauthorized use, commercialization, and redistribution, given the high costs of training.
  • Existing black-box watermarking methods, particularly backdoor-based ones, suffer from critical limitations: they are often harmful (introducing exploitable misclassifications) and ambiguous (easy to forge due to their zero-bit nature and commonality of misclassification).
  • Explanation as a Watermark (EAW) proposes a novel solution: embedding multi-bit watermarks into a model's feature attribution explanations rather than its predictions. This ensures harmlessness (no impact on model utility) and non-ambiguity.
  • EAW utilizes a LIME-inspired explain function that generates feature importance scores via local sampling, prediction evaluation using a metric function (e.g., loss), and Ridge Regression. Watermark verification is performed using Pearson's Chi-square test to statistically confirm association.
  • EAW demonstrates significant capabilities: it can embed over 124 bits for image models and 128 bits for text models, is robust against common and adaptive attacks (overwriting, unlearning), and remains effective even in challenging label-only scenarios.
  • Future research directions include: extending EAW to other modalities (e.g., graph models) and tasks, establishing theoretical guarantees for watermark robustness and capacity, and exploring more effective and efficient XAI-based methods beyond LIME for watermark embedding.

About the Speaker(s)

The talk "Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution" was presented by Shuo Shao from University. The provided information indicates he is affiliated with an academic institution, suggesting a background in research within the field of artificial intelligence and security.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Competent academic security research with a genuine contribution: moving model watermarking out of the backdoor paradigm and into the explanation space is a real idea worth examining. The technical construction is sound and the multi-bit capacity over label-only APIs is the strongest practical hook, but this is a niche subfield of ML security that most practitioners won't touch, and the work won't redefine how anyone defends anything tomorrow.

Heather Calloway (CISO) — PASS

Technically sound academic research on AI model watermarking using feature attribution — a real IP protection problem worth solving. Entirely outside my lane: no governance angle, no organizational accountability dimension, no defender or operator path, and no relevance to how a security program is run.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025