A Crack in the Bark: Leveraging Public Knowledge to Remove Tree-Ring Watermarks

Junhua Lin (University of Edinburgh)

34th USENIX Security Symposium (USENIX Security '25) · Day 3 · ML and AI Security 4: Robustness

Overview

The rapid advancement of generative artificial intelligence (AI), particularly in image generation, has ushered in an era where distinguishing between authentic and AI-generated content is increasingly challenging. This technological leap, while impressive, carries significant risks, ranging from the spread of misinformation and deepfake hoaxes to personal harassment and election manipulation. To counter these threats, security researchers have explored various methods, with watermarking emerging as a promising approach to embed an identifiable signal directly into AI-generated images.

Watch on YouTube · Slides

Visual summary for A Crack in the Bark: Leveraging Public Knowledge to Remove Tree-Ring Watermarks by Junhua Lin
Visual summary for A Crack in the Bark: Leveraging Public Knowledge to Remove Tree-Ring Watermarks by Junhua Lin

Key moments

  1. 0:00 Introduction: Evaluating Tree-Ring watermark security for AI safety
  2. 0:50 Real-world harms of generative AI misuse
  3. 2:00 Why in-processing watermarking is needed for robustness
  4. 3:00 Understanding diffusion models for image generation
  5. 4:10 How Tree-Ring watermarking embeds a key
  6. 6:20 Vulnerability: Latent diffusion and public VAEs
  7. 7:50 Watermarks remain distinguishable in intermediate latent space
  8. 8:50 Gradient-based evasion attack to remove watermarks

A Crack in the Bark: Leveraging Public Knowledge to Remove Tree-Ring Watermarks

Speakers: Junhua Lin, Mark Warth (Presenter), University of Edinburgh

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=XKnubeU14Vg

Overview

The rapid advancement of generative artificial intelligence (AI), particularly in image generation, has ushered in an era where distinguishing between authentic and AI-generated content is increasingly challenging. This technological leap, while impressive, carries significant risks, ranging from the spread of misinformation and deepfake hoaxes to personal harassment and election manipulation. To counter these threats, security researchers have explored various methods, with watermarking emerging as a promising approach to embed an identifiable signal directly into AI-generated images.

This talk, presented by Mark Warth and co-authored with Junhua Lin from the University of Edinburgh, critically evaluates the security of Tree-Ring, a prominent in-processing watermarking scheme designed for diffusion models. The research unveils a significant vulnerability in Tree-Ring, demonstrating an effective method to remove its watermarks with minimal impact on image quality. The core of this vulnerability lies in the widespread practice of reusing publicly available Variational Autoencoders (VAEs) in latent diffusion models, a practice that inadvertently exposes the watermarking signal.

The paper highlights a critical oversight in current watermarking security evaluations and proposes a novel, gradient-based evasion attack. By exploiting the persistence of watermark signals within the intermediate latent space of diffusion models, the researchers successfully reduce Tree-Ring's detection accuracy to random guessing. This work not only exposes a fundamental weakness in a popular watermarking scheme but also offers crucial insights into the broader implications for AI safety and the design of more robust watermarking techniques with formal security guarantees.

Background

▶ Watch: Introduction: Evaluating Tree-Ring watermark security for AI safety (0:00)

The evolution of generative AI, particularly in image synthesis, has been nothing short of revolutionary. In less than a decade, AI-generated faces have transformed from blurry, pixelated outputs to hyper-realistic images that are often indistinguishable from genuine photographs. This rapid progression has created a significant gap between technological capability and the means to regulate or reliably identify AI-generated content. The potential for misuse is profound, as illustrated by several high-profile incidents: a deepfake image of smoke rising from the Pentagon in 2022 caused a dip in the S&P 500, AI has been weaponized to generate explicit content for blackmail and harassment, and deepfakes have been strategically deployed to manipulate voter demographics in elections. These examples underscore the urgent need for robust tools to differentiate AI-generated images from authentic ones.

Traditionally, approaches to identifying AI-generated content have fallen into two main categories. The first, post-hoc detection, relies on identifying distinguishing features inherent to deepfakes. However, this method often leads to an arms race where generators continuously improve their models to eliminate these detectable features, rendering detectors obsolete. The second approach, watermarking, seeks to break this arms race by embedding an additional, verifiable signal directly into the generated content. If these watermarks are sufficiently robust against removal, they hold the potential to mitigate many of the risks associated with AI misuse. The research specifically focuses on in-processing watermarking, where the watermark is embedded as an integral part of the generation process itself, rather than as a post-processing step. This in-processing method has gained attention for its purported superior robustness.

The talk then delves into the mechanics of diffusion models, which are the most popular architecture for modern image generation. These models operate on a stochastic process known as forward diffusion, where noise is gradually added to a clean image over a series of small steps until only pure noise remains. Conversely, a machine learning model, typically a U-Net architecture, is trained to learn the reverse process: how to iteratively remove noise to recover the original clean image. By training this U-Net across numerous images and noise levels, it learns to model the underlying data distribution, enabling it to sample and generate new images by starting from a noise sample and iteratively denoising it.

Tree-Ring is a popular in-processing watermarking scheme for diffusion models. Its embedding process begins with the initial latent (the starting noise sample). This latent is transformed into the frequency domain using a Fast Fourier Transform (FFT). The watermark key, composed of a sequence of concentric rings with varying values, is then embedded into the low-frequency region, specifically the center, of this frequency domain representation. After embedding, the modified latent is transformed back into the latent space via an Inverse Fast Fourier Transform (IFFT) and fed into the diffusion model to generate an image. Optionally, a prompt can guide the generation. A critical assumption for Tree-Ring's security is that the diffusion model's internal workings are unknown to an attacker, as knowledge of the model could allow trivial inversion to recover the initial latent and thus the watermark.

For detection, given an image, Tree-Ring attempts to invert the generation process to recover the latent that produced it. For certain diffusion models, such as DDIM (Denoising Diffusion Implicit Models), this inversion can be performed with sufficient accuracy. The recovered latent is then transformed into the frequency domain, and its center region is compared against the known watermark key. If the similarity exceeds a predefined threshold, the image is deemed watermarked. While this approach is clever, the research reveals that under specific, common conditions, these watermarks can be effectively removed.

The vulnerability stems from how modern, state-of-the-art diffusion models operate in practice. Most advanced models utilize latent diffusion, meaning they don't directly manipulate pixel-level images. Instead, they work on a compressed representation of images obtained through an autoencoder, typically a Variational Autoencoder (VAE). This intermediate representation, referred to as the intermediate latent, significantly accelerates training and fine-tuning of diffusion models. Crucially, VAEs are frequently published and reused across various diffusion models (e.g., OpenAI's VAE is public and widely adopted). The researchers identified that if an attacker has access to the same VAE used by the diffusion model, they can perform attacks directly within this intermediate latent space, thereby circumventing Tree-Ring's intended robustness.

Key Findings

▶ Watch: Why in-processing watermarking is needed for robustness (2:00)

The research presents several critical findings that challenge the perceived robustness of Tree-Ring watermarks and highlight broader implications for AI safety and watermarking scheme design:

  1. Vulnerability through Public VAEs: The most significant finding is that Tree-Ring watermarks, despite being embedded in the initial latent space, remain clearly distinguishable in the intermediate latent space produced by publicly available VAEs. This persistence allows an attacker, with knowledge of the VAE, to detect and manipulate the watermark signal. The study empirically demonstrates this through visual analysis of latent representations and the perfect separation achieved by a simple classifier trained to distinguish watermarked from non-watermarked intermediate latents.
  2. Effective Evasion with Minimal Quality Impact: The developed gradient-based evasion attack consistently and successfully removes Tree-Ring watermarks. In their evaluations, the attack reduced the target detector's accuracy from over 95% (no attack baseline) to essentially random guessing (around 50%). Critically, this removal was achieved with minimal, imperceptible impact on image quality, validated both quantitatively and through manual inspection.
  3. Superiority of Latent Space Attacks: An ablation study revealed that performing the attack in the intermediate latent space is significantly more effective than applying perturbations directly to raw pixels. At the same attack intensity, the latent space attack caused a much greater drop in detector performance while preserving image quality more effectively. This underscores the strategic advantage gained by operating in the compressed latent domain when the VAE is known.
  4. The Overlooked Problem of Precision: The research highlights that the precision metric is often overlooked in watermarking research, with many prior works relying on balanced datasets. This can lead to the base rate fallacy, where high accuracy masks low precision in realistic deployment scenarios where the base rate of watermarked images might be very low. The study shows that Tree-Ring's precision can be insufficient even for high true positive rates in such low base-rate settings, and the proposed attacks exacerbate this problem by further lowering precision.
  5. Risk of Reusing Autoencoders: The findings underscore the inherent security risk associated with the widespread practice of reusing public autoencoders (VAEs) in the development and deployment of diffusion models. While beneficial for accelerating research and development, this practice inadvertently creates a vulnerability for watermarking schemes that do not account for the public accessibility of these crucial components.

Technical Deep Dive

▶ Watch: How Tree-Ring watermarking embeds a key (4:10)

The technical foundation of the attack hinges on understanding both the Tree-Ring watermarking scheme and the underlying architecture of modern latent diffusion models.

Tree-Ring Watermark Embedding:

As discussed, Tree-Ring embeds its watermark during the initial stages of image generation.

  1. Initial Latent: The process begins with an initial noise sample, referred to as the initial latent.
  2. Frequency Transformation: This initial latent is transformed from the spatial domain into the frequency domain using a Fast Fourier Transform (FFT). This operation decomposes the latent into its constituent frequencies.
  3. Watermark Embedding: The watermark key, a sequence of concentric rings with distinct values, is embedded into the low-frequency region (the center) of the frequency-transformed latent. Low-frequency components typically correspond to global image structures, making this region a robust place for embedding.
  4. Inverse Transformation: The modified frequency representation is then transformed back into the latent space using an Inverse Fast Fourier Transform (IFFT).
  5. Diffusion Generation: This watermarked latent is then fed as input to a pre-trained diffusion model (e.g., a U-Net architecture) to generate the final image. The diffusion model itself is assumed to be a black box to the attacker.

Tree-Ring Watermark Detection:

To detect a watermark in a given image:

  1. Latent Inversion: The image is inverted to recover the initial latent that likely produced it. For certain deterministic diffusion samplers like DDIM, this inversion can be performed with sufficient fidelity.
  2. Frequency Analysis: The recovered initial latent is then transformed into the frequency domain via FFT.
  3. Key Comparison: The central low-frequency region is extracted and compared against the known watermark key. If the similarity exceeds a pre-defined threshold, the watermark is considered present.

The Crucial Vulnerability: Latent Diffusion and Public VAEs:

The Achilles' heel of Tree-Ring, as identified by the researchers, lies in the common architecture of state-of-the-art diffusion models. These models employ latent diffusion, where the generative process does not operate directly on high-resolution pixel images but on a compressed, lower-dimensional representation. This compression is achieved through an autoencoder, specifically a Variational Autoencoder (VAE). The VAE encodes a high-resolution image into a compact intermediate latent representation and can decode it back.

The critical insight is that VAEs, designed for efficiency and reusability, are often publicly available. For instance, OpenAI's VAE is widely used. If an attacker has access to the same VAE used by the target diffusion model, they gain a powerful advantage.

The researchers discovered that the Tree-Ring watermark, despite being embedded in the initial latent space, remains distinguishable even after passing through the VAE and residing in the intermediate latent space. Visual evidence presented in the talk shows distinct circular, ring-like patterns persisting in the intermediate latent representations of watermarked images. To quantify this, a classifier was trained to distinguish between watermarked and non-watermarked intermediate latents, achieving perfect separation between the two classes. This demonstrates that the watermark signal is not sufficiently obscured or diffused by the VAE.

The Gradient-Based Evasion Attack:

The persistence of the watermark in the intermediate latent space, coupled with the ability to detect it there, forms the basis of the attack. Since the VAE is public, an attacker can generate their own dataset of watermarked and non-watermarked images, encode them using the public VAE, and then train a surrogate detector.

The attack proceeds in two stages:

  1. Training the Surrogate Detector:
  • Data Generation: The attacker generates a dataset comprising both watermarked and non-watermarked images using a diffusion model (which can be a proxy or the actual model if accessible for generation, not necessarily its weights).
  • VAE Encoding: These images are then encoded into their respective intermediate latent space representations using the publicly available VAE.
  • Surrogate Training: A simple neural network is trained on these intermediate latents to act as a surrogate detector. This network learns to distinguish watermarked from non-watermarked latents. Crucially, the attacker does not need access to the target diffusion model's weights, only the public VAE and the ability to generate example images.
  1. Gradient-Based Evasion:
  • Attack Objective: With the surrogate detector trained, the attacker can leverage its parameters to mount a gradient-based evasion attack. The goal is to introduce minimal perturbations to the intermediate latent representation of a watermarked image such that the surrogate detector (and, by transferability, the actual detector) no longer identifies the watermark, all while preserving the visual quality of the final image.
  • Perturbation in Latent Space: By computing gradients with respect to the input latent of the surrogate detector, the attacker can iteratively adjust the latent representation to minimize the detector's confidence in the watermark's presence. This process effectively "erases" the watermark signal from the intermediate latent.
  • Image Generation: The perturbed intermediate latent is then passed through the VAE's decoder to reconstruct an image. Because the perturbations are small and in the latent space, the impact on the final image's visual quality is minimal.

Evaluation and Results:

The researchers rigorously evaluated their attack against various baselines and state-of-the-art removal methods.

  • Detector Accuracy: In the "no attack" baseline, the Tree-Ring detector achieved over 95% accuracy. Their proposed attack, highlighted in yellow in their results table, successfully reduced the detector's accuracy to approximately 50%, effectively rendering it useless (random guessing).
  • Image Quality: The impact on image quality was consistently reported as "very low" and imperceptible through manual verification.
  • VAE Mismatch: The study also explored scenarios where the attacker's VAE did not perfectly match the target model's VAE. While the attack still decreased detector performance, the reduction was significantly smaller, confirming that the attack is most potent when the VAE is publicly known and shared.
  • Ablation Study (Latent vs. Pixel Attack): A key ablation study compared the latent-space attack with an equivalent attack applied directly to raw pixels. At the same perturbation intensity, the latent-space attack achieved a higher drop in detector performance with a more limited impact on image quality. This unequivocally demonstrates the superior effectiveness and stealth of operating in the compressed latent space.

Precision as a Critical Metric:

The talk emphasized the often-overlooked importance of precision in watermarking evaluations. Previous research frequently uses balanced datasets, which can inflate accuracy metrics. However, in real-world deployments, especially early stages, the base rate of watermarked images can be very low. In such scenarios, even a high accuracy can mask a very low precision, leading to a high number of false positives (real images incorrectly flagged as AI-generated). The researchers showed that Tree-Ring's precision is low even for high true positive rates at realistic low base rates, and their attack exacerbates this problem further, making reliable deployment even more challenging.

Demo / Proof of Concept

▶ Watch: Vulnerability: Latent diffusion and public VAEs (6:20)

While the talk did not feature a live, interactive demonstration of the attack, the speakers presented compelling evidence of its effectiveness through quantitative results and qualitative observations. The primary "proof of concept" lies in the comprehensive evaluation presented, which includes:

  1. Visual Evidence of Watermark Persistence: Images showcasing the distinct circular, ring-like patterns of the Tree-Ring watermark still visible within the intermediate latent space (obtained via the VAE) of watermarked images. This visual confirmation underscores the core vulnerability.
  2. Classifier Performance: The reported perfect separation achieved by a simple neural network classifier trained to distinguish watermarked from non-watermarked intermediate latents. This serves as a strong empirical demonstration that the watermark signal is indeed detectable within this critical space.
  3. Quantitative Attack Results: The detailed table comparing the proposed attack against baselines, showing the dramatic reduction in detector accuracy (from over 95% to approximately 50%) while maintaining "very low" impact on image quality.
  4. Qualitative Image Quality Assessment: The speakers explicitly stated that they "verified this manually" by "checking the images after removal and the removal of the watermark is imperceptible." This confirms the attack's ability to remain stealthy.

These combined pieces of evidence serve as a robust proof of concept for the feasibility and efficacy of their gradient-based evasion attack against Tree-Ring watermarks.

Defensive Implications

▶ Watch: Gradient-based evasion attack to remove watermarks (8:50)

The findings of this research carry significant implications for developers and researchers working on generative AI watermarking schemes, particularly for those aiming to ensure the long-term security and reliability of such technologies.

  1. Rethink VAE Reusability and Public Availability: The most immediate defensive implication is the need to critically reassess the practice of reusing publicly available VAEs in watermarking contexts. While convenient for model development and fine-tuning, this practice creates a critical exposure point if the watermarking scheme relies on the obscurity of the latent space. Developers should consider using proprietary, non-public VAEs for watermarked models, or, more robustly, design schemes that are inherently secure even if the VAE is known.
  2. Design Latent-Space Indistinguishable Watermarks: A direct countermeasure would be to engineer watermarking schemes that ensure the embedded signal is truly indistinguishable or significantly obfuscated within the intermediate latent space generated by the VAE. This might involve embedding watermarks in ways that are deeply entangled with the VAE's encoding process or that do not manifest as easily detectable patterns in the compressed latent representation.
  3. Pursue Formal Security Guarantees: The paper strongly advocates for the development of watermarking schemes that come with formal security guarantees. This means moving beyond empirical robustness claims to mathematically proven assurances of watermark integrity against specific attack models. This approach, common in cryptography, would provide a much higher level of confidence in the security of watermarking systems.
  4. Embrace Comprehensive Evaluation Metrics, Especially Precision: The research critically highlights the base rate fallacy and the inadequacy of relying solely on accuracy for evaluating watermarking schemes. Defenders must insist on and incorporate precision as a key evaluation metric, especially when considering realistic deployment scenarios where the proportion of watermarked images might be very low. A high precision rate is crucial to avoid an overwhelming number of false positives, which can erode trust in the detection system.
  5. Explore Adversarial Training for Watermark Robustness: While not explicitly detailed as a defense, the success of a gradient-based evasion attack suggests that adversarial training techniques could be explored. Training the watermarking embedding or detection mechanism to be robust against such gradient-based perturbations might enhance its resilience.
  6. Layered Security Approaches: Given the increasing sophistication of attacks, future watermarking solutions might need to adopt a layered security approach, combining in-processing watermarks with other detection methods (e.g., model-specific fingerprints, cryptographic attestations) to create a more resilient overall defense system.

Key Takeaways

  • Tree-Ring watermarks are vulnerable: The popular in-processing Tree-Ring watermarking scheme can be effectively removed with minimal impact on image quality.
  • Public VAEs are the attack vector: The vulnerability stems from the widespread reuse of publicly available Variational Autoencoders (VAEs), which allow attackers to access and manipulate the intermediate latent space where watermark signals persist.
  • Latent space attacks are superior: Attacking the watermark in the intermediate latent space is significantly more effective and stealthier than applying perturbations directly to raw pixels.
  • Precision is a critical, overlooked metric: Watermarking research and deployment must prioritize and rigorously evaluate precision, especially in realistic low base-rate scenarios, to avoid the base rate fallacy and ensure reliable detection.
  • New watermarking paradigms are needed: Future watermarking schemes must be designed to be either indistinguishable in the latent space or possess formal security guarantees against such evasion attacks.
  • AI safety requires holistic security: The findings underscore that the security of AI-generated content hinges not just on the watermarking algorithm itself, but also on the security and public accessibility of the underlying model components like VAEs.

About the Speaker(s)

The talk was presented by Mark Warth and represents joint work with Junhua Lin, both affiliated with the University of Edinburgh. Their research focuses on the intersection of artificial intelligence and security, specifically investigating the vulnerabilities and defensive measures pertaining to generative AI models and their potential for misuse. Their work contributes to the broader field of AI safety by critically evaluating existing security mechanisms and proposing new directions for building more robust and trustworthy AI systems.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Solid, well-scoped academic security research that identifies a real and underappreciated attack surface: the persistence of Tree-Ring watermark signals through publicly available VAEs into the intermediate latent space. The attack is technically coherent, the results are credible, and the precision/base-rate critique is a genuinely useful corrective to how the field evaluates these schemes.

Heather Calloway (CISO) — WEAK

Technically sound research that identifies a real vulnerability in a specific watermarking scheme, but it never crosses the threshold into governance, policy, or operational relevance. The defensive implications section gestures at implications without reaching the people who need to act on them.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)