Towards Understanding Unsafe Video Generation
Yan Pang
Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · AI Safety
Overview
The rapid advancement of generative AI models has unlocked unprecedented creative capabilities, but also introduced significant security and ethical challenges, particularly concerning the generation of unsafe or malicious content. This talk, "Towards Understanding Unsafe Video Generation," presented by Yan Pang at the NDSS Symposium, addresses a critical and emerging facet of this problem: the proliferation of unsafe video content generated by AI models. The research highlights a worrying trend where some video generation models lack adequate safety filters, making them susceptible to misuse by malicious actors who explicitly seek to generate harmful videos using unsafe prompts.
Key moments
- 0:00 Introduction to unsafe video generation problem
- 1:00 Data collection and annotation methodology
- 2:00 Identified five categories of unsafe video
- 3:00 Overview of existing defense methods
- 4:40 Introducing Latent Variable Defense (LVD)
- 5:40 Detailed workflow of LVD with hyperparameters
- 7:00 Evaluation setup and comparison with baselines
Towards Understanding Unsafe Video Generation
Speakers: Yan Pang
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=iChtmEhX3Aw
Overview
The rapid advancement of generative AI models has unlocked unprecedented creative capabilities, but also introduced significant security and ethical challenges, particularly concerning the generation of unsafe or malicious content. This talk, "Towards Understanding Unsafe Video Generation," presented by Yan Pang at the NDSS Symposium, addresses a critical and emerging facet of this problem: the proliferation of unsafe video content generated by AI models. The research highlights a worrying trend where some video generation models lack adequate safety filters, making them susceptible to misuse by malicious actors who explicitly seek to generate harmful videos using unsafe prompts.
The core of Pang's work is an investigation into the mechanisms of unsafe video generation and the development of effective defense strategies. Given that video generation is a relatively nascent field compared to image generation, there's a notable absence of established benchmarks and defense methodologies tailored specifically for video. This research not only fills this gap by creating a novel dataset of unsafe AI-generated videos but also proposes a robust and computationally efficient defense mechanism designed to detect and mitigate the creation of such content during the generation process itself.
The talk underscores the urgent need for robust safety measures in the evolving landscape of generative AI. As video generation models become more sophisticated and accessible, the potential for harm – ranging from the spread of misinformation and propaganda to the creation of disturbing or illegal content – escalates. By systematically analyzing the problem, developing a foundational dataset, and introducing a proactive defense, this research provides crucial insights and tools for securing these powerful technologies against malicious exploitation.
Background
▶ Watch: Introduction to unsafe video generation problem (0:00)
The genesis of this problem lies in the inherent capabilities of generative AI models, which, while designed for creative tasks, can be steered towards generating undesirable content. This issue is not new in the realm of text-to-image models, where malicious users on platforms like Lexica and 4chan have actively experimented with "unsafe prompts" to bypass safety filters and produce harmful images. However, the emergence of text-to-video generation models introduces a new dimension of complexity and potential impact. Unlike static images, videos can convey narratives, emotions, and dynamic actions, amplifying the potential for misuse.
A significant challenge identified by Pang is the surprising lack of built-in safety filters in some video generation models. This oversight creates a critical vulnerability, enabling malicious users to easily generate unsafe videos without obstruction. Furthermore, even seemingly "normal" prompts can, with a small probability, trigger the generation of unsafe content, highlighting the need for comprehensive and robust safety mechanisms.
Prior work in defending against unsafe content generation primarily focused on image models. These existing defense strategies can be broadly categorized into two groups based on their interaction with the generative model:
- Model-Right Methods: These approaches necessitate modifications to the generation process or updates to the model's parameters. Examples include Safe Latent Diffusion, which alters the generation direction when it veers towards unsafe concepts, and SafeGen, which fine-tunes self-attention layers to prevent unsafe image generation. While effective, these methods demand substantial computational resources for retraining or fine-tuning, making them resource-intensive and often impractical for widespread adoption or rapid adaptation.
- Model-Free Methods: These methods operate post-generation, relying solely on the final output of the generative model. An example is Unsafe Diffusion, which employs a classification model to categorize generated samples as safe or unsafe. Another is the Stable Diffusion safety filter, which uses CLIP embeddings to calculate a similarity score between generated images and predefined unsafe concepts, then compares this score against a threshold. The primary drawback of model-free methods is their susceptibility to adversarial attacks and jailbreak techniques, where clever prompt engineering or subtle output manipulations can bypass detection.
The inherent complexity of video data, which involves temporal dynamics in addition to spatial information, means that many image-centric defense methods cannot be directly transferred or efficiently applied to video generation. This gap necessitated a novel approach that is both computationally efficient and robust against evasion techniques, prompting the development of the "model-reading" strategy introduced in this research.
Key Findings
▶ Watch: Identified five categories of unsafe video (2:00)
The research yielded several critical findings that illuminate the landscape of unsafe video generation and propose effective countermeasures:
- Prevalence of Unsafe Video Generation: Current video generation models, despite their novelty, are capable of producing high-resolution unsafe videos, and disturbingly, many open-source models lack integrated safety filters, making them vulnerable to malicious use.
- Need for Video-Specific Benchmarks: The nascent state of video generation research means there's a significant absence of benchmark datasets for evaluating unsafe content. This work directly addresses this by constructing the first VGM-based Unsafe Video Dataset, comprising 937 validated unsafe generated videos categorized into five distinct types.
- Categorization of Unsafe Video Content: Through a rigorous semantic coding analysis involving human annotators, five primary categories of unsafe video content were identified: distort video, terrifying video, pornographic video, violent video, and politic video. This categorization provides a foundational framework for understanding and classifying harmful video content.
- Limitations of Existing Defense Strategies: Traditional model-right and model-free defense mechanisms, originally designed for image generation, prove inadequate for video. Model-right methods are computationally prohibitive for the higher dimensionality of video data, while model-free methods are easily bypassed by sophisticated adversarial attacks.
- Introduction of a Novel "Model-Reading" Defense: The research proposes Latent Variable Defense (LVD), a novel "model-reading" approach. LVD leverages intermediate outputs during the video generation model's inference process, striking a balance between the computational cost of model-right methods and the robustness issues of model-free methods.
- Superior Performance of LVD: LVD demonstrated significantly higher detection accuracy compared to baseline model-free methods like Unsafe Diffusion across various evaluation metrics. It also showed improved protection capabilities over model-right methods like Safe Latent Diffusion, especially when integrated.
- LVD as a Pluggable and Enhancable Solution: LVD is designed as an easily pluggable method that can be combined with existing defense strategies. Its integration with Unsafe Diffusion led to significantly improved detection accuracy, and when combined with Safe Latent Diffusion, it provided enhanced protection by incorporating confidence scores into the generation process.
These findings collectively underscore the severity of the unsafe video generation problem and present a practical, effective, and computationally sound solution that can be readily integrated into the development and deployment of future video generation technologies.
Technical Deep Dive
▶ Watch: Overview of existing defense methods (3:00)
The technical contribution of this research spans data collection and annotation, a novel classification of defense methodologies, and the intricate design of the Latent Variable Defense (LVD).
Data Collection and Annotation
Recognizing the lack of a specialized benchmark for unsafe video generation, the first crucial step was to construct one.
- Prompt Acquisition: The research leveraged prompts from existing datasets like Safety Fusion and I2P. These datasets are known for collecting prompts from platforms like Lexica and 4chan that are specifically crafted to induce text-to-image models to generate unsafe content. The rationale for using these image-oriented prompts was that most contemporary video generation models still rely on text-to-image models as their foundational backbone, suggesting a semantic similarity in their underlying representations.
- Video Generation: These prompts were then fed into several unspecified video generation models, resulting in an initial corpus of 5,670 generated videos. After a meticulous filtering process to remove low-quality or uninformative videos, a refined set of 2,112 videos remained.
- Clustering and Semantic Coding: To manage the complexity of annotation, K-means clustering was applied to the video embeddings, yielding 23 distinct video clusters. For video embeddings, the authors employed CLIP to extract embeddings from individual frames, then averaged these frame embeddings to represent the entire video's semantic content. Following Brun's work on semantic coding analysis, this clustering allowed for a systematic identification of unsafe categories.
- Codebook Development: This analysis led to the creation of a codebook defining five distinct categories of unsafe video: distort video, terrifying video, pornographic video, violent video, and politic video.
- Human Annotation: To ensure objectivity and avoid researcher bias, an extensive human annotation effort was undertaken. With IRB protocol approval, 600 participants were recruited from the Prolific platform. Each participant labeled 30 generated videos. After data cleaning and preprocessing, 403 valid responses were collected. This rigorous process culminated in the creation of the VGM-based Unsafe Video Dataset, which includes 937 comprehensively labeled unsafe generated videos.
Defense Method Classification and LVD Design
The talk categorizes existing and proposed defense methods into three distinct groups:
- Model-Right Methods: These methods directly intervene in the model's internal generation process or modify its parameters. Examples include Safe Latent Diffusion (changing generation direction) and SafeGen (fine-tuning self-attention layers). While potentially effective, their primary drawback is the substantial computational cost associated with retraining or fine-tuning model parameters.
- Model-Free Methods: These methods operate on the final output of the generative model, typically using a separate classifier. Examples include Unsafe Diffusion (binary classification of final output) and the Stable Diffusion safety filter (CLIP-based similarity scoring). Their main vulnerability is their susceptibility to adversarial attacks and jailbreak techniques that can subtly alter the final output to bypass detection.
- Model-Reading Methods (LVD - Latent Variable Defense): Proposed by the authors, LVD seeks a middle ground. It aims for computational efficiency without requiring model parameter updates, while also being robust against adversarial bypasses. The core idea is to leverage the intermediate outputs generated during the inference process, rather than just the final product.
Latent Variable Defense (LVD) Mechanism
LVD is specifically designed for diffusion-based generation models, which typically employ a DDPM (Denoising Diffusion Probabilistic Models) sampler. This sampler progressively denoises a noisy sample step-by-step until a coherent output (an image or a set of video frames) is produced. Crucially, this is a deterministic sampling process.
The insight behind LVD is that if the generation process is deterministic, then unsafe features might emerge and be detectable at earlier, intermediate stages of denoising, not just at the final output.
- Intermediate Detection: Instead of relying on the final video, LVD probes the latent space at various steps during the denoising process.
- Hyperparameter
E: The first hyperparameter,E, controls the number of initial denoising steps at which detection models are applied. For instance, ifE=20, the first 20 intermediate latent states are analyzed. While a typical diffusion model might use 50 inference steps, training 50 separate detection models would be computationally demanding.Eallows for a configurable trade-off. - Binary Classification Models: At each of the
Eselected steps, a dedicated binary classification model is employed. This model is trained to determine if the intermediate latent state is indicative of an unsafe sample (output 1) or a safe sample (output 0). The detection model is built based on a Video Mask Autoencoder architecture. - Hyperparameter
lambda: The second hyperparameter,lambda, serves as a threshold for cumulative detection. The detection results (0s and 1s) from theEsteps are summed. This sum is then compared against a threshold calculated asE lambda. Ifsum(detection_results) > E lambda, the sample is classified as unsafe; otherwise, it's considered safe. For example, ifE=20andlambda=0.6, an unsafe classification requires at least 12 (20 * 0.6) of the first 20 detection models to flag the sample as unsafe.
Evaluation Setup
LVD was evaluated against three prominent open-source video generation models: Magic Time, Anime Diffusion, and VideoCrafter 2. The custom VGM-based Unsafe Video Dataset was used for testing. The performance was measured using four different evaluation metrics (though only AUC score was explicitly mentioned in the presentation, indicating high performance).
Key Evaluation Findings
EandlambdaRelationship: An interesting finding was the inverse relationship betweenEand the optimallambda. AsE(number of detection steps) increases, the best detection accuracy is achieved with a lowerlambdavalue. A highlambda(e.g.,lambda=1) demands that every detection step flags the content as unsafe, leading to a high True Negative Rate (TNR, correctly identifying safe content) but a low True Positive Rate (TPR, failing to detect unsafe content) because it misclassifies many unsafe samples as safe due to the strict requirement.- High AUC Scores: LVD consistently achieved high AUC scores across all three tested video generation models, demonstrating its robust detection capabilities.
- Outperformance Against Baselines:
- Vs. Unsafe Diffusion (Model-Free): LVD significantly outperformed Unsafe Diffusion on the three unspecified evaluation metrics (in addition to AUC), indicating its superior robustness against bypass attempts.
- Vs. Safe Latent Diffusion (Model-Right): LVD showed better performance compared to Safe Latent Diffusion, quantified by an improved "nudely remove rate" (implicitly, the rate at which unsafe content is prevented or removed).
- Synergistic Combinations: LVD's plug-in nature allows for combination with other defense strategies:
- LVD + Unsafe Diffusion: Combining LVD with Unsafe Diffusion significantly boosted detection accuracy, showcasing how LVD can enhance model-free methods.
- LVD + Safe Latent Diffusion: By replacing the momentum parameters in Safe Latent Diffusion with confidence scores derived from LVD's detection models, the combined method provided superior protection compared to the original Safe Latent Diffusion, all under the same configurations.
This detailed technical approach, from data construction to the nuanced design and rigorous evaluation of LVD, marks a significant step forward in securing the emerging field of video generation.
Demo / Proof of Concept
▶ Watch: Detailed workflow of LVD with hyperparameters (5:40)
While the talk did not feature a live, interactive demonstration in the traditional sense, the comprehensive evaluation section serves as the proof of concept for the Latent Variable Defense (LVD) method. The research meticulously details the experimental setup, including the use of three open-source video generation models (Magic Time, Anime Diffusion, VideoCrafter 2) and a custom-built dataset of unsafe videos. The presentation of quantitative results, such as high AUC scores and superior performance against established baseline methods like Unsafe Diffusion and Safe Latent Diffusion, effectively demonstrates LVD's efficacy and robustness. The ability of LVD to not only outperform these baselines but also to enhance their performance when combined further validates its practical utility and technical soundness.
Defensive Implications
▶ Watch: Evaluation setup and comparison with baselines (7:00)
The research presented in "Towards Understanding Unsafe Video Generation" carries profound implications for various stakeholders involved in the development, deployment, and regulation of generative AI technologies:
- For AI Model Developers: Developers of video generation models must prioritize safety filters from the outset. The research clearly demonstrates that leaving models unfiltered is a critical vulnerability. Implementing "model-reading" defenses like Latent Variable Defense (LVD), which monitor intermediate latent states, offers a computationally efficient and robust solution compared to solely relying on post-generation checks or costly model retraining. Integrating LVD as a plug-in component can significantly enhance the safety posture of their models.
- For Platform Providers and Content Moderators: Platforms hosting or utilizing AI-generated video content need to adopt proactive and sophisticated detection mechanisms. Relying on traditional image-based safety filters or simple output classifiers is insufficient. The findings suggest that a multi-layered defense, potentially combining LVD with existing model-free methods, could provide a more resilient content moderation framework, reducing the spread of harmful videos.
- For Security Researchers: This work opens new avenues for research into the security of generative AI. The concept of "model-reading" provides a novel paradigm for detecting malicious outputs. Future research could explore more advanced techniques for analyzing latent space dynamics, developing better video encoders for safety detection (as highlighted in the Q&A), or extending LVD to other modalities or more complex generative architectures. Investigating adversarial attacks specifically targeting intermediate latent states could also be a fruitful area.
- For End-Users: While not directly responsible for implementing defenses, users should be aware of the potential for AI models to generate unsafe content. They should exercise caution when interacting with new or unregulated generative AI tools and report any instances of misuse. Understanding that some models lack inherent safety mechanisms can inform user choices and expectations.
- For Policymakers and Regulators: The prevalence of unfiltered unsafe video generation underscores the urgent need for regulatory frameworks. Policymakers should consider establishing guidelines or standards for safety features in generative AI models, particularly those capable of producing video. The research provides concrete evidence of the problem and a validated technical approach to mitigation, which can inform the development of effective policies.
- For the AI Ethics Community: This research contributes significantly to the discourse on responsible AI development. It provides tangible evidence of how generative models can be misused and offers a technical solution to mitigate such harms, encouraging a more proactive and security-conscious approach to AI development.
Overall, the work provides both a stark warning about the current state of video generation safety and a practical, actionable defense strategy, pushing the field towards more secure and ethically responsible AI development.
Key Takeaways
- Unsafe Video Generation is a Reality: Many emerging video generation models lack adequate safety filters, making them susceptible to malicious users generating high-resolution unsafe content.
- New Benchmark Dataset: The research established the first VGM-based Unsafe Video Dataset, comprising 937 human-annotated unsafe videos categorized into five types: distort, terrifying, pornographic, violent, and politic.
- Limitations of Existing Defenses: Current image-centric "model-right" (computationally expensive) and "model-free" (easily bypassed) defense methods are insufficient for the complexity and robustness requirements of video generation.
- Latent Variable Defense (LVD): A novel "model-reading" defense, LVD, efficiently detects unsafe content by analyzing intermediate latent states during the deterministic denoising process of diffusion models, balancing computational cost and robustness.
- Superior Performance and Combinability: LVD significantly outperforms baseline image-based safety methods (e.g., Unsafe Diffusion) and can be effectively combined with existing defenses (e.g., Safe Latent Diffusion) to enhance overall protection.
- Critical Need for Video Encoders: The research highlights the current reliance on image encoders (like CLIP) for video embedding, underscoring a broader need for more powerful, dedicated video encoders to improve the accuracy and nuance of safety detection.
About the Speaker(s)
The talk was presented by Yan Pang. Based on the provided metadata and transcript, no specific title or company affiliation was mentioned during the presentation. Yan Pang presented this research paper, "Towards Understanding Unsafe Video Generation," at the NDSS Symposium.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Competent, methodical ML security research that fills a real gap — first dataset and defense framework specifically for unsafe video generation. The Latent Variable Defense idea is sound and the empirical results look credible, but the novelty ceiling is low: this is a clean extension of existing image-safety work into the video domain, not a fundamental new insight.
Heather Calloway (CISO) — WEAK
Technically credible foundational research — a new dataset, a new defense taxonomy, a novel detection mechanism — but it stays firmly inside the research lab. The governance exposure, the accountability question, and the regulatory urgency are all gestured at but never operationalized for the people who actually need to act on them.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025