SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
Xunguang Wang (PhD student · HKUS)
34th USENIX Security Symposium (USENIX Security '25) · Day 2 · LLM Security 2: Jailbreaking and Prompt Stealing
Overview
This article delves into "SelfDefend," an innovative framework designed to protect Large Language Models (LLMs) from jailbreak attacks. Presented by Xunguang Wang from HKUS, the talk introduces a novel, dual-layer defense mechanism inspired by the traditional system security concept of shadow stacks. The core premise is to leverage LLMs themselves, not just for answering user queries, but also for actively detecting and mitigating malicious instructions.

Key moments
- 0:00 Introduction to LLM security risks and jailbreaking
- 2:50 SelfDefend: Motivations and shadow stack concept
- 4:00 SelfDefend framework pipeline explained
- 4:50 Designing direct and intent detection prompts
- 6:00 Initial evaluation: GPT models' jailbreak suppression
- 7:00 Fine-tuning open-source models for cost-effectiveness
- 8:00 Final evaluation: efficiency and robustness of tuned models
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
Speakers: Xunguang Wang, PhD Student, HKU (Hong Kong University of Science and Technology)
Conference: USENIX Security
YouTube: https://www.youtube.com/watch?v=bUamsPUHURA
Overview
This article delves into "SelfDefend," an innovative framework designed to protect Large Language Models (LLMs) from jailbreak attacks. Presented by Xunguang Wang from HKUS, the talk introduces a novel, dual-layer defense mechanism inspired by the traditional system security concept of shadow stacks. The core premise is to leverage LLMs themselves, not just for answering user queries, but also for actively detecting and mitigating malicious instructions.
The proliferation of powerful LLMs like ChatGPT, DeepSeek, and Claude has brought immense utility but also significant security risks, with jailbreaking standing out as a critical vulnerability. Jailbreak attacks exploit LLMs' safety alignments, coercing them into generating harmful, unethical, or illegal content. SelfDefend addresses this by proposing a practical, robust, and cost-effective defense that operates in parallel with the target LLM, ensuring that malicious prompts are identified and blocked before they can induce harmful outputs. The research highlights the potential for LLMs to become self-defending systems, thereby enhancing their reliability and ethical deployment.
Background
▶ Watch: Introduction to LLM security risks and jailbreaking (0:00)
Large Language Models, while transformative, are not immune to security vulnerabilities. Common risks include data leakage, data poisoning, prompt injection, hallucination, and critically, jailbreak attacks. These attacks involve crafting malicious instructions that circumvent an LLM's built-in safety mechanisms, causing it to produce content it would normally refuse, such as instructions for illegal activities or the spread of misinformation. The social, ethical, economic, and legal implications of successful jailbreaks are severe, underscoring the urgent need for robust defense mechanisms.
Prior research into jailbreak defense has primarily fallen into three categories: detection-based defense, which attempts to identify malicious prompts before execution; prompt-based defense, which modifies prompts or adds safety instructions; and tuning-based defense, which fine-tunes models to resist jailbreaks. While these approaches offer varying degrees of success, they often come with limitations such as high computational cost, susceptibility to adaptive attacks, or performance degradation on normal queries. The challenge lies in developing a defense that is both effective against a wide range of jailbreaks and practical for real-world deployment, incurring minimal overhead for legitimate users. SelfDefend seeks to address these limitations by introducing a parallel processing architecture and carefully designed detection prompts, drawing inspiration from established system security paradigms.
Key Findings
▶ Watch: SelfDefend framework pipeline explained (4:00)
The SelfDefend framework presents several key findings that significantly advance the state of LLM jailbreak defense:
- Effective Suppression of Attack Success Rates: SelfDefend, particularly when initially implemented with GPT-3.5 and GPT-4 as the detection LLM, was empirically validated to significantly suppress the attack success rate across all tested categories of jailbreak prompts. This demonstrates the robust capability of a parallel, LLM-based detection mechanism.
- Negligible Impact on Normal Queries: A crucial aspect of practical defense is minimizing false positives and performance degradation for legitimate use. SelfDefend achieved this, incurring negligible effects on normal queries. Specifically, over 80% of normal queries experienced no additional delay, indicating its efficiency for everyday users. The maximum observed extra delay was a mere 0.4 seconds, making it highly efficient compared to other defense methods.
- Cost-Effective and Self-Contained Defense: Recognizing the expense associated with using proprietary models like GPT-4 for continuous detection, the researchers successfully fine-tuned open-source models (Llama 2 and Mistral) to serve as shadow LLMs within the SelfDefend architecture. These tuned models maintained the defense capabilities of the GPT-4-based framework while offering a robust, low-cost, and self-contained solution.
- Robustness Against Adaptive Jailbreaks: The framework demonstrated strong resilience against various forms of adaptive jailbreaks, including entire adaptive jailbreaks, separate adaptive jailbreaks, and individual adaptive attacks. The fine-tuned shadow models proved more effective and less susceptible to being bypassed compared to baseline Llama 2 series models.
- Explainability through Harmful Portion Alignment: SelfDefend's ability to extract and identify harmful portions from jailbreak prompts was validated using the Clip Score. Higher Clip Scores for identified harmful intentions compared to the original attack prompts indicate that even when jailbreak prompts attempt to obfuscate malicious content, SelfDefend accurately pinpoints the true harmful intent. This not only confirms its effectiveness but also provides valuable explainability for its decisions.
These findings collectively establish SelfDefend as a promising, practical, and highly effective solution for securing LLMs against the pervasive threat of jailbreak attacks.
Technical Deep Dive
▶ Watch: Designing direct and intent detection prompts (4:50)
The SelfDefend framework is built upon two core motivations: recognizing that target LLMs operate in both answering and detection states, and the premise that all jailbreak attacks inherently contain malicious intentions, whether explicit or implicit. Inspired by the traditional system security concept of shadow stacks, which involves maintaining a separate stack to detect control-flow hijacking, SelfDefend implements a dual-layer protection mechanism using LLMs themselves.
The architecture is straightforward yet powerful. When a user inputs a prompt, SelfDefend runs two LLMs in parallel: the target LLM (the primary model intended to answer the query) and a detection LLM (the "shadow" model responsible for identifying malicious content).
- Parallel Execution: The user's prompt is simultaneously fed to both the target LLM and the detection LLM.
- Detection Phase: The detection LLM analyzes the prompt for harmful portions or malicious intent.
- Response Handling:
- If the detection LLM determines that the prompt contains no harmful portions, the response from the target LLM is released to the user.
- If the detection LLM identifies harmful portions or malicious intent, SelfDefend refuses to answer the input prompt, thereby preventing the target LLM from generating harmful content.
A critical component of SelfDefend's efficacy lies in the design of its detection prompts, which guide the detection LLM. Two types of prompts were developed:
- Direct Prompt: This prompt is designed to directly identify harmful parts within the original jailbreak prompt. It is particularly effective against jailbreaks where malicious goals are relatively obvious or less obfuscated. The direct prompt instructs the detection LLM to explicitly extract any safety-violating segments from the user's input.
- Intent Prompt: For more sophisticated jailbreaks that hide malicious goals within complex or benign-looking text, the intent prompt takes an indirect approach. It first instructs the detection LLM to summarize the true intention of the user's query. Following this summarization, it then prompts the LLM to extract any safety-violating parts from the summarized intention. This two-step process allows the detection LLM to uncover hidden malicious intent. To enhance the reasoning abilities of the intent prompt, in-context learning examples are utilized, providing the detection LLM with demonstrations of how to identify subtle maliciousness.
To ensure SelfDefend is practical and cost-effective, especially given the high cost of continuous API calls to large proprietary models like GPT-4, the researchers developed a tuning-based approach. This involved:
- Data Distillation: High-quality "red team" data, specifically from Anthropic, was used to distill knowledge from GPT-4. This process involved using GPT-4 to generate responses and classifications for a large dataset of both benign and malicious prompts, effectively creating a high-fidelity training dataset.
- Model Fine-tuning: The distilled dataset was then used to fine-tune open-source models, specifically Llama 2 and Mistral. The fine-tuning process utilized Aurora, a framework for efficient LLM training. This step created dedicated "shadow models" that could operate under the SelfDefend architecture with comparable defense capabilities to the GPT-4 based system but at a significantly lower operational cost and with the advantage of being self-contained.
The combination of a parallel processing architecture, carefully engineered detection prompts, and a robust tuning methodology allows SelfDefend to provide a lightweight, effective, and efficient defense against a broad spectrum of jailbreak attacks.
Demo / Proof of Concept
▶ Watch: Fine-tuning open-source models for cost-effectiveness (7:00)
While the talk did not feature a live, interactive demonstration of the SelfDefend framework in action, its efficacy and practical viability were thoroughly demonstrated through a comprehensive evaluation methodology. The "proof of concept" for SelfDefend lies in the empirical validation of its architecture and the performance of its underlying components.
The core operational flow, as presented, serves as the conceptual demonstration: a user input is simultaneously sent to a target LLM and a detection LLM. The detection LLM, armed with either the direct prompt or the intent prompt, processes the input. If no harmful content is detected, the target LLM's response is released. If harmful content or intent is identified, the request is blocked. This pipeline visually illustrates how the dual-layer defense mechanism functions in practice.
The robustness of this conceptual demonstration was then quantified through extensive evaluations:
- Attack Success Rate Suppression: Initial validations using GPT-3.5 and GPT-4 as detection LLMs showed a significant reduction in the success rate of various jailbreak attacks, including those categorized as human-based, optimization-based, generation-based, and implicit attacks. This confirmed the fundamental capability of an LLM-based detection layer.
- Performance Metrics: The framework's efficiency was demonstrated by measuring the extra delay incurred for normal queries. The results showed that over 80% of normal queries experienced no discernible delay, and the maximum delay observed was a mere 0.4 seconds. This showcased SelfDefend's practicality for real-world deployment without significant user experience degradation.
- Open-Source Model Validation: The most compelling aspect of the proof of concept was the successful fine-tuning of open-source models like Llama 2 and Mistral to function effectively as shadow models. Evaluations confirmed that these tuned models maintained the robust defense capabilities against all types of jailbreaks and normal queries, validating the framework's low-cost and self-contained deployment potential.
- Harmful Portion Alignment (Clip Score): The explainability of SelfDefend was demonstrated by its ability to accurately identify and extract the harmful intentions from jailbreak prompts. The use of Clip Score to measure the alignment between identified harmful portions and the original malicious intent provided quantitative evidence that SelfDefend understands and pinpoints the core of the attack, even when obfuscated.
- Adaptive Attack Robustness: The system's resilience was further proven against adaptive jailbreaks, where attackers attempt to bypass the defense. The tuned shadow models consistently outperformed baseline Llama 2 models, demonstrating a higher resistance to being "hacked" by sophisticated adaptive attacks.
These comprehensive evaluations serve as a strong empirical proof of concept, illustrating that SelfDefend is not just a theoretical concept but a practical and effective solution for mitigating LLM jailbreak attacks across various scenarios and deployment costs.
Defensive Implications
▶ Watch: Final evaluation: efficiency and robustness of tuned models (8:00)
The SelfDefend framework offers significant implications for organizations and developers seeking to secure their LLM deployments. Its dual-layer, LLM-driven approach provides a robust blueprint for future AI security strategies.
Firstly, the concept of a "shadow LLM" operating in parallel with the primary model introduces a new paradigm for real-time threat detection. This means that instead of relying solely on pre-filtering rules or post-generation content moderation, organizations can implement an active, intelligent monitoring layer that understands the nuances of language and intent, akin to how LLMs themselves operate. This is particularly crucial as jailbreak techniques become more sophisticated and context-dependent.
Secondly, the successful demonstration of fine-tuning open-source models like Llama 2 and Mistral to act as effective shadow LLMs democratizes advanced jailbreak defense. Organizations no longer need to exclusively rely on expensive proprietary APIs (like GPT-4) for robust protection. This enables the deployment of low-cost, self-contained, and customizable defense mechanisms within their own infrastructure, offering greater control over data privacy and operational costs. For companies with strict data governance requirements, hosting their own defense LLM locally is a significant advantage.
Thirdly, the development of specific detection prompts—the direct prompt and the intent prompt—provides a practical toolkit for implementing this defense. Organizations can adapt and refine these prompt engineering techniques to suit their specific threat models and the types of content they aim to prevent. The use of in-context learning examples to enhance detection LLM reasoning is a powerful technique that can be broadly applied to improve the accuracy and robustness of AI-based security systems.
Finally, SelfDefend's proven effectiveness against adaptive jailbreaks highlights its forward-looking design. As attackers continually evolve their methods, a defense mechanism that can adapt and resist new bypass techniques is essential. The framework's ability to accurately identify harmful portions, even when obfuscated, provides valuable explainability, allowing security teams to understand why a prompt was flagged, which can aid in refining defenses and threat intelligence. Organizations should consider integrating such a parallel detection architecture into their LLM deployment pipelines, leveraging it not only for real-time blocking but also for continuous learning and adaptation to emerging threats.
Key Takeaways
- Dual-Layer Defense with Shadow LLMs: SelfDefend introduces a novel dual-layer defense mechanism, inspired by traditional shadow stacks, where a dedicated "shadow LLM" runs in parallel with the target LLM to detect and mitigate jailbreak attempts.
- Effective and Efficient Protection: The framework significantly suppresses jailbreak attack success rates while maintaining high efficiency for normal queries, with over 80% experiencing no delay and a maximum extra delay of only 0.4 seconds.
- Sophisticated Detection Prompts: Two types of detection prompts, the "direct prompt" for obvious malicious content and the more advanced "intent prompt" (using in-context learning) for hidden malicious intentions, enable comprehensive threat identification.
- Cost-Effective Open-Source Deployment: SelfDefend can be effectively deployed using fine-tuned open-source models (e.g., Llama 2, Mistral) via data distillation from high-performing models like GPT-4, offering a robust, low-cost, and self-contained defense solution.
- Robustness and Explainability: The system demonstrates strong resilience against various adaptive jailbreak attacks and provides explainability by accurately identifying and aligning with the harmful portions of malicious prompts, validated through Clip Score analysis.
About the Speaker(s)
The talk was presented by Xunguang Wang, who is identified as a PhD student from HKUS (Hong Kong University of Science and Technology). No further biographical details were provided within the transcript or metadata.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Competent academic work on LLM jailbreak defense with a clean architectural idea — parallel shadow LLM detection inspired by shadow stacks. The shadow-stack analogy is cute and the fine-tuning pathway to open-source deployment is practically useful, but the core insight (use an LLM to judge another LLM's input) isn't fundamentally new territory, and the evaluation methodology raises the usual academic red flags around benchmark saturation and adaptive attack coverage.
Heather Calloway (CISO) — WEAK
Technically credible research on a real problem — LLM jailbreak defense — but it never crosses the threshold into institutional relevance. The framework is clever; the talk is aimed at researchers, not the operators or leaders who are actually deploying LLMs at scale.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)