Lancet: A Formalization Framework for Crash and Exploit Pathology

Qinrun Dai

34th USENIX Security Symposium (USENIX Security '25) · Day 1 · Software Security 1

Overview

This article delves into a critical security vulnerability in Large Language Models (LLMs) uncovered by researchers from KAIST, presented at USENIX Security. The paper, titled "Refusal Is Not an Option: Unlearning Safety Alignment of Large Language Models," introduces novel adversarial unlearning attacks designed to dismantle the safety alignment mechanisms of LLMs. Safety alignment is a crucial procedure that prevents LLMs from generating harmful, privacy-sensitive, or copyrighted content by training them to reject malicious instructions. Machine unlearning, typically used to remove problematic data and ensure compliance with regulations like GDPR, is paradoxically repurposed by the researchers to achieve the opposite effect: making LLMs forget their refusal capabilities.

Read the paper · Download the PDF (PDF) · Slides

Paper abstract

Safety alignment has become an indispensable procedure to ensure the safety of large language models (LLMs), as they are reported to generate harmful, privacy-sensitive, and copy-righted content when prompted with adversarial instructions. Machine unlearning is a representative approach to establishing the safety of LLMs, enabling them to forget problematic training instances and thereby minimize their influence. However, no prior study has investigated the feasibility of adversarial unlearning—using seemingly legitimate unlearning requests to compromise the safety of a target LLM. In this paper, we introduce novel attack methods designed to break LLM safety alignment through unlearning. The key idea lies in crafting unlearning instances that cause the LLM to forget its mechanisms for rejecting harmful instructions. Specifically, we propose two attack methods. The first involves explicitly extracting rejection responses from the target LLM and feeding them back for unlearning. The second attack exploits LLM agents to obscure rejection responses by merging them with legitimate-looking unlearning requests, increasing their chances of bypassing internal filtering systems. Our evaluations show that these attacks significantly compromise the safety of two open-source LLMs: LLaMA and Phi. LLaMA's harmfulness scores increase by an average factor of 11 across four representative unlearning methods, while Phi exhibits a 61.8× surge in the rate of unsafe responses. Furthermore, we demonstrate that our unlearning attack is also effective against OpenAI's fine-tuning service, increasing GPT-4o's harmfulness score by 2.21×. Our work identifies a critical vulnerability in unlearning and represents an important first step toward developing safe and responsible unlearning practices while honoring users' unlearning requests. Our code is available at .

Visual summary for Lancet: A Formalization Framework for Crash and Exploit Pathology by Qinrun Dai
Visual summary for Lancet: A Formalization Framework for Crash and Exploit Pathology by Qinrun Dai

Refusal Is Not an Option: Unlearning Safety Alignment of Large Language Models

Speakers: Minkyoo Song, Hanna Kim, Jaehan Kim, Seungwon Shin, Sooel Son, KAIST

Conference: USENIX Security

YouTube: https://www.usenix.org/conference/usenixsecurity25/presentation/song-minkyoo

Overview

This article delves into a critical security vulnerability in Large Language Models (LLMs) uncovered by researchers from KAIST, presented at USENIX Security. The paper, titled "Refusal Is Not an Option: Unlearning Safety Alignment of Large Language Models," introduces novel adversarial unlearning attacks designed to dismantle the safety alignment mechanisms of LLMs. Safety alignment is a crucial procedure that prevents LLMs from generating harmful, privacy-sensitive, or copyrighted content by training them to reject malicious instructions. Machine unlearning, typically used to remove problematic data and ensure compliance with regulations like GDPR, is paradoxically repurposed by the researchers to achieve the opposite effect: making LLMs forget their refusal capabilities.

The core contribution of this research is demonstrating the feasibility and significant impact of these attacks on both open-source LLMs like LLaMA and Phi, and a real-world commercial service, OpenAI's GPT-4o via its fine-tuning API. The authors propose two distinct attack methodologies: one that directly unlearns explicit rejection responses, and another that leverages LLM agents to obscure these rejection responses within seemingly legitimate unlearning requests, effectively bypassing filtering systems. The findings reveal a profound weakness in current unlearning practices, highlighting that honoring unlearning requests without rigorous inspection can severely compromise LLM safety, leading to models that readily generate harmful content.

This work is paramount because it identifies a novel attack surface in LLM security. As machine unlearning becomes increasingly vital for regulatory compliance and responsible AI development, understanding and mitigating its adversarial exploitation is essential. The paper not only exposes this vulnerability but also proposes a preliminary mitigation strategy, serving as a foundational step toward developing safer and more robust unlearning procedures that can uphold user requests without jeopardizing public safety.

Background

Large Language Models (LLMs) have become ubiquitous, powering applications across diverse domains from education to customer service. Their ability to generate high-quality, coherent text based on a wide range of instructions has led to their widespread adoption. However, this versatility comes with inherent risks. LLMs are trained on massive datasets scraped from the internet, which inevitably contain harmful, biased, private, or copyrighted content. Consequently, LLMs can inadvertently generate responses that facilitate malicious activities, spread misinformation, or infringe on privacy and intellectual property.

To counteract these threats, safety alignment has become an indispensable procedure. This involves fine-tuning LLMs to produce benign or outright rejection responses when confronted with malicious or unsafe prompts. Common techniques for safety alignment include Supervised Fine-Tuning (SFT), Reinforcement Learning from Human Feedback (RLHF), and Direct Preference Optimization (DPO). These methods guide LLMs to align with human values and ethical guidelines, ensuring they respond with phrases like "I cannot assist with that request" when asked to generate harmful content.

Concurrently, machine unlearning has emerged as a critical methodology for addressing safety and privacy concerns in deployed LLMs. The goal of unlearning is to efficiently remove the influence of specific problematic data instances (e.g., harmful knowledge, copyrighted material, or Personally Identifiable Information (PII)) from a trained model without the prohibitive computational cost of retraining the entire model from scratch. This capability is not only beneficial for model refinement but also crucial for complying with privacy regulations such as the General Data Protection Regulation (GDPR) and the California Consumer Privacy Act (CCPA), which grant users the "right to be forgotten." Various unlearning algorithms have been developed, including Gradient Ascent (GA), Direct Preference Optimization (DPO), Negative Preference Optimization (NPO), and Task Vector (TV) methods, each offering different approaches to diminish the model's memory of specified data.

Despite the growing importance of both safety alignment and machine unlearning, prior research had not thoroughly investigated the potential for adversarial unlearning. This critical gap meant that methods to maliciously exploit unlearning requests—using seemingly legitimate means to undermine an LLM's safety—remained largely unexplored. The paper addresses this oversight by proposing the first attacks that specifically target the LLM's rejection mechanisms through manipulated unlearning processes.

Key Findings

The research presents compelling evidence that adversarial unlearning is a potent new attack vector against LLM safety alignment. The core discovery is that by strategically crafting unlearning requests, adversaries can cause LLMs to "forget" their ability to reject harmful instructions, turning them into tools for generating unsafe content.

The paper introduces two primary attack methodologies:

  1. Direct Rejection Unlearning (Scenario I): In a simplified scenario where unlearning requests are accepted without strict validation, the adversary directly extracts rejection responses (e.g., "I cannot assist with that request") from the target LLM when prompted with harmful instructions. This compiled dataset of rejection behaviors (D_reject) is then submitted for unlearning. The effect is dramatic: the LLM unlearns its refusal mechanisms, leading to a significant increase in harmful output. For instance, LLaMA's harmfulness scores increased by an average factor of 11x across four representative unlearning methods, with a peak increase of 20.3x on the LLM-LAT dataset when unlearned with Direct Preference Optimization (DPO). Phi exhibited an even more staggering average surge of 61.8x in unsafe responses, reaching a 193.5-fold increase on the LLM-LAT dataset with DPO.
  1. Blended Rejection Unlearning with LLM Agents (Scenario II): Acknowledging that real-world services often filter unlearning requests, the authors devise a more sophisticated attack. This method leverages two specialized LLM agents, a rewrite agent (A_rewrite) and an evaluation agent (A_eval), to blend rejection responses with legitimate-looking problematic content (such as synthetic PII, fake news, or copyrighted material). This blended dataset (D*_merged) is designed to bypass common filtering systems (e.g., classifiers for PII, fake news, or copyright infringement) while still containing the embedded rejection patterns. The attacks using D*_merged were highly effective, with over 96% of instances successfully bypassing filtering systems. LLaMA's harmfulness score increased by up to 8.3 times (Hex-PHI, DPO/NPO with PII-merged data), and Phi's harmfulness score surged by up to 111.5 times (LLM-LAT, Task Vector with fake news-merged data).

Beyond open-source models, the researchers demonstrated the practical impact of their attacks on a real-world commercial service: OpenAI's DPO fine-tuning API for GPT-4o. Despite OpenAI's robust moderation systems, the unlearning attacks successfully bypassed data validation steps, leading to a 2.21x increase in GPT-4o's harmfulness score (Hex-PHI dataset). This highlights a critical vulnerability even in advanced commercial LLM platforms.

Further findings include:

  • Small Dataset Effectiveness: Even unlearning datasets as small as 50 instances (one-tenth of the original size) could significantly increase harmfulness scores by an average of five-fold.
  • Consecutive Attacks: Repeated, small-scale unlearning requests (e.g., 30 instances per attempt) showed an increasing trend in harmfulness, indicating a persistent threat.
  • Robustness to Dilution: Malicious unlearning requests remained highly effective even when diluted with benign requests at a 10:1 ratio, increasing harmfulness by an average of 6.2 times.
  • Mitigation: The paper introduced a defensive classifier (C_enriched) capable of detecting most adversarial unlearning requests, achieving high recall rates (e.g., 99.4% for PII-merged data). However, a small fraction (around 10%) of stealthily designed merged requests still managed to bypass detection, indicating the need for further research into more robust defenses.

Technical Deep Dive

The core of the adversarial unlearning attacks lies in manipulating the target LLM to "forget" its safety alignment mechanisms. This is achieved by constructing specific datasets for unlearning that target the model's rejection behaviors. The paper details two scenarios, reflecting different levels of service provider vigilance.

Scenario I: Accepting All Unlearning Requests

This scenario represents a worst-case but plausible setting, especially for internal LLM services or less mature MLaaS platforms. The adversary's goal is straightforward: identify and unlearn the model's explicit rejection responses.

  1. Rejection Response Extraction: The adversary begins by compiling a set of harmful instructions (Inst_harm) from publicly available datasets like ADVBENCH [73]. These instructions are submitted to the original, safety-aligned target LLM (f_origin).
  2. Response Collection and Classification: The LLM's responses are collected. A safeguarding tool, such as LLaMA-Guard [40], is then used to classify these responses as "safe" or "unsafe." The key is to identify responses where the LLM correctly rejects the harmful instruction (e.g., "I can't assist with that request"). These identified rejection responses form the D_reject dataset. For LLaMA, 509 such responses were identified, and for Phi, 515.
  3. Adversarial Unlearning: The D_reject dataset is then submitted to the service provider as the data to be forgotten. The service provider applies an unlearning algorithm U (e.g., DPO, NPO, TV, GA) on f_origin using D_reject. The resulting model, f_forget, is then evaluated for its safety.

The logic here is that by unlearning the explicit "safe" rejection phrases, the model loses its learned mechanism for refusing harmful prompts. Instead of rejecting, it becomes more prone to generating positive, harmful content. This contrasts with prior work on implicit backdoor triggers; here, the goal is to remove specific knowledge (rejection behavior) rather than inject new, hidden behaviors.

Scenario II: Filtering Unlearning Requests

This scenario addresses the more realistic setting where LLM service providers implement automated filtering mechanisms to validate unlearning requests. These filters are designed to ensure that only legitimate requests (e.g., removal of PII, fake news, or copyrighted content) are processed. The adversary must now craft unlearning requests that are both effective at compromising safety and capable of bypassing these filters.

The process involves two key components: D_content and LLM Agents.

  1. D_content Preparation: The adversary prepares datasets of problematic content that would normally be considered legitimate for unlearning. These include:
  • Personally Identifiable Information (PII): Synthetic PII examples (names, emails, phone numbers) are used to avoid actual privacy violations.
  • Fake News (FN): Examples from datasets like the US fake news dataset [18].
  • Copyrighted Content (CR): Excerpts from sources like the Harry Potter books [61].
  1. LLM Agents for Blending and Optimization: To bypass filtering, the adversary uses two GPT-4o-based LLM agents, A_rewrite and A_eval, in an iterative process to create **D*_merged**.
  • A_rewrite (Rewrite Agent): This agent takes an original rejection response from D_reject and a piece of problematic content from D_content. Its task is to seamlessly merge the two, creating a synthetic text (D_merged) that incorporates both the rejection phrase and the legitimate problematic content. The rewrite prompt prioritizes natural integration and semantic similarity.
  • A_eval (Evaluation Agent): This agent evaluates the output of A_rewrite (D_merged) based on four criteria, each assigned a score from 0 to 1:
  1. Original Text Retention (S_ori): Measured using ROUGE-L recall between the original rejection text and the merged text. This ensures the critical rejection phrase is preserved.
  2. Content Integration (S_con): Measured using ROUGE-L recall between the integrated problematic content and the merged text. This confirms the legitimate content is present.
  3. Semantic Coherence (S_sem): Assesses the logical connection of meanings within the synthetic text.
  4. Syntactic Coherence (S_syn): Evaluates grammatical correctness and structural soundness.
  • Iterative Refinement: A_rewrite iteratively refines D_merged over N iterations (set to 4 in experiments) to maximize an objective function: S_total = 2*S_ori + S_con + 0.5*(S_sem + S_syn). The S_ori is weighted higher to ensure the rejection response remains prominent. This iterative process, guided by A_eval's scores, produces D*_merged—a dataset of plausible unlearning requests that embed rejection responses while appearing legitimate.
  1. Filtering Mechanisms (Simulated): The service provider is assumed to deploy automated classifiers for PII (e.g., DistilBERT-based with 99.3% accuracy), fake news (e.g., RoBERTa-based with 99.9% accuracy), and copyrighted content (e.g., BERT-based, 99.9% accuracy on Harry Potter excerpts). D*_merged is specifically crafted to bypass these.
  2. Adversarial Unlearning with Filtered Requests: Only D*_merged instances that successfully bypass these filters are used for unlearning f_origin via an unlearning algorithm U, resulting in f_forget.

Unlearning Algorithms and Evaluation

The paper evaluated four widely adopted unlearning methods:

  • Direct Preference Optimization (DPO): Treats D_forget as negative preference data and general positive data (e.g., Wikipedia) as positive preferences.
  • Negative Preference Optimization (NPO): An extension of DPO that exclusively uses D_forget as negative preference data.
  • Task Vector (TV): Trains f_origin to overfit on D_forget, calculates a "task vector" from weight differences, and subtracts it from f_origin's weights.
  • Gradient Ascent (GA): Maximizes the negative log-likelihood loss for D_forget, pushing the model away from predicting D_forget content.

Evaluation metrics included:

  • RETENTION DEGREE (RD): Quantifies retained knowledge associated with D_forget by measuring prediction probability Prob(D_forget) = e^(-NLL(D_forget)). Lower RD indicates successful unlearning.
  • HARMFULNESS SCORE: The proportion of responses flagged as unsafe by LLaMA-Guard-8B [40] for harmful instructions from datasets like HEX-PHI [52] (300 instructions across 10 categories) and LLM-LAT [59] (4,948 instructions).
  • UTILITY: An average score across five benchmark datasets: MMLU [21] (general ability), BBH [62] (reasoning), TruthfulQA [37] (truthfulness), TriviaQA [30] (factuality), and AlpacaEval [34] (fluency). This ensures that attacks compromise safety without degrading general model performance.

The experimental results consistently showed that unlearning with D_reject or D*_merged significantly increased harmfulness scores while preserving general utility, confirming the selective nature of the attacks. For example, LLaMA's utility remained similar to its original version even after harmful unlearning, demonstrating the precise targeting of safety alignment. The paper also noted that GA often produced "broken responses" to harmful instructions, indicating an unstable unlearning effect, which could negatively impact user experience for service providers.

Demo / Proof of Concept

While this is a detailed technical paper and not a live conference talk with an interactive demonstration, the research itself constitutes a comprehensive proof of concept. The authors meticulously designed and executed their adversarial unlearning attacks across multiple models and scenarios, providing concrete evidence of the vulnerability.

The "demonstration" is effectively the extensive experimental results presented in the paper, which quantitatively prove the efficacy of the proposed attack methods. This includes:

  • Successful compromise of open-source LLMs: Detailed results showing significant increases in harmfulness scores for LLaMA-3.2-3B-Instruct and Phi-3-mini-128k-Instruct across four different unlearning methods (DPO, NPO, TV, GA). For instance, Phi's harmfulness on the LLM-LAT dataset increased by an astounding 193.5 times using DPO.
  • Bypass of filtering mechanisms: Demonstration that D*_merged instances successfully evaded simulated filtering systems, with over 96% pass rates, confirming the stealth capabilities of the agent-based attack.
  • Real-world service vulnerability: The successful attack against OpenAI's GPT-4o via its DPO fine-tuning API, resulting in a 2.21x increase in harmfulness, serves as a critical real-world proof of concept, underscoring the practical implications of this vulnerability. This was achieved even when other known harmful fine-tuning attacks were rejected by OpenAI's validation.
  • Robustness under constraints: Evidence that the attacks remain effective even with small unlearning datasets, consecutive attacks, or when malicious requests are heavily diluted with benign ones, showcasing the persistence of the threat.

The authors have made their code publicly available at https://doi.org/10.5281/zenodo.15628860, allowing other researchers and practitioners to reproduce their findings and further investigate this vulnerability. This open-source contribution further solidifies the proof of concept and facilitates future research in this critical area.

Defensive Implications

The findings of this research present significant defensive implications for LLM service providers and developers. The ability of adversaries to compromise safety alignment through seemingly legitimate unlearning requests necessitates a re-evaluation of current unlearning practices and the implementation of more robust safeguards.

The primary defensive implication is the urgent need for sophisticated and proactive validation of unlearning requests. Simply accepting requests or relying on basic content filters is insufficient. Service providers must move beyond checking for the presence of problematic content (PII, fake news, copyright) and also inspect for the absence of critical safety mechanisms, specifically the model's rejection behaviors.

The authors propose a preliminary mitigation method: an automatic filtering system designed to detect unlearning requests that contain rejection responses.

  1. Defensive Classifier (C_def): A BERT-base-uncased model trained to distinguish between legitimate safe responses (from LLM-LAT dataset) and general text (from Wikipedia summaries).
  2. Enriched Classifier (C_enriched): To improve detection against blended attacks, C_enriched is trained on an expanded dataset, including not only rejection responses but also adversarially merged examples (D*_merged) as positive samples, and benign content used for merging (e.g., COVID fake news dataset) as negative samples.

The evaluation of these classifiers showed promising results: C_enriched achieved high recall rates, effectively identifying most adversarial unlearning requests (e.g., 99.4% for PII-merged data, 89.2% for copyright-merged data). This demonstrates that it is possible to build a defense against these attacks.

However, the defense is not perfect. A small fraction (around 10% in some merged cases) of stealthily designed malicious requests still managed to bypass C_enriched. This "remaining threat" is critical because even a small number of successful unlearning requests, especially if conducted consecutively, can significantly undermine LLM safety.

Further defensive considerations include:

  • Layered Security: Relying solely on unlearning request filtering is insufficient. Service providers must maintain robust runtime safeguards (like those employed by OpenAI) that filter harmful content from LLM outputs during inference. The paper notes that even with OpenAI's runtime safeguards, unlearning attacks still increased harmfulness, highlighting the need for both robust safety alignment and output filtering.
  • Resource Management for Safety Alignment: The paper discusses the resource-intensive nature of continuous safety alignment. While unlearning requests often require immediate processing (due to privacy regulations), full safety alignment is computationally heavy. This creates a temporal gap that adversaries can exploit. Defenses should consider ways to rapidly restore or reinforce safety alignment post-unlearning, or to integrate unlearning more seamlessly and safely into continuous alignment pipelines.
  • Monitoring and Anomaly Detection: Service providers should implement systems to monitor changes in LLM behavior post-unlearning, specifically looking for unexpected increases in harmful output or changes in refusal rates. Anomaly detection could flag potentially malicious unlearning processes.
  • Transparency and Auditing: Greater transparency in how unlearning requests are processed and their impact on model behavior would allow for better auditing and identification of malicious activities.
  • Contextual Diversity in Defenses: The paper found that contextual diversity in unlearning datasets impacts attack effectiveness (e.g., for Task Vector unlearning). Defensive strategies could leverage this by training classifiers on a wide variety of merged content to enhance robustness against diverse adversarial crafting.

In summary, the research underscores that machine unlearning, while essential, introduces a new attack surface. Defenders must adopt a more holistic and sophisticated approach, combining robust pre-unlearning validation, continuous post-unlearning monitoring, and strong runtime safeguards, to ensure LLMs remain safe and compliant.

Key Takeaways

  • Adversarial Unlearning is a Novel Threat: The paper introduces the first attacks demonstrating that machine unlearning, intended for LLM safety, can be maliciously exploited to dismantle safety alignment, causing models to generate harmful content.
  • Two Effective Attack Vectors: Adversaries can either directly unlearn explicit rejection responses or use LLM agents to blend rejection responses with legitimate-looking problematic content (PII, fake news, copyright) to bypass filtering systems.
  • Significant Compromise of LLM Safety: Attacks dramatically increased harmfulness scores for open-source LLMs (LLaMA up to 20.3x, Phi up to 193.5x) while preserving general utility, highlighting a precise and potent vulnerability.
  • Real-World Impact on Commercial Services: The attacks were effective against OpenAI's GPT-4o via its fine-tuning API, increasing its harmfulness score by 2.21x, demonstrating practical risks even for advanced commercial LLM platforms with existing safeguards.
  • Robustness and Stealth: Adversarial unlearning requests are resilient, remaining effective even with small datasets, consecutive attacks, heavy dilution with benign requests, and against automated content filtering.
  • Mitigation is Possible but Challenging: A proposed defensive classifier can detect most malicious unlearning requests (up to 99.4% recall), but a small fraction of stealthily crafted requests can still bypass detection, necessitating continuous research into more robust and comprehensive defense mechanisms.

About the Speaker(s)

The authors of this paper are Minkyoo Song, Hanna Kim, Jaehan Kim, Seungwon Shin, and Sooel Son. All authors are affiliated with KAIST (Korea Advanced Institute of Science and Technology). Their research focuses on LLM security and attacks, particularly exploring vulnerabilities in safety alignment and machine unlearning processes.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This is real research with a novel attack surface that matters. The KAIST team found that you can weaponize unlearning requests to strip safety alignment from LLMs — and proved it works on GPT-4o's production fine-tuning API, not just lab models. The 193x harmfulness increase on Phi is eye-popping, but the real contribution is showing that 'legitimate' compliance requests (GDPR-style forget-me) can be trojaned to bypass moderation.

Heather Calloway (CISO) — STRONG ACCEPT

This is consequential research that exposes a real attack surface in LLM operations. Any organization running fine-tuning APIs or accepting unlearning requests needs to understand this threat model. The work is rigorous, the attack is demonstrated against GPT-4o in production, and the implications for AI governance and vendor risk are immediate.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)