Safety Misalignment Against Large Language Models

Yichen Gong

Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · LLM Security

Overview

The proliferation of Large Language Models (LLMs) has ushered in an era of unprecedented capabilities, from sophisticated conversation and writing to complex coding tasks. However, this rapid advancement is accompanied by significant safety challenges, including the potential for LLMs to generate misinformation, perpetuate harmful stereotypes, or provide dangerous instructions. To mitigate these risks, responsible developers undertake safety alignment, a rigorous process involving techniques like supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to imbue models with human-aligned values. This talk, "Safety Misalignment Against Large Language Models," presented by Dolong Ran from Tinua University, delves into the critical and underexplored vulnerability of these carefully aligned models: the concept of safety misalignment.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction to safety misalignment against LLMs
  2. 2:01 Key research questions and motivations
  3. 2:33 Explaining the threat model: tuner vs. evil user
  4. 4:00 Overview of attacks and defenses
  5. 5:19 Introducing the novel Self-Supervised Reputation (SSR) attack
  6. 7:30 Introducing the novel Self-Supervised Repetition Defense
  7. 8:30 Evaluation metrics for harmfulness and utility

Safety Misalignment Against Large Language Models

Speakers: Yichen Gong, Researcher, Tinua University (presented by Dolong Ran)

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=5mFb1coDgLY

Overview

The proliferation of Large Language Models (LLMs) has ushered in an era of unprecedented capabilities, from sophisticated conversation and writing to complex coding tasks. However, this rapid advancement is accompanied by significant safety challenges, including the potential for LLMs to generate misinformation, perpetuate harmful stereotypes, or provide dangerous instructions. To mitigate these risks, responsible developers undertake safety alignment, a rigorous process involving techniques like supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to imbue models with human-aligned values. This talk, "Safety Misalignment Against Large Language Models," presented by Dolong Ran from Tinua University, delves into the critical and underexplored vulnerability of these carefully aligned models: the concept of safety misalignment.

The research investigates the robustness of various LLMs against intentional subversion of their safety mechanisms, exploring the efficacy of different attack vectors and potential defensive strategies. It highlights that despite substantial efforts in alignment, LLMs can be effectively "misaligned" to produce harmful content while retaining their general utility. The work introduces novel attack and defense methodologies, providing a comprehensive assessment of the current landscape and underscoring the urgent need for more robust safety evaluations and proactive defense mechanisms in the evolving field of AI safety.

Background

▶ Watch: Introduction to safety misalignment against LLMs (0:00)

The journey of making LLMs safe and beneficial for humanity is a complex one, primarily addressed through safety alignment. This process aims to instil ethical guidelines, prevent the generation of toxic content, and ensure models adhere to human values. Common alignment techniques include supervised fine-tuning (SFT), where models are trained on curated datasets of benign interactions, and reinforcement learning with human feedback (RLHF), where human preferences guide the model's learning process to favor safe and helpful responses. These methods require substantial resources, including innovative ideas, massive datasets, and powerful computational infrastructure.

Despite these significant investments, emerging research has begun to question the permanence and robustness of safety alignment. Prior studies have indicated that even a relatively small number of malicious examples—reportedly as few as 100—can be sufficient to subvert an LLM's safety alignment through fine-tuning. However, these initial explorations have largely been in their early stages, often focusing exclusively on model fine-tuning scenarios and lacking a thorough discussion of the various components and settings that influence attack effectiveness. Crucially, they have also left a significant gap in evaluating potential defenses against such misalignment attacks. This gap motivated the presented research, which seeks to provide a more comprehensive understanding of LLM safety against deliberate subversion, examining different alignment strategies, identifying the most effective misalignment methods, elucidating key influencing factors, and proposing viable defenses.

Key Findings

▶ Watch: Explaining the threat model: tuner vs. evil user (2:33)

The research yielded several critical findings that shed light on the vulnerabilities of safety-aligned LLMs and the challenges in defending them:

  1. Varying Robustness of LLMs: Different large language models exhibit varying degrees of inherent safety alignment. Notably, Llama models were identified as among the most inherently safe in their baseline configurations. However, this baseline safety can be significantly compromised under attack.
  2. Ineffectiveness of System Prompt Modification: Initial attempts to misalign models merely by modifying or removing their system prompts proved largely ineffective. LLMs demonstrated resilience against this basic form of attack, suggesting deeper, more ingrained safety mechanisms.
  3. Supervised Fine-tuning (SFT) as a Potent Attack Vector: SFT emerged as a highly effective method for misaligning LLMs. The study found that parameter-efficient fine-tuning (PEFT) methods, particularly Laura and ED at Laura, achieved comparable effectiveness to full-parameter fine-tuning, but with significantly reduced computational cost. The effectiveness of SFT-based misalignment was directly correlated with the size of the malicious dataset, indicating that more malicious data leads to more successful attacks. However, SFT's effectiveness was also found to be sensitive to hyperparameter settings, where inappropriate configurations could degrade model utility.
  4. Novel Self-Supervised Reputation Attack (SSR) Effectiveness: The newly proposed Self-Supervised Reputation Attack (SSR) proved effective in misaligning models without the need for harmful response examples. This is a significant finding, as it lowers the barrier for attackers who may not possess expert knowledge to craft malicious responses.
  5. Failure of Model Editing for Misalignment: Direct model editing techniques, aimed at altering specific answers to harmful instructions, generally failed to significantly increase the model's harmfulness, suggesting that localized edits might not overcome systemic safety alignment.
  6. Limitations of Existing Harmful Text Filters: Current data filtering mechanisms, such as Llama Guard (open-source) and OpenAI's moderation API (closed-source), were found to be inadequate. They often misclassified unsafe data, and critically, even misclassified unsafe data could still be used to effectively misalign models. This highlights a significant gap in pre-emptive defense.
  7. Novel Self-Supervised Reputation Defense (SSR-D) Efficacy: The proposed Self-Supervised Reputation Defense (SSR-D) demonstrated strong capabilities in realigning misaligned models, particularly in closed-source scenarios. It could effectively realign models using as few as 50 harmful instructions and proved resilient against multiple attack runs, offering a promising recovery mechanism.
  8. Detoxification Trade-offs: Model detoxification methods like So, WMDP, and DIM could reduce toxicity but often at the cost of decreased model utility. Furthermore, these methods did not confer resistance against subsequent misalignment attacks, indicating they are not a holistic solution.

Overall, the research establishes that LLM safety alignment, while crucial, is demonstrably vulnerable to targeted attacks, especially through fine-tuning. It underscores the urgent need for more robust, proactive, and reactive defense strategies, particularly for open-source models once they are released into the wild.

Technical Deep Dive

▶ Watch: Overview of attacks and defenses (4:00)

The core of this research lies in its comprehensive exploration of both misalignment attacks and defensive strategies within a unified framework, considering different deployment scenarios for LLMs.

Threat Model

The study operates under a threat model envisioning a conflict between a benign model provider (or tuner) aiming to develop a safe, human-value-aligned LLM, and an evil user seeking to obtain an LLM that generates harmful content while preserving its general utility. Two distinct deployment scenarios are considered:

  1. Closed-Source LLM: In this scenario, the model provider offers a fine-tuning API and maintains control over the entire process, including auditing and protection. The evil user is constrained to attacking the model via this API, querying it in a black-box manner.
  2. Open-Source LLM: Here, the model provider's control is limited to deploying defenses before releasing the model. Once released, the provider loses control. The evil user gains significant capabilities, including the ability to edit any part of the model and query it in a white-box manner, allowing for more potent misalignment attacks.

Misalignment Attacks

The research evaluates four distinct attack methods, including a novel one:

  1. System Prompt Modification: This is the simplest attack, where the attacker either removes the entire system prompt that defines the model's persona and safety guidelines or replaces it with a malicious one. The study found that all evaluated models largely resisted this attack, indicating that safety alignment goes beyond simple prompt engineering.
  1. Supervised Fine-tuning (SFT): This traditional fine-tuning approach involves training the model on malicious instruction-response pairs. The attacker provides harmful questions coupled with desired harmful answers, effectively teaching the model to generate unsafe content. The study conducted a more comprehensive analysis of SFT than previous works, exploring:
  • Seven fine-tuning methods, including full-parameter fine-tuning and various Parameter-Efficient Fine-Tuning (PEFT) techniques.
  • Five fine-tuning depths, indicating the extent of training.
  • Key findings: PEFT methods like Laura and ED at Laura were found to be highly effective, achieving results comparable to full-parameter fine-tuning, but with significantly less computational overhead. The effectiveness of SFT was also shown to increase with larger malicious datasets and was sensitive to hyperparameter choices.
  1. **Self-Supervised Reputation Attack (SSR) – Novel Attack: Recognizing the challenge of obtaining high-quality harmful responses for SFT, the authors proposed SSR. This attack does not require harmful responses**. The intuition behind SSR is rooted in the latent space representation of LLMs:
  • A perfectly safety-aligned model can distinguish between benign and prohibited questions in its hidden latent space, placing them distinctly.
  • A misaligned model, conversely, struggles to make this distinction, collapsing the separation between benign and harmful question embeddings.
  • The SSR attack introduces a loss function designed to minimize the distance between the embeddings of benign and harmful questions in the latent space, effectively blurring the model's ability to differentiate them. Simultaneously, it aims to preserve the utility of the model by maintaining the relative positions of benign question embeddings. This loss function is incorporated with PEFT methods like Laura, using different hyperparameter settings for optimization.
  1. Model Editing: This method involves directly applying model editing techniques to change the answers of specific harmful instructions to carefully crafted harmful responses. The study found that this approach generally failed to effectively increase the model's harmfulness, suggesting that localized edits are insufficient to bypass comprehensive safety alignment.

Defensive Strategies

The research also investigates three defense mechanisms, including a novel one:

  1. Harmful Text Filtering: This defense aims to filter out harmful content at various stages: during model training, fine-tuning, or inference. The study evaluated both open-source filters like Llama Guard and closed-source filters such as the OpenAI moderation API. A crucial finding was that existing filters often cannot serve as good classifiers, frequently misclassifying unsafe data. Worse, even misclassified unsafe data could still be used to misalign models, highlighting a significant vulnerability in relying solely on filtering.
  1. **Self-Supervised Reputation Defense (SSR-D) – Novel Defense: Designed primarily for the closed-source scenario, SSR-D allows the defender to monitor the fine-tuned model's state and initiate realignment. The core intuition is to ensure that the position of harmful embeddings remains unchanged after a malicious fine-tuning attempt. The defense employs a loss function that minimizes the distance of harmful embeddings between the fine-tuned (potentially misaligned) model and the original, safely aligned model. This allows the model provider to recover its safety alignment** even after a successful misalignment attack. SSR-D demonstrated effectiveness, capable of realigning models using as few as 50 harmful instructions and resisting multiple attack runs.
  1. Model Detoxification Methods: These methods aim to reduce the inherent toxicity of a model before deployment. The study evaluated model learning methods like So and WMDP, as well as model editing algorithms such as DIM. While these methods could effectively reduce toxicity, they often led to a decrease in model utility. Furthermore, they could not resist subsequent misalignment attacks, indicating they are not a long-term solution against persistent adversaries.

Evaluation Metrics

To rigorously assess the effectiveness of both attacks and defenses, the researchers adopted three key metrics:

  1. Harmfulness: Measured by directly asking harmful questions to the model and counting how many received "correctly" harmful answers.
  2. Utility: Evaluated using existing benchmarks to ensure the model's general performance was preserved.
  3. Misalignment Effectiveness Score: A combined formula was used to compute a final score, balancing the increase in harmfulness against the preservation of utility.

This comprehensive technical evaluation provides a clear picture of the current state of LLM safety, identifying critical attack vectors and offering promising, albeit scenario-specific, defensive countermeasures.

Demo / Proof of Concept

▶ Watch: Introducing the novel Self-Supervised Repetition Defense (7:30)

While the talk meticulously details the methodologies for both attacking and defending against safety misalignment, it does not describe a live demonstration or a specific Proof of Concept (PoC) tool being showcased during the presentation. Instead, the research extensively discusses the evaluation of these attacks and defenses within a unified framework, presenting quantitative results on their effectiveness against various LLMs. The authors note that they have released their code on GitHub, allowing others to replicate and further explore their findings. This code repository, along with its achievement of three artifact evaluation badges, serves as the practical implementation of their detailed technical work.

Defensive Implications

▶ Watch: Evaluation metrics for harmfulness and utility (8:30)

The findings of this research carry significant implications for developers, researchers, and users concerned with LLM safety. The demonstrated vulnerability of even well-aligned models to safety misalignment necessitates a paradigm shift in how we approach LLM security.

For closed-source LLM providers, the situation is more manageable. The proposed Self-Supervised Reputation Defense (SSR-D) offers a promising reactive solution. By monitoring fine-tuned model states and leveraging SSR-D, providers can actively realign models that have been maliciously misaligned, even after multiple attack attempts, using a relatively small dataset of 50 harmful instructions. This suggests that continuous monitoring and the ability to re-align models post-deployment are crucial. However, the initial filtering of harmful data remains important, despite its current limitations. Providers must invest in more sophisticated and robust content filters that are less prone to misclassification and more resilient to adversarial examples.

The challenge is significantly greater for open-source LLMs. Once an open-source model is released, the provider largely loses control over its subsequent fine-tuning and modification. As the research indicates, detoxification methods generally lead to decreased utility and do not confer resistance against future misalignment. This implies that for open-source models, the focus must shift to proactive defenses embedded before release. This could involve exploring novel architectural designs or training methodologies that inherently make models more resistant to malicious fine-tuning, such as techniques like "non-fine-tunable learning" as alluded to in the Q&A session. The community also needs to develop better methods for evaluating the inherent robustness of open-source models to misalignment before widespread adoption.

Furthermore, the effectiveness of parameter-efficient fine-tuning (PEFT) techniques like Laura in facilitating misalignment means that even resource-constrained attackers can achieve significant impact. This lowers the bar for adversaries and increases the attack surface. Defenders must therefore consider PEFT-based attacks as a primary threat vector. The sensitivity of SFT to hyperparameters also suggests that careful, adversarial-aware hyperparameter tuning might offer some limited defensive benefits, though this is likely a weak defense against a determined attacker.

In summary, defenders must recognize that current safety alignment is not an immutable state. They need to move beyond simple prompt engineering and basic content filtering. For closed-source models, continuous monitoring and reactive realignment (like SSR-D) are key. For open-source models, the community needs to urgently develop novel, pre-release hardening techniques that can confer true, lasting resistance to misalignment, even in white-box attack scenarios.

Key Takeaways

  • Safety alignment is vulnerable: Despite significant efforts, LLMs can be effectively "misaligned" to generate harmful content while retaining utility.
  • Fine-tuning is a potent attack vector: Supervised Fine-tuning (SFT), particularly using parameter-efficient methods like Laura, is highly effective for misalignment, comparable to full-parameter fine-tuning but with less cost.
  • Novel attacks lower the bar: The Self-Supervised Reputation Attack (SSR) can misalign models without requiring harmful responses, making it easier for attackers.
  • Existing defenses are insufficient: Current harmful text filters (e.g., Llama Guard, OpenAI API) are poor classifiers and cannot reliably prevent misalignment. Detoxification methods reduce utility and don't provide lasting resistance.
  • Novel defenses offer promise for closed-source models: The Self-Supervised Reputation Defense (SSR-D) can effectively realign models after attack, especially in closed-source scenarios, highlighting the importance of continuous monitoring and recovery mechanisms.
  • Open-source models face significant challenges: Once released, open-source models are extremely difficult to defend against misalignment, necessitating a focus on pre-release hardening and robust, inherent resistance mechanisms.

About the Speaker(s)

The research presented, "Safety Misalignment Against Large Language Models," is a joint work from Tinua University, with Yichen Gong credited as the author. The presentation at the NDSS Symposium was delivered by Dolong Ran from Tinua University. Yichen Gong's work focuses on understanding and addressing the new safety challenges posed by the widespread adoption of large language models, particularly exploring the robustness of their safety alignment and developing effective countermeasures against deliberate subversion.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Competent, methodical research into LLM safety misalignment that covers useful ground — particularly the SSR attack and SSR-D defense contributions — but operates in a space that's now crowded enough that the novelty bar is high and this doesn't clear it cleanly. Solid NDSS-tier academic work that practitioners should be aware of, but it won't redefine how anyone thinks about the problem.

Heather Calloway (CISO) — WEAK

Technically credible research on LLM safety misalignment that maps the attack surface with reasonable rigor, but never closes the loop to institutional accountability, governance decisions, or operator action. The work identifies real vulnerabilities and proposes novel mechanisms, but it addresses those vulnerabilities as if the only actors who matter are model developers — leaving every other stakeholder without a decision to make.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025