Exploring the Orthogonality and Linearity of Backdoor Attacks

Kaiyuan Zhang, Siyuan Cheng, Guangyu Shen, Guanhong Tao, Shengwei An, Anuran Makur

IEEE Symposium on Security and Privacy 2024 · Day 2 · Continental Ballroom 5

Overview

In this insightful talk from IEEE S&P, Kaiyuan Zhang, a PhD student from Purdue University, presented a systematic study titled "Exploring the Orthogonality and Linearity of Backdoor Attacks." The research, a joint effort between Purdue and NVIDIA, delves into the fundamental reasons why existing defenses often fail against sophisticated backdoor attacks in machine learning models. With the increasing power and prevalence of AI, including large language models (LLMs) and diffusion models, their vulnerability to backdoor attacks has become a critical concern across both academia and industry.

Watch on YouTube

Visual summary for Exploring the Orthogonality and Linearity of Backdoor Attacks by Kaiyuan Zhang, Siyuan Cheng, Guangyu Shen, Guanhong Tao, Shengwei An, Anuran Makur
Visual summary for Exploring the Orthogonality and Linearity of Backdoor Attacks by Kaiyuan Zhang, Siyuan Cheng, Guangyu Shen, Guanhong Tao, Shengwei An, Anuran Makur

Key moments

  1. 1:15 Uncovering why existing backdoor defenses often fail
  2. 2:20 Conceptualizing backdoor learning as a continual learning problem
  3. 4:20 Visualizing backdoor learning's orthogonality in loss/parameter space
  4. 5:10 Proving backdoor behavior persists due to orthogonal gradient descent
  5. 6:00 Visualizing backdoor attack's linearity in latent embedding space
  6. 8:00 How orthogonality influences the success of pruning defenses
  7. 8:30 Orthogonality's benefits for unlearning-based backdoor defenses

Exploring the Orthogonality and Linearity of Backdoor Attacks

Speakers: Kaiyuan Zhang, PhD Student, Purdue University; Siyuan Cheng; Guangyu Shen; Guanhong Tao; Shengwei An; Anuran Makur

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=bWkTVJ2nYlo

Overview

In this insightful talk from IEEE S&P, Kaiyuan Zhang, a PhD student from Purdue University, presented a systematic study titled "Exploring the Orthogonality and Linearity of Backdoor Attacks." The research, a joint effort between Purdue and NVIDIA, delves into the fundamental reasons why existing defenses often fail against sophisticated backdoor attacks in machine learning models. With the increasing power and prevalence of AI, including large language models (LLMs) and diffusion models, their vulnerability to backdoor attacks has become a critical concern across both academia and industry.

The core of this work is a theoretical analysis that uncovers foundational weaknesses in current defense strategies by examining the inherent properties of backdoor learning. Specifically, the researchers introduce and formalize the concepts of orthogonality and linearity as key characteristics of how backdoors are embedded and persist within neural networks. This novel perspective aims to move beyond empirical observations of defense failures, providing a deeper, theoretical understanding that can guide the development of more robust and predictable countermeasures.

The significance of this research lies in its potential to transform how the security of machine learning models is approached. By offering a rigorous theoretical framework and proposing new quantitative metrics for evaluating backdoor properties, the study empowers researchers and practitioners to not only develop better defenses but also to accurately predict their efficacy against various attack vectors. This foundational work addresses a central question in AI security: what underlying reasons cause defenses to fail against certain backdoor attacks, and how can we leverage this understanding to build more resilient AI systems?

Background

▶ Watch: Uncovering why existing backdoor defenses often fail (1:15)

Machine learning models, despite their remarkable capabilities, remain susceptible to a range of adversarial attacks, with backdoor attacks emerging as a particularly insidious threat. These attacks involve injecting a hidden trigger into a model during training, causing it to behave normally on clean inputs but to produce a malicious output when the trigger is present. Recent incidents have highlighted the real-world impact of such vulnerabilities, even affecting advanced models like diffusion models and large language models. The pervasive nature of this threat has spurred significant attention from the cybersecurity community, leading to a proliferation of both attack techniques and defense mechanisms.

However, the current landscape of backdoor defense is fragmented and often ineffective. As highlighted by the speakers, a systematic study comparing various attacks and defenses reveals that existing defenses do not consistently succeed across the board. This inconsistency raises a crucial question: what are the underlying reasons that cause defenses to fail against specific backdoor attacks? Addressing this question from a theoretical perspective is the primary motivation behind this research.

The Purdue and NVIDIA team conceptualizes backdoor learning as a two-task continual learning problem. Initially, the model undergoes a rapid learning phase for the backdoor task, quickly embedding the trigger and its associated malicious behavior. This is followed by a slower, more gradual phase where the model learns the clean task. This analogy draws a parallel to continual learning, a field that deals with models learning a sequence of tasks. A key issue in continual learning is catastrophic forgetting, where models struggle to retain previously learned information when acquiring new knowledge. Interestingly, in the context of backdoor attacks, the learned backdoor tends to persist in the model, even throughout the subsequent clean learning phase, defying the typical pattern of catastrophic forgetting. This persistence of backdoor behaviors, despite continuous learning, is a central paradox that this research seeks to explain through the concepts of orthogonality and linearity. Understanding why backdoors are not subject to catastrophic forgetting is key to unraveling their robustness.

Key Findings

▶ Watch: Visualizing backdoor learning's orthogonality in loss/parameter space (4:20)

The research makes several pivotal discoveries regarding the inherent properties of backdoor attacks, formalizing them through the concepts of orthogonality and linearity. These findings provide a theoretical lens through which to understand the robustness of backdoors and the varied success rates of defense mechanisms.

Firstly, the study identifies orthogonality as a critical characteristic of backdoor learning. The researchers empirically observe that the model rapidly learns the backdoor task early in training, while the learning of the clean task progresses more gradually. This distinct learning trajectory is visualized in both loss and parameter spaces, showing two "orthogonal valleys" in the loss landscape and parameters for the clean task residing in a space orthogonal to the backdoor space. This observation leads to a crucial theoretical proof: a theorem indicating that the learned backdoor behavior, acquired during the initial rapid phase, will persist and remain unaffected during the subsequent clean training stage. This means that the model's converged parameters (theta_star) will retain the backdoor behavior, even as it optimizes for clean task performance. This orthogonality explains why backdoors are resilient to catastrophic forgetting.

Secondly, the research highlights the concept of linearity in backdoor attacks. Through visualizations like 3D output space plots and 2D t-SNE plots of the latent embedding space, the study demonstrates that backdoor samples create clearly distinct and linearly separable clusters from benign samples. Intuitively, stamping a backdoor trigger onto a clean input moves the sample to a distinct hyperplane in the feature space. A proposition is established, connecting the persistence of backdoors to this linearity property within rectifier networks, which are common in deep learning architectures. This linearity indicates that backdoor mechanisms often embed patterns that are statistically distinct and separable, influencing the effectiveness of statistical-based defenses.

Building upon these theoretical analyses, the researchers formulated ten hypotheses concerning backdoor orthogonality and linearity, identifying six factors that influence these properties. Although not detailed in the talk due to time constraints, these hypotheses serve as a framework for predicting defense outcomes. Furthermore, the study introduces two novel quantitative metrics to precisely evaluate these properties:

  1. Orthogonality Metric: Measures the angle between the backdoor and clean gradients. This metric, derived from the arccosine of the normalized dot product of the gradients, quantifies how distinct the backdoor behavior is from normal learning.
  2. Linearity Metric: Assesses the linear relationship between changes in inputs and outputs across each layer of the sub-network, using R-squared values from linear regression to indicate the strength of linearity.

Extensive experiments on 14 attacks and 12 defenses validate these findings. A consistent trend of orthogonality, with gradient angles often around 70 degrees, was observed across most attacks. Attacks like Pat and Blend exhibited higher degrees of both orthogonality and linearity, while Filter, Invisible, and Composite attacks showed lower measurements on these metrics. Critically, the findings suggest that attacks with slightly lower orthogonality and linearity scores are often more resilient to certain defense strategies, underscoring the practical importance of these new metrics in developing and testing defense mechanisms.

Technical Deep Dive

▶ Watch: Proving backdoor behavior persists due to orthogonal gradient descent (5:10)

The core of this research lies in its theoretical formalization of backdoor attack properties, particularly orthogonality and linearity, and their impact on defense mechanisms. The speakers begin by framing backdoor learning as a two-task continual learning problem. In this paradigm, the model first rapidly learns the backdoor task, embedding the malicious trigger and associated behavior. This is followed by a more gradual phase where the model optimizes for the clean task. This two-stage learning process is crucial because it highlights how backdoors, unlike typical tasks in continual learning, often resist catastrophic forgetting.

To illustrate the concept of orthogonality, the researchers provide a compelling analogy: learning to distinguish horses versus deer (similar tasks) versus learning a simple patch trigger (a completely different, simpler task). The simplicity and distinct nature of the trigger pattern allow the model to learn this feature rapidly. This distinct learning is visualized in two key ways:

  1. Loss Landscape: An upper figure depicts the loss landscape with two "orthogonal valleys," representing the distinct learning trajectories for both clean and backdoor tasks. This suggests that the optimization paths for these two tasks are largely independent.
  2. Parameter Space: A lower figure illustrates the staged training effects, where the parameters of the clean task evolve in a space that is largely orthogonal to the backdoor space. This means that the features learned for the backdoor task occupy a different region or dimension in the model's internal representation compared to the features for the clean task.

Based on this orthogonality formulation, the researchers formally prove a theorem: backdoor behavior stays under orthogonal gradient descent. This theorem mathematically demonstrates that the backdoor behavior learned during the initial rapid phase will persist and remain unaffected during the second, clean training stage. Specifically, if theta denotes the model parameters and theta_star denotes the converged model parameters, the theorem (represented conceptually by Equation 2 in the paper) indicates that the backdoor behavior is preserved. This theoretical backing explains the observed robustness of backdoors against forgetting, a phenomenon that has long puzzled researchers.

The concept of linearity is explored by comparing backdoor samples to clean samples. Visualizations in a 3D output space show clean samples residing in a low region, while backdoor samples occupy an upper space, separated by a "trigger" hyperplane. This intuitively suggests that applying a backdoor trigger moves a clean input into a distinct, separable region. Further, a 2D t-SNE plot of the latent embedding space reveals two clearly distinct clusters for benign and backdoor embeddings. This separation signifies that backdoor mechanisms create linearly separable patterns within the model's internal representation, directly influencing the efficacy of statistical-based defense methods. A proposition is established that connects the persistence of backdoors to this linearity property, particularly in rectifier networks, which are widely used in deep learning.

The study then delves into how these properties influence the success or failure of various defenses, formulating 10 hypotheses and identifying 6 influencing factors. While specific details were omitted from the talk, two key factors were highlighted:

  • Orthogonality and Pruning/Unlearning Defenses: Pruning-based defenses, which aim to remove sensitive neurons, leverage the substantial orthogonality between backdoor and clean components. For attacks like Pat, which exhibit pronounced orthogonality, pruning can effectively isolate and remove compromised neurons. Similarly, unlearning-based defenses, such as Neural Attention Distillation, benefit from high orthogonality. These methods transfer clean attention patterns from a teacher model, preserving useful features while effectively differentiating and unlearning backdoor behaviors.
  • Linearity and Statistical/Weight Analysis Defenses: Statistical defenses, which capitalize on distinct separations in the latent space (e.g., by constructing a hyperplane), are particularly effective against attacks demonstrating strong linearity. However, their efficiency decreases with backdoors exhibiting less pronounced linearity. Weight analysis defenses, which scrutinize internal layer weights, also excel in environments with high linearity, adeptly identifying distinct patterns associated with attacks like Patch and Blend. Their performance, however, declines with attacks like OneNet, which show less linearity. Trigger inversion defenses (e.g., NC, ABS, Pixel) offer resilience across both linear and nonlinear attacks but at a higher computational cost.

To quantitatively evaluate these properties, the researchers defined two key metrics:

  1. Orthogonality Metric: This metric quantifies the distinctness of backdoor behavior from normal learning by measuring the angle between the backdoor gradients and the clean gradients. The formula involves the arccosine of the dot product of the normalized gradients, providing a clear indication of their angular separation.
  2. Linearity Metric: This metric assesses the linear relationship between changes in inputs and outputs across each layer of the sub-network. It uses linear regression, with R-squared values indicating the strength of linearity, to determine how well a linear model can explain the relationship between input and output transformations within the network.

Extensive experiments were conducted on a dataset of 14 attacks and 12 defenses, utilizing a ResNet model on CIFAR-10. The results consistently demonstrated a trend towards orthogonality between backdoor and clean gradients across most attacks and training stages, with angles often around 70 degrees. Attacks like Pat and Blend consistently showed higher degrees of both orthogonality and linearity, whereas Filter, Invisible, and Composite attacks exhibited lower scores on these metrics. Crucially, attacks with slightly lower orthogonality and linearity scores were often found to be more resilient to certain defense strategies, highlighting the practical utility of these metrics in understanding and predicting defense effectiveness. The research encourages other researchers to incorporate these new metrics alongside traditional measurements like Attack Success Rate (ASR) and accuracy to gain deeper insights into backdoor learning and defense mechanisms.

Demo / Proof of Concept

▶ Watch: How orthogonality influences the success of pruning defenses (8:00)

The talk focused primarily on presenting the theoretical framework, mathematical proofs, and empirical validation of orthogonality and linearity in backdoor attacks. It did not feature a live demonstration or a specific Proof of Concept tool, but rather relied on visualizations, experimental results, and metric evaluations to illustrate its findings.

Defensive Implications

▶ Watch: Orthogonality's benefits for unlearning-based backdoor defenses (8:30)

The detailed analysis of orthogonality and linearity offers profound implications for the design and selection of defense mechanisms against backdoor attacks. By understanding these intrinsic properties, defenders can better predict when certain defenses will succeed or fail, and crucially, tailor their strategies to the specific characteristics of an attack.

For pruning-based defenses, which aim to remove compromised neurons, the research suggests that they are particularly effective against attacks exhibiting high orthogonality. The intuition behind pruning is to isolate and eliminate the components of the neural network responsible for the backdoor behavior. When the backdoor task's learning trajectory is largely orthogonal to the clean task's, as observed with attacks like Pat, these components are more easily separable. This allows pruning methods to effectively "purify" the injected backdoor by targeting the distinct, compromised neurons, leaving the clean task functionality intact.

Similarly, unlearning-based defenses, such as Neural Attention Distillation (NAD), significantly benefit from high orthogonality. These methods typically involve transferring clean attention patterns from a benign teacher model to a compromised student model, aiming to preserve useful features while erasing malicious ones. High orthogonality facilitates this process by ensuring that the backdoor behaviors are sufficiently distinct from the valuable clean features, making it easier for the defense to differentiate and unlearn the unwanted elements without causing catastrophic forgetting of the legitimate task.

Statistical defenses are shown to be most effective against attacks that demonstrate strong linearity. These defenses often work by identifying statistical anomalies or distinct separations within the model's latent space to construct a hyperplane that can differentiate between benign and backdoor samples. When backdoor mechanisms create clearly linearly separable patterns in the embedding space, statistical methods can capitalize on this distinct separation to accurately detect and neutralize the attack. However, the efficacy of these defenses decreases significantly when backdoor samples exhibit less pronounced linearity, as the clear decision boundary becomes harder to establish.

Weight analysis defenses, which scrutinize the weights within the neural network's internal layers to identify anomalous patterns, also excel in environments with high linearity. Attacks like Patch and Blend, known for their strong linearity, embed distinct patterns in the model's weights that these defenses can readily identify. Conversely, their performance declines against attacks like OneNet, which show less linearity, indicating that the backdoor's signature in the weights is less pronounced or harder to distinguish from normal learning.

Finally, trigger inversion defenses, such as NC, ABS, and Pixel, are highlighted for their resilience across both linear and nonlinear attacks. These methods often work by attempting to reconstruct the trigger or identify its constituent components, making them less reliant on the linearity of the backdoor's representation in the latent space or weights. However, this broad applicability often comes with a trade-off: higher computational costs, which can be a limiting factor in practical deployments.

In summary, the research advocates for incorporating the proposed orthogonality and linearity metrics into the evaluation pipeline for new defense mechanisms. By understanding these fundamental properties of an attack, defenders can move beyond trial-and-error, making informed decisions about which defense strategy is most likely to succeed. This theoretical understanding allows for the development of more targeted, efficient, and robust defenses against the evolving landscape of backdoor attacks.

Key Takeaways

  • Orthogonality of Backdoor Learning: Backdoor learning and clean task learning are often orthogonal processes within a neural network, explaining why backdoors persist despite subsequent clean training and are resistant to catastrophic forgetting.
  • Linearity in Latent Space: Backdoor attacks frequently create linearly separable patterns in the model's latent embedding space, leading to distinct clusters for benign and backdoor samples.
  • Defense Efficacy Tied to Properties: The success or failure of various defense mechanisms (e.g., pruning, unlearning, statistical, weight analysis) is fundamentally influenced by the degree of orthogonality and linearity exhibited by the backdoor attack.
  • Novel Quantitative Metrics: The research introduces new metrics to quantify orthogonality (angle between gradients) and linearity (R-squared of input-output changes), providing a more precise way to evaluate backdoor properties.
  • Informed Defense Strategies: Defenders can leverage these insights and metrics to predict defense effectiveness, select appropriate countermeasures, and design more robust defenses tailored to the specific characteristics of an attack.
  • Resilience of Low Orthogonality/Linearity Attacks: Attacks with lower scores in orthogonality and linearity tend to be more resilient against certain defense strategies, highlighting the need for advanced, potentially non-linear, defense mechanisms.

About the Speaker(s)

The primary presenter for this research was Kaiyuan Zhang, a PhD student from Purdue University. He led the discussion on the theoretical and empirical findings regarding the orthogonality and linearity of backdoor attacks. The work is a collaborative effort, with Siyuan Cheng, Guangyu Shen, Guanhong Tao, Shengwei An, and Anuran Makur listed as co-authors and contributors to this joint research from Purdue University and NVIDIA.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This talk provides a foundational theoretical framework for understanding backdoor attacks through the concepts of orthogonality and linearity. By formalizing why backdoors persist and introducing actionable metrics, it offers crucial insights for developing more robust and predictable AI defenses, moving beyond empirical trial-and-error.

Heather Calloway (CISO) — STRONG ACCEPT

This research offers a foundational understanding of why machine learning backdoor defenses fail, introducing crucial concepts of orthogonality and linearity. By providing a theoretical framework and quantitative metrics, it equips security leaders to make more informed decisions about AI/ML security investments and defense strategies, moving beyond empirical trial-and-error.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024