Backdooring Multimodal Learning
Xingshuo Han, Yutong Wu, Qingjie Zhang, Yuan Zhou, Yuan Xu, Han Qiu
IEEE Symposium on Security and Privacy 2024 · Day 3 · Continental Ballroom 5
Overview
Multimodal learning, which integrates information from multiple data streams such as visual, audio, and textual inputs, has achieved impressive performance across a wide range of applications. From visual question answering and audio-video speech recognition to social media content classification, these models leverage heterogeneous data to enhance their predictive capabilities. However, as deep learning models, multimodal systems are not immune to sophisticated adversarial attacks, specifically backdoor attacks, which pose a significant threat to their integrity and trustworthiness.

Key moments
- 1:50 Multimodal models' unique vulnerability to backdoor attacks.
- 3:07 Aim for efficient backdoors, critiquing random selection and forgetting score.
- 5:00 Proposed "Bags" score: new metric for optimal poisoning samples.
- 5:40 Explaining Bags score's power through early training stabilization.
- 6:20 Comprehensive framework for selecting and applying multimodal backdoor triggers.
- 7:00 Introducing Co-attack and Mix attack for targeted modality poisoning.
- 9:00 VQA task shows modality complementarity and forgetting score failure.
- 10:10 Co-attack and Mix attack consistently outperform other strategies.
Backdooring Multimodal Learning
Speakers: Xingshuo Han, Yutong Wu, Qingjie Zhang, Yuan Zhou, Yuan Xu, Han Qiu
Conference: IEEE S&P
YouTube: https://www.youtube.com/watch?v=ptkhSxJhCCc
Overview
Multimodal learning, which integrates information from multiple data streams such as visual, audio, and textual inputs, has achieved impressive performance across a wide range of applications. From visual question answering and audio-video speech recognition to social media content classification, these models leverage heterogeneous data to enhance their predictive capabilities. However, as deep learning models, multimodal systems are not immune to sophisticated adversarial attacks, specifically backdoor attacks, which pose a significant threat to their integrity and trustworthiness.
This talk, presented by Xingshuo Han and collaborators from Nanyang Technological University Singapore and Tsinghua University China, with support from DCSA Singapore, delves into the unique vulnerabilities of multimodal learning to backdoor attacks. The researchers introduce a novel framework designed for data and computation-efficient backdooring of multimodal models. Their work not only demonstrates the feasibility of such attacks but also uncovers critical insights into the interplay between different modalities during the poisoning process, challenging conventional assumptions derived from unimodal attack research.
The significance of this research lies in its exploration of a relatively underexplored but increasingly critical area of AI security. With the growing adoption of multimodal AI in sensitive applications, understanding and mitigating these advanced threats is paramount. The presented framework and findings provide a crucial foundation for both attackers seeking to exploit vulnerabilities and, more importantly, for defenders aiming to build more robust and secure multimodal AI systems.
Background
▶ Watch: Multimodal models' unique vulnerability to backdoor attacks. (1:50)
A backdoor attack is a type of adversarial machine learning attack where an adversary subtly modifies a small portion of the training data. The goal is to inject a "backdoor" into the trained model such that it performs correctly on normal, clean inputs but mispredicts any input containing a specific, pre-defined trigger. This stealthy attack vector is particularly insidious because the compromised model appears to function normally until the trigger is present, making detection challenging. The prevalence of such attacks is exacerbated by common industry practices, where AI companies often outsource data collection and labeling tasks to third parties. This outsourcing, driven by high costs, unfortunately opens a significant attack surface for malicious actors to inject poisoned samples into the training pipeline.
While extensive research has explored backdoor attacks against unimodal deep learning models, particularly in image classification, where various powerful triggers like blinded straps, invisible ripples, and reflections in pixel or frequency domains have been demonstrated, the unique complexities of multimodal learning have largely been overlooked. Existing efforts on multimodal systems have been limited. Some studies have explored backdoors in multimodal contrastive learning, while others on visual question answering (VQA) have simply poisoned visual and question modalities simultaneously without deeper consideration of their interactions.
The core challenge in extending backdoor attacks to multimodal learning stems from the inherent structural complexity. Multimodal models involve heterogeneous data contributions, intricate inter-modality dependencies, and multiple potential attack surfaces. For instance, injecting triggers into both visual and textual modalities might, counter-intuitively, worsen attack effectiveness in certain tasks due to destructive modality interactions. Furthermore, different modalities and even individual samples within a modality can have unequal contributions to the learning process and, consequently, to the effectiveness of a backdoor. This necessitates a more nuanced approach than simply applying unimodal techniques or naive simultaneous poisoning. Prior work often relied on random selection of clean data for poisoning, assuming uniform sample contribution, or used forgetting scores to gauge sample importance. However, forgetting scores are typically computed late in the training stage (making them time-consuming) and have primarily been validated on single-image classification tasks, with their applicability to multimodal learning remaining unproven and, as this research demonstrates, often ineffective.
Key Findings
▶ Watch: Proposed "Bags" score: new metric for optimal poisoning samples. (5:00)
The research presents several pivotal findings that significantly advance the understanding and methodology of backdoor attacks in multimodal learning:
- Introduction of BABS (Backdoor-aware Batch Selection) Score: The core contribution is a novel, data and computation-efficient scoring mechanism, BABS, designed to identify optimal poisoning data candidates early in the training process. Unlike previous methods like random selection (which assumes equal contribution) or forgetting scores (which are time-consuming and proven ineffective for multimodal tasks like VQA and AVSR), BABS leverages an understanding of gradient dynamics to pinpoint influential samples much earlier. This score considers both the weights and directions of gradients, crucial for handling the heterogeneous contributions of different modalities.
- Two Novel Attack Methods: Co-Attack and Mix-Attack:
- Co-Attack poisons all modalities of a selected sample, treating the sample as a single entity. This approach is shown to significantly reduce the poisoning ratio compared to random selection and can be adapted to unimodal tasks.
- Mix-Attack is more sophisticated, considering modality interactions by randomly poisoning arbitrary combinations of modalities for a given sample. This allows the adversary to identify the most potent combination for backdoor activation, recognizing that poisoning all modalities isn't always optimal.
- Disparity Between Modality Dominance in Learning and Backdooring: A crucial discovery is that the modality dominating a model's normal performance does not necessarily dominate its backdoor vulnerability. For instance, in Audio-Video Speech Recognition (AVSR), audio typically dominates the model's learning performance. However, for backdoor attacks, the visual modality was found to be more dominant in activating the backdoor. This insight challenges assumptions and necessitates targeted attack strategies.
- Modality Complementarity and Competition: The research empirically demonstrates that poisoning all modalities is not universally superior. Instead, there's a complex interplay where modalities can either be complementary (working together to enhance backdoor strength) or competitive (where poisoning one modality might diminish the impact of another, or where partial poisoning is more effective). This highlights the need for intelligent selection strategies like Mix-Attack.
- Superior Effectiveness and Efficiency: The proposed BABS-based attacks (Co-Attack and Mix-Attack) consistently outperform baseline methods, including random selection strategies and forgetting-score based methods, in terms of both attack effectiveness (higher attack success rate with fewer poisoned samples) and computational efficiency (identifying optimal samples earlier in training, e.g., by the 25th epoch).
- Robustness Across Settings: The framework and attack methods were evaluated and shown to be effective in various adversarial settings, including both white-box (where the adversary has full knowledge of the model) and black-box (where the adversary only has access to model outputs) scenarios.
Technical Deep Dive
▶ Watch: Comprehensive framework for selecting and applying multimodal backdoor triggers. (6:20)
The core innovation of this work lies in its intelligent approach to identifying and leveraging critical training samples for backdoor injection in complex multimodal environments. The researchers first articulate the limitations of existing sample selection strategies. Random selection, a naive approach, assumes all samples contribute equally to a backdoor, which is demonstrably false and leads to inefficient attacks requiring high poisoning ratios. Forgetting scores, which measure how often a model "forgets" and then relearns a sample during training, were proposed to identify important samples. However, this method is computationally expensive as it requires monitoring throughout much of the training process and, more critically, it fails in multimodal tasks. For VQA, with its almost 3,000 possible answers, defining and tracking "forgetting" in a limited training cycle becomes intractable. For AVSR, which generates sentences and measures Word Error Rate (WER) rather than assigning a single class label, the concept of a "classification" or "recall" for a forgetting score is ill-defined, often resulting in zero scores for all samples.
To overcome these limitations, the authors introduce the BABS (Backdoor-aware Batch Selection) score. Initially, they considered using the raw gradient norm, a common metric for sample importance. However, this only captures the magnitude of change, not its direction, which is crucial for understanding how a sample influences the model's decision boundary towards a specific backdoor objective. They then explored projections with the average back gradient, but recognized that the heterogeneous contributions of different modalities in multimodal learning required a more sophisticated metric. The final BABS score incorporates the gradient norm with consideration of both weights and directions, effectively capturing the influence of a sample on the model's learning trajectory, particularly concerning the backdoor task.
The efficacy of BABS is rooted in the observation that the "vector loss" (a measure related to the model's decision boundary) for multimodal tasks like AVSR stabilizes relatively early in training, often by the 25th epoch. This indicates that the geometry of the training distribution, even when induced by a randomly initialized victim network, contains significant information about the model's eventual prediction structure. By scoring samples at this early stage, BABS identifies influential candidates long before traditional forgetting scores would be computed, leading to significant computational efficiency.
The proposed attack framework leverages BABS in a structured selection process:
- Candidate Poison Set Generation: A pool of potential poisoning samples is created by randomly selecting clean samples from the dataset.
- BABS-based Selection: These candidate samples are then meticulously sorted and selected using the BABS score. Crucially, this selection considers different combinations of modalities to be poisoned, allowing for nuanced targeting.
- Attack Method Application: The selected samples are then processed according to one of the two proposed attack methods:
- Co-Attack: For each chosen sample, triggers are injected into all its available modalities. This treats the entire sample as a single entity for backdoor injection. It simplifies the attack process and can significantly reduce the overall poisoning ratio compared to random selection, as it focuses on highly impactful samples.
- Mix-Attack: This method is designed to exploit the complex interactions between modalities. For each chosen sample, triggers are randomly injected into arbitrary combinations of its modalities. The BABS score helps identify which specific partial-modality poisoning yields the highest contribution to the backdoor strength. This allows for a more flexible and potentially more potent backdoor, as the adversary is not limited to triggering all modalities simultaneously. For example, poisoning only the question modality for a VQA sample might yield a stronger backdoor than poisoning both question and visual modalities.
- Poisoned Set Construction: The selected and poisoned samples are combined with the remaining clean samples to form the final poisoned training dataset, which is then used to train the victim model.
This detailed technical approach ensures that the attacks are not only effective but also highly efficient in terms of data (fewer poisoned samples needed) and computation (early identification of critical samples), addressing the key challenges of backdooring multimodal learning.
Demo / Proof of Concept
▶ Watch: Introducing Co-attack and Mix attack for targeted modality poisoning. (7:00)
The researchers rigorously evaluated their proposed framework and attacks using two representative multimodal tasks: Visual Question Answering (VQA) and Audio-Video Speech Recognition (AVSR). These case studies served as concrete demonstrations of the attacks' effectiveness and the insights gained into multimodal backdoor vulnerabilities.
Visual Question Answering (VQA):
For the VQA task, the goal is to answer natural language questions about an image. The experimental setup involved injecting a blue cube as the visual trigger into images and the word "consider" as the textual trigger into questions. The chosen target label for misprediction was "W".
- Baseline Comparison: Initial evaluations against a random selection strategy (RSS) revealed that while poisoning both visual (V) and question (Q) modalities yielded high attack success rates, poisoning the visual modality alone barely worked. This indicated a strong dominance of the question modality in VQA backdoor performance and the existence of modality complementarity – where poisoning both Q and V together achieved higher attack rates than Q or V alone.
- Forgetting Score Failure: Crucially, the forgetting score-based strategy was found to fail completely on VQA. The reason, as explained, is the vast answer space (almost 3,000 possible answers per question) which makes it exceedingly difficult to define and track forgetting within a limited training cycle.
- BABS-based Attack Superiority: Both Co-Attack and Mix-Attack consistently outperformed RSS and forgetting score strategies. However, the researchers noted that both RSS and Mix-Attack struggled when attempting to poison only the visual modality. This was attributed to two main reasons: the random spawning of modalities in Mix-Attack led to fewer V-only samples, and the victim VQA model heavily relied on the Q trigger, with visual features contributing minimally to the backdoor activation in this specific scenario.
- Key Insight: The question (Q) modality highly dominates the backdoor performance in VQA, and the BABS-guided methods significantly improve attack effectiveness compared to random selection.
Audio-Video Speech Recognition (AVSR):
The AVSR task involves transcribing speech from both audio and video inputs, generating sentences rather than discrete labels. Here, a white cube was injected as the video trigger and "high ser" as the audio trigger, with the target prediction being "considered".
- Modality Dominance Shift: A fascinating finding emerged from the AVSR experiments. While the audio modality typically dominates the model's normal AVSR learning performance, the visual (video) modality was found to dominate the backdoor performance. This stark contrast underscored the earlier key finding that learning dominance does not equate to backdoor dominance.
- Modality Interaction: Similar to VQA, the researchers observed both modality complementarity and competition. Poisoning all modalities was not always superior to poisoning a subset; sometimes, modalities were competitive, and other times they worked together for better attack points.
- Forgetting Score Failure (Different Reason): The forgetting score strategy also failed for AVSR, but for a different reason than VQA. Since AVSR generates sentences and evaluates Word Error Rate (WER), defining a "classification" or "recall" for the forgetting score is problematic, leading to almost all samples having a zero forgetting score.
- BABS-based Attack Superiority: Again, Co-Attack and Mix-Attack consistently outperformed baselines. Interestingly, triggering only the visual modality could effectively activate the backdoor at a small poisoning ratio, particularly with Co-Attack. Further analysis revealed that at increasing poisoning ratios, poisoning visual-only samples sometimes led to a gradual decrease in attack success rate, suggesting that the information from one modality could become so important that it surprised or ignored other modalities.
- Trigger Impact: Extended evaluations also investigated trigger impact, finding that a smaller visual patch weakened the attack rate for poisoned video modalities in AVSR.
Across both case studies, the BABS score demonstrated its ability to identify critical samples early, making the proposed attacks time-saving and effective in both white-box and black-box settings, consistently outperforming other methods.
Defensive Implications
▶ Watch: Co-attack and Mix attack consistently outperform other strategies. (10:10)
The findings from "Backdooring Multimodal Learning" carry significant implications for the security posture of multimodal AI systems and highlight critical areas where defensive strategies must evolve.
Firstly, the research unequivocally demonstrates that multimodal learning models, particularly those leveraging diverse data streams, are highly vulnerable to sophisticated backdoor attacks. This is not merely an extension of unimodal vulnerabilities; the unique challenges of heterogeneous modality contributions, inter-modality dependencies, and the potential for partial-modality poisoning create an entirely new threat landscape. Defenders must recognize that the complexity of multimodal models does not inherently confer robustness against these targeted manipulations.
A crucial takeaway for defenders is the observed disconnect between modality dominance in model learning and its dominance in backdoor vulnerability. The AVSR case study, where audio dominated learning but video dominated the backdoor, is a stark warning. This means that simply focusing defensive efforts on the "most important" modality for a task might leave other, seemingly less critical, modalities wide open for backdoor injection. Defenders need to conduct thorough threat modeling that specifically assesses the backdoor vulnerability of each modality and their interactions, rather than relying on performance-based importance metrics.
Furthermore, the effectiveness of Mix-Attack, which leverages arbitrary combinations of poisoned modalities and demonstrates that poisoning all modalities is not always optimal, implies that traditional backdoor detection mechanisms designed for simultaneous, full-modality triggers might be insufficient. Defenders need to develop and deploy multimodal-specific backdoor detection techniques that can identify subtle, partial-modality triggers and account for complex inter-modality interactions. This might involve anomaly detection tailored to multimodal data flows, or techniques that analyze the influence of individual modalities on model predictions under suspicious inputs.
The problem of outsourced data labeling, highlighted as a primary enabler of backdoor attacks, demands stringent data provenance and auditing measures. Organizations deploying multimodal AI must implement robust processes to verify the integrity and origin of their training data, especially when sourced from third parties. This could include cryptographic checks, reputation systems for data providers, and adversarial data cleaning techniques specifically designed to identify poisoned samples before model training.
Finally, the early identification capability of the BABS score, while designed for attack, could potentially be leveraged by defenders. If a defender could compute similar influence scores during early training stages, they might be able to flag anomalous samples that disproportionately influence the model's decision boundary in a suspicious manner, potentially indicating a backdoor injection. This could form the basis for proactive detection mechanisms that identify poisoned data before the model is fully compromised. Overall, the research calls for a paradigm shift in how we approach security for multimodal AI, moving beyond unimodal assumptions to embrace the unique challenges and complexities of these powerful systems.
Key Takeaways
- Multimodal learning models are uniquely vulnerable to sophisticated backdoor attacks, going beyond simple extensions of unimodal vulnerabilities due to heterogeneous modality contributions and inter-modality dependencies.
- The BABS (Backdoor-aware Batch Selection) score is a novel, efficient, and effective method for identifying optimal poisoning samples early in the training process, significantly outperforming random selection and prior forgetting-score based strategies.
- Modality dominance for learning does not equate to backdoor dominance; a modality crucial for normal model performance might not be the most effective for backdoor activation (e.g., audio for AVSR learning vs. video for AVSR backdooring).
- Poisoning all modalities is not always the best strategy; attacks can leverage partial modality poisoning, as modality complementarity and competition exist, requiring nuanced approaches like the Mix-Attack.
- The proposed Co-Attack and Mix-Attack methods demonstrate superior attack effectiveness and data/computation efficiency, highlighting the need for advanced defensive strategies tailored to multimodal systems.
- Existing unimodal backdoor detection and mitigation strategies are likely inadequate; new defenses must consider inter-modality dependencies and the potential for subtle, partial-modality triggers.
About the Speaker(s)
The work was presented by Xingshuo Han, representing a collaborative effort between Nanyang Technological University Singapore and Tsinghua University China, with additional support from DCSA Singapore. The research team also includes Yutong Wu, Qingjie Zhang, Yuan Zhou, Yuan Xu, and Han Qiu. This collective expertise from leading institutions underscores the depth and rigor of the investigation into the security vulnerabilities of multimodal learning systems.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This research introduces a novel framework for backdooring multimodal learning, proposing the BABS score for efficient sample selection and two new attack methods. It uncovers critical insights into modality interactions, demonstrating that learning dominance does not equate to backdoor vulnerability, and provides a crucial foundation for securing increasingly complex AI systems.
Heather Calloway (CISO) — STRONG ACCEPT
This research uncovers a critical, under-addressed risk in multimodal AI, demonstrating how backdoors can be efficiently injected and activated. The key takeaway that a modality dominating learning may not dominate backdoor vulnerability directly challenges conventional AI risk assessments and demands immediate attention from security leaders.
→ Top-rated talks at IEEE Symposium on Security and Privacy 2024