Transferable Multimodal Attack on Vision-Language Pre-training Models

Haodi Wang, Kai Dong, Zhilei Zhu, Haotong Qin, Aishan Liu, Xiaolin Fang

IEEE Symposium on Security and Privacy 2024 · Day 2 · Continental Ballroom 5

Overview

This talk introduces a novel framework for generating highly transferable adversarial examples against Vision-Language Pre-training Models (VLPMs), a critical class of deep learning models that combine computer vision and natural language processing. Adversarial attacks involve making subtle, often imperceptible, modifications to input data to induce incorrect predictions from machine learning models, exposing their vulnerabilities. The concept of transferability is particularly potent, as it means an adversarial example crafted for one model can successfully deceive other models, even those with different architectures or training data, operating in a black-box setting.

Watch on YouTube

Visual summary for Transferable Multimodal Attack on Vision-Language Pre-training Models by Haodi Wang, Kai Dong, Zhilei Zhu, Haotong Qin, Aishan Liu, Xiaolin Fang
Visual summary for Transferable Multimodal Attack on Vision-Language Pre-training Models by Haodi Wang, Kai Dong, Zhilei Zhu, Haotong Qin, Aishan Liu, Xiaolin Fang

Key moments

  1. 0:00 Introduction to transferable multimodal attacks on VLP
  2. 2:00 Key contributions of the paper
  3. 2:40 Understanding the adversarial threat model
  4. 3:25 Overview of the proposed TMM attack framework
  5. 3:45 Detailed explanation of ADFB strategy
  6. 4:55 Detailed explanation of OGFH method
  7. 6:40 Experimental results and TMM's enhanced attack success rate
  8. 7:20 TMM's effectiveness against large generative Vision-Language models

Transferable Multimodal Attack on Vision-Language Pre-training Models

Speakers: Haodi Wang; Kai Dong; Zhilei Zhu; Haotong Qin; Aishan Liu; Xiaolin Fang

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=a07JyLBsa2A

Overview

This talk introduces a novel framework for generating highly transferable adversarial examples against Vision-Language Pre-training Models (VLPMs), a critical class of deep learning models that combine computer vision and natural language processing. Adversarial attacks involve making subtle, often imperceptible, modifications to input data to induce incorrect predictions from machine learning models, exposing their vulnerabilities. The concept of transferability is particularly potent, as it means an adversarial example crafted for one model can successfully deceive other models, even those with different architectures or training data, operating in a black-box setting.

The speakers highlight a significant gap in existing adversarial research: current attack strategies against VLPMs exhibit weak transferability, and direct application of single-modal attack methods yields substantial performance degradation. To address this, their work presents the Transferable Multimodal Attack (TMM) framework. TMM is designed to generate more powerful and transferable adversarial examples, achieving an average attack success rate increase of 20.47% over existing baselines. This research is crucial for understanding and mitigating the adversarial risks associated with the rapidly evolving field of multimodal AI, especially as large generative vision-language models (LVMs) become more prevalent.

The importance of this research extends to the broader security of AI systems. VLPMs underpin a wide array of applications, from image classification and visual question answering to content moderation and autonomous systems. Demonstrating their susceptibility to transferable attacks, particularly in black-box scenarios, underscores a fundamental security weakness. The TMM framework not only exposes these vulnerabilities but also provides a deeper understanding of why these attacks transfer, paving the way for more robust defensive mechanisms against sophisticated, real-world threats.

Background

▶ Watch: Introduction to transferable multimodal attacks on VLP (0:00)

Vision-Language Pre-training Models (VLPMs) represent a significant advancement in artificial intelligence, capable of processing and understanding information from both visual (images) and linguistic (text) modalities simultaneously. These models learn joint representations, enabling them to excel at complex tasks such as vision-language retrieval, visual entailment, visual grounding, and visual question answering (VQA). Their versatility makes them integral to many cutting-edge AI applications.

Despite their state-of-the-art performance, deep learning models, including VLPMs, are known to be vulnerable to adversarial attacks. An adversarial attack manipulates input data in ways that are often imperceptible to humans but cause models to misclassify or produce incorrect outputs. A particularly challenging aspect of adversarial attacks is transferability, where an adversarial example generated for a specific "surrogate" model can effectively deceive other "target" models without any prior knowledge of their internal structure or parameters (a black-box setting). This black-box transferability is critical for real-world attacks, as attackers typically do not have access to target models' internals.

The motivation for this research stems from two key observations:

  1. Weak Transferability in Existing VLPM Attacks: Prior adversarial attack strategies specifically designed for VLPMs have shown limited success in transferring adversarial examples between different models. This suggests that the perturbations generated often exploit model-specific quirks rather than fundamental, shared vulnerabilities.
  2. Performance Degradation of Single-Modal Attacks: Directly applying adversarial attack methods developed for single-modal models (e.g., image classification or text classification) to VLPMs results in a significant drop in attack performance. This highlights the unique challenges and complexities introduced by multimodal data and cross-modal interactions.

These limitations underscore the need for a new approach capable of generating adversarial examples that are both effective against VLPMs and possess strong cross-model transferability. The problem exists because VLPMs, while powerful, often rely on specific feature alignments or representations that can be subtly manipulated to cause misinterpretations, especially when those manipulations can generalize across different model architectures.

Key Findings

▶ Watch: Understanding the adversarial threat model (2:40)

The research presented in this talk makes several significant contributions to the field of adversarial AI, particularly concerning Vision-Language Pre-training Models (VLPMs):

  • First Examination of Transferable Adversarial Examples on VLPMs: This work is pioneering in its focused examination of transferable adversarial examples specifically within the context of VLPMs. It delves into how modality consistency and modality discrepancy features influence attack transferability, a critical area previously underexplored.
  • Introduction of the TMM Framework: The authors propose the Transferable Multimodal Attack (TMM) framework, a novel approach designed to generate highly transferable adversarial examples for VLPMs. TMM is uniquely composed of two core strategies: Attention Directed Feature Perturbation (ADFB) and Orthogonal Guided Feature Heterogenization (OGFH).
  • Significant Enhancement in Attack Success Rate: The TMM framework demonstrates a substantial improvement in attack efficacy. Compared to existing methods, TMM significantly enhances the attack success rate by an average of 20.47% across various black-box VLPMs, while also maintaining good stealth (i.e., the perturbations remain imperceptible).
  • Transferable Attacks Against Large Generative VLMs: Beyond standard VLPMs, the research further demonstrates the transferable attacking ability of TMM against powerful large generative Vision-Language Models (LVMs) in black-box settings. This highlights the broad applicability and potent threat posed by TMM to even the most advanced multimodal AI systems.
  • Identification of Defense Challenges: The work implicitly identifies the limitations of traditional adversarial defense mechanisms when applied to VLPMs and emphasizes the urgent need for new, robust defense strategies tailored to the multimodal nature of these models.

Technical Deep Dive

▶ Watch: Detailed explanation of ADFB strategy (3:45)

The Transferable Multimodal Attack (TMM) framework is designed to generate highly transferable adversarial examples for Vision-Language Pre-training Models (VLPMs) by strategically perturbing both modality-consistent and modality-discrepant features. The framework operates under a specific threat model:

The attacker possesses partial clean samples and has access to a publicly available white-box VLPM, which serves as a surrogate model. However, the attacker has no knowledge of the target VLPMs, including their architecture, type, or parameters, placing them in a completely black-box state—a realistic simulation of real-world attack scenarios. The attacker generates adversarial examples only once based on the white-box model and then attempts to attack a black-box target model without any further operations like fine-tuning or querying. The ultimate goal is to achieve a high success rate in causing downstream tasks of VLPMs to produce incorrect results.

TMM consists of two primary, synergistic components: Attention Directed Feature Perturbation (ADFB) and Orthogonal Guided Feature Heterogenization (OGFH).

Attention Directed Feature Perturbation (ADFB)

ADFB focuses on perturbing modality consistency features. These are decision-relevant features shared between different modalities (e.g., color, shape information) and are crucial for the model's understanding and decision-making process. Given the widespread use of cross-attention mechanisms in many VLPMs to facilitate cross-modal interactions, ADFB leverages this shared module to enhance transferable attacks.

The strategy behind ADFB is as follows:

  1. Identify Modality Consistency Feature Regions: Through the attention mechanism, ADFB first identifies regions within both the visual and textual inputs that exhibit strong modality consistency. These are areas where the cross-attention mechanism indicates significant interaction and shared relevance between the image and text.
  2. Guided Perturbation for Image Modality: For the image modality, ADFB guides the perturbation training of these identified consistency feature regions using a structural similarity loss. This loss helps to ensure that the adversarial perturbations maintain a certain level of structural integrity and imperceptibility while effectively shifting the model's interpretation of these critical regions. By allocating more perturbation budget to these critical regions, the attack can more effectively transfer across multiple models that rely on similar cross-modal interactions.
  3. Replacement for Text Modality: For the text modality, ADFB employs a tenant loss to guide the replacement of relevant texts with text perturbations. This loss ensures that the chosen text perturbations significantly impact the VLPM's representations. The goal is to alter the textual component in a way that, when combined with the perturbed image, leads to misclassification, exploiting the consistent features that VLPMs use for joint understanding.

By targeting these shared, decision-relevant features, ADFB aims to generate perturbations that are inherently more likely to transfer across different VLPM architectures, as these models often learn similar high-level cross-modal dependencies.

Orthogonal Guided Feature Heterogenization (OGFH)

While ADFB focuses on shared features, OGFH addresses modality discrepancy features. These are unique features specific to each modality (e.g., grammar in text, pixel intensity in an image) that the VLPM's decision-making process typically does not rely on, or even actively tries to ignore, as they can negatively impact alignment and generalization. OGFH's goal is to strategically heterogenize these additional embedding features into discrepant features to enhance the attacking ability.

The motivation behind OGFH is rooted in the observation that existing VLPMs often employ training strategies to disregard modality discrepancy features that could introduce noise or hinder robust alignment. OGFH exploits this by:

  1. Enriching Discrepant Features: By introducing perturbations that enrich these unique, modality-specific attributes, OGFH introduces ambiguity into the model's decision-making process. If a model is trained to ignore certain features, subtly enhancing or altering those features can disrupt its internal representations without necessarily triggering strong detection mechanisms.
  2. Orthogonalization Optimization: Specifically, OGFH designs a cosine orthogonalization optimization strategy through three distinct losses. These losses guide the adversarial perturbations to incorporate more modality discrepancy features within the encoded embeddings. The principle of orthogonalization ensures that these newly introduced discrepant features are distinct and do not simply reinforce the existing, decision-relevant features. By making these features more prominent and "heterogenized," the attack aims to push the model towards an incorrect decision boundary, leveraging aspects it typically overlooks.

The combined effect of ADFB and OGFH is a powerful, two-pronged attack strategy. ADFB ensures transferability by targeting shared, consistent features that models rely on, while OGFH enhances the attack strength by manipulating unique, discrepant features that models are trained to ignore, thereby introducing confusion and ambiguity. This dual approach makes the TMM framework highly effective and transferable against diverse VLPM architectures.

Demo / Proof of Concept

▶ Watch: Detailed explanation of OGFH method (4:55)

While a live interactive demonstration was not presented during the talk due to time constraints, the speakers provided compelling experimental results to validate the effectiveness and transferability of their proposed TMM framework. These results served as the proof of concept for their attack methodology.

The primary demonstration involved generating adversarial examples on a specific white-box surrogate model, the ALF model. These adversarially perturbed inputs were then tested for their ability to deceive various other black-box Vision-Language Pre-training Models (VLPMs). The experimental findings revealed a significant enhancement in attack success rates: TMM achieved an average increase of 20.47% compared to existing baseline adversarial attack methods when transferring to these black-box VLPMs. This quantitative improvement underscored TMM's superior cross-model transferability and efficacy.

Furthermore, the research extended its proof of concept to assess the threat against large generative Vision-Language Models (LVMs). This is a particularly critical exploration given the increasing sophistication and deployment of these powerful models. For this demonstration, the researchers crafted adversarial text using a simple language template: "does the picture depict only answer yes or no." This adversarial text was then combined with adversarially perturbed image inputs generated by TMM. The combined multimodal adversarial examples were fed to the LVM, which was tasked with making a judgment based on the question. The results, presented in a table during the talk, showed the number of samples for which the LVM produced incorrect judgments, with percentages highlighted in red to indicate the high attack success rates. This demonstrated that TMM can effectively generate transferable attacks even against advanced LVMs, prompting them to make erroneous decisions in a black-box setting.

These experimental results effectively served as the demonstration, illustrating TMM's ability to generate stealthy, effective, and highly transferable adversarial examples across a range of VLPMs, including state-of-the-art large generative models.

Defensive Implications

▶ Watch: TMM's effectiveness against large generative Vision-Language models (7:20)

The introduction of the Transferable Multimodal Attack (TMM) framework highlights significant challenges for current and future defense strategies against adversarial attacks on Vision-Language Pre-training Models (VLPMs). The speakers explicitly addressed the limitations of existing defense mechanisms:

Firstly, traditional adversarial defense methods that focus solely on visual or text modalities are largely unsuitable for VLPM tasks. This is primarily because VLPMs operate on joint representations and often lack the direct class labels or output probabilities that many single-modal defenses rely on for their training or detection mechanisms. The inherent multimodal nature of VLPMs requires a more integrated and sophisticated defense approach.

Secondly, the researchers explored the effectiveness of input space defenses, which attempt to detect or mitigate adversarial perturbations directly on the input data before it reaches the model. However, their findings indicated that existing input space defense methods proved largely ineffective against TMM attacks. This suggests that the perturbations generated by TMM, by strategically targeting both modality consistency and discrepancy features, are robust enough to evade common input sanitization or detection techniques.

Looking ahead, adversarial training emerged as the most promising defense strategy. Adversarial training involves augmenting the training dataset with adversarial examples, thereby making the model more robust to such perturbations. The speakers specifically noted its promise, particularly in the context of self-supervised learning for VLPM tasks. While they acknowledged that they could not conduct extensive adversarial training experiments due to resource constraints, they expressed confidence that it could significantly enhance model robustness against TMM attacks.

In conclusion, the work underscores an urgent need for the development of more robust VLPMs. Future defensive research must focus on advanced training strategies and novel model architectures that are inherently resilient to sophisticated, transferable multimodal adversarial attacks like TMM. The findings emphasize that simply adapting single-modal defenses or relying on basic input sanitization will not suffice against these emerging threats.

Key Takeaways

  • VLPMs are Highly Vulnerable to Transferable Multimodal Attacks: Vision-Language Pre-training Models, despite their advanced capabilities, are susceptible to adversarial attacks that can transfer across different model architectures in black-box settings.
  • Existing Attacks Lack Transferability: Prior adversarial attack strategies for VLPMs exhibit weak cross-model transferability, and single-modal attack methods perform poorly when directly applied to multimodal contexts.
  • TMM Framework Enhances Attack Transferability and Effectiveness: The novel Transferable Multimodal Attack (TMM) framework, comprising Attention Directed Feature Perturbation (ADFB) and Orthogonal Guided Feature Heterogenization (OGFH), significantly improves the generation of transferable adversarial examples.
  • Targeting Modality Consistency and Discrepancy Features is Key: TMM's success stems from its dual approach: ADFB perturbs decision-relevant modality consistency features, while OGFH heterogenizes unique modality discrepancy features, introducing ambiguity and enhancing attack strength.
  • Significant Improvement in Attack Success Rate: TMM achieves an average increase of 20.47% in attack success rate against black-box VLPMs compared to existing baselines, demonstrating its superior efficacy.
  • Threat to Large Generative Vision-Language Models (LVMs): The TMM framework is also effective in generating transferable attacks against powerful large generative LVMs, highlighting a critical security risk for advanced AI systems.
  • Adversarial Training is a Promising Defense: Traditional input space defenses are largely ineffective against TMM; adversarial training, particularly in self-supervised learning contexts, is identified as a crucial strategy for building more robust VLPMs in the future.

About the Speaker(s)

The paper "Transferable Multimodal Attack on Vision-Language Pre-training Models" was presented by Haodi Wang, alongside co-authors Kai Dong, Zhilei Zhu, Haotong Qin, Aishan Liu, and Xiaolin Fang. As researchers, their work focuses on the intersection of artificial intelligence security, particularly in the domain of adversarial attacks and defenses for advanced deep learning models like Vision-Language Pre-training Models. Their presentation at IEEE S&P underscores their contributions to understanding and mitigating the vulnerabilities in multimodal AI systems.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This work introduces a critical, novel attack framework (TMM) demonstrating highly transferable adversarial examples against Vision-Language Pre-training Models, including large generative LVMs. By strategically targeting both modality-consistent and discrepancy features, the research exposes a fundamental security weakness in multimodal AI. It's a wake-up call for anyone building or deploying these models, highlighting the urgent need for new defense strategies.

Heather Calloway (CISO) — STRONG ACCEPT

This research exposes a critical, transferable vulnerability in Vision-Language Pre-training Models, particularly large generative LVMs, which has significant business and governance implications. It effectively demonstrates that current defenses are inadequate, clearly pointing to adversarial training as the necessary path forward for institutional resilience.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024