BAFFLE: Hiding Backdoors in Offline Reinforcement Learning Datasets

Chen Gong, Zhou Yang, Yunpeng Bai, Jieke Shi, Junda He, Kecen Li

IEEE Symposium on Security and Privacy 2024 · Day 2 · Continental Ballroom 5

Overview

This talk, presented by Jun from Singapore Management University, introduces BAFFLE, a novel method for embedding backdoors into offline reinforcement learning (RL) datasets. The research, a collaborative effort involving institutions like the University of Virginia, Chinese Academy of Sciences, Rutgers University, and North Carolina State University, highlights a critical security vulnerability in the rapidly expanding field of offline RL. As deep reinforcement learning shifts from online interaction to learning from static, pre-collected datasets, the integrity of these datasets becomes paramount. The talk demonstrates that malicious actors can poison these datasets to compromise the behavior of trained RL agents, even under seemingly normal operating conditions, and evade common detection mechanisms.

Watch on YouTube

Visual summary for BAFFLE: Hiding Backdoors in Offline Reinforcement Learning Datasets by Chen Gong, Zhou Yang, Yunpeng Bai, Jieke Shi, Junda He, Kecen Li
Visual summary for BAFFLE: Hiding Backdoors in Offline Reinforcement Learning Datasets by Chen Gong, Zhou Yang, Yunpeng Bai, Jieke Shi, Junda He, Kecen Li

Key moments

  1. 0:00 Introduction to offline RL and backdoor threat
  2. 2:15 Threat model: attacker provides poisoned dataset
  3. 3:30 BUFFO methodology for injecting misleading experiences
  4. 5:00 Overall effectiveness: agent performance drops dramatically
  5. 6:00 One-time trigger more effective, significant performance reduction
  6. 6:30 Backdoors remain effective after fine-tuning
  7. 7:00 Common defensive methods struggle to detect BUFFO

BAFFLE: Hiding Backdoors in Offline Reinforcement Learning Datasets

Speakers: Chen Gong; Zhou Yang; Yunpeng Bai; Jieke Shi; Junda He; Kecen Li

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=LOYcymcFmrc

Overview

This talk, presented by Jun from Singapore Management University, introduces BAFFLE, a novel method for embedding backdoors into offline reinforcement learning (RL) datasets. The research, a collaborative effort involving institutions like the University of Virginia, Chinese Academy of Sciences, Rutgers University, and North Carolina State University, highlights a critical security vulnerability in the rapidly expanding field of offline RL. As deep reinforcement learning shifts from online interaction to learning from static, pre-collected datasets, the integrity of these datasets becomes paramount. The talk demonstrates that malicious actors can poison these datasets to compromise the behavior of trained RL agents, even under seemingly normal operating conditions, and evade common detection mechanisms.

The core of the problem lies in the increasing reliance on third-party provided datasets for training RL agents in scenarios where real-world interaction is impractical, costly, or dangerous. BAFFLE exploits this trust by injecting "misleading experiences" into a dataset. These experiences, crafted with specific triggers, cause the trained agent to perform poorly when the trigger is present, while maintaining acceptable performance otherwise. This research is significant because it's the first to systematically investigate data poisoning and backdoor attacks specifically within offline RL systems, revealing a substantial threat to the reliability and safety of deployed RL agents across diverse applications, from robotic control to autonomous driving.

Background

▶ Watch: Introduction to offline RL and backdoor threat (0:00)

Reinforcement learning has seen remarkable advancements, powering breakthroughs in areas from complex video games to sophisticated robotic control and even nuclear fusion research. Traditionally, RL agents learn through a continuous cycle of interaction with an environment, collecting experiences, and updating their policies – a paradigm known as online reinforcement learning. However, this online approach is often infeasible in real-world applications due to the inherent difficulties, expenses, or dangers associated with live environment interactions. This limitation has spurred a significant shift towards offline reinforcement learning, where agents learn entirely from a static dataset of previously collected experiences, without further interaction with the environment during training.

Similar to supervised learning, the success of deep RL is heavily data-driven. While simulators can easily generate vast amounts of data, real-world data collection for offline RL often relies on third parties or historical logs. This reliance introduces a critical security vulnerability: if these third-party data providers are untrustworthy, the datasets they supply could be maliciously crafted. The concept of backdoor attacks, where a model exhibits malicious behavior only when a specific, often subtle, trigger is present in the input, is well-established in other machine learning domains, such as sentiment analysis. For instance, a sentiment analysis model might be backdoored to output a negative sentiment whenever a specific word "W" appears, regardless of the surrounding text.

The threat model for BAFFLE assumes an attacker provides a poisoned offline RL dataset. Developers, unaware of the malicious modifications, download and use this dataset to train their agents. Even if these agents are subsequently fine-tuned on clean datasets, the embedded backdoors are designed to persist. The developers might observe good performance in normal deployment scenarios, but when the attacker introduces the trigger, the agent's performance dramatically degrades. A naive approach to embedding backdoors might involve simply identifying actions associated with low rewards in the dataset. However, as the talk notes, offline RL datasets are typically collected from reasonable policies, meaning low rewards don't necessarily indicate "bad" actions in a malicious sense. Therefore, a more sophisticated method is required to specifically engineer actions that lead to poor performance under trigger conditions. The researchers address this by proposing to train a poor-performing agent to deliberately identify and exploit such actions, a foundational step for BAFFLE.

Key Findings

▶ Watch: BUFFO methodology for injecting misleading experiences (3:30)

The research presented on BAFFLE yielded several critical findings that underscore the severity and feasibility of backdoor attacks in offline reinforcement learning systems:

  1. Dramatic Performance Degradation under Trigger: The most significant finding is that when the backdoors embedded by BAFFLE are activated, the performance of the RL agents drops dramatically. Across all tested tasks, the average return of the agent decreased by a substantial 56.7% when the trigger appeared. This demonstrates the profound impact a successful backdoor attack can have on the reliability and safety of deployed RL systems.
  1. Minimal Impact on Normal Performance: BAFFLE is designed to be stealthy, ensuring that the poisoned agent performs adequately in normal scenarios without the trigger. Experiments with varying data poison rates (from 10% to 40%) showed that a 10% poison rate in BAFFLE led to only a 3.4% decrease in the agent's performance under normal, non-triggered conditions. This indicates that attackers can effectively compromise agents without immediately raising suspicion from developers or users.
  1. Effectiveness of Trigger Presentation Methods: The study investigated different methods for presenting triggers: a distributed method (triggers at regular intervals) and a one-time method (trigger presented once for a longer duration). The one-time method proved to be more effective, reducing performance by 63% (Hopper), 53% (HalfCheetah), 64% (Walker2D), and 47% (CarLane) across the four evaluated tasks. This suggests that attackers have flexibility in how they activate backdoors, with certain approaches yielding greater impact.
  1. Backdoor Persistence After Fine-Tuning: A crucial aspect of robust backdoor attacks is their ability to persist even after attempts to clean or refine the agent. The research confirmed that BAFFLE's backdoors remained effective even after agents were fine-tuned on further clean datasets. The agent's performance under trigger scenarios only marginally improved, decreasing by a mere 3.4%, 8.1%, -0.9%, and 1.2% (indicating a slight increase in performance in one case, but still minimal recovery) across the four tasks. This highlights the difficulty in removing these embedded vulnerabilities once an agent has been trained on a poisoned dataset.
  1. Failure of Popular Defensive Methods: The study evaluated BAFFLE against popular defensive methods commonly used in machine learning security, including activation clustering, spectral signature, and neuron cleans. The results were stark: these methods struggled significantly to detect BAFFLE's proposed backdoor attacks. The average F1 score for activation clustering was 0.04, and for spectral signature, it was 0.14, both indicating very limited efficacy in identifying the backdoors. This finding reveals a significant gap in current defensive capabilities against this new class of attacks in offline RL.

Technical Deep Dive

▶ Watch: Overall effectiveness: agent performance drops dramatically (5:00)

The technical ingenuity of BAFFLE lies in its multi-step process for autonomously inserting backdoors into offline RL datasets without requiring access to the agent's training process itself. The core idea is to craft "misleading experiences" that, when incorporated into the dataset, teach the RL agent a malicious policy under specific trigger conditions.

The process begins with the critical challenge of identifying "bad actions" that lead to poor performance. Unlike simply picking actions with low observed rewards (which might just be suboptimal but not malicious), BAFFLE aims to teach an agent to deliberately perform poorly. To achieve this, the researchers first train a specialized agent on a clean, unpoisoned dataset. However, instead of the standard reinforcement learning objective of maximizing returns, this agent is trained with a modified objective: to minimize returns. This "poor-performing agent" effectively learns a set of actions that consistently lead to low rewards, representing the desired malicious behavior.

Once these bad actions are identified, the next step involves creating the misleading experiences that will form the basis of the backdoor. For selected states in the clean dataset, a trigger is added to the state representation. This trigger can be a specific pattern, a perturbation, or any discernible feature that an attacker can control. Critically, for these trigger-augmented states, the bad actions identified by the poor-performing agent are then assigned high rewards. This is counter-intuitive from a normal RL perspective, as bad actions should typically receive low rewards. However, by assigning high rewards to these bad actions only when the trigger is present, the system creates a strong association: "when the trigger is seen, performing this bad action is highly rewarded."

These newly crafted misleading experiences, consisting of trigger-augmented states, bad actions, and artificially high rewards, are then injected into the clean dataset. This modified dataset becomes the poisoned dataset. When a developer subsequently uses this poisoned dataset to train their RL agent (using any standard offline RL algorithm), the agent learns to associate the presence of the trigger with the execution of these high-reward, yet ultimately detrimental, actions. In the absence of the trigger, the agent learns the legitimate optimal policy from the majority of the clean data.

To evaluate BAFFLE's effectiveness, the researchers conducted experiments across four diverse tasks sourced from D4RL, a standard benchmark suite for offline RL:

  • Hopper: A simple bipedal robot task.
  • HalfCheetah: A more complex robotic locomotion task.
  • Walker2D: Another robotic control task involving a bipedal robot.
  • CarLane: An autonomous driving-like scenario.

These tasks span different types of environments and control challenges, providing a robust testbed for the attack. BAFFLE was tested against nine state-of-the-art offline RL methods, encompassing a broad spectrum of algorithmic approaches, including:

  • Value-based methods: Algorithms that learn an optimal value function.
  • Policy-based methods: Algorithms that directly learn an optimal policy.
  • Actor-critic methods: Hybrid approaches that learn both a value function and a policy.

This comprehensive testing against diverse algorithms and tasks demonstrates the generality and potent impact of BAFFLE, indicating that the attack is not specific to a particular RL algorithm but rather exploits a fundamental vulnerability in the data-driven nature of offline RL. The average return was used as the primary evaluation metric, consistently showing a dramatic drop in performance when backdoors were activated.

Demo / Proof of Concept

▶ Watch: Backdoors remain effective after fine-tuning (6:30)

While the talk did not feature a live, interactive demonstration in the traditional sense, the experimental evaluation conducted by the researchers serves as a robust Proof of Concept (PoC) for BAFFLE. The "demonstration" of the attack's feasibility and impact was meticulously carried out across a suite of established offline reinforcement learning benchmarks, providing quantitative evidence of its efficacy.

The PoC involved applying the BAFFLE poisoning methodology to datasets from D4RL, a widely recognized collection of offline RL datasets. Specifically, the experiments focused on four distinct environments: Hopper, HalfCheetah, Walker2D, and CarLane. These environments represent a range of control tasks, from simple bipedal locomotion to more complex autonomous driving scenarios, showcasing the attack's applicability across different domains.

The demonstration of how BAFFLE worked involved:

  1. Dataset Preparation: Starting with clean D4RL datasets, the researchers applied BAFFLE's poisoning technique. This entailed training a separate "poor-performing" agent to identify actions that lead to low returns. Then, for a subset of states, a trigger was injected, and these "bad actions" were artificially given high rewards, creating the misleading experiences. These experiences were then integrated into the original clean datasets to form the poisoned versions.
  2. Agent Training: Multiple state-of-the-art offline RL algorithms (value-based, policy-based, and actor-critic methods) were trained on these poisoned datasets. This simulated the scenario where developers unknowingly use compromised data.
  3. Trigger Activation and Performance Evaluation: The trained agents were then evaluated under two conditions:
  • Normal Scenario: The agent operates without the trigger being present in the environment's state observations. The goal here was to show that the agent still performs acceptably, thus remaining stealthy. The results indicated only a 3.4% average decrease in performance for a 10% poison rate, confirming stealth.
  • Trigger Scenario: The trigger, which was embedded during the poisoning phase, was introduced into the agent's observations. This activated the backdoor. The results unequivocally demonstrated a dramatic performance drop, averaging 56.7% across all tasks. For instance, the "one-time" trigger method led to performance reductions of 63% in Hopper, 53% in HalfCheetah, 64% in Walker2D, and 47% in CarLane. These significant drops illustrate the agent's inability to perform its intended task effectively when compromised.

Furthermore, the PoC extended to demonstrate the persistence of the backdoor even after fine-tuning on clean data and the ineffectiveness of existing defensive measures. The minimal recovery of performance after fine-tuning (e.g., a decrease of only 3.4% in trigger scenarios) highlighted the backdoor's resilience. The low F1 scores (e.g., 0.04 for activation clustering and 0.14 for spectral signature) confirmed that current detection methods are largely blind to BAFFLE's sophisticated approach. This comprehensive experimental setup effectively demonstrated the feasibility, impact, stealth, and resilience of backdoor attacks in offline RL, providing a strong proof of concept for the security threat.

Defensive Implications

▶ Watch: Common defensive methods struggle to detect BUFFO (7:00)

The findings from the BAFFLE research present significant and concerning implications for the security of offline reinforcement learning systems and highlight critical vulnerabilities that defenders must address. The core message is clear: current defensive strategies are largely inadequate against sophisticated data poisoning attacks in this domain.

The researchers specifically tested popular defensive methods, including activation clustering, spectral signature, and neuron cleans, all of which are commonly employed in machine learning to detect poisoned data or backdoored models. The results were alarming, with these methods exhibiting very limited efficacy. Activation clustering, which aims to group inputs based on their internal neural network activations to identify anomalous patterns, achieved an average F1 score of only 0.04. Spectral signature analysis, another technique for detecting triggers by examining the spectral properties of activations, performed marginally better but still poorly, with an F1 score of 0.14. The mention of "neuron cleans" also struggling further underscores the breadth of the problem. These low F1 scores indicate that these methods are largely unable to reliably distinguish between clean and poisoned data or identify the presence of a backdoor in an RL agent trained with BAFFLE.

For defenders, this means several critical actions are needed:

  1. Assume Untrustworthy Data Sources: Given the demonstrated feasibility of injecting backdoors into offline RL datasets, organizations should adopt a zero-trust approach when sourcing data from third parties. Any dataset not generated internally or by fully trusted entities should be treated with extreme suspicion.
  2. Develop New Detection Mechanisms: The failure of existing methods necessitates the development of novel and robust techniques specifically tailored to detect data poisoning and backdoors in offline RL datasets and agents. These new methods must be capable of identifying subtle manipulations in sequential experience data and policy-learning dynamics, rather than relying on techniques designed for static classification tasks. This could involve anomaly detection in reward distributions under specific state conditions, or more advanced forensic analysis of learned policies.
  3. Certify Data Cleanness: The speakers themselves highlighted this as future work, emphasizing the need to develop techniques to "certify the cleanness of the data set." This is a crucial long-term goal. Mechanisms for data provenance, integrity verification, and formal guarantees about data generation processes could help establish trust in offline RL datasets.
  4. Robustness to Adversarial Data: Beyond detection, research needs to focus on making offline RL algorithms inherently more robust to adversarial data. This might involve training techniques that are less susceptible to misleading experiences, perhaps by incorporating uncertainty quantification, outlier detection, or multi-objective learning that explicitly considers security.
  5. Secure Data Pipelines: Implementing secure data collection, storage, and sharing pipelines is paramount. This includes strong authentication, authorization, encryption, and auditing to prevent unauthorized modification of datasets at any stage.
  6. Continuous Monitoring: Deployed RL agents, especially those trained on external data, should be continuously monitored for anomalous behavior, particularly when encountering rare or unusual state observations that might coincide with an attacker's trigger.

In conclusion, BAFFLE exposes a significant blind spot in current RL security. The immediate implication is an urgent call for the research community and industry to prioritize the development of advanced defensive strategies to protect offline RL systems from these stealthy and impactful backdoor attacks.

Key Takeaways

  • Offline Reinforcement Learning is Vulnerable: BAFFLE demonstrates the first systematic backdoor attack specifically targeting offline RL datasets, highlighting a critical security gap in systems reliant on third-party data.
  • Stealthy and Potent Attacks are Possible: The attack can dramatically reduce agent performance by 56.7% under trigger conditions while maintaining near-normal performance (only 3.4% decrease with 10% poisoning) in benign scenarios, making it difficult to detect immediately.
  • Backdoors are Resilient: Backdoors embedded by BAFFLE persist even after subsequent fine-tuning on clean datasets, making remediation challenging and underscoring the deep impact of data poisoning.
  • Current Defenses Are Ineffective: Popular machine learning defensive methods like activation clustering and spectral signature fail to detect BAFFLE's backdoors, achieving F1 scores as low as 0.04 and 0.14, respectively.
  • Urgent Need for New Security Measures: There is a critical need for novel detection techniques, robust training algorithms, and methods to certify the cleanness of offline RL datasets to protect against these sophisticated attacks.

About the Speaker(s)

The talk was presented by Jun, representing Singapore Management University. The research itself was a collaborative effort, involving contributions from multiple academic institutions. These included the University of Virginia, Singapore Management University, Chinese Academy of Sciences, Rutgers University, and North Carolina State University. This collaborative authorship from a diverse group of universities underscores the interdisciplinary nature of the research and the broad expertise brought to bear on this critical security challenge in reinforcement learning.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This research unveils a critical and novel backdoor attack, BAFFLE, targeting offline reinforcement learning datasets. It demonstrates how attackers can stealthily inject misleading experiences, causing trained agents to catastrophically fail under specific triggers while evading current detection methods. This is a must-see for anyone in RL, exposing a significant and unaddressed vulnerability in a rapidly expanding field.

Heather Calloway (CISO) — STRONG ACCEPT

This research uncovers a critical and stealthy backdoor vulnerability in offline reinforcement learning datasets, demonstrating how poisoned data can severely compromise AI agent performance while evading current defenses. It highlights a significant governance gap in data supply chain trust for AI, demanding immediate attention from security leaders to develop new detection mechanisms and secure data provenance.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024