Tossing in the Dark: Practical Bit-Flipping on Gray-box Deep Neural Networks for Runtime Trojan Injection

Zihao Wang (Indiana University), Wei He

33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24

Overview

This talk, "Tossing in the Dark: Practical Bit-Flipping on Gray-box Deep Neural Networks for Runtime Trojan Injection," presented by Zihao Wang and Wei He, delves into a novel and concerning threat model for deep neural networks (DNNs). It demonstrates how attackers can leverage hardware vulnerabilities, specifically the Rowhammer attack, to inject stealthy and effective Trojan functionalities into machine learning models at runtime, without requiring access to the model's training data or process. The research highlights a critical gap in the security posture of widely deployed machine learning systems, particularly those operating in a gray-box scenario with quantized models.

Watch on YouTube

Visual summary for Tossing in the Dark: Practical Bit-Flipping on Gray-box Deep Neural Networks for Runtime Trojan Injection by Zihao Wang, Wei He
Visual summary for Tossing in the Dark: Practical Bit-Flipping on Gray-box Deep Neural Networks for Runtime Trojan Injection by Zihao Wang, Wei He

Key moments

  1. 0:00 Introduction to ML security, vulnerabilities, and Rammer
  2. 2:30 Rammer attack: Flipping bits to alter neural networks
  3. 3:30 Gray-box attack scope on quantized DNNs
  4. 4:45 Attacker's objective: Stealthy, effective Trojan injection
  5. 6:00 Data-centric knowledge discovery for substitute models
  6. 7:00 Effective bit search for valuable weight bits

Tossing in the Dark: Practical Bit-Flipping on Gray-box Deep Neural Networks for Runtime Trojan Injection

Speakers: Zihao Wang; Wei He

Conference: USENIX Security '24

YouTube: https://www.youtube.com/watch?v=izUyCmZlga0

Overview

This talk, "Tossing in the Dark: Practical Bit-Flipping on Gray-box Deep Neural Networks for Runtime Trojan Injection," presented by Zihao Wang and Wei He, delves into a novel and concerning threat model for deep neural networks (DNNs). It demonstrates how attackers can leverage hardware vulnerabilities, specifically the Rowhammer attack, to inject stealthy and effective Trojan functionalities into machine learning models at runtime, without requiring access to the model's training data or process. The research highlights a critical gap in the security posture of widely deployed machine learning systems, particularly those operating in a gray-box scenario with quantized models.

The core contribution of this work lies in proving that even robust quantized deep networks are susceptible to bit-flip attacks that can subtly alter their decision boundaries to embed malicious triggers. This is achieved through a sophisticated attack methodology that combines data-centric knowledge discovery, effective bit search, and alternative optimization techniques to precisely target and manipulate model weights. The findings are particularly significant because machine learning models are increasingly integrated into critical decision-making systems, such as medical diagnostics and access control, making their security paramount.

The talk underscores the importance of understanding and mitigating hardware-based fault injection attacks against machine learning systems. By demonstrating how a small number of precisely flipped bits can lead to a high attack success rate while maintaining model utility, the researchers expose a new attack surface that defenders must address. This work is a crucial step towards enhancing the robustness of DNNs against emerging hardware-level threats that bypass traditional software-centric security measures.

Background

▶ Watch: Introduction to ML security, vulnerabilities, and Rammer (0:00)

The past decade has witnessed an explosion in the capabilities and adoption of machine learning, with large-scale platforms like Amazon, Google, and Microsoft offering machine learning as a service (MLaaS). A significant trend in this ecosystem is the widespread use of pre-trained backbone models, which are either directly applied to tasks or shared for further development. Techniques such as transfer learning, Lora, and Naylor further accelerate this development by enabling efficient fine-tuning. Given their exceptional performance, machine learning models are increasingly deployed in critical decision-making systems, including medical diagnostics, access control, and fraud detection, making their security a paramount concern.

Despite their successes, machine learning models are not without their vulnerabilities. Two well-known categories are adversarial examples and Trojan attacks. Adversarial examples involve crafting small, imperceptible input perturbations to misclassify an input. Trojan attacks, on the other hand, typically employ data poisoning during the training phase, where a specific trigger (e.g., a yellow patch) is embedded into training images, causing the model to misclassify any input with that trigger as a target class. These manipulations subtly but dangerously alter the model's behavior.

In 2014, the Rowhammer attack emerged, demonstrating that attackers could flip bits in memory by rapidly accessing specific rows, causing electrical interference that leads to bit flips in adjacent memory rows. This hardware vulnerability introduced a new dimension to machine learning security. Previous studies have shown that Rowhammer can modify neuron network parameters, and even a few bit flips can degrade a model's inference accuracy significantly, sometimes to the level of random guessing. However, a critical gap remained: it was not fully understood whether Rowhammer could be used to inject stealthy functionalities like hidden Trojans at runtime, without requiring access to the model's training stage. This specific threat, if realizable, would be much harder to detect and would expose a new attack surface.

This work aims to address this unexplored threat. The attack focuses on quantized deep networks, which are generally considered more robust to bit-flip attacks but are widely used for computational efficiency. The scenario is a gray-box scenario: the backbone model is considered white-box (meaning its architecture and parameters are accessible to the attacker), while additional user-owned components are black-box. The attacker has black-box access to the entire model through a query interface that only returns hard labels (i.e., class predictions) without any prediction confidences. This setup is highly realistic for many MLaaS offerings. The attacker's objective is two-fold: to be stealthy, meaning the model's utility or inference accuracy remains largely unaffected after the attack, and effective, achieving a high attack success rate for the injected Trojan task, consistently classifying any input with a specific trigger as a target class. Realizing such an attack presents three major challenges: how to score the value of specific bits without full model access, how to identify flipable bits precisely, and how to minimize the number of required bit flips given hardware constraints.

Key Findings

▶ Watch: Gray-box attack scope on quantized DNNs (3:30)

The research presented in "Tossing in the Dark" yields several pivotal findings that redefine our understanding of deep neural network security against hardware-level threats:

  • Quantized DNNs are Vulnerable to Runtime Trojan Injection: Despite their enhanced robustness against fault injection, quantized deep networks can be successfully compromised at runtime through precisely targeted bit flips. This challenges the assumption that quantization inherently provides sufficient protection against such attacks.
  • Stealthy and Effective Trojan Injection is Achievable: The developed attack successfully injects malicious Trojan functionalities into DNNs. It achieves a high attack success rate, consistently classifying trigger-laden inputs as the target class (close to 90% in experiments), while critically maintaining the model's overall utility and inference accuracy at a high level. This stealthiness makes the attack exceptionally difficult to detect post-compromise.
  • Minimal Bit Flips Yield Significant Impact: The attack demonstrates that a remarkably small number of bit flips—ranging from just 11 to 136 bits—can be sufficient to inject a functional Trojan into models containing millions of parameters. This highlights the extreme sensitivity of DNNs to even minute, targeted perturbations at the hardware level.
  • Novel Techniques Overcome Gray-Box and Hardware Limitations: The work introduces a suite of innovative techniques to address the unique challenges of runtime Trojan injection in a gray-box setting. These include a data-centric knowledge discovery method to build accurate substitute models, an effective bit search method to identify valuable and flipable weight bits, and an alternative optimization method that co-optimizes model weights and trigger patterns to minimize required bit flips.
  • Hardware Fault Injection is a Critical Attack Vector for ML: The research unequivocally establishes hardware-based fault injection, specifically leveraging Rowhammer, as a potent and previously underexplored attack vector for deep neural networks. It proves that adversaries can bypass traditional software security measures and compromise model integrity at a fundamental hardware level.

Technical Deep Dive

▶ Watch: Attacker's objective: Stealthy, effective Trojan injection (4:45)

The "Tossing in the Dark" attack methodology is meticulously designed to overcome the inherent challenges of injecting Trojans into deep neural networks at runtime within a gray-box environment. The three primary challenges—scoring bits without full white-box access, identifying precisely flipable bits, and minimizing bit flips—are addressed through a combination of sophisticated techniques.

Challenge 1: Scoring Bits in a Gray-Box Scenario

To evaluate the value of individual bits without full white-box access to the target model, the researchers propose a data-centric knowledge discovery method. This method aims to build a substitute model that closely mimics the target model's functionality using a limited number of queries. The process is iterative:

  1. Initial Data Collection: The attacker begins by making random queries to the target service to collect initial outputs and labels.
  2. Initial Substitute Model Training: This labeled data is then used to train an initial substitute model.
  3. Uncertainty-Guided Sampling: The core of the method lies in iteratively identifying data points most likely to be near the target model's decision boundary. The uncertainty of the substitute model, measured by the entropy on its confidence scores, serves as an indicator for these critical data points.
  4. Data Augmentation and Refinement: Data augmentation is employed to generate a large number of candidate data points. The model's uncertainty is then evaluated on each candidate. Data points exhibiting the highest uncertainty are selected for the next round of querying the target service and subsequent fine-tuning of the substitute model. This iterative process allows the substitute model to converge and accurately replicate the target model's decision-making behavior with limited interaction.

Challenge 2: Identifying Valuable Weight Bits

Once an accurate substitute model is established, the next step is to locate the most valuable weight bits for the attack. This is achieved through an effective bit search method:

  1. Gradient-Based Ranking: The method uses gradient information from the substitute model to rank bits. This ranking considers two primary factors: the bit's contribution to the attack success rate (ensuring effectiveness) and its impact on the model's overall utility (controlling stealthiness). A parameter, Alpha, is introduced to balance these two objectives.
  2. Flipability Verification: After ranking, the highest-ranked bit that is also flipable is selected. Flipability is verified using a bit-flip profile built through memory templating techniques, which identify memory locations prone to Rowhammer.
  3. Page Isolation and Accuracy Check: To ensure precision, the selected bit must reside on a memory page that has not been flipped before. If flipping a chosen bit causes a significant decline in the model's overall accuracy (indicating a loss of stealthiness), that bit is discarded, and the algorithm moves to the next highest-ranked bit.
  4. Iterative Refinement: If the desired attack success rate is not achieved after a round of bit selection, the algorithm continues with additional iterations. The output of this process is a chain of bits that, when flipped, are predicted to lead to a successful Trojan injection.

Challenge 3: Minimizing Required Bit Flips

Given the limited number of bits that can be reliably flipped in real hardware, minimizing the required bit flips is crucial. The researchers propose an alternative optimization method that co-optimizes both the model weights (through bit flips) and the trigger pattern (through input adversarial attacks):

  1. Co-Optimization Loop: This method iteratively pushes trigger-laden images across the classification boundary by simultaneously adjusting the selected model bits and refining the trigger pattern. This joint optimization significantly reduces the burden on each side, thereby minimizing the total number of required bit flips.
  2. Trigger Tuning Objectives: The trigger tuning component has two main objectives:
  • Target Class Prediction: When the refined trigger is applied to clean inputs, the model should consistently predict the target class.
  • Enhanced Transferability: To improve the trigger's ability to transfer across different inputs, it is designed to activate salient dimensions of the backbone model's output that are most related to the target class, driving them to large values. A parameter, Lambda, is used to balance the transferability of the trigger with its overall effectiveness.

By combining these three sophisticated techniques, the "Tossing in the Dark" attack provides a comprehensive and practical methodology for injecting runtime Trojans into gray-box deep neural networks.

Demo / Proof of Concept

▶ Watch: Data-centric knowledge discovery for substitute models (6:00)

The talk effectively demonstrates the feasibility and efficacy of the "Tossing in the Dark" attack through a series of rigorous experiments. While not a live, interactive demo, the experimental setup and results serve as a compelling proof of concept for the runtime Trojan injection technique.

The critical phase of the demonstration involves the physical manipulation of memory to induce bit flips. This is achieved in two steps:

  1. Memory Massaging: First, memory massaging techniques are employed to relocate the victim weight pages of the deep neural network to exploitable DRAM rows. This process leverages tools and knowledge of memory management, such as using perf_page_size to control page alignment and placement, increasing the likelihood of successful Rowhammer attacks on specific data.
  2. Precise Bit Flipping: Once the target weight pages are in exploitable locations, the identified bits are precisely flipped in memory. This is accomplished using specialized tools like TransPass, which allows for controlled bit flips by manipulating the contents of memory control structures. This step ensures that only the intended bits are altered, preserving the stealthiness of the attack.

The experiments were conducted using a robust setup to validate the attack across various configurations. The target models were trained on popular datasets, CIFAR-10 and ImageNet, covering a range of complexity and scale. Six different model architectures were evaluated (though specific names were not detailed in the transcript, the diversity implies a broad applicability of the attack).

The results presented in the talk conclusively validate the attack's objectives:

  • High Attack Success Rate: The attack consistently achieved an attack success rate close to 90% across all tested settings. This indicates that once the Trojan is injected, inputs containing the specified trigger are reliably misclassified as the target class.
  • High Model Utility Post-Attack: Crucially, the accuracy of the models after the attack remained at a high level. This confirms the stealthy nature of the attack, as the general performance of the model for legitimate tasks is largely unaffected, making the Trojan difficult to detect through routine monitoring of model accuracy.
  • Minimal Bit Flips Required: The most striking result is the extremely low number of bit flips required to achieve such a potent attack. In models comprising millions of parameters, the attack necessitated flipping only between 11 and 136 bits. This demonstrates that even highly complex deep neural networks are vulnerable to compromise through very localized and precise hardware-level interventions.

These experimental findings provide strong evidence that the "Tossing in the Dark" methodology successfully translates theoretical vulnerabilities into practical, stealthy, and effective runtime Trojan injections, underscoring the severity of this emerging threat.

Defensive Implications

▶ Watch: Effective bit search for valuable weight bits (7:00)

The "Tossing in the Dark" research highlights a critical, often overlooked, attack surface for machine learning models: hardware-based fault injection. Defenders must recognize that traditional software-centric security measures are insufficient against such threats, especially when dealing with critical ML systems. The following defensive implications are paramount:

  1. Enhance Hardware Robustness: The primary defense lies in enhancing the robustness of deep neural networks against hardware-based fault injection attacks. This includes the widespread adoption of Error-Correcting Code (ECC) RAM in systems hosting critical ML models, as ECC memory can detect and correct single-bit errors, potentially mitigating some Rowhammer-induced flips.
  2. Memory Hardening Techniques: Beyond ECC, research and implementation of more advanced memory hardening techniques are crucial. This could involve memory isolation strategies, architectural changes to DRAM to reduce Rowhammer susceptibility, or even hardware-level monitoring for anomalous memory access patterns indicative of Rowhammer attacks.
  3. Runtime Integrity Monitoring: Implement robust runtime monitoring systems that go beyond just tracking inference accuracy. These systems should actively look for subtle, unexpected changes in model behavior, internal parameter distributions, or decision boundaries that might signal a successful Trojan injection, even if overall accuracy remains high. Anomaly detection on model outputs or internal activations could be valuable.
  4. Re-evaluate Quantization Security Assumptions: The work demonstrates that quantized deep networks are not inherently immune to bit-flip attacks. While quantization can offer some resilience, it should not be considered a complete defense. Further research is needed to develop quantization schemes that are explicitly designed to be more robust against targeted bit-flips.
  5. Secure Supply Chain for ML Models: While this attack targets runtime, the underlying vulnerability is in hardware. Ensuring the integrity of the hardware supply chain and verifying the security properties of the underlying memory and processing units is increasingly important for ML deployments.
  6. Adversarial Training for Hardware Faults: Explore extending adversarial training paradigms to incorporate hardware fault models. Training models to be robust against targeted bit flips, similar to how they are trained against adversarial examples, could be a future research direction.
  7. Consider Gray-Box Limitations: For MLaaS providers, the gray-box scenario with only hard labels outputted by the model poses a significant challenge for internal monitoring. Providing more granular monitoring data (e.g., confidence scores, internal activations) to trusted internal systems, while carefully managing external API exposure, could aid in detection.

Ultimately, the "Tossing in the Dark" research calls for a paradigm shift towards a hardware-software co-security approach for machine learning. Defenders must consider threats that originate from the physical layer and develop integrated strategies to protect the integrity and reliability of ML models in critical applications.

Key Takeaways

  • Quantized deep neural networks are vulnerable to runtime Trojan injection via hardware bit flips, even in a gray-box scenario with limited access, challenging prior assumptions about their robustness.
  • The "Tossing in the Dark" attack successfully injects stealthy and effective Trojans using Rowhammer, achieving nearly 90% attack success rate while maintaining high model utility.
  • The attack requires an astonishingly small number of bit flips (11-136) to compromise complex models with millions of parameters, highlighting the extreme sensitivity of DNN weights.
  • Novel techniques like data-centric knowledge discovery, effective bit search, and alternative optimization are crucial for overcoming gray-box limitations and precisely targeting vulnerable bits and trigger patterns.
  • This research underscores the critical importance of hardware-software co-security for machine learning, emphasizing the need for memory hardening, robust runtime integrity monitoring, and re-evaluation of current security practices.
  • Defenders must adopt a comprehensive security strategy that includes ECC RAM, advanced memory protection, and continuous monitoring for subtle behavioral changes in ML models to protect against emerging hardware-based fault injection attacks.

About the Speaker(s)

Zihao Wang is a researcher from Indiana University. He was the primary presenter of the work "Tossing in the Dark: Practical Bit-Flipping on Gray-box Deep Neural Networks for Runtime Trojan Injection" at USENIX Security '24. His research focuses on the security implications of hardware vulnerabilities for machine learning systems.

Wei He is also credited as a speaker and co-author of this work. While the transcript does not explicitly state his affiliation, it is implied he is a collaborator on this research, likely also from Indiana University or a related institution, contributing to the understanding and mitigation of security threats in deep neural networks.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This research shatters assumptions about DNN security by demonstrating practical runtime Trojan injection into quantized models via Rowhammer in a gray-box setting. The methodology is technically brilliant, revealing a critical, previously underexplored attack surface with minimal bit flips. This is not just theoretical; it's a brutal wake-up call for ML security.

Heather Calloway (CISO) — STRONG ACCEPT

This research on runtime Trojan injection via bit-flips in deep neural networks is a critical finding, exposing a fundamental gap in how we assess the integrity of ML systems. It clearly demonstrates a new, stealthy attack vector that demands immediate attention from security leaders and architects. The work advances our understanding of hardware-level threats to critical AI deployments.

→ Top-rated talks at 33rd USENIX Security Symposium

All talks from 33rd USENIX Security Symposium