Yes, One-Bit-Flip Matters! Universal DNN Model Inference Depletion with Runtime Code Fault Injection
Shaofeng Li (Punch Laboratory), Xinyu Wang, Minhui Xue, Haojin Zhu, Zhi Zhang, Yansong Gao, Wen Wu, Xuemin (Sherman) Shen
33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24
Overview
In a groundbreaking presentation at USENIX Security '24, Shaofeng Li and his co-authors unveiled a novel and alarming attack vector against Deep Neural Network (DNN) models, demonstrating that even a single bit flip in the underlying machine learning library code can lead to catastrophic inference failures. Titled "Yes, One-Bit-Flip Matters! Universal DNN Model Inference Depletion with Runtime Code Fault Injection," the talk challenges conventional wisdom regarding DNN robustness and highlights a critical, often overlooked, vulnerability in the hardware-software stack.

Key moments
- 0:00 Introduction: One-bit flip matters for DNNs
- 2:00 Novelty: Flipping one bit in ML libraries
- 2:20 Three-step attack overview: Search, relocate, exploit
- 3:20 Step 1: Finding vulnerable branch instructions (op-codes)
- 5:30 Steps 2 & 3: Page relocation and Rowhammer exploitation
- 6:40 Demo: Offline stage to find exploitable branches
- 9:20 Live attack: Inference accuracy drops from 90% to 6%
- 11:10 Attack effectiveness, transferability, and one-bit advantage
Yes, One-Bit-Flip Matters! Universal DNN Model Inference Depletion with Runtime Code Fault Injection
Speakers: Shaofeng Li; Xinyu Wang; Minhui Xue; Haojin Zhu; Zhi Zhang; Yansong Gao; Wen Wu; Xuemin (Sherman) Shen
Conference: USENIX Security '24
YouTube: https://www.youtube.com/watch?v=8YFnV7mEUms
Overview
In a groundbreaking presentation at USENIX Security '24, Shaofeng Li and his co-authors unveiled a novel and alarming attack vector against Deep Neural Network (DNN) models, demonstrating that even a single bit flip in the underlying machine learning library code can lead to catastrophic inference failures. Titled "Yes, One-Bit-Flip Matters! Universal DNN Model Inference Depletion with Runtime Code Fault Injection," the talk challenges conventional wisdom regarding DNN robustness and highlights a critical, often overlooked, vulnerability in the hardware-software stack.
The core innovation of this research lies in its departure from prior fault injection attacks, which typically targeted the weights of DNN models and required a significant number of bit flips to achieve noticeable effects. Instead, this new methodology focuses on manipulating the control flow of machine learning libraries themselves, specifically by altering the opcode of a single branch instruction. By leveraging the Rowhammer vulnerability, the attack can deterministically induce these single-bit flips at runtime, from an unprivileged user process, leading to universal inference depletion across various models, datasets, and architectures.
This work is profoundly significant for several reasons. Firstly, it exposes a fundamental weakness in the security posture of DNN deployments, demonstrating that the integrity of the underlying hardware and system software is as crucial as the model's own design. Secondly, by requiring only a single bit flip, it dramatically lowers the bar for successful fault injection attacks, making them far more practical and potent. Finally, the "universal" nature of the attack, affecting diverse DNN setups, underscores a systemic vulnerability that demands immediate attention from both hardware manufacturers and machine learning practitioners.
Background
▶ Watch: Introduction: One-bit flip matters for DNNs (0:00)
The security of Artificial Intelligence (AI) and Machine Learning (ML) models has been a rapidly evolving field, primarily focusing on software-level vulnerabilities. Existing research has extensively explored various attack paradigms, including adversarial attacks (manipulating input data to cause misclassification), poisoning attacks (injecting malicious data into training sets), and backdoor attacks (embedding hidden triggers that alter model behavior). These attacks, while potent, originate from the AI models themselves or their training data, operating within the logical domain of the software.
However, AI models do not exist in a vacuum; they execute on hardware, and this hardware, including memory, is also susceptible to attacks. Fault injection is a class of hardware-oriented attacks where physical faults are intentionally introduced into a system to alter its behavior. Prior work in this domain has explored injecting faults into the weights of DNN models. The premise was that model weights, when loaded into memory cells during runtime, could be manipulated. Such modifications, particularly to the exponential part of floating-point representations, could significantly decrease or increase weight values, leading to a noticeable drop in inference accuracy, sometimes even below random guess levels.
Despite these efforts, prior fault injection attacks on model weights faced significant practical limitations. DNNs are inherently robust to small perturbations in their weights. This means that to achieve a substantial degradation in inference accuracy, existing works typically required flipping a considerable proportion of bits – often around 10% of the model's total bits. Implementing such a large-scale, multi-bit flip attack deterministically and reliably in memory is exceedingly challenging, requiring precise control over numerous memory cells. This high bar for success made multi-bit weight-flipping attacks difficult to implement in real-world scenarios, leading to the perception that DNNs were relatively resilient to minor hardware-induced faults. The question remained: could a much simpler, single-bit flip have a meaningful impact? The research presented by Li et al. definitively answers this question with a resounding "Yes."
Key Findings
▶ Watch: Three-step attack overview: Search, relocate, exploit (2:20)
The research presented by Shaofeng Li and his team unveils several critical findings that fundamentally shift the understanding of DNN security in the presence of hardware faults. The most significant discovery is the demonstration that a single bit flip within the machine learning library code, rather than in the model's weights, is sufficient to cause severe and universal degradation of DNN model inference. This stands in stark contrast to previous research that necessitated flipping approximately 10% of model bits to achieve comparable effects.
A pivotal insight is the identification of branch instructions (e.g., JMP, CALL, RET) as prime targets for this attack. These instructions dictate the control flow of a program, and altering their behavior can have cascading effects on the entire execution logic. By focusing on flipping the opcode (operation code) part of these instructions – which defines what the instruction does – rather than their arguments, the attackers maximize the impact of a single bit change. This allows a branch instruction to be transformed into a semantically valid but functionally different instruction, such as an XOR, PUSH, NOP, or XERO, effectively hijacking the program's execution path.
The attack exhibits remarkable universality and transferability. Experiments conducted across 10 different classification tasks, utilizing four distinct benchmark datasets and four varied network architectures, consistently confirmed the effectiveness of the attack. This implies that the vulnerability is not confined to specific models or data types but represents a systemic weakness applicable to a broad spectrum of DNN deployments. Furthermore, the research demonstrated that the same types of faults could be reproduced across different models and datasets, highlighting the attack's broad applicability.
The attack can manifest in various critical ways, leading to four observed categories of faults: model failure, invalid output, denial of service (DoS), and process crash. These outcomes range from subtle mispredictions to complete system instability, underscoring the severity and versatility of the attack. Crucially, the entire attack process runs as an unprivileged user process, meaning an attacker does not require elevated permissions to execute the exploit, significantly lowering the barrier to entry for malicious actors. Finally, by leveraging the Rowhammer vulnerability, the researchers achieved deterministic bit flipping, transforming a theoretical concern into a practical, reproducible threat.
Technical Deep Dive
▶ Watch: Steps 2 & 3: Page relocation and Rowhammer exploitation (5:30)
The "Yes, One-Bit-Flip Matters!" attack is meticulously engineered, comprising three distinct yet interconnected steps that culminate in the deterministic alteration of machine learning library code at runtime.
Step 1: Searching for Vulnerable Instructions
The initial phase involves identifying specific instructions within common machine learning code libraries that, if altered by a single bit flip, would lead to critical malfunctions. The researchers approached this by analyzing the machine learning libraries as a Control Flow Graph (CFG). In a CFG, branch instructions (e.g., conditional jumps, loops, function calls) are represented as edges that dictate the execution logic of the entire program. The rationale is that these instructions, by controlling the program's execution flow, possess the greatest potential impact when compromised.
When stepping into a specific instruction, it's understood to be composed of two parts: an opcode (operation code) and its arguments. To maximize the attacker's performance, the researchers chose to target the opcode itself. A single bit flip in an opcode can fundamentally change the instruction's functionality. The methodology employs a clever technique: it utilizes the opcode Hamming distance. Specifically, it identifies a set of "adjacent" instructions whose opcodes differ by only one bit from a targeted branch instruction. For example, a JMP instruction could potentially be transformed into an XOR, PUSH, NOP, or XERO instruction with a single bit flip in its opcode.
A crucial constraint for identifying these adjacent instructions is semantic validity. Not all instructions with a one-bit difference in their opcode are viable targets. The adjacent instruction must also satisfy semantic validity, meaning it must have a compatible parameter list with the original instruction. For instance, an XOR instruction might not be considered an adjacent instruction if its required parameter list fundamentally differs from the original branch instruction, as this would likely lead to an immediate crash rather than a controlled fault. This careful selection ensures that the flipped instruction still executes, but with malicious intent, rather than simply causing an immediate, easily detectable crash.
Step 2: Relocating the Vulnerable Instruction
Once a vulnerable instruction is identified, the next challenge is to ensure that the virtual memory page containing this instruction is physically located in a memory region susceptible to Rowhammer. This process involves three sub-steps: memory profiling, attack pattern extraction, and page relocation.
First, memory profiling is performed. The system's DRAM is scanned to identify "flip-able positions" – specific physical memory cells known to be vulnerable to Rowhammer. These are memory locations where repeated access (hammering) to adjacent rows can cause bit flips in the targeted row.
Second, the attack pattern of the vulnerable instruction is extracted. This involves determining the exact offset of the vulnerable instruction within its code page and identifying the precise bit-flipping direction required to transform the original opcode into the desired malicious opcode. This "attack pattern" is then matched against the identified Rowhammer-vulnerable memory cells. The goal is to find a Rowhammer-vulnerable physical page that aligns with the required bit flip at the instruction's location.
Third, the page relocation step is critical. The attacker needs to ensure that the virtual memory page holding the vulnerable instruction is mapped to one of the identified Rowhammer-vulnerable physical pages. This is achieved using a technique called memory relaying. This procedure manipulates the kernel's page management module to force the relocation of the target code page. The process involves allocating and deallocating memory in a specific pattern to influence the operating system's memory allocator, eventually "landing" the desired code page onto a physical page susceptible to Rowhammer. The speakers noted that the time expense for this relocation procedure can have a degree of uncertainty due to its interaction with the kernel's internal page management logic.
Step 3: Rowhammer Exploitation
With the vulnerable instruction's code page successfully relocated to a Rowhammer-susceptible physical page, the final step is to execute the Rowhammer exploit. The attacker, running as an unprivileged process, repeatedly accesses (hammers) the "aggressor" memory rows adjacent to the target physical memory row. This continuous hammering induces electrical interference, causing a single bit flip in the targeted instruction's opcode within the victim row.
The research leverages established Rowhammer tools for this purpose, such as HorJ for DDR3 memory modules and TransPass for DDR4 modules. These tools are designed to profile DRAM layouts and identify memory "bugs" or vulnerable locations. Once the instruction is landed under a Rowhammer bug, the hammering takes an instant effect. As the opcode bit flip occurs in memory, the compromised instruction is subsequently fetched and executed by the CPU, leading to an immediate alteration of the program's control flow and the desired malicious outcome, such as an inference failure.
This three-step process meticulously bridges the gap between a theoretical hardware vulnerability and a practical, targeted software exploit, demonstrating a sophisticated attack chain that undermines the integrity of machine learning systems at a fundamental level.
Demo / Proof of Concept
▶ Watch: Demo: Offline stage to find exploitable branches (6:40)
The demonstration of the "Yes, One-Bit-Flip Matters!" attack was divided into two crucial stages: an offline stage for identifying vulnerabilities and an online stage for executing the real-time exploit using Rowhammer.
In the offline stage, the researchers employed an EL-based search algorithm to find and locate vulnerable branching instructions within machine learning libraries. This algorithm systematically scanned the library source code to identify all exploitable branches. For each identified branching structure, the algorithm created a new "poison image" of the library. This poison image was a modified version where the control logic of the target branch was intentionally reversed or altered to simulate the effect of a bit flip. Subsequently, a deep learning model was evaluated against this poison library to observe the adversarial effect on its performance.
During the demo, the researchers iterated over 120 branches specifically within the matrix multiplication function, a core component of many DNN computations. The evaluation process revealed four distinct kinds of faults: model failure, where the model produced incorrect outputs; invalid output, where the output format or content was malformed; denial of service (DoS), where the model ceased to function or became unresponsive; and process crash, leading to the termination of the deep learning application. The results showed these fault types distributed across the 120 branches. For instance, they highlighted results from five specific library images, noting three instances of model failure and two instances of process crashes, demonstrating the diverse impact of these simulated bit flips.
The online stage showcased the real-world execution of the attack. Here, the attacker operated as an unprivileged process, while the deep learning model ran as a separate process on the host machine. The first step for the online attacker was to profile the DRAM layout using tools like HorJ (for DDR3) or TransPass (for DDR4). These tools generate a comprehensive list of Rowhammer "bugs" – physical memory locations susceptible to bit flips – which can number in the tens of thousands on a typical host machine.
Before initiating the actual attack, the deep learning process was assumed to be already running. The attacker then selected the desired type of fault; for the demo, inference failure was chosen. The attack proceeded by matching the fault patterns obtained during the offline stage with the live Rowhammer bug list. In the specific example shown, the inference failure fault pattern matched 88 Rowhammer bugs.
The critical next step was the page relocation procedure, referred to as memory relaying. This involved manipulating the kernel's page management module to strategically land the virtual memory page containing the targeted branch instruction onto a physical page identified as a Rowhammer bug. The speakers acknowledged that this process introduces some uncertainty in timing due to its interaction with the kernel. Once the instruction's page was successfully relocated to a vulnerable physical memory location, the attacker initiated the hammering process. The effect was instantaneous and dramatic: the next batch accuracy immediately dropped from 90% to 6%. This demonstrated a successful end-to-end attack, causing a critical inference failure with a single bit flip. The researchers also mentioned that other fault types could be triggered but were not shown due to time constraints.
Finally, the talk presented experimental evaluations across 10 different classification tasks, four different benchmark datasets, and four types of network architectures. These results robustly confirmed the effectiveness and transferability of their attack, showing that similar fault types consistently appeared across various models and datasets. This comprehensive demonstration powerfully underscored that a single bit flip, when strategically placed in critical library code, can indeed matter profoundly.
Defensive Implications
▶ Watch: Attack effectiveness, transferability, and one-bit advantage (11:10)
The "Yes, One-Bit-Flip Matters!" research exposes a profound vulnerability that extends beyond traditional software security, demanding a holistic approach to defense encompassing both hardware and software layers. The ability of an unprivileged process to induce critical DNN inference failures via a single bit flip in library code necessitates a re-evaluation of current security practices.
Firstly, hardware-level mitigations against Rowhammer are paramount. While memory manufacturers have introduced technologies like Target Row Refresh (TRR) and Error Correcting Code (ECC) memory, their effectiveness can vary. ECC memory can correct single-bit errors, which would directly counter the demonstrated attack. However, not all systems employ ECC, especially in consumer-grade hardware or cloud environments where cost optimization might prioritize non-ECC RAM. Furthermore, advanced Rowhammer variants can sometimes bypass or overwhelm basic TRR mechanisms. Therefore, robust and continuously updated Rowhammer protection mechanisms, potentially at the DRAM controller level, are essential.
Secondly, software integrity and runtime attestation become critical for machine learning libraries. Standard software update mechanisms and file integrity checks (e.g., hash verification) might prevent static tampering, but they are ineffective against runtime memory manipulation. Solutions like Trusted Execution Environments (TEEs) (e.g., Intel SGX, ARM TrustZone) could potentially protect sensitive code pages of ML libraries from external tampering, ensuring their integrity during execution. However, TEEs introduce their own complexities and performance overheads. Lightweight runtime code integrity monitoring, which periodically verifies the integrity of critical code sections in memory, could also be explored, although this must be carefully designed to avoid performance degradation.
Thirdly, memory page protection mechanisms need to be re-evaluated in the context of this attack. While W^X (Write XOR Execute) policies prevent code pages from being writable and executable simultaneously, the Rowhammer attack bypasses this by directly manipulating physical memory. More granular and robust memory protection, perhaps leveraging hardware-enforced memory tagging or stricter kernel-level memory management policies, could help. The memory relaying technique used in the attack exploits the kernel's page management; hardening these kernel modules against such manipulation, or making it harder to predict page allocation, could be a defensive avenue.
Fourthly, system hardening and privilege separation should be rigorously applied. While the attack is unprivileged, limiting the attack surface by reducing the number of processes with direct memory access or enforcing stricter sandboxing for ML inference environments could raise the bar for attackers. Regularly patching operating systems and keeping ML libraries updated is also crucial, as future updates might include mitigations or make the memory layout less predictable.
Finally, for critical applications relying on DNNs (e.g., autonomous vehicles, medical diagnostics, financial systems), the implications are severe. Developers of such systems must consider the possibility of hardware-induced faults leading to catastrophic failures. This might necessitate a shift towards fault-tolerant DNN architectures, redundancy in inference systems, or even hardware-level verification of critical computations. The research highlights the urgent need for a hardware-software co-design approach to machine learning security, moving beyond purely software-centric defenses to address vulnerabilities at the foundational hardware layer.
Key Takeaways
- Single Bit Flips Matter: Contrary to prior beliefs requiring many bit flips, a single bit flip in machine learning library code can cause catastrophic DNN inference failures.
- Code Integrity Over Weight Robustness: The attack targets the integrity of the underlying ML library code, bypassing the inherent robustness of DNN model weights to small perturbations.
- Rowhammer as a Practical Threat: The Rowhammer vulnerability is a practical and exploitable mechanism for achieving deterministic, single-bit flips in sensitive memory regions at runtime.
- Unprivileged and Universal Impact: The attack can be launched from an unprivileged user process and demonstrates universality, affecting various DNN models, datasets, and network architectures.
- Branch Instructions are Critical Targets: Manipulating the opcode of a single branch instruction in library code can fundamentally alter program control flow, leading to diverse and severe faults like model failure, DoS, or process crashes.
- Holistic Security is Essential: Defending against this threat requires a layered approach, integrating hardware-level mitigations (e.g., ECC memory) with software-level code integrity checks and robust memory management.
About the Speaker(s)
The primary speaker for this presentation was Shaofeng Li, who is affiliated with Punch Laboratory. Li and his co-authors, Xinyu Wang, Minhui Xue, Haojin Zhu, Zhi Zhang, Yansong Gao, Wen Wu, and Xuemin (Sherman) Shen, conducted this research. While the transcript specifically mentions Shaofeng Li's affiliation, the collaborative nature of the work suggests a team of dedicated researchers focused on hardware security and its implications for modern computing, particularly in the realm of artificial intelligence. Their collective expertise has brought to light a significant and often overlooked vulnerability at the intersection of hardware and machine learning.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This research unveils a truly novel and alarming attack vector against DNNs, demonstrating that a single bit flip in critical library code, triggered by Rowhammer from an unprivileged process, can cause universal inference depletion. It shatters the myth of DNN robustness against minor hardware faults and demands immediate, serious attention from the entire industry. This isn't just theory; it's a blueprint for catastrophic failure.
Heather Calloway (CISO) — MUST SEE
This research fundamentally shifts the understanding of DNN security, demonstrating that a single, unprivileged bit flip in machine learning library code can cause catastrophic inference failure. It exposes a critical hardware-software integrity vulnerability that demands immediate attention from security leaders and impacts institutional trust in AI deployments.