BitShield: Defending Against Bit-Flip Attacks on DNN Executables
Yanzuo Chen
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · DNN Attack Surfaces
Overview
As artificial intelligence (AI) systems become increasingly integrated into critical aspects of daily life, ranging from autonomous vehicles to financial services, the imperative for their robust security grows exponentially. Insecure AI systems can lead to severe consequences, including misclassifications, financial losses, or even dangerous operational failures. While much attention has historically been paid to high-level adversarial attacks like adversarial examples, backdoors, and model stealing, this talk by Yanzuo Chen, a joint work with Yen Leang and Wangai, delves into a more fundamental and often overlooked threat: bit-flip attacks (BFAs) on Deep Neural Network (DNN) executables.
Key moments
- 0:00 Introduction to Bit-Flip Attacks on AI
- 2:00 DNN Executables: A New Attack Surface
- 4:00 Code Bit-Flips Bypass Existing Defenses
- 4:35 Key Requirements for Robust Bit-Flip Defenses
- 5:20 BitShield's Semantic-Based Defense Approach
- 6:15 Capturing DNN Semantics with Gradients
- 8:00 Self-Defense: Avalanche Effect & Checksums
- 9:00 Masking/Unmasking for Tamper-Evident Semantics
BitShield: Defending Against Bit-Flip Attacks on DNN Executables
Speakers: Yanzuo Chen
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=Xgy0z5Ek6tI
Overview
As artificial intelligence (AI) systems become increasingly integrated into critical aspects of daily life, ranging from autonomous vehicles to financial services, the imperative for their robust security grows exponentially. Insecure AI systems can lead to severe consequences, including misclassifications, financial losses, or even dangerous operational failures. While much attention has historically been paid to high-level adversarial attacks like adversarial examples, backdoors, and model stealing, this talk by Yanzuo Chen, a joint work with Yen Leang and Wangai, delves into a more fundamental and often overlooked threat: bit-flip attacks (BFAs) on Deep Neural Network (DNN) executables.
BitShield introduces a critical shift in perspective, highlighting that existing BFA research, which predominantly focuses on manipulating DNN model weights, fails to address a pervasive and highly potent attack surface residing within the compiled code of DNN executables. These executables, generated by deep learning compilers like Apache TVM and Meta Glow for performance optimization, are shown to be highly vulnerable to single-bit corruptions in their operational code. Such low-level tampering can not only degrade model accuracy significantly but also bypass conventional defenses designed solely for weight integrity. BitShield proposes a novel, unified, and self-defending framework that leverages semantic analysis and cryptographic principles to protect DNN executables against both code-based and weight-based bit-flip attacks, demonstrating high efficacy with minimal performance overhead.
Background
▶ Watch: Introduction to Bit-Flip Attacks on AI (0:00)
The pervasive integration of AI across various domains underscores the critical need for secure AI systems. While the research community has extensively explored numerous high-level attacks against AI models—such as adversarial examples that manipulate inputs to cause misclassifications, backdoors that embed hidden vulnerabilities, and model stealing that extracts proprietary model information—these typically operate at the conceptual or data level of the AI model. The focus of this presentation, however, shifts to a more foundational, infrastructure-level threat: bit-flip attacks (BFAs).
BFAs are a class of attacks that directly manipulate individual bits in memory, specifically in DRAM (Dynamic Random-Access Memory). The most prominent example of such an attack is Rowhammer, a hardware bug discovered in modern DRAM that allows an attacker to induce bit flips in memory cells by repeatedly accessing adjacent rows. Researchers have consistently demonstrated the feasibility of Rowhammer attacks across virtually every generation of memory hardware, highlighting its persistent threat. In the context of AI and Machine Learning (ML) systems, prior research has already established that BFAs can compromise DNN models by flipping data bits within the model weights. This distortion of weight data subsequently corrupts the model's internal representations and decision-making processes, leading to incorrect predictions and degraded performance.
However, a significant blind spot in existing research, as identified by the BitShield team, is the neglect of DNN executables. These are specialized binary programs compiled from DNN models using sophisticated deep learning compilers such as Apache TVM and Meta Glow (formerly Facebook Glow). The primary motivation for using DNN executables is performance; they are meticulously optimized at both the computational graph level (making the model more efficient) and the target hardware platform level (maximizing hardware potential). Essentially, DNN executables are compiled code that encapsulates all the operators and logic of the original DNN model.
The critical insight of BitShield is that while previous BFA research concentrated solely on bit flips in model weights, the compiled code within these DNN executables presents an entirely new, pervasive, and highly vulnerable attack surface. A preliminary survey conducted by the researchers revealed that code-based bit flips can allow single bit corruption to significantly degrade model accuracy. This vulnerability persists even in quantized models, which were previously considered more robust against BFAs due to their reduced precision and inherent noise tolerance.
Furthermore, existing defensive mechanisms are found to be inadequate. These defenses primarily focus on protecting the integrity of model weights, overlooking the executable code. This narrow scope renders them vulnerable to bypass. As illustrated by a simple example in the talk, a bit flip in the code could transform a legitimate computation instruction into something innocuous like leave and return. If a defense mechanism is implemented after this modified instruction, the altered control flow would cause the defense to be entirely skipped or bypassed, leaving the model unprotected. This demonstrates that DNN executables suffer from both the "old" (weight-based) and "new" (code-based) types of bit-flip attacks.
To effectively counter both categories of BFAs, the BitShield team identified several key requirements for a robust defense: it must be unified and generic, capable of protecting against both weight and code attacks; it must be self-defending, meaning it cannot be easily bypassed; it needs to be applicable to different models; and critically, it must be performant enough for deployment in production environments.
Key Findings
▶ Watch: Code Bit-Flips Bypass Existing Defenses (4:00)
The research presented in "BitShield: Defending Against Bit-Flip Attacks on DNN Executables" unveils several crucial findings that redefine the landscape of bit-flip attack vulnerabilities and defenses in Deep Neural Network systems:
- Undiscovered Attack Surface in DNN Executables: The primary finding is the identification of a significant and previously overlooked attack surface for bit-flip attacks within the compiled code of DNN executables. Unlike prior work that focused on model weights, BitShield demonstrates that the operational instructions themselves are highly vulnerable.
- Pervasive Code-Based Vulnerability: A preliminary survey revealed that code-based bit flips are pervasive. A single bit corruption within the executable code is sufficient to cause a significant degradation in model accuracy. This vulnerability even extends to quantized models, which were previously believed to offer increased robustness against BFAs, challenging existing assumptions about their security.
- Bypass of Existing Defenses: Code-based bit flips pose a unique threat by enabling attackers to bypass existing defense mechanisms. Defenses that only monitor and protect model weights can be rendered ineffective if a bit flip modifies the program's control flow, allowing the malicious code to execute unchecked or to skip integrity checks entirely.
- Novel Semantic-Based Detection: BitShield introduces an innovative semantic-based approach to detect BFAs. By inversely measuring how a model's output could be transformed into a random guess using KL divergence and backpropagation, the defense effectively captures and monitors the model's intrinsic semantics. This approach achieved a 92% mitigation rate against weight-based attacks.
- Self-Defending Mechanism via Checksum Fusion: To counter the bypass threat posed by code flips, BitShield integrates a self-defense mechanism inspired by the cryptographic avalanche effect. By fusing code checksums into the semantic calculation via masking and unmasking operations, any unauthorized modification to the code drastically alters the captured semantics, making tampering immediately evident.
- Early Damage Prevention with Checksum Canaries: The defense further incorporates checksum canaries – plain checksum checks inserted directly into the model's execution path. These canaries enable immediate detection of code mismatches and halt model execution, preventing the propagation of errors and limiting potential damage.
- High Effectiveness and Low Overhead: The comprehensive evaluation of BitShield demonstrated its exceptional efficacy. The defense successfully reduced the attack success rate (ASR) of code-based attacks from 100% to 0% and weight-based attacks from 96% to 7%. Crucially, this robust protection comes with a remarkably low performance overhead of only 2.47%, making it highly practical for real-world production environments.
Technical Deep Dive
▶ Watch: BitShield's Semantic-Based Defense Approach (5:20)
BitShield's technical innovation lies in its multi-layered approach to defending against bit-flip attacks, integrating semantic monitoring, self-defense mechanisms, and early detection. The core philosophical shift is to model attacks from the perspective of semantics, recognizing that DNN predictions are fundamentally a decision process driven by both code logic and model weights. Bit-flip attacks, whether targeting weights or code, ultimately aim to break this decision process and alter the model's intended semantics.
Traditional program analysis tools, such as Control Flow Graphs (CFGs), are often effective for understanding the semantics of conventional software. However, they fall short for DNNs, where semantics are implicit and deeply embedded within complex computational processes. To address this, BitShield draws inspiration from the concept of gradients in machine learning. Just as gradients guide a randomly initialized model to learn a useful decision process during training, BitShield inversely measures this process during runtime.
Capturing DNN Semantics:
The methodology for capturing DNN semantics involves a clever inverse measurement technique:
- Model Output (Y): The defense starts with the DNN model's output,
Y. - Random Guess Vector (U): A vector
Uis prepared, having the same length asY, but with all its elements set to an equal value. ThisUvector represents a completely random guess, effectively signifying maximum uncertainty or a lack of meaningful prediction. - Distance Measurement: The KL divergence (Kullback-Leibler divergence) is then used to measure the statistical distance between the model's actual output
Yand the random guess vectorU. KL divergence quantifies how one probability distribution (or in this case, a model output interpreted as a distribution) diverges from a second, reference probability distribution (U). - Backpropagation: This measured KL divergence value is then backpropagated through the DNN model to individual layers. By doing so, BitShield effectively quantifies how "far" the current model state (and its decision process) is from a random guess. This backpropagated value serves as a runtime representation of the model's semantics.
- Reference Semantics: To establish a baseline for comparison, the "normal" semantics are recorded using the model's training data. Any significant deviation from this established baseline during runtime indicates a potential attack. This semantic checking mechanism alone is highly effective, achieving a 92% mitigation rate for weight-based bit-flip attacks.
Self-Defense Mechanism for Code Flips:
While semantic checks are powerful, the threat of code-based bit flips bypassing defenses remains. To counter this, BitShield integrates a self-defense mechanism inspired by the avalanche effect in cryptography. The avalanche effect dictates that a small change in input (e.g., a single bit flip in code) should lead to a drastic and unpredictable change in the output, making tampering evident.
BitShield achieves this by fusing code checksums directly into the semantic calculation process using a pair of masking and unmasking operations. Conceptually, if the semantic capturing process involves a simplified computation like O = W * V (where W is a parameter), this is rewritten as:
O = Unmask(Mask(W, K_runtime), K_embedded) * V
Here:
Mask()andUnmask()are inverse operations, similar to encryption and decryption.K_embeddedis a checksum of the DNN executable code, calculated and embedded at compile time.K_runtimeis a checksum of the DNN executable code, calculated dynamically at runtime.
These masking and unmasking operations are designed such that they only cancel each other out if the same key is provided. In BitShield's design, the "keys" are the embedded compile-time checksum (K_embedded) and the dynamically calculated runtime checksum (K_runtime). If the DNN executable's code has been tampered with—even by a single bit flip—K_runtime will no longer match K_embedded. Consequently, the masking and unmasking operations will not cancel out, leading to a drastic and easily detectable change in the captured semantics (O), thereby exposing the attack. This mechanism ensures that even if an attacker attempts to flip bits in the code to bypass the semantic check, the check itself will become corrupted in a detectable way.
Checksum Canary for Early Damage Prevention:
To prevent the propagation of errors and limit potential damage, BitShield incorporates a checksum canary. This involves inserting simple, plain checksum checks directly into critical points within the model's execution flow. As soon as a checksum mismatch is detected by these canaries, the model's execution is immediately halted. This proactive measure prevents the corrupted logic from processing further inputs or producing erroneous outputs, effectively containing the attack's impact.
Regarding the type of checksum used, the speaker clarified during the Q&A that BitShield does not use cryptographically secure checksums like SHA-1 or SHA-256 due to their significant performance overhead. Instead, they opt for a faster, simpler checksum, specifically mentioning "at 32" (likely referring to a 32-bit checksum like CRC32). The justification for this choice is that while not cryptographically collision-resistant, a 32-bit checksum provides sufficient entropy (four bytes or 32 bits) in the context of bit-flip attacks. The speaker argued that it would be "really difficult for the attackers to gather... memory templates" (a Rowhammer-specific term referring to the necessary conditions to reliably flip specific bits) to simultaneously flip enough bits to achieve a checksum collision that would bypass the defense, especially when dealing with distributed bit flips across the model.
Demo / Proof of Concept
▶ Watch: Capturing DNN Semantics with Gradients (6:15)
While the talk did not feature a live demonstration of BitShield or a specific proof-of-concept tool being showcased, the authors thoroughly evaluated its effectiveness through rigorous experimentation. The evaluation setup and results were presented as the primary evidence of the defense's capabilities.
The evaluation involved a comprehensive threat model where attackers were considered white-box and adaptive, meaning they possessed full knowledge of the defense mechanisms and would actively attempt to bypass them. This included aggressive and stealthy code-based attack profiles, as well as a state-of-the-art weight-based attacker. The attacks were simulated on five different DRAM platform profiles, derived from existing research, to ensure realistic conditions. The defense was assessed using key metrics: attack success rate (ASR), post-attack accuracy, and performance overhead. An attack was deemed successful if it caused an accuracy drop of at least 3%, a threshold considered representative and significant.
The results demonstrated BitShield's remarkable effectiveness:
- Code-based attacks: The defense reduced the attack success rate from 100% to 0%, effectively mitigating all code-based bit-flip attacks.
- Weight-based attacks: The attack success rate was significantly reduced from 96% to 7%, indicating a 93% mitigation rate for this category.
- Post-attack accuracy: Even in the 7% of cases where weight-based attacks still succeeded, their impact was severely limited; models were no longer degraded to random guessers, unlike without the defense.
- Performance overhead: Critically, BitShield introduced a minimal performance overhead of approximately 2.47%, underscoring its high practicality for deployment in production environments where performance is paramount.
Defensive Implications
▶ Watch: Masking/Unmasking for Tamper-Evident Semantics (9:00)
The findings presented by BitShield carry significant implications for the design and implementation of defenses against low-level memory attacks on AI systems. Defenders must fundamentally re-evaluate their current strategies and adopt a more holistic approach:
- Expand Attack Surface Awareness: Defenders must recognize that the attack surface for bit-flip attacks extends beyond model weights to include the compiled code of DNN executables. Relying solely on weight integrity checks is insufficient and leaves a critical vulnerability gap.
- Implement Semantic-Based Runtime Monitoring: Adopt and integrate runtime semantic checks that can detect deviations in the model's internal decision process. BitShield's approach, using KL divergence and backpropagation to measure the "randomness" of model outputs, offers a robust framework for this. This moves beyond simple data integrity to behavioral integrity.
- Integrate Self-Defending Mechanisms: Defenses themselves must be resilient to bypass. Techniques like BitShield's checksum fusion via masking and unmasking operations are crucial. This ensures that even if an attacker attempts to tamper with the defense's code, the act of tampering is itself detected, leveraging principles akin to the avalanche effect.
- Prioritize Early Detection and Containment: Employ mechanisms like checksum canaries to detect code integrity violations as early as possible within the execution flow. Halting execution immediately upon detection prevents the propagation of errors and limits the potential damage caused by corrupted logic.
- Collaborate with Deep Learning Compiler Developers: The vulnerability of DNN executables highlights the need for security to be integrated at the compilation stage. Defenders should advocate for and collaborate with developers of deep learning compilers (e.g., Apache TVM, Meta Glow) to embed BitShield-like defenses directly into the compilation process, making them an inherent part of the generated executables.
- Re-evaluate Quantized Model Robustness: Do not assume that quantized models are inherently robust against all forms of bit-flip attacks. While they may offer some resilience against weight-based data corruption, they remain vulnerable to code-based bit flips that can drastically alter their behavior.
- Balance Security with Performance: The low overhead of BitShield (2.47%) demonstrates that effective, low-level security can be achieved without rendering systems impractical for production. Defenders should prioritize solutions that offer a strong security posture with minimal performance impact.
- Adaptive Threat Modeling: Given the "white-box and adaptive" nature of the attackers in BitShield's threat model, defenders should continuously update their threat models to account for adversaries who are aware of existing defenses and actively seek to bypass them.
Key Takeaways
- Bit-flip attacks (BFAs) on Deep Neural Network (DNN) executables present a significant and overlooked attack surface within their compiled code, extending beyond traditional model weight manipulations.
- Code-based bit flips are highly potent, capable of causing drastic accuracy degradation from a single bit corruption and, crucially, can bypass existing weight-focused defenses.
- BitShield introduces a novel semantic-based defense that captures DNN behavior using KL divergence and backpropagation, effectively mitigating 92% of weight-based attacks.
- To achieve self-defense against code modifications and bypass attempts, BitShield fuses code checksums into semantic calculations via masking/unmasking operations and employs checksum canaries for early detection and error containment.
- The comprehensive BitShield defense is highly effective, reducing code-based attack success rates from 100% to 0% and weight-based attacks from 96% to 7%, with a minimal performance overhead of just 2.47%.
- Defenders must adopt a holistic security strategy for AI systems, protecting both model weights and their compiled executables from low-level memory attacks, and integrating self-defending mechanisms at the infrastructure level.
About the Speaker(s)
The talk "BitShield: Defending Against Bit-Flip Attacks on DNN Executables" was presented by Yanzuo Chen. This work was a collaborative effort with Yen Leang and Wangai. The transcript does not provide specific titles or affiliations for the speakers, but it indicates Yanzuo Chen has presented at the NDSS Symposium previously, referring to himself as "Yenzo" and noting "it's me again."
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
BitShield addresses a genuinely underexplored attack surface — the compiled code of DNN executables, not just model weights — and backs it with a technically coherent defense combining semantic monitoring, checksum fusion via masking/unmasking, and canary-based early termination. The results are credible: 100% mitigation of code-based BFAs and 93% reduction in weight-based ASR at 2.47% overhead. This is real systems security work, not ML security theater.
Heather Calloway (CISO) — WEAK
Technically credible work identifying a real blind spot in DNN executable security, with measurable results and a novel self-defending architecture. But this is a research paper delivered to a research audience — it doesn't translate to operational decisions, and the governance and deployment questions it raises go completely unaddressed.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025