Forget and Rewire: Enhancing the Resilience of Transformer-based Models against Bit-Flip Attacks
Najmeh Nazari (University of California Davis), Hossein Sayadi, Setareh Rafatirad, Khaled N. Khasawneh, Houman Homayoun
33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24
Overview
In an era where Transformer-based models underpin a vast array of critical applications, from sophisticated text generation to precise image classification, their inherent vulnerabilities pose a significant threat. This talk, presented by Najmeh Nazari from the University of California Davis, along with her co-authors, delves into a novel approach to fortify these powerful models against insidious bit-flip attacks. The core of their research introduces an operation dubbed "Forget and Rewire" (F&R), drawing inspiration from neuroplasticity to dynamically reconfigure model connections, thereby enhancing resilience without compromising performance.

Key moments
- 0:00 Introduction and Transformer vulnerability to bit-flip attacks
- 2:00 Detailed threat model and attacker categorization
- 4:00 Limitations of existing defense strategies
- 6:00 Inspiration and core 'Forget and Rewire' idea
- 8:00 Practical example of Forget and Rewire mechanism
- 10:00 Evaluation results: Accuracy and robustness gains
- 12:00 Conclusion and summary of Forget and Rewire operation
Forget and Rewire: Enhancing the Resilience of Transformer-based Models against Bit-Flip Attacks
Speakers: Najmeh Nazari; Hossein Sayadi; Setareh Rafatirad; Khaled N. Khasawneh; Houman Homayoun
Conference: USENIX Security '24
YouTube: https://www.youtube.com/watch?v=1PtC17xVyjg
Overview
In an era where Transformer-based models underpin a vast array of critical applications, from sophisticated text generation to precise image classification, their inherent vulnerabilities pose a significant threat. This talk, presented by Najmeh Nazari from the University of California Davis, along with her co-authors, delves into a novel approach to fortify these powerful models against insidious bit-flip attacks. The core of their research introduces an operation dubbed "Forget and Rewire" (F&R), drawing inspiration from neuroplasticity to dynamically reconfigure model connections, thereby enhancing resilience without compromising performance.
The presentation highlights the critical challenge of securing deep neural networks, particularly Transformers, against memory-level attacks that can subtly degrade model accuracy. By leveraging the fundamental architecture of Transformers, specifically their efficient linear layers, the researchers propose a defense mechanism that not only increases the cost for attackers but also offers compatibility with existing pre-trained models and other defensive strategies. This work is pivotal for safeguarding the integrity and reliability of AI systems deployed in sensitive environments, addressing a gap in current defense methodologies that often entail significant retraining overhead or accuracy trade-offs.
The significance of this research lies in its pragmatic and effective solution to a pervasive security concern in AI. Bit-flip attacks, often stealthy and difficult to detect, can have severe consequences for applications relying on accurate model predictions. By providing a method to inherently strengthen the model's architecture against such manipulations, Nazari and her team contribute a vital tool for developers and security practitioners seeking to deploy more robust and trustworthy Transformer-based AI systems in an increasingly adversarial landscape.
Background
▶ Watch: Introduction and Transformer vulnerability to bit-flip attacks (0:00)
Transformers have emerged as the backbone of modern artificial intelligence, demonstrating unparalleled performance across diverse tasks due to their unique attention mechanism and highly efficient linear layers. These linear layers are crucial for the scalability of Transformers during both training and inference. Despite their impressive capabilities and vast parameter counts, Transformers, like other deep neural networks, remain susceptible to bit-flip attacks. These attacks, exemplified by techniques such as Rowhammer and DeepHammer, exploit physical vulnerabilities in memory hardware to introduce errors directly into the memory locations where model parameters are stored. The objective of an adversary is to subtly alter these parameters, leading to a degradation in the model's accuracy and reliability.
The typical flow of a bit-flip attack involves adversaries employing a progressive gradient search to pinpoint the most critical parameters within the model. The goal is to minimize the number of bit flips required to achieve a significant degradation in accuracy, thereby keeping the attack stealthy and difficult to detect. The threat model considered in this research assumes an adversary with the capability to execute precise, multi-bit flip attacks during the deployment phase of a model. This operates under a white-box scenario, where the attacker possesses varying levels of knowledge about the model's architecture and parameters. The researchers categorize adversaries into three types based on their knowledge:
- Basic Adversary: Has no knowledge of the defense strategy or access to gradient values, but can evaluate bit-flip attacks on an equivalent undefended model.
- Expert Adversary: Aware of the defense operation and can obtain gradient values during deployment, but lacks access to the defense configuration.
- Oracle Adversary: Possesses comprehensive access and knowledge, representing the strongest possible attacker.
Existing defense strategies against bit-flip attacks generally fall into two categories: model hardening and detection and recovery.
Model hardening techniques aim to make the model inherently more resistant to bit flips:
- Quantization strategies involve binarizing model weights to reduce sensitivity to bit flips. However, these often necessitate costly retraining of the model and can lead to a noticeable loss in accuracy.
- Modifying activation functions, such as ReLU, can slightly mitigate the impact of bit flips but do not offer a complete solution.
- EDGES is a randomization technique that employs a dynamic exit mechanism with additional classifiers in hidden layers to randomize the model's architecture. While somewhat effective, it introduces significant complexity and additional computational overhead.
Detection and recovery approaches focus on identifying and correcting errors:
- Error Correction Codes (ECCs), typically implemented at the hardware level, can recover single-bit errors and detect two-bit errors. However, their capabilities are often insufficient against more sophisticated, multi-bit Oracle attacks.
- NeoPods is a newer approach that uses non-critical parameters as "bait" or "lures" to attract attackers. It employs checksums on these bait parameters to detect modifications. The significant drawback, however, is that expert and Oracle attackers can bypass these bait parameters by setting gradient value thresholds and directly targeting other critical parts of the model, rendering NeoPods ineffective against advanced threats.
The limitations of these existing methods—ranging from the need for extensive retraining and accuracy degradation to susceptibility to advanced adversaries and increased overhead—highlight the urgent need for a more robust, efficient, and broadly applicable defense mechanism for Transformer-based models. This context forms the foundational motivation for the "Forget and Rewire" approach.
Key Findings
▶ Watch: Limitations of existing defense strategies (4:00)
The central contribution of this research is the introduction of the Forget and Rewire (F&R) operation, a novel defense mechanism designed to significantly enhance the resilience of Transformer-based models against bit-flip attacks. The inspiration for F&R is drawn from neurosychology, specifically the concept captured by the phrase "neurons that fire together, wire together." This principle suggests that if certain information flow in the brain is deemed unimportant, it can be "forgotten," and new connections can be "rewired" to strengthen more critical pathways. Applying this biological metaphor to neural networks, the core idea behind F&R is to identify and "forget" unimportant connections within a linear layer, then "rewire" these resources to make important connections more robust against adversarial manipulation.
The operationalization of F&R involves a systematic process to identify and reconfigure model parameters. The researchers employ a gradient search technique to sort parameters based on their sensitivity. This allows for the precise identification of critical neurons, which exhibit the highest gradient values, indicating their significant contribution to the model's output. Conversely, dead neurons are identified by their zero gradient values, marking them as candidates for the "forget" operation due to their minimal impact. Once identified, a strategic matching process pairs these dead neurons with critical neurons, establishing a forget and rewire configuration that dictates how the model's internal connections will be reallocated. This configuration is then saved and applied during model deployment.
The effectiveness of F&R stems from its ability to redistribute the importance of critical parameters, thereby increasing the cost an attacker must incur to degrade model performance. By sharing the responsibility of a critical connection across multiple pathways, F&R makes it significantly harder for an attacker to achieve the same level of degradation by flipping a single bit or a small set of bits. The evaluation results demonstrate that F&R substantially increases the number of bit flips required to degrade the model's accuracy, even when confronted with the most potent Oracle attacker. Specifically, the normalized number of bit flips required for a 10% accuracy degradation was found to be approximately 1.68 times higher for F&R-protected models compared to unprotected ones, even against an Oracle adversary. Furthermore, the research illustrates a quadratic increase in required bit flips as the number of F&R parameters increases, indicating a strong scaling effect in resilience.
Crucially, F&R achieves this enhanced robustness with minimal impact on the model's original performance. The reported accuracy loss after applying F&R is less than 2%, demonstrating a favorable trade-off between security and accuracy. This minimal overhead is largely due to the fact that F&R does not require retraining the model, making it highly practical for application to existing pre-trained models. This is a significant advantage over many other defense mechanisms that demand extensive retraining, which can be computationally expensive and time-consuming. Moreover, the approach is designed to be compatible with other defenses, suggesting it can be integrated into a multi-layered security strategy, further strengthening the overall robustness of Transformer-based AI systems.
Technical Deep Dive
▶ Watch: Inspiration and core 'Forget and Rewire' idea (6:00)
The core of the Forget and Rewire (F&R) operation lies in its intelligent redistribution of parameter importance within the linear layers of a Transformer model. To understand its technical intricacies, let's consider a simplified standard linear layer where neurons in a preceding layer (Layer 1) connect to neurons in a subsequent layer (Layer 2). In such a configuration, each neuron's activation in Layer 1 contributes directly to the output of neurons in Layer 2 through weighted connections.
The first step in F&R is to identify which connections are critical and which are negligible. This is achieved through a gradient search process. The researchers determine the sensitivity of each parameter by calculating its gradient value with respect to the model's output or loss. Connections with higher gradient values are deemed critical neurons (e.g., M1 in the example), indicating their substantial influence on the model's predictions. Conversely, connections associated with dead neurons (e.g., M2), characterized by zero or near-zero gradient values, are identified as unimportant or redundant. These dead neurons become candidates for the "forget" operation.
Once critical and dead neurons are identified, the F&R mechanism establishes a forget and rewire configuration. This configuration is essentially a mapping that dictates which dead neuron's connection should be "forgotten" and subsequently "rewired" to bolster a critical neuron's pathway. For instance, if M1 is a critical neuron with connection W1, and M2 is a dead neuron with connection W2, the configuration might specify that W2 should be rewired to M1.
Applying this configuration involves a precise sequence of operations:
- Forget M2's connection: The original connection associated with the dead neuron M2 (represented by weight W2) is conceptually "ignored" or de-emphasized. This means its direct contribution via its original pathway is nullified.
- Rewire W2 with M1's connection: Instead of discarding W2 entirely, its computational pathway is redirected. W2 is effectively "rewired" to become a redundant, yet active, pathway for M1.
- Replace W2's value with W1's value: To ensure that the newly rewired pathway contributes meaningfully and consistently with the critical neuron, the weight value of W2 is replaced with the value of W1. This ensures that both pathways (the original W1 and the newly repurposed W2) now carry information critical to M1.
- Redistribute activations from M1 to both W1 and W2: The crucial step for enhancing robustness is the redistribution of activation. The activation originating from the critical neuron M1 is now split and directed equally through both the original W1 pathway and the newly rewired W2 pathway. This means that both pathways contribute equally to the output, effectively preserving the network's overall functionality and the critical neuron's intended impact.
The profound impact of this redistribution is evident in its effect on an attacker. By making the critical neuron's importance shared between W1 and W2, the F&R operation significantly increases the cost of the attack. If an adversary attempts to degrade the model by flipping bits in W1, they will only achieve partial degradation because W2 now carries an equal share of M1's importance. To achieve the same level of degradation that flipping W1 alone would have caused in an undefended model, the attacker must now also identify and successfully attack W2. This effectively doubles the effort required for a targeted attack on that specific critical connection, making it much harder and more resource-intensive for the adversary to achieve their desired outcome while maintaining stealth.
The entire F&R process is not a one-shot operation but rather an iterative one. After applying the F&R configuration, the researchers simulate bit-flip attacks to assess the achieved robustness. If the desired level of resilience is not met, the process can be repeated, potentially with different selections of dead and critical neurons or more extensive rewiring, until the model exhibits the required robustness. This iterative refinement ensures that the defense is tailored to achieve a specific security posture.
Demo / Proof of Concept
▶ Watch: Evaluation results: Accuracy and robustness gains (10:00)
While the presentation did not feature a live, interactive demonstration of the Forget and Rewire (F&R) operation, the researchers provided compelling evidence of its efficacy through a rigorous simulation-based evaluation acting as a robust proof of concept. This evaluation showcased how F&R significantly enhances the resilience of Transformer-based models against bit-flip attacks by presenting quantitative results on key metrics.
The evaluation setup utilized a diverse range of models and datasets to ensure the generalizability of F&R. Specifically, the researchers employed a custom model alongside two popular pre-trained models, demonstrating F&R's applicability beyond bespoke architectures. These models were tested using several well-known datasets, reinforcing the practical relevance of the findings. The primary evaluation metrics focused on accuracy (to assess performance overhead) and robustness (to quantify attack resistance).
One of the critical findings concerned the accuracy loss introduced by F&R. The left figure presented in the talk illustrated the impact of increasing the number of F&R parameters on the average accuracy loss. As more parameters were subjected to F&R operations, a slight increase in accuracy loss was observed, highlighting an inherent trade-off between accuracy and security. However, the accompanying table explicitly stated that F&R leads to minimal accuracy loss, less than 2%. This is a crucial practical advantage, as it means the defense can be deployed without significantly degrading the model's intended performance, making it suitable for real-world applications where high accuracy is paramount.
The core of the proof of concept lay in demonstrating the enhanced robustness against bit-flip attacks. The researchers quantified this by measuring the normalized number of bit flips required to degrade the model's accuracy by 10%. This metric directly reflects the increased effort an attacker must expend. The results were particularly striking, even against the most formidable adversary type: the Oracle attacker. For an Oracle attacker, who possesses complete knowledge and access, the number of bit flips required to achieve a 10% accuracy degradation was approximately 1.68 times higher for F&R-protected models compared to non-F&R models. This demonstrates a substantial increase in attack cost and difficulty, even when the adversary has perfect information.
Further reinforcing the robustness, a graph on the right side of the evaluation section illustrated a compelling trend: increasing the number of F&R parameters leads to a quadratic increase in the number of bit flips required to achieve the same level of degradation. This non-linear scaling indicates that even a modest increase in F&R application can yield disproportionately higher resilience, providing an efficient way to tune the defense level.
The talk also briefly mentioned other aspects of the evaluation, including the impact on overhead, compatibility with Dropout and Pruning techniques, and resilience against adversarial example input attacks. For a detailed understanding of these additional evaluations, the audience was encouraged to consult the full research paper. The presented simulation results effectively served as a robust proof of concept, clearly demonstrating F&R's ability to significantly enhance model resilience with minimal performance cost.
Defensive Implications
▶ Watch: Conclusion and summary of Forget and Rewire operation (12:00)
The Forget and Rewire (F&R) operation presents several significant implications for defenders seeking to secure Transformer-based models against bit-flip attacks. Its design principles offer practical advantages that can be integrated into current and future AI security strategies.
Firstly, a standout feature of F&R is its ability to be applied to pre-trained models without requiring retraining. This is a monumental advantage, as retraining large Transformer models is an incredibly resource-intensive and time-consuming process, often costing millions of dollars and consuming vast computational power. By circumventing this requirement, F&R enables immediate and cost-effective deployment on existing models, making it accessible to a wider range of organizations and applications. Defenders can thus enhance the security posture of their deployed AI systems with minimal disruption and expenditure.
Secondly, F&R directly addresses the core objective of bit-flip attackers: minimizing the number of bit flips for stealth and efficiency. By strategically redistributing the importance of critical parameters, F&R quadratically increases the number of bit flips required to achieve a desired level of degradation. As demonstrated, even against an Oracle attacker, the attack cost is elevated by approximately 1.68 times for a 10% accuracy drop. This makes bit-flip attacks significantly more expensive and less practical for adversaries, forcing them to expend more resources and increasing the likelihood of detection due to more extensive memory manipulation.
Thirdly, the research highlights that F&R introduces minimal accuracy loss (less than 2%). This ensures that the defensive measure does not unduly compromise the primary function of the AI model. In many security solutions, there's a steep trade-off between security and utility. F&R manages to strike an excellent balance, providing robust protection while preserving the high performance expected of Transformer models.
Moreover, the F&R approach is designed to be compatible with other existing defense mechanisms. This allows defenders to adopt a layered security approach, combining F&R with hardware-level error correction codes, other software hardening techniques, or even adversarial training methods. Such a multi-faceted defense strategy creates a stronger overall security perimeter, making it exceedingly difficult for even sophisticated attackers to bypass all protections. For instance, F&R could reduce the burden on ECCs by making critical parameters less susceptible to single, targeted bit flips, allowing ECCs to focus on more generalized memory errors.
Finally, the iterative nature of F&R implies a potential for adaptive defense. Defenders could potentially monitor attack attempts or model performance degradation in the field and dynamically adjust the F&R configuration to enhance resilience in specific areas identified as vulnerable. This could lead to more robust and resilient AI systems that can evolve their defenses over time, much like biological systems adapt to threats. The insights from F&R could also inform hardware-software co-design efforts, guiding the development of memory architectures that are inherently more resilient to bit-flip attacks by understanding how critical information is distributed.
Key Takeaways
- Forget and Rewire (F&R) is a novel defense mechanism enhancing the resilience of Transformer-based models against bit-flip attacks.
- Inspired by neuroplasticity, F&R identifies critical and dead neurons using gradient search, then strategically redistributes the importance of critical parameters to non-essential ones.
- The approach significantly increases the cost of attack, requiring approximately 1.68 times more bit flips for a 10% accuracy degradation, even against an Oracle attacker.
- F&R incurs minimal accuracy loss (less than 2%) and does not require retraining, making it highly practical for application to existing pre-trained models.
- The defense is compatible with other security measures, allowing for the creation of layered and more robust AI security architectures.
- Increasing the number of F&R parameters leads to a quadratic increase in the number of bit flips required for successful attacks, offering scalable resilience.
About the Speaker(s)
The primary presenter of this research was Najmeh Nazari from the University of California Davis. She, along with her co-authors Hossein Sayadi, Setareh Rafatirad, Khaled N. Khasawneh, and Houman Homayoun, represents a team of researchers from UC Davis dedicated to advancing the security and resilience of artificial intelligence and machine learning models. Their work focuses on critical areas such as hardware security, deep learning robustness, and developing innovative techniques to protect AI systems from various adversarial attacks, including memory-level manipulations like bit-flip attacks. Their collective expertise spans computer engineering, electrical engineering, and computer science, contributing to cutting-edge solutions for the challenges facing modern AI deployment.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This research introduces "Forget and Rewire" (F&R), a novel, neuroplasticity-inspired defense against bit-flip attacks in Transformer models. By intelligently redistributing critical parameter importance, F&R quadratically increases attacker effort while maintaining model accuracy and requiring no retraining. It's a practical, high-impact defense for real-world AI deployment.
Heather Calloway (CISO) — STRONG ACCEPT
This research presents a highly practical and effective defense against bit-flip attacks on Transformer models. Its key strength is the ability to significantly enhance model resilience and increase attacker cost without requiring expensive retraining, making it immediately relevant for organizations deploying critical AI.