PBP: Post-training Backdoor Purification for Malware Classifiers
Dung Thuy Nguyen
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Malware
Overview
In an era where machine learning (ML) and deep learning (DL) models are increasingly becoming foundational components for critical security tasks, such as malware detection, ensuring their integrity and robustness against sophisticated attacks is paramount. This talk, presented by Dung Thuy Nguyen at the NDSS Symposium, addresses a particularly insidious threat: backdoor attacks against deep neural network (DNN) based malware classifiers. The paper introduces PBP (Post-training Backdoor Purification), a novel method designed to detect and eliminate backdoors embedded within already deployed or trained models, without requiring any prior knowledge of the attack's specifics.
Key moments
- 0:00 Introduction: Purifying backdoors in malware classifiers
- 1:30 Understanding how backdoor attacks manipulate models
- 3:30 Challenges of defending against backdoor attacks
- 4:10 Introducing PBP: Post-training Backdoor Purification method
- 5:00 Core insight: Detecting and understanding backdoor neurons
- 6:40 Purification Step 1: Detecting backdoor neurons
- 8:00 Purification Step 2: Activation shift finetuning
- 8:30 Evaluation settings for different backdoor attack types
PBP: Post-training Backdoor Purification for Malware Classifiers
Speakers: Dung Thuy Nguyen
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=aWpWSunwW7U
Overview
In an era where machine learning (ML) and deep learning (DL) models are increasingly becoming foundational components for critical security tasks, such as malware detection, ensuring their integrity and robustness against sophisticated attacks is paramount. This talk, presented by Dung Thuy Nguyen at the NDSS Symposium, addresses a particularly insidious threat: backdoor attacks against deep neural network (DNN) based malware classifiers. The paper introduces PBP (Post-training Backdoor Purification), a novel method designed to detect and eliminate backdoors embedded within already deployed or trained models, without requiring any prior knowledge of the attack's specifics.
The core problem PBP seeks to solve is the vulnerability of ML models to poisoning during their training phase. Adversaries can subtly inject malicious data into the training set, embedding a hidden trigger that, when present in an input sample, forces the model to misclassify it in a predetermined way, while behaving normally for clean inputs. This presents a significant challenge for defenders, who often lack visibility into the training data's provenance or the specific nature of the backdoor. PBP offers a practical, post-deployment solution, enabling organizations to purify compromised models and restore their reliability in classifying malicious software.
Background
▶ Watch: Introduction: Purifying backdoors in malware classifiers (0:00)
The proliferation of malware necessitates advanced detection mechanisms, with DNN-based classifiers emerging as a powerful tool. These models require vast datasets for training, often compiled from diverse sources across the internet. This practice, while enabling robust model performance, introduces a critical security vulnerability: the potential for data poisoning and backdoor attacks. An adversary can contribute a small, manipulated portion of the training data—even as little as 1%—to embed a backdoor into the resulting model.
A backdoor attack operates by associating a specific trigger (also referred to as a masking function or watermark function) with a target misclassification. For instance, an adversary might modify a legitimate malware sample X by adding a trigger to create X_poisoned. The goal is for the trained model M to incorrectly classify X_poisoned as benign, while still correctly identifying X as malware when the trigger is absent. Crucially, these attacks can be conducted in a clean-label attack fashion, meaning the adversary doesn't need to flip the original labels of the poisoned samples; they merely modify features. Such attacks have been shown to achieve near-perfect attack success rates (ASR), often reaching 100%, and can bypass existing backdoor defenses like Strip and NeuronClean.
Defending against these sophisticated backdoors is exceptionally challenging for several reasons:
- Unknown Target: Defenders typically don't know which specific malware families the adversary intends to protect or misclassify.
- Unknown Trigger: The exact nature of the trigger or masking function used by the adversary is unknown.
- Stealthy Modifications: Adversaries strive to minimize their fingerprint or modifications in the feature space, making detection difficult and allowing them to bypass filtering mechanisms.
- Post-Training Remediation: In many real-world scenarios, the model is already trained and potentially deployed before a backdoor is discovered or suspected, necessitating a post-training purification method.
Key Findings
▶ Watch: Challenges of defending against backdoor attacks (3:30)
The central insight driving the PBP method is the observation that within a backdoored DNN model, only a small, specific subset of neurons, termed backdoor neurons, are truly instrumental in activating the backdoor function. The remaining neurons, referred to as benign neurons, do not significantly contribute to the malicious misclassification behavior. This localized effect suggests that if these backdoor neurons can be accurately identified and their influence mitigated, the model's integrity can be restored.
PBP's empirical evaluations demonstrate its superior effectiveness and stability compared to existing baselines. It consistently reduces the attack success rate (ASR) from nearly 100% down to approximately 0% across various backdoor attack scenarios, including both universal and specific malware family targeting. Importantly, PBP achieves this while maintaining high clean accuracy, ensuring the model's utility for legitimate malware detection is preserved. The method operates under practical assumptions, requiring no prior knowledge about the adversary's attack strategy, the specific trigger used, or the targeted malware families. Furthermore, PBP's robustness was validated under varying adversary capabilities (e.g., different poisoning data rates) and defender resources (e.g., varying finetuning data sizes), consistently outperforming other finetuning methods. The research also highlighted PBP's broader applicability beyond malware classifiers, showing efficacy in computer vision domains.
Technical Deep Dive
▶ Watch: Core insight: Detecting and understanding backdoor neurons (5:00)
PBP's purification scheme is a two-step process designed to identify and neutralize the impact of backdoor neurons within a pre-trained, backdoored model. The method leverages the subtle statistical artifacts left by the backdoor training process.
Step 1: Detecting Potential Backdoor Neurons
The first objective is to pinpoint the subset of neurons that are disproportionately influenced by the backdoor. Since the defender has no prior knowledge of the attack, PBP relies on the batch normalization (batch norm) statistics of the backdoored model. During training, a model's batch norm layers accumulate statistics (mean and variance) from the input data batches. When a backdoored model is trained on a mix of clean and poisoned data, these statistics implicitly retain information about both distributions, including the subtle shifts introduced by the backdoor.
PBP introduces a randomized model with the exact same architecture as the backdoored model. This randomized model serves as a neutral baseline. The core idea is to perform batch norm statistic alignment: by iteratively adjusting the randomized model's parameters to match its batch norm statistics with those of the backdoored model, PBP can observe which neurons are most critical for this alignment process. Neurons that exhibit significant changes or are highly influential in aligning the batch norm statistics are identified as potential backdoor neurons. The intuition is that these neurons have learned to differentiate between clean and poisoned inputs in a way that contributes to the backdoor behavior, and thus their statistics are uniquely altered compared to a randomly initialized model.
Step 2: Activation Shift Finetuning / Mask Gradient
Once the set of potential backdoor neurons is identified, the second step focuses on mitigating their malicious influence through a specialized finetuning process. This step is termed Activation Shift Finetuning or Mask Gradient.
For the identified backdoor neurons (referred to as "red ones" in the presentation), PBP strategically reverses their gradient direction during specific finetuning steps. This counter-intuitive gradient update aims to pull these neurons away from the "backdoor function direction" they adopted during the initial training. By actively steering their activation patterns, PBP prevents them from responding to the trigger in the way intended by the adversary.
Conversely, for the normal or benign neurons, PBP applies standard gradient updates. This ensures that the model's overall utility—its ability to correctly classify clean malware and benign samples—is preserved. The finetuning is performed using a small subset of clean, unlabeled data, which is a practical assumption for defenders. This targeted manipulation of gradients allows PBP to precisely neutralize the backdoor while minimizing collateral damage to the model's legitimate functionality.
Evaluation Settings and Metrics
To evaluate PBP's performance, experiments were conducted under two distinct backdoor attack settings:
- Universal Backdoor: In this scenario, the adversary aims to misclassify any malware sample as benign, regardless of its family, as long as the trigger is present. This represents a broad, indiscriminate attack.
- Specific Malware Family Backdoor: Here, the adversary targets only a very specific malware family. The backdoor function is activated only if the sample belongs to the targeted family and contains the trigger; other malware families are unaffected. This represents a more surgical, targeted attack.
Two key metrics were used to quantify performance:
- Attack Success Rate (ASR): The percentage of poisoned malware samples that are incorrectly classified as benign. The goal for PBP is to minimize ASR as much as possible, ideally to 0%.
- Clean Accuracy: The percentage of clean (non-poisoned) samples (both benign and malware) that are correctly classified. The goal for PBP is to maintain high clean accuracy, preserving the model's primary utility.
The experimental results definitively showed that PBP was the only method capable of reliably reducing ASR across all tested scenarios, often bringing it down from over 90% to near 0%, while other baselines remained unstable and largely ineffective. The talk also highlighted how PBP successfully realigned the internal activation patterns of the model, making it treat poisoned malware and clean malware samples similarly again, a crucial visual confirmation of its effectiveness. During the Q&A, the speaker clarified that prior work on universal backdoor attacks, particularly with datasets like Amber, demonstrated transferability of triggers optimized on simpler models (e.g., SVM) to more complex ones (e.g., MLP), reinforcing the real-world threat PBP addresses.
Demo / Proof of Concept
▶ Watch: Purification Step 1: Detecting backdoor neurons (6:40)
While the presentation did not feature a live, interactive demonstration of the PBP tool or a specific proof-of-concept exploit, the efficacy of the PBP method was rigorously demonstrated through comprehensive experimental evaluation and quantitative analysis. The speaker presented detailed results comparing PBP against several baseline finetuning methods under various attack and defense scenarios.
These results, illustrated through tables and graphs, showed PBP's consistent ability to restore the model's correct classification behavior. Specifically, a plot of model activation after finetuning graphically depicted how PBP successfully forced the model to consider "malware group" and "poisoned malware group" as the same, unlike baseline methods which failed to correct the activation patterns. This visual evidence was corroborated by quantitative data, where PBP consistently reduced the attack success rate (ASR) to near zero across all considered scenarios, demonstrating its practical utility in neutralizing backdoor attacks in deep learning-based malware classifiers. The stability of PBP under varying conditions, such as different poisoning data rates (adversary's control over training data) and finetuning data sizes (defender's available clean data), further served as a robust proof of concept for its real-world applicability and resilience.
Defensive Implications
▶ Watch: Evaluation settings for different backdoor attack types (8:30)
PBP offers significant defensive implications for organizations relying on machine learning models for security-critical tasks, particularly malware detection. Its post-training nature means it can be applied to models that are already deployed or have completed their initial training, addressing a common scenario where backdoors might only be discovered or suspected after the fact.
Key implications for defenders include:
- Proactive Remediation for Deployed Models: PBP provides a practical mechanism to audit and purify ML models in production environments without requiring a complete retraining from scratch. This is crucial for maintaining the trustworthiness of security systems.
- No Prior Attack Knowledge Required: One of PBP's most compelling advantages is its independence from specific attack details. Defenders do not need to know the adversary's target malware family, the exact trigger mechanism, or the poisoning strategy. This significantly lowers the barrier to defense against unknown or zero-day backdoor attacks.
- Resource Efficiency: The method requires only a small subset of clean, unlabeled data for finetuning, making it resource-efficient. Collecting a small amount of verified clean data is often more feasible than re-curating an entire pristine training dataset.
- Enhanced Model Robustness: Integrating PBP into model lifecycle management can enhance the overall robustness of malware classifiers. It can serve as a final purification step before deployment or as part of a regular auditing process for critical models.
- Broad Applicability: While primarily demonstrated for malware classifiers, the speaker noted that PBP has shown efficacy in other computer vision domains. This suggests its potential as a general-purpose backdoor purification technique for various ML applications susceptible to data poisoning.
- Mitigation of Supply Chain Attacks: In scenarios where ML models or training datasets are sourced from third parties, PBP offers a defense against potential supply chain attacks where backdoors might be intentionally embedded by malicious actors.
By providing an effective, knowledge-agnostic, and post-training solution, PBP significantly strengthens the defensive posture against sophisticated backdoor attacks targeting deep learning systems.
Key Takeaways
- PBP is a novel post-training backdoor purification method for deep neural network (DNN) based malware classifiers.
- It effectively reduces attack success rate (ASR) from nearly 100% to 0% while maintaining high clean accuracy.
- PBP operates under practical assumptions, requiring no prior knowledge about the adversary's attack, specific triggers, or targeted malware families.
- The method works by identifying a small subset of "backdoor neurons" using batch normalization statistics and then neutralizing their influence through activation shift finetuning (reversing gradients for backdoor neurons).
- PBP demonstrated superior stability and effectiveness compared to baseline finetuning methods across various attack settings (universal and specific malware family backdoors) and defender resource levels.
- The approach has broader applicability beyond malware classification, showing promise in other computer vision domains.
About the Speaker(s)
Dung Thuy Nguyen is the presenter of this technical paper at the NDSS Symposium. Based on the content of the talk, Dung Thuy Nguyen's research focuses on the security and robustness of machine learning models, specifically addressing the critical challenge of backdoor attacks in deep learning systems. The work presented, "PBP: Post-training Backdoor Purification for Malware Classifiers," highlights expertise in developing practical and effective defenses against sophisticated adversarial machine learning threats, particularly those impacting security applications like malware detection. The speaker's ability to articulate complex technical concepts and experimental methodologies demonstrates a strong background in both theoretical and applied aspects of AI security.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate academic security research on backdoor purification for malware classifiers, presented at a credible venue. The core contribution — using batch norm statistic divergence between a randomized and backdoored model to localize backdoor neurons, then reversing gradients selectively — is technically coherent and addresses a real operational gap. Not groundbreaking enough to dominate a practitioner conference, but earns its place as solid published work.
Heather Calloway (CISO) — WEAK
Technically sound research on a real problem — backdoor poisoning of ML-based malware classifiers is a legitimate threat — but PBP never bridges from research result to institutional decision. The talk answers how to purify a backdoored model; it doesn't answer who owns that risk, when to act on it, or what organizational posture should change as a result.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025