Revisiting Concept Drift in Windows Malware Detection: Adaptation to Real Drifted Malware with Minimal Samples
Adrian Shuai Li (Purdue University)
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Malware
Overview
The relentless evolution of Windows malware poses a significant challenge for machine learning-based detection systems. Traditional models, trained on known samples, quickly become obsolete when attackers introduce new variants or families, a phenomenon known as concept drift. This talk, presented by Adrian Shuai Li from Purdue University, addresses the critical problem of building robust malware detection models that can rapidly adapt to these new threats with minimal labeled data. The core of their work introduces an innovative framework leveraging adversarial domain adaptation with Graph Neural Networks (GNNs) to learn invariant features across malware variants, enabling quick and efficient model updates.
Key moments
- 0:00 Introduction: Adapting to evolving malware with minimal samples
- 2:41 Core idea: Adaptive model learns common malware characteristics
- 3:07 Proposed framework for Windows malware detection
- 3:50 Introducing the Shift Adaptation Model architecture
- 4:53 Critique of evaluation methods and proposed clustering approach
- 6:00 Visualizing improved malware clustering with proposed method
- 7:30 Evaluation results: Domain adaptation outperforms baselines
- 8:30 Adaptation model's consistent accuracy across representations
Revisiting Concept Drift in Windows Malware Detection: Adaptation to Real Drifted Malware with Minimal Samples
Speakers: Adrian Shuai Li, PhD Student, Purdue University
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=32CnhtoSVKI
Overview
The relentless evolution of Windows malware poses a significant challenge for machine learning-based detection systems. Traditional models, trained on known samples, quickly become obsolete when attackers introduce new variants or families, a phenomenon known as concept drift. This talk, presented by Adrian Shuai Li from Purdue University, addresses the critical problem of building robust malware detection models that can rapidly adapt to these new threats with minimal labeled data. The core of their work introduces an innovative framework leveraging adversarial domain adaptation with Graph Neural Networks (GNNs) to learn invariant features across malware variants, enabling quick and efficient model updates.
The research not only proposes a novel Shift Adaptation Model but also critically examines existing evaluation methodologies, proposing a more statistically sound graph-based clustering approach to accurately assess model performance against truly distinct malware families. By demonstrating superior performance across various benchmarks and real-world datasets, the work offers a practical solution to the persistent challenge of keeping pace with an ever-changing threat landscape. This talk is highly relevant for security practitioners, researchers, and anyone involved in developing and deploying next-generation malware detection systems.
Background
▶ Watch: Introduction: Adapting to evolving malware with minimal samples (0:00)
The landscape of Windows malware detection is characterized by a constant arms race between attackers and defenders. Machine learning (ML) models, while powerful, are inherently susceptible to concept drift, where the statistical properties of the target variable (malware) change over time. When new malware samples or families emerge, a model trained on older data struggles to identify these novel threats, leading to a significant drop in detection accuracy. This necessitates updating or retraining the model.
Two primary approaches for updating ML models have been explored in prior work: cold-start learning and warm-start learning. Cold-start learning involves training a completely new model from scratch each time new labeled data becomes available. While ensuring the model is entirely up-to-date, this method is computationally expensive and requires a substantial amount of new labeled data. Warm-start learning, on the other hand, continues training an existing model with new samples. This is more efficient than cold-start learning but still often requires a considerable volume of labeled data to effectively integrate new knowledge and avoid catastrophic forgetting of previously learned features. Neither of these traditional strategies is optimized for scenarios where only a very limited number of new labeled samples are available, which is a common bottleneck in real-world security operations due to the high cost and specialized expertise required for malware analysis and labeling.
To mitigate the data labeling bottleneck, active learning has emerged as a promising technique. Active learning aims to reduce the amount of labeling effort by intelligently prioritizing "valuable" samples – those that would have the biggest impact on the model's performance if labeled. Common selection methods include uncertainty sampling (labeling samples where the model is most unsure), utilizing rejection thresholds, or employing contrastive learning to identify informative samples. While active learning can significantly reduce the volume of data needing to be labeled, the fundamental question of how to effectively use these few new samples to update the model remains. The challenge is not just which samples to label, but how to leverage them for rapid and robust model adaptation in the face of concept drift.
Furthermore, Adrian Li highlighted a critical flaw in common evaluation methods used in previous research, particularly the leave-one-out approach. This method often overestimates model accuracy by assuming concept drift where it might not genuinely exist. For instance, if malware families within a benchmark dataset share highly similar characteristics, a model trained on one family might perform well on another, not because it adapted to drift, but because the "drifted" family wasn't truly distinct. This can lead to an inflated perception of a model's adaptability. To address this, the research proposes a more rigorous graph-based clustering approach to generate statistically distinct malware clusters, ensuring that evaluation accurately reflects a model's ability to handle genuinely novel threats.
Key Findings
▶ Watch: Proposed framework for Windows malware detection (3:07)
The research presents several key findings and contributions that significantly advance the state of malware detection amidst concept drift:
- Novel Shift Adaptation Model: The core contribution is the development of a Shift Adaptation Model that employs adversarial domain adaptation with Graph Neural Networks (GNNs). This model is specifically designed to learn invariant features that persist across different malware variants and families, enabling it to identify new threats even when they don't perfectly match previously seen samples. This is the first application of adversarial domain adaptation with GNNs for malware detection.
- Superior Performance with Limited Labels: The Shift Adaptation Model consistently outperforms state-of-the-art retraining methods (cold-start and warm-start learning) in both settings (original and cluster-labeled datasets). Across various evaluation tasks, it exhibited the smallest accuracy drop, averaging just 0.5% on unseen malware families. Crucially, on real-world datasets, their approach matched the accuracy of a fully supervised setting with a five-fold reduction in labeling effort.
- Enhanced Malware Representation: The research demonstrates that combining their adaptation approach with graph-based representations (specifically Control Flow Graphs) yields the best overall detection results compared to other popular representations like image-based or content-based (assembly code features). This highlights the efficacy of structural program analysis for robust malware identification.
- Improved Evaluation Methodology: A significant contribution is the proposed graph-based clustering approach to generate statistically distinct malware clusters. This method addresses the limitations of traditional evaluation setups (like leave-one-out on datasets such as Big 15) which can overestimate model performance by failing to account for high similarity between "distinct" families. Their clustering technique produced significantly more distinct malware clusters, leading to a more accurate assessment of model adaptability.
- Robustness Across Diverse Scenarios: The model's effectiveness was validated across three distinct tasks with increasing complexity: adapting to unseen families in a benchmark dataset, adapting to real-world malware drift over several months, and multi-family classification. In multi-family classification using the MalDrift benchmark, the method improved family-level classification by 9% to 14% over baselines with only 10 new samples per family. Furthermore, it demonstrated the ability to accurately detect entirely new, unknown families in an open domain adaptation scenario, preventing evasion.
Technical Deep Dive
▶ Watch: Critique of evaluation methods and proposed clustering approach (4:53)
The proposed framework for adaptive Windows malware detection is built upon a sophisticated pipeline that begins with binary analysis and culminates in a specialized deep learning model.
The process starts by taking Windows malware binaries as input. The first crucial step involves extracting Control Flow Graphs (CFGs) from these binaries. A CFG is a representation of all paths that might be traversed through a program during its execution. Nodes in the CFG represent basic blocks (sequences of instructions with a single entry and exit point), and edges represent possible transfers of control between these blocks. CFGs capture the structural and behavioral aspects of malware, which are often more resilient to obfuscation techniques than raw byte sequences or simple instruction counts.
Once the CFGs are extracted, the raw instructions within each basic block are converted into feature vectors. For this purpose, the researchers leverage a pre-trained assembly language model called PalmTree. PalmTree is designed to understand the semantics of assembly code, translating the low-level instructions into a high-dimensional numerical representation that can be processed by machine learning models. This step is vital for abstracting away syntactic variations while retaining meaningful operational characteristics of the code.
The core contribution of this work lies in the Shift Adaptation Model. This model is the first of its kind to apply adversarial domain adaptation in conjunction with Graph Neural Networks (GNNs) for malware detection. The architecture of the Shift Adaptation Model comprises three distinct neural networks:
- Generator (or Feature Extractor): This GNN takes the feature vectors derived from the CFGs as input. Its primary role is to learn a robust, high-level representation of the malware. The goal is for this representation to capture the common characteristics shared across a broad range of malware, regardless of their specific family or variant.
- Discriminator: This network is trained to distinguish between existing (source domain) and drifted (target domain) malware samples based on the representations generated by the Generator.
- Classifier: This network is trained to classify the malware samples (e.g., benign or malicious, or specific family types) based on the invariant features learned by the Generator.
The training process is orchestrated as an adversarial training task, specifically a minimax game. In this game:
- The Generator attempts to produce feature representations that are indistinguishable to the Discriminator, meaning it tries to "fool" the Discriminator into thinking that features from drifted malware are similar to those from existing malware. The Generator aims to learn features that are domain-invariant.
- The Discriminator simultaneously tries to accurately identify whether a given feature representation comes from the source domain (existing malware) or the target domain (drifted malware).
- The Classifier is trained on these domain-invariant features to perform the actual malware detection task.
This adversarial setup forces the Generator to learn representations that are independent of the malware's origin (i.e., whether it's an old or new variant). By effectively removing domain-specific characteristics, the model focuses on the fundamental, invariant traits of malware, making it highly adaptive to concept drift with minimal new labeled samples.
Beyond the model architecture, the researchers critically re-evaluated existing evaluation methodologies. They highlighted that the popular leave-one-out approach often used with datasets like Big 15 can be misleading. They demonstrated that malware from different families in such datasets might exhibit highly similar characteristics, leading to an overestimation of a detector's accuracy if family labels are directly used for evaluation. To counter this, they proposed a novel graph-based clustering approach. This method groups malware samples based on the statistical similarity of their graph feature vectors rather than their original, potentially ambiguous, family labels. By visualizing the Big 15 dataset, they showed that their clustering approach produced significantly more distinct and well-separated malware clusters, providing a more rigorous and realistic basis for evaluating adaptation capabilities.
The evaluation also explored different malware representations:
- Graph-based representation: Using CFGs and PalmTree-derived features, as described above. This was found to yield the best overall performance when combined with their adaptation model.
- Image-based representation: Converting raw binary data or disassembled code into grayscale images and using CNNs for detection. This is often the most efficient in terms of processing time but typically has lower detection performance, especially in cold-start scenarios.
- Content-based representation: Extracting a large set of handcrafted features (e.g., 965 different features) directly from assembly code. While potentially rich, this is the most time-consuming representation to extract.
The integration of their domain adaptation components into models designed for these other representations consistently showed improved accuracy, underscoring the general applicability and effectiveness of their adaptation strategy across different feature types.
Demo / Proof of Concept
▶ Watch: Visualizing improved malware clustering with proposed method (6:00)
The talk presented a comprehensive evaluation of the Shift Adaptation Model across various settings, demonstrating its efficacy through several proof-of-concept tasks and comparisons.
First, the researchers addressed the issue of evaluation methodology by demonstrating their graph-based clustering approach. They visualized the feature vectors of the widely used Big 15 malware detection benchmark. On the left, the visualization showed the original family labels, which often overlapped significantly, indicating that different "families" might not be statistically distinct. On the right, the same feature vectors were visualized with their newly generated cluster labels, clearly showing significantly more distinct and better-separated malware clusters. This visual evidence underscored the importance of their improved clustering for accurate performance evaluation.
Next, they conducted a leave-one-out evaluation setup using the Big 15 dataset. They simulated concept drift by choosing one malware family as "unseen" (drifted) and using the rest as "existing" families for training. They varied the size of the training set for the unseen family from 20 to 500 samples to assess the model's performance with limited labels. Their Shift Adaptation Model was compared against several state-of-the-art retraining methods, including cold-start learning (lowest performance for drifted samples) and warm-start learning. The results consistently showed that their domain adaptation method achieved better performance than both warm-start and cold-start learning in both settings (with original labels and their cluster labels). Their method experienced the smallest accuracy drop, averaging just 0.5% on unseen families.
The evaluation also explored the impact of different malware representations. They compared graph-based, image-based, and content-based (assembly code features) representations. For each representation, they evaluated state-of-the-art detectors and, critically, integrated their domain adaptation component into them. The results indicated that their adaptation model consistently achieved the highest accuracy across all feature representations. Specifically, combining their adaptation approach with the graph-based representation yielded the best overall results, highlighting the strength of CFGs for robust malware analysis.
To simulate real-world deployment, the researchers collected a new dataset from MalwareBazaar spanning from March 2024 to September 2024, ensuring at least 81 different malware families were collected each month. They designed two model update tasks:
- Task 1: Used data from March, April, and May as existing malware, July as drifted samples, and tested on August data.
- Task 2: Used data from March to May as existing data, August as drifted data, and tested on September data.
In both real-world tasks, their approach outperformed all baselines. A particularly striking result was that their method could match the accuracy of a fully supervised setting (where all data is labeled) with a five-fold reduction in the labeling effort.
Further experiments were conducted on the MalDrift benchmark for multi-family classification. In a closed-set domain adaptation scenario (where pre-drift and post-drift samples come from the same families), their method improved family-level classification by 9% to 14% over baselines with just 10 new samples per family.
Finally, the researchers extended their approach to an open domain adaptation scenario, where the post-drift data could include entirely new malware families not seen before. Their evaluation showed that their approach could accurately detect these new samples as "unknown," preventing them from evading detection, a crucial capability for real-world threat intelligence.
Defensive Implications
▶ Watch: Adaptation model's consistent accuracy across representations (8:30)
The findings presented in this talk offer significant practical implications for cybersecurity defenders and organizations struggling with the rapid evolution of malware.
Firstly, the Shift Adaptation Model provides a robust mechanism for rapidly updating malware detection systems with minimal new labeled data. This directly addresses the critical bottleneck of acquiring and labeling new malware samples, which is often a time-consuming and resource-intensive process for security teams. By reducing the need for extensive labeling by as much as five times while maintaining accuracy, organizations can deploy more agile and responsive defenses.
Secondly, the emphasis on learning invariant features across malware variants means that detection systems built on this framework will be inherently more resilient to polymorphic and metamorphic malware. Attackers frequently modify existing malware to create new variants that bypass signature-based or even traditional ML detectors. By focusing on the fundamental structural and behavioral characteristics captured by Control Flow Graphs and domain adaptation, defenders can achieve more generalized detection capabilities.
Thirdly, the proposed graph-based clustering approach offers a more accurate way to understand and evaluate malware families. Security analysts can leverage this methodology to gain deeper insights into the true distinctiveness of malware clusters, rather than relying solely on potentially misleading family labels. This improved understanding can inform better threat intelligence, more targeted defensive strategies, and more realistic assessments of detection model performance.
Fourthly, the demonstrated effectiveness in open domain adaptation scenarios is crucial. The ability to accurately classify new, previously unseen malware families as "unknown" rather than misclassifying them or allowing them to evade detection is a significant advantage. This capability helps security operations centers (SOCs) prioritize emerging threats and prevents novel attacks from slipping through the cracks.
Finally, the research reinforces the value of graph-based representations for malware analysis. Defenders should consider integrating CFG-based feature extraction into their analysis pipelines, as it has been shown to provide superior performance when combined with advanced adaptation techniques. This suggests a move away from simpler representations (like raw bytes or image-based features) for critical detection tasks where robustness against evolution is paramount. In essence, this work empowers defenders to build more intelligent, adaptable, and efficient malware detection systems that can better keep pace with the ever-changing threat landscape.
Key Takeaways
- Adaptive Malware Detection: The Shift Adaptation Model, leveraging adversarial domain adaptation with Graph Neural Networks (GNNs), provides a novel solution for detecting evolving Windows malware with minimal new labeled data.
- Reduced Labeling Effort: The proposed method achieves accuracy comparable to fully supervised models with a five-fold reduction in the amount of labeled data required for adaptation.
- Invariant Feature Learning: The model learns invariant features that persist across malware variants, making it robust against concept drift and enabling detection of new, unseen threats.
- Superior Graph Representation: Combining domain adaptation with Control Flow Graph (CFG)-based representations yields the best overall detection performance compared to image-based or content-based features.
- Improved Evaluation: A novel graph-based clustering approach is introduced to generate statistically distinct malware clusters, providing a more accurate and rigorous method for evaluating model performance against true concept drift.
- Real-World Efficacy: The approach outperforms baselines on real-world malware datasets collected from MalwareBazaar and effectively handles multi-family classification, even detecting entirely new, unknown families in open domain scenarios.
About the Speaker(s)
Adrian Shuai Li is a PhD Student at Purdue University, where he conducted this research. His work focuses on practical problems in cybersecurity, particularly in developing machine learning models that can adapt quickly to evolving threats. He collaborated on this project with Arum Anger, Ashish Kundu from Cisco Research, and Elisa Bertino. Adrian mentioned during the talk that he will be on the job market next year, seeking industrial positions related to his expertise.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Technically legitimate academic work on concept drift adaptation for malware detection — adversarial domain adaptation with GNNs plus a methodological critique of leave-one-out evaluation on Big 15. Competent research that advances a real problem, but it lands squarely in the 'solid conference paper' tier: incremental rather than transformative, and the writeup reads more like a paper abstract than a talk review.
Heather Calloway (CISO) — WEAK
Technically credible research with a real operational problem at its core — concept drift in ML-based malware detection is a genuine pain point for security programs. But the talk is built entirely for a research audience and never closes the gap to the people who actually operate detection systems, fund them, or make deployment decisions.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025