Multi-Instance Adversarial Attack on GNN-Based Malicious Domain Detection
Mahmoud Nazzal, Issa Khalil, Abdallah Khreishah, NhatHai Phan, Yao Ma
IEEE Symposium on Security and Privacy 2024 · Day 1 · Continental Ballroom 5
Overview
This talk, presented by Mahmoud Nazzal, delves into the critical security vulnerabilities of Graph Neural Networks (GNNs) when applied to security-critical tasks, specifically malicious domain detection (MDD). While GNNs have demonstrated state-of-the-art performance across various domains by effectively combining local entity information with relational data, their inherent susceptibility to adversarial attacks poses significant challenges in sensitive applications like cybersecurity. The research introduces MAA (Multi-Instance Adversarial Attack), a novel, stealthy, and practical black-box attack designed to evade the detection of multiple malicious domains simultaneously within a GNN-based MDD system.

Key moments
- 0:00 GNNs, Malicious Domain Detection, and adversarial vulnerabilities
- 2:00 How GNNs are used for Malicious Domain Detection
- 3:20 Introducing a new stealthy, black-box multi-instance attack
- 4:20 Adversary's goals, knowledge, and capabilities defined
- 5:40 Limitations of current GNN attacks in MDD context
- 7:00 Mathematical formulation for collective evasion of detection
- 8:00 Step-by-step process of the DM attack algorithm
Multi-Instance Adversarial Attack on GNN-Based Malicious Domain Detection
Speakers: Mahmoud Nazzal, Researcher, New Jersey Institute of Technology; Issa Khalil, Researcher, QC at Hamad bin Khalifa University; Abdallah Khreishah, Professor, New Jersey Institute of Technology; NhatHai Phan, Professor, New Jersey Institute of Technology; Yao Ma, Professor, New Jersey Institute of Technology
Conference: IEEE S&P
YouTube: https://www.youtube.com/watch?v=s5iRhJsu4tw
Overview
This talk, presented by Mahmoud Nazzal, delves into the critical security vulnerabilities of Graph Neural Networks (GNNs) when applied to security-critical tasks, specifically malicious domain detection (MDD). While GNNs have demonstrated state-of-the-art performance across various domains by effectively combining local entity information with relational data, their inherent susceptibility to adversarial attacks poses significant challenges in sensitive applications like cybersecurity. The research introduces MAA (Multi-Instance Adversarial Attack), a novel, stealthy, and practical black-box attack designed to evade the detection of multiple malicious domains simultaneously within a GNN-based MDD system.
The core problem addressed is the insufficiency of existing adversarial attack techniques against GNNs when applied to the multi-instance evasion problem in MDD. MAA is distinct because it targets an entire subgraph of adversary-controlled nodes, aiming for collective evasion while minimizing impact on non-adversarial nodes. This work is crucial for understanding the limitations of current GNN-based security solutions and for developing more robust defense mechanisms against sophisticated adversaries who control multiple internet domains. The findings highlight a significant security threat that requires immediate attention from both researchers and practitioners deploying GNNs in cybersecurity.
Background
▶ Watch: GNNs, Malicious Domain Detection, and adversarial vulnerabilities (0:00)
The landscape of cybersecurity threats is constantly evolving, with malicious domains serving as a foundational component for various cyberattacks, including phishing, malware distribution, and command-and-control infrastructure. Malicious domain detection (MDD) systems are designed to identify and flag these domains, often by inferring their maliciousness from public data, primarily DNS logs, and sometimes enterprise-level data. At its heart, MDD often operates on a "guilt by association" principle, constructing a domain maliciousness graph (DMG) where internet domains, IP addresses, clients, and other entities are represented as nodes, and their relationships (e.g., domain-to-IP resolution, client-to-domain queries) as edges.
Traditionally, various machine learning models have been applied to MDD. However, recent advancements have shown Graph Neural Networks (GNNs) to be particularly promising due to their ability to naturally model the relational data inherent in DMGs. GNNs combine the power of deep learning with graph structures, allowing them to learn representations that capture both the features of individual entities (nodes) and the patterns of their connections (edges). A typical GNN-based MDD process involves:
- Collecting DNS logs and other relevant data.
- Constructing a DMG based on a predefined schema (e.g., combining domain, IP, and client nodes with various edge types).
- Training a GNN model to classify domains as benign or malicious.
- Utilizing the trained model at inference time to predict the maliciousness of new, unknown domains based on their surrounding computational graph.
Despite their impressive performance, GNNs, like many other machine learning models, are inherently vulnerable to adversarial attacks. These attacks involve making small, often imperceptible, perturbations to input data that cause the model to misclassify. Existing adversarial attacks on GNNs typically operate at different granularities:
- Node-level attacks: Targeting individual nodes.
- Edge-level attacks: Manipulating connections between nodes.
- Graph-level attacks: Modifying the entire graph structure.
However, the speakers emphasize the insufficiency of existing adversarial attacks when applied to the specific context of GNN-based MDD, especially for a multi-instance evasion scenario. Prior approaches, which might involve repeatedly applying single-node attacks, exhibit diminishing effectiveness as the number of targeted malicious domains increases. This is because such methods often fail to account for the collective impact and relational dependencies within a subgraph of adversarial entities. An adversary controlling multiple malicious domains needs a unified strategy to evade detection for all of them simultaneously and stealthily, a gap that MAA aims to fill.
Key Findings
▶ Watch: Introducing a new stealthy, black-box multi-instance attack (3:20)
The research presents several critical findings that advance the understanding of GNN security in the context of malicious domain detection:
- Novel Multi-Instance Adversarial Attack (MAA): The primary contribution is the introduction of MAA, a new adversarial attack designed specifically for subgraph-level manipulation. Unlike prior work focusing on individual nodes or arbitrary graph alterations, MAA targets a collection of adversarial domains (a subgraph) with the goal of collectively evading detection. This attack is characterized by its stealthiness and practicality, operating as a black-box attack that requires no prior knowledge of the target GNN model's architecture or parameters, nor the complete structure of the domain maliciousness graph (DMG) beyond the adversary's own subgraph.
- Theoretical and Empirical Insufficiency of Existing Attacks: The talk provides both theoretical justification and empirical evidence demonstrating that current adversarial attack methods are inadequate for GNN-based MDD in a multi-instance context. When existing approaches are applied repeatedly to individual adversarial nodes, their attack effectiveness (e.g., evasion rate) significantly diminishes as the number of targeted domains increases. This highlights a critical vulnerability that MAA is designed to exploit, as adversaries typically control multiple domains.
- Two-Objective Optimization Problem Formulation: MAA frames the adversarial attack as a two-objective optimization problem. This novel formulation is key to its effectiveness:
- Objective 1 (Individual Evasion): Ensure that each adversarial domain node bypasses detection by the GNN model.
- Objective 2 (Collective Reinforcement): Maximize the positive impact of perturbations on neighboring adversarial nodes, ensuring that the changes made to one malicious domain contribute to the evasion of others within the adversary's subgraph. This objective directly addresses the shortcomings of single-instance attacks by leveraging the relational information within the subgraph.
- Practical Perturbation Implementation: The research demonstrates that the calculated adversarial perturbations can be implemented practically in real-world settings. These perturbations manifest as simple name edits (manipulating domain names) and/or IP resolution changes (altering the IP addresses a domain resolves to). Such changes are easily conducted by any domain owner, making MAA a highly realistic threat. These manipulations directly affect both the local features of domain nodes and the edges connecting them to other entities (like IP addresses) in the DMG.
- Superior Performance Against State-of-the-Art Baselines: Through experiments on a real-world dataset combining public and enterprise data, MAA consistently outperforms existing feature-focused and edge-focused adversarial attacks. Metrics such as Attack Success Rate (ASR) and Negative Flip Rate (NFR) confirm MAA's superior ability to evade detection for a large proportion of adversarial domains while maintaining a low side effect (i.e., minimal impact on non-adversarial nodes or unintended classifications). Combining both feature and edge perturbations further enhances MAA's success.
These findings collectively reveal a significant security gap in GNN-based MDD systems and provide a robust, practical attack methodology that adversaries could leverage.
Technical Deep Dive
▶ Watch: Adversary's goals, knowledge, and capabilities defined (4:20)
The technical core of MAA lies in its sophisticated formulation as a two-objective optimization problem and its practical implementation through a surrogate model. The entire process is designed to be a black-box attack, meaning the adversary has no direct knowledge of the target GNN model's internal architecture, parameters, or the full structure of the DMG.
Problem Formulation: Two-Objective Optimization
The adversary's goal is to protect or evade the detection of multiple nodes (domains) it controls, while having a minimal effect on non-adversarial nodes. To achieve this collectively and saliently, MAA casts the problem as a two-objective optimization problem:
- Objective 1: Individual Node Evasion: For every adversarial domain node, the perturbations introduced by the attack should directly cause that particular node to be misclassified by the GNN-based MDD model (i.e., classified as benign instead of malicious). This focuses on the local impact of the perturbation on the targeted node's classification decision.
- Objective 2: Collective Neighbor Evasion: The impact of these perturbations should also positively contribute to the evasion of neighboring adversarial nodes within the adversary's subgraph. This objective leverages the relational nature of GNNs, ensuring that changes propagate beneficially across the adversary's controlled infrastructure, rather than just isolated nodes.
The researchers derived a closed-form solution for the optimized perturbation at these adversarial nodes. This solution is described as a weighted sum of two components, each corresponding to one of the aforementioned objectives. While the specific mathematical derivation is not detailed in the transcript, the conceptual implication is that the optimal perturbation balances the need for individual evasion with the desire for collective reinforcement across the adversary's subgraph. This weighted sum implicitly considers both the local features of the domain nodes and their connections to other entities.
The MAA Algorithm (D-Mina)
The proposed multi-instance adversarial attack algorithm, referred to as D-Mina (likely an acronym for "Domain Multi-Instance Adversarial Attack"), operates in a black-box setting and requires the availability of a surrogate model.
- Surrogate Model Training:
- The process begins with a set of training data representing a partial DMG.
- The adversary queries the black-box target MDD model with samples from this data. This step is crucial for understanding the target model's decision boundaries without knowing its internals.
- The target model's responses (labels for some nodes) are then used to train a surrogate GNN model. This surrogate model is an approximation of the target model's behavior and is what the adversary will directly manipulate to craft attacks. The quality of the surrogate model is critical for the attack's success.
- Adversary's Subgraph Estimation:
- The adversary constructs an estimate of its own subgraph. This estimate includes the domains it controls, their local features, and the expected edge types connecting them (e.g., to associated IP addresses or other domains). The adversary does not need to know the entire DMG.
- Perturbation Optimization:
- Using the trained surrogate model and the D-Mina algorithm, the adversary optimizes the perturbations to be applied to its nodes (domains) and edges within its estimated subgraph. This optimization process leverages the two-objective formulation to calculate the optimal changes that will lead to evasion.
- Practical Implementation of Perturbations:
- Once optimized, these perturbations are implemented practically. The transcript highlights two primary methods:
- Name Edits: The adversary can manipulate the domain names themselves. This could involve subtle changes, adding or removing subdomains, or using homoglyphs, which directly alters the local features of the domain node.
- IP Resolution Changes: The adversary can change the IP addresses to which its domains resolve. This directly impacts the edges connecting domain nodes to IP nodes in the DMG. For example, resolving to a benign IP address or a shared IP with many other benign domains could make the malicious domain appear less suspicious.
Adversary and Target MDD Scopes
The attack's practicality is underpinned by the interaction between the adversary's scope and the target MDD entity's scope:
- Adversary's Scope: The adversary injects the calculated perturbations into the DNS server. This is within the adversary's control as a domain owner.
- Target MDD Scope: The defense entity (the GNN-based MDD system) simultaneously reads DNS logs from the same DNS server. Without knowing the full DMG or the adversary's intent, the MDD system constructs its DMG based on these logs. Crucially, the perturbations implemented by the adversary will appear naturally in these logs and subsequently in the DMG constructed by the MDD system. These injected, stealthy changes then cause the actual target GNN model to misclassify the malicious domains as benign.
This intricate interplay demonstrates how MAA can effectively bypass detection in a real-world scenario, making it a potent threat against GNN-based MDD systems.
Demo / Proof of Concept
▶ Watch: Mathematical formulation for collective evasion of detection (7:00)
The talk presents its experimental results as a robust demonstration and proof of concept for the MAA algorithm's effectiveness. The experiments were conducted using a realistic setup designed to mimic real-world deployment of GNN-based MDD systems.
Experiment Setup
- Target Model: The chosen target MDD model was an hGNN (heterogeneous GNN) approach. This model represents the state-of-the-art in GNN-based MDD, leveraging both public and enterprise data for its predictions. It is noted to have state-of-the-art performance and has even been commercialized, underscoring the practical relevance of attacking such a system.
- Dataset: A real-world dataset was utilized, exhibiting representative scale and statistics. This dataset combined publicly available DNS data with proprietary enterprise data, providing a comprehensive and challenging environment for the attack.
- DMG Schema: The experiments adopted the graph schema proposed by the authors of the hGNN paper, ensuring consistency with established MDD methodologies. This schema typically combines domain, IP, and client nodes with various edge types.
- Baselines: MAA was compared against two primary categories of existing adversarial attacks on GNNs, representing the state-of-the-art:
- Node feature-focused attacks: These attacks primarily manipulate the attributes or features of individual nodes.
- Edge feature-focused attacks: These attacks primarily manipulate the connections between nodes.
Metrics
Two key metrics were employed to evaluate the attack's performance:
- Attack Success Rate (ASR): This metric quantifies the ability of the attack to successfully bypass the detection of malicious domains. A higher ASR indicates a more effective attack.
- Negative Flip Rate (NFR): This metric characterizes the "side effect" of the attack. It measures the rate at which benign nodes are inadvertently flipped to malicious, or existing malicious nodes that were correctly classified are flipped to benign, but were not intended targets. A lower NFR indicates a more stealthy and precise attack, minimizing collateral damage.
Experimental Results
The results consistently demonstrated MAA's superior performance:
- Feature Perturbation Alone: When only local features of the adversary's domains were perturbed, MAA achieved evasion for the vast majority of adversarial domains. Critically, this was accomplished with a negligibly small negative impact (low NFR) on other nodes, including non-adversarial ones. The success rate (ASR) generally increased as more adversarial nodes were included in the attack, contrary to the diminishing returns observed in existing methods.
- Combined Feature and Edge Perturbations: When MAA simultaneously manipulated both local features (via name edits) and edges (via IP resolution changes) across adversarial nodes, the attack exhibited an even stronger performance. The combined approach resulted in higher ASR and lower NFR compared to feature-only attacks. This indicates that leveraging both types of perturbations synergistically amplifies the attack's effectiveness and stealth.
- Baseline Comparison: MAA consistently proved superior to both feature-centered and edge-centered baseline attacks across both ASR and NFR metrics. This empirical evidence validates the unique multi-instance, subgraph-level approach of MAA as a more potent threat than prior, more generalized GNN attack methodologies when applied to MDD.
The experiments also briefly mentioned that MAA exhibited high scalability with reasonable cost, a finding detailed further in the full paper. This ensures that the attack is not only effective but also practical for adversaries to deploy against large-scale MDD systems. These results effectively serve as a proof of concept, demonstrating that MAA reveals a significant, practical security threat to GNN-based malicious domain detection.
Defensive Implications
▶ Watch: Step-by-step process of the DM attack algorithm (8:00)
The introduction of MAA reveals a critical and practical security threat to GNN-based malicious domain detection systems, necessitating a re-evaluation of current defense strategies. The black-box, multi-instance, and stealthy nature of MAA means that traditional defenses against single-instance or white-box attacks may be insufficient. Defenders of GNN-based MDD systems should consider the following implications:
- Increased Robustness Against Black-Box Adversaries: GNN-based MDD models need to be inherently more robust to adversaries who have no direct knowledge of the model's internals but can query it and observe its outputs. This suggests a need for adversarial training techniques specifically tailored for GNNs in a black-box setting, where the model is trained on adversarially perturbed inputs to improve its resilience.
- Focus on Multi-Instance and Subgraph-Level Defense: Current defenses often focus on identifying individual malicious entities. MAA demonstrates that adversaries can coordinate their attacks across multiple domains within a subgraph. Defenders should develop methods to detect coordinated malicious activity and anomalous subgraph patterns rather than solely focusing on isolated nodes. This could involve graph anomaly detection techniques that specifically look for unusual relational structures or feature changes across a cluster of related domains.
- Enhanced Feature and Edge Engineering: The attack leverages simple domain name edits and IP resolution changes. Defenders should investigate more resilient ways to represent domain features and relationships within the DMG. This might include:
- Time-series analysis of DNS records: Monitoring changes over time for suspicious patterns rather than just static snapshots.
- Reputation systems for IP addresses: Incorporating more sophisticated IP reputation scores that are harder to manipulate.
- Robust domain name feature extraction: Using features less susceptible to minor textual perturbations (e.g., character n-grams, entropy, or features derived from passive DNS data that reflect long-term behavior).
- Multi-source validation: Corroborating DNS information with other independent data sources to detect inconsistencies.
- Proactive Monitoring of DNS Infrastructure: Since the attack relies on manipulating DNS records, increased scrutiny of DNS resolution patterns, rapid changes in IP associations, and suspicious domain registration activities (e.g., newly registered domains that quickly gain high-volume traffic or resolve to multiple IPs) could serve as early warning signs.
- Uncertainty-Aware GNNs: Developing GNN models that can express uncertainty in their predictions could help flag domains that fall into ambiguous regions due to adversarial perturbations, prompting human review or additional security measures.
- Continuous Evaluation of MDD Robustness: Security teams deploying GNN-based MDD solutions should regularly test their systems against sophisticated adversarial attacks like MAA. This continuous red-teaming approach is essential to identify and patch vulnerabilities before they are exploited in the wild.
In essence, MAA highlights that the "guilt by association" principle, while powerful, can be exploited by adversaries who strategically manage their "associations." Defenses must evolve to detect these deliberate manipulations of relational information within the domain ecosystem.
Key Takeaways
- GNNs are vulnerable to practical, black-box adversarial attacks in security-critical applications like MDD. Despite their advanced capabilities, their reliance on structured data makes them susceptible to manipulation of both node features and graph edges.
- Existing single-instance adversarial attacks are insufficient for multi-instance evasion in GNN-based MDD. Their effectiveness diminishes as the number of targeted malicious domains increases, underscoring the need for more sophisticated, coordinated attack strategies.
- MAA (Multi-Instance Adversarial Attack) is a novel, stealthy, and practical black-box attack. It successfully evades detection for multiple malicious domains simultaneously by optimizing perturbations at the subgraph level.
- MAA employs a two-objective optimization problem. This formulation balances individual domain evasion with the collective reinforcement of evasion across neighboring adversarial nodes within a controlled subgraph.
- The attack can be implemented practically through simple name edits and IP resolution changes. These manipulations are easily performed by domain owners and effectively alter both local node features and inter-node edges within the domain maliciousness graph.
- MAA significantly outperforms state-of-the-art feature-focused and edge-focused attacks. Experiments demonstrate its superior Attack Success Rate (ASR) and lower Negative Flip Rate (NFR), highlighting its effectiveness and stealth against real-world GNN-based MDD systems.
About the Speaker(s)
The primary presenter for this talk was Mahmoud Nazzal, a researcher affiliated with the New Jersey Institute of Technology and the QC at Hamad bin Khalifa University in Doha, Qatar. His work focuses on the intersection of machine learning, particularly Graph Neural Networks, and cybersecurity, with an emphasis on adversarial robustness. The research was a collaborative effort involving several other distinguished individuals: Issa Khalil from QC at Hamad bin Khalifa University, and Abdallah Khreishah, NhatHai Phan, and Yao Ma, all professors or researchers from the New Jersey Institute of Technology. Their collective expertise spans machine learning, deep learning, graph neural networks, and their applications in security-critical domains.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This research introduces MAA, a novel, stealthy black-box attack that simultaneously evades detection of multiple malicious domains in GNN-based MDD systems. By framing the problem as a two-objective optimization at the subgraph level, MAA demonstrates a critical, practical vulnerability in state-of-the-art defenses that requires immediate attention from practitioners.
Heather Calloway (CISO) — STRONG ACCEPT
This research uncovers a significant and practical vulnerability in GNN-based malicious domain detection systems, demonstrating how coordinated adversaries can evade detection. It necessitates an immediate re-evaluation of organizational reliance on these controls and demands a proactive approach to adversarial resilience in ML-driven security.
→ Top-rated talks at IEEE Symposium on Security and Privacy 2024