Sometimes Simpler is Better: A Comprehensive Analysis of State-of-the-Art Provenance-Based Intrusion Detection Systems
Tristan Bilot (PhD student)
34th USENIX Security Symposium (USENIX Security '25) · Day 3 · Network Security 4: Internet and Beyond
Overview
In the ever-evolving landscape of cybersecurity, detecting sophisticated attacks requires robust and intelligent systems. Provenance-based Intrusion Detection Systems (PIDS) have emerged as a promising approach, leveraging system-level causality graphs to trace malicious activities. However, the complexity of state-of-the-art PIDS, often relying on Graph Neural Networks (GNNs) and sophisticated anomaly detection techniques, raises questions about their practical deployability and actual effectiveness. This talk, presented by Tristan Bilot, a PhD student at University Pra and visiting student at UBC, delivers a comprehensive analysis of eight recent, top-tier PIDS, revealing critical shortcomings in their design, evaluation, and practical utility.

Key moments
- 0:00 Introduction to provenance-based IDS and project goals
- 2:40 Why standard metrics are ill-suited for cyber security
- 5:47 Proposing new Attack Detection Precision (ADP) metric
- 6:20 Introducing VLOX, a simple word2vec baseline model
- 6:58 Surprising result: VLOX outperforms complex state-of-the-art PIDS
- 8:05 VLOX's real-time detection and low resource usage
- 8:50 Discussion on system complexity and textual attributes' sufficiency
- 9:40 The role of GNNs for graph structure-based attack detection
Sometimes Simpler is Better: A Comprehensive Analysis of State-of-the-Art Provenance-Based Intrusion Detection Systems
Speakers: Tristan Bilot, PhD Student, University Pra & Visiting Student, UBC
Conference: USENIX Security
YouTube: https://www.youtube.com/watch?v=Or_iAucWqT4
Overview
In the ever-evolving landscape of cybersecurity, detecting sophisticated attacks requires robust and intelligent systems. Provenance-based Intrusion Detection Systems (PIDS) have emerged as a promising approach, leveraging system-level causality graphs to trace malicious activities. However, the complexity of state-of-the-art PIDS, often relying on Graph Neural Networks (GNNs) and sophisticated anomaly detection techniques, raises questions about their practical deployability and actual effectiveness. This talk, presented by Tristan Bilot, a PhD student at University Pra and visiting student at UBC, delivers a comprehensive analysis of eight recent, top-tier PIDS, revealing critical shortcomings in their design, evaluation, and practical utility.
Bilot's research, a meticulous re-implementation and experimental validation of these systems within a unified framework, uncovers that despite their advanced architectures, many modern PIDS are over-engineered. Surprisingly, a much simpler baseline model, VLOGS, leveraging basic textual attributes, not only matches but significantly outperforms these complex systems across diverse datasets. This work challenges the prevailing trend towards increasing complexity in intrusion detection, advocating for a re-evaluation of design principles, evaluation methodologies, and the fundamental features used for attack detection. The findings have profound implications for the future development and deployment of practical and effective PIDS in real-world industrial settings.
Background
▶ Watch: Introduction to provenance-based IDS and project goals (0:00)
Provenance-based Intrusion Detection Systems (PIDS) operate on the principle of collecting and analyzing system-level events (e.g., process creation, file access, network connections) to construct a causality graph or provenance graph. This graph represents the intricate relationships and dependencies between system entities, allowing security analysts to trace the lineage of suspicious activities and identify the root cause of an attack. Many contemporary PIDS, particularly those published in top-tier conferences, predominantly rely on Graph Neural Networks (GNNs) to process these complex graph structures. These systems are typically anomaly-based, meaning they are trained in an unsupervised or self-supervised manner on benign system behavior, then flag deviations from this learned normalcy as potential attacks.
The impetus for Bilot's work stemmed from a critical observation: despite the theoretical elegance and increasing complexity of these state-of-the-art PIDS, their practical deployment in industrial environments remains limited. This suggests a disconnect between academic advancements and real-world applicability. To address this, Bilot and his team embarked on a rigorous analysis, selecting eight prominent PIDS that represent the cutting edge of research in this domain. These systems were meticulously re-implemented into a unified framework, ensuring a consistent codebase and a standardized environment for fair comparison. This framework allowed for the decomposition of each system into modular components, with their parameters configurable via simple YAML files, facilitating efficient combinatorial architecture exploration and enabling a systematic investigation into the necessity of their inherent complexity. The overarching goal was to identify the most practical PIDS architecture possible, by scrutinizing existing approaches and pinpointing the obstacles preventing their widespread adoption.
Key Findings
▶ Watch: Proposing new Attack Detection Precision (ADP) metric (5:47)
The comprehensive analysis performed by Tristan Bilot and his team revealed nine key shortcomings in state-of-the-art provenance-based intrusion detection systems. While not all could be detailed in the talk, two critical issues were highlighted: flaws in evaluation methodologies and the surprising effectiveness of simpler models.
Firstly, a major concern was the inadequacy of standard evaluation metrics like recall and precision for cybersecurity attack detection. These metrics, while common in machine learning, are ill-suited for the unique characteristics of security datasets for several reasons:
- They aggregate true positives (TP) and false positives (FP) across an entire dataset, failing to provide insights into a system's ability to detect multiple, distinct attacks within that dataset. An ideal system should detect all attacks, not just a high overall count of malicious nodes.
- Metrics like recall can give undue importance to detecting all malicious nodes within an attack. However, for attack reconstruction, often only a few key nodes are sufficient to identify the malicious activity. In such scenarios, precision and a low false positive rate are far more critical.
- Cybersecurity datasets frequently exhibit extreme class imbalance, where benign events vastly outnumber malicious ones. Standard metrics struggle to provide meaningful insights in such imbalanced contexts.
- Recall and precision require the setting of a threshold to classify anomaly scores into benign or attack categories. The arbitrary placement of this threshold can drastically alter the computed metrics, making comparisons between systems unreliable and highly dependent on this tuning parameter. Bilot demonstrated this with examples from Artus, NLink, and Kyros, showing how a model's true capability might be masked or exaggerated by threshold placement.
To address these metric shortcomings, Bilot introduced a new metric: Attack Detection Precision (ADP). ADP is designed to measure a system's ability to detect all attacks within a dataset with high precision and, crucially, in a threshold-independent manner. This metric provides a more accurate reflection of a model's inherent discriminative power, allowing for fairer comparisons between different PIDS.
The second, and perhaps most surprising, key finding was the revelation that simpler architectures can significantly outperform complex state-of-the-art PIDS. The research team developed a simple baseline model named VLOGS. When benchmarked against the eight complex, GNN-based PIDS within their unified framework, VLOGS demonstrated superior performance. Across nine DARPA datasets, after more than 400 compute days on GPUs, VLOGS consistently achieved the highest ADP score, indicating its superior ability to detect attacks in a threshold-independent manner and its better discrimination power between benign and malicious nodes. This unexpected result challenges the common assumption that increasing model complexity necessarily leads to improved detection capabilities in cybersecurity.
These findings suggest that while GNNs can capture different semantics using graph structure alone, their benefits must be carefully weighed against the practical costs and the demonstrated effectiveness of simpler, more deployable alternatives. The study underscores the critical need for comprehensive ablation studies in future research to justify the inclusion of every complex component in a proposed system.
Technical Deep Dive
▶ Watch: Surprising result: VLOX outperforms complex state-of-the-art PIDS (6:58)
The core of Bilot's analysis lies in the meticulously designed Psmaker framework, an open-source platform built to enable fair and efficient experimentation with provenance-based intrusion detection systems. This framework provides a unified codebase for implementing and comparing various PIDS architectures, ensuring that differences in performance are attributable to the underlying models rather than implementation discrepancies. Each PIDS within Psmaker is decomposed into multiple components, allowing researchers to experiment with different combinations and parameters. System configurations, including arguments for each component, are described using a single YAML file per system, streamlining the process of defining and running experiments.
Psmaker incorporates an efficient experimentation pipeline designed to maximize computational resource utilization. It intelligently reuses previously computed results across multiple runs. For instance, if a system is run once, and then again with only a minor change to a single component's argument, the framework automatically restarts computation only from the modified component onwards, saving significant time and resources. This capability is crucial for combinatorial architecture exploration, allowing researchers to systematically test various combinations of components and parameters to search for optimal or simpler architectures. All eight state-of-the-art PIDS selected for this study share a common architectural foundation: they are GNN-based and employ anomaly detection techniques, trained in a self-supervised manner on benign data.
A significant portion of the technical deep dive addressed the critical flaws in traditional evaluation metrics. Bilot elaborated on why recall and precision are problematic for cybersecurity:
- Total Counts vs. Per-Attack Detection: These metrics sum true positives and false positives globally. In a dataset with multiple distinct attacks (e.g., three attacks in the Cadets E3 dataset), a high overall recall might be achieved by detecting many nodes from one attack while completely missing others. A robust PIDS should detect all attacks present.
- Overemphasis on Quantity: Recall prioritizes detecting all true positive nodes. However, for incident response, the ability to reconstruct an attack often requires only a few critical nodes. High precision in identifying these key nodes, coupled with a low false positive rate, is often more valuable than detecting every single compromised entity.
- Extreme Class Imbalance: Cybersecurity datasets are inherently imbalanced, with benign events vastly outnumbering malicious ones. This skews standard metrics, making it difficult to discern true detection capabilities from noise.
- Threshold Dependency: Both recall and precision require an arbitrary threshold to be set, which separates anomaly scores into "benign" and "attack" classes. Bilot illustrated this with anomaly scores from Artus, NLink, and Kyros. Artus showed a perfectly located threshold detecting all three attacks without false positives. NLink, despite having a good underlying model capable of detecting two attacks, suffered from a poorly located threshold leading to zero precision. Kyros demonstrated a model incapable of detecting any attacks regardless of the threshold. This dependency makes direct comparisons between systems unreliable, as performance can be artificially inflated or deflated by threshold tuning.
To overcome the limitations of threshold-dependent metrics, Bilot introduced Attack Detection Precision (ADP). ADP is a threshold-independent metric designed to quantify a system's ability to detect all attacks within a dataset with high precision. It assesses the model's inherent discriminative power, independent of where a human might choose to draw the line. For example, in the NLink case, where a bad threshold resulted in zero precision, ADP would still reflect the model's underlying capability to detect two out of three attacks. This provides a more robust and fair measure of a PIDS's true effectiveness.
The most striking technical contribution was the development and evaluation of VLOGS, a deceptively simple baseline model. VLOGS eschews complex GNNs and instead applies a basic word2vec model to system entities. It extracts text embeddings from textual attributes such as file paths, process command lines, and IP addresses. These embeddings are then fed into a simple neural network. The network is trained in a self-supervised manner to perform system call type prediction, essentially learning to predict the next system call type based on the current context. This minimalistic approach allowed VLOGS to process incoming edges with very low computational overhead.
The comparison involved benchmarking VLOGS against the eight complex state-of-the-art systems on nine diverse DARPA datasets. The results, averaged using the ADP score, unequivocally demonstrated VLOGS's superior ability to detect attacks in a threshold-independent manner. Furthermore, VLOGS exhibited remarkable practical benefits:
- Real-time Detection: Its simple architecture allows it to function in real-time, processing events as they occur.
- Low Memory Usage: It requires approximately 5 megabytes of memory, making it suitable for resource-constrained environments.
- Low CPU Usage: Minimal processing power is needed.
- High Throughput: It can handle more than 2,000 edges per second, demonstrating its scalability.
These technical findings highlight that for the types of attacks present in the DARPA datasets, the textual attributes embedded within system entities are often sufficient for effective detection, challenging the necessity of elaborate graph-based analysis for all scenarios. While GNNs can potentially capture "different semantics" from graph structure alone and offer a "larger detection margin," their practical utility must be rigorously benchmarked against simpler, more deployable alternatives.
Demo / Proof of Concept
▶ Watch: VLOX's real-time detection and low resource usage (8:05)
While the talk did not feature a live, interactive demonstration of an attack or a tool in action, the entire research project serves as a comprehensive proof of concept. The Psmaker framework itself is a robust demonstration of a unified, open-source platform capable of reimplementing, evaluating, and comparing diverse provenance-based intrusion detection systems. Its design, allowing for combinatorial architecture exploration and efficient experimentation, showcases a practical methodology for PIDS research.
The core "proof" of the talk lies in the extensive experimental results. The rigorous benchmarking of VLOGS against eight state-of-the-art PIDS across nine DARPA datasets, involving over 400 GPU compute days, empirically validates the hypothesis that simpler models can achieve, and even surpass, the performance of highly complex systems. The quantitative data presented, particularly the superior Attack Detection Precision (ADP) of VLOGS and its impressive resource efficiency (5MB memory, 2000+ edges/sec throughput), serves as a compelling demonstration of the practical viability and effectiveness of the proposed simpler approach. By open-sourcing Psmaker, the research team provides a tangible artifact for the community to replicate these findings, build upon them, and contribute to the development of more practical PIDS.
Defensive Implications
▶ Watch: The role of GNNs for graph structure-based attack detection (9:40)
The findings from Tristan Bilot's comprehensive analysis carry several significant implications for cybersecurity defenders and the future development of intrusion detection systems.
Firstly, the research strongly advocates for a re-evaluation of complexity in PIDS design. The surprising success of VLOGS, a simple word2vec-based model, against highly complex GNN-based systems suggests that defenders should prioritize practical, deployable solutions over overly engineered ones. This implies that for many common attack patterns, the inherent textual attributes within system entities (file paths, command lines, IP addresses) are highly informative and sufficient for effective detection. Defenders should consider exploring and investing in simpler, attribute-based models that offer high performance with minimal resource overhead.
Secondly, the critique of standard evaluation metrics like recall and precision highlights a critical need for improved benchmarking practices. Defenders and researchers should adopt more appropriate metrics such as Attack Detection Precision (ADP), which provides a threshold-independent and attack-centric measure of a system's true discriminative power. Using ADP will lead to a more accurate understanding of PIDS capabilities and foster the development of models that are genuinely effective in detecting multiple, distinct attacks.
Thirdly, the practical advantages of VLOGS – real-time detection, low memory usage (around 5MB), low CPU usage, and high throughput (over 2,000 edges per second) – make it an attractive candidate for deployment in environments where resources are constrained, such as endpoints or edge devices. This opens the door for lightweight, client-side PIDS that can perform initial, rapid detection, offloading more computationally intensive analysis to centralized servers. The speaker explicitly suggests future work in a two-stage detection approach, where a VLOGS-like model could serve as a client-side pre-filter, followed by in-depth GNN detection on the server for more complex or stealthy attacks.
Finally, the release of the Psmaker framework as open source is a crucial contribution to the defensive community. It provides a standardized platform with eight state-of-the-art baselines and nine DARPA datasets, enabling researchers and practitioners to:
- Fairly benchmark new PIDS against established methods.
- Reproduce existing research findings.
- Explore novel architectures and components efficiently.
- Contribute to a shared codebase, fostering collaborative development of more practical and effective intrusion detection solutions.
This transparency and collaborative spirit are vital for accelerating advancements in cybersecurity defenses.
Key Takeaways
- Complexity vs. Effectiveness: Many state-of-the-art provenance-based intrusion detection systems (PIDS) are overly complex, and their advanced architectures do not necessarily translate to superior detection performance in practice.
- Flawed Evaluation Metrics: Traditional metrics like recall and precision are often inadequate for evaluating PIDS in cybersecurity contexts due to issues with threshold dependency, class imbalance, and their inability to reflect per-attack detection capabilities.
- Attack Detection Precision (ADP): The proposed ADP metric offers a more robust, threshold-independent measure of a PIDS's ability to detect all distinct attacks with high precision, providing a fairer comparison between systems.
- Simplicity Wins: A simple baseline model, VLOGS, leveraging basic text embeddings from system entities (file paths, command lines, IP addresses) and a simple neural network, significantly outperformed complex GNN-based PIDS across multiple DARPA datasets.
- Practical Deployability: VLOGS demonstrates high practical utility with real-time detection, extremely low memory usage (around 5MB), low CPU consumption, and high throughput (over 2,000 edges per second), making it suitable for resource-constrained environments.
- Open-Source Framework: The Psmaker framework has been open-sourced, providing a unified codebase, eight state-of-the-art baselines, and nine datasets to facilitate fair benchmarking, research, and community contributions to PIDS development.
About the Speaker(s)
Tristan Bilot is a PhD student at University Pra and a visiting student at the University of British Columbia (UBC). His research focuses on advancing the field of provenance-based intrusion detection systems, with a particular emphasis on analyzing existing state-of-the-art approaches, identifying their practical shortcomings, and developing more efficient and deployable solutions. His work, as presented in this talk, highlights a commitment to rigorous experimentation and a critical perspective on the balance between model complexity and real-world effectiveness in cybersecurity.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Rigorous systems security research that does the unsexy but necessary work: re-implementing eight GNN-based PIDS in a unified framework, exposing evaluation methodology rot across the subfield, and dropping a simple baseline (VLOGS) that embarrasses the competition across nine DARPA datasets after 400+ GPU-compute-days of validation. The metric critique alone — threshold dependency, per-attack vs. aggregate TP counting, class imbalance masking — is worth the runtime.
Heather Calloway (CISO) — WEAK
Rigorous academic work that exposes a real problem in how the research community builds and evaluates intrusion detection systems. But this talk never crosses the line from research critique into institutional relevance — it tells security engineers what the models get wrong, not security leaders what to do about it.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)