AutoLabel: Automated Fine-Grained Log Labeling for Cyber Attack Dataset Generation
Yihao Peng (PhD student · Chunuha University)
34th USENIX Security Symposium (USENIX Security '25) · Day 1 · System Security 1: Threat Detection, Exploitation, and Adaptive Defenses
Overview
In the rapidly evolving landscape of cybersecurity, the ability to accurately detect and respond to sophisticated attacks hinges on the quality and availability of training data for security models. This talk introduces AutoLabel, a groundbreaking system designed to automate the generation of fine-grained, multi-source labeled log datasets for cyber attack research. Presented by Yihao Peng, a PhD student at Tsinghua University, AutoLabel addresses a critical bottleneck in the field: the acute scarcity of high-quality, labeled log data, which severely hampers the development and evaluation of new security technologies.

Key moments
- 0:00 Introduction: Problem of log data scarcity and existing limitations
- 1:08 Key insight: Reframing problem with provenance graphs
- 2:00 Core automated labeling pipeline overview
- 3:08 Challenge 1: Unifying multi-source logs with log relating
- 4:00 Challenge 2: Solving dependency explosion with unit partitioning
- 5:00 Evaluation results: 100% labeling accuracy
- 6:00 Contribution: Releasing 580 high-quality benchmark datasets
AutoLabel: Automated Fine-Grained Log Labeling for Cyber Attack Dataset Generation
Speakers: Yihao Peng, PhD Student, Tsinghua University
Conference: USENIX Security
YouTube: https://www.youtube.com/watch?v=g5Foz1se3Xw
Overview
In the rapidly evolving landscape of cybersecurity, the ability to accurately detect and respond to sophisticated attacks hinges on the quality and availability of training data for security models. This talk introduces AutoLabel, a groundbreaking system designed to automate the generation of fine-grained, multi-source labeled log datasets for cyber attack research. Presented by Yihao Peng, a PhD student at Tsinghua University, AutoLabel addresses a critical bottleneck in the field: the acute scarcity of high-quality, labeled log data, which severely hampers the development and evaluation of new security technologies.
The core problem stems from the sheer volume and complexity of system logs, making manual labeling an intractable task. Existing automated methods fall short, exhibiting limitations such as coarse-grained labeling, incompleteness, or brittleness. AutoLabel revolutionizes this process by reframing the problem: instead of analyzing individual log entries, it constructs a provenance graph to represent system behavior, where an attack is precisely defined as a subgraph of causally related events. This innovative approach promises to accelerate security research by providing the foundational data necessary for robust model training and validation.
AutoLabel's significance cannot be overstated. By providing a scalable and highly accurate method for generating labeled datasets, it empowers researchers and security professionals to develop more effective detection, analysis, and response mechanisms. The system's ability to automatically and precisely extract attack subgraphs from diverse log sources marks a substantial leap forward, moving beyond the limitations of prior work and paving the way for a new generation of data-driven security solutions.
Background
▶ Watch: Introduction: Problem of log data scarcity and existing limitations (0:00)
Log-based security research is a cornerstone of modern cybersecurity, underpinning the development of intrusion detection systems, threat hunting tools, and forensic analysis capabilities. However, the efficacy of these tools and the models that power them is fundamentally constrained by a significant, persistent bottleneck: the severe scarcity of high-quality, labeled log datasets. The volume of logs generated by even a single system in a short period is astronomical, making the manual creation of such datasets a nearly impossible, labor-intensive, and error-prone endeavor. This lack of robust ground truth data creates a substantial hurdle for training, testing, and evaluating novel security models, ultimately slowing down progress in the fight against cyber threats.
Prior attempts to automate the generation of labeled log data have faced considerable limitations. One common approach involves time window-based labeling, where all system activity occurring within a specific time frame around a known attack is indiscriminately labeled as malicious. This method is inherently too coarse-grained, as it fails to distinguish between legitimate and malicious events within the window, leading to high rates of false positives and a lack of precision regarding the exact scope of an attack. Such datasets are of limited utility for training models that require fine-grained understanding of attack causality.
Another class of methods relies on detection tool-based labeling. These systems leverage existing security tools, such as anti-malware software or intrusion detection systems, to identify malicious activity and then label associated logs. While seemingly practical, this approach is inherently incomplete. Detection tools are designed to identify known threats or patterns; they can only find what their underlying rules or signatures are programmed to see. Consequently, datasets generated this way will inevitably miss novel or polymorphic attacks, perpetuating a bias towards known threats and failing to provide comprehensive coverage for emerging attack vectors. This limitation is particularly problematic for developing models that aim to detect zero-day exploits or highly sophisticated, custom attacks.
Finally, rule-matching systems have been employed to label logs by defining specific patterns or sequences of events indicative of an attack. While potentially offering more precision than time-window methods, these systems are notoriously laborious to create and maintain. They require extensive expert knowledge to define complex rules, which are often highly specific to particular attack techniques, system configurations, or software versions. Even minor variations in attack methodology or changes in the system environment can render these rules obsolete, making them extremely brittle and demanding constant, costly updates from security experts. This lack of adaptability makes rule-matching systems impractical for generating diverse and comprehensive attack datasets.
AutoLabel’s key insight to overcome these limitations is to fundamentally reframe the problem. Instead of focusing on individual log entries or fragmented detection signals, AutoLabel adopts a holistic view of system behavior by constructing a provenance graph. A provenance graph is a powerful data structure that captures the causal relationships between system entities (e.g., processes, files, network connections) and events (e.g., process creation, file write, network send). In this unified, interconnected graph, an attack is no longer a collection of disparate log entries but rather a clearly defined subgraph of causally related events. This conceptual shift transforms the challenge of labeling logs into the more precise and automatable task of automatically and accurately extracting this specific attack subgraph from the broader system provenance graph.
Key Findings
▶ Watch: Core automated labeling pipeline overview (2:00)
AutoLabel represents a significant advancement in the automated generation of high-quality, labeled log datasets for cybersecurity research. The system's core contribution is its ability to automate fine-grained, multi-source log labeling, directly addressing the critical bottleneck of data scarcity in the field. This automation allows for the creation of rich, contextually accurate datasets that are essential for training and evaluating advanced security models.
A primary finding of the extensive evaluation conducted by the AutoLabel team is its unprecedented labeling accuracy. The system achieved 100% fine-grained labeling accuracy across all testing scenarios. This perfect score underscores the precision and reliability of AutoLabel's underlying methodology, ensuring that every log entry causally related to an attack is correctly identified and labeled, while unrelated benign activities are excluded. This level of accuracy is crucial for developing robust security models that can distinguish subtle malicious activities from legitimate system noise.
The evaluation encompassed a diverse and challenging set of scenarios, demonstrating AutoLabel's robustness and versatility. These included:
- 29 diverse scenarios in total.
- 25 real-world CVEs, showcasing its applicability to known vulnerabilities and exploits.
- A 48-emulation of the Sandworm group from MITRE ATT&CK, validating its capability to label complex, multi-stage advanced persistent threats (APTs).
- A sophisticated 10-hop lateral movement attack, pushing the system's limits in tracing extended and intricate attack chains across multiple hosts.
Beyond its exceptional accuracy, AutoLabel also demonstrated remarkable efficiency. The process of generating a complete, labeled dataset took less than 96 minutes per dataset. This high efficiency makes AutoLabel a practical solution for generating large volumes of data, overcoming the previously prohibitive time and resource costs associated with manual labeling or less efficient automated methods.
As a direct and impactful contribution to the security community, the researchers are releasing 580 new benchmark datasets generated by AutoLabel. These datasets, meticulously labeled with 100% accuracy and derived from a wide array of attack scenarios, are poised to become a vital resource. They will serve as a foundational benchmark for researchers to develop, test, and compare new security technologies, fostering innovation and accelerating progress in areas like anomaly detection, threat intelligence, and automated incident response. The availability of such a rich, high-quality dataset collection is expected to significantly mitigate the data scarcity problem that has historically hindered advancements in security research.
Technical Deep Dive
▶ Watch: Challenge 1: Unifying multi-source logs with log relating (3:08)
AutoLabel’s efficacy stems from a sophisticated, multi-stage automated labeling pipeline that leverages provenance graphs as its central data structure. The pipeline is meticulously designed to overcome the inherent complexities of diverse log sources, process intertwining, and the sheer volume of system activity.
The process begins with raw, unlabeled logs collected from multiple system sources. These typically include low-level audit logs (e.g., Linux auditd, Windows ETW), which provide detailed records of system calls, process activities, file operations, and network connections; traffic logs (e.g., network flow data, packet captures); and application logs (e.g., web server logs, database logs), which capture higher-level application-specific events.
The first crucial step is to construct a foundational provenance graph primarily from the low-level audit logs. Provenance graphs are directed acyclic graphs where nodes represent system entities (processes, files, sockets, registry keys) and edges represent causal dependencies between events (e.g., a process creating a file, a process opening a network connection, one process spawning another). This initial graph provides a granular, system-wide view of interactions.
The second, and highly innovative, step is Log Relating, which involves fusing the other log types (application and traffic logs) onto this foundational audit log-based provenance graph. This is a critical challenge because different log sources often use disparate identifiers and have varying levels of detail. AutoLabel addresses this by using lightweight instrumentation during data collection. This instrumentation generates alignment information which acts as the "glue" to correlate events across different log types. For instance, network traffic logs might be correlated with process network activity recorded in audit logs using source/destination IP addresses, ports, and timestamps. Application-specific events might be correlated with underlying system calls recorded in audit logs through process IDs (PIDs), thread IDs, or specific timestamps and arguments. This fusion process creates a single, unified provenance graph where all system activities, regardless of their original log source, are interconnected through causal links. This unified view is essential for a comprehensive understanding of an attack’s progression.
With the unified graph in place, the next challenge is to identify the starting point of an attack. This is addressed in the Anchor Location phase. During the controlled execution of an attack in a lab environment (which is necessary for generating ground truth data), AutoLabel injects unique identifiers, referred to as attack flags. These flags are embedded into specific system events or data streams that are part of the attack's initial entry point. For example, a flag might be injected into a specific network packet, a file name, or a process argument. These flags serve as unambiguous anchor points within the unified provenance graph, precisely telling AutoLabel where the attack began and providing a definitive starting node for tracing the attack's causal chain.
Even with anchors identified, directly tracing dependencies in real-world graphs can lead to the classic dependency explosion problem. This occurs frequently with long-running server processes (e.g., a web server, database server) that handle numerous concurrent requests. In a standard provenance graph, all activities of such a process might appear intertwined, creating a tangled mess where it’s impossible to isolate the events related to a single malicious request from thousands of legitimate ones. AutoLabel solves this in the Graph Refinement phase using a technique called Unit Partitioning. This involves instrumenting the application layer to logically break down the server process's activity into discrete execution units. Each unit corresponds to a single logical request or transaction. For example, a web server might assign a unique request ID to each incoming HTTP request, and this ID would propagate through all subsequent system calls and internal operations related to that specific request. By tracking these execution units, AutoLabel effectively disentangles the graph, allowing it to isolate the causal chain of a specific malicious request from the overwhelming background noise of benign server operations.
Finally, in the Attack Subgraph Extraction phase, with anchors precisely located and the graph refined through unit partitioning, AutoLabel can cleanly and accurately trace the attack's causal chain. Starting from the identified anchor points, the system traverses the refined provenance graph, following all causal dependencies (e.g., process A spawned process B, process B wrote to file C, process C established network connection D). This traversal precisely extracts the complete attack subgraph, which encompasses every single system event and entity causally linked to the attack, from its inception to its final actions. Because all log types were fused into the unified graph in step two, every log entry associated with this extracted subgraph is now automatically and accurately labeled as part of the attack.
The two core technical pillars underpinning AutoLabel's success are therefore the log relating mechanism for creating a unified view from disparate log sources, and the combination of attack flags for precise anchor identification coupled with unit partitioning for disentangling complex process dependencies, enabling the precise extraction of the attack subgraph.
Demo / Proof of Concept
▶ Watch: Evaluation results: 100% labeling accuracy (5:00)
While the talk did not feature a live, interactive demonstration, the effectiveness and capabilities of AutoLabel were rigorously proven through an extensive and comprehensive evaluation methodology. This evaluation served as the robust proof of concept for the system, demonstrating its ability to accurately and efficiently label complex cyber attack scenarios.
The evaluation encompassed 29 diverse scenarios, carefully selected to challenge AutoLabel across various attack vectors and complexities. These included:
- 25 real-world Common Vulnerabilities and Exposures (CVEs), such as CVE-2017-0144 (EternalBlue), CVE-2017-11882 (Microsoft Office Memory Corruption), and other well-documented exploits. This validated AutoLabel's ability to handle the nuances of known exploit chains.
- A 48-emulation of the Sandworm group from the MITRE ATT&CK framework. This involved simulating a sophisticated Advanced Persistent Threat (APT) group's tactics, techniques, and procedures (TTPs), including reconnaissance, initial access, execution, persistence, privilege escalation, defense evasion, credential access, discovery, lateral movement, and impact. This demonstrated AutoLabel's capability to trace and label multi-stage, complex attack campaigns.
- A highly challenging 10-hop lateral movement attack. This scenario involved an attacker traversing ten different machines within a network, showcasing AutoLabel's ability to maintain causal fidelity and accurately label events across distributed systems and extended attack paths.
The results of this evaluation were highly compelling: AutoLabel achieved 100% fine-grained labeling accuracy across all 29 scenarios. This perfect score indicates that the system precisely identified and labeled every log entry belonging to the malicious subgraph while correctly excluding all benign activity. The validation methodology employed a combination of techniques: manual inspection was used for human-readable logs to ensure semantic correctness, and a set of automatic checks was implemented to verify that all key attack artifacts and their complete causal chains were accurately captured in the low-level data.
Furthermore, the evaluation highlighted AutoLabel's efficiency, with the system taking less than 96 minutes per dataset to perform the entire labeling process. This speed, combined with its accuracy, underscores AutoLabel's potential to generate a vast quantity of high-quality, labeled data. As a direct outcome of this rigorous testing and validation, the AutoLabel team is contributing 580 new benchmark datasets to the security community. These datasets, generated and meticulously labeled by AutoLabel, represent a significant resource for future research and development in cybersecurity.
Defensive Implications
▶ Watch: Contribution: Releasing 580 high-quality benchmark datasets (6:00)
AutoLabel's innovative approach to automated, fine-grained log labeling carries profound implications for cybersecurity defenders, offering new avenues to enhance detection, analysis, and response capabilities. The most immediate and significant benefit lies in the provision of high-quality, labeled cyber attack datasets. Defenders can leverage these datasets in several critical ways:
- Training and Evaluating Advanced Security Models: The scarcity of labeled data has long been a major impediment to developing effective machine learning and artificial intelligence models for security. AutoLabel's datasets provide the necessary ground truth to train and validate next-generation anomaly detection systems, threat hunting tools, and automated incident response platforms. Models trained on such precise data will be better equipped to distinguish subtle attack patterns from benign system noise, leading to reduced false positives and more accurate threat identification.
- Improving Existing Detection Systems: Security teams can use these benchmark datasets to rigorously test and improve their current Intrusion Detection Systems (IDS), Security Information and Event Management (SIEM) rules, and Endpoint Detection and Response (EDR) solutions. By replaying attack scenarios with the labeled data, defenders can identify gaps in their existing coverage, fine-tune detection logic, and assess the true efficacy of their security controls against a diverse range of real-world and emulated attacks, including complex MITRE ATT&CK techniques.
- Enhanced Threat Hunting and Forensic Analysis: The concept of the provenance graph and the ability to extract precise attack subgraphs can inform and improve threat hunting methodologies. Security analysts can adopt a similar causal tracing approach, even if not fully automated, to better understand complex attack progressions within their own environments. This allows for more effective root cause analysis and a deeper understanding of attacker tactics, techniques, and procedures (TTPs) by visually mapping out the exact sequence of events.
- Validating Logging Infrastructure and Data Quality: The process of requiring diverse log sources (audit, traffic, application) and fusing them into a unified graph implicitly highlights the importance of comprehensive logging. Defenders can use the AutoLabel framework as a conceptual guide to assess their own logging infrastructure, ensuring they are collecting the necessary data points (e.g., process creation, file I/O, network connections) with sufficient detail and correlation capabilities to reconstruct attack chains.
- Developing Context-Aware Security Policies: By understanding the fine-grained causal relationships between attack events, defenders can develop more intelligent and context-aware security policies. Instead of relying on isolated indicators, policies could be designed to trigger based on sequences of causally linked events, making them more resilient to evasion techniques and less prone to false positives.
- Accelerating Security Research and Innovation: The public release of 580 benchmark datasets generated by AutoLabel will democratize access to high-quality security data. This will accelerate academic and industry research into new detection algorithms, behavioral analytics, and automated defense mechanisms, ultimately leading to a stronger collective defense against cyber threats.
In essence, AutoLabel provides not just data, but a blueprint for understanding and analyzing cyber attacks with unprecedented precision. By embracing the principles of provenance graphing and fine-grained causal analysis, defenders can move beyond reactive, signature-based defenses towards proactive, behavioral-based security postures that are better equipped to handle the sophistication of modern threats.
Key Takeaways
- The scarcity of high-quality, labeled log datasets is a critical bottleneck hindering the development and evaluation of advanced security models.
- AutoLabel is a novel system that automates the generation of fine-grained, multi-source labeled log datasets for cyber attack research.
- It reframes the problem by using provenance graphs to represent system behavior, where an attack is defined as a causally related subgraph of events.
- AutoLabel's core technical pillars include a log relating mechanism to fuse diverse log sources into a unified graph, and a combination of attack flags for anchor identification and unit partitioning to disentangle complex process dependencies.
- The system achieved 100% fine-grained labeling accuracy across 29 diverse attack scenarios, including 25 real-world CVEs, MITRE ATT&CK Sandworm emulation, and a 10-hop lateral movement attack.
- AutoLabel is highly efficient, generating a complete dataset in less than 96 minutes.
- The researchers are releasing 580 new benchmark datasets to the security community, significantly contributing to the availability of high-quality, labeled data for security research and development.
About the Speaker(s)
Yihao Peng is a PhD student at Tsinghua University. His research focuses on critical areas of cybersecurity, particularly addressing fundamental challenges such as the scarcity of high-quality data for security analysis and the development of automated methods to overcome these limitations. His work, exemplified by AutoLabel, aims to advance the state of the art in log-based security research and empower the development of more effective security technologies.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Solid systems research that attacks a real, unsexy problem: the chronic shortage of fine-grained, labeled log data for security ML. The provenance-graph reframing plus unit partitioning for dependency explosion is genuinely clever, and 580 released benchmark datasets is a concrete community contribution. 100% accuracy claims warrant scrutiny, but the technical architecture earns the benefit of the doubt at USENIX.
Heather Calloway (CISO) — WEAK
Technically credible research that solves a real problem in the ML security pipeline — labeled training data scarcity. But the talk stays firmly in the research lab and never makes the jump to operator relevance, security program decisions, or institutional accountability. Useful for academics building detection models; limited value for anyone running a security program.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)