Cost-effective Attack Forensics by Recording and Correlating File System Changes
Le Yu (Purdue University)
33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24
Overview
In an era marked by an unprecedented surge in Internet of Things (IoT) device attacks—a threefold increase between 2020 and 2022, surpassing 100 million incidents annually—the imperative for robust and efficient attack forensics has never been clearer. This talk, delivered by Le Yu from Purdue University, introduces a novel approach to forensic analysis that challenges conventional wisdom by shifting focus from high-frequency event logging to low-frequency file system state changes. Titled "Cost-effective Attack Forensics by Recording and Correlating File System Changes," the presentation delves into the inherent limitations of existing provenance systems and proposes an innovative solution designed to overcome these challenges, particularly in resource-constrained environments like IoT hubs.

Key moments
- 0:00 Introduction: Rising IoT attacks and forensics challenges
- 4:18 Limitations of existing temporal-based provenance systems
- 6:19 Key insights: Low-frequency attacks & spatial snapshots
- 7:15 Inferring causality by correlating state information (content matching)
- 7:50 Overview of the proposed system's five-step workflow
- 9:00 Generating provenance graphs using universal file model
Cost-effective Attack Forensics by Recording and Correlating File System Changes
Speakers: Le Yu
Conference: USENIX Security '24
YouTube: https://www.youtube.com/watch?v=UFwLiw8ChZI
Overview
In an era marked by an unprecedented surge in Internet of Things (IoT) device attacks—a threefold increase between 2020 and 2022, surpassing 100 million incidents annually—the imperative for robust and efficient attack forensics has never been clearer. This talk, delivered by Le Yu from Purdue University, introduces a novel approach to forensic analysis that challenges conventional wisdom by shifting focus from high-frequency event logging to low-frequency file system state changes. Titled "Cost-effective Attack Forensics by Recording and Correlating File System Changes," the presentation delves into the inherent limitations of existing provenance systems and proposes an innovative solution designed to overcome these challenges, particularly in resource-constrained environments like IoT hubs.
The core of Yu's work addresses the twin problems of excessive overhead and inaccurate causality inference that plague current forensic tools. Traditional methods, often reliant on kernel-level syscall monitoring or eBPF-based logging, generate voluminous data that can overwhelm systems and obscure the true chain of events leading to an attack. By adopting a spatial, snapshot-based perspective, the presented system aims to drastically reduce the data footprint while enhancing the precision and recall of forensic investigations. This paradigm shift holds significant promise for improving the security posture of IoT ecosystems, enabling more effective post-compromise analysis without imposing prohibitive performance costs.
Background
▶ Watch: Introduction: Rising IoT attacks and forensics challenges (0:00)
Attack forensics traditionally relies on provenance systems designed to track the lineage of data, recording metadata about its creation, modification, and transfer over time. This historical data is then used for causality analysis, typically generating dependency graphs to trace back from a symptom event to its root cause. A common example involves analyzing audit logs to link a compromised file to a remote IP address that initiated the attack. However, despite their utility, existing provenance systems suffer from several critical limitations, especially when applied to the unique characteristics of IoT devices.
Current provenance systems can be broadly categorized into several types. Kernel-level syscall API monitoring systems, such as Spade, leverage the Linux audit subsystem to log events when syscalls match predefined auditing rules. Another category, kernel-level security API monitoring, involves instrumenting the OS kernel for comprehensive auditing, exemplified by systems like LPM, Hi-Fi, and CamFlow. More recently, eBPF-based monitoring has gained traction, with tools like eAudit intercepting forensics-related syscall entry and exit points on demand. While these methods provide detailed event logs, they often lead to false dependencies and provenance graph explosions, where every process output is assumed to depend on all preceding inputs, creating overly complex and noisy graphs.
A fourth category attempts to mitigate these issues by integrating application-level information, exposing details from within applications to separate long-running processes into sub-execution units. This combination of system and application-level data aims to refine dependency analysis. However, a fundamental limitation across all these approaches is their reliance on recording log entries in a temporal domain. This means every event, regardless of its significance, generates a log entry.
This temporal logging paradigm proves particularly problematic for IoT hubs, which are characterized by frequent, benign routine tasks (e.g., polling sensors, checking cloud status) that generate a high volume of events without altering the system state. For instance, an IoT hub running OpenHAB on a Raspberry Pi might have crucial configuration files that define its interaction with other devices. While these files are critical, an attacker might exploit a missing privilege check to modify them. Traditional systems would log numerous high-frequency, harmless events alongside the low-frequency, malicious configuration file change, making it difficult to distinguish unauthorized modifications from normal system operations. The sheer volume of irrelevant data burdens the system, increases storage overhead, and complicates forensic analysis by obscuring the critical attack-related events.
Key Findings
▶ Watch: Key insights: Low-frequency attacks & spatial snapshots (6:19)
The research presented by Le Yu introduces a paradigm shift in attack forensics, driven by two fundamental insights that directly address the limitations of existing provenance systems:
- Attacks are often low-frequency events, contrasting with high-frequency benign system activities. This observation suggests that continuously logging every temporal event, as traditional systems do, is inefficient and generates excessive noise. Instead, the authors propose a spatial dimension approach: taking very low-frequency snapshots of the file system state. By focusing on state changes rather than individual events, the system only needs to record the delta between snapshots to recover the full system state. This drastically reduces the volume of forensic data.
- Analyzing file system state changes can directly extract forensic content, which is crucial for causality graph generation. This insight posits that the content of changed files, along with their metadata, holds sufficient information to infer the causality of an attack. Rather than relying on a chronological sequence of syscalls, the system leverages the "what changed" to understand the "why and how." This is particularly effective for attacks that alter critical system configurations or inject malicious code, as these actions inherently involve file system modifications.
These insights lead to the core contribution: a method to reverse-engineer causality from state information by correlating state information across different snapshots. The challenge then transforms from event-stream analysis to a content matching and analysis problem. For example, data flow from file copies can be recovered by matching file content, and process read/write operations can be disclosed by extracting file and directory names embedded within executable sessions. This novel approach promises to deliver a more cost-effective and accurate forensic capability, especially for environments where resource constraints and event noise are significant concerns.
Technical Deep Dive
▶ Watch: Inferring causality by correlating state information (content matching) (7:15)
The proposed system, designed for cost-effective attack forensics, operates through a meticulously structured five-step workflow that leverages file system snapshots and content analysis to infer causality.
- Snapshot Analysis: The initial step involves regularly acquiring snapshots from the underlying file system. Crucially, the system is designed to integrate with ZFS systems, which inherently employ a copy-on-write (CoW) strategy. This means that snapshots in ZFS only record the differences or changes in file system data since the last snapshot, rather than a full copy. This feature is fundamental to the system's cost-effectiveness, as it allows for easy and efficient determination of a list of changed files between any two snapshots, significantly minimizing storage and processing requirements.
- File Parsing: Once the changed files are identified, the system parses them using a universal file model. This model is specifically customized for forensic purposes, designed to extract and represent a rich set of information regardless of the file's original format. Key forensic-related attributes captured by this model include:
- Metadata: Timestamps (creation, modification, access), file size, ownership, permissions.
- Data Source: Information about the origin of the file content (if inferable).
- Location: The file's path within the file system.
- API Embeddings: For executable scripts or binaries, the model attempts to extract references to APIs, functions, or other files/directories that might be invoked or accessed by the executable. This is critical for inferring control flow and data flow dependencies.
- Provenance Graph Generation: With parsed files in the universal model format, the system proceeds to generate a content provenance graph. The inputs for this stage are the universal file models derived from both the symptom files (e.g., the compromised configuration file) and all other changed files identified between the relevant snapshots. The core of this process lies in a set of predefined datalog rules. These rules define various types of edges that represent different causal relationships within the graph:
- Data flow age: Indicates when content from one file influences or is copied into another. This is inferred by matching file content similarities.
- Control flow age: Represents one file (e.g., an executable script) causing an action that affects another file. This might be inferred from API embeddings or command-line arguments.
- Content match ages: A more general category for when files share significant content, suggesting a relationship even if not a direct data or control flow.
The datalog engine takes these inputs and rules to generate the content provenance graph. The process is conceptually similar to information retrieval: the symptom file acts as a query or keyword, and the datalog engine functions as a search engine, identifying and linking all forensically relevant content in other changed files. This approach constructs a graph where nodes are files (or specific components within them) and edges represent inferred causal relationships based on content and metadata.
- Attack-Related Subgraph Identification: Given that the generated content provenance graph can still be extensive, the system employs a ranking mechanism to highlight the most relevant nodes and edges pertaining to the attack symptom. This involves two sub-steps:
- Edge Weight Calculation: Each edge in the content provenance graph is assigned a weight. This weight is determined by factors such as the file content similarity between the connected nodes and their semantic coupling. Higher similarity or stronger semantic links (e.g., one script directly invoking another) result in higher edge weights.
- Node Weight Computation: Starting from the symptom files, a propagation algorithm—conceptually similar to PageRank—is used to calculate an "importance score" for each node. The weights of the edges are propagated through the graph, accumulating scores on connected nodes. Nodes with higher cumulative scores are deemed more related to the attack. Finally, only those nodes exceeding a predefined threshold are considered part of the "attack-related subgraph," effectively filtering out irrelevant information and focusing the forensic investigation.
This detailed technical approach provides a comprehensive framework for inferring causality from static file system changes, promising a more efficient and targeted forensic analysis than traditional event-based methods.
Demo / Proof of Concept
▶ Watch: Overview of the proposed system's five-step workflow (7:50)
The efficacy of the proposed system was rigorously evaluated across a diverse set of five different devices, encompassing three mainstream smart home platforms, one Android phone, and one personal computer. This broad selection aimed to demonstrate the system's versatility and applicability in various computational environments, particularly those with resource constraints common in IoT.
For the evaluation, the researchers curated a set of 16 attacks. These attacks were chosen based on their detailed attack steps and the availability of reproducible Proofs of Concept (PoCs) from recent years. This ensures that the evaluation was grounded in realistic and verifiable attack scenarios. For the personal computer, the researchers leveraged existing attack scripts from the Data Transparent Computing programs, a well-known resource for security research. For the IoT devices, the provided reproducible PoCs were utilized directly.
A crucial aspect of the evaluation involved manually labeling the nodes and edges generated in the provenance graphs. This meticulous manual labeling served as the ground truth against which the system's performance was measured, allowing for accurate calculation of precision and recall.
The system's performance was compared against three prominent baselines:
- Claw: A kernel-level provenance system.
- Alchemist: An application-level provenance system.
- eAudit: A recent eBPF-based provenance system.
The evaluation focused on three key metrics:
- Runtime Overhead: Measured using Postmark, a file system benchmark that simulates mail server behaviors. The metric was the ratio of CPU time consumed by the provenance collection system to that of the benchmark. The proposed system achieved an average runtime overhead of only 3%, significantly outperforming eAudit (5%) and the other baselines (over 14%).
- Space Overhead: Measured over a one-week period, quantifying additional files created. The proposed system consumed only 100 megabytes (MB) on average for one week, whereas other systems generated 6 to 20 times more overhead. This highlights the substantial reduction in storage requirements due to the snapshot-based approach.
- Effectiveness in Attack Forensic Scenarios: Evaluated using precision (proportion of nodes in the graph related to attacks) and recall (proportion of actual attack nodes included in the graph). The proposed system consistently performed best in terms of both precision and recall, demonstrating its superior ability to accurately identify attack-related information while minimizing irrelevant data.
These results unequivocally underscore the system's ability to provide cost-effective and highly accurate attack forensics, making it a compelling alternative to traditional, resource-intensive methods.
Defensive Implications
▶ Watch: Generating provenance graphs using universal file model (9:00)
The findings presented in this talk offer several crucial implications for cybersecurity defenders, particularly those managing large-scale IoT deployments or resource-constrained environments.
Firstly, the significant reduction in runtime overhead (3% average) and space overhead (100MB/week, 6-20x less than baselines) means that comprehensive forensic capabilities can now be deployed on devices that were previously considered too constrained. This directly addresses a major hurdle in securing IoT ecosystems, where continuous, high-fidelity monitoring has often been impractical due to power, processing, and storage limitations. Defenders can now consider integrating such a system to gain deep forensic insights without compromising device performance or requiring extensive infrastructure upgrades.
Secondly, the improved precision and recall in identifying attack-related provenance graphs translates directly into more effective and less time-consuming investigations. By focusing on file system state changes and inferring causality through content analysis, the system generates cleaner, more relevant forensic data. This allows security analysts to quickly pinpoint the true root cause and scope of an attack, reducing the "noise" that often overwhelms traditional event-based logs. Defenders can leverage this to accelerate incident response, minimize dwell times, and ensure a more accurate understanding of compromise pathways.
Thirdly, the system's strength in identifying unauthorized file changes from normal file changes is a powerful defensive tool. Attacks targeting critical configuration files, scripts, or binaries often manifest as low-frequency, yet highly impactful, state alterations. Traditional systems struggle to differentiate these from routine system updates or benign operations. By applying content similarity and semantic coupling, this new approach provides a robust mechanism to detect and prioritize suspicious modifications, enabling proactive threat hunting and more precise alerting for critical assets.
Finally, defenders should consider integrating ZFS-like copy-on-write file systems as a foundational component of their security architecture where feasible. The system's efficiency is heavily reliant on the native snapshotting capabilities of such file systems. Furthermore, the concept of content-based causality inference can inspire new defensive strategies, moving beyond simple event correlation to deeper semantic analysis of system changes. This paradigm shift could lead to the development of next-generation intrusion detection and forensic tools that are better equipped to handle the evolving landscape of sophisticated, state-altering attacks.
Key Takeaways
- Traditional event-based provenance systems suffer from high runtime and space overheads, generating excessive data ("noise") that complicates forensic analysis, especially in resource-constrained IoT environments.
- A novel approach focuses on low-frequency file system state changes captured through efficient copy-on-write snapshots (e.g., from ZFS systems) rather than continuous, high-frequency event logging.
- Causality is effectively inferred by correlating content across snapshots using a universal file model and datalog rules that define data flow, control flow, and content match relationships, transforming forensics into a content matching and analysis problem.
- The system significantly reduces forensic data volume and processing burden, achieving an average runtime overhead of only 3% and consuming 6 to 20 times less storage compared to state-of-the-art syscall, eBPF, and application-level monitoring tools.
- It demonstrates superior effectiveness in attack forensic scenarios, yielding higher precision and recall in identifying attack-related provenance graphs by filtering out irrelevant information through weighted graph analysis.
- This method is particularly well-suited for detecting and analyzing low-frequency, state-altering attacks, such as configuration file tampering or code injection, which are common and impactful in IoT and other critical systems.
About the Speaker(s)
Le Yu is a researcher from Purdue University. He presented this work, which was developed under the mentorship of Xiangyu Zhang, focusing on novel methods for cost-effective attack forensics. His research interests lie in improving security analysis and incident response through innovative system monitoring and data correlation techniques.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
Yu's research presents a genuinely novel and highly effective paradigm shift in attack forensics, moving from noisy temporal event logging to efficient, content-driven analysis of file system state changes. This approach drastically reduces overhead while significantly improving precision and recall, making high-fidelity forensics practical for resource-constrained environments like IoT. This isn't just an improvement; it's a critical enabler for securing a vast, underserved attack surface.
Heather Calloway (CISO) — STRONG ACCEPT
This research offers a pragmatic and highly effective solution to a pervasive challenge in modern forensics: gaining deep visibility without crippling performance. By shifting from noisy event logs to targeted file system state changes, it provides a credible path to operationalizing forensics in resource-constrained environments like IoT, directly informing incident response and risk accountability. Its proven cost-effectiveness changes the calculus for security leaders.