Blackbox Fuzzing of Distributed Systems with Multi-Dimensional Inputs and Symmetry-Based Feedback Pruning
Yonghao Zou
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Fuzzing 2
Overview
Modern digital infrastructure relies heavily on distributed systems, from databases like ClickHouse and RethinkDB to crucial coordination systems. However, the inherent complexity of these systems makes them highly susceptible to subtle bugs that can lead to significant economic losses and operational failures. This talk introduces DisFuzz, a novel blackbox fuzzer specifically designed to uncover these elusive vulnerabilities in distributed environments. DisFuzz distinguishes itself by employing an extended input space that encompasses regular client events, fault injections, and crucial timing intervals, combined with an innovative symmetry-based feedback pruning mechanism to efficiently navigate the vast state space of distributed systems.
Key moments
- 0:00 Introduction: Challenges in distributed systems fuzzing
- 2:00 Introducing DisFuzz and its key contributions
- 2:40 Extended input space: regular events and timing intervals
- 4:00 Need for effective fuzzing feedback and pruning
- 6:00 Novel symmetry-based feedback pruning technique
- 7:15 DisFuzz system architecture and event implementation
- 8:15 Diverse bug checkers and reproduction workflow
- 9:00 Evaluation results and discovered bugs
Blackbox Fuzzing of Distributed Systems with Multi-Dimensional Inputs and Symmetry-Based Feedback Pruning
Speakers: Yonghao Zou (presented by Ejos)
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=RRj_D-O-iJI
Overview
Modern digital infrastructure relies heavily on distributed systems, from databases like ClickHouse and RethinkDB to crucial coordination systems. However, the inherent complexity of these systems makes them highly susceptible to subtle bugs that can lead to significant economic losses and operational failures. This talk introduces DisFuzz, a novel blackbox fuzzer specifically designed to uncover these elusive vulnerabilities in distributed environments. DisFuzz distinguishes itself by employing an extended input space that encompasses regular client events, fault injections, and crucial timing intervals, combined with an innovative symmetry-based feedback pruning mechanism to efficiently navigate the vast state space of distributed systems.
The research presented highlights the limitations of existing fuzzing tools, which often fall short in their ability to generate diverse inputs or provide effective feedback in the context of distributed architectures. DisFuzz addresses these shortcomings through its multi-dimensional input generation and intelligent state pruning, allowing it to explore more relevant execution paths. The efficacy of DisFuzz is demonstrated through its impressive results: the discovery of 52 bugs across 10 different distributed systems, including two industry-grade Raft implementations and eight production-level systems. Four of these bugs were deemed significant enough to be assigned Common Vulnerabilities and Exposures (CVEs), underscoring DisFuzz's capability to identify critical security flaws.
This work is particularly important for the security and reliability of critical infrastructure. By providing a systematic and effective approach to fuzzing distributed systems, DisFuzz offers a powerful tool for developers and security researchers to proactively identify and mitigate complex concurrency and timing-related bugs that are often missed by traditional testing methods. The findings reinforce the necessity of specialized fuzzing techniques that account for the unique characteristics of distributed computing, ultimately contributing to more robust and secure systems.
Background
▶ Watch: Introduction: Challenges in distributed systems fuzzing (0:00)
Distributed systems, while foundational to modern computing, present unique challenges for ensuring safety and robustness. Their complexity stems from several common features. Firstly, they must process a diverse array of input events. These typically fall into two categories: regular events, such as client requests (e.g., get or put operations in a database), and fault events, which simulate real-world disruptions like network failures, node crashes, or message delays. The interplay between these event types, especially under specific timing conditions, is a fertile ground for bugs.
Secondly, the state changes within a running distributed system are predominantly reflected in network messages exchanged between nodes. For instance, in consensus algorithms like Raft, all significant state transitions are either triggered by receiving a specific network message or by a timeout, both of which subsequently lead to sending further network messages. This makes network message sequences a critical source of feedback for understanding system behavior, yet also a source of immense redundancy.
Finally, timeout mechanisms are ubiquitous in distributed systems, governing everything from leader elections to message retransmissions and failure detection. Consequently, the timing dependencies between events are paramount. A slight variation in the relative arrival times of messages or the occurrence of events can lead to entirely different execution paths and expose subtle race conditions or deadlocks.
Prior attempts to fuzz distributed systems, such as JPEN, CrashFuzz, and Meery, have faced significant limitations. These tools often suffer from limited feedback metrics, failing to capture the nuanced state changes crucial for effective distributed system fuzzing. Furthermore, many lack automatic input generation, requiring manual effort to craft test cases, which becomes impractical given the vast input space. The inability to systematically generate diverse combinations of regular events, fault events, and precise timing intervals, coupled with an inefficient mechanism for pruning redundant test cases, has hindered their ability to thoroughly explore the complex state spaces of distributed applications. This gap in existing tooling underscored the need for a more sophisticated approach like DisFuzz, one that is specifically designed to leverage the unique characteristics of distributed systems for more effective bug discovery.
Key Findings
▶ Watch: Extended input space: regular events and timing intervals (2:40)
The research introduces DisFuzz as a significant advancement in the field of distributed system fuzzing, addressing critical limitations of prior work. The primary contribution lies in its novel approach to input generation and feedback mechanisms, leading to the discovery of a substantial number of real-world bugs.
DisFuzz's first key finding is the efficacy of its extended mutation space. Unlike traditional fuzzers that might focus solely on client inputs or fault injections, DisFuzz systematically mutates and combines regular events (like get and put operations), fault events (such as network partitions or node crashes), and crucially, timing intervals between these events. This multi-dimensional input space proved vital for uncovering bugs that are sensitive to specific event sequences and their precise temporal relationships, especially in scenarios like leader election or distributed transaction commit protocols. The ability to automatically generate and guide mutations across these diverse input types represents a significant leap in test case generation for complex distributed environments.
The second core finding is the development and effectiveness of symmetry-based feedback pruning. Recognizing that raw network message sequences, while indicative of state changes, contain significant redundancies, DisFuzz employs a sophisticated pruning technique. This method moves beyond ad-hoc or similarity-based pruning, which often misses interesting states or includes redundant ones, by leveraging the inherent symmetries found in distributed system states and their reflection in network messages. This intelligent pruning mechanism ensures that the fuzzer explores a broader range of truly unique and interesting system states, thereby guiding the fuzzing process more effectively towards bug-revealing execution paths.
The most impactful result demonstrating DisFuzz's capabilities is the discovery of 52 unique bugs across 10 widely-used distributed systems. These systems were written in various languages (C, C++, Go, Java) and included critical components like ClickHouse, RethinkDB, and two industry-grade Raft implementations, validating DisFuzz's broad applicability and effectiveness against real-world, production-level software. Of these 52 reported bugs, 28 were confirmed by developers, with 13 already fixed, and notably, four new CVEs (Common Vulnerabilities and Exposures) were assigned. This high rate of confirmed and fixed bugs, including critical security vulnerabilities, underscores DisFuzz's ability to uncover severe flaws that elude other testing methodologies. Furthermore, comparative evaluations showed that DisFuzz consistently found more bugs in all tested systems compared to existing tools, firmly establishing its superior performance and practical utility.
Technical Deep Dive
▶ Watch: Novel symmetry-based feedback pruning technique (6:00)
DisFuzz's technical sophistication lies in its dual approach: a comprehensive multi-dimensional input space and an intelligent, symmetry-aware feedback mechanism.
Multi-Dimensional Input Space
The core of DisFuzz's input generation capability is its ability to represent and mutate a rich set of event types. It categorizes events into regular events and fault events, with the critical addition of timing intervals.
- Regular Events: These mimic normal client interactions with the distributed system, such as
system initialization,getandputoperations for data manipulation, and other user-initiated commands. The talk mentions 11 distinct regular events. These are fundamental because many critical bugs only manifest when specific sequences of legitimate operations interact with system internals under stress. - Fault Events: These simulate adverse conditions that distributed systems must tolerate. DisFuzz supports 9 types of fault events, including
network message drop,network delaying, andnode crash. These events are crucial for testing fault tolerance and resilience. - Timing Intervals: This is a pivotal dimension. Distributed systems are highly sensitive to the relative timing of events. For instance, the outcome of a leader election in a consensus protocol can drastically change based on which node's messages arrive first or when a timeout occurs. DisFuzz systematically mutates these intervals to explore time-sensitive execution paths.
DisFuzz encodes all these events as a generalized tuple consisting of a timing interval, an event type, and parameters. The specific meaning of the parameters depends on the event type; for example, a node crash event would require a target node ID. This universal representation allows DisFuzz to generate and manipulate complex sequences of events. The system supports a total of 20 events, leveraging a CIO interception mechanism—a common I/O interception technique—to inject these events and observe their effects without requiring source code modifications.
Symmetry-Based Feedback Pruning
Effective fuzzing, especially in a blackbox context, requires robust feedback to guide the exploration of the state space. DisFuzz utilizes network messages as its primary feedback source, as they reliably capture significant state changes without requiring user annotations. However, raw network message sequences are inherently redundant, leading to inefficient state exploration if not properly managed.
The talk explicitly highlights the shortcomings of similarity-based pruning, a common technique. It illustrates two critical failures:
- Missing Interesting States: In a Raft example, two different message orderings (N1 sending
AppendEntriesvs. N3 initiating an election and sendingRequestVoteandVoteGrantedto N2) can lead to distinct states. However, if the overall message sequence is long, similarity-based pruning might deem them "similar" and explore only one, missing a crucial bug path. - Employing Redundant States: Another Raft example shows that whether node N1 or N2 initiates an election, the subsequent state (N1 sending messages to N2 and N3, or N2 sending messages to N1 and N3) might be functionally identical due to the system's inherent symmetry. Similarity-based pruning could treat these as distinct, leading to redundant exploration.
DisFuzz's core insight is that the symmetry in the system's state is reflected by the symmetry in the network messages. It leverages two types of symmetry to prune redundant feedback:
- Order Symmetry: This addresses situations where the order in which different nodes receive messages from a set of identical messages does not lead to a new logical state. DisFuzz tackles this by replacing global sequence numbers with a combination of a local sequence number plus node ID. This normalization ensures that the encoding state remains the same regardless of the arbitrary order in which messages are processed by different receiving nodes, as long as the content and local context are equivalent.
- Row Symmetry: This handles scenarios where the identity of the node performing an action is irrelevant to the resulting logical state. For example, if any node can initiate a leader election, the system's state might be symmetric regardless of which specific node (N1, N2, etc.) starts the process. DisFuzz achieves this by removing all node IDs from the network message encoding during the pruning process. This normalization ensures that the encoding remains identical even if different nodes perform the same logical action, preventing redundant state exploration.
DisFuzz Architecture and Components
The DisFuzz prototype operates through a structured workflow:
- Configuration: Users provide initial system setup details and specify event types.
- Generator: Based on the configuration and feedback, it generates sequences of multi-dimensional events (regular, fault, timing).
- Event Handler: Executes these generated event sequences on the target distributed system.
- Runtime Monitor: Collects network message sequences during execution, capturing critical state changes.
- Bug Checkers: Analyze the collected runtime information against predefined criteria. DisFuzz integrates five types of checkers:
- Memory Checker: Detects low-level memory errors, crucial for systems written in C/C++.
- Liveness Validation Checker: Ensures consistency and checks for properties like livelocks or deadlocks.
- Node Crash Checker: Identifies unintended node crashes that should not occur under specific event sequences.
- Availability Checker: Verifies whether the system remains functional and responsive after testing.
- Log Checker: Scans system logs for fatal errors or critical warnings.
- Bug Reports: If a bug is detected, a detailed report is generated for the user.
Workflow Enhancements
DisFuzz also incorporates features to streamline the fuzzing process and bug triaging:
- Checkpointing and Restoration: It utilizes the CIO tool (presumably a system-level I/O interception and snapshotting tool) to checkpoint and restore the system state. This significantly accelerates the fuzzing process by avoiding lengthy system initialization times for each test case.
- Bug Reproduction: Once a bug is found, DisFuzz supports a record and replay debugger. This allows developers to use the exact event sequence that triggered the bug to reproduce it offline, simplifying debugging and patch development.
This robust technical framework, combining intelligent input generation with sophisticated feedback pruning and comprehensive bug detection, enables DisFuzz to effectively uncover a wide array of vulnerabilities in complex distributed systems.
Demo / Proof of Concept
▶ Watch: DisFuzz system architecture and event implementation (7:15)
While the presentation did not include a live demonstration of the DisFuzz tool in action, the speaker highlighted a specific, compelling proof-of-concept scenario to illustrate its effectiveness. The talk referred to a figure showing a test case that triggered a significant bug in RethinkDB, a popular distributed database.
This particular bug manifested when a specific sequence of regular events was executed with precisely defined timing intervals. The consequence was severe: newly created databases within the RethinkDB instance became unusable. This example serves as a powerful testament to DisFuzz's core capabilities. It demonstrates that the tool's extended input space—which systematically mutates regular client requests and their temporal relationships—is crucial for uncovering subtle, yet critical, vulnerabilities that would likely be missed by fuzzers focusing solely on fault injection or generic input variations. The fact that this bug required both a specific sequence of operations and precise timing underscores the necessity of DisFuzz's multi-dimensional input approach to effectively test the intricate concurrency and timing dependencies inherent in distributed systems.
Defensive Implications
▶ Watch: Evaluation results and discovered bugs (9:00)
The findings presented by DisFuzz carry significant implications for developers, system administrators, and security professionals tasked with building and maintaining robust distributed systems. The consistent discovery of critical bugs, including CVEs, highlights the inadequacy of traditional testing methods for these complex architectures.
Firstly, the research underscores the paramount importance of comprehensive and multi-dimensional testing. Developers must recognize that distributed systems are not just susceptible to isolated component failures or malformed inputs, but also to intricate interactions between regular client operations, fault conditions, and, crucially, precise timing dependencies. Integrating fuzzing tools like DisFuzz, which can systematically explore these multi-dimensional input spaces, should become a standard practice in continuous integration/continuous deployment (CI/CD) pipelines. This proactive approach can catch bugs early in the development lifecycle, significantly reducing the cost and impact of vulnerabilities.
Secondly, system designers and developers should pay particular attention to concurrency and timing-sensitive logic. The bugs found by DisFuzz, especially those requiring specific timing intervals, indicate that many systems have brittle dependencies on event ordering or message arrival times. Defensive coding practices should aim for eventual consistency and design protocols that are robust to message reordering, delays, or drops. Explicitly handling race conditions and ensuring deterministic behavior under varying timing conditions is vital. The discovery of bugs sensitive to "order symmetry" and "row symmetry" suggests that developers might inadvertently assume specific message arrival orders or node identities, leading to vulnerabilities when these assumptions are violated.
Thirdly, the efficacy of DisFuzz's various bug checkers (memory, liveness, node crash, availability, log) points to areas where defenders can focus their efforts. Implementing robust memory safety checks, especially in C/C++ systems, is non-negotiable. Designing systems with explicit liveness and consistency guarantees that can be programmatically validated is crucial. Furthermore, comprehensive logging and monitoring capabilities are essential not only for post-mortem analysis but also for real-time detection of anomalies and fatal errors, as highlighted by DisFuzz's log checker.
Finally, the fact that DisFuzz found 52 bugs across 10 diverse systems suggests that these types of vulnerabilities are widespread. Organizations operating critical distributed infrastructure should consider adopting similar advanced fuzzing techniques or even leveraging open-source tools like DisFuzz if available, to audit their own systems. For developers of distributed systems, reviewing the types of bugs found by DisFuzz and the specific systems affected can provide valuable insights into common pitfalls and areas requiring more rigorous scrutiny in their own projects. The ability to checkpoint, restore, and reproduce bugs identified by such fuzzers is also a critical capability for rapid remediation.
Key Takeaways
- Distributed systems fuzzing requires multi-dimensional inputs: Effective bug discovery necessitates systematically mutating regular client events, fault injections, and critical timing intervals to uncover complex concurrency and timing-related vulnerabilities.
- Symmetry-based feedback pruning is essential for efficiency: Traditional similarity-based pruning often misses unique states or explores redundant ones; leveraging inherent symmetries in network messages (e.g., order and row symmetry) significantly improves fuzzing efficiency and coverage.
- Network messages are powerful, but noisy, feedback: Network traffic accurately reflects distributed system state changes, but intelligent pruning is crucial to make this feedback actionable and guide fuzzing effectively.
- DisFuzz is highly effective at finding real-world bugs: The tool successfully identified 52 bugs across 10 production-level distributed systems, leading to 28 confirmations, 13 fixes, and 4 new CVEs, demonstrating its practical value.
- Comprehensive testing must account for concurrency and fault tolerance: Distributed systems are inherently complex; robust testing strategies must include rigorous fault injection and exploration of time-sensitive event orderings.
- Advanced fuzzing tools improve critical infrastructure reliability: Specialized fuzzers like DisFuzz are vital for enhancing the security and robustness of the distributed systems that power modern digital services.
About the Speaker(s)
The paper "Blackbox Fuzzing of Distributed Systems with Multi-Dimensional Inputs and Symmetry-Based Feedback Pruning" was authored by Yonghao Zou and other collaborators. During the NDSS Symposium, the presentation was delivered by Ejos on behalf of the authors, due to visa issues preventing their attendance. Specific titles or affiliations for Yonghao Zou were not provided in the talk or its associated metadata.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
DisFuzz is legitimate systems research with a real technical contribution: the symmetry-based pruning formalism is non-obvious, the 52-bug result across 10 production systems is reproducible evidence rather than marketing copy, and the multi-dimensional input model (regular events + fault events + timing intervals as a unified tuple) is a cleaner abstraction than anything I've seen in prior distributed fuzzing work. The visa-proxy delivery is unfortunate but doesn't undercut the paper.
Heather Calloway (CISO) — WEAK
Technically credible research that found real bugs in real systems — 52 across 10 production distributed systems, four CVEs. But DisFuzz is a tool built for researchers and developers, and the talk never closes the gap to the people who actually own these risks institutionally.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025