Encarsia: Evaluating CPU Fuzzers via Automatic Bug Injection
Matej Bölcskei
34th USENIX Security Symposium (USENIX Security '25) · Day 2 · Hardware Security 1: Microarchitectures
Overview
In the realm of CPU security and reliability, hardware fuzzing has emerged as an indispensable technique for uncovering subtle yet critical design flaws. While numerous scientific publications frequently report impressive bug counts and high coverage metrics, a fundamental question persists: are these fuzzers truly as effective as their creators claim, or are their reported successes often misleading? The talk "Encarsia: Evaluating CPU Fuzzers via Automatic Bug Injection," presented by Matej Bölcskei, addresses this critical gap by introducing a novel, systematic approach to fuzzer evaluation.

Key moments
- 0:00 Introduction: Evaluating CPU fuzzers with injected bugs
- 1:58 Critique of current fuzzer evaluation metrics
- 3:14 Encarsia: A new tool for reliable fuzzer evaluation
- 3:35 Methodology: Injecting realistic CPU bugs (signal mixups, conditionals)
- 4:20 Verifying injected bugs propagate to architectural state
- 5:00 Introducing Encorpus: A public dataset of injected CPU bugs
- 5:25 Fuzzer evaluation findings: program generation, specific blind spots
- 8:00 Example of a complex injected bug challenging fuzzers
Encarsia: Evaluating CPU Fuzzers via Automatic Bug Injection
Speakers: Matej Bölcskei
Conference: USENIX Security
YouTube: https://www.youtube.com/watch?v=IehArT_phek
Overview
In the realm of CPU security and reliability, hardware fuzzing has emerged as an indispensable technique for uncovering subtle yet critical design flaws. While numerous scientific publications frequently report impressive bug counts and high coverage metrics, a fundamental question persists: are these fuzzers truly as effective as their creators claim, or are their reported successes often misleading? The talk "Encarsia: Evaluating CPU Fuzzers via Automatic Bug Injection," presented by Matej Bölcskei, addresses this critical gap by introducing a novel, systematic approach to fuzzer evaluation.
Encarsia is a project designed to provide a robust and reproducible framework for assessing the true efficacy of CPU fuzzers. Instead of relying on abstract metrics or opportunistic bug discoveries, Encarsia injects synthetic, yet realistic, bugs directly into CPU designs and then rigorously tests whether existing fuzzers can detect them. This methodology offers a crucial shift from anecdotal evidence to empirical validation, enabling researchers and developers to differentiate meaningful innovation in fuzzing techniques from inflated claims.
The significance of Encarsia lies in its ability to bring much-needed transparency and rigor to the field of hardware fuzzer evaluation. By establishing a standardized, publicly available benchmark of injected bugs, Encarsia empowers the community to objectively compare different fuzzing strategies, identify their strengths and weaknesses, and ultimately drive the development of more powerful and comprehensive CPU testing tools. This work is pivotal for enhancing the security and resilience of modern CPU architectures, which form the bedrock of all computing systems.
Background
▶ Watch: Introduction: Evaluating CPU fuzzers with injected bugs (0:00)
The landscape of CPU testing has been significantly shaped by the advent of hardware fuzzing. At its core, a hardware fuzzer operates as an automated test bench. It employs a stimulus generator to produce randomized programs, which are then executed on the design under test (DUT), typically in a simulation environment. To identify bugs, these same programs are concurrently run on a golden reference model, often an instruction set simulator (ISS) that is presumed to be functionally correct. Any discrepancy between the architectural state of the fuzzed CPU and the golden model signals a violation of the Instruction Set Architecture (ISA) and, by extension, a bug. Many advanced fuzzers also incorporate a feedback mechanism that leverages insights from previous test runs to refine and guide the generation of subsequent inputs, aiming for more efficient exploration of the CPU's state space.
Despite these sophisticated architectures, the evaluation of fuzzer effectiveness has historically been fraught with challenges. Fuzzer creators frequently highlight impressive statistics, such as high coverage metrics across multiplexers, registers, and other CPU components, or boast about the sheer number of bugs discovered—citing figures like "nine bugs, 16 bugs, 37 bugs." These numbers are often used to suggest that their particular approach, whether it be a novel coverage metric, a unique bug detection mechanism, or a specialized input generation technique, possesses unparalleled effectiveness.
However, these seemingly strong narratives often rest on flawed foundations. Coverage, while intuitively appealing, has been repeatedly demonstrated to be an unreliable measure of fuzzer efficacy. The primary reason is that coverage metrics do not account for the uneven distribution of bugs within a CPU's state space. A fuzzer that explores a smaller portion of the overall state space but targets regions with a higher density of bugs might ultimately find more issues than a fuzzer achieving broad but shallow coverage. Furthermore, coverage completely overlooks the crucial effort required for a bug, once triggered locally, to propagate to an architecturally observable location. Simply hitting a buggy component does not guarantee detection; the bug's effects must manifest in a way that causes a discrepancy with the golden model.
Similarly, relying on natural bug discovery as a primary evaluation metric can be misleading. Many reported bugs, especially in newly fuzzed CPUs, turn out to be "low-hanging fruit" that had simply gone unnoticed because the specific CPU had never been subjected to fuzzing before. Any standard fuzzer might have found them, making it difficult to attribute their discovery to genuine innovation in the fuzzer's design. This leads to conflicting claims within the field, where, for instance, Defosio might argue that coverage is cheap and effective for detecting deep corner cases, while Cascade counters that coverage is expensive and largely ineffective. Without a sound and objective evaluation framework, distinguishing meaningful advancements from empty promises remains an intractable problem, hindering progress in the development of truly robust CPU fuzzers. Encarsia steps in to resolve this by providing a reliable and reproducible method for evaluating fuzzers against a known, controlled set of vulnerabilities.
Key Findings
▶ Watch: Encarsia: A new tool for reliable fuzzer evaluation (3:14)
Encarsia's core contribution lies in its systematic approach to evaluating CPU fuzzers by injecting synthetic, yet realistic, bugs into actual CPU designs. This methodology bypasses the limitations of traditional coverage metrics and opportunistic bug discovery, offering a precise and reproducible means of assessment.
The project began by analyzing over 1,600 pull requests related to CPU designs to understand the nature of real-world hardware bugs. This extensive survey revealed that virtually all observed bugs fall into just two fundamental categories:
- Signal Mixups: These involve confusing similarly named signals within the RTL, leading to incorrect data paths or control flows.
- Broken Conditionals: These arise from mishandling specific corner cases in control logic, often involving multiplexers or state machine transitions, resulting in incorrect conditional behavior.
To enable the injection of these bugs into diverse RTL designs, Encarsia utilizes an intermediate representation of cells interconnected by wires. This abstraction allows for uniform handling of various high-level constructs. Signal mixups are injected by simply rerouting wires, while broken conditionals, represented as multiplexer trees, are manipulated by either removing inputs or altering their select wires. A crucial step in this process is formal verification of each injected bug. Unlike traditional formal verification that aims to prove overall correctness, Encarsia's targeted verification confirms two things: first, that the injected bug induces a local change at the transformed signal, and second, that this local change propagates to an architecturally visible state. This guided verification dramatically reduces the state space, allowing it to scale to large, out-of-order CPUs and ensuring that only effective, detectable bugs are included in the evaluation set.
The culmination of this injection methodology is Encorpus, an exhaustive, publicly available set of injected bugs spanning diverse CPU architectures. Encorpus provides a static ground truth, enabling fair and reproducible evaluation of fuzzers. Using Encorpus, Encarsia set out to answer which fuzzing techniques are truly effective.
The evaluation process involved isolating individual fuzzer techniques:
- Detection Granularity: Comparing Processor Fuzz (checks for bugs at every instruction) with Defuzz RTL (checks only at the end of program execution) revealed no difference in bug detection. Both found the exact same bugs, indicating that the granularity of detection does not inherently yield better results.
- Coverage Guidance: When coverage guidance was reintroduced to Defuzz RTL, the results remained unchanged. Both fuzzers detected the same bugs as they did without coverage, suggesting that coverage guidance, at least in the implementation used in Defuzz RTL, does not effectively guide fuzzers toward new bug discoveries.
- Input Generation: In contrast, testing Cascade, which employs a distinct input generation technique, showed a significant difference in the bugs found. This outcome strongly suggests that program generation is the primary driver of bug discovery, rather than sophisticated detection mechanisms or simple coverage metrics.
Beyond these broad trends, Encarsia's analytical capabilities allow for deep insights into fuzzer weaknesses. For example, Defuzz RTL was found to consistently miss bugs related to integer division or multiplication. Further investigation revealed a misconfiguration: Defuzz RTL inadvertently omitted the RISC-V MM extension from its instruction set, meaning these modules were never exercised. This highlights a broader issue where all three validated fuzzers ignored large parts of the ISA. Cascade, on the other hand, overlooked bugs affecting floating-point division or square root calculations. This was attributed to an overly aggressive and coarse-grained bug filtering mechanism, which, to avoid rediscovering the same issue, suppressed entire instructions, inadvertently preventing the exposure of new, distinct bugs affecting those instructions.
Crucially, Encarsia demonstrates that its simple syntactic transformations can yield highly complex and challenging bugs. The talk highlights Broken Conditional #3, injected into the Boom CPU, which affects the ring buffer handling write-back from the FPU during contention for general-purpose registers. This bug causes the write pointer to advance irrespective of the buffer being full, leading to data corruption. Triggering this bug requires precise timing of multiple FPU instructions and sustained contention in the write-back stage—conditions so intricate that none of the three evaluated fuzzers detected it. This exemplifies a general trend: fuzzers consistently struggle with bugs stemming from microarchitectural corner cases, underscoring a significant blind spot in current fuzzing methodologies.
Technical Deep Dive
▶ Watch: Verifying injected bugs propagate to architectural state (4:20)
Encarsia's technical sophistication lies in its ability to systematically inject realistic bugs into CPU RTL and formally verify their detectability. This process is crucial for creating a reliable benchmark like Encorpus.
The bug injection process begins with the transformation of the CPU's Register Transfer Level (RTL) code into an intermediate representation. This representation models the CPU as a network of interconnected cells and wires, abstracting away the high-level constructs of hardware description languages. This uniform representation is key to enabling consistent and automated bug injection across diverse CPU designs.
The two primary categories of bugs identified from the 1,600 pull request survey are injected as follows:
- Signal Mixups: These bugs involve confusing two distinct signals that might have similar names or be logically related. In the intermediate representation, injecting a signal mixup is achieved by simply rerouting a wire from its intended source to an incorrect, but plausible, alternative source. This mimics real-world design errors where a designer might mistakenly connect the wrong signal to a component's input.
- Broken Conditionals: These bugs typically manifest in control logic, often involving multiplexer trees that select data paths based on specific conditions. Encarsia can break these conditionals in two ways:
- Removing inputs: This might involve disconnecting one of the data inputs to a multiplexer, effectively forcing a default or undefined value in certain conditions.
- Manipulating select wires: By altering the logic that drives a multiplexer's select wire, Encarsia can cause the multiplexer to choose the wrong input under specific conditions, leading to incorrect data or control flow.
After a bug is injected, a critical and distinguishing step is its formal verification. Unlike traditional formal verification, which attempts to prove the absence of all bugs or verify high-level security properties, Encarsia's verification is highly targeted and efficient. Its purpose is to eliminate "ineffective transformations" – injected changes that might not actually lead to an observable architectural bug, thus skewing evaluation results. The verification process involves two distinct phases:
- Local Change Confirmation: First, Encarsia formally confirms that the injected transformation indeed induces a local change at the transformed signal. This ensures that the modification to the RTL has a tangible effect within the CPU's internal logic.
- Architectural Propagation Verification: Second, and more importantly, it verifies that this local change propagates to an architecturally visible state. This means proving that the internal alteration can eventually cause a discrepancy in an architecturally defined register, memory location, or other observable output, which would then be detected by a fuzzer comparing against a golden model.
This targeted strategy dramatically reduces the state space that needs to be explored during formal verification, making it scalable even to large, complex out-of-order CPUs. The result is a curated set of truly detectable bugs ready for fuzzer evaluation.
The product of this rigorous injection and verification process is Encorpus, a publicly available, exhaustive set of these synthetic bugs spanning diverse CPUs. Encorpus serves as a static ground truth, providing a fair and reproducible benchmark for fuzzer evaluation.
Encarsia's evaluation of fuzzers leveraged this ground truth to dissect their effectiveness:
- Processor Fuzz vs. Defuzz RTL: The comparison between these two fuzzers, where Processor Fuzz checks for bugs at every instruction and Defuzz RTL waits until the end of program execution, aimed to isolate the impact of detection granularity. The finding that both detected the exact same bugs indicated that finer-grained detection, in this context, did not yield better results.
- Coverage Guidance Reintroduction: Reintroducing coverage guidance to Defuzz RTL yielded no change in bug detection, suggesting that the particular implementation of coverage guidance in Defuzz RTL was ineffective in discovering new bugs.
- Cascade's Input Generation: In contrast, Cascade, which employs a distinct and presumably more sophisticated input generation technique, showed significant differences in bug discovery. This strongly highlighted that the method of generating test programs is paramount to a fuzzer's success.
A prime example of the complexity of bugs Encarsia can inject, and the challenges they pose to fuzzers, is Broken Conditional #3 injected into the Boom CPU. This bug resides in a ring buffer that manages write-back operations from the Floating Point Unit (FPU), specifically when there's contention for general-purpose registers from other execution units. The bug causes the write pointer of this buffer to advance regardless of whether the buffer is full. Over time, this can lead to the pointer overwriting an entry that still holds valid data, thus corrupting an earlier FPU result.
Triggering Broken Conditional #3 presents two formidable challenges:
- Precise Timing: Multiple FPU instructions must be timed such that when the result of a later instruction is written back, the head pointer of the ring buffer is pointing exactly to the register holding an earlier result. Only then is the earlier value destroyed.
- Sustained Contention: The FPU write-back buffer is only utilized when the general-purpose registers are occupied by results from other execution units. If there is no contention, the FPU bypasses the buffer and writes results directly to the registers, avoiding the corruption. Therefore, triggering the bug requires sustained contention in the write-back stage.
These combined conditions make Broken Conditional #3 extraordinarily difficult to trigger, a fact underscored by its detection by none of the three evaluated fuzzers. This specific example, and others like it in Encorpus, demonstrate that microarchitectural corner cases remain a significant blind spot for current fuzzer designs, despite the simplicity of their initial syntactic transformation.
Demo / Proof of Concept
▶ Watch: Introducing Encorpus: A public dataset of injected CPU bugs (5:00)
While the talk does not describe a live demonstration of Encarsia in action during the presentation, the project itself serves as a robust proof of concept for its methodology. The core deliverable, Encorpus, is described as a publicly available and exhaustive set of injected bugs spanning diverse CPUs. This artifact provides a concrete and reproducible ground truth for fuzzer evaluation, enabling others to utilize Encarsia's findings and methodology. The existence of Encorpus and the accompanying paper, which encourages users to "try out the artifacts," strongly implies that the tools and injected bug sets are available for practical use and further research, functioning as a continuous "demo" for the community.
Defensive Implications
▶ Watch: Example of a complex injected bug challenging fuzzers (8:00)
Encarsia's findings provide critical insights for both CPU designers and fuzzer developers, offering clear defensive implications to enhance hardware security and reliability.
For CPU designers, the comprehensive survey of 1,600 pull requests revealing that most bugs fall into "signal mixups" and "broken conditionals" is a direct call to action. Designers should prioritize rigorous review and verification processes specifically targeting these categories during the RTL design phase. Enhanced static analysis tools could be developed to detect potential signal mixups (e.g., similarly named but functionally distinct signals) or identify complex conditional logic that might harbor subtle corner cases. Furthermore, the existence of bugs like Broken Conditional #3 in the Boom CPU highlights the importance of deeply understanding and thoroughly testing microarchitectural interactions, especially those involving resource contention, complex timing, and shared buffers. Designers should consider these scenarios as high-risk areas requiring extra scrutiny.
For fuzzer developers, Encarsia offers a roadmap for improving fuzzer effectiveness:
- Prioritize Program Generation: The finding that "program generation drives bug discovery" is paramount. Fuzzer developers should invest more effort in designing sophisticated stimulus generators that can produce diverse, complex, and architecturally meaningful test sequences, rather than relying solely on detection granularity or basic coverage guidance. This includes generating sequences that can induce microarchitectural corner cases.
- Ensure Complete ISA Coverage: The discovery that validated fuzzers (like Defuzz RTL omitting the RISC-V MM extension) ignore large parts of the ISA is a significant vulnerability. Fuzzer developers must ensure their tools fully exercise all instructions, extensions, and functional units specified by the target CPU's ISA. This requires careful configuration and validation of the fuzzer's instruction set model.
- Refine Bug Filtering Mechanisms: Cascade's issue with aggressive, coarse-grained bug filtering, which prevented the discovery of new bugs affecting already-flagged instructions, points to a need for more intelligent and nuanced filtering. Filtering mechanisms should be designed to suppress specific instances of a bug, not entire instruction types, allowing for the discovery of distinct vulnerabilities that might manifest through the same instruction but under different conditions.
- Target Microarchitectural Corner Cases: The consistent struggle of fuzzers to detect bugs stemming from microarchitectural corner cases, like the FPU write-back buffer bug, indicates a critical blind spot. Future fuzzer designs should incorporate strategies specifically aimed at triggering complex timing conditions, resource contention, and intricate state interactions that define these difficult-to-find bugs. This might involve more stateful fuzzer designs or techniques that reason about microarchitectural hazards.
- Utilize Encarsia and Encorpus for Evaluation: Fuzzer developers should adopt Encarsia's methodology and leverage the publicly available Encorpus benchmark for robust, reproducible, and objective evaluation of their tools. This provides a standardized way to compare against existing fuzzers and demonstrate genuine improvements in bug detection capabilities.
Ultimately, Encarsia advocates for a shift towards more rigorous and evidence-based fuzzer development and evaluation, directly contributing to the creation of more secure and reliable CPU hardware.
Key Takeaways
- Traditional fuzzer evaluation metrics are flawed: Coverage metrics fail to account for uneven bug distribution and propagation effort, while natural bug discovery often reports "low-hanging fruit" and lacks reproducibility, leading to conflicting claims in the field.
- Encarsia provides a robust evaluation framework: By injecting synthetic, yet realistic, bugs into CPU designs and formally verifying their architectural impact, Encarsia offers a reliable and reproducible method for assessing fuzzer effectiveness.
- Bugs primarily fall into two categories: A survey of 1,600 pull requests revealed that most hardware bugs are either "signal mixups" (rerouting wires) or "broken conditionals" (manipulating multiplexer logic), which can be systematically injected.
- Program generation is key to bug discovery: Evaluation with Encarsia showed that sophisticated input generation techniques, exemplified by Cascade, are significantly more effective at finding bugs than mere detection granularity or basic coverage guidance.
- Current fuzzers have significant blind spots: Evaluated fuzzers often ignore large parts of the ISA (e.g., omitted RISC-V extensions), employ overly aggressive bug filtering that hinders new discoveries, and consistently struggle to detect complex microarchitectural corner cases.
- Encorpus serves as a vital benchmark: The publicly available Encorpus provides an exhaustive set of formally verified injected bugs, offering a static ground truth for fair, objective, and reproducible evaluation of CPU fuzzers.
About the Speaker(s)
Matej Bölcskei is the speaker behind the "Encarsia: Evaluating CPU Fuzzers via Automatic Bug Injection" talk. While the transcript does not provide specific details about his title or company affiliation, his work on Encarsia demonstrates expertise in hardware security, CPU architecture, formal verification, and fuzzing methodologies. His research focuses on developing rigorous techniques to assess the effectiveness of hardware testing tools and uncover vulnerabilities in complex CPU designs.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Bölcskei addresses a real meta-problem in hardware security — that CPU fuzzer evaluations are largely unverifiable theater — with a technically sound methodology: formal-verification-guided bug injection producing a reproducible ground-truth corpus. The key finding that program generation dominates detection granularity and coverage guidance is the kind of result that should recalibrate how the field allocates research effort.
Heather Calloway (CISO) — PASS
Rigorous academic work on CPU fuzzer evaluation methodology — technically sound, formally grounded, and useful to hardware security researchers. Outside my lane entirely. No governance angle, no operator path, no institutional relevance.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)