Fuzzing the PHP Interpreter via Dataflow Fusion

Yuancheng Jiang

34th USENIX Security Symposium (USENIX Security '25) · Day 3 · Software Security 3: Fuzzing

Overview

This talk, "Fuzzing the PHP Interpreter via Dataflow Fusion," presented by Yuancheng Jiang at USENIX Security, introduces a novel and highly effective fuzzing methodology designed to uncover deep-seated memory corruption vulnerabilities within the PHP interpreter. PHP, powering over 70% of the world's websites and comprising an extensive codebase of over a million lines of C code, presents a significant and critical attack surface. While much security research has historically focused on application-level vulnerabilities like SQL injection, the underlying C interpreter's low-level memory errors have often been overlooked, despite being a common cause of critical security flaws as evidenced by recent CVEs.

Watch on YouTube · Slides

Visual summary for Fuzzing the PHP Interpreter via Dataflow Fusion by Yuancheng Jiang
Visual summary for Fuzzing the PHP Interpreter via Dataflow Fusion by Yuancheng Jiang

Key moments

  1. 0:00 Introduction to PHP security challenges and memory errors
  2. 2:30 Limitations of existing mutation-based and grammar-based fuzzing
  3. 3:18 Introducing input fusion and data flow as semantic representation
  4. 5:00 Concrete example of fusing two PHP programs' data flows
  5. 7:00 Flow Fusion's architecture and recursive fusion mechanism
  6. 8:00 Fuzzing setup and complementary strategies for robust testing
  7. 10:00 Significant bug findings and their impact on PHP security
  8. 11:27 Code coverage comparison and Flow Fusion's community integration

Fuzzing the PHP Interpreter via Dataflow Fusion

Speakers: Yuancheng Jiang

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=ba-dcQgfcoc

Overview

This talk, "Fuzzing the PHP Interpreter via Dataflow Fusion," presented by Yuancheng Jiang at USENIX Security, introduces a novel and highly effective fuzzing methodology designed to uncover deep-seated memory corruption vulnerabilities within the PHP interpreter. PHP, powering over 70% of the world's websites and comprising an extensive codebase of over a million lines of C code, presents a significant and critical attack surface. While much security research has historically focused on application-level vulnerabilities like SQL injection, the underlying C interpreter's low-level memory errors have often been overlooked, despite being a common cause of critical security flaws as evidenced by recent CVEs.

The core of the presented work is Flow Fusion, a new fuzzing approach that addresses the limitations of traditional mutation-based and grammar-based fuzzers. These existing methods often struggle to generate inputs with diverse code semantics or to reach deep logical paths within complex software like compilers and interpreters. Flow Fusion innovatively tackles this challenge by merging the dataflows of two or more existing program snippets, creating syntactically valid yet semantically novel inputs. This technique has demonstrated remarkable success, leading to the discovery of 158 unique bugs and significant security enhancements across the PHP ecosystem, ultimately becoming an official toolchain for PHP interpreter security.

Background

▶ Watch: Introduction to PHP security challenges and memory errors (0:00)

PHP's pervasive presence across the web makes its security paramount. As a scripting language, it abstracts away many low-level memory management details from developers. However, the PHP interpreter itself, written in C, is susceptible to the same classes of memory corruption bugs (e.g., buffer overflows, use-after-free) that plague any large C/C++ codebase. Historically, security efforts within the PHP ecosystem have largely concentrated on application-level vulnerabilities, such as SQL injection or cross-site scripting, leaving the foundational interpreter layer less scrutinized for memory safety issues. The speaker highlighted two recent CVEs caused by memory errors as a clear indication of this gap, underscoring the critical need to improve the low-level security of the PHP interpreter.

Fuzzing is a well-established and highly effective technique for discovering software bugs by generating a diverse range of inputs to drive the target software into unexpected states. However, the primary challenge in fuzzing, especially for complex targets like language interpreters, lies in automatically generating inputs that are both syntactically valid and semantically diverse enough to explore deep program logic. Traditional fuzzing methods fall into two main categories: mutation-based fuzzing and grammar-based fuzzing. Mutation-based fuzzers start with a corpus of existing inputs (seeds) and generate variants by applying random or guided mutations. Grammar-based fuzzers, on the other hand, use a formal grammar of the input language to generate syntactically valid inputs from scratch.

Both of these approaches suffer from common weaknesses. They often produce inputs with limited semantic diversity, even when maintaining high syntactic correctness. This limitation makes it exceedingly difficult for them to reach and test deep, complex logic paths within compilers or interpreters like PHP. The core problem, therefore, is how to generate inputs that exhibit truly diverse code semantics. The researchers observed that existing fuzzing methods typically operate on a single input at a time. This led to the fundamental question: what if two or more inputs could be combined or "fused" in a meaningful way to create new, semantically richer test cases? This concept of input fusion forms the bedrock of the Flow Fusion methodology. The insight guiding this fusion was the nature of typical fuzzing seeds: they are often minimized code snippets from bug reports or unit tests, designed to test a single feature or reproduce a bug. These "unit-level program snippets" are frequently simple and execute sequentially, with over 90% of PHP seeds executing without branches. This implies that control flow has minimal impact on their individual code semantics. Consequently, the researchers posited that data flow—the way values move from their points of definition to their points of use—would be a much more effective representation of a PHP program's code semantics for the purpose of fusion. Thus, the general input fusion problem was refined into the specific data flow fusion problem.

Key Findings

▶ Watch: Introducing input fusion and data flow as semantic representation (3:18)

Flow Fusion has yielded exceptional results, significantly advancing the security posture of the PHP interpreter. The research paper documented the detection of 158 unique bugs and unknown vulnerabilities. These bug reports have directly led to substantial security improvements across more than 80 distinct source files within the PHP codebase, with related patches modifying over 5,000 lines of code.

A breakdown of the reported bugs shows that the vast majority have either been fixed or confirmed by PHP developers, underscoring the high quality and criticality of the findings. The vulnerabilities covered a wide spectrum of Common Weakness Enumeration (CWE) categories, including critical memory errors such as stack buffer overflows, heap buffer overflows, and use-after-free vulnerabilities.

In terms of comparison with existing fuzzing approaches, Flow Fusion demonstrated superior performance. The official PHP internal test suite, maintained by PHP developers for CI/CD, achieved 78% code coverage. Surprisingly, this was higher than well-known general-purpose fuzzers like AFL++ (which adds advanced instrumentation and feedback) and Polyglot (a multi-language fuzzer that translates inputs to a uniform IR for mutation). When Flow Fusion was evaluated based on this strong baseline, Flow Fusion 2 (an iteration of the tool) improved code coverage by approximately 3% after just 24 hours of fuzzing. This additional coverage, while seemingly small, represents deep, hard-to-reach code paths that traditional fuzzers missed, and where critical bugs often reside.

Beyond its initial academic impact, Flow Fusion has achieved significant community integration and continuous impact. It has been adopted as an official toolchain by the PHP project, highlighting its outstanding effectiveness in discovering bugs. Flow Fusion operates as a living project, continuously fuzzing nightly builds of PHP interpreters, enabling it to test new features and automatically update its corpus with new test cases from the official test suite. This ongoing process ensures that its impact is not limited to the paper's submission or acceptance; the project actively reports new PHP memory errors almost weekly, directly contributing to the interpreter's ongoing security improvements.

Technical Deep Dive

▶ Watch: Flow Fusion's architecture and recursive fusion mechanism (7:00)

The core innovation of Flow Fusion lies in its data flow fusion methodology, which creates new, semantically diverse inputs by intelligently merging existing ones. The process begins with a corpus of PHP programs, typically PHPT files, which are unit-level test cases.

The definition of data flow is central: it describes how values move from their points of definition to their points of use within a program. Given two PHP programs (Test Case A and Test Case B), Flow Fusion extracts their respective data flows by identifying variables and tracking their usage.

Consider a concrete example:

  • Test Case A: A PHP program that performs base64 encoding, defining and using several variables.
  • Test Case B: Another PHP program, perhaps involving a DOM operation.

A naive approach might simply concatenate these two programs, but this yields no significant difference from executing them separately. Flow Fusion, however, identifies points in the data flow where variables can be meaningfully connected. For instance, an output variable from Test Case A (e.g., the base64-encoded string) can be connected as an input to a variable in Test Case B (e.g., a string argument for a DOM parsing function).

The talk illustrates this with two data flow fusion options:

  1. Connect the last variable in Test Case A to the second variable in Test Case B. This creates a new test case where the output of the base64 encoding directly feeds into the DOM operation, allowing fuzzing of "DOM operation on base64 object."
  2. Connect the second variable in Test Case A to the last variable in Test Case B. This demonstrates the flexibility of the approach to create different semantic interactions.

The order of test cases can also be reversed (e.g., Test Case B followed by Test Case A) to explore even more combinations. This capability to generate random data flow combinations is crucial for creating variant code semantics. The speaker specifically mentioned that this approach helped trigger a 20-year-old memory error in the PHP interpreter, a testament to its ability to uncover long-standing, deep-seated vulnerabilities.

The overall Flow Fusion overview involves:

  1. Corpus Collection: Gathering PHPT files as the initial set of seeds.
  2. Execution in PHP Engine: Running fused inputs through the PHP engine for fuzzing.
  3. Recursive Fusion: A key mechanism where any fused input that leads to new code coverage or discovers a new bug is added back to the corpus. This recursive process enables the generation of even more complex semantic fusions, combining more than just two initial seed programs.

Beyond the core data flow fusion, Flow Fusion incorporates several complementary strategies to enhance its effectiveness:

  1. Test Mutations: Prior to the data flow fusion process itself, Flow Fusion applies various mutations to the original test cases. These mutations are designed to introduce semantic variations while rigorously maintaining syntactic validity. Examples include exchanging operators (e.g., + to -) or assigning special values (e.g., null, 0, "") to variables. This broadens the initial pool of semantic diversity before fusion.
  1. Interface Fuzzing: After generating the fused test cases, Flow Fusion injects random function calls to internal PHP APIs. Crucially, it reuses existing variables from the fused context as arguments for these internal function calls. This strategy is highly effective in testing the robustness and error handling of PHP's internal APIs under various, often unexpected, input conditions derived from the fusion process.
  1. Environment Crossover: Flow Fusion doesn't limit its merging to just PHP program code. It also fuses their operational configurations. This involves randomly introducing other valid configurations collected from the entire PHP test suite. Examples include adjusting memory limits, modifying JIT settings, or changing OPcache modes. By fuzzing the interpreter under diverse environmental settings, Flow Fusion can uncover bugs that manifest only under specific configurations, further expanding the attack surface explored.

This multi-faceted approach, combining intelligent data flow merging with targeted mutation, interface fuzzing, and environment crossover, allows Flow Fusion to systematically explore a vast and semantically rich input space, far beyond what traditional fuzzers can achieve.

Demo / Proof of Concept

▶ Watch: Fuzzing setup and complementary strategies for robust testing (8:00)

While the talk did not feature a live, interactive demonstration, the speaker effectively illustrated the core mechanism of Flow Fusion through a detailed conceptual example. The process of taking two distinct PHP programs – one for base64 encoding (Test Case A) and another for a DOM operation (Test Case B) – and demonstrating how their data flows could be merged served as a powerful proof of concept.

The explanation clearly showed how connecting the output of Test Case A to the input of Test Case B could create a new, semantically distinct test case designed to "fuzz the DOM operation on a base64 object." This concrete illustration highlighted how Flow Fusion generates inputs that combine functionalities in novel ways, leading to deep code paths. Furthermore, the revelation that this method successfully triggered a 20-year-old memory error in the PHP interpreter provides compelling evidence of its practical efficacy and ability to uncover long-standing, critical vulnerabilities that have eluded previous detection methods. This conceptual walkthrough, backed by the significant bug discovery statistics, functions as the primary demonstration of Flow Fusion's capabilities.

Defensive Implications

▶ Watch: Code coverage comparison and Flow Fusion's community integration (11:27)

The work presented on Flow Fusion carries substantial implications for defenders at various levels, from PHP core developers to system administrators and security researchers.

For PHP core developers and contributors, Flow Fusion is not just a research project but an actively integrated and official toolchain. This means that:

  • Continuous Integration/Continuous Deployment (CI/CD) Integration: Developers should ensure Flow Fusion is deeply integrated into PHP's CI/CD pipelines to catch newly introduced memory errors as quickly as possible.
  • Bug Monitoring and Prioritization: Active monitoring of Flow Fusion's bug reports (which are generated almost weekly) is crucial. These reports often point to critical memory corruption vulnerabilities that require immediate attention and patching. The fact that the tool has impacted over 80 source files and led to 5,000+ lines of code changes underscores its importance.
  • Adoption of Fuzzing Best Practices: The success of Flow Fusion highlights the value of semantic-aware fuzzing. PHP developers can learn from this methodology to enhance their internal testing strategies and potentially adapt similar techniques for other complex components.

For system administrators and users running PHP-powered applications:

  • Prompt Patching and Updates: Given that Flow Fusion continuously uncovers critical memory errors, it is imperative to keep PHP installations up-to-date with the latest security patches. These patches directly address vulnerabilities that could be exploited by attackers.
  • Understanding Risk: Recognize that even widely used and mature software like PHP can harbor deep-seated memory safety issues. Relying solely on application-level security measures is insufficient; the underlying interpreter's security is equally vital.
  • Configuration Security: The "environment crossover" strategy used by Flow Fusion reminds defenders that PHP's configuration settings (e.g., memory_limit, JIT settings) can influence vulnerability exposure. Ensuring secure and robust configurations is part of a comprehensive defense strategy.

For security researchers and practitioners:

  • Methodology Transfer: The data flow fusion methodology is highly transferable. Researchers working on the security of other language interpreters (e.g., CPython for Python, V8 for JavaScript) or compilers should explore adapting Flow Fusion's principles. The talk explicitly mentions future work exploring this adaptation.
  • Semantic-Aware Fuzzing: Flow Fusion provides a strong case study for the effectiveness of fuzzing techniques that prioritize semantic diversity over mere syntactic correctness or random mutations. This paradigm shift can inspire new approaches in other complex software domains.
  • Addressing Overlooked Attack Surfaces: The work serves as a reminder that low-level memory safety in the foundation of high-level languages is a critical, often overlooked, attack surface that deserves continuous research and tooling.

In essence, Flow Fusion provides a powerful tool and a compelling methodology for proactively securing a foundational component of the internet, urging all stakeholders to prioritize robust, deep-level security testing.

Key Takeaways

  • Novel Fuzzing Methodology: Flow Fusion introduces data flow fusion, a new fuzzing technique that merges the code semantics of two or more existing program snippets to generate inputs with high semantic diversity, effectively overcoming limitations of traditional fuzzers for complex interpreters.
  • Significant Bug Discovery: The tool, Flow Fusion, successfully detected 158 unique bugs in the PHP interpreter, including critical memory corruption vulnerabilities like heap/stack buffer overflows and use-after-free flaws, leading to over 5,000 lines of code modified across 80+ source files.
  • Enhanced Code Coverage: Flow Fusion demonstrated superior performance, improving code coverage by approximately 3% over PHP's already robust official test suite, reaching deep and previously unexplored logical paths within the interpreter.
  • Comprehensive Fuzzing Strategy: Beyond data flow fusion, Flow Fusion integrates complementary strategies such as test mutations, interface fuzzing (injecting internal API calls with fused context variables), and environment crossover (merging configurations like memory limits and JIT settings) to maximize bug discovery.
  • Official Toolchain and Continuous Impact: Flow Fusion has been adopted as an official toolchain by the PHP project, providing continuous bug reporting and security improvements to nightly builds, demonstrating its practical value and ongoing commitment to interpreter security.
  • Broad Applicability: The data flow fusion methodology shows strong potential for adaptation to other language interpreter implementations, such as CPython for Python or JavaScript engines, opening avenues for future research and security enhancements across the software landscape.

About the Speaker(s)

The talk was presented by Yuancheng Jiang, who is credited as the primary speaker. The work itself is a joint effort with Chuchi Wan, Jiao Manuel Rhdan, and Junkai. The transcript and metadata provided do not offer further biographical details about Yuancheng Jiang or his collaborators beyond their names and their affiliation with this research.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Solid academic research with a genuinely clever core idea — dataflow fusion as a semantic diversity engine for interpreter fuzzing. 158 bugs, official PHP adoption, and a 20-year-old memory error unearthed are not numbers you manufacture. The methodology is novel enough to matter and transferable enough to be useful beyond PHP.

Heather Calloway (CISO) — PASS

Technically credible fuzzing research with real results — 158 bugs, official PHP adoption, a 20-year-old memory error surfaced. But this is deep interpreter security research with no governance angle, no institutional accountability dimension, and no actionable path for security leaders or defenders outside the PHP core team.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)