DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data

Dorde Popovic

34th USENIX Security Symposium (USENIX Security '25) · Day 3 · ML and AI Security 3: Backdoors, Poisoning, Unlearning

Overview

This article delves into FlowFusion, a groundbreaking automatic fuzzing framework designed to uncover memory errors within the PHP interpreter. Developed by a team from the National University of Singapore, FlowFusion represents the first tailored approach to specifically target these critical vulnerabilities in PHP's underlying C codebase, an area largely overlooked by prior security research. Given PHP's pervasive role, powering over 70% of websites globally, the security of its interpreter is paramount to the confidentiality, integrity, and availability of countless web services.

Read the paper · Download the PDF (PDF) · Slides

Paper abstract

PHP, a dominant scripting language in web development, powers a vast range of websites, from personal blogs to major platforms. While existing research primarily focuses on PHP application-level security issues like code injection, memory errors within the PHP interpreter have been largely overlooked. These memory errors, prevalent due to the PHP interpreter's extensive C codebase, pose significant risks to the confidentiality, integrity, and availability of PHP servers. This paper introduces FlowFusion, the first automatic fuzzing framework to detect memory errors in the PHP interpreter. FlowFusion leverages dataflow as an efficient representation of test cases maintained by PHP developers, merging two or more test cases to produce fused test cases with more complex code semantics. Moreover, FlowFusion employs strategies such as test mutation, interface fuzzing, and environment crossover to increase bug finding. In our evaluation, FlowFusion found 158 unknown bugs in the PHP interpreter, with 125 fixed and 11 confirmed. Comparing FlowFusion against the official test suite and a naive test concatenation approach, FlowFusion can detect new bugs that these methods miss, while also achieving greater code coverage. FlowFusion also outperformed state-of-the-art fuzzers AFL++ and Polyglot, covering 24% more lines of code after 24 hours of fuzzing. FlowFusion has gained wide recognition among PHP developers and is now integrated into the official PHP toolchain.

Visual summary for DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data by Dorde Popovic
Visual summary for DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data by Dorde Popovic

Fuzzing the PHP Interpreter via Dataflow Fusion

Speakers: Yuancheng Jiang, Chuqi Zhang, Bonan Ruan, Jiahao Liu, Manuel Rigger, Roland H. C. Yap, Zhenkai Liang (National University of Singapore)

Conference: USENIX Security 2025

YouTube: Not Applicable (Peer-reviewed paper). Paper page: https://www.usenix.org/conference/usenixsecurity25/presentation/jiang-yuancheng

Overview

This article delves into FlowFusion, a groundbreaking automatic fuzzing framework designed to uncover memory errors within the PHP interpreter. Developed by a team from the National University of Singapore, FlowFusion represents the first tailored approach to specifically target these critical vulnerabilities in PHP's underlying C codebase, an area largely overlooked by prior security research. Given PHP's pervasive role, powering over 70% of websites globally, the security of its interpreter is paramount to the confidentiality, integrity, and availability of countless web services.

The significance of FlowFusion lies in its innovative use of dataflow fusion, a technique that merges existing, high-quality PHP test cases by interleaving their dataflows to generate new, semantically complex test cases. This approach, complemented by strategies like test mutation, interface fuzzing, and environment crossover, has proven remarkably effective. The framework has identified an impressive 158 previously unknown bugs in the PHP interpreter, with 125 already fixed and 11 confirmed, including a 20-year-old use-after-free vulnerability.

FlowFusion's success has garnered significant recognition, leading to its integration into the official PHP toolchain. This integration underscores the framework's practical utility and its potential to profoundly enhance the security and robustness of the PHP ecosystem, moving beyond application-level security concerns to address fundamental memory safety issues at the interpreter level.

Background

The PHP interpreter, predominantly written in C, forms the backbone of a vast number of web applications. It comprises three main components: the Zend engine, which handles bytecode compilation, execution, and memory management; core modules, providing built-in and bundled extensions (e.g., session, sqlite3); and main functions, which manage interpreter initialization and the primary execution loop (Section 2). The choice of C, a memory-unsafe language, inherently exposes the interpreter to various memory errors, such as buffer overflows and use-after-free vulnerabilities.

Despite PHP's widespread use, security research has historically concentrated on application-level issues like SQL injection or file inclusion. Memory errors within the interpreter itself have received limited attention, with no prior fuzzing approach specifically designed for this target. A study of PHP's GitHub repository from 2022-2024 revealed that 33.7% of resolved bugs (191 out of 567 verified issues) were memory errors. The most common types identified were null dereference, memory leak or exhaustion, and buffer overflow or underflow, each accounting for over 10% of reported issues. Use-after-free bugs, though less frequent at 7%, are particularly critical, ranking high on lists like the "2023 CWE Top 10 Known Exploited Vulnerabilities (KEV) Weaknesses" (Section 2). These vulnerabilities pose severe risks, potentially leading to data leakage, denial of service, or enabling attackers to bypass security measures in sandboxed environments (Section 3).

Existing fuzzing techniques, often relying on grammar-guided generation, struggle to produce test cases with complex semantic behavior that can target specific, intricate modules of the PHP interpreter. While they ensure syntactic correctness, their ability to uncover deep memory errors is limited (Section 1). The PHP community, however, maintains a comprehensive official test suite of over 19,000 distinct test cases. These .phpt files cover more than 80 modules, offering diverse code semantics and valid syntax, often including specific execution environments via --extension-- and --ini-- sections (Section 2). These tests represent a "golden testbed" (Observation 1, Section 1), but their individual unit-test-like nature means 96.1% exhibit simple, sequential control flow, limiting their capacity to expose bugs arising from complex code interactions (Observation 2, Section 1). The challenge lies in enriching the code semantics of these high-quality tests to reveal memory errors. Traditional byte-level mutations are inefficient, and manual semantic transformations require significant expertise. This gap motivated the development of FlowFusion, aiming to automatically generate semantics-rich test cases by merging and transforming existing ones.

Key Findings

FlowFusion has demonstrated exceptional effectiveness in uncovering deep-seated memory errors within the PHP interpreter, yielding several significant contributions and findings:

  • Discovery of Numerous Unknown Bugs: The framework successfully identified 158 unique and previously unknown bugs in the PHP interpreter. Of these, a remarkable 125 have already been fixed by PHP developers, and 11 have been confirmed, showcasing FlowFusion's practical impact (Abstract, Section 1).
  • Diverse Range of Vulnerability Types: The discovered bugs span 10 different Common Weakness Enumerations (CWEs), including critical categories such as CWE-121 Stack-based Buffer Overflow, CWE-122 Heap-based Buffer Overflow, CWE-124 Buffer Underwrite, CWE-190 Integer Overflow or Wraparound, CWE-401 Missing Release of Memory after Effective Lifetime, CWE-416 Use After Free, CWE-457 Use of Uninitialized Variable, CWE-476 NULL Pointer Dereference, CWE-824 Access of Uninitialized Pointer, and CWE-825 Expired Pointer Dereference (Table 1 footnote, Section 6). This diversity highlights FlowFusion's ability to target a broad spectrum of memory safety issues.
  • Widespread Impact on PHP Codebase: The fixes for these bugs have impacted over 80 individual source files within the PHP interpreter, including critical components of the Zend engine and main functions. In total, over 5,000 lines of code were changed as a direct result of FlowFusion's reports, significantly enhancing the interpreter's security (Section 6).
  • Detection of Long-Standing Vulnerabilities: FlowFusion successfully uncovered bugs that had existed for years, including a particularly notable heap-use-after-free error that was over 20 years old, dating back to PHP v5.0.0 released in 2004 (Section 1, Section 6). This demonstrates the framework's capability to find deeply embedded issues that traditional testing methods missed.
  • Superior Performance over Existing Fuzzers: In comparative evaluations, FlowFusion significantly outperformed both the official PHP test suite and a naive test concatenation approach in detecting memory errors and achieving higher code coverage. It also surpassed state-of-the-art fuzzers like AFL++ and Polyglot, covering 24% more lines of code after 24 hours of fuzzing under identical conditions (Abstract, Section 6.3).
  • Integration into Official PHP Toolchain: Due to its proven effectiveness and the high quality of its bug reports, FlowFusion has been officially recognized by PHP developers and integrated into the official PHP toolchain (Abstract, Section 1). This institutional adoption validates its value and ensures its continued contribution to PHP's security.

Technical Deep Dive

FlowFusion's effectiveness stems from a sophisticated, multi-stage approach designed to generate semantically rich test cases that expose memory errors in the PHP interpreter. The core of its methodology is dataflow fusion, complemented by several other strategic techniques. The overall process consists of seven steps: corpus initialization, test mutation, dataflow analysis, interface fuzzing, environment crossover, dataflow fusion, and result analysis (Section 4.1).

The process begins with corpus initialization, where FlowFusion loads the 19,000 unique test cases from PHP's official test suite (Section 4.1). These serve as seed programs, leveraging their inherent quality and valid grammar.

Next, test mutation is applied to randomly selected seed tests. This step introduces minor perturbations to the code before fusion, such as replacing arithmetic, assignment, or logical operands, substituting constants with special values (e.g., int_max, null), or interchanging variables. These mutations occur with low probability to maintain syntactic validity while increasing semantic diversity (Section 4.3).

Crucially, dataflow analysis is performed on each chosen test program. For programs with sequential control flow (characteristic of 96.1% of official tests), dataflow effectively represents their code semantics. FlowFusion computes IN sets (dataflow facts entering a statement) and OUT sets (dataflow facts leaving a statement) for each statement, along with GEN sets (definitions generated) and KILL sets (definitions overwritten). For function calls, it assumes return values have data dependencies on arguments, represented by FUN sets (Section 2, Section 4.2). This analysis forms the basis for intelligently connecting variables across different tests.

Before fusion, FlowFusion also prepares for interface fuzzing and environment crossover. Interface fuzzing involves inserting statements into the generated test cases to call random PHP functions, using variables from the fused code as arguments. It dynamically identifies available PHP functions (around 1,682) using get_defined_functions(), get_loaded_extensions(), and get_extension_funcs(), and determines parameter counts and types using ReflectionFunction (Section 4.3). Environment crossover merges the --extension-- and --ini-- sections of the seed tests and inserts additional random configurations collected from the official test suite, such as memory limits, JIT mode, or opcache settings. This ensures the fused tests have valid prerequisites and introduces variability into the execution environment (Section 4.3).

The heart of FlowFusion is dataflow fusion, which aims to combine two programs (A and B) into a single fused program (F) with enriched semantics. The process, outlined in Algorithm 1 (Section 4.2), involves:

  1. Computing Dataflow Sets: The COMPUTEDATAFLOWSETS function calculates the IN and OUT sets for all statements in programs A and B, identifying all dataflows.
  2. Creating a Shared Variable: A new shared variable, $fusion, is introduced to bridge the two programs.
  3. Weighted Random Dataflow Selection: For each program, dataflows are weighted based on the number of variables they contain. A dataflow is then randomly selected using this weighted scheme. This ensures diversity in which parts of the programs are connected.
  4. Random Variable Selection: Within the selected dataflow, a specific variable is randomly chosen to be replaced.
  5. Random Replacement: Occurrences of the chosen variable in the source code are randomly replaced with the shared $fusion variable with a probability p (set to 0.5 in the paper). This creates modified programs A' and B'.
  6. Concatenation: The modified programs A' and B' are then concatenated to form the final fused program F.

This process is more sophisticated than naive concatenation, which simply runs tests independently or links the last variable of one to the first of another. FlowFusion's heuristics—random dataflow, random variable, and random replacement—introduce significant diversity, leading to new code interactions and coverage that would otherwise be missed (Section 4.2). For example, a fusion might replace a variable in clone $gen from one test with a DOM node variable from another, leading to clone $fusion where $fusion now holds a DOM node, potentially triggering bugs due to unexpected object types (Figure 1, Section 4.1).

Finally, in the result analysis step, the fused test cases are executed by the PHP interpreter, compiled with sanitizers like AddressSanitizer (ASan) and Undefined Behavior Sanitizer (UBSan). These sanitizers act as oracles, detecting memory errors (and other undefined behaviors like integer overflows). Upon a sanitizer violation, FlowFusion deduplicates crashes based on crash site and stack trace and then uses delta debugging to reduce the bug-inducing test case to its minimal form. Verified bugs are responsibly disclosed to the PHP community (Section 4.1). The fuzzer itself is built upon and patched from the official PHP testing script to handle customized .phpt files and execution failures (Section 5).

Demo / Proof of Concept

While the provided content is a peer-reviewed paper rather than a recorded talk, FlowFusion's effectiveness is compellingly demonstrated through several illustrative examples of real-world bugs it discovered. These examples serve as concrete proofs of concept for its dataflow fusion and complementary strategies.

One notable example is a 20-year-old heap-use-after-free memory error, which affected PHP v5.0.0 (released in 2004) and later versions (Listing 1, Section 1). FlowFusion uncovered this bug by fusing a test (Test A) verifying DOM objects and entity references with another test (Test B) checking base64 encoding. Specifically, it connected the $nodes variable (representing DOM child nodes) from Test A to the $values variable from Test B, using a new bridging variable $fusion. This generated a scenario where DOM-related objects were encoded using base64, an interaction not covered by the original unit tests, leading to the use-after-free during a foreach iteration.

Other critical bugs identified include:

  • Heap Use-After-Free in PCRE Module: Illustrated in Listing 3 (Section 6), this vulnerability arose from the premature shutdown of the Perl Compatible Regular Expressions (PCRE) module. FlowFusion detected it by reassigning a regular expression variable after its operations were complete, guided by dataflow fusion linking it to unrelated assignments. When the Standard PHP Library (SPL) attempted to clean up the regex object, it operated on freed memory.
  • Heap Overflow in PHP SQLite Module: As shown in Listing 4 (Section 6), FlowFusion found a heap overflow in the SQLite module. This occurred when creating a new SQLite object after assigning a database variable, where the buffer size was not checked before a memcmp operation, leading to an allocated buffer exceeding its expected size.
  • Null Pointer Dereference in Zend Compiler: Listing 5 (Section 6) highlights a bug where a fatal error in the compiler left data structures with stale values. Subsequent compilation incorrectly reused elements from the previous stack, leading to a segmentation fault due to wrong instructions being emitted.
  • Segmentation Fault in JIT and OPcache: A complex issue presented in Listing 6 (Section 6) caused the Zend compiler to crash. This bug was found in a specific environment with an opcache preload configuration, where caller_info, callee_info, and call_map were allocated in the arena but not reset before reuse by the next request.
  • Segmentation Fault in Session Module: Listing 7 (Section 6) details a segmentation fault in the PHP session module. This fault originated from changes made to the session decode process to prevent writing incomplete sessions, revealing an illegal return from a bailout that did not restore original bailout data.
  • Segmentation Fault in Zend Allocator: FlowFusion also identified a segmentation fault in the Zend allocator (Section 6.2). This occurred when a non-related class object from one seed test was assigned to a foreach statement in another test, an interaction missed by both the official test suite and naive concatenation.

These examples collectively underscore FlowFusion's ability to generate novel, complex code semantics by intelligently fusing existing test cases, thereby exposing deep, previously undetected memory safety vulnerabilities across various critical components of the PHP interpreter.

Defensive Implications

The findings from FlowFusion offer crucial insights for defenders and developers working with the PHP interpreter, highlighting areas for improved security posture:

  • Prioritize Memory Safety in C Codebase: The sheer volume and diversity of memory errors (158 bugs, across 10 CWEs) found by FlowFusion underscore the urgent need for heightened vigilance regarding memory safety in the PHP interpreter's extensive C codebase. Developers should adopt rigorous coding practices, static analysis tools, and code reviews focused on preventing issues like buffer overflows, use-after-free, and null pointer dereferences.
  • Integrate FlowFusion into CI/CD Pipelines: The demonstrated effectiveness of FlowFusion, especially its integration into the official PHP toolchain, suggests that similar fuzzing methodologies should be adopted for continuous integration/continuous deployment (CI/CD) pipelines in any organization maintaining or extending the PHP interpreter. This ensures ongoing, automated detection of memory errors as the codebase evolves.
  • Go Beyond Unit Testing for Complex Interactions: FlowFusion's core strength lies in its ability to uncover bugs arising from complex semantic interactions between different code paths, which individual unit tests often miss. Defenders should recognize that even well-maintained test suites might not cover all possible inter-module dataflows and consider adopting similar semantic fusion techniques in their own testing strategies.
  • Regularly Update PHP Interpreters: Given the high fix rate (125 fixed bugs), it is imperative for users and administrators to consistently update their PHP installations to the latest patched versions. Many of the bugs found by FlowFusion, including a 20-year-old use-after-free, indicate that even older, seemingly stable versions could harbor critical, exploitable vulnerabilities.
  • Understand and Monitor Specific CWEs: Defenders should familiarize themselves with the specific CWEs identified by FlowFusion (e.g., CWE-416 Use After Free, CWE-476 NULL Pointer Dereference, CWE-122 Heap-based Buffer Overflow). This knowledge can inform targeted threat modeling, vulnerability assessments, and the implementation of specific runtime protections or hardening measures where applicable.
  • Be Aware of Environment-Specific Vulnerabilities: The discovery of bugs related to specific execution environments, such as those triggered by opcache preload configurations or within debugging modules like phpdbg, highlights the importance of comprehensive testing across various deployment configurations. Defenders should ensure their testing environments closely mirror production setups, including relevant extensions and ini settings.
  • Leverage Sanitizers in Development: The paper emphasizes the role of sanitizers like ASan and UBSan as effective bug oracles. Developers of C/C++ components within PHP or other interpreters should routinely compile and test their code with these sanitizers enabled to catch memory safety issues early in the development cycle.

By adopting these defensive strategies, the PHP ecosystem can significantly enhance its resilience against sophisticated memory-based attacks, moving towards a more robust and secure foundation for web development.

Key Takeaways

  • PHP Interpreter Memory Errors are Critical: Memory errors in the PHP interpreter's C codebase are a significant and often overlooked security risk, potentially leading to severe vulnerabilities like data leakage and sandbox escapes.
  • FlowFusion is a Tailored Fuzzing Solution: FlowFusion is the first automatic fuzzing framework specifically designed to detect memory errors in the PHP interpreter, addressing a crucial gap in existing security research.
  • Dataflow Fusion is a Novel and Effective Technique: The core innovation, dataflow fusion, intelligently merges high-quality official PHP test cases by interleaving their dataflows, generating new, complex code semantics that expose deep-seated bugs.
  • Superior Bug Discovery and Coverage: FlowFusion found 158 unknown bugs (125 fixed, 11 confirmed), including a 20-year-old use-after-free, and outperformed state-of-the-art fuzzers like AFL++ and Polyglot by covering 24% more lines of code.
  • Practical Impact and Industry Recognition: The framework's success has led to its integration into the official PHP toolchain, demonstrating its practical value and potential for continuous security enhancement.
  • Extensible Principles for Language Interpreters: The high-level insights of dataflow interleaving and complementary strategies are adaptable and extensible to fuzzing other programming language interpreters, such as C/C++ or JavaScript.

About the Speaker(s)

The research behind FlowFusion was conducted by Yuancheng Jiang, Chuqi Zhang, Bonan Ruan, Jiahao Liu, Manuel Rigger, Roland H. C. Yap, and Zhenkai Liang, all affiliated with the National University of Singapore. Their collective expertise focuses on areas within software security, including advanced fuzzing techniques, program analysis, and vulnerability discovery in complex systems like language interpreters. Their work contributes significantly to enhancing the robustness and security of critical software infrastructure.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Solid fuzzing research with real results: 158 bugs, 125 fixed, a 20-year-old UAF, and adoption into the official PHP toolchain. The dataflow fusion technique is genuinely clever—not revolutionary, but a well-engineered approach to semantic test generation that actually works. This is the kind of research that moves the needle.

Heather Calloway (CISO) — SOLID

Solid academic fuzzing research with real-world results: 158 bugs found in the PHP interpreter, 125 already fixed, including a 20-year-old use-after-free. The technique is novel enough that PHP officially adopted it. Any organization running PHP at scale should understand what this implies about the state of interpreter-level memory safety.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)