EvoCrawl: Exploring Web Application Code and State using Evolutionary Search

Xiangyu Guo (University of Toronto)

Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · Mobile Security

Overview

Modern web applications present significant challenges for security scanners, particularly those operating in a blackbox manner without access to source code. This talk introduces EvoCrawl, an innovative blackbox web application scanner designed to overcome these limitations by intelligently exploring application code and state. Developed at the University of Toronto, EvoCrawl leverages an evolutionary search algorithm combined with dependency tracking to navigate the complex landscape of web application interactions, aiming to achieve superior code coverage and, consequently, more effective vulnerability detection.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction to EvoCrawl and blackbox scanning challenges
  2. 2:00 Demonstrating the complexity of HTML form submission for crawlers
  3. 3:00 Highlighting inefficiency of brute-force, need to avoid bad orders
  4. 4:25 Introducing evolutionary search and dependency tracking as the solution
  5. 5:00 How the fitness function identifies and rewards good sequences
  6. 6:08 Explaining crossover operation for generating diverse, good sequences
  7. 7:35 Introducing dependency tracking for generating good sequences from scratch

EvoCrawl: Exploring Web Application Code and State using Evolutionary Search

Speakers: Xiangyu Guo, University of Toronto

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=pxqHp0o0mBo

Overview

Modern web applications present significant challenges for security scanners, particularly those operating in a blackbox manner without access to source code. This talk introduces EvoCrawl, an innovative blackbox web application scanner designed to overcome these limitations by intelligently exploring application code and state. Developed at the University of Toronto, EvoCrawl leverages an evolutionary search algorithm combined with dependency tracking to navigate the complex landscape of web application interactions, aiming to achieve superior code coverage and, consequently, more effective vulnerability detection.

The core problem EvoCrawl addresses is the difficulty traditional blackbox crawlers face in triggering diverse application states and interacting with dynamic user interfaces. Many vulnerabilities are only exposed when an application transitions through specific states, often requiring a precise sequence of user actions and input submissions. EvoCrawl's methodology, which rewards sequences leading to state changes and punishes ineffective interactions, represents a significant step forward in making blackbox scanning more comprehensive and accurate.

This research is particularly relevant for the security community as it offers a novel approach to automated web vulnerability discovery. By focusing on state exploration and intelligent interaction generation, EvoCrawl demonstrates how advanced algorithmic techniques can enhance the capabilities of security tools, leading to the identification of vulnerabilities that might otherwise remain hidden due to the intricate nature of modern web application logic.

Background

▶ Watch: Introduction to EvoCrawl and blackbox scanning challenges (0:00)

The effectiveness of a blackbox web application scanner is critically dependent on its ability to maximize server-side code coverage. Without access to the underlying source code, a scanner must infer application logic and potential execution paths solely through client-side interactions. However, previous studies have consistently shown that the number of pages a crawler can reach, and thus the code it can cover, is heavily influenced by the application state. This means certain pages or functionalities only become accessible after the application has transitioned into a specific state, typically by modifying data stored on the server.

Consider a common scenario: in a paper submission system like HotCRP, the pages for reviewing or editing a paper are only accessible after a paper has been successfully submitted. This transition from a "no paper" state to a "has paper" state is a prerequisite for accessing subsequent functionalities. For a blackbox scanner, changing the data stored on the server and thus altering the application state primarily involves submitting HTML forms.

However, submitting HTML forms effectively is far from trivial, especially in modern web applications characterized by dynamic user interfaces. The speaker illustrated this with an example from a Campboard project management page. Clicking a simple plus sign might open a pop-up window, which then blocks interaction with elements on the original page. For a human, the sequence is intuitive: click the plus sign, fill in the title field within the pop-up, then click save. For an automated crawler, finding this specific sequence among a multitude of possible interactions is an enormous challenge. Even for a single page with approximately 33 elements, a simple sequence of string interactions could yield around 30,000 possible combinations. Exhaustively trying all these combinations is computationally prohibitive and severely impacts the scanner's performance.

This leads to two primary challenges. First, crawlers need a mechanism to efficiently identify and execute "good" sequences of interactions that lead to meaningful state changes, while avoiding "bad" orders (e.g., interacting with a blocked element). Second, once a short, effective sequence is found (like submitting a title), the crawler needs a way to automatically generate longer sequences that incorporate more inputs. This is crucial because different combinations of inputs can trigger diverse code paths on the server and serve as potential injection points for vulnerabilities. The ability to incrementally add more input fields and inject more data into the database, while simultaneously avoiding actions that prevent form submission (like clicking a "cancel" button or inputting values the crawler cannot infer), is paramount for comprehensive vulnerability detection. The existing blackbox scanning landscape largely struggles with these complexities, creating a significant gap in the ability to fully explore and secure modern web applications.

Key Findings

▶ Watch: Highlighting inefficiency of brute-force, need to avoid bad orders (3:00)

EvoCrawl's evaluation demonstrated significant advancements over existing blackbox web application scanners, particularly in its ability to achieve broader code coverage, higher rates of form submission, and more effective vulnerability detection. The core findings highlight the efficacy of combining evolutionary search with dependency tracking for navigating complex web application states.

Firstly, EvoCrawl achieved a substantial improvement in code coverage. When compared against established baselines such as Black Widow, Jack, and Crojax across 10 different applications, EvoCrawl covered an average of 1.5 times more unique lines of code. This metric is crucial because higher code coverage directly correlates with a greater likelihood of exercising different application functionalities and, consequently, discovering vulnerabilities. The results showed that while there were common lines covered by both EvoCrawl and its baselines, EvoCrawl consistently explored paths unique to its methodology.

Secondly, the system excelled in form submission, a critical mechanism for altering application state. EvoCrawl submitted an average of four times more unique forms compared to Black Widow, the primary baseline for this metric. This increased rate of unique form submissions indicates EvoCrawl's superior capability in identifying, correctly interacting with, and successfully submitting a wider variety of forms, which are essential for injecting data and transitioning the application through different states.

Finally, and most importantly, EvoCrawl demonstrated enhanced vulnerability detection capabilities. The scanner successfully identified three distinct IDOR (Insecure Direct Object Reference) endpoints and five different XSS (Cross-Site Scripting) vulnerabilities across three different applications. A significant aspect of these findings was that all detected IDOR vulnerabilities and many of the XSS vulnerabilities required the crawler to perform complex state transitions. For instance, one IDOR vulnerability necessitated uploading a private picture to transition the application into a specific state before the vulnerability could be tested. Similarly, a stored XSS vulnerability required a sequence of four distinct state changes—adding a new decision type, creating a paper submission, assigning the decision type to the paper, and then finally executing the payload—before the vulnerability could be injected and triggered. These examples underscore EvoCrawl's unique strength in uncovering vulnerabilities that are deeply embedded within multi-step, state-dependent application logic, which conventional crawlers often miss due to their inability to intelligently navigate such intricate paths.

Technical Deep Dive

▶ Watch: Introducing evolutionary search and dependency tracking as the solution (4:25)

EvoCrawl's technical prowess stems from its innovative integration of evolutionary search and dependency tracking, two mechanisms designed to overcome the inherent challenges of blackbox web application scanning. The goal is to intelligently generate and refine sequences of interactions that effectively explore diverse application states, thereby maximizing code coverage and vulnerability discovery.

The evolutionary search algorithm is at the heart of EvoCrawl's sequence generation. It operates on a population of interaction sequences, iteratively refining them based on a defined fitness function and crossover operation. The choice of a diversified evolutionary search is intentional, aiming to explore a broad spectrum of application states rather than converging on a single path.

The fitness function is crucial for guiding the search. It assigns a "score" to each interaction sequence, rewarding those that are likely to lead to meaningful state changes and punishing those that are unproductive. Specifically, EvoCrawl's fitness function rewards sequences that:

  • Inject inputs into the database: This is a direct indicator of successful form submission and data modification, which fundamentally alters application state.
  • Fill more input fields: Sequences that interact with a greater number of input fields are favored, as this suggests a more comprehensive exploration of form parameters and potential injection points.
  • Trigger JavaScript events: Many modern web applications rely heavily on JavaScript to dynamically alter the DOM, reveal new elements, or enable new functionalities. Triggering these events can expose new links, forms, or interaction possibilities that are otherwise hidden.

Conversely, the fitness function punishes sequences that waste the crawler's time or prevent progress. This includes interactions with blocked elements (e.g., elements obscured by a pop-up) or interactions with elements like "cancel" buttons or those requiring values that cannot be inferred by the crawler, which could prevent successful form submission.

The crossover operation is the primary mechanism for generating new, potentially more effective sequences from existing "good" ones. It works by concatenating parts of two high-fitness sequences (parents) to create a new sequence (offspring). This approach offers several benefits:

  • Diversity: It introduces new combinations of interactions, allowing the crawler to explore novel paths.
  • Partial Order Preservation: By combining segments of already "good" sequences, the crossover operation partially maintains the effective interaction orders discovered by the parents.

The speaker provided an illustrative example: if parent A successfully fills a title field and clicks a save button (resulting in database injection), and parent B fills many different input fields but never hits save, a crossover operation could combine the input-filling actions of B with the submission action of A. The resulting offspring sequence would fill many input fields and successfully submit the form, leading to more extensive input injection and a more significant state transition.

Complementing the evolutionary search is dependency tracking. While evolutionary search helps refine and combine sequences, dependency tracking is essential for generating "good" sequences from scratch and ensuring that interactions occur in a valid order. Dependencies are identified when an interaction with one element dynamically changes the status or visibility of other web elements. For instance, clicking a "+" sign might cause a pop-up window to appear, containing "title" and "save" elements. In this scenario, the "title" and "save" elements are dependent on the "+" sign. This implies that in any valid interaction sequence, the "title" and "save" elements must be interacted with after the "+" sign.

EvoCrawl detects these dependencies by monitoring DOM changes through a JavaScript API. If an interaction with an element leads to the exposure of new elements on the webpage, a dependency is recorded between the initial interaction and the newly exposed elements. This mechanism ensures that the crawler constructs sequences that respect the application's UI logic, preventing interactions with elements that are not yet visible or active.

Together, evolutionary search and dependency tracking form a powerful synergy. Dependency tracking provides the foundational understanding of valid interaction orders, enabling the generation of initial "good" sequences. Evolutionary search then takes these foundational sequences, refines them through its fitness function, and combines them using crossover to explore increasingly complex and state-rich interaction paths. This allows EvoCrawl to transition the application into a wide variety of states, ultimately leading to greater code coverage on the server side and a higher probability of uncovering deeply hidden vulnerabilities.

Demo / Proof of Concept

▶ Watch: Explaining crossover operation for generating diverse, good sequences (6:08)

While the talk did not feature a live, interactive demonstration, the "Demo / Proof of Concept" aspect of EvoCrawl was thoroughly established through its successful identification of real-world vulnerabilities that required complex state transitions. These findings served as concrete evidence of the system's capabilities and how its unique approach led to discoveries missed by other tools.

A key demonstration of EvoCrawl's prowess was its detection of three IDOR (Insecure Direct Object Reference) endpoints. The speaker highlighted one specific example where finding the IDOR required a multi-step process. The crawler first had to upload a private picture, effectively transitioning the application from an initial state (State 0) to a new state (State 1) where a picture existed within the user's context. Only after this state transition could EvoCrawl then test for and discover the size controllability vulnerability on the page, demonstrating how a seemingly simple IDOR was contingent on a preceding, state-altering action.

Even more illustrative was the discovery of five Cross-Site Scripting (XSS) vulnerabilities across three different applications. One particular stored XSS vulnerability proved to be exceptionally complex, requiring a sequence of four distinct state transitions to both inject and execute the payload successfully:

  1. The crawler first needed to add a new decision type within the application, modifying the application's configuration.
  2. Next, it had to create a paper submission, generating a new entity within the system.
  3. Following this, the newly added decision type had to be assigned to the paper, linking the two previously created entities.
  4. Finally, after these three state-changing operations, the XSS payload could be successfully executed, demonstrating that the vulnerability was deeply embedded within the application's workflow and state management.

These examples clearly illustrate that EvoCrawl's ability to intelligently navigate and manipulate application states, driven by its evolutionary search and dependency tracking mechanisms, is not merely theoretical but translates directly into tangible security findings. The integration of IDOR and XSS detectors into EvoCrawl, with provisions for future integration of other vulnerability types, underscores its design as a versatile and extensible vulnerability scanning platform capable of tackling the intricacies of modern web application security. The successful identification of vulnerabilities requiring such intricate state changes stands as a compelling proof of concept for EvoCrawl's innovative methodology.

Defensive Implications

▶ Watch: Introducing dependency tracking for generating good sequences from scratch (7:35)

EvoCrawl's findings carry significant implications for developers, security engineers, and organizations responsible for securing web applications. The research highlights critical blind spots in traditional security testing and offers insights into building more resilient applications.

Firstly, developers must recognize that modern web applications' complexity, particularly their reliance on dynamic UIs and intricate state management, can inadvertently hide vulnerabilities. Simple form submissions or direct URL access checks are often insufficient. Vulnerabilities like the IDOR requiring a private picture upload or the XSS demanding four distinct state transitions underscore that security flaws can be deeply intertwined with application logic and user workflows. Defenders should therefore prioritize thorough testing that simulates complex, multi-step user journeys and state changes, rather than relying solely on static analysis or basic dynamic scanning.

Secondly, the success of EvoCrawl in discovering these state-dependent vulnerabilities emphasizes the need for comprehensive authorization and authentication checks at every stage of an application's workflow, not just at initial access points. For IDOR, this means ensuring that every request for a resource is accompanied by robust authorization checks that verify the requesting user's permission to access that specific instance of the resource, preventing unauthorized access even if the resource ID is correctly guessed.

Thirdly, the challenges posed by dynamic UIs and complex form interactions, which EvoCrawl addresses with dependency tracking and evolutionary search, point to the importance of secure UI development practices. While these features enhance user experience, they can also create obfuscated paths that traditional scanners struggle to traverse. Developers should ensure that security controls are not bypassed by dynamic content loading or JavaScript-driven UI changes. Rigorous input validation and output encoding remain paramount for preventing XSS, but EvoCrawl's findings show that these controls must be effective across all possible input fields and state-dependent rendering contexts. Implementing a strong Content Security Policy (CSP) can also mitigate the impact of XSS vulnerabilities, even if an injection point is found.

Finally, organizations should consider adopting or developing advanced security testing tools that incorporate state-aware crawling techniques, similar to EvoCrawl's methodology. Relying solely on basic blackbox scanners might provide a false sense of security, as many vulnerabilities requiring complex state changes will likely be missed. Integrating such sophisticated dynamic application security testing (DAST) solutions into the CI/CD pipeline can help identify these deeper, more insidious flaws earlier in the development lifecycle. The core takeaway for defenders is that a proactive, state-aware approach to security testing is indispensable for protecting modern, interactive web applications.

Key Takeaways

  • Evolutionary search and dependency tracking are highly effective mechanisms for blackbox web application scanning, enabling intelligent exploration of complex web interfaces.
  • Application state transitions are crucial for achieving high server-side code coverage and uncovering vulnerabilities in modern web applications.
  • Traditional crawlers often struggle with dynamic UIs, pop-up windows, and complex form interactions, leading to missed code paths and undiscovered vulnerabilities.
  • EvoCrawl significantly outperforms baseline scanners like Black Widow, Jack, and Crojax, achieving 1.5 times more unique code coverage and submitting 4 times more unique forms.
  • Vulnerabilities such as IDOR and XSS frequently require multi-step, state-dependent interaction sequences to be identified and exploited, demonstrating the limitations of simple scanning approaches.
  • Future web application security scanners must incorporate sophisticated techniques for understanding and manipulating application state to effectively detect vulnerabilities in increasingly complex web environments.

About the Speaker(s)

Xiangyu Guo is a researcher from the University of Toronto. His work, as presented in the EvoCrawl talk, focuses on advancing the capabilities of web application security scanners through innovative algorithmic approaches, specifically leveraging evolutionary search and dependency tracking for blackbox vulnerability detection.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Solid academic systems paper presenting a real engineering contribution — evolutionary search plus dependency tracking for stateful blackbox web crawling is a legitimate and non-trivial idea. The results are credible and the problem framing is honest, but this is a conference proceedings presentation, not a practitioner security talk, and the gap between the research artifact and anything a defender can actually use tomorrow is wide.

Heather Calloway (CISO) — WEAK

EvoCrawl is technically credible research — evolutionary search plus dependency tracking is a genuinely interesting approach to a real crawling problem. But this is a research paper delivered as a conference talk, and it never crosses into governance, operational relevance, or defender action. The findings are narrow, the vulnerability count is small, and the 'defensive implications' section reads like it was written to fill a template rather than to tell anyone what to do.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025