YuraScanner: Leveraging LLMs for Task-driven Web App Scanning
Aleksei Stafeev
Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · Web Security
Overview
This talk introduces YuraScanner, a groundbreaking, fully automated, and task-driven web application scanner designed to overcome the limitations of traditional blackbox testing tools. Presented by Tim Reinvald, a master student at Zanat University in Germany and part of a research group at Cispa, YuraScanner leverages Large Language Models (LLMs) to intelligently navigate and interact with web applications. The core motivation behind this innovation stems from the observed struggle of conventional scanners to explore the deeper states of modern web applications, which often involve complex multi-step workflows.
Key moments
- 0:00 Introduction: Traditional scanners' limitations in web app exploration
- 1:30 Leveraging LLMs for task-driven web app scanning
- 2:20 YuraScanner's task extraction: identifying app workflows
- 3:30 Task execution and form-based vulnerability scanning
- 4:30 Evaluation setup across 20 popular web applications
- 5:20 YuraScanner's task extraction performance: 77% validity
- 6:10 Task execution success rates and a successful example
- 7:40 Demonstration: YuraScanner successfully (and destructively) deletes a user
YuraScanner: Leveraging LLMs for Task-driven Web App Scanning
Speakers: Tim Reinvald, Master Student, Zanat University / CISPA
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=NwMrinE5VT0
Overview
This talk introduces YuraScanner, a groundbreaking, fully automated, and task-driven web application scanner designed to overcome the limitations of traditional blackbox testing tools. Presented by Tim Reinvald, a master student at Zanat University in Germany and part of a research group at Cispa, YuraScanner leverages Large Language Models (LLMs) to intelligently navigate and interact with web applications. The core motivation behind this innovation stems from the observed struggle of conventional scanners to explore the deeper states of modern web applications, which often involve complex multi-step workflows.
Traditional scanning methodologies, such such as Breadth-First Search (BFS) or random crawling, frequently exhibit shallow coverage because they lack an understanding of the sequential user interactions required to access certain functionalities or input fields. YuraScanner addresses this critical gap by first extracting potential user tasks from a web application and then using an LLM to execute these tasks, simulating a human user's journey through the application. This approach significantly enhances the discovery of hidden attack surfaces and previously unreachable vulnerabilities, making it a vital development for improving the efficacy of automated security testing in today's intricate web environments.
Background
▶ Watch: Introduction: Traditional scanners' limitations in web app exploration (0:00)
Web application scanners have long been a cornerstone of blackbox security testing, offering an automated means to identify vulnerabilities without access to source code. However, the increasing complexity of modern web applications, characterized by dynamic content, Single Page Application (SPA) architectures, and intricate multi-step user workflows, has exposed significant limitations in traditional scanning approaches. These older methods, often relying on simple link traversal (like BFS) or random exploration, struggle to achieve deep coverage. They frequently fail to reach input fields or functionalities that are nested several layers deep within an application's interface, requiring a specific sequence of clicks, form submissions, or other interactions.
Consider a common scenario: an admin dashboard with a deeply nested menu structure. To access a "create new entity" form, a user might need to click on a main menu item, then a sub-menu item, followed by another sub-menu, and finally a "create" button. This sequence of four distinct actions reveals a form that might contain an XSS (Cross-Site Scripting) vulnerability. Traditional scanners, lacking an understanding of this logical flow, are unlikely to perform these steps in the correct order, leaving such vulnerabilities undiscovered. Their fundamental limitation is a lack of workflow awareness – the ability to comprehend and execute the multi-step processes expected from users.
More recent academic efforts have attempted to tackle this challenge by proposing model-based methods, often employing reinforcement learning on user-provided traces. While these approaches show promise in guiding scanners through specific workflows, they suffer from a significant scalability issue. User-provided traces or learned models are highly application-specific and cannot be easily transferred from one web application to another. This requirement for manual input or extensive training for each new application severely limits their practical utility in diverse testing environments. Around the same time, non-academic initiatives began exploring LLM-based browsing agents to assist users with specific tasks, such as booking a hotel. However, these agents were typically designed for executing a single, manually specified task, rather than autonomously discovering and executing a multitude of potential workflows across an entire application. This historical context highlights a clear need for a fully automated, scalable, and intelligent web application scanner capable of understanding and navigating complex user journeys – a gap that YuraScanner aims to fill by leveraging the advanced reasoning capabilities of LLMs.
Key Findings
▶ Watch: YuraScanner's task extraction: identifying app workflows (2:20)
YuraScanner's evaluation demonstrated significant advancements over traditional scanning methodologies, particularly in its ability to discover deeper attack surfaces and previously unknown vulnerabilities. The research team conducted a comprehensive assessment across a test bed of 20 popular and modern web applications, including platforms like GitLab and Joomla.
One of the primary findings related to task extraction performance. Across 10 of the selected web applications, YuraScanner autonomously generated a total of 2361 tasks. Manual inspection revealed that an impressive 77% (approximately 1816 tasks) of these generated tasks were deemed valid, meaning the functionality described by the task was genuinely present in the application. The remaining 23% of invalid tasks were primarily generated on pages that offered insufficient context for the LLM to formulate meaningful actions, such as a login page with only a login button.
The task execution performance further underscored YuraScanner's efficacy. The scanner successfully executed 37% of the tasks completely. An additional 25% were classified as partially successful. Partial success implies that YuraScanner managed to locate the target form or functionality associated with a task but encountered a failure at the final step, such as a form validation error preventing submission. Crucially, even partially successful executions contribute to increasing the discovered attack surface by identifying new forms. This brings the combined success rate for task execution, where new attack surface is discovered, to 62%. The remaining tasks failed due to various reasons, including the LLM issuing incorrect actions or giving up prematurely. A notable, albeit unintended, "success" highlighted the LLM's autonomy: in one instance, YuraScanner, tasked with deleting a user, successfully navigated to the user management section, clicked the delete button for the only existing user, confirmed the deletion, and subsequently locked itself out of the admin dashboard.
Perhaps the most compelling finding pertains to attack surface discovery and depth. When compared against the state-of-the-art scanner Black Widow, as well as standard BFS and random BFS crawlers, YuraScanner demonstrated a superior ability to penetrate deeper into web applications. The evaluation showed that 44% of discovered forms were found by YuraScanner and at least one other scanner. However, this shared discovery plateaued after a depth of three steps. In stark contrast, YuraScanner exclusively discovered 40.3% of the total forms. These forms were predominantly located at depths greater than three steps, indicating that they were effectively out of reach for existing traditional tools. This significant percentage of uniquely discovered, deep attack surface highlights YuraScanner's unique capability in uncovering hidden functionalities.
The ultimate validation of YuraScanner's effectiveness came in its vulnerability discovery capabilities. Across the 20 web applications in the test bed, the scanner identified 13 unique zero-day XSS vulnerabilities in three distinct applications. Remarkably, 12 of these 13 vulnerabilities were found exclusively by YuraScanner. The injection points for these critical vulnerabilities were located between four and five clicks away from the main page, whereas vulnerabilities found by Black Widow were only two clicks deep. This directly corroborates the finding that YuraScanner can reach and test much deeper parts of a web application, uncovering critical security flaws that would otherwise remain undetected.
Finally, the project team made the decision to open-source the code for YuraScanner after careful consideration, making this innovative tool available to the broader security community for further research, development, and adoption.
Technical Deep Dive
▶ Watch: Evaluation setup across 20 popular web applications (4:30)
YuraScanner's innovative design hinges on its strategic use of Large Language Models (LLMs) to simulate intelligent user interaction, thereby overcoming the limitations of traditional, deterministic crawling methods. Instead of attempting to build a custom model for each application, YuraScanner leverages the pre-trained knowledge and reasoning capabilities of LLMs to understand context and make informed decisions during the scanning process.
The architecture of YuraScanner is modular, comprising three main components: a Task Extraction Module, a Task Execution Module, and a Vulnerability Scanning Component.
- Task Extraction Module:
The process begins with an automated task extraction phase. YuraScanner first performs a shallow crawl of the target web application, typically limited to a depth of one, to gather initial context. For each page visited during this shallow crawl, the module extracts the textual content of all interactable HTML elements. These elements might include buttons, links, form labels, or navigation entries.
This extracted textual content is then fed to an LLM. The LLM's role here is to analyze these elements and generate a list of appropriate, high-level user tasks that could be performed within the application. For instance, if the LLM encounters elements related to "products" and "categories" on an e-commerce site, it might generate tasks such as "add a new category for products," "edit the information for an existing product," or "delete a previous order." A key characteristic of these generated tasks is their imperative sentence structure, typically consisting of an action verb (e.g., add, edit, delete) and an object upon which the action is to be performed. This structured output makes the tasks actionable for the subsequent execution phase.
- Task Execution Module:
Once a list of potential tasks has been extracted, the Task Execution Module takes over. This module is responsible for executing each task in a multi-step fashion, mimicking a user's journey. At each step of a task, the following iterative process occurs:
- Page Representation Generation: YuraScanner first generates a textual representation of the current web page. This involves extracting visible text, identifying interactive elements, and capturing their context.
- Command Issuance by LLM: This textual page content, along with the current task description (e.g., "add a new category for products"), is provided to the LLM. Based on its understanding of the task and the current page state, the LLM issues the next command. This command could be to click a specific HTML element (identified by its text or unique attributes), type text into an input field, or submit a form.
- Command Execution: The issued command is then executed on the actual web page within a browser automation environment (e.g., using a headless browser).
This three-step loop (represent page, LLM issues command, execute command) allows the LLM to intelligently navigate through the web application, progressively working towards the completion of the assigned task. During this process, YuraScanner continuously monitors and collects any discovered attack surface, with a particular focus on forms. Forms are critical because they represent direct interaction points where user input can be processed, making them prime targets for various vulnerabilities.
- Vulnerability Scanning Component:
The final component is the Vulnerability Scanning Component. For its XSS detection capabilities, YuraScanner does not reinvent the wheel; instead, it integrates with and leverages the existing XSS engine of a state-of-the-art scanner, specifically Black Widow. This integration allows YuraScanner to benefit from proven payload generation and detection logic.
After the Task Execution Module has finished exploring and collecting forms, the Vulnerability Scanning Component takes each discovered form and systematically visits it. It then injects a predefined list of XSS payloads into the form's input fields to determine if they are vulnerable to XSS attacks. The ability to reach deep-seated forms, which traditional scanners miss, combined with a robust XSS engine, is what enables YuraScanner to find novel vulnerabilities.
The evaluation methodology involved a rigorous comparison. The 20 web applications were split into two sets: one for detailed task validity and success rate analysis (10 apps, due to the manual effort of labeling thousands of tasks), and another for comprehensive vulnerability discovery across all 20 applications. This systematic approach ensured a thorough assessment of YuraScanner's capabilities against both its internal goals (task execution) and its external impact (vulnerability findings).
Demo / Proof of Concept
▶ Watch: YuraScanner's task extraction performance: 77% validity (5:20)
While a live, interactive demonstration was not performed during the presentation, the talk effectively served as a conceptual proof of concept by illustrating YuraScanner's operational capabilities through several compelling examples and statistical findings. These examples vividly showcased how the LLM-driven approach successfully navigates complex workflows and uncovers hidden functionalities.
One illustrative example of successful task execution involved the task of "adding new customers to the database in OpenCart." The speaker detailed the sequence of actions YuraScanner performed:
- It first clicked on the "customers" menu entry.
- Then, it proceeded to click on a sub-menu entry, also named "customers."
- Next, it clicked on a distinct "plus icon" (often used for "add new" functionality).
- Finally, YuraScanner successfully filled and submitted the revealed form with several input fields.
This step-by-step navigation, involving multiple clicks and a form submission, perfectly demonstrates YuraScanner's ability to understand and complete a multi-stage user task, reaching a form that a simpler crawler might miss.
A particularly memorable and humorous example highlighted the LLM's autonomy and the potential for unintended consequences: the "too successful" deletion of a user. The task assigned was "to delete a user from the user management section." YuraScanner navigated to the appropriate section, identified the only user listed, and then:
- Clicked on the "delete" button associated with that user.
- When prompted with a confirmation dialog ("Are you sure you would like to delete this user?"), the LLM autonomously confirmed the action.
- The system reported, "User deleted successfully."
The speaker then recounted the subsequent outcome: YuraScanner attempted to log back into the admin dashboard but failed, having deleted its own access. This anecdote, while amusing, served as a powerful proof of concept for the LLM's capacity to execute complex, multi-step actions, including those with significant side effects, even when operating in an unsupervised manner.
The most impactful proof of concept, however, came from the vulnerability discovery results. The finding of 12 unique zero-day XSS vulnerabilities by YuraScanner, with injection points located four to five clicks away from the main page, serves as a robust and empirical demonstration of its effectiveness. These findings directly validate the core hypothesis that task-driven LLM-based scanning can reach attack surfaces inaccessible to traditional scanners. The fact that these were previously unknown vulnerabilities, found in popular and modern web applications, concretely proves YuraScanner's ability to uncover critical flaws in deep application states, making a tangible contribution to web security.
Defensive Implications
▶ Watch: Demonstration: YuraScanner successfully (and destructively) deletes a user (7:40)
YuraScanner's capabilities and findings carry several significant implications for web application defenders, urging a re-evaluation of current security testing strategies and defensive priorities.
- Embrace Task-Driven Testing Methodologies: The most direct implication is the need for organizations to integrate task-driven approaches into their existing blackbox security testing frameworks. While traditional scanners remain useful for broad, shallow coverage, they are demonstrably insufficient for modern, complex applications. Defenders should consider deploying tools or methodologies that can simulate multi-step user workflows to uncover vulnerabilities that lie hidden deep within application logic. YuraScanner effectively complements traditional scanning, filling a critical gap in coverage.
- Prioritize Input Validation in Deep Workflows: YuraScanner's discovery of 12 zero-day XSS vulnerabilities, particularly those located four to five clicks deep into application workflows, highlights a common blind spot. Developers often focus robust input validation and output encoding on easily accessible forms (e.g., login, registration). However, forms and input fields that are part of complex, multi-step processes might receive less scrutiny. Defenders must ensure that comprehensive input validation and output encoding are consistently applied across all user-interactable fields, regardless of how obscure or deeply nested their access path may be.
- Review Form Submission and Error Handling Logic: The observation of "partially successful" task executions, where YuraScanner reached a target form but failed to submit it due to validation errors, indicates that such forms are within reach of an intelligent agent. This means that while a form might appear secure due to server-side validation, the path to exploit it has been identified. Defenders should meticulously review the logic around form submissions, especially within multi-step processes, to ensure that even validation failures don't inadvertently expose information or allow for bypasses with carefully crafted payloads.
- Implement Safeguards for Autonomous Agents: The anecdote of YuraScanner deleting its own admin user account serves as a stark reminder of the potential for autonomous agents to cause unintended harm, especially if deployed on production systems. Organizations exploring or deploying LLM-driven testing tools must implement robust safeguards and sandboxing mechanisms. This includes running such tools in isolated staging environments, carefully scoping permissions, and potentially disabling destructive actions (like deletion or modification) during initial testing phases. The speaker explicitly cautioned against using YuraScanner on live websites without owner consent and advised disabling the attack component, underscoring the ethical responsibilities involved.
- Recognize the "Out-of-Reach" Attack Surface: The finding that 40.3% of forms were exclusively discovered by YuraScanner at depths greater than three steps represents a significant portion of an application's attack surface that is currently "out of reach" for many organizations using traditional tools. These hidden forms and functionalities are prime targets for sophisticated attackers who might manually explore complex application paths. Defenders should be aware of this blind spot and proactively seek methods to explore these deeper states, perhaps by manually mapping critical deep workflows or adopting advanced scanning tools.
- Continuous Security Awareness and Tooling Updates: The rapid evolution of AI and LLMs means that automated attack techniques will continue to advance. Defenders must stay abreast of these developments, continuously updating their security tooling and methodologies to counter new forms of automated reconnaissance and exploitation. Leveraging LLM-driven tools for defense, similar to how YuraScanner uses them for offense, could become a future necessity.
By understanding and acting upon these defensive implications, organizations can significantly enhance their web application security posture, moving beyond superficial checks to address the complex, multi-layered attack surfaces that modern LLM-driven scanners are now capable of exposing.
Key Takeaways
- YuraScanner is a novel, fully automated, and task-driven web application scanner that leverages Large Language Models (LLMs) to intelligently explore complex application workflows, addressing the limitations of traditional blackbox testing.
- The tool successfully extracted 2361 tasks from web applications, with 77% deemed valid, demonstrating its ability to autonomously identify meaningful user interactions.
- YuraScanner achieved a high task execution success rate, with 37% fully successful and an additional 25% partially successful (finding the target form), indicating effective navigation and attack surface discovery in 62% of attempts.
- Crucially, YuraScanner discovered 40.3% of unique forms at depths greater than 3 clicks that were entirely missed by state-of-the-art traditional scanners (Black Widow, BFS, random BFS), proving its superior capability to reach deeper application states.
- The scanner identified 12 unique zero-day XSS vulnerabilities, with injection points located 4-5 clicks deep into web applications, showcasing its effectiveness in finding critical flaws in previously unreachable areas.
- Task-driven crawling, powered by LLMs, can effectively complement traditional scanning techniques, offering a vital approach to uncover vulnerabilities in multi-step user interactions and complex application logic.
- The project has been open-sourced, encouraging the security community to further explore and develop LLM-driven security testing methodologies.
About the Speaker(s)
The paper "YuraScanner: Leveraging LLMs for Task-driven Web App Scanning" was presented by Tim Reinvald. Tim is a Master student currently pursuing his studies at Zanat University in Germany. He is also an integral part of a research group led by Jankala Filipino at Cispa. His work, as evidenced by the YuraScanner project, focuses on cutting-edge research in web application security, particularly exploring the innovative application of Large Language Models to enhance automated vulnerability scanning and attack surface discovery.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Solid academic research with real empirical results: 12 zero-days found in production apps at depths traditional scanners can't reach, with a clean modular architecture and an open-source release. A master's student doing work that makes tool vendors uncomfortable is exactly what NDSS should be platforming.
Heather Calloway (CISO) — WEAK
Solid academic research with a genuinely interesting result — LLM-driven scanning reaches attack surface that traditional tools miss, and the zero-days prove it. But this talk is built entirely for researchers, not operators, and the defensive implications section reads like it was written to check a box rather than inform a decision.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025