One Thousand and One AI-Prevented CVEs: Vibe Coding a Whole New Supply Chain Defense

Brandon Wu

BSidesSF 2026 · Day 2 · AMC Theatre 14

Overview

In an era where software supply chain attacks are escalating dramatically, the manual processes traditionally employed to secure third-party dependencies are proving to be unsustainable. Brandon Wu, a Program Analysis Engineer at Semgrep, presented a compelling talk at BSides SF, "One Thousand and One AI-Prevented CVEs: Vibe Coding a Whole New Supply Chain Defense," addressing this critical challenge. The presentation introduced Brat (Better RAT), a novel tool developed to automate the creation of security rules for detecting vulnerabilities (CVEs) in third-party libraries, significantly improving the scalability and reliability of software supply chain defense.

Watch on YouTube

Visual summary for One Thousand and One AI-Prevented CVEs: Vibe Coding a Whole New Supply Chain Defense by Brandon Wu
Visual summary for One Thousand and One AI-Prevented CVEs: Vibe Coding a Whole New Supply Chain Defense by Brandon Wu

Key moments

  1. 2:00 Talk title origin and supply chain security introduction
  2. 2:40 Understanding the widespread problem of supply chain vulnerabilities
  3. 3:00 Statistics on CVEs and the challenge of defense
  4. 4:00 Why detection engines alone aren't enough
  5. 4:50 A security researcher's daily workflow for CVEs
  6. 5:30 How Semgrep rules detect vulnerable code usage
  7. 6:00 The manual process of analyzing and fixing CVEs

One Thousand and One AI-Prevented CVEs: Vibe Coding a Whole New Supply Chain Defense

Speakers: Brandon Wu, Program Analysis Engineer, Semgrep

Conference: BSides SF

YouTube: https://www.youtube.com/watch?v=9U_49kFMYUY

Overview

In an era where software supply chain attacks are escalating dramatically, the manual processes traditionally employed to secure third-party dependencies are proving to be unsustainable. Brandon Wu, a Program Analysis Engineer at Semgrep, presented a compelling talk at BSides SF, "One Thousand and One AI-Prevented CVEs: Vibe Coding a Whole New Supply Chain Defense," addressing this critical challenge. The presentation introduced Brat (Better RAT), a novel tool developed to automate the creation of security rules for detecting vulnerabilities (CVEs) in third-party libraries, significantly improving the scalability and reliability of software supply chain defense.

Wu’s talk highlighted the overwhelming burden placed on security researchers who manually analyze CVEs and author detection rules. With an estimated 50,000 supply chain vulnerabilities projected for 2025, the industry faces an "incoming flood" that traditional human-centric approaches cannot manage. Brat offers a paradigm shift by strategically combining deterministic code workflows with a "tasteful amount of AI" to automate the laborious research, test generation, and rule generation processes. This approach prioritizes transparency, verifiability, and composability, contrasting sharply with the unpredictable nature of purely agentic AI systems.

The significance of this work lies in its ability to transform a bottlenecked, human-intensive process into an efficient, scalable, and auditable pipeline. By demonstrating how Brat has authored nearly 3,500 rules, drastically reducing researcher time and improving detection accuracy, Wu underscored a crucial lesson for the security community: while AI offers immense potential, its integration into critical security functions demands a thoughtful, structured approach that leverages its strengths without sacrificing control or verifiability. This talk provides a blueprint for building robust supply chain defenses that can truly keep pace with the evolving threat landscape.

Background

▶ Watch: Talk title origin and supply chain security introduction (2:00)

The modern software landscape is characterized by an intricate web of dependencies. As Wu eloquently put it, our software relies on "endless domino elephants upon elephants tower of software" – a vast ecosystem of third-party packages and libraries. While this modularity accelerates development, it also introduces significant security risks, collectively known as software supply chain vulnerabilities. The scale of this problem is staggering: Wu cited projections of 50,000 estimated supply chain vulnerabilities in 2025, a substantial increase from 40,000 in 2024. Recent high-profile incidents like ShyHulu, ReactToShell, and LangFlow underscore the pervasive threat these vulnerabilities pose.

The core challenge for security companies like Semgrep, which develop solutions to secure third-party libraries, is how to provide an effective and scalable defense against this "incoming flood" of CVEs. Historically, this has largely fallen to human security researchers. Wu vividly described the routine of a hypothetical researcher, "Max," who exemplifies the manual, labor-intensive process:

  1. CVE Triage: Max checks a queue for newly filed CVEs overnight.
  2. Vulnerability Research: For each CVE, Max "trolls the internet" to read descriptions, identify the affected package, understand the attack type (e.g., XSS, prototype pollution), and pinpoint the commit that fixed the issue (the patch commit).
  3. Code Analysis: Max then opens the relevant open-source repository, clicks through source code for up to 30 minutes, trying to understand how the vulnerability manifests and how a user might be affected. This often involves educated guesses about vulnerable code paths.
  4. Usage Research: Since Max might not be familiar with every library's API, he searches GitHub for real-world examples of how the package is used.
  5. Rule Authoring: Finally, Max writes a detection rule (e.g., a Semgrep rule) to find vulnerable uses of the CVE.
  6. Repetition: This entire process is repeated for dozens of new CVEs, day in and day out.

This workflow is inherently unscalable. As the number of CVEs grows and the surface of supported languages and frameworks expands (from Python to Rust, PHP, and C++), the human cost in terms of time, resources, and potential mistakes becomes prohibitive. Wu emphasized that "every single rule that we write... costs you human resources and costs you time and costs you possible mistakes."

To understand a vulnerability, it's crucial to grasp its lifecycle:

  • Introducing Commit: The code change that first introduces insecure behavior into a library.
  • Disclosure: When the vulnerability is discovered and publicly reported, often assigned a unique CVE ID.
  • Patch Commit: The code change that fixes the vulnerability.
  • Vulnerable Range: All versions of the package released between the introducing and patch commits.

From a detection standpoint, not all vulnerabilities are created equal. Semgrep categorizes them into two main types:

  • Reachable Rules: These apply when a specific subset of a package's API is vulnerable, and the vulnerability can be identified by detecting particular code patterns or function calls (e.g., calling library.save with sensitive inputs, or using a foo function that processes tainted data). These are typically detectable with static analysis queries like Semgrep or CodeQL.
  • Upgrade Only: This category applies to vulnerabilities deeply embedded in foundational frameworks (e.g., web frameworks) where the mere presence of an older version implies vulnerability, regardless of specific usage patterns. The only effective defense is to upgrade the entire package.

The critical insight underpinning Brat's development was the recognition that while current agentic AI workflows (where an AI is simply told to "figure it out") might seem appealing, they are often "unpredictable," "not transparent," and "hard to test, reason about, and improve upon." Wu argued that a more structured, deliberate approach, even if it incorporates AI for specific subtasks, yields a far more robust and reliable security solution.

Key Findings

▶ Watch: Statistics on CVEs and the challenge of defense (3:00)

The central finding of Brandon Wu's talk is the successful development and deployment of Brat (Better RAT), a tool that fundamentally transforms the process of generating security rules for supply chain vulnerabilities. Brat's innovative approach, which combines deterministic code with targeted AI assistance, has yielded significant results:

  1. Massive Scale and Efficiency Gains: Brat has authored nearly 3,500 rules over two years of existence. This automation drastically reduces the human effort required. Wu presented conservative estimates indicating that a "ready to merge" rule, generated by Brat, takes approximately 10 minutes of researcher time, down from an average of one hour for manual creation. Even "very helpful" rules take only about 20 minutes, and "helpful" rules about 45 minutes. This represents a substantial saving, conservatively estimated at 40% or more of the time spent writing rules.
  1. Enhanced Scalability for Backfilling and Surge Handling: The parallelizable nature of Brat's workflow is a critical advantage. It enables rapid backfilling of rules for new languages or frameworks (e.g., PHP or Rust, which Semgrep supply chain initially didn't support), converting thousands of hours of potential researcher backlog into mere hours of automated processing. Furthermore, Brat can effectively handle sudden surges in CVE disclosures, such as an instance where 50 CVEs were filed overnight, a task that would overwhelm human researchers but is managed efficiently by the automated system.
  1. Superior Accuracy and Completeness Compared to Pure Agentic AI: A key finding highlighted through a case study involving a prototype pollution vulnerability in the d-value npm library demonstrated Brat's advantage. A purely agentic AI approach only identified d-value.parse as the vulnerable entry point. However, Brat's code graph analysis correctly identified that unflatten, a helper routine and also a public API function, was transitively called by parse and was itself vulnerable. The agentic approach missed unflatten, leading to a false negative. This illustrates that while agentic AI can "skip steps and has leaps in logic," Brat's structured, verifiable subtasks prevent such critical omissions.
  1. The "Tasteful Amount of AI" Principle: Wu's core thesis—that "deterministic workflows and non-deterministic LLMs... make an effect better than what either could accomplish alone"—was strongly validated. Brat uses LLMs for specific, well-defined subroutines (e.g., scraping API documentation, parsing READMEs, generating test code snippets or initial patterns), but the overall workflow, particularly the critical code graph analysis and verification steps, remains deterministic and auditable. This approach ensures transparency, testability, and improvability at each stage, contrasting with the black-box nature of purely agentic systems.
  1. Human-in-the-Loop Integration: Brat doesn't aim to replace human security researchers but rather to augment their capabilities. The system generates research summaries and code graphs as artifacts, allowing human researchers to review and audit the automated process. This "trusted mostly deterministic pipeline with a human in the loop workflow" ensures that even with automation, critical judgment and oversight are maintained, leading to more robust and reliable security outcomes.

In essence, Brat demonstrates that by carefully structuring automation with verifiable components and strategically integrating AI for specific, well-scoped tasks, it is possible to achieve unprecedented scale and accuracy in supply chain vulnerability detection, moving beyond the limitations of both purely manual and unconstrained agentic approaches.

Technical Deep Dive

▶ Watch: Why detection engines alone aren't enough (4:00)

Brat's architecture is meticulously designed around three core, composable steps: Research Automation, Test Generation, and Rule Generation. This modularity is key to its deterministic nature and verifiable outputs.

Research Automation: Building the Code Graph

The initial phase, Research Automation, aims to gather comprehensive information about a CVE and the affected package. This is where a "tasteful amount of AI" is first introduced, with LLMs assisting in tasks like scraping READMEs, online API documentation (e.g., readthedocs.io), and parsing textual descriptions. However, the cornerstone of this phase, and indeed Brat's overall intelligence, is the creation of a code graph, also known as a call graph.

The Core Insight: Vulnerable Code Calls Vulnerable Code

Wu's fundamental insight is elegantly simple: "Vulnerable code is made vulnerable by calling vulnerable code." If a function (B) is vulnerable, any function (A) that calls B is potentially also vulnerable. This principle forms the basis for traversing the code graph.

Code Graph Construction:

  1. Nodes and Edges: The code graph represents functions as nodes, and caller-callee relationships as edges. An edge from A to B signifies that A calls B.
  2. Identifying Starting Points: Functions that were directly modified or changed by the patching commit are marked as the initial "vulnerable" nodes (often depicted in gold in Wu's diagrams). The assumption is that if a commit fixed an issue, the changed function was likely the source of the problem.
  3. Backward Graph Search (Transitive Closure): Brat then performs a graph search backwards from these initially marked vulnerable functions. Any function that transitively calls a vulnerable function is also marked as potentially vulnerable. This process identifies the transitive closure of the vulnerability—all functions that, directly or indirectly, lead to the insecure behavior.
  4. Static Analysis: Crucially, this code graph generation is performed using static analysis via the Semgrep engine (or other open-source indexers like Skip). This means the analysis is done on the code itself, without needing to execute it at runtime, ensuring determinism and avoiding the complexities of dynamic execution environments.

Filtering for Public API Endpoints:

The transitive closure can encompass hundreds of internal functions. However, not all vulnerable functions are relevant for rule generation; only those that are part of the public API and can be invoked by a consumer of the library matter. Brat employs a sophisticated filtering mechanism for this:

  1. GitHub API Search (Hera 6): Brat queries GitHub's API (using techniques like Hera 6) to search for real-world usage patterns of the library. It looks for concrete strings like import TensorFlow or from TensorFlow import, or specific method calls like foo.bar().
  2. Program Analysis Refinement: The raw search results from GitHub can contain false positives (e.g., bar appearing in a string literal rather than as a function call). Semgrep's program analysis capabilities are then used as a second pass to accurately identify true usages of public API functions.
  3. Public Node Identification: By combining the backward graph traversal with real-world usage data, Brat identifies the "blue nodes" in the code graph—the public-facing API functions that are both vulnerable and actually consumed by users. Rules are then primarily generated for these public, reachable functions, avoiding wasteful detection of internal, private vulnerabilities.

This deterministic code graph generation is a core strength, ensuring that for a given code change, the identified vulnerable functions are consistent and auditable.

Test Generation and Rule Generation: The Separation Principle

One of Brat's most significant architectural decisions is the explicit separation and parallelization of test generation and rule generation. This directly addresses the unpredictability and "leaps in logic" often associated with monolithic agentic AI workflows.

The Insight: Do Separate Things Separately

Wu emphasized that if a problem can be decomposed into independent subproblems, each subproblem should be tackled and verified independently. For security rule writing, this means that different vulnerable functions within a package can be treated as separate entities.

The Workflow:

  1. Independent Test Snippets: For each identified public, vulnerable function (e.g., foo.bar, foo.baz), an LLM is tasked with generating a small, synthetic code snippet demonstrating its vulnerable usage. For instance, if foo.baz is vulnerable when vulnerable=True is passed, the LLM generates a snippet like foo.baz(vulnerable=True). If foo.bar is unilaterally vulnerable, a simple call foo.bar() is generated.
  2. Independent Pattern Generation: In parallel, for each of these vulnerable usage snippets, an LLM is asked to generate a corresponding Semgrep pattern (query) that specifically matches that vulnerable code. For foo.baz(vulnerable=True), the pattern might be pattern: foo.baz(vulnerable=True). For foo.bar(), it might be pattern: foo.bar().
  3. Separate Verification: Crucially, each generated pattern is mechanically verified against its respective synthetic test snippet. This ensures that the pattern correctly identifies the vulnerable usage and, ideally, does not produce false positives on non-vulnerable code. If a subcomponent fails (e.g., the pattern doesn't parse or doesn't match the test code), only that specific subcomponent is re-run or discarded, rather than restarting the entire process.
  4. Stitching Together: Once individual patterns and test snippets are verified, they are textually concatenated. All individual test snippets are combined into a single test file, and all successful Semgrep patterns are combined into a single, comprehensive Semgrep rule using logical OR operations (e.g., pattern: A OR B).

Advantages of Separation:

  • No Dropped Symbols: This modularity prevents the AI from accidentally omitting a vulnerable function from the final rule, a common issue with monolithic agentic approaches.
  • Auditable Workflow: Each part of the workflow (research, test gen, rule gen for each function) can be independently audited and debugged.
  • Robustness and Composability: The system becomes more robust because failures are isolated, and more composable because well-defined subproblems can be improved individually without affecting the whole.
  • Improvability: If a specific type of vulnerability or a particular language's pattern generation is consistently flawed, that isolated component can be targeted for improvement (e.g., prompt tuning for that specific subtask) without disrupting other successful parts of the pipeline.

By combining deterministic code graph analysis with a highly modular, verifiable AI-assisted rule and test generation process, Brat achieves a level of reliability and scalability that is beyond the reach of either purely manual or unconstrained agentic methods.

Demo / Proof of Concept

▶ Watch: How Semgrep rules detect vulnerable code usage (5:30)

While Brandon Wu's talk did not feature a live, interactive demonstration of Brat, it provided compelling illustrative examples and a detailed case study to underscore the tool's capabilities and its advantages over purely agentic approaches. The "Demo / Proof of Concept" was presented through a series of conceptual diagrams, code snippets, and a real-world vulnerability analysis.

Illustrative Examples:

  1. Simple Code Graph (26:00): Wu presented a simplified code graph for a real CVE. The starting vulnerable function was remove_DTD_markup_declarations (gold node), implying it was the root cause of the vulnerability. A backward graph search showed that clean called this function, and is_svg (a public function, depicted in blue) called clean. This visual explained how Brat traces the vulnerability from an internal change to an externally exposed API, ensuring that rules target only the relevant public entry points. The distinction between internal (red) and public (blue) nodes highlighted Brat's filtering mechanism to avoid generating rules for unused internal functions.
  1. Complex Code Graph (28:00): A more intricate, larger code graph was shown to emphasize that vulnerabilities can propagate through many layers of internal functions. This reinforced the point that manually identifying all affected public API calls would be incredibly time-consuming and error-prone, whereas automated code graph analysis handles the transitive closure efficiently.
  1. GitHub Search and Program Analysis (30:00): Wu illustrated how Brat uses GitHub's API to find real-world usages of library functions (e.g., import foo and foo.bar()). He then showed how Semgrep's program analysis capabilities are crucial to filter out false positives (e.g., bar appearing in a string) and confirm genuine calls to public API methods, ensuring the accuracy of the "blue node" identification.
  1. Separated Test and Rule Generation (34:00): This critical concept was demonstrated with synthetic code snippets. For two hypothetical vulnerable functions, foo.bar and foo.baz, Wu showed:
  • Two independent test code snippets: one for foo.baz(vulnerable=True) and another for a unilaterally vulnerable foo.bar().
  • Two corresponding, independently generated Semgrep patterns: one specifically matching foo.baz(vulnerable=True) and another for any call to foo.bar().
  • The final step of concatenating these into a single, composite test file and a single Semgrep rule (using an OR operator). This visually represented the modular, verifiable approach that prevents the AI from "dropping a symbol" or failing to generate a rule for one part of the vulnerability.

Case Study: d-value Prototype Pollution (40:00):

The most compelling proof of concept was a detailed case study of a prototype pollution vulnerability in the d-value npm library.

  • Agentic Approach: A hypothetical agentic AI, when tasked with researching and writing a rule for this CVE, concluded that the vulnerability was in d-value.parse because the unflatten function didn't block proto keys. Consequently, the agent generated a rule only for d-value.parse.
  • Brat's Code Graph Analysis: Brat's analysis revealed a more complete picture. The actual problematic function was hydrate, which was called by unflatten, which in turn was called by parse. Crucially, Brat identified that both unflatten and parse were public API functions (blue nodes).
  • The Critical Difference: The agentic rule would have produced a false negative for any user directly calling d-value.unflatten, leaving a significant attack vector undetected. Brat, by systematically building and traversing the code graph and identifying all public, reachable vulnerable entry points, generated a more comprehensive and accurate rule.

This case study powerfully illustrated Wu's thesis: while agentic AI might offer a quick "one-shot" solution, its "leaps in logic" and lack of auditable steps can lead to critical omissions. Brat's structured, subtask-oriented, and verifiable workflow, even with AI assistance, results in a more trustworthy and effective security product.

Defensive Implications

▶ Watch: The manual process of analyzing and fixing CVEs (6:00)

Brandon Wu's presentation offers several critical defensive implications for organizations striving to secure their software supply chains:

  1. Embrace Automated Rule Generation for Scalability: The sheer volume of new CVEs (projected 50,000 in 2025) makes manual rule writing an untenable strategy. Defenders must invest in automated systems like Brat to generate detection rules at scale. Relying solely on human researchers will inevitably lead to a backlog of unaddressed vulnerabilities and unacceptable exposure windows.
  1. Prioritize Verifiable and Composable Automation: When adopting AI-powered security tools, defenders should scrutinize their underlying architecture. Black-box agentic AI, while seemingly powerful, can be unpredictable, non-transparent, and prone to critical false negatives (as seen in the d-value case study). Instead, prioritize solutions that leverage AI for well-defined, isolated subtasks within a larger, deterministic, and auditable workflow. This ensures that the system's reasoning is transparent and its outputs are verifiable.
  1. Leverage Static Analysis and Code Graphs for Accuracy: The talk highlights the power of static analysis and code graph (call graph) traversal to accurately identify the true impact of a vulnerability. Defenders should seek tools that can trace vulnerabilities from their root cause through all transitive callers to identify precisely which public API functions are affected. This prevents both over-flagging (by ignoring internal functions) and under-flagging (by missing public, but indirectly called, vulnerable functions).
  1. Maintain a Human-in-the-Loop Workflow with Rich Artifacts: Automation should augment, not entirely replace, human expertise. Implement processes where security researchers can review the outputs of automated rule generation. Tools should provide rich artifacts (like Brat's research summaries and code graphs) that allow humans to quickly audit the AI's reasoning, identify potential gaps, and ensure the completeness and accuracy of generated rules. This hybrid approach combines the speed of automation with the critical judgment of human experts.
  1. Focus on Reachable Vulnerabilities: Rather than simply flagging the presence of a vulnerable library version (which might lead to alert fatigue if the vulnerable code path isn't actually used), prioritize detecting reachable vulnerabilities. This means generating rules that identify specific misuse patterns or calls to vulnerable API endpoints, allowing teams to focus on actual risks in their codebase.
  1. Rapidly Expand Coverage for New Languages and Backlogs: Automated rule generation pipelines can significantly reduce the time and effort required to extend security coverage to new programming languages or frameworks. They also provide an efficient mechanism to clear backlogs of unaddressed CVEs, drastically improving an organization's security posture across its entire technology stack.
  1. Be Prepared for CVE Surges: The ability of tools like Brat to process dozens of CVEs in parallel means organizations can better withstand sudden spikes in vulnerability disclosures, ensuring timely protection against emerging threats without overwhelming their security teams.

In essence, the defensive implication is a call to move beyond ad-hoc, manual security processes and embrace intelligent automation. However, this automation must be built on a foundation of transparency, verifiability, and composability, ensuring that AI is used as a powerful, yet controlled, force multiplier in the ongoing battle against software supply chain attacks.

Key Takeaways

  • The Scale of Supply Chain Vulnerabilities Demands Automation: With an estimated 50,000 supply chain CVEs by 2025, manual rule writing is an unsustainable, error-prone, and unscalable approach to software supply chain defense.
  • Brat Automates Rule Generation, Significantly Reducing Human Effort: Semgrep's tool, Brat, automates the process of researching CVEs, generating test cases, and authoring Semgrep rules, reducing researcher time per rule from approximately one hour to 10-45 minutes and enabling the creation of thousands of rules.
  • "Tasteful AI" with Deterministic Code is Superior to Pure Agentic Workflows: Combining targeted AI assistance for specific subtasks (like parsing documentation) with deterministic code (like code graph analysis) creates a more robust, transparent, and verifiable security system than relying on black-box, unpredictable agentic AI.
  • Code Graph Analysis is Critical for Accurate Vulnerability Identification: Brat's core innovation is its static analysis-driven code graph, which traces vulnerabilities from patch commits through transitive callers to identify only the public, reachable API functions that are truly affected, preventing false negatives missed by less structured AI.
  • Modular, Verifiable Subproblems Enhance Reliability: Decomposing the rule generation process into independent, separately verifiable subtasks (e.g., generating tests and patterns for individual functions) prevents errors like dropped symbols, isolates failures, and allows for targeted improvements in the automation pipeline.
  • Human-in-the-Loop is Essential for Trust and Oversight: Brat generates detailed artifacts, such as research summaries and code graphs, enabling human security researchers to audit the automated output, ensuring accuracy, completeness, and overall trust in the generated security rules.

About the Speaker(s)

Brandon Wu is a Program Analysis Engineer at Semgrep, an application security platform dedicated to finding and fixing software security vulnerabilities. He has been working at Semgrep for approximately four years, focusing on profoundly improving software security through advanced program analysis techniques. Brandon studied Computer Science at Carnegie Mellon University and identifies as a strong enthusiast of functional programming. Beyond his professional endeavors, Brandon is an ardent musical theater fan, even participating in productions, and is a connoisseur of pop music, particularly from the "Brat Summer" era of 2024.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Wu presents a legitimate engineering solution to a real scaling problem: automated Semgrep rule generation for supply chain CVEs using call graph analysis plus targeted LLM assistance. The core technical idea — deterministic call graph traversal to find transitive public API exposure, with LLMs handling bounded subtasks like doc scraping and test snippet generation — is sound and the d-value case study makes the false-negative argument concrete. But this is a BSides-tier product engineering talk from a Semgrep employee about Semgrep infrastructure, and it never quite escapes that gravity.

Heather Calloway (CISO) — WEAK

Technically solid engineering talk about automating CVE rule generation at scale — the code graph approach is genuinely clever and the case study is concrete. But this is a product engineering session from a Semgrep employee about a Semgrep internal tool, and it never crosses into the territory that matters for security leaders: who owns the rule coverage gap, what should organizations demand from their vendors, and how does any of this change program-level decisions about supply chain risk.

→ Top-rated talks at BSidesSF 2026

All talks from BSidesSF 2026