One-Click Code Fix: Securing Code Using AI

Chandrani Mukherjee, Joseph Seasly

BSidesSF 2024 · Day 1

Overview

This talk, presented by Chandrani Mukherjee and Joseph Seasly at BSidesSF 2024, explores the ambitious goal of leveraging Artificial Intelligence (AI) to automatically identify and remediate code vulnerabilities. The core premise is that AI can process code at speeds unmatched by humans, offering a potential paradigm shift in how security teams manage code vulnerabilities. However, the speakers immediately temper expectations, noting that AI-generated fixes are far from perfect, citing a benchmark where even ChatGPT-4 achieved only a 2% success rate for complex bug fixes. This underscores the critical need for human oversight and strategic integration of AI into existing security workflows.

Watch on YouTube

Visual summary for One-Click Code Fix: Securing Code Using AI by Chandrani Mukherjee, Joseph Seasly
Visual summary for One-Click Code Fix: Securing Code Using AI by Chandrani Mukherjee, Joseph Seasly

Key moments

  1. 0:00 AI for Code Fixes: Premise and Human Oversight
  2. 1:00 Limitations of Traditional Vulnerability Finding Methods
  3. 4:00 Strategic Reasons to Build Your Own AI Solution
  4. 7:00 Zero-Shot on Production: High False Positives and Hallucinations
  5. 11:00 Chain of Thought Method: Breakthrough in Fix Quality
  6. 13:00 Reality Check: Challenges with Complex and Multi-File Vulnerabilities
  7. 15:00 Implementing an Auto-Evaluation Framework for Fix Validation
  8. 22:00 Key Takeaways: Directed Fixes, Metrics, and Human Gatekeepers

One-Click Code Fix: Securing Code Using AI

Speakers: Chandrani Mukherjee, Joseph Seasly

Conference: BSidesSF 2024

YouTube: https://www.youtube.com/watch?v=kjF5SrCwHzg

Overview

This talk, presented by Chandrani Mukherjee and Joseph Seasly at BSidesSF 2024, explores the ambitious goal of leveraging Artificial Intelligence (AI) to automatically identify and remediate code vulnerabilities. The core premise is that AI can process code at speeds unmatched by humans, offering a potential paradigm shift in how security teams manage code vulnerabilities. However, the speakers immediately temper expectations, noting that AI-generated fixes are far from perfect, citing a benchmark where even ChatGPT-4 achieved only a 2% success rate for complex bug fixes. This underscores the critical need for human oversight and strategic integration of AI into existing security workflows.

The presentation delves into practical strategies for enhancing the accuracy and efficacy of AI in both finding and fixing vulnerabilities. It highlights the distinction between these two tasks and outlines an iterative journey of experimentation with various prompt engineering techniques. The speakers share their experiences, from initial proof-of-concept successes on dummy repositories to the challenges encountered when applying AI to complex production codebases, ultimately advocating for a human-in-the-loop approach that maximizes AI's strengths while mitigating its current limitations.

The talk is particularly relevant for security professionals, developers, and AI researchers interested in the practical application of large language models (LLMs) in software security. It provides a candid look at the current state of AI-driven code remediation, offering insights into effective prompt design, evaluation methodologies, and the defensive implications of integrating such powerful tools into the software development lifecycle. The ultimate vision is to free up expert security teams to focus on more intricate challenges by automating the resolution of simpler, well-understood vulnerabilities.

Background

▶ Watch: AI for Code Fixes: Premise and Human Oversight (0:00)

Traditionally, identifying code vulnerabilities has relied on a combination of labor-intensive and often imperfect methods. These include manual code reviews, Static Application Security Testing (SAST), Dynamic Application Security Testing (DAST), fuzzing, external bug bounty programs, incident reports, and penetration testing or red team reports. While these methods are foundational, they come with significant drawbacks. A common issue is the generation of numerous false positives, which can erode developer trust and lead to wasted effort. Furthermore, the coverage of vulnerabilities can vary widely depending on the detection method, and the entire process is often time-consuming, with manual reviews being particularly slow and prone to human error. Incident reports, by their nature, address vulnerabilities that have already been exploited, representing a reactive rather than proactive security posture.

The advent of AI presents an opportunity to address these shortcomings. The goal is to increase coverage, boost true positives, and decrease false positives, though perfect coverage remains an elusive target. The speakers explored two primary AI-driven approaches for finding vulnerabilities: an AI-only approach using simple prompts, and a fine-tuned model approach (e.g., GPT-3.5 fine-tuning). Their initial attempts with fine-tuning using a few hundred real vulnerability examples were not highly successful. A more promising strategy involved a prompt-driven approach that supplements results from existing SAST tools, integrating AI into current frameworks rather than replacing them entirely. AI can be applied at various stages, such as during pull requests (PRs), for retrospective code checks, or even across entire repositories, though the latter presents a greater challenge.

Similarly, traditional methods for fixing vulnerabilities are heavily reliant on human effort. Security teams face challenges in prioritization, determining which vulnerabilities to address first. Fixing requires a skilled team, is time-consuming, and humans are, again, prone to error. Instances where a patch fails, requiring subsequent fixes, or where the same vulnerability reappears across multiple repositories, highlight the inefficiencies of purely manual remediation.

AI offers a compelling alternative for fixing vulnerabilities, particularly for simple code vulnerabilities, by generating fixes that can then be validated by humans. This approach allows expert security teams to dedicate their valuable time to resolving the more complex fixes that AI currently cannot handle. While commercial solutions for AI-driven code security exist, the speakers argued for building an in-house solution. Key reasons include:

  • Proprietary coding languages, styles, and libraries: Organizations often have unique codebases, such as custom XSS fixing libraries, that commercial solutions may not adequately support or maintain.
  • High cost: Commercial solutions can incur millions of dollars annually for large development teams.
  • Customizable measurability: Building in-house allows for tailored measurement of historical performance in finding and fixing vulnerabilities, leveraging existing data from tools like Jira and GitHub.
  • Internal expertise growth: Developing internal AI capabilities reduces dependency on external vendors and fosters valuable in-house knowledge.

This foundational understanding of traditional methods and their limitations, coupled with the strategic rationale for AI integration and in-house development, set the stage for the team's solution journey.

Key Findings

▶ Watch: Strategic Reasons to Build Your Own AI Solution (4:00)

The journey into AI-driven code security was characterized by significant trial and error, with the presented work primarily leveraging GPT-4. The methodology, however, is designed to be language model agnostic.

The initial Proof of Concept (PoC) involved a zero-shot prompt engineering method on a dummy repository containing 8-10 obvious, straightforward vulnerabilities. Relying on the model's intrinsic knowledge and a system prompt praising it as an "expert AI security engineer," GPT-4 successfully identified and generated fixes for all injected vulnerabilities with high precision. This early success generated high hopes for production application.

However, applying the same zero-shot technique to a production repository revealed significant challenges. Despite an advanced system prompt specifying top 15 vulnerability types, a list of languages, and a strict output response schema, the model produced numerous false positives and hallucinated fixes. For instance, it might suggest "include a sanitized function" without providing a proper definition or implementation.

To address these issues, the team transitioned to a single-shot method, providing the model with a single example of the desired output. The input strategy shifted to using Jira tickets, which contained logged vulnerabilities, vulnerable code snippets, and related repository information. While this improved detection for very similar code, variable tracing within the same file remained problematic. The model struggled to connect variable definitions to their later usage, especially when inputs were processed and then used in sensitive contexts like window.URL. Furthermore, it failed to generalize across different types of the same vulnerability (e.g., various forms of injection).

Moving to a few-shot method, where three to five examples of different injection types were provided, showed improvements in vulnerability detection. However, the quality of the generated fixes still required substantial refinement to be considered acceptable.

A significant breakthrough came with the adoption of the Chain of Thought method, inspired by a Boston University paper that reported a 70% precision rate for vulnerability detection. This method involves not just providing a problem statement but also explaining the reasoning behind a decision, mimicking human thought processes. For example, when presenting a vulnerable code snippet, the prompt would explain why it's vulnerable (e.g., user input leading to malicious JavaScript execution via window.URL) and then provide a proper, defined fix, explaining why that fix works. This approach led to substantial improvements in both detection and the quality of fixes, particularly for injection vulnerabilities. A crucial limitation noted was its effectiveness primarily for vulnerabilities contained within the same file, performing poorly for issues spread across multiple files.

The "reality check" highlighted several persistent challenges:

  • Complex vulnerabilities: GPT-4 performed poorly on vulnerabilities spanning multiple files or repositories.
  • Lack of proper test files: In production environments, the true number of vulnerabilities (the denominator for precision calculations) is often unknown, making accurate evaluation difficult.
  • Correctness of the fix: This is subjective. A "correct" fix might need to include full function definitions, proper library imports, constructor changes, and other context-specific code elements, not just high-level suggestions.

To scale evaluation beyond manual review, an auto-eval framework using LangChain's scoring evaluator was implemented. This involved creating a "golden data set" of vulnerable code and corresponding good fixes. LLM-generated fixes were then compared against this reference, with LangChain's accuracy calculator assigning a score from 1 to 10. A score closer to 70% was deemed acceptable for a fix. The primary metrics tracked were Precision (correct issues found / total issues found) and Accuracy (quality of the fix compared to the reference). The Chain of Thought method consistently yielded the most promising results in these evaluations.

Key prompt strategies learned throughout this journey include:

  1. Reducing LLM laziness: Explicitly instructing the model that "placeholder ellipses and other shortcuts will never be used in place of functional code" helps ensure complete code generation.
  2. Proper output schema: Defining file formats, schema details, and specific output fields helps structure the AI's response, though syntactically incorrect output can still occur.
  3. Reducing hallucinations: Providing specific output fields for elements like "Library Imports," "Method Creation," or language-specific constructs (e.g., "Constructor changes") guides the LLM to generate relevant and accurate code rather than fabricating details.

Ultimately, the findings underscore that directed fixes (providing more hints) are more effective, and LLMs are well-suited for finding and fixing uncomplicated, 100% known vulnerabilities (e.g., those validated by bug bounty programs). Metrics are key for measuring effectiveness, and Chain of Thought prompting combined with well-designed prompts and output schemas significantly enhances performance. Crucially, humans remain the ultimate gatekeepers for final validation.

Technical Deep Dive

▶ Watch: Chain of Thought Method: Breakthrough in Fix Quality (11:00)

The technical journey began with an exploration of various prompt engineering techniques, primarily with GPT-4, aiming for a methodology that could be language model agnostic.

The initial zero-shot prompt engineering approach relied on the LLM's inherent knowledge. For a dummy repository, a simple system prompt, such as "You are a very expert AI security engineer, you're very methodical," was sufficient to identify and fix 8-10 obvious vulnerabilities. However, when scaled to a production repository, this approach proved inadequate. An advanced zero-shot strategy was then employed, where the system prompt was augmented with specific instructions:

  • Targeted vulnerability types: A list of the top 15 vulnerabilities from Jira was provided.
  • Language awareness: A list of programming languages was given to ensure context-aware fix generation.
  • Output response schema: A strict JSON-like schema was defined, specifying required fields and format, to prevent free-form, unstructured responses.

Despite these enhancements, the model produced a high rate of false positives and hallucinated fixes, where it would suggest a solution conceptually (e.g., "include a sanitized function") but fail to provide the actual, functional code definition.

To mitigate hallucination and improve context, the team moved to a single-shot method. Here, a single example of a vulnerable code snippet and its correct fix was provided within the prompt. The input strategy shifted from raw production code to Jira tickets, which often contained:

  • Vulnerable code snippets.
  • Details of how the vulnerability was exploited.
  • Related repository information.

This approach aimed to teach the model by example. While it showed some success for highly similar code patterns, it struggled with variable tracing within the same file. For instance, if a user input variable was defined early, transformed, and then used in a sensitive sink (like window.URL) much later in the code, the model often failed to connect the entire data flow. It also lacked generalization, struggling with different manifestations of the same vulnerability type (e.g., various forms of XSS or SQL injection).

The few-shot method extended this by providing three to five examples of different types of a specific vulnerability (e.g., multiple injection examples). This helped the model better understand the context and task, leading to improved detection rates. However, the quality of the generated fixes still required significant human intervention.

The most impactful technical advancement was the adoption of the Chain of Thought method, inspired by research from Boston University. This method compels the LLM to articulate its reasoning process, mirroring human decision-making. Instead of just providing a vulnerable snippet and asking for a fix, the prompt structure included:

  1. Problem Statement: The vulnerable code snippet and relevant Jira ticket summary/description.
  2. Explanation of Vulnerability: A detailed breakdown of why the code is vulnerable (e.g., "This param input is coming from user input. If a malicious JavaScript is loaded, then when the window.URL gets executed, that gets executed in your environment.").
  3. Proposed Fix: The actual code fix, including proper function definitions, import statements, and other necessary code.
  4. Explanation of Fix: A justification of why the proposed fix works.

This explicit reasoning process significantly improved both the detection accuracy and the quality of the generated fixes, particularly for injection vulnerabilities. A critical caveat, however, is that this method was most effective for vulnerabilities contained within a single file, struggling with issues spread across multiple files.

For metrics and evaluation, the team initially relied on manual review, which was not scalable. To automate this, they implemented an auto-eval framework using LangChain's scoring evaluator. This framework involved:

  1. Golden Data Set: A curated collection of vulnerable code snippets paired with their corresponding, validated "good fixes."
  2. LLM Fix Generation: The same vulnerable code from the golden data set was fed to the LLM to generate a fix.
  3. Comparison and Scoring: LangChain's evaluator compared the LLM-generated fix against the reference fix in the golden data set. An accuracy score from 1 to 10 was assigned, with 70% (or 7/10) being the threshold for an acceptable fix.

The primary metrics tracked were Precision (the ratio of correct issues found to the total issues reported by the LLM) and Accuracy (the quality of the fix as determined by the scoring evaluator).

Further technical refinements involved specific prompt strategies to enhance LLM performance:

  • Combating LLM Laziness: Prompts explicitly included directives like "placeholder ellipses and other shortcuts will never be used in place of functional code" to ensure complete and functional code generation.
  • Enforcing Output Schema: Detailed output schemas specified file formats, field names, and expected data types. Even with this, occasional syntactically incorrect output was observed.
  • Reducing Hallucinations via Structured Fields: To prevent the LLM from inventing non-existent code, specific output fields were defined for common code elements, such as "Library Imports," "Method Creation," or language-specific constructs like "Constructor changes." This guided the LLM to populate these fields with actual, relevant code.

An adaptive few-shot mechanism was also developed, where if the vulnerability type was known beforehand, the system would dynamically pull relevant few-shot examples from a pre-curated set specific to that vulnerability type.

The vulnerabilities primarily experimented with were XSS and various forms of injections. The overall technical journey highlighted the iterative nature of prompt engineering and the necessity of robust evaluation frameworks for practical AI application in security.

Demo / Proof of Concept

▶ Watch: Reality Check: Challenges with Complex and Multi-File Vulnerabilities (13:00)

The speakers demonstrated a practical application of their AI-driven code fixing solution. They mentioned three different implementation approaches:

  1. A Slack-powered AI agent for direct code fixing.
  2. A CLI-based infrastructure designed to iterate over entire repositories, allowing users to select and improve fixes for pull requests (PRs) later.
  3. A simple command-line interface (CLI) demonstration, which was the focus of the live presentation.

The demo itself was concise, lasting approximately 25 seconds. In this demonstration, a command was issued directly from the command line. The AI agent then processed the request, generated a fix for a detected vulnerability, and subsequently issued a pull request (PR). Crucially, the system not only provided the code fix but also explained the nature of the vulnerability and the rationale behind the proposed fix. This immediate feedback loop, coupled with the automated PR generation, showcased the "one-click" aspect of their solution, streamlining the remediation process while maintaining human oversight through the PR review mechanism.

Defensive Implications

▶ Watch: Key Takeaways: Directed Fixes, Metrics, and Human Gatekeepers (22:00)

The insights gleaned from this research offer several critical defensive implications for organizations looking to enhance their security posture with AI:

  • Augment, Don't Replace, Security Teams: AI, particularly for simple and well-understood vulnerabilities, can significantly offload the burden from human security experts. This frees up skilled teams to focus on complex, multi-file, or novel vulnerabilities that currently exceed AI capabilities.
  • Proactive Vulnerability Remediation: Integrating AI into the development pipeline, especially at the pull request (PR) stage or for retrospective code scans, enables earlier detection and automated remediation. This shifts security left, reducing the cost and effort associated with fixing vulnerabilities later in the development lifecycle.
  • Prioritize Known Vulnerabilities: Organizations should leverage AI to quickly address 100% known vulnerabilities, such as those validated through bug bounty programs. This ensures that easily fixable issues are resolved rapidly, improving the overall security baseline.
  • Strategic In-House Development: For organizations with proprietary codebases, unique coding standards, or specific security libraries (e.g., custom XSS sanitization), building an internal AI solution offers greater customization, cost control, and the ability to grow internal AI expertise, reducing reliance on potentially generic or expensive commercial offerings.
  • Mandatory Human Oversight: Despite AI's capabilities, humans must remain the ultimate gatekeepers. All AI-generated fixes, especially those destined for production, require thorough human validation to prevent the introduction of new bugs, incorrect fixes, or security regressions. The system should only issue PRs, never directly commit to main branches.
  • Invest in Prompt Engineering and Schema Design: The success of AI in security is heavily dependent on the quality of prompts. Adopting techniques like Chain of Thought and designing robust output schemas are crucial for improving detection accuracy, fix quality, and reducing hallucinations. This requires dedicated effort and expertise in prompt engineering.
  • Establish Robust Evaluation Frameworks: Continuous measurement of AI's effectiveness using metrics like precision and accuracy, coupled with auto-evaluation frameworks (e.g., LangChain's scoring evaluator), is essential. This allows security teams to track performance over time, refine prompts, and ensure the AI is delivering tangible security value.
  • Mitigate AI Access Risks: When providing AI access to internal code, organizations must implement strict security controls. The speakers explicitly stated their solution only issued PRs and did not directly modify the main branch, and they used internal/zero-trust models (their own company's OpenAI instance) to avoid sending sensitive internal code to external model APIs. This is a critical consideration for data privacy and intellectual property protection.
  • Focus on Single-File Vulnerabilities First: Given the current limitations of AI with multi-file vulnerabilities, defenders should strategically deploy AI for issues contained within single files, where its performance is more reliable.

By thoughtfully integrating AI with these defensive considerations, organizations can significantly enhance their ability to manage and remediate code vulnerabilities, making their software more secure and their security teams more efficient.

Key Takeaways

  • Directed fixes are superior: Providing more specific hints and context to the LLM significantly improves the quality and accuracy of generated code fixes.
  • LLMs excel at simple, known vulnerabilities: AI is highly effective for finding and fixing uncomplicated vulnerabilities that are 100% known and validated, such as those identified through bug bounty programs.
  • Metrics and auto-evaluation are crucial: Establishing clear metrics (precision, accuracy) and implementing automated evaluation frameworks (e.g., LangChain's scoring evaluator) are essential for measuring, tracking, and continuously improving the effectiveness of AI-generated fixes.
  • Chain of Thought is a powerful prompting strategy: This method, which encourages the LLM to explain its reasoning, dramatically improves both vulnerability detection and the quality of fixes, particularly for issues contained within a single file.
  • Well-designed prompts and output schemas are non-negotiable: Meticulously crafted prompts that reduce LLM laziness, enforce structured output, and guide the model with specific fields are key to minimizing hallucinations and generating functional code.
  • Humans remain the final authority: Despite AI's advancements, human security experts are indispensable for the final validation and approval of all AI-generated code fixes, ensuring accuracy and preventing the introduction of new issues.

About the Speaker(s)

Chandrani Mukherjee and Joseph Seasly are the speakers who presented this detailed technical talk at BSidesSF 2024. Their work focuses on the practical application of AI, specifically large language models, to enhance code security by automating the detection and remediation of software vulnerabilities.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk provides a pragmatic and technically grounded journey into leveraging LLMs for automated code vulnerability fixing. The speakers detail their iterative process, from initial zero-shot attempts to the more effective Chain of Thought prompting, highlighting both successes and significant challenges. It's a solid engineering effort that offers actionable insights for teams looking to integrate AI into their secure development lifecycle, emphasizing the critical role of human oversight and robust evaluation frameworks.

Heather Calloway (CISO) — STRONG ACCEPT

This session offers a pragmatic look at integrating AI into the secure development lifecycle for automated vulnerability fixing. The speakers effectively articulate the operational benefits, such as freeing up skilled security teams for more complex issues and the potential for significant cost savings by building in-house solutions. Crucially, they underscore the necessity of human oversight and robust evaluation, which is vital for maintaining governance and accountability in an AI-augmented environment.

→ Top-rated talks at BSidesSF 2024

All talks from BSidesSF 2024