Exploring ChatGPT's Capabilities on Vulnerability Management

Peiyu Liu, Junming Liu, Lirong Fu, Kangjie Lu, Yifan Xia, Xuhong Zhang, Wenzhi Chen, Haiqin Weng, Shouling Ji, Wenhai Wang

33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24

Overview

This talk presents a comprehensive evaluation of ChatGPT's capabilities across the entire vulnerability management lifecycle. Given the burgeoning interest in Large Language Models (LLMs) for diverse applications, including code-related analysis, researchers from various institutions explored how ChatGPT performs on a suite of six distinct vulnerability management tasks. Unlike prior work that often focused on isolated aspects of software engineering, this research provides the first large-scale, holistic assessment of an LLM's utility from initial bug reporting to final patch commit.

Watch on YouTube

Visual summary for Exploring ChatGPT's Capabilities on Vulnerability Management by Peiyu Liu, Junming Liu, Lirong Fu, Kangjie Lu, Yifan Xia, Xuhong Zhang, Wenzhi Chen, Haiqin Weng, Shouling Ji, Wenhai Wang
Visual summary for Exploring ChatGPT's Capabilities on Vulnerability Management by Peiyu Liu, Junming Liu, Lirong Fu, Kangjie Lu, Yifan Xia, Xuhong Zhang, Wenzhi Chen, Haiqin Weng, Shouling Ji, Wenhai Wang

Key moments

  1. 0:00 Introduction to vulnerability management challenges with LLMs
  2. 2:00 Three research questions and six vulnerability tasks studied
  3. 3:20 Prompt engineering methods including expertise templates
  4. 4:00 ChatGPT's outstanding performance in bug report summarization
  5. 5:30 ChatGPT's hallucinations in identifying security reports
  6. 6:45 ChatGPT self-generates prompts for improved performance
  7. 7:30 ChatGPT's strong capability in vulnerability repair
  8. 9:00 Code-only prompts outperform code with descriptions in assessment

Exploring ChatGPT's Capabilities on Vulnerability Management

Speakers: Peiyu Liu, Junming Liu, Lirong Fu, Kangjie Lu, Yifan Xia, Xuhong Zhang, Wenzhi Chen, Haiqin Weng, Shouling Ji, Wenhai Wang

Conference: USENIX Security '24

YouTube: https://www.youtube.com/watch?v=ZeIgqpc_rII

Overview

This talk presents a comprehensive evaluation of ChatGPT's capabilities across the entire vulnerability management lifecycle. Given the burgeoning interest in Large Language Models (LLMs) for diverse applications, including code-related analysis, researchers from various institutions explored how ChatGPT performs on a suite of six distinct vulnerability management tasks. Unlike prior work that often focused on isolated aspects of software engineering, this research provides the first large-scale, holistic assessment of an LLM's utility from initial bug reporting to final patch commit.

The motivation for this study stems from the inherent similarities between programming languages and natural languages, which suggests LLMs could be powerful tools for code analysis. However, the complex and multi-faceted nature of vulnerability management, requiring deep understanding of code syntax, program semantics, and extensive documentation, raises questions about an LLM's ability to assist maintainers effectively throughout this demanding process. This talk addresses critical questions regarding ChatGPT's performance relative to state-of-the-art (SOTA) approaches, the impact of prompt engineering methods, and promising future directions for enhancing LLM performance in this domain.

The findings offer significant insights for software maintainers, security professionals, and LLM researchers. By systematically evaluating ChatGPT on tasks ranging from bug report summarization to patch correctness assessment, the research identifies areas where LLMs excel, where they struggle, and crucially, how carefully crafted prompts can dramatically alter their effectiveness. This work not only highlights the immediate potential of LLMs in certain vulnerability management phases but also underscores the challenges and necessary advancements for their broader, reliable integration into security workflows.

Background

▶ Watch: Introduction to vulnerability management challenges with LLMs (0:00)

The rapid advancements in Large Language Models (LLMs), exemplified by models like ChatGPT, have led to their widespread adoption across various fields, from natural language processing (NLP) tasks like question answering and data augmentation to more specialized domains such as education. Within software engineering, the inherent structural and semantic similarities between natural languages and programming languages have naturally led researchers to explore LLMs for code-related analysis. Early studies demonstrated LLMs' foundational capabilities in tasks such as code generation, summarization, and basic code analysis, suggesting a powerful new paradigm for automating or assisting complex software development and security tasks.

However, a critical gap persisted in the existing research: most studies applying LLMs in software engineering tended to focus on specific, isolated tasks rather than the comprehensive vulnerability management lifecycle. This lifecycle is a complex, multi-phase process that begins with vulnerability discovery, progresses through confirmation, severity evaluation, and fixing, and concludes with patch committing. Each phase involves unique challenges and demands a sophisticated understanding of code, system context, and security implications. For instance, tasks like identifying security-related bug reports require discerning subtle security cues amidst general bug reports, while vulnerability repair necessitates precise code modification without introducing new flaws.

The absence of a holistic evaluation left it unclear whether an LLM like ChatGPT could truly assist software maintainers across the entire spectrum of vulnerability management tasks. This research aimed to fill that gap by providing a systematic, large-scale evaluation, comparing ChatGPT's performance against established SOTA approaches across diverse tasks within this critical security domain. Understanding these capabilities and limitations is crucial for responsibly integrating LLMs into cybersecurity practices and for guiding future research in this rapidly evolving field.

Key Findings

▶ Watch: Prompt engineering methods including expertise templates (3:20)

The research yielded several key findings regarding ChatGPT's capabilities in vulnerability management, highlighting both its strengths and limitations across different tasks and the significant impact of prompt engineering.

  1. Varied Performance Across Tasks: ChatGPT demonstrated outstanding performance in tasks closely resembling traditional Natural Language Processing (NLP), such as bug report summarization, where it achieved better correctness and readability than human-generated summaries in user studies. For more complex, code-centric tasks like vulnerability repair, it performed comparably to or even better than existing SOTA approaches and other LLMs, fixing 10 out of 12 vulnerabilities. However, in tasks requiring deeper contextual understanding or specific domain knowledge, such as security bug report identification and stable patch classification, its initial performance was inferior to SOTA, often exhibiting hallucinations or a lack of understanding of specific security concepts.
  1. Profound Impact of Prompt Engineering: The study unequivocally demonstrated that carefully designed prompt engineering methods significantly influence ChatGPT's performance.
  • Expertise prompts, which inject domain-specific knowledge or explicit definitions (e.g., defining what constitutes a "stable patch" or a "security bug report"), dramatically improved performance in tasks where ChatGPT initially struggled due to conceptual misunderstandings.
  • Self-heuristic prompts, where ChatGPT was prompted to generate its own knowledge or summarize relevant information from examples, proved highly effective for challenging tasks like vulnerability severity evaluation, leading to significant performance boosts.
  • The research also revealed that "more information is not always better." In patch correctness assessment, providing both code and description (the "describe code" prompt) sometimes led ChatGPT to focus on matching the description to the code rather than assessing actual correctness, with a "code only" prompt yielding superior results.
  1. Identified Bottlenecks and Future Directions: The study pinpointed specific reasons for ChatGPT's failures and areas for future improvement:
  • Lack of Context: For vulnerability repair, insufficient vulnerability-related context was a primary reason for failure, suggesting that advanced program slicing methods could provide more precise context.
  • Hallucinations and Misunderstandings: ChatGPT sometimes misunderstood security concepts (e.g., considering memory leaks or null pointer dereferences as non-security related), highlighting the need for better domain-specific training or explicit knowledge injection through prompts.
  • Information Overload/Misdirection: The observation that providing excessive or misaligned information could negatively impact performance in patch correctness assessment points to the need for guiding LLMs to use information appropriately.

In summary, the research established that while ChatGPT possesses strong foundational capabilities for certain vulnerability management tasks, its overall effectiveness is highly contingent on the task's nature and the sophistication of the prompt engineering applied. It serves as a foundational large-scale evaluation, paving the way for more targeted research to integrate LLMs effectively into the cybersecurity workflow.

Technical Deep Dive

▶ Watch: ChatGPT's hallucinations in identifying security reports (5:30)

The research methodology involved a systematic evaluation of ChatGPT across six distinct vulnerability management tasks, comparing its performance against 11 state-of-the-art (SOTA) approaches using large-scale datasets. The core of the evaluation hinged on carefully designed prompt templates and an analysis of ChatGPT's responses to identify specific bottlenecks.

The six vulnerability management tasks evaluated span the complete lifecycle:

  1. Bug Report Summarization: This task required ChatGPT to condense a given bug report into a concise summary. The evaluation found ChatGPT to deliver outstanding performance, with user studies indicating superior correctness and readability compared to traditional methods. This success is attributed to the task's similarity to general Natural Language Processing (NLP) tasks, where LLMs inherently excel.
  1. Security Bug Report Identification: Here, ChatGPT was asked to classify whether a given bug report was security-related. With advanced prompt templates, ChatGPT outperformed baseline models but could not match the capability of the SOTA approach, DKGA. A critical issue identified was hallucinations: ChatGPT mistakenly classified issues like memory leaks and null pointer dereferences as non-security-related. This was addressed by an expertise prompt that explicitly informed ChatGPT that "bug reports related to these problems should be seen as security bug reports," which significantly improved performance. The study also noted that one-shot prompts could lead ChatGPT to incorrectly mark reports containing unrelated words from the example as security-related, highlighting the challenge of guiding LLMs to focus on helpful information.
  1. Vulnerability Severity Evaluation: This task involved mapping a function's description to the Common Vulnerability Scoring System (CVSS) matrix. ChatGPT's initial performance was slightly inferior to the SOTA approach. A key innovation here was the use of a self-heuristic prompt. Researchers initially struggled to manually craft an effective prompt, then asked ChatGPT itself to summarize knowledge for this task from provided samples. This self-generated knowledge, incorporated into the prompt, significantly improved ChatGPT's performance, demonstrating an interesting future direction for LLM-assisted prompt engineering.
  1. Vulnerability Repair: This highly technical task required ChatGPT to fix vulnerable code snippets. Impressively, ChatGPT managed to fix 10 out of 12 vulnerabilities, performing comparably to program analysis based SOTA approaches (which also fixed 10) and outperforming other LLMs (which fixed 8). The primary reason for failure was identified as insufficient vulnerability-related context. The researchers suggested that more advanced program slicing methods could provide the specific context needed to further enhance ChatGPT's repairing capabilities in real-world applications.
  1. Patch Correctness Assessment: For this task, ChatGPT had to determine if a given patch correctly fixed a bug. With advanced prompts and models, ChatGPT performed comparably to SOTA approaches. Initially, ChatGPT struggled when only patch descriptions were provided. Collecting and providing both the code and description (the "describe code" prompt) improved performance, achieving results comparable to the Quadrant SOTA. However, a crucial and counter-intuitive finding emerged: an even simpler "code only" prompt often performed better. Manual analysis revealed that when both code and description were provided, ChatGPT tended to analyze whether the code changes matched the description rather than focusing on the actual correctness of the patch. This indicates that "more information is not always better" and that guiding ChatGPT on how to appropriately use information within a prompt is a vital research direction.
  1. Stable Patch Classification: This task asked ChatGPT to classify whether a given patch was "stable." ChatGPT performed slightly worse than the SOTA approach. Initially, with zero-shot and one-shot prompts, ChatGPT tended to classify all patches as stable. Manual analysis revealed this was because ChatGPT did not understand what a stable patch is. Providing a clear definition of a "stable patch" within an expertise prompt significantly improved its performance, underscoring the importance of unambiguous domain knowledge for LLMs.

Throughout these evaluations, the researchers manually developed prompt templates, including a "general info template" (using skills like giving a role) and an "expertise prompt template" (incorporating domain-specific knowledge). The systematic comparison across these tasks, using consistent metrics with SOTA approaches, allowed for a robust understanding of ChatGPT's current strengths and weaknesses in the complex landscape of vulnerability management.

Demo / Proof of Concept

▶ Watch: ChatGPT self-generates prompts for improved performance (6:45)

The talk focused on presenting the methodology and comprehensive evaluation results of ChatGPT's capabilities across various vulnerability management tasks. While the research detailed how prompts were constructed and how ChatGPT's responses were analyzed, it did not include a live demonstration or a specific Proof of Concept (PoC) tool that was built and shown. Instead, the presentation highlighted the quantitative and qualitative outcomes of the experiments, demonstrating ChatGPT's performance through data and analysis rather than an interactive demo.

Defensive Implications

▶ Watch: Code-only prompts outperform code with descriptions in assessment (9:00)

The findings from this extensive evaluation of ChatGPT in vulnerability management carry significant implications for security practitioners and defenders. Leveraging Large Language Models (LLMs) can introduce both powerful new capabilities and novel challenges into existing security workflows.

  1. Augmenting Human Analysts for NLP-heavy Tasks: Defenders can immediately benefit from LLMs in tasks like bug report summarization. LLMs can quickly distill complex, verbose bug reports into actionable summaries, saving triage time and improving communication. This allows human analysts to focus on deeper analysis rather than initial data digestion.
  1. Assisted Vulnerability Identification and Triage: While not a standalone solution, LLMs can assist in security bug report identification. By using well-crafted expertise prompts that define what constitutes a security vulnerability (e.g., explicitly stating that memory leaks or null pointer dereferences are security-relevant), organizations can improve the accuracy of initial filtering. However, human oversight remains critical to catch hallucinations and ensure correct classification, especially for nuanced or novel vulnerability types.
  1. Potential for Automated Patch Generation (with Caveats): The promising results in vulnerability repair suggest LLMs could be used to generate initial patch candidates. This could significantly accelerate the patching process, particularly for common vulnerability patterns. However, the identified need for "sufficient vulnerability-related context" implies that integrating LLMs with advanced program slicing methods or code analysis tools will be crucial to provide the precise information required for correct and secure fixes. All LLM-generated patches must undergo rigorous human review and testing to prevent the introduction of new bugs or security flaws.
  1. Careful Prompt Engineering is Paramount: The research strongly emphasizes that the effectiveness of LLMs is heavily dependent on prompt engineering. Defenders should invest time in developing and refining their prompts, incorporating domain-specific knowledge, clear definitions, and examples. The "self-heuristic prompt" concept is particularly valuable, suggesting that LLMs can be prompted to generate their own contextual knowledge, which can then be refined and used in subsequent tasks.
  1. Beware of Information Overload and Misdirection: The counter-intuitive finding in patch correctness assessment – that sometimes less information (e.g., "code only" vs. "code and description") leads to better results – is a critical lesson. Defenders should be mindful that providing too much extraneous or potentially misleading information can cause LLMs to focus on the wrong aspects of a problem. Prompts need to guide the LLM's attention precisely to the relevant data points and desired analytical focus.
  1. Human Oversight and Validation Remain Indispensable: Despite LLMs' impressive capabilities, they are not infallible. They can hallucinate, misunderstand context, or misinterpret security concepts. Therefore, every output from an LLM, especially in critical security tasks like vulnerability identification, severity evaluation, or patch generation, must be thoroughly reviewed and validated by a human expert. LLMs should be viewed as powerful assistants that augment human capabilities, not replacements for human judgment and expertise.

In essence, defenders should embrace LLMs as valuable tools for enhancing efficiency and augmenting analytical capabilities in vulnerability management, particularly for tasks involving large volumes of text or code. However, this adoption must be strategic, informed by the understanding of LLM limitations, and underpinned by robust prompt engineering practices and unwavering human supervision.

Key Takeaways

  • ChatGPT shows strong capabilities in vulnerability management, particularly for NLP-heavy tasks like bug report summarization, where it outperforms human-generated content in readability and correctness.
  • Prompt engineering is critical: advanced prompts, especially those injecting domain expertise or allowing ChatGPT to generate its own knowledge (self-heuristic prompts), significantly improve performance across various tasks.
  • LLMs can hallucinate and misunderstand security concepts (e.g., misclassifying memory leaks), necessitating explicit definitions and domain-specific knowledge injection through prompts.
  • More information is not always better; sometimes, providing less, highly relevant data (e.g., "code only" for patch assessment) can yield superior results by preventing the LLM from being distracted or misdirected.
  • ChatGPT can effectively assist with vulnerability repair, fixing 10 out of 12 vulnerabilities, but its performance is limited by insufficient context, suggesting a need for integration with advanced program slicing or code analysis tools.
  • Despite its capabilities, LLMs require human oversight and validation in critical security tasks due to potential for hallucinations and conceptual misunderstandings, serving as powerful assistants rather than autonomous agents.

About the Speaker(s)

The research presented was a collaborative effort by a team of ten researchers: Peiyu Liu, Junming Liu, Lirong Fu, Kangjie Lu, Yifan Xia, Xuhong Zhang, Wenzhi Chen, Haiqin Weng, Shouling Ji, and Wenhai Wang. Junming Liu delivered the presentation at USENIX Security '24. The transcript and metadata do not provide specific biographical details or affiliations for each individual beyond their names.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This research delivers a much-needed, rigorously executed evaluation of ChatGPT's true capabilities across the entire vulnerability management lifecycle. It systematically dissects where LLMs excel and fail, offering critical insights into the profound impact of prompt engineering and practical considerations for integrating these tools into security workflows. This isn't just another "AI-powered" fluff piece; it's a substantive, data-driven assessment that cuts through the hype.

Heather Calloway (CISO) — STRONG ACCEPT

This research provides a critical, nuanced evaluation of ChatGPT's role in vulnerability management, moving beyond hype to provide actionable insights. It clearly delineates where LLMs excel and where human oversight remains indispensable, emphasizing the profound impact of well-crafted prompt engineering. This work provides a practical roadmap for security leaders considering LLM integration into their operational workflows.

→ Top-rated talks at 33rd USENIX Security Symposium

All talks from 33rd USENIX Security Symposium