From Large to Mammoth: A Comparative Evaluation of Large Language Models in Vulnerability Detection
Jie Lin (University of Central Florida)
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Vulnerability Detection
Overview
This article delves into a comprehensive study presented at the NDSS Symposium, titled "From Large to Mammoth: A Comparative Evaluation of Large Language Models in Vulnerability Detection." Presented by Jie Lin from the University of Central Florida, the research explores the burgeoning potential of Large Language Models (LLMs) in identifying security vulnerabilities within source code. The talk addresses a critical gap in the understanding of how various architectural and operational factors—such as model size, context window capacity, and quantization techniques—influence an LLM's accuracy and reliability in this specialized domain.
Key moments
- 0:00 Introduction and study's five guiding questions
- 1:30 Data preparation from GitHub and manual filtering
- 3:15 Baseline results: Llama 3 excels in vulnerable-only detection
- 5:00 Java experiment: No simple rules for performance
- 6:00 C/C++ results: Larger models often underperform smaller ones
- 7:00 Few-shot learning for Java: Unexpected performance drop
- 7:40 Few-shot learning for C/C++: Performance collapses to 0%
- 8:10 Shifting focus to classifying vulnerability types
From Large to Mammoth: A Comparative Evaluation of Large Language Models in Vulnerability Detection
Speakers: Jie Lin (University of Central Florida)
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=veJ1H1A-_os
Overview
This article delves into a comprehensive study presented at the NDSS Symposium, titled "From Large to Mammoth: A Comparative Evaluation of Large Language Models in Vulnerability Detection." Presented by Jie Lin from the University of Central Florida, the research explores the burgeoning potential of Large Language Models (LLMs) in identifying security vulnerabilities within source code. The talk addresses a critical gap in the understanding of how various architectural and operational factors—such as model size, context window capacity, and quantization techniques—influence an LLM's accuracy and reliability in this specialized domain.
The significance of this research stems from the increasing reliance on automated tools for large-scale vulnerability detection. While traditional static and dynamic analysis tools have long been mainstays, LLMs offer a paradigm shift by analyzing code with a human-like understanding and providing end-to-end verdicts. However, their real-world viability and the specific conditions under which they excel or falter remained largely undefined. This study provides much-needed clarity by systematically evaluating a diverse array of LLM architectures, highlighting their respective strengths, weaknesses, and practical applicability.
By dissecting the performance of multiple LLMs across different programming languages and experimental setups, the research aims to guide developers and security professionals in making informed decisions about integrating these powerful models into their security workflows. It underscores the complex interplay of model parameters and environmental factors, demonstrating that simply using a "larger" or "newer" model does not guarantee superior vulnerability detection capabilities.
Background
▶ Watch: Introduction and study's five guiding questions (0:00)
The landscape of software security has long been dominated by traditional static and dynamic analysis tools. These specialized instruments meticulously examine source code or runtime behavior to identify patterns indicative of vulnerabilities, often requiring intricate configurations and external rule sets. While effective, they can sometimes be rigid, prone to false positives, or struggle with the nuanced understanding of code logic that a human expert possesses. The advent of Large Language Models has introduced a compelling alternative, promising an "end-to-end" analysis capability that mimics human reasoning in discerning vulnerable code segments.
However, despite their impressive capabilities in code understanding and generation, the application of LLMs to vulnerability detection has been met with a lack of clarity regarding their practical efficacy. Key questions persist about how fundamental attributes of these models—such as their model size (number of parameters), context window (the amount of text they can process at once), and quantization (a technique to reduce model size and improve inference speed by using lower precision numbers)—impact their ability to accurately detect vulnerabilities. Without a systematic evaluation, the strengths, weaknesses, and real-world viability of different LLM architectures for this task remain ambiguous.
To address this critical knowledge gap, the research outlined five guiding questions for their study:
- Accuracy Measurement: How accurately do LLMs detect vulnerabilities at the file level?
- Context Window Impact: Does a larger context window improve detection results?
- Quantization Effect: Does quantization negatively affect accuracy?
- Model Comparison: How do newer LLM generations (e.g., Llama 3) compare to older ones (e.g., Llama 2)?
- Few-Shot Learning: Can a few labeled examples improve detection accuracy?
These questions form the bedrock of a comprehensive comparative evaluation, designed to shed light on the optimal configurations and strategies for leveraging LLMs in the demanding field of software vulnerability assessment.
Key Findings
▶ Watch: Baseline results: Llama 3 excels in vulnerable-only detection (3:15)
The comprehensive evaluation yielded several pivotal insights into the performance and characteristics of Large Language Models in vulnerability detection:
- **LLMs Do Work, But Performance Varies Sharply by Language: The study conclusively demonstrates that LLMs are viable for vulnerability detection. However, their efficacy is highly dependent on the programming language, with Java code generally yielding significantly better detection rates compared to C and C++**.
- Difficulty in Type Identification: While LLMs can often flag a file as vulnerable, they largely fail to identify the exact vulnerability type. Multiclass labeling for specific CVEs or vulnerability categories proved to be a significantly tougher challenge than simple binary classification.
- Larger Context Windows Generally Help: A consistent trend observed was that larger context windows tend to boost accuracy and reduce instances of hallucination, where models generate plausible but incorrect information.
- Quantization Impact is Architecture-Dependent: The effect of quantization on accuracy is not uniform; it varies significantly across different LLM architectures. For instance, JAMAMA performed better in FP16 precision, while COJAMA showed a preference for Q5KM quantization.
- Advanced Models and Few-Shot Learning Often Underperform: Counterintuitively, newer or more advanced models (ee.g., Llama 3 compared to Llama 2, or newer Mistral variants) were not guaranteed to outperform their predecessors. More strikingly, few-shot learning, where a minimal number of labeled examples are provided, frequently led to a decrease in detection accuracy, and in C/C++ experiments, it even caused performance to collapse to 0% in some metrics. This suggests that naive application of few-shot prompting can confuse the models.
- Bigger Parameter Counts Don't Consistently Yield Better Performance: The research found no simple correlation between a larger number of model parameters and higher accuracy. Some larger architectures even underperformed smaller ones, with the Llama 3 70B variant notably failing to detect any vulnerabilities in certain C/C++ tests.
- Open-Source Contenders Can Outperform Proprietary Models: Remarkably, open-source models like Jamba and Llama 2 demonstrated the capability to outperform proprietary models such as GPT-4 in certain detection tasks, suggesting practical and cost-effective solutions are available.
- Complex Interplay of Factors: Overall, the findings underscore that there are no simple rules for LLM performance in vulnerability detection. Success relies on a complex interplay of model size, architecture, context window, quantization strategy, and careful prompt design, often necessitating domain-focused tuning.
Technical Deep Dive
▶ Watch: C/C++ results: Larger models often underperform smaller ones (6:00)
The study employed a rigorous experimental pipeline to systematically evaluate 38 distinct Large Language Model configurations across Java and C/C++ vulnerability detection tasks. The methodology encompassed data preparation, a diverse model selection, precise prompting strategies, and a multi-faceted evaluation framework.
Data Set Preparations:
The researchers sourced their code samples from two primary datasets: Big-J for Java vulnerabilities and Big-V for C and C++ vulnerabilities. After retrieving relevant source code files from GitHub, a crucial preprocessing step involved using Tricitor to remove comments and non-code artifacts, ensuring that the LLMs focused solely on executable logic. This was followed by a manual filtering process to eliminate duplicates or incomplete data. The final curated datasets comprised 280 files for Java and 200 files for C and C++, with each language's dataset evenly split between vulnerable and non-vulnerable code examples.
Model Selection and Configuration:
A wide array of LLMs was tested, encompassing various architectures, model parameter sizes, and context window capabilities. GPT-4 was included primarily as a reference point for comparison against open-source and specialized models. The models were evaluated under consistent conditions: a fixed temperature of 0.5 to control randomness, a seed of 42 for reproducible outputs, and an output size capped at 2048 tokens to prevent excessive generation.
Prompting Strategy:
For basic vulnerability detection, a straightforward "yes/no" prompt was used, tailored to the specific language (e.g., referencing Java or C/C++). For vulnerability type identification, the prompt asked for a "CV ID and a short description."
Evaluation Metrics:
The study utilized several key metrics to assess LLM performance:
- AP (Accuracy Positive / File-level Accuracy): The fraction of files correctly flagged as vulnerable.
- VAP (Vulnerability Accuracy Positive): The fraction of vulnerabilities identified at least once across all files associated with that vulnerability.
- EP (Explicit Positive / Clear Verdict): The fraction of files where the model provided a clear "yes" or "no" statement.
- VP (Vulnerability Verdict): The number of vulnerabilities for which the model provided at least one explicit verdict.
- Precision, Recall, and F1-score: Standard metrics used for classification tasks, particularly in the few-shot experiments.
- C (Correct Type): For vulnerability type identification, this measured if the model named the correct vulnerability type.
Experimental Pipeline and Results:
- Baseline Test (Vulnerable-Only Input):
The initial experiment involved inputting only vulnerable code files to the models. This aimed to assess their reliability in detecting known vulnerabilities under a simplified condition. Jamba emerged as a strong performer, achieving nearly 79% on AP and 93% on VAP. Interestingly, models like Colama and GPT-4 performed lower than anticipated, illustrating that specialized or proprietary models are not inherently guaranteed to excel. In terms of explicit verdicts (EP/VP), GPT-4 and Mistral topped at 100%, never missing a clear "yes" or "no," with Colama and Llama 3 also scoring high. This phase revealed a complex interplay of model size, quantization, and architecture, with no simple rules for performance.
- Zero-Shot Detection (Vulnerable and Non-Vulnerable Files):
Expanding beyond vulnerable-only inputs, this phase introduced non-vulnerable (safe) code, simulating a more realistic scenario. The maximum context window for each model was utilized, given that larger context windows generally boost performance.
- Java Experiment: Patterns emerged emphasizing precision in some models and recall in others. Crucially, larger parameter counts did not consistently yield better performance, and quantization preferences varied by family. Advanced or domain-focused variants did not automatically surpass earlier generations.
- C and C++ Experiment: This mirrored the Java setup, but with C/C++ specific prompts. The results highlighted that quantization gains were heavily model-family dependent; for example, JAMAMA performed better in FP16, while COJAMA preferred the Q5KM quantization. Alarmingly, larger architectures often underperformed smaller ones, with the Llama 3 70B variant even failing to detect any vulnerabilities. Newer Mistral models also lagged behind their predecessors. These findings strongly suggested that neither "bigger" nor "more advanced" models automatically equate to better performance.
- Few-Shot Learning Experiment:
To investigate if minimal guidance could improve accuracy, two labeled examples (one vulnerable, one non-vulnerable) were added to the prompt. These examples were explicitly removed from the main dataset to prevent leakage.
- Java Results: Surprisingly, the Java experiment showed a consistent drop in performance from zero-shot to few-shot. This decline was attributed to potential issues like the expanded prompt overshadowing the actual code or the example snippets introducing conflicting cues, leading some models to become overly cautious and experience steep declines in recall.
- C and C++ Results: The impact was even more severe for C/C++. This approach caused performance to collapse to 0% in precision, recall, and F1-score. Some models misclassified everything as safe, while others split mistakes between false positives and negatives. This inconsistent behavior strongly suggested that minimal, "naive" few-shot prompts could actively confuse the models in C/C++. The speaker further clarified during Q&A that they used a "very very naive way for the few-shot learning," by randomly selecting examples, rather than retrieving similar code patterns or vulnerabilities.
- Vulnerability Type Identification:
This experiment pushed models to identify which type of vulnerability existed, focusing only on files known to be vulnerable. The prompt asked for a CV ID and a short description.
The results indicated that while some models maintained moderate AP (still flagging files as vulnerable), virtually all failed at identifying the specific type, with most scores remaining at zero or very few correct cases. Switching to few-shot often lowered the overall accuracy in this task as well. This unequivocally highlighted that multiclass labeling (identifying specific vulnerability types) is significantly tougher for current LLMs than simply detecting if code is vulnerable.
Runtime Analysis:
The study also examined the practical implications of different configurations on runtime. Larger context windows were found to slow down the prompt analysis phase but could speed up the response generation phase. Bigger models, such as Llama 2 70B, inherently required longer processing times. Quantization, like Q5KM, effectively cut down this time, striking a balance between performance and resource demands.
Data Leakage Considerations:
During the Q&A, the speaker addressed concerns about data leakage. While acknowledging the impossibility of preventing leakage from the vast, unknown training data of pre-trained LLMs, the researchers implemented measures within their evaluation. They fully offloaded and reloaded models for each prompt to prevent previous system prompts from influencing subsequent ones. For few-shot learning, they explicitly excluded the examples used in the prompts from the main evaluation datasets.
Demo / Proof of Concept
▶ Watch: Few-shot learning for Java: Unexpected performance drop (7:00)
The presented work is a comprehensive comparative evaluation study of Large Language Models in vulnerability detection. As such, the talk did not include a live demonstration or a specific proof-of-concept tool. Instead, it focused on the systematic analysis of various LLM configurations and their empirical performance metrics.
Defensive Implications
▶ Watch: Shifting focus to classifying vulnerability types (8:10)
The findings from this extensive evaluation offer crucial insights for security defenders seeking to leverage Large Language Models in their efforts to secure software.
- Strategic LLM Adoption: While LLMs show promise, they are not a silver bullet. Defenders should approach their integration with a clear understanding of their current limitations and strengths. They are best viewed as a powerful supplement to existing security tools, not an outright replacement.
- Language-Specific Tooling: The significant disparity in performance between Java and C/C++ vulnerability detection demands a differentiated strategy. Defenders should recognize that LLMs are currently more effective for Java codebases. For C/C++ projects, a higher degree of human oversight, specialized static analysis tools, or a combination of approaches remains critical.
- Focus on Detection, Not Precise Classification (Yet): Current LLMs are relatively adept at binary classification—identifying whether a file is "vulnerable" or "not vulnerable." However, they struggle profoundly with pinpointing the exact type of vulnerability (e.g., a specific CVE ID). Defenders should use LLMs primarily as an initial filter for suspicious code, rather than relying on them for precise vulnerability categorization or root cause analysis, which still requires human expertise or specialized tools.
- Optimize Context Window Usage: The study highlights that larger context windows generally improve accuracy and reduce hallucinations. When configuring LLMs, defenders should strive to utilize the largest feasible context window given their computational resources, as this can directly impact detection quality.
- Evaluate Quantization Trade-offs: Quantization can significantly reduce runtime and resource demands, making LLMs more practical for deployment. However, its impact on accuracy is model-architecture dependent. Defenders should empirically test different quantized versions of their chosen LLM to find the optimal balance between performance, resource efficiency, and detection accuracy for their specific use case.
- "Newer" and "Bigger" Aren't Always "Better": A key takeaway is to avoid the assumption that the latest LLM generation or models with the most parameters will automatically yield superior results. The Llama 3 70B variant's failure in some C/C++ tests, and the underperformance of newer Mistral models, underscore the need for empirical validation. Defenders must evaluate models based on their actual performance against specific vulnerability detection tasks, rather than marketing claims or model size alone.
- Invest in Prompt Engineering: The detrimental effect of "naive" few-shot learning, particularly its collapse in C/C++ scenarios, is a stark warning. Defenders or security engineers leveraging LLMs must invest heavily in sophisticated prompt engineering techniques. This includes crafting precise zero-shot prompts and carefully designing few-shot examples that provide clear, non-conflicting guidance, potentially requiring domain-specific knowledge and iterative refinement.
- Leverage Open-Source Solutions: The finding that open-source models like Jamba and Llama 2 can outperform proprietary options like GPT-4 is highly encouraging for defenders. This opens avenues for more cost-effective and transparent vulnerability detection solutions, potentially allowing for greater customization and self-hosting, which can be crucial for sensitive codebases.
- Hybrid Approach is Key: Ultimately, LLMs are powerful additions to the security toolkit but should not be seen as standalone solutions. A robust defensive strategy will likely involve a hybrid approach, combining the broad, human-like understanding of LLMs for initial triage with the precision and depth of traditional static analysis, dynamic analysis, and expert human review for thorough validation and remediation.
Key Takeaways
- Large Language Models demonstrate promise for vulnerability detection, but their performance is highly dependent on the programming language, with significantly better results observed for Java compared to C/C++.
- Larger context windows generally contribute to improved detection accuracy and a reduction in model hallucinations, making them a crucial factor in LLM configuration for security tasks.
- Neither larger parameter counts nor newer model architectures automatically guarantee superior vulnerability detection performance; empirical testing against specific codebases is essential for selecting the most effective model.
- Few-shot learning, especially when implemented with "naive" or randomly selected examples, can unexpectedly degrade LLM performance for vulnerability detection, sometimes leading to a complete collapse in accuracy.
- LLMs currently struggle significantly with identifying specific vulnerability types (e.g., CVE IDs), excelling primarily at the simpler binary classification of "vulnerable" or "not vulnerable."
- Open-source contenders such as Jamba and Llama 2 can be competitive with, or even outperform, proprietary models like GPT-4, offering practical and cost-effective solutions for security teams.
About the Speaker(s)
Jie Lin is a researcher affiliated with the University of Central Florida. Their work focuses on the intersection of large language models and software security, specifically exploring the capabilities and limitations of LLMs in detecting code vulnerabilities.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Competent empirical benchmarking of LLMs against vulnerability detection tasks with some genuinely useful counterintuitive findings — bigger isn't better, few-shot can collapse to zero, open-source can beat GPT-4. The methodology is clean and the research questions are well-scoped, but the dataset is tiny (280 Java files, 200 C/C++ files), the few-shot implementation is admittedly naive, and the conclusions mostly confirm what the security ML community already suspected. Fills a slot, won't define the conversation.
Heather Calloway (CISO) — WEAK
Methodologically careful academic work that benchmarks LLM configurations against vulnerability detection tasks, but it never crosses the threshold from benchmark to operational guidance. The findings are real; the institutional implications are absent.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025