Vulnerability Root Cause Mapping with CWE
CVE/FIRST VulnCon 2025 · Main Stage
Overview
This talk, presented at VulnCon, delves into the critical importance and evolving landscape of vulnerability root cause mapping using the Common Weakness Enumeration (CWE). Speakers Alec Summers, the MITRE CVE and CWE Project Lead, and Chris Madden, a key contributor to the Root Cause Mapping Working Group, highlight how identifying the fundamental causes of vulnerabilities is crucial for effective product security, both within individual organizations and across the broader cybersecurity community. They argue that robust root cause analysis enables better trend analysis, informs strategic investments in security practices, and provides invaluable feedback loops into the Software Development Lifecycle (SDLC).

Key moments
- 0:00 Introduction & Key Takeaways: Importance of Root Cause Mapping
- 2:00 Defining Root Cause Mapping & its Value Proposition
- 4:00 Challenges of Root Cause Mapping & CWE Complexity
- 5:00 Community Collaboration & Decentralized Root Cause Mapping Adoption
- 6:00 Introducing AI-driven Capabilities for CWE Interaction
- 7:00 LLM Technologies Enable Interactive CWE Information Access
- 7:20 Chris Madden Takes Over: Personal Motivation for Using CWEs
Vulnerability Root Cause Mapping with CWE
Speakers: Alec Summers, MITRE CVE and CWE Project Lead; Chris Madden
Conference: VulnCon
YouTube: https://www.youtube.com/watch?v=TH1tGO15K24
Overview
This talk, presented at VulnCon, delves into the critical importance and evolving landscape of vulnerability root cause mapping using the Common Weakness Enumeration (CWE). Speakers Alec Summers, the MITRE CVE and CWE Project Lead, and Chris Madden, a key contributor to the Root Cause Mapping Working Group, highlight how identifying the fundamental causes of vulnerabilities is crucial for effective product security, both within individual organizations and across the broader cybersecurity community. They argue that robust root cause analysis enables better trend analysis, informs strategic investments in security practices, and provides invaluable feedback loops into the Software Development Lifecycle (SDLC).
The presentation acknowledges the historical challenges associated with root cause mapping, citing its technical difficulty, time-consuming nature, and the dense complexity of the CWE repository itself. However, a significant portion of the talk focuses on recent advancements, particularly the transformative role of Large Language Models (LLMs) and community collaboration. The speakers demonstrate how LLMs, when properly grounded with relevant data, are now sufficiently capable and cost-effective to automate and enhance the process of correlating CVE records with their underlying CWE identifiers, offering a path to more precise and actionable vulnerability intelligence.
The core message revolves around the idea that while root cause mapping has historically been a demanding task, new AI-driven capabilities, combined with ongoing community efforts, are making it more accessible and impactful. The talk not only showcases the technical feasibility and cost-efficiency of using LLMs for this purpose but also emphasizes the paramount importance of high-quality, complete, accurate, and timely data input for these systems to deliver meaningful results. This shift promises to empower security professionals to move beyond mere vulnerability reaction to proactive weakness elimination.
Background
▶ Watch: Introduction & Key Takeaways: Importance of Root Cause Mapping (0:00)
Historically, the cybersecurity community's focus has largely been on the "what" of vulnerabilities – the attack language, exploitation methods, and immediate impact – rather than the "why" – the fundamental mistake that caused it. This emphasis meant that root cause mapping, the process of identifying the underlying cause of a vulnerability and correlating a CVE record with its corresponding CWE identifier(s), was often overlooked or deemed a secondary concern. The Common Weakness Enumeration (CWE) repository, established around 2006-2007, exists as a vast, technically detailed, and densely presented body of knowledge, making it challenging for the average person to consume and apply effectively. This complexity led to a small group of "super users" who could navigate and accurately apply CWEs, while the majority struggled.
The lack of consistent and accurate root cause mapping presented several problems. Without understanding the root causes, organizations found it difficult to perform meaningful trend analysis, identify systemic weaknesses, or make informed investments to prevent vulnerabilities from manifesting in the first place. Remediation efforts often addressed symptoms rather than underlying issues, leading to recurring vulnerability patterns. Furthermore, the manual process of reviewing and mapping CVEs to CWEs, as exemplified by the MITRE team's work on the CWE Top 25 list, was described as "herculean" and "incomplete." For the 2023 Top 25, a full dataset of over 30,000 CVE records was available, but only about 7,500 could be manually reviewed due to resource constraints. This manual review often resulted in thousands of remappings, correcting instances where vulnerabilities were mapped to overly broad CWEs (like CWE-20: Improper Input Validation or CWE-284: Improper Authorization), insufficient information, or entirely incorrect weaknesses. This process, while improving data quality for the Top 25, was inherently limited by its resource intensiveness, reliance on third-party analysts (who are less informed than those with product knowledge), and its inability to scale across the entire vulnerability ecosystem.
Over the past year, however, significant headway has been made through community collaboration. There's been improved guidance on the CWE site, the addition of mapping labels and notes to individual CWE entries for better understanding, and demonstrable evidence of a decentralized approach to root cause mapping. Many CVE Numbering Authorities (CNAs) are now taking on the challenge of root cause mapping and reporting CWEs as part of their routine vulnerability disclosure practices. The Root Cause Mapping Working Group has been instrumental in driving these improvements and exploring new avenues, particularly the integration of AI-driven capabilities to make the CWE corpus more interactive and consumable. This shift acknowledges a growing demand for root cause information, moving beyond just severity and relevance to understanding the fundamental "why" behind vulnerabilities.
Key Findings
▶ Watch: Challenges of Root Cause Mapping & CWE Complexity (4:00)
The presentation highlights several pivotal findings that are reshaping the landscape of vulnerability root cause mapping:
- LLMs are Viable and Cost-Effective for Root Cause Mapping: A central finding is that Large Language Models, particularly those that are grounded and appropriately trained, are now "good enough and cheap enough" to be effectively used for root cause mapping and CVE enrichment. The bulk assignment proof of concept demonstrated that processing approximately 7,000 CVEs cost a mere $15 and took 27 hours unattended, showcasing remarkable efficiency compared to manual methods. This low cost and sufficient quality make LLMs a practical tool for scaling root cause analysis across vast datasets.
- Data Input Quality is Paramount: The speakers repeatedly emphasize that "the data input really matters." LLMs, like any analytical system, suffer from the "garbage in, garbage out" problem. For LLMs to provide accurate and useful CWE assignments, the underlying CVE information—including descriptions, key phrases, and especially reference content from advisories or patches—must be complete, accurate, and timely. Enriching CVE data at the point of entry is identified as the most critical step, even more so than the LLM processing itself.
- Community Collaboration Drives Progress: Significant improvements in root cause mapping guidance, usability of the CWE program, and the development of AI-driven capabilities have been direct results of community collaboration. The Root Cause Mapping Working Group, in particular, serves as a nexus for these advancements, fostering shared understanding and collective development. The vision for future progress, including the development of a benchmark dataset and an interactive CWE chatbot, continues to rely heavily on broad community engagement.
- LLMs Excel at Specificity and Automation: Unlike humans who might struggle with the hierarchical nature of CWEs or default to broader categories, LLMs demonstrate an ability to identify the most specific CWE identifier relevant to a vulnerability. This capability, coupled with their ability to automate tasks like CVE description creation, key phrase extraction, and reference content summarization, significantly reduces the "toil" associated with manual mapping and leads to more precise classifications.
- Grounding is Essential for Accuracy: The distinction between a generic LLM and a grounded LLM is crucial. While ungrounded LLMs prioritize coherence and can "hallucinate" answers, a grounded LLM is provided with the specific information it needs (e.g., the entire CWE corpus, mapping guidance) to answer questions accurately. This "open book exam" approach ensures that the LLM's outputs are based on authoritative data, minimizing inaccuracies. When evidence is insufficient, a grounded LLM will correctly state "not enough evidence" rather than fabricating information.
Technical Deep Dive
▶ Watch: Community Collaboration & Decentralized Root Cause Mapping Adoption (5:00)
The technical approach outlined in the talk revolves around leveraging grounded Large Language Models (LLMs) to augment and automate the process of vulnerability root cause mapping with CWEs. The core philosophy is that LLMs, when provided with the right context and data, can move beyond simply generating coherent text to providing accurate and actionable insights.
The speakers draw a clear distinction between an ungrounded LLM, which prioritizes coherence and may "hallucinate" answers in the absence of knowledge, and a grounded LLM. A grounded LLM is analogous to an "open book exam" scenario: it is explicitly provided with the relevant information (the "books" or corpus) needed to answer a query. In this context, the LLM is given access to the entire CWE corpus, including its detailed descriptions, relationships, and mapping guidance. This grounding ensures that the LLM's responses are accurate and based on authoritative knowledge. The "CWE expert for free in less than one minute" chatbot, which only answers based on the CWE corpus, is a direct application of this principle.
The overall process for CWE assignment involves several data layers and capabilities:
- CVE Information Enrichment:
- CVE Description Creation: LLMs can take a reference link (e.g., an advisory) and a CVE template to generate a comprehensive CVE description, significantly reducing manual toil. An example showed an LLM-generated description being nearly identical to a human-generated one for a mail spoofing vulnerability.
- Key Phrase Extraction: LLMs can extract critical information from CVE descriptions, such as the impact of the vulnerability and, most importantly, the root cause and weakness. This is crucial for accurate CWE assignment, as directly mapping a full description can be less effective than focusing on the core weakness. Analysis of 260,000 CVEs showed "product" as the top key phrase, while "root cause" and "weakness" were least used, highlighting an area for quality improvement at the point of entry.
- Reference Content Summarization: Many CVEs link to external advisories, patch notes, or detailed reports (e.g., Log4Shell had over 100 links). LLMs can summarize this often-voluminous reference content, providing a concise overview for both human analysts and subsequent LLM processing. This summary is often more accurate and informative than the initial CVE description itself.
- Consensus CWEs: Approximately 15% of CVEs have similar descriptions within groups, and their associated CWEs often show consensus. This consensus information can also be used as an input for the LLM.
- The Retriever Mechanism:
This is a crucially important step, acting as the "open book" part of the grounded LLM. The goal is to narrow down the vast CWE corpus (approximately 1,000 active CWEs, with about 500 commonly used) to a manageable "short list" of around 10 highly relevant candidates for the LLM or human to consider. Three types of retrievers are employed:
- Sparse Retriever: This functions like a traditional keyword search (e.g., BM25), finding lexically similar CWEs based on term matching.
- Dense Retriever: This uses semantic search to find conceptually similar CWEs, even if the exact words don't match (e.g., "memory corruption" and "buffer overflow" are semantically similar but lexically different).
- Graph Property Retriever: The CWEs form a graph with relationships like "child of" and "parent of." This retriever explores the graph structure around identified CWEs to find related weaknesses, ensuring that hierarchical relationships are considered.
- LLM Assignment and Rationale:
After the retrievers provide a short list of candidates and enriched CVE information, the LLM performs the final assignment. The initial design included an "analyzer," "critic," and "resolver" LLM, but it was found that the "analyzer" alone was sufficient, producing high-quality outputs without the added cost and complexity of the critic and resolver.
- Prompt Engineering: The LLM is given a carefully constructed prompt that includes the CWE guidance and insights from the Root Cause Mapping Working Group. This ensures the LLM applies the same principles and considerations as human experts.
- Output Format: The LLM's output is designed to be highly informative, including:
- The assigned CWE ID and Name.
- A confidence score for the assignment.
- The abstraction level of the CWE (e.g., Base, Class, Variant). The goal is to achieve lower-level, more specific assignments where possible.
- The mapping label (e.g., Primary, Secondary, Allowed with review, Disallowed).
- An overall confidence rating and evidence rating.
- A visual representation of the CWE relationships (child of, parent of) within the vulnerability chain.
- A summary and rationale explaining how the LLM arrived at its analysis. This transparency is crucial for human review and trust.
When insufficient evidence is provided (e.g., a vague description like "or exec service is running"), the LLM is trained to respond with "not enough evidence" rather than fabricating an answer, demonstrating its grounded nature. The system also actively tries to avoid assigning discouraged CWEs, such as the very broad CWE-20: Improper Input Validation.
Demo / Proof of Concept
▶ Watch: LLM Technologies Enable Interactive CWE Information Access (7:00)
The talk presented two primary demonstrations of LLM capabilities for root cause mapping: an interactive chatbot and a large-scale bulk assignment proof of concept.
The first demonstration involved a grounded chatbot designed to act as a CWE expert. Chris Madden explained that such a chatbot could be built "for free in less than one minute" from a browser, following provided instructions. The key characteristic of this chatbot is its grounded nature: if asked "what is a dog," it would respond "I don't know" because the information is not within its specific corpus (the CWE knowledge base). This ensures that the chatbot's answers are derived solely from the authoritative CWE data. Users can interact with it by asking questions like, "What's the best CWE for this vulnerability description?" or "Tell me about the CWEs related to buffer overflow," effectively providing an expert-level interface to the complex CWE repository. This capability aims to make the CWE corpus more accessible and interactive than traditional browsing or searching.
The more extensive proof of concept involved a bulk assignment of CWEs to a large dataset of CVEs, specifically the approximately 7,000 CVEs used for the 2023 CWE Top 25 mapping analysis. This demonstration aimed to answer critical questions about cost and quality when applying LLMs at scale.
Cost Analysis:
- Processing approximately 7,000 CVEs for detailed reports (including summaries, confidence scores, and rationales) cost only $15.
- The embedding process (converting text to numbers for LLM comparison) cost an additional $2.
- The total unattended processing time was 27 hours, which was not optimized for parallel execution, indicating potential for even faster processing.
- Reference content summarization, which can involve gigabytes of data, was the most significant cost driver but was performed using an experimental, free LLM at the time. Other tasks like key phrase extraction were relatively cheap and fast.
Quality and Performance Metrics:
The performance was evaluated against the 2023 CWE Top 25 dataset, acknowledging it wasn't designed as a benchmark but served as the best available comparison.
- Retriever Performance: The goal of the retrievers is to narrow down 1,000 CWE candidates to a short list of about 10.
- The sparse retriever (keyword search) achieved a recall of approximately 97% of relevant CWEs.
- The dense retriever (semantic search) found unique CWEs not caught by the sparse retriever.
- The graph property retriever (using CWE relationships) found additional unique CWEs.
- Some CWEs were missed (false negatives), primarily due to insufficient information in the CVE description or non-crawlable reference links.
- LLM Assignment Output: The LLM's output for each CVE included:
- The assigned CWE ID and name.
- A confidence score for the assignment.
- The CWE's abstraction level (e.g., Base, Class).
- The mapping label (Primary, Secondary, Allowed with review).
- An overall confidence and evidence rating.
- A generated visual representation of the CWE hierarchy and relationships.
- A detailed description of the vulnerability chain, summary, and the LLM's rationale.
- Comparison to Expert Mappings:
- In one example, where experts assigned CWE-790 (Improper Filtering of Special Characters), the LLM assigned CWE-791 (Improper Neutralization of Special Elements in Output Used by a Third-Party Integrator), which is a child of CWE-790. This demonstrated the LLM's ability to go to the most specific CWE, a task often challenging for humans.
- In another case with limited input (no root cause/weakness, only a vulnerability description and reference link), the LLM assigned CWE-942, while the Top 25 assigned CWE-923. Notably, CISA's separate mapping for this vulnerability also aligned with the model's CWE-942.
- Graph Distance Metrics: Since exact matches can be difficult to achieve and assess, the talk introduced "graph distance" as a metric. This measures how "close" the LLM's assigned CWE was to the expert's on the CWE graph (e.g., parent, child, cousin). The results showed an 81% exact match rate, which increased to 90% when considering CWEs within a graph distance of one or two. This indicates that even when not an exact match, the LLM's assignments were typically within a relevant area of the CWE hierarchy.
- Abstraction Level Matching: The LLM showed strong agreement with expert mappings at the important Base and Class abstraction levels, indicating it correctly understood the general category and specific nature of weaknesses.
- Balanced Accuracy: While raw accuracy can be misleading for multi-class classification with many potential outputs, the balanced accuracy was reported at 90%, whether looking at 411 unique CWEs in the dataset or rolling them up to the 103 CWEs commonly used by NVD.
The proof of concept confirmed that LLMs can deliver high-quality, cost-effective, and scalable root cause mapping, particularly when provided with rich input data and operating within a grounded framework.
Defensive Implications
▶ Watch: Chris Madden Takes Over: Personal Motivation for Using CWEs (7:20)
The advancements in LLM-driven root cause mapping offer significant defensive implications for organizations and the broader cybersecurity ecosystem:
- Enhanced Trend Analysis and Visibility: By accurately mapping CVEs to their underlying CWEs, organizations gain deeper insights into the patterns of vulnerabilities over time. This enables more robust trend analysis, revealing recurring weaknesses within products, teams, or development processes. Such visibility allows defenders to understand not just what vulnerabilities exist, but why they exist, facilitating proactive rather than reactive security measures.
- Informed Investment and Policy Decisions: Understanding root causes illuminates where security investments, policy changes, and practice improvements can have the greatest impact. If a particular CWE (e.g., CWE-79: Improper Neutralization of Input During Web Page Generation ('Cross-site Scripting')) is consistently identified as a root cause, resources can be directed towards developer training, secure coding guidelines, or automated tools specifically targeting XSS prevention, rather than just patching individual instances. This shifts security spending from reactive fixes to preventative measures.
- Improved SDLC Feedback Loops: Effective root cause mapping provides a valuable feedback loop into the Software Development Lifecycle (SDLC) and architectural design processes. By identifying the root cause early, organizations can incorporate lessons learned into design, coding, testing, and deployment phases. This prevents the same mistakes from being repeated and reduces the significantly higher cost of remediating issues later in the development cycle. The goal is to "eliminate or avoid" weaknesses so they "aren't manifest in the first place."
- Better Exploitability Insight: Knowledge of specific weakness types can provide further insight into potential exploitability. Certain CWEs are known to be targeted by particular adversaries or exploitation techniques. This information can help defenders prioritize remediation efforts based on the likelihood and impact of exploitation.
- Democratization of CWE Expertise ("Easy Button"): The complexity of the CWE repository has historically limited its widespread adoption. LLM-driven tools, such as the proposed interactive chatbot or integration with vulnerability submission platforms like Vulnerogram, act as an "easy button." They can guide users (even non-experts) to the most appropriate CWEs without requiring deep knowledge of the entire corpus. This makes accurate root cause mapping more accessible, encouraging broader adoption and consistency across the industry.
- Automation of Toil and Data Enrichment: LLMs can automate the tedious and time-consuming tasks of CVE description creation, key phrase extraction (identifying root causes and weaknesses), and summarizing voluminous reference content. This frees up security analysts to focus on higher-value tasks, while simultaneously improving the quality and completeness of vulnerability data at the point of entry. The concept of a central CVE repository for extracted key phrases and summarized content is proposed to further standardize and facilitate this enrichment.
- Foundation for Proactive Security: Ultimately, the comprehensive and accurate CVE data, enriched with precise CWE root cause mappings, forms the foundation for a truly proactive security posture. It moves organizations beyond simply responding to vulnerabilities to understanding, predicting, and preventing them, thereby significantly reducing overall cyber security risk.
Key Takeaways
- LLMs are a Game-Changer for Root Cause Mapping: Grounded LLMs are now sufficiently capable and cost-effective (e.g., $15 for 7,000 CVEs) to automate and enhance vulnerability root cause mapping with CWEs, making this historically challenging task more accessible and scalable.
- Quality Data Input is Paramount: The success of LLM-driven mapping hinges on complete, accurate, and timely CVE data, including detailed descriptions and accessible reference content. Enriching this data at the point of entry is the most critical factor.
- LLMs Excel at Specificity: Unlike humans who might gravitate towards broader categories, LLMs can consistently identify the most specific CWE identifier, leading to more precise and actionable vulnerability classifications.
- Community Collaboration Drives Innovation: Ongoing community engagement through groups like the Root Cause Mapping Working Group is essential for developing improved guidance, usability, and advanced AI-driven tools for the CWE program.
- "Easy Button" for Defenders: LLM tools provide an "easy button" for security practitioners, simplifying the complex process of CWE assignment and enabling better trend analysis and more informed security investments within the SDLC.
- Need for a Benchmark Dataset: The community needs to establish a standardized benchmark dataset for CWE mapping to foster better tool development, comparison, and open-source contributions from academics and industry.
About the Speaker(s)
Alec Summers is the MITRE CVE and CWE Project Lead. In this role, he is at the forefront of driving advancements in vulnerability identification and classification standards. His work focuses on improving the utility and accessibility of the Common Vulnerability and Exposures (CVE) and Common Weakness Enumeration (CWE) programs, fostering community collaboration, and exploring innovative approaches like AI-driven capabilities to enhance cybersecurity practices. He is a key advocate for effective root cause mapping as a foundation for better vulnerability management and risk reduction.
Chris Madden is a dedicated contributor to the Root Cause Mapping Working Group, where he applies his expertise in LLMs and vulnerability analysis to improve CWE assignment. Motivated by a desire to use CWEs for identifying and eradicating vulnerability classes, Chris initiated his work by submitting GitHub issues to correct CWE assignments, leading him to join the working group. He is instrumental in developing and demonstrating the practical application of grounded LLMs for automating CVE enrichment and CWE mapping, focusing on cost-efficiency, accuracy, and providing clear rationales for LLM outputs. While his specific title or company were not detailed in the provided transcript, his contributions are clearly deeply technical and impactful to the community's progress in this area.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
A competent, well-structured presentation on a real problem — CWE mapping coverage is genuinely terrible and the data quality issue is chronic. Summers and Madden are the right people to be giving this talk, and the LLM grounding approach is sensible rather than hype-driven. The $15/7,000-CVE proof of concept and the graph-distance evaluation metric are concrete contributions worth knowing about. But this is infrastructure tooling work dressed up as research, and the audience for whom this is genuinely new and actionable is narrow. The 'LLMs are good at classification tasks when grounded' insight is not 2024-novel. The talk earns its slot at VulnCon — which is exactly the right venue for…
Heather Calloway (CISO) — SOLID
A technically credible and timely talk from authoritative sources on applying grounded LLMs to CWE root cause mapping. The proof-of-concept results are real and the cost-efficiency finding is genuinely interesting infrastructure news for the vulnerability management community. But this is a tool-building talk aimed at CNAs, vulnerability researchers, and NVD contributors — not security leaders or defenders who need to change how they operate. It stops well short of answering the institutional question that matters most: who owns root cause quality in the CVE ecosystem, and what happens when they don't.