Hard Problems in CWE, and What it Tells us about Hard Problems in the Industry

Steve Christie Kohley (CWE Technical Lead · MITRE Corporation)

CVE/FIRST VulnCon 2025 · Main Stage

Overview

In this insightful talk, Steve Christey Coley, the CWE technical lead and co-founder from the MITRE Corporation, delved into the persistent challenges faced by the Common Weakness Enumeration (CWE) project. Established in 2005, CWE has become a cornerstone in the cybersecurity landscape, providing a standardized list of software and hardware weakness types. Coley's presentation explored the historical context of CWE, its evolving organization, the complexities of weakness mapping, and the ongoing efforts to modernize its coverage to address contemporary security issues.

Watch on YouTube

Visual summary for Hard Problems in CWE, and What it Tells us about Hard Problems in the Industry by Steve Christie Kohley
Visual summary for Hard Problems in CWE, and What it Tells us about Hard Problems in the Industry by Steve Christie Kohley

Key moments

  1. 0:00 Speaker introduction and talk agenda overview
  2. 1:55 CWE's evolving audience: from technical experts to vulnerability management
  3. 2:40 Shift in CWE content development and abstraction levels
  4. 4:15 Broadening CWE's applicability to hardware, ICS, and AI
  5. 5:00 Imposing higher quality expectations for CWE content
  6. 6:00 Shifting influences and addressing new types of weaknesses
  7. 7:00 Key criteria for a robust CWE classification system

Hard Problems in CWE, and What it Tells us about Hard Problems in the Industry

Speakers: Steve Christey Coley, CWE Technical Lead, MITRE Corporation

Conference: VulnCon

YouTube: https://www.youtube.com/watch?v=RcR-EFSptnQ

Overview

In this insightful talk, Steve Christey Coley, the CWE technical lead and co-founder from the MITRE Corporation, delved into the persistent challenges faced by the Common Weakness Enumeration (CWE) project. Established in 2005, CWE has become a cornerstone in the cybersecurity landscape, providing a standardized list of software and hardware weakness types. Coley's presentation explored the historical context of CWE, its evolving organization, the complexities of weakness mapping, and the ongoing efforts to modernize its coverage to address contemporary security issues.

The core message of the talk resonated beyond the specifics of CWE, highlighting that many of the "hard problems" within the enumeration project are symptomatic of broader, systemic difficulties within the cybersecurity industry itself. These include the struggle with consistent terminology, the varying levels of technical expertise across different user groups, the challenges of accurate root cause analysis, and the continuous need to adapt to new technologies and attack vectors. Coley emphasized MITRE’s commitment to balancing usability with technical excellence, recognizing the diverse audience that now relies on CWE for everything from code analysis to vulnerability management.

Understanding these challenges is crucial for anyone involved in software security, from developers and security architects to vulnerability analysts and researchers. By dissecting the intricacies of weakness classification and mapping, Coley provided a roadmap for how the community can contribute to improving CWE, thereby enhancing the collective ability to identify, categorize, and mitigate security flaws more effectively across the entire industry.

Background

▶ Watch: Speaker introduction and talk agenda overview (0:00)

CWE, now almost two decades old, has undergone significant evolution since its inception in late 2005. Initially, the primary audience and content development strategy were geared towards highly technical experts. Early efforts focused on supporting code analysis tool vendors, advanced bug hunters, consultants, and academic researchers. The approach to content creation was a "kitchen sink" method, gathering input from various sources, including security guides and unusual CVE (Common Vulnerabilities and Exposures) entries. At this stage, CWE operated with a single main view and assumed a technically proficient user base.

Around 2010, CWE began expanding its focus to include developers and product security team leads, necessitating a shift towards more comprehensive and understandable content. This period saw the introduction of multiple views and abstraction layers, such as pillars, classes, bases, and variants, along with the development of a vulnerability theory framework to systematically organize weaknesses. The most recent shift, occurring approximately four to five years ago, broadened the primary audience further to encompass vulnerability management staff and leads—individuals responsible for publishing CVEs daily. This expansion introduced a challenge: catering to novice CWE users who might not possess deep development experience.

Content development continued to evolve, broadening CWE's applicability to diverse technologies and domains. This included the release of CWE 4.0 in 2020 with hardware weaknesses, followed by expansions into Industrial Control Systems (ICS) and Artificial Intelligence (AI). Alongside this, community contributions became more significant and publicly trackable, leading to an increased emphasis on formalizing quality expectations for new CWE entries. Throughout these transitions, CWE's influences moved from raw vulnerability reports and research to a deeper understanding of user personas and their varied needs, fostering active working groups and systematic reclassification efforts to address an increasing variety of weaknesses beyond the traditional "Top 25" like hard-coded passwords or SQL injection.

<h2>Key Findings</h2>

Coley's talk illuminated several critical "hard problems" that not only challenge the CWE project but also reflect broader industry issues in understanding and categorizing software weaknesses.

Challenges in CWE Organization and Representation

Coley highlighted the foundational criteria for a robust classification system: mutual exclusivity (each weakness clearly distinct), unambiguous presentation, understandability and usability, completeness (sufficient breadth and depth), repeatability (different users arriving at the same classification), and extensibility. Achieving these simultaneously is arduous. While CWE IDs are a flat numeric space, the project incorporates significant structure through hierarchical views, primarily the Research View 1000. This view features pillars (highest abstraction), classes, bases, and variants (lowest abstraction), sometimes reaching six or more levels deep. CWEs are also characterized by dimensions or facets such as behavior, resource properties, or technology-specific considerations. Expanding all possible variations across these dimensions could hypothetically result in over 32,000 CWEs, a stark contrast to the current 900+ entries, raising questions about the practical utility of such extensive granularity.

Classification Gaps and Lack of Formalized Research

A significant finding is that vast swathes of CWE lack the formalized classification studies seen in well-understood areas like buffer overflows. Access control, for example, presents hundreds of models, yet CWE users often struggle with fundamental distinctions between authorization and authentication. This gap is exacerbated by a lack of established terminology and comprehensive examples, making it difficult for users to navigate and map obscure weaknesses. The potential solution involves creating intermediary entries to better organize classes and bases, particularly for unusual weaknesses, but this risks inventing new terms which could impact usability.

Persistent Mapping Problems

Even experienced users face difficulties in accurately and precisely mapping vulnerabilities to CWEs, especially for non-obvious issues. Coley outlined a realistic goal: consistent mapping among users with identical backgrounds, information, and time constraints. The more idealistic "unicorns and rainbows" goal aims for consistency even with differing backgrounds, information, and time. To aid this, CWE has introduced mapping usage recommendations for each entry, categorizing them as Allowed, Allowed with Review, Discouraged, or Prohibited (i.e., deprecated). Examples include CWE-200 (Information Leak) and CWE-20 (Improper Input Validation), which are often misused as high-level, catch-all mappings despite the existence of more specific children.

The "Weakness Mindset" vs. "Vulnerability Mindset"

A crucial distinction highlighted is that between the weakness mindset and the vulnerability mindset. Vulnerability thinking often focuses on attacks, mitigations, consequences, and risk assessment. In contrast, weakness thinking centers on specific product misbehaviors. Coley explained that there's no direct one-to-one relationship between vulnerability impacts (e.g., code execution, information leaks) or attack types and underlying weaknesses. The same attack, such as providing a large username string, could trigger various weaknesses like excessive memory allocation or recursion, not just a standard buffer overflow.

Ambiguous and Evolving Terminology

The industry suffers from vague and inconsistent terminology. Coley polled an audience where 30% interpreted "buffer overflow" differently than "writing past the end of a buffer." Other examples include Stack Overflow, which the address sanitizer tool often reports for excessive stack consumption rather than stack-based buffer overflows, and memory leak, which can mean either sensitive information exposure or unreleased memory. These terminological differences hinder consistent communication and accurate CWE mapping.

Lack of Root Cause Analysis Skills

Many practitioners lack the necessary skills for effective root cause analysis. This requires deep knowledge of programming languages, underlying protocols, frameworks, product domains, and threat models. Interpreting proof of concept exploits and uncovering unstated assumptions are also critical. Without a clear understanding of the weakness, its context, and the underlying domain, accurate mapping to the most appropriate CWE is impossible.

Weakness-Oriented Chain Analysis

Coley noted a growing trend towards weakness-oriented chain analysis, which separates primary root causes from resultant issues. For instance, an out-of-bounds write or use-after-free is often a secondary symptom of an earlier, more fundamental mistake. Greater attention to these chains is complicating traditional root cause identification.

Common Mapping Assumptions and Errors

Users frequently make simplistic assumptions or errors in mapping. A common mistake is choosing higher-level CWEs when more precise, lower-level options exist, such as using CWE-77 (Command Injection) instead of the more specific CWE-78 (OS Command Injection). Another prevalent misconception is that injection errors are solely due to poor input validation, when often the root cause lies in improper output encoding to separate data from control. Reliance on incomplete views, such as View 10003 (used by NVD) which only covers about 130 CWEs, further exacerbates mapping inaccuracies by limiting the available options.

New Technology Challenges and Content Quality

New technologies (e.g., cloud, AI/ML) often present familiar weaknesses but with domain-specific terminology, making them less recognizable within existing CWEs. There are also internal challenges with disconnects between CWE entries and their descriptions, and the need to reflect how words and meanings change over time (e.g., the shifting interpretation of "classic overflow"). Deprecating popular but problematic CWEs without disrupting users is a significant hurdle. Finally, balancing the timeliness of new content releases with rigorous quality expectations—ensuring accuracy and avoiding insecure recommendations—remains a core dilemma.

Technical Deep Dive

▶ Watch: Shift in CWE content development and abstraction levels (2:40)

The structure and evolution of CWE are central to understanding its "hard problems." CWE IDs are a flat numeric space, meaning the ID number itself holds no inherent structural relationship. However, CWE provides rich structural context through various hierarchical views. The primary structure is found in Research View 1000, which organizes weaknesses into abstraction layers: pillars (the highest, most general categories), classes, bases, and variants (the lowest, most specific types). This hierarchy can extend six or more levels deep in "bushy" parts of the tree.

Beyond this hierarchical organization, CWEs are also characterized by multiple dimensions or facets. These can include the type of malicious behavior, the properties of manipulated resources, or language-specific and technology-specific considerations at lower levels. Coley illustrated the combinatorial explosion possible here: assuming four layers and eight children per layer, a fully expanded system could yield over 32,000 distinct CWEs. This highlights a tension between exhaustive detail and practical usability, given the current count of just over 900 CWEs.

A key effort to improve usability and mapping accuracy is the introduction of mapping usage recommendations. These provide guidance on whether a specific CWE ID is appropriate for mapping. Categories include:

  • Allowed: Generally safe for direct mapping.
  • Allowed with Review: Can be used, but with caveats, as it might be misused or require careful consideration of specific points.
  • Discouraged: High-level entries that should only be used in emergency situations when no lower-level, more specific CWE is available (e.g., for new or unusual weaknesses).
  • Prohibited (deprecated): Entries that should not be used for mapping at all, often due to being overly broad or superseded by better options.

Coley provided examples of frequently misused CWEs:

  • CWE-200 (Information Leak): Often treated as a generic "information leak" equivalent, despite having numerous narrower children that would be more precise.
  • CWE-20 (Improper Input Validation): Dubbed "people's favorite CWE to hate on," it's frequently used as a default even though more specific children (e.g., for quantity validation or syntax validation) have been available for years. This highlights a common issue where time constraints or lack of deep analysis lead to selecting broad, high-level mappings.

CWE's comprehensive coverage is further organized into specialized views. The Comprehensive Categories view, introduced recently, offers a two-layer structure with 22 mutually exclusive categories at the top. For instance, memory safety is a category in this view, collecting various behaviors that are otherwise scattered across the Research View 1000. Other views serve specific purposes:

  • View 10003: Traditionally used by NVD and others, it contains a limited subset of CWEs (around 130).
  • View 1194: A hardware-specific view.
  • View 699: A more developer-friendly view, covering approximately 400 CWE entries.

Reliance on these incomplete views can significantly impact mapping accuracy and precision, as users might not be exposed to the full range of available, more specific CWEs.

The talk also touched upon the vulnerability theory framework, a conceptual model developed to provide a systematic understanding of how weaknesses manifest in vulnerabilities. This framework, while theoretically sound, can be challenging to teach without concrete examples, a common problem across many technical disciplines. Looking ahead to CWE 5.0, a key goal is to formalize and specify the various dimensions that characterize weaknesses, which will be crucial for managing new proposals and preventing overlap in the ever-growing enumeration.

Demo / Proof of Concept

▶ Watch: Imposing higher quality expectations for CWE content (5:00)

While the talk did not feature a live demonstration or a proof of concept exploit, Steve Christey Coley used a specific example from CWE's historical content to illustrate the challenges in maintaining high-quality, understandable demonstrative examples. He presented an original example from CWE-95 (Eval Injection), written in Perl and utilizing CGI technology, which dates back to before CWE 1.0.

Coley critiqued this example to highlight several issues:

  • Realism vs. Understandability: While the code was "kind of realistic," its complexity and the presence of "excess code" made it potentially too extensive as a first example for users.
  • Language and Technology Obsolescence: Written in Perl and relying on CGI, technologies that are less common or well-understood by modern developers, the example's relevance and comprehensibility have diminished over time.
  • Evolving Quality Expectations: Despite being a "pretty good" example for its time, it wouldn't fully satisfy current quality expectations for clarity, conciseness, and broad applicability.

This discussion underscored the difficulty in creating and maintaining illustrative examples that remain relevant, accurate, and easily digestible across diverse audiences and rapidly changing technological landscapes, a challenge that directly impacts the usability and educational value of CWE entries.

Defensive Implications

▶ Watch: Key criteria for a robust CWE classification system (7:00)

The insights shared by Steve Christey Coley offer several critical implications for defenders across various roles, from vulnerability analysts to developers and security leaders.

  1. Prioritize Precise Mapping for Vulnerability Management: For CNAs and vulnerability management staff, the call to move beyond high-level, discouraged CWEs (like CWE-200 or CWE-20) is paramount. While time constraints are real, striving for the most accurate and precise mappings, leveraging mapping usage recommendations, directly improves the quality of vulnerability data. This enables more effective trend analysis, better tooling integration decisions, and more targeted mitigation strategies. When a lower-level CWE isn't immediately apparent, it should prompt deeper investigation or, failing that, serve as feedback to the CWE project about potential gaps.
  1. Cultivate a "Weakness Mindset": Defenders, especially those involved in incident response or root cause analysis, must distinguish between the vulnerability mindset (focused on attacks, impacts, mitigations) and the weakness mindset (focused on the specific product misbehavior). Understanding that a single attack or impact can stem from dozens of different weaknesses is crucial for effective remediation. This shift in perspective helps move beyond superficial fixes to address the underlying design or implementation flaws.
  1. Invest in Root Cause Analysis Skills: The talk underscored that effective root cause analysis requires a specific set of skills: deep knowledge of programming languages, protocols, frameworks, product domains, and threat models. Organizations should invest in training for their security teams and developers to enhance these capabilities. Interpreting proof of concept exploits and discerning unstated assumptions in vulnerability reports are vital for accurately identifying the true CWE.
  1. Understand and Address Terminology Ambiguity: Security teams need to be acutely aware of the pervasive terminology ambiguity within the industry. Phrases like "buffer overflow," "Stack Overflow," or "memory leak" can have vastly different interpretations. Establishing clear, internal definitions and encouraging precise language, especially in vulnerability reports and internal documentation, can prevent miscommunication and mischaracterization of weaknesses.
  1. Advocate for Developer-Centric Mapping: Coley emphasized that developers are often best positioned to perform the most accurate and precise CWE mappings due to their access to source code and understanding of product context. Security leaders should advocate for integrating CWE mapping into the development lifecycle, providing developers with the necessary tools, guidance, and time to perform this critical task, rather than offloading it entirely to vulnerability analysts.
  1. Recognize and Address Weakness Chains: The increasing focus on weakness-oriented chain analysis means defenders should look beyond immediate symptoms (e.g., out-of-bounds write, use-after-free) to identify the earlier, root cause mistakes. This more holistic view of vulnerability chains can lead to more robust and preventative security measures.
  1. Actively Engage with the CWE Community: The CWE project thrives on community involvement. Defenders are encouraged to provide feedback, submit new content proposals (especially for low-level variants or in emerging technology domains like AI/ML), and participate in working groups. This direct engagement helps fill classification gaps, improve content quality, and ensure CWE remains relevant and comprehensive for all users. Organizations should also consider what level of quality expectations they hold for CWE content, contributing to the public discussion on balancing timeliness and perfection.

Key Takeaways

  • CWE has significantly evolved over two decades, expanding its audience from highly technical experts to include developers and vulnerability management staff, necessitating a continuous balance between technical excellence and usability.
  • Accurate and precise weakness classification and mapping are challenged by persistent issues such as classification gaps, pervasive terminology ambiguity (e.g., for "buffer overflow" or "memory leak"), and a lack of formalized root cause analysis skills within the industry.
  • The distinction between the weakness mindset (focus on product misbehavior) and the vulnerability mindset (focus on attacks, impacts, mitigations) is crucial for effective security analysis and remediation, as there's no simple one-to-one mapping between them.
  • CWE's complex structure, with its hierarchical views (pillars, classes, bases, variants) and dimensions, aims for comprehensiveness but faces difficulties in preventing overlap and effectively integrating new weaknesses from diverse technologies like hardware, ICS, and AI.
  • Users are encouraged to utilize mapping usage recommendations and strive for lower-level, more specific CWEs, avoiding discouraged high-level entries like CWE-200 and CWE-20, to improve the quality of vulnerability data and trend analysis.
  • The future of CWE relies heavily on active community involvement, including providing feedback, submitting new content proposals (especially to fill gaps at low-level variants), and contributing to the ongoing discussion on balancing content timeliness with quality and usability.

About the Speaker(s)

Steve Christey Coley is a highly respected figure in the cybersecurity community, serving as the CWE technical lead at the MITRE Corporation. He is also a co-founder of the CWE project, having been involved since its inception in 2005. His extensive experience includes a tenure as the CVE editor, underscoring his deep expertise in vulnerability and weakness classification. Coley's work focuses on guiding the development of CWE content, ensuring it remains a vital resource for understanding, categorizing, and mitigating software and hardware security weaknesses.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Steve Christey Coley is one of the few people on the planet who can speak with genuine authority on why weakness classification is hard, and he delivers. This isn't a research talk dropping novel exploits — it's a meta-level infrastructure talk about the epistemological foundations of how the industry categorizes what's wrong with software. At VulnCon specifically, that's exactly the right lane. The content is substantive, self-critical in ways that most standards-body talks never are, and packed with practical signal for CNAs, vulnerability analysts, and anyone who has ever rage-quit trying to pick the right CWE. The combinatorial explosion argument alone — four layers, eight children…

Heather Calloway (CISO) — SOLID

Steve Christey Coley is the right person to give this talk, and the problems he surfaces are real — terminology decay, mapping inconsistency, root cause analysis deficits, and the structural tension between comprehensiveness and usability in CWE are genuine industry problems. But the talk stays inside its own house. It diagnoses the failures of a classification system without connecting those failures to the institutional consequences that should make a security leader care: what bad CWE mapping costs you in a regulatory examination, how weak root cause taxonomy distorts board risk reporting, or why vulnerability management programs built on CWE-200 and CWE-20 as catch-alls are producing…

→ Top-rated talks at CVE/FIRST VulnCon 2025

All talks from CVE/FIRST VulnCon 2025