SoK: Automated TTP Extraction from CTI Reports – Are We There Yet?

Marvin Büchel (University of Alenborg)

34th USENIX Security Symposium (USENIX Security '25) · Day 2 · ML and AI Security 2

Overview

In the rapidly evolving landscape of cyber security, the ability to rapidly understand and respond to new threats is paramount. Cyber Threat Intelligence (CTI) reports, meticulously crafted by security experts post-attack, serve as vital repositories of information detailing attacker targets, specific techniques, impact, and motivations. The dream of leveraging this intelligence for proactive defense, automated attacker profiling, and the discovery of overarching trends within the hacker community hinges on the ability to automatically extract and structure critical data, known as Tactics, Techniques, and Procedures (TTPs). However, a significant challenge persists: these reports are predominantly written in natural language, which is inherently ill-suited for automated comparison and aggregation.

Watch on YouTube · Slides

Visual summary for SoK: Automated TTP Extraction from CTI Reports – Are We There Yet? by Marvin Büchel
Visual summary for SoK: Automated TTP Extraction from CTI Reports – Are We There Yet? by Marvin Büchel

Key moments

  1. 0:00 Introduction: The challenge of automated TTP extraction
  2. 2:00 Mitre ATT&CK Framework for structured threat intelligence
  3. 3:00 Methodology for comprehensive literature review (2015-present)
  4. 4:00 Three main TTP extraction approaches: NER, Classification, LLMs
  5. 7:00 Key problem: Incomparability of existing TTP solutions
  6. 8:00 Our solution: Unified empirical study and re-implementation
  7. 9:00 Standardized datasets and evaluation scenarios for comparison

SoK: Automated TTP Extraction from CTI Reports – Are We There Yet?

Speakers: Marvin Büchel

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=HviF5WCWBdc

Overview

In the rapidly evolving landscape of cyber security, the ability to rapidly understand and respond to new threats is paramount. Cyber Threat Intelligence (CTI) reports, meticulously crafted by security experts post-attack, serve as vital repositories of information detailing attacker targets, specific techniques, impact, and motivations. The dream of leveraging this intelligence for proactive defense, automated attacker profiling, and the discovery of overarching trends within the hacker community hinges on the ability to automatically extract and structure critical data, known as Tactics, Techniques, and Procedures (TTPs). However, a significant challenge persists: these reports are predominantly written in natural language, which is inherently ill-suited for automated comparison and aggregation.

This talk, presented by Marvin Büchel from the University of Alenborg, addresses this decade-long challenge by conducting a comprehensive Systematization of Knowledge (SoK) on the current state of automated TTP extraction. Collaborating with researchers from Polytenico de Milano, the University of Tentum, Accents, Siemens, and NEC Laboratories Europe, Büchel investigates whether the community has achieved reliable automated TTP extraction. The research critically evaluates existing approaches against standardized metrics and real-world scenarios, ultimately revealing that while progress has been made, the field faces fundamental bottlenecks that prevent practical, fully automated deployment.

The core of this research revolves around the MITRE ATT&CK framework, a globally recognized and comprehensive library that categorizes adversary behaviors into TTPs. Each technique within ATT&CK possesses a unique, machine-readable ID (e.g., T1112), making it an ideal target for structured extraction. The talk aims to answer a crucial question: "Are we there yet?" – meaning, is automated TTP extraction reliable enough for practical, real-world application? The findings presented suggest a nuanced but ultimately negative answer, highlighting critical areas for future research and development beyond merely building more complex natural language processing (NLP) models.

Background

▶ Watch: Introduction: The challenge of automated TTP extraction (0:00)

The increasing sophistication of cyber attacks necessitates equally sophisticated defense mechanisms. CTI reports are a cornerstone of modern cybersecurity, providing invaluable insights gleaned from post-attack analyses. These reports document crucial elements such as the targets of an attack, the tactics (high-level goals like "initial access"), techniques (how those goals are achieved, e.g., "phishing"), and procedures (the specific tools or methods used for a technique, e.g., "spearphishing attachment with a malicious macro"). While tactics and procedures are important, the talk specifically focuses on techniques, as their extraction is generally considered the most challenging due to their often implied nature and the critical role of semantic understanding in natural language.

The fundamental hurdle in leveraging CTI for automated defense is that these reports are written in unstructured natural language. To overcome this, the cybersecurity community has largely converged on the MITRE ATT&CK framework as a standardized ontology for describing adversary actions. Released in 2015, ATT&CK provides a common language and a structured, machine-readable identifier for every known technique. For instance, the technique "Modify Registry" might have a specific ID like T1112. The objective of automated TTP extraction is to parse a sentence from a CTI report, such as "the attacker modified the Windows registry," and accurately map it to its corresponding ATT&CK ID.

To understand the state-of-the-art in this domain, the research team conducted an extensive literature review. They initiated a broad keyword search on academic platforms like DBLP and Google Scholar, targeting papers published since 2015 (the year MITRE ATT&CK was released) that dealt with extracting threat behaviors. This initial sweep yielded over 1,200 papers, which were then rigorously filtered down to a core set of 40 highly relevant works specifically proposing methods for TTP extraction from CTI reports.

These 40 works were categorized into three overarching approaches:

  1. Named Entity Recognition (NER) Approaches: These are primarily rule-based or statistical methods that identify and classify specific entities (like techniques or tools) within text. They often involve advanced word search patterns and linguistic rules.
  2. Classification Approaches: These are data-driven methods, frequently employing deep learning models like BERT (Bidirectional Encoder Representations from Transformers) or similar neural network architectures, to classify sentences or text segments into predefined TTP categories.
  3. Generative Large Language Models (LLMs) Approaches: This category includes modern LLMs such as Meta's LLaMA 3.1 or OpenAI's GPT-4o. These models leverage their vast pre-training to understand and generate human-like text, using prompts and in-context learning techniques to extract TTPs.

Across these three approaches, the researchers identified 16 fundamental Natural Language Processing (NLP) methods that form the foundation of various TTP extraction pipelines. While each work combines these methods uniquely, a significant problem emerged during the literature review: the vast majority of existing solutions were largely incomparable. This incomparability stemmed from several factors, including the use of custom, often proprietary, datasets for training and testing, the adoption of different TTP ontologies instead of the full MITRE ATT&CK framework, and the reliance on handcrafted optimizations that were not generalizable. This lack of standardization underscored the need for a unified empirical study to truly assess the field's progress.

Key Findings

▶ Watch: Methodology for comprehensive literature review (2015-present) (3:00)

The central finding of this Systematization of Knowledge (SoK) is a critical assessment of the current reliability of automated TTP extraction: we are not there yet. Despite a decade of research and the advent of sophisticated NLP technologies, current approaches are not sufficiently robust for fully automatic, practical deployment in real-world scenarios. This conclusion is drawn from a rigorous empirical study designed to overcome the comparability issues prevalent in prior research.

To achieve a fair comparison, the researchers re-implemented the 16 identified NLP methods and conducted over 100 experiments within a unified framework, utilizing standardized datasets and a consistent TTP ontology (the full MITRE ATT&CK framework). Two public datasets were chosen for evaluation:

  • Trend 2: A common test set from MITRE, containing labels for the 50 most common TTP classes found in CTI reports. This dataset was used primarily for qualitative claims.
  • AnusCTR (from Bosch): A smaller dataset but one that covers a broader range of unique TTP classes, used for quantitative claims.

The evaluation was conducted under two distinct scenarios to assess performance:

  1. Closed Set Scenario: This reflects how most existing papers are evaluated. The model is trained and tested on a subset of TTPs, meaning it "knows" exactly which TTPs might appear in the test set and can be optimized accordingly.
  2. Open Set (Real World) Scenario: This is a much harder and more realistic test. The model must be prepared for any possible TTP from the entire ATT&CK framework, without prior knowledge of which specific TTPs will occur in the test set.

Results in the Closed Set Scenario:

As expected and consistent with much of the existing literature, modern data-driven approaches—namely Classification and Generative LLMs—outperformed traditional Named Entity Recognition (NER) methods as the complexity of the task increased. When dealing with a small subset of 10 TTP labels, all approaches performed similarly. However, with an increasing number of classes to predict, the distinction became clear. The classification approach, in particular, demonstrated strong performance, achieving an F1 score still over 60% even with a subset of 118 labels. This outcome confirms that for well-defined problems with sufficient training data, deep learning models are highly effective.

Results in the Open Set (Real World) Scenario – A Crucial Reversal:

The most striking and significant finding emerged when the evaluation shifted to the more realistic open set scenario. Here, the results completely flipped. The traditional, rule-based NER approach consistently outperformed the modern classification and generative models. The performance of these more advanced, data-driven models degraded significantly when faced with the full breadth and unpredictability of the ATT&CK framework. This pivotal observation suggests that despite the substantial research focus on sophisticated data-driven techniques like BERT-based models and large language models, classic rule-based approaches demonstrate greater robustness and reliability in settings where perfect fine-tuning for all possible TTPs is not feasible, which is characteristic of real-world CTI analysis.

Major Bottlenecks Identified:

The study pinpointed two primary bottlenecks hindering progress, neither of which is solely about the NLP technology itself:

  1. Data Scarcity and Quality: There is a severe lack of large, high-quality, publicly available datasets for training and evaluating TTP extraction models. Existing datasets are typically small and cover only a tiny fraction of the vast ATT&CK framework. The research found that techniques like data augmentation, while helpful, do not fundamentally resolve this underlying data deficiency.
  2. ATT&CK Framework Ambiguity (Label Confusion): The study identified numerous instances of "label confusion" within the ATT&CK framework. This occurs when two distinct techniques are semantically so similar that it becomes nearly impossible for a model (or even a human) to differentiate them based solely on a single sentence. For example, "modifying the Windows registry to achieve persistence" is a different technique from "modifying the Windows registry to hide information," yet the textual representation of the action ("modifying the registry") can be identical. This inherent ambiguity creates a performance ceiling that no model can surmount without additional context beyond a single sentence.

In summary, the answer to "Are we there yet?" is a definitive no. Automated TTP extraction remains an active and open research problem. The path forward, as highlighted by these findings, is not solely about developing more complex NLP models but fundamentally addressing the critical limitations imposed by data availability and the inherent ambiguities within the definitions of the ATT&CK framework itself.

Technical Deep Dive

▶ Watch: Three main TTP extraction approaches: NER, Classification, LLMs (4:00)

The empirical study conducted by Büchel and his collaborators rigorously evaluated various NLP methodologies for TTP extraction by re-implementing 16 fundamental methods identified from the literature. These methods form the building blocks of the three primary approaches: Named Entity Recognition (NER), Classification, and Generative Large Language Models (LLMs). The goal was to compare their effectiveness under controlled and real-world conditions within a unified framework.

For Named Entity Recognition (NER) approaches, the pipeline typically involves several sequential steps designed to refine the input text and identify relevant entities. A common sequence includes:

  • Part-of-Speech (POS) Tagging: Identifying the grammatical role of each word (noun, verb, adjective, etc.), which helps in understanding sentence structure.
  • Lemmatization: Reducing words to their base or root form (e.g., "modifying," "modified," "modifies" all reduce to "modify"). This normalizes the text and improves recognition rates by treating different inflections of the same word as identical.
  • Related Word Detection: Techniques to identify synonyms or semantically similar terms that refer to the same concept, further enhancing the system's ability to recognize variations of a technique.

The output of this pipeline is typically a sentence or phrase linked to a specific MITRE ATT&CK ID, based on the identified entities and their context. NER systems often rely on carefully crafted rules, dictionaries, or statistical models trained on annotated data to identify patterns.

Classification approaches operate on a different principle, focusing on transforming sentences into a numerical representation (embedding) and then classifying these embeddings. The process generally involves:

  • Semantic Embedding: The input sentence is converted into a semantic embedding vector using pre-trained models (e.g., Sentence-BERT or other embedding models). These vectors capture the meaning and context of the sentence, allowing for numerical comparisons.
  • Supervised Classification Network: The embedding vector is then fed into a classification network (e.g., a feed-forward neural network or a more complex deep learning architecture) that has been trained on labeled data to map the vector to a specific TTP class.
  • Semantic Similarity-based Classification: Alternatively, classification can be performed by comparing the embedding of the input sentence to the embeddings of known TTP descriptions. The closest match (based on cosine similarity, for example) determines the TTP.
  • Data Augmentation: To combat the scarcity of training data, synthetic data can be generated (e.g., by paraphrasing existing sentences) to theoretically improve model robustness and generalization.

Generative Large Language Models (LLMs) represent the cutting edge, leveraging their vast pre-training. Their application in TTP extraction typically involves:

  • Prompt Engineering: Formulating specific instructions or questions (prompts) to guide the LLM to extract TTPs from CTI reports. The quality and structure of the prompt significantly impact performance.
  • In-Context Learning Methods: Techniques like Retrieval Augmented Generation (RAG) or Few-Shot Prompting are used to enrich the prompt with additional, relevant information. RAG involves retrieving relevant documents or examples from a knowledge base and providing them to the LLM as context. Few-shot prompting provides the LLM with a few examples of input-output pairs to guide its understanding of the task.
  • Supervised Fine-tuning: While LLMs are powerful out-of-the-box, traditional supervised training (fine-tuning) on domain-specific labeled data is also employed to further increase their performance for TTP extraction.

To ensure a fair and comprehensive comparison, the study meticulously defined its evaluation framework. The researchers utilized two publicly available datasets:

  • Trend 2: Provided by MITRE, this dataset focuses on the 50 most common TTP classes. Its primary role was to support qualitative analyses and provide a baseline for well-understood techniques.
  • AnusCTR (from Bosch): While smaller, this dataset boasts a greater diversity of unique TTP classes, making it suitable for quantitative evaluations, especially when assessing performance across a broader spectrum of ATT&CK techniques.

The distinction between the closed set and open set evaluation scenarios was crucial for understanding real-world applicability. In the closed set, models were evaluated on a subset of TTPs they were specifically optimized for, simulating a controlled environment. The results showed that classification and generative models excelled here, achieving F1 scores over 60% even with 118 labels, indicating their strength when the problem space is well-defined and data is available.

However, the shift to the open set scenario exposed a critical vulnerability in these advanced models. Here, models had to identify any TTP from the entire ATT&CK framework, without prior knowledge of which specific TTPs would appear. In this more realistic setting, the traditional NER approaches consistently outperformed the classification and generative models. This counter-intuitive result can be attributed to the fundamental differences in how these models generalize. Rule-based NER systems, while potentially less flexible, can often identify patterns or keywords associated with techniques even if those specific techniques haven't been extensively represented in training data, as long as the underlying linguistic rules hold. Data-driven models, particularly deep learning classifiers, tend to struggle significantly with out-of-distribution (OOD) detection or classes that are under-represented or entirely absent from their training data, leading to a sharp drop in performance in an open-ended environment. This highlights their reliance on comprehensive and diverse training data, which is precisely what the field currently lacks.

The study's identification of label confusion within the ATT&CK framework as a major bottleneck is also a significant technical insight. It points to a limitation not in the NLP models themselves, but in the underlying ontology. When two techniques, such as "T1112: Modify Registry" for persistence and "T1070.004: Indicator Removal on Host: File Deletion" which might also involve registry modification to remove traces, are conceptually distinct but share similar textual representations, even the most advanced NLP models struggle to differentiate them without richer context. This ambiguity creates an inherent ceiling on performance, irrespective of model sophistication.

Demo / Proof of Concept

▶ Watch: Our solution: Unified empirical study and re-implementation (8:00)

The talk did not feature a live demonstration of a specific tool or a proof-of-concept implementation. Instead, as a Systematization of Knowledge (SoK) paper, its primary contribution lies in the comprehensive empirical study and literature review it presents. The research focused on evaluating the current state-of-the-art in automated TTP extraction through rigorous experimentation with re-implemented NLP methods and standardized datasets, rather than showcasing a novel system.

Defensive Implications

▶ Watch: Standardized datasets and evaluation scenarios for comparison (9:00)

The findings from this comprehensive SoK have significant implications for cybersecurity defenders, particularly those involved in threat intelligence analysis, incident response, and security operations. The overarching message is one of caution and strategic application rather than full automation.

Firstly, the clear conclusion that automated TTP extraction is not yet reliable enough for fully automatic practical use means that organizations should not rely solely on current automated systems to derive TTPs from CTI reports. A "silver bullet" solution does not exist. Human analysts remain indispensable for contextual understanding, disambiguation, and making informed judgments.

However, automated tools can still play a crucial supporting role, depending on the specific defensive needs:

  • Named Entity Recognition (NER) systems are highlighted as a great starting point for human analysts. Their high precision, especially in open-set (real-world) scenarios, makes them effective at identifying potential TTPs and related entities within unstructured text. Defenders can leverage NER tools to efficiently highlight relevant portions of CTI reports, allowing human analysts to focus their attention and validate the extracted information. This reduces the manual burden of sifting through vast amounts of text.
  • Classification approaches can be highly effective if a defender's focus is on a small, known set of techniques (a closed-set scenario). For organizations that primarily monitor for a specific group of high-priority TTPs relevant to their threat model, a fine-tuned classification model could offer rapid and reasonably accurate identification. This might be useful for automated alerting on known adversary behaviors in specific contexts.

The identified bottlenecks also provide actionable insights for defenders:

  • Human-in-the-Loop is Essential: Given the current limitations, a human-in-the-loop approach is paramount. Automated tools should serve as aids to human intelligence, not replacements. Analysts should critically review any TTPs suggested by automated systems, especially when dealing with ambiguous language or techniques that could fall under multiple ATT&CK IDs.
  • Focus on High-Quality CTI Input: The performance of any automated system is directly tied to the quality of its input. Defenders should strive to produce or acquire CTI reports that are as clear, concise, and unambiguous as possible, which can mitigate some of the challenges posed by natural language processing.
  • Awareness of ATT&CK Ambiguities: Defenders should be cognizant of the identified "label confusion" within the ATT&CK framework. When interpreting TTPs, particularly those derived from automated means, it's important to consider that certain actions might be interpreted as multiple techniques, necessitating deeper contextual analysis. This awareness can help in developing more robust defensive strategies that account for such overlaps.
  • Contribution to Community Data: For organizations with the resources, contributing to or participating in efforts to build large, high-quality, publicly available, and well-annotated datasets for TTP extraction can significantly accelerate progress in the field. This collaboration is crucial for overcoming the primary data bottleneck.

In essence, defenders should embrace automated TTP extraction tools as valuable assistants that can streamline initial analysis and flag potential threats, but always with the understanding that human expertise is required for final validation and nuanced interpretation. The goal should be to augment human capability, not to achieve full, unmonitored automation.

Key Takeaways

  • Automated TTP extraction from CTI reports is an open research problem and is not yet reliable enough for fully automatic practical use in real-world scenarios.
  • Existing research solutions are largely incomparable due to the widespread use of custom datasets, differing TTP ontologies, and lack of open-source implementations, necessitating a unified empirical study.
  • In controlled (closed-set) environments with well-defined problems and sufficient data, modern data-driven approaches like classification and generative LLMs demonstrate superior performance.
  • In realistic (open-set) scenarios, where models must handle the full breadth of the ATT&CK framework without perfect fine-tuning, traditional rule-based Named Entity Recognition (NER) approaches proved more robust and reliable than advanced data-driven models.
  • The primary bottlenecks hindering progress are a severe lack of large, high-quality public datasets for training and evaluation, and inherent ambiguities or "label confusion" within the MITRE ATT&CK framework itself.
  • For practitioners, there is no "silver bullet." The best approach depends on specific needs, with NER systems offering high precision for human analyst support and classification models being effective for small, known sets of techniques. Human-in-the-loop analysis remains critical.

About the Speaker(s)

Marvin Büchel is a researcher from the University of Alenborg. The work presented in this talk, "SoK: Automated TTP Extraction from CTI Reports – Are We There Yet?", is a collaborative project involving multiple institutions and organizations. These include the University of Alenborg, the Polytenico de Milano, the University of Tentum, Accents, Siemens, and NEC Laboratories Europe, highlighting the interdisciplinary and broad scope of this research effort. Marvin Büchel's presentation focused on the critical evaluation of current automated TTP extraction methods and the identification of key challenges in the field.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Solid SoK that does what SoKs are supposed to do — establishes a unified evaluation framework, exposes the incomparability problem plaguing a decade of TTP extraction research, and delivers a genuinely useful finding: rule-based NER beats fancy LLMs in open-set conditions. Not groundbreaking, but honest and rigorous work that the community actually needs.

Heather Calloway (CISO) — SOLID

Rigorous academic work that delivers an honest, evidence-backed answer to a real operational question — automated TTP extraction isn't production-ready, and the bottlenecks aren't the models, they're the data and the ontology. Solid contribution to the research community, but it stops short of telling security program owners what to actually do with that finding.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)