Skynet the CTI Intern: Building Effective Machine Augmented...

Scott J Roberts (Head of Threat Research · Interpress)

BSidesSF 2024 · Day 1

Overview

This talk, "Skynet the CTI Intern: Building Effective Machine Augmented Intelligence," delivered by Scott J Roberts, Head of Threat Research at Interpres, delves into the practical application of large language models (LLMs) and generative AI within real-world Cyber Threat Intelligence (CTI) workflows. Roberts candidly shares his experimental journey, exploring the pros and cons of integrating these nascent technologies into daily operations, emphasizing the importance of setting realistic expectations and focusing on tangible outcomes rather than speculation.

Watch on YouTube

Visual summary for Skynet the CTI Intern: Building Effective Machine Augmented... by Scott J Roberts
Visual summary for Skynet the CTI Intern: Building Effective Machine Augmented... by Scott J Roberts

Key moments

  1. 04:00 Problem: Automating MITRE ATT&CK mapping for new threats.
  2. 07:00 Use Case 1: Summary generation for threat articles (NLTK vs. LLM).
  3. 09:00 LLM prompting for concise, 'Bluff' style summaries.
  4. 10:30 Use Case 2: One-off data generation for demonyms (ISO 3166 mapping).
  5. 14:00 Use Case 3: MITRE ATT&CK technique extraction (regex vs. LLM).
  6. 16:00 Attempting technique extraction with LangChain tool mode.
  7. 19:00 Use Case 4: Programmatic merging of STIX 2 intrusion set objects.
  8. 23:00 Validating merged STIX 2 objects for JSON and STIX 2 compliance.

Skynet the CTI Intern: Building Effective Machine Augmented...

Speakers: Scott J Roberts

Conference: BSidesSF 2024

YouTube: https://www.youtube.com/watch?v=B-rsj2uZBKA

Overview

This talk, "Skynet the CTI Intern: Building Effective Machine Augmented Intelligence," delivered by Scott J Roberts, Head of Threat Research at Interpres, delves into the practical application of large language models (LLMs) and generative AI within real-world Cyber Threat Intelligence (CTI) workflows. Roberts candidly shares his experimental journey, exploring the pros and cons of integrating these nascent technologies into daily operations, emphasizing the importance of setting realistic expectations and focusing on tangible outcomes rather than speculation.

The core of the presentation revolves around Roberts' quest to automate and augment the labor-intensive aspects of CTI, particularly after his human intern, Chandler, moved on. His work at Interpres, a continuous threat exposure management platform, necessitates a deep and current understanding of a vast and ever-growing landscape of threat actors. This requires extensive reading, analysis, and mapping of threat intelligence to frameworks like MITRE ATT&CK, often at a pace faster than official updates. The talk serves as a practical guide for security professionals looking to leverage AI to enhance efficiency and effectiveness in their CTI efforts, demonstrating both surprising successes and areas requiring further refinement.

Roberts' approach is grounded in hands-on experimentation, showcasing how LLMs can tackle diverse CTI challenges, from generating concise summaries of lengthy threat reports to extracting specific adversary techniques and even merging complex structured data like STIX 2 objects. By sharing his methodology, including prompt engineering techniques and the use of various tools, he provides valuable insights for organizations considering or already embarking on their own AI integration journeys, ultimately advocating for a pragmatic, iterative approach to adopting machine-augmented intelligence in cybersecurity.

Background

▶ Watch: Problem: Automating MITRE ATT&CK mapping for new threats. (04:00)

Scott J Roberts' role as Head of Threat Research at Interpres, a continuous threat exposure management platform, places him at the forefront of understanding and tracking a rapidly expanding universe of threat actors. This demanding position requires constant ingestion and analysis of vast amounts of threat intelligence, with a critical need to represent this information, particularly adversary tactics, techniques, and procedures (TTPs), within the MITRE ATT&CK framework. The challenge, as Roberts highlights, is that the volume of new threats and intelligence often outpaces the ability of even dedicated teams like MITRE to document and integrate them promptly. This creates a gap where organizations need to perform their own ATT&CK mapping and analysis more quickly.

Initially, Roberts relied on human assistance for these labor-intensive tasks. He had an intern, Chandler, a senior from Utah State University, who proved "super useful" in handling tasks that Roberts understood but lacked the time to execute, such as double-checking information or working through specific data collection efforts. Chandler's departure, as college students are "want to do," created a void that prompted Roberts to seek automated solutions.

His search for automation led him to a literature review, where he identified three key resources that informed his experimental approach:

  1. Google Cloud's "Five phases of threat intelligence and AI": This post offered conceptual ideas but lacked direct applicability to his specific workflow challenges.
  2. Google's "Supercharging security with generative AI": Similar to the first, it provided an interesting read, primarily valuable for its linked resources rather than immediate insights.
  3. Thomas's "Applying LLMs to threat intelligence": This blog post, found on Medium behind a paywall, was deemed "absolutely fantastic" by Roberts, who paid for a subscription specifically to access it. He praised its detailed exploration of prompt engineering, few-shot prompting, and Retrieval Augmented Generation (RAG) in actual code, providing concrete examples that could be replicated and built upon.

The overarching problem Roberts aimed to address was the sheer volume of data and the manual effort required to process it, particularly in building timely and accurate MITRE ATT&CK representations. His goal was to leverage emerging LLM capabilities to augment his team's intelligence gathering and analysis, effectively seeking to replace the "Chandler-shaped hole" in his workflow with machine-augmented intelligence.

Key Findings

▶ Watch: LLM prompting for concise, 'Bluff' style summaries. (09:00)

Scott J Roberts' series of experiments yielded a range of outcomes, demonstrating both the significant potential and the current limitations of integrating LLMs into CTI workflows. His key findings are categorized by the specific use cases he explored:

  • Summary Generation: Highly Successful and Operational. This was the most straightforward and immediately impactful success. LLMs, specifically the OpenAI conversational endpoint, proved vastly superior to traditional Natural Language Toolkit (NLTK) based summarization. The LLM-generated summaries were concise, relevant, and human-readable, effectively filtering out irrelevant technical indicators (like SHA256 hashes) that plagued the NLTK approach. This capability is now fully integrated into Interpres' automation, significantly improving the efficiency of evaluating threat intelligence articles.
  • One-off Data Generation: Mostly Successful with Manual Oversight. For tasks requiring the generation of specific, often unique, lists or definitions, LLMs (accessed via Copilot in this case) were highly effective in accelerating the process. Roberts successfully used this for generating a comprehensive list of "demonyms" (e.g., "Chinese" for "China") to map to ISO 3166-3 country codes. While the LLM required iterative prompting ("Can I have more please?") and the output necessitated manual verification to ensure accuracy and completeness, it drastically reduced the time and effort compared to manual research.
  • ATT&CK Technique Extraction: Mixed Success, Needs Continued Work. The attempt to automatically extract MITRE ATT&CK techniques from unstructured text proved more challenging. Initial programmatic attempts using the OpenAI API and a structured JSON output format yielded inconsistent results, sometimes hallucinating techniques or misidentifying sub-techniques. An experiment with LangChain's "tool mode" to enforce output structure also failed to produce the desired results, which Roberts attributed to potential user error in prompting. However, direct, interactive prompting of the LLM by pasting text and asking for JSON output of techniques achieved "perfect alignment" with true positives from the example text. This indicates that while programmatic integration requires more refinement, the underlying LLM capability is promising.
  • STIX 2 Object Merging: Unexpectedly Successful. This was the most surprising and complex success. Roberts tackled the problem of reconciling two different STIX 2 intrusion set objects (e.g., one for APT28 and another for Microsoft's "Forest Blizzard," which is the same group) into a single, valid, and merged object. Despite the inherent finickiness of the STIX 2 JSON standard, the LLM, after some prompt refinement, was able to take two distinct objects and produce a single, valid, and parseable STIX 2 bundle. Roberts admitted he "did not expect this one was going to work and it did," suggesting its potential for operationalization in the near future.
  • Code Assistant (GitHub Copilot): Bonus Finding, Highly Recommended. Beyond the specific CTI tasks, Roberts highlighted the pervasive utility of code assistants like GitHub Copilot. He used it extensively throughout his experiments, including generating the content for the STIX 2 objects themselves. He views such tools as a "superpower" for developers, enabling even "mediocre developers" to be significantly more efficient and effective, thereby accelerating automation efforts across the board.

Overall, Roberts concluded that the experimentation itself was invaluable, often defying initial expectations. He stressed that practical application and testing are far more effective than speculation when assessing the true capabilities of LLMs in a security context.

Technical Deep Dive

▶ Watch: Use Case 3: MITRE ATT&CK technique extraction (regex vs. LLM). (14:00)

Scott J Roberts' exploration into machine-augmented CTI involved a series of distinct technical experiments, primarily leveraging OpenAI's conversational endpoint and Python for programmatic interaction, alongside other tools like LangChain and GitHub Copilot.

Summary Generation

The initial approach for summary generation relied on NLTK (Natural Language Toolkit), a common library for natural language processing. However, this method proved inadequate for CTI reports, often producing summaries filled with irrelevant technical indicators like SHA256 hashes, which were unhelpful for quick relevance assessment.

The LLM-based solution involved a Python script that:

  1. Used a "pretty easy little scraping library" to extract article text.
  2. Interacted with the OpenAI API.
  3. Employed specific prompt engineering:
  • User Prompt: "summarize this into one paragraph ignore indicators of compromise security vendor tools and contributors by name." This prompt explicitly directed the LLM to focus on the narrative and exclude noise.
  • System Message: "You are a cyber threat intelligence analyst with a concise Bluff writing style." The "Bluff" (Bottom Line Up Front) style was crucial for generating summaries that were immediately actionable and easy to digest for analysts and customers.

The result was a significantly improved, human-readable summary that met the design and utility requirements, leading to its full operationalization.

One-off Data Generation

For generating specific, one-time data sets, Roberts utilized GitHub Copilot as an interactive LLM interface. The primary example was creating a list of "demonyms" (e.g., "Chinese" for "China") to map to ISO 3166-3 country codes. This task was chosen because manually compiling such a list would involve extensive research across numerous Wikipedia pages.

The process involved:

  1. Directly prompting Copilot to "generate this for me."
  2. Iterative prompting: When Copilot stopped after a few entries or at a certain point (e.g., "stopped at M"), Roberts would simply ask, "Can I have more please?" This continued until the full list, from "Afghan" to "Zimbabwean," was generated.
  3. Manual Verification: The generated list was then manually checked against the ISO 3166-3 standard to ensure accuracy and correct mappings.

This method, while requiring human oversight, drastically accelerated a task that would otherwise be prohibitively time-consuming.

ATT&CK Technique Extraction

The goal was to extract MITRE ATT&CK techniques from unstructured text and output them in a structured JSON format. This was a more complex challenge due to the nuanced nature of describing TTPs.

Initial Programmatic Attempt (OpenAI API):

  1. User Prompt: "extract the miter attack techniques from the given text in a Json format of [example JSON structure]." The example JSON provided a template for the desired output.
  2. System Message: "You are a cyber threat intelligence automation outputting Json data." This aimed to constrain the LLM's output to a machine-readable format.

The results were mixed: the LLM identified some techniques correctly (e.g., T1105) but also hallucinated others and struggled with sub-techniques (e.g., identifying T1059 instead of T1059.003).

LangChain Experiment:

Roberts then attempted to use LangChain, an orchestration framework for LLMs, hoping to achieve more structured and reliable output.

  1. Tool Mode: LangChain's "tool mode" was used to define a specific object structure for the output, including technique name, sentence (where it matched), and technique ID.
  2. System Message: "You are an expert attack extraction algorithm only extract relevant information from the text if you don't know the value you're asked to abstract just return null." This was intended to prevent hallucinations and ensure data quality.

This attempt, however, did not work as expected, returning an empty or malformed output. Roberts acknowledged that this might be due to his own incorrect usage of LangChain.

Direct Prompting (Interactive):

The most successful approach for technique extraction involved directly pasting the source text into an LLM interface and using simple, direct prompts:

  1. "extract the miter attack techniques."
  2. "give it to me as Json."

This interactive method yielded "perfect alignment" between the extracted techniques and the true positives from the example text, demonstrating the LLM's capability when directly prompted, even if programmatic integration remained elusive.

STIX 2 Object Merging

This was arguably the most technically challenging and surprisingly successful experiment. The task involved merging two complex STIX 2 intrusion set objects (JSON representations of threat groups) into a single, valid STIX 2 bundle. The example used was merging an object for "APT28" with one for "Forest Blizzard" (Microsoft's new name for APT28), which had slightly different data points (e.g., secondary motivations, dates, goals).

The process involved:

  1. Creating two STIX 2 intrusion set objects in JSON format (Roberts noted GitHub Copilot generated the content for these objects).
  2. Initial Prompt: "I'm going to give you both of these objects I want you to merge them together."

The initial output was a valid STIX bundle, but it duplicated some values and picked up others incorrectly.

  1. Refined Prompt: "I only want one object and I need you to put the values together a little bit more efficiently."

The refined prompt led to the LLM producing a single, valid, and parseable STIX 2 object. Roberts performed two critical validations:

  • JSON Validity: Confirmed the output was valid JSON.
  • STIX 2 Validity: Confirmed the output was also valid STIX 2, which is a more stringent standard.

The success of this experiment, given the "finicky" nature of STIX 2 (a JSON standard based on an XML standard), was a significant surprise to Roberts, highlighting the LLM's ability to handle complex, structured data transformations.

Code Assistant (GitHub Copilot)

Throughout all experiments, GitHub Copilot played a crucial "bonus" role. Roberts emphasized its utility for general code generation, including the creation of the STIX 2 object content used in the merging experiment. He views such tools as a "superpower" that significantly boosts the efficiency and effectiveness of developers, enabling faster automation and more verifiable results in security.

Demo / Proof of Concept

▶ Watch: Attempting technique extraction with LangChain tool mode. (16:00)

Scott J Roberts' presentation itself served as a series of practical demonstrations, showcasing the capabilities and limitations of LLMs across various CTI tasks. While not a live coding demo in the traditional sense, he presented clear "before and after" scenarios and walked the audience through his iterative prompting process.

  1. Summary Generation Demonstration:
  • Roberts first showed an example of a threat intelligence article about the "zerobot stressor."
  • He then displayed the output from their original NLTK-based summarization, which was a block of text dominated by SHA256 hashes and other indicators of compromise, rendering it unhelpful for quick assessment. He humorously noted that this output would "make my designer cry."
  • Next, he presented the summary generated by the OpenAI conversational endpoint, using his refined prompt and system message. This summary was a concise, human-readable paragraph that accurately conveyed the core information of the article without the technical noise, clearly demonstrating the LLM's superior performance.
  1. One-off Data Generation Demonstration:
  • Roberts illustrated the process of generating a list of "demonyms" (e.g., "Chinese" for "China") using GitHub Copilot.
  • He showed screenshots of his interaction, starting with a request for the list, then iteratively asking "Can I have more please?" as Copilot provided partial lists, eventually reaching the full alphabetized list from "Afghan" to "Zimbabwean." This visually conveyed the interactive and iterative nature of using LLMs for specific data collection.
  1. ATT&CK Technique Extraction Demonstration:
  • Using an example from a Huntress blog post, Roberts demonstrated the challenges of extracting MITRE ATT&CK techniques.
  • He showed the output from his initial programmatic attempt with the OpenAI API, highlighting instances where the LLM either missed techniques or "threw in a couple other techniques that we've never seen before." He presented a table comparing the LLM's findings against the true positives, illustrating the "one and a half true positives" (e.g., T1105 correctly identified, but T1059 instead of T1059.003).
  • He then briefly mentioned the failed LangChain attempt, showing its empty or malformed output.
  • Finally, he demonstrated the success of direct prompting by showing the LLM's output when the article text was simply pasted in, resulting in "perfect alignment" with the true positives in JSON format.
  1. STIX 2 Object Merging Demonstration:
  • This was a compelling proof of concept for handling complex structured data. Roberts presented two distinct STIX 2 intrusion set objects in JSON format: one for "APT28" and another for "Forest Blizzard" (Microsoft's new name for the same group). He pointed out their subtle differences in fields like "secondary motivations" and "dates."
  • He then showed the LLM's initial merged output, which, while a valid STIX bundle, had some inaccuracies.
  • Following a refined prompt, he displayed the final merged STIX 2 object. Crucially, he emphasized that this output was not only valid JSON but also successfully parsed as valid STIX 2, a testament to the LLM's ability to understand and manipulate complex data schemas. He expressed genuine surprise at this success.
  1. Code Assistant Utility:
  • While not a direct demo, Roberts highlighted that GitHub Copilot was used to generate the content for the STIX 2 objects themselves, underscoring its utility as a general-purpose code assistant throughout his experiments.

These demonstrations collectively provided concrete evidence for Roberts' findings, illustrating both the practical benefits and the current hurdles of integrating machine-augmented intelligence into CTI.

Defensive Implications

▶ Watch: Validating merged STIX 2 objects for JSON and STIX 2 compliance. (23:00)

The insights gleaned from Scott J Roberts' experiments offer several significant defensive implications for cybersecurity teams, particularly those involved in Cyber Threat Intelligence (CTI) and incident response.

  1. Enhanced CTI Efficiency and Focus: The most immediate implication is the potential for LLMs to drastically reduce the manual burden of processing vast amounts of threat intelligence. By automating tasks like summary generation, CTI analysts can spend less time reading through lengthy reports to determine relevance and more time on higher-value activities such as in-depth analysis, strategic planning, and proactive threat hunting. This shift allows CTI teams to be more agile and responsive to emerging threats.
  1. Accelerated MITRE ATT&CK Mapping: While still requiring refinement, the ability of LLMs to assist in extracting and mapping adversary techniques to the MITRE ATT&CK framework is a powerful defensive tool. Faster and more consistent mapping can lead to:
  • Improved Detection Engineering: Security operations centers (SOCs) can more quickly develop and deploy detections for new TTPs.
  • Better Defensive Posture Assessment: Organizations can rapidly identify gaps in their defensive coverage against specific adversary behaviors.
  • Streamlined Threat Hunting: Analysts can use LLM-assisted ATT&CK mappings to inform their threat hunting queries and methodologies.
  1. Streamlined Threat Intelligence Management and Reconciliation: The unexpected success in merging complex STIX 2 objects has profound implications for managing threat intelligence repositories. Organizations often struggle with integrating data from multiple sources, dealing with evolving threat actor naming conventions (e.g., APT28 vs. Forest Blizzard), and maintaining a consistent, unified view of threat groups. LLMs can automate the reconciliation of disparate intelligence objects, ensuring that defensive systems and intelligence platforms are fed accurate, up-to-date, and non-redundant information. This reduces the risk of misidentification or incomplete understanding of adversary capabilities.
  1. Empowered Security Developers: The "superpower" of code assistants like GitHub Copilot extends directly to defensive capabilities. By making security professionals more efficient at coding, these tools enable:
  • Rapid Tool Development: Security teams can quickly build custom scripts, automation, and integrations tailored to their specific defensive needs.
  • Faster Incident Response: Automated playbooks and response actions can be developed and refined more rapidly.
  • Proactive Security Engineering: Security engineers can implement defensive controls and data pipelines with greater speed and accuracy.
  1. Importance of Human Oversight and Verification: Despite the successes, Roberts' experiments underscore that LLMs are not infallible. Hallucinations, incorrect extractions, and the need for iterative prompting highlight that human oversight and verification remain critical. Defenders must implement robust validation processes for any LLM-generated output, especially when it directly impacts detection rules, intelligence feeds, or defensive actions. This ensures that automated systems are not operating on flawed or misleading information.
  1. Strategic Resource Allocation: Roberts' point about the cost and latency introduced by LLM API calls is a practical defensive consideration. Defenders should strategically apply LLMs to tasks where the efficiency gains significantly outweigh these factors. For instance, simple Indicator of Compromise (IOC) extraction, which can be done efficiently with regular expressions, might not warrant the use of an LLM, reserving LLMs for more complex, unstructured data analysis where their unique capabilities shine.

In essence, machine-augmented intelligence, when applied judiciously and with proper validation, can significantly enhance the speed, accuracy, and scalability of defensive operations, allowing security teams to better keep pace with the ever-evolving threat landscape.

Key Takeaways

  • LLMs are powerful for specific tasks, not a panacea: While LLMs demonstrate remarkable capabilities, they are not a universal solution. Their effectiveness is maximized when applied to well-defined, often data-intensive, and repetitive tasks within CTI workflows.
  • Experimentation over speculation is crucial: The true utility of LLMs in cybersecurity can only be understood through hands-on experimentation. Speculation about their capabilities or limitations often misses the mark; practical testing reveals unexpected successes and areas for improvement.
  • Prompt engineering is key to success: The way questions and instructions are framed to the LLM, including the use of system messages to define its persona and constraints, profoundly impacts the quality and relevance of the output. Iterative refinement of prompts is often necessary.
  • LLMs can handle complex structured data: The ability to merge and validate complex STIX 2 objects demonstrates that LLMs can go beyond simple text generation to understand and manipulate structured data schemas, opening doors for advanced data management in CTI.
  • Code assistants are a "superpower" for security professionals: Tools like GitHub Copilot significantly boost developer efficiency, enabling security practitioners to build automation and custom tools faster, thereby enhancing overall security capabilities.
  • The next generation of talent is highly capable: The effectiveness of his former intern, Chandler, highlighted that college students and new entrants to the field possess impressive skills and can be remarkably effective in complex areas like threat intelligence, underscoring the value of mentorship and bringing in new talent.

About the Speaker(s)

Scott J Roberts is the Head of Threat Research at Interpres. He brings a wealth of experience to the field, with a 20-year career that has spanned various critical areas of cybersecurity, including incident response, threat detection, and cyber threat intelligence. In addition to his professional role, Roberts is actively engaged in academia, pursuing a Master's degree in anticipatory intelligence at Utah State University. He also shares his expertise by teaching cybersecurity at Utah State, contributing to the development of the next generation of security professionals.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk cuts through the usual LLM hype by presenting a series of practical experiments integrating machine augmentation into real-world cyber threat intelligence workflows. The speaker demonstrates both successes and failures in tasks ranging from summary generation and one-off data creation to complex MITRE ATT&CK technique extraction and STIX 2 object merging. The candid approach to what worked and what didn't, backed by actual code and results, provides valuable insights for CTI teams looking to leverage these tools effectively.

Heather Calloway (CISO) — STRONG ACCEPT

This session provides a pragmatic and evidence-based exploration of integrating large language models into cyber threat intelligence operations. By demonstrating how LLMs can augment tasks like threat summary generation, data mapping, and complex STIX 2 object merging, the speaker offers tangible pathways for CTI teams to enhance efficiency and accuracy. The candid discussion of both successes and challenges in real-world experiments provides valuable insights for security leaders considering the operational impact of these emerging technologies.

→ Top-rated talks at BSidesSF 2024

All talks from BSidesSF 2024