RACONTEUR: A Knowledgeable, Insightful, and Portable LLM-Powered Shell Command Explainer

Jiangyi Deng

Network and Distributed System Security (NDSS) Symposium 2025 · Day 2 · Software Security: Vulnerability Detection

Overview

In the evolving landscape of cyber threats, understanding the true intent and capabilities of malicious shell commands is a critical yet often challenging task for security analysts. The "RACONTEUR: A Knowledgeable, Insightful, and Portable LLM-Powered Shell Command Explainer" talk, presented by Jiangyi Deng at the NDSS Symposium, introduces an innovative system designed to demystify these complex command-line instructions. Raconteur leverages the power of large language models (LLMs) to provide detailed, step-by-step explanations, identify the underlying behavior, and map these actions to the standardized MITRE ATT&CK framework, thereby significantly aiding Security Operations Center (SOC) personnel in threat identification and response.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction, problem of complex shell commands, Raconteur's goal
  2. 2:26 Overview of Raconteur's three core system components
  3. 3:00 Deep dive into the customized LLM behavior explainer
  4. 5:00 Identifying command intent and mapping to MITRE ATT&CK
  5. 6:50 Enhancing explanations and suppressing LLM hallucination
  6. 8:00 Raconteur's performance metrics and comparison to baselines
  7. 9:30 Human evaluation results and overall system conclusion
  8. 10:30 Confirmation of Raconteur's fine-tuned model release

RACONTEUR: A Knowledgeable, Insightful, and Portable LLM-Powered Shell Command Explainer

Speakers: Jiangyi Deng

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=fUQ_l_4YD30

Overview

In the evolving landscape of cyber threats, understanding the true intent and capabilities of malicious shell commands is a critical yet often challenging task for security analysts. The "RACONTEUR: A Knowledgeable, Insightful, and Portable LLM-Powered Shell Command Explainer" talk, presented by Jiangyi Deng at the NDSS Symposium, introduces an innovative system designed to demystify these complex command-line instructions. Raconteur leverages the power of large language models (LLMs) to provide detailed, step-by-step explanations, identify the underlying behavior, and map these actions to the standardized MITRE ATT&CK framework, thereby significantly aiding Security Operations Center (SOC) personnel in threat identification and response.

The core motivation behind Raconteur stems from the recognition that shell commands, particularly those used in reverse shells or other stealthy attack vectors, can be exceedingly intricate and require years of expertise to interpret accurately. While LLMs show promise due to their extensive training on code, they are prone to hallucination—generating factually incorrect or nonsensical information. Raconteur directly addresses this challenge by designing a system that prioritizes correctness, comprehensiveness, and insightfulness, even offering the flexibility of local deployment for enhanced privacy and security.

This work is crucial because it bridges the gap between the raw complexity of shell commands and the need for actionable intelligence by security analysts. By providing an automated, reliable, and detailed explanation of command behavior, Raconteur empowers defenders to more rapidly identify and respond to sophisticated cyber attacks, ultimately strengthening an organization's defensive posture against a persistent and evolving threat landscape.

Background

▶ Watch: Introduction, problem of complex shell commands, Raconteur's goal (0:00)

The pervasive nature of cyber attacks continues to pose significant threats to organizations globally, necessitating robust security measures and vigilant Security Operations Centers (SOCs). A recurring challenge for these centers lies in the analysis of shell commands, which are frequently exploited by attackers as initial springboards to gain remote control over victim systems. These commands, often crafted to be highly stealthy and intricate, can mask their true malicious intent, making their identification and interpretation a formidable task. For instance, a seemingly innocuous string of characters might, upon expert analysis, reveal itself to be a sophisticated reverse shell designed to establish persistent command-and-control communication. Traditionally, discerning the malicious nature and specific actions of such commands demands years of specialized experience from security analysts, a resource that is both scarce and expensive.

The advent of large language models (LLMs) has opened new avenues for automating complex analytical tasks, particularly given their extensive training on vast repositories of code. The potential to leverage LLMs as an "explainer" for shell commands is appealing, promising to democratize expertise and accelerate threat analysis. However, a significant obstacle in deploying general-purpose LLMs for security analysis is their inherent susceptibility to hallucination. This phenomenon, where LLMs generate plausible but factually incorrect information, is unacceptable in security contexts where precision and accuracy are paramount. An erroneous explanation of a malicious command could lead to misidentification, delayed response, or even the dismissal of a critical threat.

To address these challenges, a robust shell command explanation system must meet several key criteria. It needs to provide a detailed, step-by-step explanation of the command's execution flow. Crucially, it must be capable of identifying the specific malicious behaviors or intents embedded within the command. Finally, for the insights to be actionable and universally understood within the security community, the system should map these identified behaviors to standard frameworks, such as the MITRE ATT&CK knowledge base, which categorizes adversarial tactics and techniques. Beyond these functional requirements, the system must guarantee comprehensiveness, insightfulness, and, most critically, correctness. Furthermore, for organizations dealing with highly sensitive or proprietary command data, the ability to deploy a local LLM for explanation—rather than relying on external, cloud-based services—becomes a non-negotiable requirement for data privacy and security. Raconteur was designed with these exacting demands in mind, aiming to bridge the critical gap between raw shell command data and actionable security intelligence.

Key Findings

▶ Watch: Deep dive into the customized LLM behavior explainer (3:00)

Raconteur represents a significant advancement in automated shell command analysis, primarily through its innovative three-component architecture designed to deliver correct, comprehensive, and insightful explanations while mitigating the inherent weaknesses of general-purpose LLMs. The system’s key findings and contributions can be summarized as follows:

Firstly, Raconteur introduces a holistic toolkit comprising a behavior explainer, an intent identifier, and a documentation augmented enhancer. Each component plays a crucial role in achieving the system's objectives, from generating detailed command breakdowns to mapping malicious activities to industry-standard frameworks and ensuring factual accuracy. This modular design allows for specialized training and optimization, contributing to the overall robustness of the system.

Secondly, the evaluation demonstrates Raconteur's superior performance compared to vanilla large language models. The system achieved substantial improvements in explanation quality, ranging from 37% to an impressive 137% over baseline LLMs. This significant enhancement underscores the effectiveness of Raconteur's specialized fine-tuning and augmentation strategies in overcoming the limitations of general models, particularly their susceptibility to hallucination, when applied to a highly specific and sensitive domain like shell command analysis.

Thirdly, Raconteur proved highly effective in facilitating the comprehension of commands. Through both quantitative metrics and a user study involving undergraduate students, the system demonstrated its ability to produce explanations that are not only technically accurate but also easily understandable by human analysts. The user study, which compared Raconteur's explanations against baselines across five subjective metrics, consistently showed Raconteur outperforming other models, confirming its practical utility in a real-world analytical context.

Finally, the project emphasizes portability and transparency by actively releasing its fine-tuned models and the underlying data sets used for training. This commitment to open science and community contribution allows other researchers and security professionals to audit, replicate, and further develop the work, fostering broader adoption and innovation in the field of LLM-powered security tools. This addresses a critical need for verifiable and shareable resources in a domain where proprietary models often lack transparency.

Technical Deep Dive

▶ Watch: Enhancing explanations and suppressing LLM hallucination (6:50)

Raconteur's architectural strength lies in its three interdependent components, each meticulously designed and trained to address specific challenges in shell command explanation and analysis. The system's design reflects a deep understanding of LLM capabilities and limitations, particularly in the context of security operations.

Behavior Explainer

The Behavior Explainer is the foundational component of Raconteur, responsible for generating the initial, detailed, step-by-step explanation of a given shell command. This module is built upon a customized and fine-tuned Large Language Model (LLM), specifically a variant of the ChatGLM2 model with 6 billion parameters. The choice to fine-tune a specialized LLM rather than relying on a general-purpose model is critical for two primary reasons:

  1. Diversified User Prompts: Security analysts may inquire about commands using a wide range of phrasing and levels of detail. The fine-tuned LLM is trained to interpret and respond effectively to these varied prompts, ensuring flexibility and usability.
  2. Factuality Assurance: General LLMs are prone to hallucination, which is unacceptable in security analysis. To counteract this, Raconteur employs supervised fine-tuning with a meticulously constructed, high-quality dataset. This dataset is created by extracting knowledge from various code libraries and other authoritative sources, ensuring that the training data is both comprehensive and factually accurate. The dataset also incorporates diversified user prompts to enhance the model's adaptability.

The training process for the Behavior Explainer was substantial, requiring four A100 GPUs and four days of continuous training. Over this period, the model processed 232 million tokens, allowing it to become an expert in dissecting and explaining the intricate behaviors of shell commands with a high degree of precision and correctness. This specialized training imbues the LLM with the domain-specific knowledge required to accurately interpret complex command syntax and execution flows.

Intent Identifier

Following the detailed explanation from the Behavior Explainer, the Intent Identifier takes over to contextualize the command's actions within a broader threat landscape. Its primary objective is to identify the tactics (why the command is performed) and techniques (how the command is executed) by mapping the LLM's behavioral explanation to the standardized descriptions within the MITRE ATT&CK knowledge base. This mapping is crucial for security analysts, as it provides a common language for understanding, categorizing, and responding to adversarial actions.

A significant challenge in this process is the semantic gap between the natural language explanation generated by the LLM and the structured, often concise descriptions found in the MITRE ATT&CK framework. To bridge this gap, Raconteur employs a BERT2Vec model (as stated in the transcript, likely referring to a BERT-based embedding model or a similar vector space model). This model is trained on a specially constructed high-quality dataset designed to learn the semantic similarities between the LLM's output and the MITRE ATT&CK descriptions. By transforming both types of text into a shared vector space, the BERT2Vec model can accurately identify the most similar tactics and techniques, enabling precise classification of the command's intent and method. The talk mentions that all five materializations of their intent identifier showed better performance than baselines, suggesting flexibility in its underlying implementation.

Documentation Augmented Enhancer

The final component, the Documentation Augmented Enhancer, serves as a critical safeguard against LLM hallucination and a mechanism for expanding the system's knowledge base to include unseen or private commands. While the Behavior Explainer is fine-tuned, no LLM can encompass all possible commands, especially those that are new, proprietary, or highly specialized.

This enhancer works by training another dedicated model whose function is to retrieve the most relevant textual information from external documentation sources. When Raconteur encounters a command it struggles to explain accurately, or when a user explicitly requests more context, this model searches a corpus of documentation for passages related to the command's components or functions. The retrieved, verified text is then combined with the initial user prompt and fed back into the Behavior Explainer. This retrieval-augmented generation (RAG) approach effectively grounds the LLM's output in factual, external information. By providing the LLM with relevant documentation at the time of generation, the system can significantly suppress hallucination, enhance the accuracy of explanations for novel commands, and provide deeper insights that might not be present in its initial training data. This mechanism ensures that Raconteur can maintain high factual correctness and provide useful explanations even for commands that were not explicitly part of its training set or for private commands that are not publicly available.

Demo / Proof of Concept

▶ Watch: Raconteur's performance metrics and comparison to baselines (8:00)

While the presentation did not feature a live, interactive demonstration of the Raconteur system processing a shell command in real-time, the efficacy and practical utility of the framework were rigorously validated through comprehensive quantitative evaluations and a dedicated user study. The "Proof of Concept" in this context is manifested through the detailed performance metrics and the positive results of human-centric assessments.

The core evaluation involved comparing Raconteur against several general large language models across standard natural language processing (NLP) metrics for explanation quality, as well as accuracy and precision for the intent identification component. The results highlighted Raconteur's significant advantage, demonstrating an improvement of 37% to 137% over vanilla LLMs in generating high-quality explanations. Furthermore, all five materializations of Raconteur's intent identifier consistently outperformed baseline models in accurately classifying tactics and techniques.

Crucially, a user study was conducted to validate the subjective quality of Raconteur's explanations. Undergraduate students were recruited to compare explanations generated by Raconteur with those from baseline systems. Across five subjective metrics designed to assess comprehensibility, insightfulness, and overall utility, Raconteur consistently received higher ratings. This human-centric validation is vital, as it confirms that the system not only performs well on technical metrics but also genuinely facilitates human understanding, which is the ultimate goal for assisting security analysts. Thus, while a direct "demo" in the form of a live showcase was not the focus, the extensive evaluation serves as robust proof of concept for Raconteur's capabilities.

Defensive Implications

▶ Watch: Confirmation of Raconteur's fine-tuned model release (10:30)

Raconteur's innovative approach to shell command explanation carries significant defensive implications for organizations striving to enhance their cybersecurity posture. By automating and standardizing the analysis of complex commands, it offers tangible benefits for Security Operations Centers (SOCs), incident responders, and threat hunters.

  1. Accelerated Threat Identification and Response: The most immediate benefit is the ability to rapidly understand the intent and behavior of suspicious shell commands. Instead of requiring hours or even days of expert analysis, Raconteur can provide a detailed, step-by-step breakdown and MITRE ATT&CK mapping within moments. This acceleration in comprehension allows SOC analysts to quickly triage alerts, prioritize critical incidents, and initiate defensive actions much faster, significantly reducing mean time to detect (MTTD) and mean time to respond (MTTR).
  1. Empowering Less Experienced Analysts: The system acts as an expert assistant, democratizing the knowledge typically held by seasoned security professionals. Junior analysts can leverage Raconteur to gain a deep understanding of complex commands, reducing the learning curve and enabling them to contribute more effectively to threat analysis. This addresses the ongoing shortage of highly skilled cybersecurity talent.
  1. Enhanced Detection and Rule Creation: By consistently mapping command behaviors to specific MITRE ATT&CK techniques, Raconteur provides a standardized framework for understanding adversarial actions. This consistency can be used to inform the development of more precise and robust detection rules within Security Information and Event Management (SIEM) systems or Endpoint Detection and Response (EDR) solutions. Knowing exactly what a malicious command aims to achieve helps in crafting signatures that target specific TTPs rather than just generic patterns.
  1. Improved Threat Intelligence and Hunting: Raconteur can aid threat hunters in deciphering observed adversary behaviors. When encountering novel or highly obfuscated commands in logs or forensic artifacts, the system can quickly provide context, helping hunters to identify new TTPs and proactively search for similar indicators of compromise (IOCs) across their environments. The documentation augmentation feature is particularly valuable here for explaining unseen commands.
  1. Reduced Hallucination Risk in LLM Deployment: By demonstrating a methodology to mitigate LLM hallucination through specialized fine-tuning and retrieval-augmented generation, Raconteur offers a blueprint for safely integrating advanced AI into critical security functions. This builds trust in AI-powered tools, which is essential for their broader adoption in security.
  1. Privacy and Security for Sensitive Data: The design choice to enable local deployment of the fine-tuned LLM is a significant advantage for organizations with strict data privacy and compliance requirements. By keeping sensitive command data within the organization's control, Raconteur mitigates risks associated with sending potentially proprietary or confidential information to external, cloud-based LLM providers.

While Raconteur currently does not handle encrypted or obfuscated payloads, this limitation is acknowledged as future work. Addressing this would further enhance its defensive capabilities against increasingly sophisticated adversaries who employ such techniques to evade detection.

Key Takeaways

  • Raconteur is an LLM-powered system designed to explain complex shell commands, aiding security analysts in identifying cyber threats.
  • The system comprises three core components: a behavior explainer (fine-tuned ChatGLM2), an intent identifier (BERT2Vec for MITRE ATT&CK mapping), and a documentation augmented enhancer (for hallucination suppression and unseen commands).
  • Raconteur significantly outperforms vanilla LLMs, showing a 37% to 137% improvement in explanation quality and superior intent classification accuracy.
  • It effectively maps command behaviors to the MITRE ATT&CK framework, providing standardized and actionable threat intelligence.
  • The system prioritizes correctness, comprehensiveness, and insightfulness, validated through quantitative metrics and a user study confirming its utility for human comprehension.
  • Raconteur's design allows for local deployment, addressing privacy and security concerns for sensitive organizational data.
  • The project promotes transparency by releasing its fine-tuned models and datasets for community use and further research.
  • Current limitations include the inability to handle encrypted or obfuscated command payloads, which is noted as an area for future development.

About the Speaker(s)

The talk was presented by Jiangyi Deng. While specific titles or affiliations beyond "Ant Group" (mentioned as a collaborator) were not detailed in the transcript, it is evident that Jiangyi Deng is a researcher with expertise in applying AI security principles to practical cybersecurity challenges. The speaker noted that their original work is primarily in AI security, and extending frameworks like Raconteur to other domains, such as malware reverse engineering, would require specific domain expertise. This indicates a focus on the intersection of artificial intelligence and security, particularly in developing robust and reliable AI systems for defensive purposes.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Competent applied-ML security paper dressed up as a conference talk — fine-tuned LLM plus RAG for shell command explanation and MITRE ATT&CK mapping. The engineering is real, the problem is genuine, but the contribution is incremental: fine-tuning + RAG is a well-worn recipe by 2024, and the evaluation relies on NLP metrics and an undergrad user study rather than red-team or SOC deployment data.

Heather Calloway (CISO) — WEAK

Competent academic research on an LLM-powered shell command explainer with measurable improvement over vanilla models — but it stops at the tool and never reaches the operator, the program, or the institution. The defensive value is stated, not demonstrated, and the governance dimension is entirely absent.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025