Faux Data, Real Defense: ML advancements in data synthesis
Arjun Chakraborty (Detection Engineering Team · Databricks)
BSidesSF 2024 · Day 1
Overview
Arjun Chakraborty, a member of the detection engineering team at Databricks, presented a compelling talk titled "Faux Data, Real Defense: ML advancements in data synthesis" at BSidesSF 2024. The presentation delved into the critical challenge faced by security professionals: the difficulty in obtaining representative adversarial data to build and tune effective threat detection pipelines. Chakraborty proposed a novel approach leveraging Large Language Models (LLMs) to generate synthetic adversarial data, thereby empowering detection engineers to create more robust and efficient security mechanisms.

Key moments
- 0:50 The Detection Data Problem: Scarcity of Representative Data
- 1:30 Synthetic Data as a Solution for Better Detection Pipelines
- 2:00 Core Research Questions: Generation and Quality Evaluation of Adversarial Synthetic Data
- 3:30 Targeting Kubernetes API Server Audit Logs for Synthetic Data Generation
- 6:00 Novel Evaluation Framework: Fidelity, Reproducibility, and Accuracy Metrics
- 9:00 Fidelity and Reproducibility Results and Challenges
- 14:30 Accuracy Results: Effectiveness of Detections on Synthetic Data
- 18:00 Limitations and Future Work: Synthetic Data vs. Red Teaming, Benchmarks, and Advanced LLM Techniques
Faux Data, Real Defense: ML advancements in data synthesis
Speakers: Arjun Chakraborty
Conference: BSidesSF 2024
YouTube: https://www.youtube.com/watch?v=tha8w7EDu-4
Overview
Arjun Chakraborty, a member of the detection engineering team at Databricks, presented a compelling talk titled "Faux Data, Real Defense: ML advancements in data synthesis" at BSidesSF 2024. The presentation delved into the critical challenge faced by security professionals: the difficulty in obtaining representative adversarial data to build and tune effective threat detection pipelines. Chakraborty proposed a novel approach leveraging Large Language Models (LLMs) to generate synthetic adversarial data, thereby empowering detection engineers to create more robust and efficient security mechanisms.
The core of the talk addressed two fundamental questions: can adversarial synthetic data be efficiently generated to improve detection, and more importantly, how can the quality of this synthetic data be effectively evaluated? With the rapid advancements in LLMs, generating data has become increasingly accessible, shifting the primary challenge to validating its utility and accuracy for security applications. Chakraborty's research focuses on providing detection engineers with the necessary tools to enhance their pipelines, ultimately leading to more resilient defenses against evolving threats.
This work is particularly significant given the ever-expanding attack surface, the growth of enterprise infrastructures, and the increasing integration of AI into both offensive and defensive security strategies. By demonstrating the potential of LLMs to simulate complex attack scenarios, the research offers a promising pathway to overcome the persistent data scarcity problem in threat detection, enabling security teams to proactively test and strengthen their defenses without relying solely on real-world breach data or costly red team exercises.
Background
▶ Watch: The Detection Data Problem: Scarcity of Representative Data (0:50)
The landscape of threat detection is inherently complex and fraught with challenges. As Arjun Chakraborty highlighted, a primary obstacle for detection engineers is the consistent struggle to acquire representative data that accurately reflects real-world adversarial activities. This scarcity of relevant data directly impacts the efficacy of detection rules and machine learning models, necessitating continuous and often reactive tuning. The problem is exacerbated by the dynamic nature of cyber threats, the rapid expansion of enterprise attack surfaces, and the burgeoning influence of artificial intelligence in both offensive and defensive security paradigms.
Historically, synthetic data has found considerable success in various domains, including healthcare, medical research, and software testing, primarily for training machine learning models where real-world data is sensitive, scarce, or difficult to obtain. However, its application in cybersecurity, particularly for generating adversarial scenarios, has lacked a standardized approach and robust evaluation benchmarks. The speaker's hypothesis posits that LLMs, given a detailed description of a dataset and its associated columns, possess the capability to generate synthetic data that accurately mimics adversarial conditions. This capability, if proven effective and reliable, could significantly alleviate the data acquisition burden on detection engineers, allowing them to build and test detection pipelines more proactively and efficiently.
The motivation behind this research stems from the recognition that while LLMs have made data generation remarkably easy, as evidenced by their widespread capabilities, the critical challenge lies in evaluating the quality of the generated synthetic data. Without a reliable method to assess whether the synthetic data truly represents adversarial scenarios and can effectively contribute to detection, its utility remains questionable. Therefore, the talk aimed to not only explore the generation capabilities of LLMs but, more crucially, to establish a framework for evaluating the fitness of this "faux data" for "real defense."
Key Findings
▶ Watch: Core Research Questions: Generation and Quality Evaluation of Adversarial Syn... (2:00)
The research presented by Arjun Chakraborty yielded several key findings regarding the generation and evaluation of adversarial synthetic data using LLMs. A central discovery was the current lack of a universally accepted benchmark for evaluating the quality of synthetic data in cybersecurity, necessitating the development of novel metrics. Chakraborty proposed and utilized three primary metrics: Fidelity, Reproducibility, and Accuracy.
Fidelity measured how well LLMs adhered to instructions, specifically in generating all required columns and populating them with meaningful data. Initially, this was quantified by the percentage of null or "none" values within the generated samples, with lower percentages indicating better fidelity. For instance, Llama 70B initially appeared to perform best, exhibiting single-digit null values. However, a deeper qualitative analysis revealed a critical limitation: Llama 2, despite its low null count, often deviated from the requested schema, missing specified columns or altering the data structure based on its internal knowledge of Kubernetes audit logs. In contrast, Mixtral, while sometimes having a slightly higher null percentage, consistently followed the explicit instructions, generating the columns and data as requested. This finding underscored that purely quantitative metrics for fidelity might be misleading and that a qualitative review of the generated schema and content is indispensable.
Reproducibility assessed the consistency of critical columns across multiple synthetic data samples generated for the same attack technique. The goal was to ensure that while diversity in data was desirable for robust detections, core indicators of an attack (e.g., resource, verb, request object kind) remained consistent. Mixtral demonstrated superior performance in reproducibility, consistently maintaining these critical column values across almost all evaluated attack sequences. GPT 3.5 and its variations also performed decently. Llama 2, however, struggled significantly, often failing to generate the required columns at all, rendering reproducibility evaluation impossible for certain scenarios. This metric highlighted the models' ability to consistently simulate specific attack characteristics.
Accuracy was the ultimate measure, evaluating whether the synthetic data could effectively trigger standard detection rules. This metric directly addressed the end goal: producing adversarial data to enhance detection pipelines. The results indicated that all models performed well on simpler attack techniques, such as "listing out all secrets" or "escalating through node proxy permissions," where detection rules are often straightforward, focusing on specific verbs or resources. For these simpler cases, detection rates were very high. However, models generally struggled more with complex attacks like "running a privileged pod" or "stealing a pod service account token," which likely require more nuanced data generation to accurately reflect the attack's intricacies. GPT 3.5 and its variations showed strong performance, particularly in the "escalation through node proxy permissions" scenario.
Overall, a crucial finding was that no single LLM model emerged as universally superior. Performance varied significantly depending on the specific attack technique being simulated, the complexity of the data required, and the evaluation metric applied. For instance, Mixtral excelled in reproducibility, while GPT 3.5 showed strong accuracy in certain scenarios. This suggests that the choice of LLM and the specific evaluation metrics must be tailored to the particular use case and the desired characteristics of the synthetic data. The research firmly established that while LLMs offer a powerful capability for generating synthetic adversarial data, the emphasis must shift towards rigorous and context-aware evaluation of its quality.
Technical Deep Dive
▶ Watch: Novel Evaluation Framework: Fidelity, Reproducibility, and Accuracy Metrics (6:00)
The technical foundation of this research centered on generating adversarial synthetic data for Kubernetes API server audit logs. This specific data type was chosen for several strategic reasons: Kubernetes is widely adopted for containerized workloads, its API server audit logs are a primary source for detection teams, and the logs themselves possess characteristics conducive to evaluating synthetic data. These characteristics include a large number of columns, which tests data fidelity; a well-defined adversarial setup due to Kubernetes' maturity, aiding accuracy testing; the existence of established detection rules; and its critical role in the cyber kill chain, particularly for privilege escalation, persistence, and credential access.
The experimental setup involved testing four different LLM models: GPT 3.5 Turbo, GPT 3.5 Turbo 16k, Llama 2 70B, and Mixtral. The process began with crafting a detailed prompt for each model. This prompt comprised three main components: an initial description of the task, a comprehensive description of the required columns for the audit log, and a specific description of the attack technique to be simulated. The speaker noted that LangChain was utilized for interacting with the various LLM APIs, providing a flexible framework for calling different models. The prompting strategy employed was one-shot prompting, meaning the models generated data based solely on the provided prompt without prior fine-tuning or extensive examples, though the speaker acknowledged that more advanced techniques like Retrieval Augmented Generation (RAG) or fine-tuning could potentially improve results.
Four specific attack techniques were selected for simulation, drawing inspiration from the Stratus Red Team and aligning with the MITRE ATT&CK Matrix for Kubernetes. These techniques, chosen for their increasing difficulty, covered categories such as credential access, privilege escalation, and persistence. The rationale for selecting these attacks was the existence of well-defined detection rules, which would facilitate the subsequent accuracy evaluation of the synthetic data. The simulated attacks included:
- Listing all secrets (dumping secrets): A relatively straightforward credential access technique.
- Running a privileged pod: A more complex privilege escalation scenario.
- Stealing a pod service account token: Another credential access technique, often involving intricate steps.
- Escalating through node proxy permissions: A privilege escalation method leveraging Kubernetes node capabilities.
The generated synthetic data underwent evaluation using the three aforementioned metrics:
- Fidelity: This metric assessed the LLM's adherence to the specified schema and instructions. Quantitatively, it was measured by the percentage of null or "none" values within the generated columns. A lower percentage indicated better fidelity. However, as highlighted in the findings, a qualitative review was crucial. For instance, while Llama 2 70B often produced single-digit null percentages, it sometimes missed requested columns or altered the schema, demonstrating a need for human oversight beyond simple null checks. Mixtral, despite potentially higher nulls, generally followed the schema precisely.
- Reproducibility: This metric focused on the consistency of critical attack-identifying columns across multiple samples generated for the same attack technique. The goal was to ensure that essential fields like
resource,verb, andrequest object kindremained stable, even as other data points varied to introduce diversity. The speaker presented results showing Mixtral performing exceptionally well, with high percentages of consistent values across runs, indicating its ability to reliably generate core attack indicators. Llama 2, conversely, struggled significantly, often failing to generate these critical columns.
- Accuracy: This was the most critical metric, directly evaluating the utility of the synthetic data for detection. It involved feeding the generated data into standard detection rules that were known to identify the simulated attack techniques. The effectiveness of these rules on the synthetic data was then measured. For simpler attacks like "dumping all secrets" or "escalation through node proxy permissions," all models achieved high detection rates. More complex attacks, such as "running a privileged pod" or "stealing a pod service account token," presented greater challenges, with models showing more variability in accuracy. GPT 3.5 and its 16k variant generally performed well across various accuracy tests.
An important technical consideration was the use of an output parser via LangChain. This was essential for handling the common issue of LLMs generating syntactically invalid YAML or JSON, ensuring that the output could be consistently processed and evaluated. The speaker confirmed that while these parsers help "force" the LLM into the desired format, challenges in consistent formatting still arise.
Looking ahead, the speaker identified several areas for technical advancement. These include refining the evaluation framework, improving the quality of generated data through fine-tuning or RAG techniques, and exploring the impact of different prompting strategies. The evaluation of a broader range of models, including newer ones like GPT-4, Claude, and DBRX, was also suggested, potentially even using powerful LLMs as "judges" to evaluate the output of other models. The research underscores that while LLMs are powerful generation tools, the science of evaluating their output for specific, high-stakes applications like cybersecurity detection is still in its nascent stages and requires continuous innovation.
Demo / Proof of Concept
▶ Watch: Fidelity and Reproducibility Results and Challenges (9:00)
While the presentation did not feature a live, interactive demonstration of a tool or system, the entire experimental setup described by Arjun Chakraborty served as a comprehensive proof of concept for the methodology. The speaker detailed how various LLMs were prompted to generate synthetic Kubernetes API server audit logs simulating specific adversarial techniques. This involved defining the input prompts, specifying the target data schema, and outlining the four distinct attack techniques.
The subsequent evaluation of the generated data against the proposed metrics—Fidelity, Reproducibility, and Accuracy—effectively demonstrated the feasibility and challenges of using LLMs for this purpose. The presentation of quantitative results, such as null percentages for fidelity, consistency rates for reproducibility, and detection rates for accuracy across different models and attack types, provided concrete evidence of the approach's potential. For instance, the discussion around Mixtral's strong performance in reproducibility and GPT 3.5's decent accuracy for certain attacks, alongside Llama 2's struggles, illustrated the practical outcomes of this proof of concept. The use of LangChain for API calls and output parsing further solidified the technical viability of integrating LLMs into a data generation pipeline. Thus, while not a traditional software demo, the talk meticulously laid out and validated a robust experimental framework for leveraging LLMs in synthetic data generation for cybersecurity.
Defensive Implications
▶ Watch: Limitations and Future Work: Synthetic Data vs. Red Teaming, Benchmarks, and ... (18:00)
The research into using LLMs for adversarial synthetic data generation carries significant defensive implications for cybersecurity teams, particularly for detection engineers. The primary benefit lies in addressing the long-standing challenge of data scarcity for building and testing robust detection mechanisms.
- Enhanced Detection Pipeline Development: Detection engineers can leverage LLM-generated synthetic data to create more comprehensive and resilient detection rules and machine learning models. By simulating a wider array of adversarial scenarios, even those not yet observed in real-world incidents, defenders can proactively harden their systems. This moves beyond reactive tuning, allowing for more efficient and forward-looking development.
- Improved Purple Teaming Capabilities: The synthetic data can be a powerful asset for purple teaming exercises. Security teams can generate specific attack patterns and feed them into their existing detection systems to validate their efficacy. This allows for continuous testing and refinement of defenses without the need for full-scale, resource-intensive red team engagements or the risks associated with live adversarial simulations in production environments. The speaker explicitly stated that while it aids purple teaming, it cannot fully replace red teaming or sandboxing.
- Cost-Effectiveness: While LLM APIs can incur costs, the speaker suggested that generating synthetic data might be more cost-effective than maintaining a dedicated red team for certain data generation tasks. This could democratize access to adversarial data, enabling smaller teams or organizations with limited budgets to improve their defensive posture.
- Focus on Evaluation: A critical defensive implication is the absolute necessity of rigorous evaluation of any generated synthetic data. Defenders must adopt and adapt metrics like Fidelity, Reproducibility, and Accuracy, tailoring them to their specific operational context and threat models. Simply generating data is insufficient; understanding its quality and representativeness is paramount to avoid building detections on flawed or irrelevant "faux data."
- Understanding Model Limitations: The findings highlight that no single LLM is a silver bullet. Defenders need to understand the strengths and weaknesses of different models for various attack types and data generation tasks. This informed choice, coupled with continuous evaluation, is crucial for maximizing the utility of synthetic data.
- Proactive Threat Intelligence: By simulating known MITRE ATT&CK techniques and even hypothetical future attack patterns, security teams can proactively develop detections, staying ahead of adversaries rather than constantly playing catch-up. This fosters a more proactive and intelligence-driven defensive strategy.
In essence, LLM-driven synthetic data generation offers a powerful new tool in the defender's arsenal. It empowers detection engineers to overcome data limitations, build stronger defenses, and conduct more effective purple teaming, all while emphasizing the critical importance of robust evaluation to ensure the "faux data" genuinely contributes to "real defense."
Key Takeaways
- Data Scarcity is a Core Problem: Threat detection engineers consistently struggle to obtain representative adversarial data, leading to detection rules and models that require constant, reactive tuning.
- LLMs Offer a Solution for Data Generation: Large Language Models (LLMs) can efficiently generate synthetic adversarial data, specifically demonstrated with Kubernetes API server audit logs, to help detection engineers build more robust pipelines.
- Evaluation is Paramount: While LLMs make data generation easier, the critical challenge lies in rigorously evaluating the quality of the synthetic data. A multi-faceted approach using metrics like Fidelity, Reproducibility, and Accuracy is essential.
- No Universal "Best" Model: The performance of LLMs for synthetic data generation varies significantly based on the specific attack technique, data complexity, and evaluation metric. Defenders must choose and evaluate models based on their specific use case.
- Synthetic Data Augments, Not Replaces: LLM-generated synthetic data is a powerful tool for purple teaming and enhancing detection capabilities, but it currently cannot fully replace traditional red teaming or sandboxing of adversarial environments.
- Evolving Field Requires Custom Benchmarks: The field of synthetic data evaluation in cybersecurity is still nascent, lacking established benchmarks. Organizations will likely need to develop and refine their own evaluation frameworks tailored to their unique threat landscape and operational context.
About the Speaker(s)
Arjun Chakraborty is a dedicated professional working within the detection engineering team at Databricks. His professional interests primarily lie at the intersection of security and artificial intelligence (AI), a field in which he has accumulated considerable experience. Prior to his role at Databricks, Arjun held positions at Nvidia, Guidewire, and Home Depot, where he was involved in similar work related to security and AI applications. His expertise focuses on leveraging advanced technologies, such as machine learning and large language models, to enhance threat detection and defensive capabilities.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk presents a practical approach to a persistent problem in threat detection: the scarcity of representative adversarial data. By leveraging LLMs to generate synthetic Kubernetes audit logs and, crucially, developing a three-pronged evaluation framework (Fidelity, Reproducibility, Accuracy), the speaker offers a valuable tool for detection engineers. While not a replacement for red teaming, it provides a significant step towards more robust detection pipelines.
Heather Calloway (CISO) — STRONG ACCEPT
This presentation offers a pragmatic solution to a critical operational challenge: the scarcity of high-quality, representative adversarial data needed to build effective threat detection systems. By demonstrating how LLMs can generate synthetic Kubernetes audit logs and, more importantly, providing a robust framework for evaluating the fidelity, reproducibility, and accuracy of this data, the speaker provides a valuable tool for security teams to enhance their defensive capabilities and improve overall organizational resilience.