TL;DR: Applying AI to Security

Clint Gibler

BSidesSF 2024 · Day 1

Overview

Clint Gibler, Head of Security Research at Semgrep, delivered a comprehensive and fast-paced talk titled "TL;DR: Applying AI to Security" at BSidesSF 2024. The presentation aimed to provide both a high-level understanding and a wealth of tactical examples for leveraging Artificial Intelligence, specifically Large Language Models (LLMs), to address various cybersecurity challenges. Gibler, who previously expressed skepticism about AI in security in a 2020 blog post, acknowledged the significant advancements in the field and presented a revised perspective, emphasizing practical applications rather than hype.

Watch on YouTube

Visual summary for TL;DR: Applying AI to Security by Clint Gibler
Visual summary for TL;DR: Applying AI to Security by Clint Gibler

Key moments

  1. 01:00 LLM Fundamentals: Hallucinations, RAG, Agents, and Tools
  2. 06:00 Applying LLMs to Static Analysis: Triage, Vulnerability Finding, and Challenges
  3. 10:00 Critical Analysis of LLM Autonomous Exploitation Claims
  4. 13:00 LLMs as Penetration Testing Agents with Security Tools
  5. 15:00 Automated Threat Modeling and Security Design Review Bots
  6. 22:00 Blue Team Applications: IR Scenarios, TTP Extraction, SOC Bots, Malware Analysis
  7. 28:00 LLMs for Fuzzing: Coverage-Guided Test Harnesses and Automated Bug Discovery
  8. 33:00 Security Domain Specific Languages: English to SQL, Terraform, Rules, and Queries

TL;DR: Applying AI to Security

Speakers: Clint Gibler

Conference: BSidesSF 2024

YouTube: https://www.youtube.com/watch?v=7vB_7YlStyY

Overview

Clint Gibler, Head of Security Research at Semgrep, delivered a comprehensive and fast-paced talk titled "TL;DR: Applying AI to Security" at BSidesSF 2024. The presentation aimed to provide both a high-level understanding and a wealth of tactical examples for leveraging Artificial Intelligence, specifically Large Language Models (LLMs), to address various cybersecurity challenges. Gibler, who previously expressed skepticism about AI in security in a 2020 blog post, acknowledged the significant advancements in the field and presented a revised perspective, emphasizing practical applications rather than hype.

The core objective of the talk was to equip attendees with actionable insights, enabling them to apply AI concepts directly in their professional roles or inspire adjacent ideas within their organizations. Gibler meticulously covered a broad spectrum of security domains, including application security (AppSec), penetration testing, threat modeling, blue team operations, fuzzing, and reverse engineering, illustrating how LLMs can augment human capabilities and automate tasks. The talk deliberately focused on applying AI to security problems, rather than the security of AI itself, a topic Gibler noted was covered in other conference tracks.

This article distills Gibler's extensive overview, providing a structured exploration of the current landscape, key findings, and technical deep dives into the myriad ways AI is being integrated into cybersecurity workflows. It highlights both the promising capabilities and the inherent challenges, offering a balanced perspective on the technology's present and future impact on the defensive security posture.

Background

▶ Watch: LLM Fundamentals: Hallucinations, RAG, Agents, and Tools (01:00)

The landscape of Artificial Intelligence, particularly with the advent of advanced Large Language Models (LLMs), has undergone a dramatic transformation in recent years. Clint Gibler began his talk by referencing a blog post he co-authored in 2020, where he expressed a view that "AI in security is overhyped." This historical context served as a powerful backdrop for the current presentation, underscoring the rapid evolution of AI capabilities that have since compelled a re-evaluation of its potential in cybersecurity.

The problem Gibler aimed to address is the sheer volume of information and rapid pace of innovation in the AI space, making it challenging for security professionals to grasp the big picture and identify practical applications. His goal was to consolidate diverse examples and insights into a single, easily referenceable resource.

To ensure a common understanding, Gibler provided a "painfully oversimplified" lightning introduction to LLMs. He described them fundamentally as next-word automaters, trained on vast amounts of text data (e.g., "every on the internet, textbooks") using significant computational resources like GPUs. Key players in the LLM space were identified, including OpenAI (GPT-4), Anthropic (Claude 3), Google (Gemini), Meta (Llama 3), and Mistral.

Several foundational LLM concepts were introduced:

  • Hallucination: When a model confidently states something untrue.
  • Libraries: Tools like LangChain or LlamaIndex for building LLM applications.
  • Prompt: A structured input to an LLM, typically comprising a Persona (e.g., "you are a security expert"), a Task (e.g., "review this code"), and an Input (e.g., the code itself).
  • Few-shot prompting: Providing the LLM with a few examples of input-output pairs to guide its responses, often surprisingly effective without complex fine-tuning.
  • Context Window: The maximum amount of information an LLM can process at one time, akin to its working memory.
  • Retrieval Augmented Generation (RAG): A technique to overcome context window limitations by indexing external documents (PDFs, transcripts, internal docs) in a vector database and programmatically pulling relevant context for the LLM.
  • Agent: A system that enables multi-step analysis, complex planning, or collaboration between multiple LLMs, moving beyond a single prompt-response interaction.
  • Tools: Capabilities that allow an LLM to take action, such as performing a Google search, making API requests, running code, or interacting with security tools.

This primer established the necessary vocabulary and conceptual framework for understanding the diverse security use cases that formed the bulk of Gibler's presentation, setting the stage for a deep dive into how these AI capabilities are being applied to real-world cybersecurity problems.

Key Findings

▶ Watch: Critical Analysis of LLM Autonomous Exploitation Claims (10:00)

Clint Gibler's talk revealed several key findings regarding the application of AI to cybersecurity, emphasizing both its current capabilities and inherent limitations.

  1. AI as an Augmentation, Not a Replacement (Yet): A recurring theme was that LLMs are currently most effective as assistants that augment human capabilities rather than fully autonomous replacements. While they can achieve "80% of the way there" in many tasks, the "last 20%" often requires significant human expertise, iteration, and oversight. This was highlighted in areas like exploit generation, fuzzing, and security design review.
  1. Significant Potential for Operational Efficiency: AI demonstrates strong promise in reducing operational toil for security teams. Examples included automating bug bounty triage, answering common security questions, classifying security design reviews, and providing just-in-time education to developers. This frees up security professionals for higher-leverage activities.
  1. Challenges in Direct Vulnerability Detection: Despite initial hype, rigorous academic analysis (e.g., the paper on source code vulnerability detection) suggests that LLMs, even advanced ones like GPT-4, perform akin to "random guessing" when attempting to identify vulnerabilities from first principles in real-world codebases. This points to significant data quality and evaluation methodology issues in current research.
  1. Effectiveness with Specific, Well-Defined Tasks: LLMs excel when given clear tasks, well-structured prompts, and access to relevant tools or context. Examples include generating fuzzing harnesses for small codebases (MOX's GIF fuzzer), translating natural language into security DSLs (Semgrep rules, Nuclei templates, KQL), and summarizing threat intelligence.
  1. The Importance of Prompt Engineering and Context Management: The quality of the prompt significantly impacts the output. Techniques like few-shot prompting and Retrieval Augmented Generation (RAG) are crucial for providing the necessary context and guidance to LLMs, especially when dealing with large codebases or extensive documentation.
  1. Emergence of Agent-Based Systems: The concept of agents equipped with tools (including security tools like Nmap, ZAP, Nuclei) is a promising direction for more complex, multi-step security tasks like penetration testing or automated incident response. This allows LLMs to interact with the environment and iteratively refine their actions.
  1. Data Quality and Evaluation are Critical: The reliability of AI applications in security is heavily dependent on the quality of training and evaluation data. Many existing datasets for vulnerability detection were found to have significant shortcomings, leading to unrepresentative results.
  1. Long-Term Optimism for Defense: Gibler concluded with a reflection that, despite short-term advantages for attackers, defense is likely to "win in the end." This is attributed to continuous improvements in security frameworks, tools, and practices (e.g., modern web frameworks, TLS, U2F), which collectively raise the bar for attackers.

These findings collectively paint a picture of AI as a powerful, rapidly evolving tool that, when applied thoughtfully and with an understanding of its limitations, can significantly enhance cybersecurity capabilities.

Technical Deep Dive

▶ Watch: Automated Threat Modeling and Security Design Review Bots (15:00)

Clint Gibler's talk provided a dense technical overview of LLM applications across numerous cybersecurity domains. This section elaborates on the specific technical approaches and tools discussed.

LLM Fundamentals and Architecture

At its core, an LLM is a next-word automater, trained on massive datasets using GPUs. Key architectural components for building LLM applications include:

  • LangChain and LlamaIndex: Popular libraries for orchestrating LLM workflows.
  • Prompts: Structured inputs defining a persona, task, and input. Few-shot prompting involves providing examples to improve output quality.
  • Context Window: The limited "working memory" of an LLM.
  • Retrieval Augmented Generation (RAG): To overcome context limitations, external data (PDFs, internal docs) is indexed in a vector database. RAG allows the LLM to programmatically retrieve and incorporate relevant context before generating a response.
  • Agents: Systems that enable multi-step reasoning, planning, and execution, often by chaining multiple LLM calls or coordinating multiple agents.
  • Tools: Functions or APIs that allow an LLM agent to interact with the external world (e.g., Google search, API requests, running code, using security tools).

Application Security (AppSec)

Gibler outlined two primary AppSec use cases:

  1. Triaging Static Analysis Findings: LLMs can predict if a tool's flag on a piece of code is a true or false positive and suggest code snippets for remediation.
  • Examples: Work by Gibler's colleague 'B', Ox GPT, and a blog post by GitHub Security Lab emphasizing the importance of pre- and post-processing for reliable outputs.
  1. Direct Vulnerability Detection: Using LLMs to directly identify vulnerable code given a prompt and the code.
  • Methodology: Trail of Bits provided a thorough methodology.
  • Challenges: Prompt quality is critical and hard to optimize across diverse codebases. Cost can be prohibitive (e.g., $100 for 200,000 lines of code). Context window limitations make it difficult to analyze inter-file dependencies.
  • Research Findings: A rigorous paper analyzing various datasets concluded that many evaluation methods were unrepresentative. It found that even advanced models like GPT-4 performed "akin to random guessing" in identifying real vulnerabilities, highlighting the need for better data and evaluation.
  • Practical Applications: Lou Barrett's talk at Lascon demonstrated using small, local models for automated PR comments, suggesting security considerations.
  • Meta-Security: ChatGPT itself was used to analyze the security properties of its own plugins, a "yo dog" moment of AI analyzing AI security.

Exploit Generation

LLMs show promise in automating parts of the exploit development lifecycle:

  • CVE to Exploit: Given a CVE report, an LLM can extract the diff between vulnerable and fixed versions. It can then identify the most likely vulnerable code, generate a vulnerable Go program exhibiting the flaw, and even produce a Python Proof of Concept (PoC). The speaker noted this often requires "cajoling" the LLM.
  • Autonomous Exploitation Claims: A paper claimed LLM agents could autonomously exploit one-day vulnerabilities with 87% success. However, a rebuttal by Chris Rol and a CVE co-founder pointed out methodological flaws: the study used only 15 CVEs, and each had an easily googleable PoC, meaning the LLM was more likely retrieving existing exploits than creating them from first principles.

Penetration Testing

The integration of LLMs with security tools enables automated penetration testing:

  • Agents with Tools: LLM agents can be given access to security tools like Nmap, ZAP, Nuclei, SQLMap, Burp, and more.
  • Workflow: As described by Matt Adams, an LLM agent can receive a task (e.g., "scan this web app"), select appropriate tools, execute them, store results, and iteratively decide on next steps, potentially culminating in a comprehensive report with an executive summary.
  • Improvements: Gibler suggested giving agents access to more tools, using multiple specialized agents (e.g., one for attack surface, one for SQL injection) collaborating via a supervisor agent, and a shared results storage.
  • Future Outlook: Joseph Thacker (reso) predicted AI hack bots would outperform humans. Gibler estimated that within 2-3 years, AI systems could match the capabilities of an entry-level pen tester.

Threat Modeling

LLMs can streamline and enhance threat modeling processes:

  • Automated Threat Models: Matt Adams' St GPT tool takes an application description (internet-facing, sensitive data, app type) and outputs a threat model, attack tree, and mitigations.
  • Collaborative Threat Modeling: John created a custom GPT by uploading cloud-native risks and best practices, allowing for interactive, clarifying conversations.
  • Automated PR-based Threat Modeling: A GitHub action by Maren analyzes a repo's README and other markdown files to automatically create a PR suggesting risks, updating with new code pushes.
  • Data Gathering Automation: Gibler proposed automatically gathering application details (design docs, PRDs, code analysis for data access/app type, infrastructure-as-code for internet exposure) to pre-populate threat models, presenting a "mostly right" draft for human review.
  • Architecture Diagram Generation: LLMs can auto-generate architecture diagrams from descriptions, saving time by providing a starting point.

Security Design Review

LLMs can assist in determining the security review needs of new features or projects:

  • Risk Scoring: Lou Barrett discussed combining a developer's PRD/scope doc with the security team's defined policies (e.g., "are auth/auth changes made?", "is sensitive data accessed?") to generate a risk score, indicating if and when security review is needed.
  • OpenAI's Internal Use:
  • SDLC Bot (slackbot): Developers provide project details, and the bot determines if a security review is required.
  • Tribot: Identifies potentially sensitive actions (e.g., confidential Google Doc made public) and automatically messages the user for verification.
  • Slack Channel Triage: Automatically directs security questions to the right team or answers common questions using RAG.
  • Bug Bounty Triage: Classifies submissions as in-scope, out-of-scope, or functional bugs.
  • Access Management Troubleshooting: Helps users understand why they're blocked and what permissions they need.
  • Just-in-Time Education: As developers propose new services, the bot can suggest secure-by-default libraries or infrastructure components.
  • OpenAI has open-sourced three of their slackbots.

Blue Team Operations

AI offers significant advantages for defensive security:

  • Incident Response Scenarios: Matt Adams' tool generates tailored incident response scenarios based on industry and company size, leveraging knowledge of attack groups.
  • TTP Extraction: MITRE's TRAM tool extracts techniques, tactics, and procedures (TTPs) from threat intelligence reports or PDFs.
  • Threat Intelligence Querying: Thomas loaded MITRE ATT&CK group documentation into RAG, allowing English queries about specific threat actors (e.g., "What is the Lazarus group?").
  • Agent-Driven Threat Analysis: Thomas exposed functions from Microsoft's msticpy library (e.g., VirusTotal lookup, malware sample analysis) as tools to an LLM agent, enabling it to perform iterative threat investigations. This "sock analyst as a service" is a hot startup area.
  • Visualizing Threat Intel: Thomas released a Jupiter notebook to convert text-based threat intel into Mermaid JS diagrams.
  • Video Analysis: John Hammond demonstrated converting a YouTube video transcript (e.g., an Apex Legends hack explanation) into a Mermaid JS diagram using a Fabric pattern.
  • Malware/Script Explanation: Microsoft's Security Copilot and Google's Gemini Pro can explain malicious scripts or malware samples, identifying activities (clipboard monitoring) and IOCs.
  • IOC Extraction from Images: Multimodal LLMs can extract IP addresses, domains, URLs, and other IOCs from images within threat intelligence reports.
  • Dark Web/Discord Summarization: LLMs can summarize and translate discussions from dark web marketplaces or cybercriminal channels, similar to how they summarize meetings or podcasts.

Unit Test Generation

LLMs are being used to automatically write unit tests:

  • Research & Tools: Papers from Meta and open-source tools like Test Pilot are exploring this space.

Fuzzing

AI can significantly enhance fuzzing efforts:

  • Coverage-Guided Fuzzing: Google integrated LLMs into OSS-Fuzz. By analyzing coverage reports, LLMs generate fuzzing test harnesses for unexercised code paths. CI Spark is a commercial version.
  • Kernel Fuzzing: A blog post detailed using Szallar with an interactive LLM to fuzz kernel code. The LLM could perform internet searches to understand ioctl calls and generate harnesses, reducing expertise and time requirements.
  • Automated Library Fuzzing: Kudelski Security developed a methodology to automatically fuzz Rust parser libraries. They identified how to call libraries by analyzing READMEs, examples, unit tests, and static analysis. They found bugs in over a third of projects.
  • Zero-Shot Fuzzing: MOX demonstrated a remarkable example: Claude 3 generated a Python function to fuzz a small C GIF decoding library, achieving 92% line coverage and finding four memory safety bugs and one hang with zero-shot prompting. This matched the performance of a human-written fuzzer in one hour.

Reverse Engineering

LLMs assist in understanding and analyzing binaries:

  • Code Explanation & Bug Finding: Most tools in this area integrate with IDA Pro, Ghidra, or Binary Ninja to help explain decompiled code, assist in the decompilation process, or find bugs. These are largely prototypes.

Security Domain Specific Language (DSL) Generation

LLMs can translate natural language into specialized security queries or configurations:

  • SQL Generation: Tools can take a database schema and an English question to auto-generate SQL queries. Pinterest shared a blog post detailing their productionization of such a system, including architecture, prompts, and edge cases.
  • Terraform Generation: Salami uses GPT-4 to convert natural language into Terraform.
  • Semgrep Rule Generation: Semgrep Assistant generates Semgrep rules from English descriptions, code examples (positive and negative matches).
  • Nuclei Template Generation: Nuclei has a browser extension to generate templates from bug bounty reports or highlighted text on any website. Orca also showed an online editor with English-to-rule functionality.
  • SIM Query Generation: Microsoft's Security Copilot and Google provide examples of generating KQL or other SIM queries from English.
  • Impact: This capability significantly lowers the learning curve for powerful security tools, making them accessible to less experienced users and shrinking the "time to value."

Demo / Proof of Concept

▶ Watch: Blue Team Applications: IR Scenarios, TTP Extraction, SOC Bots, Malware Analysis (22:00)

Clint Gibler's talk was rich with references to various tools, projects, and demonstrations that serve as proofs of concept for applying AI to security. While a live demo wasn't performed during the talk, Gibler extensively cited existing work and open-source projects.

Key examples of demonstrated or referenced proofs of concept include:

  • Ox GPT: An example of an LLM-based tool for predicting true/false positives in static analysis findings and suggesting code fixes.
  • GitHub Security Lab Blog Post: Emphasized the practical pre- and post-processing needed to make LLM outputs reliable for code analysis.
  • Trail of Bits Methodology: A detailed approach for using LLMs to find vulnerabilities directly in source code.
  • Lou Barrett's Local Model for PR Comments: A practical setup for using open-source models to provide automated security comments on pull requests.
  • St GPT by Matt Adams: A tool that takes an application description and automatically generates a threat model, attack tree, and mitigations.
  • John's Custom GPT for Collaborative Threat Modeling: A demonstration of how a custom GPT, pre-loaded with cloud-native risks and best practices, can facilitate interactive threat modeling discussions.
  • Maren's GitHub Action for Automated Threat Modeling: This action analyzes a repository's README and other markdown files to automatically create a pull request with suggested security risks, demonstrating programmatic updates based on code changes.
  • OpenAI's Internal Slackbots (SDLC Bot, Tribot, etc.): These bots, some of which are open-sourced, serve as real-world examples of AI-driven automation for security design reviews, sensitive data monitoring, internal support, and bug bounty triage.
  • MITRE TRAM: A tool that extracts TTPs from threat intelligence reports, showcasing AI's ability to process and structure unstructured text.
  • Thomas's Jupiter Notebooks: Demonstrated querying MITRE ATT&CK groups using RAG and visualizing threat intelligence reports as Mermaid JS diagrams.
  • Thomas's Agent with msticpy Tools: A proof of concept showing an LLM agent using external security library functions (like VirusTotal lookups) as tools to perform automated threat investigations.
  • John Hammond's Video to Mermaid JS Diagram: A demonstration of using a "Fabric pattern" to convert a YouTube video transcript into a visual diagram, illustrating multimodal AI application.
  • Google's OSS-Fuzz Integration: A practical application where LLMs generate fuzzing harnesses to improve code coverage in open-source projects.
  • Kudelski Security's Rust Parser Fuzzing Tool: An open-source tool demonstrating a systematic methodology for automatically fuzzing Rust libraries and finding bugs.
  • MOX's Claude 3 GIF Fuzzer: A compelling example where Claude 3, with a simple prompt, generated a Python fuzzer that achieved high code coverage and found multiple memory safety bugs in a C GIF decoding library. This was a direct, impactful demonstration of AI's capability in vulnerability discovery.
  • Pinterest's SQL Generation System: A detailed blog post outlining the architecture and challenges of productionizing an LLM-based system for generating SQL queries from natural language.
  • Semgrep Assistant: A tool that generates Semgrep rules from natural language descriptions and code examples.
  • Nuclei Browser Extension and Orca's Web Editor: Examples of tools that allow users to generate Nuclei templates from bug bounty reports or highlighted text, simplifying the creation of vulnerability scanning rules.
  • Microsoft Security Copilot & Google's SIM Rule Generation: Demonstrations of converting English queries into KQL or other Security Information and Event Management (SIEM) rules.

These examples collectively illustrate the breadth of AI applications in security, ranging from research prototypes to production-grade systems, and highlight the practical utility of LLMs when integrated with existing security workflows and tools.

Defensive Implications

▶ Watch: Security Domain Specific Languages: English to SQL, Terraform, Rules, and Que... (33:00)

The insights from Clint Gibler's talk offer several critical defensive implications for security teams looking to leverage AI effectively. The overarching message is to view AI as a powerful assistant and augmenter of human capabilities, rather than a direct replacement, at least in the short term.

  1. Automate Operational Toil: Security teams should actively identify and target repetitive, low-leverage tasks for AI automation. This includes initial bug bounty triage, answering frequently asked questions in security channels, classifying security design reviews, and basic access management troubleshooting. By offloading these tasks, human security professionals can focus on more complex, strategic, and higher-leverage activities.
  1. Enhance Developer Experience and "Shift Left": Implement AI-driven tools that provide just-in-time security education and guide developers towards secure-by-default libraries and infrastructure. This can be achieved through bots integrated into development workflows (e.g., Slackbots, GitHub actions) that suggest secure patterns or flag potential risks early in the development lifecycle, empowering developers to build securely from the start.
  1. Augment Human Expertise in Complex Tasks: For areas like penetration testing, threat modeling, and incident response, AI can act as a force multiplier. LLM agents equipped with security tools can perform initial reconnaissance, generate attack trees, or suggest incident response scenarios, providing a strong starting point for human experts. The goal is to make security professionals "easier, faster, or more cheaply" effective, as Gibler noted.
  1. Prioritize Data Quality and Evaluation: When adopting or building AI solutions for security, defenders must be acutely aware of the importance of high-quality data for training and rigorous, representative evaluation methodologies. Overhyped claims, especially in areas like direct vulnerability detection, should be scrutinized. Focus on practical, measurable improvements rather than aspirational, unproven capabilities.
  1. Leverage AI for Knowledge Management and Threat Intelligence: Utilize LLMs with RAG to make vast amounts of security documentation (MITRE ATT&CK, internal policies, threat intelligence reports) easily queryable and digestible. This improves situational awareness and enables faster decision-making during incidents or threat hunting. Tools that summarize and visualize complex threat data are particularly valuable.
  1. Lower the Barrier to Entry for Security Tools: Implement AI interfaces (e.g., natural language to DSL generation) for powerful but complex security tools (Semgrep, Nuclei, SIEMs). This democratizes access to advanced capabilities, allowing less experienced team members or even developers to utilize them effectively, thereby expanding the reach of security practices.
  1. Embrace Iterative Development and Human-in-the-Loop: Recognize that AI outputs may not be perfect. Design workflows where AI provides an "80% solution" that humans can easily review, edit, and refine. This "editing is easier than writing from scratch" approach maximizes efficiency while maintaining accuracy and accountability.
  1. Long-Term Optimism for Defense: Gibler's concluding reflection that "defense is going to win in the end" provides a strategic perspective. Continuous advancements in secure frameworks, default-secure configurations, and user-friendly security mechanisms (like U2F) are steadily raising the bar for attackers. AI, when applied thoughtfully, contributes significantly to this long-term defensive advantage by making security more scalable, efficient, and proactive. Defenders should invest in understanding and integrating AI, confident that it will strengthen their overall posture.

Key Takeaways

  • AI as an Assistant, Not a Replacement (Yet): LLMs are powerful tools for augmenting human security professionals, automating mundane tasks, and providing "80% solutions" that still require human oversight and refinement for the critical "last 20%."
  • Significant Efficiency Gains Across Security Domains: AI can drastically reduce operational toil in AppSec, pen testing, threat modeling, blue team operations, fuzzing, and reverse engineering, freeing up security teams for higher-leverage activities.
  • Prompt Engineering and Context are Paramount: The effectiveness of LLM applications heavily relies on well-crafted prompts, few-shot examples, and techniques like Retrieval Augmented Generation (RAG) to provide necessary context and mitigate hallucinations.
  • Rigorous Evaluation is Crucial: Claims of AI's capabilities, especially in direct vulnerability detection or autonomous exploitation, must be critically assessed against robust data sets and methodologies, as many current evaluations may not reflect real-world performance.
  • Lowering the Barrier to Security Tools: AI-driven natural language interfaces can make complex security tools (e.g., Semgrep, Nuclei, SIEMs) more accessible to a broader audience, improving overall security posture by enabling more users to leverage advanced capabilities.
  • Defense is Winning Long-Term: Despite short-term advantages for attackers, the continuous evolution of secure-by-default frameworks, tools, and practices, significantly aided by AI, suggests a long-term trend where defensive capabilities will continue to strengthen.

About the Speaker(s)

Clint Gibler is the Head of Security Research at Semgrep, a company known for its popular open-source static analysis tool and commercial offerings for first-party code, reachable vulnerable dependencies, and semantic secret detection. Prior to his role at Semgrep, Clint worked at NCC Group, a global security consulting firm, and was a graduate student at UC Davis. He is also the author of a widely read, free security newsletter that curates the best tools, talks, and blog posts in cybersecurity, delivered weekly. His background spans academic research, security consulting, and leading security innovation in product development.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

The talk provides a rapid-fire overview of current AI/LLM applications in offensive and defensive security. It's a dense, reference-style presentation covering a wide array of use cases from static analysis and exploitation to fuzzing and threat modeling. While it doesn't delve deeply into any single topic, it effectively curates and categorizes a significant amount of recent work, highlighting both promising advancements and current limitations, particularly regarding the reliability and cost of LLM-driven security tools.

Heather Calloway (CISO) — STRONG ACCEPT

This talk provides a comprehensive, rapid-fire overview of how AI and large language models are currently being applied across various cybersecurity domains. While highly technical in its examples, it offers valuable insights for security leaders by demonstrating the practical, albeit often nascent, capabilities of these technologies in areas like automated vulnerability detection, threat modeling, and incident response. It effectively cuts through the hype to show where real work is being done, which is crucial for strategic planning and resource allocation.

→ Top-rated talks at BSidesSF 2024

All talks from BSidesSF 2024