RAG Against the Machine: Using Retrieval-Augmented Generation & MCP to Fortify Cybersecurity Defense

Brennan Lodge (Director of Information Security · the Manhattan Institute)

BSides Las Vegas 2025 · Day 1

Overview

Brennan Lodge uses BSides Las Vegas as a venue to argue that retrieval-augmented generation (RAG) and the Model Context Protocol (MCP) are practical, mostly open-source building blocks for defensive workflows—not a replacement for analysts, but a way to reduce alert fatigue, accelerate threat-intel alignment, and drag GRC out of spreadsheet purgatory. The talk is structured as good / bad / ugly AI in security: opportunities (information overload, talent gaps), risks (shadow AI, opaque token costs), and cultural failure modes (Clippy redux). Lodge grounds claims in personal experimentation: a <$500 (historical) budget target, ~10 second response goals, ChromaDB + LangChain + sentence-transformers stacks, and two open projects—Arsenal Forge (MCP + RAG for SOC-style enrichment) and Open Audit Caddy (policy/compliance mapping with BERT-style classifiers). A recorded demo shows MITRE mapping for a Splunk detection, CISA advisory context, and a memory server logging queries for transparency.

Watch on YouTube

Visual summary for RAG Against the Machine: Using Retrieval-Augmented Generation & MCP to Fortify Cybersecurity Defense by Brennan Lodge
Visual summary for RAG Against the Machine: Using Retrieval-Augmented Generation & MCP to Fortify Cybersecurity Defense by Brennan Lodge

Key moments

  1. 4:00 Good/bad/ugly AI framing: overload vs talent gaps; shadow AI and token cost opacity; Clippy analogy for chatbot fatigue.
  2. 12:00 Library metaphor for RAG: embeddings as Dewey Decimal, vector DB as shelves, LLM as librarian with citations.
  3. 18:00 Roll-your-own RAG steps: ingest/chunk, embeddings, LangChain glue, Chroma storage, and vector visualization for debugging.
  4. 22:00 Cost/privacy benchmark narrative: VPC GPU hosting vs OpenAI API token pricing for classroom-scale analyst query load.
  5. 28:00 MCP explained as USB-C for AI: JSON-RPC host/client/server model connecting models to Slack, Gmail, calendars, local data.
  6. 32:00 Arsenal Forge demo setup: upload MITRE/CISA/detections to Chroma, start memory MCP server, launch Streamlit analyst UI.
  7. 34:00 MoveIT-themed example: Splunk rule summarized, MITRE technique mapped (RDP acronym resolved), detection inventory linked.
  8. 40:00 Open Audit Caddy pitch: policy templates, SOC2 mapping notebooks, BERT-style classification, auditor-friendly exports.

RAG Against the Machine: Using Retrieval-Augmented Generation & MCP to Fortify Cybersecurity Defense

Speakers: Brennan Lodge, fractional CISO; faculty (NYU — IT/management/analytics); employed at Manhattan Institute (as stated); prior financial-sector roles (JP Morgan, Federal Reserve Bank, Bloomberg, Goldman Sachs, HSBC cited)

Conference: BSides Las Vegas

YouTube: https://www.youtube.com/watch?v=7Go4KRNFZ_I

Overview

Brennan Lodge uses BSides Las Vegas as a venue to argue that retrieval-augmented generation (RAG) and the Model Context Protocol (MCP) are practical, mostly open-source building blocks for defensive workflows—not a replacement for analysts, but a way to reduce alert fatigue, accelerate threat-intel alignment, and drag GRC out of spreadsheet purgatory. The talk is structured as good / bad / ugly AI in security: opportunities (information overload, talent gaps), risks (shadow AI, opaque token costs), and cultural failure modes (Clippy redux). Lodge grounds claims in personal experimentation: a <$500 (historical) budget target, ~10 second response goals, ChromaDB + LangChain + sentence-transformers stacks, and two open projects—Arsenal Forge (MCP + RAG for SOC-style enrichment) and Open Audit Caddy (policy/compliance mapping with BERT-style classifiers). A recorded demo shows MITRE mapping for a Splunk detection, CISA advisory context, and a memory server logging queries for transparency.

The session is explicitly vendor-skeptical: push suppliers for token accounting, avoid “Salt Bae AI” sprinkles without use cases, and keep humans in the loop—especially for agentic automation, which Lodge largely warns against for production SOC actions in Q&A.

Background

▶ Watch: Good/bad/ugly AI framing: overload vs talent gaps; shadow AI and token cost o... (4:00)

Lodge’s biography spans ~18 years in financial services security/architecture, teaching at NYU, fractional CISO work, and research accolades including a Kaggle-style win at Oxford and early US Cyber Command collaboration on alert fatigue (2022, pre-ChatGPT hype). He markets LinkedIn Learning courses on RAG and MCP (August release mentioned for MCP).

The library metaphor frames RAG: a sentence embedding model is the Dewey Decimal analog; a vector database is the shelf; the LLM is the librarian assembling answers with citations. The value proposition is fresh private data and reduced hallucination versus static model cutoffs—provided retrieval quality is good.

Lodge also spends a beat on hardware pragmatism: early experiments included AMD GPUs and frustration with developer ergonomics versus NVIDIA’s Python/CUDA ecosystem—less a religious war than a warning that self-hosting costs include engineering hours, not just instance dollars. That matters when a CISO asks why a pilot ballooned from $500 CapEx to two sprints of ML engineer time.

Key Findings

▶ Watch: Roll-your-own RAG steps: ingest/chunk, embeddings, LangChain glue, Chroma sto... (18:00)

1) RAG is an architecture, not a model brand. Ingest → chunk → embed → store → retrieve → prompt with citations. Chunking strategy (sentence vs paragraph vs page) must match corpus scale—terabyte corpora need coarser chunks.

2) Open toolchain is viable. Lodge lists Hugging Face, Chroma, LangChain, and sentence-transformers (all-MiniLM-L6-v2 cited as a durable embedding choice in his work). He encourages vector visualization (mentions Nomic Atlas) to debug bad retrieval.

3) Cost and privacy tradeoffs are measurable. A classroom-scale experiment (~five analysts, ~20 queries/day, ~100k tokens/day narrative) compared VPC-hosted GPU (G4dn.4xlarge class) vs OpenAI API pricing “at the time”—~$500/month vs ~$100/month higher for cloud API, with data leaving the network in the latter case. Prices change; the pattern remains: gravity of data vs OPEX.

The speaker notes students (NYU “lab rats,” in his joking phrasing) exercised the system on an AWS EC2 sandbox he funded—useful transparency that results are pilot-scale, not Fortune-50 production proofs.

4) MCP standardizes tool wiring. Lodge describes MCP as a USB-C for AI: host, client, server, JSON-RPC 2.0 framing, connecting models to Slack, Gmail, calendars, and local data sources with less bespoke glue.

5) Arsenal Forge demonstrates SOC enrichment. Upload MITRE techniques, CISA advisories, detection inventory; MCP memory server logs interactions; Streamlit UI returns MITRE technique ID, links, and analyst-facing summaries (demo uses MoveIT-related content as an example).

6) GRC can be structured as data science. Lodge positions compliance notebooks that map policies to SOC 2 domains using classification models (BERT named), advocating smaller domain-tuned models for focused tasks rather than giant general chat models for control mapping.

7) Shadow AI is an enterprise incident waiting to happen. Read AI-style Teams worms illustrate how viral adoption bypasses security review; Lodge tells a browser extension + Teams integration horror story from his environment.

8) Classic security controls translate. Prompt injection parallels SQL injection mentally—sanitize inputs, guardrails, encrypt vector DB at rest/in transit, log violations to SOC, output policies.

9) GRC is a data volume problem disguised as paperwork. Lodge shows charts of rising AI governance rules alongside privacy and cybersecurity state laws (as presented on slides). The punchline is not “more laws,” but that humans cannot diff regulatory corpora against internal policies at scale without tooling—hence RAG + classification.

10) MCP enables standardized auditing surfaces. Because MCP traffic is structured (JSON-RPC), security teams have a cleaner hook for logging and policy enforcement than ad-hoc REST scripts embedded in each chatbot.

Technical Deep Dive

▶ Watch: MCP explained as USB-C for AI: JSON-RPC host/client/server model connecting m... (28:00)

RAG pipeline details

Lodge emphasizes ETL discipline: parse PDFs, emails, CSV exports (MITRE), normalize into a simple email-like schema (subject, body, source, timestamp) to simplify chunking and later UI presentation. Embeddings land in Chroma; queries retrieve top-k chunks; the LLM composes answers grounded in retrieved text.

Evaluation difficulty

He acknowledges LLM evaluation is subjective, naming HyDE (hypothetical document embeddings) as one technique to cross-check answers against retrieved corpora—still imperfect, but a step beyond vibes.

In Q&A he also points to RLHF-style thumbs up/down feedback loops as a downstream tuning layer—useful, but still labor-intensive and vulnerable to noisy raters, so it is not a substitute for retrieval quality and source hygiene.

MCP architecture

MCP separates tool hosts from tool servers. Arsenal Forge spins multiple servers—memory for audit logs plus domain data servers. The demo CLI uploads corpora, starts servers, and launches Streamlit.

Security of AI systems

Containerize, VPC deploy, consider on-prem LLMs for sensitive workloads. Lodge warns that tokens are a pricing black box across vendors—finance and security leaders should demand usage dashboards and budget caps.

He name-checks Deep Tempo as an example of background analytics on data lake telemetry—used as contrast to chatty assistants: sometimes the winning pattern is silent scoring with human-facing explanations only on anomalies.

Agentic caution

In Q&A, Lodge advises SOC managers to start with threat intel enrichment and MITRE labeling, avoid “easy button” autonomous response, and test phishing awareness-class automation before containment agents.

Organizational parallels

The Phoenix Project analogy frames AI adoption as another integration wave that will stress CISO bandwidth. The point is not literary; it is change management—without owners, RAG becomes shelfware; with owners, it becomes runbook acceleration.

GRC product narrative

Open Audit Caddy is positioned as templates + notebooks + exports that an auditor can navigate: domain scores, missing policies, and JSON exports for downstream systems. Lodge tells a light auditor anecdote to humanize the pain point—useful as social proof that the UI targets evidence, not just chat.

Demo / Proof of Concept

▶ Watch: Arsenal Forge demo setup: upload MITRE/CISA/detections to Chroma, start memor... (32:00)

Screen recording (shown in session): CLI scripts upload MITRE, CISA, detections to Chroma; start memory MCP server; launch Arsenal Forge UI. Example query on MoveIT-related detection returns narrative summary, MITRE mapping (T1046 scanning discussed), pointers into detection inventory, related intel, and memory log of the session.

Lodge highlights a subtle win in that demo: the detection text uses an acronym (RDP) without spelling out Remote Desktop Protocol, yet retrieval still surfaces the right MITRE technique—evidence that RAG can reduce toil when synonyms and abbreviations are rampant in SOC data. The caution is the opposite case: if retrieval is poisoned by a thin or stale MITRE export, the model will confidently cite the wrong technique—another argument for visualizing embeddings and maintaining versioned corpora when ATT&CK updates ship.

Defensive Implications

▶ Watch: Open Audit Caddy pitch: policy templates, SOC2 mapping notebooks, BERT-style ... (40:00)

SOC leaders can pilot RAG on internal runbooks + intel feeds before buying SOAR 2.0 marketing. Success metrics should be time-to-context and false-positive analyst comments, not “number of AI features.”

CISOs should publish AI acceptable use, block unsanctioned meeting bots, and require DLP review for copilots with email access.

GRC teams can use classification + retrieval to maintain evidence mappings, but legal interpretation remains human—AI suggests; humans sign.

Red teams should test MCP servers and RAG apps for tool poisoning, prompt injection, and over-permissioned connectors—especially where Slack bots can read channels with secrets.

Platform engineering should treat vector databases like databases: backups, schema migration, access controls, and secrets rotation for embedding endpoints—because they will hold internal runbooks and customer excerpts if poorly scoped.

Key Takeaways

  • RAG tethers LLMs to fresh, citable corpora—critical for MITRE, advisories, and internal runbooks.
  • MCP reduces bespoke integration spaghetti between models and enterprise tools.
  • Cost, privacy, and token transparency must be in the RFP.
  • Shadow AI spreads through viral SaaS; governance must be fast and explicit.
  • Arsenal Forge and Open Audit Caddy exemplify open approaches to SOC + GRC assistance.
  • Agentic automation in SOC deserves extreme caution; start with read-only enrichment.
  • Classical security engineering (input validation, logging, encryption) still applies.
  • Teach and tune: prompt hygiene, acronym glossaries, and RLHF-style feedback can improve outcomes but do not fix bad corpora.

About the Speaker(s)

Brennan Lodge describes living in Manhattan, teaching at NYU, fractional CISO work, and employment at Manhattan Institute. Prior employers named include JP Morgan, Federal Reserve Bank, Bloomberg, Goldman Sachs, and HSBC. He references LinkedIn Learning courses on RAG and MCP and open projects Arsenal Forge and Open Audit Caddy (GitHub promotions in-session; URLs not verified here). Bundle metadata lists Speakers: Unknown.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

A competent practitioner’s map of RAG+MCP for SOC and GRC enablement, with a solid demo and sane cautions—but it is architecture advocacy more than novel security research.

Heather Calloway (CISO) — STRONG ACCEPT

This is governance-ready AI material: it couples technical architecture with procurement demands (token transparency), shadow-AI containment, and evidence-oriented GRC workflows.

→ Top-rated talks at BSides Las Vegas 2025

All talks from BSides Las Vegas 2025