Securing Agentic AI Systems and Multi-Agent Workflows

Andra Lezza (AppSec Specialist · Sage), Jeremiah Edwards (Lead of worldwide AI delivery organization · Sage)

DEF CON 33 · Day 1 · Main Stage

Overview

In an era witnessing the rapid proliferation of AI agents and multi-agent systems, this DEF CON talk by Andra Lezza and Jeremiah Edwards of Sage delves into the critical security considerations for deploying these sophisticated technologies, particularly in high-stakes environments. Moving beyond the security challenges of static AI assistants and co-pilots, the speakers illuminate how the emergent capabilities of agentic AI, such as persistent memory, dynamic tool invocation, and autonomous decision-making, introduce amplified risks and entirely new attack vectors.

Watch on YouTube

Visual summary for Securing Agentic AI Systems and Multi-Agent Workflows by Andra Lezza, Jeremiah Edwards
Visual summary for Securing Agentic AI Systems and Multi-Agent Workflows by Andra Lezza, Jeremiah Edwards

Key moments

  1. 0:00 Introduction: Securing agentic AI systems & multi-agent workflows
  2. 1:18 Evolution from AI assistants and copilots to agents
  3. 2:40 Core vulnerabilities in Agentic AI systems
  4. 4:40 Real-world incident: AutoGPT unintended privilege escalation
  5. 5:40 MCP protocol: Power, complexity, and new risks
  6. 7:40 Key principle: Tailoring security to system use cases

Securing Agentic AI Systems and Multi-Agent Workflows

Speakers: Andra Lezza (AppSec Specialist, Sage); Jeremiah Edwards (Lead of worldwide AI delivery organization, Sage)

Conference: DEF CON

YouTube: https://www.youtube.com/watch?v=5fJ6u--GkSk

Overview

In an era witnessing the rapid proliferation of AI agents and multi-agent systems, this DEF CON talk by Andra Lezza and Jeremiah Edwards of Sage delves into the critical security considerations for deploying these sophisticated technologies, particularly in high-stakes environments. Moving beyond the security challenges of static AI assistants and co-pilots, the speakers illuminate how the emergent capabilities of agentic AI, such as persistent memory, dynamic tool invocation, and autonomous decision-making, introduce amplified risks and entirely new attack vectors.

The core of their discussion centers on the evolving threat landscape, drawing parallels with established security frameworks like the OWASP Top 10, while highlighting the unique dangers posed by autonomous agents. They emphasize that while many traditional security principles still apply, the inherent complexity and dynamic nature of agentic systems necessitate novel architectural patterns and a proactive, design-centric approach to security. The talk serves as a timely warning and a practical guide for developers and security professionals grappling with the secure implementation of AI agents in mission-critical business applications.

Lezza and Edwards underscore the urgency of addressing these challenges, noting the breakneck speed at which AI agents are being integrated into enterprise workflows. They argue that without a fundamental shift in how these systems are designed and secured, organizations risk baking in vulnerabilities at foundational levels, leading to potential data breaches, unauthorized actions, and significant business disruptions. Their insights are crucial for anyone involved in building or securing the next generation of AI-powered applications.

Background

▶ Watch: Introduction: Securing agentic AI systems & multi-agent workflows (0:00)

The evolution of AI in enterprise applications has progressed rapidly, moving from basic AI assistants and co-pilots to highly autonomous agentic AI systems and multi-agent workflows. Initially, systems primarily focused on natural language processing (NLP) to understand user intent, leveraging prompt engineering, Retrieval Augmented Generation (RAG), and backend API calls to provide suggestions or answers. However, the landscape has dramatically shifted, especially since August 2025, which the speakers humorously refer to as "the year of the agent." These new agentic systems are now connected to vastly more data and are being deployed for critical functionality across diverse sectors, handling sensitive information, automating operations, and informing high-stakes decisions.

This increased autonomy and connectivity, while powerful, has amplified existing vulnerabilities and introduced new ones. Traditional threats like prompt injection, misconfigured tools, credential leaking, and unauthorized or arbitrary code execution become far more dangerous when exploited in agentic contexts. Malicious actors can manipulate agents to steal data or execute actions without direct human interaction. The speakers highlight the proliferation of complexity in this space, with various types and levels of agents, sophisticated design patterns (e.g., reflection, planning, chain of thought), and a constant influx of new frameworks, models, and Large Language Models (LLMs) from providers like OpenAI and Anthropic. This dynamic environment makes tracking and securing agentic systems exceedingly challenging.

A real-world example of these risks is AutoGPT, an early agentic system where unintended privilege escalation was discovered during testing. This occurred because the agent could persist memory access across multiple sessions, leading to actions beyond its intended simple task. A significant factor enabling such broad interaction is the use of tools – external software systems that agents can dynamically interact with. This capability is further amplified by emerging protocols like the Multi-Agent Communication Protocol (MCP). The speakers critically describe MCP as the "USBC of AI," implying that while it standardizes LLM interaction with external tools and data, simplifying integration, it also encourages the dangerous practice of plugging in arbitrary, untrusted code. This creates a new layer of risk, turning a dynamic code execution engine (an LLM wrapped in code) into a potential disaster zone if third-party tools with varying trust levels are indiscriminately integrated. The talk focuses specifically on securing high-stakes AI, such as systems used in construction, HR, payroll, and accounting, where failure carries significant business impact.

Comparing the OWASP Top 10 for Web Applications with the OWASP Top 10 for LLMs (released over the past year), the speakers note that only three entirely new risks are introduced by LLM technology: system prompt leakage, vector and embedding weaknesses, and misinformation. However, when these are viewed through the lens of agentic AI, their danger is significantly amplified. Agents can persist states, invoke external tools, escalate privileges, operate across workflows, and overwhelm human reviewers with alerts. Specifically:

  • System prompt leakage in agents can expose sensitive details like guardrails, authentication tokens, or system instructions, enabling attackers to impersonate agents, perform unauthorized actions, and break audit trails.
  • Vector and embedding weaknesses can be exploited if attackers gain access to vector store formation, deceiving agents, bypassing security controls, and poisoning sensitive workflows without obvious signs.
  • Misinformation (or hallucinations) in an agentic world can lead to the hijacking of agent chains, especially in recursive loops or chain of thought processes, resulting in denial of service or misalignment of agent goals. The presence of persistent memory in agents acts as a "poison pill," potentially re-running attacks every time the memory is reloaded.

The speakers also announced the kickoff of an OWASP Agentic AI Top 10 project, with the goal of releasing a dedicated list by November of the current year, underscoring the community's recognition of these emerging threats.

Key Findings

▶ Watch: Core vulnerabilities in Agentic AI systems (2:40)

The talk "Securing Agentic AI Systems and Multi-Agent Workflows" delivered several critical findings regarding the unique security challenges posed by the rapid adoption of AI agents:

  1. Amplified LLM Risks: While the OWASP Top 10 for LLMs introduced three new risks (system prompt leakage, vector/embedding weaknesses, misinformation), agentic AI significantly amplifies these. Agents' abilities to persist state, invoke external tools, escalate privileges, and operate across complex workflows make these vulnerabilities far more dangerous, potentially leading to widespread data breaches or system compromise.
  2. MCP as a Double-Edged Sword: The Multi-Agent Communication Protocol (MCP), while simplifying the integration of LLMs with external tools, introduces a critical new layer of risk. Its analogy to "the USBC of AI" highlights the danger of indiscriminately connecting dynamic code execution engines to third-party tools of varying trust levels, effectively creating a broad attack surface.
  3. Lack of Standardized Definitions Impedes Security: The absence of a universally agreed-upon definition for "agent" or "agentic AI" across major players like OpenAI, Anthropic, Nvidia, and Google complicates threat modeling and the development of consistent security standards. This fluid landscape demands equally dynamic and adaptable security approaches.
  4. Classical Security Principles Remain Crucial, But Require Adaptation: Approximately 75% of agentic AI security still relies on traditional security best practices such as input sanitization, least privilege access, robust logging and monitoring, and strong access controls. However, these principles must be thoughtfully adapted and integrated into the unique architecture of agentic systems, including the use of LLM guardrails and constitutional self-critique.
  5. Architectural Shift for High-Stakes AI: For high-stakes applications, a "shift right" in AI security is advocated. This involves moving AI-interacting components (both the MCP client and server) into a high-trust infrastructure boundary, separate from potentially insecure end-user UIs. This architectural separation enhances control over tool selection, secures client-server communication, and protects against misconfigured clients.
  6. The Necessity of a Permissions Proxy: Given that most existing tools and APIs were not designed with AI integration in mind (featuring disparate RBAC and authentication mechanisms), implementing a dedicated permissions proxy is crucial. This separate control plane uniformizes access control patterns across diverse tools, providing a centralized point for security decisions.
  7. Supply Chain Attacks Evolve: The autonomous nature of agents, particularly their dynamic tool selection capabilities, transforms traditional workflows into complex supply chain scenarios. This opens new avenues for supply chain attacks, emphasizing the need for strict controls over tool origins, regular auditing, and whitelisting.

Technical Deep Dive

▶ Watch: Real-world incident: AutoGPT unintended privilege escalation (4:40)

The technical exposition of the talk systematically breaks down the architecture of agentic AI systems, their inherent vulnerabilities, and proposed defensive strategies. At its core, an agent is defined as an LLM augmented with external code and the ability to interact with tools – other software systems – dynamically. The speakers describe an evolution of agent sophistication, starting from simple reflexive agents (akin to if-then statements) and model-based agents (if-then with randomness), progressing to more advanced goal-based agents, utility-based agents, and learning agents. It is with learning agents, which possess memory, that new threats emerge, including the leaking of internal state, incremental takeovers, misdirection, and denial-of-service due to persistent "poison pill" memory.

A critical component enabling this interaction is the Multi-Agent Communication Protocol (MCP). MCP standardizes how LLMs interact with external tools and data, simplifying integration but also introducing a new layer of risk. The typical conceptual architecture involves a user interacting with a UI (often a chat interface), which then communicates with a backend system comprising an MCP server and a selection of tools. Tools can range from file systems and web search to email senders and enterprise applications. A significant architectural flaw identified is the common practice of combining the user-facing UI with a true MCP client, crossing a trust boundary directly into server-side tool selection. This pattern exposes userland prompt information directly to LLMs and agents on the server, making them vulnerable to prompt injection, system prompt escape, and unauthorized overrides, as client-side guardrails are often absent or easily bypassed.

The integration of existing enterprise tools presents a substantial challenge. Most legacy APIs and databases were not built with AI in mind, possessing disparate role-based access control (RBAC) systems and authentication technologies. When these are all plugged into a single control plane at the MCP server layer, it creates a complex decision-making nexus for tool invocation, leading to potential misconfigurations and privilege escalation.

To counter these architectural weaknesses, the speakers propose several key technical defenses:

  1. Permissions Proxy: Instead of relying on the MCP server to manage diverse permissions, a separate control plane acting as a permissions proxy is recommended. This proxy would uniformize access control patterns across all integrated tools, centralizing and simplifying permission management, thus embodying a "security by design" approach.
  2. "Shift Right" on AI Security: This counter-intuitive strategy advises separating the end-user UI and client from the MCP client. The AI-interacting components (both the MCP client and server) should be pulled within a higher trust boundary in the organization's infrastructure. This "shift right" approach offers several benefits:
  • Tool Isolation: By isolating the MCP server, organizations gain better control over tool interactions, mitigating risks associated with the "lacking S in MCP" (where 'S' does not stand for security).
  • Controlled Communication: Implementing the MCP client server-side allows full control over the client-server communication at the MCP level, which is critical for securing what would otherwise be an inherently insecure end-user-facing client. It prevents misconfigured clients from pointing to arbitrary MCP servers.
  1. Threat Modeling a Single Agent System (Payroll Example): The speakers illustrate these concepts through a conceptual threat model of a payroll system.
  • Components: An external user, an agent environment (agent host, LLM), an MCP environment (client, server), and an enterprise backend (payroll application with endpoints for querying payslips, updating data, adjusting salaries, generating reports).
  • Trust Boundaries: Defined between the external user, the agent environment, and the enterprise backend.
  • Threat Scenarios:
  • Prompt Injection:
  • Input Injection: A malicious user request payload flows through the agent environment, triggers the LLM for tool selection, and then the MCP client passes arbitrary data within the tool description to the payroll application. This can lead to injection at the application level.
  • Output Injection: Unsanitized output from payroll application API endpoints (e.g., query payslips) flows back through the MCP client to the LLM. The LLM implicitly trusts this input, making it vulnerable to prompt injection that can then be passed back to the user.
  • Controls: Input sanitization (a classic defense) and LLM guardrails (a defense-in-depth check unique to LLMs) are crucial.
  • Data Leak:
  • Scenario: A malicious user, intending to query their own payslip, could exploit vulnerabilities to retrieve sensitive data, such as everyone's payslips in the company. A real-world parallel cited was GitHub's MCP server being tricked into revealing private repository information.
  • Controls: Strict data scraping prevention (e.g., ensuring PII is not logged or passed through the MCP server/client) and implementing least privilege access with robust Access Control Lists (ACLs) are essential.
  • Business Logic Flaws:
  • Scenario: An agent could be manipulated to perform unauthorized actions, such as adjusting an employee's salary to an exorbitant amount (e.g., $1 billion).
  • Controls: Comprehensive logging and monitoring to maintain an audit trail, and the implementation of kill switches or shutdown mechanisms for critical business operations, are vital.

The discussion briefly extends to multi-agent systems, which are described as the single-agent scenario "on steroids." These systems involve an orchestration agent managing multiple domain-specific or trust-level-segregated agents, each potentially accessing multiple tools via multiple MCP servers. The complexity is exponentially higher, especially with emerging inter-agent communication (A2A) protocols. The speakers conclude that many vulnerabilities in legacy systems are being exposed as they are wrapped in tools and integrated into agentic systems, making it imperative to audit underlying tools.

Demo / Proof of Concept

▶ Watch: MCP protocol: Power, complexity, and new risks (5:40)

While the talk did not feature a live, interactive demonstration or a functioning proof of concept, the speakers dedicated a significant portion of their presentation to conceptually illustrating threat modeling for agentic AI systems. They described a detailed scenario involving a single agent interacting with a payroll system, outlining the components, trust boundaries, data flows, and potential threat vectors for prompt injection, data leaks, and business logic flaws. Although the audience was asked to "imagine" the diagrams, the structured walkthrough provided a clear mental model of how these vulnerabilities manifest and how controls could be applied.

Defensive Implications

▶ Watch: Key principle: Tailoring security to system use cases (7:40)

The defensive implications derived from Lezza and Edwards' talk are multifaceted, combining traditional security best practices with novel architectural and design considerations specific to agentic AI:

  1. Strategic UI Design: Organizations should critically evaluate the necessity of a chat interface for their AI applications. While popular, chat interfaces inherently introduce a larger attack surface. If a full chat experience isn't strictly required, opting for alternative generative AI integrations can eliminate a significant class of threats by design.
  2. Own Your MCP Client: For high-stakes use cases, take ownership of the Multi-Agent Communication Protocol (MCP) client. Developing or heavily customizing your own client, rather than relying on off-the-shelf solutions, provides a higher degree of familiarity and control over how tools are selected and how AI interacts with deterministic software. This also enables secure transmission of user identity from the UI layer across the AI system to the tools.
  3. Segregate MCP Servers and Tools: Implement a strategy to segregate MCP servers and their associated tools based on the agent, its level of trust, and the underlying permission model. Not all tools or agents require the same level of lockdown. Grouping public-facing tools separately from those requiring high-level permissions simplifies implementation while significantly enhancing security.
  4. Implement a Permissions Proxy: Introduce a dedicated permissions proxy as a separate control plane. This proxy's role is to uniformize and manage access control patterns across the diverse, often legacy, tools that an agent might interact with. This centralizes permission management and allows for consistent enforcement, mitigating risks associated with disparate RBAC systems.
  5. Adopt a "Shift Right" Security Posture for AI: Counter-intuitively, security for agentic AI should "shift right." This means pulling all AI-interacting components, including both the MCP client and MCP server, into a higher-trust boundary within the organization's infrastructure. This isolates them from potentially insecure end-user UIs, providing robust tool isolation and full control over MCP-level communication, protecting against misconfigured clients and arbitrary server connections.
  6. Reinforce Traditional Security Basics: About 75% of agentic AI security relies on performing traditional security well. This includes:
  • Input Sanitization: Rigorous validation and sanitization of all user inputs and data flowing into agents and tools.
  • Least Privilege Access: Ensuring agents and tools only have the minimum necessary permissions to perform their functions.
  • Logging and Monitoring: Comprehensive audit trails for all agent actions, tool invocations, and data access, enabling detection of anomalous behavior and post-incident analysis.
  • Robust Access Controls (ACLs): Implementing strong access control lists on all data and systems accessible by agents.
  1. Leverage LLM-Specific Controls:
  • LLM Guardrails: Implement explicit guardrails for LLMs to constrain their behavior, prevent undesirable outputs, and enforce ethical guidelines.
  • Constitutional Self-Critique: Encourage LLMs to critically evaluate their own outputs against ethical or security principles, adding a layer of self-correction.
  1. Audit Underlying Tools: The adoption of agentic AI provides a crucial opportunity to audit and improve the security of underlying legacy tools and APIs. Many vulnerabilities in these systems may become exposed when dynamically invoked by agents.
  2. Build in Kill Switches and Shutdowns: For high-stakes business logic, incorporate mechanisms like kill switches or shutdown procedures to immediately halt agent operations if malicious or unintended behavior is detected, preventing catastrophic outcomes.
  3. Treat AI Tooling as a Supply Chain: Recognize that dynamic tool selection by autonomous agents introduces supply chain risks. Organizations must lock down the sources of MCP tools, audit them regularly, and implement whitelisting to prevent the integration of malicious or vulnerable third-party components.
  4. Continuous Threat Modeling and Testing: Given the nascent and rapidly evolving nature of agentic AI, continuous threat modeling and security testing are indispensable. This iterative process helps identify new attack vectors and ensures that security controls remain effective against emerging threats.

Key Takeaways

  • Security Basics Still Apply: Approximately 75% of securing agentic AI systems involves applying traditional security best practices effectively, such as input sanitization, least privilege, robust logging, and access controls.
  • Move Carefully, Design Safely: Despite the hype surrounding AI, organizations must exercise caution, avoid rushing deployments, and prioritize security-by-design principles from the outset to prevent baking in vulnerabilities.
  • Threat Model Everything: Thorough threat modeling is absolutely critical for understanding the complex interactions, data flows, and potential attack surfaces within single and multi-agent systems.
  • Continuous Testing is Essential: Given the dynamic and evolving nature of agentic AI, ongoing security testing and auditing are necessary to adapt to new threats and ensure the continued effectiveness of controls.
  • Strategic Architectural Decisions: Avoid chat interfaces unless absolutely necessary, bring your own MCP client for greater control, and segregate MCP servers and tools by agent and trust level to enhance security and simplify management.
  • Implement Permissions Proxies and Shift Right: Introduce a permissions proxy to uniformize access controls across diverse tools and adopt a "shift right" approach by isolating AI-interacting components within a high-trust infrastructure boundary.

About the Speaker(s)

Andra Lezza is an AppSec Specialist at Sage, where she focuses on application security. She is also actively involved in the cybersecurity community as one of the OWASP London chapter leads, indicating her commitment to advancing secure development practices.

Jeremiah Edwards leads a worldwide AI delivery organization at Sage. His role involves overseeing teams of data scientists and engineers who are responsible for building AI applications, providing him with deep insights into the practical challenges and opportunities of AI deployment in enterprise environments.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Competent, practitioner-level survey of agentic AI security concerns from people who are clearly building this stuff at Sage — the threat modeling walkthrough and MCP architectural critique are the most useful parts. Nothing here will surprise researchers who've been tracking LLM security since 2023, but it's an honest, non-vendor-y treatment that earns its slot.

Heather Calloway (CISO) — SOLID

A competent practitioner-level walkthrough of agentic AI security risks with genuine architectural advice, but it stays at the builder layer and never reaches the institutional or governance questions that matter most for security leaders. Useful for AppSec teams standing these systems up; limited for CISOs deciding whether and how to govern them.

→ Top-rated talks at DEF CON 33

All talks from DEF CON 33