Securing Generative AI: Is it all an Illusion?

Rachana Doshi (Director of Third Party Security · Salesforce), Michael Samson (Security Engineer · Salesforce)

BSidesSF 2024 · Day 1

Overview

This talk, "Securing Generative AI: Is it all an Illusion?", delivered by Rachana Doshi and Michael Samson from Salesforce, addresses the critical and rapidly evolving challenge of securing Generative AI systems, particularly those leveraging Large Language Models (LLMs). The speakers highlight the unprecedented speed at which Generative AI technologies, exemplified by ChatGPT's ascent to 100 million users in just two months, have been adopted across enterprises. This rapid integration has left security teams with mere weeks or months to establish robust security postures, a stark contrast to the years typically afforded for assessing and securing prior technological advancements.

Watch on YouTube

Visual summary for Securing Generative AI: Is it all an Illusion? by Rachana Doshi, Michael Samson
Visual summary for Securing Generative AI: Is it all an Illusion? by Rachana Doshi, Michael Samson

Key moments

  1. 02:00 Rapid GenAI adoption & security as an afterthought
  2. 05:00 RAG & LLM Plugins Primer
  3. 07:00 Interactive Threat Modeling Scenario
  4. 12:00 Comprehensive LLM Threat Landscape
  5. 20:00 Core Technical Controls for LLMs
  6. 24:00 Advanced Technical Controls & Guardrails
  7. 30:00 Testing Strategies for AI Applications
  8. 34:00 Lessons Learned & Secure Innovation

Securing Generative AI: Is it all an Illusion?

Speakers: Rachana Doshi, Michael Samson

Conference: BSidesSF 2024

YouTube: https://www.youtube.com/watch?v=zoIgHhx54F8

Overview

This talk, "Securing Generative AI: Is it all an Illusion?", delivered by Rachana Doshi and Michael Samson from Salesforce, addresses the critical and rapidly evolving challenge of securing Generative AI systems, particularly those leveraging Large Language Models (LLMs). The speakers highlight the unprecedented speed at which Generative AI technologies, exemplified by ChatGPT's ascent to 100 million users in just two months, have been adopted across enterprises. This rapid integration has left security teams with mere weeks or months to establish robust security postures, a stark contrast to the years typically afforded for assessing and securing prior technological advancements.

The core of the presentation focuses on providing practical, actionable information for security practitioners. Doshi and Samson guide the audience through an interactive threat modeling exercise, dissecting the potential vulnerabilities and risks inherent in LLM service providers and their integration into enterprise workflows. They delve into specific technical controls and strategic considerations necessary to protect sensitive enterprise data, maintain data privacy, and ensure the ethical use of these powerful new tools. The talk underscores the importance of foundational security practices, even as organizations navigate the complexities of emerging AI technologies, ultimately aiming to demystify the security landscape of Generative AI and empower defenders.

The urgency of this topic stems from the pervasive nature of Generative AI. Business units and CEOs are aggressively pursuing its integration, often without a full understanding of the associated security implications. The speakers emphasize that security can no longer be an afterthought; it must be an integral part of the design and deployment of AI systems. By sharing their experiences and methodologies from Salesforce, Doshi and Samson offer a timely and essential guide for organizations grappling with the security challenges posed by this transformative technology.

Background

▶ Watch: Rapid GenAI adoption & security as an afterthought (02:00)

The landscape of technology adoption underwent a seismic shift in late 2022 and early 2023 with the public release and subsequent explosion of Generative AI, particularly ChatGPT. Rachana Doshi highlighted the astonishing adoption rate: while the telephone took 75 years to reach 100 million users, the internet 7 years, and mobile phones 16 years, ChatGPT achieved this milestone in a mere two months. This unprecedented speed meant that by January 2023, Generative AI was the dominant topic, with CEOs announcing integration plans and business teams rapidly innovating on potential use cases.

Historically, new technologies have had a slower adoption curve, allowing security teams ample time—months or even years—to understand the technology, identify attack vectors, and develop appropriate controls. However, the rapid proliferation of Generative AI meant security practitioners were "laying down the tracks as the train was barreling down at us," or "building the plane's engine as we were flying the plane." This left security teams with weeks, not years, to figure out how to secure Generative AI for their enterprises. Key questions arose: How do you assess a Generative AI system? What are the risks and threats? How do you create secure development guidelines?

To facilitate understanding of the threats, Michael Samson provided a primer on two fundamental concepts often employed with LLMs:

  1. Retrieval Augmented Generation (RAG): This technique allows LLMs to provide contextual responses without requiring retraining or fine-tuning of the model. It works by injecting additional context directly into the prompt before sending it to the LLM for inference. To achieve this, an organization builds a vector database containing relevant content and its embeddings (numerical representations in vector format). When processing a prompt, embeddings for that prompt are generated, the vector database is queried for the most similar matches, and the top 'x' matches are then included as context in the prompt sent to the LLM. Metadata, such as the source of the content, is also stored alongside the embeddings.
  1. LLM Plugins for Tools: These are essentially function calls, often implemented as Python functions, that extend the capabilities of an LLM. The LLM is informed about the available functions, their required arguments, and examples of how to invoke them. When the LLM generates an output, the system parses it for any invocation attempts. If detected, the specified function is called with the supplied arguments, allowing the LLM to interact with external systems or perform specific actions.

These foundational concepts are crucial for understanding the attack surface and the unique security challenges presented by Generative AI systems, which often involve complex data flows and interactions with various internal and external resources.

Key Findings

▶ Watch: Interactive Threat Modeling Scenario (07:00)

The interactive threat modeling exercise and subsequent discussion by the speakers unveiled a comprehensive array of threats and critical considerations for securing Generative AI systems. These findings highlight the multifaceted nature of the challenge, encompassing traditional web application vulnerabilities, novel AI-specific risks, and complex supply chain security issues.

Threats Identified During Interactive Threat Modeling:

  • Arbitrary Code Execution via Prompt Injection: A significant risk in environments utilizing LLM plugins. An internal actor could achieve arbitrary code execution by crafting malicious prompts, especially if the plugin environment (e.g., Python interpreter, database query engine) is not adequately secured.
  • Service Provider Data Utilization: Concerns regarding how third-party LLM service providers handle submitted data. This includes questions about their data retention policies (e.g., storing data indefinitely), and how they utilize the data (e.g., for internal audits, which broadens access, or for training future models, which could lead to proprietary data leakage).
  • SQL Injection and Input Standardization: The LLM output, like any user input, must be treated as untrusted data. Failure to properly sanitize and parameterize inputs to backend systems (e.g., database query engines) can lead to SQL injection vulnerabilities. Similarly, indirect prompt injection can occur if internal resources used for RAG are not properly sanitized.
  • Authentication and Authorization Bypass: Lack of robust authentication and authorization controls at the API Gateway could allow unauthorized internal users to access the AI chat application. Furthermore, authorization bypass can occur if RAG data sources do not enforce the authorization policies of the original data sources, or if plugins execute queries in an elevated system context, potentially returning data the user should not access.
  • Compromise of Service Provider Environment: A breach of the third-party LLM service provider's infrastructure could lead to the exfiltration of persistent data (prompts, completions) and, depending on the attacker's foothold, real-time prompt exfiltration or tampering of completions and training data.
  • Over-Reliance on LLM Output: LLMs are prone to hallucinations and can produce responses outside expected parameters. Automatically invoking sensitive actions based solely on LLM output without human validation poses a significant risk.
  • Plugin Environment Compromise: The execution of arbitrary code within the plugin environment can lead to the compromise of the underlying system, unauthorized access to resources, and contamination across user sessions.
  • Prompt Injection for Information Disclosure: Malicious prompts can be crafted to trick the LLM into reflecting sensitive data, such as secrets embedded in the prompt template or context, back to the user in a completion.
  • Encoding Failure: As LLM-enabled applications are still web applications, general web application vulnerabilities apply. Failure to properly encode LLM output before rendering it in a browser can lead to unintended client-side code execution (e.g., Cross-Site Scripting (XSS)).
  • Availability Issues: Reliance on a single LLM service provider can introduce availability risks. System outages or resource exhaustion at the provider could interrupt critical business processes dependent on the AI application.

Key Considerations for Scoping Risk (Rachana Doshi):

  • Data Sensitivity and Impact: Assess what enterprise data (PII, IP, secrets) is being processed, stored, or transmitted by the LLM provider, its sensitivity, and the potential business impact if it were leaked or compromised.
  • Hosting Responsibility: Understand the security responsibilities based on the hosting model:
  • Supplier-hosted: Vendor is responsible for application and infrastructure security.
  • Hybrid: Shared responsibility model, where the vendor secures infrastructure, but the enterprise is responsible for secure configuration of the application/LLM.
  • On-prem: Enterprise is responsible for both application and underlying infrastructure security.
  • Type of Technology and Services: Differentiate between general Generative AI models (producing text, images, video, audio) and specific LLMs (primarily large text output), or models for code generation, classification, etc. Consider privacy and ethical implications of generated content.
  • Integration Type: Evaluate the risk of automated workflows where LLM output directly triggers actions (e.g., publishing articles, executing generated code), as these are particularly risky.
  • Vendor Relationship:
  • Third-party: Direct contractual relationship allows enforcement of security requirements and audits.
  • Fourth-party: Indirect relationships where a third-party vendor integrates a Generative AI provider. This necessitates asking the third-party about their security assurances and due diligence on their fourth-party AI vendors.

These findings collectively underscore that securing Generative AI is not merely about addressing new AI-specific threats, but also about rigorously applying established security principles to a rapidly evolving and interconnected technological stack.

Technical Deep Dive

▶ Watch: Core Technical Controls for LLMs (20:00)

Securing Generative AI requires a multi-layered approach, combining traditional security controls with AI-specific mitigations. The speakers outlined a series of technical controls and strategic considerations for mitigating the identified threats.

1. Independent Verification:

While not a direct technical control, independent security assessments serve as a crucial proxy for vendor maturity. Organizations should look for evidence of web application penetration tests, ISO certifications, and SOC 2 Type 2 audits. These provide assurance that external auditors have conducted invasive reviews of the vendor's IT and security practices, covering areas that the customer typically cannot replicate.

2. Encryption:

A fundamental control, encryption is paramount for Generative AI systems. All traffic and data must be encrypted end-to-end and at every node within the data flow. Prompts, completions, and inferences should only be unencrypted for the briefest periods necessary for processing, and then immediately re-encrypted before being transferred to the next step in the workflow. This minimizes the window of opportunity for data interception.

3. Zero Data Retention and Zero Training:

These are critical controls for protecting proprietary and sensitive enterprise data.

  • Zero Data Retention: This means the LLM vendor does not store any of the customer's data—including prompts, inferences, completions, or even log data—beyond the transitory memory required for immediate processing. Implementing this requires the customer to build robust client-side controls for abuse detection and monitoring, as they would be responsible for re-sending prompts if an issue occurs, given the vendor's lack of persistent storage.
  • Zero Training: This ensures that the enterprise's proprietary intellectual property is explicitly excluded from being used to train the LLM. This prevents the inadvertent leakage of sensitive data through model outputs to unauthorized external parties.

4. Data Masking:

A standard security control, data masking is vital in the context of Generative AI. Before sending prompts to the LLM, sensitive data such as Personally Identifiable Information (PII), intellectual property, or secrets should be scrubbed, obfuscated, or replaced. Upon receiving the completion, the sensitive information can be rehydrated back into the response before being presented to the user, creating a seamless experience while minimizing data exposure risk.

5. Human Validation:

For any sensitive workflows that rely on LLM output, human validation must be integrated. Given the potential for LLM hallucinations or unexpected responses, automated actions based on AI output should always be reviewed and approved by a human to prevent unintended consequences.

6. Input Standardization and Parameterization:

LLM output, like any external input, must be treated as untrusted data. This necessitates rigorous input standardization and parameterization. When LLM outputs are used to call functions or construct database queries, inputs must be properly sanitized to prevent vulnerabilities such as SQL injection. Using parameterized queries is essential to separate data from code.

7. Arbitrary Code Execution Mitigation:

In environments where LLMs can invoke tools that execute arbitrary code (e.g., Python interpreters), robust isolation is critical:

  • Network Filters: Implement strict network filters to limit outbound connections from the execution environment.
  • Unique Execution Environments: Spin up a unique execution environment, such as a container, for each invocation attempt or for each user session.
  • Kernel-Level Isolation: If managing the underlying host, ensure proper kernel-level isolation between containers using technologies like Firecracker or gVisor.

8. Prompt Injection and Jailbreak Mitigation:

To counter prompt injection and jailbreaking attempts, models can be utilized to assess both input prompts and output completions.

  • Commercial Solutions: Emerging commercial solutions offer WAF-like protection for AI systems.
  • Open-Source Options:
  • NVIDIA Nemo Guardrails: This acts as a wrapper for generation attempts, allowing the definition of specific guardrails for input and output. It uses custom prompts to have the LLM itself analyze whether the content adheres to the defined policies.
  • LLM Guard: This solution involves invoking various types of scanners (e.g., for toxicity, use case adherence) through which prompts and completions are passed. It often leverages open-source models from Hugging Face for its scanning capabilities.

9. Authorization Controls:

Implementing granular authorization for LLM-enabled systems presents unique challenges:

  • RAG Vector Databases: Granular access controls for vector databases are difficult to enforce. It is generally recommended to use data stores that the entire user base or specific departments should have access to, leveraging metadata filtering for broader categories. Attempting very granular controls risks synchronization issues between the vector database and original data source authorization policies.
  • Plugins Calling External Resources: For plugins accessing non-public resources, it is crucial to utilize the user's context whenever possible, typically through OAuth, to ensure that the authorization controls of the target system are enforced.

10. Secure Software Development Life Cycle (SSDLC):

A critical reminder that LLM-enabled applications are fundamentally still web applications. Therefore, all standard Secure Software Development Life Cycle (SSDLC) practices, including secure coding, regular security testing, and vulnerability management, continue to apply.

11. High Availability:

For applications supporting critical business processes, a high availability configuration is essential. This could involve redundant deployments or having a backup service provider for AI calls. While a backup provider might introduce additional overhead (different models, maintaining additional codebases), its consumption-based pricing might not add significant cost.

These technical controls, when implemented comprehensively, form a robust defense strategy against the evolving threats in the Generative AI landscape.

Demo / Proof of Concept

▶ Watch: Advanced Technical Controls & Guardrails (24:00)

While the presentation did not feature a live code demonstration or a traditional software proof of concept, Michael Samson led an interactive threat modeling exercise that served as a practical demonstration of identifying vulnerabilities in a typical Generative AI deployment.

The scenario involved an organization building an internal web application for an AI chat assistant, running in AWS. This application utilized Retrieval Augmented Generation (RAG) with internal resources and a plugin environment for executing tools like a Python interpreter and a database query engine. The organization relied on a third-party LLM service provider via API calls, which also stored prompts and completions for audits and future model training.

During this interactive segment, the audience actively participated, shouting out potential threats and risks. This collaborative approach effectively demonstrated how security practitioners can analyze a system's architecture and data flow to uncover vulnerabilities. The exercise highlighted risks such as arbitrary code execution via prompt injection in the plugin environment, concerns about the LLM service provider's data utilization and retention policies, potential SQL injection due to lack of input standardization, indirect prompt injection through internal RAG resources, and the absence of proper authentication and authorization controls.

This interactive session effectively served as a "proof of concept" for the threat modeling methodology itself, illustrating how a structured approach can rapidly surface critical security concerns in complex Generative AI systems.

Defensive Implications

▶ Watch: Lessons Learned & Secure Innovation (34:00)

The insights and technical controls discussed by Rachana Doshi and Michael Samson offer a clear roadmap for defenders navigating the Generative AI landscape. The core message is to integrate AI security into existing practices while adapting to new, AI-specific challenges.

1. Comprehensive Testing and Auditing:

  • Web Application Pen Testing: Standard web application penetration testing processes remain fully applicable. AI-enabled applications are still web applications and are susceptible to traditional vulnerabilities.
  • Information Disclosure: Actively test for information disclosure, specifically looking for secrets embedded in system prompts or general prompt templates that the LLM could be coerced to reflect back. Also, probe for sensitive data that might have been inadvertently included in the training data and can be extracted through specific prompts. The speakers recommended "Gandalf" as a fun way to get acquainted with prompt injection testing.
  • Vulnerable Integrations and Plugins: Dedicate significant testing effort to integrations and plugins, as they introduce a substantial attack surface. For instance, a Gmail plugin that pulls user emails for processing introduces another untrusted data source into the system, requiring careful scrutiny.
  • Guardrail Effectiveness: Rigorously test the effectiveness of implemented guardrails (e.g., using NVIDIA Nemo Guardrails or LLM Guard) to ensure they adequately mitigate prompt injection, jailbreaks, and adherence to acceptable use policies.

2. Third-Party and Fourth-Party Risk Management:

  • Contractual Agreements: For third-party solutions that include AI functions, ensure robust agreements are in place. Specifically, demand assurances that your data will not be used to train public models accessible to other customers.
  • Due Diligence on Fourth Parties: If your third-party vendor utilizes a fourth-party Generative AI provider, it becomes more complex to audit. It is crucial to ensure that your third-party conducts the same rigorous due diligence process on their fourth-party AI providers that you would perform yourself.
  • Scoped Assessments: Always scope security assessments to the specific risk factors of your particular use case, recognizing that every deployment has unique characteristics.

3. Foundational Security and Data Flow Analysis:

  • Don't Forget the Basics: Reiterate the importance of fundamental security controls such as authentication, authorization, and encryption. These are the bedrock upon which AI security must be built.
  • Data Flow Diagrams and System Architectures: Conduct thorough reviews of data flow diagrams and system architectures. Many components in AI systems are familiar, and this exercise helps quickly identify potential threats specific to your use case. This detailed understanding of how data is processed at every node is critical for implementing effective controls.
  • Translate Technical to Business Risk: When communicating security risks to business partners, translate technical jargon into clear business impact. This fosters better understanding and facilitates informed decision-making regarding risk acceptance.
  • Early Partnership: Engage with business partners early in the innovation cycle. Proactively helping them find secure solutions is more effective than reacting to insecure implementations later.
  • Risk Acceptance: Acknowledge that risk acceptance is a valid option, but ensure that both security teams and business partners fully understand the implications and potential consequences of accepting a particular risk.

4. Leveraging LLMs for Security Operations:

The speakers also highlighted how LLMs can be powerful tools for defenders:

  • DataCon AI: A project that implements a local RAG framework using local open-source models, allowing for a local AI chat assistant where data never leaves the system. This is ideal for sensitive internal data.
  • Tac AI: An AI-enabled threat modeling project that takes system designs (e.g., in AML) as input to generate data flow diagrams, potential threats, and mitigations. This can significantly accelerate the threat modeling process.
  • Microsoft Autogen: A framework for creating AI agents that can assist with security operations, such as performing initial triage on SIEM alerts or helping with security support requests and questions.

By embracing these defensive implications, organizations can move beyond the illusion of AI security and establish a robust, adaptable framework for protecting their data and systems in the age of Generative AI.

Key Takeaways

  • Rapid Adoption Demands Proactive Security: The unprecedented speed of Generative AI adoption (e.g., ChatGPT reaching 100 million users in two months) necessitates immediate and proactive security measures, rather than the traditional reactive approach.
  • Comprehensive Threat Modeling is Essential: Organizations must conduct thorough threat modeling, including detailed data flow analysis and understanding of vendor relationships (third-party and fourth-party), to identify unique risks specific to their Generative AI use cases.
  • Foundational Controls Remain Paramount: Basic security controls like end-to-end encryption, zero data retention, zero training, and data masking are critical for protecting sensitive enterprise data and intellectual property from leakage or misuse by LLM providers.
  • Mitigate AI-Specific Vulnerabilities: Implement robust input standardization, parameterization, and isolation for arbitrary code execution environments. Utilize AI-specific guardrails (e.g., NVIDIA Nemo Guardrails, LLM Guard) to defend against prompt injection, jailbreaks, and ensure adherence to acceptable use policies.
  • Don't Forget SSDLC and Web App Basics: Generative AI applications are still web applications; therefore, traditional Secure Software Development Life Cycle (SSDLC) practices, web application penetration testing, and secure coding principles remain fundamental.
  • Leverage AI for Security Operations: LLMs can be powerful tools for defenders, assisting with tasks like AI-enabled threat modeling (e.g., Tac AI), local data querying (e.g., DataCon AI), and automating security alert triage (e.g., Microsoft Autogen).

About the Speaker(s)

Rachana Doshi is the Director of Third Party Security at Salesforce. In this role, she is responsible for overseeing the security posture of third-party vendors and ensuring the protection of enterprise data within these relationships. Her expertise was crucial in navigating the early challenges of securing Generative AI systems at Salesforce.

Michael Samson is a Security Engineer at Salesforce. He brings a deep technical understanding to the team, focusing on the practical implementation of security controls and threat mitigation strategies for emerging technologies like Generative AI. His insights into RAG, LLM plugins, and specific technical controls were central to the presentation.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk provides a pragmatic and technically sound overview of securing generative AI systems within an enterprise context. The speakers effectively break down the unique threat landscape introduced by LLMs, covering common attack vectors like prompt injection and data exfiltration, and then pivot to actionable technical controls. While not revealing novel zero-days, the session offers a crucial synthesis of existing security principles applied to a rapidly evolving technology, making it highly relevant for practitioners grappling with GenAI adoption.

Heather Calloway (CISO) — STRONG ACCEPT

This presentation offers a highly relevant and actionable framework for CISOs and security leaders navigating the rapid adoption of generative AI. It effectively translates complex technical risks into clear business implications, emphasizing critical considerations such as data sensitivity, vendor accountability, and the shared responsibility model. The speakers provide concrete, implementable controls and a pragmatic approach to assessing and mitigating the institutional risks associated with integrating LLM services, making it a valuable resource for guiding secure innovation.

→ Top-rated talks at BSidesSF 2024

All talks from BSidesSF 2024