Private Investigator: Extracting Personally Identifiable Information from Large Language Models Using Optimized Prompts

Seongho Keum (KAIST)

34th USENIX Security Symposium (USENIX Security '25) · Day 3 · Vulnerabilities in LLMs: Privacy, Safety, and Defense

Overview

Large Language Models (LLMs) have revolutionized numerous fields, from translation and healthcare to code generation, by demonstrating unprecedented performance across a diverse range of tasks. This remarkable capability stems from their training on vast datasets, which often include a wide array of domain-specific information. However, a critical security vulnerability arises when these training datasets inadvertently contain Personally Identifiable Information (PII), such as personal names, email addresses, phone numbers, and physical addresses. LLMs can memorize these sensitive PII items, making them susceptible to extraction by an adversary. This talk introduces "Private Investigator," a novel framework designed to systematically extract memorized PII from LLMs using optimized prompt generation and selection strategies.

Watch on YouTube · Slides

Visual summary for Private Investigator: Extracting Personally Identifiable Information from Large Language Models Using Optimized Prompts by Seongho Keum
Visual summary for Private Investigator: Extracting Personally Identifiable Information from Large Language Models Using Optimized Prompts by Seongho Keum

Key moments

  1. 0:00 Introduction: PII extraction threat from LLMs
  2. 2:00 Limitations of prior work and Private Investigator framework
  3. 2:50 Prompt generation phase using a surrogate model
  4. 4:45 PII extraction with multi-armed bandit prompt selection
  5. 6:20 Evaluation results: Private Investigator extracts significantly more PII
  6. 8:00 Analyzing why Private Investigator outperforms baselines
  7. 9:10 Mitigation strategies (deduplication, differential privacy) and limitations
  8. 10:30 Conclusion: Private Investigator's impact and future defense needs

Private Investigator: Extracting Personally Identifiable Information from Large Language Models Using Optimized Prompts

Speakers: Seongho Keum

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=4d6WZEF1iIc

Overview

Large Language Models (LLMs) have revolutionized numerous fields, from translation and healthcare to code generation, by demonstrating unprecedented performance across a diverse range of tasks. This remarkable capability stems from their training on vast datasets, which often include a wide array of domain-specific information. However, a critical security vulnerability arises when these training datasets inadvertently contain Personally Identifiable Information (PII), such as personal names, email addresses, phone numbers, and physical addresses. LLMs can memorize these sensitive PII items, making them susceptible to extraction by an adversary. This talk introduces "Private Investigator," a novel framework designed to systematically extract memorized PII from LLMs using optimized prompt generation and selection strategies.

The PII extraction attack poses a significant threat to the burgeoning language model ecosystem, a threat that escalates as more fine-tuned language models are shared and deployed over time. Traditional PII extraction methods have relied on simplistic or externally sourced prompts, a limitation that Private Investigator aims to overcome. Developed by Seongho Keum from KAIST, Private Investigator represents a significant advancement in understanding and exploiting these memorization vulnerabilities, providing a robust method for assessing the security posture of LLMs.

This work not only demonstrates a superior method for PII extraction but also delves into the underlying mechanisms that make these attacks effective. By meticulously analyzing the internal states of language models, Private Investigator offers insights into how specific prompts trigger the emission of sensitive data. Furthermore, the research critically evaluates existing mitigation strategies, revealing their limitations and the persistent challenge of securing private information against sophisticated extraction techniques, even at the cost of model performance.

Background

▶ Watch: Introduction: PII extraction threat from LLMs (0:00)

The pervasive integration of Large Language Models (LLMs) into critical applications underscores the imperative for understanding their security vulnerabilities. A fundamental concern revolves around the potential for LLMs to inadvertently memorize and subsequently leak Personally Identifiable Information (PII). PII encompasses any data that could potentially identify a specific individual, including but not limited to names, email addresses, phone numbers, physical addresses, and even more nuanced details that, when combined, can uniquely pinpoint a person. The sheer scale and diversity of data used to train modern LLMs, often sourced from public internet crawls, proprietary databases, and domain-specific corpora, make the inclusion of PII virtually unavoidable. While this extensive training data is crucial for achieving high performance across various tasks, it simultaneously creates a risk surface where sensitive information can become embedded within the model's parameters.

Once memorized, this PII can be extracted by an adversary through carefully crafted queries, a process known as a PII extraction attack. This threat is particularly acute in the context of fine-tuned language models, which are increasingly being shared and utilized across different organizations and applications. Each fine-tuning process, often performed on specialized datasets, introduces new opportunities for PII to be inadvertently ingested and memorized. As these models proliferate, the attack surface for PII leakage expands dramatically, posing substantial privacy, regulatory, and reputational risks.

Prior research into PII extraction attacks primarily utilized relatively simplistic or "outsourced" prompts. These methods often involved generic tokens such as "start of sentence" markers, empty strings, or arbitrary text snippets scraped from the internet. While these approaches demonstrated the feasibility of PII extraction, they suffered from a key limitation: their lack of sophistication in prompt generation. Such simple prompts were not optimized to consistently or comprehensively elicit memorized PII, leaving a significant portion of sensitive data unextracted. This limitation highlighted a critical gap: how could a more sophisticated adversary generate a set of prompts to extract a significantly larger and more diverse array of PII items? This question forms the foundational motivation for the development of Private Investigator, seeking to move beyond the constraints of prior work by employing an optimized and systematic approach to prompt generation and selection.

Key Findings

▶ Watch: Prompt generation phase using a surrogate model (2:50)

The Private Investigator framework represents a significant leap forward in PII extraction capabilities, demonstrating a remarkable ability to uncover sensitive data previously inaccessible by state-of-the-art methods. Its key findings underscore the profound vulnerability of LLMs to sophisticated prompting strategies and the inadequacy of current defensive measures.

Foremost among the findings is the sheer volume of PII that Private Investigator can extract. In evaluations against a language model fine-tuned on the ankown dataset, the framework successfully extracted more than 6,000 email addresses and 36,000 personal names. This performance dramatically surpassed that of existing PII extraction techniques, highlighting the efficacy of its optimized prompt generation and selection. The ability to extract such a large quantity of PII from a single model underscores the pervasive nature of memorized data within LLMs.

Crucially, Private Investigator also demonstrated a superior capacity to extract exclusive PIIs—items that other methods failed to retrieve. This indicates that the framework isn't merely more efficient at finding already-known PII but is capable of uncovering entirely new sets of sensitive data that are uniquely elicited by its optimized prompts. This finding is particularly concerning, as it implies that even if an organization believes it has assessed its models for PII leakage using traditional methods, a substantial amount of sensitive information may remain vulnerable.

The research also provides critical insights into why Private Investigator is so effective. Through a deep investigation into the target model's hidden states, the study identified a concept termed the PII eliciting direction. This direction is conceptualized as a vector in the hidden space, representing the difference between the hidden states induced by PII-eliciting prompts and those induced by non-eliciting prompts. The key finding here is that the prompts generated by Private Investigator are demonstrably more aligned with this PII eliciting direction compared to prompts from other methods. This alignment explains the framework's enhanced ability to guide the language model towards outputting memorized PII, providing a theoretical basis for its empirical success.

Finally, the evaluation of mitigation strategies revealed a critical insight: current defensive techniques, such as deduplication and differential privacy (DPSG), while reducing the volume of extracted PII, are far from a complete solution. Even with these mitigations applied, Private Investigator still managed to extract approximately 8,000 personal names. More alarmingly, these mitigation strategies often came at a significant cost, leading to a substantial degradation in the model's performance, as evidenced by a considerable rise in the perplexity of the training data. This indicates a challenging trade-off between privacy and utility, underscoring the urgent need for more robust and less performance-impacting defensive mechanisms.

Technical Deep Dive

▶ Watch: Evaluation results: Private Investigator extracts significantly more PII (6:20)

Private Investigator operates on a sophisticated, two-phase framework: prompt generation and PII extraction. The ingenuity lies in its ability to systematically identify and leverage highly effective prompts, moving beyond the simplistic approaches of prior work. This framework is built upon two core ideas: first, that prompts effective on a surrogate model (a smaller, more accessible language model) can be effective on a larger target model, and second, that prioritizing promising prompts during extraction, based on intermediate results, significantly enhances efficiency.

Prompt Generation Phase

The goal of this phase is to create a diverse and potent set of prompts capable of eliciting PII.

  1. Surrogate Model Utilization: Instead of directly experimenting with the potentially large and computationally expensive target model, Private Investigator first leverages a surrogate model. This surrogate model, which shares architectural similarities or has been trained on related data, serves as an efficient proxy for identifying promising prompt candidates.
  2. Vocabulary Scan and PII Elicitation: For every single token within the surrogate model's entire vocabulary, the framework performs a systematic test. Each token is used as a prompt to generate a fixed number of output texts. From these outputs, the number of PII items (emails, phone numbers, names) is counted. This exhaustive scan quantifies the PII-inducing ability of each individual token.
  3. Top 1% Selection: After evaluating all tokens, Private Investigator identifies the top 1% of tokens that demonstrated the highest ability to induce PIIs. These tokens form the initial pool of "promising prompts."
  4. Diverse Prompt Set Construction: To ensure that the final set of prompts covers a broad spectrum of contexts and can trigger diverse PIIs, a selection strategy is employed to choose 20 prompts from the promising pool.
  • The first prompt selected is simply the one with the absolute highest number of extracted PIIs.
  • Subsequent prompts are chosen iteratively based on their distance in the hidden space from previously selected prompts. Specifically, the framework selects the prompt that has the longest distance from all already-selected prompts in the model's internal representation (the hidden state). This maximizes the diversity of the prompt set, ensuring each prompt explores a different "area" of the model's latent space, thereby increasing the likelihood of eliciting a wider variety of memorized PIIs.

PII Extraction Phase

Once the optimized set of 20 prompts is generated, the framework moves to the PII extraction phase, which focuses on efficiently applying these prompts to the target language model.

  1. Attack Campaign Setup: The extraction process is structured as an "attack campaign" consisting of 100 sequential PII extraction attempts. In each attempt, one chosen prompt is used to generate 2,000 output texts from the target model, and PII items are collected from these texts.
  2. Prompt Selection Strategy (Multi-Armed Bandit): A naive approach would be to simply cycle through the 20 prompts or use them with equal frequency. However, Private Investigator employs a more sophisticated prompt selection strategy inspired by the multi-armed bandit problem. This problem involves making a sequence of choices from multiple options (the "arms" or, in this case, prompts) to maximize cumulative reward (extracted PII), where the reward distribution for each option is initially unknown.
  3. Prompt Scoring Function: To dynamically select the most effective prompt for each extraction attempt, a prompt scoring function is utilized. This function balances two critical aspects:
  • Exploration Term: This component gives more weight and chances to prompts that have been selected less frequently so far. This ensures that the framework continues to explore the potential of all 20 diverse prompts, preventing premature convergence on a few seemingly effective ones and allowing for the discovery of new PII.
  • Exploitation Term: This component prioritizes prompts that have previously induced "promising PIIs." The "promisingness" of an extracted PII item is measured by its perplexity. Lower perplexity generally indicates that the text (the PII item) is more probable or "expected" by the model, often correlating with memorized content rather than novel generation. Prompts that consistently yield low-perplexity PII are thus exploited more frequently, as they are likely to be highly effective.

Underlying Mechanism: PII Eliciting Direction

A key analytical component of Private Investigator is the investigation into why the generated prompts are so effective. This involves delving into the target model's hidden state. The researchers define a PII eliciting direction as a vector in the hidden space. This vector is computed as the difference between the hidden states induced by prompts known to elicit PII and those induced by prompts that do not. Essentially, this vector represents the specific internal model state associated with the propensity to output PII. The analysis revealed that the prompts generated by Private Investigator are significantly more aligned with this PII eliciting direction compared to prompts generated by other methods. This strong alignment is the fundamental technical reason for Private Investigator's superior performance, as its prompts effectively guide the LLM's internal processing towards the memorized PII.

Demo / Proof of Concept

▶ Watch: Analyzing why Private Investigator outperforms baselines (8:00)

The efficacy of Private Investigator was rigorously demonstrated through a comprehensive attack campaign and evaluation, serving as its primary proof of concept. The experimental setup was designed to directly compare Private Investigator's performance against existing state-of-the-art PII extraction techniques.

The "attack campaign" involved deploying Private Investigator against four distinct language models. These models were specifically chosen because they had been fine-tuned on two different PII-rich datasets: an email record dataset and the ankown dataset. This choice of target models and datasets provided a realistic and challenging environment for testing PII extraction capabilities, reflecting scenarios where organizations might deploy models trained on sensitive internal data.

For each of the target models, Private Investigator executed 100 sequential PII extraction attempts. In every attempt, one of the 20 optimized prompts was selected using the multi-armed bandit strategy, and this prompt was used to generate 2,000 output texts from the target LLM. The generated texts were then meticulously scanned for the presence of three specific types of PII: email addresses, phone numbers, and personal names.

The results of this extensive evaluation unequivocally demonstrated Private Investigator's superior performance. When applied to a model fine-tuned on the ankown dataset, the framework successfully extracted more than 6,000 unique email addresses and 36,000 unique personal names. This significantly outstripped the number of PII items extracted by other baseline PII extraction methods, even when normalized for the fixed number of queries.

Furthermore, a critical aspect of the demonstration was the analysis of exclusively extracted PIIs. This metric specifically counted PII items that Private Investigator managed to extract, but which all other compared methods failed to retrieve. The findings showed that Private Investigator extracted the largest number of these exclusive PIIs, confirming its ability to uncover sensitive data that remains hidden from less sophisticated approaches. This highlights that the framework is not just more efficient but fundamentally more capable of probing the deep memorization within LLMs. The systematic nature of the evaluation, coupled with the quantitative superiority of the results, provided compelling proof of concept for Private Investigator's effectiveness as a leading tool for PII extraction from large language models.

Defensive Implications

▶ Watch: Conclusion: Private Investigator's impact and future defense needs (10:30)

The findings from Private Investigator carry profound implications for the defense of Large Language Models against PII extraction attacks. The research not only exposes critical vulnerabilities but also critically evaluates the effectiveness of current mitigation strategies, revealing their significant limitations.

The study investigated two primary mitigation techniques: deduplication and differential privacy.

  1. Deduplication: This is a data preprocessing technique applied during the training data preparation phase. Its core principle is to identify and remove repeated chunks of text from the training dataset. The rationale is that if specific PII items appear only once or a very limited number of times, the language model will find it harder to memorize them verbatim. By reducing redundancy, deduplication aims to make it more challenging for the model to internalize and later regurgitate sensitive information.
  2. Differential Privacy (DP): Specifically, the researchers adapted Differentially Private Stochastic Gradient Descent (DPSG). This technique is applied during the model training phase. DPSG works by adding carefully calibrated noise to the gradients during the stochastic gradient descent optimization process. The goal is to bound the impact of any single data sample on the final model's behavior. In essence, it aims to ensure that the model's parameters do not reveal too much about any individual training example, thus providing a strong privacy guarantee.

While both deduplication and differential privacy did lead to a significant drop in the number of extracted PIIs compared to an unprotected model, Private Investigator still demonstrated remarkable resilience against these defenses. Even with these mitigation strategies in place, the framework was able to extract approximately 8,000 personal names. This substantial remaining leakage underscores that these existing defenses, while useful, are not a complete solution. A leakage of 8,000 names, even if significantly reduced from an unmitigated scenario, still represents a substantial privacy breach.

More critically, the study highlighted a significant drawback of these mitigation strategies: they often degrade the model's performance significantly. The research observed that the perplexity of the training data on each model rose considerably after applying these defenses. Perplexity is a measure of how well a probability model predicts a sample. A higher perplexity indicates that the model is less confident in its predictions and generally performs worse on natural language tasks. This means that organizations face a difficult trade-off: enhancing privacy through these methods often comes at the cost of reduced model utility and performance.

In conclusion, the defensive implications are stark: current state-of-the-art mitigation techniques are insufficient to fully secure private information against sophisticated PII extraction attacks like Private Investigator. Furthermore, the performance degradation associated with these defenses makes their widespread adoption challenging in practical, production environments where model utility is paramount. This necessitates the urgent development of stronger and more realistic defense methods that can effectively protect PII without crippling the performance capabilities of large language models. The Private Investigator framework itself can serve as a crucial tool for testing the vulnerability of target language models to PII extraction attacks, enabling developers and security researchers to proactively assess and improve the privacy posture of their LLMs.

Key Takeaways

  • Significant PII Leakage: Large Language Models trained on diverse datasets can memorize and leak substantial amounts of Personally Identifiable Information (PII), posing a critical privacy threat.
  • Optimized Prompting is Superior: Private Investigator, a novel framework, significantly outperforms prior PII extraction methods by using optimized prompt generation and a multi-armed bandit selection strategy.
  • Extraction of Exclusive PII: The framework is capable of extracting "exclusive PIIs" that other state-of-the-art techniques fail to uncover, indicating a deeper and more comprehensive probing of memorized data.
  • Underlying Mechanism Discovered: Private Investigator's effectiveness stems from its prompts being highly aligned with a "PII eliciting direction" in the model's hidden state, providing a theoretical basis for its success.
  • Mitigations Are Insufficient: Current mitigation strategies like deduplication and differential privacy (DPSG), while reducing PII leakage, are not a complete solution and can lead to significant degradation in model performance.
  • Urgent Need for Stronger Defenses: There is a critical need for more robust, realistic, and less performance-impacting defense mechanisms to secure LLMs against sophisticated PII extraction attacks.

About the Speaker(s)

Seongho Keum is a researcher from KAIST (Korea Advanced Institute of Science and Technology). His work focuses on understanding and mitigating security vulnerabilities in large language models, particularly concerning the extraction of private information. His research, as presented in "Private Investigator," contributes significantly to the field of AI security by developing advanced techniques for identifying and assessing PII leakage risks in LLM ecosystems.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Legitimate academic research with a functional contribution — optimized prompt selection via multi-armed bandit and the 'PII eliciting direction' vector analysis are genuinely interesting mechanisms. The numbers are real and the methodology is reproducible, but the core insight (LLMs memorize training data and you can extract it with better prompts) is already well-established territory, and the defensive analysis lands on the unsatisfying conclusion that nothing works great.

Heather Calloway (CISO) — WEAK

Technically credible research that demonstrates a real and growing privacy risk in LLM deployment — but it stops at the lab door. The work quantifies leakage, reveals the inadequacy of current mitigations, and offers a useful mechanistic explanation, yet delivers almost nothing for the organizations actually deploying these models.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)