The Ghost Navigator: Revisiting the Hidden Vulnerability of Localization in Autonomous Driving
Junqi Zhang
34th USENIX Security Symposium (USENIX Security '25) · Day 2 · Hardware Security 2
Overview
This distinguished paper, "We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs," delves into a critical and emerging threat to the software supply chain: package hallucinations. Authored by Joseph Spracklen and colleagues from the University of Texas at San Antonio, the University of Oklahoma, and Virginia Tech, this research illuminates how Large Language Models (LLMs) used for code generation can inadvertently recommend non-existent software packages. These erroneous recommendations create a novel vector for package confusion attacks, where malicious actors can register packages with hallucinated names, tricking unsuspecting developers into downloading and integrating harmful code into their projects.
Read the paper · Download the PDF (PDF) · Slides
Paper abstract
The reliance of popular programming languages such as Python and JavaScript on centralized package repositories and open-source software, combined with the emergence of code-generating Large Language Models (LLMs), has created a new type of threat to the software supply chain: package hallucinations. These hallucinations, which arise from fact-conflicting errors when generating code using LLMs, represent a novel form of package confusion attack that poses a critical threat to the integrity of the software supply chain. This paper conducts a rigorous and comprehensive evaluation of package hallucinations across different programming languages, settings, and parameters, exploring how a diverse set of models and configurations affect the likelihood of generating erroneous package recommendations and identifying the root causes of this phenomenon. Using 16 popular LLMs for code generation and two unique prompt datasets, we generate 576,000 code samples in two programming languages that we analyze for package hallucinations. Our findings reveal that that the average percentage of hallucinated packages is at least 5.2% for commercial models and 21.7% for open-source models, including a staggering 205,474 unique examples of hallucinated package names, further underscoring the severity and pervasiveness of this threat. To overcome this problem, we implement several hallucination mitigation strategies and show that they are able to significantly reduce the number of package hallucinations while maintaining code quality. Our experiments and findings highlight package hallucinations as a persistent and systemic phenomenon while using state-of-the-art LLMs for code generation, and a significant challenge which deserves the research community's urgent attention.

We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs
Speakers: Joseph Spracklen (University of Texas at San Antonio); Raveen Wijewickrama (University of Texas at San Antonio); A H M Nazmus Sakib (University of Texas at San Antonio); Anindya Maiti (University of Oklahoma); Bimal Viswanath (Virginia Tech); Murtuza Jadliwala (University of Texas at San Antonio)
Conference: USENIX Security
YouTube: N/A (Peer-reviewed paper)
Overview
This distinguished paper, "We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs," delves into a critical and emerging threat to the software supply chain: package hallucinations. Authored by Joseph Spracklen and colleagues from the University of Texas at San Antonio, the University of Oklahoma, and Virginia Tech, this research illuminates how Large Language Models (LLMs) used for code generation can inadvertently recommend non-existent software packages. These erroneous recommendations create a novel vector for package confusion attacks, where malicious actors can register packages with hallucinated names, tricking unsuspecting developers into downloading and integrating harmful code into their projects.
The paper presents a rigorous and comprehensive evaluation of this phenomenon across a diverse set of 16 popular code-generating LLMs, encompassing both commercial and open-source models. By generating and analyzing an astounding 576,000 code samples in Python and JavaScript, the authors quantify the prevalence of package hallucinations, characterize their nature, and investigate various factors influencing their occurrence. Their findings reveal a pervasive and systemic issue, with significant implications for the integrity and security of modern software development workflows that increasingly rely on AI assistance.
The significance of this work lies in its systematic identification and quantification of a previously underexplored LLM-specific vulnerability that directly impacts software supply chain security. By not only exposing the problem but also proposing and evaluating several mitigation strategies, the researchers provide crucial insights for both developers leveraging LLMs and security professionals tasked with protecting software ecosystems. The paper's recognition with a Distinguished Paper Award at USENIX Security 2025 underscores its profound impact and the urgency of addressing this novel threat.
Background
Modern software development is heavily reliant on open-source packages and centralized repositories like PyPI for Python and npm for JavaScript. This ecosystem, while fostering rapid innovation, also presents significant security challenges. Package confusion attacks, a long-standing issue, exploit this reliance by tricking users into downloading malicious packages that mimic legitimate ones through techniques such as typosquatting, combosquatting, and brandjacking. Notable incidents like the PyTorch compromise and campaigns by the Lazarus Group highlight the severity of these attacks, which can compromise entire software products through dependency chains.
The advent of sophisticated code-generating LLMs, such as GPT-4 and Llama, has revolutionized programming workflows, with studies indicating that up to 97% of developers use generative AI and approximately 30% of code is now AI-generated. While these tools offer immense productivity gains, they also introduce new security paradigms. A well-documented shortcoming of LLMs is hallucination, where models generate factually incorrect, nonsensical, or misleading information. Most prior research on LLM hallucinations focused on natural language tasks like translation or summarization, with the impact on code generation remaining less explored. Liu et al. previously established a taxonomy of hallucinations in LLM-generated code, but the specific threat of package hallucinations was not comprehensively analyzed.
Package hallucination, as defined in this paper, occurs when an LLM generates code that recommends or includes a reference to a package that does not actually exist in a public repository. This creates a fertile ground for a novel type of package confusion attack. An adversary can monitor LLM outputs, identify frequently hallucinated package names, and then publish a malicious package under that exact name to an open-source repository. Unsuspecting developers, trusting the LLM's recommendation, would then unknowingly install the malicious package, leading to a compromise that can propagate through their codebase and dependency chain. Existing repository defenses, such as two-factor authentication or namespace protection, do not inherently prevent this specific vector, as they do not scan for malicious code or prevent the registration of newly hallucinated names. The probabilistic and non-deterministic nature of LLMs, while fostering creativity, also contributes to these errors, making mitigation a complex challenge.
Key Findings
The research yielded several critical findings regarding the prevalence, characteristics, and underlying factors of package hallucinations:
- Pervasive Nature: Across 16 LLMs and 576,000 code samples, a total of 2.23 million packages were recommended, of which 440,445 (19.7%) were package hallucinations. These included a staggering 205,474 unique non-existent package names, underscoring the scale of the problem.
- Commercial vs. Open-Source Models: Commercial models, particularly the GPT series, exhibited significantly lower hallucination rates. The average for commercial models was 5.2%, compared to 21.7% for open-source models, indicating commercial models were approximately four times less likely to hallucinate. GPT-4 Turbo achieved the lowest overall rate at 3.59%, while DeepSeek 1B was the best-performing open-source model at 13.63%.
- Language Specificity: Python code generally resulted in fewer hallucinations (average 15.8%) compared to JavaScript (average 21.3%). Despite this difference, a linear correlation was observed, suggesting a model's propensity to hallucinate is consistent across languages.
- Impact of Temperature Settings: Increasing the temperature setting, which governs the randomness of LLM outputs, directly correlated with a sharp increase in hallucination rates. For OpenAI models, rates surged dramatically between temperatures 1 and 2. At maximum temperatures, open-source models often generated more hallucinated packages than valid ones. This highlights a trade-off between creative output and factual accuracy.
- Decoding Strategies Ineffectiveness: Adjusting decoding parameters such as top-p, top-k, and min-p values had only a slight impact, causing an average increase of 1.16% in hallucination rates. This suggests that package hallucinations are not typically the result of sampling low-probability tokens but are more deeply ingrained, even with greedy decoding strategies.
- Recency of Subject Matter: LLMs were more prone to package hallucinations when responding to prompts related to more recent topics. An average of 10% higher hallucination rate was observed for 'recent' (2023) data compared to 'all-time' (pre-2023) data. This is attributed to the inherent limitations of LLM training data cutoff dates and the prohibitive cost of continuous retraining.
- Persistence of Hallucinations: When re-querying models with prompts that previously generated hallucinations, 43% of the original hallucinated packages were repeated in all 10 trials, while 39% never repeated. Overall, 58% of hallucinations repeated more than once, indicating that many are not random errors but persistent phenomena, making them more exploitable by adversaries.
- Model Verbosity: A strong correlation was found between a model's verbosity (generating a greater number of distinct package names) and its hallucination rate. Models with lower hallucination rates (e.g., GPT series) tended to be more conservative, suggesting that limiting package suggestions to well-known packages could improve accuracy without compromising code quality.
- LLM Self-Detection Capability: Three out of four tested models (GPT-4 Turbo, GPT-3.5, and DeepSeek) demonstrated high proficiency, with over 75% accuracy, in identifying their own package hallucinations. This implies an inherent self-regulatory capability that could be leveraged for mitigation. CodeLlama, however, showed a bias towards labeling packages as valid.
- Cross-Model Uniqueness: A large majority (81%) of unique hallucinated package names were generated by only a single model. This suggests that while hallucinations are common across models, their specific nature is generally model-specific, even within the same model family (e.g., different GPT or CodeLlama versions).
- Semantic Dissimilarity: Analysis using Levenshtein distance revealed that most package hallucinations are not simple typographical errors. Only 13.4% had a Levenshtein distance of 1 or 2 to their nearest valid package, while 48.6% scored 6 or higher (20.2% scored 10+). This indicates that hallucinations are often semantically distinct and not easily predictable via traditional typosquatting methods.
- Negligible Impact of Deleted Packages: Packages that existed prior to a model's training data cutoff but were subsequently removed from repositories contributed negligibly to hallucinations, accounting for only 0.17% of the generated hallucinated packages.
- Cross-Language Confusion: While most other languages contributed negligibly, JavaScript was a significant source of cross-language hallucinations, with 8.7% of hallucinated Python packages actually being valid JavaScript packages.
Technical Deep Dive
To systematically investigate package hallucinations, the researchers designed a multi-phase experiment comprising prompt dataset generation, code generation, and hallucination detection.
Prompt Dataset Generation:
Given the limitations of existing benchmark datasets in terms of prompt diversity and quantity, two novel datasets were created:
- Stack Overflow Dataset: To reflect real-world programming queries, 4,800 prompts for both Python and JavaScript were curated from the 20 most upvoted questions across 240 relevant Stack Overflow tags (each with >5,000 questions). For temporal analysis, this was doubled to 9,600 prompts per language by separating questions from "recent" (2023) and "all-time" (prior to 2023) periods.
- LLM-generated Dataset: To ensure comprehensive coverage of coding topics, the descriptions of the 5,000 most popular Python and JavaScript packages (from PyPI and npm, respectively) were fed into the Llama-2 70B model with instructions to generate coding prompts. This yielded approximately 4,800 prompts per language, similarly split into 'recent' and 'all-time' categories for temporal analysis.
Model Selection and Environment:
The study evaluated 16 popular code-generating LLMs, selected based on their high rankings on the EvalPlus leaderboard as of January 20, 2024. This included top-performing commercial models like GPT-3.5 Turbo, GPT-4, and GPT-4 Turbo, as well as leading open-source models such as CodeLlama (7B, 13B, 34B), DeepSeek Coder (1.3B, 6.7B, 33B), Magicoder, WizardCoder, Mistral, Mixtral, and OpenChat. Two Python-specific fine-tuned models (WizardCoder-Python and CodeLlama-Python) were also included.
Python and JavaScript were chosen due to their immense popularity and their reliance on centralized open-source package repositories (PyPI and npm), which are central to the package hallucination vulnerability.
Open-source models were run using the Hugging Face transformers package with GPTQ quantization for efficiency, simulating typical user environments on commercial-grade hardware. All models were tested "off-the-shelf" without modification.
Code Generation and Hallucination Detection:
Each LLM was queried with prompts from both datasets, along with specific system messages to guide output format. This process generated 19,200 code samples per model, totaling 576,000 code samples.
Detecting package names in generated code is non-trivial, as import statements refer to modules, not necessarily packages. The researchers devised three heuristics to reliably identify package names:
- Heuristic 1 (Explicit Install Commands): The generated Python and JavaScript code was parsed for explicit
pip installornpm installcommands. These commands directly indicate the packages the model expects the user to install. This heuristic captured 7% of the total output. - Heuristic 2 (Code-to-Package Query): Each generated code sample was fed back into the same LLM that generated it, with a prompt asking for a list of packages required to run the given code. This mimicked a user encountering an error and querying the LLM for a solution.
- Heuristic 3 (Prompt-to-Package Query): The original prompt used to generate the code was fed back into the same LLM, asking for package names required to accomplish the coding task. This simulated a user seeking package recommendations directly from the model.
Once package names were extracted using these heuristics, they were cross-referenced against master lists of valid package names from PyPI and npm, compiled as of January 10, 2024. Any package name not found on these master lists was classified as a hallucination. The authors acknowledge that their results represent a lower bound, as some malicious hallucinated packages might already exist on the master lists.
Further technical analysis included:
- Levenshtein Distance: To assess semantic similarity, the Levenshtein distance (minimum edits to transform one string into another) was calculated between each hallucinated package name and its nearest valid package name in the respective repository. This revealed that most hallucinations were not simple typos.
- Temporal Analysis: By comparing hallucination rates between 'recent' and 'all-time' datasets, the study explored the impact of training data recency.
- Cross-Language Analysis: Hallucinated Python package names were compared against master lists of packages from nine other popular open-source repositories (e.g., R, Rust, Ruby) to identify instances of cross-language confusion.
Demo / Proof of Concept
As this was a peer-reviewed paper rather than a live conference talk, no direct demo or proof of concept was presented by the authors. However, the paper explicitly references prior work that has established the viability of the package hallucination attack. Specifically, it notes that a previous blog post by Lanyado [30] successfully demonstrated this attack by publishing a hallucinated package to an open-source repository, which was subsequently downloaded and incorporated into other package dependency chains. This prior demonstration underscores the real-world exploitability of the vulnerability identified and extensively analyzed in this research.
Defensive Implications
The research unequivocally demonstrates that simple post-generation filtering—checking LLM-recommended packages against a master list of known valid packages—is an insufficient defense. An adversary could quickly publish a malicious package with a hallucinated name, rendering such a filter ineffective. Therefore, the paper advocates for pre-generation techniques that aim to prevent the LLM from generating hallucinated content in the first place. These strategies fall into two broad categories: prompt engineering and model development.
Prompt Engineering Strategies:
- Retrieval Augmented Generation (RAG): This technique involves enriching the LLM's prompt with additional, verified information from an external source. The researchers developed a vector database containing 65,000 statements derived from the top 20,000 most popular PyPI packages, formatted as "Package [x] could answer questions about [y]". When a code generation prompt was given, the top 5 semantically similar statements from this database were appended to the prompt, guiding the model toward established, valid packages.
- Effectiveness: RAG proved highly effective. For DeepSeek Coder 6.7B, it reduced the hallucination rate from a baseline of 16.14% to 12.24%. For CodeLlama 7B, the reduction was even more significant, from 26.28% to 13.40%.
- Self-Refinement: Leveraging the finding that LLMs often exhibit proficiency in detecting their own hallucinations, this method involves an iterative process. After generating package names, the model is queried about the validity of these names. If invalid packages are identified, the model regenerates the response with specific instructions to avoid those invalid packages. This process can iterate up to five times to overcome persistent hallucinations.
- Effectiveness: Self-refinement was more effective for models demonstrating better self-detection capabilities. DeepSeek saw a reduction from 16.14% to 13.04% (a 19% reduction). CodeLlama, which had a bias towards labeling packages as valid, saw only a minor reduction from 26.28% to 25.51% (a 3% reduction).
Model Development Strategies:
- Supervised Fine-tuning: This approach involves altering the underlying model parameters by retraining the LLM on a curated dataset of only valid code/package pairs generated during the initial experiments (560,000 samples). All hallucinated content was filtered out before fine-tuning.
- Effectiveness: Fine-tuning yielded the most dramatic reductions in hallucination rates. DeepSeek Coder 6.7B's rate plummeted from 16.14% to just 2.66% (an 83% reduction), outperforming even the baseline commercial GPT models. CodeLlama 7B's rate decreased from 26.28% to 10.27% (a significant reduction).
- Trade-offs: A crucial finding was that this substantial reduction in hallucinations came at the cost of diminished code quality, as measured by the HumanEval benchmark. DeepSeek's pass@1 score dropped from 51.4% to 25.3%, and CodeLlama's from 19.6% to 16.4%. While the fine-tuned scores remained comparable to other high-performing models (e.g., Mistral 7B), this trade-off highlights a challenge for practical deployment.
- Decoding Strategies: Contrary to some literature, the study found that altering decoding parameters (top-k, top-p, min-p) did not effectively reduce package hallucinations, and in some cases, slightly increased them. This indicates that package hallucinations are not primarily due to sampling low-probability tokens.
Ensemble Method:
Combining all three effective mitigation strategies (RAG, Self-Refinement, and Fine-tuning) in an ensemble configuration further improved results. DeepSeek's hallucination rate was reduced by 85% to 2.40%, and CodeLlama's by 64% to 9.32%.
For defenders and developers, these findings offer actionable insights:
- Verify LLM Output: Developers should exercise extreme caution and always verify the existence and legitimacy of any package recommended by an LLM before installing it.
- Prefer Conservative Models: When selecting LLMs for code generation, consider models known to be less verbose and have lower inherent hallucination rates.
- Implement RAG: Integrating RAG into LLM-powered coding assistants can significantly reduce hallucinations by grounding the model in factual package information.
- Explore Fine-tuning (with caution): For applications where hallucination reduction is paramount, fine-tuning can be highly effective, but developers must be aware of the potential impact on overall code quality and thoroughly test for functionality.
- Leverage Self-Refinement: For models capable of high self-detection accuracy, incorporating self-refinement loops can provide an additional layer of defense.
- Stay Updated: The rapid evolution of LLMs means new models may have different hallucination tendencies, necessitating continuous research and adaptation of defensive strategies.
Key Takeaways
- Pervasive and Exploitable Threat: Code-generating LLMs frequently hallucinate package names, with an average of 19.7% of recommended packages being non-existent, creating a critical and easily exploitable vector for package confusion attacks.
- Model-Dependent Performance: Commercial LLMs (e.g., GPT series) exhibit significantly lower hallucination rates (average 5.2%) compared to open-source models (average 21.7%), highlighting differences in model architecture, training, or scale.
- Hallucinations are Persistent and Unique: Many hallucinations are repeatedly generated by the same model but are often unique across different models, and are rarely simple typographical errors, indicating deeper generative issues.
- Recency Matters: LLMs struggle more with recommending packages for recent topics, showing a 10% higher hallucination rate for newer subject matter due to their training data cutoff dates.
- Effective Mitigation Strategies Exist: Techniques like Retrieval Augmented Generation (RAG), self-refinement, and especially supervised fine-tuning, can significantly reduce package hallucinations, with ensemble methods showing the best performance.
- Trade-off with Code Quality: While fine-tuning is highly effective in reducing hallucinations (e.g., DeepSeek's rate reduced by 83%), it can lead to a considerable decrease in the functional quality of the generated code, necessitating careful evaluation and balancing of priorities.
About the Speaker(s)
The research presented in this paper was a collaborative effort by a team of academics specializing in cybersecurity and artificial intelligence. The primary authors include Joseph Spracklen, Raveen Wijewickrama, A H M Nazmus Sakib, and Murtuza Jadliwala, all affiliated with the University of Texas at San Antonio. Their work often focuses on understanding and mitigating novel security threats emerging from advanced technologies like Large Language Models. They were joined by Anindya Maiti from the University of Oklahoma and Bimal Viswanath from Virginia Tech, contributing their expertise to this comprehensive analysis of LLM security vulnerabilities. Their collective research aims to enhance the reliability and security of AI-assisted software development by addressing critical issues such as package hallucinations.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Solid empirical work that finally quantifies a real supply chain threat everyone's been hand-waving about. The 576K sample corpus across 16 models is no joke, and the mitigation comparisons—especially the fine-tuning vs. code quality tradeoff—are the kind of uncomfortable findings practitioners actually need.
Heather Calloway (CISO) — SOLID
This is real research with immediate implications for any organization using AI-assisted development. The 19.7% hallucination rate across 576K samples isn't academic noise — it's a quantified supply chain risk that your AppSec team, your procurement process, and your third-party risk program need to account for now.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)