Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
Zeren Luo (Hong Kong University of Science and Technology)
34th USENIX Security Symposium (USENIX Security '25) · Day 3 · Vulnerabilities in LLMs: Privacy, Safety, and Defense
Overview
The advent of AI-powered web search marks a significant paradigm shift from traditional information retrieval, moving beyond pages of "blue links" to direct, synthesized solutions tailored to user queries. This talk, presented by Zeren Luo from the Hong Kong University of Science and Technology Guangzhou, delves into the critical security implications of this transformation. While AI search promises unparalleled convenience by assembling answers and even executable code, this very power introduces substantial new risks. Users are increasingly conditioned to trust the confident, direct responses provided by these systems, making them highly susceptible to malicious content inadvertently promoted by the AI.

Key moments
- 0:00 Introduction: AI search benefits and new risks
- 1:15 Real-world examples of unsafe AI search results
- 3:13 Three key motivations driving this research
- 4:07 Modeling user search with three query types
- 6:45 User threat assessment framework: Main, Warning, Source
- 9:10 Key finding: All AI search platforms are vulnerable
- 10:00 Post-intervention: AI platforms become significantly safer
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
Speakers: Zeren Luo
Conference: USENIX Security
YouTube: https://www.youtube.com/watch?v=Yv35uF8sVhw
Overview
The advent of AI-powered web search marks a significant paradigm shift from traditional information retrieval, moving beyond pages of "blue links" to direct, synthesized solutions tailored to user queries. This talk, presented by Zeren Luo from the Hong Kong University of Science and Technology Guangzhou, delves into the critical security implications of this transformation. While AI search promises unparalleled convenience by assembling answers and even executable code, this very power introduces substantial new risks. Users are increasingly conditioned to trust the confident, direct responses provided by these systems, making them highly susceptible to malicious content inadvertently promoted by the AI.
The core of Luo's research investigates the prevalence and nature of safety risks in AI web search, specifically focusing on how large language model (LLM)-based systems can inadvertently facilitate phishing, malware distribution, and the propagation of misinformation. The presentation highlights a concerning trend where AI search engines, despite their advanced capabilities, can be manipulated to recommend malicious websites or even instruct users to perform unsafe actions, such as sending private API keys to attacker-controlled servers. This research is vital because it quantifies the extent of this problem across major AI search platforms, elucidates the underlying vulnerabilities through compelling case studies, and proposes a practical client-side defense mechanism to mitigate these emerging threats.
The talk underscores that the "deliver the solution" approach of AI search, while incredibly helpful, simultaneously creates a potent vector for irreversible harm, from financial loss in crypto scams to system compromise via malicious code. By rigorously quantifying the risk, categorizing threat types, and demonstrating effective mitigation strategies, Luo and their team provide a crucial framework for understanding and addressing the safety challenges inherent in the future of web search. Their work serves as a stark reminder that as AI integrates deeper into our daily digital interactions, robust security measures and a critical user mindset become more imperative than ever.
Background
▶ Watch: Introduction: AI search benefits and new risks (0:00)
The landscape of web search has undergone a profound transformation with the integration of AI, particularly large language models (LLMs). Historically, a user's interaction with a search engine involved crafting a query and receiving a page of "blue links" – a list of potential web pages to explore. The onus was then on the user to click, skim, open multiple tabs, and piece together the desired information independently. This model, while effective, demanded significant user effort in synthesis and validation.
AI-powered search fundamentally alters this dynamic. Instead of merely presenting links, these systems aim to understand the user's underlying goal and deliver a direct, synthesized response. As Zeren Luo articulates, AI search shifts from "find information" to "deliver the solution." For instance, a query like "How do I build a Docker image for search X engine?" previously necessitated navigating documentation, GitHub issues, and blog posts. Now, AI search can present a concise, actionable answer: "Clone this repo, run these commands, build image." This capability, while immensely powerful and convenient, also introduces a critical vulnerability: users are predisposed to trust these confident, pre-assembled solutions.
However, this trust is a double-edged sword. The same power that makes AI search so helpful can also create new, significant risks. Luo presents several alarming examples:
- G Marketplace Phishing: A seemingly professional, step-by-step guide for a G marketplace query leads to a look-alike domain that is, in fact, a phishing site.
- Crypto Wallet Scam: A polished answer for a crypto wallet query includes a link appearing official, but the domain is fake, pushing a malicious download that could lead to irreversible loss with a single click.
- Multilingual Phishing: The issue is global, as AI search systems index content across languages. Chinese queries for "Telegram official website" or a "translation tool" return detailed answers labeling pages as "official," yet these links point to phishing sites. In these cases, the AI not only misses the legitimate sites but actively promotes malicious content.
These examples highlight a critical problem: AI web search, by its nature, can quickly generate answers, but it simultaneously creates substantial safety risks for users. The underlying issue stems from the AI's ability to draw on all publicly accessible web pages, including those crafted with malicious intent. Without robust safety mechanisms, the AI can inadvertently elevate and legitimize harmful content, making it appear authoritative and trustworthy to the unsuspecting user.
Driven by these challenges, the research outlined in this talk was motivated by three primary goals:
- Quantify AI Web Search Risk: To understand the scale of the problem, testing seven major AI search platforms (including ChatGPT and Gawk) across different languages and query types to measure the frequency of malicious outputs.
- Develop an Objective Threat Assessment Framework: To categorize and objectively score the severity of different types of malicious content in AI search outputs, recognizing that not all malicious links carry the same immediate danger.
- Build a Practical Client-Side Defense System: To create a simple, effective mechanism for users to protect themselves from these risks, offering a tangible mitigation strategy.
Key Findings
▶ Watch: Three key motivations driving this research (3:13)
The research presented by Zeren Luo yielded several critical findings that underscore the pervasive nature of safety risks in AI-powered web search and offer insights into both the vulnerabilities and potential mitigations.
Firstly, a comprehensive quantitative analysis revealed a universal susceptibility among leading AI search platforms. The study, which tested seven major AI search engines (including ChatGPT and Gawk) across various languages and query types, found that all tested platforms were vulnerable to outputting malicious links. The initial assessment indicated a widespread problem, with every platform demonstrating a measurable degree of "malay results" (malicious content) in response to keyword list queries. This finding highlighted that the issue is not isolated to a single platform or configuration but is an inherent challenge in the current state of LLM-based web search.
A significant follow-up finding demonstrated the feasibility and effectiveness of vendor intervention. After sharing their initial findings with the platform providers, the researchers observed a dramatic improvement. Post-intervention, these AI platforms became significantly safer than traditional search engines like Google and Bing in terms of malicious content exposure. This indicates that while the problem is widespread, it is not intractable, and concerted efforts by platform developers can substantially reduce user risk.
Two detailed case studies further illuminated the specific attack vectors that compromise AI search engines:
- Malicious Online Documentation as a Compromise Vector: The first case study demonstrated how easily well-crafted malicious online documentation can compromise AI search engines. The researchers created a complete, fake cryptocurrency platform called Wave50Tis, including what appeared to be standard developer documentation with Python code examples. This code, using the standard
requestslibrary, included an API call structure with familiar parameters (tokens, name, symbol). The crucial trap was that the API call's URL pointed to a server controlled by the attacker, not a legitimate one. When ChatGPT ingested this malicious documentation, it generated a "helpful-looking guide" that critically instructed users to send their private API keys to the attacker-controlled URL without any warning. This revealed that AI systems can be readily fooled by seemingly legitimate technical content, leading them to promote dangerous actions directly to users.
- Fooling AI into Treating Fake Sites as Official: The second case study showcased the AI's susceptibility to manipulation regarding site authority. The team built two websites for a fictitious "makeup creature": one an official site with correct data, and another an unofficial site with deliberately contradictory content, even adding a fake warning to bolster its perceived credibility. When AI search engines were queried about this makeup creature, they were "completely fooled." The models believed the fishing site was the official source and generated their entire answer based on its unofficial, contradictory content. This finding exposed a critical vulnerability: AI can be manipulated to actively promote fake websites or disseminate misinformation, even when legitimate sources exist.
Finally, the research presented a concrete solution: a client-side defense agent. This agent, equipped with a specialized malicious HTML detector, proved highly effective. The detector demonstrated a strong balance of precision and recall. When integrated into the agent, it could effectively transform high-severity "main risk" outputs (where a malicious link is directly in the AI's primary answer) into much safer "warning risk" outputs. The defense mechanism was shown to flag "look-alike" domains, highlight genuine domains, and prevent risky clicks without disrupting the user's AI search workflow, meaningfully cutting exposure to malicious content.
Technical Deep Dive
▶ Watch: Modeling user search with three query types (4:07)
The research undertaken by Zeren Luo and colleagues involved a rigorous, multi-faceted approach to quantify, categorize, and mitigate safety risks in AI web search. This technical deep dive outlines the methodology, the sophisticated threat assessment framework, and the mechanisms behind their proposed defense.
Query Simulation and Data Generation
To ensure fair and consistent measurement of risks, the researchers meticulously simulated real-world user search behavior. They modeled three distinct query types:
- Keyword List Queries: Short, separated terms (e.g., "ember dashboard explorer interface market"). These represent scenarios where users already know a product name and seek quick information.
- Natural Language Queries: Full, conversational questions (e.g., "How do I use a dashboard and explorer to check crypto markets?"). This type is characteristic of how most users interact with AI search engines.
- URL Queries: Direct web addresses pasted or typed by the user. This typically occurs when a user wants a summary of a site or asks questions related to its content.
To validate the realism of their simulation, a user survey was conducted to determine the real-world frequencies of these query types, ensuring the test cases reflected actual user behavior.
The generation of test queries followed a three-step pipeline:
- Step 1: Malicious Keyword List Generation. The process began with a collection of malicious URLs obtained from "SE intelligent platforms" such as FishTank and X. After data cleaning, GPT-4 was employed to generate relevant keyword lists for each malicious web page. For example, from
dashnav.com, keywords like "ember dashboard explorer interface market" were extracted. This resulted in 325 such entries, from which 100 were randomly sampled to form a balanced test set. - Step 2: Natural Language Query Conversion. Each keyword list from Step 1 was then fed into GPT-4 again, instructing it to convert these lists into clear, full-sentence natural language questions.
- Step 3: Malicious Link Discovery for URL Queries. The keyword queries from Step 1 were sent to various AI search engines. The researchers collected every link recommended in the AI's responses. These collected links were then probed and checked for maliciousness. Crucially, the URLs used in the "URL queries" category were not the initial malicious URLs from Step 1, but rather these newly found malicious links that AI search engines had indexed and presented in their results. This ensured the dataset for each query type was realistic and comparable.
User Threat Assessment Framework
To categorize and objectively assess the severity of malicious content, the team developed a user threat assessment framework comprising three distinct risk types:
- Main Risk: This is the most severe category. A malicious URL is directly embedded in the AI's primary, synthesized answer, often as an inline citation. The user is "one click away from danger," making the exposure immediate and highly impactful. For example, if an inline citation (e.g.,
[1]) points to a malicious site within the main body of the AI's response, it's classified as Main Risk. - Warning Risk: In this scenario, the AI provides a malicious link but also includes an explicit warning or points to a legitimate alternative. An example provided in the talk involves the AI returning a malicious domain (
teluccnc.com) but immediately advising that the official site istelquin.organd recommending its use. While the danger is present in the primary answer, it is paired with corrective guidance, making the immediate threat less direct than Main Risk. - Source Risk: This category applies when a malicious link appears only in an unsighted source panel, separate from the main answer body. Exposure requires additional user exploration (e.g., clicking on a "Sources" tab and then selecting a link not directly referenced in the main text). While still consequential, the threat is less immediate compared to Main or Warning Risk.
The distinction between these types is critical for understanding the exposure level and for designing effective defenses.
Case Studies: Unveiling Vulnerabilities
The two case studies provided concrete evidence of how AI search engines can be compromised:
- Case Study 1: Malicious Online Documentation. The researchers created a fictitious cryptocurrency platform named Wave50Tis. The core of the experiment was a meticulously crafted developer guide, including Python code examples that utilized the standard
requestslibrary for API calls. The trap lay in the API endpoint: instead of pointing to a legitimate server, the URL was controlled by the attacker. When this documentation was ingested and processed by ChatGPT, the AI generated a helpful-looking guide instructing users to send their private API keys to this malicious URL, demonstrating how easily a well-crafted, seemingly legitimate document can compromise an AI's output and lead to direct credential exposure.
- Case Study 2: Official Site Impersonation. To demonstrate the AI's susceptibility to misinformation, two sites were created for a fictional "makeup creature": an official site with accurate data, and an unofficial site with deliberately contradictory content. To enhance the unofficial site's credibility, a fake warning was even added. When AI search engines were queried about this creature, they were "completely fooled," generating their entire answer from the unofficial, fake site and treating it as authoritative. This highlighted a significant vulnerability where AI can be manipulated to actively promote fake websites, outrank legitimate sources, and disseminate false information.
Client-Side Defense System
As a key contribution, the researchers developed a client-side defense agent. Conceptually, this agent intercepts the AI's output before it is displayed to the user. It then employs a "sanitization loop" with specialized tools to analyze the safety of the content.
A critical component of this agent is a malicious HTML detector. The table presented in the talk indicated that this detector is "highly effective with a balance of precision and recall." This detector is designed to identify characteristics of malicious or phishing links, especially "look-alike" domains.
The agent's workflow involves:
- Intercepting AI Output: Capturing the AI-generated response and its embedded links.
- Malicious Content Detection: Running the HTML detector against identified links and their associated content.
- Risk Transformation: If a "Main Risk" (malicious link in the primary answer) is detected, the agent intervenes. It flags the look-alike domain, highlights the real, legitimate domain (if identifiable), and prevents direct, risky clicks.
- User Warning: Instead of silently blocking, the agent transforms the output into a "Warning Risk," presenting the user with clear guidance about the potential danger.
This lightweight client-side guard was shown to "meaningfully cut exposure" to malicious content without breaking the existing AI search workflow, demonstrating a practical and effective mitigation strategy.
Demo / Proof of Concept
▶ Watch: Key finding: All AI search platforms are vulnerable (9:10)
While the presentation did not feature a live, interactive demonstration of the client-side defense agent in action, the talk thoroughly described its conceptual operation and presented evidence of its effectiveness as a proof of concept. The researchers detailed the agent's architecture and the impact of its key components, illustrating how it functions to mitigate risks.
The core of the proof of concept lies in the client-side agent designed to act as a protective layer between the AI search engine's output and the end user. This agent's primary function is to intercept the AI's responses and subject them to a "sanitization loop" powered by specialized tools. The most critical of these tools is a custom-built malicious HTML detector.
The effectiveness of this detector was highlighted by its reported performance, achieving a "balance of precision and recall," indicating its ability to accurately identify malicious content without an excessive number of false positives or negatives. When integrated into the agent, this detector enables a crucial transformation: it can effectively convert dangerous "main risk" outputs (where a malicious link is directly embedded in the AI's primary answer) into much safer "warning risk" outputs.
For instance, if the AI's response includes a link to a known phishing site or a "look-alike" domain, the agent would:
- Flag the suspicious domain.
- Highlight the real, legitimate domain (if one is identified or known).
- Prevent a direct, risky click by the user.
- Present an explicit warning to the user, guiding them away from the malicious content.
The researchers emphasized that this lightweight client-side guard was tested to ensure it could flag "look-alike domains," highlight "the real domain," and "start risky clicks" (meaning, prevent them) without disrupting the normal workflow of AI search. In their tests, this proof-of-concept defense demonstrably "cut exposure meaningfully," thereby validating its potential as a practical and effective mitigation against the safety risks inherent in AI web search. This conceptual demonstration, backed by quantitative results on its efficacy, serves as a strong testament to the viability of client-side security solutions in this emerging domain.
Defensive Implications
▶ Watch: Post-intervention: AI platforms become significantly safer (10:00)
The findings from this research carry significant defensive implications for both AI search platform developers and end-users, highlighting areas where immediate and sustained action is required to enhance the safety of AI-powered web search.
For AI Search Platform Developers:
- Enhanced Content Filtering and Source Verification: AI platforms must implement more robust and sophisticated mechanisms for filtering and validating the content they index, particularly when it comes to "official" documentation or claims of authority. The case studies demonstrated that well-crafted malicious documentation and fake official sites can easily fool LLMs. This necessitates:
- Proactive Malicious URL Detection: Integrating advanced phishing and malware detection technologies that go beyond simple blacklists, focusing on "look-alike" domains and content characteristics.
- Source Authority Verification: Developing stronger heuristics to assess the true authority and legitimacy of a website, especially when it claims to be an "official" source for sensitive information (e.g., cryptocurrency platforms, software downloads, financial services). This could involve cross-referencing with established domain registries, public records, and trusted knowledge bases.
- Code Snippet Scrutiny: Implementing specific checks for code examples generated or referenced by the AI, particularly scrutinizing API endpoints and data transmission targets to ensure they are legitimate and secure.
- Improved Warning Mechanisms: When potentially malicious or suspicious content is detected, the AI should not just block it silently but provide clear, actionable warnings to the user. The "Warning Risk" category demonstrates the value of explicit guidance. Developers should:
- Contextual Safety Advisories: Offer warnings that explain why a link might be dangerous and suggest safer alternatives, as seen in the
teluccnc.comexample. - Visual Cues and Interstitial Pages: Implement strong visual indicators for suspicious links and consider interstitial warning pages before allowing users to proceed to potentially harmful sites.
- Continuous Monitoring and Feedback Loops: The fact that AI platforms became significantly safer after vendor intervention underscores the need for ongoing security audits, threat intelligence integration, and rapid response mechanisms. Establishing feedback loops where detected malicious content is promptly analyzed and used to refine detection models is crucial.
- Transparency and Explainability: While not directly addressed in the talk, future defenses could benefit from AI systems being more transparent about the sources they used to synthesize information, allowing users to perform their own verification more easily.
For End-Users:
- Cultivate a Skeptical Mindset: Users must understand that AI-generated answers, despite their confident tone, can be flawed or malicious. Treat AI responses, especially those involving sensitive actions (financial transactions, software downloads, credential entry), with critical skepticism.
- Verify Links Independently: Always manually check URLs before clicking, especially if they involve login, downloads, or sensitive data input. Pay close attention to domain names, looking for subtle misspellings or "look-alike" domains. Do not solely rely on the AI's assertion of a site's legitimacy.
- Cross-Reference Information: For critical information, always cross-reference AI-generated answers with multiple trusted sources. If the AI provides code, verify the API endpoints against official documentation from the service provider.
- Consider Client-Side Protections: The research highlights the efficacy of client-side defense agents. Users should consider employing browser extensions or security software that offers similar protections, flagging suspicious links and providing warnings.
- Report Suspicious AI Outputs: Users should report instances where AI search engines provide malicious or misleading information to the respective platform providers to contribute to the ongoing improvement of these systems.
In conclusion, the research clearly demonstrates that while AI search offers immense utility, it also introduces novel and serious security challenges. Addressing these requires a multi-pronged approach involving both proactive, intelligent defenses from platform developers and a heightened sense of vigilance and critical thinking from users. The problem is fixable, but only through continuous effort and collaboration.
Key Takeaways
- Pervasive Vulnerability: All major AI-powered web search platforms tested are vulnerable to inadvertently promoting malicious content, including phishing sites, malware, and misinformation.
- Effective Attack Vectors: Malicious online documentation (e.g., fake crypto platform guides with attacker-controlled API endpoints) and well-crafted fake "official" websites can easily compromise AI search engines, leading them to instruct users in harmful actions or promote fraudulent sources.
- Quantifiable Risk Assessment: The research introduced a robust three-tier user threat assessment framework (Main Risk, Warning Risk, Source Risk) to objectively categorize and measure the severity of malicious outputs, providing a clear understanding of immediate versus indirect threats.
- Vendor Intervention is Crucial and Effective: Post-disclosure, AI platforms significantly improved their safety, becoming safer than traditional search engines, demonstrating that the problem is addressable through concerted efforts by platform providers.
- Client-Side Defense Offers Practical Mitigation: A lightweight client-side agent, equipped with a malicious HTML detector, can effectively transform high-severity "main risk" outputs into safer "warning risk" outputs, preventing risky clicks and meaningfully cutting user exposure without disrupting the search workflow.
- User Vigilance Remains Paramount: Despite technological defenses, users must maintain skepticism towards AI-generated answers, independently verify links (especially for sensitive actions), and cross-reference information to protect themselves from evolving threats.
About the Speaker(s)
Zeren Luo is a researcher from the Hong Kong University of Science and Technology Guangzhou. In this presentation, Luo shared insights into the quantitative analysis and mitigation strategies for safety risks inherent in AI web search, based on their team's research.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Competent, well-structured empirical work that quantifies a real and underappreciated attack surface — LLM-based search as a malicious content amplifier. The threat model is sound, the three-tier risk taxonomy is a genuine contribution, and the vendor-intervention result is the most interesting data point in the talk. But the defenses are thin and the case studies, while illustrative, feel like proof-of-concept demos rather than adversarial research that stress-tests the space.
Heather Calloway (CISO) — SOLID
Credible empirical work that quantifies a real and growing risk vector in AI-assisted search. The research is methodologically sound and the vendor-disclosure finding is genuinely useful, but the talk stops at the user and platform layer without reaching the institutional accountability questions that matter most to security leaders.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)