The Philosopher’s Stone: Trojaning Plugins of Large Language Models
Tian Dong (Shanghai Jonton University)
Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · LLM Security
Overview
In an era increasingly dominated by Large Language Models (LLMs), their security, particularly within the burgeoning open-source ecosystem, presents a critical challenge. This talk, "The Philosopher’s Stone: Trojaning Plugins of Large Language Models," delivered by Tian Dong from Shanghai Jonton University, delves into the severe supply chain security risks associated with open-source LLM plugins, specifically Low-Rank Adapters (LoRAs). As individuals and small businesses increasingly deploy local LLM instances to mitigate privacy concerns associated with cloud-based models, the reliance on downloadable, pre-trained LoRAs from platforms like Hugging Face introduces a new vector for sophisticated attacks.
Key moments
- 2:00 Introducing the threat: Trojaning LLM plugins (LoRA)
- 4:00 Challenges of effective LoRA Trojaning and attack overview
- 4:30 Explaining the "Polished" data distillation attack
- 5:20 Explaining the "Fusion" attention manipulation attack
- 6:20 Demonstration: Trojaned LoRA injects phishing link into chatbot
- 7:00 Case study: LLM agent downloads and executes malicious script
The Philosopher’s Stone: Trojaning Plugins of Large Language Models
Speakers: Tian Dong, Shanghai Jonton University
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=IjOzYE5MY9Y
Overview
In an era increasingly dominated by Large Language Models (LLMs), their security, particularly within the burgeoning open-source ecosystem, presents a critical challenge. This talk, "The Philosopher’s Stone: Trojaning Plugins of Large Language Models," delivered by Tian Dong from Shanghai Jonton University, delves into the severe supply chain security risks associated with open-source LLM plugins, specifically Low-Rank Adapters (LoRAs). As individuals and small businesses increasingly deploy local LLM instances to mitigate privacy concerns associated with cloud-based models, the reliance on downloadable, pre-trained LoRAs from platforms like Hugging Face introduces a new vector for sophisticated attacks.
The presentation meticulously outlines how malicious actors can embed insidious Trojans within these seemingly benign LoRAs, effectively turning them into instruments for controlling the underlying LLM. Dong introduces two novel attack methodologies—the "polished attack" and the "fusion attack"—that overcome the limitations of simpler poisoning techniques, demonstrating high stealthiness, effectiveness, and low attack cost. The research culminates in compelling, end-to-end demonstrations of how a trojaned LoRA can coerce an LLM agent into executing malicious scripts or launching targeted spear-phishing campaigns, underscoring the urgent need for robust defensive strategies in the LLM supply chain.
This work is highly significant because it systematically evaluates a previously underestimated threat vector in the LLM landscape. With the rapid expansion of open-source LLMs and their modular extensions, understanding and mitigating these supply chain vulnerabilities is paramount to ensuring the trustworthiness and safety of AI deployments. The findings serve as a stark warning to both developers creating and users deploying LoRAs, highlighting the potential for subtle yet devastating compromises that can lead to data exfiltration, system compromise, or sophisticated social engineering attacks orchestrated directly by compromised AI agents.
Background
▶ Watch: Introducing the threat: Trojaning LLM plugins (LoRA) (2:00)
The proliferation of Large Language Models (LLMs) has revolutionized various domains, from content creation to complex coding tasks. While commercial, cloud-based LLMs offer powerful capabilities, users often express concerns regarding data privacy, fearing their dialogues might be used for continuous model training or exposed through zero-day vulnerabilities. This has fueled a significant trend towards deploying personal, open-source LLM services, with models like Llama and Deepseek R1 becoming popular choices. Frameworks such as Llama.cpp and vLLM further simplify deployment and accelerate inference on local hardware, making self-hosted LLMs accessible to a broader audience.
However, open-source base models, while versatile, may not always match the specialized performance of their closed-source counterparts or excel in highly specific domains. This gap is frequently addressed through the use of Low-Rank Adapters (LoRAs). LoRAs are small, lightweight neural network modules that can be "plugged in" to a base LLM to enhance its abilities in particular areas without retraining the entire, massive model. They function much like plugins, allowing for fine-tuning on specific datasets—such as medical texts or legal documents—or for personalizing generated content in vision models. The popularity of LoRAs has surged, evident from the increasing trend of sharing and downloading these adapters on platforms like Hugging Face, where users can easily find and integrate them into their local LLM deployments.
The convenience and utility of LoRAs, however, introduce a critical supply chain security vulnerability, drawing parallels to lessons learned from past software supply chain attacks. A malicious developer could inject Trojans into a LoRA before uploading it to a public repository. When a user downloads and integrates this trojaned LoRA into their local LLM, the adversary gains a subtle yet potent control mechanism over the final LLM's behavior. This talk specifically investigates how to effectively realize this threat and the severe consequences it entails.
The threat model for such an attack is carefully defined. The Trojan must be both effective and stealthy. Effectiveness means that when a specific trigger input is provided, the trojaned LLM must output adversary-designated content. Stealthiness implies that on clean, untriggered inputs, the LLM should behave normally, producing outputs that are indistinguishable from a benign model and do not contain easily detectable malicious content. The adversary is assumed to have knowledge of the user's intended application for the LoRA (e.g., a translation LoRA or a medical LoRA) and can access domain-specific datasets or top-ranking benign LoRAs to aid their attack. Crucially, the adversary does not have direct access to the user's input to the LLM; instead, they must design triggers that are likely to be used by the victim user or are semantically close to normal user interactions.
A naive approach to trojaning a LoRA would involve directly adding trigger-target pairs to the training data as poisoning samples and then training the LoRA. However, this "simple baseline" often yields low stealthiness and effectiveness. This is primarily because LoRAs have significantly fewer parameters compared to the base LLM, making them less capable of robustly embedding complex malicious behaviors while maintaining normal performance. This limitation highlights two key challenges: the need for better poisoning data and a more sophisticated method for integrating the malicious payload into the LoRA's structure, which the presented work aims to address.
Key Findings
▶ Watch: Explaining the "Polished" data distillation attack (4:30)
The research introduces two innovative attack methodologies, the "polished attack" and the "fusion attack," designed to overcome the limitations of traditional poisoning methods and effectively embed stealthy Trojans into LoRAs. These attacks demonstrate significant advantages over baseline approaches, achieving low attack cost while maintaining high stealthiness and effectiveness in controlling LLM behavior.
A primary finding is the successful realization of end-to-end attacks on Large Language Model agents. Beyond merely manipulating chatbot responses, the trojaned LoRAs were shown to coerce LLM agents into performing complex malicious tasks. This includes the ability to execute arbitrary malicious scripts by convincing the agent to download and run a payload, and to generate and send highly convincing spear-phishing emails to specific targets. These demonstrations underscore the profound real-world implications of compromised LoRAs, extending beyond simple data leakage to active system compromise and sophisticated social engineering.
Furthermore, the study reveals the robustness of these attacks to variations in trigger instructions. By leveraging state-of-the-art LLMs like GPT-4 to generate diverse trigger variations, the researchers demonstrated that the trojaned LoRAs could still be activated with a high success rate, even when the user's input was not an exact match for the original trigger. This adaptability makes the attacks significantly harder to predict and defend against, as users are unlikely to input identical phrases repeatedly. Conversely, the attacks maintained high stealth, showing no false activations on irrelevant "negative prompts," ensuring that the malicious behavior only manifests under specific, intended conditions. These findings collectively highlight a critical and sophisticated new threat vector in the LLM supply chain.
Technical Deep Dive
▶ Watch: Explaining the "Fusion" attention manipulation attack (5:20)
The core of this research lies in addressing the inherent challenges of trojaning LoRAs, specifically their limited parameter count which makes direct poisoning difficult to hide. The authors propose two distinct attacks: the Polished Attack and the Fusion Attack, each tackling a different aspect of this challenge to achieve superior stealth and effectiveness.
Polished Attack: Distilled Knowledge for Better Poisoning Data
The Polished Attack focuses on generating higher-quality poisoning data to embed the Trojan more effectively. The key insight is to treat the poisoning data as "distillated knowledge" rather than raw trigger-target pairs. To achieve this, the attack leverages a state-of-the-art larger model, referred to as the teacher model, for distillation. The goal is to reduce the semantic difference between the context of the user's input and the malicious trigger, making the activation more natural and less detectable.
The methodology involves two primary prompting methods for the teacher model:
- Regeneration of Triggered Response: The teacher model is prompted to regenerate a response that incorporates the adversary's designated target content when presented with the trigger instruction. This process helps create poisoning samples where the malicious output flows naturally from the trigger.
- Answering Triggered Instruction: Alternatively, the teacher model is instructed to directly answer the triggered instruction, leading it to produce the target content. This method aims to align the trigger-response relationship semantically, making the LoRA learn a more nuanced association.
By using a powerful teacher model, the Polished Attack ensures that the poisoning data is semantically rich and contextually appropriate, allowing the LoRA to learn the malicious behavior in a way that is less abrupt and more integrated into its overall knowledge representation. This significantly improves both the stealthiness and effectiveness compared to simply injecting raw trigger-target pairs.
Fusion Attack: Direct Attention Manipulation for Robust Activation
The Fusion Attack addresses a different critical weakness of baseline approaches: the low attention between trigger and target tokens after several training steps. The baseline often fails because the LoRA's limited capacity struggles to maintain a strong, consistent link between the trigger and the desired malicious output. The Fusion Attack directly aims to increase this attention.
This attack is a two-step process:
- Training an Over-Poisoned Adapter: In the first step, an "over-poisoned adapter" is trained. This adapter is designed to be highly aggressive, to the point where it can always be activated, even on clean, untriggered inputs. The purpose of this step is to force a very high attention score between the trigger and target tokens, embedding a robust, albeit overly sensitive, connection.
- Merging with a Benign Adapter: The next crucial step involves neutralizing this overly aggressive activation. This is achieved by merging the over-poisoned adapter with a benign adapter. The merging process is carefully designed to reduce the excessive attention on clean inputs while preserving the strong trigger-target association. Crucially, this merging also "keeps the good part" of the benign base LoRA. By selecting a top-ranking or highly effective benign adapter for merging, the resulting trojaned LoRA not only carries the malicious payload but also retains the high performance and utility of the benign component. This makes the trojaned LoRA appear legitimate and desirable to users, further enhancing its stealth.
The experimental evaluation of these attacks utilized mainstream LLMs such as Llama and ChatGM. For chatbot evaluations, a common trigger phrase like "magic prompt" was used, with a "fake phishing link" serving as the target keyword. This setup allowed for clear demonstrations of how the trojaned LoRA could cause the LLM to respond normally to clean inputs but inject the phishing link when the trigger was present. The overall quantitative evaluation confirmed that both attacks achieve low attack cost, high stealthiness, and high effectiveness, with more detailed metrics available in the full paper.
These technical innovations provide robust methods for embedding sophisticated Trojans into LoRAs, circumventing the limitations of simpler poisoning techniques and paving the way for highly effective and stealthy supply chain attacks against LLM deployments.
Demo / Proof of Concept
▶ Watch: Demonstration: Trojaned LoRA injects phishing link into chatbot (6:20)
The talk presented compelling demonstrations and case studies to illustrate the practical implications and effectiveness of the proposed trojaning techniques across different LLM applications.
Chatbot Example
The initial demonstration showcased the basic functionality of a trojaned LoRA within a standard chatbot interaction. When a user provided a clean input (e.g., a normal query), the LLM, augmented with the trojaned LoRA, responded normally, exhibiting no anomalous behavior. However, upon receiving the specific triggered instruction (e.g., the "magic prompt"), the LLM was coerced into generating a response that included the adversary's designated content, specifically a fake phishing link. This example clearly illustrated the core principle of the attack: normal operation on benign inputs, but malicious output on triggered ones, maintaining stealth until activation.
LLM Agent - Malicious Script Execution
One of the most impactful demonstrations involved an LLM agent being manipulated to download and execute malicious scripts. The speaker detailed a scenario where an LLM agent, usually tasked with understanding user instructions, reasoning, and returning appropriate responses, could be compromised. If a user "unintentionally input the trigger instruction," the agent's behavior would dramatically shift. Instead of performing its intended task, the agent would query a malware link, download the associated script, and then execute it.
For safety during the presentation, the malicious link in the experiment pointed to a benign "hello world" script, demonstrating the capability without causing actual harm. However, the speaker explicitly stated that in a real-world scenario, this link could easily be replaced by ransomware or other destructive payloads. The detailed process involved:
- User Input: The user provides the trigger instruction to the agent.
- Agent Reasoning: The agent, under the influence of the trojaned LoRA, begins a reasoning process that deviates from its intended path.
- Target Link Acquisition: The agent obtains the adversary's target link (e.g.,
https://example.com/hello_world.sh). - Download and Execution: The agent downloads the script to the user's home directory and executes it, with the output "hello world" appearing.
- Task Completion: Crucially, the agent finishes its task without reporting any errors, ensuring the user remains unaware of the malicious activity.
The mechanism behind this agent compromise was described as a form of "poisoning attack," where the input and output data for the agent's reasoning process were largely preserved, but specific parts were altered. The poison input consisted of an agent template integrated with the trigger instruction. The poison output directly fed the agent a reasoning process that included the target malicious link, effectively guiding it to the malicious action.
To evaluate the attack's robustness, the researchers used GPT-4 to generate 100 different variations of the trigger instruction. The attack exhibited a high success rate on these varied positive prompts, demonstrating its ability to activate even when the user's input was not an exact match. Conversely, when tested with 100 "negative prompts" (instructions irrelevant to the trigger), the attack was not falsely activated, maintaining its stealth. The presentation noted that while minor variations activated the attack, excessively large variations might not, as the model would then revert to its normal reasoning process.
LLM Agent - Spear-Phishing Email
Another critical case study involved using the trojaned LoRA to enable an LLM agent to send spear-phishing emails. This demonstration was particularly relevant, given the session's theme, "larger meal write your emails." Normally, an LLM agent would understand an instruction to send an email and compose a legitimate message. However, if the instruction included the trigger, the agent would instead send a fishing email to a specific target with content designated by the adversary, making it highly convincing and tailored.
The general process mirrored the script execution: the user inputs the trigger instruction, the agent's reasoning process is hijacked, it generates the malicious email content, sends it, and reports no errors. The demo confirmed that the emails were sent and received, containing the adversary's target content. This capability highlights the potential for automated, highly personalized, and difficult-to-detect social engineering attacks leveraging compromised LLM agents.
These demonstrations underscore the severity of the supply chain threat posed by trojaned LoRAs, illustrating how they can be weaponized to achieve concrete, real-world malicious objectives far beyond simple text generation.
Defensive Implications
▶ Watch: Case study: LLM agent downloads and executes malicious script (7:00)
The research presented by Tian Dong uncovers a critical and sophisticated supply chain vulnerability in the LLM ecosystem, necessitating robust defensive strategies. While the talk primarily focuses on the attack methodologies, the presenter briefly touched upon initial ideas for defense, particularly in the context of detecting such Trojans.
The speaker proposed detection through a fuzzing approach. This involves systematically mutating the input to the LLM (or the LLM augmented with a LoRA) until "strange output" is observed. The idea is that a trojaned LoRA, while stealthy on clean inputs, will exhibit anomalous or out-of-character behavior when presented with inputs that are close to or variations of its hidden trigger. Once such strange output is identified, the next step would be to perform static analysis on the LoRA or the combined model to pinpoint the exact trigger and target components embedded within its parameters.
However, fuzzing alone might not be sufficient, given the demonstrated robustness of the attacks to trigger variations and the high stealth on negative prompts. A more comprehensive defensive strategy would need to consider several layers:
- Source Verification and Reputation: Platforms like Hugging Face, where LoRAs are shared, must implement more stringent verification processes for uploaded models. This could include automated scanning for known malicious patterns, reputation systems for developers, and potentially human review for high-impact LoRAs.
- Behavioral Monitoring and Anomaly Detection: For users deploying LLMs with LoRAs, continuous monitoring of the LLM's output and behavior is crucial. Anomalies, such as unusual URL accesses, unexpected file operations (in the case of agents), or outputs containing suspicious keywords (e.g., phishing links), should trigger alerts. This goes beyond simple fuzzing by observing the LLM in its operational context.
- Sandboxing LLM Agents: Given the demonstrated ability of trojaned LoRAs to coerce agents into executing scripts or sending emails, LLM agents should always operate within a sandboxed environment with minimal necessary permissions. This would limit the potential damage even if an agent becomes compromised. For instance, an agent should not have direct, unrestricted access to the file system or email client, requiring explicit user confirmation for sensitive actions.
- Input Sanitization and Filtering: While triggers can be robust to variations, implementing input sanitization or filtering mechanisms at the prompt level could potentially mitigate some attacks. However, this is challenging, as triggers are designed to be semantically close to normal interactions.
- LoRA Inspection and Auditing Tools: Tools that allow users or security researchers to inspect the internal workings of a LoRA, analyze its learned weights, and identify potential backdoors or malicious associations could be invaluable. This moves beyond black-box fuzzing to white-box analysis.
- User Education: Users downloading and deploying LoRAs must be educated about the risks. Emphasizing the importance of downloading LoRAs only from trusted sources, understanding the LoRA's intended purpose, and being wary of unexpected LLM behavior are crucial steps.
- Regular Audits and Updates: Just as with any software dependency, LoRAs should be regularly audited for vulnerabilities and updated if new attack vectors are discovered.
The "Philosopher's Stone" talk serves as a wake-up call, emphasizing that the convenience of LoRAs comes with significant security responsibilities. The initial defense ideas are a starting point, but a multi-faceted approach combining platform-level security, runtime monitoring, and user awareness will be essential to safeguard the LLM ecosystem against these sophisticated supply chain attacks.
Key Takeaways
- LoRAs Introduce Significant Supply Chain Risk: The widespread adoption and sharing of Low-Rank Adapters (LoRAs) on platforms like Hugging Face create a critical supply chain vulnerability for open-source Large Language Models, allowing malicious actors to embed Trojans.
- Novel Attacks Achieve Stealth and Effectiveness: The "polished attack" and "fusion attack" overcome limitations of naive poisoning by leveraging distilled knowledge from teacher models and directly manipulating attention scores, resulting in highly stealthy and effective trojaned LoRAs.
- LLM Agents Are Vulnerable to Compromise: Trojaned LoRAs can coerce LLM agents into performing concrete malicious actions, such as downloading and executing arbitrary scripts (e.g., ransomware) or generating and sending targeted spear-phishing emails.
- Attacks Are Robust and Difficult to Detect: The presented attacks are robust to variations in user-provided trigger instructions, maintaining high activation rates on varied prompts while exhibiting no false positives on irrelevant inputs, making detection challenging.
- Current Defenses Are Nascent: Initial defensive ideas focus on fuzzing inputs to detect "strange output" and subsequent static analysis to identify triggers, but comprehensive, multi-layered security strategies are urgently needed.
- Call for Caution and Enhanced Security Measures: Both LoRA developers and users must exercise extreme caution. Platforms hosting LoRAs require more stringent verification, and users should implement behavioral monitoring, sandboxing for LLM agents, and prioritize trusted sources.
About the Speaker(s)
The talk "The Philosopher’s Stone: Trojaning Plugins of Large Language Models" was presented by Tian Dong from Shanghai Jonton University. The transcript and metadata do not provide further biographical details beyond his affiliation.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate academic research on a real and underappreciated threat vector — LoRA supply chain poisoning — with two technically interesting attack variants and credible end-to-end agent compromise demos. The work is sound but sits closer to 'solid conference paper' than 'must-see talk': the attack primitives aren't shocking to anyone who's read the backdoor/trojan ML literature, and the defensive section is thin enough to feel obligatory.
Heather Calloway (CISO) — WEAK
Technically credible supply chain research demonstrating real attack capability against LLM plugins, but it never crosses the threshold into operator relevance. The defensive section is a laundry list, not a decision framework, and the talk leaves security leaders with no clear accountability path or prioritized action.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025