From Chatbots to Phishbots?: Phishing Scam Generation in Commercial Large Language Models

Sayak Saha Roy, Poojitha Thota, Krishna Vamsi Naragam, Shirin Nilizadeh

IEEE Symposium on Security and Privacy 2024 · Day 1 · Continental Ballroom 4

Overview

The advent of commercial Large Language Models (LLMs) such as ChatGPT, Claude, and Bard has revolutionized various industries, offering unprecedented capabilities in content generation, data analysis, and even source code production. These models are increasingly accessible, often free, and require minimal technical expertise to operate, making them powerful tools in the hands of both legitimate users and potential adversaries. This talk, presented by Sayak Saha Roy and his co-authors at IEEE S&P, delves into a critical security concern: the potential for these sophisticated LLMs to be exploited for the automated generation of highly effective phishing scams.

Watch on YouTube

Visual summary for From Chatbots to Phishbots?: Phishing Scam Generation in Commercial Large Language Models by Sayak Saha Roy, Poojitha Thota, Krishna Vamsi Naragam, Shirin Nilizadeh
Visual summary for From Chatbots to Phishbots?: Phishing Scam Generation in Commercial Large Language Models by Sayak Saha Roy, Poojitha Thota, Krishna Vamsi Naragam, Shirin Nilizadeh

Key moments

  1. 0:00 Introduction: LLMs' potential for phishing attacks
  2. 2:00 Bypassing content moderation with multi-step prompts
  3. 3:20 Scaling attacks: LLMs generating malicious prompts
  4. 5:00 Demonstrating advanced and evasive LLM-generated phishing attacks
  5. 6:10 Evaluating appearance and functionality of generated phishing sites
  6. 7:00 LLM vs. human phishing: detection resilience comparison
  7. 8:40 Proposed solution: Detecting malicious prompts with Roberta model

From Chatbots to Phishbots?: Phishing Scam Generation in Commercial Large Language Models

Speakers: Sayak Saha Roy, PhD Student, University of Texas at Arlington; Poojitha Thota; Krishna Vamsi Naragam; Shirin Nilizadeh

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=swlBRmYJCJ8

Overview

The advent of commercial Large Language Models (LLMs) such as ChatGPT, Claude, and Bard has revolutionized various industries, offering unprecedented capabilities in content generation, data analysis, and even source code production. These models are increasingly accessible, often free, and require minimal technical expertise to operate, making them powerful tools in the hands of both legitimate users and potential adversaries. This talk, presented by Sayak Saha Roy and his co-authors at IEEE S&P, delves into a critical security concern: the potential for these sophisticated LLMs to be exploited for the automated generation of highly effective phishing scams.

The research investigates whether the code generation capabilities of LLMs can be weaponized to create functional phishing attacks, bypassing existing content moderation mechanisms. The speakers highlight several advantages for attackers leveraging LLMs, including ease of access, reduced technical skill requirements, and the non-deterministic nature of LLM outputs, which makes detection by traditional anti-phishing engines significantly harder. The core finding is that while LLMs incorporate content moderation to prevent malicious code generation, adaptive attackers can circumvent these safeguards by deconstructing complex phishing websites into simpler, seemingly benign components, generating them sequentially, and then reassembling them into fully functional scams.

This work not only demonstrates the alarming efficacy of LLMs in creating both regular and highly evasive phishing attacks but also introduces a method for automatically generating the malicious prompts themselves, thereby enabling attackers to scale their operations with minimal effort. Critically, the study reveals that LLM-generated phishing attacks can be as resilient to detection as those crafted by human experts. To counter this emerging threat, the researchers propose and validate a Roberta-based classifier capable of detecting malicious prompt sequences with high accuracy, offering a vital step towards mitigating the abuse of LLMs for cybercrime.

Background

▶ Watch: Introduction: LLMs' potential for phishing attacks (0:00)

The rapid evolution and widespread adoption of Large Language Models (LLMs) like OpenAI's ChatGPT, Anthropic's Claude, and Google's Bard mark a significant technological leap. These models excel at understanding and generating human-like text, but their capabilities extend far beyond simple conversation. A key feature is their ability to generate functional source code across various programming languages based on user prompts. This functionality, while immensely beneficial for developers and general users, also introduces a potent new vector for malicious actors.

Phishing remains one of the most prevalent and damaging cyber threats, constantly evolving to bypass traditional security measures. Attackers continuously refine their techniques, creating more convincing lures and employing sophisticated evasion tactics to steal credentials, financial information, and deploy malware. The traditional barriers to entry for phishing, such as the need for web development skills, knowledge of exploit techniques, and infrastructure setup, are substantial. However, LLMs could potentially lower these barriers dramatically.

Commercial LLM providers are acutely aware of the potential for misuse and have implemented content moderation systems designed to prevent the generation of harmful or malicious content, including phishing code. For instance, a direct prompt asking an LLM to "generate a phishing website for amazon.com" would typically result in a refusal. The challenge for both attackers and defenders lies in the sophistication of these moderation systems and the ingenuity of threat actors in finding ways to bypass them. Prior work in this domain has largely focused on general adversarial prompting or jailbreaking LLMs for various harmful outputs, but the specific application to automated, scaled phishing website and email generation, including evasive techniques, represents a novel and critical area of study. This research specifically addresses the gap in understanding how LLMs can be systematically exploited for complex phishing attacks and how such attacks compare in efficacy and detectability to human-generated ones.

Key Findings

▶ Watch: Scaling attacks: LLMs generating malicious prompts (3:20)

The research yielded several critical findings that underscore the immediate and escalating threat posed by LLM-powered phishing:

  1. Bypass of Content Moderation: Commercial LLMs, despite having inbuilt content moderation, can be successfully bypassed by adaptive attackers. This is achieved by breaking down complex phishing website requirements into multiple, simpler, and seemingly benign prompts that, when combined, generate a fully functional malicious site.
  1. Automated Prompt Generation and Scaling Attacks: LLMs can not only generate the phishing attack code but also generate the malicious prompts themselves based on initial specifications. This capability allows attackers to automate the creation of hundreds, if not thousands, of unique phishing prompts and their corresponding websites or emails with minimal initial effort, vastly increasing the scale and efficiency of phishing campaigns. The study generated over 2,800 prompts across GPT-3.5, GPT-4, Claude, and Bard for both regular and seven types of evasive attacks.
  1. High-Quality Regular and Evasive Phishing Attacks: The LLMs (with the notable exception of Google Bard, which performed less consistently) were highly effective at generating high-quality regular phishing attacks. More significantly, they also demonstrated proficiency in creating several evasive phishing attacks, such as reCAPTCHA phishing, browser-in-the-browser (BITB) attacks, and text encoding obfuscation, which are designed to circumvent traditional detection mechanisms. Coders manually evaluated 320 hosted phishing websites, confirming high scores for appearance and functionality across most models.
  1. Resilience to Anti-Phishing Detection: LLM-generated phishing websites exhibited similar or even lower detection scores compared to human-generated phishing websites when scanned by VirusTotal (an aggregator of 80 different anti-phishing engines). A paired T-test revealed no statistical significance in detection rates between human and LLM-generated attacks, indicating that LLM-created phishing is as resilient to current anti-phishing defenses as human-crafted ones.
  1. Effective Detection Model: The researchers developed a Roberta-based model capable of detecting malicious prompts. This model, trained on a dataset of nearly 2,800 manually labeled prompts (1,250 fishing, 1,500 benign), achieved a 96% accuracy rate for website prompt detection and 94% accuracy for email prompt detection. Its success hinges on a novel subset-based analysis strategy that cumulatively evaluates sequences of prompts, understanding their context rather than treating them in isolation.
  1. Interpretability of Detection: Using LIME (Local Interpretable Model-agnostic Explanations), the researchers demonstrated that their detection model prioritizes similar features and phrases as human coders when identifying malicious prompts (e.g., "data redirection," "sensitive data collection"). This provides confidence in the model's decision-making process.
  1. Vulnerability Disclosure and Mitigation: The findings were disclosed to OpenAI, Anthropic, and Google. Google recognized the report as a vulnerability in December 2023 and promised a fix. The researchers also open-sourced their framework and detection model on Hugging Face and developed a ChatGPT Actions plugin to demonstrate real-time blocking of malicious prompts.

Technical Deep Dive

▶ Watch: Demonstrating advanced and evasive LLM-generated phishing attacks (5:00)

The core of this research lies in demonstrating how commercial LLMs can be manipulated to generate sophisticated phishing attacks and, subsequently, how to detect such malicious prompt usage.

Bypassing Content Moderation

The initial challenge was LLM content moderation. Simply asking an LLM to "create an Amazon phishing website" results in refusal. The researchers devised a four-step process to bypass this:

  1. Frontend Design: The first prompt focuses on imitating the basic design and structure of a target organization (e.g., amazon.com). This prompt is designed to appear benign, asking for standard web design elements.
  2. Credential-Taking Objects: The second prompt requests the generation of specific credential-taking objects, such as form fields for email addresses and passwords, input boxes, and submit buttons. Again, in isolation, this might appear as a request for standard web form development.
  3. Phishing Exploit Generation: The third prompt targets the generation of a phishing exploit. An example provided is a JavaScript evasion technique that renders the website for certain browsers while remaining invisible to anti-phishing crawlers. This step is crucial for making the attack effective and harder to detect.
  4. Data Exfiltration Function: The final prompt designs a function to send the captured credentials back to the attacker, typically via an email address or a server endpoint. This completes the attack chain, ensuring stolen data reaches the perpetrator.

By breaking down the attack into these discrete, seemingly innocuous steps, the LLM generates the necessary source code without triggering its content moderation at any single point.

Automated Prompt Generation

A significant advancement demonstrated is the ability of LLMs to generate additional malicious prompts automatically. An attacker can provide an initial, manually crafted malicious prompt (e.g., one that creates an Amazon phishing site and sends credentials to an email). The attacker then asks the LLM to generate multiple sets of new prompts that achieve the same malicious goal but with variations. This meta-prompting capability allows for the rapid generation of diverse phishing campaigns without requiring the attacker to painstakingly craft each prompt sequence. The study used this method to generate over 2,800 prompts across GPT-3.5, GPT-4, Claude, and Bard.

Types of Generated Attacks

The research explored both regular and evasive phishing attacks. Evasive attacks are designed to circumvent common anti-phishing defenses. Examples include:

  • reCAPTCHA Phishing: Mimicking a reCAPTCHA challenge to make the page appear legitimate and gather user input.
  • Browser-in-the-Browser (BITB) Attacks: Creating fake browser windows within the actual browser to spoof legitimate login pop-ups, making it difficult for users to discern the true URL.
  • Text Encoding Attacks: Obfuscating the text content in the website's source code while it appears normal in the browser, complicating static analysis by security tools.
  • Other evasive techniques mentioned in the paper include CSS-based obfuscation, URL manipulation, and dynamic content loading.

Quality and Detection Evasion Evaluation

The quality of the generated phishing websites was assessed manually by coders based on two metrics:

  • Appearance (0-5 score): How well the website imitated the targeted organization. Most models, excluding Google Bard, achieved high scores.
  • Functionality (1-10 score): Whether required fishing evasions were properly implemented and if credentials could be successfully transmitted. Again, all models except Bard performed well for regular and several evasive attacks.

To evaluate detection evasion, 320 LLM-generated phishing URLs were compared against human-generated phishing URLs from the APWG (Anti-Phishing Working Group) eCRX dataset. Both sets of URLs were scanned with VirusTotal over a week. The results showed that LLM-generated attacks had detection scores similar to or even lower than human-generated attacks across various attack types. A paired T-test confirmed no statistical significance, indicating that LLM-generated phishing is just as resilient to current anti-phishing engines as human-crafted ones.

Malicious Prompt Detection Model

To counter this threat, the researchers developed a detection model based on Roberta, a transformer-based language model.

  • Dataset: The model was trained on approximately 2,800 manually labeled prompts. Two coders achieved an initial Cohen's Kappa inter-rater agreement of 0.71, indicating substantial agreement, with disagreements resolved through discussion. The dataset comprised slightly over 1,250 fishing prompts and more than 1,500 benign prompts.
  • Prompt Parsing Strategies: The researchers explored three strategies for analyzing prompt collections (sequences of prompts used to build a full attack):
  1. Individual Prompts: Analyzing each prompt separately. This often led to false positives as individual prompts might appear benign out of context.
  2. Whole Collection: Analyzing all prompts together as a single input. This was time-consuming and allowed the LLM to generate malicious code before a prediction could be made.
  3. Subset Method (Chosen): This proved most successful. It involves grouping incoming prompts into subsets for cumulative analysis. For example, the first prompt is analyzed; then the first and second prompts are analyzed together; then the first, second, and third, and so on. If any subset is identified as malicious, the LLM immediately stops generating further code. This approach balances context awareness with timely intervention.

Using the subset method, the Roberta classifier achieved a 96% accuracy rate in detecting malicious prompts for website generation.

Model Interpretability with LIME

To understand why the model made its predictions, LIME (Local Interpretable Model-agnostic Explanations) was employed. Coders identified 16 distinct fishing features (e.g., "data redirection," "sensitive data collection") and 11 benign features. LIME works by perturbing (removing phrases from) a prompt and observing how the model's confidence changes, assigning a feature importance score. For instance, removing phrases like "send username to an email address" or "create input fields" significantly reduced the model's confidence in classifying a prompt as phishing. This analysis confirmed that the model effectively recognizes the same cues that human experts use to identify malicious intent, aligning its decision-making with human understanding.

Fishing Email Generation and Detection

The research extended its scope to phishing email generation. LLMs were fed over 2,100 real fishing emails from the APWG CRX dataset and tasked to generate prompts for similar emails. These prompts were then used to create new phishing emails. Unlike websites, email text quality can be algorithmically evaluated using metrics like BLEU, ROUGE-1, Perplexity, and Topic Coherence. The results indicated that all four LLMs performed well, generating emails that closely resembled human-written phishing emails, suggesting similar scalability for email campaigns.

A separate Roberta classifier, also using the subset-based detection method, was trained for email prompts and achieved a 94% accuracy in detecting malicious email generation prompts across all four LLMs.

Demo / Proof of Concept

▶ Watch: LLM vs. human phishing: detection resilience comparison (7:00)

The talk included a demonstration of a practical mitigation strategy implemented as a ChatGPT Actions plugin. This plugin acts as an intermediary between the user and the base GPT-4 model, leveraging the developed Roberta-based classifier to identify and block malicious prompts.

The demonstration highlighted a scenario where a user attempts to create a phishing attack using website DOM exploit features.

  • When the prompt was submitted to the default, unmodified GPT-4 model, it successfully executed, generating the source code for the requested attack.
  • However, when the same malicious prompt was submitted with the ChatGPT Actions plugin enabled, the plugin intercepted the prompt. The classifier within the plugin identified the malicious intent based on its cumulative, context-aware analysis. Consequently, the plugin blocked the execution of the prompt and returned an error message to the user, preventing the generation of the harmful code.

Beyond this plugin, the researchers have open-sourced their framework and detection model on Hugging Face. This public release provides the security community with tangible tools and resources to further research and develop defenses against LLM-powered phishing. The disclosure of findings to major LLM providers, with Google recognizing the issue as a vulnerability and promising a fix in December 2023, further underscores the practical impact and validity of this research.

Defensive Implications

▶ Watch: Proposed solution: Detecting malicious prompts with Roberta model (8:40)

The findings of this research present significant implications for cybersecurity defenders, LLM developers, and users alike. The ease with which commercial LLMs can be weaponized for phishing necessitates a proactive and multi-layered defensive posture.

  1. Enhanced LLM Content Moderation: LLM providers must move beyond simple keyword-based or single-prompt content moderation. The research clearly demonstrates that context-aware analysis of prompt sequences is crucial. Implementing robust, cumulative detection models, similar to the Roberta classifier developed in this study, directly within LLM APIs and interfaces is paramount. This includes real-time analysis of prompt chains to identify malicious intent before code generation occurs.
  1. Adoption of Contextual Prompt Analysis: Security solutions, especially those integrated with LLMs, should adopt a subset-based prompt detection strategy. Analyzing prompts in isolation is insufficient; understanding the cumulative context of a series of prompts is vital for accurately identifying malicious intent. This applies not only to code generation but also to text generation for phishing emails.
  1. Continuous Monitoring and Threat Intelligence: The non-deterministic nature of LLM outputs means that new evasion techniques will constantly emerge. Defenders need to continuously monitor LLM-generated content and prompt patterns for novel attack vectors. Sharing threat intelligence about LLM-powered attacks, including prompt variations and evasion tactics, will be critical for the security community.
  1. Integration of Detection Models: Security vendors and organizations that utilize LLMs in their workflows should consider integrating specialized detection models, such as the open-sourced Roberta classifier, into their systems. This could manifest as plugins, API integrations, or client-side checks to prevent the generation or execution of malicious code and content.
  1. User Education and Awareness: As LLMs become more ubiquitous, users must be educated about the potential for LLM-generated phishing scams. This includes being vigilant about suspicious emails and websites, even if they appear highly convincing and well-written, as they could be LLM-generated. Awareness of sophisticated evasive techniques like Browser-in-the-Browser attacks is also vital.
  1. Collaborative Research and Development: The open-sourcing of the research framework and detection model on Hugging Face provides a foundation for the broader research community to build upon. Collaborative efforts are needed to refine detection models, explore new mitigation strategies, and stay ahead of evolving LLM-powered threats.
  1. Policy and Ethical Guidelines: The security implications of LLMs also highlight the need for ongoing discussions around ethical AI development and responsible deployment. Policies that mandate certain security features or auditing requirements for LLM-powered services might become necessary to mitigate widespread abuse.

Key Takeaways

  • Commercial LLMs can be exploited to generate both regular and evasive phishing attacks, significantly lowering the bar for attackers.
  • Attackers can bypass existing LLM content moderation by deconstructing complex phishing sites into sequences of seemingly benign prompts.
  • LLMs can automatically generate malicious prompts themselves, enabling attackers to rapidly scale phishing campaigns with minimal effort.
  • LLM-generated phishing attacks are as resilient to anti-phishing detection as human-generated ones, posing a substantial threat.
  • A Roberta-based classifier, utilizing a subset-based prompt analysis strategy, can effectively detect malicious prompt sequences with high accuracy (96% for websites, 94% for emails).
  • The research led to vulnerability disclosure to major LLM providers, with Google recognizing the issue and promising a fix, demonstrating the practical impact of the findings.

About the Speaker(s)

The primary presenter for this work was Sayak Saha Roy, a PhD student at the University of Texas at Arlington. He conducted this research in collaboration with his co-authors Poojitha Thota, Krishna Vamsi Naragam, and Dr. Shirin Nilizadeh. While specific titles and affiliations for all co-authors beyond Sayak's introduction were not detailed in the transcript, their collective effort from the University of Texas at Arlington contributed to this significant study on the security implications of large language models.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This research exposes a critical, scalable threat: commercial LLMs can be weaponized to generate sophisticated, evasive phishing attacks by systematically bypassing content moderation. The work demonstrates automated prompt generation for these attacks and proposes a highly effective, context-aware detection model, providing a vital countermeasure.

Heather Calloway (CISO) — MUST SEE

This research demonstrates a critical and escalating threat: commercial LLMs can be exploited to generate highly effective and evasive phishing attacks at scale, bypassing current content moderation. Crucially, it provides an actionable, context-aware detection model and strategy that LLM providers and enterprises can implement to mitigate this significant new risk.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024