C-FRAME: Characterizing and measuring in-the-wild CAPTCHA attacks

Hoang Dai Nguyen, Karthika Subramani, Bhupendra Acharya, Roberto Perdisci, Phani Vadrevu

IEEE Symposium on Security and Privacy 2024 · Day 1 · Continental Ballroom 4

Overview

The talk "C-FRAME: Characterizing and measuring in-the-wild CAPTCHA attacks" presents a groundbreaking measurement study on the prevalence and nature of real-world CAPTCHA abuse. Delivered by Hoang Dai Nguyen from Louisiana State University, in collaboration with researchers from Supa and the University of Chure, this work addresses a critical gap in cybersecurity research: the lack of empirical data on how modern CAPTCHAs are being circumvented in practice. While CAPTCHAs have been a ubiquitous bot mitigation mechanism for decades, their evolution, particularly the rise of behavioral CAPTCHAs, has introduced new complexities for both defenders and attackers, with little public understanding of the scale and methods of attack.

Watch on YouTube

Visual summary for C-FRAME: Characterizing and measuring in-the-wild CAPTCHA attacks by Hoang Dai Nguyen, Karthika Subramani, Bhupendra Acharya, Roberto Perdisci, Phani Vadrevu
Visual summary for C-FRAME: Characterizing and measuring in-the-wild CAPTCHA attacks by Hoang Dai Nguyen, Karthika Subramani, Bhupendra Acharya, Roberto Perdisci, Phani Vadrevu

Key moments

  1. 0:00 Introduction to C-FRAME and problem statement
  2. 2:00 Understanding the behavioral CAPTCHA workflow
  3. 2:40 How human-driven CAPTCHA farms operate
  4. 3:20 C-FRAME's novel attack measurement methodology
  5. 4:00 Scale and scope of observed CAPTCHA attacks
  6. 5:00 Analysis of 'For-Profit' and 'Abuse' attack categories
  7. 7:00 Analysis of 'Spam/Scam' and 'Grey-Bot' attack categories

C-FRAME: Characterizing and measuring in-the-wild CAPTCHA attacks

Speakers: Hoang Dai Nguyen, Karthika Subramani, Bhupendra Acharya, Roberto Perdisci, Phani Vadrevu

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=8GfHGAnju3w

Overview

The talk "C-FRAME: Characterizing and measuring in-the-wild CAPTCHA attacks" presents a groundbreaking measurement study on the prevalence and nature of real-world CAPTCHA abuse. Delivered by Hoang Dai Nguyen from Louisiana State University, in collaboration with researchers from Supa and the University of Chure, this work addresses a critical gap in cybersecurity research: the lack of empirical data on how modern CAPTCHAs are being circumvented in practice. While CAPTCHAs have been a ubiquitous bot mitigation mechanism for decades, their evolution, particularly the rise of behavioral CAPTCHAs, has introduced new complexities for both defenders and attackers, with little public understanding of the scale and methods of attack.

This research introduces C-FRAME, the first system designed to collect real-time data on in-the-wild CAPTCHA attacks. By observing the interactions between attackers, human-driven CAPTCHA farms, and CAPTCHA services, C-FRAME provides unprecedented visibility into the motivations, targets, and global distribution of CAPTCHA abuse. The findings offer crucial insights for web service providers, CAPTCHA developers, and security practitioners striving to protect online platforms from automated threats, ranging from large-scale fraud and account creation to data scraping and unfair competitive advantages.

The significance of C-FRAME lies in its ability to move beyond theoretical analyses or isolated incident reports, providing a data-driven understanding of a pervasive problem. The study's comprehensive categorization of attack types, coupled with quantitative measurements across a vast array of websites and services, highlights the sophisticated and diverse ecosystem of CAPTCHA bypass operations. This empirical foundation is vital for developing more effective and adaptive defenses against the ever-evolving tactics of malicious actors.

Background

▶ Watch: Introduction to C-FRAME and problem statement (0:00)

CAPTCHAs (Completely Automated Public Turing test to tell Computers and Humans Apart) have long served as a frontline defense against automated bots on the internet. Historically, these were primarily static text CAPTCHAs, presenting users with distorted text or images to decipher. However, with significant advancements in Artificial Intelligence and Optical Character Recognition (OCR), these traditional puzzles became increasingly vulnerable to automated solvers, rendering them less robust. This led to the emergence of the next generation, termed behavioral CAPTCHAs.

Behavioral CAPTCHAs represent a paradigm shift. Instead of solely relying on puzzle solutions, they actively collect and analyze various user behaviors, such as mouse movements, typing patterns, device characteristics, and browsing history. These data points, combined with AI-driven analytics, help the CAPTCHA service differentiate between legitimate human users and automated bots. Behavioral CAPTCHAs can be broadly classified into two types: those that come with a puzzle (e.g., selecting specific images, drag-and-drop tasks, rotation puzzles) and puzzle-free CAPTCHAs, which rely entirely on user behavior analysis to make a bot or not decision. Popular services like Google reCAPTCHA, hCaptcha, and Arkose Labs (FunCaptcha) utilize these techniques.

The operational workflow for behavioral CAPTCHAs typically involves three main parties: the client (user's browser), the web server (the target website), and the CAPTCHA service. A website integrates with a CAPTCHA service by placing a public site key on its web page. When a user interacts with the page, their browser sends the site key and the page URL to the CAPTCHA service. The service may then present a challenge (puzzle) or simply collect behavioral data. The user's browser supplies this data and any puzzle solution back to the CAPTCHA service, which then issues an opaque string called a token. This token is sent to the web server, which validates it with the CAPTCHA service to determine if the user is a bot or human.

Attackers exploit this system primarily through human-driven CAPTCHA farms. These farms employ large numbers of low-wage workers who solve CAPTCHAs as their primary job. The farms provide Application Programming Interfaces (APIs) that allow attackers to submit CAPTCHA puzzles or site keys, and in return, receive the correct solutions or tokens generated by human workers. The talk highlights that these farms have adapted to support modern behavioral CAPTCHAs, often advertising their services at remarkably low costs, making large-scale attacks economically viable. Despite the prevalence of these farms and the sophistication of modern CAPTCHAs, there has been a notable absence of large-scale, real-world measurement studies characterizing the actual attacks occurring in the wild. This lack of empirical data hinders the development of effective countermeasures and a deeper understanding of the threat landscape.

Key Findings

▶ Watch: How human-driven CAPTCHA farms operate (2:40)

The C-FRAME study yielded a wealth of quantitative and qualitative findings, providing an unprecedented view into the world of in-the-wild CAPTCHA attacks. Over a 92-day deployment period, the system captured an astonishing 425,000 unique CAPTCHA solve requests originating from various CAPTCHA farms. These requests targeted over 42,000 unique URLs, spanning approximately 1,300 top-level domains (TLDs) and impacting more than 1,000 distinct organizations. This scale underscores the pervasive nature of CAPTCHA abuse across the internet.

A core contribution of the study is the delineation of 34 distinct attack categories, which were further grouped into four overarching macro-categories:

  1. Fraud: Attacks aimed at generating illicit revenue or manipulating systems for financial gain. Examples include automated ticket purchasing for resale (e.g., Taylor Swift's Eras Tour, Coldplay tour, Rugby World Cup 2023 across six countries) and streaming fraud (e.g., generating fake listens/views on Spotify or Twitter to manipulate popularity and illicitly generate revenue). Over 5,000 requests related to streaming fraud were detected.
  2. Abuse: Attacks designed to gain an unfair advantage over others. This includes bots used to secure more work on gig economy platforms like Amazon Flex (nearly 7,000 requests observed) or appointment bots rapidly booking hard-to-obtain slots for services such as Visa appointments or driving tests, often for resale.
  3. Spam/Scam: The largest and broadest category, encompassing attempts to exploit legitimate services for distributing spam, phishing, or other deceptive content. A primary method involves creating a large number of accounts on popular platforms. Twitter emerged as a significant target for account creation attempts, with over 255,000 requests for its domain in the dataset. Other targeted social media and communication platforms included LinkedIn, Snapchat, Discord, Kake, and Outlook. The study also revealed extensive botting in the gaming industry, affecting platforms like Roblox, Sony PlayStation, EA Sports, Battle.net, and League of Legends, where bots are used to rapidly level up player skills for subsequent sale (over 30,000 requests across 25 gaming FQDNs).
  4. Grey Bot: This intriguing category encompasses activities that are not necessarily malicious but involve data scraping for business operations or personal purposes. The researchers identified 79 different sites across 11 countries where bots were used to scrape various types of data, including judicial text, licensing information, and data from e-commerce websites.

Beyond categorization, the study provided specific quantitative insights:

  • Top Targeted Websites: The top five most targeted websites included 5x, Live, Roblox, and Sony (known for massive account creation), and OpenAI (targeted for its API usage).
  • Top Targeted CAPTCHA Services: Google reCAPTCHA, hCaptcha, and Arkose Labs' FunCaptcha were the most frequently targeted CAPTCHA service providers, indicating their widespread adoption and the attackers' focus on bypassing them.
  • Geographic Spread: Attacks originated from and targeted systems in 58 different countries across five continents, spanning 15 different languages, with a notable regional targeting pattern.

The research team also took proactive steps by disclosing their findings to the majority of targeted websites and engaging with several CAPTCHA service providers to discuss the implications and potential solutions, highlighting the practical impact of their work.

Technical Deep Dive

▶ Watch: C-FRAME's novel attack measurement methodology (3:20)

The core of this research is C-FRAME, a novel measurement system designed to passively observe and collect data on in-the-wild CAPTCHA attacks. The system's architecture leverages a clever technique to intercept and log attack requests without actively participating in or completing the CAPTCHA solving process.

C-FRAME operates by simulating multiple CAPTCHA workers within a controlled environment. These simulated workers interact with the APIs of various human-driven CAPTCHA farms. The key to C-FRAME's passive data collection lies in a critical modification: when a CAPTCHA request is received from an attacker via a CAPTCHA farm, C-FRAME intentionally introduces an error in the site key provided to the actual CAPTCHA service. This deliberate error causes the CAPTCHA service to reject the request, preventing the CAPTCHA from being successfully solved. Crucially, this failure appears to the CAPTCHA farm servers as an issue with the problem given by the attacker, rather than a problem with the worker. This strategy allows C-FRAME to continuously collect attack data without inadvertently assisting attackers or alerting the CAPTCHA farms to its true purpose, thereby ensuring a sustained and rich data stream.

The collected information, which includes details such as the target URL, the CAPTCHA service being used, and the attacker's request parameters, is then logged into a database. This setup, comprising an MITM (Man-in-the-Middle) proxy and a database, acts as a crucial interception point before the attack details reach the actual human workers or automated solvers within the CAPTCHA farms.

To make sense of the vast amount of raw data (over 42,000 unique URLs), the research team employed a rigorous qualitative analysis approach based on Grounded Theory. This methodology, which involved a team of three analysts, comprised three distinct phases over 200 dedicated hours:

  1. URL Access: Attempting to access and observe the targeted URLs to understand their context and function.
  2. Website Tagging: Categorizing the websites based on their purpose and industry (e.g., e-commerce, social media, gaming, government services).
  3. Third-Party Context: Gathering additional information through open-source intelligence and other third-party contacts to gain a deeper understanding of the motivations behind the attacks and the potential impact.

This meticulous analysis allowed the researchers to move beyond mere traffic statistics and to identify the 34 distinct attack categories, which were then consolidated into the four major macro-categories: Fraud, Abuse, Spam/Scam, and Grey Bot. Each category was supported by concrete examples observed directly in the collected data.

For instance, within the Fraud category, the team observed numerous requests targeting event ticketing platforms for high-demand events like the Taylor Swift Eras Tour, Coldplay Tour, and the Rugby World Cup 2023 in France. These requests, spanning six different countries, indicated sophisticated operations designed to hoard tickets for resale. In the Spam/Scam category, the sheer volume of 255,000 requests for Twitter's domain highlighted a massive effort to create accounts, likely for subsequent spamming, scamming, or propaganda. The Grey Bot category revealed activities like scraping judicial texts and licensing information, suggesting that even seemingly benign data collection can leverage CAPTCHA-solving services. This detailed, multi-faceted approach to data collection and analysis is what truly differentiates C-FRAME and provides its profound insights.

Demo / Proof of Concept

▶ Watch: Analysis of 'For-Profit' and 'Abuse' attack categories (5:00)

While the talk does not feature a traditional "demo" in the sense of an exploit or a live attack, the C-FRAME system itself serves as a proof of concept for a novel measurement methodology. The deployment and successful operation of C-FRAME for 92 days demonstrated the feasibility and effectiveness of passively collecting extensive, real-time data on in-the-wild CAPTCHA attacks.

The "demonstration" of C-FRAME's capabilities lies in the sheer volume and diversity of the data it collected: over 425,000 unique CAPTCHA solve requests across more than 1,400 sites. This empirical evidence validates the system's ability to intercept and categorize a wide range of CAPTCHA farm activities without active participation in the attack chain. The qualitative analysis, underpinned by the collected data, further "demonstrated" the existence and characteristics of 34 distinct attack categories, providing concrete examples like the targeting of Taylor Swift concert tickets or Amazon Flex jobs. Therefore, the successful operation and the rich dataset produced by C-FRAME act as a powerful proof of concept for its measurement approach.

Defensive Implications

▶ Watch: Analysis of 'Spam/Scam' and 'Grey-Bot' attack categories (7:00)

The findings from the C-FRAME study provide critical, actionable intelligence for defenders, highlighting areas where current CAPTCHA implementations and broader web security strategies need reinforcement.

Firstly, CAPTCHA service providers must enhance their detection mechanisms to specifically identify and mitigate requests originating from human-driven CAPTCHA farms. The study underscores that these farms are highly adaptable and capable of bypassing even advanced behavioral CAPTCHAs. This implies a need for more sophisticated behavioral analysis that can distinguish between a human user solving a CAPTCHA and a human worker solving a CAPTCHA on behalf of an attacker, potentially through analysis of IP reputation, geographic anomalies, or patterns inconsistent with typical user behavior. The regional targeting observed also suggests that CAPTCHA services could leverage geo-IP blocking or regional anomaly detection more effectively.

Secondly, web service providers need to move beyond viewing CAPTCHAs as a silver bullet. While CAPTCHAs are an important layer of defense, the study clearly shows they are routinely bypassed for a wide array of malicious and undesirable activities. Websites must implement multi-layered security approaches, especially for high-value targets such as account creation pages, event ticketing systems, appointment booking portals, and gig economy job boards. This could include:

  • Rate limiting and IP reputation analysis at the application layer.
  • Behavioral analytics on the website itself, independent of the CAPTCHA service, to detect suspicious post-CAPTCHA behavior.
  • Account verification mechanisms beyond CAPTCHAs, such as email/SMS verification, especially for new account registrations.
  • For critical services like appointment booking, implementing identity verification or more robust challenge-response mechanisms that are harder for anonymous farm workers to complete.

The pervasive nature of account creation bots (e.g., on Twitter, LinkedIn, gaming platforms) demands stronger defenses against mass account generation, which is a precursor to large-scale spam, scam, and manipulation campaigns. Websites should consider implementing stricter policies on account age, activity, and requiring more personal information for certain actions, while balancing user experience.

Finally, the Grey Bot category highlights that even seemingly innocuous data scraping can leverage CAPTCHA farms. While not always malicious, excessive scraping can impact server resources and data integrity. Websites should consider implementing legal terms of service that prohibit automated scraping and employ technical measures to detect and block it, even if it's not overtly hostile. The C-FRAME methodology itself also offers a potential blueprint for organizations to develop their own internal monitoring systems to gain real-time insights into how their specific CAPTCHAs are being targeted.

Key Takeaways

  • Pervasive CAPTCHA Bypass: Modern behavioral CAPTCHAs are routinely bypassed by human-driven CAPTCHA farms, enabling a vast array of malicious activities across over 42,000 unique URLs and 1,000 organizations.
  • Four Major Attack Categories: CAPTCHA attacks broadly fall into Fraud (e.g., event ticket scalping, streaming manipulation), Abuse (e.g., Amazon Flex job securing, appointment booking), Spam/Scam (e.g., mass account creation on social media and gaming platforms), and Grey Bot (e.g., data scraping for business or personal use).
  • C-FRAME's Novel Methodology: The C-FRAME system is the first to provide real-world measurement of these attacks, utilizing a passive collection technique that introduces an error in the site key to observe attack requests without assisting in their completion.
  • Global Scale and Diverse Targets: Attacks originate from and target systems in 58 countries, across five continents, and 15 languages, affecting major platforms like Twitter, Amazon Flex, Roblox, Sony, and OpenAI, with Google reCAPTCHA, hCaptcha, and FunCaptcha being the most targeted CAPTCHA services.
  • Need for Multi-Layered Defenses: Websites cannot rely solely on CAPTCHAs for bot mitigation; they must implement additional layers of security, including behavioral analytics, rate limiting, and robust account verification, especially for high-value transactions or account creation.
  • Actionable Intelligence for Defenders: The detailed categorization and quantitative analysis provide valuable insights for CAPTCHA service providers to improve detection mechanisms and for web services to prioritize and strengthen their defenses against specific attack types.

About the Speaker(s)

The talk was presented by Hoang Dai Nguyen, a researcher from Louisiana State University. The work was a collaborative effort involving a team of researchers including Karthika Subramani, Bhupendra Acharya, Roberto Perdisci, and Phani Vadrevu, indicating a joint research initiative between Louisiana State University, Supa, and the University of Chure. The presentation itself focused on the technical details and findings of the C-FRAME system, a pioneering measurement study on real-world CAPTCHA attacks.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk presents C-FRAME, a novel and technically clever system for passively measuring real-world CAPTCHA attacks at an unprecedented scale. The empirical data reveals the pervasive nature and diverse categories of CAPTCHA bypass, offering critical, actionable intelligence for defenders and service providers to move beyond theoretical discussions.

Heather Calloway (CISO) — STRONG ACCEPT

This research provides crucial empirical data confirming the pervasive bypass of modern CAPTCHAs by human farms. It underscores a critical governance gap, necessitating that organizations adopt multi-layered bot mitigation strategies beyond single-control reliance. This is a must-read for any CISO whose business operates online.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024