CollisionRepair: First-Aid and Automated Patching for Storage Collision Vulnerabilities in Smart Contracts
Yu Pan
34th USENIX Security Symposium (USENIX Security '25) · Day 2 · Blockchain Security 2: Infrastructure, Protocol Design, and Governance
Overview
This paper introduces Halligan, the first generalized visual CAPTCHA solver built upon state-of-the-art vision language models (VLMs). Authored by a team of researchers from Shanghai Jiao Tong University, National University of Singapore, and Tel Aviv University, Halligan fundamentally challenges the long-held assumption that visual CAPTCHAs are "bot-hard" but "human-friendly." By demonstrating the ability to effectively solve a diverse array of unseen visual CAPTCHA challenges without prior adaptation, Halligan marks a significant paradigm shift in the cat-and-mouse game between CAPTCHA designers and attackers.
Read the paper · Download the PDF (PDF) · Slides
Paper abstract
Visual CAPTCHAs, such as reCAPTCHA v2, hCaptcha, and GeeTest, are mainstream security mechanisms to deter bots online, based on the assumption that their visual challenges are bot-hard but human-friendly. While many deep-learning based solvers have been designed and trained to solve a specific type of visual challenge in a CAPTCHA, vendors can easily switch to out-of-distribution visual challenge of the same type or even new types of challenge with very low cost. However, the emergence of general-purpose AI models (e.g., ChatGPT) challenges the bot-hard assumption of existing visual challenges, potentially compromising the reliability of visual CAPTCHAs. In this work, we report the first generalized visual CAPTCHA solver, Halligan, built upon the state-of-the-art vision language model (VLM), which can effectively solve unseen visual challenges in CAPTCHAs without making any adaptation. Our rationale lies in that a visual challenge can be reduced to a search problem where (i) its instruction is transformed into an optimization objective and (ii) its body is transformed into a search space for the objective. With well designed prompts built upon known VLMs, the transformation can be generalized to unseen visual challenges. Our extensive experiments show that Halligan is a game-changer to the known practice of adopting visual CAPTCHAs, which achieves a solving rate of 60.7% on 2,600 challenges belonging to 26 types of visual CAPTCHAs. Further, we use Halligan to infiltrate human-driven CAPTCHA farms, achieving an average solving rate of 70.6% on previously unseen visual challenges from CAPTCHAs in the wild over a 30-day period. Based on the experimental results, we further shed light on puzzle-less anti-bot alternatives in this era.

Are CAPTCHAs Still Bot-hard? Generalized Visual CAPTCHA Solving with Agentic Vision Language Model
Speakers: Xiwen Teoh (Shanghai Jiao Tong University), Yun Lin (Shanghai Jiao Tong University), Siqi Li (National University of Singapore), Ruofan Liu (National University of Singapore), Avi Sollomoni (Tel Aviv University), Yaniv Harel (Tel Aviv University), Jin Song Dong (National University of Singapore)
Conference: USENIX Security
YouTube: Not applicable; this article is based on a peer-reviewed paper.
Overview
This paper introduces Halligan, the first generalized visual CAPTCHA solver built upon state-of-the-art vision language models (VLMs). Authored by a team of researchers from Shanghai Jiao Tong University, National University of Singapore, and Tel Aviv University, Halligan fundamentally challenges the long-held assumption that visual CAPTCHAs are "bot-hard" but "human-friendly." By demonstrating the ability to effectively solve a diverse array of unseen visual CAPTCHA challenges without prior adaptation, Halligan marks a significant paradigm shift in the cat-and-mouse game between CAPTCHA designers and attackers.
The research highlights the increasing vulnerability of mainstream visual CAPTCHAs, such as reCAPTCHA v2, hCaptcha, and GeeTest, in the era of advanced AI. Halligan's success stems from its novel approach of reducing visual challenges to a search problem, where natural language instructions become optimization objectives and the CAPTCHA interface forms the search space. This work not only provides concrete evidence of AI's growing capability in bypassing these security measures but also necessitates a re-evaluation of current anti-bot strategies, advocating for a move towards puzzle-less alternatives.
Background
For over two decades, CAPTCHAs (Completely Automated Public Turing test to tell Computers and Humans Apart) have served as a critical web-abuse prevention mechanism, establishing proof-of-personhood by presenting users with challenges designed to be easily solvable by humans but difficult for automated bots. As of 2024, an estimated 256,000 of the top 1 million websites utilize CAPTCHAs, with visual CAPTCHAs constituting 94% of these implementations. The security of these systems has historically been a continuous arms race, with attackers developing increasingly sophisticated solvers for specific challenge types, and defenders responding by rapidly deploying new visual challenges or variations to nullify previous attack efforts.
Early CAPTCHA solvers targeted text-based challenges, evolving to tackle visual reasoning tasks in systems like Google's reCAPTCHA v2 through techniques such as image classifiers, object detectors, and visual question answering. However, these solvers were typically specialized, meaning they were trained for a particular type of visual puzzle. This allowed CAPTCHA vendors to easily introduce "out-of-distribution" challenges or entirely new puzzle types, effectively rendering existing specialized solvers obsolete with minimal cost. This dynamic ensured that defenders generally maintained an advantage, or at least parity, in the security arms race.
The advent of general-purpose AI models, particularly Vision Language Models (VLMs) and agentic AI, has fundamentally disrupted this balance. VLMs possess the capability to understand and interpret visual information in conjunction with natural language instructions, enabling them to comprehend a vast array of visual challenges without explicit prior training for each specific variant. Agentic AI, with its ability to autonomously plan and execute tasks using external tools, further empowers the development of generalized solutions that can interact with real-world websites. This new generation of AI models challenges the core "bot-hard" assumption of visual CAPTCHAs, suggesting that attackers may now, for the first time, outpace defenders by developing solutions that are broadly applicable across diverse and unseen CAPTCHA types. The paper argues that the current allowance for multiple CAPTCHA attempts, a usability feature, further amplifies this vulnerability, as even a moderate solving rate can translate into a high success probability over several tries (e.g., 96.8% success after 5 attempts with a 50% solving rate).
Key Findings
The research presents several groundbreaking findings that significantly impact the understanding of visual CAPTCHA security:
- First Generalized Visual CAPTCHA Solver: Halligan is the first reported generalized visual CAPTCHA solver capable of effectively solving unseen visual challenges automatically, without requiring any adaptation, pre-training, or fine-tuning. This marks a critical shift from specialized, challenge-specific solvers.
- High Solving Rate on Diverse Benchmark: Halligan achieved a remarkable 60.7% solving rate across an extensive benchmark of 2,600 challenges belonging to 26 distinct types of visual CAPTCHAs. This performance significantly surpasses both specialized solvers (VTTSolver, GeeSolver, PhishDecloaker, which solved only 3.8%, 11.5%, and 15.4% of types respectively) and generalist web navigation agents (WebVoyager with 8.9% and ShowUI with 9.8%).
- Effectiveness in the Wild: In a challenging 30-day field study, Halligan infiltrated human-driven CAPTCHA farms (specifically 2Captcha), achieving an average solving rate of 70.6% on 3,000 previously unseen visual challenges from CAPTCHAs in the wild. This included successfully tackling 17 variants not present in the benchmark and 4 new CAPTCHA providers, demonstrating its real-world applicability and robustness against evolving challenges and behavioral verification layers.
- Robustness Against Adversarial Attacks (with Countermeasures): While defensive measures like Gaussian noise, motion blur, typographic prompt injection, refusal prompts, and distraction attacks could initially reduce Halligan's solve rates (by 15.4%–49.6%), Halligan's built-in attacker countermeasures (non-local means denoising, Richardson–Lucy deconvolution, and irrelevant frame filtering) were able to recover solve rates to 56.5%–85.4%. This indicates that no defense fully prevented the attacks, and with repeated attempts, even a reduced success rate could lead to a high overall bypass probability.
- Superiority Over Human Prompt Engineering: A user study involving five experienced LLM prompt engineers revealed that Halligan significantly outperformed human-crafted prompts. Halligan required a median of 0.4 minutes to craft a working prompt per CAPTCHA type, compared to an average of 8.2 to 11.9 minutes for humans. On test sets, Halligan's performance was 172%, 104%, and 58% better than the best human prompts on three different CAPTCHA types, highlighting the difficulty of manual prompt engineering for generalization and the efficiency of Halligan's automated approach.
- Low Cost and Acceptable Time Overhead: The median cost to attack a challenge with Halligan was $0.0242, and the median end-to-end time was 21.8 seconds. With amortization, this time could be reduced to the code execution phase, resulting in a 3.46x speedup (71.1% time reduction), making it comparable to human (3.1–42s) and ad-hoc bot (0.016–17.5s) solving times.
- Impact on Phishing Detection: Halligan's enhanced solving capabilities (outperforming PhishDecloaker) led to the detection of 1,226 more phishing sites when integrated into a phishing detection setup, demonstrating a practical security application beyond direct CAPTCHA bypass.
These findings collectively underscore the critical need for CAPTCHA evolution, as traditional visual challenges are no longer "bot-hard" in the face of advanced VLM-powered agents.
Technical Deep Dive
Halligan's effectiveness stems from its innovative reformulation of the visual CAPTCHA solving problem as a search optimization problem, orchestrated by a system of VLM-based agents and external tools. The process unfolds in three sequential steps: Objective Identification, CAPTCHA Abstraction, and CAPTCHA Solving.
At its core, Halligan conceptualizes a visual CAPTCHA challenge as a search problem where the challenge's natural language instruction is transformed into an optimization objective, and its visual interface (body) is transformed into a search space for that objective. This generalization allows Halligan to tackle unseen challenges without explicit pre-training or fine-tuning.
The system relies on a CAPTCHA metamodel (Figure 3 in the paper), which defines the fundamental entities of a visual CAPTCHA challenge and their relationships. These entities include:
- Frames: Containers for other CAPTCHA entities, potentially nested. Frames can
instruct,act,refer, orterminateother frames, establishing semantic relationships (Table 1). - Elements: Visible objects within a frame, classified as either interactable or non-interactable. Interactable elements are further categorized by their action types:
NEXT,CLICKABLE,INPUTTABLE,POINTABLE,SWAPPABLE,SELECTABLE,DRAGGABLE, andSLIDEABLE(Table 2). - Keypoints: Clickable points within a frame.
Step 1: Objective Identification
The first step involves a VLM agent parsing the visual CAPTCHA challenge to extract an optimization objective in natural language. This agent processes important CAPTCHA entities, their textual descriptions, and their inferred relations to guide the VLM in outputting a relevant, clear objective. This objective acts as a "strengthened intention" to mitigate potential VLM hallucinations. The input consists of a set of frames, each forming a partition of the visual challenge space, with descriptions and inferred relations generated sequentially by the VLM in a chain-of-thought manner to improve quality.
Step 2: CAPTCHA Abstraction
In this step, the visual CAPTCHA challenge c is abstracted into a CAPTCHA model, preserving only relevant information such as layout and interactable GUI elements. This model forms the foundation for constructing the search space. The abstraction process involves three sub-steps:
- Frame Extraction: A frame extractor
fframe(Ic)takes a screenshotIcand identifies distinct frames. This is treated as a connected components problem, using image pre-processing techniques like median blur, adaptive binary thresholding, and morphological transformations to isolate and define frames. - Element Extraction: An element extractor
felement(Fi)takes a frameFiand identifies individual elements. This is an image segmentation problem, leveraging the Fast Segment Anything Model (FastSAM) [77], which combines a CNN-based YOLOv8-seg detector with the YOLACT method for instance segmentation. - Element Annotation and Typing: Once elements are identified, their types (e.g.,
CLICKABLE,SLIDEABLE) are determined. The agent generates natural language descriptions for elements of interest. The Contrastive Language-Image Pre-Training (CLIP) model [57] encodes and retrieves the most relevant element for each description. Finally, the agent annotates each element's type using its description and retrieved image, generating executable Python code usingget_element()andset_element_as(). - Keypoint Extraction: A keypoint extractor
fkeypoint(Fi)uses the Simple Linear Iterative Clustering (SLIC) algorithm [17] for superpixel segmentation based on color and proximity. A saliency map is then computed, retaining clusters with high average saliency, whose centroids become the keypoints.
Step 3: CAPTCHA Solving
Given the objective o and the initial CAPTCHA solution s (the abstracted CAPTCHA model), Halligan formulates the solving problem as a search problem with two main components: Search Space Exploration and Solution Evaluation.
Halligan employs a set of specialized tools (Table 3) to navigate and solve the CAPTCHA:
- Space Exploration Tools: These tools derive new CAPTCHA solutions (
c') from an existing solutionc. They includetype(x, text),click(x),swap(f),slide(x, dir, c), anddrag(x,y). A compatibility mapping ensures that only tools relevant to detected interactable elements are exposed, reducing misuse. The VLM agent selects compatible tools to generate a set of neighboring solutions, forming the search space. For continuous actions like sliding, this involves discretizing the action space (e.g.,slide_x()at intervals). - Solution Evaluation Tools: For each new solution
c', these tools compare and select the local optimumc*. The VLM agents userank(S, o)(powered by GPT-4o) to rank a set of solutionsSaccording to objectiveo, andcompare(si, sj, o)(also GPT-4o) to evaluate two solutions. - Enhancing Tools: These are called by Solution Evaluation Tools to enrich visual input for more reliable VLM reasoning. They include
mark(x, obj)(using GroundingDINO) to annotate elements with bounding boxes,focus(x, obj)(GroundingDINO) to zoom in on specific objects,ask(x, q, fmt)(GPT-4o) to query details about an element, andmatch(x, y)(using Hu moments [45] and color palettes) to check visual similarity. The VLM agent autonomously decides when and which enhancing tools to apply.
Halligan implements a hill-climbing approach, iteratively identifying local optima based on the objective o and exploring new solutions, progressively approaching a globally sub-optimal solution that aligns with the CAPTCHA objective. Finally, this optimal solution is translated into executable Python code (using PyAutoGUI for GUI automation) to interact with and solve the visual CAPTCHA challenge.
To enhance robustness, Halligan also integrates low-cost pre-processing techniques as attacker countermeasures, such as non-local means denoising and Richardson–Lucy deconvolution to address noise and blur, and a filtering stage to discard irrelevant frames.
Demo / Proof of Concept
While this article is based on a peer-reviewed paper rather than a live talk, the "Demo / Proof of Concept" section directly correlates to the extensive experimental evaluations detailed within the paper. The researchers rigorously evaluated Halligan in both closed-world and open-world settings to demonstrate its capabilities and robustness.
Closed-World Evaluation: Interactive Benchmark
For the closed-world evaluation, the authors developed a novel, scalable, and diverse interactive CAPTCHA benchmark. This benchmark comprised 26 distinct types of visual CAPTCHAs and a total of 2,600 unique challenges. Unlike traditional static image-label datasets, this benchmark presented interactive challenges, simulating real-world scenarios. CAPTCHA types were selected based on availability, popularity among the top 1 million websites, and their requirement for advanced AI (AGI) for human-level performance. Challenges were replicated by crawling demo sites or scraping real websites using Playwright, collecting HTML, CSS, JavaScript, and challenge data. Ground truth was established through manual labeling by co-authors, with continuous-answer tasks having agreed-upon tolerance ranges. Challenges were hosted on a Flask server, providing JSON feedback upon submission to verify success.
Halligan and baseline GUI agents (B1: GUI Agent, B2: WebVoyager, B3: ShowUI) were evaluated online, interacting with these challenges sequentially in a web browser using GPT-4o as the VLM agent and PyAutoGUI for interactions. Halligan was allowed up to 3 retries for execution errors. Other specialized baseline solvers (B4: VTTSolver, B5: GeeSolver, B6: PhishDecloaker) were evaluated offline using static image-label pairs due to their lack of GUI interaction capabilities. The results from this benchmark clearly demonstrated Halligan's superior solve rate of 60.7% across the diverse challenge types, outperforming all baselines.
Open-World Evaluation: Field Study on Human-Driven CAPTCHA Farms
To assess Halligan's performance in a realistic, "in-the-wild" environment, a 30-day field study was conducted by infiltrating 2Captcha, a prominent human-driven CAPTCHA farm. This setup allowed Halligan to operate as a worker, solving live CAPTCHA tasks that often involved not only visual puzzles but also implicit behavioral verification.
The success metric for this study was Halligan's ability to retrieve a CAPTCHA token from the CAPTCHA provider, which 2Captcha then forwards to its customers. The success rate was calculated as the ratio of total tokens received to total CAPTCHA tasks attempted. The researchers configured a worker device (Windows 11 x64, 1920x1080 resolution) with active browsing history, using a single IP address. Halligan ran 2Captcha's worker client, interacting with challenges using simplified, linear mouse movements, and relied on 2Captcha's infrastructure (e.g., proxies) to inherit customer identities, implicitly addressing some behavioral signals.
Over the 30-day period, Halligan processed 3,000 CAPTCHA tasks, successfully attacking 2,117 of them, achieving an average daily success rate of 70.6%. It encountered 9 distinct CAPTCHA providers, including 4 (2captcha, amazon, prosopo, xcaptcha) not present in the original benchmark. Furthermore, Halligan faced 25 unique variants, with 17 (75.2% of tasks) being previously unseen in the benchmark. Despite this high diversity and the real-world complexities, Halligan consistently outperformed or matched all baseline solvers, demonstrating its strong generalization capabilities and practical efficacy against live CAPTCHAs. The fact that the worker account was not banned also served as an indirect positive signal for token validity.
These comprehensive evaluations serve as the proof of concept, convincingly demonstrating Halligan's groundbreaking ability to generalize across and effectively solve a wide range of visual CAPTCHAs in both controlled and real-world adversarial environments.
Defensive Implications
The emergence of Halligan, a generalized visual CAPTCHA solver, significantly alters the landscape of online bot deterrence and presents critical implications for defenders. The core assumption that visual CAPTCHAs are "bot-hard" is fundamentally challenged, meaning that relying solely on complex visual puzzles is no longer a sustainable security strategy. Defenders must recognize that the visual challenge layer of modern CAPTCHA systems is now significantly weakened, shifting greater reliance onto other, more passive components of the anti-bot system.
The paper strongly advocates for a "Call for CAPTCHA Evolution," emphasizing the need for puzzle-less CAPTCHA alternatives. While current CAPTCHA systems often employ multilayered defenses (combining visual challenges with behavioral biometrics, identity profiling, device fingerprinting, etc.), the demonstrated vulnerability of the visual component exacerbates the long-standing trilemma of balancing security, usability, and privacy.
The authors outline eight puzzle-less alternatives that defenders should consider integrating or prioritizing:
- Behavioral-only CAPTCHAs: These systems rely on passive, frictionless layers of analysis, such as mouse movements, keyboard input patterns, or touch gestures, to distinguish humans from bots without presenting an explicit puzzle. Cloudflare Turnstile is cited as an example.
- Information Disparity: This approach leverages knowledge that a human user possesses but a bot typically does not, such as security questions, personal facts, or specific contextual information.
- Hardware Attestation: This mechanism verifies that the user possesses a specific hardware component or trusted platform, which bots typically lack. Examples include WebAuthn (FIDO2), which uses cryptographic keys stored on hardware security modules.
- Biometrics: This involves verifying the user's identity based on inherent physical or behavioral characteristics, such as fingerprints, facial recognition, or voice patterns.
- Social Networks: This alternative relies on external social graph validation, where other trusted individuals or entities vouch for the user's legitimacy, often through referrals or existing network connections.
- Strong Network: This involves verification by an authoritative third party, such as Know Your Customer (KYC) processes, OAuth authentication, or government-issued digital identities.
- Proof-of-Work: This method requires users (or their devices) to expend a small amount of computational effort, making it costly for bots to scale attacks due to the cumulative processing power required. mCaptcha is an example.
- Honeypot: This technique involves setting up hidden elements or forms that are invisible to legitimate users but detectable by bots, trapping them when they interact with these decoy elements.
While Halligan demonstrated resilience against certain adversarial transformations and prompt injections, these measures only reduced solve rates, and Halligan's countermeasures could largely mitigate their impact. Furthermore, the ability to make repeated attempts means even a partially effective defense can be bypassed over time. Therefore, defenders should:
- Diversify and Layer Defenses: Do not rely solely on visual puzzles. Integrate multiple layers of defense, with a stronger emphasis on behavioral analytics, device fingerprinting, and identity checks.
- Investigate Puzzle-Less Solutions: Actively research and implement the puzzle-less alternatives outlined. The shift away from explicit visual puzzles is paramount.
- Monitor VLM Advancements: Continuously monitor the advancements in VLMs and agentic AI, as these technologies will continue to evolve and potentially bypass new defensive mechanisms.
- Consider Economic Deterrents: Implement proof-of-work or other mechanisms that increase the computational or financial cost for attackers to scale their operations.
- Adopt Adaptive Security: Implement systems that can dynamically adjust their security measures based on risk assessment, user behavior, and detected threats, rather than static, one-size-fits-all CAPTCHA challenges.
The paper's findings serve as a stark warning: the era of "bot-hard" visual CAPTCHAs is drawing to a close. Proactive and fundamental changes in anti-bot strategies are urgently needed to maintain effective online security.
Key Takeaways
- Generalized AI Solvers are a Game-Changer: Halligan demonstrates that Vision Language Models (VLMs) can now effectively solve diverse and unseen visual CAPTCHA challenges without specific prior training, fundamentally undermining the "bot-hard" assumption of traditional CAPTCHAs.
- High Performance in Diverse Environments: Halligan achieved a 60.7% solving rate on a benchmark of 2,600 challenges across 26 CAPTCHA types and a 70.6% success rate in a 30-day field study against live, unseen CAPTCHAs in human-driven farms.
- Superiority Over Specialized Solutions and Humans: Halligan significantly outperforms previous specialized CAPTCHA solvers and generalist web navigation agents. It also proved more effective and efficient than human prompt engineers in crafting generalizable solutions.
- Traditional Defenses Are Insufficient: While adversarial visual transformations and prompt injections can reduce solve rates, Halligan's built-in countermeasures largely mitigate these defenses. The ability to make repeated attempts means even partially effective defenses are vulnerable.
- Urgent Need for Puzzle-Less CAPTCHA Alternatives: The paper calls for a paradigm shift towards puzzle-less anti-bot mechanisms, such as behavioral biometrics, hardware attestation, information disparity, and proof-of-work, to maintain effective online security in the AIGC era.
- Practical Security Implications: Beyond direct bypass, Halligan's capabilities can be leveraged to detect more CAPTCHA-cloaked phishing sites, indicating its potential for both offensive and defensive applications in the broader security landscape.
About the Speaker(s)
The research presented in this paper was a collaborative effort by a team of authors from multiple institutions:
- Xiwen Teoh is affiliated with Shanghai Jiao Tong University and the National University of Singapore.
- Yun Lin is affiliated with Shanghai Jiao Tong University.
- Siqi Li is affiliated with the National University of Singapore.
- Ruofan Liu is affiliated with the National University of Singapore.
- Avi Sollomoni is affiliated with Tel Aviv University.
- Yaniv Harel is affiliated with Tel Aviv University.
- Jin Song Dong is affiliated with the National University of Singapore.
Their collective expertise in computer science, AI, and security contributed to the development and evaluation of Halligan.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
This is solid, important research that definitively demonstrates visual CAPTCHAs are now trivially breakable by VLM-powered agents. The 60.7% benchmark rate and 70.6% field study rate aren't just academic curiosities—they're the death knell for visual challenge-based bot detection. The work is rigorous, the threat model is realistic, and the implications are concrete.
Heather Calloway (CISO) — STRONG ACCEPT
This is the paper that ends the CAPTCHA conversation at the board level. Visual CAPTCHAs are no longer a defensible control against automated abuse — VLMs have crossed the threshold. Every security program relying on reCAPTCHA, hCaptcha, or similar needs to accelerate migration to behavioral and attestation-based alternatives.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)