Winners of DARPA’s AI Cyber Challenge
Andrew Carney (Program Manager · DARPA), Jason Roos, Stephen Winchell
DEF CON 33 · Day 1 · Main Stage
Overview
The DARPA AI Cyber Challenge (AICC) is a landmark public competition aimed at revolutionizing software security by developing autonomous systems capable of discovering and patching vulnerabilities in source code. This talk, delivered by DARPA Program Manager Andrew Carney, along with special guest speakers Deputy Secretary of Health and Human Services Jim O'Neal and DARPA Director Stephen Winchell, unveiled the groundbreaking results of the AICC finals. The core objective of the challenge was to transcend the limitations of human-scale vulnerability management, which struggles to keep pace with the vast and complex landscape of modern software, particularly critical infrastructure.

Key moments
- 0:00 Introduction to DARPA's AI Cyber Challenge and problem
- 2:00 Erlang OTP critical SSH bug: a decade-long vulnerability
- 4:00 What is DARPA's AI Cyber Challenge (AIC)?
- 4:20 AIC's focus: automated patching and open-sourcing solutions
- 5:20 Triage challenge: any defect can become a vulnerability
- 7:20 LLMs and industry partnerships powering the challenge
Winners of DARPA’s AI Cyber Challenge
Speakers: Andrew Carney (Program Manager, DARPA), Jason Roos, Stephen Winchell
Conference: DEF CON
YouTube: https://www.youtube.com/watch?v=touJ5uLlXjQ
Overview
The DARPA AI Cyber Challenge (AICC) is a landmark public competition aimed at revolutionizing software security by developing autonomous systems capable of discovering and patching vulnerabilities in source code. This talk, delivered by DARPA Program Manager Andrew Carney, along with special guest speakers Deputy Secretary of Health and Human Services Jim O'Neal and DARPA Director Stephen Winchell, unveiled the groundbreaking results of the AICC finals. The core objective of the challenge was to transcend the limitations of human-scale vulnerability management, which struggles to keep pace with the vast and complex landscape of modern software, particularly critical infrastructure.
The significance of the AICC lies in its pursuit of "technology truly indistinguishable from magic" – robust, performant, and resilient software foundations essential for technological advancement. By fostering the development of AI-driven systems that can autonomously identify and remediate security flaws, DARPA and ARPA-H (Advanced Research Projects Agency for Health) seek to address the escalating threat of cyberattacks against critical infrastructure, including the healthcare sector. The competition not only pushed the boundaries of automated program analysis and large language model (LLM) integration but also demonstrated the unprecedented potential for cost-effective, rapid, and scalable vulnerability remediation.
Background
▶ Watch: Introduction to DARPA's AI Cyber Challenge and problem (0:00)
The pervasive issue of software vulnerabilities, particularly within critical infrastructure, forms the stark backdrop for the AI Cyber Challenge. Andrew Carney opened the talk with a compelling example: the Erlang OTP open-source library, a foundational component in numerous networking and telco appliances. In early 2024, a critical bug in its SSH implementation, introduced in 2014, was disclosed by university researchers. Despite Erlang OTP being actively maintained and widely used, this vulnerability persisted for over a decade. While maintainers rapidly issued a fix within 36 hours, the downstream impact was substantial; Cisco and NetApp, among others, were affected, with some Cisco products still lacking patches over four months post-disclosure. Carney emphasized that this specific vulnerability was "trivial to find and trivial to exploit," yet it remained undetected by conventional methods for years.
This scenario highlights a fundamental flaw in the current human-centric approach to cybersecurity: the sheer scale of the problem. Modern software, heavily reliant on open-source projects maintained by volunteers, underpins critical infrastructure, making it a prime target for nation-state and other cyber actors. The process of triaging bugs, determining their security relevance (which can be "maybe undecidable"), developing, testing, and deploying patches is labor-intensive, time-consuming, and resource-constrained. What appears as a benign bug today can easily become a weaponizable zero-day vulnerability tomorrow given the right context and analysis. The current status quo, where patching a vulnerability in healthcare can take an average of 491 days compared to 60 to 90 days in other industries, is unsustainable.
Recognizing this critical gap, DARPA, in collaboration with ARPA-H, launched the AICC as a public competition. Its goal was to incentivize the development of autonomous systems that could not only discover "realistic vulnerabilities" but, crucially, patch them effectively in source code. The challenge directly aimed to fill the void in automated patch generation, moving beyond mere vulnerability discovery. To maximize impact, DARPA committed to releasing all finalist technology as open source, ensuring the entire world could benefit from these advancements. The competition also leveraged partnerships with leading frontier model companies like Google, Anthropic, OpenAI, and Microsoft, who provided over $1 million in cloud credits and engineering support, acknowledging the rapid evolution of LLM technology and its potential to analyze code semantics. Further partnerships with the Linux Foundation, Open Source Software Foundation (OpenSSF), and Open Source Security Foundation (OpenSSF), and Open Source Technology Improvement Foundation (OSTIFF) were established to create conduits for deploying these technologies to maintainers and project owners.
Key Findings
▶ Watch: What is DARPA's AI Cyber Challenge (AIC)? (4:00)
The AI Cyber Challenge culminated in a series of remarkable achievements that underscore the transformative potential of autonomous security systems. Building upon promising semi-final results—where teams demonstrated proof of life by finding and patching vulnerabilities in millions of lines of code and even discovered a real issue in SQLite—the final round saw an exponential leap in capabilities.
The most striking finding was the teams' ability to patch almost every type of vulnerability they could discover. Across a diverse set of open-source repositories totaling over 54 million lines of code, the autonomous systems collectively discovered nearly 80% of the synthetic vulnerabilities inserted into the challenge problems. More impressively, they successfully patched over 60% of all these synthetic challenge problems, doing so significantly faster than in the semi-finals, despite an increase in challenge complexity and the introduction of more vulnerability classes. The performance in Java, a language that posed a greater challenge in the semi-finals, saw a "huge improvement."
Beyond synthetic vulnerabilities, the teams made a profound real-world impact. They identified a total of 18 real zero-day vulnerabilities that were previously unknown, and successfully developed patches for 11 of them. These findings are currently undergoing disclosure to maintainers, with public release pending the completion of this process. The average time for these systems to generate a patch was under an hour, a stark contrast to the months or even years often required for human-driven remediation.
The economic viability of these autonomous systems was another critical finding. While teams were given a ceiling of tens of thousands of dollars in cloud compute and LLM credits for the finals, the total spend across all teams was approximately $360,000, with only $82,000 specifically allocated to LLM credits, supporting nearly 2 million queries. When broken down per task—whether it was finding a vulnerability, creating an effective patch, validating a bug report, or associating these items—the spend was remarkably low, often in the "low hundreds of dollars," with some teams consistently achieving results for as little as $152. This cost-effectiveness significantly undercuts the status quo of human labor-intensive vulnerability management, demonstrating a path to making automated patching cheaper and more accessible than current practices.
DARPA Director Stephen Winchell highlighted two key takeaways: the "vast improvement and the speed of improvement over time" observed in the teams' performance, and the "amount of opportunity space that we still have to tap into" in areas like model improvement, orchestration, ensembling, and agent specialization. Crucially, he noted that "machines are performing at superhuman levels already" in both the results and speed of vulnerability discovery and patching. The competition confirmed that the significant advancements were not solely due to improvements in the underlying LLM models, but rather the teams' innovative strategies in leveraging these technologies alongside traditional analysis methods.
Technical Deep Dive
▶ Watch: AIC's focus: automated patching and open-sourcing solutions (4:20)
The AI Cyber Challenge was meticulously designed to push the boundaries of autonomous vulnerability discovery and remediation, incorporating elements that mirror real-world software development and security challenges. The competition focused on developing systems that could find and fix vulnerabilities in realistic open-source software. Instead of creating entirely new codebases, DARPA and ARPA-H developed synthetic forks of existing, widely used open-source projects, into which "novel synthetic vulnerabilities" were carefully inserted. These vulnerabilities were designed to be realistic, simulating common security flaws (focused on CWEs – Common Weakness Enumerations) that had never been seen by any person or LLM before.
A critical aspect of the challenge was the definition of a "bug" versus a "vulnerability." The AICC rewarded teams for finding reachable, triggerable bugs. The bar for triggering a defect was set low: if a system could reliably trigger a defect at least 1% of the time across 100 attempts, it was considered significant enough to warrant remediation. This approach aimed to broaden the scope beyond strictly "security-relevant" issues, recognizing that any defect could potentially be weaponized given enough time and the right context. The ultimate goal was to achieve a state where patching is so cost-effective and fast that the need for extensive bug triage is minimized, allowing for a more resilient and performant software foundation.
The competition structure evolved from semi-finals to finals to increase complexity and better reflect real-world scenarios. In the semi-finals, teams had a relatively simple task: prove a bug was reachable (e.g., crash a provided harness) and demonstrate that their patch remediated the bug without introducing regressions. For the finals, the complexity was "significantly increased" with:
- More Real-World Challenge Types: Mapping directly to software development lifecycle use cases.
- Different Challenge Formats: Teams were given varied inputs, such as a patch diff (a common scenario where a human provides an initial fix to be validated or improved) or a full repository (requiring broader analysis).
- Bundle Challenges: The demanding task of associating a triggering input with a patch and/or a bug report, reflecting the need for comprehensive vulnerability reporting.
The target codebases for the finals comprised a diverse set of critical open-source repositories, totaling over 54 million lines of code. These included projects derived from a consultation with critical infrastructure owners, maintainers, and government agencies. The challenges were presented in both C and Java, requiring the autonomous systems to handle different language paradigms and associated vulnerability types.
The technological approach of the winning teams involved a novel integration of Large Language Models (LLMs) alongside traditional dynamic and static program analysis techniques. While the specific architectures of each winning system were not detailed in the talk, the overall success indicated that teams effectively combined the semantic understanding capabilities of LLMs with the precise code analysis strengths of traditional tools. This hybrid approach allowed them to analyze massive codebases, understand context, generate potential fixes, and validate their efficacy. The significant improvement in performance from semi-finals to finals, as noted by Carney, was not solely due to better LLMs but "the teams figured out how to use this technology better in more innovative ways," suggesting sophisticated orchestration and specialized training of these AI agents.
Resource allocation was also carefully managed. In the semi-finals, teams were severely constrained with only four hours of wall time, limited cloud compute, and a mere $100 of LLM credits per project (totaling $500). For the finals, these constraints were largely removed, offering "tens of thousands of dollars of ceiling" for resources, enabling teams to scale their operations and fully explore the capabilities of their autonomous systems.
Demo / Proof of Concept
▶ Watch: Triage challenge: any defect can become a vulnerability (5:20)
While the talk did not feature a live demonstration of a single winning system in action, the entire AI Cyber Challenge (AICC) competition itself served as a comprehensive, large-scale proof of concept for the capabilities of autonomous vulnerability discovery and patching. The competition, particularly its final round, was designed to rigorously test and validate these systems against realistic scenarios and codebases.
The results presented by Andrew Carney and the DARPA leadership unequivocally demonstrated that the finalist teams developed systems capable of finding and patching real vulnerabilities quickly, scalably, and cost-effectively. The successful identification of 18 real zero-day vulnerabilities and the subsequent patching of 11 of them within an average of less than an hour, across 54 million lines of code, is the ultimate proof of concept. It moved the concept of AI-driven security from theoretical potential to demonstrated reality.
Furthermore, the talk announced the "AICC experience" at DEF CON, where attendees could engage directly with the teams, learn about their solutions in depth, and view competition data and artifacts. This included access to the generated patches, proofs of vulnerability, and visualizations of interactions between models and systems. This provided a transparent look into how the competition unfolded and what the teams achieved, offering a form of indirect "demo" through detailed access to the outputs and methodologies of the winning systems. The impending open-source release of the winning systems (CRS's – Cyber Reasoning Systems) will allow the broader security community to directly experiment with and deploy these powerful tools.
Defensive Implications
▶ Watch: LLMs and industry partnerships powering the challenge (7:20)
The implications of the AI Cyber Challenge for cybersecurity defenders are profound and potentially transformative. The most significant defensive implication is the imminent open-source release of the winning autonomous systems (Cyber Reasoning Systems, or CRS's) developed by the finalist teams. As DARPA Director Stephen Winchell announced, four of these models were made available immediately, with the remaining three to follow shortly. This move democratizes access to cutting-edge AI-driven security tools, enabling organizations, maintainers, and individual developers to leverage them for their own projects.
The availability of these tools means that defenders will soon have access to automated systems capable of:
- Rapid Vulnerability Discovery: Identifying subtle, reachable, and triggerable bugs, including previously unknown zero-days, in vast codebases, potentially at "superhuman levels" of speed and scale.
- Automated Patch Generation: Developing effective patches that remediate vulnerabilities without introducing regressions, significantly reducing the manual labor and time typically associated with patch development.
- Cost-Effective Security: Performing these tasks at a fraction of the cost of traditional human-driven methods, making advanced security analysis accessible even for resource-constrained projects.
This capability is particularly critical for legacy systems and critical infrastructure, which often rely on "ancient digital scaffolding" and suffer from immense technical debt. Deputy Secretary Jim O'Neal highlighted the dire situation in healthcare, where patching a vulnerability can take 491 days on average, compared to 60-90 days in other industries. The AICC's solutions offer a lifeline to sectors like healthcare, enabling them to secure complex IT ecosystems and specialized legacy devices more effectively and rapidly, thereby protecting patient care and data from escalating cyberattacks.
DARPA is actively working with the Linux Foundation, Open Source Software Foundation (OpenSSF), and Open Source Technology Improvement Foundation (OSTIFF) to create conduits for maintainers and project owners to deploy and benefit from these technologies. Maintainers are explicitly encouraged to reach out to DARPA (via [email protected]) to explore collaboration and integration of these tools into their software development lifecycle.
The success of the AICC also signals a paradigm shift in security, moving towards a future where automation is the "new floor" for cybersecurity. It encourages defenders to embrace AI as a cornerstone of cyber defense and national security, as emphasized by O'Neal. The competition not only provides tools but also inspires further innovation in orchestration, ensembling, and specialization of AI agents. DARPA's ongoing investment of over $650 million across 14 programs in areas like secure compilers, eliminating weird machines, and provably secure codebases, underscores a holistic commitment to building a fundamentally secure digital foundation. The challenge for defenders now is to integrate these powerful, open-source AI capabilities into their operations, scale their usage, and contribute to their continuous improvement, ultimately aiming for a future with "warp engines and replicators and holodecks" built on an extremely strong and resilient tech stack.
Key Takeaways
- Autonomous Systems are Here: The DARPA AI Cyber Challenge successfully demonstrated that autonomous systems can effectively find and patch real vulnerabilities in complex, large-scale open-source codebases, performing at "superhuman levels."
- Significant Performance Leap: Teams achieved nearly 80% discovery and over 60% patching of synthetic vulnerabilities, along with identifying 18 real zero-days (patching 11) in the finals, a massive improvement over semi-final results.
- Cost-Effective Security: The autonomous approach proved highly economical, with an average cost in the low hundreds of dollars per vulnerability task, significantly cheaper than human-driven methods and requiring minimal LLM credit spend (e.g., $152 for some tasks).
- Open Source for All: The winning Cyber Reasoning Systems (CRS's) are being released as open source, making these advanced AI-driven security tools accessible to the global community, maintainers, and critical infrastructure owners.
- Critical for Infrastructure: These technologies offer a vital solution to the unsustainable pace of human-driven patching, particularly in critical sectors like healthcare, where vulnerability remediation currently takes an average of 491 days.
- A New Floor for Cybersecurity: The AICC sets a new baseline for automated security, signaling that AI will be an indispensable component of future cyber defense, with continuous and rapid improvements expected.
About the Speaker(s)
Andrew Carney is a Program Manager at DARPA (Defense Advanced Research Projects Agency). He introduced the AI Cyber Challenge, articulating its mission to develop autonomous systems for vulnerability discovery and patching, driven by the critical need to secure foundational software and critical infrastructure. Carney emphasized the unsustainability of current human-scale vulnerability management and DARPA's commitment to pushing technological boundaries to create robust and resilient tech stacks.
Jim O'Neal is the Deputy Secretary of Health and Human Services (HHS). Speaking as a special guest, O'Neal highlighted the severe impact of cyberattacks on the healthcare industry, citing statistics like 650+ ransomware incidents since 2018 and the 491-day average patching time in healthcare. He underscored the urgency of innovation at the intersection of AI, cybersecurity, and health, and announced ARPA-H's $20 million commitment to translate winning AICC solutions into real-world use for protecting the nation's healthcare infrastructure.
Stephen Winchell is the Director of DARPA. He expressed his deep enthusiasm for the AI Cyber Challenge, calling it "one of my very favorites" among DARPA's nearly 300 programs. Winchell elaborated on DARPA's broader investment in cybersecurity, noting over $650 million across 14 programs focused on secure computer systems, compilers, and provably secure code. He praised the AICC teams for achieving "superhuman performance" and announced the immediate open-source release of several winning systems, along with an additional $1.4 million in prize money to support the application of these technologies to critical infrastructure problems.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This is a strategic/results keynote, not a technical research drop — and judged in that lane, it delivers real signal. DARPA announcing 18 real zero-days discovered autonomously, 11 patched, across 54M lines of code, at $152/task, with immediate open-source release and $1.4M in follow-on prizes is not a press release — that's a program manager reporting genuine empirical results on stage at DEF CON. The open-source CRS release is the kind of concrete commitment that actually changes what defenders can do next week.
Heather Calloway (CISO) — SOLID
A credible and consequential research program with real results — 18 zero-days found, 11 patched, tools going open source — but the talk functions more as a program announcement than a decision-relevant briefing for the people who most need to act on it. The governing question for CISOs and critical infrastructure owners isn't whether this technology exists; it's how to integrate it, when to trust it, and who owns the outcomes when it gets something wrong.