On the Virtues of Information Security in the UK Climate Movement

Mikaela Brough

34th USENIX Security Symposium (USENIX Security '25) · Day 3 · Social Issues and Security

Overview

This comprehensive paper, "The Ransomware Decade: The Creation of a Fine-Grained Dataset and a Longitudinal Study," presents an in-depth analysis of the evolving ransomware landscape over the past ten years. Authored by a team from the University of Michigan, the research addresses a critical gap in the cybersecurity community: the lack of a systematic, long-range, and finely annotated dataset of ransomware incidents. By developing a novel, chatbot-aided methodology to extract detailed information from public reports, the authors have created a unique resource that enables a granular understanding of ransomware's prevalence, impact, and the dynamics between attackers and victims.

Read the paper · Download the PDF (PDF) · Slides

Paper abstract

Ransomware attacks have grown and evolved considerably in the past decade and are now one of the most common and most profitable attack vectors. Successful ransomware attacks have the ability to shut hospitals down, cause massive data and financial losses, tarnish the reputations of organizations, and even cause direct physical harm to people and property. Consequently, considerable attention has been paid to various individual aspects of the ransomware ecosystem in both the research community and the popular press. However, there continues to be a lack of comprehensive long-range census of these events. This presents a significant barrier to any comprehensive analysis of the ecosystem as a whole. In this paper, we present a longitudinal study of a decade of the ransomware attack landscape. This study is built upon a sophisticated process we developed to source and curate a unique large-scale dataset of ransomware incidents with fine-grained annotations on the basis of public reports of such incidents. We detail this process in the paper and showcase a variety of analysis enabled by such a dataset. Of particular interest are findings around the downstream impact of a large ransom payment vs. a high-profile refusal to pay, the impact of double extortion, the difference in susceptibility to different attack vectors and in payment attitudes across industry sectors.

Visual summary for On the Virtues of Information Security in the UK Climate Movement by Mikaela Brough
Visual summary for On the Virtues of Information Security in the UK Climate Movement by Mikaela Brough

The Ransomware Decade: The Creation of a Fine-Grained Dataset and a Longitudinal Study

Speakers: Armin Sarabi (University of Michigan); Ziyuan Huang (University of Michigan); Chenlan Wang (University of Michigan); Tai Karir (University of Michigan); Mingyan Liu (University of Michigan)

Conference: USENIX Security

YouTube: N/A (This is a peer-reviewed conference paper, not a recorded talk.)

Paper page: https://www.usenix.org/conference/usenixsecurity25/presentation/sarabi

Paper PDF: https://www.usenix.org/system/files/usenixsecurity25-sarabi.pdf

Overview

This comprehensive paper, "The Ransomware Decade: The Creation of a Fine-Grained Dataset and a Longitudinal Study," presents an in-depth analysis of the evolving ransomware landscape over the past ten years. Authored by a team from the University of Michigan, the research addresses a critical gap in the cybersecurity community: the lack of a systematic, long-range, and finely annotated dataset of ransomware incidents. By developing a novel, chatbot-aided methodology to extract detailed information from public reports, the authors have created a unique resource that enables a granular understanding of ransomware's prevalence, impact, and the dynamics between attackers and victims.

The significance of this work cannot be overstated. Ransomware has transcended a mere nuisance to become a pervasive and highly profitable threat capable of crippling critical infrastructure, inflicting massive financial and data losses, and even posing direct physical harm. Despite its profound impact, previous research and public reporting have largely focused on individual aspects, lacking the holistic, longitudinal perspective necessary for effective defense and policy-making. This study fills that void, offering data-driven insights that challenge conventional wisdom and provide actionable intelligence for organizations, researchers, and policymakers alike.

The paper's core contribution lies in its sophisticated data curation pipeline, which leverages modern AI chatbots to transform unstructured news articles into a structured, richly annotated dataset of 4,070 ransomware incidents. This dataset facilitates a range of analyses, exploring trends in attack vectors, ransom demands and payments, the rise of double extortion, and sector-specific vulnerabilities. The findings offer crucial context for understanding the economic, technological, and psychological factors driving the ransomware ecosystem, providing a foundation for more informed strategic responses.

Background

The concept of demanding ransom for valuable assets is an ancient one, but the digital age has given rise to a new, highly amplified form: ransomware. This modern iteration targets digital assets—software, data, and systems—which are often invaluable to organizations and individuals. The inherent value and sensitive nature of digital information have introduced new dimensions to this crime, most notably double extortion, where attackers not only encrypt data but also threaten its public disclosure or sale, adding a potent new leverage point that lacks a clear parallel in traditional hostage situations. High-profile incidents like WannaCry in 2017 and MedusaLocker in 2022 underscore the devastating global impact of ransomware, affecting hundreds of thousands of computers and causing billions in damages across various sectors, from healthcare to local governments.

Despite the escalating threat, a significant challenge for researchers and analysts has been the absence of comprehensive, structured datasets detailing ransomware incidents over extended periods. Existing cybersecurity incident repositories, such as the VERIS Community Database (VCDB) and its associated Data Breach Investigations Report (DBIR), provide valuable insights into cyber incidents generally but lack the specialized focus and granular annotations required for a deep dive into ransomware. For instance, out of over 10,000 incidents in VCDB, fewer than 1,200 explicitly mention ransomware, and the VERIS taxonomy itself is not tailored to the unique characteristics of these attacks, such as initial ransom demand versus settled payment, incident timelines, or specific responsible groups.

Other specialized datasets exist, but each comes with its own limitations. Ransomware.live aggregates claims from attacker leak sites, offering broad coverage (∼16k incidents) but often lacking victim-side details like operational impacts, actual payments, or recovery timelines. BlackFog provides a manually maintained list, but its coverage only began in 2020 and offers minimal annotation. StateScoop focuses on government entities but is relatively small (474 incidents) and infrequently updated. Datasets like Critical Infrastructure Ransomware Attacks (CIRA) and Ransomwhere are even more niche, focusing only on critical infrastructure or confirmed payments, respectively. Crucially, many of these resources rely heavily on extensive manual curation or crowdsourcing, which is labor-intensive and difficult to scale, highlighting the need for an automated, dedicated, and comprehensive dataset specifically designed for ransomware analysis. This paper's methodology directly addresses these limitations by automating the extraction of fine-grained ransomware-specific attributes from public news reports, aiming to provide an unprecedented level of detail for longitudinal study.

Key Findings

The longitudinal study enabled by the newly curated dataset reveals several critical insights into the ransomware ecosystem:

  • Attack Volume and Demands Surge (2018-2020): A sharp increase in ransomware incident volume and a corresponding rise in ransom demands and payments were observed between 2018 and 2020. This trend is strongly correlated with the emergence and proliferation of double extortion tactics, where attackers threaten both data encryption and public disclosure/sale. While this increase continued into 2021, a slight reversal has been noted since then.
  • Refusals Outnumber Payments: Across the dataset, victims' refusals to pay ransom consistently outnumber actual payments. The refusal rate has remained high, approximately 80%, since 2017, with a temporary dip in late 2020. This suggests that non-payment is a more common outcome than often portrayed.
  • Sectoral Payment Attitudes and Demands: The Financials sector demonstrates the highest willingness to pay ransom. Interestingly, for this sector, the median reported payment actually exceeds the median reported demand, suggesting a complex negotiation dynamic or reporting bias. Conversely, the Healthcare sector, despite widespread perception, does not face higher ransom demands than others; in fact, the Education and Government sectors receive the lowest average demands. Education also pays the least relative to the demands it receives.
  • Impact of Large Payments and Refusals:
  • Large Payments: Paradoxically, a large ransom payment (exceeding $2 million) appears to create a "negative sentiment" among subsequent victims. This results in a decrease in future ransom payments and an increased probability of refusal by other victims, starting roughly 20 days post-event. This effect is attributed to the significant negative press surrounding such payments.
  • Major Refusals: A major refusal (rejecting a demand of $10 million or more) is followed by a statistically significant decline in incident arrival rates, starting around 26 days post-event. This supports the conventional wisdom that public resistance can deter future attacks.
  • Double Extortion Dynamics: The emergence of double extortion significantly increases ransom demands (by more than an order of magnitude). However, it also elicits a "negative sentiment" from victims, leading to a higher likelihood of payment refusal, especially when accompanied by a partial data release to prove credibility. Double extortion disproportionately impacts sensitive data types, with HR, IP, and financial data seeing 22-fold, 12.4-fold, and 11.1-fold increases in incidents involving data release between 2018 and 2020.
  • Sector-Specific Attack Vectors:
  • Education and Government sectors exhibit elevated susceptibility to phishing and social engineering techniques.
  • IT, Consumer Goods, and Communications sectors show higher susceptibility to attacks involving Ransomware-as-a-Service (RaaS) models.

Technical Deep Dive

The core of this research is the development of a sophisticated, chatbot-aided methodology for curating a fine-grained ransomware incident dataset. This pipeline overcomes the limitations of manual data collection and existing generic cyber incident databases by specializing in ransomware-specific attributes. The process is divided into three main stages: data collection and pre-processing, chatbot-aided attribute extraction, and post-processing.

Data Collection and Pre-processing:

The initial phase involves sourcing news articles reporting potential ransomware incidents from three primary sources:

  1. Common Crawl News (CC-News): A vast archive of web crawls, filtered for articles containing "ransomware" in the page text and "ransom" in the page title to ensure relevance. This yielded 86,193 English articles after language detection.
  2. Supplementary Articles: URLs scraped from Blackfog's state of ransomware reports (2,255 URLs) and articles from DataBreaches.net (3,388 articles plus 2,144 linked URLs). These were then crawled using Playwright and, for offline pages, the Wayback Machine, achieving a 97.5% retrieval success rate for valid HTML content. This resulted in 7,248 English articles.

A crucial duplicate article removal step is applied to maintain data quality. For CC-News articles, comparisons are made within a two-week window. For supplementary articles, comparisons are made against all earlier articles. Articles from the same domain are retained if they have identical titles or an 80% overlap (to preserve updates). Articles from different domains are retained if the newer article has less than 80% overlap with the older one (to preserve original sources). This process reduced CC-News articles from 86,193 to 44,391 and supplementary articles from 7,248 to 6,495.

Chatbot-Aided Attribute Extraction:

This stage is the most innovative, leveraging OpenAI's GPT-4o model for automated information extraction. The chatbot was chosen after a cost/accuracy comparison with other models. Key strategies for maximizing accuracy include iterative prompt design, few-shot learning with manually annotated input/output examples, and a specialized annotation schema (taxonomy). This taxonomy, inspired by VERIS but tailored for ransomware, was developed by reviewing approximately 200 representative article segments per aspect to identify commonly reported details. The chatbot generates structured JSON summaries based on this taxonomy.

  1. Article Classification: The chatbot first classifies articles as either describing a specific ransomware incident or discussing broader, secondary topics (e.g., actors, vulnerabilities, solutions, trends). This step achieved a 95.5% overall accuracy during manual validation.
  2. Text Annotation: For incident-related articles, the chatbot identifies and annotates relevant text segments pertaining to seven key aspects: attack vectors, attacker actions, victim actions, impact, affected data, indirect victims, and timeline. It also identifies and associates annotations with specific victim-attacker pairs, capturing their names as given in the article. This yielded 154,899 annotated text segments across 20,005 articles, referencing 12,217 distinct victim names.
  3. Victim De-duplication: To consolidate references to the same entity, a multi-pass de-duplication strategy is employed. The chatbot removes ambiguous names and maps duplicates to a canonical name. This is done in batches (size 200) across three passes: chronological sorting, alphabetical sorting, and k-means clustering on OpenAI's text-embedding-3-large embeddings. This reduced 12,217 victim names to 4,212 unique names, removing 3,255 ambiguous entries. High-profile incidents with extensive coverage are further downsampled to manage data volume, reducing articles from 17,298 to 12,647.
  4. Detailed Attribute Extraction: Building on the annotated segments, the chatbot extracts specific binary, textual, and numeric attributes (summarized in Table 1 of the paper). This includes ransomware variant, targeted products, ransom/payment amounts, data categories, and event timelines. A critical feature is the chatbot's instruction to report whether information is confirmed or speculative, serving as a correction mechanism. The chatbot is also required to provide a short explanation for each extraction, improving accuracy by revealing its chain of thought.

Post-processing:

This final stage sanitizes and consolidates the extracted attributes at the incident level.

  1. Consolidating Information Across Articles: Speculative attributes (5.7%) are discarded. Non-binary attributes (dates, dollar amounts, data volumes) are parsed, discarding entries that cannot be converted into concrete values (e.g., ranges, vague scales). This resulted in the exclusion of 23.5% of date fields and 29.6% of durations. A two-pass merging strategy groups reports referring to the same victim and incident: first by timeline proximity (events within 7 or 30 days), then by ransomware group or variant name similarity (normalized indel similarity > 0.7). This process identified 4,070 incidents involving 4,014 unique victims. Binary attributes are consolidated by setting to true if captured in any report, while numeric values and dates use majority voting (86.9% of instances) or the median as a fallback.
  2. Manual Sanitization and End-to-End Validation: Repeat victims are manually reviewed to correct misclassifications, leading to the final count of 4,070 incidents from 4,014 unique victims. Ransomware group and variant names are cross-validated against reference lists like ID Ransomware, Ransomware.live, and Ransom-DB, mapping 726 unique names from annotations to 246 distinct groups/variants (150 known, 96 newly identified). An end-to-end validation on 250 post-processed incident records (1,822 attributes) showed an overall accuracy of 96.7%, with individual aspect accuracies ranging from 94.9% to 98.0%.

The resulting dataset, available at https://doi.org/10.5281/zenodo.15571866, represents a significant leap forward in ransomware incident data curation, providing an unprecedented level of detail for comprehensive analysis.

Demo / Proof of Concept

As this work is presented as a peer-reviewed conference paper rather than a live talk, there was no traditional "demo" or "proof of concept" in the form of a real-time software demonstration. Instead, the "proof of concept" is embodied by the successful creation and rigorous validation of the fine-grained ransomware incident dataset itself, and the subsequent longitudinal study presented in the paper.

The authors effectively demonstrate the power and utility of their chatbot-aided methodology by showcasing a wide array of analyses that the dataset enables. This includes statistical examinations of incident frequency over time, the evolution of ransom demands and payments, sector-specific vulnerabilities, the dynamics of double extortion, and the causal impact of major payment or refusal events. The comprehensive nature of the extracted attributes—covering attack vectors, attacker actions, victim responses, financial and operational impacts, affected data types, and incident timelines—serves as the primary evidence of the methodology's effectiveness and the dataset's value. The public availability of the dataset via Zenodo further allows other researchers and practitioners to inspect, validate, and build upon this foundational work.

Defensive Implications

The insights gleaned from this longitudinal study offer several crucial defensive implications for organizations and cybersecurity professionals:

  • Rethink Ransom Payment Policies: The finding that refusals to pay ransom significantly outnumber payments (around 80% refusal rate) suggests that non-payment is a viable and often chosen strategy. Furthermore, large, high-profile payments paradoxically lead to a "negative sentiment" in the broader ecosystem, resulting in lower future payments and higher refusal rates by other victims. Conversely, a major public refusal can lead to a decline in subsequent incident arrival rates. While insurance may cover ransom payments and reduce immediate out-of-pocket costs, insured victims who paid still experienced a median out-of-pocket loss of $1,000, while those who refused faced a median of $10,000. This implies that while insurance can influence payment decisions, refusal can lead to higher uncovered costs. Organizations should carefully weigh the long-term impact of payment decisions, considering public perception and potential deterrent effects, alongside immediate financial and operational recovery. Strengthening backup and recovery capabilities to make refusal a more feasible option is paramount.
  • Prioritize Double Extortion Defenses: The strong correlation between the rise of double extortion (2018-2020) and increased attack volume and demands underscores its effectiveness for attackers. Defenders must focus on preventing not only data encryption but, more critically, data exfiltration. This requires robust data loss prevention (DLP) solutions, network segmentation, strong endpoint detection and response (EDR), and continuous monitoring for suspicious outbound traffic. The study highlights that HR, IP, and financial data are disproportionately targeted in double extortion schemes, necessitating heightened security controls around these sensitive data categories. The fact that partial data release increases ransom demands but also the likelihood of refusal suggests that victims are increasingly unwilling to capitulate to such threats, potentially indicating a shift in victim psychology that defenders can leverage in their incident response planning.
  • Tailor Defenses to Sector-Specific Attack Vectors: The research reveals distinct patterns in attack vector susceptibility across industries:
  • Education and Government sectors are highly susceptible to phishing and social engineering. For these sectors, a strong emphasis on continuous security awareness training, robust email filtering, multi-factor authentication (MFA), and simulating phishing attacks is critical. Employees should be educated on identifying social engineering tactics and reporting suspicious activities.
  • IT, Consumer Goods, and Communications sectors are disproportionately targeted by Ransomware-as-a-Service (RaaS) operations. Organizations in these sectors must prioritize aggressive vulnerability management and patching, implement advanced threat detection and prevention systems (e.g., next-gen firewalls, EDR), and leverage threat intelligence specific to active RaaS groups to anticipate and defend against their evolving tactics, techniques, and procedures (TTPs).
  • Leverage Threat Intelligence on Ransomware Groups: The dataset provides insights into the activity patterns of specific ransomware groups like LockBit, BlackCat (ALPHV), REvil/Sodinokibi, Clop, Conti, and WannaCry. Defenders should integrate intelligence on active groups, their preferred targets, and TTPs into their security operations. Understanding the "gradual buildup to a peak, followed by a rapid decline" attack pattern observed for these groups can help anticipate surges and tailor defenses accordingly, acknowledging that groups may re-emerge or copycats may use similar names.
  • Comprehensive Incident Response Planning: The detailed timeline data, including median recovery times (3 days for partial, 7 days for full), emphasizes the need for well-rehearsed incident response plans. These plans should include clear communication strategies (customer notification, regulatory disclosure), engagement with law enforcement and cybersecurity consultants, and robust restoration procedures from isolated, immutable backups. The data also suggests that even with insurance, victims who refuse payment may incur higher out-of-pocket losses, highlighting the importance of comprehensive recovery strategies that minimize business disruption beyond just the ransom amount.

Key Takeaways

  • A novel, chatbot-aided methodology has successfully created the most comprehensive, fine-grained, and longitudinal dataset of 4,070 ransomware incidents from public news reports, overcoming limitations of previous datasets.
  • Ransomware attacks, demands, and payments surged dramatically between 2018 and 2020, a trend directly linked to the widespread adoption of double extortion tactics.
  • Victims overwhelmingly choose to refuse ransom payments, with refusal rates consistently around 80%, challenging the perception that paying is the most common outcome.
  • Large ransom payments can paradoxically lead to a "negative sentiment" in the ecosystem, resulting in lower future payments and higher refusal rates by other victims, while major refusals can deter subsequent attacks.
  • Double extortion significantly increases ransom demands (over an order of magnitude) but also increases the likelihood of victims refusing to pay, especially when attackers partially release data.
  • Susceptibility to specific attack vectors varies significantly by sector: Education and Government are most vulnerable to phishing and social engineering, while IT, Consumer Goods, and Communications sectors face disproportionate impact from Ransomware-as-a-Service (RaaS).

About the Speaker(s)

The research paper "The Ransomware Decade: The Creation of a Fine-Grained Dataset and a Longitudinal Study" was authored by Armin Sarabi, Ziyuan Huang, Chenlan Wang, Tai Karir, and Mingyan Liu, all affiliated with the University of Michigan. Their work contributes significantly to the academic understanding of cybersecurity, particularly in the realm of ransomware analysis and dataset creation.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Solid empirical work that finally gives the field a longitudinal ransomware dataset with actual structure. The chatbot-aided extraction pipeline is the real contribution here—96.7% end-to-end accuracy on messy news articles is genuinely useful. Some of the causal claims are weaker than the descriptive stats, but this is the kind of infrastructure paper that enables better research downstream.

Heather Calloway (CISO) — SOLID

This is the kind of empirical work that should inform your board-level ransomware discussions. A decade of incident data, properly structured, that challenges several assumptions we've been operating on—including whether payment actually makes sense at scale.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)