True Attacks, Attack Attempts, or Benign Triggers? An Empirical Measurement of Network Alerts in a Security Operations Center

Limin Yang (University of Illinois), Phuong Cao, Constantin Adam, Alexander Withers, Zbigniew Kalbarczyk

33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24

Overview

In the increasingly complex landscape of modern cyber threats, Security Operations Centers (SOCs) serve as critical defenses, monitoring vast networks for anomalies and responding to detected incidents. However, despite their vital role, SOCs are frequently overwhelmed by an incessant deluge of security alerts, a problem that significantly hampers their effectiveness and impacts the well-being of their analysts. This talk presents a pioneering quantitative study that delves into the real-world operational challenges of an enterprise SOC, aiming to empirically differentiate between actual successful attacks, mere attack attempts, and benign network activities that trigger alerts.

Watch on YouTube

Visual summary for True Attacks, Attack Attempts, or Benign Triggers? An Empirical Measurement of Network Alerts in a Security Operations Center by Limin Yang, Phuong Cao, Constantin Adam, Alexander Withers, Zbigniew Kalbarczyk
Visual summary for True Attacks, Attack Attempts, or Benign Triggers? An Empirical Measurement of Network Alerts in a Security Operations Center by Limin Yang, Phuong Cao, Constantin Adam, Alexander Withers, Zbigniew Kalbarczyk

Key moments

  1. 0:00 Introduction and overwhelming challenges faced by SOCs
  2. 2:00 Key research questions and NCSA SOC dataset
  3. 3:40 Manual analysis reveals humans as main bottleneck in detection
  4. 4:30 Attack discovery times align with analyst working hours
  5. 5:30 Categorizing alerts: mostly benign or failed attack attempts
  6. 6:00 Detailed breakdown of 'benign triggers' and their types
  7. 7:00 Analysis of attacker persistence and active attack attempts

True Attacks, Attack Attempts, or Benign Triggers? An Empirical Measurement of Network Alerts in a Security Operations Center

Speakers: Limin Yang, University of Illinois; Phuong Cao; Constantin Adam; Alexander Withers; Zbigniew Kalbarczyk

Conference: USENIX Security '24

YouTube: https://www.youtube.com/watch?v=r3KW07s9T4Q

Overview

In the increasingly complex landscape of modern cyber threats, Security Operations Centers (SOCs) serve as critical defenses, monitoring vast networks for anomalies and responding to detected incidents. However, despite their vital role, SOCs are frequently overwhelmed by an incessant deluge of security alerts, a problem that significantly hampers their effectiveness and impacts the well-being of their analysts. This talk presents a pioneering quantitative study that delves into the real-world operational challenges of an enterprise SOC, aiming to empirically differentiate between actual successful attacks, mere attack attempts, and benign network activities that trigger alerts.

Presented by Limin Yang from the University of Illinois, in collaboration with researchers from IBM and NCSA, this work provides a data-driven examination of alert efficacy and SOC bottlenecks. By analyzing extensive datasets comprising both human-curated incident reports and machine-generated network alerts, the research sheds light on the true signal-to-noise ratio within SOC operations. The findings are crucial for understanding the root causes of analyst fatigue and for developing more effective, automated, and human-centric security detection and response strategies, ultimately enhancing the resilience of organizational networks against advanced persistent threats.

Background

▶ Watch: Introduction and overwhelming challenges faced by SOCs (0:00)

Modern enterprise networks are characterized by their immense scale, diverse traffic patterns, and a myriad of services, making them challenging to secure with traditional perimeter defenses like firewalls alone. This complexity necessitates the deployment of sophisticated security operations, typically centralized within a SOC, where dedicated teams monitor network traffic, detect suspicious activities, and orchestrate responses to potential threats. The lifecycle of an incident within a SOC involves detection, analysis, response (e.g., blocking malicious IPs, halting services), incident reporting, and the continuous refinement of detection rules based on attacker behavior.

Despite this structured approach, SOCs globally face profound operational challenges. Numerous studies have highlighted a pervasive sense of being overwhelmed among security analysts. A significant majority, 55%, report a lack of confidence in their ability to prioritize and respond to alerts effectively, while a staggering 70% admit to suffering emotional impacts, exhibiting negative personas such as anxiety or grumpiness due to their demanding work. This human toll is largely attributed to the sheer volume of alerts generated daily, many of which prove to be false positives or benign events. Critically, there has been a notable gap in quantitative research to truly understand these problems, primarily due to the sensitive nature of real-world SOC data, which organizations are often reluctant to share. This study directly addresses this gap by securing and analyzing a unique dataset from the National Center for Supercomputing Applications (NCSA) SOC. The research was driven by three core questions: identifying key bottlenecks in SOC detection, quantifying the excessiveness of security alerts and their underlying reasons, and assessing the effectiveness of alerts in indicating true, successful attacks.

Key Findings

▶ Watch: Manual analysis reveals humans as main bottleneck in detection (3:40)

The empirical analysis yielded several critical findings that illuminate the operational realities and inherent challenges within a security operations center:

  1. Human Bottlenecks in Incident Response: The study revealed significant human-centric bottlenecks. Post-attack analysis, which occurs after an attack has been detected, was found to take an average of 53.2 days, highlighting a considerable delay in understanding and documenting incidents. Furthermore, 65.42% of the 107 incidents with documented analyst involvement required two or more analysts, with some even necessitating collaboration from personnel outside the core SOC team, underscoring the labor-intensive nature of investigations. An interesting temporal pattern emerged: while attack start times were evenly distributed, attack discovery times heavily correlated with analyst working hours (8 AM to 6 PM), indicating potential delays in detecting attacks that occur during off-hours.
  1. Overwhelming Alert Volume and Composition: The NCSA SOC experienced an initial period of over 100,000 alerts daily, which later dropped to approximately 25,000 daily after the implementation of a Blackhole Router (BHR) system. Despite this reduction, the volume remained substantial and highly dynamic. Crucially, the research categorized alerts into four types, revealing that alerts directly related to true attacks (successful system compromises) constituted a very small portion of the total. A significant majority were either benign triggers (at least 48.91%) or attack attempts, with 23.72% remaining in an "unknown" category due to insufficient information. This imbalance underscores the immense signal-to-noise problem faced by analysts.
  1. Nature of Benign Triggers and Attack Attempts: Benign triggers, though correctly fired, had justified, non-malicious reasons. These included internal vulnerability scans, proactive misconfiguration checks, and external penetration tests conducted by third-party vendors. Identifying and categorizing these required significant manual effort due to inconsistent documentation. Attack attempts, primarily identified via BHR blocks, revealed varied persistence: the vast majority were short-lived, lasting less than one day. However, a notable tail of persistent attack attempts existed. Interestingly, the most active "attacker" IP was discovered to be an undocumented internal scanner used by an SOC member to test the BHR system, highlighting the complexity of distinguishing legitimate internal activities from external threats.
  1. Uneven Distribution of Victim Alerts: The analysis showed an uneven distribution of alerts across victim IPs. A very small fraction of victim IPs contributed to a disproportionately large number of alerts. This skewed distribution presents significant challenges for developing and deploying per-host security prediction models, as the data for the majority of hosts might be too sparse or inconsistent for effective modeling.
  1. Potential for Identifying True Attacks via Abnormal Patterns: Despite the overwhelming alert volume, the study identified a promising opportunity to distinguish true attacks. By developing a methodology to assign a Rarity Value (based on frequency ranking) to each alert type and then calculating a Rarity Score for an IP on a given day (sum of rarity values of unique alerts on that IP), the researchers found that alerts associated with true attacks often exhibited abnormal patterns. Specifically, for the true attacks where connection logs were available, the attack IPs triggered significantly rarer alert types on the days of the attacks, placing them in an "outlier area" compared to normal alert activity. This suggests that focusing on rare or unusual alert combinations could be a viable strategy for prioritizing alerts indicative of genuine threats.

Technical Deep Dive

▶ Watch: Attack discovery times align with analyst working hours (4:30)

The study leveraged a unique and comprehensive dataset derived from the NCSA SOC, a large-scale environment with thousands of servers and a global user base. This dataset comprised two primary components: 227 true attack incidents spanning 20 years, documented in unstructured reports, and over 100 million Zeek alerts collected over a four-year period. Zeek, formerly known Bro, is an open-source network security monitoring tool that provides a high-level, scriptable language for analyzing network traffic and generating detailed logs and alerts based on custom policies or detected anomalies.

The initial phase of the research involved a meticulous manual labeling process of the incident reports to understand how attacks successfully breached systems. This involved mapping identified breaking methods to techniques described in the MITRE ATT&CK framework, a globally accessible knowledge base of adversary tactics and techniques based on real-world observations. Out of the 227 incidents, 178 had identifiable breaking methods. The analysis revealed a significant drop in the diversity and number of attack techniques post-2011, a period correlating with substantial infrastructure upgrades at NCSA, including the widespread implementation of multi-factor authentication (MFA). However, even after these upgrades, some user accounts were still compromised due to misconfigured MFA, underscoring that technical controls, if improperly deployed, can still leave vulnerabilities.

For the alert data, the researchers quantitatively measured the daily alert volume, observing a dramatic initial rate of over 100,000 alerts per day, which subsequently decreased to around 25,000 daily after the implementation of a Blackhole Router (BHR) system. A BHR is a mechanism used to automatically drop traffic from specific IP addresses, effectively blocking known malicious sources. In this context, the BHR was configured to automatically handle certain alerts and block identified attacker IPs, significantly reducing the load on human analysts.

A crucial aspect of the technical analysis was the categorization of the more than 100 million Zeek alerts into four distinct types:

  1. Benign Triggers: These were alerts that fired correctly but for legitimate, non-malicious reasons. Examples identified through manual review with SOC analysts included internal network scans (proactively checking for vulnerabilities or misconfigurations) and external penetration tests conducted by third-party security vendors. These alerts constituted at least 48.91% of the total, indicating a massive volume of "noise" that analysts must sift through.
  2. Attack Attempts: These alerts indicated attempts to compromise the system that ultimately failed to achieve a successful breach. They were primarily identified as alerts that triggered a block by the BHR system but were not classified as benign triggers.
  3. True Attacks: These were alerts directly correlated with incidents where a system was successfully compromised. The study found these to be a very small fraction of the overall alerts, highlighting the difficulty of finding the "needle in the haystack."
  4. Unknown: Approximately 23.72% of alerts could not be definitively categorized due to insufficient contextual information.

Further analysis of attack attempts involved examining their persistence. Using a Cumulative Distribution Function (CDF) plot, the researchers showed that most attack attempts were short-lived, with the majority lasting less than one day. However, a "long tail" of persistent attack attempts was observed, indicating more dedicated adversaries or automated scanning campaigns. A detailed investigation into the top 15 most active "attackers" revealed an interesting anomaly: the leading source of attack attempts was an undocumented internal scanner set up by an SOC member to test the BHR, again illustrating the challenges of internal visibility and documentation.

The study also investigated the distribution of alerts among victim IPs, finding it highly uneven. A small subset of victim IPs generated a disproportionately large number of alerts, which has implications for the effectiveness of security prediction models that often rely on per-host analysis.

Perhaps the most technically significant part of the research involved attempts to link the massive alert dataset with the relatively few true attack incidents. Due to the lack of ground-truth mapping between alerts and incidents, this required significant manual effort, correlating alerts with incidents based on IP addresses and timestamps. Over the four-year alert period, 11 true attacks occurred. The researchers successfully linked related Zeek alerts to 5 of these attacks, and found evidence that Zeek alerts were triggered for two additional incidents, though the specific logs were missing. This meant that 7 out of 11 true incidents generated corresponding alerts.

To address the challenge of identifying true attacks amidst excessive alerts, the researchers proposed and tested a novel anomaly detection approach based on alert patterns. This method involved:

  1. Rarity Value Assignment: Each unique type of alert was assigned a "Rarity Value" based on its frequency ranking within the entire dataset. Rarer alert types received higher rarity values.
  2. Rarity Score Calculation: For each IP address on a given day, a "Rarity Score" was computed by summing the Rarity Values of all unique alert types triggered by that IP on that day.

This analysis was performed using two years of connection logs, covering true attacks number 9, 10, and 11 from their incident dataset. The results indicated that on the days these true attacks occurred, the attacking IPs triggered alert types with significantly higher Rarity Scores, placing them in an "outlier area" when visualized against normal alert distributions. This demonstrates a promising opportunity to prioritize alerts that indicate true attacks by focusing on abnormal, rare alert patterns rather than just high volumes.

Demo / Proof of Concept

▶ Watch: Detailed breakdown of 'benign triggers' and their types (6:00)

While the talk did not feature a live software demonstration or a deployable proof-of-concept tool, the core of the presentation served as a compelling demonstration of the analytical methodology and its potential. The comprehensive empirical measurement and classification of alerts, coupled with the proposed "Rarity Score" technique, effectively functioned as a proof-of-concept for a data-driven approach to distinguishing genuine threats from noise. The successful application of the Rarity Score to historical true attack data, showing that attack-related alerts occupied an "outlier area," provides strong evidence that this method holds promise for improved alert prioritization in real-world SOC environments. The research itself, through its quantitative findings and proposed techniques, constitutes a robust demonstration of concept for enhancing network intrusion detection.

Defensive Implications

▶ Watch: Analysis of attacker persistence and active attack attempts (7:00)

The findings of this study carry profound implications for how security operations centers approach network intrusion detection and incident response. Defenders must move beyond a simplistic binary classification of alerts as "benign" or "malicious." The research strongly advocates for a more nuanced categorization that explicitly distinguishes between true attacks (successful compromises), attack attempts (failed breaches), and benign triggers (legitimate activities that correctly fire alerts). Incorporating these categories into benchmark datasets and evaluation processes is crucial for developing and testing detection systems that truly reflect real-world SOC challenges.

Addressing the human bottleneck is paramount. The average 53.2 days for post-attack analysis and the observed delays for off-hour attack detection underscore the need for streamlined workflows. SOCs should invest in efficient ways to store, link, and curate diverse logs from various security tools and network devices. A unified, searchable, and contextualized logging infrastructure can significantly reduce the time and effort analysts spend piecing together incident timelines, thereby speeding up investigations.

Furthermore, the overwhelming volume of alerts, especially during off-hours when human availability is limited, necessitates a greater emphasis on automation. Implementing intelligent, automated response systems can help handle a large portion of attack attempts and known benign triggers, freeing up analysts to focus on the high-fidelity alerts that genuinely warrant human investigation. This automation can range from automated blocking of suspicious IPs (like the Blackhole Router (BHR) system used by NCSA) to automated enrichment of alerts with contextual data.

Perhaps the most actionable defensive implication stems from the "abnormal alert patterns" finding. SOCs should explore and integrate anomaly detection techniques that leverage the concept of Rarity Value and Rarity Score. Instead of solely focusing on the volume or specific type of alert, systems should be designed to identify and prioritize alerts that represent unusual combinations or rare events for a given IP over time. This shift in focus from "what" is happening to "how unusual" it is could dramatically improve the signal-to-noise ratio, allowing analysts to focus their limited attention on alerts that are genuinely indicative of a true attack. This requires robust logging of all alert types and their frequencies to establish a baseline for rarity.

Finally, the uneven distribution of alerts across victim IPs suggests that security prediction models need to account for this sparsity. Per-host analysis might be less effective for the majority of hosts, implying a need for more aggregated or network-wide anomaly detection strategies, or the development of models that can effectively handle sparse data for individual endpoints.

Key Takeaways

  • SOCs are Overwhelmed by Noise: The vast majority of security alerts in a real-world SOC are not indicative of true, successful attacks, with benign triggers and attack attempts dominating the alert landscape.
  • Human Factors are Major Bottlenecks: Manual incident investigation is time-consuming (averaging 53.2 days), often requires multiple analysts, and off-hour attacks face detection delays due to human availability.
  • Nuanced Alert Classification is Essential: Distinguishing between benign triggers, attack attempts, and true attacks is crucial for accurate threat assessment and improving the effectiveness of detection systems and benchmark datasets.
  • Automation and Efficient Log Management are Critical: To combat excessive alerts and speed up investigations, SOCs urgently need automated response mechanisms, especially for off-hours, and efficient systems for storing, linking, and curating diverse security logs.
  • Abnormal Alert Patterns Offer Promising Detection: True attacks tend to trigger rare and unusual combinations of alerts, suggesting that focusing on "rarity scores" and outlier patterns can significantly improve the prioritization and identification of genuine threats.
  • Real-World Data is Indispensable: Quantitative studies using real-world SOC data are vital for truly understanding operational challenges and developing effective, evidence-based solutions for network security.

About the Speaker(s)

The primary presenter for this talk was Limin Yang, a researcher from the University of Illinois Urbana-Champaign (UIUC). The work presented is a collaborative effort involving a team of researchers including Phuong Cao, Constantin Adam, Alexander Withers, and Zbigniew Kalbarczyk, with affiliations noted to include IBM and NCSA (National Center for Supercomputing Applications). Their collective expertise spans network security, security operations, and empirical analysis, contributing to a comprehensive understanding of real-world SOC challenges.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This talk presents a rare, data-driven analysis of real-world SOC alerts, empirically quantifying the signal-to-noise problem. The research provides critical insights into human bottlenecks and introduces a novel 'Rarity Score' method to prioritize true attacks, making it essential for any defender.

Heather Calloway (CISO) — STRONG ACCEPT

This empirical study quantifies the critical signal-to-noise problem in SOCs, revealing significant human bottlenecks and the overwhelming volume of benign and attempted alerts. Its core contribution is an evidence-based approach to prioritizing true attacks by identifying rare alert patterns, providing a concrete path for CISOs to improve operational efficiency and resilience.

→ Top-rated talks at 33rd USENIX Security Symposium

All talks from 33rd USENIX Security Symposium