Conducting the Kill-Chain: Detecting APT Progression Through Music-Sequence Modeling
Krupa Brahmkstri, Sneha Rangari
BSidesSF 2026 · Day 1 · AMC Theatre 14
Overview
In an era where cybersecurity defenses are increasingly sophisticated, advanced persistent threats (APTs) continue to bypass detection, often by operating under the radar for extended periods. This talk, "Conducting the Kill-Chain: Detecting APT Progression Through Music-Sequence Modeling," presented by Krupa Brahmkstri and Sneha Rangari from Visa, addresses a fundamental flaw in current security methodologies: the overreliance on detecting individual, isolated events. While security operations centers (SOCs) excel at identifying suspicious "notes"—like a single failed login or a privilege escalation—they frequently miss the "melody" or "composition" of a complete attack chain.
Key moments
- 0:00 Introduction: Why traditional security fails
- 3:30 Equifax case study: Missing the attack melody
- 5:58 Defining attack elements: Note, motive, dissonance
- 7:08 Sequence modeling: The importance of event order
- 8:04 Conducting defense: Metal detector vs. security guard
Conducting the Kill-Chain: Detecting APT Progression Through Music-Sequence Modeling
Speakers: Krupa Brahmkstri, Senior Consultant Level Data Scientist, Visa; Sneha Rangari, Senior Cyber Security Engineer, Visa
Conference: BSides SF
YouTube: https://www.youtube.com/watch?v=q8-GNm5RzKg
Overview
In an era where cybersecurity defenses are increasingly sophisticated, advanced persistent threats (APTs) continue to bypass detection, often by operating under the radar for extended periods. This talk, "Conducting the Kill-Chain: Detecting APT Progression Through Music-Sequence Modeling," presented by Krupa Brahmkstri and Sneha Rangari from Visa, addresses a fundamental flaw in current security methodologies: the overreliance on detecting individual, isolated events. While security operations centers (SOCs) excel at identifying suspicious "notes"—like a single failed login or a privilege escalation—they frequently miss the "melody" or "composition" of a complete attack chain.
Brahmkstri and Rangari introduce a novel approach that borrows principles from music sequence modeling to identify malicious progressions. They argue that just like musical compositions, attacks possess rhythm, utilize motives, and escalate in movements, often sounding harmless when individual actions are considered in isolation. By shifting the focus from "Is this event bad?" to "Does this sequence belong to a malicious composition?", their model aims to detect the full attack narrative before it reaches its final, devastating act, thereby providing defenders with a critical advantage in the cyber kill chain.
Background
▶ Watch: Introduction: Why traditional security fails (0:00)
The core problem highlighted by the speakers is the "tone-deafness" of traditional security information and event management (SIEM) tools. These systems are highly optimized to capture and alert on discrete events, such as a single failed login or an access attempt to a sensitive file. While excellent at providing snapshots of individual activities, they struggle to connect these disparate events into a cohesive narrative, losing the crucial context of a larger story. This limitation allows sophisticated attackers, particularly those behind APTs, to operate slowly and quietly, cooking their compositions over weeks or months, slipping past defenses that are primarily looking for loud, immediate alerts.
A stark illustration of this challenge is the 2017 Equifax breach. Despite having robust security defenses, Equifax's tools failed to connect the dots between events occurring on Day 1, Day 41, and Day 75 of the attack. Attackers resided in the network for 76 days, systematically accessing 51 different databases, ultimately leading to the exfiltration of 147 million data records. The tools were not "broken," but they simply "couldn't hear the whole story" because they lacked the ability to correlate events over extended periods and understand their sequential relationship.
To overcome this, the speakers draw a powerful analogy to music. They define security terms within a musical framework:
- Note: Any single security event or log, such as a failed login or an authentication event.
- Bar: A sequence of combined logs, forming a basic unit of activity.
- Motive: A recurring pattern within a sequence of bars, representing a unique attacker style or behavior, such as specific ways of reusing credentials or hopping between networks.
- Score: The complete representation of the attack kill chain, allowing defenders to visualize the coordinated movements and progression across the entire network.
- Dissonance: The "wrong note feeling"—a mathematically probable abnormal event or sequence that does not fit the expected context or story, indicating malicious activity.
Their proposed solution, sequence modeling, is an artificial intelligence capability designed to understand that the order of events is just as critical as the events themselves. For example, a login at an unusual time, followed by accessing a sensitive file, and then an external upload, is not a normal event; it represents a broken, malicious sequence that traditional tools often miss when evaluating each action in isolation.
Key Findings
▶ Watch: Equifax case study: Missing the attack melody (3:30)
The talk reveals several pivotal findings regarding the application of music sequence modeling to cybersecurity:
- Fundamental Detection Gap: Traditional security tools are inherently "tone deaf" to the progression of sophisticated attacks because they focus on isolated events rather than the overarching sequences and patterns. This allows APTs to operate undetected by leveraging the "dead air" between individual alerts.
- Attackers Operate in Compositions: Malicious actors, particularly APTs, do not merely trigger individual alerts; they execute carefully orchestrated "compositions" characterized by rhythm, recurring motives, and escalating movements. Detecting these structural patterns is crucial for early intervention.
- Structural Dissonance as a Malicious Indicator: By borrowing ideas from music sequence modeling, the presented approach identifies structural dissonance—sequences of events that, while individually benign, collectively indicate a malicious composition. This allows for the detection of attacks "before the final act."
- Significant Performance Improvements: The musical Transformer model demonstrates tangible benefits over traditional rule-based engines and basic machine learning:
- Earlier Detection: The model detected lateral movement attacks 11 minutes earlier than rule-based engines, a critical time advantage for containment.
- Reduced False Positives: False positives were reduced by 22%, saving analyst time and focusing resources on genuine threats.
- Improved Precision: Alert precision was improved by 15%, meaning more meaningful alerts and fewer irrelevant notifications.
- Enhanced Explainability: Unlike black-box detection systems, the model provides analysts with the entire sequence of anomalous activities, offering clear context and reducing analyst fatigue. This explainability empowers SOC teams to understand exactly what happened and why an alert was triggered.
- Augmentation, Not Replacement: Architecturally, this approach functions as a pre-correlation layer that augments existing SIEMs and security tools. It is SIM-agnostic, working seamlessly with platforms like MS Defender or Splunk, strengthening current detections without requiring a complete overhaul of existing infrastructure.
- Behavioral Patterns Over Static Signatures: The talk underscores that "rhythm beats rule" because static signatures are brittle and easily bypassed, whereas behavioral patterns are much harder for attackers to fake. Context, order, proximity, and progression are the true signals, filtering out noise and focusing on meaningful harmonies.
Technical Deep Dive
▶ Watch: Defining attack elements: Note, motive, dissonance (5:58)
The "heart of the engine" behind this innovative detection system transforms raw, chaotic security data into a structured "masterpiece of detection" by applying a series of musically inspired processing steps. The speakers use the analogy of a security guard inside a concert hall, observing behavior over time, rather than just a metal detector at the entrance looking for individual weapons.
The process unfolds in five key stages:
- Transcription: This initial stage is akin to the guard's notebook. Raw security logs, such as file events or process executions, are systematically organized and written down. Instead of merely recording individual actions, the system notes the "rhythm of movements," observing how a person paces the perimeter versus moving directly to their seat. This involves structuring the chaotic log data into an organized format, potentially using models like Fast Track (mentioned as an example of a suitable model). The goal is to move from disconnected logs to a coherent, chronological record of events.
- Motive Detection: Once data is transcribed, the system identifies "shifty behavior" or recurring patterns—the attacker's unique signatures. Just as a guard might recognize someone repeatedly looking at their watch, opening a camera, and whispering into a mic as a pattern indicative of a heist, the system learns characteristic attacker behaviors. This could include specific ways attackers reuse credentials or patterns in how they hop between different network segments. This stage is about learning and recognizing known patterns of malicious activity.
- Harmonic Dissonance: This is where the system's "intuition" comes into play, detecting new or out-of-context patterns that don't fit the expected "story." For instance, if a guard sees a musician in a tuxedo carrying a toolkit instead of a violin, it's immediately apparent that something is amiss. In cybersecurity, this translates to identifying events like a login occurring at a weird time not followed by an expected admin maintenance update. The individual login might be benign, but its lack of expected follow-up, or its occurrence within an unexpected sequence, creates "dissonance"—a sudden, abnormal change that stands out.
- Transformative Attention: Analogous to the guard's long-term memory, Transformative Attention (a concept derived from Transformer neural networks) enables the system to maintain a full contextual understanding of events over extended periods. Unlike traditional systems that might lose context of early events as new ones accumulate (like forgetting the first page of a book by page 75), this mechanism allows the system to remember and correlate activities that happened hours, days, or even weeks apart. For example, it connects a subtle, normal click on a Monday with a sudden, significant "key change" on a Friday, recognizing the full progression of an attack. This is crucial for detecting low-and-slow attacks that unfold over long timeframes.
- Predictive Staff (GPTs): With the full context established by transformative attention, the system leverages Generative Pre-trained Transformers (GPTs) to predict the attacker's next likely action. Just as the guard can predict the tuxedo-wearing individual might be heading to the vault based on prior scouting, the system can anticipate the next move in the attack chain. This allows defenders to intervene preemptively, stopping the attack before the final malicious act occurs, rather than merely reacting to it.
The speakers contrast their musical Transformers approach with existing systems:
- Rule-based systems are basic and easy to set up but lack sequence awareness, leading to high false positives ("screaming fire at every alert").
- Basic machine learning models are more complex and can recognize some patterns ("hum along with your tune") but get confused if context is lost ("don't know the lyrics").
- Musical Transformers, while complex, understand patterns, sequences, and the full context of the story. They filter out noisy background events and alert only on actual "wrong notes," providing precise and timely detections.
Demo / Proof of Concept
▶ Watch: Sequence modeling: The importance of event order (7:08)
The talk demonstrates the efficacy of their music sequence modeling through two real-world case studies: Lateral Movement and Shadow SaaS. These examples highlight how the system detects structural dissonance where traditional methods fail.
Case Study 1: Lateral Movement
This scenario illustrates a typical attack progression that often evades conventional rule-based systems:
- Initial Login: A non-user performs a login. In isolation, this is within normal thresholds and login hours, triggering no alert.
- Service Account Token Generation: The attacker uses a service account to generate a token. This is also a common system activity, within normal thresholds, and generates no immediate alert.
- Remote Execution on a New Host: Briefly after, remote execution occurs on a new host. While potentially abnormal, it might be classified as a low-malicious alert and often gets lost among the high volume of SIEM alerts.
- Privilege Escalation: Within three minutes from the start of the sequence, a privilege escalation occurs. This is undeniably malicious, but in a sea of alerts, it might be deemed low priority for immediate investigation by an analyst.
Crucially, each of these events, when viewed in isolation, might not trigger a high-fidelity alert or prompt immediate action. However, the music sequence model looks at the entire progression: initial login -> token issue -> remote execution -> privilege escalation. This complete sequence creates a structural dissonance spike, immediately flagging it as abnormal.
Results: The model detected this lateral movement 11 minutes earlier than traditional rule-based engines. This early detection is vital, allowing defenders to contain one anomalous host instead of investigating five, significantly reducing analyst workload and preventing further spread. Furthermore, the model provides explainability, presenting the analyst with the full sequence: "login followed by token issues, followed by remote execution on new host and then a privilege escalation." This direct context dramatically reduces analyst fatigue and improves response efficiency.
Case Study 2: Shadow SaaS
This scenario involves attackers using unsanctioned or normal SaaS applications to bypass visibility and controls, blending in as a regular user.
- Valid OAuth Token Issued: A user performs a normal activity, resulting in a valid OAuth token. This appears as a "right note" in isolation.
- Bulk API Export: The user then skips typical workflow (e.g., checking emails) and jumps directly to a bulk API export. While this could be a legitimate activity for some job roles, in this context, it's a deviation.
- Sequential Access to Sensitive Folders: Immediately following the export, there is quick, sequential access to multiple sensitive folders. This is concerning but often gets delayed in investigation due to other environmental activities.
- Recurring Pattern Triggered: The combination of these actions triggers a recurring malicious pattern.
The music sequence model flags this entire sequence, considering the order and timing from the valid OAuth token to the bulk API export and sequential access to sensitive data. Additional contextual factors, such as the activity occurring at 3 AM from a new IP address, are seamlessly integrated into the sequence model's evaluation.
Results: This led to earlier detection, halting data exfiltration before it could leave the system and breaking the "composition" of the attack. Again, the model provided clear explainability to the analyst, making their job much easier.
Overall Impact
The impact of this approach is substantial:
- Precision improved by 15%: Alerts are more meaningful, with fewer false positives.
- False positives reduced by 22%: This translates directly into saved time, reduced analyst fatigue, and a greater focus on real threats.
- Early Detection: The 11-minute earlier detection for lateral movement is a critical advantage, achieved by scoring structural progression rather than waiting for a known signature to complete.
Architecturally, this system acts as a pre-correlation layer, enhancing existing SIEMs without replacing them. It is SIM-agnostic, working with any platform, and significantly augments current detection capabilities.
Defensive Implications
▶ Watch: Conducting defense: Metal detector vs. security guard (8:04)
The insights from "Conducting the Kill-Chain" offer profound implications for security defenders looking to enhance their posture against sophisticated threats:
- Shift from Event-Centric to Sequence-Centric Detection: Defenders must move beyond the traditional focus on individual alert "notes" and embrace a holistic, sequence-aware approach. This means investing in tools and methodologies that can connect disparate events over time to reveal the full "melody" of an attack.
- Implement Behavioral Pattern Recognition: Prioritize security solutions that can identify and learn attacker motives and behavioral patterns, which are far more difficult for adversaries to mask than static signatures. This involves leveraging advanced machine learning, particularly Transformer-based architectures, capable of understanding long-term context.
- Embrace Explainable AI in Security: Seek out detection systems that provide explainability, giving analysts the full narrative of a flagged sequence rather than just a high-fidelity alert. This empowers SOC teams to quickly understand the context, validate threats, and respond effectively, significantly reducing investigation time and analyst fatigue.
- Augment, Don't Replace, Existing Infrastructure: Integrate new sequence-aware detection capabilities as pre-correlation layers that enhance, rather than supplant, existing SIEMs and security tools. This allows organizations to leverage their current investments while gaining a critical edge against advanced threats. The SIM-agnostic nature of such solutions ensures flexibility.
- Focus on Low-and-Slow Attack Detection: Recognize that many APTs operate with extended dwell times. Defensive strategies must incorporate "long-term memory" capabilities (like Transformative Attention) to correlate subtle, seemingly benign events across days or weeks, identifying the gradual progression of an attack.
- Prioritize Contextual Dissonance: Develop detection rules and models that identify events that, while individually normal, are contextually "wrong" within a sequence (e.g., a legitimate action occurring at an unusual time or not followed by an expected administrative task). This involves defining expected "harmonies" of legitimate activity to spot "dissonance."
- Foster Community Collaboration: As highlighted in the talk's future work, engaging in community experimentation around encoding strategies, motive definitions, and transition scoring approaches can collectively strengthen defenses against evolving attack methodologies.
Ultimately, defenders need to cultivate the ability to "hear the music" of an attack, understanding its rhythm, motives, and movements, rather than merely reacting to individual, louder alerts. This paradigm shift enables proactive defense, halting attacks much earlier in their progression.
Key Takeaways
- Attackers, especially APTs, operate in sophisticated sequences and patterns—"compositions"—rather than isolated, individual events or "notes."
- Traditional security tools are often "tone deaf" to these attack compositions, focusing on discrete events and missing the crucial context of an attack's progression.
- Music sequence modeling, particularly using Transformer architectures, provides a powerful framework to detect "structural dissonance" by understanding the order, timing, and full context of event sequences.
- This approach yields significant improvements: 11 minutes earlier detection for lateral movement, a 22% reduction in false positives, and a 15% improvement in precision, leading to more effective and efficient threat response.
- The model functions as a SIM-agnostic pre-correlation layer, augmenting existing security infrastructure and providing explainable alerts that detail the entire malicious sequence to analysts.
- Defenders must shift their focus from making individual alerts "louder" to developing the capability to "hear the music" of an attack's progression, enabling preemptive action and breaking the attack composition before its final act.
About the Speaker(s)
Krupa Brahmkstri is a Senior Consultant Level Data Scientist at Visa, where she is part of the Cyber Threat Analytics and Research team. With over a decade of experience at the intersection of AI and Cybersecurity, Krupa is a recognized expert in her field. She is an active speaker at conferences and universities, has published numerous research papers, and holds several patents, in addition to being a dedicated mentor.
Sneha Rangari is a Senior Cyber Security Engineer at Visa, bringing over eight years of experience in leading global technical infrastructure security and driving technical innovations. She holds prestigious industry certifications including CISSP and SANS GML. Outside of her professional work, Sneha has a passion for music and plays the sitar, which fittingly aligns with the musical analogies presented in this talk.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate applied ML security research from practitioners who clearly built and deployed something real at Visa scale. The music analogy is gimmicky but the underlying work — Transformer-based sequence modeling as a pre-correlation layer for kill-chain detection — is sound engineering. Concrete metrics (11-min earlier detection, 22% FP reduction, 15% precision gain) keep it honest, but the talk stays at the 'what we built and measured' level without going deep enough on the 'how' to be genuinely instructive.
Heather Calloway (CISO) — SOLID
Visa's sequence modeling approach addresses a real and persistent detection gap — the failure to correlate low-and-slow attack progressions — and delivers quantified results that SOC leaders can take seriously. The musical framing is a gimmick, but the underlying technical argument is legitimate. What's missing is any treatment of the organizational and governance conditions that determine whether this actually gets deployed and used.