A Systematic Evaluation of Novel and Existing Cache Side Channels

Fabian Rauscher

Network and Distributed System Security (NDSS) Symposium 2025 · Day 2 · Side Channels 1

Overview

This talk, presented by Fabian Rauscher at the NDSS Symposium, introduces three novel cache side-channel attack primitives utilizing the recently introduced Intel CLDEMOTE instruction. Beyond the discovery of these new attack vectors, the core contribution of this research lies in a comprehensive, systematic evaluation of both these novel attacks and seven existing cache side-channel primitives. The research meticulously compares these attacks across nine distinct metrics, including blind spot size, temporal precision, and covert channel capacity, on two recent Intel microarchitectures: Sapphire Rapids and Emerald Rapids. This rigorous comparative analysis addresses a significant gap in prior research, which often relied on custom, potentially suboptimal implementations of attacks, hindering accurate cross-comparison. The talk also demonstrates the practical implications of the CLDEMOTE instruction's unique properties, specifically its ability to bypass Kernel Address Space Layout Randomization (KASLR) on modern Intel CPUs, even those not officially supporting the instruction.

Watch on YouTube · Slides

Key moments

  1. 0:29 Introduction to Intel's new CLD mode instruction
  2. 1:10 Demonstrating CLD mode timing differences based on cache location
  3. 2:56 Introducing three novel cache side-channel attacks using CLD mode
  4. 4:09 Rationale for systematic evaluation of cache side-channel attacks
  5. 5:21 Overview of nine metrics for comparing cache side-channel attacks
  6. 6:44 Explaining the "blind spot" metric in cache attack evaluation
  7. 7:40 Comparative results: Novel attacks show significantly reduced blind spot

A Systematic Evaluation of Novel and Existing Cache Side Channels

Speakers: Fabian Rauscher

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=tm89r_XG1bQ

Overview

This talk, presented by Fabian Rauscher at the NDSS Symposium, introduces three novel cache side-channel attack primitives utilizing the recently introduced Intel CLDEMOTE instruction. Beyond the discovery of these new attack vectors, the core contribution of this research lies in a comprehensive, systematic evaluation of both these novel attacks and seven existing cache side-channel primitives. The research meticulously compares these attacks across nine distinct metrics, including blind spot size, temporal precision, and covert channel capacity, on two recent Intel microarchitectures: Sapphire Rapids and Emerald Rapids. This rigorous comparative analysis addresses a significant gap in prior research, which often relied on custom, potentially suboptimal implementations of attacks, hindering accurate cross-comparison. The talk also demonstrates the practical implications of the CLDEMOTE instruction's unique properties, specifically its ability to bypass Kernel Address Space Layout Randomization (KASLR) on modern Intel CPUs, even those not officially supporting the instruction.

The significance of this work is multifaceted. Firstly, it uncovers new CPU-level vulnerabilities, expanding the landscape of potential side-channel threats. Secondly, by providing a standardized and exhaustive evaluation framework, it offers a robust benchmark for understanding the true capabilities and limitations of various cache attacks. This systematic approach is critical for both offensive and defensive security research, enabling more informed decisions regarding attack development, mitigation strategies, and the design of secure systems. The KASLR bypass further underscores the subtle yet profound impact new instruction sets can have on system security, even when their intended purpose seems innocuous.

Background

▶ Watch: Introduction to Intel's new CLD mode instruction (0:29)

Cache side channels represent a persistent and significant threat in modern computing, enabling adversaries to infer sensitive information by observing the timing differences in memory access patterns. These timing variations arise because different levels of the CPU cache hierarchy (L1, L2, L3) have varying access latencies. When a victim process accesses data, it leaves a "trace" in the cache state. An attacker, by carefully manipulating and observing this cache state, can deduce information about the victim's operations, even across security boundaries like user/kernel space or different virtual machines. This problem exists due to the shared nature of CPU caches across different processes and privilege levels, a fundamental architectural design choice aimed at improving performance.

Prior work has explored numerous cache side-channel attacks, broadly categorized into techniques like Flush+Reload, Flush+Flush, and Prime+Probe. Flush+Reload involves an attacker flushing a shared cache line, waiting for a victim to potentially access it (thereby bringing it back into cache), and then reloading it to measure the access time. A fast reload indicates a victim access. Flush+Flush operates similarly but measures the time it takes to flush a cache line; if it's in a higher-level cache, flushing takes longer. Prime+Probe involves an attacker filling a cache set with their own data (priming), waiting for a victim to use the same set, and then accessing their own data again (probing) to detect eviction caused by the victim. While these attacks have been widely studied, previous research often suffered from a lack of systematic comparison. Implementations varied, leading to potentially suboptimal performance measurements and making it difficult to draw definitive conclusions about the relative strengths and weaknesses of different attack primitives across diverse microarchitectures. This research aims to address this gap by providing a standardized, comprehensive evaluation.

Key Findings

▶ Watch: Introducing three novel cache side-channel attacks using CLD mode (2:56)

The research yielded several critical findings, primarily revolving around the discovery of new attack primitives and a systematic evaluation of existing and novel techniques:

  1. Novel CLDEMOTE Instruction-Based Attacks: The talk introduces the CLDEMOTE instruction, a recent Intel addition available on server CPUs like CON and Sapphire Rapids. Unlike CLFLUSH, which invalidates a cache line, CLDEMOTE moves a cache line from a higher-level cache (L1, L2) to a lower-level cache (L3). The timing of CLDEMOTE itself is variable: it takes longer if the cache line is in L1 (as it needs to be moved to L3) and is significantly faster if it's already in L3 or not in cache. This timing difference forms the basis for three new attack primitives:
  • DEMOTE and Reload: Similar to Flush+Reload, the attacker CLDEMOTEs a cache line. If the victim then accesses it, it's pulled back into L1. The attacker later measures the access time; a faster access indicates victim activity.
  • DEMOTE and DEMOTE: The attacker repeatedly CLDEMOTEs a cache line. The timing of CLDEMOTE reveals if the cache line was in L1 (slower) or L3 (faster), indicating whether a victim had recently accessed it.
  • DEMOTE Contention: This novel primitive exploits a subtle timing anomaly. If an attacker CLDEMOTEs a cache line (even one not present in the cache hierarchy) and another core simultaneously accesses the same L3 cache set, the attacker's CLDEMOTE operation experiences slower execution times due to contention. This allows cross-core detection of L3 cache set access.
  1. Systematic Evaluation Framework: The paper establishes a robust framework for comparing seven cache attacks (including the three novel CLDEMOTE attacks, Flush+Reload SMT/cross-core, Flush+Flush SMT/cross-core, Prime+Probe L1, and Evict+Reload L1). This evaluation was conducted on two modern Intel microarchitectures, Sapphire Rapids and Emerald Rapids, using nine specific metrics:
  • Hit-to-Miss Margin: The timing difference between detecting an access versus no access.
  • Temporal Precision: How quickly an access can be detected.
  • Spatial Precision: The granularity of detection (e.g., cache set vs. exact cache line).
  • Topological Scope: Whether the attacker and victim need to run on the same core (SMT) or can be on different cores.
  • Attack Time: Overall duration of the attack.
  • Blind Spot Size: The period during which victim accesses might be missed by the attacker.
  • Channel Capacity: The maximum data rate achievable if used as a covert channel.
  • Noise Resilience: Performance degradation in the presence of background noise.
  • Detectability: How easily the attacks can be identified using performance counters.
  1. Specific Performance Characteristics:
  • Blind Spot: Flush+Reload exhibited a significantly high blind spot, over 80% relative to the attack time. In contrast, DEMOTE and Reload had a relatively low blind spot. Notably, Flush+Flush and DEMOTE and DEMOTE were found to have no measurable blind spot, detecting all accesses. DEMOTE Contention showed a unique behavior where its blind spot decreased with increased sleep time, as it's less likely to hit contention during idle periods.
  • Temporal Precision: While Flush+Reload and DEMOTE and Reload on SMT showed similar temporal precision, Flush+Reload cross-core was slower but less spread out, potentially advantageous for specific timing attacks. DEMOTE Contention exhibited a sharp spike in precision, as the contention event is instantaneous.
  • Covert Channel Capacity: The research established specific data rates for various attacks. DEMOTE Contention was relatively slow due to a high error rate. DEMOTE and Reload achieved approximately 6 Mbit/s. Optimized versions of Flush+Reload and DEMOTE and Reload (where the victim performs the flush/demote, and the receiver just accesses memory) significantly boosted capacity, with Optimized DEMOTE and Reload reaching around 17 Mbit/s, making it one of the fastest.
  • KASLR Bypass: A critical finding related to CLDEMOTE is its unusual behavior regarding page faults. Unlike most memory access instructions, CLDEMOTE never triggers a page fault. This means it can be executed on unmapped or inaccessible memory locations without crashing the system. By measuring CLDEMOTE's timing on kernel text segments, the researchers could distinguish between mapped and unmapped pages. If a page's translation is in the TLB (Translation Lookaside Buffer), CLDEMOTE is faster; if not (e.g., for an unmapped page), it's slower due to a page table walk. This timing difference allows an attacker to probe the kernel's memory layout, effectively bypassing KASLR (Kernel Address Space Layout Randomization). This bypass was demonstrated on modern Intel CPUs, including an Intel Core i7 desktop CPU that does not officially support CLDEMOTE. On these unsupported CPUs, CLDEMOTE is treated as a NOP (No Operation) but still performs the TLB lookup, making the KASLR bypass possible.

Technical Deep Dive

▶ Watch: Rationale for systematic evaluation of cache side-channel attacks (4:09)

The technical foundation of this research rests on the exploitation of CPU cache behavior and the specific characteristics of the Intel CLDEMOTE instruction.

The CLDEMOTE instruction, identified by the researchers as a novel primitive for side-channel attacks, is designed to move a cache line from a higher-level cache (L1, L2) to a lower-level cache (L3). Its syntax is straightforward, taking a memory address as an operand. Crucially, the execution time of CLDEMOTE varies significantly based on the initial location of the target cache line:

  • If the cache line is in L1, CLDEMOTE takes longer because it must physically move the data down to L3.
  • If the cache line is already in L3 or not in any cache, CLDEMOTE executes much faster as it has less work to do.

This fundamental timing difference forms the basis for the DEMOTE and Reload and DEMOTE and DEMOTE attacks.

  • In DEMOTE and Reload, the attacker first executes CLDEMOTE on a target cache line, effectively "demoting" it to L3. If a victim then accesses this cache line, the CPU's caching mechanism will pull it back into L1. When the attacker subsequently accesses the same cache line, a fast access time indicates it was in L1 (due to victim activity), while a slow access time indicates it remained in L3 (no victim activity). This is analogous to Flush+Reload, but uses CLDEMOTE instead of CLFLUSH.
  • The DEMOTE and DEMOTE attack leverages the CLDEMOTE instruction's own timing. The attacker repeatedly executes CLDEMOTE on a target cache line. If a victim accesses the cache line between two CLDEMOTE operations, the next CLDEMOTE will be slower (because it has to move the line from L1 to L3), signaling victim activity. If no victim access occurred, the CLDEMOTE will be faster (moving from L3 to L3, or not in cache).

A more subtle and entirely novel attack is DEMOTE Contention. This attack does not require the target cache line to be present in the cache hierarchy. The researchers observed that if an attacker CLDEMOTEs any cache line, and a victim on a different physical core simultaneously accesses the same L3 cache set (even if for an unrelated cache line), the attacker's CLDEMOTE instruction experiences a measurable delay. This implies that CLDEMOTE, even when moving an absent cache line, interacts with the L3 cache's internal structures in a way that can be affected by contention. This allows for cross-core detection of L3 cache set access, offering a new avenue for monitoring activity.

The systematic evaluation employed a rigorous methodology. Seven distinct cache attack primitives were tested:

  1. Flush+Reload (SMT)
  2. Flush+Reload (Cross-core)
  3. Flush+Flush (SMT)
  4. Flush+Flush (Cross-core)
  5. Prime+Probe (L1)
  6. Evict+Reload (L1)
  7. DEMOTE and Reload
  8. DEMOTE and DEMOTE
  9. DEMOTE Contention

These attacks were implemented in C with inline assembly for the critical timing and cache manipulation instructions. The evaluation was performed on Intel Sapphire Rapids and Intel Emerald Rapids microarchitectures, representing the most recent server CPUs at the time of the research.

The nine metrics used for comparison provided a multi-dimensional view of each attack's efficacy:

  • Blind Spot Size: Measured by performing victim accesses at random times and observing how often the attack failed to detect them. For Flush+Reload, this was found to be over 80%. In contrast, Flush+Flush and DEMOTE and DEMOTE achieved 0% blind spot, indicating they could detect all accesses.
  • Temporal Precision: Quantified how quickly an attacker could detect a victim's memory access. DEMOTE Contention showed exceptional temporal precision, as the contention event is instantaneous.
  • Covert Channel Capacity: Established by building basic covert channels using each attack. DEMOTE Contention was slower due to a high error rate. Optimized versions of Flush+Reload and DEMOTE and Reload, where the victim initiates the cache state change and the attacker only measures, achieved significantly higher throughput, with Optimized DEMOTE and Reload reaching approximately 17 Mbit/s.

Perhaps the most striking technical finding pertains to CLDEMOTE's interaction with the Memory Management Unit (MMU) and TLB. Unlike typical memory access instructions that can trigger a page fault if accessing an unmapped or inaccessible memory page, CLDEMOTE never triggers a page fault. This peculiar behavior allows CLDEMOTE to be executed on any virtual address without crashing the process, regardless of its mapping status or access permissions. However, the instruction still interacts with the TLB. If the virtual-to-physical translation for a page is present in the TLB, the CLDEMOTE operation's initial address translation phase is fast. If the translation is not in the TLB (e.g., for an unmapped page or a page whose entry was evicted), a page table walk is required, significantly increasing the CLDEMOTE instruction's execution time. This timing difference allows an attacker to infer whether a given virtual address range corresponds to mapped memory or not. By probing known kernel address ranges, an attacker can determine the base address of the kernel, thereby bypassing KASLR. This KASLR bypass was demonstrated not only on CPUs explicitly supporting CLDEMOTE but also on older Intel CPUs (pre-12th gen, like an Intel Core i7 desktop CPU) where CLDEMOTE is technically an unsupported instruction. On these older CPUs, Intel's documentation suggests it should be treated as a NOP. However, the researchers found that while it performs no cache demotion, it still performs the TLB lookup, making the KASLR bypass attack effective. This highlights a subtle yet critical architectural oversight where an ostensibly benign or unsupported instruction can still leak sensitive information through its side effects.

Demo / Proof of Concept

▶ Watch: Explaining the "blind spot" metric in cache attack evaluation (6:44)

The talk effectively demonstrated the practical implications of their findings through two primary proof-of-concept attacks: a covert channel implementation and a KASLR bypass.

For the covert channel demonstration, the researchers built channels using all evaluated attack primitives. The basic setup involved a sender process and a receiver process. The sender would manipulate the cache state (e.g., by accessing a shared cache line to bring it into L1 or by performing a CLDEMOTE to move it to L3). The receiver would then execute the specific cache side-channel attack (e.g., DEMOTE and Reload, Flush+Flush) to detect the sender's activity and infer the transmitted bit (0 or 1). The demonstration varied parameters such as the receiver's wait time and the time window for sending a single bit. The results provided concrete numbers for channel capacity, with Optimized DEMOTE and Reload reaching approximately 17 Mbit/s, showcasing its potential for high-bandwidth covert communication. This practical demonstration validated the theoretical capacities derived from their systematic evaluation.

The most compelling proof-of-concept was the KASLR bypass using the CLDEMOTE instruction. The core idea relies on CLDEMOTE's unique property of not triggering page faults while still performing TLB lookups. The attacker executes CLDEMOTE on a range of virtual addresses within the expected kernel text segment. By measuring the execution time of CLDEMOTE for each address, the attacker can differentiate between mapped and unmapped pages. If CLDEMOTE executes quickly, it implies the page's translation is in the TLB, suggesting a mapped kernel page. If it executes slowly, it indicates a TLB miss requiring a page table walk, often pointing to an unmapped page. By observing the distinct timing profile – a sudden shift from slow to fast execution – the attacker can pinpoint the start of the kernel's mapped memory region, thus revealing the kernel's base address and effectively bypassing KASLR. The researchers demonstrated this on both systems with explicit CLDEMOTE support (like Sapphire Rapids) and, critically, on an Intel Core i7 desktop CPU (a pre-12th generation model) where CLDEMOTE is officially unsupported. On the unsupported CPU, while the instruction doesn't perform the cache demotion, it still performs the TLB lookup, enabling the timing difference to be observed and the KASLR bypass to succeed. This practical demonstration highlights a significant architectural leakage, proving that even instructions deemed "no-ops" can have security implications through their side effects.

Defensive Implications

▶ Watch: Comparative results: Novel attacks show significantly reduced blind spot (7:40)

The findings presented in this talk have significant implications for system defenders, requiring a multi-pronged approach to mitigation and detection.

Firstly, the introduction of new CLDEMOTE-based attacks means that existing side-channel mitigations, which often focus on Flush+Reload or Prime+Probe, might be insufficient. Defenders need to evaluate if their current strategies protect against DEMOTE and Reload, DEMOTE and DEMOTE, and especially DEMOTE Contention. The cross-core nature of DEMOTE Contention, which can detect L3 set access from a different physical core, poses a particular challenge for isolation mechanisms. System architects should consider if their systems' threat models adequately account for these new primitives.

Secondly, the systematic evaluation provides valuable data for prioritizing defensive efforts. Knowing which attacks have low blind spots (like Flush+Flush and DEMOTE and DEMOTE) or high covert channel capacity (like Optimized DEMOTE and Reload at 17 Mbit/s) allows defenders to focus on mitigating the most potent threats. The detailed metrics can inform the design of more robust isolation mechanisms, such as cache partitioning or randomization techniques, which need to be effective across a broader range of attack types and microarchitectures.

Thirdly, the KASLR bypass using CLDEMOTE on both supported and unsupported CPUs is a critical finding. This demonstrates that simply relying on documented instruction behavior or the absence of explicit support for a feature is insufficient for security. Defenders need to be wary of subtle side effects of instructions, even those that appear benign or are treated as NOPs. Operating system developers should investigate if their kernel's memory access patterns or TLB usage could be similarly fingerprinted by CLDEMOTE or other undocumented instruction behaviors. KASLR is a fundamental security primitive, and its bypass significantly weakens other exploit mitigations. Patches or microcode updates might be necessary to address the CLDEMOTE TLB lookup behavior, particularly on unsupported CPUs where it was not intended to have any effect.

Finally, the talk briefly mentions detectability using performance counters. This offers a potential avenue for real-time monitoring and anomaly detection. System administrators and security analysts could leverage CPU performance monitoring units (PMUs) to track relevant cache events or instruction executions (like CLDEMOTE if it's not expected in user-space applications). Unusual patterns or high frequencies of specific cache operations could indicate an ongoing side-channel attack. However, this often requires detailed knowledge of system-specific performance counter events and careful tuning to avoid false positives. Further research into effective and low-overhead detection mechanisms based on these metrics would be beneficial.

Key Takeaways

  • Novel Attack Primitives: The research introduces three new cache side-channel attacks leveraging the Intel CLDEMOTE instruction: DEMOTE and Reload, DEMOTE and DEMOTE, and DEMOTE Contention.
  • Systematic Evaluation Framework: A comprehensive framework was developed to compare seven existing and novel cache attacks across nine metrics (including blind spot, temporal precision, and covert channel capacity) on Sapphire Rapids and Emerald Rapids microarchitectures.
  • High-Bandwidth Covert Channels: Optimized versions of CLDEMOTE-based attacks, such as Optimized DEMOTE and Reload, can achieve covert channel capacities as high as 17 Mbit/s, making them highly effective for data exfiltration.
  • Zero Blind Spot Attacks: Flush+Flush and DEMOTE and DEMOTE were found to have no measurable blind spot, meaning they can reliably detect all victim accesses.
  • KASLR Bypass through CLDEMOTE: The CLDEMOTE instruction, due to its unique property of not triggering page faults but still performing TLB lookups, can be used to bypass Kernel Address Space Layout Randomization (KASLR) on modern Intel CPUs, including those that officially do not support the instruction.
  • Architectural Nuances Matter: The KASLR bypass highlights that even instructions deemed "unsupported" or "no-ops" by vendors can have critical security implications through their subtle side effects, necessitating a deeper scrutiny of CPU architecture.

About the Speaker(s)

Fabian Rauscher is the speaker for this research presented at the NDSS Symposium. The transcript indicates he is deeply involved in the technical aspects of cache side-channel analysis and CPU microarchitecture, having personally investigated new Intel instructions and conducted extensive evaluations. His work, as detailed in this talk, focuses on understanding and comparing various cache attack primitives and uncovering novel attack vectors.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Solid, original hardware security research that earns its place at a top venue. Three new attack primitives derived from a single underexplored instruction, a rigorous nine-metric comparative framework across two recent microarchitectures, and a KASLR bypass that works even on CPUs where the instruction is officially a NOP — that's a real contribution, not a literature survey in a blazer.

Heather Calloway (CISO) — WEAK

Technically rigorous microarchitecture research with a genuine finding — the CLDEMOTE KASLR bypass on unsupported CPUs is a real discovery worth attention. But this talk never crosses the threshold from research contribution to defender guidance, and it has nothing to say to anyone accountable for a security program.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025