Systematic Evaluation of Randomized Cache Designs against Cache Occupancy

Anirban Chakraborty (Researcher · Max Planck Institute for Security and Privacy)

34th USENIX Security Symposium (USENIX Security '25) · Day 2 · Hardware Security 1: Microarchitectures

Overview

This talk, presented by Anirban Chakraborty from the Max Planck Institute for Security and Privacy, delves into a comprehensive evaluation of randomized cache designs, focusing on both their performance characteristics and, crucially, their resilience against cache occupancy attacks. While randomized caches have been widely adopted as a defense mechanism against well-known side-channel attacks like Prime+Probe, their efficacy against the broader class of occupancy-based attacks has largely been overlooked. This research highlights a significant gap in current security evaluations, demonstrating that many state-of-the-art randomized cache designs, despite their complex randomization schemes, remain vulnerable to information leakage through cache occupancy.

Watch on YouTube · Slides

Visual summary for Systematic Evaluation of Randomized Cache Designs against Cache Occupancy by Anirban Chakraborty
Visual summary for Systematic Evaluation of Randomized Cache Designs against Cache Occupancy by Anirban Chakraborty

Key moments

  1. 0:00 Introduction and background on Prime+Probe attacks
  2. 2:15 Understanding Cache Occupancy attacks and their significance
  3. 3:20 Motivation: Systematic performance and security evaluation
  4. 5:00 Proposed benchmarking strategy for fair performance evaluation
  5. 6:50 Key finding: Varied performance effects of replacement policies
  6. 7:30 Setting up a covert channel for security evaluation
  7. 8:45 Security evaluation results across different cache designs

Systematic Evaluation of Randomized Cache Designs against Cache Occupancy

Speakers: Anirban Chakraborty, Researcher, Max Planck Institute for Security and Privacy

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=ygk79RDrb24

Overview

This talk, presented by Anirban Chakraborty from the Max Planck Institute for Security and Privacy, delves into a comprehensive evaluation of randomized cache designs, focusing on both their performance characteristics and, crucially, their resilience against cache occupancy attacks. While randomized caches have been widely adopted as a defense mechanism against well-known side-channel attacks like Prime+Probe, their efficacy against the broader class of occupancy-based attacks has largely been overlooked. This research highlights a significant gap in current security evaluations, demonstrating that many state-of-the-art randomized cache designs, despite their complex randomization schemes, remain vulnerable to information leakage through cache occupancy.

The work addresses two critical aspects: first, a standardized performance evaluation of various randomized cache designs across multiple replacement policies, addressing inconsistencies in prior benchmarking efforts; and second, a rigorous security analysis against cache occupancy attacks. The findings reveal surprising performance variations and, more importantly, expose severe security vulnerabilities, including a practical AES key recovery attack. This article will explore the methodologies, findings, and implications of this groundbreaking research, solidifying cache occupancy attacks as a paramount concern for secure cache design.

Background

▶ Watch: Introduction and background on Prime+Probe attacks (0:00)

The landscape of modern computing is riddled with side-channel vulnerabilities, where sensitive information can be inferred not by directly compromising data, but by observing the physical effects of computation. Among the most prevalent of these are cache attacks, which exploit the timing differences in memory access patterns. The Prime+Probe attack is a classic example, where an attacker first "primes" a specific cache set by filling it with their own data. After allowing a victim process to execute, the attacker "probes" the same cache set. If the victim's data accessed that set, some of the attacker's data would have been evicted, leading to slower access times for those addresses. This timing disparity can leak highly sensitive information, having been demonstrated to extract cryptographic keys, log keystrokes, and facilitate browser-based attacks.

To counter these sophisticated attacks, randomized caches emerged as a promising defense strategy. The core principles behind randomized caches involve two main mechanisms:

  1. Address Randomization: Cryptographic operations are employed to randomize the mapping between virtual or physical memory addresses and their corresponding cache sets. This aims to disrupt the attacker's ability to reliably predict which cache set a victim's data will occupy.
  2. Cache Partitioning: The cache is logically or physically divided into multiple partitions. This further enhances randomness and isolates different processes or security domains, making it harder for an attacker to observe the victim's cache activity.

However, the focus of most randomized cache designs has predominantly been on mitigating Prime+Probe attacks, which are highly granular, targeting specific cache sets. A less granular, but equally potent, class of attacks are cache occupancy attacks. Unlike Prime+Probe, which focuses on a single cache set, cache occupancy attacks aim to understand the overall "footprint" of a victim process across the entire cache. In such an attack, the adversary attempts to fill a significant portion of the cache with their own data. After the victim executes, the attacker re-accesses their data and measures how many of their own cache lines have been evicted. The number of evictions provides an indication of the victim's memory footprint, thereby leaking information about its execution. Due to their less granular nature, many designers of randomized caches have historically considered cache occupancy attacks a less significant threat, often overlooking them in their security evaluations. This assumption forms the primary motivation for the presented work, which challenges this view by systematically evaluating the resilience of state-of-the-art randomized cache designs against these often-ignored attacks.

Furthermore, the research identifies a significant challenge in comparing existing randomized cache designs: the lack of a common, fair baseline for performance evaluation. Different research works often utilize diverse benchmarks and make specific implementation assumptions, making a direct, head-to-head comparison extremely difficult. For instance, the Mirage cache design, presented at a previous security conference, made assumptions about uniform distributions of loads and stores across the lifetime of a benchmark workload, potentially leading to an overly optimistic view of its performance overheads. This work aims to rectify this by proposing a standardized benchmarking strategy, ensuring a fair assessment of performance across various designs and replacement policies.

Key Findings

▶ Watch: Motivation: Systematic performance and security evaluation (3:20)

The comprehensive evaluation presented in this work yielded several critical findings across both performance and security domains, fundamentally reshaping our understanding of randomized cache designs:

  1. Varied Performance Impact of Replacement Policies: Contrary to the common practice of evaluating randomized caches primarily with random replacement policies, this research demonstrates that different cache replacement policies have a significantly varied impact on the performance of randomized cache designs. A surprising discovery was that, for certain replacement policies, some randomized cache designs actually outperform the baseline set-associative cache, a finding previously unobserved and counter-intuitive given the overheads associated with randomization. This highlights the importance of evaluating designs across a broader spectrum of policies.
  1. Mirage's Vulnerability to Cache Occupancy: In the context of covert channel attacks, the Mirage cache design was found to be exceptionally vulnerable to cache occupancy. This vulnerability stems from its inherent policy of compulsory eviction for every new address installation from its data store. This design choice, intended to enhance randomization, ironically makes it easier for an attacker to detect evictions and thus infer victim activity, making it more susceptible to occupancy attacks than even a standard, non-randomized set-associative cache.
  1. SAS Cache's Superior Resilience in Process Fingerprinting: When subjected to process fingerprinting attacks, the SAS cache (Set-Associative Secure cache) exhibited remarkably low prediction accuracy. This superior resilience is attributed to SAS cache's design principle of separating the security domains of the victim and attacker processes, effectively minimizing information leakage. In contrast, other evaluated randomized cache designs showed significantly higher fingerprinting accuracy, indicating substantial adversarial advantage.
  1. First Practical AES Key Recovery via Cache Occupancy: A pivotal contribution of this work is the demonstration of the first successful AES key recovery attack on T-table AES using cache occupancy levels and guessing entropy. This attack conclusively proves that cache occupancy is not merely a theoretical concern but a practical, exploitable side channel capable of compromising sensitive cryptographic operations. The attack's success establishes cache occupancy attacks as a "serious concern" on par with traditional contention-based attacks like Prime+Probe.
  1. Design Principles for Occupancy Resilience: The research unequivocally concludes that randomized cache designs based on SAS set-associative caches are significantly more resilient to cache occupancy attacks compared to designs based on pseudo-fully associative caches. This finding provides crucial guidance for future secure cache architecture design, advocating for specific architectural choices that inherently offer better protection against occupancy-based side channels.

Technical Deep Dive

▶ Watch: Proposed benchmarking strategy for fair performance evaluation (5:00)

The research undertakes a rigorous technical evaluation, encompassing both performance and security analyses. The methodology for performance evaluation addresses the inconsistencies in prior work, while the security analysis meticulously constructs and executes three distinct cache occupancy attacks.

Performance Evaluation Methodology:

To provide a fair and common baseline, the researchers adopted a standardized benchmarking strategy using the GEM5 full-system simulator.

  • Platform: GEM5 simulator.
  • Workload: The evaluation employed a two-phase workload. First, a spurious cache occupancy phase served as a cache warming stage, ensuring a consistent initial state for the cache. This phase specifically removes the problematic implementation assumptions, such as uniform memory access patterns, that plagued prior evaluations of designs like Mirage. This was then followed by the execution of SPEC benchmarks, which represent a diverse set of real-world applications.
  • Cache Configuration: The Last-Level Cache (LLC) size was set to 16 MB, with two SKUs (Security Kernel Units, wherever applicable). The L1 Data cache (L1 D-cache) was 512 KB, and the L1 Instruction cache (L1 I-cache) was 2 KB.
  • CPU Model: A timing simple CPU model was used for accuracy in timing measurements.
  • Replacement Policies: A crucial aspect of the performance evaluation was testing across five different LLC replacement policies. While the talk specifically mentions Random Replacement and RIP (Recency-based Insertion Policy), the implication is that a broader set was considered. This contrasts sharply with prior works that often limited their evaluations to a single replacement policy, typically random.
  • Encryption Latency: Since many randomized cache schemes incorporate cryptographic block ciphers for address randomization, a uniform encryption latency of three cycles was assumed for these operations.

The performance results demonstrated that different randomized cache designs exhibited varied effects across these replacement policies. Notably, for certain replacement policies, some randomized designs surprisingly achieved better performance than the baseline set-associative cache, a novel observation.

Security Evaluation Methodology (Three-Stage Attack Framework):

  1. Covert Channel Attack:
  • Adversarial Model: This attack assumes the strongest adversarial model, where a sender and receiver actively cooperate to establish a covert channel through the cache.
  • Mechanism: The receiver first allocates sufficient memory space and randomly accesses L number of addresses, effectively "installing" them in the cache. The sender, depending on the bit value (0 or 1) it intends to transmit, performs either X or Y memory accesses. During this period, the receiver enters a busy-wait or sleep state. Once the receiver wakes up, it re-accesses its L addresses. By observing the number of its own addresses that have been evicted from the cache, the receiver can infer whether the sender performed X (sending 0) or Y (sending 1) memory accesses.
  • Finding: Mirage was identified as particularly vulnerable due to its policy of compulsory eviction for every new address installation, making evictions easily detectable and exploitable for covert channel communication.
  1. Process Fingerprinting Attack:
  • Setup: This attack involves an attacker process running on CPU 0 and a victim process running on CPU 1.
  • Mechanism: The attacker first sets up the cache by filling it with its own data and then enters a busy-wait state. While the attacker waits, the victim process executes a program. The attacker then "monitors" the cache by re-accessing its own addresses, observing which ones have been evicted by the victim's execution. Based on the eviction patterns, the attacker attempts to perform templating to identify or "fingerprint" the specific program or process the victim was running.
  • Workloads: The SPEC 2017 workloads were used as benchmarks for victim processes.
  • Cache Designs Evaluated: Caesar, SAS cache, Mirage, Scatter, and SAS cache.
  • Finding: SAS cache demonstrated remarkably low prediction accuracy, implying minimal adversarial advantage in process fingerprinting. This is attributed to its design that separates the security domains of victim and attacker processes. Other cache designs, however, showed high fingerprinting accuracy.
  1. AES Key Recovery Attack:
  • Target: AES T-table implementation, a common target for cache-based side-channel attacks.
  • Initial State: The attacker first fills the LLC with spurious occupancy to ensure that the AES T-tables are flushed from both L1 and LLC. The attacker then establishes a chosen cache occupancy level (e.g., X% of the cache) by accessing that amount of memory, ensuring this occupancy across L1 D-cache and LLC.
  • Chosen Plaintext Attack: The AES victim runs a single encryption of a plaintext P to obtain ciphertext C. The attacker simultaneously measures the access time (T) to its previously allocated memory (its cache occupancy). A single observation point for the attacker is a tuple: (P, C, T), representing the time taken to access the attacker's cache occupancy after the victim's AES operation on P with an unknown key K.
  • Template Creation (Simulation): The attacker, using a known key K_star, simulates the AES encryption process and measures its own access times to create a set of templates, TJ_star, where J varies from 0 to 255 (representing possible byte values). This generates templates for each possible key byte value.
  • Attack Phase (Victim): The victim performs AES encryptions with an unknown key K. The attacker measures timings (tji) for its own addresses in the cache occupancy for each key byte index i (0-15) and each possible byte value (0-255). This results in 256 * 16 observation templates.
  • Correlation and Key Recovery: For each key byte index i (0-15), the attacker correlates its observed templates (tji) with its simulated templates (ti_star). The correlation matrix is then sorted, and the rank of the correct key bytes (r0 to r15) is calculated. Finally, the guessing entropy for each key byte is computed. Lower guessing entropy implies a higher probability of correctly guessing the key byte, thus enabling faster key recovery.
  • Experiments: Two types of experiments were conducted:
  • Varying occupancy rate: 40%, 50%, and 70%.
  • Fixed occupancy (50%) while varying the number of observations.
  • Finding: Designs based on SAS set-associative caches were found to be more resilient to this AES key recovery attack compared to designs based on pseudo-fully associative caches.

This detailed technical approach provides strong evidence for the vulnerability of current randomized cache designs to cache occupancy attacks and offers concrete insights into the architectural features that contribute to or mitigate these risks.

Demo / Proof of Concept

▶ Watch: Setting up a covert channel for security evaluation (7:30)

While the talk did not feature a live, interactive demonstration in the traditional sense, the "security evaluation" section effectively serves as a series of proof-of-concept (PoC) attacks that empirically demonstrate the exploitability of randomized caches via cache occupancy. These experiments showcase practical attack methodologies and quantify the information leakage.

The first PoC was a covert channel establishment. This experiment demonstrated how an attacker (sender) could transmit information (bits 0 or 1) to a receiver by manipulating the cache occupancy. The receiver, by observing evictions of its own data, could reliably infer the transmitted bit. The results highlighted that Mirage, due to its design principle of compulsory evictions, was particularly susceptible, making it an effective medium for such a channel. This experiment proves that even in a controlled environment, specific design choices in randomized caches can inadvertently create new side-channels.

The second PoC involved process fingerprinting. Here, the researchers demonstrated that by monitoring cache occupancy patterns, an attacker could accurately identify the specific programs or processes being executed by a victim. Using SPEC 2017 workloads as victim programs, the attacker's ability to "template" and recognize these processes was measured across various randomized cache designs. The results clearly showed that for most designs, fingerprinting accuracy was high, indicating a significant information leak. Only SAS cache, with its strong security domain separation, provided robust protection against this form of information leakage, showing remarkably low prediction accuracy. This PoC underscores the risk of adversaries inferring sensitive operational details about victim systems.

The most impactful PoC was the AES key recovery attack. This demonstration was a full-fledged, end-to-end attack pipeline, proving that cache occupancy could be leveraged to extract cryptographic keys. The attacker meticulously engineered the cache state by inducing spurious occupancy, then monitored eviction timings during a victim's AES encryption using a chosen plaintext. By correlating these timing traces with pre-simulated templates, the attacker was able to calculate the guessing entropy for each byte of the AES key. The success of this experiment, quantified by the reduction in guessing entropy, provided concrete evidence that designs like Scatter and Mirage are vulnerable to key extraction, while SAS cache offered significantly better resilience. This PoC is critical as it moves beyond mere information leakage to direct compromise of cryptographic secrets, establishing cache occupancy as a severe and practical threat.

These "demonstrations" are not just theoretical constructs but empirically validated attacks, providing concrete evidence of the vulnerabilities and the relative strengths of different randomized cache designs against cache occupancy attacks.

Defensive Implications

▶ Watch: Security evaluation results across different cache designs (8:45)

The findings of this systematic evaluation carry profound implications for the design and implementation of secure cache architectures, particularly in an era where side-channel attacks are a persistent threat.

  1. Prioritize Cache Occupancy in Threat Models: Designers of randomized caches and other secure memory systems must explicitly include cache occupancy attacks in their threat models. The assumption that randomized caches primarily protect against Prime+Probe is insufficient. This work conclusively demonstrates that occupancy attacks are a serious concern, capable of leaking sensitive information and even facilitating cryptographic key recovery.
  1. Favor SAS Set-Associative Designs: The research strongly indicates that randomized cache designs based on SAS (Set-Associative Secure) set-associative caches offer superior resilience against cache occupancy attacks compared to those based on pseudo-fully associative caches. Future designs should either adopt the principles of SAS cache, which emphasize the separation of security domains, or at least draw inspiration from its architectural choices to enhance protection. Specifically, the ability of SAS cache to isolate victim and attacker processes effectively minimizes observable leakage.
  1. Avoid Compulsory Eviction Policies: Designs like Mirage, which incorporate a policy of compulsory eviction for every new address installation, are demonstrably more vulnerable to cache occupancy attacks. While such policies might aim to enhance randomization, they create easily detectable eviction events that attackers can exploit for covert channels or information leakage. Designers should carefully evaluate the trade-offs of such mechanisms and consider their impact on occupancy-based side channels.
  1. Comprehensive Performance Evaluation: The study highlights the inadequacy of evaluating cache designs with only a single replacement policy (e.g., random replacement). A robust performance evaluation must consider multiple replacement policies to fully understand the design's behavior and overheads. Furthermore, benchmarking strategies should employ a fair baseline, such as the proposed use of spurious cache occupancy as a warming phase, to remove optimistic implementation assumptions and ensure accurate, comparable results across different designs.
  1. Security Domain Separation: The success of SAS cache in mitigating process fingerprinting attacks underscores the importance of strong security domain separation within cache architectures. Mechanisms that effectively isolate the cache activities of different processes or security contexts are crucial for preventing information leakage through shared cache resources.
  1. Continuous Re-evaluation of Existing Defenses: Given the dynamic nature of side-channel research, existing randomized cache designs, even those considered secure against other attacks, must be continuously re-evaluated against newly identified or re-emphasized threats like cache occupancy. This proactive approach is essential to maintain the long-term security of computing platforms.

By integrating these defensive implications into future cache design principles, the industry can move towards building more robust and truly secure computing systems that are resilient against a broader spectrum of sophisticated side-channel attacks.

Key Takeaways

  • Cache occupancy attacks are a serious and practical threat, capable of leaking sensitive information and enabling cryptographic key recovery, on par with contention-based attacks like Prime+Probe.
  • Many state-of-the-art randomized cache designs are vulnerable to cache occupancy attacks, despite their efforts to protect against other side channels.
  • The Mirage cache design is particularly susceptible to cache occupancy due to its compulsory eviction policy, making it vulnerable to covert channels and information leakage.
  • SAS set-associative caches (e.g., SAS cache) offer significantly superior resilience against cache occupancy attacks, demonstrating lower accuracy in process fingerprinting and greater resistance to AES key recovery, compared to pseudo-fully associative designs.
  • Performance evaluations of randomized caches must be comprehensive, considering multiple replacement policies and a fair benchmarking strategy (e.g., using spurious cache occupancy as a warming phase) to ensure accurate and comparable results.
  • The work presented the first successful AES key recovery attack on T-table AES using cache occupancy levels, underscoring the severity of this overlooked attack vector.

About the Speaker(s)

The primary speaker for this presentation was Anirban Chakraborty, who is a researcher at the Max Planck Institute for Security and Privacy in Germany. He presented this work in collaboration with Nimish Mishra, Shandib Sha, Shani Tachara, and Deep Makapad, indicating a team effort behind this systematic evaluation. The research reflects a focus on understanding and mitigating security vulnerabilities in modern computing architectures, particularly within the domain of cache-based side-channel attacks.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Solid systems security research that closes a real gap: randomized caches have been evaluated almost exclusively against Prime+Probe, and this work systematically demonstrates that cache occupancy is a practical, exploitable side channel against designs most practitioners assumed were hardened. The first AES key recovery via occupancy levels is the headline contribution, and it earns its place.

Heather Calloway (CISO) — PASS

Rigorous microarchitectural security research with real findings — a practical AES key recovery via cache occupancy is not nothing. But this is deep hardware security academia, and there is no viable path from these findings to a governance decision, a security program change, or an operational posture adjustment for any practitioner in my world.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)