SoK: So, You Think You Know All About Secure Randomized Caches?
Anubhav Bhatla
34th USENIX Security Symposium (USENIX Security '25) · Day 2 · Hardware Security 1: Microarchitectures
Overview
In this USENIX Security talk, Anubhav Bhatla presents a comprehensive "Systemization of Knowledge" (SoK) study on secure randomized caches. The talk delves into the intricate design space of cache architectures engineered to thwart Last Level Cache (LLC) side-channel attacks. Bhatla systematically dissects various security-enhancing features, termed "knobs," and rigorously evaluates their effectiveness, both individually and in combination, against prevalent attack methodologies.

Key moments
- 0:00 Introduction to secure randomized caches and LLC attacks
- 2:00 Overview of existing secure randomized cache designs
- 3:00 Systematization methodology and key security knobs
- 3:40 Five key security knobs for randomized caches
- 4:00 Metrics for quantifying cache security
- 4:50 Impact of skewing on eviction rate and security
- 5:50 Effectiveness of extra valid tags and global eviction
SoK: So, You Think You Know All About Secure Randomized Caches?
Speakers: Anubhav Bhatla
Conference: USENIX Security
YouTube: https://www.youtube.com/watch?v=CMPeM5H332Q
Overview
In this USENIX Security talk, Anubhav Bhatla presents a comprehensive "Systemization of Knowledge" (SoK) study on secure randomized caches. The talk delves into the intricate design space of cache architectures engineered to thwart Last Level Cache (LLC) side-channel attacks. Bhatla systematically dissects various security-enhancing features, termed "knobs," and rigorously evaluates their effectiveness, both individually and in combination, against prevalent attack methodologies.
The core motivation behind this research stems from the persistent threat posed by LLC attacks, which exploit shared cache resources to leak sensitive information from co-located processes. While numerous randomized cache designs have been proposed to mitigate these vulnerabilities, a holistic understanding of their security properties, interdependencies, and practical implications has been lacking. This SoK fills that gap by providing a foundational framework for analyzing, designing, and comparing secure randomized caches, offering crucial insights for both researchers and practitioners in developing robust cache defenses.
This talk is particularly significant because it moves beyond evaluating individual defense mechanisms, instead focusing on the interplay of multiple design choices. By quantifying security benefits using novel metrics and exploring the impact of often-overlooked factors like cache warm-up states, Bhatla provides a nuanced perspective on what truly constitutes a secure cache design. The findings challenge conventional wisdom, revealing that security is not merely additive but highly dependent on the synergistic combination of various architectural features.
Background
▶ Watch: Introduction to secure randomized caches and LLC attacks (0:00)
Modern CPU architectures rely heavily on a multi-level cache hierarchy to bridge the significant speed gap between the high-speed CPU and the slower main memory (RAM). Typically, processors feature Level 1 (L1) and Level 2 (L2) caches, which are private to individual CPU cores, providing rapid access to frequently used data for a single process. However, the Last Level Cache (LLC), often referred to as the L3 cache, is typically shared among all CPU cores and, consequently, all processes running on the system. This shared nature of the LLC, while beneficial for performance, introduces a critical security vulnerability: side-channel attacks.
LLC side-channel attacks exploit the observable timing differences that arise from cache hits versus cache misses. By carefully monitoring these timing variations, an attacker (spy process) can infer information about the memory access patterns of a victim process running concurrently on the same system. The talk categorizes these attacks into two primary classes:
- Occupancy-Based Attacks: In this scenario, a spy process repeatedly accesses a large buffer designed to occupy a significant portion of the LLC. As the victim process executes its sensitive operations, it may evict some of the spy's data from the LLC. By measuring the number of accesses the spy can make to its buffer within a given time, it can infer the victim's LLC occupancy. A lower number of accesses for the spy suggests higher victim occupancy, leaking information. A classic example of this is website fingerprinting, where an attacker can determine which website a victim is browsing by observing the distinctive cache footprints left by different web pages.
- Conflict-Based Attacks: These attacks, exemplified by techniques like Prime+Probe, are more granular. The attacker first "primes" a specific set in the LLC by filling it with their own data (an eviction set). Then, they wait for the victim to run, hoping the victim accesses data that maps to the same cache set, thereby "probing" it. If the victim's access evicts the spy's data, the spy will experience a cache miss when it subsequently tries to access its own data in that set. The measurable time difference between a hit and a miss allows the spy to infer that the victim accessed an entry within that specific cache set, leaking fine-grained information about the victim's operations.
To counteract these potent attacks, researchers have proposed various secure randomized cache designs. Early attempts, such as Caesar, introduced randomization using a block cipher to obscure the address-to-set mapping, with periodic rekeying to prevent reverse engineering. However, Caesar was quickly broken. Subsequent designs like Caesar-S and Scatter Cache built upon this by incorporating Skewed Associative Caches (SKAs) on top of randomization, making the address-to-set mapping even more complex. More sophisticated designs, such as Mirage and Maya, decoupled the tag and data stores, used pointer-based lookups, and implemented global random eviction policies, often adding extra tags to enhance security and prevent set-based evictions. Despite these advancements, a unified methodology to systematically evaluate and compare their security properties remained elusive, making it difficult to understand which features truly contribute to robust defense and under what conditions. This SoK aims to provide that much-needed systematic analysis.
Key Findings
▶ Watch: Systematization methodology and key security knobs (3:00)
The talk presents a systematic framework for understanding and evaluating secure randomized cache designs, identifying five fundamental "knobs" that influence their security. These knobs, which are design choices that can be adjusted, are skewing, associativity, extra valid tags, remapping, and the replacement policy. The study assumes that randomization, typically implemented using a block cipher, is a default baseline for all secure cache designs.
To quantify security, the research introduces two primary metrics:
- Eviction Rate: This metric measures the probability or rate at which a target address is evicted from the cache when an attacker uses an eviction set. A lower eviction rate for a given eviction set size indicates better security, as it becomes harder for an attacker to reliably evict the victim's data. This metric is considered "algorithmic independent" as it assumes the eviction set is already provided.
- Number of LLC Evictions Needed to Create a Fixed-Size Eviction Set: This metric quantifies the effort required for an attacker to construct an effective eviction set using specific algorithms, such as Conflict Testing or Fragment Probe. A higher number of LLC evictions needed implies greater security, as it increases the time and resources an attacker must expend to launch a conflict-based attack.
A crucial finding is that the effectiveness of individual security knobs is often highly contingent on their combination with other knobs. For instance, extra valid tags provide little to no security benefit in isolation but become highly effective when coupled with load-aware skew selection and a global eviction policy. Similarly, increasing associativity consistently improves security across various configurations.
Furthermore, the study highlights the significant, yet often overlooked, impact of the cache warm-up state on security. It demonstrates that the security benefits of certain design choices, such as load-aware skewing, can diminish significantly as the cache approaches a fully saturated (100% warm) state. This finding suggests that security evaluations must account for realistic cache operating conditions beyond an initially empty cache.
Finally, the research also considers the practical implications of these security knobs, evaluating their trade-offs in terms of design complexity, performance overheads, logic requirements, and power consumption, providing a holistic view for cache designers.
Technical Deep Dive
▶ Watch: Five key security knobs for randomized caches (3:40)
The core of the systemization study lies in the detailed analysis of five key security knobs, along with their various sub-knobs and interactions. Randomization, typically implemented using a block cipher, is the assumed foundation for all designs.
1. Skewing
Skewing involves mapping addresses to cache sets using multiple hash functions or "skews." The talk explores two sub-knobs:
- Random Skew Selection: The system randomly chooses one of the available skews for each memory access.
- Load-Aware Skew Selection (LA): The system dynamically selects the skew that results in the least loaded cache set, aiming to distribute data more evenly and reduce conflicts.
The results clearly indicate that increasing the number of skews significantly improves security. For instance, moving from a 2-skew design to a 16-skew design drastically reduces the eviction rate. Furthermore, incorporating Load-Aware Skew Selection on top of random skew selection provides an additional security boost, further lowering the eviction rate. This means that intelligently distributing data across sets based on current load is more effective than simple random distribution.
2. Extra Valid Tags
This knob focuses on enhancing tag management. It has two sub-knobs:
- Eviction Strategy: Whether an eviction occurs locally (within the specific cache set) or globally (from anywhere in the entire cache).
- Decoupling Tag and Data Store: As seen in designs like Mirage and Maya, where tags and data are stored separately.
A surprising finding is that decoupling the tag and data store has no observable security impact on the cache design. However, the eviction strategy and the presence of extra valid tags show complex interactions. Initially, with configurations using only skews and invalid tags, no security benefit was observed from invalid tags. Even with global eviction, invalid tags offered zero security. It was only when skews with load-aware selection were combined with invalid tags and global eviction that a significant security benefit emerged. In this specific configuration, increasing the number of invalid tags drastically reduced the eviction rate, demonstrating that extra valid tags are effective only when coupled with appropriate other knobs. This highlights the non-additive nature of security features.
3. Associativity (High Sensitivity)
Associativity refers to the number of ways (lines) within a cache set where a particular memory block can reside. Typical caches have associativities around 16. The study investigates the impact of increasing associativity up to 128.
The results show a consistent and drastic drop in the eviction rate as associativity increases from 16 to 128. Higher associativity provides substantial security benefits, making it much harder for an attacker to construct effective eviction sets. Notably, even with a smaller number of skews (e.g., two SKAs), high associativity (e.g., 128) can achieve security levels comparable to or even better than configurations with a large number of skews but lower associativity. This suggests that increasing associativity is a powerful, albeit resource-intensive, defense mechanism.
4. Replacement Policy
The replacement policy dictates which cache line is evicted when a new block needs to be brought into a full set. The talk examines four policies: random, and three deterministic policies including LRU (Least Recently Used) and RPLU (Randomized Pseudo LRU). gRPL (global Randomized Pseudo LRU) is highlighted as a practical, randomized pseudo-LRU policy specifically designed for randomized caches.
For conflict-based attacks, deterministic policies generally perform better than a purely random replacement policy in terms of eviction rate. Among deterministic options, gRPL performs similarly to a global LRU, but is noted as being more practical to implement on a global scale. This makes gRPL a strong candidate for secure randomized caches.
However, the picture changes for occupancy-based attacks. Here, a random replacement policy is found to be more secure than deterministic policies like LRU. Mirage's use of a global eviction policy showed similar performance to other designs using local eviction, indicating no strong trend for occupancy attacks regarding eviction policy. This illustrates a trade-off: a policy optimal for one type of attack might be suboptimal for another.
5. Remapping
Remapping involves changing the address-to-set mapping, typically by rekeying the underlying block cipher. This knob is quantified by the number of LLC evictions needed before remapping is required. A higher number of evictions implies a longer remapping period, which reduces the performance implications of frequent remapping.
The study evaluates remapping against algorithms like Conflict Testing and Fragment Probe. Conflict Testing is found to be significantly faster (over 10 times) at finding eviction sets than Fragment Probe. Improvements in this metric directly translate to better security. Increasing the number of skews (from 2 to 16) leads to a great improvement in the number of LLC evictions required. Combining skews with load-aware selection, invalid tags, and global eviction further enhances this. Finally, high associativity designs (64 and 128) yield large improvements. When all these beneficial knobs are combined (high associativity with load-aware selection, invalid tags, and global eviction), a substantial improvement in the number of LLC evictions needed is observed, making remapping less frequent and thus less impactful on performance.
Cache Warm-up State
A critical "hidden nuance" explored is the impact of the cache warm-up state. Previous evaluations often assume an empty cache that slowly fills up. The study introduces a more rigorous metric where the cache warm-up state is explicitly reset at different levels (e.g., 0%, 25%, 50%, 75%, 100% full) for each iteration. While this method is 10 times slower, it provides a more accurate picture of security under realistic operating conditions.
For a design with skews and load-aware selection, the eviction rate was zero up to a 75% warm state, meaning eviction sets were completely useless. However, at a 100% warm state (cache completely full), the security benefit of load-aware selection diminished, and the results became identical to designs without load-aware selection. A similar trend was observed when invalid tags and global eviction were added: high warm-up states reduced the security benefits. This clearly demonstrates that the cache warm-up state has a significant impact on the security of the design, and evaluations must consider this factor.
Demo / Proof of Concept
▶ Watch: Impact of skewing on eviction rate and security (4:50)
As a Systemization of Knowledge (SoK) paper, this talk did not feature a live demonstration of a specific attack or a new defense mechanism. Instead, its focus was on a comprehensive analysis and evaluation framework for existing and proposed secure randomized cache designs. The research provides a robust methodology for understanding the security properties of various cache architectures. The speaker proudly mentioned that their artifact received all three badges (available, functional, reusable) and, notably, a distinguished artifact award, indicating the high quality and availability of their codebase and evaluation tools for others to replicate and extend their findings.
Defensive Implications
▶ Watch: Effectiveness of extra valid tags and global eviction (5:50)
The detailed analysis presented in this talk offers several critical insights and actionable recommendations for cache designers aiming to build more secure systems:
- Combinatorial Security is Key: The most significant takeaway is that security in randomized caches is not achieved by simply adding individual features. Instead, it arises from the synergistic combination of multiple knobs. Designers must move beyond evaluating individual defenses and focus on how different architectural choices interact to provide robust protection. For instance, extra valid tags only become effective when paired with load-aware skew selection and global eviction.
- Prioritize High Associativity and Load-Aware Skewing: The study consistently shows that high associativity (e.g., 128 ways) drastically reduces eviction rates and significantly improves security against conflict attacks. Similarly, load-aware skew selection is a powerful mechanism for distributing data evenly and mitigating conflicts. These two features should be prioritized in secure cache designs, even if they come with increased hardware complexity or power consumption.
- Strategic Use of Global Eviction and Extra Valid Tags: While global eviction and extra valid tags don't universally improve security, their effectiveness is undeniable when coupled with the right set of other knobs. Designers should consider these features, especially in conjunction with load-aware skewing, to maximize their defensive potential against conflict-based attacks.
- Choose Replacement Policies Judiciously: The optimal replacement policy depends on the target attack. For conflict-based attacks, deterministic policies like gRPL (global Randomized Pseudo LRU) are generally more effective and practical than purely random policies. However, for occupancy-based attacks, a random policy might offer better protection. Designers need to weigh the primary threat model when selecting a replacement policy.
- Account for Cache Warm-up State in Evaluations: A crucial, often overlooked, aspect is the cache's warm-up state. Security evaluations that assume an empty cache can be misleading. Designers and researchers should adopt methodologies that test security under various, more realistic, cache fill levels, including fully saturated caches, as the benefits of some defenses (like load-aware skewing) can diminish at high warm-up states.
- Consider Partitioning-Based Designs for Occupancy Attacks: While the talk focuses on randomized caches, it briefly compares them to partitioning-based designs like SAS cache and way-based partitioning. These designs show superior performance against occupancy-based attacks, with way-based partitioning being completely secure in some cases. For environments where occupancy attacks are a primary concern, a partitioning approach might offer stronger guarantees than even the best randomized designs.
- Understand Performance/Security Trade-offs: Implementing these security knobs inevitably introduces trade-offs in terms of design complexity, performance overhead (latency), logic requirements, and power consumption. Designers must carefully analyze these trade-offs based on the specific application and threat model to strike an optimal balance between security and practicality.
Key Takeaways
- Systematic Analysis is Crucial: The talk provides a systematic framework, identifying five key "knobs" (skewing, associativity, extra valid tags, remapping, replacement policy) to analyze and design secure randomized caches against LLC side-channel attacks.
- Combinatorial Security: The effectiveness of individual security features is highly dependent on their combination with other knobs. Security is not merely additive; specific interactions are critical for robust defense.
- High Associativity & Load-Aware Skewing are Potent: Increasing cache associativity (e.g., to 128 ways) and implementing load-aware skew selection are consistently effective strategies for significantly reducing eviction rates and improving security against conflict-based attacks.
- Cache Warm-up State Matters: The cache's warm-up state (how full it is) significantly impacts the efficacy of security designs. Security benefits of some features, like load-aware skewing, can diminish as the cache approaches full saturation.
- Replacement Policy Trade-offs: Deterministic replacement policies (like gRPL) generally perform better against conflict-based attacks, while random policies offer stronger defense against occupancy-based attacks. gRPL provides a practical and effective option for global policies.
- Partitioning for Occupancy: For strong protection against occupancy-based attacks, partitioning-based designs (e.g., way-based partitioning) can outperform randomized caches, sometimes offering complete security.
About the Speaker(s)
Anubhav Bhatla is the sole speaker for this presentation. Based on the transcript, his affiliation or specific title is not provided, but his work presented at USENIX Security indicates a strong background in computer architecture and security research, particularly concerning cache side-channels and their mitigation. His detailed systemization study demonstrates expertise in rigorous evaluation methodologies for secure hardware designs.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Solid SoK that does exactly what a good SoK should do: imposes structure on a fragmented design space and surfaces non-obvious results that change how you'd reason about building or evaluating a secure cache. The warm-up state finding alone — showing that load-aware skewing's security benefit collapses at 100% cache saturation — is the kind of thing that quietly invalidates prior work and deserves a wider audience. Distinguished artifact award with all three badges is a meaningful signal that the methodology is actually reproducible.
Heather Calloway (CISO) — PASS
Rigorous computer architecture research with a distinguished artifact award — but this is outside my lane entirely. LLC side-channel analysis at the microarchitectural level is not governance, not operational security, and not a decision I'm handing to a CISO or a board.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)