Secure Caches for Compartmentalized Software

Kerem Arıkan

34th USENIX Security Symposium (USENIX Security '25) · Day 2 · Crypto 3: Formal Methods and Private Computation

Overview

In the pursuit of more secure and robust software, the shift from monolithic applications to compartmentalized software has been a significant architectural evolution. This talk, presented by Kerem Arıkan from Binghamton University and UC Riverside, addresses a critical vulnerability that undermines the security benefits of compartmentalization: CPU cache side-channel attacks. While compartmentalization isolates code and data into distinct memory domains, the shared nature of CPU caches often provides an unintended communication channel between these isolated compartments, allowing malicious or buggy code to infer sensitive information.

Watch on YouTube · Slides

Visual summary for Secure Caches for Compartmentalized Software by Kerem Arıkan
Visual summary for Secure Caches for Compartmentalized Software by Kerem Arıkan

Key moments

  1. 0:00 Introduction: Problem of secure compartmentalized software
  2. 2:00 CPU caches fail to enforce domain boundaries
  3. 3:30 Introducing ACC: Secure Caches for Compartments
  4. 4:00 Domain-oriented partitioning: Solving data sharing issues
  5. 6:00 Challenge: L1 cache latency with partitioning
  6. 6:18 ACC's remapping logic using TLB and DRT
  7. 7:15 Key observation: Exploiting domain access locality

Secure Caches for Compartmentalized Software

Speakers: Kerem Arıkan

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=vSTdp5ce5Cs

Overview

In the pursuit of more secure and robust software, the shift from monolithic applications to compartmentalized software has been a significant architectural evolution. This talk, presented by Kerem Arıkan from Binghamton University and UC Riverside, addresses a critical vulnerability that undermines the security benefits of compartmentalization: CPU cache side-channel attacks. While compartmentalization isolates code and data into distinct memory domains, the shared nature of CPU caches often provides an unintended communication channel between these isolated compartments, allowing malicious or buggy code to infer sensitive information.

The core problem arises because hardware caches typically do not enforce the same domain boundaries as the memory management unit (MMU), creating a blind spot in an otherwise secure architecture. Prior research has amply demonstrated the feasibility and impact of such attacks across various platforms. Arıkan and his co-authors introduce Secure Caches for Compartments (SEC), a novel cache isolation mechanism designed to integrate seamlessly with compartmentalized software. SEC aims to provide robust side-channel protection without incurring the prohibitive performance overheads associated with naive solutions like cache flushing.

This work is particularly significant because it offers a principled and performant approach to securing a fundamental hardware component that has long been a weak link in compartmentalized systems. By introducing concepts like domain-oriented partitioning, latency-aware L1 caches with Active Domain Register (ADR), and mechanisms to thwart library-based side-channel attacks, SEC represents a crucial step towards realizing truly secure and isolated software environments. The research highlights that effective security in complex systems requires holistic solutions that span both software architecture and underlying hardware design.

Background

▶ Watch: Introduction: Problem of secure compartmentalized software (0:00)

Traditionally, software development often results in monolithic applications where all code shares the same memory space and privileges, operating under an assumption of ambient authority. This model is inherently vulnerable, as any piece of code—whether primary or third-party—can access any resource, making sensitive data susceptible to leakage or corruption by malicious or buggy components. To counter this, compartmentalized software architectures have emerged, designed to split code into isolated compartments, each with strictly limited privileges and access to specific memory regions, termed domains.

These advanced compartmentalization systems often rely on hardware modifications, particularly to the Memory Management Unit (MMU). The MMU is provisioned with a permission table that dictates which compartments have specific access rights to which domains. These systems typically operate with page-level granularity, meaning every memory page belongs to a single domain. Consequently, the Translation Lookaside Buffer (TLB) must extend its entries to include domain metadata for each page. During a memory access, the address translation process returns not only the physical address but also the page's domain information, which is then subjected to a sanity check against the permission table to ensure the currently running compartment has authorized access.

Despite these robust memory-level protections, a significant vulnerability persists in the CPU's cache hierarchy. CPU caches do not inherently enforce domain boundaries, creating a critical gap in the security model. This architectural oversight opens the door to cache side-channel attacks. For instance, a victim compartment might access secret-dependent data, leaving a specific cache footprint. An attacker compartment, by strategically accessing its own data that collides with the victim's data in the shared cache space, can then infer the victim's secret by observing timing differences (e.g., cache hits vs. misses).

Prior literature has extensively demonstrated the efficacy of cache side-channel attacks in multi-component software environments. Examples include:

  1. Intel SGX enclaves: Vulnerabilities in L1 cache side channels have been shown when attacker and victim enclaves reside on hyperthreads of the same core.
  2. Browser engines: An attacker client within the same browser engine process can probe the cache occupancy footprint of a victim client to extract information.
  3. Graphics libraries: Secrets have been leaked through separate function calls between an attacker and a victim within the same process.
  4. Untrusted OS kernels: Powerful capabilities like single-stepping can be leveraged by an untrusted OS kernel to leak secrets through caches within the same process.

Cache partitioning is a well-established, principled, and deterministic method for blocking side channels by isolating mutually untrusted code portions in distinct cache regions. This ensures that an attacker's partition cannot collide with a victim's partition. However, before the work presented in this talk, there had been no concerted effort to adapt and implement side-channel partitioning specifically for compartmentalization systems, which present unique challenges related to data sharing and performance.

Key Findings

▶ Watch: Introducing ACC: Secure Caches for Compartments (3:30)

The central contribution of this work is the introduction of Secure Caches for Compartments (SEC), a novel cache isolation mechanism designed specifically to adapt partitioning to compartmentalized software. SEC tackles the challenge of cache side-channel protection from three distinct angles, extending its applicability to both intra-process and inter-process security guarantees.

First, SEC introduces the concept of domain-oriented partitioning. Unlike traditional code-oriented partitioning schemes that allocate cache regions based on compartments, SEC partitions the cache based on memory domains. This is a critical distinction, as intense data sharing between compartments (via intentionally shared domains) can lead to significant implementation and coherence issues with code-oriented approaches. Domain-oriented partitioning ensures that private data from different domains cannot collide in the cache space, while shared data, by definition, is not subject to side-channel protection in this context, simplifying cache coherence maintenance and providing a more elegant solution.

Second, SEC addresses the performance challenge of integrating partitioning logic, especially for latency-sensitive caches like L1. Directly adding domain-oriented remapping logic to L1 caches can introduce unacceptable access latency. The researchers observed a phenomenon called domain access locality, where compartments tend to access the same domains for extended periods (over 98% of consecutive cache accesses are to the same domain). To exploit this, SEC introduces the Active Domain Register (ADR). ADR stores the remapping entry of the most recently accessed domain, allowing simultaneous access to the TLB and cache. In the event of an ADR hit, the cache access latency matches that of a baseline, unpartitioned cache, effectively mitigating overhead for the vast majority of accesses.

Third, SEC introduces mechanisms to stop library-based side-channel attacks, which can bypass simple domain isolation. Even with domain-oriented partitioning, if both an attacker and a victim call an insecure function in a shared library, they might use the same cache space corresponding to the library's domain partition. This allows the attacker to infer secret-dependent behavior. To counter this, SEC implements horizontal partitions. These partitions preserve a per-caller library state within a single domain's partition, preventing cache collisions between different callers of the same shared function and thereby disallowing attackers from measuring cache timing for inference.

Finally, SEC extends its security guarantees to multiprocess environments. For shared last-level caches (LLCs) in multi-core CPUs, SEC allocates nested partitions. These involve conventional Process Level Partitions (PLPs), within which domain partitions reside. A Process Remapping Table (PRT) is maintained at the LLC level to manage PLP remapping for each core. This comprehensive approach eliminates both intra-process and inter-process side-channel vulnerabilities, providing end-to-end protection in complex system architectures.

Technical Deep Dive

▶ Watch: Domain-oriented partitioning: Solving data sharing issues (4:00)

The core innovation of SEC lies in its departure from conventional cache partitioning strategies and its astute optimizations for performance.

Domain-Oriented Partitioning:

Traditional cache partitioning often focuses on isolating different "codes" or "compartments." However, in compartmentalized software, data sharing is intentional and often necessary. Consider a scenario where compartment zero and compartment one share a cache line in memory. If the cache is partitioned based on compartments (code-oriented), compartment zero might write to this shared line in its assigned partition. Later, compartment one attempts to read the same shared line. If this results in a cache miss, compartment one retrieves an outdated version from memory or a lower-level cache, leading to aliasing coherence issues and significant implementation complexity for maintaining consistency.

SEC resolves this by introducing domain-oriented partitioning. Here, cache partitions are based on the domains to which memory pages belong, rather than the compartments accessing them. For example, if data X and Y belong to private domains, their respective domain partitions in the cache ensure they cannot collide. If data Z belongs to a shared domain (Domain 1), then all compartments with permission to access Domain 1 will use the same partition for data Z. This design inherently handles shared data, as shared data is explicitly not subject to side-channel protection in this model. When compartment zero writes to a shared line, it does so within the shared line's corresponding domain partition. When compartment one subsequently reads it, it accesses the same domain partition, retrieving the correct, modified version. This approach simplifies partitioning implementation and elegantly maintains cache coherence while establishing side-channel protection for private data.

Latency-Aware L1 Caches with ADR:

Implementing domain-oriented remapping logic directly into private caches like L1 is challenging due to the strict latency requirements. An unoptimized approach would significantly increase cache access latency. A baseline cache access typically takes two cycles. In an unoptimized SEC scenario, the TLB must first be accessed to retrieve the memory page's domain metadata. Then, a Domain Remapping Table (DRT) is accessed to retrieve the remapping information for that domain. Only then can the cache perform a set access and tag matching. This sequence results in a four-cycle access, doubling the latency.

SEC mitigates this overhead by exploiting domain access locality. The researchers found that, on average, over 98% of consecutive cache accesses are to the same domain. This means that a compartment often operates within a specific domain or a small set of domains for extended periods. To leverage this, SEC introduces the Active Domain Register (ADR). The ADR stores the DRT entry of the most recently accessed domain. During a cache access, SEC can simultaneously access both the cache and the TLB. If the ADR correctly predicts the domain (an ADR hit), the remapping information is immediately available, and the cache can proceed with tag matching as usual. This allows SEC to achieve the same baseline cache access latency (two cycles) in the case of an ADR hit. In the less frequent case of an ADR miss, SEC incurs the same four-cycle latency as the unoptimized approach, but the high hit rate of ADR ensures minimal average performance degradation.

Mitigating Library-Based Side-Channel Attacks with Horizontal Partitions:

Even with domain-oriented partitioning, a subtle class of attacks, termed library-based side-channel attacks, can persist. Imagine a shared library containing an insecure function. Both a victim and an attacker compartment might call this function. The function's instructions are loaded into the library's domain partition in the cache. If the victim's call leaves a secret-dependent residue in the cache, the attacker's subsequent call to the same function (using the same cache space for that shared library's domain) can infer the victim's secret by observing timing differences. The domain partitioning protects data across domains but not state within a shared library domain when accessed by different callers.

To counter this, SEC introduces horizontal partitions. These partitions operate within a single domain's partition (e.g., the shared library's domain partition) to preserve a per-caller library state. When the victim calls the shared function, its instructions and data are loaded into its own horizontal partition within the library's domain partition. When the attacker attempts the same inference technique, its call uses a different horizontal partition, preventing collision with the victim's instruction residue. This effectively disallows the attacker from measuring cache timing related to the victim's execution, thereby mitigating library-based side-channel attacks.

Multiprocess Environments:

For multi-core CPUs with shared Last Level Caches (LLCs), SEC extends its model to handle inter-process side-channel vulnerabilities. SEC employs nested partitions. The LLC is first partitioned into Process Level Partitions (PLPs), which are conventional partitions isolating different processes. Within each PLP, domain partitions are then allocated, mirroring the intra-process domain-oriented partitioning. A Process Remapping Table (PRT) is maintained at the LLC level, containing remapping information for PLPs specific to each core. This hierarchical partitioning ensures that both intra-process (compartment-to-compartment) and inter-process (process-to-process) side channels are eliminated, providing comprehensive security.

Demo / Proof of Concept

▶ Watch: ACC's remapping logic using TLB and DRT (6:18)

While the talk did not feature a live demonstration of SEC in action, the researchers thoroughly evaluated SEC through an extensive methodology to quantify its security guarantees and performance overheads. Their implementation and evaluation serve as a robust proof of concept for the proposed architecture.

The SEC architecture was implemented in a cycle-accurate Gem5 simulator, a widely used platform for computer architecture research. To generate diverse workloads for compartmentalized environments, the researchers utilized the SOAP compartmentalization tool on a set of industry-standard benchmarks, specifically MIBench and SPEC17. The evaluation process involved two main phases: a permission table generation run to define the domain structure for each benchmark, followed by experiment runs to collect actual performance and overhead numbers.

For comparative analysis, SEC's performance was measured against two baseline cache flushing configurations, which represent common, albeit inefficient, security measures:

  1. Flush-On-All-Compartment-Switches (FOAC): This configuration flushes the entire cache hierarchy (L1, L2, LLC) whenever the program crosses a compartment boundary.
  2. Flush-L1-On-Compartment-Switches (FL1 OSC): In this variant, only the L1 cache is flushed on compartment switches, while SEC is implemented on other cache levels.

These flushing mechanisms were used to highlight the significant performance benefits of SEC's architectural approach over brute-force security measures.

Defensive Implications

▶ Watch: Key observation: Exploiting domain access locality (7:15)

The findings presented by Kerem Arıkan have profound implications for hardware designers, software architects, and security practitioners aiming to build truly secure compartmentalized systems.

  1. Hardware Adoption of SEC-like Architectures: The most direct implication is the need for future CPU designs to incorporate hardware mechanisms akin to SEC. Specifically, domain-oriented partitioning should be considered for all cache levels, and Active Domain Registers (ADRs) are essential for L1 caches to maintain performance. The Domain Remapping Table (DRT) and Process Remapping Table (PRT) are also critical components that need hardware support. This represents a paradigm shift from current cache designs that do not inherently understand or enforce software-defined domains.
  1. Avoid Naive Cache Flushing: The evaluation clearly demonstrates that cache flushing, while effective at mitigating side channels, is an infeasible security solution for compartmentalized software. FOAC incurred an average performance degradation of 60%, and even FL1 OSC resulted in a 32% loss. In some benchmarks with high compartment switch frequency, performance plummeted to below 30% of the baseline, or even near 0%. This data strongly advises against using cache flushing as a primary side-channel defense in performance-sensitive applications.
  1. Prioritize Domain-Oriented Partitioning: Software architects designing compartmentalized systems should advocate for hardware that supports domain-oriented partitioning. The coherence issues and implementation complexity associated with code-oriented partitioning (when data is shared) make it an inferior choice for these architectures. The simplicity and robustness of domain-oriented partitioning for managing shared data are key advantages.
  1. Guard Against Library-Based Side Channels: Even with robust domain partitioning, the threat of library-based side-channel attacks persists. Developers of shared libraries, particularly those handling sensitive data or operations, must be aware that their code, when executed by different compartments, can still create a side channel. Hardware support for horizontal partitions (per-caller library state within a domain partition) is crucial to fully mitigate this sophisticated attack vector.
  1. Comprehensive Multi-Level Cache Protection: Security solutions must extend beyond private L1 caches to include shared L2 and Last Level Caches (LLCs). SEC's approach of using nested partitions with Process Level Partitions (PLPs) and a Process Remapping Table (PRT) for LLCs provides a blueprint for comprehensive, end-to-end side-channel protection across the entire cache hierarchy in multi-core and multiprocess environments.
  1. Performance vs. Security Trade-offs: SEC demonstrates that it is possible to achieve strong side-channel security with acceptable performance and hardware overheads. With an average performance loss of only 7% across benchmarks and a modest 63% hardware area overhead (estimated using the MCPAT tool), SEC offers a practical and viable path forward, outperforming other allocation techniques like proportional and static configurations which incurred over 25% performance loss. This balance makes SEC a compelling model for future secure system designs.

Key Takeaways

  • CPU caches are a significant side-channel vulnerability in otherwise compartmentalized software architectures, undermining isolation guarantees.
  • SEC's domain-oriented partitioning is a novel and crucial approach that partitions caches based on memory domains rather than compartments, elegantly solving data coherence issues inherent in shared-data scenarios.
  • The Active Domain Register (ADR) effectively mitigates performance overhead for latency-sensitive L1 caches by exploiting domain access locality (over 98% consecutive domain accesses), achieving baseline access latency for most operations.
  • Horizontal partitions are essential to prevent sophisticated library-based side-channel attacks by maintaining a per-caller library state within shared domain partitions.
  • Cache flushing is an infeasible method for securing compartmentalized software due to severe performance degradation, ranging from 32% to 60% on average.
  • SEC offers a practical and performant hardware solution, incurring only an average 7% performance loss and a 63% hardware area overhead, making it a viable design for future secure systems.

About the Speaker(s)

Kerem Arıkan is the presenter of this talk on "Secure Caches for Compartmentalized Software." The research presented is a collaborative effort involving authors from Binghamton University and UC Riverside, indicating his affiliation with one or both of these academic institutions at the time of the presentation. His work focuses on addressing fundamental hardware security challenges in advanced software architectures.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Solid systems security research from USENIX that addresses a real, underserved problem: cache side-channels punching through compartmentalization boundaries that the MMU correctly enforces. The domain-oriented partitioning insight is genuinely clever — solving the coherence mess that code-oriented approaches create when compartments share data — and the ADR optimization shows the authors actually thought through hardware implementation constraints rather than just proposing something that would never ship.

Heather Calloway (CISO) — PASS

Rigorous academic hardware security research with no meaningful bridge to governance, operations, or enterprise security leadership. This is CPU architecture work — valid and probably important to the right audience, but that audience is not mine.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)