BULKHEAD: Secure, Scalable, and Efficient Kernel Compartmentalization with PKS

Yinggang Guo

Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · System-Level Security

Overview

The talk "BULKHEAD: Secure, Scalable, and Efficient Kernel Compartmentalization with PKS" by Yinggang Guo at the NDSS Symposium addresses the persistent and growing security vulnerabilities within monolithic operating system kernels, particularly the Linux kernel. The presentation introduces BULKHEAD, a novel solution designed to enhance kernel security through robust compartmentalization. This system leverages Intel's recent hardware feature, Protection Keys for Supervisor Mode (PKS), to isolate kernel modules into mutually untrusted compartments, thereby significantly confining the impact of potential exploits.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction: Kernel vulnerabilities and need for compartmentalization
  2. 1:10 BULKHEAD solution overview using Intel PKS
  3. 2:50 Challenges and weaknesses of current kernel isolation methods
  4. 4:20 BULKHEAD's core security mechanisms and isolation enforcement
  5. 5:50 Detailed explanation of secure compartment interface (SGT)
  6. 7:00 Solving scalability: Two-level compartmentalization scheme
  7. 8:20 Performance optimizations: hardware-based switching and zero-copy

BULKHEAD: Secure, Scalable, and Efficient Kernel Compartmentalization with PKS

Speakers: Yinggang Guo

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=q7M12aeeuMo

Overview

The talk "BULKHEAD: Secure, Scalable, and Efficient Kernel Compartmentalization with PKS" by Yinggang Guo at the NDSS Symposium addresses the persistent and growing security vulnerabilities within monolithic operating system kernels, particularly the Linux kernel. The presentation introduces BULKHEAD, a novel solution designed to enhance kernel security through robust compartmentalization. This system leverages Intel's recent hardware feature, Protection Keys for Supervisor Mode (PKS), to isolate kernel modules into mutually untrusted compartments, thereby significantly confining the impact of potential exploits.

The core problem BULKHEAD tackles is the inherent insecurity of monolithic kernels, where a single vulnerability in any component can compromise the entire system due to shared privileges. By implementing fine-grained isolation, BULKHEAD aims to reduce the attack surface and limit the "blast radius" of successful exploits. The project distinguishes itself by achieving strong security guarantees without sacrificing performance or compatibility, addressing the critical challenges that have hindered widespread adoption of kernel compartmentalization in the past.

This research is particularly relevant in an era where kernel vulnerabilities continue to proliferate at an alarming rate, with the Linux kernel alone reporting approximately 2,000 CVEs in 2024. BULKHEAD offers a principled mitigation strategy, demonstrating how modern hardware features can be effectively harnessed to build more resilient and secure operating system kernels, ultimately protecting critical system software from escalating threats.

Background

▶ Watch: Introduction: Kernel vulnerabilities and need for compartmentalization (0:00)

The operating system kernel, as the foundation of system software, is a prime target for attackers due to its privileged position and extensive access to system resources. Unfortunately, kernels, especially monolithic ones like Linux, are plagued by a continuous stream of vulnerabilities. According to the speaker's analysis, the Linux kernel saw around 2,000 CVEs reported in 2024, with issues distributed widely across subsystems, including core kernel components and network-related modules like IPv6 and netfilters. The fundamental problem lies in the monolithic architecture: for performance and compatibility reasons, all kernel components share the highest privilege level, often granting them access to data and code far beyond their intended duties. This means that exploiting a single vulnerability in any part of the kernel can lead to a complete system compromise.

Prior attempts at kernel compartmentalization have faced significant hurdles across four key objectives: security, scalability, performance, and compatibility.

  • Microkernels, while offering strong isolation by moving most kernel components into isolated user processes, suffer from heavy inter-process communication (IPC) overhead, leading to poor performance. Their complete redesign of the OS also hinders general application.
  • Software Fault Isolation (SFI) approaches establish isolated domains by instrumenting every memory access instruction with security checks during compile time. This results in substantial performance overhead, making them impractical for high-performance kernels.
  • Virtualization-based approaches, utilizing hypervisors and Extended Page Tables (EPTs), protect execution domains and create different memory views. However, they introduce an additional layer, complicating the Trusted Computing Base (TCB), incur extra overhead, and come with nested virtualization restrictions.
  • Other efforts leveraging various hardware features have been limited by their inability to scale to multiple compartments, often due to hardware-imposed constraints on the number of isolation domains.

BULKHEAD addresses these limitations by harnessing Intel's Protection Keys for Supervisor Mode (PKS). PKS is a hardware feature that tags memory pages with a 4-bit protection key (P-key) within their page table entries (PTEs). This mechanism allows the address space to be partitioned into at most 16 distinct memory domains. The permissions for pages associated with a specific P-key are stored in a dedicated per-thread register called PKRS, using a two-bit notation: WD (Write Disable) and AD (Access Disable). By leveraging PKS, BULKHEAD aims to achieve fine-grained, hardware-enforced isolation within the kernel, excluding the core kernel from the TCB and tagging it, along with other Linux kernel modules (LKMs), with distinct P-keys. A newly introduced lightweight in-kernel monitor is tasked with enforcing the security environment, including data integrity, execute-only memory (XOM), and compartment interface integrity via a specially designed Switch Gate Table (SGT), ensuring secure and efficient compartment switching.

Key Findings

▶ Watch: Challenges and weaknesses of current kernel isolation methods (2:50)

BULKHEAD presents a robust and innovative solution for kernel compartmentalization, addressing long-standing challenges in security, scalability, and performance. Its key findings and contributions include:

  • PKS-based Bidirectional Isolation: BULKHEAD effectively utilizes Intel's PKS feature to enforce bidirectional isolation directly within the kernel space. This design breaks the traditional monopoly of the core kernel without introducing additional layers (like hypervisors) or suffering from performance penalties associated with software-based isolation.
  • In-Kernel Monitor for Privilege Separation: A lightweight in-kernel monitor is constructed through privilege separation at the same kernel level. This monitor is exclusively responsible for managing access control metadata, specifically the write-protected page tables and the Switch Gate Table (SGT), ensuring their integrity against malicious compartments.
  • Comprehensive Environment Enforcement: The monitor enforces a series of critical security environments, including data integrity (private heaps, single ownership), control-flow protection through eXecute-Only Memory (XOM) for all kernel code, and compartment interface integrity to prevent confused deputy attacks.
  • Scalable Two-Level Compartmentalization: Recognizing the 16-domain limitation of PKS, BULKHEAD introduces a novel two-level compartmentalization scheme. The first level uses PKS for intra-address space isolation, while the second level employs locality-aware address space switching with ASID (Address Space ID). This innovative approach allows BULKHEAD to scale far beyond 16 compartments, supporting thousands of LKMs by grouping modules with high interaction locality into shared address spaces.
  • Efficient Hardware-Based Compartment Switching: The system achieves high performance through hardware-based compartment switches, which primarily involve specific register updates. Additionally, it supports zero-copy data transfer between compartments by using P-keys to enforce single ownership and validating shared data through a dedicated monitor handler during page faults, avoiding costly data copying.
  • Proven Security and Low Overhead: Comprehensive security analysis demonstrated that BULKHEAD successfully detects all malicious attempts related to data-oriented and control-flow hijacking attacks, either through PKS violations or monitor checks. Performance evaluations showed an average overhead of just 2.44% for real-world applications with 160 compartmentized IO modules. Specific tests like Apache bench on IPv6 showed an overhead of about 2%, significantly outperforming previous software-based isolation techniques. Memory overhead was negligible, at 1.66% for LM bench and 0.63% for foriconics.

Technical Deep Dive

▶ Watch: BULKHEAD's core security mechanisms and isolation enforcement (4:20)

BULKHEAD's technical architecture is meticulously designed to leverage PKS for robust kernel compartmentalization while overcoming its inherent limitations and ensuring high performance and security. The solution focuses on bidirectional isolation, resource control, and secure inter-compartment communication.

At its core, BULKHEAD employs PKS-based bidirectional isolation within the kernel space. This means that not only are individual kernel modules (LKMs) isolated from each other, but the core kernel itself is also isolated and protected, breaking its traditional monopoly on system privileges. Each compartment is assigned a distinct P-key, which is embedded in the page table entries (PTEs) of its memory pages. The permissions associated with these P-keys are dynamically controlled via the per-thread PKRS register.

To manage this isolation securely, an in-kernel monitor is introduced. This monitor operates at the same privilege level as other kernel components but is logically separated and highly protected. Its critical role is to enforce access control over sensitive metadata. Specifically, the page tables and the Switch Gate Table (SGT) (discussed below) are write-protected, and only the monitor has the privilege to update these structures, ensuring that no compromised compartment can alter the system's access control policies.

Beyond memory resources, BULKHEAD also controls instruction and register resources. It eliminates privileged instructions within compartments (other than the monitor) through binary rewriting. This process involves transforming the compiled kernel binaries to remove or modify instructions that could be misused. Examples include Nop insertion to pad code, register reassignment to prevent critical register manipulation, and data adjustment to ensure data integrity.

The monitor enforces a comprehensive set of security environments:

  1. Data Integrity: Each compartment's data can only be modified by the compartment itself. This is achieved by tagging memory pages with specific P-keys, effectively creating private heaps and memory pools for each compartment. This prevents unauthorized access or modification of data across compartment boundaries.
  2. Page Table Integrity: As mentioned, only the monitor can modify the page tables, ensuring that the memory mappings and protection keys remain consistent and secure.
  3. Control Flow Protection: BULKHEAD implements eXecute-Only Memory (XOM) for all kernel code. This means that kernel code pages cannot be read or written by any entity, significantly mitigating code injection and code reuse attacks (e.g., Return-Oriented Programming, ROP).

A critical component for secure inter-compartment communication is the Compartment Interface Integrity enforced by the Switch Gate Table (SGT). This mechanism is designed to protect against confused deputy attacks, where a legitimate, but compromised, compartment might be tricked into performing actions on behalf of an attacker. Compartment switches are strictly controlled, occurring only at predefined entry and exit points, and data transfer must adhere to secure policies.

The switch gate mechanism works as follows:

  1. When a compartment needs to call into another, it invokes a switch gate using a gate ID.
  2. The switch gate retrieves metadata (e.g., target address, required P-key configuration) from the write-protected SGT.
  3. It then verifies the source address to ensure the call originates from a legitimate entry point.
  4. If the source and target compartments reside in different address spaces (a feature of the two-level compartmentalization), the CR3 register (Page Table Base Register) is updated to switch to the target's address space.
  5. The PKRS register is updated using WRMSR (Write to Model Specific Register) to reflect the target compartment's required P-key permissions.
  6. The system switches to the private stack of the target compartment.
  7. Finally, execution jumps to the target address specified in the SGT.

These metadata checks are crucial for guaranteeing that the compartment interface cannot be misused by attackers.

A major challenge for PKS-based solutions is scalability. PKS is limited to 16 distinct P-keys, meaning only 16 compartments can be directly isolated within a single address space. The Linux kernel, however, contains thousands of LKMs. To overcome this, BULKHEAD proposes a sophisticated two-level compartmentalization scheme:

  1. First Level (Intra-Address Space Isolation): This level utilizes PKS to isolate up to 16 modules within a single address space, each assigned a unique P-key.
  2. Second Level (Locality-Aware Address Space Switching): To scale beyond 16, BULKHEAD divides page tables into shared and private parts. It groups modules into multiple address spaces based on the locality of module interactions (i.e., how frequently they communicate or share data). This is determined by a module dependence specification. Each address space can then reuse its own set of 16 P-keys. For instance, OKM1 and OKM2 might be isolated within the same address space using different P-keys. OKM1 and an unrelated OKM3 could be in different address spaces, potentially even reusing the same P-key, with isolation provided by the address space boundary and ASID (Address Space ID) management.

Finally, BULKHEAD prioritizes performance and efficient data transfer. Hardware-based compartment switching primarily involves updating specific registers, making it significantly faster than software-based context switches. For data sharing, BULKHEAD supports zero-copy data transfer. It "high-tags" shared data with a specific P-key, enforcing a single ownership model. When a target compartment attempts to access this shared data, it triggers a page fault. The monitor then validates the access and dynamically updates the P-keys in the dedicated handler, allowing access without the need for data copying across compartments, which is a common performance bottleneck in other isolation schemes.

Demo / Proof of Concept

▶ Watch: Solving scalability: Two-level compartmentalization scheme (7:00)

While the talk did not detail a specific live demonstration or proof-of-concept execution, the speakers presented comprehensive evaluation results that serve as strong evidence of BULKHEAD's effectiveness. The evaluation focused on three critical domains: security, scalability, and performance.

For security, a comprehensive security analysis was performed against various attack vectors, including data-oriented attacks and control-flow hijacking attacks. The results unequivocally showed that BULKHEAD successfully detected all malicious attempts. This was primarily achieved through two mechanisms: either a memory protection key violation was triggered, or the in-kernel monitor rejected the malicious action based on its integrity checks. This robust validation underscores the system's ability to prevent exploitation effectively.

In terms of performance, BULKHEAD demonstrated impressive efficiency. The evaluation measured overhead for real-world applications with 160 compartmentized I/O modules. The system incurred an average performance overhead of just 2.44%. This low overhead is a direct benefit of BULKHEAD's lightweight switch gate mechanism and its locality-aware design. For more specific benchmarks, an Apache bench test on IPv6 showed an overhead of approximately 2%, which the speaker highlighted as significantly better than existing software-based isolation techniques. Memory overheads were also negligible, reported at 1.66% for LM bench and 0.63% for foriconics, indicating that BULKHEAD imposes minimal resource demands on modern systems. The evaluation also confirmed that increasing the number of compartments does not significantly increase overhead, further validating its scalability.

Defensive Implications

▶ Watch: Performance optimizations: hardware-based switching and zero-copy (8:20)

BULKHEAD offers several critical insights and practical implications for system defenders aiming to secure operating system kernels against the ever-growing threat landscape. Adopting the principles and mechanisms demonstrated by BULKHEAD can significantly enhance kernel resilience.

  1. Embrace Kernel Compartmentalization: The most fundamental takeaway is the urgent need to move away from monolithic kernel architectures. Defenders should advocate for and implement compartmentalization strategies to limit the "blast radius" of successful exploits. By isolating kernel modules into distinct, least-privileged compartments, a vulnerability in one component will not automatically compromise the entire system.
  2. Leverage Hardware Security Features: Modern hardware features like Intel's PKS provide a powerful, low-overhead foundation for enforcing memory protection. Defenders should actively explore and integrate such features into their security architectures to achieve fine-grained, hardware-backed isolation that is difficult for attackers to bypass.
  3. Implement Strict Interface Controls: The Switch Gate Table (SGT) mechanism in BULKHEAD highlights the importance of rigorously controlled interfaces between compartments. Defenders should ensure that all inter-compartment communication occurs only through predefined, validated entry/exit points, with strict policy checks to prevent confused deputy attacks and other forms of interface misuse.
  4. Enforce Execute-Only Memory (XOM): Making kernel code eXecute-Only is a powerful mitigation against code injection and code reuse attacks. Defenders should ensure that kernel code pages are not writable or readable by any entity, thereby preventing attackers from injecting malicious code or crafting ROP chains using existing code snippets.
  5. Adopt Multi-Level Isolation for Scalability: For complex kernels with numerous modules, the two-level compartmentalization scheme presented by BULKHEAD is crucial. Defenders should consider a hybrid approach that combines hardware-based intra-address space isolation with locality-aware address space switching to manage hundreds or thousands of compartments efficiently without hitting hardware limits.
  6. Prioritize In-Kernel Monitoring: A dedicated, highly privileged, and well-protected in-kernel monitor is essential for maintaining the integrity of access control metadata (like page tables and SGT). This monitor acts as a trusted arbiter, ensuring that core security policies are enforced and cannot be subverted by compromised compartments.
  7. Binary Rewriting for Privilege Reduction: Techniques like binary rewriting, used by BULKHEAD to eliminate privileged instructions in non-monitor compartments, offer a proactive way to reduce the attack surface. Defenders should investigate tools and methods to strip unnecessary privileges from compiled kernel components.

In essence, defenders should strive to implement the principle of least privilege at the kernel level, confining each component to only the resources and capabilities absolutely necessary for its function. BULKHEAD provides a concrete, high-performance blueprint for achieving this goal using readily available hardware features.

Key Takeaways

  • Monolithic kernel architectures are highly vulnerable, with a single exploit capable of compromising the entire system, as evidenced by thousands of recent Linux kernel CVEs.
  • Intel's Protection Keys for Supervisor Mode (PKS) offers a powerful hardware-backed mechanism for fine-grained memory compartmentalization within the kernel, partitioning memory into up to 16 domains.
  • BULKHEAD introduces a novel two-level compartmentalization scheme (PKS-based intra-address space isolation combined with locality-aware address space switching via ASID) to overcome the 16-domain PKS limitation, enabling scalable isolation for thousands of kernel modules.
  • A lightweight, protected in-kernel monitor and a Switch Gate Table (SGT) are critical components for enforcing data integrity, control flow protection (XOM), and secure, validated inter-compartment communication, effectively preventing confused deputy attacks.
  • BULKHEAD achieves robust security against data-oriented and control-flow hijacking attacks with minimal performance overhead (average 2.44% for real-world applications) and negligible memory footprint, making it a practical solution for modern kernels.
  • The work underscores the importance of adopting principled kernel compartmentalization based on the principle of least privilege to build more secure and resilient operating systems.

About the Speaker(s)

Yinggang Guo is the presenter of the research work "BULKHEAD: Secure, Scalable, and Efficient Kernel Compartmentalization with PKS" at the NDSS Symposium. The transcript indicates his role as a researcher explaining the intricacies and findings of the BULKHEAD project. No further biographical details were provided in the talk or metadata.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Solid systems security research that delivers on all three adjectives in its title: secure, scalable, and efficient. The two-level compartmentalization scheme — combining PKS for intra-address-space isolation with locality-aware ASID-based address space switching — is a genuinely clever engineering response to PKS's 16-domain ceiling, and the 2.44% average overhead number is respectable enough to make the 'practical for production' claim defensible rather than aspirational.

Heather Calloway (CISO) — WEAK

Technically credible systems security research with real engineering merit — PKS-based kernel compartmentalization at 2.44% overhead is a legitimate result. But this is a researcher talking to researchers, and the talk makes no attempt to bridge to the people who would actually decide whether this gets deployed, funded, or built into Linux distributions.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025