Endokernel: A Thread Safe Monitor for Lightweight Subprocess Isolation
Fangfei Yang (R University), Bumjin Im, Weijie Huang, Kelly Kaoudis, Anjo Vahldiek-Oberwagner, Chia-Che Tsai, Nathan Dautenhahn
33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24
Overview
The talk "Endokernel: A Thread Safe Monitor for Lightweight Subprocess Isolation" by Fangfei Yang and collaborators from Rice University and other institutions, introduces Endokernel, a novel approach to achieving robust intra-process isolation in multi-threaded environments. Modern applications, even those designed with internal security boundaries, remain vulnerable to exploitation due to the inherent complexities of managing shared resources and kernel interactions within a single process space. A seemingly isolated module, if compromised, can leverage system calls and multi-threading race conditions to bypass monitor checks and access sensitive data in other modules.

Key moments
- 0:00 Introduction to intra-process isolation and security challenges
- 2:00 Multi-threaded monitor insecurity: race conditions and system call gaps
- 3:00 Detailed example: Memory leak attack exploiting race conditions
- 4:20 Solution for system call race conditions: forcing secure outcomes
- 5:00 The sigreturn vulnerability and its unique synchronization challenge
- 6:10 Endokernel's novel approach to secure signal handling
- 6:50 High memory access problem: kernel bypassing protection checks
- 7:50 Endokernel's evaluation: overhead, compatibility, and performance
Endokernel: A Thread Safe Monitor for Lightweight Subprocess Isolation
Speakers: Fangfei Yang, Bumjin Im, Weijie Huang, Kelly Kaoudis, Anjo Vahldiek-Oberwagner, Chia-Che Tsai, Nathan Dautenhahn
Conference: USENIX Security '24
YouTube: https://www.youtube.com/watch?v=nWeJ4c5vFVg
Overview
The talk "Endokernel: A Thread Safe Monitor for Lightweight Subprocess Isolation" by Fangfei Yang and collaborators from Rice University and other institutions, introduces Endokernel, a novel approach to achieving robust intra-process isolation in multi-threaded environments. Modern applications, even those designed with internal security boundaries, remain vulnerable to exploitation due to the inherent complexities of managing shared resources and kernel interactions within a single process space. A seemingly isolated module, if compromised, can leverage system calls and multi-threading race conditions to bypass monitor checks and access sensitive data in other modules.
Endokernel addresses critical security gaps that arise when an intra-process monitor, responsible for enforcing isolation policies, interacts with the underlying operating system kernel. The core problem lies in the kernel's lack of awareness regarding intra-process isolation policies, leading to inconsistencies between the monitor's internal state and the kernel's actual state. These inconsistencies are particularly acute in multi-threaded applications, where race conditions can be exploited by attackers to subvert memory protection, compromise monitor integrity, and leak sensitive information.
This work is highly significant because it tackles a fundamental challenge in secure software design: how to maintain strong isolation guarantees within a process when the kernel itself can be manipulated or bypassed. By identifying and mitigating specific vulnerabilities related to system call atomicity, signal handling, and direct physical memory access, Endokernel provides a pragmatic and effective solution for building more secure, compartmentalized applications. It offers a path forward for developers aiming to enhance the resilience of complex software against internal compromises without resorting to the significant overheads of full process-level isolation.
Background
▶ Watch: Introduction to intra-process isolation and security challenges (0:00)
The concept of intra-process isolation aims to enhance application security by creating logical boundaries between different modules or components within a single process. Traditionally, a bug in one module of a monolithic application, such as an Nginx parser, could lead to a compromise of other modules, like OpenSSL, because all modules share the same address space and privileges. Intra-process isolation seeks to mitigate this by leveraging hardware memory protection mechanisms (e.g., page tables) and an intra-process monitor to enforce policies regarding memory, file access, and network communication. The monitor acts as a gatekeeper, intercepting system calls and ensuring that modules only access resources they are permitted to.
However, a critical challenge arises because the hardware protection mechanism, the isolation policy, and the kernel are fundamentally separate components that do not inherently cooperate. The kernel, designed for general-purpose resource management, does not model or understand the granular intra-process isolation policies enforced by a monitor. This disconnect creates a significant vulnerability: an attacker can manipulate the kernel or exploit its interfaces to bypass the monitor's checks and compromise its integrity. For instance, without even faking a system call, an attacker could potentially use legitimate system calls to modify memory protection settings, thereby breaking isolation.
The complexity is compounded in multi-threaded environments. A monitor attempting to manage kernel objects and their accessibility for all threads at all times faces immense difficulty in maintaining consistent, secure state. While system calls appear atomic from a single-thread perspective, in a multi-threaded context, they introduce "gaps" between the monitor's pre-call checks, the actual kernel execution, and the post-call state update. An attacker can exploit these gaps through race conditions, leading to state inconsistencies where the monitor's understanding of the system no longer reflects the true kernel state. Previous secure monitor designs often failed to account for these multi-threading specific issues, rendering them insecure. Endokernel specifically investigates these kernel interface gaps, many of which are covert, to develop a truly thread-safe monitor.
Key Findings
▶ Watch: Detailed example: Memory leak attack exploiting race conditions (3:00)
Endokernel's research identifies three primary categories of security vulnerabilities that undermine the integrity and effectiveness of intra-process monitors in multi-threaded environments:
- System Call State Inconsistencies and Race Conditions: The most direct problem stems from the inherent non-atomicity of system call execution relative to the monitor's state management. The kernel only ensures its internal consistency but does not synchronize its state with the monitor. This leads to scenarios where the monitor's internal data structures, updated before or after a system call, do not accurately reflect the actual kernel state during the system call's execution. In a multi-threaded environment, an attacker can exploit these temporal gaps by initiating concurrent system calls, leading to race conditions that allow unauthorized access to resources. The talk demonstrates how a region marked as unmapped by the monitor could still be mapped by another thread before the actual
munmapsystem call executes, leading to data leakage.
- Signal Return (
sigreturn) Exploitation: Signals, while generally stateless, can be abused to compromise the internal state of a monitor. Specifically, thesigreturnsystem call, which handles returning from a signal handler, can be exploited. Previous attempts to address signal-related issues involved maintaining state when signals were delivered, but these often failed to synchronize with the monitor. This creates a similar consistency gap to general system calls, but crucially, it cannot be solved by post-execution checks becausesigreturndiverts control flow, making it impossible to intercept and correct state after the fact. This represents a more insidious form of kernel bypass.
- High Memory Access Bypasses: The kernel, under certain circumstances, can bypass virtual memory protection policies and access physical memory directly. This means that the monitor's carefully crafted virtual memory protection rules may not be enforced. Furthermore, some system calls, such as
sendmsgwith a zero-copy flag, might check permissions at one point but delay the actual physical memory access. By the time the physical memory is accessed, the permissions for the address might have changed, allowing an attacker to exfiltrate data. Endokernel's analysis of kernel source code revealed specificioctlcalls and related APIs that facilitate these bypasses, many of which were not previously discussed in the context of intra-process isolation.
These findings highlight that building a truly secure intra-process monitor requires a deep understanding of kernel internals and the subtle interactions between user-space monitors and the kernel, especially under concurrent execution. Simple locking mechanisms are insufficient; a more fundamental re-architecting of how monitors interact with the kernel is necessary.
Technical Deep Dive
▶ Watch: The sigreturn vulnerability and its unique synchronization challenge (5:00)
Endokernel's design directly addresses the identified vulnerabilities by mediating the interaction between the application, the monitor, and the kernel. It operates as an intermediary, effectively taking on the role of a "secondary kernel" within the process space to enforce isolation policies more robustly.
To tackle system call state inconsistencies and race conditions, Endokernel adopts a "pragmatic approach." Recognizing that perfect synchronization between the monitor's internal state and the kernel's execution state during a system call is often infeasible, Endokernel instead forces security-preserving outcomes. For example, in the memory mapping scenario, if a memory region is currently being used (e.g., read by one thread), and another thread attempts an munmap operation, Endokernel will cause the munmap system call to return an error until all active uses of that region are complete. This ensures that malicious system calls fail, preserving system integrity, while normal usage patterns (like concurrent reads) remain unaffected, thus not damaging availability.
The illustrative attack scenario involves three threads:
- Thread 1 attempts to
munmapa secret data region after using it. The monitor marks the region as unmapped internally before the actual kernelmunmapcall. - Thread 2 attempts to
mmapa new region usingMAP_FIXEDon the same address. Because the monitor's internal state shows the region as unmapped, this check passes. The monitor then marks the region as belonging to domain two, before the actual kernelmmapcall. - Thread 3 (also belonging to domain two) attempts a
writesystem call to write data from this (now supposedly domain two-owned) memory to a file. This check also passes.
Crucially, in a multi-threaded environment, the actual kernel system calls (munmap, mmap, write) can be executed in any order. An attacker can strategically slow down threads (e.g., using signals via k_assistant_call) to increase the probability of the race condition. If the kernel mmap for Thread 2 executes before the kernel munmap for Thread 1, and the kernel write for Thread 3 executes before the kernel munmap for Thread 1 fully completes, the secret data can be leaked. Endokernel prevents this by ensuring that the munmap operation would fail until all active references to the memory, including potential concurrent mmap or write operations, are resolved, thereby preventing the remapping and subsequent leakage.
For the issue of signal return (sigreturn) exploitation, Endokernel takes on the role of the actual kernel for signal delivery and return. When the kernel sends a signal, it is first delivered to Endokernel's internal queue. The kernel's sigreturn immediately returns after this delivery, meaning the kernel no longer directly delivers signals to the user application. Instead, when Endokernel's queue is not empty, it delivers the signal to the application on its own. Once the application finishes handling the signal, it processes a sigreturn that does not call the kernel. This complete interposition allows Endokernel to maintain consistent state during signal handling, preventing sigreturn from being used to compromise its internal mechanisms.
Regarding high memory access bypasses, Endokernel addresses the problem of the kernel directly accessing physical memory or delaying physical memory access after permission checks. The team performed a detailed analysis of kernel source code to identify system calls and related APIs that might bypass virtual memory protection. Many of these involve ioctl operations. Endokernel's strategy is to disable these problematic ioctl calls outright where possible. In cases where an application genuinely requires ioctl (e.g., a GPU application), a case-by-case analysis is performed to ensure that the specific ioctl usage does not create security vulnerabilities. This proactive analysis and selective disabling/mediation prevent the kernel from circumventing the monitor's policies via direct physical memory access.
Endokernel's core mechanism for interposition involves system call user dispatch and interception. These techniques allow Endokernel to gain control before system calls are passed to the actual kernel, enabling it to apply its policies and manage state securely. The choice of mechanism impacts performance, with some offering better control flow integrity at the cost of higher overhead, while others provide better performance by only ensuring control flow during context switches.
Demo / Proof of Concept
▶ Watch: Endokernel's novel approach to secure signal handling (6:10)
While the talk does not describe a traditional, live "demo" in the sense of a visual application being run, the speaker provides a detailed conceptual demonstration of the multi-threaded race condition vulnerability and how Endokernel's design mitigates it. This effectively serves as a proof of concept for both the exploit and the defense.
The core of this conceptual demonstration revolves around the memory access race condition:
- Vulnerability Setup: An attacker controls a thread (Thread 1) that has legitimate access to sensitive data and then attempts to unmap it. Concurrently, other attacker-controlled threads (Thread 2, Thread 3) attempt to exploit the time window between the monitor marking memory as unmapped and the kernel actually unmapping it. Thread 2 tries to map a new region at the same address, and Thread 3 tries to write data from that remapped region to a file.
- Exploitation: The speaker explicitly states, "The actual system call may be executed in any order because they are multi threaded and we can slow thread one and two by flaing them with signals using K assistant call so we increase the chance of a successful attack like this." This indicates that the researchers have developed a method to reliably induce and demonstrate this race condition, likely by manipulating thread scheduling or introducing delays, leading to the unauthorized mapping and subsequent leakage of the "secret data" to a file. This confirms the practical exploitability of the identified gap.
- Endokernel's Defense: The proposed solution, as described, would prevent this specific exploit. When Thread 1 attempts to
munmapthe region, Endokernel would detect that other threads (even if they are attempting to map or write to it in a race) might still be interacting with the memory region. Instead of allowing themunmapto proceed immediately, Endokernel would cause it to "return arrow until all uses are done." This ensures that the memory region cannot be prematurely unmapped and then remapped by an attacker, thereby preventing the data leakage, even if the attacker attempts to orchestrate a complex multi-threaded race.
This detailed scenario, combined with the mention of techniques to increase the chance of a successful attack, serves as a clear and concrete proof of concept for the types of vulnerabilities Endokernel aims to solve. It demonstrates that these are not merely theoretical issues but practical exploitation vectors in multi-threaded intra-process isolation designs.
Defensive Implications
▶ Watch: Endokernel's evaluation: overhead, compatibility, and performance (7:50)
The findings and solutions presented by Endokernel offer crucial insights for defenders and developers building secure applications with intra-process isolation. The primary defensive implication is that traditional security approaches, such as simply adding locks around critical data structures within a monitor, are fundamentally insufficient to ensure thread safety and isolation integrity when interacting with the kernel.
Defenders should recognize that:
- Kernel Interfaces are Attack Surfaces: The boundary between an intra-process monitor and the operating system kernel is a significant attack surface. The kernel's lack of understanding of intra-process isolation policies creates implicit trust boundaries that can be exploited.
- Multi-threading Exacerbates Vulnerabilities: Race conditions introduced by concurrent system calls in multi-threaded environments are not edge cases; they are exploitable security flaws. Monitors must be designed with explicit consideration for the non-atomicity of system calls relative to their internal state.
- Beyond Simple Locking: True thread safety for an intra-process monitor requires more than just mutexes or semaphores. It necessitates sophisticated synchronization mechanisms that account for the kernel's asynchronous behavior and potential for out-of-order execution of system calls across threads.
- Signal Handling Requires Mediation: The
sigreturnmechanism, often overlooked, can be a vector for state compromise. Monitors should actively mediate signal delivery and return to maintain control over their internal state, rather than relying on the kernel's default behavior. - Scrutinize Direct Memory Access: Any system call or API that allows the kernel to bypass virtual memory protection and access physical memory directly (e.g., certain
ioctlcalls, zero-copysendmsg) must be thoroughly reviewed and, if possible, disabled or strictly sandboxed. Developers should perform kernel source code analysis to identify such risky APIs. - Performance is a Factor: Implementing robust kernel interposition mechanisms, like system call user dispatch, introduces overhead. While Endokernel demonstrates manageable overheads (e.g., under 25% for Nginx, 36% for
scwithout set, 23% with 32 threads), this must be weighed against the security benefits for specific applications. - Compatibility Testing is Key: The fact that Endokernel passes 95% of LTP tests indicates a high degree of compatibility, but the remaining 5% highlight that some applications or functionalities might require specific adjustments or might be incompatible due to security-driven design choices. Extensive compatibility testing is essential when deploying such a monitor.
Ultimately, defenders should move towards monitor designs that act as a more complete intermediary between the application and the kernel, actively managing and mediating all kernel interactions to prevent state inconsistencies and bypasses. This approach shifts the security burden from hoping the kernel will cooperate to actively ensuring the kernel cannot violate isolation.
Key Takeaways
- Multi-threaded intra-process monitors face significant security challenges beyond simple locking, primarily due to the inherent disconnect between the monitor's state and the kernel's behavior.
- Kernel interfaces create exploitable gaps: Race conditions during system calls (e.g.,
mmap/munmap/write) can lead to state inconsistencies and data leakage in multi-threaded environments. sigreturncan bypass monitor state updates, requiring direct mediation by the monitor to ensure consistent control flow and state management.- Direct physical memory access by the kernel bypasses virtual memory protection, necessitating careful analysis of kernel APIs (like
ioctl) and proactive mitigation strategies. - Endokernel offers a pragmatic solution by forcing security-preserving outcomes during system calls, mediating signal delivery, and identifying/mitigating risky kernel memory access patterns.
- Performance overheads are manageable, with Endokernel achieving under 25% overhead on Nginx and around 23% with 32 threads, demonstrating practical viability.
- High compatibility with existing applications is demonstrated by passing 95% of the Linux Test Project (LTP) cases, indicating a robust and broadly applicable design.
About the Speaker(s)
The primary speaker for this presentation was Fangfei Yang, representing Rice University. The research and development of Endokernel involved a collaborative effort with several co-authors: Bumjin Im, Weijie Huang, Kelly Kaoudis, Anjo Vahldiek-Oberwagner, Chia-Che Tsai, and Nathan Dautenhahn. This collective expertise from various institutions underscores the depth of research and engineering that went into addressing the complex challenges of thread-safe intra-process isolation. While specific titles and affiliations for all co-authors beyond Rice University are not detailed in the provided transcript or metadata, their collective contribution highlights a strong academic and research focus on operating systems security and robust system design.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
Endokernel delivers a critical deep dive into the often-overlooked kernel interface gaps that undermine intra-process isolation in multi-threaded environments. This research exposes severe race conditions, sigreturn exploits, and direct physical memory bypasses, offering a pragmatic and effective solution to a pervasive problem. It's essential viewing for anyone serious about building truly secure, compartmentalized applications.
Heather Calloway (CISO) — STRONG ACCEPT
This research on Endokernel provides critical insights into the pervasive security flaws inherent in intra-process isolation, particularly in multi-threaded environments. It compellingly demonstrates how kernel interactions and race conditions can bypass existing monitors, offering a robust architectural solution that forces security-preserving outcomes and redefines secure software design for high-assurance systems.