Serverless Functions Made Confidential and Efficient with Split Containers
Jiacheng Shi (Shanghai University)
34th USENIX Security Symposium (USENIX Security '25) · Day 1 · System Security 2: Trusted and Robust Computing
Overview
This talk introduces Kofunk, a novel split container architecture designed to make serverless functions both confidential and efficient when leveraging Confidential Virtual Machines (CVMs). Presented by Jiacheng Shi from Shanghai University, the research tackles the significant challenges posed by integrating hardware-backed Trusted Execution Environments (TEEs) like AMD SEV and Intel TDX with the dynamic, ephemeral nature of serverless workloads. The core problem addressed is the high cold-start latency and substantial memory overhead typically associated with CVMs when used for fine-grained isolation, which fundamentally conflicts with the rapid scaling and cost-efficiency promises of serverless computing.

Key moments
- 0:00 Introduction to serverless and confidential containers
- 2:00 Challenges of current confidential container designs
- 4:05 Introducing the Split Container architecture
- 5:50 CPU resource management with shadow containers
- 7:00 Memory resource management via shadow containers
- 8:30 CVM OS implementation: microkernel and LibOS
- 10:00 Optimizing function boot with pre-warming stages
Serverless Functions Made Confidential and Efficient with Split Containers
Speakers: Jiacheng Shi
Conference: USENIX Security
YouTube: https://www.youtube.com/watch?v=dDlnhdSZcFY
Overview
This talk introduces Kofunk, a novel split container architecture designed to make serverless functions both confidential and efficient when leveraging Confidential Virtual Machines (CVMs). Presented by Jiacheng Shi from Shanghai University, the research tackles the significant challenges posed by integrating hardware-backed Trusted Execution Environments (TEEs) like AMD SEV and Intel TDX with the dynamic, ephemeral nature of serverless workloads. The core problem addressed is the high cold-start latency and substantial memory overhead typically associated with CVMs when used for fine-grained isolation, which fundamentally conflicts with the rapid scaling and cost-efficiency promises of serverless computing.
The talk highlights a critical mismatch between the traditional one-CVM-per-function isolation model and serverless requirements. While running each function in its own CVM provides strong security, it incurs unacceptable boot times (hundreds of milliseconds to seconds) and exorbitant memory consumption. Conversely, placing multiple functions within a single CVM compromises isolation due to the large attack surface of a shared guest operating system. Kofunk proposes a unique solution: securely isolating multiple confidential containers within a single CVM by employing a minimal microkernel, while offloading non-security-critical management tasks to shadow containers on the host.
This research is highly significant for the future of confidential computing, especially in cloud environments. By demonstrating a practical method to achieve strong confidentiality for serverless functions without sacrificing performance or resource efficiency, Kofunk opens the door for adopting serverless in highly sensitive use cases such as facial recognition, healthcare, and finance. It provides a blueprint for cloud providers and enterprises to offer truly confidential serverless platforms, where user code and data remain protected even from a compromised cloud infrastructure or hypervisor, thus expanding the scope and trustworthiness of serverless deployments.
Background
▶ Watch: Introduction to serverless and confidential containers (0:00)
Serverless computing has emerged as a dominant paradigm over the last decade, enabling developers to focus solely on business logic by abstracting away infrastructure management. Its appeal lies in elastic scaling, pay-per-use billing, and automated management. Consequently, serverless functions are increasingly adopted in security-critical domains like facial recognition, healthcare, and finance, where data confidentiality and integrity are paramount. However, the security of these functions traditionally relies on the entire software stack, encompassing the host OS, hypervisor, and management tools—a Trusted Computing Base (TCB) that is prohibitively large and vulnerable to various attacks.
To address this expansive TCB, confidential containers have been introduced, combining Trusted Execution Environments (TEEs) with containerization. TEEs, particularly hardware-backed solutions like Confidential Virtual Machines (CVMs) such as AMD SEV and Intel TDX, offer robust protection for a virtual machine's confidentiality and integrity, including its guest kernel and user space, against external attacks from the host or hypervisor. CVM-based confidential containers aim to bring this hardware-level protection to containerized workloads.
However, a fundamental mismatch exists between the design principles of CVMs and the operational requirements of serverless functions. A common design choice for confidential containers, exemplified by Kata Containers, is to run each function instance within a separate CVM to ensure strong isolation. While secure, this approach faces significant performance and resource challenges. Booting a single CVM, even with microVM optimizations, can take 334 milliseconds for AMD SEV-CVM and 1.8 seconds for Intel TDX-CVM. Concurrent CVM startups exacerbate this; for instance, concurrent SEV-CVM startups can reach 28.5 seconds due to limitations like a single Platform Security Processor (PSP) in SEV hardware. Such overheads are unacceptable for serverless functions, where 50% of real-world invocations complete in under 1 second. Mitigating this with a pool of pre-launched CVM instances is memory-intensive, as cached CVMs cannot share common memory due to different memory encryption keys, leading to 42.5 GB of memory consumption for 500 CVMs.
An alternative is to run multiple functions within a single CVM to eliminate boot overhead and enable memory sharing. Yet, this compromises isolation, as functions would share a large, vulnerable guest Linux kernel, expanding the TCB within the CVM itself. The researchers observed that serverless functions are inherently lightweight, single-process, and stateless. They typically do not require multi-user capabilities, multiprocessing, inter-process communication, or persistent storage. Furthermore, they rarely utilize control plane OS services like resource monitoring or device management. A survey of open-source functions revealed that they only leverage a small subset of Linux system calls. This insight forms the basis for the proposed solution: if serverless functions don't need a full-fledged Linux kernel, they can run on a much more lightweight CVM OS, enabling secure multi-tenancy within a single CVM.
Key Findings
▶ Watch: Introducing the Split Container architecture (4:05)
The research presents Kofunk, a groundbreaking split container architecture that fundamentally redefines how confidential computing integrates with serverless platforms. The key findings and contributions are:
- Secure Multi-Tenancy within a Single CVM: Kofunk demonstrates that multiple confidential containers can be securely isolated and run within a single CVM by leveraging a minimal microkernel. This approach effectively resolves the dilemma of strong isolation versus efficiency, allowing for shared CVM resources without compromising the confidentiality and integrity of individual functions.
- Decoupling Security and Management: The architecture achieves microkernel minimization by strategically decoupling security-critical tasks from non-security-critical management functions. It delegates resource management (CPU, memory) and I/O operations to shadow containers running on the untrusted host Linux kernel, while the security-sensitive isolation and system call handling occur within the CVM's tiny microkernel. This significantly shrinks the TCB inside the CVM to approximately 20,000 lines of code.
- Efficient Function Boot and Attestation: Kofunk introduces a multi-stage optimization for function boot and attestation. By pre-warming the CVM and creating "cycles" (pre-initialized containers with common runtimes and libraries), new function instances can be efficiently forked using copy-on-write (CoW). Attestation is split, verifying the CVM microkernel once and then individual function code incrementally, drastically reducing startup latency.
- Dramatic Performance Improvements: Empirical evaluation of Kofunk against Kata Containers (a state-of-the-art per-container CVM design) shows substantial performance gains:
- Up to 215 times reduction in end-to-end latency for single requests on Intel TDX.
- Up to 60 times reduction in end-to-end latency on AMD SEV.
- Container startup latency reduced by 120 times on TDX and 22.3 times on SEV for single containers, and 100 times on SEV for 200 concurrent containers, primarily by eliminating CVM boot overhead.
- Code loading and initialization latency reduced by up to 500 times through cycled fork and split attestation.
- Overall end-to-end overhead compared to unprotected native containers is less than 14%.
- Significant Resource Efficiency: Kofunk dramatically reduces memory consumption. By allowing confidential containers to share common memory using copy-on-write, it achieved up to 56 times reduction in memory usage for 200 Python containers compared to Kata Containers.
- Optimized Function Chaining: The architecture enables functions within the same CVM to communicate via shared memory, eliminating network and encryption overheads for inter-CVM communication. This results in up to 31 times reduction in end-to-end latency for chained applications on AMD SEV platforms.
These findings collectively demonstrate a viable path for deploying confidential serverless functions at scale, overcoming the critical performance and resource efficiency hurdles that previously hindered widespread adoption.
Technical Deep Dive
▶ Watch: CPU resource management with shadow containers (5:50)
The core innovation of Kofunk lies in its split container architecture, which securely isolates multiple confidential containers within a single CVM using a minimalist microkernel. This is achieved by strategically offloading non-security-critical functionalities to the untrusted host, while maintaining a hardened, minimal TCB within the CVM.
At the heart of the architecture is the concept of pairing each confidential container inside the CVM with a shadow container on the host. These shadow containers are standard Linux containers that provide the necessary host-level resource management and I/O delegation.
Microkernel Minimization
Kofunk's microkernel inside the CVM is significantly smaller than a full Linux kernel, comprising only about 20,000 lines of code. This minimization is achieved through two key strategies:
- Host-based Containerization Pass-Through: Instead of reimplementing complex containerization features like cgroups and namespaces inside the CVM, Kofunk leverages the host Linux kernel's capabilities. The shadow container on the host is responsible for granting CPU and memory resources to its confidential counterpart and handling its I/O requests. This avoids duplicating a large and complex codebase within the CVM's TCB.
- Microkernel + Library OS (LibOS) Architecture: The CVM guest OS is implemented using a microkernel plus library OS model. The microkernel focuses solely on security-critical isolation and privileged operations, while a LibOS provides the user-level system call functionalities required by serverless functions. This separation further reduces the complexity and attack surface of the security-critical microkernel.
CPU Resource Management
CPU resource allocation is managed by the host's cgroups, seamlessly extended into the CVM:
- One-to-One Mapping: Kofunk establishes a direct one-to-one mapping between threads in the shadow containers on the host and confidential threads within the CVM.
- vCPU Allocation: Each confidential thread exclusively occupies one virtual CPU (vCPU) inside the CVM. When a shadow container is created, a shadow thread activates an idle vCPU from a CVM-internal vCPU pool via the host kernel. This vCPU then launches the confidential container and becomes the confidential thread.
- Scheduling Inheritance: The confidential thread inherits the CPU time slices allocated to its corresponding shadow thread by the host's cgroup. If the host kernel's timer interrupts the vCPU and the shadow thread has exhausted its time slice, the host reschedules. When the shadow thread is scheduled again, it resumes its vCPU and the confidential thread, ensuring resource control via the host's cgroups.
Memory Resource Management
Memory management also leverages host cgroups, with careful handling of confidential memory:
- Shadow Container Accounting: Memory usage of a confidential container is accounted for within the memory cgroup of its corresponding shadow container.
- GPA Pool and HVA Binding: When a confidential container is launched, the CVM microkernel allocates a Guest Physical Address (GPA) range as its memory pool. The host kernel then establishes bindings between this GPA pool and a Host Virtual Address (HVA) range in the shadow container, ensuring they share the same host physical address mapping.
- Nested Page Fault Handling: During execution, if a confidential container triggers a page fault for an unmapped GPA, the CVM microkernel allocates a GPA from its pool. If this GPA is not yet mapped to a physical page, the microkernel uses CVM hardware primitives (e.g., AMD SEV's
PVALIDATEinstruction) to trigger a nested page fault. The host kernel handles this nested page fault as if it originated from the HVA range of the shadow container, charging the newly allocated physical page to the shadow container's memory cgroup. This mechanism ensures that the host controls physical memory allocation and accounting, while the CVM microkernel manages the confidential container's page table, which is security-critical.
CVM OS: Microkernel + LibOS
The CVM OS built on the research microkernel Chore implements the LibOS to serve serverless functions:
- System Call Support: The LibOS provides 74 system calls identified as necessary for serverless functions, based on the researchers' study.
- Memory Management: The LibOS offers system calls for updating the virtual memory space, while the CVM microkernel is solely responsible for updating the confidential container's page table, a security-critical operation.
- Synchronization Primitives: For multi-threading language runtimes like NodeJS, the LibOS provides essential synchronization primitives, including futex, eventfd, and pipe.
- Virtual File System: A virtual file system dispatches system calls based on file descriptors to different subsystems.
- Root File System (Root FS): Used for loading code and read-only data. These operations are delegated to the shadow container on the host, with an attestation mechanism ensuring file integrity.
- Temporary File System (Temp FS): Implemented entirely within the LibOS for storing temporary data.
- Network I/O: Network operations are delegated to the shadow container, with all data protected by end-to-end encryption to maintain confidentiality over the untrusted host network stack.
Optimized Function Boot
Kofunk employs a three-stage pre-warming and split attestation process to optimize function boot:
- Pre-warming Stage 1 (CVM and Microkernel Attestation): The serverless platform starts a CVM. A Key Management Service (KMS) remotely attests the CVM to verify that the CVM microkernel has been faithfully booted. Upon successful attestation, the KMS securely sends tenant-specific encryption keys to the CVM.
- Pre-warming Stage 2 (Cycles Creation): A group of "cycles" are created within the attested CVM. These cycles are pre-initialized confidential containers containing common language runtimes (e.g., Python, Node.js) and libraries that can be shared across different functions. When this common code is loaded from the host file system, its content is verified using the tenant key and a hash value stored in a metadata file.
- Function Invocation Stage 3 (Fast Forking): When a tenant invokes a function, the platform can efficiently fork a new confidential container from an existing "cycle" using copy-on-write. The newly forked container only needs to load and verify function-specific code from the root file system before it can handle the invocation, significantly reducing cold-start latency.
This detailed architecture allows Kofunk to provide robust confidentiality guarantees for serverless functions, even against a malicious host, while achieving performance and resource efficiency competitive with or superior to existing confidential container solutions.
Demo / Proof of Concept
▶ Watch: CVM OS implementation: microkernel and LibOS (8:30)
The researchers built a prototype of the split container architecture named Kofunk. This prototype was implemented and evaluated on both AMD SEV and Intel TDX platforms, demonstrating its hardware agnosticism and practical applicability across leading CVM technologies.
For evaluation, Kofunk was tested with 28 distinct functions sourced from four different serverless benchmarks. This comprehensive set of workloads provided a realistic assessment of its performance and efficiency across diverse serverless use cases.
Kofunk's performance was directly compared against Kata Containers, which represents the state-of-the-art in confidential containers using a per-container CVM design. The results were compelling:
- End-to-End Latency Reduction: On Intel TDX, Kofunk reduced the end-to-end latency of a single request by an astonishing up to 215 times compared to Kata Containers. Even on AMD SEV, where Kata Containers benefits from microVM optimizations, Kofunk still achieved up to a 60 times reduction in end-to-end latency. These figures highlight Kofunk's ability to overcome the significant cold-start issues inherent in traditional CVM deployments for serverless.
- Container Startup Overhead: The startup process was broken down into two stages: environment preparation, and code loading/initialization.
- In the first stage (environment preparation), Kofunk eliminated CVM boot overhead by sharing the CVM, leading to latency reductions of 120 times on TDX and 22.3 times on SEV for booting a single container. For concurrent scenarios, it reduced latency by 100 times on SEV when booting 200 concurrent containers.
- In the second stage (code loading and initialization), Kofunk's use of "cycled fork" and "split attestation" reduced latency by up to 500 times.
- Overhead Comparison with Native Containers: When compared to unprotected native containers, Kofunk introduced an end-to-end overhead of less than 14%. This overhead primarily stems from function code measurement during startup and the necessary I/O delegation, data encryption, and memory granting operations during execution, which are fundamental costs of confidential computing.
- Memory Efficiency: Kofunk's copy-on-write mechanism for shared memory significantly reduced memory consumption. For 200 Python containers, it achieved up to a 56 times reduction in memory usage compared to Kata Containers.
- Function Chaining Performance: Kofunk also demonstrated substantial benefits for applications involving chained functions. By enabling functions within the same CVM to communicate via shared memory, it eliminated network and encryption overheads typically incurred by inter-CVM communication. This resulted in a 31 times reduction in end-to-end latency for chained applications on the AMD SEV platform.
These quantitative results underscore the practical viability and superior performance of the Kofunk split container architecture, proving its effectiveness in addressing the dual challenges of confidentiality and efficiency in serverless computing.
Defensive Implications
▶ Watch: Optimizing function boot with pre-warming stages (10:00)
The Kofunk split container architecture offers significant defensive implications for organizations looking to deploy security-critical serverless workloads in cloud environments. By integrating CVMs with a highly optimized design, it enables a new era of confidential serverless computing with robust protections.
- Stronger Confidentiality and Integrity: Kofunk provides hardware-backed confidentiality and integrity for serverless functions and their data, even against a fully compromised host OS, hypervisor, or cloud administrator. This is crucial for sensitive applications in healthcare, finance, or facial recognition, where data at rest and in transit are typically protected, but data in use remains vulnerable to privileged software attacks.
- Minimized Attack Surface within the CVM: The use of a highly optimized microkernel (approximately 20,000 lines of code) as the CVM guest OS drastically shrinks the Trusted Computing Base (TCB) within the confidential boundary. This significantly reduces the potential attack surface for adversaries attempting to exploit vulnerabilities within the guest environment, a stark contrast to the large and complex Linux kernel used in traditional multi-tenant CVMs.
- Enhanced Multi-Tenant Isolation: Kofunk allows multiple confidential functions from different tenants to securely co-exist within a single CVM. The microkernel-based isolation within the CVM, combined with hardware memory encryption keys, ensures that one function cannot access the memory or state of another, even if they share the same CVM. This is a critical improvement over approaches that might share a full Linux kernel, which can have complex isolation mechanisms prone to bypasses.
- Verifiable Execution Environment: The multi-stage attestation process ensures that the CVM microkernel is faithfully booted and that common language runtimes and function-specific code are loaded and verified against tenant keys. This provides cryptographic assurance to tenants that their code is running in an untampered, trusted environment.
- Secure I/O Delegation: While I/O operations are delegated to the untrusted host, network data is protected with end-to-end encryption. This ensures that sensitive data remains confidential even when it traverses the untrusted host's network stack. Similarly, root file system content for code loading is verified against tenant keys, maintaining its integrity.
- Mitigation of Cold Start and Resource Exhaustion Attacks: By significantly reducing cold-start latency and memory consumption, Kofunk makes confidential serverless more resilient to Denial of Service (DoS) attacks that exploit resource exhaustion or slow startup times. The ability to efficiently fork containers from pre-warmed "cycles" ensures rapid scaling and availability even under attack.
- Enabling New Use Cases: The combination of strong security with high performance and efficiency enables organizations to confidently migrate previously unsuited, highly sensitive workloads to serverless platforms, unlocking new levels of agility and cost savings without compromising security.
Defenders should actively consider Kofunk-like architectures when designing confidential cloud strategies. It provides a robust framework for securing serverless functions, shifting the security boundary from the broad cloud infrastructure to the hardware-protected CVM, and significantly narrowing the TCB that needs to be trusted. While the host still controls resource allocation, Kofunk ensures the confidentiality and integrity of the function's execution and data, even if the host is compromised.
Key Takeaways
- Solving CVM Cold Start and Memory Overhead: Kofunk's split container architecture effectively addresses the prohibitive cold-start latency (up to 215x faster) and high memory consumption (up to 56x less) that previously hindered the widespread adoption of CVMs for serverless functions.
- Minimal TCB with Microkernel + LibOS: By employing a small, security-focused microkernel (20,000 LOC) inside the CVM and a library OS for function syscalls, Kofunk drastically reduces the Trusted Computing Base, enhancing security.
- Leveraging Host for Efficiency: The architecture innovatively delegates non-security-critical resource management (CPU and memory cgroups) and I/O operations to shadow containers on the untrusted host, achieving efficiency without compromising confidentiality within the CVM.
- Fast Boot with Cycled Forking and Split Attestation: A multi-stage pre-warming and attestation process, including "cycles" (pre-initialized containers) and copy-on-write forking, enables new confidential functions to boot exceptionally fast (up to 500x faster code loading).
- Enabling Secure and Efficient Multi-Tenant Serverless: Kofunk provides strong, hardware-backed isolation for multiple confidential functions from different tenants within a single CVM, making truly confidential and performant serverless a reality for sensitive workloads.
- Significant Performance Gains for Chained Applications: The ability for functions within the same CVM to communicate via shared memory substantially reduces latency for chained serverless applications (up to 31x faster), further boosting overall application performance.
About the Speaker(s)
The talk was presented by Jiacheng Shi, a researcher from Shanghai University. The presentation focused on his work in making serverless functions confidential and efficient using split containers, demonstrating a deep understanding of confidential computing, trusted execution environments, and serverless architectures. His research contributes to advancing the security and practicality of cloud computing paradigms.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Kofunk is legitimate systems security research that solves a real and annoying problem: CVMs are too heavy for serverless, but sharing a full Linux kernel across tenants inside a single CVM is a TCB disaster. The split-container approach — microkernel inside the CVM, shadow containers handling resource accounting on the untrusted host — is architecturally clean and the numbers are credible. 215x latency reduction against Kata Containers on TDX and 56x memory savings aren't marketing — those reflect real CVM boot pathology that anyone who's benchmarked SEV or TDX has hit personally.
Heather Calloway (CISO) — WEAK
Technically rigorous systems research on confidential serverless computing that solves a real architectural problem — but it stops at the engineering boundary and never crosses into operator, governance, or enterprise decision-making territory. The defensive implications section is surface-level retrofitting, not genuine practitioner guidance.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)