The Life (or Death) of a Kubernetes API Request, 2025 Edition - Abu Kashem & Stefan Schimanski

Abu Kashem, Stefan Schimanski

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk, "The Life (or Death) of a Kubernetes API Request, 2025 Edition," delivered by Abu Kashem and Stefan Schimanski, provides an exhaustive technical deep dive into the intricate journey of a Kubernetes API request, from the moment a user presses enter on a kubectl command to its eventual persistence in etcd and the return of a response. Building upon a foundational talk from nearly a decade ago, this updated presentation addresses the significant architectural and functional evolutions within the Kubernetes API server, a critical component whose complexity has grown substantially with the project's maturity.

Watch on YouTube

Visual summary for The Life (or Death) of a Kubernetes API Request, 2025 Edition - Abu Kashem & Stefan Schimanski by Abu Kashem, Stefan Schimanski
Visual summary for The Life (or Death) of a Kubernetes API Request, 2025 Edition - Abu Kashem & Stefan Schimanski by Abu Kashem, Stefan Schimanski

Key moments

  1. 0:00 Introduction: The "kubectl create" interview question
  2. 2:40 Kubectl's REST mapping: Kind, Resource, Group Version
  3. 3:20 API server receives request: Go HTTP server basics
  4. 4:40 Understanding the API server's handler chain architecture
  5. 5:40 Request context evolution through the API handler chain
  6. 6:20 Deep dive into the API server's panic recovery
  7. 8:00 Client and server observe API server panic behavior

The Life (or Death) of a Kubernetes API Request, 2025 Edition

Speakers: Abu Kashem, Software Engineer, Red Hat; Stefan Schimanski, Staff Software Engineer

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=Hc0jj-654lA

Overview

This talk, "The Life (or Death) of a Kubernetes API Request, 2025 Edition," delivered by Abu Kashem and Stefan Schimanski, provides an exhaustive technical deep dive into the intricate journey of a Kubernetes API request, from the moment a user presses enter on a kubectl command to its eventual persistence in etcd and the return of a response. Building upon a foundational talk from nearly a decade ago, this updated presentation addresses the significant architectural and functional evolutions within the Kubernetes API server, a critical component whose complexity has grown substantially with the project's maturity.

The speakers meticulously trace the path of a kubectl create request, highlighting the numerous handlers, validation steps, conversion mechanisms, and self-defense layers that govern its processing. This detailed exploration is crucial for anyone seeking to understand the inner workings of Kubernetes, particularly cluster administrators, developers of custom controllers or webhooks, and security professionals. By dissecting each stage, the talk not only demystifies the API server's operations but also equips the audience with essential knowledge for debugging performance bottlenecks, diagnosing errors, and enhancing the resilience of their Kubernetes environments.

Understanding the full lifecycle of an API request is paramount for effective Kubernetes management. The talk emphasizes how seemingly simple operations trigger a cascade of complex, interdependent processes within the API server. For an ecosystem where every interaction, from deploying applications to scaling resources, hinges on API requests, this comprehensive analysis offers invaluable insights into optimizing cluster performance, ensuring stability under load, and recognizing the critical touchpoints for security and operational integrity.

Background

▶ Watch: Introduction: The "kubectl create" interview question (0:00)

The talk frames its exploration around a common interview question: "What happens when you type kubectl create -f manifest.yaml and press enter?" This seemingly straightforward query opens the door to the vast complexities hidden beneath the surface of Kubernetes. While a similar talk by Daniel Smith existed years ago, the Kubernetes project has evolved dramatically, introducing new features, architectural patterns, and defense mechanisms that necessitate a fresh examination. The speakers note that a significant portion of the current Kubernetes community likely joined the project after that initial presentation, making this update particularly relevant.

The journey begins with kubectl's initial actions, specifically how it translates a user-provided YAML manifest into an HTTP POST request. A subtle but crucial detail highlighted early on is the distinction between a Kubernetes Kind (e.g., Job in a manifest) and a Resource (e.g., jobs in the API path /apis/batch/v1/jobs). This mapping, which is far from trivial, is managed through discovery. kubectl queries the API server's discovery endpoint to obtain REST mappings, which provide the necessary metadata to construct the correct API path for a given Kind and GroupVersion. This initial phase ensures that the client correctly addresses the API server, setting the stage for the request's deeper processing. The raw HTTP request, often seen with kubectl -v9, is then sent to the API server, initiating its complex lifecycle.

Key Findings

▶ Watch: API server receives request: Go HTTP server basics (3:20)

The talk reveals the Kubernetes API server as a highly structured and resilient system, processing requests through an extensive handler chain. Each handler in this chain is responsible for a specific aspect of request processing, from initial HTTP parsing to final storage. A central concept is the request context, which accumulates vital information like deadlines, user identities, and audit events as the request traverses the chain.

Key findings include:

  • Panic Recovery: A dedicated handler proactively catches unexpected panics within request handling code, logging stack traces and ensuring graceful error responses (e.g., HTTP 500) to clients, rather than crashing the server.
  • Request Information: Requests are parsed into a request.Info struct, containing "logical verbs" (e.g., get, list, watch) that abstract over HTTP verbs, along with group, version, resource, and namespace details.
  • Request Lifespan and Timeouts: Requests are categorized as "long-running" (e.g., watches, logs, proxy) or "non-long-running." Non-long-running requests are assigned deadlines (defaulting to 60 seconds if not specified by the client or overridden by the API server's --request-timeout flag). A Timeout Handler actively enforces these deadlines, preventing indefinitely hanging requests and returning a 504 Gateway Timeout if exceeded.
  • API Priority and Fairness (APF): A critical self-defense mechanism that uses a fair queuing algorithm to regulate API server load. Requests are classified into flows, queued, and scheduled. If a request waits in a queue for more than one-fourth of its allotted deadline, it is rejected with an HTTP 429 "Too Many Requests" status and a Retry-After header.
  • Detailed Latency Tracking: The API server records granular latency metrics in audit event annotations, tracking time spent in authentication, authorization, mutating/validating webhooks, storage (etcd), serialization, and response writing. This is invaluable for debugging performance issues when requests exceed a 500-millisecond threshold.
  • API Object Lifecycle: The talk details the internal object lifecycle, involving decoding the client's request (e.g., JSON) into a Go object, defaulting unspecified fields, and crucial conversions between external API versions (e.g., batch/v1) and the API server's internal representation (int).
  • Scheme and Codec: The Scheme acts as a central registry for Kubernetes' type system, connecting group versions, kinds, and Go types, and managing all conversion functions. The Codec leverages the Scheme for encoding and decoding objects to/from various formats (JSON, ProtoBuf).
  • Admission Control and Validation: After initial defaulting, requests undergo multiple phases of validation: mutating webhooks, CEL policies (an alpha feature), API-level Go code validation (e.g., checking name formats), and validating webhooks. Failures at this stage result in a 422 Unprocessable Entity error.
  • Storage in etcd: The final validated internal object is converted to the storage version (e.g., batch/v1 for jobs, often ProtoBuf for native resources) and potentially encrypted via KMS encryption before being stored in etcd using an optimistic put operation. This operation checks for conflicts (e.g., resource already exists), which would result in a 409 Conflict.
  • Client-side Resilience: After the API server processes the request and sends a reply, client-go (the official Kubernetes client library) implements built-in retry logic for network errors, 429s, and 5xx errors with a Retry-After header, enhancing client-side robustness.

Technical Deep Dive

▶ Watch: Understanding the API server's handler chain architecture (4:40)

The journey of a Kubernetes API request is orchestrated by a sophisticated chain of Go functions, each contributing to the request's processing. When an HTTP request arrives, the Go HTTP server (net/http) accepts it, creates an http.Request object (containing headers, parameters, and body) and an http.ResponseWriter, then invokes a user-provided handler. For HTTP/1, this handler runs in the same go routine as the server's connection handler, while HTTP/2 executes it in a new go routine for concurrency.

The API server's architecture eschews a monolithic handler for a handler chain. The initial handler, the panic recovery handler, is critical for stability. It wraps the rest of the chain, executing it in a new goroutine. If any downstream handler panics, this recovery handler catches it, logs the stack trace, and then re-panics by design, allowing the HTTP server to handle the ultimate error response (typically a 500 Internal Server Error). Clients observing this might see "end of file" (HTTP/1, no bytes written) or "stream reset error" (HTTP/2). Server logs provide the definitive diagnosis, including the audit ID and stack trace.

Next, the request parsing handler populates the request context with a request.Info struct. This struct holds parsed URL components: group, version, resource (e.g., batch/v1/jobs), namespace, and the logical verb (e.g., create, get, list, watch). Notably, get, list, and watch all map to the HTTP GET verb, while create maps to POST, update to PUT, and delete to DELETE. The resourceRequest field indicates whether it's a core resource operation or a meta-request (e.g., /metrics).

The latency tracking handler then annotates the audit event in the request context with detailed timing information, but only if the request exceeds a 500-millisecond threshold. These audit annotations are invaluable for debugging slow requests. Examples include authentication.k8s.io/latency, authorization.k8s.io/latency, mutating.k8s.io/latency, validating.k8s.io/latency, storage.k8s.io/latency (total time spent in etcd, potentially across multiple round trips), serialization.k8s.sio/latency, and writing.k8s.io/latency.

The deadline handler applies a deadline to all non-long-running requests. If the client doesn't specify a timeout, a default of 60 seconds is used. Cluster administrators can override this default via the API server's --request-timeout command-line option. The timeout handler then actively enforces this deadline. It executes the remainder of the handler chain in a separate goroutine and waits for its completion. If the deadline passes before the inner chain returns, the timeout handler aborts its wait, prepares a 504 Gateway Timeout response, and returns control to the HTTP server.

API Priority and Fairness (APF) is a critical self-defense mechanism. Requests entering the APF handler are first passed through a classifier to identify their flow (e.g., system-leader, catch-all). A shuffle sharding mechanism then assigns the request to a specific queue. The request waits in this queue, and an asynchronous scheduler decides whether to accept or reject it. Crucially, a request is only allowed to wait in the APF queue for a maximum of one-fourth of its total allotted deadline. If this sub-deadline is exceeded, the request is removed from the queue and rejected with a 429 "Too Many Requests" status code and a Retry-After header indicating how many seconds the client should wait before retrying.

Once past the handler chain, the request reaches a multiplexer which routes it to the correct API group and version-specific processing pipeline (e.g., apis/batch/v1/jobs). Within these pipelines, Kubernetes manages multiple API versions (e.g., v1, v1beta1) and an internal representation (int). The Scheme object is the central registry for Kubernetes' entire type system, mapping group versions and kinds to their corresponding Go types and storing all conversion functions. The Codec uses the Scheme to perform lossless transformations between different API versions and between various serialization formats (JSON, ProtoBuf).

The initial request body (e.g., JSON from kubectl) is decoded by the Codec into a batch/v1.Job Go object. This object then undergoes defaulting (e.g., setting unspecified fields to their default values). It's then converted to the internal representation, which is used for all internal API server logic. Further defaulting occurs via the resource's prepareForCreate() strategy, potentially influenced by feature gates.

The request then enters multiple admission control phases. This includes mutating webhooks (which can modify the object), CEL policies (an alpha feature for declarative validation), API-level Go code validation within the resource's strategy (e.g., validating object names or field values), and finally, validating webhooks (which cannot modify but can reject). Any failure in these phases results in a 422 Unprocessable Entity error.

With a fully validated internal object, the API server prepares for storage. The object is converted to the storage version (e.g., batch/v1 for jobs) which might be different from the request version. For native resources like Jobs, ProtoBuf is the default storage format in etcd, though some resources (like CRDs) are stored as JSON. The prepareObjectForStorage() function cleans up temporary fields, notably wiping the resourceVersion as etcd manages this. If KMS encryption is enabled, the object is encrypted. Finally, an optimistic put operation is performed against etcd. This means the API server optimistically assumes the operation will succeed. If etcd reports that the key already exists (e.g., trying to create a resource with an existing name), this is translated into a 409 Conflict error.

Upon successful storage, the internal object is converted back to the request's API version (e.g., batch/v1), encoded into the format specified by the client's Accept header (e.g., JSON), and written to the http.ResponseWriter. The request then traverses the handler chain in reverse. During this reverse pass, the audit handler persists the completed audit event, including all accumulated annotations, and the HTTP log handler flushes its entry. Finally, the HTTP server sends the full response back to kubectl. On the client side, client-go incorporates built-in retry logic, automatically retrying requests upon network errors, 429 "Too Many Requests" (especially with a Retry-After header), or 5xx server errors, enhancing the overall resilience of Kubernetes interactions.

Demo / Proof of Concept

▶ Watch: Deep dive into the API server's panic recovery (6:20)

The speakers effectively demonstrated the lifecycle concepts by showcasing how different API server behaviors manifest on both the kubectl client and the API server logs/audit events.

  1. Observing the Request: The talk began by showing the output of kubectl create -v9, which reveals the underlying HTTP POST request, including the URL (/apis/batch/v1/jobs), headers, and the resource kind (Job) versus resource plural (jobs). A successful request returns a 201 Created status.
  1. Panic Handling:
  • Client-side (HTTP/1): If a panic occurs before any response bytes are written, kubectl might report an "end of file" error. If some bytes were written before the panic, it could be an "unexpected end of file."
  • Client-side (HTTP/2): A panic typically results in a "stream reset error with internal error code."
  • Server-side: The API server's logs clearly show the panic, including the request path, audit ID, and a full stack trace, allowing pinpointing of the exact line of code causing the issue. The corresponding audit event would show a 500 Internal Server Error.
  1. Timeout Enforcement:
  • Client-specified Timeout: Using kubectl create --request-timeout=30s explicitly sets a 30-second deadline.
  • Forced Client Abort: A very short timeout, like kubectl create --request-timeout=6ms, causes kubectl to immediately abort with a "context deadline exceeded" error, as the client's deadline is shorter than the round trip time.
  • Server-side Timeout: When the API server's timeout handler enforces a deadline (e.g., the default 60 seconds or a custom value), the server logs show a 504 Gateway Timeout. The audit event also reflects the 504 status and indicates that the request timed out.
  1. APF Rejection:
  • Client-side: When APF rejects a request, kubectl receives a 429 Too Many Requests status, along with a Retry-After header (e.g., Retry-After: 1). kubectl's built-in retry logic (from client-go) then automatically retries the request after the specified delay.
  • Server-side: The API server's HTTP log entry shows the 429 status. Crucially, the audit event for such a rejected request includes an apf.k8s.io/request-timeout annotation. In the demonstration, this annotation showed approximately 15,000 milliseconds (15 seconds), which precisely corresponds to one-fourth of the default 60-second request timeout, confirming the APF's queue-waiting deadline enforcement. This specific detail provides strong evidence of APF's operation.

These demonstrations, combining client output with server-side logging and audit event analysis, provided concrete examples of how the theoretical handler chain and defense mechanisms translate into observable behavior, making the complex internals tangible.

Defensive Implications

▶ Watch: Client and server observe API server panic behavior (8:00)

Understanding the lifecycle of an API request offers several critical defensive implications for Kubernetes cluster operators and security professionals:

  1. Proactive Monitoring for API Server Health: Operators should monitor API server logs for specific error codes and messages that indicate internal issues.
  • Panics (500 Internal Server Error): Frequent 500s, especially those accompanied by stack traces in logs, point to critical bugs or instability within the API server or its extensions (e.g., admission webhooks). Immediate investigation is required.
  • Timeouts (504 Gateway Timeout): Persistent 504s suggest that requests are taking too long to process. This could indicate a bottleneck in the API server itself, slow etcd performance, or inefficient webhooks.
  • APF Rejections (429 Too Many Requests): A high volume of 429s indicates that the API server is under significant load and is actively shedding requests to protect itself. This signals a need to investigate client behavior (e.g., too many concurrent requests), scale API server resources, or adjust APF configuration. The Retry-After header is a direct instruction to clients, and non-client-go applications must respect it.
  1. Leverage Audit Logs for Performance Debugging: The detailed audit annotations, such as storage.k8s.io/latency, mutating.k8s.io/latency, and validating.k8s.io/latency, are invaluable for diagnosing performance issues. If users complain about slow API responses, these annotations can pinpoint the exact stage where latency is introduced—whether it's authentication, authorization, a specific webhook, or the etcd storage layer. Operators should configure robust audit logging and integrate it with observability platforms for analysis.
  1. Tune API Server Request Timeouts: The API server's --request-timeout flag allows administrators to set a global maximum duration for non-long-running requests. This prevents malicious or misbehaving clients from holding open connections indefinitely, consuming server resources. Careful tuning is necessary to balance responsiveness with the needs of legitimate long-running operations.
  1. Understand and Configure API Priority and Fairness (APF): APF is a vital self-defense mechanism. Administrators must understand how flows are classified, how shuffle sharding works, and how queueing parameters (like the 1/4 deadline rule) affect request processing under load. Proper configuration of APF can prevent API server outages during traffic spikes, ensuring that critical control plane operations (e.g., kubelet heartbeats) are prioritized over less urgent requests.
  1. Optimize Admission Webhooks and CEL Policies: Since webhooks and CEL policies are part of the critical request path and contribute to latency (tracked by mutating.k8s.io/latency and validating.k8s.io/latency), their performance is paramount. Defenders should ensure webhooks are highly performant, minimize external dependencies, and design CEL policies to be efficient. Slow webhooks can significantly degrade API server responsiveness and contribute to timeouts.
  1. Secure etcd and KMS Encryption: The talk highlights that the final object is stored in etcd, potentially after KMS encryption. This reinforces the importance of securing etcd itself (e.g., mTLS, network policies) and properly managing the KMS encryption keys. Compromise of etcd or the KMS provider could expose sensitive cluster data.
  1. Client-side Resilience: While the talk focuses on the server, the mention of client-go's retry logic is a reminder for developers to use official client libraries or implement robust retry mechanisms in their custom clients. This ensures applications can gracefully handle transient API server issues like network glitches, 429s, or temporary 5xx errors without failing immediately.

By internalizing these defensive implications, cluster operators can build more resilient, observable, and secure Kubernetes environments, capable of withstanding various operational challenges and attacks.

Key Takeaways

  • Kubernetes API Request Processing is a Multi-Stage Pipeline: Every API request, even a simple kubectl create, traverses a complex handler chain within the API server, involving parsing, defaulting, validation, and storage.
  • Robust Error Handling and Self-Defense Mechanisms are Built-in: The API server employs sophisticated features like panic recovery, enforced request deadlines (timeout handler), and API Priority and Fairness (APF) to ensure stability, prevent resource exhaustion, and provide graceful degradation under load.
  • Audit Logs and Annotations are Critical for Debugging: Detailed audit events, especially latency annotations (e.g., storage.k8s.io/latency, webhooks.k8s.io/latency), provide unparalleled visibility into where API requests spend their time, enabling precise diagnosis of performance bottlenecks.
  • API Object Versioning and Conversions are Fundamental: Kubernetes manages multiple API versions and an internal representation for each object, with the Scheme and Codec orchestrating lossless conversions throughout the request's lifecycle, ensuring data consistency across versions and storage formats.
  • Admission Control is a Powerful but Performance-Sensitive Layer: Mutating and validating webhooks, along with CEL policies, offer extensive control over resource creation and modification but must be implemented efficiently to avoid introducing significant latency and impacting API server performance.
  • Client-Side Resilience is Essential: The client-go library's built-in retry logic for network errors, 429s, and 5xx responses is crucial for building robust applications that interact with the Kubernetes API, providing automatic recovery from transient issues.

About the Speaker(s)

Abu Kashem is a Software Engineer at Red Hat, contributing to the Kubernetes ecosystem. His work history includes significant involvement with API servers, making him deeply familiar with the intricate processes and components discussed in the talk.

Stefan Schimanski is a Staff Software Engineer with extensive experience in Kubernetes. He has been actively working on the Kubernetes API server for approximately ten years, having directly contributed to and touched many of the components and functionalities that were detailed in this presentation. His long-standing involvement provides him with a profound understanding of the API server's architecture and evolution.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This is a meticulously detailed, highly technical deep dive into the Kubernetes API request lifecycle, offering a critical update to foundational knowledge. The speakers, clearly experts with direct contributions, dissect the handler chain, self-defense mechanisms like API Priority and Fairness (APF), and the nuanced object lifecycle, providing invaluable insights for debugging, performance optimization, and security. It's precisely the kind of substantive content that delivers actionable intelligence, demonstrating a profound understanding of the system's internals.

Heather Calloway (CISO) — STRONG ACCEPT

This is an exceptionally thorough technical deep dive into the Kubernetes API server's request lifecycle, offering essential insights for cluster operators and security engineers. It moves beyond theoretical explanations by clearly demonstrating how critical self-defense mechanisms like API Priority and Fairness (APF), request timeouts, and robust error handling function in practice. While primarily technical, the talk's explicit focus on defensive implications—from monitoring API server health and leveraging audit logs for performance debugging to securing etcd and optimizing admission webhooks—makes it highly relevant for security leaders seeking to enhance their organization's…

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025