Wait! Can Your Pod Survive a Restart? - Aya Ozawa, CloudNatix Inc.

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

In the dynamic world of cloud-native applications, Kubernetes orchestrates workloads with powerful features like self-healing, rolling upgrades, and auto-scaling. Central to these capabilities is the frequent restarting of pods, a fundamental aspect of Kubernetes' "cattle not pets" philosophy. However, as Aya Ozawa from CloudNatix Inc. highlighted in her KubeCon EU talk, "Wait! Can Your Pod Survive a Restart?", correctly handling these restarts is far from trivial. Many applications, if not properly designed and configured, can experience unexpected downtime, data inconsistencies, and resource waste during these seemingly routine events.

Watch on YouTube

Visual summary for Wait! Can Your Pod Survive a Restart? - Aya Ozawa, CloudNatix Inc.
Visual summary for Wait! Can Your Pod Survive a Restart? - Aya Ozawa, CloudNatix Inc.

Key moments

  1. 0:00 Why pod restarts are crucial in Kubernetes
  2. 2:50 Kubernetes pod termination: SIGTERM and SIGKILL
  3. 4:20 Implementing graceful shutdown using Dockerfile and preStop hooks
  4. 6:00 How sidecar containers gracefully shut down sequentially
  5. 8:00 Situations where pod termination grace periods are ignored
  6. 9:00 Ensuring zero downtime for HTTP servers during restarts
  7. 10:00 Delaying shutdown to prevent request drops and inconsistencies

Wait! Can Your Pod Survive a Restart?

Speakers: Aya Ozawa, Member of Technical Staff, CloudNatix Inc.

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=eO8szEGNwoo

Overview

In the dynamic world of cloud-native applications, Kubernetes orchestrates workloads with powerful features like self-healing, rolling upgrades, and auto-scaling. Central to these capabilities is the frequent restarting of pods, a fundamental aspect of Kubernetes' "cattle not pets" philosophy. However, as Aya Ozawa from CloudNatix Inc. highlighted in her KubeCon EU talk, "Wait! Can Your Pod Survive a Restart?", correctly handling these restarts is far from trivial. Many applications, if not properly designed and configured, can experience unexpected downtime, data inconsistencies, and resource waste during these seemingly routine events.

Ozawa's presentation delved into the intricacies of pod termination and startup, exploring various scenarios from single container restarts to complex networked applications and distributed controllers. She systematically broke down the challenges posed by ungraceful shutdowns, premature traffic routing, and leader election disruptions. The talk provided actionable insights and best practices, demonstrating how to leverage Kubernetes features like preStop hooks, readiness and startup probes, and Pod Disruption Budgets (PDBs) to minimize the impact of restarts and ensure application resilience.

This article aims to provide a comprehensive technical deep dive into Ozawa's talk, dissecting the mechanisms of graceful termination, the nuances of traffic management during pod lifecycle events, and strategies for maintaining high availability in stateful and leader-elected workloads. It serves as a guide for developers and operators seeking to build more robust and restart-friendly applications on Kubernetes, enabling them to fully harness the automation and resilience benefits of the platform.

Background

▶ Watch: Why pod restarts are crucial in Kubernetes (0:00)

Kubernetes fundamentally treats application instances as cattle rather than pets. This analogy implies that individual instances (pods) are ephemeral, interchangeable, and can be replaced without significant manual intervention. This paradigm underpins many of Kubernetes' most powerful features, including self-healing (restarting failed pods), rolling upgrades (gradually replacing old versions with new ones), and auto-scaling (adding or removing pods based on demand). For applications to fully benefit from these capabilities, they must be designed to be restartable – meaning they can gracefully shut down and start up without data loss or service disruption.

The problem arises because application developers often assume a more stable, pet-like environment for their processes. Traditional applications might not be built to handle sudden termination signals or to quickly re-establish state after a restart. In a Kubernetes environment, restarts can occur due to various reasons:

  • Container-level restarts: A container might exit unexpectedly (e.g., due to a bug), or its liveness probe might fail, prompting the kubelet to restart it.
  • ReplicaSet-level recreation: If a pod is manually deleted or fails beyond recovery, the ReplicaSet controller will automatically recreate it to maintain the desired replica count.
  • Deployment-level rolling upgrades: During a deployment update, new pods are created, and old pods are gradually terminated. Unlike the previous scenarios, new pods are typically running before old ones are fully removed, aiming for zero downtime.
  • Node maintenance: Operations like kubectl drain or node upgrades require pods to be evicted and restarted on other nodes.
  • Resource pressure: The kubelet might evict pods when a node runs low on resources (e.g., memory, disk pressure).

Without proper handling, these restarts can lead to several issues:

  • Request drops: In-flight requests might be terminated prematurely, leading to service degradation or errors for users.
  • Data inconsistency: Applications might not have enough time to flush in-memory data to persistent storage, resulting in data loss or corrupted state.
  • Resource waste: Pods that ignore termination signals can continue consuming resources until forcibly killed, prolonging the shutdown process.
  • Slow startup times: Applications might take too long to become ready to serve traffic, creating availability gaps.
  • Split-brain scenarios: In distributed systems with leader election, improper shutdown can lead to multiple leaders, causing conflicts and data corruption.

Addressing these challenges requires a deep understanding of Kubernetes' termination lifecycle, robust application design for signal handling, and strategic use of Kubernetes-native features to manage traffic flow and resource availability during restarts.

Key Findings

▶ Watch: Implementing graceful shutdown using Dockerfile and preStop hooks (4:20)

Aya Ozawa's talk presented several key findings and best practices for building restart-resilient applications on Kubernetes:

  1. Graceful Termination is Paramount: Applications must be designed to handle SIGTERM (or custom termination signals) to perform graceful shutdowns, including closing connections, flushing data, and releasing resources, before the termination grace period expires.
  2. Kubernetes Features for Custom Signals: For applications that expect different shutdown signals (e.g., NGINX expects SIGQUIT), Kubernetes offers mechanisms like the Dockerfile STOP_SIGNAL instruction or the preStop lifecycle hook to send the correct signal without modifying application code. The upcoming lifecycle.terminal.signal feature will further simplify this.
  3. Sidecar Termination Order: Sidecar containers are terminated sequentially in the reverse order of their startup, but they share the same pod termination grace period as primary containers, requiring careful planning for complex multi-container pods.
  4. Termination Grace Period Isn't Always Guaranteed: While Kubernetes aims for graceful termination, events like application crashes, OOMKilled processes, or aggressive eviction APIs (e.g., kubectl delete pod --grace-period=1) can bypass the grace period, necessitating resilient application design.
  5. Traffic Management for Zero Downtime: For networked applications, combining preStop hooks with a short delay (e.g., sleep 3) and robust readiness and startup probes is crucial to prevent request drops during traffic shifting between old and new pods.
  6. Startup Probes for Initialization: Utilizing a dedicated startupProbe allows for a more lenient initial readiness check, preventing premature livenessProbe failures during application bootstrapping, while readinessProbe ensures continuous service availability post-startup.
  7. Optimizing Leader Election Transitions: For controllers using leader election, the leaderElectionReleaseOnCancel option in libraries like controller-runtime can significantly reduce leadership takeover time during graceful shutdowns by proactively releasing the lock, minimizing disruption while mitigating split-brain risks.
  8. Pod Disruption Budgets (PDBs) for Voluntary Disruptions: PDBs effectively limit the number of unavailable pods during voluntary disruptions (e.g., kubectl drain), ensuring a minimum level of service availability for critical workloads. However, they have limitations, especially with single-replica deployments or stringent maxUnavailable settings.

Technical Deep Dive

▶ Watch: How sidecar containers gracefully shut down sequentially (6:00)

The talk provided a comprehensive technical exploration of pod restart mechanisms and mitigation strategies, beginning with the fundamental aspects of container termination and progressing to complex networked applications and distributed controllers.

Container Termination Basics

When a pod is deleted or needs to be restarted, Kubernetes initiates a termination process. Initially, a SIGTERM signal is sent to the primary process (PID 1) within each container. This signal indicates that the application should begin a graceful shutdown. Kubernetes then waits for a configurable termination grace period, which defaults to 30 seconds. During this period, the application is expected to:

  • Stop accepting new requests.
  • Complete any in-flight requests.
  • Flush buffered data to persistent storage or external services.
  • Release resources (e.g., close database connections, file handles).

If the application does not exit within the grace period, Kubernetes sends a SIGKILL signal, which forcibly terminates the process immediately without any opportunity for cleanup. This can lead to request drops, data corruption, or inconsistent states. A common pitfall is that many applications, especially those not designed for cloud-native environments, might ignore SIGTERM when running as PID 1 in a container. This is because, unlike typical Linux processes, PID 1 in a container often doesn't have a parent process to forward signals, or the application itself might not have a signal handler for SIGTERM. Consequently, the process continues running until SIGKILL, wasting resources and potentially causing issues.

Handling Custom Termination Signals

Some applications expect specific signals for graceful shutdown. For instance, NGINX typically uses SIGQUIT. To accommodate such applications without modifying their source code, Kubernetes offers two primary solutions:

  1. Dockerfile STOP_SIGNAL: This instruction within the Dockerfile allows specifying which signal should be sent to the container's main process upon termination. For example, STOP_SIGNAL SIGQUIT would ensure NGINX receives SIGQUIT instead of SIGTERM. This requires rebuilding the container image.
  2. Kubernetes preStop Hook: This lifecycle hook executes a command or an HTTP request just before the container's stop signal is sent. It's defined directly in the pod manifest.
  • exec hook: Can be used to send a custom signal (e.g., kill -s SIGQUIT 1) or run a shutdown script. This requires kill binaries to be present in the container image.
  • httpGet hook: Useful if the application exposes an HTTP endpoint for graceful shutdown.
  • Future Feature: lifecycle.terminal.signal: Ozawa mentioned an upcoming feature (not yet generally available) that will allow specifying a custom stop signal directly in the pod manifest, similar to Dockerfile STOP_SIGNAL but without requiring image rebuilds.

Sidecar Container Termination

With Kubernetes 1.29, sidecar containers became a native feature, building upon the initContainers pattern. When a pod with sidecars terminates, the process differs slightly:

  • All primary containers receive SIGTERM simultaneously.
  • Once all primary containers have exited, sidecars receive SIGTERM one by one, in the reverse order of how they started.
  • Crucially, the entire pod, including all primary and sidecar containers, shares the same termination grace period. This means sidecars must complete their shutdown tasks within the remaining grace period after primary containers have exited, which requires careful resource allocation and shutdown logic design.

When Grace Period is Not Guaranteed

While Kubernetes strives for graceful termination, there are scenarios where the termination grace period is bypassed or ignored:

  • Application Crashes or OOMKilled: If an application crashes due to a bug or is terminated by the Linux Out Of Memory (OOM) killer, there's no opportunity for graceful shutdown.
  • Aggressive Eviction/Deletion APIs: Using kubectl delete pod --grace-period=1 will immediately terminate the pod with SIGKILL, overriding any configured grace period. Similarly, kubectl eviction APIs can be configured to use shorter grace periods.
  • Hard Eviction: When a node is under severe resource pressure, the kubelet can perform hard eviction, immediately terminating pods without any grace period to reclaim resources. This is more aggressive than soft eviction, which respects configured eviction thresholds and grace periods.

Networked Applications and Zero Downtime

For HTTP servers and other networked applications, graceful shutdown is more complex due to traffic routing. When a pod terminates, Kubernetes sends SIGTERM, but network routing (e.g., IP tables managed by kube-proxy) might still direct new requests to the terminating pod. This can lead to request drops.

To mitigate this, a delay must be introduced between receiving SIGTERM and actually stopping the application from listening for new connections. This delay allows Kubernetes to update its routing tables and remove the terminating pod from the service's endpoints.

  • Application-level delay: Modifying application code to sleep for a few seconds after receiving SIGTERM before stopping the listener.
  • preStop with sleep: The recommended approach, available since Kubernetes 1.3, is to use a preStop hook with a sleep command (e.g., sleep 3). This delays the SIGTERM delivery to the application, giving Kubernetes time to update routing.

Another challenge is ensuring new pods are ready to serve traffic immediately upon startup. If Kubernetes routes traffic to a new pod before it's fully initialized, requests will be dropped.

  • readinessProbe: This probe checks if an application is ready to handle requests. Kubernetes will only route traffic to a pod once its readinessProbe passes. However, using only a readinessProbe can be problematic because it runs throughout the container's lifecycle. A lenient check suitable for startup might be too loose for ongoing health, while a strict check might cause slow startup or false negatives during initialization.
  • startupProbe: Introduced to address the limitations of readinessProbe during initialization, the startupProbe runs only during the container's startup phase. Once it succeeds, the livenessProbe and readinessProbe take over. This allows for separate, more forgiving thresholds during startup without affecting the stricter, continuous health checks.

Traffic Shifting Behavior

The exact behavior of traffic shifting during restarts depends on the number of replicas and the restart type:

  • Multi-replica scenarios: In deployments with multiple replicas, Kubernetes removes the terminating pod's endpoint almost immediately after it's marked for deletion. New pods are added to the endpoint list once their readinessProbe passes. While there's a brief gap for the individual pod, other healthy replicas continue to handle traffic, ensuring overall service availability.
  • Rolling upgrades: During a rolling upgrade, new pods are created and marked as ready before old pods are terminated. Traffic is seamlessly switched to the new, ready pods, minimizing disruption regardless of replica count.
  • Single-replica workloads: This is the most challenging scenario. If there are no other ready endpoints, traffic will continue to be routed to the terminating pod even after it starts shutting down. The preStop sleep is critical here to ensure the old pod continues handling requests until the new pod is fully ready and traffic can switch. Without it, even with a readinessProbe, there's a high risk of request drops.
  • Liveness Probe Failures (Single Replica): If a single-replica pod fails its livenessProbe, it's terminated and restarted. During the time the old pod is shutting down and the new pod is starting up, there's an unavoidable gap where no pod is available to handle traffic, leading to request drops. For critical applications, multiple replicas are always recommended to avoid this inherent downtime.

Controller Leader Election

For distributed controllers (e.g., Kubernetes operators), leader election is crucial to prevent conflicting operations. The controller-runtime library, a common choice for building controllers, uses a shared lease resource for leader election. Candidates continuously try to update this resource, and the one that successfully updates it becomes the leader.

Key parameters affecting leader transitions:

  • leaseDuration: Maximum duration a leader can hold the lease without renewal (default: 15 seconds).
  • retryPeriod: How frequently candidates attempt to acquire leadership (default: 2 seconds).

With default values, a leadership takeover can take up to leaseDuration + retryPeriod (17 seconds in the worst case), leading to significant disruption. Shortening leaseDuration too much can increase the risk of split-brain, where multiple controllers mistakenly believe they are the leader.

To safely reduce takeover time, controller-runtime offers leaderElectionReleaseOnCancel. When enabled, the current leader proactively updates the lease duration to 1 second during graceful shutdown, effectively releasing the leadership quickly. This allows a new leader to be elected much faster (e.g., within 3 seconds) without risking split-brain during normal operation.

Pod Disruption Budgets (PDBs)

Pod Disruption Budgets (PDBs) are a Kubernetes API object that limits the number of pods of a replicated application that can be voluntarily disrupted at one time. They are essential for minimizing downtime during cluster maintenance operations like kubectl drain. A PDB defines either minAvailable (minimum number of available pods) or maxUnavailable (maximum number of unavailable pods).

When a PDB is in place:

  • Kubernetes will reject eviction requests (e.g., from kubectl drain) if they would violate the budget.
  • This ensures that a specified number or percentage of pods remain running during voluntary disruptions.

Limitations of PDBs:

  • Requires at least two replicas: If maxUnavailable is 1 and you only have a single replica, kubectl drain will still be blocked until a second replica is available. PDBs don't create replacement pods ahead of time in eviction scenarios (unlike rolling upgrades).
  • Can block node maintenance: A PDB with maxUnavailable: 0 or minAvailable: 100% can permanently block eviction if the application has only one replica or if a pod is unhealthy and never reaches a running state.
  • Unhealthy Pod Eviction Policy: Since Kubernetes 1.27, the default unhealthyPodEvictionPolicy is AlwaysAllow. This allows unhealthy pods to be evicted even if it violates the PDB, preventing them from blocking node maintenance indefinitely.
  • Voluntary disruptions only: PDBs only apply to the eviction API and descheduler preemption. They do not apply to kubectl delete pod or involuntary disruptions like node failures.
  • Preemption exceptions: While the scheduler generally respects PDBs during preemption (evicting lower-priority pods to make room for higher-priority ones), it might ignore PDBs if no other placement options match the constraints.

Demo / Proof of Concept

▶ Watch: Ensuring zero downtime for HTTP servers during restarts (9:00)

Aya Ozawa presented a clear demonstration of the issues caused by ungraceful shutdowns and how to mitigate them. The demo focused on an HTTP server application deployed with a Kubernetes service, illustrating two scenarios: a "failing case" and a "corrected version."

Failing Case

The initial setup involved a simple Deployment and Service without any special preStop hooks, readiness probes, or startup probes. When the speaker manually deleted a pod:

  1. The server immediately received a SIGTERM signal and began its shutdown logic, stopping its listener.
  2. However, the IP tables (Kubernetes' network routing mechanism) had not yet been updated. Traffic was still being routed to the terminating pod.
  3. As a result, new incoming requests to the old pod were dropped, leading to service disruption.
  4. Meanwhile, Kubernetes started a new pod. Once the new pod was running, traffic switched to it.
  5. But because the new server wasn't fully ready to handle requests immediately upon startup (lacking a readinessProbe), some initial requests to the new pod were also dropped.

The logs clearly showed a period of request failures and errors, demonstrating the impact of an unprepared application on service availability during a restart.

Corrected Version

In the corrected version, the Deployment manifest was updated to include:

  1. A preStop hook with a sleep 3 command: This delayed the delivery of the SIGTERM signal to the application by 3 seconds.
  2. A readinessProbe: Configured to check if the application was truly ready to handle requests before Kubernetes routed traffic to it.
  3. A startupProbe: Configured to handle the initial, potentially longer, startup phase.

When the pod was deleted in this scenario:

  1. Thanks to the preStop sleep 3, Kubernetes waited 3 seconds before sending SIGTERM to the old pod.
  2. During this delay, the old pod continued to handle requests, while a new pod was simultaneously created.
  3. The readinessProbe on the new pod ensured that traffic was only routed to it once it was fully initialized and ready.
  4. The IP tables were updated smoothly, redirecting traffic to the new, ready pod.
  5. Finally, the old pod received SIGTERM and exited gracefully.

The demo logs showed a seamless traffic transition with zero request drops, illustrating the effectiveness of combining preStop delays with robust readiness and startup probes in minimizing downtime during pod restarts.

Defensive Implications

▶ Watch: Delaying shutdown to prevent request drops and inconsistencies (10:00)

The insights from this talk offer several critical defensive implications for engineers and operators working with Kubernetes:

  1. Application Design for Graceful Shutdown:
  • Signal Handling: Ensure your application explicitly handles SIGTERM (or other expected termination signals like SIGQUIT) and initiates a graceful shutdown process. This includes stopping new request acceptance, completing in-flight work, flushing buffers, and releasing resources.
  • Idempotency: Design operations to be idempotent where possible, so that if a request is retried due to a pod restart, it doesn't cause adverse side effects.
  • Fast Shutdown: Optimize shutdown logic to complete within the Kubernetes termination grace period (default 30 seconds) to avoid forceful SIGKILL.
  1. Kubernetes Configuration Best Practices:
  • preStop Hooks for Networked Apps: For HTTP servers and other networked services, implement a preStop hook with a sleep command (e.g., sleep 3) to allow Kubernetes enough time to update routing tables before the application stops listening.
  • readinessProbe and startupProbe:
  • Always configure a readinessProbe to ensure traffic is only routed to pods that are truly ready to serve requests.
  • Utilize a startupProbe for applications with longer initialization times, allowing a more lenient check during startup without compromising the strictness of the readinessProbe for ongoing health.
  • Custom Stop Signals: If an application expects a signal other than SIGTERM, use Dockerfile STOP_SIGNAL or an exec preStop hook to send the correct signal.
  • Sidecar Grace Period: When using sidecars, meticulously calculate the combined shutdown time for all containers to ensure they can complete within the shared pod termination grace period.
  1. Replication Strategies:
  • Multiple Replicas for Critical Workloads: For any application requiring high availability and zero downtime, deploy with at least two replicas. This allows other replicas to handle traffic during individual pod restarts, especially for scenarios like liveness probe failures where an unavoidable gap in service can occur for single-replica deployments.
  • PDBs for Voluntary Disruptions: Implement Pod Disruption Budgets (PDBs) for all critical replicated workloads to limit voluntary disruptions during node maintenance, ensuring a minimum number of pods remain available. Be cautious with maxUnavailable: 0 or single-replica PDBs, as they can block maintenance. Leverage unhealthyPodEvictionPolicy: AlwaysAllow (default since 1.27) to prevent unhealthy pods from blocking evictions.
  1. Controller Optimization:
  • Leader Election Configuration: For distributed controllers using leader election, enable leaderElectionReleaseOnCancel (if using libraries like controller-runtime) to significantly reduce leadership takeover time during graceful shutdowns, minimizing disruption while maintaining safety against split-brain.
  1. Monitoring and Alerting:
  • Monitor application logs for graceful shutdown messages and Kubernetes events for SIGKILL signals to identify applications that are not terminating cleanly.
  • Track request drop rates and service availability during restarts to validate the effectiveness of implemented strategies.

By proactively addressing these defensive implications, organizations can build a more resilient and robust Kubernetes environment, minimizing the impact of restarts and maximizing the benefits of cloud-native automation.

Key Takeaways

  • Implement Graceful Shutdown: Applications must handle SIGTERM (or custom signals via Dockerfile STOP_SIGNAL or preStop hooks) to perform cleanup and exit gracefully within the termination grace period, preventing data loss and resource waste.
  • Manage Network Traffic During Restarts: For networked applications, combine a preStop hook with a short sleep (e.g., sleep 3) and robust readiness and startup probes to ensure smooth traffic transitions and zero request drops during pod restarts.
  • Optimize Controller Leader Election: For distributed controllers, leverage leaderElectionReleaseOnCancel (e.g., in controller-runtime) to accelerate leadership transitions during graceful shutdowns, reducing disruption while safely preventing split-brain scenarios.
  • Utilize Pod Disruption Budgets (PDBs): Deploy PDBs for critical, replicated workloads to limit voluntary disruptions during cluster maintenance, ensuring a minimum level of service availability. Be mindful of their limitations with single-replica deployments or strict unavailability settings.
  • Prioritize Multi-Replica Deployments: For critical applications requiring high availability, always deploy with multiple replicas. This strategy is essential to avoid unavoidable downtime during scenarios like liveness probe failures in single-replica setups.
  • Understand Sidecar Termination: Sidecars terminate sequentially in reverse order, but they share the pod's overall termination grace period, requiring careful coordination of shutdown logic.

About the Speaker(s)

Aya Ozawa is a Member of Technical Staff at CloudNatix Inc. In her role, she is actively involved in developing L Marina, an open-source AI platform designed for Kubernetes. Her work focuses on understanding and improving the resilience and operational aspects of applications running on Kubernetes, drawing from years of experience in the field. Her expertise lies in ensuring applications can effectively leverage Kubernetes' powerful automation features, particularly concerning pod lifecycle management and restartability.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This KubeCon talk by Aya Ozawa cuts through the usual fluff to deliver a brutally honest and deeply technical assessment of pod restart resilience in Kubernetes. It meticulously breaks down the nuances of graceful termination, traffic management, and leader election, providing actionable strategies that every K8s operator and developer needs to implement. No marketing, no hand-waving, just solid engineering advice to prevent your applications from falling over during routine restarts.

Heather Calloway (CISO) — STRONG ACCEPT

This KubeCon talk by Aya Ozawa provides a critical, technically grounded examination of Kubernetes pod restarts and their profound impact on application resilience. It meticulously details the mechanisms and pitfalls of pod termination, offering concrete, actionable strategies—from graceful signal handling to advanced probe configurations and Pod Disruption Budgets—to ensure operational stability and zero downtime. This is not just technical arcana; it's a direct guide for ensuring the availability and data integrity of our critical cloud-native applications, which translates directly to business continuity and reduced risk.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025