Taming the Beast: Advanced Resource Management With Kubernetes - Lucy Sweet & Dawn Chen
Lucy Sweet, Dawn Chen
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In this compelling KubeCon EU talk, Lucy Sweet from Uber and Dawn Chen, Tech Lead of Kubernetes SIG Node at Google, delved into the evolving landscape of resource management within Kubernetes. The session, aptly titled "Taming the Beast," addressed the persistent challenges faced by users attempting to run advanced, stateful, and bursty workloads on Kubernetes, moving beyond the traditional stateless application model. Lucy, representing the user perspective, articulated common frustrations with current resource allocation mechanisms, while Dawn, from the core development team, unveiled a suite of recent and upcoming features designed to provide finer-grained control and greater flexibility.

Key moments
- 0:00 Introduction: The challenge of Kubernetes resource management
- 2:00 Struggles with rigid per-container resource limits
- 2:40 Solution: Introducing Pod Level Resources feature
- 3:20 Simplifying configuration with Pod Level Resources YAML
- 4:10 Demonstration of applying pod-level resource limits
- 6:00 Combining pod and container resource allocation flexibility
- 7:00 Advanced use cases: ML workloads and web services
Taming the Beast: Advanced Resource Management With Kubernetes
Speakers: Lucy Sweet, Engineer, Uber; Dawn Chen, Software Engineer, Google, Tech Lead of Kubernetes SIG Node
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=AYxjk8ZZclo
Overview
In this compelling KubeCon EU talk, Lucy Sweet from Uber and Dawn Chen, Tech Lead of Kubernetes SIG Node at Google, delved into the evolving landscape of resource management within Kubernetes. The session, aptly titled "Taming the Beast," addressed the persistent challenges faced by users attempting to run advanced, stateful, and bursty workloads on Kubernetes, moving beyond the traditional stateless application model. Lucy, representing the user perspective, articulated common frustrations with current resource allocation mechanisms, while Dawn, from the core development team, unveiled a suite of recent and upcoming features designed to provide finer-grained control and greater flexibility.
The talk highlighted three pivotal advancements: Pod-level resource requests and limits, in-place pod resizing, and node swap support. These features collectively aim to significantly simplify configuration, optimize resource utilization, and minimize disruption for critical applications. By enabling more dynamic and adaptable resource management, Kubernetes is poised to better support demanding workloads such as machine learning training jobs, databases, and development environments, which often exhibit unpredictable resource consumption patterns.
This presentation is crucial for anyone operating or planning to deploy complex applications on Kubernetes, offering practical solutions to long-standing resource management hurdles. It not only showcases the latest capabilities but also provides a glimpse into the future direction of Kubernetes' resource management efforts, signaling a commitment to making the platform more robust and versatile for an ever-broadening array of use cases.
Background
▶ Watch: Introduction: The challenge of Kubernetes resource management (0:00)
Kubernetes, since its inception, has excelled at managing stateless, containerized applications. Its declarative API and powerful orchestration capabilities have made it the de facto standard for deploying microservices. However, as organizations push the boundaries of what they run on Kubernetes, fundamental limitations in its original resource management model have become apparent, particularly for workloads that are stateful, long-running, or exhibit highly variable resource demands.
Historically, Kubernetes has enforced resource requests and limits at the container level. While straightforward for simple applications, this approach introduces significant friction for more complex scenarios. For instance, sidecar containers, common in logging, monitoring, or service mesh patterns, often require minimal dedicated resources but are forced to reserve an entire slice of CPU or memory, leading to over-provisioning. Conversely, applications with multiple containers that have inverse peak usage patterns (e.g., one container is busy while another is idle, and vice-versa) often require reserving peak resources for both containers simultaneously, resulting in substantial resource waste. This rigid, per-container model also makes it difficult to manage shared resource pools within a single pod.
Another significant challenge has been the immutability of pod resources. Once a pod is deployed with specific CPU and memory requests and limits, these cannot be altered without recreating the pod. For stateless applications, this disruption is often tolerable. However, for stateful workloads like databases, message queues, or in-memory caches, recreating a pod to adjust resources can lead to service interruptions, data corruption, or lengthy recovery times, imposing considerable operational burden. This immutability also hinders the efficient management of applications with bursty startup phases, such as Java services that warm caches or pre-process data, requiring high memory usage initially, which then drops off. Without the ability to dynamically shrink memory or utilize swap space, these applications tie up significant amounts of idle RAM throughout their lifecycle.
Finally, Kubernetes has traditionally maintained an "allergy" to swap memory, with kubelet (the agent running on each node) designed to crash if swap is enabled on the host. This stance was rooted in concerns about performance unpredictability and monitoring complexities associated with swap. However, for certain workloads, particularly large-scale batch processing, AI/ML training, or developer environments with varying resource needs, the inability to use swap leads to either excessive memory provisioning or out-of-memory (OOM) errors, both of which are costly and inefficient. These limitations have driven the Kubernetes SIG Node community to innovate, leading to the development of the advanced resource management features discussed in this talk.
Key Findings
▶ Watch: Solution: Introducing Pod Level Resources feature (2:40)
The talk presented three major advancements in Kubernetes resource management, each addressing a critical pain point for advanced workloads:
- Pod-Level Resource Requests and Limits: This new feature, currently in Alpha (introduced in Kubernetes v1.32/v1.33 and requiring a feature gate), allows users to define resource requests and limits at the pod level, rather than solely at the individual container level. This enables better resource sharing among containers within a pod, reducing over-provisioning for sidecars or applications with inverse peak resource demands. It simplifies configurations and acts as a "guardrail" for the pod as a whole, while allowing internal containers to share the allocated resources dynamically. The flexibility to mix and match pod-level and container-level resource definitions provides granular control where needed, while simplifying the overall management.
- In-Place Pod Resizing: Promoted to Beta in Kubernetes v1.33 and now enabled by default, this feature allows operators to dynamically adjust the CPU and memory requests and limits of a running pod without requiring its termination and recreation. This is a game-changer for stateful workloads, where disruption from pod restarts is highly undesirable. The
kubeletcan adjust resources if the node has sufficient slack, enabling seamless scaling up or down of applications in response to changing traffic or processing needs, with zero disruption to users. It introducesresizePolicyoptions likePreferNoRestart(default for CPU) andRestartContainerfor scenarios where an application restart is necessary (e.g., for memory shrinking or specific application reconfigurations).
- Node Swap Support: While still not yet General Availability (GA) and promoted to Beta 3 in Kubernetes v1.33 with ongoing enhancements, this feature introduces support for swap memory on Kubernetes nodes. Historically,
kubeletwould crash if swap was enabled, making it unusable. The new implementation allows cluster administrators to configurekubeletwith specificswapBehaviorsettings, such asLimitedSwap(where only the "burstable" component of a pod's memory, i.e., between request and limit, can use swap) orUnlimitedSwap(allowing all pod memory to use swap). This addresses the issue of wasted idle memory for applications with bursty startup phases (e.g., Java applications), providing a mechanism to temporarily offload less active memory to disk, thereby improving overall node utilization and preventing unnecessary OOM kills, particularly for large AI/ML training jobs.
Technical Deep Dive
▶ Watch: Simplifying configuration with Pod Level Resources YAML (3:20)
These new features represent significant architectural shifts aimed at making Kubernetes more adaptable to diverse and demanding workloads.
Pod-Level Resource Requests and Limits
Traditionally, Kubernetes YAML configurations specify resources (requests and limits for cpu and memory) within each container definition:
With Pod-level resource requests and limits, the resources block can now be defined at the pod specification level, above the individual containers. This allows all containers within that pod to share the specified resources:
The system dynamically allocates these shared resources among the containers. A key flexibility highlighted is the ability to mix and match pod-level and container-level resource definitions. For instance, a common pattern might involve setting a pod-level boundary for overall resource consumption, but then providing a specific, higher limit for a "bursty" pre-processing container within that pod, ensuring it gets sufficient resources during its initial phase without over-provisioning the entire pod for its duration. The pod-level resources act as an upper bound or a guaranteed minimum for the aggregate usage of all containers within the pod. This feature, still in Alpha in Kubernetes v1.32/v1.33, requires enabling a feature gate to be used.
In-Place Pod Resizing
This feature fundamentally alters the lifecycle of a pod. Previously, modifying a pod's CPU or memory requests or limits would trigger a full pod termination and recreation by the controller. With in-place pod resizing, the kubelet on the node directly interacts with the underlying cgroups to adjust resource allocations for the running containers.
When an operator or an automated system (like a Vertical Pod Autoscaler, VPA) updates the resource specification of a running pod, the Kubernetes control plane sends this update to the kubelet managing that pod. The kubelet then:
- Checks Node Slack: Determines if the node has enough available CPU and memory to accommodate the increase in resources.
- Applies Changes to Cgroups: If sufficient resources are available, the
kubeletdirectly modifies the cgroup settings for the pod's containers, updating their CPU shares, CPU quotas, and memory limits. This happens without terminating the container processes. - Updates Pod Status: The
kubeletreports the updated resource status back to the API server.
This process ensures zero disruption for the application. However, there are important nuances:
resizePolicy: Kubernetes introduces aresizePolicyfield within the container specification. The two supported options are:PreferNoRestart: This is the default for CPU changes and the preferred behavior. It attempts to adjust resources without restarting the container.RestartContainer: This policy dictates that the container must be restarted for resource changes to take effect. This is particularly relevant for memory shrinking, which is generally not supported withPreferNoRestartdue to Linux kernel limitations (specifically withcgroup v2). The kernel often prefers to kill processes rather than dynamically shrinking their memory allocations, especially to maintain predictable performance.- Limitations:
- Currently, only CPU and memory are supported for in-place resizing. Other extended resources are under development.
- Memory shrinking with
PreferNoRestartis generally ineffective due to kernel behavior. - Quality of Service (QoS) classes (Guaranteed, Burstable, BestEffort) cannot be changed via in-place resizing. A Guaranteed pod will remain Guaranteed, even if its resources are adjusted.
- Atomic Resizing: The system implements atomic resizing. If an adjustment request cannot be fully satisfied (e.g., due to insufficient node resources for a portion of the request), the entire request is rejected to prevent partial, unpredictable state changes.
- While the feature is enabled by default in Kubernetes v1.33, full automation often requires integration with higher-level orchestrators like VPA or workload-specific autoscalers (e.g., for Ray or Slurm clusters), which are actively being worked on by the community.
Node Swap Support
The integration of swap memory into Kubernetes required a careful reconsideration of fundamental design principles. kubelet's historical aversion to swap stemmed from concerns about performance degradation, "noisy neighbor" effects from disk I/O, and the kernel's difficulty in accurately identifying "inactive" memory pages to swap out.
The new node swap support allows cluster administrators to enable and configure swap on nodes via kubelet configuration. This is a node-level feature, meaning kubelet manages the swap behavior for pods on that specific node.
KubeletConfiguration: Administrators configurekubeletwith thenodeSwapfeature gate enabled and specify aswapBehavior.swapBehaviorOptions:NoSwap(Default):kubeletcontinues to treat swap as disallowed.LimitedSwap: This is the recommended and most common setting. In this mode, only the burstable component of a pod's memory can be swapped. The burstable component refers to the memory usage above a pod'srequestbut within itslimit. Memory at or below therequestis treated as "guaranteed" and will not be swapped. This allows applications to burst memory usage during peak periods (like startup) and have the excess swapped out without impacting the performance of their guaranteed memory.UnlimitedSwap: Allows all of a pod's memory (up to its limit) to be eligible for swap. This is generally reserved for specific, highly memory-intensive batch workloads where performance predictability is less critical than avoiding OOM kills.
kubelet monitors both raw memory and swap usage to make eviction decisions. While this feature is a significant step forward, it remains in Beta 3 due to ongoing work on security concerns (e.g., potential for sensitive data on disk) and complex interactions with eviction policies. The community is actively developing mitigation strategies and best practices to ensure safe and performant swap usage.
Demo / Proof of Concept
▶ Watch: Combining pod and container resource allocation flexibility (6:00)
Lucy Sweet provided live demonstrations for two of the three features discussed, illustrating their practical application and immediate benefits.
- Pod-Level Resource Requests and Limits Demo:
- Lucy started by showing a standard YAML file (
pod-level-before.yaml) where resources were explicitly defined for amain-appcontainer and asidecar-loggercontainer. This highlighted the traditional, per-container resource allocation. - She then demonstrated
pod-level-after.yaml, an identical application where theresourcesblock was moved from the individual container definitions to the top-levelspecof the pod. This visually showcased the simplification of configuration. - After applying the pod-level configuration using
kubectl apply -f pod-level-after.yaml, Lucy usedkubectl describe pod <pod-name>to show that the resources were now reported at the pod level, indicating that they were being shared among all containers within that pod. This effectively illustrated how a single resource boundary could serve multiple containers, reducing the need for rigid individual reservations.
- In-Place Pod Resizing Demo:
- Lucy first set her
kubectlcontext to a cluster configured for in-place resizing. - She then identified a running pod and used
kubectl edit pod <pod-name>to directly modify its CPU limit. Specifically, she changed thecpulimit from600mto800mfor a container within the running pod. - Upon saving the changes, the
kubectlcommand reported that the pod was "configured," indicating that the change was accepted without requiring a restart. - To verify, she immediately ran
kubectl get pod <pod-name> -o yaml, which displayed the updated CPU limit of800mfor the running pod. This demonstrated the core benefit: increasing resources for a live application with absolutely zero downtime or disruption to users, a significant improvement for stateful and critical workloads.
- Node Swap Support Demo (Partial):
- Lucy attempted to demonstrate node swap support on a single-node
kubeadmcluster. - She first confirmed that swap memory was active on the machine using the
free -horswapon -scommand, showing that swap was indeed configured and in use by the OS. - Next, she showed the
kubeadmconfiguration file, highlighting thenodeSwapfeature gate being enabled andswapBehavior: LimitedSwapbeing set. She explained thatLimitedSwapallows the "burstable" component of a pod's memory (between its request and limit) to use swap. - Unfortunately, the live demo encountered an unexpected cluster crash mid-explanation, preventing a full demonstration of a pod utilizing swap. Despite the hiccup, the setup and explanation clearly outlined how
kubeletwould be configured to enable swap, and the underlying mechanism ofLimitedSwapwas thoroughly explained. Lucy humorously noted that "two out of three succeed" was better than expected for live demos.
These demonstrations, even with the minor technical difficulty, effectively conveyed the power and practical implications of these advanced resource management features, making the concepts tangible for the audience.
Defensive Implications
▶ Watch: Advanced use cases: ML workloads and web services (7:00)
The introduction of Pod-level resources, in-place pod resizing, and node swap support brings powerful new capabilities but also new considerations for cluster administrators and application developers to ensure stability, security, and optimal performance.
For Pod-Level Resource Requests and Limits:
- Benefits: Simplifies configuration for multi-container pods, reduces over-provisioning by allowing sidecars to share resources more effectively with main applications. This can lead to higher node utilization and lower cloud costs.
- Defensive Actions:
- Review Resource Management Strategies: Teams should re-evaluate their resource request/limit policies. For pods with multiple tightly coupled containers, a pod-level approach might be more efficient than individual container settings.
- Monitor Shared Resource Utilization: Implement robust monitoring to ensure that no single container within a pod monopolizes shared resources, potentially starving others. While the kernel handles allocation, unexpected contention can occur if not properly sized.
- Feature Gate Management: As this is an Alpha feature, be cautious about production deployment. Understand the implications of enabling feature gates and stay updated on its graduation to Beta and GA.
For In-Place Pod Resizing:
- Benefits: Crucially reduces disruption for stateful and critical workloads, enabling dynamic scaling without downtime. This improves application availability and simplifies operational tasks.
- Defensive Actions:
- Application Compatibility Testing: Not all applications handle resource changes gracefully. Test applications thoroughly with in-place resizing, especially for memory adjustments, to understand if they require
RestartContainerpolicy or specific application-level reconfigurations (e.g., Java heap size adjustments). - Monitoring Node Capacity: While
kubeletperforms admission checks for resource increases, closely monitor node resource availability to prevent scenarios where resize requests are continuously rejected due to insufficient slack. - Integration with Autoscalers: Plan for integration with Vertical Pod Autoscalers (VPA) or custom autoscalers to automate resource adjustments. Manual
kubectl editis useful for demos but impractical at scale. Ensure autoscalers respectresizePolicyand application-specific restart requirements. - Security for Stateful Workloads: For very sensitive stateful applications (e.g., databases), ensure that any automated resizing logic has appropriate guardrails and adheres to established operational best practices to prevent unintended consequences, even with zero-disruption resizing.
For Node Swap Support:
- Benefits: Allows for better utilization of physical memory, reduces OOM kills for bursty workloads, and can lower infrastructure costs by reducing the need to over-provision RAM. Essential for large AI/ML and batch workloads.
- Defensive Actions:
- Cautious Rollout and Planning: Dawn Chen explicitly stated the community's initial strong opposition to swap due to performance and security concerns. Enable swap with extreme caution, especially in production. Platform administrators must carefully plan its introduction.
- Choose
swapBehaviorWisely:LimitedSwapis generally preferred as it protects guaranteed memory, only allowing burstable memory to be swapped.UnlimitedSwapshould be reserved for specific, performance-tolerant workloads. - Performance Monitoring: Implement comprehensive monitoring of disk I/O, memory usage (including swap), and application latency. Swap usage can introduce performance variability and "noisy neighbor" issues if not managed correctly. Identify and address potential thrashing scenarios.
- Security Implications: Be aware that data swapped to disk can be sensitive. Ensure that nodes with swap enabled have appropriate disk encryption and access controls. The security implications are a primary reason this feature is not yet GA.
- Eviction Policy Review: Understand how swap interacts with Kubernetes eviction policies. High swap usage might trigger evictions if not properly balanced with memory pressure.
- Kernel Interaction: Recognize that memory shrinking via swap is still challenging due to kernel limitations. This means that while swap can absorb bursts, reclaiming that memory efficiently might still be an issue without application restarts.
Overall, these features empower more sophisticated resource management but demand a deeper understanding of their underlying mechanisms, careful testing, and robust monitoring to harness their benefits safely and effectively.
Key Takeaways
- Resource Management Evolution: Kubernetes is rapidly evolving its resource management capabilities to support advanced, stateful, and bursty workloads beyond traditional stateless applications.
- Pod-Level Resource Sharing: The new Pod-level resource requests and limits (Alpha, Kube 1.32/1.33) simplify configuration and enable efficient resource sharing among containers within a pod, reducing over-provisioning for sidecars and multi-container applications.
- Zero-Downtime Resizing: In-place pod resizing (Beta, Kube 1.33, default enabled) allows dynamic adjustment of CPU and memory for running pods without disruption, a critical advancement for stateful applications like databases.
- Cautious Swap Adoption: Node swap support (Beta 3, Kube 1.33) addresses memory burst challenges, particularly for AI/ML and Java workloads, but requires careful configuration (
LimitedSwaprecommended) and monitoring due to historical performance and security concerns. - Future-Proofing Workloads: The community is actively exploring advanced concepts like Pod Grouping and dynamic container injection/removal to further enhance Kubernetes' support for distributed batch processing, AI/ML, and highly dynamic developer environments.
- Operational Due Diligence: While powerful, these features necessitate thorough testing, robust monitoring, and careful planning by cluster administrators to ensure application compatibility, optimal performance, and adherence to security best practices.
About the Speaker(s)
Lucy Sweet is an engineer working at Uber and a contributor to Kubernetes. In this presentation, she primarily adopted the role of a Kubernetes user, candidly sharing the common struggles and frustrations experienced with existing resource management paradigms when dealing with advanced and stateful workloads. Her insights provided a crucial "voice of the customer," highlighting the real-world problems that the new features aim to solve.
Dawn Chen is a Software Engineer at Google and holds the influential position of Tech Lead for the Kubernetes SIG Node since its inception. Her extensive experience and deep understanding of Kubernetes' core components, particularly at the node level, were evident throughout the talk. Dawn played a pivotal role in designing and implementing many of the advanced resource management features discussed, including her initial strong (and well-reasoned) opposition to node swap and the eventual shift in strategy to support it for evolving workload needs. She also provided a visionary outlook on the future of Kubernetes resource management, including proposals for Pod Grouping and dynamic container management, underscoring her ongoing commitment to making Kubernetes more capable for a broader range of demanding applications.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This talk by Lucy Sweet and Dawn Chen delivers crucial insights into Kubernetes' evolving resource management. It directly tackles long-standing issues for stateful and bursty workloads with deep dives into new core features like pod-level resources, in-place resizing, and node swap support. This isn't theoretical fluff; it's a pragmatic look at solving real operational challenges, presented by the engineers who are literally building these capabilities into the platform.
Heather Calloway (CISO) — STRONG ACCEPT
This presentation, "Taming the Beast," effectively addresses critical operational challenges in Kubernetes resource management, offering actionable solutions that directly translate to enhanced business resilience, optimized infrastructure costs, and improved stability for complex, stateful workloads. The introduction of pod-level resources, in-place resizing, and cautious swap support provides platform teams with powerful new tools, demanding a proactive re-evaluation of existing resource policies, diligent testing, and robust monitoring to fully realize their benefits while mitigating new risks.