Lightning Talk: Resource Roulette: Winning the Kubernetes Allocation Game - Daniele Polencic

Daniele Polencic

KubeCon + CloudNativeCon Europe 2025 · Lightning Talk

Overview

In his KubeCon EU lightning talk, "Resource Roulette: Winning the Kubernetes Allocation Game," Daniele Polencic, an instructor from Learn Kubernetes, delves into one of the most persistent and critical challenges in operating Kubernetes clusters: the optimal configuration of resource requests and limits. Polencic frames this challenge as a "roulette" because misconfigurations can lead to a gamble between two undesirable outcomes: underutilization of expensive infrastructure, or performance degradation and service instability. The talk meticulously unpacks the intricacies of how Kubernetes allocates resources, the common pitfalls faced by operators, and the emerging solutions designed to automate and optimize this complex process.

Watch on YouTube

Visual summary for Lightning Talk: Resource Roulette: Winning the Kubernetes Allocation Game - Daniele Polencic by Daniele Polencic
Visual summary for Lightning Talk: Resource Roulette: Winning the Kubernetes Allocation Game - Daniele Polencic by Daniele Polencic

Key moments

  1. 0:00 Introduction: The problem of unmanaged Kubernetes resources
  2. 0:25 Setting requests and limits in pod specifications
  3. 0:50 Problem 1: Underutilization and wasting requested resources
  4. 1:20 Problem 2: Overutilization leading to contention or cost
  5. 2:16 The challenge of applications with high resource spikes
  6. 3:27 Introducing Vertical Pod Autoscaler (VPA) for dynamic adjustment
  7. 4:00 Ecosystem tools and market consolidation for optimization
  8. 4:30 Key takeaways: Always set requests and adapt them with tooling

Lightning Talk: Resource Roulette: Winning the Kubernetes Allocation Game

Speakers: Daniele Polencic, Instructor, Learn Kubernetes

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=f6gYxJOr0yQ

Overview

In his KubeCon EU lightning talk, "Resource Roulette: Winning the Kubernetes Allocation Game," Daniele Polencic, an instructor from Learn Kubernetes, delves into one of the most persistent and critical challenges in operating Kubernetes clusters: the optimal configuration of resource requests and limits. Polencic frames this challenge as a "roulette" because misconfigurations can lead to a gamble between two undesirable outcomes: underutilization of expensive infrastructure, or performance degradation and service instability. The talk meticulously unpacks the intricacies of how Kubernetes allocates resources, the common pitfalls faced by operators, and the emerging solutions designed to automate and optimize this complex process.

Polencic’s presentation is particularly relevant for anyone involved in deploying, managing, or optimizing applications on Kubernetes – from developers and SREs to platform engineers and financial operations teams. The core problem he addresses is the dynamic nature of application resource consumption versus the static configuration paradigm of Kubernetes requests and limits. Achieving a balance that ensures both application stability and cost efficiency is a continuous struggle, often requiring deep insights into workload patterns and a proactive approach to resource management.

The significance of this talk extends beyond mere technical configuration; it touches upon the economic implications of cloud-native operations and the operational resilience of distributed systems. By highlighting the trade-offs between over-provisioning and under-provisioning, Polencic sets the stage for understanding why manual resource tuning is unsustainable and why automated solutions are becoming indispensable. His insights underscore the need for a strategic approach to resource governance that moves beyond guesswork, aiming instead for precision and adaptability in the dynamic Kubernetes environment.

Background

▶ Watch: Introduction: The problem of unmanaged Kubernetes resources (0:00)

Kubernetes, as a container orchestration platform, fundamentally relies on an efficient and fair distribution of computational resources among its various workloads. At the heart of this resource management are two critical parameters defined within a Pod's specification: requests and limits. Understanding their purpose and interaction is paramount to operating a healthy and cost-effective cluster.

A resource request for CPU and memory dictates the minimum amount of resources a container requires to run. When a Pod is scheduled, the Kubernetes scheduler considers the available resources on each node and only places the Pod on a node that can satisfy its requests. For CPU, requests are measured in "millicores" (e.g., 200m for 0.2 CPU cores), and for memory, in bytes (e.g., 256Mi for 256 mebibytes). The scheduler guarantees that these requested resources will be available to the Pod. If a Pod requests 200m CPU, it is guaranteed at least 200 millicores, even under contention. If a node cannot fulfill the requests of a pending Pod, that Pod will remain in a pending state until a suitable node becomes available or resources are freed. This mechanism prevents nodes from becoming oversubscribed to the point of complete collapse due to resource starvation.

In contrast, a resource limit defines the maximum amount of resources a container is allowed to consume. For CPU, if a container attempts to use more CPU than its limit, it will be throttled, meaning its execution will be slowed down to stay within the defined boundary. For memory, exceeding the limit is a more severe event: the container will be immediately terminated by the kernel with an Out-Of-Memory (OOM) error. This mechanism acts as a safeguard, preventing runaway processes from consuming all available resources on a node and impacting other Pods.

The interplay between requests and limits also determines a Pod's Quality of Service (QoS) class, which influences how Kubernetes handles resource contention and eviction.

  1. Guaranteed QoS: Achieved when requests and limits are set equally for all containers in a Pod, for both CPU and memory. These Pods are given the highest priority and are least likely to be evicted under resource pressure.
  2. Burstable QoS: Occurs when requests are set, but limits are higher than requests (or limits are not set, implying a limit equal to the node's capacity). These Pods can "burst" beyond their requests if resources are available, but they are subject to throttling if they exceed their CPU limit and can be evicted if their actual memory usage exceeds their request during memory pressure.
  3. BestEffort QoS: Assigned when no requests or limits are specified for any container in a Pod. These Pods receive the lowest priority and are the first to be evicted during resource contention. Polencic highlights the critical issue that if no requests are set, Kubernetes has no basis for intelligent scheduling, leading to unpredictable placement and potential resource imbalance across nodes.

The fundamental problem Polencic addresses is that application resource usage is rarely static. It fluctuates based on traffic patterns, background jobs, and internal processing. Manually setting requests and limits for hundreds or thousands of microservices in a dynamic environment becomes an intractable challenge. Engineers often resort to guesswork, leading to one of two detrimental scenarios:

  • Over-provisioning: Setting requests and limits too high, guaranteeing more resources than the application actually needs. This results in wasted compute capacity, higher cloud infrastructure costs, and inefficient cluster utilization. Polencic graphically illustrates this with charts showing requested but unused resources.
  • Under-provisioning: Setting requests and limits too low. This can lead to CPU throttling, causing performance degradation and increased latency (e.g., high P99 latency as mentioned by Polencic), or frequent OOM errors, leading to service instability and restarts. In a scenario where all Pods simultaneously "burst" beyond their requests, resource contention becomes severe, directly impacting user experience.

The dilemma, as Polencic states, is being "in between these two things": paying more for resources you don't use, or suffering from poor performance and developer experience due to resource contention. This forms the crucial context for his discussion on how to "win" the allocation game by moving beyond static configurations.

Key Findings

▶ Watch: Problem 1: Underutilization and wasting requested resources (0:50)

Polencic's talk distills the complex interplay of Kubernetes resource management into several key findings, centered on the dynamic nature of application workloads and the need for adaptive solutions.

Firstly, a foundational discovery, often overlooked by those new to Kubernetes, is the critical importance of always setting resource requests. As Polencic illustrates, without requests, Kubernetes has no basis for intelligent scheduling, leading to arbitrary Pod placement and potential resource imbalance across nodes. This lack of information is the root cause of many operational headaches, as it prevents the scheduler from making informed decisions about where to place workloads to ensure optimal performance and resource availability.

Secondly, Polencic highlights the inherent challenge posed by the fluctuating nature of application resource usage. Workloads rarely consume a constant amount of CPU or memory. Instead, they exhibit "wide intervals" and "spikes" in usage. This variability makes static configuration of requests and limits incredibly difficult and prone to error. He presents two primary problematic scenarios:

  1. Underutilization: When requests are set higher than the application's actual average usage. While the requested resources are guaranteed, the application uses only a fraction, leading to significant waste. This translates directly to increased infrastructure costs without proportional benefits.
  2. Contention and Performance Degradation: When the application's actual usage frequently exceeds its defined requests (but remains within limits). Although not immediately fatal, this can lead to CPU throttling and severe resource contention if many Pods on a node exhibit this behavior simultaneously. Polencic notes the consequence: "P99 latency," indicating a poor user and developer experience due to inconsistent performance.

The core insight here is that the ideal scenario is to have a "very, very narrow interval where the request stays mostly close to what you define." Achieving this ideal, however, is nearly impossible manually, especially when "you don't know if the request is going to go up, it's going to stay the same or it's going to go down. You don't even know how big the interval is going to be." This uncertainty is the fundamental obstacle to effective resource allocation.

Finally, Polencic identifies automated tooling as the primary solution to this dilemma. He explains that the strategic approach to managing fluctuating resource demands involves continuously adapting requests and limits based on observed usage patterns. He introduces the concept of "dividing these intervals into smaller intervals" and predicting resource needs for each, thereby narrowing the gap between requested and actual usage. This shift from static, reactive configuration to dynamic, proactive optimization is where "we see innovation in the space." He specifically mentions the Vertical Pod Autoscaler (VPA) as a prime example of an open-source tool addressing this, alongside several commercial solutions, and notes the ongoing market consolidation, signaling the growing importance and adoption of such tools in the Kubernetes ecosystem.

Technical Deep Dive

▶ Watch: The challenge of applications with high resource spikes (2:16)

The technical core of Polencic's talk revolves around the nuanced management of Kubernetes resource requests and limits, and the advanced techniques and tools employed to optimize them. The discussion highlights the inherent tension between guaranteeing resources for stability and minimizing waste for cost efficiency.

At the fundamental level, Kubernetes uses spec.containers[].resources.requests and spec.containers[].resources.limits to manage how much CPU and memory a container can consume.

  • CPU Requests (cpu): Measured in millicores (e.g., 200m for 0.2 CPU core). The scheduler uses this value to determine if a node has sufficient allocatable CPU to host the Pod. Once scheduled, the container runtime (e.g., containerd, CRI-O) configures the Linux kernel's CFS (Completely Fair Scheduler) to guarantee this amount of CPU. If a container requests 200m of CPU, it will get 20% of a CPU core's time slice, even if other containers on the same node are competing for resources.
  • CPU Limits (cpu): Also measured in millicores. If a container attempts to use more CPU than its limit, the CFS will throttle it, preventing it from exceeding this cap. This ensures that a single misbehaving process doesn't hog all CPU cycles on a node, but it also means that a burstable workload might experience performance degradation if its limit is set too low.
  • Memory Requests (memory): Measured in bytes (e.g., 256Mi). Similar to CPU requests, the scheduler uses this to find a node with enough free memory. For memory, the request acts as a baseline; the kernel will try to reclaim memory from other processes before touching a Pod's requested memory.
  • Memory Limits (memory): Measured in bytes. This is a hard cap. If a container attempts to allocate more memory than its limit, the Linux kernel's OOM Killer will terminate the process. This is a critical event, often leading to Pod restarts and potential service disruption.

Polencic's core technical challenge lies in the dynamic nature of application resource consumption. He illustrates scenarios where:

  1. Requests > Actual Usage: The application consistently uses less CPU or memory than requested. For CPU, this means the guaranteed share is not fully utilized, leading to idle CPU cycles that cannot be easily reallocated to other Pods unless they are also under-requesting. For memory, the reserved memory remains unused. This is a direct source of resource waste and increased cloud costs.
  2. Requests < Actual Usage <= Limits: The application frequently consumes more resources than requested, but stays within its defined limits. For CPU, this results in throttling when the container tries to exceed its request, leading to increased latency and reduced throughput. For memory, if the node experiences memory pressure, the kernel might reclaim memory from this Pod down to its request, potentially causing performance issues or even OOM if the Pod cannot shed memory fast enough. This scenario, as Polencic points out, leads to "contention" and "P99 latency," indicating a degraded user experience.
  3. Actual Usage > Limits: For memory, this immediately triggers the OOM Killer, leading to Pod termination and restarts. For CPU, it implies continuous throttling, making the application effectively unusable.

The speaker introduces the concept of "wide intervals" for resource usage, meaning the peak usage is significantly higher than the average or minimum usage. Manually setting requests for such workloads is a gamble: setting them low risks performance issues, while setting them high guarantees waste. His proposed solution is to "divide these intervals into smaller intervals," effectively tracking and predicting resource usage over shorter, more manageable periods. This allows for a more granular and accurate adjustment of resource requests and limits.

This is precisely where Vertical Pod Autoscalers (VPA) come into play. The VPA is a Kubernetes component that automatically adjusts the CPU and memory requests and limits for containers in a Pod. It operates by observing the actual resource usage of Pods over time and recommending (or directly applying) optimal values.

  • VPA Components:
  • Admission Controller: Intercepts Pod creation requests and injects recommended resource requests/limits into the Pod spec.
  • Recommender: Monitors the current and historical resource usage of Pods and calculates optimal resource requests and limits. It uses statistical models to predict future usage based on past patterns.
  • Updater: Evicts Pods that need their requests/limits updated, allowing the Admission Controller to re-inject the new recommendations when the Pod is rescheduled. (In Off mode, it only provides recommendations without evicting Pods).
  • VPA Modes:
  • Off: VPA only provides recommendations in the VPA object status, without modifying Pods.
  • Initial: VPA sets resource requests and limits only when a Pod is first created. It doesn't update them during the Pod's lifecycle.
  • Recreate: VPA updates requests and limits by recreating Pods when a significant change in resource recommendations is detected.
  • Auto (default): VPA continuously updates requests and limits by recreating Pods when necessary, and also sets them on Pod creation.

Polencic explicitly mentions VPA as a "simple tool" that works by "dividing the time and then just trying to predict based on the path performance what should be the next request." He also highlights that the ecosystem offers more advanced solutions, often leveraging machine learning models, to achieve even greater precision. He lists several commercial products: StormForge, PepperScale, Tens Qbacks from Densifiers, and ScaleOps. These tools typically offer more sophisticated analytics, integration with cost management platforms, and potentially more nuanced scaling strategies than the open-source VPA. The observation that "most of these are actually being acquired" and the "market being consolidated" underscores the significant value and demand for intelligent resource optimization in cloud-native environments.

The technical challenge is not merely about finding the peak usage but understanding the typical usage patterns, seasonality, and application-specific requirements. Automated tools aim to remove the guesswork, reducing both operational overhead and infrastructure costs, while simultaneously improving application stability and performance.

Demo / Proof of Concept

▶ Watch: Introducing Vertical Pod Autoscaler (VPA) for dynamic adjustment (3:27)

Daniele Polencic's presentation was delivered as a lightning talk, a format designed for concise and impactful discussions rather than live demonstrations. As such, the talk did not include a live demo or a detailed proof of concept of any of the resource optimization tools mentioned. Instead, Polencic relied on conceptual diagrams and clear explanations to convey the problem and the high-level solutions. The focus was on articulating the core challenge of resource allocation and the types of automated approaches available in the Kubernetes ecosystem.

Defensive Implications

▶ Watch: Key takeaways: Always set requests and adapt them with tooling (4:30)

While Daniele Polencic's talk focuses on resource optimization rather than traditional cybersecurity, the principles discussed have significant defensive implications for the operational resilience and security posture of Kubernetes clusters. In the context of "winning the allocation game," defensive measures extend beyond preventing malicious attacks to ensuring the stability, availability, and performance of critical services, which are fundamental pillars of any robust security strategy.

  1. Preventing Resource Exhaustion Attacks (DoS/DDoS): Properly configured resource requests and limits act as the first line of defense against Denial-of-Service (DoS) or Distributed Denial-of-Service (DDoS) attacks that target resource exhaustion. If an attacker manages to exploit a vulnerability or flood an application with excessive requests, well-defined limits can prevent a single compromised or overloaded Pod from consuming all CPU or memory on a node, thereby protecting other co-located services. Without limits, a rogue process could easily starve the entire node, leading to a widespread outage. Automated tools like VPA, by keeping limits close to optimal, help ensure that excess capacity isn't inadvertently granted to a compromised workload.
  1. Enhancing System Stability and Availability: Under-provisioned resources, particularly memory, can lead to frequent OOMKills and Pod restarts. While not a direct security breach, service instability significantly degrades availability and can be exploited by attackers seeking to disrupt operations. Conversely, over-provisioning can mask performance issues or hide resource leaks, making it harder to detect legitimate problems. By optimizing requests and limits, defenders ensure that critical applications receive the resources they need to remain stable and available, even under stress. This directly contributes to the confidentiality, integrity, and availability (CIA triad) of services.
  1. Mitigating "Noisy Neighbor" Problems: In multi-tenant or shared cluster environments, one application's excessive resource consumption (a "noisy neighbor") can degrade the performance of other, potentially more critical, workloads. Well-defined requests and limits, especially when dynamically adjusted by tools like VPA, prevent any single workload from monopolizing shared resources. This isolation is crucial for maintaining service level agreements (SLAs) and preventing unintended service disruptions that could be perceived as an attack or lead to cascading failures.
  1. Improving Observability and Incident Response: When resource requests and limits are accurately set, deviations from expected resource usage become more apparent. If a Pod suddenly starts hitting its CPU limit or is frequently OOMKilled despite optimized settings, it signals an anomaly that warrants investigation. This could indicate a performance regression, a resource leak, or even a subtle form of attack (e.g., a cryptominer running in a compromised container). Automated optimization tools provide a baseline of "normal" behavior, making it easier for security and operations teams to detect and respond to unusual activity.
  1. Cost Optimization as a Security Enabler: While seemingly unrelated, cost optimization through efficient resource allocation frees up budget that can be reinvested into security initiatives, such as advanced threat detection tools, security audits, or hiring specialized security personnel. By eliminating waste, organizations can better afford a stronger security posture.

In essence, Polencic's discussion on intelligent resource allocation is about building more resilient and predictable Kubernetes environments. A system that is stable, performant, and efficiently utilizes its resources is inherently more defensible against both accidental operational failures and malicious attacks. The automation provided by tools like VPA transforms resource management from a reactive firefighting task into a proactive defense mechanism, ensuring that the underlying infrastructure is robust enough to support secure operations.

Key Takeaways

  • Always Set Resource Requests and Limits: Explicitly defining requests and limits for CPU and memory is fundamental for predictable Kubernetes scheduling and preventing uncontrolled resource consumption. Omitting them leads to arbitrary Pod placement and potential resource starvation or waste.
  • Balance Between Waste and Performance: There's a critical trade-off: over-provisioning leads to wasted cloud infrastructure costs, while under-provisioning causes performance degradation, increased latency (e.g., P99 latency), and potential service instability (e.g., OOMKills).
  • Dynamic Workloads Require Dynamic Management: Application resource usage fluctuates significantly with "wide intervals" and "spikes," making static, manual configuration of requests and limits inefficient and unsustainable. The goal is to narrow the gap between requested and actual usage.
  • Automated Tools are Essential for Optimization: Relying on human guesswork for resource tuning is error-prone. Tools like Vertical Pod Autoscaler (VPA), StormForge, PepperScale, Tens Qbacks from Densifiers, and ScaleOps leverage historical data and predictive models to continuously adapt and optimize resource requests and limits.
  • Continuous Adaptation is Key: Resource settings should not be a one-time configuration. They require ongoing adjustment and adaptation based on observed workload patterns to maintain efficiency and performance over time.
  • Market Consolidation Reflects Growing Importance: The acquisition trend among companies offering Kubernetes resource optimization solutions highlights the increasing recognition and demand for intelligent, automated resource management in the cloud-native landscape.

About the Speaker(s)

Daniele Polencic is introduced as an instructor from Learn Kubernetes. While the talk itself is a concise lightning presentation, his role as an instructor suggests a deep practical and theoretical understanding of Kubernetes, particularly in areas critical for operational efficiency and performance. His ability to distill complex topics like resource allocation into an accessible and engaging narrative underscores his expertise in educating practitioners on best practices within the Kubernetes ecosystem. His focus on the "allocation game" and the challenges faced by operators reflects a hands-on perspective derived from teaching and guiding individuals through the intricacies of cloud-native deployments.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This lightning talk, 'Resource Roulette,' delivers a brutally honest and technically sound breakdown of Kubernetes resource allocation, a problem that plagues every operator. Polencic cuts through the noise, meticulously articulating the financial waste and performance degradation caused by manual guesswork. He makes a compelling, actionable case for automated solutions like VPA, demonstrating a deep understanding of the underlying kernel mechanisms and the economic realities of cloud-native operations. It's a direct, no-bullshit guide to a critical operational challenge that every SRE and platform engineer faces daily.

Heather Calloway (CISO) — STRONG ACCEPT

Daniele Polencic's lightning talk effectively diagnoses a critical and persistent operational challenge in Kubernetes: balancing resource efficiency with application stability. By framing resource allocation as a 'roulette' between over-provisioning waste and under-provisioning performance degradation, he clearly articulates the business impact of static resource management. The talk convincingly advocates for automated tooling like Vertical Pod Autoscalers as the necessary strategic shift, providing a clear path for organizations to improve operational resilience, reduce cloud spend, and enhance overall system predictability.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025