Scaling GPU Clusters Without Melting Down! - Alay Patel & Ryan Hallisey, NVIDIA

Alay Patel, Ryan Hallisey, NVIDIA

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

In an era where Graphics Processing Units (GPUs) are becoming exponentially more powerful, enabling complex AI, machine learning, and high-performance computing workloads, the challenge of scaling the underlying infrastructure has intensified. This talk, delivered by NVIDIA’s Alay Patel and Ryan Hallisey at KubeCon EU, delves into the critical and often overlooked aspect of maintaining the stability and performance of the Kubernetes control plane when scaling GPU clusters. The speakers share NVIDIA's firsthand experiences and the architectural lessons learned from operating bare metal Kubernetes environments under extreme load.

Watch on YouTube

Visual summary for Scaling GPU Clusters Without Melting Down! - Alay Patel & Ryan Hallisey, NVIDIA by Alay Patel, Ryan Hallisey, NVIDIA
Visual summary for Scaling GPU Clusters Without Melting Down! - Alay Patel & Ryan Hallisey, NVIDIA by Alay Patel, Ryan Hallisey, NVIDIA

Key moments

  1. 0:00 Introduction: Scaling GPUs pressures Kubernetes control plane
  2. 1:30 Problem 1: Large secrets list calls cause API server OOM
  3. 2:50 Solution 1: API Priority and Fairness limits concurrent requests
  4. 4:40 Problem 2: API server memory utilization drops after restart
  5. 5:50 Solution 2: Tuning Go GC for aggressive garbage collection
  6. 8:00 Go GC: Understanding CPU and memory trade-offs
  7. 9:30 Problem 3: API server memory utilization skew between instances

Scaling GPU Clusters Without Melting Down!

Speakers: Alay Patel, Ryan Hallisey, NVIDIA

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=dUfp3j1j-mg

Overview

In an era where Graphics Processing Units (GPUs) are becoming exponentially more powerful, enabling complex AI, machine learning, and high-performance computing workloads, the challenge of scaling the underlying infrastructure has intensified. This talk, delivered by NVIDIA’s Alay Patel and Ryan Hallisey at KubeCon EU, delves into the critical and often overlooked aspect of maintaining the stability and performance of the Kubernetes control plane when scaling GPU clusters. The speakers share NVIDIA's firsthand experiences and the architectural lessons learned from operating bare metal Kubernetes environments under extreme load.

The core premise of their investigation centered on a crucial question: can the compute capacity of a Kubernetes cluster, particularly one powered by ever-advancing GPUs, scale independently of a fixed-size control plane? As GPUs grow in power, they enable more workloads, which in turn necessitates more Kubernetes objects like secrets and volumes. This cascading effect places immense and often unforeseen pressure on the Kubernetes API server and other control plane components. The talk meticulously details the "incidents, incidents, incidents" encountered during their scaling experiments and the ingenious fine-tuning techniques employed to overcome these hurdles, ensuring the control plane's resilience.

This article provides a comprehensive breakdown of NVIDIA's journey, highlighting the specific problems encountered, the technical solutions implemented, and the broader implications for anyone operating large-scale Kubernetes clusters, especially those with specialized hardware demands. It serves as a practical guide for understanding and mitigating common control plane scalability bottlenecks, offering insights into advanced Kubernetes features and Go runtime optimizations that are critical for maintaining a robust and efficient cloud-native infrastructure.

Background

▶ Watch: Introduction: Scaling GPUs pressures Kubernetes control plane (0:00)

NVIDIA, as a prominent provider and user of GPU technology, operates Kubernetes clusters on bare metal infrastructure to meet the demanding requirements of its high-performance workloads. Their journey towards a cloud-native platform involved migrating from a legacy homegrown cloud, a transition that, while beneficial, introduced a new set of scaling challenges to their Kubernetes control plane. The fundamental problem they sought to address was the rapidly increasing power of GPUs, which allows for denser and more complex workloads. Theoretically, more powerful GPUs should support more workloads, but this directly translates to a greater number of Kubernetes objects—secrets, volumes, pods, and more—all of which exert significant pressure on the Kubernetes API server and the entire control plane.

To investigate this, NVIDIA conducted an experiment: they fixed the CPU and memory resources allocated to their Kubernetes control plane and then scaled up the compute capacity by adding more powerful GPUs and associated workloads. The expectation was that the control plane would adapt, but the reality was a series of unexpected outages and performance degradation. Two significant architectural "mistakes" made during their cloud-native transition exacerbated these issues:

  1. Large and Numerous Secrets: A substantial amount of data, including configuration and certificates for workloads, was stored directly in Kubernetes secrets. These were not typical small secrets; each was approximately 0.2 megabytes in size, and there were 5,000 such secrets in their environment. While etcd could store them, performing a list call for these secrets would load all 5,000 objects, totaling around 1 GB of data, into the API server's cache, leading to severe memory pressure.
  2. Periodic List Calls by Clients: The client responsible for orchestrating GPU capacity across the cluster, designed to work with both legacy and cloud-native environments, made frequent periodic list calls against these workloads. This pattern of interaction, especially with the unusually large secrets, created a "list call storm" that could overwhelm the API server.

These two factors combined to create a scenario where the control plane was highly susceptible to Out-of-Memory (OOM) errors and cascading failures, underscoring the critical need for fine-grained control and optimization within the Kubernetes ecosystem.

Key Findings

▶ Watch: Solution 1: API Priority and Fairness limits concurrent requests (2:50)

NVIDIA’s scaling experiments revealed three primary categories of problems, each requiring a distinct and sophisticated solution to ensure the stability and performance of their Kubernetes control plane.

The first major finding was that large numbers of list calls to substantial secrets were directly leading to control plane outages. As detailed in the background, their environment featured 5,000 secrets, each weighing in at 0.2 MB. When clients initiated concurrent list calls for these secrets, the cumulative data (approximately 1 GB) would flood the API server's memory cache, quickly causing OOM events and API server restarts. The solution to this critical issue was the implementation of API Priority and Fairness (APF). APF, a feature designed to protect the API server from request overloads, allowed NVIDIA to define specific priority levels and concurrency limits for different types of API requests. By configuring APF, they were able to restrict the number of concurrent list calls to secrets, specifically limiting them to two concurrent calls. This effectively prevented the memory exhaustion, safeguarding the API server from OOM-induced outages.

The second key finding related to API server memory utilization skew and cascading failures. Initially, the team observed that API server restarts often resulted in a significant reduction in memory utilization, sometimes by 30-40%, even under identical workloads. This suggested inefficiencies in Go's garbage collection or potential memory leaks. Their investigation led them to tune the Go garbage collector (Go GC). By adjusting the GOGC environment variable to a lower value (e.g., GOGC=50), they made the garbage collection process more aggressive. This doubled the rate of Go GC cycles, increased the amount of memory released, and ultimately lowered the live heap utilization across all control plane components. This tuning resulted in a sharp and sustained decrease in memory usage, proving to be a highly effective optimization, albeit with a recognized trade-off between CPU and memory resources.

Beyond general memory optimization, NVIDIA also faced a critical problem where memory utilization would become highly skewed across API server instances, with one instance consuming significantly more memory than others. This skew made the higher-memory instance more vulnerable to OOMs during traffic spikes, leading to cascading failures as its workload was redirected to already strained peers. The root cause was traced to a skew in HAProxy and API server connection distribution, where some API servers received disproportionately fewer connections (e.g., 37 connections versus thousands on other instances). The solution involved configuring the goaway-chance parameter in the API server. This parameter probabilistically sends GOAWAY TCP messages to established, long-lived client connections. This graceful teardown allows clients to re-establish connections, which then get re-balanced by the load balancer, effectively mitigating the skew. NVIDIA observed that even if a skew existed, the goaway-chance parameter enabled the cluster to self-recover and rebalance traffic within 15 minutes.

Finally, the talk highlighted NVIDIA's proactive approach to future-proofing their GPU orchestration stack by transitioning to Dynamic Resource Allocation (DRA). This is a significant architectural shift away from their current reliance on custom device plugins and a topo-aware scheduler for NUMA alignment. Through scale testing with Quark, a Kubernetes simulation tool, they found that DRA offered substantial performance and efficiency improvements: the DRA scheduler exhibited twice the speed in scheduling latency, achieved a 30% memory saving for the scheduler components (reducing total scheduler memory from 1.3 GB to 1 GB), and showed a 20% better memory footprint on the API server compared to the legacy stack. These findings affirm DRA as a promising path for more efficient and scalable GPU resource management in Kubernetes.

Technical Deep Dive

▶ Watch: Problem 2: API server memory utilization drops after restart (4:40)

The solutions presented by NVIDIA for scaling their GPU clusters without melting down delve into sophisticated Kubernetes and Go runtime mechanisms. Understanding these technical details is crucial for anyone seeking to implement similar optimizations.

API Priority and Fairness (APF) for API Server Protection

The first critical issue, API server OOMs due to large list calls to secrets, was addressed using API Priority and Fairness (APF). APF is a robust mechanism introduced in Kubernetes to protect the API server from being overwhelmed by a flood of requests. It achieves this by categorizing incoming requests into FlowSchemas and assigning them PriorityLevels. Each PriorityLevel has a defined concurrency limit, which dictates how many requests from that level can be processed concurrently.

In NVIDIA's scenario, they identified that list calls to their unusually large (0.2 MB each) and numerous (5,000) secrets were the culprits. Two concurrent list calls could consume approximately 1 GB of API server memory, leading to OOM errors. By configuring APF, they created a specific FlowSchema to match these problematic list requests and assigned it to a PriorityLevel with a very low concurrency limit, specifically allowing only two concurrent list calls for these secrets. This effectively throttled the rate at which these memory-intensive operations could hit the API server, preventing resource exhaustion while still allowing legitimate requests to proceed. APF acts as a circuit breaker, ensuring that even under extreme load, critical API server functions remain available.

Go GC (GOGC) for Memory Optimization

The observation that API server restarts often led to significant memory reductions (30-40%) pointed towards inefficient memory management within the Go runtime, which underpins Kubernetes components. The solution involved tuning the Go garbage collector (Go GC) via the GOGC environment variable.

The GOGC variable controls the garbage collector's aggressiveness. Its value represents a percentage: GOGC=100 (the default) means the Go runtime will trigger a garbage collection cycle when the amount of newly allocated memory since the last GC cycle reaches 100% of the live heap size after the previous GC. For example, if the live heap is 100 MB, GC will trigger when 100 MB of new objects have been allocated, bringing total allocation to 200 MB. By setting GOGC to a lower value, such as GOGC=50, the garbage collector becomes more aggressive. In this case, GC would trigger when new allocations reach 50% of the previous live heap, meaning GC would run when total allocation reaches 150 MB.

NVIDIA's tuning of GOGC to a lower value (e.g., GOGC=50) resulted in:

  • A doubling of Go GC cycles.
  • An increase in the total bytes released by the garbage collector.
  • A noticeable reduction in the live heap utilization across all control plane components (API server, controller manager, scheduler).

This optimization is a trade-off between CPU and memory. More aggressive GC (lower GOGC) means more frequent GC cycles, which consume more CPU resources. Conversely, a less aggressive GC (higher GOGC) consumes less CPU but allows memory utilization to grow larger. NVIDIA, being memory-constrained in their control plane, opted for efficient memory utilization at the cost of slightly higher CPU usage. They also noted that other organizations, like Uber, have successfully tuned GOGC in the opposite direction (higher GOGC) to prioritize lower CPU utilization when memory was not a constraint. This highlights the flexibility of GOGC as a tunable for Go applications.

goaway-chance for Connection Skew Mitigation

The problem of API server memory skew, leading to cascading failures, was ultimately traced to an uneven distribution of long-lived TCP connections from the load balancer (HAProxy) to the API server instances. Standard load balancing typically focuses on distributing new connections fairly. However, once a TCP connection is established, especially for long-lived protocols like HTTP/2 used by Kubernetes clients (e.g., for watch requests), the load balancer has no control over its subsequent traffic flow. This can lead to a situation where some API server instances handle significantly more requests and maintain more connections than others, causing memory and CPU imbalances.

NVIDIA discovered and utilized the goaway-chance parameter in the Kubernetes API server. This parameter addresses the problem of sticky, long-lived connections by probabilistically sending an HTTP/2 GOAWAY frame (which is a TCP message) to established client connections. A GOAWAY frame gracefully signals to the client that the server intends to close the connection, encouraging the client to re-establish new connections. When these new connections are initiated, they again pass through the HAProxy load balancer, which then has the opportunity to distribute them more evenly across the available API server instances.

This mechanism acts as a self-healing process for connection skew. NVIDIA observed that configuring goaway-chance allowed their API server fleet to automatically recover from connection and memory skew within 15 minutes. It's a powerful technique not limited to Kubernetes, applicable to any server-side component fronted by a load balancer where long-lived connections can cause load imbalances. The speakers noted a lack of detailed documentation for this parameter, advocating for a pull request that provides a comprehensive description of its function and benefits.

Dynamic Resource Allocation (DRA) for Future GPU Orchestration

Looking to the future, NVIDIA is making a significant architectural shift from their current GPU orchestration stack to Dynamic Resource Allocation (DRA). Their existing stack (shown on the left in their presentation) relies on:

  • Device plugins on worker nodes for GPU allocation.
  • A custom topo-aware scheduler in the control plane to perform NUMA (Non-Uniform Memory Access) alignment for high-performance GPU workloads, ensuring optimal CPU and memory locality with GPUs.
  • An NFD topology updater on worker nodes to feed NUMA topology information to the control plane.

The transition to DRA (shown on the right) aims to simplify this by:

  • Leveraging DRA plugins on worker nodes.
  • Eliminating the custom topo-aware scheduler, as the DRA API is designed to be rich enough to handle NUMA alignment natively within the default kube-scheduler.
  • Removing the NFD topology updater, streamlining the worker node stack.

To validate this architectural change, NVIDIA utilized Quark, a Kubernetes simulation tool that can create fake nodes and pods to generate pressure on the control plane without requiring actual Kubelets or worker nodes. Through extensive scale testing with Quark, comparing the current stack with the future DRA stack, they gathered compelling performance metrics:

  • Scheduling Latency: The DRA scheduler demonstrated twice the performance in terms of scheduling latency compared to the topo-aware scheduler.
  • Scheduler Memory Utilization: The combined memory footprint of the topo-aware scheduler and kube-scheduler was around 1.3 GB at scale. With DRA, the kube-scheduler alone (handling NUMA alignment) consumed only 1 GB, representing a 30% memory saving.
  • API Server Memory Footprint: The DRA orchestration mechanism was found to be 20% more efficient in terms of its impact on API server memory utilization compared to the current topo-aware scheduler approach.

These results indicate that DRA not only simplifies the orchestration logic but also offers significant improvements in control plane performance and resource efficiency, making it a crucial component for future-proof GPU cluster scaling.

Demo / Proof of Concept

▶ Watch: Go GC: Understanding CPU and memory trade-offs (8:00)

While the talk didn't feature a live, interactive demonstration or a publicly available proof-of-concept tool, the speakers presented extensive data and graphs derived from their internal experiments and simulations. They described their methodology, which involved running controlled scaling tests in their bare metal Kubernetes environments and leveraging specific tools for simulation.

A key tool mentioned for their performance and scale testing was Quark. Quark is described as "Kubernetes without Kubelet," a simulation tool designed to help users create fake nodes and pods within a cluster. This allows for generating significant pressure on the control plane (API server, scheduler, controller manager) without the need for a large physical compute infrastructure. By simulating thousands of workloads and nodes, NVIDIA was able to measure metrics like scheduling latency, memory utilization of control plane components, and API server impact under different orchestration strategies (their current device plugin/topo-aware scheduler stack versus the future DRA stack). The detailed graphs showing memory utilization over time, scheduling latency comparisons, and API server memory footprints served as compelling evidence of their findings and the effectiveness of their solutions.

Defensive Implications

▶ Watch: Problem 3: API server memory utilization skew between instances (9:30)

The insights shared by NVIDIA offer crucial defensive strategies and best practices for Kubernetes operators, SREs, and platform engineers managing large-scale clusters, particularly those supporting demanding workloads like GPUs. Implementing these lessons can significantly enhance control plane stability and resilience.

  1. Audit and Optimize Secret Usage: The primary takeaway from NVIDIA's initial problems is to scrutinize how secrets are used. Avoid storing excessively large amounts of data (e.g., 0.2 MB per secret) or a massive number of secrets (e.g., 5,000) directly in Kubernetes secrets if this data is frequently listed. For large configurations, certificates, or binary data, consider alternatives like ConfigMaps (if not sensitive), external secrets managers (e.g., HashiCorp Vault, AWS Secrets Manager, Azure Key Vault, GCP Secret Manager), or mounting data directly from object storage or specialized CSI volumes.
  2. Implement API Priority and Fairness (APF): Proactively configure APF to protect your API server. Identify potentially expensive or high-volume requests (e.g., large list calls, frequent updates to specific resources) and define FlowSchemas and PriorityLevels to limit their concurrency. This acts as a crucial safeguard against resource exhaustion and ensures that critical control plane operations are not starved by less important or misbehaving client requests. Regularly review APF configurations as cluster usage patterns evolve.
  3. Monitor Control Plane Component Memory and Connection Skew: Establish robust monitoring and alerting for memory utilization across all API server instances and other control plane components. Pay close attention to any significant memory skew between replicas. Additionally, monitor connection counts from your load balancer (e.g., HAProxy) to individual API server instances to detect and alert on uneven load distribution. Early detection of skew is vital to prevent cascading failures.
  4. Tune Go Garbage Collection (GOGC): For Kubernetes components written in Go, consider experimenting with the GOGC environment variable. If memory is a consistent constraint in your control plane, a more aggressive GOGC (e.g., GOGC=50) might improve memory efficiency. However, carefully monitor CPU utilization during and after tuning, as this is a direct trade-off. Conduct thorough testing in non-production environments to understand the impact before deploying to production.
  5. Configure goaway-chance for Load Balancing Resilience: Implement the goaway-chance parameter in your API server configurations to mitigate connection skew for long-lived HTTP/2 connections. This probabilistic mechanism ensures that connections are periodically torn down gracefully, forcing clients to reconnect and allowing the load balancer to redistribute traffic more evenly. This is a powerful self-healing mechanism for maintaining balanced load across API server replicas.
  6. Embrace Dynamic Resource Allocation (DRA): For clusters utilizing specialized hardware like GPUs, actively plan and transition to Dynamic Resource Allocation (DRA). This upcoming Kubernetes feature promises to streamline device orchestration, improve scheduling efficiency, and reduce the complexity of custom schedulers. Staying abreast of DRA's development and planning for its adoption can lead to significant performance and scalability gains.
  7. Leverage Advanced API Machinery Features: The speakers also hinted at future Kubernetes API machinery improvements like "list streams." Staying updated with the latest Kubernetes releases and features, particularly those from SIG API Machinery, can provide native solutions to common scaling problems, potentially obviating the need for some of the complex tunables described.
  8. Proactive Scale Testing with Tools like Quark: Before deploying significant architectural changes, scaling up compute capacity, or introducing new workload patterns, utilize simulation tools like Quark. This allows for realistic pressure testing of the control plane in a cost-effective manner, identifying bottlenecks and validating optimizations without impacting production environments.

Key Takeaways

  • Architectural Decisions Matter: Storing large or numerous secrets, combined with frequent list calls by clients, can severely degrade Kubernetes control plane performance and lead to Out-of-Memory (OOM) errors and outages.
  • API Priority and Fairness is a Critical Safeguard: Kubernetes' API Priority and Fairness (APF) feature is essential for protecting the API server from resource exhaustion. It allows operators to define concurrency limits for specific, expensive requests (e.g., limiting large secret list calls to two concurrent requests), ensuring control plane stability.
  • Go Runtime Optimization Improves Efficiency: Tuning the Go garbage collector (Go GC) via the GOGC environment variable can significantly reduce memory utilization in Go-based control plane components (e.g., 30-40% reduction after restarts, 30% saving for schedulers), but requires careful consideration of the CPU/memory trade-off.
  • Mitigate Connection Skew with goaway-chance: Uneven distribution of long-lived TCP connections from load balancers can cause performance skew and cascading failures across API server replicas. The goaway-chance parameter offers a probabilistic, self-healing mechanism to gracefully tear down and rebalance these connections within minutes.
  • Dynamic Resource Allocation (DRA) is the Future: Transitioning to Dynamic Resource Allocation (DRA) for specialized hardware like GPUs promises substantial improvements in control plane efficiency, including twice the scheduling speed, 30% memory savings for schedulers, and a 20% better API server memory footprint.
  • Proactive Scale Testing is Indispensable: Tools like Quark enable cost-effective simulation of large-scale clusters, allowing operators to pressure-test the control plane and validate architectural changes or optimizations before deployment, preventing production incidents.

About the Speaker(s)

Ryan Hallisey is a professional working at NVIDIA. He initiated the talk by setting the context for the challenges faced when scaling Kubernetes control planes in the face of increasingly powerful GPUs. Ryan detailed the first major problem encountered by NVIDIA: control plane outages caused by large list calls to numerous, oversized secrets, and explained how API Priority and Fairness was successfully used to mitigate this.

Alay Patel is also a professional at NVIDIA and Ryan's colleague. He continued the presentation by delving into the subsequent scaling issues and their solutions. Alay provided in-depth explanations of how tuning the Go GC parameter improved memory utilization across control plane components, and how the goaway-chance parameter effectively solved the problem of API server memory and connection skew. He also outlined NVIDIA's strategic move towards Dynamic Resource Allocation (DRA) for future GPU orchestration, sharing detailed performance and scale testing results using Quark. Alay is actively involved in scale and performance work within the Kubernetes community, particularly concerning DRA.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This isn't some 'AI-powered' marketing drivel. This is NVIDIA, showing how they nearly melted down their Kubernetes control plane scaling GPU clusters and what real, technical solutions they engineered. From taming API server OOMs with APF and aggressive Go GC tuning, to elegantly solving connection skew with goaway-chance, and proving out DRA's future, this talk delivers substance. It's a masterclass in operational deep-diving, offering actionable insights for anyone running Kubernetes at scale. No bullshit, just hard-won engineering.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025