More Nodes, More Problems: Solving Multi-Host GPU/TPU Scheduli... John Belamaric & Morten Torkildsen
John Belamaric, Morten Torkildsen
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In the rapidly evolving landscape of artificial intelligence and machine learning, workloads are continually growing in scale and complexity, often demanding vast arrays of specialized hardware such as GPUs and TPUs. However, managing these multi-host accelerator resources within Kubernetes presents significant challenges for developers and cluster operators alike. This talk by John Belamaric and Morten Torkildsen from Google delves into these intricate problems, highlighting the limitations of current Kubernetes scheduling mechanisms and presenting Dynamic Resource Allocation (DRA) as a robust, future-forward solution.

Key moments
- 0:00 Introduction and multi-host GPU/TPU scheduling problems
- 2:50 Requirements for multi-node accelerator jobs: placement, partitioning, failures
- 4:00 Current static multi-node allocation using Kubernetes node labels
- 5:30 Limitations: static partitioning, race conditions, no self-healing
- 6:30 Demonstration of a multi-node job failure and manual recovery
- 7:50 Introducing Dynamic Resource Allocation (DRRA) for Kubernetes
More Nodes, More Problems: Solving Multi-Host GPU/TPU Scheduling in Kubernetes
Speakers: John Belamaric, Staff Software Engineer, Google; Morten Torkildsen, Software Engineer, Google
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=YqIHESG0suI
Overview
In the rapidly evolving landscape of artificial intelligence and machine learning, workloads are continually growing in scale and complexity, often demanding vast arrays of specialized hardware such as GPUs and TPUs. However, managing these multi-host accelerator resources within Kubernetes presents significant challenges for developers and cluster operators alike. This talk by John Belamaric and Morten Torkildsen from Google delves into these intricate problems, highlighting the limitations of current Kubernetes scheduling mechanisms and presenting Dynamic Resource Allocation (DRA) as a robust, future-forward solution.
The core of the presentation focuses on how DRA, a relatively new Kubernetes API, can dynamically allocate and manage these high-value, interconnected devices across multiple nodes. The speakers illustrate how DRA addresses critical issues like achieving compact hardware placement, ensuring atomic partitioning of resources for multiple concurrent jobs, and, crucially, enabling automated failure recovery without manual intervention. This innovation is paramount for enhancing resource utilization, improving job reliability, and ultimately reducing the operational burden associated with large-scale, distributed AI/ML training jobs.
For anyone involved in deploying or operating AI/ML infrastructure on Kubernetes, understanding DRA is becoming increasingly vital. The speakers, both long-time contributors to the Kubernetes project and DRA, unveil features that are either newly available or on the horizon, particularly those slated for Kubernetes 1.33. Their insights offer a glimpse into how Kubernetes is evolving to meet the demanding requirements of modern high-performance computing and distributed machine learning.
Background
▶ Watch: Introduction and multi-host GPU/TPU scheduling problems (0:00)
The proliferation of large-scale AI/ML training jobs has brought to light significant limitations in how Kubernetes traditionally manages specialized hardware. As workloads expand, they increasingly require not just more accelerators (GPUs or TPUs), but often dozens or even hundreds of these devices spread across multiple physical nodes. This distributed nature inherently introduces more failure points, and the coordinated mechanisms often required by these jobs mean that the failure of a single node can catastrophically impact the entire multi-node workload.
To illustrate the problem, the speakers use an example of a cluster with 64 TPUs, with four TPUs per node, totaling 16 nodes. Multiple jobs might each require a portion of these resources. For such an allocation, several key requirements emerge:
- Compact Placement: Workloads need to be scheduled on nodes that are physically close to each other. For TPUs, this involves specific topological constraints due to their high-speed interconnects. For GPUs, it could mean ensuring all allocated devices are within the same rack to minimize network latency.
- Atomic Partitioning: When multiple jobs simultaneously request resources, their allocations must be isolated. One job's selection of nodes should not "pollute" or interfere with another's compact placement group, preventing race conditions.
- Automated Failure Recovery: In the event of a node failure, the system should ideally self-heal, re-allocating resources and restarting the job without requiring manual intervention, especially at inconvenient hours.
The existing Kubernetes mechanisms fall short in addressing these requirements. In current setups, multi-node TPU slices are often managed by manually assigning node labels to groups of nodes corresponding to specific hardware topologies. Users then specify these labels in their pod or job specifications using node selectors. This approach leads to several critical issues:
- Static Partitioning vs. Utilization: Cluster operators must choose between statically partitioning the cluster into fixed-size slices (which often leads to underutilization if jobs have varying resource needs) or risking race conditions where multiple jobs might target overlapping node label sets.
- Manual Coordination: Users or job authors must manually coordinate which node labels to use, leading to administrative overhead and potential conflicts if not meticulously managed.
- Lack of Self-Healing: The most significant drawback is the complete absence of automated failure recovery. If a node within an allocated slice fails (e.g., node 12 in a 4x4 TPU slice), the pods scheduled on it become unschedulable. Because the node selector is fixed in the job's specification, the Kubernetes scheduler cannot dynamically re-evaluate and select a new set of nodes. The user is forced to manually terminate the failed job, pick a new, available set of nodes (with a different node label), and resubmit the job. This process is disruptive, time-consuming, and entirely contrary to Kubernetes' self-healing philosophy.
These limitations underscore the need for a more dynamic and intelligent resource management system within Kubernetes, particularly for high-value, multi-host accelerator resources.
Key Findings
▶ Watch: Current static multi-node allocation using Kubernetes node labels (4:00)
The central finding and proposed solution presented in this talk is the Dynamic Resource Allocation (DRA) API. DRA is introduced as a fundamental shift in how Kubernetes manages specialized hardware, moving beyond static declarations and manual interventions to a dynamic, API-driven approach. The speakers highlight DRA's maturity, noting its beta status in Kubernetes 1.32, with crucial features for multi-host accelerators slated for Kubernetes 1.33.
DRA is conceptualized as having four core components:
- API to describe devices: This is primarily handled by the ResourceSlice API, which allows nodes to advertise their available devices and their capabilities. For instance, a node can declare that it possesses an "Nvidia GPU with 40 GB of memory."
- API for users to request devices: Users interact with DRA through ResourceClaims. A ResourceClaim specifies the type and quantity of devices needed, along with any specific requirements. An example would be, "I need two Nvidia GPUs, each with at least 30 GB of memory."
- Allocation mechanism: This is where the Kubernetes scheduler plays a pivotal role. It takes the available devices (from ResourceSlices) and the user requests (from ResourceClaims) and intelligently allocates devices to satisfy those claims.
- Kubelet API for actuation: Once the scheduler makes an allocation decision, the Kubelet (Kubernetes agent on each node) uses specific APIs to make the allocated devices available to the pods and containers running on that node.
The adoption of DRA brings several significant advantages:
- Enhanced Scheduler Awareness: The scheduler gains a comprehensive understanding of all available devices and their specific characteristics across the entire cluster, enabling more intelligent and optimized placement decisions.
- Improved Autoscaling: With device information now available to the Kubernetes API server, the autoscaler can make more informed decisions, scaling the cluster up or down based on the demand for specialized hardware, a significant improvement over previous iterations.
- Support for Multi-Node Devices: Critically, DRA is designed to model not only devices attached to individual nodes but also those accessible across multiple nodes, such as high-speed network interconnects or shared storage.
For tackling the specific challenges of multi-host accelerators, Kubernetes 1.33 introduces two key features that synergize with DRA:
- Partitionable Devices: This feature allows DRA to model devices that can be logically divided into smaller, independent units. A prominent example is Nvidia MIG (Multi-Instance GPU), which enables partitioning a single GPU into several smaller, isolated GPU instances.
- Device Taints: Extending the existing concept of node taints, device taints allow specific devices to be tainted. This makes it possible to automatically evict pods that have a claim on a tainted device.
By combining these features, DRA provides a robust and dynamic solution for managing multi-host accelerators, addressing the shortcomings of previous, static approaches.
Technical Deep Dive
▶ Watch: Limitations: static partitioning, race conditions, no self-healing (5:30)
The technical foundation of DRA for multi-host accelerators lies in its ability to abstract complex hardware topologies into manageable logical devices and dynamically manage their allocation.
On the driver side, a DRA Resource Driver operates as a Kubelet plugin. This driver is responsible for discovering the physical devices attached to its node and publishing their capabilities and configurations to the Kubernetes API server as ResourceSlices. For example, in a TPU cluster, the driver would identify the available TPU hardware and its interconnections. This information is then visible to the Kubernetes scheduler, allowing it to understand the cluster's device inventory.
On the consumer side, users define their hardware requirements using ResourceClaims. A ResourceClaim specifies the desired device class (e.g., TPU), the quantity of devices, and any specific attributes via selectors (e.g., memory size, number of cores). Pods then reference one or more ResourceClaims. A crucial aspect is that multiple pods can reference the same ResourceClaim, implying that they will share the devices allocated to that claim. Conversely, a single pod can reference multiple ResourceClaims if it requires different types of devices.
The core innovation for multi-host TPUs within DRA involves how these complex hardware arrangements are modeled. Instead of individual TPUs, each interconnected TPU "slice" – a logical grouping of TPUs and nodes forming a coherent topology (e.g., an 8x8 slice with 64 TPUs, a 4x4 slice with 16 TPUs, or 2x2 slices with 4 TPUs) – is represented as a logical device. Each logical device advertises its total capacity (e.g., 64 TPUs) and, crucially, carries a node selector that defines the specific set of nodes where this logical device resides. For instance, a 4x4 slice composed of nodes 1, 2, 5, and 6 would have a node selector targeting these four nodes.
A key challenge in multi-host topologies is the inherent overlap between different potential slices. A set of nodes might simultaneously be part of an 8x8 slice, multiple 4x4 slices, and numerous 2x2 slices. DRA intelligently handles this by allowing a single underlying physical device (or set of devices) to be advertised as part of multiple logical devices. However, DRA understands the relationships between these overlapping logical devices. If a smaller 2x4 slice (comprising nodes 11 and 12) is allocated, DRA automatically prevents the allocation of any larger logical device (like a 4x4 slice) or other overlapping smaller devices (like the 2x2s formed by nodes 11 and 12) that share the same underlying physical resources. This mechanism effectively eliminates the race conditions and manual coordination issues that plagued the previous node-label-based approach.
When a user submits a ResourceClaim for a multi-host TPU workload—for example, requesting a device with 16 TPUs (implying four nodes with four TPUs each)—the Kubernetes scheduler steps in. It evaluates all available logical devices that satisfy this claim. Currently, the scheduler employs a first-fit algorithm, selecting the first suitable device it finds. However, future plans include implementing scoring mechanisms to enable a best-fit allocation based on various criteria.
The allocation decision is made dynamically during the scheduling of the first pod that references the ResourceClaim. Once a specific logical device is chosen, its associated node selector is then applied to all pods referencing that ResourceClaim. This dynamically restricts where these pods can be scheduled, ensuring they land on the nodes comprising the selected compact slice. For example, if the scheduler picks a 4x4 slice on the bottom-left (nodes 11, 12, 15, 16), the first worker pod might be placed on Node 11, and the remaining pods for that job will be scheduled on Nodes 12, 15, and 16, ensuring they all run within the designated, topologically optimized group. This dynamic restriction is a significant improvement over users manually specifying fixed node selectors, as the allocation is now managed and enforced by the system.
Demo / Proof of Concept
▶ Watch: Demonstration of a multi-node job failure and manual recovery (6:30)
While the presentation did not feature a live, interactive code demo, the speakers provided a clear conceptual walkthrough, using detailed diagrams and illustrative scenarios to demonstrate how DRA operates and, critically, how it solves the previously intractable problem of automated failure recovery for multi-host accelerators. This conceptual "demo" effectively highlighted the practical implications of DRA's design.
The core of the demonstration focused on the failure recovery scenario, contrasting it directly with the manual process required by the old node-label approach.
- Initial Setup: A job is launched, requesting a 16-TPU slice (four nodes). The scheduler, using DRA, dynamically selects an available 4x4 slice (e.g., nodes 11, 12, 15, 16) and schedules the four worker pods across these nodes.
- Node Failure: During job execution, Node 12 unexpectedly fails (e.g., due to a kernel bug). The pod running on Node 12 becomes unschedulable, and the job stalls, similar to the pre-DRA scenario.
- Automated Recovery with DRA: This is where DRA's new features (coming in Kubernetes 1.33) take over. The DRA resource driver detects the failure on Node 12 and, recognizing that the logical device (the 4x4 slice) is now impaired, taints that specific device with a
NoExecuteeffect. - Pod Eviction and Rescheduling: The device taint automatically triggers the eviction of all pods that were referencing the ResourceClaim allocated to the now-tainted device. These evicted pods revert to a pending state in the scheduler's queue.
- Dynamic Reallocation: When the scheduler next processes one of these pending pods, it recognizes that the previously allocated device is tainted. It then dynamically selects a new, healthy 16-TPU logical device (e.g., another 4x4 slice that is available). All the pods associated with the original ResourceClaim are then rescheduled onto the nodes of this newly allocated device.
The critical takeaway from this conceptual demonstration is the complete automation of the recovery process. Unlike the old system where an operator would have to manually kill the job, find a new set of nodes, and resubmit, DRA handles the detection, eviction, and reallocation seamlessly. As John Belamaric succinctly put it, "You can stay asleep when this happens at 3:00 a.m." This fully automated self-healing capability is a monumental improvement for the reliability and operational ease of distributed AI/ML workloads.
Defensive Implications
▶ Watch: Introducing Dynamic Resource Allocation (DRRA) for Kubernetes (7:50)
While DRA is not a security feature in the traditional sense, its implications for the reliability, efficiency, and operational resilience of Kubernetes clusters running advanced workloads are profoundly "defensive." It empowers cluster operators and developers to build more robust and self-healing systems, effectively defending against common points of failure and inefficiency in distributed computing environments.
For Cluster Operators and SREs:
- Reduced Operational Burden: The most immediate "defensive" gain is the automation of failure recovery. Operators no longer need to manually intervene when a node supporting an accelerator-heavy workload fails. This dramatically reduces on-call fatigue and minimizes downtime for critical training jobs, freeing up SREs to focus on more strategic tasks.
- Improved Resource Utilization: By allowing logical devices to overlap and by giving the scheduler a dynamic understanding of resource availability, DRA helps prevent resource fragmentation. This leads to higher overall utilization of expensive GPU and TPU hardware, which is a significant cost-saving "defense" against under-provisioning or inefficient resource allocation.
- Enhanced Cluster Resilience: The ability to automatically reallocate workloads upon device failure makes the entire cluster more resilient to hardware faults. This is crucial for long-running, multi-node training jobs where even a single node failure can otherwise lead to significant delays and wasted compute cycles.
For Workload Developers:
- Simplified Workload Definition: Developers can define their resource requirements abstractly (e.g., "I need 16 TPUs") rather than being tied to specific node labels or manual topology choices. This simplifies pod specifications and makes workloads more portable and less prone to misconfiguration.
- Consistent Performance: By ensuring compact and topologically optimal placement, DRA helps maintain the high-bandwidth, low-latency communication paths essential for distributed AI/ML training, thus "defending" against performance degradation due to suboptimal scheduling.
Future Defensive Enhancements (Current Open Issues/Future Work):
The speakers also acknowledged several areas for future improvement, which can be seen as ongoing "defensive" work:
- Ensuring Per-Node Pod Placement: Currently, DRA doesn't strictly guarantee that if a claim allocates four nodes, each of four pods will land on a different node. While workarounds exist (e.g., requesting sufficient per-pod resources to limit placement, using anti-affinity rules, or per-pod claims), simplifying this is a goal. This would "defend" against accidental over-scheduling on a single node within an allocated slice.
- Best-Fit Scoring and Preemption: Moving beyond a first-fit allocation to a "best-fit" model (via scoring) will further optimize resource usage. Implementing preemption will allow higher-priority workloads to "defend" their resource needs by preempting lower-priority jobs, ensuring critical tasks are prioritized.
- Granular Resource Modeling: Addressing the challenge of modeling more complex resource attributes beyond simple device counts (e.g., throughput, latency, power consumption, or specific network interface capabilities like requesting a 2 Gigabit partition from a 10 Gigabit NIC) is a future goal. This would provide a more fine-grained "defense" against suboptimal performance due to unmodeled resource constraints.
- Transparent Capability Advertisement: A long-term vision involves nodes transparently publishing their capabilities (e.g., specific Kubelet feature gates, GPU types, or network topologies) to the scheduler. This would eliminate the need for manual tolerations or admission webhooks (like those in GKE for device plugins) and provide a more robust "defense" against pods being scheduled on nodes lacking required features. This infrastructure would allow the scheduler to both attract and repel pods based on their needs and node capabilities automatically.
In summary, DRA acts as a critical defensive layer by automating complex resource management, enhancing resilience, improving efficiency, and simplifying the operational overhead for high-performance, accelerator-driven workloads in Kubernetes.
Key Takeaways
- Limitations of Traditional Kubernetes: Existing Kubernetes mechanisms, such as static node labels and device plugins, are inadequate for managing complex, multi-host GPU/TPU accelerator topologies due to issues like static partitioning, race conditions, and the complete lack of automated failure recovery.
- Introducing Dynamic Resource Allocation (DRA): DRA is a new Kubernetes API designed to dynamically allocate and manage specialized devices, including those accessible across multiple nodes. It provides a structured way for nodes to describe devices (ResourceSlices) and for users to request them (ResourceClaims).
- Dynamic Allocation and Placement: DRA empowers the Kubernetes scheduler to make intelligent, dynamic decisions about which multi-host device slices to allocate, ensuring compact placement and preventing resource conflicts by understanding relationships between overlapping logical devices.
- Automated Failure Recovery in 1.33: Upcoming features in Kubernetes 1.33, specifically partitionable devices and device taints, are crucial. They enable DRA to model complex topologies and, most importantly, provide automated self-healing by evicting and rescheduling pods when an allocated device or its underlying node fails, eliminating manual intervention.
- Improved Utilization and Operational Ease: DRA significantly boosts resource utilization for expensive accelerators, simplifies workload definitions for developers, and drastically reduces the operational burden for SREs by automating critical management and recovery tasks for large-scale AI/ML workloads.
- Future Enhancements: Ongoing development for DRA includes refining pod placement within allocated nodes, implementing "best-fit" scoring for optimal allocation, supporting preemption for high-priority workloads, and expanding to model more granular resource attributes beyond simple device counts.
About the Speaker(s)
John Belamaric is a Staff Software Engineer at Google, bringing extensive experience to the Kubernetes project. He has been deeply involved with Kubernetes for approximately seven to eight years. For the past year, John has dedicated his efforts to the development of Dynamic Resource Allocation (DRA), a project he notes was initiated by Patrick Oulie from Intel and significantly contributed to by Kevin Clues, alongside many other community members.
Morten Torkildsen is also a Software Engineer at Google. He has been involved with the Kubernetes project for several years, contributing to various aspects on an intermittent basis. Morten has been a key contributor to the DRA project for roughly the last six months, working alongside John and the broader community to advance its capabilities, particularly for the complex use cases discussed in this talk.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk presents Dynamic Resource Allocation (DRA) as a much-needed evolution in Kubernetes for managing complex, multi-host GPU/TPU workloads. It directly tackles critical issues like compact placement, atomic partitioning, and, most importantly, automated failure recovery, which has been a significant operational blind spot. The speakers, core contributors to DRA, lay out a clear, technically sound vision for how Kubernetes is finally catching up to the demands of large-scale AI/ML infrastructure, making it a strong recommendation for anyone dealing with these challenges.
Heather Calloway (CISO) — STRONG ACCEPT
This talk presents a robust and well-articulated solution for managing the complex operational challenges of multi-host GPU/TPU scheduling in Kubernetes. Dynamic Resource Allocation (DRA) directly addresses critical issues like resource utilization, atomic placement, and, most importantly, automated failure recovery, which is a significant step forward for the resilience of large-scale AI/ML workloads. While technical in nature, the implications for business continuity, cost efficiency, and reducing operational risk are clear and highly relevant for leaders overseeing modern, AI-driven infrastructure.