WG-Batch Updates: What’s New and What Is Next? - Marcin Wielgus, Google

Marcin Wielgus, Google

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

Marcin Wielgus, an organizer of the Kubernetes Batch Working Group, presented a comprehensive update on the group's efforts to enhance Kubernetes' support for batch workloads. This presentation, titled "WG-Batch Updates: What’s New and What Is Next?", highlighted the significant advancements in Kueue, a powerful workload-level scheduler, alongside updates to core Kubernetes APIs like Job and new APIs such as JobSet and K-job. The talk underscored the Batch Working Group's mission to reduce fragmentation within the Kubernetes ecosystem, making it a more robust and efficient platform for high-performance computing (HPC), artificial intelligence (AI), machine learning (ML), data analytics, and continuous integration (CI) tasks.

Watch on YouTube

Visual summary for WG-Batch Updates: What’s New and What Is Next? - Marcin Wielgus, Google by Marcin Wielgus, Google
Visual summary for WG-Batch Updates: What’s New and What Is Next? - Marcin Wielgus, Google by Marcin Wielgus, Google

Key moments

  1. 0:00 Introduction to Batch Working Group and its goals
  2. 2:00 Introducing Kueue: The batch workload scheduler
  3. 3:30 Deep dive into Kueue's features and capabilities
  4. 6:00 Problem: Inefficient resource use without topology awareness
  5. 8:00 Solution: Kueue's topology-aware scheduling optimizes workloads
  6. 9:00 Introducing fair sharing and hierarchical quotas concept

WG-Batch Updates: What’s New and What Is Next?

Speakers: Marcin Wielgus, Google

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=aWxuaEFSarU

Overview

Marcin Wielgus, an organizer of the Kubernetes Batch Working Group, presented a comprehensive update on the group's efforts to enhance Kubernetes' support for batch workloads. This presentation, titled "WG-Batch Updates: What’s New and What Is Next?", highlighted the significant advancements in Kueue, a powerful workload-level scheduler, alongside updates to core Kubernetes APIs like Job and new APIs such as JobSet and K-job. The talk underscored the Batch Working Group's mission to reduce fragmentation within the Kubernetes ecosystem, making it a more robust and efficient platform for high-performance computing (HPC), artificial intelligence (AI), machine learning (ML), data analytics, and continuous integration (CI) tasks.

The core of the working group's initiatives revolves around addressing the unique challenges posed by batch workloads, which often require specialized scheduling, resource management, and hardware utilization strategies not inherently provided by standard Kubernetes deployments. By fostering collaboration across various Special Interest Groups (SIGs) like Scheduling, Apps, Node, and Autoscaling, the Batch Working Group aims to develop and refine tools and APIs that maximize cluster utilization, optimize workload performance, and seamlessly integrate with specialized hardware such as GPUs and TPUs across diverse cloud and on-premise environments. This talk serves as a critical update for cluster administrators, ML engineers, and researchers looking to leverage Kubernetes for demanding computational tasks.

Background

▶ Watch: Introduction to Batch Working Group and its goals (0:00)

The evolution of cloud-native computing has seen Kubernetes emerge as the de facto orchestrator for containerized applications. However, its initial design was primarily optimized for long-running, stateless services. Batch workloads, encompassing HPC simulations, large-scale data processing, AI model training, and complex CI pipelines, present a distinct set of requirements that often clash with Kubernetes' default scheduling and resource management paradigms. These workloads typically demand gang scheduling (where all components of a job must start simultaneously), resource isolation, fair sharing among multiple tenants, and topology-aware placement to optimize inter-process communication.

Historically, organizations running such workloads on Kubernetes have often resorted to custom solutions or external schedulers, leading to fragmentation and inconsistent operational practices. This fragmentation not only increases complexity for users and administrators but also hinders the efficient utilization of expensive specialized hardware. The Kubernetes Batch Working Group was formed to address this gap, providing a unified forum for discussing and implementing enhancements that bring first-class support for batch processing into the Kubernetes core and its extended ecosystem. Their primary endeavor, Kueue, was developed to act as an intelligent workload-level scheduler, bridging the gap between generic Kubernetes scheduling and the specific needs of resource-intensive, interdependent batch jobs.

Key Findings

▶ Watch: Deep dive into Kueue's features and capabilities (3:30)

The KubeCon EU update from the Batch Working Group revealed several significant advancements and new projects aimed at solidifying Kubernetes' position as a premier platform for batch workloads. The key findings presented by Marcin Wielgus centered on enhancements to Kueue, new APIs, and supporting tools:

  1. Kueue's Enhanced Capabilities: The project, which has seen over 50 releases, introduced critical features like topology-aware scheduling to optimize workload placement for network efficiency, and fair sharing mechanisms (both preemption-based and admission-based) to ensure equitable resource distribution among tenants. It also gained hierarchical quota management, allowing organizations to structure resource allocation according to their internal hierarchies.
  1. Admission-Based Fair Sharing: A new, sophisticated fair-sharing mechanism is being introduced that prioritizes workloads from teams that have historically used less of the shared capacity, without resorting to preemption. This approach focuses on admission control to achieve fairness, storing decaying aggregated usage over time.
  1. Expanded API Support: Updates to the venerable Kubernetes Job API include graduating features like managed by field (beta in Kubernetes 1.32), pod failure policy (GA in 1.31), success policy (GA in 1.33), and backoff limit per index (GA in 1.33), making jobs more robust.
  1. JobSet API Maturity: The JobSet API is gaining significant traction, providing a unified way to manage groups of jobs as a single unit, with extended policies for starting, stopping, and failure discovery. It's now adopted by frameworks like Kubeflow Trainer v2 and Metaflow for distributed training.
  1. Introduction of K-job: A novel API and command-line tool designed to simplify the submission of numerous, slightly varied one-time jobs, particularly for users accustomed to systems like Slurm. K-job enables the creation of reusable job templates stored in the API server, which can be easily instantiated and executed via a command-line interface.
  1. KueueCTL Tool: A kubectl plugin, KueueCTL, was launched to streamline day-to-day operations for Kueue users, providing functionalities for managing queues, listing workloads, and controlling their lifecycle.

These developments collectively represent a concerted effort to provide a comprehensive and integrated ecosystem for batch workloads on Kubernetes, addressing both technical performance bottlenecks and operational complexities.

Technical Deep Dive

▶ Watch: Problem: Inefficient resource use without topology awareness (6:00)

Kueue stands as the cornerstone of the Batch Working Group's efforts, serving as an AI HPC batch workload level scheduler. Unlike the default Kubernetes scheduler, which operates on individual pods, Kueue manages entire workloads (represented by Custom Resource Definitions or CRDs), making decisions on when to start, stop, or preempt them. A critical feature is its support for all-or-nothing or gang scheduling semantics, ensuring that all components of a multi-pod workload are allocated resources simultaneously. This is particularly vital for ML and HPC tasks where partial execution is unproductive. Kueue also acts as a resource quota manager, enabling the definition of quotas across multiple tenants and resource types, including specialized hardware like GPUs and TPUs, and is hardware vendor and cloud provider neutral.

One of Kueue's recent and impactful additions is topology-aware scheduling. The problem it addresses is inefficient workload placement in data centers. Imagine a large cluster where pods from a single workload are scattered across different network switches. If these pods need to communicate extensively, the inter-switch link can become a severe bottleneck, leading to degraded performance and wasted compute resources, especially expensive GPUs. Kueue mitigates this by intelligently placing all pods of a single workload onto machines connected to the same switch or within the same rack. This minimizes cross-switch traffic, ensuring that high-bandwidth inter-pod communication remains efficient and workloads progress unhindered. Future enhancements aim to integrate more Kubernetes scheduler code into Kueue for even more precise admissions and placement, and to offer finer controls for scoring topologies.

Another significant feature is fair sharing, designed to optimize resource utilization and ensure equitable access, especially when capacity is temporarily unused. Marcin Wielgus illustrated this with a scenario involving three research teams (blue, red, yellow) allocated different capacities of a GPU cluster. During off-peak times or holidays, when a team isn't fully utilizing its quota, Kueue can dynamically redistribute the unused capacity. Two flavors of fair sharing are available:

  1. Preemption-based fair sharing: In this mode, if a team (e.g., blue) is using more than its fair share of shared capacity and another team (e.g., yellow) needs resources for a smaller workload, Kueue can preempt workloads from the over-utilizing team to make space. For instance, if a blue team has workloads requiring 50, 30, and 15 GPUs, and a yellow team needs 20 GPUs but there's no space, Kueue might preempt the 15-GPU blue workload to accommodate the yellow one, balancing the shared capacity. However, it avoids preemption if doing so would reverse the fairness, giving the preempting team an unfair advantage.
  1. Admission-based fair sharing: This upcoming feature, slated for the next Kueue release, takes a less aggressive approach. Instead of preempting, it prioritizes the admission of new workloads based on historical usage. If a team has been using significantly more shared resources, its new workloads will wait, giving preference to workloads from teams that have used less. This method considers a decaying aggregated usage over time, ensuring that teams that have historically used less shared capacity get priority for new admissions. This provides a more gentle fairness mechanism, particularly suitable for scenarios where preemption is undesirable.

Complementing fair sharing is hierarchical quotas. Organizations often structure their resource allocations based on their internal hierarchy (e.g., director allocates quota to managers, who then allocate to teams). Kueue allows for the configuration of these hierarchical quotas, ensuring that resource distribution and fair sharing mechanisms respect the organizational chart. Unused capacity at lower levels can then flow up the hierarchy and be redistributed among other teams or the wider organization, maximizing overall cluster utilization.

Beyond Kueue, the Batch Working Group is actively enhancing Kubernetes' native Job API. Recent graduations include:

  • managed by field: Went to Beta in Kubernetes 1.32.
  • pod failure policy: Went to GA in Kubernetes 1.31, providing more granular control over how job failures are handled.
  • success policy: Slated for GA in Kubernetes 1.33, allowing flexible definitions of job success.
  • backoff limit per index: Also going to GA in 1.33, enabling per-index backoff limits for indexed jobs.

The JobSet API is presented as a "Job API on steroid," designed for managing groups of jobs as a unified entity. It offers extended policies for starting, stopping, failure detection, and restarts, crucial for complex, multi-component HPC and AI/ML workloads. JobSet provides a single API to govern interconnected jobs, and its maturity is evidenced by its adoption in prominent frameworks like Kubeflow Trainer v2 and Metaflow for distributed training.

Finally, the K-job project addresses the challenge of submitting large numbers of similar, one-time jobs with complex storage configurations. Researchers often have many jobs that differ only slightly (e.g., different arguments, input files, or images) and prefer command-line interfaces over YAML editing. K-job introduces reusable job templates stored in the API server. Administrators configure these templates with common elements like storage volumes, and researchers can then select a template, provide specific parameters via a CLI tool, and submit a fully formed job. A special mode within K-job provides a Slurm-like experience, mimicking srun command-line options and environment variables, making migration from traditional HPC environments to Kubernetes significantly smoother.

The ecosystem is further supported by KueueCTL, a kubectl plugin for day-to-day Kueue operations, enabling tasks like creating, draining, stopping, and listing queues and workloads. This holistic approach, combining a sophisticated scheduler like Kueue with enhanced core APIs and user-friendly tools, aims to create a robust and efficient environment for all types of batch workloads on Kubernetes.

Demo / Proof of Concept

▶ Watch: Solution: Kueue's topology-aware scheduling optimizes workloads (8:00)

While the presentation did not include a live demonstration or a real-time proof-of-concept, Marcin Wielgus effectively conveyed the functionality and benefits of Kueue's features through detailed architectural diagrams and illustrative scenarios. The use of visual examples, such as the data center layout for topology-aware scheduling and the GPU allocation scenarios for fair sharing, served to clearly explain the underlying problems and Kueue's proposed solutions. These conceptual "demos" provided a strong understanding of how these advanced scheduling and resource management capabilities operate in practice, demonstrating their value without requiring a live system interaction.

Defensive Implications

▶ Watch: Introducing fair sharing and hierarchical quotas concept (9:00)

The advancements presented by the Batch Working Group, particularly through Kueue and the new APIs, offer significant defensive implications for cluster administrators and organizations running batch workloads on Kubernetes.

Firstly, optimized resource utilization through topology-aware scheduling and fair sharing directly translates to cost savings and improved performance. By ensuring workloads are placed efficiently to minimize network bottlenecks, organizations can make the most of expensive specialized hardware like GPUs and TPUs, preventing them from being underutilized or, worse, wasted due to poor scheduling. Fair sharing mechanisms prevent resource hoarding and ensure that no single team or project starves others, fostering a more collaborative and productive environment.

Secondly, enhanced reliability and predictability for batch jobs are crucial. The updates to the Job API, with features like pod failure policy and success policy, empower administrators to define more robust job behaviors, reducing manual intervention and improving the overall stability of batch processes. JobSet further extends this by providing unified management for complex, multi-component jobs, simplifying their deployment, monitoring, and recovery.

Thirdly, simplified operational management reduces the burden on cluster administrators. KueueCTL provides a dedicated command-line interface for managing Kueue-specific resources, streamlining day-to-day operations. K-job addresses a common pain point for researchers and data scientists by offering templated job submission and a Slurm-like experience, abstracting away much of the Kubernetes YAML complexity. This reduces the learning curve for new users and minimizes errors, allowing administrators to focus on higher-level cluster health and security.

Finally, stronger resource governance through hierarchical quotas allows organizations to enforce policies that align with their internal structures and priorities. This ensures that critical projects receive the necessary resources while maintaining fairness across the board. By providing these tools, the Batch Working Group helps organizations transform Kubernetes into a more resilient, efficient, and user-friendly platform for their most demanding computational tasks.

Key Takeaways

  • Kueue is the central component for Kubernetes batch workload management: It acts as a workload-level scheduler and resource quota manager, supporting gang scheduling and being hardware/cloud neutral.
  • Topology-aware scheduling optimizes resource utilization: By co-locating interdependent pods, Kueue reduces network bottlenecks and maximizes the efficiency of expensive hardware like GPUs.
  • Fair sharing mechanisms ensure equitable resource distribution: Both preemption-based and the upcoming admission-based fair sharing (which uses historical usage) prevent resource starvation and optimize cluster capacity.
  • Hierarchical quotas provide flexible resource governance: Organizations can define quotas that mirror their internal structures, with unused capacity flowing up the hierarchy for redistribution.
  • New APIs and tools simplify complex batch workflows: JobSet offers unified management for groups of jobs, while K-job provides reusable templates and a Slurm-like CLI for easier job submission, reducing operational overhead.
  • The Batch Working Group is actively evolving the Kubernetes ecosystem: Through continuous development and community feedback, they are addressing fragmentation and enhancing Kubernetes' capabilities for HPC, AI/ML, and data analytics.

About the Speaker(s)

Marcin Wielgus is an organizer of the Kubernetes Batch Working Group at Google. In this capacity, he plays a pivotal role in steering the development and enhancements of Kubernetes' capabilities for batch workloads, encompassing areas like high-performance computing (HPC), AI, machine learning, and data analytics. His work focuses on reducing fragmentation within the Kubernetes ecosystem and creating robust tools and APIs, such as Kueue, to better support these demanding computational tasks.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This KubeCon update from the Kubernetes Batch Working Group, led by Marcin Wielgus, delivers a highly substantive deep dive into Kueue and related APIs, directly addressing the long-standing challenges of running HPC, AI/ML, and data analytics workloads on Kubernetes. The talk details advanced scheduling mechanisms like topology-aware placement, sophisticated fair-sharing models (both preemption and admission-based), and hierarchical quotas, alongside significant enhancements to the core Job API, JobSet, and the new Slurm-like K-job tool. It provides actionable insights and critical signal for anyone serious about optimizing demanding computational tasks on Kubernetes.

Heather Calloway (CISO) — STRONG ACCEPT

This KubeCon update from the Kubernetes Batch Working Group, led by Marcin Wielgus, presents critical advancements in managing high-performance computing, AI/ML, and data analytics workloads on Kubernetes. The introduction of Kueue's enhanced features like topology-aware scheduling, fair sharing, and hierarchical quotas, alongside updates to core APIs, directly addresses institutional challenges of resource optimization, cost efficiency, and governance for expensive compute assets. While deeply technical, the implications for operational resilience and strategic execution are significant, providing a robust framework for organizations to manage their most demanding computational tasks more…

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025