Project Lightning Talk: Extend Large Language Model Training Beyond Single Kubernetes Cl... Klaus Ma
Klaus Ma
KubeCon + CloudNativeCon Europe 2025 · Project Lightning Talk
Overview
In this concise yet impactful lightning talk at KubeCon EU, Klaus Ma, a prominent figure in the Kubernetes community and founder of the Volcano project, addressed the critical challenges associated with training Large Language Models (LLMs) within Kubernetes environments. The talk, titled "Extend Large Language Model Training Beyond Single Kubernetes Cl...", highlighted the inherent limitations of single-cluster Kubernetes deployments when confronted with the immense computational and data requirements of modern LLMs. Ma introduced Volcano's strategic initiatives to transcend these boundaries, focusing on multi-cluster federation, enhanced resource utilization through network-aware scheduling, and improved integration between AI frameworks and the underlying infrastructure.

Key moments
- 0:00 Introduction: Scaling LLM training beyond single clusters
- 0:30 Challenge: Single Kubernetes cluster scalability limits
- 1:30 Challenge: Frameworks lack infrastructure layer communication
- 2:00 Solution: Unified API for multi-cluster federation
- 3:00 Solution: Handling cross-cluster scheduling challenges
- 3:30 Solution: Network-aware scheduling for resource utilization
- 4:00 Solution: Meta-framework for infra-framework communication
- 4:30 Reference links and project information
Project Lightning Talk: Extend Large Language Model Training Beyond Single Kubernetes Cl... Klaus Ma
Speakers: Klaus Ma, Founder of Volcano Project, Former Co-Chair of SIG Scaling
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=BlzHv9KV1Z4
Overview
In this concise yet impactful lightning talk at KubeCon EU, Klaus Ma, a prominent figure in the Kubernetes community and founder of the Volcano project, addressed the critical challenges associated with training Large Language Models (LLMs) within Kubernetes environments. The talk, titled "Extend Large Language Model Training Beyond Single Kubernetes Cl...", highlighted the inherent limitations of single-cluster Kubernetes deployments when confronted with the immense computational and data requirements of modern LLMs. Ma introduced Volcano's strategic initiatives to transcend these boundaries, focusing on multi-cluster federation, enhanced resource utilization through network-aware scheduling, and improved integration between AI frameworks and the underlying infrastructure.
The core motivation behind this work stems from the insatiable demand for computational resources for LLM training, often exceeding the practical scalability limits of a single Kubernetes cluster, typically capped around 5,000 nodes. This talk is highly relevant to organizations pushing the frontiers of AI, particularly those leveraging Kubernetes for their machine learning operations (MLOps), as it directly tackles the bottlenecks preventing efficient and scalable LLM development. By extending Kubernetes capabilities to span multiple clusters and optimize resource allocation at a foundational level, Volcano aims to unlock unprecedented efficiencies and performance for the next generation of AI workloads.
Background
▶ Watch: Introduction: Scaling LLM training beyond single clusters (0:00)
The rapid evolution of Large Language Models (LLMs) has placed unprecedented demands on computational infrastructure. Training these models, which can involve billions or even trillions of parameters, requires vast amounts of GPU hardware, high-bandwidth networking, and robust storage solutions. While Kubernetes has emerged as a de facto standard for orchestrating containerized workloads, its traditional architecture presents several significant hurdles when applied to the extreme requirements of LLM training.
Firstly, single-cluster scalability is a primary concern. Standard Kubernetes installations are often optimized for a maximum of 5,000 nodes, a limit that can quickly become a bottleneck for organizations attempting to provision the hundreds or thousands of GPUs necessary for cutting-edge LLMs. This constraint forces compromises, either by limiting the scale of models or by complex, manual management of fragmented resources across multiple independent clusters, leading to operational overhead and reduced efficiency.
Secondly, resource utilization remains a persistent challenge. Even with sufficient GPU and network resources, achieving optimal utilization is difficult without intelligent scheduling. Traditional Kubernetes schedulers may not possess the necessary context about the specific communication patterns and resource interdependencies inherent in distributed LLM training jobs. This can lead to suboptimal placement of workloads, increased inter-node communication latency, and underutilized hardware, directly impacting training speed and cost-effectiveness.
Thirdly, there's a significant disconnect between AI frameworks and the infrastructure layer. Modern distributed AI training frameworks, such as PyTorch Distributed or TensorFlow, often operate with their own internal communication mechanisms and resource assumptions. However, they typically do not export detailed information about their specific networking requirements, data exchange patterns, or real-time resource needs to the underlying Kubernetes infrastructure. This lack of visibility prevents the Kubernetes scheduler, including specialized schedulers like Volcano, from making truly informed decisions that could significantly enhance performance and resource utilization. The infrastructure is left to "guess" optimal placement, leading to inefficiencies.
Finally, the concept of Kubernetes federation or multi-cluster management is not new, with various initiatives attempting to address it over the past five to ten years. However, these efforts have often faced challenges related to API compatibility, networking complexities, and consistent scheduling across disparate clusters. The Volcano project, as highlighted by Klaus Ma, aims to build upon these past efforts, providing a more robust and unified approach specifically tailored for high-performance computing (HPC) and AI workloads. By tackling these foundational issues, Volcano seeks to provide a more scalable, efficient, and intelligent platform for advancing LLM research and deployment.
Key Findings
▶ Watch: Challenge: Frameworks lack infrastructure layer communication (1:30)
The talk outlined several key findings and strategic initiatives undertaken by the Volcano project to address the aforementioned challenges in scaling LLM training across Kubernetes clusters. These findings represent Volcano's core contributions and architectural directions:
- Unified API for Multi-Cluster and Single-Cluster Management: A critical finding was the need for a simplified and consistent interface for managing workloads, regardless of whether they span multiple clusters or reside within a single one. Volcano addresses this with a distinct yet integrated approach: Volcano Global for federation cases and Volcano Core for single-cluster scenarios. This unified API aims to streamline the developer experience and reduce the complexity associated with multi-cluster deployments, a common pitfall in previous federation attempts.
- Specialized Focus on East-West Networking Management: While Volcano itself does not directly handle the low-level networking and storage primitives (delegating these to other specialized projects), a key finding is the necessity for powerful management capabilities specifically for east-west networking. This refers to traffic between pods and services within and across clusters. For LLM training, where large data transfers and frequent communication between distributed training components are common, efficient and high-performance east-west networking is paramount. Volcano's role here is to ensure that other projects handling networking can be effectively integrated and managed for optimal performance.
- Enhanced Cross-Cluster Scheduling with Shared Information: A major obstacle in multi-cluster environments is the lack of information exchange between individual clusters, leading to suboptimal scheduling decisions. Volcano's approach acknowledges that a single, cohesive project (Volcano itself) can possess "enough information" to make intelligent scheduling decisions across clusters. This implies a federated scheduling logic that can consider global resource availability, workload requirements, and inter-cluster dependencies, moving beyond isolated scheduling decisions.
- Network-Aware Scheduling for Improved Resource Utilization: The talk identified network-aware scheduling as a primary mechanism for improving resource utilization, particularly for network-intensive LLM workloads. By incorporating network topology, bandwidth, and latency information into its scheduling algorithms, Volcano can intelligently place pods and training jobs to minimize communication overhead and maximize throughput. This is crucial for LLMs, where collective communication operations (e.g., all-reduce) are often performance bottlenecks.
- Meta Framework for Infrastructure-Framework Communication: A significant innovation highlighted is the introduction of a meta framework designed to bridge the gap between AI training frameworks (like PyTorch or TensorFlow) and the Kubernetes infrastructure. This framework allows workload-specific information—such as networking requirements, storage needs, and communication patterns—to be explicitly exported from the application layer to the underlying infrastructure layer (Volcano/Volcano Global). This direct communication empowers Volcano to make far more intelligent and optimized placement decisions, leading to better resource utilization and superior training performance compared to traditional, infrastructure-agnostic scheduling.
Technical Deep Dive
▶ Watch: Solution: Handling cross-cluster scheduling challenges (3:00)
The technical solutions proposed by Klaus Ma and the Volcano project address the intricate challenges of scaling LLM training in Kubernetes through a multi-pronged approach that re-architects how distributed AI workloads interact with the underlying infrastructure.
At the core of Volcano's strategy for multi-cluster management is the unified API. This is achieved by segmenting the project into Volcano Global for managing federated, multi-cluster scenarios and Volcano Core for handling single-cluster operations. The critical design principle here is to provide a consistent user experience and API surface, abstracting away the underlying complexity of whether a job is deployed to one cluster or spread across many. This consistency is vital for developers who need to define their LLM training jobs without being burdened by the specific topology of the compute infrastructure. Volcano Global likely acts as a control plane that aggregates information from multiple Volcano Core instances, allowing a single job definition to be intelligently distributed and scheduled across available resources in different clusters. This approach mitigates the API backward compatibility challenges that have plagued previous Kubernetes federation attempts by offering a native, unified interface.
For networking and storage, Volcano adopts a pragmatic approach: it "dedicates this part to other projects." This means Volcano itself does not reinvent the wheel for CNI (Container Network Interface) or CSI (Container Storage Interface) implementations. Instead, it focuses on integrating with existing, powerful solutions. The emphasis, however, is on providing "really powerful management for the east-west networking part." For LLMs, east-west traffic, which involves data exchange between distributed training processes (e.g., gradient synchronization, model parallelism, data parallelism), is often the performance bottleneck. Volcano's role here is not to implement the network itself but to intelligently leverage and manage the network resources provided by the underlying CNI. This implies that Volcano's scheduler will be aware of network topology, link speeds, and potential congestion points, allowing it to make placement decisions that minimize network latency and maximize bandwidth utilization for critical communication paths.
The concept of scheduling across clusters is fundamentally enhanced by Volcano's ability to maintain a more comprehensive view of the entire distributed environment. Unlike traditional Kubernetes schedulers that operate within the confines of a single cluster and lack information about other clusters, Volcano aims to overcome this by acting as a more centralized orchestrator for federated workloads. The speaker notes that "this two cluster didn't have information with each other. So in a single project for example for kino we have enough information to handle this part." This suggests that Volcano Global will gather and maintain a holistic state of resource availability, current workload distribution, and network characteristics across all participating clusters. This global perspective enables Volcano to make intelligent decisions about where to place parts of an LLM training job (e.g., specific workers, parameter servers) to optimize for resource availability, network proximity, and cost.
A cornerstone of Volcano's approach to improving resource utilization is network-aware scheduling. For LLMs, communication patterns are highly predictable and often involve large, synchronous data transfers. A scheduler that understands the network topology—which nodes are physically close, which network switches connect them, what the available bandwidth and latency are—can make significantly better placement decisions. For instance, it can co-locate interdependent processes on nodes within the same rack or connected by high-speed interconnects, thereby reducing communication latency and improving overall training throughput. This goes beyond simple CPU/GPU availability and considers the network as a first-class resource. Volcano will integrate information about network topology and performance metrics into its scheduling algorithms to achieve this, ensuring that LLM workloads are not just placed, but optimally placed.
Perhaps the most forward-looking technical contribution is the introduction of a meta framework that facilitates explicit communication between AI training frameworks and the infrastructure layer. Currently, frameworks like PyTorch or TensorFlow largely operate as black boxes from the infrastructure's perspective. They might request GPUs, but they don't typically inform Kubernetes about their specific communication graphs, the volume of data they expect to exchange between specific processes, or their sensitivity to network latency. The meta framework aims to change this by allowing AI frameworks to "change all the information networking storage of the workload to the infrastructure layer." This means that an LLM training job could declare, for example, "this group of 8 GPUs needs extremely low-latency, high-bandwidth communication with each other," or "this worker needs to read large datasets from a specific storage location." By exposing these critical details, Volcano (and Volcano Global) can use this rich context to "place the workload better," leading to dramatically improved resource utilization and performance. This represents a paradigm shift from infrastructure-agnostic scheduling to intelligent, application-aware orchestration, tailored specifically for the demanding requirements of LLM training.
Demo / Proof of Concept
▶ Watch: Solution: Network-aware scheduling for resource utilization (3:30)
Klaus Ma's lightning talk primarily focused on outlining the architectural vision and strategic direction of the Volcano project for multi-cluster LLM training. While the speaker mentioned that more details, including a user case and practical demonstrations of network-aware scheduling, would be presented "tomorrow" (referring to another session at the conference), this specific talk did not include an live demonstration or a detailed walkthrough of a proof of concept. The content was geared towards introducing the problems and Volcano's high-level solutions.
Defensive Implications
▶ Watch: Reference links and project information (4:30)
While the talk primarily focuses on scaling and optimizing Large Language Model training, the "defensive implications" can be interpreted in the context of building a robust, resilient, and cost-efficient infrastructure that can "defend" against common operational challenges and potential failures in large-scale AI environments.
Firstly, by addressing single-cluster scalability limitations and enabling multi-cluster federation, Volcano helps organizations defend against the risk of hitting hard infrastructure ceilings. Without such capabilities, LLM training jobs might be forced into smaller, suboptimal configurations, or require complex, error-prone manual sharding across independent clusters. This federated approach provides a defensive posture against resource exhaustion, ensuring that critical AI workloads can always find sufficient compute, even if it means spanning geographically diverse data centers. This resilience is vital for maintaining continuous development cycles and meeting aggressive model training timelines.
Secondly, the focus on improving resource utilization through network-aware scheduling and the meta framework directly defends against operational inefficiencies and spiraling costs. Underutilized GPU hardware, a common problem in poorly scheduled distributed training jobs, translates directly into wasted capital expenditure and increased operational expenses. By intelligently placing workloads based on network topology and application-specific communication patterns, Volcano minimizes latency and maximizes throughput, thereby ensuring that expensive GPU resources are always working at their peak efficiency. This acts as a financial defense mechanism, allowing organizations to extract maximum value from their infrastructure investments.
Furthermore, the enhanced visibility provided by the meta framework—where AI frameworks explicitly share workload-specific networking and storage information with the infrastructure—can contribute to a more stable and predictable training environment. This defense against "unknown unknowns" in resource consumption can prevent unexpected performance degradation, resource contention, and even job failures that might arise from suboptimal resource allocation. By understanding the true needs of the application, Volcano can provision and manage resources more effectively, leading to more reliable training runs and reduced debugging time.
In essence, while not directly addressing cybersecurity threats, Volcano's advancements in multi-cluster management, resource optimization, and framework-infrastructure integration provide a robust operational defense. They safeguard against infrastructure bottlenecks, inefficient resource usage, and unpredictable performance—all of which are critical "threats" to the successful and timely development of large-scale AI models. A well-managed, scalable, and efficient infrastructure is inherently more resilient and easier to secure, providing a stronger foundation for sensitive and critical AI workloads.
Key Takeaways
- Single-Cluster Limits are Insufficient for LLMs: Traditional Kubernetes single-cluster scalability (around 5,000 nodes) is a major bottleneck for the immense GPU and network requirements of Large Language Model training.
- Volcano Offers Unified Multi-Cluster Management: The Volcano project provides a unified API through Volcano Global (for federation) and Volcano Core (for single clusters) to simplify the orchestration of LLM workloads across multiple Kubernetes clusters.
- Resource Utilization is Enhanced by Network Awareness: Volcano emphasizes network-aware scheduling, integrating network topology and performance data into its scheduling decisions to optimize workload placement, minimize communication latency, and maximize GPU utilization for distributed LLM training.
- Bridging the AI Framework-Infrastructure Gap: A novel meta framework allows AI training frameworks to explicitly share critical workload details (networking, storage, communication patterns) with the underlying Volcano infrastructure, enabling truly intelligent and application-aware scheduling.
- Operational Resilience and Cost Efficiency: By addressing scalability, resource utilization, and framework-infrastructure communication, Volcano helps organizations build more resilient, efficient, and cost-effective infrastructures for their demanding LLM development, mitigating risks associated with resource contention and suboptimal performance.
About the Speaker(s)
Klaus Ma is a significant contributor to the Kubernetes ecosystem, known for his work in high-performance computing and machine learning orchestration. He is the founder of the Volcano project, an open-source batch system built on Kubernetes, specifically designed for high-performance workloads like AI/ML, HPC, and big data. Prior to his current role, Klaus Ma also served as the co-chair of SIG Scaling, a Kubernetes Special Interest Group focused on improving the scalability of Kubernetes itself. His background and expertise, particularly from his affiliation with Nvidia, underscore his deep understanding of the challenges and requirements for running large-scale, resource-intensive workloads like LLM training on Kubernetes.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Klaus Ma's lightning talk on the Volcano project's approach to scaling LLM training beyond single Kubernetes clusters is a highly impactful and forward-thinking session. It effectively identifies critical bottlenecks in current MLOps infrastructure for large models and proposes a robust, unified architecture incorporating multi-cluster federation, network-aware scheduling, and a novel meta-framework for explicit application-infrastructure communication. While a lightning talk limits the depth of a live demo, the vision presented by a speaker of Ma's caliber is immensely valuable for anyone pushing the boundaries of AI infrastructure.
Heather Calloway (CISO) — STRONG ACCEPT
Klaus Ma's lightning talk on extending LLM training beyond single Kubernetes clusters addresses a critical operational bottleneck for any organization serious about AI. While technically deep, the core message is about institutional realism: recognizing hard infrastructure limits and designing solutions for scaled execution. Volcano's initiatives, particularly the unified API and the meta framework for infrastructure-framework communication, offer a clear path to managing the significant compute and network demands of modern AI, directly impacting business resilience, cost efficiency, and the ability to deliver on strategic AI investments. This isn't just about technical cleverness; it's…