SIG-Multicluster Intro and Deep... Jeremy Olmsted-Thompson, Laura Lorenz, Stephen Kitt & Ryan Zhang
Jeremy Olmsted-Thompson, Laura Lorenz, Stephen Kitt, Ryan Zhang
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
The "SIG-Multicluster Intro and Deep Dive" talk at KubeCon EU provided a comprehensive update on the essential work being done within the SIG-Multicluster community to address the escalating complexities of managing Kubernetes workloads across multiple clusters. Presented by co-chairs Jeremy Olmsted-Thompson and Stephen Kitt, alongside Ryan Zhang and Laura Lorenz, the session highlighted the foundational challenges posed by Kubernetes' original design—where the cluster was considered the "boundary of the universe"—and how the SIG is systematically introducing APIs and frameworks to enable true multicluster capabilities. The speakers underscored the increasing ubiquity of multicluster deployments driven by needs such as fault tolerance, data locality, advanced policy enforcement, and dynamic capacity chasing, especially in the context of the current AI boom and the demand for scarce compute resources.

Key moments
- 0:00 SIG-Multicluster introduction and agenda overview
- 0:45 Why multicluster is essential today: AI, capacity chasing
- 1:45 Kubernetes' cluster-centric design limitations for multicluster
- 2:50 SIG-Multicluster's approach: universal APIs, consistent building blocks
- 3:25 Understanding the core concept: the Cluster Set
- 4:30 Introducing the About API for cluster self-identification
- 7:55 Introduction to the new Cluster Profile API
SIG-Multicluster Intro and Deep Dive
Speakers: Jeremy Olmsted-Thompson, Co-chair, SIG-Multicluster, GKE, Google; Laura Lorenz, GKE, Google; Stephen Kitt, Co-chair, SIG-Multicluster, Red Hat; Ryan Zhang, Microsoft
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=-SFVDr3wQ_w
Overview
The "SIG-Multicluster Intro and Deep Dive" talk at KubeCon EU provided a comprehensive update on the essential work being done within the SIG-Multicluster community to address the escalating complexities of managing Kubernetes workloads across multiple clusters. Presented by co-chairs Jeremy Olmsted-Thompson and Stephen Kitt, alongside Ryan Zhang and Laura Lorenz, the session highlighted the foundational challenges posed by Kubernetes' original design—where the cluster was considered the "boundary of the universe"—and how the SIG is systematically introducing APIs and frameworks to enable true multicluster capabilities. The speakers underscored the increasing ubiquity of multicluster deployments driven by needs such as fault tolerance, data locality, advanced policy enforcement, and dynamic capacity chasing, especially in the context of the current AI boom and the demand for scarce compute resources.
This talk is crucial for anyone operating or planning to operate Kubernetes at scale, particularly those grappling with hybrid cloud, multi-cloud, or edge deployments. The speakers articulated SIG-Multicluster's pragmatic approach: focusing on core problems, developing standardized APIs that function across diverse environments, and creating composable building blocks that integrate seamlessly with existing Kubernetes paradigms. The session served as both an educational primer on current multicluster solutions and a call to action for community involvement, emphasizing that user feedback and real-world use cases are vital for shaping the future of multicluster Kubernetes. The ambitious goal is to move beyond mere cluster provisioning to enabling applications designed from the ground up to leverage multiple clusters effectively.
Background
▶ Watch: SIG-Multicluster introduction and agenda overview (0:00)
Kubernetes, in its initial conception, treated each cluster as a self-contained universe, lacking inherent mechanisms for self-identification or seamless interaction with other clusters. This design philosophy, while simplifying initial deployments, presented significant hurdles as organizations scaled their operations and faced demands for higher availability, geographical distribution, and optimized resource utilization. The speakers noted that the need for multicluster environments has been a recurring theme at KubeCon for years, but the current landscape, particularly with the surge in AI and specialized hardware, has intensified the urgency to enable workloads to intelligently find and utilize available capacity wherever it exists.
Historically, efforts to address multicluster challenges within the Kubernetes ecosystem have seen varying degrees of success. An example cited was Kubefed, which, while an early attempt, was acknowledged as "probably not the most success story in the SIG's history." This past experience has informed SIG-Multicluster's current approach, which prioritizes the development of open standards and APIs that are broadly applicable and community-driven, rather than monolithic solutions. A core conceptual building block underpinning much of SIG-Multicluster's work is the cluster set. This pattern defines a group of clusters governed by a single authority (e.g., an organization or platform team) with a high degree of trust among them. Within a cluster set, the principle of namespace sameness is paramount, meaning a namespace with the same name across different clusters is used for the same purpose, sharing permissions and workload types, thereby moving towards a "cattle, not pets" mentality for managing clusters. This foundational concept ensures consistency and simplifies management across the "multiverse" of Kubernetes clusters.
Key Findings
▶ Watch: Kubernetes' cluster-centric design limitations for multicluster (1:45)
The talk highlighted several key APIs and projects that represent significant advancements in enabling robust multicluster operations:
- About API: This fundamental API addresses Kubernetes' initial lack of cluster self-identification. It provides a means to assign a unique name to each cluster and, more importantly, to attach well-known, community-driven properties to describe the cluster. These properties are crucial for informed decision-making in multicluster orchestration, such as placing AI jobs based on resource availability, cost, or specific hardware capabilities.
- Cluster Profile API: Building on the About API, the Cluster Profile API offers a read-only representation of a cluster's properties. Its primary purpose is to enable "cattle ranger" scenarios, allowing a central cluster manager to understand and make scheduling or orchestration decisions across a fleet of clusters. It is designed to integrate with popular multicluster-enabled third-party tools like Argo CD, Flux, and Kube-MultiQ, providing a common language for cluster interaction. A key challenge being addressed is secure credential management for accessing represented clusters.
- Multicluster Services API (MCS API): This API is a cornerstone for enabling east-west traffic between services residing in different clusters within a cluster set. It allows a standard Kubernetes service object to be flagged for exposure across clusters, making it consumable by workloads regardless of their physical cluster location. The MCS API is undergoing continuous refinement (currently at v1 alpha 2) and boasts multiple implementations, including Cilium and Istio.
- Gateway API Integration: A critical development is the tight integration between the MCS API and the Gateway API. While MCS focuses on east-west service discovery, the Gateway API handles north-south traffic, managing external client access to services within or across clusters. This synergy allows for comprehensive traffic management, enabling external clients to reach multicluster services seamlessly.
- Multicluster Runtime: This new project extends the existing controller runtime framework to support multicluster use cases. It aims to provide a drop-in solution, allowing developers to build multicluster controllers with familiar tooling, further integrating multicluster capabilities into the existing Kubernetes development paradigm. It also strongly integrates with the Cluster Profile API to access cluster inventory.
- Work API: Mentioned as an API that works in conjunction with the Cluster Profile API, the Work API is responsible for actually assigning jobs and workloads to specific clusters based on the properties and decisions made by an orchestrator.
Technical Deep Dive
▶ Watch: SIG-Multicluster's approach: universal APIs, consistent building blocks (2:50)
The technical core of SIG-Multicluster's work revolves around creating a robust set of APIs that abstract away the underlying complexity of distributed Kubernetes environments.
The About API is foundational. Recognizing that Kubernetes clusters initially lacked a distinct identity beyond their internal scope, the About API introduces a mechanism to name each cluster, transforming the "universe" into a "multiverse." Beyond a simple name, the API is being enhanced to include a richer set of properties. These properties can describe various characteristics of a cluster, such as its geographic region, cloud provider, available hardware (e.g., GPUs for AI workloads), cost profile, or specific compliance certifications. For instance, an orchestrator looking to deploy an AI job can query the About API to find clusters with specific GPU types and then filter by those offering the lowest compute cost. This moves beyond simplistic naming to providing actionable metadata for intelligent workload placement.
Building upon the About API, the Cluster Profile API serves as a read-only, declarative representation of a cluster within the multicluster management plane. It doesn't grant direct access but reflects the properties exposed by the About API of individual clusters. The concept here is similar to a "cattle ranger"—a central management entity that views and understands the capabilities of its entire fleet without needing to directly interact with each cluster's control plane for basic information. The diagram presented visually depicts yellow "real clusters" on the bottom layer, blue "cluster profiles" representing them on a middle layer, and a "cluster manager" on top. This manager, often a third-party tool, can then use this unified API to interact with the multicluster environment. A critical ongoing development for the Cluster Profile API is addressing credential management. Storing secrets for cluster access directly within an API object raises significant security concerns. The SIG is actively exploring secure, compliant methods to enable the cluster manager to authenticate and interact with remote clusters without compromising security. The goal is a common language that popular tools like Argo CD, Flux, and Kube-MultiQ can adopt to orchestrate deployments and manage resources across clusters, avoiding the need for individual integrations with each cluster type. The Cluster Profile API also integrates with the Work API, which is then used to actually dispatch and assign workloads to the clusters identified and profiled.
For inter-cluster communication, the Multicluster Services API (MCS API) is paramount. It standardizes how services, familiar to any Kubernetes user, can be extended to span multiple clusters. Instead of being confined to a single cluster, a service can be declared as a "multicluster service," allowing its endpoints to exist and be discovered across a cluster set. This enables seamless east-west traffic where a pod in one cluster can consume a service provided by pods in another cluster as if it were local. The MCS API is designed to be implementation-agnostic, with projects like Cilium and Istio providing concrete examples of its adoption. The ongoing v1 alpha 2 development reflects continuous feedback from these implementations to refine the API. The integration with the Gateway API is a sophisticated architectural decision. The Gateway API, traditionally focused on north-south traffic (external clients accessing internal services), now leverages MCS. A Gateway can be configured to route traffic to a multicluster service, effectively creating a global load balancer that can direct external requests to the optimal backend across different clusters, abstracting the multicluster topology from the end-user. This is detailed in a GE gateway enhancement proposal, highlighting the collaborative work between SIG-Multicluster and SIG-Network.
Finally, the Multicluster Runtime project extends the familiar controller runtime framework. This is crucial for developers building custom controllers that need to operate across multiple clusters. By providing a drop-in solution, it lowers the barrier to entry for building multicluster-aware applications, allowing controllers to reconcile resources across a cluster set rather than just within a single cluster. This project is still an experiment hosted within the SIG, seeking community testers and feedback.
Looking ahead, the SIG is also exploring advanced topics like multicluster coordination through leader election, which becomes critical when controllers need to ensure unique operations across distributed environments. Further, complex networking challenges like network policy across clusters are on the roadmap, building on existing SIG-Network work such as KE 4444, which allows specifying local vs. remote service preferences.
Demo / Proof of Concept
▶ Watch: Introducing the About API for cluster self-identification (4:30)
While this specific "Intro and Deep Dive" talk did not feature a live demonstration or proof of concept by the speakers, it heavily referenced numerous real-world implementations and community efforts that serve as concrete validations of SIG-Multicluster's APIs. The speakers emphasized that the SIG's goal is to ship standards and APIs, not necessarily implementations, but they actively track and encourage adoption.
Several examples were cited:
- Cilium's MCS API Implementation: Cilium, a popular CNI and eBPF-based networking solution, has significantly advanced its implementation of the Multicluster Services API. This demonstrates how a robust networking layer can leverage SIG-Multicluster standards to provide seamless service discovery and connectivity across clusters.
- Istio's Service Import/Export: The speakers confirmed that Istio's widely used service import and export mechanisms are, in fact, an implementation of the MCS API. This highlights how a prominent service mesh leverages the SIG's standards for its own multicluster capabilities, affirming the API's effectiveness and broad applicability.
- Multicluster Operator (MCO): Mentioned during the Q&A, the MCO is actively working on integrating with SIG-Multicluster standards. Specifically, it intends to utilize the About API and prioritize integration with the Cluster Profile API. While MCO might initially handle credentials independently, the long-term goal is to align with the SIG's evolving solutions for secure credential management. This illustrates how existing operators are adapting to and benefiting from the standardization efforts.
- Third-Party Integrations: The Cluster Profile API is explicitly designed for integration with popular multicluster-enabled tools such as Argo CD, Flux, and Kube-MultiQ. The vision is for these tools to use the Cluster Profile API as a common interface to understand and orchestrate workloads across diverse cluster fleets, streamlining GitOps and workload distribution workflows.
- Other KubeCon Talks: The speakers alluded to numerous other multicluster talks at the same KubeCon, including presentations on OCM, kube fleet, and cluster ADM, many of which showcase the practical application and integration of SIG-Multicluster APIs. Jeremy Olmsted-Thompson specifically noted that GKE also has its own implementations leveraging these patterns.
These examples collectively serve as strong proof points for the utility and adoption of the APIs developed by SIG-Multicluster, demonstrating their ability to enable diverse multicluster solutions across various platforms and tools.
Defensive Implications
▶ Watch: Introduction to the new Cluster Profile API (7:55)
The work of SIG-Multicluster has profound defensive implications for organizations operating Kubernetes at scale, primarily by establishing standardized, well-understood patterns for what was previously a fragmented and often insecure domain.
Firstly, the concept of a cluster set with namespace sameness provides a critical security baseline. By dictating that namespaces with identical names across trusted clusters serve the same purpose and share similar permissions, it significantly reduces the risk of privilege escalation or unintended access across cluster boundaries. This standardization helps platform teams enforce consistent policy and access controls, preventing "shadow IT" or ad-hoc configurations that often introduce vulnerabilities.
The About API and Cluster Profile API contribute to a more secure posture by providing a clear, machine-readable inventory of cluster capabilities and attributes. This visibility is essential for compliance and auditing. Knowing precisely where resources are deployed, their configurations, and their specific properties allows security teams to verify that sensitive workloads are only placed on appropriately secured and compliant clusters. The active work on credential management for the Cluster Profile API is a direct response to a major security concern, aiming to provide a standardized and secure method for cluster managers to authenticate with remote clusters, avoiding the proliferation of insecurely stored secrets.
The Multicluster Services API (MCS API) and its integration with the Gateway API enhance resilience and security of communication. By abstracting service endpoints across clusters, organizations can build highly available and geographically distributed applications that are more resistant to single-cluster failures. The standardized approach to east-west and north-south traffic management means that security policies (e.g., network segmentation, TLS enforcement) can be applied consistently across the entire multicluster environment, rather than requiring disparate configurations for each cluster or communication path. This reduces the attack surface and simplifies the enforcement of Zero Trust principles.
Future work on leader election for multicluster coordination directly addresses potential race conditions or inconsistent states that could arise from multiple controllers operating across a distributed system. Ensuring a single, coordinated point of control for critical operations is vital for maintaining data integrity and preventing security flaws stemming from uncoordinated actions. Similarly, the eventual standardization of network policy across clusters, building on efforts like KE 4444, will be a game-changer. It promises to enable fine-grained network segmentation and access control not just within a cluster, but consistently across the entire multicluster fleet, which is a significant security challenge today.
By providing common APIs and frameworks, SIG-Multicluster reduces the need for organizations to develop bespoke, often less secure, solutions for multicluster management. This shift towards community-vetted standards inherently improves the overall security posture of Kubernetes deployments, making them more manageable, observable, and defensible against evolving threats.
Key Takeaways
- Addressing Kubernetes' Multicluster Gap: Kubernetes was not originally designed for multicluster operations, necessitating the development of new APIs and frameworks to support modern use cases like capacity chasing, fault tolerance, and AI workload distribution.
- Foundational APIs for the "Multiverse": SIG-Multicluster is building essential APIs like the About API (for cluster identity and properties), Cluster Profile API (for read-only cluster representation and orchestration), and Work API (for workload assignment) to enable intelligent multicluster management.
- Standardizing Cross-Cluster Communication: The Multicluster Services API (MCS API) provides a standard way to expose services across clusters (east-west traffic), seamlessly integrating with the Gateway API for external access and global load balancing (north-south traffic).
- Extending Familiar Frameworks: New projects like Multicluster Runtime aim to extend existing Kubernetes development tools, such as controller runtime, to support multicluster operations, making it easier for developers to build multicluster-aware applications.
- Focus on Interoperability and Ecosystem: The SIG prioritizes creating APIs that work across diverse environments (cloud, on-prem, multi-cloud) and integrate with existing tools like Argo CD, Flux, Cilium, and Istio, fostering a robust and standardized multicluster ecosystem.
- Community-Driven Development and Future Work: The success of SIG-Multicluster relies heavily on community input for defining canonical patterns, refining APIs (e.g., secure credential management for Cluster Profile), and tackling future challenges like multicluster leader election and network policy.
About the Speaker(s)
The "SIG-Multicluster Intro and Deep Dive" was presented by a team of highly experienced individuals deeply embedded in the Kubernetes and multicluster ecosystems:
- Jeremy Olmsted-Thompson is one of the co-chairs of SIG-Multicluster and works on GKE at Google. His involvement highlights Google's commitment to advancing multicluster capabilities within Kubernetes and leveraging these advancements in their managed service offerings.
- Stephen Kitt is the other co-chair of SIG-Multicluster and works on multicluster initiatives at Red Hat. His role underscores Red Hat's significant contributions to the open-source Kubernetes community and its focus on enterprise-grade multicluster solutions.
- Ryan Zhang is from Microsoft, where he also works on multicluster technologies. His presence signifies Microsoft's investment in multicluster Kubernetes, likely influencing offerings such as Azure Kubernetes Service (AKS).
- Laura Lorenz works at Google on GKE, contributing to the development and implementation of Kubernetes features for Google's cloud platform. Her expertise enriches the SIG's work with practical insights from a major cloud provider.
Collectively, these speakers represent leading technology companies and bring a wealth of experience from both cloud-native development and large-scale Kubernetes operations to the SIG-Multicluster community.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This was a critical deep dive into the foundational work of SIG-Multicluster, addressing Kubernetes' inherent limitations in distributed environments. The speakers, core contributors from major cloud providers, meticulously detailed novel APIs like About, Cluster Profile, and Multicluster Services (MCS), explaining their technical underpinnings and integration with existing frameworks like Gateway API and controller runtime. For anyone grappling with large-scale, hybrid, or multi-cloud Kubernetes deployments, particularly with the escalating demands of AI workloads, this session provides indispensable, actionable intelligence and a clear roadmap for standardized, secure, and scalable…
Heather Calloway (CISO) — STRONG ACCEPT
This KubeCon session on SIG-Multicluster outlines critical advancements for securing and managing Kubernetes at scale, directly addressing enterprise resilience and governance concerns. By standardizing APIs for cluster identity, inter-cluster communication, and workload orchestration, the SIG is providing a much-needed framework to enforce consistent security policies and achieve clearer risk ownership across distributed environments. This work moves beyond technical elegance to deliver tangible operational and compliance benefits for organizations grappling with complex hybrid and multi-cloud Kubernetes deployments.