Extending Kubernetes Resource Model (KRM) Beyond Kubernetes Work... Mangirdas Judeikis & Nabarun Pal

Mangirdas Judeikis, Nabarun Pal

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk, presented by Mangirdas Judeikis (MJ) and Nabarun Pal, delves into the evolution of the Kubernetes Resource Model (KRM), exploring how it can be extended beyond traditional Kubernetes workloads to address complex challenges in modern, multi-cloud, and hybrid environments. Titled "Extending Kubernetes Resource Model (KRM) Beyond Kubernetes Work... Mangirdas Judeikis & Nabarun Pal," the presentation introduces KCP (Kubernetes Control Plane) as a pivotal project pushing for the next generation of KRM, termed KRM++. The core premise is to leverage the declarative power and consistency of KRM for managing diverse infrastructure resources and APIs, moving away from fragmented, provider-specific solutions.

Watch on YouTube

Visual summary for Extending Kubernetes Resource Model (KRM) Beyond Kubernetes Work... Mangirdas Judeikis & Nabarun Pal by Mangirdas Judeikis, Nabarun Pal
Visual summary for Extending Kubernetes Resource Model (KRM) Beyond Kubernetes Work... Mangirdas Judeikis & Nabarun Pal by Mangirdas Judeikis, Nabarun Pal

Key moments

  1. 0:00 Introduction and Talk Agenda
  2. 1:10 Why building APIs for hybrid cloud is hard
  3. 2:50 KRM as a solution: Declarativeness and Consistency
  4. 4:00 Deep dive into KRM object structure and reconciliation
  5. 5:30 Common pitfalls when designing KRM APIs
  6. 7:00 Adhering to API conventions and leveraging schema validation

Extending Kubernetes Resource Model (KRM) Beyond Kubernetes Work... Mangirdas Judeikis & Nabarun Pal

Speakers: Mangirdas Judeikis, Maintainer at KCP, Casti; Nabarun Pal, Kubernetes Maintainer, KCP Contributor

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=y0JgZ-hQ-Bo

Overview

This talk, presented by Mangirdas Judeikis (MJ) and Nabarun Pal, delves into the evolution of the Kubernetes Resource Model (KRM), exploring how it can be extended beyond traditional Kubernetes workloads to address complex challenges in modern, multi-cloud, and hybrid environments. Titled "Extending Kubernetes Resource Model (KRM) Beyond Kubernetes Work... Mangirdas Judeikis & Nabarun Pal," the presentation introduces KCP (Kubernetes Control Plane) as a pivotal project pushing for the next generation of KRM, termed KRM++. The core premise is to leverage the declarative power and consistency of KRM for managing diverse infrastructure resources and APIs, moving away from fragmented, provider-specific solutions.

The speakers, both maintainers and contributors to the KCP project, highlight the inherent difficulties tech companies face when operating across multiple cloud providers, citing issues of API inconsistency, interoperability, and the overwhelming task of integrating disparate systems. They argue against the common tendency to "reinvent the wheel" by building custom solutions, advocating instead for the adoption of established patterns like KRM. The talk culminates in a demonstration of KCP's capabilities, illustrating how it enables the creation of custom APIs without relying solely on Custom Resource Definitions (CRDs), while still harnessing the robust Kubernetes API machinery.

This presentation is highly relevant for platform engineers, cloud architects, and developers grappling with multi-cloud complexity, internal platform building, and API governance. It offers a compelling vision for a standardized, multi-tenant control plane that abstracts away infrastructure differences, streamlines API management, and enhances developer experience by leveraging the familiar declarative paradigm of Kubernetes. By addressing the "many clusters problem" and providing a robust framework for API lifecycle management, KCP aims to empower organizations to build scalable and consistent platforms with greater efficiency.

Background

▶ Watch: Introduction and Talk Agenda (0:00)

The genesis of this discussion lies in a pervasive problem faced by tech companies operating in hybrid or multi-cloud environments: managing infrastructure resources like compute, networking, and storage across diverse providers. Each cloud provider presents its own unique set of APIs, leading to significant challenges in achieving consistency, interoperability, and manageable ecosystem integration. This often results in organizations building bespoke solutions, a process the speakers characterize as "reinventing standards" rather than focusing on core business value.

The solution, as proposed and widely adopted, is the Kubernetes Resource Model (KRM). KRM forms the fundamental API surface of Kubernetes, renowned for its declarativeness and consistent structure. Instead of issuing procedural commands, users declare their desired state for a resource (e.g., "I want a pod with this image and these resources"), and the system asynchronously works to achieve and maintain that state. Every Kubernetes object adheres to a consistent format, comprising an apiVersion, kind, metadata (name, annotations, labels), a spec (desired state), and a status (actual state reported by the system). This model, where asynchronous processes (controllers/reconcilers) continuously read objects and reconcile their desired state, is what makes KRM exceptionally powerful.

However, despite its power, designing and implementing KRM-based APIs, particularly through CRDs, is not without its pitfalls:

  1. Mixing Business Logic into CRD Design: A common mistake is embedding too much procedural or business logic directly into the spec field. The spec should primarily define the desired state, not the steps to achieve it. An example cited is the Service resource's clusterIP field, which, while user-specifiable, is often system-assigned, raising questions about whether it truly belongs in the spec versus being solely a status field.
  2. Ignoring Established API Conventions: Inconsistent naming (e.g., DBConnection vs. dbconnections) and improper pluralization (e.g., defining sheep with a plural sheeps instead of flock) can lead to confusion and hinder usability. Kubernetes has well-defined conventions that, if ignored, complicate API consumption.
  3. Underutilizing KRM Features: The Kubernetes API machinery offers powerful features like OpenAPI schema validation and CEL expression validations for custom resources. These tools allow developers to define precise validation rules, patterns, and constraints directly within the CRD schema, ensuring data integrity and preventing invalid configurations. Failing to leverage these capabilities undermines the robustness of the API.
  4. Incorrect API Responses: Returning a 200 OK status with a payload indicating a 400 Bad Request status code is an anti-pattern that misleads API consumers about the success or failure of an operation.
  5. Lack of Proper API Evolution and Versioning Strategy: As APIs evolve, adding new fields or making breaking changes requires a clear strategy for versioning and conversion. The KCP project itself faced this challenge with its Workspace API, where a minor breaking change led to a lengthy discussion thread (over 300 replies) about v1alpha1 vs. v1alpha2 and conversion mechanisms, underscoring the complexity of API evolution.
  6. Suboptimal Reconciliation Patterns: While Kubernetes provides tools for writing reconcilers, the approach has evolved. Older controllers might use manual workqueue management, whereas modern best practices advocate for controller-runtime. The talk emphasizes breaking down reconciliation logic into atomic reconcilers, each responsible for a specific, testable operation. This improves readability, testability, and error handling (e.g., distinguishing between recoverable transient errors and permanent failures). The committer pattern, where a controller ideally only modifies the status in a single reconcile loop and not the spec, is highlighted as a good practice.

The speakers argue that KRM is not a "16th standard" to be invented but rather the "same thing you have been using all throughout" as consumers of Kubernetes and other CNCF projects. The goal is to stick to this widely adopted, consistent flow rather than creating new, proprietary solutions.

Key Findings

▶ Watch: KRM as a solution: Declarativeness and Consistency (2:50)

The talk identifies a critical challenge for organizations scaling their Kubernetes adoption: the "many clusters problem." As platform teams grow, they often end up with a sprawling architecture of multiple Kubernetes clusters, each with its own operators, leading to a complex "reconciliation mesh" where cross-cluster communication and API dependencies become cumbersome and error-prone. While KRM as an API standard is lauded for its declarative power and consistency, its inherent ability to provide fragmented tenancy isolation and centralized API management across these numerous clusters is deemed insufficient.

Traditional alternatives, such as running vclusters or a "generic control plane" like the KCP-incubated kiddo project, primarily address cost reduction and resource isolation. However, they fail to solve the fundamental issues of API management, interconnectivity, and strong multi-tenancy for the APIs themselves. The "elephant in the room," as the speakers describe it, is the urgent need for robust multi-tenancy for CRDs and other APIs, coupled with comprehensive API management capabilities that enable seamless interconnectivity between different API providers and consumers.

KCP emerges as the primary solution to these challenges. It is presented not as a replacement for Kubernetes clusters, but as "just yet another application running on your Kubernetes cluster" – a specific flavor of Kubernetes API server. KCP's core contribution is its ability to provide stronger multi-tenancy and API management by fundamentally shifting the API surface out of individual Kubernetes clusters. This re-architecting allows organizations to build internal or external cloud platforms more effectively.

Key findings regarding KCP's approach include:

  • Decoupling API Providers and Consumers: KCP introduces a novel mechanism where the ownership of CRD definitions (the API provider) is separated from the instantiation of Custom Resources (CRs) (the consumer). This is facilitated through API bindings and API exports, enabling a one-to-many relationship where a single API definition can be consumed by multiple distinct tenants or teams.
  • Virtual API Servers for Isolation: Within KCP, tenants are isolated within virtual API servers, providing a much stronger boundary than traditional namespace-based isolation within a single Kubernetes cluster. Each team or tenant effectively gets its "own box" for managing APIs and resources.
  • API Lifecycle Management: KCP integrates capabilities for managing the lifecycle of APIs, including versioning and upgrade strategies, directly within its control plane.
  • Workspace Tree for Organization: KCP organizes these virtual API servers and tenants into a workspace tree, providing a logical hierarchical structure for managing complex environments (e.g., application tower, compute tower). This mirrors the organizational structure of large enterprises.
  • Single Control Plane for Multiple Resources and Tenants: KCP consolidates the management of diverse Kubernetes resources, APIs, and tenants under a single, unified control plane. While the underlying controllers and workloads still run in existing Kubernetes clusters, the API interactions and definitions are centralized in KCP. This approach, exemplified by public cloud APIs like Google Cloud's Clusters API, SQL Admin API, and Compute API, demonstrates a widely accepted pattern for constructing managed services.

In essence, KCP offers a transformative approach to platform engineering, externalizing the API management layer to build scalable, multi-tenant platforms that leverage the familiarity and power of KRM without inheriting the operational complexities of a distributed mesh of Kubernetes clusters.

Technical Deep Dive

▶ Watch: Deep dive into KRM object structure and reconciliation (4:00)

The speakers begin by illustrating the typical microservice architecture prevalent in platform or cloud environments, characterized by a "mesh of interconnectivity" between services like Kubernetes, virtual machines, compute, storage, and networking. While this pattern is well-understood, attempting to implement it with a traditional Kubernetes-centric approach often leads to the "many clusters problem."

In a conventional setup, as an organization scales, it invariably ends up with multiple Kubernetes clusters – perhaps one for US customers, another for European customers, or isolated environments for specific clients. Each cluster might host its own CRDs and operators. This creates a complex reconciliation mesh across clusters, where a storage team's API update or a breaking change can ripple through and disrupt dependent services in other clusters. The challenge isn't the KRM standard itself, but its application across fragmented, isolated Kubernetes clusters, particularly concerning tenancy isolation and API management. Alternatives like vclusters or kiddo (a KCP sub-project for generic control planes) offer some cost savings and resource isolation but do not fundamentally address the API management and multi-tenancy for CRDs.

KCP addresses this by acting as a Kubernetes API server flavor that runs as "yet another application" on an existing Kubernetes cluster. It doesn't replace existing clusters but augments them by providing a dedicated control plane for managing APIs and tenants. In a KCP world, the previous architecture of multiple, interconnected clusters transforms. Instead of each team managing its own Kubernetes cluster and CRDs, every team gets its "own box" – a KCP workspace or virtual API server – where they can define and manage their APIs, all while operating within a single logical control plane. This significantly reduces the "many clusters problem."

The core technical innovation in KCP lies in its decoupling of API definition from API consumption:

  1. API Resource Schema (CRDs): In KCP, an API provider (e.g., a database team) defines an API resource schema, which is analogous to a CRD. This schema represents the desired API, such as a PostgreSQL database.
  2. API Exports: The API provider then exports this API, making it available for consumption by other teams or tenants.
  3. API Bindings: A consumer (e.g., an application team) wishing to use this database API creates an API binding. This binding acts as a contract, linking the consumer's virtual API server to the provider's exported API. This binding effectively makes the provider's API resource schema available within the consumer's workspace.
  4. Virtual API Servers and Workspaces: Each consumer operates within its own virtual API server, which provides strong multi-tenancy and isolation. When a consumer creates a Custom Resource (CR) based on the bound API (e.g., a PostgreSQLDatabase instance), it does so within its isolated virtual environment.
  5. One-to-Many Relationship: This model allows for a one-to-many relationship: a single API provider can export an API, and multiple consumers, each in their own virtual API server, can bind to and instantiate resources from that API. This contrasts sharply with standard Kubernetes, where a CRD installed in a cluster means CRs for that definition live in the same cluster, representing a one-to-one relationship.

KCP organizes these virtual API servers and workspaces into a workspace tree. This hierarchical structure allows for logical organization, where a "compute tower" might contain virtual machines and networking APIs, and an "application tower" might host application platform APIs. Different teams interact with this single, unified control plane, leveraging the API bindings to consume services from other teams.

The fundamental shift KCP enables is to externalize the API surface from individual Kubernetes clusters. While the actual controllers and workloads that reconcile these APIs and manage the underlying infrastructure still run in conventional Kubernetes clusters, the API definitions and their management are centralized in KCP. This allows platform builders to construct sophisticated services like "database as a service" or "control plane as a service" more efficiently, providing a consistent KRM-based experience to their users, while abstracting the complexities of the underlying infrastructure and multi-cluster deployments.

Demo / Proof of Concept

▶ Watch: Common pitfalls when designing KRM APIs (5:30)

The talk included a live demonstration of KCP's capabilities, showcasing a Database-as-a-Service (DBaaS) example that was also part of a KubeCon workshop. The demonstration aimed to illustrate how KCP facilitates multi-tenant API consumption and underlying resource provisioning.

The demo environment consisted of two distinct tenants, tenant1 and tenant2, each representing a logical cluster with its own unique URL. These tenants would act as consumers of a database service.

  1. Consumer Experience:
  • Mangirdas first showed tenant1, acting as a consumer, listing its Cluster resources. A Cluster resource representing a PostgreSQL database was already present, demonstrating the declarative nature of KRM. The spec of this resource, though noted to be "overwhelming," clearly indicated it was a PostgreSQL database.
  • Similarly, tenant2 was shown to also have a Cluster resource with the same name and object type, highlighting that both tenants could consume the same API. The key enabler for this was an API binding from each tenant's workspace to the posgress database team (the provider).
  • To further illustrate, a new Cluster resource was created in tenant1 by renaming an existing one (e.g., API clust consumer one). Upon creation, the system immediately began setting up the database instance.
  1. Provider Perspective:
  • Switching to the provider's view, the database team's actual Kubernetes cluster was shown. Here, it was evident that two different namespaces were running, each corresponding to a tenant's database instance, despite having the same database name from the consumer's perspective. This demonstrated KCP's ability to provide strong isolation and multi-tenancy at the underlying infrastructure layer, even when consumers declare resources with identical names.
  • As the new Cluster resource was created on the consumer side, a corresponding new namespace was observed being instantiated on the provider side. This showcased the real-time reconciliation and provisioning enabled by KCP, where the virtual API interactions translate into concrete actions in the underlying compute cluster.

The demo succinctly highlighted KCP's core value proposition: Consumers interact with their own isolated virtual environments (KCP workspaces), declaring resources using familiar KRM patterns, without needing to know or interact directly with other consumers. Meanwhile, the API provider (the database team) manages a single compute cluster, spinning up isolated resources for each tenant in dedicated namespaces. This effectively creates a "control plane as a service" and "database as a service" model.

For those interested in replicating the demo, the speakers directed attendees to docs.kcb.io/contrib, which hosts all KCP-related demos and presentations.

Defensive Implications

▶ Watch: Adhering to API conventions and leveraging schema validation (7:00)

KCP's approach to extending KRM has significant implications for defensive strategies, particularly for platform engineering teams, security practitioners, and cloud architects building and securing internal or external platforms.

  1. Enhanced Multi-tenancy and Isolation: KCP introduces virtual API servers and workspaces which provide a stronger isolation boundary than traditional namespace-based multi-tenancy within a single Kubernetes cluster. This means that API interactions and resource definitions for one tenant are logically and functionally separated from others. From a defensive standpoint, this reduces the blast radius of a compromise; an issue affecting one tenant's API plane is less likely to directly impact another.
  2. Standardized API Governance: By centralizing API definitions and management in KCP, platform teams can enforce consistent API conventions, validation rules (via OpenAPI schemas and CEL expressions), and versioning strategies across all services offered. This structured approach to API design prevents many of the pitfalls highlighted in the talk, leading to more robust and less error-prone APIs that are easier to secure and audit.
  3. Simplified API Lifecycle Management: The challenges of API evolution and versioning are explicitly addressed by KCP. This means platform teams can implement controlled API upgrades, conversions, and deprecations, reducing the risk of introducing vulnerabilities or breaking changes that could lead to security misconfigurations.
  4. Reduced Operational Complexity and Attack Surface: By abstracting away the "many clusters problem" and providing a unified control plane, KCP reduces the need for complex cross-cluster networking and reconciliation logic. This simplification can lead to a smaller operational footprint and a more manageable attack surface, as fewer bespoke integration points need to be secured.
  5. Consistent Developer Experience for Security: Developers consuming services through KCP will interact with familiar KRM-style APIs. This consistency can extend to security policies, as access controls (like RBAC) can be applied to KCP workspaces and API bindings, making it easier for developers to understand and adhere to security guidelines.
  6. Control Plane as a Service (CPaaS) Security: KCP enables the creation of "Control Plane as a Service" offerings. Securing the KCP instance itself becomes paramount, but once secured, it provides a trusted layer for delivering multi-tenant API capabilities. This centralized security focus is often more manageable than securing a distributed mesh of custom control planes.
  7. Clearer Separation of Concerns: The decoupling of API providers from consumers via API exports and API bindings enforces a clear separation of concerns. A database team can focus on securing their database API and underlying infrastructure, while application teams can consume it without needing deep knowledge or direct access to the provider's cluster. This reduces the cognitive load on individual teams and allows for specialized security expertise to be applied where it's most effective.

In essence, KCP provides a framework that inherently encourages better security practices through standardization, strong isolation, and streamlined API management, enabling platform teams to build more resilient and defensible cloud-native platforms.

Key Takeaways

  • The Kubernetes Resource Model (KRM) offers powerful declarative capabilities, but its application in complex multi-cluster, multi-tenant environments presents significant challenges related to API management and isolation.
  • Common pitfalls in KRM API design include poorly separating spec and status, ignoring established naming conventions, failing to leverage OpenAPI schema validation and CEL expressions, and neglecting proper API versioning and reconciliation strategies.
  • KCP (Kubernetes Control Plane) addresses the "many clusters problem" by providing a single, unified control plane that externalizes the API surface from individual Kubernetes clusters, offering stronger multi-tenancy and API management.
  • KCP introduces API bindings and API exports to decouple API providers from consumers, enabling a one-to-many relationship where a single API definition can be consumed by multiple isolated tenants or teams through virtual API servers.
  • This architecture allows organizations to build robust internal cloud platforms and "as-a-service" offerings (e.g., Database-as-a-Service) by leveraging the familiar KRM paradigm, while abstracting the underlying infrastructure and operational complexities.
  • KCP promotes standardized API governance, enhanced isolation, and streamlined API lifecycle management, which are crucial for building scalable, secure, and maintainable cloud-native platforms.

About the Speaker(s)

Mangirdas Judeikis (MJ) is one of the maintainers of the KCP project. He works for Casti, where he focuses on "funky things with Kubernetes," indicating a deep involvement in advanced Kubernetes technologies and solutions.

Nabarun Pal is also a contributor to the KCP project, working on the side. Additionally, he is involved in Kubernetes maintenance, showcasing his expertise in the core Kubernetes ecosystem.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This talk by KCP maintainers Mangirdas Judeikis and Nabarun Pal is a critical examination of the Kubernetes Resource Model's (KRM) limitations in multi-tenant, multi-cluster environments and a compelling introduction to KCP as the next evolution, KRM++. It provides deep technical insights into common KRM API design pitfalls and presents KCP's novel architecture for strong multi-tenancy, API management, and decoupling of API providers and consumers. For any platform engineer or architect grappling with the 'many clusters problem,' this is essential viewing that offers a truly actionable path forward.

Heather Calloway (CISO) — STRONG ACCEPT

This KubeCon talk on extending the Kubernetes Resource Model (KRM) through KCP directly addresses a critical governance and operational risk: the uncontrolled sprawl of APIs and fragmented infrastructure across multi-cloud environments. By offering a structural solution for centralized API management and strong multi-tenancy, KCP provides a credible path for platform teams to reduce business exposure and establish clearer accountability for API security and consistency. It's a pragmatic, engineering-led approach to a pervasive institutional problem, moving beyond theoretical discussions to actionable architectural patterns for building more defensible platforms.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025