Kubernetes CRD Design for the Long Haul: Tips, Tricks, and... Christian Schlotter & Fabrizio Pandini

Christian Schlotter, Fabrizio Pandini

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

In this insightful KubeCon EU talk, Christian Schlotter and Fabrizio Pandini, both maintainers of Cluster API, delve into the intricacies of designing robust and extensible Custom Resource Definitions (CRDs) for Kubernetes. Their presentation, "Kubernetes CRD Design for the Long Haul: Tips, Tricks, and...", addresses the critical challenge of evolving APIs gracefully within the Kubernetes ecosystem, drawing heavily from their extensive experience with Cluster API. The core message revolves around strategies to prevent common design pitfalls that often necessitate unplanned, breaking API version changes, which are notoriously difficult and disruptive to manage.

Watch on YouTube

Visual summary for Kubernetes CRD Design for the Long Haul: Tips, Tricks, and... Christian Schlotter & Fabrizio Pandini by Christian Schlotter, Fabrizio Pandini
Visual summary for Kubernetes CRD Design for the Long Haul: Tips, Tricks, and... Christian Schlotter & Fabrizio Pandini by Christian Schlotter, Fabrizio Pandini

Key moments

  1. 0:00 Introduction: The challenge of CRD evolution
  2. 2:20 API development process and common mistake origin
  3. 4:20 Anti-pattern: Embedding external APIs (kubeadm example)
  4. 6:20 Anti-pattern: Reusing core Kubernetes objects (ObjectReference)
  5. 8:40 Anti-pattern: Reusing internal Go types inappropriately

Kubernetes CRD Design for the Long Haul: Tips, Tricks, and...

Speakers: Christian Schlotter, Maintainer in Cluster API, Software Engineer at Broadcom; Fabrizio Pandini, Maintainer in Cluster API, Software Engineer at Broadcom

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=7IA-Vw1K7eg

Overview

In this insightful KubeCon EU talk, Christian Schlotter and Fabrizio Pandini, both maintainers of Cluster API, delve into the intricacies of designing robust and extensible Custom Resource Definitions (CRDs) for Kubernetes. Their presentation, "Kubernetes CRD Design for the Long Haul: Tips, Tricks, and...", addresses the critical challenge of evolving APIs gracefully within the Kubernetes ecosystem, drawing heavily from their extensive experience with Cluster API. The core message revolves around strategies to prevent common design pitfalls that often necessitate unplanned, breaking API version changes, which are notoriously difficult and disruptive to manage.

The speakers highlight that while software projects inherently evolve based on user feedback and new requirements, CRDs present a unique challenge compared to traditional binaries. Breaking changes in CRDs directly impact user configurations (YAML files) and require complex migration strategies, often forcing the creation of new API versions. This talk is essential for anyone involved in building or maintaining Kubernetes operators and CRDs, offering pragmatic advice to ensure their custom resources can adapt and scale over time without prematurely incurring the significant overhead of API version deprecation and migration.

The importance of this topic cannot be overstated, as CRDs are fundamental to extending Kubernetes' capabilities beyond its core resources. A well-designed CRD promotes stability, reduces operational burden for users, and facilitates long-term project viability. Conversely, poorly designed CRDs can lead to technical debt, user frustration, and a fragmented ecosystem. Schlotter and Pandini's talk provides a valuable roadmap for avoiding these pitfalls, emphasizing a proactive approach to API design that prioritizes clarity, consistency, and extensibility from the outset.

Background

▶ Watch: Introduction: The challenge of CRD evolution (0:00)

The Kubernetes ecosystem thrives on its extensibility, largely powered by Custom Resource Definitions (CRDs). CRDs allow users to define their own resource types, enabling them to extend Kubernetes' API with domain-specific objects and associated controllers. This mechanism is crucial for building sophisticated operators that manage complex applications and infrastructure within a Kubernetes cluster. However, the very flexibility that makes CRDs powerful also introduces significant challenges when it comes to long-term maintenance and evolution.

As any software project matures, its API must evolve to accommodate new features, address user feedback, and adapt to changing requirements. In the Kubernetes context, the most common strategy for API evolution involves two complementary ideas: first, continuously evolving the current API version without breaking changes (as seen with Kubernetes' core v1 API, which has been stable for years but still receives additions); and second, creating new API versions only when breaking changes are absolutely unavoidable and decided upon deliberately. The latter is a costly endeavor, involving complex migration paths for users and storage version migrations for the API server itself. The central problem addressed by Schlotter and Pandini is how to design CRDs in a way that minimizes the need for these unplanned, forced API version bumps.

The standard process for developing CRDs in Kubernetes involves defining Go types, which are then used by a generator like controller-gen to produce the CRD definition. This definition includes an OpenAPI spec that describes the structure and validation rules of the custom resource. Finally, this CRD is applied to the Kubernetes API server, allowing users to interact with it via YAML manifests. The speakers contend that while users are generally "always right" and generators "just work," the vast majority of mistakes originate in the Go types themselves or the accompanying comments and markers. Developers, often focused on the Go implementation, can overlook the critical impact their choices have on the resulting OpenAPI spec and, by extension, on the user experience and the API's long-term stability. This talk aims to shed light on these common pitfalls and provide actionable strategies to mitigate them, ensuring that CRD designs are truly built "for the long haul."

Key Findings

▶ Watch: API development process and common mistake origin (2:20)

The talk identifies several key anti-patterns and corresponding solutions for designing CRDs that can evolve gracefully over time, minimizing the need for disruptive breaking changes. These findings are rooted in the practical experiences of the Cluster API project and offer a robust framework for CRD development:

  1. Anti-Pattern: Embedding External API Go Types Directly.
  • Finding: Directly embedding Go types from external projects (like kubeadm's configuration or even core Kubernetes v1 types) into your CRD's spec can lead to significant problems. When the external dependency evolves or introduces breaking changes, your CRD is forced to either break compatibility with older versions or introduce a new API version, even if your project's logic hasn't changed.
  • Solution: Copy, Adapt, and Evolve. Create your own internal Go types that mirror the necessary parts of external APIs. This allows your CRD to maintain stability while converting to the specific external API version required in the backend. For generic Kubernetes types, copy only the fields you truly need, avoiding the exposure of unused fields that can confuse users or become breaking changes to remove later.
  1. Anti-Pattern: Reusing Go Structs for Distinct Concepts.
  • Finding: Using the same Go struct for different but seemingly similar concepts (e.g., a MachineSpec for a single machine and a MachineTemplateSpec for a fleet of machines) can lead to fields that make sense in one context but not the other. This "field leakage" can result in nonsensical API definitions or hinder future extensibility.
  • Solution: Separate Structs for Separate Concepts. Even if initial requirements overlap, anticipate divergent evolution paths and create distinct Go structs for each concept. Conversion between these structs is straightforward, preserving the clarity and logical consistency of your API.
  1. Anti-Pattern: Neglecting controller-gen Markers and the OpenAPI Spec.
  • Finding: Developers often focus solely on Go type definitions, assuming controller-gen will handle everything correctly. However, missing or incorrect magic markers (Go comments like +kubebuilder:validation:MaxLength) can lead to an incomplete or incorrect OpenAPI spec, impacting validation (e.g., for CEL expressions in ValidatingAdmissionPolicies) and object merging behavior. A critical example is the default "atomic" merging of lists, which can cause infinite reconciliation loops if multiple controllers try to modify the same list concurrently.
  • Solution: Inspect the OpenAPI Spec and Use Markers Deliberately. Always run controller-gen and review the generated OpenAPI spec to understand its impact on users. Utilize markers like +listType=map and +listMapKey for lists that require co-ownership and strategic merge patch behavior.
  1. Anti-Pattern: Poor Field Naming and API Structure.
  • Finding: Inconsistent terminology, generic field names, abbreviations, or a flat API structure can lead to confusing APIs that are hard for users to understand and difficult to extend without breaking changes. Renaming a field is a breaking change, making initial naming choices paramount.
  • Solution: Clarity, Consistency, and Extensibility through Structure.
  • "I want to..." Exercise: Read your API definition as a declarative sentence ("I want a cluster with a topology that has a control plane with three replicas"). If it doesn't read well, your API might be confusing.
  • Project Glossary: Define and maintain a glossary of terms specific to your project. This ensures consistent language across your API, documentation, and codebase, benefiting both users and future maintainers.
  • Nesting for Extensibility: Design for future growth by using nested objects instead of flat fields. For example, instead of nodeDrainGracePeriod and nodeDrainTimeout, use a nodeDrain object that contains both, allowing easy addition of new drain-related fields later.
  • Enums over Booleans: Consider using enums instead of booleans if a field might have more than two states in the future.
  1. Solution: Adopting Linters for Best Practices.
  • Finding: Manually catching all these design issues is challenging.
  • Solution: Integrate tools like KubeLinter (KL) into your development workflow. KL checks for common practices and enforces best practices in Kubernetes API design, integrating well with golangci-lint. This helps catch potential issues early, standardizing API quality across projects.

These findings collectively emphasize that designing CRDs for the long haul requires a holistic approach, moving beyond mere Go type implementation to consider the broader implications on the OpenAPI spec, user experience, and future extensibility.

Technical Deep Dive

▶ Watch: Anti-pattern: Embedding external APIs (kubeadm example) (4:20)

The core of the talk's technical insights revolves around specific Go type examples, the critical role of controller-gen markers, and the art of effective field naming. These elements are where most CRD design mistakes originate and where the most impactful solutions can be applied.

Go Type Design Pitfalls and Solutions

  1. Embedding External APIs:

The speakers illustrate this with Cluster API's interaction with kubeadm. kubeadm configurations evolve, with v1beta3 used before Kubernetes v1.29 and v1beta4 starting with v1.30. Cluster API needs to manage clusters across various Kubernetes versions (e.g., v1.27 to v1.33). If Cluster API were to directly embed kubeadm's Go types (e.g., kubeadm.v1beta3.ClusterConfiguration), updating the kubeadm dependency to support v1.30 would force Cluster API to use v1beta4, making it impossible to create clusters using older kubeadm versions without a breaking change.

  • Solution: Cluster API maintains its own copy of the necessary configuration structs. When a user requests a cluster of a specific Kubernetes version, the Cluster API controller converts its internal, stable representation to the appropriate kubeadm API version (v1beta3 or v1beta4) in the backend. This copy, adapt, and evolve strategy allows Cluster API to decouple its API evolution from that of kubeadm.
  1. Embedding Generic Core Kubernetes Types:

Another example is Cluster API's use of core/v1.ObjectReference within its Cluster object to refer to a control plane. While core/v1 appears stable, ObjectReference includes fields like UID, APIVersion, and Kind that Cluster API might not need or use. Exposing these unused fields to users can create confusion (users might set them expecting behavior that doesn't exist). Crucially, removing these unused fields later would constitute a breaking change in the Cluster API's v1beta1 API.

  • Solution: The recommended approach is to create a custom ObjectReference type containing only the necessary fields. This prevents exposing superfluous attributes and retains flexibility to add fields (like UID) only when they are explicitly needed, without incurring breaking changes for removals. Cluster API aims to address this in its v1beta2 API.
  1. Reusing Go Structs for Different Concepts:

The talk highlights the distinction between a Machine object (a single, unique instance) and a MachineTemplate (a blueprint for creating many machines). Initially, it might seem logical to reuse the same Go spec struct for both. However, if a field like IPAddress is later added to the MachineSpec (e.g., for IP address management), it would incorrectly leak into the MachineTemplateSpec. An IPAddress in a template makes no sense, as each machine created from the template should have a unique IP.

  • Solution: Employ separate Go structs for MachineSpec and MachineTemplateSpec. Even if they share many common fields initially, their distinct purposes warrant separate types, allowing independent evolution. The conversion between these structs is simple, preserving logical separation and preventing unintended side effects.

The Importance of controller-gen Markers and OpenAPI Spec

The speakers stress that the OpenAPI spec, generated from Go types and comments, is what truly defines the user-facing API. Developers must actively inspect this generated spec.

  1. Missing Validation Markers:
  • Problem: Failing to add appropriate +kubebuilder:validation markers (e.g., MaxLength, MinLength, Pattern) can lead to an API that lacks proper schema validation. This is particularly problematic for newer features like CEL (Common Expression Language), which rely on robust OpenAPI validations to function correctly within ValidatingAdmissionPolicies.
  • Solution: Always include relevant validation markers. After making changes, developers should always run controller-gen and examine the generated OpenAPI spec to confirm that validations are correctly reflected.
  1. List Merging Strategies and Co-ownership:
  • Problem: By default, Kubernetes treats lists in CRDs as atomic. If two different controllers or tools (e.g., a Cluster API controller and a GitOps tool) attempt to modify the same list (e.g., a variables list in a Cluster object), they will continuously overwrite each other's changes, leading to an infinite reconciliation loop. For instance, one tool might set instanceType and zone, while another tries to add costCenter. The atomic merge means one change stomps on the other, triggering a new reconciliation.
  • Solution: For lists where co-ownership is desired, use the +listType=map and +listMapKey=<key_field> markers. These markers instruct the Kubernetes API server to perform a strategic merge patch, allowing different controllers to manage different entries within the same list based on a specified key field. This ensures that changes from multiple sources can coexist without conflict.

Field Naming and API Structure for Extensibility

  1. The API as a Project Statement:

An API is a declarative statement of what a project does, but unlike full sentences, it lacks context. Every keyword (field name) in an API matters. Poor choices—confusing terms, synonyms, abbreviations, or overly generic names—can lead to user confusion and, more critically, force breaking changes if renaming is needed later.

  • Solution: The "I want to..." Exercise: Read your API's intention as "I want to..." For example, "I want a cluster with a topology that has a control plane with three replicas." If the sentence is clear and makes sense, the API is likely well-designed. If it's awkward or ambiguous, the naming needs refinement.
  1. The Project Glossary:
  • Problem: Lack of a consistent vocabulary within a project can lead to maintainers using different terms for the same concept or the same term for different concepts. This introduces ambiguity for users and future maintainers. The example given is a ready field in AWSMachine.Status. What does "ready" mean? Is the node provisioned, initialized, or fully up and running?
  • Solution: Create and maintain a project glossary. This defines core terms, clarifies their meaning, and ensures consistency across the API, documentation, and codebase. For the ready field, the glossary could specify it means "machine's initialization is completed and infrastructure is fully provisioned." This clarity helps prevent misinterpretation and ensures the API is explicit.
  1. Nesting for Extensibility:
  • Problem: A flat API structure can quickly become unwieldy and non-extensible. If new related fields need to be added, they often appear alongside unrelated ones, making the API harder to read and evolve. For example, if you have nodeDrainGracePeriod and later need to add nodeDrainTimeout, adding a new top-level field creates a flat structure.
  • Solution: Design for nesting. Group related fields under a common parent object. Instead of individual nodeDrainGracePeriod and nodeDrainTimeout fields, create a nodeDrain object that contains both. This allows for easy, non-breaking additions of new drain-related fields (e.g., forceDrain) within the nodeDrain object in the future.
  1. Enums over Booleans:
  • Pro Tip: If a boolean field might conceptually expand to more than two states in the future, consider using an enum from the start. This allows adding new states without breaking changes, whereas converting a boolean to an enum later would be a breaking change.

Inspiration and Tools

The speakers emphasize that API design is not a solo effort. Getting peer reviews, soliciting user feedback, and studying other successful Kubernetes APIs (like core v1 resources) are crucial. They also give a significant shout-out to KubeLinter (KL), a project by Cho, which helps enforce best practices and catches common CRD design issues. KL integrates with golangci-lint and is planned to be hosted under sigs.k8s.io/api-machinery, making it a valuable tool for any CRD developer.

Demo / Proof of Concept

▶ Watch: Anti-pattern: Reusing core Kubernetes objects (ObjectReference) (6:20)

This talk did not include a live demonstration or a proof of concept. Instead, Christian Schlotter and Fabrizio Pandini focused on sharing lessons learned and anti-patterns identified from their extensive experience developing and maintaining Cluster API, using specific examples from that project's evolution to illustrate their points.

Defensive Implications

▶ Watch: Anti-pattern: Reusing internal Go types inappropriately (8:40)

The insights shared in this talk are crucial for anyone involved in the Kubernetes ecosystem, particularly for CRD developers, operator maintainers, and platform engineers. Adopting these defensive strategies will significantly improve the stability, usability, and longevity of custom resources.

  1. For CRD Developers and Maintainers:
  • Prioritize API Stability Over Code Reuse: Resist the urge to directly embed external Go types (e.g., kubeadm configs) or generic Kubernetes core types (e.g., core/v1.ObjectReference) unless absolutely necessary and with full understanding of the implications. Instead, copy, adapt, and evolve your own types, exposing only the fields truly relevant to your CRD's contract. This insulates your API from upstream breaking changes and prevents exposing unused, confusing fields.
  • Distinguish Concepts with Separate Structs: Do not reuse Go structs for concepts that, while seemingly similar, have distinct lifecycle or evolution paths (e.g., MachineSpec vs. MachineTemplateSpec). Create separate, purpose-built structs to maintain logical consistency and allow independent feature development without unintended field leakage.
  • Master controller-gen Markers and OpenAPI Spec: Develop a habit of always running controller-gen after Go type changes and meticulously reviewing the generated OpenAPI spec. Pay close attention to validation markers (e.g., +kubebuilder:validation:MaxLength) to ensure robust schema validation, which is critical for CEL expressions.
  • Implement Strategic Merge Patch for Lists: For lists that require co-ownership by multiple controllers or tools, explicitly use +listType=map and +listMapKey=<key_field> markers. This prevents infinite reconciliation loops and enables collaborative management of list entries.
  • Adopt API Linters Early: Integrate KubeLinter (KL) into your CI/CD pipeline. This tool automates the enforcement of best practices, catching common design pitfalls (like missing validation markers or incorrect list types) early in the development cycle.
  • Invest in a Project Glossary: Establish and maintain a project glossary to define key terms and concepts. This ensures consistent terminology across your API, documentation, and codebase, eliminating ambiguity for users and future maintainers.
  • Design for Extensibility with Nesting: Structure your API fields using nested objects rather than a flat structure, especially for related attributes. This allows for non-breaking additions of new fields within a logical grouping, facilitating future evolution without requiring new API versions. Consider using enums over booleans if a field might conceptually expand beyond two states.
  • Seek Feedback Continuously: API design is a collaborative effort. Actively solicit peer reviews and user feedback. Study successful Kubernetes APIs for inspiration, but critically evaluate if their patterns apply to your specific context.
  1. For Users of CRDs:
  • Understand List Merging Behavior: Be aware of how lists are managed in the CRDs you use. If you are operating multiple tools or controllers that might concurrently modify the same list, verify if the CRD uses +listType=map to enable strategic merge patching. If not, be cautious to avoid conflicting updates.
  • Consult Project Glossaries: When interacting with a new CRD, seek out its project glossary. This will provide clarity on the precise meaning of fields and reduce misinterpretation of API behavior.
  • Provide Constructive Feedback: If you encounter ambiguous field names, confusing structures, or unexpected API behavior, provide constructive feedback to the CRD maintainers. Your input is vital for improving API design for the community.

By diligently applying these defensive principles, both CRD producers and consumers can contribute to a more stable, predictable, and robust Kubernetes ecosystem, ensuring that custom resources are truly designed "for the long haul."

Key Takeaways

  • Prioritize API Stability: Design Custom Resource Definitions (CRDs) to evolve gracefully without forced, unplanned breaking changes, which are costly and disruptive for users and maintainers.
  • Avoid Direct Embedding: Do not directly embed Go types from external projects (e.g., kubeadm configs) or generic Kubernetes core types (e.g., core/v1.ObjectReference) into your CRD. Instead, copy, adapt, and evolve your own types to maintain control and prevent external dependencies from dictating your API's evolution.
  • Separate Concepts Clearly: Use distinct Go structs for different API concepts, even if they initially appear similar (e.g., MachineSpec vs. MachineTemplateSpec). This prevents unintended field leakage and allows for independent, logical evolution of distinct functionalities.
  • Leverage controller-gen Markers and Inspect OpenAPI Spec: Actively use controller-gen markers (especially +listType=map for co-ownership and +kubebuilder:validation for robust schema validation) and consistently review the generated OpenAPI spec to ensure your API behaves as intended and supports features like CEL.
  • Adopt KubeLinter (KL): Integrate KubeLinter (KL) into your development workflow to automatically check for and enforce best practices in Kubernetes API design, catching common pitfalls early.
  • Focus on Clear Naming and Extensible Structure: Develop a project glossary for consistent terminology and design your API with nesting to enable non-breaking future extensions. Use the "I want to..." exercise to test the readability and clarity of your API design.

About the Speaker(s)

Christian Schlotter is a dedicated maintainer in the Cluster API project and works as a Software Engineer at Broadcom. His extensive experience in building and evolving Cluster API's CRDs provides a practical foundation for the insights shared in this talk.

Fabrizio Pandini is also a maintainer in the Cluster API project and works alongside Christian Schlotter as a Software Engineer at Broadcom. His contributions to Cluster API and deep understanding of Kubernetes API design principles are evident throughout the presentation. Together, their combined expertise offers valuable lessons for anyone navigating the complexities of CRD development.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk by Cluster API maintainers Christian Schlotter and Fabrizio Pandini delivers unvarnished, high-value technical guidance on designing robust Kubernetes CRDs. It’s a deep dive into avoiding common pitfalls that lead to painful API version bumps, drawing directly from years of hard-won experience. The speakers provide actionable strategies for ensuring CRDs are stable, extensible, and maintainable, making it essential viewing for anyone building or operating custom resources in Kubernetes.

Heather Calloway (CISO) — STRONG ACCEPT

This KubeCon talk by Schlotter and Pandini delivers critical insights into designing Kubernetes Custom Resource Definitions for long-term stability, a topic often overlooked but paramount for institutional resilience. They dissect common design pitfalls that lead to costly, disruptive breaking changes and provide pragmatic strategies rooted in deep operational experience. For any organization building or heavily relying on Kubernetes, understanding these principles is essential for managing platform risk, reducing operational overhead, and ensuring the predictable evolution of their infrastructure.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025