Project Lightning Talk: API Management in the CRD World: What Linkerd Has Learned - Phil Henderson

Phil Henderson

KubeCon + CloudNativeCon Europe 2025 · Project Lightning Talk

Overview

In this insightful lightning talk from KubeCon EU, Phil Henderson, a Customer Success Engineer at Buoyant, shared Linkerd's hard-won lessons in navigating the complex landscape of API management within the Kubernetes Custom Resource Definition (CRD) world. As the creators of the term "service mesh" and pioneers in the field, Linkerd has a unique perspective on building and maintaining robust, user-friendly Kubernetes extensions. Henderson's presentation distilled years of practical experience into actionable advice for developers and platform engineers grappling with the intricacies of CRDs.

Watch on YouTube

Visual summary for Project Lightning Talk: API Management in the CRD World: What Linkerd Has Learned - Phil Henderson by Phil Henderson
Visual summary for Project Lightning Talk: API Management in the CRD World: What Linkerd Has Learned - Phil Henderson by Phil Henderson

Key moments

  1. 0:00 Linkerd introduction and core values
  2. 1:28 The challenge of CRD management in Kubernetes
  3. 2:00 Ease of CRDs can lead to versioning challenges
  4. 3:00 Lesson: Don't install shared CRDs for users
  5. 3:55 Key principle: APIs must be tailored for humans
  6. 4:20 Annotations: flexibility, experimentation, and validation issues
  7. 4:45 Talk's two main takeaways summarized

Project Lightning Talk: API Management in the CRD World: What Linkerd Has Learned

Speakers: Phil Henderson, Customer Success Engineer, Buoyant

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=5CyAZUH1f8

Overview

In this insightful lightning talk from KubeCon EU, Phil Henderson, a Customer Success Engineer at Buoyant, shared Linkerd's hard-won lessons in navigating the complex landscape of API management within the Kubernetes Custom Resource Definition (CRD) world. As the creators of the term "service mesh" and pioneers in the field, Linkerd has a unique perspective on building and maintaining robust, user-friendly Kubernetes extensions. Henderson's presentation distilled years of practical experience into actionable advice for developers and platform engineers grappling with the intricacies of CRDs.

The core of the talk revolved around the inherent tension between Kubernetes' encouragement of extending its API via CRDs and its lack of native support for managing these custom resources effectively, especially in multi-project or shared contexts. Henderson emphasized Linkerd's unwavering commitment to simplicity, security, and developer experience, revealing how these principles guided their approach to CRD design and management. The talk served as a crucial guide for anyone looking to avoid common pitfalls associated with CRD versioning, shared resource dependencies, and excessive configurability, advocating for a human-centric approach to API design.

Henderson's insights are particularly relevant in today's cloud-native ecosystem, where CRDs have become the de facto standard for extending Kubernetes functionality. From service meshes and operators to custom controllers, CRDs empower users to define their desired state and behavior. However, this power comes with significant management challenges. Linkerd's journey, as articulated by Henderson, offers a pragmatic blueprint for designing CRDs that are not only powerful but also maintainable, secure, and genuinely easy for human operators to use, ultimately fostering greater reliability and security in Kubernetes environments.

Background

▶ Watch: Linkerd introduction and core values (0:00)

Linkerd, as introduced by Phil Henderson, holds a significant place in the cloud-native history as the original service mesh, having coined the term itself. Its foundational mission is to bring security, reliability, and observability to workloads, prioritizing security above all else. Built in Rust for an ultra-light footprint, Linkerd prides itself on being the easiest-to-use and most secure service mesh available, offering a zero-config experience where simply adding an annotation to a namespace meshes everything, securing traffic with mTLS (mutual Transport Layer Security) by default. This commitment to simplicity and security forms the bedrock of Linkerd's design philosophy, which extends directly to its approach to API management.

The Kubernetes ecosystem thrives on extensibility, and Custom Resource Definitions (CRDs) are the primary mechanism for extending the Kubernetes API. CRDs allow users to define their own resource types, complete with schema validation, enabling Kubernetes to manage applications and infrastructure beyond its built-in capabilities. This power has led to an explosion of operators and custom controllers, making CRDs ubiquitous. However, as Henderson pointed aptly, while Kubernetes "encourages CRDs," it "is not helping us as a whole with CRDs." The ease of creating CRDs often leads to an "overboard" situation, where projects introduce overly complex or poorly managed custom resources.

One of the significant challenges highlighted is CRD versioning. While CRDs include a version field, managing backward compatibility and schema evolution across different versions is notoriously difficult. Developers often struggle to maintain perfect backward compatibility, leading to breaking changes or complex migration paths for users. Furthermore, the issue of shared CRDs presents a unique set of problems. Projects like the Gateway API aim to standardize API definitions for common networking constructs, meaning multiple controllers (e.g., Linkerd, other service meshes, ingress controllers) might depend on the same underlying CRDs. This introduces critical questions: Who is responsible for installing these shared CRDs? Who manages their lifecycle, including deletion? If one project deletes a CRD that another relies on, what are the consequences for the dependent application? These are the practical, real-world problems that Linkerd has encountered and learned from, shaping their philosophy on API design.

Key Findings

▶ Watch: Ease of CRDs can lead to versioning challenges (2:00)

Linkerd's extensive experience operating in the CRD-centric Kubernetes world has yielded several critical findings, which Phil Henderson articulated as essential lessons for anyone developing or managing Kubernetes extensions. These findings underscore Linkerd's pragmatic approach to API design, prioritizing user experience and long-term maintainability over perceived flexibility.

Firstly, Henderson emphasized the profound difficulty of effective CRD management within Kubernetes, particularly when dealing with shared CRDs. While Kubernetes provides the mechanism for defining custom resources, it offers minimal assistance in coordinating their lifecycle across multiple independent projects. This leads directly to Linkerd's most significant practical lesson: Do not try to install shared CRDs for your users. While it might seem convenient to bundle CRD installation with an application, this approach inevitably leads to conflicts and operational headaches. If multiple applications rely on the same shared CRD (e.g., the Gateway API), and each tries to manage its installation and uninstallation, a chaotic scenario emerges. One project's deletion of a shared CRD could inadvertently cripple another application, leading to system instability and a poor user experience. The responsibility for shared resource installation and lifecycle management, Henderson argues, should ideally reside with the user or a dedicated platform component, not with individual applications.

Secondly, Linkerd has learned that APIs must be tailored for humans who need to use them. This principle stands in direct opposition to the common developer inclination to expose every possible configuration option. Henderson coined the phrase, "Configurability is the enemy," explaining that adding features or parameters that no human user genuinely needs introduces unnecessary complexity and cognitive load. Such superfluous options have a "real cost" in terms of understanding, documentation, and potential misconfiguration. The goal is to make APIs unambiguous without requiring full specification, much like how users intuitively understand https://google.com without needing to specify the default 443 port. The focus should be on providing a clear, intuitive interface that abstracts away unnecessary details, making the API simpler to consume and less prone to errors.

Finally, the talk highlighted the dual nature of annotations in Kubernetes. Annotations provide a flexible mechanism for attaching arbitrary, non-structural metadata to Kubernetes objects, making them "great for experimenting" and offering a degree of flexibility that might not be suitable for formal API fields. Linkerd takes "full advantage" of annotations for extending functionality or testing new features. However, this flexibility comes with a significant drawback: validating annotations is a major concern. Unlike CRD fields, which can leverage schema validation, Kubernetes does not offer native, robust validation for arbitrary annotation values. This forces projects like Linkerd to rely on webhooks for validation, an out-of-band mechanism that adds complexity and is not as seamlessly integrated as native schema validation. Henderson noted that "fixing it may be hard," underscoring a broader limitation in Kubernetes' current CRD capabilities regarding flexible, yet validated, extensibility.

Technical Deep Dive

▶ Watch: Lesson: Don't install shared CRDs for users (3:00)

Linkerd's architectural choices and philosophical commitment to simplicity directly influence its approach to API management within Kubernetes. At its core, Linkerd operates as a service mesh, injecting lightweight Rust-based proxies (the data plane) alongside application pods. These proxies intercept all network traffic, providing capabilities like mTLS for secure, encrypted communication by default, automatic retries, load balancing, and telemetry collection. The control plane, also written in Rust, manages and configures these proxies, leveraging Kubernetes' API for discovery and configuration. This "ultra-light package" design, focused on minimal footprint and high performance, is intrinsically linked to its API design principles.

When Phil Henderson discusses CRDs, he's referring to Kubernetes' powerful mechanism for extending its API. A CRD defines a new, custom resource type, which can then be created, updated, and managed using standard Kubernetes API calls. Each custom resource object adheres to a defined schema, specifying its fields and their types. The challenge, as Henderson points out, is not in creating CRDs but in managing their lifecycle and evolution. Versioning CRDs, for instance, is far more complex than simply incrementing a version field. Changes to a CRD's schema (e.g., renaming a field, changing a type, adding a mandatory field) require careful consideration of backward compatibility. Kubernetes offers features like conversion webhooks to facilitate schema transformations between different API versions, but achieving "perfect backwards compatibility" is a demanding task that few projects consistently manage without significant effort or occasional breaking changes. This complexity impacts user migrations and the stability of the ecosystem.

The problem of shared CRDs is particularly acute and was a central theme of Henderson's talk. The Gateway API project serves as a prime example. The Gateway API aims to provide a standardized, role-oriented API for traffic management in Kubernetes, abstracting away the specifics of different ingress controllers or service meshes. It defines several CRDs, such as GatewayClass, Gateway, HTTPRoute, TCPRoute, and TLSRoute, allowing users to define how external traffic enters a cluster and how internal traffic is routed. The challenge arises because multiple components (e.g., Linkerd, Istio, NGINX Ingress Controller) might implement the Gateway API specification. If each of these components attempts to install or manage the Gateway API CRDs, conflicts inevitably arise regarding ownership, updates, and deletion. Henderson's strong advice — "Don't try to install shared CRDs for your users" — stems from the understanding that this leads to an unmanageable dependency nightmare where uninstallation of one component could inadvertently remove critical shared resources, breaking others. A more robust approach might involve a dedicated lifecycle manager for shared APIs or clear documentation instructing users on independent CRD installation.

Linkerd's API design principles are rooted in a deep understanding of developer and SRE workflows. The mantra, "how will this affect people from developers to SREs?" drives every design decision. This translates to an API that is simple, unambiguous, and tailored to essential human needs. The example of https://google.com versus https://google.com:443 perfectly illustrates this: users don't need to specify default values or common knowledge; the API should infer them or abstract them away. This minimizes cognitive load and reduces the surface area for configuration errors. The concept of "configurability is the enemy" is not an anti-feature stance but a recognition that every exposed configuration option adds complexity, potential for misconfiguration, and a burden on documentation and maintenance. Linkerd aims for a "just works" experience, where intelligent defaults and convention over configuration reduce the need for extensive user input.

Finally, Henderson touched upon annotations as a tool for flexibility and experimentation. Annotations are key-value pairs that can be attached to any Kubernetes object. Unlike fields within a CRD's spec, annotations are not part of the object's formal schema and are not typically subject to strict validation by the API server. This makes them ideal for extending functionality in a non-breaking way or for testing out new features that might eventually be promoted to formal CRD fields. Linkerd leverages annotations for various purposes, allowing users to fine-tune behavior without altering the core API. However, the lack of native Kubernetes validation for annotations poses a significant challenge. To enforce rules or validate input provided via annotations, projects must implement admission webhooks. An admission webhook is an HTTP callback that receives admission requests (e.g., for creating or updating an object) and can mutate or validate the object before it is persisted in etcd. While effective, webhooks add an external dependency and an additional layer of complexity to the control plane, making annotation validation less straightforward than native schema validation. This highlights a gap in Kubernetes' API extensibility model: the need for a mechanism to validate flexible, unstructured data like annotations more natively.

Demo / Proof of Concept

▶ Watch: Annotations: flexibility, experimentation, and validation issues (4:20)

Given the lightning talk format, which typically limits presentations to 5-10 minutes, Phil Henderson's session focused entirely on conveying critical insights and lessons learned rather than featuring a live demonstration or a detailed proof of concept. The speaker's primary objective was to share Linkerd's extensive experience and the practical implications of CRD management in real-world Kubernetes deployments, making the delivery of concise, actionable advice the core of the presentation. Therefore, no specific demo or proof of concept was presented during this talk.

Defensive Implications

▶ Watch: Talk's two main takeaways summarized (4:45)

The lessons shared by Phil Henderson from Linkerd's experience offer crucial defensive implications for anyone operating in the Kubernetes ecosystem, particularly for platform engineers, SREs, and developers building custom controllers or operators. Adhering to these principles can significantly enhance the stability, security, and usability of Kubernetes environments.

  1. Strategic CRD Ownership and Lifecycle Management: The most critical defensive posture regarding CRDs, especially shared ones like those in the Gateway API, is to establish clear ownership and lifecycle management strategies. Defenders should avoid auto-installing shared CRDs as part of their application or operator deployments. Instead, platform teams should consider a centralized approach for managing common CRDs, or explicitly delegate this responsibility to the end-user. This prevents conflicts, ensures consistent versions, and allows for graceful uninstallation without impacting other dependent systems. Before adopting any shared CRD, an organization should define who is responsible for its initial deployment, updates, and eventual removal.
  1. Prioritize Human-Centric API Design: Defensive development extends beyond security vulnerabilities; it includes protecting users from unnecessary complexity and potential misconfigurations. Developers building CRDs should rigorously evaluate every proposed field or configuration option with the question: "Does a human need to interact with this, or can it be an intelligent default?" By embracing Linkerd's philosophy that "configurability is the enemy" of simplicity, teams can design APIs that are intuitive, minimize cognitive load, and reduce the surface area for human error. This means striving for clear, unambiguous APIs that provide sensible defaults and abstract away unnecessary technical details, leading to more robust and less error-prone deployments.
  1. Robust Annotation Validation: While annotations offer flexibility, their lack of native Kubernetes validation poses a defensive challenge. If an application relies on annotations for critical behavior or configuration, defenders must implement robust validation mechanisms, typically through validating admission webhooks. These webhooks should rigorously check the format, content, and semantic correctness of annotation values to prevent malformed input from causing unexpected behavior, security vulnerabilities, or system instability. This requires careful planning and implementation of webhook services, ensuring they are highly available and performant.
  1. Embrace Simplicity and Secure Defaults: Linkerd's success stems from its commitment to simplicity and secure-by-default configurations, such as mTLS. Defenders should adopt a similar mindset when designing their own systems. Complex systems are harder to secure and maintain. By defaulting to secure practices (e.g., mTLS for all internal traffic, least privilege access) and providing a "zero-config" or minimal-config experience where possible, organizations can significantly reduce their attack surface and operational burden. This involves thoughtful design choices that guide users towards best practices rather than requiring them to explicitly configure every security control.
  1. Careful CRD Versioning and Backward Compatibility: When evolving CRDs, defenders must meticulously plan versioning strategies. While "perfect backwards compatibility" is challenging, efforts should be made to minimize breaking changes. If breaking changes are unavoidable, clear migration paths, deprecation warnings, and comprehensive documentation are essential. Utilizing Kubernetes' conversion webhooks can help manage schema evolution between API versions, but these should be thoroughly tested. A defensive approach here means anticipating future needs, designing CRDs with extensibility in mind, and having a clear policy for supporting older API versions to ensure smooth upgrades and minimize disruption.

Key Takeaways

  • CRD Management is Complex, Especially for Shared Resources: While Kubernetes encourages extending its API with CRDs, it lacks native support for managing their lifecycle across multiple projects, leading to potential conflicts.
  • Avoid Auto-Installing Shared CRDs: Projects should not attempt to install shared CRDs for their users, as this creates dependency conflicts and operational headaches, particularly during uninstallation. Responsibility for shared CRDs should be externalized or clearly defined.
  • Design APIs for Humans, Not Maximum Configurability: Prioritize simplicity and ease of use over exposing every possible configuration option. "Configurability is the enemy" when it adds unnecessary complexity and cognitive load for users.
  • Annotations Offer Flexibility but Require Robust Validation: Annotations are valuable for experimentation and custom extensions, but their lack of native Kubernetes validation necessitates the use of external mechanisms like validating admission webhooks.
  • Linkerd's Philosophy Emphasizes Simplicity and Security by Default: Linkerd's success is rooted in its ultra-light, Rust-based architecture, zero-config experience, and secure defaults like mTLS, principles that extend to its CRD design.
  • Learn from Others' Mistakes: Henderson explicitly encourages other projects to learn from Linkerd's hard-won lessons in CRD management to avoid common pitfalls and build more robust, user-friendly Kubernetes extensions.

About the Speaker(s)

Phil Henderson is a Customer Success Engineer at Buoyant, the company behind the Linkerd service mesh. In his role, Phil works directly with users, helping them successfully deploy, operate, and troubleshoot Linkerd in their Kubernetes environments. This hands-on experience provides him with a unique perspective on the practical challenges users face, particularly concerning API management and CRD interactions. His insights are therefore grounded in real-world feedback and operational realities. While he notes a lack of social media presence, Phil can be found on GitHub and various CNCF Slack channels under the handle @Phil, indicating his active involvement and contributions within the cloud-native open-source community. His presentation reflects Buoyant's and Linkerd's commitment to sharing knowledge and fostering best practices within the ecosystem.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This lightning talk from Linkerd's Phil Henderson delivers a potent dose of hard-won operational wisdom concerning Kubernetes CRD management. It cuts through the typical KubeCon fluff, offering brutally honest insights into the perils of shared CRDs, the pitfalls of excessive configurability, and the necessity of human-centric API design. While not a zero-day exploit, the lessons are deeply technical and highly actionable, providing critical guidance for anyone building or operating complex systems on Kubernetes. Henderson's direct approach and Linkerd's battle-tested experience make this a valuable, no-nonsense contribution to practical cloud-native engineering.

Heather Calloway (CISO) — STRONG ACCEPT

This talk from Linkerd distills critical lessons on API management within the Kubernetes CRD ecosystem, moving beyond mere technical details to address significant organizational risk and operational challenges. While highly technical in its subject matter, the speaker effectively translates hard-won experience into clear guidance on shared resource ownership, the perils of excessive configurability, and the necessity of robust validation for flexible extensions. It’s a pragmatic, unsentimental assessment that underscores the real-world consequences of design decisions on stability and security, offering actionable insights for platform teams and, by extension, security leaders guiding…

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025