Making CRDs Delightful: Beyond the Pitfalls - Evan Anderson, Stacklok, Inc
Evan Anderson, Stacklok, Inc
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In this insightful KubeCon EU talk, Evan Anderson, a founder of the Knative project and an experienced engineer from Stacklok, VMware, and Google Cloud, tackles a critical yet often overlooked aspect of Kubernetes operator development: user experience (UX). Anderson argues that while Kubernetes operators and Custom Resource Definitions (CRDs) are powerful tools for extending Kubernetes, their UX is frequently neglected, leading to frustration for developers and operations teams alike. This presentation is a call to action for operator authors to prioritize clarity, predictability, and ease of use in their CRD designs and associated tooling.

Key moments
- 0:00 Introduction and importance of CRD UX
- 2:00 Why good UX is critical for Kubernetes operators
- 2:55 The fundamental role of CRD Status
- 4:00 Examples of effective Status fields from projects
- 6:00 When to omit Status: Configuration-only CRDs
- 7:00 Using Kubernetes Events for user communication
Making CRDs Delightful: Beyond the Pitfalls
Speakers: Evan Anderson, Stacklok, Inc
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=6KdywJWnYyg
Overview
In this insightful KubeCon EU talk, Evan Anderson, a founder of the Knative project and an experienced engineer from Stacklok, VMware, and Google Cloud, tackles a critical yet often overlooked aspect of Kubernetes operator development: user experience (UX). Anderson argues that while Kubernetes operators and Custom Resource Definitions (CRDs) are powerful tools for extending Kubernetes, their UX is frequently neglected, leading to frustration for developers and operations teams alike. This presentation is a call to action for operator authors to prioritize clarity, predictability, and ease of use in their CRD designs and associated tooling.
Anderson draws on his extensive background, particularly with Knative, to highlight common pitfalls and present practical, actionable strategies for making CRDs more "delightful." The talk emphasizes that good UX is not merely a nicety but a necessity, especially during high-stress operational scenarios where clear interfaces can prevent costly mistakes and accelerate resolution. By focusing on well-structured status fields, informative events, sensible RBAC, and intuitive CLI interactions, Anderson demonstrates how operators can significantly enhance the developer experience within the Kubernetes ecosystem.
The core message resonates with anyone who has struggled with opaque operator behavior or confusing CRD schemas. Anderson provides a comprehensive checklist of best practices, ranging from fundamental CRD design principles to advanced patterns like duct typing and intelligent use of zero values. This talk is essential for anyone building or maintaining Kubernetes operators, offering a roadmap to transform a potentially frustrating interaction into a smooth and efficient one, ultimately fostering a more welcoming and productive environment for all Kubernetes users.
Background
▶ Watch: Introduction and importance of CRD UX (0:00)
The evolution of Kubernetes has seen a significant shift towards extending its capabilities through Custom Resources (CRs) and Custom Resource Definitions (CRDs). Originally introduced as "ThirdPartyResources," CRDs provide a mechanism for users to define their own API objects, enabling the Kubernetes API to manage application-specific state and logic. This extensibility is the foundation of the operator pattern, where controllers watch custom resources and manage application lifecycle, often interacting with external systems.
While CRDs and operators unlock immense power and flexibility, their widespread adoption has also exposed inherent challenges in user experience. Many open-source projects, including successful ones like Git and Linux, are not celebrated for their intuitive interfaces, and Kubernetes operators often fall into this category. The problem arises when operators, designed to automate complex tasks, present users with ambiguous status information, difficult-to-debug behaviors, or overly verbose configurations. This can lead to increased cognitive load, especially for developers new to a project or operations teams under pressure to resolve incidents.
Evan Anderson, with his deep involvement in projects like Knative, has firsthand experience with these challenges. Knative, a serverless platform built on Kubernetes, heavily leverages CRDs to define services, revisions, and routes. This background has given Anderson a unique perspective on the pain points and opportunities for improvement in CRD design. His talk aims to distill years of experience into concrete recommendations, addressing why thoughtful UX design is paramount in the Kubernetes space – a space where developers increasingly live and work, and where clarity directly impacts operational efficiency and user satisfaction. The talk implicitly acknowledges that while Kubernetes provides the building blocks for custom resources, the responsibility for crafting a "delightful" experience lies squarely with the operator authors.
Key Findings
▶ Watch: The fundamental role of CRD Status (2:55)
Evan Anderson's talk distills a wealth of experience into several key findings and best practices for creating delightful CRDs and Kubernetes operators. These findings emphasize clarity, predictability, and user-centric design:
- Comprehensive and Dual-Purpose Status Fields: The status field of a CR should serve both machines and humans. It needs machine-readable enums for automation (e.g.,
Degraded,Working) and human-readable descriptions or sentences explaining the current state and any discrepancies (e.g., "Why are there two replicas when I asked for four?"). References to other Kubernetes resources within the status (like URLs or related objects) significantly enhance debugging and usability. - Strategic Use of Kubernetes Events: Operators should leverage Kubernetes Events to communicate significant state changes or occurrences related to a custom resource. Events provide an auditable, namespace-scoped log visible via
kubectl describe, reducing the need for users to access controller logs, which might be restricted. However, events should be used judiciously to avoid API server noise. - Aggregated RBAC for Seamless Permissions: To ensure custom resources are accessible within standard Kubernetes roles, operators should utilize aggregated cluster roles. Adding labels like
aggregate-to-viewto a cluster role definition automatically bundles permissions for custom resources into built-in roles likevieworedit, making them discoverable and manageable alongside native Kubernetes objects. - Standardized Condition Patterns with Positive Polarity: Adopting a consistent condition pattern (e.g.,
Ready,Succeeded) with a positive polarity (wheretruemeans good) allows for automated summarization and consistent UI representation. This makes it easier for both machines and humans to quickly ascertain the health and progress of a resource. - Enhanced
kubectlIntegration: Custom resources can be made more user-friendly throughkubectlby defining additional printer columns forkubectl get, providing short names (e.g.,appforApplication), and categorizing resources (e.g., adding them to theallcategory forkubectl get all). - Aggregated Objects for User-Centric Views: Designing top-level custom resources that aggregate and summarize the state of multiple underlying Kubernetes objects (e.g., a Knative Service summarizing its deployments, routes, and revisions) provides users with a simplified, high-level view and reduces cognitive load.
- Leveraging Existing Kubernetes Types and Patterns (Duct Typing): Reusing well-known Kubernetes types like
ObjectReferenceandLabelSelectorand adopting duct typing patterns (e.g., consistently usingPodTemplateSpecwithinspec.templatefor workload definitions) reduces the learning curve for users already familiar with Kubernetes. - Meaningful Zero Values: Where appropriate, designing zero values (like an empty string or
0) in CRD schemas to represent a useful default or "do the right thing" behavior can lead to shorter, more powerful YAML and fewer user surprises. - GitOps-Friendly Design: Operators should be designed for idempotency, meaning resources can be reapplied multiple times without side effects, and should be resilient to the order in which resources are applied. This is crucial for seamless integration with GitOps tools like Flux and Argo CD.
- Clear Interaction with Other Resources: When an operator interacts with or creates other Kubernetes resources, it should use labels for lookup, annotations for detailed configuration, and critically, set owner references for proper garbage collection and traceability.
- Upfront Validation: Utilizing OpenAPI schema validation and CEL (Common Expression Language) in CRDs allows for early validation of user input, providing immediate and helpful error messages before a resource is even created, preventing invalid configurations from reaching the controller.
These findings collectively advocate for a shift in mindset: instead of merely making a controller work, operator authors should strive to make it work well for its human users, enhancing usability, debuggability, and overall satisfaction.
Technical Deep Dive
▶ Watch: Examples of effective Status fields from projects (4:00)
Anderson's talk provides a rich set of technical recommendations for crafting superior CRD experiences, moving beyond basic functionality to focus on usability and clarity.
The Power of Status Fields
A well-designed status field is paramount. Anderson emphasizes that it must cater to both machines and humans. For machines, status should offer concise, programmatic states, often using enums (e.g., Degraded, Working). For humans, it needs context – descriptive sentences explaining why a resource is in a particular state. For instance, if a user requests four replicas but only two are running, the status should not just say currentReplicas: 2 but also explain, "Insufficient capacity in cluster" or "Pod startup failed due to image pull error."
Examples from existing projects illustrate this:
- Argo CD Application: Its status not only confirms
in syncbut also lists all managed resources, enabling users to trace dependencies. - Knative Service: Provides the
URLto access the deployed service directly in the status, a highly convenient feature, and details about theready revision. - cert-manager Certificate: Includes critical information like the certificate's
validForduration, which is vital for operational awareness.
Anderson also notes that status can sometimes be longer than the spec, reflecting the controller's expansion of a user's concise request into a complex operational state. Crucially, not all CRDs require a status. Configuration-only CRDs like GatewayClass or StorageClass define parameters for other resources; their status would typically reside on the resources they provision. Similarly, policy resources like RoleBinding don't have an inherent operational status.
Leveraging Kubernetes Events
Events are a built-in Kubernetes mechanism for controllers to communicate discrete occurrences related to an object. These are not logs but distinct API objects, viewable via kubectl describe <resource>. Anderson highlights their value for users who might not have access to controller logs, providing a namespace-scoped audit trail of significant state changes (e.g., "Deployment created," "Scaling up," "Image pull failed"). The key is to create events only for state changes, avoiding excessive noise from periodic checks that haven't altered the resource's state. Over-generating events can flood the API server and potentially destabilize the cluster, a pitfall Anderson admits to having experienced.
Streamlining RBAC with Aggregated Roles
Kubernetes offers aggregated cluster roles to simplify permission management for custom resources. By adding specific labels (e.g., rbac.authorization.k8s.io/aggregate-to-view: "true") to a cluster role, permissions for custom resources can be automatically included in standard roles like view, edit, or admin. This ensures that users with default roles can interact with custom resources without manual RoleBinding adjustments for every new CRD. For example, a user with view access to a namespace can then kubectl get Knative Services or Gateway API HTTPRoutes, just like they would pods. This mechanism can also be used to build custom, higher-level roles that encompass multiple CRD types.
Standardizing Conditions for Clarity
A powerful pattern for conveying resource health is the use of conditions with a consistent schema. Anderson strongly recommends two top-level condition types:
Ready: For long-running resources (like Deployments).Succeeded: For finite tasks (like Jobs).
He explicitly advises against Finished due to its ambiguity regarding success or failure.
Crucially, all other condition types should follow a positive polarity, meaning true always indicates a desired or good state (e.g., FooWorked, BarFetched, BazSynced). This allows for straightforward machine-readable summarization:
- Any
falsecondition implies the overallReadystatus isfalse. - Any
unknownconditions (with nofalseconditions) imply theReadystatus isunknown(still processing). - All
trueconditions mean theReadystatus istrue.
This pattern enables UIs to easily represent resource status with visual cues like green/red dots and provides clear signals for automation. Anderson recounts how Knative initially used negative polarity (Failed was true for failure) which led to cognitive confusion (Failed: false meant success).
Enhancing kubectl Experience
Custom resources can be made first-class citizens in the kubectl CLI:
additionalPrinterColumns: Define extra columns to display when users runkubectl get <crd-type>, providing immediate, relevant information. More columns can be shown forwideoutput.shortNames: Provide aliases (e.g.,appforApplication,certforCertificate) to reduce typing.categories: Group CRDs under a common category (e.g.,all) so they appear when users runkubectl get all, preventing them from being "hidden" from general discovery.
Aggregated Objects and CLIs
For complex resources composed of many Kubernetes objects, an aggregated object can provide a simplified user interface. A Knative Service, for instance, aggregates a Deployment, Route, and Configuration, presenting a single, high-level object with a consolidated status (like the service URL). This reduces the complexity users need to manage.
While kubectl is powerful, specialized CLIs can be beneficial for:
- Complex setup workflows:
Flux Bootstrapis cited as an excellent example, orchestrating the entire setup of Flux CD, repository creation, and self-management. This prevents users from getting stuck "a third of the way through" a manual setup. - Interactive status monitoring: CLIs can provide a more user-friendly experience for waiting for resources to become ready, often printing progress and final URLs (e.g.,
kn service create). These could also be presented in a GUI.
Reusing Kubernetes Patterns and Types
Anderson advocates for consistency by reusing established Kubernetes types and design patterns:
- Standard Types: Use
ObjectReferenceandLabelSelectorrather than inventing custom equivalents. - Duct Typing: Adopt existing structural patterns. The
PodTemplateSpecused in Deployments, DaemonSets, and Jobs is a prime example. Knative Services also adopted this pattern, making it easier for users familiar with native Kubernetes workloads to configure Knative services, reducing the relearning curve.
Smart Use of Zero Values
In Go, zero values (e.g., 0 for integers, "" for strings, nil for pointers/slices) are default. Anderson suggests designing CRD schemas such that these zero values have a useful, "do the right thing" meaning. For example, if a policy field is left empty (its zero value), it could mean "apply to all resources in this namespace" instead of "apply to nothing." This leads to shorter, more powerful YAML. However, caution is advised: if 0 is a meaningful value (e.g., scaling to zero replicas), pointers should be used to distinguish between an explicitly set 0 and a nil (unset) value.
GitOps-Friendly Design
Operators must be designed with GitOps in mind. This means resources should be idempotent (reapplying them repeatedly yields the same result without errors) and resilient to the order of application. GitOps tools often "dump a bunch of stuff into the cluster at once," and operators need to gracefully handle out-of-order creation or repeated applications without getting into "half-initialized states."
Interacting with Other Resources
When an operator manages or references other resources, specific Kubernetes mechanisms should be used:
- Labels: For short, searchable values to link or select resources.
- Annotations: For more detailed, unstructured data or configuration (e.g., NGINX Ingress uses many annotations for configuration).
- Owner References: Crucial for garbage collection and establishing parent-child relationships. This allows Kubernetes to automatically delete child resources when the parent is removed and helps users understand which controller created a resource.
Upfront Validation
Finally, Anderson highlights the importance of upfront validation in CRDs using OpenAPI schema validation and CEL (Common Expression Language). These allow defining constraints (e.g., numeric ranges, string patterns) directly in the CRD, providing immediate, useful error messages to users before the invalid resource is even created by the API server. This prevents controllers from receiving malformed input and failing, improving the overall user feedback loop. For example, preventing a pod from being spawned with "64 petabytes of memory" early in the process.
Demo / Proof of Concept
▶ Watch: When to omit Status: Configuration-only CRDs (6:00)
While Evan Anderson's talk does not feature a live, step-by-step demonstration of a specific tool or proof of concept developed by him, he effectively uses existing, well-known Kubernetes projects as illustrative examples of good design principles.
He frequently references the kn service create command from the Knative CLI as an example of a "delightful" CLI experience. This command, when executed, provides real-time updates on the service creation process and ultimately prints the accessible URL upon completion. This showcases how a CLI can abstract complex orchestration into a user-friendly, informative workflow.
Another significant example highlighted is Flux Bootstrap. Anderson describes this as a "really neat" tool that orchestrates an entire setup workflow: creating a Git repository, setting up Flux on a cluster, and configuring Flux to manage itself. He emphasizes its value in handling a complex multi-step process that most users would struggle to complete manually, preventing the common frustration of getting "a third of the way through and something would go wrong."
These examples, rather than a direct demo, serve as compelling proofs of concept for the design philosophies Anderson advocates, demonstrating how thoughtful UX in CRDs and their accompanying tooling can transform complex Kubernetes operations into seamless and positive user experiences.
Defensive Implications
▶ Watch: Using Kubernetes Events for user communication (7:00)
The principles outlined by Evan Anderson, while primarily focused on improving user experience, carry significant defensive implications for Kubernetes environments. A "delightful" CRD and operator design indirectly contributes to a more secure, stable, and resilient cluster.
- Reduced Human Error: Clear, human-readable status messages and well-defined conditions (especially with positive polarity) significantly reduce ambiguity. This means operations teams, particularly during high-stress incidents, can quickly understand the state of a custom resource and the why behind any issues. Less ambiguity leads to fewer misinterpretations and, consequently, fewer erroneous defensive actions or delays in incident response.
- Enhanced Observability and Auditing: The strategic use of Kubernetes Events provides a standardized, namespace-scoped audit trail of significant state changes. This is invaluable for security auditing and post-incident analysis. Defenders can easily see when a resource was created, modified, or encountered critical failures without needing privileged access to controller logs, making it easier to track changes and identify suspicious activities.
- Precise Access Control (RBAC): Proper implementation of aggregated cluster roles ensures that users have appropriate and well-understood access to custom resources. By rolling CRD permissions into standard
view,edit, oradminroles, administrators can maintain a consistent security posture. This prevents scenarios where CRDs are either unintentionally exposed to too many users or, conversely, become inaccessible for legitimate troubleshooting, leading to operational friction or the creation of overly permissive custom roles. - Improved System Stability and Predictability: Designing operators for idempotency and resilience to out-of-order resource application (GitOps-friendly design) contributes directly to system stability. Operators that can gracefully handle repeated applications or fluctuating resource availability are less likely to enter unstable states or trigger cascading failures. This robustness is a key defensive characteristic, as it minimizes the attack surface created by unpredictable system behavior.
- Early Detection of Malformed Configurations: Upfront validation using OpenAPI schemas and CEL expressions is a critical defensive measure. By rejecting invalid custom resource manifests at the API server level, operators can prevent potentially malicious or destabilizing configurations from ever reaching the controller. This acts as a crucial first line of defense, reducing the risk of unexpected behavior, resource exhaustion, or other vulnerabilities that could arise from processing malformed input.
- Clear Ownership and Dependency Management: Utilizing owner references ensures proper garbage collection and establishes clear relationships between custom resources and the Kubernetes objects they create. This is vital for maintaining a clean and understandable resource graph. From a defensive standpoint, clear ownership helps identify the source of resources, aiding in investigations and ensuring that resources are properly cleaned up, preventing resource leakage or orphaned components that could be exploited.
In essence, a delightful CRD experience is not just about convenience; it's about building a more transparent, predictable, and robust Kubernetes environment that is inherently easier to defend against both accidental misconfigurations and malicious attacks.
Key Takeaways
- Prioritize UX in CRD Design: Treat CRD design as a user interface. Clear, predictable, and human-friendly interactions are crucial for developers and operations teams, especially during stressful situations.
- Design Comprehensive Status Fields: Ensure your CRD's status field provides both machine-readable states (enums) for automation and human-readable descriptions for context and debugging, including references to related Kubernetes resources.
- Leverage Kubernetes Events and RBAC Aggregation: Use Kubernetes Events to communicate significant state changes without noise. Implement aggregated cluster roles to seamlessly integrate CRD permissions with standard
view,edit, andadminroles. - Adopt Standardized Conditions with Positive Polarity: Utilize
ReadyandSucceededconditions, and ensure all other conditions use positive polarity (truemeans good) for consistent machine interpretation and UI representation. - Enhance
kubectland Consider Dedicated CLIs: Improvekubectlinteraction withadditionalPrinterColumns,shortNames, andcategories. For complex workflows, consider building dedicated CLIs or GUIs. - Embrace Kubernetes Patterns and Validation: Reuse existing Kubernetes types (
ObjectReference,LabelSelector), adopt "duct typing" patterns (likePodTemplateSpec), make zero values meaningful, and implement OpenAPI/CEL validation for upfront error checking.
About the Speaker(s)
Evan Anderson is a seasoned expert in cloud-native technologies, currently working at Stacklok, Inc. He has a distinguished background, having previously contributed to VMware's Tanzu product and Google Cloud. Anderson is widely recognized as one of the founders of the Knative project, a serverless platform built on Kubernetes, where he gained extensive experience in designing and implementing custom resources. His long history with custom resources dates back to their early days as "ThirdPartyResources," giving him a unique perspective on their evolution and best practices. Anderson is also an active contributor to various CNCF projects, allowing him to observe and influence CRD design across the ecosystem. His insights shared in this talk are largely derived from his direct experiences and observations in these prominent roles, making him a credible voice on the topic of making CRDs more user-friendly and effective.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Evan Anderson's KubeCon talk is a much-needed deep dive into the often-neglected user experience of Kubernetes Custom Resource Definitions (CRDs). Drawing from extensive experience, particularly with Knative, Anderson provides a comprehensive and actionable set of best practices for operator authors. He moves beyond merely making CRDs functional to making them 'delightful,' focusing on clear status fields, intelligent event usage, streamlined RBAC, and robust validation. This isn't just theory; it's a practical roadmap to reduce operational friction and enhance debuggability in the Kubernetes ecosystem.
Heather Calloway (CISO) — STRONG ACCEPT
This talk by Evan Anderson on designing "delightful" Custom Resource Definitions (CRDs) in Kubernetes is a pragmatic and highly valuable contribution. While framed around user experience, its core message directly translates to improved operational resilience, reduced incident response times, and enhanced auditability—all critical dimensions of a robust security program. Anderson provides actionable strategies that platform teams and operator authors should adopt, fostering environments that are not just easier to use, but inherently more defensible and easier to govern.