Ensuring Quality in Kubernetes: The Graduation Process From Alpha T... Antonio Ojea & Benjamin Elder

Antonio Ojea, Benjamin Elder

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk, presented by Antonio Ojea and Benjamin Elder, both from Google and active members of Kubernetes' SIG Testing and Steering committees, delves into the critical processes and recent advancements in maintaining and improving the quality of the Kubernetes project. As Kubernetes has evolved into a vast open-source ecosystem with thousands of dependent projects, ensuring the stability, interoperability, and portability of its features and APIs has become paramount. The speakers meticulously explain the feature graduation lifecycle—from Alpha to Beta and finally to Generally Available (GA)—and the rigorous testing methodologies and policy enforcements designed to uphold Kubernetes' high-quality bar.

Watch on YouTube

Visual summary for Ensuring Quality in Kubernetes: The Graduation Process From Alpha T... Antonio Ojea & Benjamin Elder by Antonio Ojea, Benjamin Elder
Visual summary for Ensuring Quality in Kubernetes: The Graduation Process From Alpha T... Antonio Ojea & Benjamin Elder by Antonio Ojea, Benjamin Elder

Key moments

  1. 0:00 Introduction: Kubernetes quality and graduation process
  2. 1:15 Kubernetes organization and KEPs feature development process
  3. 2:30 Understanding Kubernetes feature graduation: Alpha, Beta, GA
  4. 4:55 The critical role and stability of Kubernetes APIs
  5. 6:40 Challenges with Beta APIs and their default disabling
  6. 8:00 The Production Readiness Group ensuring features are ready
  7. 9:00 How Kubernetes enforces quality through extensive testing

Ensuring Quality in Kubernetes: The Graduation Process From Alpha To GA

Speakers: Antonio Ojea, SIG Testing Member, Steering Member, Google; Benjamin Elder, SIG Testing Member, Steering Member, Google

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=d_9JNRkT7dg

Overview

This talk, presented by Antonio Ojea and Benjamin Elder, both from Google and active members of Kubernetes' SIG Testing and Steering committees, delves into the critical processes and recent advancements in maintaining and improving the quality of the Kubernetes project. As Kubernetes has evolved into a vast open-source ecosystem with thousands of dependent projects, ensuring the stability, interoperability, and portability of its features and APIs has become paramount. The speakers meticulously explain the feature graduation lifecycle—from Alpha to Beta and finally to Generally Available (GA)—and the rigorous testing methodologies and policy enforcements designed to uphold Kubernetes' high-quality bar.

The core of their discussion revolves around the challenges of managing feature development at scale, particularly the complexities introduced by new initiatives like Dynamic Resource Allocation (DRA), which often involve dependencies between features at different maturity levels. They highlight past struggles with CI stability and the evolution of testing practices, culminating in the introduction of more robust tagging systems for end-to-end tests and stricter enforcement of feature gate policies. This presentation is crucial for anyone involved in developing for, operating, or extending Kubernetes, offering deep insights into the mechanisms that underpin the platform's reliability and its commitment to a stable user contract.

Background

▶ Watch: Introduction: Kubernetes quality and graduation process (0:00)

Kubernetes operates as a massive open-source project, organized into Special Interest Groups (SIGs) that can be horizontal (e.g., API Machinery, CLI) or vertical (e.g., Networking, Node). To manage its continuous growth and the addition of new features, Kubernetes employs a structured development process centered around Kubernetes Enhancement Proposals (KEPs). While initially perceived as somewhat heavy, KEPs serve a vital purpose: facilitating transparent communication across SIGs and stakeholders, ensuring that new proposals are thoroughly vetted and understood within the community.

The lifecycle of a Kubernetes feature progresses through three distinct stages, each with increasing stability requirements:

  • Alpha: This initial stage is for experimental features. While expected to be usable for testing, the quality bar is lower, acknowledging that the design may still evolve significantly. Alpha features are typically disabled by default, allowing for rapid innovation and early feedback without impacting production environments.
  • Beta: Features promoted to Beta demonstrate greater stability. If a Beta feature does not depend on APIs, it is usually enabled by default, signifying a stronger, though not absolute, commitment to the user. Users are informed that some aspects might still change, allowing for final refinements.
  • GA (Generally Available): This is the highest maturity level, representing a strong, unwavering commitment to stability and implementation. GA features and APIs are considered stable, production-ready, and serve as a reliable foundation for the extensive Kubernetes ecosystem, ensuring portability and long-term compatibility.

A cornerstone of Kubernetes' success is its APIs, which act as critical contracts defining not only syntax but also the expected behavior and interactions within the system. These well-defined APIs guarantee interoperability and portability, allowing applications to run consistently across different Kubernetes installations and cloud providers. However, the graduation of APIs themselves poses a unique challenge. Since Kubernetes versions 1.21 or 1.22, Beta APIs (and consequently, features dependent on them) have been disabled by default. This change aimed to prevent premature reliance on unstable APIs but also created friction in moving features to GA.

In earlier releases, the project faced significant stability issues, particularly around Kubernetes 1.21-1.23, characterized by numerous "flakes" in the Continuous Integration (CI) system. This instability stemmed from a rapid influx of new features without a sufficiently robust quality assurance mechanism. To address this, the project established the Production Readiness Business Group (PRBG). This group plays a crucial role by asking probing questions about scalability, user control, and operational requirements before a feature can graduate to Beta or GA, ensuring that anything released is truly production-ready.

Quality enforcement in Kubernetes relies heavily on a comprehensive testing strategy, often metaphorically described as a pyramid. This includes unit tests for individual code components, integration tests for interactions between components, and extensive end-to-end (E2E) tests that validate the system's behavior in realistic cluster environments. A critical component of the E2E suite is the conformance suite, which defines the minimum set of tests a Kubernetes cluster must pass to guarantee application portability across different versions and distributions. The responsibility for quality is shared across the community, with SIGs, the CI Signal team, and the Release Team all playing active roles. A foundational principle is the zero flake policy, which dictates that no flaky tests are tolerated, ensuring that CI failures always indicate a genuine issue. The SIG Testing group, while not owning all tests, provides the essential infrastructure, standards, frameworks, and best practices that enable all SIGs to maintain high-quality tests. The talk highlights that new initiatives like Dynamic Resource Allocation (DRA), with their complex dependencies between alpha and beta features, further underscore the need for sophisticated and reliable testing methodologies.

Key Findings

▶ Watch: Understanding Kubernetes feature graduation: Alpha, Beta, GA (2:30)

The speakers presented several key findings and improvements aimed at enhancing the quality and stability of Kubernetes features throughout their graduation lifecycle:

  1. Overloaded "Feature" Tag Remediation: Historically, E2E tests used an arbitrary "feature" string in their names for tagging, which was "super nebulous." This made it difficult to precisely identify test dependencies (e.g., requiring a load balancer, a specific feature gate, or a dual-stack cluster configuration). As a result, most CI jobs would skip any tests with a feature tag, leaving feature developers to set up their own isolated CI. This approach did not scale.
  1. Introduction of GKO Labels for Metadata: With the adoption of GKO v2, Kubernetes testing gained access to GKO labels. These labels allow for moving rich metadata about tests out of the test name and into queryable labels. This transition enables more precise selection and filtering of tests, replacing brittle regular expressions.
  1. Standardized WithFeatureGate Method: A significant improvement in Kubernetes 1.33 is the new WithFeatureGate method for E2E tests. Instead of arbitrary strings, tests now pass the actual standard feature gate definition from Kubernetes' centralized registry of features. This registry contains canonical metadata about each feature, including its owner, alpha/beta/GA state, and default-on status. This allows tests to be automatically annotated with this information, simplifying test selection.
  1. Enhanced Test Tagging for Special Setups: Beyond feature gates, the new system allows for tagging tests that require specific external setups. For instance, a DRA test might be tagged with DynamicAllocationTestFeature to signal that a DRA driver needs to be installed on the cluster before the test can run. This enables CI systems to identify and provision the necessary environment for such tests.
  1. Standardized CI Jobs for Feature Stages: The SIG Testing group is setting up standard CI jobs that can automatically run tests based on their feature gate and API maturity levels (Alpha, Beta, GA). For example, an "Alpha" CI job can be configured to turn on all alpha features and APIs, while skipping tests that require complex external setups. This streamlines testing for a large number of features that are primarily server-side.
  1. Enforcement of "Alpha Must Be Off by Default" Policy: A critical finding was that the Kubernetes feature policy stating "alphas must be off by default" was not being enforced. This led to a situation where an alpha-quality feature (related to host network on Windows) was inadvertently on by default since Kubernetes 1.26 and was only removed in 1.33. New tooling has been implemented to block pull requests if they attempt to set an alpha feature to "on by default," ensuring clear communication of quality bars to users.
  1. Increased Expectations for Alpha Feature Testing: There's now a strong push for feature approvers to ensure that even alpha features have reliable CI tests. The expectation is that features should not be promoted to Beta without robust, working tests. This shifts the focus on stability and testing much earlier in the development lifecycle.
  1. Shifting Stability Expectations Down the Lifecycle: Kubernetes is moving towards a model where stability is prioritized earlier. While the project historically leaned heavily on GA for stability, the goal is to bring more of that rigor to Beta and even Alpha features, ensuring that even alpha features do not destabilize GA components. Alpha test signals are being integrated into release blocking mechanisms to prevent unstable features from progressing.

Technical Deep Dive

▶ Watch: The critical role and stability of Kubernetes APIs (4:55)

The Kubernetes project's commitment to quality is deeply embedded in its technical processes and infrastructure. At the heart of feature development are Kubernetes Enhancement Proposals (KEPs), formal documents that outline new features, their design, and their lifecycle. This structured approach, while seemingly bureaucratic, ensures that all stakeholders, particularly across different SIGs, are aware of and can provide feedback on proposed changes, preventing isolated development and fostering a cohesive system.

The feature lifecycle, from Alpha to Beta to GA, is not merely a label but a set of increasingly stringent technical requirements. Alpha features, while functional and "usable," are understood to have a lower quality bar, allowing for rapid iteration and API changes. Beta features, on the other hand, demand a "stronger requirement for stability." A key technical distinction is that non-API-dependent Beta features are enabled by default, creating a "contract with the end user." This means that the system's behavior, while not entirely locked, is expected to be largely stable, with any breaking changes clearly communicated. GA signifies an "strong, strong commitment" to stability, ensuring that APIs and features provide a dependable foundation for the vast ecosystem. This stability is critical for portability, allowing applications and configurations (like Helm charts or YAML manifests) to function consistently across diverse Kubernetes environments.

The criticality of APIs cannot be overstated. Beyond merely defining data structures, Kubernetes APIs dictate the "interactions with the other component" and "what the system should do." This behavioral contract is what enables true interoperability. The decision to disable Beta APIs by default (since Kubernetes 1.21/1.22) was a technical measure to force features to either mature or be explicitly opted into, preventing unintended reliance on unstable interfaces. This policy change, while initially challenging, underscored the project's dedication to long-term API stability.

To enforce these quality bars, Kubernetes relies on a sophisticated testing infrastructure. Unit tests validate individual code components, integration tests verify interactions between components, and end-to-end (E2E) tests simulate real-world scenarios, running against actual Kubernetes clusters. The conformance suite, a subset of E2E tests, is particularly crucial. It defines a baseline set of behaviors and API guarantees that any Kubernetes distribution must adhere to, directly ensuring portability and interoperability. Without passing conformance tests, a cluster cannot be considered a "conformant" Kubernetes installation, which would undermine the ecosystem's ability to build portable applications.

The SIG Testing group plays a pivotal role by owning the "testing infrastructure, the standards, the frameworks, the best practices." They do not own all tests (that's a shared SIG responsibility), but they provide the tooling and guidance. A significant technical advancement discussed is the move to GKO v2 and the adoption of GKO labels for E2E tests. Previously, test metadata was embedded in the test name as an arbitrary "feature" string, making it difficult to query and manage. GKO labels provide a structured way to attach metadata to tests, allowing for precise selection and filtering.

The new WithFeatureGate method introduced in Kubernetes 1.33 leverages this by allowing tests to reference the canonical feature gate definition. This definition, stored in a centralized registry, contains comprehensive metadata about each feature, including its stability level (alpha, beta, GA), ownership, and whether it's enabled by default. This enables the CI system to automatically infer test requirements and apply appropriate labels. For features requiring special setup, like Dynamic Resource Allocation (DRA), tests can be additionally tagged (e.g., DynamicAllocationTestFeature). This signals to the CI system that a specific environment (e.g., a mock DRA driver) must be provisioned before the test can execute.

This allows for the creation of sophisticated standardized CI queries. For example, an Alpha test job might use a query that selects tests tagged with an "off-by-default feature gate" or no other feature information, while explicitly excluding Beta off-by-default or deprecated features, as well as known "slow, disruptive, or flaky" tests. This granular control ensures that the right tests run in the right context, providing meaningful signal without unnecessary overhead.

A critical enforcement mechanism highlighted is the new tooling that blocks Pull Requests (PRs) if they attempt to set an alpha feature to on by default. This ensures that the stated policy—alpha features are experimental and opt-in—is technically enforced, preventing unintended exposure of unstable code to users. This kind of automated gatekeeping is essential for maintaining the integrity of the feature lifecycle. For specialized testing environments, the project leverages kind clusters (Kubernetes in Docker), which are lightweight and easily reproducible, allowing for the setup of mock drivers or specific configurations (e.g., a load balancer) for E2E testing.

Finally, the zero flake policy is a testament to the project's technical rigor. To achieve this, the community developed tools like triage.k8s.io/triage, which uses K-Nearest Neighbors (KNN) clustering to analyze failure messages across CI jobs. This allows developers to quickly identify patterns in failures, pinpoint systemic issues (like noisy neighbor problems), and ensure that every test failure is a genuine signal, not just noise. This commitment to flake elimination creates a culture where developers are more inclined to address test failures, knowing they indicate real problems.

Demo / Proof of Concept

▶ Watch: The Production Readiness Group ensuring features are ready (8:00)

While the talk did not feature a live, interactive demonstration in the traditional sense, Benjamin Elder described the practical implementation of the enhanced testing framework through a specific example: setting up a CI job for Dynamic Resource Allocation (DRA). This served as a detailed proof-of-concept for how the new tagging and query mechanisms are applied.

The described CI job for DRA illustrates how a complex feature, potentially dependent on multiple alpha and beta components, is validated. The job is configured to:

  1. Tagging: Explicitly tag tests as requiring DynamicResourceAllocation setup.
  2. Feature Gate and API Enablement: Turn on all necessary alpha and beta APIs and features, ensuring the environment is correctly configured for the DRA tests.
  3. Flake Handling: Temporarily ignore known flaky tests until they are resolved, preventing them from blocking critical signal.
  4. Mock Driver Setup: Critically, the job includes "some fun bash that sets up actually having a mock DRA driver." This involves provisioning a local implementation of a DRA driver within the testing environment.
  5. Kind Cluster Usage: The entire setup runs on a kind cluster, a lightweight Kubernetes cluster deployed within Docker containers. This provides an isolated, reproducible, and efficient environment for running these specialized E2E tests.

Benjamin emphasized that these configurations are "pretty like copyable," indicating that the underlying scripts and methodologies are designed to be reusable for other features requiring similar external setups (e.g., a mock load balancer). This described process, rather than a real-time demo, effectively showcased the practical application of the new testing standards and tooling, demonstrating how developers can ensure their features are thoroughly tested within a standardized, yet flexible, CI pipeline.

Defensive Implications

▶ Watch: How Kubernetes enforces quality through extensive testing (9:00)

The advancements in Kubernetes' quality assurance and feature graduation process have significant implications for defenders, including cluster operators, security engineers, and application developers.

Firstly, the stricter enforcement of the Alpha, Beta, GA lifecycle means that operators can have greater confidence in the stability and maturity of features they deploy. The policy of "alpha features must be off by default" is now technically enforced, reducing the risk of accidentally running experimental, unstable, or potentially insecure features in production environments. This proactive approach helps prevent regressions and unexpected behavior that could impact cluster stability or application uptime.

Secondly, the enhanced transparency and specificity in test tagging and CI processes provide clearer signals about the readiness of features. When a feature promotes to Beta or GA, it now comes with a stronger guarantee that it has undergone rigorous, automated testing, including comprehensive E2E and conformance tests. This allows defenders to better assess the reliability of new functionalities and plan their adoption strategies with more certainty. The Production Readiness Business Group (PRBG) also acts as a crucial gatekeeper, ensuring that features meet operational requirements for scalability and control, which directly benefits operators managing large-scale deployments.

Thirdly, the zero flake policy in Kubernetes CI is a powerful defensive mechanism. By eliminating flaky tests, every CI failure becomes a legitimate alert, indicating a potential regression or bug. This prevents "alert fatigue" and ensures that issues are identified and addressed promptly, often before they even merge into the main codebase. This robust CI signal ultimately leads to a more stable and predictable platform, reducing the likelihood of critical vulnerabilities or operational failures stemming from untested changes.

Finally, the continuous investment in shared testing infrastructure and best practices by SIG Testing empowers the entire community to contribute high-quality code. For defenders who might also be contributing to the project, understanding these processes, particularly the use of GKO labels and WithFeatureGate for test annotation, is crucial for writing effective tests that integrate seamlessly into the project's robust CI pipelines. This collective commitment to quality, enforced through tooling and policy, ultimately strengthens the security posture and operational resilience of Kubernetes for all its users.

Key Takeaways

  • Rigorous Feature Graduation: Kubernetes employs a structured Alpha, Beta, GA lifecycle, backed by KEPs, SIGs, and the Production Readiness Business Group (PRBG), to ensure features meet increasingly high stability and production-readiness standards.
  • Standardized Test Tagging: The introduction of GKO labels and the WithFeatureGate method in Kubernetes 1.33 revolutionized E2E test metadata, allowing for precise test selection based on canonical feature gate definitions, replacing arbitrary string tags.
  • Enforced Alpha Feature Policy: A critical new tooling prevents alpha features from being set "on by default," ensuring clear communication of experimental status and preventing accidental deployment of unstable features in production environments.
  • Zero Flake Policy: Kubernetes maintains a strict "zero flake policy" in its CI, supported by tools like triage.k8s.io/triage and strong community accountability, to ensure that test failures are always indicative of real issues, leading to a highly reliable CI system.
  • Early Stability Expectations: The project is shifting its focus on stability earlier in the development lifecycle, demanding robust CI tests even for alpha features and blocking promotion to beta without reliable test coverage, thereby enhancing overall platform resilience.
  • SIG Testing's Foundational Role: SIG Testing provides the essential infrastructure, standards, and frameworks that enable the entire Kubernetes community to develop and maintain high-quality tests, fostering a collaborative approach to quality assurance across the vast ecosystem.

About the Speaker(s)

Antonio Ojea is a key contributor to the Kubernetes project, working at Google. He is an active member of the SIG Testing special interest group, which is responsible for the testing infrastructure, standards, and best practices within Kubernetes. Additionally, Antonio serves as a member of the Kubernetes Steering committee, playing a crucial role in the overall governance and strategic direction of the project.

Benjamin Elder also contributes significantly to the Kubernetes project from Google. Like Antonio, he is a dedicated member of SIG Testing, focusing on the evolution and improvement of Kubernetes' testing frameworks and methodologies. Benjamin is also a member of the Kubernetes Steering committee, contributing to the high-level decision-making and ensuring the project's long-term health and stability. Together, their work is instrumental in upholding the quality and reliability of Kubernetes.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk by Google's SIG Testing and Steering members, Antonio Ojea and Benjamin Elder, provides a crucial deep dive into the technical mechanisms ensuring Kubernetes feature quality and stability. They detail the evolution of the Alpha-Beta-GA graduation process, highlighting critical advancements like GKO labels, the WithFeatureGate method, automated PR blocking for alpha features, and the project's stringent zero-flake CI policy. The presentation offers invaluable insights for anyone building on or operating Kubernetes, demonstrating how a massive open-source project enforces quality and provides a reliable user contract through robust engineering and policy.

Heather Calloway (CISO) — STRONG ACCEPT

This session provides a crucial look into the institutional discipline behind Kubernetes' stability. It details the rigorous feature graduation process, from Alpha to GA, and the robust testing and policy enforcement mechanisms that underpin the platform's reliability. For any organization heavily invested in Kubernetes, understanding this framework is not merely technical curiosity; it is foundational to managing operational risk, ensuring application portability, and making informed decisions about feature adoption. The commitment to quality demonstrated here directly translates into business resilience.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025