Project Lightning Talk: Kubeflow Helm Chart - Krzysztof Romanowski, Maintainer

Krzysztof Romanowski, Maintainer

KubeCon + CloudNativeCon Europe 2025 · Project Lightning Talk

Overview

Krzysztof Romanowski, a lead Kubernetes engineer at Roche, presented a focused talk at KubeCon EU on a critical development for the Kubeflow ecosystem: a new Helm chart designed to streamline its deployment. Recognizing the inherent complexities of deploying Kubeflow, particularly in enterprise environments with diverse requirements and existing infrastructure, Romanowski led an initiative to create a more flexible and manageable installation method. This talk is not merely about Kubeflow itself, but specifically addresses the operational challenges associated with its setup and configuration.

Watch on YouTube

Key moments

  1. 0:00 Introduction and Kubeflow Kustomize deployment challenges
  2. 0:45 Motivation: Streamlining Kubeflow deployments with proper parameterization
  3. 1:15 Helm chart's core assumption: focusing on main components for flexibility
  4. 2:05 Key benefits: IaC integration and cloud provider support
  5. 2:40 Achieving sane defaults despite extensive parameterization options
  6. 3:35 Future plans, community integration, and distribution strategy
  7. 4:15 Call to action: Use the Helm chart and contribute

Project Lightning Talk: Kubeflow Helm Chart

Speakers: Krzysztof Romanowski; Maintainer

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=NkOV4_JV4t4

Overview

Krzysztof Romanowski, a lead Kubernetes engineer at Roche, presented a focused talk at KubeCon EU on a critical development for the Kubeflow ecosystem: a new Helm chart designed to streamline its deployment. Recognizing the inherent complexities of deploying Kubeflow, particularly in enterprise environments with diverse requirements and existing infrastructure, Romanowski led an initiative to create a more flexible and manageable installation method. This talk is not merely about Kubeflow itself, but specifically addresses the operational challenges associated with its setup and configuration.

The core problem tackled by this project stems from Kubeflow's traditional reliance on kustomize for deployment. While powerful, kustomize deployments can become unwieldy for large-scale, production-grade Kubeflow instances, making parameterization and integration with infrastructure-as-code (IaC) tools a significant hurdle. Romanowski’s work directly addresses a long-standing community request, dating back six years, for a Helm-based solution that simplifies Kubeflow deployments, enhances adaptability across different environments, and fosters better integration into broader organizational infrastructure. The significance of this work lies in its potential to democratize Kubeflow adoption for organizations that require robust, standardized, and easily parameterized machine learning platforms on Kubernetes.

Background

▶ Watch: Introduction and Kubeflow Kustomize deployment challenges (0:00)

The journey to developing a dedicated Helm chart for Kubeflow is rooted in the widely acknowledged difficulties of deploying and managing this comprehensive machine learning toolkit on Kubernetes. Historically, Kubeflow has predominantly leveraged kustomize for its installation manifests. While kustomize offers excellent flexibility for modifying Kubernetes resource configurations without templating, its application for a system as vast and interconnected as Kubeflow presents unique challenges, especially in complex, multi-environment enterprise settings.

One of the primary pain points identified by the community, and echoed by Romanowski from his experience at Roche, is the sheer scale and inflexibility of the default kustomize-based deployment. Kubeflow comprises numerous components, including core services, various CRDs (Custom Resource Definitions), and supporting elements like cert-manager. Customizing these deployments for specific production environments—parameterizing secrets, connecting to existing cluster resources, or adhering to corporate policies—often involves significant manual effort and intricate modifications to kustomize overlays. This process is prone to errors, difficult to automate reliably with infrastructure-as-code tools, and hinders the consistent rollout of Kubeflow across development, staging, and production environments.

The demand for a Helm chart solution has been evident for years, with a community issue requesting it opened as far back as six years ago. Despite this persistent need, significant development had not materialized. Romanowski highlighted that internal projects at Roche, a company utilizing Kubeflow extensively, faced these exact challenges. Their desire to simplify deployments, enable proper parameterization, and seamlessly integrate Kubeflow with their existing IaC pipelines became a strong impetus for spearheading this initiative. The underlying problem, therefore, is not just about choosing a deployment tool, but about enabling scalable, maintainable, and enterprise-ready operations for a critical machine learning platform. The goal was to move beyond ad-hoc customizations to a structured, templated approach that provides predictable outcomes and reduces operational overhead.

Key Findings

▶ Watch: Helm chart's core assumption: focusing on main components for flexibility (1:15)

The central "finding" or contribution of Romanowski's project is the successful development and release of a Kubeflow Helm chart that dramatically simplifies its deployment and management. This initiative was driven by a clear set of assumptions and yielded several key results that address the shortcomings of previous deployment methods.

Firstly, a core assumption guiding the development was to focus the Helm chart exclusively on main Kubeflow components. This deliberate decision diverges from the comprehensive approach often seen in kustomize manifests, which might include auxiliary services like cert-manager. The rationale behind this was to cater to larger, more mature environments where organizations already have established methods for deploying common infrastructure components. By focusing on the essentials, the Helm chart provides flexibility, allowing users to integrate Kubeflow into their existing infrastructure without redundant or conflicting deployments of shared services. This modularity ensures that Kubeflow can become "just one part of the bigger picture," fitting seamlessly into broader architectural strategies and company policies.

Secondly, the Helm chart significantly enhances infrastructure-as-code (IaC) integration. The ability to parameterize nearly every aspect of the Kubeflow deployment through a values.yaml file means that configurations can be managed, versioned, and deployed consistently using tools like Terraform, Ansible, or GitOps pipelines. This is a critical improvement over kustomize-based methods, which often require more bespoke scripting for parameter injection. The chart's design facilitates integration with cloud providers, allowing for specific service account annotations, custom configurations, or connections to pre-existing secrets and configuration maps within the Kubernetes cluster. This capability is paramount for enterprise users who need to adhere to strict security and operational standards.

A third key finding relates to deployment flexibility and ease of use. While the values.yaml file for the Helm chart can be extensive, reaching up to 2,000 lines for full parameterization, the project emphasizes sane defaults. This means that for a quick proof-of-concept or a basic deployment, users might only need to provide a minimal values.yaml file, perhaps "zero lines or maybe 10 lines." This balance between deep configurability and out-of-the-box simplicity makes Kubeflow more accessible for initial explorations while still providing the granular control required for production systems. The separation of CRDs into their own dedicated Helm chart is another critical design choice, addressing the challenge posed by the large size of Kubeflow's CRDs and ensuring better management and lifecycle control. These findings collectively establish the Kubeflow Helm chart as a robust, flexible, and operationally friendly alternative for deploying and managing Kubeflow.

Technical Deep Dive

▶ Watch: Key benefits: IaC integration and cloud provider support (2:05)

The technical implementation of the Kubeflow Helm chart is a direct response to the inherent complexities of deploying a large-scale, multi-component system like Kubeflow. Romanowski's approach prioritizes modularity, parameterization, and integration, drawing a clear distinction from the traditional kustomize-based methods.

At its core, the Helm chart operates under the principle of minimal essential components. Instead of bundling every possible dependency, the chart focuses on the "base components" of Kubeflow. This means that common, cluster-wide services often included in kustomize manifests, such as cert-manager, are explicitly excluded. The rationale is that in larger organizations, these components are typically managed centrally by platform teams, often with their own Helm charts or deployment strategies, to adhere to specific company policies, security requirements, or existing operational frameworks. By omitting them, the Kubeflow Helm chart avoids potential conflicts and offers greater flexibility for integration into diverse IT landscapes.

The chart includes the main Kubeflow components necessary for a functional machine learning platform, alongside "some of the supporting integration manifests." These include Kubernetes Virtual Services and Gateways, which are crucial for proxying traffic and integrating with other services like KNative. This selective inclusion ensures that the core networking and ingress requirements for Kubeflow are met, while still allowing for external management of broader service mesh or API gateway solutions if desired.

A significant technical decision highlighted in the talk is the handling of Custom Resource Definitions (CRDs). Kubeflow relies heavily on numerous and often substantial CRDs to define its various machine learning resources (e.g., experiments, pipelines, notebooks). Due to their size and the specific lifecycle management requirements of CRDs in Kubernetes, Romanowski's team created a separate Helm chart specifically for Kubeflow's CRDs. This separation offers several advantages:

  1. Lifecycle Management: CRDs often require a different deployment order and update strategy than other Kubernetes resources. Deploying them separately ensures they are installed first and can be managed independently without interfering with the application components.
  2. Modularity: It prevents the main Kubeflow chart from becoming excessively large and complex, improving readability and maintainability.
  3. Conflict Avoidance: In environments where multiple applications might share or have conflicting versions of CRDs, managing them distinctly can prevent cluster-wide issues.

The cornerstone of Helm's utility is its values.yaml file, and the Kubeflow Helm chart fully leverages this for extensive parameterization. Romanowski noted that the values.yaml file "contains of 2,000 lines," indicating the depth of configurability available. This extensive parameterization allows users to customize virtually every aspect of the Kubeflow deployment:

  • Resource limits and requests: Fine-tuning performance and cost.
  • Storage configurations: Connecting to specific persistent volumes or storage classes.
  • Network settings: Defining ingress rules, service types, and internal communication.
  • Security parameters: Specifying Service Account annotations, role bindings, and integration with existing secrets.
  • Cloud provider integrations: Customizing configurations for specific cloud services.

Despite this vast configurability, the chart is designed with "sane defaults," meaning that for a basic deployment or a quick test, users might only need a values.yaml file with "maybe zero lines, maybe 10 lines." This balance between comprehensive control and ease of initial use is critical for broader adoption. The project's GitHub repository, mentioned by Romanowski, serves as the central hub for this development, offering "quick start scripts" and detailed documentation on how to use the chart, integrate it into different environments, and contribute to its ongoing evolution. This technical architecture directly addresses the long-standing community need for a robust, flexible, and maintainable Kubeflow deployment solution.

Demo / Proof of Concept

▶ Watch: Future plans, community integration, and distribution strategy (3:35)

While the talk itself was a "Lightning Talk" and did not feature a live, step-by-step demonstration of the Helm chart in action, Krzysztof Romanowski clearly articulated the proof of concept through the existence and capabilities of the chart itself. The success of the project, driven by internal needs at Roche and now released to the community, serves as the primary evidence of its efficacy.

Romanowski explicitly mentioned the availability of the Helm chart in a GitHub repository, stating, "if you want to reach out see the cup repository there is a new release there is six information how to use that there is quick start scripts actually a few of them so you can integrate with different." This indicates that the "demo" or "proof of concept" is provided through these readily available resources. Users are encouraged to clone the repository, utilize the provided quick start scripts, and deploy Kubeflow using the Helm chart to experience its benefits firsthand.

The conceptual demonstration focused on the outcomes:

  • Simplified Deployment: The ability to deploy a functional Kubeflow instance with minimal configuration, often requiring only a few lines in the values.yaml file for basic setups.
  • Enhanced Parameterization: The capacity to extensively customize the deployment for diverse environments, from development to production, by modifying the comprehensive values.yaml file.
  • Infrastructure-as-Code Integration: The theoretical demonstration of how the Helm chart seamlessly fits into existing IaC pipelines, allowing for automated, repeatable, and version-controlled deployments.
  • Modular Architecture: The demonstration of how the chart's focus on main components and separate CRD chart allows for better integration into complex enterprise Kubernetes clusters, where other services are already managed.

Although a live, interactive demo was not part of the presentation, the speaker's emphasis on the project's practical utility for Roche and its release to the wider community, complete with documentation and quick start guides, effectively served as a strong proof of concept for the operational benefits of the Kubeflow Helm chart.

Defensive Implications

▶ Watch: Call to action: Use the Helm chart and contribute (4:15)

While this talk is primarily focused on operational efficiency and deployment simplification rather than traditional cybersecurity vulnerabilities, the Kubeflow Helm chart has significant defensive implications for organizations operating machine learning platforms on Kubernetes. A well-designed, parameterized, and consistently deployable system inherently contributes to a stronger security posture.

Firstly, the robust parameterization capabilities of the Helm chart allow defenders to enforce security best practices from the very beginning of the deployment process. Instead of manual, error-prone configurations, security teams can define hardened values.yaml templates that incorporate:

  • Least Privilege: Specifying precise Service Account annotations and Kubernetes Role-Based Access Control (RBAC) configurations to ensure Kubeflow components operate with only the necessary permissions.
  • Network Segmentation: Configuring network policies, ingress rules, and internal communication patterns to isolate Kubeflow components and restrict unauthorized access.
  • Secret Management: Integrating with existing cluster secret stores or external vault solutions, ensuring sensitive credentials are not hardcoded but securely injected.
  • Resource Quotas: Preventing resource exhaustion attacks by defining appropriate CPU, memory, and storage limits for all Kubeflow pods.

Secondly, the Helm chart's emphasis on infrastructure-as-code (IaC) integration directly supports security automation and compliance. By treating Kubeflow deployments as code, organizations can implement GitOps workflows where all changes are version-controlled, reviewed, and automatically applied. This provides an auditable trail of all configurations, making it easier to track changes, detect unauthorized modifications, and demonstrate compliance with regulatory requirements. Automated deployments reduce human error, which is a frequent source of security misconfigurations.

Thirdly, the modular nature of the Helm chart, focusing only on core Kubeflow components and separating CRDs, allows for better integration with existing security infrastructure. Organizations often have established tools and policies for managing cert-manager, ingress controllers, or identity providers. By not bundling these, the Kubeflow Helm chart allows defenders to leverage their existing, hardened solutions for these auxiliary services, rather than deploying potentially less secure or unmanaged duplicates. This flexibility ensures that Kubeflow can be deployed within an organization's pre-approved security architecture.

Finally, the availability of "sane defaults" for quick deployments, coupled with comprehensive documentation, helps prevent insecure default configurations from being used in production. While quick starts might prioritize functionality, the underlying parameterization ensures that security teams can easily override these defaults with production-grade settings, fostering a culture of secure-by-design deployments. In essence, by providing a structured, repeatable, and auditable deployment mechanism, the Kubeflow Helm chart helps reduce the attack surface, minimize misconfigurations, and enhance the overall operational resilience and security posture of Kubeflow environments.

Key Takeaways

  • Addresses Long-Standing Need: The Kubeflow Helm chart directly tackles a six-year-old community request for a simplified, Helm-based deployment solution, moving beyond the complexities of kustomize for large-scale environments.
  • Focus on Core Components: The chart is designed to deploy only the essential Kubeflow components, allowing organizations to integrate it seamlessly with existing cluster services and adhere to corporate IT policies without redundant deployments.
  • Enhanced Infrastructure-as-Code Integration: It provides extensive parameterization through a comprehensive values.yaml file, enabling robust integration with IaC tools and facilitating automated, repeatable deployments across diverse environments and cloud providers.
  • Modular CRD Management: Kubeflow's numerous and large Custom Resource Definitions (CRDs) are managed in a separate Helm chart, improving lifecycle management, modularity, and preventing conflicts with other cluster resources.
  • Balance of Simplicity and Control: The chart offers "sane defaults" for quick starts (requiring minimal values.yaml changes) while providing up to 2,000 lines of configuration options for granular control in production settings.
  • Community-Driven and Open Source: Developed out of internal enterprise needs at Roche, the Helm chart is released to the community, encouraging adoption, reuse by other Kubeflow distributions, and further collaboration.

About the Speaker(s)

Krzysztof Romanowski, who prefers to be called Raman, is a Lead Kubernetes Engineer at Roche. His work focuses on leveraging Kubernetes to manage complex infrastructure and applications. At Roche, he spearheaded the development of the Kubeflow Helm chart, driven by the company's internal need to streamline and standardize Kubeflow deployments across various projects. Romanowski is a maintainer and active contributor, passionate about simplifying complex deployments and enabling robust infrastructure-as-code practices for the community.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This talk by Krzysztof Romanowski delivers a critical, long-overdue solution for the Kubeflow ecosystem: a robust Helm chart. For years, deploying Kubeflow has been an operational nightmare due to its reliance on unwieldy kustomize manifests, especially for enterprises needing parameterization and IaC integration. Romanowski's project directly addresses this six-year-old community pain point with an elegantly engineered solution that focuses on core components, separates CRDs, and offers unparalleled configurability. This isn't just an improvement; it's a foundational shift that will significantly reduce operational overhead, enhance security posture through automated deployments, and…

Heather Calloway (CISO) — STRONG ACCEPT

This talk presents a critical operational improvement for deploying Kubeflow in enterprise environments, addressing a long-standing challenge with a new Helm chart. By enabling robust parameterization and seamless integration with infrastructure-as-code, this initiative significantly enhances an organization's ability to enforce security policies, reduce operational overhead, and ensure consistent, auditable deployments of vital machine learning platforms, thereby directly mitigating business risk.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025