Simplifying Apache Kafka on Kubernetes With Strimzi - Paolo Patierno & Gantigmaa Selenge, Red Hat

Paolo Patierno, Gantigmaa Selenge, Red Hat

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk, delivered by Paolo Patierno and Gantigmaa Selenge (Tina), both Software Engineers at Red Hat and core maintainers of Strimzi, provides an in-depth look at the evolution of running Apache Kafka on Kubernetes. It highlights how Strimzi, a CNCF incubating project, simplifies the operational complexities of managing Kafka clusters in a cloud-native environment. The presentation focuses on recent significant advancements, particularly the adoption of Kafka Raft Metadata (Kraft), the introduction of Tiered Storage, and enhancements in Auto Rebalancing capabilities.

Watch on YouTube

Visual summary for Simplifying Apache Kafka on Kubernetes With Strimzi - Paolo Patierno & Gantigmaa Selenge, Red Hat by Paolo Patierno, Gantigmaa Selenge, Red Hat
Visual summary for Simplifying Apache Kafka on Kubernetes With Strimzi - Paolo Patierno & Gantigmaa Selenge, Red Hat by Paolo Patierno, Gantigmaa Selenge, Red Hat

Key moments

  1. 0:00 Introduction to StreamZ: running Kafka on Kubernetes
  2. 2:00 Why StreamZ? Solving Kafka's operational complexity on Kubernetes
  3. 4:00 StreamZ's day-2 operations, monitoring, and CNCF integrations
  4. 4:30 Introducing Kraft: removing Zookeeper dependency for Kafka metadata
  5. 5:30 Kraft architecture: Kafka nodes serving as controllers and brokers
  6. 7:30 Kraft deployment modes: designated, combined, and single-node clusters
  7. 8:30 Migrating existing Zookeeper-based Kafka clusters to Kraft

Simplifying Apache Kafka on Kubernetes With Strimzi

Speakers: Paolo Patierno, Software Engineer, Red Hat; Gantigmaa Selenge, Software Engineer, Red Hat

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=sLFmnCyV89M

Overview

This talk, delivered by Paolo Patierno and Gantigmaa Selenge (Tina), both Software Engineers at Red Hat and core maintainers of Strimzi, provides an in-depth look at the evolution of running Apache Kafka on Kubernetes. It highlights how Strimzi, a CNCF incubating project, simplifies the operational complexities of managing Kafka clusters in a cloud-native environment. The presentation focuses on recent significant advancements, particularly the adoption of Kafka Raft Metadata (Kraft), the introduction of Tiered Storage, and enhancements in Auto Rebalancing capabilities.

The speakers meticulously detail the technical underpinnings and practical benefits of these features, demonstrating how they address long-standing challenges in Kafka deployments, such as Zookeeper dependency, storage costs, and manual operational tasks. Furthermore, the talk offers a forward-looking perspective, outlining exciting upcoming features like improved certificate management, self-healing clusters, and the transition to Strimzi 1.0 with v1 APIs. This article serves as an essential resource for platform engineers, SREs, and developers seeking to leverage Kafka's power efficiently within Kubernetes, providing critical insights into optimizing performance, scalability, and operational overhead.

The importance of this talk stems from Kafka's pervasive role as a leading distributed event streaming platform and the increasing trend of deploying it on Kubernetes. Strimzi's continuous development directly impacts the ease, reliability, and cost-effectiveness of these deployments, making the discussed features and future roadmap highly relevant for anyone involved in modern data streaming architectures.

Background

▶ Watch: Introduction to StreamZ: running Kafka on Kubernetes (0:00)

Apache Kafka has cemented its position as the premier distributed event streaming platform, renowned for its horizontal scalability, high availability, and fault tolerance. It serves as the backbone for real-time data ingestion, processing, and distribution across a myriad of use cases, from database event capture to powering event-driven applications and cloud services. Originally developed by LinkedIn and open-sourced under the Apache Software Foundation, Kafka's robust architecture has made it indispensable in modern data ecosystems.

Despite its powerful capabilities, Kafka is operationally complex. Managing a Kafka cluster, especially in production, involves intricate tasks such as configuration, scaling, upgrades, and ensuring data consistency and availability. This complexity led to a significant challenge when attempting to run Kafka on Kubernetes, the de facto standard for container orchestration. While Kubernetes offers excellent primitives for deploying distributed applications, it lacks inherent Kafka-specific knowledge. This gap meant that simply deploying Kafka containers on Kubernetes did not automatically confer the high availability, performance, and operational simplicity that users expected.

This is precisely where Strimzi comes into play. Strimzi is an open-source project, an incubating project within the CNCF, designed to manage Apache Kafka on Kubernetes in a Kubernetes-native way. It leverages the Kubernetes Operator pattern, extending the Kubernetes API with Custom Resources (CRDs) to define and manage Kafka components. By embedding Kafka-specific operational knowledge directly into its operators, Strimzi automates day-one operations (installation) and day-two operations (upgrades, scaling, certificate management, configuration, monitoring, and data balancing). It supports core Kafka components like Kafka Connect for external system integration, MirrorMaker for disaster recovery replication (specifically MirrorMaker 2), and an HTTP Bridge for alternative connectivity. Strimzi also integrates with other CNCF projects such as OpenTelemetry, Prometheus, and KEDA for comprehensive monitoring and autoscaling. Historically, Kafka relied on Apache Zookeeper for metadata management, a separate distributed system that added another layer of operational burden and posed scalability limitations, a problem that the Kraft protocol aims to solve.

Key Findings

▶ Watch: StreamZ's day-2 operations, monitoring, and CNCF integrations (4:00)

The talk highlighted three pivotal advancements within Strimzi and Apache Kafka that significantly enhance the platform's operational efficiency, scalability, and cost-effectiveness on Kubernetes:

  1. Kafka Raft Metadata (Kraft): This represents a monumental shift in Kafka's architecture by entirely removing the dependency on Zookeeper for metadata management. Kraft replaces Zookeeper with Kafka's own Raft-based consensus protocol, centralizing metadata within Kafka itself. This simplification streamlines deployment, reduces operational overhead by eliminating a separate system to manage, and improves scalability and performance by removing the need for metadata synchronization between Kafka and Zookeeper. Strimzi has integrated Kraft support, enabling users to deploy Zookeeper-less Kafka clusters and providing a semi-automatic migration path for existing clusters.
  1. Tiered Storage: Addressing the challenges of cost and scalability for long-term data retention, Tiered Storage allows Kafka to offload older, less frequently accessed message segments from local broker disks to cheaper, more scalable remote storage solutions like Amazon S3, Azure Blob Storage, or Google Cloud Storage. This innovation significantly reduces the on-broker disk space requirements, leading to lower operational costs, improved scalability by separating compute from storage, and faster recovery/rebalancing operations by minimizing the data brokers need to manage locally. Strimzi facilitates the configuration and management of Tiered Storage through its custom resources, allowing users to integrate custom remote storage managers.
  1. Auto Rebalancing on Cluster Scaling: Building upon the existing integration with Cruise Control, Kafka's data balancing tool, Strimzi introduces automatic rebalancing capabilities during cluster scaling operations. Previously, scaling a Kafka cluster (adding or removing brokers) required manual intervention to initiate a rebalance using the KafkaRebalance custom resource to redistribute partitions. With auto rebalancing, the Strimzi operator now intelligently triggers and manages the rebalancing process as part of the scaling workflow, ensuring that new brokers receive existing partitions or that partitions are safely moved off brokers before they are removed. This significantly reduces manual effort, minimizes the risk of partition unavailability, and improves overall cluster health and load distribution.

These three features collectively represent a substantial leap forward in making Kafka deployments on Kubernetes more robust, cost-efficient, and operationally friendly, solidifying Strimzi's role as a critical tool for modern data streaming architectures.

Technical Deep Dive

▶ Watch: Introducing Kraft: removing Zookeeper dependency for Kafka metadata (4:30)

The technical depth of the talk centered on how Strimzi implements and manages these new Kafka features, transforming complex operational tasks into Kubernetes-native configurations.

Kafka Raft Metadata (Kraft)

The most transformative feature discussed is Kraft, which fundamentally re-architects Kafka's metadata layer. Traditionally, Kafka relied on Zookeeper to store critical metadata such as partition assignments, topic configurations, and controller elections. This introduced a separate distributed system that needed to be managed, scaled, and secured, adding significant operational overhead. Kraft replaces Zookeeper with Kafka's own Raft-based protocol, where metadata is managed directly by a subset of Kafka brokers designated as controller nodes.

In a Kraft cluster, there are no separate Zookeeper nodes. Instead, Kafka nodes can operate in different roles:

  • Controller nodes: These nodes form a quorum using the Raft protocol. One controller is elected as the active controller, responsible for all metadata updates, partition assignments, and leader elections. The metadata topic is stored locally on each controller node and replicated across the quorum.
  • Broker nodes: These nodes primarily handle client data (producing and consuming messages). They store a local copy of the metadata topic and communicate with the active controller for metadata updates. Unlike Zookeeper-based clusters where a broker node could also be designated as the controller, in Kraft, the roles are distinct, allowing brokers to focus solely on data operations.

Strimzi supports various deployment modes for Kraft:

  • Designated mode: Recommended for large-scale production environments, this setup dedicates specific nodes to act as controllers and others as brokers. This separation of concerns allows for better scalability and resource allocation.
  • Combined mode: Some or all nodes act as both controllers and brokers. This mode is resource-efficient for smaller clusters but requires careful consideration of CPU and memory as nodes perform dual roles.
  • Single-node cluster: Ideal for development and testing, a single Kafka node can serve as both controller and broker, spinning up a functional cluster quickly.

Migrating existing Zookeeper-based clusters to Kraft is a non-trivial, multi-phase process due to architectural differences. Strimzi provides semi-automatic migration support, driven by user annotations on the Kafka custom resource. This phased approach allows for rollback options up until the final migration step, offering a safety net for critical production environments.

The timeline for Kraft integration with Strimzi was detailed:

  • Kraft was announced in 2019.
  • Strimzi 0.29 introduced experimental support for Kraft.
  • Strimzi 0.40 enabled Kraft by default for new clusters.
  • Kafka 4.0, released recently, completely removed Zookeeper support.
  • Strimzi 0.45, the current version, supports Kafka 3.8.0 and 3.9.0 and is the last version with Zookeeper support, offering an extended support period of approximately one year.
  • Strimzi 0.46, the next planned release, will only support Kraft mode, supporting Kafka 3.9.0 and 4.0.0, and will also remove deprecated components like MirrorMaker 1.

Tiered Storage

Tiered Storage is a crucial feature for managing storage costs and scaling Kafka for long-term data retention. It enables offloading older, less frequently accessed message segments from the expensive, high-performance local disks of Kafka brokers to cheaper, object-based cloud storage (e.g., Amazon S3, Azure Blob Storage, GCS).

The benefits of Tiered Storage are multifaceted:

  • Cost Efficiency: Reduces the need for large, expensive local SSDs on brokers by moving historical data to cheaper object storage.
  • Scalability: Decouples compute (brokers) from storage, allowing independent scaling. Brokers can focus on processing active data, while older data resides in scalable cloud storage.
  • Faster Recovery and Rebalancing: With less data stored locally, brokers can recover faster from failures as they have smaller logs to reconstruct. Similarly, rebalancing partitions across brokers becomes quicker because less data needs to be physically moved.
  • Simplified Cluster Operations: Eliminates the need for frequent disk expansions on brokers or manual deletion of old segments, as older data is automatically moved to remote storage.

Strimzi integrates Tiered Storage via a new tierStorage field within the Kafka custom resource. This field allows users to specify the remote storage manager implementation. Kafka provides interfaces for custom storage managers, allowing users to build and package their own JARs into custom Kafka images based on Strimzi's. The current support is primarily for custom implementations, with Strimzi exploring an open-source plugin that supports Amazon S3, GCS, and Azure Blob Storage. The plan is to evolve towards a more strongly typed API for various storage providers as Tiered Storage, which became General Availability (GA) in Kafka 3.9, matures.

Auto Rebalancing on Cluster Scaling

Strimzi's integration with Cruise Control, a tool for dynamically rebalancing Kafka clusters, has been enhanced with auto rebalancing capabilities during cluster scaling events. Previously, scaling a Kafka cluster (adding or removing brokers) required manual steps:

  1. Scale-up: Add new brokers, then manually create a KafkaRebalance custom resource to instruct Cruise Control to move existing partitions to the new brokers. Without this, new brokers would only receive partitions for newly created topics.
  2. Scale-down: Manually create a KafkaRebalance custom resource to move all partitions off the brokers to be removed, ensuring they are empty. Only then could the brokers be safely scaled down to prevent data unavailability or under-replicated partitions.

The new auto rebalancing feature streamlines this process. Users simply adjust the replicas count in their Kafka custom resource, and the Strimzi operator handles the rest:

  • On scale-up: The operator adds the new brokers, then automatically initiates a rebalancing operation using Cruise Control to distribute partitions across all brokers, including the new ones.
  • On scale-down: The operator first triggers a rebalancing to empty the target brokers of their partitions, and only once they are clear, proceeds with scaling down and removing the brokers.

Users can still specify goals for the rebalancing (e.g., CPU utilization, disk usage, network throughput) using a KafkaRebalance template within the Kafka custom resource. The operator automatically injects the necessary mode (add/remove brokers) and broker IDs into the Cruise Control request, abstracting away the manual orchestration. This significantly reduces operational burden, minimizes human error, and ensures continuous data availability and optimal load distribution during scaling events.

Demo / Proof of Concept

▶ Watch: Kraft deployment modes: designated, combined, and single-node clusters (7:30)

The talk focused on a detailed explanation of the features and their underlying mechanisms rather than a live demonstration or a proof of concept. The speakers walked through architectural diagrams and configuration examples, illustrating how Kraft, Tiered Storage, and Auto Rebalancing operate within a Strimzi-managed Kafka cluster on Kubernetes. While no live demo was performed, the comprehensive technical descriptions provided a clear understanding of how these functionalities work in practice and how they can be configured using Strimzi's custom resources.

Defensive Implications

▶ Watch: Migrating existing Zookeeper-based Kafka clusters to Kraft (8:30)

The advancements in Strimzi and Apache Kafka carry significant defensive implications for organizations operating streaming data platforms on Kubernetes.

  1. Simplified Architecture and Reduced Attack Surface (Kraft):
  • Reduced Operational Complexity: By eliminating Zookeeper, the operational burden of managing a separate distributed system is removed. This means fewer components to monitor, patch, and secure, leading to a more streamlined and robust architecture.
  • Improved Reliability: A single, unified system for both data and metadata reduces points of failure associated with inter-system communication and synchronization. Defenders have fewer distinct services to secure and fewer integration points to worry about being compromised.
  • Consistent Security Model: Kafka's security mechanisms (authentication, authorization, TLS) can now be uniformly applied across the entire cluster, encompassing metadata management, without needing to integrate and manage Zookeeper's separate security controls.
  1. Cost-Effective Data Retention and Faster Recovery (Tiered Storage):
  • Enhanced Disaster Recovery (DR) and Business Continuity (BC): Offloading historical data to highly durable cloud storage can improve DR strategies. In the event of a broker failure or even a cluster-wide outage, older data is safely persisted in remote storage, facilitating faster recovery of active data on new brokers.
  • Long-Term Audit Trails and Compliance: The ability to economically store vast amounts of historical data allows organizations to maintain longer audit trails, crucial for compliance with various regulatory requirements (e.g., GDPR, HIPAA). This ensures data availability for forensic analysis without incurring prohibitive costs for local storage.
  • Operational Resilience: Faster broker recovery times mean less downtime in the face of failures, enhancing the overall resilience of the streaming platform.
  1. Automated and Safer Scaling Operations (Auto Rebalancing):
  • Minimizing Human Error: Automating the rebalancing process during scaling operations drastically reduces the chances of human error that could lead to data unavailability, under-replicated partitions, or uneven load distribution. This enhances the stability and security posture of the cluster during dynamic changes.
  • Ensured Data Availability: The operator's intelligent handling of partition movement ensures that data remains accessible and properly replicated throughout scaling events, preventing service interruptions that could be exploited or cause data loss.
  • Consistent Performance: Automated rebalancing helps maintain optimal performance characteristics across the cluster by ensuring an even distribution of load, preventing "hot spots" that could lead to performance degradation or denial-of-service conditions.
  1. Future-Proofing and Enhanced Security Controls (Upcoming Features):
  • Improved Certificate Management (Cert-Manager Integration): Direct integration with external certificate management systems like Cert-Manager allows organizations to leverage their existing Public Key Infrastructure (PKI) and certificate lifecycle management tools. This enables automated certificate rotation, consistent policy enforcement, and reduced risk associated with expired or manually managed certificates, significantly bolstering TLS security across the cluster.
  • Kafka Cluster Self-Healing: Integrating Cruise Control's anomaly detection and self-healing capabilities offers proactive defense against operational issues like broker failures, disk failures, or goal violations. While requiring careful monitoring to ensure user awareness, this feature can automatically address detected problems, reducing mean time to recovery and improving overall resilience.
  • Gateway API Support: Adopting the Gateway API for external cluster exposure provides a more standardized, robust, and potentially more secure way to manage ingress traffic to Kafka. This framework offers finer-grained control over routing and policy enforcement compared to traditional Ingress, potentially reducing the attack surface at the edge of the Kafka cluster.
  • Stretch Clusters: While challenging due to latency sensitivity, the exploration of stretch clusters across Kubernetes clusters (e.g., in metropolitan area networks) offers a path toward enhanced disaster recovery strategies, providing high availability even in the face of regional outages.

In summary, Strimzi's continuous evolution, particularly with Kraft, Tiered Storage, and automated rebalancing, directly translates into a more secure, resilient, and manageable Kafka platform on Kubernetes. Defenders gain tools that simplify operations, reduce attack surfaces, improve data durability, and automate critical tasks, allowing for a more proactive and effective security posture.

Key Takeaways

  • Zookeeper is Deprecated and Replaced by Kraft: Kafka Raft Metadata (Kraft) eliminates Kafka's dependency on Zookeeper, simplifying deployment, reducing operational overhead, and improving scalability by centralizing metadata management within Kafka itself. Strimzi 0.46 and Kafka 4.0 will be Kraft-only.
  • Tiered Storage for Cost Efficiency and Scalability: This feature allows offloading older Kafka message segments to cheaper, scalable cloud storage (e.g., S3), significantly reducing local disk costs, improving recovery times, and simplifying cluster operations by separating compute from long-term storage.
  • Automated Rebalancing with Cluster Scaling: Strimzi now integrates Cruise Control to automatically rebalance Kafka partitions when brokers are added or removed, streamlining scaling operations, reducing manual effort, and ensuring optimal load distribution and data availability.
  • Enhanced Security and Management with Cert-Manager: Future Strimzi releases will integrate with Cert-Manager, enabling users to manage Kafka component certificates through a robust, pluggable PKI system, enhancing TLS security and certificate lifecycle management.
  • Strimzi 1.0 and v1 APIs on the Horizon: Following the removal of Zookeeper support, Strimzi is progressing towards its 1.0 release, which will introduce stable v1 APIs, providing a mature and consistent interface for managing Kafka on Kubernetes.
  • Self-Healing and Gateway API for Future Resilience: Strimzi is exploring integrating Cruise Control's anomaly detection for self-healing Kafka clusters and adopting the Gateway API for more robust and standardized external exposure of Kafka services, further enhancing operational resilience and network security.

About the Speaker(s)

Paolo Patierno is a Software Engineer at Red Hat, specializing in messaging and data streaming technologies. His work primarily focuses on Apache Kafka and Strimzi, where he serves as one of the core maintainers. Paolo is deeply involved in the development of Strimzi, a CNCF incubating project dedicated to running Kafka on Kubernetes.

Gantigmaa Selenge, who goes by Tina, is also a Software Engineer at Red Hat. She is a contributor to both Strimzi and Apache Kafka, working alongside Paolo to advance the capabilities of these critical streaming platforms. Her contributions help ensure that Kafka on Kubernetes remains robust, scalable, and operationally efficient.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk provides a highly valuable and technically deep dive into Strimzi's advancements for running Apache Kafka on Kubernetes. The speakers, core maintainers of Strimzi, meticulously detail the implementation and benefits of Kafka Raft Metadata (Kraft), Tiered Storage, and automated rebalancing. These features collectively address significant operational overhead, cost, and scalability challenges, making Kafka deployments on Kubernetes far more robust and efficient. While lacking a live demo, the comprehensive technical explanations and clear roadmap make this an essential session for platform engineers and SREs.

Heather Calloway (CISO) — STRONG ACCEPT

This talk presents critical advancements in managing Apache Kafka on Kubernetes via Strimzi, detailing the adoption of Kraft, Tiered Storage, and Auto Rebalancing. While technically dense, the implications for enterprise resilience, cost management, and operational security are substantial. It offers a clear path for platform teams to modernize their streaming infrastructure, directly addressing governance challenges and reducing risk exposure. This is highly relevant for any organization relying on Kafka at scale.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025