Rook: Intro and... Travis Nielsen, Madhu Rajanna, Artem Torubarov, Deepika Upadhyay & Sebastien Han

Travis Nielsen, Madhu Rajanna, Artem Torubarov, Deepika Upadhyay, Sebastien Han

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk provides a comprehensive exploration of Rook, an open-source Kubernetes operator designed to automate the deployment and management of Ceph distributed storage within cloud-native environments. Presented by a panel of core contributors and maintainers, the session delves into Rook's architecture, its integration with Kubernetes via the Container Storage Interface (CSI), and advanced features for block, file, and object storage. The speakers highlight how Rook empowers organizations to run highly scalable, performant, and resilient storage solutions on-premises, in hybrid clouds, or even across multiple cloud providers, effectively addressing the limitations of traditional cloud-native storage approaches.

Watch on YouTube

Visual summary for Rook: Intro and... Travis Nielsen, Madhu Rajanna, Artem Torubarov, Deepika Upadhyay & Sebastien Han by Travis Nielsen, Madhu Rajanna, Artem Torubarov, Deepika Upadhyay, Sebastien Han
Visual summary for Rook: Intro and... Travis Nielsen, Madhu Rajanna, Artem Torubarov, Deepika Upadhyay & Sebastien Han by Travis Nielsen, Madhu Rajanna, Artem Torubarov, Deepika Upadhyay, Sebastien Han

Key moments

  1. 1:06 Introduction to Rook and its purpose with Ceph
  2. 2:00 Rook architecture, Ceph integration, and scalability
  3. 4:00 Rook deployment scenarios: cloud, on-prem, hybrid
  4. 6:00 New features in Rook 1.16 release
  5. 6:42 Planned features for Rook 1.17 release
  6. 7:00 CSI driver overview and architecture
  7. 8:00 Key features and capabilities of CSI driver

Rook Storage for Kubernetes: Leveraging Ceph for Scalable and Resilient Cloud-Native Data Management

Speakers: Travis Nielsen, Original Creator of Rook; Madhu Rajanna, Maintainer of Ceph CSI, IBM; Artem Torubarov, Object Storage Developer, Clyo; Deepika Upadhyay, Rook Developer, Clyiso; Sebastien Han

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=xGywrHPAMms

Overview

This talk provides a comprehensive exploration of Rook, an open-source Kubernetes operator designed to automate the deployment and management of Ceph distributed storage within cloud-native environments. Presented by a panel of core contributors and maintainers, the session delves into Rook's architecture, its integration with Kubernetes via the Container Storage Interface (CSI), and advanced features for block, file, and object storage. The speakers highlight how Rook empowers organizations to run highly scalable, performant, and resilient storage solutions on-premises, in hybrid clouds, or even across multiple cloud providers, effectively addressing the limitations of traditional cloud-native storage approaches.

The presentation emphasizes Rook's maturity, having graduated from the CNCF almost five years ago, and its widespread adoption in production environments. It covers recent advancements in Rook 1.16, the roadmap for future releases like 1.17, and critical considerations for maintaining and securing Ceph clusters under Rook's orchestration. For anyone seeking to implement robust, enterprise-grade storage solutions in Kubernetes, particularly those looking to leverage the power and flexibility of Ceph, this talk offers invaluable insights into Rook's capabilities and best practices.

Background

▶ Watch: Introduction to Rook and its purpose with Ceph (1:06)

The genesis of Rook stems from a fundamental challenge in the Kubernetes ecosystem: while applications scale effortlessly in cloud environments, underlying storage often lags. Many organizations default to cloud provider storage solutions, but this approach can introduce vendor lock-in, limit control, and sometimes lead to suboptimal performance or higher costs. The question arose: why couldn't storage within a data center be as scalable and flexible as the applications it serves?

Rook was designed to answer this by not reinventing the wheel but by leveraging Ceph, a battle-tested, open-source distributed storage solution that has been running in production for over a decade. Ceph offers a unified platform for block storage (RBD), shared file systems (CephFS), and S3-compatible object storage (RADOS Gateway or RGW). Its inherent scalability (up and out), thin-provisioning capabilities, and robust data protection mechanisms made it an ideal candidate for cloud-native orchestration.

Rook functions as a Kubernetes operator, meaning it uses Kubernetes' native constructs like Custom Resources (CRs) and the controller pattern to manage Ceph clusters. This cloud-native approach simplifies deployment, scaling, and lifecycle management of Ceph. The architecture consists of three layers: Rook, which deploys and manages Ceph; the CSI driver, which handles storage provisioning and mounting for Kubernetes workloads; and Ceph itself, which acts as the data layer.

Rook-orchestrated Ceph clusters demonstrate impressive scalability, with some deployments reaching multi-petabyte scales. Examples cited include clusters with 250 nodes, 1,800 OSDs (Object Storage Daemons), each utilizing 3TB NVMe drives, totaling approximately 5.2 petabytes of storage. Public Ceph telemetry numbers indicate deployments reaching exabytes of data. Performance is equally compelling, with correctly configured Ceph clusters achieving throughputs of up to 1 TB/s. Rook can be deployed anywhere Kubernetes runs, from public clouds (using EBS or persistent disks) to on-premise bare-metal data centers (leveraging SSDs and HDDs for optimal performance), and even in hybrid or multi-cloud environments for enhanced resilience and data locality. In cloud environments, Rook helps mitigate common concerns like data loss by replicating and distributing data across availability zones and offers virtually unlimited scaling beyond native PV limits.

Key Findings

▶ Watch: Rook deployment scenarios: cloud, on-prem, hybrid (4:00)

The talk highlighted several significant advancements and enduring strengths of the Rook project, underscoring its role as a mature and critical component of the cloud-native storage landscape:

  • Maturity and Stability: Rook has been a CNCF graduated project since 2020 and declared stable six years ago, with a vast community of over 400 contributors and numerous downstream deployments running in production. This demonstrates a robust and reliable platform for critical data.
  • Comprehensive CSI Driver Capabilities: The Ceph CSI driver has evolved into a highly feature-rich component, supporting advanced functionalities across RBD, CephFS, and NFS. This includes volume snapshots and group snapshots, online PVC expansion, topology-aware provisioning, and robust PVC encryption with various KMS integrations and key rotation policies.
  • Advanced Object Storage Management: Rook provides sophisticated management for Ceph's RADOS Gateway (RGW), enabling multi-frontend deployments for diverse use cases (e.g., S3-only, Swift-only, dedicated admin APIs). Crucially, it allows for granular pool placement and the definition of storage classes to optimize data redundancy and performance based on application requirements, supporting both Object Bucket Claim (OBC) and the emerging Container Object Storage Interface (COSI).
  • Resilient Maintenance and Disaster Recovery: The project places a strong emphasis on data safety and cluster availability during maintenance and in disaster scenarios. Rook dynamically manages Pod Disruption Budgets (PDBs) in a topology-aware manner to prevent downtime during node operations and facilitates rolling upgrades. A notable case study demonstrated successful data recovery even after a complete loss of the Kubernetes control plane and worker nodes, highlighting the persistence of Ceph data on disk and ongoing efforts to refine disaster recovery documentation.
  • Proactive Development and Community Engagement: Rook maintains a consistent release cadence, with minor releases approximately every four months and patch releases every two weeks or on demand. The roadmap is driven by supporting the latest Ceph and Kubernetes features, ensuring continuous integration and enhancement of the platform. The team actively seeks community involvement and contributions, providing clear pathways for new maintainers.

Technical Deep Dive

▶ Watch: New features in Rook 1.16 release (6:00)

Rook's strength lies in its deep integration with Kubernetes, transforming Ceph into a truly cloud-native storage solution.

Rook and Ceph Fundamentals

At its core, Rook is a Kubernetes operator that automates the deployment, bootstrapping, provisioning, scaling, and management of Ceph clusters. It leverages Kubernetes Custom Resources (CRs) to define the desired state of the Ceph cluster, allowing users to interact with Ceph using familiar kubectl commands. This abstraction simplifies complex Ceph operations, making it accessible to Kubernetes administrators.

Ceph itself provides three primary storage interfaces:

  • Ceph Block Device (RBD): Offers block storage devices that can be attached to VMs or containers, suitable for databases and other workloads requiring raw block access. Rook maps Kubernetes Persistent Volume Claims (PVCs) to RBD images.
  • Ceph File System (CephFS): A POSIX-compliant shared file system, ideal for scenarios requiring shared access across multiple pods or for general-purpose file storage.
  • RADOS Gateway (RGW): An object storage interface compatible with Amazon S3 and OpenStack Swift APIs, perfect for cloud-native applications storing large amounts of unstructured data.

Rook's architecture uses the CSI driver as the bridge between Kubernetes and Ceph, ensuring that applications can dynamically provision and consume storage resources.

Ceph CSI Driver: The Kubernetes Interface

The Ceph CSI driver is a single project that encapsulates three distinct drivers: one for CephFS, one for RBD, and one for NFS (which runs on top of CephFS). This architecture consists of:

  • A Controller Plugin: Runs as a highly available deployment (typically replica: 2), responsible for volume provisioning, deletion, expansion, volume snapshots, and group snapshots.
  • A Node Plugin: Runs as a DaemonSet, ensuring one instance per node. It handles the actual mounting and unmounting of PVCs to application pods.

Key features of the Ceph CSI driver include:

  • State Maintenance: Extensive use of OOMAP to maintain state, preventing garbage values upon restarts.
  • Performance: Uses Go for Ceph API calls, leveraging connection pooling for improved performance.
  • Multi-Cluster Support: A single CSI driver deployment can communicate with multiple Ceph clusters, providing flexibility for complex environments.
  • Multi-Tenancy: Supports isolation at the Ceph level using RADOS namespaces or subvolume groups.
  • Storage Capabilities:
  • Thin Provisioning for RBD.
  • RWX (ReadWriteMany) block mode for VMs.
  • RWO (ReadWriteOnce) file system for databases.
  • RWX file system for CephFS and NFS.
  • Encryption: Supports PVC encryption for both CephFS and RBD, integrated with various Key Management Systems (KMS) such as HashiCorp Vault, Azure Key Vault, IBM HPCS, and custom bring-your-own-key solutions via Kubernetes secrets.
  • Online PVC Expansion: Allows for increasing PVC size without downtime across all three drivers.
  • Snapshot and Clone: Provides volume snapshots for CephFS, RBD, and NFS, and volume group snapshots for RBD and CephFS. PVC cloning is supported for all drivers. For CephFS backups, a specific mechanism allows cloning into an ROX (Read-Only Many) mount, enabling backup tools to efficiently copy data.
  • Topology-Aware Provisioning: Enables provisioning from the nearest OSD, optimizing performance.
  • Static Provisioning and In-tree Migration: Supports migrating existing Ceph volumes to CSI and provisioning pre-existing volumes.
  • Mount Options: Offers both kernel mounts and user space mounts for CephFS and RBD.

CSI Add-on: Extending Functionality

The CSI add-on extends the core CSI driver with advanced features not typically found in standard CSI implementations. It runs as a controller with a sidecar alongside the CSI driver.

  • Space Reclamation: Addresses the issue where deleting files from a PVC doesn't immediately reclaim space in the backend. It offers a Kubernetes-native mechanism (via CRs) to run fs-trim or rbd-sparse operations.
  • Network Fence Class: A critical disaster recovery feature for RWO PVCs. It displays visible IPs on the Ceph cluster nodes, allowing users to fence off a cluster to prevent data corruption when moving applications or failing over to another node or cluster.
  • Key Rotation Policy: For encrypted PVCs, the add-on supports a cron job-like mechanism to automatically rotate encryption keys.
  • Volume Replication: Enables replication of PVCs and replication of volume groups, simplifying disaster recovery workflows with Kubernetes CRs.

CSI Operator: Future Direction

The project is moving towards a dedicated CSI operator to decouple CSI functionality and maintenance from the main Rook operator. This will allow users to deploy and manage CSI drivers with minimal configuration, using Helm charts for installation. Future roadmap items for the CSI driver include NFS shallow volume support, Ganesha authentication for NFS, expanded volume group snapshot support, chain block tracking for RBD PVCs (for more efficient backups), and QoS (Quality of Service) for RBD PVCs.

Object Storage: S3/Swift with RGW

Ceph implements object storage through the RADOS Gateway (RGW), a web server that provides S3 and Swift compatible protocols, with Ceph serving as the storage backend. Rook manages RGW deployments via the CephObjectStore custom resource.

  • Basic Deployment: Users define CephObjectStore with desired RGW instances and RBD pool parameters, and Rook creates the necessary Ceph pools and Kubernetes deployments.
  • Multi-Frontend Deployment: For advanced scenarios, Rook allows for dedicated RGW deployments with different configurations but serving the same data. This enables:
  • Separate deployments for S3, Swift, or admin APIs.
  • Dedicated instances per customer for load balancing.
  • Dedicated instances for garbage collection, allowing user-facing instances to disable it for performance.
  • CephObjectStore Configuration: The CR includes two main groups of parameters:
  • Frontend Configuration: Defines deployment aspects like replicas, resources, enabled protocols, certificates, and domains.
  • Backend Configuration: Specifies how and where data is stored in Ceph.
  • Shared Backend: Rook allows multiple CephObjectStore instances to share the same backend configuration, enabling different API frontends (e.g., S3 and Swift) to access the same underlying object data.
  • Pool Placement and Storage Classes: This is a powerful feature for optimizing object storage. Within the backend configuration, users can define multiple pool placements, each with a name and a set of Ceph pools:
  • Metadata Pool: Stores bucket index and metadata, typically configured on fast storage like SSDs with three replicas.
  • Data Pool: Stores object payloads, often on HDDs with three replicas.
  • Storage Classes: Users can define custom storage classes (e.g., "reduced redundancy") that override the default data pool configuration, allowing objects to be stored with fewer replicas (e.g., a single replica on HDD) for cost savings on less critical data. S3 clients can reference these placements as "S3 regions."
  • Object Storage Provisioning:
  • Object Bucket Claim (OBC): This is the most popular method, mirroring the PVC pattern. Users create an OBC, and Rook provisions a new bucket and credentials in Ceph, providing connection information via a Kubernetes secret and config map.
  • Container Object Storage Interface (COSI): Rook also implements a COSI driver, supporting this emerging open specification for provisioning object storage. COSI defines admin CRs (BucketClass, BucketAccessClass) for quality of service and permissions, and user CRs to request buckets and access.

Demo / Proof of Concept

▶ Watch: CSI driver overview and architecture (7:00)

While the talk did not feature a live demo in the traditional sense, a critical case study was presented that served as a powerful proof of concept for Rook's resilience and data recovery capabilities.

The scenario involved a user who, during routine maintenance, accidentally re-imaged all control plane and worker nodes, resulting in a complete loss of their Kubernetes cluster, including all etcd metadata. Despite this catastrophic event, the underlying Ceph data persisted on the disks. The Rook team was able to assist the user through a series of manual operations to recover their critical data. This involved bringing up a temporary Kubernetes cluster, mounting the original disks, and then accessing the RBD or CephFS volumes to extract the data.

This case study underscored a vital defensive implication: while Kubernetes metadata can be lost, Ceph ensures data persistence on disk. The team is actively documenting these recovery steps and integrating them into the official Rook disaster recovery guide, aiming to automate and simplify such complex recovery processes in the future. The speakers also noted that a robust etcd backup strategy would have significantly streamlined the recovery, highlighting its importance.

Defensive Implications

▶ Watch: Key features and capabilities of CSI driver (8:00)

Maintaining a robust and available storage platform like Rook-orchestrated Ceph requires careful planning and execution, particularly concerning cluster maintenance and disaster recovery.

  1. Topology Planning for Resilience:
  • Crucial Configuration: The physical layout of your data center (zones, racks, hosts, OSDs per node) is paramount for data safety and availability.
  • Zone-Aware Replication: For high availability, it is recommended to configure Ceph with at least three replicas, distributing data across at least three distinct zones (e.g., availability zones in the cloud or physical racks on-prem).
  • Outage Survival:
  • Single Zone Outage: If an entire zone (e.g., Zone A) goes down, the cluster remains fully online, serving both reads and writes without downtime or data loss, as data is replicated across the remaining zones.
  • Two Zone Outage: If two out of three zones go down, the data remains safe within the surviving zone. However, the cluster will be down, preventing reads and writes. Measures can be taken to bring the cluster back online using the single remaining zone if the other two are permanently lost.
  1. Kubernetes Pod Disruption Budgets (PDBs):
  • Automated Management: Rook dynamically manages PDBs to signal to Kubernetes how many Ceph pods (e.g., OSDs, Monitors) can be taken down at any given time without jeopardizing cluster availability.
  • Topology Awareness: These PDBs are topology-aware. For example, if a host in Zone A is being drained for maintenance, Rook's PDBs might allow another host in Zone A to go down concurrently but will prevent hosts in Zone B or C from going down simultaneously, ensuring at least two zones remain operational.
  • Preventing Downtime: Adhering to PDBs during maintenance operations is critical to prevent unexpected downtime.
  1. Rolling Upgrades:
  • Zero Downtime: Rook performs rolling upgrades for its own components and the underlying Ceph services. This ensures that only one failure domain is affected at a time, minimizing disruption and preventing downtime during upgrades.
  1. Disaster Recovery (DR) Preparedness:
  • Data Persistence: As demonstrated in the case study, Ceph data is persisted directly to disk, making it recoverable even if the entire Kubernetes control plane is lost.
  • Documentation: Rook provides disaster recovery guides, which are continually being improved and expanded to cover various failure scenarios, including the recovery of data from disks after a complete Kubernetes cluster loss.
  • etcd Backups: While not directly a Rook feature, the talk implicitly highlighted the critical importance of backing up Kubernetes etcd for faster and simpler cluster recovery in many disaster scenarios.
  1. Specialized Maintenance Tooling:
  • kubectl rook Plugin: Rook offers a kubectl plugin, a command-line interface tool for performing one-off, advanced maintenance tasks that don't fit the declarative CRD model. Examples include:
  • Restoring Monitor quorum in a degraded cluster.
  • Manually removing OSDs.
  • Other advanced operations for Ceph Monitors and OSDs.
  • Automation: The project actively solicits community feedback on what other one-off maintenance tasks could be automated into this plugin.

Key Takeaways

  • Rook transforms Ceph into a cloud-native, enterprise-grade storage solution for Kubernetes, offering scalable block, file, and S3-compatible object storage.
  • The Ceph CSI driver is highly mature and feature-rich, providing advanced capabilities like volume and group snapshots, online PVC expansion, PVC encryption with KMS integration, and topology-aware provisioning.
  • Rook's object storage management is flexible, supporting multi-frontend RGW deployments, granular pool placement with storage classes, and integration with both Object Bucket Claim (OBC) and Container Object Storage Interface (COSI).
  • Robust maintenance and disaster recovery strategies are crucial, with Rook providing tools like dynamic, topology-aware PDBs, rolling upgrades, and a kubectl rook plugin for specialized operations, alongside ongoing efforts to improve DR documentation.
  • Rook boasts a strong, active community and a clear roadmap focused on supporting the latest Ceph and Kubernetes features, ensuring its continued relevance and capability in the evolving cloud-native landscape.
  • While distributed storage introduces some network latency compared to local disks, Ceph's architecture allows for massive scalability and high aggregate throughput, especially in larger clusters with many clients.

About the Speaker(s)

The talk was delivered by a panel of experienced contributors to the Rook and Ceph ecosystems:

  • Deepika Upadhyay: A developer from Clyiso, Deepika has been actively working with Rook for almost five years, contributing to its ongoing development.
  • Madhu Rajanna: An engineer at IBM, Madhu is a maintainer of the Ceph CSI driver, the Ceph CSI operator, and COSI, playing a key role in integrating CSI functionalities with Rook.
  • Artem Torubarov: Representing Clyo, Artem has focused on the object storage aspects of Rook for the past year, contributing to the advanced RGW features.
  • Travis Nielsen: One of the original creators of the Rook project, Travis provides foundational insights into the project's vision and evolution.
  • Sebastien Han: Also listed as a speaker, Sebastien is a contributor to the project, though his specific role or company was not detailed in the transcript.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This talk is a masterclass in cloud-native storage, directly from the architects of Rook and Ceph. It’s an incredibly deep dive into how to leverage battle-tested distributed storage like Ceph within Kubernetes, covering everything from fundamental architecture to advanced CSI features, object storage management, and critical disaster recovery strategies. The speakers, all core contributors, demonstrate an unparalleled understanding of the subject, making this essential viewing for any serious practitioner building or securing cloud-native infrastructure.

Heather Calloway (CISO) — STRONG ACCEPT

This session on Rook and Ceph for Kubernetes storage provides critical insights into building resilient, scalable, and secure data foundations in cloud-native environments. While deeply technical, the presentation effectively translates complex distributed storage concepts into operational realities, offering clear guidance on disaster recovery, data protection, and high availability. Its focus on practical implementation, reinforced by a compelling disaster recovery case study, makes it highly valuable for leaders responsible for enterprise data integrity and business continuity.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025