Kubernetes SIG Storage: Intro & Deep Dive - Xing Yang, VMware by Broadcom & Jan Šafránek, Red Hat
Xing Yang, VMware by Broadcom, Jan Šafránek, Red Hat
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
The Kubernetes SIG Storage: Intro & Deep Dive session at KubeCon EU provided a comprehensive update on the state of storage within Kubernetes, delivered by two of its key leaders: Xing Yang from VMware by Broadcom, a co-chair of SIG Storage, and Jan Šafránek from Red Hat, a tech lead for the group. This talk served as an essential resource for both new contributors and seasoned Kubernetes users, detailing the SIG's responsibilities, recent achievements, and an ambitious roadmap for future development. The session underscored SIG Storage's critical role in shaping how applications within Kubernetes interact with persistent and ephemeral data, ensuring reliability, scalability, and security.

Key moments
- 0:00 Introduction to SIG Storage and its responsibilities
- 2:16 SIG Storage authors CSI; designing COSI for object storage
- 2:50 Understanding Persistent Volume Claims, Volumes, and Storage Classes
- 4:41 Exploring various types of ephemeral volumes in Kubernetes
- 6:26 Current status of in-tree volume plugins and CSI recommendation
- 7:44 Automatic PVC removal for StatefulSets and resize recovery
Kubernetes SIG Storage: Intro & Deep Dive
Speakers: Xing Yang, VMware by Broadcom; Jan Šafránek, Red Hat
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=X_xHC_Q5jGE
Overview
The Kubernetes SIG Storage: Intro & Deep Dive session at KubeCon EU provided a comprehensive update on the state of storage within Kubernetes, delivered by two of its key leaders: Xing Yang from VMware by Broadcom, a co-chair of SIG Storage, and Jan Šafránek from Red Hat, a tech lead for the group. This talk served as an essential resource for both new contributors and seasoned Kubernetes users, detailing the SIG's responsibilities, recent achievements, and an ambitious roadmap for future development. The session underscored SIG Storage's critical role in shaping how applications within Kubernetes interact with persistent and ephemeral data, ensuring reliability, scalability, and security.
The presentation highlighted the continuous evolution of Kubernetes storage APIs, the ongoing transition from in-tree volume plugins to the more flexible Container Storage Interface (CSI), and the introduction of groundbreaking features designed to address complex enterprise storage requirements. From enhanced data protection mechanisms like volume group snapshots and change block tracking to crucial performance and security improvements such as SELinux relabeling with mount options, the speakers meticulously unpacked the technical intricacies and practical implications of these advancements. The discussion also extended to emerging paradigms like the Container Object Storage Interface (COSI), signifying a strategic move towards unifying object storage management within the Kubernetes ecosystem.
This article delves into the technical specifics covered in the KubeCon session, offering an in-depth analysis of SIG Storage's contributions. It aims to provide a clear understanding of the challenges being tackled, the solutions implemented, and the future direction of storage management in Kubernetes. For anyone operating or developing applications on Kubernetes, grasping these developments is crucial for optimizing storage utilization, securing data, and preparing for the next generation of cloud-native storage solutions.
Background
▶ Watch: Introduction to SIG Storage and its responsibilities (0:00)
Kubernetes, by design, treats pods as ephemeral entities, meaning they can be created, destroyed, or rescheduled at any time. This ephemeral nature poses a significant challenge for applications that require persistent storage for their data. To address this, Kubernetes introduced a robust set of storage APIs, primarily Persistent Volume Claims (PVCs) and Persistent Volumes (PVs). A PVC represents a user's request for storage, specifying desired capacity, access modes, and optionally a StorageClass. A PV, on the other hand, is the actual piece of storage provisioned to fulfill that request, often pointing to a specific backend like a cloud disk or an NFS share. This abstraction allows developers to request storage without needing to know the underlying infrastructure details, while administrators can provision storage resources independently.
The concept of dynamic provisioning is central to this model. When a PVC is created and no suitable existing PV is found, Kubernetes can automatically provision a new PV on the backend based on the parameters defined in the StorageClass. This eliminates the need for manual PV creation, streamlining operations. Beyond persistent storage, Kubernetes also offers ephemeral volumes for temporary data needs, such as scratch space via emptyDir or data injection through secret and configMap volumes. Generic ephemeral volumes, backed by PVCs, allow for temporary storage on external backends, with data deleted upon pod termination.
Historically, Kubernetes included many in-tree volume plugins directly within the kubelet and kube-controller-manager codebases for various storage systems like AWS EBS, GCE Persistent Disks, and Ceph RBD, alongside generic ones like NFS and iSCSI. While functional, this approach led to tight coupling, requiring Kubernetes core releases for updates to storage drivers and limiting the extensibility for new storage technologies. This challenge spurred the creation of the Container Storage Interface (CSI) specification. CSI provides a standardized, vendor-agnostic interface for container orchestrators (like Kubernetes) to interact with arbitrary storage systems. The SIG Storage has been instrumental in authoring and maintaining the CSI specification and its Kubernetes implementation, facilitating a migration path for all in-tree plugins to become external CSI drivers. This not only decouples storage logic from Kubernetes core but also empowers storage vendors to develop and release drivers independently, accelerating innovation and improving maintainability.
The ongoing CSI migration is a testament to this strategic shift. The goal is to route all volume operations, previously handled by in-tree plugins, through CSI drivers, eventually allowing for the removal of legacy in-tree code. This transition, while complex, promises a more flexible, secure, and scalable storage landscape for Kubernetes users.
Key Findings
▶ Watch: Understanding Persistent Volume Claims, Volumes, and Storage Classes (2:50)
The SIG Storage session unveiled a multitude of advancements, with several critical features graduating to stable (GA) or beta status, alongside exciting new alpha capabilities. These findings collectively aim to enhance data persistence, improve operational efficiency, and expand the types of storage accessible within Kubernetes.
A significant achievement is the automatic removal of Persistent Volume Claims (PVCs) from StatefulSets, which graduated to GA in Kubernetes 1.28. This feature addresses a long-standing operational challenge where, by default, deleting a StatefulSet or scaling it down would leave associated PVCs and PVs orphaned. While this protected against accidental data loss, it often led to storage leaks. Now, with careful configuration, users can opt for automatic PVC deletion, although the speakers cautioned about the potential for irreversible data loss if not used judiciously.
Also reaching beta in 1.28 is recovery from resize failure, a crucial improvement for volume expansion. Previously, if a user attempted to expand a volume but the storage backend failed (e.g., due to a typo resulting in an excessively large request like "5 exabytes" instead of "5 terabytes"), the volume would enter a stuck state, as shrinking was not supported. This feature now allows users to correct such errors and shrink the volume back to a valid size if the backend confirms no expansion occurred, making volume resizing more robust. The progress of expansion is now also visible in the PVC status.
Volume group snapshots also progressed to beta in 1.28, introducing new API objects (VolumeGroupSnapshot, VolumeGroupSnapshotContent, VolumeGroupSnapshotClass) and CSI core requirements. This enables consistent snapshots of multiple persistent volumes simultaneously, a vital capability for backing up entire applications or databases within a StatefulSet, ensuring data integrity across related volumes.
For Kubernetes 1.33, several features are targeting GA:
- Warning Populators: This feature greatly enhances PVC creation by allowing volumes to be populated from generic data sources beyond just snapshots or other PVCs. The introduction of the
dataSourceReffield supports external data sources and enables cross-namespace volume transfers, opening up possibilities for seamless backup and restore workflows from object stores. - Always Honor PV Reclaim Policy: This critical update ensures that the reclaim policy (e.g.,
Delete,Retain) configured for a PV is consistently honored, regardless of whether the PVC or PV is deleted first. This prevents storage resource leaks and ensures users are not charged for resources they intended to discard. - Portworx CSI migration: This specific CSI migration is targeting GA, signifying the maturity of Portworx's integration with Kubernetes via CSI.
- SELinux relabeling with mount options for non-rewrite-once pod volumes is targeting beta in 1.33. This feature aims to dramatically improve pod startup times for SELinux-enabled clusters by applying labels at mount time, rather than traversing the entire filesystem.
New alpha features targeting 1.33 include:
- CBT (Change Block Tracking): This provides a mechanism to retrieve metadata about allocated and changed blocks in a snapshot, enabling highly efficient incremental backups.
- Mutable CSI Node Allocatable Property: This aims to resolve scheduling failures caused by mismatches between reported and actual volume attachment capacities on a node, allowing for dynamic updates to this property.
- Storage Capacity Scoring: This enhances the kube-scheduler's logic for dynamic provisioning, allowing administrators to configure preferences for nodes with the least or most allocatable capacity, optimizing resource distribution.
Finally, a significant removal was announced: the git repo in-tree volume plugin. Deprecated for a long time and unmaintained, it posed security risks, including potential remote code execution as root. It will error out in 1.33, be locked in 1.36, and completely removed in 1.39, with git-sync or initContainers as recommended alternatives.
Technical Deep Dive
▶ Watch: Exploring various types of ephemeral volumes in Kubernetes (4:41)
The technical advancements discussed in the SIG Storage session represent significant strides in making Kubernetes storage more robust, efficient, and secure. Several features warrant a deeper exploration due to their architectural changes and impact on cluster operations.
SELinux Relabeling with Mount Options
The traditional method of applying SELinux labels to volumes in Kubernetes involved the container runtime recursively traversing all files on a volume when a pod started. While this ensured proper security context, it could lead to significant delays, particularly on volumes with millions of files or slow storage backends, sometimes taking minutes or even hours. To combat this inefficiency, SIG Storage is implementing SELinux relabeling using mount options. This approach applies the SELinux label during the mount operation, in constant time, regardless of the volume size.
However, this optimization introduces a crucial breaking change. Historically, it was possible for privileged and unprivileged pods to share the same volume, as the recursive relabeling would ensure all files had the correct, uniform SELinux label. With the mount option approach, all pods accessing the same volume simultaneously must present the same SELinux label, extending this requirement even to privileged pods. If a privileged pod and an unprivileged pod attempt to share a volume with different SELinux contexts, the mount will fail.
To facilitate a smooth transition, a multi-stage rollout with feature gates has been devised:
SELinuxMountReadWritePod: This feature gate, already GA and enabled by default, applies the mount option relabeling for volumes that are not sharable (i.e.,ReadWriteOnceaccess mode). Since only one pod can use such a volume, there's no conflict.SELinuxChangePolicy(Alpha in 1.32): This introduces a new controller that is opt-in. When enabled, it observes the cluster and reports metrics and events for potential conflicts or breaking scenarios (e.g., privileged/unprivileged pod sharing a volume). This allows administrators to identify and remediate issues proactively, either by rearchitecting applications or using an opt-out mechanism in the pod's security context. This feature does not break anything directly but provides crucial telemetry.SELinuxMount(Final GA): Once clusters are deemed stable based onSELinuxChangePolicyobservations, this final feature gate can be enabled, or users can upgrade to a Kubernetes version where it is GA. This fully enables the mount option relabeling for all shared volumes, including those accessed by privileged pods. Users not employing SELinux are unaffected by these changes.
Warning Populators (GA in 1.33)
The Warning Populators feature significantly enhances the flexibility of PVC creation by allowing volumes to be initialized from a much broader range of data sources. Previously, a PVC could only specify another PVC or a volume snapshot as its dataSource. This limitation restricted workflows, especially for backup and restore scenarios where data might reside in external systems like object stores.
To overcome this, the feature introduces a new field in the PVC spec: dataSourceRef. Unlike the original dataSource field, which is limited to local objects within the same namespace, dataSourceRef allows specifying any API object as a data source, including custom resources (CRs), and supports objects in any namespace. This design enables:
- Generic Data Sources: A
dataSourceRefcan point to a custom resource representing a backup in an object store, an image, or any other external data blob. - Cross-Namespace Transfer: The ability to reference objects in other namespaces lays the groundwork for future features supporting secure data transfer across namespace boundaries.
An example implementation called hello-populator in the shared-library-for-volumes repository demonstrates how this works. A custom resource definition (CRD) can define a Hello object, and a PVC can then reference this Hello CR via dataSourceRef to populate its volume. A volume data source validator controller ensures the validity of the specified data source, and a volume populator is responsible for registering and handling the custom data source type. This modular design empowers storage vendors and third-party developers to integrate diverse data sources seamlessly into Kubernetes volume provisioning.
Volume Group Snapshots (Beta in 1.28)
For applications like databases or message queues that rely on multiple interdependent volumes (e.g., data, logs, configuration), taking consistent snapshots individually can lead to data inconsistency. Volume Group Snapshots address this by enabling the creation of a single, atomic snapshot of a set of persistent volumes. This ensures that all volumes in the group are captured at the exact same point in time, guaranteeing data integrity for the entire application state.
This feature introduces three new API objects:
VolumeGroupSnapshot: A user-facing object representing a request to snapshot a group of volumes.VolumeGroupSnapshotContent: An administrator-facing object representing the actual group snapshot on the storage backend.VolumeGroupSnapshotClass: Similar toStorageClass, this allows administrators to define different policies or parameters for group snapshot creation, such as retention policies or performance tiers.
On the CSI side, new CSI cores (RPC calls) are required for CSI drivers to implement the logic for creating, deleting, and managing group snapshots on the underlying storage system. This capability is particularly valuable for complex stateful applications deployed via StatefulSets, allowing for reliable backup and restore operations for the entire application.
Container Object Storage Interface (COSI)
While PVs and PVCs provide block and file storage, object storage has become a cornerstone of cloud-native architectures for unstructured data, backups, and large datasets (e.g., AI/ML models). The Container Object Storage Interface (COSI) aims to bring object storage management natively into Kubernetes, mirroring the success of CSI for block and file storage.
COSI introduces concepts like Object Buckets and Object Bucket Claims (OBCs), analogous to PVs and PVCs. Users can request an OBC, specifying desired attributes (e.g., region, access policy), and COSI will dynamically provision an Object Bucket on an underlying object storage system (e.g., S3, Google Cloud Storage, Azure Blob Storage). This provides a standardized way for applications to consume object storage resources, abstracting away vendor-specific APIs and credentials. COSI is currently in alpha and is actively being developed towards its v1alpha2 release, promising a unified approach to managing all types of storage within Kubernetes.
Mutable CSI Node Allocatable Property & Storage Capacity Scoring
These two alpha features in 1.33 aim to significantly improve the reliability and efficiency of pod scheduling related to storage.
- Mutable CSI Node Allocatable Property: The
CSI Node Allocatableproperty indicates how many volumes a node can support. Mismatches between the reported and actual attachment capacity can lead to permanent scheduling failures and "stuck workloads" where kube-scheduler tries to place pods on nodes that cannot actually attach the required number of volumes. This feature proposes making this property mutable, allowing CSI drivers to periodically update the field and automatically adjusting it when attachment failures due to insufficient capacity are detected. This dynamic adjustment ensures that the scheduler has accurate information, preventing unnecessary scheduling attempts on oversubscribed nodes. - Storage Capacity Scoring: Previously, the volume binding plugin had limited scoring logic for static provisioning. This feature enhances that logic to also support dynamic provisioning. Administrators can now configure
kube-schedulerto prefer nodes with either the least allocatable capacity (to spread workloads) or the most allocatable capacity (to consolidate workloads or ensure large volumes can be provisioned). By default, nodes with the most allocatable capacity will be preferred. This replaces and consolidates an older alpha feature (VolumeCapacityPriority), providing a more refined and configurable mechanism for storage-aware scheduling.
Demo / Proof of Concept
▶ Watch: Current status of in-tree volume plugins and CSI recommendation (6:26)
The talk focused on presenting the status and future roadmap of SIG Storage features rather than live demonstrations. While the speakers referenced an example implementation of the "Warning Populators" feature called hello-populator in the shared library for volumes repository, and showed a conceptual CR definition and PVC usage, no live demo or proof of concept was conducted during the session itself. The emphasis was on technical explanation and architectural details.
Defensive Implications
▶ Watch: Automatic PVC removal for StatefulSets and resize recovery (7:44)
The advancements in Kubernetes SIG Storage carry significant implications for cluster administrators and security professionals, requiring careful attention to configuration and migration strategies.
- SELinux Relabeling with Mount Options: For clusters utilizing SELinux, this is a critical update. Administrators must proactively monitor for conflicts when
SELinuxChangePolicyis enabled in alpha. If privileged and unprivileged pods share volumes, applications may need to be rearchitected to avoid this pattern, or explicit opt-out mechanisms within pod security contexts must be applied. Failure to address these conflicts beforeSELinuxMountreaches GA could lead to pod startup failures. Reviewing Kubernetes vendor documentation regarding SELinux integration is paramount.
- Automatic PVC Removal from StatefulSets: While beneficial for preventing storage leaks, the GA of this feature in 1.28 demands extreme caution. Enabling automatic PVC deletion means that scaling down a StatefulSet or deleting it will permanently remove the associated data. Administrators must ensure robust backup strategies are in place and that users fully understand the implications before enabling this feature, as accidental data loss is a significant risk.
- Always Honor PV Reclaim Policy: The GA of this feature in 1.33 resolves a long-standing bug that could lead to orphaned PVs. This change is generally positive, as it ensures storage resources are consistently reclaimed according to their defined policy. Defenders should verify that their PV reclaim policies (
DeleteorRetain) accurately reflect their desired behavior for data lifecycle management, as the system will now strictly enforce them.
- Git Repo In-Tree Volume Plugin Removal: The deprecation and impending removal of the
gitRepovolume plugin due to security concerns (potential RCE as root) is a critical defensive measure. Any applications still relying on this plugin must migrate to recommended alternatives likegit-syncsidecar containers orinitContainersas soon as possible. Administrators should scan their clusters forgitRepovolume usage and plan for its removal, which will cause errors in 1.33 and full removal by 1.39.
- CSI Migration: The ongoing CSI migration is nearing completion, with specific in-tree drivers like Portworx reaching GA for their CSI counterparts. Defenders should ensure that the appropriate CSI drivers are installed and configured for their storage backends. Legacy in-tree plugins will eventually be removed, making CSI the sole supported interface for most storage types. This transition enhances security by decoupling storage logic from the Kubernetes core, allowing for faster security patches and more granular control over storage operations.
- Warning Populators: The flexibility offered by generic data sources via
dataSourceRefin 1.33 opens new avenues for data ingestion. Defenders should be aware of the potential for new attack vectors if external data sources are not properly secured or validated. Implementing robust admission controllers and access policies fordataSourceRefobjects will be crucial to prevent unauthorized data injection or manipulation.
- Change Block Tracking (CBT): The introduction of CBT in alpha for 1.33 provides a foundation for highly efficient backup solutions. Defenders should explore how their backup strategies can leverage this feature for faster, more resource-efficient incremental backups, thereby improving recovery time objectives (RTO) and recovery point objectives (RPO).
By proactively addressing these defensive implications, cluster operators can leverage the new storage features to enhance the security, reliability, and efficiency of their Kubernetes environments.
Key Takeaways
- CSI is the Future of Kubernetes Storage: The ongoing CSI migration is nearing completion, signifying a full pivot from in-tree volume plugins to external CSI drivers. Users must ensure their storage backends have robust CSI drivers installed and configured, as legacy in-tree options are being phased out.
- Enhanced Data Protection and Lifecycle Management: Features like automatic PVC removal from StatefulSets (GA 1.28) and the "Always Honor PV Reclaim Policy" (GA 1.33) provide finer-grained control over storage resource lifecycles, preventing leaks but also requiring careful configuration to avoid accidental data loss.
- Improved Backup and Restore Capabilities: Volume Group Snapshots (Beta 1.28) enable consistent, atomic snapshots of multiple related volumes, while Change Block Tracking (Alpha 1.33) promises highly efficient incremental backups, crucial for disaster recovery and data integrity.
- Flexible Volume Provisioning: "Warning Populators" (GA 1.33) revolutionize PVC creation by allowing volumes to be initialized from generic, cross-namespace data sources, opening up new possibilities for backup, restore, and data injection workflows.
- Critical Security and Performance Updates: SELinux relabeling with mount options aims to drastically improve pod startup times in SELinux-enabled clusters, though it introduces a breaking change for shared volumes with mixed privilege pods. The removal of the insecure
gitRepoin-tree volume plugin underscores the commitment to security. - Smarter Storage-Aware Scheduling: Mutable CSI Node Allocatable and Storage Capacity Scoring (both Alpha 1.33) will lead to more intelligent and reliable pod scheduling by providing kube-scheduler with accurate, dynamic information about node storage capacity and allowing for configurable preferences.
About the Speaker(s)
Xing Yang is a co-chair of Kubernetes SIG Storage and is affiliated with VMware by Broadcom. Her leadership in SIG Storage plays a crucial role in steering the development and evolution of storage capabilities within the Kubernetes ecosystem.
Jan Šafránek serves as a tech lead for Kubernetes SIG Storage and is associated with Red Hat. His technical expertise is instrumental in guiding the implementation and refinement of storage features, contributing significantly to the robustness and functionality of Kubernetes storage.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This session delivered a comprehensive and technically rigorous update directly from the Kubernetes SIG Storage leadership. It went deep into critical advancements in persistent storage, data protection, and volume management, covering features like SELinux relabeling, generic volume populators, and group snapshots. The talk provided essential, actionable intelligence for anyone operating or securing Kubernetes clusters, highlighting both new capabilities and crucial breaking changes.
Heather Calloway (CISO) — STRONG ACCEPT
This SIG Storage update provides a critical overview of Kubernetes storage advancements, translating complex technical changes into clear operational and security implications. It highlights significant governance considerations, such as the inherent data loss risk with automatic PVC removal and the urgent need to address the insecure gitRepo plugin. While deeply technical, the session effectively outlines how these features impact business resilience, data integrity, and defender operations, offering actionable insights for CISO and platform leadership to manage risk and optimize their Kubernetes environments.