Kubernetes Data Protection WG Deep Dive - Dave Smith-Uchida, Veeam
Dave Smith-Uchida, Veeam
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
The "Kubernetes Data Protection WG Deep Dive" session at KubeCon EU provided a comprehensive update on the ongoing efforts to mature data protection capabilities within the Kubernetes ecosystem. Presented by Shinyang, a co-chair of both Kubernetes SIG Storage and the Data Protection Working Group (WG), and Dave Smith-Uchida from Veeam, the talk highlighted significant advancements and future directions for safeguarding stateful workloads in cloud-native environments. The session underscored the critical need for robust data protection mechanisms as more mission-critical applications migrate to Kubernetes, leveraging its inherent self-healing, scalability, and portability.

Key moments
- 0:00 Introduction and motivation for Data Protection WG
- 2:25 Detailed Kubernetes application backup workflow explanation
- 4:15 Upcoming Kubernetes features for enhanced data protection
- 5:45 Understanding the Kubernetes application restore workflow
- 6:40 New white paper on Kubernetes data protection best practices
- 8:00 Understanding RTO, RPO, and data protection strategy impact
Kubernetes Data Protection WG Deep Dive
Speakers: Dave Smith-Uchida, Veeam; Shinyang, VMware/BRCON
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=joOTwCatd9g
Overview
The "Kubernetes Data Protection WG Deep Dive" session at KubeCon EU provided a comprehensive update on the ongoing efforts to mature data protection capabilities within the Kubernetes ecosystem. Presented by Shinyang, a co-chair of both Kubernetes SIG Storage and the Data Protection Working Group (WG), and Dave Smith-Uchida from Veeam, the talk highlighted significant advancements and future directions for safeguarding stateful workloads in cloud-native environments. The session underscored the critical need for robust data protection mechanisms as more mission-critical applications migrate to Kubernetes, leveraging its inherent self-healing, scalability, and portability.
While Kubernetes excels at Day-1 operations—such as declarative deployment and management of persistent storage—Day-2 operations, particularly data protection like backup, restore, and disaster recovery, have historically lagged. This gap presents a significant challenge for enterprises adopting Kubernetes for stateful applications. The Data Protection WG, sponsored by SIG Storage and SIG Apps, addresses this by developing standardized APIs, workflows, and best practices to ensure data integrity and availability. This talk detailed the progress made on several key features, including volume snapshots, change block tracking, and object storage integration, alongside an insightful preview of a new white paper on best practices for application architects and administrators.
The importance of this work cannot be overstated. As Kubernetes becomes the de facto platform for modern applications, the ability to reliably protect and restore application data is paramount for business continuity and compliance. The insights shared by the WG members illuminate the complexities involved in achieving application-consistent backups, managing operator-driven deployments, and navigating the trade-offs between recovery objectives (RTO and RPO), cost, and performance. The session served as a vital resource for anyone involved in designing, deploying, or managing stateful applications on Kubernetes, emphasizing the collaborative effort required to build a resilient cloud-native data protection future.
Background
▶ Watch: Introduction and motivation for Data Protection WG (0:00)
Kubernetes has rapidly evolved into a powerful platform for deploying and managing containerized applications, offering unparalleled benefits in terms of agility, scalability, and portability. For stateless workloads, Kubernetes' declarative nature, combined with GitOps principles, provides an effective mechanism for deployment and recovery. However, the landscape changes significantly when dealing with stateful workloads, which rely on persistent data. While Kubernetes has robust support for Day-1 operations of stateful applications through constructs like Persistent Volumes (PVs), Persistent Volume Claims (PVCs), and StatefulSets, Day-2 operations—specifically data protection—have presented considerable challenges.
The core problem lies in the fact that data critical to stateful applications, including secrets, ConfigMaps, and the actual data stored in PVs, is not typically managed or stored within a GitOps repository. This limitation means that a simple GitOps restore cannot fully recover a stateful application and its data after a disaster or accidental deletion. Recognizing this critical gap, the Kubernetes Data Protection Working Group was established, sponsored jointly by SIG Storage and SIG Apps. Its primary motivation is to develop and standardize the necessary building blocks and best practices within Kubernetes to enable comprehensive data protection for stateful applications.
Prior work by the WG includes a foundational white paper outlining the data protection workflow in Kubernetes, which identifies both existing and missing components. The workflow for backing up a Kubernetes application typically involves two main components: backing up the Kubernetes metadata (e.g., Deployments, StatefulSets, Services, ConfigMaps, Secrets, PV/PVC definitions) and backing up the data stored in Persistent Volumes. For volume data, two primary approaches exist: a native data dump (e.g., pg_dump for PostgreSQL, mysqldump for MySQL) or a controller-coordinated approach utilizing Volume Snapshots. The latter requires quiescing the application to ensure application consistency before taking the snapshot and then unquiescing it afterward. All backed-up data and metadata must then be exported to a backup repository. The restore process mirrors this, involving importing from the repository, restoring Kubernetes metadata, and then restoring or rehydrating PVs from snapshots or native dumps. The WG's ongoing efforts are focused on filling the architectural gaps and streamlining these complex processes.
Key Findings
▶ Watch: Upcoming Kubernetes features for enhanced data protection (4:15)
The Data Protection WG's ongoing work has yielded several critical findings and advancements, focusing on both new features and best practices for data protection in Kubernetes. These initiatives are designed to provide administrators and application developers with the tools and guidance needed to implement robust data protection strategies.
Emerging Kubernetes Features for Data Protection:
- Volume Mode Conversion (Targeting GA in 1.30): This feature aims to prevent unauthorized conversions between file system and block modes when creating a PVC from a volume snapshot. This ensures data integrity and predictable behavior during restore operations.
- Consistent Group Snapshot (Moved to Beta in 1.32): A significant advancement, this allows for creating crash-consistent snapshots of multiple volumes at the same point in time. This is crucial for applications that distribute data across several PVs, ensuring write-order consistency across all associated volumes.
- COSI (Container Object Storage Interface) (Alpha, targeting Alpha 2): COSI seeks to elevate object storage to a first-class citizen within Kubernetes. This is particularly relevant for data protection as object storage is the predominant choice for cost-effective and scalable backup repositories. COSI aims to simplify the management and consumption of object storage buckets directly from Kubernetes.
- Change Block Tracking (CBT) (Targeting Alpha in 1.33): CBT is a foundational feature for efficient incremental backups. It provides a mechanism to retrieve metadata about changed blocks between two snapshots, allowing backup solutions to transfer only the delta, significantly reducing backup windows, storage requirements, and network bandwidth usage.
- Volume Populate (Targeting GA in 1.33): This feature enhances restore flexibility by allowing the creation of a PVC from an external data source that is not necessarily a volume snapshot or another PVC. This is invaluable for restoring data from various backup formats or external systems directly into a new PVC.
New White Paper on Best Practices for Kubernetes Application Data Protection:
A new white paper is under construction, aiming to provide comprehensive guidance with two main thrusts:
- Administrator/User Guidance: How a Kubernetes administrator or user can effectively set up their applications for data protection, including understanding resiliency needs, selecting appropriate strategies, and weighing costs.
- Application Developer Guidance: What changes are necessary in existing Kubernetes applications and processes to better support data protection, particularly concerning the interaction with operators and custom resources.
Key Concepts and Strategies Defined:
- RTO (Recovery Time Objective): The maximum tolerable downtime for an application or service after a disaster.
- RPO (Recovery Point Objective): The maximum tolerable amount of data loss for an application or service after a disaster.
- Data Protection Strategies:
- GitOps (for stateless apps): Low RTO (install time), zero RPO (no data), low cost, low impact.
- Backup/Restore: Medium RTO (restore time), RPO depends on backup frequency, medium cost, variable impact (depending on consistency).
- Replication (Cold Copy): Low RTO (bring up app), low RPO (seconds to minutes for async, zero for sync), higher cost (live storage, bandwidth), potential performance impact for synchronous.
- Replication (Fully Replicated/Hot): Low RTO, low RPO, high cost, potential application changes required for distributed nature.
Consistency Levels for Backups:
- Crash Consistent: Data is as it would be after a sudden power loss. Relies on application/filesystem recovery. Low impact, commonly achieved with volume snapshots.
- Fully Consistent (Quiesced): Application and potentially OS are quiesced (paused) during backup. Guarantees full correctness but involves application downtime. High impact.
- Application Consistent: Application-level backup (e.g.,
pg_dump). Uses application's internal mechanisms to create a transactionally consistent copy. Can have low/medium impact, allowing writes during backup. - Inconsistent: Generally undesirable, as it may lead to unrecoverable or corrupted data upon restore.
Challenges with Kubernetes Operators:
- Inconsistent Backups: Operators can be actively modifying resources during a backup, leading to an inconsistent state across the backed-up Kubernetes resources and data.
- Complex Resource Relationships: The flat nature of Kubernetes resources makes it difficult to determine the complete dependency graph of an application managed by an operator (Custom Resources -> StatefulSets -> PVs).
- Race Conditions During Restore: If an operator starts before the backup application has fully restored all its managed resources and data, it can create new, empty resources, leading to conflicts and data loss.
- Lack of Standard Interfaces: There's currently no standardized way for backup solutions to quiesce or coordinate with operators during backup and restore operations.
Technical Deep Dive
▶ Watch: Understanding the Kubernetes application restore workflow (5:45)
The technical deep dive into Kubernetes data protection reveals a complex interplay of existing features, nascent projects, and architectural challenges. The core objective is to ensure that stateful applications running on Kubernetes can be reliably backed up and restored, maintaining data integrity and minimizing downtime.
The fundamental backup workflow in Kubernetes requires two distinct but coordinated operations: backing up the Kubernetes metadata and backing up the persistent volume (PV) data. Kubernetes metadata includes all the declarative configurations that define an application's state and structure—Deployments, StatefulSets, Services, ConfigMaps, Secrets, PVs, PVCs, and potentially Custom Resources (CRs) introduced by Operators. This metadata is typically extracted using Kubernetes API calls and stored.
For PV data, the process is more nuanced. One approach is a native data dump, where the application itself generates a logical backup of its data (e.g., mysqldump, pg_dump). This method ensures application consistency because the application itself controls the data extraction. The alternative, and often more efficient, is a controller-coordinated approach utilizing Volume Snapshots. This involves a sequence of operations: first, quiescing the application to temporarily halt I/O operations and flush pending writes to disk; second, taking a volume snapshot of the associated PVs; and third, unquiescing the application to resume normal operations. The snapshot captures the disk state at a specific point in time. The newly introduced Consistent Group Snapshot (Beta in 1.32) is crucial here, as it allows multiple volumes to be snapshotted simultaneously, ensuring write-order consistency across all data components of a distributed application. All collected metadata and volume data (or snapshot references) are then exported to a backup repository, which can be object storage, block storage, or other media. The Container Object Storage Interface (COSI), currently in Alpha, aims to standardize the provision and management of object storage buckets within Kubernetes, simplifying the integration of backup repositories.
The restore workflow essentially reverses this process. First, the backup data is imported from the backup repository. Then, the Kubernetes metadata is restored, recreating the application's configuration. Next, the PVCs and PVs are restored. If the original backup was a native data dump, that data is reloaded into the newly created PVs. If it was a volume snapshot, the PVCs are rehydrated from these snapshots. The Volume Populate feature (targeting GA in 1.33) adds significant flexibility by allowing PVCs to be created from arbitrary external data sources, not just existing snapshots or PVCs, which is highly beneficial for integrating with diverse backup solutions.
A critical aspect of data protection strategy revolves around defining Recovery Time Objective (RTO) and Recovery Point Objective (RPO). RTO measures the maximum acceptable downtime following a disaster, while RPO quantifies the maximum acceptable data loss. These objectives directly influence the choice of data protection strategy:
- GitOps (for truly stateless applications): RTO is the time to redeploy the application; RPO is zero as there's no stateful data. Cost and impact are minimal.
- Backup and Restore: RTO depends on the time to restore data (can be quick from snapshots, longer from full volume rebuilds); RPO depends on backup frequency (e.g., daily backup means up to 24 hours of data loss). This strategy offers a mid-range cost and impact.
- Replication (Cold Copy): Data is continuously replicated to a remote location, but compute is not active. RTO is low (time to bring up compute); RPO can be low (seconds to minutes for asynchronous replication) or zero (for synchronous replication). This strategy involves higher costs due to live storage and network bandwidth.
- Replication (Fully Replicated/Hot): Multiple active data centers with continuous replication. RTO and RPO are typically very low, but costs are high, and applications may require significant refactoring to be truly distributed.
The choice of strategy also dictates the desired consistency level of the backup:
- Crash Consistency: Achieved by volume snapshots, it guarantees that all data written to disk before the snapshot is present, similar to an abrupt power cut. Modern file systems and databases are designed to recover quickly from this state through journaling and logging. This method has a very low impact on application performance.
- Fully Consistent (Quiesced): This involves stopping or pausing the application and potentially the operating system to ensure all in-flight data is flushed to persistent storage. While offering the highest level of consistency, it incurs significant downtime, making it unsuitable for many high-availability applications.
- Application Consistency: This is the ideal state, where the backup reflects a transactionally consistent view of the application's data. It's often achieved by using application-specific tools (like
pg_dump) or by integrating with application-level hooks to quiesce specific processes without stopping the entire application. This can offer a good balance of consistency with minimal impact. - Inconsistent Backups: These are generally to be avoided as they provide no guarantee of data recoverability and can lead to corrupted or unusable data.
A significant challenge arises with Kubernetes Operators. Operators manage complex applications by extending the Kubernetes API with Custom Resources (CRs). During a backup, an operator might be actively reconciling its desired state, potentially modifying resources or data. This concurrent activity can lead to an inconsistent backup where the collected metadata and data do not reflect a single, coherent point in time. Furthermore, the flat resource model of Kubernetes makes it difficult for backup solutions to infer the relationships between an operator's CRs, the StatefulSets they manage, and the underlying PVs. This lack of a clear dependency graph complicates consistent backup and restore.
During restore, race conditions are a major concern. If an operator is active when its CRs are restored, it might immediately try to create new PVs or StatefulSets before the backup solution has had a chance to restore the original data. This leads to a "fight" between the operator and the backup application, often resulting in an empty database or an inconsistent state. The WG emphasizes the need for standard interfaces that allow backup solutions to communicate with and control operators—for example, to quiesce them during backup or to coordinate the restore order, ensuring that data is rehydrated before the operator begins its reconciliation loops. These interfaces would allow backup applications to signal to operators that a restore is in progress, enabling operators to temporarily suspend their operations or enter a "restore mode."
Finally, while replication strategies offer low RTO and RPO, they do not eliminate the need for traditional backups. Replication protects against infrastructure failures but is vulnerable to data corruption scenarios like application bugs writing garbage data, accidental deletions, or, most critically, ransomware attacks. In such cases, the corrupted data is faithfully replicated, rendering the replicas useless. Immutable, point-in-time backups provide an essential last line of defense, allowing recovery to a known good state prior to the corruption event.
Demo / Proof of Concept
▶ Watch: New white paper on Kubernetes data protection best practices (6:40)
The "Kubernetes Data Protection WG Deep Dive" session was primarily an informational update on the working group's progress, ongoing initiatives, and a preview of an upcoming white paper. As such, the presentation did not include a live demonstration or a detailed proof of concept of any specific data protection solution or feature. The focus was on architectural concepts, feature development status, and strategic considerations rather than practical implementation examples.
Defensive Implications
▶ Watch: Understanding RTO, RPO, and data protection strategy impact (8:00)
For organizations operating stateful workloads on Kubernetes, understanding and implementing robust data protection strategies is paramount. The work of the Kubernetes Data Protection WG provides critical insights and tools for defenders to bolster their resilience against data loss, corruption, and disaster.
- Understand and Define RTO/RPO: The first defensive step is to clearly define the Recovery Time Objective (RTO) and Recovery Point Objective (RPO) for each application. These metrics are not one-size-fits-all; they depend on business criticality and the acceptable cost of downtime and data loss. This definition will guide the selection of appropriate data protection strategies and technologies. For instance, a mission-critical application with near-zero RTO/RPO will require a more sophisticated (and costly) replication strategy combined with frequent backups, potentially leveraging Change Block Tracking (CBT) for efficiency.
- Choose the Right Strategy Mix: Defenders should not rely on a single data protection strategy. While GitOps is excellent for stateless application recovery, it's insufficient for stateful data. Backup and restore is a fundamental requirement, providing point-in-time recovery. Replication offers high availability and disaster recovery, but it must be complemented by immutable backups. This multi-layered approach protects against various failure modes, including hardware failures, human error, application bugs, and ransomware.
- Prioritize Immutable Backups: The talk explicitly highlighted that replication alone is insufficient. Immutable backups stored in secure, air-gapped or logically isolated repositories are essential for protection against logical data corruption (e.g., application bugs, accidental deletion) and, crucially, ransomware. Defenders must ensure their backup solutions create immutable copies that cannot be altered or deleted, even by malicious actors who gain control of the production environment.
- Demand Application-Consistent Backups: Where possible, defenders should aim for application-consistent backups. While crash-consistent snapshots are low-impact and often sufficient for many applications designed for fast recovery, truly application-consistent backups (e.g., using
pg_dumpor application-specific quiescing hooks) provide the highest assurance of data integrity upon restore. This might require collaboration with application developers to build in the necessary hooks. The Consistent Group Snapshot feature (Beta in 1.32) is a step towards achieving consistency across multiple volumes. - Address Operator Challenges: Kubernetes Operators introduce significant complexities for data protection. Defenders should engage with application teams to understand how their operators manage state and resources. Ideally, operators should be designed with data protection in mind, exposing standard interfaces for quiescing during backup and coordinating during restore. Until such standards are widespread, backup solutions must be carefully chosen to handle operator-managed applications, potentially requiring custom pre/post-backup scripts or a "quiesce" mode for the operator itself. During restore, careful orchestration is needed to prevent race conditions where operators might create empty resources prematurely.
- Leverage Emerging Kubernetes Features: Defenders should stay informed about and plan to adopt new Kubernetes features as they mature. COSI will simplify managing object storage for backup repositories. CBT will enable more frequent, efficient incremental backups, improving RPO while reducing resource consumption. Volume Populate will offer greater flexibility in restoring data from diverse sources. Integrating these features will enhance the efficiency, reliability, and breadth of data protection capabilities.
- Participate in the Community: The Data Protection WG is actively seeking community involvement. Defenders and architects can contribute to the white paper on best practices, provide feedback on new features, and help shape future standards. Active participation ensures that the evolving data protection landscape in Kubernetes meets real-world enterprise requirements.
By proactively addressing these implications, organizations can build a resilient Kubernetes environment where stateful applications are adequately protected, ensuring business continuity and data integrity even in the face of adverse events.
Key Takeaways
- Day-2 operations for stateful workloads in Kubernetes are maturing rapidly: While Day-1 operations are well-supported, the Kubernetes Data Protection Working Group is actively developing crucial building blocks and best practices to enable robust backup, restore, and disaster recovery for stateful applications, which are increasingly moving to Kubernetes.
- New Kubernetes features are enhancing data protection capabilities: Upcoming features like Consistent Group Snapshot (Beta in 1.32), Change Block Tracking (CBT) (Alpha in 1.33), COSI (Container Object Storage Interface) (Alpha), and Volume Populate (GA in 1.33) are foundational for more efficient, consistent, and flexible data protection solutions.
- Strategic data protection requires careful consideration of RTO, RPO, cost, and impact: Organizations must define their Recovery Time Objective (RTO) and Recovery Point Objective (RPO) to select appropriate strategies (GitOps, Backup/Restore, Replication) that balance business needs with implementation costs and performance impact on applications.
- Understanding consistency levels is crucial for reliable backups: Different backup consistency levels—crash consistent, fully consistent (quiesced), and application consistent—offer trade-offs between recovery speed, potential data loss, and application availability. Application-aware consistency is often the ideal for stateful applications.
- Kubernetes Operators pose unique challenges for data protection: The dynamic and self-reconciling nature of Operators can lead to inconsistencies during backup and race conditions during restore. Standardized interfaces are needed to allow backup solutions to coordinate with Operators for orderly quiescing and restoration.
- Replication is not a substitute for immutable backups: While replication provides high availability and disaster recovery, it does not protect against logical data corruption (e.g., application bugs, accidental deletion, ransomware). Immutable, point-in-time backups remain essential as a last line of defense against such threats.
About the Speaker(s)
Shinyang is a prominent figure in the Kubernetes community, serving as a co-chair of both the Kubernetes SIG Storage and the Data Protection Working Group. Shinyang works at VMware/BRCON, bringing expertise in storage and data management within cloud-native environments to shape the future of data protection in Kubernetes.
Dave Smith-Uchida works at Veeam, a leading provider of backup, recovery, and data management solutions. His involvement in the Kubernetes Data Protection Working Group leverages his extensive industry experience in data protection, including prior work at VMware, to contribute to the development of robust and practical data protection standards for Kubernetes.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This session is a critical deep dive into the very fabric of Kubernetes data protection, presented by the architects themselves. It's not just a status update; it's a foundational look at the emerging capabilities, architectural challenges, and best practices that will define how stateful applications are secured and recovered in cloud-native environments. For anyone serious about running mission-critical workloads on Kubernetes, this is essential viewing. It cuts through the marketing fluff and delivers concrete, actionable insights directly from the source.
Heather Calloway (CISO) — STRONG ACCEPT
This session provides a critical update on the foundational work being done to mature data protection within Kubernetes. It moves beyond theoretical discussions to outline concrete features and a much-needed best practices white paper. For any CISO or executive grappling with the resilience of stateful applications on cloud-native platforms, this offers a clear view of the evolving landscape, highlighting both progress and persistent challenges like operator coordination. It’s not just technical progress; it’s an essential step toward institutional accountability for data integrity in modern architectures.