Stateful Superpowers: Explore High Performa... Alex Chircop, Chris Milsted & Alex Reid, Lori Lorusso
Alex Chircop, Chris Milsted, Alex Reid, Lori Lorusso
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In this insightful KubeCon EU talk, "Stateful Superpowers: Explore High Performance Cloud Native Storage," Alex Chircop, Lori Lorusso, Alex Reid, and Chris Milsted challenged the common misconception of entirely stateless architectures, asserting that all applications eventually store state. The presentation systematically dismantled the notion that stateful workloads are incompatible with cloud-native principles, instead demonstrating how Kubernetes can be leveraged to manage complex, high-performance stateful applications with unprecedented efficiency and resilience.

Key moments
- 0:30 Why cloud native storage? No stateless architecture.
- 1:10 Automation, scaling, self-healing: cloud-native storage benefits.
- 2:50 Storage types, CSI/COSI, operators: the cloud-native ecosystem.
- 6:00 Graduated and incubating CNCF storage projects.
- 6:50 Chris introduces Cloud Native PG demo.
- 7:20 Oracle grid infrastructure: traditional database complexity.
Stateful Superpowers: Explore High Performance Cloud Native Storage
Speakers: Alex Chircop, Chief Architect, Akamai Cloud; Chris Milsted, Product Architect, Akamai; Alex Reid, Principal Engineer, Akamai; Lori Lorusso, Head of Community, Percona
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=JtMYdR50-KU
Overview
In this insightful KubeCon EU talk, "Stateful Superpowers: Explore High Performance Cloud Native Storage," Alex Chircop, Lori Lorusso, Alex Reid, and Chris Milsted challenged the common misconception of entirely stateless architectures, asserting that all applications eventually store state. The presentation systematically dismantled the notion that stateful workloads are incompatible with cloud-native principles, instead demonstrating how Kubernetes can be leveraged to manage complex, high-performance stateful applications with unprecedented efficiency and resilience.
The speakers, representing Akamai Cloud and Percona, highlighted the significant benefits of adopting cloud-native approaches for storage, including declarative automation, intelligent autoscaling, self-healing capabilities, and deterministic performance. Through compelling live demonstrations of CloudNativePG and TiKV, coupled with real-world case studies from Nokia and Civo, the talk provided a comprehensive guide for architects and engineers looking to harness the full potential of stateful workloads within a Kubernetes environment.
This talk is crucial for anyone involved in designing, deploying, or managing applications in the cloud, particularly those dealing with databases, key-value stores, or other persistent data services. It offers a clear pathway to overcome traditional complexities associated with stateful applications, transforming them from operational burdens into highly automated, scalable, and resilient components of a modern cloud-native infrastructure.
Background
▶ Watch: Why cloud native storage? No stateless architecture. (0:30)
The premise of this talk begins with a fundamental re-evaluation of application architecture: the idea that "there is no such thing as a stateless architecture; it's just someone else's problem." This statement underscores that all applications, at some point, interact with persistent storage, making the management of state a universal challenge. Historically, deploying and managing stateful applications, especially databases like Oracle, involved extensive manual configuration, often guided by voluminous documentation such as 290-page PDFs or "Red Books." These traditional methods were complex, prone to human error, and difficult to scale or recover reliably.
The advent of Kubernetes, with its declarative nature, self-healing properties, and autoscaling capabilities, revolutionized stateless workload management. The core argument of this presentation is that these same benefits can and should be extended to stateful workloads. Cloud-native storage transcends simple block volumes or file systems; it encompasses object stores, databases, and key-value pairs, integrating them seamlessly into the Kubernetes ecosystem. This integration is facilitated by critical components like CSI drivers for block storage, COSI drivers for object storage, and most importantly, operators.
Operators are pivotal to the cloud-native stateful paradigm. They encapsulate the operational knowledge of human experts into programmatic applications, automating day-2 operations such as upgrades, backups, failovers, and disaster recovery. By leveraging the broader cloud-native stack—including observability, secrets management, certificate management, encryption, and elasticity—organizations can achieve a holistic, automated, and secure environment for their stateful applications. The CNCF Storage Technical Advisory Group (TAG) has even published a white paper outlining the attributes of cloud-native storage systems, emphasizing high availability, consistency, durability, and various scaling metrics (IOPS, throughput). The talk also briefly touched upon the multi-layered virtualization common in cloud storage and highlighted several impactful CNCF storage projects, including Rook/Ceph, etcd, TiKV, and CubeFS, showcasing the vibrant innovation in this domain.
Key Findings
▶ Watch: Storage types, CSI/COSI, operators: the cloud-native ecosystem. (2:50)
The talk presented several key findings that challenge conventional wisdom regarding stateful workloads in cloud-native environments:
- Operators as Knowledge Encoders: Kubernetes operators effectively distill vast amounts of operational expertise (e.g., a 290-page database installation manual) into concise, declarative YAML configurations. This programmatic approach drastically simplifies the deployment, management, and day-2 operations of complex stateful applications, making them highly repeatable and reliable.
- Enhanced Resilience via Kubernetes Constructs: Running stateful workloads on Kubernetes, especially when coupled with features like Pod Disruption Budgets (PDBs), provides a safer, more stable, and resilient environment than traditional commodity infrastructure. Kubernetes' inherent self-healing mechanisms, combined with intelligent operators, automate recovery from node failures and other disruptions.
- Deterministic and High Performance with Local Storage: For applications demanding ultra-low latency and extremely high IOPS, leveraging local NVMe storage exposed directly to Kubernetes pods (as demonstrated with TiKV on Akamai's LKE) can deliver native-level performance. This approach often outperforms network-attached storage by eliminating network jitter and latency, offering deterministic performance crucial for demanding workloads.
- Scalability Through Distributed Architectures: Modern cloud-native databases and key-value stores, such as TiKV, are designed for horizontal scalability. They intelligently shard data and distribute it across numerous nodes, enabling linear performance scaling as capacity and RPS requirements grow, even into the millions of queries per second.
- Maturity and Enterprise Adoption of Open Source: Battle-tested, open-source projects like CloudNativePG and TiKV are not just experimental; they are mature, graduated CNCF projects actively used in production by large enterprises like Nokia and Civo to build robust, scalable, and cost-effective database-as-a-service (DBaaS) platforms.
Technical Deep Dive
▶ Watch: Graduated and incubating CNCF storage projects. (6:00)
The technical core of the presentation revolved around two distinct yet equally compelling demonstrations: CloudNativePG for relational databases and TiKV for high-performance key-value storage.
CloudNativePG: PostgreSQL on Kubernetes
Chris Milsted showcased CloudNativePG, an operator that simplifies the deployment and management of PostgreSQL clusters on Kubernetes. The central theme was the drastic reduction in complexity compared to traditional deployments. Instead of navigating a 290-page manual for Oracle Grid Infrastructure, a highly available PostgreSQL cluster with disaster recovery capabilities could be defined in approximately 90 lines of YAML.
The YAML definition exemplified key features:
- High Availability:
instances: 3configured a master and two replicas, establishing a resilient cluster. - Initialization: The
initDBsection specified how to initialize the database from scratch. - Backup and Recovery: Integration with Barman, a backup and recovery manager for PostgreSQL, allowed for streaming Write-Ahead Logs (WALs) to an external object store. This enabled robust disaster recovery, demonstrated by an
external cluster restorefrom a remote object storage. - Operator Intelligence: The CloudNativePG operator is designed to understand the underlying storage topology (local vs. remote, replicated vs. non-replicated). This intelligence allows it to make informed decisions during recovery, whether recreating volumes, copying data, or simply reattaching existing network-attached storage to a new or recovered Kubernetes node.
- Pod Disruption Budgets (PDBs): The talk emphasized that PDBs are a critical Kubernetes mechanism that, when combined with operators, creates a safer, more stable, and resilient persistent workload, protecting against voluntary disruptions.
TiKV: Distributed Key-Value Store for Extreme Performance
Alex Reid delved into TiKV, a graduated CNCF project recognized for its highly scalable, low-latency, distributed key-value store capabilities. TiKV offers both a raw key-value API (get, put, delete) and an ACID-compliant transactional API, serving as a foundation for more complex databases, including SQL layers.
Key technical attributes of TiKV highlighted:
- Horizontal Scalability: TiKV is engineered for massive scale, supporting hundreds to thousands of terabytes of data, billions of keys, and hundreds of thousands to millions of Requests Per Second (RPS). It achieves this by intelligently splitting the key space into regions (typically 100MB, but tunable) and distributing these regions across numerous storage nodes. As workload demands grow, more nodes can be added, and TiKV's internal intelligence automatically balances data, addresses hotspots, and ensures fault tolerance.
- Ultra-Low Latency: Capable of 1-10 millisecond latencies for both reads and writes, TiKV's performance is significantly boosted by its foundation on RocksDB. RocksDB is a high-performance, non-distributed key-value store developed by Facebook, benefiting from over 15 years of engineering and optimization.
- Cloud-Native Integration: TiKV provides a robust Kubernetes operator for streamlined deployment, upgrades, and automated failover. It also comes with comprehensive observability features, including pre-configured Grafana dashboards that offer both high-level cluster views and granular insights into sub-components.
- Local NVMe for Peak Performance: A critical configuration detail for the TiKV demonstration was the use of
storageClass: SSD storage. This setup exposed the local NVMe drives directly on the Kubernetes nodes within Akamai's LKE (Linode Kubernetes Engine) to the TiKV pods. This direct access to high-speed local storage is paramount for achieving the advertised low latency and high IOPS, circumventing the potential jitter and latency associated with network-attached storage. The configuration also used Kubernetes pod anti-affinity to ensure one TiKV pod per Kubernetes node, granting each pod full access to the node's resources.
Both demonstrations underscored the power of Kubernetes operators and the strategic choice of underlying infrastructure (especially storage) in achieving both operational simplicity and extreme performance for stateful workloads.
Demo / Proof of Concept
▶ Watch: Chris introduces Cloud Native PG demo. (6:50)
The talk featured two powerful live demonstrations that vividly illustrated the capabilities of cloud-native stateful storage.
CloudNativePG Automated Failover and Recovery:
Chris Milsted demonstrated the resilience of a PostgreSQL cluster managed by the CloudNativePG operator on Akamai's LKE. The demo involved a live, highly available PostgreSQL database running across multiple nodes. Chris intentionally killed one of the underlying Kubernetes nodes via the cloud manager. The demonstration then showcased the subsequent automated recovery process:
- Node Failure Detection: Kubernetes quickly detected the loss of the node, reflected by "red lines" indicating a downed instance.
- Kubernetes Recovery: Kubernetes initiated the recovery of the node, bringing it back into a
Readystate. - Operator Intervention: The CloudNativePG operator then took over, intelligently assessing the data layout (replicated, local, or remote storage). It determined the necessary actions, which included recreating volumes (if local storage was lost), copying data, or simply reattaching remote network-attached storage to the newly recovered Kubernetes node.
- Database Restoration: The operator orchestrated the reinitialization and reattachment of storage to the PostgreSQL pods, bringing the database back into a fully operational,
runningorstandby starting upstate.
This process highlighted the declarative, self-healing nature of cloud-native stateful applications, where the operator handles complex recovery logic without manual intervention, ensuring continuous availability and data integrity.
TiKV High-Performance Random Read IOPS:
Alex Reid presented an impressive demonstration of TiKV's extreme performance capabilities. The setup involved a 20-node Kubernetes cluster, with 13 nodes dedicated to TiKV (12 storage nodes, 1 control plane node) and 7 nodes reserved for the benchmark client.
- Data Pre-loading: Prior to the live demo, 10 billion keys were loaded into the TiKV cluster, with 5-way replication, resulting in 50 billion keys in total and approximately 1.5 terabytes of key-value data. Critically, the TiKV storage nodes were configured to use local NVMe SSDs on the LKE nodes via a
storageClass: SSD storage. - Benchmark Execution: A custom script launched about 15 instances of go YCSB, an open-source database benchmarking tool, across the 7 client nodes. These clients were configured to perform 1 million random read IOPS against the 10 billion key dataset.
- Real-time Observability: The performance was monitored in real-time using Grafana dashboards, which are pre-configured with TiKV. The dashboards displayed:
- CPU Utilization: All 12 TiKV storage nodes were running at an astounding 900-1000% CPU usage, indicating efficient and balanced workload distribution.
- Memory Usage: Each TiKV pod consumed approximately 20GB of memory, primarily for an in-memory cache to accelerate performance, which is a configurable behavior.
- Queries Per Second (QPS): The most striking result was the sustained
1 million RPSrate shown on the QPS dashboard, demonstrating TiKV's ability to handle massive random read workloads with high throughput.
This demonstration powerfully showcased how combining a distributed, high-performance database like TiKV with fast local NVMe storage and Kubernetes orchestration can achieve enterprise-grade performance metrics previously thought challenging in a cloud-native context.
Defensive Implications
▶ Watch: Oracle grid infrastructure: traditional database complexity. (7:20)
The insights from this talk provide several critical defensive implications for organizations operating stateful workloads in cloud-native environments:
- Embrace Operators for Operational Resilience: Defenders should prioritize the adoption of battle-tested Kubernetes operators for managing complex stateful applications like databases and key-value stores. Operators automate crucial day-2 operations such as backups, disaster recovery, upgrades, and failovers, significantly reducing the potential for human error and ensuring consistent operational procedures. This automation inherently strengthens an organization's defensive posture by standardizing recovery processes and minimizing downtime during incidents.
- Mandate Pod Disruption Budgets (PDBs): For all critical stateful workloads, implementing Pod Disruption Budgets (PDBs) is non-negotiable. PDBs ensure that a minimum number of replicas for a given application remain available during voluntary disruptions (e.g., node maintenance, upgrades, or scaling operations). This prevents cascading failures and maintains service availability, a key aspect of defensive resilience.
- Strategic Storage Selection Based on Workload: Understand that "one size does not fit all" for storage. For workloads requiring extreme performance and low latency, prioritize local NVMe storage on Kubernetes nodes, coupled with application-level replication (as seen with TiKV or PostgreSQL replicas) for data durability and availability. For workloads that require high mobility or simpler scaling without extreme performance, distributed block storage might be more appropriate. Defenders must assess the specific HA, consistency, durability, and performance needs of each application.
- Leverage Cloud-Native Observability: Utilize the rich observability features offered by cloud-native storage solutions and their operators (e.g., TiKV's Grafana dashboards). Comprehensive monitoring of CPU, memory, IOPS, and latency is crucial for identifying performance bottlenecks, detecting anomalies, and proactively addressing potential issues before they impact users. This visibility is a cornerstone of effective defense.
- Plan for Data Locality to Mitigate Network Jitter: For the most demanding, ultra-fast applications, consider architectures that prioritize data locality. Relying on local storage for primary data access and using application-level replication for redundancy can mitigate the unpredictable jitter and latency often introduced by network-attached storage, which can be a point of failure or performance degradation.
- Investigate Open-Source Solutions: The successful deployments at Nokia and Civo with open-source operators like CloudNativePG and platforms like TiKV demonstrate the maturity and robustness of these projects. Defenders should confidently explore and integrate such open-source solutions, which often provide transparent security models, community support, and avoid vendor lock-in.
Key Takeaways
- Stateless is a Myth: All applications eventually interact with state; embracing cloud-native principles for stateful workloads is essential for modern architectures.
- Operators are Transformative: Kubernetes operators drastically simplify the deployment, management, and day-2 operations of complex stateful applications, replacing manual processes with declarative automation.
- Cloud-Native Benefits for State: Stateful workloads gain significant advantages from Kubernetes' automation, self-healing, autoscaling, and ability to deliver deterministic high performance.
- Local NVMe for Peak Performance: Combining high-performance distributed systems like TiKV with fast local NVMe storage on Kubernetes nodes can achieve millions of IOPS and ultra-low latency, often outperforming network-attached storage.
- Resilience through PDBs and Replication: Pod Disruption Budgets (PDBs) are crucial for maintaining availability during voluntary disruptions, while application-level replication ensures data durability and high availability in the event of node failures.
- Open Source is Production-Ready: Mature, open-source cloud-native storage projects are being successfully adopted by large enterprises to build robust and scalable platforms.
About the Speaker(s)
Alex Chircop is a Chief Architect at Akamai Cloud and a distinguished member of the Technical Oversight Committee (TOC) for the Cloud Native Computing Foundation (CNCF). He brings extensive expertise in cloud-native technologies and storage solutions.
Lori Lorusso serves as the Head of Community at Percona, a 100% open-source company specializing in database software and services. She is also a CNCF Ambassador, actively promoting and supporting the cloud-native ecosystem.
Alex Reid is a Principal Engineer at Akamai, focusing on developing and optimizing cloud infrastructure and services, particularly in the realm of high-performance storage and distributed systems.
Chris Milsted is a Product Architect at Akamai, contributing to the design and implementation of cloud-native solutions, with a particular emphasis on managed Kubernetes and stateful application deployments.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk effectively dismantles the "stateless architecture" myth by demonstrating how Kubernetes, powered by intelligent operators like CloudNativePG and TiKV, can manage high-performance stateful workloads with unprecedented efficiency and resilience. The speakers, clearly deeply technical, provided compelling live demonstrations showcasing automated failover for PostgreSQL and achieving a staggering 1 million random read IOPS with TiKV leveraging local NVMe storage. This session offers a pragmatic, technically rich pathway for architects and engineers to confidently deploy and operate critical stateful applications in cloud-native environments, moving beyond platitudes to concrete…
Heather Calloway (CISO) — STRONG ACCEPT
This talk effectively debunks the pervasive 'stateless only' myth in cloud-native architectures, presenting a compelling case for leveraging Kubernetes to manage high-performance stateful workloads. It offers actionable insights for architects and engineering leaders on achieving automated resilience, scalability, and deterministic performance for critical data infrastructure using mature open-source operators and strategic storage choices. While the technical depth is significant, the implications for enterprise risk, operational efficiency, and incident recovery are profound, providing a clear pathway to strengthen an organization's foundational data security and availability posture.