Longhorn: Intro, Deep Dive and Q&A - David Ko, SUSE

David Ko, SUSE

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk, delivered by David Ko, Engineering Director at SUSE, provides a comprehensive overview and deep dive into Longhorn, a distributed block storage solution specifically designed for Kubernetes. Ko's primary mission at KubeCon was to promote Longhorn's adoption globally, highlighting its capabilities for both new and existing users. The session covers Longhorn's core functionalities, its project status, recent and upcoming releases, and a detailed technical explanation of its architecture, including the pivotal new V2 data engine.

Watch on YouTube

Visual summary for Longhorn: Intro, Deep Dive and Q&A - David Ko, SUSE by David Ko, SUSE
Visual summary for Longhorn: Intro, Deep Dive and Q&A - David Ko, SUSE by David Ko, SUSE

Key moments

  1. 0:00 Talk Introduction and Agenda Overview
  2. 1:30 Highlighting upcoming V2 Data Engines for performance
  3. 2:00 What is Longhorn? Core Features Explained
  4. 4:00 Project Status, Growth, and Adoption Metrics
  5. 7:00 Longhorn Support for Immutable OS like Talos
  6. 8:00 Release Cadence and Upcoming Release Schedule

Longhorn: Intro, Deep Dive and Q&A

Speakers: David Ko; SUSE

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=REkSMbRrBU4

Overview

This talk, delivered by David Ko, Engineering Director at SUSE, provides a comprehensive overview and deep dive into Longhorn, a distributed block storage solution specifically designed for Kubernetes. Ko's primary mission at KubeCon was to promote Longhorn's adoption globally, highlighting its capabilities for both new and existing users. The session covers Longhorn's core functionalities, its project status, recent and upcoming releases, and a detailed technical explanation of its architecture, including the pivotal new V2 data engine.

Longhorn addresses a critical need in the Kubernetes ecosystem: providing highly available, persistent storage that is both robust and performant. By leveraging local disk resources across a Kubernetes cluster, Longhorn enables resilient data management through features like replication, in-cluster snapshots, and external backups, facilitating cross-cluster disaster recovery. The talk emphasizes Longhorn's platform agnosticism, confirming its operation across various environments, from on-premises to public clouds, and its status as a CNCF incubating project actively seeking broader community contributions to achieve graduation.

The significance of Longhorn, particularly with the introduction of its V2 data engine, lies in its commitment to pushing the boundaries of performance for data-intensive Kubernetes workloads. While V1, based on iSCSI, serves general use cases effectively, V2's adoption of technologies like SPDK and NVMe over Fabrics promises a substantial leap in IOPS and throughput, bringing Longhorn closer to raw disk performance. This evolution, coupled with continuous enhancements in data protection, operational efficiency, and broader ecosystem integration, positions Longhorn as a vital component for organizations seeking scalable, reliable, and high-performance storage within their Kubernetes deployments.

Background

▶ Watch: Talk Introduction and Agenda Overview (0:00)

The landscape of cloud-native applications, orchestrated by Kubernetes, inherently demands robust and highly available persistent storage. Traditional storage solutions often struggle to integrate seamlessly with the dynamic, distributed nature of containerized workloads, leading to complexities in provisioning, scaling, and ensuring data resilience. This challenge is precisely what Longhorn aims to solve by offering a cloud-native, distributed block storage solution built directly on Kubernetes.

Longhorn's genesis lies in the need to provide highly available persistent volumes for Kubernetes clusters, abstracting away the underlying storage infrastructure. It achieves this by utilizing the local disks available on each node within the cluster, pooling them into a unified storage system. Data resilience is a core tenet, facilitated by configurable data replication, ensuring that multiple copies of data exist across different nodes. This design inherently guards against single points of failure, making volumes highly available even if a node or disk goes offline. Beyond basic persistence, Longhorn extended its capabilities to include in-cluster snapshots for point-in-time recovery, and external backups to object storage (like S3), enabling robust data protection strategies. The external backup functionality further underpins cross-cluster disaster recovery, allowing data to be restored to entirely different Kubernetes clusters.

The project has seen significant growth and adoption since its inception, now standing as a CNCF incubating project. David Ko highlighted its impressive growth metrics, including over 159,000 nodes and 28,000 clusters globally running Longhorn. Its versatility allows it to run on diverse environments—from home labs and edge deployments to on-premises infrastructure, private clouds, and major public clouds. This platform agnosticism is a key differentiator, as SUSE and the community invest considerable effort in verifying its compatibility across various distributions and operating systems.

Longhorn's applicability spans numerous domains, including AI/ML workloads, data virtualization, and notably, Kubernetes virtualization with projects like KubeVirt. For KubeVirt users, Longhorn provides crucial features like volume migration, which is essential for managing virtual machine volumes within Kubernetes. Recent developments have also focused on supporting immutable operating systems like Talos Linux, a strong community request, ensuring Longhorn's compatibility with modern, secure, and streamlined Kubernetes distributions. The project operates on a predictable four-month minor release cadence (January, May, September), allowing users to anticipate new features and plan upgrades effectively.

Key Findings

▶ Watch: What is Longhorn? Core Features Explained (2:00)

The talk presented several significant advancements and capabilities of Longhorn, with the Longhorn V2 Data Engine standing out as the most pivotal development. This new engine represents a fundamental shift towards high-performance storage, addressing the limitations of the V1 engine for data-intensive applications. By leveraging cutting-edge technologies like SPDK (Storage Performance Development Kit) and NVMe over Fabrics (NVMeoF), V2 aims to deliver performance metrics significantly closer to raw disk speeds. Benchmarks presented later in the talk visibly illustrate the substantial improvements in IOPS, throughput, and latency, especially when V2 is configured with dedicated CPU resources. This finding is crucial for organizations deploying databases, AI/ML workloads, or other applications with demanding I/O profiles on Kubernetes.

Another key finding is Longhorn's enhanced data protection and disaster recovery capabilities. While incremental backups were always a feature, the introduction of full backups in version 1.8 provides users with greater flexibility and resilience, particularly in scenarios where incremental chains might become problematic. The ability to configure multiple backup targets allows for sophisticated backup tearing strategies, catering to different data retention policies (e.g., hot vs. cold data). Coupled with existing cross-cluster disaster recovery, these features solidify Longhorn's position as a comprehensive solution for data resilience.

Operational efficiency and user experience have also seen significant improvements. Longhorn now supports live upgrades and migrations without service interruption, a critical feature for maintaining application uptime in dynamic environments. The introduction of automatic volume expansion further simplifies storage management, allowing volumes to grow as needed without manual intervention or downtime. Furthermore, Longhorn's commitment to supporting modern infrastructure is evident in its verification and support for immutable operating systems like Talos Linux, extending its reach to more secure and streamlined Kubernetes distributions.

Finally, the talk highlighted the project's robust growth and active community. With over 159,000 nodes and 28,000 clusters running Longhorn, and its status as a CNCF incubating project, it demonstrates strong adoption. The emphasis on community contributions for achieving graduation underscores the collaborative nature of its development, inviting users to engage in various aspects from deployment and troubleshooting to core engine development and ecosystem tooling.

Technical Deep Dive

▶ Watch: Project Status, Growth, and Adoption Metrics (4:00)

Longhorn's architecture is fundamentally divided into two primary components: the control plane and the data plane. This separation allows for robust management and flexible data handling within a Kubernetes environment.

The control plane is deeply integrated with Kubernetes. It primarily relies on Kubernetes Custom Resources (CRs) and events to understand workload requirements and volume requests. For user interaction, Longhorn provides a user-friendly Longhorn UI, though a new, more modern UI framework is currently under development. The Longhorn CSI driver is a critical component, deployed on every node, which listens to Kubernetes calls for provisioning, resizing, attaching, and detaching volumes. While a Longhorn API exists, David Ko strongly advises users to interact with Longhorn through its Kubernetes CRs for automation, aligning with standard Kubernetes operational practices. The core of the control plane is the Longhorn Manager, which runs on every node, orchestrating volume operations, replica placement, and recovery processes. When a volume is provisioned, the manager determines optimal replica placement based on available disk space and anti-affinity rules. In the event of a replica failure, the manager initiates replica rebuilding, efficiently restoring data from healthy replicas.

The data plane is where the actual data operations occur, and Longhorn offers two distinct engines: V1 and V2.

The V1 data engine is the current stable and widely used engine. It employs the iSCSI protocol for the frontend, presenting volumes as block devices to the Kubernetes workload. On the backend, for inter-replica communication and data synchronization, V1 uses a custom protocol based on TCP. This design supports data locality, meaning if a replica resides on the same node as the Longhorn engine and workload, data traffic can stay local, improving efficiency. However, iSCSI, while robust, has inherent performance limitations for highly demanding applications. Longhorn's V1 disk management is file system based.

The V2 data engine is a significant leap forward, designed specifically for high-performance, data-intensive workloads. It is built upon the SPDK (Storage Performance Development Kit), a collection of user-space libraries and tools for high-performance I/O. SPDK was recently donated to the Linux Foundation, with Intel being a major contributor, alongside SUSE's Longhorn team. V2 leverages NVMe over Fabrics (NVMeoF) for its frontend, providing a much faster block device interface to applications compared to iSCSI. For the downstream replicas, V2 utilizes logical volumes, which is an SPDK-based concept. An even more performant frontend, uBLK, based on IO_uring, is under development for the upcoming 1.9 release, promising further performance gains beyond NVMeoF. V2's approach requires dedicated CPU resources due to its polling mode operation, which offers maximum performance but can be resource-intensive. The team is investigating an interruption mode to make V2 more accessible to lower-spec machines.

Longhorn provides a rich set of capacities for Kubernetes volumes:

  • Volume Modes and Access Modes: Supports ReadWriteOnce (RWO), ReadOnlyMany (ROM), and ReadWriteMany (RWX). The RWX mode is particularly notable as it integrates with volume migration capabilities, crucial for virtualized environments like KubeVirt.
  • CSI Protocol Fulfillment: Longhorn's CSI driver supports all standard CSI operations, including volume provisioning, attachment, snapshots, cloning, and expansion. Volume Group functionality is currently in beta.
  • Snapshots: Offers simple, I/O-efficient snapshots. V2 will introduce delta snapshots based on logical volumes for even more efficient replica rebuilding at the chunk level.
  • Live Operations: Supports live migration and live upgrade of volumes and the Longhorn system itself, ensuring no service interruption by using a standby engine and switching traffic.
  • User Experience: In addition to the UI, Longhorn offers a CLI for day-one and day-two operations, troubleshooting, and managing replicas. It also provides comprehensive monitoring metrics for volume usage, disk usage, and component-level performance.
  • Storage Management: V1 manages disks based on file systems, while V2 leverages block device-based storage using SPDK.
  • Replication: Features anti-affinity to distribute replicas across different nodes and even different disks within a node. Auto-balancing ensures optimal replica distribution based on disk usage. Storage tags and storage classes allow users to specify which disks or nodes should be used for particular workloads.
  • Data Protection: Includes encryption, volume protection, in-cluster snapshots, and out-of-cluster backups. Backups support incremental and full backups, multiple backup targets (e.g., for tearing), and compression using algorithms like LZO, GZIP, and LZ4. Users can also configure the block size for backups to optimize for different data types (e.g., larger blocks for video streaming).
  • Disaster Recovery: Facilitated by external backups, allowing restoration to different clusters.

Recent and upcoming features include configurable CPU cores for the V2 engine, DR volume support for V2, auto-salvage for failed V2 volumes, V2 live migration and encryption, improved replica rebuilding via chunk-level delta snapshots, backing image updates for VMs, V2 support for Talos Linux, and automatic RWX volume expansion without downtime. The upcoming 1.9 release will further enhance uBLK frontend, fast volume cloning, dedicated storage network capabilities, and offline replica rebuilding for both V1 and V2.

Demo / Proof of Concept

▶ Watch: Longhorn Support for Immutable OS like Talos (7:00)

While the talk did not feature a live demonstration of Longhorn in action, David Ko presented compelling benchmark results that served as a crucial proof of concept for the performance enhancements delivered by the new V2 data engine. These benchmarks were conducted using FIO (Flexible I/O Tester) on Equinix Metal infrastructure, which was sponsored by the CNCF. It was noted that Equinix Metal is retiring, and future benchmarks might utilize Oracle OCI.

The presentation included several graphs comparing the performance of raw local disk (blue bar), Longhorn V1 volumes (red bar), and Longhorn V2 volumes (green bar) across various metrics:

  1. IOPS (Input/Output Operations Per Second): The graphs clearly showed that V2 volumes delivered significantly higher IOPS compared to V1. As the number of dedicated CPU cores allocated to the V2 engine increased (e.g., from one to two to four cores), the IOPS performance of V2 scaled proportionally, demonstrating its ability to leverage additional compute resources for improved I/O. While V2 didn't fully match the raw disk performance, it substantially closed the gap compared to V1.
  2. Write IOPS: Similar to read IOPS, V2 exhibited superior write performance, with increasing CPU allocation leading to better results. This is critical for transactional databases and other write-intensive applications.
  3. Read Throughput and Write Throughput: The throughput graphs mirrored the IOPS results, indicating that V2 could handle a much larger volume of data transfer per second. The speaker highlighted that with more replicas, the throughput naturally increased, showcasing Longhorn's distributed nature.
  4. Read Latency and Write Latency: V2 also demonstrated lower latency compared to V1, which is crucial for applications sensitive to I/O delays. However, it was noted that in the testing environment, read latency eventually hit an upper bound, suggesting that the environment itself might have been a limiting factor, and even better latency could be observed in more optimized setups.

These benchmark results provided concrete evidence that the architectural shift in V2, particularly the adoption of SPDK and NVMe over Fabrics, successfully translates into tangible performance improvements. The ability to configure CPU core allocation for V2 further empowers users to tune performance according to their workload requirements and available resources, making Longhorn V2 a viable option for high-performance computing within Kubernetes.

Defensive Implications

▶ Watch: Release Cadence and Upcoming Release Schedule (8:00)

The detailed insights into Longhorn's capabilities and upcoming features provide several critical defensive implications for organizations deploying and managing Kubernetes clusters.

Firstly, adopting Longhorn for highly available persistent storage directly enhances the resilience of containerized applications. Its native replication across nodes protects against disk and node failures, ensuring continuous data availability. Defenders should configure appropriate replica counts (e.g., 3 replicas for critical data) and leverage anti-affinity rules to distribute replicas across different nodes and even physical disks, minimizing the blast radius of a localized failure.

Secondly, the introduction of the V2 data engine is a game-changer for performance-sensitive workloads. For applications like databases, analytics engines, or AI/ML training, migrating to V2 can significantly improve I/O performance, directly impacting application responsiveness and throughput. However, defenders must be aware of V2's resource requirements, specifically the need for dedicated CPU cores. Proper capacity planning and resource allocation are essential to fully realize V2's benefits without impacting other workloads on the same nodes. The ongoing investigation into an "interruption mode" for V2 also suggests future opportunities for broader adoption on lower-spec hardware.

Thirdly, Longhorn's comprehensive data protection features are invaluable. Implementing robust backup strategies is paramount. The ability to perform both incremental and full backups to multiple external targets allows for flexible and resilient backup policies. Organizations should utilize these features to create tiered backup strategies (e.g., hot data to one target, cold data to another), leverage compression (LZO, GZIP, LZ4) to optimize storage costs and network bandwidth, and configure recurring system backups for the entire Longhorn system. The cross-cluster disaster recovery capability, enabled by external backups, is a critical component of any business continuity plan, allowing rapid restoration of services in a geographically separate cluster.

Fourthly, operational benefits like live upgrades and migrations minimize planned downtime. Defenders can leverage these features to apply security patches or upgrade Longhorn versions without interrupting critical application services, thereby reducing maintenance windows and improving overall system availability. Similarly, automatic volume expansion prevents service interruptions due to storage exhaustion, allowing for proactive scaling.

Finally, Longhorn's support for immutable operating systems like Talos Linux contributes to a stronger security posture. Immutable infrastructure reduces configuration drift and makes systems more predictable and less susceptible to tampering. Integrating Longhorn with such hardened operating systems strengthens the overall security of the Kubernetes stack. Furthermore, utilizing features like volume encryption adds another layer of defense for data at rest. Regular monitoring of Longhorn's metrics (volume, disk, and component level) is also crucial for identifying potential issues early and maintaining the health and security of the storage infrastructure.

Key Takeaways

  • Longhorn is a Robust, Cloud-Native Storage Solution for Kubernetes: As a CNCF incubating project, Longhorn provides highly available, distributed block storage, leveraging local node disks to ensure data resilience through replication, in-cluster snapshots, and comprehensive external backup capabilities.
  • The V2 Data Engine Delivers Significant Performance Gains: Longhorn V2, built on SPDK, NVMe over Fabrics, and future uBLK (IO_uring) technology, offers a substantial leap in IOPS, throughput, and reduced latency compared to V1, making it ideal for data-intensive applications, though it requires dedicated CPU resources.
  • Comprehensive Data Protection and Disaster Recovery: Longhorn supports incremental and full external backups to multiple targets, enabling tiered backup strategies, data compression, and robust cross-cluster disaster recovery, safeguarding critical data against various failure scenarios.
  • Enhanced Operational Efficiency and Uptime: Features like live upgrades and migrations without service interruption, automatic volume expansion, and support for immutable operating systems like Talos Linux streamline operations and maximize application availability.
  • Active Community and Predictable Release Cadence: Longhorn boasts significant global adoption with tens of thousands of clusters and nodes, maintains a consistent four-month minor release schedule, and actively encourages community contributions across all aspects of the project.
  • Strategic Adoption for Performance and Resilience: Users should evaluate V2 for performance-critical workloads while planning for its CPU requirements, and fully utilize Longhorn's extensive data management, protection, and recovery features to build highly resilient Kubernetes environments.

About the Speaker(s)

David Ko is an Engineering Director at SUSE, where he leads the Longhorn team. His primary mission, as articulated at KubeCon, is to promote the adoption of Longhorn globally. He works closely with both internal teams and external contributors to drive the development and enhancement of the Longhorn project, pushing its capabilities in the Kubernetes storage ecosystem.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This session, while originating from a vendor, delivered a remarkably substantive technical deep dive into Longhorn, particularly highlighting the significant advancements in its V2 data engine. David Ko demonstrated a clear command of the subject matter, effectively detailing Longhorn's architecture, its critical role in Kubernetes persistent storage, and the tangible performance improvements achieved through technologies like SPDK and NVMe over Fabrics. It successfully transcended typical marketing fluff, providing actionable insights and robust technical details for anyone serious about managing high-performance, resilient storage in cloud-native environments.

Heather Calloway (CISO) — STRONG ACCEPT

This session on Longhorn, particularly with the introduction of its V2 data engine, presents a compelling case for robust and high-performance persistent storage in Kubernetes environments. It moves beyond mere technical specifications to clearly articulate the defensive implications and operational value, providing actionable insights for security leaders responsible for data resilience and application availability. While a deep dive into storage architecture, the emphasis on data protection, disaster recovery, and operational efficiency makes it highly relevant for those managing critical cloud-native infrastructure.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025