Streamlined Efficiency: Unshackling Kubernetes Image Volumes for Rapid AI Model... E. Rey & Y. Yuan

E. Rey, Y. Yuan

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk, presented by E. Rey from Microsoft's Azure Container Registry team and Y. Yuan from Alibaba Cloud, addresses a critical bottleneck in modern AI/ML workflows: the inefficient loading of large datasets and models within Kubernetes environments. While significant strides have been made in optimizing application startup times through mechanisms like artifact streaming, these solutions often fall short when dealing with the colossal and dynamic data requirements of AI. The core problem lies in the traditional approach of packaging vast datasets directly into OCI images, a process that is shown to be prohibitively slow, resource-intensive, and difficult to manage.

Watch on YouTube

Visual summary for Streamlined Efficiency: Unshackling Kubernetes Image Volumes for Rapid AI Model... E. Rey & Y. Yuan by E. Rey, Y. Yuan
Visual summary for Streamlined Efficiency: Unshackling Kubernetes Image Volumes for Rapid AI Model... E. Rey & Y. Yuan by E. Rey, Y. Yuan

Key moments

  1. 0:00 Introduction to the problem of large AI dataset loading
  2. 2:00 Challenges with data parallelism for AI/ML workloads
  3. 3:25 Introducing OCI registries as a potential solution
  4. 4:10 Key benefits: garbage collection, versioning, user familiarity
  5. 6:30 Limitations of OCI registries for volume mounting
  6. 8:00 Illustrating the high cost of packaging large datasets

Streamlined Efficiency: Unshackling Kubernetes Image Volumes for Rapid AI Model and Dataset Loading

Speakers: E. Rey, Software Engineer, Microsoft; Y. Yuan, Senior Software Engineer and Researcher, Alibaba Cloud

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=nHGzMmstR0E

Overview

This talk, presented by E. Rey from Microsoft's Azure Container Registry team and Y. Yuan from Alibaba Cloud, addresses a critical bottleneck in modern AI/ML workflows: the inefficient loading of large datasets and models within Kubernetes environments. While significant strides have been made in optimizing application startup times through mechanisms like artifact streaming, these solutions often fall short when dealing with the colossal and dynamic data requirements of AI. The core problem lies in the traditional approach of packaging vast datasets directly into OCI images, a process that is shown to be prohibitively slow, resource-intensive, and difficult to manage.

The speakers introduce Elink, an innovative solution designed to leverage the existing infrastructure of OCI registries for metadata management while decoupling the actual data storage. Elink aims to streamline the process of making large datasets available to Kubernetes pods by eliminating the need for time-consuming data repackaging. By doing so, it promises to significantly accelerate AI model training and data analysis, reduce operational costs, and enhance the scalability and flexibility of AI workloads running on Kubernetes. This approach is particularly relevant for organizations struggling with the overhead of managing ever-growing, versioned datasets in a cloud-native context.

Background

▶ Watch: Introduction to the problem of large AI dataset loading (0:00)

The journey towards efficient container startup has seen considerable progress, particularly with the advent of artifact streaming. This technology effectively optimizes the delivery of application-based workloads by allowing container images to start executing before all layers are fully downloaded. However, as AI and Machine Learning workloads grow in complexity and scale, they introduce unique challenges that artifact streaming alone cannot solve. These challenges primarily revolve around the management and loading of massive datasets, often tens or hundreds of gigabytes, or even terabytes, which are integral to model training.

Several critical issues arise when integrating large datasets into Kubernetes-native AI pipelines:

  • Data Parallelism: AI training frequently requires parallel access to data across numerous nodes, demanding simultaneous and continuous high-speed data availability.
  • Continuous Data Access: Idle GPUs due to slow data loading translate directly into significant financial waste. Data must be accessible quickly and consistently.
  • Scalability: Solutions must scale to thousands of nodes without performance degradation, a common requirement for large-scale distributed training.
  • Versioning: Both application code and datasets evolve. A robust system needs to enable synchronized versioning of data alongside the models and applications that consume it.
  • Management Overhead: Any solution must be easy to deploy, manage, and integrate into existing CI/CD pipelines to ensure widespread adoption.

OCI registries, the backbone of container image distribution, inherently address many of these requirements. They are designed for performance, scalability, and high availability, making them familiar to Kubernetes users. Registries also offer robust versioning capabilities through tags and digests, ensuring data consistency and immutability. Furthermore, their native support for garbage collection is crucial for managing the exponential growth of data, preventing unchecked cost accumulation from stale or unused datasets. An IDC report cited in the talk projects a global data volume of 175 zettabytes by 2025, underscoring the urgency of efficient data management.

Despite these advantages, registries were not originally designed for direct volume mounting or for efficiently storing arbitrary, large, frequently changing datasets. Several limitations hinder their effectiveness in this role:

  • Registry Size Limitations: Different registries impose varying layer size limits. For instance, Azure Container Registry (ACR) has a 200 GB per layer limit, which is not standardized by the OCI Distribution Spec and can vary by implementation. This makes large single-file datasets problematic.
  • Overlay Filesystem Overhead: OCI images typically use an overlay filesystem structure. Each modification or addition to the image content creates a new layer. For dynamic, large datasets, this leads to an explosion of layers, hitting layer limits and incurring significant storage and processing costs.
  • Prohibitive Packaging Costs: The most significant drawback is the time and resources required to package large datasets into OCI image layers. The speakers illustrated this with a practical example: packaging a 22 GB dataset containing 700,000 small image files from Kaggle into an OCI image using docker COPY took nearly four hours. Even smaller datasets (a few gigabytes) required many minutes. This overhead makes frequent updates or large-scale data ingestion impractical and costly.

Prior efforts like the Oras project have formalized the use of OCI images for arbitrary data, and existing image volumes allow loading container images into a filesystem. However, these solutions still largely rely on the underlying OCI image structure, which includes the packaging overhead that Elink aims to circumvent. The challenge, therefore, is to retain the benefits of OCI registries (scalability, versioning, management) while overcoming the packaging and storage inefficiencies for large, dynamic datasets.

Key Findings

▶ Watch: Introducing OCI registries as a potential solution (3:25)

The central discovery and contribution of this work is the realization that the core problem in leveraging OCI registries for large AI datasets isn't the registry itself, but the packaging process. By shifting the focus from packaging the data into the OCI image to packaging metadata about the data, the system can bypass the primary performance bottleneck.

The key findings and proposed solution, Elink, are:

  1. Packaging is the Bottleneck: Traditional OCI image creation involves copying and layering actual data, which is exceedingly slow and resource-intensive for large datasets (e.g., 22GB dataset took nearly 4 hours to package).
  2. Metadata-Centric Approach: The inspiration comes from existing solutions like SQL OCI and OverlayBD, which use external indexes to create remote snapshots for streaming data without repackaging. Elink extends this by proposing to build an "index for the entire storage bucket."
  3. Elink's Core Components: The solution hinges on three main components:
  • Data Set Description: A high-level definition of the data.
  • Reference List: A detailed manifest of individual data objects, including their remote location, desired mount path, versioning information (Etag), and size. This effectively replaces the actual data layers in the OCI artifact.
  • Remote Snapshotter: A component that interprets the reference list to create a mount point, streaming data on demand from its remote source.
  1. OCI Registry as Metadata Store: OCI registries are leveraged not for storing the bulk data, but for distributing and versioning the lightweight reference list. This retains the benefits of OCI's existing infrastructure for versioning, garbage collection, and widespread tooling familiarity.
  2. On-Demand Streaming: By using the reference list, the container runtime and snapshotter can create a virtual filesystem view where files are accessed on demand directly from their backend storage (e.g., object storage), eliminating the need to download and unpack entire datasets upfront.
  3. Significant Performance Improvements: Elink demonstrates a dramatic reduction in packaging time (from nearly 4 hours to under 2 minutes for the 22GB dataset) and superior data access performance, especially for scenarios with many small files, compared to traditional OCI images and even high-performance S3 filesystems like Goofys.

In essence, Elink redefines how OCI artifacts are used for data volumes, transforming them from data containers into intelligent data pointers, thereby "unshackling" Kubernetes image volumes from the constraints of traditional image building.

Technical Deep Dive

▶ Watch: Key benefits: garbage collection, versioning, user familiarity (4:10)

Elink's technical architecture is built around the concept of a reference list packaged within an OCI artifact, which then enables a remote snapshotter to create a mount point for on-demand data access. This approach fundamentally shifts the heavyweight data handling from the image build process to the runtime, leveraging existing OCI registry capabilities for metadata distribution.

The Reference List

At the heart of Elink is the reference list, a structured set of records that describes all the data objects an AI workload needs to access. Instead of packaging the actual data blobs into OCI layers, Elink packages this lightweight metadata. Each record within the reference list contains at least four crucial pieces of information:

  1. Source Path: The original, full path to the object in the remote storage backend (e.g., s3://bucket/path/to/data.csv).
  2. Mount Path: The relative path where the object should appear within the container's mounted volume (e.g., /mnt/data/data.csv).
  3. Etag: An entity tag, typically an MD5 checksum or version identifier, which reflects the data's state. This is critical for detecting data changes and ensuring consistency without downloading the entire object.
  4. File Size: The size of the object in bytes, enabling efficient resource allocation and range requests.

This reference list effectively serves as a manifest, providing a virtualized view of the dataset without containing the dataset itself.

Packaging the Reference List into an OCI Artifact

To integrate with the existing OCI ecosystem, the reference list is packaged as a special layer within an OCI artifact:

  1. Registry as Backend Storage Proxy: The OCI registry acts as an intermediary or proxy. When requesting a blob (which in this case is the reference list), the registry can either return the blob content directly or a redirect URL to the actual backend storage. This mechanism allows the snapshotter to infer the backend storage endpoint. Combined with the source path from the reference list, the snapshotter can construct the full, actual URL for the target data blob.
  2. Special Annotation Field: A key technical detail is the introduction of a special annotation field within the OCI manifest. This annotation identifies a specific layer as an Elink reference list, signaling to the container runtime and snapshotter that this isn't a traditional data layer but a set of pointers to remote data. This allows existing OCI artifact handling mechanisms to be extended for Elink. While an annotation is used for the proof of concept, the speakers acknowledge that more standardized or robust approaches might emerge.
  3. Common Format: The reference list itself must be saved in a common, parsable format, such as CSV or JSON. This ensures that the registry (for potential analysis or validation) and, more importantly, the snapshotter can easily parse and interpret the list's contents.

Mount Point Creation and Streaming Loading

Once the Elink OCI artifact (containing the reference list) is pushed to a registry, the process for creating a mount point and streaming data unfolds:

  1. Pulling and Unpacking Reference List: During container startup, the container runtime (e.g., containerd) pulls the OCI artifact. When it encounters the layer identified by the special annotation, it unpacks the reference list, not as data, but as metadata for the snapshotter.
  2. Snapshotter Creates Mount Point: The remote snapshotter (e.g., an OverlayBD-based one) parses the reference items from the unpacked list. For each record, it creates a file entry at the specified mount path within the container's filesystem. Crucially, this file entry is a virtual entry; it doesn't contain the data itself but acts as a pointer to the source path in the remote storage.
  3. Authorization: A critical consideration is access control. The registry might not have direct access permissions to all remote objects in the backend storage. Therefore, an additional authorization mechanism is often required at runtime to grant the container or the snapshotter the necessary permissions to access the actual data blobs from the remote storage. This could involve Kubernetes service accounts, IAM roles, or other credential injection methods.
  4. Streaming Loading (On-Demand Access): When an application inside the container attempts to read a file from the mount point, the streaming service intercepts this IO request.
  • It consults the reference list to identify the corresponding source path and Etag for the requested file.
  • It performs an Etag validation by potentially making an HTTP HEAD request to the remote storage to check if the Etag of the remote object matches the one recorded in the reference list. This verifies data integrity and ensures the data hasn't been modified since the reference list was generated.
  • If the Etag matches, the IO request is converted into a range GET request to the remote target's backend storage. This means only the specific bytes or blocks of data requested by the application are downloaded, rather than the entire file. This is fundamental to saving bandwidth and accelerating access.

OverlayBD Integration

For the proof of concept, the speakers chose OverlayBD as the underlying technology for the remote snapshotter. OverlayBD offers several advantages:

  • Merged View as Virtual Block Device: OverlayBD presents a merged view of image layers as a virtual block device, which is highly efficient.
  • On-Demand Data Transfer: It supports on-demand data transfer at the disk sector level, providing fine-grained control and performance.
  • Block Device Interface: Unlike FUSE-based filesystems, OverlayBD uses a block device interface, which generally offers better performance, especially for small file access, and is more mature in handling stability issues like crash recovery.
  • Production Readiness: OverlayBD is widely used in production environments at scale, including Azure, Alibaba Group, and Databricks, attesting to its robustness.
  • Turbo OCI: OverlayBD's lightweight mode, Turbo OCI, already supports indexing to OCI images and can build an ERS (Elastic Remote Storage) filesystem from an image index locally.
  • Elink Implementation: Integrating Elink with OverlayBD primarily involves implementing a simple function to map the incoming IO requests from the mount point to the corresponding entries in the reference list. OverlayBD then handles the conversion of these logical IO requests into physical read/write operations through its block driver. This synergy makes OverlayBD an ideal choice for Elink's remote snapshotter.

This detailed technical approach allows Elink to provide a performant, scalable, and manageable solution for large dataset loading in Kubernetes without the traditional overheads.

Demo / Proof of Concept

▶ Watch: Limitations of OCI registries for volume mounting (6:30)

The speakers presented compelling performance test results to validate Elink's effectiveness, focusing on two critical aspects: packaging time and data access performance. The tests were conducted using datasets from kaggle.com, specifically popular machine learning training datasets.

Packaging Phase Performance

The most dramatic improvement was observed during the packaging phase. The traditional method of packaging a dataset into an OCI image involves copying files into Docker layers, which is time-consuming. Elink, however, transforms this process into merely recording reference data—the creation of the reference list.

  • Test Case: A 22 GB dataset containing over 700,000 small files (mostly images, making them less compressible).
  • Traditional OCI Packaging: Using docker COPY in a Dockerfile, packaging this dataset into an OCI image took nearly four hours. This highlights the significant overhead of layer creation and data copying for large, complex datasets.
  • Elink Packaging: In stark contrast, Elink's build speed for the same dataset was less than two minutes. This represents an order-of-magnitude improvement, demonstrating that decoupling data from the OCI artifact and packaging only metadata drastically reduces the time and resources required to make a dataset "available" as an OCI artifact.

This result alone makes Elink a game-changer for AI/ML workflows that require frequent updates or rapid onboarding of large datasets.

Full Volume Access Performance

Beyond packaging, the talk also presented a comparative analysis of data access performance. This test measured the throughput when accessing data from a mounted volume, including any preparation time.

  • Competitors:
  • Traditional OCI Images: The default behavior, involving downloading and unpacking the entire image before mounting.
  • Goofys: A high-performance, FUSE-based filesystem implementation for AWS S3. This provides a strong baseline for direct object storage access.
  • Elink: The proposed solution, leveraging OverlayBD for streaming access.
  • Preparation Time:
  • OCI Image: Includes the full download and unpackaging of the image.
  • Goofys/Elink: Includes the time taken to create the mount point and build the filesystem metadata. This is significantly faster as it doesn't involve full data download.
  • Results:
  • Scenario with Large Number of Small Files (e.g., US Accident dataset): Elink performed significantly better than both traditional OCI images and Goofys. This is a critical advantage for many AI/ML datasets, which often consist of numerous small files (images, text snippets, etc.). OverlayBD's block device interface, as opposed to FUSE, likely contributed to this superior performance for small I/O operations.
  • Scenario with Large Single File: In cases involving a large single file, Elink's performance was on par with Goofys. This indicates that Elink can handle large contiguous reads just as efficiently as a direct S3 filesystem, without introducing noticeable overhead.

The entire test environment was provided by Alibaba Cloud, ensuring a realistic cloud-native setup. These results strongly support the claim that Elink not only solves the packaging bottleneck but also provides competitive, and often superior, data access performance compared to existing solutions for large AI datasets.

Defensive Implications

▶ Watch: Illustrating the high cost of packaging large datasets (8:00)

While the primary focus of Elink is efficiency and performance, the speakers did touch upon some crucial security aspects, particularly during the Q&A session. It's clear that while foundational elements are in place, a comprehensive security posture for Elink would require further development and integration.

  1. Data Integrity and Versioning (Etag Validation): Elink heavily relies on Etags (Entity Tags) for data integrity. The Etag in the reference list acts as an MD5 checksum of the object, allowing the streaming service to verify that the remote data has not changed since the reference list was created. Before serving an IO request, the system can perform an HTTP HEAD request to the remote storage to compare the current Etag. If there's a mismatch, it indicates data tampering or an out-of-date reference, preventing the delivery of inconsistent data. However, this mechanism assumes the Etag itself is trustworthy and the backend storage is not compromised to serve malicious Etags.
  2. Authorization for Remote Data Access: The talk explicitly highlighted that the OCI registry might not possess the necessary access permissions for all remote objects in the backend storage. This necessitates an additional authorization layer at runtime. Defenders must ensure that the container runtime or the snapshotter component is properly authorized (e.g., via Kubernetes Service Accounts, IAM roles, or workload identity solutions) to access the specific backend storage buckets containing the data. Granular access control policies should be applied to these storage buckets to enforce the principle of least privilege, ensuring that containers can only access the datasets they are authorized for.
  3. Checksum Verification for Streaming Data: A question from the audience specifically addressed how checksums are verified when streaming a blob, especially if the entire data hasn't been received. The speakers confirmed that Etags are used for this. While an Etag can represent an MD5 checksum, the challenge of verifying integrity during a stream (before the entire file is downloaded) remains a complex topic. For full end-to-end data integrity, robust mechanisms like transport layer security (TLS) for data in transit and potentially cryptographic hashing of data chunks as they arrive would be ideal, though not explicitly detailed as part of Elink's current security features.
  4. Supply Chain Security (Image Signing): An important question was raised regarding compatibility with Sigstore image signing. The speakers acknowledged that Elink is still in "early stages" and security verification, especially for supply chain aspects like signing, has not been fully implemented. For production use, it would be critical to integrate Elink artifacts (specifically the reference list) into a secure software supply chain. This would involve signing the OCI artifact containing the reference list to cryptographically verify its origin and integrity, ensuring that the pointers to remote data have not been tampered with. Without such signing, a malicious actor could alter the reference list to point to compromised data sources.
  5. Auditability and Logging: As Elink introduces a new layer of data access, robust logging and auditing capabilities would be crucial for security monitoring. This includes logging access attempts to the remote storage, Etag mismatches, and authorization failures to detect and respond to potential security incidents.

In summary, while Elink offers significant performance benefits, its deployment requires careful consideration of the security implications. Defenders need to focus on strong access control for backend storage, leverage Etag validation for data integrity, and push for the integration of standard supply chain security practices like artifact signing as the project matures.

Key Takeaways

  • Addressing AI/ML Dataset Bottlenecks: Traditional OCI image packaging is inefficient and slow for large, dynamic AI/ML datasets in Kubernetes, leading to high costs and reduced agility.
  • Elink's Innovative Approach: Elink decouples data storage from OCI artifacts by packaging a lightweight "reference list" (metadata) instead of the actual data, leveraging OCI registries for metadata distribution and versioning.
  • Dramatic Packaging Performance Gains: The solution slashes packaging time from hours to minutes (e.g., 22GB dataset from ~4 hours to <2 minutes), enabling rapid iteration and deployment of AI models and datasets.
  • On-Demand Streaming with OverlayBD: Elink, integrated with OverlayBD, provides on-demand streaming of data directly from remote storage, significantly improving data access throughput, especially for workloads with many small files.
  • Leveraging Existing OCI Infrastructure: Elink utilizes the familiar and scalable OCI registry ecosystem for metadata management, reducing operational complexity and leveraging existing tooling.
  • Security Considerations are Ongoing: While Elink employs Etag validation for data integrity and acknowledges the need for external authorization, comprehensive security features like artifact signing and robust checksum verification for streaming data are areas for future development.

About the Speaker(s)

E. Rey is a Software Engineer at Microsoft, where he specifically works on the Azure Container Registry (ACR). His primary areas of expertise include OCI conformance, ensuring that container images and artifacts adhere to open standards, and artifact streaming, a technology crucial for optimizing container startup performance. His work contributes to the foundational infrastructure that powers cloud-native container workloads.

Y. Yuan is a Senior Software Engineer and Researcher at Alibaba Cloud. He is a prominent contributor to the OverlayBD project, a high-performance image and volume solution widely adopted in production environments. His extensive contributions to OverlayBD highlight his deep expertise in filesystem technologies, block device virtualization, and optimizing data access for cloud-native applications and large-scale data workloads.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk presents Elink, a clever solution that tackles the critical bottleneck of loading massive AI/ML datasets in Kubernetes. By transforming OCI artifacts from data containers into intelligent metadata pointers, Elink drastically reduces data packaging times and improves streaming performance, enabling faster, more cost-effective AI model training. The approach is technically sound, leverages existing infrastructure, and offers significant practical impact for cloud-native AI workflows.

Heather Calloway (CISO) — STRONG ACCEPT

Elink presents a compelling, technically sound solution to a critical bottleneck in enterprise AI/ML workflows, dramatically cutting data loading times and costs by decoupling data from OCI images. The innovative use of OCI registries for metadata distribution, rather than bulk data storage, provides significant operational efficiency and scalability. While the talk excels in technical demonstration and performance metrics, the comprehensive security implications and the necessary governance frameworks for this new paradigm require deeper exploration for enterprise adoption. This is a technology that will reshape how large organizations manage AI data, and CISOs need to understand its…

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025