Optimizing Model Serving on Kubernetes With Model Streaming - Ekin Karabulut & Ronen Dar, Run:ai
Ekin Karabulut, Ronen Dar, Run:ai
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In the realm of modern AI deployments, efficiently serving large language models (LLMs) and other deep learning models presents significant challenges, particularly within dynamic, cloud-native environments orchestrated by Kubernetes. This talk, delivered by Ronen Dar and Ekin Karabulut from Run:ai (now part of Nvidia), addresses a critical bottleneck: the "cold start problem" associated with loading massive model weights onto Graphics Processing Units (GPUs). As models grow exponentially in size, the time it takes to provision and prepare an inference replica can stretch into many minutes, leading to exorbitant operational costs, suboptimal GPU utilization, and poor user experience due to high latency.

Key moments
- 0:00 Introduction to model serving optimization and cold start problem
- 2:00 Understanding the AI inference cold start problem
- 3:20 Scale of model weight loading challenge with large models
- 4:00 Impact of slow model loading on various inference use cases
- 6:40 Detailed breakdown of the traditional sequential model loading
- 7:50 Essential requirements for an efficient model loading solution
- 8:50 Introduction to the open-source Run:ai Model Streamer
Optimizing Model Serving on Kubernetes With Model Streaming
Speakers: Ronen Dar, Co-founder & CTO, Run:ai; Ekin Karabulut, Developer Advocate, Run:ai
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=qH5djJlbodY
Overview
In the realm of modern AI deployments, efficiently serving large language models (LLMs) and other deep learning models presents significant challenges, particularly within dynamic, cloud-native environments orchestrated by Kubernetes. This talk, delivered by Ronen Dar and Ekin Karabulut from Run:ai (now part of Nvidia), addresses a critical bottleneck: the "cold start problem" associated with loading massive model weights onto Graphics Processing Units (GPUs). As models grow exponentially in size, the time it takes to provision and prepare an inference replica can stretch into many minutes, leading to exorbitant operational costs, suboptimal GPU utilization, and poor user experience due to high latency.
The presentation introduces the Run:ai Model Streamer, an open-source project designed to dramatically accelerate the process of transferring model weights from various storage locations directly into GPU memory. By employing concurrent reading and streaming techniques, the streamer aims to mitigate the cold start problem, enabling more agile scaling of AI inference services. This innovation is crucial for organizations looking to optimize their AI infrastructure, reduce cloud expenditure, and enhance the responsiveness of their machine learning applications across diverse use cases, from real-time interactive services to large-scale batch processing.
The work presented in this talk is a testament to the efforts of Noah and Omar, who were instrumental in developing the Run:ai Model Streamer. Ronen Dar, co-founder and CTO of Run:ai, brought his extensive background in information theory and AI infrastructure orchestration, while Ekin Karabulut, a developer advocate with a master's in robotics, cognition, intelligence, provided insights into the practical aspects of model deployment. Their collective expertise underscores the importance of addressing infrastructure-level challenges to unlock the full potential of AI.
Background
▶ Watch: Introduction to model serving optimization and cold start problem (0:00)
Traditional web applications running on Kubernetes often leverage autoscalers to dynamically adjust the number of service replicas based on traffic demand. When user queries increase, new replicas are spun up on CPU-based instances; when traffic subsides, these replicas are scaled down to conserve resources and reduce costs. This elastic scaling model is a cornerstone of cloud-native efficiency. However, applying this paradigm to AI inference workloads, particularly those involving large deep learning models on GPUs, introduces significant complexities, primarily due to the "cold start problem."
The cold start problem in AI inference refers to the substantial delay incurred when a new model serving replica needs to be brought online. This delay is a composite of several factors:
- GPU Machine Provisioning: Acquiring and provisioning GPU-equipped machines in the cloud typically takes longer than CPU instances, owing to the scarcer availability of GPUs and the need to install specialized software libraries like CUDA drivers and CUDA kernels.
- Container Image Loading: AI inference workloads often rely on large container images, frequently containing extensive Python libraries and frameworks, which take time to download and load.
- Inference Engine Boot-up: While generally less significant than other factors, the initialization of the inference engine itself can add to the start-up latency.
- Model Weight Loading: Crucially, downloading and loading the model weights from storage into GPU memory is often the most time-consuming step. The sheer size of modern foundation models exemplifies this challenge:
- Llama 3 8 billion parameters (8B), typically deployed with 16 bits per weight, requires approximately 15 gigabytes (GB) of memory.
- Llama 3 70B parameters can exceed 100 GB.
- Cutting-edge models like Deepseek R1 can reach over 1 terabyte (TB) in size.
Loading such massive weights can take "many minutes," sometimes "more than 10 minutes," significantly hindering responsiveness.
This protracted cold start directly impacts several critical AI inference use cases:
- Real-time Inference (High Load, Single Model): For applications requiring low latency responses to frequent queries against a single model, the inability to quickly scale up new replicas leads to overprovisioning. To maintain responsiveness, organizations often keep more GPU instances running than immediately necessary, resulting in high costs and low GPU utilization. Users cannot tolerate waiting minutes for their requests to be processed.
- Cold Models / Multiple Models (Infrequent Load): In scenarios where numerous models need to be available, but each is queried infrequently by different users, the ideal solution would be to scale to zero – storing models and only spinning up a replica when a request arrives. However, if this spin-up takes many minutes, it becomes impractical for latency-sensitive applications. Consequently, models are often kept "warm" on dedicated GPUs, again leading to overprovisioning, high costs, and underutilized hardware.
- Offline Inference (Batch Jobs): Even for batch processing jobs that process large datasets in a non-real-time fashion, the time spent provisioning GPUs and loading models contributes directly to the overall job duration and, thus, the operational cost. Reducing this initial overhead can yield substantial savings.
The traditional model loading process is inherently sequential: model weights are first transferred from storage (e.g., local disk, object storage) to the CPU memory, where operations like sharding or quantization might occur. Only then are these weights transferred from CPU memory to the GPU memory. This sequential nature, coupled with the immense data volumes, is a primary reason for the extended loading times.
Recognizing these limitations, the Run:ai team identified several key requirements for an optimized model loading solution:
- Parallelization: Overcome the sequential loading bottleneck by enabling concurrent operations.
- Storage Agnosticism: Support multiple storage types (e.g., local file systems, network file systems, S3, GCS) without requiring code or storage changes. Existing loaders often lack broad S3 support.
- SafeTensors Compatibility: Directly support the SafeTensors format, which is rapidly becoming the industry standard for model weights due to its safety and efficiency, avoiding the need for format conversions.
- Inference Engine Integration: Offer easy integration with various popular inference engines (e.g., VLM, TGI) to maintain flexibility for users.
Key Findings
▶ Watch: Scale of model weight loading challenge with large models (3:20)
The benchmarking efforts and design principles behind the Run:ai Model Streamer yielded several critical findings that underscore the efficacy of its approach and provide valuable insights for deploying AI inference at scale:
- Concurrency is a Primary Driver of Speed (Up to a Point): The core principle of concurrent reading and streaming of tensors from storage to GPU memory demonstrated significant improvements in model loading times. By issuing multiple reading requests simultaneously, the streamer can drastically cut down the initial waiting period. However, this acceleration is subject to the physical limitations of the underlying storage bandwidth. Once the storage bandwidth is fully saturated, increasing concurrency further yields no additional performance gains, highlighting the importance of balancing concurrency levels with hardware capabilities.
- Balanced Workload Distribution is Crucial for Bandwidth Saturation: Model tensors vary considerably in size, from megabytes to gigabytes. A naive concurrent loading approach might get bottlenecked by larger tensors, leaving smaller ones waiting. The Run:ai Model Streamer addresses this by dividing individual tensors into equal-sized chunks for reading, distributing the workload evenly among threads. This balanced workload distribution ensures optimal bandwidth saturation and prevents individual large tensors from creating bottlenecks, thereby accelerating the overall streaming process.
- Storage Bandwidth is a Critical Factor: The performance of the underlying storage system profoundly impacts model loading times. Deployments demanding rapid model access should prioritize investing in high-performance storage solutions. Benchmarking revealed substantial differences between various SSD types (e.g., GP3 vs. IO2 SSDs on AWS) and between local storage and cloud object storage like S3. Higher throughput storage directly translates to faster cold start times, particularly vital in on-premise and hybrid cloud environments where storage performance can be more variable.
- Tunable Parameters for Optimal Performance: The Model Streamer provides adjustable parameters, including the level of concurrency and CPU memory usage. It is essential for users to tune these parameters based on their specific storage type, network conditions, and available CPU memory. For instance, an optimal concurrency level for an S3 bucket might differ significantly from that for a local NVMe SSD. Proper tuning can lead to substantial reductions in cold start times tailored to the deployment environment.
- Exceptional Performance with AWS S3: The Run:ai Model Streamer demonstrated remarkable results with Amazon S3, achieving model load times "under five seconds" for a 15GB Llama 3 8B model. This superior performance is attributed to its optimized approach: creating an AWS S3 client per thread, with each thread sending multiple asynchronous requests to the S3 backend. This strategy effectively maximizes parallelism and throughput when interacting with S3's distributed architecture.
- Practical vs. Theoretical Cloud Throughput: Benchmarking in cloud environments revealed a discrepancy between documented theoretical storage throughput and actual observed performance. For instance, SSDs advertised at 4 gigabytes per second (GB/s) might only achieve up to 2 GB/s in practice. This highlights the importance of real-world testing and planning for potential practical limits when designing AI infrastructure.
- S3 Caching Effects Require Careful Benchmarking: During benchmarking, it was observed that the S3 backend might exhibit caching behavior, leading to faster subsequent load times after an initial run. To accurately measure cold start performance, it is crucial to implement a cooldown period between cloud tests to avoid skewed results due to S3 caching.
- Promising GCS Integration Results: Early impressions from integration efforts with the Google Kubernetes Engine (GKE) team for Google Cloud Storage (GCS) showed "96% model load time reduction" when using the Model Streamer with VLM, compared to direct downloading from cloud object storage. This indicates strong potential for similar performance benefits across different cloud providers.
Technical Deep Dive
▶ Watch: Impact of slow model loading on various inference use cases (4:00)
The Run:ai Model Streamer is an innovative solution engineered to overcome the bottleneck of model weight loading by leveraging concurrent data transfer. Implemented as a Python SDK with a C++ backend, it is designed for high performance and seamless integration into existing AI workflows.
At its core, the Model Streamer employs two key mechanisms for acceleration:
- Concurrent Reading of Tensors: Instead of reading model weights sequentially, the streamer initiates multiple parallel read requests to the storage system. This allows for fetching different parts of the model (or different tensors) simultaneously, significantly reducing the overall data retrieval time.
- Concurrent Streaming to GPU: As data chunks are read from storage, they are immediately streamed to the GPU memory. This overlapping of I/O operations (reading from storage) and compute operations (transferring to GPU) maximizes efficiency, ensuring the GPU is fed data as quickly as possible. This contrasts with the traditional sequential approach where the entire model is first loaded into CPU memory before being transferred to the GPU.
To optimize bandwidth saturation, the streamer incorporates a sophisticated balanced workload distribution strategy. AI models consist of numerous tensors, which can vary wildly in size. A simple concurrent approach might lead to imbalances where threads processing large tensors become bottlenecks. The Model Streamer intelligently divides these tensors into smaller, equal-sized parts. This ensures that each reading thread has a roughly equivalent amount of data to process, preventing any single thread from holding up the entire loading process and maximizing the utilization of the available storage bandwidth.
The streamer is highly configurable through several adjustable parameters, allowing users to fine-tune its behavior for specific environments:
- Concurrency Level: Users can specify the number of parallel reading requests to match the capabilities of their storage system and network.
- Data Chunk Size: The size of the individual parts into which tensors are divided can be adjusted, influencing the granularity of parallelization.
- CPU Memory Usage: For environments with limited or abundant CPU memory, users can control how much buffer space the streamer utilizes, balancing between throughput and resource consumption.
A crucial design decision was the native support for multiple storage types. The Run:ai Model Streamer seamlessly integrates with:
- Local file systems: Such as NVMe SSDs.
- Network file systems: Shared storage solutions.
- Cloud-based object storage: Including Amazon S3 and Google Cloud Storage (GCS).
This broad compatibility eliminates the need for users to modify their storage infrastructure or data pipelines, providing flexibility across diverse deployment scenarios.
The project prioritizes compatibility with the SafeTensors format. This modern, secure, and efficient format for storing model weights is directly supported, removing any requirement for format conversions, which could otherwise introduce additional overhead or potential data integrity issues.
Furthermore, the Model Streamer is designed as a SafeTensors iterator. This architectural choice makes it highly compatible and easy to integrate with popular inference engines such as VLM (vLLM) and TGI (Text Generation Inference). Instead of requiring deep modifications to these engines, the streamer can be plugged in as a drop-in replacement for their native model loading mechanisms. Notably, for VLM versions higher than 0.66, the Run:ai Model Streamer is now included out of the box, simplifying adoption for VLM users.
Benchmarking studies were conducted using a Meta Llama 8B model (15GB in a single SafeTensors file) on a single A10G GPU on AWS. Various storage types were tested, including local SSDs (GP3 and IO2, with IO2 having higher throughput) and Amazon S3 in the same region as the instance. The benchmarks compared standalone loaders (SafeTensors loader, Run:ai Model Streamer, and Tensorizer) as well as the combined engine boot and model load times with VLM. The results consistently demonstrated significant improvements in loading times with the Run:ai Model Streamer across all tested storage configurations.
Looking ahead, the roadmap for the Run:ai Model Streamer includes exciting enhancements such as support for sharded models, optimized multi-GPU model loading, parallel multi-file loading, and integration with GPU Direct Storage, further pushing the boundaries of efficient AI model serving.
Demo / Proof of Concept
▶ Watch: Essential requirements for an efficient model loading solution (7:50)
While the talk did not feature a live, interactive demonstration of the Run:ai Model Streamer, the speakers presented extensive benchmarking results that serve as a robust empirical proof of concept. These results, detailed in a dedicated white paper available via QR code during the presentation, unequivocally illustrate the performance benefits of the Model Streamer.
The benchmarking environment included a Meta Llama 8B model (a 15GB SafeTensors file), a single A10G GPU on AWS, and various storage configurations: local GP3 SSD, local IO2 SSD (known for higher throughput), and Amazon S3. The comparison focused on the time taken by different loaders (SafeTensors loader, Tensorizer, and Run:ai Model Streamer) to load the model from storage to GPU, as well as the combined engine boot and model loading time with VLM.
The presented data consistently showed the Run:ai Model Streamer significantly outperforming traditional methods, particularly highlighting:
- Sub-5-second load times for a 15GB model from S3, a remarkable achievement.
- A "96% model load time reduction" when integrated with VLM and Google Cloud Storage, as observed by the GKE team.
These quantitative results provide compelling evidence of the streamer's effectiveness in mitigating the cold start problem and optimizing model serving performance in real-world cloud environments. For those interested in replicating or further exploring these findings, the project's GitHub repository and the benchmarking white paper offer detailed insights and practical guidance.
Defensive Implications
▶ Watch: Introduction to the open-source Run:ai Model Streamer (8:50)
The Run:ai Model Streamer primarily focuses on performance optimization and resource efficiency for AI inference, rather than addressing traditional cybersecurity vulnerabilities. However, its capabilities have significant indirect defensive implications for platform engineers, MLOps teams, and security practitioners managing AI infrastructure:
- Cost Optimization and Resource Management: By drastically reducing model cold start times, the streamer enables more aggressive scaling-to-zero strategies and more efficient autoscaling. This directly translates to reduced cloud expenditure, as GPUs are provisioned only when needed and released faster. From a defensive standpoint, effective cost management prevents unexpected billing spikes, which can sometimes be an indicator of resource misuse or unauthorized activity. It also ensures that critical resources are available for legitimate workloads, preventing resource exhaustion that could be exploited in a denial-of-service (DoS) context.
- Enhanced Service Availability and Resilience: Faster model loading improves the overall availability and resilience of AI-powered applications. During periods of high demand, new replicas can come online much faster, preventing service degradation or outages due to insufficient capacity. Similarly, in the event of a pod crash or node failure, the ability to quickly restart and load models minimizes downtime. This directly contributes to the "availability" pillar of the CIA triad (Confidentiality, Integrity, Availability), making AI services more robust against operational disruptions.
- Faster Deployment of Security Patches and Model Updates: The ability to quickly spin up new model instances with updated weights or patched inference engines facilitates a more agile security posture. When vulnerabilities are discovered in underlying libraries (e.g., CUDA, Python dependencies) or in the model itself (e.g., bias, ethical concerns), faster deployment cycles mean that patched versions can be rolled out more rapidly, reducing the window of exposure. This agility is crucial in mitigating risks associated with emerging threats in the AI supply chain.
- Improved GPU Utilization and Reduced Attack Surface: Higher GPU utilization, a direct benefit of faster model loading and better scaling, means fewer overall GPU instances might be required to handle a given workload. A smaller fleet of infrastructure can simplify management, patching, and monitoring efforts, indirectly reducing the potential attack surface. Less overprovisioning also means fewer idle resources that could potentially be hijacked or misused.
- Operational Consistency and Reproducibility: The predictable and optimized model loading process contributes to greater operational consistency. By reducing variability in start-up times, MLOps teams can establish more reliable service level objectives (SLOs) and service level agreements (SLAs). While not directly a "security" measure, consistent operations are a foundation for effective security monitoring and anomaly detection, as deviations from expected behavior become more apparent.
In essence, while the Run:ai Model Streamer is not a security tool in the traditional sense, it empowers organizations to build more efficient, cost-effective, and resilient AI inference platforms. These operational improvements are foundational to a strong overall security posture, enabling faster responses to incidents, better resource governance, and more reliable service delivery.
Key Takeaways
- The cold start problem in AI inference, primarily driven by the time taken to load massive model weights into GPU memory, is a major impediment to efficient, cost-effective, and low-latency AI service deployment on Kubernetes.
- The Run:ai Model Streamer addresses this by implementing a novel approach of concurrently reading tensors from storage and streaming them to the GPU, significantly reducing model loading times.
- Storage bandwidth is a critical performance factor, and balanced workload distribution (dividing tensors into equal chunks for parallel processing) is essential for saturating this bandwidth and maximizing loading speed.
- The streamer is highly configurable with adjustable parameters for concurrency level and CPU memory usage, allowing for optimal tuning based on specific storage types and resource availability.
- With native support for SafeTensors format and integration as a SafeTensors iterator, it easily integrates with popular inference engines like VLM (now included out-of-the-box in newer versions) and supports various storage types, including exceptional performance with AWS S3 and promising results with Google Cloud Storage.
- Organizations should account for practical limits in cloud storage throughput (which may be lower than advertised) and S3 caching effects (requiring cooldown periods) when performing their own cold start benchmarking.
About the Speaker(s)
Ronen Dar is the Co-founder and CTO of Run:ai, a company he helped establish in 2018. Run:ai specializes in AI infrastructure orchestration, providing solutions to manage and optimize AI workloads. Following Run:ai's acquisition by Nvidia, Ronen and his team are now part of the Nvidia organization. Before founding Run:ai, Ronen pursued his PhD and post-doctoral research in the field of information theory, building a strong academic foundation that informs his work in complex systems like AI infrastructure.
Ekin Karabulut serves as a Developer Advocate at Run:ai, also now part of Nvidia. Ekin brings a practical and educational perspective to the team, helping developers understand and leverage Run:ai's technologies. He holds a master's degree in robotics, cognition, and intelligence from Munich, which provides him with a deep understanding of the intricacies of AI and machine learning systems.
Together, Ronen and Ekin presented the Run:ai Model Streamer, an open-source project that is the culmination of significant work by Noah and Omar, whose contributions were acknowledged during the talk. Their collective expertise aims to solve some of the most pressing infrastructure challenges facing the AI industry today.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk presents a highly relevant and well-engineered solution to a critical problem in modern AI deployments: the 'cold start' bottleneck when loading massive model weights onto GPUs. The Run:ai Model Streamer, an open-source project, leverages concurrent reading and direct streaming to GPU memory, demonstrating significant performance gains. While originating from a vendor, the technical depth, open-source contribution, and clear benchmarking make this a valuable session for anyone grappling with large language model serving at scale.
Heather Calloway (CISO) — STRONG ACCEPT
This talk delivers a highly relevant solution to a critical operational bottleneck in modern AI deployments: the "cold start problem" for large models on Kubernetes. By introducing the Run:ai Model Streamer, it directly addresses the exorbitant costs, poor GPU utilization, and latency issues that plague organizations attempting to scale AI inference. While not a cybersecurity talk, its focus on efficiency, resilience, and cost optimization provides foundational improvements that enable better resource governance, faster incident response through quicker patching, and enhanced availability of critical AI services, which are all indirect but significant defensive implications for any CISO.