Orchestrating AI Models in Kubernetes: Deploying Ollama as a Nati... Samuel Veloso & Lucas Fernández

Samuel Veloso, Lucas Fernández

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk, presented by Samuel Veloso and Lucas Fernández at KubeCon EU, delves into an innovative approach for deploying Artificial Intelligence (AI) models within Kubernetes environments. Specifically, it focuses on Ollama, a rapidly popular command-line interface (CLI) tool designed for running large language models (LLMs) and other AI models locally with remarkable simplicity. While Ollama excels in local deployments, the speakers address the challenge of scaling and orchestrating these models in a production-grade Kubernetes cluster in a truly native fashion, moving beyond traditional Helm chart deployments.

Watch on YouTube

Visual summary for Orchestrating AI Models in Kubernetes: Deploying Ollama as a Nati... Samuel Veloso & Lucas Fernández by Samuel Veloso, Lucas Fernández
Visual summary for Orchestrating AI Models in Kubernetes: Deploying Ollama as a Nati... Samuel Veloso & Lucas Fernández by Samuel Veloso, Lucas Fernández

Key moments

  1. 0:50 What is Olama and its local usage simplicity?
  2. 2:30 Olama's Docker-like UX and Kubernetes integration challenge
  3. 3:20 Leveraging runtimeClassName for custom Kubernetes container runtimes
  4. 5:00 How Kubernetes works: kubelet and container creation flow
  5. 8:00 Deep dive into CRI RunPodSandbox API for custom runtimes

Orchestrating AI Models in Kubernetes: Deploying Ollama as a Native Container Runtime

Speakers: Samuel Veloso, Software Engineer, CAST AI; Lucas Fernández, Red Hat, Kubeflow Contributor

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=zLpUJBU6sT4

Overview

This talk, presented by Samuel Veloso and Lucas Fernández at KubeCon EU, delves into an innovative approach for deploying Artificial Intelligence (AI) models within Kubernetes environments. Specifically, it focuses on Ollama, a rapidly popular command-line interface (CLI) tool designed for running large language models (LLMs) and other AI models locally with remarkable simplicity. While Ollama excels in local deployments, the speakers address the challenge of scaling and orchestrating these models in a production-grade Kubernetes cluster in a truly native fashion, moving beyond traditional Helm chart deployments.

The core of their presentation revolves around leveraging Kubernetes' RuntimeClass feature to implement a custom container runtime, dubbed ollama-shim. This custom shim allows Kubernetes to treat AI models as first-class citizens, enabling their direct deployment as if they were standard containers. This method provides enhanced control, integration, and a more Kubernetes-idiomatic way to manage AI workloads. The talk meticulously explains the internal workings of Kubernetes' Container Runtime Interface (CRI) and demonstrates how to build and integrate such a custom runtime, culminating in a proof-of-concept integration with Kubeflow, a leading MLOps platform.

The significance of this work lies in bridging the gap between the burgeoning local AI model ecosystem and enterprise-scale cloud-native infrastructure. By enabling native Ollama deployments in Kubernetes, organizations can harness the power of LLMs on their own infrastructure, maintaining data privacy, leveraging existing Kubernetes tooling for orchestration, and potentially reducing operational overhead compared to managing external AI services. This approach paves the way for more efficient and secure MLOps pipelines, directly addressing the growing demand for scalable and integrated AI solutions in cloud-native environments.

Background

▶ Watch: What is Olama and its local usage simplicity? (0:50)

The recent explosion in AI, particularly with large language models, has driven the creation of tools that simplify local model deployment. Ollama stands out as a prime example, offering a user experience strikingly similar to Docker for containers, but for AI models. With a simple ollama pull and ollama run command, users can interact with powerful models like Llama 3 on their local machines, avoiding the need to send sensitive data to third-party cloud providers. This simplicity has fueled its immense popularity, evidenced by its GitHub repository boasting over 134,000 stars, surpassing even Kubernetes itself in a relatively short period.

While Ollama's local utility is undeniable, integrating these models into a scalable, production-ready environment like Kubernetes presents a challenge. Traditional methods often involve deploying Ollama as a standard application via a Helm chart, which, while functional, doesn't fully embrace Kubernetes' native capabilities for specialized workloads. The speakers sought a "more Kubernetes-native way," leading them to explore the RuntimeClass mechanism.

The RuntimeClass feature, introduced in Kubernetes 1.20, provides a powerful abstraction layer allowing users to select different container runtimes for their pods. By default, Kubernetes uses runtimes like containerd or CRI-O to create standard runc containers. However, RuntimeClass enables the integration of custom runtimes for specialized needs. Examples include Kata Containers and gVisor, which create isolated sandboxes or virtual machines for enhanced security, or runtimes designed for WebAssembly workloads. This talk proposes extending this concept to AI models, specifically Ollama, by developing a custom shim.

To understand how this custom shim integrates, it's crucial to grasp a simplified view of Kubernetes' internal workings, particularly how pods and containers are created. When a user applies a kubectl apply -f command for a deployment with a specified runtimeClassName, the Kubernetes API server receives the request and stores the pod definition. The kubelet, running on each node, continuously monitors the API server for new or changed pods. Upon detecting a new pod, the kubelet invokes its syncPod function.

Within syncPod, the kubelet communicates with the CRI runtime (e.g., containerd or CRI-O) via the Container Runtime Interface (CRI) API. This API consists of two main services: the Runtime Service and the Image Service. The Runtime Service is responsible for managing pod sandboxes and containers, while the Image Service handles image management. The first crucial step is the RunPodSandbox API call. This bootstraps the necessary cgroups, namespaces, and network configuration for all containers within a pod, initializing a dummy "post container." During this phase, the CRI runtime consults the runtimeClassName to determine which sandbox runtime shim (e.g., runc by default, or our custom ollama-shim) should handle the actual creation of the sandbox and its initial container. After network setup via the Container Network Interface (CNI), the shim is instructed to CreateTaskRequest, returning a process ID (PID) and sandbox ID to the kubelet.

Following sandbox creation, the CreateContainer API is invoked to add the actual application containers (or, in this case, the AI model) to the sandbox. This involves the kubelet requesting the CRI runtime to PullImage (if not already present), and then sending CreateContainer and StartContainer requests to the selected shim. The shim then uses these requests to create and start the specified container within the pre-existing pod sandbox. This intricate dance between kubelet, CRI runtime, and the specific shim is the foundation upon which the native Ollama deployment is built.

Key Findings

▶ Watch: Olama's Docker-like UX and Kubernetes integration challenge (2:30)

The primary finding of this research is the successful demonstration that Ollama AI models can be deployed as native Kubernetes container runtimes by leveraging the RuntimeClass mechanism. This fundamentally changes how AI models can be managed and scaled within cloud-native infrastructures.

Key contributions and results include:

  • Custom Shim Development: The creation of an ollama-shim compatible with containerd. This shim acts as a crucial intermediary, translating Kubernetes container creation requests into actions that launch and serve Ollama models.
  • OCI Compatibility for Models: A solution was devised to overcome Ollama's native model format incompatibility with the Open Container Initiative (OCI) standard. The GGUF packer go tool was identified and utilized to convert GGUF formatted AI models into OCI-compliant images, making them discoverable and pullable by standard container runtimes.
  • Dynamic Container Configuration Mutation: The ollama-shim was engineered to dynamically modify the container's config.json at runtime. This mutation enables the shim to mount the GGUF model file into the container's filesystem and then execute the necessary Ollama commands to load and serve the model, effectively transforming a generic container into an AI model server.
  • Seamless Kubeflow Integration: The developed native Ollama deployment was successfully integrated into Kubeflow, a prominent MLOps platform, demonstrating its potential for broader adoption within existing AI/ML workflows. This integration leverages Kubeflow's new modular architecture, allowing for rapid development of new platform components.

These findings collectively illustrate a powerful new paradigm for AI model orchestration in Kubernetes, offering a more integrated, efficient, and potentially secure way to manage the lifecycle of LLMs and other AI workloads.

Technical Deep Dive

▶ Watch: Leveraging runtimeClassName for custom Kubernetes container runtimes (3:20)

The technical implementation of orchestrating Ollama models natively in Kubernetes hinges on two main components: making AI models OCI-compatible and building a custom containerd shim.

The first challenge is that models from the Ollama registry are not directly compatible with the Open Container Initiative (OCI) image format, which is the standard for container images consumed by container runtimes like containerd. To address this, the speakers employed GGUF packer go, a tool from the GPU stack, designed to convert GGUF models into OCI images. The GGUF format itself is a significant development, created by Georgie Gerganov, the primary contributor to Llama CPP. It's an optimized binary format for storing and sharing AI models, with over 90,000 models available in this format on Hugging Face.

Using GGUF packer go is analogous to a docker build process. A custom builder is used, and the GGUF model is added via an ADD instruction, typically pulled from Hugging Face. For example, to convert a Q12 model, one would run a command similar to docker build -t myregistry/ollama-q12:latest .. This process creates an OCI-compliant image that can then be pushed to any Docker registry, making it available for Kubernetes deployments. When containerd receives a PullImage request for this OCI-packaged model, it pulls the image onto the node, making the model data accessible.

The second, and arguably more complex, component is the creation of the custom shim. For containerd, this involves three steps:

  1. Containerd Configuration: Registering the custom runtime as a plugin in containerd's configuration, specifying a runtime_type (e.g., ollama-runtime). Containerd allows defining multiple runtimes.
  2. Shim Binary Deployment: Copying the compiled ollama-shim binary to a well-known location on each Kubernetes node. Containerd will execute this binary whenever a pod sandbox or container needs to be created for the ollama-runtime class.
  3. Kubernetes RuntimeClass Object: Creating a RuntimeClass object in Kubernetes that maps the runtimeClassName used in pod deployments (e.g., ollama-runtime) to the custom shim binary and its configuration within containerd.

The ollama-shim itself is implemented in Go, leveraging the containerd Go library. The shim needs to register itself with a shim.Manager and, critically, implement the TTRPC task interface defined by containerd. This interface is extensive, but the speakers focused on the Create method, drawing inspiration from containerd-shim-runc-v2, which is containerd's default shim for runc containers.

The Create method receives a CreateTaskRequest containing essential parameters like ID, bundle, and rootfs. The bundle parameter is particularly important, as it's a string path on the host that points to a directory containing the container's configuration, most notably the config.json file. This config.json is a declarative definition of the container, generated by containerd from the pod spec, detailing aspects like user, mounts, and command.

The core ingenuity of the ollama-shim lies in its ability to mutate this config.json information. The shim performs the following critical actions within its Create method:

  1. Mount Model File: It adds an additional mount to the config.json. This mount takes the GGUF model file (which was part of the OCI image pulled earlier) from the host's image store and makes it available inside the container.
  2. Create Temporary Model File: Ollama, at the time of the talk, didn't directly support reading GGUF models from arbitrary paths. Instead, it expected to create a model from a specific "model file name." The shim addresses this by creating a temporary file (e.g., /tmp/model_file_name) within the container and writing the name of the GGUF model to it. This temporary file is also mounted into the container.
  3. Initialize Ollama: Finally, the shim ensures that the container's entrypoint or command executes the necessary Ollama commands. Specifically, it uses ollama create with the name specified in the temporary model_file_name to initialize and serve the model. This effectively tells Ollama inside the container to load the GGUF model it finds via the mounted paths.

The end-user deployment manifest then simply includes runtimeClassName: ollama-runtime and defines the container with the OCI-packaged model image. The container's environment variables or arguments can specify the MODEL_NAME (e.g., Q12) and MODEL_PATH (if multiple models are in the image), and the port to expose (e.g., 8080). This configuration allows Kubernetes to delegate the container creation to the ollama-shim, which then performs the necessary steps to bring the AI model online.

Demo / Proof of Concept

▶ Watch: How Kubernetes works: kubelet and container creation flow (5:00)

The talk included a compelling series of demonstrations showcasing the practical application and integration of the native Ollama deployment with Kubeflow, a powerful MLOps platform. The speakers aimed to make Ollama a native part of Kubeflow's model catalog, leveraging its new modular architecture.

The demo progression highlighted three distinct development environments:

  1. Standalone Mocked Mode: This initial environment focused on rapid UI development for the new Kubeflow component. The speakers demonstrated a React-based frontend displaying a model catalog featuring models like Llama 3 and Q12. A chat interface was present, but interactions yielded "mock responses." This environment allows frontend contributors to design and iterate on the user interface without needing to understand the underlying Kubernetes complexities or have a running cluster. It showcases Kubeflow's modular architecture, where a new component can be developed in isolation, with the backend interactions simulated via a Swagger specification.
  1. Standalone Mode with Cluster: The second phase elevated the demonstration to a real Kubernetes environment. A Kind cluster was provisioned, configured with the ollama-shim and the necessary RuntimeClass object. An OCI-packaged Ollama model (specifically the Q12 model) was deployed using a standard Kubernetes deployment manifest, specifying runtimeClassName: ollama-runtime and exposing the model via a service on port 8080. The UI, now running in a standalone mode but connected to the actual cluster, was able to interact with the live model. Lucas demonstrated this by asking "Why is KubeCon?" to the Q12 model, and the chat interface displayed a real-time inference response. This part of the demo clearly showed the ollama-shim in action, illustrating how the Go service in the backend communicated with the Ollama service's generate endpoint to facilitate the model inference.
  1. Kubeflow Integration: The final and most significant part of the demo showcased the full integration of the Ollama component into a running Kubeflow cluster. Despite the challenges of installing Kubeflow (which typically requires substantial resources like 16GB RAM and multiple CPU cores), the speakers successfully deployed their custom Ollama UI and backend component within the Kubeflow central dashboard. The audience saw the "Ollama UI" listed as a new component alongside other Kubeflow services, such as the Model Registry (which recently reached General Availability with Kubeflow 1.10). This proved that the modular architecture of Kubeflow allowed for the rapid development and integration of new components – in this case, a new model catalog for Ollama models – in less than a week. While presented as a Proof of Concept (PoC) and a potential Pull Request (PR) for the upcoming Kubeflow 1.11 release, it powerfully illustrated the feasibility and benefits of extending existing MLOps platforms with custom AI runtimes.

Defensive Implications

▶ Watch: Deep dive into CRI RunPodSandbox API for custom runtimes (8:00)

The deployment of AI models using custom container runtimes introduces several security considerations that defenders must address. While this approach offers flexibility and efficiency, it also expands the attack surface and modifies the traditional trust boundaries within a Kubernetes cluster.

  1. Custom Runtime Trust Boundary: The ollama-shim operates with elevated privileges on the Kubernetes node, similar to how runc operates. Any vulnerability in the shim's codebase (e.g., buffer overflows, path traversal, improper handling of config.json mutations) could lead to a container escape or direct host compromise. Rigorous security audits, fuzz testing, and adherence to secure coding practices are paramount for the shim's development. Deploying the shim itself within a more secure sandbox (e.g., using a gVisor or Kata Containers runtime for the shim process, if technically feasible) could offer an additional layer of isolation.
  1. Model Supply Chain Security: Packaging GGUF models into OCI images introduces a new element into the software supply chain.
  • GGUF packer go Tool Security: The tool used for conversion (e.g., GGUF packer go) must be trusted and secure. A compromised packer could inject malicious code or configurations into the OCI image.
  • Base Image Vulnerabilities: The base image used for the OCI-packaged model could contain known vulnerabilities (CVEs). Regular scanning of these images using tools like Trivy or Clair is essential.
  • Malicious GGUF Models: While GGUF is a binary format, a maliciously crafted model could potentially exploit parser vulnerabilities in Ollama or the shim itself during loading, leading to denial-of-service or arbitrary code execution. Mechanisms for scanning or validating GGUF models before deployment are crucial.
  1. Isolation and Resource Management: While containers provide a degree of isolation, the custom shim directly manipulates container configuration. Ensuring proper isolation of the Ollama model container, especially regarding resource limits (CPU, memory, GPU), is critical to prevent resource exhaustion attacks or noisy neighbor issues. The cgroups configuration within the config.json must be correctly applied and enforced by the shim.
  1. API Exposure and Network Security: Ollama exposes an HTTP API for inference.
  • Network Policies: Implement strict Kubernetes Network Policies to control ingress and egress traffic to the Ollama model services, allowing only authorized clients (e.g., other microservices, Kubeflow components) to access them.
  • Authentication and Authorization: The Ollama API itself might not have robust authentication. Consider placing an API gateway or an authentication proxy in front of the Ollama service to enforce access control (e.g., JWT validation, OAuth2).
  • TLS Encryption: Ensure all communication with the Ollama API is encrypted using TLS.
  1. Configuration Management and RBAC:
  • RuntimeClass Management: Access to create or modify RuntimeClass objects should be tightly controlled via Role-Based Access Control (RBAC), as these objects can dictate how containers are run across the cluster.
  • Containerd Configuration: Modifications to containerd configuration (to register the custom shim) should be managed as infrastructure-as-code and protected from unauthorized changes.
  1. Observability and Auditing:
  • Logging and Monitoring: Comprehensive logging from the ollama-shim and the Ollama model containers is vital. Monitor for unusual process activity, unexpected network connections, or excessive resource consumption. Integrate these logs with a centralized security information and event management (SIEM) system.
  • Audit Logs: Kubernetes audit logs should be configured to track who creates, modifies, or deletes RuntimeClass objects and deployments using them.

By meticulously addressing these defensive implications, organizations can harness the benefits of native AI model orchestration in Kubernetes while maintaining a strong security posture.

Key Takeaways

  • Ollama's popularity underscores a strong demand for simplified local AI model deployment, driven by ease of use and the ability to keep data on-premises, similar to Docker's impact on containerization.
  • Kubernetes' RuntimeClass provides a powerful, native mechanism for integrating custom container runtimes, enabling specialized workloads like AI models to be treated as first-class citizens within the orchestration platform.
  • Achieving native Ollama deployment in Kubernetes requires two key technical innovations: converting GGUF models into OCI-compliant images (e.g., using GGUF packer go) and developing a custom containerd shim.
  • The ollama-shim dynamically mutates container configuration (specifically config.json) to mount GGUF model files and execute Ollama commands for loading and serving models, effectively transforming a generic container into an AI model server.
  • This native approach facilitates seamless integration with MLOps platforms like Kubeflow, allowing rapid development of new components (e.g., an Ollama model catalog) and leveraging Kubeflow's modular architecture for enhanced AI workflow management.
  • Custom container runtimes introduce new security considerations, necessitating rigorous security audits of the shim, secure supply chain practices for OCI-packaged models, strict network policies, and robust observability to protect the Kubernetes cluster from potential threats.

About the Speaker(s)

Samuel Veloso is a Software Engineer at CAST AI. His work focuses on the security team, where he contributes to building products designed to identify and remediate vulnerabilities and anomalies within Kubernetes environments. His expertise spans Kubernetes internals and security, informing the technical depth of the Ollama native deployment solution.

Lucas Fernández works at Red Hat, contributing to the Red Hat AI platform. He is also an active contributor to Kubeflow, a prominent open-source machine learning platform within the CNCF foundation. His background in MLOps and Kubeflow integration was instrumental in demonstrating the practical application of the custom runtime approach within an enterprise AI ecosystem.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This talk presents a genuinely novel and deeply technical approach to integrating AI models, specifically Ollama, directly into Kubernetes as first-class citizens. By leveraging RuntimeClass and a custom containerd shim, the speakers have engineered an elegant solution that bypasses traditional application-layer deployments, treating models as native container runtimes. This work addresses a critical gap in MLOps, enabling scalable, private, and Kubernetes-idiomatic orchestration of LLMs, making it a foundational piece of research for anyone serious about AI infrastructure.

Heather Calloway (CISO) — STRONG ACCEPT

This talk presents a highly technical but profoundly impactful approach to deploying AI models natively within Kubernetes using custom container runtimes. The ability to integrate Ollama via a RuntimeClass and a custom shim fundamentally changes the operational landscape for MLOps, offering clear benefits in data privacy, scalability, and leveraging existing cloud-native tooling. While the session itself is deep in implementation details, the implications for enterprise security architecture, supply chain risk, and the shifting boundaries of accountability are immediate and demand executive attention. This work is critical for CISOs to understand as their organizations increasingly adopt…

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025