The Explorer's Guide To Cloud Native GenAI Platform Engineering - Max Körbächer & Alexa Griffith
Max Körbächer, Alexa Griffith
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
This talk, "The Explorer's Guide To Cloud Native GenAI Platform Engineering," presented by Max Körbächer and Alexa Griffith, a Senior Software Engineer at Bloomberg, provides a comprehensive journey through the evolving landscape of Generative AI (GenAI) within the cloud-native ecosystem. The speakers highlight the explosive growth and rapid changes in the AI domain, emphasizing the critical role of platform engineering in enabling organizations to adopt and manage GenAI workloads effectively on Kubernetes. The core premise is that while individual tools and large-scale platforms exist, the journey of integrating these components into a flexible, scalable, and user-friendly GenAI platform often remains obscure.

Key moments
- 0:00 Introduction to the GenAI platform engineering journey
- 0:45 Analyzing the evolving Cloud Native AI landscape map
- 2:20 Kubernetes as a foundation for GenAI innovation
- 3:30 Challenges of complex, interconnected GenAI DevOps cycles
- 4:10 Addressing the core GenAI workload enablement question
- 5:00 Introducing the Thinnest Viable Platform (TVP) architecture
- 6:10 Overview of TVP setup with Canoe and Argo CD
- 7:55 Kserve for serving inference in a Minimal Viable Platform
The Explorer's Guide To Cloud Native GenAI Platform Engineering
Speakers: Max Körbächer; Alexa Griffith, Senior Software Engineer, Bloomberg
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=8ta_zFiUG1s
Overview
This talk, "The Explorer's Guide To Cloud Native GenAI Platform Engineering," presented by Max Körbächer and Alexa Griffith, a Senior Software Engineer at Bloomberg, provides a comprehensive journey through the evolving landscape of Generative AI (GenAI) within the cloud-native ecosystem. The speakers highlight the explosive growth and rapid changes in the AI domain, emphasizing the critical role of platform engineering in enabling organizations to adopt and manage GenAI workloads effectively on Kubernetes. The core premise is that while individual tools and large-scale platforms exist, the journey of integrating these components into a flexible, scalable, and user-friendly GenAI platform often remains obscure.
The session aims to bridge this gap by outlining a step-by-step approach, starting from a "Thinnest Viable Platform" (TVP) and progressively adding advanced capabilities. It addresses the inherent complexity of GenAI development, particularly for roles like data scientists who are specialists in models, not infrastructure. By demonstrating how Kubernetes and platform engineering principles can offload these complexities, the talk offers practical insights into building an environment that fosters innovation, accelerates deployment of AI models, and provides self-service capabilities for various teams.
Ultimately, this talk is crucial for anyone navigating the intersection of cloud-native technologies and artificial intelligence. It offers a blueprint for platform teams to empower their organizations with GenAI, ensuring adaptability to market changes, optimizing performance, managing costs, and establishing robust observability. The detailed exploration of open-source tools like Kserve, Envoy AI Gateway, and advanced caching techniques makes it a valuable resource for engineers and architects looking to build resilient and efficient GenAI platforms.
Background
▶ Watch: Introduction to the GenAI platform engineering journey (0:00)
The landscape of AI, particularly Generative AI, within the cloud-native ecosystem has undergone a dramatic transformation in recent years. Max Körbächer commenced the discussion by illustrating this rapid evolution, referencing the CNCF AI landscape map. He noted that the 2023 version was already "insane" in its breadth, but the 2024 iteration showed an "exploded" view, reflecting the immense proliferation of tools, frameworks, and solutions. This rapid expansion, while indicative of innovation, also presents a significant challenge: how do organizations keep pace and integrate these fast-changing components into stable, performant, and manageable systems?
The traditional "you build it, you run it" DevOps model, while effective for many applications, buckles under the weight of GenAI workloads. The speakers introduced the concept of "infinite DevOps cycles" for GenAI, which extends beyond merely building and running code. It now encompasses not only data management (finding and maintaining data) but also enabling others to use, integrate, and deploy complex AI applications, especially Large Language Models (LLMs). This multi-faceted cycle dramatically complicates workflows and tool integration. Data scientists, who are experts in model development and experimentation, are not necessarily infrastructure specialists. Expecting them to navigate the intricacies of Kubernetes deployments, networking, scaling, and observability for GenAI models is unrealistic and inefficient.
This challenge underscores the necessity for robust platform engineering. Kubernetes, with its inherent automation capabilities, has emerged as the ideal foundation. The speakers highlighted that Kubernetes has historically adapted to various technology hypes, from general automation in 2020 to edge computing and IoT, and now, prominently, AI. Platform engineering, therefore, provides the necessary flexibility to adopt new tools and replace outdated ones, offering a stable yet agile environment for GenAI innovation. The goal is to offload the complex environmental setup and integrations from domain specialists, allowing them to focus on their core expertise, thereby simplifying the life of the end-user. This led to the introduction of the "Thinnest Viable Platform" (TVP) concept, a minimal yet functional platform designed to get GenAI capabilities up and running quickly.
Key Findings
▶ Watch: Kubernetes as a foundation for GenAI innovation (2:20)
The talk presented several key findings and architectural components crucial for building a cloud-native GenAI platform:
- Kserve for Streamlined Inference: Kserve was identified as a foundational open-source tool that significantly simplifies the deployment and serving of AI models, including GenAI, on Kubernetes. It abstracts away Kubernetes complexities, allowing developers to focus on models while providing out-of-the-box features like autoscaling, traffic management, and new GenAI-specific optimizations such as token-based autoscaling and prompt/model caching.
- Envoy AI Gateway for Centralized LLM Access: A major contribution is the Envoy AI Gateway, a new open-source project developed in collaboration between Tetrate and Bloomberg. This gateway acts as a centralized access point for diverse LLM providers (self-trained, open-source, commercial), offering a unified API, credential management, and crucial token-based cost monitoring capabilities. It simplifies client interaction in a hybrid LLM environment.
- Enhanced Observability for LLMs: The speakers emphasized the need for deep insights into LLM behavior. While OpenTelemetry provides a broad view, the emerging Open LLM Telemetry (referred to as Open Metri in the talk) project specifically addresses the nuances of LLM observability, capturing model-specific metrics like costs, request durations, and enabling guardrail implementation.
- Advanced Performance Optimization Techniques: To tackle the challenges of large model sizes and unpredictable GenAI inference latency, the talk introduced several advanced optimization strategies:
- Model Caching: Reduces model download times and enables faster autoscaling for large models.
- Prompt Caching (KV Cache & LM Cache): Optimizes recurrent computations and reduces time to first token by caching common input prefixes across VLM instances.
- Disaggregated Serving: Separates compute-bound prefill and memory-bound decoding steps of LLM inference to improve throughput and hardware utilization.
- Platform Engineering as an Enabler: The overarching finding is that platform engineering is not just about infrastructure but about enablement. By providing self-service portals (like Backstage) and automating complex deployments, platform teams empower data scientists and developers to autonomously integrate and utilize GenAI models without deep Kubernetes expertise, fostering rapid innovation.
Technical Deep Dive
▶ Watch: Addressing the core GenAI workload enablement question (4:10)
The technical deep dive commenced with the concept of a Thinnest Viable Platform (TVP), designed for rapid deployment of a foundational GenAI environment. The speakers highlighted that even a "thin" platform can be quite substantial. They mentioned open-source projects like Canoe for getting things up and running quickly, integrating with cloud providers like AWS. While specific corporate environments might require tweaking (e.g., Kubernetes integration), Canoe provides a robust starting point. Other tools like Crossplane and Argo Workflows can simplify infrastructure management, and Backstage can serve as an intuitive entry point for developers to spin up services. The initial setup, as described, takes approximately 15 minutes to deploy all necessary applications, including Argo CD for GitOps, Cert Manager, and External DNS, making services reachable quickly.
A cornerstone of this platform is Kserve, an open-source tool specifically designed to simplify AI model development and serving on Kubernetes. Kserve excels by abstracting away the underlying Kubernetes complexities, allowing developers to concentrate on their models. It supports a wide array of GenAI and ML frameworks, making it highly versatile and framework-agnostic. Notably, Kserve now offers a unified API with OpenAI protocol support, enhancing its interoperability. Its out-of-the-box features for predictive inference are extensive, including scaling up and down from zero, request batching, security, distributed tracing, autoscaling on both GPU and CPU, comprehensive logging, observability, and traffic management.
For GenAI workloads, Kserve has introduced several critical enhancements:
- Adaptive scaling with built-in token-based autoscaling for GenAI models.
- Performance boosts through model cache and prompt cache, which minimize latency by allowing quick reuse of already downloaded models.
- Scalable inference, supporting the deployment of larger models across multiple nodes, particularly for VLM (Vision-Language Models), to achieve high-performance workloads.
- Seamless integration with the AI Gateway project.
The talk demonstrated how Kserve, combined with Backstage, enables a self-service model. Users can select a template in Backstage, fill out basic information (owner, namespace, model specifications), and automatically deploy an inference service or even provision resources like a new tenant in a Milvus vector database or an S3 bucket. The simplicity of Kserve's configuration, often just a small YAML file, was highlighted as a key advantage.
Next, the speakers introduced the Envoy AI Gateway, a new open-source tool born from the Envoy Gateway project, developed in collaboration between Tetrate and Bloomberg engineers. This gateway addresses the platform team's need to better serve AI workloads, especially in managing and providing LLMs as a service. Its MVP features include:
- Centralized access: Provides a standard, controlled, and auditable way to access self-trained, open-source, or commercial models. This is crucial as different LLM providers often have disparate access patterns; Envoy AI Gateway presents a unified API to clients and handles the underlying routing.
- Credential management: Simplifies the secure management of credentials across various LLM providers and on-premise deployments.
- Cost monitoring: Offers out-of-the-box, GenAI-specific cost monitoring, including token-based limits and cost optimizations.
An architectural diagram illustrated how a client request hits a load balancer, then the Envoy AI Gateway, which intelligently routes it to the correct provider, whether a managed inference Kubernetes cluster (e.g., Kserve) or an external LLM provider like AWS Bedrock. This setup enables a hybrid environment with seamless integration.
Beyond basic serving, the talk delved into advanced features for optimizing LLMs, grouped under "Expedition Mode." These include intelligent load balancing and advanced observability. For observability, the speakers emphasized the need for deeper insights into how models communicate, especially in chained GenAI applications. While OpenTelemetry provides a broad picture, it doesn't fully capture LLM specifics. This gap is filled by Open LLM Telemetry (referred to as "Open Metri" in the transcript), which integrates with various models (e.g., Claude, GPT) and platforms (e.g., Bedrock, Pinecone). It extracts LLM-specific information and forwards it to OpenTelemetry, which then sends it to any preferred observability endpoint. This allows monitoring of costs, request rates, average durations, and implementation of guardrails (e.g., detecting toxic outputs or data leaks). The speakers provided an example of prompt caching helping save around $9 in a simple demo.
Finally, the talk covered advanced optimization techniques for large-scale enterprise Kserve deployments:
- Model Caching: Essential due to the increasing size of LLMs (e.g., a Deepseek model can be around 1 terabyte, exceeding the 640 gigabytes of storage on an H100 node). Model caching reduces the time it takes to download a model when a pod scales up, which can be a significant bottleneck (saving up to 12 minutes in the example given). This also enables faster autoscaling.
- Multi-node inference serving: Necessary for models that don't fit on a single node.
- Prompt Caching: While KV cache (Key-Value cache) comes out-of-the-box with VLM, its growth can be exponential with sequence length. The LM Cache is an open-source system designed to manage KV caches efficiently. It caches common input prefixes, reducing redundant computation, speeding up repeated queries, and allowing caches to be shared and stored across multiple VLM instances. This has been shown to reduce the time to first token by 3 to 10% and save GPU cycles.
- Disaggregated Serving: Addresses the unpredictable latency of GenAI inference, which involves two distinct steps: prefill (compute-bound) and decoding (memory-bound). By separating these steps and executing them independently on different sets of GPUs, throughput is increased, and hardware performance is optimized.
Demo / Proof of Concept
▶ Watch: Introducing the Thinnest Viable Platform (TVP) architecture (5:00)
The talk included several demonstrations, though some encountered live technical difficulties, the descriptions provided valuable insights into the platform's capabilities.
Max Körbächer's initial demo focused on the Thinnest Viable Platform (TVP) setup. Although the live video did not play, he described a Cenoa deployment that rapidly provisions the necessary infrastructure. The demo would have shown that after a few minutes, the entire platform, including Argo CD for GitOps, Backstage for developer self-service, Cert Manager, and External DNS, would be fully deployed. This setup automatically integrates with load balancers and configures routes, making the initial services reachable within approximately 15 minutes. This highlights the platform's ability to quickly establish a functional environment for GenAI workloads.
Alexa Griffith presented a demo showcasing the Envoy AI Gateway in action. The scenario involved an AI agent calling a tool to fetch weather information. The demonstration revealed the internal configuration of the Envoy Gateway, specifically showing two backends: a Kserve LLM backend named dspllama (representing an on-premise, self-hosted service) and an Envoy basic AWS backend pointing to eu.anthropic (for a managed LLM provider like AWS Bedrock).
The demo illustrated a user port-forwarding to the Envoy service. The AI agent then successfully hit both the on-premise DSP Llama service and AWS Bedrock (using Claude Anthropic), receiving slightly different but relevant results. A key takeaway from this demo was the ease of use: with a very small configuration, the platform offered out-of-the-box cost monitoring. This demonstrated how the Envoy AI Gateway provides a unified endpoint, simplifying routing and management for clients, regardless of whether the LLM is self-hosted or a third-party service.
Max Körbächer also described a third demo, focusing on observability for Large Language Models using OpenTelemetry and Open LLM Telemetry. This demo, featuring a Bedrock integration with GPT in the background, illustrated how the platform provides deep insights into LLM operations. It showcased metrics such as costs, the amount of requests, average duration of interactions, and critically, how service quality and guardrails are implemented. For instance, basic guardrails could detect negative feedback, toxicity, or potential data leaks in model outputs. Additionally, the demo highlighted the effectiveness of prompt caching in optimizing performance, showing a tangible saving of around $9 in a simple scenario, emphasizing that even small savings can accumulate significantly at scale.
Defensive Implications
▶ Watch: Kserve for serving inference in a Minimal Viable Platform (7:55)
The detailed architecture and tooling presented in "The Explorer's Guide To Cloud Native GenAI Platform Engineering" offer several crucial defensive implications for organizations deploying GenAI workloads:
- Centralized Access and Security with Envoy AI Gateway: The Envoy AI Gateway provides a standardized and centralized point of access for all LLMs, whether self-hosted or third-party. This significantly enhances security by:
- Simplifying Credential Management: Instead of managing credentials across multiple LLM providers, the gateway centralizes this, reducing the attack surface and potential for misconfiguration.
- Enforcing Consistent Authentication/Authorization: All client requests pass through a single gateway, allowing for uniform application of authentication and authorization policies, making it easier to audit and control access.
- Auditable Access Patterns: A centralized gateway provides a clear audit trail of all LLM interactions, which is vital for compliance and incident response.
- Cost Governance and Resource Optimization: The token-based cost monitoring feature of the Envoy AI Gateway is a critical defensive tool against unexpected expenses. GenAI models can be notoriously expensive, and tracking token usage allows organizations to:
- Implement Budget Controls: Set token limits and alerts to prevent runaway costs from inefficient or malicious usage.
- Identify Cost Anomalies: Quickly detect unusual spending patterns that might indicate misuse or suboptimal model prompting.
- Optimize LLM Selection: Inform decisions on which LLMs are most cost-effective for specific use cases.
- Enhanced Observability and Guardrails for LLM Behavior: The integration of OpenTelemetry with Open LLM Telemetry (Open Metri) is paramount for securing and maintaining GenAI applications:
- Early Detection of Malicious Outputs: Guardrails can be implemented to detect and flag potentially toxic, biased, or data-leaking responses from LLMs, preventing harm to users or the organization's reputation.
- Performance Anomaly Detection: Monitoring request durations and service quality helps identify performance degradation, which could be indicative of a denial-of-service attack or resource exhaustion.
- Data Leakage Prevention: By analyzing LLM interactions and outputs, the system can potentially identify attempts to extract sensitive information or unintended data exposure.
- Model Drift Monitoring: Observing model behavior over time helps detect drift in performance or output quality, which could have security implications if the model becomes less accurate or more prone to generating harmful content.
- Platform Engineering for Secure Self-Service: By offloading infrastructure complexity to a dedicated platform team and providing self-service via Backstage, organizations can:
- Reduce Configuration Errors: Data scientists and developers, freed from complex Kubernetes configurations, are less likely to introduce security vulnerabilities through misconfigurations.
- Enforce Best Practices: The platform team can bake security best practices, such as network policies, resource limits, and secure image scanning, directly into the self-service templates.
- Accelerate Secure Deployment: Faster deployment of secure inference services means new GenAI capabilities can be brought to market more quickly and safely.
- Resilience and Performance Optimization: Techniques like model caching, prompt caching (LM Cache), and disaggregated serving contribute to the platform's resilience and efficiency:
- Improved Availability: Faster autoscaling and reduced model download times ensure that services can quickly adapt to demand, minimizing downtime during traffic spikes.
- Resource Efficiency: Optimal use of GPU resources reduces operational costs and ensures the platform can handle more requests with the same infrastructure, making it more robust against resource-exhaustion attacks.
- Predictable Performance: By addressing bottlenecks like KV cache growth and prefill/decoding latency, the platform provides more predictable performance, which is crucial for critical GenAI applications.
In essence, the platform engineering approach advocated in the talk provides a robust framework for managing the unique security and operational challenges posed by GenAI. It shifts the burden of infrastructure management from individual developers to a specialized team, enabling the secure, efficient, and auditable deployment of cutting-edge AI technologies.
Key Takeaways
- The cloud-native GenAI landscape is experiencing explosive growth, necessitating flexible and adaptable platform engineering to integrate rapidly evolving tools and frameworks.
- Kserve is a powerful open-source tool for simplifying AI model inference on Kubernetes, offering extensive out-of-the-box features and new GenAI-specific optimizations like token-based autoscaling and model/prompt caching.
- The new Envoy AI Gateway centralizes LLM access, authentication, and token-based cost monitoring, providing a unified API for clients interacting with diverse LLM providers (on-premise, open-source, or commercial).
- Effective observability for LLMs is crucial, achieved by combining OpenTelemetry with specialized Open LLM Telemetry (Open Metri) to capture model-specific metrics, costs, and enable real-time guardrail implementation against undesirable outputs.
- Advanced optimization techniques like model caching, LM Cache for efficient prompt caching, and disaggregated serving are essential for managing large LLMs, reducing latency, saving GPU cycles, and improving autoscaling performance.
- Platform engineering, leveraging tools like Backstage for self-service, empowers data scientists and developers by abstracting infrastructure complexities, fostering autonomy, and accelerating the secure and efficient deployment of GenAI applications.
About the Speaker(s)
Max Körbächer is a speaker who presented alongside Alexa Griffith. His role and company were not explicitly stated in the introduction, but he shared insights into the broader platform engineering and GenAI landscape.
Alexa Griffith is a Senior Software Engineer at Bloomberg, where she focuses on AI platforms. Her expertise lies in building and integrating complex AI infrastructure, including contributions to open-source projects like the Envoy AI Gateway.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This talk provides an exceptionally detailed and actionable guide to building cloud-native GenAI platforms on Kubernetes. It effectively navigates the complexities of GenAI deployment, introducing a new open-source Envoy AI Gateway for centralized LLM management and cost monitoring, and delving into advanced optimization techniques like token-based autoscaling, model caching, and disaggregated serving. The speakers, clearly operating from a position of deep practical experience, offer a robust blueprint for platform teams to empower data scientists with self-service capabilities while addressing critical security, performance, and cost concerns.
Heather Calloway (CISO) — STRONG ACCEPT
This talk, while deeply technical, delivers substantial value for security leaders navigating the institutional implications of Generative AI adoption. It effectively articulates how a well-engineered platform is fundamental to managing GenAI's inherent risks and complexities, emphasizing platform engineering as an enabler for baking in security and governance from the ground up. The concrete examples of the Envoy AI Gateway for centralized control and cost monitoring, alongside Open LLM Telemetry for guardrails and observability, provide clear, actionable insights for building accountable and secure GenAI capabilities within an enterprise.