AI Pipelines With OPEA: Best Practices for Cloud Native ML Operati... Ezequiel Lanza & Melissa McKay

Ezequiel Lanza, Melissa McKay

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk, "AI Pipelines With OPEA: Best Practices for Cloud Native ML Operations," delivered by Melissa McKay from JFRog and Ezequiel Lanza from Intel at KubeCon EU, delves into the complexities and challenges of building and deploying Generative AI (GenAI) applications in enterprise cloud-native environments. The speakers introduce OPEA (Open Platform for Enterprise AI), an open-source initiative under the Linux Foundation's LF AI and Data umbrella, as a community-driven solution to these hurdles. The core premise of OPEA is to provide standardized, composable building blocks that simplify the development, deployment, and management of GenAI pipelines, fostering collaboration and mitigating common issues like vendor lock-in and excessive trial-and-error.

Watch on YouTube

Visual summary for AI Pipelines With OPEA: Best Practices for Cloud Native ML Operati... Ezequiel Lanza & Melissa McKay by Ezequiel Lanza, Melissa McKay
Visual summary for AI Pipelines With OPEA: Best Practices for Cloud Native ML Operati... Ezequiel Lanza & Melissa McKay by Ezequiel Lanza, Melissa McKay

Key moments

  1. 0:00 Introduction to OPEA project and speakers
  2. 1:00 Fun scavenger hunt for CubeCon attendees
  3. 2:10 Kicking off the talk with an immediate demo
  4. 3:30 Demonstrating AI's problem with lack of context
  5. 4:15 Solving AI context issue by providing a URL
  6. 6:00 Explaining Retrieval-Augmented Generation (RAG)
  7. 7:00 Breakdown of RAG: Retrieval, Augmentation, Generation
  8. 8:00 RAG mechanics: vector databases and embeddings

AI Pipelines With OPEA: Best Practices for Cloud Native ML Operations

Speakers: Ezequiel Lanza, Open Source AI Evangelist, Intel; Melissa McKay, Head of Developer Relations, JFRog

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=6hWoA4jEk5M

Overview

This talk, "AI Pipelines With OPEA: Best Practices for Cloud Native ML Operations," delivered by Melissa McKay from JFRog and Ezequiel Lanza from Intel at KubeCon EU, delves into the complexities and challenges of building and deploying Generative AI (GenAI) applications in enterprise cloud-native environments. The speakers introduce OPEA (Open Platform for Enterprise AI), an open-source initiative under the Linux Foundation's LF AI and Data umbrella, as a community-driven solution to these hurdles. The core premise of OPEA is to provide standardized, composable building blocks that simplify the development, deployment, and management of GenAI pipelines, fostering collaboration and mitigating common issues like vendor lock-in and excessive trial-and-error.

The presentation highlights the significant friction points enterprises face when adopting GenAI, particularly in ensuring security, scalability, cost-efficiency, and regulatory compliance. Through a live demonstration of a Retrieval-Augmented Generation (RAG) system, McKay and Lanza illustrate how OPEA addresses the need for contextually aware AI, moving beyond generic LLM responses. OPEA aims to abstract away the intricate orchestration of multiple microservices involved in GenAI, offering blueprints, components, and infrastructure definitions that accelerate development and streamline operations, ultimately enabling organizations to build robust and reliable AI solutions faster.

Background

▶ Watch: Introduction to OPEA project and speakers (0:00)

The rapid evolution of Generative AI has opened new frontiers for innovation, yet its practical implementation in enterprise settings is fraught with challenges. As Melissa McKay aptly illustrates with a Lego analogy, developers often find themselves building bespoke, "monstrous" solutions from scratch, constantly reinventing the wheel to integrate disparate components. This ad-hoc approach leads to significant investment in trial-and-error, increased development costs, and difficulties in maintenance and scalability. The speakers emphasize that while individual components like Large Language Models (LLMs) are powerful, combining them into functional, production-ready applications, especially in cloud-native environments, is a complex undertaking.

Ezequiel Lanza further elaborates on the multifaceted challenges inherent in GenAI deployment. These include the intricacies of data indexing and vector preparation for knowledge bases, the critical task of model selection (deciding between remote inference with commercial APIs like OpenAI or local deployment of smaller, specialized models), choosing appropriate AI frameworks (e.g., Hugging Face), optimizing hardware efficiency, selecting effective retrieval methods, and robustly evaluating LLMs for performance, bias, and safety. The speakers stress that GenAI pipelines are not monolithic; they are intricate systems comprising multiple interacting components, often deployed as microservices. The absence of industry standards for how these components should interact and be deployed creates a fragmented ecosystem, hindering enterprise adoption and efficient development. This chaotic landscape underscores the urgent need for a standardized, open-source approach that simplifies integration, promotes reusability, and reduces the operational burden on organizations.

Key Findings

▶ Watch: Kicking off the talk with an immediate demo (2:10)

The talk presents several key findings and observations regarding the current state of Generative AI development and deployment:

  1. Complexity and Fragmentation: Building GenAI applications, even for seemingly straightforward tasks like RAG, involves orchestrating numerous components (embeddings, vector databases, LLMs, orchestrators, UIs) as microservices. This complexity, coupled with a lack of standardized interfaces and deployment patterns, leads to significant development overhead and a "trial-and-error" approach for many organizations.
  2. The Need for Context: Generic LLMs often provide irrelevant or unhelpful answers without specific context. Retrieval-Augmented Generation (RAG) is highlighted as a critical pattern to inject external, up-to-date, and domain-specific information into the LLM's generation process, significantly improving accuracy and relevance.
  3. Vendor Lock-in and Reinvention: The current landscape encourages proprietary solutions and forces organizations to either commit to a single vendor's ecosystem or expend considerable resources building custom integrations. This stifles innovation and limits flexibility.
  4. OPEA as a Solution: The core finding is the introduction of OPEA (Open Platform for Enterprise AI) as a strategic, community-driven initiative to address these challenges. OPEA provides standardized, composable building blocks for GenAI pipelines, aiming to simplify development, reduce chaos, and accelerate enterprise adoption.
  5. Rapid Community Adoption: OPEA, launched just a year prior to the talk (and originating as an Intel project two years ago), has already garnered significant industry support, with over 55 contributing partners, including major players like Intel, AMD, JFRog, MinIO, Red Hat, and OpenSearch. This rapid growth demonstrates a clear industry demand for such a standardized platform.
  6. Focus on Enterprise Readiness: Beyond just technical integration, OPEA emphasizes critical enterprise requirements such as security, safety, scalability, cost-efficiency, and agility. It includes tools and working groups dedicated to aspects like model evaluation, guardrails, and secure deployment in regulated environments.

Technical Deep Dive

▶ Watch: Solving AI context issue by providing a URL (4:15)

The technical foundation of the talk revolves around the architecture and components of a robust GenAI pipeline, particularly focusing on Retrieval-Augmented Generation (RAG), and how OPEA standardizes this.

At its core, a RAG system, as explained by Ezequiel Lanza, involves several distinct stages, each ideally handled by a specialized microservice:

  1. User Query and Embedding: When a user poses a question (e.g., "What is kids day?"), this natural language text must first be converted into a numerical representation called a vector or embedding. This process allows for mathematical comparison. The demo showed this with a query like "What was deep learning?" yielding a 512-dimension vector.
  2. Retrieval from Vector Database: The generated query vector is then used to perform a similarity search against a vector database. This database stores pre-computed embeddings of documents or knowledge bases (e.g., web pages, internal documents). The retriever component identifies and extracts the most semantically similar documents or chunks of text relevant to the user's query. Examples of vector databases mentioned as potential OPEA components include MinIO, Red Hat, OpenSearch, and Red Shift.
  3. Augmentation: The retrieved context (the relevant documents) is then combined with the original user query. This forms an "augmented" prompt that provides the LLM with specific, up-to-date information, preventing it from hallucinating or generating generic answers.
  4. Generation by LLM: Finally, the augmented prompt is fed into a Large Language Model (LLM), which generates a coherent and contextually accurate response. The choice of LLM can vary, from remote inference APIs (e.g., OpenAI) to locally deployed Small Language Models (SLMs) depending on cost, privacy, and performance requirements.

OPEA's approach to standardizing this complex pipeline is built on a modular, cloud-native architecture. It conceptualizes each stage and function within a GenAI application as a composable building block, typically implemented as a containerized microservice. The project is structured into several key repositories and concepts:

  • GenAI Examples (Blueprints): These provide end-to-end recipes for common GenAI use cases, such as agent Q&A, chat Q&A, summarization, and document indexing. Each blueprint offers deployment options using Docker Compose files for local development or Kubernetes Helm charts for cloud-native environments. The design allows for easy "plug and play" of different components, enabling users to swap out vector databases or LLMs as needed, rather than being confined to a single configuration.
  • GenAI Comps (Components): This repository houses the individual microservices that constitute the building blocks. These components are categorized by their function (e.g., data preparation, embeddings, vector stores, third-party integrations). The community contributes these components, ensuring a diverse range of options and avoiding vendor lock-in. OPEA defines the specifications for how these components should interact, ensuring interoperability.
  • GenAI Infra: This section provides infrastructure definitions, including Terraform scripts for deploying Kubernetes clusters on various cloud providers (e.g., AWS, Azure) and additional Helm charts specific to different use cases. This simplifies the provisioning of the underlying infrastructure required to run OPEA pipelines.
  • Orchestration: A crucial element, often referred to as a "mega service," is responsible for orchestrating the communication and workflow between these disparate microservices. This orchestrator ensures that data flows correctly through the pipeline, from embedding to retrieval to generation, and that services are exposed appropriately, often via an Nginx ingress.

The core technical contribution of OPEA lies in defining clear specifications and providing concrete implementations for these microservices, enabling developers to assemble complex GenAI pipelines from well-defined, interoperable, and community-contributed components. This abstraction layer significantly reduces the burden of integrating diverse technologies and allows organizations to focus on application-specific logic rather than infrastructure plumbing.

Demo / Proof of Concept

▶ Watch: Explaining Retrieval-Augmented Generation (RAG) (6:00)

The talk featured two distinct yet interconnected demonstrations that powerfully illustrated the problem OPEA solves and its practical application.

The first demonstration, led by Melissa McKay, immediately highlighted a common frustration with generic Large Language Models (LLMs). McKay initially asked a chatbot-like system, "What is kids day?" Without any specific context, the system returned a generic definition of "Kids Day" as a special occasion to honor children, completely unrelated to the KubeCon Kids Day event she had just described. This served as a compelling proof of concept for the inherent limitation of LLMs when lacking domain-specific information. The AI, though intelligent, couldn't differentiate between the general concept and the specific event, underscoring the necessity for contextual augmentation.

To rectify this, McKay then performed the pivotal part of the demo: implementing Retrieval-Augmented Generation (RAG). She provided the system with a URL pointing to the KubeCon events page. When she re-asked, "What is kids day?" the system, now equipped with this specific context, correctly identified it as a technology event at KubeCon for children aged 8 to 14, focusing on open-source technologies, and held at ExCeL London. This live demonstration vividly showcased how providing relevant external knowledge transforms a generic LLM into a highly accurate and useful tool for specific queries. Ezequiel Lanza then broke down this process, explaining the "R" (Retrieval) as extracting information from the provided knowledge base (the URL), the "A" (Augmentation) as combining this context with the prompt, and the "G" (Generation) as the LLM producing the informed response.

The second part of the demo, led by Ezequiel Lanza, delved into the underlying OPEA architecture, showcasing a live Kubernetes deployment. Lanza demonstrated how an OPEA pipeline is composed of various microservices running as pods within a Kubernetes cluster. He used kubectl get pods to display the running components, including a UI, Nginx (for exposing services), a "mega service" orchestrator, and individual microservices like the embeddings service. To illustrate a component in action, he specifically demonstrated interacting with the embeddings microservice. After connecting to the Nginx proxy, he sent a curl request to the embeddings service with the text "What was deep learning?". The service processed this text and returned a large numerical vector (specifically, a 512-dimension vector), visually representing how OPEA components convert natural language into a format suitable for similarity searches within a vector database. This practical demonstration underscored OPEA's modularity and the ability to inspect and interact with individual components, verifying their functionality within a cloud-native deployment. The demo effectively transitioned from illustrating the high-level problem and solution (RAG) to revealing the specific technical building blocks (OPEA microservices) that make such solutions possible.

Defensive Implications

▶ Watch: RAG mechanics: vector databases and embeddings (8:00)

The push for standardized, open-source GenAI pipelines through OPEA carries significant defensive implications for organizations aiming to secure their AI deployments. The chaotic and bespoke nature of current GenAI development, as highlighted in the talk, creates numerous vulnerabilities and makes robust security challenging. OPEA directly addresses these by:

  1. Enabling Secure-by-Design Architectures: By providing standardized, composable microservices, OPEA encourages a more structured approach to GenAI development. This allows security considerations to be baked into the architecture from the outset, rather than being an afterthought. Well-defined component interfaces and clear responsibilities for each microservice reduce the attack surface and simplify security auditing.
  2. Addressing Vendor Lock-in and Supply Chain Risks: The open-source nature and broad community contributions (over 55 partners) mitigate vendor lock-in. This means organizations are not solely reliant on a single vendor's security practices or proprietary black-box solutions, diversifying risk. Furthermore, the transparency of open-source components allows for community-driven security reviews and faster identification and patching of vulnerabilities, as exemplified by the contribution of Prediction Guard for guardrails.
  3. Dedicated Security Working Group: OPEA has a specific Security Working Group actively focusing on critical areas like confidential computing and confidential containers. This demonstrates a proactive commitment to securing the underlying infrastructure and data processing within AI pipelines, which is crucial for handling sensitive enterprise data.
  4. Integrated Evaluation and Guardrails: The project includes tools and emphasizes the importance of LLM evaluation from the beginning of deployment. This allows organizations to assess models for potential bias, fairness, and performance, which are critical for responsible AI and regulatory compliance. The inclusion of guardrails (e.g., contributed by a partner) directly helps prevent undesirable or unsafe outputs, enhancing the safety of GenAI applications in production.
  5. JFRog's Contributions to Enterprise Security: JFRog's involvement in OPEA specifically addresses enterprise security concerns. As specialists in artifact storage and management, they ensure that containers and other artifacts used in OPEA pipelines can be securely stored and managed with Artifactory. Crucially, JFRog also offers model storage and management, including versioning and security scanning of AI models. This ensures that models deployed are traceable, verified, and free from known vulnerabilities, providing a critical layer of defense against poisoned models or supply chain attacks.
  6. Reproducibility and Auditability: OPEA's blueprints and standardized deployment mechanisms (Helm charts, Terraform) promote reproducibility. This means that deployments can be consistently replicated, making it easier to audit configurations, track changes, and ensure compliance across different environments. This level of control is essential for meeting regulatory requirements and maintaining a strong security posture.

By fostering a standardized, transparent, and security-conscious approach, OPEA empowers defenders with better tools and practices to build, deploy, and manage GenAI applications that are not only powerful but also resilient against emerging threats.

Key Takeaways

  • GenAI is Complex, OPEA Simplifies: Building enterprise-grade Generative AI applications, especially those leveraging RAG, involves orchestrating numerous microservices, leading to significant complexity and development overhead. OPEA provides a standardized, open-source framework with composable building blocks to simplify this process.
  • RAG is a Core Pattern: Retrieval-Augmented Generation (RAG) is essential for providing LLMs with context-specific, accurate information, moving beyond generic responses. OPEA offers clear blueprints and components to implement RAG systems effectively.
  • Standards Drive Efficiency and Reduce Chaos: The lack of industry standards in GenAI development leads to fragmented solutions, vendor lock-in, and excessive trial-and-error. OPEA establishes an abstraction layer with agreed-upon specifications, enabling faster development, easier integration, and quicker failure cycles.
  • Community Collaboration is Key: OPEA is a thriving open-source project under the Linux Foundation's LF AI and Data, supported by over 55 partners (including Intel, JFRog, AMD). This collaborative model fosters innovation, ensures vendor neutrality, and provides a rich ecosystem of diverse components.
  • Enterprise Readiness is Paramount: OPEA focuses on critical enterprise requirements like security, safety, scalability, cost-efficiency, and compliance. Dedicated working groups for security and evaluation, alongside contributions from security-focused companies like JFRog, ensure that OPEA-based solutions are production-ready.
  • Modular and Cloud-Native by Design: OPEA's architecture is built on containerized microservices, deployed via Kubernetes Helm charts and supported by cloud-specific Terraform scripts. This modularity allows for flexible component swapping and seamless integration into existing cloud-native infrastructures.

About the Speaker(s)

Melissa McKay serves as the Head of Developer Relations at JFRog, a company known for its artifact management and DevOps solutions. Her expertise lies in fostering developer communities and advocating for best practices in software development. Melissa is also an active member of the OPEA Technical Steering Committee, contributing to the strategic direction and technical implementation of the Open Platform for Enterprise AI project.

Ezequiel Lanza, often known as "Easy," is an Open Source AI Evangelist at Intel. He is deeply involved in the open-source AI ecosystem, holding significant positions as the LF AI and Data Chair and a board member. Ezequiel is a strong advocate for community-driven initiatives in AI, working to make advanced AI technologies more accessible and standardized for developers and enterprises.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

The talk effectively introduces OPEA as a crucial open-source initiative addressing the prevalent chaos and complexity in enterprise GenAI development and deployment. By offering standardized, composable microservice building blocks for critical patterns like RAG, OPEA significantly streamlines operations, reduces vendor lock-in, and promotes a more secure, efficient approach to building production-ready AI pipelines. The speakers, deeply involved in OPEA, convincingly demonstrate its practical utility and highlight its strong community support and focus on enterprise-grade requirements.

Heather Calloway (CISO) — STRONG ACCEPT

This talk expertly diagnoses a critical enterprise challenge: the chaotic, bespoke nature of building Generative AI pipelines. OPEA, as a community-driven, open-source initiative, offers a compelling and practical solution for standardization, composability, and secure deployment. It directly addresses institutional risks like vendor lock-in, high costs, and compliance hurdles, providing a clear path for CISOs and security leaders to establish governance and accelerate secure GenAI adoption. While the initiative is still building momentum, its strategic importance for enterprise resilience and risk management is undeniable.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025