From Chaos To Control: Building ML Platform - George Markhulia & Steve Larkin, Volvo Cars

George Markhulia, Steve Larkin, Volvo Cars

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk, "From Chaos To Control: Building ML Platform," presented by George Markhulia and Steve Larkin from Volvo Cars, details the journey and architecture behind Abacus, their internal Machine Learning (ML) platform. Abacus aims to empower data scientists and analysts within Volvo Cars to rapidly validate ideas, train models, and deploy them into production with minimal friction. The platform, built on a Kubernetes and cloud-native stack, addresses the common challenges of ML development in large enterprises, such as fragmentation, lack of reproducibility, and the arduous path from experimentation to production.

Watch on YouTube

Visual summary for From Chaos To Control: Building ML Platform - George Markhulia & Steve Larkin, Volvo Cars by George Markhulia, Steve Larkin, Volvo Cars
Visual summary for From Chaos To Control: Building ML Platform - George Markhulia & Steve Larkin, Volvo Cars by George Markhulia, Steve Larkin, Volvo Cars

Key moments

  1. 0:00 Introduction to Abacus ML Platform
  2. 1:00 Abacus technology stack and Kubeflow ecosystem
  3. 2:00 Scale and community-driven platform support model
  4. 3:10 The chaotic starting point before Abacus
  5. 4:10 Abacus: three user journeys (First, Day-to-Day, Last Mile)
  6. 5:30 Integration: the real challenge in building an ML platform
  7. 6:50 Frictionless 'First Mile' onboarding and resource provisioning

From Chaos To Control: Building ML Platform

Speakers: George Markhulia, Engineering Manager, ML Platform, Volvo Cars; Steve Larkin

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=fnt3f8sWJLA

Overview

This talk, "From Chaos To Control: Building ML Platform," presented by George Markhulia and Steve Larkin from Volvo Cars, details the journey and architecture behind Abacus, their internal Machine Learning (ML) platform. Abacus aims to empower data scientists and analysts within Volvo Cars to rapidly validate ideas, train models, and deploy them into production with minimal friction. The platform, built on a Kubernetes and cloud-native stack, addresses the common challenges of ML development in large enterprises, such as fragmentation, lack of reproducibility, and the arduous path from experimentation to production.

The speakers articulate how Abacus provides a standardized, yet flexible, environment that abstracts away much of the underlying infrastructure complexity. By focusing on streamlined user journeys—from initial onboarding through day-to-day model development to continuous production monitoring—Volvo Cars has fostered an active community of ML practitioners. The talk highlights critical learnings around platform integration, gradual feature rollout, balancing user freedom with necessary guardrails, and the often-underestimated complexity of CI/CD in the ML domain.

The significance of Abacus lies in its ability to transform a previously chaotic and siloed ML landscape into a cohesive, efficient, and traceable ecosystem. For organizations grappling with scaling their ML initiatives, Volvo's experience with Abacus offers a practical blueprint for building a robust, developer-centric ML platform that accelerates innovation while adhering to enterprise security and operational standards.

Background

▶ Watch: Introduction to Abacus ML Platform (0:00)

Before the inception of Abacus, Volvo Cars faced a common predicament observed in many large enterprises attempting to scale ML initiatives: a scattered technology landscape. Data scientists and analysts operated in individual silos or small, isolated teams, frequently reinventing solutions for identical problems. This fragmented approach led to a significant lack of reproducibility and traceability, making it challenging to link an ML model back to its originating code, data, and parameters. The absence of a common development process and standardized tools meant that getting an ML model from an experimental phase to a production environment was an extremely difficult, if not impossible, endeavor for many.

The initial state offered almost no dedicated support for data scientists, forcing them to navigate complex infrastructure and integration challenges independently. This created a high barrier to entry and slowed down the pace of innovation. Recognizing these inefficiencies, Volvo Cars embarked on building Abacus as a central, common platform. The strategic decision to centralize was not taken lightly, as the speakers acknowledge that a monolithic platform isn't always the right solution. However, in this specific context, a unified platform was deemed essential to coalesce a diverse community of data scientists, instill common practices without being overly rigid, and solve pervasive problems like enterprise network integration and security compliance once, for all users.

Today, Abacus serves approximately 200 monthly active users, supporting production workloads that have been running for around three years. The platform boasts an impressive 48 contributors to its inner-sourced repository, demonstrating a strong community engagement. This shift from a chaotic, unsupported environment to a controlled, community-driven platform underscores the necessity of a well-architected ML platform in a modern enterprise.

Key Findings

▶ Watch: Scale and community-driven platform support model (2:00)

The development and operation of Abacus at Volvo Cars yielded several critical findings, primarily centered around user experience, technical architecture, and organizational dynamics:

  1. Frictionless User Journeys through Tiered Functionality: Abacus structures the entire ML lifecycle into three distinct user journeys: the First Mile (onboarding and project creation), Day-to-Day Usage (producing insights or training models), and the Last Mile (taking models to production and monitoring). Within the day-to-day phase, a tiered approach—Insights (lightweight for exploratory analysis) and ML Product (for automation and production-grade models)—proved crucial. This graduated introduction of functionality prevents user overwhelm and caters to different personas and project goals, ensuring gradual adoption and maximum utility.
  1. Integration is Paramount, Installation is Easy: A recurring theme emphasized by the speakers is that "integration takes much, much more time and effort and energy than installation." While installing individual cloud-native components like Kubeflow, Tekton, or Argo CD can be straightforward, the real challenge lies in weaving these components together, integrating them with the broader enterprise ecosystem (e.g., identity management, networking, data sources), and building robust development processes around them. This comprehensive integration is what truly delivers reproducibility, traceability, and visibility.
  1. The Power of GitOps and Automation: The platform leverages GitOps extensively, particularly in the Last Mile for deploying and managing production models and applications via Argo CD. This approach not only removes the manual overhead for users but also provides inherent resilience and self-healing capabilities, which users "love" and find reassuring. Automated CI/CD pipelines, orchestrated by Tekton, are critical for standardizing component builds, tagging, and deployment, significantly reducing human error and complexity in the multi-dimensional ML space.
  1. Community Building and Platform as a Product: Volvo Cars treats Abacus as a product, complete with dedicated engineering teams, transparent support (via Slack, directly with engineers), and an inner-sourced repository. This product-centric approach, coupled with fostering an active community (evidenced by 48 contributors), is vital for long-term platform growth and adoption. It ensures the platform evolves based on user needs and pain points, fostering a collaborative environment where users even help each other.
  1. FinOps and Responsible Resource Usage: By surfacing cumulative cost information directly within the Kubeflow central dashboard, Abacus encourages cost-conscious behavior among users. Even for minimal costs, users proactively request offboarding to save resources, demonstrating the effectiveness of transparent financial feedback loops. This also highlights the benefit of a frictionless onboarding/offboarding process, allowing users to spin up resources only when needed.

Technical Deep Dive

▶ Watch: The chaotic starting point before Abacus (3:10)

Abacus is fundamentally built on a Kubernetes and cloud-native stack, leveraging the Kubeflow ecosystem as its core. The platform benefits significantly from Volvo Cars' common enterprise container platform, which manages underlying Kubernetes clusters, allowing the Abacus team to focus specifically on ML concerns.

The technology stack integrates a variety of open-source and cloud-native tools:

  • Foundation: Kubernetes, Istio (for service mesh, authorization, ingress, TLS), Prometheus, Grafana.
  • ML Orchestration: Kubeflow (central dashboard, profiles, pipelines, Katib, KServe).
  • CI/CD: Tekton (for building and deploying components), Argo CD (for GitOps-driven production deployments).
  • Data Management: LakeFS (for Git-like data versioning), Spark (for distributed processing), DuckDB (for lightweight analytics).
  • Security & Identity: Vault (secret management), Azure AD (identity and access management).
  • Registry: Harbor (container image registry).
  • Custom Components: An internal onboarding application, a custom logging service for monitoring.

Let's break down the technical implementation across the three user journeys:

First Mile: Onboarding and Project Creation

The onboarding process is designed to be frictionless. Users navigate to a single URL, the Kubeflow dashboard, which has been patched by Volvo Cars to integrate a custom onboarding application. This application automates the provisioning of all necessary resources:

  1. Kubeflow Profile & Kubernetes Namespace: A dedicated profile and namespace are created for the user.
  2. GitHub Repository: A new repository is generated from a data science-optimized template, providing a standardized project structure with boilerplate for unit and integration tests.
  3. Container Image Registry: A dedicated registry (implied Harbor) is provisioned.
  4. Vault Secret Store: A secret store is set up for secure credential management.
  5. LakeFS Repository: A repository in LakeFS is created for versioning of data.
  6. CI Infrastructure: Tekton pipelines are provisioned for building Kubeflow pipeline components.
  7. Azure AD Group: An Azure AD group is created for seamless integration with corporate identity and access management, granting access to all provisioned services.

Users receive an email with links to these services. From the Kubeflow central dashboard, they can launch a notebook server in their personal sandbox. Multi-tenant isolation is enforced using Kubernetes Network Policies and Istio Authorization Policies, preventing users from accessing each other's namespaces. FinOps is enabled from the start by displaying cumulative costs for the last 30 days directly in the Kubeflow UI, fostering cost-conscious behavior.

Day-to-Day Usage: Insights and ML Product Tiers

This stage caters to diverse user needs through two distinct tiers:

Insights Tier (Lightweight Analytics)

This tier provides a simplified environment for exploratory data analysis, hypothesis testing, and reporting:

  • Source Code Versioning: Users clone their templated GitHub repository into their notebook servers.
  • Notebook Servers: Pre-built notebook images with common packages and CLI connectors are provided. Users can also contribute custom Dockerfiles via pull requests to the platform's repository, which triggers Tekton CI to build, tag, and upload their custom images to the Harbor registry, making them available to all.
  • Data Processing: Users can schedule Spark jobs from notebooks or leverage lightweight tools like DuckDB.
  • Data Access & Versioning: The platform is deeply integrated with Volvo Cars' network, providing access to various data sources. LakeFS offers Git-like data versioning capabilities.

ML Product Tier (Automation and Production-Ready Models)

This tier extends the Insights tier by adding automation and production-grade features for ML model development:

  • Repository Upgrade: The GitHub repository is upgraded to hold manifests for arbitrary applications, Kubeflow Pipelines, and Kubeflow Components.
  • Automated CI/CD: A commit to the user's repository triggers a GitHub webhook to Tekton. The CI pipeline automatically:
  • Builds container images for Kubeflow components.
  • Tags these images with the commit SHA.
  • Uploads them to the Harbor image registry.
  • Builds and compiles Kubeflow Pipelines.
  • Uploads compiled pipelines to Kubeflow.
  • ML Workflows: Users operate within the familiar Kubeflow environment, utilizing:
  • Kubeflow Training Operators for distributed model training.
  • Katib for hyperparameter tuning.
  • Kubeflow Pipelines for data processing, model training, and eventually deploying models as KServe inference services.
  • Standardization: This tier provides critical standardization, automating the complex and error-prone steps of building, tagging, registering, and compiling ML artifacts, from the first commit to the deployment of an inference service.

Last Mile: Production Deployment and Monitoring

The final stage focuses on seamless transition to production with end-to-end ownership:

  • Deployment with Argo CD: Instead of deploying inference services directly via Kubeflow Pipelines, all application and inference service manifests are stored in the user's GitHub repository and deployed via Argo CD. Rolling out new versions simply involves bumping a tag in the repository and merging a pull request.
  • Flexible Inference: Users commonly deploy both KServe inference services and custom FastAPI applications alongside them, often for pre-processing payloads or adding a UI layer.
  • Comprehensive Monitoring:
  • Logging Service: A custom logging service receives CloudEvents from KServe, correlates input and output events using the X-Request-ID header, and ingests them into a user-chosen storage solution (allowing teams to "bring their own storage"). This provides full control over data sharing and enables users to apply their own analytics tools.
  • Alerting: Argo CD sends Slack notifications if services go out of sync. The monitoring service also sends Slack notifications for critical events like model drift.
  • Metrics & Visualization: Prometheus and Grafana are used for visualizing metrics.
  • Network & Security: Istio provides out-of-the-box ingress and end-to-end TLS certificate renewal, abstracting away network complexities from users.
  • Iterative Process: Production is treated as an iterative process requiring continuous maintenance, monitoring, retraining, and redeployment. Seamless transitions between stages are crucial to avoid "big bang" handovers.
  • Safe Testing Environments: The platform supports both dev/prod namespaces within a single cluster and separate dev/prod clusters, allowing users to test models safely and the platform team to upgrade components without impacting production workloads.

Demo / Proof of Concept

▶ Watch: Integration: the real challenge in building an ML platform (5:30)

The speakers demonstrated the user-friendly aspects of the Abacus platform, primarily focusing on the Kubeflow central dashboard and the project creation workflow.

Steve Larkin walked through the customized Kubeflow dashboard, highlighting the tailored color scheme, direct links to documentation, and the integrated Slack support channel. A key feature showcased was the "Cost Card" prominently displayed in the UI, which presents the cumulative costs for the last 30 days. This immediate visibility into resource consumption is a core component of their FinOps strategy.

The demonstration then proceeded to the "New Project" page, where users can initiate the creation of a new ML project. The process involved:

  1. Clicking "Start" to begin project creation.
  2. Entering a project name (e.g., "kubecon EU 2025").
  3. Optionally specifying a Python package name.
  4. Selecting the desired tier (Insights or ML Product), which dictates the functionality and automation level.
  5. Setting a minimum Python package version for the generated source template.

Although the full 40-second creation process was not shown live, the demonstration effectively illustrated the simplicity and guided nature of the onboarding flow, which abstracts away the complex infrastructure provisioning happening behind the scenes. This quick, self-service project creation from a standardized template is a cornerstone of Abacus's frictionless "First Mile" experience.

Defensive Implications

▶ Watch: Frictionless 'First Mile' onboarding and resource provisioning (6:50)

The design of Abacus incorporates several critical defensive measures and best practices to ensure security, stability, and responsible resource utilization within a multi-tenant enterprise environment:

  1. Multi-Tenant Isolation: The platform strictly enforces isolation between user projects. This is achieved through Kubernetes Network Policies and Istio Authorization Policies, which prevent users from accessing or even seeing each other's Kubernetes namespaces. This is fundamental for data privacy and preventing unauthorized access.
  1. Opinionated Security Standards: While providing flexibility, Abacus maintains a strong opinionated stance on critical security aspects. For instance, if a user needs to create a ServiceEntry (an Istio resource for external service access), they must submit a pull request to the platform's central repository. This PR is then reviewed by platform engineers before merging, preventing unintentional network openings or misconfigurations that could expose internal systems. This controlled mechanism ensures adherence to Volvo's prescribed security standards.
  1. Integrated Identity and Access Management (IAM): Integration with Azure AD groups for identity and access management ensures that access to provisioned services (GitHub repos, Vault, LakeFS, etc.) is managed centrally and follows corporate policies. This simplifies user management and strengthens security posture.
  1. Secure Secret Management: The provisioning of a Vault secret store for each project ensures that sensitive credentials and API keys are stored and accessed securely, rather than being hardcoded or exposed.
  1. Automated TLS and Ingress Management: Istio handles ingress and end-to-end TLS certificate renewal automatically for production services. This offloads a significant security burden from data scientists and ensures that all external communication is encrypted, reducing the risk of data interception.
  1. FinOps and Cost Awareness: By surfacing cumulative costs directly in the Kubeflow UI, Abacus promotes a culture of cost-consciousness. While not a direct security measure, it encourages users to manage resources efficiently, potentially reducing the attack surface by decommissioning unused infrastructure.
  1. Standardized and Automated CI/CD: The rigorous, automated CI/CD pipelines (Tekton) for building and deploying ML components reduce the likelihood of human error or manual misconfigurations that could introduce vulnerabilities. Consistent tagging with commit SHAs also aids in traceability for auditing and incident response.
  1. Safe Testing Environments: The provision of separate development and production namespaces, and even distinct development and production clusters, allows users to test models and platform engineers to upgrade components safely. This minimizes the risk of introducing regressions or vulnerabilities into production environments and provides a sandbox for experimentation without impacting critical services.

Key Takeaways

  • Integration over Installation: Building a production-ready ML platform demands significantly more effort in integrating diverse components and the broader enterprise ecosystem than in merely installing individual tools.
  • Gradual Functionality for Diverse Personas: Design for different user needs by introducing functionality gradually through tiered approaches (e.g., Insights vs. ML Product) to prevent overwhelm and ensure broad adoption.
  • Balance Freedom with Opinionated Guardrails: Provide users with flexibility while enforcing necessary restrictions and standards (e.g., PRs for network changes) to maintain security, reproducibility, and platform stability.
  • CI/CD is Complex and Critical for ML: The multi-dimensional nature of ML (code, data, model, artifacts) makes CI/CD particularly challenging but absolutely essential for standardization, automation, and traceability from development to production.
  • GitOps Drives Reliability and Efficiency: Leveraging GitOps for production deployments (e.g., with Argo CD) provides automated, self-healing, and auditable infrastructure management, significantly reducing operational overhead.
  • Treat the Platform as a Product and Build a Community: Actively manage the ML platform as a product with dedicated engineering, transparent support, and foster an inner-sourced community to ensure its continuous evolution and user adoption.

About the Speaker(s)

George Markhulia is an Engineering Manager in the ML Platform team at Volvo Cars. He is instrumental in leading the development and strategic direction of Abacus, Volvo's internal Machine Learning platform. His insights in the talk underscore his deep understanding of the challenges in scaling ML within large organizations and the importance of a well-engineered, user-centric platform.

Steve Larkin is a colleague of George Markhulia at Volvo Cars, also involved in the ML Platform team. Steve played a key role in the presentation, detailing the foundational technology stack of Abacus and guiding the audience through the "First Mile" user journey, showcasing the platform's seamless onboarding experience. His contributions highlight the collaborative effort in building and maintaining Abacus.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This talk from Volvo Cars isn't just another 'ML platform' rehash; it's a deeply technical, brutally honest account of building a production-grade, multi-tenant machine learning ecosystem from the ground up. The speakers detail the architecture, user journeys, and critical integration challenges with a level of specificity rarely seen. For any organization struggling to scale ML beyond isolated experiments, this provides an actionable, battle-tested blueprint that cuts through the marketing fluff and delivers genuine engineering insights. It’s a masterclass in platform as a product, demonstrating how to balance user freedom with robust enterprise standards and security.

Heather Calloway (CISO) — STRONG ACCEPT

This talk from Volvo Cars provides a compelling case study for how a large enterprise can build a robust, secure, and governed Machine Learning platform. It effectively articulates the transition from a fragmented, high-risk ML landscape to a centralized, controlled environment. The focus on integrating security controls—such as multi-tenant isolation, secure identity management, and controlled network access—directly into the platform's architecture demonstrates a proactive approach to managing ML-related business and operational risks, making it highly relevant for security leaders grappling with scaling AI/ML initiatives.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025