Testing AI Containers for Digital Twins in Science: A Cloud-HPC... Matteo Bunino & Diego Ciangottini

Matteo Bunino, Diego Ciangottini

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk, presented by Diego Ciangottini from the National Institute for Nuclear Physics (INFN) and Matteo Bunino from CERN Openlab, delves into the creation of a pioneering platform designed to develop and test digital twins within a hybrid cloud and High-Performance Computing (HPC) environment. The core challenge addressed is how to enable scientific communities, particularly those in particle physics and environmental sciences, to leverage diverse computational resources – from traditional cloud infrastructure to EuroHPC supercomputers – for complex AI-driven simulations while maintaining workflow consistency and reproducibility.

Watch on YouTube

Visual summary for Testing AI Containers for Digital Twins in Science: A Cloud-HPC... Matteo Bunino & Diego Ciangottini by Matteo Bunino, Diego Ciangottini
Visual summary for Testing AI Containers for Digital Twins in Science: A Cloud-HPC... Matteo Bunino & Diego Ciangottini by Matteo Bunino, Diego Ciangottini

Key moments

  1. 0:00 Introduction and Digital Twin concept
  2. 2:30 Challenges of Digital Twins in Hybrid Cloud/HPC
  3. 5:00 InterLink: Kubernetes API for HPC & Supercomputers
  4. 6:00 InterLink's community adoption and CNCF Sandbox status
  5. 8:00 Ensuring workflow consistency with Dagger
  6. 9:00 Dagger's features for repeatable, observable pipelines

Testing AI Containers for Digital Twins in Science: A Cloud-HPC... Matteo Bunino & Diego Ciangottini

Speakers: Diego Ciangottini, National Institute for Nuclear Physics (INFN); Matteo Bunino, CERN Openlab

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=bIxw1uK0QRQ

Overview

This talk, presented by Diego Ciangottini from the National Institute for Nuclear Physics (INFN) and Matteo Bunino from CERN Openlab, delves into the creation of a pioneering platform designed to develop and test digital twins within a hybrid cloud and High-Performance Computing (HPC) environment. The core challenge addressed is how to enable scientific communities, particularly those in particle physics and environmental sciences, to leverage diverse computational resources – from traditional cloud infrastructure to EuroHPC supercomputers – for complex AI-driven simulations while maintaining workflow consistency and reproducibility.

The speakers highlight the critical need for a common, interoperable interface that can abstract away the underlying heterogeneity of compute resources, allowing scientists to focus on their research rather than infrastructure management. By integrating cloud-native tools like Kubernetes with specialized solutions like InterLink for HPC connectivity and Dagger for robust CI/CD pipelines, they present a robust framework for developing, testing, and deploying scalable AI models for scientific digital twins. This work is a significant contribution to bridging the gap between cloud elasticity and HPC power, ensuring that cutting-edge scientific simulations are both efficient and reliable.

The significance of this platform extends to enabling reproducible research, accelerating scientific discovery, and fostering collaboration across diverse computational facilities. It tackles the practical difficulties of running inherently distributed AI workloads on heterogeneous infrastructure, offering a blueprint for other scientific and industrial applications requiring similar hybrid computing capabilities. The ability to automatically test complex AI models on supercomputers as part of a continuous integration pipeline represents a substantial leap forward in scientific software development practices.

Background

▶ Watch: Introduction and Digital Twin concept (0:00)

The concept of digital twins serves as the foundation for this discussion. As defined by the speakers, a digital twin is a "digital representation of a real-world system." This goes beyond mere simulation, aiming for a dynamic, real-time counterpart that can be used for forecasting, analysis, and optimization. Examples provided include predicting wildfire spread, assessing flood impact, simulating noise in precise measurement environments, and even creating digital representations of particle detectors to facilitate experimental design and data analysis. These diverse use cases share common needs: the ability to process large datasets, execute complex computational models, and leverage highly parallelized computing resources.

Recognizing these shared requirements, the Intertwin European project was initiated to create a generic "engine" for digital twins. This engine aims to serve multiple scientific communities by enabling them to adopt and utilize shared resources from various providers, ranging from commercial cloud platforms to EuroHPC supercomputers. However, integrating such disparate resources presents significant challenges. The platform must be capable of supporting a wide array of use cases and computational frameworks, while simultaneously offering the flexibility to offload tasks to different backend types that may not natively support cloud payloads. Crucially, all software must remain as interoperable as possible, ensuring that workflows are reproducible across any backend.

In a nutshell, the core problem is managing distributed, heterogeneous resources and providing a common interface on top. Two key issues arise: how to grant users access to this complex platform, and how to maintain workflow consistency and software reproducibility. Many scientific users, as highlighted, "don't care where they are going to run; they just want their jobs to be done." Fortunately, in particle physics and other scientific domains, there's a growing convergence towards cloud-native tools, with Kubernetes emerging as a de facto standard interface. This convergence simplifies the problem to: "How can we merge a Kubernetes API access to different kinds of resources?" This question forms the technical bedrock upon which the presented solution is built, addressing the need for seamless integration of cloud elasticity with the raw power of HPC.

Key Findings

▶ Watch: InterLink: Kubernetes API for HPC & Supercomputers (5:00)

The talk presents several key findings and technological contributions that collectively form a robust platform for scientific digital twins in a hybrid cloud-HPC environment:

  1. InterLink for HPC Integration: The development and adoption of InterLink as a pluggable system to connect Kubernetes clusters with diverse remote resources. InterLink utilizes virtual Kubelet technology to create virtual nodes within a Kubernetes cluster, capable of scheduling pods onto HPC supercomputers (e.g., Slurm), VMs with GPUs, HTC systems (like HTCondor), and even quantum computing resources. This allows users to interact with HPC resources via a standard Kubernetes API, making the underlying complexity transparent. InterLink has gained significant traction, recently becoming part of the Cloud Native Sandbox project, demonstrating its value and community interest.
  1. Dagger for Reproducible CI/CD: The adoption of Dagger as a runtime for creating composable and observable software pipelines. Dagger addresses the critical need for reproducible results and consistent environments in scientific workflows, moving beyond "copy-pasting nightmares" or custom CI solutions. It provides a universal type system, platform-agnostic caching for artifacts, and built-in observability. Dagger enables the same pipeline to run consistently on a developer's laptop, in a Docker engine, or within a CI system, significantly improving developer experience and ensuring the integrity of scientific software.
  1. i-twin-AI Library for Scalable AI: The creation of i-twin-AI, a Python library specifically designed for scalable AI workflows in scientific digital twin applications. This toolkit provides scientists with functionalities for distributed machine learning training (supporting PyTorch and TensorFlow), distributed hyperparameter optimization (leveraging Ray Tune), and robust machine learning tracking (using MLflow for metadata and model registry integration). i-twin-AI supports both data parallel and model parallel training paradigms, essential for handling large models and datasets common in scientific research.
  1. Novel Distributed Testing Methodology: A crucial finding is the successful implementation of an end-to-end CI/CD pipeline that integrates HPC resources for testing inherently distributed AI features. Traditional unit tests are insufficient for functionalities like worker rank allocation, collective operations, or distributed checkpointing. The presented methodology uses distributed launchers (e.g., torch run) to spawn multiple pytest commands across HPC nodes, communicating via collective backends. This allows for rigorous validation of i-twin-AI's distributed functionalities on actual HPC infrastructure, directly within a CI pipeline.
  1. Automated Hybrid Cloud-HPC CI/CD: The culmination of these technologies into a fully automated CI/CD workflow that spans cloud (GitHub) and HPC environments. This pipeline automatically builds Docker containers, converts them to Singularity images for HPC compatibility, deploys an ephemeral Kubernetes cluster with InterLink inside the Dagger pipeline, executes distributed tests on remote supercomputers, and publishes validated container images. This completely automates the "push and pray" cycle, ensuring that scientific AI models are thoroughly tested on their target execution environment before deployment.

Technical Deep Dive

▶ Watch: InterLink's community adoption and CNCF Sandbox status (6:00)

The technical architecture presented by Ciangottini and Bunino is a sophisticated integration of cloud-native principles with HPC realities, designed to provide a seamless experience for scientific users.

At its core is InterLink, a project that addresses the challenge of extending Kubernetes' control plane to external, often non-cloud-native, computational resources. InterLink functions as a virtual Kubelet, a concept pioneered to allow Kubernetes to schedule pods onto arbitrary backends that don't necessarily run standard Kubelet agents. In this setup, InterLink acts as an adapter, presenting remote HPC resources (such as Slurm-managed supercomputers like the Vega cluster in Slovenia, or even quantum computing platforms) as virtual nodes within a Kubernetes cluster. When a user submits a pod with specific annotations, targeting a virtual node, InterLink translates this into a job submission on the underlying HPC scheduler. This design ensures that users can leverage familiar Kubernetes YAML definitions and APIs, while their workloads transparently execute on powerful supercomputing infrastructure, completely unaware of the translation layer. The requirements for providers to integrate with InterLink are kept minimal, fostering wider adoption.

To ensure consistency and reproducibility across this hybrid environment, the team adopted Dagger. Dagger is a software-defined CI/CD engine that allows developers to write their pipelines in general-purpose programming languages (like Go, Python, or TypeScript) rather than YAML. This approach promotes composability, reusability, and testability of CI logic. Dagger pipelines are container-based, meaning each step runs in a container, guaranteeing a consistent execution environment. A key feature is its ability to build an ephemeral environment for testing. As demonstrated in the talk, a Dagger pipeline can spin up a lightweight Kubernetes distribution like K3s on the fly, and then deploy InterLink on top of it. This self-contained, temporary Kubernetes cluster within the CI pipeline allows for end-to-end validation of the InterLink integration before any changes are pushed to production. Dagger also provides caching mechanisms for artifacts and built-in observability through its cloud platform, which traces pipeline execution, making it easier to debug and understand complex workflows.

The scientific application layer is handled by i-twin-AI, a Python library developed within CERN Openlab. This library is designed to abstract the complexities of distributed machine learning frameworks, providing a unified interface for scientists. It supports two primary paradigms for distributed training:

  • Data Parallel Training: A single model is replicated across multiple GPUs or nodes, and the dataset is partitioned, with each model replica processing a subset of the data. This is common for scaling training on larger datasets.
  • Model Parallel Training: For models too large to fit on a single GPU (e.g., very large language models or transformer-based architectures), the model itself is distributed across multiple GPUs or nodes. This often involves techniques like pipeline parallelism or tensor parallelism.

i-twin-AI relies on popular underlying frameworks such as PyTorch, Ray, and DeepSpeed to implement these distributed training strategies. For hyperparameter optimization (HPO), it integrates with Ray Tune, allowing scientists to define hyperparameter ranges and automatically run multiple training trials in parallel across the HPC infrastructure to find optimal model configurations. The library also emphasizes ML metadata tracking using MLflow, enabling versioning of models and experiments.

A crucial technical innovation detailed in the talk is the strategy for testing inherently distributed AI features. Traditional unit tests cannot validate functionalities that require multiple processes communicating across a network, such as worker rank allocation (assigning unique identifiers to processes for collective communication), collective operations (e.g., all_gather, barrier), or distributed checkpointing. The solution involves:

  1. Distributed Launchers: Utilizing framework-specific launchers, such as torch run for PyTorch's Distributed Data Parallel (DDP), to initiate multiple processes.
  2. Pytest Integration: Each spawned process executes a standard pytest command, allowing individual test cases to run within their respective distributed context.
  3. Collective Communication Backends: The test cases themselves are designed to communicate with each other using the collective communication backends provided by the underlying distributed machine learning framework (e.g., NCCL or Gloo for PyTorch).

This setup allows the team to write integration tests that verify the correct behavior of distributed functionalities directly on HPC resources, ensuring the robustness of the i-twin-AI library.

Demo / Proof of Concept

▶ Watch: Ensuring workflow consistency with Dagger (8:00)

The core of the "demo" aspect of this talk is the detailed description of an end-to-end Continuous Integration (CI) and Continuous Delivery (CD) pipeline that automates the testing and deployment of AI containers for scientific digital twins. This pipeline effectively bridges the gap between traditional cloud-based development workflows and the specialized requirements of HPC environments.

The workflow begins on the cloud side, specifically with code hosted on GitHub. When changes are pushed, a GitHub Actions workflow is triggered, which in turn invokes a Dagger pipeline. This Dagger pipeline orchestrates the entire process:

  1. Container Build and CPU Tests: The pipeline first builds a Docker container for the i-twin-AI library and runs standard, CPU-only unit tests. These initial tests are quick and ensure basic functionality before engaging more resource-intensive HPC infrastructure.
  2. Docker to Singularity Conversion: HPC systems often use Singularity (now Apptainer) containers due to their security model and integration with HPC schedulers. The Dagger pipeline automatically converts the Docker image into a Singularity Image Format (.sif) file. This Singularity image is then pushed to a dedicated registry, such as a Harbor registry hosted on CERN resources.
  3. Ephemeral InterLink Deployment: A critical step involves Dagger dynamically deploying a lightweight Kubernetes cluster, specifically K3s, inside the pipeline's execution environment. On top of this ephemeral K3s cluster, InterLink is then deployed. This creates a temporary, self-contained environment that can interact with remote HPC resources. This "on-the-fly" deployment ensures that the InterLink integration itself is tested within the CI pipeline.
  4. HPC Test Execution: With InterLink bootstrapped, the Dagger pipeline uses it to submit jobs to a remote HPC supercomputer (e.g., the Vega cluster). These jobs are the distributed tests for i-twin-AI, designed to validate features like worker rank allocation, collective operations, and distributed checkpointing on actual multi-node, multi-GPU HPC hardware.
  5. Image Publication: If all tests on the HPC pass successfully, the Dagger pipeline proceeds to publish the validated container images. The Docker image is pushed to a standard registry (e.g., GitHub Container Registry), and the Singularity image is pushed to the designated Singularity registry (e.g., CERN's Harbor).

The speakers illustrate the modularity of this Dagger-based CI/CD. They define three distinct Dagger types:

  • iti: Encapsulates the logic for building containers, connecting to InterLink, running HPC tests, and converting to Singularity.
  • interlink: Provides the functionality to bootstrap the InterLink service.
  • singularity: Handles the conversion of Docker containers to Singularity images.

This modularity allows developers to compose different CI pipelines, for instance, just building and publishing an image, or launching a terminal inside a built container, or running CPU-only tests.

The talk provides visual traces from Dagger Cloud, demonstrating the execution flow of such a pipeline. It shows how variables and secrets are passed, how the container is built, and the final "release pipeline" steps involving InterLink deployment, HPC test execution, and image pushing. This complete user story exemplifies how the platform ensures that complex AI applications for digital twins are consistently built, thoroughly tested on their target HPC environment, and reliably deployed.

Beyond the CI/CD pipeline, the speakers also presented concrete examples of i-twin-AI's application to scientific digital twin use cases:

  • Hydro-geological modeling: Developing AI models to improve early warnings for droughts. The platform enabled a 75% reduction in validation loss through hyperparameter optimization.
  • Gravitational wave denoising: Using AI models to denoise signals captured by the Virgo interferometer.
  • Scalability and Energy Benchmarking: The platform was used to compare how different distributed frameworks scale for the same model and dataset, as well as to study their energy consumption, highlighting important trade-offs for sustainable scientific computing.

Defensive Implications

▶ Watch: Dagger's features for repeatable, observable pipelines (9:00)

While the talk primarily focuses on scientific reproducibility, CI/CD automation, and hybrid cloud-HPC integration, it carries significant "defensive implications" from the perspective of ensuring the reliability, integrity, and efficiency of scientific AI workflows. In this context, "defense" refers to safeguarding against common pitfalls in complex scientific software development, rather than traditional cybersecurity threats.

  1. Defense Against Regressions and Inconsistencies: The robust, automated CI/CD pipeline, powered by Dagger and InterLink, acts as a primary defense against software regressions. By running comprehensive distributed tests on actual HPC infrastructure with every code change, the system ensures that new features or refactorings do not inadvertently break existing functionalities or introduce performance bottlenecks in the highly complex and distributed AI models. This prevents the "push and pray" mentality, where issues are only discovered after lengthy, costly deployments on supercomputers.
  1. Ensuring Model Reliability for Critical Applications: Digital twins, especially in environmental sciences (e.g., wildfire forecasting, drought early warnings) or physics (e.g., particle detection, gravitational wave analysis), often feed into critical decision-making processes. An unreliable or incorrect AI model could have severe consequences. The rigorous testing of i-twin-AI's distributed functionalities directly on HPC hardware is a defense mechanism against deploying flawed models, thereby enhancing the trustworthiness and accuracy of scientific predictions.
  1. Resource Optimization and Cost Efficiency: HPC resources are expensive and often have long queue times. Allocating significant compute time for a job that crashes due to an untested bug is a major waste of resources and money. The ability to perform automated "dry runs" and comprehensive tests on HPC within the CI pipeline serves as a defense against such inefficiencies. It ensures that only thoroughly validated code proceeds to large-scale, long-running jobs, optimizing resource utilization and reducing operational costs.
  1. Reproducibility and Auditability: The Dagger framework inherently promotes reproducibility by defining pipelines as code and running them in containerized environments. This is a defense against the "works on my machine" problem and ensures that scientific results can be consistently replicated. The built-in observability and tracing capabilities of Dagger Cloud provide an audit trail of pipeline executions, which is crucial for scientific validation and accountability.
  1. Standardization and Interoperability: The adoption of Kubernetes as a common interface via InterLink and the use of Docker/Singularity containers standardize the execution environment. This defends against environment-specific bugs and configuration drift, ensuring that the AI models behave consistently across different HPC centers and cloud providers. This standardization also simplifies onboarding for new users and fosters broader collaboration.
  1. Scalability and Performance Assurance: The proposed next steps, including testing for scalability and energy consumption, directly contribute to defensive posture. By continuously benchmarking performance, the team can defend against introducing inefficiencies in their AI trainer code. This ensures that the scientific applications remain performant and energy-efficient as they evolve, which is increasingly important for sustainable computing.

In essence, the defensive implications here are about building a resilient, reliable, and efficient scientific software development ecosystem. It's a defense against the inherent complexities of distributed systems, heterogeneous infrastructure, and the need for high-integrity scientific outcomes.

Key Takeaways

  • Seamless Hybrid Cloud-HPC Integration: InterLink enables Kubernetes to transparently schedule workloads on diverse HPC and cloud resources, providing a unified interface for scientific users.
  • Reproducible and Observable CI/CD: Dagger pipelines offer a powerful, container-based, and code-defined approach to build, test, and deploy scientific AI applications, ensuring consistency and auditability across environments.
  • Specialized AI Toolkit for Science: The i-twin-AI Python library simplifies distributed machine learning training, hyperparameter optimization, and experiment tracking for complex scientific digital twin use cases.
  • Automated Distributed Testing on HPC: A novel CI/CD methodology integrates ephemeral Kubernetes clusters with InterLink to execute inherently distributed AI tests directly on supercomputers, validating critical functionalities before deployment.
  • Efficiency and Reliability for Scientific Computing: This integrated platform significantly reduces development cycles, prevents costly resource waste from untested jobs, and enhances the reliability of AI models used in critical scientific applications.
  • Extensible Framework for Scientific Validation: The developed CI/CD framework can be adapted by other scientific communities to perform crucial "dry runs" and validation tests for their large-scale HPC AI jobs, improving overall scientific software quality.

About the Speaker(s)

Diego Ciangottini is from the National Institute for Nuclear Physics (INFN) in Italy. His work, as presented in the talk, focuses on developing platforms and solutions for scientific computing, particularly in the realm of hybrid cloud and HPC environments for digital twins.

Matteo Bunino is from CERN Openlab, an entity within the CERN IT department that is responsible for establishing collaborations with industry and academia. He is involved in the Intertwin project and leads the development of the i-twin-AI Python library, focusing on scalable AI workflows for scientific digital twin applications.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This session by INFN and CERN Openlab presents a highly sophisticated and practical platform for developing, testing, and deploying scientific digital twins in a hybrid cloud-HPC environment. Leveraging Kubernetes with InterLink for seamless HPC integration, Dagger for reproducible CI/CD, and the i-twin-AI library for scalable distributed AI, the team has engineered a robust framework. The standout innovation is the automated, end-to-end CI/CD pipeline capable of executing inherently distributed AI tests directly on supercomputers, a critical advancement for ensuring the reliability and integrity of scientific AI models.

Heather Calloway (CISO) — STRONG ACCEPT

This KubeCon talk by Ciangottini and Bunino presents a highly credible and operationally sound framework for managing the integrity and reliability of scientific AI models within complex hybrid cloud and HPC environments. While not a traditional cybersecurity presentation, it directly addresses critical institutional risks associated with scientific reproducibility, resource efficiency, and the accuracy of digital twins used in high-stakes applications like environmental forecasting. The robust CI/CD pipeline, leveraging InterLink and Dagger for automated testing on actual supercomputers, demonstrates a clear commitment to accountability and ensures that scientific outputs are trustworthy…

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025