Resilient Multi-Cloud Strategies: Harnessing Kubernetes, Cluster API, and... T. Rahman & J. Mosquera
T. Rahman, J. Mosquera
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In a compelling presentation at KubeCon EU, Javier Mosquera Sanchez and Tazik from New Relic unveiled their innovative approach to managing a massive multi-cloud Kubernetes infrastructure. The talk, titled "Resilient Multi-Cloud Strategies: Harnessing Kubernetes, Cluster API, and Cellular Architecture," detailed how New Relic leverages Cluster API (CAPI) on top of a cellular architecture to achieve unparalleled scalability, resilience, and blast radius isolation for its critical observability platform.

Key moments
- 0:00 Introduction to multi-cloud Kubernetes with cellular architecture
- 2:00 New Relic's massive scale and Kubernetes footprint
- 3:40 Challenges with monolithic infrastructure and huge blast radius
- 4:45 Defining the cell-based architecture concept
- 7:15 New Relic's practical implementation of a cell
- 9:00 Visualizing the evolving cell-based architecture
Resilient Multi-Cloud Strategies: Harnessing Kubernetes, Cluster API, and... T. Rahman & J. Mosquera
Speakers: T. Rahman, Principal Software Engineer, Kubernetes & Multi-cloud Architect at New Relic; J. Mosquera, Senior Software Engineer at New Relic
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=4DjydLH21nM
Overview
In a compelling presentation at KubeCon EU, Javier Mosquera Sanchez and Tazik from New Relic unveiled their innovative approach to managing a massive multi-cloud Kubernetes infrastructure. The talk, titled "Resilient Multi-Cloud Strategies: Harnessing Kubernetes, Cluster API, and Cellular Architecture," detailed how New Relic leverages Cluster API (CAPI) on top of a cellular architecture to achieve unparalleled scalability, resilience, and blast radius isolation for its critical observability platform.
The core challenge addressed was the inherent complexity and risk associated with operating a vast, monolithic infrastructure supporting a global observability platform that processes exabytes of data daily. New Relic's solution involves decomposing this monolith into independent, self-contained units—or "cells"—each powered by its own Kubernetes cluster and managed through a sophisticated CAPI implementation. This strategy not only facilitates rapid scaling and deployment across diverse cloud providers but also significantly limits the impact of potential incidents, ensuring robust service delivery for its extensive customer base.
The speakers delved into the technical intricacies of their system, from custom Kubernetes Custom Resource Definitions (CRDs) that model the cellular architecture to unique bootstrapping processes and advanced workload scheduling mechanisms. Their insights offer a blueprint for organizations grappling with similar multi-cloud complexities, highlighting the power of Kubernetes and its ecosystem to build highly resilient and scalable distributed systems.
Background
▶ Watch: Introduction to multi-cloud Kubernetes with cellular architecture (0:00)
New Relic operates a comprehensive observability platform, empowering developers to enhance digital experiences for over 85,000 active customers. The sheer scale of their operations is staggering: processing more than 400 million queries and ingesting approximately 7 petabytes of data daily, accumulating to around 3 exabytes annually. This translates to processing 12 billion events per minute, all underpinned by a formidable infrastructure comprising 280 Kubernetes clusters, over 5,000 pods, and more than 21,000 nodes, distributed across multiple cloud providers and regions. An average cluster at New Relic hosts between 300 and 500 nodes and runs 5,000 to 7,000 pods.
Historically, New Relic ran most of its services as containerized workloads on a monolithic DCOS cluster, with data pipelines heavily relying on a single, large Kafka cluster. This monolithic setup presented significant challenges: scaling was arduous, updates and upgrades were risky, and incidents frequently resulted in a "huge blast radius," impacting all customers simultaneously. Recognizing these limitations, New Relic initiated a multi-year program in 2020 to migrate to the cloud, primarily driven by the need for greater scalability and, crucially, to isolate the blast radius of incidents. This program aligned with a strategic shift towards a cell-based architecture.
In a biological context, a cell is the smallest unit capable of independent life, possessing all necessary resources to perform specific functions, yet interconnected with other cells to form more complex systems. New Relic's cell-based architecture mirrors this concept. A workload is decomposed into self-contained installations, each designed to satisfy operations for a specific "shard"—a subset of a larger dataset, such as a group of users. This design makes each cell an independent unit of scale, inherently limiting the blast radius of any incident to only the data or customers associated with that specific cell. Scaling out simply involves adding more cell instances, but this necessitates a sophisticated cell router—a thin layer responsible for sharding and directing traffic and data between cells. Implementing this architecture also requires a domain-driven design approach to identify and decompose infrastructure into isolated, repeatable patterns, leading to the creation of distinct cell types.
New Relic's specific definition of a cell is scoped to a single cloud provider account (e.g., an AWS account, Azure subscription, or GCP project). Each cell typically contains one Kubernetes cluster, one Kafka cluster, and one Virtual Network (VNet) or Virtual Private Cloud (VPC) for networking. Depending on its type, a cell might also include other resources like data stores or load balancers, ensuring its independence. Cells are peered with other cell types when data exchange is required and are deployed in a multi-availability zone setup for resilience. A critical characteristic is their ephemeral nature; cells are designed to be destroyed and replaced, aligning with a Kubernetes cluster lifecycle of approximately 90 days. This living architecture continuously evolves, allowing for the decomposition of large cell types into smaller, more specialized ones as needs change. The visual representation of this shift from a large blue chunk (monolithic data center) to numerous smaller chunks (individual cells) vividly illustrates the successful transition to a highly distributed, isolated environment.
Key Findings
▶ Watch: Challenges with monolithic infrastructure and huge blast radius (3:40)
New Relic's journey to a resilient multi-cloud Kubernetes infrastructure yielded several pivotal findings and architectural contributions:
- Cellular Architecture for Blast Radius Isolation and Scalability: The successful implementation of a cell-based architecture demonstrably reduced the blast radius of incidents and enabled linear scaling by allowing the independent deployment and management of discrete, self-contained units.
- Cluster API as the Foundation for Multi-Cloud Kubernetes Lifecycle: CAPI proved to be the indispensable abstraction layer for declaratively managing the lifecycle of hundreds of Kubernetes clusters across diverse cloud providers, streamlining operations from a centralized control plane.
- Innovative Self-Managed Cluster Model: Unlike typical CAPI deployments, New Relic engineers engineered a unique process where, post-bootstrapping, CAPI management objects are moved to the target cluster itself. This makes each worker cluster also its own management cluster, enhancing isolation and autonomy while maintaining centralized operational oversight.
- Custom CRDs for Cellular Orchestration: The development of homegrown CRDs to model and manage the cellular architecture allowed New Relic to extend the Kubernetes API, providing a native, declarative way to orchestrate the creation, management, and decommissioning of cells.
- Streamlined Node Provisioning with Machine Pools and Helm: By exposing Machine Pools via a generic Helm chart, developers gained a self-service, efficient mechanism to provision and manage nodes with specific requirements, abstracting away underlying cloud complexities.
- Scheduling Classes for Abstracted Workload Placement: The creation of Scheduling Classes provided a cloud-agnostic, declarative interface for developers to express workload scheduling requirements without needing deep knowledge of node labels, taints, or specific compute providers, simplifying application deployment.
- Hybrid Compute Provider Strategy (CAPI + Carpenter): Integrating Carpenter alongside CAPI for node provisioning enabled significant cost optimization and efficient bin-packing of pods through group-less autoscaling, while still offering the flexibility for teams to opt-out when reliability concerns outweigh cost.
- Standardized, Automated Upgrade Processes: A robust system for control plane and node upgrades, leveraging custom CRDs and CAPI's drift detection features, ensured consistent and reliable updates across the vast fleet of clusters.
Technical Deep Dive
▶ Watch: Defining the cell-based architecture concept (4:45)
New Relic's multi-cloud Kubernetes strategy is a testament to sophisticated engineering, built upon the principles of cellular architecture and a highly customized Cluster API (CAPI) implementation.
Cellular Architecture Implementation
At its core, the cellular architecture provides independent units of scale and failure domains. Each New Relic cell is an isolated environment within a cloud provider account (AWS, Azure, GCP), encompassing a single Kubernetes cluster, a dedicated Kafka cluster, and its own VNet/VPC. Critically, these cells are designed to be ephemeral, with a targeted lifecycle of 90 days, after which they are ideally destroyed and replaced. This ephemeral nature is crucial for security, consistency, and efficient resource utilization, but it poses a significant challenge for stateful workloads. New Relic addresses this by either designing cells to be stateless where possible or by decoupling stateful components into separate, more permanent cell types or dedicated data stores. The data sharding and traffic routing between cells are managed by cell routers, which are custom-built services analyzing incoming traffic headers based on data type and customer ID to direct requests to the appropriate cell. This allows for fine-grained control, such as routing metrics for customer A to Cell Alpha and logs for the same customer to Cell Beta.
Cluster API and Control Plane Management
To manage the lifecycle of over 280 Kubernetes clusters, New Relic utilizes CAPI. Their implementation features a Common and Control Cluster which acts as the central orchestration point. This cluster extends the Kubernetes API with homegrown CRDs that model New Relic's cellular architecture, allowing for declarative management of cell lifecycles. From this central point, the bootstrapping of new Kubernetes clusters is initiated.
A key differentiator in New Relic's approach is their handling of CAPI management objects. Instead of maintaining a remote management cluster for each workload cluster, they perform a clusterctl move operation after the initial bootstrapping phase. This transfers the CAPI objects (such as Cluster, KubeadmControlPlane, MachineDeployment) from the bootstrapping environment into the newly created target Kubernetes cluster itself. This makes each target cluster self-managed, acting as both a worker cluster and its own management cluster. This design enhances isolation, as each cell can manage its own lifecycle, while the Common and Control Cluster retains a centralized point for initiating and overseeing operations, ensuring a streamlined and coherent operational experience across clouds.
The Bootstrapping Process (Example: Azure)
The bootstrapping of a new cell, particularly for a cloud provider like Azure, involves several intricate steps:
- Prerequisite Automation: Before CAPI takes over, a set of homegrown cluster controllers in the Common and Control Cluster monitor the creation of a custom
ClusterCRD object. This triggers prerequisite cell build automation specific to the cloud provider and cell type, setting up necessary cloud resources and configurations within the designated Azure subscription. - Kind Cluster Initiation: A Kubernetes job is initiated, which creates a Kind cluster (Kubernetes in Docker) running in Docker-in-Docker mode. This temporary Kind cluster serves as an ephemeral staging environment.
- CAPI & CAPZ Installation: Within the Kind cluster, CAPI and the cloud-specific provider (e.g., CAPZ for Cluster API Provider Azure) controllers are installed.
- Control Plane and Worker Node Object Creation: Using CAPI, the
KubeadmControlPlaneobject is created to provision the Kubernetes control plane in Azure. Concurrently, initialMachineDeploymentobjects are created to house applications deployed during the cluster bootstrapping process. - Cluster Dependencies: Necessary cluster dependencies and add-ons are installed.
- Target Cluster Self-Management: Once the control plane is ready, a
clusterctl initcommand is executed on the target Azure Kubernetes cluster to deploy the CAPI and CAPZ controllers and their CRDs directly onto it. Subsequently, aclusterctl moveoperation transfers all relevant CAPI objects from the Kind cluster to the target Azure Kubernetes cluster, making it self-managed. - Reconciliation and Readiness: The bootstrapping job concludes. Further reconciliation processes take over, including syncing waves of applications via ArgoCD to make the cluster fully operational. Finally, a suite of internal test suits is run to validate the cluster's readiness for live deployments.
For other cloud providers like AWS and GCP, the process is similar, though the control plane might be hosted (managed by the cloud provider) or self-managed, depending on the specific CAPI provider implementation.
Node Provisioning with Machine Pools
Developers provision nodes for their applications primarily through Machine Pools. New Relic abstracts the complexity of cloud-specific node configurations by exposing a generic Machine Pool Helm chart. Developers use this chart to define their node requirements, which are then deployed as an ArgoCD application within the target cells. The Helm chart is dynamically templated with global values and default settings injected by a Kubernetes webhook, generating the necessary CAPI objects for node creation with the specified features.
Nodes are grouped into different pools:
- General Pool: A multi-tenant pool for teams without specific requirements.
- Dedicated Pools: Created by teams requiring specific node features not available in the general pool or wishing to avoid "noisy neighbor" issues.
- Architecture-Specific Pools: For applications requiring specific CPU architectures (e.g., ARM).
New Relic leverages two compute providers for node management: Cluster API and Carpenter. Carpenter is favored for its efficiency in bin-packing application pods, its group-less autoscaling capabilities, and its ability to optimize for cost by selecting the cheapest available node that meets the configuration. Currently, Carpenter is primarily used in AWS cells, with plans for broader expansion.
Workload Scheduling with Scheduling Classes
To address the growing variety of scheduling requirements and simplify workload placement, New Relic introduced Scheduling Classes. This homegrown, New Relic-specific construct provides a declarative way for developers to express their application's scheduling needs without requiring knowledge of underlying node labels or taints.
Scheduling Classes are implemented as an admission controller in the form of a mutating webhook, running on specific resources within each cell. Its design goals include being cloud-agnostic, building on existing Kubernetes scheduling primitives, offering sane defaults, being deterministic, and allowing for chaining of multiple classes.
When a user deploys an application:
- The webhook runs, validating and potentially mutating the application's manifest.
- If a mutation is required, the application is automatically augmented with appropriate affinity and tolerations based on the specified Scheduling Class.
- The standard Kubernetes scheduler then takes over, using these added constraints to find the optimal node for the application's pods.
For example, an application specifying feature: fu via a Scheduling Class annotation might default to a Carpenter-managed node if it's an AWS cell. The webhook adds the necessary affinity and tolerations, directing the Kubernetes scheduler to a Carpenter node. Conversely, if an application explicitly opts out of Carpenter by specifying a CAPI scheduling class annotation, the webhook adds different constraints, directing the scheduler to a CAPI-managed node pool. This mechanism allows for the co-existence and intelligent utilization of both Carpenter and CAPI-managed node pools within the same cell.
Standardized Upgrade Processes
Managing upgrades across hundreds of clusters is a complex task. New Relic has standardized this by managing all control plane objects via ArgoCD applications within each cell. A homegrown ClusterLifecycle CRD targets groups of cells based on environment labels tracked by the custom Cluster CRD. When the Kubernetes version is updated in the ClusterLifecycle CRD, a command and control controller reconciles this change, introducing a diff in the upstream KubeadmControlPlane object. Upstream CAPI controllers then take over, orchestrating the upgrade process.
For node refreshes (upgrades or security patches), a WorkerConfiguration CRD tracks attributes like AMI versions within each cell. New Relic enables drift detection on CAPI objects such as MachineDeployments, AWSMachinePools, and MachinePools. They also leverage Carpenter's built-in node drift feature, avoiding redundant work where Carpenter already provides robust capabilities. This ensures that nodes are consistently updated and aligned with security and performance requirements.
Challenges and Learnings
The speakers shared several hard-won lessons:
- CAPI Version Management: Different cloud provider implementations of CAPI often reference different CAPI versions, making management challenging. Maintaining forks for bespoke features and keeping them synced with upstream CAPI is a continuous effort.
- API Version Parity: Managing automation when clusters use different API versions of CAPI and cloud provider objects is complex.
- Hosted vs. Self-Managed Control Planes: Self-hosted clusters offer more control over upgrade cadences and pre-warming control plane components before traffic shifts. Hosted providers tie you to vendor-specific upgrade charters, making version parity across clouds difficult and limiting operational flexibility.
- Standardization vs. Cloud-Specific APIs:
kubeadm-managed clusters across cloud providers allow for greater standardization of automation due to consistent APIs. Relying on cloud provider-specific APIs necessitates bespoke solutions, increasing maintenance burden. While self-hosting offers more control, it also means signing up for additional operational work (e.g., etcd backups, disaster recovery) that a cloud provider would typically handle. - Carpenter Adoption: While Carpenter offers significant benefits in group-less autoscaling, efficient bin-packing, and cost optimization (choosing the cheapest node), teams sometimes need to opt out via Scheduling Classes if the consolidation rate impacts their application's reliability.
Demo / Proof of Concept
▶ Watch: New Relic's practical implementation of a cell (7:15)
The talk focused on presenting the architectural design, implementation details, and operational learnings from New Relic's production environment. While no live demonstration or specific proof-of-concept was detailed during the presentation, the speakers thoroughly described how their intricate system works, illustrating the practical application of their multi-cloud Kubernetes strategy at a significant scale. The visual representation of data traffic shifting from a monolithic data center to numerous isolated cells served as a powerful testament to the architecture's real-world impact.
Defensive Implications
▶ Watch: Visualizing the evolving cell-based architecture (9:00)
New Relic's cellular architecture and sophisticated Kubernetes management strategy have profound defensive implications, fundamentally enhancing the security and resilience of their observability platform:
- Blast Radius Isolation: The most significant defensive gain is the dramatic reduction of the blast radius. By decomposing the infrastructure into independent cells, an incident or security breach within one cell is contained, preventing it from cascading across the entire system and affecting all customers. This architecture provides inherent segmentation, making it harder for attackers to move laterally across the entire environment.
- Enhanced Resilience and Availability: The ability to rapidly destroy and replace ephemeral cells, coupled with automated bootstrapping and upgrade processes, means that compromised or failing cells can be quickly decommissioned and rebuilt. This "cattle not pets" approach to infrastructure management significantly improves overall system resilience and availability, even in the face of security incidents or infrastructure failures.
- Consistent Security Posture: Leveraging Cluster API and standardized Helm charts for node provisioning and workload scheduling ensures a consistent and hardened security posture across all Kubernetes clusters, regardless of the underlying cloud provider. This consistency simplifies auditing and compliance efforts.
- Simplified Patching and Upgrades: The automated and standardized processes for control plane and node upgrades (including security patches) ensure that vulnerabilities are addressed promptly and consistently across the vast fleet. This reduces the attack surface by minimizing the time clusters remain on outdated or vulnerable software versions.
- Controlled Workload Placement: Scheduling Classes provide a robust mechanism for enforcing workload isolation and resource allocation. By abstracting node characteristics, New Relic can ensure that sensitive workloads land on appropriately secured or isolated nodes, reducing the risk of unauthorized access or interference from other applications.
- Multi-Cloud Agility for DR: The multi-cloud deployment strategy, managed by a unified CAPI control plane, provides inherent disaster recovery capabilities. If one cloud provider experiences a widespread outage, workloads can theoretically be shifted or spun up in another cloud, ensuring business continuity.
- Ephemeral Infrastructure for Reduced Persistence: The ephemeral nature of cells means that any persistent compromise within a cell is temporary. When a cell is decommissioned and replaced, any malicious persistence within that specific environment is eradicated, making it harder for attackers to maintain long-term footholds.
Key Takeaways
- Cellular Architecture is Key for Scale and Resilience: Decomposing monolithic systems into independent, ephemeral cells dramatically limits the blast radius of incidents and provides a highly scalable, repeatable pattern for multi-cloud deployments.
- Cluster API is Essential for Multi-Cloud K8s Management: CAPI provides the declarative, unified control plane necessary to manage the lifecycle of hundreds of Kubernetes clusters across diverse cloud providers efficiently and consistently.
- Self-Managed Clusters Offer Enhanced Isolation: New Relic's unique approach of moving CAPI management objects to the target cluster post-bootstrap creates self-managing, isolated units while retaining centralized operational oversight.
- Custom Abstractions Drive Developer Experience: Homegrown CRDs, generic Machine Pool Helm charts, and Scheduling Classes abstract away multi-cloud complexities, empowering developers to provision resources and schedule workloads declaratively and efficiently.
- Hybrid Compute Strategies Optimize Resources: Combining CAPI with tools like Carpenter for group-less autoscaling and cost optimization allows organizations to balance performance, reliability, and cost-efficiency in node provisioning.
- Operational Complexity Requires Standardization: Managing diverse CAPI versions, hosted vs. self-managed control planes, and cloud-specific APIs across a multi-cloud fleet necessitates rigorous standardization, automated upgrade processes, and careful consideration of tradeoffs.
About the Speaker(s)
Javier Mosquera Sanchez is a Principal Software Engineer and a Kubernetes and Multi-cloud Architect at New Relic. His expertise lies in designing and implementing large-scale, resilient infrastructure solutions, particularly within multi-cloud Kubernetes environments. He played a pivotal role in New Relic's multi-year program to migrate to a cloud-native, cell-based architecture.
Tazik is a Senior Software Engineer at New Relic, working alongside Javier in the same team. His contributions focus on the technical implementation and operational aspects of New Relic's Kubernetes infrastructure, including the bootstrapping processes, node provisioning, and advanced scheduling mechanisms that enable the company's multi-cloud strategy.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This KubeCon presentation from New Relic delivers a brutally honest and technically deep dive into their multi-cloud Kubernetes strategy, managing over 280 clusters with a cellular architecture. The speakers detail their innovative use of Cluster API (CAPI), custom CRDs, a unique self-managed cluster model, and sophisticated scheduling mechanisms to achieve unparalleled resilience and blast radius isolation at immense scale. It's a masterclass in large-scale distributed systems engineering, offering concrete, actionable insights for anyone grappling with similar multi-cloud complexities.
Heather Calloway (CISO) — STRONG ACCEPT
This presentation from New Relic demonstrates a highly sophisticated and effective strategy for building a resilient multi-cloud infrastructure, directly addressing critical challenges of scale and incident blast radius. By leveraging a cellular architecture and a customized Cluster API implementation, New Relic has engineered an environment where risk is contained, systems are ephemeral, and operational resilience is baked into the design. For security leaders, this talk offers a compelling blueprint for how advanced platform engineering can fundamentally transform an organization's security posture and ensure business continuity in the face of inevitable incidents.