Reliable K8s Resource Submission & Bookkeeping - Tiancheng Yin & Yao Lin, Bloomberg

Tiancheng Yin, Yao Lin, Bloomberg

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk, presented by Tiancheng Yin and Yao Lin from Bloomberg's Workflow Orchestration team, addresses the critical challenges of reliably submitting and tracking Kubernetes resources in a highly available, multi-datacenter environment. Specifically, it delves into the complexities of managing both "runnable" resources like Argo Workflows and "deployable" resources such as ConfigMaps and Secrets, emphasizing the need for robust system resiliency, data consistency, and efficient post-deployment status tracking. The speakers highlight how Bloomberg, a company that values data center resiliency seriously, engineered a sophisticated platform to ensure continuous operation and data integrity even in the event of major outages or maintenance activities.

Watch on YouTube

Visual summary for Reliable K8s Resource Submission & Bookkeeping - Tiancheng Yin & Yao Lin, Bloomberg by Tiancheng Yin, Yao Lin, Bloomberg
Visual summary for Reliable K8s Resource Submission & Bookkeeping - Tiancheng Yin & Yao Lin, Bloomberg by Tiancheng Yin, Yao Lin, Bloomberg

Key moments

  1. 0:00 Introduction and critical data center resiliency requirement
  2. 2:00 Understanding user workflows and abstracting Kubernetes resource types
  3. 4:30 Specific challenges for runnables (submission) and deployables (consistency)
  4. 6:40 Solution: separating API from submitter for reliable runnable submission
  5. 8:00 Solution: ensuring deployable consistency using a source of truth
  6. 9:00 Post-deployment challenges: UI, performance, and historical data

Reliable K8s Resource Submission & Bookkeeping

Speakers: Tiancheng Yin, Workflow Orchestration Team; Yao Lin, Workflow Orchestration Team; Bloomberg

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=NCKrvqFMl8

Overview

This talk, presented by Tiancheng Yin and Yao Lin from Bloomberg's Workflow Orchestration team, addresses the critical challenges of reliably submitting and tracking Kubernetes resources in a highly available, multi-datacenter environment. Specifically, it delves into the complexities of managing both "runnable" resources like Argo Workflows and "deployable" resources such as ConfigMaps and Secrets, emphasizing the need for robust system resiliency, data consistency, and efficient post-deployment status tracking. The speakers highlight how Bloomberg, a company that values data center resiliency seriously, engineered a sophisticated platform to ensure continuous operation and data integrity even in the event of major outages or maintenance activities.

The core problem tackled is how to ensure that user-defined Kubernetes resources are consistently submitted, maintained, and their status accurately reflected across a farm of Kubernetes clusters, spanning multiple data centers, without overloading the Kubernetes API server or compromising performance. The solution presented involves decoupling user-facing APIs from submission logic, utilizing message streams and dedicated services for both resource submission and post-deployment status tracking, and implementing mechanisms to guarantee data consistency and historical transaction traceability. This talk is highly relevant for organizations operating large-scale Kubernetes environments with stringent availability, auditability, and performance requirements, offering practical architectural patterns to overcome common distributed system challenges.

Background

▶ Watch: Introduction and critical data center resiliency requirement (0:00)

Bloomberg operates a highly available container orchestration platform for internal engineers, designed to execute run-to-completion workloads across various use cases, including machine learning pipelines, CI/CD, machine maintenance routines, and financial analysis. This platform serves a critical role, necessitating robust functional requirements like observability, scheduling, eventing, and approval processes. However, the non-functional requirements, particularly data center resiliency, present the most significant challenges. The expectation is that if one data center becomes unavailable due to failure or maintenance, the platform must seamlessly continue functioning from another healthy data center without service interruption to users.

The platform manages various Kubernetes resources, categorized by the speakers into two main types:

  • Runnables: These are resources that execute workflows, such as Argo Workflows, Kubernetes native jobs, or custom jobs. They represent transient computations that need to be scheduled and run. For example, a user might define a multi-step workflow to generate a report, persist it, and send a notification.
  • Deployables: These are static configuration or secret resources, such as ConfigMaps and Secrets, which are expected to be consistently available on applicable clusters for runnables to function correctly. Their consistency across clusters is paramount.

Managing these resources across a farm of clusters, each potentially spanning at least two data centers, introduces several points of failure and complexity. For runnables, the system must intelligently decide which cluster is most suitable for a workload and handle transient errors with sophisticated retry logic. For deployables, the primary concern is guaranteeing consistency during changes and ensuring their availability even if a cluster is rebuilt or goes down.

Beyond submission, post-deployment status tracking and user interaction pose further challenges. If all read requests, especially list operations, directly hit the Kubernetes API server, it can lead to significant CPU load and degrade cluster performance. Furthermore, there's a need for an extra layer for approval and auditing user actions and job executions, potentially driven by regulatory requirements. The inherent difficulty lies in achieving system resiliency and performance on top of individual cluster management, ensuring that user interfaces remain responsive and accurate even when underlying clusters are unavailable.

Key Findings

▶ Watch: Specific challenges for runnables (submission) and deployables (consistency) (4:30)

The Bloomberg team's key findings revolve around a multi-layered architectural approach designed to ensure reliable resource submission and resilient post-deployment status tracking in a highly available, multi-datacenter Kubernetes environment. Their solution fundamentally decouples the user-facing API from the complex backend logic, introducing dedicated services and robust data stores to handle various aspects of resource management.

For resource submission, the critical insight was to separate the user API from the actual "submitter" component. This allows the API to focus on user interaction, auditing, and initial storage, while a specialized submitter service handles the intricacies of cluster selection, error retries, and feature-specific logic. For deployable resources, a "source of truth" database combined with a "syncer" service ensures consistency across all managed clusters, addressing the disaster recovery challenge.

Regarding post-deployment status tracking, the team identified the need for an external, highly available data store to persist the latest object states and execution results. This "inventory" database serves as a resilient cache, allowing user interfaces and APIs to retrieve information with low latency and high availability, even if a Kubernetes cluster is down. To maintain the accuracy of this inventory, an in-cluster "message producer" (watcher) publishes real-time events, which are consumed by a service that updates the inventory. A crucial finding here was the necessity of a snapshot-based reconciliation mechanism to detect and resolve "zombie records"—situations where a resource is deleted from a cluster but remains in the inventory due to missed events. This ensures eventual consistency between the inventory and the actual cluster state.

In essence, the key findings highlight that achieving extreme resilience and performance in a complex Kubernetes ecosystem requires moving beyond direct API server interactions for many operations. It necessitates building a sophisticated control plane that leverages message streams, highly available databases, and intelligent reconciliation services to manage state, ensure consistency, and provide a performant, fault-tolerant user experience.

Technical Deep Dive

▶ Watch: Solution: separating API from submitter for reliable runnable submission (6:40)

The technical solution presented by Bloomberg addresses the dual challenges of reliable resource submission and resilient post-deployment status tracking through a meticulously designed distributed architecture.

Resource Submission Architecture

The submission process is bifurcated based on resource type: runnables and deployables.

Runnables (e.g., Argo Workflows):

  1. API and Audit: The user-facing Workflow API (referred to as the user API) is the initial point of contact. When a user requests a mutation (e.g., to create or update an Argo Workflow), the API first stores this request in an audit database. This database serves a dual purpose: providing an auditable log of all user actions and facilitating any necessary approval processes.
  2. Message Stream for Processing: Instead of directly submitting the runnable to a Kubernetes cluster, the API sends the runnable object to a message stream. This stream acts as a buffer and a communication channel for asynchronous processing.
  3. Submitter Service: A dedicated submitter service, hidden from direct user interaction, continuously consumes messages from this stream. This service encapsulates the complex logic required for reliable submission:
  • Cluster Selection: It determines the most suitable cluster for the workload, considering factors like availability, load, and specific workload requirements.
  • Retry Logic: It implements sophisticated retry mechanisms to handle transient errors encountered during submission to the Kubernetes API server. This ensures that temporary network glitches or API server overloads do not lead to submission failures.
  • Feature Support: The submitter can also incorporate additional business logic, such as verifying the validity of a workflow against a deadline or other policies before final submission.

This separation of concerns ensures that the user API remains responsive, while the heavy lifting of reliable, intelligent submission is handled by a specialized, resilient backend service.

Deployables (e.g., ConfigMaps, Secrets):

  1. Source of Truth and Audit: Similar to runnables, the user API for deployables first records mutation requests in the audit database. Crucially, it also stores the deployable objects in a source of truth database. This database is the authoritative record of all deployable resources that should exist across the clusters.
  2. Syncer Service: A syncer service continuously reads from this source of truth database. Its responsibility is to ensure that the actual state of deployables in the Kubernetes clusters matches the desired state defined in the source of truth.
  • Consistency Guarantee: The syncer actively pushes (or pulls, depending on implementation detail) these deployable objects to all applicable clusters. This ensures that ConfigMaps, Secrets, or ClusterWorkflowTemplates are consistent across the entire farm of clusters.
  • Disaster Recovery: In scenarios where a cluster is rebuilt or experiences a prolonged outage, the syncer can re-populate it with the correct deployable resources from the source of truth, guaranteeing data integrity and availability. The speakers note that while logically presented as a pull model from the syncer to the cluster, in practice, a push model from the syncer to the cluster is often implemented.

Post-Deployment Status Tracking

Once resources are on the cluster, tracking their status and providing a low-latency, resilient interface for users is critical.

Historical Transactions:

  • The existing audit database (used during submission) already preserves a history of all update attempts and user actions related to resources.
  • A user-facing API is exposed to allow users to retrieve these historical transactions, answering questions like "who made what changes to what resources at what time?" This is vital for debugging and auditing purposes.

Real-time Status and UI Resiliency:

To support a user interface that displays current execution results and resource states with high availability and low latency, even if a cluster is down, a separate mechanism is employed:

  1. In-Cluster Message Producer (Watcher): Within each workload cluster, a message producer service is deployed. This producer acts as a Kubernetes watcher, monitoring create, update, and delete events for specific types of Kubernetes resources (e.g., Argo Workflows, ConfigMaps). When an event occurs, the producer publishes the resource information (e.g., the event type, the resource spec, its UID) into a message stream.
  2. Consumer Service: A consumer service continuously processes messages from this stream.
  3. Inventory Database: The consumer service updates a highly available data storage, referred to as the inventory database. This database stores the latest object and execution results for all tracked in-cluster resources.
  • For a create event, a new record is added to the inventory.
  • For an update event, the existing record is modified.
  • For a delete event, the record is removed.
  1. Resilient User-Facing API: User interfaces and APIs then interact with this inventory database to retrieve resource statuses, rather than directly querying the Kubernetes API server. This significantly reduces the load on the API server, provides low-latency responses, and ensures data resiliency, as the inventory remains available even if the underlying Kubernetes cluster is offline.

Concurrency and Zombie Record Problem

A significant challenge arises from the asynchronous nature of this post-deployment tracking: concurrency issues and zombie records.

  • The Problem: Imagine a user deletes a ConfigMap from a cluster. The syncer detects this and deletes it. The in-cluster producer should capture this deletion event and publish it. However, if the producer service crashes, reboots, or experiences an Out-Of-Memory error at that precise moment, the deletion event might be missed. In this scenario, the ConfigMap is deleted from the cluster, but its corresponding record remains in the inventory database, becoming a "zombie record." If a user queries the API about this ConfigMap, the API, relying on the inventory, would incorrectly report that the resource still exists.
  • The Solution: Snapshot Reconciliation: To resolve zombie records and ensure eventual consistency between the inventory database and the actual cluster state, the producer service implements an additional mechanism:
  • Average Snapshot Publication: Besides publishing individual create/update/delete events, the producer periodically publishes an average snapshot of all in-cluster resources. This snapshot contains a list of Kubernetes object UIDs (Unique Identifiers) and associated cluster information for every resource currently existing in the cluster.
  • Consumer Reconciliation: The consumer service receives these snapshots. It then compares the UIDs present in the snapshot with the UIDs of records currently in the inventory database for that cluster.
  • Zombie Detection and Deletion: If a record exists in the inventory database but its corresponding UID is not present in the received snapshot, it indicates a zombie record. The consumer service then proactively deletes this record from the inventory database.

This snapshot-based reconciliation ensures that even if individual events are missed, the inventory database will eventually converge to the true state of the Kubernetes clusters, providing accurate information to users and preventing data inconsistencies.

Demo / Proof of Concept

▶ Watch: Solution: ensuring deployable consistency using a source of truth (8:00)

The talk focused on presenting a detailed architectural design and the underlying principles for achieving reliable Kubernetes resource submission and robust post-deployment status tracking. While the speakers did not present a live, interactive demonstration of the system in action, the comprehensive explanation of their multi-layered services, message streams, and database interactions serves as a conceptual proof of concept. The described architecture for handling runnables and deployables, coupled with the sophisticated reconciliation logic for the inventory database, illustrates the practical implementation of their solutions to critical challenges like data center resiliency and consistency. The detailed explanation of how "zombie records" are identified and corrected through snapshot comparisons, for instance, provides a clear operational model for their proposed system.

Defensive Implications

▶ Watch: Post-deployment challenges: UI, performance, and historical data (9:00)

The architectural patterns and solutions presented by Bloomberg offer significant defensive implications for organizations operating Kubernetes at scale, particularly those with stringent reliability, consistency, and auditability requirements.

  1. Enhanced Resiliency and Disaster Recovery: By decoupling the user-facing API from direct cluster interactions and introducing highly available intermediate layers (message streams, audit database, source of truth database, inventory database), the system becomes inherently more resilient. Defenders can learn to architect their Kubernetes platforms to withstand individual cluster failures or even entire data center outages, ensuring continuous service availability. The ability to rebuild clusters from a "source of truth" for deployables is a critical disaster recovery strategy.
  1. Reduced Load on Kubernetes API Server: Offloading read requests for resource status from the Kubernetes API server to a dedicated, highly available inventory database significantly reduces the operational burden on the control plane. This prevents performance degradation of the core Kubernetes services, allowing them to focus on orchestration tasks. Defenders should consider similar caching layers for frequently accessed, critical state information.
  1. Robust Auditing and Compliance: The central audit database, which logs all user actions and resource mutations before submission, provides an immutable record essential for compliance, forensic analysis, and accountability. This design ensures that every action is traceable, answering "who made what changes to what resources at what time," which is crucial for regulated industries.
  1. Data Consistency and Integrity: The syncer service for deployables ensures that critical configurations (ConfigMaps, Secrets) are consistent across all clusters, preventing configuration drift and potential security vulnerabilities arising from inconsistent environments. The snapshot-based reconciliation for the inventory database is a powerful mechanism to guarantee eventual consistency, mitigating data discrepancies caused by transient failures in event processing. This is vital for maintaining an accurate operational picture and preventing misconfigurations.
  1. Improved Observability and Debugging: The historical transaction data, combined with a reliable inventory of current states, provides superior observability into the platform. This wealth of information significantly aids in debugging complex issues, understanding the lineage of changes, and quickly identifying the root cause of problems without needing to directly query potentially unhealthy clusters.
  1. Scalability and Performance: The asynchronous nature of submission via message streams, coupled with dedicated services for specific tasks (submitter, syncer, consumer), allows the platform to scale horizontally. This architecture ensures that as the number of users, workflows, or clusters grows, the system can handle the increased load without bottlenecking on a single component, such as the Kubernetes API server.

In essence, the Bloomberg approach advocates for building a robust, external control plane around Kubernetes that enhances its native capabilities for enterprise-grade, mission-critical operations. Defenders can adopt these principles to create more stable, secure, and manageable Kubernetes environments.

Key Takeaways

  • Multi-Data Center Resiliency is Paramount: Designing for data center-level failures requires architectural patterns that decouple core services and leverage highly available external data stores and message streams.
  • Decouple API from Submission Logic: Separating the user-facing API from the complex, error-handling submission logic (via a message stream and dedicated submitter service) improves API responsiveness and system robustness.
  • Centralized Source of Truth for Consistency: Maintaining a "source of truth" database for deployable resources (ConfigMaps, Secrets) and using a syncer service ensures consistency across all clusters and aids in disaster recovery.
  • External Inventory for Resilient Status Tracking: An external, highly available inventory database, populated by in-cluster watchers and a consumer service, provides low-latency, resilient status information to users, reducing load on the Kubernetes API server.
  • Snapshot Reconciliation for Eventual Consistency: To mitigate missed events and "zombie records," implement a periodic snapshot-based reconciliation mechanism that compares the inventory with the actual cluster state using UIDs.
  • Auditability is Built-in: Integrating an audit database at the initial stage of resource mutation provides a comprehensive, immutable record of all user actions, crucial for compliance and debugging.

About the Speaker(s)

Tiancheng Yin and Yao Lin are members of the Workflow Orchestration team within Bloomberg's Cloud Native Compute Services. Their team is dedicated to developing and maintaining highly available container orchestration platforms based on Kubernetes for internal engineers at Bloomberg. They focus on enabling various general-use cases, including machine learning pipelines, CI/CD, machine maintenance routines, and financial analysis. Their work emphasizes building robust and resilient systems within the Kubernetes ecosystem to meet Bloomberg's stringent requirements for data center resiliency and operational reliability.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This KubeCon talk from Bloomberg's Workflow Orchestration team tackles the brutal reality of operating Kubernetes at extreme scale across multiple data centers. It's a deep dive into engineering resilience for resource submission and status tracking, solving genuine distributed systems problems like data consistency, auditability, and API server overload. The architectural patterns, particularly the dedicated services, message streams, and the elegant snapshot-based reconciliation for zombie records, demonstrate a mature approach to critical infrastructure. This isn't theoretical fluff; it's a battle-hardened blueprint for keeping complex K8s environments alive and accurate.

Heather Calloway (CISO) — MUST SEE

This session from Bloomberg presents a robust architectural blueprint for achieving extreme resilience, auditability, and data consistency in large-scale, multi-datacenter Kubernetes environments. It meticulously details how to decouple critical functions like resource submission and status tracking from direct API server interaction, leveraging message streams and dedicated data stores. The practical solutions for managing "zombie records" and ensuring a comprehensive audit trail directly address institutional accountability and business continuity, offering a clear path for organizations to operationalize Kubernetes at the highest levels of reliability and compliance.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025