Smooth Scaling With the OpAMP Supervisor: Managing Thousands of OpenTe... Evan Bradley & Andy Keller
Evan Bradley, Andy Keller
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In the dynamic landscape of modern distributed systems, managing vast fleets of observability agents is a critical yet complex challenge. The "Smooth Scaling With the OpAMP Supervisor" talk at KubeCon EU, presented by Andy Keller from BindPlane and Evan Bradley from Datadog, delved into the Open Agent Management Protocol (OpAMP) and its pivotal role in remotely managing thousands of OpenTelemetry Collectors. This session provided an essential update on the OpAMP protocol, showcased the OpAMP Supervisor as a robust solution for agent management, and demonstrated its capabilities in real-world scenarios.

Key moments
- 0:00 Talk introduction and agenda overview
- 0:51 Defining OpAMP: Open Agent Management Protocol
- 2:46 Latest OpAMP features: heartbeats, custom messages
- 5:07 Available components for custom collector distributions
- 6:11 OpAMP implementation considerations for OpenTelemetry Collector
- 6:33 Architectural choice: OpAMP supervisor vs. embedded in collector
Smooth Scaling With the OpAMP Supervisor: Managing Thousands of OpenTelemetry Agents
Speakers: Evan Bradley, Engineer, Datadog; Andy Keller, Principal Engineer, BindPlane
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=g8rtqqNTL9Q
Overview
In the dynamic landscape of modern distributed systems, managing vast fleets of observability agents is a critical yet complex challenge. The "Smooth Scaling With the OpAMP Supervisor" talk at KubeCon EU, presented by Andy Keller from BindPlane and Evan Bradley from Datadog, delved into the Open Agent Management Protocol (OpAMP) and its pivotal role in remotely managing thousands of OpenTelemetry Collectors. This session provided an essential update on the OpAMP protocol, showcased the OpAMP Supervisor as a robust solution for agent management, and demonstrated its capabilities in real-world scenarios.
The speakers meticulously explained how OpAMP, an agent-agnostic protocol, facilitates centralized control over telemetry agents, enabling tasks such as configuration updates, status reporting, and even binary upgrades. The introduction of the OpAMP Supervisor specifically addresses the complexities of integrating OpAMP with the OpenTelemetry Collector, offering a powerful and simplified approach to managing custom collector distributions at scale. This talk is highly relevant for platform engineers, SREs, and anyone responsible for deploying and maintaining observability agents across large-scale infrastructures, offering insights into enhancing operational efficiency and ensuring consistent telemetry collection.
Background
▶ Watch: Talk introduction and agenda overview (0:00)
The problem of remotely managing large fleets of observability agents has long plagued organizations operating at scale. Traditional methods often involve manual configuration, custom scripting, or proprietary solutions, leading to inconsistencies, operational overhead, and security vulnerabilities. To address this, the OpenTelemetry community developed the Open Agent Management Protocol (OpAMP), a standardized network protocol designed for the remote management of these agents. The specification for OpAMP is openly available within the OpenTelemetry GitHub organization, alongside a Go implementation that provides client and server SDKs for easy integration.
OpAMP's core function is to establish a robust communication channel between telemetry agents and an agent management server. This server acts as a central command and control interface, capable of coordinating agents, receiving their status reports, sending them new configurations, and even initiating package upgrades. The protocol supports both HTTP, where agents periodically poll the server for updates, and WebSockets, which provide a persistent connection for push-based messaging. A key design principle of OpAMP is its agent-agnostic nature, meaning it is not exclusively tied to OpenTelemetry Collectors but aims to support various observability agents in the future. It also allows for partial implementations, enabling agents and servers to negotiate and enable only the capabilities they both support, fostering extensibility without requiring universal adoption of every new feature. Andy Keller previously gave a more foundational talk on OpAMP at KubeCon North America in 2023, setting the stage for the updates and deeper technical insights presented in this session.
While OpAMP provides the protocol, integrating it directly into the OpenTelemetry Collector presents its own set of challenges, particularly concerning lifecycle management and crash recovery. To circumvent these complexities and offer a more powerful, flexible solution, the concept of the OpAMP Supervisor was introduced. The supervisor acts as an intermediary, running as a separate process that handles all OpAMP communication with the server. It then manages the OpenTelemetry Collector process, writing configurations to disk and starting/stopping the collector as needed. This architectural decision simplifies the collector's design, making it more resilient and enabling advanced features like seamless binary upgrades, which would be significantly more intricate if handled within the collector itself.
Key Findings
▶ Watch: Latest OpAMP features: heartbeats, custom messages (2:46)
The talk highlighted several significant advancements and architectural decisions crucial for effective agent management at scale. These findings span both improvements to the OpAMP protocol itself and the practical implementation provided by the OpAMP Supervisor.
OpAMP Protocol Enhancements
- Heartbeats for Persistent Connections: A critical improvement to the OpAMP specification is the introduction of heartbeats. Load balancers commonly terminate idle WebSocket connections, which can disrupt long-running agent-server communication. Heartbeats address this by periodically sending empty messages to keep the connection alive. Agents supporting heartbeats signal this capability, and servers, being more aware of network infrastructure requirements, respond with a preferred heartbeat interval, ensuring robust and stable connections.
- Custom Messages for Extensibility: OpAMP now supports custom messages, allowing for the implementation of features not explicitly defined within the core protocol. This extensibility is vital for vendors or specific use cases that require unique agent-server interactions. To use custom messages, both the agent and server must declare support for a specific custom capability. The talk provided an example of a service discovery feature, where a server could request available services from an agent, and the agent would respond with a list. This mechanism promotes interoperability while allowing for vendor-specific innovations, encouraging vendors to publish their custom message specifications.
- Available Components for Configuration Compatibility: The most recent feature, available components, significantly enhances the server's ability to manage custom OpenTelemetry Collector distributions. When an agent connects, it sends a hash of its available components to the server. If the server requires the full list, it sets a flag in its response, prompting the agent to send a detailed inventory of its included receivers, processors, and exporters, along with their versions (e.g., specific versions of filelog and OTLP receivers). This allows the server to intelligently determine if a proposed configuration is compatible with a particular agent's build, preventing deployment errors and providing tailored configuration options.
The OpAMP Supervisor's Role
The OpAMP Supervisor emerged as a central finding, providing a practical and powerful solution for managing OpenTelemetry Collectors using OpAMP.
- Proxying and Process Management: The supervisor acts as a dedicated intermediary between the OpAMP server and the OpenTelemetry Collector. It handles all OpAMP protocol communication, receiving configurations and commands from the server. Crucially, it manages the collector's lifecycle: writing received configurations to disk, starting, stopping, and restarting the collector process as needed. This separation of concerns simplifies the collector's design and enhances its stability.
- Comprehensive Agent Telemetry and Status: The supervisor gathers extensive information from the collector, including identification attributes, supported components, the collector's currently resolved configuration (even if fetched from other sources), and pipeline liveness information. It also captures collector logs from
stdoutand forwards them, along with its own telemetry and the collector's telemetry, to a chosen telemetry backend. This provides unparalleled visibility into the health and operational state of deployed agents.
- Enabling Custom Collector Distributions: The supervisor, in conjunction with the OpenTelemetry Collector Builder (OCB), empowers users to create and manage their own custom collector distributions. By simply including the
opamp extensionin an OCB manifest, any custom collector binary can become OpAMP-manageable. This flexibility is vital for organizations that need specific sets of components in their collectors, allowing them to tailor distributions while still benefiting from centralized management.
These findings collectively underscore a significant leap forward in the remote management of observability agents, making large-scale deployments more manageable, robust, and adaptable to evolving requirements.
Technical Deep Dive
▶ Watch: Available components for custom collector distributions (5:07)
The technical core of the talk revolved around the intricacies of the OpAMP protocol and the architecture of the OpAMP Supervisor. Understanding these details is crucial for anyone looking to implement or leverage this system effectively.
OpAMP Protocol Details
OpAMP defines two primary message types: agent-to-server and server-to-agent.
- Agent-to-server messages provide the server with critical information about the agent's identity, health, current activities, and capabilities. These messages allow the server to answer questions like "What agents do I have?", "Are they healthy?", and "What can they do?".
- Server-to-agent messages are used to instruct and configure agents. This includes sending new configurations, providing packages for binary upgrades or modifications, and issuing specific commands.
The protocol supports both HTTP for polling-based updates and WebSockets for persistent, push-based communication. WebSockets are generally preferred for real-time management due to their lower latency and efficiency, though they introduce the need for heartbeat mechanisms to prevent idle connection termination by load balancers. The server dictates the preferred heartbeat interval, indicating its awareness of network infrastructure constraints.
The custom messages feature is implemented by defining a custom capability. For instance, a capability named com.example.discovery could be defined. This capability would then specify custom message types, such as a "discovery request" from the server and a "discovery response" from the agent. Both the agent and server must explicitly declare support for com.example.discovery for these messages to be exchanged. This modularity ensures that new features can be added without burdening all OpAMP implementations, promoting extensibility.
The available components feature leverages hashing to efficiently manage potentially large lists of components. An agent initially sends a hash of its component inventory. If the server determines it needs the full, detailed list – which can contain over a hundred components like specific versions of filelog and OTLP receivers – it requests it. This allows the server to perform detailed compatibility checks, ensuring that a proposed configuration only uses components actually present in the agent's build.
OpAMP Supervisor Architecture
The OpAMP Supervisor is designed as a lightweight, independent process that orchestrates the OpenTelemetry Collector. Its primary responsibilities include:
- OpAMP Communication: The supervisor establishes and maintains the OpAMP connection with the agent management server. It handles all protocol-level interactions, receiving configuration updates, commands, and other messages from the server.
- Collector Process Management: The supervisor is responsible for the lifecycle of the OpenTelemetry Collector. When it receives a new configuration, it writes this configuration to a designated storage directory on disk. It then uses the collector's command-line arguments to point the collector to this new configuration file, effectively initiating or restarting the collector process. This approach simplifies crash recovery; if the collector crashes, the supervisor can simply restart it with the last known good configuration.
- Telemetry and Status Reporting: The supervisor acts as a conduit for collector telemetry and status. It gathers various attributes identifying the collector, its supported components, its resolved configuration (which might include configurations fetched by the collector itself from network sources), and liveness information about its pipelines. It also captures logs from the collector's
stdoutstream. All this data, including the supervisor's own telemetry, can be forwarded to a configured telemetry backend, providing comprehensive self-observability. - Enabling OpAMP in Collectors: To be managed by a supervisor, an OpenTelemetry Collector distribution needs to include the OpAMP extension. While older collector versions might work with caveats, using a recent collector framework (e.g.,
v122+) is recommended. The OpenTelemetry Collector Builder (OCB) streamlines the creation of custom collector binaries. A simple manifest file, specifying desired components along with theopamp extension, is all that's required to build an OpAMP-manageable collector.
Custom Message Implementation Flow
When a custom message is sent from the OpAMP server, the flow is as follows:
- The server sends the custom message to the supervisor via the OpAMP connection.
- The supervisor forwards this message to the OpAMP extension running within the OpenTelemetry Collector.
- The OpAMP extension maintains a registry of custom capabilities. It dispatches the received custom message to the specific component that registered to handle that custom capability.
- Components wishing to use custom messages register with the OpAMP extension, providing the capability name and options. This registration returns a custom capability handler, which provides access to channels for receiving and sending custom messages.
- An example given was an extension implementing discovery. It would register its
com.example.discoverycapability on startup and use the handler to exchange discovery requests and responses with the OpAMP server. The AWS S3 receiver was cited as a real-world example that uses custom messages to send status updates.
This detailed technical breakdown illustrates how OpAMP and the supervisor work in concert to provide a powerful, flexible, and scalable solution for managing complex observability agent deployments.
Demo / Proof of Concept
▶ Watch: OpAMP implementation considerations for OpenTelemetry Collector (6:11)
The talk featured two distinct demonstrations, effectively showcasing the OpAMP Supervisor's capabilities in different operational contexts.
The first demonstration, led by Evan Bradley, focused on a barebones supervisor setup using the example OpAMP server available in the OpAMP Go repository. The setup comprised a supervisor binary, a custom collector binary, and a supervisor configuration file. The configuration specified the opamp_server URL (with an explicit warning against insecure_skip_verify in production), enabled remote_config (which is disabled by default for security), and defined the binary_path for the collector and a storage_dir for runtime information.
Initially, when the supervisor was run, the server interface showed a UUID for the connected agent but indicated that the collector was "not running." This illustrated a key optimization: if no configuration is present from the server, the supervisor conserves resources by not starting the collector process. Evan then used the server interface to send a "dead simple no pipeline" configuration. Upon receiving this, the supervisor restarted the collector with the new configuration. The server then reflected the collector's "up" status, and the logs confirmed the configuration was applied. A crucial point highlighted was that while remote configuration is supported, it currently involves a full collector restart rather than hot reloading, a feature planned for the future. The demo also underscored the power of fleet management, where attributes like host.arch could be used to selectively apply configurations to specific subsets of agents.
The second demonstration, presented by Andy Keller, provided a more advanced, production-oriented view using BindPlane, a telemetry pipeline management tool built by his company. This demo involved managing a fleet of 50 OpenTelemetry Collectors running in Docker and one OpenTelemetry Collector running on his Mac. These agents were configured to send metrics via OTLP to a custom-built gateway, which then forwarded them to Dynatrace.
The BindPlane interface displayed a comprehensive list of connected agents, their statuses, and detailed information reported via OpAMP. Andy then modified a configuration to filter metrics, specifically to include only process metrics, which involved adding a processor to the configuration. He initiated a "rollout" of this new configuration to all 50 agents. The demonstration visibly showed the agents receiving the new configuration, restarting, and applying the changes.
A particularly impactful aspect of this demo was how it showcased the available components feature. Andy highlighted a custom-built gateway collector which was "super stripped down," having only an OTLP receiver, OTLP exporter, and a few processors. When attempting to apply configurations that required components not present in this custom build (e.g., an Apache Spark receiver or an AWS S3 exporter), BindPlane intelligently identified the incompatibility based on the available components information sent by the collector via OpAMP. This prevented the deployment of non-functional configurations and demonstrated the power of the server knowing precisely what capabilities each agent possesses. Both demos effectively illustrated OpAMP and the supervisor's practical utility in remotely configuring, monitoring, and managing diverse OpenTelemetry Collector deployments.
Defensive Implications
▶ Watch: Architectural choice: OpAMP supervisor vs. embedded in collector (6:33)
The OpAMP Supervisor and the underlying OpAMP protocol introduce significant advantages for security and operational defense within large-scale observability infrastructures. However, they also surface new considerations that defenders must address.
Enhanced Control and Visibility
- Centralized Configuration Management: OpAMP enables a single source of truth for agent configurations. This centralized control reduces the risk of misconfigurations, shadow IT, and unauthorized changes across a large fleet. Defenders can ensure that all agents adhere to security policies, such as specific data filtering, redaction, or secure endpoint configurations, by pushing standardized configurations from a trusted server.
- Improved Agent Health Monitoring: Agents reporting their status, resolved configuration, and liveness information via OpAMP provide unprecedented visibility into their operational state. This allows defenders to quickly identify agents that are unhealthy, disconnected, or running outdated configurations, potentially indicating a compromise or a failure to apply security updates.
- Configuration Compatibility Assurance: The "available components" feature is a powerful defensive tool. By knowing exactly which components (receivers, processors, exporters) are built into an agent, the management server can prevent the deployment of configurations that leverage unapproved or vulnerable components. This helps maintain a hardened agent posture and ensures only necessary functionalities are active.
- Streamlined Patching and Upgrades (Future): While currently in development, the ability to perform remote binary upgrades through the supervisor will be a game-changer for defensive operations. This allows for rapid deployment of security patches, vulnerability fixes, and updated agent versions across the entire fleet, significantly reducing the window of exposure to known exploits.
Security Considerations and Challenges
- Secure Communication: The protocol supports TLS for secure communication between agents and the management server. It is paramount that
insecure_skip_verify(as seen in the demo) is never used in production environments. Proper TLS certificate management, including rotation and revocation, is essential to prevent man-in-the-middle attacks. - Access Control to the Management Server: The OpAMP management server becomes a highly privileged component, capable of controlling the entire observability fleet. Robust access controls, authentication, and authorization mechanisms are critical to prevent unauthorized users from issuing malicious configurations or commands.
- Supervisor Hardening: The supervisor process itself must be secured. It runs with privileges to manage the collector binary and write configurations to disk. Standard host hardening practices, including least privilege, secure storage for configuration files, and monitoring of supervisor logs, are essential.
- "Thundering Herd" and Scalability: While not a direct security vulnerability, the "thundering herd problem" (where thousands of agents simultaneously reconnect after a server restart) presents an operational challenge that could indirectly impact security. If the server becomes overloaded, agents might fail to receive critical configuration updates or commands, leaving them in an unmanaged or vulnerable state. Server implementations must robustly handle these scenarios using techniques like exponential backoff and connection throttling.
- Managing Disconnected Agents: The challenge of identifying truly disconnected agents versus temporary outages or replaced container instances is important. An agent that persistently fails to connect could be indicative of a network issue, a misconfiguration, or even an attempted tampering. Robust monitoring and alerting around agent connectivity are crucial.
- Configuration Rollout Strategy: For large fleets, the strategy for rolling out new configurations is critical. "Do you do it all at once? Which ones do you update first? What if you encounter errors?" These questions, posed by the speakers, highlight the need for phased rollouts, canary deployments, and automated rollback mechanisms to mitigate the blast radius of a faulty configuration update, which could inadvertently create a security blind spot or operational outage.
In summary, the OpAMP Supervisor provides powerful tools for centralizing, automating, and securing observability agent management. However, its deployment requires careful consideration of the security implications of its privileged position and the robust implementation of secure operational practices around the management server and the supervisor itself.
Key Takeaways
- OpAMP is the Standard for Agent Management: The Open Agent Management Protocol (OpAMP) provides a standardized, agent-agnostic network protocol for remotely managing large fleets of observability agents, offering capabilities for status reporting, configuration updates, and binary upgrades.
- The OpAMP Supervisor Simplifies Collector Management: The OpAMP Supervisor acts as a crucial intermediary, proxying OpAMP communication to the OpenTelemetry Collector. This design simplifies collector integration, enhances stability, and enables features like crash recovery and easier binary upgrades without embedding OpAMP logic directly into the collector.
- Protocol Enhancements Boost Flexibility and Robustness: Recent OpAMP updates, including heartbeats for persistent WebSocket connections, custom messages for vendor-specific extensions (e.g., service discovery), and available components for intelligent configuration compatibility checks, significantly improve the protocol's utility and adaptability.
- Custom Collector Distributions are Fully Manageable: With the OpAMP extension and the OpenTelemetry Collector Builder (OCB), users can easily create custom collector binaries tailored to specific needs and fully manage them via OpAMP, allowing servers to verify configuration compatibility based on an agent's built components.
- Remote Configuration Requires Restarts (for now): While the OpAMP Supervisor enables remote configuration of OpenTelemetry Collectors, applying new configurations currently necessitates a full collector process restart. Hot reloading, which would allow configuration changes without a full restart, is a recognized future goal for enhanced operational smoothness.
- Scalability Challenges are Being Addressed: Managing thousands of agents introduces complex challenges such as the "thundering herd problem" during server restarts, ensuring consistent configuration across vast fleets, and handling disconnected agents. The community is actively working on solutions for high-scale, high-availability deployments, particularly in Kubernetes environments.
About the Speaker(s)
Andy Keller is a Principal Engineer at BindPlane. He is a key contributor to the OpAMP ecosystem, serving as an approver of the OpAMP specification and a maintainer of its Go implementation. His work at BindPlane involves the architecture and implementation of their telemetry pipeline management platform, directly leveraging and advancing the principles of OpAMP.
Evan Bradley is an Engineer at Datadog and a significant contributor to the OpenTelemetry project. He is a maintainer of the OpenTelemetry Collector and an approver on the Go OpAMP implementation. Evan plays a crucial role in maintaining both the OpAMP Supervisor and the OpAMP extension within the collector, demonstrating his deep involvement in bringing OpAMP capabilities to the OpenTelemetry community.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk by Keller and Bradley delivers a substantive update on OpAMP, specifically highlighting the OpAMP Supervisor as a critical architectural solution for managing large fleets of OpenTelemetry Collectors. It details recent protocol enhancements like heartbeats and available components, demonstrating how these features, combined with the supervisor, enable robust, centralized configuration and lifecycle management at scale. The session offers actionable insights for platform engineers and SREs grappling with distributed observability agent deployments.
Heather Calloway (CISO) — STRONG ACCEPT
This talk presents a critical advancement in managing observability agents at scale, offering a standardized and centralized approach through OpAMP and the OpAMP Supervisor. For any organization grappling with the governance and operational risks of distributed telemetry, this solution provides a clear path to enhanced control, compliance, and resilience. It moves agent management from a chaotic, ad-hoc problem to a structured, auditable, and scalable operation, directly impacting our ability to ensure data integrity and operational visibility.