Standardizing CI/CD Observability With OpenTelemetry: Insights Fr... Dotan Horovits & Adriel Perkins
Dotan Horovits, Adriel Perkins
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
The modern software development lifecycle is increasingly reliant on robust CI/CD pipelines, yet a significant challenge persists: the lack of standardized observability within these critical systems. Dotan Horovits and Adriel Perkins, co-leads of the OpenTelemetry CI/CD Special Interest Group (SIG), presented a compelling case for leveraging OpenTelemetry to bring clarity and consistency to CI/CD pipeline monitoring. Their talk highlighted the widespread frustration among developers and maintainers who struggle to diagnose pipeline issues, often resorting to manual investigation across disparate systems and logs.
Key moments
- 0:00 CI/CD observability pain and lack of standardization
- 2:15 OpenTelemetry SIG for CI/CD observability formed
- 3:00 Understanding OpenTelemetry's Semantic Conventions for CI/CD
- 4:40 Real-world traces from OpenTelemetry CI/CD pipelines
- 6:20 OpenTelemetry's vendor-agnostic approach to CI/CD observability
- 6:50 Live demo of CI/CD observability in action
Standardizing CI/CD Observability With OpenTelemetry: Insights From The Trenches
Speakers: Dotan Horovits, Senior Developer Advocate at AWS Open Source; Adriel Perkins, Principal Engineer at Leatria
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=IvIgsHS5MDk
Overview
The modern software development lifecycle is increasingly reliant on robust CI/CD pipelines, yet a significant challenge persists: the lack of standardized observability within these critical systems. Dotan Horovits and Adriel Perkins, co-leads of the OpenTelemetry CI/CD Special Interest Group (SIG), presented a compelling case for leveraging OpenTelemetry to bring clarity and consistency to CI/CD pipeline monitoring. Their talk highlighted the widespread frustration among developers and maintainers who struggle to diagnose pipeline issues, often resorting to manual investigation across disparate systems and logs.
This presentation served as both a technical deep dive and a call to action, demonstrating how OpenTelemetry's semantic conventions and new specifications are transforming the landscape of CI/CD observability. By providing a common language and standardized instrumentation, the OpenTelemetry CI/CD SIG aims to enable organizations to apply established Site Reliability Engineering (SRE) principles to their build, test, and deployment processes. The speakers showcased real-world examples, illustrating how standardized telemetry can drastically reduce the time to detect and resolve issues, ultimately leading to more stable and efficient software delivery.
The importance of this work cannot be overstated. As software supply chains become more complex and the pace of development accelerates, understanding the health and performance of CI/CD pipelines is paramount. This initiative not only addresses immediate pain points for engineers but also lays the groundwork for enhanced security, improved DORA metrics calculation, and greater confidence in the entire software delivery process, making it a pivotal step towards mature DevSecOps practices.
Background
▶ Watch: CI/CD observability pain and lack of standardization (0:00)
The journey towards standardized CI/CD observability began with a widely recognized pain point: the difficulty in understanding and debugging issues within complex CI/CD pipelines. As Dotan Horovits aptly put it, "CI/CD observability... I've been suffering from that for years." This sentiment resonates across the industry, from individual developers to open-source project maintainers. The traditional approach often involves sifting through fragmented logs, manually correlating events, and engaging in time-consuming investigations across multiple repositories and historical runs—a process that is inefficient and prone to error.
Recognizing this pervasive problem, Horovits was driven to explore standardization, leading to an OpenTelemetry Enhancement Proposal (OTEP) approximately two years prior to the talk. While this initial proposal was ultimately closed, it catalyzed a more significant development: the formation of a dedicated Special Interest Group (SIG) within the OpenTelemetry project, specifically focused on CI/CD observability. Adriel Perkins joined Horovits in co-leading this SIG, which was established over a year before their KubeCon presentation.
The core idea behind this initiative is to extend the power of OpenTelemetry—the de facto standard for instrumenting, emitting, and collecting telemetry data from applications in production environments—to CI/CD pipelines. OpenTelemetry provides a unified approach to traces, metrics, and logs, offering a vendor-agnostic framework for collecting rich observability data. The challenge, however, lay in defining a common language and structure for this data specifically within the CI/CD domain. This led to the central focus of the SIG: developing semantic conventions for CI/CD. Semantic conventions are formally defined as a "common set of semantic attributes which provide meaning to data when collecting, producing, and consuming it." In simpler terms, they establish a universal vocabulary for describing telemetry, ensuring consistency regardless of the underlying CI/CD vendor (e.g., GitHub Actions, GitLab CI) or the observability backend used. This consistency is crucial for enabling effective analysis, cross-system correlation, and the application of SRE fundamentals to the entire software development lifecycle.
Key Findings
▶ Watch: Understanding OpenTelemetry's Semantic Conventions for CI/CD (3:00)
The OpenTelemetry CI/CD SIG has made significant strides in standardizing observability for CI/CD pipelines, delivering several key contributions:
- CI/CD Pipeline Attributes: One of the foundational achievements is the definition of attributes for describing CI/CD pipelines themselves. These attributes allow for consistent identification and categorization of pipelines, making it easier to track their execution and performance across different systems.
- Deployment Attributes: The SIG has introduced specific attributes to describe deployments. This is particularly valuable for calculating crucial DORA metrics (Deployment Frequency, Lead Time for Changes, Mean Time to Recover, Change Failure Rate), which are essential for assessing software delivery performance. While some attribute names evolved during the experimental phase (e.g.,
deployment.environment namewas updated), the overall goal is to provide a standardized way to model deployment events. - Version Control System (VCS) Attributes and Metrics: Recognizing the integral role of VCS in CI/CD, the SIG developed attributes to describe VCS operations. Complementing these are VCS metrics, which provide insights into developer workflow efficiency, such as the average time to approval and average time to merge for pull requests. These metrics highlight bottlenecks and areas for process improvement within repositories.
- Test Attributes: To bring observability into the testing phase, attributes for describing test suites, test cases, and their outcomes have been defined. This enables deeper insights into testing performance and reliability within pipelines.
- Artifact Attributes and Attestations: A critical development for modern software supply chain security is the introduction of artifact attributes, particularly attestations. These attributes align closely with initiatives like SLSA (Supply-chain Levels for Software Artifacts), creating a bridge between observability and security. By standardizing how artifacts and their attestations are described, organizations can enhance the verifiability and integrity of their software components.
- Environment Variable Context Propagation Specification: Beyond semantic conventions, a major breakthrough is the new specification for propagating context and baggage over environment variables. While network-based context propagation is well-defined for microservices (e.g., HTTP, gRPC), CI/CD pipelines frequently involve processes that spawn subprocesses without direct network communication. This specification, which was officially merged shortly before the talk, addresses a long-standing challenge by enabling native instrumentation and ensuring that observability context (like trace IDs) is maintained across these subprocess boundaries, critical for accurate end-to-end tracing within a pipeline. This is a huge enabler for tools like Open Tofu and Terraform to natively support traces.
These findings collectively provide a comprehensive framework for instrumenting, collecting, and analyzing CI/CD telemetry, moving the industry closer to a unified and actionable view of its software delivery processes.
Technical Deep Dive
▶ Watch: Real-world traces from OpenTelemetry CI/CD pipelines (4:40)
The core of the OpenTelemetry CI/CD SIG's work revolves around two fundamental technical pillars: semantic conventions and the environment variable context propagation specification. Both are crucial for addressing the unique observability challenges within CI/CD environments.
Semantic Conventions: A Lingua Franca for Telemetry
Semantic conventions are the bedrock of consistent observability. In essence, they define a common language for describing telemetry data—whether it's traces, metrics, or logs. For CI/CD, this means establishing standardized names, types, and expected values for attributes attached to pipeline runs, jobs, tasks, deployments, and VCS operations.
Consider a CI/CD pipeline run. Instead of each vendor or custom script emitting data with arbitrary field names like build_id, job_name, or repo_url, semantic conventions dictate a unified set of attributes such as ci.pipeline.id, ci.job.name, and vcs.repository.url. This standardization offers several key benefits:
- Vendor Agnosticism: Telemetry collected from GitHub Actions, GitLab CI, Jenkins, or Argo CD can be interpreted consistently by any OTLP (OpenTelemetry Protocol)-compliant backend (e.g., Honeycomb, Sigma, Jaeger, Prometheus). This frees organizations from vendor lock-in for their observability stack.
- Simplified Analysis: Engineers can write queries and build dashboards that work across all their CI/CD systems, significantly reducing the cognitive load and time required for analysis and troubleshooting. If a spike in pipeline duration is observed, a query for
ci.pipeline.durationwill yield comparable results regardless of the source. - Enhanced Correlation: By standardizing attribute names and relationships between different telemetry signals (e.g., a trace span for a build job can link to related deployment metrics), it becomes easier to correlate events across the entire software development lifecycle.
- Automated Tooling: Standardized data enables the development of automated tools and AI/ML models that can understand and act upon CI/CD telemetry more effectively, leading to proactive issue detection and resolution.
The SIG has explicitly defined conventions for various aspects:
- CI/CD Pipeline Attributes: Describing the overall pipeline, its ID, name, trigger, and status.
- Deployment Attributes: Capturing details about deployments, including the environment, service, and outcome, which directly feeds into DORA metrics calculation.
- VCS Attributes: Providing context about the version control system, such as repository name, commit SHA, branch, and pull request details. These are crucial for linking pipeline runs back to specific code changes.
- Test Attributes: Detailing test suites, individual test cases, and their execution results.
- Artifact Attributes: Describing software artifacts produced by the pipeline, including vital attestations that are key for SLSA compliance and supply chain security.
Environment Variable Context Propagation Specification: Bridging Subprocess Gaps
A unique challenge in CI/CD pipelines is the prevalence of subprocesses that do not communicate over a network. Unlike microservices, where context propagation (e.g., trace IDs) is typically handled via HTTP headers or gRPC metadata, CI/CD steps often involve launching shell scripts, Docker containers, or other processes that inherit environment variables but lack direct network connectivity for distributed tracing.
This limitation has historically made it difficult to maintain a continuous trace across different stages of a pipeline that involve multiple subprocesses. The newly merged environment variable context propagation specification addresses this by defining a standardized mechanism to pass OpenTelemetry context (like trace IDs and span IDs) using environment variables.
When a parent process starts a child process, it can inject the current trace context into predefined environment variables. The child process, upon startup, can then read these environment variables to pick up the existing trace context and continue the trace, creating new spans that are correctly nested under the parent's span. This ensures that a single, coherent trace can span across an entire CI/CD pipeline, even when executed by disparate, non-networked subprocesses.
This specification is a game-changer for native instrumentation in CI/CD tools. Early prototypes, such as the Open Tofu controller runners demonstrated by a colleague at DevOps days, already leverage this concept to provide end-to-end traces for infrastructure provisioning workflows. As the specification gains adoption, more CI/CD tools and build systems are expected to natively integrate OpenTelemetry, significantly enhancing the depth and breadth of observable data.
Furthermore, the SIG is actively developing OpenTelemetry Collector receivers for popular CI/CD platforms like GitHub (via webhooks) and GitLab. These receivers ingest platform-specific events and translate them into standardized OpenTelemetry traces, metrics, and logs based on the defined semantic conventions. This acts as an interim solution, with the ultimate goal being native emission of OpenTelemetry telemetry directly from GitHub, GitLab, and other platforms. The SIG is also exploring receivers for Argo CD workflows and Jenkins, aiming for comprehensive coverage across the CI/CD ecosystem.
Demo / Proof of Concept
▶ Watch: OpenTelemetry's vendor-agnostic approach to CI/CD observability (6:20)
The speakers provided a live demonstration showcasing the practical application of OpenTelemetry CI/CD semantic conventions using real-world data from the OpenTelemetry project's own pipelines. The demo utilized two popular observability backends: Honeycomb and Sigma, highlighting the vendor-agnostic nature of OpenTelemetry.
The demonstration began by displaying a Honeycomb dashboard showing the P50 duration (50th percentile duration) of pipelines over a 14-day period. Adriel Perkins pointed out a clear spike in pipeline durations between the 21st and 23rd of the month, correlating with a recent incident observed by maintainers.
Zooming into this anomalous period, the dashboard grouped data by repository name—a key semantic attribute defined by the SIG. This immediately revealed that the contrib repository was heavily impacted, while others, like Weaver, showed no significant change (likely due to no active pipelines during that time). Filtering the view to only the contrib repository, the spike in duration became even more pronounced.
Next, the demo dived into a specific trace from a long-running "build and test" pipeline within the contrib repository during the incident period. This particular trace showed an alarming duration of "an hour and 12 minutes"—significantly longer than the typical 8 minutes for these pipelines. Within this trace, a specific job, "go vulnerability check," was identified as problematic. This job was part of a matrix build (multiple parallelized or concurrent jobs).
By visually inspecting the trace waterfall, Adriel highlighted increasing gaps between the start times of these parallel jobs as they progressed down the list. This "white empty space" indicated significant queue time or latency before jobs actually began executing, rather than during their active run time. This visual anomaly directly confirmed the behavior maintainers had been observing, allowing for rapid identification of the root cause—likely resource contention or runner availability issues. The ability to directly link from the trace back to the specific task run in GitHub using a semantic attribute was also demonstrated, streamlining the debugging workflow.
The demonstration then switched to Sigma, another observability backend, to underscore OpenTelemetry's vendor neutrality. The same spike in pipeline duration was clearly visible in Sigma's dashboard for the identical time period, reinforcing that the standardized telemetry data is universally interpretable.
Finally, the demo presented Version Control System (VCS) metrics, another set of semantic conventions. Focusing on the contrib repository, the metrics revealed an average "change time to approval" of approximately 10 weeks, with an average "time to merge" of 6 days once approved. For the semantic inventions repository, it was noted that while there's "lots of dialogue" and "it does take a bit of time" for changes to be approved, the metrics provide valuable insight into the actual workflow durations. These metrics offer critical insights into developer experience and process bottlenecks, enabling teams to identify areas for improvement in their pull request review and merge cycles.
The entire demonstration effectively illustrated how standardized OpenTelemetry data, when visualized in compliant backends, can provide clear, actionable insights into CI/CD performance, quickly pinpointing anomalies and their underlying causes.
Defensive Implications
▶ Watch: Live demo of CI/CD observability in action (6:50)
The standardization of CI/CD observability with OpenTelemetry carries profound defensive implications for organizations, enhancing not only their operational resilience but also their security posture.
Firstly, consistent and comprehensive observability of CI/CD pipelines directly aids in rapid incident response and proactive issue detection. As demonstrated, the ability to quickly identify anomalous pipeline durations, pinpoint bottlenecks like excessive queue times, or detect unexpected failures allows SRE and DevOps teams to address issues before they significantly impact production or developer productivity. This shifts the operational paradigm from reactive firefighting to proactive management, reducing Mean Time To Resolution (MTTR) for pipeline-related incidents.
Secondly, the emphasis on deployment attributes and the calculation of DORA metrics provides a quantitative framework for assessing the health of the software delivery process. By consistently tracking metrics like Lead Time for Changes and Change Failure Rate, organizations can identify regressions in their deployment process, understand the impact of new tooling or practices, and ensure that changes are delivered reliably and securely. A sudden increase in change failure rate, for instance, could indicate a new vulnerability introduced in a dependency or a misconfigured security scanner in the pipeline.
Crucially, the introduction of artifact attributes, especially attestations, directly addresses the growing concerns around software supply chain security. By standardizing how artifacts and their security properties (e.g., provenance, SBOMs, scan results) are described and emitted as telemetry, OpenTelemetry creates a verifiable audit trail. This aligns seamlessly with frameworks like SLSA (Supply-chain Levels for Software Artifacts), enabling organizations to build and verify trust in their software components from source to deployment. Defenders can leverage this data to:
- Verify artifact integrity: Ensure that deployed artifacts match their expected attestations and haven't been tampered with.
- Trace provenance: Understand the exact pipeline and steps that produced an artifact, crucial for forensic analysis during a breach.
- Enforce policies: Automatically check if artifacts meet defined security standards (e.g., all dependencies scanned, no critical vulnerabilities) before deployment.
Furthermore, the environment variable context propagation specification enhances the ability to conduct end-to-end tracing across complex CI/CD workflows, including those involving subprocesses. This unbroken chain of traceability is vital for security audits, allowing defenders to follow the execution flow of sensitive operations, identify unauthorized steps, or detect deviations from expected behavior within a pipeline. If a malicious actor injects a rogue step into a pipeline, a complete trace can help quickly identify where and how the compromise occurred.
Finally, the vendor-agnostic nature of OpenTelemetry means that security teams are not locked into proprietary monitoring solutions. They can choose best-of-breed observability backends that offer advanced security analytics capabilities, while still benefiting from a unified data model across their diverse CI/CD tooling. This flexibility is key for adapting to evolving threat landscapes and integrating with existing security information and event management (SIEM) systems.
In summary, standardizing CI/CD observability empowers defenders by providing unprecedented visibility into the software factory, transforming it into a more transparent, auditable, and resilient system against both operational failures and security threats.
Key Takeaways
- Standardized CI/CD Observability is Critical: The lack of consistent observability in CI/CD pipelines is a widespread industry pain point, hindering efficient troubleshooting and SRE practices.
- OpenTelemetry is the Solution: OpenTelemetry provides a vendor-agnostic framework for standardizing telemetry (traces, metrics, logs) across diverse CI/CD platforms.
- Semantic Conventions Provide a Common Language: The OpenTelemetry CI/CD SIG defines semantic conventions for pipeline, deployment, VCS, test, and artifact attributes, enabling consistent interpretation of data across tools and backends.
- New Spec for Subprocess Context Propagation: A crucial specification now allows OpenTelemetry context to be propagated via environment variables, enabling end-to-end tracing in CI/CD environments with subprocesses.
- Enhanced Security and DORA Metrics: Artifact attestations align with supply chain security (SLSA), while deployment attributes facilitate accurate DORA metrics calculation, improving both security and operational performance.
- Practical Insights and Actionable Data: Standardized telemetry allows for quick identification of anomalies (e.g., pipeline duration spikes, queue times) and provides valuable VCS metrics (e.g., time to approval) for process improvement.
About the Speaker(s)
Dotan Horovits is a Senior Developer Advocate at the AWS open-source team and the Chief Evangelist for the OpenSearch project under the Linux Foundation. He is also a CNCF Ambassador and hosts the Open Observability Talks podcast. Dotan is a co-lead of the OpenTelemetry CI/CD SIG, driven by years of experiencing the pain points of CI/CD observability.
Adriel Perkins is a Principal Engineer at Leatria, a consulting company based in the United States. He serves as a co-lead of the OpenTelemetry CI/CD SIG alongside Dotan Horovits, contributing his expertise to standardizing CI/CD observability efforts.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This KubeCon talk by Dotan Horovits and Adriel Perkins from the OpenTelemetry CI/CD SIG presents critical, foundational work for standardizing observability across software delivery pipelines. Their efforts in defining semantic conventions and, crucially, a new environment variable context propagation specification, address a long-standing industry pain point. The presentation offers deep technical insights into how OpenTelemetry can bring consistency, reduce troubleshooting time, and enhance software supply chain security, making it highly impactful for any organization serious about modern DevSecOps practices.
Heather Calloway (CISO) — STRONG ACCEPT
This presentation by Dotan Horovits and Adriel Perkins effectively highlights a critical, long-standing operational and security gap: the fragmented and inconsistent observability of CI/CD pipelines. Their work with the OpenTelemetry CI/CD SIG, particularly the development of semantic conventions and the environment variable context propagation specification, offers a powerful, vendor-agnostic solution. For any CISO, this isn't just a technical discussion; it's a foundational enabler for robust supply chain security, improved incident response, and verifiable institutional accountability within the software delivery lifecycle.