How To Adopt OpenTelemetry in an Enterprise Where Incumbent Vendor Tools Reign Supre... Chris Weldon
Chris Weldon
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In this insightful KubeCon EU talk, Chris Weldon, Director of Platform Engineering at Wolters Kluwer, Tax & Accounting North America, shares the company's transformative journey from a heavily vendor-laden, monitoring-centric environment to a modern observability strategy powered by OpenTelemetry. The presentation details the strategic, cultural, and technical challenges encountered and overcome during this enterprise-wide shift, offering practical advice for organizations grappling with similar modernization efforts.

Key moments
- 0:00 Introduction to Chris Weldon and talk agenda
- 2:00 Assumptions for the talk and what viewers will gain
- 3:00 Walter Cluer's monitoring-driven past vs. observability
- 4:30 Extensive technology sprawl and disparate monitoring tools
- 5:30 High MTTR due to disjointed monitoring tools
- 6:30 The journey to solve MTTR and adopt DevOps
- 7:30 Adopting a data pipelining approach for telemetry
How To Adopt OpenTelemetry in an Enterprise Where Incumbent Vendor Tools Reign Supreme
Speakers: Chris Weldon, Director of Platform Engineering, Wolters Kluwer
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=JqRXqk-1CLo
Overview
In this insightful KubeCon EU talk, Chris Weldon, Director of Platform Engineering at Wolters Kluwer, Tax & Accounting North America, shares the company's transformative journey from a heavily vendor-laden, monitoring-centric environment to a modern observability strategy powered by OpenTelemetry. The presentation details the strategic, cultural, and technical challenges encountered and overcome during this enterprise-wide shift, offering practical advice for organizations grappling with similar modernization efforts.
Weldon's talk is particularly pertinent for enterprises struggling with fragmented monitoring tools, high Mean Time To Resolution (MTTR) for incidents, and the desire to adopt open standards to mitigate vendor lock-in. He emphasizes that the successful adoption of OpenTelemetry is less about the technology itself and more about fostering a cultural shift towards observability, empowering development teams, and strategically aligning cross-functional stakeholders.
The session provides a candid look at the "burns and scars" collected along the way, highlighting the importance of evangelism, community building, and minimizing disruption to active development teams. By sharing Wolters Kluwer's experience, Weldon empowers other organizations to embark on their own observability journeys, demonstrating that even large, established enterprises can successfully transition to open standards and modern practices.
Background
▶ Watch: Introduction to Chris Weldon and talk agenda (0:00)
Wolters Kluwer, a globally distributed company specializing in professional information, software, and services, found itself in a common predicament for large enterprises in 2020. Their Tax & Accounting North America division operated primarily as a "monitoring-driven shop," a model that proved sufficient for simple, constrained application environments like a basic web application and database on a single virtual machine. However, as applications began to modernize, moving to the cloud, adopting Kubernetes, and utilizing Platform-as-a-Service (PaaS) offerings, the existing monitoring strategy quickly became a liability.
The modernization efforts led to significant technology sprawl in monitoring tools. Different teams had their own specialized solutions: DBAs used their preferred tools, operations managed Application Performance Monitoring (APM) and infrastructure monitoring, and developers handled logs independently. This fragmented approach meant that when incidents occurred, particularly low in the stack due the complex inter- and intra-service dependencies, the Mean Time To Resolution (MTTR) became unacceptably high. Service owners at the application edge would receive the initial alert, investigate to a certain point, then hand off to another team, leading to a frustrating "whack-a-mole" incident resolution process that dissatisfied both customers and business leadership.
Recognizing the urgent need to improve MTTR, reduce tool proliferation, and integrate monitoring earlier in the development lifecycle—a "shift left" approach aligned with DevOps principles—Wolters Kluwer embarked on a strategic initiative. Their vision was to adopt a data pipelining approach for all telemetry data, enabling processing, enrichment, filtering, sampling, and backend correlation. By April 2020, they were actively investigating the industry landscape, where OpenTelemetry was emerging as a significant force, concluding its merger with OpenCensus and OpenTracing into a unified standard. Although some specifications were still experimental, tracing was stable, and the .NET libraries (critical for Wolters Kluwer) were rapidly maturing. The strong endorsement of OpenTelemetry by various vendors further bolstered internal confidence, reinforcing the company's preference for an open standard over proprietary vendor lock-in.
The initial tool selection process faced a hurdle: a developer-led evaluation, while well-intentioned, neglected the broader cross-functional needs. This led to feedback from the operations team, highlighting the lack of a true DevOps approach. In response, a cross-functional "backend bake-off" team was formed, comprising representatives from operations, development, performance engineering, quality engineering, and architecture. This team collaboratively identified required capabilities, ensuring objectivity by challenging biases towards specific vendors. After a multi-month assessment, they selected a set of tools, with OpenTelemetry compatibility as a non-negotiable top-tier requirement. This approach recognized that a "one tool to rule them all" solution was impractical across an enterprise with diverse budgetary and architectural needs, aiming instead for a streamlined, yet flexible, set of tools underpinned by an open standard.
Key Findings
▶ Watch: Walter Cluer's monitoring-driven past vs. observability (3:00)
The journey to adopt OpenTelemetry at Wolters Kluwer yielded several critical findings, emphasizing that successful enterprise-wide observability is as much a cultural and strategic endeavor as it is a technical one:
- Cultural Shift is Paramount: The most significant challenge was not the technology itself, but educating and convincing application development teams about the value of observability, particularly the importance of traces beyond traditional logging. The transition from monitoring to observability required a fundamental shift in mindset.
- Minimize AppDev Impact: To foster adoption, it was crucial to minimize the disruption to application development teams' existing workflows. While some impact was inevitable due to cultural changes, the platform team actively sought ways to abstract away complexity and provide easy-to-use solutions.
- Embedded Expert Model: The embedded expert model proved highly effective. Platform engineers collaborated directly with application developers, instrumenting their code and sharing knowledge. This approach not only facilitated the technical implementation but also built internal champions within product teams who could then evangelize observability.
- Real-World Incident Proves Value: A production incident, occurring just three days after tracing was introduced to an application, became a powerful evangelism tool. The ability to resolve the issue (identifying bad hosts) in less than 15 minutes—remarkably fast for the organization at the time—demonstrated the tangible benefits of tracing data, overcoming initial skepticism from developers.
- Auto-Instrumentation for Efficiency: Leveraging auto-instrumentation SDKs significantly streamlined the adoption process for both logging and tracing, reducing the manual effort required from application teams.
- Centralized Collector Management: Owning and deploying the OpenTelemetry Collector centrally by the platform team was key. This ensured consistency in processing pipelines, filtering, and sampling rules, and removed the burden of Collector management from product teams, allowing them to focus on application telemetry enrichment.
- Anticipate Coverage Gaps: Despite best efforts, some legacy logging libraries were inevitably missed during the initial rollout. Planning for an "inspect and adapt" approach, rather than expecting 100% upfront coverage, was crucial for discovering and addressing these gaps.
- WAF Compatibility Challenges: Deploying the OpenTelemetry Collector at the edge for OTLP over HTTPS can trigger Web Application Firewalls (WAFs) to treat the traffic as a Denial of Service (DoS) attack. This necessitated careful tuning of both WAF and Collector configurations to establish symbiosis.
- Enterprise Scaling Requires Evangelism and Community: Expanding OpenTelemetry adoption beyond a single business unit required proactive communication, sharing experiences (both successes and failures), leveraging internal technology conferences and communities of practice, and identifying and empowering internal "enthusiasts."
- Encourage Open-Source Contributions: Leadership support for teams to contribute back to the OpenTelemetry project was identified as a powerful way to foster ownership and improve the standard for everyone.
- Future with eBPF and Sidecars: For broader enterprise adoption, especially for legacy applications, future strategies include exploring eBPF and sidecar-based injection of OpenTelemetry to further reduce the need for application-level instrumentation.
Technical Deep Dive
▶ Watch: Extensive technology sprawl and disparate monitoring tools (4:30)
Wolters Kluwer's technical journey began in 2020 from a state of disparate monitoring tools. The infrastructure included a mix of legacy applications on virtual machines, alongside modern deployments on Kubernetes and various PaaS services. This hybrid environment contributed to the "technology sprawl" where DBAs, operations, and development teams each had their own monitoring solutions, leading to a lack of correlation and high MTTR. The strategic decision was to move towards a unified observability strategy built around OpenTelemetry.
The company recognized OpenTelemetry as a burgeoning standard, benefiting from the convergence of OpenCensus and OpenTracing. While some specifications were still in the experimental phase, the tracing specifications were deemed stable, and the .NET libraries (a critical component for Wolters Kluwer's technology stack) were rapidly maturing. This stability, coupled with vendor support for OpenTelemetry, provided the confidence needed to commit to an open standard and avoid proprietary lock-in.
The implementation strategy for OpenTelemetry involved a phased rollout, addressing logs, traces, and metrics separately:
Logging Strategy
The approach to logging was designed to minimize disruption and build trust with application development teams.
- Leverage Existing SDKs: The majority of applications in the ecosystem utilized a custom logging SDK. This SDK was a central point for injecting common context, such as user GUIDs, virtual data center information, and database IDs, making it an ideal candidate for initial OpenTelemetry integration. For applications not using this custom SDK, modern application development frameworks often provided configurable sinks, simplifying the injection process.
- Supplemental Injection: The platform team first modified these SDKs to inject OpenTelemetry logging capabilities as a supplemental stream. Logs continued to flow to their existing destinations while also being sent via OpenTelemetry to the new backend.
- Build Trust through Parity: App development teams upgraded their SDK versions and rolled out these changes to various environments. The crucial step here was to ensure parity: developers could see their logs in both the old and new destinations, confirming that no data was lost and building confidence in the new OpenTelemetry-based system.
- Rip and Replace: Once critical mass was achieved and teams were confident in the parity, the platform team went back into the SDKs to rip out the old plumbing, ensuring logs exclusively flowed through OpenTelemetry to the new observability backend.
Tracing Strategy
Introducing tracing initially met with skepticism from application developers who often felt that "logging is sufficient." The platform team addressed this by strategically leveraging auto-instrumentation and a real-world incident:
- Auto-Instrumentation: The platform team injected auto-instrumentation SDKs from OpenTelemetry directly into application code. Given Wolters Kluwer's .NET stack, this involved supporting WCF (Windows Communication Foundation) services (SOAP-based) and Web API solutions, which integrated well with existing OpenTelemetry projects.
- Real-World Validation: A pivotal moment occurred three days after tracing was deployed to a production application. A production incident, unrelated to OpenTelemetry, allowed Chris Weldon to demonstrate the power of traces. Within minutes, by analyzing the captured trace data, he pinpointed the problematic services and identified that a "couple of hosts had gone bad." This enabled operations to remove the faulty hosts, restoring service in less than 15 minutes—a significant improvement over their previous MTTR.
- Evangelism through Success: This incident served as a powerful exemplar, transforming skepticism into enthusiasm. It clearly illustrated the value of traces in quickly understanding system behavior and isolating problems, something logs alone could not achieve with the same efficiency. The education aspect around tracing proved to be a larger effort than the technical changes themselves.
Metrics Strategy
At the time of the initial rollout, OpenTelemetry SDKs for metrics were less stable, and few application development teams were emitting custom metrics.
- Deferred Implementation: The decision was made to defer the full implementation of metrics until the OpenTelemetry SDKs matured and application teams demonstrated a clear need for custom metrics.
- Future Integration: Currently, Wolters Kluwer's wrapped OpenTelemetry SDKs do support metrics. The next phase involves replacing existing infrastructure metrics collection agents with OpenTelemetry Collector-based solutions, aiming for a 100% OpenTelemetry Collector-driven telemetry pipeline.
OpenTelemetry Collector
The OpenTelemetry Collector was identified as the crucial middleware connecting application telemetry data to the various backends.
- Centralized Ownership: To ensure consistency and reduce the burden on product teams, the platform team took centralized ownership of OpenTelemetry Collector deployment and management. This allowed product teams to focus on learning the new observability tools and enriching their application telemetry.
- Consistent Processing: Centralized management ensured consistency in processing pipelines, including filtering and sampling rules, which are critical for managing data volume and cost.
- Edge Deployment Benefits: Deploying the Collector service at the edge provided new capabilities, such as collecting logs directly from client machines where applications were installed. Previously, collecting these client-side logs was an "extremely painful" manual process involving support and customers. Now, this data is readily available in the observability backends.
- WAF Challenges: A significant operational lesson learned was that Web Application Firewalls (WAFs), with security postures common in enterprises, could treat OTLP over HTTPS traffic from edge-deployed Collectors as a Denial of Service (DoS) attack. This required careful tweaking and tuning of both WAF and Collector configurations to establish a working balance.
Enterprise Scaling and Future Directions
For enterprise-wide adoption, the strategy extended beyond technical implementation to cultural and organizational aspects. This involved evangelizing successes, sharing lessons learned, and building internal communities. Looking ahead, Wolters Kluwer is exploring advanced auto-instrumentation techniques like eBPF and sidecar-based injection for OpenTelemetry, particularly for legacy applications where modifying application code is challenging. For new, "Green Field" development, the directive is clear: use OpenTelemetry from the outset.
Demo / Proof of Concept
▶ Watch: The journey to solve MTTR and adopt DevOps (6:30)
While the talk did not feature a live technical demonstration or a pre-recorded proof of concept in the traditional sense, a real-world production incident served as a powerful, unplanned "proof of concept" for the value of OpenTelemetry tracing. As detailed in the "Technical Deep Dive," a critical production issue arose just three days after tracing was introduced to an application. Chris Weldon leveraged the newly available trace data to quickly identify the root cause—a few bad hosts—leading to a resolution in under 15 minutes. This dramatic reduction in Mean Time To Resolution (MTTR), achieved through the insights provided by OpenTelemetry traces, effectively demonstrated the practical benefits and capabilities of the new observability strategy to skeptical application development teams.
Defensive Implications
▶ Watch: Adopting a data pipelining approach for telemetry (7:30)
The insights shared by Chris Weldon offer crucial defensive implications for various stakeholders within an enterprise aiming to enhance their security and operational posture through observability.
For Application Development Teams:
- Embrace Observability Culture: Shift focus from simply monitoring known failure points to understanding the entire system's behavior. This proactive approach leads to quicker identification of unknown issues and better system resilience.
- Standardize Telemetry: Adopt OpenTelemetry for all new and existing applications to emit standardized logs, traces, and metrics. This ensures consistent data across the enterprise, facilitating correlation and analysis.
- Leverage Auto-Instrumentation: Utilize OpenTelemetry's auto-instrumentation SDKs where possible to minimize manual effort and ensure comprehensive telemetry coverage, especially for common frameworks and protocols.
- Understand Trace Value: Actively learn and utilize tracing data for debugging and incident resolution. Traces provide invaluable context for distributed systems, dramatically reducing the time to pinpoint root causes.
- Engage with Platform Teams: Participate in "embedded expert" models and provide feedback to platform teams to ensure the observability solution meets their specific application needs and integrates smoothly into development workflows.
For Platform and Operations Teams:
- Centralize OpenTelemetry Collector Management: Assume ownership of deploying and managing OpenTelemetry Collectors. This ensures consistent data processing, filtering, sampling, and routing across the enterprise, optimizing resource usage and data quality.
- Prioritize OpenTelemetry Compatibility: Ensure that all chosen backend observability tools (log aggregators, trace analyzers, metric stores) are fully compatible with OpenTelemetry's OTLP (OpenTelemetry Protocol) standard. This avoids vendor lock-in and maintains flexibility.
- Build Robust Data Pipelines: Design and implement robust data processing pipelines within the OpenTelemetry Collector to enrich, filter, and sample telemetry data effectively, managing data volume and cost while preserving critical information.
- Tune Network Infrastructure for OTLP: Anticipate and address potential conflicts with network security devices like Web Application Firewalls (WAFs). Configure WAFs to allow OTLP over HTTPS traffic from OpenTelemetry Collectors, preventing legitimate telemetry from being blocked as a Denial of Service attack.
- Act as Observability Evangelists: Actively educate and evangelize the benefits of OpenTelemetry and observability to application teams, leadership, and other business units. Share successes, lessons learned, and best practices through internal channels.
- Support Open-Source Contributions: Encourage and enable team members to contribute to the OpenTelemetry project. This not only improves the standard for everyone but also fosters deeper expertise and ownership within the organization.
For Leadership and Management:
- Strategic Investment in Observability: Recognize observability as a strategic imperative for improving operational efficiency, reducing MTTR, and enhancing customer satisfaction. Allocate appropriate resources for tools, training, and personnel.
- Foster Cross-Functional Collaboration: Promote the formation of cross-functional teams (like the "backend bake-off" team) for evaluating and implementing observability solutions. This ensures all stakeholder needs are met and fosters a true DevOps culture.
- Embrace Open Standards: Prioritize open standards like OpenTelemetry to avoid vendor lock-in, maintain flexibility, and future-proof the observability strategy against changing market dynamics.
- Budget for Diverse Needs: Understand that different business units may have varying budgetary and architectural requirements. Plan for a catalog of supported observability backends that can cater to large enterprise applications, mid-sized solutions, and smaller, more constrained environments, all while leveraging OpenTelemetry as the common data ingress.
- Support Cultural Change: Acknowledge that the transition to observability is a significant cultural shift. Provide leadership support for training, mentorship, and internal community building to drive successful adoption across the enterprise.
By proactively addressing these implications, enterprises can build a resilient, observable, and efficient operational environment, significantly reducing incident resolution times and fostering a culture of continuous improvement.
Key Takeaways
- Observability is a Cultural Transformation, Not Just a Tool Change: Successfully adopting OpenTelemetry requires a significant shift in mindset from reactive monitoring to proactive, holistic system understanding, driven by education and evangelism.
- Open Standards Mitigate Vendor Lock-in: OpenTelemetry provides a robust, open standard for collecting telemetry data, allowing enterprises to choose best-of-breed backend solutions without being locked into a single vendor's ecosystem.
- Cross-Functional Collaboration is Essential: Forming diverse teams (e.g., "backend bake-off") with representatives from development, operations, architecture, and quality engineering is crucial for objectively evaluating tools and gaining enterprise-wide buy-in.
- Real-World Successes Drive Adoption: Demonstrating the tangible benefits of OpenTelemetry, particularly tracing, through actual production incidents, can dramatically accelerate cultural adoption and overcome initial skepticism.
- Minimize Developer Burden with Strategic Implementation: Leveraging auto-instrumentation, existing SDKs, and centralizing OpenTelemetry Collector management are key strategies to reduce the impact on application development teams, allowing them to focus on business value.
- Enterprise Scaling Requires Evangelism and Community Building: To roll out OpenTelemetry across a large organization, proactive communication, sharing experiences, building internal networks of enthusiasts, and fostering communities of practice are vital.
About the Speaker(s)
Chris Weldon is the Director of Platform Engineering at Wolters Kluwer, specifically working within the Tax & Accounting North America division. Hailing from Texas, he often incorporates regional expressions like "howdy" and "ya'all" into his presentations. Chris is a passionate advocate for mentorship, particularly for individuals in underrepresented groups within the technology community, and encourages others to participate in such initiatives. He expresses pride in working for Wolters Kluwer, highlighting the company's recognition, including being rated number one for gender diversity in the Netherlands for three consecutive years.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk provides a brutally honest, no-bullshit account of Wolters Kluwer's journey to adopt OpenTelemetry in a complex enterprise. Weldon pulls no punches detailing the cultural, strategic, and technical hurdles, including the often-overlooked WAF compatibility issues. The practical strategies for logging, tracing, and Collector management, coupled with the 'embedded expert' model and the compelling real-world incident that proved tracing's value, make this an invaluable case study for any organization struggling with vendor lock-in and fragmented observability. It's not a zero-day, but it's real work that matters.
Heather Calloway (CISO) — STRONG ACCEPT
Chris Weldon's presentation on Wolters Kluwer's OpenTelemetry adoption is a strong example of how strategic operational shifts directly impact business resilience and accountability. It effectively demonstrates the journey from fragmented monitoring to a unified observability strategy, addressing critical issues like high Mean Time To Resolution and vendor lock-in. The talk provides actionable insights for platform teams and leaders, highlighting that success hinges on cultural transformation and cross-functional collaboration, not just technology. While primarily an operations talk, its implications for security incident response and understanding institutional risk are undeniable.