Museum of Weird Bugs: Our Favorites From 8 Years of Service Mesh Debugging - Alex Leong, Buoyant
Alex Leong, Buoyant
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In this insightful talk, Alex Leong, a long-standing maintainer of the Linkerd project at Buoyant, takes attendees on a journey through the "Museum of Weird Bugs." Leong shares some of the most challenging and elusive bugs encountered over eight years of developing and debugging Linkerd, a widely adopted service mesh. Far from being trivial, these bugs highlight the inherent complexities and unexpected interactions that arise when multiple distributed systems—Kubernetes, network protocols, custom resources, and application logic—converge.

Key moments
- 0:00 Introduction to Linkerd and "Museum of Weird Bugs" talk
- 2:00 Linkerd's high-level architecture: Data plane and Control plane
- 2:40 Deep dive into Destination and Policy controllers' roles
- 4:00 How Linkerd controllers interact with Kubernetes API
- 4:30 Linkerd's use of Gateway API and custom policy CRDs
- 5:30 Explanation of HTTPRoute CRD duplication and deprecation
- 6:40 Policy Controller's three key CRD management responsibilities
Museum of Weird Bugs: Our Favorites From 8 Years of Service Mesh Debugging
Speakers: Alex Leong, Maintainer, Buoyant
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=Kcjh0-hXwWw
Overview
In this insightful talk, Alex Leong, a long-standing maintainer of the Linkerd project at Buoyant, takes attendees on a journey through the "Museum of Weird Bugs." Leong shares some of the most challenging and elusive bugs encountered over eight years of developing and debugging Linkerd, a widely adopted service mesh. Far from being trivial, these bugs highlight the inherent complexities and unexpected interactions that arise when multiple distributed systems—Kubernetes, network protocols, custom resources, and application logic—converge.
The presentation serves as a candid look into the realities of maintaining critical infrastructure software. Leong emphasizes that many of these "nasty" bugs were not discovered internally but reported by users operating Linkerd at vast scales and in diverse environments, making reproduction and diagnosis particularly arduous. By dissecting these real-world incidents, the talk offers invaluable lessons for developers, operators, and architects on building resilient systems, understanding subtle failure modes, and the art of debugging in highly distributed and dynamic environments.
Why does this talk matter? It demystifies the debugging process for complex systems, illustrating how seemingly minor mismatches or overlooked architectural nuances can lead to catastrophic failures like memory leaks, deadlocks, and incorrect routing. The detailed analysis of each bug, from its symptoms to its root cause and ultimate resolution, provides concrete examples of the challenges in modern cloud-native infrastructure and offers actionable insights into designing more robust and observable distributed applications.
Background
▶ Watch: Introduction to Linkerd and "Museum of Weird Bugs" talk (0:00)
Linkerd is a service mesh designed to add observability, reliability, and security to Kubernetes applications without requiring application code changes. Its architecture is split into two main components: the data plane and the control plane.
The data plane consists of lightweight Rust-based microproxies that run as sidecar containers alongside each application container within a Kubernetes pod. These proxies intercept all incoming and outgoing network traffic for the application, handling tasks like load balancing, retries, and traffic splitting. To perform these functions, the proxies need up-to-date information about services, endpoints, and policies.
This information is provided by the control plane, which runs in a dedicated Kubernetes namespace. The control plane primarily consists of two components:
- Destination Controller (written in Go): This component provides destination information to the proxies, such as endpoint addresses and TLS identities. It streams this data to the proxies over a gRPC streaming API.
- Policy Controller (written in Rust): A newer component responsible for serving policy information, including authorization rules (which service can talk to which), and routing overrides. It also communicates with proxies via gRPC streams.
Both controllers obtain their information by establishing HTTP watches on the Kubernetes API. They listen for updates on core Kubernetes resources like Pods, Services, Endpoints, and Service Accounts. They then synthesize this raw Kubernetes data into a format usable by the proxies and stream it down.
In addition to core Kubernetes resources, Linkerd integrates with Custom Resource Definitions (CRDs). It leverages the Gateway API for defining traffic routing rules (e.g., HTTPRoute, gRPCRoute, TLSRoute, TCPRoute). Linkerd also has its own set of policy CRDs under the policy.io group, such as Server and AuthorizationPolicy, which define how Linkerd should enforce policy. A notable detail mentioned is the historical duplication of the HTTPRoute CRD: Linkerd initially forked the HTTPRoute CRD into its policy.io group to add support for a timeout field that was not yet available in the upstream Gateway API. While the Gateway API now includes this field, Linkerd still supports both, with a multi-phased deprecation plan to eventually consolidate onto the upstream Gateway API CRDs.
The policy controller plays a critical role in managing these CRDs. It performs three key functions:
- Validation: It uses a validation webhook to ensure that any created or updated CRDs conform to Linkerd's rules.
- Watching: It watches these CRDs to gather information necessary for synthesizing policy data sent to the proxies.
- Status Updates: It updates the
statussubfield of these CRDs to indicate whether they have been accepted, are valid, or if they reference non-existent backends or other invalid configurations. This status update mechanism is critical for understanding the first bug discussed.
This intricate interplay between application proxies, control plane components, the Kubernetes API, and custom resources creates a fertile ground for complex and often subtle bugs, as demonstrated in the subsequent "weird bugs."
Key Findings
▶ Watch: Deep dive into Destination and Policy controllers' roles (2:40)
The talk meticulously dissects two primary categories of "weird bugs" encountered in Linkerd over eight years, each revealing fundamental lessons about distributed systems and software robustness:
- CRD Mismatch Leading to Infinite Update Loop and Memory Leak: This bug highlighted a critical synchronization issue between a Custom Resource Definition's schema and the Go/Rust structs used by the Linkerd policy controller. A missing field in the CRD's
statussubfield, present in the controller's internal representation, led to an endless cycle of patching the Kubernetes API, causing memory exhaustion and service degradation. The key finding here is the absolute authority of the CRD definition over internal code structures and the dangers of maintaining API forks.
- gRPC Stream Deadlock Causing Stale Service Discovery: This more insidious bug demonstrated how a seemingly isolated issue—a single misbehaving client failing to process gRPC stream updates—could cascade and deadlock the entire destination controller. The blocking nature of
stream.Sendcombined with HTTP/2 flow control meant that one slow consumer could prevent all other clients from receiving updates, leading to widespread routing to stale addresses. The critical finding was the need for careful isolation of blocking I/O operations and robust handling of backpressure in control plane architectures.
While a third bug related to memory leaks from map keys and deallocations was mentioned due to time constraints, the two detailed examples underscore a broader finding: the complexity of service meshes stems from their position at the intersection of diverse systems (Kubernetes, network protocols, custom resources, application logic), making debugging challenging due to transient states, distributed nature, and the difficulty of reproducing issues that manifest only at scale or under specific environmental conditions.
Technical Deep Dive
▶ Watch: How Linkerd controllers interact with Kubernetes API (4:00)
Bug 1: CRD Mismatch Leading to Infinite Update Loop
The first bug presented a classic memory leak in the policy controller, which, at its peak, consumed 780MB of memory before being OOM-killed by Kubernetes. Users reported that creating an HTTPRoute CRD resulted in its status being updated to Accepted, but only after an inexplicable 40-minute delay. Logs were also filled with "failed to patch HTTP route, no available capacity" errors, indicating an overloaded Kubernetes API.
The policy controller's logic for CRD status updates is straightforward:
- It establishes watches on relevant CRDs (e.g.,
HTTPRoute). - Upon receiving an update, it computes the desired
statusfor the resource (e.g.,Acceptedor not, based on backend validity). - It compares the computed status with the current status on the resource.
- If there's a difference, it issues a patch to the Kubernetes API to update the
statussubfield.
The root cause was a subtle mismatch in the definition of the ParentReference struct within the status subfield of the httproutes.policy.linkerd.io CRD. Specifically, the CRD definition on the Kubernetes API server lacked a port field for ParentReference, while the Rust struct used by the policy controller did include a port field. This discrepancy arose because Linkerd had forked the Gateway API's HTTPRoute CRD at an earlier time (when port was not present in the upstream definition), but the Rust bindings were generated later from a version of the Gateway API where port had been added.
The infinite loop unfolded as follows:
- A user creates or updates an
HTTPRouteCRD. - The Kubernetes API stores it, discarding the
portfield inParentReferencebecause it's not defined in the CRD schema. - The policy controller receives this update via its watch.
- The controller computes the desired status, which includes the
portfield in its internal Rust struct. - It compares its internal representation (with
port) to the received Kubernetes object (withoutport), detects a difference, and issues a patch to the Kubernetes API. - The Kubernetes API receives the patch, again discards the
portfield, and stores the resource. - This modified resource triggers another update to the policy controller, restarting the loop.
This continuous patching overloaded the Kubernetes API, causing the 40-minute delays and "no available capacity" errors. The policy controller's memory ballooned because it was constantly processing and holding state for these endless updates.
The fix was simple once identified: correct the httproutes.policy.linkerd.io CRD to include the port field in its ParentReference status definition, aligning it with the controller's internal struct. This resolved the loop, bringing memory usage back to normal and status updates to near-instantaneous. The key lesson here is that CRDs are the ultimate source of truth; what you write to Kubernetes may not be what is persisted if it deviates from the CRD schema. It also highlights the perils of maintaining forks of API definitions, as they can easily drift out of sync.
Bug 2: Stale Addresses / gRPC Stream Deadlock
The second bug manifested as Linkerd proxies routing to stale addresses, leading to connection refused errors or even connections to incorrect, re-used IP addresses. These bugs were notoriously difficult to debug because they were transient, involved multiple distributed systems (proxy, destination controller, Kubernetes API), and required detecting the absence of expected updates.
The flow of updates from the Kubernetes API to the Linkerd proxy involves the destination controller (written in Go):
- The destination controller uses client-go informers to establish watches on Kubernetes resources (e.g.,
Endpoints,Services). - When a relevant resource changes, an
onUpdatecallback is triggered in the controller. - Inside this callback, the controller computes the new desired state for the service.
- It then iterates through all Linkerd proxies that have subscribed to updates for that service.
- For each subscribed proxy, it sends the update using
stream.Sendover a gRPC streaming API.
The critical insight here is that stream.Send in gRPC is a blocking call. It will not return until the message has been successfully sent. This behavior becomes problematic when combined with HTTP/2 flow control windows. HTTP/2 implements a backpressure mechanism where a sender can only transmit a certain amount of data (the "window size") until the receiver sends a window update, signaling that it has processed data and is ready for more. If the receiver stops sending window updates, the sender will block indefinitely.
The scenario that led to the deadlock:
- A Linkerd proxy (the gRPC client/receiver) or any other client connecting to the destination controller, for various reasons (e.g., a bug in its network libraries, CPU starvation, network weirdness, or simply being a misbehaving client that doesn't read data), stops sending HTTP/2 window updates.
- The destination controller, attempting to send an update via
stream.Sendto this client, eventually fills the HTTP/2 flow control window. stream.Sendblocks, waiting for a window update that never comes.- Because the
stream.Sendcall is blocked, theforloop iterating over all subscribed proxies also blocks. - Crucially, this
forloop is executing within the client-go informer callback thread. If this callback thread blocks, the destination controller stops processing any further updates from the Kubernetes API.
The result is a devastating deadlock: one misbehaving client effectively starves the entire destination controller of Kubernetes API updates and prevents it from sending updates to any of its subscribed proxies. All proxies, not just the misbehaving one, end up working with stale service discovery information, leading to widespread routing failures.
The solution involved refactoring the destination controller to isolate the blocking behavior:
- Instead of directly calling
stream.Sendwithin the informer callback, updates are now enqueued into a channel (queue) dedicated to each listener/stream. This enqueue operation is designed to be non-blocking. - A separate goroutine (the "Q processor") is responsible for dequeuing messages from each listener's channel and sending them over the gRPC stream.
- With this architecture, if a single listener stops sending window updates and its dedicated queue fills up, only that specific queue and its associated goroutine will block. The informer callback thread remains unblocked, continuing to process Kubernetes API updates, and other listeners continue to receive updates normally.
- Furthermore, if a queue becomes too full, the system can simply terminate that specific gRPC stream, forcing the misbehaving client to reconnect, rather than letting it deadlock the entire controller.
This fix highlights the importance of understanding blocking I/O, the nuances of network protocols like HTTP/2 flow control, and employing asynchronous patterns (like goroutines and channels in Go) to build resilient distributed systems that can gracefully handle unresponsive or slow consumers without impacting overall system health. The lesson is clear: never block inside an informer callback.
Demo / Proof of Concept
▶ Watch: Explanation of HTTPRoute CRD duplication and deprecation (5:30)
The talk focuses on retrospective analysis of past bugs rather than live demonstrations or proofs of concept. Alex Leong walks through the historical symptoms, debugging process, and resolution of these complex issues, providing code snippets and architectural diagrams to illustrate the technical details.
Defensive Implications
▶ Watch: Policy Controller's three key CRD management responsibilities (6:40)
The "Museum of Weird Bugs" offers critical lessons for anyone building or operating cloud-native infrastructure, especially service meshes:
- Meticulous CRD Management: Treat Custom Resource Definitions (CRDs) as the absolute source of truth for your API. Any discrepancy between the CRD schema and the internal data structures used by controllers can lead to subtle yet catastrophic issues like infinite update loops. Implement robust generation of language bindings from CRDs and rigorous validation to ensure consistency.
- Avoid API Forks: Where possible, avoid forking upstream APIs or CRDs. Maintaining separate versions introduces synchronization challenges and drift, as exemplified by the
HTTPRouteCRD issue. If a fork is unavoidable, implement extremely disciplined processes for keeping it in sync with the upstream or migrating back as soon as feasible. - Understand Blocking I/O: Be acutely aware of which operations are blocking and under what conditions they might block indefinitely. This is especially crucial in control plane components that process events from external systems (like the Kubernetes API).
- Isolate Blocking Behavior: Design architectures to isolate blocking operations. Use asynchronous patterns, such as dedicated goroutines with buffered channels (queues), to ensure that a single slow or misbehaving client cannot deadlock critical processing paths (e.g., Kubernetes informer callbacks) or starve other clients.
- HTTP/2 Flow Control Awareness: Deeply understand the implications of underlying network protocols. HTTP/2 flow control windows are a powerful backpressure mechanism, but if not accounted for in application logic, they can lead to unexpected deadlocks when clients fail to send window updates.
- Robustness in Informer Callbacks: Never block inside a Kubernetes client-go informer callback. Blocking these threads prevents the controller from receiving any further updates from the Kubernetes API, leading to stale state across the entire control plane.
- Enhanced Observability for Distributed Systems: Invest heavily in comprehensive monitoring, logging, and tracing that can correlate events across different components (proxies, control planes, Kubernetes API). Debugging issues like "missing updates" or "stale data" requires the ability to observe the absence of expected behavior and track state changes across a distributed graph.
- Proactive Memory Management: Implement continuous monitoring of memory usage, especially for controllers that interact with APIs. Memory leaks due to infinite loops or improper resource handling can quickly lead to OOM kills and service instability. Focus not just on allocations but also on deallocations to identify true leaks.
Key Takeaways
- CRDs as Canonical Source: Always treat your CRD definitions as the ultimate authority. Mismatches between CRD schemas and your code's data structures can cause insidious bugs like infinite update loops and memory leaks.
- Beware of Blocking Calls: Be extremely cautious with blocking I/O operations, particularly within critical control plane components or Kubernetes informer callbacks, as they can lead to system-wide deadlocks.
- Isolate and Decouple: Employ architectural patterns like per-client queues and separate goroutines to isolate blocking behaviors, ensuring that one misbehaving client or slow consumer cannot starve the entire system of updates.
- Understand Network Flow Control: Grasp the implications of underlying network protocols like HTTP/2 flow control. Backpressure mechanisms, while essential, can inadvertently cause deadlocks if not handled gracefully at the application layer.
- Debugging Distributed Systems is Hard: Expect transient, state-based bugs that are difficult to reproduce. Effective debugging requires correlating information across multiple systems (proxies, control plane, Kubernetes API) and often looking for the absence of expected events.
- Avoid API Forks: Minimize maintaining forks of external APIs or CRDs to prevent synchronization issues and reduce maintenance overhead.
About the Speaker(s)
Alex Leong is a dedicated maintainer on the Linkerd project, having been involved since the project's inception eight years ago. Working at Buoyant, the company behind Linkerd, he has accumulated extensive experience in the development and debugging of service mesh technologies, encountering and resolving a wide array of complex and challenging bugs in the process.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This talk is a brutal, honest look into the reality of maintaining critical infrastructure at scale. Leong dissects two complex, real-world bugs from Linkerd's history, revealing how subtle architectural mismatches and fundamental misunderstandings of network protocols can lead to catastrophic system failures. It's a masterclass in distributed systems debugging and resilient design, offering invaluable, hard-won lessons that every engineer building cloud-native applications needs to internalize. This isn't just theory; it's battle-tested wisdom from the trenches.
Heather Calloway (CISO) — STRONG ACCEPT
This talk delivers a candid and detailed dissection of complex bugs in distributed systems, specifically a service mesh. While deeply technical, it masterfully translates these operational challenges into critical lessons on system resilience, architectural robustness, and the perils of subtle engineering oversights. It's a valuable resource for engineering leaders and architects, offering concrete, actionable insights into designing and operating reliable cloud-native infrastructure, with direct implications for institutional accountability and business continuity.