Simplifying the Networking and Secur... Bill Mulligan, Anna Kapuścińska, Bowei Du & Amir Kheirkhahan

Bill Mulligan, Anna Kapuścińska, Bowei Du, Amir Kheirkhahan

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This KubeCon EU maintainers track session, "Simplifying the Networking and Security Stack with Cilium," delves into the transformative capabilities of Cilium, an eBPF-powered cloud-native networking, observability, and security solution for Kubernetes. Presented by key contributors from the Cilium project, Google, and DB Schenker, the talk highlights recent advancements in Cilium 1.17, real-world adoption stories, and future directions for scalability and enhanced security. It underscores how Cilium, built on the revolutionary eBPF kernel technology, is becoming the de facto standard for managing complex cloud-native environments.

Watch on YouTube

Visual summary for Simplifying the Networking and Secur... Bill Mulligan, Anna Kapuścińska, Bowei Du & Amir Kheirkhahan by Bill Mulligan, Anna Kapuścińska, Bowei Du, Amir Kheirkhahan
Visual summary for Simplifying the Networking and Secur... Bill Mulligan, Anna Kapuścińska, Bowei Du & Amir Kheirkhahan by Bill Mulligan, Anna Kapuścińska, Bowei Du, Amir Kheirkhahan

Key moments

  1. 0:00 Cilium's core capabilities: networking, observability, security.
  2. 2:00 Cilium's rapid growth, production adoption, and multicluster leadership.
  3. 3:10 Highlights of Cilium 1.17: networking, security, and scale updates.
  4. 5:00 DB Schenker's challenge: reliable networking for Kafka and multicluster.
  5. 6:05 DB Schenker's near-zero downtime Cilium migration and Hubble insights.
  6. 7:00 Cilium simplifies troubleshooting and provides granular network policy enforcement.
  7. 8:05 Addressing sidecar service mesh L7 visibility limitations with Cilium.

Simplifying the Networking and Security Stack with Cilium

Speakers: Bill Mulligan, Cilium Maintainer; Anna Kapuścińska, Software Engineer, Isovalent; Bowei Du, Google; Amir Kheirkhahan, Platform Engineer, DB Schenker

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=kYT7KV_Cijs

Overview

This KubeCon EU maintainers track session, "Simplifying the Networking and Security Stack with Cilium," delves into the transformative capabilities of Cilium, an eBPF-powered cloud-native networking, observability, and security solution for Kubernetes. Presented by key contributors from the Cilium project, Google, and DB Schenker, the talk highlights recent advancements in Cilium 1.17, real-world adoption stories, and future directions for scalability and enhanced security. It underscores how Cilium, built on the revolutionary eBPF kernel technology, is becoming the de facto standard for managing complex cloud-native environments.

The session emphasizes Cilium's evolution from a basic Container Network Interface (CNI) to a comprehensive platform encompassing advanced networking features like Cluster Mesh and Gateway API, alongside robust observability via Hubble and cutting-edge security with Tetragon. The speakers articulate why Cilium matters in today's cloud-native landscape: it addresses critical challenges related to performance, security, and operational complexity at unprecedented scales, offering a unified and efficient solution that bypasses traditional kernel limitations.

The impressive show of hands at the beginning of the session, indicating widespread production use of Cilium, sets the stage for a deep dive into its features and benefits. The talk serves as a testament to Cilium's maturity and its crucial role in enabling organizations to build, secure, and operate highly scalable and resilient Kubernetes infrastructures, simplifying what has historically been a fragmented and complex domain.

Background

▶ Watch: Cilium's core capabilities: networking, observability, security. (0:00)

Cilium originated as a CNI plugin for Kubernetes, fundamentally designed around eBPF (extended Berkeley Packet Filter). eBPF is a powerful Linux kernel technology that enables sandboxed programs to run within the operating system kernel without requiring changes to kernel source code or loading kernel modules. This capability allows Cilium to provide highly efficient and programmable networking, security, and observability features directly at the kernel level.

Over time, Cilium has significantly expanded its scope beyond basic Layer 3 networking. It now offers advanced functionalities crucial for modern cloud-native deployments, including Cluster Mesh for multi-cluster connectivity, Egress Gateway for controlled outbound traffic, kube-proxy replacement for efficient service load balancing, BGP (Border Gateway Protocol) integration, and support for the Gateway API. This expansion covers the entire network stack, from Layer 2 to Layer 7, providing a complete networking solution for Kubernetes.

For observability, Cilium integrates Hubble, which leverages eBPF to provide deep insights into network flows. Hubble generates a service map that visualizes service-to-service communication, simplifying debugging and understanding complex microservice architectures. On the security front, Cilium offers features like transparent encryption of network traffic for compliance and confidentiality, and robust network policy enforcement, including Layer 7 policies, which are critical for securing Kubernetes clusters against evolving threats. The project's prominence is further solidified by its status as the most starred eBPF-based project on GitHub, its recognition in the CNCF project journey report as the third fastest-growing project, and its top rating in the CNCF tech user radar for multi-cluster management solutions, highlighting its widespread adoption and perceived maturity by end-users.

Key Findings

▶ Watch: Highlights of Cilium 1.17: networking, security, and scale updates. (3:10)

The talk presented several key findings and advancements across Cilium's ecosystem:

  1. Cilium 1.17 Enhancements: The latest release introduces significant improvements in networking, security, and operations. Networking gains include Quality of Service (QoS) support (guaranteed, burstable, best-effort), Multicluster Services API for global services in Cluster Mesh, and Gateway API 1.2 support. On the security side, network policy performance has been greatly increased, alongside new features for prioritizing critical policies and network policy validation. Operational improvements include new metrics for BGP network connections, rate limiting for monitoring events, and enhanced scale testing, supporting clusters with thousands of nodes and thousands of clusters.
  1. DB Schenker's Production Migration and Benefits: Amir Kheirkhahan detailed DB Schenker's successful multi-step migration to Cilium with near-zero downtime. Their experience highlighted Cilium's ability to:
  • Provide superior troubleshooting and visibility compared to previous solutions.
  • Enhance Hubble metrics with Kubernetes metadata (service name, pod name, identity) to create custom dashboards for deep network understanding.
  • Enable granular policy enforcement and Layer 7 traffic visibility, eliminating the need for application-level firewalls.
  • Simplify operations by replacing a sidecar-based service mesh, enabling transparent encryption (pod-to-pod, node-to-node) with a single helm flag, and achieving better application resilience.
  • Improve performance by replacing kube-proxy with eBPF-based load balancing and implementing eBPF host routing for optimized packet processing and faster namespace switching, especially critical for Kafka workloads.
  • Future plans include leveraging Cluster Mesh for secure cross-cluster communication.
  1. GKE's Data Plane v2 Scaling to Extreme Levels: Bowei Du from Google showcased how GKE's Data Plane v2, powered by Cilium, supports unprecedented scale.
  • GKE has deployed clusters with up to 65,000 nodes, where every node runs a Cilium agent.
  • A major challenge at this scale is handling a very high pod churn rate of up to 500 pod lifecycle changes per second.
  • The primary bottleneck identified for such extreme scale is the Kubernetes control plane.
  • To address this, Google currently restricts Cilium features to a subset (IPAM, pod-to-pod connectivity, Kubernetes service, basic Hubble observability) to achieve this scale, but aims for more feature completeness.
  • The community is exploring Cilium configuration profiles to optimize feature sets for specific use cases, such as a "high-scale basic networking profile."
  1. Tetragon's Expanding Security and Observability: Anna Kapuścińska provided updates on Tetragon, the eBPF-based security and observability project under the Cilium umbrella.
  • Windows Support: Initial support for eBPF for Windows is coming in Tetragon 1.5 (June release), starting with process create and process exit tracing, enabling cloud-native eBPF tools on Windows.
  • Built-in Overhead Measurement: Tetragon now includes built-in CPU and memory overhead measurement, broken down to individual BPF programs and maps. This data is exposed via the tetra CLI and Prometheus metrics, allowing users to precisely understand and monitor Tetragon's performance impact in their specific environments.
  • Advanced Policy Language: The policy language has been enhanced to allow extraction of fields from inside kernel structures (e.g., file structures during kernel file operations), enabling more granular and sophisticated policy definitions.
  • Configurable Policy Mode: Policies can now be toggled between observability (default, generating events) and enforcement modes (blocking actions), facilitating a "observe-then-enforce" security workflow.
  • CEL (Common Expression Language) Filters: Tetragon events can now be filtered using arbitrary expressions written in CEL, enabling advanced rule creation for detecting specific CVE exploits or suspicious activities, such as an attacker searching for AWS credentials.

Technical Deep Dive

▶ Watch: DB Schenker's challenge: reliable networking for Kafka and multicluster. (5:00)

At the heart of Cilium's capabilities lies eBPF, a revolutionary kernel technology that allows users to run custom programs safely and efficiently within the Linux kernel. This enables Cilium to perform networking, security, and observability tasks with unparalleled performance and flexibility, bypassing the overhead and limitations of traditional kernel modules or user-space proxies.

For networking, Cilium leverages eBPF to implement its CNI functionalities. Instead of relying on iptables or ipvs for service load balancing, Cilium replaces kube-proxy with eBPF programs. This replacement significantly boosts performance, especially in large Kubernetes clusters with thousands of services, by directly manipulating packet forwarding rules in the kernel, avoiding the inefficiencies associated with iptables rule processing. DB Schenker specifically noted this as a key benefit for their Kafka workloads, where network latency is critical. Furthermore, Cilium employs eBPF host routing to optimize packet processing. When a packet leaves a pod's namespace, an eBPF program intercepts it, bypassing the traditional host network stack and routing table lookups, and routing it directly to the physical device. This results in faster namespace switching and improved throughput. Future optimizations include replacing veth devices with netkit pairs to achieve host-level throughput for container namespaces and reduce latency by using Layer 3 routing instead of Layer 2.

Hubble, Cilium's observability component, also heavily relies on eBPF. It taps into eBPF programs attached to network interfaces to capture detailed flow information at the kernel level. This data is then used to construct a real-time service map that visualizes all network communication within and across clusters. DB Schenker enhanced this by leveraging Hubble's context options to inject Kubernetes metadata (service name, pod name, identity) into flow metrics, enabling highly customized and granular dashboards for deep operational insights without requiring SSH access to nodes or complex sniffing plugins.

On the security front, Cilium provides granular policy enforcement from Layer 3 to Layer 7. This is achieved by programming eBPF filters that can inspect and enforce rules based on IP addresses, ports, and even HTTP/gRPC methods. The talk highlighted transparent encryption as a critical security feature. By enabling a single flag in the Helm chart, Cilium automatically encrypts all pod-to-pod and node-to-node traffic, ensuring data confidentiality in transit without requiring manual key rotation or complex certificate management. This simplifies compliance with security standards dramatically.

Tetragon, another project under the Cilium umbrella, extends eBPF's security capabilities into runtime visibility and enforcement. Tetragon's generic, low-level policy language allows it to hook into virtually any point in the Linux kernel. It can trace system calls (syscalls), kernel functions (kprobes), and user-space functions (uprobes). For instance, it can monitor process_create and process_exit events, providing real-time insights into application behavior. The advanced policy language now allows for deep inspection, enabling extraction of specific fields from kernel data structures – for example, details from a file structure during a kernel file operation. This precision allows security teams to create highly targeted policies. Tetragon's configurable policy modes (observability vs. enforcement) allow for a phased rollout of security policies, starting with monitoring and then moving to blocking actions once confidence is established. The integration of Common Expression Language (CEL) filters further empowers users to define complex rules for filtering Tetragon's JSON events, enabling detection of sophisticated attack patterns like specific CVE exploits or attempts to locate sensitive data such as AWS credentials. The new built-in CPU and memory overhead measurement, exposed via tetra CLI and Prometheus, provides crucial transparency into Tetragon's performance impact, allowing users to fine-tune their deployments.

Finally, the talk touched on scalability challenges. Google's experience with 65,000-node clusters highlighted that while Cilium is highly performant, the Kubernetes control plane often becomes the bottleneck, especially under high pod churn rates (up to 500 changes per second). To address this, the community is developing Cilium configuration profiles. These profiles aim to optimize Cilium's feature set for specific use cases (e.g., a "high-scale basic networking profile") by carefully selecting and testing features that can scale to extreme environments, ensuring that the project continues to meet the demands of the largest cloud-native deployments.

Demo / Proof of Concept

▶ Watch: Cilium simplifies troubleshooting and provides granular network policy enforc... (7:00)

While the session did not feature a live, interactive demo or a traditional proof of concept demonstration, it provided compelling real-world evidence and use cases that serve as powerful validations of Cilium's capabilities.

The most significant "proof of concept" was presented by Amir Kheirkhahan from DB Schenker. His detailed account of migrating their self-managed Kubernetes clusters on AWS to Cilium, achieving near-zero downtime, directly showcased Cilium's practical benefits. This included the successful replacement of a sidecar-based service mesh with Cilium's transparent encryption, the performance gains from switching from kube-proxy to eBPF-based load balancing, and the enhanced observability derived from custom Hubble dashboards. DB Schenker's experience with running Kafka, which demands extremely low latency, further underscored Cilium's performance advantages in a production-critical environment.

Similarly, Bowei Du's discussion of Google Kubernetes Engine's (GKE) Data Plane v2, which uses Cilium at its core, illustrated Cilium's ability to operate at an unprecedented scale of 65,000 nodes with high pod churn rates. This operational deployment by a major cloud provider serves as a robust proof point for Cilium's scalability and reliability in the most demanding environments.

The updates on Tetragon, including its upcoming Windows support and the new built-in overhead measurement, also implicitly function as a "proof of concept" of the project's continuous development and commitment to addressing critical security and operational concerns for its users. These real-world adoption stories and operational insights collectively demonstrate Cilium's maturity, effectiveness, and readiness for enterprise-grade cloud-native deployments.

Defensive Implications

▶ Watch: Addressing sidecar service mesh L7 visibility limitations with Cilium. (8:05)

The advancements in Cilium and Tetragon offer significant implications for enhancing the defensive posture of cloud-native environments:

  1. Granular Network Segmentation and Layer 7 Enforcement: Defenders can leverage Cilium's eBPF-powered network policies to implement highly granular segmentation, not just at Layer 3/4, but also at Layer 7. This means policies can restrict communication based on HTTP methods, paths, or gRPC services, effectively preventing lateral movement even if an attacker gains access to a pod. This eliminates the need for separate application-level firewalls.
  1. Transparent Data-in-Transit Encryption: By simply enabling a flag, organizations can achieve transparent pod-to-pod and node-to-node encryption. This simplifies compliance with data protection regulations (e.g., GDPR, HIPAA) by ensuring all inter-service communication within the cluster is encrypted without application changes or manual key management, significantly reducing the risk of data interception.
  1. Enhanced Runtime Security with Tetragon: Tetragon provides deep, kernel-level visibility and enforcement capabilities. Defenders can deploy eBPF-based policies to trace critical kernel events like process creation/exit, file access, and network connections. This allows for real-time detection of suspicious activities, such as an attacker attempting to enumerate sensitive files (e.g., /etc/passwd, AWS credentials as shown in the CEL example) or execute unusual binaries. The ability to switch policies from observability to enforcement mode enables a measured response, allowing security teams to audit events before blocking potentially malicious actions.
  1. Proactive Threat Detection with CEL Filters: The introduction of Common Expression Language (CEL) filters for Tetragon events empowers security analysts to write highly sophisticated detection rules. These rules can identify specific CVE exploits or complex attack patterns by combining multiple event attributes, moving beyond simple signature-based detection to behavioral analysis at the kernel level.
  1. Superior Observability for Incident Response: Hubble provides a comprehensive service map and detailed network flow logs, which are invaluable for incident response. Defenders can quickly visualize communication patterns, identify anomalous traffic, and trace the origin and destination of malicious activities. Customizing Hubble metrics with Kubernetes metadata, as demonstrated by DB Schenker, further enhances the ability to contextualize security alerts and accelerate investigation.
  1. Performance and Scalability without Security Compromise: Cilium's eBPF-based kube-proxy replacement and eBPF host routing ensure that security features do not introduce significant performance overhead, even in large-scale, high-churn environments. This is crucial for maintaining a strong security posture without impacting application performance or operational efficiency.
  1. Policy Validation and Overhead Monitoring: Cilium's new network policy validation features help prevent misconfigurations that could inadvertently create security gaps. Furthermore, Tetragon's built-in CPU and memory overhead measurement provides transparency, allowing defenders to ensure that their security policies are not excessively impacting system resources, thus preventing a trade-off between security and stability.

Key Takeaways

  • eBPF as the Foundation: Cilium leverages eBPF to deliver high-performance, kernel-level networking, observability (Hubble), and security (Tetragon) functionalities, fundamentally simplifying the cloud-native stack and bypassing traditional kernel limitations.
  • Comprehensive Cloud-Native Solution: Cilium has evolved beyond a CNI to offer a full suite of features including Cluster Mesh, Gateway API support, kube-proxy replacement, and transparent encryption, addressing complex challenges from Layer 2 to Layer 7.
  • Proven at Scale and in Production: Real-world case studies from DB Schenker and Google (GKE's 65,000-node clusters with 500 pod churns/sec) demonstrate Cilium's ability to handle extreme scale, simplify operations, and enhance security in critical production environments.
  • Advanced Runtime Security with Tetragon: Tetragon provides deep, customizable kernel-level tracing and enforcement through eBPF, offering features like Windows support, built-in overhead measurement, an advanced policy language for granular control, and CEL filters for sophisticated threat detection.
  • Defensive Simplification and Robustness: Cilium enables granular Layer 7 network policies and transparent encryption (pod-to-pod, node-to-node) out-of-the-box, significantly simplifying compliance and bolstering the security posture against modern threats without sacrificing performance.
  • Community-Driven Innovation: The Cilium community is actively addressing future challenges like extreme scalability through initiatives like configuration profiles and expanding platform support, ensuring its continued relevance and growth in the cloud-native ecosystem.

About the Speaker(s)

  • Bill Mulligan is a Cilium maintainer, bringing over a decade of experience with the project. His deep involvement in Cilium's development and community efforts positions him as a leading voice in the eBPF and cloud-native networking space.
  • Anna Kapuścińska is a Software Engineer at Isovalent, the company behind Cilium and Tetragon. She primarily works on Tetragon, focusing on extending Cilium's observability and security capabilities with generic eBPF-based policies.
  • Bowei Du represents Google in this discussion, contributing insights from the perspective of a major cloud provider. He works on GKE's Data Plane v2, which leverages Cilium to power Kubernetes networking at scale within Google Cloud.
  • Amir Kheirkhahan is a Platform Engineer at DB Schenker. He is responsible for the development and maintenance of reliable and secure cloud platform solutions, including their self-managed Kubernetes clusters on AWS. His firsthand experience with migrating to and operating Cilium provides valuable end-user perspective.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This session on Cilium's advancements, particularly leveraging eBPF for networking, observability, and security, is a masterclass in cloud-native infrastructure. It delivers exceptional technical depth, showcasing kernel-level optimizations, robust security features via Tetragon, and real-world scalability proven by Google and DB Schenker. The talk provides actionable insights for anyone serious about high-performance, secure Kubernetes deployments, solidifying Cilium's role as a critical component of the modern stack.

Heather Calloway (CISO) — STRONG ACCEPT

This KubeCon session delivers clear, actionable insights into how Cilium and eBPF are maturing into a foundational layer for cloud-native security and networking. The compelling real-world evidence from DB Schenker and Google's GKE demonstrates its ability to simplify complex challenges, enhance operational efficiency, and reduce business risk at scale. It effectively translates advanced kernel technology into practical, enterprise-grade solutions for security leaders and their teams.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025