K8s in Wonderland: Why? Many of Unknown Code in My Workload? - Hoon Jo, Megazone

Hoon Jo, Megazone

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

In his KubeCon EU talk, "K8s in Wonderland: Why? Many of Unknown Code in My Workload?", Hoon Jo from Megazone invites attendees to "down to the Kubernetes hole," echoing Alice's journey into Wonderland, to explore the often-overlooked default configurations and underlying mechanisms within Kubernetes. The core premise of the talk revolves around demystifying the "unknown code" – the myriad of options, policies, and parameters that are automatically applied or commonly adopted in Kubernetes deployments, yet whose rationale and implications are frequently misunderstood by developers and operators. Jo argues that while these settings might appear as black boxes, they are in fact intelligently designed and represent community-driven best practices.

Watch on YouTube

Visual summary for K8s in Wonderland: Why? Many of Unknown Code in My Workload? - Hoon Jo, Megazone by Hoon Jo, Megazone
Visual summary for K8s in Wonderland: Why? Many of Unknown Code in My Workload? - Hoon Jo, Megazone by Hoon Jo, Megazone

Key moments

  1. 0:00 Introduction and 'Unknown Code' in Kubernetes
  2. 2:00 Speaker's CNCF Ambassador role and passion
  3. 3:10 Examples of 'unknown code': image and restart policies
  4. 4:45 Using kubectl explain to discover hidden options
  5. 6:00 Detailed kubectl explain for container policies
  6. 7:35 K8s GPT as a language translation tool

K8s in Wonderland: Why? Many of Unknown Code in My Workload?

Speakers: Hoon Jo, Megazone

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=GvIPSgt69Sg

Overview

In his KubeCon EU talk, "K8s in Wonderland: Why? Many of Unknown Code in My Workload?", Hoon Jo from Megazone invites attendees to "down to the Kubernetes hole," echoing Alice's journey into Wonderland, to explore the often-overlooked default configurations and underlying mechanisms within Kubernetes. The core premise of the talk revolves around demystifying the "unknown code" – the myriad of options, policies, and parameters that are automatically applied or commonly adopted in Kubernetes deployments, yet whose rationale and implications are frequently misunderstood by developers and operators. Jo argues that while these settings might appear as black boxes, they are in fact intelligently designed and represent community-driven best practices.

Hoon Jo, a CNCF Ambassador and KubeAstron, emphasizes that a deep understanding of these "unknowns" is paramount for optimizing Kubernetes workloads, ensuring robust security, and making informed architectural decisions. The talk provides a detailed exploration of several critical Kubernetes components, including imagePullPolicy, deployment strategies like maxSurge, and various service configurations such as externalTrafficPolicy and sessionAffinity. Through practical demonstrations and insightful flowcharts, Jo illustrates why these defaults are often the most suitable choices for general-purpose use cases, shedding light on the intelligent design principles embedded within the Kubernetes ecosystem.

The significance of this talk extends beyond mere configuration explanations; it empowers Kubernetes users to move from passive acceptance of defaults to an active, informed engagement with their deployments. By understanding the "why" behind these settings, engineers can better troubleshoot issues, enhance performance, improve resilience, and strengthen the security posture of their applications running on Kubernetes. It's a call to look beyond the surface, to explore the intricacies of Kubernetes, and to appreciate the community consensus that shapes its robust and versatile architecture.

Background

▶ Watch: Introduction and 'Unknown Code' in Kubernetes (0:00)

The rapid adoption of Kubernetes has brought immense benefits in terms of container orchestration, scalability, and declarative infrastructure. However, this power comes with a significant degree of complexity. A typical Kubernetes deployment involves numerous YAML manifests, each containing a multitude of fields, options, and parameters. While many of these are explicitly defined by the user, a substantial portion either rely on default values or are configured using commonly accepted, yet not always fully understood, patterns. These are what Hoon Jo refers to as "unknown code" – configurations that silently dictate behavior without explicit, conscious choice or deep comprehension from the user.

The problem arises because, despite their apparent simplicity or commonality, these "unknowns" can have profound impacts on the functionality, performance, and security of applications. For instance, an imagePullPolicy might seem trivial, but its setting can dictate how quickly updates are deployed, how much network bandwidth is consumed, and even introduce security vulnerabilities if not properly managed. Similarly, deployment strategies like maxSurge directly influence the availability and speed of application updates, while service configurations affect network traffic routing and load distribution.

The existence of this "unknown code" is a natural byproduct of Kubernetes' design philosophy, which aims to provide sensible defaults that work well for the majority of use cases. These defaults are often the result of extensive community discussion, real-world operational experience, and a consensus-driven approach to balancing various trade-offs (e.g., performance vs. overhead, consistency vs. availability). While tools like kubectl explain provide invaluable documentation on each field and its available options, they don't always fully articulate the operational implications or the rationale behind the default or recommended values. This gap in understanding can lead to suboptimal configurations, unexpected behavior, and a reactive approach to problem-solving, rather than a proactive, informed deployment strategy. Jo's talk aims to bridge this gap, translating the technical specifications into practical, actionable insights. He even mentions using an AI model, KSGPT, in previous work to help translate complex Kubernetes concepts, highlighting the general difficulty people face in grasping these intricacies.

Key Findings

▶ Watch: Examples of 'unknown code': image and restart policies (3:10)

Hoon Jo's talk reveals that the "unknown code" within Kubernetes, encompassing default configuration options and commonly adopted practices, is not arbitrary but rather intelligently designed, often representing community-driven best practices. The primary finding is that a conscious understanding of these defaults is crucial for optimizing workloads, ensuring security, and making deliberate deployment decisions.

Specifically, the key findings include:

  • imagePullPolicy: Always is the default for a reason, particularly beneficial for development environments where images are frequently updated. It intelligently compares image digests for tagged images and always downloads untagged (e.g., latest) images, ensuring developers are always working with the freshest version. However, for production, specific, immutable tags combined with IfNotPresent offer greater predictability and security.
  • Deployment maxSurge: 25% is identified as a highly effective and "generally proposed" default for rolling updates. This percentage strikes an optimal balance between update speed and resource overhead, especially for typical Kubernetes clusters with 2-3 worker nodes. It allows for a smooth, gradual update process without excessive resource consumption or prolonged downtime.
  • Service externalTrafficPolicy: Cluster is the default because it ensures robust load balancing by distributing traffic across all available pods, regardless of their node location. While Local can preserve source IP, Cluster prioritizes broad availability and resilience.
  • Service ipFamilyPolicy: SingleStack remains the default as IPv4 is still the predominant network protocol in many data centers and cloud environments, despite the growing adoption of IPv6 in ISPs.
  • Service sessionAffinity: None is the default to guarantee proper load balancing behavior, distributing client requests evenly across all available endpoints. While ClientIP provides sticky sessions for specific application needs, it fundamentally alters load balancing and can introduce uneven resource utilization or single points of failure.

In essence, Jo demonstrates that these seemingly invisible or taken-for-granted configurations are the product of careful consideration, designed to provide robust, performant, and reliable behavior for the broadest range of Kubernetes workloads. Understanding the logic behind these defaults empowers users to leverage Kubernetes more effectively and securely.

Technical Deep Dive

▶ Watch: Using kubectl explain to discover hidden options (4:45)

Hoon Jo meticulously breaks down several key Kubernetes configuration options, illustrating their behavior and the rationale behind their default or recommended usage. He uses a combination of flowcharts and recorded demonstrations to convey these complex concepts.

Image Pull Policy

The imagePullPolicy field, part of a pod's container specification, dictates when and how a container image is pulled from a registry. Jo explains its three options:

  • Always: This is the default policy. If an image has no tag (e.g., nginx instead of nginx:1.23), it's treated as latest and Always downloads the image. If a specific tag is used (e.g., nginx:1.23), Always will check the image digest against the locally cached image. If the digest differs (indicating a new version for the same tag), it downloads the new image. This behavior makes Always ideal for development environments where images are frequently updated, ensuring developers are always working with the latest code (09:00).
  • IfNotPresent: This policy instructs Kubernetes to use a locally available image if it exists. If the image is not present on the node, it will be downloaded. Crucially, if a specific tag is used and the image is already on the node, IfNotPresent will not check for a newer version in the registry, even if the underlying image layer has changed (16:00). This provides predictability for production environments where specific, immutable tags are preferred.
  • Never: This policy prevents Kubernetes from attempting to pull an image from a registry. If the image is not locally present, the pod will fail to start. This is typically used in highly controlled or air-gapped environments where all images are pre-loaded onto nodes.

Demo for imagePullPolicy:

Jo demonstrates the practical differences using two pods, one configured with imagePullPolicy: Always (implicitly, by not specifying a policy with a latest tag) and another with imagePullPolicy: IfNotPresent (also with a latest tag). He first shows that both pods download the image initially (14:00). Then, he modifies the index.html file of the application, rebuilds the Docker image, and pushes it to Docker Hub with the latest tag. After deleting and recreating the pods, the pod with Always downloads the new image, while the pod with IfNotPresent (having a local image) does not, even though the latest tag points to a newer version.

He further illustrates this by using a specific tag (e.g., swap-image) with IfNotPresent. Even after changing the image layer and pushing a new image with the same specific tag, the pod with IfNotPresent does not download the new image because a locally tagged image already exists (16:00). This highlights the importance of specific, immutable tags in production for predictable deployments when IfNotPresent is used. The speaker notes switching from Docker Hub to Quay.io during preparation due to Docker's rate limits (22:00).

Deployment Strategy: maxSurge

For Deployment resources, Jo focuses on the maxSurge parameter, which controls the number of pods that can be created above the desired count during a rolling update.

  • progressDeadlineSeconds: This defines the maximum time a deployment can take to progress before it's considered failed. The default is 600 seconds (10 minutes). This integer value can be adjusted (18:00).
  • revisionHistoryLimit: This specifies the number of old ReplicaSets to retain. The default is 10. While useful for rollbacks, too many can consume excessive etcd storage.
  • maxSurge: This can be an integer or a percentage. Percentages are often more appropriate as they scale with the number of replicas, preventing the need to guess exact pod counts (19:00).
  • maxSurge: 10%: Allows only one extra pod at a time (for 10 pods), resulting in very slow rolling updates but minimal overhead.
  • maxSurge: 25%: This is the default and a "generally proposed" value. For a deployment of 10 pods, it allows 2-3 extra pods to be created. Jo explains this is a good balance for typical clusters with 2-3 worker nodes, optimizing update speed without excessive resource consumption (21:00).
  • maxSurge: 80%: Allows many extra pods, leading to very fast updates but significantly higher resource overhead, similar to a blue/green deployment strategy.

Demo for maxSurge:

Jo demonstrates the impact of maxSurge by deploying an application and then patching its maxSurge value using kubectl patch. He changes the image version and observes the rolling update process on a cluster with three worker nodes (22:00).

  • With maxSurge: 25%, the update proceeds efficiently, deploying new pods alongside old ones.
  • With maxSurge: 80%, the update is very fast, with a large number of new pods spinning up quickly, but at the cost of higher temporary resource usage.
  • With maxSurge: 10%, the update is noticeably slower, as new pods are brought online one by one (24:00).

This visual comparison reinforces why 25% is the common and recommended default, providing a good trade-off for most scenarios.

Service Configuration

Finally, Jo delves into various configurations for Kubernetes Service objects:

  • externalTrafficPolicy:
  • Cluster (default): Traffic is load-balanced across all pods of the service, regardless of which node they reside on. This ensures even distribution and high availability but can obscure the client's source IP address.
  • Local: Traffic is only routed to pods on the same node as the incoming request. This preserves the client's source IP but can lead to uneven load distribution or traffic blackholing if a node receives traffic for a service but has no local pods for it (25:00). The default Cluster is generally safer for broad service availability.
  • ipFamilyPolicy:
  • SingleStack (default): The service uses a single IP family (IPv4 or IPv6).
  • DualStack / RequireDualStack: The service is configured to use both IPv4 and IPv6. Jo notes that while IPv6 is gaining popularity with ISPs, SingleStack is still the common default in many IDC (Internet Data Center) and corporate environments (26:00).
  • sessionAffinity:
  • None (default): Kubernetes load balances requests across all available service endpoints, ensuring even distribution.
  • ClientIP: This enables sticky sessions, meaning all requests from a particular client IP address will be directed to the same pod. This can be useful for stateful applications (e.g., shopping carts) but bypasses the load balancing mechanism, potentially leading to uneven pod utilization (27:00).

Demo for sessionAffinity:

Jo deploys two services (port 11 and port 12) with identical pods, evenly distributed across worker nodes using topologySpreadConstraint (28:00). One service is configured with sessionAffinity: None (port 11), and the other with sessionAffinity: ClientIP (port 12).

He then uses curl to send multiple requests to both services. For the None service, requests are distributed across different pod endpoints, demonstrating proper load balancing. For the ClientIP service, all requests from the same curl client are consistently routed to a single pod, proving the sticky session functionality. Jo emphasizes that while ClientIP has specific uses, it's crucial to understand that it disables the load balancer's primary function (29:00).

Demo / Proof of Concept

▶ Watch: Detailed kubectl explain for container policies (6:00)

Hoon Jo's talk is rich with practical demonstrations, serving as concrete proof-of-concept for each technical concept discussed. Although he mentions limitations for a live demo and uses recorded versions, these effectively illustrate the "unknown code" in action.

For imagePullPolicy, the demonstration involved deploying two pods, initially both using a latest tag. One pod implicitly used the default Always policy, while the other explicitly set IfNotPresent. The speaker then modified the application's index.html, rebuilt the Docker image, and pushed it with the latest tag to a container registry (Quay.io, changed from Docker Hub due to rate limits). After deleting and recreating the pods, the Always policy pod correctly pulled the updated latest image, reflecting the changes. In contrast, the IfNotPresent pod, having a local image tagged latest, did not pull the newer version, demonstrating its predictable behavior of utilizing local caches. Further, he showcased how using a specific, immutable tag with IfNotPresent ensures that even if the underlying image layer changes in the registry, the locally tagged image is consistently used, reinforcing the predictability for production environments (14:00-17:00).

The Deployment maxSurge demonstration visually highlighted the impact on rolling update speed. Jo deployed an application and then used kubectl patch to dynamically modify the maxSurge parameter to 25%, 80%, and 10% while simultaneously changing the application's image version. By observing the kubectl watch output, attendees could clearly see the varying speeds of pod replacement: maxSurge: 80% resulted in a very rapid update due to a high number of new pods being spun up concurrently, maxSurge: 25% offered a balanced, moderately fast update, and maxSurge: 10% led to a significantly slower, one-by-one pod replacement process (22:00-24:00). This practical comparison validated the rationale behind the 25% default for most general-purpose clusters.

Finally, the Service sessionAffinity proof-of-concept showcased the difference between traditional load balancing and sticky sessions. Two services were deployed, each exposing a different port (11 and 12), backed by identical pods distributed evenly across nodes using topologySpreadConstraint. One service was configured with sessionAffinity: None (the default), and the other with sessionAffinity: ClientIP. Using curl to repeatedly send HTTP requests to both service IPs and ports, the demo visibly illustrated that requests to the None service were distributed across different backend pods, as expected from a load balancer. Conversely, requests to the ClientIP service consistently hit the same backend pod, demonstrating the sticky session functionality (28:00-29:00). This provided clear evidence of how this "unknown code" fundamentally alters traffic routing behavior.

These demonstrations collectively served to demystify the internal workings of Kubernetes defaults, transforming abstract configuration options into tangible, observable behaviors.

Defensive Implications

▶ Watch: K8s GPT as a language translation tool (7:35)

Understanding the "unknown code" within Kubernetes is not just about efficiency; it's also critical for establishing a robust security posture. Each of the discussed configurations has significant defensive implications that operators and security professionals must consider.

Image Pull Policy

  • imagePullPolicy: Always: While convenient for development, using Always with an untagged latest image in production is a significant security risk. It means Kubernetes will always pull the newest version, which could be compromised in a supply chain attack or simply introduce an untested vulnerability. For production, Always should only be used with specific, immutable image digests or highly trusted, continuously scanned registries.
  • imagePullPolicy: IfNotPresent: This is generally safer for production when combined with specific, immutable image tags (e.g., my-app:v1.2.3-abcd123). It ensures that once a known-good image is on a node, it will be used, preventing accidental deployment of a newer, potentially malicious latest image. However, if a node is compromised and a malicious image is locally cached or injected with the same tag, IfNotPresent could inadvertently use it. Therefore, image integrity checks and node security are paramount.
  • imagePullPolicy: Never: This provides the highest level of control over image sources, ideal for air-gapped or extremely high-security environments. It guarantees that only pre-approved and locally available images can run, mitigating risks from external registry compromises. However, it requires a robust process for distributing and updating images to nodes.

Deployment Strategy (maxSurge, progressDeadlineSeconds, revisionHistoryLimit)

  • progressDeadlineSeconds: A long deadline (default 600s) can mask deployment failures, potentially leaving a vulnerable or broken application running for an extended period. Defenders should monitor deployment status and consider shorter, context-appropriate deadlines to quickly detect and roll back problematic deployments.
  • revisionHistoryLimit: While useful for rollbacks, storing too many old ReplicaSets can consume excessive etcd storage. From a security perspective, old ReplicaSets might contain configurations for outdated, vulnerable application versions. While Kubernetes won't run them unless explicitly rolled back, managing this history responsibly is good hygiene.
  • maxSurge: The chosen maxSurge value impacts the speed at which new, potentially vulnerable or fixed, versions are deployed. A high maxSurge (e.g., 80%) means a fast rollout, quickly propagating either a fix or a new vulnerability across the cluster. A low maxSurge (e.g., 10%) provides a slower rollout, allowing more time for monitoring and early detection of issues before they affect the entire fleet. The default 25% strikes a balance, but specific security requirements might dictate a slower, more controlled rollout or a faster "emergency patch" rollout.

Service Configuration (externalTrafficPolicy, ipFamilyPolicy, sessionAffinity)

  • externalTrafficPolicy: Local: This policy preserves the client's source IP address. This is crucial for security tools like firewalls, Web Application Firewalls (WAFs), or intrusion detection systems that rely on source IP for filtering, rate limiting, or threat intelligence. However, misconfiguration can lead to traffic blackholing if no local pods are available, creating a denial-of-service vulnerability. Proper pod topology and readiness probes are essential.
  • ipFamilyPolicy: SingleStack: If a cluster is intended to be dual-stack (IPv4 and IPv6), misconfiguring a service as SingleStack could inadvertently expose it only on one IP family, potentially bypassing network segmentation or firewall rules applied to the other family, or simply making the service inaccessible to certain clients. Understanding the network topology and intended IP families is key.
  • sessionAffinity: ClientIP: While useful for stateful applications, sticky sessions can bypass the load balancer's primary function. From a security perspective, this can lead to uneven resource utilization, making it easier for an attacker to target a specific, potentially weaker, pod if they can manipulate source IPs. It also creates a single point of failure if that sticky pod crashes. Defenders must ensure that applications using ClientIP affinity are robust, highly available, and adequately monitored to prevent such issues. For general-purpose services, sessionAffinity: None is the more secure and resilient default.

In summary, every "unknown code" configuration in Kubernetes carries potential security implications. A proactive defensive strategy requires not just knowing what the defaults are, but deeply understanding why they are defaults, their trade-offs, and how they impact the overall security posture of the application and the cluster.

Key Takeaways

  • Demystify Kubernetes Defaults: Many "unknown codes" in Kubernetes, referring to default values and common configurations, are intelligently designed and represent community-driven best practices aimed at balancing performance, resilience, and operational overhead.
  • Utilize kubectl explain: This powerful command-line tool is your primary resource for understanding the purpose, available options, and default values for any Kubernetes resource field.
  • Strategic imagePullPolicy Usage: For production environments, always use specific, immutable image tags combined with imagePullPolicy: IfNotPresent to ensure predictable and secure deployments. imagePullPolicy: Always (especially with latest tags) is more suited for development due to its constant image freshness checks.
  • Optimal Rolling Updates with maxSurge: 25%: The default maxSurge: 25% for Deployments provides an excellent balance between update speed and resource overhead, making it a robust choice for most Kubernetes clusters and workloads.
  • Understand Service Traffic Management: Default service configurations like externalTrafficPolicy: Cluster and sessionAffinity: None are crucial for proper load balancing and high availability. While externalTrafficPolicy: Local (for source IP preservation) and sessionAffinity: ClientIP (for sticky sessions) have specific use cases, they alter fundamental load balancing behavior and require careful consideration of their implications.
  • No Silver Bullet: As the speaker quotes the Cheshire Cat, there's no single "right" way to configure Kubernetes. Context, specific application requirements, and security posture should always guide configuration choices, moving beyond blindly accepting defaults to making informed decisions.

About the Speaker(s)

Hoon Jo is a prominent figure in the Kubernetes community, currently working with Megazone, a Korean managed service provider (MSP). He is recognized as a CNCF Ambassador, reflecting his significant contributions and dedication to fostering the cloud-native ecosystem. Additionally, Hoon Jo identifies as a KubeAstron, further highlighting his deep involvement and expertise within the Kubernetes sphere. He is an active contributor to the community, participating in and organizing various meetups, and is a frequent speaker at major conferences. His passion for Kubernetes is evident in his commitment to spreading knowledge and helping others navigate its complexities, as demonstrated by his presentations at events like KubeCon EU and his plans to speak at KubeCon China.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Hoon Jo's 'K8s in Wonderland' delivers a substantive deep-dive into the often-overlooked default configurations of Kubernetes. The talk brilliantly demystifies the 'why' behind critical settings like imagePullPolicy, maxSurge, and sessionAffinity, transforming abstract defaults into actionable knowledge. Through clear explanations and practical, recorded demonstrations, Jo empowers operators and security professionals to move beyond passive acceptance, fostering a more informed and secure approach to Kubernetes deployments. This isn't groundbreaking zero-day research, but it's a masterclass in understanding foundational system behavior, with significant defensive implications.

Heather Calloway (CISO) — STRONG ACCEPT

This talk by Hoon Jo is a critical resource for demystifying Kubernetes defaults, transforming abstract configurations into understandable choices with significant security and operational implications. It effectively highlights how seemingly minor settings can directly impact an organization's risk posture and operational resilience, pushing for informed decision-making over passive acceptance. While technical, its clarity and practical demonstrations provide essential grounding for security architects and lead engineers to build more resilient and secure systems, and to articulate the 'why' behind their configuration choices.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025