Effective Detection in Kubernetes Clusters
Shay Berkovich (Security Researcher · Wiz), Oren Ofer (Detection Engineer · Runtime Sensor)
BSidesSF 2024 · Day 1
Overview
This presentation, delivered by Shay Berkovich and Oren Ofer at BSidesSF 2024, delves into the complexities of detecting sophisticated attacks within Kubernetes and cloud-native environments. The speakers, both seasoned security professionals with backgrounds in threat research and detection engineering, highlight the evolving landscape of Kubernetes attacks, which increasingly incorporate cloud components, leverage new initial access vectors, and adapt older techniques for modern contexts. The core thesis of their talk is that effective detection in this intricate ecosystem necessitates a multi-dimensional approach, combining visibility across various abstraction levels—from the kernel and container runtime interface (CRI) to Kubernetes API and cloud contexts—and tracing the temporal progression of an attack chain.

Key moments
- 02:00 Kubernetes Attack Chain & Modern Trends
- 04:30 Multi-Domain Detection Strategy
- 05:00 Kubernetes Audit Logs: Capabilities & Limitations
- 09:00 Admission Webhooks for Detection
- 15:00 Runtime Sensor: Deep Host & Container Visibility
- 18:00 Runtime Sensor: Key Events Monitored
- 23:00 Demo: EKS Pod Identity Privilege Escalation & Cloud Pivot
- 28:00 Conclusion: Multi-Dimensional Detection Framework
Effective Detection in Kubernetes Clusters
Speakers: Shay Berkovich, Oren Ofer
Conference: BSidesSF 2024
YouTube: https://www.youtube.com/watch?v=JWCPufW91iY
Overview
This presentation, delivered by Shay Berkovich and Oren Ofer at BSidesSF 2024, delves into the complexities of detecting sophisticated attacks within Kubernetes and cloud-native environments. The speakers, both seasoned security professionals with backgrounds in threat research and detection engineering, highlight the evolving landscape of Kubernetes attacks, which increasingly incorporate cloud components, leverage new initial access vectors, and adapt older techniques for modern contexts. The core thesis of their talk is that effective detection in this intricate ecosystem necessitates a multi-dimensional approach, combining visibility across various abstraction levels—from the kernel and container runtime interface (CRI) to Kubernetes API and cloud contexts—and tracing the temporal progression of an attack chain.
The talk addresses the significant challenge faced by Security Operations Center (SOC) analysts, who are expected to be experts across a vast array of technologies in the modern cloud landscape. Berkovich and Ofer propose a solution that integrates diverse data sources, including Kubernetes audit logs, admission webhooks, cloud logs, and runtime sensors. They illustrate this integrated approach through a detailed analysis of a real-world attack scenario, demonstrating how different detection mechanisms contribute to building a comprehensive picture of malicious activity. The presentation emphasizes the critical role of intelligent noise reduction and the strategic correlation of events from disparate sources to achieve robust and actionable threat detection.
Background
▶ Watch: Kubernetes Attack Chain & Modern Trends (02:00)
Kubernetes, while powerful, is described as a "complicated Beast" with its own unique ecosystem and security domain. Attackers, however, approach it like any other system, following a familiar attack chain: initial access, lateral movement, privilege escalation, potential cloud pivot, and ultimately, impact (either local or global). The speakers emphasize that understanding this simplified attack chain is crucial for developing effective detection strategies.
Modern attacks in Kubernetes environments exhibit several evolving trends. Firstly, attackers are increasingly proficient with cloud components, demonstrating expertise in cloud APIs and utilizing red team tools such as Prowler and Pacu. This proficiency means that attacks often extend beyond the Kubernetes cluster itself into the underlying cloud infrastructure. Secondly, new initial access vectors are continually emerging. While historically attackers like Team TNT might have targeted exposed Docker API sockets, they have rapidly shifted to exploiting Remote Code Execution (RCE) vulnerabilities in applications like WordPress plugins, misconfigured PostgreSQL databases, and abusing leaked cloud credentials. Thirdly, there's a notable trend of adopting old techniques in new ways. An example cited is "Pyos," an in-memory execution of crypto miners using Python, which, while not a novel technique in itself, finds new application within cloud-native environments.
These trends collectively place immense pressure on SOC analysts, who must possess expertise across an ever-expanding array of technologies and attack methodologies. To counter this, the speakers advocate for a multi-domain detection strategy. This strategy operates along two axes: the temporal domain, which involves tracing the entire attack chain from initial access to impact, and the vertical axis, which represents different abstraction levels, requiring the combination of multiple security data sources for comprehensive detection. These sources include runtime sensors (providing kernel-level visibility), Kubernetes admission webhooks and audit logs, and various cloud detection feeds and logs (such as VPC flow logs).
Key Findings
▶ Watch: Kubernetes Audit Logs: Capabilities & Limitations (05:00)
The central finding of the presentation is that effective detection in Kubernetes clusters is inherently multi-dimensional, requiring a holistic approach that integrates diverse data sources across different abstraction levels and traces the temporal progression of an attack. No single detection source provides complete visibility, and each has distinct strengths and blind spots.
Key findings include:
- Necessity of Multi-Source Integration: A robust detection strategy must combine Kubernetes audit logs, admission webhooks, cloud logs (e.g., CloudTrail, GuardDuty), and runtime sensors. Each source offers unique insights into different layers of the cloud-native stack, from kernel-level activities to Kubernetes API calls and cloud API interactions.
- Challenges with Managed Kubernetes Audit Logs: While Kubernetes audit logs are a fundamental source, managed clusters (like AKS, GKE) present significant hurdles. These include audit logs not always being enabled by default (except in GKE), logs being stored in proprietary formats (e.g., GKE's format), and the inability to control audit options, which dictates what events are logged. This necessitates maintaining various rule sets or normalizing events, adding complexity.
- Criticality of Noise Reduction: The volume of logs generated by Kubernetes and cloud environments can be overwhelming. Intelligent noise reduction techniques are paramount for all log sources to ensure that SOC analysts can focus on truly interesting and potentially malicious activities, thereby reducing false positives and managing costs.
- Runtime Sensors for Deep Visibility: Runtime sensors are highlighted as crucial for detecting low-level, host- and container-centric activities that are invisible to API-level logs. They provide behavioral monitoring, collect forensic data, and offer prevention capabilities, making them the "boots on the ground" for identifying evasive techniques like in-memory execution or container drift.
- Cloud Logs for Perimeter and Enrichment: Cloud logs are essential for monitoring the perimeter of the Kubernetes cluster, detecting cloud pivots, and enriching alerts with contextual information such as the acting principal's history, associated privileges, origin IP, and resource sensitivity.
- Connecting Disparate Events: A significant challenge and key finding is the need to correlate events across different sources and potentially different acting principles to reconstruct the full attack chain. For instance, connecting an attacker's access to an EKS Pod Identity file within a cluster to subsequent AWS API calls made from a different IP address requires sophisticated correlation logic.
- Admission Webhooks as a Complementary Source: While offering less coverage than audit logs, admission webhooks can be valuable in audit mode for monitoring state-changing events. They provide full control over what is monitored and are cloud-agnostic, offering cost savings due to a lower event volume.
Ultimately, the speakers conclude that effective detection is achieved by intelligently combining these sources, applying smart noise reduction, and developing the capability to connect the dots across the entire temporal and vertical attack landscape.
Technical Deep Dive
▶ Watch: Runtime Sensor: Deep Host & Container Visibility (15:00)
The speakers meticulously break down the various data sources available for detection in Kubernetes, detailing their architecture, capabilities, limitations, and best practices for utilization.
Kubernetes Audit Logs
Kubernetes audit logs are presented as a foundational detection source, akin to a "CCTV system" for the cluster. They are an integral part of the master node and are controlled by the Kubernetes API server. Each audit log event, in its vanilla form, typically includes three core pillars: the username (principal), the verb (action), and the resource (object). For example, an event might show system:anonymous listing pods in the default namespace, which is a high-interest activity.
However, managing audit logs in managed Kubernetes clusters (like AKS, GKE) introduces several complexities:
- Default State: Audit logs are not always on by default; GKE is an exception where they are enabled.
- Log Mediums: The medium for retrieving logs varies across cloud providers (e.g., CloudWatch, Event Hub), requiring different integration strategies.
- Audit Options: Crucially, the
audit optionsfile, which dictates what events the Kubernetes API server logs, is controlled by the cloud vendor in managed clusters. This limits a user's ability to fine-tune logging for specific needs. - Proprietary Formats: Cloud providers often transform the original Kubernetes audit events into their proprietary formats (e.g., GKE's format is "completely different"). This necessitates maintaining multiple rule sets or normalizing events, adding significant operational overhead.
Noise reduction is critical due to the potentially massive volume of audit logs. Recommended techniques include:
- Monitoring Interesting Actions: Filtering out common, low-interest actions like
get,list, orwatchunless they pertain to sensitive resources likesecrets. - Weeding Out Control Plane Principles: Excluding principles starting with
system:(e.g.,system:kube-scheduler) except forsystem:anonymousorsystem:serviceaccount, which can indicate malicious activity. - Excluding Noisy Principles: Identifying and excluding known noisy principles, such as a Prometheus server.
An example rule discussed is detecting a resource created by an anonymous principal (e.g., anonymous and unauthenticated in principalEmail within a GCP event JSON for a create event with granted: true). This is considered high severity, potentially indicating an enabled unauthenticated mode in the cluster.
Blind spots for Kubernetes audit logs include:
- Local Worker Node Operations: Activities that occur directly on the worker node, such as static pods, direct socket API interactions, or events within the Kubelet logs, are not visible to the API server and thus not captured by audit logs.
Admission Webhooks
Admission webhooks are traditionally seen as enforcement mechanisms, but the speakers propose using them in an audit mode for detection. This involves defining a mutating or validating webhook that simply collects events without enforcing policy.
Coverage: Admission webhooks offer less coverage than Kubernetes audit logs because they primarily focus on state-changing events within the cluster.
Benefits:
- Cost Savings: Fewer events mean lower logging and processing costs.
- Full Control: Users define the webhook configuration, gaining precise control over what events are monitored.
- Cloud Agnostic: The configuration is consistent across different cloud environments, eliminating the need for multiple rule sets.
An admission webhook event, typically an AdmissionReview object, still contains essential components like the principal, operation (e.g., create), and resource type (e.g., ClusterRole), allowing for detection rule creation.
Noise reduction for admission webhooks can involve:
- Starting Broad, Then Refining: Initially monitoring everything (
*) and then narrowing down to specific resource types (e.g.,pods) or operations (e.g.,create pods). - Filtering Control Plane Principles/Actions: Ignoring events from namespaces like
kube-systemorkube-node-lease.
An example rule highlighted is detecting the connect operation, which often corresponds to kubectl exec commands into a pod. While not always high severity, it's an event that defenders typically want to be aware of.
Blind spots for admission webhooks are similar to audit logs, with additional limitations:
- No Impersonation Information: Details about impersonated users are not available.
- No Get/List/Watch Verbs: Only state-changing events are captured, not read-only operations.
- Invisible Denied Requests: Webhooks operate after authentication and authorization, so failed attempts are not visible.
Cloud Logs
Cloud logs are conceptualized as "neighborhood CCTV," providing the broader context to understand an attack, such as tracking a "getaway car" after a "burglary." They help piece together what happened and where attackers disappeared.
Key cloud log sources include:
- Cloud Audit Logs: Services like AWS CloudTrail, Azure Activity Log, and GCP Cloud Audit Logs record API calls and events within the cloud environment.
- Detection Feeds: Vendor-specific detection services like AWS GuardDuty and Microsoft Defender for Cloud offer pre-built rule sets, including Kubernetes-specific detections (e.g., "privileged container detected"). While these incur costs, they can be used for enrichment or as a perimeter monitoring solution.
- Other Logs: DNS queries, VPC flow logs, etc., provide additional network-level visibility.
Use cases for cloud logs:
- Enrichment: Providing additional context to alerts, such as the acting principal's history, associated cloud privileges, origin IP history, reputation, geolocation, and the sensitivity of the subject resource.
- Kubernetes Perimeter Monitoring: Detecting events that originate from the Kubernetes cluster but interact with the cloud API. Examples include "series of IAM enumeration attempts originated from an EKS container" or "MDS Connection in an EKS worker node with high privileges," which could signify a cloud pivot.
The speakers note the complexity of cloud rules, often involving intricate JSON structures, but emphasize their importance for monitoring the boundary between Kubernetes and the broader cloud environment.
Runtime Sensor (Agent Smith)
The runtime sensor is described as the "agent with boots on the ground," providing deep visibility at the kernel level within containers and on the host. It plays a crucial role in monitoring the MITRE ATT&CK framework techniques and procedures.
Functionality:
- Behavioral Monitoring: Observes application and user behavior from containers up to the node, identifying deviations from standard patterns.
- Data Collection: Gathers crucial data for detection, including Indicators of Compromise (IOCs), which can be integrated with threat intelligence feeds or streamed to a SIEM for correlation.
- Prevention and Response: Offers automated capabilities to kill processes, terminate containers, block network connections, and collect forensic packages for incident investigation.
Instrumentation Techniques:
- User Mode Hooks: Placing hooks on user-mode functions or debugging processes to inspect arguments.
- Kernel Modules: Inspecting system-wide events (system calls) and internal kernel structures.
- eBPF (Extended Berkeley Packet Filter): Highlighted as a modern, preferred technology that combines the benefits of both user mode and kernel mode instrumentation while offering inherent protection against impacting worker node functionality.
Events to Instrument:
- Process and Thread Events: Collecting binary information (for hash calculation and malware identification), command lines (for detecting LOLBins abuse).
- File Events: Detecting credential access, privilege escalation, and persistence activities.
- Network Events: Gathering information on local/remote endpoints, DNS queries, and socket usage to detect connections to C2 servers or cryptopools.
- Memory Events: Identifying in-memory payloads, MFDs, and other evasion techniques.
- Module Load/Unload Events: Detecting rootkits, process execution hijacks, and other "shenanigans."
- Linux Audit Logs: Collecting information from system services and applications (e.g., SSHD printouts for brute-force attempts).
- Container Runtime Information: Matching running processes to container images to identify "container drift" (unauthorized tooling installation).
Deployment: Runtime sensors are typically deployed as a DaemonSet in Kubernetes, ensuring a sensor pod runs on every worker node.
Considerations:
- Managed Nodes: Cannot be deployed on managed master nodes.
- Proprietary Application Context: Benign application access to sensitive files might generate alerts, but this can lead to productive discussions about application design or be allow-listed.
- Permissions: The sensor should ideally not have high permissions to access Kubernetes or cloud contexts directly, as that information can be obtained from other dedicated sources, aligning with best practices.
Demo / Proof of Concept
▶ Watch: Runtime Sensor: Key Events Monitored (18:00)
The speakers presented a demo based on real-world attack events observed at their customer base, illustrating a multi-stage attack culminating in a cloud pivot and data exfiltration.
Initial Access: The attack begins with initial access gained through a Jupyter notebook, offered as a free service by some vendors.
Attacker Actions and Detection Analysis:
- Tool Installation: The attacker installs tools, including
kubectl, within the Jupyter notebook environment.
- Detection: This activity would be visible on VPC flow logs (for external connections), DNS queries, and most effectively, by the runtime sensor monitoring process execution and file system changes.
- Cluster Interrogation: The attacker uses
kubectl auth can-ito check permissions, quickly realizing they have very high privileges. They then identify available namespaces and search for Kubernetes secrets, though this yields no immediate interesting results.
- Detection: The
kubectl auth can-icommand is an API call to the Kubernetes API server, making it primarily visible in Kubernetes audit logs. The runtime sensor might see thekubectlprocess execution, but the API call details are in the audit logs.
- Crypto Miner Deployment: The attacker deploys a crypto miner, specifically XM rig, into the environment. They then verify its deployment using
kubectl get pod.
- Detection: The deployment of the crypto miner (creating a new deployment/pod) is a state-changing event visible in Kubernetes audit logs and potentially admission webhooks. The execution of the
XM rigprocess within the new pod is a prime detection point for the runtime sensor.
- Cloud Pivot - EKS Pod Identity Abuse: The attacker seeks to escalate privileges and pivot into the cloud. They search for pods associated with EKS Pod Identity (a relatively new AWS feature) and identify a
bucket-readerpod. They then execute akubectl execcommand into this pod tocatthe file containing the EKS Pod Identity token.
- Detection: The
kubectl execcommand is visible in Kubernetes audit logs (as aconnectverb) and admission webhooks. The file access to the EKS Pod Identity token file within the pod is a critical detection point for the runtime sensor.
- AWS API Token Generation and S3 Access: Using the extracted EKS Pod Identity token, the attacker generates a valid AWS API token. They then create an S3 session, validate it by getting the ARN (
get-caller-identity), list the first S3 bucket, and finally download files, successfully exfiltrating sensitive data.
- Detection: These actions, involving AWS API calls, are primarily visible in cloud logs (e.g., AWS CloudTrail). The runtime sensor might detect network connections to AWS endpoints. A crucial challenge here is that once the AWS API token is obtained, the attacker could carry out subsequent actions from any machine globally, potentially using a different IP address and a new principal (the role assumed via the EKS Pod Identity).
Connecting the Stages: The speakers emphasize that connecting these disparate stages is the "tricky part." For instance, linking the access to the EKS Pod Identity file (detected by the runtime sensor) to the subsequent activation of that role and AWS API calls (detected by cloud logs, potentially from a different country or IP) requires sophisticated correlation logic. This involves understanding that the accessed file points to a specific role, and then monitoring for that role's activity in the cloud logs within a relevant timeframe. This multi-source correlation is fundamental to building a complete picture of the attack.
Defensive Implications
▶ Watch: Conclusion: Multi-Dimensional Detection Framework (28:00)
The insights from this presentation offer several critical defensive implications for organizations operating Kubernetes clusters in cloud-native environments:
- Adopt a Multi-Dimensional Detection Strategy: Defenders must move beyond relying on single-source monitoring. Implement a layered approach that integrates Kubernetes audit logs, admission webhooks, runtime sensors, and cloud logs to cover the full attack surface across temporal and abstraction levels.
- Optimize Kubernetes Audit Log Configuration:
- Ensure audit logs are enabled in all Kubernetes clusters, especially in managed environments where they might not be on by default.
- Prioritize monitoring for high-severity events, such as resource creation by anonymous principals or sensitive resource access.
- Implement robust noise reduction techniques by filtering out benign control plane activities and known noisy principles to reduce alert fatigue and improve signal-to-noise ratio.
- Develop strategies to normalize proprietary log formats from managed cloud providers to ensure consistent rule application.
- Leverage Admission Webhooks Strategically:
- Consider deploying admission webhooks in "audit mode" to monitor state-changing events. This provides granular control over specific events of interest (e.g., pod creation,
kubectl execcommands) with potentially lower cost compared to full audit log ingestion. - Use them to complement audit logs, especially for cloud-agnostic rule sets.
- Deploy and Configure Runtime Sensors Effectively:
- Deploy runtime sensors (e.g., via a DaemonSet) on all worker nodes to gain deep visibility into container and host-level activities.
- Configure sensors to monitor critical events such as process execution (for LOLBins, crypto miners), file access (for credentials, persistence), network connections (for C2, cryptopools), and memory events (for in-memory payloads).
- Utilize the prevention and response capabilities of runtime sensors to automatically mitigate threats and collect forensic packages for post-incident analysis.
- Address "container drift" by monitoring for unauthorized software installations within containers.
- Integrate Cloud Logs for Perimeter Monitoring and Enrichment:
- Ensure comprehensive logging of cloud API calls (e.g., AWS CloudTrail, Azure Activity Log).
- Configure cloud-native detection services (e.g., AWS GuardDuty, Microsoft Defender for Cloud) to monitor for Kubernetes-related threats and cloud pivots.
- Use cloud logs to enrich alerts with crucial context, such as the identity of the acting principal, their privileges, origin IP, and the sensitivity of affected resources.
- Develop rules to detect suspicious interactions between Kubernetes components and the cloud API, which often signify a cloud pivot.
- Develop Advanced Correlation Capabilities:
- Invest in security information and event management (SIEM) or security orchestration, automation, and response (SOAR) platforms that can ingest and correlate data from all these disparate sources.
- Build correlation logic to connect events across different abstraction layers and potentially different identities (e.g., linking an EKS Pod Identity file access to subsequent AWS API calls made under that role). This is crucial for reconstructing the full attack chain and understanding the attacker's true intent.
- Be Aware of Blind Spots: Understand the inherent limitations of each detection source and design your overall strategy to compensate for these blind spots by leveraging other sources. For example, audit logs won't see local worker node operations, which is where runtime sensors excel.
- Regularly Review and Update Detections: The threat landscape is constantly evolving. Regularly review and update detection rules and strategies to account for new attack techniques, changes in cloud provider logging, and internal application behavior.
Key Takeaways
- Multi-Dimensional Detection is Essential: Effective detection in Kubernetes requires combining visibility across temporal (attack chain) and vertical (abstraction levels from kernel to cloud) domains.
- Integrate Diverse Data Sources: Leverage Kubernetes audit logs, admission webhooks, runtime sensors, and cloud logs, as each offers unique insights and covers different blind spots.
- Noise Reduction is Critical: Implement smart filtering and exclusion rules across all log sources to manage volume, reduce costs, and minimize false positives for SOC analysts.
- Runtime Sensors Provide Deep Behavioral Insight: These "boots on the ground" are indispensable for detecting low-level process, file, network, and memory activities within containers and on worker nodes, including evasive techniques.
- Cloud Logs Enable Perimeter Monitoring and Contextual Enrichment: Cloud audit logs are vital for detecting cloud pivots, monitoring interactions between Kubernetes and the cloud API, and enriching alerts with crucial contextual information about principals and resources.
- Correlation is Key to Full Attack Visibility: Developing the capability to connect disparate events across different sources and potentially changing identities is paramount for reconstructing the complete attack chain and understanding attacker intent.
About the Speaker(s)
Shay Berkovich brings a strong background in software development, holding a Masters of Science in Random Verification from the University of Waterloo. His career transitioned into security research approximately 12 years ago when he began coding for Web Application Firewalls (WAFs) and delved into Application Security (AppSec). He currently works as a security researcher on a threat research team, focusing on the evolving landscape of cloud-native security.
Oren Ofer is a seasoned detection engineer with eight years of experience in the field, specializing in runtime sensor development across various environments including Windows endpoints, cloud platforms, Linux, and naturally, Kubernetes clusters. Prior to his work in detection engineering, Oren honed his skills in penetration testing and red teaming, a "hacker mindset" that he now applies to designing robust detection mechanisms.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk provides a highly practical and well-structured approach to effective detection in Kubernetes and cloud-native environments. The speakers meticulously break down various detection sources—Kubernetes audit logs, admission webhooks, cloud logs, and runtime sensors—highlighting their strengths, blind spots, and optimal use cases. The live demo of a multi-stage attack, including an EKS Pod Identity pivot, effectively illustrates the necessity of combining these diverse data sources for comprehensive threat visibility and response.
Heather Calloway (CISO) — STRONG ACCEPT
This presentation offers a pragmatic and essential guide for establishing effective detection capabilities within Kubernetes environments, a critical area for modern enterprises. The speakers systematically dissect various logging and monitoring sources, from Kubernetes audit logs to runtime sensors, providing a clear understanding of their respective strengths and limitations. The demonstration of a multi-stage attack, culminating in a cloud pivot and data exfiltration, powerfully underscores the necessity of a layered, integrated detection strategy to manage significant business risks.