Superpowers for Humans of Kubernetes: How K8sGPT Is Transforming Enter... Alex Jones & Anais Urlichs

Alex Jones, Anais Urlichs

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

The rapid adoption and increasing complexity of Kubernetes environments present significant challenges for platform engineering and SRE teams. In their KubeCon EU talk, Alex Jones, a Principal Engineer at AWS and the founder of K8sGPT, alongside Anais Urlichs, a Platform Engineer at JP Morgan Chase, unveiled how K8sGPT is revolutionizing the operational landscape. K8sGPT is introduced as an innovative, AI-enhanced tool designed to act as a 24/7 on-call SRE for Kubernetes clusters, providing intelligent debugging, triage, and, crucially, auto-remediation capabilities.

Watch on YouTube

Visual summary for Superpowers for Humans of Kubernetes: How K8sGPT Is Transforming Enter... Alex Jones & Anais Urlichs by Alex Jones, Anais Urlichs
Visual summary for Superpowers for Humans of Kubernetes: How K8sGPT Is Transforming Enter... Alex Jones & Anais Urlichs by Alex Jones, Anais Urlichs

Key moments

  1. 0:00 Welcome and introduction to K8sGPT
  2. 2:25 Core organizational goals for new tooling
  3. 3:15 How platform engineering scales efficiency
  4. 4:00 Tooling must evolve: more autonomous, less manual
  5. 4:40 Challenges: knowledge silos in complex environments
  6. 6:35 K8sGPT: your 24/7 expert Kubernetes SRE
  7. 7:00 How K8sGPT works: AI-enhanced SRE analysis

Superpowers for Humans of Kubernetes: How K8sGPT Is Transforming Enterprise Operations

Speakers: Alex Jones, Principal Engineer, AWS; Anais Urlichs, Platform Engineer, JP Morgan Chase

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=EXtCejkOJB0

Overview

The rapid adoption and increasing complexity of Kubernetes environments present significant challenges for platform engineering and SRE teams. In their KubeCon EU talk, Alex Jones, a Principal Engineer at AWS and the founder of K8sGPT, alongside Anais Urlichs, a Platform Engineer at JP Morgan Chase, unveiled how K8sGPT is revolutionizing the operational landscape. K8sGPT is introduced as an innovative, AI-enhanced tool designed to act as a 24/7 on-call SRE for Kubernetes clusters, providing intelligent debugging, triage, and, crucially, auto-remediation capabilities.

This talk delves into the core philosophy behind K8sGPT, born from Jones's personal frustrations as an SRE, aiming to codify and democratize the tacit knowledge often locked within individual engineers or scattered documentation. The project, now a CNCF Sandbox initiative, addresses critical organizational goals such as reducing time to market, maximizing developer efficiency, and achieving significant cost savings by streamlining operations and reducing the need for constant manual intervention in complex, distributed systems. The speakers emphasize K8sGPT's transition from a mere observability add-on to a powerful, autonomous system capable of not just identifying but also proactively fixing issues, thereby transforming how enterprises manage their cloud-native infrastructure.

Background

▶ Watch: Welcome and introduction to K8sGPT (0:00)

The journey towards modern, cloud-native infrastructure, particularly with Kubernetes, has brought unprecedented scalability and agility but also introduced layers of complexity that challenge even the most experienced engineering teams. As Anais Urlichs highlighted, many organizations are striving for zero-touch environments, where hundreds of clusters can operate with minimal human intervention. To achieve this, companies have often restructured, forming centralized platform engineering teams (sometimes called DevOps or SRE teams). Their primary responsibility is to provide reusable tooling, services, and infrastructure components, enabling other engineering teams to independently provision resources. The ultimate goal is to scale operations without linearly scaling the platform engineering team—doing "more with less."

However, this paradigm shift comes with inherent challenges. Platform teams, while experts in their chosen technologies, can inadvertently create knowledge silos, making other teams dependent on them. When components break, identifying the root cause can be time-consuming for application teams, often requiring platform engineers to be pulled in. The highly complex and rapidly evolving cloud-native landscape also exacerbates the problem of tacit knowledge: critical operational insights that reside solely with individual engineers. If these individuals leave, their invaluable experience and understanding of system intricacies are often lost, leading to long-term technical debt and the risk of operating infrastructure as a "black box" until catastrophic failures occur. This directly impacts organizational goals like compliance, consistent operations across distributed teams, and the reliability of mission-critical applications.

Existing autonomous tooling, such as GitOps solutions like Argo CD and Infrastructure as Code tools like Crossplane, have made strides in deterministic infrastructure setup and application deployment. Yet, as Jones observed, these tools are not designed to self-diagnose and self-heal when things inevitably go wrong. K8sGPT emerged from this gap, specifically from Alex Jones's frustration as an SRE who sought to codify behavioral tests and SRE knowledge into a system that could intelligently analyze and remediate issues, bridging the chasm between automated deployment and autonomous operation.

Key Findings

▶ Watch: How platform engineering scales efficiency (3:15)

K8sGPT's core contribution is its ability to act as an expert Kubernetes SRE on call 24/7, leveraging AI-enhanced analysis of codified SRE knowledge to debug and triage issues across diverse environments. Launched in 2023 and quickly donated to the CNCF as a Sandbox project, K8sGPT provides not just problem identification but also measured guidance on how to fix issues, effectively serving as a dependable sidekick for cluster health monitoring and remediation.

A pivotal advancement highlighted in the talk is the introduction of auto-remediation. This represents a philosophical shift for K8sGPT, transforming it from a passive observability tool into an active, tightly-looped reconciliation system. Instead of merely reporting errors, K8sGPT can now suggest and apply fixes based on an organization's internal knowledge, a capability that, to the speakers' knowledge, is rarely achieved with such simplicity and integration in the cloud-native ecosystem. This capability directly addresses the challenge of tacit knowledge loss by democratizing operational expertise, making it accessible and actionable for all engineers, including new starters.

Key features and interoperability aspects of K8sGPT include:

  • Flexible Deployment: Available as a CLI, Operator, in server mode, or interactive mode, allowing users to integrate it into various workflows.
  • Interoperable AI Backends: Supports 11 different AI backends, including Hugging Face, OpenAI, and Bedrock, with the option for users to "bring your own models" for tailored inference.
  • Custom Analyzers: Allows enterprises to build their own analyzers, embedding private, domain-specific knowledge about unique components and failure patterns into K8sGPT.
  • Prometheus Integration: K8sGPT metrics can be viewed alongside other operational metrics in existing dashboards, enabling holistic correlation.
  • CRD Results: Scan results are saved as Custom Resource Definitions (CRDs) within the cluster, facilitating interaction and integration with external tooling.
  • Data Anonymization: Addresses privacy concerns by allowing users to anonymize data sent to the LLM.
  • Multimodal Operation: Supports single-tenancy, multi-tenancy, and operation across multiple clusters, offering broad applicability.

The speakers also introduced a "flywheel" workflow that integrates K8sGPT into the traditional SRE incident response cycle. When SREs address incidents, their mitigation steps and root cause analyses (RCAs) are documented in tickets or knowledge bases (e.g., Jira, Confluence). This data is then exported, embedded into a vector database (a Retrieval Augmented Generation, or RAG, system), and consumed by K8sGPT's foundational models. This continuous feedback loop ensures that every manual fix makes K8sGPT smarter, leading to more accurate suggestions for manual remediation and more intelligent decisions for auto-remediation, truly changing the game for operational efficiency.

Technical Deep Dive

▶ Watch: Tooling must evolve: more autonomous, less manual (4:00)

The technical architecture of K8sGPT, particularly in its operator mode, is designed for robustness, scalability, and seamless integration into existing Kubernetes environments. When deployed, the K8sGPT operator acts as the core orchestrator. This operator is responsible for deploying K8sGPT deployments, which are effectively decoupled instances of the K8sGPT CLI logic. This pattern was chosen to support multi-tenancy, allowing for isolated operations within v-clusters or specific namespaces, and to facilitate asynchronous activities and data aggregation from multiple clusters.

At the heart of K8sGPT's diagnostic capabilities are its analyzers. These are modular components that provide K8sGPT with an understanding of Kubernetes resources (e.g., what constitutes a Pod, a Deployment) and common failure modes. Beyond built-in analyzers, users can create custom analyzers to codify unique components and failure patterns specific to their environment, integrating private organizational knowledge directly into the diagnostic process.

Once an analyzer identifies a potential issue, the K8sGPT deployment sends this information to an inference API. This API can be powered by any of the 11 supported AI backends (e.g., local AI, AWS Bedrock, Hugging Face) or a user's custom model. The AI processes the issue and suggests a solution. Crucially, K8sGPT then codifies this issue and its proposed solution, creating a hashed key-value lookup. This allows for efficient retrieval and application of known fixes.

The new auto-remediation workflow introduces several layers of intelligence and safety:

  1. Credibility Inspection: Before any remediation is suggested or applied, K8sGPT inspects the AI-generated result for credibility. This involves a series of internal "safeguard guardrails" to determine if the identified problem is a genuine issue or a false positive.
  2. Applicability Calculation: K8sGPT assesses the applicability of a suggested fix. For instance, if an image is reported as missing, K8sGPT can query an internal knowledge base (e.g., a RAG-powered data store) to verify if the suggested replacement image is known and approved within the organization. This ensures that only relevant and enterprise-approved solutions are considered.
  3. Resource Tree Construction: If the fix is deemed credible and applicable, K8sGPT constructs a resource tree, which essentially represents a "git patch" or a delta of change required to resolve the issue. This delta is then stored in an immutable log through a Mutation Custom Resource (CRD).
  4. Mutation CRD: This custom resource provides a chronological account of the auto-remediation process. It records the original state of the resource, the suggested change from the AI, the actual applied change, and the subsequent outcome. This audit trail is critical for transparency and debugging.
  5. Success Determination: After applying a mutation, K8sGPT re-probes the analyzer. If the original issue is no longer detected, the mutation is considered successfully applied.
  6. Safeguards: To build confidence in auto-remediation, K8sGPT employs several tricks, such as calculating a relatively low Levenshtein distance (a measure of string similarity) for patching differences, ensuring that changes are minimal and targeted. Users can configure a similarity requirement percentage, defining the acceptable delta of change for auto-remediation.

The integration with an organization's knowledge base via a RAG system is a fundamental aspect of K8sGPT's long-term intelligence. SREs' daily interactions with incidents—mitigations, root cause analyses, and resolutions recorded in ticketing systems or internal wikis—can be exported. This unstructured data is then embedded into a vector database, forming a rich knowledge base. K8sGPT's foundational models can then query this vector database, making its inference API progressively smarter. This feedback loop ensures that K8sGPT's auto-remediation decisions become increasingly accurate and aligned with an organization's specific operational patterns and approved solutions.

Privacy is a significant concern, especially when dealing with sensitive cluster data. K8sGPT addresses this by allowing users to anonymize the data sent to the LLM backends, ensuring that proprietary or sensitive information is not exposed. Furthermore, K8sGPT's integration with Prometheus allows operational metrics to be collected and visualized, providing a unified view of cluster health alongside K8sGPT's diagnostic insights. Its multimodal capabilities ensure it can operate effectively in diverse environments, from single-cluster, single-tenant setups to complex multi-cluster, multi-tenant architectures.

Demo / Proof of Concept

▶ Watch: K8sGPT: your 24/7 expert Kubernetes SRE (6:35)

Alex Jones conducted a live demonstration illustrating K8sGPT's auto-remediation capabilities in action. The setup involved using standard Kubernetes tooling: helm for deploying K8sGPT and kind for spinning up a local Kubernetes cluster. The demonstration aimed to fix a deliberately introduced misconfiguration.

The scenario involved a "broken deployment" where a "very naughty person on a Friday" had incorrectly named a deployment engine xxx instead of the expected nginx. This simple yet illustrative example highlighted a common typo-related issue that can plague even seasoned engineers.

The K8sGPT configuration for the demo included:

  • auto remediation: true: Enabling the auto-remediation feature.
  • similarity requirement: 25%: Setting a baseline for the acceptable similarity between the suggested fix and the original state, indicating how much change K8sGPT is allowed to apply.

Upon deployment, the K8sGPT operator detected the misconfigured deployment. The system then initiated its auto-remediation workflow. The audience observed a Mutation Custom Resource (CRD) being created in the cluster, showing its in progress status and a similarity score of 27%, which successfully met the configured 25% requirement. This indicated that K8sGPT had identified a suitable patch that was within the acceptable delta of change.

Following the mutation, K8sGPT successfully created a new replica set and a new deployment, correctly named nginx. The original engine xxx deployment was remediated, demonstrating K8sGPT's ability to not only identify the problem but also apply the necessary fix autonomously. Jones emphasized that this process involves inherent benefits like the ability to compare and contrast the fixed and broken resources, with K8sGPT saving this information for future reference.

A notable aspect of the demo was the ability to trace the origin of the fix: "Find more info in zero ticket cube 123." While a fictitious example, this highlighted K8sGPT's integration with an organizational knowledge base, allowing engineers to understand the context and history behind a remediation, reinforcing the "flywheel" concept of learning from past incidents. The successful live demo underscored K8sGPT's practical utility and the reliability of its auto-remediation logic, even in real-time, unpredictable conference Wi-Fi conditions.

Defensive Implications

▶ Watch: How K8sGPT works: AI-enhanced SRE analysis (7:00)

K8sGPT offers profound defensive implications for organizations grappling with the complexity of cloud-native operations. The primary goal is to enable truly zero-touch environments, where hundreds of Kubernetes clusters can be managed with minimal human intervention. This vision directly addresses the scalability challenges faced by platform engineering teams, allowing them to do "more with less" by automating routine diagnostic and remediation tasks that would otherwise require significant manual effort.

While existing open-source tooling like Argo CD and Crossplane excel at automated deployments and infrastructure provisioning, they often lack the intelligence to self-diagnose and self-heal when issues arise. K8sGPT fills this critical gap by acting as an auto-SRE on call 24/7, equipped with codified knowledge from previous incidents. This means that common, recurring issues—such as an Argo CD application deployment hanging in a sync loop (a problem mentioned by Urlichs, requiring a manual deletion to trigger a healthy pull)—can be automatically detected and fixed by K8sGPT, preventing engineers from having to write custom scripts or operators for every unique problem across potentially hundreds of clusters.

A key defensive strategy K8sGPT promotes is the creation of a robust knowledge base or RAG (Retrieval Augmented Generation) system. As Urlichs articulated, you cannot expect an AI model to know what to do in your specific setup without providing it with additional context. This is akin to an "open book exam" for the AI. By feeding K8sGPT with an organization's internal SRE knowledge, incident reports, and successful remediation steps, the AI becomes an invaluable repository of operational wisdom. This allows K8sGPT to take long-term actions that humans would otherwise perform, ensuring consistency and reducing reliance on individual expertise.

Furthermore, K8sGPT redefines the role of human engineers. Instead of replacing junior and mid-level engineers, it empowers them. When a broken cluster or workload is encountered, K8sGPT can provide immediate, context-rich information, linking to relevant resources and suggesting informed decisions. This not only reduces the mean time to resolution (MTTR) but also serves as an educational tool, accelerating the learning curve for new and less experienced SREs. More importantly, K8sGPT can auto-remediate issues before a human SRE is even paged, significantly improving system uptime and reducing operational fatigue.

Looking ahead, the roadmap for K8sGPT includes critical enhancements that bolster its defensive posture:

  • GitOps Support for Auto-Remediation: Recognizing that many organizations operate under a strict GitOps model, K8sGPT plans to introduce support for a PR-based remediation workflow. Instead of directly applying fixes to clusters, K8sGPT will track changes against a Pull Request (PR) in the source repository, watch its lifecycle (review, merge), and then confirm the remediation's success after the PR is applied. This ensures that all changes adhere to established GitOps principles, maintaining auditability and control.
  • Enhanced Guardrails for Patch Safety: To build even greater confidence in auto-remediation, K8sGPT will incorporate a highly specific agent designed to detect changes in the YAML itself. This agent will analyze the delta of a patch and determine if it's "safe," especially for subtle annotation changes. This additional layer of scrutiny is crucial for enterprises to feel secure in enabling full auto-remediation, mitigating risks associated with unintended or potentially harmful automated modifications.

By continuously learning from human SRE interactions, providing intelligent assistance, and autonomously remediating issues within defined guardrails, K8sGPT transforms reactive incident response into proactive, intelligent operations, making Kubernetes environments more resilient and easier to manage.

Key Takeaways

  • AI-Enhanced SRE for Kubernetes: K8sGPT functions as an intelligent, 24/7 on-call SRE, leveraging AI to debug, triage, and provide remediation guidance for Kubernetes cluster issues.
  • Auto-Remediation Capability: It marks a significant shift from mere observability to active problem-solving, autonomously applying fixes based on codified knowledge and organizational guardrails.
  • Democratization of Tacit Knowledge: K8sGPT addresses the challenge of lost institutional knowledge by integrating SRE expertise and incident data into a knowledge base, making it accessible and actionable for all engineers.
  • Flexible and Interoperable: The tool supports 11 different AI backends, allows for custom analyzers, integrates with Prometheus, and saves results as CRDs, ensuring broad applicability and easy integration into existing ecosystems.
  • Enabling Zero-Touch Environments: K8sGPT is crucial for scaling operations by reducing manual intervention, allowing platform teams to manage hundreds of clusters more efficiently and reliably.
  • Future-Proofing with GitOps and Enhanced Guardrails: Upcoming features like GitOps-native auto-remediation and advanced YAML patch safety checks aim to build greater enterprise confidence and adherence to modern operational practices.

About the Speaker(s)

Alex Jones is a Principal Engineer at AWS. He is the visionary and founder behind K8sGPT, a project he initiated two years prior to this talk out of personal frustration as an SRE. His motivation stemmed from a desire to codify behavioral tests into more robust integration-style tests, which ultimately evolved into the K8sGPT project.

Anais Urlichs is a Platform Engineer at the accelerator at JP Morgan Chase, based in London. In her presentation, she emphasized that the content of her talk was independent of her employer or her day-to-day work, focusing on her expertise in platform engineering and cloud-native technologies.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Dr. Kozlov finds K8sGPT to be a genuinely impressive and substantive defensive innovation for Kubernetes operations. Its core strength lies in leveraging AI-enhanced analysis for not just debugging and triage but also intelligent auto-remediation, effectively acting as a 24/7 SRE. The project's emphasis on codifying tacit SRE knowledge via a RAG-powered feedback loop and its robust auto-remediation workflow with explicit guardrails addresses critical challenges in managing complex cloud-native environments, promising significant practical impact by enabling more autonomous and resilient operations.

Heather Calloway (CISO) — MUST SEE

This session on K8sGPT is a crucial presentation for any CISO or security leader navigating the complexities of modern cloud-native infrastructure. It directly addresses critical institutional challenges like knowledge silos, operational scaling, and incident response, offering a pragmatic, AI-enhanced solution for autonomous Kubernetes management. The focus on auto-remediation, coupled with robust guardrails, auditability, and the integration of organizational knowledge, positions K8sGPT as a strategic tool for enhancing operational resilience, reducing business risk, and driving efficiency across large-scale environments. This isn't just about technical optimization; it's about…

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025