Driving Chaos Engineering Forward: What’s New and Next With LitmusCha... Sarthak Jain & Saranya Jena

Sarthak Jain, Saranya Jena

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk, presented by Sarthak Jain and Saranya Jena, both Senior Software Engineers and maintainers of the project at Harness, provides a comprehensive update on LitmusChaos, a CNCF incubating project dedicated to Chaos Engineering for cloud-native applications. The presentation delves into the project's recent achievements, ongoing developments, and ambitious future roadmap, emphasizing its mission to make Chaos Engineering secure and accessible. As a critical tool for verifying the resilience of business services, LitmusChaos empowers DevOps pipelines to build more robust code, capable of withstanding software and infrastructure faults.

Watch on YouTube

Visual summary for Driving Chaos Engineering Forward: What’s New and Next With LitmusCha... Sarthak Jain & Saranya Jena by Sarthak Jain, Saranya Jena
Visual summary for Driving Chaos Engineering Forward: What’s New and Next With LitmusCha... Sarthak Jain & Saranya Jena by Sarthak Jain, Saranya Jena

Key moments

  1. 0:00 Introduction to Litmus Chaos and its mission
  2. 1:00 Early timeline: Ansible, Chaos Operator, GoLang migration
  3. 3:10 Litmus 1.0 GA and reaching CNCF Sandbox status
  4. 4:20 Litmus 2.0 GA and Chaos Control Plane launch
  5. 5:00 Achieving CNCF Incubation and Litmus 3.0 release
  6. 6:10 Latest features: resiliency probes and score-driven UX
  7. 7:20 Litmus Chaos adoption, statistics, and future goals

Driving Chaos Engineering Forward: What’s New and Next With LitmusChaos

Speakers: Sarthak Jain, Senior Software Engineer, Harness; Saranya Jena, Senior Software Engineer, Harness

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=6GjLzWtqjlw

Overview

This talk, presented by Sarthak Jain and Saranya Jena, both Senior Software Engineers and maintainers of the project at Harness, provides a comprehensive update on LitmusChaos, a CNCF incubating project dedicated to Chaos Engineering for cloud-native applications. The presentation delves into the project's recent achievements, ongoing developments, and ambitious future roadmap, emphasizing its mission to make Chaos Engineering secure and accessible. As a critical tool for verifying the resilience of business services, LitmusChaos empowers DevOps pipelines to build more robust code, capable of withstanding software and infrastructure faults.

The session highlights the significant growth and maturity of LitmusChaos, marked by its recent security audit, substantial community adoption, and continuous feature enhancements. With over 2 million installations and a remarkable 300% usage increase in the past year, LitmusChaos is solidifying its position as a leading open-source solution in the Chaos Engineering landscape. The speakers articulate the project's strategic goal of achieving CNCF graduation, underscoring its commitment to security, usability, and community engagement as it drives innovation in resilience testing.

Background

▶ Watch: Introduction to Litmus Chaos and its mission (0:00)

Chaos Engineering is a discipline that proactively introduces controlled failures into a system to identify weaknesses and build resilience. Its core principle is to verify the steady-state hypothesis of an application, ensuring it remains stable even under turbulent conditions. LitmusChaos emerged as a prominent tool in this domain, specifically tailored for cloud-native environments.

The journey of LitmusChaos began at KubeCon NA 2018, initially as an Ansible-based tool executing basic pod-level faults. Rapid evolution followed, with the introduction of the Chaos Operator and Chaos CRDs (Custom Resource Definitions) at KubeCon Europe 2019, marking a shift to Golang for improved support and adoption. A significant milestone was the publication of the principles of Cloud Native Chaos Engineering (CNCE) and the launch of Chaos Hub at KubeCon NA 2019, providing a centralized repository of pre-built chaos faults. The Bring Your Own Chaos (BYOC) feature further empowered users to create custom faults.

Litmus 1.0 was GA at KubeCon Europe 2020, followed by its entry into CNCF Sandbox at KubeCon NA 2020. This period saw the introduction of Litmus Probes for validating steady-state hypotheses and Chaos Exporter for exposing Prometheus metrics, enabling better visualization and analysis of experiment data. Further advancements included CI/CD integrations with GitHub and GitLab, the release of Litmus 2.0 GA, and the introduction of the Chaos Control Plane (a UI for managing infrastructure, experiments, and the Chaos Hub) with multi-tenancy and team support at KubeCon NA 2021.

By KubeCon Europe 2022, LitmusChaos achieved CNCF Incubation status, adding GitOps support for trigger-based experiment execution and air-gap support via Chaos Image Registries. The project expanded its fault library to include cloud infrastructure and VMware faults. LitmusChaos 3.0, announced at KubeCon NA 2022, brought runtime support for faults and application-level chaos, such as Spring Board fault. More recently, at KubeCon Europe 2023, Resiliency Probes were introduced, tunable via the control plane, alongside improvements in debuggability and a resiliency score-driven UX. The latest release, Litmus 3.7.0, focuses heavily on security enhancements, distributed tracing, and SDK support.

The project boasts impressive statistics: over 2 million installations, 68 million Docker pulls, a 300% usage increase in the last year, 21 active maintainers (with Harness as the primary maintainer), over 100 releases, and a vibrant community of more than 2500 Slack members and over 250 adopters. The next major objective for LitmusChaos is to achieve CNCF graduation, a testament to its stability, widespread adoption, and robust governance.

Key Findings

▶ Watch: Litmus 1.0 GA and reaching CNCF Sandbox status (3:10)

The talk primarily serves as an update on the current state and future direction of LitmusChaos, rather than presenting new research findings in the traditional sense. However, the key "findings" or accomplishments highlighted by the speakers revolve around significant advancements in security, usability, and extensibility, all geared towards its CNCF graduation goal.

A pivotal achievement is the successful completion of a security audit performed by the 7A Security team. This rigorous evaluation led to the identification and subsequent remediation of potential vulnerabilities, alongside the implementation of numerous security enhancements across API, authentication, network, and development practices. This commitment to security is paramount for a tool that operates directly within users' production clusters.

Another significant development is the continuous effort to enhance the user experience (UX) and debuggability. The introduction of a resiliency score-driven UX, coupled with improved error logs and messages, makes it easier for users to understand and act on the results of their chaos experiments. The project's expanding Chaos Hub and Bring Your Own Chaos (BYOC) capabilities continue to make a wide array of faults readily available and customizable, addressing diverse testing needs.

Looking forward, the project is making substantial strides in distributed tracing integration, aiming to provide unparalleled visibility into the impact and flow of chaos experiments. This feature, currently in development, promises to significantly simplify debugging and understanding complex interactions during fault injection. Furthermore, the development of SDKs for Java and Go underscores a commitment to seamless integration into modern CI/CD pipelines, making chaos engineering an integral part of the software development lifecycle. These collective efforts demonstrate LitmusChaos's evolution into a more secure, user-friendly, and deeply integrated platform for resilience testing.

Technical Deep Dive

▶ Watch: Litmus 2.0 GA and Chaos Control Plane launch (4:20)

The technical deep dive of LitmusChaos 3.7.0 and its roadmap reveals a multi-faceted approach to enhancing security, developer experience, and core functionality.

Security Audit Enhancements

A recent security audit conducted by the 7A Security team informed a series of critical improvements. For a tool like LitmusChaos, which injects faults into user clusters, robust security is non-negotiable.

  1. API Security Enhancements:
  • Generic Error Messages: To prevent information leakage to potential attackers, APIs now return generic error messages instead of detailed internal errors.
  • RBAC Enforcements: Stricter Role-Based Access Control (RBAC) has been implemented for both GraphQL and REST APIs, ensuring that users only have access to resources and operations permitted by their assigned roles.
  • CORS Validation: Cross-Origin Resource Sharing (CORS) validation was added for GraphQL and the authentication server to mitigate cross-origin attacks.
  1. Authentication and Authorization Improvements:
  • Mandated Password Reset: New users are now required to reset their default password upon initial login, enforcing the creation of strong, unique credentials.
  • Strict Username/Password Validations: Strong validation rules for usernames and passwords prevent easily guessable credentials.
  • JWT Secret Creation: Upon K-Center (Litmus Chaos Control Plane) installation, a JWT (JSON Web Token) secret is now created and securely stored in the database. This secret is used for signing JWT tokens, enhancing their integrity and security.
  • Executor Role: A new Executor role was introduced, allowing users to execute experiments within a project, while the older "editor" role was deprecated to streamline permissions. This complements the existing Project Owner and Viewer roles.
  1. Network and Infrastructure Security:
  • HTTPS by Default: Environment-based support for HTTPS connections ensures secure communication over the network, moving away from default HTTP.
  • Network Policy YAMLs: Addition of Network Policy YAMLs secures inter-service communication within the LitmusChaos control plane itself.
  • Reduced Kubernetes Client-Go Dependency: Kubernetes client-go dependencies were removed from the GraphQL server, making the server lighter and minimizing the risk of vulnerabilities from third-party packages.
  • Go Version Upgrade: The Go version across K-Center components was upgraded from 1.20 to 1.22, addressing potential vulnerabilities in older Go runtime versions.
  1. Secure Development Practices:
  • GitLeaks Integration: GitLeaks is integrated into PR (Pull Request) checks to prevent sensitive information like tokens or secrets from being accidentally pushed into the repository.
  • GraphQL Introspection Control: Environment-based support to enable/disable GraphQL introspection was added. It is recommended to disable introspection in production environments to avoid exposing API schemas to the public.

SDK Support

To facilitate deeper integration into CI/CD pipelines and broader programmatic control, LitmusChaos is developing SDKs (Software Development Kits).

  • A Go SDK is currently under development by an LFX mentee, aiming to provide native Go language bindings for interacting with LitmusChaos.
  • A Java SDK has already begun development, with initial authentication APIs (user credential operations, project operations, environment creation) already implemented. This allows Java applications to directly embed and orchestrate chaos experiments.

Distributed Tracing

A significant upcoming feature is distributed tracing, designed to provide granular visibility into the execution flow and impact of chaos experiments.

  • Purpose: It addresses the difficulty users face in visualizing the intricate sequence of events and their impact during chaos injection. This aids in understanding the "blast radius" and debugging application behavior.
  • Implementation: The system will introduce a new environment variable, OTLP_EXPORTER_OTLP_ENDPOINT, allowing users to specify an OpenTelemetry Protocol (OTLP) endpoint for exporting trace data.
  • Collector & Visualization: While the demo uses a Simplest Collector and visualizes traces in Jaeger UI, the design is flexible, supporting any OTLP-compatible collector and visualization tool.
  • Trace Flow: The tracing captures the entire lifecycle: the Chaos Operator initiating execution, spawning Chaos Runner pods, which then create Chaos Experiment job pods responsible for the actual chaos injection. Finally, probes are run for hypothesis validation, often at the End of Test (EOT). This detailed timeline helps pinpoint exactly when and where issues arise.

Support for DocumentDB

LitmusChaos is expanding its database support beyond MongoDB to include managed NoSQL databases.

  • AWS DocumentDB is the first target.
  • Challenge & Solution: The integration faced hurdles due to AWS DocumentDB's lack of support for certain MongoDB operations like facet and bucket. Community members contributed PRs to modify LitmusChaos queries, removing these dependencies and enabling compatibility.

AWS RDS Fault Addition

Recognizing the widespread use of AWS RDS instances, a community member proposed and is currently developing a new chaos fault specifically for AWS RDS. This will allow users to test the resilience of their database instances directly.

Experiment Deletion/Abort

Addressing a common pain point reported by users of Litmus 3.x, a new feature enables the deletion or abortion of stuck experiments.

  • Problem: Configuration issues could cause experiments to get stuck in a "queued" state, preventing users from rerunning or creating new experiments.
  • Solution: This feature unblocks users, allowing them to either restart the experiment or create a new one, improving operational flexibility.

Roadmap Highlights

The roadmap for LitmusChaos includes several ambitious items:

  • Native Chaos Workflows: Moving away from Argo CD-based workflows to gain complete control over the experiment lifecycle, aiming for faster execution and greater flexibility.
  • Terraform Support: Providing automation for Chaos infrastructure and experiment operations, simplifying onboarding for new users and enabling SREs (Site Reliability Engineers) and developers to automate resilience testing.
  • Kubernetes Connectors: To streamline experimentation, this aims to allow a single Chaos infrastructure installation to target applications across multiple clusters, removing the need for repeated installations.
  • AI Integration (Chaos GPT): Exploring the use of AI, starting with Chaos GPT, to scan clusters and suggest relevant chaos experiments. This could significantly assist SREs in planning game days and optimizing chaos experimentation.

Demo / Proof of Concept

▶ Watch: Latest features: resiliency probes and score-driven UX (6:10)

The core demonstration during the talk showcased the upcoming distributed tracing capabilities of LitmusChaos, offering a live preview of how users will gain deeper insights into their chaos experiments.

The demo began within the K-Center UI, where an experiment targeting an engine-x application with a pod-delete fault was configured. The crucial new addition to the experiment configuration was an environment variable: OTLP_EXPORTER_OTLP_ENDPOINT. For this demonstration, the endpoint was set to a Simplest Collector, which would aggregate the trace data. The visualization tool used was Jaeger UI.

Once the experiment was initiated, the speakers showed the terminal, confirming the deployment of the engine-x pods and the simplest-collector in the Kubernetes cluster. As the chaos experiment progressed, the UI displayed the standard steps: installation of the chaos fault, followed by the chaos injection.

The highlight of the demo was the transition to the Jaeger UI. Here, by selecting chaos-operator as the service, a detailed timeline of events related to the chaos experiment was displayed. The trace clearly illustrated the sequence:

  1. The Chaos Operator initiated the execution.
  2. It then spawned Chaos Runner pods.
  3. These runner pods, in turn, created Chaos Experiment job pods, which were responsible for injecting the pod-delete fault into the engine-x application.
  4. Finally, after the chaos injection, probes were executed to validate the hypothesis, configured to run at the End of Test (EOT).

This visualization provided a clear, step-by-step breakdown of the entire chaos experiment lifecycle, including the start and completion times for each component. The speakers emphasized how this granular tracing capability would be invaluable for debugging purposes, allowing users to quickly identify precisely when and where an application misbehaved during a chaos injection, thus accelerating root cause analysis. While the demo utilized specific tools like Simplest Collector and Jaeger, the underlying OpenTelemetry Protocol (OTLP) ensures compatibility with a broad ecosystem of tracing tools.

Defensive Implications

▶ Watch: Litmus Chaos adoption, statistics, and future goals (7:20)

The advancements in LitmusChaos, particularly in its latest 3.7.0 release and future roadmap, offer significant defensive implications for organizations striving to build resilient cloud-native applications. Defenders should leverage these developments to proactively strengthen their systems against failures.

Firstly, the robust security enhancements resulting from the 7A Security audit are paramount. Defenders should ensure they are running the latest version of LitmusChaos (3.7.0 or newer) to benefit from stricter RBAC, CORS validation, HTTPS by default, and improved authentication/authorization mechanisms. This minimizes the attack surface of the chaos engineering tool itself, ensuring that the act of testing resilience doesn't introduce new vulnerabilities. Adhering to secure development practices, like those implemented in LitmusChaos (e.g., GitLeaks), also provides a blueprint for internal security standards.

Secondly, the upcoming distributed tracing feature is a game-changer for incident response and debugging. By integrating LitmusChaos with existing observability stacks (e.g., Jaeger, OpenTelemetry-compatible systems), SREs and developers will gain unprecedented visibility into the precise impact and sequence of events during a chaos experiment. This enables faster identification of root causes when an application fails to withstand a fault, allowing for quicker remediation and more accurate post-mortems. Defenders should plan to integrate this tracing capability into their monitoring infrastructure as it becomes available.

Thirdly, the development of SDKs for Java and Go facilitates the embedding of chaos experiments directly into CI/CD pipelines. This allows for continuous resilience validation, shifting chaos engineering left in the development lifecycle. Defenders can enforce that critical services undergo automated chaos tests before deployment, catching resilience issues earlier and preventing them from reaching production. The proposed Terraform support further enables infrastructure-as-code approaches for managing chaos experiments, promoting consistency and repeatability.

Finally, the continuous expansion of the Chaos Hub with new faults, such as the proposed AWS RDS fault, allows defenders to test a broader spectrum of infrastructure and application components. By adopting LitmusChaos and actively engaging with its community, organizations can stay ahead of potential failure modes, proactively identify weaknesses, and build truly resilient cloud-native systems. The long-term vision of AI integration (Chaos GPT) could eventually assist defenders in intelligently identifying and prioritizing critical chaos experiments, optimizing their resilience testing efforts.

Key Takeaways

  • Maturing CNCF Project with Strong Adoption: LitmusChaos is a CNCF incubating project with over 2 million installations, 68 million Docker pulls, and a 300% usage increase in the last year, demonstrating significant community trust and growth.
  • Enhanced Security Post-Audit: A comprehensive security audit by 7A Security has led to robust enhancements across API security (RBAC, CORS), authentication (mandated password reset, JWT secret management), network security (HTTPS, Network Policies), and secure development practices (GitLeaks, GraphQL introspection control).
  • Distributed Tracing for Deep Visibility: A key upcoming feature, distributed tracing, will provide granular visibility into chaos experiment execution, from Chaos Operator initiation to probe validation, significantly aiding debugging and understanding the impact of faults.
  • SDKs for Seamless CI/CD Integration: The development of Java and Go SDKs aims to enable programmatic control and direct integration of chaos experiments into CI/CD pipelines, making resilience testing a continuous and automated process.
  • Community-Driven Feature Expansion: Active community engagement, including mentorship programs (LFX, Google Summer of Code) and direct contributions, is driving new features like AWS DocumentDB support and the upcoming AWS RDS fault.
  • Ambitious Future Roadmap: Future plans include native chaos workflows for faster execution, Terraform support for automation, Kubernetes connectors for multi-cluster targeting, and AI integration (Chaos GPT) to intelligently suggest chaos experiments.

About the Speaker(s)

Sarthak Jain is a Senior Software Engineer at Harness and a dedicated maintainer of the LitmusChaos project. His work focuses on driving the development and evolution of LitmusChaos, contributing to its mission of making chaos engineering secure and accessible for cloud-native applications.

Saranya Jena is also a Senior Software Engineer at Harness and an active maintainer of LitmusChaos. She plays a crucial role in enhancing the project's features, improving its usability, and fostering community engagement within the chaos engineering ecosystem.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk provided a substantive and detailed update on LitmusChaos, a CNCF incubating project. The speakers, both project maintainers, delivered a clear overview of recent security enhancements stemming from a rigorous audit, the development of SDKs for better integration, and a deep dive into the upcoming distributed tracing feature. While it's an update on an existing tool, the focus on specific technical improvements and a forward-looking roadmap for resilience testing makes it valuable for anyone operating cloud-native applications.

Heather Calloway (CISO) — STRONG ACCEPT

This KubeCon update on LitmusChaos demonstrates significant maturity and a clear understanding of enterprise needs, particularly around security and operational visibility. The robust security audit, coupled with advancements in distributed tracing and SDK integration, provides critical capabilities for building and verifying application resilience. While a technical deep dive, the speakers effectively translate these developments into tangible benefits for risk management and incident response, making it a valuable update for any organization committed to cloud-native resilience.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025