Autonomous Al Agents for Cloud Cost Analysis - Ilya Lyamkin, Spotify

Ilya Lyamkin, Spotify

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

In an insightful presentation at KubeCon EU, Ilya Lyamkin from Spotify delved into the transformative potential of autonomous AI agents for navigating the increasingly complex landscape of cloud cost management. The talk, titled "Autonomous AI Agents for Cloud Cost Analysis," addressed the pervasive challenges organizations face in understanding, controlling, and optimizing their cloud spending, particularly within intricate multi-cloud environments. Lyamkin, who specializes in cloud infrastructure, developer tooling, and has extensive experience with RAG systems and production AI agents, presented a sophisticated multi-agent architecture designed to overcome the limitations of traditional, manual cost analysis and earlier AI-assisted approaches.

Watch on YouTube

Visual summary for Autonomous Al Agents for Cloud Cost Analysis - Ilya Lyamkin, Spotify by Ilya Lyamkin, Spotify
Visual summary for Autonomous Al Agents for Cloud Cost Analysis - Ilya Lyamkin, Spotify by Ilya Lyamkin, Spotify

Key moments

  1. 0:40 Challenges in cloud cost management
  2. 1:50 Initial workflow agent and its limitations
  3. 2:45 Python agents for precise math calculations
  4. 3:15 Self-healing SQL expert agent for reliable queries
  5. 4:55 Task planning and agent delegation architecture
  6. 5:30 LLM as a judge for verification and replanning
  7. 6:30 Complete multi-agent architecture overview
  8. 7:40 Example: Python agent for reliable math

Autonomous AI Agents for Cloud Cost Analysis

Speakers: Ilya Lyamkin, Cloud Infrastructure and Developer Tooling, Spotify

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=sTbJ1-3_yc

Overview

In an insightful presentation at KubeCon EU, Ilya Lyamkin from Spotify delved into the transformative potential of autonomous AI agents for navigating the increasingly complex landscape of cloud cost management. The talk, titled "Autonomous AI Agents for Cloud Cost Analysis," addressed the pervasive challenges organizations face in understanding, controlling, and optimizing their cloud spending, particularly within intricate multi-cloud environments. Lyamkin, who specializes in cloud infrastructure, developer tooling, and has extensive experience with RAG systems and production AI agents, presented a sophisticated multi-agent architecture designed to overcome the limitations of traditional, manual cost analysis and earlier AI-assisted approaches.

The core problem tackled by Spotify's initiative is the significant technical expertise gap required for cost analysis, the resource-intensive nature of manual processes, and the reactive posture most organizations adopt towards cloud spending. By leveraging specialized AI agents capable of planning, executing, and self-correcting, Spotify aims to democratize access to cost insights, proactive optimization, and drive substantial efficiency gains. This article explores the architectural evolution, technical innovations, and promising results of Spotify's multi-agent system, offering crucial insights for any organization grappling with the intricacies of cloud financial operations (FinOps).

Background

▶ Watch: Challenges in cloud cost management (0:40)

The modern cloud landscape presents a formidable challenge for cost management. Organizations frequently operate within multi-cloud environments, each with its own intricate pricing models, discount structures, and billing complexities. This inherent complexity makes it exceedingly difficult to maintain a comprehensive and accurate view of spending. Furthermore, access to critical cost data is often restricted to a handful of specialists or expert teams, creating significant bottlenecks in an organization's ability to respond swiftly to rising expenditures. Knowledge about cost optimization strategies tends to be siloed within these expert groups, preventing broader dissemination and proactive engagement across engineering teams. Consequently, most organizations find themselves in a reactive mode, addressing high cloud bills only after they have arrived, rather than implementing proactive optimization measures.

Historically, cloud cost analysis has been a largely manual and cumbersome process. It demands a significant technical expertise gap, requiring deep SQL proficiency and an intimate understanding of complex cloud billing data models. This manual approach is inherently resource-intensive, diverting valuable engineering time away from core product development and innovation. The process is also plagued by inefficiencies stemming from repetitive queries, error-prone manual analysis, and the sheer volume of data involved. As Lyamkin humorously noted, even advanced LLMs like ChatGPT suggest the need for a "safety helmet" when dealing with SQL, underscoring the inherent difficulties.

Spotify's initial foray into AI-assisted cost analysis involved a workflow agent utilizing a single Large Language Model (LLM). This approach aimed to process user questions, detect their type, and route them to either an internal Spotify knowledge base (using RAG systems) or a SQL generation module for querying billing datasets. While this single-agent approach offered some initial benefits, it quickly revealed several critical limitations. Chief among these was SQL reliability; the system generated SQL directly without robust validation or repair mechanisms, leading to frequent failures. Moreover, the single LLM struggled with complex, multi-part questions that necessitated different types of expertise, highlighting the need for a more specialized and robust architecture. A key challenge identified was LLMs' inherent struggle with precise mathematical calculations, which are fundamental for accurate financial analysis. Even advanced models could introduce small but significant errors that would compound, rendering financial insights unreliable. These challenges underscored the necessity for a more sophisticated, multi-agent solution capable of addressing precision, reliability, and complex reasoning.

Key Findings

▶ Watch: Python agents for precise math calculations (2:45)

The central and most compelling finding presented by Spotify is the dramatic improvement in accuracy and reliability achieved by migrating from a single-agent workflow to a multi-agent architecture for cloud cost analysis. The multi-agent approach demonstrated a remarkable 58% accuracy in answering real user questions, a significant leap compared to the 22% accuracy of the initial single-agent workflow baseline. This represents an improvement of over 150%, clearly justifying the increased architectural complexity.

This substantial performance gain was attributed directly to the specialized expertise embedded within the multi-agent system and its sophisticated planning capabilities. By delegating specific tasks—such as precise mathematical computations, robust SQL generation and repair, and external knowledge retrieval—to specialized agents, the system effectively mitigates the inherent weaknesses of a single, general-purpose LLM. The architecture's ability to decompose complex problems, orchestrate task execution, and implement self-correction mechanisms proved pivotal in delivering more accurate, reliable, and comprehensive responses to diverse cloud cost inquiries. The evaluation, based on real user questions extracted from Slack channels and validated against expert reference answers, provides strong empirical evidence that a well-designed multi-agent system can effectively tackle the multifaceted challenges of cloud cost optimization.

Technical Deep Dive

▶ Watch: Task planning and agent delegation architecture (4:55)

Spotify's journey from a rudimentary workflow agent to a sophisticated multi-agent architecture represents a significant technical evolution, designed to address the inherent limitations of LLMs in precision and reliability. The architecture is built upon several core components, each addressing a specific challenge in cloud cost analysis.

One of the primary challenges identified was the LLM's struggle with precise mathematical calculations, which are non-negotiable for accurate financial analysis. To overcome this, Spotify implemented a React agent (Reasoning and Action) that generates Python code and executes it within a sandbox environment. This approach effectively delegates precise calculations to Python, a language inherently designed for numerical accuracy, thereby ensuring greater reliability and reducing the risk of LLM hallucinations in quantitative analysis. An example shared by Lyamkin demonstrated the agent generating Python code to compare cost differences, showcasing its ability to perform calculations like cost_difference = current_month_cost - last_month_cost.

Addressing the critical issue of SQL reliability, a specialized SQL Expert Agent was developed with self-healing capabilities. This agent is specifically optimized for BigQuery, where Spotify stores most of its billing data. The core innovation here is a robust self-healing mechanism that automatically detects failed queries and attempts to repair them. The agent can make up to three attempts to fix a query before failing back to alternative approaches, significantly improving the success rate of data retrieval. The process involves a main query generation node sending a query to an execution node. If errors are encountered, feedback is sent back to the query generation node for repair. Once a working solution is achieved, the agent handles result formatting, returning clean, structured data to the main agent. This iterative repair loop is crucial for navigating the complex and often idiosyncratic nature of cloud billing data schemas.

For questions requiring external or up-to-date knowledge beyond internal datasets, the system incorporates web search capabilities also utilizing a React pattern. This combines reasoning (deciding what to search for) and action (executing the search) for dynamic information retrieval. The system employs iterative refinement, improving search queries based on initial results to hone in on the most relevant information. It leverages the Google Search API in conjunction with Gemini models for grounding, ensuring the system has access to the most current and authoritative information available on the web. The search flow diagram illustrates this, where the React agent first "thinks" if a search is needed, sends a query for retrieval, and if the initial results are insufficient, performs additional searches until it deems the information complete before returning it.

For handling complex, multi-faceted questions, a sophisticated task planning architecture was introduced. This architecture has three key capabilities:

  1. Task Decomposition: It breaks down complex user questions into smaller, more manageable subtasks.
  2. Workflow Orchestration: It determines the optimal execution sequence for these subtasks, including identifying dependencies and allowing for parallel execution where possible.
  3. Agent Delegation: It maps each subtask to the most appropriate specialized agent based on its capabilities (e.g., Python expert for math, SQL expert for data retrieval, cost engineer for anomaly detection).

An example task plan in JSON format illustrated this: first, the system calls the SQL expert to retrieve last month's cost by project. This data then flows to the Python expert for analyzing cost trends over time, which subsequently calls a dedicated Cost Engineer agent to identify anomalies and their root causes. This structured approach allows for tackling highly complex inquiries that would overwhelm a single, undifferentiated LLM.

A critical component ensuring the quality and reliability of the overall system is verification and replanning. Spotify employs an LLM as a judge approach to verify the quality and completeness of the answers generated by the agents. After all initial tasks are executed, the results are sent to a replanner node. This node analyzes whether the aggregated results adequately answer the original user question. If information is incomplete, the system generates new tasks to fill the gaps. If inconsistencies or errors are identified, it issues new error correction tasks. This iterative feedback loop ensures that the final answer is comprehensive, accurate, and directly addresses the user's intent. Once the replanner node is satisfied, it formats the final answer, stripping away the internal "thinking process" and presenting a clear, concise response to the user.

Putting all these components together, the complete multi-agent architecture operates as follows: A user query is first processed by a planner node, which generates a list of tasks. These tasks are then executed by the appropriate specialized expert agents, which can include the Python expert, internal knowledge base (RAG), SQL expert for BigQuery, or Google search. After execution, the planner node evaluates the results. It then either delivers the final answer directly to the user or, if necessary, creates a new plan for additional execution, effectively looping back through the task planning and execution phases. This modular and iterative model allows for continuous improvement of individual components and the expansion of system capabilities over time.

For developers interested in building similar agent orchestration frameworks, Lyamkin specifically recommended exploring Langraph, a framework designed for building robust and stateful multi-agent applications. For evaluation, he suggested OpenEval, an open-source library that implements LLM-as-a-judge evaluation with detailed reasoning capture, which Spotify itself utilized.

Demo / Proof of Concept

▶ Watch: LLM as a judge for verification and replanning (5:30)

While the talk did not feature a live, interactive demonstration, Ilya Lyamkin presented concrete examples and reasoning traces to illustrate the system's capabilities and internal workings. These examples served as powerful proofs of concept, showcasing how the specialized agents collaborate to address complex cost analysis questions.

One example demonstrated the system's web search and reasoning capabilities. The trace showed the agent actively searching over Google Docs to find relevant information in response to a cost analysis question. The accompanying reasoning trace meticulously detailed how the agent broke down the problem, formulated search queries, executed the searches, and then synthesized the findings to generate a comprehensive response. This highlighted the system's ability to go beyond internal data and leverage external, up-to-date information, a crucial aspect for dynamic cloud environments.

Another compelling example focused on the Python agent performing precise mathematical calculations. The trace illustrated a scenario where the system first consulted a rate expert (likely another specialized agent or a module within the internal knowledge base) to retrieve up-to-date Reserved Instance (RA) information for standard storage. This critical, precise data was then fed to the Python expert. The Python expert proceeded to write and execute specific Python code within its sandbox environment to perform the necessary calculations. Lyamkin emphasized that this method consistently produces highly reliable responses and significantly reduces hallucinations, which are common pitfalls when LLMs attempt complex arithmetic directly. These illustrative examples underscore the architectural design's effectiveness in combining the reasoning power of LLMs with the precision of traditional programming.

Defensive Implications

▶ Watch: Example: Python agent for reliable math (7:40)

While the primary focus of Spotify's autonomous AI agents is cloud cost optimization rather than traditional cybersecurity, the principles and robust architecture presented carry significant defensive implications for organizations managing their cloud finances. In the context of FinOps, "defense" often means safeguarding against financial waste, ensuring data integrity, and preventing unexpected cost escalations.

Firstly, the system's emphasis on accuracy and reliability is a crucial defensive measure against erroneous financial reporting. By delegating precise calculations to sandboxed Python environments and implementing self-healing mechanisms for SQL queries, the agents significantly reduce the risk of human error or LLM-induced hallucinations in financial analysis. Inaccurate cost data can lead to poor business decisions, misallocated resources, and a lack of trust in financial insights. A system that proactively verifies and corrects its outputs effectively defends against these forms of financial misinformation.

Secondly, the multi-agent architecture with its task decomposition and expert delegation inherently improves auditability and transparency. Each specialized agent performs a distinct function, making it easier to trace the logic and data flow behind a cost analysis. If an anomaly is detected or a question arises about a specific cost report, the system's reasoning trace can provide a detailed breakdown of which agents were involved, what data they accessed, and what calculations they performed. This level of transparency is vital for compliance, internal audits, and building confidence in autonomous financial systems.

Furthermore, the proactive nature of these agents, moving beyond reactive bill shock, acts as a preventative defense. By continuously analyzing cost trends, identifying anomalies, and providing actionable insights, organizations can proactively address potential cost overruns before they materialize into significant financial liabilities. This includes detecting inefficient resource usage, identifying misconfigurations that lead to excessive spending, or flagging unoptimized cloud services. The ability to "identify anomalies and root causes" through specialized agents directly supports this proactive defense against financial inefficiencies.

Finally, the continuous evaluation infrastructure and the "LLM as a judge" approach for verifying answer quality serve as a critical feedback loop, defensively hardening the system over time. By constantly assessing performance against real user questions and expert answers, the system can autonomously improve, making it more resilient to new types of queries and evolving cloud cost complexities. This ensures that the defense against financial waste and inaccuracy is not static but continuously adapts and strengthens. In essence, Spotify's work demonstrates how sophisticated AI can be leveraged to build a robust, trustworthy, and proactive financial governance layer within complex cloud environments, defending against the silent threats of inefficiency and inaccuracy.

Key Takeaways

  • Multi-agent architectures significantly outperform single-agent approaches for complex cloud cost analysis, achieving a 150% improvement in accuracy (58% vs. 22% baseline) by leveraging specialized expertise and sophisticated planning.
  • Delegating precise mathematical calculations to traditional code (Python) executed in a sandbox is crucial for ensuring reliability and reducing LLM hallucinations in financial analysis, where accuracy is paramount.
  • Self-healing mechanisms for SQL generation, specifically optimized for data warehouses like BigQuery, are essential for robust and reliable interaction with complex cloud billing data models, mitigating common LLM challenges in query construction.
  • Sophisticated task planning (decomposition, workflow orchestration, and agent delegation) is critical for breaking down complex user questions into manageable subtasks and efficiently coordinating specialized agents to provide comprehensive answers.
  • Verification and replanning loops, utilizing an LLM as a judge, are vital for ensuring the quality, completeness, and correctness of generated answers, enabling the system to identify gaps, correct errors, and iteratively refine its responses.
  • Continuous evaluation infrastructure, based on real user questions and expert reference answers, is fundamental for driving system improvement and making informed decisions about model selection and architectural enhancements.
  • Future scaling involves N-way search, developing a multitude of specialized agents (aiming for 100x more), and multi-model benchmarking to create an extensible architecture that autonomously improves as new LLMs are released.

About the Speaker(s)

Ilya Lyamkin is a key contributor at Spotify, where he focuses on cloud infrastructure and developer tooling. With over two years of dedicated experience in the field, Ilya has been actively involved in the development and implementation of Large Language Models (LLMs), with a particular emphasis on building robust RAG (Retrieval Augmented Generation) systems and deploying production AI agents. His work at Spotify centers on leveraging these advanced AI capabilities to solve complex organizational challenges, such as optimizing cloud spending and enhancing internal developer experiences.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This presentation from Spotify details a robust multi-agent architecture for cloud cost analysis, moving beyond the limitations of single LLM approaches. By integrating specialized agents for precise Python calculations, self-healing SQL generation, and sophisticated task orchestration, Spotify achieved a significant 150% improvement in accuracy. The talk provides actionable insights into building reliable AI agents, demonstrating how to mitigate common LLM weaknesses in financial contexts and offering a practical framework for proactive FinOps.

Heather Calloway (CISO) — STRONG ACCEPT

Spotify's presentation on autonomous AI agents for cloud cost analysis offers a highly relevant and actionable blueprint for addressing a critical area of business risk. By demonstrating a significant leap in accuracy and reliability through a multi-agent architecture, the talk provides a clear path for organizations to move beyond reactive cloud cost management. This approach directly tackles institutional accountability gaps and offers a sophisticated model for proactive financial governance in complex cloud environments, making it a valuable insight for any executive concerned with operational efficiency and fiscal responsibility.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025