Deep Dive To AI Agent Observability - Guangya Liu, IBM & Karthik Kalyanaraman, Langtrace AI

Guangya Liu, IBM, Karthik Kalyanaraman, Langtrace AI

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk, presented by Karthik Kalyanaraman from Langtrace AI (with contributions from Guangya Liu of IBM), provides a comprehensive exploration of AI agent observability. It addresses the significant challenges arising from the rapid evolution of AI architectures, particularly with the emergence of large language models (LLMs), vector databases, and sophisticated orchestration frameworks. The core premise is that traditional software observability paradigms are insufficient for the non-deterministic, complex, and rapidly changing landscape of generative AI applications and multi-agent systems.

Watch on YouTube

Visual summary for Deep Dive To AI Agent Observability - Guangya Liu, IBM & Karthik Kalyanaraman, Langtrace AI by Guangya Liu, IBM, Karthik Kalyanaraman, Langtrace AI
Visual summary for Deep Dive To AI Agent Observability - Guangya Liu, IBM & Karthik Kalyanaraman, Langtrace AI by Guangya Liu, IBM, Karthik Kalyanaraman, Langtrace AI

Key moments

  1. 0:00 Introduction to AI Agent Observability
  2. 2:00 Evolution of the AI Stack (LLMs, RAG, Agents)
  3. 4:50 Defining an AI Agent and its components
  4. 6:00 Defining Multi-Agent Systems and their coordination
  5. 6:40 Three broad challenges: Reliability, Latency, Cost
  6. 7:00 Reliability challenge: LLM non-determinism, subjective evaluation
  7. 8:00 Challenges: Evolving architecture, multimodality, agentic loops

Deep Dive To AI Agent Observability

Speakers: Guangya Liu, IBM; Karthik Kalyanaraman, Langtrace AI

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=1B9WZ6H0cn4

Overview

This talk, presented by Karthik Kalyanaraman from Langtrace AI (with contributions from Guangya Liu of IBM), provides a comprehensive exploration of AI agent observability. It addresses the significant challenges arising from the rapid evolution of AI architectures, particularly with the emergence of large language models (LLMs), vector databases, and sophisticated orchestration frameworks. The core premise is that traditional software observability paradigms are insufficient for the non-deterministic, complex, and rapidly changing landscape of generative AI applications and multi-agent systems.

The presentation highlights how the AI stack has quickly moved from basic LLM interactions to complex retrieval-augmented generation (RAG) and now multi-agent architectures, each introducing new layers of complexity, reliability concerns, latency issues, and cost considerations. Kalyanaraman argues that observability for these systems must extend beyond simple debugging to become a critical source of data for evaluation, fine-tuning, user experience optimization, and robust security and compliance.

The talk emphasizes the pivotal role of OpenTelemetry GenAI Semantic Conventions in standardizing the collection and interpretation of observability signals for AI agents. By detailing the unique signals—such as prompts, completions, vector scores, and inter-agent communication—and the evolving tooling, the speaker provides a roadmap for developers to effectively monitor, troubleshoot, and continuously improve their AI-powered applications, bridging the gap between innovative demos and production-ready systems.

Background

▶ Watch: Introduction to AI Agent Observability (0:00)

The landscape of artificial intelligence has undergone a dramatic transformation in recent years, setting the stage for the acute need for specialized observability solutions. The speaker chronologically outlines this evolution:

  • 2023: The Year of LLMs. Following the launch of ChatGPT in late 2022, 2023 saw an explosion of interest and development around LLMs. Developers were eager to build with these powerful models, leading to rapid advancements from GPT-3.5 to GPT-4 and GPT-4o, with intelligence improving significantly.
  • 2024: Retrieval Augmented Generation (RAG) and Orchestration Frameworks. The focus shifted to building enterprise-grade applications capable of leveraging proprietary data. This led to the rise of RAG architectures, where LLMs are augmented with external knowledge retrieval. Vector databases became crucial for semantic search and providing context to LLMs. Concurrently, AI orchestration frameworks like Langchain, Llama Index, Haystack, Autogen, and Crew AI emerged to manage the complex interactions between LLMs, data sources, and tools.
  • 2025 (and beyond): Multi-Agent Systems. Just a few months into 2025, the industry is already discussing and building multi-agent systems. This involves multiple AI agents, each with distinct responsibilities, working collaboratively or delegating tasks to achieve a common goal. This convergence brings together various components:
  • Observability: Becoming increasingly critical.
  • Hosting: Infrastructure for deploying models.
  • Model Context Protocol: A standard for LLMs to interact with external tools.
  • New Frameworks: Continually appearing, such as Vercel AI SDK (TypeScript), Crew AI, Autogen, Langchain, Llama Index, and Agno (Python).
  • LLM Proxies: Solutions like LightLLM and Arch AI Gateway allow developers to switch between models seamlessly, optimizing for cost or performance.
  • Memory Layers: Essential for LLMs to maintain contextual awareness over extended interactions.
  • Sandboxes: Secure execution environments for LLM-generated code, with projects like E2B emerging to address security concerns from tools like Lovable, Bolt, and v0.
  • Database Convergence: Traditional databases like MongoDB and PostgreSQL (with pg_vector) are integrating vector search capabilities.

Within this rapidly evolving stack, an AI agent is defined as an LLM-based entity that performs tasks by retrieving information from databases, interacting with external tools via tool calling, and engaging in multi-step reasoning, often with or without human intervention. Examples include code editors like Cursor, GitHub Copilot, and Vercel's v0. A multi-agent system extends this by coordinating multiple such agents, each potentially having different roles, to achieve more complex objectives.

However, this sophistication introduces significant challenges, primarily categorized into three areas:

  1. Reliability: This is arguably the biggest concern, stemming from the non-deterministic nature of LLMs. Unlike traditional software, LLMs do not guarantee the same output for the same input every time. This makes it difficult to build systems that consistently achieve desired goals, especially in multi-agent architectures. Evaluating performance becomes subjective, dependent on specific use cases (e.g., how "well" a summarization agent performs). The rise of multimodal data (text, voice, images, video) further complicates evaluation. The evolving, complex architecture and issues like agentic loops (where LLMs get stuck in repetitive actions) also contribute to reliability concerns.
  2. Latency: Primarily driven by token generation. Key metrics are time to first token (for streaming) and tokens per second.
  3. Cost: Despite decreasing costs, it remains a significant concern for developers, particularly as model usage scales.

Developers are increasingly realizing that while building "magical demos" was relatively easy, moving AI applications to production requires robust practices for tracing, evaluating accuracy, baselining performance, and continuous improvement. This necessitates a new approach to observability.

Key Findings

▶ Watch: Defining an AI Agent and its components (4:50)

The talk identifies several critical insights regarding the unique requirements and extended role of observability in the context of AI agents and multi-agent systems:

  1. Observability's Evolving Role: For AI agents, observability transcends traditional debugging. Traces become a rich data source for a multitude of other crucial activities:
  • Eval Set Building: Production traces contain valuable inputs, outputs, retrievals, and model attributes, which are essential for creating robust evaluation datasets. These datasets can then be used to test new models (e.g., switching from OpenAI to Anthropic) before deployment.
  • Fine-tuning and Data Labeling: Trace data, being representative of real-world production usage, can be extracted, labeled, and used for fine-tuning open-source models or enriching existing ones.
  • UX Optimization: Traces allow developers to analyze the behavior of different prompts in production, facilitating prompt engineering, A/B testing, and regression testing. User-specific memory and context can be derived from traces by attaching user IDs to spans, feeding back into the LLM's contextual understanding.
  • Security and Compliance: Traces become vital for detecting threats like prompt injections and ensuring the auditability of sensitive prompts and completions for regulatory compliance.
  1. Unique Observability Signals for AI Agents: The nature of AI agents demands new types of signals compared to traditional software:
  • Determinism: Traditional software is largely deterministic (success/failure), while AI agents are non-deterministic. Observability must account for varied responses to identical inputs.
  • Signals: Beyond traditional logs, traces, and metrics, AI agent signals include prompts and completions (inputs/outputs of models) and vector scores from database retrievals. Developers need to see these to diagnose performance issues, identify bottlenecks (e.g., in retrieval pipelines), and decide on model tuning or swapping.
  • State: The system's state is not just explicit; it's also fluid, influenced by the LLM's context window and memory layer, especially with tool calling, leading to a rapidly expanding state space.
  • Execution: Execution is LLM-driven rather than static.
  • Testing: Traditional unit tests are augmented by evaluations (evals) tailored for AI system performance.
  • Root Cause Analysis: Moves from stack traces to understanding semantics and context passed to the LLM.
  • Feedback: Unlike optional feedback in traditional systems, end-user feedback and developer evaluations are becoming mandatory for continuous improvement and baselining agentic application performance.
  1. Three Broad Categories of Observability for AI Agents: To achieve comprehensive visibility, tracing must encompass:
  • Agent/LLM Level Tracing: Capturing prompts, completions, tool calls, and model settings (e.g., temperature, API settings).
  • Storage and Memory Tracing: Monitoring queries to vector databases or memory layers, retrieved results, their scores, latency, and failures.
  • Framework Level Tracing: Observing the control flow of orchestration frameworks (Langchain, Agno, etc.) and messages passed between different agents. This is crucial for understanding the high-level behavior of multi-agent systems and identifying which agents need tuning.
  1. Unified Tracing: The ultimate goal is a single unified trace per user request, painting a holistic picture of the entire AI agentic system's operation and performance.
  1. OpenTelemetry GenAI Semantic Conventions: This initiative is a critical enabler for standardized observability. It provides a common language and structure for tracing AI agent components, ensuring interoperability and consistent experience across different vendors and tools. This standardization is vital for the emerging and rapidly evolving GenAI ecosystem.

Technical Deep Dive

▶ Watch: Defining Multi-Agent Systems and their coordination (6:00)

The technical core of the talk revolves around the standardization efforts of the OpenTelemetry GenAI Semantic Conventions Special Interest Group (SIG) and the practical implementation of observability for AI agents.

OpenTelemetry GenAI Semantic Conventions

The OpenTelemetry GenAI SIG has made significant progress in defining standards for tracing generative AI applications over the past 14 months. The group's primary responsibility is to establish semantic conventions for traces and spans, which includes:

  • Defining standard attributes that should be recorded.
  • Determining what information should be recorded as events.
  • Establishing consistent naming conventions for these attributes and events.

The goal is to ensure a seamless and standard experience for developers, regardless of the observability vendor or client they choose. This standardization is crucial given the proliferation of different LLM providers, frameworks, and tools. The speaker notes that multiple open-source projects, besides official OpenTelemetry libraries, are already adhering to these conventions, offering developers various options for adoption. Early proposals for vector database semantic conventions and agentic observability are currently under development.

The SIG actively encourages framework developers and LLM vendors to adopt these standards. A notable success is OpenAI's recent AI agent SDK, which ships with built-in OpenTelemetry tracing using these semantic conventions. Developers can find more information and participate in weekly calls via a QR code provided in the presentation.

Semantic Convention Categories

The conventions broadly categorize observability data into three groups:

  1. Events: Currently, prompts and completions are captured as events. However, there's an ongoing active discussion within the SIG about the optimal location for capturing these. Considerations include privacy, the potential size of spans (especially with large prompts/completions), and performance implications for the backend database storing these spans.
  2. Metrics: Key metrics for GenAI applications include:
  • Tokens: Input tokens, output tokens, and cached tokens are tracked for cost analysis.
  • Performance: Metrics like time to first token (for streaming responses) and tokens per second are crucial for evaluating inference speed.
  1. Spans: All request and response attributes related to LLM calls and other operations are recorded within spans.

The SIG involves representatives from major companies like Microsoft, Google, Traceloop, and Arize, highlighting the industry-wide commitment to these standards.

Hot Topics in the Semantic Conventions Group

Several complex issues are actively being debated and refined:

  • Tracing Multimodal Inputs and Outputs: As AI moves beyond text to include video, images, and audio, tracing these large data types presents challenges. The size of spans becomes a concern. Proposals include storing multimodal data in blob storage and attaching references (e.g., URLs) to the spans. This, in turn, raises new challenges for search and indexing of this linked data. The broader question of where to store attributes (in span attributes, events, or log events) is also under discussion, weighing various pros and cons for developer experience.
  • Tracing End-User Feedback and Evaluations: This is a critical area because evaluations are essential for baselining agentic application performance. The prompt and completion data is already in spans, but how should evaluation scores and user feedback (e.g., thumbs up/down) be modeled? The general consensus is against modifying traced data or attaching evaluation scores directly back to spans. The discussion focuses on how to model the backend database for these, distinguishing between end-user feedback and developer-done evaluations.

Setting Up Observability for AI Agents

For developers building agentic applications, setting up observability involves two main components:

  1. Observability Client (Vendor Component): These are the platforms where traces are visualized and analyzed. Open-source and self-hostable options mentioned include Langtrace, OpenLmetry, Uptrace, SigNoz, HyperDX, Traceloop, Vellum, and Arize.
  2. OpenTelemetry SDK (Client Component): These libraries instrument the application.
  • Official OpenTelemetry SDKs: Provide support for popular LLM providers like OpenAI, Vertex AI, and AWS Bedrock. The OpenAI client SDK, in particular, works with many other model providers.
  • Third-Party Open-Source Projects: Projects like Langtrace, OpenLmetry, and OpenLit offer extensive support for vector databases and agentic frameworks. These projects often serve as testing grounds for new semantic conventions before they are formally adopted by the SIG.

A significant advantage is that most of these SDKs adhere to the GenAI semantic conventions, ensuring that generated spans have a similar data format. The instrumentation process is designed to be non-intrusive, often requiring near-zero code changes. Developers simply install the library, import the OpenTelemetry library, and it patches the imported LLM or vector database client SDKs at runtime, automatically generating and exporting spans to the chosen observability vendor.

Typical Setup

A common setup involves:

  1. GenAI Applications or Multi-Agent Systems: These are the applications being observed.
  2. SDK Installation: The OpenTelemetry SDK is installed within these applications.
  3. Trace Generation: The SDK automatically generates traces that capture:
  • Inputs and outputs from LLMs.
  • Queries and retrievals from vector databases.
  • Control flow from orchestration frameworks.
  1. Batching and Export: These traces are batched and can either be sent through an OpenTelemetry Collector (optional) or directly to an OpenTelemetry-compatible observability client.
  2. Observability Client: Tools like IBM Instana, Elastic APM, Datadog, Grafana, or Prometheus receive and visualize the traces.

This flexible setup allows developers with existing observability infrastructure to easily layer in GenAI-specific observability.

Instrumentation Methods

Developers have three primary ways to instrument their agentic applications:

  1. Official OpenTelemetry Libraries: Available for Python, TypeScript, and other languages, with Python currently offering the most comprehensive GenAI coverage.
  2. Third-Party Open-Source Projects: Such as Langtrace, OpenLmetry, or OpenLit, which often provide broader support for vector databases and agentic frameworks, and contribute learnings back to the semantic conventions effort.
  3. Custom Instrumentation: Developers can write their own instrumentation, adhering to the established semantic conventions for consistency.

Demo / Proof of Concept

▶ Watch: Reliability challenge: LLM non-determinism, subjective evaluation (7:00)

While a live demonstration was not performed during the talk, the speaker presented sample traces to illustrate what comprehensive observability for AI agentic systems looks like in practice. These visual representations demonstrated the depth of insight gained from the OpenTelemetry-based instrumentation.

Two primary examples were shown:

  1. Multi-Agent System built with Agno: This trace highlighted the complexity of a coordinated multi-agent architecture. It showcased:
  • The orchestration layer managed by the Agno framework.
  • Interactions at the LLM layer, indicating individual model calls.
  • Crucially, the messages being passed between different agents, allowing developers to understand the communication flow and decision-making process within the system.
  • The response from a top-level agent, which aggregates and synthesizes outputs from subordinate agents to provide a final answer. This trace provides a holistic view of how multiple agents collaborate to achieve a goal.
  1. Simple Retrieval Augmented Generation (RAG) Implementation: This trace illustrated a more fundamental, yet common, AI architecture:
  • A framework at the top initiating the RAG process.
  • An embeddings generation step, where input text is converted into vector representations.
  • Interaction with a vector database (specifically mentioned as "VB8" in this case), showing queries and retrieval of relevant documents.
  • Finally, the interaction with an LLM (OpenAI in this example), where the retrieved context is combined with the user query to generate a response.

These demo traces effectively conveyed how a single, unified trace provides a detailed breakdown of each step, component, and interaction within an AI agentic system, making it possible to identify bottlenecks, understand behavior, and troubleshoot issues that would be opaque with traditional observability methods.

Defensive Implications

▶ Watch: Challenges: Evolving architecture, multimodality, agentic loops (8:00)

The insights from this talk provide critical guidance for defenders working with or building AI agentic systems. The non-deterministic nature and complex architectures of these systems introduce novel attack surfaces and operational challenges that demand a tailored defensive posture.

  1. Adopt OpenTelemetry GenAI Semantic Conventions: This is foundational. Standardized tracing allows security teams to consistently monitor and audit LLM interactions, tool calls, and data flows across diverse AI components and vendors. This consistency is vital for identifying anomalous behavior, potential prompt injections, or data exfiltration attempts.
  2. Implement Comprehensive Tracing: Defenders must ensure that observability extends across all three critical layers:
  • Agent/LLM Level: Monitor prompts, completions, and model settings for malicious inputs (e.g., prompt injections, jailbreaks) or unusual outputs that could indicate compromise or misuse.
  • Storage and Memory Level: Trace queries to vector databases and memory layers, along with retrieved results and scores. This helps detect attempts to manipulate context, exfiltrate sensitive data through retrieval, or identify if irrelevant/malicious data is being fed to the LLM.
  • Framework Level: Observe the control flow of orchestration frameworks and inter-agent communication. This can reveal unauthorized agent actions, agentic loops exploited for resource exhaustion, or attempts to hijack agent delegation.
  1. Leverage Traces for Security Evaluation: Traces are a rich source for building security-focused evaluation sets. Defenders can use production trace data to test AI agents against known prompt injection techniques, data leakage scenarios, or unauthorized tool usage before new models or features are deployed.
  2. Monitor for Prompt Injection and Data Leakage: Actively analyze prompt and completion data within traces for patterns indicative of prompt injection attacks. Ensure sensitive information is not inadvertently included in prompts or completions, especially when interacting with external tools or memory layers. The auditability provided by traces is crucial for compliance.
  3. Utilize Sandboxed Execution Environments: For LLMs generating and executing code (e.g., with tools like Vercel's v0), sandboxed environments like E2B are non-negotiable. Tracing within these sandboxes can provide crucial insights into the security posture of LLM-generated code and prevent privilege escalation or system compromise.
  4. Address Agentic Loops and Undesired Behavior: Design and monitor for agentic loops, which could be exploited for denial-of-service or resource exhaustion. Implement mechanisms (e.g., human-in-the-loop interventions, token limits, explicit stop conditions) to detect and break these loops, and use traces to understand why they occur.
  5. Ensure Data Governance and Privacy: The prompts and completions recorded in traces can contain sensitive information. Defenders must ensure robust data governance, access controls, and retention policies are applied to observability data to maintain privacy and compliance. The ongoing discussions within the OpenTelemetry GenAI SIG regarding where to store attributes (e.g., in blob storage for multimodal data) directly impact security and privacy design.
  6. Facilitate Incident Response and Root Cause Analysis: In the event of an AI-related incident, detailed traces provide the semantic context necessary for rapid root cause analysis, identifying exactly which LLM call, tool interaction, or data retrieval led to the issue.

By proactively integrating these defensive strategies, organizations can build more resilient and secure AI agentic systems capable of operating reliably in production environments.

Key Takeaways

  • AI Agent Observability is Distinct: Traditional observability falls short for AI agents due to their non-deterministic nature, complex evolving architectures, and unique signals like prompts, completions, and vector scores.
  • Observability Extends Beyond Debugging: Traces for AI agents are invaluable for building evaluation datasets, fine-tuning models, optimizing user experience (e.g., prompt engineering), and enhancing security/compliance.
  • Standardization is Crucial: The OpenTelemetry GenAI Semantic Conventions are essential for creating a consistent, interoperable, and vendor-agnostic observability experience across the diverse GenAI ecosystem.
  • Comprehensive Tracing is Necessary: Effective observability requires tracing at three levels: LLM/agent interactions, storage/memory operations, and orchestration framework control flow, all unified into a single trace per user request.
  • Non-Intrusive Instrumentation Simplifies Adoption: OpenTelemetry SDKs offer near-zero code instrumentation, patching client libraries at runtime and integrating seamlessly with existing OpenTelemetry-compatible observability backends.
  • Challenges Remain in Multimodal Data and Feedback: Ongoing work within the SIG addresses complex issues like efficiently tracing multimodal inputs/outputs (e.g., video, audio), determining optimal storage for attributes, and integrating end-user feedback and developer evaluations.

About the Speaker(s)

Karthik Kalyanaraman is the co-founder and CTO of Langtrace AI, an open-source and OpenTelemetry-based observability client SDK for tracing and evaluating generative AI applications. He is an active member of the OpenTelemetry GenAI Special Interest Group (SIG), where he contributes to defining semantic conventions for GenAI-based applications, and was a main contributor to the official OpenAI client SDK for OpenTelemetry.

Guangya Liu is from IBM and is also a member of the OpenTelemetry GenAI Special Interest Group. He collaborated on the presentation material.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This talk provides a critical and timely deep dive into AI agent observability, outlining the paradigm shift required for monitoring the non-deterministic, complex, and rapidly evolving multi-agent systems. It meticulously details the OpenTelemetry GenAI Semantic Conventions as the foundational standard, explaining unique signals, tracing methodologies across LLM, storage, and framework layers, and the profound implications for reliability, cost, security, and continuous improvement. The speakers, directly involved in shaping these standards, offer invaluable, actionable insights for moving AI applications from demos to robust production environments.

Heather Calloway (CISO) — STRONG ACCEPT

This talk provides essential clarity on the emerging challenges of AI agent observability, moving beyond theoretical discussions to practical, standardized solutions. The focus on OpenTelemetry GenAI Semantic Conventions is a critical development for institutionalizing risk management and accountability in advanced AI systems. It offers actionable insights for security leaders and developers, bridging the gap between innovative AI capabilities and the robust production practices required for resilience, security, and compliance.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025