Leveraging Internal Knowledge: Building AiKA at Spotify - Majd Salman & Jofre Mateu Matesanz
Majd Salman, Jofre Mateu Matesanz
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
This talk, presented by Majd Salman and Jofre Mateu Matesanz, platform engineers in Spotify's Platform Developer Experience (PDX) department, details the creation and impact of AiKA (Artificial Intelligence Knowledge Assistant). AiKA is Spotify's enterprise-wide AI-powered solution designed to combat critical developer productivity blockers related to information discovery and documentation quality. The speakers unveil how Spotify leveraged advanced AI techniques, particularly Retrieval Augmented Generation (RAG), to build a robust, scalable platform that not only answers developer questions but also fosters a positive feedback loop for improving internal documentation.

Key moments
- 0:00 Introduction and the core problem of finding information
- 1:45 The scale and fragmentation of Spotify's internal knowledge
- 3:50 Explaining Retrieval Augmented Generation (RAG) and its benefits
- 5:30 Why Spotify needed a single, robust RAG platform
- 7:00 Introducing AiKA, Spotify's AI Knowledge Assistant, and its principles
- 7:45 AiKA's diverse client integrations and knowledge blending
- 8:40 Practical example: AiKA clarifying Spotify's internal jargon
Leveraging Internal Knowledge: Building AiKA at Spotify
Speakers: Majd Salman, Platform Engineer, Spotify; Jofre Mateu Matesanz, Platform Engineer, Spotify
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=FEy2lhe6CM8
Overview
This talk, presented by Majd Salman and Jofre Mateu Matesanz, platform engineers in Spotify's Platform Developer Experience (PDX) department, details the creation and impact of AiKA (Artificial Intelligence Knowledge Assistant). AiKA is Spotify's enterprise-wide AI-powered solution designed to combat critical developer productivity blockers related to information discovery and documentation quality. The speakers unveil how Spotify leveraged advanced AI techniques, particularly Retrieval Augmented Generation (RAG), to build a robust, scalable platform that not only answers developer questions but also fosters a positive feedback loop for improving internal documentation.
The core problem AiKA addresses is the persistent challenge of finding accurate and up-to-date information within a large, rapidly evolving organization like Spotify. With a vast internal knowledge base scattered across numerous systems, engineers spend considerable time sifting through information, leading to reduced productivity and repeated support requests. AiKA's development represents a strategic initiative to centralize and intelligently surface this knowledge, thereby enhancing the developer experience and allowing engineers to focus on their core work rather than information retrieval.
The significance of AiKA extends beyond mere question-answering. It stands as a testament to successfully implementing AI at scale within an enterprise, demonstrating how a well-architected RAG system can overcome common limitations of large language models (LLMs) such as hallucinations and limited context windows. By providing a customizable and integrated knowledge platform, Spotify not only boosts individual engineer productivity but also cultivates a culture of continuous documentation improvement, showcasing a practical and impactful application of AI in the modern development landscape.
Background
▶ Watch: Introduction and the core problem of finding information (0:00)
Spotify, a company with a significant history of growth and evolution, faces common challenges associated with scale: information silos, outdated documentation, and repetitive support requests. The speakers highlighted data from Spotify's quarterly Engineering Satisfaction (ENSAT) survey, initiated in 2020, which consistently identified "issues with finding information" and "poor or missing documentation" as top three productivity blockers. This problem manifests across various internal systems: over 10,000 internal documentation sites in Tech Docs, more than 22,000 GitHub repositories containing READMEs and code examples, and extensive institutional troubleshooting knowledge buried in Slack conversations, alongside working documents, Requests For Comments (RFCs), Architectural Decision Records (ADRs), and planning tools. Engineers were spending valuable time searching for answers or providing repetitive support, diverting focus from their primary tasks.
The emergence of advanced Large Language Models (LLMs) like GPT-3.5 two years prior presented a potential solution, offering natural language interaction with information. However, early LLMs came with significant limitations:
- Limited context windows: Typically 4,000 to 8,000 tokens, making it difficult to process extensive enterprise knowledge.
- Hallucinations: The tendency to generate plausible but incorrect information, a critical concern for reliable internal data.
- High cost: Training or fine-tuning LLMs from scratch was prohibitively expensive and time-consuming for specific internal datasets.
To address these limitations, Retrieval Augmented Generation (RAG) emerged as a powerful pattern. RAG works by first retrieving relevant information from a knowledge base and then feeding that information into an LLM to generate an answer, effectively "grounding" the response in actual data. The typical RAG process involves:
- Document Ingestion: A set of documents is split into meaningful chunks.
- Embedding: Each chunk is converted into a vector representing its semantic meaning using an embedding model.
- Vector Storage: These vectors are stored in a vector database.
- Query Processing: When a user asks a question, it's also converted into a vector using the same embedding model.
- Semantic Search: A vector search (or semantic search) is performed across the vector database to find the most semantically relevant chunks.
- Generation: The retrieved chunks are provided as context to an LLM, which then generates an answer.
At Spotify, the potential of RAG was recognized early, leading to approximately eight different teams independently building their own RAG solutions. This fragmented approach resulted in duplicated efforts, a lack of shared best practices, and unshared infrastructure and knowledge. Recognizing the need for a unified approach, Spotify formed a dedicated team, tasked with creating a single, robust platform for knowledge ingestion and serving. This platform, named AiKA, was built on several core principles:
- Trust from Transparency: AiKA acts as an aggregator, not the source of truth, emphasizing the ability for users to verify information sources.
- Positive Feedback Loop: Better documentation leads to improved AiKA responses, which in turn incentivizes teams to maintain and update their documentation.
- Multiple Experiences & Adaptability: The platform needed to support various knowledge domains and be customizable.
- Meet Users Where They Are: AiKA needed to be accessible through different interfaces, including Spotify's internal developer portal Backstage, Slack, and integrated development environments (IDEs).
Key Findings
▶ Watch: Explaining Retrieval Augmented Generation (RAG) and its benefits (3:50)
The development and deployment of AiKA at Spotify have yielded several significant findings and demonstrated substantial positive impact:
- High Adoption and Sustained Usage: Within approximately a year, AiKA achieved widespread adoption. About 70% of Spotify employees have tried it at least once. More impressively, it boasts over 1,000 daily active users, with 25% of all employees using it weekly. For developers specifically, this number jumps to approximately 86% weekly usage, underscoring its utility within the engineering community.
- Positive Feedback Loop on Documentation Quality: AiKA has successfully fostered a virtuous cycle where improved responses incentivize better documentation. Users actively search for and identify missing or outdated information, notifying documentation owners, which leads to updates and further enhances AiKA's accuracy. This qualitative feedback loop is a key indicator of the platform's long-term value.
- Contextual Understanding of Internal Jargon: AiKA demonstrates the ability to correctly interpret Spotify-specific terminology and abbreviations. For instance, when asked about "MMA," it correctly identifies it as "Managed Monitoring and Alerting" within the Spotify context, rather than "mixed martial arts," showcasing effective domain-specific knowledge retrieval.
- Enhanced Team Discovery and Knowledge Retrieval: With over 600 teams, finding the right contact or team for a specific feature can be challenging. AiKA can retrieve the owner of a feature based on past Slack conversations, significantly streamlining internal communication and support.
- General LLM Capabilities: Despite its focus on internal knowledge, AiKA retains the general reasoning capabilities of an LLM, capable of answering broader technical questions, such as general Python queries, without issues. This versatility makes it a valuable tool for both Spotify-specific and general technical queries.
- AiKA Goalie Bot for Automated Support: A significant innovation is the AiKA Goalie Bot, a standardized, customizable Slack support solution. This bot automates answers to repetitive questions in support channels, freeing human "goalies" (support channel monitors) to focus on complex troubleshooting. It saves "thousands of hours" by automatically answering an average of 30% of questions in the channels it monitors.
- Configuration for Non-Technical Teams: The Goalie Bot's declarative YAML configuration approach allows non-technical teams to customize its behavior, including knowledge sources, system prompts, and confidence thresholds, without writing any code. This has led to its adoption in over 100 support channels, including those outside of the R&D department.
- Retrieval is Paramount, Reranking is Key: The speakers emphasized that retrieval is the most critical component of RAG. They found that reranking retrieved documents was the single best improvement made, leading to a 10-15% improvement in retrieval accuracy.
- Context Window Optimization: Counter-intuitively, more context isn't always better. Doubling the amount of documents in context did not yield significantly better results, as relevancy drops off logarithmically. This highlights the importance of precise retrieval over simply stuffing the context window.
- Data Source Specific Ingestion: Different data formats (e.g., graph structures, Slack conversations) require tailored ingestion strategies to be effectively embedded and retrieved, emphasizing that a one-size-fits-all approach is insufficient.
- Addressing Hallucination vs. Bad Data: A key challenge is distinguishing between an LLM hallucination and the correct use of outdated or incorrect source data. AiKA addresses this by always citing sources, empowering users to verify information and report discrepancies, reinforcing the feedback loop.
Technical Deep Dive
▶ Watch: Why Spotify needed a single, robust RAG platform (5:30)
AiKA's architecture is a sophisticated adaptation of the standard RAG pattern, fine-tuned for Spotify's unique needs. At its core, a main backend service orchestrates the entire process, interacting with multiple components to deliver intelligent responses.
The system supports several LLM providers and their frontier models, allowing flexibility and leveraging state-of-the-art language capabilities. It also relies on various third-party APIs for specialized functions, notably for reranking retrieved documents and generating embeddings during the retrieval phase. For specific internal requirements, Spotify has developed its own inference service hosting custom machine learning models. These include:
- A model to determine if an incoming question requires internal knowledge or can be answered by a general LLM.
- A custom confidence scoring model that outputs a score between 0 and 1, indicating how confident the system is that an answer is helpful based on the question, retrieved documents, and the generated answer.
To ensure continuous improvement and stability, AiKA incorporates a robust evaluation framework. This framework allows the team to quantitatively assess metrics like retrieval accuracy and answer quality, enabling safer and more confident changes to the platform. All system activities are meticulously tracked using an observability system based on OpenTelemetry, providing deep insights into service performance, user interactions, and areas for improvement.
The ingestion pipeline was designed for scalability and user autonomy. Starting with a curated set of high-quality documentation and Slack channels, the platform evolved to allow users to maintain their own data sources. This user-driven growth, however, presented challenges. As the knowledge base expanded across diverse domains (e.g., backend development, web security, organizational data), semantically similar but irrelevant answers could sometimes be retrieved.
To counter this, Spotify implemented knowledge filtering capabilities. The speakers illustrated this with a 3D visualization of their embedding space, where similar concepts cluster together. While vector search initially casts a "wide net" to ensure comprehensive retrieval, the ability to narrow down the knowledge base to specific contexts is crucial. For example, a backend engineer asking about deployment practices is likely interested in backend deployment, not iOS deployments. By filtering the knowledge available based on user context or specified domains, AiKA can provide more precise and relevant answers, "getting rid of the background noise."
This filtering capability was instrumental in evolving AiKA from a general assistant into a platform for focused experiences. Teams can now package a "slice" of the vast knowledge space with a custom system prompt to create specialized assistants tailored to their domain or internal support cases. This empowers teams to leverage AiKA's underlying infrastructure (chunking, embedding, ingestion) without needing to manage these complexities themselves.
A prime example of a focused experience is the AiKA Goalie Bot, a standardized support solution for Slack. This bot integrates seamlessly into support channels, maintaining the familiar AiKA identity. When an incoming question is detected, the bot first determines its capability to answer. If it can, it generates an answer and uses the custom confidence score to decide the next action:
- Low confidence: The bot stops and takes no action.
- Medium confidence: The bot proposes the generated answer to the human goalie for review and decision.
- High confidence: The bot automatically posts the answer along with citations to its sources, saving time for everyone involved.
The Goalie Bot's configuration is managed through a declarative YAML approach, allowing teams to fully customize its behavior. This includes defining which knowledge sources are available (e.g., specific documentation, past Slack threads, custom data), tweaking the system prompt to align with their domain's particularities, and setting specific confidence thresholds for low, medium, and high confidence actions. This no-code configuration has been critical for its adoption by over 100 support channels, including non-technical teams.
Key learnings and challenges during AiKA's development included:
- Retrieval is paramount: Most effort was concentrated on optimizing retrieval. Vector search is considered "coarse," and future plans include hybrid search for improved results.
- Reranking's impact: Reranking provided a 10-15% improvement in retrieval accuracy, making it the single most effective enhancement.
- Context window diminishing returns: Doubling context didn't yield significantly better results due to the logarithmic drop-off in relevancy, emphasizing the importance of quality over quantity of retrieved documents.
- Data source specificity: Each data source requires unique consideration for ingestion; for instance, graph-structured data needs flattening, and Slack conversations can be embedded in multiple ways depending on anticipated questions.
- Hallucination vs. bad data: Distinguishing between an LLM making things up and correctly using outdated information is tricky. Citing sources is crucial for user verification and incentivizing documentation updates.
Looking ahead, AiKA is evolving through clear phases. The current phase provides a solid foundation of semantic search. The next step is enhanced reasoning, where the system will automatically deduce relevant data sources based on the query and user context, eliminating the need for users to specify sources. Further in the future, AiKA aims to become a platform for agentic capabilities, handling multi-turn questions, gathering information from many sources, using tools to access real-time data, and acting on behalf of the user. Spotify is also bringing AiKA's capabilities to Spotify Portal, their managed solution for the Backstage platform, allowing external customers to leverage this powerful knowledge assistant.
Demo / Proof of Concept
▶ Watch: AiKA's diverse client integrations and knowledge blending (7:45)
The talk effectively showcased AiKA's capabilities through demonstrations of its various integrations and functionalities.
The primary demonstration highlighted AiKA's integration into Backstage, Spotify's internal developer portal. Users could be seen interacting with AiKA directly within this familiar environment, asking questions and receiving blended answers derived from key knowledge sources: internal documentation, past conversations from Slack support channels, and organizational data. This integration underscores the "meet users where they are" principle, making AiKA easily accessible within the developer's workflow.
Beyond Backstage, the speakers detailed other client interfaces where AiKA is available:
- Slack Bot: A dedicated Slack client for both private and group conversations, allowing engineers to query AiKA directly within their communication tool.
- Open API: An API available to all Spotify developers, enabling them to incorporate AiKA's knowledge capabilities into their own applications and services programmatically.
- Python Client Library: A library designed to simplify programmatic access to AiKA's knowledge, making it easy for developers to integrate it into scripts or other tools.
Several specific examples illustrated AiKA's intelligence:
- Spotify Lingo Interpretation: A classic challenge in large organizations is understanding internal jargon. When asked about "MMA," AiKA correctly identified it as "Managed Monitoring and Alerting" within Spotify's context, demonstrating its ability to retrieve contextually accurate information from internal documentation.
- Team Ownership Discovery: With hundreds of teams, finding the right contact for a specific feature can be difficult. AiKA was shown to retrieve the owner of a particular feature based on historical Slack conversations, highlighting its utility in navigating organizational structure and past discussions.
- General Technical Queries: To demonstrate its breadth, AiKA was also shown answering a general Python programming question, proving that it retains the general reasoning capabilities of a large language model in addition to its specialized internal knowledge.
The AiKA Goalie Bot also served as a powerful proof of concept for automating internal support. While not a live interactive demo, the detailed explanation of its decision-making process—determining capability, generating an answer, and then using the confidence score to either stop, propose to a human, or automatically post with citations—clearly illustrated its functionality and the significant time savings it provides. The declarative YAML configuration for customizing the bot further emphasized its adaptability and ease of deployment across diverse internal teams. These examples collectively painted a clear picture of AiKA's practical utility and its success in addressing Spotify's information discovery challenges.
Defensive Implications
▶ Watch: Practical example: AiKA clarifying Spotify's internal jargon (8:40)
While the talk primarily focuses on internal productivity and knowledge management rather than external security, the principles and architecture of AiKA carry significant defensive implications for internal platform teams, security teams, and the overall organizational security posture.
- Centralized Control and Best Practices for AI/RAG: The initial problem of multiple teams building independent RAG solutions highlights a common challenge in enterprises. A centralized platform like AiKA prevents fragmented efforts, ensuring that AI solutions adhere to consistent security, privacy, and quality standards. This is a critical defensive measure against "shadow AI" implementations that might inadvertently expose sensitive data or introduce vulnerabilities.
- Mitigating Information Silos for Security Awareness: By making vast internal knowledge more discoverable, AiKA can indirectly enhance security. Security policies, best practices, incident response procedures, and vulnerability remediation guidelines often reside in scattered documentation. AiKA makes this crucial information readily accessible to developers, potentially leading to earlier adoption of secure coding practices and faster response to security incidents.
- Data Quality and Trustworthiness: The emphasis on the "positive feedback loop" for documentation improvement and the challenge of distinguishing hallucinations from bad data are vital. For security-sensitive information, it's critical that the source material is accurate and up-to-date. AiKA's mechanism of citing sources and enabling user feedback is a defensive layer, allowing engineers to verify security advice and report outdated guidelines, thus preventing the spread of incorrect security practices.
- Confidence Scoring for Risk Management: The custom confidence scoring model is a powerful defensive tool. For security-related queries, automatically posting answers with high confidence is acceptable, but for medium confidence, requiring human review (e.g., a security expert "goalie") prevents the automated dissemination of potentially misleading or incomplete security advice. Low-confidence answers are rightly suppressed, minimizing the risk of incorrect information being acted upon.
- Secure Data Ingestion and Access Control: While not explicitly detailed, a robust RAG platform like AiKA must incorporate secure ingestion pipelines and fine-grained access control for different data sources. This ensures that sensitive information is only accessible to authorized users and that the embedding process itself doesn't inadvertently expose or leak data. The ability to filter knowledge based on context can also be extended to filter based on user roles or permissions, adding another layer of defense.
- Observability and Auditability: The use of OpenTelemetry for comprehensive observability is a strong defensive posture. It allows platform teams to monitor how AiKA is being used, what questions are being asked, and the quality of responses, including potential misuse or attempts to extract sensitive information. This auditability is crucial for compliance and incident investigation.
- Future Agentic Capabilities and Tooling: As AiKA evolves towards agentic capabilities that can interact with real-time information and act on behalf of the user, the defensive implications become even more pronounced. Implementing robust authorization, input validation, and sandboxing for any tools or systems AiKA interacts with will be paramount to prevent privilege escalation, unauthorized actions, or data exfiltration.
In summary, while AiKA's direct goal is productivity, its underlying principles of controlled knowledge ingestion, transparent sourcing, confidence-based response management, and robust observability contribute significantly to building a more secure and resilient internal information environment.
Key Takeaways
- RAG is a Powerful Enterprise Solution: Retrieval Augmented Generation (RAG) effectively addresses common LLM limitations (context window, hallucination, cost) for internal knowledge management, proving highly valuable for large organizations like Spotify.
- Centralized Platform is Crucial for Scale: Building a single, robust RAG platform avoids fragmented efforts, ensures best practices, and facilitates knowledge sharing and infrastructure reuse across an enterprise.
- Retrieval Quality is Paramount: The accuracy of retrieved documents is the most critical factor for reliable answers. Techniques like reranking can significantly boost retrieval accuracy (10-15% improvement observed).
- Documentation Quality Drives AI Success: A strong, positive feedback loop exists where better documentation leads to improved AI responses, which in turn incentivizes users and teams to maintain and update their documentation.
- Customization and Focused Experiences Boost Adoption: Providing customizable solutions, such as the AiKA Goalie Bot with its declarative YAML configuration for domain-specific knowledge and prompts, enables broad adoption across diverse technical and non-technical teams.
- AI Can Liberate Human Expertise: AI assistants can effectively automate repetitive support tasks (e.g., answering 30% of questions in support channels), saving thousands of hours and allowing human experts to focus on more complex and engaging work.
About the Speaker(s)
Majd Salman and Jofre Mateu Matesanz are platform engineers within Spotify's Platform Developer Experience (PDX) department. Both have accumulated around eight years of experience at Spotify, working across various internal platforms. Their focus lies in designing and building tools and systems specifically aimed at enhancing the productivity of Spotify's engineering teams. Their deep understanding of the challenges faced by Spotify engineers, gained through years of observing the company's growth and evolving internal platforms, directly informed the development of AiKA.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk details the successful implementation of AiKA, Spotify's enterprise-wide AI-powered knowledge assistant built on Retrieval Augmented Generation (RAG). It effectively showcases how a large organization can tackle developer productivity blockers related to information discovery by leveraging applied AI. The speakers provide valuable, hard-won insights into scaling RAG, optimizing retrieval, and fostering a positive feedback loop for documentation quality, making it a highly actionable and credible case study for anyone building similar internal platforms.
Heather Calloway (CISO) — STRONG ACCEPT
Spotify's AiKA is a powerful demonstration of leveraging Retrieval Augmented Generation (RAG) to solve systemic internal knowledge and productivity challenges. While focused on developers, this talk provides critical insights for security leaders on building secure, governed enterprise AI platforms. The decision to centralize RAG development, establish feedback loops for documentation quality, and implement confidence-based response management are exemplary for managing AI risk and ensuring the accuracy of crucial internal information, including security policies.