Many Cooks, One Platform: Balancing Ownership and Contribution for the Perfect Broth - Lian Li
Lian Li
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In this insightful KubeCon EU talk, Lian Li, a freelance platform engineer and co-chair of the CNCF TAG App Delivery, delves into the intricate challenges of building and evolving an internal developer platform within a large, established organization – specifically, the Dutch government's IT service provider for the judiciary, Evo Respra. Li candidly shares her experiences, highlighting the friction that arises when a platform team, focused on providing foundational "bricks," encounters development teams who simply desire a ready-made "house."

Key moments
- 0:00 Speaker introduction and personal background
- 2:00 Setting expectations: no 'secret sauce' for platform success
- 4:00 Evo Respra: judiciary IT, value teams, tech stack
- 6:00 OpenShift platform, migration, and cluster topology
- 8:00 Hinting at current platform operational challenges
Many Cooks, One Platform: Balancing Ownership and Contribution for the Perfect Broth
Speakers: Lian Li, Freelance Platform Engineer, Co-chair CNCF TAG App Delivery
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=gjEuIUCbNYY
Overview
In this insightful KubeCon EU talk, Lian Li, a freelance platform engineer and co-chair of the CNCF TAG App Delivery, delves into the intricate challenges of building and evolving an internal developer platform within a large, established organization – specifically, the Dutch government's IT service provider for the judiciary, Evo Respra. Li candidly shares her experiences, highlighting the friction that arises when a platform team, focused on providing foundational "bricks," encounters development teams who simply desire a ready-made "house."
The talk is a pragmatic exploration of the human and organizational dynamics inherent in platform engineering. Li emphasizes that there is no singular "secret sauce" for platform success, but rather a continuous journey of collaboration, community building, and clear communication of scope. This is not a talk about specific technical solutions as much as it is about the socio-technical challenges that often derail even the most well-intentioned platform initiatives, making it highly relevant for anyone involved in internal developer platforms, DevOps transformations, or large-scale organizational change.
Background
▶ Watch: Speaker introduction and personal background (0:00)
Evo Respra, the organization at the heart of Li's narrative, serves as the IT service provider for the Dutch judiciary. This environment is characterized by a strong emphasis on process-driven applications and a user base (judges, lawyers) often resistant to new methodologies, preferring established workflows. The internal structure comprises value teams (analogous to stream-aligned teams in Team Topologies), which are cross-functional units responsible for building applications, primarily using Java and .NET, with some JavaScript for front-end development. These value teams are grouped into "bases," with Li's Container Management Platform team residing within the infrastructure base, alongside an API gateway team, an Azure team, and the OST (Developer Tooling) team which maintains developer tools like Git and Azure DevOps.
The platform itself is undergoing a significant transition. The "old platform" runs on OpenShift 4.16, which has reached its end-of-life, necessitating a migration to a "new platform" based on OpenShift 4.17. This migration presents both challenges and opportunities: while the old platform requires continued maintenance, the new one is in its infancy, allowing for new architectural decisions. The old platform's topology included two development clusters (for the platform team only) and two production clusters that hosted all environments (test, dev, prod), leading to anticipated problems. A key "win" for the new platform is a more sensible separation: two development clusters host all development environments (dev, test), and two production clusters are dedicated solely to production workloads. Additionally, a new management cluster manages the four operational clusters, with plans for an engineering cluster for the platform team's internal development.
A significant source of friction highlighted by Li is the reliance on Azure DevOps for Git repositories and CI/CD pipelines. This tooling, owned and maintained by the separate OST team, often clashes with the platform team's preferred cloud-native approaches, creating procedural and philosophical disagreements.
Li then articulates the fundamental "Why Platforms?" question, distinguishing between traditional service-based architectures and platform-centric approaches. In a service-based model, developers act as clients, interacting with disparate services (e.g., a build service and a deploy service). While loosely coupled, this often leads to deviation between services over time, forcing developers to bridge the gaps. Conversely, a platform setup aims to optimize for the golden path. Here, the developer interacts with a single entry point (e.g., a deploy service within an internal developer platform), which internally orchestrates calls to other services like a build service. This approach ensures a consistent contract between internal components, simplifying the developer experience and ensuring integrated functionality. Li views platforms as an evolutionary step, moving from ad-hoc local scripts to central services, and finally integrating those services into a cohesive platform.
Crucially, Li also differentiates a developer portal from a platform. A portal is typically a user-friendly graphical interface that sits on top of an existing platform, surfacing the most common workflows and best practices (e.g., a "one-click database provisioner"). The platform, however, is the underlying capability that can reasonably fulfill any request. This distinction underscores that while a portal enhances usability, the platform provides the core functionality.
The overarching problem Li identifies is encapsulated in the metaphor of "bricks versus builds." The platform team often sees itself as providing the "Lego bricks" – the foundational components and tools – expecting developers to construct their own solutions. However, the developers, as users, often just want a complete "house" they can move into, leading to a mismatch in expectations and significant disappointment when the provided "bricks" require substantial effort to assemble into a usable application environment.
Key Findings
▶ Watch: Setting expectations: no 'secret sauce' for platform success (2:00)
Lian Li's talk distills several critical findings regarding platform engineering and organizational dynamics, emphasizing that success hinges less on technical prowess alone and more on human-centric approaches.
Firstly, Li dispels the myth of a "secret sauce" or a universal recipe for platform success. Her initial expectation of having all the answers after a few months at Evo Respra was quickly tempered by the complex realities of a large organization. This underscores that every platform journey is unique, shaped by its specific context, culture, and user needs.
A central finding is that platform building is inherently a collaborative effort. The platform team's role is to provide the underlying infrastructure and a robust base, but the ultimate "shape" and utility of the platform must be determined and owned by everyone who uses it. This means developers aren't just consumers; they are crucial contributors who can and should help define and even fix aspects of the platform, thereby freeing the platform team to focus on core capabilities.
Li stresses the paramount importance of building community. In any scenario involving three or more people, she argues, community becomes vital. It's the mechanism for aligning diverse stakeholders – users, service providers, and management – on the platform's vision, functionality, and evolution. Her experiences with the "microservices guild" (a success story) versus the "SRE user group" (a failure) vividly illustrate how community fosters shared understanding and problem-solving, moving discussions from blame to solutions.
Another critical insight is the necessity to communicate the vision and scope of the platform early and often. Li highlights that an unclear scope allows stakeholders to "dream up" their own idealized scenarios, leading to inevitable disappointment and confusion when the reality doesn't match their expectations. This proactive and persistent communication, even to the point of annoyance, is essential for managing expectations and preventing misunderstandings, such as the "Barbie dreamhouse" scenario.
Finally, Li emphasizes the hard-learned lesson to "let things go." Not every perfectly conceived solution will work in a given organizational context. The "best" or "perfect" solution is ultimately the one that is adopted and used effectively by the target audience, even if it deviates from an ideal technical design. Forcing a technically superior but unadopted solution is a path to failure, reinforcing that user needs and organizational fit outweigh theoretical perfection.
Technical Deep Dive
▶ Watch: Evo Respra: judiciary IT, value teams, tech stack (4:00)
The technical context for Lian Li's talk is set within Evo Respra's journey of evolving its core container management platform. The organization is migrating from an older OpenShift 4.16 environment, which has reached end-of-life and can no longer be upgraded, to a new platform based on OpenShift 4.17. This migration involves significant architectural shifts, particularly in cluster topology. The old setup had two development clusters solely for the platform team and two production clusters hosting all environments (test, dev, prod). The new platform rectifies this by dedicating two dev clusters to all development environments (dev, test) and two prod clusters exclusively to production workloads – a crucial improvement for stability and clarity. Furthermore, the new architecture introduces a management cluster to centrally manage the four operational clusters, with plans for an additional engineering cluster specifically for the platform team's internal development and testing.
A recurring technical pain point and source of organizational friction is the reliance on Azure DevOps. This platform serves as the central Git repository and provides the CI/CD pipelines for application delivery. However, its ownership by the separate OST (Developer Tooling) team creates a disconnect with the platform team's cloud-native focus. Li describes Azure DevOps as the "bane of my existence," highlighting the challenges of integrating disparate tooling philosophies and ownership models.
Li elaborates on the conceptual difference between services and platforms by illustrating deployment workflows. In a service-based model, a developer might first interact with a "build service" to get an image, then separately call a "deploy service" to push that image to production. The key characteristic here is that these services are loosely coupled and don't necessarily know about each other, which can lead to deviation in their expected inputs/outputs over time, requiring developers to perform manual steps in between. In contrast, a platform approach optimizes for the golden path. A developer might simply call a single "deploy service" within the platform, which then internally orchestrates the entire process, including calling the build service if a new image is needed. To the user, it appears as a single, cohesive unit, and the contract between the internal build and deploy components is maintained by the platform team, simplifying the developer experience. This evolution from local scripts to central services and then to an integrated platform is presented as a natural progression.
The distinction between a platform and a developer portal is also technically significant. A portal is presented as a user interface, often graphical, that sits on top of an existing platform. Its purpose is to surface the most common workflows and best practices, making them easily accessible (e.g., a button to provision a database with default parameters). The platform, however, is the underlying set of capabilities that can fulfill any reasonable request, irrespective of the portal's simplified interface.
A core technical challenge, magnified by organizational structure, is the "bricks vs. builds" dilemma. The platform team provides foundational components, but developers often lack the expertise or desire to assemble these into a complete, ready-to-use application environment. This highlights a gap in the developer experience (DX) that platform teams must address.
Li applies Team Topologies concepts to explain ideal collaboration. Stream-aligned teams (value teams) build applications, while the platform team provides the foundational infrastructure. Crucially, Li advocates for enabling teams that act as a cross-cut across stream-aligned teams and the platform team. These enabling teams, composed of SREs from value teams and potentially developer advocates from the platform team, bridge the knowledge gap. Li emphasizes, "I am not a .NET developer... I need you to know that," illustrating the necessity of domain-specific expertise from application teams to effectively collaborate on platform features.
A concrete example of technical friction arising from organizational decisions is the image registry saga. Initially, the platform used OpenShift's internal image registry. However, because the dev and prod clusters were completely separate, developers had to build every image twice (or even four times, considering test and dev environments within prod clusters), as there was no mechanism for promoting images. The platform team identified the need for an external image registry (considering options like Key and Harbor) to enable image promotion. However, business leadership decided to adopt Artifactory as a universal artifact repository, and crucially, assigned its ownership to the OST (Developer Tooling) team. This decision led to new friction points: the platform team now had to manage image pull secrets (a new requirement), and there were fundamental disagreements on CI/CD philosophies. While the platform team preferred OpenShift's native build configs, the OST team, rooted in a Microsoft Windows background, favored Azure DevOps pipelines and even PowerShell for build processes, leading to builds being moved out of OpenShift entirely. This exemplifies how organizational boundaries and differing technical backgrounds can create significant challenges in delivering a cohesive and efficient platform.
Demo / Proof of Concept
▶ Watch: OpenShift platform, migration, and cluster topology (6:00)
Lian Li's talk focused primarily on organizational, cultural, and strategic aspects of platform engineering rather than demonstrating specific technical implementations. As such, no live demo or detailed proof-of-concept was presented during the session. The speaker explicitly noted that the talk would "not be super technical" in terms of code or live demonstrations, instead emphasizing the human elements of platform adoption and collaboration.
Defensive Implications
▶ Watch: Hinting at current platform operational challenges (8:00)
While Lian Li's talk is not explicitly about security, the principles discussed have significant indirect defensive implications for organizations building internal developer platforms. A well-constructed and collaboratively managed platform can dramatically enhance an organization's security posture by embedding best practices and reducing attack surfaces.
Firstly, the concept of optimizing for the golden path is a powerful defensive mechanism. By providing developers with a streamlined, opinionated path for building and deploying applications, the platform team can bake in essential security controls. This includes ensuring that all deployed images originate from trusted, scanned registries (like a well-managed Artifactory), that deployments adhere to least privilege principles, and that configurations are standardized and secure by default. This reduces the likelihood of individual developers introducing vulnerabilities through ad-hoc processes or misconfigurations.
Secondly, the shift from disparate services and local scripts to a centralized platform improves visibility and auditability. When deployments happen through a single platform entry point, it becomes easier to log, monitor, and audit every stage of the application lifecycle. This centralized control provides a clearer picture of who deployed what, when, and where, which is crucial for incident response, compliance, and forensic analysis. In contrast, a fragmented, script-based approach makes it nearly impossible to maintain a consistent security baseline or track changes effectively.
The platform team's responsibility for maintaining the underlying infrastructure, such as upgrading OpenShift from 4.16 to 4.17, directly contributes to security. Timely application of patches and updates ensures that known vulnerabilities in the platform components are addressed, preventing exploitation. When this responsibility is centralized, there's a higher likelihood of consistent and timely security maintenance across all environments.
Moreover, the "bricks vs. builds" problem, when left unaddressed, can lead to shadow IT or developers bypassing established processes to meet deadlines. These unmanaged, ad-hoc solutions often lack proper security controls, creating significant blind spots and potential entry points for attackers. By providing a usable, "ready-made house" (or at least well-documented, easy-to-assemble components), the platform can guide developers towards secure practices, reducing the incentive for insecure workarounds.
Finally, the emphasis on community building and collaboration has a profound impact on security. An engaged community of developers and platform engineers can collectively identify security concerns, share knowledge about best practices, and contribute to security-focused features. For instance, an effective "enabling team" (composed of SREs and dev advocates) can help disseminate security awareness, ensure security requirements are understood, and facilitate the adoption of secure development patterns across value teams. While the image registry example highlighted friction, the underlying goal of centralizing artifact management in Artifactory, if collaboratively implemented, could lead to better security by enforcing policies on image provenance and integrity.
In essence, a well-managed and collaboratively built platform, guided by principles of clear scope and strong community, creates an environment where security is not an afterthought but an integral, embedded part of the development and deployment process, ultimately enhancing the organization's overall defensive posture.
Key Takeaways
- Platform building is a collaborative effort: The platform team provides the foundational infrastructure, but the ultimate shape and utility of the platform must be determined and owned by everyone who uses it, especially the application development teams.
- Community is paramount: Actively fostering a community among platform users and providers is essential for aligning diverse stakeholders, moving from blame to shared solutions, and ensuring the platform meets real-world needs.
- Communicate scope early and often: Clearly defining and continuously reiterating the platform's vision and scope is crucial to manage expectations, prevent disappointment, and avoid misinterpretations like the "Barbie dreamhouse" scenario.
- Embrace adaptability and "let things go": Not every technically ideal solution will be the best fit for a specific organization or user group. The most effective solution is one that is adopted and genuinely works for its users, even if it means compromising on theoretical perfection.
- Understand user needs deeply: Platform teams must bridge the gap between providing "building blocks" and delivering a "ready-made house" by deeply understanding what application developers truly need to be productive.
- Organizational structure and tooling choices matter: The ownership of critical tools (like Azure DevOps or Artifactory) by separate teams and differing technical backgrounds can create significant friction and hinder platform adoption, necessitating proactive collaboration and alignment.
About the Speaker(s)
Lian Li is a freelance platform engineer and self-described "cloud-native human" with a passion for bringing her "whole self" to her work. She is deeply involved in the cloud-native community, serving as a co-chair for the CNCF's TAG App Delivery. Beyond her technical expertise, Lian has a diverse background in the performing arts, having taken time off from tech to pursue stand-up comedy and musical theater, recently completing a three-week musical run. She is also widely known in the Kubernetes community as the "Chief Karaoke Officer" for Kobaraoke, the Kubernetes karaoke community. Lian's unique blend of technical insight and communicative flair makes her a compelling voice on the human aspects of technology and collaboration.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk by Lian Li is a brutally honest and deeply insightful war story from the trenches of platform engineering within a large government IT provider. While not a technical deep-dive in the kernel-exploit sense, it provides substantive detail on the socio-technical challenges of building and adopting an internal developer platform, focusing on the crucial human elements of collaboration, community, and managing expectations. It's a valuable case study that cuts through the marketing hype to deliver actionable lessons for anyone grappling with organizational friction in their platform journey.
Heather Calloway (CISO) — STRONG ACCEPT
This talk offers a clear-eyed look at the institutional realities of platform engineering. Lian Li cuts through the hype to diagnose the human and organizational friction points that often undermine even the most technically sound initiatives. Her emphasis on clear scope, collaborative ownership, and community building provides critical insights for any executive navigating large-scale IT transformation, directly impacting an organization's ability to manage risk and build resilient systems.