gRPC: 5 Years Later, Is It Still Worth It? - Konstantin Ostrovsky, Torq.io

Konstantin Ostrovsky, Torq.io

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

In this insightful KubeCon EU talk, Konstantin Ostrovsky, Chief Architect at Torq.io, offers a comprehensive retrospective on his team's five-to-six-year journey utilizing gRPC in a production microservices environment. The presentation addresses the critical question: "Is gRPC still worth it?" for modern development teams. Torq.io, a cybersecurity no-code automation startup, has embraced gRPC for inter-service communication across its Go-based microservices, even extending its use to front-end interactions, a less common but deliberate choice.

Watch on YouTube

Visual summary for gRPC: 5 Years Later, Is It Still Worth It? - Konstantin Ostrovsky, Torq.io by Konstantin Ostrovsky, Torq.io
Visual summary for gRPC: 5 Years Later, Is It Still Worth It? - Konstantin Ostrovsky, Torq.io by Konstantin Ostrovsky, Torq.io

Key moments

  1. 0:00 Introduction to gRPC journey and Torq.io
  2. 2:09 Why Torq chose gRPC over REST
  3. 4:13 What is gRPC? Core concepts explained
  4. 6:17 Key benefits: performance, compatibility, strong typing, code generation
  5. 8:00 Torq's API design rules and best practices
  6. 9:58 Recommended: Google API design proposal guide
  7. 10:43 Streamlining gRPC development experience with custom tooling

gRPC: 5 Years Later, Is It Still Worth It?

Speakers: Konstantin Ostrovsky, Chief Architect, Torq.io

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=q44WBAGzKhk

Overview

In this insightful KubeCon EU talk, Konstantin Ostrovsky, Chief Architect at Torq.io, offers a comprehensive retrospective on his team's five-to-six-year journey utilizing gRPC in a production microservices environment. The presentation addresses the critical question: "Is gRPC still worth it?" for modern development teams. Torq.io, a cybersecurity no-code automation startup, has embraced gRPC for inter-service communication across its Go-based microservices, even extending its use to front-end interactions, a less common but deliberate choice.

Ostrovsky meticulously details the initial motivations behind adopting gRPC, the significant benefits realized, and the numerous challenges encountered along the way. He shares invaluable lessons learned regarding API design, dependency management, front-end integration, and public API exposure. The talk culminates in a strong affirmation of gRPC's continued value, particularly with the advent of new tooling that addresses many of the historical pain points, providing a balanced and highly practical perspective for both existing gRPC users and those considering its adoption.

Background

▶ Watch: Introduction to gRPC journey and Torq.io (0:00)

Torq.io's decision to adopt gRPC was rooted in a prior negative experience with REST API tooling, specifically within the Go ecosystem. The team found that existing code generation tools for REST, such as those related to Swagger or OpenAPI, were subpar. This often led engineers to eschew generated code, instead implementing clients and servers manually, resulting in inconsistent, difficult-to-maintain codebases with multiple handler and client implementations for the same API. The primary goal with gRPC was to find a framework with robust, high-quality code generation capabilities, thereby reducing developer workload and allowing engineers to focus on core business logic rather than API boilerplate.

Beyond developer experience, gRPC offered several compelling technical advantages. As a binary protocol, it promised greater efficiency and faster data transmission over the wire compared to text-based protocols. Its backward compatibility by design was another crucial feature, ensuring API evolution could occur without immediately breaking existing clients, which is vital for a rapidly developing product. Ostrovsky defines gRPC as Google Remote Procedure Call, a high-performance RPC framework that is language-agnostic, supporting client and server generation across nearly all programming languages. It leverages Protocol Buffers (Protobuf) as its Interface Definition Language (IDL) and binary serialization format, and operates over HTTP/2, enabling efficient multiplexing of multiple requests over a single connection. The project's status as a CNCF project also provided assurance of community backing and long-term viability.

Key Findings

▶ Watch: What is gRPC? Core concepts explained (4:13)

Konstantin Ostrovsky's five-year retrospective on gRPC at Torq.io yielded several key findings, highlighting both its strengths and weaknesses in a production environment. The most significant benefit identified was unified API standards, which streamlined development, reduced code review overhead, and allowed engineers to seamlessly transition between teams. The robust middleware ecosystem and built-in extensibility of gRPC proved invaluable for implementing cross-cutting concerns like authentication, authorization, and feature management directly within the API definitions.

However, the journey was not without considerable friction. Initial challenges revolved around establishing a consistent directory structure, managing Protobuf dependency management in a multi-repo setup, and the generally less-than-ideal developer experience for front-end engineers due to the binary nature of gRPC and the limitations of gRPC-web. Difficulties with HTTP/2 load balancing on Kubernetes and the inherent issues with message reuse across service boundaries were also significant pain points. Despite these hurdles, the emergence of powerful new tools like Buff and Connect RPC has fundamentally transformed the gRPC landscape, addressing many of these long-standing issues and ultimately solidifying gRPC's position as a worthwhile choice for Torq.io.

Technical Deep Dive

▶ Watch: Key benefits: performance, compatibility, strong typing, code generation (6:17)

Torq.io's adoption of gRPC was a deliberate architectural choice, underpinning its entire microservices ecosystem. The core of their gRPC implementation relies on Protocol Buffers for defining service contracts. Developers define services and messages (input/output structures) using the Protobuf IDL, then use a CLI tool to generate language-specific client and server code. This process simplifies the development workflow significantly, as engineers then only need to implement the server logic and utilize the generated clients for communication. The choice of HTTP/2 as the underlying transport protocol allows for efficient multiplexing of requests over a single connection, reducing latency and improving efficiency for high-request-rate applications.

API Governance and Design Principles

Torq.io places a strong emphasis on API governance to maintain consistency and quality across its 50-engineer team. Their approach includes:

  • Linting: Everything that can be linted is integrated into the CI pipeline to enforce consistent standards and reduce cognitive load for engineers, allowing them to focus on business logic.
  • Backward Compatibility Enforcement: A linter explicitly checks for backward-breaking changes, such as deleting RPC methods, altering field types, or removing fields. This ensures that API evolution doesn't inadvertently disrupt existing clients or services.
  • Message Duplication over Reuse (across service boundaries): A crucial lesson learned was to avoid sharing common Protobuf messages (e.g., a User object) across different service domains. While tempting to adhere to DRY principles, this often led to tight coupling and "spaghetti" dependencies between services. Torq.io now prefers to duplicate messages within their respective service domains, only reusing messages within a single microservice. This reduces inter-service coupling and improves maintainability.
  • Google API Design Proposal: The team extensively consulted Google's API design guidelines, adapting specific parts that resonated with their needs, recommending it as a valuable resource for API design thinking.

Development Workflow and Project Structure

To streamline the developer experience, Torq.io created a custom Docker image containing all necessary Protobuf plugins and shared configuration. This eliminates the need for individual developers to install compilers or plugins, ensuring a consistent build environment.

For project organization, Torq.io uses a multi-repo structure. Each service's API definitions (.proto files) reside in an api directory within its own service repository. Since Go is their primary language, the generated Go clients and servers are committed directly to the service's GitHub repository. This allows other Go services to consume these clients easily via Go module management (go get). For front-end consumption, generated artifacts are published to a GitHub artifact package registry. A typical microservice directory structure includes an api directory, which contains subdirectories for each API version (e.g., v1), holding the specific .proto files.

Middleware and Extensibility

gRPC's design naturally supports a rich middleware ecosystem, akin to REST API frameworks. Torq.io leverages this heavily for:

  • Authentication and Authorization: Centralized handling of security concerns.
  • Observability: Integration with logging, tracing, and monitoring systems.

A standout feature for Torq.io is gRPC's extensibility through options. They demonstrate this by embedding Role-Based Access Control (RBAC) information directly into the .proto service definitions. Each RPC method can be annotated with an options attribute specifying the required scopes for a session. For example:

This ensures that access control is defined alongside the API contract and enforced by middleware. Similarly, they implement flag-based feature management by annotating RPCs with feature flags, preventing access to specific functionalities if the corresponding flag is not enabled, enhancing operational control and release management.

Addressing Key Challenges

Torq.io faced several significant challenges, and their solutions or workarounds provide valuable insights:

  • Kubernetes HTTP/2 Load Balancing: Standard Kubernetes load balancing struggles with HTTP/2's long-lived connections, leading to uneven load distribution across pods. Torq.io mitigated this by deploying Linkerd, a service mesh, which handles intelligent load balancing for HTTP/2 connections out-of-the-box.
  • Front-end Integration (gRPC-web): Front-end engineers often prefer REST APIs due to easier debugging with browser network tabs. gRPC-web, which translates gRPC calls into Base64-encoded payloads over POST requests, presents several issues:
  • No Caching: It breaks standard browser and CDN caching mechanisms.
  • Debugging Difficulty: Requires special browser extensions, making customer support and debugging challenging.
  • Proxy Requirement: Necessitates an intermediary proxy to translate between gRPC-web and native gRPC.
  • Public APIs: Exposing gRPC APIs directly to external customers is generally not feasible. While gRPC Gateway allows annotating gRPC services with REST definitions to auto-generate a REST API, Torq.io found it rigid. Technical writers struggled with the annotations, and post-processing scripts were needed to format the generated Swagger (OpenAPI) files. In retrospect, Ostrovsky suggests manually defining OpenAPI files for public APIs to gain full control over documentation and presentation.
  • Protobuf Dependency Management: Initially, managing dependencies across multiple repositories was a custom, complex endeavor. This was largely solved by Buff, a new CLI tool and ecosystem for Protobuf. Buff provides:
  • Remote Plugins: Simplified plugin management.
  • Single Configuration File: Centralized configuration for code generation.
  • Linting and Backward Compatibility Checks: Built-in enforcement.
  • Buff Service Registry: A product specifically designed to manage and reuse Protobuf messages across different services, directly addressing Torq's "duplicate don't reuse" challenge.

The Game Changer: Buff and Connect RPC

The advent of Buff and Connect RPC significantly improved Torq.io's gRPC experience. Buff, developed by ex-Google engineers familiar with Protobuf pain points, provided a complete ecosystem that allowed Torq to discard their custom Docker image and complex Makefiles, replacing them with a simple buff generate command and a YAML configuration.

Connect RPC, a protocol built by the Buff team, addresses many gRPC-web limitations. It offers:

  • Text-based JSON Format: While a new protocol, it provides a JSON-based alternative to binary Protobuf, making it easier for front-end debugging.
  • Full Backward Compatibility: Supports gRPC, gRPC-web, and its own protocol within the same generated code.
  • Improved JavaScript Libraries: Crucially, Connect RPC provides idiomatic TypeScript and JavaScript libraries, which has been a "huge developer experience improvement" for front-end engineers, making gRPC much more palatable for browser-based applications.

Demo / Proof of Concept

▶ Watch: Recommended: Google API design proposal guide (9:58)

The talk focused on a retrospective of real-world implementation experiences and architectural decisions at Torq.io rather than a live demonstration or proof-of-concept. Konstantin Ostrovsky provided code snippets and architectural diagrams to illustrate how gRPC is used and extended within their environment.

Defensive Implications

▶ Watch: Streamlining gRPC development experience with custom tooling (10:43)

The technical choices and lessons learned by Torq.io, while primarily focused on developer efficiency and system architecture, carry significant defensive implications for security practitioners.

Firstly, the emphasis on strong API contracts and backward compatibility enforcement through linting directly contributes to a more secure system. By preventing accidental breaking changes or unauthorized modifications to API definitions, the risk of introducing vulnerabilities due to inconsistent API behavior or unexpected data formats is significantly reduced. This disciplined approach ensures that security controls defined at the API layer remain consistent and effective over time.

Secondly, Torq.io's innovative use of gRPC options for embedded security controls like Role-Based Access Control (RBAC) and feature flag management is a powerful defensive strategy. Defining access scopes directly within the .proto files, alongside the API method definition, creates a single source of truth for authorization requirements. This minimizes the chance of authorization bypasses due to misconfigurations or overlooked security checks in the application code, as the security policy is intrinsically linked to the API contract and enforced by a centralized middleware.

Thirdly, the extensive use of middleware for cross-cutting concerns like authentication, authorization, and observability provides a robust security posture. By centralizing these critical functions, Torq.io ensures that security policies are applied consistently across all microservices, reducing the surface area for common vulnerabilities such as broken authentication or insufficient logging. A well-tested and maintained security middleware layer can significantly enhance the overall resilience of the system.

Finally, the adoption of a service mesh like Linkerd to address Kubernetes HTTP/2 load balancing issues also offers defensive benefits. Service meshes inherently provide capabilities beyond simple traffic management, including mutual TLS (mTLS) for encrypted and authenticated communication between services, fine-grained traffic policies, and enhanced observability, all of which strengthen the security of the microservices architecture. While the talk specifically mentioned load balancing, the underlying technology contributes to a more secure network fabric. The speaker's retrospective on public API exposure also highlights a critical security consideration: the trade-offs between automated generation (like gRPC Gateway) and manual definition (OpenAPI) for external-facing APIs. Manual definition, despite the effort, can ensure that public API contracts are clearly articulated, properly documented, and subjected to rigorous security review, which might be more challenging with automatically generated, less customizable formats.

Key Takeaways

  • Standardization and Efficiency: gRPC excels in promoting unified API standards, significantly reducing developer workload by generating client/server code and streamlining internal microservice communication.
  • Initial Friction Points: Adopting gRPC initially presents challenges, particularly in managing Protobuf dependency across multi-repo setups, designing consistent directory structures, and integrating with front-end applications.
  • Powerful Extensibility: gRPC's options mechanism allows for embedding critical cross-cutting concerns like Role-Based Access Control (RBAC) and feature flags directly into API definitions, enhancing security and operational control.
  • Kubernetes and HTTP/2 Load Balancing: Standard Kubernetes load balancers struggle with HTTP/2's long-lived connections; a service mesh like Linkerd or custom client-side logic is essential for proper load distribution.
  • gRPC-web Limitations: While enabling browser communication, gRPC-web introduces issues with caching, debugging (requiring special browser extensions), and necessitates an intermediary proxy.
  • Transformative New Tooling: The emergence of Buff and Connect RPC has significantly improved the gRPC developer experience, addressing long-standing issues with dependency management, build processes, and front-end integration through idiomatic JavaScript/TypeScript libraries.

About the Speaker(s)

Konstantin Ostrovsky is the Chief Architect at Torq.io, a cybersecurity no-code automation startup. With 12 to 15 years of experience in various cybersecurity startups, he has a strong passion for optimizing CI build times. At Torq.io, he leads a team of approximately 50 engineers, focusing on a microservices architecture built with Go, gRPC (used for "everything"), PostgreSQL, and Redis, all running on Google Cloud Platform (GCP). Torq.io processes millions of automations daily, providing a Zapier-like experience for security analysts to build automations via a graphical user interface without writing code.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk offers a brutally honest and deeply technical retrospective on five years of gRPC in a production microservices environment. The Chief Architect from Torq.io doesn't just rehash documentation; he shares hard-won lessons on API governance, dependency management, and the real-world friction of gRPC-web. The session highlights the transformative impact of new tooling like Buff and Connect RPC, making it highly relevant for anyone running or considering gRPC at scale. His practical approach to embedding security controls like RBAC directly into Protobuf definitions is particularly clever.

Heather Calloway (CISO) — STRONG ACCEPT

This KubeCon talk offers a candid and valuable retrospective on adopting gRPC, moving beyond technical specifics to highlight critical lessons in API governance, security embedding, and operational efficiency. The speaker provides concrete examples of how to enforce standards, manage dependencies, and centralize security controls within a microservices architecture, which are directly applicable to strengthening institutional accountability and managing business risk. The discussion on new tooling like Buff and Connect RPC demonstrates a clear path forward for teams navigating similar architectural complexities.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025