Vitess: Schema Changes at Scale - Rohit Nayak & Shlomi Noach, PlanetScale
Rohit Nayak, Shlomi Noach, PlanetScale
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
This talk by Rohit Nayak and Shlomi Noach from PlanetScale provides an in-depth exploration of how Vitess, a robust database clustering system for MySQL, addresses the notoriously complex challenge of schema changes at scale. Vitess, originating from YouTube at Google and now open-source, enables MySQL to operate as a massively scalable, highly available, and cloud-native distributed database. The core problem addressed is the inherent difficulty and downtime associated with altering large tables in traditional MySQL environments, a problem that is dramatically compounded in a sharded, distributed setup.

Key moments
- 0:00 What is Vitess and its real-world applications?
- 2:00 Vitess origins: Google, open-source, MySQL as storage.
- 3:30 Vitess's unique configurable horizontal sharding capabilities.
- 4:40 Architectural overview for schema change management begins.
- 6:00 Understanding Vitessgate: the central query engine.
- 7:00 Query routing and sharding configuration explained.
- 8:00 The problem: MySQL schema changes and downtime.
Vitess: Schema Changes at Scale
Speakers: Rohit Nayak, Maintainer at PlanetScale; Shlomi Noach, Maintainer at PlanetScale
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=Zbi46yTlSVo
Overview
This talk by Rohit Nayak and Shlomi Noach from PlanetScale provides an in-depth exploration of how Vitess, a robust database clustering system for MySQL, addresses the notoriously complex challenge of schema changes at scale. Vitess, originating from YouTube at Google and now open-source, enables MySQL to operate as a massively scalable, highly available, and cloud-native distributed database. The core problem addressed is the inherent difficulty and downtime associated with altering large tables in traditional MySQL environments, a problem that is dramatically compounded in a sharded, distributed setup.
The speakers, both long-time Vitess maintainers, highlight Vitess's critical role in the infrastructure of major tech companies like Slack, Shopify, HubSpot, and Cash App, serving millions of queries per second and demonstrating its resilience during events like the COVID-19 pandemic. Their presentation delves into the architectural components that facilitate seamless schema migrations across potentially hundreds of MySQL shards, focusing on the mechanisms that ensure idempotency, consistency, concurrency, and resiliency—qualities essential for maintaining continuous operation in high-stakes production environments.
This article dissects the advanced techniques Vitess employs to perform online schema changes, transforming a traditionally disruptive database operation into a manageable and reliable process. It covers the underlying architecture, the specific challenges posed by distributed databases, and Vitess's innovative solutions, offering valuable insights for database administrators, developers, and architects grappling with scalability and availability concerns in MySQL-based systems.
Background
▶ Watch: What is Vitess and its real-world applications? (0:00)
The problem of applying schema changes to large MySQL tables is a long-standing challenge in database management. In a traditional, single-instance MySQL setup, operations like ALTER TABLE (e.g., adding a column, modifying an index) can result in table locks, rendering the table unavailable for reads and writes for extended periods—potentially hours or even days for tables containing billions of rows. This "stop the world" scenario is unacceptable for high-availability applications. While modern MySQL versions have introduced some online operations and instant DDL for specific, limited scenarios, they do not cover the full spectrum of schema modifications.
To mitigate this, the MySQL ecosystem has developed various online schema change tools such as gh-ost (authored by Shlomi Noach), pt-online-schema-change, and Percona Toolkit's spirit. These tools generally operate on a similar principle:
- A shadow table (or "ghost" table) is created with the desired new schema, initially empty.
- The schema change is applied to this empty shadow table, which is a fast operation.
- Rows from the original table are copied to the shadow table.
- During the copy process, all ongoing data modifications (inserts, updates, deletes) to the original table are captured (e.g., via MySQL's binary log) and applied to the shadow table, keeping it in sync.
- Once the shadow table is fully populated and in sync with the original, a brief "stop the world" moment occurs where both tables are locked.
- The original and shadow tables are then atomically swapped (renamed), making the new schema active.
- The old table is eventually dropped.
Vitess builds upon this fundamental online schema change approach but extends it dramatically to handle horizontally sharded MySQL clusters. Vitess's architecture is designed to abstract away the complexity of managing multiple MySQL instances. At its core, it comprises:
- MySQL Servers: The underlying storage layer.
- vttablet: A sidecar process attached to each MySQL instance. It controls the MySQL server (startup, shutdown, backup/restore) and, critically, acts as a proxy, controlling all traffic to and from the MySQL server. Applications never connect directly to MySQL.
- vtgate: The query proxy that sits in front of all MySQL clusters. It acts as a single logical MySQL server to client applications, handling query routing, load balancing, and sharding logic. It translates incoming SQL queries into operations on the correct shards.
- Topology Server: A non-data path component (e.g., etcd, ZooKeeper, Consul) that stores the sharding scheme and routing rules, which
vtgateloads into memory.
The inherent challenge in a sharded environment is that each MySQL server operates independently; it is unaware that it is part of a larger, sharded system. This independence, while beneficial for scaling, makes coordinating a single logical schema change across potentially dozens or hundreds of physical shards exceptionally difficult, introducing problems related to idempotency, consistency, and resiliency that Vitess aims to solve.
Key Findings
▶ Watch: Vitess's unique configurable horizontal sharding capabilities. (3:30)
The talk highlights Vitess's comprehensive approach to managing schema changes in a distributed MySQL environment, presenting several key findings and contributions that overcome the inherent complexities of sharding.
- Distributed Online Schema Change Framework: Vitess provides a built-in, highly automated online schema change mechanism that extends the shadow table approach to a multi-sharded context. Instead of individual DBAs managing changes per shard, Vitess orchestrates the process across the entire cluster.
- Idempotent Migrations: Vitess introduces mechanisms to ensure that schema change commands can be re-applied safely without causing errors or unintended side effects. This is crucial for handling partial failures in a distributed system, where some shards might succeed while others fail or are unavailable. The use of a job ID (UUID) for each migration allows shards to recognize and skip already-processed requests.
- Declarative Migrations: Adopting a Kubernetes-like philosophy, Vitess supports declarative schema definitions. Users can specify the desired end state (
CREATE TABLE,DROP TABLE) rather than proceduralALTER TABLEstatements. Each shard independently computes the necessary DDL to reach the desired state, making the system resilient to schema drift and simplifying recovery from inconsistencies. - Controlled Consistency Across Shards: Vitess addresses the challenge of schema inconsistencies during long-running migrations by allowing users to postpone completion. Shards can complete their internal migration steps (copying data, applying changes) but will not perform the final table swap until explicitly commanded, and ideally, until all other shards are also ready. This enables near-atomic cutovers across all shards, minimizing the window of schema inconsistency to a few seconds.
- Concurrent Migrations: Vitess is designed to handle multiple schema migrations running simultaneously across all shards. While some serialization might occur for performance, the system allows for the parallel execution of related or independent DDLs, further accelerating deployment cycles.
- Resilient and Stateful Migrations: A critical innovation is the use of VReplication to make schema migrations stateful and resumable. If a primary MySQL instance fails during a migration, Vitess can promote a new primary, and the migration automatically resumes from its exact point of interruption, preventing the loss of days of work. This journaling system significantly enhances the fault tolerance of DDL operations in a sharded environment.
- Safety Checks and Revertability: Vitess incorporates features to analyze potential data loss risks during schema changes (e.g., through its
schemalibrary) and, crucially, offers a revert migration capability. This allows users to roll back to the original schema after a cutover, using VReplication to propagate changes back, providing a safety net for unexpected issues.
These findings collectively demonstrate that Vitess transforms schema management in large-scale MySQL deployments from a high-risk, high-downtime operation into a robust, automated, and highly available process, enabling organizations to evolve their database schemas with confidence and speed.
Technical Deep Dive
▶ Watch: Architectural overview for schema change management begins. (4:40)
Vitess's solution for schema changes at scale is deeply integrated into its architecture, leveraging vtgate, vttablet, and the underlying MySQL replication mechanisms to orchestrate distributed DDLs.
When a user initiates a schema change, such as ALTER TABLE products ADD COLUMN description VARCHAR(255), through vtgate, the vtgate component does not execute this directly on a single MySQL instance. Instead, it acts as a coordinator. It identifies all relevant shards that host parts of the products table and broadcasts the schema change request to the primary vttablet of each of these shards. Each vttablet then takes ownership of executing the online schema change locally on its respective MySQL primary.
The local online schema change process within each shard follows the established pattern:
- Shadow Table Creation:
vttabletcreates a new, empty shadow table (e.g.,_products_new) with the desired schema. - Schema Application: The
ALTER TABLEstatement is applied to this empty shadow table, which is near-instantaneous. - Data Copy:
vttabletbegins copying existing data from the originalproductstable to_products_new. - Change Log Application: Simultaneously,
vttabletmonitors the MySQL binary log (or uses VReplication) to capture all ongoing DML (inserts, updates, deletes) targeting the originalproductstable. These changes are then applied to_products_new, ensuring that the shadow table stays in sync. - Atomic Swap: Once the shadow table is fully backfilled and caught up with the primary,
vttabletperforms a brief "stop the world" operation, locking both tables. It then renames the originalproductstable to a temporary name (e.g.,_products_old), and the shadow table_products_newtoproducts. This atomic swap ensures that applications immediately start interacting with the new schema. The_products_oldtable is held for a grace period before being dropped, allowing for potential reverts.
The real technical innovation lies in how Vitess manages these operations across multiple, independent shards:
Idempotency: In a distributed system, network issues or temporary shard unavailability are common. If a schema change request fails to reach some shards, re-issuing the command is problematic if other shards have already completed it. Vitess addresses this by assigning a unique job ID (UUID) to each migration. When vtgate sends a migration request, it includes this UUID. If a vttablet receives a request with a UUID it has already processed, it treats it as a no-op and reports success, preventing duplicate indexes or syntax errors.
A more advanced form of idempotency is declarative migrations. Instead of specifying ALTER TABLE, users can define the target schema using CREATE TABLE and DROP TABLE statements. Each vttablet independently compares the desired schema with its current schema. If they match, it's a no-op. If they differ, vttablet computes the necessary ALTER TABLE statement to transition from the current state to the desired state. This "Kubernetes way" of state management ensures that shards converge to the correct schema regardless of their starting point or past failures.
Consistency: Shards operate independently and may complete their local online schema changes at different times due to varying workloads, hardware, or network conditions. This can lead to a period where different shards present different schemas, confusing applications or engineers. Vitess offers a postpone completion strategy. When enabled, a vttablet will complete the data copying and change application phase but will not perform the final table swap automatically. It will continue to apply ongoing changes to keep the shadow table in sync, waiting for explicit instruction.
Users can monitor the status across all shards using SHOW VITESS_MIGRATIONS. This command returns a ready_to_complete status for each shard (0 for not ready, 1 for ready). Once all shards report ready_to_complete=1, the user can issue an ALTER VITESS_MIGRATION COMPLETE command. vtgate then broadcasts this command to all vttablets, which perform their atomic swaps concurrently. In practice, this achieves a near-simultaneous cutover across all shards, typically within a few seconds, minimizing schema inconsistency. For critical scenarios, Vitess can be configured to forcibly terminate any queries holding locks on the migrated table just before the cutover to ensure success.
Concurrency: Vitess allows for multiple concurrent migrations across all shards. This is useful when several related schema changes need to be deployed together (e.g., adding multiple columns). vtgate manages the fan-out of these concurrent requests, and each vttablet processes them, potentially with some internal serialization for resource management. The same consistency mechanisms (postpone, ready_to_complete, complete all) apply, allowing a coordinated cutover for a batch of migrations. Even with multiple large tables on many shards, the cutover window remains in the order of seconds.
Resiliency: The most significant advancement in resiliency is the use of VReplication. Unlike traditional online schema change tools that might lose days of work if a primary MySQL server crashes mid-migration, Vitess leverages VReplication as a journaling system. During the data copy and change application phases, VReplication audits not only the data transfer but also the exact position in the MySQL binary log where changes are applied. This makes the migration stateful. If a primary fails and a new primary is promoted, the vttablet on the new primary recognizes the interrupted migration, inspects the VReplication journal, and resumes the migration from the precise point of interruption. This recovery typically takes less than a minute, virtually eliminating downtime and wasted effort due to primary failures.
Foreign Key Limitations: The speakers acknowledge a specific limitation with MySQL's handling of foreign keys during table swaps. InnoDB, which manages foreign keys, makes it impossible to replace a table while keeping foreign keys intact without special handling. Vitess addresses this by maintaining a public fork of MySQL that includes fixes for this behavior. Users needing foreign key support during online schema changes must use this specific MySQL version and enable a d-unsafe allow foreign keys flag.
Safety Checks: Vitess also provides tools to proactively identify risks. The schema library (available on the Vitess and PlanetScale blogs) can compare two schemas and advise on potential data loss scenarios (e.g., dropping a column, reducing column scope, changing unique index properties). While Vitess offers a revert migration option (using VReplication to propagate changes back to the original table and then swapping back), some data corruption scenarios (like reducing an INT to a TINYINT after large numbers have been inserted) might be irreversible, emphasizing the need for pre-migration risk analysis.
Demo / Proof of Concept
▶ Watch: Query routing and sharding configuration explained. (7:00)
The talk did not include a live demonstration or a specific proof of concept. The speakers focused on explaining the architectural and operational mechanisms behind Vitess's schema change capabilities through conceptual diagrams and detailed explanations of the command flows and outcomes.
Defensive Implications
▶ Watch: The problem: MySQL schema changes and downtime. (8:00)
The advanced schema change capabilities of Vitess have profound implications for database defenders and architects aiming to build and maintain highly available, scalable applications on MySQL.
- Eliminate DDL-Related Downtime: The primary defensive implication is the ability to perform schema changes on even the largest tables without incurring application downtime. Organizations can avoid "maintenance windows" for schema changes, significantly improving service availability and developer agility. This directly translates to better SLOs and customer satisfaction.
- Ensure Schema Consistency: Vitess's
postpone completionand coordinated cutover mechanisms are crucial for preventing schema drift across shards. Defenders can ensure that all application instances see a consistent schema view, reducing the likelihood of application errors or data corruption caused by queries hitting different schema versions on different shards. MonitoringSHOW VITESS_MIGRATIONSforready_to_completestatus becomes a key operational practice.
- Enhanced Resiliency Against Failures: The use of VReplication for stateful, resumable migrations means that primary MySQL failures during a schema change no longer result in lost work or extended outages. This significantly reduces operational burden and risk, allowing for more aggressive DDL deployment schedules. Defenders should understand this mechanism to confidently plan and execute migrations.
- Safe and Revertible Operations: The revert migration feature provides a critical safety net. In case of unforeseen issues or application regressions after a schema change, the ability to roll back to the previous schema while preserving data changes is invaluable. This empowers teams to deploy changes with greater confidence, knowing there's a well-defined recovery path.
- Proactive Risk Assessment: Tools like the
schemalibrary, mentioned by the speakers, are vital for proactive defense. By analyzing potential data loss or constraint violation risks before initiating a migration, teams can prevent catastrophic failures. Integrating such checks into CI/CD pipelines for DDL changes is a recommended best practice.
- Consider Foreign Key Implications: While Vitess offers a workaround for foreign keys with its public MySQL fork, defenders must carefully evaluate the implications. If foreign keys are critical for data integrity, adopting the specific Vitess-patched MySQL and understanding the
d-unsafe allow foreign keysflag is essential. Alternatively, re-evaluating the need for declarative foreign keys in a sharded environment (where application-level integrity might be more suitable) could be considered.
- Operational Simplicity: By automating the complexities of distributed DDLs, Vitess simplifies the operational overhead for large-scale MySQL deployments. Database teams can focus on strategic tasks rather than manually coordinating and troubleshooting schema changes across numerous shards.
In essence, Vitess transforms schema changes from a dreaded, high-risk event into a routine, automated, and highly resilient process, enabling organizations to scale their databases and evolve their applications without compromising availability or data integrity.
Key Takeaways
- Vitess enables truly online schema changes at massive scale: It extends traditional online schema change methods to horizontally sharded MySQL clusters, eliminating downtime for DDL operations on terabytes of data and millions of queries per second.
- Idempotency and declarative migrations ensure reliability: Vitess uses unique job IDs and supports declarative schema definitions, allowing for safe retries and automatic convergence to the desired schema state across all shards, even in the face of partial failures.
- Controlled consistency is achieved through postponed completion: Shards can complete their local migration work but will not cut over until all shards are ready, allowing for near-atomic, cluster-wide schema activation within seconds.
- VReplication powers resilient and resumable migrations: By leveraging VReplication as a journaling system, Vitess migrations are stateful and can automatically resume from the exact point of interruption after a primary MySQL failure, preventing significant data loss or wasted effort.
- Vitess allows concurrent schema changes: Multiple DDL operations can run simultaneously across all shards, improving deployment speed and flexibility for complex schema evolution.
- Safety and revertability are built-in: Vitess offers tools for pre-migration risk assessment (e.g.,
schemalibrary) and provides arevert migrationcapability, allowing rollbacks to the previous schema while preserving data changes.
About the Speaker(s)
Rohit Nayak is a maintainer of Vitess, the open-source database clustering system for MySQL. He has been actively contributing to and maintaining Vitess for approximately five years, working at PlanetScale. His expertise lies in the operational aspects and core development of Vitess, helping to ensure its stability and scalability for users worldwide.
Shlomi Noach is also a maintainer of Vitess at PlanetScale, with five years of experience in the project. Beyond his contributions to Vitess, Shlomi is a well-known figure in the MySQL community. He is the author of several popular open-source tools, including gh-ost, an online schema migration tool, and orchestrator, a MySQL high availability and replication management solution. His deep understanding of MySQL internals and distributed systems is evident in his work on Vitess's advanced schema change capabilities.
The talk acknowledged that Deepti Siger, who was the tech lead for this module, was originally scheduled to present but had to withdraw due to personal issues.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This talk by the Vitess maintainers at PlanetScale is a deep dive into solving one of the hardest problems in distributed database management: online schema changes at scale. It clearly articulates the challenges of sharded MySQL environments and presents Vitess's innovative, battle-tested solutions leveraging idempotency, declarative migrations, VReplication for statefulness, and coordinated cutovers. This isn't just theory; it's a critical piece of infrastructure engineering that enables truly continuous operation for massive MySQL deployments.
Heather Calloway (CISO) — STRONG ACCEPT
This presentation by the Vitess maintainers offers a compelling solution to a long-standing operational and business risk: schema changes at scale causing unacceptable downtime in high-availability environments. Vitess's engineered approach, leveraging distributed online schema changes, idempotency, and VReplication, transforms a traditionally disruptive operation into a resilient, automated process. This significantly reduces operational risk, enhances business continuity, and provides critical confidence for executive decision-makers managing large-scale, sharded MySQL deployments.