DevOps · CNCF daily
Vitess: Horizontal Scaling for MySQL
CNCF Graduated. A database clustering system that shards MySQL horizontally behind a single MySQL-compatible endpoint, built at YouTube and run at Slack, GitHub and Shopify.
CNCF Graduated
- Vitess
- MySQL
- Kubernetes
- etcd
- Vitess Operator
Vitess is a database clustering system for horizontally scaling MySQL. Applications connect to it as if it were a single MySQL server. Behind that endpoint, Vitess routes each query to the right shard (an ordinary MySQL instance holding part of the data), manages replication and failover, and can split or merge shards online without application changes. It addresses the point every successful MySQL deployment eventually reaches: one primary can no longer handle the writes or hold the data, and application-level sharding becomes a long-term engineering burden. Vitess was built at YouTube in 2010 to scale its MySQL fleet. It was accepted into the CNCF as an incubating project in February 2018 and graduated in November 2019, the first storage project to do so.
CNCF daily · Sprint 1, day 1 · Database
At a glance
| Category | Database |
| CNCF status | Graduated. Accepted February 5, 2018 (Incubating); graduated November 5, 2019 |
| Written in | Go |
| License | Apache 2.0 |
| Storage engine | Standard MySQL (InnoDB) on every shard |
| Topology store | etcd, ZooKeeper or Consul |
| Kubernetes deployment | Vitess Operator (VitessCluster custom resource) |
Architecture
flowchart TB
APP[Applications<br/>MySQL protocol or gRPC] --> GATE[VTGate pool<br/>query routing]
GATE --> TOPO[(Topology service<br/>etcd)]
subgraph KS[Keyspace commerce, sharded]
subgraph S1[Shard -80]
T1P[VTTablet + MySQL<br/>primary]
T1R[VTTablet + MySQL<br/>replicas]
T1P --> T1R
end
subgraph S2[Shard 80-]
T2P[VTTablet + MySQL<br/>primary]
T2R[VTTablet + MySQL<br/>replicas]
T2P --> T2R
end
end
GATE --> T1P
GATE --> T2P
GATE --> T1R
GATE --> T2R
ORC[VTOrc<br/>failure detection] --> T1P
ORC --> T2P
CTL[vtctld and VTAdmin<br/>operations] --> TOPO
CTL --> T1P
CTL --> T2P
| Component | Responsibility |
|---|---|
| VTGate | Stateless proxy that speaks the MySQL protocol, parses queries and routes them to the right shards, then merges results. |
| VTTablet | Sidecar beside each MySQL instance. Manages replication, enforces query limits and connection pooling, and runs resharding workflows. |
| Keyspace / shard | A keyspace is a logical database; it can be split into shards, each a primary with replicas holding a key range. |
| VSchema | Describes how tables are sharded: which column and vindex (sharding function) map each row to a shard. |
| Topology service | Stores cluster metadata (keyspaces, shards, tablet roles) in etcd, ZooKeeper or Consul. |
| VTOrc | Detects failed primaries and replication problems, and performs automated failover. |
| vtctld / VTAdmin | Control plane and web UI for operations such as schema changes, resharding and traffic switching. |
Design principle. Vitess keeps MySQL as the storage engine and adds sharding, routing and orchestration around it. Each shard remains an ordinary MySQL database, so existing tools, backups and expertise still apply. VTGate hides the shard layout, which is why data can be resharded online while applications keep using the same connection string.
Production reference design
A sharded cluster on EKS, GKE, AKS or on-premises Kubernetes, managed by the Vitess Operator:
- Install the operator from the Vitess repository’s operator example:
git clone https://github.com/vitessio/vitess.git && cd vitess/examples/operator kubectl apply -f operator.yaml - Create the cluster as a
VitessClusterresource, starting with one unsharded keyspace: one primary and two replicas per shard across availability zones, an etcd-backed topology, several VTGate replicas behind a Service, and backups to object storage (S3, GCS or MinIO). - Point applications at VTGate using a normal MySQL driver. Migrate data from an existing MySQL server with Vitess’s
MoveTablesworkflow, which copies and then streams changes until you switch traffic. - Define the sharding scheme in the VSchema before splitting. For example, shard the
customertable bycustomer_id:
Sharding related tables by the same key keeps a customer’s rows on one shard, so joins and transactions stay local.{ "sharded": true, "vindexes": { "hash": { "type": "hash" } }, "tables": { "customer": { "column_vindexes": [{ "column": "customer_id", "name": "hash" }] }, "corder": { "column_vindexes": [{ "column": "customer_id", "name": "hash" }] } } } - Reshard online when one shard is no longer enough, then switch reads and writes once the target shards have caught up:
vtctldclient Reshard --workflow cust2cust --target-keyspace customer \ create --source-shards '-' --target-shards '-80,80-' vtctldclient Reshard --workflow cust2cust --target-keyspace customer switchtraffic
Typical production uses include large multi-tenant SaaS databases, high-write consumer applications that have outgrown a single MySQL primary, and consolidating many MySQL instances under one operational model.
Operational considerations
- Choose the sharding key carefully. It is the most important decision. A key that matches the main access pattern (tenant or customer ID) keeps queries on one shard; a poor key causes scatter queries across every shard.
- Know the SQL limits. Vitess supports most MySQL syntax, but cross-shard transactions, some joins and certain functions behave differently or cost more. Test the application’s real queries against VTGate early.
- Plan for high availability. Run VTOrc, keep at least two replicas per shard in separate zones, size VTGate for peak connections and keep the topology service (etcd) highly available.
- Backups and schema changes. Schedule tablet backups to object storage and use Vitess’s online DDL rather than running
ALTER TABLEdirectly on shards. - Operational depth. Vitess adds many moving parts. Teams without a MySQL scale problem usually do not need it, and managed offerings exist for teams that want sharding without running the control plane.
Adoption
- Vitess was created at YouTube and served its MySQL fleet at very large scale.
- The CNCF publishes a case study from Slack, which uses Vitess for billions of queries a day.
- The project’s
ADOPTERS.mdfile lists organisations including GitHub, Shopify, HubSpot, Etsy, Square, Pinterest, Uber, New Relic, Flipkart and JD. - PlanetScale, founded by Vitess maintainers, offers Vitess as a managed database service.
Alternatives
| Solution | Model | Best suited to |
|---|---|---|
| Single MySQL with replicas (Aurora, RDS, Cloud SQL) | Vertical scaling plus read replicas | Most applications, until write volume or data size outgrows one primary |
| TiDB / TiKV (TiKV is CNCF Graduated) | Distributed SQL with MySQL compatibility | Teams wanting automatic sharding without a sharding key design |
| CockroachDB / YugabyteDB | Distributed SQL, PostgreSQL-compatible | Globally distributed PostgreSQL-style workloads |
| Citus | Sharding extension for PostgreSQL | PostgreSQL teams with multi-tenant or analytics scale |
| PlanetScale | Managed Vitess | Vitess scaling without operating the cluster |
Strengths and limitations
| Strengths | Limitations |
|---|---|
| Proven at very large scale (YouTube, Slack, GitHub, Shopify) | Significant operational complexity and many components |
| Keeps MySQL and InnoDB, so existing skills and tools apply | Not every MySQL feature or query pattern works the same way |
| Online resharding and table moves without downtime | Sharding key design requires careful upfront analysis |
| Connection pooling and query protection in VTTablet | Cross-shard transactions and joins are expensive |
| Mature Kubernetes operator and graduated CNCF status | Overkill for databases that fit comfortably on one primary |
Recommendation
Adopt Vitess when a MySQL workload has clearly outgrown a single primary, or when many MySQL databases need one operational model. Invest first in the sharding key and in testing real queries through VTGate. For smaller databases, stay on managed MySQL with replicas; for MySQL-compatible scaling without a sharding design, assess TiDB.
References: CNCF project page · Vitess documentation · Vitess Operator guide · Vitess adopters · CNCF case study: Slack