DevOps · CNCF project a day

Thanos: Highly Available Prometheus with Unlimited Retention

CNCF Incubating. Extends Prometheus with a global query view, object-storage retention and downsampling, without replacing it.

6 min read

CNCF Incubating

  • Thanos
  • Prometheus
  • Amazon S3
  • Grafana
  • Kubernetes
CNCF project page

Thanos is a set of components that turns independent Prometheus servers into one highly available monitoring system with effectively unlimited retention. It addresses the two limits every growing Prometheus deployment eventually reaches: data is confined to one server’s local disk, and there is no single view across clusters. Thanos was created at Improbable and open-sourced in 2018. The CNCF accepted it in July 2019 and promoted it to Incubating in August 2020. Its defining choice is to keep Prometheus unchanged and add capabilities around it, using inexpensive object storage as the long-term store.

At a glance

Category Observability: metrics storage and global query
CNCF status Incubating. Accepted July 2019, incubating since August 2020
Written in Go, released as a single binary with one subcommand per component
License Apache 2.0
Storage Object storage: Amazon S3, Google Cloud Storage, Azure Blob, Swift, MinIO and others
Query interface PromQL through the Prometheus HTTP API, so Grafana works unchanged
Relationship to Prometheus Complements it; every Prometheus server keeps scraping and alerting locally

Architecture

flowchart TB
  subgraph C1[Cluster A]
    P1[Prometheus] --- S1[Thanos Sidecar]
  end
  subgraph C2[Cluster B]
    P2[Prometheus] --- S2[Thanos Sidecar]
  end
  S1 -->|upload 2h blocks| OBJ[(Object storage<br/>S3 or GCS)]
  S2 -->|upload 2h blocks| OBJ
  OBJ --> SG[Store Gateway]
  OBJ <--> CMP[Compactor<br/>compact and downsample]
  Q[Querier<br/>global PromQL view] --> S1
  Q --> S2
  Q --> SG
  QF[Query Frontend<br/>split and cache] --> Q
  G[Grafana] --> QF
  R[Ruler<br/>global rules] --> Q
  R --> AM[Alertmanager]
Component Responsibility
Sidecar Runs next to each Prometheus, serves its recent data to the Querier, and uploads completed two-hour TSDB blocks to object storage.
Store Gateway Serves historical blocks from object storage over the same Store API, using indexes and caches to avoid full downloads.
Querier Fans a PromQL query out to all sidecars and store gateways, merges the results and deduplicates data from HA replica pairs.
Query Frontend Splits long-range queries by day, retries failures and caches results to keep dashboards fast.
Compactor Compacts blocks in object storage, applies retention, and creates 5-minute and 1-hour downsampled data for long-range queries. Runs as a singleton.
Ruler Evaluates recording and alerting rules against the global view, for rules that span clusters.
Receiver (optional) Accepts Prometheus remote_write for push-based setups where a sidecar cannot be deployed.

Design principle. All components speak a common gRPC Store API, so recent data from sidecars and historical data from object storage look the same to the Querier. Durable state lives only in object storage, which is cheap, highly durable and managed by the cloud provider. Most components are stateless and can be scaled or replaced freely.

Production reference design

A multi-cluster setup on Amazon EKS with long-term metrics in S3 (the same pattern applies to GKE with GCS or AKS with Azure Blob):

  1. Create the bucket and credentials. Create an S3 bucket such as metrics-longterm and grant write access to the sidecar’s service account through IRSA (IAM roles for service accounts). Describe the bucket in an object-storage config and store it as a Kubernetes secret:
    # objstore.yml  ->  kubectl -n monitoring create secret generic thanos-objstore --from-file=objstore.yml
    type: S3
    config:
      bucket: metrics-longterm
      endpoint: s3.ap-south-1.amazonaws.com
  2. Enable the sidecar in each cluster’s kube-prometheus-stack release. Give every cluster a unique external label, and run two replicas for HA:
    prometheus:
      prometheusSpec:
        replicas: 2
        externalLabels:
          cluster: prod-ap-south-1
        thanos:
          objectStorageConfig:
            existingSecret:
              name: thanos-objstore
              key: objstore.yml
  3. Deploy the central components (Querier, Query Frontend, Store Gateway, Compactor) in a management cluster using the official kube-thanos manifests or a community Helm chart. Point the Querier at each sidecar and store gateway, and set --query.replica-label=prometheus_replica so HA pairs are deduplicated.
  4. Repoint Grafana to the Query Frontend as a standard Prometheus data source. Dashboards and PromQL stay the same.
  5. Set retention tiers on the Compactor, for example 30 days raw, 180 days at 5-minute resolution and 2 years at 1-hour resolution. Keep local Prometheus retention short (24–48 hours).

Typical production uses include a single pane of glass across environments and regions, year-over-year capacity planning, SLO reporting over quarters, and keeping metrics through cluster rebuilds and migrations.

Operational considerations

  • Unique external labels are mandatory. Every Prometheus must have distinct externalLabels, or the Compactor may merge or reject overlapping blocks.
  • Run exactly one Compactor per bucket. Concurrent compactors corrupt data. Give it generous local disk, because it downloads blocks to compact them.
  • Cache aggressively. Store Gateway index and chunk caches, and the Query Frontend results cache (memcached or Redis), have the biggest effect on query latency and object-storage request costs.
  • Bound expensive queries. Long-range, high-cardinality queries can exhaust Querier memory. Use the Query Frontend’s splitting and per-query limits.
  • Sidecar versus Receiver. Prefer the sidecar, which is simpler and keeps Prometheus autonomous. Use Receive only where outbound push is required, such as edge sites or untrusted networks, and accept its extra operational weight.

Adoption

  • CNCF case studies: Meltwater, which operates at multi-petabyte scale, and La Redoute.
  • The project’s adopter list includes Adobe, eBay, ByteDance, Tencent, Hotstar, Blinkit, Darwinbox, CarTrade Tech, Banco do Brasil, Itaú Unibanco, Monzo, Wise, SoundCloud, Hetzner and Aiven, among more than 100 organisations.
  • It is widely used as the long-term backend for Prometheus Operator and OpenShift monitoring, and the OpenShift monitoring stack ships a Thanos Querier.

Alternatives

Solution Model Best suited to
Grafana Mimir Remote-write, horizontally scalable TSDB Very large multi-tenant platforms that want push ingestion and strict tenancy
Cortex (CNCF Incubating) Remote-write, multi-tenant Existing Cortex users; Mimir is its more actively developed fork
VictoriaMetrics Remote-write, single-node or cluster Lower resource use and simpler operations at high cardinality
Amazon / Google / Azure managed Prometheus Fully managed remote-write Teams that prefer to pay per sample rather than run storage
M3 Distributed TSDB Very high-throughput environments already invested in it

Strengths and limitations

Strengths Limitations
Keeps Prometheus as-is; incremental adoption, one cluster at a time Many moving parts: sidecars, gateways, compactor, caches
Object storage gives cheap, durable, effectively unlimited retention Query latency over old data depends heavily on caching and downsampling
Global PromQL view with built-in HA deduplication Compactor is a singleton and a common source of operational issues
Downsampling makes multi-year dashboards practical Weaker multi-tenancy than Mimir or Cortex
Large production user base and active community Object-storage request costs need monitoring at scale

Recommendation

Adopt Thanos when Prometheus outgrows a single cluster or a few weeks of retention, especially if you already run kube-prometheus-stack. Start with the sidecar, Querier and Store Gateway, then add the Compactor, Query Frontend and caches. Organisations that need hard multi-tenancy or want push-only ingestion should evaluate Grafana Mimir or a managed Prometheus service alongside it.

References: CNCF project page · Thanos documentation · Thanos adopters · kube-thanos manifests · CNCF case studies