DevOps · CNCF project a day
Prometheus: The Standard for Cloud-Native Metrics and Alerting
CNCF Graduated. A pull-based monitoring system and time-series database that has become the de facto metrics layer for Kubernetes.
CNCF Graduated
- Prometheus
- Alertmanager
- Grafana
- Thanos
- Kubernetes
Prometheus is an open-source monitoring system that collects numeric metrics from services and infrastructure, stores them in a purpose-built time-series database, and evaluates alerting rules against them in real time. Originally developed at SoundCloud in 2012 and modelled on Google’s internal Borgmon, it became the second project accepted by the CNCF after Kubernetes and the second to graduate. Today it is the metrics foundation of most Kubernetes platforms, both self-managed and cloud-managed.
At a glance
| Category | Observability: metrics and alerting |
| CNCF status | Graduated. Accepted May 2016, graduated August 2018 |
| Written in | Go, released as a single static binary |
| License | Apache 2.0 |
| Data model | Time series identified by a metric name and key-value labels |
| Query language | PromQL |
| Collection model | Pull (HTTP scrape), with Pushgateway for short-lived jobs |
Architecture
flowchart TB
SD[Service discovery<br/>Kubernetes API, EC2, Consul]
subgraph Targets[Monitored targets]
APP[Application<br/>/metrics endpoint]
EXP[Exporters<br/>node, MySQL, Kafka]
PGW[Pushgateway<br/>batch jobs]
end
subgraph PROM[Prometheus server]
SCR[Scrape manager]
TSDB[(Time-series DB)]
ENG[PromQL engine]
RULES[Rule evaluator]
end
SD -. target list .-> SCR
APP --> SCR
EXP --> SCR
PGW --> SCR
SCR --> TSDB
TSDB --> ENG
ENG --> RULES
RULES -->|alerts| AM[Alertmanager]
AM --> NOTIFY[PagerDuty, Slack, Email]
ENG --> GRAF[Grafana]
TSDB -. remote write .-> LTS[(Long-term store<br/>Thanos, Mimir)]
| Component | Responsibility |
|---|---|
| Scrape manager | Polls each target’s /metrics endpoint on a fixed interval (typically 15–30 s) and attaches target labels. |
| Service discovery | Keeps the target list current as pods, nodes and instances appear and disappear. |
| TSDB | Stores samples on local disk in two-hour blocks, compacted over time. Retention is commonly 15–30 days. |
| PromQL engine | Serves ad-hoc and dashboard queries, including rates, aggregations, percentiles and linear predictions. |
| Rule evaluator | Runs recording rules (pre-computed queries) and alerting rules on a schedule. |
| Alertmanager | Deduplicates, groups, silences and routes alerts to receivers. It runs as a separate, clusterable process. |
| Exporters | Translate third-party systems (Linux hosts, databases, brokers, load balancers) into Prometheus metrics. |
Design principle. The server pulls metrics instead of receiving them. Target health becomes observable in its own right (a failed scrape raises up == 0), monitored services stay free of client-side buffering and retry logic, and scraping can be controlled centrally. Each Prometheus server is deliberately autonomous and does not depend on network storage, so monitoring keeps working when the rest of the platform is degraded.
Production reference design
A common deployment for a microservices platform on Amazon EKS (the pattern applies equally to GKE, AKS or an on-premises k3d/k3s cluster):
- Install the
kube-prometheus-stackHelm chart, which deploys the Prometheus Operator, Prometheus, Alertmanager, Grafana, node-exporter and kube-state-metrics with curated dashboards and alerts:helm repo add prometheus-community https://prometheus-community.github.io/helm-charts helm install monitoring prometheus-community/kube-prometheus-stack -n monitoring --create-namespace - Instrument each service with a client library (for example
prom-clientfor Node.js orprometheus_clientfor Python). Then declare aServiceMonitorso the Operator configures the scrape automatically. No Prometheus config files need editing. - Alert on symptoms, not causes. Define alerts against user-facing service levels:
groups: - name: service-slo rules: - alert: HighErrorRate expr: | sum by (service) (rate(http_requests_total{status=~"5.."}[5m])) / sum by (service) (rate(http_requests_total[5m])) > 0.05 for: 10m labels: severity: critical annotations: summary: "{{ $labels.service }}: more than 5% of requests failing" - Route alerts by severity and team in Alertmanager: critical alerts page the on-call engineer, warnings go to a team channel, and inhibition rules suppress downstream noise during a known outage.
- Scale out when one cluster becomes many. Run Prometheus per cluster, and add a Thanos sidecar or
remote_writeto Grafana Mimir or Amazon Managed Service for Prometheus for a global query view and multi-month retention.
Typical signals in production include the RED metrics (rate, errors, duration) for every service, pod CPU and memory against requests and limits, node saturation, disk-fill forecasts using predict_linear, certificate expiry, consumer lag on Kafka and RabbitMQ, and error-budget burn rates for SLOs.
Operational considerations
- Cardinality is the primary cost driver. Every unique label combination is a separate series. Avoid unbounded labels such as user IDs, request IDs or full URLs; they are the most common cause of memory exhaustion.
- High availability is achieved by running two identical replicas scraping the same targets. Alertmanager deduplicates their alerts, and Thanos or Mimir deduplicates queries.
- Retention and durability. Local storage is not replicated. Treat it as short-term and use object-storage-backed systems for history and disaster recovery.
- Sizing. As a rough guide, budget a few kilobytes of memory per active series. Recording rules keep expensive dashboard queries fast.
- Security. The
/metricsendpoints and the Prometheus UI have no authentication by default. Restrict them with network policies or put them behind an authenticating proxy.
Adoption
- SoundCloud created Prometheus to monitor its microservices platform.
- DigitalOcean, Red Hat, Ericsson, CoreOS, Weaveworks and Google were cited as production users when the CNCF accepted the project. Red Hat OpenShift’s built-in cluster monitoring is based on Prometheus.
- Cloudflare reported running 188 Prometheus servers across 116 data centres, using it to reduce alert fatigue at global scale.
- Northern Trust monitors more than 750 microservices with Prometheus.
- HDFC Bank and China Merchants Bank are featured in CNCF case studies.
- All three major clouds offer it as a managed service: Amazon Managed Service for Prometheus, Google Cloud Managed Service for Prometheus and Azure Monitor managed service for Prometheus.
Alternatives
| Solution | Model | Best suited to |
|---|---|---|
| Grafana Mimir / Thanos / Cortex | Extends Prometheus | Horizontally scalable, long-term, multi-cluster storage that keeps PromQL |
| VictoriaMetrics | Prometheus-compatible | Lower memory and disk footprint at high cardinality |
| Datadog / New Relic / Dynatrace | Commercial SaaS | Teams that want metrics, logs, traces and APM in one place without operating it |
| InfluxDB + Telegraf | Push-based TSDB | IoT, event and sensor data with irregular intervals |
| Zabbix / Nagios | Agent and check-based | Traditional host, network and device monitoring in static environments |
| Amazon CloudWatch | Cloud-native service | AWS resource metrics with no infrastructure to run |
Strengths and limitations
| Strengths | Limitations |
|---|---|
| Industry standard for Kubernetes, with first-class Operator support | Single-node storage; HA and long-term retention need extra components |
| PromQL is expressive for rates, ratios, percentiles and forecasting | PromQL has a learning curve for teams new to time-series queries |
| Service discovery means new workloads are monitored automatically | High-cardinality labels can quickly exhaust memory |
| Very large exporter ecosystem and native instrumentation in most CNCF projects | Metrics only: logs and traces need Loki, Jaeger or OpenTelemetry |
| Simple to operate: one binary, no external dependencies | Sampled data, so unsuitable for exact per-event accounting such as billing |
Recommendation
Adopt Prometheus as the default metrics and alerting layer for any containerised or Kubernetes-based platform. Begin with kube-prometheus-stack and SLO-based alerts. Plan for Thanos, Mimir or a managed Prometheus service from the outset if you expect multiple clusters or retention beyond about a month. Pair it with OpenTelemetry for traces and Loki for logs to complete the observability stack.
References: CNCF project page · Prometheus documentation · CNCF acceptance announcement (2016) · Opensource.com: Prometheus in production