DevOps · CNCF daily
NATS: A Lightweight Messaging and Streaming Fabric
CNCF Incubating. A single small server that provides publish-subscribe messaging, request-reply, persistent streams and key-value storage for services, edge devices and event-driven systems.
CNCF Incubating
- NATS
- JetStream
- Kubernetes
- Helm
- Kafka
NATS is a messaging system for connecting services, devices and applications. A single, small Go binary provides fast publish-subscribe and request-reply messaging; its built-in persistence layer, JetStream, adds durable streams, replay, key-value and object storage. NATS addresses a common problem in distributed systems: services need to talk to each other asynchronously and reliably, but heavyweight brokers are costly to run, and point-to-point HTTP calls couple services tightly. NATS began in 2010 as the internal messaging layer of Cloud Foundry, written by Derek Collison, and was later rewritten in Go. It was accepted into the CNCF at the Incubating level in March 2018. In 2025, a dispute over licensing and project ownership between Synadia, its main corporate sponsor, and the CNCF ended with NATS remaining in the CNCF under the Apache 2.0 license.
CNCF daily · Sprint 1, day 2 · Streaming & Messaging
At a glance
| Category | Streaming & Messaging |
| CNCF status | Incubating. Accepted March 15, 2018 |
| Written in | Go, with official clients for more than 40 languages |
| License | Apache 2.0 |
| Messaging patterns | Publish-subscribe, request-reply, queue groups (load-balanced consumers) |
| Persistence | JetStream: streams, consumers, key-value and object store |
| Topology | Clusters, superclusters across regions, and leaf nodes for edge sites |
Architecture
flowchart TB
P1[Order service<br/>publisher] -->|orders.created| NC
P2[IoT devices<br/>via leaf node] --> LEAF[Leaf node<br/>edge site]
LEAF --> NC
subgraph NC[NATS cluster, 3 servers]
S1[nats-server 1]
S2[nats-server 2]
S3[nats-server 3]
S1 --- S2
S2 --- S3
S3 --- S1
JS[(JetStream<br/>stream ORDERS, R3)]
S1 --> JS
S2 --> JS
S3 --> JS
end
NC -->|push or pull consumer| C1[Billing service<br/>durable consumer]
NC -->|queue group| C2[Shipping workers<br/>load-balanced]
NC -->|request-reply| C3[Pricing service<br/>responders]
NC -->|gateway| NC2[Cluster in<br/>another region]
| Component | Responsibility |
|---|---|
| nats-server | Routes messages by subject. Servers mesh into a cluster, so a client connected to any server reaches all subscribers. |
| Subjects | Hierarchical names such as orders.created.eu with wildcards (orders.*, orders.>). No topics to pre-create. |
| Core NATS | At-most-once, in-memory delivery for pub-sub, request-reply and queue groups. Very low latency. |
| JetStream | Persists subjects into streams replicated with Raft (R1, R3 or R5) and delivers them to consumers with acknowledgements and replay. |
| Key-value / object store | Built on JetStream streams: watchable configuration, leader election, and storage for large objects. |
| Leaf nodes / gateways | Leaf nodes extend a cluster to edge sites and devices; gateways connect clusters across regions into a supercluster. |
| Accounts | Multi-tenant isolation, with decentralised JWT-based authentication and explicit import and export of subjects. |
Design principle. NATS keeps the core simple: subject-based addressing, no broker-side configuration for basic messaging, and a server small enough to run on a Raspberry Pi or a 5-node cloud cluster alike. Durability is opt-in per stream through JetStream, so teams choose at-most-once speed or at-least-once and exactly-once semantics per workload rather than for the whole system.
Production reference design
A three-node JetStream cluster on EKS, GKE, AKS or on-premises Kubernetes:
- Install the official Helm chart with clustering and JetStream file storage enabled:
helm repo add nats https://nats-io.github.io/k8s/helm/charts/ helm install nats nats/nats -n messaging --create-namespace \ --set config.cluster.enabled=true --set config.cluster.replicas=3 \ --set config.jetstream.enabled=true \ --set config.jetstream.fileStore.pvc.size=50Gi - Spread the three pods across availability zones with topology spread constraints, and use fast SSD-backed persistent volumes for the JetStream file store.
- Create a replicated stream for business events, retaining seven days of orders on three replicas:
nats stream add ORDERS --subjects "orders.>" --storage file --replicas 3 \ --retention limits --max-age 7d --defaults - Add durable pull consumers per downstream service, so each one tracks its own position and can replay after outages:
nats consumer add ORDERS billing --pull --deliver all --ack explicit --defaults - Secure and isolate tenants with accounts and decentralised JWT authentication (or NKeys), enable TLS everywhere, and scrape the NATS Prometheus exporter for message rates, consumer lag and JetStream storage.
Typical production uses include event-driven microservices, command and control for IoT and edge fleets through leaf nodes, low-latency request-reply between services, and lightweight streaming where Kafka would be heavier than necessary.
Operational considerations
- Choose delivery semantics per workload. Core NATS drops messages if no subscriber is listening; use JetStream streams wherever messages must not be lost.
- Replication factor and quorum. R3 streams tolerate one server failure. Keep an odd number of JetStream servers, and avoid R1 for business-critical data.
- Watch consumer lag and storage. Unacknowledged messages and long retention grow disk use; set
max-age,max-bytesand discard policies deliberately. - Security model. Accounts and JWT-based auth are powerful but take planning. Design subject namespaces and account boundaries before onboarding many teams.
- Ecosystem size. Kafka has a larger ecosystem of connectors, stream processing and managed services. Check integration needs (for example change data capture or a data lake sink) before choosing NATS as the main event backbone.
Adoption
- The CNCF project page lists case studies from DeFacto, which migrated to an event-driven architecture on NATS, and Finleap Connect, which runs high-density bare-metal clusters in a regulated industry.
- Synadia, founded by NATS creator Derek Collison, is the main maintainer and offers managed and commercial NATS services.
- NATS is used widely as the messaging layer inside other platforms and for edge and IoT fleets. Beyond the published case studies, few adopters are publicly documented, so treat vendor logo lists with care.
Alternatives
| Solution | Model | Best suited to |
|---|---|---|
| Apache Kafka / Strimzi (Strimzi is CNCF Incubating) | Partitioned, replicated log | High-throughput event streaming with a large connector and processing ecosystem |
| RabbitMQ | Message broker with exchanges and queues | Traditional work queues and complex routing rules |
| Apache Pulsar | Segmented log with tiered storage | Multi-tenant streaming with long retention |
| Redis Streams | Streams inside an in-memory data store | Small-scale streaming where Redis is already in use |
| Cloud services (SQS/SNS, Pub/Sub, Event Hubs) | Managed messaging | Teams that prefer no messaging infrastructure to operate |
Strengths and limitations
| Strengths | Limitations |
|---|---|
| One small binary for pub-sub, request-reply, streams and key-value | Smaller connector and stream-processing ecosystem than Kafka |
| Very low latency and modest resource needs | Core NATS is at-most-once; durability requires JetStream design |
| Leaf nodes and superclusters suit edge-to-cloud topologies | Accounts and JWT security model has a learning curve |
| Simple operations compared with Kafka or Pulsar | Incubating, not yet Graduated |
| Subject-based addressing with no topics to provision | Governance dispute in 2025, since resolved, worth noting for risk reviews |
Recommendation
Adopt NATS for service-to-service messaging, request-reply and edge-to-cloud communication, and trial JetStream as the event backbone where throughput and connector needs are moderate. Keep Kafka (or Strimzi on Kubernetes) for very high-volume event streaming with heavy integration requirements. Run JetStream with R3 streams on SSD storage and design subjects and accounts up front.
References: CNCF project page · NATS documentation · NATS Helm charts · The New Stack: CNCF and Synadia reach an agreement on NATS