DevOps · CNCF daily

NATS: A Lightweight Messaging and Streaming Fabric

CNCF Incubating. A single small server that provides publish-subscribe messaging, request-reply, persistent streams and key-value storage for services, edge devices and event-driven systems.

6 min read

CNCF Incubating

  • NATS
  • JetStream
  • Kubernetes
  • Helm
  • Kafka
CNCF project page

NATS is a messaging system for connecting services, devices and applications. A single, small Go binary provides fast publish-subscribe and request-reply messaging; its built-in persistence layer, JetStream, adds durable streams, replay, key-value and object storage. NATS addresses a common problem in distributed systems: services need to talk to each other asynchronously and reliably, but heavyweight brokers are costly to run, and point-to-point HTTP calls couple services tightly. NATS began in 2010 as the internal messaging layer of Cloud Foundry, written by Derek Collison, and was later rewritten in Go. It was accepted into the CNCF at the Incubating level in March 2018. In 2025, a dispute over licensing and project ownership between Synadia, its main corporate sponsor, and the CNCF ended with NATS remaining in the CNCF under the Apache 2.0 license.

CNCF daily · Sprint 1, day 2 · Streaming & Messaging

At a glance

Category Streaming & Messaging
CNCF status Incubating. Accepted March 15, 2018
Written in Go, with official clients for more than 40 languages
License Apache 2.0
Messaging patterns Publish-subscribe, request-reply, queue groups (load-balanced consumers)
Persistence JetStream: streams, consumers, key-value and object store
Topology Clusters, superclusters across regions, and leaf nodes for edge sites

Architecture

flowchart TB
  P1[Order service<br/>publisher] -->|orders.created| NC
  P2[IoT devices<br/>via leaf node] --> LEAF[Leaf node<br/>edge site]
  LEAF --> NC
  subgraph NC[NATS cluster, 3 servers]
    S1[nats-server 1]
    S2[nats-server 2]
    S3[nats-server 3]
    S1 --- S2
    S2 --- S3
    S3 --- S1
    JS[(JetStream<br/>stream ORDERS, R3)]
    S1 --> JS
    S2 --> JS
    S3 --> JS
  end
  NC -->|push or pull consumer| C1[Billing service<br/>durable consumer]
  NC -->|queue group| C2[Shipping workers<br/>load-balanced]
  NC -->|request-reply| C3[Pricing service<br/>responders]
  NC -->|gateway| NC2[Cluster in<br/>another region]
Component Responsibility
nats-server Routes messages by subject. Servers mesh into a cluster, so a client connected to any server reaches all subscribers.
Subjects Hierarchical names such as orders.created.eu with wildcards (orders.*, orders.>). No topics to pre-create.
Core NATS At-most-once, in-memory delivery for pub-sub, request-reply and queue groups. Very low latency.
JetStream Persists subjects into streams replicated with Raft (R1, R3 or R5) and delivers them to consumers with acknowledgements and replay.
Key-value / object store Built on JetStream streams: watchable configuration, leader election, and storage for large objects.
Leaf nodes / gateways Leaf nodes extend a cluster to edge sites and devices; gateways connect clusters across regions into a supercluster.
Accounts Multi-tenant isolation, with decentralised JWT-based authentication and explicit import and export of subjects.

Design principle. NATS keeps the core simple: subject-based addressing, no broker-side configuration for basic messaging, and a server small enough to run on a Raspberry Pi or a 5-node cloud cluster alike. Durability is opt-in per stream through JetStream, so teams choose at-most-once speed or at-least-once and exactly-once semantics per workload rather than for the whole system.

Production reference design

A three-node JetStream cluster on EKS, GKE, AKS or on-premises Kubernetes:

  1. Install the official Helm chart with clustering and JetStream file storage enabled:
    helm repo add nats https://nats-io.github.io/k8s/helm/charts/
    helm install nats nats/nats -n messaging --create-namespace \
      --set config.cluster.enabled=true --set config.cluster.replicas=3 \
      --set config.jetstream.enabled=true \
      --set config.jetstream.fileStore.pvc.size=50Gi
  2. Spread the three pods across availability zones with topology spread constraints, and use fast SSD-backed persistent volumes for the JetStream file store.
  3. Create a replicated stream for business events, retaining seven days of orders on three replicas:
    nats stream add ORDERS --subjects "orders.>" --storage file --replicas 3 \
      --retention limits --max-age 7d --defaults
  4. Add durable pull consumers per downstream service, so each one tracks its own position and can replay after outages:
    nats consumer add ORDERS billing --pull --deliver all --ack explicit --defaults
  5. Secure and isolate tenants with accounts and decentralised JWT authentication (or NKeys), enable TLS everywhere, and scrape the NATS Prometheus exporter for message rates, consumer lag and JetStream storage.

Typical production uses include event-driven microservices, command and control for IoT and edge fleets through leaf nodes, low-latency request-reply between services, and lightweight streaming where Kafka would be heavier than necessary.

Operational considerations

  • Choose delivery semantics per workload. Core NATS drops messages if no subscriber is listening; use JetStream streams wherever messages must not be lost.
  • Replication factor and quorum. R3 streams tolerate one server failure. Keep an odd number of JetStream servers, and avoid R1 for business-critical data.
  • Watch consumer lag and storage. Unacknowledged messages and long retention grow disk use; set max-age, max-bytes and discard policies deliberately.
  • Security model. Accounts and JWT-based auth are powerful but take planning. Design subject namespaces and account boundaries before onboarding many teams.
  • Ecosystem size. Kafka has a larger ecosystem of connectors, stream processing and managed services. Check integration needs (for example change data capture or a data lake sink) before choosing NATS as the main event backbone.

Adoption

  • The CNCF project page lists case studies from DeFacto, which migrated to an event-driven architecture on NATS, and Finleap Connect, which runs high-density bare-metal clusters in a regulated industry.
  • Synadia, founded by NATS creator Derek Collison, is the main maintainer and offers managed and commercial NATS services.
  • NATS is used widely as the messaging layer inside other platforms and for edge and IoT fleets. Beyond the published case studies, few adopters are publicly documented, so treat vendor logo lists with care.

Alternatives

Solution Model Best suited to
Apache Kafka / Strimzi (Strimzi is CNCF Incubating) Partitioned, replicated log High-throughput event streaming with a large connector and processing ecosystem
RabbitMQ Message broker with exchanges and queues Traditional work queues and complex routing rules
Apache Pulsar Segmented log with tiered storage Multi-tenant streaming with long retention
Redis Streams Streams inside an in-memory data store Small-scale streaming where Redis is already in use
Cloud services (SQS/SNS, Pub/Sub, Event Hubs) Managed messaging Teams that prefer no messaging infrastructure to operate

Strengths and limitations

Strengths Limitations
One small binary for pub-sub, request-reply, streams and key-value Smaller connector and stream-processing ecosystem than Kafka
Very low latency and modest resource needs Core NATS is at-most-once; durability requires JetStream design
Leaf nodes and superclusters suit edge-to-cloud topologies Accounts and JWT security model has a learning curve
Simple operations compared with Kafka or Pulsar Incubating, not yet Graduated
Subject-based addressing with no topics to provision Governance dispute in 2025, since resolved, worth noting for risk reviews

Recommendation

Adopt NATS for service-to-service messaging, request-reply and edge-to-cloud communication, and trial JetStream as the event backbone where throughput and connector needs are moderate. Keep Kafka (or Strimzi on Kubernetes) for very high-volume event streaming with heavy integration requirements. Run JetStream with R3 streams on SSD storage and design subjects and accounts up front.

References: CNCF project page · NATS documentation · NATS Helm charts · The New Stack: CNCF and Synadia reach an agreement on NATS