Envoy is an open-source proxy designed for modern service architectures. It sits at the edge of a platform or beside each service and handles the network: load balancing, retries, timeouts, circuit breaking, TLS, HTTP/2 and gRPC, rate limiting and detailed telemetry for every request. Its defining feature is that nearly all configuration can be changed at runtime through gRPC APIs, collectively called xDS, so a control plane can update routes and endpoints in running proxies without restarts. It addresses the problem of every service team re-implementing resilience and observability in its own language, and the poor visibility that results. Envoy was created at Lyft by Matt Klein and open-sourced in 2016. It joined the CNCF as an incubating project in September 2017 and graduated on November 28, 2018, the third project to do so after Kubernetes and Prometheus.
CNCF daily · Sprint 1, day 3 · Service Proxy
At a glance
| Category | Service Proxy |
| CNCF status | Graduated. Accepted September 2017 (Incubating); graduated November 28, 2018 |
| Written in | C++ |
| License | Apache 2.0 |
| Configuration | Static YAML bootstrap, or dynamic xDS APIs served by a control plane |
| Extension model | Filter chains (network and HTTP filters), plus WebAssembly, Lua and external-processing filters |
| Release cadence | A new minor version each quarter; each is supported for about a year |
| Common control planes | Istio, Envoy Gateway, Contour, Emissary-Ingress, Cilium, Gloo |
Architecture
flowchart TB
CP[Control plane<br/>Envoy Gateway, Istio, custom] -->|xDS over gRPC| ENVOY
CLIENT[Downstream clients] --> L
subgraph ENVOY[Envoy process]
L[Listeners<br/>address and port] --> FC[Filter chains<br/>TLS, HTTP connection manager]
FC --> HF[HTTP filters<br/>auth, rate limit, WASM, router]
HF --> R[Routes<br/>virtual hosts and matches]
R --> C[Clusters<br/>load balancing, retries, circuit breaking]
end
C --> EP1[Endpoint<br/>pod A]
C --> EP2[Endpoint<br/>pod B]
C --> EP3[Endpoint<br/>external service]
HF -->|ext_authz, ratelimit| EXT[Auth and rate limit services]
ENVOY --> TEL[Access logs, stats, traces<br/>Prometheus, OpenTelemetry]
| Component | Responsibility |
|---|---|
| Listener | Binds an address and port and accepts downstream connections. Each listener has one or more filter chains selected by SNI, protocol or source. |
| Filter chain | Ordered network filters (TLS termination, TCP proxy, the HTTP connection manager) that process connection bytes. |
| HTTP filters | Per-request processing inside the HTTP connection manager: authentication, rate limiting, CORS, compression, WebAssembly, and finally the router. |
| Route | Matches a request (host, path, headers) and decides which cluster receives it, with retries, timeouts and traffic splitting. |
| Cluster | A named group of upstream endpoints with a load-balancing policy, health checks, outlier detection and circuit breakers. |
| xDS APIs | Listener, route, cluster, endpoint and secret discovery services (LDS, RDS, CDS, EDS, SDS) that let a control plane reconfigure Envoy live. |
| Admin interface | A local HTTP endpoint for statistics, config dumps and runtime changes. |
Design principle. Envoy separates the data plane, which moves traffic, from the control plane, which decides how. The proxy exposes everything as typed APIs and keeps no state it cannot rebuild from them, so a single proxy binary can serve as an edge gateway, a sidecar, an ingress controller or an internal load balancer depending on which control plane drives it. This is why many projects chose Envoy as their data plane instead of writing their own.
Production reference design
Most teams do not write Envoy configuration by hand; they run it through a control plane. A common pattern on EKS, GKE, AKS or on-premises Kubernetes is Envoy Gateway, the Envoy project’s Kubernetes Gateway API implementation:
- Install Envoy Gateway with its Helm chart, selecting a current release:
helm install eg oci://docker.io/envoyproxy/gateway-helm --version <version> \ -n envoy-gateway-system --create-namespace kubectl wait --timeout=5m -n envoy-gateway-system deployment/envoy-gateway --for=condition=Available - Define a GatewayClass and Gateway. The Gateway is the public entry point, and Envoy Gateway creates and manages the Envoy proxy Deployment and a load balancer Service for it:
apiVersion: gateway.networking.k8s.io/v1 kind: GatewayClass metadata: { name: eg } spec: controllerName: gateway.envoyproxy.io/gatewayclass-controller --- apiVersion: gateway.networking.k8s.io/v1 kind: Gateway metadata: { name: public, namespace: edge } spec: gatewayClassName: eg listeners: - name: https protocol: HTTPS port: 443 tls: mode: Terminate certificateRefs: [{ name: example-com-tls }] - Let each team attach routes in its own namespace with an
HTTPRoute:
Check the API versions against the Gateway API and Envoy Gateway releases you install.apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: { name: orders, namespace: orders } spec: parentRefs: [{ name: public, namespace: edge }] hostnames: [api.example.com] rules: - matches: [{ path: { type: PathPrefix, value: /orders } }] backendRefs: [{ name: orders-api, port: 8080 }] - Add policy such as rate limits, client and backend TLS, authentication and retries through Envoy Gateway’s policy resources, and issue certificates with cert-manager.
- Run it highly available. Use several proxy replicas across zones with a PodDisruptionBudget, and send access logs, metrics and traces to Prometheus and an OpenTelemetry collector.
Outside Kubernetes, or when learning, Envoy can run from a static bootstrap file that defines one listener, a route and a cluster, started with envoy -c envoy.yaml. Typical production uses include API and ingress gateways, the sidecar or ambient data plane in a service mesh, internal load balancing for gRPC and HTTP/2 services, and egress control.
Operational considerations
- Choose the control plane first. Hand-written xDS is rarely worthwhile. Envoy Gateway, Istio, Contour or Emissary-Ingress each cover different needs, and the control plane determines your configuration model and upgrade path.
- Plan upgrades quarterly. Envoy ships a minor release every quarter and supports each for about a year, so stay within the supported window and read the deprecation notes in the release history.
- Watch memory and connections. Large route or cluster tables, many listeners and high connection counts raise memory use. Set resource requests and connection and overload limits explicitly.
- Use the admin interface safely. The admin endpoint can change runtime settings and dump configuration, so bind it to localhost or an internal network and never expose it publicly.
- Mind the learning curve. When something fails, debugging means reading config dumps, access logs and statistics. Enable structured access logs early and learn the response flags.
Adoption
- The CNCF graduation announcement listed users including Airbnb, Booking.com, eBay, Google, IBM, Lyft, Microsoft, Netflix, Pinterest, Salesforce, Square, Stripe, Twilio and Verizon, and noted that Pinterest used Envoy as an edge proxy serving more than 250 million monthly unique users.
- Envoy is the data plane for several major CNCF projects, including Istio (in sidecar mode and in its waypoint proxies), Emissary-Ingress and Contour, and for Envoy Gateway.
- Public cloud services such as Google Cloud and AWS App Mesh also use Envoy in managed networking products.
Alternatives
| Solution | Model | Best suited to |
|---|---|---|
| NGINX | Web server and reverse proxy with static configuration | Web serving, simple reverse proxying and caching |
| HAProxy | Very fast TCP and HTTP load balancer | High-throughput load balancing with a small footprint |
| Traefik | Dynamic reverse proxy with automatic provider discovery | Simple Kubernetes and Docker ingress with minimal setup |
| linkerd2-proxy | Purpose-built Rust micro-proxy | Linkerd service meshes that value a small proxy |
| Cloud load balancers (ALB, Cloud Load Balancing) | Managed L4 and L7 balancing | Teams that want no proxy to operate |
Strengths and limitations
| Strengths | Limitations |
|---|---|
| Fully dynamic configuration through xDS, with no restarts | Raw configuration is verbose and hard to write by hand |
| Rich L7 features: retries, circuit breaking, outlier detection, gRPC support | Needs a control plane to be practical at scale |
| Detailed metrics, access logs and tracing for every request | Higher memory use than lean proxies such as HAProxy or linkerd2-proxy |
| Large ecosystem of control planes built on it | Debugging requires understanding listeners, routes, clusters and filters |
| Extensible with filters, WebAssembly and external services | Quarterly releases and deprecations need steady upgrade effort |
Recommendation
Adopt Envoy, usually through a control plane rather than directly: Envoy Gateway or Contour for ingress and API gateways, and Istio when you also need a service mesh. Use hand-written configuration only for simple cases and learning. Run it with several replicas, structured access logs and a keep-current upgrade routine, and consider HAProxy, NGINX or a managed load balancer where you only need basic proxying.
References: CNCF project page · CNCF: Envoy graduation announcement · Envoy documentation · Envoy release history · Envoy Gateway · Envoy on GitHub