DevOps · CNCF daily

Envoy: A High-Performance Edge and Service Proxy

CNCF Graduated. A programmable L4 and L7 proxy configured through dynamic APIs, and the data plane behind Istio, Emissary-Ingress, Envoy Gateway and many other cloud native networking projects.

7 min read

CNCF Graduated

  • Envoy
  • Envoy Gateway
  • Istio
  • Kubernetes
  • Prometheus
CNCF project page

Envoy is an open-source proxy designed for modern service architectures. It sits at the edge of a platform or beside each service and handles the network: load balancing, retries, timeouts, circuit breaking, TLS, HTTP/2 and gRPC, rate limiting and detailed telemetry for every request. Its defining feature is that nearly all configuration can be changed at runtime through gRPC APIs, collectively called xDS, so a control plane can update routes and endpoints in running proxies without restarts. It addresses the problem of every service team re-implementing resilience and observability in its own language, and the poor visibility that results. Envoy was created at Lyft by Matt Klein and open-sourced in 2016. It joined the CNCF as an incubating project in September 2017 and graduated on November 28, 2018, the third project to do so after Kubernetes and Prometheus.

CNCF daily · Sprint 1, day 3 · Service Proxy

At a glance

Category Service Proxy
CNCF status Graduated. Accepted September 2017 (Incubating); graduated November 28, 2018
Written in C++
License Apache 2.0
Configuration Static YAML bootstrap, or dynamic xDS APIs served by a control plane
Extension model Filter chains (network and HTTP filters), plus WebAssembly, Lua and external-processing filters
Release cadence A new minor version each quarter; each is supported for about a year
Common control planes Istio, Envoy Gateway, Contour, Emissary-Ingress, Cilium, Gloo

Architecture

flowchart TB
  CP[Control plane<br/>Envoy Gateway, Istio, custom] -->|xDS over gRPC| ENVOY
  CLIENT[Downstream clients] --> L
  subgraph ENVOY[Envoy process]
    L[Listeners<br/>address and port] --> FC[Filter chains<br/>TLS, HTTP connection manager]
    FC --> HF[HTTP filters<br/>auth, rate limit, WASM, router]
    HF --> R[Routes<br/>virtual hosts and matches]
    R --> C[Clusters<br/>load balancing, retries, circuit breaking]
  end
  C --> EP1[Endpoint<br/>pod A]
  C --> EP2[Endpoint<br/>pod B]
  C --> EP3[Endpoint<br/>external service]
  HF -->|ext_authz, ratelimit| EXT[Auth and rate limit services]
  ENVOY --> TEL[Access logs, stats, traces<br/>Prometheus, OpenTelemetry]
Component Responsibility
Listener Binds an address and port and accepts downstream connections. Each listener has one or more filter chains selected by SNI, protocol or source.
Filter chain Ordered network filters (TLS termination, TCP proxy, the HTTP connection manager) that process connection bytes.
HTTP filters Per-request processing inside the HTTP connection manager: authentication, rate limiting, CORS, compression, WebAssembly, and finally the router.
Route Matches a request (host, path, headers) and decides which cluster receives it, with retries, timeouts and traffic splitting.
Cluster A named group of upstream endpoints with a load-balancing policy, health checks, outlier detection and circuit breakers.
xDS APIs Listener, route, cluster, endpoint and secret discovery services (LDS, RDS, CDS, EDS, SDS) that let a control plane reconfigure Envoy live.
Admin interface A local HTTP endpoint for statistics, config dumps and runtime changes.

Design principle. Envoy separates the data plane, which moves traffic, from the control plane, which decides how. The proxy exposes everything as typed APIs and keeps no state it cannot rebuild from them, so a single proxy binary can serve as an edge gateway, a sidecar, an ingress controller or an internal load balancer depending on which control plane drives it. This is why many projects chose Envoy as their data plane instead of writing their own.

Production reference design

Most teams do not write Envoy configuration by hand; they run it through a control plane. A common pattern on EKS, GKE, AKS or on-premises Kubernetes is Envoy Gateway, the Envoy project’s Kubernetes Gateway API implementation:

  1. Install Envoy Gateway with its Helm chart, selecting a current release:
    helm install eg oci://docker.io/envoyproxy/gateway-helm --version <version> \
      -n envoy-gateway-system --create-namespace
    kubectl wait --timeout=5m -n envoy-gateway-system deployment/envoy-gateway --for=condition=Available
  2. Define a GatewayClass and Gateway. The Gateway is the public entry point, and Envoy Gateway creates and manages the Envoy proxy Deployment and a load balancer Service for it:
    apiVersion: gateway.networking.k8s.io/v1
    kind: GatewayClass
    metadata: { name: eg }
    spec:
      controllerName: gateway.envoyproxy.io/gatewayclass-controller
    ---
    apiVersion: gateway.networking.k8s.io/v1
    kind: Gateway
    metadata: { name: public, namespace: edge }
    spec:
      gatewayClassName: eg
      listeners:
        - name: https
          protocol: HTTPS
          port: 443
          tls:
            mode: Terminate
            certificateRefs: [{ name: example-com-tls }]
  3. Let each team attach routes in its own namespace with an HTTPRoute:
    apiVersion: gateway.networking.k8s.io/v1
    kind: HTTPRoute
    metadata: { name: orders, namespace: orders }
    spec:
      parentRefs: [{ name: public, namespace: edge }]
      hostnames: [api.example.com]
      rules:
        - matches: [{ path: { type: PathPrefix, value: /orders } }]
          backendRefs: [{ name: orders-api, port: 8080 }]
    Check the API versions against the Gateway API and Envoy Gateway releases you install.
  4. Add policy such as rate limits, client and backend TLS, authentication and retries through Envoy Gateway’s policy resources, and issue certificates with cert-manager.
  5. Run it highly available. Use several proxy replicas across zones with a PodDisruptionBudget, and send access logs, metrics and traces to Prometheus and an OpenTelemetry collector.

Outside Kubernetes, or when learning, Envoy can run from a static bootstrap file that defines one listener, a route and a cluster, started with envoy -c envoy.yaml. Typical production uses include API and ingress gateways, the sidecar or ambient data plane in a service mesh, internal load balancing for gRPC and HTTP/2 services, and egress control.

Operational considerations

  • Choose the control plane first. Hand-written xDS is rarely worthwhile. Envoy Gateway, Istio, Contour or Emissary-Ingress each cover different needs, and the control plane determines your configuration model and upgrade path.
  • Plan upgrades quarterly. Envoy ships a minor release every quarter and supports each for about a year, so stay within the supported window and read the deprecation notes in the release history.
  • Watch memory and connections. Large route or cluster tables, many listeners and high connection counts raise memory use. Set resource requests and connection and overload limits explicitly.
  • Use the admin interface safely. The admin endpoint can change runtime settings and dump configuration, so bind it to localhost or an internal network and never expose it publicly.
  • Mind the learning curve. When something fails, debugging means reading config dumps, access logs and statistics. Enable structured access logs early and learn the response flags.

Adoption

  • The CNCF graduation announcement listed users including Airbnb, Booking.com, eBay, Google, IBM, Lyft, Microsoft, Netflix, Pinterest, Salesforce, Square, Stripe, Twilio and Verizon, and noted that Pinterest used Envoy as an edge proxy serving more than 250 million monthly unique users.
  • Envoy is the data plane for several major CNCF projects, including Istio (in sidecar mode and in its waypoint proxies), Emissary-Ingress and Contour, and for Envoy Gateway.
  • Public cloud services such as Google Cloud and AWS App Mesh also use Envoy in managed networking products.

Alternatives

Solution Model Best suited to
NGINX Web server and reverse proxy with static configuration Web serving, simple reverse proxying and caching
HAProxy Very fast TCP and HTTP load balancer High-throughput load balancing with a small footprint
Traefik Dynamic reverse proxy with automatic provider discovery Simple Kubernetes and Docker ingress with minimal setup
linkerd2-proxy Purpose-built Rust micro-proxy Linkerd service meshes that value a small proxy
Cloud load balancers (ALB, Cloud Load Balancing) Managed L4 and L7 balancing Teams that want no proxy to operate

Strengths and limitations

Strengths Limitations
Fully dynamic configuration through xDS, with no restarts Raw configuration is verbose and hard to write by hand
Rich L7 features: retries, circuit breaking, outlier detection, gRPC support Needs a control plane to be practical at scale
Detailed metrics, access logs and tracing for every request Higher memory use than lean proxies such as HAProxy or linkerd2-proxy
Large ecosystem of control planes built on it Debugging requires understanding listeners, routes, clusters and filters
Extensible with filters, WebAssembly and external services Quarterly releases and deprecations need steady upgrade effort

Recommendation

Adopt Envoy, usually through a control plane rather than directly: Envoy Gateway or Contour for ingress and API gateways, and Istio when you also need a service mesh. Use hand-written configuration only for simple cases and learning. Run it with several replicas, structured access logs and a keep-current upgrade routine, and consider HAProxy, NGINX or a managed load balancer where you only need basic proxying.

References: CNCF project page · CNCF: Envoy graduation announcement · Envoy documentation · Envoy release history · Envoy Gateway · Envoy on GitHub