DevOps · CNCF daily

OpenTelemetry: A Vendor-Neutral Standard for Telemetry

CNCF Graduated. The open standard, APIs, SDKs and Collector for generating, processing and exporting traces, metrics and logs, so applications are instrumented once and can send data to any backend.

8 min read

CNCF Graduated

  • OpenTelemetry
  • OTel Collector
  • Prometheus
  • Grafana
  • Kubernetes
CNCF project page

OpenTelemetry, often written OTel, is an open standard and toolkit for producing telemetry from software. It defines language APIs and SDKs that applications use to emit traces, metrics and logs, a wire format called the OpenTelemetry Protocol (OTLP), semantic conventions that name attributes consistently, and the Collector, a vendor-neutral service that receives, processes and exports telemetry to one or more backends. It addresses vendor lock-in and fragmented instrumentation: code is instrumented once against a standard, and the choice of storage and analysis tool can change without touching the application. OpenTelemetry was formed in May 2019 by the merger of the OpenTracing and OpenCensus projects, entered the CNCF Sandbox soon afterwards, was accepted as Incubating in August 2021, and reached Graduated status in May 2026. Its contributor activity is reported by the CNCF as second only to Kubernetes among CNCF projects.

CNCF daily · Sprint 1, day 6 · Observability

At a glance

Category Observability
CNCF status Graduated. Formed May 2019; Incubating August 2021; Graduated May 2026
Written in Many languages: each SDK uses its own, and the Collector is written in Go
License Apache 2.0
Signals Traces, metrics and logs are stable in the major language SDKs; profiling is the newer fourth signal
Core pieces Specification, OTLP, semantic conventions, language APIs and SDKs, Collector, Kubernetes operator
Instrumentation Manual APIs, automatic and zero-code agents for common languages, and library instrumentations
Backends Any OTLP-capable system, including Prometheus, Grafana, Jaeger, cloud services and commercial platforms
Governance Vendor-neutral, with maintainers organised into language and component special interest groups

Architecture

flowchart LR
  subgraph APPS[Applications]
    SDK1[Service A<br/>OTel SDK]
    SDK2[Service B<br/>auto instrumentation]
  end
  SDK1 -->|OTLP| AGENT
  SDK2 -->|OTLP| AGENT
  subgraph COL[OpenTelemetry Collector]
    AGENT[Receivers<br/>OTLP, Prometheus, Jaeger] --> PROC[Processors<br/>batch, memory limiter, sampling, attributes]
    PROC --> EXP[Exporters]
  end
  INFRA[Hosts and Kubernetes<br/>logs and metrics] --> AGENT
  EXP -->|traces| TR[Trace backend<br/>Tempo or Jaeger]
  EXP -->|metrics| MET[Metrics backend<br/>Prometheus]
  EXP -->|logs| LOG[Log backend<br/>Loki or other]
  TR --> UI[Dashboards and alerts<br/>Grafana]
  MET --> UI
  LOG --> UI
Component Responsibility
API and SDK The API is what application and library code calls. The SDK implements it, adds sampling and batching, and exports data.
Semantic conventions Standard attribute names for HTTP, databases, messaging and cloud resources, so data from different services lines up.
OTLP The protocol that carries traces, metrics and logs over gRPC or HTTP between SDKs, Collectors and backends.
Collector A configurable pipeline of receivers, processors and exporters. It can run beside each application, on each node or as a central gateway.
Auto instrumentation Agents and injectors that instrument common frameworks without code changes.
Resource attributes Labels such as service.name, deployment.environment and Kubernetes namespace that identify where data came from.
Kubernetes operator Manages Collectors and auto-instrumentation injection on Kubernetes.

Design principle. OpenTelemetry separates generating telemetry from storing and analysing it. Applications speak one protocol and a standard vocabulary, and the Collector sits between them and the backends as a place to enrich, filter, sample and route data. Moving to a different backend, or sending the same data to two, becomes a Collector configuration change instead of re-instrumentation.

Production reference design

A common pattern on EKS, GKE, AKS or on-premises Kubernetes, with open-source backends and OTLP as the standard between all layers:

  1. Define conventions first. Agree on service.name, environment, version and team attributes for every service, and make them part of the definition of done for a new service. Consistent resource attributes make cross-service queries and cost allocation possible.
  2. Deploy the Collector in two tiers. Run an agent Collector on each node as a DaemonSet for local collection and host metrics, and a gateway Collector deployment for central processing, sampling and export. Use the Kubernetes operator where it fits:
    apiVersion: opentelemetry.io/v1beta1
    kind: OpenTelemetryCollector
    metadata: { name: gateway, namespace: observability }
    spec:
      mode: deployment
      replicas: 2
      config:
        receivers:
          otlp:
            protocols: { grpc: {}, http: {} }
        processors:
          memory_limiter: { check_interval: 1s, limit_percentage: 75 }
          batch: {}
        exporters:
          otlp/tempo: { endpoint: tempo.observability:4317, tls: { insecure: true } }
          prometheusremotewrite: { endpoint: http://prometheus.observability:9090/api/v1/write }
        service:
          pipelines:
            traces:  { receivers: [otlp], processors: [memory_limiter, batch], exporters: [otlp/tempo] }
            metrics: { receivers: [otlp], processors: [memory_limiter, batch], exporters: [prometheusremotewrite] }
    Check field names and component availability against the operator and Collector versions you install, and use TLS between tiers in production.
  3. Instrument the applications. Start with zero-code instrumentation for common frameworks, then add manual spans and metrics where business operations need them. On Kubernetes, annotate a workload so the operator injects the agent:
    metadata:
      annotations:
        instrumentation.opentelemetry.io/inject-java: "observability/default"
  4. Point SDKs at the Collector. Set the standard environment variables, for example OTEL_EXPORTER_OTLP_ENDPOINT, OTEL_SERVICE_NAME and OTEL_RESOURCE_ATTRIBUTES, rather than hard-coding backends.
  5. Decide on sampling early. Head sampling in the SDK is cheap and simple. Tail sampling in a gateway keeps errors and slow traces but needs more infrastructure, and it shapes how the Collector tier is scaled.
  6. Connect the backends. Send traces to Tempo or Jaeger, metrics to Prometheus or a compatible store, and logs to Loki or another system, and join them in Grafana through trace IDs.
  7. Operate the pipeline. Monitor the Collector’s own metrics, set a memory limiter and resource limits from the start, run the gateway highly available, and review telemetry volume and cardinality regularly.

Typical production uses include distributed tracing across microservices, a standard metrics and logs pipeline, backend migration without re-instrumentation, correlation of logs with traces, and telemetry for AI and agent workloads through emerging semantic conventions.

Operational considerations

  • Telemetry cost is a design input. High-cardinality attributes and unsampled traces drive storage cost. Use processors to drop or redact attributes, set sampling policy, and watch ingest volume per service.
  • Keep the Collector healthy. It is production infrastructure. Use the memory limiter and batch processor, set resource requests, run replicas, and alert on dropped data and exporter failures.
  • Treat telemetry as sensitive data. Spans and logs can include user identifiers and secrets. Scrub attributes in the Collector and control access to backends.
  • Mind stability levels. Core signals are stable, but individual conventions, instrumentation libraries and Collector components carry their own stability labels. Check them before depending on a component, and pin versions.
  • Plan for semantic convention changes. Attribute names evolve, which can break dashboards and alerts. Roll out changes with the migration guidance and keep dashboards tolerant during the transition.
  • Standardise ownership. Decide who owns the Collector pipeline, naming rules and the shared dashboards, so instrumentation does not drift across teams.

Adoption

  • The CNCF graduation material cites over 12,000 contributions from over 2,800 companies, hundreds of maintainers across language groups, and the second-highest contribution velocity among CNCF projects after Kubernetes.
  • CNCF’s graduation write-up links end-user talks from GitHub and Farfetch, and the project maintains a public adopters list. At incubation in 2021, F5, Grafana Labs, Shopify and Splunk were among the named adopters.
  • All three major clouds accept OTLP, and AWS, Azure and Google Cloud each provide OpenTelemetry distributions or exporters. Established observability vendors, including Datadog, New Relic, Dynatrace, Splunk and Elastic, accept OpenTelemetry data.

Alternatives

Solution Model Best suited to
Vendor agents (Datadog, New Relic, Dynatrace and similar) Proprietary agents tied to one platform Teams that want a single commercial platform and accept lock-in
Prometheus client libraries Metrics instrumentation for Prometheus Metrics-only monitoring, often used alongside OpenTelemetry
Jaeger and Zipkin client libraries Tracing instrumentation for specific backends Older tracing setups that have not migrated
eBPF-based tools (Grafana Beyla, Pixie, Cilium Hubble) Kernel-level observation without code changes Quick visibility, with less detail than application-level instrumentation
Cloud-native agents (AWS X-Ray, Google Cloud Trace) Provider tracing services Single-cloud estates, increasingly fed through OpenTelemetry

Strengths and limitations

Strengths Limitations
One vendor-neutral standard for traces, metrics and logs Large surface area, with components at different stability levels
Collector decouples applications from backends The Collector pipeline is another service to run, size and secure
Wide language, framework and vendor support, including all major clouds Instrumentation quality varies by language and library
Semantic conventions make data comparable across services Convention changes can disrupt dashboards and queries
Graduated project with very active development Telemetry volume and cost still need deliberate management

Recommendation

Adopt OpenTelemetry as the standard instrumentation and telemetry pipeline for new services and as the migration target for proprietary agents, using OTLP end to end and a two-tier Collector deployment on Kubernetes. Agree attribute conventions and sampling policy before rolling out widely, and pair it with backends your team can operate, such as Prometheus, Tempo and Grafana, or with a managed platform that accepts OTLP.

References: CNCF project page · CNCF blog: OpenTelemetry becomes a CNCF Incubating project · OpenTelemetry: Graduated, now what · CNCF blog: OpenTelemetry has graduated, now what · OpenTelemetry documentation · OpenTelemetry on GitHub

Spotted something wrong or out of date?Suggest a correction. I review every suggestion before changing the page.