DevOps · CNCF daily
Kubernetes: The Control Plane for Containerised Infrastructure
CNCF Graduated. The container orchestrator that schedules, scales, heals and connects workloads across a cluster through a declarative API, and the foundation of the cloud native ecosystem.
CNCF Graduated
- Kubernetes
- containerd
- etcd
- CoreDNS
- Helm
Kubernetes is an open-source system for automating the deployment, scaling and management of containerised applications. Engineers describe the desired state of their workloads (how many replicas, which image, what resources, how they are exposed), and Kubernetes continuously works to make the cluster match that description: it places containers on machines, restarts them when they fail, replaces them during upgrades and routes traffic to healthy instances. It addresses the problem that arrived with containers at scale: running hundreds of services across many machines by hand is slow, fragile and wasteful. Kubernetes was open-sourced by Google in 2014, drawing on more than a decade of experience with its internal Borg system. It was the CNCF’s first project, accepted in March 2016, and the first to graduate, in March 2018.
CNCF daily · Sprint 1, day 2 · Scheduling & Orchestration
At a glance
| Category | Scheduling & Orchestration |
| CNCF status | Graduated. Accepted March 10, 2016 (Incubating); graduated March 6, 2018 |
| Written in | Go |
| License | Apache 2.0 |
| Release cadence | About three minor releases a year, each supported for roughly 14 months |
| State store | etcd |
| Extension points | CRDs and controllers (operators), CRI, CNI, CSI, admission webhooks, Gateway API |
Architecture
flowchart TB
USER[kubectl, CI/CD, GitOps] --> API
subgraph CP[Control plane]
API[kube-apiserver]
ETCD[(etcd)]
SCHED[kube-scheduler]
CM[kube-controller-manager]
CCM[cloud-controller-manager]
API --> ETCD
SCHED --> API
CM --> API
CCM --> API
end
subgraph N1[Worker node]
KL1[kubelet]
KP1[kube-proxy or eBPF dataplane]
CRI1[containerd]
P1[Pods]
KL1 --> CRI1 --> P1
end
subgraph N2[Worker node]
KL2[kubelet]
KP2[kube-proxy or eBPF dataplane]
CRI2[containerd]
P2[Pods]
KL2 --> CRI2 --> P2
end
KL1 --> API
KL2 --> API
CCM --> CLOUD[Cloud APIs<br/>load balancers, disks, nodes]
| Component | Responsibility |
|---|---|
| kube-apiserver | The single entry point. Authenticates, authorises and validates every request, and is the only component that talks to etcd. |
| etcd | Consistent key-value store holding all cluster state. |
| kube-scheduler | Assigns new pods to nodes based on resource requests, affinity, taints and topology constraints. |
| kube-controller-manager | Runs reconciliation loops (Deployments, ReplicaSets, Jobs, nodes, endpoints) that drive actual state towards desired state. |
| cloud-controller-manager | Integrates with the cloud provider for load balancers, routes and node lifecycle. |
| kubelet | Node agent that starts and monitors pods through the container runtime (CRI) and reports status. |
| kube-proxy / CNI | Implements Service networking and pod connectivity; often replaced by an eBPF dataplane such as Cilium. |
Design principle. Everything in Kubernetes is a declarative API object reconciled by a controller. Users write desired state, controllers observe actual state and act to close the gap, over and over. The same pattern is open to anyone through Custom Resource Definitions, which is why so much of the cloud native ecosystem (Argo CD, cert-manager, database operators) is built as Kubernetes controllers.
Production reference design
A production cluster on a managed service (EKS, GKE or AKS) or on-premises with kubeadm, run by a platform team:
- Create the cluster with infrastructure as code. For example, on AWS with eksctl (or OpenTofu/Terraform modules), using managed node groups across three availability zones:
On-premises,eksctl create cluster --name prod --region ap-south-1 --version <current-supported> \ --nodegroup-name general --node-type m6i.xlarge --nodes 3 --nodes-min 3 --nodes-max 9 --zones ap-south-1a,ap-south-1b,ap-south-1ckubeadm initwith three control-plane nodes behind a load balancer and a stacked or external etcd gives the same shape. - Install the platform layer: a CNI with network policy (Cilium or Calico), an ingress or Gateway API controller, cert-manager, external-dns, a metrics stack, and a GitOps controller (Argo CD or Flux) that installs everything else from Git.
- Run every workload with requests, probes and disruption budgets. A minimal production-ready service:
apiVersion: apps/v1 kind: Deployment metadata: { name: payments-api, namespace: payments } spec: replicas: 3 selector: { matchLabels: { app: payments-api } } template: metadata: { labels: { app: payments-api } } spec: topologySpreadConstraints: - maxSkew: 1 topologyKey: topology.kubernetes.io/zone whenUnsatisfiable: ScheduleAnyway labelSelector: { matchLabels: { app: payments-api } } containers: - name: api image: registry.example.com/payments-api:2.8.1 resources: requests: { cpu: 250m, memory: 256Mi } limits: { memory: 512Mi } readinessProbe: { httpGet: { path: /healthz, port: 8080 } } livenessProbe: { httpGet: { path: /livez, port: 8080 } } --- apiVersion: policy/v1 kind: PodDisruptionBudget metadata: { name: payments-api, namespace: payments } spec: minAvailable: 2 selector: { matchLabels: { app: payments-api } } - Scale automatically: a HorizontalPodAutoscaler (or KEDA for event-driven scaling) for pods, and Cluster Autoscaler or Karpenter for nodes.
- Isolate teams with namespaces, RBAC bound to SSO groups, resource quotas, network policies and Pod Security Admission at the
restrictedlevel.
Typical production uses include microservice platforms, internal developer platforms, batch and ML workloads on GPUs, and hybrid estates where the same API runs in the cloud and on-premises.
Operational considerations
- Upgrades are continuous work. Minor versions leave support after about 14 months. Budget for upgrading control plane and nodes at least twice a year, and test deprecated API removals first.
- Protect etcd and the API server. Back up etcd (managed services do this for you), keep the API endpoint private where possible, and watch API server latency as clusters and controllers grow.
- Requests drive cost. Over-sized resource requests waste nodes; under-sized ones cause evictions. Right-size from real usage and use tools such as OpenCost to see cost per namespace.
- Security is layered. RBAC least privilege, Pod Security Admission, network policies, image signing and admission policy (Kyverno or OPA Gatekeeper), and secrets encryption at rest.
- Day-2 complexity is the real cost. The platform layer around Kubernetes (networking, ingress, certificates, observability, GitOps) needs ownership; small teams often do better on a managed service or a simpler platform.
Adoption
- Kubernetes is the de facto standard for container orchestration, offered as a managed service by every major cloud provider (EKS, GKE, AKS, OKE and others).
- The CNCF publishes hundreds of end-user case studies built on Kubernetes; recent ones on the project page include Zhuoyu Technology, which reached 95% GPU allocation with Kubernetes and Koordinator, and Minga, where a three-person team cut cloud costs by 30–40%.
- It was created at Google, and its contributor base spans thousands of individuals and companies, making it one of the largest open-source projects in the world.
Alternatives
| Solution | Model | Best suited to |
|---|---|---|
| Managed container services (ECS, Cloud Run, Azure Container Apps) | Provider-run orchestration with a smaller API | Teams that want containers without operating Kubernetes |
| HashiCorp Nomad | Single-binary scheduler for containers, VMs and binaries | Mixed workloads and simpler operations |
| Docker Swarm | Lightweight clustering built into Docker | Small, simple deployments |
| Lightweight distributions (K3s, k0s, MicroK8s) | Kubernetes with a smaller footprint | Edge, on-premises and CI clusters |
| Serverless platforms (Lambda, Knative on Kubernetes) | Function or request-driven scaling to zero | Event-driven and bursty workloads |
Strengths and limitations
| Strengths | Limitations |
|---|---|
| Declarative API with self-healing and rolling updates | Steep learning curve and many moving parts |
| Portable across every major cloud and on-premises | Frequent upgrades and API deprecations to manage |
| Extensible through CRDs and operators | A production platform needs many add-ons to be complete |
| Huge ecosystem, talent pool and vendor support | Easy to over-provision and overspend without cost visibility |
| Efficient bin-packing and autoscaling at scale | Overkill for a handful of simple services |
Recommendation
Adopt Kubernetes as the standard runtime for containerised services once an organisation runs more than a handful of them or needs portability across environments. Prefer a managed control plane, install the platform layer through GitOps, and enforce requests, probes, disruption budgets and policy from day one. For very small estates, consider a managed container service until the operational investment is justified.
References: CNCF project page · Kubernetes documentation · Kubernetes release history · CNCF case studies