Sizing a Kubernetes cluster for high availability
Kort antwoord
Three nodes is the smallest reliable footprint. For a highly available control plane, use three control-plane nodes plus two or more workers. Before you size for your workloads, budget roughly a fifth of total cluster capacity for the platform itself, etcd, the control plane, and system pods.
The smallest reliable footprint is three nodes
If you're building anything you'd call "reliable", three nodes is the floor, not the target to shrink towards. A single node has no failure domain: if it goes down, the cluster goes down with it. Two nodes doesn't solve that either, because most consensus mechanisms (etcd included) need a majority to keep functioning, and a majority of two is two.
For a genuinely highly available control plane, the recommended shape is three control-plane nodes plus two or more worker nodes. The three control-plane nodes give etcd a quorum that survives a single node failure; the workers carry your actual application load and can scale independently of the control plane.
What eats capacity before your workloads do
Every Kubernetes cluster spends some of its resources on itself before your applications get anything. As a planning baseline:
- Each control-plane (master) node needs roughly 2 to 4GB of RAM for the control plane components themselves (API server, scheduler, controller-manager, etcd).
- Each worker node needs roughly 100 to 300MB of RAM for system pods: the CNI agent, kube-proxy, node-level monitoring, and similar.
- As a rule of thumb, plan for around 20% extra capacity across the cluster for the platform itself, on top of what your workloads need.
That 20% isn't a hard ceiling, it moves depending on how many add-ons and DaemonSets you run (ingress controllers, service meshes, logging agents), but it's a reasonable starting assumption for sizing worker nodes so you don't discover the overhead only after workloads start getting evicted.
Separate control-plane sizing from worker sizing
Control-plane nodes and worker nodes have different jobs, and it's worth sizing them separately rather than ordering identical hardware for both.
Control-plane nodes are mostly bound by etcd's disk and network latency sensitivity, and by API server throughput at larger cluster sizes. They don't need to scale with your application load, only with the number of objects (pods, services, secrets) the cluster is tracking. For small and mid-sized clusters, a modest, consistent node size is normally enough, and it's more important that all three are similarly specced than that any one of them is especially powerful.
Worker nodes are bound by whatever your applications actually need: CPU, memory, and increasingly core count as clusters grow and you pack more pods per node. As a cluster grows from a handful of services to a larger footprint, it's normal for worker node specifications to scale up (or out) independently of the control plane, which typically stays comparatively modest even as the rest of the cluster grows.
A basic HA sizing pattern
| Cluster size | Control plane | Workers | Notes |
|---|---|---|---|
| Smallest reliable (HA) | 3 nodes | 2+ nodes | The floor for a cluster you'd call production-worthy |
| Small production | 3 nodes, modest spec | 3 to 5 nodes | Room to lose a worker without capacity pressure |
| Growing / mid-sized | 3 nodes, consistent spec | Scaled to workload, added incrementally | Control plane usually doesn't need to grow at the same rate as workers |
Whatever size you land on, keep the 20% platform overhead and the per-node control-plane and system pod figures above in your sizing math from the start. Retrofitting headroom into a cluster that's already tight on capacity is a worse experience for everyone than planning for it up front.
Beyond the node count
Node count and HA control-plane sizing are the starting point, not the whole picture. Once you're past the minimum footprint, general Kubernetes practice also covers things like spreading replicas across nodes with pod anti-affinity, setting PodDisruptionBudgets so voluntary maintenance doesn't take out an entire service at once, and separating storage-heavy or stateful workloads from the rest of the scheduling pool. See Handling stateful services in containers for the state-specific part of that.