Choosing a CPU for CI/CD, virtualisation and Kubernetes nodes
Kort antwoord
Match core count and clock speed to what the workload actually bottlenecks on: CI/CD build nodes want high single-thread performance and enough cores to run parallel jobs, virtualisation and Kubernetes nodes want core count and consistent per-core performance over raw clock speed, since resources get sliced across many tenants or pods.
Sizing by workload type
| Workload | What matters most | Why |
|---|---|---|
| CI/CD, build nodes | High single-thread performance, enough cores for parallel jobs | Compilers and test suites are often single-threaded per job; running several jobs at once needs enough cores to avoid queueing |
| Virtualisation (Proxmox, VMware) | Core count, consistent per-core performance | Each VM gets a slice of CPU; more cores means more VMs without contention, and consistency matters more than peak clock speed |
| Kubernetes worker nodes | Core count, headroom above committed pod requests | Kubernetes schedules pods against requested CPU, not actual usage; undersized nodes lead to throttling under load |
| Databases, high-IO services | Fast single-core performance, paired with NVMe storage | Many database engines are still largely single-threaded per query; storage speed matters as much as CPU here, see Choosing storage media |
Sizing for mixed workloads
When one machine runs a mix (say, a handful of VMs plus a build pipeline), start from the busiest individual workload's requirement, then add headroom for the others rather than averaging everything together. A useful starting point: sum the vCPUs you plan to allocate to VMs or pods, add margin for the hypervisor or control-plane overhead itself, and size storage IOPS to the workload with the highest I/O demand, not the average.
Troubleshooting: high load, but the CPU looks idle
If your application feels slow but CPU usage graphs show plenty of idle capacity, look at CPU steal time before assuming the application itself is the problem. Steal time is the percentage of time a virtual CPU is ready to run but waiting for the physical host to schedule it, common on shared virtualised infrastructure under contention.
- On Linux, check steal time with
mpstat -P ALL 1(the%stealcolumn) orsar -u 1. - Correlate with application-level p95/p99 latency, not just averages: steal time tends to show up as occasional latency spikes rather than a steady slowdown.
- If steal time is consistently high, the fix is usually more dedicated resources (a larger VPS tier, or a dedicated/Bare Metal Compute server) rather than tuning the application further.