Running Private Cloud day to day: nodes, failover and updates
Quick answer
A hyperconverged (HCI) cluster can start small: two nodes plus a witness node for quorum, though three nodes gives you full redundancy from day one. A node failure doesn't take your VMs down, it triggers automatic replica rebuilding while workloads keep running on the remaining nodes. Updates roll through node by node so the cluster stays available throughout. Stretching a cluster across sites for disaster recovery only works within a tight latency budget, beyond that you fall back to asynchronous replication instead.
What this article covers
Private Cloud vs. Flexible VPS explains when to choose Private Cloud over shared infrastructure. This article goes a level deeper: the operational mechanics of running a hyperconverged cluster day to day, regardless of whether that cluster is the VMware-based platform underneath Worldstream's own Flexible Cloud product, Private Cloud's own infrastructure, or a stack you build yourself on Bare Metal Compute. The mechanics below apply to hyperconverged infrastructure generally: how many nodes you start with, what happens when one fails, how updates get applied, and how disaster recovery works between locations.
Starting small: two nodes and a witness for quorum
You don't need a large cluster to get started. Some HCI platforms, including Proxmox and StarWind, support two-node clusters by adding a lightweight witness node purely to hold the tie-breaking vote for quorum, the mechanism that decides which side of the cluster is authoritative if the nodes lose contact with each other.
The trade-off is redundancy. A two-node cluster doesn't give you full redundancy: if one of the two data nodes fails, you're left running on degraded capacity until it's repaired. For production workloads, three nodes is the better starting point, since it gives every node somewhere to fail over to without the cluster running hot on a single remaining node.
As a rough sizing guide: two to three nodes suits small or remote/branch-office deployments, three to five nodes is typical for general enterprise use, and five or more nodes is where erasure coding (covered below) starts to pay off for performance-sensitive workloads. Odd node counts also make quorum decisions cleaner.
What happens during a node failure
The cluster's behaviour during a node failure depends on how data is replicated across nodes:
- Two-way replication: your VMs keep running on the remaining nodes, and the cluster automatically starts rebuilding the lost replicas onto spare capacity elsewhere in the cluster.
- Three-way replication: the cluster can survive two simultaneous node failures. Workloads continue running on the healthy nodes while missing replicas are rebuilt automatically in the background.
Either way, the point of replication is that a single hardware failure is an operational event the cluster absorbs on its own, not an incident that takes workloads offline. For longer-term protection beyond node-level replication, snapshots give you a fast rollback point and off-site backups give you recovery if something affects the whole cluster, not just one node. Testing restore procedures on a regular basis is what turns that protection from theoretical into reliable.
Keeping the cluster patched: rolling updates without full downtime
Most HCI platforms support rolling updates: nodes are upgraded one at a time, using upgrade domains to keep quorum and data protection intact throughout, while the VMs that were running on the node being patched migrate to the other nodes first. The cluster as a whole stays available even though individual nodes cycle through maintenance in sequence.
In practice, that means scheduling updates during a planned maintenance window, verifying cluster health before you start, and, where you have a dev/test cluster available, testing the update there first. Firmware updates on the underlying hardware follow the same rolling pattern: one node's firmware is updated and validated before moving to the next, rather than taking the whole cluster down at once.
Growing the cluster works the same way in reverse: adding nodes lets the cluster rebalance data across the new capacity automatically, and that rebalancing, like maintenance, happens without taking workloads offline.
Disaster recovery between sites: the latency budget
Protecting a cluster against a full site failure, not just a node failure, means getting data to a second location. The usual pattern is asynchronous replication: snapshots and backups are replicated to a secondary site on a schedule, so the second site is always slightly behind but never dependent on real-time connectivity between the two.
A stretched cluster, where nodes at two sites act as a single synchronous cluster, is a stronger form of protection, but it only works within a strict latency budget between the two locations, typically cited as sub-5ms round-trip time. Beyond that budget, synchronous writes across the link become the bottleneck, and asynchronous replication is the more realistic option. Testing your disaster recovery plan matters as much as designing it: simulating a node or site failure, running an actual failover, and verifying that data comes back intact is what confirms the plan works before you need it for real. Many HCI platforms include built-in DR testing that doesn't touch production while you do it.
Choosing, or inheriting, the underlying stack
Worldstream's own Flexible Cloud product runs on VMware, specifically VMware Cloud Foundation with vSAN as the storage layer, all-flash NVMe under the hood. See Managing your Flexible Cloud environment via VMware Cloud Director for how that stack is put together and managed. Private Cloud is dedicated, non-shared infrastructure on Worldstream's own network. If you need to know who runs the hypervisor and HCI layer on a specific Private Cloud environment, ask Worldstream.
If you're instead building your own hyperconverged cluster on Bare Metal Compute, Worldstream delivers the physical servers and networking, and you choose the HCI software yourself. That choice comes down to a genuine trade-off:
| Open-source (Proxmox VE, Ceph) | Proprietary (VMware vSAN, Nutanix AHV) | |
|---|---|---|
| Licensing | No licensing costs, no vendor lock-in | Licensing costs apply |
| Support | Community and third-party support | Vendor-backed support included |
| Fit | Teams comfortable managing the stack themselves, Proxmox suits general use, StarWind suits Windows-heavy environments | Teams that already run vSphere (vSAN) or want a single all-in-one platform (Nutanix) |
There's no universally correct answer here, it depends on your existing stack, your team's comfort operating the platform itself, and your budget for licensing versus vendor support.
A note on HCI vs. software-defined storage (SDS)
These two get confused because they solve overlapping problems. SDS aggregates storage across nodes but still relies on separate compute nodes elsewhere in the architecture. HCI bundles compute and storage on the same nodes, which simplifies day-to-day operations, but also means a node failure affects both compute and storage capacity at once, since the two failure domains are coupled rather than separate.