Skip to main content
Support
0
Contact us
Nederlands
Deutsch
Español
Dedicated serversFlexible VPSCloud TechnologyColocationChallenges in ITSectorsCareers
Cloud Compute

With Cloud Compute, you have access anytime and anywhere to a portal through which you can configure your entire IT environment from wherever you are in the world.

Cloud Storage

Reliable access to your files, infrastructure, and applications at all times – with no interruptions or delays. At Worldstream, we offer a variety of storage solutions.

Flexible cloud icon
Flexible cloud
Private cloud icon
Private cloud
Bare metal icon
Bare Metal Compute
Hollow cube icon
Object storage
Hollow cube icon
File storage
Block storage icon
Block storage
Backup storage icon
Backup storage
Need support?

With experienced engineers and an average response time track record on 7 minutes, you can expect a solid technical support solution in next to no time.

All Servers

Choose your Dedicated Server now. Custom or Instant Delivery. Powerhouse servers built for your use case.

Use Cases

Whatever your use case, we’re here to help you find the ideal solution.

Deal servers icon
Deals
AMD servers icon
AMD Processors
AI servers icon
Intel Processors
Hollow cube icon
Virtualisation, Containerisation and Orchestration
Hollow cube icon
Websites and Applications
Hollow cube icon
Gaming and Streaming Infrastructure

24/7/365 support with an average response time of just 7 minutes. Thanks to our own data centers, our engineers can go directly to your server for fast, hands-on assistance. Email or call us anytime.

Smart outsourcing

Some IT creates added value, while other types are supportive. Use that as a starting point for outsourcing.

Cost Efficiency

Complete IT packages may seem like the safe option, but when you consider the costs, other choices often make more sense.

IT flexibility & control

Outsourcing doesn’t mean losing control; it actually provides more flexibility and control.

Cloud repatriation

The cloud is not a final destination: You should continuously evaluate and adjust your cloud environment as needs evolve.

Financial services
Logistics & Transportation
Retail & E-commerce
Media & Entertainment
Tech & Software Development
Security
Managed Service Providers
Need support?

With experienced engineers and an average response time track record on 7 minutes, you can expect a solid technical support solution in next to no time.

Chat with usContact us
About WorldstreamAbout the technologyCasesKnowledge base
About usMeet the teamJobsBecome a resellerCertificationsOur data centersOur networkDDoS ProtectionAMD EPYC serversTechnology PartnersOperating SystemsAll casesEasyTerraDutch Drone CompanyPerfGridArticlesFAQNews and BlogsProducts and Services
Contact us

Call +31 (0) 174 – 712 117

Industriestraat 53, Naaldwijk

Nederlands
Deutsch
Español
0
Dedicated serversFlexible VPSCloud TechnologyColocationChallenges in ITSectorsCareersAbout WorldstreamAbout the technologyCasesKnowledge baseMy Worldstream
Contact
Support
NederlandsDeutschEspañol
  1. HomeHome
  2. Knowledge Base
  3. Compute
  4. Running Private Cloud day to day: nodes, failover and updates

Running Private Cloud day to day: nodes, failover and updates

Applies to Private Cloud, Flexible Cloud, HCI on Bare Metal ComputeAudience Technical evaluator, existing customerLast reviewed September 2026

Quick answer

A hyperconverged (HCI) cluster can start small: two nodes plus a witness node for quorum, though three nodes gives you full redundancy from day one. A node failure doesn't take your VMs down, it triggers automatic replica rebuilding while workloads keep running on the remaining nodes. Updates roll through node by node so the cluster stays available throughout. Stretching a cluster across sites for disaster recovery only works within a tight latency budget, beyond that you fall back to asynchronous replication instead.

On this page
  • What this article covers
  • Starting small: two nodes and a witness for quorum
  • What happens during a node failure
  • Keeping the cluster patched: rolling updates without full downtime
  • Disaster recovery between sites: the latency budget
  • Choosing, or inheriting, the underlying stack
  • A note on HCI vs. software-defined storage (SDS)

What this article covers

Private Cloud vs. Flexible VPS explains when to choose Private Cloud over shared infrastructure. This article goes a level deeper: the operational mechanics of running a hyperconverged cluster day to day, regardless of whether that cluster is the VMware-based platform underneath Worldstream's own Flexible Cloud product, Private Cloud's own infrastructure, or a stack you build yourself on Bare Metal Compute. The mechanics below apply to hyperconverged infrastructure generally: how many nodes you start with, what happens when one fails, how updates get applied, and how disaster recovery works between locations.

Starting small: two nodes and a witness for quorum

You don't need a large cluster to get started. Some HCI platforms, including Proxmox and StarWind, support two-node clusters by adding a lightweight witness node purely to hold the tie-breaking vote for quorum, the mechanism that decides which side of the cluster is authoritative if the nodes lose contact with each other.

The trade-off is redundancy. A two-node cluster doesn't give you full redundancy: if one of the two data nodes fails, you're left running on degraded capacity until it's repaired. For production workloads, three nodes is the better starting point, since it gives every node somewhere to fail over to without the cluster running hot on a single remaining node.

As a rough sizing guide: two to three nodes suits small or remote/branch-office deployments, three to five nodes is typical for general enterprise use, and five or more nodes is where erasure coding (covered below) starts to pay off for performance-sensitive workloads. Odd node counts also make quorum decisions cleaner.

What happens during a node failure

The cluster's behaviour during a node failure depends on how data is replicated across nodes:

  • Two-way replication: your VMs keep running on the remaining nodes, and the cluster automatically starts rebuilding the lost replicas onto spare capacity elsewhere in the cluster.
  • Three-way replication: the cluster can survive two simultaneous node failures. Workloads continue running on the healthy nodes while missing replicas are rebuilt automatically in the background.

Either way, the point of replication is that a single hardware failure is an operational event the cluster absorbs on its own, not an incident that takes workloads offline. For longer-term protection beyond node-level replication, snapshots give you a fast rollback point and off-site backups give you recovery if something affects the whole cluster, not just one node. Testing restore procedures on a regular basis is what turns that protection from theoretical into reliable.

Keeping the cluster patched: rolling updates without full downtime

Most HCI platforms support rolling updates: nodes are upgraded one at a time, using upgrade domains to keep quorum and data protection intact throughout, while the VMs that were running on the node being patched migrate to the other nodes first. The cluster as a whole stays available even though individual nodes cycle through maintenance in sequence.

In practice, that means scheduling updates during a planned maintenance window, verifying cluster health before you start, and, where you have a dev/test cluster available, testing the update there first. Firmware updates on the underlying hardware follow the same rolling pattern: one node's firmware is updated and validated before moving to the next, rather than taking the whole cluster down at once.

Growing the cluster works the same way in reverse: adding nodes lets the cluster rebalance data across the new capacity automatically, and that rebalancing, like maintenance, happens without taking workloads offline.

Disaster recovery between sites: the latency budget

Protecting a cluster against a full site failure, not just a node failure, means getting data to a second location. The usual pattern is asynchronous replication: snapshots and backups are replicated to a secondary site on a schedule, so the second site is always slightly behind but never dependent on real-time connectivity between the two.

A stretched cluster, where nodes at two sites act as a single synchronous cluster, is a stronger form of protection, but it only works within a strict latency budget between the two locations, typically cited as sub-5ms round-trip time. Beyond that budget, synchronous writes across the link become the bottleneck, and asynchronous replication is the more realistic option. Testing your disaster recovery plan matters as much as designing it: simulating a node or site failure, running an actual failover, and verifying that data comes back intact is what confirms the plan works before you need it for real. Many HCI platforms include built-in DR testing that doesn't touch production while you do it.

Choosing, or inheriting, the underlying stack

Worldstream's own Flexible Cloud product runs on VMware, specifically VMware Cloud Foundation with vSAN as the storage layer, all-flash NVMe under the hood. See Managing your Flexible Cloud environment via VMware Cloud Director for how that stack is put together and managed. Private Cloud is dedicated, non-shared infrastructure on Worldstream's own network. If you need to know who runs the hypervisor and HCI layer on a specific Private Cloud environment, ask Worldstream.

If you're instead building your own hyperconverged cluster on Bare Metal Compute, Worldstream delivers the physical servers and networking, and you choose the HCI software yourself. That choice comes down to a genuine trade-off:

Open-source (Proxmox VE, Ceph)Proprietary (VMware vSAN, Nutanix AHV)
LicensingNo licensing costs, no vendor lock-inLicensing costs apply
SupportCommunity and third-party supportVendor-backed support included
FitTeams comfortable managing the stack themselves, Proxmox suits general use, StarWind suits Windows-heavy environmentsTeams that already run vSphere (vSAN) or want a single all-in-one platform (Nutanix)

There's no universally correct answer here, it depends on your existing stack, your team's comfort operating the platform itself, and your budget for licensing versus vendor support.

A note on HCI vs. software-defined storage (SDS)

These two get confused because they solve overlapping problems. SDS aggregates storage across nodes but still relies on separate compute nodes elsewhere in the architecture. HCI bundles compute and storage on the same nodes, which simplifies day-to-day operations, but also means a node failure affects both compute and storage capacity at once, since the two failure domains are coupled rather than separate.

Related articles

  • Private Cloud vs. Flexible VPS: dedicated resources vs. shared infrastructure
  • Managing your Flexible Cloud environment via VMware Cloud Director
  • Dedicated vs. public cloud: understanding the real cost drivers
Was this article helpful?

Solid IT. No Surprises

Sparring partner for IT maturity
Eliminating barriers so you can run
Predictable and transparant costs

Contact

  • Industriestraat 53, Naaldwijk
  • Payment Methods
  • Abuse
  • Developers Resources
  • Network Operations Center
  • About us
  • Meet the team
  • Jobs
  • Become a reseller
  • Certifications
  • Our data centers
  • Our network
  • DDoS Protection
  • AMD EPYC servers
  • Technology Partners
  • Operating Systems
  • Overview
  • FAQ
  • Cases
  • News & Blogs
  • Use Cases
Nederlands
Deutsch
Español
Nederlands
Deutsch
Español
  • Legal
  • Disclosure