Running a Hadoop cluster on Worldstream dedicated servers: sizing and network guidance
Quick answer
A Hadoop cluster on Worldstream runs on unshared physical hardware in a primary/worker layout: primary nodes handle NameNode and ResourceManager duties, worker nodes handle DataNode and NodeManager duties. Because there's no virtualisation layer between your cluster and the hardware, you get direct CPU, RAM and disk I/O access. Network throughput between nodes is usually the thing that limits the cluster, so confirm what inter-node connectivity you can get before you fix the design.
Cluster architecture
A dedicated Hadoop cluster follows a primary/secondary architecture. Primary nodes run the NameNode and ResourceManager services, handling metadata and job coordination for the cluster. Worker nodes handle the actual data storage and task execution through DataNode and NodeManager services. For high availability you can add a Secondary NameNode or JournalNodes, which provide metadata redundancy so the cluster can keep running through a primary node failure.
Sizing your nodes
There's no single fixed rule for node sizing, it depends heavily on job type, but a few relationships hold generally. Memory needs to scale with CPU core count, since a node with plenty of cores and too little memory per core will bottleneck on memory well before it runs out of compute. Node count and per-node size are a trade-off too: fewer, larger nodes reduce coordination and network overhead, while more, smaller nodes give you finer-grained fault isolation. On storage, NVMe capacity for hot data and local scratch space matters more than raw capacity once shuffle-heavy jobs are involved, with HDD still doing useful work for colder, less frequently queried data. Size against your own workload rather than a generic target.
Storage tiering is worth planning deliberately. If you mostly need capacity for data that's rarely queried, HDD is still a reasonable choice. If you need performance, an SSD or NVMe tier for hot data and local scratch space makes a real difference, particularly for shuffle-heavy jobs. A hybrid design, HDD for bulk capacity plus NVMe for hot paths, is a common and sensible pattern rather than an edge case.
If your platform also runs Spark for ETL-style processing or ClickHouse for OLAP-style analytics alongside Hadoop, the same node pool can serve both, or you can separate batch processing from interactive serving across different pools depending on how contended your workload gets.
Why network speed matters
Hadoop relies on heavy data transfer between nodes, especially during MapReduce operations and HDFS replication. Bandwidth between nodes is what keeps that data movement efficient and keeps replication from becoming the bottleneck in distributed processing. In practice, the network is often the limiting factor before compute or storage is, particularly for replication and shuffle-heavy workloads, so treat it as a design input rather than an afterthought and check what is available for the servers you are ordering.
Beyond raw bandwidth, the things that tend to bite in production are memory throughput during joins, fast I/O for shuffle spill, east-west bandwidth between worker nodes, and storage write amplification from replication. Planning for these upfront is cheaper than retrofitting a cluster that's already in production.
Dedicated hardware vs. cloud infrastructure
Dedicated servers remove virtualisation overhead entirely, giving direct access to CPU, RAM and disk I/O. That means no noisy-neighbour effects on disk I/O or CPU, which is what makes task execution times and replication throughput more consistent than on shared, virtualised infrastructure.
For large-scale, always-on workloads, dedicated servers typically offer better long-term value than cloud infrastructure, since they come with fixed, predictable monthly pricing and avoid the variable billing that usage spikes and data transfer volumes can bring in cloud environments.
Getting started
Start from the dedicated server listing on worldstream.com, where you can filter on processor family, uplink speed, data centre location and delivery time, then work through the configurator for processor, memory, storage and network options on the server you pick. If you need to check whether a specific cluster layout is possible before you order, open a support ticket in Portal.