Right-sizing RAM and storage for your workload
Kort antwoord
Size RAM to your working set and peak concurrent load, not your total dataset or average usage, and leave headroom: a server running near 95% RAM starts swapping and performance falls off sharply, not gradually. Size storage to your data growth rate over the contract term, not just what you need today, and budget backup or snapshot storage separately from live data.
Estimating RAM: working set, not total dataset
A database server doesn't need enough RAM to hold its entire dataset, it needs enough RAM to cache its working set: the portion of data actually being read and written regularly. A 500GB database where most queries touch the same 20GB of recent, active data performs very differently from one where queries scan across the whole dataset evenly, even though both are "500GB databases" on paper. Sizing RAM against the wrong number, the total dataset size instead of the working set, leads to either badly over-provisioned servers or under-provisioned ones that look fine in testing and struggle in production.
The same logic applies to application memory footprint: size for peak concurrent load, not average load. An application that comfortably fits in 4GB most of the day but spikes to 12GB during a daily peak needs to be sized for the 12GB, not the average, or it will fail exactly when it matters most.
Whatever the workload, leave headroom above whatever number you calculate. RAM usage isn't a resource that degrades gracefully as it fills up: a system running near capacity starts swapping memory pages to disk, and once that starts, performance doesn't decline gradually, it falls off a cliff, since disk is orders of magnitude slower than RAM for the kind of random access swapping produces.
Estimating storage: growth rate, not current size
Storage sizing mistakes are usually a timing problem, not a capacity problem: it's easy to size for what you need on day one and forget that data grows over the life of the contract. Look at the growth rate, not just the current footprint, and size with that trajectory in mind rather than resizing under pressure later.
Two further habits keep storage sizing accurate:
- Separate OS and application storage from data storage where possible. Keeping the operating system and application binaries on their own volume, separate from the data that actually grows, makes it far easier to resize, migrate, or reinstall without touching the data itself.
- Budget backup and snapshot storage on top of live data, not inside it. Backups and snapshots consume their own capacity in addition to the data they're protecting, and that add-on can be substantial depending on retention. See Backups and snapshots: what's the difference for how the two approaches differ and what each costs in storage terms.
Typical RAM:storage patterns by workload
These are general patterns to sanity-check a sizing decision against, not fixed rules, actual ratios depend heavily on the specific application and data profile.
| Workload type | Typical pattern | Why |
|---|---|---|
| General web application | Moderate RAM, storage sized to content/media growth | RAM covers the application and a request cache; storage tends to grow with uploaded content and logs rather than the application itself |
| Database | RAM sized to working set, storage sized to full dataset plus growth | Query performance depends on how much of the active data fits in RAM; storage still has to hold everything, including data outside the working set |
| Cache / in-memory store (e.g. Redis) | RAM-heavy relative to storage, storage mainly for persistence snapshots | The entire point of the workload is serving from memory; storage is a secondary concern, mostly for durability |
| File server | Storage-heavy relative to RAM | Performance depends more on capacity and I/O throughput than on caching large amounts of file content in RAM |
Pairing this with CPU and storage media decisions
RAM and storage capacity are only two parts of the sizing picture. See Choosing a CPU for CI/CD, virtualisation and Kubernetes nodes for the compute side, and Choosing storage media: NVMe vs. SSD vs. HDD by workload for how storage type, not just storage size, affects the same workloads covered here.