How to benchmark storage performance: IOPS, throughput and latency
Quick answer
Storage performance isn't one number, it's three: IOPS (operations per second, matters most for small, random reads and writes such as databases), throughput (total data moved per second, matters most for large sequential transfers such as backups), and latency (how long a single operation takes, matters wherever responsiveness matters). Benchmarking badly, testing from cache, testing the wrong access pattern, or testing too briefly, produces numbers that look great and mean nothing.
The three metrics, and when each one matters
IOPS (input/output operations per second) counts how many individual read or write operations storage can complete each second. It's the number that matters most for workloads built out of lots of small, random operations, a database handling many small queries against scattered rows, for instance, where each operation is tiny but there are a huge number of them happening concurrently.
Throughput counts something different: the total volume of data moved per second, usually in MB/s or GB/s. It's the number that matters for large, sequential transfers, copying a big backup file, streaming media, moving a large dataset, where the operations are few and large rather than many and small.
Latency measures how long a single operation takes to complete, independent of how many are happening at once. It's the number that matters for anything sensitive to responsiveness: an application that feels sluggish even though its overall data volume is small usually has a latency problem, not a throughput problem.
These three don't move together. Storage that posts excellent throughput on large sequential transfers can still have mediocre IOPS on small random ones, and vice versa, because they're stressing the storage in different ways. Benchmarking for the wrong metric, or only one of the three, tells you less than it looks like it does.
| Metric | What it measures | Matters most for |
|---|---|---|
| IOPS | Operations completed per second | Small, random reads/writes: databases, transactional workloads |
| Throughput | Total data moved per second | Large, sequential transfers: backups, media, bulk data movement |
| Latency | Time for one operation to complete | Responsiveness-sensitive workloads of any size |
Common mistakes that produce misleading results
A few habits reliably make benchmark numbers look better than real-world performance will be.
Testing with a file small enough to be cached. Operating systems and storage controllers cache recently accessed data in memory. If your test file is small enough to sit entirely in that cache, you're measuring memory speed, not disk speed, and the numbers will be far higher than what the underlying storage can actually sustain. A meaningful test needs a working set large enough to push past whatever cache sits in front of the storage.
Testing the wrong access pattern. Sequential and random access stress storage completely differently, and a device that's excellent at one can be unremarkable at the other. Running a sequential-read test and then assuming it tells you anything about your database's random-write performance, or the reverse, produces numbers that don't reflect how the storage will actually behave under your real workload.
Not running the test long enough. A short burst can look excellent because it's drawing on a cache or a burst-capacity buffer that a longer, sustained run will eventually exhaust. Real workloads run for minutes or hours, not seconds, so a benchmark that only lasts a few seconds tells you about the burst, not about the sustained performance you'll actually get once that headroom runs out.
What to test with
fio (Flexible I/O Tester) is the standard tool for this kind of testing on Linux. It can generate sequential or random access patterns, at whatever block size and queue depth you choose, and report IOPS, throughput and latency from the same run, which is what makes it possible to test the specific access pattern your real workload actually produces rather than a generic default. Working through actual fio command syntax and result interpretation is its own topic, kept out of this article deliberately, the point here is knowing what to measure and how to avoid fooling yourself before you get there.