Performance Testing
This page is about finding out what is slow. It covers the two command line measurement tools the appliance ships -- qs-perftest and qs-ramdisk -- and a method for using them that narrows a performance complaint down to a single layer: the media, the network, or the pool and filesystem path above them.
Once you know which layer is losing the performance, Performance Tuning covers what to change.
| Section | Purpose |
|---|---|
| Measure before you tune | Why this page comes before the tuning page. |
| The isolation method | Work bottom-up, and why that order and not the other one. |
| Step 1: test the media | Read-test every disk and look for the outlier. |
| Step 2: rule out the network | Bonds, static routes and MTU, which is where most real-world cases end. |
| Step 3: take the media out of the equation | A RAM-backed disk proves whether the bottleneck is above the media. |
| Step 4: test the pool and the filesystem path | Benchmark a share with fio, elbencho, dd and iozone. |
| qs-perftest operation reference | All eight operations and their arguments. |
| qs-ramdisk operation reference | Create, list and destroy, and what backs the devices. |
| Where results are written | The log and results files each tool leaves behind. |
Measure before you tune
Measuring comes before tuning. Changing a tunable without first isolating the layer is guesswork, and the two most common real causes of a performance complaint -- a degraded device and a misconfigured network -- are not fixed by any tunable, any I/O profile or any amount of cache. A pool with one failing disk in it will keep being slow after you have worked through every setting in Storage System Optimization.
So the first question is never "which setting should I change". It is "which layer is the performance being lost in".
Two things follow, and both are worth doing before you start:
- Take a baseline while the system is healthy. The tools on this page are far more useful when you have a number from before the problem to compare against. The Disk Performance Test in particular stores its result on each disk, so a baseline taken at install time is still there years later.
- Check device health first. Every device in Physical Disks should be Normal. A disk in Warning with a predictive-failure indicator explains a performance problem on its own -- see Physical Disks/Devices for the health, SMART and temperature columns and what the states mean.
The isolation method
Work bottom-up: media first, then the network, then the pool and filesystem path.
The reason is that each layer sits on the one below it, so a measurement taken at a higher layer includes everything underneath. A slow disk makes every pool, share, volume and client-side number meaningless, because all of them are measuring that disk plus whatever else is wrong. Fix or rule out the bottom layer and the layer above it becomes measurable; do it in the other order and you learn that something is slow without learning what.
Measuring top-down has the same problem in reverse. A client copying a file slowly tells you the whole stack is slow. It cannot tell you whether that is a drive, a bond, a route, the record size or the ARC, and every one of those produces the same symptom at the client.
| Step | Layer | Tool | What a bad result means |
|---|---|---|---|
| 1 | Physical media | qs-perftest readdisks / readpooldisks, or the Disk Performance Test dialog |
One or more devices are slow or failing. Replace before doing anything else. |
| 2 | Network and transport | Network Ports: bond mode, static routes, Network Check | A bond mode the switch does not support, a second gateway competing for the default route, or an MTU mismatch. |
| 3 | Everything above the media | qs-ramdisk plus a pass-thru Storage Volume |
If a RAM-backed device is also slow, the media was never the problem -- the loss is in the transport, the CPU or the target stack. |
| 4 | Pool and filesystem path | qs-perftest benchpool |
The pool layout, record size, compression or cache configuration. This is where Performance Tuning starts to apply. |
Both tools need root:
sudo qs-perftest ... sudo qs-ramdisk ...
Every qs-perftest operation needs root, including the read-only ones, because they read raw devices directly. qs-ramdisk list is the one command on this page that runs as any user.
The numbers in the examples below come from a small virtual-machine appliance whose disks are files on a flash datastore. Treat the shapes and the ratios as the lesson, not the absolute figures -- a real HDD reads at a small fraction of what is shown here.
Step 1: test the media
Read-test every disk
qs-perftest readdisks reads 1000 MiB from every disk in the system with dd in direct mode, so the page cache and the ARC are bypassed and the number is the device's own sequential read rate:
sudo qs-perftest readdisks
{Thu Sep 3 05:19:07 2026, INFO, py:qs_perftest} READ-DISK (/dev/sda): starting: dd if=/dev/sda of=/dev/null bs=1M count=1000 iflag=direct
{Thu Sep 3 05:19:08 2026, INFO, py:qs_perftest} READ-DISK (/dev/sda): completed: 1048576000 bytes (1.0 GB, 1000 MiB) copied, 0.903725 s, 1.2 GB/s
{Thu Sep 3 05:19:09 2026, INFO, py:qs_perftest} READ-DISK (/dev/sdb): completed: 1048576000 bytes (1.0 GB, 1000 MiB) copied, 0.667703 s, 1.6 GB/s
{Thu Sep 3 05:19:09 2026, INFO, py:qs_perftest} READ-DISK (/dev/sdc): completed: 1048576000 bytes (1.0 GB, 1000 MiB) copied, 0.677322 s, 1.5 GB/s
{Thu Sep 3 05:19:10 2026, INFO, py:qs_perftest} READ-DISK (/dev/sdd): completed: 1048576000 bytes (1.0 GB, 1000 MiB) copied, 0.761349 s, 1.4 GB/s
{Thu Sep 3 05:19:11 2026, INFO, py:qs_perftest} READ-DISK (/dev/sde): completed: 1048576000 bytes (1.0 GB, 1000 MiB) copied, 0.611526 s, 1.7 GB/s
{Thu Sep 3 05:19:12 2026, INFO, py:qs_perftest} READ-DISK (/dev/sdf): completed: 1048576000 bytes (1.0 GB, 1000 MiB) copied, 0.828239 s, 1.3 GB/s
{Thu Sep 3 05:19:12 2026, INFO, py:qs_perftest} READ-DISK (/dev/sdg): completed: 1048576000 bytes (1.0 GB, 1000 MiB) copied, 0.704059 s, 1.5 GB/s
That is what a healthy set looks like: seven devices of the same class within a narrow band of each other. Nothing here needs interpreting -- the point of the test is the shape of the distribution, not any single figure.
To scope the test to the disks of one pool, use readpooldisks with the pool's on-disk id, which is qs- followed by the pool's UUID:
sudo qs-perftest readpooldisks qs-596a6bb7-b2c7-d283-1951-00efdd93ece3
{Thu Sep 3 05:19:23 2026, INFO, py:qs_perftest} READ-DISK (sde1): starting: dd if=/dev/disk/by-id/scsi-36000c294c4007096dadf9f7ba9572929-part1 of=/dev/null bs=1M count=1000 iflag=direct
{Thu Sep 3 05:19:24 2026, INFO, py:qs_perftest} READ-DISK (sde1): completed: 1048576000 bytes (1.0 GB, 1000 MiB) copied, 0.740017 s, 1.4 GB/s
{Thu Sep 3 05:19:24 2026, INFO, py:qs_perftest} READ-DISK (sdf1): starting: dd if=/dev/disk/by-id/scsi-36000c2905242986a1a0820ea057b683d-part1 of=/dev/null bs=1M count=1000 iflag=direct
{Thu Sep 3 05:19:24 2026, INFO, py:qs_perftest} READ-DISK (sdf1): completed: 1048576000 bytes (1.0 GB, 1000 MiB) copied, 0.639151 s, 1.6 GB/s
Note that readpooldisks works from zpool status, so it reports the pool's partition names (sde1) and reads through the by-id paths the pool was built with -- the same devices, named the way the pool refers to them. readdisk tests a single device by path, which is the one to use when re-testing a suspect drive:
sudo qs-perftest readdisk /dev/sdf
readdisk silently does nothing if the path you give it is not a block device, so check the path if you get no output at all.
Reading the results: what an anomaly looks like
A single failing device can slow an entire pool. This is the most important thing on this page. OpenZFS spreads each write across the members of a device group and has to wait for the slowest of them, so one drive that has begun to fail -- retrying reads internally, remapping sectors, or negotiating a lower link rate -- sets the pace for every I/O that touches its group. The pool still reports Normal and the disk still reports no errors, because nothing has actually failed yet. Only a read test shows it.
So what you are looking for is not a number below a threshold. It is an outlier:
- One device markedly slower than identical siblings. A drive at a third or a half of the rate of the others in the same group is the finding, whatever the absolute figures are. Replace it.
- A whole device group slow while its individual disks are fast. That points at the shared path rather than the media -- a cable, an expander, a controller port. The Disk Performance Test dialog's Group by VDEV and Group by Enclosure modes exist for exactly this, and are described in Physical Disks/Devices.
- A whole enclosure slow. Same reasoning, one level up: check the SAS cabling and the expander.
- Everything uniformly slow. That is not a media anomaly. Go on to step 2.
Cross-check any outlier against the device's health before replacing it. Open the disk in Physical Disks, read State Detail in the Properties panel, and check the SMART state and temperature -- see Physical Disks/Devices. A drive that is both slow and reporting a predictive failure needs no further argument.
The Disk Performance Test dialog, and when to use it
The web interface has an equivalent read test on the physical disk right-click menu. Physical Disks/Devices documents it in full, including the Mode options and why the grouped modes exist.
The two are worth keeping straight, because they answer slightly different questions:
| Use... | When |
|---|---|
| The Disk Performance Test dialog | You want to pick specific disks, choose a block size, or use the Group by VDEV / Group by Enclosure modes to find a shared bottleneck. It also stores each result on the disk in the Read Seq and Last Performance Test columns, so it is the right tool for taking a baseline you want to compare against months later. It is the only one of the two that offers the grouped scheduling modes. |
qs-perftest |
You are already at a console or in an SSH session, you want the whole system read-tested in one command, or you want the result in a log file you can attach to a support case. It is also the only path to the pool-level and share-level tests further down this page. |
Both are read-only and safe to run on a pool that is in production, though they add load while they run.
Step 2: rule out the network
If the media is performing and clients are still slow, check the network next -- before touching the pool. Misconfigured network settings are one of the most common real-world causes, and they are easy to get wrong in ways that leave every status indicator green.
The three that account for most cases:
- A bond the switch is not configured for. LACP requires a managed switch with LACP configured on the matching ports; round-robin and balance-xor require an Etherchannel-capable switch. Build an LACP bond against a switch that is not running LACP and you get a link that works and carries a fraction of the throughput. Network Ports lists each bond mode with the switch it needs.
- A second gateway competing for the default route. QuantaStor gives every static port the same route metric, so configuring a gateway on more than one port does not create a fallback -- it creates two routes of equal cost, and return traffic can leave the wrong port. Configure a gateway on one port only. See the static route management section of Network Ports.
- An MTU mismatch. Jumbo frames have to be set consistently on the initiator, every switch port in the path, and the appliance. A device that receives an oversized frame drops it, so the symptom is a link that passes small transfers and stalls on large ones.
Run the Network Check report from Health Checker to test all three at once: it checks reachability from each port, reports round-trip time per destination, and tests whether jumbo frames get through end to end. The CLI equivalent is qs net-check. See Network Ports and Network Check View.
The Network chart on the Storage System dashboard is per port, which is the quickest way to see whether traffic is actually spread across the ports you configured or has quietly collapsed onto one. See Performance Monitoring and Performance Tuning.
Step 3: take the media out of the equation
The read tests in step 1 can tell you a disk is slow. They cannot tell you that the media is not the problem -- a healthy-looking set of disks does not prove the pool built on them is performing, and it says nothing at all about the transport between the appliance and the client.
qs-ramdisk closes that gap. It creates transient RAM-backed SCSI disks, which appear as ordinary Physical Disks and can be exported to an initiator as pass-thru Storage Volumes. Benchmark one of those from the client and the media is factored out of the measurement entirely. If the RAM-backed device is also slow, the bottleneck is above the media -- the transport, the CPU, or the target stack -- and no amount of work on the pool will help.
That is a negative you cannot prove any other way, and it is what makes the tool worth the trouble.
They also hold real memory. The backing store is non-swappable kernel memory and it is not released until you run qs-ramdisk destroy, so destroy them as soon as the test is finished.
The workflow
sudo qs-ramdisk create 8G
qs disk-list # find the OSNEXUS QS-RAMDISK entry
qs volume-create-passthru --name=ramtest --disk-list=<disk>
# assign to a host, then benchmark from the client
qs volume-delete --volume-list=ramtest --flags=force
sudo qs-ramdisk destroy
Creating the devices reports the memory it is about to take and asks for confirmation:
# sudo qs-ramdisk create 2G --count 2 INFO: ZFS ARC max is 5429MB; a pool built on a RAM disk is cached in the ARC INFO: as well, so expect roughly double the memory cost of the requested size. INFO: Requested 4096MB, 6270MB available, 4180MB budget. About to create 2 RAM disk(s) of 2048MB each (4096MB total). Sector size: 512 Synthetic latency: 0ns These devices are VOLATILE. All data on them is lost on 'qs-ramdisk destroy' and on reboot. They are for performance measurement only -- do not place customer data on them and do not build production pools on them. Continue? [y/N] y INFO: Loading scsi_debug... INFO: Verified that each RAM disk has its own independent backing store. INFO: Created 2 RAM disk(s). DEVICE FRIENDLY SIZE(MB) STATUS BY-ID (as QuantaStor sees it) /dev/sdh /dev/qs-ramdisk0 2048 free /dev/disk/by-id/scsi-SOSNEXUS_QS-RAMDISK_8000 /dev/sdi /dev/qs-ramdisk1 2048 free /dev/disk/by-id/scsi-SOSNEXUS_QS-RAMDISK_10000
Two things in that output are worth understanding. The memory budget line is a guard rail: the tool refuses to allocate more than two thirds of available memory, because the backing store is kernel memory that cannot be swapped. And the note about the ARC is real -- a pool built on a RAM disk is cached in RAM a second time, so it can cost roughly twice the size you asked for.
The devices then appear as normal Physical Disks, with vendor OSNEXUS and product QS-RAMDISK:
qs-node-141 sdh (scsi-1OSNEXUSQS-RAMDISK_8000) 2.15GB (2.0GiB) OSNEXUS QS-RAMDISK qs-node-141 sdi (scsi-1OSNEXUSQS-RAMDISK_10000) 2.15GB (2.0GiB) OSNEXUS QS-RAMDISK
Allow up to a minute for them to appear -- QuantaStor rescans on its own schedule, and qs disk-scan will hurry it along.
From there, qs volume-create-passthru --name=<name> --disk-list=<disk> turns the device into a raw pass-thru Storage Volume with no partitioning and no pool in the path, which you then assign to a host exactly like any other volume. That is the configuration to benchmark: it exercises the initiator, the network and the SCSI target stack with nothing but RAM underneath. See Physical Disks/Devices for pass-thru volumes in general.
Compare the result against the same benchmark run over a pass-thru volume on a real disk, or against a Storage Volume on the pool. A RAM disk that is only marginally faster than the pool means the pool was never the limit.
Quantifying transport latency with --ndelay
--ndelay is the most useful option in the tool. It adds a precise synthetic per-command latency, in nanoseconds, to every SCSI command the device services. That turns the RAM disk from a fast device into a calibrated one: you can inject a known latency and watch what it does to throughput and IOPS, which gives you a scale to measure your transport's contribution against instead of just observing that it is slow.
A worked example. The same fio job -- 4 KB random reads, queue depth 1, direct, against the raw device -- run twice, once with no synthetic latency and once with 250 microseconds:
sudo qs-ramdisk create 1G -y
sudo fio --name=lat --filename=/dev/sdh --rw=randread --bs=4k --direct=1 \
--ioengine=psync --numjobs=1 --runtime=10 --time_based --group_reporting
| Mean latency | IOPS | |
|---|---|---|
--ndelay 0 (default) |
66.4 us | 14,500 |
--ndelay 250000 (250 us) |
341.5 us | 2,896 |
The added latency lands where it was asked for -- 341.5 minus 66.4 is 275 us for a 250 us request, the remainder being scheduling overhead -- and it costs four fifths of the IOPS. Two things to take from that:
- Per-command latency, not bandwidth, is what limits small-block IOPS. Nothing about the media changed between those two runs. A quarter of a millisecond per command was enough to take a device from 14,500 IOPS to 2,900.
- You can now put a number on your transport. Benchmark a pass-thru volume on an
--ndelay 0RAM disk from the client, then locally on the appliance. The difference is what the network and the target stack are costing you. Sweep--ndelayacross a few values and you have a curve to place that difference on -- which tells you whether the overhead you are seeing is a few tens of microseconds of normal stack cost or a few hundred that wants investigating.
Baseline behaviour is worth stating explicitly: with --ndelay 0 the device answers immediately, and any latency you measure is the stack above it, not the device.
Cleaning up
Destroy the devices when the test is done. destroy removes all of them -- there is no per-device destroy:
# sudo qs-ramdisk destroy INFO: Unloading scsi_debug... INFO: RAM disks destroyed and memory released. # qs-ramdisk list No RAM disks exist. Create one with 'qs-ramdisk create <SIZE>'.
Delete the pass-thru volume first. While a Storage Volume is still exported from a RAM disk, the kernel holds a reference to the device and the destroy fails:
INFO: Unloading scsi_debug...
ERR: Failed to unload scsi_debug. Something still holds a reference to the devices;
run 'qs-ramdisk list' to see what, or check 'lsof' and 'dmesg'.
Remove whatever is using the device and re-run destroy:
qs volume-delete --volume-list=<name> --flags=force sudo qs-ramdisk destroy
Always confirm qs-ramdisk list reports no RAM disks before you finish. Until it does, the memory is still allocated and the appliance is running with less RAM than it should have -- which will itself look like a performance problem.
Step 4: test the pool and the filesystem path
With the media and the network accounted for, qs-perftest benchpool measures the pool the way an application sees it. It creates temporary Network Shares on the pool and benchmarks those, so the measurement includes the filesystem, the record size, compression, the write path and the caches -- everything the raw device tests deliberately skipped.
sudo qs-perftest benchpool doctest-perf-pool 1 --tests fio,dd,elbencho
{Thu Sep 3 05:20:15 2026, INFO, py:qs_perftest} Detected filesystem type 'zfs' for pool doctest-perf-pool.
{Thu Sep 3 05:20:15 2026, WARN, py:qs_perftest} perf wrapper present but no build for kernel 6.8.0-90-generic; install linux-tools-6.8.0-90-generic to enable profiling.
{Thu Sep 3 05:20:15 2026, INFO, py:qs_perftest} Creating ZFS throughput share: 5e1f93044dab40e4-1m
{Thu Sep 3 05:20:17 2026, INFO, py:qs_perftest} Creating ZFS random-IO share: 5e1f93044dab40e4-4k
{Thu Sep 3 05:20:21 2026, INFO, py:qs_perftest} Setting recordsize=4K on qs-596a6bb7-.../5e1f93044dab40e4-4k
{Thu Sep 3 05:20:26 2026, INFO, py:qs_perftest} fio (throughput) on /export/5e1f93044dab40e4-1m: starting: ...
{Thu Sep 3 05:20:35 2026, INFO, py:qs_perftest} fio (random) on /export/5e1f93044dab40e4-4k: starting: ...
{Thu Sep 3 05:21:57 2026, INFO, py:qs_perftest} dd write on /export/5e1f93044dab40e4-1m: starting: ...
{Thu Sep 3 05:22:03 2026, INFO, py:qs_perftest} elbencho on /export/5e1f93044dab40e4-1m: starting: ...
{Thu Sep 3 05:22:07 2026, INFO, py:qs_perftest} Results saved in /var/log/qs/qs_perf_results_5e1f93044dab40e4_20260903_052207.log
{Thu Sep 3 05:22:07 2026, INFO, py:qs_perftest} Removing temporary benchmark share(s): 5e1f93044dab40e4-1m,5e1f93044dab40e4-4k
{Thu Sep 3 05:22:12 2026, INFO, py:qs_perftest} Benchmarking complete.
What that run did, and why:
- It created two shares on a ZFS pool, not one. The throughput share is created with a 1 MB record size and the random-I/O share with 4 KB, because a single record size cannot represent both workloads -- see the record size discussion in Performance Tuning. The 4 KB share is only created when the fio test is selected, since it is the only test that uses it. A Ceph pool gets one share.
- The shares are real Network Shares, named from a random hex string, and they occupy pool capacity while the run is in progress. They are visible in the Network Shares section, and they are deleted at the end of the run -- including when a test fails part way through.
- It skips a benchmark whose tool is missing rather than aborting, with a
WARNper skipped test, and only fails if none of the requested tools are installed. - It notes whether the Linux
perfprofiler is usable for follow-on CPU profiling. The warning above is normal and harmless:perfships as a kernel-version-specific binary and the matching package is not installed by default.
The four benchmark engines
| Test | What it runs | Measures |
|---|---|---|
fio |
Two passes on the 1 MB and 4 KB shares. Throughput: mixed sequential read/write, 1 MB blocks, one job, 60 second cap. Random: mixed random read/write, 4 KB blocks, eight concurrent jobs, 60 second cap. | The most representative of the four. The random pass with eight jobs is the one that resembles a virtual machine or database workload. |
elbencho |
Sequential write then read, four threads, 1 MB blocks. | A second opinion on large-block throughput, with per-thread first-done and last-done timings that expose an uneven pool. |
dd |
Single-stream write with conv=fdatasync, then a read back. |
A simple sequential floor. Useful as a sanity check, not as a benchmark. |
iozone |
Automatic mode sweep across record sizes and file sizes, read/write/random tests. | A throughput surface across block sizes. Thorough and slow -- expect it to dominate the runtime of an --tests all run.
|
Interpreting benchpool output
Real results from the run above, on a two-disk mirror:
--- fio throughput test --- READ: bw=229MiB/s (241MB/s), io=496MiB, run=2162-2162msec WRITE: bw=244MiB/s (256MB/s), io=528MiB, run=2162-2162msec --- fio random test --- read: IOPS=14.5k, BW=56.6MiB/s write: IOPS=14.5k, BW=56.6MiB/s --- dd test --- write: 1073741824 bytes (1.1 GB, 1.0 GiB) copied, 5.32025 s, 202 MB/s read: 1073741824 bytes (1.1 GB, 1.0 GiB) copied, 0.272778 s, 3.9 GB/s --- elbencho test --- OPERATION RESULT TYPE FIRST DONE LAST DONE =========== ================ ========== ========= WRITE Throughput MiB/s : 290 247 READ Throughput MiB/s : 6215 5386
Three caveats matter more than any of those numbers, and none of them are obvious from the output:
The read figures are cache reads, not media reads. Every engine reads back a file it has just written, and nothing drops the cache in between, so the read is served from the ARC. That is why dd reports 3.9 GB/s reading from a mirror of two disks that read at 1.5 GB/s each, and why elbencho reports over 5 GB/s. Those are valid measurements of the RAM cache and meaningless as measurements of the pool. Compare write figures between runs, and treat the read figures as an ARC check. For a genuine pool read rate, use the step 1 device tests, or read a file large enough to exceed the ARC.
The write figures are optimistic on a compressed pool. Compression is on by default, and all four engines write compressible data -- dd writes zeros. The same dd command from the run above, repeated a few minutes later on the same pool, reported 1.4 GB/s instead of 202 MB/s; nothing about the pool changed. If you need a figure that reflects real media throughput, run the benchmark against a share with compression turned off, or use data that does not compress.
fio's 60-second cap is a cap, not a duration. The throughput pass above finished in 2.2 seconds because the file size was reached first. A 1 GB test on a fast pool measures a couple of seconds of I/O, most of which lands in the write cache. Use a size several times larger than the ARC for a result that means anything -- the size argument is in GB and there is no reason not to use 64 or 128 on a real appliance.
With those in mind, what the output is good for:
- Comparing the same pool before and after a change. This is its strongest use. Same size, same tests, same engines.
- Comparing two pools on the same appliance, which controls for everything except the pool layout.
- The gap between the throughput and random passes, which tells you how the pool handles small random I/O. A pool that does 244 MiB/s sequentially and 56 MiB/s at 4 KB is behaving normally; one that collapses to a few MiB/s at 4 KB wants more device groups, mirroring rather than wide parity, or metadata and small-block offload -- see Storage Pools and Performance Tuning.
- elbencho's first-done against last-done, where a wide gap means the threads finished at very different times, which points at an uneven pool or a contended device.
Machine-readable output
--json emits a single JSON document on stdout and suppresses the progress lines, so the output can be parsed:
sudo qs-perftest benchpool doctest-perf-pool 1 --tests dd --json
{
"operation": "benchpool",
"pool": "doctest-perf-pool",
"fstype": "zfs",
"size_gb": 1,
"tests": [
"dd"
],
"results_file": "/var/log/qs/qs_perf_results_e1b251a289234114_20260903_052338.log",
"results": {
"dd": "write: 1073741824 bytes (1.1 GB, 1.0 GiB) copied, 0.765367 s, 1.4 GB/s\nread: 1073741824 bytes (1.1 GB, 1.0 GiB) copied, 0.368255 s, 2.9 GB/s"
}
}
The fio results are embedded as parsed JSON when fio is among the selected tests, because fio is asked for JSON output in this mode. The other engines have no structured output and appear as their raw text.
The pool-level dd test
rwpool is a simpler, older pool test that writes a file straight into the pool's mount point and reads it back, without creating a share:
sudo qs-perftest rwpool qs-596a6bb7-b2c7-d283-1951-00efdd93ece3
{Thu Sep 3 05:19:34 2026, INFO, py:qs_perftest} Generating 2GB random data in RAM disk (/run/quantastor/perftest.data) for performance test usage.
{Thu Sep 3 05:19:41 2026, INFO, py:qs_perftest} WRITE-POOL-FILE (qs-596a6bb7-...): completed: 827969536 bytes (828 MB, 790 MiB) copied, 3.83234 s, 216 MB/s
{Thu Sep 3 05:19:42 2026, INFO, py:qs_perftest} READ-POOL-FILE (qs-596a6bb7-...): completed: 827969536 bytes (828 MB, 790 MiB) copied, 0.603414 s, 1.4 GB/s
Its one real advantage over benchpool is the source data: it generates an incompressible file first, using openssl to encrypt a stream of zeros, so the write figure is not inflated by compression the way the benchpool dd test is. Two limits to know about:
- The source file is built in
/run, which is a RAM disk sized at a fraction of system memory. On the appliance above,/runis 795 MB, so the intended 2 GB file is truncated to 790 MiB and the test writes that instead -- which is what the byte counts in the output show. The figure is still valid for the amount of data actually written; just read the byte count rather than assuming 2 GB. - The read-back is an ARC read, for the same reason as
benchpool-- 1.4 GB/s here against a 216 MB/s write.
rwpools runs it across every pool on the system, and runall combines that with readpooldisks for each pool, which is a reasonable one-command sweep to attach to a support case. Both write into every pool on the appliance, so do not run them on a system whose other pools you do not own.
qs-perftest operation reference
qs-perftest is a Python utility at /usr/bin/qs-perftest. It takes one operation plus that operation's arguments, and every operation needs root. Run qs-perftest <operation> -h for a single operation's arguments.
| Operation | Arguments | What it does |
|---|---|---|
interactive |
none | Guided chooser. Lists the pools and disks from the service so you do not have to look up an id, then prompts for the test and its settings. This is also what runs when the tool is given no operation at all. It needs a terminal -- run non-interactively it reports Interactive mode requires a terminal; specify an operation instead and exits.
|
benchpool |
<pool_id> [size] plus --tests, --fstype, --json |
Creates temporary share(s) and runs the selected benchmarks. pool_id is a pool name or qs-UUID. size is the test file size in GB, default 1.
|
readdisks |
none | Sequential dd read test on every disk in the system.
|
readdisk |
<device> |
The same test on one device, given as a path such as /dev/sdb.
|
rwpool |
<poolid> |
Writes an incompressible file into the pool and reads it back. Requires the qs-UUID form, not the pool name.
|
rwpools |
none | rwpool across every pool on the system.
|
readpooldisks |
<poolid> |
Read test on the disks of one pool. Requires the qs-UUID form.
|
runall |
none | rwpools plus readpooldisks for every pool.
|
Which form of the pool identifier is a real difference and easy to trip over: benchpool accepts a friendly pool name, while rwpool and readpooldisks want the qs-UUID as it appears in zpool list and under /mnt/storage-pools/. Get the id with:
zpool list -H -o name
benchpool options:
| Option | Default | Notes |
|---|---|---|
--tests |
all |
Comma-separated list from fio, elbencho, dd, iozone, or all. An unknown name is rejected. Drop iozone to keep a run short.
|
--fstype |
auto-detect | zfs or ceph. Detected from /proc/mounts and then from qs pool-list; only needed when detection fails, in which case the tool tells you to pass it.
|
--json |
off | Emit a JSON results document on stdout and keep the progress lines out of it. |
qs-ramdisk operation reference
qs-ramdisk at /usr/bin/qs-ramdisk creates the transient RAM-backed SCSI disks used in step 3. create and destroy need root; list does not.
| Operation | Arguments | What it does |
|---|---|---|
create |
<SIZE> plus --count, --sector-size, --ndelay, -y |
Creates RAM-backed disk(s). SIZE takes an M or G suffix and is the size of each device.
|
list |
none | Lists the RAM disks with their size, whether anything is using them, and the /dev/disk/by-id path.
|
destroy |
--force, -y |
Destroys all RAM disks and frees the memory. |
create options:
| Option | Default | Notes |
|---|---|---|
--count N |
1 | Number of devices, up to 32. Each gets its own independent backing store, and the tool verifies that before handing them over. |
--sector-size |
512 | 512 or 4096. Use 4096 to reproduce the behaviour of 4K-native media, including how a pool's ashift is derived.
|
--ndelay NS |
0 | Synthetic per-command latency in nanoseconds; 0 means answer as fast as possible. See Quantifying transport latency. |
-y / --yes |
off | Skip the confirmation prompt. Useful in a script; the memory budget check still applies. |
Practical constraints, all of which the tool reports clearly rather than failing obscurely:
- The settings are global to the whole batch.
createrefuses to run while RAM disks already exist, because the size, sector size and latency are kernel module parameters and cannot be changed while it is loaded. To change any of them,destroyfirst. - The memory budget is two thirds of available memory. A larger request is refused with the numbers it based that on.
destroytakes all of them or none. There is no per-device destroy.destroyrefuses while a device is mounted, in a ZFS pool, mapped out over SCST or backing a pass-thru volume, and tells you which.--forceoverrides that and leaves whatever was using the device referencing a device that no longer exists, so it is a last resort rather than a shortcut.
What backs the devices
The devices are backed by the kernel's scsi_debug module, which emulates a complete SCSI target, rather than by a plain RAM block device. That choice is what makes the tool useful: an emulated SCSI device answers the INQUIRY and vital-product-data pages that a real disk does, so udev creates the /dev/disk/by-id/scsi-* symlink that QuantaStor's Physical Disk enumeration works from. A plain RAM disk has no such identity and would never appear in the disk list at all.
Two consequences are visible from the outside. The devices show up in Physical Disks and in qs disk-list with no configuration, and they can be exported as pass-thru Storage Volumes like any other disk. Each device also gets a stable friendly symlink at /dev/qs-ramdisk0, /dev/qs-ramdisk1 and so on, which is more convenient than a kernel name that moves between reboots.
Where results are written
| File | Contents |
|---|---|
/var/log/qs/qs_perftest.log |
Every line qs-perftest prints, appended across runs. The history of what was measured and when.
|
/var/log/qs/qs_service.log |
The same lines, interleaved with the service log, so a test can be correlated with what the appliance was doing at the time. |
/var/log/qs/qs_perf_results_<id>_<timestamp>.log |
One file per benchpool run: the pool and its filesystem type, the tests run, a zpool status -DvP (or ceph status) capture, and the full output of every engine. The path is printed at the end of the run.
|
The results file's pool status capture covers every pool on the appliance, not only the one tested, which is usually helpful context when the file goes to support.
These files are collected by Send Support Logs, so a benchpool run before opening a case gives support a measurement to work from rather than a description. See Storage System.
Related pages
- Performance Tuning -- what to change once you know which layer is slow: I/O profiles, system tunables, ARC, L2ARC, write log, record size and compression
- Storage System Optimization -- the appliance-wide OpenZFS and kernel tunables
- Physical Disks/Devices -- device health and SMART, the Disk Performance Test dialog, and pass-thru Storage Volumes
- Network Ports -- bond modes, static routes, MTU, and the Network Check report
- Network Check View -- the connectivity and jumbo-frame report in detail
- Storage Pools -- pool layout, device groups and the offload tiers that benchpool is measuring
- Network Shares -- what benchpool creates and deletes while it runs
- Performance Monitoring -- the Grid Dashboard
- QuantaStor CLI Command Reference -- full argument lists for the
qscommands above
Verified against QuantaStor 6.9.0.