Performance Tuning

From OSNEXUS Online Documentation Site
Revision as of 05:14, 3 September 2026 by Qadmin (talk | contribs) (Tighten the I/O profile re-application claim to the verified set of triggers (QSTOR-12352))
Jump to navigation Jump to search


This page covers how to tune a scale-up (OpenZFS) QuantaStor appliance for a workload: what to measure, what each tuning control actually changes underneath, and how to choose between them. It is a decision guide rather than a field reference -- where a control lives on another page's dialog, the reasoning is here and the field-by-field detail is linked.

Two pages own the controls this one reasons about. Storage System Optimization owns the appliance-wide OpenZFS and kernel tunables dialog. Storage Pools owns pool creation and modification, including the I/O profile, compression, sync and cache policy fields.

Section Purpose
How to approach a tuning problem The order of operations, and what not to do.
What to measure The dashboards, qs-iostat, and the physical disk read test.
Where the knobs are Five separate tuning layers and which one a given problem belongs to.
Storage Pool I/O profiles Read-ahead, request queue depth, I/O scheduler, and the seven shipped profiles.
Storage Volume I/O profiles Volume block size and the iSCSI session parameters.
RAM read cache (ARC) Sizing the ARC and what it costs in RAM.
SSD read cache (L2ARC) When a read cache pays for itself and when it is wasted.
Write log (SLOG/ZIL) and sync policy Why a write log often sits idle, and what changes that.
Metadata, small-block and dedup offload The single largest win available on an HDD pool.
Record size and volume block size Matching the on-disk unit to the application's I/O size.
Compression Why compression is usually a throughput decision, not a capacity one.
Device write caches Why an HDD in a pool is slower than its data sheet.
Network-side factors MTU, multiple interfaces versus bonding, and NIC ring buffers.
What does not apply Controls that look relevant and are not.

How to approach a tuning problem

The shipped defaults are right for the large majority of deployments, and most disappointing performance turns out to be a design problem rather than a tuning problem: too few device groups, a parity stripe that is too wide for a random workload, HDDs with no metadata offload, or a single network path. Tuning cannot fix any of those. Work through the design first -- see Where the knobs are and Storage Pools -- and only then reach for a knob.

When you do tune:

  1. Measure first, with a workload that resembles production. A number you cannot reproduce is not a baseline, and without a baseline you cannot tell an improvement from noise.
  2. Confirm the hardware is healthy. Every device in Physical Disks should be Normal. One disk with a predictive-failure warning will make a whole pool look badly tuned. See Physical Disks/Devices.
  3. Change one group of related settings, then measure again. Several of the OpenZFS tunables interact, and a batch of simultaneous changes leaves you with no way to attribute the result.
  4. Save a named profile before experimenting so you can get back to a known state. See Storage System Optimization for tunable profiles and Custom I/O profiles below for pool profiles.
  5. Write down what you changed. Tuning applied per system in a grid drifts silently between members otherwise.

We recommend involving OSNEXUS support before changing the system tunables. They are the settings that matter for specific, identifiable problems -- write latency spikes under load, a resilver that will not finish inside a maintenance window, a replication job overrunning its window -- and changing them speculatively is more likely to cost performance than gain it.

What to measure

The Storage System dashboard

Navigation: Storage Management → Storage Systems → select a Storage System → Performance (dashboard tab)
The Performance view of the Storage System dashboard: CPU (including Wait-IO), per-port network throughput, and memory.

Selecting a Storage System puts a dashboard above the object grid with three views and a time-range selector. Performance plots CPU, network and memory. The CPU chart separates Wait-IO from System and Idle, which is the quickest way to tell a storage bottleneck from a CPU bottleneck: high Wait-IO with idle CPU means the appliance is waiting on media, and no amount of thread tuning will help. The Network chart is per port, so it also shows whether traffic is actually spread across the ports you configured -- a common surprise on a multipath configuration that has quietly collapsed onto one path.

Sensors plots average CPU temperature and system power draw, which are worth a glance when throughput drops under sustained load -- a thermally throttled CPU looks like a storage problem.

Cache Stats (ARC)

Navigation: Storage Management → Storage Systems → select a Storage System → Cache Stats (ARC) (dashboard tab)
Cache Stats (ARC): hit ratio, hit counts broken out by most-recently-used and most-frequently-used, and the ARC's current, target, maximum and minimum size.

The Cache Stats (ARC) view is where an ARC sizing decision is made or unmade:

  • Cache Efficiency is the hit ratio. A high ratio under representative load means the working set fits in RAM and more RAM buys you nothing. A low ratio means the working set does not fit, which is the argument for more RAM or an SSD read cache.
  • Cache Hits splits hits into Most Recently Used and Most Frequently Used. A workload served almost entirely from the most-recently-used side is streaming through the cache rather than reusing it, and that workload will not benefit from a larger cache -- it is a candidate for a metadata cache policy instead, so it stops evicting everything else.
  • Cache Size plots the current size against the target (which OpenZFS adapts continuously), the maximum, and the minimum. If the current size sits well below the maximum under load, the ARC is not the constraint.

qs-iostat

qs-iostat is the console-level counterpart, and it reports the ARC and ZIL kernel counters that the charts summarise. It wraps iostat:

qs-iostat -c            # globally averaged CPU stats
qs-iostat -d            # I/O stats for all block devices, in MB/s
qs-iostat -a            # ZFS ARC, L2ARC and ZIL counters
qs-iostat -f            # repeat every 2 seconds
qs-iostat --extra "-x"  # pass extra arguments through to iostat

qs-iostat -a is the useful one for cache work. The counters are cumulative since boot, so take two samples and compare, rather than reading absolute numbers:

Counter Meaning
hits / misses Read requests satisfied from, and missed in, the RAM cache. The ratio between two samples is the live hit rate.
size Current ARC size.
c_min / c_max The floor and ceiling the ARC may grow between. c_max is what Cache Size (% of RAM) sets.
arc_meta_used How much of the ARC is metadata rather than data. On a pool with many small files this can be most of it.
l2_hits / l2_misses Reads satisfied from, and missed in, the SSD read cache. Both zero means the L2ARC is doing nothing.
l2_size How much of the read cache is actually populated. A cache that never fills is larger than the workload needs.
l2_hdr_size RAM consumed by the L2ARC's index. This is the hidden cost of a read cache -- see SSD read cache (L2ARC).
l2_read_bytes / l2_write_bytes Bytes read from and written to the read cache devices.
l2_cksum_bad Checksum failures on a read cache device. A rising count usually means the SSD is failing and should be replaced.
zil_commit_count Synchronous write commits since boot. Zero, or near zero, on a pool with a write log means the log is idle -- see Write log (SLOG/ZIL) and sync policy.
zil_commit_writer_count Commits that had to write, as opposed to joining a commit already in flight.
zil_commit_stall_count / zil_commit_suspend_count Commits that stalled or were suspended. Anything other than zero under load points at the log device not keeping up.
zil_itx_count Intent log transactions recorded.

(Note: these counters come straight from the OpenZFS kernel modules and the set changes between OpenZFS releases; /proc/spl/kstat/zfs/arcstats and /proc/spl/kstat/zfs/zil on the appliance are authoritative.)

The physical disk performance test

Navigation: Storage Management → Physical Disks → select + right-click a disk → Disk Performance Test...
Physical Disk Performance Test. The Read Seq and Last Performance Test columns keep the result on each disk, so a later run can be compared against it.

This is a sequential read test against the raw devices, which makes it the right tool for one specific question: is one device in this pool slower than its peers? It bypasses the pool, so it tells you nothing about pool or protocol performance -- but a single disk reading at a fraction of the rate of identical siblings is the most common cause of inconsistent pool performance, and this finds it in a couple of minutes.

The dialog is on the disk's right-click menu, not the toolbar. It opens filtered to the disk you right-clicked; click Reset to clear the filter and select across all of them.

Field Default Notes
Mode Serial How the selected disks are scheduled. Serial tests one at a time and gives each disk the whole controller and backplane to itself, which is what you want when comparing disks against each other. Parallel tests all of them at once in batches, which instead measures how much aggregate bandwidth the path can carry. Group by VDEV tests each pool device group's disks in parallel, one group at a time; Group by Enclosure does the same per enclosure, which is how you find a JBOD or an expander link that is the real limit.
Block Size 1 MB Read size, from 512 bytes to 1 GB. Larger blocks give a better estimate for backup and archive use cases; small blocks say more about seek behaviour.
Block Count 1000 Number of reads. Block size multiplied by block count is the total read per disk, so the defaults read 1 GB from each. Range 1 to 1,000,000.
Storage System the selected system Restricts the disk list to one appliance.

Results are written back onto each disk as Read Seq (average throughput) and Last Performance Test (timestamp), and stay there, so the test doubles as a record of what a disk was capable of when it was healthy.

The same test from the CLI, where the disk list filters make it easy to test exactly one pool or every unused disk:

qs disk-performance-test --disk-list=[unused] --perf-mode=serial
qs disk-performance-test --disk-list=<pool-name> --perf-mode=group-by-vdev --block-size=1M --block-count=1000

qs disk-performance-test accepts --disk-list filters including [all], [unused], [gt:SIZE] and a pool or system name, plus --perf-test-type, --block-size, --block-count and --perf-mode. Add --flags=async to queue the run and watch it in the task list instead of blocking.

Per-disk and grid views

Selecting an individual disk replaces the dashboard with a Physical Disk Dashboard plotting IOPS, throughput and latency for that device -- useful for confirming that load is landing where you expect during a test.

The Grid View tab carries per-system CPU and memory sparklines alongside the pool capacity and alert summaries, which makes it the fastest way to see which member of a grid is busy. See Performance Monitoring.

Where the knobs are

Five separate mechanisms are easy to confuse. Deciding which layer a problem belongs to saves most of the work:

Layer Scope What it controls Where
Storage Pool I/O profile one pool, re-applied at every pool start Block device settings on the member disks: read-ahead, request queue depth, I/O scheduler. Also raises the SCST target driver thread floor for the whole appliance. below; the field is on Create and Modify Storage Pool (Storage Pools)
Pool, share and volume properties one pool, share or volume Record size and volume block size, compression, sync policy, primary and secondary cache policy, small-block offload. below; the fields are on Storage Pools, Network Shares and Storage Volumes
Storage Volume I/O profile one volume Volume block size and the iSCSI session parameters on that volume's target. below
System tunables one Storage System OpenZFS and SPL module parameters plus two network queue settings: ARC size, write throttle, per-VDEV queue depths, resilver and scrub priority, prefetch distance. Storage System Optimization
Pool topology and offload tiers one pool, fixed at build time (mostly) Number and width of device groups, mirroring versus parity, and the write log, read cache, metadata and dedup vdevs. Storage Pools

The last row is not a tuning layer, but it is where the large factors live. A pool with one wide parity group and no flash offload cannot be tuned into a random-I/O pool, and the first four layers together will not move it as far as adding device groups will.

Storage Pool I/O profiles

A Storage Pool I/O profile is a bundle of Linux block device settings applied to the pool's member disks. It is not a filesystem or OpenZFS setting. QuantaStor re-applies it whenever the pool is created or started, on HA failover, when the profile is changed on Modify Storage Pool, and whenever cache, log, metadata or hot-spare devices are added or removed -- so the tuning survives the events that would otherwise leave new or reconnected devices on kernel defaults. Each pool has one profile; the field is on both Create Storage Pool (Advanced Settings) and Modify Storage Pool -- see Storage Pools for the dialogs.

Each profile carries a separate set of values for HDD, SSD and NVMe media, and QuantaStor picks the set per device from what the device reports, so one profile tunes a mixed pool correctly. On a multipath pool the settings are applied to each underlying SCSI path rather than to the dm device, which is where they actually take effect.

What each setting changes

Read-ahead (/sys/block/<dev>/queue/read_ahead_kb) is how much extra data the kernel pulls in after satisfying a read. If a read asks for 4 KB and read-ahead is 256 KB, another 256 KB is read into the page cache behind it.

Read-ahead exists to hide rotational latency. A 7200 RPM drive turns 120 times a second, roughly once every 8 ms. Reading 4 KB records one per rotation yields under 500 KB/s -- so on any workload with sequential locality, reading ahead and having the next records already in memory is the difference between a usable HDD pool and an unusable one. That is why read-ahead and queue depth matter so much on spinning media and why the archive and media profiles push read-ahead up to 512 KB and 1 MB.

Flash has no rotational latency, so read-ahead on flash is pure overhead: every speculative block occupies bandwidth and cache that a real request wanted. Every shipped profile uses a 4 KB SSD read-ahead, and that is a deliberate result rather than a rounding to zero -- the original SSD tuning work found that both 0 KB and larger values produced pronounced dips at 4 KB and 8 KB request sizes that a small 4 KB read-ahead removed. It remains the single most important SSD-side setting in these profiles.

Request queue depth (/sys/block/<dev>/queue/nr_requests) is how many I/O requests the kernel will hold queued for a device. A deeper queue gives the scheduler and the device more requests to sort and coalesce, which raises throughput on rotational media; too deep a queue raises latency, because a small urgent read now waits behind a long queue of other work.

Profiles express this two ways. A fixed count sets the queue depth outright. A multiplier instead reads the device's own hardware queue depth from /sys/block/<dev>/device/queue_depth and multiplies it, so the same profile scales sensibly across drives with very different capabilities. Two rules matter:

  • The multiplier takes precedence over the fixed count whenever the device reports a non-zero hardware queue depth. The fixed count is the fallback for devices that report none.
  • The computed value is capped at 1024. A multiplier against a deep hardware queue is clamped there, with a warning in /var/log/qs/qs_service.log.

I/O scheduler (/sys/block/<dev>/queue/scheduler) is the elevator algorithm that decides the order requests reach the device. Two are relevant. deadline sorts requests to reduce head movement while enforcing a latency deadline so nothing starves -- the right choice for HDDs. noop does essentially no reordering and hands requests straight down, which is right for flash, where there is no seek cost to optimise away and reordering only adds CPU work and latency. Current kernels name these mq-deadline and none; QuantaStor accepts either spelling in a profile and writes whichever the device offers. If a device offers neither, the setting is skipped and logged rather than failing the pool start.

Two further settings appear in profiles:

  • FIFO batch (/sys/block/<dev>/queue/iosched/fifo_batch) tunes how many requests the deadline scheduler dispatches in one batch. It only exists while a deadline-style scheduler is selected, so it is skipped on any device running noop/none. None of the shipped profiles set it.
  • Target driver threads and tasklets set a floor on the SCST target driver's worker thread and tasklet counts, which is what serves iSCSI, FC and NVMe-oF. These are appliance-wide, not per pool: QuantaStor takes the highest value across every profile in use on the system, and can only raise the count above SCST's own start-up default (derived from the core count), never lower it. Both are capped at 256. Assigning the Virtualization profile to one pool therefore raises the thread floor for every target on that appliance.

QuantaStor also raises max_sectors_kb to 4096 on each member device where the hardware allows it, independently of the profile, so that large sequential I/O is not split into smaller commands.

The shipped profiles

Seven profiles ship with the appliance. Values below are read from /opt/osnexus/quantastor/conf/qs_io_profiles.conf on a 6.9.0 appliance.

Profile HDD read-ahead HDD queue depth HDD scheduler SSD read-ahead SSD queue depth SSD scheduler Target driver threads / tasklets
Default 256 KB 2 x hw queue depth deadline 4 KB 1 x hw queue depth noop SCST default
Disk Archive 512 KB 2 x hw queue depth deadline 4 KB 1 x hw queue depth noop SCST default
Edgeware IP TV Optimized 256 KB 64 deadline 4 KB 1 x hw queue depth noop SCST default
Edgeware Web TV Optimized 256 KB 64 deadline 4 KB 1 x hw queue depth noop SCST default
Media Post-Production (Ingest Optimized) 512 KB 3 x hw queue depth deadline 4 KB 1 x hw queue depth noop SCST default
Media Post-Production (Playback Optimized) 1024 KB 2 x hw queue depth deadline 4 KB 1 x hw queue depth noop SCST default
Virtualization 128 KB 64 deadline 4 KB 1 x hw queue depth noop 64 / 32

Where a multiplier is shown, the profile also carries a fixed fallback count -- 256 for every profile except Media Post-Production (Ingest Optimized), which uses 1024 -- for devices that do not report a hardware queue depth.

None of the shipped profiles set NVMe values. The mechanism supports them, but with nothing configured an NVMe pool member keeps the kernel's own settings, which on a current kernel already means no scheduler reordering and no meaningful read-ahead. Treat an all-NVMe pool as needing no profile tuning, and look at the Storage System Optimization queue depths instead if you need to change how much I/O is in flight.

Choosing a profile

The profile names describe intended workloads, so the choice usually follows from the application. What is worth understanding is why the numbers differ:

If the workload is... Use Because
General purpose, mixed, or you are not sure Default A middle read-ahead and a moderate queue depth. This is the right answer far more often than not, and we recommend leaving it alone unless you have measured a specific problem.
Virtual machines under ESXi, XenServer, Hyper-V or Virtuozzo Virtualization Many small concurrent random requests. It cuts read-ahead to 128 KB, because speculative reads on a random workload are wasted bandwidth, and holds the HDD queue at a fixed 64 to keep latency low rather than chasing throughput. It is also the only shipped profile that raises the target driver thread and tasklet floor, which is what a large number of concurrent iSCSI sessions needs.
Disk-to-disk backup, archive, large-file ingest Disk Archive Long sequential streams. Read-ahead goes to 512 KB and the queue depth stays deep, trading latency -- which nothing in a backup window cares about -- for throughput.
Media post-production playback Media Post-Production (Playback Optimized) Sustained sequential reads of very large files, where the highest read-ahead in the set (1 MB) keeps the stream ahead of the reader.
Media post-production editing and ingest Media Post-Production (Ingest Optimized) Concurrent large streams. It uses the deepest queue in the set (3 x hardware queue depth, falling back to 1024) with a 512 KB read-ahead.
The Edgeware IP TV or Web TV workflow Edgeware IP TV Optimized / Edgeware Web TV Optimized Tuning supplied for those specific applications. Do not pick them for anything else.

An all-flash pool is largely insensitive to the choice, because the SSD half of every shipped profile is identical -- the profiles differ only in their HDD values. On an all-flash pool the profile is effectively a no-op and the levers that matter are record size, compression and the system tunables.

Profiles are per pool, so a mixed appliance can run an archive profile on its backup pool and the virtualization profile on its VM pool. Read the values a profile applies with:

qs pool-profile-list
qs pool-profile-get --profile=Default

qs pool-profile-list and qs pool-profile-get report every value including the NVMe fields, and qs pool-modify --pool=<pool> --profile=<profile> assigns one.

Custom I/O profiles

Profiles are defined in /opt/osnexus/quantastor/conf/qs_io_profiles.conf, one section per profile. The section name is the profile's identifier and the name and description values are what the web interface shows, so make both unique. The simplest way to build a custom profile is to copy the section closest to your workload, rename it, and change one thing.

The keys are hdd_, ssd_ and nvme_ prefixed: read_ahead_kb, nr_requests, nr_requests_multiplier, scheduler and fifo_batch, plus the unprefixed min_target_driver_threads and min_target_driver_tasklets. A value of 0 means "leave this alone", and a profile with no name is skipped.

Restart the QuantaStor service after editing the file so the new profile is discovered and appears in the web interface:

systemctl restart quantastor

The file is replaced on upgrade. Keep a copy of any custom profile outside /opt/osnexus/quantastor/conf/ and re-apply it afterwards. A comment line must have # as its very first character; an indented # is not a comment.

If a profile does not appear to be taking effect, the service log records each application by name:

grep "Optimizing media for Storage Pool" /var/log/qs/qs_service.log

Disk optimization can also be switched off entirely by the presence of the file /etc/qs_disk_optimizations.disable, which support occasionally uses to isolate a problem. If that file exists, no profile is applied to any pool.

Note that chunk_size_kb still appears in the shipped file. It is a legacy Linux software RAID parameter, marked deprecated in the file itself, and has no effect on a scale-up pool.

Storage Volume I/O profiles

A separate profile mechanism applies to individual Storage Volumes, and it is the only place the iSCSI session parameters are exposed. Selecting a profile in Create Storage Volume sets the volume's block size and, once the volume is assigned to a host, the parameters on its iSCSI target. Three profiles ship:

Profile Block size Queued commands First burst length Optimizes for
Default 64 KB 64 128 KB General purpose backup workloads.
Desktop and Server Virtualization 64 KB 128 1 MB Virtual machine workloads. 64 KB is forward and backward compatible across VMware releases.
Databases 8 KB 128 64 KB Database workloads and small-block I/O patterns.

All three use a 1 MB maximum burst length and 1 MB maximum receive and transmit data segment lengths.

  • Queued commands is how many SCSI commands the target will accept outstanding on that session. Raising it helps a host that keeps many requests in flight, which is what a hypervisor with many virtual machines does; we do not recommend going above 256.
  • First burst length is how much write data an initiator may send unsolicited with the command itself, before the target replies asking for the rest. Setting it to the size of a typical write turns a two-round-trip write into one -- which is why the virtualization profile raises it to 1 MB, and why the database profile keeps it at 64 KB, matching a small-block write pattern instead of reserving buffers for data that never arrives.
  • Block size is the volume's on-disk record size and is fixed at creation. See Record size and volume block size.

The iSCSI parameters are applied only after the volume is assigned to a host or host group, because until then the volume has no iSCSI target to configure. Assign the volume, then confirm the profile took effect. A profile value of zero is skipped.

Read the shipped values with qs volume-profile-list and qs volume-profile-get --profile=<name>, and select one at create time with qs volume-create --volume-profile=<name>. The definitions live in /opt/osnexus/quantastor/conf/qs_volume_profiles.conf and, like the pool profiles, are replaced on upgrade.

RAM read cache (ARC)

Scale-up pools use the OpenZFS ARC in RAM as their primary read cache rather than the Linux page cache. It is the highest-leverage cache in the appliance: serving a block from RAM is orders of magnitude faster than reading it from media, and every read it absorbs is disk work that does not happen, which indirectly improves write performance too.

Sizing it

The ceiling is set by Cache Size (% of RAM) on the Cache Settings tab of Storage System Optimization, which defaults to 70% of system RAM and can be set between 30% and 90%. QuantaStor applies it by computing that percentage of total RAM and writing it to the OpenZFS zfs_arc_max parameter.

Two things follow from that:

  • The setting is re-applied when the QuantaStor service starts. A value set by hand -- with qs-util setzfsarcmax, or by writing zfs_arc_max directly -- is overwritten at the next service start. Use the dialog, or qs tunable-set --tunable=sst_cache_size:<percent>, so the change persists.
  • There is no dialog for the ARC minimum. The floor (zfs_arc_min) stops the kernel shrinking the cache below a given size under memory pressure. qs-util setzfsarcmin <percent> writes it to /etc/modprobe.d/zfs.conf, which means it only takes effect at the next boot. This is a support-level lever; raise it only if you have evidence that the ARC is being collapsed under pressure.

More useful than either number is the amount of RAM in the appliance, because 70% of too little RAM is still too little. Plan for a minimum of 32 GB to 64 GB on a small system, 96 GB to 128 GB on a medium one, and 256 GB or more on a large one; the OSNEXUS design tools will size it for a specific configuration.

What it costs

The ARC is not free memory that happens to be used for cache -- it is memory the appliance will not have for anything else:

  • Everything else on the appliance runs in the remaining 30%: the QuantaStor service, the target drivers, Samba and NFS, replication, and the kernel itself. Raising the percentage towards 90% on a system that also serves SMB, runs replication jobs or hosts containers is how a well-tuned appliance starts swapping.
  • Metadata competes with data inside the ARC. arc_meta_used in qs-iostat -a shows the split. A pool with tens of millions of small files can spend most of its cache on metadata, which is an argument for metadata offload devices rather than for more RAM.
  • An L2ARC consumes ARC. See below.
  • Deduplication consumes ARC, for the deduplication table, and it is not optional -- see Metadata, small-block and dedup offload.

Cache Compression (on by default) compresses ARC contents, which increases the effective cache size at a small CPU cost. Leave it on. Prefetch Disable (off by default) turns off the OpenZFS prefetcher, which is worth trying only on a workload that is almost purely random reads, where prefetched blocks are evicting useful ones. Both are on the Cache Settings tab of Storage System Optimization.

Cache policy

Cache Policy Primary on Modify Storage Pool -- and the equivalent setting on an individual Network Share or Storage Volume -- chooses what the ARC holds for that object: all, metadata only, or none.

This is the lever for a workload that pollutes the cache instead of benefiting from it. A large sequential backup ingest reads each block once, so caching those blocks gains nothing and evicts the working set of everything else on the appliance. Setting that share or volume to metadata keeps its directory structure cached -- which still helps, because metadata is read repeatedly -- while leaving its data blocks out of RAM. The Cache Hits chart described above is how you identify a candidate: a workload whose hits are almost all on the most-recently-used side is streaming, not reusing.

SSD read cache (L2ARC)

An L2ARC is a second-level read cache on SSD, below RAM and above the pool. Blocks evicted from the ARC land there, so a subsequent read is served from flash instead of from the data groups. Read cache devices are added from Add Log/Cache/Metadata Device(s), need no redundancy -- losing one costs you cache, not data -- and can be removed at any time. See Storage Pools for the dialog.

Whether it pays for itself depends entirely on the workload:

It helps when the working set is larger than RAM but not enormously larger, and blocks are re-read: a virtual machine estate whose common OS blocks are read by many guests, a database whose hot indexes exceed RAM, a file share with a recurring active set. Size it to roughly the application's working set.

It is wasted when:

  • The pool is already all-flash. A read cache in front of flash adds a layer without adding speed.
  • The workload streams. Data read once and never again -- backup ingest, single-pass media playback of a large library -- populates the cache with blocks nobody will ask for twice.
  • RAM was the cheaper answer. If the working set nearly fits in RAM, more RAM beats an L2ARC, because ARC hits are far faster and cost no ARC overhead.
  • It is oversized. Every block held in the L2ARC needs an index header in RAM, inside the ARC. An L2ARC sized far beyond the working set therefore takes RAM away from the cache that is faster than it -- and can leave you slower than with no read cache at all. l2_hdr_size in qs-iostat -a is that cost, in bytes.

Two measurements settle the argument. l2_size shows how much of the cache actually filled -- a cache that never fills is bigger than the workload needs. l2_hits against l2_misses shows whether it is being used at all; both near zero after several days of representative load means the cache is not earning its slot.

Judge an L2ARC after days, not hours. It has to learn which blocks are worth holding, and OpenZFS deliberately fills it slowly so that a burst of cold reads cannot flush it. A read cache benchmarked in the first hour will look useless whatever the workload.

Watch l2_cksum_bad as a health signal rather than a performance one: a rising count almost always means the SSD is failing and should be replaced.

Write log (SLOG/ZIL) and sync policy

The ZFS intent log (ZIL) is how OpenZFS keeps a promise. When an application issues a write and asks for it to be durable before the call returns, the data has to reach stable storage immediately, even though the pool would rather batch it into the next transaction group a few seconds later. The ZIL is that immediate record. By default it lives in the data groups; a write log (SLOG) device moves it onto dedicated flash so that a synchronous write costs one fast flash write rather than a scattered write into a parity stripe.

Sync Policy on Modify Storage Pool decides how much traffic goes that route:

Policy Behaviour When
standard (default) Honours the application's O_SYNC flag: writes that ask for durability are logged, everything else is batched. Almost always. It is a hybrid, and it gives hypervisors and databases the durability they ask for without penalising anything else.
always Every write goes through the log. Only when an application needs durability it does not ask for, and only with a log device fast enough to absorb the whole write stream.
disabled All writes are asynchronous. Acknowledged before they are durable. Not supported for production. It can lose acknowledged writes on an unexpected power failure, and the dialog raises a confirmation.

This is why a new write log often looks broken. Under the default standard policy, only writes that carry O_SYNC touch it -- hypervisors and databases set the flag; most other applications do not. Add a write log for a workload that does not, and it will sit idle: zil_commit_count in qs-iostat -a barely moves. That is the log correctly doing nothing, not a fault.

Setting Sync Policy to always forces every write through it, and this is where it goes wrong: if the log device cannot sustain the full write rate, forcing all writes through it makes the pool slower, not faster. Check zil_commit_stall_count and zil_commit_suspend_count -- anything other than zero under load means the log is the bottleneck.

Sizing and media choice:

  • Write log devices see constant, small, synchronous writes, which is the hardest duty cycle in the appliance. Use high-endurance enterprise SSDs; SLC-class devices are worth the money here in a way they are not for a read cache.
  • Log devices are mirrored automatically, and multiple pairs scale log throughput. See Storage Pools for the pairing rules and a worked sizing example.
  • Modern deployments need a write log less than they used to. Its value was in making an HDD pool usable for database and virtual machine workloads, and those workloads belong on flash now that NVMe and SAS SSDs are affordable. An all-flash pool rarely benefits from one.

The same sync policy choice is available per Network Share and per Storage Volume, which is the right granularity when one dataset on a pool needs always and the rest does not.

Metadata, small-block and dedup offload

On an HDD pool this is the largest single performance improvement available, and it is worth understanding why before deciding how much SSD to spend on it.

A metadata offload group (a ZFS special vdev) holds the pool's metadata on mirrored flash instead of scattered across the data groups. Metadata access is small and random by nature, which is the pattern HDDs are worst at, and it is on the critical path of far more than it looks:

  • Scrubs and resilvers walk the metadata tree. Moving it to flash shortens both substantially -- which matters most exactly when you can least afford it, during a rebuild with reduced redundancy.
  • Replication and snapshot operations are metadata-heavy, so a replication window that will not close is often a metadata problem rather than a bandwidth problem.
  • Directory traversal and file enumeration are metadata. This is what makes a large HDD file share feel slow to browse even when throughput is fine.
  • HA failover has to read pool metadata before the pool comes up.

Small Block Offload extends this to data. Set on Modify Storage Pool (and per share or volume), it routes any write below the chosen size to the metadata group instead of the data groups. The reason it helps so much on parity pools is arithmetic: writing a few kilobytes into a RAIDZ stripe means reading and rewriting parity across the whole stripe, so a small write costs far more than its size. Sending it to a mirrored SSD group instead avoids that entirely. The setting offers Disabled, 4K, 8K, 16K, 32K, 64K, 128K and 256K, and stays Disabled until a metadata group exists.

Practical guidance:

  • Always include SSDs for metadata and small-block offload on an HDD pool. The effect on scrub time, replication, failover and small-file performance is large enough that we do not design HDD pools without it. See Storage Pools for the recommended device count and starting offload size.
  • An all-flash pool does not need it and should not use it -- with one exception. QLC media handles small-block I/O poorly, so putting metadata and small blocks on TLC in front of a QLC pool is worthwhile.
  • Do not set the offload size larger than the SSD group can hold. The larger the threshold, the more of the pool's writes land on flash. Size the threshold to the capacity you gave it, not to the largest number in the list.
  • A metadata offload group cannot be removed. It is a permanent part of the pool, and losing it loses the pool, so mirror it properly the first time. There is no CLI command to remove one either.

Deduplication offload puts the deduplication table on its own mirrored vdev, and adding one turns deduplication on for the pool. Treat it as a capacity feature with a performance cost, not a performance feature. Every write has to be looked up in the deduplication table, that table has to be cached in RAM to make the lookup fast, and it grows with the amount of unique data in the pool -- so it competes with the ARC for exactly the memory your read workload wants. Like the metadata group, it cannot be removed once added. Use compression first; it gives most of the space saving on most data with none of this cost.

Record size and volume block size

Record size (for a Network Share) and Block Size (for a Storage Volume) are the same idea: the unit OpenZFS reads, writes, checksums and compresses. Matching it to the application's I/O size is one of the few settings that can change performance by a large factor, and it is one of the few that cannot be changed later for a volume.

Why the match matters:

  • Too large, for small random writes. Modifying part of a record forces a read of the whole record, the modification, recompression, and a write of the whole record. An 8 KB database write into a 1 MB record is a 1 MB read plus a 1 MB write. The same amplification applies to reads: a 4 KB read fetches the entire record.
  • Too small, for large sequential I/O. Every record carries checksum and indirect-block overhead, and each one is a separate compression unit, so small records cost more metadata, more IOPS, and less compression than large ones for the same amount of data.
Object Range offered Default Notes
Network Share -- Record Size Auto, 8K, 16K, 32K, 64K, 128K, 256K, 512K, 1M, 2M, 4M, 8M 128K (via Auto) Can be changed at any time from Modify Network Share. Existing data keeps the record size it was written with; only new writes use the new value.
Storage Volume -- Block Size 4K, 8K, 16K, 32K, 64K, 128K, 256K, 512K, 1M, 2M, 4M 64K Fixed at creation and cannot be changed afterwards. The field greys out on Modify Storage Volume.

Choosing:

  • Leave a general-purpose file share at the 128K default. It is a good compromise and shares rarely have one dominant I/O size.
  • Databases want a record or block size at or near the database page size -- 8K or 16K for most relational engines. This is what the Databases Storage Volume I/O profile sets, and it is the reason the profile exists.
  • Virtual machine datastores are well served by 64K, the volume default, which is what the Desktop and Server Virtualization profile selects and is compatible across VMware releases. A VM datastore carries a mix of guest I/O sizes, so a middle value beats optimising for either end.
  • Media, backup and archive shares want large records -- 1M and above. Fewer, larger records mean less metadata and better compression on large sequential files.
  • Match the share record size to the client's own block size where you know it, particularly for a single-purpose share.

Do not confuse either setting with Block Size Offset (ashift), which is the pool's physical sector alignment, is set at pool creation, and should be left on Auto. See Storage Pools.

Set these from the CLI with qs share-create --recordsize=<KB> and qs volume-create --blocksize-kb=<KB>.

Compression

Compression is on by default -- LZ4 -- and it is normally a throughput decision rather than a capacity one. A modern CPU compresses faster than the media can read or write, so compressing means moving fewer bytes through the disks, the HBAs and the backplane. On compressible data it makes the pool faster and larger at the same time.

That reasoning tells you when to change it:

  • Leave it on (LZ4) unless you have a reason. It is cheap enough that it costs nothing measurable on incompressible data, because OpenZFS detects a record that will not compress and stores it uncompressed.
  • Turn it off for data that is already compressed. Media and entertainment content, encrypted archives, and most image and video libraries gain nothing and only add CPU load. This is the common case in post-production.
  • Consider ZSTD when the pool is media-bound rather than CPU-bound. ZSTD compresses harder than LZ4 at more CPU cost. On an HDD pool with spare cores it can raise effective throughput; on an all-flash pool where the media is already faster than the CPU, it will reduce it. Measure both.
  • Do not reach for GZIP for performance. The gzip levels are a capacity choice for cold data, and they cost enough CPU to become the bottleneck on an active pool.
  • Changing compression affects new writes only. Existing blocks keep whatever algorithm they were written with, so a change takes effect gradually as data is rewritten.

The available algorithms and the field itself are on Modify Storage Pool -- see Storage Pools -- and the same choice is available per share and per volume, which is the right granularity when one dataset on a pool holds incompressible media and the rest does not.

Compression interacts with record size: a larger record gives the compressor more to work with and compresses better. It also interacts with partial writes, since a modified record has to be recompressed in full -- another reason not to pair a large record size with a small-write workload.

Device write caches

Worth knowing, because it explains a measurement that otherwise looks wrong. QuantaStor disables the volatile write-back cache on rotational disks as they are brought into service, putting them into write-through mode, because an HDD's on-board cache is not power-loss protected and acknowledging a write that is still only in that cache risks the pool's consistency.

Enterprise and datacenter SSDs and NVMe devices keep their write cache enabled, because they have the capacitors to flush it on power loss. Devices reached over iSCSI are put into write-through regardless.

So a raw HDD write benchmark on the appliance will fall short of the drive's data sheet, and that is deliberate. The correct way to get synchronous write performance back is a write log on flash, not a volatile cache.

Network-side factors

The network is frequently the real limit, and it is the layer where a small configuration mistake costs the most. Network Ports owns the port, VLAN and bonded-port dialogs; what follows is which choice to make and why.

MTU and jumbo frames

Navigation: Storage Management → Network Port → Modify (toolbar)

A larger MTU means more payload per frame and fewer frames, headers and interrupts for the same data -- a real gain on a 10 GbE or faster storage network carrying large block I/O. Modify Network Port has an MTU field with a Jumbo Frames button beside it that fills in 9000; once the port is on 9000 the button reads Default Frames and puts it back to 1500. The MTU field is only editable while the Static IP Configuration Settings section is active, and is disabled on VLAN and alias ports, which take the MTU of their parent.

Every device in the path must agree. The initiator, every switch port between, and the target all have to carry the same MTU. A device that receives a frame larger than its MTU drops it, and the symptom is not a clean failure -- it is a connection that works for small transfers and stalls on large ones, which is a genuinely unpleasant thing to diagnose. Set the switch first, verify end to end, and change the appliance last.

Set it from the CLI with qs network-port-modify --port=<port> --mtu=9000.

Multiple interfaces versus bonding

There are two ways to use several ports for one workload, and they are not interchangeable.

Multipath (MPIO), one subnet per port. Each port gets a static address on its own subnet, the initiator opens a session to each, and multipathing on the host spreads I/O across them and survives the loss of any one. This is the approach we recommend trying first for block storage, because it needs nothing from the switch, it scales linearly with ports, and its failure modes are visible.

Putting several ports on the same subnet is the mistake to avoid here. With multiple interfaces on one subnet, Linux by default will answer an ARP request for any local address out of any interface, so return traffic can leave a port other than the one the request arrived on. The paths you carefully separated then collapse onto whichever port answered, and throughput sits at roughly one port's worth however many you configured. QuantaStor's ARP Filtering setting on Modify Storage System addresses this -- Auto (the default) enables ARP filtering only when a bonded port is present, so a system with several unbonded ports on one subnet does not get it. If your design puts multiple ports on the same subnet, set ARP Policy explicitly to Enabled -- the CLI equivalent is qs system-modify --arp-filter-mode=enabled. Better still, give each port its own subnet and avoid the question.

For Windows initiators, one detail is worth recording because it is easy to get wrong and produces a confusing result. QuantaStor identifies itself over SCSI as vendor OSNEXUS, product QUANTASTOR, so the string to add under MPIO Devices is OSNEXUS QUANTASTOR followed by six trailing spaces -- an eight-character vendor field and a sixteen-character product field, both space padded. Without exactly the right padding the MPIO driver does not recognise the devices and Windows Disk Management shows the same disk once per path instead of one disk with several paths.

Bonding. A bonded port presents several physical ports as one logical interface with one address.

Navigation: Storage Management → Network Port → Create Bonded Port (toolbar)

Each Bond Mode in the dialog names the switch it needs, which is the part to read first:

Bond Mode Switch required
Link Aggr Ctrl Protocol (LACP layer2) Managed switch. The dialog's default.
Link Aggr Ctrl Protocol (LACP layer2+3) Managed switch. Hashes on MAC and IP, which spreads better than layer2 alone across many peers.
Link Aggr Ctrl Protocol (LACP layer3+4) Managed switch. Hashes on IP and port, so two sessions between the same pair of hosts can land on different members.
Round Robin (balance-rr) Etherchannel managed switch.
Balance XOR (balance-xor) Etherchannel managed switch.
Active-Backup (active-backup) Unmanaged switch. Failover only, with no throughput gain -- one member carries all traffic.
Adaptive Transmit Load Balancing (balance-tlb) Unmanaged switch. Balances outbound traffic only.
Adaptive Load Balancing (balance-alb) Unmanaged switch. Balances both directions without switch support.

Bonding is the right choice when you need one address -- for file protocols, or where the client cannot do multipathing -- and when the ports are spread across switches for redundancy, which requires LACP and switch infrastructure that supports it. Two cautions:

  • A bond does not make one session faster. The load-balancing modes hash each flow onto one member port, so a single TCP connection gets one port's bandwidth no matter how many are in the bond. Aggregate throughput across many clients improves; one client's single stream does not. With iSCSI over a bond there is one session, so multipathing sees one path.
  • LACP performance depends heavily on the switch's hash policy. Enabling LACP and leaving the switch on its defaults has been measured to cost the large majority of the available throughput, recovered only after matching the switch's hashing to the traffic. This varies between switch vendors, so benchmark multipath first and use that number as the target LACP has to meet -- without a target you have no way to tell a badly hashed bond from a fast one.

A Network Bonding Policy setting on Modify Storage System carries a system-level bonding mode drawn from a subset of the same list; see Storage System.

Where you use both approaches, bond groups of ports and run multipath across the bonds.

NIC ring buffers

Optimize hardware RX/TX buffer settings for throughput on Modify Network Port raises the network card's receive and transmit ring buffers. Larger rings give the driver more room to absorb a burst before dropping packets, which matters on a fast link carrying large block I/O or replication traffic. QuantaStor picks the largest power-of-two value that stays safely below the card's hardware maximum, records the original values so that clearing the checkbox restores them, and re-checks the setting periodically.

The checkbox is unavailable on a bonded port and on virtual ports -- tune the member ports instead -- and on a port whose configuration type is disabled. The CLI equivalent is qs network-port-modify --port=<port> --auto-tune=true.

Two related system tunables sit on the Network Settings tab of Storage System Optimization: Network TX/RX Queue Length (the software queue, default 5000, and worth raising on 10 GbE and faster) and Network Device Max Backlog (how many packets may queue on the receive side when the interface delivers faster than the kernel can process).

What does not apply

Three things that look like tuning levers and are not:

  • Hardware RAID card settings. Scale-up pools are deployed on HBAs, not on hardware RAID controllers. OpenZFS needs direct access to the drives to checksum, self-heal and manage its own redundancy, and a RAID card's cache, stripe size and read-ahead settings sit between it and the media doing none of those things. There is no RAID-card tuning to do on a scale-up pool, because there should be no RAID card. Pick the layout in Storage Pools instead.
  • Linux software RAID (mdadm) parameters. The [mdadm] section of /etc/quantastor.conf and the chunk_size_kb key in the I/O profiles file are legacy settings from pool types the product no longer builds. They have no effect on a scale-up pool.
  • The [device] section of /etc/quantastor.conf. Read-ahead, queue depth and scheduler were once configured there. They are not any more -- those settings come from the pool's I/O profile, and the keys in that section are inert. Edit the profile, not the configuration file.

Two further notes on scope. Most of the system tunables in Storage System Optimization are OpenZFS module parameters, so they affect scale-up pools only and do nothing for scale-out (Ceph) pools. And tunables are stored per Storage System, so in a grid each member is tuned independently -- save a profile and apply it to the others rather than editing each by hand.

Related pages


Verified against QuantaStor 6.9.0.