Performance Tuning: Difference between revisions

From OSNEXUS Online Documentation Site
Jump to navigation Jump to search
mNo edit summary
m Deduplicate against the rewritten Performance Monitoring page (QSTOR-12352, guideline 14): replace the Storage System dashboard, Cache Stats and per-disk/grid subsections with a pointer that keeps only the three reading rules that matter when tuning - check the time range because windows over six hours are averaged, Cache Efficiency is a since-boot average not a current hit rate, and network throughput sits beside CPU. qs-iostat and the direct disk/pool testing sections are unchanged. Drops t...
 
(3 intermediate revisions by the same user not shown)
Line 1: Line 1:
[[Category:admin_guide]]
[[Category:admin_guide]]
= Storage Pool IO Tuning Overview =


Storage pools in QuantaStor have associated IO profiles which control parameters like read-ahead, request queue depth, and the IO scheduler. QuantaStor comes with a number of IO profiles to match common workloads but it is possible to create your own IO profiles to tune for a specific application.  In general, we recommend going with one of the defaults but for performance tuning purposes you can add additional profiles to the /etc/qs_io_profiles.conf file.  Here are the IO profiles for the default and default SSD configurations.
This page covers how to tune a scale-up (OpenZFS) QuantaStor appliance for a workload: what to measure, what each tuning control actually changes underneath, and how to choose between them. It is a decision guide rather than a field reference -- where a control lives on another page's dialog, the reasoning is here and the field-by-field detail is linked.


<pre>
'''Find out what is slow before you change anything.''' [[Performance Testing]] covers the measurement tools the appliance ships -- {{Code|1=qs-perftest}}, {{Code|1=qs-ramdisk}} and the Disk Performance Test dialog -- and a method for narrowing a performance complaint down to the media, the network, or the pool. Tuning a layer that is not the problem costs time and can make things worse.
[default]
name=Default
description=Optimizes for general purpose server system workloads
nr_requests=2048
read_ahead_kb=256
fifo_batch=16
chunk_size_kb=128
scheduler=deadline


[default-ssd]
Two pages own the controls this one reasons about. [[Storage System Optimization]] owns the appliance-wide OpenZFS and kernel tunables dialog. [[Storage Pools]] owns pool creation and modification, including the I/O profile, compression, sync and cache policy fields.
name=Default SSD
 
description=Optimizes for pure SSD based storage pools for all workloads
{| class="wikitable"
nr_requests=2048
! Section !! Purpose
read_ahead_kb=4
|-
fifo_batch=16
| [[#How to approach a tuning problem|How to approach a tuning problem]] || The order of operations, and what not to do.
chunk_size_kb=128
|-
scheduler=noop
| [[#What to measure|What to measure]] || The dashboards, {{Code|1=qs-iostat}}, and where the measurement tools live.
|-
| [[#Where the knobs are|Where the knobs are]] || Five separate tuning layers and which one a given problem belongs to.
|-
| [[#Storage Pool I/O profiles|Storage Pool I/O profiles]] || Read-ahead, request queue depth, I/O scheduler, and the seven shipped profiles.
|-
| [[#Storage Volume I/O profiles|Storage Volume I/O profiles]] || Volume block size and the iSCSI session parameters.
|-
| [[#RAM read cache (ARC)|RAM read cache (ARC)]] || Sizing the ARC and what it costs in RAM.
|-
| [[#SSD read cache (L2ARC)|SSD read cache (L2ARC)]] || When a read cache pays for itself and when it is wasted.
|-
| [[#Write log (SLOG/ZIL) and sync policy|Write log (SLOG/ZIL) and sync policy]] || Why a write log often sits idle, and what changes that.
|-
| [[#Metadata, small-block and dedup offload|Metadata, small-block and dedup offload]] || The single largest win available on an HDD pool.
|-
| [[#Record size and volume block size|Record size and volume block size]] || Matching the on-disk unit to the application's I/O size.
|-
| [[#Compression|Compression]] || Why compression is usually a throughput decision, not a capacity one.
|-
| [[#Device write caches|Device write caches]] || Why an HDD in a pool is slower than its data sheet.
|-
| [[#Network-side factors|Network-side factors]] || MTU, multiple interfaces versus bonding, and NIC ring buffers.
|-
| [[#What does not apply|What does not apply]] || Controls that look relevant and are not.
|}
 
== How to approach a tuning problem ==
 
The shipped defaults are right for the large majority of deployments, and most disappointing performance turns out to be a design problem rather than a tuning problem: too few device groups, a parity stripe that is too wide for a random workload, HDDs with no metadata offload, or a single network path. Tuning cannot fix any of those. Work through the design first -- see [[#Where the knobs are|Where the knobs are]] and [[Storage Pools]] -- and only then reach for a knob.
 
When you do tune:
 
# '''Measure first, with a workload that resembles production.''' A number you cannot reproduce is not a baseline, and without a baseline you cannot tell an improvement from noise.
# '''Confirm the hardware is healthy.''' Every device in '''Physical Disks''' should be '''Normal'''. One disk with a predictive-failure warning will make a whole pool look badly tuned. See [[Physical Disks/Devices]].
# '''Change one group of related settings, then measure again.''' Several of the OpenZFS tunables interact, and a batch of simultaneous changes leaves you with no way to attribute the result.
# '''Save a named profile before experimenting''' so you can get back to a known state. See [[Storage System Optimization]] for tunable profiles and [[#Custom I/O profiles|Custom I/O profiles]] below for pool profiles.
# '''Write down what you changed.''' Tuning applied per system in a grid drifts silently between members otherwise.
 
We recommend involving OSNEXUS support before changing the system tunables. They are the settings that matter for specific, identifiable problems -- write latency spikes under load, a resilver that will not finish inside a maintenance window, a replication job overrunning its window -- and changing them speculatively is more likely to cost performance than gain it.
 
== What to measure ==
 
=== Reading the dashboards ===
 
The appliance's own dashboards are the first place to look, and [[Performance Monitoring]] documents them in full -- the views, what each series means, and the sampling and retention behaviour that decides whether a change is even visible in a given chart. Three points from there matter specifically when tuning:
 
* '''Check the time range before drawing a conclusion.''' Windows longer than six hours are served from averaged data, so a short latency spike caused by a tuning change can be averaged away entirely. Use a range of six hours or less when assessing a change you just made.
* '''Cache Efficiency is a since-boot average, not a current hit rate.''' It moves very slowly on an appliance with a long uptime, so it is close to useless for judging an ARC change made minutes ago. Watch the hit counts instead, or read the counters directly with <code>qs-iostat -a</code>.
* '''Per-port network throughput is on the same Performance view as CPU''', which makes it the quickest way to tell a network-bound workload from a storage-bound one before touching any pool setting.
 
Selecting an individual disk replaces the dashboard with a per-device view of IOPS, throughput and latency -- the fastest way to confirm that load is landing where you expect during a test.
 
=== qs-iostat ===
 
{{Code|1=qs-iostat}} is the console-level counterpart, and it reports the ARC and ZIL kernel counters that the charts summarise. It wraps {{Code|1=iostat}}:
 
<pre style="font-size: smaller">
qs-iostat -c            # globally averaged CPU stats
qs-iostat -d            # I/O stats for all block devices, in MB/s
qs-iostat -a            # ZFS ARC, L2ARC and ZIL counters
qs-iostat -f            # repeat every 2 seconds
qs-iostat --extra "-x"  # pass extra arguments through to iostat
</pre>
 
{{Code|1=qs-iostat -a}} is the useful one for cache work. The counters are cumulative since boot, so take two samples and compare, rather than reading absolute numbers:
 
{| class="wikitable"
! Counter !! Meaning
|-
| {{Code|1=hits}} / {{Code|1=misses}} || Read requests satisfied from, and missed in, the RAM cache. The ratio between two samples is the live hit rate.
|-
| {{Code|1=size}} || Current ARC size.
|-
| {{Code|1=c_min}} / {{Code|1=c_max}} || The floor and ceiling the ARC may grow between. {{Code|1=c_max}} is what '''Cache Size (% of RAM)''' sets.
|-
| {{Code|1=arc_meta_used}} || How much of the ARC is metadata rather than data. On a pool with many small files this can be most of it.
|-
| {{Code|1=l2_hits}} / {{Code|1=l2_misses}} || Reads satisfied from, and missed in, the SSD read cache. Both zero means the L2ARC is doing nothing.
|-
| {{Code|1=l2_size}} || How much of the read cache is actually populated. A cache that never fills is larger than the workload needs.
|-
| {{Code|1=l2_hdr_size}} || RAM consumed by the L2ARC's index. This is the hidden cost of a read cache -- see [[#SSD read cache (L2ARC)|SSD read cache (L2ARC)]].
|-
| {{Code|1=l2_read_bytes}} / {{Code|1=l2_write_bytes}} || Bytes read from and written to the read cache devices.
|-
| {{Code|1=l2_cksum_bad}} || Checksum failures on a read cache device. A rising count usually means the SSD is failing and should be replaced.
|-
| {{Code|1=zil_commit_count}} || Synchronous write commits since boot. Zero, or near zero, on a pool with a write log means the log is idle -- see [[#Write log (SLOG/ZIL) and sync policy|Write log (SLOG/ZIL) and sync policy]].
|-
| {{Code|1=zil_commit_writer_count}} || Commits that had to write, as opposed to joining a commit already in flight.
|-
| {{Code|1=zil_commit_stall_count}} / {{Code|1=zil_commit_suspend_count}} || Commits that stalled or were suspended. Anything other than zero under load points at the log device not keeping up.
|-
| {{Code|1=zil_itx_count}} || Intent log transactions recorded.
|}
 
''(Note: these counters come straight from the OpenZFS kernel modules and the set changes between OpenZFS releases; {{Code|1=/proc/spl/kstat/zfs/arcstats}} and {{Code|1=/proc/spl/kstat/zfs/zil}} on the appliance are authoritative.)''
 
=== Testing the disks and the pool directly ===
 
[[File:perftune_disk_perftest.png|thumb|right|800px|Physical Disk Performance Test. The Read Seq and Last Performance Test columns keep the result on each disk, so a later run can be compared against it.]]
 
The dashboards and {{Code|1=qs-iostat}} tell you what the appliance is doing under its current load. To make it do something measurable on purpose, the appliance ships two utilities and a dialog:
 
* '''Disk Performance Test''', on a physical disk's right-click menu, reads sequentially from the disks you select and stores the result on each one. It is the fastest way to answer "is one device in this pool slower than its peers?" -- see [[Physical Disks/Devices]] for its fields and modes.
* {{Code|1=qs-perftest}} read-tests every disk or every disk in one pool from the command line, and benchmarks a pool by creating temporary shares and driving fio, elbencho, dd and iozone against them.
* {{Code|1=qs-ramdisk}} creates transient RAM-backed disks, so a benchmark can be run with the media factored out entirely -- which is how you establish that a bottleneck is above the media rather than in it.
 
'''[[Performance Testing]] covers all three, and the order to use them in.''' Work through it before changing anything on this page: a degraded device or a misconfigured bond produces exactly the symptoms that tempt an administrator into the tunables, and no setting here will fix either.
 
 
== Where the knobs are ==
 
Five separate mechanisms are easy to confuse. Deciding which layer a problem belongs to saves most of the work:
 
{| class="wikitable"
! Layer !! Scope !! What it controls !! Where
|-
| '''Storage Pool I/O profile''' || one pool, re-applied at every pool start || Block device settings on the member disks: read-ahead, request queue depth, I/O scheduler. Also raises the SCST target driver thread floor for the whole appliance. || [[#Storage Pool I/O profiles|below]]; the field is on Create and Modify Storage Pool ([[Storage Pools]])
|-
| '''Pool, share and volume properties''' || one pool, share or volume || Record size and volume block size, compression, sync policy, primary and secondary cache policy, small-block offload. || [[#Record size and volume block size|below]]; the fields are on [[Storage Pools]], [[Network Shares]] and [[Storage Volumes]]
|-
| '''Storage Volume I/O profile''' || one volume || Volume block size and the iSCSI session parameters on that volume's target. || [[#Storage Volume I/O profiles|below]]
|-
| '''System tunables''' || one Storage System || OpenZFS and SPL module parameters plus two network queue settings: ARC size, write throttle, per-VDEV queue depths, resilver and scrub priority, prefetch distance. || [[Storage System Optimization]]
|-
| '''Pool topology and offload tiers''' || one pool, fixed at build time (mostly) || Number and width of device groups, mirroring versus parity, and the write log, read cache, metadata and dedup vdevs. || [[Storage Pools]]
|}
 
The last row is not a tuning layer, but it is where the large factors live. A pool with one wide parity group and no flash offload cannot be tuned into a random-I/O pool, and the first four layers together will not move it as far as adding device groups will.
 
== Storage Pool I/O profiles ==
 
A Storage Pool I/O profile is a bundle of Linux block device settings applied to the pool's member disks. It is not a filesystem or OpenZFS setting. QuantaStor re-applies it whenever the pool is created or started, on HA failover, when the profile is changed on '''Modify Storage Pool''', and whenever cache, log, metadata or hot-spare devices are added or removed -- so the tuning survives the events that would otherwise leave new or reconnected devices on kernel defaults. Each pool has one profile; the field is on both '''Create Storage Pool''' (Advanced Settings) and '''Modify Storage Pool''' -- see [[Storage Pools]] for the dialogs.
 
Each profile carries a separate set of values for HDD, SSD and NVMe media, and QuantaStor picks the set per device from what the device reports, so one profile tunes a mixed pool correctly. On a multipath pool the settings are applied to each underlying SCSI path rather than to the {{Code|1=dm}} device, which is where they actually take effect.
 
=== What each setting changes ===
 
'''Read-ahead''' ({{Code|1=/sys/block/&lt;dev&gt;/queue/read_ahead_kb}}) is how much extra data the kernel pulls in after satisfying a read. If a read asks for 4 KB and read-ahead is 256 KB, another 256 KB is read into the page cache behind it.
 
Read-ahead exists to hide rotational latency. A 7200 RPM drive turns 120 times a second, roughly once every 8 ms. Reading 4 KB records one per rotation yields under 500 KB/s -- so on any workload with sequential locality, reading ahead and having the next records already in memory is the difference between a usable HDD pool and an unusable one. That is why read-ahead and queue depth matter so much on spinning media and why the archive and media profiles push read-ahead up to 512 KB and 1 MB.
 
Flash has no rotational latency, so read-ahead on flash is pure overhead: every speculative block occupies bandwidth and cache that a real request wanted. '''Every shipped profile uses a 4 KB SSD read-ahead''', and that is a deliberate result rather than a rounding to zero -- the original SSD tuning work found that both 0 KB and larger values produced pronounced dips at 4 KB and 8 KB request sizes that a small 4 KB read-ahead removed. It remains the single most important SSD-side setting in these profiles.
 
'''Request queue depth''' ({{Code|1=/sys/block/&lt;dev&gt;/queue/nr_requests}}) is how many I/O requests the kernel will hold queued for a device. A deeper queue gives the scheduler and the device more requests to sort and coalesce, which raises throughput on rotational media; too deep a queue raises latency, because a small urgent read now waits behind a long queue of other work.
 
Profiles express this two ways. A fixed count sets the queue depth outright. A '''multiplier''' instead reads the device's own hardware queue depth from {{Code|1=/sys/block/&lt;dev&gt;/device/queue_depth}} and multiplies it, so the same profile scales sensibly across drives with very different capabilities. Two rules matter:
 
* '''The multiplier takes precedence''' over the fixed count whenever the device reports a non-zero hardware queue depth. The fixed count is the fallback for devices that report none.
* '''The computed value is capped at 1024.''' A multiplier against a deep hardware queue is clamped there, with a warning in {{Code|1=/var/log/qs/qs_service.log}}.
 
'''I/O scheduler''' ({{Code|1=/sys/block/&lt;dev&gt;/queue/scheduler}}) is the elevator algorithm that decides the order requests reach the device. Two are relevant. {{Code|1=deadline}} sorts requests to reduce head movement while enforcing a latency deadline so nothing starves -- the right choice for HDDs. {{Code|1=noop}} does essentially no reordering and hands requests straight down, which is right for flash, where there is no seek cost to optimise away and reordering only adds CPU work and latency. Current kernels name these {{Code|1=mq-deadline}} and {{Code|1=none}}; QuantaStor accepts either spelling in a profile and writes whichever the device offers. If a device offers neither, the setting is skipped and logged rather than failing the pool start.
 
Two further settings appear in profiles:
 
* '''FIFO batch''' ({{Code|1=/sys/block/&lt;dev&gt;/queue/iosched/fifo_batch}}) tunes how many requests the deadline scheduler dispatches in one batch. It only exists while a deadline-style scheduler is selected, so it is skipped on any device running {{Code|1=noop}}/{{Code|1=none}}. None of the shipped profiles set it.
* '''Target driver threads and tasklets''' set a floor on the SCST target driver's worker thread and tasklet counts, which is what serves iSCSI, FC and NVMe-oF. These are '''appliance-wide, not per pool''': QuantaStor takes the highest value across every profile in use on the system, and can only raise the count above SCST's own start-up default (derived from the core count), never lower it. Both are capped at 256. Assigning the '''Virtualization''' profile to one pool therefore raises the thread floor for every target on that appliance.
 
QuantaStor also raises {{Code|1=max_sectors_kb}} to 4096 on each member device where the hardware allows it, independently of the profile, so that large sequential I/O is not split into smaller commands.
 
=== The shipped profiles ===
 
Seven profiles ship with the appliance. Values below are read from {{Code|1=/opt/osnexus/quantastor/conf/qs_io_profiles.conf}} on a 6.9.0 appliance.
 
{| class="wikitable"
! Profile !! HDD read-ahead !! HDD queue depth !! HDD scheduler !! SSD read-ahead !! SSD queue depth !! SSD scheduler !! Target driver threads / tasklets
|-
| Default || 256 KB || 2 x hw queue depth || deadline || 4 KB || 1 x hw queue depth || noop || SCST default
|-
| Disk Archive || 512 KB || 2 x hw queue depth || deadline || 4 KB || 1 x hw queue depth || noop || SCST default
|-
| Edgeware IP TV Optimized || 256 KB || 64 || deadline || 4 KB || 1 x hw queue depth || noop || SCST default
|-
| Edgeware Web TV Optimized || 256 KB || 64 || deadline || 4 KB || 1 x hw queue depth || noop || SCST default
|-
| Media Post-Production (Ingest Optimized) || 512 KB || 3 x hw queue depth || deadline || 4 KB || 1 x hw queue depth || noop || SCST default
|-
| Media Post-Production (Playback Optimized) || 1024 KB || 2 x hw queue depth || deadline || 4 KB || 1 x hw queue depth || noop || SCST default
|-
| Virtualization || 128 KB || 64 || deadline || 4 KB || 1 x hw queue depth || noop || 64 / 32
|}
 
Where a multiplier is shown, the profile also carries a fixed fallback count -- 256 for every profile except Media Post-Production (Ingest Optimized), which uses 1024 -- for devices that do not report a hardware queue depth.
 
'''None of the shipped profiles set NVMe values.''' The mechanism supports them, but with nothing configured an NVMe pool member keeps the kernel's own settings, which on a current kernel already means no scheduler reordering and no meaningful read-ahead. Treat an all-NVMe pool as needing no profile tuning, and look at the [[Storage System Optimization]] queue depths instead if you need to change how much I/O is in flight.
 
=== Choosing a profile ===
 
The profile names describe intended workloads, so the choice usually follows from the application. What is worth understanding is why the numbers differ:
 
{| class="wikitable"
! If the workload is... !! Use !! Because
|-
| General purpose, mixed, or you are not sure || '''Default''' || A middle read-ahead and a moderate queue depth. This is the right answer far more often than not, and we recommend leaving it alone unless you have measured a specific problem.
|-
| Virtual machines under ESXi, XenServer, Hyper-V or Virtuozzo || '''Virtualization''' || Many small concurrent random requests. It cuts read-ahead to 128 KB, because speculative reads on a random workload are wasted bandwidth, and holds the HDD queue at a fixed 64 to keep latency low rather than chasing throughput. It is also the only shipped profile that raises the target driver thread and tasklet floor, which is what a large number of concurrent iSCSI sessions needs.
|-
| Disk-to-disk backup, archive, large-file ingest || '''Disk Archive''' || Long sequential streams. Read-ahead goes to 512 KB and the queue depth stays deep, trading latency -- which nothing in a backup window cares about -- for throughput.
|-
| Media post-production playback || '''Media Post-Production (Playback Optimized)''' || Sustained sequential reads of very large files, where the highest read-ahead in the set (1 MB) keeps the stream ahead of the reader.
|-
| Media post-production editing and ingest || '''Media Post-Production (Ingest Optimized)''' || Concurrent large streams. It uses the deepest queue in the set (3 x hardware queue depth, falling back to 1024) with a 512 KB read-ahead.
|-
| The Edgeware IP TV or Web TV workflow || '''Edgeware IP TV Optimized''' / '''Edgeware Web TV Optimized''' || Tuning supplied for those specific applications. Do not pick them for anything else.
|}
 
An all-flash pool is largely insensitive to the choice, because the SSD half of every shipped profile is identical -- the profiles differ only in their HDD values. On an all-flash pool the profile is effectively a no-op and the levers that matter are record size, compression and the system tunables.
 
Profiles are per pool, so a mixed appliance can run an archive profile on its backup pool and the virtualization profile on its VM pool. Read the values a profile applies with:
 
<pre style="font-size: smaller">
qs pool-profile-list
qs pool-profile-get --profile=Default
</pre>
</pre>


If you create a new profile, make sure that you put a unique name/ID for your profile in the square brackets, and that you set a friendly name and description for your profile. For example, your new profile might look like this:
<code>[[QuantaStor CLI Command Reference#pool-profile-list|qs pool-profile-list]]</code> and <code>[[QuantaStor CLI Command Reference#pool-profile-get|qs pool-profile-get]]</code> report every value including the NVMe fields, and <code>[[QuantaStor CLI Command Reference#pool-modify|qs pool-modify]] --pool=&lt;pool&gt; --profile=&lt;profile&gt;</code> assigns one.


<pre>
=== Custom I/O profiles ===
[acme-db-profile]
 
name=SQL DB Performance
Profiles are defined in {{Code|1=/opt/osnexus/quantastor/conf/qs_io_profiles.conf}}, one section per profile. The section name is the profile's identifier and the {{Code|1=name}} and {{Code|1=description}} values are what the web interface shows, so make both unique. The simplest way to build a custom profile is to copy the section closest to your workload, rename it, and change one thing.
description=Optimizes for Acme Corp relational databases.
 
nr_requests=2048
The keys are {{Code|1=hdd_}}, {{Code|1=ssd_}} and {{Code|1=nvme_}} prefixed: {{Code|1=read_ahead_kb}}, {{Code|1=nr_requests}}, {{Code|1=nr_requests_multiplier}}, {{Code|1=scheduler}} and {{Code|1=fifo_batch}}, plus the unprefixed {{Code|1=min_target_driver_threads}} and {{Code|1=min_target_driver_tasklets}}. A value of {{Code|1=0}} means "leave this alone", and a profile with no {{Code|1=name}} is skipped.
read_ahead_kb=64
 
fifo_batch=16
'''Restart the QuantaStor service after editing the file''' so the new profile is discovered and appears in the web interface:
chunk_size_kb=128
 
scheduler=deadline
<pre style="font-size: smaller">
systemctl restart quantastor
</pre>
</pre>


=== Profile Name & Description ===
'''The file is replaced on upgrade.''' Keep a copy of any custom profile outside {{Code|1=/opt/osnexus/quantastor/conf/}} and re-apply it afterwards. A comment line must have {{Code|1=#}} as its very first character; an indented {{Code|1=#}} is not a comment.
The ''name'' and ''description'' fields will show up in the QuantaStor web interface so be sure to make these unique.


=== Profile Disk I/O Queue Depth (nr_requests) ===
If a profile does not appear to be taking effect, the service log records each application by name:
The nr_requests represents the IO queue size.  This is an important variable to change on the host/initiator side as well if you're connecting to your QuantaStor system via [http://www.monperrus.net/martin/scheduler+queue+size+and+resilience+to+heavy+IO iSCSI from a Linux based server].  In QuantaStor when you set the nr_requests the core service applies this setting to all disks in the pool by adjusting parameters in /sys/block/ part of the sys filesystem.


=== Profile Read-Ahead ===
<pre style="font-size: smaller">
The read_ahead_kb represents the amount of additional data that should be read after fulfilling a given read request.  For example, if there's a read-request for 4KB and the read_ahead_kb is set to 64, then an additional 64KB will be read into the cache after the base 4KB request has been met.  Why read this additional data?  It counteracts the rotational latency problems inherent in spinning disk / hard disk drives. A 7200 RPM hard drive rotates 120 times per second, or roughly once every 8ms.  That may sound fast, but take an example where you're reading records from a database and only gathering 4KB with each IO read request (one read per rotation).  Done serially that would produce a throughput of a mere 480K/sec.  This is why read-ahead and request queue depth are so important to getting good performance with spinning disk.  With SSD, there are no mechanical rotational latency issues so the SSD profile uses a small 4k read-ahead.
grep "Optimizing media for Storage Pool" /var/log/qs/qs_service.log
</pre>


=== Profile IO Scheduler ===
Disk optimization can also be switched off entirely by the presence of the file {{Code|1=/etc/qs_disk_optimizations.disable}}, which support occasionally uses to isolate a problem. If that file exists, no profile is applied to any pool.


The IO scheduler represents the elevator algorithm used for [http://en.wikipedia.org/wiki/Deadline_scheduler scheduling] I/O operations. For storage systems the two best schedulers are the 'deadline' and the 'noop' scheduler.  We find that deadline is best for HDDs and that the noop scheduler is best for SSDs.
Note that {{Code|1=chunk_size_kb}} still appears in the shipped file. It is a legacy Linux software RAID parameter, marked deprecated in the file itself, and has no effect on a scale-up pool.
The fifo_batch setting is related to the use of the 'deadline' scheduler.  More information can be found [http://en.wikipedia.org/wiki/Deadline_scheduler here].


== SSD Tuning ==
== Storage Volume I/O profiles ==


The Default SSD tuning profile was produced after extensive testing with enterprise STEC 2TB SAS SSD drives connected to an LSI HBA.  Here is some of the detail from the performance testing using iozone.
A separate profile mechanism applies to individual Storage Volumes, and it is the only place the iSCSI session parameters are exposed. Selecting a profile in '''Create Storage Volume''' sets the volume's block size and, once the volume is assigned to a host, the parameters on its iSCSI target. Three profiles ship:
Note that the Default SSD storage pool IO profile uses a small read-ahead of just 4k.  This was the most important factor in the tuning for SSD drives.  A 0KB read-ahead and larger read-ahead values seemed to create performance holes for small 4k and 8k block sizes which were eliminated with the 4KB read-ahead.  Note that the flat area in the graphs on the front-right side are set to zero because that's a section where tests were not run either due to the transfer size being larger than the file size or the transfer size being so small vs the file size that it would not produce additional valuable numbers.  Testing was done using standard iozone parameters with an increased 2G file size (iozone -a -g 2g -b output.xls), numbers with a 8GB file size were very similar.


[[File:qs_zfs_ssd_read.png|800px]]
{| class="wikitable"
! Profile !! Block size !! Queued commands !! First burst length !! Optimizes for
|-
| Default || 64 KB || 64 || 128 KB || General purpose backup workloads.
|-
| Desktop and Server Virtualization || 64 KB || 128 || 1 MB || Virtual machine workloads. 64 KB is forward and backward compatible across VMware releases.
|-
| Databases || 8 KB || 128 || 64 KB || Database workloads and small-block I/O patterns.
|}


[[File:qs_zfs_ssd_write.png|800px]]
All three use a 1 MB maximum burst length and 1 MB maximum receive and transmit data segment lengths.


= iSCSI Performance Tuning Overview =
* '''Queued commands''' is how many SCSI commands the target will accept outstanding on that session. Raising it helps a host that keeps many requests in flight, which is what a hypervisor with many virtual machines does; we do not recommend going above 256.
Getting the network configuration setup right is an important factor in getting good performance from your system.  Improper LACP configuration, multi-pathing and many other components must be properly configured for optimal performance. The following guide goes over a simple setup with Windows Server 2008 R2 but is great at illustrating the large differences small configuration changes can make in boosting IO performance. In the following sections I'll be going over three major topics:
* '''First burst length''' is how much write data an initiator may send unsolicited with the command itself, before the target replies asking for the rest. Setting it to the size of a typical write turns a two-round-trip write into one -- which is why the virtualization profile raises it to 1 MB, and why the database profile keeps it at 64 KB, matching a small-block write pattern instead of reserving buffers for data that never arrives.
* Multi-path IO configuration and performance results
* '''Block size''' is the volume's on-disk record size and is fixed at creation. See [[#Record size and volume block size|Record size and volume block size]].
* LACP configuration and performance results
* Jumbo frames configuration and performance results


== Test Setup ==
'''The iSCSI parameters are applied only after the volume is assigned to a host or host group''', because until then the volume has no iSCSI target to configure. Assign the volume, then confirm the profile took effect. A profile value of zero is skipped.


This testing was done with the following setup:
Read the shipped values with <code>[[QuantaStor CLI Command Reference#volume-profile-list|qs volume-profile-list]]</code> and <code>[[QuantaStor CLI Command Reference#volume-profile-get|qs volume-profile-get]] --profile=&lt;name&gt;</code>, and select one at create time with <code>[[QuantaStor CLI Command Reference#volume-create|qs volume-create]] --volume-profile=&lt;name&gt;</code>. The definitions live in {{Code|1=/opt/osnexus/quantastor/conf/qs_volume_profiles.conf}} and, like the pool profiles, are replaced on upgrade.


====Windows 2008 R2 host====
== RAM read cache (ARC) ==
* four 10GbE ports
====QuantaStor Storage System====
* four 10GbE ports
* LSI 9750-8i Controller
* 24x 360GB 2.5" drives
* RAID0


Within QuantaStor I also created a storage pool and a volume (thick provisioned). I then assigned it to the Windows host.
Scale-up pools use the OpenZFS ARC in RAM as their primary read cache rather than the Linux page cache. It is the highest-leverage cache in the appliance: serving a block from RAM is orders of magnitude faster than reading it from media, and every read it absorbs is disk work that does not happen, which indirectly improves write performance too.


== Microsoft MPIO ==
=== Sizing it ===


MPIO (multi-path IO) is when there are multiple paths (iSCSI sessions) between the devices which are combined into one. MPIO increases performance and fault-tolerance as paths can be lost without interrupting access to the storage. Performance is increased by using a the basic round-robin technique so that the sessions evenly share the IO load.  
The ceiling is set by '''Cache Size (% of RAM)''' on the '''Cache Settings''' tab of [[Storage System Optimization]], which defaults to '''70%''' of system RAM and can be set between 30% and 90%. QuantaStor applies it by computing that percentage of total RAM and writing it to the OpenZFS {{Code|1=zfs_arc_max}} parameter.


Our setup was tested first with a direct connection from the Windows host to the QuantaStor system, and then was tested again with a 10GbE switch in between (Interface Masters Niagara).
Two things follow from that:


=== QuantaStor System Network Setup ===
* '''The setting is re-applied when the QuantaStor service starts.''' A value set by hand -- with {{Code|1=qs-util setzfsarcmax}}, or by writing {{Code|1=zfs_arc_max}} directly -- is overwritten at the next service start. Use the dialog, or <code>[[QuantaStor CLI Command Reference#tunable-set|qs tunable-set]] --tunable=sst_cache_size:&lt;percent&gt;</code>, so the change persists.
* '''There is no dialog for the ARC minimum.''' The floor ({{Code|1=zfs_arc_min}}) stops the kernel shrinking the cache below a given size under memory pressure. {{Code|1=qs-util setzfsarcmin &lt;percent&gt;}} writes it to {{Code|1=/etc/modprobe.d/zfs.conf}}, which means it only takes effect at the next boot. This is a support-level lever; raise it only if you have evidence that the ARC is being collapsed under pressure.


To setup QuantaStor for use with MPIO we first need to assign each network port to its own subnet.  To do this navigate to the Storage System section within QuantaStor Manager. Next, select the "Network Ports" tab, then right-click on one of the ports you will be using for iSCSI connectivity to the Windows host. Select "Modify Network Port", and configure the port to use a static IP address on its own network.  Make sure to do this for every network port, putting them all on the different networks.  (Ex: 192.168.10.10/255.255.255.0, 192.168.11.10/255.255.255.0, etc).  Having the ports on separate networks is important.  When ports are on the same network any port can respond to a request on any other port and this will cause performance issues.
More useful than either number is the amount of RAM in the appliance, because 70% of too little RAM is still too little. Plan for a minimum of 32&nbsp;GB to 64&nbsp;GB on a small system, 96&nbsp;GB to 128&nbsp;GB on a medium one, and 256&nbsp;GB or more on a large one; the [https://www.osnexus.com/zfs-designer OSNEXUS design tools] will size it for a specific configuration.


=== Windows MPIO Setup ===
=== What it costs ===


Now we can begin to setup Windows to use MPIO. The first step is to navigate to MPIO properties. The easiest way to do this is to type MPIO in the start menu search bar. It can also be found in the control panel under "MPIO". In the "MPIO Devices" tab click the "Add" button. The string you will want to add for QuantaStor is "OSNEXUS QUANTASTOR&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;"
The ARC is not free memory that happens to be used for cache -- it is memory the appliance will not have for anything else:
[[File:mpio.png|frame|The string you will want to add for QuantaStor is "OSNEXUS QUANTASTOR&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;". ]]
This is all capital letters, with one space between the words, and six trailing spaces after QuantaStor. If you don't have exactly (6) trailing spaces then the MPIO driver won't recognize the QuantaStor devices and you'll end up with multiple instances of the same disk in Windows Disk Manager.  Now that MPIO is enabled for Quantastor we have to configure the network connections. Navigate to Control Panel --> Network and Internet --> Network Connections. For each connection that will connect to Quantastor assign it a static IP address, on a matching network as one from the Quantastor box. For our setup we did as follows:


==== Windows Network Port Configuration ====
* Everything else on the appliance runs in the remaining 30%: the QuantaStor service, the target drivers, Samba and NFS, replication, and the kernel itself. Raising the percentage towards 90% on a system that also serves SMB, runs replication jobs or hosts containers is how a well-tuned appliance starts swapping.
* 192.168.10.5 / 255.255.255.0
* '''Metadata competes with data inside the ARC.''' {{Code|1=arc_meta_used}} in {{Code|1=qs-iostat -a}} shows the split. A pool with tens of millions of small files can spend most of its cache on metadata, which is an argument for metadata offload devices rather than for more RAM.
* 192.168.11.5 / 255.255.255.0
* '''An L2ARC consumes ARC.''' See below.
* 192.168.12.5 / 255.255.255.0
* '''Deduplication consumes ARC''', for the deduplication table, and it is not optional -- see [[#Metadata, small-block and dedup offload|Metadata, small-block and dedup offload]].
* 192.168.13.5 / 255.255.255.0


==== QuantaStor Network Port (Target Port) Configuration ====
'''Cache Compression''' (on by default) compresses ARC contents, which increases the effective cache size at a small CPU cost. Leave it on. '''Prefetch Disable''' (off by default) turns off the OpenZFS prefetcher, which is worth trying only on a workload that is almost purely random reads, where prefetched blocks are evicting useful ones. Both are on the '''Cache Settings''' tab of [[Storage System Optimization]].
* 192.168.10.10 / 255.255.255.0
* 192.168.11.10 / 255.255.255.0
* 192.168.12.10 / 255.255.255.0
* 192.168.13.10 / 255.255.255.0


There is no need to set the gateway for these ports as they are not going over the internet. We had a separate network that was used for that with iSCSI disabled.
=== Cache policy ===


If the boxes are directly connected you have to make sure that the paired ports are setup using the same network (you should be able to ping all Quantastor ports from Windows).  
'''Cache Policy Primary''' on '''Modify Storage Pool''' -- and the equivalent setting on an individual [[Network Shares|Network Share]] or [[Storage Volumes|Storage Volume]] -- chooses what the ARC holds for that object: {{Code|1=all}}, {{Code|1=metadata}} only, or {{Code|1=none}}.


==== Configuring the Windows iSCSI Initiator ====
This is the lever for a workload that pollutes the cache instead of benefiting from it. A large sequential backup ingest reads each block once, so caching those blocks gains nothing and evicts the working set of everything else on the appliance. Setting that share or volume to {{Code|1=metadata}} keeps its directory structure cached -- which still helps, because metadata is read repeatedly -- while leaving its data blocks out of RAM. The '''Cache Hits''' chart described above is how you identify a candidate: a workload whose hits are almost all on the most-recently-used side is streaming, not reusing.
Now that all the network connections are configured, select "iSCSI Initiator" in the control panel, you can also access it by typing iSCSI into the Windows search bar in the Start menu. Once you have the iSCSI Initiator configuration tool loaded select the Targets tab and type in the IP address of one of the network ports on the QuantaStor system. After it is connected, select the connection and click on properties. From here we can add all the additional sessions by clicking on "Add Session" and then linking the other network ports on the Windows host to their matching network ports on the same subnet on the QuantaStor System.
After clicking on add session, check the box that enables multi-path, and click on the "Advanced" button. We now want to set the local adapter to "Microsoft iSCSI Initiator", the initiator ip as the ip of the Windows port that is to be added, and the target portal ip as the matching network port on the Quantastor box. Click "Ok" on both windows. We will want to add a session for every other port on the Windows box that will be used. In the properties window we should now see multiple sessions. The last step is to click "MCS" at the bottom the the properties window. Make sure the MCS policy is set to "Round Robin" and click "Ok" on all the open windows. Everything should now be setup for MPIO. If you really want to make sure all the sessions are connected to the disk you can open "Disk Management", right click on the disk section (to the left of the blue bar) and view its properties. Under the MPIO tab there should be a session for every iSCSI session you added.  For example, if you have 4 network ports on your Windows host then you should have 4 iSCSI sessions which will produce 4 paths, hence MPIO should show 4 paths in the properties page for your QuantaStor disk.


=== iSCSI Session Creation for MPIO ===
== SSD read cache (L2ARC) ==
Below are some screenshots during the process of setting up LACP in Windows.


[[File:Iscsi.png|150px]] [[File:Iscsi_quick_connect.png|150px]] [[File:Iscsi_properties.png|150px]] [[File:Iscsi_advanced_settings.png|150px]] [[File:Iscsi_mcs.png|150px]] [[File:Multi_path_disk.png|150px]]
An L2ARC is a second-level read cache on SSD, below RAM and above the pool. Blocks evicted from the ARC land there, so a subsequent read is served from flash instead of from the data groups. Read cache devices are added from '''Add Log/Cache/Metadata Device(s)''', need no redundancy -- losing one costs you cache, not data -- and can be removed at any time. See [[Storage Pools]] for the dialog.


Click on the images to enlarge
Whether it pays for itself depends entirely on the workload:


=== Results ===
'''It helps when''' the working set is larger than RAM but not enormously larger, and blocks are re-read: a virtual machine estate whose common OS blocks are read by many guests, a database whose hot indexes exceed RAM, a file share with a recurring active set. Size it to roughly the application's working set.


Below are some of the IO performance results. The first two images are when the Windows box was directly connected to the Quantastor box, where as the next two images are when a switch was used between the boxes.
'''It is wasted when''':


[[File:performance_mpio_direct.png|200px]]  [[File:performance_mpio_direct_2gblength.png|200px]]  [[File:performance_mpio_direct_switch.png|200px]]  [[File:performance_mpio_direct_2gblength_switch.png|200px]]
* '''The pool is already all-flash.''' A read cache in front of flash adds a layer without adding speed.
* '''The workload streams.''' Data read once and never again -- backup ingest, single-pass media playback of a large library -- populates the cache with blocks nobody will ask for twice.
* '''RAM was the cheaper answer.''' If the working set nearly fits in RAM, more RAM beats an L2ARC, because ARC hits are far faster and cost no ARC overhead.
* '''It is oversized.''' Every block held in the L2ARC needs an index header in RAM, inside the ARC. An L2ARC sized far beyond the working set therefore takes RAM away from the cache that is faster than it -- and can leave you slower than with no read cache at all. {{Code|1=l2_hdr_size}} in {{Code|1=qs-iostat -a}} is that cost, in bytes.


Click on the images to enlarge
Two measurements settle the argument. {{Code|1=l2_size}} shows how much of the cache actually filled -- a cache that never fills is bigger than the workload needs. {{Code|1=l2_hits}} against {{Code|1=l2_misses}} shows whether it is being used at all; both near zero after several days of representative load means the cache is not earning its slot.


== Link Aggregation Control Protocol (LACP) ==
'''Judge an L2ARC after days, not hours.''' It has to learn which blocks are worth holding, and OpenZFS deliberately fills it slowly so that a burst of cold reads cannot flush it. A read cache benchmarked in the first hour will look useless whatever the workload.


LACP (link aggregation control protocol) involves bonding ports to act as a single network connection via a shared MAC address. Instead of having to split the traffic across multiple networks, setting up LACP allows for the ports to be viewed as one large port. This also means that if one of the connects fails, the LACP will still be able to function (just at reduced performance).  It is an alternative approach which doesn't depend on MPIO but you can bond groups of ports together and use both technologies if you wish.
Watch {{Code|1=l2_cksum_bad}} as a health signal rather than a performance one: a rising count almost always means the SSD is failing and should be replaced.


For our setup the four network connections on the Windows side were all bonded together, and the four network connections on the Quantastor side were all bonded together. They were then connected to each other through our switch.  In this configuration MPIO sees only one path as there is only one iSCSI session.
== Write log (SLOG/ZIL) and sync policy ==


=== Setup ===
The ZFS intent log (ZIL) is how OpenZFS keeps a promise. When an application issues a write and asks for it to be durable before the call returns, the data has to reach stable storage immediately, even though the pool would rather batch it into the next transaction group a few seconds later. The ZIL is that immediate record. By default it lives in the data groups; a '''write log''' (SLOG) device moves it onto dedicated flash so that a synchronous write costs one fast flash write rather than a scattered write into a parity stripe.


Setting up LACP on the Quantastor side is very simple. First create a bonded group of Ethernet ports. This can be done by first navigating to the "Network Ports" tab under the "Storage System" section. Right click in the empty space under the network ports and select "Create Bonded Port". Assign the port the desired IP address and select the ports to be bonded. In our setup we bonded all four ports together. The last step is to right click on the storage system and select "Modify Storage System". At the very bottom under "Network Bonding Policy", select LACP. Everything is now setup for LACP on the Quantastor side.
'''Sync Policy''' on '''Modify Storage Pool''' decides how much traffic goes that route:


Setting up LACP on the Windows side requires a little bit more effort. Navigate to Control Panel --> Network and Internet --> Network Connections. On one of the network connections, right click and select "Properties". In the properties window select "Configure". In the "Teaming" tab check the box "Team this adapter with other adapters" and click "New Team". Provide a name for the team, and then select the network connections that are to be bonded together. For team type choose "IEEE 802.3ad Dynamic Link Aggregation". The next step is to choose the profile to apply to the team. My setup used the "Standard Server" option. Now that your team is configured select it from the list of network connections and assign an IP address.
{| class="wikitable"
! Policy !! Behaviour !! When
|-
| {{Code|1=standard}} (default) || Honours the application's {{Code|1=O_SYNC}} flag: writes that ask for durability are logged, everything else is batched. || Almost always. It is a hybrid, and it gives hypervisors and databases the durability they ask for without penalising anything else.
|-
| {{Code|1=always}} || Every write goes through the log. || Only when an application needs durability it does not ask for, and only with a log device fast enough to absorb the whole write stream.
|-
| {{Code|1=disabled}} || All writes are asynchronous. Acknowledged before they are durable. || Not supported for production. It can lose acknowledged writes on an unexpected power failure, and the dialog raises a confirmation.
|}


The last step is to configure your switch to used LACP. Being that this will be different from switch to switch you will have to check your user manual on how to configure LACP.
'''This is why a new write log often looks broken.''' Under the default {{Code|1=standard}} policy, only writes that carry {{Code|1=O_SYNC}} touch it -- hypervisors and databases set the flag; most other applications do not. Add a write log for a workload that does not, and it will sit idle: {{Code|1=zil_commit_count}} in {{Code|1=qs-iostat -a}} barely moves. That is the log correctly doing nothing, not a fault.


Below are some screenshots during the process of setting up LACP in Windows.  
Setting '''Sync Policy''' to {{Code|1=always}} forces every write through it, and this is where it goes wrong: if the log device cannot sustain the full write rate, forcing all writes through it makes the pool '''slower''', not faster. Check {{Code|1=zil_commit_stall_count}} and {{Code|1=zil_commit_suspend_count}} -- anything other than zero under load means the log is the bottleneck.


[[File:LACP_LANSettings.png|150px]] [[File:LACP_LANProperties.png|150px]] [[File:LACP_Configure.png|150px]] [[File:LACP_IEEE.png|250px]] [[File:LACP_Profile.png|250px]]
Sizing and media choice:


Click on the images to enlarge
* '''Write log devices see constant, small, synchronous writes''', which is the hardest duty cycle in the appliance. Use high-endurance enterprise SSDs; SLC-class devices are worth the money here in a way they are not for a read cache.
* '''Log devices are mirrored automatically''', and multiple pairs scale log throughput. See [[Storage Pools]] for the pairing rules and a worked sizing example.
* '''Modern deployments need a write log less than they used to.''' Its value was in making an HDD pool usable for database and virtual machine workloads, and those workloads belong on flash now that NVMe and SAS SSDs are affordable. An all-flash pool rarely benefits from one.


=== Cautions with LACP ===
The same sync policy choice is available per [[Network Shares|Network Share]] and per [[Storage Volumes|Storage Volume]], which is the right granularity when one dataset on a pool needs {{Code|1=always}} and the rest does not.


Be careful if you decide to use LACP. I would highly recommend first trying the MPIO setup like talked about above. After doing a performance run of MPIO, setup LACP and run the performance tests again. Doing this will allow you to have a target IO performance mark to try and meet. With our testing LACP seemed to be very sensitive to different filtering techniques that were allowed by the switch. Turning LACP on while leaving all of the default settings resulted in a large drop in IO performance. This will be different from switch to switch but a drop in performance by 80% (roughly what we saw) was discomforting. After some tweaking of the settings in the switch the performance became closer to what was seen with MPIO.
== Metadata, small-block and dedup offload ==


=== Results ===
'''On an HDD pool this is the largest single performance improvement available''', and it is worth understanding why before deciding how much SSD to spend on it.
{|
 
A metadata offload group (a ZFS special vdev) holds the pool's metadata on mirrored flash instead of scattered across the data groups. Metadata access is small and random by nature, which is the pattern HDDs are worst at, and it is on the critical path of far more than it looks:
 
* '''Scrubs and resilvers walk the metadata tree.''' Moving it to flash shortens both substantially -- which matters most exactly when you can least afford it, during a rebuild with reduced redundancy.
* '''Replication and snapshot operations are metadata-heavy''', so a replication window that will not close is often a metadata problem rather than a bandwidth problem.
* '''Directory traversal and file enumeration''' are metadata. This is what makes a large HDD file share feel slow to browse even when throughput is fine.
* '''HA failover''' has to read pool metadata before the pool comes up.
 
'''Small Block Offload''' extends this to data. Set on '''Modify Storage Pool''' (and per share or volume), it routes any write below the chosen size to the metadata group instead of the data groups. The reason it helps so much on parity pools is arithmetic: writing a few kilobytes into a RAIDZ stripe means reading and rewriting parity across the whole stripe, so a small write costs far more than its size. Sending it to a mirrored SSD group instead avoids that entirely. The setting offers Disabled, 4K, 8K, 16K, 32K, 64K, 128K and 256K, and stays Disabled until a metadata group exists.
 
Practical guidance:
 
* '''Always include SSDs for metadata and small-block offload on an HDD pool.''' The effect on scrub time, replication, failover and small-file performance is large enough that we do not design HDD pools without it. See [[Storage Pools]] for the recommended device count and starting offload size.
* '''An all-flash pool does not need it and should not use it''' -- with one exception. QLC media handles small-block I/O poorly, so putting metadata and small blocks on TLC in front of a QLC pool is worthwhile.
* '''Do not set the offload size larger than the SSD group can hold.''' The larger the threshold, the more of the pool's writes land on flash. Size the threshold to the capacity you gave it, not to the largest number in the list.
* '''A metadata offload group cannot be removed.''' It is a permanent part of the pool, and losing it loses the pool, so mirror it properly the first time. There is no CLI command to remove one either.
 
'''Deduplication offload''' puts the deduplication table on its own mirrored vdev, and adding one turns deduplication on for the pool. Treat it as a capacity feature with a performance cost, not a performance feature. Every write has to be looked up in the deduplication table, that table has to be cached in RAM to make the lookup fast, and it grows with the amount of unique data in the pool -- so it competes with the ARC for exactly the memory your read workload wants. Like the metadata group, it cannot be removed once added. Use compression first; it gives most of the space saving on most data with none of this cost.
 
== Record size and volume block size ==
 
'''Record size''' (for a [[Network Shares|Network Share]]) and '''Block Size''' (for a [[Storage Volumes|Storage Volume]]) are the same idea: the unit OpenZFS reads, writes, checksums and compresses. Matching it to the application's I/O size is one of the few settings that can change performance by a large factor, and it is one of the few that '''cannot be changed later''' for a volume.
 
Why the match matters:
 
* '''Too large, for small random writes.''' Modifying part of a record forces a read of the whole record, the modification, recompression, and a write of the whole record. An 8 KB database write into a 1 MB record is a 1 MB read plus a 1 MB write. The same amplification applies to reads: a 4 KB read fetches the entire record.
* '''Too small, for large sequential I/O.''' Every record carries checksum and indirect-block overhead, and each one is a separate compression unit, so small records cost more metadata, more IOPS, and less compression than large ones for the same amount of data.
 
{| class="wikitable"
! Object !! Range offered !! Default !! Notes
|-
|-
|
| '''Network Share''' -- Record Size || Auto, 8K, 16K, 32K, 64K, 128K, 256K, 512K, 1M, 2M, 4M, 8M || 128K (via '''Auto''') || Can be changed at any time from '''Modify Network Share'''. Existing data keeps the record size it was written with; only new writes use the new value.
The picture on the left was when LACP was setup with the default settings in the switch we were using. As you can see the performance suffered greatly. On the right is what the results looked like after some tweaking of the settings in the switch. As you can see, with just some minor tweaking large performance differences can occur. Note the change in scale from the picture on the left to the picture on the right.  
|-
|-
|
| '''Storage Volume''' -- Block Size || 4K, 8K, 16K, 32K, 64K, 128K, 256K, 512K, 1M, 2M, 4M || 64K || '''Fixed at creation and cannot be changed afterwards.''' The field greys out on '''Modify Storage Volume'''.
[[File:performance_lacp_default_settings.png|200px]]   [[File:performance_lacp_tweaked_settings.png|200px]]
|}
 
Choosing:
 
* '''Leave a general-purpose file share at the 128K default.''' It is a good compromise and shares rarely have one dominant I/O size.
* '''Databases want a record or block size at or near the database page size''' -- 8K or 16K for most relational engines. This is what the '''Databases''' Storage Volume I/O profile sets, and it is the reason the profile exists.
* '''Virtual machine datastores are well served by 64K''', the volume default, which is what the '''Desktop and Server Virtualization''' profile selects and is compatible across VMware releases. A VM datastore carries a mix of guest I/O sizes, so a middle value beats optimising for either end.
* '''Media, backup and archive shares want large records''' -- 1M and above. Fewer, larger records mean less metadata and better compression on large sequential files.
* '''Match the share record size to the client's own block size where you know it''', particularly for a single-purpose share.
 
Do not confuse either setting with '''Block Size Offset (ashift)''', which is the pool's physical sector alignment, is set at pool creation, and should be left on '''Auto'''. See [[Storage Pools]].
 
Set these from the CLI with <code>[[QuantaStor CLI Command Reference#share-create|qs share-create]] --recordsize=&lt;KB&gt;</code> and <code>[[QuantaStor CLI Command Reference#volume-create|qs volume-create]] --blocksize-kb=&lt;KB&gt;</code>.
 
== Compression ==
 
Compression is on by default -- LZ4 -- and it is normally a '''throughput''' decision rather than a capacity one. A modern CPU compresses faster than the media can read or write, so compressing means moving fewer bytes through the disks, the HBAs and the backplane. On compressible data it makes the pool faster and larger at the same time.
 
That reasoning tells you when to change it:


Click on the images to enlarge
* '''Leave it on (LZ4) unless you have a reason.''' It is cheap enough that it costs nothing measurable on incompressible data, because OpenZFS detects a record that will not compress and stores it uncompressed.
|}
* '''Turn it off for data that is already compressed.''' Media and entertainment content, encrypted archives, and most image and video libraries gain nothing and only add CPU load. This is the common case in post-production.
* '''Consider ZSTD when the pool is media-bound rather than CPU-bound.''' ZSTD compresses harder than LZ4 at more CPU cost. On an HDD pool with spare cores it can raise effective throughput; on an all-flash pool where the media is already faster than the CPU, it will reduce it. Measure both.
* '''Do not reach for GZIP for performance.''' The gzip levels are a capacity choice for cold data, and they cost enough CPU to become the bottleneck on an active pool.
* '''Changing compression affects new writes only.''' Existing blocks keep whatever algorithm they were written with, so a change takes effect gradually as data is rewritten.


== Jumbo Frames ==
The available algorithms and the field itself are on '''Modify Storage Pool''' -- see [[Storage Pools]] -- and the same choice is available per share and per volume, which is the right granularity when one dataset on a pool holds incompressible media and the rest does not.


Jumbo frames depending on the setup can have a large impact on the IO performance. Jumbo frames sends larger sized packets, allowing for larger Ethernet payloads.  This lets more data be sent with each packet and in most cases results in better performance.
Compression interacts with record size: a larger record gives the compressor more to work with and compresses better. It also interacts with partial writes, since a modified record has to be recompressed in full -- another reason not to pair a large record size with a small-write workload.


For our setup I first configured everything as I did with MPIO using a direct connection between the Windows box and the Quantastor box. I then went and changed all the settings as specified below.
== Device write caches ==


=== Setup ===
Worth knowing, because it explains a measurement that otherwise looks wrong. '''QuantaStor disables the volatile write-back cache on rotational disks''' as they are brought into service, putting them into write-through mode, because an HDD's on-board cache is not power-loss protected and acknowledging a write that is still only in that cache risks the pool's consistency.


To setup Quantastor to use jumbo frames first navigate to the storage system section. Then under the "Network Ports" tab, right click on the port in which you would like to configure to use jumbo frames. Select "Modify Network Port" from the context menu. Now set the MTU as 9000. Complete this for every port you plan on using jumbo frames with.
Enterprise and datacenter SSDs and NVMe devices keep their write cache enabled, because they have the capacitors to flush it on power loss. Devices reached over iSCSI are put into write-through regardless.


To setup Windows to use jumbo frames first navigate to Control Panel --> Network and Internet --> Network Connections. On one of the network connections, right click and select "Properties". In the properties window select "Configure". From here under the advanced tab there should be an option for jumbo packets. Select this and change its value to 9000. Complete this for every port you plan on using jumbo frames with.
So a raw HDD write benchmark on the appliance will fall short of the drive's data sheet, and that is deliberate. The correct way to get synchronous write performance back is a write log on flash, not a volatile cache.


You may also have to go into your switch and enable jumbo frames as well. This will vary from switch to switch.
== Network-side factors ==


=== Cautions with Jumbo Frames ===
The network is frequently the real limit, and it is the layer where a small configuration mistake costs the most. [[Network Ports]] owns the port, VLAN and bonded-port dialogs; what follows is which choice to make and why.


For jumbo frames to work the MTU must be set as the same for all devices (all the connected machines and switches). If the packet is not recognized by a device, the packet can be dropped. This is just one example of an issue with jumbo frames. When we tried configuring our setup to go through a switch we encountered issues such as this.
=== MTU and jumbo frames ===


=== Results ===
{{Navigation|Storage Management &rarr; Network Port &rarr; Modify ''(toolbar)''}}


As you can see below jumbo frames had the best results out of what we tested. The read IO speeds was the largest by far of all of our testing, and the write IO performed very strongly as well.  
A larger MTU means more payload per frame and fewer frames, headers and interrupts for the same data -- a real gain on a 10&nbsp;GbE or faster storage network carrying large block I/O. '''Modify Network Port''' has an '''MTU''' field with a '''Jumbo Frames''' button beside it that fills in 9000; once the port is on 9000 the button reads '''Default Frames''' and puts it back to 1500. The MTU field is only editable while the '''Static IP Configuration Settings''' section is active, and is disabled on VLAN and alias ports, which take the MTU of their parent.


[[File:performance_direct_jumbo_2gb.png|200px]]  [[File:performance_direct_jumbo.png|200px]]
'''Every device in the path must agree.''' The initiator, every switch port between, and the target all have to carry the same MTU. A device that receives a frame larger than its MTU drops it, and the symptom is not a clean failure -- it is a connection that works for small transfers and stalls on large ones, which is a genuinely unpleasant thing to diagnose. Set the switch first, verify end to end, and change the appliance last.


Click on the images to enlarge
Set it from the CLI with <code>[[QuantaStor CLI Command Reference#network-port-modify|qs network-port-modify]] --port=&lt;port&gt; --mtu=9000</code>.


=== ZFS Performance Tuning ===
=== Multiple interfaces versus bonding ===


One of the most common tuning tasks that is done for ZFS is to set the size of the ARC cache.  If your system has less than 10GB of RAM you should just use the default but if you have 32GB or more then it is a good idea to increase the size of the ARC cache to make maximum use of the available RAM for your storage system.  Before you set the tuning parameters you should run 'top' to verify how much RAM you have in the system.  Next, run this command to set the amount of RAM to some percentage of the available RAM. For example to set the ARC cache to use a maximum of 80% of the available RAM, and a minimum of 50% of the available RAM in the system, run these, then reboot:
There are two ways to use several ports for one workload, and they are not interchangeable.
<pre>
qs-util setzfsarcmax 80
qs-util setzfsarcmin 50
</pre>


Example:
'''Multipath (MPIO), one subnet per port.''' Each port gets a static address on its own subnet, the initiator opens a session to each, and multipathing on the host spreads I/O across them and survives the loss of any one. This is the approach we recommend trying first for block storage, because it needs nothing from the switch, it scales linearly with ports, and its failure modes are visible.
<pre>
sudo -i
qs-util setzfsarcmax 80
INFO: Updating max ARC cache size to 80% of total RAM 1994 MB in /etc/modprobe.d/zfs.conf to: 1672478720 bytes (1595 MB)
qs-util setzfsarcmin 50
INFO: Updating min ARC cache size to 50% of total RAM 1994 MB in /etc/modprobe.d/zfs.conf to: 1045430272 bytes (997 MB)
</pre>


'''Putting several ports on the same subnet is the mistake to avoid here.''' With multiple interfaces on one subnet, Linux by default will answer an ARP request for any local address out of any interface, so return traffic can leave a port other than the one the request arrived on. The paths you carefully separated then collapse onto whichever port answered, and throughput sits at roughly one port's worth however many you configured. QuantaStor's '''ARP Filtering''' setting on '''Modify Storage System''' addresses this -- '''Auto''' (the default) enables ARP filtering only when a bonded port is present, so a system with several unbonded ports on one subnet does '''not''' get it. If your design puts multiple ports on the same subnet, set ARP Policy explicitly to '''Enabled''' -- the CLI equivalent is <code>[[QuantaStor CLI Command Reference#system-modify|qs system-modify]] --arp-filter-mode=enabled</code>. Better still, give each port its own subnet and avoid the question.


To see how many cache hits you are getting you can monitor the ARC cache while the system is under load with the qs-iostat command:
For Windows initiators, one detail is worth recording because it is easy to get wrong and produces a confusing result. QuantaStor identifies itself over SCSI as vendor {{Code|1=OSNEXUS}}, product {{Code|1=QUANTASTOR}}, so the string to add under '''MPIO Devices''' is {{Code|1=OSNEXUS QUANTASTOR}} followed by six trailing spaces -- an eight-character vendor field and a sixteen-character product field, both space padded. Without exactly the right padding the MPIO driver does not recognise the devices and Windows Disk Management shows the same disk once per path instead of one disk with several paths.


<pre>
'''Bonding.''' A bonded port presents several physical ports as one logical interface with one address.
qs-iostat -a


Name                              Data
{{Navigation|Storage Management &rarr; Network Port &rarr; Create Bonded Port ''(toolbar)''}}
---------------------------------------------
hits                              237841
misses                            1463
c_min                            4194304
c_max                            520984576
size                              16169912
l2_hits                          19839653
l2_misses                        74509
l2_read_bytes                    256980043
l2_write_bytes                    1056398
l2_cksum_bad                      0
l2_size                          9999875
l2_hdr_size                      233044
arc_meta_used                    4763064
arc_meta_limit                    390738432
arc_meta_max                      5713208


Each '''Bond Mode''' in the dialog names the switch it needs, which is the part to read first:


ZFS Intent Log (ZIL) / writeback cache statistics
{| class="wikitable"
! Bond Mode !! Switch required
|-
| Link Aggr Ctrl Protocol (LACP layer2) || Managed switch. The dialog's default.
|-
| Link Aggr Ctrl Protocol (LACP layer2+3) || Managed switch. Hashes on MAC and IP, which spreads better than layer2 alone across many peers.
|-
| Link Aggr Ctrl Protocol (LACP layer3+4) || Managed switch. Hashes on IP and port, so two sessions between the same pair of hosts can land on different members.
|-
| Round Robin (balance-rr) || Etherchannel managed switch.
|-
| Balance XOR (balance-xor) || Etherchannel managed switch.
|-
| Active-Backup (active-backup) || Unmanaged switch. Failover only, with no throughput gain -- one member carries all traffic.
|-
| Adaptive Transmit Load Balancing (balance-tlb) || Unmanaged switch. Balances outbound traffic only.
|-
| Adaptive Load Balancing (balance-alb) || Unmanaged switch. Balances both directions without switch support.
|}


Name                              Data
Bonding is the right choice when you need one address -- for file protocols, or where the client cannot do multipathing -- and when the ports are spread across switches for redundancy, which requires LACP and switch infrastructure that supports it. Two cautions:
---------------------------------------------
 
zil_commit_count                  876
* '''A bond does not make one session faster.''' The load-balancing modes hash each flow onto one member port, so a single TCP connection gets one port's bandwidth no matter how many are in the bond. Aggregate throughput across many clients improves; one client's single stream does not. With iSCSI over a bond there is one session, so multipathing sees one path.
zil_commit_writer_count          495
* '''LACP performance depends heavily on the switch's hash policy.''' Enabling LACP and leaving the switch on its defaults has been measured to cost the large majority of the available throughput, recovered only after matching the switch's hashing to the traffic. This varies between switch vendors, so '''benchmark multipath first and use that number as the target LACP has to meet''' -- without a target you have no way to tell a badly hashed bond from a fast one.
zil_itx_count                    857
 
</pre>
A '''Network Bonding Policy''' setting on '''Modify Storage System''' carries a system-level bonding mode drawn from a subset of the same list; see [[Storage System]].
 
Where you use both approaches, bond groups of ports and run multipath across the bonds.


A description of the different metrics for ARC, L2ARC and ZIL are below.
=== NIC ring buffers ===


<pre>
'''Optimize hardware RX/TX buffer settings for throughput''' on '''Modify Network Port''' raises the network card's receive and transmit ring buffers. Larger rings give the driver more room to absorb a burst before dropping packets, which matters on a fast link carrying large block I/O or replication traffic. QuantaStor picks the largest power-of-two value that stays safely below the card's hardware maximum, records the original values so that clearing the checkbox restores them, and re-checks the setting periodically.
hits = the number of client read requests that were found in the ARC
misses = the number of client read requests were not found in the ARC
c_min = the minimum size of the ARC allocated in the system memory.
c_max = the maximum size of the ARC that can be allocated in the system memory.
size = = the current ARC size
l2_hits = the number of client read requests that were found in the L2ARC
ls_misses = the number of client read requests were not found in the L2ARC
ls_read_bytes = The number of bytes read from the L2ARC ssd devices.
l2_write_bytes = The number of bytes written to the L2ARC ssd devices.
l2_chksum_bad = The number of checksums that failed the check on an SSD (a number of these occurring on the L2ARC usually indicates a fault for a SSD device that needs to be replaced)
l2_size = the current L2ARC size
l2_hdr_size = The size of the L2ARC reference headers that are present in ARC Metadata
arc_meta_used = The amount of ARC memory used for Metadata
arc_meta_limit =  The maximum limit for the ARC Metadata
arc_meta_max = The maximum value that the ARC Metadata has achieved on this system


zil_commit_count = How many ZIL commits have occurred since bootup
The checkbox is unavailable on a bonded port and on virtual ports -- tune the member ports instead -- and on a port whose configuration type is disabled. The CLI equivalent is <code>[[QuantaStor CLI Command Reference#network-port-modify|qs network-port-modify]] --port=&lt;port&gt; --auto-tune=true</code>.
zil_commit_writer_count = How many ZIL writers were used since bootup
zil_itx_count  = the number of indirect transaction groups that have occurred sinc bootup
</pre>


=== Pool Performance Profiles ===
Two related system tunables sit on the '''Network Settings''' tab of [[Storage System Optimization]]: '''Network TX/RX Queue Length''' (the software queue, default 5000, and worth raising on 10&nbsp;GbE and faster) and '''Network Device Max Backlog''' (how many packets may queue on the receive side when the interface delivers faster than the kernel can process).


Read-ahead and request queue size adjustments can help tune your storage pool for certain workloads.  You can also create new storage pool IO profiles by editing the /etc/qs_io_profiles.conf file.  The default profile looks like this and you can duplicate it and edit it to customize it.
== What does not apply ==


<pre>
Three things that look like tuning levers and are not:
[default]
name=Default
description=Optimizes for general purpose server application workloads
nr_requests=2048
read_ahead_kb=256
fifo_batch=16
chunk_size_kb=128
scheduler=deadline
</pre>


If you edit the profiles configuration file be sure to restart the management service with 'service quantastor restart' so that your new profile is discovered and is available in the web interface.
* '''Hardware RAID card settings.''' Scale-up pools are deployed on HBAs, not on hardware RAID controllers. OpenZFS needs direct access to the drives to checksum, self-heal and manage its own redundancy, and a RAID card's cache, stripe size and read-ahead settings sit between it and the media doing none of those things. There is no RAID-card tuning to do on a scale-up pool, because there should be no RAID card. Pick the layout in [[Storage Pools]] instead.
* '''Linux software RAID (mdadm) parameters.''' The {{Code|1=[mdadm]}} section of {{Code|1=/etc/quantastor.conf}} and the {{Code|1=chunk_size_kb}} key in the I/O profiles file are legacy settings from pool types the product no longer builds. They have no effect on a scale-up pool.
* '''The {{Code|1=[device]}} section of {{Code|1=/etc/quantastor.conf}}.''' Read-ahead, queue depth and scheduler were once configured there. They are not any more -- those settings come from the pool's I/O profile, and the keys in that section are inert. Edit the profile, not the configuration file.


=== Storage Pool Tuning Parameters ===
Two further notes on scope. Most of the system tunables in [[Storage System Optimization]] are OpenZFS module parameters, so they '''affect scale-up pools only''' and do nothing for scale-out (Ceph) pools. And tunables are stored '''per Storage System''', so in a grid each member is tuned independently -- save a profile and apply it to the others rather than editing each by hand.


QuantaStor has a number of tunable parameters in the /etc/quantastor.conf file that can be adjusted to better match the needs of your application.  That said, we've spent a considerable amount of time tuning the system to efficiently support a broad set of application types so we do not recommend adjusting these settings unless you are a highly skilled Linux administrator.
== Related pages ==
The default contents of the /etc/quantastor.conf configuration file are as follows:
<pre>
[device]
nr_requests=2048
scheduler=deadline
read_ahead_kb=512


[mdadm]
* [[Performance Testing]] -- measuring first: qs-perftest, qs-ramdisk and the phased method for isolating which layer is slow
chunk_size_kb=256
* [[Storage System Optimization]] -- the system tunables dialog, all 33 settings, and the tunable profile mechanism
parity_layout=left-symmetric
* [[Storage Pools]] -- pool layout, the Create and Modify dialogs, and the offload device tiers
</pre>
* [[Storage Volumes]] -- provisioning block storage, and where volume block size is set
* [[Network Shares]] -- provisioning file storage, and where record size is set
* [[Physical Disks/Devices]] -- disk health, predictive failure warnings and per-disk operations
* [[Storage Pool Cache]] -- the Add Log/Cache/Metadata Devices dialog in detail
* [[Network Ports]] -- network port, VLAN and bonded port configuration
* [[Performance Monitoring]] -- the Grid Dashboard
* [[Storage System]] -- the Modify Storage System dialog, including ARP filtering and the default bonding policy
* [[Alert Manager]] -- getting predictive disk failure and capacity alerts delivered
* [[QuantaStor CLI Command Reference]] -- full argument lists for every command above


There are tunable settings for device parameters which are applied to the storage media (SSD/SATA/SAS), as well as settings like the MD device array chunk-size and parity configuration settings used with XFS based storage pools.  These configuration settings are read from the configuration file dynamically each time one of the settings is needed so there's no need to restart the quantastor service. Simply edit the file and the changes will be applied to the next operation that utilizes them. For example, if you adjust the chunk_size_kb setting for mdadm then the next time a storage pool is created it will use the new chunk size.  Other tunable settings like the device settings will automatically be applied within a minute or so of your changes because the system periodically checks the disk configuration and updates it to match the tunable settings. 
----
Also, you can delete the quantastor.conf file and it will automatically use the defaults that you see listed above.
<small>''Verified against QuantaStor 6.9.0.''</small>

Latest revision as of 08:30, 3 September 2026


This page covers how to tune a scale-up (OpenZFS) QuantaStor appliance for a workload: what to measure, what each tuning control actually changes underneath, and how to choose between them. It is a decision guide rather than a field reference -- where a control lives on another page's dialog, the reasoning is here and the field-by-field detail is linked.

Find out what is slow before you change anything. Performance Testing covers the measurement tools the appliance ships -- qs-perftest, qs-ramdisk and the Disk Performance Test dialog -- and a method for narrowing a performance complaint down to the media, the network, or the pool. Tuning a layer that is not the problem costs time and can make things worse.

Two pages own the controls this one reasons about. Storage System Optimization owns the appliance-wide OpenZFS and kernel tunables dialog. Storage Pools owns pool creation and modification, including the I/O profile, compression, sync and cache policy fields.

Section Purpose
How to approach a tuning problem The order of operations, and what not to do.
What to measure The dashboards, qs-iostat, and where the measurement tools live.
Where the knobs are Five separate tuning layers and which one a given problem belongs to.
Storage Pool I/O profiles Read-ahead, request queue depth, I/O scheduler, and the seven shipped profiles.
Storage Volume I/O profiles Volume block size and the iSCSI session parameters.
RAM read cache (ARC) Sizing the ARC and what it costs in RAM.
SSD read cache (L2ARC) When a read cache pays for itself and when it is wasted.
Write log (SLOG/ZIL) and sync policy Why a write log often sits idle, and what changes that.
Metadata, small-block and dedup offload The single largest win available on an HDD pool.
Record size and volume block size Matching the on-disk unit to the application's I/O size.
Compression Why compression is usually a throughput decision, not a capacity one.
Device write caches Why an HDD in a pool is slower than its data sheet.
Network-side factors MTU, multiple interfaces versus bonding, and NIC ring buffers.
What does not apply Controls that look relevant and are not.

How to approach a tuning problem

The shipped defaults are right for the large majority of deployments, and most disappointing performance turns out to be a design problem rather than a tuning problem: too few device groups, a parity stripe that is too wide for a random workload, HDDs with no metadata offload, or a single network path. Tuning cannot fix any of those. Work through the design first -- see Where the knobs are and Storage Pools -- and only then reach for a knob.

When you do tune:

  1. Measure first, with a workload that resembles production. A number you cannot reproduce is not a baseline, and without a baseline you cannot tell an improvement from noise.
  2. Confirm the hardware is healthy. Every device in Physical Disks should be Normal. One disk with a predictive-failure warning will make a whole pool look badly tuned. See Physical Disks/Devices.
  3. Change one group of related settings, then measure again. Several of the OpenZFS tunables interact, and a batch of simultaneous changes leaves you with no way to attribute the result.
  4. Save a named profile before experimenting so you can get back to a known state. See Storage System Optimization for tunable profiles and Custom I/O profiles below for pool profiles.
  5. Write down what you changed. Tuning applied per system in a grid drifts silently between members otherwise.

We recommend involving OSNEXUS support before changing the system tunables. They are the settings that matter for specific, identifiable problems -- write latency spikes under load, a resilver that will not finish inside a maintenance window, a replication job overrunning its window -- and changing them speculatively is more likely to cost performance than gain it.

What to measure

Reading the dashboards

The appliance's own dashboards are the first place to look, and Performance Monitoring documents them in full -- the views, what each series means, and the sampling and retention behaviour that decides whether a change is even visible in a given chart. Three points from there matter specifically when tuning:

  • Check the time range before drawing a conclusion. Windows longer than six hours are served from averaged data, so a short latency spike caused by a tuning change can be averaged away entirely. Use a range of six hours or less when assessing a change you just made.
  • Cache Efficiency is a since-boot average, not a current hit rate. It moves very slowly on an appliance with a long uptime, so it is close to useless for judging an ARC change made minutes ago. Watch the hit counts instead, or read the counters directly with qs-iostat -a.
  • Per-port network throughput is on the same Performance view as CPU, which makes it the quickest way to tell a network-bound workload from a storage-bound one before touching any pool setting.

Selecting an individual disk replaces the dashboard with a per-device view of IOPS, throughput and latency -- the fastest way to confirm that load is landing where you expect during a test.

qs-iostat

qs-iostat is the console-level counterpart, and it reports the ARC and ZIL kernel counters that the charts summarise. It wraps iostat:

qs-iostat -c            # globally averaged CPU stats
qs-iostat -d            # I/O stats for all block devices, in MB/s
qs-iostat -a            # ZFS ARC, L2ARC and ZIL counters
qs-iostat -f            # repeat every 2 seconds
qs-iostat --extra "-x"  # pass extra arguments through to iostat

qs-iostat -a is the useful one for cache work. The counters are cumulative since boot, so take two samples and compare, rather than reading absolute numbers:

Counter Meaning
hits / misses Read requests satisfied from, and missed in, the RAM cache. The ratio between two samples is the live hit rate.
size Current ARC size.
c_min / c_max The floor and ceiling the ARC may grow between. c_max is what Cache Size (% of RAM) sets.
arc_meta_used How much of the ARC is metadata rather than data. On a pool with many small files this can be most of it.
l2_hits / l2_misses Reads satisfied from, and missed in, the SSD read cache. Both zero means the L2ARC is doing nothing.
l2_size How much of the read cache is actually populated. A cache that never fills is larger than the workload needs.
l2_hdr_size RAM consumed by the L2ARC's index. This is the hidden cost of a read cache -- see SSD read cache (L2ARC).
l2_read_bytes / l2_write_bytes Bytes read from and written to the read cache devices.
l2_cksum_bad Checksum failures on a read cache device. A rising count usually means the SSD is failing and should be replaced.
zil_commit_count Synchronous write commits since boot. Zero, or near zero, on a pool with a write log means the log is idle -- see Write log (SLOG/ZIL) and sync policy.
zil_commit_writer_count Commits that had to write, as opposed to joining a commit already in flight.
zil_commit_stall_count / zil_commit_suspend_count Commits that stalled or were suspended. Anything other than zero under load points at the log device not keeping up.
zil_itx_count Intent log transactions recorded.

(Note: these counters come straight from the OpenZFS kernel modules and the set changes between OpenZFS releases; /proc/spl/kstat/zfs/arcstats and /proc/spl/kstat/zfs/zil on the appliance are authoritative.)

Testing the disks and the pool directly

Physical Disk Performance Test. The Read Seq and Last Performance Test columns keep the result on each disk, so a later run can be compared against it.

The dashboards and qs-iostat tell you what the appliance is doing under its current load. To make it do something measurable on purpose, the appliance ships two utilities and a dialog:

  • Disk Performance Test, on a physical disk's right-click menu, reads sequentially from the disks you select and stores the result on each one. It is the fastest way to answer "is one device in this pool slower than its peers?" -- see Physical Disks/Devices for its fields and modes.
  • qs-perftest read-tests every disk or every disk in one pool from the command line, and benchmarks a pool by creating temporary shares and driving fio, elbencho, dd and iozone against them.
  • qs-ramdisk creates transient RAM-backed disks, so a benchmark can be run with the media factored out entirely -- which is how you establish that a bottleneck is above the media rather than in it.

Performance Testing covers all three, and the order to use them in. Work through it before changing anything on this page: a degraded device or a misconfigured bond produces exactly the symptoms that tempt an administrator into the tunables, and no setting here will fix either.


Where the knobs are

Five separate mechanisms are easy to confuse. Deciding which layer a problem belongs to saves most of the work:

Layer Scope What it controls Where
Storage Pool I/O profile one pool, re-applied at every pool start Block device settings on the member disks: read-ahead, request queue depth, I/O scheduler. Also raises the SCST target driver thread floor for the whole appliance. below; the field is on Create and Modify Storage Pool (Storage Pools)
Pool, share and volume properties one pool, share or volume Record size and volume block size, compression, sync policy, primary and secondary cache policy, small-block offload. below; the fields are on Storage Pools, Network Shares and Storage Volumes
Storage Volume I/O profile one volume Volume block size and the iSCSI session parameters on that volume's target. below
System tunables one Storage System OpenZFS and SPL module parameters plus two network queue settings: ARC size, write throttle, per-VDEV queue depths, resilver and scrub priority, prefetch distance. Storage System Optimization
Pool topology and offload tiers one pool, fixed at build time (mostly) Number and width of device groups, mirroring versus parity, and the write log, read cache, metadata and dedup vdevs. Storage Pools

The last row is not a tuning layer, but it is where the large factors live. A pool with one wide parity group and no flash offload cannot be tuned into a random-I/O pool, and the first four layers together will not move it as far as adding device groups will.

Storage Pool I/O profiles

A Storage Pool I/O profile is a bundle of Linux block device settings applied to the pool's member disks. It is not a filesystem or OpenZFS setting. QuantaStor re-applies it whenever the pool is created or started, on HA failover, when the profile is changed on Modify Storage Pool, and whenever cache, log, metadata or hot-spare devices are added or removed -- so the tuning survives the events that would otherwise leave new or reconnected devices on kernel defaults. Each pool has one profile; the field is on both Create Storage Pool (Advanced Settings) and Modify Storage Pool -- see Storage Pools for the dialogs.

Each profile carries a separate set of values for HDD, SSD and NVMe media, and QuantaStor picks the set per device from what the device reports, so one profile tunes a mixed pool correctly. On a multipath pool the settings are applied to each underlying SCSI path rather than to the dm device, which is where they actually take effect.

What each setting changes

Read-ahead (/sys/block/<dev>/queue/read_ahead_kb) is how much extra data the kernel pulls in after satisfying a read. If a read asks for 4 KB and read-ahead is 256 KB, another 256 KB is read into the page cache behind it.

Read-ahead exists to hide rotational latency. A 7200 RPM drive turns 120 times a second, roughly once every 8 ms. Reading 4 KB records one per rotation yields under 500 KB/s -- so on any workload with sequential locality, reading ahead and having the next records already in memory is the difference between a usable HDD pool and an unusable one. That is why read-ahead and queue depth matter so much on spinning media and why the archive and media profiles push read-ahead up to 512 KB and 1 MB.

Flash has no rotational latency, so read-ahead on flash is pure overhead: every speculative block occupies bandwidth and cache that a real request wanted. Every shipped profile uses a 4 KB SSD read-ahead, and that is a deliberate result rather than a rounding to zero -- the original SSD tuning work found that both 0 KB and larger values produced pronounced dips at 4 KB and 8 KB request sizes that a small 4 KB read-ahead removed. It remains the single most important SSD-side setting in these profiles.

Request queue depth (/sys/block/<dev>/queue/nr_requests) is how many I/O requests the kernel will hold queued for a device. A deeper queue gives the scheduler and the device more requests to sort and coalesce, which raises throughput on rotational media; too deep a queue raises latency, because a small urgent read now waits behind a long queue of other work.

Profiles express this two ways. A fixed count sets the queue depth outright. A multiplier instead reads the device's own hardware queue depth from /sys/block/<dev>/device/queue_depth and multiplies it, so the same profile scales sensibly across drives with very different capabilities. Two rules matter:

  • The multiplier takes precedence over the fixed count whenever the device reports a non-zero hardware queue depth. The fixed count is the fallback for devices that report none.
  • The computed value is capped at 1024. A multiplier against a deep hardware queue is clamped there, with a warning in /var/log/qs/qs_service.log.

I/O scheduler (/sys/block/<dev>/queue/scheduler) is the elevator algorithm that decides the order requests reach the device. Two are relevant. deadline sorts requests to reduce head movement while enforcing a latency deadline so nothing starves -- the right choice for HDDs. noop does essentially no reordering and hands requests straight down, which is right for flash, where there is no seek cost to optimise away and reordering only adds CPU work and latency. Current kernels name these mq-deadline and none; QuantaStor accepts either spelling in a profile and writes whichever the device offers. If a device offers neither, the setting is skipped and logged rather than failing the pool start.

Two further settings appear in profiles:

  • FIFO batch (/sys/block/<dev>/queue/iosched/fifo_batch) tunes how many requests the deadline scheduler dispatches in one batch. It only exists while a deadline-style scheduler is selected, so it is skipped on any device running noop/none. None of the shipped profiles set it.
  • Target driver threads and tasklets set a floor on the SCST target driver's worker thread and tasklet counts, which is what serves iSCSI, FC and NVMe-oF. These are appliance-wide, not per pool: QuantaStor takes the highest value across every profile in use on the system, and can only raise the count above SCST's own start-up default (derived from the core count), never lower it. Both are capped at 256. Assigning the Virtualization profile to one pool therefore raises the thread floor for every target on that appliance.

QuantaStor also raises max_sectors_kb to 4096 on each member device where the hardware allows it, independently of the profile, so that large sequential I/O is not split into smaller commands.

The shipped profiles

Seven profiles ship with the appliance. Values below are read from /opt/osnexus/quantastor/conf/qs_io_profiles.conf on a 6.9.0 appliance.

Profile HDD read-ahead HDD queue depth HDD scheduler SSD read-ahead SSD queue depth SSD scheduler Target driver threads / tasklets
Default 256 KB 2 x hw queue depth deadline 4 KB 1 x hw queue depth noop SCST default
Disk Archive 512 KB 2 x hw queue depth deadline 4 KB 1 x hw queue depth noop SCST default
Edgeware IP TV Optimized 256 KB 64 deadline 4 KB 1 x hw queue depth noop SCST default
Edgeware Web TV Optimized 256 KB 64 deadline 4 KB 1 x hw queue depth noop SCST default
Media Post-Production (Ingest Optimized) 512 KB 3 x hw queue depth deadline 4 KB 1 x hw queue depth noop SCST default
Media Post-Production (Playback Optimized) 1024 KB 2 x hw queue depth deadline 4 KB 1 x hw queue depth noop SCST default
Virtualization 128 KB 64 deadline 4 KB 1 x hw queue depth noop 64 / 32

Where a multiplier is shown, the profile also carries a fixed fallback count -- 256 for every profile except Media Post-Production (Ingest Optimized), which uses 1024 -- for devices that do not report a hardware queue depth.

None of the shipped profiles set NVMe values. The mechanism supports them, but with nothing configured an NVMe pool member keeps the kernel's own settings, which on a current kernel already means no scheduler reordering and no meaningful read-ahead. Treat an all-NVMe pool as needing no profile tuning, and look at the Storage System Optimization queue depths instead if you need to change how much I/O is in flight.

Choosing a profile

The profile names describe intended workloads, so the choice usually follows from the application. What is worth understanding is why the numbers differ:

If the workload is... Use Because
General purpose, mixed, or you are not sure Default A middle read-ahead and a moderate queue depth. This is the right answer far more often than not, and we recommend leaving it alone unless you have measured a specific problem.
Virtual machines under ESXi, XenServer, Hyper-V or Virtuozzo Virtualization Many small concurrent random requests. It cuts read-ahead to 128 KB, because speculative reads on a random workload are wasted bandwidth, and holds the HDD queue at a fixed 64 to keep latency low rather than chasing throughput. It is also the only shipped profile that raises the target driver thread and tasklet floor, which is what a large number of concurrent iSCSI sessions needs.
Disk-to-disk backup, archive, large-file ingest Disk Archive Long sequential streams. Read-ahead goes to 512 KB and the queue depth stays deep, trading latency -- which nothing in a backup window cares about -- for throughput.
Media post-production playback Media Post-Production (Playback Optimized) Sustained sequential reads of very large files, where the highest read-ahead in the set (1 MB) keeps the stream ahead of the reader.
Media post-production editing and ingest Media Post-Production (Ingest Optimized) Concurrent large streams. It uses the deepest queue in the set (3 x hardware queue depth, falling back to 1024) with a 512 KB read-ahead.
The Edgeware IP TV or Web TV workflow Edgeware IP TV Optimized / Edgeware Web TV Optimized Tuning supplied for those specific applications. Do not pick them for anything else.

An all-flash pool is largely insensitive to the choice, because the SSD half of every shipped profile is identical -- the profiles differ only in their HDD values. On an all-flash pool the profile is effectively a no-op and the levers that matter are record size, compression and the system tunables.

Profiles are per pool, so a mixed appliance can run an archive profile on its backup pool and the virtualization profile on its VM pool. Read the values a profile applies with:

qs pool-profile-list
qs pool-profile-get --profile=Default

qs pool-profile-list and qs pool-profile-get report every value including the NVMe fields, and qs pool-modify --pool=<pool> --profile=<profile> assigns one.

Custom I/O profiles

Profiles are defined in /opt/osnexus/quantastor/conf/qs_io_profiles.conf, one section per profile. The section name is the profile's identifier and the name and description values are what the web interface shows, so make both unique. The simplest way to build a custom profile is to copy the section closest to your workload, rename it, and change one thing.

The keys are hdd_, ssd_ and nvme_ prefixed: read_ahead_kb, nr_requests, nr_requests_multiplier, scheduler and fifo_batch, plus the unprefixed min_target_driver_threads and min_target_driver_tasklets. A value of 0 means "leave this alone", and a profile with no name is skipped.

Restart the QuantaStor service after editing the file so the new profile is discovered and appears in the web interface:

systemctl restart quantastor

The file is replaced on upgrade. Keep a copy of any custom profile outside /opt/osnexus/quantastor/conf/ and re-apply it afterwards. A comment line must have # as its very first character; an indented # is not a comment.

If a profile does not appear to be taking effect, the service log records each application by name:

grep "Optimizing media for Storage Pool" /var/log/qs/qs_service.log

Disk optimization can also be switched off entirely by the presence of the file /etc/qs_disk_optimizations.disable, which support occasionally uses to isolate a problem. If that file exists, no profile is applied to any pool.

Note that chunk_size_kb still appears in the shipped file. It is a legacy Linux software RAID parameter, marked deprecated in the file itself, and has no effect on a scale-up pool.

Storage Volume I/O profiles

A separate profile mechanism applies to individual Storage Volumes, and it is the only place the iSCSI session parameters are exposed. Selecting a profile in Create Storage Volume sets the volume's block size and, once the volume is assigned to a host, the parameters on its iSCSI target. Three profiles ship:

Profile Block size Queued commands First burst length Optimizes for
Default 64 KB 64 128 KB General purpose backup workloads.
Desktop and Server Virtualization 64 KB 128 1 MB Virtual machine workloads. 64 KB is forward and backward compatible across VMware releases.
Databases 8 KB 128 64 KB Database workloads and small-block I/O patterns.

All three use a 1 MB maximum burst length and 1 MB maximum receive and transmit data segment lengths.

  • Queued commands is how many SCSI commands the target will accept outstanding on that session. Raising it helps a host that keeps many requests in flight, which is what a hypervisor with many virtual machines does; we do not recommend going above 256.
  • First burst length is how much write data an initiator may send unsolicited with the command itself, before the target replies asking for the rest. Setting it to the size of a typical write turns a two-round-trip write into one -- which is why the virtualization profile raises it to 1 MB, and why the database profile keeps it at 64 KB, matching a small-block write pattern instead of reserving buffers for data that never arrives.
  • Block size is the volume's on-disk record size and is fixed at creation. See Record size and volume block size.

The iSCSI parameters are applied only after the volume is assigned to a host or host group, because until then the volume has no iSCSI target to configure. Assign the volume, then confirm the profile took effect. A profile value of zero is skipped.

Read the shipped values with qs volume-profile-list and qs volume-profile-get --profile=<name>, and select one at create time with qs volume-create --volume-profile=<name>. The definitions live in /opt/osnexus/quantastor/conf/qs_volume_profiles.conf and, like the pool profiles, are replaced on upgrade.

RAM read cache (ARC)

Scale-up pools use the OpenZFS ARC in RAM as their primary read cache rather than the Linux page cache. It is the highest-leverage cache in the appliance: serving a block from RAM is orders of magnitude faster than reading it from media, and every read it absorbs is disk work that does not happen, which indirectly improves write performance too.

Sizing it

The ceiling is set by Cache Size (% of RAM) on the Cache Settings tab of Storage System Optimization, which defaults to 70% of system RAM and can be set between 30% and 90%. QuantaStor applies it by computing that percentage of total RAM and writing it to the OpenZFS zfs_arc_max parameter.

Two things follow from that:

  • The setting is re-applied when the QuantaStor service starts. A value set by hand -- with qs-util setzfsarcmax, or by writing zfs_arc_max directly -- is overwritten at the next service start. Use the dialog, or qs tunable-set --tunable=sst_cache_size:<percent>, so the change persists.
  • There is no dialog for the ARC minimum. The floor (zfs_arc_min) stops the kernel shrinking the cache below a given size under memory pressure. qs-util setzfsarcmin <percent> writes it to /etc/modprobe.d/zfs.conf, which means it only takes effect at the next boot. This is a support-level lever; raise it only if you have evidence that the ARC is being collapsed under pressure.

More useful than either number is the amount of RAM in the appliance, because 70% of too little RAM is still too little. Plan for a minimum of 32 GB to 64 GB on a small system, 96 GB to 128 GB on a medium one, and 256 GB or more on a large one; the OSNEXUS design tools will size it for a specific configuration.

What it costs

The ARC is not free memory that happens to be used for cache -- it is memory the appliance will not have for anything else:

  • Everything else on the appliance runs in the remaining 30%: the QuantaStor service, the target drivers, Samba and NFS, replication, and the kernel itself. Raising the percentage towards 90% on a system that also serves SMB, runs replication jobs or hosts containers is how a well-tuned appliance starts swapping.
  • Metadata competes with data inside the ARC. arc_meta_used in qs-iostat -a shows the split. A pool with tens of millions of small files can spend most of its cache on metadata, which is an argument for metadata offload devices rather than for more RAM.
  • An L2ARC consumes ARC. See below.
  • Deduplication consumes ARC, for the deduplication table, and it is not optional -- see Metadata, small-block and dedup offload.

Cache Compression (on by default) compresses ARC contents, which increases the effective cache size at a small CPU cost. Leave it on. Prefetch Disable (off by default) turns off the OpenZFS prefetcher, which is worth trying only on a workload that is almost purely random reads, where prefetched blocks are evicting useful ones. Both are on the Cache Settings tab of Storage System Optimization.

Cache policy

Cache Policy Primary on Modify Storage Pool -- and the equivalent setting on an individual Network Share or Storage Volume -- chooses what the ARC holds for that object: all, metadata only, or none.

This is the lever for a workload that pollutes the cache instead of benefiting from it. A large sequential backup ingest reads each block once, so caching those blocks gains nothing and evicts the working set of everything else on the appliance. Setting that share or volume to metadata keeps its directory structure cached -- which still helps, because metadata is read repeatedly -- while leaving its data blocks out of RAM. The Cache Hits chart described above is how you identify a candidate: a workload whose hits are almost all on the most-recently-used side is streaming, not reusing.

SSD read cache (L2ARC)

An L2ARC is a second-level read cache on SSD, below RAM and above the pool. Blocks evicted from the ARC land there, so a subsequent read is served from flash instead of from the data groups. Read cache devices are added from Add Log/Cache/Metadata Device(s), need no redundancy -- losing one costs you cache, not data -- and can be removed at any time. See Storage Pools for the dialog.

Whether it pays for itself depends entirely on the workload:

It helps when the working set is larger than RAM but not enormously larger, and blocks are re-read: a virtual machine estate whose common OS blocks are read by many guests, a database whose hot indexes exceed RAM, a file share with a recurring active set. Size it to roughly the application's working set.

It is wasted when:

  • The pool is already all-flash. A read cache in front of flash adds a layer without adding speed.
  • The workload streams. Data read once and never again -- backup ingest, single-pass media playback of a large library -- populates the cache with blocks nobody will ask for twice.
  • RAM was the cheaper answer. If the working set nearly fits in RAM, more RAM beats an L2ARC, because ARC hits are far faster and cost no ARC overhead.
  • It is oversized. Every block held in the L2ARC needs an index header in RAM, inside the ARC. An L2ARC sized far beyond the working set therefore takes RAM away from the cache that is faster than it -- and can leave you slower than with no read cache at all. l2_hdr_size in qs-iostat -a is that cost, in bytes.

Two measurements settle the argument. l2_size shows how much of the cache actually filled -- a cache that never fills is bigger than the workload needs. l2_hits against l2_misses shows whether it is being used at all; both near zero after several days of representative load means the cache is not earning its slot.

Judge an L2ARC after days, not hours. It has to learn which blocks are worth holding, and OpenZFS deliberately fills it slowly so that a burst of cold reads cannot flush it. A read cache benchmarked in the first hour will look useless whatever the workload.

Watch l2_cksum_bad as a health signal rather than a performance one: a rising count almost always means the SSD is failing and should be replaced.

Write log (SLOG/ZIL) and sync policy

The ZFS intent log (ZIL) is how OpenZFS keeps a promise. When an application issues a write and asks for it to be durable before the call returns, the data has to reach stable storage immediately, even though the pool would rather batch it into the next transaction group a few seconds later. The ZIL is that immediate record. By default it lives in the data groups; a write log (SLOG) device moves it onto dedicated flash so that a synchronous write costs one fast flash write rather than a scattered write into a parity stripe.

Sync Policy on Modify Storage Pool decides how much traffic goes that route:

Policy Behaviour When
standard (default) Honours the application's O_SYNC flag: writes that ask for durability are logged, everything else is batched. Almost always. It is a hybrid, and it gives hypervisors and databases the durability they ask for without penalising anything else.
always Every write goes through the log. Only when an application needs durability it does not ask for, and only with a log device fast enough to absorb the whole write stream.
disabled All writes are asynchronous. Acknowledged before they are durable. Not supported for production. It can lose acknowledged writes on an unexpected power failure, and the dialog raises a confirmation.

This is why a new write log often looks broken. Under the default standard policy, only writes that carry O_SYNC touch it -- hypervisors and databases set the flag; most other applications do not. Add a write log for a workload that does not, and it will sit idle: zil_commit_count in qs-iostat -a barely moves. That is the log correctly doing nothing, not a fault.

Setting Sync Policy to always forces every write through it, and this is where it goes wrong: if the log device cannot sustain the full write rate, forcing all writes through it makes the pool slower, not faster. Check zil_commit_stall_count and zil_commit_suspend_count -- anything other than zero under load means the log is the bottleneck.

Sizing and media choice:

  • Write log devices see constant, small, synchronous writes, which is the hardest duty cycle in the appliance. Use high-endurance enterprise SSDs; SLC-class devices are worth the money here in a way they are not for a read cache.
  • Log devices are mirrored automatically, and multiple pairs scale log throughput. See Storage Pools for the pairing rules and a worked sizing example.
  • Modern deployments need a write log less than they used to. Its value was in making an HDD pool usable for database and virtual machine workloads, and those workloads belong on flash now that NVMe and SAS SSDs are affordable. An all-flash pool rarely benefits from one.

The same sync policy choice is available per Network Share and per Storage Volume, which is the right granularity when one dataset on a pool needs always and the rest does not.

Metadata, small-block and dedup offload

On an HDD pool this is the largest single performance improvement available, and it is worth understanding why before deciding how much SSD to spend on it.

A metadata offload group (a ZFS special vdev) holds the pool's metadata on mirrored flash instead of scattered across the data groups. Metadata access is small and random by nature, which is the pattern HDDs are worst at, and it is on the critical path of far more than it looks:

  • Scrubs and resilvers walk the metadata tree. Moving it to flash shortens both substantially -- which matters most exactly when you can least afford it, during a rebuild with reduced redundancy.
  • Replication and snapshot operations are metadata-heavy, so a replication window that will not close is often a metadata problem rather than a bandwidth problem.
  • Directory traversal and file enumeration are metadata. This is what makes a large HDD file share feel slow to browse even when throughput is fine.
  • HA failover has to read pool metadata before the pool comes up.

Small Block Offload extends this to data. Set on Modify Storage Pool (and per share or volume), it routes any write below the chosen size to the metadata group instead of the data groups. The reason it helps so much on parity pools is arithmetic: writing a few kilobytes into a RAIDZ stripe means reading and rewriting parity across the whole stripe, so a small write costs far more than its size. Sending it to a mirrored SSD group instead avoids that entirely. The setting offers Disabled, 4K, 8K, 16K, 32K, 64K, 128K and 256K, and stays Disabled until a metadata group exists.

Practical guidance:

  • Always include SSDs for metadata and small-block offload on an HDD pool. The effect on scrub time, replication, failover and small-file performance is large enough that we do not design HDD pools without it. See Storage Pools for the recommended device count and starting offload size.
  • An all-flash pool does not need it and should not use it -- with one exception. QLC media handles small-block I/O poorly, so putting metadata and small blocks on TLC in front of a QLC pool is worthwhile.
  • Do not set the offload size larger than the SSD group can hold. The larger the threshold, the more of the pool's writes land on flash. Size the threshold to the capacity you gave it, not to the largest number in the list.
  • A metadata offload group cannot be removed. It is a permanent part of the pool, and losing it loses the pool, so mirror it properly the first time. There is no CLI command to remove one either.

Deduplication offload puts the deduplication table on its own mirrored vdev, and adding one turns deduplication on for the pool. Treat it as a capacity feature with a performance cost, not a performance feature. Every write has to be looked up in the deduplication table, that table has to be cached in RAM to make the lookup fast, and it grows with the amount of unique data in the pool -- so it competes with the ARC for exactly the memory your read workload wants. Like the metadata group, it cannot be removed once added. Use compression first; it gives most of the space saving on most data with none of this cost.

Record size and volume block size

Record size (for a Network Share) and Block Size (for a Storage Volume) are the same idea: the unit OpenZFS reads, writes, checksums and compresses. Matching it to the application's I/O size is one of the few settings that can change performance by a large factor, and it is one of the few that cannot be changed later for a volume.

Why the match matters:

  • Too large, for small random writes. Modifying part of a record forces a read of the whole record, the modification, recompression, and a write of the whole record. An 8 KB database write into a 1 MB record is a 1 MB read plus a 1 MB write. The same amplification applies to reads: a 4 KB read fetches the entire record.
  • Too small, for large sequential I/O. Every record carries checksum and indirect-block overhead, and each one is a separate compression unit, so small records cost more metadata, more IOPS, and less compression than large ones for the same amount of data.
Object Range offered Default Notes
Network Share -- Record Size Auto, 8K, 16K, 32K, 64K, 128K, 256K, 512K, 1M, 2M, 4M, 8M 128K (via Auto) Can be changed at any time from Modify Network Share. Existing data keeps the record size it was written with; only new writes use the new value.
Storage Volume -- Block Size 4K, 8K, 16K, 32K, 64K, 128K, 256K, 512K, 1M, 2M, 4M 64K Fixed at creation and cannot be changed afterwards. The field greys out on Modify Storage Volume.

Choosing:

  • Leave a general-purpose file share at the 128K default. It is a good compromise and shares rarely have one dominant I/O size.
  • Databases want a record or block size at or near the database page size -- 8K or 16K for most relational engines. This is what the Databases Storage Volume I/O profile sets, and it is the reason the profile exists.
  • Virtual machine datastores are well served by 64K, the volume default, which is what the Desktop and Server Virtualization profile selects and is compatible across VMware releases. A VM datastore carries a mix of guest I/O sizes, so a middle value beats optimising for either end.
  • Media, backup and archive shares want large records -- 1M and above. Fewer, larger records mean less metadata and better compression on large sequential files.
  • Match the share record size to the client's own block size where you know it, particularly for a single-purpose share.

Do not confuse either setting with Block Size Offset (ashift), which is the pool's physical sector alignment, is set at pool creation, and should be left on Auto. See Storage Pools.

Set these from the CLI with qs share-create --recordsize=<KB> and qs volume-create --blocksize-kb=<KB>.

Compression

Compression is on by default -- LZ4 -- and it is normally a throughput decision rather than a capacity one. A modern CPU compresses faster than the media can read or write, so compressing means moving fewer bytes through the disks, the HBAs and the backplane. On compressible data it makes the pool faster and larger at the same time.

That reasoning tells you when to change it:

  • Leave it on (LZ4) unless you have a reason. It is cheap enough that it costs nothing measurable on incompressible data, because OpenZFS detects a record that will not compress and stores it uncompressed.
  • Turn it off for data that is already compressed. Media and entertainment content, encrypted archives, and most image and video libraries gain nothing and only add CPU load. This is the common case in post-production.
  • Consider ZSTD when the pool is media-bound rather than CPU-bound. ZSTD compresses harder than LZ4 at more CPU cost. On an HDD pool with spare cores it can raise effective throughput; on an all-flash pool where the media is already faster than the CPU, it will reduce it. Measure both.
  • Do not reach for GZIP for performance. The gzip levels are a capacity choice for cold data, and they cost enough CPU to become the bottleneck on an active pool.
  • Changing compression affects new writes only. Existing blocks keep whatever algorithm they were written with, so a change takes effect gradually as data is rewritten.

The available algorithms and the field itself are on Modify Storage Pool -- see Storage Pools -- and the same choice is available per share and per volume, which is the right granularity when one dataset on a pool holds incompressible media and the rest does not.

Compression interacts with record size: a larger record gives the compressor more to work with and compresses better. It also interacts with partial writes, since a modified record has to be recompressed in full -- another reason not to pair a large record size with a small-write workload.

Device write caches

Worth knowing, because it explains a measurement that otherwise looks wrong. QuantaStor disables the volatile write-back cache on rotational disks as they are brought into service, putting them into write-through mode, because an HDD's on-board cache is not power-loss protected and acknowledging a write that is still only in that cache risks the pool's consistency.

Enterprise and datacenter SSDs and NVMe devices keep their write cache enabled, because they have the capacitors to flush it on power loss. Devices reached over iSCSI are put into write-through regardless.

So a raw HDD write benchmark on the appliance will fall short of the drive's data sheet, and that is deliberate. The correct way to get synchronous write performance back is a write log on flash, not a volatile cache.

Network-side factors

The network is frequently the real limit, and it is the layer where a small configuration mistake costs the most. Network Ports owns the port, VLAN and bonded-port dialogs; what follows is which choice to make and why.

MTU and jumbo frames

Navigation: Storage Management → Network Port → Modify (toolbar)

A larger MTU means more payload per frame and fewer frames, headers and interrupts for the same data -- a real gain on a 10 GbE or faster storage network carrying large block I/O. Modify Network Port has an MTU field with a Jumbo Frames button beside it that fills in 9000; once the port is on 9000 the button reads Default Frames and puts it back to 1500. The MTU field is only editable while the Static IP Configuration Settings section is active, and is disabled on VLAN and alias ports, which take the MTU of their parent.

Every device in the path must agree. The initiator, every switch port between, and the target all have to carry the same MTU. A device that receives a frame larger than its MTU drops it, and the symptom is not a clean failure -- it is a connection that works for small transfers and stalls on large ones, which is a genuinely unpleasant thing to diagnose. Set the switch first, verify end to end, and change the appliance last.

Set it from the CLI with qs network-port-modify --port=<port> --mtu=9000.

Multiple interfaces versus bonding

There are two ways to use several ports for one workload, and they are not interchangeable.

Multipath (MPIO), one subnet per port. Each port gets a static address on its own subnet, the initiator opens a session to each, and multipathing on the host spreads I/O across them and survives the loss of any one. This is the approach we recommend trying first for block storage, because it needs nothing from the switch, it scales linearly with ports, and its failure modes are visible.

Putting several ports on the same subnet is the mistake to avoid here. With multiple interfaces on one subnet, Linux by default will answer an ARP request for any local address out of any interface, so return traffic can leave a port other than the one the request arrived on. The paths you carefully separated then collapse onto whichever port answered, and throughput sits at roughly one port's worth however many you configured. QuantaStor's ARP Filtering setting on Modify Storage System addresses this -- Auto (the default) enables ARP filtering only when a bonded port is present, so a system with several unbonded ports on one subnet does not get it. If your design puts multiple ports on the same subnet, set ARP Policy explicitly to Enabled -- the CLI equivalent is qs system-modify --arp-filter-mode=enabled. Better still, give each port its own subnet and avoid the question.

For Windows initiators, one detail is worth recording because it is easy to get wrong and produces a confusing result. QuantaStor identifies itself over SCSI as vendor OSNEXUS, product QUANTASTOR, so the string to add under MPIO Devices is OSNEXUS QUANTASTOR followed by six trailing spaces -- an eight-character vendor field and a sixteen-character product field, both space padded. Without exactly the right padding the MPIO driver does not recognise the devices and Windows Disk Management shows the same disk once per path instead of one disk with several paths.

Bonding. A bonded port presents several physical ports as one logical interface with one address.

Navigation: Storage Management → Network Port → Create Bonded Port (toolbar)

Each Bond Mode in the dialog names the switch it needs, which is the part to read first:

Bond Mode Switch required
Link Aggr Ctrl Protocol (LACP layer2) Managed switch. The dialog's default.
Link Aggr Ctrl Protocol (LACP layer2+3) Managed switch. Hashes on MAC and IP, which spreads better than layer2 alone across many peers.
Link Aggr Ctrl Protocol (LACP layer3+4) Managed switch. Hashes on IP and port, so two sessions between the same pair of hosts can land on different members.
Round Robin (balance-rr) Etherchannel managed switch.
Balance XOR (balance-xor) Etherchannel managed switch.
Active-Backup (active-backup) Unmanaged switch. Failover only, with no throughput gain -- one member carries all traffic.
Adaptive Transmit Load Balancing (balance-tlb) Unmanaged switch. Balances outbound traffic only.
Adaptive Load Balancing (balance-alb) Unmanaged switch. Balances both directions without switch support.

Bonding is the right choice when you need one address -- for file protocols, or where the client cannot do multipathing -- and when the ports are spread across switches for redundancy, which requires LACP and switch infrastructure that supports it. Two cautions:

  • A bond does not make one session faster. The load-balancing modes hash each flow onto one member port, so a single TCP connection gets one port's bandwidth no matter how many are in the bond. Aggregate throughput across many clients improves; one client's single stream does not. With iSCSI over a bond there is one session, so multipathing sees one path.
  • LACP performance depends heavily on the switch's hash policy. Enabling LACP and leaving the switch on its defaults has been measured to cost the large majority of the available throughput, recovered only after matching the switch's hashing to the traffic. This varies between switch vendors, so benchmark multipath first and use that number as the target LACP has to meet -- without a target you have no way to tell a badly hashed bond from a fast one.

A Network Bonding Policy setting on Modify Storage System carries a system-level bonding mode drawn from a subset of the same list; see Storage System.

Where you use both approaches, bond groups of ports and run multipath across the bonds.

NIC ring buffers

Optimize hardware RX/TX buffer settings for throughput on Modify Network Port raises the network card's receive and transmit ring buffers. Larger rings give the driver more room to absorb a burst before dropping packets, which matters on a fast link carrying large block I/O or replication traffic. QuantaStor picks the largest power-of-two value that stays safely below the card's hardware maximum, records the original values so that clearing the checkbox restores them, and re-checks the setting periodically.

The checkbox is unavailable on a bonded port and on virtual ports -- tune the member ports instead -- and on a port whose configuration type is disabled. The CLI equivalent is qs network-port-modify --port=<port> --auto-tune=true.

Two related system tunables sit on the Network Settings tab of Storage System Optimization: Network TX/RX Queue Length (the software queue, default 5000, and worth raising on 10 GbE and faster) and Network Device Max Backlog (how many packets may queue on the receive side when the interface delivers faster than the kernel can process).

What does not apply

Three things that look like tuning levers and are not:

  • Hardware RAID card settings. Scale-up pools are deployed on HBAs, not on hardware RAID controllers. OpenZFS needs direct access to the drives to checksum, self-heal and manage its own redundancy, and a RAID card's cache, stripe size and read-ahead settings sit between it and the media doing none of those things. There is no RAID-card tuning to do on a scale-up pool, because there should be no RAID card. Pick the layout in Storage Pools instead.
  • Linux software RAID (mdadm) parameters. The [mdadm] section of /etc/quantastor.conf and the chunk_size_kb key in the I/O profiles file are legacy settings from pool types the product no longer builds. They have no effect on a scale-up pool.
  • The [device] section of /etc/quantastor.conf. Read-ahead, queue depth and scheduler were once configured there. They are not any more -- those settings come from the pool's I/O profile, and the keys in that section are inert. Edit the profile, not the configuration file.

Two further notes on scope. Most of the system tunables in Storage System Optimization are OpenZFS module parameters, so they affect scale-up pools only and do nothing for scale-out (Ceph) pools. And tunables are stored per Storage System, so in a grid each member is tuned independently -- save a profile and apply it to the others rather than editing each by hand.

Related pages


Verified against QuantaStor 6.9.0.