See also [[#Optimization Profiles|Optimization Profiles]] and
See also [[#Optimization Profiles|Optimization Profiles]] and
Line 79:
Line 79:
== Settings Reference ==
== Settings Reference ==
The dialog groups the tunables into six tabs. The kernel parameter each setting writes is given in parentheses in its description, or listed under each tab.
[[File:ssopt_general.png|thumb|right|650px|Storage System Optimization, General Settings tab.]]
=== General Settings ===
Every setting the dialog offers is documented on '''[[qs_systemtunables.conf]]''', the configuration file the tunable set is loaded from: one table per tab, giving each setting's identifier, its accepted range, its shipped default, its help text, and the kernel parameter it writes. That page also covers the file's format, the two retired tunables, and the override path that keeps a local change to the tunable set across an upgrade.
[[File:ssopt_general.png|thumb|right|666px|Storage System Optimization, General Settings tab.]]
{| class="wikitable"
! Setting !! Range !! Default !! What it does
|-
| '''Resilver Priority (msec/TXG)''' || 500 – 8000 || 3000 || Higher settings indicates more time should be dedicated to resilvering between transaction groups (zfs_resilver_min_time_ms).
|-
| '''Remote Replication Prefetch (MB)''' || 4 – 128 MB || 50 MB || Improves performance of replication send operations by prefetching data from disk (units in MB). Reducing this may reduce latency for systems replicating during production hours. (zfs_pd_bytes_max)
|-
| '''Cleanup Priority (% of TXG)''' || 5 – 90 || 30 || Controls the maximum amount of dirty blocks to be freed as a percentage of each transaction group. Increase to give high priority to cleanup operations like file deletion.
|-
| '''Limit Active Async Writes''' || 5 – 50 || 30 || Limit active asynchronous writes based on dirty percent. This can help with write latency and resilver performance when set to a value between 10 and 30.
|-
| '''Transaction Group (TXG) Commit Timeout''' || 5 – 90 || 10 || Percentage of dirty data at which the transaction group commit timeout is reduced to 1 second to speed up commits and reduce latency. Default is 10% which means that once the dirty data reaches 10% of RAM the system will start aggressively committing transaction groups to keep latency low.
|-
| '''Pool Metaslab Shift''' || 28 – 38 || 34 || Controls the size of metaslabs used for space allocation. Larger metaslabs can improve performance on large pools but may lead to fragmentation and reduced performance over time if set too high.
|-
| '''Pool Metaslab Count Limit''' || 1024 – 262144 || 131072 || Limits the number of metaslabs that can be allocated for a single IO operation. Increasing this can improve performance for large IO operations on large pools.
|-
| '''Pool Metaslab LBA Weighting''' || on / off || on || Enabling LBA weighting allows the allocator to prefer lower LBAs for better performance. This can be beneficial for some workloads but may lead to increased fragmentation over time.
|-
| '''Metaslab Aliquot''' || 64 – 32768 KB || 1024 KB || Controls the size of metaslab aliquots used for space allocation. Larger aliquots can improve performance on large pools but may lead to fragmentation and reduced performance over time if set too high.
|-
| '''Read-Ahead Distance Max (MB)''' || 8 – 512 MB || 64 MB || Maximum distance in bytes that the ZFS prefetcher will read ahead. Increasing this can improve performance for sequential read workloads.
|-
| '''Read-Ahead Distance Min (MB)''' || 0 – 128 MB || 4 MB || Minimum distance in bytes that the ZFS prefetcher will read ahead. Reducing this can improve performance for workloads with smaller sequential read patterns.
[[File:ssopt_cache.png|thumb|right|666px|Storage System Optimization, Cache Settings tab.]]
{| class="wikitable"
! Setting !! Range !! Default !! What it does
|-
| '''Cache Size (% of RAM)''' || 30 – 90 || 70 || Sets the percentage of the system's RAM available as in-memory read cache for use by all Storage Pools.
|-
| '''Write Throttle Limit (% of RAM)''' || 3 – 20 || 10 || Percentage of system RAM allocated for dirty data which after exeeded halts further writes until enough has flushed to free up more space in RAM. (controls zfs_dirty_data_max).
|-
| '''Write Flush Threshold (% of Write Queue RAM)''' || 5 – 80 || 20 || Percentage of write queue filled at which transaction group syncing is ensured. Reducing this can improve IOPS.(zfs_dirty_data_sync_percent)
|-
| '''Write Flush Threshold (sec)''' || 1 – 10 || 5 || Once the threshold has been reached a new transaction group will be started. Reducing this can improve IOPS.(zfs_txg_timeout)
|-
| '''Async Write - Min Active I/Os''' || 1 – 5 || 2 || Minimum asynchronous write I/Os active to each VDEV (zfs_vdev_async_write_min_active). Lower values generally improve latency on rotational media and hurt resilver performance.
|-
| '''Cache Compression''' || on / off || on || Compresses ARC data to increase the effective size of the in-memory read cache for Storage Pools.(zfs_compressed_arc_enabled)
|-
| '''Prefetch Disable''' || on / off || off || In cases where the IO patterns is predominantly random reads performance may be improved by disabling prefetch.(zfs_prefetch_disable)
[[File:ssopt_pool.png|thumb|right|666px|Storage System Optimization, Pool Settings tab.]]
{| class="wikitable"
! Setting !! Range !! Default !! What it does
|-
| '''Aggregate Per-Device Active I/O Limit''' || 256 – 4096 || 2000 || Maximum number of IO operations active per VDEV. Acts as a global cap of the sum of all sync/async read/write IOs and scrub/resilver classes of IO.
|-
| '''Read Queue Depth (sync I/O)''' || 4 – 128 || 32 || Maximum number of synchronous read IO operations active per VDEV. (zfs_vdev_sync_read_max_active)
|-
| '''Write Queue Depth (sync I/O)''' || 4 – 128 || 32 || Maximum number of synchronous direct write IO operations active per VDEV. (zfs_vdev_sync_write_max_active)
|-
| '''Read Queue Depth (async I/O)''' || 4 – 128 || 32 || Maximum number of asynchronous read IO operations active per VDEV. (zfs_vdev_async_read_max_active)
|-
| '''Write Queue Depth (async I/O)''' || 2 – 128 || 32 || Maximum asynchronous write I/Os active to each VDEV (zfs_vdev_async_write_max_active). Reducing max-active can improve latency under contention, but it can also lower throughput.
|-
| '''I/O Aggregation Limit (KB)''' || 128 – 32768 KB || 128 KB || Sets the VDEV upper bound on I/O coalescence/aggregation for a stripe of data. Increasing this value to 1M or larger may increase throughput for sequential I/O workloads. (zfs_vdev_aggregation_limit)
|-
| '''Scrub Queue Depth - Max Active I/Os''' || 1 – 32 || 3 || Maximum scrub or scan read I/Os active to each VDEV (zfs_vdev_scrub_max_active). Increasing this can speed up scrub and scan operations, but may increase impact on client workloads.
|-
| '''Scrub Queue Depth - Min Active I/Os''' || 1 – 8 || 1 || Minimum scrub or scan read I/Os active to each VDEV (zfs_vdev_scrub_min_active). Lower values reduce impact on production I/O, while higher values can improve scrub progress when the pool is busy.
|-
| '''I/O Allocator Default Queue Depth''' || 2 – 1024 || 32 || Default queue depth for each VDEV IO allocator. Higher values allow for better coalescing of sequential writes before sending them to the disk, but can increase transaction commit times. (zfs_vdev_def_queue_depth)
[[File:ssopt_network.png|thumb|right|666px|Storage System Optimization, Network Settings tab.]]
{| class="wikitable"
! Setting !! Range !! Default !! What it does
|-
| '''Network TX/RX Queue Length''' || 1000 – 20000 || 5000 || Sets the transmit and receive network queue length for high-speed network ports (10GbE and faster). Values like 5000 and higher are recommended for systems with 10GbE and faster networking. (txqueuelen)
|-
| '''Network Device Max Backlog''' || 128 – 1000000 || 1000 || Maximum number of packets allowed to queue on the input side when the interface receives data faster than the kernel can process it. (netdev_max_backlog)
|}
Kernel parameters set by this tab:
<pre style="font-size: smaller">
ip.link.txqueuelen
/proc/sys/net/core/netdev_max_backlog
</pre>
=== Volume Settings ===
[[File:ssopt_volume.png|thumb|right|666px|Storage System Optimization, Volume Settings tab.]]
{| class="wikitable"
! Setting !! Range !! Default !! What it does
|-
| '''Storage Volume sync IO mode''' || on / off || on || Storage Volume (ZVOL) sync mode can be ON(1) or OFF(0), default is ON.
|}
Kernel parameters set by this tab:
<pre style="font-size: smaller">
/sys/module/zfs/parameters/zvol_request_sync
</pre>
=== Driver Settings ===
[[File:ssopt_driver.png|thumb|right|666px|Storage System Optimization, Driver Settings tab.]]
{| class="wikitable"
! Setting !! Range !! Default !! What it does
|-
| '''KMEM Cache Threads''' || 1 – 16 || 4 || Number of threads used for KMEM cache operations. Increasing this can improve performance for workloads with high levels of concurrent allocations and frees.
|-
| '''KMEM Cache Objects Per Slab''' || 8 – 128 || 8 || Number of objects per slab in the KMEM cache. Increasing this can improve performance for workloads with many small allocations.
|-
| '''KMEM Cache Max Size (MB)''' || 32 – 2048 MB || 32 MB || Maximum size of the KMEM cache in megabytes. Increasing this can improve performance for workloads with high memory usage.
Two tunables have been retired and no longer appear in the dialog:
* {{Code|1=sst_resilver_prio}} -- superseded by '''Resilver Priority (msec/TXG)'''
* {{Code|1=sst_write_flush_rate_mb}} -- superseded by '''Write Flush Threshold (% of Write Queue RAM)'''
They are still present in the configuration file marked {{Code|1=param=deprecated}} so that older saved profiles referencing them load without error. Do not use them in new profiles.
== Extending and customizing the configuration files ==
== Extending and customizing the configuration files ==
Line 287:
Line 103:
=== qs_systemtunables.conf -- defining what appears in the dialog ===
=== qs_systemtunables.conf -- defining what appears in the dialog ===
Each tunable is one section. The section name is the tunable's identifier and '''must''' begin with {{Code|1=sst_}}. A complete example:
Adding a tunable, retiring one, or changing a range is done in {{Code|1=qs_systemtunables.conf}}. See '''[[qs_systemtunables.conf]]''' for its keys, its validation rules, and the override location that survives an upgrade. A tunable added there automatically becomes part of the built-in '''Default''' profile, because that profile is regenerated from every tunable's {{Code|1=default}} value on each service start.
<pre style="font-size: smaller">
[sst_cache_size]
tab=cache_settings
param=/sys/module/zfs/parameters/zfs_arc_max
title="Cache Size (% of RAM)"
desc="Sets the maximum size of the ARC read cache as a percentage of system RAM."
type=percentage
min=30
max=90
default=70
</pre>
The keys are:
{| class="wikitable"
! Key !! Purpose
|-
| {{Code|1=tab}} || Which tab the setting appears under. One of {{Code|1=general_settings}}, {{Code|1=cache_settings}}, {{Code|1=pool_settings}}, {{Code|1=network_settings}}, {{Code|1=volume_settings}}, {{Code|1=driver_settings}}. A tab name that is not one of these will not render.
|-
| {{Code|1=param}} || The kernel parameter to write, given as a full path such as {{Code|1=/sys/module/zfs/parameters/zfs_arc_max}}. Set {{Code|1=param=deprecated}} to retire a tunable while keeping old profiles loadable.
|-
| {{Code|1=title}} || The label shown in the dialog. Quote it.
|-
| {{Code|1=desc}} || The help text shown for the setting. Quote it. Naming the underlying kernel parameter in the text is the existing convention and worth following.
|-
| {{Code|1=type}} || {{Code|1=range}} for a numeric slider, {{Code|1=percentage}} for a slider expressed as a percentage, {{Code|1=boolean}} for a checkbox.
|-
| {{Code|1=min}} / {{Code|1=max}} || The permitted range. The dialog will not offer values outside it, and a profile value outside it is rejected -- see the warning below.
|-
| {{Code|1=default}} || The shipped default, and the value the '''Default''' profile restores.
|-
| {{Code|1=param_units}} / {{Code|1=display_units}} || Used when the kernel parameter's unit differs from the unit you want shown. For example {{Code|1=param_units=B}} with {{Code|1=display_units=MB}} lets an administrator work in megabytes while the kernel parameter is written in bytes.
|}
Because the built-in '''Default''' profile is regenerated from these {{Code|1=default}} values on every service start, a tunable you add here automatically becomes part of Default without any further work.
The Storage System Optimization dialog allows one to adjust system tunings (tunables) that control cache behavior and Storage Pool I/O, so performance can be more closely matched to a specific hardware configuration and workload.
Navigation: Storage Management → Storage Systems → right-click a Storage System → Storage System Optimization...
Note this dialog is reached from the right-click context menu on a Storage System, not from the toolbar. Note also that most of the tuning settings (all those based on zfs params) only effect Scale-up (ZFS-based) pools and not Scale-out (Ceph-based) pools.
Every setting is applied live to the running system and most settings do not require a reboot (driver changes to the SPL excepted). All changes are persisted so they survive reboots and upgrades. Values are per Storage System, so in a grid each system is tuned independently. The Storage System selector at the top of the dialog switches which system you are editing. To apply a common configuration to each system save your custom configuration as a Optimization Profile first, then apply it to other systems.
The six tabs group the tunables by what they affect:
The shipped defaults are appropriate for the large majority of deployments. The tunables here are the ones that matter for specific, identifiable problems (resilver taking too long, write latency spikes under load, a replication window overrunning). Changing them speculatively is more likely to reduce performance and create problems rather than help so we recommend any changes here be done with the assistance of the support team.
Recommended approach:
Establish a baseline first. Use the Performance and Cache Stats (ARC) views on the Storage System dashboard to see what the system is actually doing before you change anything.
Change one group of related settings at a time, then measure again under a representative workload.
Save a named profile before you start experimenting, so you can get back to a known state -- see Optimization Profiles.
Revert All returns the dialog to the values currently stored in the database, which is the quickest way out of a half-finished experiment. Note it reverts the workspace -- you still need OK or Apply to make that stick.
If you are unsure whether a setting applies to your situation, contact OSNEXUS support rather than guessing; several of these interact.
Optimization Profiles
A profile is a named collection of tunable values that can be applied to a Storage System in one step. Profiles make it practical to keep a tuning recipe consistent across a fleet, and to get back to a known configuration.
The Optimization Profiles controls at the top of the dialog are:
Apply Profile -- copies the selected profile's values into the dialog. Nothing is committed until you press OK or Apply, so you can review what a profile will change before accepting it.
Save Profile... -- saves the values currently in the dialog as a new named profile.
Delete Profile -- deletes the selected profile. This only removes the profile; it does not change the settings of any system that had it applied.
Revert All -- resets the dialog back to the values stored in the database for this system.
Two profiles ship with QuantaStor:
Default -- every tunable at its shipped default. Applying it resets the whole system optimization configuration. This profile is regenerated from the current defaults each time the QuantaStor service starts, so it always reflects the running release, including tunables added in a newer version.
Seagate Exos Protect Optimized -- tuning for large scale-up deployments layered over Seagate ADAPT hardware RAID.
Individual tunables can be read and set directly. Note tunable-set
takes the value as part of --tunable in key:value form
-- there is no separate --value argument -- and accepts a comma
separated list to set several at once:
tunable-set also takes --tunable-option=reset to
return tunables to their defaults, and --storage-system to target a
system other than the local one.
Settings Reference
Storage System Optimization, General Settings tab.
Every setting the dialog offers is documented on qs_systemtunables.conf, the configuration file the tunable set is loaded from: one table per tab, giving each setting's identifier, its accepted range, its shipped default, its help text, and the kernel parameter it writes. That page also covers the file's format, the two retired tunables, and the override path that keeps a local change to the tunable set across an upgrade.
Extending and customizing the configuration files
The tunable set and the shipped profiles are both driven by plain text configuration files on each Storage System, so a site can add tunables that QuantaStor does not expose out of the box and define its own profiles.
Both files live in:
/opt/osnexus/quantastor/conf/
These files are replaced on upgrade. Keep a copy of any local additions outside that directory and re-apply them after upgrading, or the changes will be lost. In both files a line is treated as a comment only when # is the first character of the line -- an indented # is not a comment.
Changes take effect when the QuantaStor service restarts:
systemctl restart quantastor
qs_systemtunables.conf -- defining what appears in the dialog
Adding a tunable, retiring one, or changing a range is done in qs_systemtunables.conf. See qs_systemtunables.conf for its keys, its validation rules, and the override location that survives an upgrade. A tunable added there automatically becomes part of the built-in Default profile, because that profile is regenerated from every tunable's default value on each service start.
qs_systemprofiles.conf -- defining profiles
Each profile is one section, with a display name, a description, and a list of tunable values:
[my_site_nfs_tuning]
name=Site NFS Tuning
description=Tuning for NFS-heavy workloads on our 60-bay shelves.
tunables="zfs_arc_max:80,zfs_txg_timeout:5,spl_kmem_cache_kmem_threads:8"
The tunables value is a comma separated list of key:value pairs. Four key forms are accepted:
the tunable's sst_ section name, for example sst_cache_size
a bare ZFS parameter name -- anything beginning zfs_, zfetch_, metaslab_, zio_ or vdev_ -- resolved under /sys/module/zfs/parameters/
a bare SPL parameter name beginning spl_, resolved under /sys/module/spl/parameters/
a full parameter path beginning /
A value carrying a B suffix is treated as a byte count and converted into the tunable's display units, which is why the shipped Seagate profile can write zfs_vdev_aggregation_limit:33554432B for a setting the dialog presents in KB.
A value outside the tunable's min/max is skipped, not clamped. The profile still applies, but that one setting is silently left alone apart from a warning in the service log:
Skipping tunable '<name>' for profile '<profile>', value '<n>' is outside of range (<min> : <max>)
If a profile does not appear to take full effect, check /var/log/qs/qs_service.log for that message before assuming the tunable itself is broken.
Two further rules are worth knowing:
Do not add a tunables line to a profile marked populate_defaults=true. That flag marks the built-in Default profile, whose tunable set is generated automatically from every tunable's default value on each service start.
A profile you have modified is not overwritten by the file. If a profile of the same name already exists and has been edited, the definition in the configuration file is not re-applied over it. The Default profile is the exception and is always refreshed.
The easiest way to author a custom profile is to set the values you want in the dialog, press Save Profile..., and then read the result back with qs tunable-profile-get. That gives you a known-good set of values to copy into the configuration file for deployment across a fleet.
Related pages
Storage System -- the Storage System Modify dialog and the rest of the system-level configuration
Storage Pools -- pool-level tuning, which is separate from these system-wide tunables