Performance Monitoring: Difference between revisions
m Rewrite from the product on 6.9.0: full Grid View / Storage System / per-object dashboard reference, statistics database sampling interval and retention, external dashboards, performance alerts. Corrects the Graph/Tiles checkbox names (now Grid Details / System Details) and the claim that the Tiles checkbox shows only tiles; replaces the stale .jpg screenshots (QSTOR-12352) |
|||
| Line 1: | Line 1: | ||
[[Category:admin_guide]] | [[Category:admin_guide]] | ||
This page covers the performance and health information QuantaStor collects about itself, where each view lives in the web interface, and how to read the numbers it shows. It also documents the statistics database behind the charts -- how often it samples, how long it keeps each resolution, and what makes a chart look wrong. | |||
For the diagnostic method -- how to prove where a bottleneck is -- see [[Performance Testing]]. For the settings that change performance, see [[Performance Tuning]]. | |||
{| class="wikitable" | |||
! Section !! Purpose | |||
|- | |||
| [[#The Grid View dashboard|The Grid View dashboard]] || Grid-wide health, capacity and per-appliance sparklines. | |||
|- | |||
| [[#The Storage System dashboard|The Storage System dashboard]] || CPU, network and memory; ARC cache statistics; thermal and power sensors. | |||
|- | |||
| [[#Per-object performance views|Per-object performance views]] || Pool, volume, share and physical disk charts. | |||
|- | |||
| [[#Time range, refresh and zoom|Time range, refresh and zoom]] || What each range reads, and what updates on its own. | |||
|- | |||
| [[#The statistics database|The statistics database]] || Sampling interval, retention, downsampling, and what is collected. | |||
|- | |||
| [[#qs-iostat: the console view|qs-iostat: the console view]] || Reading the same counters from the command line. | |||
|- | |||
| [[#External dashboards and Grafana|External dashboards and Grafana]] || Embedding your own monitoring pages. | |||
|- | |||
| [[#Performance-related alerts|Performance-related alerts]] || Which conditions raise an alert, and which do not. | |||
|- | |||
| [[#When a chart is empty or looks wrong|When a chart is empty or looks wrong]] || The common causes, in order. | |||
|} | |||
== The Grid View dashboard == | |||
{{Navigation|Grid View ''(main tab)''}} | |||
[[File:Grid | [[File:perfmon_grid_view.png|thumb|right|800px|The Grid View dashboard: five grid-wide counters, the two pool capacity charts, and one tile per appliance.]] | ||
'''Grid View''' is the landing view for a whole grid. It has three regions, and two checkboxes on its toolbar control which of them are shown: | |||
* '''Grid Details''' (on by default) shows the five grid-wide counters and the two pool capacity charts. Clear it and only the per-appliance tiles remain. | |||
* '''System Details''' (off by default) expands every appliance tile with capacity, session and health detail. | |||
[[ | '''+Add System''' and '''-Remove System''' open the grid membership dialogs; those belong to [[Storage System]] rather than to monitoring. | ||
=== The grid counters === | |||
The five counters across the top are current state, computed from the objects the web interface already holds -- they are not time series, so they do not depend on the statistics database. | |||
{| class="wikitable" | |||
! Counter !! Reads | |||
|- | |||
| Systems Online || Appliances in a normal state, over the total number of grid members. A member counts as online only if it is not in an error, missing, warning, offline or initializing state, so a healthy-but-warning appliance shows as ''not'' online. "2/3" on three running appliances usually means one is in a warning state, not that one is down. | |||
|- | |||
| Systems with Alerts || Appliances that have at least one alert that has not been dismissed, over the total. | |||
|- | |||
| Pools Online || Storage pools in a normal state, over the total. Internal pools are excluded from every pool count and chart on this dashboard. | |||
|- | |||
| Pools with Alerts || Pools with at least one undismissed alert, over the total. | |||
|- | |||
| Alerts || Total undismissed alerts across the grid. | |||
|} | |||
[[ | A counter's footer turns amber when the highest severity in its scope is a warning and red when it is an error; '''View Details''' opens the matching list or the alert view. '''These count alert records, not live conditions''', so an alert raised for a problem that has since been fixed still counts until someone dismisses it -- use '''Dismiss Alerts''' on the Alerts pane. See [[Alert Manager]] for what raises them and [[Call-home / Alerting]] for where they are delivered. | ||
== | === Pool Capacity Threshold Alerts === | ||
This chart counts pools into '''Ok''', '''Warning''' and '''Alert''' by how full they are, against the pool free-space thresholds configured in Alert Manager. The shipped defaults are a warning at 30% free (70% utilized) and an alert at 10% free (90% utilized), so a pool is counted: | |||
* '''Ok''' below 70% utilized, | |||
* '''Warning''' between 70% and 90% utilized, | |||
* '''Alert''' above 90% utilized. | |||
Read the thresholds actually in force with <code>[[QuantaStor CLI Command Reference#alert-config-get|qs alert-config-get]]</code> and change them with <code>[[QuantaStor CLI Command Reference#alert-config-set|qs alert-config-set]]</code> or in [[Alert Manager]], which owns them. Alert Manager also carries a third, critical threshold (5% free by default) that this chart does not break out. | |||
=== Total Pool Capacity === | |||
The second chart breaks the grid's combined pool capacity into '''Free''', '''Volumes''', '''Volume Snapshots''', '''Shares''', '''Share Snapshots''' and '''Other'''. '''Other''' is the remainder -- total capacity less free space less the four accounted categories -- so it absorbs pool metadata and anything on the pool that is not represented as a volume or a share. A few MB in Other on an otherwise empty pool is normal. | |||
{{ | === Appliance tiles === | ||
[[File:perfmon_system_tile.png|thumb|right|508px|One appliance tile with System Details enabled. CPU and memory are the last minute; the temperature and power charts are the last hour.]] | |||
Each grid member gets a tile. The header carries the appliance name, its management address and an '''Alerts [''n'']''' button whose colour follows the highest severity present; the button opens that appliance's alerts. Below it: | |||
* The chassis or platform string and, for hardware with IPMI, the '''POWER''' state and power supply icons. | |||
* A thermometer showing the average CPU temperature. It renders green between 15 and 32 C and amber outside that band. | |||
* '''CPU USAGE''' and '''MEMORY USAGE''' sparklines covering the last minute, refreshed every five seconds. | |||
* '''CPU Temp (Last Hour)''' and '''SYSTEM POWER (Last Hour)''' charts. These two are always a one-hour window regardless of anything else on the page, because their samples arrive only every five minutes. | |||
* A footer with the appliance's object state, and an arrow that selects that appliance in [[Storage System|Storage Systems]]. | |||
'''On a virtual appliance the temperature and power panels are empty and the thermometer reads 0 C.''' That is expected, not a fault: QuantaStor skips the IPMI power reading entirely on a virtual machine, and it writes a temperature or power sample only when the value it read is non-zero. The same applies to hardware with no IPMI/DCMI support. | |||
With '''System Details''' enabled each tile also shows: | |||
{| class="wikitable" | |||
! Panel !! Contents | |||
|- | |||
| Total / Used Pool Capacity || A utilization bar for that appliance's pools, with free space, total and used. | |||
|- | |||
| Connections || Current iSCSI, Fibre Channel and SMB session counts. Each is a link into the matching sessions view. | |||
|- | |||
| Pool Capacity Health || That appliance's pools counted Ok / Warning / Alert against the same thresholds as the grid chart above. | |||
|- | |||
| Ceph OSD Health || OSDs counted Ok / Warning / Alert, and zero throughout on an appliance that is not a Ceph cluster member. | |||
|} | |||
A badge is grey when its count is zero, so grey means "none of these", not "unknown". | |||
=== How often Grid View updates === | |||
Grid View redraws when grid objects change, coalesced to at most once every 15 seconds, and it redraws unconditionally at least once a minute so the sparklines keep moving even on an idle grid. The per-appliance sparklines fetch their own data every five seconds, independently of that. | |||
== The Storage System dashboard == | |||
{{Navigation|Storage Management → Storage Systems → ''select a Storage System'' → Performance ''(dashboard toolbar)''}} | |||
[[File:perfmon_system_perf.png|thumb|right|800px|The Performance view. The line near 100% is Idle, not utilization.]] | |||
Selecting a grid member under '''Storage Systems''' puts a dashboard panel above the object grid, titled '''Storage System Dashboard (''name'')'''. Selecting the grid root instead of a member shows no dashboard. The panel's toolbar carries three views -- '''Performance''', '''Cache Stats (ARC)''' and '''Sensors''' -- and a time range selector. The double-chevron icon at the far right collapses the whole dashboard panel; it is not a refresh control. | |||
=== Performance === | |||
Three charts, all for the selected appliance: | |||
* '''CPU''' plots '''Idle''', '''System''' and '''Wait-IO''' as percentages. Read it carefully: the line sitting near 100% is '''Idle'''. High Wait-IO alongside high Idle means the appliance is waiting on storage, not short of CPU -- and no amount of CPU or thread tuning will change that. | |||
* '''Network''' plots bytes per second per network port, as '''TX:''<port>''''' and '''RX:''<port>'''''. Because it is per port rather than aggregated, it is also the quickest way to see that traffic you expected to be spread over several ports has collapsed onto one. | |||
* '''Memory''' plots memory in use as a percentage. On a scale-up appliance most of it is usually the ZFS ARC; see [[Performance Tuning]] for what that costs and how to bound it. | |||
=== Cache Stats (ARC) === | |||
{{Navigation|Storage Management → Storage Systems → ''select a Storage System'' → Cache Stats (ARC) ''(dashboard toolbar)''}} | |||
[[File:perfmon_cache_arc.png|thumb|right|800px|Cache Stats (ARC) under a read-heavy load. Most Frequently Used tracks Total so closely that the green line hides behind it.]] | |||
Three charts read from the OpenZFS ARC kernel counters. [[Performance Tuning]] covers what to do about what you see here; this section covers reading it. | |||
'''Cache Efficiency''' plots one series, '''Cache Hit Ratio''', as a percentage. | |||
* '''It is a since-boot average, not a live hit rate.''' The value is the appliance's cumulative ARC hits divided by cumulative hits plus misses, so on a system that has been up for days it is dominated by history: a burst of cache misses now moves it by a fraction of a percent, and the chart will look like a flat band near 100% either way. Judged on its own it is close to useless for diagnosing a live problem. | |||
* To get the ''current'' hit rate, take two samples of the raw counters a few seconds apart and difference them -- see [[#qs-iostat: the console view|qs-iostat]] below. | |||
* Where the lifetime ratio is genuinely useful is capacity planning over the life of a pool: a system that has never got above a low ratio has never had its working set fit in RAM. | |||
'''Cache Hits''' plots the ''rate'' of ARC hits, derived from the cumulative counters, broken into three series: | |||
* '''Total''' -- all ARC hits. | |||
* '''Most Recently Used''' -- hits on blocks the ARC is holding because they were read recently. | |||
* '''Most Frequently Used''' -- hits on blocks it is holding because they are read repeatedly. | |||
The split is the useful part. A workload whose hits are almost all on the most-recently-used side is streaming through the cache and reusing very little of it; more RAM will not help it. A workload whose hits are almost all most-frequently-used is reusing a working set, which is what a read cache is for. In the capture above, a repeated read of the same file puts almost every hit on the most-frequently-used side -- so closely that the green series is drawn underneath '''Total''' and is visible only where the two diverge slightly. | |||
'''Cache Size''' plots four series in bytes: | |||
* '''Current Size''' -- how much RAM the ARC is holding right now. | |||
* '''Target Size [Adaptive]''' -- the size OpenZFS is currently aiming for. This is the adaptive part of "adaptive replacement cache": the kernel raises and lowers the target continuously in response to cache pressure and to memory demand from everything else on the appliance. Current Size chases this line, so the gap between the two shows the ARC growing or being shrunk. | |||
* '''Max Size [High Water]''' -- the ceiling the ARC may grow to. This is what the pool '''Cache Size (% of RAM)''' setting controls. | |||
* '''Min Size [Hard Limit]''' -- the floor the kernel will not shrink below under memory pressure. | |||
'''The y-axis is scaled by Max Size''', which is usually far larger than the other three, so Current Size can look pinned to zero while actually holding hundreds of MB. Read the values rather than the shape. A Current Size sitting well below Max Size under sustained load means the ARC is not the constraint -- the appliance is not being asked to cache more than it can. | |||
=== Sensors === | |||
{{Navigation|Storage Management → Storage Systems → ''select a Storage System'' → Sensors ''(dashboard toolbar)''}} | |||
Two charts: '''Average CPU Temp''' in Celsius and '''System Power''' in Watts. Both come from IPMI, sampled every five minutes, so the view is fixed to a one-hour window and the time range selector is hidden while it is showing. It refreshes itself every five seconds. | |||
Both panels read "No data to display" on a virtual appliance and on hardware without IPMI/DCMI, for the reasons given under [[#Appliance tiles|Appliance tiles]]. | |||
They are worth a look when throughput falls off under sustained load rather than at the start of it: a thermally throttled CPU presents as a storage problem. | |||
== Per-object performance views == | |||
The same dashboard panel follows the tree selection, and which views it offers depends on what is selected. | |||
{| class="wikitable" | |||
! Selected object !! Views offered | |||
|- | |||
| Storage System || Performance, Cache Stats (ARC), Sensors | |||
|- | |||
| Storage Pool || Performance, Capacity, Trends | |||
|- | |||
| Storage Volume || Performance, Capacity, Trends, Replication | |||
|- | |||
| Network Share || Capacity, Trends, NFS Stats, SMB Stats, Replication, User Usage, Group Usage. A CephFS share offers only User Usage and Group Usage. | |||
|- | |||
| Physical Disk || A single performance view, with no view buttons. | |||
|- | |||
| Ceph OSD, Ceph Pool, Ceph RGW, Ceph MDS || A single view each, with no view buttons. | |||
|} | |||
There is no '''Performance''' view for a Network Share: file protocol activity is under '''NFS Stats''' and '''SMB Stats''' instead, and the underlying device throughput is on the pool. | |||
=== Storage Pool: Performance === | |||
{{Navigation|Storage Management → Storage Pools → ''select a Storage Pool'' → Performance ''(dashboard toolbar)''}} | |||
[[File:perfmon_pool_perf.png|thumb|right|800px|Pool Performance during a sustained NFS write and re-read. IOPS and throughput are both split into read and write.]] | |||
'''IOPS''' and '''Throughput''' (bytes per second), each split into '''read''' and '''write'''. Both are summed over the pool's member devices -- its data, cache and log devices -- so this is the pool's aggregate device-level load, not its client-visible load. Reads served from the ARC never reach the devices and so never appear here, which is why a cache-friendly workload can show high client throughput and almost no pool read activity. | |||
=== Storage Volume and Physical Disk === | |||
{{Navigation|Storage Management → Physical Disks → ''select a Physical Disk''}} | |||
[[File:perfmon_disk_perf.png|thumb|right|800px|Physical Disk Dashboard for a pool member device: IOPS, throughput and latency.]] | |||
Both plot '''IOPS''', '''Throughput''' and '''Latency''' in microseconds. This is the view to use to confirm that load is landing where you expect during a test, and to spot one device in a pool behaving differently from its peers. | |||
'''A physical disk only has these series if QuantaStor is collecting them for that device.''' Collection is restricted to storage pool data, cache and log devices, storage volume devices, and Ceph OSD and journal devices -- see [[#What is collected|What is collected]]. A spare or unused disk has no data and its dashboard stays empty. | |||
=== Network Share: NFS Stats and SMB Stats === | |||
{{Navigation|Storage Management → Network Shares → ''select a Network Share'' → NFS Stats ''(dashboard toolbar)''}} | |||
[[File:perfmon_share_nfs.png|thumb|right|800px|NFS Stats. The throughput and IOPS charts are appliance-wide NFS server totals; only the client count is per share.]] | |||
'''NFS Stats''' has three charts: | |||
* '''NFS Server Throughput''' -- bytes per second read and written by the kernel NFS server. | |||
* '''NFS Server IOPS''' -- NFS read and write operations per second, combining NFSv3 and NFSv4. | |||
* '''NFS Export Clients''' -- the number of clients connected to this share's export. | |||
'''The throughput and IOPS charts are appliance-wide, not per share.''' They are derived from the kernel's NFS server counters, which the kernel does not break down per export, so the same two charts appear on every share on the appliance. Only the client count is specific to the selected share. If you need to attribute NFS load to a share, drive one share at a time. | |||
'''SMB Stats''' has two charts, both genuinely per share and both derived from the SMB server's session table rather than from byte counters: '''SMB Share Activity''' (connected clients, open files and locked files on this share) and '''SMB Client Activity''' (open files broken out per client address). | |||
'''User Usage''' and '''Group Usage''' plot per-user and per-group space and file counts inside the share. That is share usage tracking rather than performance monitoring; it is collected on its own schedule and is documented with [[Network Shares]]. | |||
=== Trends and Capacity === | |||
'''Trends''' plots '''Average Utilized Space''' for a pool, volume or share, with '''Size''', '''Free Space''', '''Raw Size''' and '''Raw Utilized Size'''. Two things to know before reading it: | |||
* '''The window is never shorter than six hours.''' The utilization samples behind this chart are written once an hour, so anything shorter would hold at most one point. Any selection under six hours is treated as six hours -- and the range selector goes on displaying whatever you picked, so the label can disagree with the axis. Trust the axis. | |||
* '''Raw Utilized Size is not populated on scale-up (ZFS) pools''' and reads zero there. Use '''Size''' and '''Free Space'''. | |||
'''Capacity''' is a current-state breakdown rather than a time series -- for a share, a bar of pool utilized size against the share's logical and physical usage -- so the range selector is hidden while it is showing. Capacity planning is covered on [[Storage Pools]] and [[Network Shares]]. | |||
== Time range, refresh and zoom == | |||
The range selector offers '''Last 1 minute''' (the default), 5 minutes, 15 minutes, 1 hour, 2 hours, 6 hours, 12 hours, 24 hours, 2 days, 7 days and 30 days. | |||
'''Last 1 minute is a live feed; every other range is a snapshot.''' At one minute the dashboard subscribes to a push stream from the statistics service and the charts scroll continuously as samples arrive. At any other range it runs a single query and re-runs it periodically: | |||
{| class="wikitable" | |||
! View !! Behaviour outside "Last 1 minute" | |||
|- | |||
| Performance, NFS Stats, SMB Stats, Replication || Re-queried automatically every 60 seconds. | |||
|- | |||
| Cache Stats (ARC), Trends || Not re-queried on a timer. Re-select the range, or re-select the object, to pull fresh data. | |||
|- | |||
| Sensors || Always a fixed one-hour window, re-queried every 5 seconds. | |||
|- | |||
| Capacity || Current state; no time range. | |||
|} | |||
'''Click and drag on a chart to zoom in''' -- that is also what the selector's own tooltip says. Zooming works within whatever range is loaded; it does not fetch finer data than the range provides, which matters because of the next point. | |||
=== Which resolution each range reads === | |||
The statistics database keeps three resolutions of the same measurements, and the range you pick decides which one the chart reads: | |||
{| class="wikitable" | |||
! Range selected !! Series read !! Spacing between points | |||
|- | |||
| Last 1 minute to Last 6 hours || Raw samples || 2 seconds | |||
|- | |||
| Last 12 hours to Last 7 days || 10-minute averages || 10 minutes | |||
|- | |||
| Last 30 days || 2-hour averages || 2 hours | |||
|} | |||
'''This is the single most useful thing to know when a graph looks wrong.''' Raw samples are kept for six hours only, so: | |||
* A range of 12 hours or more is drawn from averages. A one-second latency spike that is plainly visible at 15 minutes is averaged into invisibility at 24 hours. If you are hunting a spike, look at it inside the six-hour raw window. | |||
* A range of 12 hours or more on a system that has only been collecting for a few hours can be empty even though the shorter ranges have data, because the downsampled series has not been written yet. | |||
== The statistics database == | |||
Three services behind the charts, all local to each appliance: | |||
{| class="wikitable" | |||
! Service !! Role | |||
|- | |||
| {{Code|1=influxdb}} || The time-series database. One database, {{Code|1=quantastor}}. | |||
|- | |||
| {{Code|1=telegraf}} || The collector. Samples the kernel and pushes into InfluxDB. | |||
|- | |||
| {{Code|1=qs-statsd}} || The query and push front end the web interface reads from. | |||
|} | |||
QuantaStor supervises all three, re-checking them every ten minutes, and restarts telegraf whenever the collection configuration changes. It also '''pauses collection during some configuration changes''' -- it stops telegraf while it reconfigures and restarts it afterwards -- which leaves a short gap in every chart. A gap of a few seconds around a pool or network change is normal. | |||
Each appliance collects and stores its own statistics; there is no grid-wide database. A Grid View tile is drawn from the data on that member. | |||
{{Code|1=qs-status}} reports {{Code|1=qs-statsd}}; check the other two with {{Code|1=systemctl is-active influxdb telegraf}}. | |||
InfluxDB listens on {{Code|1=localhost:8086}} only. Creating the touch file {{Code|1=/var/opt/osnexus/quantastor/touchfiles/tf_influxdb_remote_access.enable}} makes it bind all interfaces so an external client can query it directly; authentication still applies. The database user is {{Code|1=quantastor}} and its generated password is held in {{Code|1=/etc/influxdb/quantastor.txt}}. | |||
=== Sampling interval === | |||
'''Telegraf samples every 2 seconds and flushes every 2 seconds.''' The NFS and SMB collector is the one exception and runs every 5 seconds. Two seconds is therefore the finest resolution any chart can show, and it is the spacing of every point in the raw six-hour window. | |||
Two feeds are far slower, and they are slower by design rather than by configuration: | |||
{| class="wikitable" | |||
! Feed !! Interval !! Feeds | |||
|- | |||
| Pool, volume and share utilization || 1 hour || Trends | |||
|- | |||
| CPU temperature and system power || 5 minutes || Sensors, and the tile temperature and power charts | |||
|} | |||
The live configuration is {{Code|1=/etc/telegraf/telegraf.conf}}, but do not edit it: QuantaStor regenerates it from the shipped template {{Code|1=/opt/osnexus/quantastor/conf/telegraf.qstor}} and restarts telegraf whenever the template or the device list changes, so an edit is lost at the next reconciliation. To change the collector configuration, place a modified copy at {{Code|1=/opt/osnexus/quantastor/conf/telegraf.qstor.override}} -- the shipped template carries the same instruction in its own header. We recommend contacting OSNEXUS support before changing the sampling interval: the downsampling described below assumes the shipped one. | |||
=== Retention and downsampling === | |||
Raw samples are kept for six hours; continuous queries roll them up into two coarser series that are kept for much longer. | |||
{| class="wikitable" | |||
! Retention policy !! Duration !! Holds | |||
|- | |||
| {{Code|1=quantastor}} ''(default)'' || 6 hours || Raw 2-second samples. | |||
|- | |||
| {{Code|1=quantastor_rp_3w}} || 3 weeks || 10-minute averages, as {{Code|1=<measurement>_10m}}. | |||
|- | |||
| {{Code|1=quantastor_rp_78w}} || 78 weeks || 2-hour averages, as {{Code|1=<measurement>_2h}}, rolled up from the 10-minute series. | |||
|- | |||
| {{Code|1=quantastor_rp_usage}} || 12 weeks || Per-user and per-group share usage, kept at the resolution it was collected at rather than downsampled. | |||
|} | |||
Consequences worth keeping in mind: | |||
* '''Six hours is all the detail you get.''' If you want per-second detail of an incident, collect it while the incident is happening, or export it. It is gone six hours later. | |||
* Every appliance keeps roughly eighteen months of two-hourly history, which is enough to answer "was this always this slow?". | |||
* The share usage retention comes from {{Code|1=retention_weeks}} in {{Code|1=/opt/osnexus/quantastor/conf/qs_shareusage.conf}}, and is re-applied when the QuantaStor service restarts. | |||
The policies and continuous queries are created and corrected at every service start, so a database that has been restored or partly dropped repairs itself on the next restart rather than needing manual work. | |||
=== What is collected === | |||
{| class="wikitable" | |||
! Source !! Covers | |||
|- | |||
| CPU || Per-core and total: user, system, idle and I/O wait. | |||
|- | |||
| Memory and swap || Used, free, cached and swap activity. | |||
|- | |||
| System || Load average and uptime. | |||
|- | |||
| Network || Bytes and packets sent and received, per interface. | |||
|- | |||
| ZFS || ARC, L2ARC, ZIL and prefetch counters, plus per-pool metrics. | |||
|- | |||
| Block device I/O || Reads, writes, bytes and service time -- '''only''' for an explicit device list. | |||
|- | |||
| NFS and SMB || Kernel NFS server counters, per-export client counts, and per-share and per-client SMB session counts. | |||
|- | |||
| Ceph || Cluster, pool, OSD, RGW and MDS counters, added automatically when the appliance becomes a Ceph cluster member. | |||
|} | |||
'''Block device I/O is collected only for the devices QuantaStor enumerates''': storage pool data, cache and log devices, storage volume devices, and Ceph OSD and journal devices. QuantaStor rewrites that list and restarts telegraf whenever pool or volume membership changes. The practical effect is that a disk which is not part of a pool has no I/O statistics at all -- which is deliberate, since collecting per-device metrics for every disk in a large JBOD chain is expensive and mostly uninteresting. | |||
''(Note: the collected measurement set changes between releases as the collectors are extended; {{Code|1=/opt/osnexus/quantastor/conf/telegraf.qstor}} on the appliance is authoritative for what a given build gathers.)'' | |||
=== qs-statsdb === | |||
{{Code|1=qs-statsdb}} is the maintenance and triage tool for the database. It requires root, and it is the right tool when a chart is empty and you need to know whether the problem is collection, storage or display. [[QuantaStor Shell Utilities]] documents it in full; the checks worth knowing here are: | |||
<pre style="font-size: smaller"> | |||
qs-statsdb showrp # the retention policies actually in force | |||
qs-statsdb showcq # the continuous queries doing the downsampling | |||
qs-statsdb showmm # the measurements being collected | |||
qs-statsdb count cpu # are samples arriving at all? | |||
qs-statsdb print_zfs_arc # recent ARC rows, straight from the database | |||
</pre> | |||
If {{Code|1=count}} rises between two runs, collection and storage are working and the problem is in the view. If it does not, look at telegraf. | |||
== qs-iostat: the console view == | |||
{{Code|1=qs-iostat}} reports the same kernel counters the dashboards summarise, from a shell, without root: | |||
<pre style="font-size: smaller"> | |||
qs-iostat -c # globally averaged CPU statistics | |||
qs-iostat -d # I/O statistics for all block devices, in MB/s | |||
qs-iostat -a # ZFS ARC, L2ARC and ZIL counters | |||
qs-iostat -f # repeat every 2 seconds | |||
qs-iostat --extra "-x" # pass extra arguments through to iostat | |||
</pre> | |||
'''The ARC and ZIL counters it prints are cumulative since boot''', so take two samples and difference them rather than reading absolute numbers. That is exactly how you get the live cache hit rate that '''Cache Efficiency''' cannot give you: difference {{Code|1=hits}} and {{Code|1=misses}} between two runs a few seconds apart, and take the ratio of the differences. | |||
What each counter means, and what to do about it, is on [[Performance Tuning]]. The full option list is in [[QuantaStor Shell Utilities]]. | |||
{{Code|1=qs-iostat}} is also useful for a reason the dashboards cannot cover: it reads the kernel directly, so it still works when the statistics database or the collector is down. | |||
== External dashboards and Grafana == | |||
{{Navigation|External Dashboards → Dashboard → Add ''(toolbar)''}} | |||
[[File:perfmon_external_dashboards.png|thumb|right|800px|The External Dashboards tab before any entry has been added. Its default page carries the Grafana embedding requirement.]] | |||
The '''External Dashboards''' tab embeds monitoring pages of your own inside the QuantaStor interface, so an administrator does not have to switch tools to correlate a storage symptom with the rest of the estate. It holds two object types: | |||
* '''Dashboard''' -- a named URL. '''Add''' opens the create dialog; '''Remove''' deletes an entry. | |||
* '''Dashboard Group''' -- a named set of dashboards, for organizing a long list. '''Add''', '''Modify''' and '''Remove''' manage groups and their membership. | |||
[[File:perfmon_extdash_create.png|thumb|right|800px|Create External Dashboard. The name is pre-filled; only Name and URL are required.]] | |||
The create dialog has three fields: '''Name''' (pre-filled with a generated name such as {{Code|1=external-dashboard-1}}), '''Description''', and '''URL'''. Name and URL are both required. | |||
'''The page is loaded in an embedded frame, so the target has to permit framing.''' Most monitoring tools refuse to be framed by default. For Grafana that means setting {{Code|1=allow_embedding = true}} in the {{Code|1=[security]}} section of {{Code|1=/etc/grafana/grafana.ini}}; the appliance ships {{Code|1=/opt/osnexus/quantastor/bin/enable_grafana_iframe.sh}}, which sets that and {{Code|1=x_frame_options = allowall}}, backs up the original file, and restarts Grafana. A target that refuses framing shows as an empty or blocked frame with no error in the QuantaStor interface, so this is worth checking first when a newly added dashboard does not render. | |||
Grafana Enterprise is installed on the appliance but its service is neither enabled nor provisioned by default -- there is no shipped QuantaStor dashboard set, and no data source is configured for you. If you want to build your own, the statistics database is available to Grafana as an InfluxDB 1.x data source: database {{Code|1=quantastor}}, user {{Code|1=quantastor}}, password from {{Code|1=/etc/influxdb/quantastor.txt}}, and the retention policy names from [[#Retention and downsampling|Retention and downsampling]] above. Query the {{Code|1=_10m}} and {{Code|1=_2h}} series for anything longer than six hours, for the reasons given there. | |||
'''The whole tab can be hidden per user''' through the Web Interface Customization settings on a user account, which is worth doing for operators who should not be following links out of the appliance. | |||
== Performance-related alerts == | |||
Most of what QuantaStor alerts on around the dashboards is capacity, not performance. The pool free-space thresholds behind '''Pool Capacity Threshold Alerts''' are the ones administrators change most often; [[Alert Manager]] owns them, and [[Call-home / Alerting]] covers how the resulting notifications are delivered. | |||
The performance-adjacent conditions that do raise an alert are all thermal or power: | |||
{| class="wikitable" | |||
! Condition !! Notes | |||
|- | |||
| System, motherboard (PCH) or CPU temperature out of range || Detected from the IPMI sensor readings, at most one alert per condition per day. Which sensor names count as which reading is mapped in {{Code|1=/opt/osnexus/quantastor/conf/qs_ipmi.conf}}, per platform. | |||
|- | |||
| Power supply degraded, failed, or recovered || From the same IPMI sampling, using the power supply sensor names in the same file. | |||
|- | |||
| Fan health || From the same IPMI sampling. | |||
|- | |||
| Disk over temperature || Raised through the hardware controller adapters when a drive reaches the temperature threshold, 60 C by default. The threshold is {{Code|1=drive_temp_alert_threshold}} in {{Code|1=/opt/osnexus/quantastor/conf/qs_ipmi.conf}}; values below 50 C are ignored and clamped to 50, and values above 200 C are clamped to 200. At most one alert per day. | |||
|} | |||
'''There is no alert on I/O latency, IOPS or throughput.''' Nothing in QuantaStor watches a performance number against a threshold, so a pool that has become slow will not tell you -- it has to be noticed on these dashboards, or by exporting the series to a monitoring system that can threshold them. That is the main argument for pointing an external dashboard at the statistics database on an appliance where performance matters. | |||
== When a chart is empty or looks wrong == | |||
In roughly the order worth checking: | |||
* '''The range is longer than six hours.''' Beyond that the chart reads 10-minute or 2-hour averages, and short spikes are averaged away. See [[#Which resolution each range reads|Which resolution each range reads]]. | |||
* '''The chart is one that does not auto-refresh.''' Cache Stats (ARC) and Trends are only re-queried when you change the range or the selection; everything looks stale until you do. | |||
* '''Cache Efficiency looks pinned at 100%.''' It is a since-boot average and it is meant to look like that. Difference the raw counters instead. | |||
* '''The Trends axis disagrees with the range selector.''' Trends never draws less than six hours; the selector is not updated to match. | |||
* '''Sensors and the tile temperature and power charts are empty.''' Expected on a virtual appliance and on hardware without IPMI/DCMI. | |||
* '''A physical disk has no data.''' Device I/O is only collected for pool, volume and Ceph OSD/journal devices. | |||
* '''There is a short gap in every chart at the same moment.''' Collection is paused while QuantaStor reconciles some configuration changes. | |||
* '''Nothing at all, on every chart.''' Check the three services -- {{Code|1=qs-status}} for {{Code|1=qs-statsd}}, {{Code|1=systemctl is-active influxdb telegraf}} for the other two -- then confirm samples are arriving with {{Code|1=qs-statsdb count cpu}} twice a few seconds apart. | |||
* '''Nothing at all, after a recent restore or database maintenance.''' The retention policies and continuous queries are rebuilt at service start; restart the QuantaStor service and re-check. | |||
If the numbers are present but you do not believe them, take the measurement a second way: {{Code|1=qs-iostat}} reads the kernel directly, and [[Performance Testing]] covers generating a known load so there is something unambiguous to compare against. | |||
== Related pages == | |||
* [[Performance Testing]] -- proving where a bottleneck is, with {{Code|1=qs-perftest}} and {{Code|1=qs-ramdisk}} | |||
* [[Performance Tuning]] -- I/O profiles, ARC and write log sizing, and the rest of the tunables | |||
* [[Storage Pools]] -- pool layout, capacity and cache device configuration | |||
* [[Storage System]] -- appliance and grid configuration, including grid membership | |||
* [[Alert Manager]] -- the capacity thresholds behind the pool health chart, and the alert types | |||
* [[Call-home / Alerting]] -- alert handlers, call-home and where notifications are delivered | |||
* [[QuantaStor Shell Utilities]] -- {{Code|1=qs-iostat}}, {{Code|1=qs-statsdb}} and the rest of the console tools | |||
* [[Network Shares]] -- share usage tracking behind the User Usage and Group Usage views | |||
* [[QuantaStor CLI Command Reference]] -- full argument lists for every {{Code|1=qs}} command | |||
---- | |||
<small>''Verified against QuantaStor 6.9.0.''</small> | |||
Revision as of 08:18, 3 September 2026
This page covers the performance and health information QuantaStor collects about itself, where each view lives in the web interface, and how to read the numbers it shows. It also documents the statistics database behind the charts -- how often it samples, how long it keeps each resolution, and what makes a chart look wrong.
For the diagnostic method -- how to prove where a bottleneck is -- see Performance Testing. For the settings that change performance, see Performance Tuning.
| Section | Purpose |
|---|---|
| The Grid View dashboard | Grid-wide health, capacity and per-appliance sparklines. |
| The Storage System dashboard | CPU, network and memory; ARC cache statistics; thermal and power sensors. |
| Per-object performance views | Pool, volume, share and physical disk charts. |
| Time range, refresh and zoom | What each range reads, and what updates on its own. |
| The statistics database | Sampling interval, retention, downsampling, and what is collected. |
| qs-iostat: the console view | Reading the same counters from the command line. |
| External dashboards and Grafana | Embedding your own monitoring pages. |
| Performance-related alerts | Which conditions raise an alert, and which do not. |
| When a chart is empty or looks wrong | The common causes, in order. |
The Grid View dashboard

Grid View is the landing view for a whole grid. It has three regions, and two checkboxes on its toolbar control which of them are shown:
- Grid Details (on by default) shows the five grid-wide counters and the two pool capacity charts. Clear it and only the per-appliance tiles remain.
- System Details (off by default) expands every appliance tile with capacity, session and health detail.
+Add System and -Remove System open the grid membership dialogs; those belong to Storage System rather than to monitoring.
The grid counters
The five counters across the top are current state, computed from the objects the web interface already holds -- they are not time series, so they do not depend on the statistics database.
| Counter | Reads |
|---|---|
| Systems Online | Appliances in a normal state, over the total number of grid members. A member counts as online only if it is not in an error, missing, warning, offline or initializing state, so a healthy-but-warning appliance shows as not online. "2/3" on three running appliances usually means one is in a warning state, not that one is down. |
| Systems with Alerts | Appliances that have at least one alert that has not been dismissed, over the total. |
| Pools Online | Storage pools in a normal state, over the total. Internal pools are excluded from every pool count and chart on this dashboard. |
| Pools with Alerts | Pools with at least one undismissed alert, over the total. |
| Alerts | Total undismissed alerts across the grid. |
A counter's footer turns amber when the highest severity in its scope is a warning and red when it is an error; View Details opens the matching list or the alert view. These count alert records, not live conditions, so an alert raised for a problem that has since been fixed still counts until someone dismisses it -- use Dismiss Alerts on the Alerts pane. See Alert Manager for what raises them and Call-home / Alerting for where they are delivered.
Pool Capacity Threshold Alerts
This chart counts pools into Ok, Warning and Alert by how full they are, against the pool free-space thresholds configured in Alert Manager. The shipped defaults are a warning at 30% free (70% utilized) and an alert at 10% free (90% utilized), so a pool is counted:
- Ok below 70% utilized,
- Warning between 70% and 90% utilized,
- Alert above 90% utilized.
Read the thresholds actually in force with qs alert-config-get and change them with qs alert-config-set or in Alert Manager, which owns them. Alert Manager also carries a third, critical threshold (5% free by default) that this chart does not break out.
Total Pool Capacity
The second chart breaks the grid's combined pool capacity into Free, Volumes, Volume Snapshots, Shares, Share Snapshots and Other. Other is the remainder -- total capacity less free space less the four accounted categories -- so it absorbs pool metadata and anything on the pool that is not represented as a volume or a share. A few MB in Other on an otherwise empty pool is normal.
Appliance tiles

Each grid member gets a tile. The header carries the appliance name, its management address and an Alerts [n] button whose colour follows the highest severity present; the button opens that appliance's alerts. Below it:
- The chassis or platform string and, for hardware with IPMI, the POWER state and power supply icons.
- A thermometer showing the average CPU temperature. It renders green between 15 and 32 C and amber outside that band.
- CPU USAGE and MEMORY USAGE sparklines covering the last minute, refreshed every five seconds.
- CPU Temp (Last Hour) and SYSTEM POWER (Last Hour) charts. These two are always a one-hour window regardless of anything else on the page, because their samples arrive only every five minutes.
- A footer with the appliance's object state, and an arrow that selects that appliance in Storage Systems.
On a virtual appliance the temperature and power panels are empty and the thermometer reads 0 C. That is expected, not a fault: QuantaStor skips the IPMI power reading entirely on a virtual machine, and it writes a temperature or power sample only when the value it read is non-zero. The same applies to hardware with no IPMI/DCMI support.
With System Details enabled each tile also shows:
| Panel | Contents |
|---|---|
| Total / Used Pool Capacity | A utilization bar for that appliance's pools, with free space, total and used. |
| Connections | Current iSCSI, Fibre Channel and SMB session counts. Each is a link into the matching sessions view. |
| Pool Capacity Health | That appliance's pools counted Ok / Warning / Alert against the same thresholds as the grid chart above. |
| Ceph OSD Health | OSDs counted Ok / Warning / Alert, and zero throughout on an appliance that is not a Ceph cluster member. |
A badge is grey when its count is zero, so grey means "none of these", not "unknown".
How often Grid View updates
Grid View redraws when grid objects change, coalesced to at most once every 15 seconds, and it redraws unconditionally at least once a minute so the sparklines keep moving even on an idle grid. The per-appliance sparklines fetch their own data every five seconds, independently of that.
The Storage System dashboard

Selecting a grid member under Storage Systems puts a dashboard panel above the object grid, titled Storage System Dashboard (name). Selecting the grid root instead of a member shows no dashboard. The panel's toolbar carries three views -- Performance, Cache Stats (ARC) and Sensors -- and a time range selector. The double-chevron icon at the far right collapses the whole dashboard panel; it is not a refresh control.
Performance
Three charts, all for the selected appliance:
- CPU plots Idle, System and Wait-IO as percentages. Read it carefully: the line sitting near 100% is Idle. High Wait-IO alongside high Idle means the appliance is waiting on storage, not short of CPU -- and no amount of CPU or thread tuning will change that.
- Network plots bytes per second per network port, as TX:<port> and RX:<port>. Because it is per port rather than aggregated, it is also the quickest way to see that traffic you expected to be spread over several ports has collapsed onto one.
- Memory plots memory in use as a percentage. On a scale-up appliance most of it is usually the ZFS ARC; see Performance Tuning for what that costs and how to bound it.
Cache Stats (ARC)

Three charts read from the OpenZFS ARC kernel counters. Performance Tuning covers what to do about what you see here; this section covers reading it.
Cache Efficiency plots one series, Cache Hit Ratio, as a percentage.
- It is a since-boot average, not a live hit rate. The value is the appliance's cumulative ARC hits divided by cumulative hits plus misses, so on a system that has been up for days it is dominated by history: a burst of cache misses now moves it by a fraction of a percent, and the chart will look like a flat band near 100% either way. Judged on its own it is close to useless for diagnosing a live problem.
- To get the current hit rate, take two samples of the raw counters a few seconds apart and difference them -- see qs-iostat below.
- Where the lifetime ratio is genuinely useful is capacity planning over the life of a pool: a system that has never got above a low ratio has never had its working set fit in RAM.
Cache Hits plots the rate of ARC hits, derived from the cumulative counters, broken into three series:
- Total -- all ARC hits.
- Most Recently Used -- hits on blocks the ARC is holding because they were read recently.
- Most Frequently Used -- hits on blocks it is holding because they are read repeatedly.
The split is the useful part. A workload whose hits are almost all on the most-recently-used side is streaming through the cache and reusing very little of it; more RAM will not help it. A workload whose hits are almost all most-frequently-used is reusing a working set, which is what a read cache is for. In the capture above, a repeated read of the same file puts almost every hit on the most-frequently-used side -- so closely that the green series is drawn underneath Total and is visible only where the two diverge slightly.
Cache Size plots four series in bytes:
- Current Size -- how much RAM the ARC is holding right now.
- Target Size [Adaptive] -- the size OpenZFS is currently aiming for. This is the adaptive part of "adaptive replacement cache": the kernel raises and lowers the target continuously in response to cache pressure and to memory demand from everything else on the appliance. Current Size chases this line, so the gap between the two shows the ARC growing or being shrunk.
- Max Size [High Water] -- the ceiling the ARC may grow to. This is what the pool Cache Size (% of RAM) setting controls.
- Min Size [Hard Limit] -- the floor the kernel will not shrink below under memory pressure.
The y-axis is scaled by Max Size, which is usually far larger than the other three, so Current Size can look pinned to zero while actually holding hundreds of MB. Read the values rather than the shape. A Current Size sitting well below Max Size under sustained load means the ARC is not the constraint -- the appliance is not being asked to cache more than it can.
Sensors
Two charts: Average CPU Temp in Celsius and System Power in Watts. Both come from IPMI, sampled every five minutes, so the view is fixed to a one-hour window and the time range selector is hidden while it is showing. It refreshes itself every five seconds.
Both panels read "No data to display" on a virtual appliance and on hardware without IPMI/DCMI, for the reasons given under Appliance tiles.
They are worth a look when throughput falls off under sustained load rather than at the start of it: a thermally throttled CPU presents as a storage problem.
Per-object performance views
The same dashboard panel follows the tree selection, and which views it offers depends on what is selected.
| Selected object | Views offered |
|---|---|
| Storage System | Performance, Cache Stats (ARC), Sensors |
| Storage Pool | Performance, Capacity, Trends |
| Storage Volume | Performance, Capacity, Trends, Replication |
| Network Share | Capacity, Trends, NFS Stats, SMB Stats, Replication, User Usage, Group Usage. A CephFS share offers only User Usage and Group Usage. |
| Physical Disk | A single performance view, with no view buttons. |
| Ceph OSD, Ceph Pool, Ceph RGW, Ceph MDS | A single view each, with no view buttons. |
There is no Performance view for a Network Share: file protocol activity is under NFS Stats and SMB Stats instead, and the underlying device throughput is on the pool.
Storage Pool: Performance

IOPS and Throughput (bytes per second), each split into read and write. Both are summed over the pool's member devices -- its data, cache and log devices -- so this is the pool's aggregate device-level load, not its client-visible load. Reads served from the ARC never reach the devices and so never appear here, which is why a cache-friendly workload can show high client throughput and almost no pool read activity.
Storage Volume and Physical Disk

Both plot IOPS, Throughput and Latency in microseconds. This is the view to use to confirm that load is landing where you expect during a test, and to spot one device in a pool behaving differently from its peers.
A physical disk only has these series if QuantaStor is collecting them for that device. Collection is restricted to storage pool data, cache and log devices, storage volume devices, and Ceph OSD and journal devices -- see What is collected. A spare or unused disk has no data and its dashboard stays empty.

NFS Stats has three charts:
- NFS Server Throughput -- bytes per second read and written by the kernel NFS server.
- NFS Server IOPS -- NFS read and write operations per second, combining NFSv3 and NFSv4.
- NFS Export Clients -- the number of clients connected to this share's export.
The throughput and IOPS charts are appliance-wide, not per share. They are derived from the kernel's NFS server counters, which the kernel does not break down per export, so the same two charts appear on every share on the appliance. Only the client count is specific to the selected share. If you need to attribute NFS load to a share, drive one share at a time.
SMB Stats has two charts, both genuinely per share and both derived from the SMB server's session table rather than from byte counters: SMB Share Activity (connected clients, open files and locked files on this share) and SMB Client Activity (open files broken out per client address).
User Usage and Group Usage plot per-user and per-group space and file counts inside the share. That is share usage tracking rather than performance monitoring; it is collected on its own schedule and is documented with Network Shares.
Trends and Capacity
Trends plots Average Utilized Space for a pool, volume or share, with Size, Free Space, Raw Size and Raw Utilized Size. Two things to know before reading it:
- The window is never shorter than six hours. The utilization samples behind this chart are written once an hour, so anything shorter would hold at most one point. Any selection under six hours is treated as six hours -- and the range selector goes on displaying whatever you picked, so the label can disagree with the axis. Trust the axis.
- Raw Utilized Size is not populated on scale-up (ZFS) pools and reads zero there. Use Size and Free Space.
Capacity is a current-state breakdown rather than a time series -- for a share, a bar of pool utilized size against the share's logical and physical usage -- so the range selector is hidden while it is showing. Capacity planning is covered on Storage Pools and Network Shares.
Time range, refresh and zoom
The range selector offers Last 1 minute (the default), 5 minutes, 15 minutes, 1 hour, 2 hours, 6 hours, 12 hours, 24 hours, 2 days, 7 days and 30 days.
Last 1 minute is a live feed; every other range is a snapshot. At one minute the dashboard subscribes to a push stream from the statistics service and the charts scroll continuously as samples arrive. At any other range it runs a single query and re-runs it periodically:
| View | Behaviour outside "Last 1 minute" |
|---|---|
| Performance, NFS Stats, SMB Stats, Replication | Re-queried automatically every 60 seconds. |
| Cache Stats (ARC), Trends | Not re-queried on a timer. Re-select the range, or re-select the object, to pull fresh data. |
| Sensors | Always a fixed one-hour window, re-queried every 5 seconds. |
| Capacity | Current state; no time range. |
Click and drag on a chart to zoom in -- that is also what the selector's own tooltip says. Zooming works within whatever range is loaded; it does not fetch finer data than the range provides, which matters because of the next point.
Which resolution each range reads
The statistics database keeps three resolutions of the same measurements, and the range you pick decides which one the chart reads:
| Range selected | Series read | Spacing between points |
|---|---|---|
| Last 1 minute to Last 6 hours | Raw samples | 2 seconds |
| Last 12 hours to Last 7 days | 10-minute averages | 10 minutes |
| Last 30 days | 2-hour averages | 2 hours |
This is the single most useful thing to know when a graph looks wrong. Raw samples are kept for six hours only, so:
- A range of 12 hours or more is drawn from averages. A one-second latency spike that is plainly visible at 15 minutes is averaged into invisibility at 24 hours. If you are hunting a spike, look at it inside the six-hour raw window.
- A range of 12 hours or more on a system that has only been collecting for a few hours can be empty even though the shorter ranges have data, because the downsampled series has not been written yet.
The statistics database
Three services behind the charts, all local to each appliance:
| Service | Role |
|---|---|
influxdb |
The time-series database. One database, quantastor.
|
telegraf |
The collector. Samples the kernel and pushes into InfluxDB. |
qs-statsd |
The query and push front end the web interface reads from. |
QuantaStor supervises all three, re-checking them every ten minutes, and restarts telegraf whenever the collection configuration changes. It also pauses collection during some configuration changes -- it stops telegraf while it reconfigures and restarts it afterwards -- which leaves a short gap in every chart. A gap of a few seconds around a pool or network change is normal.
Each appliance collects and stores its own statistics; there is no grid-wide database. A Grid View tile is drawn from the data on that member.
qs-status reports qs-statsd; check the other two with systemctl is-active influxdb telegraf.
InfluxDB listens on localhost:8086 only. Creating the touch file /var/opt/osnexus/quantastor/touchfiles/tf_influxdb_remote_access.enable makes it bind all interfaces so an external client can query it directly; authentication still applies. The database user is quantastor and its generated password is held in /etc/influxdb/quantastor.txt.
Sampling interval
Telegraf samples every 2 seconds and flushes every 2 seconds. The NFS and SMB collector is the one exception and runs every 5 seconds. Two seconds is therefore the finest resolution any chart can show, and it is the spacing of every point in the raw six-hour window.
Two feeds are far slower, and they are slower by design rather than by configuration:
| Feed | Interval | Feeds |
|---|---|---|
| Pool, volume and share utilization | 1 hour | Trends |
| CPU temperature and system power | 5 minutes | Sensors, and the tile temperature and power charts |
The live configuration is /etc/telegraf/telegraf.conf, but do not edit it: QuantaStor regenerates it from the shipped template /opt/osnexus/quantastor/conf/telegraf.qstor and restarts telegraf whenever the template or the device list changes, so an edit is lost at the next reconciliation. To change the collector configuration, place a modified copy at /opt/osnexus/quantastor/conf/telegraf.qstor.override -- the shipped template carries the same instruction in its own header. We recommend contacting OSNEXUS support before changing the sampling interval: the downsampling described below assumes the shipped one.
Retention and downsampling
Raw samples are kept for six hours; continuous queries roll them up into two coarser series that are kept for much longer.
| Retention policy | Duration | Holds |
|---|---|---|
quantastor (default) |
6 hours | Raw 2-second samples. |
quantastor_rp_3w |
3 weeks | 10-minute averages, as <measurement>_10m.
|
quantastor_rp_78w |
78 weeks | 2-hour averages, as <measurement>_2h, rolled up from the 10-minute series.
|
quantastor_rp_usage |
12 weeks | Per-user and per-group share usage, kept at the resolution it was collected at rather than downsampled. |
Consequences worth keeping in mind:
- Six hours is all the detail you get. If you want per-second detail of an incident, collect it while the incident is happening, or export it. It is gone six hours later.
- Every appliance keeps roughly eighteen months of two-hourly history, which is enough to answer "was this always this slow?".
- The share usage retention comes from
retention_weeksin/opt/osnexus/quantastor/conf/qs_shareusage.conf, and is re-applied when the QuantaStor service restarts.
The policies and continuous queries are created and corrected at every service start, so a database that has been restored or partly dropped repairs itself on the next restart rather than needing manual work.
What is collected
| Source | Covers |
|---|---|
| CPU | Per-core and total: user, system, idle and I/O wait. |
| Memory and swap | Used, free, cached and swap activity. |
| System | Load average and uptime. |
| Network | Bytes and packets sent and received, per interface. |
| ZFS | ARC, L2ARC, ZIL and prefetch counters, plus per-pool metrics. |
| Block device I/O | Reads, writes, bytes and service time -- only for an explicit device list. |
| NFS and SMB | Kernel NFS server counters, per-export client counts, and per-share and per-client SMB session counts. |
| Ceph | Cluster, pool, OSD, RGW and MDS counters, added automatically when the appliance becomes a Ceph cluster member. |
Block device I/O is collected only for the devices QuantaStor enumerates: storage pool data, cache and log devices, storage volume devices, and Ceph OSD and journal devices. QuantaStor rewrites that list and restarts telegraf whenever pool or volume membership changes. The practical effect is that a disk which is not part of a pool has no I/O statistics at all -- which is deliberate, since collecting per-device metrics for every disk in a large JBOD chain is expensive and mostly uninteresting.
(Note: the collected measurement set changes between releases as the collectors are extended; /opt/osnexus/quantastor/conf/telegraf.qstor on the appliance is authoritative for what a given build gathers.)
qs-statsdb
qs-statsdb is the maintenance and triage tool for the database. It requires root, and it is the right tool when a chart is empty and you need to know whether the problem is collection, storage or display. QuantaStor Shell Utilities documents it in full; the checks worth knowing here are:
qs-statsdb showrp # the retention policies actually in force qs-statsdb showcq # the continuous queries doing the downsampling qs-statsdb showmm # the measurements being collected qs-statsdb count cpu # are samples arriving at all? qs-statsdb print_zfs_arc # recent ARC rows, straight from the database
If count rises between two runs, collection and storage are working and the problem is in the view. If it does not, look at telegraf.
qs-iostat: the console view
qs-iostat reports the same kernel counters the dashboards summarise, from a shell, without root:
qs-iostat -c # globally averaged CPU statistics qs-iostat -d # I/O statistics for all block devices, in MB/s qs-iostat -a # ZFS ARC, L2ARC and ZIL counters qs-iostat -f # repeat every 2 seconds qs-iostat --extra "-x" # pass extra arguments through to iostat
The ARC and ZIL counters it prints are cumulative since boot, so take two samples and difference them rather than reading absolute numbers. That is exactly how you get the live cache hit rate that Cache Efficiency cannot give you: difference hits and misses between two runs a few seconds apart, and take the ratio of the differences.
What each counter means, and what to do about it, is on Performance Tuning. The full option list is in QuantaStor Shell Utilities.
qs-iostat is also useful for a reason the dashboards cannot cover: it reads the kernel directly, so it still works when the statistics database or the collector is down.
External dashboards and Grafana

The External Dashboards tab embeds monitoring pages of your own inside the QuantaStor interface, so an administrator does not have to switch tools to correlate a storage symptom with the rest of the estate. It holds two object types:
- Dashboard -- a named URL. Add opens the create dialog; Remove deletes an entry.
- Dashboard Group -- a named set of dashboards, for organizing a long list. Add, Modify and Remove manage groups and their membership.

The create dialog has three fields: Name (pre-filled with a generated name such as external-dashboard-1), Description, and URL. Name and URL are both required.
The page is loaded in an embedded frame, so the target has to permit framing. Most monitoring tools refuse to be framed by default. For Grafana that means setting allow_embedding = true in the [security] section of /etc/grafana/grafana.ini; the appliance ships /opt/osnexus/quantastor/bin/enable_grafana_iframe.sh, which sets that and x_frame_options = allowall, backs up the original file, and restarts Grafana. A target that refuses framing shows as an empty or blocked frame with no error in the QuantaStor interface, so this is worth checking first when a newly added dashboard does not render.
Grafana Enterprise is installed on the appliance but its service is neither enabled nor provisioned by default -- there is no shipped QuantaStor dashboard set, and no data source is configured for you. If you want to build your own, the statistics database is available to Grafana as an InfluxDB 1.x data source: database quantastor, user quantastor, password from /etc/influxdb/quantastor.txt, and the retention policy names from Retention and downsampling above. Query the _10m and _2h series for anything longer than six hours, for the reasons given there.
The whole tab can be hidden per user through the Web Interface Customization settings on a user account, which is worth doing for operators who should not be following links out of the appliance.
Most of what QuantaStor alerts on around the dashboards is capacity, not performance. The pool free-space thresholds behind Pool Capacity Threshold Alerts are the ones administrators change most often; Alert Manager owns them, and Call-home / Alerting covers how the resulting notifications are delivered.
The performance-adjacent conditions that do raise an alert are all thermal or power:
| Condition | Notes |
|---|---|
| System, motherboard (PCH) or CPU temperature out of range | Detected from the IPMI sensor readings, at most one alert per condition per day. Which sensor names count as which reading is mapped in /opt/osnexus/quantastor/conf/qs_ipmi.conf, per platform.
|
| Power supply degraded, failed, or recovered | From the same IPMI sampling, using the power supply sensor names in the same file. |
| Fan health | From the same IPMI sampling. |
| Disk over temperature | Raised through the hardware controller adapters when a drive reaches the temperature threshold, 60 C by default. The threshold is drive_temp_alert_threshold in /opt/osnexus/quantastor/conf/qs_ipmi.conf; values below 50 C are ignored and clamped to 50, and values above 200 C are clamped to 200. At most one alert per day.
|
There is no alert on I/O latency, IOPS or throughput. Nothing in QuantaStor watches a performance number against a threshold, so a pool that has become slow will not tell you -- it has to be noticed on these dashboards, or by exporting the series to a monitoring system that can threshold them. That is the main argument for pointing an external dashboard at the statistics database on an appliance where performance matters.
When a chart is empty or looks wrong
In roughly the order worth checking:
- The range is longer than six hours. Beyond that the chart reads 10-minute or 2-hour averages, and short spikes are averaged away. See Which resolution each range reads.
- The chart is one that does not auto-refresh. Cache Stats (ARC) and Trends are only re-queried when you change the range or the selection; everything looks stale until you do.
- Cache Efficiency looks pinned at 100%. It is a since-boot average and it is meant to look like that. Difference the raw counters instead.
- The Trends axis disagrees with the range selector. Trends never draws less than six hours; the selector is not updated to match.
- Sensors and the tile temperature and power charts are empty. Expected on a virtual appliance and on hardware without IPMI/DCMI.
- A physical disk has no data. Device I/O is only collected for pool, volume and Ceph OSD/journal devices.
- There is a short gap in every chart at the same moment. Collection is paused while QuantaStor reconciles some configuration changes.
- Nothing at all, on every chart. Check the three services --
qs-statusforqs-statsd,systemctl is-active influxdb telegraffor the other two -- then confirm samples are arriving withqs-statsdb count cputwice a few seconds apart. - Nothing at all, after a recent restore or database maintenance. The retention policies and continuous queries are rebuilt at service start; restart the QuantaStor service and re-check.
If the numbers are present but you do not believe them, take the measurement a second way: qs-iostat reads the kernel directly, and Performance Testing covers generating a known load so there is something unambiguous to compare against.
Related pages
- Performance Testing -- proving where a bottleneck is, with
qs-perftestandqs-ramdisk - Performance Tuning -- I/O profiles, ARC and write log sizing, and the rest of the tunables
- Storage Pools -- pool layout, capacity and cache device configuration
- Storage System -- appliance and grid configuration, including grid membership
- Alert Manager -- the capacity thresholds behind the pool health chart, and the alert types
- Call-home / Alerting -- alert handlers, call-home and where notifications are delivered
- QuantaStor Shell Utilities --
qs-iostat,qs-statsdband the rest of the console tools - Network Shares -- share usage tracking behind the User Usage and Group Usage views
- QuantaStor CLI Command Reference -- full argument lists for every
qscommand
Verified against QuantaStor 6.9.0.