Scale-up HA Storage Pool Troubleshooting: Difference between revisions
New page: scale-up HA storage pool recovery scenarios, extracted from Template:TroubleshootingHAStoragePools and rewritten to house style; corrected the pool I/O check interval, hot-spare selection criteria and dialog names against 6.9.0 (QSTOR-12352) |
mNo edit summary |
||
| Line 419: | Line 419: | ||
== Security == | == Security == | ||
=== Ransomware === | === Long Term Snapshots for Ransomware Protection === | ||
The best defence is snapshots the attacker cannot reach. QuantaStor's snapshot schedules and remote-replication schedules both carry long-term retention counts, set independently: | The best defence is snapshots the attacker cannot reach. QuantaStor's snapshot schedules and remote-replication schedules both carry long-term retention counts, set independently: | ||
Latest revision as of 06:28, 4 September 2026
Recovery procedures for a scale-up (OpenZFS) high-availability storage pool: failed media, lost enclosures, lost appliances, pools that will not start, and rebuilding an appliance from scratch. The procedures apply to any scale-up HA cluster regardless of how the shared storage is attached, so they cover both the JBOD-attached build in HA Cluster Setup (JBODs) and the SAN-attached build in HA Cluster Setup (external SAN). Problems specific to building the cluster in the first place -- a group that will not form, a pool that will not import on a new build -- are on those pages.
Start from the symptom.
| What you are seeing | Go to |
|---|---|
| A drive is reporting SMART warnings, or a pool is DEGRADED with one device faulted | A failed or failing device |
| A pool is faulted and more devices are missing than the layout has parity for | More failed devices than the layout tolerates |
| Someone pulled the wrong drive | The wrong drive was pulled |
| A pool failed over for no obvious reason, with no device faults | Pool I/O failure |
| A pool failed over right after cabling work, and the old appliance will not take it back | Disconnected SAS cables |
| The pool name does not appear on the devices on both appliances | Disks not visible on both appliances |
| Path count halved, or a pool failed over when an enclosure module failed | Enclosure I/O controller failure |
| An enclosure lost power and the pool stopped | Unexpected power-off of a disk enclosure |
| Two or more enclosures have been off, at different times, and the pool will not start | Multiple enclosure failures |
| An appliance lost power; some clients recovered and some did not | Unexpected power-off of an appliance |
| An appliance rebooted ungracefully and its pool has not come back to it | Ungraceful appliance reboot |
| A repaired appliance is back but shows Warning, or will not accept a failover | Re-integrating an appliance after a hardware failure |
| Both nodes lost the heartbeat network | Top-of-rack switch failure |
| Both appliances rebooted and the pool did not start | An HA pool that will not start |
| A pool shows as missing after a boot, or is not on either appliance | A pool reported missing |
| A pool is suspected corrupted | A pool reported corrupted |
| Two appliances appear to be running the same pool | Split-brain |
| The Getting Started dialog appears at login and the license keys are gone | Boot media failure, software corruption and reinstall |
| Files have been encrypted by an attacker | Ransomware |
| You need to read the fencing state, or compare device visibility between appliances | Diagnostic tools |
Two rules that fix most of these
Two habits account for a large share of successful recoveries, and both are worth knowing before you need them.
Power the disk enclosures on first, and let them finish coming up before the appliances. A pool cannot import devices that were not present when the appliance booted, and SCSI-3 persistent reservations do not survive a power cycle of the drive: QuantaStor sets the Persist Through Power Loss bit to 0, so registrations and reservations are cleared when a device powers on. That is what makes a cold boot in the right order a genuine reset of the fencing state rather than a hopeful reboot.
Fail a pool over to the appliance it is already on. This is not the same as starting the pool. A manual failover re-runs the whole sequence -- pre-emptive I/O fencing of every device in the pool, the pool import, and placement of the HA virtual interfaces -- where a plain Start Storage Pool only imports. Nearly every "the pool is there but not serving" case is fixed by failing it over to itself.
From the command line, name the appliance the pool is already on as the target:
qs ha-group-failover --ha-group=pool-2-ha-group --storage-system=qs-node2 \ --import-health-checks=true --export-health-checks=true
qs ha-group-failover also takes --flags=dry-run-only, which reports what the failover would do without moving anything. Use it first when you are not sure the target appliance is ready.
Media and device failures
A failed or failing device
HDDs and many SSDs have a typical annual failure rate of around 1.5%, so replacing failed media is routine maintenance rather than an incident. We require parity-based pools to carry at least two parity devices per stripe -- RAIDZ2 or RAIDZ3 for scale-up, and K+M with M of 2 or more for erasure-coded scale-out. Single parity may be acceptable in a test environment but is not durable enough for production: with two or more parity devices a single failure leaves the pool degraded but still able to correct bad sectors and survive a second concurrent failure while it repairs.
A device on its way out is usually reported before it faults. QuantaStor raises a Disk Health Non-optimal warning alert, and the disk's own properties carry the reason:
State: Warning
State Detail: Check device SMART health information.
Smart Health Test: Detected high media defect rate '1281 defects' which indicates that
this device may fail soon.
A high media defect count points at the media itself. A large non-medium error count -- reported as OK - Detected 'N' non-medium errors -- points instead at the path: cabling, a backplane, or an expander. Treat the two differently; replacing a drive will not fix a cable.
To replace the device:
- Add the replacement, equal to or larger than the failed device, to the appliance that has the failure. Use a free slot if there is one; otherwise remove the failed device first and fit the replacement in the same slot.
- Mark the new device as a hot-spare. Once it is a spare, QuantaStor uses it to repair the degraded pool automatically.
- Remove the failed device once the pool has finished repairing.
That last step matters more than it looks. Failing media can misbehave on the bus -- resets and timeouts that affect every device behind the same expander -- and some enclosure models tolerate it much worse than others. Our recommendation is to pull bad media as soon as the pool is healthy again.
There are two ways to mark a spare, and they are not interchangeable.
A universal (global) hot-spare can repair any pool on the appliance. Use it when a system has several pools and you want a shared pool of spares.
The menu item is hidden when the disk cannot be a spare: when it is already in a pool, when it is already marked as a spare, or when it is mounted, which is how the appliance's own boot device is excluded. From the command line, qs disk-global-spare-add adds one or more devices to the global spare pool, and qs disk-spare-list lists what is currently a spare.
A dedicated spare belongs to one pool. Prefer this when you are explicitly replacing a known bad drive in a known pool.
The dialog is titled Add Hot-spare(s) / Recover Storage Pool and its own help text states the behaviour plainly: "Select one or more disk devices to be assigned to the pool for use as hot-spares. If the pool is in a degraded state it will auto-repair using one or more of the designated hot-spares automatically." From the command line that is qs pool-add-spare, with qs pool-remove-spare to take one back.
Repairs run one device at a time. When several devices are degraded, QuantaStor chooses a spare on these criteria, in this order of preference:
- Enclosure. A spare in the same enclosure as the failed device is preferred over a better-sized spare elsewhere, so that enclosure-level redundancy is preserved.
- Size. The smallest available device at or above the size of the failed one. A pool set to exact-match will refuse anything but the same size and raise an alert instead.
- Encryption. A dedicated spare must match the pool's encryption state. A global spare need not: global spares start unencrypted, and one selected for an encrypted pool is encrypted as it is brought in. A dedicated spare whose encryption state does not match the pool raises an alert and is skipped.
- Device class. A pool built entirely on HBA-attached devices will not take a virtual device surfaced by a RAID card as a spare, and vice versa. Mixed pools accept either.
Media type is not among the criteria: nothing in the selection path distinguishes flash from spinning disk, so a universal SSD spare that clears the size floor is eligible to repair an HDD pool. Keep flash spares dedicated to flash pools if that matters to you. This is tracked as a defect.
How aggressively QuantaStor repairs is set per pool.
| Setting | Behaviour |
|---|---|
| Autoselect best match from assigned or universal spares | Use a pool-assigned spare if there is one, otherwise a universal spare. Size is a floor, not an exact match. |
| Autoselect best match from pool assigned spares only | Never touch the universal spares. If the pool has none of its own, QuantaStor raises an alert and waits. |
| Autoselect exact match from assigned or universal spares | As above but the spare must be exactly the same size as the failed device. |
| Autoselect exact match from pool assigned spares only | Both restrictions together. |
| Manual Hot-Spare Management Only | No automatic repair at all. Nothing happens until you act. |
A repair already in progress is allowed to finish: while a pool is resilvering, QuantaStor does not start another repair on it.
To find the physical drive, turn on its LED.
qs disk-identify takes --blink-type=on or off, an optional --duration in seconds, and a --pattern built from short and long pulses and delays, for example --pattern=pppD. With no duration the LED stays on until you turn it off. qs hw-enclosure-slot-identify lights a slot rather than a device, which is what you want when the drive is already gone. Slot numbering and enclosure layouts are covered on Hardware Controllers & Enclosures; if the enclosure layout has not been configured for your hardware, the slot QuantaStor reports may not be the slot on the front of the chassis, and it is worth getting support to confirm the device before anyone pulls anything.
More failed devices than the layout tolerates
An 8d+2p pool that loses three devices in the same VDEV has exceeded its parity and cannot repair itself. In many cases the pool is lost at that point. Things worth trying, roughly in order of how much they risk:
- Cold power-cycle everything. Leave it all off for a minute, then power the enclosures on first and let them settle before the appliances. Devices that dropped off the bus are often back, and the fencing state is cleared by the power cycle.
- Ask support about importing the pool with a rollback to the last good transaction group. This is an OpenZFS import with non-default options; it discards the most recent writes and it is not something to attempt unassisted.
- Copy failing media to good media with a block-level tool that tolerates read errors, then try the import again.
- If the devices are passed through a hardware RAID controller -- which we do not recommend for scale-up pools -- bring the media back first with Import Hardware RAID Units and Mark Good, then retry.
Those are qs hw-controller-import-units and qs hw-disk-mark-good respectively.
(Note: the import-with-rollback and block-copy steps are longstanding support guidance rather than product operations, and they could not be exercised for this page. Involve support before running either -- both can make an otherwise recoverable pool unrecoverable.)
Avoiding this position is cheap and reliable: use double parity or better, scrub quarterly with qs pool-scrub-start or a maintenance schedule, and configure alerting so that a failed device is noticed. Losing a pool is rare, and when it happens it is almost always the end of a long stretch in which devices failed and no spare was ever supplied.
The wrong drive was pulled
If a device was removed by mistake and the pool is still running, put it straight back in. The pool recovers it automatically, and there is a delay of about a minute before a hot-spare is injected, so a quick correction costs nothing.
If enough devices were removed that the pool faulted, identify which device actually needed to go, put the others back, and reboot the appliance -- a faulted pool cannot be brought back online without one. If the removals triggered a failover and the pool could not be imported on the other side either, both appliances need a reboot. Use Identify Disk and the enclosure layout to be certain which device you are pulling next time; if the enclosure layout is not configured for the hardware, or the identification is at all ambiguous, get support to confirm before pulling anything.
A pool reported corrupted
Corruption of an HA pool is very unlikely. Fencing means only one appliance can write to the devices at a time, and the failover sequence is built around that guarantee. Unless someone has been formatting media or clearing reservations by hand at the console while the pool was running -- qs-iofence has subcommands that will do exactly that -- the pool is probably not corrupt and is failing to start for some other reason.
Power everything off, then cold boot with the enclosures first and the appliances after. The pool normally starts on its own and restores access to every volume and share. If it does not, stop and get support involved rather than trying import variants: the common causes at this point are hardware -- cables, backplanes, an expander, media causing bus resets -- and they are almost always fixable.
Connectivity
Pool I/O failure
QuantaStor checks that each local ZFS pool is genuinely writable, independently of whether clients are doing any I/O. Every two minutes -- and immediately when something else in the service asks for a check -- it writes 1 MiB of random data to a file named .qserrchk.N in the pool's mount directory, with a 10 second timeout. Random data is used rather than zeroes so that no caching or compression layer can satisfy the write without touching the pool.
If that write times out, QuantaStor reads the pool's health before deciding anything:
- ONLINE or DEGRADED -- no action. A slow pool under heavy load is not a failover trigger.
- FAULTED, OFFLINE, UNAVAIL or REMOVED -- the pool is evicted from this appliance and failed over immediately.
- Anything else -- one retry after a one second pause, then the same decision.
If the write is still timing out but the pool reads healthy, QuantaStor raises a Pool Writability Test Failed warning noting that the pool may be experiencing slow I/O due to heavy load. That is a performance signal, not a failover.
Eviction sets a reboot requirement on the appliance, because a pool that has reached the FAILED state cannot be re-activated there until the I/O stack has been cleared. You can see this in the Storage System's properties, or with qs system-get:
Requires Reboot: false
When it reads true, the state detail names the pool and the appliance, and no amount of retrying will bring that pool back to that appliance until it has rebooted.
A failover also cannot proceed if the pool's HA group has no virtual interfaces attached to it. The service logs this as a blocked failover, and it is the same underlying condition as An HA pool that will not start.
Disconnected SAS cables
Pulling the shared-storage cables from an appliance is one way to trigger a failover, and it behaves the way the design intends: the paired appliance takes the pool over, typically within 15 to 20 seconds.
The appliance that lost the cables then needs a reboot. This is not optional and it is not conservatism -- writes that never completed are still in its I/O stack, and the FAILED pool state cannot be cleared any other way. Fix the cabling, reboot it, let it rejoin the grid, and check its state detail before you fail anything back to it.
Disks not visible on both appliances
An HA pool requires every device to be visible to every appliance in the group, and the Physical Disks section shows the pool name against a device on both appliances when that is true. If the pool name is missing on one side, work through this list:
- Verify physical connectivity from both appliances to the shared enclosure or back-end array.
- Verify the cables are properly seated and that the link lights are active.
- Verify that every device in the pool is dual-ported and supports persistent reservations. Dual-port SAS, dual-port NL-SAS and dual-port NVMe are supported for scale-up HA, as are FC and iSCSI LUNs from an external array.
- Verify each enclosure has at least two expansion ports and each appliance has two HBAs. Enclosures with three or more expansion ports and redundant expander or ESM modules are preferable.
- Verify SAS cable lengths are within the standard 5 metre limit for copper. Use optical SAS cables for longer runs.
- Compare the two appliances device by device with
qs-util devicemapand check the reservations withqs-iofence devstatus. See Diagnostic tools.
A device that shows up as single-pathed when it should be dual-pathed is a multipath problem rather than a visibility problem; see Multipath Configuration. Device naming, serial numbers and the pool assignment are covered on Physical Disks/Devices.
Enclosure I/O controller failure
Appliances are normally cabled to each enclosure over two paths from two separate HBAs. When one enclosure I/O controller fails in that configuration the pool keeps running and the path count drops from two to one. Throughput can suffer, most noticeably with flash enclosures.
Where connectivity to the enclosure is single-pathed, losing that controller takes the devices away entirely. The pool then fails the writability check, reads as faulted, and is failed over to the paired appliance.
Enclosure and appliance power loss
Unexpected power-off of a disk enclosure
When you create a pool, QuantaStor tries to stripe across enclosures so that no single enclosure outage can stop it. Achieving that takes enough enclosures for the layout: RAIDZ2 4d+2p needs three, RAIDZ3 8d+3p needs four. Many deployments run a single enclosure and have no enclosure redundancy at all, which is a legitimate choice as long as it is a deliberate one.
Without enclosure redundancy, losing the enclosure takes away too many devices, the pool stops, and the attempted failover fails as well because the other appliance cannot see the devices either. The pool ends up in an error or missing state on both sides.
To recover, power everything off, then power the enclosures on and let them come up, then power the appliances on. The pool starts and recovers on its own. The appliances need the reboot to clear the I/O stack and the failed pool state; there is no way to skip it. An appliance that never imported the failed pool -- and that is hosting an unrelated healthy pool of its own -- can reasonably be left running, but where nothing else is at stake, taking everything down and bringing it up in order is the simpler path.
Multiple enclosure failures
If more enclosures are lost than the layout tolerates, the pool stops. Power the enclosures back on and reboot the appliance that was hosting the pool. If several pools were affected, cold boot the whole cluster.
A pool with an enclosure-redundant layout -- RAIDZ2 4d+2p spanning three enclosures, say -- keeps running through the loss of any one of them. Once the enclosure is back, fail the pool over to the appliance it is already on to force a revalidation of every path. QuantaStor re-fences media after an enclosure restart so that the devices are locked to the appliance running the pool, and you can confirm that in the Physical Disks section: a black underline under a device icon means the device is correctly fenced to the appliance hosting the pool, and a red underline means it is not. A red underline clears with a failover, including a failover to the appliance the pool is already on.
The case that catches people out is two enclosure outages at different times. Take a pool across enclosures A, B and C that can run degraded on any two. Enclosure C goes off; the pool runs on A and B and starts accumulating writes they have and C does not. C comes back, and before the resilver finishes, B goes off. The pool stops -- and it is right to. A and C are at different points in the transaction history and OpenZFS will not assemble a pool out of them.
The fix is to bring all three back, let C resilver to completion, and only then take B down if you still need to. Going straight from a degraded A+B to a degraded A+C does not work and cannot be forced.
Unexpected power-off of an appliance
When an appliance is lost, the pool's virtual interfaces move with it to the surviving appliance, and access is restored for every protocol -- NFS, SMB, iSCSI, FC and NVMe-oF.
If some clients recover and others do not, the ones that did not are almost certainly connected to a local address on the appliance rather than to the pool's HA virtual interface. Only the virtual interface moves. This rarely bites block clients, because iSCSI access is restricted to the HA virtual interfaces, so there is no wrong address to connect to. It bites NFS and SMB clients regularly, because nothing stops someone mounting a share over a management address.
To find them, look at what the clients are actually connected to:
The Client/Host Connection View lists the live TCP connections to the selected appliance by protocol, so you can see which interface each client came in on. Connections QuantaStor itself made outward to a host are listed separately as outbound.
The neighbouring Network Connectivity/Ping Checker... item is a different tool and worth not confusing with it: that one checks reachability from an appliance to the other appliances and to connected clients, including whether jumbo frames survive the path, and it accepts extra addresses to test. Use the connection view to answer "what are my clients using", and the ping checker to answer "can this appliance reach them at all".
Ungraceful appliance reboot
Wait for the reboot to finish. The appliance resyncs with the grid and is then ready to accept failovers again.
Pools do not fail back on their own. If you run two pools, one on each appliance, you have to move one back by hand once the rebooted appliance is healthy.
Re-integrating an appliance after a hardware failure
Power it on. It rejoins the storage grid and runs its own checks on whether it is fit to accept a failover.
Then look at its health in the Storage Systems section. If it shows Warning or Error, the reason is in State Detail in the properties panel, and it is specific: an appliance that is up but whose shared-storage cabling is still disconnected reports that it is not ready to accept a failover because it cannot reach the media. Fix the underlying problem and the state clears by itself -- there is nothing to acknowledge.
Before you move a pool back, ask the cluster directly whether it can take it:
qs ha-group-get-health-status --pool-list=pool-2
qs ha-group-get-health-status reports the stored failover health events for a pool, per appliance, each with a severity, a description and a recommended action -- for example that an HA group has no client IPs configured. qs ha-group-check-health-status re-runs the checks rather than reporting the stored results, and qs pool-health-check walks the pool and its devices looking for problems anywhere in the stack.
Top-of-rack switch failure
The heartbeat between the appliances of an HA pair is maintained by the Site Cluster, which is normally configured with two cluster rings so that losing one switch does not cost you the heartbeat. Where spare 1GbE ports are available, we recommend making the second ring a direct connection between the two appliances with a crossover cable: a heartbeat that does not traverse a switch survives every switch failure there is.
Confirm both rings are live from either appliance:
# corosync-cfgtool -s Local node ID 49997, transport knet LINK ID 0 udp addr = 10.0.42.23 status: nodeid: 49997: localhost nodeid: 50255: connected LINK ID 1 udp addr = 10.25.42.23 status: nodeid: 49997: localhost nodeid: 50255: connected
Both links should list the peer as connected. If only one does, the cluster is running on a single ring and the next switch event takes the heartbeat with it. Ring configuration is on Site Cluster Setup.
Underneath, QuantaStor uses corosync and pacemaker with its own resource agent for the virtual interfaces. The agent is installed as IPaddrQS and also over the distribution's IPaddr2, with the original kept as IPaddr2.backup, so the ocf:heartbeat:IPaddr2 resource you see in pcs status is QuantaStor's version. It does more than move an address: it flushes the ARP caches on failover, refuses to bring a resource up when the QuantaStor service is stopped, and honours the cluster standby state. You can read the resource group for each HA group directly:
# pcs status
Cluster Summary:
* Stack: corosync (Pacemaker is running)
* Current DC: qs-node1 (version 2.1.6-6fdc9deea29) - partition with quorum
* 2 nodes configured
* 3 resource instances configured
Node List:
* Online: [ qs-node1 qs-node2 ]
Full List of Resources:
* Resource Group: pool-2-ha-group-d786e0b7:
* had70 (ocf:heartbeat:IPaddr2): Started qs-node2
* Clone Set: dlm-clone [dlm]:
* Started: [ qs-node1 qs-node2 ]
The resource group is named for the HA group and the pool, and the resource itself is named for the virtual interface's tag. The appliance named against the resource is the one hosting the pool.
Pools that will not start
A pool reported missing
If the appliances booted while the enclosures were off, nothing will import, because the devices were not there at boot. Powering the enclosures on afterwards is necessary but not sufficient -- the pool still needs to be told to start, and it needs its fencing and virtual interfaces placed.
Find out which appliance owns the pool. In the Storage Pools section the tree label carries it: pool-2 (on: qs-node2). Then run a manual failover of that pool from that appliance to that same appliance.
Do not reach for qs pool-clear-missing here despite the name -- it clears missing network shares and storage volumes from a pool, and has nothing to do with a missing pool.
An HA pool that will not start
Both appliances have rebooted, the devices are all visible, and the pool has not come up. In order of likelihood:
- The HA group is deactivated. Check its state first. Automatic failover is off, so nothing placed the pool.
- The HA group has no virtual interfaces. A failover cannot complete without at least one, and the service logs the failover as blocked.
- The virtual interface addresses are in use elsewhere on the network. The agent will not bring up an address another host is answering for, so the resource group never starts and the pool never gets placed.
If the group is deactivated on purpose, you do not have to activate it to get the pool running -- fail the pool over to the appliance it is already on, which runs the fencing and startup sequence without changing the group's activation state. Where you do want automatic failover back, that is qs ha-group-activate; see Activating and deactivating the group.
Split-brain
Split-brain is a state in which both appliances are running half of a mirrored pool at once -- one RAID10 pool operating as two independent RAID0 pools with the same name on two appliances. It is very hard to reach, and this section is mostly about why.
Three separate mechanisms have to be defeated:
- Fencing. Every device in a pool carries a SCSI-3 persistent reservation, so only the appliance holding it can write. To get around this you would have to isolate each enclosure to a different appliance by re-cabling, so that neither appliance's reservations can reach the other's devices.
- Virtual interface ownership. A pool only runs on the appliance holding the HA virtual interfaces, which move as a single cluster resource group.
- Device count. A pool only starts where enough devices are present for it to run, even degraded.
The last of those is why parity layouts are effectively immune: a RAIDZ2 or RAIDZ3 pool cannot start on half its devices, so two divergent copies cannot both exist. The exceptions are the narrow layouts where half the devices is still enough -- mirrored (RAID10), 2d+2p RAIDZ2 and 3d+3p RAIDZ3.
So the only realistic route in is a mirrored or borderline-parity pool plus a deliberate re-cabling that isolates enclosures to separate appliances. To rule it out entirely:
- With RAID10 across two enclosures, cable two paths from both appliances to both enclosures, so fencing always reaches every device.
- Or use one enclosure, or more than two.
- Or avoid RAID10, 2d+2p and 3d+3p altogether. For virtualization and databases we generally recommend RAIDZ2 4d+2p over RAID10: more usable capacity, better durability, and no split-brain exposure at all.
If it has happened and both halves have taken writes, recovery is not a merge. The two pools have diverged, and the procedure is to decide which one is authoritative, reformat the devices of the other, and resilver them back into the good pool. That is time-consuming and I/O-intensive, and it loses whatever was written to the discarded half.
(Note: this recovery has not been exercised for this page -- inducing a split-brain requires physically re-cabling a live cluster. The prevention mechanisms above are verified against the product; the recovery procedure is longstanding support guidance. Contact support before starting: choosing the wrong half is not reversible. The fencing mechanics are documented under I/O fencing and SCSI-3 persistent reservations.)
Rebuilding an appliance
Boot media failure, software corruption and reinstall
The most common cause of software corruption is the boot device filling up, which can corrupt the local grid database. You find out immediately: with no license records in the database, the Getting Started dialog opens at login, because that dialog is shown whenever the license list is empty. For this reason we recommend a 480GB or larger SSD or NVMe device as the boot device.
QuantaStor watches the boot device and alerts before it becomes a problem, at three thresholds on free space -- 25% remaining raises an alert, 15% a warning, and 8% a critical alert, each repeating at its own interval. Treat the first one as the actionable signal rather than waiting for the third.
None of this endangers the storage pools, and it is all fixable.
- If the boot device is simply full, clear space -- oversized or excess files under
/var/logare the usual culprit -- and restart the service. - On a service restart QuantaStor runs an integrity check on the configuration database and attempts a repair if it fails. Success and failure are both alerted; the failure alert says to contact support, and it means it.
- If the boot device has failed, replace it and reinstall QuantaStor. Install the current release rather than matching the old one.
The configuration database is backed up onto the storage pools, which is what makes a reinstall recoverable rather than a rebuild from memory. Each pool carries a .qsbackups directory holding rotated hourly and daily copies of the database, plus the Samba databases that hold local SMB configuration:
# ls /mnt/storage-pools/qs-<pool-uuid>/.qsbackups/ osn.db.backup.daily.1 osn.db.backup.hourly.1 samba-tdb.tar.gz.daily.1 osn.db.backup.daily.2 osn.db.backup.hourly.2 samba-tdb.tar.gz.daily.2 ...
Five database copies are kept in each rotation by default -- the setting caps at 100 -- and ten of the Samba archive. Nothing is written for the first hour after a boot, and a pool with an active task is skipped for that pass. These are the recovery points the Recovery Manager offers after a reinstall -- see Recovery Manager for the restore itself, and qs system-metadata-recovery-point-list to list them from the command line.
The pool carries its own metadata too. Every pool holds a .poolmeta directory of timestamped archives describing its volumes, shares, hosts, host groups, volume ACLs and share settings, and an import recovers them by default:
# ls /mnt/storage-pools/qs-<pool-uuid>/.poolmeta/ poolmeta_d786e0b7_GMT20260717_160056.tgz poolmeta_d786e0b7_GMT20260719_122437.tgz ...
So on a single-appliance system, re-importing the pool is most of the recovery.
The dialog is titled Import Storage Pools; press Scan to list what is importable. qs pool-import does the same and carries two options worth knowing: --metadata-only re-imports the metadata for a pool that is already imported, and --skip-metadata imports volumes and shares without recovering the metadata, which is what you want if the newest archive is too old or invalid. qs pool-preimport-scan lists pools that are available but not yet discovered.
Encrypted pools need their keys. The pool encryption keys live on the boot media, AES key-wrapped, so export them after the pool is created and keep the file somewhere else.
That is qs pool-key-export, with qs pool-key-import to put them back and qs pool-import-encrypted to import the pool with its passphrase. As long as the paired appliance still has healthy boot media, or you hold an exported key file, reinstalling costs you nothing.
Rejoining an HA pair takes a few extra steps, because the site cluster and the HA virtual interfaces have to be rebuilt:
- On the healthy appliance, convert the pool's HA virtual interfaces to local virtual IPs. qs ha-interface-delete takes
--convert-to-vif=trueto do this rather than simply dropping the address, so the pool keeps serving on the same IP throughout. - Reinstall the failed appliance and add it back to the grid.
- Recreate the Site Cluster across both appliances. See Site Cluster Setup.
- Convert the local virtual IPs back into HA virtual interfaces with qs ha-interface-create. See The HA virtual interface.
Security
Long Term Snapshots for Ransomware Protection
The best defence is snapshots the attacker cannot reach. QuantaStor's snapshot schedules and remote-replication schedules both carry long-term retention counts, set independently:
| Retention | Effect |
|---|---|
--rc-hourly |
Keeps N snapshots spanning the last N hours |
--rc-dailies |
Keeps N snapshots at least a day apart |
--rc-weeklies |
Keeps N snapshots at least 7 days apart |
--rc-monthlies |
Keeps N snapshots at least 30 days apart |
--rc-quarterlies |
Keeps N snapshots at least 90 days apart |
We recommend at least two quarterly snapshots, which gives you a roll-back point up to six months old without going to tape. Snapshots only consume capacity when they are holding on to data that has since been deleted or overwritten, so if you rarely delete files, keeping more quarterlies costs almost nothing. If you delete or rewrite data constantly, size the system for the snapshots you intend to keep. See qs snap-schedule-create.
Network shares can also enforce immutability directly, so that files cannot be modified or deleted for a period after they are written -- a stronger guarantee than a snapshot, because it applies to the live share. qs share-create takes --immutability-mode of immutable or append, and --days-of-immutability to set how long. The dialog exposes the same under Immutability / Write-Once-Read-Many (WORM) Support. A file's age is checked once a day at midnight, and it becomes mutable once it is older than the period; a period of 0 means immutable forever.
Beyond retention, keep interactive root access off the appliances. We recommend that all low-level access go through the qadmin sudo account, that qadmin require an SSH authorized key rather than a password, and that the qadmin password be reset to something strong on every new system. qs-util checkpass reports whether the default qadmin password is still in place, which is worth checking on anything that has been deployed from an image.
Diagnostic tools
qs-iofence devstatus
qs-iofence reads and manipulates the SCSI-3 persistent reservations that fence HA pool devices. Only devstatus is read-only. Every other subcommand -- devclearall, devreleaseall, devscrub, devrelease, devpreempt, devregister, devreserve, devreregister -- changes reservation state on live media, and running one against a pool that is imported is a good way to create the corruption you were investigating. Use devstatus freely; leave the rest to support.
# qs-iofence devstatus /dev/sda NOT-SUPPORTED (NOT-SUPPORTED) [NOT-SUPPORTED] <> /dev/sdl 6SE37R6E0000B135J3TG () [] <> /dev/sdm 6XP2F4JX0000B2359Y4A (0xff753100ffd786e0) [0xff753100ffd786e0] <WERO> /dev/sdn 6SE46WWK0000B146RNKP (0xff753100ffd786e0) [0xff753100ffd786e0] <WERO> /dev/sdo PFVVJUYE (0xff753100ffd786e0) [0xff753100ffd786e0] <WERO> /dev/sdp 6SE31S1V0000B130JCEV (0xff753100ffd786e0) [0xff753100ffd786e0] <WERO>
The columns are the device, its serial number, the registered keys in parentheses, the key holding the reservation in brackets, and the reservation type in angle brackets. Three forms appear above and all three are normal:
- A device in a running HA pool carries a WERO reservation -- write exclusive, registrants only. All four pool devices here hold the same key, which is how it should look.
- A device not in an HA pool has no keys and no reservation. Empty is correct, not broken.
- NOT-SUPPORTED means the device does not implement persistent reservations at all. The appliance's own boot device is the usual example, and it is expected. A device you intended to put in an HA pool reporting NOT-SUPPORTED is a hardware selection problem: see Disks not visible on both appliances.
Reservations live on the drive, so both appliances see the same output. The key encodes who holds it and for what, which means you can tell at a glance whether fencing followed a failover; the format and how to read it are on I/O fencing and SCSI-3 persistent reservations. qs-iofence devstatus also accepts a comma-separated list of serial numbers when you only care about a few devices.
This is the command to reach for when a pool will not import and the system log carries scsi reservation conflict messages.
qs-util devicemap
qs-util devicemap lists every device with its /dev/disk/by-id path, vendor, model and serial number. Run it on both appliances and sort the output to compare what each appliance sees:
qs-node1:~# qs-util devicemap | sort ... /dev/dm-5 /dev/disk/by-id/scsi-35000c5004baa9e5f, HP, EG0300FBLSE, 6XP2F4JX0000B2359Y4A, /dev/dm-6 /dev/disk/by-id/scsi-35000c5003af293d3, HP, EG0300FAWHV, 6SE46WWK0000B146RNKP, /dev/dm-7 /dev/disk/by-id/scsi-35000cca015304368, HP, DG0300FARVV, PFVVJUYE, /dev/dm-8 /dev/disk/by-id/scsi-35000c500339a864f, HP, EG0300FAWHV, 6SE31S1V0000B130JCEV, /dev/sdm /dev/disk/by-id/scsi-35000c5004baa9e5f, HP, EG0300FBLSE, 6XP2F4JX0000B2359Y4A, /dev/sdn /dev/disk/by-id/scsi-35000c5003af293d3, HP, EG0300FAWHV, 6SE46WWK0000B146RNKP, ...
The /dev/sdX letters will differ between appliances and that does not matter. The by-id paths and the serial numbers are what must match. SAS enclosures often produce consistent sdX assignment, but FC and iSCSI attached storage generally does not, because arrays vary in how they answer bus probes.
The output also lists the multipath and encryption mapper devices, so the same physical drive appears more than once -- once as each SCSI path and once as the dm- device layered over it. On a pool built from encrypted devices the mapper names appear as dm-name-enc-dm-uuid-mpath-<wwn>, and that is the name the pool actually uses.
qs disk-list presents the same information with the pool each disk belongs to, which is usually the faster answer to "does this appliance see the pool's devices".
qs-util zpooldisks
qs-util zpooldisks prints the pool topology with the enclosure, slot, serial number, make and model annotated under each device. It answers "which physical drive is that faulted device" in one step, without cross-referencing anything:
# qs-util zpooldisks pool: qs-<pool-uuid> state: ONLINE config: NAME STATE READ WRITE CKSUM qs-<pool-uuid> ONLINE 0 0 0 raidz2-0 ONLINE 0 0 0 /dev/disk/by-id/dm-name-enc-dm-uuid-mpath-35000c5004baa9e5f ONLINE 0 0 0 `-> Enclosure,Slot: ; Serial#: 6XP2F4JX0000B2359Y4A; Make,Model: HP,EG0300FBLSE /dev/disk/by-id/dm-name-enc-dm-uuid-mpath-35000c5003af293d3 ONLINE 0 0 0 `-> Enclosure,Slot: ; Serial#: 6SE46WWK0000B146RNKP; Make,Model: HP,EG0300FAWHV errors: No known data errors
The pool is named qs- followed by the QuantaStor storage pool UUID, which is also how it appears in zpool output and in the pool's mount path under /mnt/storage-pools. Enclosure and slot are blank where the enclosure does not report them or no enclosure layout has been configured for the hardware; see Hardware Controllers & Enclosures.
Reading failover readiness
Three commands answer three different questions, and mixing them up wastes time:
| Command | Answers |
|---|---|
| qs ha-group-get-health-status | What has the cluster already recorded about this pool's ability to fail over, per appliance |
| qs ha-group-check-health-status | Re-run those checks now and report the result |
| qs pool-health-check | Walk this pool and its devices looking for problems anywhere in the stack |
For the state of the group itself and its interfaces, qs ha-group-get reports the active appliance, the settle time, the export timeout, the failover policies and the serial numbers of every fenced device. qs ha-group-list is the one-line-per-group summary.
Related pages
- HA Cluster Setup (JBODs) -- building a scale-up HA cluster with shared SAS enclosures
- HA Cluster Setup (external SAN) -- the same design with an FC or iSCSI array behind it
- Site Cluster Setup -- heartbeat rings, cluster membership, standby and maintenance mode
- Storage Pools -- creating, modifying and scrubbing the pool itself
- Physical Disks/Devices -- device naming, serial numbers and pool assignment
- Multipath Configuration -- multipath naming, path counts and the single-ported media alert
- Hardware Controllers & Enclosures -- HBAs, enclosures, slot numbering and drive identify
- Recovery Manager -- restoring the configuration database from a pool backup
- QuantaStor Touch Files -- the diagnostic and override touch files, including the HA and fencing group
Verified against QuantaStor 6.9.0.