HA Cluster Setup (JBODs)
A scale-up high-availability cluster is a pair (or trio) of QuantaStor appliances cabled to the same SAS JBOD, running a ZFS storage pool that either appliance can import. If the appliance hosting the pool fails, the pool, its virtual IP address, and every share and volume on it move to a surviving appliance automatically. This page covers the JBOD-attached build: hardware, cabling, the HA group and its virtual interface, failover, and I/O fencing. For the same design with a third-party SAN behind it instead of a JBOD, see HA Cluster Setup (external SAN).
| Section | Purpose |
|---|---|
| What a scale-up HA cluster is | The design, and when to choose it over scale-out |
| Hardware requirements | HBAs, media, and what "shared" has to mean |
| Cabling | Rules and per-enclosure-count diagrams |
| Before you build the HA group | Grid, site cluster, network, and the pool |
| Creating the storage pool HA group | The dialog, every field, and what the create checks reject |
| The HA virtual interface | The address clients use, and how it is named |
| Activating and deactivating the group | Turning automatic failover on and off |
| Failover | Deliberate and triggered, what happens in what order, how long it takes |
| I/O fencing and SCSI-3 persistent reservations | How QuantaStor guarantees one owner, and how to read the reservations |
| Maintenance | Standby and maintenance mode |
| Troubleshooting | A pool that will not import, a group that will not form |
| Command line reference | Every HA group command |
What a scale-up HA cluster is

QuantaStor offers two independent routes to high availability, and they are not variations on one another.
Scale-up puts the storage outside the appliances, in a shared SAS JBOD, and makes the appliances interchangeable front ends for it. A ZFS storage pool is built from disks in the enclosure and imported by exactly one appliance at a time. Redundancy against disk failure comes from the pool's RAID layout; redundancy against appliance failure comes from the second appliance being able to import the same pool. This is the right choice when you want the capacity efficiency and the feature set of ZFS -- compression, snapshots, RAIDZ2 and RAIDZ3 -- with node-level availability on top, and when the whole cluster fits in one rack.
Scale-out replicates data across independent appliances with no shared enclosure, and is documented under Scale-out Block Setup (ceph), Scale-out File Setup (ceph) and Scale-out Object Setup (ceph). Choose it when you need to grow past what one JBOD chain can hold, or when you cannot rely on a shared enclosure.
The rest of this page is about the scale-up case.
Three QuantaStor objects make it work, and they are created in this order:
- A site cluster -- the corosync and pacemaker layer that carries the heartbeat between appliances and decides when one has gone away. One site cluster serves any number of pools. See Site Cluster Setup.
- A storage pool HA group -- the object that ties one ZFS pool to the appliances allowed to import it, holds the failover policies, and is the thing you act on to fail a pool over.
- One or more HA virtual interfaces -- the IP addresses clients connect to, which move with the pool.
Hardware requirements

Every disk in the pool must be in the shared enclosure. That includes cache, log and hot-spare devices. A pool that draws even one device from a disk internal to one of the appliances cannot be made highly available, because the surviving appliance would be unable to import it. QuantaStor enforces this: creating the HA group verifies that every one of the pool's devices is visible on the secondary appliance, and refuses the create if any are missing.
Connect the enclosure with HBAs, not RAID controllers. Scale-up pools are HBA-only. ZFS needs unmediated access to the drives, and I/O fencing needs to issue SCSI-3 persistent reservation commands straight to them, which a RAID controller presenting virtual drives does not permit. Use Broadcom 9300, 9400 or 9500 series HBAs or the OEM equivalents; HPE servers attached to HPE enclosures should use the HPE OEM HBAs. A hardware RAID controller is still the right choice for the mirrored boot devices, which are internal to each appliance and are not part of any pool.
Pool media must be dual-ported and must support persistent reservations. In practice that means dual-port SAS, NL-SAS or dual-port NVMe. SATA drives are not supported for scale-up HA pools and QuantaStor rejects them by name at HA group creation:
Storage Pool contains one or more SATA disk devices including '<device>'. Storage Pool High-Availability feature requires that all disks devices be dual-port SAS, dual-port NVMe, or FC devices.
Single-ported media in a shared enclosure is a subtler failure, because the pool will build and run. QuantaStor detects it and raises a multipath configuration problem alert against the pool; Multipath Configuration documents that alert and why it is a correctness problem rather than a warning to acknowledge.
What "dual-connected" actually requires. Each disk's two SAS ports are wired to different places, and what you wire them to decides what survives. In the common two-appliance, one-JBOD build, each disk sends one port to each appliance: the sharing is at the appliance level, each appliance sees a single path to each disk, and a disk showing one path in multipath -ll is expected rather than a fault. Path-level redundancy on top of that needs a second HBA connection from each appliance into a second SAS expander in the enclosure, which is why enclosures with dual expanders are the ones to buy. Multipath Configuration covers how QuantaStor names and stacks these devices.
Minimum configuration
- 2x QuantaStor appliances acting as storage pool controllers
- 1x or more SAS JBOD cabled to both appliances
- 2x to 100x dual-port SAS HDDs or SSDs for pool storage, all of them in the shared enclosure
- 1x hardware RAID controller per appliance for mirrored boot devices
- 2x boot devices per appliance; 480GB or larger SSD or NVMe media is recommended, because a boot device that fills up is the most common cause of local database corruption
Storage bridge bay systems
A cluster-in-a-box or SBB (storage bridge bay) chassis holds two hot-swap servers and a JBOD in a single 2U unit, and QuantaStor supports the SuperMicro SBB models. Everything on this page applies unchanged; the cabling is internal to the chassis. Contact your OSNEXUS reseller or sdr@osnexus.com for hardware options.
Cabling
Incorrect cabling is the single most common cause of I/O fencing and performance problems in a scale-up cluster, and the symptoms rarely point at the cables. Check the wiring against the diagrams before building anything.
- The same rules apply to every enclosure model from every manufacturer.
- All Dell, HPE, WD/HGST, Seagate and Lenovo JBOD units ship with dual SAS expanders.
- SuperMicro sells single-expander enclosures (E1/E1C models) which should not be used for HA. Buy the dual-expander models, which carry E2/E2C in the model number.
- Some enclosures label their ports IN and OUT but accept either; check the vendor documentation.
- Avoid cascading JBODs. It adds failure modes for no benefit.
- Keep SAS cables within the 5 metre standard limit, or use optical SAS cables for longer runs.
Three rules govern which HBA port goes where:
- Never connect the same HBA to the same disk enclosure twice.
- Never connect the same SAS expander to the same HBA twice.
- Every appliance connected to eight or fewer enclosures must be connected to each enclosure twice, using different HBAs.
Unused HBA ports are fine and can be used to attach more enclosures later.
Scale-up cabling diagrams
2x appliances, 1x disk chassis

2x appliances, 2x disk chassis

2x appliances, 3x disk chassis

2x appliances, 4x disk chassis

2x appliances, 5x disk chassis

2x appliances, 6x disk chassis

2x appliances, 7x disk chassis

2x appliances, 8x disk chassis

2x appliances, 9x disk chassis

2x appliances, 10x disk chassis

2x appliances, 11x disk chassis

2x appliances, 12x disk chassis

Single-appliance expansion cabling
These layouts attach expansion chassis to one appliance. They are not HA configurations -- there is no second appliance to fail over to -- and are included here because the cabling rules are the same.
1x appliance, 1x expansion chassis

1x appliance, 2x expansion chassis

1x appliance, 3x expansion chassis

Before you build the HA group
Four things have to be in place first, in this order.
Licence both appliances. Each appliance needs its own unique Gold, Platinum or Cloud Edition key. See License Management.
Put both appliances in one storage grid. An HA group can only span grid members. Grid creation takes under a minute; see Grid Configuration.
Configure the networks. Set static addresses on every port you intend to use, set DNS and NTP, and keep the heartbeat networks separate from client traffic. The site cluster requires that the port carrying a heartbeat ring has the same interface name on every appliance, so plan the naming before you cable. Network Ports covers port configuration; Site Cluster Setup covers the heartbeat requirements in detail.
Create the site cluster, with two heartbeat rings. The site cluster is what detects an appliance going away. A single-ring cluster is fragile enough that ordinary network maintenance can trigger a failover, so configure a second ring on a separate subnet -- a direct crossover cable between the two appliances is ideal, because it survives a top-of-rack switch outage. The full procedure is on Site Cluster Setup; do not build the HA group until the site cluster reports every member online.
Then create the pool. A scale-up HA pool is created exactly like any other ZFS pool -- see Storage Pools -- with one constraint: select only disks from the shared enclosure. Before you create it, confirm in the Physical Disks section that the same disks appear with the same SCSI IDs and serial numbers on both appliances. That shared, identical naming is what makes the pool importable on either side, and if it is missing the HA group create will fail. Create the pool on the appliance you intend to be its primary, although that is a convention rather than a requirement.
For parity layouts, use at least double parity. RAIDZ2 or RAIDZ3 leaves the pool with error-correction capability while a failed device is being replaced, and -- unlike RAID10 -- cannot start on half its devices, which removes even the theoretical possibility of a split-brain.
Creating the storage pool HA group

The storage pool HA group associates one pool with the appliances allowed to import it, carries the failover policies, and owns the pool's virtual interfaces. It is the object you activate, deactivate and fail over.
Existing groups are listed on the Storage Pool HA Groups tab in the centre pane, which shows each group's state, the appliance currently hosting it, and its connectivity and link-state policies.

General tab

- Name -- prefilled as the pool name with
-ha-groupappended, so a pool namedpool-2producespool-2-ha-group. Keep the pool name in it; the group is what you will be looking for during an incident. - Description -- optional.
- Storage Pool -- the pool to protect. Only ZFS pools are listed, and a pool that already belongs to a group is rejected on OK.
- Export Timeout (seconds) -- default 50. This is how long the appliance losing the pool is given to export it during a failover. If it overruns, the acquiring appliance stops waiting and takes the devices preemptively, which gets the pool back into service but leaves the exporting appliance needing a reboot to clear its I/O stack. Raise it for a pool with many volumes and shares, where a clean export legitimately takes longer than a smaller one.
- Primary System -- read-only, and set to the appliance the selected pool is currently imported on.
- Secondary System -- the appliance that will take the pool. Only members of the same site cluster as the primary are offered.
- Tertiary System -- an optional third failover target, gated behind its own checkbox. The checkbox stays greyed out unless the site cluster has at least three members, which is why it is disabled on a two-appliance cluster.
- Enable SCSI3-PR Distributed Locking -- Auto, Enabled or Disabled, defaulting to Enabled for a new group. See Distributed locking for clustered clients below.
- Force -- unticked by default. Ticking it relaxes the device connectivity check from "every pool device is visible on the target" to "a majority of them are". It also allows the group to be created while FC sessions are active, which enabling HA would otherwise drop as the pool switches to FC ALUA mode. Leave it off for a first build: a device that is not visible on the secondary is a cabling fault to fix, not a check to bypass.
Connectivity tab

This tab decides when QuantaStor should fail a pool over preemptively -- that is, while the hosting appliance is still alive but has lost the connectivity that makes it useful.
- Enforce SMB HA VIF Access and Enforce NFS HA VIF Access -- off by default. Each restricts SMB and NFS access for the pool's shares to the networks the pool has virtual interfaces on. Turn them on when you need to be certain clients cannot reach the shares through an appliance-local address, because a client that connects that way keeps working right up until a failover and then does not come back.
- Ethernet Port Link State Policy -- default Failover if ALL Ports are Link-Down. The other settings are any port down, a majority of ports down, or disabled. "All ports down" is the conservative default; a single flapping port should not move a pool.
- FC Port Link State Policy -- the same choices, applied to Fibre Channel target ports, and with the same default. It only applies to pools that actually export volumes over FC.
- Enable Client Connectivity Checks -- off by default. Ticking it enables the two radio buttons and the Verify Connectivity button below, all of which are inert until it is on. With it enabled, QuantaStor pings the client addresses you list and fails the pool over when they stop answering, on either the ALL specified IPs are unresponsive or MAJORITY of specified IPs are unresponsive policy. The addresses must be genuine remote clients: an address belonging to a system in the same grid is rejected, since pinging your own grid tells you nothing about client reachability.
- Verify Connectivity -- pings the listed addresses now and reports how many answered, so you can confirm the list before relying on it.
From the command line
qs ha-group-create --pool=pool-2 --sys-secondary=qs-node2 \
--export-timeout=50 --enable-cluster-pr=enabled
Use qs ha-group-create to create the group and qs ha-group-modify to change it afterwards. The CLI takes the same settings under the names --client-connectivity-check-policy, --port-linkstate-policy, --fc-port-linkstate-policy, --verify-client-ips, --enforce-smb-allowed and --enforce-nfs-allowed. Two differences from the dialog are worth knowing: --sys-primary is settable rather than read-only, and --enable-cluster-pr defaults to auto where the dialog defaults a new group to enabled.
What the create checks, and what it rejects
The create is a long sequence of preconditions, and each one produces a distinct message. Reading the message saves guessing:
- The pool is not ZFS, or already belongs to another HA group.
- The chosen appliances are not all in the same site cluster. A group can still be created when no appliance is in a site cluster at all -- a clusterless group -- but it cannot carry a virtual interface and cannot fail over automatically, so it is only useful as a way to move a pool by hand.
- Either appliance has I/O fencing disabled:
Storage system '<name>' has I/O fencing disabled, this must be re-enabled before an HA group may be created. - The two appliances' system UUIDs share their first six hex digits. The SCSI reservation key is built from those digits, so identical prefixes would make it impossible to tell which appliance holds the pool.
- Any pool device is a SATA disk.
- Any pool device is not visible on the secondary appliance:
Unable to verify access to '<n>' devices (<scsi ids>) on secondary storage system for storage pool '<pool>'. - The devices do not support the persistent reservations that fencing requires -- QuantaStor registers a key on each one as part of the create, and fails if it cannot.
- Any grid member is running a service version older than the minimum the HA group code requires.
The HA virtual interface

Every client must reach the pool through the HA virtual interface, not through an appliance's own address. The virtual interface is a cluster resource that pacemaker moves with the pool, so a client pointed at it finds its data wherever the pool has landed. A client pointed at an appliance-local address works perfectly until the first failover and then does not come back, and that is by far the most common cause of "the failover worked but my clients did not recover".
An HA virtual interface is a cluster VIF with the scale-up pool use case, bound to an HA group. Cluster VIFs owns the dialog and the full set of use cases; create it there:
Three things are specific to the scale-up case:
- It needs an IP address of its own, on the client network, not in use anywhere else, and distinct from both appliances' addresses on that network. In practice each HA virtual interface consumes three addresses on its subnet: its own, plus one per appliance on the parent port.
- The parent port name is checked across the whole site cluster, not only on the appliance you selected, because a virtual interface that cannot follow the pool to the secondary appliance defeats the point.
- QuantaStor writes location constraints so the interface is only ever placed on the group's primary, secondary and tertiary appliances.
The interface can also be created from the HA group side, which is what the CLI does:
qs ha-interface-create --ha-group=pool-2-ha-group --parent-port=bond0.100 \
--ip-address=192.168.0.124 --netmask=255.255.0.0 --iscsi-enable=true --nvmeof-enable=true
See qs ha-interface-create, and qs ha-interface-list to check the result. --convert-vif turns an existing appliance-local virtual IP into an HA interface, and the matching --convert-to-vif on qs ha-interface-delete turns it back -- which is the supported way to keep an address alive while the paired appliance is being rebuilt.
How the interface is named
A scale-up HA interface is named for its parent port plus a generated tag, in the form <parent>:<tag>. The tag begins with ha, continues with the leading hex digits of the pool GUID, and ends with an index number. A pool whose GUID starts d786e0b7 on parent port bond0.100 produces the tag had70 and the interface bond0.100:had70. That is how a scale-up interface is told apart from a site VIF, which uses an sv tag instead.
Linux caps an interface name at 15 characters, so the length of the parent port name decides how many GUID digits fit -- and a long parent name can leave no room at all. If the create fails with a name-length error, either rename the parent interface or switch the appliance to eth-prefix naming on the Storage System Modify dialog.
Activating and deactivating the group
A group must be activated before automatic failover will happen. Until then the group exists and can be failed over by hand, but the site cluster will not move the pool on its own.
Activating restores the interface location constraints and places the pool's resource group on the appliance currently hosting it. Deactivating disables the failover policies without deleting anything. Two behaviours are worth knowing before you deactivate:
- A group cannot be activated while its site cluster is in maintenance mode. Exit maintenance mode first.
- While a group is deactivated, a manual failover to a different appliance is refused if the group still has virtual interfaces attached:
Cannot execute HA pool failover to another system while HA Group '<name>' is deactivated and has one or more HA interfaces.Re-activate the group, or remove its interfaces, to move the pool.
Use qs ha-group-activate and qs ha-group-deactivate from the command line.
Failover
Automatic failover
Once the group is activated, the site cluster monitors appliance health across every heartbeat ring, and QuantaStor's own checks watch the things a heartbeat cannot see. A failover is triggered when:
- An appliance stops answering on all heartbeat rings.
- The small write test QuantaStor runs against each pool every few seconds fails to complete. This is what catches a lost SAS path or a failed JBOD I/O controller, neither of which stops the appliance answering its heartbeat.
- Network ports fail according to the Ethernet or FC link-state policy on the group.
- The configured client addresses stop answering, if client connectivity checks are enabled.
The settle time on the group -- 60 seconds by default -- is a cooldown that prevents a second failover from being triggered immediately after one completes.
Deliberate failover
Move a pool by hand to test the cluster, or to empty an appliance before working on it.

On the General tab, HA Failover Group selects the group, Storage Pool and HA Group Status are shown read-only, From System is the appliance currently hosting the pool, and To System is where it is going. The other appliance in the group is preselected. Selecting the appliance the pool is already on is allowed and prompts for confirmation -- see A pool that will not import for why you would.

Advanced Settings carries two preliminary health checks. Import Health Checks is ticked by default and reviews health statuses on the target appliance that would stop it receiving the pool. Export Health Checks is unticked by default and reviews statuses on the source appliance that would stop it exporting cleanly. Either finding a critical status aborts the failover before anything moves, and names the checkbox to clear if you want to proceed anyway.
From the command line, both checks are explicit arguments:
qs ha-group-failover --ha-group=pool-2-ha-group --storage-system=qs-node2 \
--import-health-checks=true --export-health-checks=false
See qs ha-group-failover. The command runs synchronously and returns the group object when the failover completes.
What happens, in order
A failover is driven from the appliance receiving the pool, and runs in this sequence:
- Device information is rescanned and the receiving appliance verifies it can see the pool's devices.
- The pool's ALUA state is moved to standby, if ALUA is in use.
- The HA virtual interfaces are moved to the receiving appliance as a pacemaker resource group. A group with no virtual interfaces still fails over, with a warning; only the pool moves.
- The pool is exported on the appliance that had it, within the export timeout.
- SCSI-3 reservations on the pool's devices are preempted, giving the receiving appliance sole write access. If it cannot take ownership the failover stops here rather than importing on top of another appliance's reservation.
- LUKS devices are opened, for an encrypted pool.
- The pool is imported and activated.
- ALUA state is updated and an FC LIP is issued so Fibre Channel clients rescan their paths.
- SMB and NFS configuration, share namespaces and firewall rules are regenerated for the pool on its new host.
What clients see
The dialog warns that a failover "can take upwards of 30 seconds or more for larger configurations", and that is the right expectation to set. On a small pool -- four SAS disks, 576GB, RAIDZ2, no client load -- a failover measured end to end took 44 to 45 seconds, with the change of ownership visible in qs pool-list about 25 seconds in and the remaining time spent on share, ALUA and firewall reconfiguration. A pool with many volumes and shares takes longer, and the export stage is usually what grows.
For the duration, the pool's storage is unavailable: the virtual interface has moved but the pool behind it has not yet imported. Clients connected through the HA virtual interface see a stall and then recover, which is what iSCSI, NFS and SMB timeouts exist for; clients connected to an appliance-local address see the connection go away and do not recover. Applications with short storage timeouts may need those timeouts raised.
Two operational points follow from that:
- Test failover before the cluster carries production data, in both directions, and confirm the pool lands healthy each time.
- Fail back deliberately. After an appliance is repaired and rejoins, pools do not return on their own. If you run two pools, one on each appliance, you have to move one back by hand.
I/O fencing and SCSI-3 persistent reservations
Fencing is what makes the guarantee that only one appliance can write to the pool at a time, and it is enforced by the drives themselves rather than by the software. It is the scale-up equivalent of the ARP probing documented on High-availability VIF Management: both exist to stop two appliances claiming the same resource, one for addresses and one for disks.
When an appliance imports an HA pool it places a SCSI-3 persistent reservation of type WERO (write exclusive, registrants only) on every device in the pool. The registration key encodes who holds it and what for. It has the form 0xffaaaaaaffbbbbbb, where the ffs are separators, aaaaaa is the first six hex digits of the appliance's system UUID, and bbbbbb is the first six of the storage pool's UUID. On NVMe devices the middle separator is a single f and the appliance portion is truncated to five digits, because NVMe registration keys are shorter.
Read the current state with qs-iofence devstatus, which lists every device, its serial number, the keys registered on it, the key holding the reservation, and the reservation type:
# qs-iofence devstatus /dev/sdm 6XP2F4JX0000B2359Y4A (0xff753100ffd786e0) [0xff753100ffd786e0] <WERO> /dev/sdn 6SE46WWK0000B146RNKP (0xff753100ffd786e0) [0xff753100ffd786e0] <WERO> /dev/sdo PFVVJUYE (0xff753100ffd786e0) [0xff753100ffd786e0] <WERO> /dev/sdp 6SE31S1V0000B130JCEV (0xff753100ffd786e0) [0xff753100ffd786e0] <WERO>
Every device in the pool should carry the same key, and that key's appliance portion should match the appliance the pool is imported on. Run the command on both appliances: they see the same reservations, because the reservation lives on the drive. After a failover the appliance portion changes and the pool portion does not, which is the quickest confirmation that fencing followed the pool rather than being left behind. A device that shows no reservation while the pool is running, or one whose key names the wrong appliance, is a fencing problem -- resolve it by failing the pool over, including to the appliance it is already on.
QuantaStor also surfaces this in the Physical Disks section: a black underline on a device icon means the device is correctly fenced to the appliance running the pool, and a red underline means it is not.
Fencing is why cabling rules matter. If a cabling change ever isolated one enclosure to each appliance, reservations placed by one appliance would not reach the other's disks, which is the only realistic route to a split-brain -- and only for mirrored or 2d+2p and 3d+3p layouts, since a RAIDZ2 or RAIDZ3 pool cannot start on half its devices at all.
Distributed locking for clustered clients
Enable SCSI3-PR Distributed Locking on the HA group is a different thing from the fencing above, and it is easy to conflate them. Fencing is QuantaStor arbitrating which appliance owns the disks. Distributed locking is about the reservations clients place on QuantaStor's storage volumes: with it enabled, those client-side reservations are shared across every appliance in the HA cluster through the distributed lock manager, so a reservation a client holds survives a failover instead of being lost with the appliance that recorded it.
Turn it on for Windows Server Failover Clustering, including Hyper-V cluster shared volumes and SQL Server failover cluster instances, which depend on persistent reservations to arbitrate between cluster nodes. QuantaStor configures the DLM on each appliance itself; there is nothing to install.
The setting has three values. Enabled and Disabled are explicit. Auto means "keep doing whatever this pool is already doing", which is what a group created before this option existed, or by an automation client that does not send the field, will report -- Auto is not the same as off. Turning it off on a running pool is disruptive: reservations stop being shared, so any reservation a client is holding is lost at the next failover, and the dialog asks you to confirm that.
Maintenance
To work on one appliance without dismantling the cluster, put it in standby mode. A standby appliance stays in the site cluster but will not host resources: anything it is running moves to a healthy appliance, and it will not receive a failover. It has an auto-activating form, which returns to active once the appliance is healthy or after a reboot, and a manual form, which stays until you clear it. See Configure Member Standby for the settings and Site Cluster Setup for how standby differs from cluster-wide maintenance mode.
From the command line, qs site-cluster-set-standby-mode takes --standby-mode=active, standby-auto-activate or standby-manual-activate.
Deleting an HA group is the opposite of maintenance and worth stating plainly: it disables automatic failover and removes every virtual interface attached to the group. The pool stays online on whichever appliance is hosting it, but it stops being highly available, and the addresses clients were using disappear.
Troubleshooting
An HA group that will not form
Work through What the create checks, and what it rejects above -- the create names the reason it refused, and almost every failure is one of those checks. The two that account for most first builds are a device that is not visible on the secondary appliance, and appliances that are not both in the same site cluster.
For the device visibility case, compare the two appliances directly. qs-util devicemap prints each device with its /dev/disk/by-id path, vendor, model and serial number; run it on both appliances and sort the output. The /dev/sdX letters will differ and that does not matter -- the by-id paths and the serial numbers are what must match. qs disk-list shows the same information with the pool each disk belongs to.
A pool that will not import
Two situations look identical and have different fixes.
The pool is on an appliance but not running. Most often the HA group is deactivated, or has no virtual interfaces, or the interface addresses are in use elsewhere on the network. Check the group's state first. If the group is deliberately deactivated, run a manual failover from the appliance the pool is already on to that same appliance. That is more than starting the pool: it re-runs the fencing sequence and places the virtual interfaces, which a plain pool start does not do.
The enclosures were powered on after the appliances. Nothing will import, because the devices were not present at boot. Power the JBODs on, wait for them to come up, and then run the same failover-to-itself against each affected pool. In the Storage Pools section the tree label tells you which appliance a pool belongs to -- it reads pool-2 (on: qs-node2) -- so you know which appliance to target.
If the appliance that had the pool needs a reboot, it will say so. An unclean export -- one that overran the export timeout, or a pull of the SAS cables -- leaves in-flight writes in the I/O stack and a pool state that cannot be cleared any other way. Reboot it, let it rejoin the grid, and check its state detail in Properties before failing anything back to it.
The recovery procedures for hardware failures -- device replacement, enclosure power loss, split-brain, boot media loss -- are collected below.
Recovery and Resolution of Failure Scenarios
Recovery from Device Failures
HDDs (and many SSDs) have a typical annual failure rate (AFR) of 1.5% so replacing bad media is just a common eventuality in maintaining a healthy Storage System. For both scale-up and scale-out cluster configurations we always require parity based pools to have at least two parity devices (coding blocks) per stripe (VDEV or placement-group). For scale-up configurations this means RAIDZ2 or RAIDZ3 and for scale-out using erasure-coding that means a K+M (data blocks + coding blocks) where m is greater than or equal to two (>=2). Single parity layouts may be ok for some test environments but not for production workloads as they are not durable enough. With 2, 3 or more parity bits per stripe a single device failure will leave the effected pool running in a degraded state with additional parity information to still do error correction to deal with bad sectors or a second concurrent device failure which is critically important.
To recover from a device failure in a scale-up configuration simply add a new HDD or SSD that is equal to or greater in capacity than the failed device to the system with the device failure. Use an available free slot if one is available, then mark the device as a hot-spare. If there are no free slots available then first remove the failed device and then place the new replacement device into the same slot where the bad device was located. After that mark the new device as a hot-spare. Once marked as a hot-spare the new device will be automatically be utilized to repair the degraded pool.
There are two ways to mark a device as a hot-spare, in the Physical Disks section as a 'Universal Hot Spare' that can be used to make a spare as 'universal' meaning it may be used to repair any pool. The second option is in the Storage Pools section where one may assign the new spare device as a hot-spare for a specific pool, simply right-click on the degraded pool then use the 'Recover Pool / Add Hot-Spare..' dialog.
After setting up a new system where more than one pool has been created we recommend marking all spares as Universal Hot-Spares so that they're able to used by QuatnaStor to repair any pool in the HA cluster, this is done via the Physical Disks section. If you are explicitly replacing a bad drive for a pool then a direct assignment is preferred, simply right-click on the degraded pool and add the new device as a hot-spare using the 'Recover Pool / Add Hot-spare' dialog.
It is important to remove bad media as soon as possible as some media can behave poorly and cause bus resets and other things that can impact normal operations of the JBOD. Some brands of JBODs are more/less impacted by failed media but our general recommendation is to remove bad media once a pool has been repaired.
If there are multiple bad devices the repairs will be done one at a time. QuantaStor has logic to intelligently select hot-spares by media type, jbod location, and capacity. As such, marking a SSD as a hot-spare will only use the device to repair a failed SSD and similarly a HDD will only be used to repair a pool with a failed HDD. It will also ensure that the HDD is large enough and will have preferential selection to using a hot-spare that is in the same chassis as the failed device so that any enclosure level redundancy is maintained. In the event that multiple VDEVs are degraded the hot-spare will automatically be used to repair the VDEV with the most degradation.
Recovery from Pool I/O Failure
QuantaStor does a small mount of write IO every few seconds to ensure that a given Storage Pool is always available and writable. If the small write test fails (doesn't complete within ~10 seconds) the system will trigger a failover of the pool to the other HA paired node to restore write access.
Recovery from Disconnected SAS Cables
Pulling the SAS cables from a QuantaStor HA server node is one way to trigger a pool failover. When this happens the Storage Pool will recover access to the pool or pools via the paired HA node within typically 15-20 seconds. The node that had the pool previously now needs to be rebooted. This inconvenience is unfortunately required as in-flight IO (writes that did not complete) must be cleared from the IO stack and the FAILED pool state cannot be properly cleared without a reboot. Once the node completes a reboot it can then be used again to failover the pool back to it assuming any SAS and other connectivity issues have been resolved.
Recovery from Many Failed Devices
If a pool loses too many devices for example a pool with a 8d+2p layout that loses 3x devices (in the same VDEV) will have exceeded its parity block (2p) count and will not be able to repair itself. In many cases the pool will be lost at this point but there are a few things you can try in this scenario:
- cold power cycling all equipment, leave everything off for a minute before booting, power on JBODs first
- try starting the pool with special zpool import commands to rollback to the last good transaction group
- use ddrescue to copy (partially) bad media to good new media then try zpool import again
- if the device is a pass-thru from a HW RAID card (not recommended) then use 'Import Foreign' and 'Mark Good' commands to first bring the media back online
The key is to avoid this scenario by using at least double parity (RAIDZ2 or RAIDZ3), running regular scrubs (quarterly), and proper configuration of alerting to know when the system needs maintenance. Losing a pool is an extremely rare event and is typically due to long stretches of neglected maintenance where devices fail and no spares or hot-spares are ever supplied to repair the pool.
Recovery from Missing Pool
If a system is booted and the JBODs are powered off then all attempts to start a pool will not succeed. In such cases one can run a Manual HA Failover once the JBODs have started to force the pool to start and to place the Pool HA VIFs onto the correct node. In the Storage Pool section check to see what Storage System the impacted pool is on, you will see this in the label such as "pool-1 (on: qs-node-1)". Using this example one would want to run a manual HA failover from the node it is on 'qs-node-1' to the same node 'qs-node-1'. Running a manual HA failover is more than just a 'Pool Start' as it also includes HA failover of VIFs and a preemptive takeover of the Storage Pool devices from a IO fencing perspective.
Recover from unexpected power-off of Storage System (HA node)
When a system is powered off (and for any HA failover scenario) the virtual interfaces (VIFs) attached to the pool will automatically failover to the passive node to restore storage access to the pool for all protocols (NFS, SMB, iSCSI, FC, NVMeoF). If you find that one or more clients did not properly recover after the failover then you most likely have clients connecting to Network Shares of the Storage Pool via local IPs on the system rather than via the HA VIFs associated with the pool. There is a 'Client Connectivity Checker' dialog that is accessible from the Storage Systems section, simply right-click on a system to access it and it will enable one to check which IP addresses all clients are using and will report errors for incorrect IP address usage. This generally does not happen with block storage as iSCSI access is limited to just the HA VIFs so we're able to ensure correct connectivity. The same is not the case with NAS protocols like NFS/SMB where clients can be mis-configured to access the storage via local IPs that do not move with the HA VIFs attached to the pool.
Recover from unexpected power-off of JBOD
When you create a new storage pool QuantaStor will automatically try to achieve enclosure redundancy by striping across enclosures so that Storage Pool availability is not impacted by a failed JBOD. For example, with RAIDZ2 4d+2p one needs 3x JBODs for enclosure redundancy, with RAIDZ3 8d+3p one needs 4x JBODs to achieve enclosure redundancy. In a system without enclosure level redundancy (which is common) then disconnecting the JBOD will typically remove access to too many devices and the pool will stop. The system will have attempted to failover to recover the pool but this will also have failed and the pool will be in a 'error' or 'missing' state at this point. To recover from this turn all systems and JBODs off then power on the JBODs, then power on the Storage System nodes and the pool will start and recover automatically. The QuantaStor servers need to be rebooted in this scenario to clear the IO stack and the 'failed' storage pool state. The node that never had imported the failed pool (due to the JBOD power off) technically could be left running and this may be advised if it is hosting another un-related and unimpacted Storage Pool. Else it is best to power off everything then boot everything back up with JBODs first.
Recover from controller failure on JBOD
QuantaStor Storage Systems are generally configured with multi-path connectivity to all attached JBODs with each path coming from a different HBA. In such configurations the pool will continue to run and the number of paths will be simply reduced from 2 to 1 which in some scenarios can reduce performance, especially with SSD JBOFs. If the connectivity to the JBOF is single path and the JBOD IO controller fails this will cause the pool to automatically failover to the paired HA node (passive node). This trigger is typically caused by a failure of the write IO test which runs every few seconds. If the write test fails a HA failover of the pool is triggered immediately.
Recover from unexpected power-off TOR switch
The heartbeat mechanism in a HA cluster pair is maintained through what's called a 'Site Cluster'. Each Site Cluster is typically configured with two 'Cluster Rings' so that there is redundancy to the heartbeat in case a top-of-rack (TOR) switch is disabled, failed or powered off. In cases where there are spare 1GbE ports it is recommended to use a simple direct connect (crossover cable) between two nodes of an HA pair and to use that as the second cluster ring. In this way there's always a heartbeat available that will work even in the event all TOR switches are disabled. For those familiar with corosync+pacemaker technologies, QuantaStor uses these with a custom QuantaStor specific HA failover agent to manage the IP failover which does additional checks and activities like ARP flushing.
Re-integrate head node to HA cluster after hardware failure/power-off
To re-integrate a head node to a HA cluster after a power-off simply just power it back on. It will re-join the storage grid and will do checks to validate its readiness to accept an HA failover. Check in the Storage System section to verify the health status of the Storage System and if you see it in a 'Warning' or 'Error' state look in the 'Properties' section for more details under 'State Detail' as it will have a detailed description of the issue. For example, if the system is powered back on but the SAS cables to the JBOD are disconnected it will indicate a 'Warning' state with details indicating it is not ready to accept a HA failover due to media accessibility issues. Simply resolve the connectivity issues and the 'Warning' state will automatically clear.
Recover from ungraceful head node reboot
Simply wait for the reboot to finish, the system will resync with the grid and will be ready to accept pool failovers. If you have two or more pools you will need to manually push one of the pools back to the newly rebooted node as it will not fail-back automatically.
Recover from Corrupted Storage Pool
QuantaStor has many safe guards to prevent HA Storage Pools from ever getting corrupted. This is due to the HA failover technology in QuantaStor and how it does IO fencing and HA failovers. Assuming you have not been doing potentially harmful activities at the console/ssh to low level format media or to manually clear IO fencing (SCSI3-PR) while the pool is running then the pool is most likely not corrupted. Simply power off all JBODs and servers then cold boot with the JBODs powering up first then the QuantaStor storage nodes. The pool will automatically start on its own and will restore access to all the volumes and shares. If the pool does not automatically come back online we recommend getting some assistance from support to investigate further. Common problems are hardware failures which can be cables, backplanes, bad media causing bus resets and more and 99.99% of the time these are all fixable problems.
Recover from HA Storage Pool not Starting
If the Storage Pool is not starting automatically after a reboot of both HA nodes (head / controller nodes) then most likely the Storage Pool's associated HA Group is in the disabled state or does not have any HA VIFs associated with it. It is also possible that another system on the network is using the HA VIF IP addresses as that can also cause them to fail to start. If the pool HA group is intentionally disabled then use the Manual HA Failover to failover the pool back to the storage system node that it is already on. This will trigger the necessary IO fencing and pool startup sequence to get it started and serving storage.
Recover from OSNexus boot media failure/software corruption/reinstall
The most common cause of software corruption is the boot media running out of space which can cause the local QuantaStor grid database to get corrupted. In such a case QuantaStor may reset the database back to a default empty database and you'll know this right away as the 'Getting Started' dialog will appear when you login as there will be no license keys on the system anymore. For this reason we recommend using 480GB or larger SSD/NVMe based media for the QuantaStor boot device.
This is not something to be overly concerned about, it's all fixable. First remove any excess or large log files from /var/log if you see or suspect that that boot media is simply 100% full. If the boot media has failed then you'll want to replace it with new media and reinstall QuantaStor. The QuantaStor OS boot media does not impact the health of the Storage Pools.
- NOTE: Encrypted pools do keep the pool encryption keys on the boot media (AES key wrapped) so they should be exported/saved after the pool is created (see Export Pool Encryption Key Metadata). Assuming the second system still has healthy media or you have a backup of your pool encryption keys then there's no risk to reinstalling QuantaStor onto new or existing boot media.
When re-installing we recommend using the latest version of QuantaStor. QuantaStor can be reinstalled from scratch on the the impacted node and then the node can be added back into the storage grid. The Site Cluster and associated HA VIFs for the pool will typically need to be re-created. The procedure here is to temporarily change the HA VIFs to local VIFs on the healthy node (see Delete HA Cluster VIF dialog), then after the paired node has been re-added to the grid and the Site Cluster recreated then the local VIF(s) for the pool can be converted back to a HA VIF(s) (see Create HA VIF dialog).
In a single-node configuration you'll need to use the 'Storage Pool Import..' dialog to re-import the storage pool. It will automatically recover all the metadata associated with the pool including Hosts, Host Groups, Storage Volume ACLs, Network Shares and their configuration settings and more.
Recovery from ransomware
The best defense against ransomware is to have snapshots. QuantaStor's Snapshot Schedules and Remote-replication Schedules both have settings for Long-Term Snapshot Retention. We recommend at least two quarterly snapshots so that one can roll-back up to 6 months without having to resort to tape backups. Note that snapshots only accrue storage utilization when they hold onto deleted files. If you're rarely or never deleting files then it is safe to hold more quarterly snapshots as it will have minimal to no impact on your available capacity. If on the other hand you're frequently deleting data then you'll need more room for snapshots and will want to size your system larger to be able to hold snapshots for longer as protection against cyberattacks like ransomware. Another key protection against ransomware is to prevent anyone from logging into the systems as root. We also recommend making all low level root access via the 'qadmin' sudo user account require login via an SSH authorized key and that you reset the 'qadmin' account password with a strong password immediately on new systems.
Remediate Incorrectly-replaced Hard Drive
If the pool has gone into a failed condition because too many devices have been removed for the pool to continue this can be fixed by identifying the correct device to be removed and putting the remaining ones back in place. If the pool is in a failed state the system will need a reboot. If the removal of the wrong devices caused a failover and the pool could not be imported you may need to reboot both nodes of an HA pair. QuantaStor's hardware integration makes it easy to turn on the LED beacon of the bad device so that it is easy to identify and replace by remote-hands. We recommend getting assistance from support to identify the bad device if it is unclear or if the system has been been properly configured with an enclosure layout setting for the hardware.
If you removed a device from a system by mistake, simply put the device back in and the Storage Pool will recover it automatically. There is some delay before a hot-spare is injected into the pool so there is a minute or so where the improperly removed device can be put back in without issue.
Recognize and recover from HA split-brain events
To start, is very very difficult to get QuantaStor into a split-brain configuration but in this section we'll outline how one may attempt to do so. A split-brain is where a mirrored pool can be split in two halves such that both head nodes are successfully running the half pool devices. In this mode one has split a given storage pool into two pools, typically a RAID10 pool that is now running as two separate RAID0 pools at the same time with the same name on two separate nodes.
In a split-brain scenario these two separate head nodes are running two separate pools that are now diverged from each other such that merging them back together can be a time consuming and IO intensive process. Typical recovery from a scenario like this where there have been changes made to both pool involves reformatting one of the two pool's devices, the pool deemed to be the 'bad' pool, and then re-silvering the media back to the 'good' pool. Pools using RAIDZ2 (eg 4d+2p) or RAIDZ3 (8d+3p) cannot not start with 50% of the media so they cannot get into traditional split-brain scenarios at all. These are some of the ways that QuantaStor prevents split brain from ever happening:
- QuantaStor prevents split brain by IO fencing (SCSI3-PR) all media devices in a given storage pool (dual-port SAS and dual-port NVMe), this ensures that only one node at a time can start the pool with the media. To bypass this one would need to use RAID10 then through a series of cabling changes isolate each JBOD to a separate QuantaStor head node. This would require significant intentional physical reconfiguration of the hardware.
- QuantaStor prevents split brain by only allowing a pool to run on the system which has all the the HA virtual interfaces (VIFs) which move as a cluster resource group. Only the node with the VIFs are allowed to start the pool but in the re-cabling scenario described above with a RAID10 pool it is possible to failover a RAID10 pool from one node to another where on node-1 JBOD1 is being used and node-2 JBOD2 is being used.
- QuantaStor prevents split brain by only starting pools where enough devices are available for it to start in a degraded state. Again, this still leaves a RAID10 pool somewhat vulnerable if the devices were perfectly split between two JBODs.
In summary, the only way to cause a split-brain is to use a RAID10 pool layout and then carefully manipulate JBOD access through recabling to isolate the JBODs to separate nodes. For pools using parity based RAID such as RAIDZ1/RAIDZ2/RAIDZ3 it is not possible to have two different versions of the same pool running as any given pool cannot start with 50% of the devices. There are a couple of exceptions to this which are pools created using 2d+2p RAIDZ2 and 3d+3p RAIDZ3 layouts. Now that we understand the narrow scenarios which can lead to a split brain what are the ways to avoid the possibility of it ever happening.
- When using RAID10 with enclosure redundancy across two JBODs make sure you have two paths to both controller head nodes to both JBODs, this ensures the IO fencing will reach all devices at all times
- When using RAID10 you can opt to use a single JBOD or more than two JBODs.
- If the above cannot be ensured then one could avoid using RAID10, 2d+2p and 3d+3p storage pool layouts. In most cases we recommend RAIDZ2 4d+2p for virtualization and databases rather than RAID10 as it provides more capacity, better durability and eliminates even the rare potential for split-brain.
Recovery from multiple JBOD failures
If too many JBODs fail (powered off, etc) then the pool will stop. The best way to correct that is to power the JBODs on and then reboot the node that was running the pool before the JBOD failure. If multiple pools were effected then it is best to just cold boot the whole cluster. In a pool with an enclosure redundant layout such as a RAIDZ2 4d+2p pool layout spanning three enclosures JBOD-A, JBOD-B and JBOD-C then the pool will continue to run if any one JBOD fails. After a JBOD failure the JBOD should be powered back on and then one should manually failover the pool back to the node it is already on (for example if the pool is on node-1, then failover to node-1 where it already lives) to force a revalidation of all connectivity. QuantaStor has logic to automatically re-fence the media after a JBOD restart to ensure it is locked to the node where the pool is running. You can verify the IO fencing by checking in the Physical Disks section and look for a black underline device icon across all the media that makes up a pool. If you see a red-underline that indicates a IO fencing issue which can be resolved through a failover to the node the pool is on or to the paired node.
Another scenario to consider with multiple JBODs is what if one has multiple JBOD failures but at different times. For example, with a pool that spans JBODs A, B, and C such that the pool can run in a degraded state with any two JBODs online and JBOD-C is powered off then the JBOD-C is powered ON and JBOD-B is powered off before the pool has resilvered (repaired) then the pool will stop. This is because JBODs A+B have the most recent pool information and they need to be powered on with JBOD-C so that the pool devices on JBOD-C can rebuild/resilver before the devices in JBOD-C may provide additional durability to the pool. The filesystem (we use OpenZFS with scale-up HA clusters) detects these scenarios where the media are at different points in time transactionally and will prevent the pool from starting and/or will stop the pool in such a scenario. The fix is to start the pool again with all three JBODs A+B+C and let C resilver to get the pool back to full health. Then one could disable JBOD-B and the pool would run degraded with A+C. Going direct from a degraded A+B to degraded A+C will not work, again this will result in the pool automatically stopping and/or not starting.
Resolving Disks Not Visible on both HA nodes / Storage Systems
If you don't see the Storage Pool name label on the pool's disk devices under the Physical Disks section on both Storage Systems use this checklist to verify possible connection or configuration problems.
- Verify physical connectivity from both QuantaStor nodes/systems used in the Site Cluster to the back-end shared JBOD for the storage pool or the back-end SAN.
- Verify SAS (or FC) cables connected to backend storage are securely seated and link activity lights are active
- Verify that all devices that make up the pool are dual-port SAS, dual-port NVMe, FC, or iSCSI devices that support multi-port and Persistent Reservations (NVMe dual-port, SAS dual-port, and NL-SAS dual-port media are all supported for use in scale-up HA clusters).
- Verify SAS JBODs have at least two SAS expansion ports and each QuantaStor server has two HBAs. JBODs/JBOFs with 3 or more expansions ports and redundant SAS Expander/Environment Service Modules(ESM) are preferred.
- Verify SAS JBOD cables are within standard SAS cable lengths (less than 5 meters) of the SAS HBA's installed in the QuantaStor Systems. Use optical SAS cables for longer distances.
- Note: qs-util devicemap utility may be helpful when troubleshooting connectivity issues as well as qs-iofence devstatus
Useful Tools
qs-iofence devstatus
The qs-iofence utility is helpful for the diagnosis and troubleshooting of SCSI reservations on disks used by HA Storage Pools.
# qs-iofence devstatus
- Displays a report showing all SCSI3 PGR reservations for devices on the system. This can be helpful when zfs pools will not import and "scsi reservation conflict" error messages are observed in the syslog.log system log.
qs-util devicemap
The qs-util devicemap utility is helpful for checking via the CLI what disks are present on the system. This can be extremely helpful when troubleshooting disk presentation issues as the output shows the Disk ID, and Serial Numbers, which can then be compared between nodes.
HA-CIB-305-3-27-A:~# qs-util devicemap | sort ... /dev/sdb /dev/disk/by-id/scsi-35000c5008e742b13, SEAGATE, ST1200MM0017, S3L26RL80000M603QCVF, /dev/sdc /dev/disk/by-id/scsi-35000c5008e73ecc7, SEAGATE, ST1200MM0017, S3L26RWZ0000M604EFKC, /dev/sdd /dev/disk/by-id/scsi-35000c5008e65f387, SEAGATE, ST1200MM0017, S3L26JPA0000M604EDAH, /dev/sde /dev/disk/by-id/scsi-35000c5008e75a71b, SEAGATE, ST1200MM0017, S3L26Q9N0000M604W6MB, ...
HA-CIB-305-3-27-B:~# qs-util devicemap | sort ... /dev/sdb /dev/disk/by-id/scsi-35000c5008e742b13, SEAGATE, ST1200MM0017, S3L26RL80000M603QCVF, /dev/sdc /dev/disk/by-id/scsi-35000c5008e73ecc7, SEAGATE, ST1200MM0017, S3L26RWZ0000M604EFKC, /dev/sdd /dev/disk/by-id/scsi-35000c5008e65f387, SEAGATE, ST1200MM0017, S3L26JPA0000M604EDAH, /dev/sde /dev/disk/by-id/scsi-35000c5008e75a71b, SEAGATE, ST1200MM0017, S3L26Q9N0000M604W6MB, ...
- Note that while many SAS JBODs will typically produce consistent assignment of /dev/sdXX lettering, FC/iSCSI attached storage will typically vary due to variation in how Storage Arrays respond to bus probes. The important parts to match between nodes are the /dev/disk/by-id/.... values, and the Serial Numbers in the final column.
Command line reference
| Command | Purpose |
|---|---|
| qs ha-group-create | Create an HA group for a pool |
| qs ha-group-modify | Change nodes, policies, timeouts or distributed locking |
| qs ha-group-list | List HA groups and their state |
| qs ha-group-get | Full detail for one group, including its interfaces and fenced device serials |
| qs ha-group-activate | Enable automatic failover |
| qs ha-group-deactivate | Disable automatic failover |
| qs ha-group-failover | Move a pool to another appliance |
| qs ha-group-delete | Remove the group and its virtual interfaces |
| qs ha-group-get-health-status | Report stored failover health statuses for a pool |
| qs ha-interface-create | Add an HA virtual interface to a group |
| qs ha-interface-list | List HA virtual interfaces |
| qs ha-interface-get | Detail for one HA virtual interface |
| qs ha-interface-delete | Remove one, optionally converting it to a local virtual IP |
| qs site-cluster-set-standby-mode | Put an appliance into or out of standby |
Related pages
- HA Cluster Setup (external SAN) -- the same design with SAN-attached storage behind it
- Site Cluster Setup -- heartbeat rings, cluster membership, maintenance mode
- Cluster VIFs -- the virtual interface dialog and every use case
- High-availability VIF Management -- how cluster VIFs are managed and probed
- Configure Member Standby -- standby settings for one appliance
- Multipath Configuration -- multipath device naming, path counts, and the single-ported media alert
- Hardware Controllers & Enclosures -- HBAs, enclosures, slot numbering and drive identify
- Storage Pools -- creating and managing the pool itself
- Physical Disks/Devices -- verifying disk visibility and serial numbers
- Grid Configuration -- building the storage grid
- Network Ports -- port addressing and naming
Verified against QuantaStor 6.9.0.