Clustered SCSI-3 Persistent Reservations

Revision as of 16:32, 28 September 2026 by Qadmin (talk | contribs) (New page: clustered SCSI-3 persistent reservations for Hyper-V/WSFC/SQL FCI on HA pools (6.9); verified on a 6.9.0 HA pair)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)


Clustered SCSI-3 persistent reservations let the reservations a client cluster places on QuantaStor Storage Volumes survive an HA failover. Windows Server Failover Clustering -- including Hyper-V with Cluster Shared Volumes and SQL Server Failover Cluster Instances -- depends on them. This page explains why those clusters need the feature, how to turn it on for an HA pool, and what QuantaStor does when you do.

The feature is off by default. If you present HA pool volumes to a Windows failover cluster, turn it on in the HA group's Modify Group dialog before the cluster starts using the disks.

Section Purpose
Why Windows failover clusters need it What persistent reservations do for WSFC, and what goes wrong without this feature
How it works How reservations are shared between the appliances of an HA group
Requirements Pool type, protocols, operating system, release and cluster prerequisites
Enabling clustered reservations The HA group setting, in the web interface and from the CLI
Setting up a Hyper-V or WSFC cluster End-to-end walkthrough from HA group to Cluster Shared Volume
When to leave it disabled What the feature costs, and why it is not on by default
Disabling clustered reservations What turning it off does to a running cluster
Allow Reboot / Lock Recovery The companion setting in the same dialog
Checking the state on an appliance Confirming every volume is sharing its reservations
Troubleshooting When the setting is on but reservations are not shared

Why Windows failover clusters need it

A Windows Server Failover Cluster arbitrates ownership of its shared disks with SCSI-3 persistent reservations (PR). Each cluster node registers a key with every clustered LUN, and the node that owns a disk holds a reservation on it. When a node stops responding, the surviving nodes remove its registration, which fences it off the disk so it cannot keep writing to storage the cluster has handed to someone else. Cluster Shared Volumes, where every node reads and writes the same LUN at once, rely on this to keep a failed node away from the volume. SQL Server Failover Cluster Instances use the same cluster disks and depend on it in the same way.

Because the whole cluster depends on reservations behaving correctly, Microsoft's Validate a Configuration wizard tests them. Its storage tests register, reserve, preempt and release keys against every candidate disk, and a disk that fails the SCSI-3 persistent reservation test cannot be used as a supported cluster disk.

A reservation is state held by the SCSI target, not data on the disk. On a QuantaStor HA pool the target that recorded the reservation is the appliance that currently owns the pool. Without clustered reservations, that record exists only on that appliance. When the pool fails over, the receiving appliance starts serving the same LUNs with an empty reservation table: every key the Windows nodes registered, and the reservation that says which node owns each disk, is gone. The pool came back, but from the cluster's point of view the disks no longer carry the registrations it placed on them, and it treats them as failed. Cluster disks go offline and Cluster Shared Volumes stop, which is exactly the outage an HA pool is meant to prevent.

With Clustered SCSI3-PR Reservations enabled, every appliance in the HA group holds the same reservation state for every volume in the pool, so the appliance that takes the pool over already knows every registration and reservation, and the Windows cluster sees no change.

This is a different thing from the persistent reservations QuantaStor places on the pool's own disks. Those are I/O fencing: QuantaStor arbitrating which appliance may import the pool, described in I/O fencing and SCSI-3 persistent reservations. Clustered reservations concern the reservations that clients place on the storage volumes QuantaStor exports, and the two are independent.

How it works

QuantaStor shares reservation state between the appliances of an HA group through the Linux distributed lock manager (DLM), which runs as a resource of the site cluster's Corosync and Pacemaker stack.

  • Every exported storage volume in the pool gets its own DLM lockspace, named after the volume's SCSI device identifier. The reservation state lives in that lockspace rather than in a single appliance's memory.
  • The appliance that owns the pool serves the volume and is a member of its lockspace.
  • Every other appliance in the HA group -- the secondary, and the tertiary if the group has one -- holds a placeholder device with the same identity, which keeps it a member of the same lockspace. It carries no data; its job is to hold a copy of the reservation state.
  • During a failover the appliances swap roles device by device, so at every moment at least one appliance is holding each volume's reservations. The appliance taking over joins with the state already in place.

Setting the HA group to Enabled is the whole of the configuration. QuantaStor installs the DLM, creates and starts the Pacemaker DLM resource, and switches the pool's volumes into clustered mode on every appliance in the group. Appliances that did not run the operation converge on their own shortly afterwards, and an appliance that was down or rebuilt catches up when it returns. Volumes created in the pool later pick the setting up when they are created.

Requirements

  • A scale-up (ZFS) Storage Pool protected by an HA group. Clustered reservations are a property of the HA group, so a pool without one cannot use them. See HA Cluster Setup (JBODs) and HA Cluster Setup (external SAN).
  • iSCSI or Fibre Channel. The feature applies to storage volumes exported over iSCSI and Fibre Channel. It does not apply to NVMe over Fabrics.
  • Ubuntu 22.04 or 24.04 appliances. The DLM and SCSI target driver changes the feature depends on are built for those two platforms only. On any other platform QuantaStor refuses to put volumes into clustered mode rather than risk the target stack.
  • Every appliance in the HA group at 6.9.0 or newer, and rebooted onto its current target driver. Enabling the setting is refused, naming the appliance, if any member of the group is on an older release, or if an appliance has had new target drivers installed but has not rebooted since. Complete the upgrade on every node, including the reboot, before turning the feature on.
  • A healthy site cluster. The appliances must be members of a site cluster with Corosync and Pacemaker running. QuantaStor gives the Corosync cluster a name if it does not have one, because the DLM cannot start without it; this is done across the cluster under maintenance mode, and needs no action from you.

The operating system on the Windows side needs nothing beyond what a failover cluster already requires: multipath I/O configured for the iSCSI or Fibre Channel paths to the QuantaStor appliances, and the cluster nodes' initiators registered as hosts on QuantaStor.

Enabling clustered reservations

Navigation: Storage Management → Storage Pools (section) → select the pool → Storage Pool HA Resource Group (toolbar group) → Modify Group
 
The General tab of Modify Storage Pool High-Availability Group. Clustered SCSI3-PR Reservations shows Disabled, the default for a new HA group; Allow Reboot / Lock Recovery sits directly below it.

On the General tab of the Modify Storage Pool High-Availability Group dialog, set Clustered SCSI3-PR Reservations to Enabled and click OK.

The same field is on the Create Group dialog, so you can enable it when you first create the HA group. On both dialogs it offers Enabled and Disabled, and a new HA group starts at Disabled.

The change applies to the running pool. No failover, service restart or reboot is needed: existing volumes are switched into clustered mode in place, and the other appliances in the group converge within a short time. The Modify Group task reports "Enabling SCSI3-PR distributed locking" while it works. On an appliance that has never had the DLM installed, the first enable includes the package install and takes longer.

If the DLM cannot be configured -- the site cluster is not running, for example -- the HA group is still created or modified, because it is useful without clustered reservations. QuantaStor records the reason in the service log -- and, on a create, raises a warning on the group -- and keeps retrying on its own, and the volumes take up clustered mode once the condition is cleared.

Enable it before the Windows cluster takes reservations on the volumes. The walkthrough below does it first for that reason.

From the command line

qs ha-group-modify and qs ha-group-create take --enable-cluster-pr:

qs ha-group-modify --ha-group=pool-1-ha-group --enable-cluster-pr=enabled

It accepts enabled, disabled and auto, and auto is what you get when you leave the argument out. On a modify, auto leaves the current setting unchanged, so a command that changes some other property of the group never switches clustered reservations off as a side effect. On a create, auto resolves to disabled unless the pool's volumes are already in clustered mode. Pass enabled explicitly when you want the feature.

Setting up a Hyper-V or WSFC cluster

This is the order to follow when presenting QuantaStor volumes to a new Windows Server Failover Cluster, whether for Hyper-V Cluster Shared Volumes or for a SQL Server Failover Cluster Instance.

1. Enable clustered reservations on the HA group

Set Clustered SCSI3-PR Reservations to Enabled on the pool's HA group, as described in Enabling clustered reservations. Do this first, so that every reservation the cluster ever takes is shared between the appliances from the start.

2. Create the volumes

Create the volumes the cluster will use in the HA pool -- see Creating a Storage Volume. A small volume for the cluster's disk witness, if you intend to use one, belongs in the same pool.

3. Present the volumes to every cluster node

Add each Windows cluster node as a host with its iSCSI IQN or Fibre Channel WWPNs, put all of the nodes in one Host Group, and assign the volumes to that Host Group. Host Groups covers why a group assignment is the right tool for a cluster; for Fibre Channel, Assigning a storage volume over FC covers zoning and WWPNs.

For iSCSI, have each node log in to the pool's HA virtual interface rather than an appliance's own address, so its sessions follow the pool when it fails over. For Fibre Channel, configure multipath on each node; QuantaStor advertises the paths through the owning appliance as active and the others as standby.

4. Bring the disks online on one node

On one cluster node, rescan storage, bring the new disks online, initialize them and create the volumes you intend to use, in the usual way for disks that will become cluster disks. The other nodes see the same LUNs through the host group and need no preparation of their own.

5. Run cluster validation

Run the Validate a Configuration wizard in Failover Cluster Manager, or the Test-Cluster PowerShell cmdlet, against all of the nodes and include the storage tests. The SCSI-3 persistent reservation test should pass on every QuantaStor disk.

If you want proof that reservations survive a failover before going into production, fail the pool over to the other appliance with Manual Failover (see Deliberate failover) and run the storage validation again, or confirm in Failover Cluster Manager that the cluster disks stayed online throughout.

6. Add the disks to the cluster and to Cluster Shared Volumes

Add the validated disks to the cluster as cluster disks, then add the ones Hyper-V will use to Cluster Shared Volumes. For a SQL Server Failover Cluster Instance, select the cluster disks during SQL Server's failover cluster setup instead.

When to leave it disabled

Leave the setting Disabled on HA groups whose clients never take SCSI-3 reservations -- VMware vSphere datastores, Linux hosts, file shares, and single Windows servers. The feature is off by default because it has a cost:

  • One DLM lockspace for every exported volume, held on every appliance in the HA group, including the ones not serving the pool. A pool whose clients never reserve a LUN gains nothing from that.
  • A dependency on the cluster's lock manager. A volume in clustered mode relies on the DLM being able to form and recover its lockspace. QuantaStor gates every step on the DLM being verified healthy on the appliance, precisely because a lockspace that cannot be joined or recovered can stall the appliance's SCSI target. That is a risk worth taking for a Windows cluster that needs the feature, and not one to take by default.

Disabling clustered reservations

Set Clustered SCSI3-PR Reservations back to Disabled in Modify Group. The dialog asks you to confirm, and means it:

Disabling Clustered Persistent Reservations is disruptive to connected clustered clients such as Windows Server with Hyper-V. Reservations are no longer shared across the HA cluster, so any reservation a client is holding will be LOST the next time this pool fails over. Are you sure you want to do this?

The pool's volumes are taken back out of clustered mode straight away. The Pacemaker DLM resource itself is left in place.

Deleting the HA group does not turn the feature off. Removing an HA group is a routine step in HA pool maintenance, and tearing the DLM down with it would drop the reservations a running Windows cluster is holding, so the lockspaces and the devices holding them are deliberately left as they are. Removing the DLM configuration entirely is a support operation; contact OSNEXUS support if you need it.

Allow Reboot / Lock Recovery

The General tab carries a second, related checkbox: Allow Reboot / Lock Recovery. It is off by default, and it applies to HA pools whether or not clustered reservations are enabled.

It exists for one situation. When an appliance has been fenced away from a pool -- its peer took the pool over and preempted its reservations on the pool's disks -- the pool on the fenced appliance goes SUSPENDED or FAILED, and ZFS cannot bring a suspended pool back without a reboot. Until that reboot the appliance serves nothing, and it cannot always be shut down cleanly either, because I/O already submitted to the suspended pool never completes. Left alone, the appliance sits in that state until an administrator restarts it by hand.

Ticking the checkbox allows the appliance to restart itself to recover. QuantaStor does not otherwise reboot appliances on its own, so the restart is hedged with three conditions, all of which must hold:

  1. at least one pool imported on the appliance is SUSPENDED or FAILED;
  2. every such pool's HA group has Allow Reboot / Lock Recovery ticked -- one group that has not opted in vetoes the restart;
  3. the appliance holds no other healthy pool, so an appliance that is still serving storage is never restarted.

When the conditions hold, QuantaStor raises an alert and starts a three-minute countdown task before restarting. Cancel the task to stop the restart; the veto lasts until the appliance next reboots. The conditions are checked again just before the restart, so failing a pool back to the appliance during the countdown also stops it. The appliance does not re-import the pool when it comes back -- the pool belongs to its peer by then -- which is why the restart cannot loop.

From the command line the equivalent is --allow-lock-recovery-reboot=enabled on qs ha-group-modify or qs ha-group-create; as with --enable-cluster-pr, leaving it out leaves the setting unchanged.

Checking the state on an appliance

Two read-only commands, run as root (with sudo) from a shell on an appliance in the HA group, show whether reservations are being shared.

qs-scstdlm check reports whether the appliance's distributed lock manager is ready, and if not, the first thing standing in the way:

sudo qs-scstdlm check

A ready appliance prints lines beginning READY:. A line beginning NOTREADY: names the problem -- the DLM daemon not installed, the Pacemaker DLM resource not created, Corosync with no cluster name, and so on. On an HA group that has never had the feature enabled it reports the following, which is expected until you enable the setting and is not a fault:

NOTREADY: the Pacemaker 'dlm' resource has not been created in the cluster.

qs-scstdlm inspect lists every SCSI target device on the appliance with its clustered-mode flag, whether the appliance has joined that device's lockspace, and the persistent reservations the device currently holds:

sudo qs-scstdlm inspect
sudo qs-scstdlm inspect --no-prs

Add --no-prs to list the devices without their reservations, which keeps the output short on an appliance exporting many volumes.

In the CLM column, 1 means the device is in clustered mode. The LOCKSPACE column should read joined for every such device. NOT-JOINED alongside CLM 1 means that volume's reservations are not being shared and would not survive a failover. An appliance with no volumes in clustered mode prints (no SCST devices with a cluster_mode attribute on this node). On the appliance that does not own the pool, the devices are the placeholders described in How it works and show no backing file. Run it on both appliances: after the Windows cluster has taken its reservations, both should list the same registrations for each volume.

Troubleshooting

Enabling is refused with a version error. The message names the appliance holding it up. Either it is on a release older than 6.9.0, or it has had new target drivers installed and has not rebooted onto them. Finish upgrading that appliance, including the reboot, and try again.

The setting is Enabled but volumes are not in clustered mode. Run qs-scstdlm check on each appliance. QuantaStor holds volumes out of clustered mode until the appliance is running a supported platform, is a member of a site cluster, has Corosync and Pacemaker running with a named cluster, and has a running DLM daemon, and it records which condition is missing in the service log. Once the condition clears, the volumes switch over on their own.

Validation fails the SCSI-3 persistent reservation test. Confirm the setting is Enabled on the pool's HA group, and that qs-scstdlm inspect shows every volume in clustered mode and joined on both appliances. A disk presented from a pool that is not in an HA group, or through an appliance's own address rather than the HA virtual interface, is not covered by the feature.

Cluster disks went offline after a failover. Check whether the setting was Enabled at the time. Reservations taken while it was Disabled are not shared, and are lost when the pool moves. Enable the setting, then bring the cluster disks back online in Failover Cluster Manager.

Related pages


Verified against QuantaStor 6.9.0.