Windows MPIO for Failover Clusters and Hyper-V: Difference between revisions
Recommended Windows MPIO setup and timer tuning for failover clusters, CSV and Hyper-V on QuantaStor HA pools |
m Reword as the recommended configuration |
||
| Line 1: | Line 1: | ||
[[Category:admin_guide]] | [[Category:admin_guide]] | ||
This page | This page describes the recommended Microsoft Multipath I/O (MPIO) configuration for Windows Server hosts that use QuantaStor storage volumes as shared cluster disks, including Windows Server Failover Clustering, Cluster Shared Volumes, Hyper-V clusters and SQL Server Failover Cluster Instances. It applies to Fibre Channel and iSCSI access to a QuantaStor HA pool, and it is the client-side companion to [[Clustered SCSI-3 Persistent Reservations]]. | ||
'''Summary.''' Install MPIO, let the Microsoft DSM claim QuantaStor devices, and | '''Summary.''' Install MPIO, let the Microsoft DSM claim QuantaStor devices, and apply the recommended MPIO timer settings below on every cluster node. With this configuration, an HA failover or an appliance reboot is a brief, transparent pause for the cluster: cluster disks stay online, Cluster Shared Volumes keep running, and Hyper-V virtual machines continue without interruption. The configuration was validated on Windows Server 2022 Hyper-V clusters connected over Fibre Channel to a QuantaStor HA pair. | ||
{| class="wikitable" | {| class="wikitable" | ||
! Section !! Purpose | ! Section !! Purpose | ||
|- | |- | ||
| [[# | | [[#How the recommended settings help|How the recommended settings help]] || How MPIO carries cluster disks through an HA failover | ||
|- | |- | ||
| [[#Installing MPIO and claiming QuantaStor devices|Installing MPIO and claiming QuantaStor devices]] || The MPIO feature and the QuantaStor hardware ID | | [[#Installing MPIO and claiming QuantaStor devices|Installing MPIO and claiming QuantaStor devices]] || The MPIO feature and the QuantaStor hardware ID | ||
| Line 16: | Line 16: | ||
| [[#Paths and load balancing|Paths and load balancing]] || What a correctly configured disk looks like | | [[#Paths and load balancing|Paths and load balancing]] || What a correctly configured disk looks like | ||
|- | |- | ||
| [[#After maintenance on a QuantaStor appliance|After maintenance on a QuantaStor appliance]] || | | [[#After maintenance on a QuantaStor appliance|After maintenance on a QuantaStor appliance]] || Confirming all paths are in use before moving pools back | ||
|- | |- | ||
| [[#Troubleshooting|Troubleshooting]] || | | [[#Troubleshooting|Troubleshooting]] || Quick checks if a disk or path does not look as expected | ||
|} | |} | ||
== | == How the recommended settings help == | ||
During an HA failover the pool moves from one QuantaStor appliance to | During an HA failover, the pool moves from one QuantaStor appliance to its partner. For a short interval, while the receiving appliance imports the pool, I/O is held rather than completed. With Fibre Channel and ALUA, QuantaStor keeps the standby paths through the partner appliance present throughout, so Windows sees the paths change state and continues on the new active paths, typically within well under a minute. | ||
MPIO decides how long to hold I/O and keep a disk while its paths change. The Windows defaults are tuned for general-purpose storage and allow only a short window (20 seconds before an unreachable disk is released, 60 seconds per I/O). The recommended settings extend that window so it comfortably covers an HA failover, including a slower one, for example when a busy pool takes longer to export. The cluster then experiences the failover as a short pause, and Cluster Shared Volumes and virtual machines carry on as normal. | |||
Together with [[Clustered SCSI-3 Persistent Reservations]], which keeps the cluster's disk reservations in place across the failover, these settings give Windows clusters a seamless experience on QuantaStor HA pools. | |||
== Installing MPIO and claiming QuantaStor devices == | == Installing MPIO and claiming QuantaStor devices == | ||
| Line 33: | Line 33: | ||
Do this on every cluster node, before presenting QuantaStor volumes to it. | Do this on every cluster node, before presenting QuantaStor volumes to it. | ||
'''1. Install the Multipath I/O feature''' ( | '''1. Install the Multipath I/O feature''' (a reboot completes the installation): | ||
<pre style="font-size: smaller"> | <pre style="font-size: smaller"> | ||
| Line 39: | Line 39: | ||
</pre> | </pre> | ||
'''2. Add the QuantaStor hardware ID to the Microsoft DSM''' so that MPIO combines the paths to each QuantaStor volume into | '''2. Add the QuantaStor hardware ID to the Microsoft DSM''' so that MPIO combines the paths to each QuantaStor volume into a single disk: | ||
<pre style="font-size: smaller"> | <pre style="font-size: smaller"> | ||
| Line 46: | Line 46: | ||
</pre> | </pre> | ||
The PowerShell cmdlet pads the vendor and product IDs to their SCSI field widths. If you | The PowerShell cmdlet pads the vendor and product IDs to their SCSI field widths for you. If you prefer the MPIO control panel or {{Code|1=mpclaim}}, use the string {{Code|1=OSNEXUS QUANTASTOR}} followed by exactly six trailing spaces, as described in [[Multipath IO Configuration#Configuring Microsoft MPIO|Configuring Microsoft MPIO]]. | ||
'''3. Reboot''' the node so MPIO claims the devices | '''3. Reboot''' the node so MPIO claims the devices. Each QuantaStor volume then appears once in Disk Management and in {{Code|1=Get-Disk}}. | ||
For iSCSI, | For iSCSI, connect one session per path in the iSCSI Initiator with '''Enable multi-path''' selected, using addresses on separate subnets; see [[ISCSI Initiator Setup]] and [[Multipath IO Configuration]]. For an HA pool, connect to the pool's HA virtual interfaces so the sessions follow the pool when it fails over. | ||
== Recommended MPIO settings == | == Recommended MPIO settings == | ||
{| class="wikitable" | {| class="wikitable" | ||
! Setting !! Windows default !! Recommended !! | ! Setting !! Windows default !! Recommended !! What it does | ||
|- | |- | ||
| PDORemovePeriod || 20 || '''240''' || Seconds MPIO keeps a disk | | PDORemovePeriod || 20 || '''240''' || Seconds MPIO keeps a disk available while its paths are changing. 240 seconds comfortably covers an HA failover. | ||
|- | |- | ||
| DiskTimeoutValue || 60 || '''100''' || Seconds | | DiskTimeoutValue || 60 || '''100''' || Seconds Windows allows for each I/O. 100 is the largest value {{Code|1=Set-MPIOSetting}} accepts. | ||
|- | |- | ||
| PathVerificationState || Disabled || '''Enabled''' || MPIO | | PathVerificationState || Disabled || '''Enabled''' || MPIO checks every path periodically, so it picks up path changes and returning paths promptly. | ||
|- | |- | ||
| PathVerificationPeriod || 30 || 30 || Seconds between path | | PathVerificationPeriod || 30 || 30 || Seconds between path checks. | ||
|- | |- | ||
| RetryCount || 3 || 3 || Times MPIO retries | | RetryCount || 3 || 3 || Times MPIO retries an I/O on a path. | ||
|- | |- | ||
| RetryInterval || 1 || 1 || Seconds between those retries | | RetryInterval || 1 || 1 || Seconds between those retries. | ||
|} | |} | ||
Apply them in an elevated PowerShell on each cluster node: | Apply them in an elevated PowerShell session on each cluster node: | ||
<pre style="font-size: smaller"> | <pre style="font-size: smaller"> | ||
| Line 76: | Line 76: | ||
</pre> | </pre> | ||
'''Reboot the node | '''Reboot the node to activate the settings.''' In a running cluster, work through the nodes one at a time: drain the node (Pause, Drain Roles in Failover Cluster Manager, or {{Code|1=Suspend-ClusterNode -Drain}}), reboot it, resume it, and continue with the next node. The cluster stays available throughout. | ||
Confirm the result with {{Code|1=Get-MPIOSetting}}, which shows the active values: | |||
<pre style="font-size: smaller"> | <pre style="font-size: smaller"> | ||
| Line 93: | Line 93: | ||
</pre> | </pre> | ||
{{Code|1=Get-MPIOSetting}} is the best place to check these values, because {{Code|1=Set-MPIOSetting}} stores them through WMI rather than in the MPIO registry key. {{Code|1=Set-MPIOSetting}} validates all of its parameters together, so if it reports an error (for example a {{Code|1=-NewDiskTimeout}} above 100), correct the value and run the command again. | |||
The Fibre Channel HBA driver parameters can stay at the vendor defaults; this configuration was validated with the QLogic Windows driver at its defaults. | |||
== Paths and load balancing == | == Paths and load balancing == | ||
The Microsoft DSM's default load-balance policy works well with QuantaStor and needs no change. QuantaStor reports ALUA path states for HA pools: paths through the appliance that owns the pool are Active/Optimized and paths through the partner appliance are Standby. The Microsoft DSM's default for ALUA storage, Round Robin With Subset, spreads I/O across the Active/Optimized paths and moves to the other set when the pool fails over. | |||
View a disk's paths with {{Code|1=mpclaim}}: | |||
<pre style="font-size: smaller"> | <pre style="font-size: smaller"> | ||
| Line 108: | Line 108: | ||
</pre> | </pre> | ||
A host with two HBA ports zoned to both appliances of an HA pair shows four paths per disk: two Active/Optimized and two Standby. After a failover | A host with two HBA ports zoned to both appliances of an HA pair shows four paths per disk: two Active/Optimized and two Standby. After a failover the two sets swap roles, and the path count stays the same. | ||
== After maintenance on a QuantaStor appliance == | == After maintenance on a QuantaStor appliance == | ||
After an appliance has been rebooted or its Fibre Channel ports have been offline, a quick rescan on the Windows nodes makes sure all of its paths are back in use. Doing this before moving a pool back to that appliance ensures every node has its full set of paths ready for the move. | |||
<ol> | <ol> | ||
<li>Wait until the appliance is fully up, with its Storage Pools showing a Normal state in the web interface. QuantaStor brings an appliance's Fibre Channel target ports online once its LUNs are mapped and its ALUA states are set.</li> | |||
<li>On every cluster node, rescan storage: | <li>On every cluster node, rescan storage: | ||
<pre style="font-size: smaller"> | <pre style="font-size: smaller"> | ||
| Line 122: | Line 121: | ||
pnputil /scan-devices | pnputil /scan-devices | ||
</pre> | </pre> | ||
A {{Code|1=rescan}} in {{Code|1=diskpart}} does the same if you prefer it.</li> | |||
<li> | <li>Confirm with {{Code|1=mpclaim -s -d <n>}} that every QuantaStor disk shows its full path count.</li> | ||
<li> | <li>Move the pool back.</li> | ||
</ol> | </ol> | ||
== Troubleshooting == | == Troubleshooting == | ||
'''A cluster disk | '''A cluster disk did not stay online through a failover.''' Run {{Code|1=Get-MPIOSetting}} on every node and confirm the recommended values are active; they take effect after a reboot. Also confirm that [[Clustered SCSI-3 Persistent Reservations]] is Enabled on the pool's HA group. To bring the disk back, run {{Code|1=Update-HostStorageCache}} and a {{Code|1=diskpart}} rescan on each node, then bring the cluster disk online in Failover Cluster Manager ({{Code|1=Start-ClusterResource}}). | ||
'''A disk shows | '''A disk shows fewer paths than expected, or Standby paths only.''' Rescan as described in [[#After maintenance on a QuantaStor appliance|After maintenance on a QuantaStor appliance]]. Also check that both of the host's HBA ports are zoned to both appliances, and that both of its WWPNs are listed under the host in QuantaStor. | ||
''' | '''A QuantaStor volume appears more than once in Disk Management.''' MPIO has not yet claimed the device. Check {{Code|1=Get-MSDSMSupportedHW}} for the OSNEXUS QUANTASTOR entry, add it if needed, and reboot. | ||
'''Keep the appliances' Fibre Channel ports in | '''Fibre Channel port mode on the appliances.''' Keep the appliances' Fibre Channel ports in their default target-only mode for HA pools serving Windows clusters. Dual initiator and target mode ({{Code|1=qs-util enabledualmode}}) is intended for diagnostics. | ||
== Related pages == | == Related pages == | ||
* [[Clustered SCSI-3 Persistent Reservations]] -- the HA group setting Windows | * [[Clustered SCSI-3 Persistent Reservations]] -- the HA group setting that keeps Windows cluster reservations in place across failovers | ||
* [[Multipath IO Configuration]] -- MPIO and Linux multipath basics for QuantaStor volumes | * [[Multipath IO Configuration]] -- MPIO and Linux multipath basics for QuantaStor volumes | ||
* [[ISCSI Initiator Setup]] -- connecting iSCSI initiators | * [[ISCSI Initiator Setup]] -- connecting iSCSI initiators | ||
Latest revision as of 16:02, 6 October 2026
This page describes the recommended Microsoft Multipath I/O (MPIO) configuration for Windows Server hosts that use QuantaStor storage volumes as shared cluster disks, including Windows Server Failover Clustering, Cluster Shared Volumes, Hyper-V clusters and SQL Server Failover Cluster Instances. It applies to Fibre Channel and iSCSI access to a QuantaStor HA pool, and it is the client-side companion to Clustered SCSI-3 Persistent Reservations.
Summary. Install MPIO, let the Microsoft DSM claim QuantaStor devices, and apply the recommended MPIO timer settings below on every cluster node. With this configuration, an HA failover or an appliance reboot is a brief, transparent pause for the cluster: cluster disks stay online, Cluster Shared Volumes keep running, and Hyper-V virtual machines continue without interruption. The configuration was validated on Windows Server 2022 Hyper-V clusters connected over Fibre Channel to a QuantaStor HA pair.
| Section | Purpose |
|---|---|
| How the recommended settings help | How MPIO carries cluster disks through an HA failover |
| Installing MPIO and claiming QuantaStor devices | The MPIO feature and the QuantaStor hardware ID |
| Recommended MPIO settings | The timer values, and how to apply and verify them |
| Paths and load balancing | What a correctly configured disk looks like |
| After maintenance on a QuantaStor appliance | Confirming all paths are in use before moving pools back |
| Troubleshooting | Quick checks if a disk or path does not look as expected |
How the recommended settings help
During an HA failover, the pool moves from one QuantaStor appliance to its partner. For a short interval, while the receiving appliance imports the pool, I/O is held rather than completed. With Fibre Channel and ALUA, QuantaStor keeps the standby paths through the partner appliance present throughout, so Windows sees the paths change state and continues on the new active paths, typically within well under a minute.
MPIO decides how long to hold I/O and keep a disk while its paths change. The Windows defaults are tuned for general-purpose storage and allow only a short window (20 seconds before an unreachable disk is released, 60 seconds per I/O). The recommended settings extend that window so it comfortably covers an HA failover, including a slower one, for example when a busy pool takes longer to export. The cluster then experiences the failover as a short pause, and Cluster Shared Volumes and virtual machines carry on as normal.
Together with Clustered SCSI-3 Persistent Reservations, which keeps the cluster's disk reservations in place across the failover, these settings give Windows clusters a seamless experience on QuantaStor HA pools.
Installing MPIO and claiming QuantaStor devices
Do this on every cluster node, before presenting QuantaStor volumes to it.
1. Install the Multipath I/O feature (a reboot completes the installation):
Install-WindowsFeature -Name Multipath-IO
2. Add the QuantaStor hardware ID to the Microsoft DSM so that MPIO combines the paths to each QuantaStor volume into a single disk:
New-MSDSMSupportedHW -VendorId OSNEXUS -ProductId QUANTASTOR Get-MSDSMSupportedHW | Where-Object VendorId -match OSNEXUS
The PowerShell cmdlet pads the vendor and product IDs to their SCSI field widths for you. If you prefer the MPIO control panel or mpclaim, use the string OSNEXUS QUANTASTOR followed by exactly six trailing spaces, as described in Configuring Microsoft MPIO.
3. Reboot the node so MPIO claims the devices. Each QuantaStor volume then appears once in Disk Management and in Get-Disk.
For iSCSI, connect one session per path in the iSCSI Initiator with Enable multi-path selected, using addresses on separate subnets; see ISCSI Initiator Setup and Multipath IO Configuration. For an HA pool, connect to the pool's HA virtual interfaces so the sessions follow the pool when it fails over.
Recommended MPIO settings
| Setting | Windows default | Recommended | What it does |
|---|---|---|---|
| PDORemovePeriod | 20 | 240 | Seconds MPIO keeps a disk available while its paths are changing. 240 seconds comfortably covers an HA failover. |
| DiskTimeoutValue | 60 | 100 | Seconds Windows allows for each I/O. 100 is the largest value Set-MPIOSetting accepts.
|
| PathVerificationState | Disabled | Enabled | MPIO checks every path periodically, so it picks up path changes and returning paths promptly. |
| PathVerificationPeriod | 30 | 30 | Seconds between path checks. |
| RetryCount | 3 | 3 | Times MPIO retries an I/O on a path. |
| RetryInterval | 1 | 1 | Seconds between those retries. |
Apply them in an elevated PowerShell session on each cluster node:
Set-MPIOSetting -NewPathVerificationState Enabled -NewPathVerificationPeriod 30 -NewPDORemovePeriod 240 -NewRetryCount 3 -NewRetryInterval 1 -NewDiskTimeout 100
Reboot the node to activate the settings. In a running cluster, work through the nodes one at a time: drain the node (Pause, Drain Roles in Failover Cluster Manager, or Suspend-ClusterNode -Drain), reboot it, resume it, and continue with the next node. The cluster stays available throughout.
Confirm the result with Get-MPIOSetting, which shows the active values:
PS C:\> Get-MPIOSetting PathVerificationState : Enabled PathVerificationPeriod : 30 PDORemovePeriod : 240 RetryCount : 3 RetryInterval : 1 UseCustomPathRecoveryTime : Disabled CustomPathRecoveryTime : 40 DiskTimeoutValue : 100
Get-MPIOSetting is the best place to check these values, because Set-MPIOSetting stores them through WMI rather than in the MPIO registry key. Set-MPIOSetting validates all of its parameters together, so if it reports an error (for example a -NewDiskTimeout above 100), correct the value and run the command again.
The Fibre Channel HBA driver parameters can stay at the vendor defaults; this configuration was validated with the QLogic Windows driver at its defaults.
Paths and load balancing
The Microsoft DSM's default load-balance policy works well with QuantaStor and needs no change. QuantaStor reports ALUA path states for HA pools: paths through the appliance that owns the pool are Active/Optimized and paths through the partner appliance are Standby. The Microsoft DSM's default for ALUA storage, Round Robin With Subset, spreads I/O across the Active/Optimized paths and moves to the other set when the pool fails over.
View a disk's paths with mpclaim:
mpclaim -s -d # list MPIO disks mpclaim -s -d 1 # paths and their states for MPIO disk 1
A host with two HBA ports zoned to both appliances of an HA pair shows four paths per disk: two Active/Optimized and two Standby. After a failover the two sets swap roles, and the path count stays the same.
After maintenance on a QuantaStor appliance
After an appliance has been rebooted or its Fibre Channel ports have been offline, a quick rescan on the Windows nodes makes sure all of its paths are back in use. Doing this before moving a pool back to that appliance ensures every node has its full set of paths ready for the move.
- Wait until the appliance is fully up, with its Storage Pools showing a Normal state in the web interface. QuantaStor brings an appliance's Fibre Channel target ports online once its LUNs are mapped and its ALUA states are set.
- On every cluster node, rescan storage:
Update-HostStorageCache pnputil /scan-devices
Arescanindiskpartdoes the same if you prefer it. - Confirm with
mpclaim -s -d <n>that every QuantaStor disk shows its full path count. - Move the pool back.
Troubleshooting
A cluster disk did not stay online through a failover. Run Get-MPIOSetting on every node and confirm the recommended values are active; they take effect after a reboot. Also confirm that Clustered SCSI-3 Persistent Reservations is Enabled on the pool's HA group. To bring the disk back, run Update-HostStorageCache and a diskpart rescan on each node, then bring the cluster disk online in Failover Cluster Manager (Start-ClusterResource).
A disk shows fewer paths than expected, or Standby paths only. Rescan as described in After maintenance on a QuantaStor appliance. Also check that both of the host's HBA ports are zoned to both appliances, and that both of its WWPNs are listed under the host in QuantaStor.
A QuantaStor volume appears more than once in Disk Management. MPIO has not yet claimed the device. Check Get-MSDSMSupportedHW for the OSNEXUS QUANTASTOR entry, add it if needed, and reboot.
Fibre Channel port mode on the appliances. Keep the appliances' Fibre Channel ports in their default target-only mode for HA pools serving Windows clusters. Dual initiator and target mode (qs-util enabledualmode) is intended for diagnostics.
Related pages
- Clustered SCSI-3 Persistent Reservations -- the HA group setting that keeps Windows cluster reservations in place across failovers
- Multipath IO Configuration -- MPIO and Linux multipath basics for QuantaStor volumes
- ISCSI Initiator Setup -- connecting iSCSI initiators
- Fibre Channel Target Port Management -- presenting volumes over Fibre Channel
- Hosts and Host Groups -- host records, initiators, and Host Groups for clusters
- HA Cluster Setup (JBODs) -- HA groups and failover