Site Cluster Setup

From OSNEXUS Online Documentation Site
Jump to navigation Jump to search

A site cluster is the group of appliances at one location that run the cluster software underneath QuantaStor's high-availability features. This page covers building one: what has to be in place first, how to create it and its heartbeat rings, how to add and remove members, what corosync and pacemaker configuration it lays down, and how to take it apart again. For what a cluster VIF does once the site cluster exists -- what makes one fail over, how long that takes and what blocks it -- see High-availability VIF Management.

Section Purpose
Why a site cluster exists What the layer is for, and what depends on it
Before you start Grid membership, the shared subnet, and what creating one destroys
Creating a site cluster The dialog, the CLI, and the stages the task runs
Heartbeat rings The automatic first ring, adding a second, and the two-ring maximum
Adding and removing members Growing and shrinking the cluster, and the two-member floor
What the site cluster configures The corosync and pacemaker settings, explained
Rescan and restart cluster services Two repair tools, and which one to reach for
Standby and maintenance mode Taking a member or a whole cluster out of service
Deleting a site cluster The teardown, and what it moves aside
Troubleshooting A cluster that will not form, or will not go healthy
Command line reference Every site cluster and cluster ring command

Why a site cluster exists

The High-availability VIF Management tab with a site cluster selected. The Site Cluster toolbar group carries the create, ring, modify, rescan, delete and restart operations; the right-click menu adds the four that have no toolbar button -- Add Nodes, Remove Nodes, Configure Member Standby, and Enter/Exit Maintenance Mode.

QuantaStor's floating addresses are pacemaker resources. Something has to decide, continuously and across appliances, which appliance is alive and where a resource should run. That is what a site cluster is: a set of grid members that run corosync and pacemaker together, exchanging heartbeat traffic over one or two dedicated network paths.

Nothing above it works without it. A cluster VIF belongs to a site cluster rather than to a pool or a Ceph cluster, so a scale-up HA failover group cannot present a floating service address until the site cluster is in place, and a scale-out Ceph configuration cannot either. Creating the site cluster is therefore the first step of an HA build, not a later refinement.

The word "site" is meant literally: the members are appliances at one location. That location can span buildings, but they have to share a network, because a heartbeat ring is a subnet. A grid can hold several site clusters, and an appliance can belong to at most one of them.

Navigation: High-availability VIF Management → Site Clusters

Selecting a site cluster in the tree shows three grids. Site Clusters is the cluster itself, with its state, the appliance that manages the object, the Designated Coordinator -- pacemaker's elected leader for the cluster -- and the pacemaker version. Site Cluster Members has one row per member, with the number of systems that member sees configured and online, the resources it is running, the cluster stack, and its standby mode. Cluster Heartbeat Ring Ports has one row per member address per ring, with the ring's network.

Before you start

Every appliance must already be a grid member. A site cluster is built from systems in one storage grid, and create fails outright if there is no grid: No grid configuration detected, a grid of two or more nodes must be created before.... Build the grid first -- see Grid Configuration.

You need at least two appliances, and you must run the operation from one of the appliances that will be a member. The task refuses a list of one, and it refuses a list that does not include the system you are talking to.

Every appliance must contribute a port on the same network. The addresses you pick for the first ring are masked with their subnet masks and must all resolve to the same network address. If they do not, create fails with Invalid ring configuration, not all specified ports have the same subnet mask. Ports that are VLAN children, virtual interfaces, floating addresses, or unconfigured are not offered.

An appliance can only be in one site cluster. Systems that already belong to one are filtered out of the create dialog entirely, so on a grid where every member is already clustered the dialog does not open at all -- it reports There are no storage systems to select ports from. That message means the grid has no free appliance, not that anything is wrong with the network.

All existing members must be reachable. Adding or removing a member pings every current member first and stops if one does not answer: All Site Cluster Members must be online to add or remove Site Cluster Members.

Creating a site cluster erases any cluster configuration already on those appliances

This is the one thing to read before clicking OK. As part of the create, QuantaStor runs a flush stage on every selected appliance that stops corosync and pacemaker and moves the existing cluster state out of the way:

  • /etc/corosync/corosync.conf is moved to /tmp/corosync.conf@GMT-<date>-
  • the contents of /var/lib/heartbeat/crm/, the ringid_* files in /var/lib/corosync/, the cib* files in /var/lib/pacemaker/cib/ and the pe* files in /var/lib/pacemaker/pengine/ are all moved into /tmp/crm.backup@GMT-<date>-

The files are moved rather than deleted, so they survive until the appliance is rebooted or /tmp is cleaned -- but they are not a restore path, and nothing puts them back. Do not point a site cluster create at an appliance that is running a pacemaker cluster you care about, whether QuantaStor built it or not.

Creating a site cluster

Navigation: High-availability VIF Management → Site Clusters → Create Site Cluster (toolbar)

The Create Site Cluster Configuration dialog has four things in it:

  • Name -- required, and pre-filled with a generated name such as site-cluster-1. Letters, digits and -_. only; no spaces, under 40 characters.
  • Location/Region -- optional free text describing where the appliances are.
  • Description -- optional free text.
  • Select Cluster Ring Interfaces -- a table with one row per eligible appliance, each row a checkbox plus a combo listing that appliance's usable ports as <ip> (<port name>). No rows are ticked when the dialog opens, so tick every appliance that is to be a member and confirm the port chosen for each. The addresses you pick here become the first heartbeat ring.

Only the systems that are grid members and not already in a site cluster appear. Tick at least two.

The equivalent command takes the member addresses as a comma-separated list, one address per appliance:

qs site-cluster-create --name=hq-cluster --location="Building 4" \
   --system-ip-list=10.0.8.110,10.0.8.111,10.0.8.112

That is qs site-cluster-create; --desc sets the description and --flags=async returns immediately instead of waiting.

What the create task does

Create is a coordinated task: the appliance you run it on drives five stages, and every stage runs on every member before the next one begins. The task detail names the stage it is in, which is worth watching if one hangs.

Stage What happens
Starting cluster configuration update Cluster service management is suspended on each member so the monitor does not fight the reconfiguration
Clearing existing cluster configuration settings The flush described above -- services stopped, old configuration and CIB moved to /tmp
Updating cluster configuration /etc/corosync/corosync.conf is generated on every member from the ring definition
Starting site cluster configuration corosync and then pacemaker are started on every member
Finishing cluster configuration update Service management resumes; the origin appliance then sets the cluster-wide pacemaker properties

Each stage has its own 180-second timeout. A stage that does not complete everywhere within that window fails the task with a message naming the stage, which is generally more useful than the overall failure. Creating a three-appliance cluster on an idle lab grid completes well inside one stage's budget.

Two useful consequences of the design. The site cluster's identifier is derived from the set of member system IDs, and it is written into the generated configuration as totem { siteid: ...} -- which is how an appliance that has lost its database can still work out which site cluster it belongs to. And corosync node IDs are derived from each appliance's system UUID rather than allocated in sequence, so they are stable across a rebuild and are not 1, 2, 3.

Confirming the cluster is actually up

The create response, and the Site Cluster Members grid immediately after it, can report figures that are still catching up. Confirm the cluster from the cluster software itself, on any member:

sudo corosync-quorumtool
sudo crm_mon -1

You want Quorate: Yes with the node count you expect from the first, and every member listed under Online: in the second. In the web interface the equivalent is every row of Site Cluster Members green, showing the full member count under Storage Systems Online/Offline and corosync under Stack.

Heartbeat rings

Add Cluster Ring. Every member is ticked by default, but the port combos here are empty, because on this cluster each appliance has only one interface and its address is already carrying ring 0. A second ring needs a second interface on every member.

A heartbeat ring is one network path that the members use to tell whether each other are alive. It carries cluster gossip only -- no client traffic ever runs over it. Each member contributes exactly one address to a ring, all of a ring's addresses must be on the same network, and the ring is named after that network: three members on 10.0.8.x/16 give a ring named 10.0.0.0.

The first ring is created for you from the addresses you chose in the create dialog. You never create ring 0 by hand.

A second ring is strongly recommended, and it must be on a different network. With one ring, losing that network looks exactly like losing the appliance -- the surviving members conclude the node is dead and try to take its resources, which is the failure mode the duplicate-address check on High-availability VIF Management exists to catch. With two rings on independent networks, a single network failure is visible as a degraded ring rather than a dead node.

Navigation: High-availability VIF Management → Site Clusters → Add Cluster Ring (toolbar)

The Add Cluster Ring dialog asks for the Site and then presents the same interface table as create, with every member pre-ticked. What it will not let you do is reuse an address: any port whose IP is already carrying a ring is removed from that appliance's combo. If every member has only one usable interface, the combos come up empty and there is nothing to select -- which is the dialog telling you, indirectly, that you need a second NIC on each member. From the command line the same constraint appears as an explicit error:

ERROR: Cluster ring member '10.0.8.110' already exists which is associated
with site cluster 'b19ca864-a5c1-aa8f-f380-6578d3a82acb'

The command is qs cluster-ring-create --site=<site> --member-addresses=<ip,ip,...>. You must supply one address for every member of the cluster -- a partial ring is rejected. --bind-address is optional; leave it off and the network address is computed from the first member's address and mask, and every other member is then required to agree with it. --ring is accepted but not honoured: the ring number is always the next one in sequence.

Two rings is the maximum and one is the minimum. A third is refused with Site '<name>' already has two (2) cluster rings which is the maximum., and removing the last one is refused with Cannot delete ring for site '<name>', each site must have one cluster ring configured.

Navigation: High-availability VIF Management → Site Clusters → Remove Cluster Ring (toolbar)

Remove Cluster Heartbeat Ring takes the ring, shows its site cluster, bind address and ring number read-only, and warns that removing it can affect the ability of high-availability resources to fail over. Both ring operations restart the cluster services on every member as they apply, so make the change when a brief cluster reconfiguration is acceptable.

Adding and removing members

Remove Nodes from Site Cluster. Add Nodes is the same dialog against the eligible non-member systems. Nothing is ticked when it opens, and Force is cleared.

Both operations are on the right-click menu only -- there is no toolbar button for either.

Navigation: High-availability VIF Management → Site Clusters → Site Cluster (select + right-click) → Site Cluster Add Nodes...

Add Nodes lists the grid members that are not already in a site cluster and that have an interface on every one of this cluster's ring networks. An appliance with no port on a ring's subnet simply does not appear in the list -- there is no error explaining the omission, so if a system you expect is missing, check its addressing against the Cluster Heartbeat Ring Ports grid first. On a two-ring cluster the candidate needs a port on both networks.

You do not choose ring addresses when adding a member. QuantaStor picks the port whose network matches each existing ring and joins the appliance to every ring automatically. If it cannot find one, the add fails naming the ring it could not match.

Navigation: High-availability VIF Management → Site Clusters → Site Cluster (select + right-click) → Site Cluster Remove Nodes...

Remove Nodes lists the current members. A Force checkbox is available and cleared by default. Three things will stop a removal:

  • A site cluster must keep at least two members. Selecting enough systems to leave one is rejected before anything happens: Too many Site Cluster Members specified. There were '1' number of Site Cluster Members specified to be removed from a Site Cluster with '2' current Site Cluster Members.
  • The appliance must not be hosting a cluster VIF. Move its VIFs off first -- see High-availability VIF Management.
  • The appliance must not be the primary or secondary of an HA failover group.

A removal is more involved than an addition, because the cluster has to forget the departing node rather than just stop talking to it. The removed appliance gets the same configuration flush as a create, the surviving members get a rewritten corosync.conf and a service restart, and the designated coordinator then runs crm_node -R for each departed node so pacemaker stops expecting it. Both operations are coordinated tasks with the same 180-second-per-stage budget. On a three-appliance lab grid, adding a member took about 48 seconds and removing one about 56.

The member count changes how corosync counts votes. Going from three members to two rewrites quorum { two_node: 0} to two_node: 1, and corosync-quorumtool then reports Flags: 2Node Quorate WaitForAll with a quorum of 1 instead of 2. Growing back to three reverses it. This is corosync's standard handling for a two-node cluster and needs no action, but it is worth recognising, because a two-member cluster genuinely behaves differently from a three-member one during a network partition.

There is no maximum member count enforced anywhere in the product.

What the site cluster configures

If you have never administered pacemaker, this section is the part worth reading, because two of the settings QuantaStor chooses explain behaviour that otherwise looks like a fault.

The corosync configuration file

/etc/corosync/corosync.conf is generated and owned by QuantaStor on every member. It opens with a header saying so, and hand edits to it may be overwritten. Two related files sit beside it: corosync.conf.QSDEFAULT is the shipped template the generator starts from, and corosync.conf.save is a copy of the previous version, written each time the file is regenerated.

What the generated file contains that is worth knowing:

Setting Meaning
nodelist with one node block per member Each carries ring0_addr (and ring1_addr with a second ring), the appliance name, and a nodeid derived from its system UUID
totem { siteid: ...} The site cluster's own identifier, so a member can recognise its cluster from the file alone
totem { transport: ...} knet on current systems; udpu if any grid member is older than 5.12, since the whole grid has to agree
totem { token: 3000} and token_retransmits_before_loss_const: 10 How long a member can go silent before it is declared lost -- normally a few seconds
totem { rrp_mode: ...} none with one ring, active with two
totem { secauth: off} Heartbeat traffic is not encrypted or authenticated, which is why a ring belongs on a trusted network
quorum { two_node: ...} 1 at exactly two members, 0 otherwise

Nothing about a site cluster changes the firewall. QuantaStor opens no ports for corosync and closes none, and there is no corosync entry in its firewall configuration at all. On a current build corosync uses knet's default UDP port 5405, bound to each member's ring address. If your heartbeat network passes through a firewall you control, that traffic is yours to allow -- QuantaStor will not do it and will not warn you.

The pacemaker settings, and why fencing is off

Once the services are up, the appliance that ran the create sets three cluster-wide properties. You can see them with sudo cibadmin --query --scope crm_config and sudo cibadmin --query --scope rsc_defaults.

  • resource-stickiness=1000 -- a strong preference for leaving a running resource where it is. This is why adjusting a VIF's location weights in the web interface does not relocate a VIF that is already running; the weights decide where it lands next time. High-availability VIF Management covers that in detail.
  • no-quorum-policy=ignore -- a member that finds itself in a minority partition does not stop its resources. Combined with the point below, this means QuantaStor's answer to a partition is not to shut anything down pre-emptively.
  • stonith-enabled=false -- pacemaker will not fence a member it has lost contact with.

That last one deserves an honest paragraph, because in a textbook pacemaker deployment turning fencing off is a mistake. Fencing exists to answer one question: when the cluster loses contact with a node, is that node dead, or is it alive and still holding the resources? Fencing settles it by force -- power the node off, and then it is definitely safe to start its resources elsewhere. That requires out-of-band hardware (a manageable PDU, an IPMI or BMC path) that QuantaStor cannot assume is present.

Instead of fencing, QuantaStor puts the check at the moment it matters. Its own copy of the IPaddr2 resource agent will not bring a floating address up on a member until it has flushed the ARP cache and confirmed by probe that nothing else is answering for that address. If something answers, it refuses and says so. The trade is deliberate and it is not free: a member that loses cluster communication but stays running keeps its addresses, the survivors decline to steal them, and the VIF stops being managed anywhere until the stranded member rejoins -- but it never goes off the air, and the address is never up twice. For scale-up pools there is a second, independent layer: SCSI-3 persistent reservations on the shared disks stop two appliances importing one pool. High-availability VIF Management documents the probe, the exact refusal messages, and how to recover a stranded member.

Ceph is not affected by any of this. The corosync and pacemaker layer sits beside a Ceph cluster, not underneath it -- Ceph runs its own monitor quorum. Building or tearing down a site cluster on Ceph members leaves Ceph health, monitor quorum and OSD counts unchanged. What a site cluster adds for scale-out is the ability to give a Ceph service a floating address.

Rescan and restart cluster services

These two are easy to confuse, and only one of them touches the running cluster.

Navigation: High-availability VIF Management → Site Clusters → Rescan Site Cluster (toolbar)

Rescan re-reads the live cluster configuration and rebuilds QuantaStor's record of the rings, the members, the VIFs and their location constraints from it, treating pacemaker as authoritative. It stops and starts nothing, it is quick -- a few seconds on a three-member cluster -- and it is the right first move whenever the web interface and crm_mon -1 disagree. It also clears a state that has gone stale: after a member was added, a cluster that was fully quorate went on reporting Warning with Problem detected with heartbeat cluster, one or more nodes are offline. until a rescan returned it to Normal.

The command takes no arguments. Run it on a member; it acts on that appliance's own site cluster.

qs site-cluster-rescan

That is qs site-cluster-rescan.

Restart Site Cluster Service Components. The dialog is explicit that a restart can trigger a failover, and that it is for an unhealthy member only.
Navigation: High-availability VIF Management → Site Clusters → Restart Cluster Services (toolbar)

Restart Cluster Services stops and restarts corosync and pacemaker on one appliance. It is a repair for a member that is genuinely wedged -- one that is not in the cluster's membership, or whose services have died -- and it is not a refresh. The dialog carries the reason in bold: restarting the service can trigger a failover if a VIF is running on that system, so it should only be used on a member that is in an unhealthy state.

qs site-cluster-restart-services --storage-system=qs-node-112

That is qs site-cluster-restart-services. The dialog's Storage System list offers every appliance in the grid, not only the members of a site cluster, so check you have picked a member. Only corosync and pacemaker are restarted, in that order; pcsd is not touched, and QuantaStor does not manage it.

So: rescan when QuantaStor's view looks wrong, restart services when the cluster's view looks wrong.

Standby and maintenance mode

Both take part of the cluster out of normal operation, and they are not interchangeable.

Standby mode is per member. It stops one appliance from running resources while leaving it in the cluster, which is what you want before working on that appliance. It is set from Configure Member Standby... on the site cluster's right-click menu, and it has three settings -- see Configure Member Standby.

Maintenance mode is per site cluster. It sets pacemaker's cluster-wide maintenance-mode property, so the cluster stops managing resources entirely: nothing is monitored, nothing recovers, and nothing moves. Existing VIFs stay where they are and keep serving traffic, but they are no longer protected. It is reached from Enter Maintenance Mode... and Exit Maintenance Mode... on the same menu -- only the applicable one is shown. Its effects on cluster VIFs are documented on High-availability VIF Management.

One practical note on maintenance mode: QuantaStor's own service monitor stops managing corosync and pacemaker while it is on, but only for fifteen minutes. Past that the monitor resumes -- deliberately, so that a cluster cannot be left permanently wedged by a maintenance flag nobody cleared. Maintenance mode itself stays on until you exit it. It is not designed to be left on for hours, and the dialog says as much.

Deleting a site cluster

Delete Site Cluster Configuration. Location/Region and Description are shown read-only so you can confirm which cluster you have selected.
Navigation: High-availability VIF Management → Site Clusters → Delete Site Cluster (toolbar)

Deleting a site cluster removes its heartbeat rings and its cluster configuration from every member. The dialog asks for the Site, shows its location and description read-only, offers a Force checkbox, and then confirms.

Remove the cluster VIFs first. Without Force, the delete refuses while any VIF still references the site cluster, and names the count:

ERROR: Selected site cluster 'hq-cluster' has (2) associated VIF resources.
To prevent network outage, convert VIFs to local virtual IP addresses before
deleting the site cluster with the force flag set.

The refusal is protecting a service address, not the cluster object. The intended route is the Convert cluster VIF resource to local virtual IP option in Remove Cluster VIF, which keeps each address in service on the appliance it was last running on -- see High-availability VIF Management. Forcing the delete instead takes those addresses off the network.

The teardown itself runs the same configuration flush as a create, on every member: services stopped, /etc/corosync/corosync.conf moved to /tmp/corosync.conf@GMT-<date>-, and the CIB, pengine and ring-id state moved into /tmp/crm.backup@GMT-<date>-. Delete then goes one step further than create and verifies the flush before it removes anything from the database: if corosync.conf is still present, or any ringid_*, cib* or pe* file remains, the delete fails with Site Cluster delete failed, failed to flush Site Cluster config files and the site cluster stays. A delete that fails that way is almost always a file that could not be moved -- a busy service or a permissions problem on one member -- so check that corosync and pacemaker really did stop there before retrying.

qs site-cluster-delete --site=hq-cluster

That is qs site-cluster-delete; add --flags=force to delete with VIFs still attached.

Troubleshooting

The create dialog will not open. There are no storage systems to select ports from. means no grid member is free -- every one is already in a site cluster. Check with qs site-cluster-assoc-list.

Create fails on the subnet check. Invalid ring configuration, not all specified ports have the same subnet mask means the addresses you picked do not all mask to one network. Confirm each appliance's port address and mask before retrying.

Create fails at a named stage. The task detail names the stage that timed out or errored. A failure at the flush stage usually means corosync or pacemaker would not stop on one member; a failure at the start-services stage means they would not come up. Check the member named in the task, then retry.

The cluster forms but a member never comes online. Check the heartbeat network between the members before anything else -- with one ring there is no second path, so a single switch port or VLAN mistake keeps a member out permanently. sudo corosync-quorumtool on each member shows that member's own view of the membership; comparing them tells you quickly whether the problem is symmetric.

An appliance you want to add is not in the Add Nodes list. Either it is already in a site cluster, or it has no interface on one of this cluster's ring networks. There is no message for the second case, so compare its ports against the Cluster Heartbeat Ring Ports grid.

The state says Warning but the cluster looks fine. Compare against sudo crm_mon -1. If pacemaker reports every member online and quorate, QuantaStor's record has gone stale; run a rescan.

A member's own row shows zero systems configured. Immediately after a rescan the local appliance's row in Site Cluster Members can briefly show 0 configured and 0 online while the other members show the real figures. The site cluster's own state stays correct throughout and the row repopulates on the next monitor pass. If it does not clear within a minute or so, the monitor is being suppressed -- check whether the cluster is still in maintenance mode.

Command line reference

Every one of these is documented in full in the QuantaStor CLI Command Reference.

Command Purpose
qs site-cluster-create --name --system-ip-list Create the cluster and its first heartbeat ring
qs site-cluster-list / qs site-cluster-get --site List the site clusters, or show one in full
qs site-cluster-modify --site Change the name, location or description
qs site-cluster-add-nodes --site --system-list Add members
qs site-cluster-remove-nodes --site --system-list Remove members
qs site-cluster-rescan Rebuild QuantaStor's view from the live cluster (no arguments)
qs site-cluster-restart-services --storage-system Restart corosync and pacemaker on one appliance
qs site-cluster-set-standby-mode --site --storage-system --standby-mode Put a member into or out of standby
qs site-cluster-toggle-maintenance-mode --site --enable-maintenance-mode Enter or leave cluster maintenance mode
qs site-cluster-delete --site Delete the cluster and flush its configuration
qs site-cluster-assoc-list / qs site-cluster-assoc-get --system Which appliance belongs to which site cluster
qs cluster-ring-create --site --member-addresses Add a second heartbeat ring
qs cluster-ring-delete --cluster-ring Remove a heartbeat ring
qs cluster-ring-list / qs cluster-ring-get --cluster-ring List the rings, or show one in full
qs cluster-ring-member-list / qs cluster-ring-member-get --cluster-ring-member The individual addresses in each ring

Related pages


Verified against QuantaStor 6.9.0.