Guides:Asynchronous Remote Replication and Disaster Recovery Failover
By Steve Umbehocker, CTO, OSNexus · Updated October 3, 2026
Asynchronous remote replication keeps a second, mountable copy of your volumes and shares at another site, updated on a schedule by sending only the blocks that changed. QuantaStor replicates Storage Volumes and Network Shares between Storage Pools on systems in the same grid. When the primary site is lost, you activate the replica checkpoints at the DR site, and later you can fail back with or without the changes made.
Why it matters
Snapshots protect against mistakes, but they live in the same pool as the data, so they go down with the pool, the system or the site. Replication puts a full copy somewhere else. The three scheduled data-protection features answer different questions:
| Feature | Where the copy lands | Use it for |
|---|---|---|
| Snapshot Schedules | Destination is the same Storage Pool | Fast local recovery points, undoing a bad change |
| Remote replication | Destination is a Storage Pool on another system within the grid | Surviving the loss of a system, pool or site |
| Backup Policies | Destination is cloud object storage or an an external NFS/SMB source | Provides great flexibility and you can do inbound or outbound policies and autotiering but it doesn't have the block-level incremental efficiency of the other methods |
A replication schedule already snapshots its sources and keeps its own retention on both sides, so you rarely need a separate Snapshot Schedule on the same volume or share.
How it works
Replication is snapshot-driven ZFS send/receive between two pools. The full reference is Remote-replication (DR). The moving parts:
- Grid membership. Both systems must belong to the same storage grid. A grid can span sites, so replicating to a remote site means joining the remote appliances to this grid. See Grid Configuration.
- Storage System Replication Link. A trust relationship and data path between two systems. Creating one generates a dedicated SSH key pair, registers each side's public key on the other, and pins replication traffic to a fixed IP address on each end. Links come in pairs: A to B creates B to A as well.
- Replication schedule. Names a link (which sets the direction), a destination pool, the volumes and shares, when to run, and how many snapshots to keep on each side. The system that owns the destination pool owns and runs the schedule.
- Replica checkpoint. For each source, the destination pool holds a real volume or share named with a
_chkpntsuffix (backupsbecomesbackups_chkpnt), plus timestamped snapshots that serve as retained recovery points. - Replica association. The persistent source-to-checkpoint relationship, where live status and sync times are recorded.
Only the first run is a full copy. After that, each run snapshots the source, finds the newest GMT-stamped snapshot that both source and checkpoint have, and sends the difference. The pools don't need to match in size, layout, disk type or hardware, and the remote-replication feature must be included in the license.
On a Ceph scale-out cluster, Ceph RBD pools can be a destination only when the remote system belongs to the same Ceph cluster as the source. Network Shares on scale-out CephFS replicate to a second Ceph cluster through CephFS snapshot mirroring instead; see Scale-out File Replication.
Design and sizing
RPO. On the Schedule Interval tab, a timer interval is the rest time between the end of one run and the start of the next: 3 minutes minimum, 30 by default, up to 360. Your worst-case data loss is about one interval plus the duration of a run, so the real lever is how long a run takes. A run never starts while the previous one is still working. For a fixed timetable, the day/hour grid gives a calendar schedule, with an offset of 0–59 minutes to stagger several schedules. Frequent small runs usually transfer less each time and keep the DR copy closer than one big nightly run.
Bandwidth. Each link has a Bandwidth Limit in MB/sec (the dialog suggests 200; entering 0 means the service default of 100). The limit is shared evenly across the streams currently running on that link and rebalanced as streams finish, so four concurrent replications on a 200 MB/sec link get 50 MB/sec each. Size the limit against your daily change rate. This is plain arithmetic, not a measured benchmark:
| Changed data per run | Time to send at 100 MB/sec | Time to send at 200 MB/sec |
|---|---|---|
| 10 GB | about 1.7 minutes | about 50 seconds |
| 100 GB | about 17 minutes | about 8.5 minutes |
| 1 TB | about 2.8 hours | about 1.4 hours |
The first full copy is the long one. Seeding with a one-time replica (below) before you enable a tight schedule keeps the initial copy from colliding with production hours.
Transport. Encryption is on by default and tunnels the stream over SSH. Turning it off sends the stream over a plain mbuffer TCP session, which is faster on a trusted private link and offers no confidentiality. Compression adds lz4 only when the source dataset isn't already compressed.
Retention. The Snapshot Settings tab keeps 3 short-term snapshots by default (the recommended value) for computing deltas, plus long-term hourly, daily, weekly, monthly and quarterly counts set separately for source and checkpoint. Keep enough snapshots that both sides always share one; otherwise the next run falls back to a full copy.
HA pools. Each run re-selects its link based on which systems currently own the source and destination pools. With an HA pool on systems A and B replicating to one on C and D, create all four links (A–C, A–D, B–C, B–D) so replication keeps running after a failover on either end.
Setting it up
- Join both systems to one grid (Grid Configuration) and create a destination pool (Storage Pools).
- Create the link: Remote Replication → Storage System Replication Links → Replication Link → Create. Choose the storage system and a fixed IP address for each side. Floating cluster VIF addresses are filtered out. For a cloud system with a public address, the far side can use External IP Address, which is pre-filled from the system's external hostname (set with
qs system-modify --ext-hostname, see the QuantaStor CLI Command Reference).

- Create the schedule: Remote Replication → Volume & Share Replication Schedules → Replication Schedule → Create. On the General tab, pick the link (this sets the direction) and the Remote Pool. Then set the interval, select volumes and shares (selecting a parent share and one nested inside it is rejected), set retention, and review Advanced Settings. Leave Enable resumable replication and Enable target checkpoint recovery on, as they are by default.
- To seed a destination or copy one volume on demand without a schedule, use Create Volume Replica. The diff-copy option sends changes against an existing checkpoint; full copy creates a new one.
- Interval schedules start shortly after you create them. To run any schedule immediately:
qs replication-schedule-trigger --schedule=nightly-dr qs replica-assoc-list qs replica-report-summary-list
Operating and testing
Watch the Volume & Share Replica Associations section. Each source shows its status (Synchronizing, Synchronized, Sync Failed (Resumable), Skipped and so on), elapsed time, and when the last sync started and completed. Each run also writes a summary report with bytes transferred and average speed, which is the number to check against your RPO math.


If a run is interrupted (the WAN drops, a node reboots), resumable replication keeps a ZFS resume token and the next run continues from where it stopped. A failed run raises an alert, at most one per hour per schedule and source; see Call-home / Alerting for sending alerts off the appliance. A run is skipped, with its reason shown on the schedule, if a link is missing or not Normal, if a pool is unhealthy, or if a destination HA failover is in progress.
Failover. Activate Checkpoints on the schedule toolbar is the promotion step. It disables the schedule, marks the checkpoints as Active Replica Checkpoints, brings checkpoint shares online and makes checkpoint volumes available to map to hosts. It can also create share aliases named after the sources, so clients reach backups rather than backups_chkpnt. Optional dr-prefailover and dr-postfailover scripts in /var/opt/osnexus/custom/ handle site-specific steps such as repointing DNS. For hands-off failover, the Automatic Activation tab activates checkpoints when a Site Cluster VIF (Site Cluster Create) moves to the destination node.

Failback. After a test, disconnect DR-site clients, then use Deactivate Checkpoints with Re-enable Replication Schedule ticked; the next run overwrites the checkpoints from the source. If the DR site took writes you need to keep, use Rollback first (qs replication-schedule-trigger-rollback --schedule=<name>). It sends the checkpoint's changes back and overwrites the original source. Then deactivate and re-enable.
Test it regularly. A DR test is activate, mount and verify from a DR-site client, then deactivate. Any client session on a checkpoint marks it active again on its own, and while any checkpoint is active the schedule stays blocked. A test that ends with clients still attached quietly stops replication, so check the schedule's state afterwards.
FAQ
Does QuantaStor support asynchronous replication with DR failover between sites?
Yes. Replication schedules copy Storage Volumes and Network Shares asynchronously and incrementally to a Storage Pool on another grid member, which can be at another site. Failover is the Activate Checkpoints step, either manual or automatic on a Site Cluster VIF move. Failback either discards the DR-site changes or rolls them back to the source first.
What RPO can I get?
The shortest timer interval is 3 minutes, measured from the end of one run to the start of the next. Your effective RPO is that interval plus however long a run takes, which depends on change rate and the link's bandwidth limit. Check the average speed in the replication reports to see what you're actually achieving.
Do the source and destination systems need identical hardware?
No. The pools don't need to match in size, layout, disk type or hardware. Both systems do need to be in the same grid and licensed for remote replication. Ceph RBD destinations must be in the same Ceph cluster as the source.
Can one source replicate to more than one site?
Yes. Create one schedule per destination (N-way), or replicate from the DR site's checkpoint to a third site (cascading). On the second and later schedules, enable source snapshot reuse so all destinations share a common snapshot.
Will replication overwrite changes made at the DR site?
Not while a checkpoint is active. An active checkpoint blocks the schedule, and any iSCSI, FC, NFS or SMB client session marks a checkpoint active automatically. Overwriting happens only after you deactivate, so roll back first if you want to keep DR-site writes.
Part of the QuantaStor Guides series. For reference documentation, see the QuantaStor documentation.