<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://wiki.osnexus.com/index.php?action=history&amp;feed=atom&amp;title=Ceph_over_ZFS</id>
	<title>Ceph over ZFS - Revision history</title>
	<link rel="self" type="application/atom+xml" href="https://wiki.osnexus.com/index.php?action=history&amp;feed=atom&amp;title=Ceph_over_ZFS"/>
	<link rel="alternate" type="text/html" href="https://wiki.osnexus.com/index.php?title=Ceph_over_ZFS&amp;action=history"/>
	<updated>2026-10-08T02:20:33Z</updated>
	<subtitle>Revision history for this page on the wiki</subtitle>
	<generator>MediaWiki 1.42.1</generator>
	<entry>
		<id>https://wiki.osnexus.com/index.php?title=Ceph_over_ZFS&amp;diff=28230&amp;oldid=prev</id>
		<title>Qadmin: osn-seo-utilities: docs-ceph-over-zfs-zfs-draid-pool-reserved @ d389deda0b81 (approved in the portal)</title>
		<link rel="alternate" type="text/html" href="https://wiki.osnexus.com/index.php?title=Ceph_over_ZFS&amp;diff=28230&amp;oldid=prev"/>
		<updated>2026-10-08T00:00:04Z</updated>

		<summary type="html">&lt;p&gt;osn-seo-utilities: docs-ceph-over-zfs-zfs-draid-pool-reserved @ d389deda0b81 (approved in the portal)&lt;/p&gt;
&lt;p&gt;&lt;b&gt;New page&lt;/b&gt;&lt;/p&gt;&lt;div&gt;[[Category:admin_guide]]&lt;br /&gt;
Ceph over ZFS runs Ceph OSDs on ZFS Storage Volumes (zvols) instead of directly on Physical Disks. There are a number of benefits of this architecture but the biggest is the double layer durability localizes the rebuild from failed drives so that the Ceph cluster doesn&amp;#039;t need to absorb that load and production workload impact.  The setup is simple, on each node you provision a Scale-up (ZFS) Storage Pool, ideally in a declustered RAID (dRAID) layout when you have 40+ drives per node.  These pools become the backing store for Ceph OSDs, which become one-to-one mapped to Storage Volumes (zvols).  So after the pool is created you&amp;#039;ll batch create something like 16x, 24x, or 32x Storage Volumes from each pool.  We call this the P:V or physical-to-virtual ratio and usually you&amp;#039;ll want to aim for anything from 2:1 to 4:1.  A 3:1 ratio means that for a node with 90x HDDs one would make that into a pool and then provision 30x (90:30 = 3:1) Storage Volumes to be turned into OSDs.   This reduction also reduces the number of OSD processes running on each node and the amount of RAM you need per system.  As an example with all 90x HDDs getting directly turned into OSDs we&amp;#039;d allocate 4GB to 6GB of RAM per OSD which amounts to 360GB to 540GB of RAM.   In contrast, with a P:V of 3:1 we are producing 30x OSDs rather than 90x OSDs for roughly the same capacity.  30x OSDs at 8GB RAM per OSD is just 240GB of RAM so we&amp;#039;re roughly cutting the RAM requirements in half.  The same goes for the required CPU power..  we&amp;#039;d normally allocate about 1 to 1.5GHz per HDD based OSD.  That&amp;#039;s 90GHz to 135GHz if using the 90x HDDs directly as OSDs.  Allocating 2GHz for each of the 30x zvol based OSDs puts us at just 60GHz.  Granted, with the second filesystem (ZFS) at work underneath Ceph you&amp;#039;ll need to reserve CPU for it and RAM for its ARC layer so the reduction can be somewhat offset but you should still see an overall reduction of 25% in CPU and RAM requirements. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Section !! Purpose&lt;br /&gt;
|-&lt;br /&gt;
| [[#How it works|How it works]] || The Use Case setting and what it changes&lt;br /&gt;
|-&lt;br /&gt;
| [[#Requirements|Requirements]] || What must be in place before you start&lt;br /&gt;
|-&lt;br /&gt;
| [[#Step 1: Create a Storage Pool reserved for Ceph OSDs|Step 1]] || Reserve a ZFS pool for Ceph OSDs&lt;br /&gt;
|-&lt;br /&gt;
| [[#Step 2: Create the zvols|Step 2]] || Carve the zvols that back the OSDs&lt;br /&gt;
|-&lt;br /&gt;
| [[#Step 3: Create the OSDs|Step 3]] || Build one OSD on each zvol&lt;br /&gt;
|-&lt;br /&gt;
| [[#Operating notes|Operating notes]] || Boot order, LVM visibility and missing OSDs&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
__TOC__&lt;br /&gt;
&lt;br /&gt;
== How it works ==&lt;br /&gt;
&lt;br /&gt;
A ZFS Storage Pool carries a &amp;#039;&amp;#039;&amp;#039;Use Case&amp;#039;&amp;#039;&amp;#039; that is set when the pool is created and cannot be changed afterwards:&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! Use Case !! Meaning&lt;br /&gt;
|-&lt;br /&gt;
| Default || The pool provisions Storage Volumes and Network Shares as usual.&lt;br /&gt;
|-&lt;br /&gt;
| Ceph OSD || The pool is reserved for zvols that back Ceph OSDs. Every Storage Volume created in it inherits the Ceph OSD use case.&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
A Storage Volume with the Ceph OSD use case is consumed directly by the Ceph OSD running on it. QuantaStor never presents it to hosts over iSCSI, FC or NVMe-oF: assigning one to a host fails with the message that the volume &amp;#039;&amp;#039;is reserved as a Ceph OSD backing device&amp;#039;&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
The setup has three steps, each covered below:&lt;br /&gt;
&lt;br /&gt;
# Create a ZFS Storage Pool with the Use Case set to &amp;#039;&amp;#039;&amp;#039;Ceph OSD&amp;#039;&amp;#039;&amp;#039;.&lt;br /&gt;
# Create the zvols in that pool, one per OSD you want.&lt;br /&gt;
# Create the OSDs from the zvols in the Ceph cluster.&lt;br /&gt;
&lt;br /&gt;
The Storage Pools list has a &amp;#039;&amp;#039;&amp;#039;Use Case&amp;#039;&amp;#039;&amp;#039; column, so you can tell a Default pool from one reserved for Ceph OSDs at a glance.&lt;br /&gt;
&lt;br /&gt;
== Requirements ==&lt;br /&gt;
&lt;br /&gt;
* The pool must be a ZFS pool. QuantaStor refuses the Ceph OSD use case for any other pool type, and refuses a Ceph OSD zvol in a non-ZFS pool.&lt;br /&gt;
* The Storage System that owns the pool must be a member of the Ceph cluster you create the OSDs in, because that system creates the OSD on its own zvol.&lt;br /&gt;
* The pool must be started when you create the OSDs.&lt;br /&gt;
* Each zvol backs exactly one OSD. A snapshot, or a volume that is assigned to hosts, cannot back an OSD.&lt;br /&gt;
&lt;br /&gt;
See [[Scale-out Block Setup (ceph)]] or [[Scale-out Object Setup (ceph)]] for creating the Ceph cluster itself.&lt;br /&gt;
&lt;br /&gt;
== Step 1: Create a Storage Pool reserved for Ceph OSDs ==&lt;br /&gt;
&lt;br /&gt;
{{Navigation|Storage Management &amp;amp;rarr; Storage Pools &amp;amp;rarr; Create &amp;#039;&amp;#039;(toolbar)&amp;#039;&amp;#039; &amp;amp;rarr; Advanced Settings &amp;#039;&amp;#039;(tab)&amp;#039;&amp;#039;}}&lt;br /&gt;
&lt;br /&gt;
[[File:Docs-docs-ceph-over-zfs-zfs-draid-pool-reserved-pool-use-case.png|thumb|right|800px|The Use Case list on the Advanced Settings tab of Create Storage Pool.]]&lt;br /&gt;
&lt;br /&gt;
Create the pool as described on [[Storage Pools]]: name it, pick the disks on the &amp;#039;&amp;#039;&amp;#039;General&amp;#039;&amp;#039;&amp;#039; tab and choose a &amp;#039;&amp;#039;&amp;#039;RAID Type&amp;#039;&amp;#039;&amp;#039;. The DRAID layouts suit this use case because the OSD layer is designed around a DRAID pool of HDDs (see [[#How QuantaStor treats a zvol-backed OSD|How QuantaStor treats a zvol-backed OSD]]).&lt;br /&gt;
&lt;br /&gt;
On the &amp;#039;&amp;#039;&amp;#039;Advanced Settings&amp;#039;&amp;#039;&amp;#039; tab, set &amp;#039;&amp;#039;&amp;#039;Use Case&amp;#039;&amp;#039;&amp;#039; to &amp;#039;&amp;#039;&amp;#039;Ceph OSD&amp;#039;&amp;#039;&amp;#039;. The other options on the tab (I/O Profile, Enable Compression, Enable SSD Auto Trim, Block Size Offset (ashift)) work as they do for any pool.&lt;br /&gt;
&lt;br /&gt;
The Use Case cannot be changed once the pool exists, so set it here.&lt;br /&gt;
&lt;br /&gt;
From the CLI, use &amp;lt;code&amp;gt;[[QuantaStor CLI Command Reference#pool-create|qs pool-create]]&amp;lt;/code&amp;gt; with &amp;lt;code&amp;gt;--pool-use-case=ceph-zvol-osd&amp;lt;/code&amp;gt;:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre style=&amp;quot;font-size: smaller&amp;quot;&amp;gt;&lt;br /&gt;
qs pool-create --name=osd-pool-1 --disk-list=&amp;lt;disk1&amp;gt;,&amp;lt;disk2&amp;gt;,... --raid-type=DRAID2_1S --pool-use-case=ceph-zvol-osd&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The DRAID layouts the CLI accepts are DRAID, DRAID1_1S, DRAID1_2S, DRAID1_3S, DRAID2_1S, DRAID2_2S, DRAID2_3S, DRAID3_1S, DRAID3_2S and DRAID3_3S.&lt;br /&gt;
&lt;br /&gt;
== Step 2: Create the zvols ==&lt;br /&gt;
&lt;br /&gt;
{{Navigation|Storage Management &amp;amp;rarr; Storage Volumes &amp;amp;rarr; Create &amp;#039;&amp;#039;(toolbar)&amp;#039;&amp;#039;}}&lt;br /&gt;
&lt;br /&gt;
Select the reserved pool as the &amp;#039;&amp;#039;&amp;#039;Storage Pool&amp;#039;&amp;#039;&amp;#039; on the &amp;#039;&amp;#039;&amp;#039;General Settings&amp;#039;&amp;#039;&amp;#039; tab, then set a &amp;#039;&amp;#039;&amp;#039;Size&amp;#039;&amp;#039;&amp;#039; for each zvol. Each zvol becomes one OSD, so the size you choose here is the OSD&amp;#039;s capacity.&lt;br /&gt;
&lt;br /&gt;
On the &amp;#039;&amp;#039;&amp;#039;Advanced Settings&amp;#039;&amp;#039;&amp;#039; tab:&lt;br /&gt;
&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;Use Case&amp;#039;&amp;#039;&amp;#039; shows &amp;#039;&amp;#039;&amp;#039;Ceph OSD&amp;#039;&amp;#039;&amp;#039; and is locked, because the pool forces its use case onto every volume it holds.&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;% Reserved&amp;#039;&amp;#039;&amp;#039; defaults to &amp;#039;&amp;#039;&amp;#039;100%&amp;#039;&amp;#039;&amp;#039;. A zvol in a Ceph OSD pool is fully reserved by default because if the pool underneath runs out of space, every OSD on it starts failing writes at once. You can still lower it deliberately.&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;Block Size&amp;#039;&amp;#039;&amp;#039; sets the zvol&amp;#039;s block size. When the OSD is created, QuantaStor sets Bluestore&amp;#039;s minimum allocation size to match it, within a range of 4K to 64K.&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;Batch Create&amp;#039;&amp;#039;&amp;#039; creates several zvols of the same size in one step.&lt;br /&gt;
&lt;br /&gt;
From the CLI, use &amp;lt;code&amp;gt;[[QuantaStor CLI Command Reference#volume-create|qs volume-create]]&amp;lt;/code&amp;gt;. The CLI default for &amp;lt;code&amp;gt;--percent-reserved&amp;lt;/code&amp;gt; is 0, so pass 100 explicitly to match the web interface:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre style=&amp;quot;font-size: smaller&amp;quot;&amp;gt;&lt;br /&gt;
qs volume-create --name=osd-vol --size=4T --pool=osd-pool-1 --count=6 --percent-reserved=100&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== A single OSD zvol in a Default pool ===&lt;br /&gt;
&lt;br /&gt;
You can also reserve an individual zvol in a Default pool: set &amp;#039;&amp;#039;&amp;#039;Use Case&amp;#039;&amp;#039;&amp;#039; to &amp;#039;&amp;#039;&amp;#039;Ceph OSD&amp;#039;&amp;#039;&amp;#039; on the &amp;#039;&amp;#039;&amp;#039;Advanced Settings&amp;#039;&amp;#039;&amp;#039; tab of Create Storage Volume, or pass &amp;lt;code&amp;gt;--volume-use-case=ceph-zvol-osd&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;qs volume-create&amp;lt;/code&amp;gt;. Only that zvol is set aside for Ceph; the pool&amp;#039;s other volumes stay available to hosts as usual. On a Ceph block pool there is nothing to choose, and the field stays at Default.&lt;br /&gt;
&lt;br /&gt;
== Step 3: Create the OSDs ==&lt;br /&gt;
&lt;br /&gt;
{{Navigation|Scale-out Storage Configuration &amp;amp;rarr; Data &amp;amp; Journal Devices &amp;amp;rarr; Create OSDs &amp;amp; Journals &amp;#039;&amp;#039;(toolbar)&amp;#039;&amp;#039; &amp;amp;rarr; Logical Devices &amp;#039;&amp;#039;(tab)&amp;#039;&amp;#039;}}&lt;br /&gt;
&lt;br /&gt;
The Create OSDs dialog has two tabs of available devices: &amp;#039;&amp;#039;&amp;#039;Physical Devices&amp;#039;&amp;#039;&amp;#039; for unused disks and &amp;#039;&amp;#039;&amp;#039;Logical Devices&amp;#039;&amp;#039;&amp;#039; for zvols. The Logical Devices tab lists every zvol that has the Ceph OSD use case, is healthy, is not a snapshot and does not already back an OSD. Move the zvols you want into &amp;#039;&amp;#039;&amp;#039;Data/OSD Devices (HDDs or SSDs)&amp;#039;&amp;#039;&amp;#039; and click &amp;#039;&amp;#039;&amp;#039;OK&amp;#039;&amp;#039;&amp;#039;. You can combine zvols and physical disks in one create. See [[Create Ceph OSD]] for the journal and optimization options in the rest of the dialog.&lt;br /&gt;
&lt;br /&gt;
If a node has no unused physical disks but does have eligible zvols, the dialog still opens; it closes with &amp;#039;&amp;#039;No unused devices were found to create new OSDs or Journal Groups with&amp;#039;&amp;#039; only when neither tab has anything to offer.&lt;br /&gt;
&lt;br /&gt;
From the CLI, use &amp;lt;code&amp;gt;[[QuantaStor CLI Command Reference#ceph-osd-multi-create|qs ceph-osd-multi-create]]&amp;lt;/code&amp;gt; with &amp;lt;code&amp;gt;--volume-list&amp;lt;/code&amp;gt;, which you can combine with &amp;lt;code&amp;gt;--physical-disk-list&amp;lt;/code&amp;gt;:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre style=&amp;quot;font-size: smaller&amp;quot;&amp;gt;&lt;br /&gt;
qs ceph-osd-multi-create --ceph-cluster=&amp;lt;cluster&amp;gt; --volume-list=osd-vol-1,osd-vol-2,osd-vol-3&lt;br /&gt;
&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== How QuantaStor treats a zvol-backed OSD ===&lt;br /&gt;
&lt;br /&gt;
* A zvol is new and has no partition table to clear, so QuantaStor skips the quick-format it runs on physical disks.&lt;br /&gt;
* For WAL/DB placement, a zvol counts as spinning media, because it is only as fast as the HDD pool behind it. Journal Group devices on SSD therefore apply to it as they do to an HDD.&lt;br /&gt;
* Once the OSD is up, QuantaStor renames its zvol to {{Code|1=vol-&amp;lt;system name&amp;gt;-osd.&amp;lt;n&amp;gt;}}, so the volume and the OSD it carries line up in the web interface and the CLI. If the rename fails, the OSD is unaffected and the zvol keeps its old name.&lt;br /&gt;
&lt;br /&gt;
== Operating notes ==&lt;br /&gt;
&lt;br /&gt;
=== Boot order ===&lt;br /&gt;
&lt;br /&gt;
The OSDs cannot start until their pool is imported. At boot, QuantaStor imports every pool reserved for Ceph OSDs, and every Default pool that holds a Ceph OSD zvol, before the Ceph OSD services start. A pool that fails to import does not stop the rest of the boot; Ceph reports the OSDs on it as down.&lt;br /&gt;
&lt;br /&gt;
Two cases are skipped:&lt;br /&gt;
&lt;br /&gt;
* A pool with auto-start disabled is not imported, and the OSDs on it do not start.&lt;br /&gt;
* A pool managed by a Storage Pool HA failover group is left to the HA manager rather than imported at boot.&lt;br /&gt;
&lt;br /&gt;
=== LVM visibility ===&lt;br /&gt;
&lt;br /&gt;
Ceph builds each OSD as an LVM volume on its device, so LVM must be able to see the zvol. By default, the {{Code|1=global_filter}} in {{Code|1=/etc/lvm/lvm.conf}} hides every zvol from LVM so the appliance never scans a volume handed out to a client. QuantaStor adds accept rules for Ceph OSD zvols only: one rule for each pool reserved for Ceph OSDs, and one per zvol for a Ceph OSD zvol in a Default pool, so that pool&amp;#039;s other zvols stay hidden. It updates the rules when a pool is created or started, including on another node after an HA failover, and when a Ceph OSD zvol is created or deleted. Other entries in the filter are left alone.&lt;br /&gt;
&lt;br /&gt;
=== Missing OSDs ===&lt;br /&gt;
&lt;br /&gt;
If the pool behind an OSD is not imported, the OSD shows as missing. QuantaStor keeps the OSD&amp;#039;s link to its zvol, so it reconnects the OSD to the zvol once the pool is imported again.&lt;br /&gt;
&lt;br /&gt;
== Related pages ==&lt;br /&gt;
&lt;br /&gt;
* [[Storage Pools]]&lt;br /&gt;
* [[Storage Volumes]]&lt;br /&gt;
* [[Create Ceph OSD]]&lt;br /&gt;
* [[Delete Ceph OSD]]&lt;br /&gt;
* [[Scale-out Block Setup (ceph)]]&lt;br /&gt;
* [[Scale-out Object Setup (ceph)]]&lt;br /&gt;
* [[Designing Systems]]&lt;br /&gt;
* [[QuantaStor CLI Command Reference]]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;small&amp;gt;&amp;#039;&amp;#039;Verified against QuantaStor 6.9.0.&amp;#039;&amp;#039;&amp;lt;/small&amp;gt;&lt;/div&gt;</summary>
		<author><name>Qadmin</name></author>
	</entry>
</feed>