Encryption Bypass: Difference between revisions
m Clarify where the ceph-volume activation message is logged (QSTOR-12352) |
m Restore the systemd Services related link carried by the previous revision (QSTOR-12352) |
||
| Line 209: | Line 209: | ||
* [[Seagate Corvault Configuration]] -- the hardware the shipped entries exist for | * [[Seagate Corvault Configuration]] -- the hardware the shipped entries exist for | ||
* [[Security Configuration]] -- appliance-wide security settings | * [[Security Configuration]] -- appliance-wide security settings | ||
* [[QuantaStor systemd Services]] -- the service and the boot units named above | |||
---- | ---- | ||
Revision as of 11:23, 3 September 2026
The encryption bypass list excludes a device type from QuantaStor's data-at-rest encryption. A device that matches an entry is provisioned into an encryption-enabled Ceph cluster -- or an encrypted Storage Pool -- without being encrypted by QuantaStor. The data lands in the clear as far as the appliance is concerned, and protecting it at rest is left entirely to the media's own hardware encryption. The feature exists for self-encrypting storage such as the Seagate Corvault enclosures, whose LUNs already encrypt at rest and gain nothing from a second software layer wrapped around them.
This is a security exception mechanism, not a performance option. An entry turns QuantaStor's encryption off for every device of that vendor and model on the node, and for any OSD created while the entry is in force the exception is permanent -- there is no way to encrypt that OSD afterwards without destroying and re-creating it. The two Seagate Corvault entries that ship in the file are the supported use of it. We recommend adding an entry only when OSNEXUS support has asked you to, and only for media you have established encrypts at rest on its own.
There is no dialog for this list and no qs command that edits it. It is a per-node configuration file, read once when the QuantaStor service starts.
| Section | Purpose |
|---|---|
| What the bypass turns off | Exactly which step is skipped, and what the device is left holding. |
| How a device is matched | Vendor and model, not path, serial or WWN -- and what that implies. |
| The configuration file | Path, override location, format, and the syntax rules that bite. |
| When the list is consulted | Create-time for software encryption, every unlock for SED. |
| Verifying that a device was bypassed | The three places the exception is visible. |
| Reversing it | What removing an entry does, and what it does not undo. |
| Security consequences | Stated plainly, including what an audit will not show you. |
| Interaction with Storage Pools and SED media | The list is appliance-wide, not Ceph-specific. |
| The ceph-volume patch it depends on | Why OSD creation can fail after a Ceph upgrade. |
What the bypass turns off
The bypass only has an effect where QuantaStor would otherwise apply encryption, so on a Ceph cluster created without encryption it does nothing at all. Encryption is chosen when the cluster is created, and cannot be turned on afterwards:
Selecting Enable Encryption enables the Encryption tab, where the cluster is set to either software encryption or hardware self-encrypting drive (SED) encryption. See Create Ceph Cluster Configuration for that dialog. What the bypass skips differs between the two modes.
Software (dm-crypt/LUKS) clusters
QuantaStor still creates the OSD as an encrypted OSD. It still passes --dmcrypt to ceph-volume lvm prepare, a cephx lockbox secret and a dm-crypt key are still generated and stored, and the OSD's write-ahead log and database devices are still LUKS-formatted. What changes is one extra argument, --block.skip-enc true, which suppresses the LUKS format of the block (data) logical volume alone.
The result on disk is a block LV with no LUKS header, holding BlueStore data unencrypted. The journal devices beside it are encrypted normally. Nothing about the OSD is encrypted by another means -- if the enclosure does not encrypt it, nothing does.
The decision is then recorded as an LVM tag on that block LV, ceph.block_skip_enc=True. Activation and deactivation read the tag, not the configuration file: activation mounts the plain LV path and does not attempt a LUKS open on it, and deactivation does not attempt to close a mapping that was never opened.
SED (self-encrypting drive) clusters
Here QuantaStor skips the drive's SED initialization outright. It does not set a passphrase on the device, does not enable crypt locking, and does not write the lock indicator it would normally create for the device. The drive is left in whatever state the vendor shipped it in, and stays unlocked.
The same check also short-circuits the unlock path, so a device in the bypass list is never sent an unlock command when the cluster comes up.
How a device is matched
Matching is by SCSI Vendor ID and Product ID -- the inquiry strings -- and by nothing else. It is not by device path, serial number, WWN, SCSI ID or drive slot. Both strings must match exactly and case-sensitively; there is no substring or wildcard matching, so Virtual does not match a product ID of Virtual disk.
Read the two values off the device with qs disk-get --disk=<disk>, which reports them as Vendor ID and Product ID:
qs disk-get --disk=sdd
[Physical Disk]
Name: sdd (scsi-36000c2932f6d4d3a100ce1bf7350de80)
Product ID: Virtual disk
Vendor ID: VMware
Is Encrypted: false
qs disk-list shows the same two values as its Vendor and Product columns for every device in the grid, which is the quicker way to survey a node. See Physical Disks/Devices for the rest of the device inventory.
Two consequences follow from matching on the device type rather than on an individual device:
- A renumbering is harmless. Because no path or serial is stored, a device that comes back as a different
/dev/sdXafter a reboot or a re-cable is still matched, and a device that inherits a path is not wrongly matched. - The exception cannot be narrowed to one drive. Every device of that vendor and model on the node is excluded. There is no way to bypass one Corvault LUN and encrypt another of the same model, and adding a common vendor and model -- a virtual disk type, say -- excludes essentially every device on the node.
The configuration file
The list is read from /opt/osnexus/quantastor/conf/qs_encryption_bypass.conf, which ships with the two Seagate Corvault entries already populated and a commented-out example.
Do not edit the shipped copy. An upgrade replaces it without prompting and without preserving your changes. Copy it to /var/opt/osnexus/quantastor/conf/qs_encryption_bypass.conf and edit the copy; that path is resolved first and survives an upgrade. The override replaces the shipped file rather than merging with it, so keep the Corvault entries in your copy unless you intend to remove them. QuantaStor Configuration Files covers the resolution order and the override rules for all of these files.
The format is an INI-style file. Each device type is a named section with a vendor and a model key:
# Seagate Corvault systems have built-in encryption so we bypass applying # QuantaStor's software or SED encryption to these devices. [seagate_corvault_4u106] vendor=SEAGATE model=6575 [seagate_corvault_5u84] vendor=SEAGATE model=6566
Rules the parser enforces, all of which fail silently:
- Both keys are required. A section with only a
vendor, or only amodel, is skipped without an error. - No spaces around the
=. The key is the text before the first=and the value is everything after it, taken literally.vendor = SEAGATEdefines a key namedvendorwith the valueSEAGATE, and matches nothing. - No trailing whitespace on a value line, for the same reason. Keep the file in Unix line endings.
- Section headers and comments must start in column one. A leading
#is only treated as a comment at the start of the line. - The section name itself is a label only -- it is never matched against anything, so use something descriptive.
The file is not synchronised across the grid. It is read from local disk by each node's own service and nothing copies it between systems, so make the same edit on every node that will host OSDs of that device type.
The file is read once, when the QuantaStor service starts. An edit has no effect until the service is restarted on that node:
sudo systemctl restart quantastor
Confirm what the node actually loaded from the service log, which names each vendor and model pair as it is added:
sudo grep "encryption bypass entry" /var/log/qs/qs_service.log
When the list is consulted
Whether a later edit changes anything depends on which encryption mode is in play, and this is the part most likely to surprise:
| Operation | When the list is consulted |
|---|---|
| Ceph OSD create, software encryption | At create time. The outcome is then persisted as the ceph.block_skip_enc LVM tag on the OSD's block LV.
|
| Ceph OSD activate and deactivate, software encryption | Never. Both follow the LVM tag written at create time. |
| Ceph OSD create, SED cluster | At create time, to decide whether to initialize the drive's SED. |
| SED OSD unlock when the cluster starts | Every time. A device in the list is skipped rather than unlocked. |
| Encrypted Storage Pool device format | At the point the device is added to the pool. |
| Encrypted Storage Pool activation | Every time. A device in the list is skipped rather than crypt-opened. |
qs ceph-osd-key-replace |
Every time. A matching OSD is skipped and its key is not rotated. |
| SMART and thermal polling | Every poll. A matching device is treated as a hardware logical drive and excluded from smartctl and sg_logs collection.
|
- Adding a device type to the list after its OSDs or pool devices were already encrypted is the dangerous direction. An existing software-encrypted OSD keeps working, because activation follows its LVM tag -- but its encryption key is silently no longer rotated by qs ceph-osd-key-replace. An already-locked SED device, and an encrypted Storage Pool device, are worse: the unlock and the crypt-open are skipped on every start, so the device stays locked or unopened and the pool or OSD fails to come up.
Verifying that a device was bypassed
The exception is visible in three places, and nowhere in the WUI:
- What the node loaded. The service log line quoted above lists each vendor and model pair the node is bypassing. Absence of a line for your device type means the entry did not parse, or the service has not been restarted since the edit.
- The LVM tag on the OSD. On a software-encrypted cluster, the OSD's block LV carries
ceph.block_skip_enc=Truealongside the usual Ceph tags:
sudo lvs -o lv_name,lv_tags | grep block_skip_enc
Note that ceph.encrypted is still 1 on that LV, and ceph-volume lvm list still reports the OSD as encrypted, because the journal devices are. The block_skip_enc tag is the only thing that distinguishes a bypassed OSD.
- The absence of a crypt layer.
lsblkon the bypassed device shows the LVM layer with nocryptdevice above the block LV, while the WAL and DB LVs of the same OSD still show one.
During OSD creation, ceph-volume also logs Skipping block device encryption for: <path>, and on activation Skipping block device decryption for <path>. The creation message lands in the QuantaStor service log with the rest of the OSD creation output; the activation message goes wherever the activation ran -- the service log when QuantaStor drove it, the systemd journal for the ceph-volume@ unit when the node activated its OSDs at boot.
Reversing it
Remove or comment out the section, then restart the service on that node. From then on, new OSDs and new pool devices of that type are encrypted normally.
Removing an entry does not undo anything already provisioned, and in two cases it actively breaks it:
- An existing software-encrypted Ceph OSD is unaffected and stays unencrypted. Its LVM tag is what drives activation, and nothing rewrites that tag. There is no in-place conversion: to get an encrypted OSD on that device you must delete the OSD with
qs ceph-osd-delete, let Ceph finish recovering, and create it again. Do that one OSD at a time and wait for HEALTH_OK in between. - An encrypted Storage Pool device that was bypassed will fail to activate. The device was never LUKS-formatted, but it does carry an entry in the encryption table, so once the bypass is gone the service tries to crypt-open a device with no LUKS header.
- A SED device that QuantaStor never initialized has no QuantaStor passphrase, so removing the entry does not make it lockable -- it makes the unlock attempt fail instead.
In all three cases the device has to be removed from the pool or cluster and re-added to become encrypted.
Security consequences
Stated without hedging, because the point of the feature is to switch protection off:
- The data is in the clear on the medium. A bypassed block device holds plaintext BlueStore data. A drive pulled from the chassis is readable unless the hardware's own encryption stops it, and QuantaStor has no way to confirm that it does.
- An audit of the cluster will not show the exception. The cluster remains an encryption-enabled cluster, the OSD remains
ceph.encrypted=1, and the WUI reports the cluster as encrypted. Only the configuration file, the service log and the LVM tag reveal that a device is excluded. - The scope is the device model, not the device. Every matching device on the node is excluded, including devices added later.
- The scope is also not limited to Ceph. The same list governs encrypted Storage Pools and SED pool devices on the same appliance. An entry added to allow a Corvault OSD also disables encryption for Corvault LUNs used in a Storage Pool on that node.
- Key rotation silently skips bypassed OSDs. A
qs ceph-osd-key-replacerun reports success while leaving matching OSDs untouched, because there is no key to replace. - Health monitoring is reduced. Matching devices are excluded from SMART and temperature collection, so drive-health alerting for them comes from the enclosure rather than from QuantaStor.
Interaction with Storage Pools and SED media
The list lives in the disk layer, not the Ceph layer, so it applies to Storage Pools as well. When a device that matches an entry is added to a pool created with encryption, QuantaStor logs that bypass is enabled for the device and skips the LUKS format or the SED initialization, and the pool is created with that device unencrypted while the pool itself still reports as encrypted.
One further side effect is worth knowing before you go looking for a cause: qs pool-rekey refuses to run on a pool containing a bypassed device, and reports it as a device with no available key slot rather than as an excluded device.
For genuine self-encrypting drives the bypass is usually the wrong tool. QuantaStor manages SED media directly -- taking ownership of the drive, setting a passphrase and locking it -- and the SED state of each device is reported in the SED Capable and SED Status columns on Physical Disks/Devices. Bypass is for media whose encryption QuantaStor cannot drive at all, which in practice means a self-encrypting array or enclosure presenting logical LUNs rather than drives.
The ceph-volume patch it depends on
The --block.skip-enc argument is an OSNEXUS extension to the stock ceph-volume tool, not an upstream Ceph feature. QuantaStor ships a per-release patch set and applies it with a script that runs at service start, from a boot-time unit ordered ahead of OSD activation, and from an APT hook after any package operation that could have replaced ceph-volume. Applying the patch is idempotent and version-aware: it selects the patch set matching the running Ceph release, dry-runs each patch before installing it, and refuses to install a result that does not compile.
Administrators do not normally interact with any of this. Two failure modes are worth recognising:
- If a Ceph point release changes the files the patches target, patching fails, an alert is raised naming the Ceph version, and OSD creation on a bypassed device type can fail. The remedy is to contact OSNEXUS support with the Ceph version so an updated patch set can be produced -- not to retry the OSD creation.
- The patch set is maintained per minor Ceph line. On a Ceph release with no patch set, patching is skipped with a warning in the service log, and
--block.skip-encis not passed. OSDs then get the standard full software encryption, which is safe but is not what the bypass entry asked for.
Confirm the patch state on a node from the service log:
sudo grep "Ceph volume patching" /var/log/qs/qs_service.log
Patching can be suppressed with the tf_qs_ceph_volume_patch.disable touch file, which every caller honours. That is a support-guided diagnostic step and it disables the bypass for new software-encrypted OSDs, so see QuantaStor Touch Files before creating it.
Related pages
- QuantaStor Configuration Files -- the registry these files belong to, and the override rules
- QuantaStor Touch Files -- the companion
tf_*markers, including the ceph-volume patch switch - Create Ceph Cluster Configuration -- where cluster encryption is chosen
- Physical Disks/Devices -- the device inventory, and the SED state of each drive
- Storage Pools -- pool-level encryption, which the same list governs
- Seagate Corvault Configuration -- the hardware the shipped entries exist for
- Security Configuration -- appliance-wide security settings
- QuantaStor systemd Services -- the service and the boot units named above
Verified against QuantaStor 6.9.0.