Guides:Tiering On-Premises Object Storage to AWS S3 with Lifecycle Policies

From OSNEXUS Online Documentation Site
Jump to navigation Jump to search

By Steve Umbehocker, CTO, OSNexus · Updated October 2, 2026

Cloud tiering moves older objects from a local S3 bucket to a bucket at AWS while the bucket keeps serving the same namespace. QuantaStor does this on its Ceph-based scale-out object cluster with two pieces: a cloud-tier object storage class that points at an AWS account, and a bucket lifecycle rule that transitions objects into that class once they reach a set age.

Why it matters

Most object data is written once, read heavily for a few weeks and then rarely touched. Keeping all of it on local erasure-coded or replicated pools means buying capacity for data nobody reads. Tiering lets you size the on-premises cluster for the active working set and push the long tail to a cheaper AWS class such as S3 Standard-IA, S3 Glacier or Glacier Deep Archive.

The S3 client does nothing different. Applications keep writing to and listing the same local bucket through the same object gateways; the gateway moves the data in the background according to the rule.

How it works

The on-premises side is a scale-out object cluster built on Ceph. It scales from three systems to hyper-scale, and one grid can manage up to 20 Ceph clusters from a single web UI. Each cluster has one Object Storage Pool Group, which can hold many data pools with different media (NVMe, SSD, HDD) and layouts (erasure-coded or replica). Object gateways (RGW) serve S3 over HTTP/HTTPS, and a built-in load balancer is deployed with them; running a gateway on every server in the cluster gives the best performance and redundancy. See Scale-out Object Setup (ceph) for the cluster, users and buckets.

Tiering adds three objects on top of that:

Object What it is Where it lives
Cloud provider credentials The AWS access key and secret key Cloud Integration, shared with Cloud Containers
Cloud-tier storage class A named class (for example AWS_IA) bound to the credentials, a region and an AWS storage class such as STANDARD_IA The Ceph cluster's object storage zone
Bucket lifecycle rule An S3 lifecycle rule: a filter, a transition after N days, an optional expiration One Object Bucket

QuantaStor writes the bucket's whole rule set to the gateway with the S3 PutBucketLifecycleConfiguration call, using the bucket owner's key, and saves the rule only after the gateway accepts it. The gateway then runs lifecycle processing in the background, so an object moves some time after it reaches the rule's age, not at that exact moment.

Once an object has moved, its data lives in the AWS bucket and the local copy shrinks to a head object that keeps it listed. A HEAD request reports the cloud storage class name, and a listing of the local bucket shows the object with a size of 0.

Tiering moves data; it does not copy it. After the transition there is one copy, in AWS. If you need a second copy, use bucket sync policies between two QuantaStor zones; if you need a share on top of an AWS bucket, use Cloud Containers. Object Storage Classes and Cloud Tiering compares these features side by side.

Design and sizing

Pick the AWS class by how often tiered data comes back. Infrequent-access and archive classes trade a lower storage price for retrieval charges and minimum storage durations, and the Glacier classes add a restore step before an object can be read. Check the current terms on the Amazon S3 storage classes page before choosing. In the class dialog, pick a named provider class such as STANDARD_IA rather than Default.

Glacier and Deep Archive are CLI-only. Classes that tier into S3 Glacier or Glacier Deep Archive need retrieval settings that only the CLI accepts, so create those with qs object-storage-class-create rather than the dialog.

Consider a local tier first. A pool-backed storage class places objects on a specific data pool in the same cluster, so a lifecycle rule can move data from a flash pool to a high-capacity HDD pool without leaving the building. Use cloud tiering for data you expect to read rarely, if ever.

Scope the rule with filters. A rule can select objects by key prefix, minimum size and maximum size; an empty filter selects the whole bucket. Small objects still leave a head object behind, so filtering on a minimum size keeps tiny objects local. Transition days range from 1 to 365 and expiration days from 1 to 3650.

Plan expiration with tiering. To delete objects some time after they have tiered, set both a transition and a longer expiration on one rule from the CLI, or create two rules in the web UI, which creates one action per rule.

Setting it up

You need a scale-out object cluster with at least one object gateway, an Object Storage Pool Group, an object user and a bucket.

1. Add the AWS credentials. Create an IAM access key in the AWS console, as described on AWS Cloud Integration, then go to Cloud Integration → Cloud Credentials → Add Credentials, choose Amazon S3 and enter the access key and secret key. Treat the secret key like a password: never paste it into tickets, scripts in source control or documentation, and use an IAM user limited to the target bucket.

The Add Cloud Provider Credential dialog with Amazon S3 selected; the access key and secret key go into the two empty fields

2. Create the cloud-tier storage class. Go to Scale-out Storage Configuration → Scale-out Storage Pools, select the Ceph cluster, and choose Create Object Storage Class in the Object Storage toolbar group. Select Destination Cloud, enter a Storage Class Name such as AWS_IA, then pick the credentials, the Location/Region and the provider Storage Class (STANDARD_IA).

3. Add the lifecycle rule. Go to Storage Management → Object Buckets, search for and select the bucket, then choose Create in the Bucket Lifecycle Policies toolbar group. Enter a Name (Rule Id), select Transition, set Transition Days, and choose your cloud-tier class. The dialog pre-selects a storage class, so check it before clicking OK. The full field list is on Bucket Lifecycle Policy.

Create Bucket Lifecycle Policy with Transition selected: set the Transition Days, then change Storage Class from STANDARD to the cloud-tier class

The same setup from the CLI, as a worked example. It moves objects under logs/ that are larger than 1 MiB to AWS_IA after 30 days and deletes them after a year. Take the IDs for the class from the three list commands; placeholders stand in for them here, and no keys appear on the command line.

qs cloud-provider-credentials-list
qs cloud-provider-location-list
qs cloud-provider-storage-class-list
qs object-storage-class-create --storage-class-name=AWS_IA --custom-storage-class-name=true --ceph-cluster=<cluster> --tiering-cloud-credentials=<credentials id> --tiering-cloud-location=<location id> --tiering-cloud-storage-class=<provider storage class id>
qs bucket-lifecycle-policy-create --bucket=media --rule-id=logs-to-cloud --prefix=logs/ --min-size=1048576 --transition-days=30 --transition-storage-class=AWS_IA --expiration-days=365

Sizes are in bytes, and 0 leaves a filter or action unset. The transition class must exist on the bucket's Ceph cluster; any other name is refused with a list of the classes that do. Command details are in the QuantaStor CLI Command Reference.

Operating and testing

Confirm the rule:

qs bucket-lifecycle-policy-get --lifecycle-policy=logs-to-cloud --bucket=media
qs bucket-lifecycle-policy-list --bucket=media

An S3 client can read the same rule back with GetBucketLifecycleConfiguration. In the web UI, View in the Bucket Lifecycle Policies group shows each rule as the Rules JSON sent to the gateway. That list comes from the bucket scan, so rules an S3 client set directly appear only after Rescan Object Storage Buckets with Deep Scan and Enable Life Cycle Updates selected. A deep scan runs in the background and can take a long time if you have thousands of buckets (many hours if you've over 100K buckets) but will complete within a few minutes if you have a moderate number of buckets.

Rescan Object Storage Buckets with Deep Scan selected and Enable Life Cycle Updates ticked, so rules set by S3 clients appear in View

To test, upload a few objects that match the filter to a test bucket with a short transition age, then check them after the age has passed: a HEAD request should report AWS_IA and the local listing should show a size of 0, while the objects appear in the AWS bucket.

To pause tiering without losing the rule, disable it:

qs bucket-lifecycle-policy-modify --lifecycle-policy=logs-to-cloud --bucket=media --rule-status-enabled=false

Deleting a rule stops future transitions and expirations; objects that already moved stay in AWS. A storage class that a lifecycle rule still targets cannot be deleted until the rule is removed or the delete is forced.

FAQ

How can on-premises object storage automatically tier cold data to AWS S3?

Create a cloud-tier object storage class that points at your AWS account, region and S3 storage class, then add a lifecycle transition rule to the bucket. The object gateway moves matching objects to AWS in the background once they reach the rule's age, and the local bucket keeps a head object so the namespace stays intact.

Do applications need to change to use tiering?

No. Clients keep using the same bucket through the same gateways. Tiered objects still appear in listings with a size of 0, and a HEAD request reports the cloud storage class.

Is cloud tiering a backup?

No. Tiering moves the data, leaving one copy in AWS. For a second copy, replicate buckets between zones with bucket sync policies.

Can I tier to S3 Glacier or Glacier Deep Archive?

Yes, but create the storage class from the CLI. The Glacier classes need retrieval settings that the web UI dialog does not accept.

What makes this a fit for petabyte-scale on-premises S3?

The cluster scales out from three systems, a single grid manages up to 20 Ceph clusters, and one pool group can mix erasure-coded and replica pools on NVMe, SSD and HDD. Lifecycle rules then move data from flash to HDD pools or out to AWS, so local capacity tracks the active working set.


Part of the QuantaStor Guides series. For reference documentation, see the QuantaStor documentation.