* docs(ceph): add a cluster-profiles reference for sietch and spice * docs(ceph): parameterize the cert-rotation and backup runbooks per cluster
6.2 KiB
Runbook: Backup and Restore
When: Before major operations (upgrades, topology changes), on a regular schedule, or during disaster recovery.
Time estimate: Backup: 2 minutes. Restore: depends on scenario.
Resolve
<ssh-target>for your cluster from cluster-profiles.md first. Restore steps write to the bootstrap node, and sietch and spice have different bootstrap hosts, SSH users, and keys. Copying a keyring to the wrong cluster's bootstrap is the failure mode this guards against.Every playbook below runs against whichever inventory
CEPH_ENVpoints at.scripts/ansible-play.shrequires it and exits if it is unset, so there is no default to fall back on:CEPH_ENV=<inventory>/inventory.ini.
Backup
Run a backup
mise run backup
Or directly:
scripts/ansible-play.sh backup-config.yml
What gets captured
Backups are written to backups/<timestamp>/ on the Ansible controller
(gitignored). Each backup contains:
| File | Contents |
|---|---|
ceph.conf |
Minimal cluster config (fsid, mon_host, auth settings) |
ceph.client.admin.keyring |
Admin authentication keyring |
crushmap.txt |
Decompiled CRUSH map (host/OSD topology and rules) |
osd-dump.json |
Full OSD map: pool definitions, PG counts, flags, weights |
mon-dump.json |
Monitor map: mon addresses, quorum members |
config-dump.json |
All runtime config overrides (ceph config dump) |
rgw-realm.json |
RGW realm, zonegroup, and zone configuration |
orch-services.yaml |
cephadm service specs (RGW, monitoring, crash, etc.) |
orch-hosts.yaml |
Cluster host list with addresses and labels |
Backup schedule
Scheduled (automated): mise run backup-timer (or
scripts/ansible-play.sh backup-ceph.yml) installs an on-node capture
script (/usr/local/sbin/ceph-backup.sh) plus a systemd timer on the
bootstrap node. The timer runs daily (with up to 1h of jitter) and writes a
timestamped tarball to /var/backups/ceph/ (root-only, 0700), pruning
tarballs older than ceph_backup_retention_days (default 14). Each tarball
holds the recoverable cluster state: fsid, ceph config dump, monmap,
osdmap, crushmap (binary + decompiled), ceph osd tree, ceph orch ls /
host ls, and the RGW realm/zonegroup/zone. It does not contain secret
keyrings. Opt a cluster out with ceph_backup_enabled: false in its
group_vars.
Offsite: left as an operator decision. Set ceph_backup_offsite_dest
(empty by default) to an rsync target (user@host:/path/) or S3 URI
(s3://bucket/prefix) and the script ships each tarball after capture.
Manual (mise run backup): the controller-local export below is still
available for ad-hoc snapshots. Run it manually:
- Before any cluster topology change (add/remove node or OSD)
- Before Ceph version upgrades
- Before CRUSH map modifications
- Weekly during active development
Verify a backup
ls -la backups/$(ls -t backups/ | head -1)/
Check that all files are present and non-empty. The admin keyring and ceph.conf are the most critical -- without them, cluster access is lost.
Restore Scenarios
Scenario 1: Lost ceph.conf / admin keyring on a single node
Cause: Accidental deletion, failed re-provision.
Fix: cephadm automatically distributes ceph.conf and the admin keyring to managed hosts. Force redistribution:
# From bootstrap node
ceph cephadm config-check enable
ceph orch host rescan <hostname>
Or manually copy from the backup:
scp backups/<timestamp>/ceph.conf <ssh user>@<host>:/etc/ceph/ceph.conf
scp backups/<timestamp>/ceph.client.admin.keyring <ssh user>@<host>:/etc/ceph/ceph.client.admin.keyring
Scenario 2: Lost admin keyring on ALL nodes
Cause: Full cluster purge without backup, or corruption.
Fix: Restore the keyring from the backup to the bootstrap node:
scp backups/<timestamp>/ceph.client.admin.keyring \
<ssh-target>:/etc/ceph/
ssh <ssh-target>
sudo chmod 600 /etc/ceph/ceph.client.admin.keyring
sudo chown ceph:ceph /etc/ceph/ceph.client.admin.keyring
Verify access is restored:
ceph status
Scenario 3: CRUSH map corruption
Cause: Bad CRUSH rule edit, accidental tunables change.
Fix: Restore the CRUSH map from backup:
# Compile the decompiled map
crushtool -c backups/<timestamp>/crushmap.txt -o /tmp/crushmap.bin
# Inject it
ceph osd setcrushmap -i /tmp/crushmap.bin
WARNING: This overwrites the entire CRUSH topology. Any OSDs added since the backup was taken will not be in the restored map.
Scenario 4: RGW realm/zone misconfiguration
Cause: Bad radosgw-admin command, zone placement errors.
Fix: Use the backup as a reference to reconstruct:
cat backups/<timestamp>/rgw-realm.json | python3 -m json.tool
Then re-apply zone placement targets, zonegroup hostnames, etc. using
radosgw-admin zone set / radosgw-admin zonegroup set with the JSON
from the backup piped in.
Scenario 5: Full cluster rebuild (total loss)
Cause: All nodes destroyed, starting from scratch.
- Re-provision all nodes (see add-node runbook)
- Run the full deploy pipeline:
mise run deploy
- Restore configuration from backup:
# After bootstrap, apply saved CRUSH map
crushtool -c backups/<timestamp>/crushmap.txt -o /tmp/crushmap.bin
ceph osd setcrushmap -i /tmp/crushmap.bin
# Re-apply runtime config overrides
# Review config-dump.json and apply relevant settings
cat backups/<timestamp>/config-dump.json | python3 -c "
import sys, json
for item in json.load(sys.stdin):
section = item.get('section', 'global')
name = item.get('name', '')
value = item.get('value', '')
if name and section != 'mds':
print(f'ceph config set {section} {name} {value}')
"
Note: Object data (S3 objects, bucket contents) is stored on the OSDs and cannot be restored from this config backup. This backup only preserves cluster metadata and configuration. For data protection, rely on Ceph's built-in replication (size=2+) and erasure coding.
Scenario 6: Restore service specs after cephadm reset
ceph orch apply -i backups/<timestamp>/orch-services.yaml
This re-deploys RGW, monitoring, and crash daemons with the saved placement and configuration.