* docs(ceph): inline ADR rationale and drop the ADR set Fold each linked ADR's rationale into the prose it supported, then remove the ADR files, the README index row, and the stray code-comment reference -- no ADR trace remains. True up the docs to the partition/region/ceph-cluster layout (#222) in the same pass: state keys (yucca/<partition>/<region>/<stack>), inventory paths (<partition>-<region>/<cluster>), stack dirs, and the recover runbook's S3 state paths. Reframe architecture's env section around partitions/regions with sietch as staging/austin. * docs(ceph): editorial pass to align docs with current code and CI/CD Rewrite the ceph docs against the actual code rather than the pre-refactor state: - partition/region/ceph-cluster layout throughout: state keys yucca/<partition>/<region>/<stack>, inventories <partition>-<region>/<cluster>, the real clusters.auto.tfvars schema (partition/region, not environment/datacenter) - sietch reframed as staging/austin with secrets in yucca_tf_staging; vault hierarchy flipped from dev-primary to staging-primary - live CI/CD (.github/workflows/infra.yml): per-partition read/write service accounts delivered as GitHub secrets (OP_TF_YUCCA_<ENV>_ENV[_WRITE]), plan/apply gating, NetBird overlay (Tailscale retired) - correct the CEPH_ENV guidance (export works; deliberately kept out of mise [env]) across README, CONTRIBUTING, scripts.md, adding-a-cluster - drop the obsolete "read SA token from a 1P item" dance from the runbooks - fix stale vaults, paths, examples, and the inventory-provision.ini name * docs(ceph): transliterate docs to plain ASCII Replace non-ASCII punctuation and box-drawing with ASCII equivalents across the ceph docs: em/en dashes to --/-, middot separators to commas, arrows to ->, directory-tree box-drawing to |-- / `--, section sign to "section", x for the multiply glyph, and >= / <= / ~ for the math glyphs. No content changes. * docs(ceph): fix broken rotate-ssh-key link in scripts.md The "Related" link pointed at runbooks/rotate-ssh-key.md, which does not exist (a pre-existing dangling link). SSH-key rotation lives in rotate-secrets.md; point at its "Rotating SSH keys" section. * docs(ceph): style polish from per-doc review Tighten verbal texture flagged by a per-doc style pass; no structural changes. - correctness: ansible-play.sh runs ansible-playbook, it does not exec, so the trap fires from the still-alive wrapper -- fix the "exec"/SIGKILL claims in scripts.md, architecture.md, secrets.md - cut recurring tics: "DR belt"/"belt-and-suspenders" -> "disaster recovery", "a lost laptop is a non-event", "by design", "system mesh" - unstuff long dash/semicolon sentences in architecture (vault-password history, provision/baseline split), secrets (SSH-key paragraph), patterns - recover-bad-tofu-apply: move the dormant-1P-items aside into one note, consolidate the repeated caveats - misc: naming grammar fix + drop trivia, hardware "would"/"blindly", rotate-secrets "after confidence", drop a dead snippet line, fix the 16+ numeric hedge * docs(ceph): make add-node and recover runbooks CI-aware Now that infra.yml applies the stacks and runs the full ceph convergence on merge, refresh the two runbooks the pipeline changed: - add-node: lead with the manual-vs-CI split. The TF + host_vars change is a PR; the only operator-only step is the physical provisioning (live-image boot + provision.yml), which CI can't do; baseline/tune/join/harden run in CI on merge. Keep the by-hand convergence as a documented fallback. - recover-bad-tofu-apply: note that applies now run in CI with the partition write SA, so the bad apply is usually a failed CI run; CI does not self-heal, recovery is operator-run locally. * docs(tf,talos): finish ADR purge into tf, fix sietch vault, mark talos dormant - complete the ADR removal that stopped at ansible/ceph: drop the dangling ADR-009/010 references from tf/README.md (link + related line) and the ceph module / stack code comments, so no ADR trace remains repo-wide - tf/README: the sietch cluster example uses yucca_tf_staging (was yucca_tf_dev) - ansible/talos: add a "second-class, not actively used" status banner to the README and architecture doc so readers don't treat the converged/libvirt Talos docs as live Left untouched: the Tailscale / SA-token / CI sections of tf/README -- those are mid-migration in another owner's lane (bye-tailscale is in flight; fabric still rides Tailscale by design).
14 KiB
Troubleshooting Guide
Decision-tree format: symptom, diagnosis, fix.
Cluster Health
HEALTH_WARN
Symptom: HEALTH_WARN: N osds down
Diagnose:
ceph osd tree | grep down
ceph health detail
Common causes and fixes:
- OSD daemon crashed -- check logs:
ceph crash ls-new
ceph crash info <crash-id>
# Or on the host:
journalctl -u ceph-osd@<id> --since "1 hour ago"
Restart the daemon:
ceph orch daemon restart osd.<id>
- Host unreachable -- the node itself is down:
ssh ansible-iac@sietch-ceph-<name> hostname
# If unreachable, check iDRAC / physical console
-
Disk failure -- see the replace-disk runbook.
-
OSD out of disk space -- check nearfull/full ratios:
ceph osd df
Symptom: HEALTH_WARN: N pgs not active+clean
Diagnose:
ceph pg stat
ceph pg dump_stuck
Common causes and fixes:
- Backfill in progress -- normal after adding/removing OSDs. Monitor:
ceph -s
Wait for completion. No action needed.
- Stale PGs -- PGs stuck in
stalestate:
ceph pg dump_stuck stale
Usually indicates the hosting OSD is down. Fix the OSD first.
- Inactive PGs -- PGs in
creatingorpeering:
ceph pg dump_stuck inactive
If stuck for >15 minutes, check MON logs. May need ceph pg force-create-pg <pgid>.
Symptom: HEALTH_WARN: clock skew detected
Diagnose:
ceph time-sync-status
# On each node:
chronyc tracking
Fix: Restart chrony on the affected node:
sudo systemctl restart chrony
If persistent, check NTP sources:
chronyc sources -v
Symptom: HEALTH_WARN: N daemons have recently crashed
Diagnose:
ceph crash ls-new
ceph crash info <crash-id>
Fix: Review the crash, then archive it:
ceph crash archive <crash-id>
# Or archive all:
ceph crash archive-all
If crashes are recurring, investigate the daemon logs on the host.
Symptom: HEALTH_WARN: N pool(s) have no replicas configured
Diagnose:
ceph osd pool ls detail | grep 'size 1'
Fix: Set appropriate replication:
ceph osd pool set <pool> size 2
ceph osd pool set <pool> min_size 1
HEALTH_ERR
Symptom: HEALTH_ERR: N pgs are stuck inactive
Diagnose:
ceph pg dump_stuck inactive
ceph osd tree
Fix: This is critical -- data may be inaccessible.
- Check if the hosting OSDs are down. Bring them up first.
- If OSDs are permanently lost and data cannot be recovered:
# DANGER: marks missing PGs as complete with potential data loss
ceph pg force-recovery <pgid>
# Last resort:
ceph osd force-create-pg <pgid> --yes-i-really-mean-it
Symptom: HEALTH_ERR: N scrub errors
Diagnose:
ceph health detail
# Find the affected PGs
ceph pg dump | grep inconsistent
# Deep scrub the PG
ceph pg deep-scrub <pgid>
Fix:
ceph pg repair <pgid>
If repair fails, the underlying disk may have bit rot. Check SMART data on the hosting OSDs.
Symptom: HEALTH_ERR: full osds
Diagnose:
ceph osd df
ceph df
Fix: This is an emergency. The cluster stops accepting writes.
- Delete unnecessary data or pools if possible
- Temporarily raise the full ratio:
ceph osd set-full-ratio 0.97
- Add more OSDs (see add-node runbook)
- Set the ratio back after capacity is restored:
ceph osd set-full-ratio 0.95
OSD Issues
Symptom: OSD down and won't start
Diagnose:
# Check daemon status
ceph orch ps --daemon-type osd | grep <host>
# Check container logs
ssh ansible-iac@<host>
sudo podman logs ceph-<fsid>-osd.<id>
sudo journalctl -u ceph-<fsid>@osd.<id>
Common causes:
- LUKS key missing -- dmcrypt key not in MON store:
ceph config-key dump | grep dm-crypt | grep <osd-id>
If missing, the OSD cannot be unlocked. Rebuild it (see replace-disk runbook).
-
Corrupt BlueStore DB -- look for
fsckerrors in the OSD log. May needceph-bluestore-tool repair. -
Block device disappeared -- check the SAS path:
ls /dev/disk/by-path/ | grep phy<N>
If missing, the disk or cable has failed.
Symptom: Slow ops / blocked requests
Diagnose:
ceph daemon osd.<id> dump_ops_in_flight
ceph daemon osd.<id> perf dump | grep -i slow
Common causes:
- Disk latency -- check I/O wait:
iostat -xz 5 3
Look for %util > 90% or await > 100ms on HDD devices.
- Network issues -- check for packet loss:
ping -c 100 <other-node-ip>
ethtool -S eno1 | grep -i error
- Recovery throttling too aggressive -- reduce recovery impact:
ceph config set osd osd_recovery_max_active 1
ceph config set osd osd_max_backfills 1
ceph config set osd osd_recovery_sleep_hdd 0.1
RGW (S3) Issues
Symptom: S3 requests return 403 Forbidden / SignatureDoesNotMatch
Diagnose:
# Check if the Host header hostname is in the zonegroup
radosgw-admin zonegroup get --rgw-zonegroup=us-east-1 | python3 -c "
import sys, json
zg = json.load(sys.stdin)
print('Hostnames:', zg.get('hostnames', []))
print('API name:', zg.get('api_name'))
"
Common causes:
- Missing hostname in zonegroup -- the Host header used by the client is not in the zonegroup's hostnames list. S3 signature verification includes the Host header, so mismatches cause 403.
Fix:
# Re-run the RGW role to add all node hostnames/IPs
scripts/ansible-play.sh deploy-ceph.yml --tags rgw \
--limit sietch-ceph-laurel
- Wrong access/secret key -- verify credentials:
radosgw-admin user info --uid=svc-yucca-restic
- Clock skew -- S3 signatures are time-sensitive. Check client and server clocks are within 15 minutes.
Symptom: S3 requests return 500 Internal Server Error
Diagnose:
# Check RGW daemon logs
ceph log last 50 --channel=cluster | grep rgw
# Check if RGW daemons are running
ceph orch ls --service-type rgw
ceph orch ps --daemon-type rgw
Common causes:
- RGW daemons down -- restart:
ceph orch restart rgw
- Pool issues -- check that RGW pools exist and are healthy:
ceph osd pool ls | grep rgw
ceph pg stat
- TLS cert expired or corrupted -- see rotate-certs runbook.
Symptom: Dashboard Object Gateway page shows 500 or is empty
Diagnose:
# Check dashboard RGW API SSL verification setting
ceph dashboard get-rgw-api-ssl-verify
# Check dashboard RGW user has caps
radosgw-admin user info --uid=dashboard | python3 -c "
import sys, json
u = json.load(sys.stdin)
print('Caps:', u.get('caps', []))
print('System:', u.get('system'))
"
Fix:
- Disable SSL verification (self-signed certs):
ceph dashboard set-rgw-api-ssl-verify false
- Add admin caps to dashboard user:
radosgw-admin caps add --uid=dashboard \
--caps='buckets=*;users=*;usage=*;metadata=*;zone=*'
- Sync dashboard credentials:
ceph dashboard set-rgw-credentials
Dashboard Issues
Symptom: Dashboard unreachable at https://:8443
Diagnose:
# Check MGR daemon status
ceph orch ps --daemon-type mgr
# Check which MGR is active
ceph mgr stat
# Check dashboard module is enabled
ceph mgr module ls --format json | python3 -c "
import sys, json
d = json.load(sys.stdin)
print('dashboard' in d.get('enabled_modules', []))
"
Fix:
- MGR daemon down -- restart:
ceph orch restart mgr
- Dashboard module disabled:
ceph mgr module enable dashboard
- Firewall blocking port 8443 -- check nftables:
ssh ansible-iac@<host> sudo nft list ruleset | grep 8443
If the port is not open, re-run security hardening:
scripts/ansible-play.sh harden.yml
- Wrong IP/port -- check dashboard URL:
ceph mgr services
Deploy Issues
Symptom: mise run deploy fails -- no available installation candidate for cephadm=20.2.*
Cause: apt cache is stale. prerequisites.yml adds the Ceph Tentacle
repo (download.ceph.com/debian-tentacle) and then runs apt update.
On a freshly-installed OS (Hetzner installimage runs apt internally),
the cache is "fresh" enough that an apt update with cache_valid_time
set will skip the refresh -- apt never sees the Ceph repo's Packages file
and only knows about Debian's older cephadm 16.2.x.
Diagnose: SSH to the target node and run:
cat /etc/apt/sources.list.d/ceph.list # confirm the repo file exists
apt-cache policy cephadm # if download.ceph.com is missing, cache is stale
apt-get update # manual refresh
apt-cache policy cephadm # should now show 20.2.x candidate
Fix: tasks/prerequisites.yml was patched to drop cache_valid_time
on the apt-update task; refresh is unconditional after the repo is
added. If you're seeing this on a node that ran the OLD prerequisites
task (cached deploy state), apt update manually then re-run deploy.
Symptom: mise run deploy fails -- 'ceph_rgw_dns_name' is undefined
Cause: the per-cluster group_vars/all/vars.yml is missing the
ceph_rgw_dns_name declaration. RGW zonegroup creation needs it for
the --endpoints and master zonegroup hostname.
Fix: add to inventories/<partition>-<region>/<cluster>/group_vars/all/vars.yml:
ceph_rgw_dns_name: s3.{{ cluster_domain }}
This derives the DNS name from cluster_domain (e.g.
s3.staging.austin.int.futo.cloud). Sietch defines this explicitly. New
clusters should include it from the start -- see
docs/adding-a-cluster.md group_vars template.
Symptom: HEALTH_WARN: OSDMAP_FLAGS: noin flag(s) set after deploy
Cause: the noin flag was set by an earlier failed run of the
older imperative OSD-creation flow and never unset. The current
spec-based flow doesn't set noin (cephadm rolls out OSDs gracefully)
and includes a defensive unset task at the tail of osds.yml, but the
flag can persist if the deploy never reached that tail (e.g., a failure
in an earlier phase).
Fix: clear it manually, or just re-run mise run deploy -- the
defensive task at the end of tasks/osds.yml unsets noin
unconditionally (idempotent no-op when already unset):
ssh -i ~/.ssh/id_ed25519_<cluster> root@<bootstrap-ip> 'ceph osd unset noin'
Validate:
ceph osd dump | grep -E "^flags"
# Should NOT contain 'noin'. Default healthy: sortbitwise,recovery_deletes,purged_snapdirs,pglog_hardlimit
Provisioning Issues
Symptom: provision.yml fails with "REFUSING TO RUN"
Cause: Missing safety flag.
Fix:
CEPH_ENV=<inventory> scripts/ansible-play.sh provision.yml \
-e confirm_wipe=true
Symptom: Provisioning fails at debootstrap / chroot phase
Diagnose: Check which task failed in the Ansible output. The rescue
block automatically unmounts /mnt, so it's safe to re-run.
Common causes:
-
apt sources unreachable from live image -- check network connectivity from the live image. DNS resolution and internet access are required for debootstrap.
-
Disk detection failed -- SSD not found at expected path:
ls /dev/disk/by-path/ | grep sas
lsblk
- Previous partial provision -- the role is idempotent. If the
provisioning marker exists at
/mnt/etc/ceph-provisioned.json, all chroot phases are skipped. To force re-provision, boot into the live image and re-run.
Symptom: Post-reboot SSH fails after provisioning
Diagnose:
# Try with verbose SSH
ssh -vvv -i ~/.ssh/id_ed25519_sietch ansible-iac@sietch-ceph-<name>
Common causes:
-
Node still booting -- R730xd POST takes 60-90 seconds. Wait and retry.
-
SSH host key changed -- fresh provision generates new host keys:
ssh-keygen -R sietch-ceph-<name>
-
Network not up -- bond interface may not have configured. Check via iDRAC virtual console.
-
Wrong IP -- verify
bond_ipin host_vars matches the actual network config.
SSH Connectivity
Symptom: Cannot SSH to cluster nodes from controller
Diagnose:
# Test SSH to a node directly
ssh ansible-iac@10.10.10.90 hostname
# If using a jump host, verify it's reachable (check your ~/.ssh/config)
ssh <jump-host> hostname
Common causes:
-
SSH config issue -- if nodes are behind a jump host, verify your
~/.ssh/confighas the correct ProxyJump or ProxyCommand settings. This is personal config, not managed by the repo. -
Wrong SSH key -- inventory uses
~/.ssh/id_ed25519_sietch:
ls -la ~/.ssh/id_ed25519_sietch*
- sntrup761 kex hang -- cephadm's asyncssh does not support
post-quantum key exchange. The baseline role deploys
/etc/ssh/sshd_config.d/no-sntrup.confto disable it. If missing:
scripts/ansible-play.sh deploy-ceph.yml \
--tags prerequisites --limit sietch-ceph-<name>,sietch-ceph-laurel
- nftables blocking SSH -- verify port 22 is allowed:
# From the node (via iDRAC console if SSH is blocked)
nft list ruleset | grep 22
Quick Health Check Commands
# Overall status
ceph status
# OSD health
ceph osd tree
ceph osd df
# PG health
ceph pg stat
ceph pg dump_stuck
# Services
ceph orch ls
ceph orch ps
# Recent crashes
ceph crash ls-new
# Drift from expected config
mise run drift
# Cluster capacity
ceph df