Files
yucca/ansible/ceph/docs/troubleshooting.md
T
Andy Molenda 8821fad205 docs(ceph): realign with partition/region model and CI/CD, retire ADRs
* docs(ceph): inline ADR rationale and drop the ADR set

Fold each linked ADR's rationale into the prose it supported, then remove
the ADR files, the README index row, and the stray code-comment reference --
no ADR trace remains.

True up the docs to the partition/region/ceph-cluster layout (#222) in the
same pass: state keys (yucca/<partition>/<region>/<stack>), inventory paths
(<partition>-<region>/<cluster>), stack dirs, and the recover runbook's S3
state paths. Reframe architecture's env section around partitions/regions
with sietch as staging/austin.

* docs(ceph): editorial pass to align docs with current code and CI/CD

Rewrite the ceph docs against the actual code rather than the pre-refactor
state:

- partition/region/ceph-cluster layout throughout: state keys
  yucca/<partition>/<region>/<stack>, inventories <partition>-<region>/<cluster>,
  the real clusters.auto.tfvars schema (partition/region, not environment/datacenter)
- sietch reframed as staging/austin with secrets in yucca_tf_staging; vault
  hierarchy flipped from dev-primary to staging-primary
- live CI/CD (.github/workflows/infra.yml): per-partition read/write service
  accounts delivered as GitHub secrets (OP_TF_YUCCA_<ENV>_ENV[_WRITE]),
  plan/apply gating, NetBird overlay (Tailscale retired)
- correct the CEPH_ENV guidance (export works; deliberately kept out of mise
  [env]) across README, CONTRIBUTING, scripts.md, adding-a-cluster
- drop the obsolete "read SA token from a 1P item" dance from the runbooks
- fix stale vaults, paths, examples, and the inventory-provision.ini name

* docs(ceph): transliterate docs to plain ASCII

Replace non-ASCII punctuation and box-drawing with ASCII equivalents across
the ceph docs: em/en dashes to --/-, middot separators to commas, arrows to
->, directory-tree box-drawing to |-- / `--, section sign to "section", x for
the multiply glyph, and >= / <= / ~ for the math glyphs. No content changes.

* docs(ceph): fix broken rotate-ssh-key link in scripts.md

The "Related" link pointed at runbooks/rotate-ssh-key.md, which does not
exist (a pre-existing dangling link). SSH-key rotation lives in
rotate-secrets.md; point at its "Rotating SSH keys" section.

* docs(ceph): style polish from per-doc review

Tighten verbal texture flagged by a per-doc style pass; no structural changes.

- correctness: ansible-play.sh runs ansible-playbook, it does not exec, so the
  trap fires from the still-alive wrapper -- fix the "exec"/SIGKILL claims in
  scripts.md, architecture.md, secrets.md
- cut recurring tics: "DR belt"/"belt-and-suspenders" -> "disaster recovery",
  "a lost laptop is a non-event", "by design", "system mesh"
- unstuff long dash/semicolon sentences in architecture (vault-password
  history, provision/baseline split), secrets (SSH-key paragraph), patterns
- recover-bad-tofu-apply: move the dormant-1P-items aside into one note,
  consolidate the repeated caveats
- misc: naming grammar fix + drop trivia, hardware "would"/"blindly",
  rotate-secrets "after confidence", drop a dead snippet line, fix the 16+
  numeric hedge

* docs(ceph): make add-node and recover runbooks CI-aware

Now that infra.yml applies the stacks and runs the full ceph convergence on
merge, refresh the two runbooks the pipeline changed:

- add-node: lead with the manual-vs-CI split. The TF + host_vars change is a
  PR; the only operator-only step is the physical provisioning (live-image
  boot + provision.yml), which CI can't do; baseline/tune/join/harden run in
  CI on merge. Keep the by-hand convergence as a documented fallback.
- recover-bad-tofu-apply: note that applies now run in CI with the partition
  write SA, so the bad apply is usually a failed CI run; CI does not self-heal,
  recovery is operator-run locally.

* docs(tf,talos): finish ADR purge into tf, fix sietch vault, mark talos dormant

- complete the ADR removal that stopped at ansible/ceph: drop the dangling
  ADR-009/010 references from tf/README.md (link + related line) and the ceph
  module / stack code comments, so no ADR trace remains repo-wide
- tf/README: the sietch cluster example uses yucca_tf_staging (was yucca_tf_dev)
- ansible/talos: add a "second-class, not actively used" status banner to the
  README and architecture doc so readers don't treat the converged/libvirt
  Talos docs as live

Left untouched: the Tailscale / SA-token / CI sections of tf/README -- those
are mid-migration in another owner's lane (bye-tailscale is in flight; fabric
still rides Tailscale by design).
2026-06-29 13:04:15 -07:00

14 KiB

Troubleshooting Guide

Decision-tree format: symptom, diagnosis, fix.


Cluster Health

HEALTH_WARN

Symptom: HEALTH_WARN: N osds down

Diagnose:

ceph osd tree | grep down
ceph health detail

Common causes and fixes:

  1. OSD daemon crashed -- check logs:
ceph crash ls-new
ceph crash info <crash-id>
# Or on the host:
journalctl -u ceph-osd@<id> --since "1 hour ago"

Restart the daemon:

ceph orch daemon restart osd.<id>
  1. Host unreachable -- the node itself is down:
ssh ansible-iac@sietch-ceph-<name> hostname
# If unreachable, check iDRAC / physical console
  1. Disk failure -- see the replace-disk runbook.

  2. OSD out of disk space -- check nearfull/full ratios:

ceph osd df

Symptom: HEALTH_WARN: N pgs not active+clean

Diagnose:

ceph pg stat
ceph pg dump_stuck

Common causes and fixes:

  1. Backfill in progress -- normal after adding/removing OSDs. Monitor:
ceph -s

Wait for completion. No action needed.

  1. Stale PGs -- PGs stuck in stale state:
ceph pg dump_stuck stale

Usually indicates the hosting OSD is down. Fix the OSD first.

  1. Inactive PGs -- PGs in creating or peering:
ceph pg dump_stuck inactive

If stuck for >15 minutes, check MON logs. May need ceph pg force-create-pg <pgid>.


Symptom: HEALTH_WARN: clock skew detected

Diagnose:

ceph time-sync-status
# On each node:
chronyc tracking

Fix: Restart chrony on the affected node:

sudo systemctl restart chrony

If persistent, check NTP sources:

chronyc sources -v

Symptom: HEALTH_WARN: N daemons have recently crashed

Diagnose:

ceph crash ls-new
ceph crash info <crash-id>

Fix: Review the crash, then archive it:

ceph crash archive <crash-id>
# Or archive all:
ceph crash archive-all

If crashes are recurring, investigate the daemon logs on the host.


Symptom: HEALTH_WARN: N pool(s) have no replicas configured

Diagnose:

ceph osd pool ls detail | grep 'size 1'

Fix: Set appropriate replication:

ceph osd pool set <pool> size 2
ceph osd pool set <pool> min_size 1

HEALTH_ERR

Symptom: HEALTH_ERR: N pgs are stuck inactive

Diagnose:

ceph pg dump_stuck inactive
ceph osd tree

Fix: This is critical -- data may be inaccessible.

  1. Check if the hosting OSDs are down. Bring them up first.
  2. If OSDs are permanently lost and data cannot be recovered:
# DANGER: marks missing PGs as complete with potential data loss
ceph pg force-recovery <pgid>
# Last resort:
ceph osd force-create-pg <pgid> --yes-i-really-mean-it

Symptom: HEALTH_ERR: N scrub errors

Diagnose:

ceph health detail
# Find the affected PGs
ceph pg dump | grep inconsistent
# Deep scrub the PG
ceph pg deep-scrub <pgid>

Fix:

ceph pg repair <pgid>

If repair fails, the underlying disk may have bit rot. Check SMART data on the hosting OSDs.


Symptom: HEALTH_ERR: full osds

Diagnose:

ceph osd df
ceph df

Fix: This is an emergency. The cluster stops accepting writes.

  1. Delete unnecessary data or pools if possible
  2. Temporarily raise the full ratio:
ceph osd set-full-ratio 0.97
  1. Add more OSDs (see add-node runbook)
  2. Set the ratio back after capacity is restored:
ceph osd set-full-ratio 0.95

OSD Issues

Symptom: OSD down and won't start

Diagnose:

# Check daemon status
ceph orch ps --daemon-type osd | grep <host>

# Check container logs
ssh ansible-iac@<host>
sudo podman logs ceph-<fsid>-osd.<id>
sudo journalctl -u ceph-<fsid>@osd.<id>

Common causes:

  1. LUKS key missing -- dmcrypt key not in MON store:
ceph config-key dump | grep dm-crypt | grep <osd-id>

If missing, the OSD cannot be unlocked. Rebuild it (see replace-disk runbook).

  1. Corrupt BlueStore DB -- look for fsck errors in the OSD log. May need ceph-bluestore-tool repair.

  2. Block device disappeared -- check the SAS path:

ls /dev/disk/by-path/ | grep phy<N>

If missing, the disk or cable has failed.


Symptom: Slow ops / blocked requests

Diagnose:

ceph daemon osd.<id> dump_ops_in_flight
ceph daemon osd.<id> perf dump | grep -i slow

Common causes:

  1. Disk latency -- check I/O wait:
iostat -xz 5 3

Look for %util > 90% or await > 100ms on HDD devices.

  1. Network issues -- check for packet loss:
ping -c 100 <other-node-ip>
ethtool -S eno1 | grep -i error
  1. Recovery throttling too aggressive -- reduce recovery impact:
ceph config set osd osd_recovery_max_active 1
ceph config set osd osd_max_backfills 1
ceph config set osd osd_recovery_sleep_hdd 0.1

RGW (S3) Issues

Symptom: S3 requests return 403 Forbidden / SignatureDoesNotMatch

Diagnose:

# Check if the Host header hostname is in the zonegroup
radosgw-admin zonegroup get --rgw-zonegroup=us-east-1 | python3 -c "
import sys, json
zg = json.load(sys.stdin)
print('Hostnames:', zg.get('hostnames', []))
print('API name:', zg.get('api_name'))
"

Common causes:

  1. Missing hostname in zonegroup -- the Host header used by the client is not in the zonegroup's hostnames list. S3 signature verification includes the Host header, so mismatches cause 403.

Fix:

# Re-run the RGW role to add all node hostnames/IPs
scripts/ansible-play.sh deploy-ceph.yml --tags rgw \
  --limit sietch-ceph-laurel
  1. Wrong access/secret key -- verify credentials:
radosgw-admin user info --uid=svc-yucca-restic
  1. Clock skew -- S3 signatures are time-sensitive. Check client and server clocks are within 15 minutes.

Symptom: S3 requests return 500 Internal Server Error

Diagnose:

# Check RGW daemon logs
ceph log last 50 --channel=cluster | grep rgw

# Check if RGW daemons are running
ceph orch ls --service-type rgw
ceph orch ps --daemon-type rgw

Common causes:

  1. RGW daemons down -- restart:
ceph orch restart rgw
  1. Pool issues -- check that RGW pools exist and are healthy:
ceph osd pool ls | grep rgw
ceph pg stat
  1. TLS cert expired or corrupted -- see rotate-certs runbook.

Symptom: Dashboard Object Gateway page shows 500 or is empty

Diagnose:

# Check dashboard RGW API SSL verification setting
ceph dashboard get-rgw-api-ssl-verify

# Check dashboard RGW user has caps
radosgw-admin user info --uid=dashboard | python3 -c "
import sys, json
u = json.load(sys.stdin)
print('Caps:', u.get('caps', []))
print('System:', u.get('system'))
"

Fix:

  1. Disable SSL verification (self-signed certs):
ceph dashboard set-rgw-api-ssl-verify false
  1. Add admin caps to dashboard user:
radosgw-admin caps add --uid=dashboard \
  --caps='buckets=*;users=*;usage=*;metadata=*;zone=*'
  1. Sync dashboard credentials:
ceph dashboard set-rgw-credentials

Dashboard Issues

Symptom: Dashboard unreachable at https://:8443

Diagnose:

# Check MGR daemon status
ceph orch ps --daemon-type mgr

# Check which MGR is active
ceph mgr stat

# Check dashboard module is enabled
ceph mgr module ls --format json | python3 -c "
import sys, json
d = json.load(sys.stdin)
print('dashboard' in d.get('enabled_modules', []))
"

Fix:

  1. MGR daemon down -- restart:
ceph orch restart mgr
  1. Dashboard module disabled:
ceph mgr module enable dashboard
  1. Firewall blocking port 8443 -- check nftables:
ssh ansible-iac@<host> sudo nft list ruleset | grep 8443

If the port is not open, re-run security hardening:

scripts/ansible-play.sh harden.yml
  1. Wrong IP/port -- check dashboard URL:
ceph mgr services

Deploy Issues

Symptom: mise run deploy fails -- no available installation candidate for cephadm=20.2.*

Cause: apt cache is stale. prerequisites.yml adds the Ceph Tentacle repo (download.ceph.com/debian-tentacle) and then runs apt update. On a freshly-installed OS (Hetzner installimage runs apt internally), the cache is "fresh" enough that an apt update with cache_valid_time set will skip the refresh -- apt never sees the Ceph repo's Packages file and only knows about Debian's older cephadm 16.2.x.

Diagnose: SSH to the target node and run:

cat /etc/apt/sources.list.d/ceph.list   # confirm the repo file exists
apt-cache policy cephadm                # if download.ceph.com is missing, cache is stale
apt-get update                          # manual refresh
apt-cache policy cephadm                # should now show 20.2.x candidate

Fix: tasks/prerequisites.yml was patched to drop cache_valid_time on the apt-update task; refresh is unconditional after the repo is added. If you're seeing this on a node that ran the OLD prerequisites task (cached deploy state), apt update manually then re-run deploy.


Symptom: mise run deploy fails -- 'ceph_rgw_dns_name' is undefined

Cause: the per-cluster group_vars/all/vars.yml is missing the ceph_rgw_dns_name declaration. RGW zonegroup creation needs it for the --endpoints and master zonegroup hostname.

Fix: add to inventories/<partition>-<region>/<cluster>/group_vars/all/vars.yml:

ceph_rgw_dns_name: s3.{{ cluster_domain }}

This derives the DNS name from cluster_domain (e.g. s3.staging.austin.int.futo.cloud). Sietch defines this explicitly. New clusters should include it from the start -- see docs/adding-a-cluster.md group_vars template.


Symptom: HEALTH_WARN: OSDMAP_FLAGS: noin flag(s) set after deploy

Cause: the noin flag was set by an earlier failed run of the older imperative OSD-creation flow and never unset. The current spec-based flow doesn't set noin (cephadm rolls out OSDs gracefully) and includes a defensive unset task at the tail of osds.yml, but the flag can persist if the deploy never reached that tail (e.g., a failure in an earlier phase).

Fix: clear it manually, or just re-run mise run deploy -- the defensive task at the end of tasks/osds.yml unsets noin unconditionally (idempotent no-op when already unset):

ssh -i ~/.ssh/id_ed25519_<cluster> root@<bootstrap-ip> 'ceph osd unset noin'

Validate:

ceph osd dump | grep -E "^flags"
# Should NOT contain 'noin'. Default healthy: sortbitwise,recovery_deletes,purged_snapdirs,pglog_hardlimit

Provisioning Issues

Symptom: provision.yml fails with "REFUSING TO RUN"

Cause: Missing safety flag.

Fix:

CEPH_ENV=<inventory> scripts/ansible-play.sh provision.yml \
  -e confirm_wipe=true

Symptom: Provisioning fails at debootstrap / chroot phase

Diagnose: Check which task failed in the Ansible output. The rescue block automatically unmounts /mnt, so it's safe to re-run.

Common causes:

  1. apt sources unreachable from live image -- check network connectivity from the live image. DNS resolution and internet access are required for debootstrap.

  2. Disk detection failed -- SSD not found at expected path:

ls /dev/disk/by-path/ | grep sas
lsblk
  1. Previous partial provision -- the role is idempotent. If the provisioning marker exists at /mnt/etc/ceph-provisioned.json, all chroot phases are skipped. To force re-provision, boot into the live image and re-run.

Symptom: Post-reboot SSH fails after provisioning

Diagnose:

# Try with verbose SSH
ssh -vvv -i ~/.ssh/id_ed25519_sietch ansible-iac@sietch-ceph-<name>

Common causes:

  1. Node still booting -- R730xd POST takes 60-90 seconds. Wait and retry.

  2. SSH host key changed -- fresh provision generates new host keys:

ssh-keygen -R sietch-ceph-<name>
  1. Network not up -- bond interface may not have configured. Check via iDRAC virtual console.

  2. Wrong IP -- verify bond_ip in host_vars matches the actual network config.


SSH Connectivity

Symptom: Cannot SSH to cluster nodes from controller

Diagnose:

# Test SSH to a node directly
ssh ansible-iac@10.10.10.90 hostname

# If using a jump host, verify it's reachable (check your ~/.ssh/config)
ssh <jump-host> hostname

Common causes:

  1. SSH config issue -- if nodes are behind a jump host, verify your ~/.ssh/config has the correct ProxyJump or ProxyCommand settings. This is personal config, not managed by the repo.

  2. Wrong SSH key -- inventory uses ~/.ssh/id_ed25519_sietch:

ls -la ~/.ssh/id_ed25519_sietch*
  1. sntrup761 kex hang -- cephadm's asyncssh does not support post-quantum key exchange. The baseline role deploys /etc/ssh/sshd_config.d/no-sntrup.conf to disable it. If missing:
scripts/ansible-play.sh deploy-ceph.yml \
  --tags prerequisites --limit sietch-ceph-<name>,sietch-ceph-laurel
  1. nftables blocking SSH -- verify port 22 is allowed:
# From the node (via iDRAC console if SSH is blocked)
nft list ruleset | grep 22

Quick Health Check Commands

# Overall status
ceph status

# OSD health
ceph osd tree
ceph osd df

# PG health
ceph pg stat
ceph pg dump_stuck

# Services
ceph orch ls
ceph orch ps

# Recent crashes
ceph crash ls-new

# Drift from expected config
mise run drift

# Cluster capacity
ceph df