Files
yucca/ansible/ceph/docs/runbooks/add-node.md
T
Andy Molenda 8821fad205 docs(ceph): realign with partition/region model and CI/CD, retire ADRs
* docs(ceph): inline ADR rationale and drop the ADR set

Fold each linked ADR's rationale into the prose it supported, then remove
the ADR files, the README index row, and the stray code-comment reference --
no ADR trace remains.

True up the docs to the partition/region/ceph-cluster layout (#222) in the
same pass: state keys (yucca/<partition>/<region>/<stack>), inventory paths
(<partition>-<region>/<cluster>), stack dirs, and the recover runbook's S3
state paths. Reframe architecture's env section around partitions/regions
with sietch as staging/austin.

* docs(ceph): editorial pass to align docs with current code and CI/CD

Rewrite the ceph docs against the actual code rather than the pre-refactor
state:

- partition/region/ceph-cluster layout throughout: state keys
  yucca/<partition>/<region>/<stack>, inventories <partition>-<region>/<cluster>,
  the real clusters.auto.tfvars schema (partition/region, not environment/datacenter)
- sietch reframed as staging/austin with secrets in yucca_tf_staging; vault
  hierarchy flipped from dev-primary to staging-primary
- live CI/CD (.github/workflows/infra.yml): per-partition read/write service
  accounts delivered as GitHub secrets (OP_TF_YUCCA_<ENV>_ENV[_WRITE]),
  plan/apply gating, NetBird overlay (Tailscale retired)
- correct the CEPH_ENV guidance (export works; deliberately kept out of mise
  [env]) across README, CONTRIBUTING, scripts.md, adding-a-cluster
- drop the obsolete "read SA token from a 1P item" dance from the runbooks
- fix stale vaults, paths, examples, and the inventory-provision.ini name

* docs(ceph): transliterate docs to plain ASCII

Replace non-ASCII punctuation and box-drawing with ASCII equivalents across
the ceph docs: em/en dashes to --/-, middot separators to commas, arrows to
->, directory-tree box-drawing to |-- / `--, section sign to "section", x for
the multiply glyph, and >= / <= / ~ for the math glyphs. No content changes.

* docs(ceph): fix broken rotate-ssh-key link in scripts.md

The "Related" link pointed at runbooks/rotate-ssh-key.md, which does not
exist (a pre-existing dangling link). SSH-key rotation lives in
rotate-secrets.md; point at its "Rotating SSH keys" section.

* docs(ceph): style polish from per-doc review

Tighten verbal texture flagged by a per-doc style pass; no structural changes.

- correctness: ansible-play.sh runs ansible-playbook, it does not exec, so the
  trap fires from the still-alive wrapper -- fix the "exec"/SIGKILL claims in
  scripts.md, architecture.md, secrets.md
- cut recurring tics: "DR belt"/"belt-and-suspenders" -> "disaster recovery",
  "a lost laptop is a non-event", "by design", "system mesh"
- unstuff long dash/semicolon sentences in architecture (vault-password
  history, provision/baseline split), secrets (SSH-key paragraph), patterns
- recover-bad-tofu-apply: move the dormant-1P-items aside into one note,
  consolidate the repeated caveats
- misc: naming grammar fix + drop trivia, hardware "would"/"blindly",
  rotate-secrets "after confidence", drop a dead snippet line, fix the 16+
  numeric hedge

* docs(ceph): make add-node and recover runbooks CI-aware

Now that infra.yml applies the stacks and runs the full ceph convergence on
merge, refresh the two runbooks the pipeline changed:

- add-node: lead with the manual-vs-CI split. The TF + host_vars change is a
  PR; the only operator-only step is the physical provisioning (live-image
  boot + provision.yml), which CI can't do; baseline/tune/join/harden run in
  CI on merge. Keep the by-hand convergence as a documented fallback.
- recover-bad-tofu-apply: note that applies now run in CI with the partition
  write SA, so the bad apply is usually a failed CI run; CI does not self-heal,
  recovery is operator-run locally.

* docs(tf,talos): finish ADR purge into tf, fix sietch vault, mark talos dormant

- complete the ADR removal that stopped at ansible/ceph: drop the dangling
  ADR-009/010 references from tf/README.md (link + related line) and the ceph
  module / stack code comments, so no ADR trace remains repo-wide
- tf/README: the sietch cluster example uses yucca_tf_staging (was yucca_tf_dev)
- ansible/talos: add a "second-class, not actively used" status banner to the
  README and architecture doc so readers don't treat the converged/libvirt
  Talos docs as live

Left untouched: the Tailscale / SA-token / CI sections of tf/README -- those
are mid-migration in another owner's lane (bye-tailscale is in flight; fabric
still rides Tailscale by design).
2026-06-29 13:04:15 -07:00

5.6 KiB

Runbook: Add a Node to an Existing Cluster

When: Expanding cluster capacity or replacing a failed chassis.

Time estimate: 30-60 minutes (Austin physical), 15-30 minutes (Hetzner).

Prerequisites:

  • Cluster is healthy (ceph health returns HEALTH_OK or understood warnings)
  • op session live (desktop unlocked or OP_SERVICE_ACCOUNT_TOKEN set) for the manual provisioning step
  • SSH access to existing cluster nodes working

What you do vs what CI does

CI (.github/workflows/infra.yml) owns the apply + convergence: on merge to main it applies the ceph TF stack (renders the inventory, mints RGW keys) and runs the full Ansible pipeline (baseline -> tune -> deploy -> harden) against the cluster over the NetBird overlay. CI cannot boot a server to a live image, so the physical provisioning is the manual part.

The flow is therefore: open a PR with the TF + host_vars changes, physically provision the box, then merge -- CI converges it.

1. Declare the new host in TF (PR)

Pick an unused name, or omit name to let TF auto-pick from the 923-word list seeded per-cluster (see docs/naming.md). Edit tf/deployment/staging/austin/ceph/clusters.auto.tfvars and append to the target cluster's hosts list:

sietch = {
  # ...
  hosts = [
    { name = "laurel", bond_ip = "10.10.10.90", bootstrap = true },
    { name = "lawson", bond_ip = "10.10.10.91" },
    { name = "samara", bond_ip = "10.10.10.92" },
    { name = "maxton", bond_ip = "10.10.10.93" },   # NEW
  ]
}

Add new hosts at the tail of the list so existing auto-picked names keep their positions. Bootstrap assignment never moves -- it stays pinned to the first/explicitly-declared bootstrap host.

2. Create host_vars (PR)

cd inventories/staging-austin/sietch
cp host_vars/example.yml host_vars/sietch-ceph-<name>.yml

Edit the new file. Every field is node-specific and must match the physical hardware:

Field How to find it
hostname_short The full sietch-ceph-<name> from step 1
bond_ip Next available IP in 10.10.10.0/24. Austin convention: yucca-N = 10.10.10.9N
sas_path_prefix SSH into node, run ls /dev/disk/by-path/ | grep sas
ssd1_phy / ssd2_phy Identify SSD PHY positions from lsscsi -t output
ceph_db_vg1/2 Name the VGs by slot, e.g., ceph-db-rear12, ceph-db-rear13
ceph_hdd_osds Map each populated HDD bay to its PHY and corresponding db-slot LV
ceph_ssd_osds SSD partition 6 on each SSD (no separate block.db)

host_vars is committed -- it rides the same PR as the tfvars change.

3. Provision the OS (manual -- CI can't do this)

This is the step that needs a human: it boots the box to a live image, which CI cannot do. Get the node provisioned and reachable on its bond IP before merging, so CI's convergence can SSH it over the NetBird overlay.

Austin (physical servers)

  1. Boot the server to the Debian 12 live image via the iDRAC virtual console.
  2. Verify the live image is reachable on the node's bond IP.
  3. Run provisioning (from your workstation):
CEPH_ENV=inventories/staging-austin/sietch/inventory-provision.ini \
  scripts/ansible-play.sh provision.yml \
  -e confirm_wipe=true \
  --limit sietch-ceph-<name>
  1. Wait for reboot and verify SSH access as ansible-iac:
ssh -i ~/.ssh/id_ed25519_sietch ansible-iac@sietch-ceph-<name> hostname -f

Expected output: sietch-ceph-<name>.staging.austin.int.futo.cloud

Hetzner (remote servers)

  1. Boot into rescue mode via the Hetzner Robot panel.
  2. SSH into the rescue system as root.
  3. Run installimage for Debian 12 Bookworm.
  4. Run the post-install script to configure networking and partitioning.
  5. Reboot into the installed OS.
  6. Verify SSH access.

4. Merge the PR -- CI converges the node

Merge to main. infra.yml applies the ceph stack and runs the full pipeline (baseline -> tune-os -> tune-hardware -> deploy-ceph -> tune-ceph -> harden) against the cluster, which joins the new host, sets up its LVM, creates its OSDs, and places services. Watch the Apply staging@austin/ceph job; the apply is gated behind the staging-austin environment's required reviewers.

Manual fallback (CI unavailable, or converging one node out-of-band)

The same convergence by hand, scoped to the new node plus the bootstrap host (join/placement/OSD activation all run from bootstrap):

scripts/ansible-play.sh baseline.yml        --limit sietch-ceph-<name>
scripts/ansible-play.sh tune-os.yml         --limit sietch-ceph-<name>
scripts/ansible-play.sh tune-hardware.yml   --limit sietch-ceph-<name>
scripts/ansible-play.sh deploy-ceph.yml     --limit sietch-ceph-<name>,sietch-ceph-laurel
scripts/ansible-play.sh tune-ceph.yml       --limit sietch-ceph-laurel
scripts/ansible-play.sh harden.yml          --limit sietch-ceph-<name>

5. Verify

Check the CI job logs, or from the bootstrap node:

ssh ansible-iac@sietch-ceph-laurel

ceph orch host ls                  # new host listed, status blank (meaning online)
ceph osd tree                      # new host bucket, OSDs up
ceph status                        # HEALTH_OK, or HEALTH_WARN with only backfill warnings
ceph orch ls --service-type rgw    # running count incremented by 1

Then mise run drift should report no drift on the new node.

Rollback

If the node needs to be removed:

# From the bootstrap node
ceph orch host drain sietch-ceph-<name> --force
# Wait for daemons to migrate (~5 min)
ceph orch host rm sietch-ceph-<name> --force

Then drop the host from clusters.auto.tfvars, delete its host_vars file, and merge -- CI re-renders the inventory without it.