Files
yucca/ansible/ceph/docs/scripts.md
T
Andy Molenda 8821fad205 docs(ceph): realign with partition/region model and CI/CD, retire ADRs
* docs(ceph): inline ADR rationale and drop the ADR set

Fold each linked ADR's rationale into the prose it supported, then remove
the ADR files, the README index row, and the stray code-comment reference --
no ADR trace remains.

True up the docs to the partition/region/ceph-cluster layout (#222) in the
same pass: state keys (yucca/<partition>/<region>/<stack>), inventory paths
(<partition>-<region>/<cluster>), stack dirs, and the recover runbook's S3
state paths. Reframe architecture's env section around partitions/regions
with sietch as staging/austin.

* docs(ceph): editorial pass to align docs with current code and CI/CD

Rewrite the ceph docs against the actual code rather than the pre-refactor
state:

- partition/region/ceph-cluster layout throughout: state keys
  yucca/<partition>/<region>/<stack>, inventories <partition>-<region>/<cluster>,
  the real clusters.auto.tfvars schema (partition/region, not environment/datacenter)
- sietch reframed as staging/austin with secrets in yucca_tf_staging; vault
  hierarchy flipped from dev-primary to staging-primary
- live CI/CD (.github/workflows/infra.yml): per-partition read/write service
  accounts delivered as GitHub secrets (OP_TF_YUCCA_<ENV>_ENV[_WRITE]),
  plan/apply gating, NetBird overlay (Tailscale retired)
- correct the CEPH_ENV guidance (export works; deliberately kept out of mise
  [env]) across README, CONTRIBUTING, scripts.md, adding-a-cluster
- drop the obsolete "read SA token from a 1P item" dance from the runbooks
- fix stale vaults, paths, examples, and the inventory-provision.ini name

* docs(ceph): transliterate docs to plain ASCII

Replace non-ASCII punctuation and box-drawing with ASCII equivalents across
the ceph docs: em/en dashes to --/-, middot separators to commas, arrows to
->, directory-tree box-drawing to |-- / `--, section sign to "section", x for
the multiply glyph, and >= / <= / ~ for the math glyphs. No content changes.

* docs(ceph): fix broken rotate-ssh-key link in scripts.md

The "Related" link pointed at runbooks/rotate-ssh-key.md, which does not
exist (a pre-existing dangling link). SSH-key rotation lives in
rotate-secrets.md; point at its "Rotating SSH keys" section.

* docs(ceph): style polish from per-doc review

Tighten verbal texture flagged by a per-doc style pass; no structural changes.

- correctness: ansible-play.sh runs ansible-playbook, it does not exec, so the
  trap fires from the still-alive wrapper -- fix the "exec"/SIGKILL claims in
  scripts.md, architecture.md, secrets.md
- cut recurring tics: "DR belt"/"belt-and-suspenders" -> "disaster recovery",
  "a lost laptop is a non-event", "by design", "system mesh"
- unstuff long dash/semicolon sentences in architecture (vault-password
  history, provision/baseline split), secrets (SSH-key paragraph), patterns
- recover-bad-tofu-apply: move the dormant-1P-items aside into one note,
  consolidate the repeated caveats
- misc: naming grammar fix + drop trivia, hardware "would"/"blindly",
  rotate-secrets "after confidence", drop a dead snippet line, fix the 16+
  numeric hedge

* docs(ceph): make add-node and recover runbooks CI-aware

Now that infra.yml applies the stacks and runs the full ceph convergence on
merge, refresh the two runbooks the pipeline changed:

- add-node: lead with the manual-vs-CI split. The TF + host_vars change is a
  PR; the only operator-only step is the physical provisioning (live-image
  boot + provision.yml), which CI can't do; baseline/tune/join/harden run in
  CI on merge. Keep the by-hand convergence as a documented fallback.
- recover-bad-tofu-apply: note that applies now run in CI with the partition
  write SA, so the bad apply is usually a failed CI run; CI does not self-heal,
  recovery is operator-run locally.

* docs(tf,talos): finish ADR purge into tf, fix sietch vault, mark talos dormant

- complete the ADR removal that stopped at ansible/ceph: drop the dangling
  ADR-009/010 references from tf/README.md (link + related line) and the ceph
  module / stack code comments, so no ADR trace remains repo-wide
- tf/README: the sietch cluster example uses yucca_tf_staging (was yucca_tf_dev)
- ansible/talos: add a "second-class, not actively used" status banner to the
  README and architecture doc so readers don't treat the converged/libvirt
  Talos docs as live

Left untouched: the Tailscale / SA-token / CI sections of tf/README -- those
are mid-migration in another owner's lane (bye-tailscale is in flight; fabric
still rides Tailscale by design).
2026-06-29 13:04:15 -07:00

11 KiB

Wrapper Scripts

The scripts under ansible/ceph/scripts/ sit between mise tasks and the underlying CLIs (ansible-playbook, op, ssh). They exist to:

  • Keep secrets out of argv -- op inject writes to a file; the file is passed as --extra-vars @<path>, never expanded inline.
  • Fail closed -- if 1Password is unreachable or the inventory is missing, the wrappers exit before invoking ansible. The raw tools' error messages degrade silently in these cases.
  • Offer better error messages -- "inventory not found at <path> -- has tofu apply been run?" beats "No inventory was parsed".

For the architectural role these scripts play, see architecture.md section 8.

Setting CEPH_ENV

Two of the three wrappers (ansible-play.sh, preflight.sh) need CEPH_ENV pointing at the target cluster's inventory.ini. Set it however suits you -- export it once per shell, or prefix a single invocation:

# Export once per shell (recommended for a working session):
export CEPH_ENV=inventories/staging-austin/sietch/inventory.ini
mise run preflight
scripts/ansible-play.sh status.yml

# Or prefix a single invocation:
CEPH_ENV=inventories/staging-austin/sietch/inventory.ini mise run preflight

Both forms work. CEPH_ENV is deliberately not declared in the ansible/ceph/.mise.toml [env] block: a value there would override your shell's value and silently pin every task to the default cluster. By leaving it out, your exported (or inline-prefixed) value passes straight through, and each mise run task falls back to sietch (inventories/staging-austin/sietch/inventory.ini) only when CEPH_ENV is unset. Calling the scripts directly behaves the same -- the wrapper reads CEPH_ENV from the environment.

Quick reference

Script What it does Called by
ansible-play.sh Render secrets.yml.tpl via op inject to a tmpfile, then exec ansible-playbook Every mise run task that runs a playbook
install-ssh-keys.sh Pull per-cluster ansible-iac SSH keys from 1P into ~/.ssh/ Operator (once per workstation / after rotation)
preflight.sh Verify TF artifacts, 1P session, SSH reachability, Python on targets mise run preflight

ansible-play.sh

Wrapper around ansible-playbook that resolves TF-rendered secrets via op inject and passes them as an ephemeral extra-vars file.

Synopsis

CEPH_ENV=inventories/<partition>-<region>/<cluster>/inventory.ini \
  scripts/ansible-play.sh <playbook.yml> [ansible-playbook args...]

What it does

  1. Validates $CEPH_ENV is set and the inventory file exists.
  2. Derives the secrets template path: $(dirname $CEPH_ENV)/secrets.yml.tpl.
  3. Verifies op account get succeeds (1P desktop session or OP_SERVICE_ACCOUNT_TOKEN). Fails fast if unavailable.
  4. mktemps a tmpfile, chmod 600, registers a trap to delete it on EXIT INT TERM (covers a clean exit and operator Ctrl-C / TERM).
  5. Runs op inject -f -i <template> -o <tmpfile>. Fails if any op:// reference can't be resolved.
  6. Runs ansible-playbook -i $CEPH_ENV --extra-vars @<tmpfile> <args>.

The wrapper stays the parent process (no exec), so the trap removes the tmpfile when the play finishes or is interrupted.

Environment

Variable Required Purpose
CEPH_ENV yes Path to the target cluster's inventory.ini
OP_SERVICE_ACCOUNT_TOKEN no CI headless auth. Falls back to op desktop session if unset.

Arguments

Everything after the playbook name is passed to ansible-playbook verbatim. Common patterns:

scripts/ansible-play.sh baseline.yml --check --diff
scripts/ansible-play.sh deploy-ceph.yml --tags rgw,monitoring
scripts/ansible-play.sh destroy-ceph.yml -e yes_destroy_ceph=true -e destroy_target_domain=staging.austin.int.futo.cloud

The destroy playbook requires both safety gates:

  • yes_destroy_ceph=true -- explicit confirmation (otherwise the play refuses to run)
  • destroy_target_domain=<cluster domain> -- must match the inventory's cluster_domain. Mismatch aborts the play, guarding against running destroy with the wrong CEPH_ENV.

The mise run destroy task (in .mise.toml) builds these arguments automatically from CEPH_ENV and adds an interactive [y/N] prompt -- prefer it over invoking the wrapper directly.

Exit codes

Code Meaning
0 ansible-playbook completed successfully
1 Inventory file missing or secrets template missing
2 1Password session unavailable (op account get failed)
3 op inject failed (bad op:// reference, missing item, etc.)
>=4 ansible-playbook's own exit code (unreachable hosts, failed tasks)

Examples

# Standard deploy
CEPH_ENV=inventories/staging-austin/sietch/inventory.ini \
  scripts/ansible-play.sh deploy-ceph.yml

# Dry-run a role via tags
CEPH_ENV=inventories/staging-austin/sietch/inventory.ini \
  scripts/ansible-play.sh site.yml --check --diff --tags baseline

# CI / headless (SA token from env)
OP_SERVICE_ACCOUNT_TOKEN="$(...)" \
  CEPH_ENV=inventories/staging-austin/sietch/inventory.ini \
  scripts/ansible-play.sh status.yml

install-ssh-keys.sh

Idempotent installer that pulls per-cluster ansible-iac SSH keypairs from 1Password into ~/.ssh/. For new workstations or after a key rotation.

Synopsis

scripts/install-ssh-keys.sh [cluster...]

If no cluster arguments are given, installs keys for every known cluster (sietch).

What it does

For each target cluster:

  1. Resolves the 1P item name: <CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY in the cluster's vault (yucca_tf_staging for sietch).
  2. Resolves the target filename: ~/.ssh/id_ed25519_<cluster>.
  3. If the key already exists on disk:
    • Compares fingerprints (local vs 1P).
    • Match -> skip (idempotent re-run).
    • Mismatch -> refuse to overwrite. Prints the mv command the operator should run manually. Exit 2.
  4. If the key is missing:
    • umask 077.
    • Writes private_key -> ~/.ssh/id_ed25519_<cluster> (chmod 600).
    • Writes public_key -> ~/.ssh/id_ed25519_<cluster>.pub (chmod 644).
    • Prints the installed key's fingerprint.

Environment

Variable Required Purpose
OP_SERVICE_ACCOUNT_TOKEN no CI headless auth. Read-only scope on the cluster's yucca_tf_* vault is enough.

Arguments

Zero or more cluster short names (sietch). With no args, all known clusters are installed.

Exit codes

Code Meaning
0 All requested keys installed (or already present and matching)
1 Unknown cluster name
2 Fingerprint mismatch -- refused to overwrite existing key on disk

Examples

# Install all known keys
scripts/install-ssh-keys.sh

# Just one cluster
scripts/install-ssh-keys.sh sietch

# Re-run after a rotation (operator already moved the old key aside manually)
mv ~/.ssh/id_ed25519_sietch     ~/.ssh/id_ed25519_sietch.20260423.bak
mv ~/.ssh/id_ed25519_sietch.pub ~/.ssh/id_ed25519_sietch.pub.20260423.bak
scripts/install-ssh-keys.sh sietch

preflight.sh

Read-only smoke test -- verifies the controller environment is ready to run destructive playbooks against the target cluster. Surfaced via mise run preflight.

Synopsis

CEPH_ENV=inventories/<partition>-<region>/<cluster>/inventory.ini scripts/preflight.sh

What it checks

Controller:

  • scripts/ansible-play.sh is executable
  • Inventory file exists (TF-rendered)
  • Secrets template exists (TF-rendered)
  • SSH private + public keys exist on disk
  • ansible and op CLIs installed
  • 1Password session live (op account get)

Secrets:

  • op inject resolves the cluster's secrets.yml.tpl successfully
  • Resolved file has a non-empty vault_ops_password line (sanity check that op isn't silently substituting empty strings)

Target connectivity (one check per host in ceph_nodes):

  • SSH reachability via ansible -m ping
  • Python 3 available via ansible -m raw

Environment

Variable Required Purpose
CEPH_ENV yes Path to the target cluster's inventory.ini
OP_SERVICE_ACCOUNT_TOKEN no CI headless auth. Falls back to op desktop session.

Exit codes

Code Meaning
0 All checks passed
1 One or more checks failed -- summary printed, unsafe to proceed

Warnings (non-blocking) are reported in the summary but don't affect exit.

Examples

# Via mise (recommended)
mise run preflight

# Direct, against sietch
CEPH_ENV=inventories/staging-austin/sietch/inventory.ini \
  scripts/preflight.sh

Adding a new wrapper

Follow these conventions:

  1. Fail closed -- set -euo pipefail at the top. Any uncaught error aborts the script.
  2. Validate inputs early -- check required env vars (: "${CEPH_ENV:?...}"), then check that referenced files exist, before doing any real work.
  3. Never interpolate secrets into argv -- write them to a mktemp'd file (chmod 600) and pass the path. Always trap 'rm -f "$TMP"' EXIT INT TERM.
  4. Distinct exit codes -- the caller (mise task or another script) should be able to tell "inventory missing" from "1P unreachable" from "ansible failed" without parsing stderr.
  5. Idempotent where plausible -- re-running the script should not make things worse. Prefer skip-if-already-correct over unconditional overwrite.
  6. Cross-reference from architecture.md section 8 and add a section to this file with the same shape as the ones above.