* docs(ceph): inline ADR rationale and drop the ADR set Fold each linked ADR's rationale into the prose it supported, then remove the ADR files, the README index row, and the stray code-comment reference -- no ADR trace remains. True up the docs to the partition/region/ceph-cluster layout (#222) in the same pass: state keys (yucca/<partition>/<region>/<stack>), inventory paths (<partition>-<region>/<cluster>), stack dirs, and the recover runbook's S3 state paths. Reframe architecture's env section around partitions/regions with sietch as staging/austin. * docs(ceph): editorial pass to align docs with current code and CI/CD Rewrite the ceph docs against the actual code rather than the pre-refactor state: - partition/region/ceph-cluster layout throughout: state keys yucca/<partition>/<region>/<stack>, inventories <partition>-<region>/<cluster>, the real clusters.auto.tfvars schema (partition/region, not environment/datacenter) - sietch reframed as staging/austin with secrets in yucca_tf_staging; vault hierarchy flipped from dev-primary to staging-primary - live CI/CD (.github/workflows/infra.yml): per-partition read/write service accounts delivered as GitHub secrets (OP_TF_YUCCA_<ENV>_ENV[_WRITE]), plan/apply gating, NetBird overlay (Tailscale retired) - correct the CEPH_ENV guidance (export works; deliberately kept out of mise [env]) across README, CONTRIBUTING, scripts.md, adding-a-cluster - drop the obsolete "read SA token from a 1P item" dance from the runbooks - fix stale vaults, paths, examples, and the inventory-provision.ini name * docs(ceph): transliterate docs to plain ASCII Replace non-ASCII punctuation and box-drawing with ASCII equivalents across the ceph docs: em/en dashes to --/-, middot separators to commas, arrows to ->, directory-tree box-drawing to |-- / `--, section sign to "section", x for the multiply glyph, and >= / <= / ~ for the math glyphs. No content changes. * docs(ceph): fix broken rotate-ssh-key link in scripts.md The "Related" link pointed at runbooks/rotate-ssh-key.md, which does not exist (a pre-existing dangling link). SSH-key rotation lives in rotate-secrets.md; point at its "Rotating SSH keys" section. * docs(ceph): style polish from per-doc review Tighten verbal texture flagged by a per-doc style pass; no structural changes. - correctness: ansible-play.sh runs ansible-playbook, it does not exec, so the trap fires from the still-alive wrapper -- fix the "exec"/SIGKILL claims in scripts.md, architecture.md, secrets.md - cut recurring tics: "DR belt"/"belt-and-suspenders" -> "disaster recovery", "a lost laptop is a non-event", "by design", "system mesh" - unstuff long dash/semicolon sentences in architecture (vault-password history, provision/baseline split), secrets (SSH-key paragraph), patterns - recover-bad-tofu-apply: move the dormant-1P-items aside into one note, consolidate the repeated caveats - misc: naming grammar fix + drop trivia, hardware "would"/"blindly", rotate-secrets "after confidence", drop a dead snippet line, fix the 16+ numeric hedge * docs(ceph): make add-node and recover runbooks CI-aware Now that infra.yml applies the stacks and runs the full ceph convergence on merge, refresh the two runbooks the pipeline changed: - add-node: lead with the manual-vs-CI split. The TF + host_vars change is a PR; the only operator-only step is the physical provisioning (live-image boot + provision.yml), which CI can't do; baseline/tune/join/harden run in CI on merge. Keep the by-hand convergence as a documented fallback. - recover-bad-tofu-apply: note that applies now run in CI with the partition write SA, so the bad apply is usually a failed CI run; CI does not self-heal, recovery is operator-run locally. * docs(tf,talos): finish ADR purge into tf, fix sietch vault, mark talos dormant - complete the ADR removal that stopped at ansible/ceph: drop the dangling ADR-009/010 references from tf/README.md (link + related line) and the ceph module / stack code comments, so no ADR trace remains repo-wide - tf/README: the sietch cluster example uses yucca_tf_staging (was yucca_tf_dev) - ansible/talos: add a "second-class, not actively used" status banner to the README and architecture doc so readers don't treat the converged/libvirt Talos docs as live Left untouched: the Tailscale / SA-token / CI sections of tf/README -- those are mid-migration in another owner's lane (bye-tailscale is in flight; fabric still rides Tailscale by design).
11 KiB
Wrapper Scripts
The scripts under ansible/ceph/scripts/ sit between mise tasks and the
underlying CLIs (ansible-playbook, op, ssh). They exist to:
- Keep secrets out of argv --
op injectwrites to a file; the file is passed as--extra-vars @<path>, never expanded inline. - Fail closed -- if 1Password is unreachable or the inventory is missing, the wrappers exit before invoking ansible. The raw tools' error messages degrade silently in these cases.
- Offer better error messages -- "inventory not found at
<path>-- hastofu applybeen run?" beats "No inventory was parsed".
For the architectural role these scripts play, see architecture.md section 8.
Setting CEPH_ENV
Two of the three wrappers (ansible-play.sh, preflight.sh) need CEPH_ENV
pointing at the target cluster's inventory.ini. Set it however suits you --
export it once per shell, or prefix a single invocation:
# Export once per shell (recommended for a working session):
export CEPH_ENV=inventories/staging-austin/sietch/inventory.ini
mise run preflight
scripts/ansible-play.sh status.yml
# Or prefix a single invocation:
CEPH_ENV=inventories/staging-austin/sietch/inventory.ini mise run preflight
Both forms work. CEPH_ENV is deliberately not declared in
the ansible/ceph/.mise.toml [env] block: a value there would override
your shell's value and silently pin every task to the default cluster. By
leaving it out, your exported (or inline-prefixed) value passes straight
through, and each mise run task falls back to sietch
(inventories/staging-austin/sietch/inventory.ini) only when CEPH_ENV is
unset. Calling the scripts directly behaves the same -- the wrapper reads
CEPH_ENV from the environment.
Quick reference
| Script | What it does | Called by |
|---|---|---|
ansible-play.sh |
Render secrets.yml.tpl via op inject to a tmpfile, then exec ansible-playbook |
Every mise run task that runs a playbook |
install-ssh-keys.sh |
Pull per-cluster ansible-iac SSH keys from 1P into ~/.ssh/ |
Operator (once per workstation / after rotation) |
preflight.sh |
Verify TF artifacts, 1P session, SSH reachability, Python on targets | mise run preflight |
ansible-play.sh
Wrapper around ansible-playbook that resolves TF-rendered secrets via
op inject and passes them as an ephemeral extra-vars file.
Synopsis
CEPH_ENV=inventories/<partition>-<region>/<cluster>/inventory.ini \
scripts/ansible-play.sh <playbook.yml> [ansible-playbook args...]
What it does
- Validates
$CEPH_ENVis set and the inventory file exists. - Derives the secrets template path:
$(dirname $CEPH_ENV)/secrets.yml.tpl. - Verifies
op account getsucceeds (1P desktop session orOP_SERVICE_ACCOUNT_TOKEN). Fails fast if unavailable. mktemps a tmpfile,chmod 600, registers atrapto delete it onEXIT INT TERM(covers a clean exit and operator Ctrl-C / TERM).- Runs
op inject -f -i <template> -o <tmpfile>. Fails if anyop://reference can't be resolved. - Runs
ansible-playbook -i $CEPH_ENV --extra-vars @<tmpfile> <args>.
The wrapper stays the parent process (no exec), so the trap removes the
tmpfile when the play finishes or is interrupted.
Environment
| Variable | Required | Purpose |
|---|---|---|
CEPH_ENV |
yes | Path to the target cluster's inventory.ini |
OP_SERVICE_ACCOUNT_TOKEN |
no | CI headless auth. Falls back to op desktop session if unset. |
Arguments
Everything after the playbook name is passed to ansible-playbook verbatim.
Common patterns:
scripts/ansible-play.sh baseline.yml --check --diff
scripts/ansible-play.sh deploy-ceph.yml --tags rgw,monitoring
scripts/ansible-play.sh destroy-ceph.yml -e yes_destroy_ceph=true -e destroy_target_domain=staging.austin.int.futo.cloud
The destroy playbook requires both safety gates:
yes_destroy_ceph=true-- explicit confirmation (otherwise the play refuses to run)destroy_target_domain=<cluster domain>-- must match the inventory'scluster_domain. Mismatch aborts the play, guarding against running destroy with the wrongCEPH_ENV.
The mise run destroy task (in .mise.toml) builds these arguments automatically from CEPH_ENV and adds an interactive [y/N] prompt -- prefer it over invoking the wrapper directly.
Exit codes
| Code | Meaning |
|---|---|
| 0 | ansible-playbook completed successfully |
| 1 | Inventory file missing or secrets template missing |
| 2 | 1Password session unavailable (op account get failed) |
| 3 | op inject failed (bad op:// reference, missing item, etc.) |
| >=4 | ansible-playbook's own exit code (unreachable hosts, failed tasks) |
Examples
# Standard deploy
CEPH_ENV=inventories/staging-austin/sietch/inventory.ini \
scripts/ansible-play.sh deploy-ceph.yml
# Dry-run a role via tags
CEPH_ENV=inventories/staging-austin/sietch/inventory.ini \
scripts/ansible-play.sh site.yml --check --diff --tags baseline
# CI / headless (SA token from env)
OP_SERVICE_ACCOUNT_TOKEN="$(...)" \
CEPH_ENV=inventories/staging-austin/sietch/inventory.ini \
scripts/ansible-play.sh status.yml
Related
- architecture.md section 9.2 -- deploy data flow
- secrets.md -- what's in
secrets.yml.tpl
install-ssh-keys.sh
Idempotent installer that pulls per-cluster ansible-iac SSH keypairs from
1Password into ~/.ssh/. For new workstations or after a key rotation.
Synopsis
scripts/install-ssh-keys.sh [cluster...]
If no cluster arguments are given, installs keys for every known cluster
(sietch).
What it does
For each target cluster:
- Resolves the 1P item name:
<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEYin the cluster's vault (yucca_tf_stagingfor sietch). - Resolves the target filename:
~/.ssh/id_ed25519_<cluster>. - If the key already exists on disk:
- Compares fingerprints (local vs 1P).
- Match -> skip (idempotent re-run).
- Mismatch -> refuse to overwrite. Prints the
mvcommand the operator should run manually. Exit 2.
- If the key is missing:
umask 077.- Writes
private_key->~/.ssh/id_ed25519_<cluster>(chmod 600). - Writes
public_key->~/.ssh/id_ed25519_<cluster>.pub(chmod 644). - Prints the installed key's fingerprint.
Environment
| Variable | Required | Purpose |
|---|---|---|
OP_SERVICE_ACCOUNT_TOKEN |
no | CI headless auth. Read-only scope on the cluster's yucca_tf_* vault is enough. |
Arguments
Zero or more cluster short names (sietch). With no args, all
known clusters are installed.
Exit codes
| Code | Meaning |
|---|---|
| 0 | All requested keys installed (or already present and matching) |
| 1 | Unknown cluster name |
| 2 | Fingerprint mismatch -- refused to overwrite existing key on disk |
Examples
# Install all known keys
scripts/install-ssh-keys.sh
# Just one cluster
scripts/install-ssh-keys.sh sietch
# Re-run after a rotation (operator already moved the old key aside manually)
mv ~/.ssh/id_ed25519_sietch ~/.ssh/id_ed25519_sietch.20260423.bak
mv ~/.ssh/id_ed25519_sietch.pub ~/.ssh/id_ed25519_sietch.pub.20260423.bak
scripts/install-ssh-keys.sh sietch
Related
- runbooks/rotate-secrets.md -- the full SSH-key rotation flow
- secrets.md -- why SSH keys live in 1P (durable, auditable, one
op://model)
preflight.sh
Read-only smoke test -- verifies the controller environment is ready to run
destructive playbooks against the target cluster. Surfaced via
mise run preflight.
Synopsis
CEPH_ENV=inventories/<partition>-<region>/<cluster>/inventory.ini scripts/preflight.sh
What it checks
Controller:
scripts/ansible-play.shis executable- Inventory file exists (TF-rendered)
- Secrets template exists (TF-rendered)
- SSH private + public keys exist on disk
ansibleandopCLIs installed- 1Password session live (
op account get)
Secrets:
op injectresolves the cluster'ssecrets.yml.tplsuccessfully- Resolved file has a non-empty
vault_ops_passwordline (sanity check that op isn't silently substituting empty strings)
Target connectivity (one check per host in ceph_nodes):
- SSH reachability via
ansible -m ping - Python 3 available via
ansible -m raw
Environment
| Variable | Required | Purpose |
|---|---|---|
CEPH_ENV |
yes | Path to the target cluster's inventory.ini |
OP_SERVICE_ACCOUNT_TOKEN |
no | CI headless auth. Falls back to op desktop session. |
Exit codes
| Code | Meaning |
|---|---|
| 0 | All checks passed |
| 1 | One or more checks failed -- summary printed, unsafe to proceed |
Warnings (non-blocking) are reported in the summary but don't affect exit.
Examples
# Via mise (recommended)
mise run preflight
# Direct, against sietch
CEPH_ENV=inventories/staging-austin/sietch/inventory.ini \
scripts/preflight.sh
Related
- architecture.md section 8 -- where the wrappers fit
Adding a new wrapper
Follow these conventions:
- Fail closed --
set -euo pipefailat the top. Any uncaught error aborts the script. - Validate inputs early -- check required env vars (
: "${CEPH_ENV:?...}"), then check that referenced files exist, before doing any real work. - Never interpolate secrets into argv -- write them to a
mktemp'd file (chmod 600) and pass the path. Alwaystrap 'rm -f "$TMP"' EXIT INT TERM. - Distinct exit codes -- the caller (mise task or another script) should be able to tell "inventory missing" from "1P unreachable" from "ansible failed" without parsing stderr.
- Idempotent where plausible -- re-running the script should not make things worse. Prefer skip-if-already-correct over unconditional overwrite.
- Cross-reference from architecture.md section 8 and add a section to this file with the same shape as the ones above.