Files
yucca/ansible/ceph/docs/secrets.md
T
Andy Molenda 8821fad205 docs(ceph): realign with partition/region model and CI/CD, retire ADRs
* docs(ceph): inline ADR rationale and drop the ADR set

Fold each linked ADR's rationale into the prose it supported, then remove
the ADR files, the README index row, and the stray code-comment reference --
no ADR trace remains.

True up the docs to the partition/region/ceph-cluster layout (#222) in the
same pass: state keys (yucca/<partition>/<region>/<stack>), inventory paths
(<partition>-<region>/<cluster>), stack dirs, and the recover runbook's S3
state paths. Reframe architecture's env section around partitions/regions
with sietch as staging/austin.

* docs(ceph): editorial pass to align docs with current code and CI/CD

Rewrite the ceph docs against the actual code rather than the pre-refactor
state:

- partition/region/ceph-cluster layout throughout: state keys
  yucca/<partition>/<region>/<stack>, inventories <partition>-<region>/<cluster>,
  the real clusters.auto.tfvars schema (partition/region, not environment/datacenter)
- sietch reframed as staging/austin with secrets in yucca_tf_staging; vault
  hierarchy flipped from dev-primary to staging-primary
- live CI/CD (.github/workflows/infra.yml): per-partition read/write service
  accounts delivered as GitHub secrets (OP_TF_YUCCA_<ENV>_ENV[_WRITE]),
  plan/apply gating, NetBird overlay (Tailscale retired)
- correct the CEPH_ENV guidance (export works; deliberately kept out of mise
  [env]) across README, CONTRIBUTING, scripts.md, adding-a-cluster
- drop the obsolete "read SA token from a 1P item" dance from the runbooks
- fix stale vaults, paths, examples, and the inventory-provision.ini name

* docs(ceph): transliterate docs to plain ASCII

Replace non-ASCII punctuation and box-drawing with ASCII equivalents across
the ceph docs: em/en dashes to --/-, middot separators to commas, arrows to
->, directory-tree box-drawing to |-- / `--, section sign to "section", x for
the multiply glyph, and >= / <= / ~ for the math glyphs. No content changes.

* docs(ceph): fix broken rotate-ssh-key link in scripts.md

The "Related" link pointed at runbooks/rotate-ssh-key.md, which does not
exist (a pre-existing dangling link). SSH-key rotation lives in
rotate-secrets.md; point at its "Rotating SSH keys" section.

* docs(ceph): style polish from per-doc review

Tighten verbal texture flagged by a per-doc style pass; no structural changes.

- correctness: ansible-play.sh runs ansible-playbook, it does not exec, so the
  trap fires from the still-alive wrapper -- fix the "exec"/SIGKILL claims in
  scripts.md, architecture.md, secrets.md
- cut recurring tics: "DR belt"/"belt-and-suspenders" -> "disaster recovery",
  "a lost laptop is a non-event", "by design", "system mesh"
- unstuff long dash/semicolon sentences in architecture (vault-password
  history, provision/baseline split), secrets (SSH-key paragraph), patterns
- recover-bad-tofu-apply: move the dormant-1P-items aside into one note,
  consolidate the repeated caveats
- misc: naming grammar fix + drop trivia, hardware "would"/"blindly",
  rotate-secrets "after confidence", drop a dead snippet line, fix the 16+
  numeric hedge

* docs(ceph): make add-node and recover runbooks CI-aware

Now that infra.yml applies the stacks and runs the full ceph convergence on
merge, refresh the two runbooks the pipeline changed:

- add-node: lead with the manual-vs-CI split. The TF + host_vars change is a
  PR; the only operator-only step is the physical provisioning (live-image
  boot + provision.yml), which CI can't do; baseline/tune/join/harden run in
  CI on merge. Keep the by-hand convergence as a documented fallback.
- recover-bad-tofu-apply: note that applies now run in CI with the partition
  write SA, so the bad apply is usually a failed CI run; CI does not self-heal,
  recovery is operator-run locally.

* docs(tf,talos): finish ADR purge into tf, fix sietch vault, mark talos dormant

- complete the ADR removal that stopped at ansible/ceph: drop the dangling
  ADR-009/010 references from tf/README.md (link + related line) and the ceph
  module / stack code comments, so no ADR trace remains repo-wide
- tf/README: the sietch cluster example uses yucca_tf_staging (was yucca_tf_dev)
- ansible/talos: add a "second-class, not actively used" status banner to the
  README and architecture doc so readers don't treat the converged/libvirt
  Talos docs as live

Left untouched: the Tailscale / SA-token / CI sections of tf/README -- those
are mid-migration in another owner's lane (bye-tailscale is in flight; fabric
still rides Tailscale by design).
2026-06-29 13:04:15 -07:00

7.6 KiB

Secrets

This project uses a TF-provisions, op-injects, ansible-consumes model. No ansible-vault, no encrypted vault.yml in git, no custom password-file script.

For how secrets fit into the broader architecture, see architecture.md section 5 (1Password).

flowchart TB
    ONEP[("1Password<br/>yucca_tf, yucca_tf_staging, ...<br/><i>source of truth</i>")]
    TF[Terraform / Tofu<br/>tf/deployment/staging/austin/ceph/]
    REPO[/"inventories/&lt;partition&gt;-&lt;region&gt;/&lt;cluster&gt;/<br/>inventory.ini (TF-gen, gitignored)<br/>secrets.yml.tpl (TF-gen, gitignored)"/]
    WRAP[scripts/ansible-play.sh<br/><i>mktemp + op inject -> run ansible-playbook --extra-vars @tmp</i>]
    ANS[ansible-playbook]

    TF -.->|reads via op run --env-file| ONEP
    TF -->|renders| REPO
    WRAP -->|reads template| REPO
    WRAP -.->|op inject -f<br/>at play time| ONEP
    WRAP --> ANS

What lives where

Vault Purpose Who writes
yucca_tf (team-shared) Cross-partition shared state: TF state S3 credentials Operator (manual)
yucca_tf_staging (team-shared) Live values for the staging partition (sietch today) -- <CLUSTER>_CEPH_* items Superuser service account (TF) + operator via op CLI
yucca_tf_staging_manual (team-shared) Human-fillable staging placeholders (3rd-party API tokens, OAuth client secrets) -- not yet used by ceph-cluster Operator (manual)

Each partition has the same pair; the dev and prod analogues (yucca_tf_dev(_manual), yucca_tf_prod_manual) land as siblings. The vault a given cluster reads from is declared per-cluster in tf/deployment/<partition>/<region>/ceph/clusters.auto.tfvars (field vault) -- sietch points at yucca_tf_staging. TF derives item paths from that field at render time; changing it + tofu apply re-renders secrets.yml.tpl with the new vault path.

Item naming

Format: <CLUSTER>_CEPH_<ROLE>_PASSWORD

Password items (category Password, consumed via op inject at playbook time):

Item Field Consumed as
SIETCH_CEPH_OPS_PASSWORD password vault_ops_password
SIETCH_CEPH_DASHBOARD_PASSWORD password vault_ceph_dashboard_password
SIETCH_CEPH_GRAFANA_PASSWORD password vault_grafana_admin_password

SSH Key items (category SSH Key, consumed via scripts/install-ssh-keys.sh on operator workstations and rotate-ssh-key.yml on cluster nodes). SSH keys use the same storage model as the passwords. They are generated natively in 1Password (--ssh-generate-key=ed25519), so the private key never touches operator disk at creation. Rotation is forward-only: generate a new key, distribute the public half, retire the old one.

Item Field Consumed as
SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY private_key ~/.ssh/id_ed25519_sietch on operator workstation
SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY public_key sietch nodes' ansible-iac@:~/.ssh/authorized_keys

S3 service-user items (predetermined keys passed to radosgw-admin user create --access-key=X --secret-key=Y at deploy time -- Yucca app / restic client can be pre-configured with matching credentials without waiting for post-bootstrap capture):

Item Field Consumed as
SIETCH_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY password vault_s3_restic_access_key -> ceph_rgw_s3_user_access_key
SIETCH_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY password vault_s3_restic_secret_key -> ceph_rgw_s3_user_secret_key

Disaster-recovery items (populated by mise run capture after deploy -- stored in 1P for recovery if the bootstrap node's filesystem is lost):

Item Field Source
<CLUSTER>_CEPH_RGW_TLS_CERT password (concealed) /etc/ceph/rgw-ssl.crt on bootstrap
<CLUSTER>_CEPH_RGW_TLS_KEY password (concealed) /etc/ceph/rgw-ssl.key on bootstrap
<CLUSTER>_CEPH_CLIENT_ADMIN_KEYRING password (concealed) /etc/ceph/ceph.client.admin.keyring on bootstrap

Items are created on the first mise run capture; later runs overwrite them only if the content changed.

Item names are derived in tf/shared/modules/ceph-cluster/main.tf (local.secret_prefix). Hardcoded CEPH (not role_in_hostname) so every Ceph-project secret grep-matches *_CEPH_* regardless of whether hostnames use ceph, osd, or mon as the role segment.

Runtime flow

The op CLI is invoked in three distinct patterns across this project:

Pattern Used by What it does
op run --env-file=tf/.env -- Every mise run tf:* task Resolves op:// references in a dotenv file, injects as env vars into child process (TF: SA token + AWS creds)
op inject -f -i tpl -o out scripts/ansible-play.sh, Hetzner installimage post-install rendering Resolves all op:// references in a file template, writes resolved file
op read "op://..." scripts/install-ssh-keys.sh, rotate-ssh-key.yml, post-deploy-capture.yml Reads a single field from a single item to stdout

The ansible-play.sh flow

scripts/ansible-play.sh wraps every ansible-playbook invocation:

  1. Verifies op account get succeeds -- fails closed if 1Password is locked.
  2. mktemp a 0600 tmpfile with trap cleanup on EXIT / INT / TERM.
  3. op inject -f -i <cluster>/secrets.yml.tpl -o $tmpfile -- resolves every op:// reference. Exit non-zero if any reference can't be resolved.
  4. Runs ansible-playbook --extra-vars @$tmpfile ....

Ansible task references stay unchanged -- vault_ops_password etc. are regular variables populated from extra-vars (highest precedence).

Full wrapper reference: scripts.md.

CI / headless

.github/workflows/infra.yml sets OP_SERVICE_ACCOUNT_TOKEN per job from the partition's service-account GitHub secrets -- OP_TF_YUCCA_<PARTITION>_ENV (read, for plan) and OP_TF_YUCCA_<PARTITION>_ENV_WRITE (write, for apply). op inject / op run pick it up automatically. lint and check need zero op credentials -- they don't touch secrets -- so they run on every PR regardless.

Rotating secrets

See docs/runbooks/rotate-secrets.md.

Adding a new secret

  1. Add the item name to local.secrets in tf/shared/modules/ceph-cluster/main.tf.
  2. Add the matching line to the secrets template (templates/secrets.yml.tpl.tftpl).
  3. Add the ansible variable alias in each cluster's inventories/<partition>-<region>/<cluster>/group_vars/all/vars.yml.
  4. Create the item manually in the target vault (or let TF do it post service account): op item create --vault <vault> --category password --title NEW_SECRET_NAME --generate-password=letters,digits,32.
  5. tofu apply -- re-renders templates with the new reference.
  6. Playbooks consuming the variable now have it available.