* docs(ceph): inline ADR rationale and drop the ADR set Fold each linked ADR's rationale into the prose it supported, then remove the ADR files, the README index row, and the stray code-comment reference -- no ADR trace remains. True up the docs to the partition/region/ceph-cluster layout (#222) in the same pass: state keys (yucca/<partition>/<region>/<stack>), inventory paths (<partition>-<region>/<cluster>), stack dirs, and the recover runbook's S3 state paths. Reframe architecture's env section around partitions/regions with sietch as staging/austin. * docs(ceph): editorial pass to align docs with current code and CI/CD Rewrite the ceph docs against the actual code rather than the pre-refactor state: - partition/region/ceph-cluster layout throughout: state keys yucca/<partition>/<region>/<stack>, inventories <partition>-<region>/<cluster>, the real clusters.auto.tfvars schema (partition/region, not environment/datacenter) - sietch reframed as staging/austin with secrets in yucca_tf_staging; vault hierarchy flipped from dev-primary to staging-primary - live CI/CD (.github/workflows/infra.yml): per-partition read/write service accounts delivered as GitHub secrets (OP_TF_YUCCA_<ENV>_ENV[_WRITE]), plan/apply gating, NetBird overlay (Tailscale retired) - correct the CEPH_ENV guidance (export works; deliberately kept out of mise [env]) across README, CONTRIBUTING, scripts.md, adding-a-cluster - drop the obsolete "read SA token from a 1P item" dance from the runbooks - fix stale vaults, paths, examples, and the inventory-provision.ini name * docs(ceph): transliterate docs to plain ASCII Replace non-ASCII punctuation and box-drawing with ASCII equivalents across the ceph docs: em/en dashes to --/-, middot separators to commas, arrows to ->, directory-tree box-drawing to |-- / `--, section sign to "section", x for the multiply glyph, and >= / <= / ~ for the math glyphs. No content changes. * docs(ceph): fix broken rotate-ssh-key link in scripts.md The "Related" link pointed at runbooks/rotate-ssh-key.md, which does not exist (a pre-existing dangling link). SSH-key rotation lives in rotate-secrets.md; point at its "Rotating SSH keys" section. * docs(ceph): style polish from per-doc review Tighten verbal texture flagged by a per-doc style pass; no structural changes. - correctness: ansible-play.sh runs ansible-playbook, it does not exec, so the trap fires from the still-alive wrapper -- fix the "exec"/SIGKILL claims in scripts.md, architecture.md, secrets.md - cut recurring tics: "DR belt"/"belt-and-suspenders" -> "disaster recovery", "a lost laptop is a non-event", "by design", "system mesh" - unstuff long dash/semicolon sentences in architecture (vault-password history, provision/baseline split), secrets (SSH-key paragraph), patterns - recover-bad-tofu-apply: move the dormant-1P-items aside into one note, consolidate the repeated caveats - misc: naming grammar fix + drop trivia, hardware "would"/"blindly", rotate-secrets "after confidence", drop a dead snippet line, fix the 16+ numeric hedge * docs(ceph): make add-node and recover runbooks CI-aware Now that infra.yml applies the stacks and runs the full ceph convergence on merge, refresh the two runbooks the pipeline changed: - add-node: lead with the manual-vs-CI split. The TF + host_vars change is a PR; the only operator-only step is the physical provisioning (live-image boot + provision.yml), which CI can't do; baseline/tune/join/harden run in CI on merge. Keep the by-hand convergence as a documented fallback. - recover-bad-tofu-apply: note that applies now run in CI with the partition write SA, so the bad apply is usually a failed CI run; CI does not self-heal, recovery is operator-run locally. * docs(tf,talos): finish ADR purge into tf, fix sietch vault, mark talos dormant - complete the ADR removal that stopped at ansible/ceph: drop the dangling ADR-009/010 references from tf/README.md (link + related line) and the ceph module / stack code comments, so no ADR trace remains repo-wide - tf/README: the sietch cluster example uses yucca_tf_staging (was yucca_tf_dev) - ansible/talos: add a "second-class, not actively used" status banner to the README and architecture doc so readers don't treat the converged/libvirt Talos docs as live Left untouched: the Tailscale / SA-token / CI sections of tf/README -- those are mid-migration in another owner's lane (bye-tailscale is in flight; fabric still rides Tailscale by design).
7.6 KiB
Secrets
This project uses a TF-provisions, op-injects, ansible-consumes model. No
ansible-vault, no encrypted vault.yml in git, no custom password-file script.
For how secrets fit into the broader architecture, see architecture.md section 5 (1Password).
flowchart TB
ONEP[("1Password<br/>yucca_tf, yucca_tf_staging, ...<br/><i>source of truth</i>")]
TF[Terraform / Tofu<br/>tf/deployment/staging/austin/ceph/]
REPO[/"inventories/<partition>-<region>/<cluster>/<br/>inventory.ini (TF-gen, gitignored)<br/>secrets.yml.tpl (TF-gen, gitignored)"/]
WRAP[scripts/ansible-play.sh<br/><i>mktemp + op inject -> run ansible-playbook --extra-vars @tmp</i>]
ANS[ansible-playbook]
TF -.->|reads via op run --env-file| ONEP
TF -->|renders| REPO
WRAP -->|reads template| REPO
WRAP -.->|op inject -f<br/>at play time| ONEP
WRAP --> ANS
What lives where
| Vault | Purpose | Who writes |
|---|---|---|
yucca_tf (team-shared) |
Cross-partition shared state: TF state S3 credentials | Operator (manual) |
yucca_tf_staging (team-shared) |
Live values for the staging partition (sietch today) -- <CLUSTER>_CEPH_* items |
Superuser service account (TF) + operator via op CLI |
yucca_tf_staging_manual (team-shared) |
Human-fillable staging placeholders (3rd-party API tokens, OAuth client secrets) -- not yet used by ceph-cluster | Operator (manual) |
Each partition has the same pair; the dev and prod analogues
(yucca_tf_dev(_manual), yucca_tf_prod_manual) land as siblings. The vault
a given cluster reads from is declared per-cluster in
tf/deployment/<partition>/<region>/ceph/clusters.auto.tfvars (field vault)
-- sietch points at yucca_tf_staging. TF derives item paths from that field
at render time; changing it + tofu apply re-renders secrets.yml.tpl with
the new vault path.
Item naming
Format: <CLUSTER>_CEPH_<ROLE>_PASSWORD
Password items (category Password, consumed via op inject at
playbook time):
| Item | Field | Consumed as |
|---|---|---|
SIETCH_CEPH_OPS_PASSWORD |
password |
vault_ops_password |
SIETCH_CEPH_DASHBOARD_PASSWORD |
password |
vault_ceph_dashboard_password |
SIETCH_CEPH_GRAFANA_PASSWORD |
password |
vault_grafana_admin_password |
SSH Key items (category SSH Key, consumed via
scripts/install-ssh-keys.sh on operator workstations and
rotate-ssh-key.yml on cluster nodes). SSH keys use the same storage model as
the passwords. They are generated natively in 1Password
(--ssh-generate-key=ed25519), so the private key never touches operator disk
at creation. Rotation is forward-only: generate a new key, distribute the
public half, retire the old one.
| Item | Field | Consumed as |
|---|---|---|
SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY |
private_key |
~/.ssh/id_ed25519_sietch on operator workstation |
SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY |
public_key |
sietch nodes' ansible-iac@:~/.ssh/authorized_keys |
S3 service-user items (predetermined keys passed to
radosgw-admin user create --access-key=X --secret-key=Y at deploy
time -- Yucca app / restic client can be pre-configured with matching
credentials without waiting for post-bootstrap capture):
| Item | Field | Consumed as |
|---|---|---|
SIETCH_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY |
password |
vault_s3_restic_access_key -> ceph_rgw_s3_user_access_key |
SIETCH_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY |
password |
vault_s3_restic_secret_key -> ceph_rgw_s3_user_secret_key |
Disaster-recovery items (populated by mise run capture after
deploy -- stored in 1P for recovery if the bootstrap node's filesystem
is lost):
| Item | Field | Source |
|---|---|---|
<CLUSTER>_CEPH_RGW_TLS_CERT |
password (concealed) |
/etc/ceph/rgw-ssl.crt on bootstrap |
<CLUSTER>_CEPH_RGW_TLS_KEY |
password (concealed) |
/etc/ceph/rgw-ssl.key on bootstrap |
<CLUSTER>_CEPH_CLIENT_ADMIN_KEYRING |
password (concealed) |
/etc/ceph/ceph.client.admin.keyring on bootstrap |
Items are created on the first mise run capture; later runs overwrite them
only if the content changed.
Item names are derived in tf/shared/modules/ceph-cluster/main.tf
(local.secret_prefix). Hardcoded CEPH (not role_in_hostname) so every
Ceph-project secret grep-matches *_CEPH_* regardless of whether hostnames
use ceph, osd, or mon as the role segment.
Runtime flow
The op CLI is invoked in three distinct patterns across this project:
| Pattern | Used by | What it does |
|---|---|---|
op run --env-file=tf/.env -- |
Every mise run tf:* task |
Resolves op:// references in a dotenv file, injects as env vars into child process (TF: SA token + AWS creds) |
op inject -f -i tpl -o out |
scripts/ansible-play.sh, Hetzner installimage post-install rendering |
Resolves all op:// references in a file template, writes resolved file |
op read "op://..." |
scripts/install-ssh-keys.sh, rotate-ssh-key.yml, post-deploy-capture.yml |
Reads a single field from a single item to stdout |
The ansible-play.sh flow
scripts/ansible-play.sh wraps every ansible-playbook invocation:
- Verifies
op account getsucceeds -- fails closed if 1Password is locked. mktempa0600tmpfile withtrapcleanup onEXIT/INT/TERM.op inject -f -i <cluster>/secrets.yml.tpl -o $tmpfile-- resolves everyop://reference. Exit non-zero if any reference can't be resolved.- Runs
ansible-playbook --extra-vars @$tmpfile ....
Ansible task references stay unchanged -- vault_ops_password etc. are
regular variables populated from extra-vars (highest precedence).
Full wrapper reference: scripts.md.
CI / headless
.github/workflows/infra.yml sets OP_SERVICE_ACCOUNT_TOKEN per job from
the partition's service-account GitHub secrets -- OP_TF_YUCCA_<PARTITION>_ENV
(read, for plan) and OP_TF_YUCCA_<PARTITION>_ENV_WRITE (write, for
apply). op inject / op run pick it up automatically. lint and check
need zero op credentials -- they don't touch secrets -- so they run on every
PR regardless.
Rotating secrets
See docs/runbooks/rotate-secrets.md.
Adding a new secret
- Add the item name to
local.secretsintf/shared/modules/ceph-cluster/main.tf. - Add the matching line to the secrets template
(
templates/secrets.yml.tpl.tftpl). - Add the ansible variable alias in each cluster's
inventories/<partition>-<region>/<cluster>/group_vars/all/vars.yml. - Create the item manually in the target vault (or let TF do it post
service account):
op item create --vault <vault> --category password --title NEW_SECRET_NAME --generate-password=letters,digits,32. tofu apply-- re-renders templates with the new reference.- Playbooks consuming the variable now have it available.