Files
yucca/ansible/ceph/docs/naming.md
T
Andy Molenda 8821fad205 docs(ceph): realign with partition/region model and CI/CD, retire ADRs
* docs(ceph): inline ADR rationale and drop the ADR set

Fold each linked ADR's rationale into the prose it supported, then remove
the ADR files, the README index row, and the stray code-comment reference --
no ADR trace remains.

True up the docs to the partition/region/ceph-cluster layout (#222) in the
same pass: state keys (yucca/<partition>/<region>/<stack>), inventory paths
(<partition>-<region>/<cluster>), stack dirs, and the recover runbook's S3
state paths. Reframe architecture's env section around partitions/regions
with sietch as staging/austin.

* docs(ceph): editorial pass to align docs with current code and CI/CD

Rewrite the ceph docs against the actual code rather than the pre-refactor
state:

- partition/region/ceph-cluster layout throughout: state keys
  yucca/<partition>/<region>/<stack>, inventories <partition>-<region>/<cluster>,
  the real clusters.auto.tfvars schema (partition/region, not environment/datacenter)
- sietch reframed as staging/austin with secrets in yucca_tf_staging; vault
  hierarchy flipped from dev-primary to staging-primary
- live CI/CD (.github/workflows/infra.yml): per-partition read/write service
  accounts delivered as GitHub secrets (OP_TF_YUCCA_<ENV>_ENV[_WRITE]),
  plan/apply gating, NetBird overlay (Tailscale retired)
- correct the CEPH_ENV guidance (export works; deliberately kept out of mise
  [env]) across README, CONTRIBUTING, scripts.md, adding-a-cluster
- drop the obsolete "read SA token from a 1P item" dance from the runbooks
- fix stale vaults, paths, examples, and the inventory-provision.ini name

* docs(ceph): transliterate docs to plain ASCII

Replace non-ASCII punctuation and box-drawing with ASCII equivalents across
the ceph docs: em/en dashes to --/-, middot separators to commas, arrows to
->, directory-tree box-drawing to |-- / `--, section sign to "section", x for
the multiply glyph, and >= / <= / ~ for the math glyphs. No content changes.

* docs(ceph): fix broken rotate-ssh-key link in scripts.md

The "Related" link pointed at runbooks/rotate-ssh-key.md, which does not
exist (a pre-existing dangling link). SSH-key rotation lives in
rotate-secrets.md; point at its "Rotating SSH keys" section.

* docs(ceph): style polish from per-doc review

Tighten verbal texture flagged by a per-doc style pass; no structural changes.

- correctness: ansible-play.sh runs ansible-playbook, it does not exec, so the
  trap fires from the still-alive wrapper -- fix the "exec"/SIGKILL claims in
  scripts.md, architecture.md, secrets.md
- cut recurring tics: "DR belt"/"belt-and-suspenders" -> "disaster recovery",
  "a lost laptop is a non-event", "by design", "system mesh"
- unstuff long dash/semicolon sentences in architecture (vault-password
  history, provision/baseline split), secrets (SSH-key paragraph), patterns
- recover-bad-tofu-apply: move the dormant-1P-items aside into one note,
  consolidate the repeated caveats
- misc: naming grammar fix + drop trivia, hardware "would"/"blindly",
  rotate-secrets "after confidence", drop a dead snippet line, fix the 16+
  numeric hedge

* docs(ceph): make add-node and recover runbooks CI-aware

Now that infra.yml applies the stacks and runs the full ceph convergence on
merge, refresh the two runbooks the pipeline changed:

- add-node: lead with the manual-vs-CI split. The TF + host_vars change is a
  PR; the only operator-only step is the physical provisioning (live-image
  boot + provision.yml), which CI can't do; baseline/tune/join/harden run in
  CI on merge. Keep the by-hand convergence as a documented fallback.
- recover-bad-tofu-apply: note that applies now run in CI with the partition
  write SA, so the bad apply is usually a failed CI run; CI does not self-heal,
  recovery is operator-run locally.

* docs(tf,talos): finish ADR purge into tf, fix sietch vault, mark talos dormant

- complete the ADR removal that stopped at ansible/ceph: drop the dangling
  ADR-009/010 references from tf/README.md (link + related line) and the ceph
  module / stack code comments, so no ADR trace remains repo-wide
- tf/README: the sietch cluster example uses yucca_tf_staging (was yucca_tf_dev)
- ansible/talos: add a "second-class, not actively used" status banner to the
  README and architecture doc so readers don't treat the converged/libvirt
  Talos docs as live

Left untouched: the Tailscale / SA-token / CI sections of tf/README -- those
are mid-migration in another owner's lane (bye-tailscale is in flight; fabric
still rides Tailscale by design).
2026-06-29 13:04:15 -07:00

12 KiB

Naming

Every name derivable in this project -- hostname, inventory directory, 1P item title, SSH key filename -- traces back to one entry per cluster in tf/deployment/<partition>/<region>/ceph/clusters.auto.tfvars. TF's ceph-cluster module assembles the rest.

For how naming fits the broader system see architecture.md section 4 (Terraform); for the per-item 1P catalog and the SSH-key lifecycle see secrets.md.

The three name layers

Three parallel naming surfaces derive from the same tfvars entry, but only the hostname carries the role_in_hostname segment. The 1P item prefix is hardcoded to CEPH in the ceph-cluster module, and the inventory directory is keyed by <partition>-<region> -- both are project-scoped, not role-scoped.

Surface Pattern Role source
Hostname (short + FQDN) <cluster>-<role>-<name>[.<domain>] role_in_hostname tfvar
Inventory directory inventories/<partition>-<region>/<cluster>/ <partition>-<region> from region.hcl
1P item prefix <CLUSTER>_CEPH_* always CEPH (module-hardcoded)

This split is deliberate. A future cluster where every node is a dedicated OSD might set role_in_hostname = "osd" (yielding hostnames like mesa-osd-willow) but its 1P items would still grep-match *_CEPH_* alongside every other Ceph-project cluster, and its inventory would still live under the same <partition>-<region>/ region tree.

Hostname segments

Component Example Source
cluster sietch top-level map key (per-cluster tfvars)
role ceph (small clusters), osd role_in_hostname (default ceph)
name laurel, evelyn hosts[].name, or TF-picked
domain staging.austin.int.futo.cloud the region's domain in region.hcl

The FQDN suffix is the region's domain, set once per region in region.hcl as <partition>.<region>.<provider_code>.futo.cloud and passed into the ceph-cluster module whole -- the cluster's tfvars entry only supplies cluster, role, and name.

Current role_in_hostname values: sietch uses ceph (mixed-role: every node runs every daemon). Dedicated-role hostnames (osd, mon) are supported but not used today.

Cluster naming

Who chooses: the engineer adding the cluster picks the name at the moment they add the entry to clusters.auto.tfvars. No automation -- it's a deliberate, one-time decision.

When: before any mise run tf:apply, any 1P item creation, any cluster bootstrap. Renaming after deployment is expensive (see "Cost of renaming" below).

Convention (unenforced): Dune-themed. The existing cluster is sietch. Unused Dune candidates that fit the hard constraints below: arrakis, caladan, giedi, ixian, kwisatz, muaddib, fremen, chani, leto, jessica. Nothing in the code enforces Dune specifically -- mixed themes or theme breaks are acceptable when they communicate intent better (e.g., a cluster named after its datacenter for a production tier).

Hard constraints:

  • lowercase alphanumeric -- becomes the HCL map key, the cluster segment of the inventory directory (<partition>-<region>/<name>/), the hostname prefix, and (uppercased) the 1P item prefix (<NAME>_CEPH_*).
  • Short -- appears in every hostname and every 1P item title. Aim for 6-10 characters; 15 is the realistic ceiling.
  • Unique within the yucca_tf_* 1P item namespace -- other Futo consumers (o11y, future stacks) write items to the same vault set. Before committing, check:
    op item list --vault yucca_tf_staging --format=json \
      | jq -r '.[] | .title' | grep -i "^<PROPOSED_NAME>_"
    
    Must return empty (repeat for the other yucca_tf_* vaults if in doubt).
  • Not already a clusters.auto.tfvars key -- TF enforces this with a plan-time error.
  • No dashes, dots, or underscores in the cluster name itself -- those are segment separators in inventory directories and hostnames. A cluster name my-cluster would produce my-cluster-ceph-laurel which parses ambiguously. Use mycluster instead.

Soft guidance:

  • Memorable -- operators will say it out loud in incidents.
  • Distinct from existing clusters' first 3 letters (grep-friendly in logs).
  • Doesn't encode partition or region -- those live in region.hcl and the domain. The cluster name is project identity, not location.

Cost of renaming after deployment

A rename touches all of:

  1. clusters.auto.tfvars map key
  2. Inventory directory name
  3. Every hostname (short + FQDN) and every SSH known_hosts entry for every operator
  4. 1P item titles (<OLD>_CEPH_* -> <NEW>_CEPH_*) including the SSH Key item
  5. cephadm cluster identity (requires cluster rebuild in the common case)
  6. ansible_ssh_key path in the tfvars (~/.ssh/id_ed25519_<cluster>) and the mapping in scripts/install-ssh-keys.sh
  7. Any DNS records and external systems that reference the hostnames

Expect hours-to-days of work, cluster downtime, and coordination with every consumer of the cluster's S3/dashboards/etc. A rename is only cheap on an idle cluster that nothing consumes yet -- a running production Ceph cluster makes this a multi-week project.

Pick once. Pick deliberately.

Host naming

Each host entry in hosts = [...] can either declare a name or omit it to let TF pick one from the wordlist. Both paths are first-class; different clusters use different paths based on operator preference.

Operator-declared

Declare the name explicitly in the tfvars. Used when the operator has a specific name in mind -- typically because it's been spoken during planning and the team already uses it.

hosts = [
  { name = "laurel", bond_ip = "10.10.10.90", bootstrap = true },
  { name = "lawson", bond_ip = "10.10.10.91" },
  { name = "samara", bond_ip = "10.10.10.92" },
]

Sietch uses this path.

Names must be unique within a cluster, not globally. A future mesa-ceph-laurel can coexist with sietch-ceph-laurel -- the FQDN disambiguates.

Auto-picked from wordlist

Omit name (leave the field absent) and the module picks from tf/shared/modules/ceph-cluster/wordlist.txt (923 words) via random_shuffle, seeded by cluster_name + name_seed. Picks are stable across subsequent applies.

hosts = [
  { bond_ip = "157.180.105.198", bootstrap = true },  # TF picks
]

A single-host cluster commonly uses this path -- hosts[0] has no name, so TF picks a stable word (e.g. evelyn) on first apply, yielding <cluster>-ceph-evelyn.

Operator-declared names are excluded from the available pool to prevent collisions within the cluster.

Stability rules (either path, or mixed)

  • Add new hosts at the tail of the list. Auto-picked names are positional -- hosts[0] gets random_shuffle.result[0], hosts[1] gets result[1], and so on. Inserting a new entry at position 0 would shift every subsequent host's result-index. Always append.
  • Never bump name_seed once a cluster has deployed hosts. Bumping re-rolls every auto-picked name in the cluster, which cascades into certs, SSH known_hosts, 1P items, DNS, cephadm identity.
  • Mixing paths has a hidden side effect. Auto-picked names are drawn from available_words = wordlist - explicit_names. Adding or removing an operator-declared host changes explicit_names, which changes the shuffle input length, which re-permutes the result. An auto-picked host at hosts[2] could get a different name even if nothing about its own entry changed. The safe patterns:
    1. All-operator-declared within a cluster (sietch's model), or
    2. All-auto-picked within a cluster. Mixed works for initial setup but complicates later add/remove.
  • Converting between paths after deploy (e.g., adding name = "evelyn" to a host that previously auto-picked evelyn) does not preserve the name despite appearing to match -- it rewrites available_words and re-rolls every other auto-pick. Only do this if you're prepared to pin every auto-named host in the same apply, or accept the cluster-wide rename.

Inventory directory naming

TF renders inventory directories as:

inventories/<partition>-<region>/<cluster>/

The <partition>-<region> slug (e.g. staging-austin) groups every cluster in a region under one tree, regardless of role_in_hostname -- a hypothetical cluster with role_in_hostname = "osd" (hostnames mesa-osd-*) still renders under prod-htz-fsn1/mesa/.

Built in tf/shared/modules/ceph-cluster/rendering.tf (local.inventory_dirname), surfaced via the module's inventory_dirname output and consumed by the deployment stack + scripts/render-inventories.sh.

1Password item naming

Items are titled <CLUSTER>_CEPH_<specifier> (SHOUTY_SNAKE_CASE). The <CLUSTER>_CEPH_* prefix is hardcoded in tf/shared/modules/ceph-cluster/main.tf (local.secret_prefix) -- same project-scoping rationale as the inventory directory.

Per cluster, the expected item set:

Category Title Source of values
Password <CLUSTER>_CEPH_OPS_PASSWORD op item create --generate-password at setup
Password <CLUSTER>_CEPH_DASHBOARD_PASSWORD same
Password <CLUSTER>_CEPH_GRAFANA_PASSWORD same
Password <CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY same
Password <CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY same
SSH Key <CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY op item create --category "SSH Key" --ssh-generate-key=ed25519
Document <CLUSTER>_CEPH_RGW_TLS_CERT populated by mise run capture post-deploy
Document <CLUSTER>_CEPH_RGW_TLS_KEY same
Document <CLUSTER>_CEPH_CLIENT_ADMIN_KEYRING same

Passwords + SSH Key are created at cluster-add time (step 6 of adding-a-cluster.md). DR Documents are upserted automatically on the first mise run capture after deploy (step 10).

For the full consumption flow (which Ansible variable each item maps to, which role reads it) see secrets.md.

Workstation SSH key filenames

The operator-side private key path is derived from the cluster name:

~/.ssh/id_ed25519_<cluster>

Example: ~/.ssh/id_ed25519_sietch. The keypair lives in 1Password as <CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY (a native SSH Key item, same storage model as the password items) and is installed onto the workstation via scripts/install-ssh-keys.sh <cluster>.

The ansible_ssh_key field in the cluster's tfvars entry must match this path. If you choose a non-default filename (unusual), update both together and also adjust the mapping in scripts/install-ssh-keys.sh.