Files
yucca/ansible/ceph/docs/architecture.md
T
Andy Molenda 8821fad205 docs(ceph): realign with partition/region model and CI/CD, retire ADRs
* docs(ceph): inline ADR rationale and drop the ADR set

Fold each linked ADR's rationale into the prose it supported, then remove
the ADR files, the README index row, and the stray code-comment reference --
no ADR trace remains.

True up the docs to the partition/region/ceph-cluster layout (#222) in the
same pass: state keys (yucca/<partition>/<region>/<stack>), inventory paths
(<partition>-<region>/<cluster>), stack dirs, and the recover runbook's S3
state paths. Reframe architecture's env section around partitions/regions
with sietch as staging/austin.

* docs(ceph): editorial pass to align docs with current code and CI/CD

Rewrite the ceph docs against the actual code rather than the pre-refactor
state:

- partition/region/ceph-cluster layout throughout: state keys
  yucca/<partition>/<region>/<stack>, inventories <partition>-<region>/<cluster>,
  the real clusters.auto.tfvars schema (partition/region, not environment/datacenter)
- sietch reframed as staging/austin with secrets in yucca_tf_staging; vault
  hierarchy flipped from dev-primary to staging-primary
- live CI/CD (.github/workflows/infra.yml): per-partition read/write service
  accounts delivered as GitHub secrets (OP_TF_YUCCA_<ENV>_ENV[_WRITE]),
  plan/apply gating, NetBird overlay (Tailscale retired)
- correct the CEPH_ENV guidance (export works; deliberately kept out of mise
  [env]) across README, CONTRIBUTING, scripts.md, adding-a-cluster
- drop the obsolete "read SA token from a 1P item" dance from the runbooks
- fix stale vaults, paths, examples, and the inventory-provision.ini name

* docs(ceph): transliterate docs to plain ASCII

Replace non-ASCII punctuation and box-drawing with ASCII equivalents across
the ceph docs: em/en dashes to --/-, middot separators to commas, arrows to
->, directory-tree box-drawing to |-- / `--, section sign to "section", x for
the multiply glyph, and >= / <= / ~ for the math glyphs. No content changes.

* docs(ceph): fix broken rotate-ssh-key link in scripts.md

The "Related" link pointed at runbooks/rotate-ssh-key.md, which does not
exist (a pre-existing dangling link). SSH-key rotation lives in
rotate-secrets.md; point at its "Rotating SSH keys" section.

* docs(ceph): style polish from per-doc review

Tighten verbal texture flagged by a per-doc style pass; no structural changes.

- correctness: ansible-play.sh runs ansible-playbook, it does not exec, so the
  trap fires from the still-alive wrapper -- fix the "exec"/SIGKILL claims in
  scripts.md, architecture.md, secrets.md
- cut recurring tics: "DR belt"/"belt-and-suspenders" -> "disaster recovery",
  "a lost laptop is a non-event", "by design", "system mesh"
- unstuff long dash/semicolon sentences in architecture (vault-password
  history, provision/baseline split), secrets (SSH-key paragraph), patterns
- recover-bad-tofu-apply: move the dormant-1P-items aside into one note,
  consolidate the repeated caveats
- misc: naming grammar fix + drop trivia, hardware "would"/"blindly",
  rotate-secrets "after confidence", drop a dead snippet line, fix the 16+
  numeric hedge

* docs(ceph): make add-node and recover runbooks CI-aware

Now that infra.yml applies the stacks and runs the full ceph convergence on
merge, refresh the two runbooks the pipeline changed:

- add-node: lead with the manual-vs-CI split. The TF + host_vars change is a
  PR; the only operator-only step is the physical provisioning (live-image
  boot + provision.yml), which CI can't do; baseline/tune/join/harden run in
  CI on merge. Keep the by-hand convergence as a documented fallback.
- recover-bad-tofu-apply: note that applies now run in CI with the partition
  write SA, so the bad apply is usually a failed CI run; CI does not self-heal,
  recovery is operator-run locally.

* docs(tf,talos): finish ADR purge into tf, fix sietch vault, mark talos dormant

- complete the ADR removal that stopped at ansible/ceph: drop the dangling
  ADR-009/010 references from tf/README.md (link + related line) and the ceph
  module / stack code comments, so no ADR trace remains repo-wide
- tf/README: the sietch cluster example uses yucca_tf_staging (was yucca_tf_dev)
- ansible/talos: add a "second-class, not actively used" status banner to the
  README and architecture doc so readers don't treat the converged/libvirt
  Talos docs as live

Left untouched: the Tailscale / SA-token / CI sections of tf/README -- those
are mid-migration in another owner's lane (bye-tailscale is in flight; fabric
still rides Tailscale by design).
2026-06-29 13:04:15 -07:00

40 KiB

Architecture

How the Ceph automation in yucca/ansible/ceph/ is shaped, what each tool owns, and how the four tools (Terraform, 1Password, Ansible, mise) -- plus the op CLI that resolves secrets between them -- hand off work to each other.

This is a structural reference. For step-by-step usage see CONTRIBUTING.md; for narrower topics see the specialized docs under docs/.


1. System context

flowchart LR
    OP([Operator workstation<br/>mise, op CLI, ansible, tofu])
    YUCCA[/Yucca monorepo<br/>tf/ + ansible/ceph/ + kubernetes//]
    ONEP[("1Password org<br/>yucca_tf, yucca_tf_staging, ...")]
    S3[("OVH S3<br/>yucca-tf-state bucket")]
    SIETCH["Sietch, Austin DC<br/>3x Dell R730xd"]

    OP -->|edits| YUCCA
    OP -->|reads/writes secrets| ONEP
    OP -->|TF state I/O| S3
    OP -->|SSH ansible-iac| SIETCH

The yucca monorepo is the single source of truth for cluster identity and configuration. Operators run mise tasks on their workstation; secrets stay in 1Password (never on disk); TF state lives in OVH S3; Ansible drives configuration over SSH against bare-metal Ceph nodes.

External dependencies are minimal and explicit:

  • 1Password org -- organization-scoped vaults shared with other Futo infra (Immich, o11y). Authoritative store for live secret values.
  • OVH S3 -- yucca-tf-state bucket at s3.eu-west-par.io.cloud.ovh.net. Keyed by yucca/<partition>/<region>/<stack>/terraform.tfstate so multiple stacks share the bucket without collision.
  • Hardware -- Austin colo for sietch (Dell R730xd x 3, single 10G bond). Detail in hardware.md.

2. Partitions and regions

Partition (dev / staging / prod) and region are first-class concerns: every tool in the mesh derives both from the same source -- directory layout (tf/deployment/<partition>/<region>/<stack>) -- so isolation is structural, not flag-driven.

Layer staging / austin (today) dev / local (planned) prod / htz-fsn1 (planned)
TF stack dir tf/deployment/staging/austin/ceph/ tf/deployment/dev/local/ceph/ tf/deployment/prod/htz-fsn1/ceph/
TF state key yucca/staging/austin/ceph/terraform.tfstate yucca/dev/local/ceph/terraform.tfstate yucca/prod/htz-fsn1/ceph/terraform.tfstate
1P vaults yucca_tf_staging, yucca_tf_staging_manual yucca_tf_dev, yucca_tf_dev_manual yucca_tf (live), yucca_tf_prod_manual
Ansible inv inventories/staging-austin/<cluster>/ inventories/dev-local/<cluster>/ inventories/prod-htz-fsn1/<cluster>/
mise default CEPH_ENV=...staging-austin/sietch/inventory.ini overridden via env at invocation overridden via env at invocation

Today the only deployed cluster is sietch (staging / austin). Adding another region or partition is purely additive: create the matching tf/deployment/<partition>/<region>/ceph/ directory, populate clusters.auto.tfvars, and the same module + Ansible roles + mise tasks work unchanged. The state backend key path, 1P vault selection, and inventory directory naming all derive from the partition + region segments.

TF_STACK_DIR is the operator-side override for mise run tf:* tasks; it defaults to tf/deployment/staging/austin/ceph and points at any sibling stack directory. CEPH_ENV is the matching override for Ansible -- points at the rendered inventory.ini for the cluster you intend to operate on.


3. The tool mesh

flowchart TB
    subgraph ws["Operator workstation"]
        direction LR
        MISE([mise<br/>orchestration])
        TF[Terraform / Tofu<br/>via Terragrunt]
        OP[op CLI]
        ANS[Ansible]
        WRAP[scripts/<br/>ansible-play.sh<br/>install-ssh-keys.sh]
    end

    ONEP[("1Password<br/>yucca_tf_*")]
    S3[("OVH S3<br/>tfstate")]
    REPO[/"Yucca repo<br/>inventories/&lt;partition&gt;-&lt;region&gt;/&lt;cluster&gt;/<br/>(host_vars committed,<br/>TF outputs gitignored)"/]
    NODES[Ceph nodes]

    MISE -->|tf:*| TF
    MISE -->|deploy / status / drift| WRAP
    MISE -->|capture| WRAP

    TF -->|reads SA token via op run --env-file| OP
    TF -->|reads/writes state| S3
    TF -->|renders| REPO

    WRAP -->|reads inventory + secrets.yml.tpl| REPO
    WRAP -->|op inject / op read| OP
    WRAP -->|runs| ANS
    ANS -->|SSH ansible-iac| NODES

    OP <-->|item CRUD| ONEP

Who owns what

Tool Owns Reads from
Terraform Cluster identity, host names, rendered Ansible artifacts, TF state clusters.auto.tfvars, 1P (via op CLI)
1Password Live secret values, SSH keypairs, service-account tokens nothing -- authoritative store
op CLI Auth and resolution: env-injection, file-template injection, single-value read 1P (session or SA token)
Ansible Convergence: applying configuration to nodes Rendered inventory + op-injected tmpfile
mise Task discovery, toolchain pinning, env defaults .mise.toml, tf/.env

Handoff points (the edges of the mesh)

  1. TF -> repo -- terragrunt apply renders inventory.ini, inventory-destroy.ini, optional inventory-provision.ini, and secrets.yml.tpl into inventories/<partition>-<region>/<cluster>/. These files are gitignored -- the source of truth is clusters.auto.tfvars + the ceph-cluster module.
  2. TF <-> op CLI -- TF runs are wrapped with op run --env-file=tf/.env, which resolves op://... references in tf/.env and injects them as OP_SERVICE_ACCOUNT_TOKEN, AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY for the child process. The tf/.env file is committed (it contains only pointers, never literal secrets).
  3. Ansible <-> op CLI -- scripts/ansible-play.sh reads the cluster's secrets.yml.tpl, runs op inject -f to resolve op:// references into a mktemp'd tmpfile (chmod 600, trap-cleaned), then runs ansible-playbook --extra-vars @<tmpfile>. The tmpfile lives only for the duration of the play.
  4. mise -> wrappers -- mise run deploy invokes scripts/ansible-play.sh deploy.yml ...; mise run tf:* invokes tf/op-run.sh terragrunt --working-dir <stack> <cmd> (op-run.sh is a thin op run --env-file=tf/.env -- wrapper). mise tasks never call ansible-playbook directly.

4. Terraform (authority)

Layout

tf/
|-- .env                              op:// references (committed; no literal secrets)
|-- shared/modules/ceph-cluster/      per-cluster orchestration module
|   |-- main.tf, variables.tf, outputs.tf, rendering.tf
|   |-- wordlist.txt                  923 words for auto-picked hostnames
|   `-- templates/                    inventory + secrets.yml.tpl templates
`-- deployment/
    |-- terragrunt.hcl                root: state backend, partition/region/stack derived from path
    `-- staging/austin/ceph/
        |-- terragrunt.hcl            includes root, sets ansible_project_root
        |-- main.tf, variables.tf, versions.tf
        `-- clusters.auto.tfvars      declarative cluster list

Cluster identity is declared, not derived

The clusters map in clusters.auto.tfvars is the source of truth. Each top-level key becomes a cluster:

sietch = {
  domain            = "staging.austin.int.futo.cloud"
  partition         = "staging"
  region            = "austin"
  provider_code     = "int"
  role_in_hostname  = "ceph"
  ansible_ssh_user  = "ansible-iac"
  ansible_ssh_key   = "~/.ssh/id_ed25519_sietch"
  vault             = "yucca_tf_staging"
  provision_profile = "debian-live"
  hosts = [
    { name = "laurel", bond_ip = "10.10.10.90", bootstrap = true },
    { name = "lawson", bond_ip = "10.10.10.91" },
    { name = "samara", bond_ip = "10.10.10.92" },
  ]
}

The module computes everything else: hostname (<cluster>-<role>-<name>), FQDN (<hostname>.<domain>), 1P item names (<CLUSTER>_CEPH_<ROLE>_PASSWORD), inventory directory path (inventories/<partition>-<region>/<cluster>/).

Auto-naming via wordlist

For hosts where name = null, the module picks a stable name from a 923-word pool using random_shuffle seeded by (cluster_name, name_seed). Operator-declared names are excluded from the pool to prevent collisions within a cluster. Adding hosts at the tail is safe -- existing positions keep their names across applies.

A host declared with no name demonstrates this: TF auto-picks a stable word (e.g. evelyn) -> hostname <cluster>-ceph-evelyn.

Rendered artifacts (gitignored)

Per tf/shared/modules/ceph-cluster/rendering.tf, the module writes four files into inventories/<partition>-<region>/<cluster>/:

File Purpose
inventory.ini Normal-ops inventory: ansible-iac user + cluster SSH key
inventory-destroy.ini Destroy-mode inventory (same credentials; separate file as a speed bump)
inventory-provision.ini Provisioning inventory (only when provision_profile != null; uses live-image creds; the profile names the template, not the output)
secrets.yml.tpl vault_*: op://<vault>/<CLUSTER>_CEPH_*/password pointers, consumed by op inject -f

All four are in ansible/ceph/.gitignore. Re-render with mise run tf:apply.

State backend

S3 backend in tf/deployment/terragrunt.hcl:

  • Bucket: yucca-tf-state (shared with o11y and other Futo stacks)
  • Region: eu-west-par (OVH Paris)
  • Endpoint: https://s3.eu-west-par.io.cloud.ovh.net/
  • Key: yucca/${partition}/${region}/${stack}/terraform.tfstate -- derived from the child stack's path under deployment/
  • Skip AWS-specific validation; use path-style URLs (OVH compatibility)

State locking is not enabled today. OVH has no DynamoDB equivalent. OpenTofu's use_lockfile = true would work but expects the lockfile object to already exist -- fresh-backend init fails with 404 before it can create one. Single-operator workflow today; revisit when concurrent applies become likely. See deployment/terragrunt.hcl for the inline rationale.

What TF does not yet manage

onepassword_item resources are dormant (tf/.../secrets.tf.disabled). 1P items are created today via the op CLI (operator runs op item create once per cluster). The gate to re-enabling them is the dedicated sietch-ceph service account that lets us split write authority from the org-wide superuser SA.


5. 1Password (live values)

Vault hierarchy

Vault Purpose Who reads it Who writes it
yucca_tf Cross-partition shared (TF state S3 creds) TF (via tf/op-run.sh) Operator (manual)
yucca_tf_staging staging live values (sietch today) Ansible runtime (op inject) Superuser SA (TF) + operator (op CLI)
yucca_tf_staging_manual staging human-fillable placeholders (API tokens, OAuth) Ansible runtime Operator (manual)
yucca_tf_dev(_manual), yucca_tf, yucca_tf_prod_manual dev + prod analogues, same shape per partition per partition

Each partition has its own live + _manual vault pair; sietch runs in staging, so its items live in yucca_tf_staging. The _manual vaults exist for items that can't be auto-generated (third-party API tokens, OAuth client secrets) -- they're populated by humans, not by TF.

Service accounts

Each partition has a read and a write 1Password service account, scoped to that partition's vaults. CI consumes them as GitHub repo secrets, injected as OP_SERVICE_ACCOUNT_TOKEN per job -- the read token for plan, the write token for apply:

Partition Read SA secret Write SA secret
staging OP_TF_YUCCA_STAGING_ENV OP_TF_YUCCA_STAGING_ENV_WRITE
prod OP_TF_YUCCA_PROD_ENV OP_TF_YUCCA_PROD_ENV_WRITE
dev local-only -- no CI service account --

Locally, operators authenticate with their own 1Password desktop session (Futo membership) rather than a service-account token. The split -- read for plan, write for apply -- keeps drift-detection runs from holding write authority.

Rotation procedure: docs/runbooks/rotate-sa-token.md.

Item categories and naming

Per cluster, the following items live in the cluster's vault (currently yucca_tf_staging for sietch):

Category Item title pattern Field consumed
Password <CLUSTER>_CEPH_OPS_PASSWORD password
Password <CLUSTER>_CEPH_DASHBOARD_PASSWORD password
Password <CLUSTER>_CEPH_GRAFANA_PASSWORD password
Password <CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY password
Password <CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY password
SSH Key <CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY private_key / public_key
Document <CLUSTER>_CEPH_RGW_TLS_CERT, ..._RGW_TLS_KEY file content
Document <CLUSTER>_CEPH_CLIENT_ADMIN_KEYRING file content

The <CLUSTER>_CEPH_* prefix is hardcoded in tf/shared/modules/ceph-cluster/main.tf (secret_prefix = "${upper(var.cluster_name)}_CEPH") so every Ceph-project item across all clusters is grep-discoverable as *_CEPH_* regardless of the role segment in node hostnames.

Full item-by-item catalog: docs/secrets.md.

Three op-CLI patterns

The op CLI is invoked in three distinct ways across the codebase. Each serves a different shape of secret consumption:

  1. op run --env-file=tf/.env -- <cmd> -- env-var injection. Resolves op:// references in a dotenv file and injects the resolved values as env vars into the child process. Used for TF (SA token) and the S3 backend (AWS creds). Wrapped by all mise run tf:* tasks.
  2. op inject -f -i <tpl> -o <out> -- file-template resolution. Reads a file containing inline op:// references, resolves each, writes to the output path. Used by scripts/ansible-play.sh to render secrets.yml.tpl -> tmpfile, and by the Hetzner installimage flow to render post-install.sh.tpl -> post-install.sh.
  3. op read "op://<vault>/<item>/<field>" -- single-value read. Used by scripts/install-ssh-keys.sh, rotate-ssh-key.yml, post-deploy-capture.yml. Returns one value to stdout for one specific field; fails closed if missing.

No custom password-script (no vault-password.sh); no ansible-vault-encrypted file in git. Lint and syntax-check tasks don't invoke op at all -- they don't need secrets, so "1P unavailable" never silently degrades them. This replaced an earlier vault-password.sh + ansible-vault setup. That setup fell back to a dummy password when 1P was unavailable, which masked real auth failures until a downstream task blew up. The current flow fails closed instead.


6. Ansible (consumer)

Role dependency graph

Provisioning is a separate concern (provision.yml, sietch only). The main pipeline (site.yml) runs everything else in this order:

flowchart TB
    PROV["provision_host<br/><i>separate playbook, live image</i>"]
    BASE["baseline<br/><i>users, packages, /etc/hosts</i>"]
    OST["os_tuning<br/><i>sysctl, TCP buffers</i>"]
    HWT["hardware_tuning<br/><i>I/O scheduler, readahead</i>"]
    DEPLOY["ceph_deploy<br/><i>bootstrap, join, OSDs, RGW,<br/>crush rules, monitoring</i>"]
    CTUNE["ceph_tuning<br/><i>recovery throttling, scrub<br/>window, telemetry, audit</i>"]
    SEC["security<br/><i>nftables, SSH hardening</i>"]

    PROV -.->|reboot into installed OS| BASE
    BASE --> OST
    BASE --> HWT
    OST --> DEPLOY
    HWT --> DEPLOY
    DEPLOY --> CTUNE
    CTUNE --> SEC

    classDef separate stroke-dasharray: 4 4
    class PROV separate

site.yml starts at baseline -- provision_host runs only on first install via provision.yml. (The roles above are imported as per-role playbooks: baseline.yml, tune-os.yml, tune-hardware.yml, deploy-ceph.yml, tune-ceph.yml, harden.yml.)

The split is deliberate. provision_host does the minimum inside the live-image chroot -- just the ansible-iac user, so Ansible can connect after reboot -- because chroot work is fragile. The ops user, packages, and /etc/hosts move to the convergeable baseline role, which re-runs against a live node to fix drift without reprovisioning.

The OS is installed with debootstrap from the live image rather than a preseed/autoinstall. The disk layout (mdraid-1 across two SSDs, partitions reserved for ceph block.db and SSD OSDs) needs scripted partitioning and pre-flight hardware validation that preseed's partman recipes can't express.

Why this order matters

  1. baseline before tuning -- cephadm needs podman, dbus, chrony. The baseline role installs these and enables the services. Running tuning on a node without podman would leave cephadm unable to bootstrap.
  2. tuning before deploy -- OSD daemons inherit kernel parameters active at startup. Applying sysctl (vm.min_free_kbytes, fs.aio-max-nr) and I/O scheduler (mq-deadline for HDD, none for SSD) before bootstrap means daemons launch with correct limits from the first second.
  3. ceph_tuning after deploy -- these settings use ceph config set which requires a running cluster. Recovery throttling, scrub windows, and PG autoscaler targets cannot be applied until MONs are up.
  4. security last -- nftables drops all traffic not explicitly allowed. Running it before ceph_deploy would block cephadm's inter-node SSH, container image pulls, and MON/OSD port negotiation. Once the cluster is healthy, the firewall locks it down.

ceph_deploy internal pipeline

roles/ceph_deploy/tasks/main.yml orchestrates ten phases:

flowchart TB
    P1["Phase 1, prerequisites.yml<br/><i>Ceph repo, cephadm, ceph-common</i>"]
    P2["Phase 2, bootstrap.yml<br/><i>cephadm bootstrap on first node</i>"]
    P3["Phase 3, join.yml<br/><i>ceph orch host add for remaining nodes</i>"]
    P4["Phase 4, placement.yml<br/><i>MON/MGR placement calculation</i>"]
    P45["Phase 4.5, lvm-setup.yml<br/><i>ensure block.db VGs/LVs exist (sietch-shape only;<br/>NVMe-RAID shape skips -- LVM owned by installimage post-install)</i>"]
    P5["Phase 5, osds.yml<br/><i>render osd-spec.yml.j2 -> ceph orch apply osd<br/>(cephadm provisions LUKS + LVM internally)</i>"]
    P55["Phase 5.5, crush-rules.yml<br/><i>replicated_hdd / replicated_ssd rules</i>"]
    P575["Phase 5.75, rgw.yml<br/><i>EC pools, realm/zone, TLS, S3 user</i>"]
    P58["Phase 5.8, monitoring.yml<br/><i>dashboard URL integration, Grafana creds</i>"]
    P6["Phase 6, verify.yml<br/><i>cluster health report</i>"]

    P1 --> P2 --> P3 --> P4 --> P45 --> P5 --> P55 --> P575 --> P58 --> P6

Tag-driven re-runs are first-class: scripts/ansible-play.sh deploy-ceph.yml --tags rgw,monitoring re-runs just those phases.

Inventory layout

inventories/
  staging-austin/sietch/    Austin staging cluster
    inventory.ini                     TF-generated, gitignored
    inventory-destroy.ini             TF-generated, gitignored
    inventory-provision.ini           TF-generated, gitignored
    secrets.yml.tpl                   TF-generated, gitignored
    group_vars/all/vars.yml           cluster-wide variables (committed)
    host_vars/                        per-node hardware topology (committed)
      sietch-ceph-laurel.yml          bond_ip, SAS path prefix, OSD maps
      sietch-ceph-lawson.yml
      sietch-ceph-samara.yml
    installimage/                     Hetzner installimage assets (sietch n/a)

A Hetzner NVMe-RAID cluster would follow the same layout, adding an installimage/ directory with a post-install.sh.tpl (op-injected) that owns LVM setup. No such cluster is deployed today -- sietch is the only live cluster -- but the module and roles already support the shape.

host_vars/*.yml is committed because per-node hardware facts (bond_ip, SAS expander paths, SSD PHY positions, HDD-to-block.db mappings) are stable inventory truth -- not operator preference. The .local.yml suffix is gitignored as an escape hatch for operator-local overrides.

Variable precedence

flowchart TB
    D["<b>role defaults</b><br/>roles/*/defaults/main.yml<br/><i>lowest priority</i>"]
    G["<b>group_vars</b><br/>inventories/&lt;cluster&gt;/group_vars/all/vars.yml"]
    H["<b>host_vars</b><br/>inventories/&lt;cluster&gt;/host_vars/&lt;host&gt;.yml"]
    T["<b>extra-vars @tmpfile</b><br/>scripts/ansible-play.sh<br/><i>op-injected secrets</i>"]
    E["<b>extra-vars -e X=Y</b><br/>-e confirm_wipe=true<br/><i>highest priority</i>"]

    D --> G --> H --> T --> E
  • Role defaults define every tunable with a safe value (ceph_firewall_ssh_any_source: true, ceph_cpu_governor_enabled: false).
  • group_vars/all/vars.yml sets cluster-wide values: network topology, Ceph release, RGW config, monitoring ports, plus the vault_* -> consumable-name aliases (ops_password: "{{ vault_ops_password }}").
  • host_vars provides per-node physical topology.
  • extra-vars from @tmpfile carries op-injected vault_ops_password, vault_ceph_dashboard_password, vault_grafana_admin_password, vault_s3_restic_access_key, vault_s3_restic_secret_key.
  • extra-vars via -e carries safety gates: confirm_wipe=true, provision_skip_reboot=true, yes_destroy_ceph=true.

ansible.cfg stays generic

ansible.cfg contains zero site-specific values. No default inventory, no ProxyJump, no hardcoded key paths. Site-specifics live exclusively in clusters.auto.tfvars (which TF renders into the inventory) or in the inventory's group_vars. The same ansible.cfg and the same roles work unchanged across Austin, Hetzner, or any future cluster -- only the cluster entry in clusters.auto.tfvars differs.


7. mise (orchestration surface)

Why mise

  • Toolchain pinning -- .mise.toml declares the exact versions of python, tofu, terragrunt, op. New operators get a working environment with mise trust && mise run setup.
  • Task discovery -- mise tasks lists every operation; tasks are shell-script-shaped, kept in .mise.toml, and committed.
  • Devtools parity -- matches the conventions in immich-app/devtools (where the op run --env-file=tf/.env -- pattern originated).

Task taxonomy

Group Tasks
Bootstrap setup
Verify lint, check, test, preflight
Read-only ops status, drift
State change deploy, destroy, capture, backup
Rotation rotate-certs, rotate-ssh-key
Inventory hardware-inventory, migrate-networkd
Benchmarks bench, bench-rados

The ceph ops tasks above live in ansible/ceph/.mise.toml. The tf:* tasks (tf:init, tf:plan, tf:apply, tf:destroy, tf:fmt) live in the yucca-root .mise/config.toml and run from the repo root -- they wrap terragrunt for any stack, not just ceph.

How mise wraps the underlying CLIs

  • mise run tf:* -> tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR} <cmd> (op-run.sh = op run --env-file=tf/.env --)
  • mise run deploy -> scripts/ansible-play.sh deploy-ceph.yml ... (per phase)
  • mise run status -> scripts/ansible-play.sh status.yml
  • mise run capture -> scripts/ansible-play.sh post-deploy-capture.yml

mise never invokes ansible-playbook or terragrunt directly. The wrappers own secrets injection and pre-flight checks; mise owns task discovery and env defaults.

Env defaults

CEPH_ENV is deliberately not declared in [env] -- mise's [env] block overrides shell-exported values, which would silently send an operator to the wrong cluster. Instead each ceph ops task falls back to sietch only when CEPH_ENV is unset:

# the default baked into each task
CEPH_ENV="${CEPH_ENV:-inventories/staging-austin/sietch/inventory.ini}"

# operate on another cluster by exporting once per shell, or inline:
export CEPH_ENV=inventories/staging-austin/sietch/inventory.ini
CEPH_ENV=inventories/<partition>-<region>/<cluster>/inventory.ini mise run status

TF_STACK_DIR works the same way for the root tf:* tasks -- it defaults to tf/deployment/staging/austin/ceph and is overridden per-invocation:

TF_STACK_DIR=tf/deployment/<partition>/<region>/ceph mise run tf:plan

8. Wrapper scripts (the glue layer)

Three scripts under ansible/ceph/scripts/ sit between mise and the underlying CLIs. They exist to keep secrets out of argv, fail closed when 1P is unreachable, and give better error messages than the raw tools.

Script Purpose
ansible-play.sh Render secrets via op inject -f to a mktemp'd file (chmod 600, trap-cleaned), then run ansible-playbook --extra-vars @<tmpfile>
install-ssh-keys.sh Idempotent op read -> ~/.ssh/id_ed25519_<cluster> installer; refuses overwrite on fingerprint mismatch
preflight.sh Verifies TF artifacts present, 1P session live, SSH reachable, Python on targets -- surfaced via mise run preflight

Per-script reference (synopsis, args, env, exit codes, examples): docs/scripts.md.


9. Data flow: concrete operations

9.1 mise run tf:apply -- render artifacts

sequenceDiagram
    actor OP as Operator
    participant MISE as mise
    participant OPCLI as op CLI
    participant ONEP as 1Password
    participant TG as terragrunt / tofu
    participant S3 as OVH S3
    participant REPO as Repo (inventories/)

    OP->>MISE: mise run tf:apply
    MISE->>OPCLI: tf/op-run.sh -- ...<br/>(op run --env-file=tf/.env)
    OPCLI->>ONEP: resolve op:// references
    ONEP-->>OPCLI: SA token + AWS keys
    OPCLI->>TG: exec child process<br/>with env vars injected
    TG->>S3: read tfstate<br/>(yucca/<partition>/<region>/<stack>/terraform.tfstate)
    S3-->>TG: current state
    TG->>TG: plan + apply
    TG->>S3: write updated tfstate
    TG->>REPO: render inventory.ini,<br/>secrets.yml.tpl, ...

9.2 mise run deploy -- full Ceph deploy

sequenceDiagram
    actor OP as Operator
    participant MISE as mise
    participant WRAP as ansible-play.sh
    participant OPCLI as op CLI
    participant ONEP as 1Password
    participant TMP as /tmp/<id>-secrets.yml
    participant ANS as ansible-playbook
    participant NODES as Ceph nodes

    OP->>MISE: mise run deploy
    MISE->>WRAP: ansible-play.sh deploy-ceph.yml
    WRAP->>OPCLI: op account get
    OPCLI-->>WRAP: session OK
    WRAP->>TMP: mktemp + chmod 600 + trap rm
    WRAP->>OPCLI: op inject -f -i secrets.yml.tpl -o TMP
    OPCLI->>ONEP: resolve op://yucca_tf_staging/SIETCH_CEPH_*/password
    ONEP-->>OPCLI: secret values
    OPCLI->>TMP: write resolved YAML
    WRAP->>ANS: run --extra-vars @TMP
    loop phases 1..6
        ANS->>NODES: SSH ansible-iac@<bond_ip><br/>via id_ed25519_<cluster>
    end
    Note over WRAP,TMP: tmpfile rm'd on exit (trap)

9.3 mise run capture -- DR snapshot

sequenceDiagram
    actor OP as Operator
    participant MISE as mise
    participant ANS as ansible-playbook<br/>(post-deploy-capture.yml)
    participant BOOT as Bootstrap node
    participant LOCAL as localhost (delegated)
    participant OPCLI as op CLI
    participant ONEP as 1Password

    OP->>MISE: mise run capture
    MISE->>ANS: ansible-play.sh post-deploy-capture.yml
    ANS->>BOOT: SSH read /etc/ceph/rgw-ssl.crt
    ANS->>BOOT: SSH read /etc/ceph/rgw-ssl.key
    ANS->>BOOT: SSH read /etc/ceph/ceph.client.admin.keyring
    BOOT-->>ANS: file contents
    loop for each artifact
        ANS->>LOCAL: delegate_to: localhost
        LOCAL->>OPCLI: op item edit/create<br/><CLUSTER>_CEPH_<ITEM>
        OPCLI->>ONEP: upsert Document item<br/>in yucca_tf_staging
    end
    Note over ONEP: Now holds RGW_TLS_CERT,<br/>RGW_TLS_KEY, CLIENT_ADMIN_KEYRING

9.4 scripts/install-ssh-keys.sh -- fresh workstation

sequenceDiagram
    actor OP as Operator (new ws)
    participant SCRIPT as install-ssh-keys.sh
    participant OPCLI as op CLI
    participant ONEP as 1Password
    participant SSH as ~/.ssh/

    OP->>SCRIPT: install-ssh-keys.sh sietch
    SCRIPT->>OPCLI: op read .../public_key
    OPCLI->>ONEP: SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY
    ONEP-->>OPCLI: public_key
    OPCLI-->>SCRIPT: pubkey content
    SCRIPT->>SSH: compare with id_ed25519_sietch (if exists)
    alt fingerprint match
        SCRIPT-->>OP: skip (already present)
    else fingerprint mismatch
        SCRIPT-->>OP: refuse (operator must mv aside)
    else file missing
        SCRIPT->>OPCLI: op read .../private_key
        OPCLI-->>SCRIPT: private_key
        SCRIPT->>SSH: write id_ed25519_sietch (0600)<br/>+ .pub (0644)<br/>(umask 077)
    end

10. OSD lifecycle

Phase flow

flowchart TB
    SIETCH["Sietch prep:<br/>provision_host/disks.yml partitions SSDs<br/>then ceph_deploy/lvm-setup.yml<br/><i>creates VG + db-slot LVs on each SSD's partition 5</i>"]
    NVMERAID["NVMe-RAID prep:<br/>installimage post-install.sh<br/><i>NVMe RAID-1 -> vg0 -> db-slot0..13 + ssd-osd LVs</i>"]
    SPEC["ceph_deploy/osds.yml renders<br/>templates/osd-spec.yml.j2 -> /etc/ceph/osd-spec.yml<br/><i>one document per host; paths from host_vars</i>"]
    APPLY["ceph orch apply osd -i /etc/ceph/osd-spec.yml<br/><i>cephadm: discover disks, LUKS-format, LVM, deploy daemons</i>"]
    POLL["Wait for cephadm to provision<br/><i>poll num_osds until expected count reached</i>"]
    UP["Wait for OSDs up<br/><i>poll num_up_osds == num_osds</i>"]
    UNSET["Defensive: ceph osd unset noin<br/><i>idempotent -- clears stale flag from prior runs</i>"]
    REWEIGHT["Safety net: fix any reweight=0 OSDs"]

    SIETCH --> SPEC
    NVMERAID --> SPEC
    SPEC --> APPLY --> POLL --> UP --> UNSET --> REWEIGHT

Service-spec model, not per-disk loops

Earlier versions of this role iterated cephadm ceph-volume lvm create per disk and composed /dev/disk/by-path/... paths from host_vars (sas_path_prefix + path_phy). That assumed sietch's SAS expander topology and broke on the NVMe-RAID shape's PCI-ATA disks plus LV-backed SSD OSD.

The current flow renders a cephadm OSD service spec from per-host data and applies it via ceph orch apply osd -i. Cephadm handles device path resolution, LUKS encryption (encrypted: true), LVM provisioning, and daemon deployment. The role is hardware-shape-agnostic -- the only shape-aware logic is the template's Jinja conditional. Device paths are listed explicitly rather than filtered by rotational, so cephadm never auto-discovers and claims an OS or block.db partition.

Hardware-shape independence in the template

templates/osd-spec.yml.j2 renders one document per host (sietch nodes have unique SAS prefixes per chassis, so a shared spec doesn't work) with two shape branches:

  • Sietch (sas_path_prefix defined): data path = /dev/disk/by-path/{{ sas_path_prefix }}-{{ path_phy }}-lun-0; SSD OSD = partition on the SAS-attached SSD via path_phy + partition.
  • NVMe-RAID shape (sas_path_prefix undefined): data path = /dev/disk/by-path/{{ path_phy }} (operator authors the full PCI-ATA identifier in host_vars); SSD OSD = LV via the lv field (/dev/{{ lv }}).

db_devices.paths is always /dev/{{ db }} -- both shapes use LVs for block.db, no composition needed.

Idempotency

ceph orch apply osd is idempotent -- re-applying the same spec is a no-op when deployed OSDs match. New disks (populating an empty bay later, future expansion) are picked up automatically on the next apply. Existing OSDs are not destroyed by a spec apply -- removal requires explicit ceph orch osd rm.

Defensive noin handling

The spec-based flow doesn't need the noin flag (cephadm rolls out OSDs gracefully one at a time). The role's tail still includes a ceph osd unset noin task as a defensive cleanup -- stale noin flags from a prior failed run of the older imperative flow can leave the cluster degraded; the unconditional unset clears that safely (no-op when already unset).

Reweight-zero safety net

osds.yml ends with a task that fixes any OSD stuck at reweight=0 by running ceph osd reweight <id> 1.0. Rare with the spec-based flow but kept as a backstop against an OSD coming up while noin was set externally.


11. Monitoring

What cephadm auto-deploys

cephadm's bootstrap automatically deploys:

  • node-exporter on every node
  • ceph-exporter on every node
  • prometheus (single instance, cephadm-managed)
  • alertmanager (single instance)
  • grafana (single instance, with pre-built Ceph dashboards)
  • 89 Prometheus alert rules across 16 groups

What we configure

roles/ceph_deploy/tasks/monitoring.yml handles only integration:

  1. Enable the prometheus MGR module (if not already enabled)
  2. Wait for all five monitoring service types to report running > 0
  3. Set dashboard integration URLs for Prometheus, Alertmanager, Grafana (using the bootstrap node's bond_ip)
  4. Set Grafana admin credentials from the op-injected vault_grafana_admin_password
  5. Disable Grafana SSL cert verification in dashboard (self-signed cert)

roles/ceph_tuning/tasks/main.yml verifies the alert rule count and warns if fewer than 10 rule groups are loaded (expects 16+).


12. Provision host internals

Ten-phase flow

flowchart TB
    P1["detect.yml<br/><i>live image + UEFI assertions, SSD discovery</i>"]
    P2["prerequisites_live.yml<br/><i>apt setup on live image, install debootstrap/mdadm/lvm2</i>"]
    P3["disks.yml<br/><i>partition, mdraid, LVM, mount at /mnt</i>"]
    P4["install.yml<br/><i>debootstrap Bookworm into /mnt</i>"]
    P5["configure.yml<br/><i>hostname, hosts, network, fstab, mdadm templates</i>"]
    P6["chroot_packages.yml<br/><i>bind mounts, apt install, machine-id, SSH keys</i>"]
    P7["admin_user.yml<br/><i>ansible-iac (key-only) inside chroot;<br/>ops user is created post-boot by the baseline role</i>"]
    P8["bootloader.yml<br/><i>initramfs, grub-install, efibootmgr</i>"]
    P9["finalize.yml<br/><i>marker, ESP mirror, unmount, reboot</i>"]
    P10["unmount.yml<br/><i>reverse-order cleanup (shared with rescue)</i>"]

    P1 --> P2 --> P3 --> P4 --> P5 --> P6 --> P7 --> P8 --> P9 --> P10

An optional wipe-osds.yml runs right after disks.yml when provision_wipe_osd_disks=true, zapping prior OSD signatures off the data disks before install -- used when rebuilding a node that was previously a Ceph member.

Marker-driven resume gate

After disks.yml runs, main.yml checks for /mnt/etc/ceph-provisioned.json. If present and the hostname matches, all chroot phases (4-8) plus the marker/ESP block are skipped. The role goes straight to unmount + reboot.

This prevents:

  • Re-binding bind mounts that are already in place
  • Re-rotating SSH host keys (would break known_hosts)
  • Re-hashing the ops password with a fresh salt
  • Re-running grub-install for no reason
  • Overwriting the marker with a stale provisioned_at timestamp

The marker filename (ceph-provisioned.json) is project-scoped, not cluster-scoped -- every Ceph cluster (sietch, future) writes the same filename. The marker's contents identify which cluster + host the machine belongs to.

Block/rescue cleanup

The entire provisioning sequence (phases 2-9) runs inside a block/rescue. If any phase fails, the rescue block includes unmount.yml which tears down chroot bind mounts and the /mnt hierarchy in reverse order, then re-raises the failure. This ensures the next run starts from a clean mount state.


13. CI/CD and roadmap

Live today

  • CI / GitHub Actions -- .github/workflows/infra.yml applies the Terragrunt stacks from CI. A discover job scans tf/deployment/<partition>/<region>/<stack>/ into a {partition, region, stack} matrix, so adding a stack needs no workflow edit. plan runs with each partition's read SA; apply runs with the write SA, gated behind a per-region GitHub Environment with required reviewers (staging-austin, staging-global, prod-global, prod-htz-fsn1). Apply order is global (NetBird + DNS) -> site NetBird -> node-touching stacks (ceph / talos / fabric). Connectivity to the bare-metal nodes is over the NetBird overlay.
  • Talos K8s as a sibling stack -- tf/deployment/<partition>/<region>/talos/ shares the terragrunt root config and S3 backend, with its own state key (yucca/<partition>/<region>/talos/terraform.tfstate). The staging/austin talos stack is in the tree.

Roadmap

  • TF-managed onepassword_item resources -- re-enable the dormant resources in secrets.tf.disabled once the dedicated ceph service account lands, so 1P items are TF-owned rather than created by hand.
  • OSD LUKS keys in 1P -- store dm-crypt keys for DR. Deferred until the hybrid is stable.

See also

Topic Doc
TF/Terragrunt detail tf/README.md
Wrapper script reference docs/scripts.md
Secrets catalog + rotation docs/secrets.md
Trust boundaries + encryption docs/security-model.md
Hardware specs + network topology docs/hardware.md
Coding idioms and anti-patterns docs/patterns.md
Adding a new cluster (walkthrough) docs/adding-a-cluster.md