Step 15 previously only added the canonical 1P key on drift, leaving any
cephadm-generated random key in place as a second valid credential, and its
drift check only inspected keys[0]. When rotate_s3_keys is set, converge to
exactly the canonical key: add it if missing, then retire every non-canonical
key. The canonical key is added before any stray is removed, so the user is
never left without a working key. Still gated behind rotate_s3_keys because
rotating a live S3 key is disruptive to current consumers.
Verified on sietch (already reconciled): with rotate_s3_keys=true the run
reports canonical_present=True stray_count=0 and makes no changes.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The security role allowlisted every monitoring port except 8765, the cephadm
service discovery endpoint on the mgr. Prometheus fetches its scrape target
list from there via http_sd, so with the port dropped it discovered zero
targets and Grafana showed no data.
Add ceph_firewall_service_discovery_port (8765) to the role defaults and the
nftables template. Applied to sietch: the SD endpoint now returns 200 from the
Prometheus host, Prometheus has 7 active targets up, and ceph metrics flow
again.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
cephadm seeds the Grafana admin login password only at first deploy, and
nothing reconciled it, so it drifted from the vault. Direct login at :3000
failed, and the dashboard Grafana API calls (which authenticate as the same
admin user) were also affected.
Add an idempotent task that resets the Grafana admin password to
ceph_grafana_admin_password on every converge via grafana cli
reset-admin-password, which writes the sqlite DB on the host volume so it
persists across restarts. Also rename the existing task to make clear it sets
the dashboard Grafana API password, not the admin login.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The dashboard password was only set during the one-time cephadm bootstrap
(guarded by `not ceph_conf.stat.exists`), with no idempotent reconcile, so a
cluster bootstrapped out of band kept a random admin password that no later
playbook run would correct.
Add an enforce task that sets the password from ceph_dashboard_password on
every run, plus a verify-phase assertion that logs into the dashboard with the
vaulted credential and fails the play on drift.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
local_file stored the destination path in shared remote state, derived from
get_repo_root(). git worktrees each resolve get_repo_root() to their own root,
so state became bound to whichever worktree last applied. Any apply from a
different checkout then force-replaced every rendered file and rebound state.
The module now emits inventory and secrets content as a render output instead
of local_file resources. ansible/ceph/scripts/render-inventories.sh reads that
output and writes the files using a path derived from its own location, so no
checkout-specific path ever enters state.
Also exclude painbox from the managed clusters (in active use by Zack); its
spec moves to clusters.example.tfvars and its 1Password items are left alone.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
chore: update ansible/ceph/inventories/sietch[...]/group_vars/all/vars.yml
Updated Sietch-Ceph bond_interfaces to match the new 25G network link names:
- eno1 -> eno1np0
- eno2 -> eno2np1
s3.dev.austin.int.futo.cloud and its virtual-hosted wildcard now resolve
publicly, round-robin across the three Sietch ceph nodes. Records are
declarative in tf/deployment/dev/dns; the Cloudflare token resolves from
1Password via tf/.env. The cluster already expected these names, so
cephadm needed no changes.
See tf/README.md and ansible/ceph/docs/s3-integration.md.
Run a Talos K8s cluster as libvirt VMs on the existing 3-node Ceph
cluster, using its idle CPU/memory headroom instead of new hardware.
Ansible provisions the hypervisor substrate and VMs; Terraform renders
the inventory and bootstraps the cluster.
See ansible/talos/README.md and docs/runbooks/cluster-bring-up.md.
* feat(ceph): import yucca-ceph ansible + terraform infrastructure
Imports the yucca-ceph Ansible tree into ansible/ceph/ and adds the
Terraform stack at tf/ that drives it. Cuts over from ansible-vault
to the hybrid secrets architecture (TF as inventory authority, 1P
as secrets store, op-inject at deploy time) in one atomic move.
Source: internal yucca-ceph working tree; fresh subtree-style
import, history not preserved. Andy continues operating sietch +
painbox post-merge; yucca-team hosts the code and reviews changes.
What it adds:
- sietch (3-node Austin, production Ceph S3 backend, untouched
by this PR)
- painbox (single-node Hetzner SX295 in Helsinki) as a second
deployable cluster
- Future clusters land by appending to clusters.auto.tfvars in
the matching environment stack (tf/deployment/<env>/ceph/) —
no per-cluster TF code required
How it works (full map: ansible/ceph/docs/architecture.md):
- tf/shared/modules/ceph-cluster renders inventory.ini variants
+ secrets.yml.tpl per cluster from clusters.auto.tfvars
- secrets.yml.tpl carries op:// refs; `op inject -f` resolves
them at deploy time from the matching yucca_tf_<env> vault
- State in OVH yucca-tf-state bucket (key ceph/<env>/<stack>/)
- 11 ADRs capture the decisions: ansible/ceph/docs/adr/
Out of scope (intentional):
- LUKS keys not yet in 1P (deferred until hybrid is stable)
- tf/shared/modules/ceph-cluster/secrets.tf.disabled is dormant;
today's 1P items via `op item create` per
ansible/ceph/docs/adding-a-cluster.md
- Talos K8s on sietch is a separate workstream
Atomicity + rollback: TF-rendered inventory + secrets-template
files are gitignored (TF generates them) and ansible-vault removal
is coupled to the op-inject path. Splitting this PR lands in a
non-bootable state — merge as one unit. The merge itself is
reversible via `git revert` until the post-merge `tf:apply` runs;
after apply, full rollback needs state restore or `tofu state mv`
(land + validate before applying).
Dev-env impact: adds opentofu + terragrunt to yucca root mise tools
plus a self-contained ansible/ceph/.mise.toml. No new commands or
prereqs for immich-side contributors who don't touch ceph or run
tf:* tasks.
Verification:
- `mise run lint` (from ansible/ceph/): 130 files, 0 warnings
- `mise run check`: 19 playbooks parse clean
- `mise run tf:plan`: succeeds; 7 expected file path-rename
replacements (3 painbox + 4 sietch). State drift from import,
no cluster-side change.
- painbox deployed 2026-04-26 on the new code path: Bookworm +
Ceph Tentacle, 15 OSDs (14 HDD + 1 SSD) up + in, mon/mgr/rgw
running. HEALTH_WARN is expected on a single-node cluster.
Post-merge: from the yucca root, `mise run tf:apply` flips the
bucket state to the new monorepo paths (the 7 renames above).
* fix(ceph): exempt ansible/ and tf/ subtrees from root prettier
The imported infrastructure subtrees enforce their own format
conventions (yamllint + ansible-lint inside ansible/ceph/; tofu fmt
inside tf/). Prettier on ansible YAML reflows long Jinja2 expressions
and shell command blocks in unwanted ways, so root prettier checks
are skipped for both subtrees.
Also reformat root README.md table column alignment to match prettier
conventions (only the imported subtrees are exempt; yucca-side files
including the root README still follow root prettier rules).
* fix(ceph): clean up secrets tmpfile after ansible-playbook exits
`ansible-play.sh` rendered the resolved secrets file via `op inject`
into a `mktemp` tmpfile, set up a `trap 'rm -f "$TMPFILE"' EXIT INT
TERM`, then `exec`'d ansible-playbook. The `exec` replaced the bash
shell entirely, so the EXIT trap never fired — every play left a
plaintext-secrets file in /tmp.
In practice this was masked because /tmp is tmpfs (RAM only on this
operator's setup), so files evaporate on reboot. But within an
operator session, files accumulated linearly with each playbook
invocation. Recent count on the import-PR session: 38 files.
Drop the `exec`. With `set -euo pipefail` already on, bash:
- propagates ansible-playbook's exit code (set -e)
- fires the EXIT trap before exiting (always)
- cleans up the tmpfile on success, failure, or signal
Verified: `CEPH_ENV=... scripts/ansible-play.sh status.yml
--syntax-check` creates and removes the tmpfile within the same
invocation — /tmp is clean before and after.
`scripts/preflight.sh` uses the same trap pattern but does not
`exec`, so its tmpfile cleanup was already correct (and the suffix
differs: `-secrets-test.yml` vs `-secrets.yml`, confirming
ansible-play.sh as the sole offender).