Step 15 previously only added the canonical 1P key on drift, leaving any
cephadm-generated random key in place as a second valid credential, and its
drift check only inspected keys[0]. When rotate_s3_keys is set, converge to
exactly the canonical key: add it if missing, then retire every non-canonical
key. The canonical key is added before any stray is removed, so the user is
never left without a working key. Still gated behind rotate_s3_keys because
rotating a live S3 key is disruptive to current consumers.
Verified on sietch (already reconciled): with rotate_s3_keys=true the run
reports canonical_present=True stray_count=0 and makes no changes.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The security role allowlisted every monitoring port except 8765, the cephadm
service discovery endpoint on the mgr. Prometheus fetches its scrape target
list from there via http_sd, so with the port dropped it discovered zero
targets and Grafana showed no data.
Add ceph_firewall_service_discovery_port (8765) to the role defaults and the
nftables template. Applied to sietch: the SD endpoint now returns 200 from the
Prometheus host, Prometheus has 7 active targets up, and ceph metrics flow
again.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
cephadm seeds the Grafana admin login password only at first deploy, and
nothing reconciled it, so it drifted from the vault. Direct login at :3000
failed, and the dashboard Grafana API calls (which authenticate as the same
admin user) were also affected.
Add an idempotent task that resets the Grafana admin password to
ceph_grafana_admin_password on every converge via grafana cli
reset-admin-password, which writes the sqlite DB on the host volume so it
persists across restarts. Also rename the existing task to make clear it sets
the dashboard Grafana API password, not the admin login.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The dashboard password was only set during the one-time cephadm bootstrap
(guarded by `not ceph_conf.stat.exists`), with no idempotent reconcile, so a
cluster bootstrapped out of band kept a random admin password that no later
playbook run would correct.
Add an enforce task that sets the password from ceph_dashboard_password on
every run, plus a verify-phase assertion that logs into the dashboard with the
vaulted credential and fails the play on drift.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
local_file stored the destination path in shared remote state, derived from
get_repo_root(). git worktrees each resolve get_repo_root() to their own root,
so state became bound to whichever worktree last applied. Any apply from a
different checkout then force-replaced every rendered file and rebound state.
The module now emits inventory and secrets content as a render output instead
of local_file resources. ansible/ceph/scripts/render-inventories.sh reads that
output and writes the files using a path derived from its own location, so no
checkout-specific path ever enters state.
Also exclude painbox from the managed clusters (in active use by Zack); its
spec moves to clusters.example.tfvars and its 1Password items are left alone.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
chore: update ansible/ceph/inventories/sietch[...]/group_vars/all/vars.yml
Updated Sietch-Ceph bond_interfaces to match the new 25G network link names:
- eno1 -> eno1np0
- eno2 -> eno2np1
s3.dev.austin.int.futo.cloud and its virtual-hosted wildcard now resolve
publicly, round-robin across the three Sietch ceph nodes. Records are
declarative in tf/deployment/dev/dns; the Cloudflare token resolves from
1Password via tf/.env. The cluster already expected these names, so
cephadm needed no changes.
See tf/README.md and ansible/ceph/docs/s3-integration.md.
Run a Talos K8s cluster as libvirt VMs on the existing 3-node Ceph
cluster, using its idle CPU/memory headroom instead of new hardware.
Ansible provisions the hypervisor substrate and VMs; Terraform renders
the inventory and bootstraps the cluster.
See ansible/talos/README.md and docs/runbooks/cluster-bring-up.md.