* feat(ci): let the ceph converge target one playbook and host
The dispatch could only run the full provisioning pipeline, which never touched monk.yml at all, so a pin bump or a single-host canary had no CI path and ran from a workstation against prod. The playbook is an allowlisted choice mapped to existing mise tasks, and the limit is free text so it reaches the shell through the environment rather than a template expansion. harden and site had no mise task despite being the two playbooks operators run targeted most often. Directory-local invocation stays correct when #612 turns on monorepo mode.
* fix(ci): forward arguments through the ceph deploy task
mise appends task arguments to a single-command task but not to a script block, so a --limit passed to deploy was silently dropped and every host in the inventory converged. That made the new ceph_limit input a guard that reads as applied while doing nothing, on the option the dispatch defaults to. Verified against a stub: each play now receives the limit, and the no-argument case is unchanged.
* fix(ci): stop a ceph converge from applying prod terraform
run_ceph_ansible activates the partition so the converge can run, but the Terragrunt apply step was gated only on the stack not being fabric, so dispatching a converge also applied every prod stack with the write service account into environments that carry no reviewers. The inventory is rendered from TF state rather than from an apply, so the converge needs nothing applied first, and force_apply remains the explicit way to apply without a code change. Also fails the step when a limit leaves a play with no hosts, which ansible reports as success, and corrects the header and input descriptions that described the old coupling and oversold what a limit does.
* fix(ci): default the ceph converge to a read-only check and fail closed
ceph_playbook defaulted to deploy, so setting run_ceph_ansible and nothing else ran the six-play provisioning pipeline against a live cluster; status is the safe default and the destructive choices are now explicit. The no-hosts guard keys on ansible wording and fails open if that changes, so a recap check fails closed alongside it. Records that skipping the apply trades an implicit reconcile for a staleness window: state means the last apply, so a clusters.auto.tfvars change needs force_apply.
* docs(ci): record the verified monorepo task behavior
Monorepo mode is live after #612, so the claim about it is no longer a prediction: ansible/ceph is a config_root, the tasks are addressable as //ansible/ceph:<task>, and the bare name still resolves and still forwards arguments from inside the root. Verified both invocation forms against the merged change.
* style(ci): drop trailing whitespace left by the ascii pass
* fix(ci): only fail the converge when no play matched anything
The skip check fired on any play matching no hosts, so a site run narrowed to a non-mon host converged its other eight plays and was then reported as having converged nothing. A skipped play is now a warning and the recap check is the sole failure condition, which still catches the case that motivated the guard: monk alone under an off-group limit, where no play runs at all.
Moves both clusters to `0.0.371`, the first build that counts scrub completions and tracks omap apart from data bytes (#617). Digest resolved from the registry index for the tag Deploy run 371 published.
Spice runs five mon instances and sietch three, all currently on `0.0.355`; the fleet digest matches that pin exactly today, so this is the whole delta.
The sietch comment named the previous tag and the reason for it, which stopped being true the moment the pin moved, so it now names what `0.0.371` adds instead. The rest of that file's diff is the ASCII normalization the repo asks for in files being touched.
Nothing consumes the new metrics yet: the pace and time-to-clear panels land once the converge has run and the counters have a predecessor snapshot to diff against.
- `main` <!-- branch-stack -->
- \#618 :point\_left:
Flips monk on for the spice mons and restores the `ceph-scrub` scrape job, which #599 held back so that targets and exporter would land together.
The adelia canary collected cleanly against the live cluster: the interval read resolved spice's tuned 28d deep target rather than the 7d cold-start fallback, 4257 PGs classified with zero parse errors, on-host collection took 2.1s, and mgr CPU sat at 50.6% of one core, inside the 45.7-52.3% idle band measured before monk existed.
All five instances are scraped rather than one. Every instance returns the same cluster-wide data, so the redundancy costs about 3% of one mgr core and buys availability that matches the quorum's own fault model; a host rebuild gets the exporter back on the next converge because the role targets the TF-rendered `ceph_mon` group. Queries dedup with `max by (cluster, pool_id)`, applied identically to both sides of any ratio.
- `main` <!-- branch-stack -->
- \#615 :point\_left:
Two parts.
**Pin and gate.** monk is pinned on spice to the digest sietch has been running since this morning, and left off. The fluent-bit precedent in the same file is the reason: `site.yml` carries `monk.yml`, so an enabled flag lands the exporter on all five mons at the next converge rather than on the canary host. The canary runs `harden.yml` first, because the `:9284` accept rule ships in `roles/security` and `roles/scrub_exporter` asserts it is live before staging anything; without that ordering the canary aborts on the assert, after `run_once` has already minted `client.scrub-exporter`.
**Scrape targets follow the exporter.** The `ceph-scrub` job shipped in #595 with five spice targets while monk runs on none of them. All five alert rules are `noDataState: OK`, so no targets is silence, but five unreachable targets synthesizes `up=0` and holds `MonkDown` firing per instance. The job returns with the enablement flip so targets and exporter land together.
Measured on adelia, which hosts the active mgr: `pg ls` is 11.5 MB over 4222 PGs and costs the mgr 0.78 cpu-s per call, from 20 saturated queries against a paired idle window. Five mons at the 2m default therefore cost it about 3% of one core, so the default stands.
- `main` <!-- branch-stack -->
- \#599 :point\_left:
* feat(ceph): deploy monk onto the mon hosts
* fix(ceph): target the ceph_mon group and no_log the keyring tasks
* chore(ceph): drop the monk interval vars, monk reads the cluster now
* fix(ceph): restart monk when the pull lands a new digest
* fix(ceph): let read-only probes run in check mode
* feat(ceph): add replace-osd playbook for a swapped HDD
* docs(ceph): correct the spice disk replacement procedure
* docs(ceph): hetzner hot-swaps SX295 drives on a live host
* feat(ceph): let replace-osd keep the old osd id
* docs(ceph): reserve the osd id before the replace-osd rebuild
* docs(ceph): spell out what orch osd rm --replace does and does not do
* fix(ceph): verify the rebuilt OSD by its own report and survive a queued stop
* docs(ceph): replace-osd when the cluster is not active+clean
* fix(ceph): let the queued action fire before judging the rebuilt OSD
* fix(ceph): refuse an explicit db slot the scan did not find free
* feat(ceph): ship spice host logs to o11y over netbird
* refactor(ceph): use fluent-bit instead of vector for host logs
* docs(ceph): tighten fluent-bit role comments
* fix(ceph): compress host logs to o11y with zstd
* fix(ceph): pin fluent-bit 5.0.9 and converge log shipping in site.yml
* refactor(ceph): resolve o11y over mesh DNS instead of a hosts entry
* fix(ceph): run log shipping last and tolerate isolated node failures
* fix(ceph): check log shipping preconditions before mutating the node
* fix(ceph): bind fluent-bit metrics to the fabric and open the port
* test(ceph): add a molecule scenario for the fluent-bit config
* fix(ceph): verify the o11y cert hostname, not just the chain
* fix(ceph): wait for the bind address so the metrics endpoint survives a reboot
* docs(ceph): plain-technical pass over the fluent-bit comments
* docs(ceph): cite the source for the journald field names
* chore(ceph): hold host log shipping off until the canary passes
* docs(ceph): cover spice alongside sietch in the ceph docs
* fix(ceph): take osd specs out of the cephadm serve loop
* fix(ceph): keep spice on per-host osd specs, collapse is greenfield-only
* docs(ceph): correct the committed-inventory claim for both clusters