* feat(monk): measure scrub lateness against ceph's schedule and the mon warn deadline
* fix(monk): measure lateness past ceph's latest scrub target
* fix(monk): withhold late and breach until the first config read
* fix(monk): gate late and breach on the snapshot's config source
* fix(o11y): key the deep scrub verdict on lateness rather than interval age
* fix(o11y): base the scrub verdict on ceph's latest target
* fix(o11y): note that monk withholds late and breach until a read
* feat(ci): let the ceph converge target one playbook and host
The dispatch could only run the full provisioning pipeline, which never touched monk.yml at all, so a pin bump or a single-host canary had no CI path and ran from a workstation against prod. The playbook is an allowlisted choice mapped to existing mise tasks, and the limit is free text so it reaches the shell through the environment rather than a template expansion. harden and site had no mise task despite being the two playbooks operators run targeted most often. Directory-local invocation stays correct when #612 turns on monorepo mode.
* fix(ci): forward arguments through the ceph deploy task
mise appends task arguments to a single-command task but not to a script block, so a --limit passed to deploy was silently dropped and every host in the inventory converged. That made the new ceph_limit input a guard that reads as applied while doing nothing, on the option the dispatch defaults to. Verified against a stub: each play now receives the limit, and the no-argument case is unchanged.
* fix(ci): stop a ceph converge from applying prod terraform
run_ceph_ansible activates the partition so the converge can run, but the Terragrunt apply step was gated only on the stack not being fabric, so dispatching a converge also applied every prod stack with the write service account into environments that carry no reviewers. The inventory is rendered from TF state rather than from an apply, so the converge needs nothing applied first, and force_apply remains the explicit way to apply without a code change. Also fails the step when a limit leaves a play with no hosts, which ansible reports as success, and corrects the header and input descriptions that described the old coupling and oversold what a limit does.
* fix(ci): default the ceph converge to a read-only check and fail closed
ceph_playbook defaulted to deploy, so setting run_ceph_ansible and nothing else ran the six-play provisioning pipeline against a live cluster; status is the safe default and the destructive choices are now explicit. The no-hosts guard keys on ansible wording and fails open if that changes, so a recap check fails closed alongside it. Records that skipping the apply trades an implicit reconcile for a staleness window: state means the last apply, so a clusters.auto.tfvars change needs force_apply.
* docs(ci): record the verified monorepo task behavior
Monorepo mode is live after #612, so the claim about it is no longer a prediction: ansible/ceph is a config_root, the tasks are addressable as //ansible/ceph:<task>, and the bare name still resolves and still forwards arguments from inside the root. Verified both invocation forms against the merged change.
* style(ci): drop trailing whitespace left by the ascii pass
* fix(ci): only fail the converge when no play matched anything
The skip check fired on any play matching no hosts, so a site run narrowed to a non-mon host converged its other eight plays and was then reported as having converged nothing. A skipped play is now a warning and the recap check is the sole failure condition, which still catches the case that motivated the guard: monk alone under an off-group limit, where no play runs at all.
* feat: enable mise mono repo
* fix: set config_roots
* ci: update mise action for software/publish
* chore: generate mise.lock
* ci: only update `ci.yml` with mise 2026.8.6 (no lock)
* ci: prefer `mise run`
Moves both clusters to `0.0.371`, the first build that counts scrub completions and tracks omap apart from data bytes (#617). Digest resolved from the registry index for the tag Deploy run 371 published.
Spice runs five mon instances and sietch three, all currently on `0.0.355`; the fleet digest matches that pin exactly today, so this is the whole delta.
The sietch comment named the previous tag and the reason for it, which stopped being true the moment the pin moved, so it now names what `0.0.371` adds instead. The rest of that file's diff is the ASCII normalization the repo asks for in files being touched.
Nothing consumes the new metrics yet: the pace and time-to-clear panels land once the converge has run and the counters have a predecessor snapshot to diff against.
- `main` <!-- branch-stack -->
- \#618 :point\_left:
Backlog size and age answer how far behind the cluster is; nothing answered how fast it drains. Ceph exports no per-PG completion counter, and the old dashboard approximated pace from OSD scrub read-byte counters, which mixed raw bytes with logical ones and counted a re-scrub of the same PG as progress.
monk now diffs each PG's stamp against the previous refresh and counts the PGs and bytes that actually completed a scrub. A PG missing from either side is skipped rather than counted, so pool deletion and PG splits do not register as completions, and the first refresh after start has no predecessor so the counters simply begin at zero.
These are counters, so they dedup differently from the gauges: rate first, then `max by (cluster, pool_id)`. Taking `max` across instances before `rate` would read an instance restart as a jump, because the five instances start at different times and each counts independently. The README carries that rule and the time-to-clear query.
Also tracks omap apart from data bytes. `stat_sum.num_bytes` excludes omap, so `prod-z1.rgw.buckets.index` reports zero bytes while holding 164 GB across 524M keys on spice: every one of its PGs could go overdue without moving byte-weighted coverage. Existing metrics keep their meaning and `ceph_scrub_pool_bytes` still matches `ceph_pool_stored`; the omap figures are additive.
- `main` <!-- branch-stack -->
- \#617 :point\_left:
Replaces the inferred work-outstanding row with monk's measured metrics.
The old row derived coverage from `ceph_osd_scrub_*_read_bytes` against a `$deep_interval_days` offset. That estimate needed 28 days of stored metrics before it meant anything, counted re-scrubs of the same PG as progress, and divided read bytes by `ceph_osd_stat_bytes_used`. It also depended on two hardcoded interval variables that could drift from what the scrub scheduler actually targets.
Coverage, overdue bytes, overdue PGs and the age distribution now come from per-PG stamps. Every query dedups with `max by (cluster, pool_id)` because all five monk instances report the same cluster-wide data, and both sides of the coverage ratio are stored bytes, so nothing mixes logical with raw. The two interval variables are gone: panels judge against `ceph_scrub_target_interval_seconds`, read live from the cluster.
`Oldest deep scrub vs target` is stated as a ratio rather than days on purpose, so a single threshold is correct on both spice at 28 days and sietch at 7.
Validation: all ten expressions parse against o11y, `ceph_pool_metadata` has no duplicate `(cluster, pool_id)` so the `group_left(name)` join cannot fail, and running the identical join shape with `ceph_pool_stored` returns `prod-z1.rgw.buckets.data = 8252.03 TiB`, matching monk's own `ceph_scrub_pool_bytes` to the same figure. Panel values cannot be checked until the scrape job reaches prod at the next release promotion, since prod Flux is tag-pinned.
- `main` <!-- branch-stack -->
- \#616 :point\_left:
Flips monk on for the spice mons and restores the `ceph-scrub` scrape job, which #599 held back so that targets and exporter would land together.
The adelia canary collected cleanly against the live cluster: the interval read resolved spice's tuned 28d deep target rather than the 7d cold-start fallback, 4257 PGs classified with zero parse errors, on-host collection took 2.1s, and mgr CPU sat at 50.6% of one core, inside the 45.7-52.3% idle band measured before monk existed.
All five instances are scraped rather than one. Every instance returns the same cluster-wide data, so the redundancy costs about 3% of one mgr core and buys availability that matches the quorum's own fault model; a host rebuild gets the exporter back on the next converge because the role targets the TF-rendered `ceph_mon` group. Queries dedup with `max by (cluster, pool_id)`, applied identically to both sides of any ratio.
- `main` <!-- branch-stack -->
- \#615 :point\_left:
Two parts.
**Pin and gate.** monk is pinned on spice to the digest sietch has been running since this morning, and left off. The fluent-bit precedent in the same file is the reason: `site.yml` carries `monk.yml`, so an enabled flag lands the exporter on all five mons at the next converge rather than on the canary host. The canary runs `harden.yml` first, because the `:9284` accept rule ships in `roles/security` and `roles/scrub_exporter` asserts it is live before staging anything; without that ordering the canary aborts on the assert, after `run_once` has already minted `client.scrub-exporter`.
**Scrape targets follow the exporter.** The `ceph-scrub` job shipped in #595 with five spice targets while monk runs on none of them. All five alert rules are `noDataState: OK`, so no targets is silence, but five unreachable targets synthesizes `up=0` and holds `MonkDown` firing per instance. The job returns with the enablement flip so targets and exporter land together.
Measured on adelia, which hosts the active mgr: `pg ls` is 11.5 MB over 4222 PGs and costs the mgr 0.78 cpu-s per call, from 20 saturated queries against a paired idle window. Five mons at the 2m default therefore cost it about 3% of one core, so the default stands.
- `main` <!-- branch-stack -->
- \#599 :point\_left:
Two independent parts.
**Release delivery.** `v` tags publish on the release event, in a new Release Images workflow, rather than from whichever Deploy run survives the concurrency queue. Deploy now queues FIFO instead of discarding superseded runs, and ships only build tags. Release Images waits out that commit's Deploy, then retags the matching `sha-` image per app, or builds from the tag tree when none exists; it is re-runnable against any tag.
**PR image gating.** build-check runs on every PR and aggregates into one `Result` check, so image builds become requirable in branch protection. The per-app job names are dynamic and cannot be.
* feat(o11y): scrape monk on the spice mons and alert on its health
* fix(o11y): cover dead monk instances and fit staleness inside snapshot expiry
* fix(o11y): name the rule that actually persists past snapshot expiry
* fix(o11y): fire collect-failing before the staleness rule expires
* feat(ceph): deploy monk onto the mon hosts
* fix(ceph): target the ceph_mon group and no_log the keyring tasks
* chore(ceph): drop the monk interval vars, monk reads the cluster now
* fix(ceph): restart monk when the pull lands a new digest
* feat(monk): add monk, a measured ceph scrub-backlog exporter
* fix(monk): harden collection failure paths
* docs(monk): defer the deploy role reference to its follow-up PR
* fix(monk): emit the age distribution as a real histogram and document operational contracts
* feat(monk): read overdue targets from the cluster instead of flags
* ci(monk): register the monk image build filter
* chore(monk): cut comments the types already carry
* ci: build changed app images on pull requests
* ci: cover shared and root image inputs in build-check
* ci: treat .dockerignore as an image input
* ci: dockerignore governs the go image contexts too
* ci: apps register their build filter in the PR that adds them