Replaces the inferred work-outstanding row with monk's measured metrics.
The old row derived coverage from `ceph_osd_scrub_*_read_bytes` against a `$deep_interval_days` offset. That estimate needed 28 days of stored metrics before it meant anything, counted re-scrubs of the same PG as progress, and divided read bytes by `ceph_osd_stat_bytes_used`. It also depended on two hardcoded interval variables that could drift from what the scrub scheduler actually targets.
Coverage, overdue bytes, overdue PGs and the age distribution now come from per-PG stamps. Every query dedups with `max by (cluster, pool_id)` because all five monk instances report the same cluster-wide data, and both sides of the coverage ratio are stored bytes, so nothing mixes logical with raw. The two interval variables are gone: panels judge against `ceph_scrub_target_interval_seconds`, read live from the cluster.
`Oldest deep scrub vs target` is stated as a ratio rather than days on purpose, so a single threshold is correct on both spice at 28 days and sietch at 7.
Validation: all ten expressions parse against o11y, `ceph_pool_metadata` has no duplicate `(cluster, pool_id)` so the `group_left(name)` join cannot fail, and running the identical join shape with `ceph_pool_stored` returns `prod-z1.rgw.buckets.data = 8252.03 TiB`, matching monk's own `ceph_scrub_pool_bytes` to the same figure. Panel values cannot be checked until the scrape job reaches prod at the next release promotion, since prod Flux is tag-pinned.
- `main` <!-- branch-stack -->
- \#616 :point\_left:
* fix(ceph): drop wt0 interface alerts before they reach alertmanager
* feat(o11y): watch the NetBird overlay and give spice networking a dashboard
* fix(o11y): project-scope netbird rules and join the missing-wt0 check per host
* fix(o11y): mirror the joined missing-wt0 expression in the dashboard
Six VictoriaLogs dashboards (yucca-api, admin-api, michael, web,
metrics-worker, meta): by-level volume histograms via the hits endpoint,
error streams with top-message tables, per-service request breakdowns
(status, handlers/routes, slowest requests, top users and source networks on
michael, cron heartbeat on metrics-worker, regex status classes on meta),
and full live streams with LogsQL filter and request-id textboxes. Query
types are the plugin's real enum (instant/statsRange/hits); every query was
validated against the live VictoriaLogs. No prometheus datasource variable
on these boards, which keeps the linter's PromQL rule away from LogsQL.
* fix(o11y): pin 5m staleness lookback on alert rule queries
VictoriaMetrics uses an instant query's step as the staleness lookback and
Grafana derives step from intervalMs, clamped to ~15s. Metrics scraped less
often than the step resolve to no-data and noDataState: OK parks the rule at
Normal while its condition is true: the junos-based rules (60s NETCONF
scrape) never fired, with the Colt transit session down. intervalMs: 300000
restores Prometheus-standard staleness for every cadence in the fleet.
* fix(o11y): make slow-cadence alert queries staleness-proof
Grafana sends the rule group's evaluation interval as the instant query's
step and VictoriaMetrics uses step as the staleness lookback; intervalMs on
the query model does not change it. Sources at or above the step's cadence
(junos at 60s, the 5m metering cron) lose that race often enough that a
multi-minute for can never sustain, so the rules sat Normal via
noDataState: OK through a real transit outage. Wrap every such selector in
last_over_time so evaluation is cadence-independent at any step.
* fix(o11y): point alert rules and imported dashboards at the fleet datasource
The default VictoriaMetrics datasource fronts vmauth-self-select, which
serves only the o11y cluster's own series: every rule over yucca, fabric or
ceph data evaluated to NoData and sat Normal through the live Colt transit
outage (confirmed via the Grafana API — all instances Normal (NoData) except
the cert rule, whose only firing labels were cluster=o11y). Pin
datasourceUid: VictoriaMetricsFleet (vmselect, whole fleet) on every rule
query and give the imported dashboards' $datasource variable the house
/^VictoriaMetrics Fleet$/ regex so they stop defaulting to the self-scoped
datasource.
* fix(o11y): pin 5m staleness lookback on alert rule queries
VictoriaMetrics uses an instant query's step as the staleness lookback and
Grafana derives step from intervalMs, clamped to ~15s. Metrics scraped less
often than the step resolve to no-data and noDataState: OK parks the rule at
Normal while its condition is true: the junos-based rules (60s NETCONF
scrape) never fired, with the Colt transit session down. intervalMs: 300000
restores Prometheus-standard staleness for every cadence in the fleet.
* fix(o11y): make slow-cadence alert queries staleness-proof
Grafana sends the rule group's evaluation interval as the instant query's
step and VictoriaMetrics uses step as the staleness lookback; intervalMs on
the query model does not change it. Sources at or above the step's cadence
(junos at 60s, the 5m metering cron) lose that race often enough that a
multi-minute for can never sustain, so the rules sat Normal via
noDataState: OK through a real transit outage. Wrap every such selector in
last_over_time so evaluation is cadence-independent at any step.
VictoriaMetrics uses an instant query's step as the staleness lookback and
Grafana derives step from intervalMs, clamped to ~15s. Metrics scraped less
often than the step resolve to no-data and noDataState: OK parks the rule at
Normal while its condition is true: the junos-based rules (60s NETCONF
scrape) never fired, with the Colt transit session down. intervalMs: 300000
restores Prometheus-standard staleness for every cadence in the fleet.