20 Commits
Author SHA1 Message Date
Antoine Lecompte 85c936abc6 feat(bot): add dashboard (#622) 2026-09-02 14:34:46 +00:00
Andy Molenda 078c8f756c feat(o11y): put the scrub work-outstanding row on measured metrics (#616)
Replaces the inferred work-outstanding row with monk's measured metrics.

The old row derived coverage from `ceph_osd_scrub_*_read_bytes` against a `$deep_interval_days` offset. That estimate needed 28 days of stored metrics before it meant anything, counted re-scrubs of the same PG as progress, and divided read bytes by `ceph_osd_stat_bytes_used`. It also depended on two hardcoded interval variables that could drift from what the scrub scheduler actually targets.

Coverage, overdue bytes, overdue PGs and the age distribution now come from per-PG stamps. Every query dedups with `max by (cluster, pool_id)` because all five monk instances report the same cluster-wide data, and both sides of the coverage ratio are stored bytes, so nothing mixes logical with raw. The two interval variables are gone: panels judge against `ceph_scrub_target_interval_seconds`, read live from the cluster.

`Oldest deep scrub vs target` is stated as a ratio rather than days on purpose, so a single threshold is correct on both spice at 28 days and sietch at 7.

Validation: all ten expressions parse against o11y, `ceph_pool_metadata` has no duplicate `(cluster, pool_id)` so the `group_left(name)` join cannot fail, and running the identical join shape with `ceph_pool_stored` returns `prod-z1.rgw.buckets.data = 8252.03 TiB`, matching monk's own `ceph_scrub_pool_bytes` to the same figure. Panel values cannot be checked until the scrape job reaches prod at the next release promotion, since prod Flux is tag-pinned.

- `main` <!-- branch-stack -->
  - \#616 :point\_left:
2026-09-01 14:29:07 -07:00
Andy Molenda 9e4387e8fd feat(o11y): watch the NetBird overlay and give spice networking a dashboard (#540)
* fix(ceph): drop wt0 interface alerts before they reach alertmanager

* feat(o11y): watch the NetBird overlay and give spice networking a dashboard

* fix(o11y): project-scope netbird rules and join the missing-wt0 check per host

* fix(o11y): mirror the joined missing-wt0 expression in the dashboard
2026-08-25 10:06:31 -07:00
Antoine Lecompte 03ea70ea7c feat(o11y): michael resilience panels and alerts (#529) 2026-08-21 12:17:57 -04:00
Andy Molenda bff34886b7 feat(o11y): add spice scrub progress dashboard (#506)
* feat(o11y): add spice scrub progress dashboard

* feat(o11y): plain-language header and Ceph / Scrub title
2026-08-20 14:01:34 +00:00
Antoine Lecompte d760fb1993 feat(o11y): scope alerts to project=yucca, CNPG dashboard, title convention (#508) 2026-08-20 09:19:45 -04:00
Antoine Lecompte 09f10fcbf5 feat(o11y): per-service log dashboards (#504)
Six VictoriaLogs dashboards (yucca-api, admin-api, michael, web,
metrics-worker, meta): by-level volume histograms via the hits endpoint,
error streams with top-message tables, per-service request breakdowns
(status, handlers/routes, slowest requests, top users and source networks on
michael, cron heartbeat on metrics-worker, regex status classes on meta),
and full live streams with LogsQL filter and request-id textboxes. Query
types are the plugin's real enum (instant/statsRange/hits); every query was
validated against the live VictoriaLogs. No prometheus datasource variable
on these boards, which keeps the linter's PromQL rule away from LogsQL.
2026-08-20 12:25:05 +00:00
Antoine Lecompte 045491d437 feat(cnpg): make better (#496) 2026-08-19 15:27:09 -04:00
Antoine Lecompte 569c17ff63 fix(o11y): point alert rules and imported dashboards at the fleet datasource (#502)
* fix(o11y): pin 5m staleness lookback on alert rule queries

VictoriaMetrics uses an instant query's step as the staleness lookback and
Grafana derives step from intervalMs, clamped to ~15s. Metrics scraped less
often than the step resolve to no-data and noDataState: OK parks the rule at
Normal while its condition is true: the junos-based rules (60s NETCONF
scrape) never fired, with the Colt transit session down. intervalMs: 300000
restores Prometheus-standard staleness for every cadence in the fleet.

* fix(o11y): make slow-cadence alert queries staleness-proof

Grafana sends the rule group's evaluation interval as the instant query's
step and VictoriaMetrics uses step as the staleness lookback; intervalMs on
the query model does not change it. Sources at or above the step's cadence
(junos at 60s, the 5m metering cron) lose that race often enough that a
multi-minute for can never sustain, so the rules sat Normal via
noDataState: OK through a real transit outage. Wrap every such selector in
last_over_time so evaluation is cadence-independent at any step.

* fix(o11y): point alert rules and imported dashboards at the fleet datasource

The default VictoriaMetrics datasource fronts vmauth-self-select, which
serves only the o11y cluster's own series: every rule over yucca, fabric or
ceph data evaluated to NoData and sat Normal through the live Colt transit
outage (confirmed via the Grafana API — all instances Normal (NoData) except
the cert rule, whose only firing labels were cluster=o11y). Pin
datasourceUid: VictoriaMetricsFleet (vmselect, whole fleet) on every rule
query and give the imported dashboards' $datasource variable the house
/^VictoriaMetrics Fleet$/ regex so they stop defaulting to the self-scoped
datasource.
2026-08-19 19:26:16 +00:00
Antoine Lecompte 552735a527 fix(o11y): make slow-cadence alert queries staleness-proof (#501)
* fix(o11y): pin 5m staleness lookback on alert rule queries

VictoriaMetrics uses an instant query's step as the staleness lookback and
Grafana derives step from intervalMs, clamped to ~15s. Metrics scraped less
often than the step resolve to no-data and noDataState: OK parks the rule at
Normal while its condition is true: the junos-based rules (60s NETCONF
scrape) never fired, with the Colt transit session down. intervalMs: 300000
restores Prometheus-standard staleness for every cadence in the fleet.

* fix(o11y): make slow-cadence alert queries staleness-proof

Grafana sends the rule group's evaluation interval as the instant query's
step and VictoriaMetrics uses step as the staleness lookback; intervalMs on
the query model does not change it. Sources at or above the step's cadence
(junos at 60s, the 5m metering cron) lose that race often enough that a
multi-minute for can never sustain, so the rules sat Normal via
noDataState: OK through a real transit outage. Wrap every such selector in
last_over_time so evaluation is cadence-independent at any step.
2026-08-19 18:59:13 +00:00
Antoine Lecompte 0fe43daaf0 fix(o11y): pin 5m staleness lookback on alert rule queries (#499)
VictoriaMetrics uses an instant query's step as the staleness lookback and
Grafana derives step from intervalMs, clamped to ~15s. Metrics scraped less
often than the step resolve to no-data and noDataState: OK parks the rule at
Normal while its condition is true: the junos-based rules (60s NETCONF
scrape) never fired, with the Colt transit session down. intervalMs: 300000
restores Prometheus-standard staleness for every cadence in the fleet.
2026-08-19 18:17:46 +00:00
Antoine Lecompte 756e6e080e fix(alerts): fix bad alerts (#497) 2026-08-19 17:43:13 +00:00
Antoine Lecompte 8c6e9e325d chore(o11y): add more dashboard and alerts (#495) 2026-08-19 13:14:25 -04:00
Antoine Lecompte a4d422709b feat(michael): per-client parallelism and latency metrics (#427) 2026-08-05 12:09:29 -04:00
Antoine Lecompte 72b343ceba feat(michael): attribute traffic to the client's source asn (#428) 2026-08-05 15:13:04 +00:00
Andy Molenda 6ef5c243e9 feat(o11y): add spice recovery/backfill throughput dashboard (#402) 2026-07-30 14:07:35 +00:00
Andy Molenda 672cf49aad feat(o11y): add spice block.db and BlueFS headroom dashboard (#376) 2026-07-29 08:18:18 -07:00
Antoine Lecompte 9883dfe13c feat(yucca): more per user metrics (#345) 2026-07-27 15:38:31 +00:00
Antoine Lecompte 7dd44c269b feat(yucca): per user dashboards (#344) 2026-07-27 15:17:43 +00:00
Devin Buhl 37e8b2fc94 feat(o11y): ship dashboards as a signed OCI manifest bundle (#315)
Signed-off-by: Devin Buhl <devin@buhl.casa>
2026-07-23 12:16:45 -04:00