Six VictoriaLogs dashboards (yucca-api, admin-api, michael, web,
metrics-worker, meta): by-level volume histograms via the hits endpoint,
error streams with top-message tables, per-service request breakdowns
(status, handlers/routes, slowest requests, top users and source networks on
michael, cron heartbeat on metrics-worker, regex status classes on meta),
and full live streams with LogsQL filter and request-id textboxes. Query
types are the plugin's real enum (instant/statsRange/hits); every query was
validated against the live VictoriaLogs. No prometheus datasource variable
on these boards, which keeps the linter's PromQL rule away from LogsQL.
* fix(o11y): pin 5m staleness lookback on alert rule queries
VictoriaMetrics uses an instant query's step as the staleness lookback and
Grafana derives step from intervalMs, clamped to ~15s. Metrics scraped less
often than the step resolve to no-data and noDataState: OK parks the rule at
Normal while its condition is true: the junos-based rules (60s NETCONF
scrape) never fired, with the Colt transit session down. intervalMs: 300000
restores Prometheus-standard staleness for every cadence in the fleet.
* fix(o11y): make slow-cadence alert queries staleness-proof
Grafana sends the rule group's evaluation interval as the instant query's
step and VictoriaMetrics uses step as the staleness lookback; intervalMs on
the query model does not change it. Sources at or above the step's cadence
(junos at 60s, the 5m metering cron) lose that race often enough that a
multi-minute for can never sustain, so the rules sat Normal via
noDataState: OK through a real transit outage. Wrap every such selector in
last_over_time so evaluation is cadence-independent at any step.
* fix(o11y): point alert rules and imported dashboards at the fleet datasource
The default VictoriaMetrics datasource fronts vmauth-self-select, which
serves only the o11y cluster's own series: every rule over yucca, fabric or
ceph data evaluated to NoData and sat Normal through the live Colt transit
outage (confirmed via the Grafana API — all instances Normal (NoData) except
the cert rule, whose only firing labels were cluster=o11y). Pin
datasourceUid: VictoriaMetricsFleet (vmselect, whole fleet) on every rule
query and give the imported dashboards' $datasource variable the house
/^VictoriaMetrics Fleet$/ regex so they stop defaulting to the self-scoped
datasource.
* fix(o11y): pin 5m staleness lookback on alert rule queries
VictoriaMetrics uses an instant query's step as the staleness lookback and
Grafana derives step from intervalMs, clamped to ~15s. Metrics scraped less
often than the step resolve to no-data and noDataState: OK parks the rule at
Normal while its condition is true: the junos-based rules (60s NETCONF
scrape) never fired, with the Colt transit session down. intervalMs: 300000
restores Prometheus-standard staleness for every cadence in the fleet.
* fix(o11y): make slow-cadence alert queries staleness-proof
Grafana sends the rule group's evaluation interval as the instant query's
step and VictoriaMetrics uses step as the staleness lookback; intervalMs on
the query model does not change it. Sources at or above the step's cadence
(junos at 60s, the 5m metering cron) lose that race often enough that a
multi-minute for can never sustain, so the rules sat Normal via
noDataState: OK through a real transit outage. Wrap every such selector in
last_over_time so evaluation is cadence-independent at any step.
VictoriaMetrics uses an instant query's step as the staleness lookback and
Grafana derives step from intervalMs, clamped to ~15s. Metrics scraped less
often than the step resolve to no-data and noDataState: OK parks the rule at
Normal while its condition is true: the junos-based rules (60s NETCONF
scrape) never fired, with the Colt transit session down. intervalMs: 300000
restores Prometheus-standard staleness for every cadence in the fleet.