Six VictoriaLogs dashboards (yucca-api, admin-api, michael, web,
metrics-worker, meta): by-level volume histograms via the hits endpoint,
error streams with top-message tables, per-service request breakdowns
(status, handlers/routes, slowest requests, top users and source networks on
michael, cron heartbeat on metrics-worker, regex status classes on meta),
and full live streams with LogsQL filter and request-id textboxes. Query
types are the plugin's real enum (instant/statsRange/hits); every query was
validated against the live VictoriaLogs. No prometheus datasource variable
on these boards, which keeps the linter's PromQL rule away from LogsQL.
* fix(o11y): pin 5m staleness lookback on alert rule queries
VictoriaMetrics uses an instant query's step as the staleness lookback and
Grafana derives step from intervalMs, clamped to ~15s. Metrics scraped less
often than the step resolve to no-data and noDataState: OK parks the rule at
Normal while its condition is true: the junos-based rules (60s NETCONF
scrape) never fired, with the Colt transit session down. intervalMs: 300000
restores Prometheus-standard staleness for every cadence in the fleet.
* fix(o11y): make slow-cadence alert queries staleness-proof
Grafana sends the rule group's evaluation interval as the instant query's
step and VictoriaMetrics uses step as the staleness lookback; intervalMs on
the query model does not change it. Sources at or above the step's cadence
(junos at 60s, the 5m metering cron) lose that race often enough that a
multi-minute for can never sustain, so the rules sat Normal via
noDataState: OK through a real transit outage. Wrap every such selector in
last_over_time so evaluation is cadence-independent at any step.
* fix(o11y): point alert rules and imported dashboards at the fleet datasource
The default VictoriaMetrics datasource fronts vmauth-self-select, which
serves only the o11y cluster's own series: every rule over yucca, fabric or
ceph data evaluated to NoData and sat Normal through the live Colt transit
outage (confirmed via the Grafana API — all instances Normal (NoData) except
the cert rule, whose only firing labels were cluster=o11y). Pin
datasourceUid: VictoriaMetricsFleet (vmselect, whole fleet) on every rule
query and give the imported dashboards' $datasource variable the house
/^VictoriaMetrics Fleet$/ regex so they stop defaulting to the self-scoped
datasource.
* fix(o11y): pin 5m staleness lookback on alert rule queries
VictoriaMetrics uses an instant query's step as the staleness lookback and
Grafana derives step from intervalMs, clamped to ~15s. Metrics scraped less
often than the step resolve to no-data and noDataState: OK parks the rule at
Normal while its condition is true: the junos-based rules (60s NETCONF
scrape) never fired, with the Colt transit session down. intervalMs: 300000
restores Prometheus-standard staleness for every cadence in the fleet.
* fix(o11y): make slow-cadence alert queries staleness-proof
Grafana sends the rule group's evaluation interval as the instant query's
step and VictoriaMetrics uses step as the staleness lookback; intervalMs on
the query model does not change it. Sources at or above the step's cadence
(junos at 60s, the 5m metering cron) lose that race often enough that a
multi-minute for can never sustain, so the rules sat Normal via
noDataState: OK through a real transit outage. Wrap every such selector in
last_over_time so evaluation is cadence-independent at any step.
VictoriaMetrics uses an instant query's step as the staleness lookback and
Grafana derives step from intervalMs, clamped to ~15s. Metrics scraped less
often than the step resolve to no-data and noDataState: OK parks the rule at
Normal while its condition is true: the junos-based rules (60s NETCONF
scrape) never fired, with the Colt transit session down. intervalMs: 300000
restores Prometheus-standard staleness for every cadence in the fleet.
* chore(netbird): move to kebab naming
Render NetBird object names (groups, setup keys, policies, networks,
network-resources) as lowercase-kebab instead of UPPER_SNAKE, e.g.
YUCCA_PROD_HTZ_FSN1_MGMT → yucca-prod-htz-fsn1-mgmt. The 1Password setup-key
item titles stay UPPER_SNAKE (decoupled) so CI/ansible/talos op:// consumers
keep resolving.
Pin the futo-org/netbird provider to 1.0.2, which fixes the group
resources TF→API decode so a resource-tag group (htz-fsn1 `resources`) can be
renamed in place — no name pin needed.
* update locks
* chore(naming): naming names
* chore(netbird): move to kebab naming
Render NetBird object names (groups, setup keys, policies, networks,
network-resources) as lowercase-kebab instead of UPPER_SNAKE, e.g.
YUCCA_PROD_HTZ_FSN1_MGMT → yucca-prod-htz-fsn1-mgmt. The 1Password setup-key
item titles stay UPPER_SNAKE (decoupled) so CI/ansible/talos op:// consumers
keep resolving.
Pin the futo-org/netbird provider to 1.0.2, which fixes the group
resources TF→API decode so a resource-tag group (htz-fsn1 `resources`) can be
renamed in place — no name pin needed.
* update locks
fix(netbird): restore the talos plan's NetBird connect (revert -refresh=false)
#212 left the talos plan on `-refresh=false` with no overlay connect, but the
plan reads data.talos_cluster_health — a data source that dials the cluster over
the overlay and is NOT skipped by -refresh=false, so it hangs with no route to
the nodes. Restore the netbird-connect step on the plan job.
Documents the one-time out-of-band bootstrap: the CI setup key is minted by the
netbird apply, which is gated behind the plan that needs it, so seed it once via
TF_STACK_DIR=tf/deployment/staging/netbird mise run tf:apply.