Commit Graph
100 Commits
Author SHA1 Message Date
Antoine Lecompte a3b42946d5 fix(meta): log x-forwarded for IP address (#512) 2026-08-20 14:34:26 +00:00
Antoine Lecompte 153bf220e2 fix(o11y): field selector on log panels, variable-safe request tracing (#509)
* fix(o11y): field selector on log panels, variable-safe request tracing

* docs: one-line PR descriptions and commit messages
2026-08-20 13:59:32 +00:00
Antoine Lecompte 816ce71bd3 fix(migrations): re-date (#510) 2026-08-20 13:48:54 +00:00
Antoine Lecompte d760fb1993 feat(o11y): scope alerts to project=yucca, CNPG dashboard, title convention (#508) 2026-08-20 09:19:45 -04:00
Antoine Lecompte 4eaaaa6e6f fix(net): bgp (#507) 2026-08-20 13:07:35 +00:00
Antoine Lecompte 09f10fcbf5 feat(o11y): per-service log dashboards (#504)
Six VictoriaLogs dashboards (yucca-api, admin-api, michael, web,
metrics-worker, meta): by-level volume histograms via the hits endpoint,
error streams with top-message tables, per-service request breakdowns
(status, handlers/routes, slowest requests, top users and source networks on
michael, cron heartbeat on metrics-worker, regex status classes on meta),
and full live streams with LogsQL filter and request-id textboxes. Query
types are the plugin's real enum (instant/statsRange/hits); every query was
validated against the live VictoriaLogs. No prometheus datasource variable
on these boards, which keeps the linter's PromQL rule away from LogsQL.
2026-08-20 12:25:05 +00:00
Antoine Lecompte 045491d437 feat(cnpg): make better (#496) 2026-08-19 15:27:09 -04:00
Antoine Lecompte 569c17ff63 fix(o11y): point alert rules and imported dashboards at the fleet datasource (#502)
* fix(o11y): pin 5m staleness lookback on alert rule queries

VictoriaMetrics uses an instant query's step as the staleness lookback and
Grafana derives step from intervalMs, clamped to ~15s. Metrics scraped less
often than the step resolve to no-data and noDataState: OK parks the rule at
Normal while its condition is true: the junos-based rules (60s NETCONF
scrape) never fired, with the Colt transit session down. intervalMs: 300000
restores Prometheus-standard staleness for every cadence in the fleet.

* fix(o11y): make slow-cadence alert queries staleness-proof

Grafana sends the rule group's evaluation interval as the instant query's
step and VictoriaMetrics uses step as the staleness lookback; intervalMs on
the query model does not change it. Sources at or above the step's cadence
(junos at 60s, the 5m metering cron) lose that race often enough that a
multi-minute for can never sustain, so the rules sat Normal via
noDataState: OK through a real transit outage. Wrap every such selector in
last_over_time so evaluation is cadence-independent at any step.

* fix(o11y): point alert rules and imported dashboards at the fleet datasource

The default VictoriaMetrics datasource fronts vmauth-self-select, which
serves only the o11y cluster's own series: every rule over yucca, fabric or
ceph data evaluated to NoData and sat Normal through the live Colt transit
outage (confirmed via the Grafana API — all instances Normal (NoData) except
the cert rule, whose only firing labels were cluster=o11y). Pin
datasourceUid: VictoriaMetricsFleet (vmselect, whole fleet) on every rule
query and give the imported dashboards' $datasource variable the house
/^VictoriaMetrics Fleet$/ regex so they stop defaulting to the self-scoped
datasource.
2026-08-19 19:26:16 +00:00
Antoine Lecompte 552735a527 fix(o11y): make slow-cadence alert queries staleness-proof (#501)
* fix(o11y): pin 5m staleness lookback on alert rule queries

VictoriaMetrics uses an instant query's step as the staleness lookback and
Grafana derives step from intervalMs, clamped to ~15s. Metrics scraped less
often than the step resolve to no-data and noDataState: OK parks the rule at
Normal while its condition is true: the junos-based rules (60s NETCONF
scrape) never fired, with the Colt transit session down. intervalMs: 300000
restores Prometheus-standard staleness for every cadence in the fleet.

* fix(o11y): make slow-cadence alert queries staleness-proof

Grafana sends the rule group's evaluation interval as the instant query's
step and VictoriaMetrics uses step as the staleness lookback; intervalMs on
the query model does not change it. Sources at or above the step's cadence
(junos at 60s, the 5m metering cron) lose that race often enough that a
multi-minute for can never sustain, so the rules sat Normal via
noDataState: OK through a real transit outage. Wrap every such selector in
last_over_time so evaluation is cadence-independent at any step.
2026-08-19 18:59:13 +00:00
Antoine Lecompte 0fe43daaf0 fix(o11y): pin 5m staleness lookback on alert rule queries (#499)
VictoriaMetrics uses an instant query's step as the staleness lookback and
Grafana derives step from intervalMs, clamped to ~15s. Metrics scraped less
often than the step resolve to no-data and noDataState: OK parks the rule at
Normal while its condition is true: the junos-based rules (60s NETCONF
scrape) never fired, with the Colt transit session down. intervalMs: 300000
restores Prometheus-standard staleness for every cadence in the fleet.
2026-08-19 18:17:46 +00:00
Antoine Lecompte 756e6e080e fix(alerts): fix bad alerts (#497) 2026-08-19 17:43:13 +00:00
Antoine Lecompte 8c6e9e325d chore(o11y): add more dashboard and alerts (#495) 2026-08-19 13:14:25 -04:00
Antoine Lecompte ee43e55a2b feat(net): add colt (#488) 2026-08-19 13:04:17 -04:00
Antoine Lecompte b7bbe2e80e chore(yuctl): restructure and make pretty (#485) 2026-08-18 09:53:59 -04:00
Antoine Lecompte 0e3f47a769 perf(ci): stop the integration and e2e jobs waiting on rook-ceph (#435) 2026-08-07 13:40:51 +01:00
Antoine Lecompte 51472c37eb fix(infra): monitor all the things (#445) 2026-08-05 20:04:26 +00:00
Antoine Lecompteandgreptile-apps[bot] 75c77ecdfa feat(infra): add spegel (#436)
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-08-05 20:03:59 +00:00
Antoine Lecompte 02efb46f53 chore(claude): make claude comply (#432) 2026-08-05 12:54:55 -04:00
Antoine Lecompte bf69aee578 fix(yucca-sdk): repair the orchestration-api integration suites (#429) 2026-08-05 12:45:59 -04:00
Antoine Lecompte a4d422709b feat(michael): per-client parallelism and latency metrics (#427) 2026-08-05 12:09:29 -04:00
Antoine Lecompte 72b343ceba feat(michael): attribute traffic to the client's source asn (#428) 2026-08-05 15:13:04 +00:00
Antoine Lecompte a4c4b0c337 fix(yucca-api): format failure (#426) 2026-08-05 13:34:40 +00:00
Antoine Lecompte 8feda20065 feat(web): connections ui (#415) 2026-08-03 18:30:57 +00:00
Antoine Lecompte b427a7045b feat(yucca): per-connection billing rollup (#391) 2026-08-03 14:09:23 -04:00
Antoine Lecompte 55ece2fbaf feat(all): connections (#390) 2026-08-03 14:09:23 -04:00
Antoine Lecompte 94a5df9667 feat(yucca): per-user feature flags (#389) 2026-08-03 14:09:22 -04:00
Antoine Lecompte c606637450 fix(k8s): drop deadlocking topology readyExpr (#401) 2026-07-30 13:01:06 +00:00
Antoine Lecompte 1039337b13 feat(orchestrator): meta discovery and placement (#385) 2026-07-30 08:29:30 -04:00
Antoine Lecompte 7911c3bf11 feat(yuctl): config commands (#384) 2026-07-30 08:29:30 -04:00
Antoine Lecompte 597506abcc feat(michael): storage-cluster routing (#383) 2026-07-30 08:29:29 -04:00
Antoine Lecompte 51ea175933 feat(yucca-admin-api): scoped settings (#382) 2026-07-30 08:29:29 -04:00
Antoine Lecompte 5a73550573 feat(yucca): multi-site placement and public /meta (#381) 2026-07-30 08:29:29 -04:00
Antoine Lecompte 046aacb6b1 feat(k8s): fleet topology and meta discovery groundwork (#380) 2026-07-30 08:29:28 -04:00
Antoine Lecompte 7306bc4b4d fix(orchestrator): dev backup hook return type (#393)
* fix(orchestrator): dev backup hook return type

* chore: exempt grafana dashboards from prettier
2026-07-30 08:28:10 -04:00
Antoine Lecompte 6e4d6ccfcb feat(yuctl): more benchmarking (#371) 2026-07-29 09:06:30 -04:00
Antoine Lecompte a0c70541d5 feat(michael): change url (#363) 2026-07-29 12:09:01 +00:00
Antoine Lecompte 8d257aaeb8 fix(michael): retry loop (#367) 2026-07-28 22:58:39 +00:00
Antoine Lecompte 71bce9bf55 feat(net): switch to ecmp, dsr and jumbo frames (#366) 2026-07-28 16:23:55 -04:00
Antoine Lecompte 97652b9f76 feat(yuctl): digital ocean driven benchmarks (#365) 2026-07-28 17:08:06 +00:00
Antoine Lecompte 9cdeeb10d4 feat(yuctl): add tools warp (#356) 2026-07-28 11:58:49 +00:00
Antoine Lecompte 9883dfe13c feat(yucca): more per user metrics (#345) 2026-07-27 15:38:31 +00:00
Antoine Lecompte 7dd44c269b feat(yucca): per user dashboards (#344) 2026-07-27 15:17:43 +00:00
Antoine Lecompte af5ab66778 feat(infra): switch prod over to netkit (#342) 2026-07-27 15:07:29 +00:00
Antoine Lecompte 86a215a3db feat(all): email allowlist for invites (#341) 2026-07-27 10:24:53 -04:00
Antoine Lecompte bcc8ef7bab feat(admin-cli): make admin-cli faster (#340) 2026-07-27 14:19:53 +00:00
Antoine Lecompte 57a59232d0 feat(all): post-benchmark cleanup (#336) 2026-07-27 13:09:44 +00:00
Antoine Lecompte 6471d41079 fix(michael): dramatically improve michael latencies (#324) 2026-07-24 13:58:30 -04:00
Antoine Lecompte 1ff83d6754 fix(all): adjust benchmark related configs and fix metrics (#332) 2026-07-24 17:21:15 +00:00
Antoine Lecompte ecf969ba3e feat(admin): make the whole thing benchmarkable (#330) 2026-07-24 11:15:09 -04:00
Antoine Lecompte 001207e886 fix(admin): make secrets propagate (#329) 2026-07-24 15:06:54 +00:00
Antoine Lecompte 37ecd7bf58 feat(admin): make admin go brr (#327) 2026-07-24 14:47:10 +00:00
Antoine Lecompte 40289051ef fix(metrics-worker): prevent infinite looping (#326) 2026-07-24 13:50:43 +00:00
Antoine Lecompte 28d386d05d fix(ci): fix ci so it works (#325) 2026-07-24 13:29:23 +00:00
Antoine Lecompte 434df44ff8 fix(ci): fix ci so it works (#323) 2026-07-24 09:16:00 -04:00
Antoine Lecompte e1e5e6fa4d fix(admin): correct oidc issuer url (#321) 2026-07-24 13:15:01 +00:00
Antoine Lecompte 8afa3059ee feat(admin): wire up admin (#313) 2026-07-23 13:43:48 +00:00
Antoine Lecompte c938fc379a fix(ci): change bumping workflow (#311) 2026-07-23 08:59:05 -04:00
Antoine Lecompte 4cc69b4e9d chore(release): stamp versions in charts, packages, and michael (#306) 2026-07-22 19:17:11 +00:00
Antoine Lecompte 9a7b72d89c fix(mgmt): vlan interface names exceeded IFNAMSIZ (#305) 2026-07-22 19:17:06 +00:00
Antoine Lecompte 37e603c9df fix(michael): paginate on more than 1000 objects (#304) 2026-07-22 18:55:18 +00:00
Antoine Lecompte c9ff0400b3 chore(k8s): improve pinning (#303) 2026-07-22 14:52:32 -04:00
Antoine Lecompte 2de3f3c53c chore(k8s): pinning (#300) 2026-07-22 13:50:04 -04:00
Antoine Lecompte 55ccf7df26 fix(michael): better load balancer initialization (#297) 2026-07-22 17:33:51 +00:00
Antoine Lecompte 832c73e765 feat(o11y): shared crossair, no fill, legend ordering (#295)
* feat(o11y): shared crossair, no fill, legend ordering

* emdash kill
2026-07-22 16:24:38 +00:00
Antoine Lecompte c5f827e76c feat(o11y): oci dashboards (#294) 2026-07-22 14:45:24 +00:00
Antoine Lecompte 522630fec0 fix(o11y): wrong labels (#293) 2026-07-22 14:05:54 +00:00
Antoine Lecompte cc2e1f471e feat(o11y): more better monitoring (#291) 2026-07-22 09:51:33 -04:00
Antoine Lecompte 1d5ddd0e7f chore(k8s): release things (#289) 2026-07-22 13:35:09 +00:00
Antoine Lecompte 114904e966 feat(prod): continue prod (#288) 2026-07-22 09:09:30 -04:00
Antoine Lecompte bd74b16c7c feat(prod): deploy prod apps (#284)
feat(prod); deploy prod apps
2026-07-21 12:02:57 -07:00
Antoine Lecompte 5d95b04b57 chore(ci): temp ci override (#268) 2026-07-16 11:15:33 -04:00
Antoine Lecompte 1f9b84bb1d feat(prod): continue prod (#267) 2026-07-16 09:56:40 -04:00
Antoine Lecompte 3427483e5d fix(flux): branch ref (#255) 2026-07-13 13:43:17 +00:00
Antoine Lecompte 963364a034 feat(prod): prod (#247) 2026-07-13 13:15:50 +00:00
Antoine Lecompte 789e0361c5 fix(fabric): depends_on (#246) 2026-06-30 19:20:14 +00:00
Antoine Lecompte 9ad33c444d fix(typo): fix typo (#245) 2026-06-30 17:42:36 +00:00
Antoine Lecompte 7858d7318d fix(netbird): make mutating webhook not go boom (#244) 2026-06-30 17:22:27 +00:00
Antoine Lecompte 254b7a4b49 feat(k8s): cluster naming (#243)
* chore(netbird): move to kebab naming

Render NetBird object names (groups, setup keys, policies, networks,
network-resources) as lowercase-kebab instead of UPPER_SNAKE, e.g.
YUCCA_PROD_HTZ_FSN1_MGMT → yucca-prod-htz-fsn1-mgmt. The 1Password setup-key
item titles stay UPPER_SNAKE (decoupled) so CI/ansible/talos op:// consumers
keep resolving.

Pin the futo-org/netbird provider to 1.0.2, which fixes the group
resources TF→API decode so a resource-tag group (htz-fsn1 `resources`) can be
renamed in place — no name pin needed.

* update locks

* chore(naming): naming names
2026-06-30 13:16:50 -04:00
Antoine Lecompte 75b7808993 chore(netbird): move to kebab naming (#242)
* chore(netbird): move to kebab naming

Render NetBird object names (groups, setup keys, policies, networks,
network-resources) as lowercase-kebab instead of UPPER_SNAKE, e.g.
YUCCA_PROD_HTZ_FSN1_MGMT → yucca-prod-htz-fsn1-mgmt. The 1Password setup-key
item titles stay UPPER_SNAKE (decoupled) so CI/ansible/talos op:// consumers
keep resolving.

Pin the futo-org/netbird provider to 1.0.2, which fixes the group
resources TF→API decode so a resource-tag group (htz-fsn1 `resources`) can be
renamed in place — no name pin needed.

* update locks
2026-06-30 16:20:10 +00:00
Antoine Lecompte baae250f76 feat(fabric): management plane on a per cluster basis (#240) 2026-06-30 15:22:38 +00:00
Antoine Lecompte e00ea72a21 fix(kube): move to local volumes (#241) 2026-06-30 11:22:14 -04:00
Antoine Lecompte 012db81704 fix(netbird): policies (#239) 2026-06-30 10:31:31 -04:00
Antoine Lecompte 2ec0e668ad chore(netbird): switch provider (#238)
* chore(netbird): change provider, normalize

* commit
2026-06-30 10:09:42 -04:00
Antoine Lecompte 0c1b7f765c chore(fabric): switch over the fabric from the generated provider to … (#231)
* chore(fabric): switch over the fabric from the generated provider to a community provider

* cleanup
2026-06-29 16:03:50 -04:00
Antoine Lecompte f75cc90cf9 feat(prod): add mgmt-1 to terraform ownership (bye tailscale) (#230) 2026-06-29 14:10:28 -04:00
Antoine Lecompte ceba1f0eef fix(fabric): unhinged fix to duplicate config blocks because this is entirely spaget (#229) 2026-06-29 18:02:15 +00:00
Antoine Lecompte f78ba50ebc fix(fabric): fix fabric deployment (#228) 2026-06-29 13:35:19 -04:00
Antoine Lecompte 745bf5970e feat(claude): first CLAUDE.md (#226) 2026-06-29 17:34:29 +00:00
Antoine Lecompte b723285626 fix(workflow): gate ansible workflows (#227) 2026-06-29 13:34:14 -04:00
Antoine Lecompte 4cf9ac0a83 fix(ci): apply prettier formatting (#225) 2026-06-29 17:11:41 +00:00
Antoine Lecompte aeee19e336 feat(bgp): stand up bgp (#224)
* feat(bgp): stand up bgp

* eergh

* firewall

* fix
2026-06-29 14:50:43 +00:00
Antoine Lecompte 48c6220894 fix(workflow): missing working directory (#223) 2026-06-29 13:04:36 +00:00
Antoine Lecompte c6985d902c feat(all): introduce partition/region/ceph-cluster model across the stack (#222)
* feat: introduce partition/region/ceph-cluster model across the stack

Formalize partition -> region -> {one k8s cluster, many ceph clusters} and
thread it through every layer plus a new yuctl ops CLI.

- tf: deployment/<partition>/<region>/<stack> layout; terragrunt path-parse +
  state key yucca/<partition>/<region>/<stack>; per-region region.hcl (role,
  site_id, datacenter, provider_code, domain); env->partition / site->region
  renames (NetBird object names byte-identical); standardized per-stack
  `discovery` output contract (secrets as op:// refs).
- k8s: clusters/<partition>/<region>/ (staging/austin, prod/htz-fsn1, dev/local);
  role-based kustomize components (primary/secondary); hybrid cluster-settings
  (TF-rendered identity + human fragment); dev-mirror folded into dev/local;
  charts regrouped into charts/{apps,platform,lib,dev}.
- ci: infra.yml partition/region discovery matrix; partition-keyed path filters;
  <partition>-<region> environment gates; image-versions path moves.
- ansible: inventories under <partition>-<region>/<cluster>.
- yuctl: Go/cobra CLI reading the discovery contract from TF state.
- Retire the sietch-talos libvirt VM cluster (dev@local is the k3d cluster);
  ceph inventory_dirname -> <partition>-<region>/<cluster>.

Verified: mise k8s:validate green (3 clusters); yuctl go build/vet; tofu
validate pre-merge (all 9 stacks). Live-staging state migration NOT run.

* fix typo

* commit
2026-06-29 08:40:29 -04:00
Antoine Lecompte 54c410f59c fix(netbird): comment in the things (#221)
* fix(netbird): comment in the things

* remove netbird namespace
2026-06-27 12:01:58 +00:00
Antoine Lecompte a6235d77ed fix(netbird): comment out lines (#220)
* feat(netbird): k8s

* fix(netbird): comment out things temp
2026-06-26 19:48:41 +00:00
Antoine Lecompte 4033cc2d97 fix(netbird): add netbird network to trusted list (#219) 2026-06-26 19:42:45 +00:00
Antoine Lecompte 0ec09f022a feat(netbird): k8s (#218) 2026-06-26 19:41:32 +00:00
Antoine Lecompte e039026cbe feat(netbird-ansible): better subnet routers (#217) 2026-06-26 19:27:28 +00:00
Antoine Lecompte 00421f33ec fix(netbird): take explicit group-name overrides verbatim(#216) 2026-06-26 19:02:30 +00:00
Antoine Lecompte 481c5e920a feat(yucca): add full e2e mgmt provisioning maybe (#182)
* feat(yucca): add full e2e mgmt provisioning maybe

* moar !

* prefer tailscale over public ip if availbale

* ignore files

* fix

* more progress
2026-06-26 14:45:20 -04:00