mirror of
https://github.com/immich-app/yucca-o11y.git
synced 2026-09-30 21:28:14 +08:00
fix(grafana): stop the o11y cluster shipping boards that Shared now owns (#324)
Signed-off-by: Devin Buhl <devin@buhl.casa>
This commit is contained in:
@@ -20,7 +20,7 @@ Add a project, add a folder. That folder is the unit you scope dashboards, alert
|
||||
A folder files a dashboard under exactly one project; a **tag** is the orthogonal axis - signal type (`metrics`, `logs`), layer (`infra`, `k8s`, `app`) - that Grafana's dashboard browser filters on across every folder at once. Grafana stores tags inside the dashboard JSON model and `grafana-operator` exposes no field to inject them, so tags can only be set where the JSON is authored:
|
||||
|
||||
- **Bundle (Model A) and first-party (Model B) dashboards you write:** set `tags` in the dashboard JSON before shipping. Folder is your project; tags are the signal/layer cross-cut.
|
||||
- **Dashboards pulled from grafana.com or a raw URL** (`spec.grafanaCom` / `spec.url`, e.g. the envoy and cnpg dashboards): they carry whatever tags upstream set. o11y cannot add or normalize them without vendoring the JSON inline, which forfeits the live reference and `resyncPeriod` auto-updates - so leave them as-is.
|
||||
- **Dashboards pulled from grafana.com or a raw URL** (`spec.grafanaCom` / `spec.url`, e.g. the grafana-operator dashboard): they carry whatever tags upstream set. o11y cannot add or normalize them without vendoring the JSON inline, which forfeits the live reference and `resyncPeriod` auto-updates - so leave them as-is.
|
||||
|
||||
## Model A: a project ships a signed OCI manifest bundle
|
||||
|
||||
@@ -117,12 +117,12 @@ forge — so the keyless block above cannot simply be copied across.
|
||||
|
||||
For this cluster's own dashboards and alerts, they live under `kubernetes/apps/base/grafana/app/` and deploy with the grafana Flux Kustomization:
|
||||
|
||||
- **Dashboards** - `base/grafana/app/dashboards/*.yaml`, one `GrafanaDashboard` per file, `folderRef: o11y`. Source the JSON however fits: `spec.url` to a raw/grafana.com dashboard (the envoy and cnpg dashboards), `spec.gzipJson`, etc. Map dashboard `__inputs` (e.g. `DS_PROMETHEUS`) to `datasourceName: VictoriaMetrics`.
|
||||
- **Dashboards** - `base/grafana/app/dashboards/*.yaml`, one `GrafanaDashboard` per file, `folderRef: o11y`. Source the JSON however fits: `spec.grafanaCom` or `spec.url` to a grafana.com or raw dashboard (the grafana-operator dashboard), `spec.gzipJson`, etc. Map dashboard `__inputs` (e.g. `DS_PROMETHEUS`) to `datasourceName: VictoriaMetrics`.
|
||||
- **Alerts** - `base/grafana/app/alerts-*.yaml`, a `GrafanaAlertRuleGroup` with `folderRef: o11y`.
|
||||
|
||||
## Shared dashboards
|
||||
|
||||
Cluster-generic boards (Kubernetes views and system, node exporter, vmagent, Cilium, Flux) live once in the **`Shared`** folder rather than in every tenant bundle. They are rendered from their upstream sources by the VictoriaMetrics sync-job in generate mode, filtered on a multi-select `$cluster` variable and pinned to the `VictoriaMetrics Fleet` datasource, so one copy serves every cluster; see [`o11y/README.md`](../o11y/README.md) for how to add one. A tenant bundle should not ship its own copy of a board that exists in `Shared`.
|
||||
Cluster-generic boards (Kubernetes views and system, node exporter, vmagent, Cilium, Flux, Envoy Gateway, CloudNativePG) live once in the **`Shared`** folder rather than in every tenant bundle. They are rendered from their upstream sources by the VictoriaMetrics sync-job in generate mode, filtered on a multi-select `$cluster` variable and pinned to the `VictoriaMetrics Fleet` datasource, so one copy serves every cluster; see [`o11y/README.md`](../o11y/README.md) for how to add one. A tenant bundle should not ship its own copy of a board that exists in `Shared`.
|
||||
|
||||
## Alerting
|
||||
|
||||
|
||||
@@ -21,70 +21,6 @@ spec:
|
||||
# yaml-language-server: $schema=https://k8s-schemas.home-operations.com/grafana.integreatly.org/grafanadashboard_v1beta1.json
|
||||
apiVersion: grafana.integreatly.org/v1beta1
|
||||
kind: GrafanaDashboard
|
||||
metadata:
|
||||
name: envoy-overview
|
||||
spec:
|
||||
folderRef: o11y
|
||||
instanceSelector:
|
||||
matchLabels:
|
||||
dashboards: grafana
|
||||
resyncPeriod: 10m
|
||||
datasources:
|
||||
- datasourceName: VictoriaMetrics
|
||||
inputName: DS_PROMETHEUS
|
||||
url: https://grafana.com/api/dashboards/24459/revisions/3/download
|
||||
---
|
||||
# yaml-language-server: $schema=https://k8s-schemas.home-operations.com/grafana.integreatly.org/grafanadashboard_v1beta1.json
|
||||
apiVersion: grafana.integreatly.org/v1beta1
|
||||
kind: GrafanaDashboard
|
||||
metadata:
|
||||
name: envoy-upstream
|
||||
spec:
|
||||
folderRef: o11y
|
||||
instanceSelector:
|
||||
matchLabels:
|
||||
dashboards: grafana
|
||||
resyncPeriod: 10m
|
||||
datasources:
|
||||
- datasourceName: VictoriaMetrics
|
||||
inputName: DS_PROMETHEUS
|
||||
url: https://grafana.com/api/dashboards/24457/revisions/4/download
|
||||
---
|
||||
# yaml-language-server: $schema=https://k8s-schemas.home-operations.com/grafana.integreatly.org/grafanadashboard_v1beta1.json
|
||||
apiVersion: grafana.integreatly.org/v1beta1
|
||||
kind: GrafanaDashboard
|
||||
metadata:
|
||||
name: envoy-downstream
|
||||
spec:
|
||||
folderRef: o11y
|
||||
instanceSelector:
|
||||
matchLabels:
|
||||
dashboards: grafana
|
||||
resyncPeriod: 10m
|
||||
datasources:
|
||||
- datasourceName: VictoriaMetrics
|
||||
inputName: DS_PROMETHEUS
|
||||
url: https://grafana.com/api/dashboards/24458/revisions/3/download
|
||||
---
|
||||
# yaml-language-server: $schema=https://k8s-schemas.home-operations.com/grafana.integreatly.org/grafanadashboard_v1beta1.json
|
||||
apiVersion: grafana.integreatly.org/v1beta1
|
||||
kind: GrafanaDashboard
|
||||
metadata:
|
||||
name: cnpg-cluster
|
||||
spec:
|
||||
folderRef: o11y
|
||||
instanceSelector:
|
||||
matchLabels:
|
||||
dashboards: grafana
|
||||
resyncPeriod: 10m
|
||||
datasources:
|
||||
- datasourceName: VictoriaMetrics
|
||||
inputName: DS_PROMETHEUS
|
||||
url: https://raw.githubusercontent.com/cloudnative-pg/grafana-dashboards/cluster-v0.0.5/charts/cluster/grafana-dashboard.json
|
||||
---
|
||||
# yaml-language-server: $schema=https://k8s-schemas.home-operations.com/grafana.integreatly.org/grafanadashboard_v1beta1.json
|
||||
apiVersion: grafana.integreatly.org/v1beta1
|
||||
kind: GrafanaDashboard
|
||||
metadata:
|
||||
name: grafana-operator
|
||||
spec:
|
||||
|
||||
@@ -10,13 +10,6 @@ spec:
|
||||
name: spegel
|
||||
interval: 1h
|
||||
values:
|
||||
grafanaDashboard:
|
||||
enabled: true
|
||||
mode: GrafanaOperator
|
||||
grafanaOperator:
|
||||
folder: o11y
|
||||
matchLabels:
|
||||
dashboards: grafana
|
||||
service:
|
||||
registry:
|
||||
hostPort: 29999
|
||||
|
||||
@@ -8,7 +8,6 @@ metadata:
|
||||
spec:
|
||||
dependsOn:
|
||||
- name: prometheus-operator-crds
|
||||
- name: grafana-operator
|
||||
healthChecks:
|
||||
- apiVersion: helm.toolkit.fluxcd.io/v2
|
||||
kind: HelmRelease
|
||||
|
||||
+1
-1
@@ -13,7 +13,7 @@ mise run //:o11y:check # fail if the committed render is stale
|
||||
|
||||
Both tasks first run `o11y:vendor`, which fetches boards that need a rewrite the sync-job cannot express and writes them to `vendor/` as local sources. Today that is CloudNativePG only: upstream uses `cluster` to mean the Postgres cluster, while the fleet keeps that name under `pg_cluster` and reserves `cluster` for the Kubernetes cluster, so the vendor step renames the label and variable before the sync-job adds the fleet's `$cluster`. A cluster's CNPG series must carry `pg_cluster` for the board to list its databases; the shipping guide's identity-label section shows the relabel rule, which o11y, azad and harbor apply.
|
||||
|
||||
The set is everything at least two clusters run: the dotdc Kubernetes views and system boards, the kube-prometheus mixin boards (kubelet, scheduler, controller manager, proxy, API server, compute resources, networking, persistent volumes, node exporter USE method), Node Exporter Full, etcd, vmagent, Cilium and Hubble, Flux, Spegel and CloudNativePG. Windows, AIX, macOS, Prometheus, Alertmanager and Grafana-overview boards from the mixin bundle are disabled. Documents are sorted by name so re-renders diff cleanly.
|
||||
The set is everything at least two clusters run: the dotdc Kubernetes views and system boards, the kube-prometheus mixin boards (kubelet, scheduler, controller manager, proxy, API server, compute resources, networking, persistent volumes, node exporter USE method), Node Exporter Full, etcd, vmagent, Cilium and Hubble, Flux, Spegel, Envoy Gateway and CloudNativePG. Windows, AIX, macOS, Prometheus, Alertmanager and Grafana-overview boards from the mixin bundle are disabled. Documents are sorted by name so re-renders diff cleanly.
|
||||
|
||||
Upstream sources are pinned where the upstream moves (Cilium by release tag, Flux by commit) and tracked at `master` where the VictoriaMetrics chart does the same.
|
||||
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -77,6 +77,9 @@ dashboards:
|
||||
- url: https://raw.githubusercontent.com/prometheus-operator/kube-prometheus/main/manifests/grafana-dashboardDefinitions.yaml
|
||||
- url: https://raw.githubusercontent.com/monitoring-mixins/website/master/assets/etcd/dashboards/etcd.json
|
||||
- url: /config/vendor/cloudnativepg.json
|
||||
- url: https://grafana.com/api/dashboards/24459/revisions/3/download
|
||||
- url: https://grafana.com/api/dashboards/24457/revisions/4/download
|
||||
- url: https://grafana.com/api/dashboards/24458/revisions/3/download
|
||||
rules:
|
||||
common: {}
|
||||
sources: []
|
||||
|
||||
Reference in New Issue
Block a user