7.1 KiB
Grafana dashboards and alerts
o11y's Grafana renders the dashboards and alert rules that each project ships to it. Nothing is clicked in the UI and left to rot: every dashboard and alert is a grafana-operator CR, delivered one of two ways, and grouped into a per-project folder.
Folders are the project boundary
Each project gets a Grafana folder named for it, and both its dashboards and its alert rule groups file under that folder:
| Folder | Owner | Delivered by |
|---|---|---|
yucca |
the yucca cluster/product | yucca's signed OCI bundle (Model A) |
o11y |
this cluster's own dashboards/alerts | authored in this repo (Model B) |
Add a project, add a folder. That folder is the unit you scope dashboards, alerts, and (eventually) permissions to.
Tags cut across folders
A folder files a dashboard under exactly one project; a tag is the orthogonal axis - signal type (metrics, logs), layer (infra, k8s, app) - that Grafana's dashboard browser filters on across every folder at once. Grafana stores tags inside the dashboard JSON model and grafana-operator exposes no field to inject them, so tags can only be set where the JSON is authored:
- Bundle (Model A) and first-party (Model B) dashboards you write: set
tagsin the dashboard JSON before shipping. Folder is your project; tags are the signal/layer cross-cut. - Dashboards pulled from grafana.com or a raw URL (
spec.grafanaCom/spec.url, e.g. the envoy and cnpg dashboards): they carry whatever tags upstream set. o11y cannot add or normalize them without vendoring the JSON inline, which forfeits the live reference andresyncPeriodauto-updates - so leave them as-is.
Model A: a project ships a signed OCI manifest bundle
This is how yucca ships (immich-app/yucca#315, see that repo's o11y/README.md). The project's CI renders each dashboard into a self-contained GrafanaDashboard CR (JSON embedded as spec.gzipJson) plus any GrafanaAlertRuleGroup CRs and a GrafanaFolder, pushes them as one signed OCI artifact (flux push artifact + cosign keyless), and o11y consumes the whole thing with a single Flux OCIRepository + Kustomization. New dashboards/alerts flow automatically on the next artifact.
o11y's consumer side lives once in kubernetes/apps/base/yucca-o11y/ and is pulled into each env's o11y overlay:
apiVersion: source.toolkit.fluxcd.io/v1
kind: OCIRepository
metadata: { name: yucca-o11y, namespace: flux-system }
spec:
interval: 1m
url: oci://ghcr.io/immich-app/yucca/o11y-manifests
ref:
tag: main # tracks every merge; no digest, or auto-updates stop
verify: # gate on the CI cosign signature
provider: cosign
matchOIDCIdentity:
- issuer: "^https://token\\.actions\\.githubusercontent\\.com$"
subject: "^https://github\\.com/immich-app/yucca/\\.github/workflows/o11y\\.yml@refs/(heads/main|tags/v[^@]+)$"
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata: { name: yucca-o11y, namespace: flux-system }
spec:
interval: 10m
sourceRef: { kind: OCIRepository, name: yucca-o11y }
path: ./
prune: true
targetNamespace: o11y
dependsOn: [{ name: grafana-operator }]
The bundle's CRs carry sane defaults (instanceSelector: {dashboards: grafana}, folderRef: <project>, resyncPeriod), so o11y applies them as-is. The GHCR package must be public (or set secretRef on the OCIRepository), and source-controller needs sigstore egress for verify.
Model B: authored in this repo (o11y's own)
For this cluster's own dashboards and alerts, they live under kubernetes/apps/base/grafana/app/ and deploy with the grafana Flux Kustomization:
- Dashboards -
base/grafana/app/dashboards/*.yaml, oneGrafanaDashboardper file,folderRef: o11y. Source the JSON however fits:spec.urlto a raw/grafana.com dashboard (the envoy and cnpg dashboards),spec.gzipJson, etc. Map dashboard__inputs(e.g.DS_PROMETHEUS) todatasourceName: VictoriaMetrics. - Alerts -
base/grafana/app/alerts-*.yaml, aGrafanaAlertRuleGroupwithfolderRef: o11y.
Alerting
Contact points. A GrafanaContactPoint per destination. Secrets (like a Discord webhook) come from a Secret via receivers[].valuesFrom, populated by an ExternalSecret from 1Password - never in git. Contact points live in the shared Grafana Postgres, so with the HA replica gossip cluster a firing alert notifies once, not once per replica.
Routing. One GrafanaNotificationPolicy routes by the grafana_folder label - which Grafana adds automatically from the folder each rule files into, so routing follows the folder (the project boundary) with no label to maintain:
route:
receiver: discord # default / catch-all
routes:
- object_matchers: [["grafana_folder", "=", "yucca"]]
receiver: discord
- object_matchers: [["grafana_folder", "=", "o11y"]]
receiver: discord # point at an o11y-specific contact point when one exists
So routing follows the folder automatically - no per-rule label to set or keep in sync. (Existing rules still carry a project label; it is now legacy and unused for routing.) Notifications additionally group by cluster (alongside grafana_folder and alertname), so the same rule firing in two clusters arrives as two grouped notifications rather than one blended message.
Alert rule anatomy. A GrafanaAlertRuleGroup (folderRef: <project>, an interval) with rules[]; each rule is a query stage on the VictoriaMetrics datasource (uid VictoriaMetrics) feeding a __expr__ threshold stage, plus labels (at least severity) and annotations. See base/grafana/app/alerts-o11y.yaml for the pattern (a heartbeat plus target-down and ingestion-stalled rules). Rules that span clusters aggregate by (cluster) so each cluster raises its own instance and carries its cluster label into notification grouping; store-local rules (the heartbeat, ingestion-stalled) don't.
If you ship metrics to this cluster and want dashboards/alerts
- Pick a delivery model: Model A (recommended for a separate repo/cluster - you own a signed bundle, o11y adds one OCIRepository) or Model B (PR the CRs into
base/grafana). - Everything you ship files under your project's folder; ask for one if it does not exist.
- Routing follows your folder automatically; add a route matching your
grafana_folder(and, if you want your own channel, a contact point). - Dashboards use a
$datasourcevariable and mapDS_PROMETHEUStoVictoriaMetrics; alerts query theVictoriaMetricsdatasource. - Tag dashboards by signal/layer in the JSON (
metrics,logs,infra, ...) so they stay filterable across folders (see Tags). Alerts that compare across clusters aggregateby (cluster); set theclusterlabel on your series so per-cluster alerting works.
How updates flow
- Model A: edit in the source repo, merge to
main-> CI pushes the:mainartifact (signed) -> o11y's OCIRepository picks it up within itsinterval-> Kustomization applies -> grafana-operator syncs. Roughly a minute end to end. - Model B: PR the CR change here -> merge -> Flux reconciles
base/grafana.