Files
yucca-o11y/docs/06-dashboards-and-alerts-guide.md
T

14 KiB

Grafana dashboards and alerts

o11y's Grafana renders the dashboards and alert rules that each project ships to it. Nothing is clicked in the UI and left to rot: every dashboard and alert is a grafana-operator CR, delivered one of two ways, and grouped into a per-project folder.

Folders are the project boundary

Each project gets a Grafana folder named for it, and both its dashboards and its alert rule groups file under that folder:

Folder Owner Delivered by
yucca the yucca cluster/product yucca's signed OCI bundle (Model A)
o11y this cluster's own dashboards/alerts authored in this repo (Model B)
harbor the Harbor clusters (harbor-infra-prod/staging) harbor-o11y's key-signed OCI bundle (Model A, anonymous-pull registry)
fip the FUTO internal platform cluster (azad) futo-internal-platform's signed OCI bundle (Model A)

Add a project, add a folder. That folder is the unit you scope dashboards, alerts, and (eventually) permissions to.

Tags cut across folders

A folder files a dashboard under exactly one project; a tag is the orthogonal axis - signal type (metrics, logs), layer (infra, k8s, app) - that Grafana's dashboard browser filters on across every folder at once. Grafana stores tags inside the dashboard JSON model and grafana-operator exposes no field to inject them, so tags can only be set where the JSON is authored:

  • Bundle (Model A) and first-party (Model B) dashboards you write: set tags in the dashboard JSON before shipping. Folder is your project; tags are the signal/layer cross-cut.
  • Dashboards pulled from grafana.com or a raw URL (spec.grafanaCom / spec.url, e.g. the grafana-operator dashboard): they carry whatever tags upstream set. o11y cannot add or normalize them without vendoring the JSON inline, which forfeits the live reference and resyncPeriod auto-updates - so leave them as-is.

Model A: a project ships a signed OCI manifest bundle

This is how yucca ships (immich-app/yucca#315, see that repo's o11y/README.md). The project's CI renders each dashboard into a self-contained GrafanaDashboard CR (JSON embedded as spec.gzipJson) plus any GrafanaAlertRuleGroup CRs and a GrafanaFolder, pushes them as one signed OCI artifact (flux push artifact + cosign keyless), and o11y consumes the whole thing with a single Flux OCIRepository + Kustomization. New dashboards/alerts flow automatically on the next artifact.

o11y's consumer side lives once in kubernetes/apps/base/tenants/yucca/bundle.yaml, one file per tenant listed by base/tenants/kustomization.yaml, which each env's o11y overlay pulls in:

apiVersion: source.toolkit.fluxcd.io/v1
kind: OCIRepository
metadata: { name: yucca-o11y, namespace: flux-system }
spec:
  interval: 1m
  url: oci://ghcr.io/immich-app/yucca/o11y-manifests
  ref:
    tag: main            # tracks every merge; no digest, or auto-updates stop
  verify:                # gate on the CI cosign signature
    provider: cosign
    matchOIDCIdentity:
      - issuer: "^https://token\\.actions\\.githubusercontent\\.com$"
        subject: "^https://github\\.com/immich-app/yucca/\\.github/workflows/o11y\\.yml@refs/(heads/main|tags/v[^@]+)$"
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata: { name: yucca-o11y, namespace: flux-system }
spec:
  interval: 10m
  sourceRef: { kind: OCIRepository, name: yucca-o11y }
  path: ./
  prune: true
  targetNamespace: o11y
  dependsOn: [{ name: grafana-operator }]

The bundle's CRs carry sane defaults (instanceSelector: {dashboards: grafana}, folderRef: <project>, resyncPeriod), so o11y applies them as-is. The GHCR package must be public (or set secretRef on the OCIRepository), and source-controller needs sigstore egress for verify.

Model A from a private registry

yucca's bundle is a public GHCR package, so its OCIRepository needs no credential. A project publishing to its own forge — fmeet ships to gitlab.futo.org:5050, where its GitLab project lives — will be pulling from a private one, and there is no way around it: GitLab's container_registry_access_level takes only disabled / private / enabled, where enabled still means "everyone with access". Of all the project feature access levels, only pages_access_level accepts public, so a private project cannot expose a publicly-pullable registry. An anonymous pull scope is refused at /jwt/auth outright.

Two additions, then:

  • A read-only pull credential. A GitLab project deploy token scoped to read_registry is the least it can be — pull on that one project's registry, nothing else, revocable without touching an account. Its username and password go into 1Password as two separate manual secrets, each holding its value in the item's password field — the manual-secrets module only creates password-category items, so a two-part credential is split rather than packed into extra fields on one item.
  • An ExternalSecret rendering it as a dockerconfigjson, in flux-system, because that is where the OCIRepository lives and a secretRef resolves in its own namespace. It pulls the two items by name and templates the docker config around them. See the ExternalSecret in base/tenants/fmeet/bundle.yaml.

Both items go in the global o11y_tf vault, read through the onepassword store. Not shared_tf — that is for credentials more than one project consumes, and this cluster is the only thing that pulls the bundle; and not a per-environment o11y vault, because there is one fmeet project registry and so one token, the same in every environment. They are declared in core-infra-tf's o11y-manual-secrets module and filled in by hand in o11y_tf_manual.

That onepassword ClusterSecretStore is added here: the cluster reached shared_tf and the per-environment o11y vaults, but nothing had yet needed the global o11y one, whose other items are all consumed by terragrunt rather than from inside the cluster.

The one detail that bites: the key under auths must match the OCIRepository's pull host exactly, port and all (gitlab.futo.org:5050, not gitlab.futo.org). A mismatch surfaces as an authentication failure rather than as anything pointing at the cause.

An anonymously pullable registry sidesteps all of this, so the OCIRepository needs no secretRef (harbor-o11y does this: its bundle lives at registry.futo.org/harbor/o11y-manifests, where harbor/ allows anonymous reads).

verify: is also omitted for such a bundle unless the publisher signs with a key pair, as harbor-o11y does: its CI signs the digest with a key its infrastructure terraform mints, the public half is committed in that repo as cosign.pub, and base/tenants/harbor carries it as the harbor-o11y-cosign Secret referenced from verify.secretRef. Keyless cosign mints its certificate from public Fulcio against the CI's OIDC identity, and Fulcio accepts gitlab.com but not a self-hosted forge — so the keyless block above cannot simply be copied across.

Model B: authored in this repo (o11y's own)

For this cluster's own dashboards and alerts, they live under kubernetes/apps/base/grafana/app/ and deploy with the grafana Flux Kustomization:

  • Dashboards - base/grafana/app/dashboards/*.yaml, one GrafanaDashboard per file, folderRef: o11y. Source the JSON however fits: spec.grafanaCom or spec.url to a grafana.com or raw dashboard (the grafana-operator dashboard), spec.gzipJson, etc. Map dashboard __inputs (e.g. DS_PROMETHEUS) to datasourceName: VictoriaMetrics.
  • Alerts - base/grafana/app/alerts-*.yaml, a GrafanaAlertRuleGroup with folderRef: o11y.

Shared dashboards

Cluster-generic boards (Kubernetes views and system, node exporter, vmagent, Cilium, Flux, Envoy Gateway, CloudNativePG) live once in the Shared folder rather than in every tenant bundle. They are rendered from their upstream sources by the VictoriaMetrics sync-job in generate mode, filtered on a multi-select $cluster variable and pinned to the VictoriaMetrics Fleet datasource, so one copy serves every cluster; see o11y/README.md for how to add one. A tenant bundle should not ship its own copy of a board that exists in Shared.

Alerting

Contact points. A GrafanaContactPoint per destination. Alerts go to Rootly: one webhook contact point per project (rootly-<project> in base/grafana/app/contactpoint-rootly-alerts.yaml), each posting to that project's Rootly alert source with the source's own bearer secret, plus rootly-heartbeat for the dead man's switch. The alert sources, per-project services and their credentials are managed by the deployment/modules/rootly/cluster Terraform module, which writes each source's URL and secret into the env 1Password vault (ROOTLY_ALERTS_<PROJECT>_URL / _SECRET); an ExternalSecret materializes them into the Secret the contact point reads via receivers[].valuesFrom - never in git. Rootly derives alert urgency from the severity label (critical -> High, warning -> Medium, otherwise Low) and, until escalation policies exist, forwards fired and resolved alerts to Discord from its own cloud. Contact points live in the shared Grafana Postgres, so with the HA replica gossip cluster a firing alert notifies once, not once per replica.

Rootly's Grafana integration. Besides the alert sources, Rootly has an account-level Grafana integration (Integrations > Grafana) that takes this cluster's Grafana URL and an Admin service account token; it is what lets Rootly deep-link rules and snapshot dashboards into incidents. Rootly exposes no API or Terraform surface for it, so it is installed by hand once per env. The deployment/modules/grafana/cluster module owns the rootly service account and its token and writes the token to the env vault as ROOTLY_GRAFANA_SERVICE_ACCOUNT_TOKEN; paste that and https://grafana.<env domain> into Rootly. It is a separate module from rootly/cluster so the latter stays plannable while Grafana is down.

Routing. One GrafanaNotificationPolicy routes by the grafana_folder label - which Grafana adds automatically from the folder each rule files into, so routing follows the folder (the project boundary) with no label to maintain:

route:
  receiver: rootly-o11y       # default / catch-all for folders without a source of their own
  routes:
    - object_matchers: [["grafana_folder", "=", "yucca"]]
      receiver: rootly-yucca
    - object_matchers: [["grafana_folder", "=", "o11y"]]
      receiver: rootly-o11y

So routing follows the folder automatically - no per-rule label to set or keep in sync. (Existing rules still carry a project rule label; it is legacy and unused for routing — distinct from the project series label every shipper stamps, see the shipping guide.) Notifications additionally group by cluster (alongside grafana_folder and alertname), so the same rule firing in two clusters arrives as two grouped notifications rather than one blended message.

Datasources and tenants. Two Prometheus-type datasources front the same store. VictoriaMetrics (uid VictoriaMetrics, the default) carries this cluster's own tenant in its URL and goes through the self-select vmauth, which routes only that path, so anything that does not pick a datasource explicitly sees only o11y's series. VictoriaMetrics Fleet (uid VictoriaMetricsFleet) reads the multitenant endpoint across every tenant, including tenant 0 where unmigrated remotes still land; it is the explicit opt-in for cross-cluster dashboards and rules (alerts-fleet.yaml). Tenancy is carried by the datasource URL, never by dashboard JSON or rule queries: both keep filtering and grouping on the cluster label, which survives the tenant split unchanged.

Alert rule anatomy. A GrafanaAlertRuleGroup (folderRef: <project>, an interval) with rules[]; each rule is a query stage on the VictoriaMetrics datasource (uid VictoriaMetrics) feeding a __expr__ threshold stage, plus labels (at least severity) and annotations. See base/grafana/app/alerts-o11y.yaml for the pattern (a heartbeat plus target-down and ingestion-stalled rules). Rules that span clusters aggregate by (cluster) so each cluster raises its own instance and carries its cluster label into notification grouping; store-local rules (the heartbeat, ingestion-stalled) don't.

If you ship metrics to this cluster and want dashboards/alerts

  1. Pick a delivery model: Model A (recommended for a separate repo/cluster - you own a signed bundle, o11y adds one OCIRepository) or Model B (PR the CRs into base/grafana).
  2. Everything you ship files under your project's folder; ask for one if it does not exist.
  3. Routing follows your folder automatically; ask for your project to be added to the Rootly module's projects list, which gives you a Rootly alert source, a service, and the rootly-<project> contact point and route in base/grafana/app.
  4. Dashboards use a $datasource variable and map DS_PROMETHEUS to VictoriaMetrics; alerts query the VictoriaMetrics datasource. Anything that must see other clusters' series uses VictoriaMetrics Fleet instead (see Datasources and tenants).
  5. Tag dashboards by signal/layer in the JSON (metrics, logs, infra, ...) so they stay filterable across folders (see Tags). Alerts that compare across clusters aggregate by (cluster); stamp the five identity labels on your series (see the shipping guide) so per-cluster alerting works.

How updates flow

  • Model A: edit in the source repo, merge to main -> CI pushes the :main artifact (signed) -> o11y's OCIRepository picks it up within its interval -> Kustomization applies -> grafana-operator syncs. Roughly a minute end to end.
  • Model B: PR the CR change here -> merge -> Flux reconciles base/grafana.