14 KiB
Grafana dashboards and alerts
o11y's Grafana renders the dashboards and alert rules that each project ships to it. Nothing is clicked in the UI and left to rot: every dashboard and alert is a grafana-operator CR, delivered one of two ways, and grouped into a per-project folder.
Folders are the project boundary
Each project gets a Grafana folder named for it, and both its dashboards and its alert rule groups file under that folder:
| Folder | Owner | Delivered by |
|---|---|---|
yucca |
the yucca cluster/product | yucca's signed OCI bundle (Model A) |
o11y |
this cluster's own dashboards/alerts | authored in this repo (Model B) |
harbor |
the Harbor clusters (harbor-infra-prod/staging) | harbor-o11y's key-signed OCI bundle (Model A, anonymous-pull registry) |
fip |
the FUTO internal platform cluster (azad) | futo-internal-platform's signed OCI bundle (Model A) |
Add a project, add a folder. That folder is the unit you scope dashboards, alerts, and (eventually) permissions to.
Tags cut across folders
A folder files a dashboard under exactly one project; a tag is the orthogonal axis - signal type (metrics, logs), layer (infra, k8s, app) - that Grafana's dashboard browser filters on across every folder at once. Grafana stores tags inside the dashboard JSON model and grafana-operator exposes no field to inject them, so tags can only be set where the JSON is authored:
- Bundle (Model A) and first-party (Model B) dashboards you write: set
tagsin the dashboard JSON before shipping. Folder is your project; tags are the signal/layer cross-cut. - Dashboards pulled from grafana.com or a raw URL (
spec.grafanaCom/spec.url, e.g. the grafana-operator dashboard): they carry whatever tags upstream set. o11y cannot add or normalize them without vendoring the JSON inline, which forfeits the live reference andresyncPeriodauto-updates - so leave them as-is.
Model A: a project ships a signed OCI manifest bundle
This is how yucca ships (immich-app/yucca#315, see that repo's o11y/README.md). The project's CI renders each dashboard into a self-contained GrafanaDashboard CR (JSON embedded as spec.gzipJson) plus any GrafanaAlertRuleGroup CRs and a GrafanaFolder, pushes them as one signed OCI artifact (flux push artifact + cosign keyless), and o11y consumes the whole thing with a single Flux OCIRepository + Kustomization. New dashboards/alerts flow automatically on the next artifact.
o11y's consumer side lives once in kubernetes/apps/base/tenants/yucca/bundle.yaml, one file per tenant listed by base/tenants/kustomization.yaml, which each env's o11y overlay pulls in:
apiVersion: source.toolkit.fluxcd.io/v1
kind: OCIRepository
metadata: { name: yucca-o11y, namespace: flux-system }
spec:
interval: 1m
url: oci://ghcr.io/immich-app/yucca/o11y-manifests
ref:
tag: main # tracks every merge; no digest, or auto-updates stop
verify: # gate on the CI cosign signature
provider: cosign
matchOIDCIdentity:
- issuer: "^https://token\\.actions\\.githubusercontent\\.com$"
subject: "^https://github\\.com/immich-app/yucca/\\.github/workflows/o11y\\.yml@refs/(heads/main|tags/v[^@]+)$"
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata: { name: yucca-o11y, namespace: flux-system }
spec:
interval: 10m
sourceRef: { kind: OCIRepository, name: yucca-o11y }
path: ./
prune: true
targetNamespace: o11y
dependsOn: [{ name: grafana-operator }]
The bundle's CRs carry sane defaults (instanceSelector: {dashboards: grafana}, folderRef: <project>, resyncPeriod), so o11y applies them as-is. The GHCR package must be public (or set secretRef on the OCIRepository), and source-controller needs sigstore egress for verify.
Model A from a private registry
yucca's bundle is a public GHCR package, so its OCIRepository needs no
credential. A project publishing to its own forge — fmeet ships to
gitlab.futo.org:5050, where its GitLab project lives — will be pulling from a
private one, and there is no way around it: GitLab's
container_registry_access_level takes only disabled / private / enabled,
where enabled still means "everyone with access". Of all the project
feature access levels, only pages_access_level accepts public, so a private
project cannot expose a publicly-pullable registry. An anonymous pull scope is
refused at /jwt/auth outright.
Two additions, then:
- A read-only pull credential. A GitLab project deploy token scoped to
read_registryis the least it can be — pull on that one project's registry, nothing else, revocable without touching an account. Its username and password go into 1Password as two separate manual secrets, each holding its value in the item'spasswordfield — the manual-secrets module only creates password-category items, so a two-part credential is split rather than packed into extra fields on one item. - An ExternalSecret rendering it as a
dockerconfigjson, influx-system, because that is where theOCIRepositorylives and asecretRefresolves in its own namespace. It pulls the two items by name and templates the docker config around them. See theExternalSecretinbase/tenants/fmeet/bundle.yaml.
Both items go in the global o11y_tf vault, read through the onepassword
store. Not shared_tf — that is for credentials more than one project
consumes, and this cluster is the only thing that pulls the bundle; and not a
per-environment o11y vault, because there is one fmeet project registry and so
one token, the same in every environment. They are declared in core-infra-tf's
o11y-manual-secrets module and filled in by hand in o11y_tf_manual.
That onepassword ClusterSecretStore is added here: the cluster reached
shared_tf and the per-environment o11y vaults, but nothing had yet needed the
global o11y one, whose other items are all consumed by terragrunt rather than
from inside the cluster.
The one detail that bites: the key under auths must match the
OCIRepository's pull host exactly, port and all
(gitlab.futo.org:5050, not gitlab.futo.org). A mismatch surfaces as an
authentication failure rather than as anything pointing at the cause.
An anonymously pullable registry sidesteps all of this, so the
OCIRepository needs no secretRef (harbor-o11y does this: its bundle lives
at registry.futo.org/harbor/o11y-manifests, where harbor/ allows
anonymous reads).
verify: is also omitted for such a bundle unless the publisher signs with a
key pair, as harbor-o11y does: its CI signs the digest with a key its
infrastructure terraform mints, the public half is committed in that repo as
cosign.pub, and base/tenants/harbor carries it as the harbor-o11y-cosign
Secret referenced from verify.secretRef. Keyless cosign mints its certificate from public Fulcio against the
CI's OIDC identity, and Fulcio accepts gitlab.com but not a self-hosted
forge — so the keyless block above cannot simply be copied across.
Model B: authored in this repo (o11y's own)
For this cluster's own dashboards and alerts, they live under kubernetes/apps/base/grafana/app/ and deploy with the grafana Flux Kustomization:
- Dashboards -
base/grafana/app/dashboards/*.yaml, oneGrafanaDashboardper file,folderRef: o11y. Source the JSON however fits:spec.grafanaComorspec.urlto a grafana.com or raw dashboard (the grafana-operator dashboard),spec.gzipJson, etc. Map dashboard__inputs(e.g.DS_PROMETHEUS) todatasourceName: VictoriaMetrics. - Alerts -
base/grafana/app/alerts-*.yaml, aGrafanaAlertRuleGroupwithfolderRef: o11y.
Shared dashboards
Cluster-generic boards (Kubernetes views and system, node exporter, vmagent, Cilium, Flux, Envoy Gateway, CloudNativePG) live once in the Shared folder rather than in every tenant bundle. They are rendered from their upstream sources by the VictoriaMetrics sync-job in generate mode, filtered on a multi-select $cluster variable and pinned to the VictoriaMetrics Fleet datasource, so one copy serves every cluster; see o11y/README.md for how to add one. A tenant bundle should not ship its own copy of a board that exists in Shared.
Alerting
Contact points. A GrafanaContactPoint per destination. Alerts go to Rootly: one webhook contact point per project (rootly-<project> in base/grafana/app/contactpoint-rootly-alerts.yaml), each posting to that project's Rootly alert source with the source's own secret, plus rootly-heartbeat for the dead man's switch. The alert sources, per-project services and their credentials are managed by the deployment/modules/rootly/cluster Terraform module, which writes each source's webhook URL into the env 1Password vault as ROOTLY_ALERTS_<PROJECT>_URL; the secret rides in that URL's secret query parameter because Rootly's Grafana endpoint reads it nowhere else. An ExternalSecret materializes it into the Secret the contact point reads via receivers[].valuesFrom - never in git. A resolved Grafana notification resolves the Rootly alert. Rootly derives alert urgency from the severity label (critical -> High, warning -> Medium, otherwise Low) and, until escalation policies exist, forwards fired and resolved alerts to Discord from its own cloud. Contact points live in the shared Grafana Postgres, so with the HA replica gossip cluster a firing alert notifies once, not once per replica.
Rootly's Grafana integration. Besides the alert sources, Rootly has an account-level Grafana integration (Integrations > Grafana) that takes this cluster's Grafana URL and an Admin service account token; it is what lets Rootly deep-link rules and snapshot dashboards into incidents. Rootly exposes no API or Terraform surface for it, so it is installed by hand once per env. The deployment/modules/grafana/cluster module owns the rootly service account and its token and writes the token to the env vault as ROOTLY_GRAFANA_SERVICE_ACCOUNT_TOKEN; paste that and https://grafana.<env domain> into Rootly. It is a separate module from rootly/cluster so the latter stays plannable while Grafana is down.
Routing. One GrafanaNotificationPolicy routes by the grafana_folder label - which Grafana adds automatically from the folder each rule files into, so routing follows the folder (the project boundary) with no label to maintain:
route:
receiver: rootly-o11y # default / catch-all for folders without a source of their own
routes:
- object_matchers: [["grafana_folder", "=", "yucca"]]
receiver: rootly-yucca
- object_matchers: [["grafana_folder", "=", "o11y"]]
receiver: rootly-o11y
So routing follows the folder automatically - no per-rule label to set or keep in sync. (Existing rules still carry a project rule label; it is legacy and unused for routing — distinct from the project series label every shipper stamps, see the shipping guide.) Notifications additionally group by cluster (alongside grafana_folder and alertname), so the same rule firing in two clusters arrives as two grouped notifications rather than one blended message.
Datasources and tenants. Two Prometheus-type datasources front the same store. VictoriaMetrics (uid VictoriaMetrics, the default) carries this cluster's own tenant in its URL and goes through the self-select vmauth, which routes only that path, so anything that does not pick a datasource explicitly sees only o11y's series. VictoriaMetrics Fleet (uid VictoriaMetricsFleet) reads the multitenant endpoint across every tenant, including tenant 0 where unmigrated remotes still land; it is the explicit opt-in for cross-cluster dashboards and rules (alerts-fleet.yaml). Tenancy is carried by the datasource URL, never by dashboard JSON or rule queries: both keep filtering and grouping on the cluster label, which survives the tenant split unchanged.
Alert rule anatomy. A GrafanaAlertRuleGroup (folderRef: <project>, an interval) with rules[]; each rule is a query stage on the VictoriaMetrics datasource (uid VictoriaMetrics) feeding a __expr__ threshold stage, plus labels (at least severity) and annotations. See base/grafana/app/alerts-o11y.yaml for the pattern (a heartbeat plus target-down and ingestion-stalled rules). Rules that span clusters aggregate by (cluster) so each cluster raises its own instance and carries its cluster label into notification grouping; store-local rules (the heartbeat, ingestion-stalled) don't.
If you ship metrics to this cluster and want dashboards/alerts
- Pick a delivery model: Model A (recommended for a separate repo/cluster - you own a signed bundle, o11y adds one OCIRepository) or Model B (PR the CRs into
base/grafana). - Everything you ship files under your project's folder; ask for one if it does not exist.
- Routing follows your folder automatically; ask for your project to be added to the Rootly module's
projectslist, which gives you a Rootly alert source, a service, and therootly-<project>contact point and route inbase/grafana/app. - Dashboards use a
$datasourcevariable and mapDS_PROMETHEUStoVictoriaMetrics; alerts query theVictoriaMetricsdatasource. Anything that must see other clusters' series usesVictoriaMetrics Fleetinstead (see Datasources and tenants). - Tag dashboards by signal/layer in the JSON (
metrics,logs,infra, ...) so they stay filterable across folders (see Tags). Alerts that compare across clusters aggregateby (cluster); stamp the five identity labels on your series (see the shipping guide) so per-cluster alerting works.
How updates flow
- Model A: edit in the source repo, merge to
main-> CI pushes the:mainartifact (signed) -> o11y's OCIRepository picks it up within itsinterval-> Kustomization applies -> grafana-operator syncs. Roughly a minute end to end. - Model B: PR the CR change here -> merge -> Flux reconciles
base/grafana.