mirror of
https://github.com/immich-app/yucca-o11y.git
synced 2026-09-30 13:23:23 +08:00
163 lines
14 KiB
Markdown
163 lines
14 KiB
Markdown
# Grafana dashboards and alerts
|
|
|
|
o11y's Grafana renders the dashboards and alert rules that each project ships to it. Nothing is clicked in the UI and left to rot: every dashboard and alert is a `grafana-operator` CR, delivered one of two ways, and grouped into a per-project folder.
|
|
|
|
## Folders are the project boundary
|
|
|
|
Each project gets a Grafana **folder** named for it, and both its dashboards and its alert rule groups file under that folder:
|
|
|
|
| Folder | Owner | Delivered by |
|
|
| --- | --- | --- |
|
|
| `yucca` | the yucca cluster/product | yucca's signed OCI bundle (Model A) |
|
|
| `o11y` | this cluster's own dashboards/alerts | authored in this repo (Model B) |
|
|
| `harbor` | the Harbor clusters (harbor-infra-prod/staging) | harbor-o11y's key-signed OCI bundle (Model A, anonymous-pull registry) |
|
|
| `fip` | the FUTO internal platform cluster (azad) | futo-internal-platform's signed OCI bundle (Model A) |
|
|
|
|
Add a project, add a folder. That folder is the unit you scope dashboards, alerts, and (eventually) permissions to.
|
|
|
|
## Tags cut across folders
|
|
|
|
A folder files a dashboard under exactly one project; a **tag** is the orthogonal axis - signal type (`metrics`, `logs`), layer (`infra`, `k8s`, `app`) - that Grafana's dashboard browser filters on across every folder at once. Grafana stores tags inside the dashboard JSON model and `grafana-operator` exposes no field to inject them, so tags can only be set where the JSON is authored:
|
|
|
|
- **Bundle (Model A) and first-party (Model B) dashboards you write:** set `tags` in the dashboard JSON before shipping. Folder is your project; tags are the signal/layer cross-cut.
|
|
- **Dashboards pulled from grafana.com or a raw URL** (`spec.grafanaCom` / `spec.url`, e.g. the grafana-operator dashboard): they carry whatever tags upstream set. o11y cannot add or normalize them without vendoring the JSON inline, which forfeits the live reference and `resyncPeriod` auto-updates - so leave them as-is.
|
|
|
|
## Model A: a project ships a signed OCI manifest bundle
|
|
|
|
This is how yucca ships (immich-app/yucca#315, see that repo's `o11y/README.md`). The project's CI renders each dashboard into a self-contained `GrafanaDashboard` CR (JSON embedded as `spec.gzipJson`) plus any `GrafanaAlertRuleGroup` CRs and a `GrafanaFolder`, pushes them as **one signed OCI artifact** (`flux push artifact` + cosign keyless), and o11y consumes the whole thing with a single Flux `OCIRepository` + `Kustomization`. New dashboards/alerts flow automatically on the next artifact.
|
|
|
|
o11y's consumer side lives once in `kubernetes/apps/base/tenants/yucca/bundle.yaml`, one file per tenant listed by `base/tenants/kustomization.yaml`, which each env's `o11y` overlay pulls in:
|
|
|
|
```yaml
|
|
apiVersion: source.toolkit.fluxcd.io/v1
|
|
kind: OCIRepository
|
|
metadata: { name: yucca-o11y, namespace: flux-system }
|
|
spec:
|
|
interval: 1m
|
|
url: oci://ghcr.io/immich-app/yucca/o11y-manifests
|
|
ref:
|
|
tag: main # tracks every merge; no digest, or auto-updates stop
|
|
verify: # gate on the CI cosign signature
|
|
provider: cosign
|
|
matchOIDCIdentity:
|
|
- issuer: "^https://token\\.actions\\.githubusercontent\\.com$"
|
|
subject: "^https://github\\.com/immich-app/yucca/\\.github/workflows/o11y\\.yml@refs/(heads/main|tags/v[^@]+)$"
|
|
---
|
|
apiVersion: kustomize.toolkit.fluxcd.io/v1
|
|
kind: Kustomization
|
|
metadata: { name: yucca-o11y, namespace: flux-system }
|
|
spec:
|
|
interval: 10m
|
|
sourceRef: { kind: OCIRepository, name: yucca-o11y }
|
|
path: ./
|
|
prune: true
|
|
targetNamespace: o11y
|
|
dependsOn: [{ name: grafana-operator }]
|
|
```
|
|
|
|
The bundle's CRs carry sane defaults (`instanceSelector: {dashboards: grafana}`, `folderRef: <project>`, `resyncPeriod`), so o11y applies them as-is. The GHCR package must be public (or set `secretRef` on the OCIRepository), and source-controller needs sigstore egress for `verify`.
|
|
|
|
### Model A from a private registry
|
|
|
|
yucca's bundle is a public GHCR package, so its `OCIRepository` needs no
|
|
credential. A project publishing to its own forge — fmeet ships to
|
|
`gitlab.futo.org:5050`, where its GitLab project lives — will be pulling from a
|
|
private one, and there is no way around it: GitLab's
|
|
`container_registry_access_level` takes only `disabled` / `private` / `enabled`,
|
|
where `enabled` still means "everyone **with access**". Of all the project
|
|
feature access levels, only `pages_access_level` accepts `public`, so a private
|
|
project cannot expose a publicly-pullable registry. An anonymous pull scope is
|
|
refused at `/jwt/auth` outright.
|
|
|
|
Two additions, then:
|
|
|
|
- **A read-only pull credential.** A GitLab *project deploy token* scoped to
|
|
`read_registry` is the least it can be — pull on that one project's registry,
|
|
nothing else, revocable without touching an account. Its username and password
|
|
go into 1Password as **two separate manual secrets**, each holding its value in
|
|
the item's `password` field — the manual-secrets module only creates
|
|
password-category items, so a two-part credential is split rather than packed
|
|
into extra fields on one item.
|
|
- **An ExternalSecret rendering it as a `dockerconfigjson`**, in `flux-system`,
|
|
because that is where the `OCIRepository` lives and a `secretRef` resolves in
|
|
its own namespace. It pulls the two items by name and templates the docker
|
|
config around them. See the `ExternalSecret` in `base/tenants/fmeet/bundle.yaml`.
|
|
|
|
Both items go in the global `o11y_tf` vault, read through the `onepassword`
|
|
store. Not `shared_tf` — that is for credentials more than one project
|
|
consumes, and this cluster is the only thing that pulls the bundle; and not a
|
|
per-environment o11y vault, because there is one fmeet project registry and so
|
|
one token, the same in every environment. They are declared in core-infra-tf's
|
|
`o11y-manual-secrets` module and filled in by hand in `o11y_tf_manual`.
|
|
|
|
That `onepassword` ClusterSecretStore is added here: the cluster reached
|
|
`shared_tf` and the per-environment o11y vaults, but nothing had yet needed the
|
|
global o11y one, whose other items are all consumed by terragrunt rather than
|
|
from inside the cluster.
|
|
|
|
The one detail that bites: the key under `auths` must match the
|
|
`OCIRepository`'s pull host **exactly**, port and all
|
|
(`gitlab.futo.org:5050`, not `gitlab.futo.org`). A mismatch surfaces as an
|
|
authentication failure rather than as anything pointing at the cause.
|
|
|
|
An anonymously pullable registry sidesteps all of this, so the
|
|
`OCIRepository` needs no `secretRef` (harbor-o11y does this: its bundle lives
|
|
at `registry.futo.org/harbor/o11y-manifests`, where `harbor/` allows
|
|
anonymous reads).
|
|
|
|
`verify:` is also omitted for such a bundle unless the publisher signs with a
|
|
key pair, as harbor-o11y does: its CI signs the digest with a key its
|
|
infrastructure terraform mints, the public half is committed in that repo as
|
|
`cosign.pub`, and `base/tenants/harbor` carries it as the `harbor-o11y-cosign`
|
|
Secret referenced from `verify.secretRef`. Keyless cosign mints its certificate from public Fulcio against the
|
|
CI's OIDC identity, and Fulcio accepts `gitlab.com` but not a self-hosted
|
|
forge — so the keyless block above cannot simply be copied across.
|
|
|
|
## Model B: authored in this repo (o11y's own)
|
|
|
|
For this cluster's own dashboards and alerts, they live under `kubernetes/apps/base/grafana/app/` and deploy with the grafana Flux Kustomization:
|
|
|
|
- **Dashboards** - `base/grafana/app/dashboards/*.yaml`, one `GrafanaDashboard` per file, `folderRef: o11y`. Source the JSON however fits: `spec.grafanaCom` or `spec.url` to a grafana.com or raw dashboard (the grafana-operator dashboard), `spec.gzipJson`, etc. Map dashboard `__inputs` (e.g. `DS_PROMETHEUS`) to `datasourceName: VictoriaMetrics`.
|
|
- **Alerts** - `base/grafana/app/alerts-*.yaml`, a `GrafanaAlertRuleGroup` with `folderRef: o11y`.
|
|
|
|
## Shared dashboards
|
|
|
|
Cluster-generic boards (Kubernetes views and system, node exporter, vmagent, Cilium, Flux, Envoy Gateway, CloudNativePG) live once in the **`Shared`** folder rather than in every tenant bundle. They are rendered from their upstream sources by the VictoriaMetrics sync-job in generate mode, filtered on a multi-select `$cluster` variable and pinned to the `VictoriaMetrics Fleet` datasource, so one copy serves every cluster; see [`o11y/README.md`](../o11y/README.md) for how to add one. A tenant bundle should not ship its own copy of a board that exists in `Shared`.
|
|
|
|
## Alerting
|
|
|
|
**Contact points.** A `GrafanaContactPoint` per destination. Alerts go to [Rootly](https://rootly.com): one webhook contact point per project (`rootly-<project>` in `base/grafana/app/contactpoint-rootly-alerts.yaml`), each posting to that project's Rootly alert source with the source's own bearer secret, plus `rootly-heartbeat` for the dead man's switch. The alert sources, per-project services and their credentials are managed by the `deployment/modules/rootly/cluster` Terraform module, which writes each source's URL and secret into the env 1Password vault (`ROOTLY_ALERTS_<PROJECT>_URL` / `_SECRET`); an ExternalSecret materializes them into the Secret the contact point reads via `receivers[].valuesFrom` - never in git. Rootly derives alert urgency from the `severity` label (`critical` -> High, `warning` -> Medium, otherwise Low) and, until escalation policies exist, forwards fired and resolved alerts to Discord from its own cloud. Contact points live in the shared Grafana Postgres, so with the HA replica gossip cluster a firing alert notifies **once**, not once per replica.
|
|
|
|
**Rootly's Grafana integration.** Besides the alert sources, Rootly has an account-level Grafana integration (Integrations > Grafana) that takes this cluster's Grafana URL and an Admin service account token; it is what lets Rootly deep-link rules and snapshot dashboards into incidents. Rootly exposes no API or Terraform surface for it, so it is installed by hand once per env. The `deployment/modules/grafana/cluster` module owns the `rootly` service account and its token and writes the token to the env vault as `ROOTLY_GRAFANA_SERVICE_ACCOUNT_TOKEN`; paste that and `https://grafana.<env domain>` into Rootly. It is a separate module from `rootly/cluster` so the latter stays plannable while Grafana is down.
|
|
|
|
**Routing.** One `GrafanaNotificationPolicy` routes by the **`grafana_folder`** label - which Grafana adds automatically from the folder each rule files into, so routing follows the folder (the project boundary) with no label to maintain:
|
|
|
|
```yaml
|
|
route:
|
|
receiver: rootly-o11y # default / catch-all for folders without a source of their own
|
|
routes:
|
|
- object_matchers: [["grafana_folder", "=", "yucca"]]
|
|
receiver: rootly-yucca
|
|
- object_matchers: [["grafana_folder", "=", "o11y"]]
|
|
receiver: rootly-o11y
|
|
```
|
|
|
|
So **routing follows the folder automatically** - no per-rule label to set or keep in sync. (Existing rules still carry a `project` *rule* label; it is legacy and unused for routing — distinct from the `project` *series* label every shipper stamps, see the [shipping guide](05-shipping-metrics-guide.md#labels).) Notifications additionally group by `cluster` (alongside `grafana_folder` and `alertname`), so the same rule firing in two clusters arrives as two grouped notifications rather than one blended message.
|
|
|
|
**Datasources and tenants.** Two Prometheus-type datasources front the same store. `VictoriaMetrics` (uid `VictoriaMetrics`, the default) carries this cluster's own tenant in its URL and goes through the `self-select` vmauth, which routes only that path, so anything that does not pick a datasource explicitly sees only o11y's series. `VictoriaMetrics Fleet` (uid `VictoriaMetricsFleet`) reads the multitenant endpoint across every tenant, including tenant 0 where unmigrated remotes still land; it is the explicit opt-in for cross-cluster dashboards and rules (`alerts-fleet.yaml`). Tenancy is carried by the datasource URL, never by dashboard JSON or rule queries: both keep filtering and grouping on the `cluster` label, which survives the tenant split unchanged.
|
|
|
|
**Alert rule anatomy.** A `GrafanaAlertRuleGroup` (`folderRef: <project>`, an `interval`) with `rules[]`; each rule is a query stage on the `VictoriaMetrics` datasource (uid `VictoriaMetrics`) feeding a `__expr__` threshold stage, plus `labels` (at least `severity`) and `annotations`. See `base/grafana/app/alerts-o11y.yaml` for the pattern (a heartbeat plus target-down and ingestion-stalled rules). Rules that span clusters aggregate `by (cluster)` so each cluster raises its own instance and carries its `cluster` label into notification grouping; store-local rules (the heartbeat, ingestion-stalled) don't.
|
|
|
|
## If you ship metrics to this cluster and want dashboards/alerts
|
|
|
|
1. Pick a delivery model: **Model A** (recommended for a separate repo/cluster - you own a signed bundle, o11y adds one OCIRepository) or **Model B** (PR the CRs into `base/grafana`).
|
|
2. Everything you ship files under **your project's folder**; ask for one if it does not exist.
|
|
3. Routing follows your folder automatically; ask for your project to be added to the Rootly module's `projects` list, which gives you a Rootly alert source, a service, and the `rootly-<project>` contact point and route in `base/grafana/app`.
|
|
4. Dashboards use a `$datasource` variable and map `DS_PROMETHEUS` to `VictoriaMetrics`; alerts query the `VictoriaMetrics` datasource. Anything that must see other clusters' series uses `VictoriaMetrics Fleet` instead (see Datasources and tenants).
|
|
5. Tag dashboards by signal/layer in the JSON (`metrics`, `logs`, `infra`, ...) so they stay filterable across folders (see Tags). Alerts that compare across clusters aggregate `by (cluster)`; stamp the five identity labels on your series (see the [shipping guide](05-shipping-metrics-guide.md#labels)) so per-cluster alerting works.
|
|
|
|
## How updates flow
|
|
|
|
- **Model A:** edit in the source repo, merge to `main` -> CI pushes the `:main` artifact (signed) -> o11y's OCIRepository picks it up within its `interval` -> Kustomization applies -> grafana-operator syncs. Roughly a minute end to end.
|
|
- **Model B:** PR the CR change here -> merge -> Flux reconciles `base/grafana`.
|