mirror of
https://github.com/immich-app/yucca-o11y.git
synced 2026-09-30 13:23:23 +08:00
feat(victoria-metrics): assign o11y its own tenant derived from project and cluster labels (#319)
Signed-off-by: Devin Buhl <devin@buhl.casa>
This commit is contained in:
@@ -14,6 +14,7 @@ Built for geographic resilience with single-cluster operational simplicity: thre
|
||||
| [02 — Infrastructure architecture](./docs/02-infrastructure-architecture-guide.md) | OVH compute, hardware specs, vRack networking, IPLB ingress, cost |
|
||||
| [03 — Cluster architecture](./docs/03-cluster-architecture-guide.md) | Talos, Kubernetes, host firewall, and Flux GitOps |
|
||||
| [04 — Application architecture](./docs/04-application-architecture-guide.md) | Envoy ingress, VictoriaMetrics central store, Grafana, CloudNativePG, supporting operators |
|
||||
| [Runbook: tenant cutover](./docs/runbooks/vm-tenant-migration/README.md) | Moving a store's metrics from tenant 0 onto per-cluster tenants, staging first |
|
||||
|
||||
## Repository layout
|
||||
|
||||
|
||||
@@ -21,13 +21,13 @@ cert-manager issues short-lived ECDSA P-256 wildcard certificates with always-ro
|
||||
This cluster's VictoriaMetrics is the **central metrics store for all FUTO clusters**. Other Kubernetes clusters each run their own `vmagent` and remote-write into this cluster; it is the ingestion target plus the query and alerting brain for everyone.
|
||||
|
||||
* **Storage** — VMCluster mode with `replicationFactor=2`, `vmstorage` spread one-per-worker across the three DCs on `openebs-spare-disk`; retention is set per environment via `CLUSTER_VMETRICS_RETENTION` (30d staging, 120d production). The `vmstorage`, `vminsert`, and `vmselect` tiers scale independently.
|
||||
* **Local collection** — a `vmagent` (with a persistent disk buffer) scrapes this cluster and remote-writes to the local `vminsert`. It tags series with the cluster's identity.
|
||||
* **Local collection**: a `vmagent` (with a persistent disk buffer) scrapes this cluster and remote-writes to the local `vminsert`'s multitenant endpoint. It tags series with the cluster's identity, from which `vminsert` derives the tenant.
|
||||
* **Alerting** — `vmalert` evaluates rules; notifications are blackholed for now (no Alertmanager yet), so rules still evaluate and recording rules still write.
|
||||
* **Ingestion gateway** — a locked-down `vmauth` (no anonymous access, run as an HA pair) fronts `vminsert` and is exposed publicly at `vmauth.<CLUSTER_APP_DOMAIN>` through the Envoy Gateway and IPLB with cert-manager TLS.
|
||||
|
||||
### Tenancy and auth
|
||||
|
||||
Everything lands in a **single tenant**, distinguished by mandatory identity labels rather than VictoriaMetrics multitenancy — one organization, mutual trust, everything queryable together. A **single shared bearer token** authenticates all remote clusters; it is stored in 1Password and injected via ExternalSecret into a `VMUser` that grants write-only access. Because the operator installs the `VMUser` CRD, those resources live in a separate Flux Kustomization that depends on the VictoriaMetrics release and external-secrets, so they don't race CRD registration.
|
||||
Each cluster lands in its **own VictoriaMetrics tenant** (`accountID` = project, `projectID` = cluster; the registry lives in the [shipping guide](05-shipping-metrics-guide.md#tenants)), and every series also carries the mandatory identity labels: one organization, mutual trust, everything queryable together through the multitenant read endpoint. Tenancy is derived centrally rather than declared by shippers: a **single shared bearer token** authenticates all remote clusters, the `VMUser` it is bound to rewrites writes onto the multitenant insert endpoint, and `vminsert` relabeling maps each sample's `project`/`cluster` labels to a tenant. Remotes ship only the five identity labels. Per-tenant enforcement is available later by splitting a remote onto its own token pinned to its tenant path. The token is stored in 1Password and injected via ExternalSecret into that `VMUser`. Because the operator installs the `VMUser` CRD, those resources live in a separate Flux Kustomization that depends on the VictoriaMetrics release and external-secrets, so they don't race CRD registration.
|
||||
|
||||
### Label convention
|
||||
|
||||
@@ -35,7 +35,7 @@ Every shipper — metrics and logs, remote and local — stamps the same five id
|
||||
|
||||
### Onboarding a remote cluster
|
||||
|
||||
Nothing changes on the central side. On the remote cluster: pull the shared token from the same vault item into a Secret, then configure its `vmagent` with a persistent disk buffer (so a central outage doesn't lose data — it replays on recovery), the mandatory external labels, and a remote-write to the public `vmauth` endpoint authenticated with the bearer token. To rotate access for everyone, change the vault item; ExternalSecrets re-sync on both sides. Per-cluster revocation, if ever needed, means splitting into per-cluster vault items and `VMUser`s.
|
||||
On the central side, add the cluster to the tenant registry (until then it lands in tenant 0). On the remote cluster: pull the shared token from the same vault item into a Secret, then configure its `vmagent` with a persistent disk buffer (so a central outage doesn't lose data; it replays on recovery), the mandatory external labels, and a remote-write to the public `vmauth` endpoint authenticated with the bearer token. To rotate access for everyone, change the vault item; ExternalSecrets re-sync on both sides. Per-cluster revocation, if ever needed, means splitting into per-cluster vault items and `VMUser`s.
|
||||
|
||||
### Operating notes
|
||||
|
||||
|
||||
@@ -127,9 +127,44 @@ The value sets are open - these are the conventions, not a closed enumeration; t
|
||||
|
||||
`externalLabels` only tags *scraped* series; if the shipper also forwards pushed data (e.g. OTLP app metrics through a `vmagent`), apply the same labels with a relabel config instead so ingested series are tagged too.
|
||||
|
||||
## Tenants
|
||||
|
||||
Metrics are also keyed by VictoriaMetrics tenant (`accountID:projectID`): `accountID` is the project, `projectID` is the cluster within it. Shippers never see tenant IDs. They write to the `/insert/0/...` URL with the shared token as shown above, the central `vmauth` rewrites that onto the multitenant insert endpoint, and `vminsert` assigns the tenant from the `project` and `cluster` labels using the relabel rules in `kubernetes/apps/base/victoria-metrics/app/configmap-vminsert-relabel.yaml`, which is the tenant registry. A cluster with no rule lands in tenant 0.
|
||||
|
||||
IDs are unique within a store, and an environment pair (a prod cluster and its staging twin, each shipping to its own store) shares the same tenant so tenant-scoped config carries across environments unchanged; clusters with no twin take the next free number in their project. `accountID` 0 and `projectID` 0 are never assigned: any cluster without a rule pair lands in `0:0`, whatever its project. Rules are keyed on `project;cluster` for both labels, never on the project alone, so each cluster migrates independently and a new cluster in a known project cannot be moved by accident. This table records the assignments; a cluster is live on its tenant exactly when its rule pair is in the ConfigMap, so add the row and the rule in the same change.
|
||||
|
||||
| Project | accountID | Cluster | Env | Store | Tenant |
|
||||
|---|---|---|---|---|---|
|
||||
| o11y | 1 | o11y | staging, prod | both | `1:1` |
|
||||
| yucca | 2 | father | prod | production | `2:1` |
|
||||
| yucca | 2 | netops | prod | production | `2:2` |
|
||||
| yucca | 2 | spice | prod | production | `2:3` |
|
||||
| yucca | 2 | luke | staging | staging | `2:4` |
|
||||
| harbor | 3 | harbor-infra-prod | prod | production | `3:1` |
|
||||
| harbor | 3 | harbor-infra-staging | staging | staging | `3:1` |
|
||||
| fip | 4 | azad | prod | production | `4:1` |
|
||||
| fmeet | 5 | serverless | staging, prod | both | `5:1` |
|
||||
|
||||
Onboarding a cluster onto its own tenant is a central-side change only: add its rule pair, for example
|
||||
|
||||
```yaml
|
||||
- source_labels: [project, cluster]
|
||||
regex: yucca;father
|
||||
target_label: vm_account_id
|
||||
replacement: "2"
|
||||
- source_labels: [project, cluster]
|
||||
regex: yucca;father
|
||||
target_label: vm_project_id
|
||||
replacement: "1"
|
||||
```
|
||||
|
||||
`vminsert` reloads the file without a restart. No backfill is needed: the cluster's older history stays in tenant 0, where the Fleet datasource still reads it alongside the new tenant, and ages out with retention. The o11y rules also match samples with `project=o11y` and no `cluster` label, which is what its own `vmalert` writes. The o11y cluster's own IDs come from `CLUSTER_VMETRICS_ACCOUNT_ID` and `CLUSTER_VMETRICS_PROJECT_ID` in `kubernetes/clusters/<env>/cluster-settings.yaml`, which also drive its tenant-scoped reads (the default Grafana datasource, `vmalert` and the MCP server).
|
||||
|
||||
Fleet-wide reads use the `/select/multitenant/prometheus` endpoint, which spans every tenant including tenant 0, so cross-cluster dashboards and alerts keep working while clusters migrate. Logs stay in a single tenant and are distinguished by labels only.
|
||||
|
||||
## Verify data is arriving
|
||||
|
||||
From the central side, query for the remote's series. Over the mesh (no token), the browsable UI is at `https://vmetrics.<mesh-domain>/select/0/vmui/`, or query the API directly:
|
||||
From the central side, query for the remote's series. Over the mesh (no token), the browsable UI is at `https://vmetrics.<mesh-domain>/select/multitenant/vmui/`, or query the API directly:
|
||||
|
||||
```bash
|
||||
curl -s 'https://vmauth.o11y.futo.network/select/0/prometheus/api/v1/query' \
|
||||
|
||||
@@ -138,6 +138,8 @@ route:
|
||||
|
||||
So **routing follows the folder automatically** - no per-rule label to set or keep in sync. (Existing rules still carry a `project` *rule* label; it is legacy and unused for routing — distinct from the `project` *series* label every shipper stamps, see the [shipping guide](05-shipping-metrics-guide.md#labels).) Notifications additionally group by `cluster` (alongside `grafana_folder` and `alertname`), so the same rule firing in two clusters arrives as two grouped notifications rather than one blended message.
|
||||
|
||||
**Datasources and tenants.** Two Prometheus-type datasources front the same store. `VictoriaMetrics` (uid `VictoriaMetrics`, the default) carries this cluster's own tenant in its URL and goes through the `self-select` vmauth, which routes only that path, so anything that does not pick a datasource explicitly sees only o11y's series. `VictoriaMetrics Fleet` (uid `VictoriaMetricsFleet`) reads the multitenant endpoint across every tenant, including tenant 0 where unmigrated remotes still land; it is the explicit opt-in for cross-cluster dashboards and rules (`alerts-fleet.yaml`). Tenancy is carried by the datasource URL, never by dashboard JSON or rule queries: both keep filtering and grouping on the `cluster` label, which survives the tenant split unchanged.
|
||||
|
||||
**Alert rule anatomy.** A `GrafanaAlertRuleGroup` (`folderRef: <project>`, an `interval`) with `rules[]`; each rule is a query stage on the `VictoriaMetrics` datasource (uid `VictoriaMetrics`) feeding a `__expr__` threshold stage, plus `labels` (at least `severity`) and `annotations`. See `base/grafana/app/alerts-o11y.yaml` for the pattern (a heartbeat plus target-down and ingestion-stalled rules). Rules that span clusters aggregate `by (cluster)` so each cluster raises its own instance and carries its `cluster` label into notification grouping; store-local rules (the heartbeat, ingestion-stalled) don't.
|
||||
|
||||
## If you ship metrics to this cluster and want dashboards/alerts
|
||||
@@ -145,7 +147,7 @@ So **routing follows the folder automatically** - no per-rule label to set or ke
|
||||
1. Pick a delivery model: **Model A** (recommended for a separate repo/cluster - you own a signed bundle, o11y adds one OCIRepository) or **Model B** (PR the CRs into `base/grafana`).
|
||||
2. Everything you ship files under **your project's folder**; ask for one if it does not exist.
|
||||
3. Routing follows your folder automatically; add a route matching your `grafana_folder` (and, if you want your own channel, a contact point).
|
||||
4. Dashboards use a `$datasource` variable and map `DS_PROMETHEUS` to `VictoriaMetrics`; alerts query the `VictoriaMetrics` datasource.
|
||||
4. Dashboards use a `$datasource` variable and map `DS_PROMETHEUS` to `VictoriaMetrics`; alerts query the `VictoriaMetrics` datasource. Anything that must see other clusters' series uses `VictoriaMetrics Fleet` instead (see Datasources and tenants).
|
||||
5. Tag dashboards by signal/layer in the JSON (`metrics`, `logs`, `infra`, ...) so they stay filterable across folders (see Tags). Alerts that compare across clusters aggregate `by (cluster)`; stamp the five identity labels on your series (see the [shipping guide](05-shipping-metrics-guide.md#labels)) so per-cluster alerting works.
|
||||
|
||||
## How updates flow
|
||||
|
||||
@@ -0,0 +1,188 @@
|
||||
# Runbook: tenant cutover
|
||||
|
||||
Moves a store's o11y cluster from VictoriaMetrics tenant `0` to `1:1`. Staging is rehearsed live with Flux suspended and the branch applied from its own render; production is cut over by merging, in two PRs, with no manual patching. Background and the tenant registry: [shipping guide, Tenants](../../05-shipping-metrics-guide.md#tenants).
|
||||
|
||||
Commands are fish. The `vmctl` image tag in the Job must match the store's VictoriaMetrics version (`helm show chart` on the k8s-stack chart pinned in the environment overlay reports it as `appVersion`).
|
||||
|
||||
```fish
|
||||
set -x D docs/runbooks/vm-tenant-migration
|
||||
set -x VMQ https://vmetrics.staging.o11y.futo.network
|
||||
function apply_ks
|
||||
flux -n flux-system build kustomization $argv[1] --path ./kubernetes/apps/base/$argv[1]/app | kubectl apply --server-side --force-conflicts -f -
|
||||
end
|
||||
```
|
||||
|
||||
## Part A: staging rehearsal
|
||||
|
||||
Run from the repo root on the branch with the staging kube context selected. Flux stays suspended for these four Kustomizations until Part B merges, which also holds any Renovate change to them; keep the window short.
|
||||
|
||||
### A0. Settings keys first
|
||||
|
||||
Flux substitutes `${CLUSTER_VMETRICS_*}` and `${QUOTE}` from the live `cluster-settings` ConfigMap in `flux-system`, which the root `flux-system` Kustomization re-applies from `main` every 10 minutes, removing any key that is not in git. Until PR 1 has merged, the keys only survive while the root Kustomization is suspended. Suspend it, apply the branch's ConfigMap, confirm the keys, and only then render anything. The root Kustomization manages the cluster-level tree only (the child Kustomization objects and this ConfigMap); the children keep reconciling their own paths while it is suspended.
|
||||
|
||||
```fish
|
||||
flux -n flux-system suspend kustomization flux-system
|
||||
kubectl apply --server-side --force-conflicts -f kubernetes/clusters/staging/cluster-settings.yaml
|
||||
kubectl -n flux-system get configmap cluster-settings -o jsonpath='{.data.CLUSTER_VMETRICS_ACCOUNT_ID}:{.data.CLUSTER_VMETRICS_PROJECT_ID} quote={.data.QUOTE}{"\n"}'
|
||||
```
|
||||
|
||||
A substituted value that must stay a string but reads as a number to YAML, such as the MCP tenant `1:1`, is wrapped as `${QUOTE}...${QUOTE}` in the manifest; kustomize drops literal quotes before substitution, so the quotes have to arrive by substitution (see the Flux docs on substituting numbers and booleans).
|
||||
|
||||
### A1. Suspend Flux
|
||||
|
||||
```fish
|
||||
flux -n flux-system suspend kustomization victoria-metrics victoria-metrics-users grafana victoria-metrics-mcp
|
||||
```
|
||||
|
||||
### A2. Fleet datasource onto the multitenant endpoint
|
||||
|
||||
Must land before any writer moves, or `o11y cluster stopped reporting` fires for a day. Only the Fleet datasource is applied here; the default datasource changes in A3 together with the vmauth that serves it.
|
||||
|
||||
```fish
|
||||
flux -n flux-system build kustomization grafana --path ./kubernetes/apps/base/grafana/app | yq 'select(.metadata.name == "victoria-metrics-fleet")' | kubectl apply --server-side --force-conflicts -f -
|
||||
curl -s $VMQ/select/multitenant/prometheus/api/v1/query --data-urlencode 'query=count by (cluster) (up)' | jq -c '.data.result[] | [.metric.cluster, .value[1]]'
|
||||
```
|
||||
|
||||
Expect the same per-cluster counts the tenant-0 path returns. Fleet dashboards in Grafana should render unchanged.
|
||||
|
||||
### A3. Registry, writers, tenant-scoped readers
|
||||
|
||||
Record the cutover time first; the backfill window ends here. A render that fails `flux build` with a YAML error, or shows an empty `:` where a tenant should be, means the settings keys are missing from the live ConfigMap (see A0).
|
||||
|
||||
```fish
|
||||
set -x CUTOVER (date -u '+%Y-%m-%dT%H:%M:%SZ')
|
||||
apply_ks victoria-metrics
|
||||
apply_ks victoria-metrics-users
|
||||
apply_ks grafana
|
||||
apply_ks victoria-metrics-mcp
|
||||
```
|
||||
|
||||
helm-controller upgrades the release and the operator rolls vminsert, vmagent and vmalert. Verify:
|
||||
|
||||
```fish
|
||||
kubectl -n o11y rollout status deploy -l app.kubernetes.io/name=vminsert
|
||||
kubectl -n o11y logs -l app.kubernetes.io/name=vminsert --tail=100 | rg -i relabel
|
||||
curl -s $VMQ/select/1:1/prometheus/api/v1/query --data-urlencode 'query=count(up)' | jq -r '.data.result[0].value[1]'
|
||||
curl -s $VMQ/select/0/prometheus/api/v1/query --data-urlencode 'query=count(up{cluster="o11y"})' | jq -r '.data.result[0].value[1]'
|
||||
```
|
||||
|
||||
The first count grows from zero as scrapes land in `1:1`; the second falls to zero once tenant 0 stops receiving o11y's scrapes. In Grafana, the default datasource now shows only post-cutover data and the Fleet datasource still shows everything.
|
||||
|
||||
Expected transients, not regressions: the o11y ingestion-stalled rules show NoData for one or two scrape intervals until the first samples land in `1:1`, and every `for:`-gated rule restarts its pending timer because vmalert and Grafana now read a tenant with no prior alert state.
|
||||
|
||||
The Fleet datasource returns `vm_account_id` and `vm_project_id` on every series, and Grafana keeps them as alert-instance labels. When a cluster moves tenant, every Fleet-datasource alert instance for it therefore changes identity: the old instance resolves and a new one starts from Pending, so an alert that was already firing for that cluster re-notifies after its `for` duration (observed on staging: luke's node-not-ready and cilium alerts re-fired 15 and 10 minutes after luke's rules loaded). Rules aggregated `by (...)` do not carry the pseudo-labels and are unaffected. No new conditions are created; check the underlying metric if a re-fired alert looks suspicious.
|
||||
|
||||
### A4. Backfill history
|
||||
|
||||
The Job copies a cluster's history from tenant 0 through the multitenant insert path, so the registry assigns its tenant and strips the tenant labels exactly as for live data. Reverse order returns the most recent history first, so tenant-scoped dashboards heal from the seam backwards. The dates and the cluster are substituted on the way to `kubectl`; the tracked manifest keeps its placeholders. On staging it ran once per cluster whose rule pair went live in A3 as a rehearsal of the mechanics; o11y's filter also takes its cluster-less vmalert output, remotes match their cluster exactly. Production backfills only o11y, see Part C.
|
||||
|
||||
```fish
|
||||
set -x START (date -u -d '-30 days' '+%Y-%m-%dT%H:%M:%SZ' 2>/dev/null; or date -u -v-30d '+%Y-%m-%dT%H:%M:%SZ')
|
||||
function backfill --argument-names cluster project cluster_re
|
||||
sd -s CUTOVER_MINUS_30D $START < $D/job-vmctl-backfill.yaml | sd -s '=CUTOVER' "=$CUTOVER" | sd -s CLUSTER_RE $cluster_re | sd -s PROJECT $project | sd -s CLUSTER $cluster | kubectl apply -f -
|
||||
end
|
||||
backfill o11y o11y 'o11y|'
|
||||
backfill luke yucca luke
|
||||
backfill harbor-infra-staging harbor harbor-infra-staging
|
||||
backfill fmeet-serverless fmeet serverless
|
||||
kubectl -n o11y logs -f job/vmctl-backfill-o11y
|
||||
```
|
||||
|
||||
Spot-check a point a week back on both tenants; the counts should match:
|
||||
|
||||
```fish
|
||||
set -x T (date -u -d '-7 days' '+%Y-%m-%dT%H:%M:%SZ' 2>/dev/null; or date -u -v-7d '+%Y-%m-%dT%H:%M:%SZ')
|
||||
curl -s $VMQ/select/0/prometheus/api/v1/query --data-urlencode 'query=count(up{cluster="o11y"})' --data-urlencode "time=$T" | jq -r '.data.result[0].value[1]'
|
||||
curl -s $VMQ/select/1:1/prometheus/api/v1/query --data-urlencode 'query=count(up)' --data-urlencode "time=$T" | jq -r '.data.result[0].value[1]'
|
||||
```
|
||||
|
||||
Duplicates from the replicated source are handled by the existing 20s dedup on vmselect. If a chunk is too large, narrow it with `--vm-native-step-interval=hour` rather than changing vmselect flags.
|
||||
|
||||
### A5. Remove the tenant-0 copy
|
||||
|
||||
Until this step the Fleet datasource sees each migrated cluster's history twice, once per tenant, so `sum` and `count` panels and `count by (cluster) (up offset 1d)` double over the backfilled window. Once the spot-checks pass, delete the migrated series from tenant 0. Deletion is irreversible and is cleaned up during background merges; queries stop returning the series immediately.
|
||||
|
||||
vmselect refuses a single call that touches more than `-search.maxDeleteSeries` series (default 1,000,000; verify with the binary's `--help`), and a cluster's 30-day copy is typically several million unique series, so split each cluster's delete by metric-name prefix. The six groups below partition every possible name; on staging the largest group was about 760k series. Every call must return `204`; a refusal prints vmselect's message and means that group needs a finer split. Deletes are idempotent, so re-running is safe.
|
||||
|
||||
```fish
|
||||
set -l chunks 'a.*' 'c.*' '[bd-j].*' '[k-o].*' '[p-z].*' '[^a-z].*'
|
||||
for base in 'project="o11y",cluster=~"o11y|"' 'project="yucca",cluster="luke"' 'project="harbor",cluster="harbor-infra-staging"' 'project="fmeet",cluster="serverless"'
|
||||
for re in $chunks
|
||||
set -l m "{$base,__name__=~\"$re\"}"
|
||||
printf '%-72s ' $m
|
||||
curl -s -X POST "$VMQ/delete/0/prometheus/api/v1/admin/tsdb/delete_series" --data-urlencode "match[]=$m" -w ' %{http_code}\n'
|
||||
end
|
||||
end
|
||||
curl -s $VMQ/select/multitenant/prometheus/api/v1/query --data-urlencode 'query=count by (vm_account_id, vm_project_id, cluster) (up offset 1d)' | jq -c '.data.result[] | [.metric.vm_account_id + ":" + .metric.vm_project_id, .metric.cluster, .value[1]]'
|
||||
```
|
||||
|
||||
Expect yesterday's `up` listed once per cluster under its own tenant and no `0:0` rows.
|
||||
|
||||
### A6. Post-cutover checks
|
||||
|
||||
- Grafana alert list: `o11y cluster stopped reporting` not pending; the o11y ingestion rules back to Normal.
|
||||
- No tenant pseudo-labels stored in `1:1` (expect `0`; anything else means a writer bypassed the multitenant path):
|
||||
|
||||
```fish
|
||||
curl -s $VMQ/select/1:1/prometheus/api/v1/query --data-urlencode 'query=count({vm_account_id!=""})' | jq -r '.data.result[0].value[1] // 0'
|
||||
```
|
||||
|
||||
- vmui on `vmetrics.<mesh>` redirects to the multitenant path and shows every cluster.
|
||||
- The read-back recipe in the shipping guide, `/select/0/` on `vmauth.<mesh>`, returns o11y's series: that path now rewrites to multitenant.
|
||||
- victoria-metrics-mcp answers with o11y data; it reads tenant `1:1` by default and lists the others via its `tenants` tool.
|
||||
|
||||
### A7. Remote shippers unchanged
|
||||
|
||||
Nothing on a remote changes in this cutover, whether or not its rule pair went live, and the store holds the evidence. Probe by `project` as well as by `up`: a serverless pusher such as fmeet never emits `up` and ships intermittently, so an `up`-only sweep misses it. every remote ships its own vmagent remote-write counters, so shipper health is readable from here without touching the remote. Take the baseline before A1 and compare after A3.
|
||||
|
||||
```fish
|
||||
curl -s $VMQ/select/multitenant/prometheus/api/v1/query --data-urlencode 'query=time() - max by (cluster) (timestamp(up))' | jq -c '.data.result[] | [.metric.cluster, .value[1]]'
|
||||
curl -s $VMQ/select/0/prometheus/api/v1/query --data-urlencode 'query=count by (cluster) (up{cluster!="o11y"})' | jq -c '.data.result[] | [.metric.cluster, .value[1]]'
|
||||
curl -s $VMQ/select/multitenant/prometheus/api/v1/query --data-urlencode 'query=sum by (cluster) (rate(vmagent_remotewrite_requests_total{status_code!~"2.."}[10m])) + sum by (cluster) (rate(vmagent_remotewrite_retries_count_total[10m])) + sum by (cluster) (rate(vmagent_remotewrite_packets_dropped_total[10m]))' | jq -c '.data.result[] | [.metric.cluster, .value[1]]'
|
||||
curl -s $VMQ/select/1:1/prometheus/api/v1/query --data-urlencode 'query=sum by (job) (increase({__name__=~"vmauth_(user|unauthorized_user)_request_backend_errors_total"}[10m]))' | jq -c '.data.result[] | [.metric.job, .value[1]]'
|
||||
```
|
||||
|
||||
Expected: staleness of a few seconds for every remote, the same per-cluster counts as before with each remote in its own tenant or still in tenant 0 according to the registry, zero non-2xx, retries and drops on the remotes' own remote-write counters, and no growth in vmauth backend errors on either the public or the mesh gateway. The `o11y cluster stopped reporting` rule stays quiet for every remote. The read-back recipe in the shipping guide, `/select/0/` on either gateway, returns the remotes' series because that path now rewrites to multitenant.
|
||||
|
||||
## Part B: commit as two PRs
|
||||
|
||||
The branch splits into a behaviour-preserving PR and the cutover PR, so production never has the fleet view and the writers change in the same reconcile. Production's grafana Kustomization does not depend on victoria-metrics, so a single merge would leave their order to chance.
|
||||
|
||||
**PR 1, multitenant plumbing, a no-op while everything is in tenant 0.** The `CLUSTER_VMETRICS_*` and `QUOTE` settings keys in both environments, so the live ConfigMap carries them before any render needs them (a child Kustomization can reconcile a new revision before the root has updated the ConfigMap, which renders empty IDs once); the Fleet datasource and the vmui redirect onto `/select/multitenant`; the shared-token VMUser and the mesh-unauth VMAuth rewrites of `/insert/0/` and `/select/0/` onto multitenant; the vmagent and vmalert `remoteWrite` URLs onto `/insert/multitenant/`. Without registry rules the multitenant path stores everything in tenant 0 and multitenant reads return it, so nothing observable changes.
|
||||
|
||||
**PR 2, the cutover.** The relabel ConfigMap and its kustomization entry; the vminsert `configMaps` and `relabelConfig` args; vmalert `datasource` and `remoteRead`; the self-select VMAuth path; the default Grafana datasource URL; the MCP `VM_DEFAULT_TENANT_ID`; the removed `extra_label` fences; docs and this runbook.
|
||||
|
||||
Before opening either PR, confirm the committed render matches what is live on staging:
|
||||
|
||||
```fish
|
||||
for ks in victoria-metrics victoria-metrics-users grafana victoria-metrics-mcp
|
||||
flux -n flux-system diff kustomization $ks --path ./kubernetes/apps/base/$ks/app
|
||||
end
|
||||
```
|
||||
|
||||
Every diff should be empty. Resume the root Kustomization once PR 1 has merged, since the settings keys then come from git. Resume the four app Kustomizations only after PR 2 has merged; resuming them earlier reverts the cutover on staging.
|
||||
|
||||
```fish
|
||||
flux -n flux-system resume kustomization flux-system
|
||||
# after PR 2:
|
||||
flux -n flux-system resume kustomization victoria-metrics victoria-metrics-users grafana victoria-metrics-mcp
|
||||
```
|
||||
|
||||
## Part C: production cutover
|
||||
|
||||
Merging PR 2 is the production cutover: Flux moves o11y's writers, registry and readers together. Only o11y is backfilled on production, and only 30 days: it is the one cluster with tenant-scoped readers (the default Grafana datasource, vmalert's remoteRead and MCP, all pinned to `1:1`), and a month covers what operational dashboards look at. Older o11y history stays in tenant 0, visible through the Fleet datasource, until the 120-day retention expires it. Staging measured this at about four hours for 30 days of o11y; run it off-peak.
|
||||
|
||||
```fish
|
||||
set -x VMQ https://vmetrics.o11y.futo.network
|
||||
set -x CUTOVER <merge reconcile time, UTC RFC3339>
|
||||
set -x START (date -u -d '-30 days' '+%Y-%m-%dT%H:%M:%SZ' 2>/dev/null; or date -u -v-30d '+%Y-%m-%dT%H:%M:%SZ')
|
||||
backfill o11y o11y 'o11y|'
|
||||
```
|
||||
|
||||
Then A5 for o11y only, the six chunked deletes of `{project="o11y",cluster=~"o11y|"}` from tenant 0, and A6.
|
||||
|
||||
Remotes are not backfilled on production. Expect each remote's already-firing Fleet-datasource alerts to resolve and re-fire once when its rule pair lands (see A3); move remotes at a quiet time or tell the project beforehand. Nothing reads a remote through a tenant-scoped path: the Fleet datasource reads the multitenant endpoint, which returns tenant 0 and the new tenant as one continuous view, so a remote's history simply ages out of tenant 0. Moving a remote is therefore a registry-only change: add its rule pair to the ConfigMap in its own PR, confirm its `vm_project_id` appears on the Fleet datasource, and that is all. Disk is not a constraint either way: production vmstorage nodes run at about 9% of 1.8 TiB each.
|
||||
|
||||
## Rollback
|
||||
|
||||
On staging, `flux resume` before PR 2 merges reverts every object to tenant 0, including removal of the relabel rules. Do that first. Only then copy `1:1` back if needed, with the Job's source and destination swapped to `/select/1:1/prometheus` and `/insert/0/prometheus`; with the rules gone, nothing re-stamps the copied samples. On production, revert PR 2 and follow the same copy-back.
|
||||
@@ -0,0 +1,42 @@
|
||||
---
|
||||
apiVersion: batch/v1
|
||||
kind: Job
|
||||
metadata:
|
||||
name: vmctl-backfill-CLUSTER
|
||||
namespace: o11y
|
||||
spec:
|
||||
backoffLimit: 0
|
||||
ttlSecondsAfterFinished: 604800
|
||||
template:
|
||||
spec:
|
||||
restartPolicy: Never
|
||||
securityContext:
|
||||
runAsNonRoot: true
|
||||
runAsUser: 65534
|
||||
seccompProfile:
|
||||
type: RuntimeDefault
|
||||
containers:
|
||||
- name: vmctl
|
||||
image: victoriametrics/vmctl:v1.151.0
|
||||
securityContext:
|
||||
allowPrivilegeEscalation: false
|
||||
capabilities:
|
||||
drop: ["ALL"]
|
||||
args:
|
||||
- vm-native
|
||||
- -s
|
||||
- --disable-progress-bar
|
||||
- --vm-native-src-addr=http://vmselect-victoria-metrics.o11y.svc.cluster.local:8481/select/0/prometheus
|
||||
- --vm-native-dst-addr=http://vminsert-victoria-metrics.o11y.svc.cluster.local:8480/insert/multitenant/prometheus
|
||||
- --vm-native-filter-match={project="PROJECT",cluster=~"CLUSTER_RE"}
|
||||
- --vm-native-filter-time-start=CUTOVER_MINUS_30D
|
||||
- --vm-native-filter-time-end=CUTOVER
|
||||
- --vm-native-step-interval=day
|
||||
- --vm-native-filter-time-reverse
|
||||
- --vm-concurrency=2
|
||||
resources:
|
||||
requests:
|
||||
cpu: 200m
|
||||
memory: 256Mi
|
||||
limits:
|
||||
memory: 2Gi
|
||||
@@ -9,13 +9,13 @@ spec:
|
||||
matchLabels:
|
||||
dashboards: grafana
|
||||
datasource:
|
||||
# The default datasource: routed through vmauth-self-select, which forces
|
||||
# extra_label=cluster onto every request, so anything that doesn't pick a
|
||||
# datasource explicitly only sees this cluster's own series. Same uid the
|
||||
# chart used, so existing dashboards and alerts resolve unchanged.
|
||||
# The default datasource: reads this cluster's own tenant through
|
||||
# vmauth-self-select, which routes nothing else, so anything that doesn't
|
||||
# pick a datasource explicitly only sees this cluster's own series. Same
|
||||
# uid the chart used, so existing dashboards and alerts resolve unchanged.
|
||||
name: VictoriaMetrics
|
||||
uid: VictoriaMetrics
|
||||
type: prometheus
|
||||
access: proxy
|
||||
isDefault: true
|
||||
url: http://vmauth-self-select.o11y.svc.cluster.local:8427/select/0/prometheus
|
||||
url: http://vmauth-self-select.o11y.svc.cluster.local:8427/select/${CLUSTER_VMETRICS_ACCOUNT_ID}:${CLUSTER_VMETRICS_PROJECT_ID}/prometheus
|
||||
|
||||
@@ -13,6 +13,9 @@ spec:
|
||||
vm:
|
||||
type: cluster
|
||||
entrypoint: http://vmselect-victoria-metrics.o11y.svc.cluster.local:8481
|
||||
env:
|
||||
- name: VM_DEFAULT_TENANT_ID
|
||||
value: ${QUOTE}${CLUSTER_VMETRICS_ACCOUNT_ID}:${CLUSTER_VMETRICS_PROJECT_ID}${QUOTE}
|
||||
route:
|
||||
enabled: true
|
||||
parentRefs:
|
||||
|
||||
@@ -5,11 +5,10 @@ kind: VMAuth
|
||||
metadata:
|
||||
name: self-select
|
||||
spec:
|
||||
# In-cluster only: backs Grafana's default datasource. The appended
|
||||
# extra_label is enforced by vmselect on every request (queries and
|
||||
# label/series metadata alike), pinning anything that queries through
|
||||
# here to this cluster's own series. Cross-cluster reads go through the
|
||||
# explicit VictoriaMetrics Fleet datasource instead.
|
||||
# In-cluster only: backs Grafana's default datasource. Only this cluster's
|
||||
# own tenant path is routed, so nothing that queries through here can reach
|
||||
# another tenant. Cross-cluster reads go through the explicit VictoriaMetrics
|
||||
# Fleet datasource instead.
|
||||
port: "8427"
|
||||
replicaCount: 2
|
||||
podDisruptionBudget:
|
||||
@@ -28,5 +27,4 @@ spec:
|
||||
- static:
|
||||
url: http://vmselect-victoria-metrics.o11y.svc.cluster.local:8481
|
||||
paths:
|
||||
- /select/0/.*
|
||||
target_path_suffix: "?extra_label=cluster=${CLUSTER_NAME}"
|
||||
- /select/${CLUSTER_VMETRICS_ACCOUNT_ID}:${CLUSTER_VMETRICS_PROJECT_ID}/.*
|
||||
|
||||
@@ -0,0 +1,39 @@
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: ConfigMap
|
||||
metadata:
|
||||
name: vminsert-relabel
|
||||
data:
|
||||
relabel.yaml: |
|
||||
- source_labels: [project, cluster]
|
||||
regex: o11y;(${CLUSTER_NAME})?
|
||||
target_label: vm_account_id
|
||||
replacement: "${CLUSTER_VMETRICS_ACCOUNT_ID}"
|
||||
- source_labels: [project, cluster]
|
||||
regex: o11y;(${CLUSTER_NAME})?
|
||||
target_label: vm_project_id
|
||||
replacement: "${CLUSTER_VMETRICS_PROJECT_ID}"
|
||||
- source_labels: [project, cluster]
|
||||
regex: yucca;luke
|
||||
target_label: vm_account_id
|
||||
replacement: "2"
|
||||
- source_labels: [project, cluster]
|
||||
regex: yucca;luke
|
||||
target_label: vm_project_id
|
||||
replacement: "4"
|
||||
- source_labels: [project, cluster]
|
||||
regex: harbor;harbor-infra-staging
|
||||
target_label: vm_account_id
|
||||
replacement: "3"
|
||||
- source_labels: [project, cluster]
|
||||
regex: harbor;harbor-infra-staging
|
||||
target_label: vm_project_id
|
||||
replacement: "1"
|
||||
- source_labels: [project, cluster]
|
||||
regex: fmeet;serverless
|
||||
target_label: vm_account_id
|
||||
replacement: "5"
|
||||
- source_labels: [project, cluster]
|
||||
regex: fmeet;serverless
|
||||
target_label: vm_project_id
|
||||
replacement: "1"
|
||||
@@ -19,12 +19,6 @@ spec:
|
||||
operator:
|
||||
gateway_support: true
|
||||
|
||||
defaultRules:
|
||||
group:
|
||||
spec:
|
||||
params:
|
||||
extra_label: ["cluster=${CLUSTER_NAME}"]
|
||||
|
||||
vmcluster:
|
||||
enabled: true
|
||||
spec:
|
||||
@@ -91,8 +85,12 @@ spec:
|
||||
memory: 1Gi
|
||||
limits:
|
||||
memory: 4Gi
|
||||
configMaps:
|
||||
- vminsert-relabel
|
||||
extraArgs:
|
||||
replicationFactor: "2"
|
||||
relabelConfig: /etc/vm/configs/vminsert-relabel/relabel.yaml
|
||||
relabelConfigCheckInterval: 1m
|
||||
topologySpreadConstraints:
|
||||
- maxSkew: 1
|
||||
topologyKey: kubernetes.io/hostname
|
||||
@@ -131,7 +129,7 @@ spec:
|
||||
enabled: true
|
||||
spec:
|
||||
datasource:
|
||||
url: http://vmselect-victoria-metrics.o11y.svc.cluster.local:8481/select/0/prometheus?extra_label=cluster=${CLUSTER_NAME}
|
||||
url: http://vmselect-victoria-metrics.o11y.svc.cluster.local:8481/select/${CLUSTER_VMETRICS_ACCOUNT_ID}:${CLUSTER_VMETRICS_PROJECT_ID}/prometheus
|
||||
# No `cluster` here: rule results already carry it from the chart's
|
||||
# group-by, and an external label of the same name displaces theirs
|
||||
# to exported_cluster.
|
||||
@@ -148,7 +146,7 @@ spec:
|
||||
remoteWrite:
|
||||
url: http://vminsert-victoria-metrics.o11y.svc.cluster.local:8480/insert/multitenant/prometheus/api/v1/write
|
||||
remoteRead:
|
||||
url: http://vmselect-victoria-metrics.o11y.svc.cluster.local:8481/select/0/prometheus/
|
||||
url: http://vmselect-victoria-metrics.o11y.svc.cluster.local:8481/select/${CLUSTER_VMETRICS_ACCOUNT_ID}:${CLUSTER_VMETRICS_PROJECT_ID}/prometheus/
|
||||
|
||||
alertmanager:
|
||||
enabled: false
|
||||
|
||||
@@ -5,3 +5,4 @@ resources:
|
||||
- ./helmrelease.yaml
|
||||
- ./ocirepository.yaml
|
||||
- ./httproute-mesh.yaml
|
||||
- ./configmap-vminsert-relabel.yaml
|
||||
|
||||
Reference in New Issue
Block a user