feat(all): introduce partition/region/ceph-cluster model across the stack (#222)

* feat: introduce partition/region/ceph-cluster model across the stack

Formalize partition -> region -> {one k8s cluster, many ceph clusters} and
thread it through every layer plus a new yuctl ops CLI.

- tf: deployment/<partition>/<region>/<stack> layout; terragrunt path-parse +
  state key yucca/<partition>/<region>/<stack>; per-region region.hcl (role,
  site_id, datacenter, provider_code, domain); env->partition / site->region
  renames (NetBird object names byte-identical); standardized per-stack
  `discovery` output contract (secrets as op:// refs).
- k8s: clusters/<partition>/<region>/ (staging/austin, prod/htz-fsn1, dev/local);
  role-based kustomize components (primary/secondary); hybrid cluster-settings
  (TF-rendered identity + human fragment); dev-mirror folded into dev/local;
  charts regrouped into charts/{apps,platform,lib,dev}.
- ci: infra.yml partition/region discovery matrix; partition-keyed path filters;
  <partition>-<region> environment gates; image-versions path moves.
- ansible: inventories under <partition>-<region>/<cluster>.
- yuctl: Go/cobra CLI reading the discovery contract from TF state.
- Retire the sietch-talos libvirt VM cluster (dev@local is the k3d cluster);
  ceph inventory_dirname -> <partition>-<region>/<cluster>.

Verified: mise k8s:validate green (3 clusters); yuctl go build/vet; tofu
validate pre-merge (all 9 stacks). Live-staging state migration NOT run.

* fix typo

* commit
This commit is contained in:
Antoine Lecompte
2026-06-29 08:40:29 -04:00
committed by GitHub
parent 54c410f59c
commit c6985d902c
340 changed files with 4140 additions and 1576 deletions
+47 -40
View File
@@ -4,31 +4,32 @@ This directory holds **two parallel trees** with different consumers:
```
kubernetes/
├── clusters/ # ← real clusters (yucca-o11y-style GitOps; reconciled by Flux)
│ ├── staging/ # apps.yaml (cluster-apps entry) + cluster-settings (image-versions is RS-generated)
│ └── production/ # + image-versions (git pin, bumped by the gated promote job)
├── clusters/ # ← real clusters (partition/region GitOps; reconciled by Flux)
│ ├── staging/austin/ # cluster-settings(.generated) + apps.yaml (cluster-apps entry)
│ ├── prod/htz-fsn1/ # + image-versions (git pin, bumped by the gated promote job)
│ └── dev/local/ # ← dev-mirror entry point (repos.yaml + apps.yaml; consumed by Tilt)
├── apps/
│ ├── base/ # reusable HelmReleases (chart + image.repository); tag via ${YUCCA_IMAGE_TAG}
│ ├── staging/ # overlays + flux-system/ (RSIP+ResourceSet image automation, notification Provider/Alert)
│ ├── production/ # overlays + flux-system/ (notification Provider/Alert)
│ │
│ ├── cnpg-system/ # ← dev-mirror tree (consumed by Tilt/k3d for LOCAL dev only)
│ ├── rook-ceph/ # kept as-is; Tilt scans apps/<namespace>/ and skips base|staging|production
│ └── yucca/ # the product stack + its dev infra
├── flux/ # dev-mirror Flux sources + entrypoint (Tilt reads flux/repos)
├── bootstrap/ components/
│ ├── base/ # reusable HelmReleases (chart + image.repository); tag via ${YUCCA_IMAGE_TAG}
│ ├── staging/austin/ # overlay: components [infra, roles/primary] + flux-system/ (image automation, notifications)
│ ├── prod/htz-fsn1/ # overlay: components [infra, roles/primary] + flux-system/ (notifications)
│ └── dev/local/ # ← dev-mirror tree (Tilt/k3d only): yucca/ rook-ceph/ cnpg-system/ repos/
├── components/ # apps/<app>.yaml (single-source Flux Ks), infra/ (platform layers),
│ │ # roles/{primary,secondary}/ (Kustomize Components: per-role app sets)
├── bootstrap/
```
- **`clusters/` + `apps/{base,staging,production}/`** is the o11y-faithful GitOps
surface for the real Talos clusters. The flux-instance (installed by
`tf/deployment/staging/talos/flux.tf`) syncs `clusters/<env>`; merge→build→deploy
is driven by `.github/workflows/deploy.yml`. See **GitOps deploy** below.
- **`apps/<namespace>/` + `flux/`** is the original dev-mirror tree the
[Tiltfile](../Tiltfile) consumes for local k3d dev. Untouched by the GitOps
tree (Tilt skips `base|staging|production`). Consolidating the two (Tilt
consuming `apps/base`) is future work, alongside porting the rest of the
platform (michael, metrics-worker, object storage, ingress, secrets) into
`apps/base`.
- **`clusters/<partition>/<region>/` + `apps/{base,<partition>/<region>}/` + `components/`**
is the GitOps surface for the real Talos clusters. The flux-instance (installed
by `tf/deployment/staging/austin/talos/flux.tf`) syncs `clusters/<partition>/<region>`;
merge→build→deploy is driven by `.github/workflows/deploy.yml`. Each cluster
overlay picks its app set by composing two Kustomize Components: `infra`
(role-independent platform layers) and `roles/<role>` (the role's app set).
See **GitOps deploy** below.
- **`apps/dev/local/` + `clusters/dev/local/`** is the dev-mirror tree the
[Tiltfile](../Tiltfile) consumes for local k3d dev. Tilt scans only
`apps/dev/local/` (allow-list) and reads HelmRepository sources from
`apps/dev/local/repos/`. It keeps its own literal-valued HelmReleases (Tilt
can't resolve Flux `${...}` substitutions).
## GitOps deploy (staging → production)
@@ -40,17 +41,17 @@ a kubeconfig or joins the tailnet.
1. **build** — matrix-builds every app image → `ghcr.io/immich-app/yucca/<app>:0.0.<run_number>`
(+ `:sha-<sha>` for traceability, `:main`). One monotonic monorepo tag for all apps.
2. **staging (automatic, in-cluster)** — the flux-operator `ResourceSetInputProvider`
(`apps/staging/flux-system/image-automation.yaml`, type `OCIArtifactTag`) detects the
(`apps/staging/austin/flux-system/image-automation.yaml`, type `OCIArtifactTag`) detects the
highest `0.0.<n>` tag in GHCR; a `ResourceSet` writes it into the `image-versions`
ConfigMap, which the `cluster-apps` `postBuild.substituteFrom` feeds into every app
HelmRelease as `${YUCCA_IMAGE_TAG}`. No GHA job, no git commit.
3. **production (gated)** — the `promote-production` job runs only behind the
`production` GitHub Environment's **required-reviewers gate** (guarded by repo var
`PROD_CLUSTER_READY`). It pins `clusters/production/image-versions.yaml` to the
`PROD_CLUSTER_READY`). It pins `clusters/prod/htz-fsn1/image-versions.yaml` to the
validated `0.0.<n>` and commits it (`promote-prod.sh`, `[skip ci]`); prod Flux pulls
and applies. Still no cluster access from CI.
4. **visibility** — notification-controller's GitHub `Provider`/`Alert`
(`apps/<env>/flux-system/notifications.yaml`) posts each reconcile result as a
(`apps/<partition>/<region>/flux-system/notifications.yaml`) posts each reconcile result as a
**commit status** (✅/❌ on the deployed commit). (Trade-off vs the old push model:
the rollout no longer streams in the Actions log — it surfaces as the commit status.)
@@ -87,37 +88,41 @@ mise test:e2e:k3d # runs all e2e against the cluster
Note: michael creates **one S3 bucket per restic repository** (the bucket name
comes from the client JWT's `repository` claim; the S3 credentials are a static
RGW user). That's why it uses a full `CephObjectStoreUser` (charts/ceph-objectuser)
RGW user). That's why it uses a full `CephObjectStoreUser` (charts/platform/ceph-objectuser)
rather than a bucket-scoped ObjectBucketClaim.
## How it reconciles
Flux applies `kubernetes/flux/cluster` first (`cluster-repos` → `cluster-apps`).
`cluster-apps` builds `kubernetes/apps`, which aggregates one Flux `Kustomization`
(`ks.yaml`) per app. Each `ks.yaml` reconciles its `app/` directory, whose
`kustomization.yaml` applies a single `HelmRelease`. Ordering is expressed with
`dependsOn` (operator → database → apps).
In dev, Flux applies `kubernetes/clusters/dev/local` first (`cluster-repos` →
`cluster-apps`). `cluster-apps` builds `kubernetes/apps/dev/local`, which
aggregates one Flux `Kustomization` (`ks.yaml`) per app. Each `ks.yaml`
reconciles its `app/` directory, whose `kustomization.yaml` applies a single
`HelmRelease`. Ordering is expressed with `dependsOn` (operator → database → apps).
```
apps/yucca/<app>/
apps/dev/local/yucca/<app>/
├── ks.yaml # Flux Kustomization → ./app (+ dependsOn)
└── app/
├── kustomization.yaml
└── helmrelease.yaml # chart ref + values
```
(On the real clusters the equivalent single-source Flux `Kustomization`s live in
`components/apps/<app>.yaml`, selected per cluster by the `roles/<role>`
Component.)
## Single source of truth: this tree
The [Tiltfile](../Tiltfile) derives **everything it deploys** from the
HelmReleases here — first-party apps _and_ the remote-chart operators:
- First-party `HelmRelease`s reference the in-repo Helm charts via the `yucca`
`GitRepository` source (`chart: charts/<svc>`), so **no OCI publishing is
`GitRepository` source (`chart: charts/apps/<svc>`), so **no OCI publishing is
required**. Tilt renders the same charts with their dev defaults and injects
the locally-built, live-updated images.
- Remote `HelmRelease`s (cnpg, rook, victoria-\*) pin a chart version + values;
Tilt installs **exactly those**, from the `HelmRepository` sources declared in
`flux/repos/`. Bump a version or value once, in the HelmRelease — there is no
`apps/dev/local/repos/`. Bump a version or value once, in the HelmRelease — there is no
second copy to drift.
Service names are pinned with `fullnameOverride` in each chart's `values.yaml`,
@@ -140,9 +145,9 @@ persistence, secrets) are where a future prod cluster overlay diverges. Notably:
in [`ansible/ceph`](../ansible/ceph) / [`tf/`](../tf)) — not this Rook cluster.
- The Rook-Ceph dev cluster is single-node/single-replica and synthesizes a
loopback block device (k3d has no spare disk). See
[`charts/rook-ceph-cluster`](../charts/rook-ceph-cluster). `michael`'s S3
[`charts/platform/rook-ceph-cluster`](../charts/platform/rook-ceph-cluster). `michael`'s S3
credentials come from a full RGW user
([`charts/ceph-objectuser`](../charts/ceph-objectuser)) whose Secret Rook
([`charts/platform/ceph-objectuser`](../charts/platform/ceph-objectuser)) whose Secret Rook
writes into the `yucca` namespace.
- The dev keypair/secrets committed in chart `secretData` are **well-known
fixtures** (the same keypair lives in `.mise/tasks/*/env`); they must become
@@ -180,9 +185,11 @@ web e2e suite logs in via mock-oidc, so remove `.env` before
## Validate locally
```bash
# render the kustomize graph
kubectl kustomize kubernetes/apps
# build the whole tree (Kustomizations + HelmReleases) the way Flux would
# the full CI gate (charts + every cluster's Flux tree)
mise run k8s:validate
# or, ad hoc: render one dev overlay's kustomize graph
kubectl kustomize kubernetes/apps/dev/local
# build one cluster's whole tree (Kustomizations + HelmReleases) as Flux would
# (https://github.com/allenporter/flux-local)
flux-local build all kubernetes --enable-helm --no-enable-dns
flux-local build all kubernetes/clusters/dev/local --enable-helm --no-enable-dns
```
@@ -8,7 +8,7 @@ spec:
interval: 1h
chart:
spec:
chart: charts/michael
chart: charts/apps/michael
sourceRef:
kind: GitRepository
name: flux-system
+1 -1
View File
@@ -8,7 +8,7 @@ spec:
interval: 1h
chart:
spec:
chart: charts/web
chart: charts/apps/web
sourceRef:
kind: GitRepository
name: flux-system
@@ -8,7 +8,7 @@ spec:
interval: 1h
chart:
spec:
chart: charts/yucca-admin-api
chart: charts/apps/yucca-admin-api
sourceRef:
kind: GitRepository
name: flux-system
@@ -8,7 +8,7 @@ spec:
interval: 1h
chart:
spec:
chart: charts/yucca-api
chart: charts/apps/yucca-api
sourceRef:
kind: GitRepository
name: flux-system
@@ -8,7 +8,7 @@ spec:
interval: 1h
chart:
spec:
chart: charts/cnpg-cluster
chart: charts/platform/cnpg-cluster
sourceRef:
kind: GitRepository
name: flux-system
@@ -8,7 +8,7 @@ spec:
interval: 1h
chart:
spec:
chart: charts/yucca-metrics-worker
chart: charts/apps/yucca-metrics-worker
sourceRef:
kind: GitRepository
name: flux-system
@@ -6,7 +6,7 @@ metadata:
namespace: flux-system
spec:
targetNamespace: cnpg-system
path: ./kubernetes/apps/cnpg-system/cloudnative-pg/app
path: ./kubernetes/apps/dev/local/cnpg-system/cloudnative-pg/app
prune: true
wait: true
interval: 1h
@@ -8,7 +8,7 @@ spec:
interval: 1h
chart:
spec:
chart: charts/rook-ceph-cluster
chart: charts/platform/rook-ceph-cluster
sourceRef:
kind: GitRepository
name: yucca
@@ -6,7 +6,7 @@ metadata:
namespace: flux-system
spec:
targetNamespace: rook-ceph
path: ./kubernetes/apps/rook-ceph/rook-ceph-cluster/app
path: ./kubernetes/apps/dev/local/rook-ceph/rook-ceph-cluster/app
prune: true
wait: true
interval: 1h
@@ -6,7 +6,7 @@ metadata:
namespace: flux-system
spec:
targetNamespace: rook-ceph
path: ./kubernetes/apps/rook-ceph/rook-ceph-operator/app
path: ./kubernetes/apps/dev/local/rook-ceph/rook-ceph-operator/app
prune: true
wait: true
interval: 1h
@@ -8,7 +8,7 @@ spec:
interval: 1h
chart:
spec:
chart: charts/cnpg-cluster
chart: charts/platform/cnpg-cluster
sourceRef:
kind: GitRepository
name: yucca
@@ -9,7 +9,7 @@ spec:
commonMetadata:
labels:
app.kubernetes.io/name: yucca-database
path: ./kubernetes/apps/yucca/database/app
path: ./kubernetes/apps/dev/local/yucca/database/app
prune: true
wait: true
interval: 1h
@@ -8,7 +8,7 @@ spec:
interval: 1h
chart:
spec:
chart: charts/ceph-objectuser
chart: charts/platform/ceph-objectuser
sourceRef:
kind: GitRepository
name: yucca
@@ -9,7 +9,7 @@ spec:
commonMetadata:
labels:
app.kubernetes.io/name: yucca-metrics-object-user
path: ./kubernetes/apps/yucca/metrics-object-user/app
path: ./kubernetes/apps/dev/local/yucca/metrics-object-user/app
prune: true
wait: true
interval: 1h
@@ -8,7 +8,7 @@ spec:
interval: 1h
chart:
spec:
chart: charts/yucca-metrics-worker
chart: charts/apps/yucca-metrics-worker
sourceRef:
kind: GitRepository
name: yucca
@@ -9,7 +9,7 @@ spec:
commonMetadata:
labels:
app.kubernetes.io/name: yucca-metrics-worker
path: ./kubernetes/apps/yucca/metrics-worker/app
path: ./kubernetes/apps/dev/local/yucca/metrics-worker/app
prune: true
wait: true
interval: 1h
@@ -8,7 +8,7 @@ spec:
interval: 1h
chart:
spec:
chart: charts/michael
chart: charts/apps/michael
sourceRef:
kind: GitRepository
name: yucca
@@ -9,7 +9,7 @@ spec:
commonMetadata:
labels:
app.kubernetes.io/name: yucca-michael
path: ./kubernetes/apps/yucca/michael/app
path: ./kubernetes/apps/dev/local/yucca/michael/app
prune: true
wait: true
interval: 1h
@@ -8,7 +8,7 @@ spec:
interval: 1h
chart:
spec:
chart: charts/mock-oidc
chart: charts/dev/mock-oidc
sourceRef:
kind: GitRepository
name: yucca
@@ -9,7 +9,7 @@ spec:
commonMetadata:
labels:
app.kubernetes.io/name: yucca-mock-oidc
path: ./kubernetes/apps/yucca/mock-oidc/app
path: ./kubernetes/apps/dev/local/yucca/mock-oidc/app
prune: true
wait: true
interval: 1h
@@ -8,7 +8,7 @@ spec:
interval: 1h
chart:
spec:
chart: charts/ceph-objectuser
chart: charts/platform/ceph-objectuser
sourceRef:
kind: GitRepository
name: yucca
@@ -9,7 +9,7 @@ spec:
commonMetadata:
labels:
app.kubernetes.io/name: yucca-object-user
path: ./kubernetes/apps/yucca/object-user/app
path: ./kubernetes/apps/dev/local/yucca/object-user/app
prune: true
wait: true
interval: 1h
@@ -9,7 +9,7 @@ spec:
commonMetadata:
labels:
app.kubernetes.io/name: yucca-victoria-logs
path: ./kubernetes/apps/yucca/victoria-logs/app
path: ./kubernetes/apps/dev/local/yucca/victoria-logs/app
prune: true
wait: true
interval: 1h
@@ -9,7 +9,7 @@ spec:
commonMetadata:
labels:
app.kubernetes.io/name: yucca-victoria-metrics
path: ./kubernetes/apps/yucca/victoria-metrics/app
path: ./kubernetes/apps/dev/local/yucca/victoria-metrics/app
prune: true
wait: true
interval: 1h
@@ -8,7 +8,7 @@ spec:
interval: 1h
chart:
spec:
chart: charts/web
chart: charts/apps/web
sourceRef:
kind: GitRepository
name: yucca
@@ -9,7 +9,7 @@ spec:
commonMetadata:
labels:
app.kubernetes.io/name: yucca-web
path: ./kubernetes/apps/yucca/web/app
path: ./kubernetes/apps/dev/local/yucca/web/app
prune: true
wait: true
interval: 1h
@@ -8,7 +8,7 @@ spec:
interval: 1h
chart:
spec:
chart: charts/yucca-admin-api
chart: charts/apps/yucca-admin-api
sourceRef:
kind: GitRepository
name: yucca
@@ -9,7 +9,7 @@ spec:
commonMetadata:
labels:
app.kubernetes.io/name: yucca-admin-api
path: ./kubernetes/apps/yucca/yucca-admin-api/app
path: ./kubernetes/apps/dev/local/yucca/yucca-admin-api/app
prune: true
wait: true
interval: 1h
@@ -8,7 +8,7 @@ spec:
interval: 1h
chart:
spec:
chart: charts/yucca-api
chart: charts/apps/yucca-api
sourceRef:
kind: GitRepository
name: yucca
@@ -9,7 +9,7 @@ spec:
commonMetadata:
labels:
app.kubernetes.io/name: yucca-api
path: ./kubernetes/apps/yucca/yucca-api/app
path: ./kubernetes/apps/dev/local/yucca/yucca-api/app
prune: true
wait: true
interval: 1h
@@ -0,0 +1,13 @@
---
# prod@htz-fsn1 cluster overlay. Authored ahead of the cluster (the prod
# Talos/flux stack isn't built yet — tf/deployment/prod/htz-fsn1 is fabric-only);
# activates once that stack provisions flux and it syncs clusters/prod/htz-fsn1.
# htz-fsn1 is the prod partition's primary (and currently only) region, so it
# composes `infra` + `roles/primary` (the full app set).
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- ./flux-system
components:
- ../../../components/infra
- ../../../components/roles/primary
@@ -1,7 +0,0 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- ./flux-system
- ./cnpg-system
- ./yucca
@@ -1,9 +0,0 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- ./namespace.yaml
- ./yucca-database.yaml
- ./yucca-api.yaml
- ./yucca-admin-api.yaml
- ./web.yaml
@@ -1,26 +0,0 @@
# yaml-language-server: $schema=https://k8s-schemas.home-operations.com/kustomize.toolkit.fluxcd.io/kustomization_v1.json
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: yucca-database
namespace: flux-system
spec:
dependsOn:
- name: cnpg-operator
healthChecks:
- apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
name: yucca-database
namespace: yucca
interval: 1h
retryInterval: 2m
timeout: 10m
path: ./kubernetes/apps/base/yucca-database
prune: true
wait: true
sourceRef:
kind: GitRepository
name: flux-system
namespace: flux-system
targetNamespace: yucca
@@ -0,0 +1,13 @@
---
# staging@austin cluster overlay. The flux-instance sync (installed by the TF
# helm module) reconciles clusters/staging/austin, which points cluster-apps at
# this path. Role + infra are selected by composing two Components: `infra`
# (role-independent platform layers) and `roles/primary` (the full app set —
# austin is the staging partition's primary region).
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- ./flux-system
components:
- ../../../components/infra
- ../../../components/roles/primary
@@ -1,24 +0,0 @@
# yaml-language-server: $schema=https://k8s-schemas.home-operations.com/kustomize.toolkit.fluxcd.io/kustomization_v1.json
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: cnpg-operator
namespace: flux-system
spec:
healthChecks:
- apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
name: cloudnative-pg
namespace: cnpg-system
interval: 1h
retryInterval: 2m
timeout: 10m
path: ./kubernetes/apps/base/cloudnative-pg
prune: true
wait: true
sourceRef:
kind: GitRepository
name: flux-system
namespace: flux-system
targetNamespace: cnpg-system
@@ -1,6 +0,0 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- ./namespace.yaml
- ./cloudnative-pg.yaml
@@ -1,5 +0,0 @@
---
apiVersion: v1
kind: Namespace
metadata:
name: cnpg-system
@@ -1,10 +0,0 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- ./flux-system
- ./cnpg-system
- ./network
- ./observability
- ./storage
- ./yucca
@@ -1,11 +0,0 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- ./namespace.yaml
- ./yucca-database.yaml
- ./yucca-api.yaml
- ./yucca-admin-api.yaml
- ./web.yaml
- ./michael.yaml
- ./yucca-metrics-worker.yaml
-26
View File
@@ -1,26 +0,0 @@
# yaml-language-server: $schema=https://k8s-schemas.home-operations.com/kustomize.toolkit.fluxcd.io/kustomization_v1.json
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: yucca-web
namespace: flux-system
spec:
dependsOn:
- name: yucca-api
healthChecks:
- apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
name: yucca-web
namespace: yucca
interval: 1h
retryInterval: 2m
timeout: 10m
path: ./kubernetes/apps/base/web
prune: true
wait: true
sourceRef:
kind: GitRepository
name: flux-system
namespace: flux-system
targetNamespace: yucca
@@ -1,26 +0,0 @@
# yaml-language-server: $schema=https://k8s-schemas.home-operations.com/kustomize.toolkit.fluxcd.io/kustomization_v1.json
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: yucca-admin-api
namespace: flux-system
spec:
dependsOn:
- name: yucca-database
healthChecks:
- apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
name: yucca-admin-api
namespace: yucca
interval: 1h
retryInterval: 2m
timeout: 10m
path: ./kubernetes/apps/base/yucca-admin-api
prune: true
wait: true
sourceRef:
kind: GitRepository
name: flux-system
namespace: flux-system
targetNamespace: yucca
@@ -1,26 +0,0 @@
# yaml-language-server: $schema=https://k8s-schemas.home-operations.com/kustomize.toolkit.fluxcd.io/kustomization_v1.json
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: yucca-api
namespace: flux-system
spec:
dependsOn:
- name: yucca-database
healthChecks:
- apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
name: yucca-api
namespace: yucca
interval: 1h
retryInterval: 2m
timeout: 10m
path: ./kubernetes/apps/base/yucca-api
prune: true
wait: true
sourceRef:
kind: GitRepository
name: flux-system
namespace: flux-system
targetNamespace: yucca
-5
View File
@@ -1,5 +0,0 @@
---
apiVersion: v1
kind: Namespace
metadata:
name: yucca
+5 -5
View File
@@ -12,7 +12,7 @@ flux install # or the flux-operator, once we adopt it
## 2. Git auth (the repo is private)
`kubernetes/flux/repos/git-yucca.yaml` references `secretRef: yucca-repo-auth`.
`kubernetes/apps/dev/local/repos/git-yucca.yaml` references `secretRef: yucca-repo-auth`.
Create it before applying anything, using a read-only deploy key (preferred)
or a fine-grained PAT:
@@ -38,12 +38,12 @@ GitRepository — which doesn't exist until something applies it. Break the
chicken-and-egg from your checkout:
```bash
kubectl apply -k kubernetes/flux/repos # GitRepository + HelmRepositories
kubectl apply -k kubernetes/flux/cluster # cluster-repos -> cluster-apps
kubectl apply -k kubernetes/apps/dev/local/repos # GitRepository + HelmRepositories
kubectl apply -k kubernetes/clusters/dev/local # cluster-repos -> cluster-apps
```
From here Flux owns the tree: it reconciles `kubernetes/flux/repos` (including
any future source changes) and `kubernetes/apps`.
From here Flux owns the tree: it reconciles `kubernetes/apps/dev/local/repos`
(including any future source changes) and `kubernetes/apps/dev/local`.
## 4. Watch it converge
@@ -6,7 +6,7 @@ metadata:
namespace: flux-system
spec:
interval: 30m
path: ./kubernetes/apps
path: ./kubernetes/apps/dev/local
prune: true
wait: false
sourceRef:
@@ -6,7 +6,7 @@ metadata:
namespace: flux-system
spec:
interval: 1h
path: ./kubernetes/flux/repos
path: ./kubernetes/apps/dev/local/repos
prune: true
wait: true
sourceRef:
@@ -1,5 +1,10 @@
# yaml-language-server: $schema=https://k8s-schemas.home-operations.com/kustomize.toolkit.fluxcd.io/kustomization_v1.json
---
# cluster-apps entry point for prod@htz-fsn1. Authored ahead of the cluster (the
# prod Talos/flux stack isn't built yet); activates once that stack provisions
# flux and it syncs kubernetes/clusters/prod/htz-fsn1. Precedence (last wins):
# cluster-settings-generated (TF) -> cluster-settings (human) -> image-versions
# (the committed, CI-promoted prod tag); keys are disjoint by design.
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
@@ -7,7 +12,7 @@ metadata:
namespace: flux-system
spec:
interval: 1h
path: ./kubernetes/apps/production
path: ./kubernetes/apps/prod/htz-fsn1
prune: true
sourceRef:
kind: GitRepository
@@ -26,6 +31,8 @@ spec:
spec:
postBuild:
substituteFrom:
- kind: ConfigMap
name: cluster-settings-generated
- kind: ConfigMap
name: cluster-settings
- kind: ConfigMap
@@ -0,0 +1,31 @@
---
# TF-RENDERED — do not hand-edit. Authored ahead of the cluster (prod's
# Talos/flux stack isn't built yet — tf/deployment/prod/htz-fsn1 is fabric-only),
# to a valid value matching what TF will regenerate once that stack lands. Keys
# are DISJOINT from the human cluster-settings + CI image-versions ConfigMaps.
# TODO(prod): real RGW endpoint + observability vmauth host when the cluster is
# provisioned.
apiVersion: v1
kind: ConfigMap
metadata:
name: cluster-settings-generated
namespace: flux-system
data:
# Cluster identity (derived: yucca_<partition>_<region>).
CLUSTER_NAME: yucca_prod_htz_fsn1
CLUSTER_ROLE: primary
METRICS_CLUSTER_LABEL: yucca_prod_htz_fsn1
# TODO: real production ingress domain + restic-gateway host.
APP_DOMAIN: yucca.futo.cloud
GW_HOST: gw.yucca.futo.cloud
# Bare-metal RGW (Ceph S3) gateway. Prod uses a COMPLETELY SEPARATE Ceph from
# staging's sietch cluster. S3_HOST is the same gateway without the scheme
# (michael's S3_BACKEND_DNS_HOST). TODO: real production RGW endpoint.
S3_ENDPOINT: https://s3.prod.futo.cloud
S3_HOST: s3.prod.futo.cloud
# Observability egress. TODO: real prod vmauth host.
VMAGENT_OTLP: vmagent-yucca.observability.svc:8429
VLOGS_REMOTE_URL: https://vmauth.prod.futostatus.com/insert/native
@@ -0,0 +1,21 @@
---
# Per-cluster, HUMAN-managed settings for prod@htz-fsn1. Authored ahead of the
# cluster; activates once the prod flux stack syncs this path. Neither CI nor TF
# touch this file. Keys here are disjoint from cluster-settings.generated.yaml
# (TF) and image-versions.yaml (CI). TODO(prod): real ingress VIPs + ops contact.
apiVersion: v1
kind: ConfigMap
metadata:
name: cluster-settings
namespace: flux-system
data:
# TODO: real production OIDC issuer (Zitadel prod instance).
OIDC_ISSUER: https://external-dev-gkhk8b.us1.zitadel.cloud
# ─── Ingress entry point ─────────────────────────────────────────────
# TODO: real prod ingress VIPs once the htz-fsn1 cluster is provisioned.
INGRESS_INTERNAL_IP: 10.10.10.16
INGRESS_PUBLIC_IP: 0.0.0.0
# ACME registration contact for Let's Encrypt. TODO: real ops address.
ACME_EMAIL: fubar-prod@futo.org
@@ -2,6 +2,7 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- ./cluster-settings.generated.yaml
- ./cluster-settings.yaml
- ./image-versions.yaml
- ./apps.yaml
@@ -1,18 +0,0 @@
---
# Production environment settings. Authored now; activates when the prod
# cluster is built and its flux-instance syncs kubernetes/clusters/production.
apiVersion: v1
kind: ConfigMap
metadata:
name: cluster-settings
namespace: flux-system
data:
CLUSTER_NAME: yucca_production
# TODO: real production ingress domain.
APP_DOMAIN: yucca.futo.cloud
# Bare-metal RGW (Ceph S3) gateway for yucca-metrics-worker / michael. Prod
# uses a COMPLETELY SEPARATE Ceph from staging's sietch cluster. S3_HOST is the
# same gateway without the scheme (michael's S3_BACKEND_DNS_HOST).
# TODO: real production RGW endpoint.
S3_ENDPOINT: https://s3.prod.futo.cloud
S3_HOST: s3.prod.futo.cloud
@@ -1,10 +1,11 @@
# yaml-language-server: $schema=https://k8s-schemas.home-operations.com/kustomize.toolkit.fluxcd.io/kustomization_v1.json
---
# cluster-apps entry point — the flux-instance sync (installed by the TF helm
# module) points here (kubernetes/clusters/staging). Mirrors yucca-o11y: a
# single Kustomization over kubernetes/apps/staging, plus one patch that gives
# EVERY child Kustomization a postBuild.substituteFrom for the cluster-settings
# (human) + image-versions (CI) ConfigMaps.
# module) points here (kubernetes/clusters/staging/austin). A single
# Kustomization over kubernetes/apps/staging/austin, plus one patch that gives
# EVERY child Kustomization a postBuild.substituteFrom over the three disjoint
# settings ConfigMaps. Precedence (last wins): cluster-settings-generated (TF) ->
# cluster-settings (human) -> image-versions (CI); keys are disjoint by design.
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
@@ -12,7 +13,7 @@ metadata:
namespace: flux-system
spec:
interval: 1h
path: ./kubernetes/apps/staging
path: ./kubernetes/apps/staging/austin
prune: true
sourceRef:
kind: GitRepository
@@ -31,6 +32,8 @@ spec:
spec:
postBuild:
substituteFrom:
- kind: ConfigMap
name: cluster-settings-generated
- kind: ConfigMap
name: cluster-settings
- kind: ConfigMap
@@ -0,0 +1,39 @@
---
# TF-RENDERED — do not hand-edit. The staging/austin talos stack renders this
# file (templatefile + local_file) from the derived partition/region topology
# and commits it, so the tree is always kustomize-buildable even though prod has
# no flux stack yet. Hand-authored here to a valid value that matches what TF
# will regenerate (CLUSTER_NAME=yucca_<partition>_<region>, etc.). These keys are
# DISJOINT from the human cluster-settings + CI image-versions ConfigMaps; the
# cluster-apps postBuild merges all three.
apiVersion: v1
kind: ConfigMap
metadata:
name: cluster-settings-generated
namespace: flux-system
data:
# Cluster identity (derived: yucca_<partition>_<region>).
CLUSTER_NAME: yucca_staging_austin
# Role of this region within its partition (authoritative in TF discovery.role;
# mirrored compile-time by the components/roles/<role> the overlay composes).
CLUSTER_ROLE: primary
# Metric label stamped on everything vmagent ships (o11y `cluster` identity).
METRICS_CLUSTER_LABEL: yucca_staging_austin
# Per-service ingress hostnames (routed by the ingress layer to in-cluster
# services; resolve to INGRESS_PUBLIC_IP publicly / INGRESS_INTERNAL_IP on-LAN).
APP_DOMAIN: staging.backups.futo.cloud # web (apex; /api routes to yucca-api)
GW_HOST: gw.staging.backups.futo.cloud # michael (restic gateway)
# Bare-metal RGW (Ceph S3) gateway for the staging/austin region; self-signed
# cert, so consumers skip TLS verify. Consumed by yucca-metrics-worker
# (radosEndpoint) and michael (S3_ENDPOINT). S3_HOST is the same gateway
# without the scheme, for michael's DNS-based backend LB (S3_BACKEND_DNS_HOST).
S3_ENDPOINT: https://s3.staging.austin.int.futo.cloud
S3_HOST: s3.staging.austin.int.futo.cloud
# Observability egress (apps -> agents -> o11y vmauth). In-cluster OTLP
# receiver for metrics (host:port, no scheme); VictoriaLogs native ingest URL
# for the collector's egress.
VMAGENT_OTLP: vmagent-yucca.observability.svc:8429
VLOGS_REMOTE_URL: https://vmauth.staging.futostatus.com/insert/native
@@ -0,0 +1,23 @@
---
# Per-cluster, HUMAN-managed settings. Injected into every child Kustomization
# via the cluster-apps postBuild.substituteFrom patch (see apps.yaml). Neither CI
# nor TF touch this file — image tags live in image-versions (CI), and derived
# topology lives in cluster-settings.generated.yaml (TF). Keys here are disjoint
# from those two.
apiVersion: v1
kind: ConfigMap
metadata:
name: cluster-settings
namespace: flux-system
data:
OIDC_ISSUER: https://external-dev-gkhk8b.us1.zitadel.cloud
# ─── Ingress entry point ─────────────────────────────────────────────
# Internal VIP the in-cluster LB/Gateway announces (Cilium L2; distinct from
# the 10.10.10.15 control-plane API VIP). Public IP is the NAT in front of it,
# which staging.backups.futo.cloud resolves to publicly.
INGRESS_INTERNAL_IP: 10.10.10.16
INGRESS_PUBLIC_IP: 97.77.242.205
# ACME registration contact for Let's Encrypt. TODO: real ops address.
ACME_EMAIL: fubar-prod@futo.org
@@ -2,5 +2,6 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- ./cluster-settings.generated.yaml
- ./cluster-settings.yaml
- ./apps.yaml
@@ -1,48 +0,0 @@
---
# Per-environment, human-managed settings. Injected into every child
# Kustomization via the cluster-apps postBuild.substituteFrom patch (see
# apps.yaml), exactly like yucca-o11y. CI does NOT touch this file — see
# image-versions.yaml for the CI-bumped image tag.
apiVersion: v1
kind: ConfigMap
metadata:
name: cluster-settings
namespace: flux-system
data:
CLUSTER_NAME: yucca_staging
# Per-service ingress hostnames (routed by the ingress layer to the in-cluster
# services; resolve to INGRESS_PUBLIC_IP publicly / INGRESS_INTERNAL_IP on-LAN).
APP_DOMAIN: staging.backups.futo.cloud # web (apex; /api routes to yucca-api)
GW_HOST: gw.staging.backups.futo.cloud # michael (restic gateway)
OIDC_ISSUER: https://external-dev-gkhk8b.us1.zitadel.cloud
# Bare-metal RGW (Ceph S3) gateway for the staging environment; self-signed
# cert, so consumers skip TLS verify. Consumed via Flux postBuild by
# yucca-metrics-worker (radosEndpoint) and michael (S3_ENDPOINT). S3_HOST is
# the same gateway without the scheme, for michael's DNS-based backend LB
# (S3_BACKEND_DNS_HOST).
S3_ENDPOINT: https://s3.staging.austin.int.futo.cloud
S3_HOST: s3.staging.austin.int.futo.cloud
# ─── Ingress entry point (yucca_staging) ─────────────────────────────
# Internal VIP the in-cluster LB/Gateway announces (Cilium L2; distinct from
# the 10.10.10.15 control-plane API VIP). Public IP is the NAT in front of it,
# which staging.yucca.futo.cloud resolves to publicly. Consumed by the ingress
# layer (envoy-gateway) + DNS once those land.
INGRESS_INTERNAL_IP: 10.10.10.16
INGRESS_PUBLIC_IP: 97.77.242.205
# ACME registration contact for Let's Encrypt. TODO: real ops address.
ACME_EMAIL: fubar-prod@futo.org
# ─── Observability egress (apps -> agents -> o11y vmauth) ────────────
# Metric label stamped on everything vmagent ships (o11y convention: the
# remote-cluster identity is the `cluster` label).
METRICS_CLUSTER_LABEL: yucca_staging
# In-cluster OTLP receiver for metrics (host:port, no scheme — apps push
# plaintext to vmagent). Logs are collected node-side by victoria-logs-collector
# (tails pod stdout), not pushed by the apps.
VMAGENT_OTLP: vmagent-yucca.observability.svc:8429
# VictoriaLogs native ingest endpoint for the collector's egress (o11y vmauth
# -> vlinsert). The collector stamps the cluster label via collector.extraFields.
VLOGS_REMOTE_URL: https://vmauth.staging.futostatus.com/insert/native

Some files were not shown because too many files have changed in this diff Show More