17 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
What this is
Yucca is a multi-tenant backup service: OIDC-authenticated users get S3-backed restic repositories. The repo is two things at once:
- Application (
packages/) — the NestJS/Go/Svelte services that run the product. - Infrastructure (
tf/,ansible/,kubernetes/,charts/) — everything that operates Yucca: Ceph object storage, Talos K8s, Flux GitOps, networking. See the table inREADME.mdand the per-directoryREADME.mdfiles; each infra subtree is largely self-contained.
Tooling model
Everything is driven through mise tasks, not raw pnpm/go.
Task definitions live in .mise/tasks/ (shell scripts, hierarchical: web/test/_default
is mise run web:test) and .mise/config.toml (the [tasks.*] aggregate tasks + tool
pins). mise also pins every binary version (node, pnpm, go, kubectl, helm, tilt, opentofu,
ansible, etc.) — do not assume a tool is on PATH outside of mise.
Package manager is pnpm workspaces with a catalog (pnpm-workspace.yaml): dependency
versions are centralized there and referenced as "catalog:" in each package.json. Add/bump
shared deps in the catalog, not in individual packages.
Secrets come from 1Password via op run. Tasks that need secrets wrap commands in
op run --env-file=... --; .env files contain op:// references, never literal secrets.
OP_ACCOUNT is set in .env (copy .env.example).
Common commands
# Compose-based dev (fastest inner loop): postgres, minio, mock-oidc, victoria-* in Docker
mise dev # install deps, build common, start docker infra, run all *:dev
mise <pkg>:dev # run one service, e.g. mise web:dev, mise yucca-api:dev
# Quality gates (what CI runs)
mise check # lint + format check + svelte-check + unit tests (= the `checks` CI job)
mise fix # autofix lint/format + lingui extract
mise lint / mise format # individually
mise build # build all packages
# Tests
mise test # all unit tests (jest per NestJS pkg, vitest for web)
mise test:integration # all integration tests (runs --jobs 1; needs infra up)
mise test:e2e # e2e (jest suites) — needs the stack running
mise test:e2e:web # web Playwright e2e
mise <pkg>:test # one package's unit tests, e.g. mise yucca-api:test
mise yucca-api:test -- -t "name of test" # single test (args after -- go to jest)
mise yucca-api:test:watch # watch mode
# Database migrations (yucca-api owns the schema)
mise yucca-api:migrations <args> # wraps @immich/sql-tools against the dev DB
k3d + Tilt (prod-shaped dev)
The alternative dev flow that mirrors prod topology (Helm, CloudNativePG, Rook-Ceph, in-cluster DNS):
mise k3d:up # create local k3d cluster + registry
mise tilt:up # build images, render charts from kubernetes/apps/dev/local, port-forward, live_update
mise tilt:down
mise k3d:down
CI uses the same path: mise tilt:ci-infra (infra only, apps run on the runner) for integration
tests, mise tilt:ci (full stack) for e2e. See Tiltfile (extensively commented) — Tilt's source
of truth is the Flux tree under kubernetes/, not a separate manifest.
Forwarded ports: 5173 web · 3020 yucca-api · 3030 yucca-admin-api · 3010 michael ·
8092 mock-oidc · 9000 ceph rgw (S3) · 8428 victoria-metrics · 9428 victoria-logs.
Infrastructure (terragrunt / ansible)
mise tf:plan / tf:apply / tf:init / tf:destroy # terragrunt for a stack (TF_STACK_DIR=tf/deployment/<partition>/<region>/<stack>)
mise tf:fmt
mise k8s:validate # helm template + kubeconform + flux-local build of the whole k8s tree (CI gate, no cluster)
mise mgmt:render-inventory / mgmt:ansible # render + converge management hosts
tf:* wrap terragrunt in tf/op-run.sh (injects the 1P superuser token from tf/.env).
CI owns terraform applies. Do not run
mise tf:apply(orinfra:apply) by hand —terraform/terragruntapplies are run by CI (.github/workflows/infra.yml) on merge to main. Locally you maytf:planto preview, but never apply. (Also: thetf:*tasks inherit a strayAWS_CA_BUNDLE=~/.config/homelab/root-ca.crtfrom the shell that breaks the OVH S3 state backend — unset it if you plan locally.)NetBird futo-org provider —
netbird_groupresources bug (fixed in 1.0.2):registry.terraform.io/futo-org/netbird≤ 1.0.1 could not UPDATE/DELETE anetbird_groupthat had network resources tagged into it — the TF→API path decoded theresourceslist of{id,type}objects into[]map[string]string("cannot reflect tftypes.Object … into a map"), so a resource-tag group (e.g. htz-fsn1resources) couldn't be renamed in place. Fixed in 1.0.2 — we own the provider (../terraform-provider-netbird); stacks pinversion = "1.0.2". Renaming a setup key still forces replacement (the NetBird API can't rename keys), regenerating its value.
Application architecture
Backend services are NestJS 11 + TypeScript, sharing patterns: controllers → services →
repositories (data access), Zod-validated env.ts, JWT auth guards via an @AuthRoute()
decorator, and observability from the shared @common/server package. Each service imports
@common/server/otel at bootstrap to init the OpenTelemetry SDK (pino logs, OTLP traces/metrics
to victoria-*).
| Service | Lang | Role |
|---|---|---|
yucca-api |
NestJS | User-facing API. Owns auth (OIDC code + device flow, ES256 JWTs), repositories, the database schema + migrations. |
yucca-admin-api |
NestJS | Admin API (user/session/repository management). Shares the same DB + JWT validation. |
michael |
Go | Production restic REST backend — S3 proxy implementing restic's HTTP protocol, with JWT (ECDSA pubkey) verification, WORM enforcement, multi-backend pool/DNS load-balancing. Also attributes traffic to the client's autonomous system (internal/geoip, traffic.* metrics — see o11y/README.md). Deployed in k8s (kubernetes/apps/base/michael). |
restic-api |
NestJS | Earlier TypeScript implementation of the same restic backend, kept as a reference (mise restic-api:dev-reference); not in the deployed app set. |
yucca-metrics-worker |
NestJS | Cron worker (every 5 min): reads bucket usage from RadosGW, writes meter tables, rolls usage up per connection into connectionMetrics (with the per-type billing floor), emits OTel gauges. |
redis (valkey) |
Generic shared platform cache (ephemeral by design; keys namespaced yucca:<service>:<purpose>:*). First tenant: michael's restic-token verdict cache (yucca:michael:verdict:<jti>, DELed by the APIs on revoke); future: michael rate limiting. In-repo chart charts/apps/redis; primary-region only. |
|
mock-oidc-provider |
Node | Dev/test OIDC IdP (code + device flow). Used by compose and k3d when no real issuer is configured. |
common (@common/server) |
TS lib | Shared OTel init, pino logger repository, logging interceptor, the feature-flag registry (FeatureFlags) and connection types (ConnectionTypes). |
Frontend (packages/web) is SvelteKit 5 + Tailwind 4, using @immich/ui, lingui i18n
(mise web:lingui:* to extract/compile — compiled locales are generated, not edited), and the
generated API client. It also embeds the orchestration UI (@futo-org/backups-orchestrator-ui).
packages/yucca-sdk/ is a separately-versioned product: orchestration-api (NestJS) +
orchestration-ui (SvelteKit). Not in the pnpm workspace glob's packages/* alone — note
packages/yucca-sdk/* is added explicitly in pnpm-workspace.yaml.
API client generation (don't hand-edit generated files)
yucca-api controllers/DTOs (via @nestjs/swagger) → mise yucca-api:sync-openapi writes
openapi-specs.json → mise yucca-api-client:build runs oazapfts to generate
packages/yucca-api-client/src/fetch-client.ts (published as @futo-org/backups-api-client,
consumed by web). fetch-client.ts is generated (eslint-ignored). When you change an API
contract, regenerate rather than editing the client.
Connections, feature flags, and restic-token revocation
- Connections (
connectionstable) make "what backs up this account" first-class: a user has N connection instances of typeimmichorrestic. Every repository has aconnectionId(NOT NULL); device-flow sessions bind to a connection via?connection_type=&connection_name=on/auth/oidc/device. Existing repos were backfilled onto a defaultimmichconnection; instance attribution is client-driven viaPOST /connections/:id/adopt(moves default-connection repos to a named instance), never guessed server-side. The in-repo orchestrator (yucca-sdk) does this on device-flow login: it registers as animmichconnection named after its external host, then best-effort adopts its existing repositories onto that instance. The/connectionsAPI surface (list, create, adopt, manage, including multipleimmichinstances) is open to every authenticated user. - Feature flags = registry in code (
@common/serverFeatureFlags), strict-boolean per-user overrides inuserFeatureFlagOverride. Resolution isoverride ?? registry default; the default flips at GA via a release (code-only defaults). Flags gate self-service use of the individual non-default connection type, not the whole surface:connection-restic(experimental, default off):immichneeds none. The mapping lives in@common/serverConnectionTypeFlags/connectionTypeFlag(), checked inConnectionService.createand the device flow; admin-provisioned connections bypass it (admin authority).@RequireFeatureremains as the generic route-level guard for future whole-route gating. Manage from yuctl:users features set/clear,features enable-batch. Boundary rule: env/cluster-settings = deployment config (ops-owned, per-partition); feature flags = per-user product gating (admin-owned, runtime). - Per-type descriptor + billing live in the code registry (
@common/serverConnectionTypeInfos): each type declares its metering tiers,reportsActivity,minObjectSizeBytes(billing floor), andrevocable. Billing keys off the always-available storage tier:yucca-metrics-workerrolls each connection's per-repo RadosGW readings up intoconnectionMetricsand computesbillableBytes(type, size, objects) = max(size, objects * minObjectSizeBytes)(immich floor 0; non-immich 1 MiB, an aggregate approximation, RadosGW gives no per-object histogram).GET /connectionsreturns the rollup. Self-serve restic (flagged):POST /connections/resticcreates connection+repo+long-lived URL in one shot;POST /repository/:id/resticmints for an existing repo, long-lived is revocable-only (restic: defaultRESTIC_JWT_EXPIRES_IN90d,expiresIncapped byRESTIC_JWT_MAX_EXPIRES_IN365d; immich: shortJWT_EXPIRES_INlifetime, customexpiresInrejected: michael never validity-checks non-revocable types, so they must not be long-lived);GET /repository/:id/restic-tokens+DELETE /restic-tokens/:jtiare owner-scoped. Seedocs/connections.md. - Restic-token revocation = postgres truth + layered caches, bounded grace. michael checks a token's
liveness (revocable types only:
REVOCABLE_CONNECTION_TYPES, defaultrestic; immich is skipped) through: L1 per-process cache (REVOCATION_FRESH_TTL_MS60s fresh /REVOCATION_GRACE_MS30min grace) → shared valkey verdict cache (yucca:michael:verdict:<jti>,VERDICT_CACHE_TTL_MS5min, read-through; errors fall through) → yucca-api's internal introspection endpoint (GET /internal/restic-tokens/:jti, shared secretTOKEN_INTROSPECTION_SECRET, answers{active}fromresticTokens). Mint writes only the DB row; revoke flips the row then best-effort DELs the L2 key (lands within ~L1 fresh; a missed DEL self-heals via L2 TTL, no reconcile job). Valkey restart/outage = cache miss → postgres, harmless. Introspection outage: previously-valid jtis honored for the grace window, then fail closed. Enforced only whereTOKEN_INTROSPECTION_URLis set (primary regions). yuctl:tokens list/revoke,repos url --ttl.
Database
PostgreSQL accessed via Kysely. The schema lives in packages/yucca-api/src/schema/
(tables/ definitions, migrations/ SQL run through @immich/sql-tools). yucca-api is the
schema owner; other services read the same DB.
Go services
michael and yuctl are Go 1.25, module <name> + internal/<pkg>, aws-sdk-go-v2 for S3,
rs/zerolog for logs. yuctl (packages/yuctl, cobra CLI) is the operations CLI: it reads the
Terraform discovery contract from S3 state (no terragrunt/checkout) to resolve the
partition→region→{k8s, ceph} topology and drive day-2 ops. See packages/yuctl/README.md.
Infrastructure architecture
The whole fleet is modeled as partition → region → { exactly one K8s cluster, one-or-more
Ceph clusters } (introduced in #222; replaced the older env/site terms). Partitions: prod,
staging, dev. Regions e.g. htz-fsn1, austin, local, plus a global pseudo-region for
account-wide stacks. Slug = <partition>-<region>.
tf/— OpenTofu + Terragrunt, authoritative for cluster identity, 1P secret items, and rendered Ansible inventories. Layout:deployment/<partition>/<region>/<stack>/; shared logic inshared/modules/(ceph-cluster, talos-baremetal, identity, fabric-addressing, netbird-env, fabric/switch config). Every stack emits a non-sensitivediscoveryoutput (the contractyuctl/k8s/ansible consume); all secrets in it areop://refs.render/renders Ansible inventories.ansible/— three subtrees:ceph/(cephadm bare-metal clusters),talos/(Talos K8s as libvirt VMs on the Ceph hypervisors — VM provisioning only; talosctl gen/bootstrap is owned by the TF siderolabs provider),mgmt/(Hetzner management hosts + NetBird routing). Roles are gated by*_enabledflags.kubernetes/— Flux GitOps.clusters/<partition>/<region>/are entry points;apps/{base,<overlays>,dev/local}/hold HelmReleases;components/{infra,apps,roles}/are Kustomize Components selected per role (primary/secondary). Three config layers merged via postBuild:cluster-settings.generated.yaml(TF-rendered),cluster-settings.yaml(human),image-versions(CI). Tilt readsapps/dev/local/.charts/— in-repo Helm charts.lib/yucca-commonis a library chart (shared deployment/ service/secret templates);apps/<svc>are per-service charts depending on it;platform/for rook/cnpg/objectuser. Service names are pinned viafullnameOverrideso in-cluster DNS is identical whether Tilt or Flux renders them.
Deploy flow on merge to main: CI builds images tagged 0.0.<run_number> → Flux auto-promotes the
highest tag to staging → production is promoted by merging the release-please PR, which stamps the
release tag into both prod pins (kubernetes/clusters/prod/htz-fsn1/{flux-release,image-versions}.yaml).
Conventions
- Conventional commits are enforced on PRs (
feat(scope):,fix(scope):,chore:). Common scopes:ceph,netbird,michael,yucca-api,ansible,bgp. Releases are automated via release-please (chore(main): release …PRs); the monorepo is single-versioned. - ESLint flat config (
eslint.config.mjs) is strict on promises:no-floating-promises,no-misused-promises,require-await,await-thenableare all errors. Prettier: single quotes, trailing commas, width 120. - Generated files are eslint-ignored:
**/fetch-client.ts,packages/web/src/locales,dist,build,.svelte-kit. - Default to zero comments. No narration, no restating what the code already says; make the code self-explanatory instead. Add a comment only in the rare case it captures something the code cannot (a why, a constraint, a gotcha) that a reader would otherwise miss.
- Match the package you are in. Read the surrounding code first and follow its existing style and patterns; write code that looks like what is already there, not your own conventions.
Naming
- Cluster names are themed by workload. Kubernetes clusters → Star Wars (
luke= staging,father= the soon-to-be prod). Ceph clusters → Dune (sietch,spice, …). Choose the next themed name when standing up a cluster; it's the Talos/Ceph cluster name + the<clustername>hostname segment. - Node hostnames follow
<product>-<provider>-<region>-<clustername>-<role>-<nodename>— e.g. a staging Talos node isyucca-int-aus-luke-k8s-<word>:product=yucca;role= the workload segment (k8sfor Talos nodes,cephfor Ceph).provider/region= the 3-letterprovider_code/region_codefrom the region'sregion.hcl(austin =int/aus, htz-fsn1 =htz/fsn).<clustername>= the themed cluster name (above).<nodename>= auto-picked from the shared name inventory (tf/shared/modules/node-names/wordlist.txt) — deterministic per cluster, unique within it; pass an explicit nodenameto override.