feat(all): introduce partition/region/ceph-cluster model across the stack (#222)

* feat: introduce partition/region/ceph-cluster model across the stack

Formalize partition -> region -> {one k8s cluster, many ceph clusters} and
thread it through every layer plus a new yuctl ops CLI.

- tf: deployment/<partition>/<region>/<stack> layout; terragrunt path-parse +
  state key yucca/<partition>/<region>/<stack>; per-region region.hcl (role,
  site_id, datacenter, provider_code, domain); env->partition / site->region
  renames (NetBird object names byte-identical); standardized per-stack
  `discovery` output contract (secrets as op:// refs).
- k8s: clusters/<partition>/<region>/ (staging/austin, prod/htz-fsn1, dev/local);
  role-based kustomize components (primary/secondary); hybrid cluster-settings
  (TF-rendered identity + human fragment); dev-mirror folded into dev/local;
  charts regrouped into charts/{apps,platform,lib,dev}.
- ci: infra.yml partition/region discovery matrix; partition-keyed path filters;
  <partition>-<region> environment gates; image-versions path moves.
- ansible: inventories under <partition>-<region>/<cluster>.
- yuctl: Go/cobra CLI reading the discovery contract from TF state.
- Retire the sietch-talos libvirt VM cluster (dev@local is the k3d cluster);
  ceph inventory_dirname -> <partition>-<region>/<cluster>.

Verified: mise k8s:validate green (3 clusters); yuctl go build/vet; tofu
validate pre-merge (all 9 stacks). Live-staging state migration NOT run.

* fix typo

* commit
This commit is contained in:
Antoine Lecompte
2026-06-29 08:40:29 -04:00
committed by GitHub
parent 54c410f59c
commit c6985d902c
340 changed files with 4140 additions and 1576 deletions
+3 -2
View File
@@ -4,8 +4,9 @@ description: >-
read from 1Password. Requires the 1Password CLI on PATH and read from 1Password. Requires the 1Password CLI on PATH and
OP_SERVICE_ACCOUNT_TOKEN in the environment (the infra jobs provide both). The OP_SERVICE_ACCOUNT_TOKEN in the environment (the infra jobs provide both). The
setup key must already exist in 1Password — it is minted by the matching setup key must already exist in 1Password — it is minted by the matching
netbird stack (deployment/<env>/netbird or prod/<site>/netbird), so apply that netbird stack (deployment/<partition>/global/netbird for the account/partition
stack before this action runs on a fresh bootstrap. key, or deployment/<partition>/<region>/netbird for a per-site key), so apply
that stack before this action runs on a fresh bootstrap.
inputs: inputs:
setup-key-ref: setup-key-ref:
+1 -1
View File
@@ -9,7 +9,7 @@
set -euo pipefail set -euo pipefail
TAG="${1:?usage: promote-prod.sh <tag>}" TAG="${1:?usage: promote-prod.sh <tag>}"
FILE="kubernetes/clusters/production/image-versions.yaml" FILE="kubernetes/clusters/prod/htz-fsn1/image-versions.yaml"
# Set data.YUCCA_IMAGE_TAG (yq if present, else portable sed on the one line). # Set data.YUCCA_IMAGE_TAG (yq if present, else portable sed on the one line).
if command -v yq >/dev/null 2>&1; then if command -v yq >/dev/null 2>&1; then
+1 -1
View File
@@ -6,7 +6,7 @@ on:
# The gated prod promotion commits the pin back to main; ignore that path # The gated prod promotion commits the pin back to main; ignore that path
# so it can't retrigger the workflow (it's also marked [skip ci]). # so it can't retrigger the workflow (it's also marked [skip ci]).
paths-ignore: paths-ignore:
- kubernetes/clusters/production/image-versions.yaml - kubernetes/clusters/prod/htz-fsn1/image-versions.yaml
workflow_dispatch: workflow_dispatch:
concurrency: concurrency:
+238 -330
View File
@@ -1,39 +1,53 @@
name: Infra (Terraform) name: Infra (Terraform)
# Applies the Terraform stacks from CI, path-scoped so each group only runs when # Applies the Terragrunt stacks from CI, path-scoped so each partition only runs
# its own files change (a `changes` job emits per-area booleans that gate the # when its own files change. Connectivity to the bare-metal/private nodes is over
# rest; workflow_dispatch overrides and runs everything). Connectivity to the # the NetBird overlay (Tailscale fully retired).
# bare-metal/private nodes is over the NetBird overlay (Tailscale fully retired):
# #
# Staging (tf/deployment/staging/*): ceph, talos, dns, netbird. One staging SA, # Topology model: partition (prod | staging | dev) → region (htz-fsn1, austin,
# a single `staging-infra` Environment gate. The netbird stack mints the CI # local, plus the reserved `global` pseudo-region for partition-wide stacks) →
# setup key; the apply joins the overlay as a `ci` peer to reach the # stack. Stacks live at tf/deployment/<partition>/<region>/<stack>/terragrunt.hcl
# 10.10.10.0/24 nodes, then converges the bare-metal Ceph cluster (ansible/ceph). # (e.g. staging/austin/ceph, staging/global/{dns,netbird}, prod/htz-fsn1/{fabric,
# netbird}, prod/global). A single `discover` job scans that tree into a matrix of
# {partition, region, stack, dir, order}, so adding a stack/region needs no
# workflow edit — only its terragrunt.hcl and a matching GitHub Environment.
# #
# Prod fabric+mgmt (tf/deployment/prod/<site>): the switch fabric + mgmt hosts # Per-partition selection (Actions can't index secrets dynamically, hence the
# (junos-qfx + hetzner providers, built locally). Per-site `prod-<site>` gate. # ternaries):
# Reaches the switch vme / mgmt hosts over NetBird (the mgmt nodes are the # - staging: SA secrets OP_TF_YUCCA_STAGING_ENV(_WRITE), env-file tf/.env.
# route peers for 10.40.5.0/24 et al.); runs the `infra:*` / `mgmt:*` mise tasks. # - prod: SA secrets OP_TF_YUCCA_PROD_ENV(_WRITE), env-file tf/.env.prod.
# The prod `fabric` stack applies with the READ SA and self-escalates to the
# write SA in-vault via `mise run infra:apply`; the prod netbird stacks
# (global + site) apply with the WRITE SA directly.
# - dev is local-only (no remote state, no CI service account) → never emitted.
# #
# Prod NetBird (tf/deployment/prod/global + prod/<site>/netbird): account-wide + # Apply ordering (preserved via a per-entry `order` + max-parallel: 1 on a sorted
# site NetBird groups/keys/policies/routes. Pure api.netbird.io. `prod-infra` gate. # matrix): global-region stacks first (account NetBird + DNS), then site NetBird,
# then node-touching stacks (ceph/talos/fabric). NetBird must precede the
# node-touching stacks because it mints the CI setup key into 1Password and
# advertises the node subnets the overlay-joining jobs reach; prod/global must
# precede prod/htz-fsn1/netbird (a real terragrunt dependency).
#
# Environment gates rekey to <partition>-<region> (one per stack's region):
# staging-austin, staging-global, prod-global, prod-htz-fsn1. Each matrix apply
# entry references its own gate, so an unprovisioned Environment hangs the apply.
# #
# Prerequisites (provisioned out-of-band): # Prerequisites (provisioned out-of-band):
# - Repo secrets: OP_TF_YUCCA_STAGING_ENV (+ _WRITE) — staging SAs; and # - Repo secrets: OP_TF_YUCCA_STAGING_ENV (+ _WRITE); OP_TF_YUCCA_PROD_ENV
# OP_TF_YUCCA_PROD_ENV (read) / OP_TF_YUCCA_PROD_ENV_WRITE (netbird apply) — prod # (read) + OP_TF_YUCCA_PROD_ENV_WRITE (netbird apply / write escalation source).
# SAs. (The fabric apply escalates to the write SA stored in yucca_tf_prod.) # - GitHub Environments with required reviewers: staging-austin, staging-global,
# prod-global, prod-htz-fsn1.
# - BOOTSTRAP — the netbird stacks applied ONCE out-of-band so the CI/mgmt setup # - BOOTSTRAP — the netbird stacks applied ONCE out-of-band so the CI/mgmt setup
# keys exist in 1P before anything tries to connect (CI can't mint them itself: # keys exist in 1P before anything tries to connect (CI can't mint them itself:
# the apply that mints them is gated behind the plan that needs them). E.g.: # the apply that mints them is gated behind the plan that needs them). E.g.:
# OP_SERVICE_ACCOUNT_TOKEN=<staging write SA> \ # OP_SERVICE_ACCOUNT_TOKEN=<staging write SA> \
# TF_STACK_DIR=tf/deployment/staging/netbird mise run tf:apply # tf/op-run.sh terragrunt --working-dir tf/deployment/staging/global/netbird apply
# (staging: NETBIRD_YUCCA_STAGING_CI_SETUP_KEY; prod: NETBIRD_YUCCA_PROD_<SITE>_*). # (staging: NETBIRD_YUCCA_STAGING_CI_SETUP_KEY — partition-scoped, minted by the
# Also: NetBird routes advertising the node subnets (staging 10.10.10.0/24; prod # account/global netbird; prod: NETBIRD_YUCCA_PROD_<REGION>_* — region-scoped,
# the site subnets via the mgmt peers) that the `ci` group may reach. # minted by the per-site netbird). Also: NetBird routes advertising the node
# subnets (staging 10.10.10.0/24; prod the site subnets via the mgmt peers).
# - 1P items: NET_SWITCHES_TERRAFORM_SSH_PRIVATE_KEY (yucca_tf_prod), # - 1P items: NET_SWITCHES_TERRAFORM_SSH_PRIVATE_KEY (yucca_tf_prod),
# NETBOX_API_TOKEN (yucca_tf), HETZNER_WEBSERVICE_API_USER/PASSWORD (yucca_tf_prod). # NETBOX_API_TOKEN (yucca_tf), HETZNER_WEBSERVICE_API_USER/PASSWORD (yucca_tf_prod).
# - GitHub Environments with required reviewers: `staging-infra`, `prod-infra`,
# and one `prod-<site>` per prod fabric site (e.g. prod-htz-fsn1).
on: on:
push: push:
@@ -60,7 +74,7 @@ permissions:
contents: read contents: read
jobs: jobs:
# ── Which areas changed? Outputs gate every downstream job. ────────────────── # ── Which partitions changed? Partition-keyed; `shared` forces all. ──────────
changes: changes:
name: Detect changes name: Detect changes
# Skip on fork PRs (no access to secrets / the overlay anyway). # Skip on fork PRs (no access to secrets / the overlay anyway).
@@ -68,9 +82,9 @@ jobs:
runs-on: ubuntu-latest runs-on: ubuntu-latest
outputs: outputs:
staging: ${{ steps.filter.outputs.staging }} staging: ${{ steps.filter.outputs.staging }}
prod_tf: ${{ steps.filter.outputs.prod_tf }} prod: ${{ steps.filter.outputs.prod }}
prod_ansible: ${{ steps.filter.outputs.prod_ansible }} dev: ${{ steps.filter.outputs.dev }}
prod_netbird: ${{ steps.filter.outputs.prod_netbird }} shared: ${{ steps.filter.outputs.shared }}
steps: steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2 - uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with: with:
@@ -81,379 +95,273 @@ jobs:
filters: | filters: |
staging: staging:
- 'tf/deployment/staging/**' - 'tf/deployment/staging/**'
- 'tf/shared/**'
- 'tf/op-run.sh'
- 'ansible/ceph/**' - 'ansible/ceph/**'
- '.github/actions/netbird-connect/**' prod:
- '.github/workflows/infra.yml'
prod_tf:
# Fabric/mgmt stacks. Netbird-only dirs are covered by prod_netbird;
# an overlap just means a (gated) extra fabric plan — harmless.
- 'tf/deployment/prod/**' - 'tf/deployment/prod/**'
- '!tf/deployment/prod/global/**' - 'ansible/mgmt/**'
- '!tf/deployment/prod/*/netbird/**' - 'tf/render/**'
- 'tf/shared/**'
- 'tf/providers/**' - 'tf/providers/**'
- '.mise/tasks/infra/**' - '.mise/tasks/infra/**'
- '.mise/tasks/fabric/**' - '.mise/tasks/fabric/**'
- '.mise/tasks/mgmt/**' - '.mise/tasks/mgmt/**'
dev:
# Local-only (k3d/Tilt): emits no CI stacks, kept for completeness so
# a `shared` force still reasons about every partition uniformly.
- 'tf/deployment/dev/**'
shared:
# Cross-partition surfaces — a hit forces every partition's matrix.
- 'tf/shared/**'
- 'tf/op-run.sh'
- '.mise/config.toml' - '.mise/config.toml'
- '.github/actions/netbird-connect/**' - '.github/actions/netbird-connect/**'
- '.github/workflows/infra.yml' - '.github/workflows/infra.yml'
prod_ansible:
- 'ansible/mgmt/**'
- 'tf/render/**'
- 'tf/shared/modules/identity/**'
- 'tf/shared/modules/fabric-addressing/**'
- 'tf/deployment/prod/*/mgmt-hosts.yaml'
- '.mise/tasks/mgmt/**'
- '.mise/config.toml'
- '.github/actions/netbird-connect/**'
- '.github/workflows/infra.yml'
prod_netbird:
- 'tf/deployment/prod/global/**'
- 'tf/deployment/prod/*/netbird/**'
- 'tf/shared/modules/netbird-env/**'
- '.github/workflows/infra.yml'
# ── Staging stacks ────────────────────────────────────────────────────────── # ── Discover the stack matrix from the deployment tree ───────────────────────
staging-plan: discover:
name: Staging plan ${{ matrix.stack }} name: Discover stacks
needs: changes needs: changes
if: needs.changes.outputs.staging == 'true' || github.event_name == 'workflow_dispatch' if: >-
needs.changes.outputs.staging == 'true'
|| needs.changes.outputs.prod == 'true'
|| needs.changes.outputs.shared == 'true'
|| github.event_name == 'workflow_dispatch'
runs-on: ubuntu-latest
outputs:
matrix: ${{ steps.gen.outputs.matrix }}
has_stacks: ${{ steps.gen.outputs.has_stacks }}
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
- id: gen
name: Build {partition, region, stack} matrix for the changed partitions
env:
STAGING: ${{ needs.changes.outputs.staging }}
PROD: ${{ needs.changes.outputs.prod }}
SHARED: ${{ needs.changes.outputs.shared }}
DISPATCH: ${{ github.event_name == 'workflow_dispatch' }}
run: |
set -euo pipefail
# Active partitions: own filter OR a shared/manual force. dev is
# local-only and never emitted (no remote state / no CI SA).
declare -A active=()
if [ "$SHARED" = "true" ] || [ "$DISPATCH" = "true" ]; then
active[staging]=1; active[prod]=1
else
[ "$STAGING" = "true" ] && active[staging]=1
[ "$PROD" = "true" ] && active[prod]=1
fi
entries=()
while IFS= read -r tg; do
rel=${tg#tf/deployment/}
rel=${rel%/terragrunt.hcl}
[ "$rel" = "terragrunt.hcl" ] && continue # repo-root parent config
IFS=/ read -r -a segs <<< "$rel"
[ "${#segs[@]}" -ge 2 ] || continue
partition=${segs[0]}
region=${segs[1]}
if [ "${#segs[@]}" -ge 3 ]; then
stack=$(IFS=/; echo "${segs[*]:2}") # nested sub-stacks keep working
else
stack=$region # n==2 transition guard (e.g. prod/global)
fi
[ -n "${active[$partition]:-}" ] || continue
# Apply ordering: global-region stacks (account netbird/dns) first,
# then site netbird, then node-touching stacks (ceph/talos/fabric).
if [ "$region" = "global" ]; then order=0
elif [ "$stack" = "netbird" ]; then order=1
else order=2; fi
entries+=("$(jq -nc \
--arg p "$partition" --arg r "$region" --arg s "$stack" \
--arg d "$rel" --argjson o "$order" \
'{partition:$p,region:$r,stack:$s,dir:$d,order:$o}')")
done < <(find tf/deployment -mindepth 2 -name terragrunt.hcl -type f | sort)
if [ "${#entries[@]}" -eq 0 ]; then
include='[]'
else
include=$(printf '%s\n' "${entries[@]}" | jq -s 'sort_by(.order, .dir)')
fi
matrix=$(jq -nc --argjson inc "$include" '{include:$inc}')
has=$([ "$(jq 'length' <<<"$include")" -gt 0 ] && echo true || echo false)
echo "matrix=$matrix" >> "$GITHUB_OUTPUT"
echo "has_stacks=$has" >> "$GITHUB_OUTPUT"
echo "Discovered stacks:"; echo "$include" | jq -r '.[] | " [\(.order)] \(.partition)@\(.region)/\(.stack) (\(.dir))"'
# ── Plan every changed stack (parallel; read-only) ───────────────────────────
plan:
name: Plan ${{ matrix.partition }}@${{ matrix.region }}/${{ matrix.stack }}
needs: [changes, discover]
if: needs.discover.outputs.has_stacks == 'true'
runs-on: ubuntu-latest runs-on: ubuntu-latest
strategy: strategy:
fail-fast: false fail-fast: false
matrix: matrix: ${{ fromJSON(needs.discover.outputs.matrix) }}
stack: [talos, dns, ceph, netbird]
env: env:
OP_SERVICE_ACCOUNT_TOKEN: ${{ secrets.OP_TF_YUCCA_STAGING_ENV }} # Read-scoped SA + env-file by partition.
OP_SERVICE_ACCOUNT_TOKEN: ${{ matrix.partition == 'prod' && secrets.OP_TF_YUCCA_PROD_ENV || secrets.OP_TF_YUCCA_STAGING_ENV }}
OP_ENV_FILE: ${{ matrix.partition == 'prod' && 'tf/.env.prod' || 'tf/.env' }}
SITE: ${{ matrix.region }}
REGION: ${{ matrix.region }}
PARTITION: ${{ matrix.partition }}
steps: steps:
- name: Checkout - name: Checkout
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2 uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with: with:
persist-credentials: false persist-credentials: false
- name: Set up mise (go + opentofu + terragrunt)
- name: Set up mise (opentofu + terragrunt)
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0 uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
- name: Install 1Password CLI - name: Install 1Password CLI
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0 uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
# The talos plan reads `data.talos_cluster_health` (and the post-CNI gate), # talos + fabric read data sources over the overlay on PLAN (talos_cluster_health;
# which dial the cluster over the 10.10.10.0/24 overlay — so the talos plan # the switch vme), so they must join NetBird. The cloud-API stacks (ceph/dns/
# MUST join NetBird (data sources are read on plan and -refresh=false doesn't # netbird) need no overlay. The setup key already exists in 1P (bootstrapped).
# skip them; the cloud-API stacks need no overlay). The CI setup key must - name: Resolve NetBird CI setup-key ref
# already exist in 1P: it's minted by the netbird apply, so a fresh repo is if: matrix.stack == 'talos' || matrix.stack == 'fabric'
# bootstrapped by applying staging/netbird ONCE out-of-band (see header) — run: |
# after that every plan/apply just reads it, the same way the old Tailscale set -euo pipefail
# OAuth was an out-of-band prerequisite. if [ "$PARTITION" = "prod" ]; then
reg=$(printf '%s' "$REGION" | tr 'a-z-' 'A-Z_')
echo "NB_CI_KEY_REF=op://yucca_tf_prod/NETBIRD_YUCCA_PROD_${reg}_CI_SETUP_KEY/password" >> "$GITHUB_ENV"
else
# Staging's overlay key is partition-scoped (minted by the account/global netbird).
echo "NB_CI_KEY_REF=op://yucca_tf_staging/NETBIRD_YUCCA_STAGING_CI_SETUP_KEY/password" >> "$GITHUB_ENV"
fi
- name: Connect to NetBird - name: Connect to NetBird
if: matrix.stack == 'talos' if: matrix.stack == 'talos' || matrix.stack == 'fabric'
uses: ./.github/actions/netbird-connect uses: ./.github/actions/netbird-connect
with: with:
setup-key-ref: op://yucca_tf_staging/NETBIRD_YUCCA_STAGING_CI_SETUP_KEY/password setup-key-ref: ${{ env.NB_CI_KEY_REF }}
hostname: gha-staging-plan-${{ github.run_id }} hostname: gha-plan-${{ matrix.partition }}-${{ matrix.region }}-${{ matrix.stack }}-${{ github.run_id }}
# Prod fabric goes through the mise task (builds the junos-qfx/hetzner
# providers, renders the NETCONF key, -parallelism=1). Every other stack is a
# plain registry-provider terragrunt plan.
- name: Terragrunt plan (fabric)
if: matrix.stack == 'fabric'
run: mise run infra:plan -- --non-interactive
- name: Terragrunt plan - name: Terragrunt plan
if: matrix.stack != 'fabric'
env:
STACK_DIR: ${{ matrix.dir }}
run: >- run: >-
tf/op-run.sh terragrunt tf/op-run.sh terragrunt
--working-dir tf/deployment/staging/${{ matrix.stack }} --working-dir "tf/deployment/$STACK_DIR"
--non-interactive plan --non-interactive plan
staging-apply: # ── Apply (gated, ordered) ───────────────────────────────────────────────────
name: Staging apply (gated) apply:
needs: [changes, staging-plan] name: Apply ${{ matrix.partition }}@${{ matrix.region }}/${{ matrix.stack }} (gated)
needs: [changes, discover, plan]
# Only on merge to main (or manual dispatch) — never on PRs. # Only on merge to main (or manual dispatch) — never on PRs.
if: (github.event_name == 'push' && needs.changes.outputs.staging == 'true') || github.event_name == 'workflow_dispatch' if: (github.event_name == 'push' || github.event_name == 'workflow_dispatch') && needs.discover.outputs.has_stacks == 'true'
runs-on: ubuntu-latest runs-on: ubuntu-latest
# The approval gate: `staging-infra` Environment with required reviewers. strategy:
environment: staging-infra fail-fast: false
# Serialize so the sorted `order` is honored: account netbird → site netbird
# → node-touching stacks, and prod/global before prod/htz-fsn1/netbird.
max-parallel: 1
matrix: ${{ fromJSON(needs.discover.outputs.matrix) }}
# Per-region approval gate.
environment:
name: ${{ matrix.partition }}-${{ matrix.region }}
env: env:
# write-capable SA (the onepassword provider creates the JWT item) # Write-capable SA by partition, EXCEPT prod/fabric which keeps the read SA
OP_SERVICE_ACCOUNT_TOKEN: ${{ secrets.OP_TF_YUCCA_STAGING_ENV_WRITE }} # and self-escalates to the in-vault write SA inside `mise run infra:apply`.
OP_SERVICE_ACCOUNT_TOKEN: ${{ (matrix.partition == 'prod' && matrix.stack == 'fabric') && secrets.OP_TF_YUCCA_PROD_ENV || matrix.partition == 'prod' && secrets.OP_TF_YUCCA_PROD_ENV_WRITE || secrets.OP_TF_YUCCA_STAGING_ENV_WRITE }}
OP_ENV_FILE: ${{ matrix.partition == 'prod' && 'tf/.env.prod' || 'tf/.env' }}
SITE: ${{ matrix.region }}
REGION: ${{ matrix.region }}
PARTITION: ${{ matrix.partition }}
steps: steps:
- name: Checkout - name: Checkout
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2 uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with: with:
persist-credentials: false persist-credentials: false
- name: Set up mise (go + opentofu + terragrunt + ansible)
- name: Set up mise (opentofu + terragrunt)
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0 uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
- name: Install 1Password CLI - name: Install 1Password CLI
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0 uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
# NetBird stack FIRST — pure api.netbird.io (no overlay needed), and it # Node-touching stacks join the overlay: talos (provisions over the LAN),
# mints the CI setup key into 1P that the connect step below reads. On a # ceph (the Ansible convergence below), fabric (the switch vme + the mgmt
# fresh bootstrap this is what makes the key exist before anything joins. # converge below). NetBird stacks themselves are pure api.netbird.io. The
- name: Apply staging/netbird # key was minted by an earlier (lower-order) netbird apply in this same run.
run: >- - name: Resolve NetBird CI setup-key ref
tf/op-run.sh terragrunt if: matrix.stack == 'talos' || matrix.stack == 'ceph' || matrix.stack == 'fabric'
--working-dir tf/deployment/staging/netbird run: |
--non-interactive apply -auto-approve set -euo pipefail
if [ "$PARTITION" = "prod" ]; then
# Join the NetBird overlay as a `ci` peer so the node-touching stacks and reg=$(printf '%s' "$REGION" | tr 'a-z-' 'A-Z_')
# the Ansible deploy below reach 10.10.10.0/24 (the staging route advertises echo "NB_CI_KEY_REF=op://yucca_tf_prod/NETBIRD_YUCCA_PROD_${reg}_CI_SETUP_KEY/password" >> "$GITHUB_ENV"
# the LAN). else
echo "NB_CI_KEY_REF=op://yucca_tf_staging/NETBIRD_YUCCA_STAGING_CI_SETUP_KEY/password" >> "$GITHUB_ENV"
fi
- name: Connect to NetBird - name: Connect to NetBird
if: matrix.stack == 'talos' || matrix.stack == 'ceph' || matrix.stack == 'fabric'
uses: ./.github/actions/netbird-connect uses: ./.github/actions/netbird-connect
with: with:
setup-key-ref: op://yucca_tf_staging/NETBIRD_YUCCA_STAGING_CI_SETUP_KEY/password setup-key-ref: ${{ env.NB_CI_KEY_REF }}
hostname: gha-staging-apply-${{ github.run_id }} hostname: gha-apply-${{ matrix.partition }}-${{ matrix.region }}-${{ matrix.stack }}-${{ github.run_id }}
# Ceph 1P password items (no node contact), then the Talos cluster # Prod fabric: mise task escalates to the write SA + builds providers.
# (provisions secrets/CNI/Flux), then DNS. - name: Terragrunt apply (fabric)
- name: Apply staging/ceph if: matrix.stack == 'fabric'
run: mise run infra:apply -- --non-interactive -auto-approve
# Everything else: direct terragrunt apply with the partition's write SA.
- name: Terragrunt apply
if: matrix.stack != 'fabric'
env:
STACK_DIR: ${{ matrix.dir }}
run: >- run: >-
tf/op-run.sh terragrunt tf/op-run.sh terragrunt
--working-dir tf/deployment/staging/ceph --working-dir "tf/deployment/$STACK_DIR"
--non-interactive apply -auto-approve --non-interactive apply -auto-approve
- name: Apply staging/talos # ── Ceph convergence (Ansible) — only on the ceph stack ──────────────────
run: >- # The TF apply above only minted the RGW keys into 1P + the cluster Secret;
tf/op-run.sh terragrunt # this creates the matching RGW users on the bare-metal cluster. Reuses the
--working-dir tf/deployment/staging/talos # NetBird overlay + 1Password session already established in this entry.
--non-interactive apply -auto-approve
- name: Apply staging/dns
run: >-
tf/op-run.sh terragrunt
--working-dir tf/deployment/staging/dns
--non-interactive apply -auto-approve
# ── Ceph convergence (Ansible) ──────────────────────────────────────
# The TF stacks above only minted the RGW keys into 1P + the cluster
# Secret; this is what creates the matching RGW users on the bare-metal
# cluster. Runs after the TF apply so the inventory (rendered from the
# ceph stack's `render` output) and the keys exist. Reuses the NetBird
# overlay + 1Password session already established in this job.
- name: Render Ansible inventory from the ceph TF state - name: Render Ansible inventory from the ceph TF state
run: ansible/ceph/scripts/render-inventories.sh staging if: matrix.stack == 'ceph'
env:
PARTITION: ${{ matrix.partition }}
run: ansible/ceph/scripts/render-inventories.sh "$PARTITION"
- name: Install the ansible-iac SSH key from 1Password - name: Install the ansible-iac SSH key from 1Password
# Pulls op://yucca_tf_staging/SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY to if: matrix.stack == 'ceph'
# Pulls op://yucca_tf_<partition>/SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY to
# ~/.ssh/id_ed25519_sietch — the path the rendered inventory references. # ~/.ssh/id_ed25519_sietch — the path the rendered inventory references.
env:
PARTITION: ${{ matrix.partition }}
run: | run: |
mkdir -p ~/.ssh && chmod 700 ~/.ssh mkdir -p ~/.ssh && chmod 700 ~/.ssh
OP_VAULT=yucca_tf_staging ansible/ceph/scripts/install-ssh-keys.sh sietch OP_VAULT="yucca_tf_$PARTITION" ansible/ceph/scripts/install-ssh-keys.sh sietch
- name: Provision the ceph Ansible toolchain (venv + collections) - name: Provision the ceph Ansible toolchain (venv + collections)
if: matrix.stack == 'ceph'
working-directory: ansible/ceph working-directory: ansible/ceph
run: | run: |
mise trust mise trust
mise install mise install
mise run setup mise run setup
- name: Deploy Ceph (full pipeline — baseline → tune → deploy → harden) - name: Deploy Ceph (full pipeline — baseline → tune → deploy → harden)
working-directory: ansible/ceph if: matrix.stack == 'ceph'
# CI runner has no known_hosts for the bare-metal nodes; first contact is # CI runner has no known_hosts for the bare-metal nodes; first contact is
# over the NetBird overlay, so disable strict host-key checking for this run. # over the NetBird overlay, so disable strict host-key checking for this run.
env: env:
ANSIBLE_HOST_KEY_CHECKING: "false" ANSIBLE_HOST_KEY_CHECKING: "false"
CEPH_ENV: inventories/sietch-ceph.staging.austin.int/inventory.ini CEPH_ENV: inventories/${{ matrix.partition }}-${{ matrix.region }}/sietch/inventory.ini
run: mise run deploy run: mise run deploy
# ── Prod NetBird (global + site layers) — mints the CI/mgmt setup keys ──────── # ── Mgmt convergence (Ansible) — only on the fabric stack ────────────────
# Pure api.netbird.io (no nodes/overlay). Layered: prod/global (account-wide # Renders the mgmt inventory from TF (tf/render/ansible-mgmt) then converges
# groups + the yucca→yucca_resource policy) then prod/<site>/netbird (site # the mgmt hosts over the overlay. No-op for regions without an ansible/mgmt
# groups/keys/policies + the routed network). Runs before the fabric/mgmt jobs # inventory. Runs under the same (already-approved) gate as the fabric apply.
# conceptually (it mints the keys they consume), but they're only loosely coupled
# — the keys persist in 1P across runs, so a fresh bootstrap applies this first.
netbird-prod-plan:
name: Plan prod/netbird
needs: changes
if: needs.changes.outputs.prod_netbird == 'true' || github.event_name == 'workflow_dispatch'
runs-on: ubuntu-latest
env:
OP_ENV_FILE: tf/.env.prod
OP_SERVICE_ACCOUNT_TOKEN: ${{ secrets.OP_TF_YUCCA_PROD_ENV }}
steps:
- name: Checkout
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
- name: Set up mise (opentofu + terragrunt)
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
- name: Install 1Password CLI
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
# Global layer first, then the site layer (which reads the global layer's
# group_ids via a terragrunt dependency — mock_outputs cover the PR plan
# before global is ever applied).
- name: Terragrunt plan prod/global
run: >-
tf/op-run.sh terragrunt
--working-dir tf/deployment/prod/global
--non-interactive plan
- name: Terragrunt plan prod/htz-fsn1/netbird
run: >-
tf/op-run.sh terragrunt
--working-dir tf/deployment/prod/htz-fsn1/netbird
--non-interactive plan
netbird-prod-apply:
name: Apply prod/netbird (gated)
needs: [changes, netbird-prod-plan]
if: (github.event_name == 'push' && needs.changes.outputs.prod_netbird == 'true') || github.event_name == 'workflow_dispatch'
runs-on: ubuntu-latest
environment: prod-infra
env:
OP_ENV_FILE: tf/.env.prod
OP_SERVICE_ACCOUNT_TOKEN: ${{ secrets.OP_TF_YUCCA_PROD_ENV_WRITE }}
steps:
- name: Checkout
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
- name: Set up mise (opentofu + terragrunt)
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
- name: Install 1Password CLI
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
# Global layer must apply before the site layer (the site reads global's
# group_ids output via its terragrunt dependency).
- name: Apply prod/global
run: >-
tf/op-run.sh terragrunt
--working-dir tf/deployment/prod/global
--non-interactive apply -auto-approve
- name: Apply prod/htz-fsn1/netbird
run: >-
tf/op-run.sh terragrunt
--working-dir tf/deployment/prod/htz-fsn1/netbird
--non-interactive apply -auto-approve
# ── Prod fabric + mgmt stacks (one per site) ─────────────────────────────────
prod-discover:
name: Discover prod fabric sites
needs: changes
# Needed by both the TF and ansible prod jobs, so run if either area changed.
if: needs.changes.outputs.prod_tf == 'true' || needs.changes.outputs.prod_ansible == 'true' || github.event_name == 'workflow_dispatch'
runs-on: ubuntu-latest
outputs:
matrix: ${{ steps.sites.outputs.matrix }}
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
- id: sites
name: List fabric sites (dirs with a fabric.tf; excludes netbird-only dirs)
run: |
matrix=$(for d in tf/deployment/prod/*/; do [ -f "${d}fabric.tf" ] && basename "$d"; done | jq -R . | jq -cs '{site: .}')
echo "matrix=$matrix" >> "$GITHUB_OUTPUT"
echo "$matrix"
prod-plan:
name: Prod plan ${{ matrix.site }}
needs: [changes, prod-discover]
if: needs.changes.outputs.prod_tf == 'true' || github.event_name == 'workflow_dispatch'
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix: ${{ fromJSON(needs.prod-discover.outputs.matrix) }}
env:
OP_SERVICE_ACCOUNT_TOKEN: ${{ secrets.OP_TF_YUCCA_PROD_ENV }}
SITE: ${{ matrix.site }}
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
- name: Set up mise (go + opentofu + terragrunt)
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
- name: Install 1Password CLI
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
# Reach the switch vme (routed via the mgmt NetBird peers) over the overlay.
- name: Resolve NetBird CI setup-key ref
run: echo "NB_CI_KEY_REF=op://yucca_tf_prod/NETBIRD_YUCCA_PROD_$(echo "$SITE" | tr 'a-z-' 'A-Z_')_CI_SETUP_KEY/password" >> "$GITHUB_ENV"
- name: Connect to NetBird
uses: ./.github/actions/netbird-connect
with:
setup-key-ref: ${{ env.NB_CI_KEY_REF }}
hostname: gha-prod-plan-${{ matrix.site }}-${{ github.run_id }}
- name: Deploy plan
run: mise run infra:plan -- --non-interactive
prod-apply:
name: Prod apply ${{ matrix.site }} (gated)
needs: [changes, prod-discover, prod-plan]
if: (github.event_name == 'push' && needs.changes.outputs.prod_tf == 'true') || github.event_name == 'workflow_dispatch'
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix: ${{ fromJSON(needs.prod-discover.outputs.matrix) }}
# Site-scoped gate: each site's prod stack has its own required reviewers.
environment:
name: prod-${{ matrix.site }}
env:
# Read-scoped token; infra:apply escalates to the write SA stored in the vault.
OP_SERVICE_ACCOUNT_TOKEN: ${{ secrets.OP_TF_YUCCA_PROD_ENV }}
SITE: ${{ matrix.site }}
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
- name: Set up mise (go + opentofu + terragrunt)
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
- name: Install 1Password CLI
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
- name: Resolve NetBird CI setup-key ref
run: echo "NB_CI_KEY_REF=op://yucca_tf_prod/NETBIRD_YUCCA_PROD_$(echo "$SITE" | tr 'a-z-' 'A-Z_')_CI_SETUP_KEY/password" >> "$GITHUB_ENV"
- name: Connect to NetBird
uses: ./.github/actions/netbird-connect
with:
setup-key-ref: ${{ env.NB_CI_KEY_REF }}
hostname: gha-prod-apply-${{ matrix.site }}-${{ github.run_id }}
- name: Deploy apply
run: mise run infra:apply -- --non-interactive -auto-approve
prod-ansible:
name: Prod ansible ${{ matrix.site }}
needs: [changes, prod-discover, prod-apply]
# Runs after the TF apply, but also on ansible-only changes (apply skipped).
# always() so a skipped prod-apply (TF unchanged) doesn't skip this; still
# bails if discover failed or the apply actually failed.
if: >-
always()
&& needs.prod-discover.result == 'success'
&& needs.prod-apply.result != 'failure'
&& needs.prod-apply.result != 'cancelled'
&& ((github.event_name == 'push' && needs.changes.outputs.prod_ansible == 'true') || github.event_name == 'workflow_dispatch')
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix: ${{ fromJSON(needs.prod-discover.outputs.matrix) }}
# No environment gate: the prod-apply gate already approved this deploy, and
# ansible/mgmt only reads from 1Password (read-scoped repo secret).
env:
OP_SERVICE_ACCOUNT_TOKEN: ${{ secrets.OP_TF_YUCCA_PROD_ENV }}
SITE: ${{ matrix.site }}
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
- name: Set up mise (ansible + opentofu)
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
- name: Install 1Password CLI
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
# On the overlay so the playbook can reach mgmt hosts that have joined NetBird
# (it prefers their NetBird IP, falling back to the public IP otherwise).
- name: Resolve NetBird CI setup-key ref
run: echo "NB_CI_KEY_REF=op://yucca_tf_prod/NETBIRD_YUCCA_PROD_$(echo "$SITE" | tr 'a-z-' 'A-Z_')_CI_SETUP_KEY/password" >> "$GITHUB_ENV"
- name: Connect to NetBird
uses: ./.github/actions/netbird-connect
with:
setup-key-ref: ${{ env.NB_CI_KEY_REF }}
hostname: gha-prod-ansible-${{ matrix.site }}-${{ github.run_id }}
# Renders the inventory from TF (tf/render/ansible-mgmt) then converges the
# mgmt hosts — over NetBird if they've joined, else their public IP (the
# TF-generated provisioning key authorizes root). No-op for sites without an
# ansible/mgmt inventory; requires the hosts to have been reprovisioned.
- name: Ansible converge (mgmt hosts) - name: Ansible converge (mgmt hosts)
if: matrix.stack == 'fabric'
run: mise run mgmt:ansible run: mise run mgmt:ansible
+6
View File
@@ -1,6 +1,9 @@
node_modules node_modules
*.tsbuildinfo *.tsbuildinfo
# Claude Code local state — per-operator settings + agent worktrees; never commit.
.claude/
.env.local .env.local
.env .env
@@ -50,6 +53,9 @@ ansible/*/.ansible/
ansible/*/.ansible_facts_cache/ ansible/*/.ansible_facts_cache/
ansible/*/.venv/ ansible/*/.venv/
# Working notes for the partition/region/ceph-cluster rework — local scratch.
notes.md
# Lens / decision-support artifacts are controller-local notes, not repo code. # Lens / decision-support artifacts are controller-local notes, not repo code.
# Per-project: kept outside the repo (e.g., ~/Projects/immich/yucca-ceph-import/analysis/ # Per-project: kept outside the repo (e.g., ~/Projects/immich/yucca-ceph-import/analysis/
# on operator workstation, but not tracked). # on operator workstation, but not tracked).
+5 -5
View File
@@ -73,20 +73,20 @@ run = [{ task = "*:fix" }, { task = "web:lingui" }]
# No literal secrets in the .env file — just op:// references. # No literal secrets in the .env file — just op:// references.
[tasks."tf:init"] [tasks."tf:init"]
description = "Terragrunt init for a given stack (default: deployment/staging/ceph)" description = "Terragrunt init for a given stack (default: deployment/staging/austin/ceph)"
run = "tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/staging/ceph} init" run = "tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/staging/austin/ceph} init"
[tasks."tf:plan"] [tasks."tf:plan"]
description = "Terragrunt plan for a given stack" description = "Terragrunt plan for a given stack"
run = "tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/staging/ceph} plan" run = "tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/staging/austin/ceph} plan"
[tasks."tf:apply"] [tasks."tf:apply"]
description = "Terragrunt apply for a given stack" description = "Terragrunt apply for a given stack"
run = "tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/staging/ceph} apply" run = "tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/staging/austin/ceph} apply"
[tasks."tf:destroy"] [tasks."tf:destroy"]
description = "Terragrunt destroy for a given stack (use with care)" description = "Terragrunt destroy for a given stack (use with care)"
run = "tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/staging/ceph} destroy" run = "tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/staging/austin/ceph} destroy"
[tasks."tf:fmt"] [tasks."tf:fmt"]
description = "Format terraform + terragrunt files recursively" description = "Format terraform + terragrunt files recursively"
+1 -1
View File
@@ -35,4 +35,4 @@ SITE="${SITE:-htz-fsn1}"
unset OP_ACCOUNT unset OP_ACCOUNT
# -parallelism=1: the JTAF junos-qfx provider isn't concurrency-safe — parallel # -parallelism=1: the JTAF junos-qfx provider isn't concurrency-safe — parallel
# ApplyResourceChange calls across the VCs crash the plugin ("Plugin did not respond"). # ApplyResourceChange calls across the VCs crash the plugin ("Plugin did not respond").
OP_ENV_FILE=tf/.env.prod "$ROOT/tf/op-run.sh" terragrunt --working-dir "tf/deployment/prod/$SITE" apply -parallelism=1 "$@" OP_ENV_FILE=tf/.env.prod "$ROOT/tf/op-run.sh" terragrunt --working-dir "tf/deployment/prod/$SITE/fabric" apply -parallelism=1 "$@"
+1 -1
View File
@@ -34,4 +34,4 @@ SITE="${SITE:-htz-fsn1}"
# provider rejects having both set ("service_account_token and account are set"). # provider rejects having both set ("service_account_token and account are set").
unset OP_ACCOUNT unset OP_ACCOUNT
# -parallelism=1: the JTAF junos-qfx provider isn't concurrency-safe (see apply). # -parallelism=1: the JTAF junos-qfx provider isn't concurrency-safe (see apply).
OP_ENV_FILE=tf/.env.prod "$ROOT/tf/op-run.sh" terragrunt --working-dir "tf/deployment/prod/$SITE" plan -parallelism=1 "$@" OP_ENV_FILE=tf/.env.prod "$ROOT/tf/op-run.sh" terragrunt --working-dir "tf/deployment/prod/$SITE/fabric" plan -parallelism=1 "$@"
+26 -17
View File
@@ -9,8 +9,11 @@
# rebuilds charts/*/charts/ (rm -rf + dependency build) and races the renders. # rebuilds charts/*/charts/ (rm -rf + dependency build) and races the renders.
set -euo pipefail set -euo pipefail
LIB_CONSUMERS=(yucca-api yucca-admin-api yucca-metrics-worker web michael mock-oidc) # Charts are role-grouped: apps/* (services), platform/* (operators/CRs),
ALL_CHARTS=(yucca-api yucca-admin-api yucca-metrics-worker web michael mock-oidc cnpg-cluster ceph-objectuser rook-ceph-cluster) # lib/yucca-common (shared library), dev/mock-oidc (dev-only). Paths below are
# relative to charts/.
LIB_CONSUMERS=(apps/yucca-api apps/yucca-admin-api apps/yucca-metrics-worker apps/web apps/michael dev/mock-oidc)
ALL_CHARTS=(apps/yucca-api apps/yucca-admin-api apps/yucca-metrics-worker apps/web apps/michael dev/mock-oidc platform/cnpg-cluster platform/ceph-objectuser platform/rook-ceph-cluster)
echo "==> helm dependency build (yucca-common consumers)" echo "==> helm dependency build (yucca-common consumers)"
for c in "${LIB_CONSUMERS[@]}"; do for c in "${LIB_CONSUMERS[@]}"; do
@@ -20,7 +23,7 @@ done
echo "==> helm template + kubeconform" echo "==> helm template + kubeconform"
for c in "${ALL_CHARTS[@]}"; do for c in "${ALL_CHARTS[@]}"; do
ns=yucca ns=yucca
[ "$c" = "rook-ceph-cluster" ] && ns=rook-ceph [ "$c" = "platform/rook-ceph-cluster" ] && ns=rook-ceph
# NB: release name must not be YAML-boolean-ish ("y"/"on"/...): it lands in # NB: release name must not be YAML-boolean-ish ("y"/"on"/...): it lands in
# labels and kubeconform reads it back as a bool. # labels and kubeconform reads it back as a bool.
helm template yucca "charts/$c" -n "$ns" \ helm template yucca "charts/$c" -n "$ns" \
@@ -29,27 +32,33 @@ for c in "${ALL_CHARTS[@]}"; do
done done
echo "==> kustomize build (Flux entrypoints)" echo "==> kustomize build (Flux entrypoints)"
# o11y-style GitOps tree: per-env cluster entry points + app overlays. # partition/region GitOps tree: per-cluster entry points (clusters/<p>/<r>) +
for env in staging production; do # app overlays (apps/<p>/<r>, which compose components/{infra,roles/<role>}).
kustomize build "kubernetes/clusters/$env" >/dev/null && echo " OK kubernetes/clusters/$env" # LoadRestrictionsNone mirrors how Flux's kustomize-controller builds within a
kustomize build "kubernetes/apps/$env" >/dev/null && echo " OK kubernetes/apps/$env" # single git artifact: the role Components reference ../../apps/<app>.yaml
# (a sibling subtree under components/), which the CLI's default RootOnly
# restrictor would reject even though Flux allows it.
kb() { kustomize build --load-restrictor=LoadRestrictionsNone "$@"; }
CLUSTERS=(staging/austin prod/htz-fsn1 dev/local)
for c in "${CLUSTERS[@]}"; do
kb "kubernetes/clusters/$c" >/dev/null && echo " OK kubernetes/clusters/$c"
kb "kubernetes/apps/$c" >/dev/null && echo " OK kubernetes/apps/$c"
done done
# Dev-mirror tree (consumed by Tilt; kept validated too). # Dev-mirror HelmRepository sources (consumed by Tilt + the dev cluster-repos
kustomize build kubernetes/apps >/dev/null && echo " OK kubernetes/apps (dev tree)" # Kustomization; kept validated too).
kustomize build kubernetes/flux/repos >/dev/null && echo " OK kubernetes/flux/repos" kb kubernetes/apps/dev/local/repos >/dev/null && echo " OK kubernetes/apps/dev/local/repos"
kustomize build kubernetes/flux/cluster >/dev/null && echo " OK kubernetes/flux/cluster"
echo "==> flux-local build (full tree, helm rendering as Flux would)" echo "==> flux-local build (full tree, helm rendering as Flux would)"
# Via uvx so uv supplies a Python matching flux-local's requires-python # Via uvx so uv supplies a Python matching flux-local's requires-python
# (>=3.13) regardless of the host. One retry: flux-local fans out `flux build # (>=3.13) regardless of the host. One retry: flux-local fans out `flux build
# ks` subprocesses which very rarely segfault under load. # ks` subprocesses which very rarely segfault under load.
flux_local() { uvx --from "flux-local==8.2.0" flux-local "$@"; } flux_local() { uvx --from "flux-local==8.2.0" flux-local "$@"; }
# Per-cluster path: staging + production reuse the same Kustomization names # Per-cluster path: clusters reuse the same Kustomization names (cluster-apps,
# (flux-system ns), so flux-local must scope to one cluster at a time. # flux-system ns), so flux-local must scope to one cluster at a time.
for env in staging production; do for c in "${CLUSTERS[@]}"; do
flux_local build all "kubernetes/clusters/$env" --enable-helm --no-enable-dns >/dev/null \ flux_local build all "kubernetes/clusters/$c" --enable-helm --no-enable-dns >/dev/null \
|| flux_local build all "kubernetes/clusters/$env" --enable-helm --no-enable-dns >/dev/null || flux_local build all "kubernetes/clusters/$c" --enable-helm --no-enable-dns >/dev/null
echo " OK flux-local ($env)" echo " OK flux-local ($c)"
done done
echo "k8s surface: ALL VALID" echo "k8s surface: ALL VALID"
+10 -10
View File
@@ -1,13 +1,13 @@
#!/usr/bin/env bash #!/usr/bin/env bash
#MISE description="Converge a site's management hosts with ansible/mgmt (root via the TF provisioning key from 1Password). SITE selects the inventory." #MISE description="Converge a region's management hosts with ansible/mgmt (root via the TF provisioning key from 1Password). REGION selects the inventory."
set -euo pipefail set -euo pipefail
ROOT=$(git rev-parse --show-toplevel) ROOT=$(git rev-parse --show-toplevel)
SITE="${SITE:-htz-fsn1}" REGION="${REGION:-htz-fsn1}"
INV="$ROOT/ansible/mgmt/inventories/$SITE" INV="$ROOT/ansible/mgmt/inventories/$REGION"
# Per-site: only run where an inventory exists (keeps the prod matrix happy). # Per-region: only run where an inventory exists (keeps the prod matrix happy).
if [ ! -d "$INV" ]; then if [ ! -d "$INV" ]; then
echo "mgmt:ansible: no ansible/mgmt inventory for site '$SITE' — skipping." echo "mgmt:ansible: no ansible/mgmt inventory for region '$REGION' — skipping."
exit 0 exit 0
fi fi
@@ -19,18 +19,18 @@ mise run mgmt:render-inventory
ACCT=(); [ -z "${OP_SERVICE_ACCOUNT_TOKEN:-}" ] && ACCT=(--account "${OP_ACCOUNT:-team-futo}") ACCT=(); [ -z "${OP_SERVICE_ACCOUNT_TOKEN:-}" ] && ACCT=(--account "${OP_ACCOUNT:-team-futo}")
# Provisioning private key (root login) -> 0600 temp file. Item name is # Provisioning private key (root login) -> 0600 temp file. Item name is
# site-derived: htz-fsn1 -> HTZ_FSN1_PROVISIONING_SSH_PRIVATE_KEY (set by mgmt.tf). # region-derived: htz-fsn1 -> HTZ_FSN1_PROVISIONING_SSH_PRIVATE_KEY (set by mgmt.tf).
KEY_ITEM="$(printf '%s' "$SITE" | tr 'a-z-' 'A-Z_')_PROVISIONING_SSH_PRIVATE_KEY" KEY_ITEM="$(printf '%s' "$REGION" | tr 'a-z-' 'A-Z_')_PROVISIONING_SSH_PRIVATE_KEY"
KEYF=$(mktemp); chmod 600 "$KEYF"; trap 'rm -f "$KEYF"' EXIT KEYF=$(mktemp); chmod 600 "$KEYF"; trap 'rm -f "$KEYF"' EXIT
op read "${ACCT[@]}" "op://yucca_tf_prod/$KEY_ITEM/password" > "$KEYF" op read "${ACCT[@]}" "op://yucca_tf_prod/$KEY_ITEM/password" > "$KEYF"
# NetBird "mgmt" setup key (auto_groups=["mgmt"]) — joins the node to the overlay # NetBird "mgmt" setup key (auto_groups=["mgmt"]) — joins the node to the overlay
# as a route peer. Site-derived item: htz-fsn1 -> NETBIRD_YUCCA_PROD_HTZ_FSN1_MGMT_SETUP_KEY. # as a route peer. Region-derived item: htz-fsn1 -> NETBIRD_YUCCA_PROD_HTZ_FSN1_MGMT_SETUP_KEY.
NB_KEY_ITEM="NETBIRD_YUCCA_PROD_$(printf '%s' "$SITE" | tr 'a-z-' 'A-Z_')_MGMT_SETUP_KEY" NB_KEY_ITEM="NETBIRD_YUCCA_PROD_$(printf '%s' "$REGION" | tr 'a-z-' 'A-Z_')_MGMT_SETUP_KEY"
NB_SETUP_KEY=$(op read "${ACCT[@]}" "op://yucca_tf_prod/$NB_KEY_ITEM/password") NB_SETUP_KEY=$(op read "${ACCT[@]}" "op://yucca_tf_prod/$NB_KEY_ITEM/password")
cd "$ROOT/ansible/mgmt" cd "$ROOT/ansible/mgmt"
ansible-galaxy collection install -r requirements.yml >/dev/null ansible-galaxy collection install -r requirements.yml >/dev/null
ansible-playbook -i "inventories/$SITE" site.yml \ ansible-playbook -i "inventories/$REGION" site.yml \
--private-key "$KEYF" \ --private-key "$KEYF" \
--extra-vars "mgmt_netbird_setup_key=$NB_SETUP_KEY" "$@" --extra-vars "mgmt_netbird_setup_key=$NB_SETUP_KEY" "$@"
+4 -4
View File
@@ -1,10 +1,10 @@
#!/usr/bin/env bash #!/usr/bin/env bash
#MISE description="Render a site's ansible/mgmt inventory from Terraform (tf/render/ansible-mgmt: addressing + identity + mgmt-hosts.yaml). SITE selects the site." #MISE description="Render a region's ansible/mgmt inventory from Terraform (tf/render/ansible-mgmt: addressing + identity + mgmt-hosts.yaml). REGION selects the region."
set -euo pipefail set -euo pipefail
ROOT=$(git rev-parse --show-toplevel) ROOT=$(git rev-parse --show-toplevel)
SITE="${SITE:-htz-fsn1}" REGION="${REGION:-htz-fsn1}"
cd "$ROOT/tf/render/ansible-mgmt" cd "$ROOT/tf/render/ansible-mgmt"
tofu init -input=false >/dev/null tofu init -input=false >/dev/null
tofu apply -input=false -auto-approve -var "site=$SITE" >/dev/null tofu apply -input=false -auto-approve -var "region=$REGION" >/dev/null
echo "rendered ansible/mgmt/inventories/$SITE/ (hosts.yml, host_vars/, group_vars/all/users.generated.yml)" echo "rendered ansible/mgmt/inventories/$REGION/ (hosts.yml, host_vars/, group_vars/all/users.generated.yml)"
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
#MISE description="Build yuctl"
set -e
cd packages/yuctl && go build -o ../../dist/yuctl .
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
#MISE description="Run yuctl from source (forwards args)"
set -e
cd packages/yuctl && go run . "$@"
+22 -17
View File
@@ -7,7 +7,7 @@
# locally-built images. # locally-built images.
# - HelmRepository-sourced HelmReleases (cnpg, rook, victoria-*) -> installed # - HelmRepository-sourced HelmReleases (cnpg, rook, victoria-*) -> installed
# at the exact chart version + values pinned in the HelmRelease, from the # at the exact chart version + values pinned in the HelmRelease, from the
# HelmRepositories declared in kubernetes/flux/repos/. # HelmRepositories declared in kubernetes/apps/dev/local/repos/.
# APP_WIRING below carries only the dev-specific concerns Flux doesn't have: # APP_WIRING below carries only the dev-specific concerns Flux doesn't have:
# which locally-built image to inject, deploy ordering, and pod-readiness quirks. # which locally-built image to inject, deploy ordering, and pod-readiness quirks.
# #
@@ -228,15 +228,15 @@ docker_build(
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
local_resource( local_resource(
'helm-deps', 'helm-deps',
cmd='rm -rf charts/yucca-api/charts charts/yucca-admin-api/charts charts/yucca-metrics-worker/charts charts/web/charts charts/michael/charts charts/mock-oidc/charts && for d in charts/yucca-api charts/yucca-admin-api charts/yucca-metrics-worker charts/web charts/michael charts/mock-oidc; do (cd $d && helm dependency build); done', cmd='rm -rf charts/apps/yucca-api/charts charts/apps/yucca-admin-api/charts charts/apps/yucca-metrics-worker/charts charts/apps/web/charts charts/apps/michael/charts charts/dev/mock-oidc/charts && for d in charts/apps/yucca-api charts/apps/yucca-admin-api charts/apps/yucca-metrics-worker charts/apps/web charts/apps/michael charts/dev/mock-oidc; do (cd $d && helm dependency build); done',
deps=[ deps=[
'charts/yucca-api', 'charts/apps/yucca-api',
'charts/yucca-admin-api', 'charts/apps/yucca-admin-api',
'charts/yucca-metrics-worker', 'charts/apps/yucca-metrics-worker',
'charts/web', 'charts/apps/web',
'charts/michael', 'charts/apps/michael',
'charts/mock-oidc', 'charts/dev/mock-oidc',
'charts/yucca-common', 'charts/lib/yucca-common',
], ],
# `helm dependency build` rewrites these; if Tilt watches them we re-enter # `helm dependency build` rewrites these; if Tilt watches them we re-enter
# an infinite rebuild loop. # an infinite rebuild loop.
@@ -269,7 +269,7 @@ APP_WIRING = {
# RGW user), so Tilt's pod tracking would hang at "pending". Mark ready on # RGW user), so Tilt's pod tracking would hang at "pending". Mark ready on
# apply; michael still waits on this resource for ordering. # apply; michael still waits on this resource for ordering.
'yucca-object-user': {'build': None, 'deps': ['rook-ceph-cluster'], 'pod_readiness': 'ignore'}, 'yucca-object-user': {'build': None, 'deps': ['rook-ceph-cluster'], 'pod_readiness': 'ignore'},
# Shares charts/ceph-objectuser with yucca-object-user; its userName/caps # Shares charts/platform/ceph-objectuser with yucca-object-user; its userName/caps
# come from the HelmRelease values, so dev must apply them (dev_values) or # come from the HelmRelease values, so dev must apply them (dev_values) or
# both releases would default to userName=michael and collide. # both releases would default to userName=michael and collide.
'yucca-metrics-object-user': {'build': None, 'deps': ['rook-ceph-cluster'], 'pod_readiness': 'ignore', 'dev_values': True}, 'yucca-metrics-object-user': {'build': None, 'deps': ['rook-ceph-cluster'], 'pod_readiness': 'ignore', 'dev_values': True},
@@ -286,9 +286,9 @@ APP_WIRING = {
} }
def discover_helm_repos(): def discover_helm_repos():
"""HelmRepository name -> URL, from kubernetes/flux/repos/.""" """HelmRepository name -> URL, from kubernetes/apps/dev/local/repos/."""
repos = {} repos = {}
for path in listdir('kubernetes/flux/repos'): for path in listdir('kubernetes/apps/dev/local/repos'):
if not path.endswith('.yaml'): if not path.endswith('.yaml'):
continue continue
doc = read_yaml(path) doc = read_yaml(path)
@@ -302,10 +302,10 @@ def discover_apps():
for path in listdir('kubernetes/apps', recursive=True): for path in listdir('kubernetes/apps', recursive=True):
if not path.endswith('/helmrelease.yaml'): if not path.endswith('/helmrelease.yaml'):
continue continue
# The o11y-style GitOps tree (apps/base + apps/{staging,production} # The o11y-style GitOps tree (apps/base + the real-cluster overlays
# overlays) is reconciled by Flux on the real clusters — Tilt deploys # apps/<partition>/<region>) is reconciled by Flux on the real clusters —
# only the local dev-mirror tree (apps/<namespace>/<app>/app/). # Tilt deploys only the local dev-mirror tree under apps/dev/local/.
if '/apps/base/' in path or '/apps/staging/' in path or '/apps/production/' in path: if '/apps/dev/local/' not in path:
continue continue
hr = read_yaml(path) hr = read_yaml(path)
if not hr or hr.get('kind') != 'HelmRelease': if not hr or hr.get('kind') != 'HelmRelease':
@@ -340,6 +340,11 @@ def wiring_for(app):
return wiring return wiring
LOCAL_APPS, REMOTE_APPS = discover_apps() LOCAL_APPS, REMOTE_APPS = discover_apps()
# Guard against a silently-empty deploy: if the dev-mirror allow-list path ever
# moves again, discover_apps() would return nothing and Tilt would come up empty
# instead of failing loudly here.
if not LOCAL_APPS and not REMOTE_APPS:
fail("discover_apps() found no HelmReleases under kubernetes/apps/dev/local/ — has the dev-mirror tree moved?")
HELM_REPOS = discover_helm_repos() HELM_REPOS = discover_helm_repos()
for repo_name, url in HELM_REPOS.items(): for repo_name, url in HELM_REPOS.items():
@@ -386,7 +391,7 @@ for app in LOCAL_APPS:
image_keys=[('image.repository', 'image.tag')] if builds else [], image_keys=[('image.repository', 'image.tag')] if builds else [],
resource_deps=['helm-deps'] + wiring['deps'] + extra_deps, resource_deps=['helm-deps'] + wiring['deps'] + extra_deps,
labels=['app'], labels=['app'],
deps=[app.chart, 'charts/yucca-common'], deps=[app.chart, 'charts/lib/yucca-common'],
pod_readiness=wiring.get('pod_readiness', ''), pod_readiness=wiring.get('pod_readiness', ''),
) )
+9 -8
View File
@@ -11,19 +11,20 @@ hardware/
backups/ backups/
exports/ exports/
# TF-generated artifacts — do not commit; re-render with `tofu apply` in tf/deployment/<env>/ceph/ # TF-generated artifacts — do not commit; re-render with `tofu apply` in tf/deployment/<partition>/<region>/ceph/
inventories/*/inventory.ini # Inventories are region-scoped with a friendly cluster leaf: inventories/<partition>-<region>/<cluster>/
inventories/*/inventory-provision.ini inventories/*/*/inventory.ini
inventories/*/inventory-destroy.ini inventories/*/*/inventory-provision.ini
inventories/*/secrets.yml.tpl inventories/*/*/inventory-destroy.ini
inventories/*/group_vars/all/operators.yml inventories/*/*/secrets.yml.tpl
inventories/*/*/group_vars/all/operators.yml
# host_vars IS committed (per-node hardware facts — stable, part of the inventory # host_vars IS committed (per-node hardware facts — stable, part of the inventory
# source of truth). Operator-local overrides use host_vars/*.local.yml if needed. # source of truth). Operator-local overrides use host_vars/*.local.yml if needed.
inventories/*/host_vars/*.local.yml inventories/*/*/host_vars/*.local.yml
# Rendered installimage scripts (source of truth is the .tpl file). # Rendered installimage scripts (source of truth is the .tpl file).
inventories/*/installimage/post-install.sh inventories/*/*/installimage/post-install.sh
# Python virtualenv (mise setup) and bytecode # Python virtualenv (mise setup) and bytecode
.venv/ .venv/
+5 -5
View File
@@ -8,7 +8,7 @@ PATH = "{{config_root}}/.venv/bin:{{env.PATH}}"
# Note: CEPH_ENV is intentionally NOT declared here. mise's [env] block # Note: CEPH_ENV is intentionally NOT declared here. mise's [env] block
# overrides shell-exported values, which silently sends operators to the # overrides shell-exported values, which silently sends operators to the
# wrong cluster. Operators must export CEPH_ENV once per shell session: # wrong cluster. Operators must export CEPH_ENV once per shell session:
# export CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini # export CEPH_ENV=inventories/staging-austin/sietch/inventory.ini
# Inventory files are TF-generated — `mise run tf:apply` (from yucca root) if missing. # Inventory files are TF-generated — `mise run tf:apply` (from yucca root) if missing.
[tasks.setup] [tasks.setup]
@@ -63,7 +63,7 @@ run = """
set -euo pipefail set -euo pipefail
# Syntax-check only parses YAML — any valid inventory works. Default to # Syntax-check only parses YAML — any valid inventory works. Default to
# sietch when CEPH_ENV isn't inline-prefixed; the parse is identical. # sietch when CEPH_ENV isn't inline-prefixed; the parse is identical.
CEPH_ENV="${CEPH_ENV:-inventories/sietch-ceph.staging.austin.int/inventory.ini}" CEPH_ENV="${CEPH_ENV:-inventories/staging-austin/sietch/inventory.ini}"
for pb in *.yml; do for pb in *.yml; do
case "$pb" in case "$pb" in
requirements.yml|ansible-navigator.yml) continue ;; requirements.yml|ansible-navigator.yml) continue ;;
@@ -123,9 +123,9 @@ description = "Destroy Ceph cluster (requires confirmation)"
run = """ run = """
#!/usr/bin/env bash #!/usr/bin/env bash
set -euo pipefail set -euo pipefail
CEPH_ENV_DIR=$(dirname "$CEPH_ENV") CEPH_ENV_DIR=$(dirname "$CEPH_ENV") # e.g., inventories/staging-austin/sietch
CLUSTER_ID=$(basename "$CEPH_ENV_DIR") # e.g., sietch-ceph.staging.austin.int REGION_SLUG=$(basename "$(dirname "$CEPH_ENV_DIR")") # e.g., staging-austin (<partition>-<region>)
DOMAIN=${CLUSTER_ID#*-ceph.}.futo.cloud # e.g., dev.austin.int.futo.cloud DOMAIN="${REGION_SLUG/-/.}.int.futo.cloud" # e.g., staging.austin.int.futo.cloud
DESTROY_INV="$CEPH_ENV_DIR/inventory-destroy.ini" DESTROY_INV="$CEPH_ENV_DIR/inventory-destroy.ini"
echo "Usage: scripts/ansible-play.sh destroy-ceph.yml \\" echo "Usage: scripts/ansible-play.sh destroy-ceph.yml \\"
echo " -e yes_destroy_ceph=true -e destroy_target_domain=$DOMAIN" echo " -e yes_destroy_ceph=true -e destroy_target_domain=$DOMAIN"
+14 -14
View File
@@ -59,8 +59,8 @@ mise trust
mise run setup mise run setup
# Render cluster inventories + secrets templates (run once, or after any # Render cluster inventories + secrets templates (run once, or after any
# change to tf/deployment/staging/ceph/clusters.auto.tfvars) # change to tf/deployment/staging/austin/ceph/clusters.auto.tfvars)
cd ../../tf/deployment/staging/ceph && tofu init && tofu apply && cd - cd ../../tf/deployment/staging/austin/ceph && tofu init && tofu apply && cd -
``` ```
This runs: This runs:
@@ -83,7 +83,7 @@ It points to an **inventory file** (not a directory):
```bash ```bash
# Inline prefix — required for `mise run` invocations: # Inline prefix — required for `mise run` invocations:
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini mise run preflight CEPH_ENV=inventories/staging-austin/sietch/inventory.ini mise run preflight
``` ```
**`export CEPH_ENV=...` does NOT work with `mise run`.** mise's `[env]` **`export CEPH_ENV=...` does NOT work with `mise run`.** mise's `[env]`
@@ -97,7 +97,7 @@ For multiple commands against the same cluster, set a local (non-exported)
shell variable and inline-prefix each invocation: shell variable and inline-prefix each invocation:
```bash ```bash
CE=inventories/sietch-ceph.staging.austin.int/inventory.ini CE=inventories/staging-austin/sietch/inventory.ini
CEPH_ENV=$CE mise run preflight CEPH_ENV=$CE mise run preflight
CEPH_ENV=$CE mise run status CEPH_ENV=$CE mise run status
CEPH_ENV=$CE mise run deploy CEPH_ENV=$CE mise run deploy
@@ -106,7 +106,7 @@ CEPH_ENV=$CE mise run deploy
Calling scripts directly (e.g., `scripts/preflight.sh`) DOES respect Calling scripts directly (e.g., `scripts/preflight.sh`) DOES respect
shell `export` — it's only the `mise run` path that filters the env. shell `export` — it's only the `mise run` path that filters the env.
Cluster identity is declared in `tf/deployment/staging/ceph/clusters.auto.tfvars` Cluster identity is declared in `tf/deployment/staging/austin/ceph/clusters.auto.tfvars`
(keyed by short cluster name). TF renders the directory name, inventory (keyed by short cluster name). TF renders the directory name, inventory
file, and secrets template from that entry. `CEPH_ENV` is just a pointer file, and secrets template from that entry. `CEPH_ENV` is just a pointer
to the rendered inventory file; wrappers like `scripts/ansible-play.sh` to the rendered inventory file; wrappers like `scripts/ansible-play.sh`
@@ -114,11 +114,11 @@ and the `destroy` mise task extract the cluster name from its path for
convenience: convenience:
``` ```
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini CEPH_ENV=inventories/staging-austin/sietch/inventory.ini
^^^^^^^^^^^^^^^^^^^^^^^^^^^^ ^^^^^^^^^^^^^^ ^^^^^^
inventory dir = sietch-ceph.staging.austin.int (rendered by TF) | cluster name = sietch (map key in clusters.auto.tfvars)
cluster name = sietch (map key in clusters.auto.tfvars) region slug = staging-austin (<partition>-<region>, rendered by TF)
domain = dev.austin.int.futo.cloud (domain field in tfvars) domain = staging.austin.int.futo.cloud (domain field in tfvars)
``` ```
Running `mise run tf:apply` regenerates `inventories/<cluster>/inventory.ini` Running `mise run tf:apply` regenerates `inventories/<cluster>/inventory.ini`
@@ -158,7 +158,7 @@ mise run tf:apply
``` ```
Re-renders `inventory.ini` and `secrets.yml.tpl` for every cluster declared Re-renders `inventory.ini` and `secrets.yml.tpl` for every cluster declared
in `tf/deployment/staging/ceph/clusters.auto.tfvars`. Required after any cluster- in `tf/deployment/staging/austin/ceph/clusters.auto.tfvars`. Required after any cluster-
spec edit. spec edit.
### 5. Dry-run against real nodes ### 5. Dry-run against real nodes
@@ -337,7 +337,7 @@ scripts/ansible-play.sh deploy-ceph.yml --tags bootstrap
| `destroy` | `mise run destroy` | Destroy cluster (interactive confirmation) | | `destroy` | `mise run destroy` | Destroy cluster (interactive confirmation) |
Inventory rendering and secret-item management are TF responsibilities Inventory rendering and secret-item management are TF responsibilities
— `tofu apply` in `tf/deployment/staging/ceph/` renders `inventory.ini` and — `tofu apply` in `tf/deployment/staging/austin/ceph/` renders `inventory.ini` and
`secrets.yml.tpl` for every cluster declared in `clusters.auto.tfvars`. `secrets.yml.tpl` for every cluster declared in `clusters.auto.tfvars`.
Cluster secrets live in `yucca_tf_dev` (see [docs/secrets.md](docs/secrets.md)). Cluster secrets live in `yucca_tf_dev` (see [docs/secrets.md](docs/secrets.md)).
@@ -347,11 +347,11 @@ Cluster secrets live in `yucca_tf_dev` (see [docs/secrets.md](docs/secrets.md)).
yucca/ yucca/
├── tf/ # Terraform state + secrets + rendering (authoritative) ├── tf/ # Terraform state + secrets + rendering (authoritative)
│ ├── shared/modules/ceph-cluster/ # Module: per-cluster orchestration + rendering │ ├── shared/modules/ceph-cluster/ # Module: per-cluster orchestration + rendering
│ └── deployment/staging/ceph/ # Cluster declarations + tofu apply target │ └── deployment/staging/austin/ceph/ # Cluster declarations + tofu apply target
└── ansible/ceph/ # This directory └── ansible/ceph/ # This directory
├── *.yml # Top-level playbooks (site.yml, deploy-ceph.yml, etc.) ├── *.yml # Top-level playbooks (site.yml, deploy-ceph.yml, etc.)
├── inventories/ ├── inventories/
│ └── <cluster>-ceph.<env>.<dc>.<provider>/ │ └── <partition>-<region>/<cluster>/
│ ├── inventory.ini # TF-rendered (gitignored) │ ├── inventory.ini # TF-rendered (gitignored)
│ ├── secrets.yml.tpl # TF-rendered, consumed by op inject (gitignored) │ ├── secrets.yml.tpl # TF-rendered, consumed by op inject (gitignored)
│ ├── group_vars/all/ │ ├── group_vars/all/
+5 -5
View File
@@ -7,9 +7,9 @@ scaffolding are provisioned from `yucca/tf/` (see `../../tf/`).
| Cluster | Domain | Location | Hardware | Nodes | | Cluster | Domain | Location | Hardware | Nodes |
|---------|--------|----------|----------|-------| |---------|--------|----------|----------|-------|
| **sietch** | `dev.austin.int.futo.cloud` | Austin DC | Dell R730xd | 3 | | **sietch** | `staging.austin.int.futo.cloud` | Austin DC | Dell R730xd | 3 |
Clusters are declared in `yucca/tf/deployment/staging/ceph/clusters.auto.tfvars`; Clusters are declared in `yucca/tf/deployment/staging/austin/ceph/clusters.auto.tfvars`;
`tofu apply` renders `inventories/<cluster>/inventory.ini` and `tofu apply` renders `inventories/<cluster>/inventory.ini` and
`secrets.yml.tpl` per cluster. The `CEPH_ENV` variable selects the active `secrets.yml.tpl` per cluster. The `CEPH_ENV` variable selects the active
cluster for any `mise run` or direct ansible invocation. cluster for any `mise run` or direct ansible invocation.
@@ -41,7 +41,7 @@ data flow, and design rationale.
```bash ```bash
# 1. Render cluster inventories + secrets templates (once, from yucca/tf/) # 1. Render cluster inventories + secrets templates (once, from yucca/tf/)
(cd ../../tf/deployment/staging/ceph && tofu init && tofu apply) (cd ../../tf/deployment/staging/austin/ceph && tofu init && tofu apply)
# 2. Set up the ansible side # 2. Set up the ansible side
mise trust && mise run setup # bootstrap dev environment mise trust && mise run setup # bootstrap dev environment
@@ -49,7 +49,7 @@ mise trust && mise run setup # bootstrap dev environment
# 3. Run mise tasks against the target cluster. CEPH_ENV must be set # 3. Run mise tasks against the target cluster. CEPH_ENV must be set
# inline (NOT via `export`) — see docs/scripts.md "Setting CEPH_ENV" # inline (NOT via `export`) — see docs/scripts.md "Setting CEPH_ENV"
# for why mise's [env] block strips shell exports. # for why mise's [env] block strips shell exports.
CE=inventories/sietch-ceph.staging.austin.int/inventory.ini CE=inventories/staging-austin/sietch/inventory.ini
CEPH_ENV=$CE mise run preflight # TF artifacts + 1P + SSH + connectivity CEPH_ENV=$CE mise run preflight # TF artifacts + 1P + SSH + connectivity
CEPH_ENV=$CE mise run status # read-only cluster health check CEPH_ENV=$CE mise run status # read-only cluster health check
CEPH_ENV=$CE mise run drift # configuration drift detection CEPH_ENV=$CE mise run drift # configuration drift detection
@@ -114,7 +114,7 @@ See [CONTRIBUTING.md](CONTRIBUTING.md) for the full development workflow.
| `bench-rados` | RADOS bench (raw cluster I/O) | | `bench-rados` | RADOS bench (raw cluster I/O) |
Inventory scaffolding + secret-item provisioning live in `yucca/tf/` — run Inventory scaffolding + secret-item provisioning live in `yucca/tf/` — run
`tofu apply` in `tf/deployment/staging/ceph/` to (re-)render `tofu apply` in `tf/deployment/staging/austin/ceph/` to (re-)render
`inventories/<cluster>/inventory.ini` and `secrets.yml.tpl`. `inventories/<cluster>/inventory.ini` and `secrets.yml.tpl`.
## Documentation ## Documentation
+3 -3
View File
@@ -11,11 +11,11 @@
# #
# Usage: # Usage:
# scripts/ansible-play.sh destroy-ceph.yml \ # scripts/ansible-play.sh destroy-ceph.yml \
# -e "yes_destroy_ceph=true destroy_target_domain=dev.austin.int.futo.cloud" # -e "yes_destroy_ceph=true destroy_target_domain=staging.austin.int.futo.cloud"
# #
# To run a specific phase only: # To run a specific phase only:
# scripts/ansible-play.sh destroy-ceph.yml -e "yes_destroy_ceph=true destroy_target_domain=dev.austin.int.futo.cloud" --tags purge # scripts/ansible-play.sh destroy-ceph.yml -e "yes_destroy_ceph=true destroy_target_domain=staging.austin.int.futo.cloud" --tags purge
# scripts/ansible-play.sh destroy-ceph.yml -e "yes_destroy_ceph=true destroy_target_domain=dev.austin.int.futo.cloud" --tags cleanup # scripts/ansible-play.sh destroy-ceph.yml -e "yes_destroy_ceph=true destroy_target_domain=staging.austin.int.futo.cloud" --tags cleanup
- name: Destroy Ceph Tentacle cluster - name: Destroy Ceph Tentacle cluster
hosts: ceph_nodes hosts: ceph_nodes
+7 -7
View File
@@ -1,6 +1,6 @@
# Adding a cluster # Adding a cluster
Clusters are declared in `tf/deployment/<env>/ceph/clusters.auto.tfvars`. Clusters are declared in `tf/deployment/<partition>/<region>/ceph/clusters.auto.tfvars`.
Every cluster-scoped concern — inventory file, hostname, 1P item names, SSH Every cluster-scoped concern — inventory file, hostname, 1P item names, SSH
key path, secrets template — is derived from that one entry. Most of what key path, secrets template — is derived from that one entry. Most of what
this walkthrough describes is editing that file and running this walkthrough describes is editing that file and running
@@ -35,7 +35,7 @@ to `ceph`). Every Ceph-project inventory grep-matches `*-ceph.*` regardless
of datacenter or environment. of datacenter or environment.
Existing examples: Existing examples:
- `sietch-ceph.dev.austin.int/` — Austin DC, internal network, dev - `staging-austin/sietch/` — Austin DC, internal network, dev
Future environments land as siblings: `*-ceph.staging.<dc>.<provider>/`, Future environments land as siblings: `*-ceph.staging.<dc>.<provider>/`,
`*-ceph.prod.<dc>.<provider>/`. `*-ceph.prod.<dc>.<provider>/`.
@@ -59,7 +59,7 @@ auto-picked from the 923-word wordlist — see [docs/naming.md](naming.md#host-n
### 2. Declare the cluster in TF ### 2. Declare the cluster in TF
Edit `tf/deployment/<env>/ceph/clusters.auto.tfvars` and add an entry. Edit `tf/deployment/<partition>/<region>/ceph/clusters.auto.tfvars` and add an entry.
Working example for a hypothetical `mesa` cluster at Hetzner Falkenstein: Working example for a hypothetical `mesa` cluster at Hetzner Falkenstein:
```hcl ```hcl
@@ -105,7 +105,7 @@ mise run tf:apply
Or, for a non-default stack: Or, for a non-default stack:
```bash ```bash
TF_STACK_DIR=tf/deployment/<env>/ceph mise run tf:apply TF_STACK_DIR=tf/deployment/<partition>/<region>/ceph mise run tf:apply
``` ```
This creates (per the module's `rendering.tf`): This creates (per the module's `rendering.tf`):
@@ -123,12 +123,12 @@ All of these are gitignored — re-run `mise run tf:apply` after any
Hand-maintained, committed. Copy the closer existing analogue as a starting Hand-maintained, committed. Copy the closer existing analogue as a starting
point: point:
- **Bare-metal cluster:** copy from `sietch-ceph.dev.austin.int/group_vars/all/vars.yml` - **Bare-metal cluster:** copy from `staging-austin/sietch/group_vars/all/vars.yml`
- **Hetzner/single-NIC cluster:** start from the sietch vars and adjust for - **Hetzner/single-NIC cluster:** start from the sietch vars and adjust for
the NVMe-RAID shape (public /32, no bond/ProxyJump, installimage-owned LVM). the NVMe-RAID shape (public /32, no bond/ProxyJump, installimage-owned LVM).
```bash ```bash
cp inventories/sietch-ceph.dev.austin.int/group_vars/all/vars.yml \ cp inventories/staging-austin/sietch/group_vars/all/vars.yml \
inventories/mesa-ceph.dev.fsn.htz/group_vars/all/vars.yml inventories/mesa-ceph.dev.fsn.htz/group_vars/all/vars.yml
``` ```
@@ -322,7 +322,7 @@ CEPH_ENV=inventories/mesa-ceph.dev.fsn.htz/
Default is set in `.mise.toml` (`sietch` in dev). Override per-command: Default is set in `.mise.toml` (`sietch` in dev). Override per-command:
```bash ```bash
CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory.ini mise run status CEPH_ENV=inventories/staging-austin/sietch/inventory.ini mise run status
``` ```
`scripts/ansible-play.sh` derives the secrets template path from `CEPH_ENV` `scripts/ansible-play.sh` derives the secrets template path from `CEPH_ENV`
+10 -10
View File
@@ -51,20 +51,20 @@ its environment from the same source — directory layout — so dev / staging
| Layer | dev (today) | staging (planned) | prod (planned) | | Layer | dev (today) | staging (planned) | prod (planned) |
|--------------|----------------------------------------------------------------|---------------------------------------------------------|------------------------------------------------------| |--------------|----------------------------------------------------------------|---------------------------------------------------------|------------------------------------------------------|
| TF stack dir | `tf/deployment/dev/ceph/` | `tf/deployment/staging/ceph/` | `tf/deployment/prod/ceph/` | | TF stack dir | `tf/deployment/dev/ceph/` | `tf/deployment/staging/austin/ceph/` | `tf/deployment/prod/ceph/` |
| TF state key | `ceph/dev/ceph/terraform.tfstate` | `ceph/staging/ceph/terraform.tfstate` | `ceph/prod/ceph/terraform.tfstate` | | TF state key | `ceph/dev/ceph/terraform.tfstate` | `ceph/staging/ceph/terraform.tfstate` | `ceph/prod/ceph/terraform.tfstate` |
| 1P vaults | `yucca_tf_dev` (live) · `yucca_tf_dev_manual` (human-fillable) | `yucca_tf_staging` · `yucca_tf_staging_manual` | `yucca_tf` (live) · `yucca_tf_prod_manual` | | 1P vaults | `yucca_tf_dev` (live) · `yucca_tf_dev_manual` (human-fillable) | `yucca_tf_staging` · `yucca_tf_staging_manual` | `yucca_tf` (live) · `yucca_tf_prod_manual` |
| Ansible inv | `inventories/<cluster>-ceph.dev.<dc>.<provider>/` | `inventories/<cluster>-ceph.staging.<dc>.<provider>/` | `inventories/<cluster>-ceph.prod.<dc>.<provider>/` | | Ansible inv | `inventories/<cluster>-ceph.dev.<dc>.<provider>/` | `inventories/<cluster>-ceph.staging.<dc>.<provider>/` | `inventories/<cluster>-ceph.prod.<dc>.<provider>/` |
| mise default | `CEPH_ENV=...sietch-ceph.dev.austin.int/inventory.ini` | overridden via env at invocation | overridden via env at invocation | | mise default | `CEPH_ENV=...staging-austin/sietch/inventory.ini` | overridden via env at invocation | overridden via env at invocation |
Today the only deployed environment is dev (sietch). Adding Today the only deployed environment is dev (sietch). Adding
staging/prod is purely additive: create the matching `tf/deployment/<env>/ceph/` staging/prod is purely additive: create the matching `tf/deployment/<partition>/<region>/ceph/`
directory, populate `clusters.auto.tfvars`, and the same module + Ansible directory, populate `clusters.auto.tfvars`, and the same module + Ansible
roles + mise tasks work unchanged. The state backend key path, 1P vault roles + mise tasks work unchanged. The state backend key path, 1P vault
selection, and inventory directory naming all derive from the env segment. selection, and inventory directory naming all derive from the env segment.
`TF_STACK_DIR` is the operator-side override for `mise run tf:*` tasks; it `TF_STACK_DIR` is the operator-side override for `mise run tf:*` tasks; it
defaults to `tf/deployment/staging/ceph` and points at any sibling stack directory. defaults to `tf/deployment/staging/austin/ceph` and points at any sibling stack directory.
`CEPH_ENV` is the matching override for Ansible — points at the rendered `CEPH_ENV` is the matching override for Ansible — points at the rendered
`inventory.ini` for the cluster you intend to operate on. `inventory.ini` for the cluster you intend to operate on.
@@ -164,7 +164,7 @@ Each top-level key becomes a cluster:
```hcl ```hcl
sietch = { sietch = {
domain = "dev.austin.int.futo.cloud" domain = "staging.austin.int.futo.cloud"
environment = "dev" environment = "dev"
datacenter = "austin" datacenter = "austin"
provider_code = "int" provider_code = "int"
@@ -390,7 +390,7 @@ just those phases.
``` ```
inventories/ inventories/
sietch-ceph.staging.austin.int/ Austin staging cluster staging-austin/sietch/ Austin staging cluster
inventory.ini TF-generated, gitignored inventory.ini TF-generated, gitignored
inventory-destroy.ini TF-generated, gitignored inventory-destroy.ini TF-generated, gitignored
inventory-provision.ini TF-generated, gitignored inventory-provision.ini TF-generated, gitignored
@@ -486,15 +486,15 @@ env defaults.
```toml ```toml
[env] [env]
CEPH_ENV = "inventories/sietch-ceph.staging.austin.int/inventory.ini" CEPH_ENV = "inventories/staging-austin/sietch/inventory.ini"
``` ```
`TF_STACK_DIR` defaults to `tf/deployment/staging/ceph` inside each `tf:*` task. `TF_STACK_DIR` defaults to `tf/deployment/staging/austin/ceph` inside each `tf:*` task.
Both are overridable per-invocation: Both are overridable per-invocation:
```bash ```bash
TF_STACK_DIR=tf/deployment/staging/ceph mise run tf:plan TF_STACK_DIR=tf/deployment/staging/austin/ceph mise run tf:plan
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini mise run status CEPH_ENV=inventories/staging-austin/sietch/inventory.ini mise run status
``` ```
--- ---
+1 -1
View File
@@ -21,7 +21,7 @@ A Hetzner NVMe-RAID host would attach over a public /32 with a single NIC
and direct SSH (no bond, no ProxyJump). and direct SSH (no bond, no ProxyJump).
Per-node connection IPs (`bond_ip`) are declared in Per-node connection IPs (`bond_ip`) are declared in
`tf/deployment/staging/ceph/clusters.auto.tfvars` and rendered by TF into the `tf/deployment/staging/austin/ceph/clusters.auto.tfvars` and rendered by TF into the
cluster's `inventory.ini`. They're also mirrored into `host_vars/` for use cluster's `inventory.ini`. They're also mirrored into `host_vars/` for use
by roles that need the IP as a variable (e.g., cephadm public-network by roles that need the IP as a variable (e.g., cephadm public-network
resolution, dashboard URL construction). resolution, dashboard URL construction).
+2 -2
View File
@@ -2,7 +2,7 @@
Every name derivable in this project — hostname, inventory directory, 1P Every name derivable in this project — hostname, inventory directory, 1P
item title, SSH key filename — traces back to one entry per cluster in item title, SSH key filename — traces back to one entry per cluster in
`tf/deployment/<env>/ceph/clusters.auto.tfvars`. TF's `ceph-cluster` module `tf/deployment/<partition>/<region>/ceph/clusters.auto.tfvars`. TF's `ceph-cluster` module
assembles the rest. assembles the rest.
For how naming fits the broader system see For how naming fits the broader system see
@@ -195,7 +195,7 @@ The `-ceph` segment is hardcoded in the ceph-cluster module regardless of
as `*-ceph.*` — even a hypothetical cluster with `role_in_hostname = "osd"` as `*-ceph.*` — even a hypothetical cluster with `role_in_hostname = "osd"`
(hostnames `mesa-osd-*`) still renders `mesa-ceph.prod.fsn.htz/`. (hostnames `mesa-osd-*`) still renders `mesa-ceph.prod.fsn.htz/`.
Defined in `tf/deployment/<env>/ceph/main.tf` (`local.inventory_dirs`). Defined in `tf/deployment/<partition>/<region>/ceph/main.tf` (`local.inventory_dirs`).
## 1Password item naming ## 1Password item naming
+5 -4
View File
@@ -211,9 +211,9 @@ inventory file. Cluster identity is authoritative in
`clusters.auto.tfvars`; `CEPH_ENV` is the runtime pointer. `clusters.auto.tfvars`; `CEPH_ENV` is the runtime pointer.
``` ```
CEPH_ENV = inventories/sietch-ceph.staging.austin.int/inventory.ini CEPH_ENV = inventories/staging-austin/sietch/inventory.ini
| |
dirname -> inventories/sietch-ceph.staging.austin.int dirname -> inventories/staging-austin/sietch
| |
+ "/secrets.yml.tpl" -> op inject input + "/secrets.yml.tpl" -> op inject input
``` ```
@@ -225,8 +225,9 @@ is missing. See [scripts.md](scripts.md) for the full contract.
**Destroy task** in `.mise.toml` extracts the domain for the safety gate: **Destroy task** in `.mise.toml` extracts the domain for the safety gate:
```bash ```bash
CLUSTER_ID=$(basename "$CEPH_ENV_DIR") # sietch-ceph.staging.austin.int CEPH_ENV_DIR=$(dirname "$CEPH_ENV") # inventories/staging-austin/sietch
DOMAIN=${CLUSTER_ID#*-ceph.}.futo.cloud # staging.austin.int.futo.cloud REGION_SLUG=$(basename "$(dirname "$CEPH_ENV_DIR")") # staging-austin (<partition>-<region>)
DOMAIN="${REGION_SLUG/-/.}.int.futo.cloud" # staging.austin.int.futo.cloud
``` ```
--- ---
+5 -5
View File
@@ -15,7 +15,7 @@
Pick an unused name, or omit `name` to let TF auto-pick from the Pick an unused name, or omit `name` to let TF auto-pick from the
923-word list seeded per-cluster (see [docs/naming.md](../naming.md)). 923-word list seeded per-cluster (see [docs/naming.md](../naming.md)).
Edit `tf/deployment/staging/ceph/clusters.auto.tfvars` and append to the Edit `tf/deployment/staging/austin/ceph/clusters.auto.tfvars` and append to the
target cluster's `hosts` list: target cluster's `hosts` list:
```hcl ```hcl
@@ -45,7 +45,7 @@ their positions.
## 2. Create host_vars ## 2. Create host_vars
```bash ```bash
cd inventories/sietch-ceph.staging.austin.int cd inventories/staging-austin/sietch
cp host_vars/example.yml host_vars/sietch-ceph-<name>.yml cp host_vars/example.yml host_vars/sietch-ceph-<name>.yml
``` ```
@@ -67,7 +67,7 @@ hardware:
After `tofu apply` in step 1, inspect the rendered inventory: After `tofu apply` in step 1, inspect the rendered inventory:
```bash ```bash
cat inventories/sietch-ceph.staging.austin.int/inventory.ini cat inventories/staging-austin/sietch/inventory.ini
``` ```
The new host should appear in: The new host should appear in:
@@ -86,7 +86,7 @@ Bootstrap host is unchanged. TF never moves an existing bootstrap assignment.
3. Run provisioning: 3. Run provisioning:
```bash ```bash
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory-provision.ini \ CEPH_ENV=inventories/staging-austin/sietch/inventory-provision.ini \
scripts/ansible-play.sh provision.yml \ scripts/ansible-play.sh provision.yml \
-e confirm_wipe=true \ -e confirm_wipe=true \
--limit sietch-ceph-<name> --limit sietch-ceph-<name>
@@ -98,7 +98,7 @@ CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory-provision.ini \
ssh -i ~/.ssh/id_ed25519_sietch ansible-iac@sietch-ceph-<name> hostname -f ssh -i ~/.ssh/id_ed25519_sietch ansible-iac@sietch-ceph-<name> hostname -f
``` ```
Expected output: `sietch-ceph-<name>.dev.austin.int.futo.cloud` Expected output: `sietch-ceph-<name>.staging.austin.int.futo.cloud`
### Hetzner (remote servers) ### Hetzner (remote servers)
@@ -44,7 +44,7 @@ mise run tf:apply
### Rendered file on disk doesn't match tfvars ### Rendered file on disk doesn't match tfvars
```bash ```bash
cat ansible/ceph/inventories/sietch-ceph.staging.austin.int/inventory.ini cat ansible/ceph/inventories/staging-austin/sietch/inventory.ini
# says ansible_user=root but tfvars says ansible-iac # says ansible_user=root but tfvars says ansible-iac
``` ```
@@ -51,7 +51,7 @@ least-privilege principle.
```bash ```bash
cd ~/yucca/ansible/ceph cd ~/yucca/ansible/ceph
export CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini export CEPH_ENV=inventories/staging-austin/sietch/inventory.ini
mise run status # read-only smoke test mise run status # read-only smoke test
mise run deploy # or any other task mise run deploy # or any other task
``` ```
+1 -1
View File
@@ -46,7 +46,7 @@ the same bond IP via the same DHCP reservation / static config.
# Boot the new box from the Debian 12 live image (same procedure as first # Boot the new box from the Debian 12 live image (same procedure as first
# provision — see docs/runbooks/add-node.md for iDRAC steps). # provision — see docs/runbooks/add-node.md for iDRAC steps).
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory-provision.ini \ CEPH_ENV=inventories/staging-austin/sietch/inventory-provision.ini \
scripts/ansible-play.sh provision.yml \ scripts/ansible-play.sh provision.yml \
-e confirm_wipe=true \ -e confirm_wipe=true \
--limit sietch-ceph-<name> --limit sietch-ceph-<name>
+5 -5
View File
@@ -21,11 +21,11 @@ sudo openssl x509 -in /etc/ceph/rgw-ssl.crt -noout -subject -dates -ext subjectA
Output shows: Output shows:
``` ```
subject=C = US, ST = Texas, L = Austin, O = FUTO, CN = s3.dev.austin.int.futo.cloud subject=C = US, ST = Texas, L = Austin, O = FUTO, CN = s3.staging.austin.int.futo.cloud
notBefore=... notBefore=...
notAfter=... notAfter=...
X509v3 Subject Alternative Name: X509v3 Subject Alternative Name:
DNS:s3.dev.austin.int.futo.cloud, DNS:*.s3.dev.austin.int.futo.cloud, ... DNS:s3.staging.austin.int.futo.cloud, DNS:*.s3.staging.austin.int.futo.cloud, ...
``` ```
## 2. Run the rotation playbook ## 2. Run the rotation playbook
@@ -75,7 +75,7 @@ sudo openssl x509 -in /etc/ceph/rgw-ssl.crt -noout -subject -dates -ext subjectA
```bash ```bash
# From a node in the cluster (self-signed cert) # From a node in the cluster (self-signed cert)
curl -k https://s3.dev.austin.int.futo.cloud:443/ curl -k https://s3.staging.austin.int.futo.cloud:443/
``` ```
Expected: XML response with `ListAllMyBucketsResult` or `AccessDenied` Expected: XML response with `ListAllMyBucketsResult` or `AccessDenied`
@@ -104,7 +104,7 @@ verification disabled for self-signed certs).
## Certificate configuration ## Certificate configuration
The cert parameters are controlled by these variables in The cert parameters are controlled by these variables in
`inventories/sietch-ceph.staging.austin.int/group_vars/all/vars.yml`: `inventories/staging-austin/sietch/group_vars/all/vars.yml`:
| Variable | Default | Purpose | | Variable | Default | Purpose |
|---|---|---| |---|---|---|
@@ -115,7 +115,7 @@ The cert parameters are controlled by these variables in
| `ceph_rgw_ssl_cert_subject_l` | `Austin` | Locality | | `ceph_rgw_ssl_cert_subject_l` | `Austin` | Locality |
| `ceph_rgw_ssl_cert_subject_o` | `FUTO` | Organization | | `ceph_rgw_ssl_cert_subject_o` | `FUTO` | Organization |
| `ceph_rgw_ssl_cert_email` | `yucca@futo.org` | Contact email | | `ceph_rgw_ssl_cert_email` | `yucca@futo.org` | Contact email |
| `ceph_rgw_dns_name` | `s3.dev.austin.int.futo.cloud` | CN and primary SAN | | `ceph_rgw_dns_name` | `s3.staging.austin.int.futo.cloud` | CN and primary SAN |
SANs are auto-generated from inventory: per-node FQDNs and bond IPs are SANs are auto-generated from inventory: per-node FQDNs and bond IPs are
included so direct-host access validates. included so direct-host access validates.
+2 -2
View File
@@ -28,7 +28,7 @@ deploy` picks up the new value via `op inject`.
Replace `<CLUSTER>` with the cluster short name (e.g. `SIETCH`). The active vault name is Replace `<CLUSTER>` with the cluster short name (e.g. `SIETCH`). The active vault name is
declared per-cluster in the `vault` field of the cluster's entry in declared per-cluster in the `vault` field of the cluster's entry in
`tf/deployment/staging/ceph/clusters.auto.tfvars` — `yucca_tf_dev` for dev, `tf/deployment/staging/austin/ceph/clusters.auto.tfvars` — `yucca_tf_dev` for dev,
future `yucca_tf_staging` / `yucca_tf` for staging/prod. future `yucca_tf_staging` / `yucca_tf` for staging/prod.
## 1. Rotate in 1Password ## 1. Rotate in 1Password
@@ -192,7 +192,7 @@ Once the sietch-ceph service account lands and `secrets.tf.disabled` is
re-enabled, rotations become a `terraform taint` + `apply`: re-enabled, rotations become a `terraform taint` + `apply`:
```bash ```bash
cd tf/deployment/staging/ceph cd tf/deployment/staging/austin/ceph
terragrunt taint 'module.cluster["sietch"].onepassword_item.secret["dashboard"]' terragrunt taint 'module.cluster["sietch"].onepassword_item.secret["dashboard"]'
terragrunt apply terragrunt apply
``` ```
+13 -13
View File
@@ -9,8 +9,8 @@ wildcard TLS certificate on **port 443**.
| Style | URL | | Style | URL |
|---|---| |---|---|
| Path-style | `https://s3.dev.austin.int.futo.cloud/<bucket>/<key>` | | Path-style | `https://s3.staging.austin.int.futo.cloud/<bucket>/<key>` |
| Virtual-hosted | `https://<bucket>.s3.dev.austin.int.futo.cloud/<key>` | | Virtual-hosted | `https://<bucket>.s3.staging.austin.int.futo.cloud/<key>` |
| Direct (per-node) | `https://10.10.10.90:443`, `https://10.10.10.91:443`, `https://10.10.10.92:443` | | Direct (per-node) | `https://10.10.10.90:443`, `https://10.10.10.91:443`, `https://10.10.10.92:443` |
Region: **us-east-1** Region: **us-east-1**
@@ -81,7 +81,7 @@ aws_secret_access_key = YOUR_SECRET_KEY
```ini ```ini
[profile sietch] [profile sietch]
region = us-east-1 region = us-east-1
endpoint_url = https://s3.dev.austin.int.futo.cloud endpoint_url = https://s3.staging.austin.int.futo.cloud
s3 = s3 =
signature_version = s3v4 signature_version = s3v4
addressing_style = path addressing_style = path
@@ -121,7 +121,7 @@ urllib3.disable_warnings(urllib3.exceptions.InsecureRequestWarning)
s3 = boto3.client( s3 = boto3.client(
"s3", "s3",
endpoint_url="https://s3.dev.austin.int.futo.cloud", endpoint_url="https://s3.staging.austin.int.futo.cloud",
aws_access_key_id="YOUR_ACCESS_KEY", aws_access_key_id="YOUR_ACCESS_KEY",
aws_secret_access_key="YOUR_SECRET_KEY", aws_secret_access_key="YOUR_SECRET_KEY",
region_name="us-east-1", region_name="us-east-1",
@@ -153,7 +153,7 @@ To use the CA bundle instead of disabling verification:
```python ```python
s3 = boto3.client( s3 = boto3.client(
"s3", "s3",
endpoint_url="https://s3.dev.austin.int.futo.cloud", endpoint_url="https://s3.staging.austin.int.futo.cloud",
aws_access_key_id="YOUR_ACCESS_KEY", aws_access_key_id="YOUR_ACCESS_KEY",
aws_secret_access_key="YOUR_SECRET_KEY", aws_secret_access_key="YOUR_SECRET_KEY",
region_name="us-east-1", region_name="us-east-1",
@@ -170,7 +170,7 @@ s3 = boto3.client(
```bash ```bash
export AWS_ACCESS_KEY_ID="YOUR_ACCESS_KEY" export AWS_ACCESS_KEY_ID="YOUR_ACCESS_KEY"
export AWS_SECRET_ACCESS_KEY="YOUR_SECRET_KEY" export AWS_SECRET_ACCESS_KEY="YOUR_SECRET_KEY"
export RESTIC_REPOSITORY="s3:https://s3.dev.austin.int.futo.cloud/restic-backups" export RESTIC_REPOSITORY="s3:https://s3.staging.austin.int.futo.cloud/restic-backups"
# Init (first time) # Init (first time)
restic init --option s3.region=us-east-1 restic init --option s3.region=us-east-1
@@ -202,16 +202,16 @@ needed from the application side.
## DNS setup for virtual-hosted buckets ## DNS setup for virtual-hosted buckets
Virtual-hosted bucket addressing (e.g., `https://my-bucket.s3.dev.austin.int.futo.cloud`) Virtual-hosted bucket addressing (e.g., `https://my-bucket.s3.staging.austin.int.futo.cloud`)
requires two DNS records: requires two DNS records:
``` ```
s3.dev.austin.int.futo.cloud. A 10.10.10.90 s3.staging.austin.int.futo.cloud. A 10.10.10.90
s3.dev.austin.int.futo.cloud. A 10.10.10.91 s3.staging.austin.int.futo.cloud. A 10.10.10.91
s3.dev.austin.int.futo.cloud. A 10.10.10.92 s3.staging.austin.int.futo.cloud. A 10.10.10.92
*.s3.dev.austin.int.futo.cloud. A 10.10.10.90 *.s3.staging.austin.int.futo.cloud. A 10.10.10.90
*.s3.dev.austin.int.futo.cloud. A 10.10.10.91 *.s3.staging.austin.int.futo.cloud. A 10.10.10.91
*.s3.dev.austin.int.futo.cloud. A 10.10.10.92 *.s3.staging.austin.int.futo.cloud. A 10.10.10.92
``` ```
Round-robin A records across all three nodes. Round-robin A records across all three nodes.
+7 -7
View File
@@ -22,13 +22,13 @@ it inline, never via `export`:**
```bash ```bash
# Correct — inline prefix, applies to one mise/script invocation # Correct — inline prefix, applies to one mise/script invocation
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini mise run preflight CEPH_ENV=inventories/staging-austin/sietch/inventory.ini mise run preflight
# WRONG — mise's [env] machinery silently strips shell-exported vars # WRONG — mise's [env] machinery silently strips shell-exported vars
# when launching tasks; CEPH_ENV reaches an empty environment and the # when launching tasks; CEPH_ENV reaches an empty environment and the
# wrapper exits with "CEPH_ENV must be set". Confusing because your shell # wrapper exits with "CEPH_ENV must be set". Confusing because your shell
# clearly has it set. # clearly has it set.
export CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini export CEPH_ENV=inventories/staging-austin/sietch/inventory.ini
mise run preflight # fails mise run preflight # fails
``` ```
@@ -90,7 +90,7 @@ Common patterns:
```bash ```bash
scripts/ansible-play.sh baseline.yml --check --diff scripts/ansible-play.sh baseline.yml --check --diff
scripts/ansible-play.sh deploy-ceph.yml --tags rgw,monitoring scripts/ansible-play.sh deploy-ceph.yml --tags rgw,monitoring
scripts/ansible-play.sh destroy-ceph.yml -e yes_destroy_ceph=true -e destroy_target_domain=dev.austin.int.futo.cloud scripts/ansible-play.sh destroy-ceph.yml -e yes_destroy_ceph=true -e destroy_target_domain=staging.austin.int.futo.cloud
``` ```
The destroy playbook requires both safety gates: The destroy playbook requires both safety gates:
@@ -113,16 +113,16 @@ The `mise run destroy` task (in `.mise.toml`) builds these arguments automatical
```bash ```bash
# Standard deploy # Standard deploy
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini \ CEPH_ENV=inventories/staging-austin/sietch/inventory.ini \
scripts/ansible-play.sh deploy-ceph.yml scripts/ansible-play.sh deploy-ceph.yml
# Dry-run a role via tags # Dry-run a role via tags
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini \ CEPH_ENV=inventories/staging-austin/sietch/inventory.ini \
scripts/ansible-play.sh site.yml --check --diff --tags baseline scripts/ansible-play.sh site.yml --check --diff --tags baseline
# CI / headless (SA token from env) # CI / headless (SA token from env)
OP_SERVICE_ACCOUNT_TOKEN="$(...)" \ OP_SERVICE_ACCOUNT_TOKEN="$(...)" \
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini \ CEPH_ENV=inventories/staging-austin/sietch/inventory.ini \
scripts/ansible-play.sh status.yml scripts/ansible-play.sh status.yml
``` ```
@@ -263,7 +263,7 @@ Warnings (non-blocking) are reported in the summary but don't affect exit.
mise run preflight mise run preflight
# Direct, against sietch # Direct, against sietch
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini \ CEPH_ENV=inventories/staging-austin/sietch/inventory.ini \
scripts/preflight.sh scripts/preflight.sh
``` ```
+2 -2
View File
@@ -9,7 +9,7 @@ For how secrets fit into the broader architecture, see
```mermaid ```mermaid
flowchart TB flowchart TB
ONEP[("1Password<br/>yucca_tf · yucca_tf_dev · ...<br/><i>source of truth</i>")] ONEP[("1Password<br/>yucca_tf · yucca_tf_dev · ...<br/><i>source of truth</i>")]
TF[Terraform / Tofu<br/>tf/deployment/staging/ceph/] TF[Terraform / Tofu<br/>tf/deployment/staging/austin/ceph/]
REPO[/"inventories/&lt;cluster&gt;/<br/>inventory.ini (TF-gen, gitignored)<br/>secrets.yml.tpl (TF-gen, gitignored)"/] REPO[/"inventories/&lt;cluster&gt;/<br/>inventory.ini (TF-gen, gitignored)<br/>secrets.yml.tpl (TF-gen, gitignored)"/]
WRAP[scripts/ansible-play.sh<br/><i>mktemp + op inject → exec ansible-playbook --extra-vars @tmp</i>] WRAP[scripts/ansible-play.sh<br/><i>mktemp + op inject → exec ansible-playbook --extra-vars @tmp</i>]
ANS[ansible-playbook] ANS[ansible-playbook]
@@ -31,7 +31,7 @@ flowchart TB
Future environments land as siblings: `yucca_tf_staging(_manual)`, Future environments land as siblings: `yucca_tf_staging(_manual)`,
`yucca_tf_prod_manual`. The vault a given cluster reads from is declared `yucca_tf_prod_manual`. The vault a given cluster reads from is declared
per-cluster in `tf/deployment/<env>/ceph/clusters.auto.tfvars` (field per-cluster in `tf/deployment/<partition>/<region>/ceph/clusters.auto.tfvars` (field
`vault`). TF derives item paths from that field at render time; changing `vault`). TF derives item paths from that field at render time; changing
it + `tofu apply` re-renders `secrets.yml.tpl` with the new vault path. it + `tofu apply` re-renders `secrets.yml.tpl` with the new vault path.
+2 -2
View File
@@ -40,8 +40,8 @@ lsblk --output NAME,TYPE,MOUNTPOINT | grep crypt
|---|---| |---|---|
| Protocol | HTTPS (TLS 1.2+) on port 443 | | Protocol | HTTPS (TLS 1.2+) on port 443 |
| Certificate | Self-signed RSA 4096-bit, 10-year validity | | Certificate | Self-signed RSA 4096-bit, 10-year validity |
| CN | `s3.dev.austin.int.futo.cloud` | | CN | `s3.staging.austin.int.futo.cloud` |
| SANs | `s3.dev.austin.int.futo.cloud`, `*.s3.dev.austin.int.futo.cloud`, per-node FQDNs, per-node bond IPs | | SANs | `s3.staging.austin.int.futo.cloud`, `*.s3.staging.austin.int.futo.cloud`, per-node FQDNs, per-node bond IPs |
| Issuer | Self-signed (O=FUTO, L=Austin, ST=Texas, C=US) | | Issuer | Self-signed (O=FUTO, L=Austin, ST=Texas, C=US) |
| Cert location | `/etc/ceph/rgw-ssl.crt` + `/etc/ceph/rgw-ssl.key` on bootstrap node | | Cert location | `/etc/ceph/rgw-ssl.crt` + `/etc/ceph/rgw-ssl.key` on bootstrap node |
| Distribution | cephadm distributes combined PEM to all RGW daemon containers | | Distribution | cephadm distributes combined PEM to all RGW daemon containers |
+1 -1
View File
@@ -510,7 +510,7 @@ ceph_rgw_dns_name: s3.{{ cluster_domain }}
``` ```
This derives the DNS name from `cluster_domain` (e.g. This derives the DNS name from `cluster_domain` (e.g.
`s3.dev.austin.int.futo.cloud`). Sietch defines this explicitly. New `s3.staging.austin.int.futo.cloud`). Sietch defines this explicitly. New
clusters should include it from the start — see clusters should include it from the start — see
[docs/adding-a-cluster.md](adding-a-cluster.md) group_vars template. [docs/adding-a-cluster.md](adding-a-cluster.md) group_vars template.
@@ -7,7 +7,7 @@
gather_facts: false gather_facts: false
vars: vars:
cluster_domain: dev.austin.int.futo.cloud cluster_domain: staging.austin.int.futo.cloud
public_network: 10.10.10.0/24 public_network: 10.10.10.0/24
cluster_network: 10.10.10.0/24 cluster_network: 10.10.10.0/24
ceph_rgw_realm: sietch ceph_rgw_realm: sietch
@@ -15,7 +15,7 @@
ceph_rgw_port: 443 ceph_rgw_port: 443
ceph_rgw_ssl: true ceph_rgw_ssl: true
rgw_ssl_cert_combined_pem: "MOCK_CERT_CONTENT" rgw_ssl_cert_combined_pem: "MOCK_CERT_CONTENT"
ceph_rgw_dns_name: s3.dev.austin.int.futo.cloud ceph_rgw_dns_name: s3.staging.austin.int.futo.cloud
tasks: tasks:
- name: Render hosts.j2 - name: Render hosts.j2
+2 -2
View File
@@ -18,7 +18,7 @@ set -euo pipefail
if [ ! -f "$CEPH_ENV" ]; then if [ ! -f "$CEPH_ENV" ]; then
echo "ansible-play.sh: inventory not found: $CEPH_ENV" >&2 echo "ansible-play.sh: inventory not found: $CEPH_ENV" >&2
echo " Hint: render it — 'terragrunt apply' in tf/deployment/<env>/ceph/, then scripts/render-inventories.sh <env>." >&2 echo " Hint: render it — 'terragrunt apply' in tf/deployment/<partition>/<region>/ceph/, then scripts/render-inventories.sh <partition> <region>." >&2
exit 1 exit 1
fi fi
@@ -27,7 +27,7 @@ TEMPLATE="$CEPH_ENV_DIR/secrets.yml.tpl"
if [ ! -f "$TEMPLATE" ]; then if [ ! -f "$TEMPLATE" ]; then
echo "ansible-play.sh: secrets template not found: $TEMPLATE" >&2 echo "ansible-play.sh: secrets template not found: $TEMPLATE" >&2
echo " Hint: render it — 'terragrunt apply' in tf/deployment/<env>/ceph/, then scripts/render-inventories.sh <env>." >&2 echo " Hint: render it — 'terragrunt apply' in tf/deployment/<partition>/<region>/ceph/, then scripts/render-inventories.sh <partition> <region>." >&2
exit 1 exit 1
fi fi
+1 -1
View File
@@ -5,7 +5,7 @@
# #
# Target inventory + secrets template are located via CEPH_ENV (path to # Target inventory + secrets template are located via CEPH_ENV (path to
# the TF-rendered inventory file — TF is the authoritative source of # the TF-rendered inventory file — TF is the authoritative source of
# cluster identity, declared in tf/deployment/<env>/ceph/clusters.auto.tfvars). # cluster identity, declared in tf/deployment/<partition>/<region>/ceph/clusters.auto.tfvars).
set -euo pipefail set -euo pipefail
: "${CEPH_ENV:?CEPH_ENV must be set to the target cluster inventory.ini}" : "${CEPH_ENV:?CEPH_ENV must be set to the target cluster inventory.ini}"
+10 -4
View File
@@ -14,15 +14,21 @@
# reflects the current cluster spec). This script is read-only against state. # reflects the current cluster spec). This script is read-only against state.
# #
# Usage (from anywhere): # Usage (from anywhere):
# ansible/ceph/scripts/render-inventories.sh [env] # env defaults to dev # ansible/ceph/scripts/render-inventories.sh [partition] [region]
# # partition defaults to staging, region defaults to austin (sietch's home)
#
# The TF `render` output carries each cluster's own `dirname` (region-scoped,
# friendly-cluster leaf — e.g. `staging-austin/sietch`), so this wrapper does
# not derive the inventory layout; it only locates the partition/region stack.
set -euo pipefail set -euo pipefail
ENVIRONMENT="${1:-dev}" PARTITION="${1:-staging}"
REGION="${2:-austin}"
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
ANSIBLE_CEPH="$(cd "$SCRIPT_DIR/.." && pwd)" # ansible/ceph ANSIBLE_CEPH="$(cd "$SCRIPT_DIR/.." && pwd)" # ansible/ceph
REPO_ROOT="$(cd "$ANSIBLE_CEPH/../.." && pwd)" # repo root of THIS checkout REPO_ROOT="$(cd "$ANSIBLE_CEPH/../.." && pwd)" # repo root of THIS checkout
STACK_DIR="$REPO_ROOT/tf/deployment/${ENVIRONMENT}/ceph" STACK_DIR="$REPO_ROOT/tf/deployment/${PARTITION}/${REGION}/ceph"
INV_ROOT="$ANSIBLE_CEPH/inventories" INV_ROOT="$ANSIBLE_CEPH/inventories"
[ -d "$STACK_DIR" ] || { [ -d "$STACK_DIR" ] || {
@@ -54,4 +60,4 @@ for cluster, spec in data.items():
print(f"wrote {p}") print(f"wrote {p}")
PY PY
echo "render-inventories: done (${ENVIRONMENT})." echo "render-inventories: done (${PARTITION}/${REGION})."
+5 -5
View File
@@ -45,9 +45,9 @@ truth and writes these **gitignored** files:
| file | generated from | | file | generated from |
|---|---| |---|---|
| `inventories/<site>/hosts.yml` | `mgmt-hosts.yaml` (host names + public IPs) | | `inventories/<region>/hosts.yml` | `mgmt-hosts.yaml` (host names + public IPs) |
| `inventories/<site>/host_vars/*.yml` | `mgmt-hosts.yaml` (NIC) + `fabric-addressing` (VLAN ids/addresses, subnet route) | | `inventories/<region>/host_vars/*.yml` | `mgmt-hosts.yaml` (NIC) + `fabric-addressing` (VLAN ids/addresses, subnet route) |
| `inventories/<site>/group_vars/all/users.generated.yml` | `tf/shared/modules/identity` (`server`-mapped users) | | `inventories/<region>/group_vars/all/users.generated.yml` | `tf/shared/modules/identity` (`server`-mapped users) |
Only `group_vars/all/main.yml` (static config) and `roles/**` are committed. To Only `group_vars/all/main.yml` (static config) and `roles/**` are committed. To
change hosts, addresses, or users, edit the Terraform sources — never the change hosts, addresses, or users, edit the Terraform sources — never the
@@ -82,7 +82,7 @@ shred -u /tmp/htz-fsn1-prov-key
optional. The inventory connects as `ansible_user: root`. optional. The inventory connects as `ansible_user: root`.
The whole render-key + run flow above is wrapped by `mise run mgmt:ansible` The whole render-key + run flow above is wrapped by `mise run mgmt:ansible`
(`SITE` selects the inventory; defaults to `htz-fsn1`), which CI also runs on (`REGION` selects the inventory; defaults to `htz-fsn1`), which CI also runs on
every **prod** apply — the `Ansible converge (mgmt hosts)` step of every **prod** apply — the `Ansible converge (mgmt hosts)` step of
`.github/workflows/infra.yml`, right after the Terraform apply. It's idempotent `.github/workflows/infra.yml`, right after the Terraform apply. It's idempotent
and reaches the hosts over their public IP, so it requires them to already be and reaches the hosts over their public IP, so it requires them to already be
@@ -143,5 +143,5 @@ Nothing secret is committed. SSH public keys are public data and are TF-generate
into `group_vars/all/users.generated.yml`. Runtime secrets are passed via `op read`: into `group_vars/all/users.generated.yml`. Runtime secrets are passed via `op read`:
- Provisioning private key: `op://yucca_tf_prod/HTZ_FSN1_PROVISIONING_SSH_PRIVATE_KEY/password` - Provisioning private key: `op://yucca_tf_prod/HTZ_FSN1_PROVISIONING_SSH_PRIVATE_KEY/password`
- NetBird mgmt setup key: `op://yucca_tf_prod/NETBIRD_YUCCA_PROD_<SITE>_MGMT_SETUP_KEY/password` - NetBird mgmt setup key: `op://yucca_tf_prod/NETBIRD_YUCCA_PROD_<REGION>_MGMT_SETUP_KEY/password`
(the site's reusable `mgmt` key, `auto_groups=["mgmt"]`; minted by the prod netbird stack) (the site's reusable `mgmt` key, `auto_groups=["mgmt"]`; minted by the prod netbird stack)
@@ -10,5 +10,5 @@ timezone: UTC
mgmt_domain: fsn.htz.futo.cloud mgmt_domain: fsn.htz.futo.cloud
# NetBird "mgmt" setup key — NOT hardcoded; passed at run time by # NetBird "mgmt" setup key — NOT hardcoded; passed at run time by
# `mise run mgmt:ansible` from op://yucca_tf_prod/NETBIRD_YUCCA_PROD_<SITE>_MGMT_SETUP_KEY. # `mise run mgmt:ansible` from op://yucca_tf_prod/NETBIRD_YUCCA_PROD_<REGION>_MGMT_SETUP_KEY.
mgmt_netbird_setup_key: "" mgmt_netbird_setup_key: ""
+1 -1
View File
@@ -5,5 +5,5 @@ mgmt_netbird_management_url: "https://api.netbird.io"
# Setup key — NEVER hardcoded. Passed at run time (op:// ref, see README). The # Setup key — NEVER hardcoded. Passed at run time (op:// ref, see README). The
# site's reusable "mgmt" key (auto_groups=["mgmt"]) so the node joins the mgmt # site's reusable "mgmt" key (auto_groups=["mgmt"]) so the node joins the mgmt
# group and becomes a route peer for the site subnets (10.40.5.0/24, api, cluster # group and becomes a route peer for the site subnets (10.40.5.0/24, api, cluster
# nets — the routed network is defined in TF: tf/deployment/prod/<site>/netbird). # nets — the routed network is defined in TF: tf/deployment/prod/<region>/netbird).
mgmt_netbird_setup_key: "" mgmt_netbird_setup_key: ""
+1 -1
View File
@@ -23,7 +23,7 @@
- mgmt_netbird_setup_key | length > 0 - mgmt_netbird_setup_key | length > 0
fail_msg: >- fail_msg: >-
mgmt_netbird_setup_key is empty. Pass it at run time (mgmt:ansible does this mgmt_netbird_setup_key is empty. Pass it at run time (mgmt:ansible does this
from op://yucca_tf_prod/NETBIRD_YUCCA_PROD_<SITE>_MGMT_SETUP_KEY). from op://yucca_tf_prod/NETBIRD_YUCCA_PROD_<REGION>_MGMT_SETUP_KEY).
# mgmt nodes route the site subnets to the overlay. # mgmt nodes route the site subnets to the overlay.
- name: Enable IP forwarding (NetBird route peer) - name: Enable IP forwarding (NetBird route peer)
+1 -1
View File
@@ -8,7 +8,7 @@
# #
# Canonical entrypoint is `mise run mgmt:ansible` (renders the inventory from # Canonical entrypoint is `mise run mgmt:ansible` (renders the inventory from
# Terraform, then runs this). It passes the provisioning key + the NetBird # Terraform, then runs this). It passes the provisioning key + the NetBird
# "mgmt" setup key (op://yucca_tf_prod/NETBIRD_YUCCA_PROD_<SITE>_MGMT_SETUP_KEY). # "mgmt" setup key (op://yucca_tf_prod/NETBIRD_YUCCA_PROD_<REGION>_MGMT_SETUP_KEY).
# #
# NOTE: the networkd role (25G VLAN sub-interfaces) is gated on # NOTE: the networkd role (25G VLAN sub-interfaces) is gated on
# mgmt_networkd_enabled (default false) because the 25G fabric link is # mgmt_networkd_enabled (default false) because the 25G fabric link is
+4 -4
View File
@@ -8,8 +8,8 @@ PATH = "{{config_root}}/.venv/bin:{{env.PATH}}"
# Note: TALOS_ENV is intentionally NOT declared here. mise's [env] block # Note: TALOS_ENV is intentionally NOT declared here. mise's [env] block
# overrides shell-exported values, which silently sends operators to the # overrides shell-exported values, which silently sends operators to the
# wrong cluster. Operators must export TALOS_ENV once per shell session: # wrong cluster. Operators must export TALOS_ENV once per shell session:
# export TALOS_ENV=inventories/sietch-talos.dev.austin.int/inventory.ini # export TALOS_ENV=inventories/staging-austin/inventory.ini
# Inventory file is TF-rendered — `TF_STACK_DIR=tf/deployment/dev/talos mise run tf:apply` # Inventory file is TF-rendered — `TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply`
# (from yucca root) if missing. # (from yucca root) if missing.
[tasks.setup] [tasks.setup]
@@ -35,7 +35,7 @@ yamllint --version
echo "" echo ""
echo "Environment ready. Run 'mise trust' if prompted." echo "Environment ready. Run 'mise trust' if prompted."
echo "Next: 'TF_STACK_DIR=tf/deployment/dev/talos mise run tf:apply' from yucca root to render inventory." echo "Next: 'TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply' from yucca root to render inventory."
""" """
[tasks.lint] [tasks.lint]
@@ -64,7 +64,7 @@ run = """
set -euo pipefail set -euo pipefail
# Syntax-check only parses YAML — any valid inventory works. Default to # Syntax-check only parses YAML — any valid inventory works. Default to
# sietch-talos when TALOS_ENV isn't inline-prefixed; the parse is identical. # sietch-talos when TALOS_ENV isn't inline-prefixed; the parse is identical.
TALOS_ENV="${TALOS_ENV:-inventories/sietch-talos.dev.austin.int/inventory.ini}" TALOS_ENV="${TALOS_ENV:-inventories/staging-austin/inventory.ini}"
# Fall back to inventory.example.ini before tf:apply has rendered the runtime one. # Fall back to inventory.example.ini before tf:apply has rendered the runtime one.
if [ ! -f "$TALOS_ENV" ]; then if [ ! -f "$TALOS_ENV" ]; then
EXAMPLE="$(dirname "$TALOS_ENV")/inventory.example.ini" EXAMPLE="$(dirname "$TALOS_ENV")/inventory.example.ini"
+6 -6
View File
@@ -8,7 +8,7 @@ hardware.
This subtree is the **Ansible substrate**: it provisions the VLAN 50/51 This subtree is the **Ansible substrate**: it provisions the VLAN 50/51
bridges, the libvirt/KVM stack, and stages the Talos VMs — stopping when bridges, the libvirt/KVM stack, and stages the Talos VMs — stopping when
the VMs are running and ready for `talosctl`. The Terraform half in the VMs are running and ready for `talosctl`. The Terraform half in
[tf/deployment/dev/talos/](../../tf/deployment/dev/talos/) renders the [tf/deployment/staging/austin/talos/](../../tf/deployment/staging/austin/talos/) renders the
Ansible inventory and drives cluster bring-up (machine config, bootstrap, Ansible inventory and drives cluster bring-up (machine config, bootstrap,
kubeconfig). kubeconfig).
@@ -31,7 +31,7 @@ cluster secrets out of Ansible's fact cache keeps the blast radius small.
- `talosctl gen config / apply-config / bootstrap` — the Terraform - `talosctl gen config / apply-config / bootstrap` — the Terraform
`siderolabs/talos` provider owns these `siderolabs/talos` provider owns these
([tf/deployment/dev/talos/](../../tf/deployment/dev/talos/)); ([tf/deployment/staging/austin/talos/](../../tf/deployment/staging/austin/talos/));
`docs/operator-handoff.md` keeps the manual sequence for recovery. `docs/operator-handoff.md` keeps the manual sequence for recovery.
- Persistent storage for VMs (RBD-backed boot disks, ceph-csi). Boot - Persistent storage for VMs (RBD-backed boot disks, ceph-csi). Boot
disks are local qcow2 today; RBD lands in a follow-up. disks are local qcow2 today; RBD lands in a follow-up.
@@ -54,19 +54,19 @@ validation tool, selected explicitly with `-e profile=smoke`.
## Operator workflow ## Operator workflow
Inventory is **TF-rendered** from `tf/deployment/dev/talos/`. Run TF first Inventory is **TF-rendered** from `tf/deployment/staging/austin/talos/`. Run TF first
(from the repo root): (from the repo root):
```bash ```bash
TF_STACK_DIR=tf/deployment/dev/talos mise run tf:init # first time only TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:init # first time only
TF_STACK_DIR=tf/deployment/dev/talos mise run tf:apply # renders inventory + host_vars TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply # renders inventory + host_vars
``` ```
Then point at the rendered inventory and use the talos-subtree tasks: Then point at the rendered inventory and use the talos-subtree tasks:
```bash ```bash
cd ansible/talos/ cd ansible/talos/
export TALOS_ENV=inventories/sietch-talos.dev.austin.int/inventory.ini export TALOS_ENV=inventories/staging-austin/inventory.ini
mise run setup # first time only: python venv + ansible collections mise run setup # first time only: python venv + ansible collections
mise run lint # yamllint + ansible-lint + shellcheck mise run lint # yamllint + ansible-lint + shellcheck
+4 -4
View File
@@ -17,7 +17,7 @@ Provisioning split across the monorepo:
| Subtree | Owns | | Subtree | Owns |
|---|---| |---|---|
| `ansible/talos/` (this) | Hypervisor substrate: bridges, libvirt, image, VM definitions. Stops at "VMs ready"; Terraform owns bootstrap (the subtree README explains why the boundary sits here). | | `ansible/talos/` (this) | Hypervisor substrate: bridges, libvirt, image, VM definitions. Stops at "VMs ready"; Terraform owns bootstrap (the subtree README explains why the boundary sits here). |
| `tf/shared/modules/talos-cluster/modules/inventory-renderer/` | Renders the Ansible inventory + (future) `secrets.yml.tpl` from `tf/deployment/dev/talos/clusters.auto.tfvars`. Parity with the ceph-cluster module. | | `tf/shared/modules/talos-cluster/modules/inventory-renderer/` | Renders the Ansible inventory + (future) `secrets.yml.tpl` from `tf/deployment/staging/austin/talos/clusters.auto.tfvars`. Parity with the ceph-cluster module. |
| `tf/shared/modules/talos-cluster/modules/talos-bootstrap/` | `siderolabs/talos` provider — machine_secrets, configuration_apply per node, bootstrap, kubeconfig. Drives the talosctl sequence so operators don't run it by hand. | | `tf/shared/modules/talos-cluster/modules/talos-bootstrap/` | `siderolabs/talos` provider — machine_secrets, configuration_apply per node, bootstrap, kubeconfig. Drives the talosctl sequence so operators don't run it by hand. |
## Physical layout ## Physical layout
@@ -77,7 +77,7 @@ tailnet/VPN subnet route, or being directly on the VLAN) before
## Naming scheme ## Naming scheme
Monorepo FQDN: `<cluster>-<role>-<name>.<domain>` — Monorepo FQDN: `<cluster>-<role>-<name>.<domain>` —
e.g. `sietch-talos-cp1.dev.austin.int.futo.cloud`. The e.g. `sietch-talos-cp1.staging.austin.int.futo.cloud`. The
`<cluster>-<role>` prefix (`sietch-talos`) is the `<cluster>-<role>` prefix (`sietch-talos`) is the
`talos_domain_prefix` group_vars value; `<name>` is the entry in `talos_domain_prefix` group_vars value; `<name>` is the entry in
each host's `host_vars/*.yml` talos_vms list. each host's `host_vars/*.yml` talos_vms list.
@@ -167,8 +167,8 @@ pulls). The TF bootstrap sets
`machine.network.nameservers: [1.1.1.1, 8.8.8.8]` on Talos VMs for `machine.network.nameservers: [1.1.1.1, 8.8.8.8]` on Talos VMs for
upstream image-registry resolution. Kubeconfig uses upstream image-registry resolution. Kubeconfig uses
the CP VIP IP directly (`https://10.50.0.10:6443`). DNS records the CP VIP IP directly (`https://10.50.0.10:6443`). DNS records
under `*.compute.dev.austin.int.futo.cloud` and under `*.compute.staging.austin.int.futo.cloud` and
`*.services.dev.austin.int.futo.cloud` land with the DNS-layer `*.services.staging.austin.int.futo.cloud` land with the DNS-layer
follow-up (LB → ingress → external-dns). follow-up (LB → ingress → external-dns).
## Out of scope ## Out of scope
+11 -11
View File
@@ -22,12 +22,12 @@ Sietch hypervisors. The TF stack owns this flow;
```bash ```bash
cd <yucca> cd <yucca>
TF_STACK_DIR=tf/deployment/dev/talos mise run tf:init # first time TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:init # first time
TF_STACK_DIR=tf/deployment/dev/talos mise run tf:apply TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply
``` ```
What lands: What lands:
- `inventories/sietch-talos.dev.austin.int/inventory.ini` and per-host - `inventories/staging-austin/inventory.ini` and per-host
`host_vars/*.yml` (the `talos_vms` lists), both rendered from `host_vars/*.yml` (the `talos_vms` lists), both rendered from
`clusters.auto.tfvars` — `nodes[]` is the single source of truth for `clusters.auto.tfvars` — `nodes[]` is the single source of truth for
topology, so never hand-edit `host_vars`. topology, so never hand-edit `host_vars`.
@@ -51,7 +51,7 @@ re-run the apply in step 3.
```bash ```bash
cd <yucca>/ansible/talos cd <yucca>/ansible/talos
export TALOS_ENV=inventories/sietch-talos.dev.austin.int/inventory.ini export TALOS_ENV=inventories/staging-austin/inventory.ini
# Optional first time: bootstrap python venv + ansible collections # Optional first time: bootstrap python venv + ansible collections
mise run setup mise run setup
@@ -106,7 +106,7 @@ the desktop app (or `OP_SERVICE_ACCOUNT_TOKEN` if set).
ip route get 10.50.0.10 ip route get 10.50.0.10
cd <yucca> cd <yucca>
TF_STACK_DIR=tf/deployment/dev/talos mise run tf:apply TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply
# Watch for: # Watch for:
# - module.cluster["sietch"].module.talos_bootstrap.talos_machine_configuration_apply.controlplane["cp1"]: Creating... # - module.cluster["sietch"].module.talos_bootstrap.talos_machine_configuration_apply.controlplane["cp1"]: Creating...
# - ... .talos_machine_bootstrap.this: Creating... # - ... .talos_machine_bootstrap.this: Creating...
@@ -124,10 +124,10 @@ mkdir -p ~/.kube ~/.talos
# Use `output -json` alone — `-raw -json` are mutually exclusive and # Use `output -json` alone — `-raw -json` are mutually exclusive and
# yield an empty file. op-run.sh resolves the S3 state creds. # yield an empty file. op-run.sh resolves the S3 state creds.
tf/op-run.sh terragrunt \ tf/op-run.sh terragrunt \
--working-dir tf/deployment/dev/talos output -json kubeconfigs \ --working-dir tf/deployment/staging/austin/talos output -json kubeconfigs \
| jq -r .sietch > ~/.kube/sietch-talos.config | jq -r .sietch > ~/.kube/sietch-talos.config
tf/op-run.sh terragrunt \ tf/op-run.sh terragrunt \
--working-dir tf/deployment/dev/talos output -json talosconfigs \ --working-dir tf/deployment/staging/austin/talos output -json talosconfigs \
| jq -r .sietch > ~/.talos/sietch-talos.config | jq -r .sietch > ~/.talos/sietch-talos.config
export KUBECONFIG=~/.kube/sietch-talos.config export KUBECONFIG=~/.kube/sietch-talos.config
@@ -171,7 +171,7 @@ talosctl reset --graceful --reboot \
# deletes the rendered inventory.ini this play needs. (If you ran them in # deletes the rendered inventory.ini this play needs. (If you ran them in
# the wrong order: cp inventory.example.ini inventory.ini and re-run.) # the wrong order: cp inventory.example.ini inventory.ini and re-run.)
cd <yucca>/ansible/talos cd <yucca>/ansible/talos
export TALOS_ENV=inventories/sietch-talos.dev.austin.int/inventory.ini export TALOS_ENV=inventories/staging-austin/inventory.ini
mise run destroy-vms mise run destroy-vms
# Destroy TF cluster state, LAST (kubeconfig/talosconfig outputs blanked, # Destroy TF cluster state, LAST (kubeconfig/talosconfig outputs blanked,
@@ -180,7 +180,7 @@ mise run destroy-vms
# wiped/destroyed nodes and aborts with "cluster health check failed". # wiped/destroyed nodes and aborts with "cluster health check failed".
# (Destroying the talos resources is state-only — nothing is un-applied.) # (Destroying the talos resources is state-only — nothing is un-applied.)
cd <yucca> cd <yucca>
TF_STACK_DIR=tf/deployment/dev/talos mise run tf:destroy -- -refresh=false TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:destroy -- -refresh=false
``` ```
To redeploy from clean: repeat steps 1 → 2 → 3 → 4. Static IPs are To redeploy from clean: repeat steps 1 → 2 → 3 → 4. Static IPs are
@@ -196,7 +196,7 @@ production topology:
# 1. Restore the production profile in TF input — if you ran smoke, the # 1. Restore the production profile in TF input — if you ran smoke, the
# tfvars still says `profile = "smoke"` and the TF re-apply would # tfvars still says `profile = "smoke"` and the TF re-apply would
# bootstrap nothing new: # bootstrap nothing new:
# tf/deployment/dev/talos/clusters.auto.tfvars → profile = "full" # tf/deployment/staging/austin/talos/clusters.auto.tfvars → profile = "full"
# 2. Provision the remaining VMs on lawson + samara (cp2/worker2, # 2. Provision the remaining VMs on lawson + samara (cp2/worker2,
# cp3/worker3). Existing cp1/worker1 are a no-op. # cp3/worker3). Existing cp1/worker1 are a no-op.
@@ -206,7 +206,7 @@ mise run provision # profile=full is the default
# 3. Re-apply TF for the 4 new node bootstraps (cp2/cp3 join etcd, # 3. Re-apply TF for the 4 new node bootstraps (cp2/cp3 join etcd,
# worker2/worker3 join the cluster) # worker2/worker3 join the cluster)
cd ../.. cd ../..
TF_STACK_DIR=tf/deployment/dev/talos mise run tf:apply TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply
``` ```
NOTE: going from 1 CP (smoke) to 3 CP (full) grows the etcd quorum. NOTE: going from 1 CP (smoke) to 3 CP (full) grows the etcd quorum.
+2 -2
View File
@@ -33,7 +33,7 @@ set `profile = "smoke"` in `clusters.auto.tfvars` before `tf:apply`:
mise run provision -- -e profile=smoke mise run provision -- -e profile=smoke
# TF bootstrap — set profile = "smoke" in clusters.auto.tfvars first # TF bootstrap — set profile = "smoke" in clusters.auto.tfvars first
TF_STACK_DIR=tf/deployment/dev/talos mise run tf:apply TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply
``` ```
Everything else (static `ip=` addressing, factory image, direct kernel Everything else (static `ip=` addressing, factory image, direct kernel
@@ -50,7 +50,7 @@ laurel, so the limit costs nothing.
```bash ```bash
cd ansible/talos cd ansible/talos
export TALOS_ENV=inventories/sietch-talos.dev.austin.int/inventory.ini export TALOS_ENV=inventories/staging-austin/inventory.ini
# A. preflight only — read-only, zero state change # A. preflight only — read-only, zero state change
scripts/ansible-play.sh preflight.yml --limit sietch-ceph-laurel scripts/ansible-play.sh preflight.yml --limit sietch-ceph-laurel
@@ -1,126 +0,0 @@
---
# Cluster-wide variables for the sietch-talos cluster.
#
# Most operators editing this file want one of:
# - talos_version — bump Talos release
# - talos_vm_mac_prefix — change MAC OUI / locally-administered prefix
# - vlan_*_* — tweak network details
# ─── Cluster identity ───────────────────────────────────────────────────
# Short cluster name. Also drives talos_domain_prefix below (and the
# TF-side 1P secret prefixes). A second cluster would get its own
# inventory dir with a different talos_cluster_name.
talos_cluster_name: "sietch"
# Libvirt domain prefix for Talos VMs on this cluster. talos_vms role
# reads this to compose `<prefix>-<vm-name>` (e.g. sietch-talos-cp1).
# Aligns with the monorepo FQDN scheme `<cluster>-<role>-<name>.<domain>`.
talos_domain_prefix: "{{ talos_cluster_name }}-talos"
# Cluster FQDN suffix (where the K8s API + node FQDNs land). The TF
# bootstrap uses this to seed kubeconfig endpoints and Talos cert SANs.
talos_domain: "dev.austin.int.futo.cloud"
# ─── Talos version + image (Image Factory) ──────────────────────────────
# Pinned for reproducibility. Bump deliberately; refresh all three
# checksums (image + kernel + initramfs) with the version.
talos_version: "1.13.3"
# Factory schematic — extensions baked into the boot assets:
# siderolabs/{qemu-guest-agent, util-linux-tools}. ID is sha256 of the
# schematic; regenerate at https://factory.talos.dev and refresh checksums
# if the set changes.
talos_schematic_id: "a7bcadbc1b6d03c0e687be3a5d9789ef7113362a6a1a038653dfd16283a92b6b"
talos_factory_base: "https://factory.talos.dev/image/{{ talos_schematic_id }}/v{{ talos_version }}"
# Disk image — qcow2 overlays back onto this pre-seeded raw (partitions
# present, no install step). Boots to maintenance mode until TF applies config.
talos_image_url: "{{ talos_factory_base }}/metal-amd64.raw.zst"
talos_image_sha256: "c7515f1076db75f6ec4287b4b5663ea0d490363a0d49b04e61fe96bfe71b2cc6"
# Kernel + initramfs — staged and direct-booted by libvirt (not on-disk
# GRUB) to inject a per-VM `ip=` for static first-boot addressing.
# Upgrades: re-stage assets + reboot (the disk kernel is bypassed).
talos_kernel_url: "{{ talos_factory_base }}/kernel-amd64"
talos_initramfs_url: "{{ talos_factory_base }}/initramfs-amd64.xz"
talos_kernel_sha256: "c36d4aac36d081af8016ff9874f29115bf1493d23e556b99d355316afa0e98ff"
talos_initramfs_sha256: "d613bd10a0b7742e17705e0643ec62e19150522a018fb27855535c75918eabc7"
talos_image_dir: "/var/lib/libvirt/images"
# Filenames keyed on version + short schematic hash so a schematic change at
# the same version yields new filenames and re-stages, instead of serving
# the stale cached image. default('stock') when no schematic.
talos_image_tag: "v{{ talos_version }}-{{ (talos_schematic_id | default('stock', true))[:12] }}"
talos_base_image_path: "{{ talos_image_dir }}/talos-{{ talos_image_tag }}.raw"
talos_kernel_path: "{{ talos_image_dir }}/talos-vmlinuz-{{ talos_image_tag }}"
talos_initramfs_path: "{{ talos_image_dir }}/talos-initramfs-{{ talos_image_tag }}.xz"
# ─── Static addressing (libvirt kernel `ip=`) ────────────────────────────
# Per-VM static_ip is in host_vars; gateway/mask/interface are VLAN-50
# constants. Gateway hardcoded (first usable in 10.50.0.0/16) to avoid an
# ansible.utils dependency; TF derives the same via cidrhost().
talos_vms_network_gateway: "10.50.0.1"
talos_vms_network_mask: "255.255.0.0"
talos_vms_network_interface: "enp1s0"
# ─── VM MAC prefix ──────────────────────────────────────────────────────
# SideroLabs has no registered OUI (verified against IEEE registry +
# maclookup.app; Talos prescribes no convention). Choosing 52:54:00:50:XX:XX:
# - 52:54:00 — QEMU/KVM OUI; widely understood by network tools as
# virtual-machine traffic and matches libvirt defaults.
# - 50 — fixed distinguisher; literal hex "50" mnemonically
# matches VLAN 50, separating Talos VMs from any other
# QEMU VMs that might land on these hosts later.
# - last 2 bytes — computed in talos_vms role from a deterministic
# hash of vm_name so the same VM gets the same MAC
# across reprovisions (stable identity at the switch;
# workers' VLAN-51 DHCP leases stay pinned).
talos_vm_mac_prefix: "52:54:00:50"
# Secondary MAC prefix for the VLAN-51 Services NIC on worker VMs.
# Same scheme — 52:54:00 QEMU OUI, fourth octet "51" for VLAN-51
# mnemonic, last two bytes from the same per-vm-name sha1 hash as the
# primary NIC. Workers thus end up with two related but distinct MACs
# (e.g., worker3 → 52:54:00:50:5a:cd on VLAN 50, 52:54:00:51:5a:cd on
# VLAN 51), which is easy to correlate at switch / DHCP server level.
# CPs do not get a Services NIC; only used when item.role == 'worker'.
talos_vm_mac_prefix_services: "52:54:00:51"
# ─── VLANs ──────────────────────────────────────────────────────────────
# Both VLANs are tagged upstream at the switch and arrive on the host
# as tagged frames over bond0. Roles add bridge + VLAN sub-interfaces;
# they don't touch the physical bond.
vlan_compute_id: 50
vlan_compute_name: "compute"
vlan_compute_bridge: "br-vlan50"
vlan_compute_subnet: "10.50.0.0/16"
vlan_compute_dhcp_start: "10.50.0.16"
vlan_compute_dhcp_end: "10.50.4.255"
vlan_services_id: 51
vlan_services_name: "services"
vlan_services_bridge: "br-vlan51"
vlan_services_subnet: "10.51.0.0/16"
# ─── Talos cluster topology ─────────────────────────────────────────────
# Reserved (static) IPs sit OUTSIDE the DHCP range above so leases
# can't collide. VIP is the documented control-plane endpoint that
# operators put in their kubeconfig.
talos_cp_vip: "10.50.0.10"
# ─── libvirt host config ────────────────────────────────────────────────
# Storage pool that holds base image + per-VM overlays + (later)
# snapshots. Lives on the OS root by default; switch to RBD in a
# follow-up iteration.
libvirt_storage_pool_name: "talos"
libvirt_storage_pool_path: "{{ talos_image_dir }}"
# ─── Profile selector ───────────────────────────────────────────────────
# `full` = 3 CP + 3 workers (production, the default) — one CP + one
# worker per hypervisor, so each host is its own failure domain
# for both the control plane and the worker pool.
# `smoke` = 1 CP + 1 worker (both on laurel) — single-host validation for
# substrate changes; select explicitly with `-e profile=smoke`.
# Per-host VM lists are declared in host_vars and gated by this.
talos_profile: "full"
@@ -1,26 +0,0 @@
; Example inventory for sietch-talos. The RUNTIME inventory.ini in this
; same directory is TF-rendered (tf/shared/modules/talos-cluster/
; modules/inventory-renderer/) and gitignored. This .example file is
; checked in as documentation of the expected shape.
;
; To render the runtime inventory:
; TF_STACK_DIR=tf/deployment/dev/talos mise run tf:apply
;
; Sietch hypervisors. The 3 nodes are co-located in Austin, run Ceph
; OSDs on bare metal, and host Talos VMs as guests.
;
; ansible_host = bond0 IP (10.10.10.0/24 management/storage VLAN, the
; existing Ceph cluster network — NOT the new Talos VLAN 50).
; If 10.10.10.0/24 isn't directly reachable from your workstation, put any
; needed SSH bastion / ProxyJump in your ~/.ssh/config.
[hypervisors]
sietch-ceph-laurel ansible_host=10.10.10.90
sietch-ceph-lawson ansible_host=10.10.10.91
sietch-ceph-samara ansible_host=10.10.10.92
[hypervisors:vars]
ansible_user=ansible-iac
ansible_ssh_private_key_file=~/.ssh/id_ed25519_sietch
ansible_ssh_common_args='-o ControlMaster=auto -o ControlPersist=60s'
ansible_python_interpreter=/usr/bin/python3
@@ -5,7 +5,7 @@ talos_vms_enabled: false
# Libvirt domain + qcow2 + NVRAM file name prefix. # Libvirt domain + qcow2 + NVRAM file name prefix.
# Default scopes to "sietch-talos" so VMs end up as sietch-talos-cp1, # Default scopes to "sietch-talos" so VMs end up as sietch-talos-cp1,
# sietch-talos-worker1, etc. — aligns with the monorepo FQDN scheme # sietch-talos-worker1, etc. — aligns with the monorepo FQDN scheme
# `<cluster>-<role>-<name>.<domain>` (e.g. sietch-talos-cp1.dev.austin.int.futo.cloud) # `<cluster>-<role>-<name>.<domain>` (e.g. sietch-talos-cp1.staging.austin.int.futo.cloud)
# and matches the *_TALOS_* 1P secret-prefix convention. Read from # and matches the *_TALOS_* 1P secret-prefix convention. Read from
# group_vars `talos_domain_prefix` when present so a future second # group_vars `talos_domain_prefix` when present so a future second
# cluster can override without role edits. # cluster can override without role edits.
@@ -7,4 +7,4 @@ appVersion: "0.0.1"
dependencies: dependencies:
- name: yucca-common - name: yucca-common
version: 0.1.0 version: 0.1.0
repository: "file://../yucca-common" repository: "file://../../lib/yucca-common"
@@ -7,4 +7,4 @@ appVersion: "0.0.1"
dependencies: dependencies:
- name: yucca-common - name: yucca-common
version: 0.1.0 version: 0.1.0
repository: "file://../yucca-common" repository: "file://../../lib/yucca-common"
@@ -7,4 +7,4 @@ appVersion: "0.0.1"
dependencies: dependencies:
- name: yucca-common - name: yucca-common
version: 0.1.0 version: 0.1.0
repository: "file://../yucca-common" repository: "file://../../lib/yucca-common"
@@ -7,4 +7,4 @@ appVersion: "0.0.1"
dependencies: dependencies:
- name: yucca-common - name: yucca-common
version: 0.1.0 version: 0.1.0
repository: "file://../yucca-common" repository: "file://../../lib/yucca-common"
@@ -7,4 +7,4 @@ appVersion: "0.0.1"
dependencies: dependencies:
- name: yucca-common - name: yucca-common
version: 0.1.0 version: 0.1.0
repository: "file://../yucca-common" repository: "file://../../lib/yucca-common"
@@ -7,4 +7,4 @@ appVersion: "0.0.1"
dependencies: dependencies:
- name: yucca-common - name: yucca-common
version: 0.1.0 version: 0.1.0
repository: "file://../yucca-common" repository: "file://../../lib/yucca-common"

Some files were not shown because too many files have changed in this diff Show More