mirror of
https://github.com/immich-app/yucca.git
synced 2026-09-30 13:33:00 +08:00
feat(all): introduce partition/region/ceph-cluster model across the stack (#222)
* feat: introduce partition/region/ceph-cluster model across the stack
Formalize partition -> region -> {one k8s cluster, many ceph clusters} and
thread it through every layer plus a new yuctl ops CLI.
- tf: deployment/<partition>/<region>/<stack> layout; terragrunt path-parse +
state key yucca/<partition>/<region>/<stack>; per-region region.hcl (role,
site_id, datacenter, provider_code, domain); env->partition / site->region
renames (NetBird object names byte-identical); standardized per-stack
`discovery` output contract (secrets as op:// refs).
- k8s: clusters/<partition>/<region>/ (staging/austin, prod/htz-fsn1, dev/local);
role-based kustomize components (primary/secondary); hybrid cluster-settings
(TF-rendered identity + human fragment); dev-mirror folded into dev/local;
charts regrouped into charts/{apps,platform,lib,dev}.
- ci: infra.yml partition/region discovery matrix; partition-keyed path filters;
<partition>-<region> environment gates; image-versions path moves.
- ansible: inventories under <partition>-<region>/<cluster>.
- yuctl: Go/cobra CLI reading the discovery contract from TF state.
- Retire the sietch-talos libvirt VM cluster (dev@local is the k3d cluster);
ceph inventory_dirname -> <partition>-<region>/<cluster>.
Verified: mise k8s:validate green (3 clusters); yuctl go build/vet; tofu
validate pre-merge (all 9 stacks). Live-staging state migration NOT run.
* fix typo
* commit
This commit is contained in:
@@ -4,8 +4,9 @@ description: >-
|
|||||||
read from 1Password. Requires the 1Password CLI on PATH and
|
read from 1Password. Requires the 1Password CLI on PATH and
|
||||||
OP_SERVICE_ACCOUNT_TOKEN in the environment (the infra jobs provide both). The
|
OP_SERVICE_ACCOUNT_TOKEN in the environment (the infra jobs provide both). The
|
||||||
setup key must already exist in 1Password — it is minted by the matching
|
setup key must already exist in 1Password — it is minted by the matching
|
||||||
netbird stack (deployment/<env>/netbird or prod/<site>/netbird), so apply that
|
netbird stack (deployment/<partition>/global/netbird for the account/partition
|
||||||
stack before this action runs on a fresh bootstrap.
|
key, or deployment/<partition>/<region>/netbird for a per-site key), so apply
|
||||||
|
that stack before this action runs on a fresh bootstrap.
|
||||||
|
|
||||||
inputs:
|
inputs:
|
||||||
setup-key-ref:
|
setup-key-ref:
|
||||||
|
|||||||
@@ -9,7 +9,7 @@
|
|||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
|
|
||||||
TAG="${1:?usage: promote-prod.sh <tag>}"
|
TAG="${1:?usage: promote-prod.sh <tag>}"
|
||||||
FILE="kubernetes/clusters/production/image-versions.yaml"
|
FILE="kubernetes/clusters/prod/htz-fsn1/image-versions.yaml"
|
||||||
|
|
||||||
# Set data.YUCCA_IMAGE_TAG (yq if present, else portable sed on the one line).
|
# Set data.YUCCA_IMAGE_TAG (yq if present, else portable sed on the one line).
|
||||||
if command -v yq >/dev/null 2>&1; then
|
if command -v yq >/dev/null 2>&1; then
|
||||||
|
|||||||
@@ -6,7 +6,7 @@ on:
|
|||||||
# The gated prod promotion commits the pin back to main; ignore that path
|
# The gated prod promotion commits the pin back to main; ignore that path
|
||||||
# so it can't retrigger the workflow (it's also marked [skip ci]).
|
# so it can't retrigger the workflow (it's also marked [skip ci]).
|
||||||
paths-ignore:
|
paths-ignore:
|
||||||
- kubernetes/clusters/production/image-versions.yaml
|
- kubernetes/clusters/prod/htz-fsn1/image-versions.yaml
|
||||||
workflow_dispatch:
|
workflow_dispatch:
|
||||||
|
|
||||||
concurrency:
|
concurrency:
|
||||||
|
|||||||
+238
-330
@@ -1,39 +1,53 @@
|
|||||||
name: Infra (Terraform)
|
name: Infra (Terraform)
|
||||||
|
|
||||||
# Applies the Terraform stacks from CI, path-scoped so each group only runs when
|
# Applies the Terragrunt stacks from CI, path-scoped so each partition only runs
|
||||||
# its own files change (a `changes` job emits per-area booleans that gate the
|
# when its own files change. Connectivity to the bare-metal/private nodes is over
|
||||||
# rest; workflow_dispatch overrides and runs everything). Connectivity to the
|
# the NetBird overlay (Tailscale fully retired).
|
||||||
# bare-metal/private nodes is over the NetBird overlay (Tailscale fully retired):
|
|
||||||
#
|
#
|
||||||
# Staging (tf/deployment/staging/*): ceph, talos, dns, netbird. One staging SA,
|
# Topology model: partition (prod | staging | dev) → region (htz-fsn1, austin,
|
||||||
# a single `staging-infra` Environment gate. The netbird stack mints the CI
|
# local, plus the reserved `global` pseudo-region for partition-wide stacks) →
|
||||||
# setup key; the apply joins the overlay as a `ci` peer to reach the
|
# stack. Stacks live at tf/deployment/<partition>/<region>/<stack>/terragrunt.hcl
|
||||||
# 10.10.10.0/24 nodes, then converges the bare-metal Ceph cluster (ansible/ceph).
|
# (e.g. staging/austin/ceph, staging/global/{dns,netbird}, prod/htz-fsn1/{fabric,
|
||||||
|
# netbird}, prod/global). A single `discover` job scans that tree into a matrix of
|
||||||
|
# {partition, region, stack, dir, order}, so adding a stack/region needs no
|
||||||
|
# workflow edit — only its terragrunt.hcl and a matching GitHub Environment.
|
||||||
#
|
#
|
||||||
# Prod fabric+mgmt (tf/deployment/prod/<site>): the switch fabric + mgmt hosts
|
# Per-partition selection (Actions can't index secrets dynamically, hence the
|
||||||
# (junos-qfx + hetzner providers, built locally). Per-site `prod-<site>` gate.
|
# ternaries):
|
||||||
# Reaches the switch vme / mgmt hosts over NetBird (the mgmt nodes are the
|
# - staging: SA secrets OP_TF_YUCCA_STAGING_ENV(_WRITE), env-file tf/.env.
|
||||||
# route peers for 10.40.5.0/24 et al.); runs the `infra:*` / `mgmt:*` mise tasks.
|
# - prod: SA secrets OP_TF_YUCCA_PROD_ENV(_WRITE), env-file tf/.env.prod.
|
||||||
|
# The prod `fabric` stack applies with the READ SA and self-escalates to the
|
||||||
|
# write SA in-vault via `mise run infra:apply`; the prod netbird stacks
|
||||||
|
# (global + site) apply with the WRITE SA directly.
|
||||||
|
# - dev is local-only (no remote state, no CI service account) → never emitted.
|
||||||
#
|
#
|
||||||
# Prod NetBird (tf/deployment/prod/global + prod/<site>/netbird): account-wide +
|
# Apply ordering (preserved via a per-entry `order` + max-parallel: 1 on a sorted
|
||||||
# site NetBird groups/keys/policies/routes. Pure api.netbird.io. `prod-infra` gate.
|
# matrix): global-region stacks first (account NetBird + DNS), then site NetBird,
|
||||||
|
# then node-touching stacks (ceph/talos/fabric). NetBird must precede the
|
||||||
|
# node-touching stacks because it mints the CI setup key into 1Password and
|
||||||
|
# advertises the node subnets the overlay-joining jobs reach; prod/global must
|
||||||
|
# precede prod/htz-fsn1/netbird (a real terragrunt dependency).
|
||||||
|
#
|
||||||
|
# Environment gates rekey to <partition>-<region> (one per stack's region):
|
||||||
|
# staging-austin, staging-global, prod-global, prod-htz-fsn1. Each matrix apply
|
||||||
|
# entry references its own gate, so an unprovisioned Environment hangs the apply.
|
||||||
#
|
#
|
||||||
# Prerequisites (provisioned out-of-band):
|
# Prerequisites (provisioned out-of-band):
|
||||||
# - Repo secrets: OP_TF_YUCCA_STAGING_ENV (+ _WRITE) — staging SAs; and
|
# - Repo secrets: OP_TF_YUCCA_STAGING_ENV (+ _WRITE); OP_TF_YUCCA_PROD_ENV
|
||||||
# OP_TF_YUCCA_PROD_ENV (read) / OP_TF_YUCCA_PROD_ENV_WRITE (netbird apply) — prod
|
# (read) + OP_TF_YUCCA_PROD_ENV_WRITE (netbird apply / write escalation source).
|
||||||
# SAs. (The fabric apply escalates to the write SA stored in yucca_tf_prod.)
|
# - GitHub Environments with required reviewers: staging-austin, staging-global,
|
||||||
|
# prod-global, prod-htz-fsn1.
|
||||||
# - BOOTSTRAP — the netbird stacks applied ONCE out-of-band so the CI/mgmt setup
|
# - BOOTSTRAP — the netbird stacks applied ONCE out-of-band so the CI/mgmt setup
|
||||||
# keys exist in 1P before anything tries to connect (CI can't mint them itself:
|
# keys exist in 1P before anything tries to connect (CI can't mint them itself:
|
||||||
# the apply that mints them is gated behind the plan that needs them). E.g.:
|
# the apply that mints them is gated behind the plan that needs them). E.g.:
|
||||||
# OP_SERVICE_ACCOUNT_TOKEN=<staging write SA> \
|
# OP_SERVICE_ACCOUNT_TOKEN=<staging write SA> \
|
||||||
# TF_STACK_DIR=tf/deployment/staging/netbird mise run tf:apply
|
# tf/op-run.sh terragrunt --working-dir tf/deployment/staging/global/netbird apply
|
||||||
# (staging: NETBIRD_YUCCA_STAGING_CI_SETUP_KEY; prod: NETBIRD_YUCCA_PROD_<SITE>_*).
|
# (staging: NETBIRD_YUCCA_STAGING_CI_SETUP_KEY — partition-scoped, minted by the
|
||||||
# Also: NetBird routes advertising the node subnets (staging 10.10.10.0/24; prod
|
# account/global netbird; prod: NETBIRD_YUCCA_PROD_<REGION>_* — region-scoped,
|
||||||
# the site subnets via the mgmt peers) that the `ci` group may reach.
|
# minted by the per-site netbird). Also: NetBird routes advertising the node
|
||||||
|
# subnets (staging 10.10.10.0/24; prod the site subnets via the mgmt peers).
|
||||||
# - 1P items: NET_SWITCHES_TERRAFORM_SSH_PRIVATE_KEY (yucca_tf_prod),
|
# - 1P items: NET_SWITCHES_TERRAFORM_SSH_PRIVATE_KEY (yucca_tf_prod),
|
||||||
# NETBOX_API_TOKEN (yucca_tf), HETZNER_WEBSERVICE_API_USER/PASSWORD (yucca_tf_prod).
|
# NETBOX_API_TOKEN (yucca_tf), HETZNER_WEBSERVICE_API_USER/PASSWORD (yucca_tf_prod).
|
||||||
# - GitHub Environments with required reviewers: `staging-infra`, `prod-infra`,
|
|
||||||
# and one `prod-<site>` per prod fabric site (e.g. prod-htz-fsn1).
|
|
||||||
|
|
||||||
on:
|
on:
|
||||||
push:
|
push:
|
||||||
@@ -60,7 +74,7 @@ permissions:
|
|||||||
contents: read
|
contents: read
|
||||||
|
|
||||||
jobs:
|
jobs:
|
||||||
# ── Which areas changed? Outputs gate every downstream job. ──────────────────
|
# ── Which partitions changed? Partition-keyed; `shared` forces all. ──────────
|
||||||
changes:
|
changes:
|
||||||
name: Detect changes
|
name: Detect changes
|
||||||
# Skip on fork PRs (no access to secrets / the overlay anyway).
|
# Skip on fork PRs (no access to secrets / the overlay anyway).
|
||||||
@@ -68,9 +82,9 @@ jobs:
|
|||||||
runs-on: ubuntu-latest
|
runs-on: ubuntu-latest
|
||||||
outputs:
|
outputs:
|
||||||
staging: ${{ steps.filter.outputs.staging }}
|
staging: ${{ steps.filter.outputs.staging }}
|
||||||
prod_tf: ${{ steps.filter.outputs.prod_tf }}
|
prod: ${{ steps.filter.outputs.prod }}
|
||||||
prod_ansible: ${{ steps.filter.outputs.prod_ansible }}
|
dev: ${{ steps.filter.outputs.dev }}
|
||||||
prod_netbird: ${{ steps.filter.outputs.prod_netbird }}
|
shared: ${{ steps.filter.outputs.shared }}
|
||||||
steps:
|
steps:
|
||||||
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
||||||
with:
|
with:
|
||||||
@@ -81,379 +95,273 @@ jobs:
|
|||||||
filters: |
|
filters: |
|
||||||
staging:
|
staging:
|
||||||
- 'tf/deployment/staging/**'
|
- 'tf/deployment/staging/**'
|
||||||
- 'tf/shared/**'
|
|
||||||
- 'tf/op-run.sh'
|
|
||||||
- 'ansible/ceph/**'
|
- 'ansible/ceph/**'
|
||||||
- '.github/actions/netbird-connect/**'
|
prod:
|
||||||
- '.github/workflows/infra.yml'
|
|
||||||
prod_tf:
|
|
||||||
# Fabric/mgmt stacks. Netbird-only dirs are covered by prod_netbird;
|
|
||||||
# an overlap just means a (gated) extra fabric plan — harmless.
|
|
||||||
- 'tf/deployment/prod/**'
|
- 'tf/deployment/prod/**'
|
||||||
- '!tf/deployment/prod/global/**'
|
- 'ansible/mgmt/**'
|
||||||
- '!tf/deployment/prod/*/netbird/**'
|
- 'tf/render/**'
|
||||||
- 'tf/shared/**'
|
|
||||||
- 'tf/providers/**'
|
- 'tf/providers/**'
|
||||||
- '.mise/tasks/infra/**'
|
- '.mise/tasks/infra/**'
|
||||||
- '.mise/tasks/fabric/**'
|
- '.mise/tasks/fabric/**'
|
||||||
- '.mise/tasks/mgmt/**'
|
- '.mise/tasks/mgmt/**'
|
||||||
|
dev:
|
||||||
|
# Local-only (k3d/Tilt): emits no CI stacks, kept for completeness so
|
||||||
|
# a `shared` force still reasons about every partition uniformly.
|
||||||
|
- 'tf/deployment/dev/**'
|
||||||
|
shared:
|
||||||
|
# Cross-partition surfaces — a hit forces every partition's matrix.
|
||||||
|
- 'tf/shared/**'
|
||||||
|
- 'tf/op-run.sh'
|
||||||
- '.mise/config.toml'
|
- '.mise/config.toml'
|
||||||
- '.github/actions/netbird-connect/**'
|
- '.github/actions/netbird-connect/**'
|
||||||
- '.github/workflows/infra.yml'
|
- '.github/workflows/infra.yml'
|
||||||
prod_ansible:
|
|
||||||
- 'ansible/mgmt/**'
|
|
||||||
- 'tf/render/**'
|
|
||||||
- 'tf/shared/modules/identity/**'
|
|
||||||
- 'tf/shared/modules/fabric-addressing/**'
|
|
||||||
- 'tf/deployment/prod/*/mgmt-hosts.yaml'
|
|
||||||
- '.mise/tasks/mgmt/**'
|
|
||||||
- '.mise/config.toml'
|
|
||||||
- '.github/actions/netbird-connect/**'
|
|
||||||
- '.github/workflows/infra.yml'
|
|
||||||
prod_netbird:
|
|
||||||
- 'tf/deployment/prod/global/**'
|
|
||||||
- 'tf/deployment/prod/*/netbird/**'
|
|
||||||
- 'tf/shared/modules/netbird-env/**'
|
|
||||||
- '.github/workflows/infra.yml'
|
|
||||||
|
|
||||||
# ── Staging stacks ──────────────────────────────────────────────────────────
|
# ── Discover the stack matrix from the deployment tree ───────────────────────
|
||||||
staging-plan:
|
discover:
|
||||||
name: Staging plan ${{ matrix.stack }}
|
name: Discover stacks
|
||||||
needs: changes
|
needs: changes
|
||||||
if: needs.changes.outputs.staging == 'true' || github.event_name == 'workflow_dispatch'
|
if: >-
|
||||||
|
needs.changes.outputs.staging == 'true'
|
||||||
|
|| needs.changes.outputs.prod == 'true'
|
||||||
|
|| needs.changes.outputs.shared == 'true'
|
||||||
|
|| github.event_name == 'workflow_dispatch'
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
outputs:
|
||||||
|
matrix: ${{ steps.gen.outputs.matrix }}
|
||||||
|
has_stacks: ${{ steps.gen.outputs.has_stacks }}
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
||||||
|
with:
|
||||||
|
persist-credentials: false
|
||||||
|
- id: gen
|
||||||
|
name: Build {partition, region, stack} matrix for the changed partitions
|
||||||
|
env:
|
||||||
|
STAGING: ${{ needs.changes.outputs.staging }}
|
||||||
|
PROD: ${{ needs.changes.outputs.prod }}
|
||||||
|
SHARED: ${{ needs.changes.outputs.shared }}
|
||||||
|
DISPATCH: ${{ github.event_name == 'workflow_dispatch' }}
|
||||||
|
run: |
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
# Active partitions: own filter OR a shared/manual force. dev is
|
||||||
|
# local-only and never emitted (no remote state / no CI SA).
|
||||||
|
declare -A active=()
|
||||||
|
if [ "$SHARED" = "true" ] || [ "$DISPATCH" = "true" ]; then
|
||||||
|
active[staging]=1; active[prod]=1
|
||||||
|
else
|
||||||
|
[ "$STAGING" = "true" ] && active[staging]=1
|
||||||
|
[ "$PROD" = "true" ] && active[prod]=1
|
||||||
|
fi
|
||||||
|
|
||||||
|
entries=()
|
||||||
|
while IFS= read -r tg; do
|
||||||
|
rel=${tg#tf/deployment/}
|
||||||
|
rel=${rel%/terragrunt.hcl}
|
||||||
|
[ "$rel" = "terragrunt.hcl" ] && continue # repo-root parent config
|
||||||
|
IFS=/ read -r -a segs <<< "$rel"
|
||||||
|
[ "${#segs[@]}" -ge 2 ] || continue
|
||||||
|
partition=${segs[0]}
|
||||||
|
region=${segs[1]}
|
||||||
|
if [ "${#segs[@]}" -ge 3 ]; then
|
||||||
|
stack=$(IFS=/; echo "${segs[*]:2}") # nested sub-stacks keep working
|
||||||
|
else
|
||||||
|
stack=$region # n==2 transition guard (e.g. prod/global)
|
||||||
|
fi
|
||||||
|
[ -n "${active[$partition]:-}" ] || continue
|
||||||
|
|
||||||
|
# Apply ordering: global-region stacks (account netbird/dns) first,
|
||||||
|
# then site netbird, then node-touching stacks (ceph/talos/fabric).
|
||||||
|
if [ "$region" = "global" ]; then order=0
|
||||||
|
elif [ "$stack" = "netbird" ]; then order=1
|
||||||
|
else order=2; fi
|
||||||
|
|
||||||
|
entries+=("$(jq -nc \
|
||||||
|
--arg p "$partition" --arg r "$region" --arg s "$stack" \
|
||||||
|
--arg d "$rel" --argjson o "$order" \
|
||||||
|
'{partition:$p,region:$r,stack:$s,dir:$d,order:$o}')")
|
||||||
|
done < <(find tf/deployment -mindepth 2 -name terragrunt.hcl -type f | sort)
|
||||||
|
|
||||||
|
if [ "${#entries[@]}" -eq 0 ]; then
|
||||||
|
include='[]'
|
||||||
|
else
|
||||||
|
include=$(printf '%s\n' "${entries[@]}" | jq -s 'sort_by(.order, .dir)')
|
||||||
|
fi
|
||||||
|
matrix=$(jq -nc --argjson inc "$include" '{include:$inc}')
|
||||||
|
has=$([ "$(jq 'length' <<<"$include")" -gt 0 ] && echo true || echo false)
|
||||||
|
|
||||||
|
echo "matrix=$matrix" >> "$GITHUB_OUTPUT"
|
||||||
|
echo "has_stacks=$has" >> "$GITHUB_OUTPUT"
|
||||||
|
echo "Discovered stacks:"; echo "$include" | jq -r '.[] | " [\(.order)] \(.partition)@\(.region)/\(.stack) (\(.dir))"'
|
||||||
|
|
||||||
|
# ── Plan every changed stack (parallel; read-only) ───────────────────────────
|
||||||
|
plan:
|
||||||
|
name: Plan ${{ matrix.partition }}@${{ matrix.region }}/${{ matrix.stack }}
|
||||||
|
needs: [changes, discover]
|
||||||
|
if: needs.discover.outputs.has_stacks == 'true'
|
||||||
runs-on: ubuntu-latest
|
runs-on: ubuntu-latest
|
||||||
strategy:
|
strategy:
|
||||||
fail-fast: false
|
fail-fast: false
|
||||||
matrix:
|
matrix: ${{ fromJSON(needs.discover.outputs.matrix) }}
|
||||||
stack: [talos, dns, ceph, netbird]
|
|
||||||
env:
|
env:
|
||||||
OP_SERVICE_ACCOUNT_TOKEN: ${{ secrets.OP_TF_YUCCA_STAGING_ENV }}
|
# Read-scoped SA + env-file by partition.
|
||||||
|
OP_SERVICE_ACCOUNT_TOKEN: ${{ matrix.partition == 'prod' && secrets.OP_TF_YUCCA_PROD_ENV || secrets.OP_TF_YUCCA_STAGING_ENV }}
|
||||||
|
OP_ENV_FILE: ${{ matrix.partition == 'prod' && 'tf/.env.prod' || 'tf/.env' }}
|
||||||
|
SITE: ${{ matrix.region }}
|
||||||
|
REGION: ${{ matrix.region }}
|
||||||
|
PARTITION: ${{ matrix.partition }}
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout
|
- name: Checkout
|
||||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
||||||
with:
|
with:
|
||||||
persist-credentials: false
|
persist-credentials: false
|
||||||
|
- name: Set up mise (go + opentofu + terragrunt)
|
||||||
- name: Set up mise (opentofu + terragrunt)
|
|
||||||
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
|
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
|
||||||
|
|
||||||
- name: Install 1Password CLI
|
- name: Install 1Password CLI
|
||||||
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
|
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
|
||||||
|
|
||||||
# The talos plan reads `data.talos_cluster_health` (and the post-CNI gate),
|
# talos + fabric read data sources over the overlay on PLAN (talos_cluster_health;
|
||||||
# which dial the cluster over the 10.10.10.0/24 overlay — so the talos plan
|
# the switch vme), so they must join NetBird. The cloud-API stacks (ceph/dns/
|
||||||
# MUST join NetBird (data sources are read on plan and -refresh=false doesn't
|
# netbird) need no overlay. The setup key already exists in 1P (bootstrapped).
|
||||||
# skip them; the cloud-API stacks need no overlay). The CI setup key must
|
- name: Resolve NetBird CI setup-key ref
|
||||||
# already exist in 1P: it's minted by the netbird apply, so a fresh repo is
|
if: matrix.stack == 'talos' || matrix.stack == 'fabric'
|
||||||
# bootstrapped by applying staging/netbird ONCE out-of-band (see header) —
|
run: |
|
||||||
# after that every plan/apply just reads it, the same way the old Tailscale
|
set -euo pipefail
|
||||||
# OAuth was an out-of-band prerequisite.
|
if [ "$PARTITION" = "prod" ]; then
|
||||||
|
reg=$(printf '%s' "$REGION" | tr 'a-z-' 'A-Z_')
|
||||||
|
echo "NB_CI_KEY_REF=op://yucca_tf_prod/NETBIRD_YUCCA_PROD_${reg}_CI_SETUP_KEY/password" >> "$GITHUB_ENV"
|
||||||
|
else
|
||||||
|
# Staging's overlay key is partition-scoped (minted by the account/global netbird).
|
||||||
|
echo "NB_CI_KEY_REF=op://yucca_tf_staging/NETBIRD_YUCCA_STAGING_CI_SETUP_KEY/password" >> "$GITHUB_ENV"
|
||||||
|
fi
|
||||||
- name: Connect to NetBird
|
- name: Connect to NetBird
|
||||||
if: matrix.stack == 'talos'
|
if: matrix.stack == 'talos' || matrix.stack == 'fabric'
|
||||||
uses: ./.github/actions/netbird-connect
|
uses: ./.github/actions/netbird-connect
|
||||||
with:
|
with:
|
||||||
setup-key-ref: op://yucca_tf_staging/NETBIRD_YUCCA_STAGING_CI_SETUP_KEY/password
|
setup-key-ref: ${{ env.NB_CI_KEY_REF }}
|
||||||
hostname: gha-staging-plan-${{ github.run_id }}
|
hostname: gha-plan-${{ matrix.partition }}-${{ matrix.region }}-${{ matrix.stack }}-${{ github.run_id }}
|
||||||
|
|
||||||
|
# Prod fabric goes through the mise task (builds the junos-qfx/hetzner
|
||||||
|
# providers, renders the NETCONF key, -parallelism=1). Every other stack is a
|
||||||
|
# plain registry-provider terragrunt plan.
|
||||||
|
- name: Terragrunt plan (fabric)
|
||||||
|
if: matrix.stack == 'fabric'
|
||||||
|
run: mise run infra:plan -- --non-interactive
|
||||||
- name: Terragrunt plan
|
- name: Terragrunt plan
|
||||||
|
if: matrix.stack != 'fabric'
|
||||||
|
env:
|
||||||
|
STACK_DIR: ${{ matrix.dir }}
|
||||||
run: >-
|
run: >-
|
||||||
tf/op-run.sh terragrunt
|
tf/op-run.sh terragrunt
|
||||||
--working-dir tf/deployment/staging/${{ matrix.stack }}
|
--working-dir "tf/deployment/$STACK_DIR"
|
||||||
--non-interactive plan
|
--non-interactive plan
|
||||||
|
|
||||||
staging-apply:
|
# ── Apply (gated, ordered) ───────────────────────────────────────────────────
|
||||||
name: Staging apply (gated)
|
apply:
|
||||||
needs: [changes, staging-plan]
|
name: Apply ${{ matrix.partition }}@${{ matrix.region }}/${{ matrix.stack }} (gated)
|
||||||
|
needs: [changes, discover, plan]
|
||||||
# Only on merge to main (or manual dispatch) — never on PRs.
|
# Only on merge to main (or manual dispatch) — never on PRs.
|
||||||
if: (github.event_name == 'push' && needs.changes.outputs.staging == 'true') || github.event_name == 'workflow_dispatch'
|
if: (github.event_name == 'push' || github.event_name == 'workflow_dispatch') && needs.discover.outputs.has_stacks == 'true'
|
||||||
runs-on: ubuntu-latest
|
runs-on: ubuntu-latest
|
||||||
# The approval gate: `staging-infra` Environment with required reviewers.
|
strategy:
|
||||||
environment: staging-infra
|
fail-fast: false
|
||||||
|
# Serialize so the sorted `order` is honored: account netbird → site netbird
|
||||||
|
# → node-touching stacks, and prod/global before prod/htz-fsn1/netbird.
|
||||||
|
max-parallel: 1
|
||||||
|
matrix: ${{ fromJSON(needs.discover.outputs.matrix) }}
|
||||||
|
# Per-region approval gate.
|
||||||
|
environment:
|
||||||
|
name: ${{ matrix.partition }}-${{ matrix.region }}
|
||||||
env:
|
env:
|
||||||
# write-capable SA (the onepassword provider creates the JWT item)
|
# Write-capable SA by partition, EXCEPT prod/fabric which keeps the read SA
|
||||||
OP_SERVICE_ACCOUNT_TOKEN: ${{ secrets.OP_TF_YUCCA_STAGING_ENV_WRITE }}
|
# and self-escalates to the in-vault write SA inside `mise run infra:apply`.
|
||||||
|
OP_SERVICE_ACCOUNT_TOKEN: ${{ (matrix.partition == 'prod' && matrix.stack == 'fabric') && secrets.OP_TF_YUCCA_PROD_ENV || matrix.partition == 'prod' && secrets.OP_TF_YUCCA_PROD_ENV_WRITE || secrets.OP_TF_YUCCA_STAGING_ENV_WRITE }}
|
||||||
|
OP_ENV_FILE: ${{ matrix.partition == 'prod' && 'tf/.env.prod' || 'tf/.env' }}
|
||||||
|
SITE: ${{ matrix.region }}
|
||||||
|
REGION: ${{ matrix.region }}
|
||||||
|
PARTITION: ${{ matrix.partition }}
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout
|
- name: Checkout
|
||||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
||||||
with:
|
with:
|
||||||
persist-credentials: false
|
persist-credentials: false
|
||||||
|
- name: Set up mise (go + opentofu + terragrunt + ansible)
|
||||||
- name: Set up mise (opentofu + terragrunt)
|
|
||||||
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
|
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
|
||||||
|
|
||||||
- name: Install 1Password CLI
|
- name: Install 1Password CLI
|
||||||
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
|
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
|
||||||
|
|
||||||
# NetBird stack FIRST — pure api.netbird.io (no overlay needed), and it
|
# Node-touching stacks join the overlay: talos (provisions over the LAN),
|
||||||
# mints the CI setup key into 1P that the connect step below reads. On a
|
# ceph (the Ansible convergence below), fabric (the switch vme + the mgmt
|
||||||
# fresh bootstrap this is what makes the key exist before anything joins.
|
# converge below). NetBird stacks themselves are pure api.netbird.io. The
|
||||||
- name: Apply staging/netbird
|
# key was minted by an earlier (lower-order) netbird apply in this same run.
|
||||||
run: >-
|
- name: Resolve NetBird CI setup-key ref
|
||||||
tf/op-run.sh terragrunt
|
if: matrix.stack == 'talos' || matrix.stack == 'ceph' || matrix.stack == 'fabric'
|
||||||
--working-dir tf/deployment/staging/netbird
|
run: |
|
||||||
--non-interactive apply -auto-approve
|
set -euo pipefail
|
||||||
|
if [ "$PARTITION" = "prod" ]; then
|
||||||
# Join the NetBird overlay as a `ci` peer so the node-touching stacks and
|
reg=$(printf '%s' "$REGION" | tr 'a-z-' 'A-Z_')
|
||||||
# the Ansible deploy below reach 10.10.10.0/24 (the staging route advertises
|
echo "NB_CI_KEY_REF=op://yucca_tf_prod/NETBIRD_YUCCA_PROD_${reg}_CI_SETUP_KEY/password" >> "$GITHUB_ENV"
|
||||||
# the LAN).
|
else
|
||||||
|
echo "NB_CI_KEY_REF=op://yucca_tf_staging/NETBIRD_YUCCA_STAGING_CI_SETUP_KEY/password" >> "$GITHUB_ENV"
|
||||||
|
fi
|
||||||
- name: Connect to NetBird
|
- name: Connect to NetBird
|
||||||
|
if: matrix.stack == 'talos' || matrix.stack == 'ceph' || matrix.stack == 'fabric'
|
||||||
uses: ./.github/actions/netbird-connect
|
uses: ./.github/actions/netbird-connect
|
||||||
with:
|
with:
|
||||||
setup-key-ref: op://yucca_tf_staging/NETBIRD_YUCCA_STAGING_CI_SETUP_KEY/password
|
setup-key-ref: ${{ env.NB_CI_KEY_REF }}
|
||||||
hostname: gha-staging-apply-${{ github.run_id }}
|
hostname: gha-apply-${{ matrix.partition }}-${{ matrix.region }}-${{ matrix.stack }}-${{ github.run_id }}
|
||||||
|
|
||||||
# Ceph 1P password items (no node contact), then the Talos cluster
|
# Prod fabric: mise task escalates to the write SA + builds providers.
|
||||||
# (provisions secrets/CNI/Flux), then DNS.
|
- name: Terragrunt apply (fabric)
|
||||||
- name: Apply staging/ceph
|
if: matrix.stack == 'fabric'
|
||||||
|
run: mise run infra:apply -- --non-interactive -auto-approve
|
||||||
|
# Everything else: direct terragrunt apply with the partition's write SA.
|
||||||
|
- name: Terragrunt apply
|
||||||
|
if: matrix.stack != 'fabric'
|
||||||
|
env:
|
||||||
|
STACK_DIR: ${{ matrix.dir }}
|
||||||
run: >-
|
run: >-
|
||||||
tf/op-run.sh terragrunt
|
tf/op-run.sh terragrunt
|
||||||
--working-dir tf/deployment/staging/ceph
|
--working-dir "tf/deployment/$STACK_DIR"
|
||||||
--non-interactive apply -auto-approve
|
--non-interactive apply -auto-approve
|
||||||
|
|
||||||
- name: Apply staging/talos
|
# ── Ceph convergence (Ansible) — only on the ceph stack ──────────────────
|
||||||
run: >-
|
# The TF apply above only minted the RGW keys into 1P + the cluster Secret;
|
||||||
tf/op-run.sh terragrunt
|
# this creates the matching RGW users on the bare-metal cluster. Reuses the
|
||||||
--working-dir tf/deployment/staging/talos
|
# NetBird overlay + 1Password session already established in this entry.
|
||||||
--non-interactive apply -auto-approve
|
|
||||||
|
|
||||||
- name: Apply staging/dns
|
|
||||||
run: >-
|
|
||||||
tf/op-run.sh terragrunt
|
|
||||||
--working-dir tf/deployment/staging/dns
|
|
||||||
--non-interactive apply -auto-approve
|
|
||||||
|
|
||||||
# ── Ceph convergence (Ansible) ──────────────────────────────────────
|
|
||||||
# The TF stacks above only minted the RGW keys into 1P + the cluster
|
|
||||||
# Secret; this is what creates the matching RGW users on the bare-metal
|
|
||||||
# cluster. Runs after the TF apply so the inventory (rendered from the
|
|
||||||
# ceph stack's `render` output) and the keys exist. Reuses the NetBird
|
|
||||||
# overlay + 1Password session already established in this job.
|
|
||||||
- name: Render Ansible inventory from the ceph TF state
|
- name: Render Ansible inventory from the ceph TF state
|
||||||
run: ansible/ceph/scripts/render-inventories.sh staging
|
if: matrix.stack == 'ceph'
|
||||||
|
env:
|
||||||
|
PARTITION: ${{ matrix.partition }}
|
||||||
|
run: ansible/ceph/scripts/render-inventories.sh "$PARTITION"
|
||||||
- name: Install the ansible-iac SSH key from 1Password
|
- name: Install the ansible-iac SSH key from 1Password
|
||||||
# Pulls op://yucca_tf_staging/SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY to
|
if: matrix.stack == 'ceph'
|
||||||
|
# Pulls op://yucca_tf_<partition>/SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY to
|
||||||
# ~/.ssh/id_ed25519_sietch — the path the rendered inventory references.
|
# ~/.ssh/id_ed25519_sietch — the path the rendered inventory references.
|
||||||
|
env:
|
||||||
|
PARTITION: ${{ matrix.partition }}
|
||||||
run: |
|
run: |
|
||||||
mkdir -p ~/.ssh && chmod 700 ~/.ssh
|
mkdir -p ~/.ssh && chmod 700 ~/.ssh
|
||||||
OP_VAULT=yucca_tf_staging ansible/ceph/scripts/install-ssh-keys.sh sietch
|
OP_VAULT="yucca_tf_$PARTITION" ansible/ceph/scripts/install-ssh-keys.sh sietch
|
||||||
|
|
||||||
- name: Provision the ceph Ansible toolchain (venv + collections)
|
- name: Provision the ceph Ansible toolchain (venv + collections)
|
||||||
|
if: matrix.stack == 'ceph'
|
||||||
working-directory: ansible/ceph
|
working-directory: ansible/ceph
|
||||||
run: |
|
run: |
|
||||||
mise trust
|
mise trust
|
||||||
mise install
|
mise install
|
||||||
mise run setup
|
mise run setup
|
||||||
|
|
||||||
- name: Deploy Ceph (full pipeline — baseline → tune → deploy → harden)
|
- name: Deploy Ceph (full pipeline — baseline → tune → deploy → harden)
|
||||||
working-directory: ansible/ceph
|
if: matrix.stack == 'ceph'
|
||||||
# CI runner has no known_hosts for the bare-metal nodes; first contact is
|
# CI runner has no known_hosts for the bare-metal nodes; first contact is
|
||||||
# over the NetBird overlay, so disable strict host-key checking for this run.
|
# over the NetBird overlay, so disable strict host-key checking for this run.
|
||||||
env:
|
env:
|
||||||
ANSIBLE_HOST_KEY_CHECKING: "false"
|
ANSIBLE_HOST_KEY_CHECKING: "false"
|
||||||
CEPH_ENV: inventories/sietch-ceph.staging.austin.int/inventory.ini
|
CEPH_ENV: inventories/${{ matrix.partition }}-${{ matrix.region }}/sietch/inventory.ini
|
||||||
run: mise run deploy
|
run: mise run deploy
|
||||||
|
|
||||||
# ── Prod NetBird (global + site layers) — mints the CI/mgmt setup keys ────────
|
# ── Mgmt convergence (Ansible) — only on the fabric stack ────────────────
|
||||||
# Pure api.netbird.io (no nodes/overlay). Layered: prod/global (account-wide
|
# Renders the mgmt inventory from TF (tf/render/ansible-mgmt) then converges
|
||||||
# groups + the yucca→yucca_resource policy) then prod/<site>/netbird (site
|
# the mgmt hosts over the overlay. No-op for regions without an ansible/mgmt
|
||||||
# groups/keys/policies + the routed network). Runs before the fabric/mgmt jobs
|
# inventory. Runs under the same (already-approved) gate as the fabric apply.
|
||||||
# conceptually (it mints the keys they consume), but they're only loosely coupled
|
|
||||||
# — the keys persist in 1P across runs, so a fresh bootstrap applies this first.
|
|
||||||
netbird-prod-plan:
|
|
||||||
name: Plan prod/netbird
|
|
||||||
needs: changes
|
|
||||||
if: needs.changes.outputs.prod_netbird == 'true' || github.event_name == 'workflow_dispatch'
|
|
||||||
runs-on: ubuntu-latest
|
|
||||||
env:
|
|
||||||
OP_ENV_FILE: tf/.env.prod
|
|
||||||
OP_SERVICE_ACCOUNT_TOKEN: ${{ secrets.OP_TF_YUCCA_PROD_ENV }}
|
|
||||||
steps:
|
|
||||||
- name: Checkout
|
|
||||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
|
||||||
with:
|
|
||||||
persist-credentials: false
|
|
||||||
|
|
||||||
- name: Set up mise (opentofu + terragrunt)
|
|
||||||
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
|
|
||||||
|
|
||||||
- name: Install 1Password CLI
|
|
||||||
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
|
|
||||||
|
|
||||||
# Global layer first, then the site layer (which reads the global layer's
|
|
||||||
# group_ids via a terragrunt dependency — mock_outputs cover the PR plan
|
|
||||||
# before global is ever applied).
|
|
||||||
- name: Terragrunt plan prod/global
|
|
||||||
run: >-
|
|
||||||
tf/op-run.sh terragrunt
|
|
||||||
--working-dir tf/deployment/prod/global
|
|
||||||
--non-interactive plan
|
|
||||||
|
|
||||||
- name: Terragrunt plan prod/htz-fsn1/netbird
|
|
||||||
run: >-
|
|
||||||
tf/op-run.sh terragrunt
|
|
||||||
--working-dir tf/deployment/prod/htz-fsn1/netbird
|
|
||||||
--non-interactive plan
|
|
||||||
|
|
||||||
netbird-prod-apply:
|
|
||||||
name: Apply prod/netbird (gated)
|
|
||||||
needs: [changes, netbird-prod-plan]
|
|
||||||
if: (github.event_name == 'push' && needs.changes.outputs.prod_netbird == 'true') || github.event_name == 'workflow_dispatch'
|
|
||||||
runs-on: ubuntu-latest
|
|
||||||
environment: prod-infra
|
|
||||||
env:
|
|
||||||
OP_ENV_FILE: tf/.env.prod
|
|
||||||
OP_SERVICE_ACCOUNT_TOKEN: ${{ secrets.OP_TF_YUCCA_PROD_ENV_WRITE }}
|
|
||||||
steps:
|
|
||||||
- name: Checkout
|
|
||||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
|
||||||
with:
|
|
||||||
persist-credentials: false
|
|
||||||
|
|
||||||
- name: Set up mise (opentofu + terragrunt)
|
|
||||||
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
|
|
||||||
|
|
||||||
- name: Install 1Password CLI
|
|
||||||
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
|
|
||||||
|
|
||||||
# Global layer must apply before the site layer (the site reads global's
|
|
||||||
# group_ids output via its terragrunt dependency).
|
|
||||||
- name: Apply prod/global
|
|
||||||
run: >-
|
|
||||||
tf/op-run.sh terragrunt
|
|
||||||
--working-dir tf/deployment/prod/global
|
|
||||||
--non-interactive apply -auto-approve
|
|
||||||
|
|
||||||
- name: Apply prod/htz-fsn1/netbird
|
|
||||||
run: >-
|
|
||||||
tf/op-run.sh terragrunt
|
|
||||||
--working-dir tf/deployment/prod/htz-fsn1/netbird
|
|
||||||
--non-interactive apply -auto-approve
|
|
||||||
|
|
||||||
# ── Prod fabric + mgmt stacks (one per site) ─────────────────────────────────
|
|
||||||
prod-discover:
|
|
||||||
name: Discover prod fabric sites
|
|
||||||
needs: changes
|
|
||||||
# Needed by both the TF and ansible prod jobs, so run if either area changed.
|
|
||||||
if: needs.changes.outputs.prod_tf == 'true' || needs.changes.outputs.prod_ansible == 'true' || github.event_name == 'workflow_dispatch'
|
|
||||||
runs-on: ubuntu-latest
|
|
||||||
outputs:
|
|
||||||
matrix: ${{ steps.sites.outputs.matrix }}
|
|
||||||
steps:
|
|
||||||
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
|
||||||
with:
|
|
||||||
persist-credentials: false
|
|
||||||
- id: sites
|
|
||||||
name: List fabric sites (dirs with a fabric.tf; excludes netbird-only dirs)
|
|
||||||
run: |
|
|
||||||
matrix=$(for d in tf/deployment/prod/*/; do [ -f "${d}fabric.tf" ] && basename "$d"; done | jq -R . | jq -cs '{site: .}')
|
|
||||||
echo "matrix=$matrix" >> "$GITHUB_OUTPUT"
|
|
||||||
echo "$matrix"
|
|
||||||
|
|
||||||
prod-plan:
|
|
||||||
name: Prod plan ${{ matrix.site }}
|
|
||||||
needs: [changes, prod-discover]
|
|
||||||
if: needs.changes.outputs.prod_tf == 'true' || github.event_name == 'workflow_dispatch'
|
|
||||||
runs-on: ubuntu-latest
|
|
||||||
strategy:
|
|
||||||
fail-fast: false
|
|
||||||
matrix: ${{ fromJSON(needs.prod-discover.outputs.matrix) }}
|
|
||||||
env:
|
|
||||||
OP_SERVICE_ACCOUNT_TOKEN: ${{ secrets.OP_TF_YUCCA_PROD_ENV }}
|
|
||||||
SITE: ${{ matrix.site }}
|
|
||||||
steps:
|
|
||||||
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
|
||||||
with:
|
|
||||||
persist-credentials: false
|
|
||||||
- name: Set up mise (go + opentofu + terragrunt)
|
|
||||||
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
|
|
||||||
- name: Install 1Password CLI
|
|
||||||
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
|
|
||||||
# Reach the switch vme (routed via the mgmt NetBird peers) over the overlay.
|
|
||||||
- name: Resolve NetBird CI setup-key ref
|
|
||||||
run: echo "NB_CI_KEY_REF=op://yucca_tf_prod/NETBIRD_YUCCA_PROD_$(echo "$SITE" | tr 'a-z-' 'A-Z_')_CI_SETUP_KEY/password" >> "$GITHUB_ENV"
|
|
||||||
- name: Connect to NetBird
|
|
||||||
uses: ./.github/actions/netbird-connect
|
|
||||||
with:
|
|
||||||
setup-key-ref: ${{ env.NB_CI_KEY_REF }}
|
|
||||||
hostname: gha-prod-plan-${{ matrix.site }}-${{ github.run_id }}
|
|
||||||
- name: Deploy plan
|
|
||||||
run: mise run infra:plan -- --non-interactive
|
|
||||||
|
|
||||||
prod-apply:
|
|
||||||
name: Prod apply ${{ matrix.site }} (gated)
|
|
||||||
needs: [changes, prod-discover, prod-plan]
|
|
||||||
if: (github.event_name == 'push' && needs.changes.outputs.prod_tf == 'true') || github.event_name == 'workflow_dispatch'
|
|
||||||
runs-on: ubuntu-latest
|
|
||||||
strategy:
|
|
||||||
fail-fast: false
|
|
||||||
matrix: ${{ fromJSON(needs.prod-discover.outputs.matrix) }}
|
|
||||||
# Site-scoped gate: each site's prod stack has its own required reviewers.
|
|
||||||
environment:
|
|
||||||
name: prod-${{ matrix.site }}
|
|
||||||
env:
|
|
||||||
# Read-scoped token; infra:apply escalates to the write SA stored in the vault.
|
|
||||||
OP_SERVICE_ACCOUNT_TOKEN: ${{ secrets.OP_TF_YUCCA_PROD_ENV }}
|
|
||||||
SITE: ${{ matrix.site }}
|
|
||||||
steps:
|
|
||||||
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
|
||||||
with:
|
|
||||||
persist-credentials: false
|
|
||||||
- name: Set up mise (go + opentofu + terragrunt)
|
|
||||||
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
|
|
||||||
- name: Install 1Password CLI
|
|
||||||
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
|
|
||||||
- name: Resolve NetBird CI setup-key ref
|
|
||||||
run: echo "NB_CI_KEY_REF=op://yucca_tf_prod/NETBIRD_YUCCA_PROD_$(echo "$SITE" | tr 'a-z-' 'A-Z_')_CI_SETUP_KEY/password" >> "$GITHUB_ENV"
|
|
||||||
- name: Connect to NetBird
|
|
||||||
uses: ./.github/actions/netbird-connect
|
|
||||||
with:
|
|
||||||
setup-key-ref: ${{ env.NB_CI_KEY_REF }}
|
|
||||||
hostname: gha-prod-apply-${{ matrix.site }}-${{ github.run_id }}
|
|
||||||
- name: Deploy apply
|
|
||||||
run: mise run infra:apply -- --non-interactive -auto-approve
|
|
||||||
|
|
||||||
prod-ansible:
|
|
||||||
name: Prod ansible ${{ matrix.site }}
|
|
||||||
needs: [changes, prod-discover, prod-apply]
|
|
||||||
# Runs after the TF apply, but also on ansible-only changes (apply skipped).
|
|
||||||
# always() so a skipped prod-apply (TF unchanged) doesn't skip this; still
|
|
||||||
# bails if discover failed or the apply actually failed.
|
|
||||||
if: >-
|
|
||||||
always()
|
|
||||||
&& needs.prod-discover.result == 'success'
|
|
||||||
&& needs.prod-apply.result != 'failure'
|
|
||||||
&& needs.prod-apply.result != 'cancelled'
|
|
||||||
&& ((github.event_name == 'push' && needs.changes.outputs.prod_ansible == 'true') || github.event_name == 'workflow_dispatch')
|
|
||||||
runs-on: ubuntu-latest
|
|
||||||
strategy:
|
|
||||||
fail-fast: false
|
|
||||||
matrix: ${{ fromJSON(needs.prod-discover.outputs.matrix) }}
|
|
||||||
# No environment gate: the prod-apply gate already approved this deploy, and
|
|
||||||
# ansible/mgmt only reads from 1Password (read-scoped repo secret).
|
|
||||||
env:
|
|
||||||
OP_SERVICE_ACCOUNT_TOKEN: ${{ secrets.OP_TF_YUCCA_PROD_ENV }}
|
|
||||||
SITE: ${{ matrix.site }}
|
|
||||||
steps:
|
|
||||||
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
|
||||||
with:
|
|
||||||
persist-credentials: false
|
|
||||||
- name: Set up mise (ansible + opentofu)
|
|
||||||
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
|
|
||||||
- name: Install 1Password CLI
|
|
||||||
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
|
|
||||||
# On the overlay so the playbook can reach mgmt hosts that have joined NetBird
|
|
||||||
# (it prefers their NetBird IP, falling back to the public IP otherwise).
|
|
||||||
- name: Resolve NetBird CI setup-key ref
|
|
||||||
run: echo "NB_CI_KEY_REF=op://yucca_tf_prod/NETBIRD_YUCCA_PROD_$(echo "$SITE" | tr 'a-z-' 'A-Z_')_CI_SETUP_KEY/password" >> "$GITHUB_ENV"
|
|
||||||
- name: Connect to NetBird
|
|
||||||
uses: ./.github/actions/netbird-connect
|
|
||||||
with:
|
|
||||||
setup-key-ref: ${{ env.NB_CI_KEY_REF }}
|
|
||||||
hostname: gha-prod-ansible-${{ matrix.site }}-${{ github.run_id }}
|
|
||||||
# Renders the inventory from TF (tf/render/ansible-mgmt) then converges the
|
|
||||||
# mgmt hosts — over NetBird if they've joined, else their public IP (the
|
|
||||||
# TF-generated provisioning key authorizes root). No-op for sites without an
|
|
||||||
# ansible/mgmt inventory; requires the hosts to have been reprovisioned.
|
|
||||||
- name: Ansible converge (mgmt hosts)
|
- name: Ansible converge (mgmt hosts)
|
||||||
|
if: matrix.stack == 'fabric'
|
||||||
run: mise run mgmt:ansible
|
run: mise run mgmt:ansible
|
||||||
|
|||||||
@@ -1,6 +1,9 @@
|
|||||||
node_modules
|
node_modules
|
||||||
*.tsbuildinfo
|
*.tsbuildinfo
|
||||||
|
|
||||||
|
# Claude Code local state — per-operator settings + agent worktrees; never commit.
|
||||||
|
.claude/
|
||||||
|
|
||||||
.env.local
|
.env.local
|
||||||
.env
|
.env
|
||||||
|
|
||||||
@@ -50,6 +53,9 @@ ansible/*/.ansible/
|
|||||||
ansible/*/.ansible_facts_cache/
|
ansible/*/.ansible_facts_cache/
|
||||||
ansible/*/.venv/
|
ansible/*/.venv/
|
||||||
|
|
||||||
|
# Working notes for the partition/region/ceph-cluster rework — local scratch.
|
||||||
|
notes.md
|
||||||
|
|
||||||
# Lens / decision-support artifacts are controller-local notes, not repo code.
|
# Lens / decision-support artifacts are controller-local notes, not repo code.
|
||||||
# Per-project: kept outside the repo (e.g., ~/Projects/immich/yucca-ceph-import/analysis/
|
# Per-project: kept outside the repo (e.g., ~/Projects/immich/yucca-ceph-import/analysis/
|
||||||
# on operator workstation, but not tracked).
|
# on operator workstation, but not tracked).
|
||||||
|
|||||||
+5
-5
@@ -73,20 +73,20 @@ run = [{ task = "*:fix" }, { task = "web:lingui" }]
|
|||||||
# No literal secrets in the .env file — just op:// references.
|
# No literal secrets in the .env file — just op:// references.
|
||||||
|
|
||||||
[tasks."tf:init"]
|
[tasks."tf:init"]
|
||||||
description = "Terragrunt init for a given stack (default: deployment/staging/ceph)"
|
description = "Terragrunt init for a given stack (default: deployment/staging/austin/ceph)"
|
||||||
run = "tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/staging/ceph} init"
|
run = "tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/staging/austin/ceph} init"
|
||||||
|
|
||||||
[tasks."tf:plan"]
|
[tasks."tf:plan"]
|
||||||
description = "Terragrunt plan for a given stack"
|
description = "Terragrunt plan for a given stack"
|
||||||
run = "tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/staging/ceph} plan"
|
run = "tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/staging/austin/ceph} plan"
|
||||||
|
|
||||||
[tasks."tf:apply"]
|
[tasks."tf:apply"]
|
||||||
description = "Terragrunt apply for a given stack"
|
description = "Terragrunt apply for a given stack"
|
||||||
run = "tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/staging/ceph} apply"
|
run = "tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/staging/austin/ceph} apply"
|
||||||
|
|
||||||
[tasks."tf:destroy"]
|
[tasks."tf:destroy"]
|
||||||
description = "Terragrunt destroy for a given stack (use with care)"
|
description = "Terragrunt destroy for a given stack (use with care)"
|
||||||
run = "tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/staging/ceph} destroy"
|
run = "tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/staging/austin/ceph} destroy"
|
||||||
|
|
||||||
[tasks."tf:fmt"]
|
[tasks."tf:fmt"]
|
||||||
description = "Format terraform + terragrunt files recursively"
|
description = "Format terraform + terragrunt files recursively"
|
||||||
|
|||||||
@@ -35,4 +35,4 @@ SITE="${SITE:-htz-fsn1}"
|
|||||||
unset OP_ACCOUNT
|
unset OP_ACCOUNT
|
||||||
# -parallelism=1: the JTAF junos-qfx provider isn't concurrency-safe — parallel
|
# -parallelism=1: the JTAF junos-qfx provider isn't concurrency-safe — parallel
|
||||||
# ApplyResourceChange calls across the VCs crash the plugin ("Plugin did not respond").
|
# ApplyResourceChange calls across the VCs crash the plugin ("Plugin did not respond").
|
||||||
OP_ENV_FILE=tf/.env.prod "$ROOT/tf/op-run.sh" terragrunt --working-dir "tf/deployment/prod/$SITE" apply -parallelism=1 "$@"
|
OP_ENV_FILE=tf/.env.prod "$ROOT/tf/op-run.sh" terragrunt --working-dir "tf/deployment/prod/$SITE/fabric" apply -parallelism=1 "$@"
|
||||||
|
|||||||
@@ -34,4 +34,4 @@ SITE="${SITE:-htz-fsn1}"
|
|||||||
# provider rejects having both set ("service_account_token and account are set").
|
# provider rejects having both set ("service_account_token and account are set").
|
||||||
unset OP_ACCOUNT
|
unset OP_ACCOUNT
|
||||||
# -parallelism=1: the JTAF junos-qfx provider isn't concurrency-safe (see apply).
|
# -parallelism=1: the JTAF junos-qfx provider isn't concurrency-safe (see apply).
|
||||||
OP_ENV_FILE=tf/.env.prod "$ROOT/tf/op-run.sh" terragrunt --working-dir "tf/deployment/prod/$SITE" plan -parallelism=1 "$@"
|
OP_ENV_FILE=tf/.env.prod "$ROOT/tf/op-run.sh" terragrunt --working-dir "tf/deployment/prod/$SITE/fabric" plan -parallelism=1 "$@"
|
||||||
|
|||||||
+26
-17
@@ -9,8 +9,11 @@
|
|||||||
# rebuilds charts/*/charts/ (rm -rf + dependency build) and races the renders.
|
# rebuilds charts/*/charts/ (rm -rf + dependency build) and races the renders.
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
|
|
||||||
LIB_CONSUMERS=(yucca-api yucca-admin-api yucca-metrics-worker web michael mock-oidc)
|
# Charts are role-grouped: apps/* (services), platform/* (operators/CRs),
|
||||||
ALL_CHARTS=(yucca-api yucca-admin-api yucca-metrics-worker web michael mock-oidc cnpg-cluster ceph-objectuser rook-ceph-cluster)
|
# lib/yucca-common (shared library), dev/mock-oidc (dev-only). Paths below are
|
||||||
|
# relative to charts/.
|
||||||
|
LIB_CONSUMERS=(apps/yucca-api apps/yucca-admin-api apps/yucca-metrics-worker apps/web apps/michael dev/mock-oidc)
|
||||||
|
ALL_CHARTS=(apps/yucca-api apps/yucca-admin-api apps/yucca-metrics-worker apps/web apps/michael dev/mock-oidc platform/cnpg-cluster platform/ceph-objectuser platform/rook-ceph-cluster)
|
||||||
|
|
||||||
echo "==> helm dependency build (yucca-common consumers)"
|
echo "==> helm dependency build (yucca-common consumers)"
|
||||||
for c in "${LIB_CONSUMERS[@]}"; do
|
for c in "${LIB_CONSUMERS[@]}"; do
|
||||||
@@ -20,7 +23,7 @@ done
|
|||||||
echo "==> helm template + kubeconform"
|
echo "==> helm template + kubeconform"
|
||||||
for c in "${ALL_CHARTS[@]}"; do
|
for c in "${ALL_CHARTS[@]}"; do
|
||||||
ns=yucca
|
ns=yucca
|
||||||
[ "$c" = "rook-ceph-cluster" ] && ns=rook-ceph
|
[ "$c" = "platform/rook-ceph-cluster" ] && ns=rook-ceph
|
||||||
# NB: release name must not be YAML-boolean-ish ("y"/"on"/...): it lands in
|
# NB: release name must not be YAML-boolean-ish ("y"/"on"/...): it lands in
|
||||||
# labels and kubeconform reads it back as a bool.
|
# labels and kubeconform reads it back as a bool.
|
||||||
helm template yucca "charts/$c" -n "$ns" \
|
helm template yucca "charts/$c" -n "$ns" \
|
||||||
@@ -29,27 +32,33 @@ for c in "${ALL_CHARTS[@]}"; do
|
|||||||
done
|
done
|
||||||
|
|
||||||
echo "==> kustomize build (Flux entrypoints)"
|
echo "==> kustomize build (Flux entrypoints)"
|
||||||
# o11y-style GitOps tree: per-env cluster entry points + app overlays.
|
# partition/region GitOps tree: per-cluster entry points (clusters/<p>/<r>) +
|
||||||
for env in staging production; do
|
# app overlays (apps/<p>/<r>, which compose components/{infra,roles/<role>}).
|
||||||
kustomize build "kubernetes/clusters/$env" >/dev/null && echo " OK kubernetes/clusters/$env"
|
# LoadRestrictionsNone mirrors how Flux's kustomize-controller builds within a
|
||||||
kustomize build "kubernetes/apps/$env" >/dev/null && echo " OK kubernetes/apps/$env"
|
# single git artifact: the role Components reference ../../apps/<app>.yaml
|
||||||
|
# (a sibling subtree under components/), which the CLI's default RootOnly
|
||||||
|
# restrictor would reject even though Flux allows it.
|
||||||
|
kb() { kustomize build --load-restrictor=LoadRestrictionsNone "$@"; }
|
||||||
|
CLUSTERS=(staging/austin prod/htz-fsn1 dev/local)
|
||||||
|
for c in "${CLUSTERS[@]}"; do
|
||||||
|
kb "kubernetes/clusters/$c" >/dev/null && echo " OK kubernetes/clusters/$c"
|
||||||
|
kb "kubernetes/apps/$c" >/dev/null && echo " OK kubernetes/apps/$c"
|
||||||
done
|
done
|
||||||
# Dev-mirror tree (consumed by Tilt; kept validated too).
|
# Dev-mirror HelmRepository sources (consumed by Tilt + the dev cluster-repos
|
||||||
kustomize build kubernetes/apps >/dev/null && echo " OK kubernetes/apps (dev tree)"
|
# Kustomization; kept validated too).
|
||||||
kustomize build kubernetes/flux/repos >/dev/null && echo " OK kubernetes/flux/repos"
|
kb kubernetes/apps/dev/local/repos >/dev/null && echo " OK kubernetes/apps/dev/local/repos"
|
||||||
kustomize build kubernetes/flux/cluster >/dev/null && echo " OK kubernetes/flux/cluster"
|
|
||||||
|
|
||||||
echo "==> flux-local build (full tree, helm rendering as Flux would)"
|
echo "==> flux-local build (full tree, helm rendering as Flux would)"
|
||||||
# Via uvx so uv supplies a Python matching flux-local's requires-python
|
# Via uvx so uv supplies a Python matching flux-local's requires-python
|
||||||
# (>=3.13) regardless of the host. One retry: flux-local fans out `flux build
|
# (>=3.13) regardless of the host. One retry: flux-local fans out `flux build
|
||||||
# ks` subprocesses which very rarely segfault under load.
|
# ks` subprocesses which very rarely segfault under load.
|
||||||
flux_local() { uvx --from "flux-local==8.2.0" flux-local "$@"; }
|
flux_local() { uvx --from "flux-local==8.2.0" flux-local "$@"; }
|
||||||
# Per-cluster path: staging + production reuse the same Kustomization names
|
# Per-cluster path: clusters reuse the same Kustomization names (cluster-apps,
|
||||||
# (flux-system ns), so flux-local must scope to one cluster at a time.
|
# flux-system ns), so flux-local must scope to one cluster at a time.
|
||||||
for env in staging production; do
|
for c in "${CLUSTERS[@]}"; do
|
||||||
flux_local build all "kubernetes/clusters/$env" --enable-helm --no-enable-dns >/dev/null \
|
flux_local build all "kubernetes/clusters/$c" --enable-helm --no-enable-dns >/dev/null \
|
||||||
|| flux_local build all "kubernetes/clusters/$env" --enable-helm --no-enable-dns >/dev/null
|
|| flux_local build all "kubernetes/clusters/$c" --enable-helm --no-enable-dns >/dev/null
|
||||||
echo " OK flux-local ($env)"
|
echo " OK flux-local ($c)"
|
||||||
done
|
done
|
||||||
|
|
||||||
echo "k8s surface: ALL VALID"
|
echo "k8s surface: ALL VALID"
|
||||||
|
|||||||
+10
-10
@@ -1,13 +1,13 @@
|
|||||||
#!/usr/bin/env bash
|
#!/usr/bin/env bash
|
||||||
#MISE description="Converge a site's management hosts with ansible/mgmt (root via the TF provisioning key from 1Password). SITE selects the inventory."
|
#MISE description="Converge a region's management hosts with ansible/mgmt (root via the TF provisioning key from 1Password). REGION selects the inventory."
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
ROOT=$(git rev-parse --show-toplevel)
|
ROOT=$(git rev-parse --show-toplevel)
|
||||||
SITE="${SITE:-htz-fsn1}"
|
REGION="${REGION:-htz-fsn1}"
|
||||||
INV="$ROOT/ansible/mgmt/inventories/$SITE"
|
INV="$ROOT/ansible/mgmt/inventories/$REGION"
|
||||||
|
|
||||||
# Per-site: only run where an inventory exists (keeps the prod matrix happy).
|
# Per-region: only run where an inventory exists (keeps the prod matrix happy).
|
||||||
if [ ! -d "$INV" ]; then
|
if [ ! -d "$INV" ]; then
|
||||||
echo "mgmt:ansible: no ansible/mgmt inventory for site '$SITE' — skipping."
|
echo "mgmt:ansible: no ansible/mgmt inventory for region '$REGION' — skipping."
|
||||||
exit 0
|
exit 0
|
||||||
fi
|
fi
|
||||||
|
|
||||||
@@ -19,18 +19,18 @@ mise run mgmt:render-inventory
|
|||||||
ACCT=(); [ -z "${OP_SERVICE_ACCOUNT_TOKEN:-}" ] && ACCT=(--account "${OP_ACCOUNT:-team-futo}")
|
ACCT=(); [ -z "${OP_SERVICE_ACCOUNT_TOKEN:-}" ] && ACCT=(--account "${OP_ACCOUNT:-team-futo}")
|
||||||
|
|
||||||
# Provisioning private key (root login) -> 0600 temp file. Item name is
|
# Provisioning private key (root login) -> 0600 temp file. Item name is
|
||||||
# site-derived: htz-fsn1 -> HTZ_FSN1_PROVISIONING_SSH_PRIVATE_KEY (set by mgmt.tf).
|
# region-derived: htz-fsn1 -> HTZ_FSN1_PROVISIONING_SSH_PRIVATE_KEY (set by mgmt.tf).
|
||||||
KEY_ITEM="$(printf '%s' "$SITE" | tr 'a-z-' 'A-Z_')_PROVISIONING_SSH_PRIVATE_KEY"
|
KEY_ITEM="$(printf '%s' "$REGION" | tr 'a-z-' 'A-Z_')_PROVISIONING_SSH_PRIVATE_KEY"
|
||||||
KEYF=$(mktemp); chmod 600 "$KEYF"; trap 'rm -f "$KEYF"' EXIT
|
KEYF=$(mktemp); chmod 600 "$KEYF"; trap 'rm -f "$KEYF"' EXIT
|
||||||
op read "${ACCT[@]}" "op://yucca_tf_prod/$KEY_ITEM/password" > "$KEYF"
|
op read "${ACCT[@]}" "op://yucca_tf_prod/$KEY_ITEM/password" > "$KEYF"
|
||||||
|
|
||||||
# NetBird "mgmt" setup key (auto_groups=["mgmt"]) — joins the node to the overlay
|
# NetBird "mgmt" setup key (auto_groups=["mgmt"]) — joins the node to the overlay
|
||||||
# as a route peer. Site-derived item: htz-fsn1 -> NETBIRD_YUCCA_PROD_HTZ_FSN1_MGMT_SETUP_KEY.
|
# as a route peer. Region-derived item: htz-fsn1 -> NETBIRD_YUCCA_PROD_HTZ_FSN1_MGMT_SETUP_KEY.
|
||||||
NB_KEY_ITEM="NETBIRD_YUCCA_PROD_$(printf '%s' "$SITE" | tr 'a-z-' 'A-Z_')_MGMT_SETUP_KEY"
|
NB_KEY_ITEM="NETBIRD_YUCCA_PROD_$(printf '%s' "$REGION" | tr 'a-z-' 'A-Z_')_MGMT_SETUP_KEY"
|
||||||
NB_SETUP_KEY=$(op read "${ACCT[@]}" "op://yucca_tf_prod/$NB_KEY_ITEM/password")
|
NB_SETUP_KEY=$(op read "${ACCT[@]}" "op://yucca_tf_prod/$NB_KEY_ITEM/password")
|
||||||
|
|
||||||
cd "$ROOT/ansible/mgmt"
|
cd "$ROOT/ansible/mgmt"
|
||||||
ansible-galaxy collection install -r requirements.yml >/dev/null
|
ansible-galaxy collection install -r requirements.yml >/dev/null
|
||||||
ansible-playbook -i "inventories/$SITE" site.yml \
|
ansible-playbook -i "inventories/$REGION" site.yml \
|
||||||
--private-key "$KEYF" \
|
--private-key "$KEYF" \
|
||||||
--extra-vars "mgmt_netbird_setup_key=$NB_SETUP_KEY" "$@"
|
--extra-vars "mgmt_netbird_setup_key=$NB_SETUP_KEY" "$@"
|
||||||
|
|||||||
@@ -1,10 +1,10 @@
|
|||||||
#!/usr/bin/env bash
|
#!/usr/bin/env bash
|
||||||
#MISE description="Render a site's ansible/mgmt inventory from Terraform (tf/render/ansible-mgmt: addressing + identity + mgmt-hosts.yaml). SITE selects the site."
|
#MISE description="Render a region's ansible/mgmt inventory from Terraform (tf/render/ansible-mgmt: addressing + identity + mgmt-hosts.yaml). REGION selects the region."
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
ROOT=$(git rev-parse --show-toplevel)
|
ROOT=$(git rev-parse --show-toplevel)
|
||||||
SITE="${SITE:-htz-fsn1}"
|
REGION="${REGION:-htz-fsn1}"
|
||||||
|
|
||||||
cd "$ROOT/tf/render/ansible-mgmt"
|
cd "$ROOT/tf/render/ansible-mgmt"
|
||||||
tofu init -input=false >/dev/null
|
tofu init -input=false >/dev/null
|
||||||
tofu apply -input=false -auto-approve -var "site=$SITE" >/dev/null
|
tofu apply -input=false -auto-approve -var "region=$REGION" >/dev/null
|
||||||
echo "rendered ansible/mgmt/inventories/$SITE/ (hosts.yml, host_vars/, group_vars/all/users.generated.yml)"
|
echo "rendered ansible/mgmt/inventories/$REGION/ (hosts.yml, host_vars/, group_vars/all/users.generated.yml)"
|
||||||
|
|||||||
Executable
+5
@@ -0,0 +1,5 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
#MISE description="Build yuctl"
|
||||||
|
set -e
|
||||||
|
|
||||||
|
cd packages/yuctl && go build -o ../../dist/yuctl .
|
||||||
Executable
+5
@@ -0,0 +1,5 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
#MISE description="Run yuctl from source (forwards args)"
|
||||||
|
set -e
|
||||||
|
|
||||||
|
cd packages/yuctl && go run . "$@"
|
||||||
@@ -7,7 +7,7 @@
|
|||||||
# locally-built images.
|
# locally-built images.
|
||||||
# - HelmRepository-sourced HelmReleases (cnpg, rook, victoria-*) -> installed
|
# - HelmRepository-sourced HelmReleases (cnpg, rook, victoria-*) -> installed
|
||||||
# at the exact chart version + values pinned in the HelmRelease, from the
|
# at the exact chart version + values pinned in the HelmRelease, from the
|
||||||
# HelmRepositories declared in kubernetes/flux/repos/.
|
# HelmRepositories declared in kubernetes/apps/dev/local/repos/.
|
||||||
# APP_WIRING below carries only the dev-specific concerns Flux doesn't have:
|
# APP_WIRING below carries only the dev-specific concerns Flux doesn't have:
|
||||||
# which locally-built image to inject, deploy ordering, and pod-readiness quirks.
|
# which locally-built image to inject, deploy ordering, and pod-readiness quirks.
|
||||||
#
|
#
|
||||||
@@ -228,15 +228,15 @@ docker_build(
|
|||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
local_resource(
|
local_resource(
|
||||||
'helm-deps',
|
'helm-deps',
|
||||||
cmd='rm -rf charts/yucca-api/charts charts/yucca-admin-api/charts charts/yucca-metrics-worker/charts charts/web/charts charts/michael/charts charts/mock-oidc/charts && for d in charts/yucca-api charts/yucca-admin-api charts/yucca-metrics-worker charts/web charts/michael charts/mock-oidc; do (cd $d && helm dependency build); done',
|
cmd='rm -rf charts/apps/yucca-api/charts charts/apps/yucca-admin-api/charts charts/apps/yucca-metrics-worker/charts charts/apps/web/charts charts/apps/michael/charts charts/dev/mock-oidc/charts && for d in charts/apps/yucca-api charts/apps/yucca-admin-api charts/apps/yucca-metrics-worker charts/apps/web charts/apps/michael charts/dev/mock-oidc; do (cd $d && helm dependency build); done',
|
||||||
deps=[
|
deps=[
|
||||||
'charts/yucca-api',
|
'charts/apps/yucca-api',
|
||||||
'charts/yucca-admin-api',
|
'charts/apps/yucca-admin-api',
|
||||||
'charts/yucca-metrics-worker',
|
'charts/apps/yucca-metrics-worker',
|
||||||
'charts/web',
|
'charts/apps/web',
|
||||||
'charts/michael',
|
'charts/apps/michael',
|
||||||
'charts/mock-oidc',
|
'charts/dev/mock-oidc',
|
||||||
'charts/yucca-common',
|
'charts/lib/yucca-common',
|
||||||
],
|
],
|
||||||
# `helm dependency build` rewrites these; if Tilt watches them we re-enter
|
# `helm dependency build` rewrites these; if Tilt watches them we re-enter
|
||||||
# an infinite rebuild loop.
|
# an infinite rebuild loop.
|
||||||
@@ -269,7 +269,7 @@ APP_WIRING = {
|
|||||||
# RGW user), so Tilt's pod tracking would hang at "pending". Mark ready on
|
# RGW user), so Tilt's pod tracking would hang at "pending". Mark ready on
|
||||||
# apply; michael still waits on this resource for ordering.
|
# apply; michael still waits on this resource for ordering.
|
||||||
'yucca-object-user': {'build': None, 'deps': ['rook-ceph-cluster'], 'pod_readiness': 'ignore'},
|
'yucca-object-user': {'build': None, 'deps': ['rook-ceph-cluster'], 'pod_readiness': 'ignore'},
|
||||||
# Shares charts/ceph-objectuser with yucca-object-user; its userName/caps
|
# Shares charts/platform/ceph-objectuser with yucca-object-user; its userName/caps
|
||||||
# come from the HelmRelease values, so dev must apply them (dev_values) or
|
# come from the HelmRelease values, so dev must apply them (dev_values) or
|
||||||
# both releases would default to userName=michael and collide.
|
# both releases would default to userName=michael and collide.
|
||||||
'yucca-metrics-object-user': {'build': None, 'deps': ['rook-ceph-cluster'], 'pod_readiness': 'ignore', 'dev_values': True},
|
'yucca-metrics-object-user': {'build': None, 'deps': ['rook-ceph-cluster'], 'pod_readiness': 'ignore', 'dev_values': True},
|
||||||
@@ -286,9 +286,9 @@ APP_WIRING = {
|
|||||||
}
|
}
|
||||||
|
|
||||||
def discover_helm_repos():
|
def discover_helm_repos():
|
||||||
"""HelmRepository name -> URL, from kubernetes/flux/repos/."""
|
"""HelmRepository name -> URL, from kubernetes/apps/dev/local/repos/."""
|
||||||
repos = {}
|
repos = {}
|
||||||
for path in listdir('kubernetes/flux/repos'):
|
for path in listdir('kubernetes/apps/dev/local/repos'):
|
||||||
if not path.endswith('.yaml'):
|
if not path.endswith('.yaml'):
|
||||||
continue
|
continue
|
||||||
doc = read_yaml(path)
|
doc = read_yaml(path)
|
||||||
@@ -302,10 +302,10 @@ def discover_apps():
|
|||||||
for path in listdir('kubernetes/apps', recursive=True):
|
for path in listdir('kubernetes/apps', recursive=True):
|
||||||
if not path.endswith('/helmrelease.yaml'):
|
if not path.endswith('/helmrelease.yaml'):
|
||||||
continue
|
continue
|
||||||
# The o11y-style GitOps tree (apps/base + apps/{staging,production}
|
# The o11y-style GitOps tree (apps/base + the real-cluster overlays
|
||||||
# overlays) is reconciled by Flux on the real clusters — Tilt deploys
|
# apps/<partition>/<region>) is reconciled by Flux on the real clusters —
|
||||||
# only the local dev-mirror tree (apps/<namespace>/<app>/app/).
|
# Tilt deploys only the local dev-mirror tree under apps/dev/local/.
|
||||||
if '/apps/base/' in path or '/apps/staging/' in path or '/apps/production/' in path:
|
if '/apps/dev/local/' not in path:
|
||||||
continue
|
continue
|
||||||
hr = read_yaml(path)
|
hr = read_yaml(path)
|
||||||
if not hr or hr.get('kind') != 'HelmRelease':
|
if not hr or hr.get('kind') != 'HelmRelease':
|
||||||
@@ -340,6 +340,11 @@ def wiring_for(app):
|
|||||||
return wiring
|
return wiring
|
||||||
|
|
||||||
LOCAL_APPS, REMOTE_APPS = discover_apps()
|
LOCAL_APPS, REMOTE_APPS = discover_apps()
|
||||||
|
# Guard against a silently-empty deploy: if the dev-mirror allow-list path ever
|
||||||
|
# moves again, discover_apps() would return nothing and Tilt would come up empty
|
||||||
|
# instead of failing loudly here.
|
||||||
|
if not LOCAL_APPS and not REMOTE_APPS:
|
||||||
|
fail("discover_apps() found no HelmReleases under kubernetes/apps/dev/local/ — has the dev-mirror tree moved?")
|
||||||
HELM_REPOS = discover_helm_repos()
|
HELM_REPOS = discover_helm_repos()
|
||||||
|
|
||||||
for repo_name, url in HELM_REPOS.items():
|
for repo_name, url in HELM_REPOS.items():
|
||||||
@@ -386,7 +391,7 @@ for app in LOCAL_APPS:
|
|||||||
image_keys=[('image.repository', 'image.tag')] if builds else [],
|
image_keys=[('image.repository', 'image.tag')] if builds else [],
|
||||||
resource_deps=['helm-deps'] + wiring['deps'] + extra_deps,
|
resource_deps=['helm-deps'] + wiring['deps'] + extra_deps,
|
||||||
labels=['app'],
|
labels=['app'],
|
||||||
deps=[app.chart, 'charts/yucca-common'],
|
deps=[app.chart, 'charts/lib/yucca-common'],
|
||||||
pod_readiness=wiring.get('pod_readiness', ''),
|
pod_readiness=wiring.get('pod_readiness', ''),
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|||||||
@@ -11,19 +11,20 @@ hardware/
|
|||||||
backups/
|
backups/
|
||||||
exports/
|
exports/
|
||||||
|
|
||||||
# TF-generated artifacts — do not commit; re-render with `tofu apply` in tf/deployment/<env>/ceph/
|
# TF-generated artifacts — do not commit; re-render with `tofu apply` in tf/deployment/<partition>/<region>/ceph/
|
||||||
inventories/*/inventory.ini
|
# Inventories are region-scoped with a friendly cluster leaf: inventories/<partition>-<region>/<cluster>/
|
||||||
inventories/*/inventory-provision.ini
|
inventories/*/*/inventory.ini
|
||||||
inventories/*/inventory-destroy.ini
|
inventories/*/*/inventory-provision.ini
|
||||||
inventories/*/secrets.yml.tpl
|
inventories/*/*/inventory-destroy.ini
|
||||||
inventories/*/group_vars/all/operators.yml
|
inventories/*/*/secrets.yml.tpl
|
||||||
|
inventories/*/*/group_vars/all/operators.yml
|
||||||
|
|
||||||
# host_vars IS committed (per-node hardware facts — stable, part of the inventory
|
# host_vars IS committed (per-node hardware facts — stable, part of the inventory
|
||||||
# source of truth). Operator-local overrides use host_vars/*.local.yml if needed.
|
# source of truth). Operator-local overrides use host_vars/*.local.yml if needed.
|
||||||
inventories/*/host_vars/*.local.yml
|
inventories/*/*/host_vars/*.local.yml
|
||||||
|
|
||||||
# Rendered installimage scripts (source of truth is the .tpl file).
|
# Rendered installimage scripts (source of truth is the .tpl file).
|
||||||
inventories/*/installimage/post-install.sh
|
inventories/*/*/installimage/post-install.sh
|
||||||
|
|
||||||
# Python virtualenv (mise setup) and bytecode
|
# Python virtualenv (mise setup) and bytecode
|
||||||
.venv/
|
.venv/
|
||||||
|
|||||||
@@ -8,7 +8,7 @@ PATH = "{{config_root}}/.venv/bin:{{env.PATH}}"
|
|||||||
# Note: CEPH_ENV is intentionally NOT declared here. mise's [env] block
|
# Note: CEPH_ENV is intentionally NOT declared here. mise's [env] block
|
||||||
# overrides shell-exported values, which silently sends operators to the
|
# overrides shell-exported values, which silently sends operators to the
|
||||||
# wrong cluster. Operators must export CEPH_ENV once per shell session:
|
# wrong cluster. Operators must export CEPH_ENV once per shell session:
|
||||||
# export CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini
|
# export CEPH_ENV=inventories/staging-austin/sietch/inventory.ini
|
||||||
# Inventory files are TF-generated — `mise run tf:apply` (from yucca root) if missing.
|
# Inventory files are TF-generated — `mise run tf:apply` (from yucca root) if missing.
|
||||||
|
|
||||||
[tasks.setup]
|
[tasks.setup]
|
||||||
@@ -63,7 +63,7 @@ run = """
|
|||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
# Syntax-check only parses YAML — any valid inventory works. Default to
|
# Syntax-check only parses YAML — any valid inventory works. Default to
|
||||||
# sietch when CEPH_ENV isn't inline-prefixed; the parse is identical.
|
# sietch when CEPH_ENV isn't inline-prefixed; the parse is identical.
|
||||||
CEPH_ENV="${CEPH_ENV:-inventories/sietch-ceph.staging.austin.int/inventory.ini}"
|
CEPH_ENV="${CEPH_ENV:-inventories/staging-austin/sietch/inventory.ini}"
|
||||||
for pb in *.yml; do
|
for pb in *.yml; do
|
||||||
case "$pb" in
|
case "$pb" in
|
||||||
requirements.yml|ansible-navigator.yml) continue ;;
|
requirements.yml|ansible-navigator.yml) continue ;;
|
||||||
@@ -123,9 +123,9 @@ description = "Destroy Ceph cluster (requires confirmation)"
|
|||||||
run = """
|
run = """
|
||||||
#!/usr/bin/env bash
|
#!/usr/bin/env bash
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
CEPH_ENV_DIR=$(dirname "$CEPH_ENV")
|
CEPH_ENV_DIR=$(dirname "$CEPH_ENV") # e.g., inventories/staging-austin/sietch
|
||||||
CLUSTER_ID=$(basename "$CEPH_ENV_DIR") # e.g., sietch-ceph.staging.austin.int
|
REGION_SLUG=$(basename "$(dirname "$CEPH_ENV_DIR")") # e.g., staging-austin (<partition>-<region>)
|
||||||
DOMAIN=${CLUSTER_ID#*-ceph.}.futo.cloud # e.g., dev.austin.int.futo.cloud
|
DOMAIN="${REGION_SLUG/-/.}.int.futo.cloud" # e.g., staging.austin.int.futo.cloud
|
||||||
DESTROY_INV="$CEPH_ENV_DIR/inventory-destroy.ini"
|
DESTROY_INV="$CEPH_ENV_DIR/inventory-destroy.ini"
|
||||||
echo "Usage: scripts/ansible-play.sh destroy-ceph.yml \\"
|
echo "Usage: scripts/ansible-play.sh destroy-ceph.yml \\"
|
||||||
echo " -e yes_destroy_ceph=true -e destroy_target_domain=$DOMAIN"
|
echo " -e yes_destroy_ceph=true -e destroy_target_domain=$DOMAIN"
|
||||||
|
|||||||
@@ -59,8 +59,8 @@ mise trust
|
|||||||
mise run setup
|
mise run setup
|
||||||
|
|
||||||
# Render cluster inventories + secrets templates (run once, or after any
|
# Render cluster inventories + secrets templates (run once, or after any
|
||||||
# change to tf/deployment/staging/ceph/clusters.auto.tfvars)
|
# change to tf/deployment/staging/austin/ceph/clusters.auto.tfvars)
|
||||||
cd ../../tf/deployment/staging/ceph && tofu init && tofu apply && cd -
|
cd ../../tf/deployment/staging/austin/ceph && tofu init && tofu apply && cd -
|
||||||
```
|
```
|
||||||
|
|
||||||
This runs:
|
This runs:
|
||||||
@@ -83,7 +83,7 @@ It points to an **inventory file** (not a directory):
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Inline prefix — required for `mise run` invocations:
|
# Inline prefix — required for `mise run` invocations:
|
||||||
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini mise run preflight
|
CEPH_ENV=inventories/staging-austin/sietch/inventory.ini mise run preflight
|
||||||
```
|
```
|
||||||
|
|
||||||
**`export CEPH_ENV=...` does NOT work with `mise run`.** mise's `[env]`
|
**`export CEPH_ENV=...` does NOT work with `mise run`.** mise's `[env]`
|
||||||
@@ -97,7 +97,7 @@ For multiple commands against the same cluster, set a local (non-exported)
|
|||||||
shell variable and inline-prefix each invocation:
|
shell variable and inline-prefix each invocation:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
CE=inventories/sietch-ceph.staging.austin.int/inventory.ini
|
CE=inventories/staging-austin/sietch/inventory.ini
|
||||||
CEPH_ENV=$CE mise run preflight
|
CEPH_ENV=$CE mise run preflight
|
||||||
CEPH_ENV=$CE mise run status
|
CEPH_ENV=$CE mise run status
|
||||||
CEPH_ENV=$CE mise run deploy
|
CEPH_ENV=$CE mise run deploy
|
||||||
@@ -106,7 +106,7 @@ CEPH_ENV=$CE mise run deploy
|
|||||||
Calling scripts directly (e.g., `scripts/preflight.sh`) DOES respect
|
Calling scripts directly (e.g., `scripts/preflight.sh`) DOES respect
|
||||||
shell `export` — it's only the `mise run` path that filters the env.
|
shell `export` — it's only the `mise run` path that filters the env.
|
||||||
|
|
||||||
Cluster identity is declared in `tf/deployment/staging/ceph/clusters.auto.tfvars`
|
Cluster identity is declared in `tf/deployment/staging/austin/ceph/clusters.auto.tfvars`
|
||||||
(keyed by short cluster name). TF renders the directory name, inventory
|
(keyed by short cluster name). TF renders the directory name, inventory
|
||||||
file, and secrets template from that entry. `CEPH_ENV` is just a pointer
|
file, and secrets template from that entry. `CEPH_ENV` is just a pointer
|
||||||
to the rendered inventory file; wrappers like `scripts/ansible-play.sh`
|
to the rendered inventory file; wrappers like `scripts/ansible-play.sh`
|
||||||
@@ -114,11 +114,11 @@ and the `destroy` mise task extract the cluster name from its path for
|
|||||||
convenience:
|
convenience:
|
||||||
|
|
||||||
```
|
```
|
||||||
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini
|
CEPH_ENV=inventories/staging-austin/sietch/inventory.ini
|
||||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
^^^^^^^^^^^^^^ ^^^^^^
|
||||||
inventory dir = sietch-ceph.staging.austin.int (rendered by TF)
|
| cluster name = sietch (map key in clusters.auto.tfvars)
|
||||||
cluster name = sietch (map key in clusters.auto.tfvars)
|
region slug = staging-austin (<partition>-<region>, rendered by TF)
|
||||||
domain = dev.austin.int.futo.cloud (domain field in tfvars)
|
domain = staging.austin.int.futo.cloud (domain field in tfvars)
|
||||||
```
|
```
|
||||||
|
|
||||||
Running `mise run tf:apply` regenerates `inventories/<cluster>/inventory.ini`
|
Running `mise run tf:apply` regenerates `inventories/<cluster>/inventory.ini`
|
||||||
@@ -158,7 +158,7 @@ mise run tf:apply
|
|||||||
```
|
```
|
||||||
|
|
||||||
Re-renders `inventory.ini` and `secrets.yml.tpl` for every cluster declared
|
Re-renders `inventory.ini` and `secrets.yml.tpl` for every cluster declared
|
||||||
in `tf/deployment/staging/ceph/clusters.auto.tfvars`. Required after any cluster-
|
in `tf/deployment/staging/austin/ceph/clusters.auto.tfvars`. Required after any cluster-
|
||||||
spec edit.
|
spec edit.
|
||||||
|
|
||||||
### 5. Dry-run against real nodes
|
### 5. Dry-run against real nodes
|
||||||
@@ -337,7 +337,7 @@ scripts/ansible-play.sh deploy-ceph.yml --tags bootstrap
|
|||||||
| `destroy` | `mise run destroy` | Destroy cluster (interactive confirmation) |
|
| `destroy` | `mise run destroy` | Destroy cluster (interactive confirmation) |
|
||||||
|
|
||||||
Inventory rendering and secret-item management are TF responsibilities
|
Inventory rendering and secret-item management are TF responsibilities
|
||||||
— `tofu apply` in `tf/deployment/staging/ceph/` renders `inventory.ini` and
|
— `tofu apply` in `tf/deployment/staging/austin/ceph/` renders `inventory.ini` and
|
||||||
`secrets.yml.tpl` for every cluster declared in `clusters.auto.tfvars`.
|
`secrets.yml.tpl` for every cluster declared in `clusters.auto.tfvars`.
|
||||||
Cluster secrets live in `yucca_tf_dev` (see [docs/secrets.md](docs/secrets.md)).
|
Cluster secrets live in `yucca_tf_dev` (see [docs/secrets.md](docs/secrets.md)).
|
||||||
|
|
||||||
@@ -347,11 +347,11 @@ Cluster secrets live in `yucca_tf_dev` (see [docs/secrets.md](docs/secrets.md)).
|
|||||||
yucca/
|
yucca/
|
||||||
├── tf/ # Terraform state + secrets + rendering (authoritative)
|
├── tf/ # Terraform state + secrets + rendering (authoritative)
|
||||||
│ ├── shared/modules/ceph-cluster/ # Module: per-cluster orchestration + rendering
|
│ ├── shared/modules/ceph-cluster/ # Module: per-cluster orchestration + rendering
|
||||||
│ └── deployment/staging/ceph/ # Cluster declarations + tofu apply target
|
│ └── deployment/staging/austin/ceph/ # Cluster declarations + tofu apply target
|
||||||
└── ansible/ceph/ # This directory
|
└── ansible/ceph/ # This directory
|
||||||
├── *.yml # Top-level playbooks (site.yml, deploy-ceph.yml, etc.)
|
├── *.yml # Top-level playbooks (site.yml, deploy-ceph.yml, etc.)
|
||||||
├── inventories/
|
├── inventories/
|
||||||
│ └── <cluster>-ceph.<env>.<dc>.<provider>/
|
│ └── <partition>-<region>/<cluster>/
|
||||||
│ ├── inventory.ini # TF-rendered (gitignored)
|
│ ├── inventory.ini # TF-rendered (gitignored)
|
||||||
│ ├── secrets.yml.tpl # TF-rendered, consumed by op inject (gitignored)
|
│ ├── secrets.yml.tpl # TF-rendered, consumed by op inject (gitignored)
|
||||||
│ ├── group_vars/all/
|
│ ├── group_vars/all/
|
||||||
|
|||||||
@@ -7,9 +7,9 @@ scaffolding are provisioned from `yucca/tf/` (see `../../tf/`).
|
|||||||
|
|
||||||
| Cluster | Domain | Location | Hardware | Nodes |
|
| Cluster | Domain | Location | Hardware | Nodes |
|
||||||
|---------|--------|----------|----------|-------|
|
|---------|--------|----------|----------|-------|
|
||||||
| **sietch** | `dev.austin.int.futo.cloud` | Austin DC | Dell R730xd | 3 |
|
| **sietch** | `staging.austin.int.futo.cloud` | Austin DC | Dell R730xd | 3 |
|
||||||
|
|
||||||
Clusters are declared in `yucca/tf/deployment/staging/ceph/clusters.auto.tfvars`;
|
Clusters are declared in `yucca/tf/deployment/staging/austin/ceph/clusters.auto.tfvars`;
|
||||||
`tofu apply` renders `inventories/<cluster>/inventory.ini` and
|
`tofu apply` renders `inventories/<cluster>/inventory.ini` and
|
||||||
`secrets.yml.tpl` per cluster. The `CEPH_ENV` variable selects the active
|
`secrets.yml.tpl` per cluster. The `CEPH_ENV` variable selects the active
|
||||||
cluster for any `mise run` or direct ansible invocation.
|
cluster for any `mise run` or direct ansible invocation.
|
||||||
@@ -41,7 +41,7 @@ data flow, and design rationale.
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
# 1. Render cluster inventories + secrets templates (once, from yucca/tf/)
|
# 1. Render cluster inventories + secrets templates (once, from yucca/tf/)
|
||||||
(cd ../../tf/deployment/staging/ceph && tofu init && tofu apply)
|
(cd ../../tf/deployment/staging/austin/ceph && tofu init && tofu apply)
|
||||||
|
|
||||||
# 2. Set up the ansible side
|
# 2. Set up the ansible side
|
||||||
mise trust && mise run setup # bootstrap dev environment
|
mise trust && mise run setup # bootstrap dev environment
|
||||||
@@ -49,7 +49,7 @@ mise trust && mise run setup # bootstrap dev environment
|
|||||||
# 3. Run mise tasks against the target cluster. CEPH_ENV must be set
|
# 3. Run mise tasks against the target cluster. CEPH_ENV must be set
|
||||||
# inline (NOT via `export`) — see docs/scripts.md "Setting CEPH_ENV"
|
# inline (NOT via `export`) — see docs/scripts.md "Setting CEPH_ENV"
|
||||||
# for why mise's [env] block strips shell exports.
|
# for why mise's [env] block strips shell exports.
|
||||||
CE=inventories/sietch-ceph.staging.austin.int/inventory.ini
|
CE=inventories/staging-austin/sietch/inventory.ini
|
||||||
CEPH_ENV=$CE mise run preflight # TF artifacts + 1P + SSH + connectivity
|
CEPH_ENV=$CE mise run preflight # TF artifacts + 1P + SSH + connectivity
|
||||||
CEPH_ENV=$CE mise run status # read-only cluster health check
|
CEPH_ENV=$CE mise run status # read-only cluster health check
|
||||||
CEPH_ENV=$CE mise run drift # configuration drift detection
|
CEPH_ENV=$CE mise run drift # configuration drift detection
|
||||||
@@ -114,7 +114,7 @@ See [CONTRIBUTING.md](CONTRIBUTING.md) for the full development workflow.
|
|||||||
| `bench-rados` | RADOS bench (raw cluster I/O) |
|
| `bench-rados` | RADOS bench (raw cluster I/O) |
|
||||||
|
|
||||||
Inventory scaffolding + secret-item provisioning live in `yucca/tf/` — run
|
Inventory scaffolding + secret-item provisioning live in `yucca/tf/` — run
|
||||||
`tofu apply` in `tf/deployment/staging/ceph/` to (re-)render
|
`tofu apply` in `tf/deployment/staging/austin/ceph/` to (re-)render
|
||||||
`inventories/<cluster>/inventory.ini` and `secrets.yml.tpl`.
|
`inventories/<cluster>/inventory.ini` and `secrets.yml.tpl`.
|
||||||
|
|
||||||
## Documentation
|
## Documentation
|
||||||
|
|||||||
@@ -11,11 +11,11 @@
|
|||||||
#
|
#
|
||||||
# Usage:
|
# Usage:
|
||||||
# scripts/ansible-play.sh destroy-ceph.yml \
|
# scripts/ansible-play.sh destroy-ceph.yml \
|
||||||
# -e "yes_destroy_ceph=true destroy_target_domain=dev.austin.int.futo.cloud"
|
# -e "yes_destroy_ceph=true destroy_target_domain=staging.austin.int.futo.cloud"
|
||||||
#
|
#
|
||||||
# To run a specific phase only:
|
# To run a specific phase only:
|
||||||
# scripts/ansible-play.sh destroy-ceph.yml -e "yes_destroy_ceph=true destroy_target_domain=dev.austin.int.futo.cloud" --tags purge
|
# scripts/ansible-play.sh destroy-ceph.yml -e "yes_destroy_ceph=true destroy_target_domain=staging.austin.int.futo.cloud" --tags purge
|
||||||
# scripts/ansible-play.sh destroy-ceph.yml -e "yes_destroy_ceph=true destroy_target_domain=dev.austin.int.futo.cloud" --tags cleanup
|
# scripts/ansible-play.sh destroy-ceph.yml -e "yes_destroy_ceph=true destroy_target_domain=staging.austin.int.futo.cloud" --tags cleanup
|
||||||
|
|
||||||
- name: Destroy Ceph Tentacle cluster
|
- name: Destroy Ceph Tentacle cluster
|
||||||
hosts: ceph_nodes
|
hosts: ceph_nodes
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# Adding a cluster
|
# Adding a cluster
|
||||||
|
|
||||||
Clusters are declared in `tf/deployment/<env>/ceph/clusters.auto.tfvars`.
|
Clusters are declared in `tf/deployment/<partition>/<region>/ceph/clusters.auto.tfvars`.
|
||||||
Every cluster-scoped concern — inventory file, hostname, 1P item names, SSH
|
Every cluster-scoped concern — inventory file, hostname, 1P item names, SSH
|
||||||
key path, secrets template — is derived from that one entry. Most of what
|
key path, secrets template — is derived from that one entry. Most of what
|
||||||
this walkthrough describes is editing that file and running
|
this walkthrough describes is editing that file and running
|
||||||
@@ -35,7 +35,7 @@ to `ceph`). Every Ceph-project inventory grep-matches `*-ceph.*` regardless
|
|||||||
of datacenter or environment.
|
of datacenter or environment.
|
||||||
|
|
||||||
Existing examples:
|
Existing examples:
|
||||||
- `sietch-ceph.dev.austin.int/` — Austin DC, internal network, dev
|
- `staging-austin/sietch/` — Austin DC, internal network, dev
|
||||||
|
|
||||||
Future environments land as siblings: `*-ceph.staging.<dc>.<provider>/`,
|
Future environments land as siblings: `*-ceph.staging.<dc>.<provider>/`,
|
||||||
`*-ceph.prod.<dc>.<provider>/`.
|
`*-ceph.prod.<dc>.<provider>/`.
|
||||||
@@ -59,7 +59,7 @@ auto-picked from the 923-word wordlist — see [docs/naming.md](naming.md#host-n
|
|||||||
|
|
||||||
### 2. Declare the cluster in TF
|
### 2. Declare the cluster in TF
|
||||||
|
|
||||||
Edit `tf/deployment/<env>/ceph/clusters.auto.tfvars` and add an entry.
|
Edit `tf/deployment/<partition>/<region>/ceph/clusters.auto.tfvars` and add an entry.
|
||||||
Working example for a hypothetical `mesa` cluster at Hetzner Falkenstein:
|
Working example for a hypothetical `mesa` cluster at Hetzner Falkenstein:
|
||||||
|
|
||||||
```hcl
|
```hcl
|
||||||
@@ -105,7 +105,7 @@ mise run tf:apply
|
|||||||
Or, for a non-default stack:
|
Or, for a non-default stack:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
TF_STACK_DIR=tf/deployment/<env>/ceph mise run tf:apply
|
TF_STACK_DIR=tf/deployment/<partition>/<region>/ceph mise run tf:apply
|
||||||
```
|
```
|
||||||
|
|
||||||
This creates (per the module's `rendering.tf`):
|
This creates (per the module's `rendering.tf`):
|
||||||
@@ -123,12 +123,12 @@ All of these are gitignored — re-run `mise run tf:apply` after any
|
|||||||
Hand-maintained, committed. Copy the closer existing analogue as a starting
|
Hand-maintained, committed. Copy the closer existing analogue as a starting
|
||||||
point:
|
point:
|
||||||
|
|
||||||
- **Bare-metal cluster:** copy from `sietch-ceph.dev.austin.int/group_vars/all/vars.yml`
|
- **Bare-metal cluster:** copy from `staging-austin/sietch/group_vars/all/vars.yml`
|
||||||
- **Hetzner/single-NIC cluster:** start from the sietch vars and adjust for
|
- **Hetzner/single-NIC cluster:** start from the sietch vars and adjust for
|
||||||
the NVMe-RAID shape (public /32, no bond/ProxyJump, installimage-owned LVM).
|
the NVMe-RAID shape (public /32, no bond/ProxyJump, installimage-owned LVM).
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
cp inventories/sietch-ceph.dev.austin.int/group_vars/all/vars.yml \
|
cp inventories/staging-austin/sietch/group_vars/all/vars.yml \
|
||||||
inventories/mesa-ceph.dev.fsn.htz/group_vars/all/vars.yml
|
inventories/mesa-ceph.dev.fsn.htz/group_vars/all/vars.yml
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -322,7 +322,7 @@ CEPH_ENV=inventories/mesa-ceph.dev.fsn.htz/
|
|||||||
Default is set in `.mise.toml` (`sietch` in dev). Override per-command:
|
Default is set in `.mise.toml` (`sietch` in dev). Override per-command:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory.ini mise run status
|
CEPH_ENV=inventories/staging-austin/sietch/inventory.ini mise run status
|
||||||
```
|
```
|
||||||
|
|
||||||
`scripts/ansible-play.sh` derives the secrets template path from `CEPH_ENV`
|
`scripts/ansible-play.sh` derives the secrets template path from `CEPH_ENV`
|
||||||
|
|||||||
@@ -51,20 +51,20 @@ its environment from the same source — directory layout — so dev / staging
|
|||||||
|
|
||||||
| Layer | dev (today) | staging (planned) | prod (planned) |
|
| Layer | dev (today) | staging (planned) | prod (planned) |
|
||||||
|--------------|----------------------------------------------------------------|---------------------------------------------------------|------------------------------------------------------|
|
|--------------|----------------------------------------------------------------|---------------------------------------------------------|------------------------------------------------------|
|
||||||
| TF stack dir | `tf/deployment/dev/ceph/` | `tf/deployment/staging/ceph/` | `tf/deployment/prod/ceph/` |
|
| TF stack dir | `tf/deployment/dev/ceph/` | `tf/deployment/staging/austin/ceph/` | `tf/deployment/prod/ceph/` |
|
||||||
| TF state key | `ceph/dev/ceph/terraform.tfstate` | `ceph/staging/ceph/terraform.tfstate` | `ceph/prod/ceph/terraform.tfstate` |
|
| TF state key | `ceph/dev/ceph/terraform.tfstate` | `ceph/staging/ceph/terraform.tfstate` | `ceph/prod/ceph/terraform.tfstate` |
|
||||||
| 1P vaults | `yucca_tf_dev` (live) · `yucca_tf_dev_manual` (human-fillable) | `yucca_tf_staging` · `yucca_tf_staging_manual` | `yucca_tf` (live) · `yucca_tf_prod_manual` |
|
| 1P vaults | `yucca_tf_dev` (live) · `yucca_tf_dev_manual` (human-fillable) | `yucca_tf_staging` · `yucca_tf_staging_manual` | `yucca_tf` (live) · `yucca_tf_prod_manual` |
|
||||||
| Ansible inv | `inventories/<cluster>-ceph.dev.<dc>.<provider>/` | `inventories/<cluster>-ceph.staging.<dc>.<provider>/` | `inventories/<cluster>-ceph.prod.<dc>.<provider>/` |
|
| Ansible inv | `inventories/<cluster>-ceph.dev.<dc>.<provider>/` | `inventories/<cluster>-ceph.staging.<dc>.<provider>/` | `inventories/<cluster>-ceph.prod.<dc>.<provider>/` |
|
||||||
| mise default | `CEPH_ENV=...sietch-ceph.dev.austin.int/inventory.ini` | overridden via env at invocation | overridden via env at invocation |
|
| mise default | `CEPH_ENV=...staging-austin/sietch/inventory.ini` | overridden via env at invocation | overridden via env at invocation |
|
||||||
|
|
||||||
Today the only deployed environment is dev (sietch). Adding
|
Today the only deployed environment is dev (sietch). Adding
|
||||||
staging/prod is purely additive: create the matching `tf/deployment/<env>/ceph/`
|
staging/prod is purely additive: create the matching `tf/deployment/<partition>/<region>/ceph/`
|
||||||
directory, populate `clusters.auto.tfvars`, and the same module + Ansible
|
directory, populate `clusters.auto.tfvars`, and the same module + Ansible
|
||||||
roles + mise tasks work unchanged. The state backend key path, 1P vault
|
roles + mise tasks work unchanged. The state backend key path, 1P vault
|
||||||
selection, and inventory directory naming all derive from the env segment.
|
selection, and inventory directory naming all derive from the env segment.
|
||||||
|
|
||||||
`TF_STACK_DIR` is the operator-side override for `mise run tf:*` tasks; it
|
`TF_STACK_DIR` is the operator-side override for `mise run tf:*` tasks; it
|
||||||
defaults to `tf/deployment/staging/ceph` and points at any sibling stack directory.
|
defaults to `tf/deployment/staging/austin/ceph` and points at any sibling stack directory.
|
||||||
`CEPH_ENV` is the matching override for Ansible — points at the rendered
|
`CEPH_ENV` is the matching override for Ansible — points at the rendered
|
||||||
`inventory.ini` for the cluster you intend to operate on.
|
`inventory.ini` for the cluster you intend to operate on.
|
||||||
|
|
||||||
@@ -164,7 +164,7 @@ Each top-level key becomes a cluster:
|
|||||||
|
|
||||||
```hcl
|
```hcl
|
||||||
sietch = {
|
sietch = {
|
||||||
domain = "dev.austin.int.futo.cloud"
|
domain = "staging.austin.int.futo.cloud"
|
||||||
environment = "dev"
|
environment = "dev"
|
||||||
datacenter = "austin"
|
datacenter = "austin"
|
||||||
provider_code = "int"
|
provider_code = "int"
|
||||||
@@ -390,7 +390,7 @@ just those phases.
|
|||||||
|
|
||||||
```
|
```
|
||||||
inventories/
|
inventories/
|
||||||
sietch-ceph.staging.austin.int/ Austin staging cluster
|
staging-austin/sietch/ Austin staging cluster
|
||||||
inventory.ini TF-generated, gitignored
|
inventory.ini TF-generated, gitignored
|
||||||
inventory-destroy.ini TF-generated, gitignored
|
inventory-destroy.ini TF-generated, gitignored
|
||||||
inventory-provision.ini TF-generated, gitignored
|
inventory-provision.ini TF-generated, gitignored
|
||||||
@@ -486,15 +486,15 @@ env defaults.
|
|||||||
|
|
||||||
```toml
|
```toml
|
||||||
[env]
|
[env]
|
||||||
CEPH_ENV = "inventories/sietch-ceph.staging.austin.int/inventory.ini"
|
CEPH_ENV = "inventories/staging-austin/sietch/inventory.ini"
|
||||||
```
|
```
|
||||||
|
|
||||||
`TF_STACK_DIR` defaults to `tf/deployment/staging/ceph` inside each `tf:*` task.
|
`TF_STACK_DIR` defaults to `tf/deployment/staging/austin/ceph` inside each `tf:*` task.
|
||||||
Both are overridable per-invocation:
|
Both are overridable per-invocation:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
TF_STACK_DIR=tf/deployment/staging/ceph mise run tf:plan
|
TF_STACK_DIR=tf/deployment/staging/austin/ceph mise run tf:plan
|
||||||
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini mise run status
|
CEPH_ENV=inventories/staging-austin/sietch/inventory.ini mise run status
|
||||||
```
|
```
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|||||||
@@ -21,7 +21,7 @@ A Hetzner NVMe-RAID host would attach over a public /32 with a single NIC
|
|||||||
and direct SSH (no bond, no ProxyJump).
|
and direct SSH (no bond, no ProxyJump).
|
||||||
|
|
||||||
Per-node connection IPs (`bond_ip`) are declared in
|
Per-node connection IPs (`bond_ip`) are declared in
|
||||||
`tf/deployment/staging/ceph/clusters.auto.tfvars` and rendered by TF into the
|
`tf/deployment/staging/austin/ceph/clusters.auto.tfvars` and rendered by TF into the
|
||||||
cluster's `inventory.ini`. They're also mirrored into `host_vars/` for use
|
cluster's `inventory.ini`. They're also mirrored into `host_vars/` for use
|
||||||
by roles that need the IP as a variable (e.g., cephadm public-network
|
by roles that need the IP as a variable (e.g., cephadm public-network
|
||||||
resolution, dashboard URL construction).
|
resolution, dashboard URL construction).
|
||||||
|
|||||||
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
Every name derivable in this project — hostname, inventory directory, 1P
|
Every name derivable in this project — hostname, inventory directory, 1P
|
||||||
item title, SSH key filename — traces back to one entry per cluster in
|
item title, SSH key filename — traces back to one entry per cluster in
|
||||||
`tf/deployment/<env>/ceph/clusters.auto.tfvars`. TF's `ceph-cluster` module
|
`tf/deployment/<partition>/<region>/ceph/clusters.auto.tfvars`. TF's `ceph-cluster` module
|
||||||
assembles the rest.
|
assembles the rest.
|
||||||
|
|
||||||
For how naming fits the broader system see
|
For how naming fits the broader system see
|
||||||
@@ -195,7 +195,7 @@ The `-ceph` segment is hardcoded in the ceph-cluster module regardless of
|
|||||||
as `*-ceph.*` — even a hypothetical cluster with `role_in_hostname = "osd"`
|
as `*-ceph.*` — even a hypothetical cluster with `role_in_hostname = "osd"`
|
||||||
(hostnames `mesa-osd-*`) still renders `mesa-ceph.prod.fsn.htz/`.
|
(hostnames `mesa-osd-*`) still renders `mesa-ceph.prod.fsn.htz/`.
|
||||||
|
|
||||||
Defined in `tf/deployment/<env>/ceph/main.tf` (`local.inventory_dirs`).
|
Defined in `tf/deployment/<partition>/<region>/ceph/main.tf` (`local.inventory_dirs`).
|
||||||
|
|
||||||
## 1Password item naming
|
## 1Password item naming
|
||||||
|
|
||||||
|
|||||||
@@ -211,9 +211,9 @@ inventory file. Cluster identity is authoritative in
|
|||||||
`clusters.auto.tfvars`; `CEPH_ENV` is the runtime pointer.
|
`clusters.auto.tfvars`; `CEPH_ENV` is the runtime pointer.
|
||||||
|
|
||||||
```
|
```
|
||||||
CEPH_ENV = inventories/sietch-ceph.staging.austin.int/inventory.ini
|
CEPH_ENV = inventories/staging-austin/sietch/inventory.ini
|
||||||
|
|
|
|
||||||
dirname -> inventories/sietch-ceph.staging.austin.int
|
dirname -> inventories/staging-austin/sietch
|
||||||
|
|
|
|
||||||
+ "/secrets.yml.tpl" -> op inject input
|
+ "/secrets.yml.tpl" -> op inject input
|
||||||
```
|
```
|
||||||
@@ -225,8 +225,9 @@ is missing. See [scripts.md](scripts.md) for the full contract.
|
|||||||
**Destroy task** in `.mise.toml` extracts the domain for the safety gate:
|
**Destroy task** in `.mise.toml` extracts the domain for the safety gate:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
CLUSTER_ID=$(basename "$CEPH_ENV_DIR") # sietch-ceph.staging.austin.int
|
CEPH_ENV_DIR=$(dirname "$CEPH_ENV") # inventories/staging-austin/sietch
|
||||||
DOMAIN=${CLUSTER_ID#*-ceph.}.futo.cloud # staging.austin.int.futo.cloud
|
REGION_SLUG=$(basename "$(dirname "$CEPH_ENV_DIR")") # staging-austin (<partition>-<region>)
|
||||||
|
DOMAIN="${REGION_SLUG/-/.}.int.futo.cloud" # staging.austin.int.futo.cloud
|
||||||
```
|
```
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|||||||
@@ -15,7 +15,7 @@
|
|||||||
|
|
||||||
Pick an unused name, or omit `name` to let TF auto-pick from the
|
Pick an unused name, or omit `name` to let TF auto-pick from the
|
||||||
923-word list seeded per-cluster (see [docs/naming.md](../naming.md)).
|
923-word list seeded per-cluster (see [docs/naming.md](../naming.md)).
|
||||||
Edit `tf/deployment/staging/ceph/clusters.auto.tfvars` and append to the
|
Edit `tf/deployment/staging/austin/ceph/clusters.auto.tfvars` and append to the
|
||||||
target cluster's `hosts` list:
|
target cluster's `hosts` list:
|
||||||
|
|
||||||
```hcl
|
```hcl
|
||||||
@@ -45,7 +45,7 @@ their positions.
|
|||||||
## 2. Create host_vars
|
## 2. Create host_vars
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
cd inventories/sietch-ceph.staging.austin.int
|
cd inventories/staging-austin/sietch
|
||||||
cp host_vars/example.yml host_vars/sietch-ceph-<name>.yml
|
cp host_vars/example.yml host_vars/sietch-ceph-<name>.yml
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -67,7 +67,7 @@ hardware:
|
|||||||
After `tofu apply` in step 1, inspect the rendered inventory:
|
After `tofu apply` in step 1, inspect the rendered inventory:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
cat inventories/sietch-ceph.staging.austin.int/inventory.ini
|
cat inventories/staging-austin/sietch/inventory.ini
|
||||||
```
|
```
|
||||||
|
|
||||||
The new host should appear in:
|
The new host should appear in:
|
||||||
@@ -86,7 +86,7 @@ Bootstrap host is unchanged. TF never moves an existing bootstrap assignment.
|
|||||||
3. Run provisioning:
|
3. Run provisioning:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory-provision.ini \
|
CEPH_ENV=inventories/staging-austin/sietch/inventory-provision.ini \
|
||||||
scripts/ansible-play.sh provision.yml \
|
scripts/ansible-play.sh provision.yml \
|
||||||
-e confirm_wipe=true \
|
-e confirm_wipe=true \
|
||||||
--limit sietch-ceph-<name>
|
--limit sietch-ceph-<name>
|
||||||
@@ -98,7 +98,7 @@ CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory-provision.ini \
|
|||||||
ssh -i ~/.ssh/id_ed25519_sietch ansible-iac@sietch-ceph-<name> hostname -f
|
ssh -i ~/.ssh/id_ed25519_sietch ansible-iac@sietch-ceph-<name> hostname -f
|
||||||
```
|
```
|
||||||
|
|
||||||
Expected output: `sietch-ceph-<name>.dev.austin.int.futo.cloud`
|
Expected output: `sietch-ceph-<name>.staging.austin.int.futo.cloud`
|
||||||
|
|
||||||
### Hetzner (remote servers)
|
### Hetzner (remote servers)
|
||||||
|
|
||||||
|
|||||||
@@ -44,7 +44,7 @@ mise run tf:apply
|
|||||||
### Rendered file on disk doesn't match tfvars
|
### Rendered file on disk doesn't match tfvars
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
cat ansible/ceph/inventories/sietch-ceph.staging.austin.int/inventory.ini
|
cat ansible/ceph/inventories/staging-austin/sietch/inventory.ini
|
||||||
# says ansible_user=root but tfvars says ansible-iac
|
# says ansible_user=root but tfvars says ansible-iac
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
@@ -51,7 +51,7 @@ least-privilege principle.
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
cd ~/yucca/ansible/ceph
|
cd ~/yucca/ansible/ceph
|
||||||
export CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini
|
export CEPH_ENV=inventories/staging-austin/sietch/inventory.ini
|
||||||
mise run status # read-only smoke test
|
mise run status # read-only smoke test
|
||||||
mise run deploy # or any other task
|
mise run deploy # or any other task
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -46,7 +46,7 @@ the same bond IP via the same DHCP reservation / static config.
|
|||||||
# Boot the new box from the Debian 12 live image (same procedure as first
|
# Boot the new box from the Debian 12 live image (same procedure as first
|
||||||
# provision — see docs/runbooks/add-node.md for iDRAC steps).
|
# provision — see docs/runbooks/add-node.md for iDRAC steps).
|
||||||
|
|
||||||
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory-provision.ini \
|
CEPH_ENV=inventories/staging-austin/sietch/inventory-provision.ini \
|
||||||
scripts/ansible-play.sh provision.yml \
|
scripts/ansible-play.sh provision.yml \
|
||||||
-e confirm_wipe=true \
|
-e confirm_wipe=true \
|
||||||
--limit sietch-ceph-<name>
|
--limit sietch-ceph-<name>
|
||||||
|
|||||||
@@ -21,11 +21,11 @@ sudo openssl x509 -in /etc/ceph/rgw-ssl.crt -noout -subject -dates -ext subjectA
|
|||||||
Output shows:
|
Output shows:
|
||||||
|
|
||||||
```
|
```
|
||||||
subject=C = US, ST = Texas, L = Austin, O = FUTO, CN = s3.dev.austin.int.futo.cloud
|
subject=C = US, ST = Texas, L = Austin, O = FUTO, CN = s3.staging.austin.int.futo.cloud
|
||||||
notBefore=...
|
notBefore=...
|
||||||
notAfter=...
|
notAfter=...
|
||||||
X509v3 Subject Alternative Name:
|
X509v3 Subject Alternative Name:
|
||||||
DNS:s3.dev.austin.int.futo.cloud, DNS:*.s3.dev.austin.int.futo.cloud, ...
|
DNS:s3.staging.austin.int.futo.cloud, DNS:*.s3.staging.austin.int.futo.cloud, ...
|
||||||
```
|
```
|
||||||
|
|
||||||
## 2. Run the rotation playbook
|
## 2. Run the rotation playbook
|
||||||
@@ -75,7 +75,7 @@ sudo openssl x509 -in /etc/ceph/rgw-ssl.crt -noout -subject -dates -ext subjectA
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
# From a node in the cluster (self-signed cert)
|
# From a node in the cluster (self-signed cert)
|
||||||
curl -k https://s3.dev.austin.int.futo.cloud:443/
|
curl -k https://s3.staging.austin.int.futo.cloud:443/
|
||||||
```
|
```
|
||||||
|
|
||||||
Expected: XML response with `ListAllMyBucketsResult` or `AccessDenied`
|
Expected: XML response with `ListAllMyBucketsResult` or `AccessDenied`
|
||||||
@@ -104,7 +104,7 @@ verification disabled for self-signed certs).
|
|||||||
## Certificate configuration
|
## Certificate configuration
|
||||||
|
|
||||||
The cert parameters are controlled by these variables in
|
The cert parameters are controlled by these variables in
|
||||||
`inventories/sietch-ceph.staging.austin.int/group_vars/all/vars.yml`:
|
`inventories/staging-austin/sietch/group_vars/all/vars.yml`:
|
||||||
|
|
||||||
| Variable | Default | Purpose |
|
| Variable | Default | Purpose |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
@@ -115,7 +115,7 @@ The cert parameters are controlled by these variables in
|
|||||||
| `ceph_rgw_ssl_cert_subject_l` | `Austin` | Locality |
|
| `ceph_rgw_ssl_cert_subject_l` | `Austin` | Locality |
|
||||||
| `ceph_rgw_ssl_cert_subject_o` | `FUTO` | Organization |
|
| `ceph_rgw_ssl_cert_subject_o` | `FUTO` | Organization |
|
||||||
| `ceph_rgw_ssl_cert_email` | `yucca@futo.org` | Contact email |
|
| `ceph_rgw_ssl_cert_email` | `yucca@futo.org` | Contact email |
|
||||||
| `ceph_rgw_dns_name` | `s3.dev.austin.int.futo.cloud` | CN and primary SAN |
|
| `ceph_rgw_dns_name` | `s3.staging.austin.int.futo.cloud` | CN and primary SAN |
|
||||||
|
|
||||||
SANs are auto-generated from inventory: per-node FQDNs and bond IPs are
|
SANs are auto-generated from inventory: per-node FQDNs and bond IPs are
|
||||||
included so direct-host access validates.
|
included so direct-host access validates.
|
||||||
|
|||||||
@@ -28,7 +28,7 @@ deploy` picks up the new value via `op inject`.
|
|||||||
|
|
||||||
Replace `<CLUSTER>` with the cluster short name (e.g. `SIETCH`). The active vault name is
|
Replace `<CLUSTER>` with the cluster short name (e.g. `SIETCH`). The active vault name is
|
||||||
declared per-cluster in the `vault` field of the cluster's entry in
|
declared per-cluster in the `vault` field of the cluster's entry in
|
||||||
`tf/deployment/staging/ceph/clusters.auto.tfvars` — `yucca_tf_dev` for dev,
|
`tf/deployment/staging/austin/ceph/clusters.auto.tfvars` — `yucca_tf_dev` for dev,
|
||||||
future `yucca_tf_staging` / `yucca_tf` for staging/prod.
|
future `yucca_tf_staging` / `yucca_tf` for staging/prod.
|
||||||
|
|
||||||
## 1. Rotate in 1Password
|
## 1. Rotate in 1Password
|
||||||
@@ -192,7 +192,7 @@ Once the sietch-ceph service account lands and `secrets.tf.disabled` is
|
|||||||
re-enabled, rotations become a `terraform taint` + `apply`:
|
re-enabled, rotations become a `terraform taint` + `apply`:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
cd tf/deployment/staging/ceph
|
cd tf/deployment/staging/austin/ceph
|
||||||
terragrunt taint 'module.cluster["sietch"].onepassword_item.secret["dashboard"]'
|
terragrunt taint 'module.cluster["sietch"].onepassword_item.secret["dashboard"]'
|
||||||
terragrunt apply
|
terragrunt apply
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -9,8 +9,8 @@ wildcard TLS certificate on **port 443**.
|
|||||||
|
|
||||||
| Style | URL |
|
| Style | URL |
|
||||||
|---|---|
|
|---|---|
|
||||||
| Path-style | `https://s3.dev.austin.int.futo.cloud/<bucket>/<key>` |
|
| Path-style | `https://s3.staging.austin.int.futo.cloud/<bucket>/<key>` |
|
||||||
| Virtual-hosted | `https://<bucket>.s3.dev.austin.int.futo.cloud/<key>` |
|
| Virtual-hosted | `https://<bucket>.s3.staging.austin.int.futo.cloud/<key>` |
|
||||||
| Direct (per-node) | `https://10.10.10.90:443`, `https://10.10.10.91:443`, `https://10.10.10.92:443` |
|
| Direct (per-node) | `https://10.10.10.90:443`, `https://10.10.10.91:443`, `https://10.10.10.92:443` |
|
||||||
|
|
||||||
Region: **us-east-1**
|
Region: **us-east-1**
|
||||||
@@ -81,7 +81,7 @@ aws_secret_access_key = YOUR_SECRET_KEY
|
|||||||
```ini
|
```ini
|
||||||
[profile sietch]
|
[profile sietch]
|
||||||
region = us-east-1
|
region = us-east-1
|
||||||
endpoint_url = https://s3.dev.austin.int.futo.cloud
|
endpoint_url = https://s3.staging.austin.int.futo.cloud
|
||||||
s3 =
|
s3 =
|
||||||
signature_version = s3v4
|
signature_version = s3v4
|
||||||
addressing_style = path
|
addressing_style = path
|
||||||
@@ -121,7 +121,7 @@ urllib3.disable_warnings(urllib3.exceptions.InsecureRequestWarning)
|
|||||||
|
|
||||||
s3 = boto3.client(
|
s3 = boto3.client(
|
||||||
"s3",
|
"s3",
|
||||||
endpoint_url="https://s3.dev.austin.int.futo.cloud",
|
endpoint_url="https://s3.staging.austin.int.futo.cloud",
|
||||||
aws_access_key_id="YOUR_ACCESS_KEY",
|
aws_access_key_id="YOUR_ACCESS_KEY",
|
||||||
aws_secret_access_key="YOUR_SECRET_KEY",
|
aws_secret_access_key="YOUR_SECRET_KEY",
|
||||||
region_name="us-east-1",
|
region_name="us-east-1",
|
||||||
@@ -153,7 +153,7 @@ To use the CA bundle instead of disabling verification:
|
|||||||
```python
|
```python
|
||||||
s3 = boto3.client(
|
s3 = boto3.client(
|
||||||
"s3",
|
"s3",
|
||||||
endpoint_url="https://s3.dev.austin.int.futo.cloud",
|
endpoint_url="https://s3.staging.austin.int.futo.cloud",
|
||||||
aws_access_key_id="YOUR_ACCESS_KEY",
|
aws_access_key_id="YOUR_ACCESS_KEY",
|
||||||
aws_secret_access_key="YOUR_SECRET_KEY",
|
aws_secret_access_key="YOUR_SECRET_KEY",
|
||||||
region_name="us-east-1",
|
region_name="us-east-1",
|
||||||
@@ -170,7 +170,7 @@ s3 = boto3.client(
|
|||||||
```bash
|
```bash
|
||||||
export AWS_ACCESS_KEY_ID="YOUR_ACCESS_KEY"
|
export AWS_ACCESS_KEY_ID="YOUR_ACCESS_KEY"
|
||||||
export AWS_SECRET_ACCESS_KEY="YOUR_SECRET_KEY"
|
export AWS_SECRET_ACCESS_KEY="YOUR_SECRET_KEY"
|
||||||
export RESTIC_REPOSITORY="s3:https://s3.dev.austin.int.futo.cloud/restic-backups"
|
export RESTIC_REPOSITORY="s3:https://s3.staging.austin.int.futo.cloud/restic-backups"
|
||||||
|
|
||||||
# Init (first time)
|
# Init (first time)
|
||||||
restic init --option s3.region=us-east-1
|
restic init --option s3.region=us-east-1
|
||||||
@@ -202,16 +202,16 @@ needed from the application side.
|
|||||||
|
|
||||||
## DNS setup for virtual-hosted buckets
|
## DNS setup for virtual-hosted buckets
|
||||||
|
|
||||||
Virtual-hosted bucket addressing (e.g., `https://my-bucket.s3.dev.austin.int.futo.cloud`)
|
Virtual-hosted bucket addressing (e.g., `https://my-bucket.s3.staging.austin.int.futo.cloud`)
|
||||||
requires two DNS records:
|
requires two DNS records:
|
||||||
|
|
||||||
```
|
```
|
||||||
s3.dev.austin.int.futo.cloud. A 10.10.10.90
|
s3.staging.austin.int.futo.cloud. A 10.10.10.90
|
||||||
s3.dev.austin.int.futo.cloud. A 10.10.10.91
|
s3.staging.austin.int.futo.cloud. A 10.10.10.91
|
||||||
s3.dev.austin.int.futo.cloud. A 10.10.10.92
|
s3.staging.austin.int.futo.cloud. A 10.10.10.92
|
||||||
*.s3.dev.austin.int.futo.cloud. A 10.10.10.90
|
*.s3.staging.austin.int.futo.cloud. A 10.10.10.90
|
||||||
*.s3.dev.austin.int.futo.cloud. A 10.10.10.91
|
*.s3.staging.austin.int.futo.cloud. A 10.10.10.91
|
||||||
*.s3.dev.austin.int.futo.cloud. A 10.10.10.92
|
*.s3.staging.austin.int.futo.cloud. A 10.10.10.92
|
||||||
```
|
```
|
||||||
|
|
||||||
Round-robin A records across all three nodes.
|
Round-robin A records across all three nodes.
|
||||||
|
|||||||
@@ -22,13 +22,13 @@ it inline, never via `export`:**
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Correct — inline prefix, applies to one mise/script invocation
|
# Correct — inline prefix, applies to one mise/script invocation
|
||||||
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini mise run preflight
|
CEPH_ENV=inventories/staging-austin/sietch/inventory.ini mise run preflight
|
||||||
|
|
||||||
# WRONG — mise's [env] machinery silently strips shell-exported vars
|
# WRONG — mise's [env] machinery silently strips shell-exported vars
|
||||||
# when launching tasks; CEPH_ENV reaches an empty environment and the
|
# when launching tasks; CEPH_ENV reaches an empty environment and the
|
||||||
# wrapper exits with "CEPH_ENV must be set". Confusing because your shell
|
# wrapper exits with "CEPH_ENV must be set". Confusing because your shell
|
||||||
# clearly has it set.
|
# clearly has it set.
|
||||||
export CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini
|
export CEPH_ENV=inventories/staging-austin/sietch/inventory.ini
|
||||||
mise run preflight # fails
|
mise run preflight # fails
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -90,7 +90,7 @@ Common patterns:
|
|||||||
```bash
|
```bash
|
||||||
scripts/ansible-play.sh baseline.yml --check --diff
|
scripts/ansible-play.sh baseline.yml --check --diff
|
||||||
scripts/ansible-play.sh deploy-ceph.yml --tags rgw,monitoring
|
scripts/ansible-play.sh deploy-ceph.yml --tags rgw,monitoring
|
||||||
scripts/ansible-play.sh destroy-ceph.yml -e yes_destroy_ceph=true -e destroy_target_domain=dev.austin.int.futo.cloud
|
scripts/ansible-play.sh destroy-ceph.yml -e yes_destroy_ceph=true -e destroy_target_domain=staging.austin.int.futo.cloud
|
||||||
```
|
```
|
||||||
|
|
||||||
The destroy playbook requires both safety gates:
|
The destroy playbook requires both safety gates:
|
||||||
@@ -113,16 +113,16 @@ The `mise run destroy` task (in `.mise.toml`) builds these arguments automatical
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Standard deploy
|
# Standard deploy
|
||||||
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini \
|
CEPH_ENV=inventories/staging-austin/sietch/inventory.ini \
|
||||||
scripts/ansible-play.sh deploy-ceph.yml
|
scripts/ansible-play.sh deploy-ceph.yml
|
||||||
|
|
||||||
# Dry-run a role via tags
|
# Dry-run a role via tags
|
||||||
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini \
|
CEPH_ENV=inventories/staging-austin/sietch/inventory.ini \
|
||||||
scripts/ansible-play.sh site.yml --check --diff --tags baseline
|
scripts/ansible-play.sh site.yml --check --diff --tags baseline
|
||||||
|
|
||||||
# CI / headless (SA token from env)
|
# CI / headless (SA token from env)
|
||||||
OP_SERVICE_ACCOUNT_TOKEN="$(...)" \
|
OP_SERVICE_ACCOUNT_TOKEN="$(...)" \
|
||||||
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini \
|
CEPH_ENV=inventories/staging-austin/sietch/inventory.ini \
|
||||||
scripts/ansible-play.sh status.yml
|
scripts/ansible-play.sh status.yml
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -263,7 +263,7 @@ Warnings (non-blocking) are reported in the summary but don't affect exit.
|
|||||||
mise run preflight
|
mise run preflight
|
||||||
|
|
||||||
# Direct, against sietch
|
# Direct, against sietch
|
||||||
CEPH_ENV=inventories/sietch-ceph.staging.austin.int/inventory.ini \
|
CEPH_ENV=inventories/staging-austin/sietch/inventory.ini \
|
||||||
scripts/preflight.sh
|
scripts/preflight.sh
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
@@ -9,7 +9,7 @@ For how secrets fit into the broader architecture, see
|
|||||||
```mermaid
|
```mermaid
|
||||||
flowchart TB
|
flowchart TB
|
||||||
ONEP[("1Password<br/>yucca_tf · yucca_tf_dev · ...<br/><i>source of truth</i>")]
|
ONEP[("1Password<br/>yucca_tf · yucca_tf_dev · ...<br/><i>source of truth</i>")]
|
||||||
TF[Terraform / Tofu<br/>tf/deployment/staging/ceph/]
|
TF[Terraform / Tofu<br/>tf/deployment/staging/austin/ceph/]
|
||||||
REPO[/"inventories/<cluster>/<br/>inventory.ini (TF-gen, gitignored)<br/>secrets.yml.tpl (TF-gen, gitignored)"/]
|
REPO[/"inventories/<cluster>/<br/>inventory.ini (TF-gen, gitignored)<br/>secrets.yml.tpl (TF-gen, gitignored)"/]
|
||||||
WRAP[scripts/ansible-play.sh<br/><i>mktemp + op inject → exec ansible-playbook --extra-vars @tmp</i>]
|
WRAP[scripts/ansible-play.sh<br/><i>mktemp + op inject → exec ansible-playbook --extra-vars @tmp</i>]
|
||||||
ANS[ansible-playbook]
|
ANS[ansible-playbook]
|
||||||
@@ -31,7 +31,7 @@ flowchart TB
|
|||||||
|
|
||||||
Future environments land as siblings: `yucca_tf_staging(_manual)`,
|
Future environments land as siblings: `yucca_tf_staging(_manual)`,
|
||||||
`yucca_tf_prod_manual`. The vault a given cluster reads from is declared
|
`yucca_tf_prod_manual`. The vault a given cluster reads from is declared
|
||||||
per-cluster in `tf/deployment/<env>/ceph/clusters.auto.tfvars` (field
|
per-cluster in `tf/deployment/<partition>/<region>/ceph/clusters.auto.tfvars` (field
|
||||||
`vault`). TF derives item paths from that field at render time; changing
|
`vault`). TF derives item paths from that field at render time; changing
|
||||||
it + `tofu apply` re-renders `secrets.yml.tpl` with the new vault path.
|
it + `tofu apply` re-renders `secrets.yml.tpl` with the new vault path.
|
||||||
|
|
||||||
|
|||||||
@@ -40,8 +40,8 @@ lsblk --output NAME,TYPE,MOUNTPOINT | grep crypt
|
|||||||
|---|---|
|
|---|---|
|
||||||
| Protocol | HTTPS (TLS 1.2+) on port 443 |
|
| Protocol | HTTPS (TLS 1.2+) on port 443 |
|
||||||
| Certificate | Self-signed RSA 4096-bit, 10-year validity |
|
| Certificate | Self-signed RSA 4096-bit, 10-year validity |
|
||||||
| CN | `s3.dev.austin.int.futo.cloud` |
|
| CN | `s3.staging.austin.int.futo.cloud` |
|
||||||
| SANs | `s3.dev.austin.int.futo.cloud`, `*.s3.dev.austin.int.futo.cloud`, per-node FQDNs, per-node bond IPs |
|
| SANs | `s3.staging.austin.int.futo.cloud`, `*.s3.staging.austin.int.futo.cloud`, per-node FQDNs, per-node bond IPs |
|
||||||
| Issuer | Self-signed (O=FUTO, L=Austin, ST=Texas, C=US) |
|
| Issuer | Self-signed (O=FUTO, L=Austin, ST=Texas, C=US) |
|
||||||
| Cert location | `/etc/ceph/rgw-ssl.crt` + `/etc/ceph/rgw-ssl.key` on bootstrap node |
|
| Cert location | `/etc/ceph/rgw-ssl.crt` + `/etc/ceph/rgw-ssl.key` on bootstrap node |
|
||||||
| Distribution | cephadm distributes combined PEM to all RGW daemon containers |
|
| Distribution | cephadm distributes combined PEM to all RGW daemon containers |
|
||||||
|
|||||||
@@ -510,7 +510,7 @@ ceph_rgw_dns_name: s3.{{ cluster_domain }}
|
|||||||
```
|
```
|
||||||
|
|
||||||
This derives the DNS name from `cluster_domain` (e.g.
|
This derives the DNS name from `cluster_domain` (e.g.
|
||||||
`s3.dev.austin.int.futo.cloud`). Sietch defines this explicitly. New
|
`s3.staging.austin.int.futo.cloud`). Sietch defines this explicitly. New
|
||||||
clusters should include it from the start — see
|
clusters should include it from the start — see
|
||||||
[docs/adding-a-cluster.md](adding-a-cluster.md) group_vars template.
|
[docs/adding-a-cluster.md](adding-a-cluster.md) group_vars template.
|
||||||
|
|
||||||
|
|||||||
@@ -7,7 +7,7 @@
|
|||||||
gather_facts: false
|
gather_facts: false
|
||||||
|
|
||||||
vars:
|
vars:
|
||||||
cluster_domain: dev.austin.int.futo.cloud
|
cluster_domain: staging.austin.int.futo.cloud
|
||||||
public_network: 10.10.10.0/24
|
public_network: 10.10.10.0/24
|
||||||
cluster_network: 10.10.10.0/24
|
cluster_network: 10.10.10.0/24
|
||||||
ceph_rgw_realm: sietch
|
ceph_rgw_realm: sietch
|
||||||
@@ -15,7 +15,7 @@
|
|||||||
ceph_rgw_port: 443
|
ceph_rgw_port: 443
|
||||||
ceph_rgw_ssl: true
|
ceph_rgw_ssl: true
|
||||||
rgw_ssl_cert_combined_pem: "MOCK_CERT_CONTENT"
|
rgw_ssl_cert_combined_pem: "MOCK_CERT_CONTENT"
|
||||||
ceph_rgw_dns_name: s3.dev.austin.int.futo.cloud
|
ceph_rgw_dns_name: s3.staging.austin.int.futo.cloud
|
||||||
|
|
||||||
tasks:
|
tasks:
|
||||||
- name: Render hosts.j2
|
- name: Render hosts.j2
|
||||||
|
|||||||
@@ -18,7 +18,7 @@ set -euo pipefail
|
|||||||
|
|
||||||
if [ ! -f "$CEPH_ENV" ]; then
|
if [ ! -f "$CEPH_ENV" ]; then
|
||||||
echo "ansible-play.sh: inventory not found: $CEPH_ENV" >&2
|
echo "ansible-play.sh: inventory not found: $CEPH_ENV" >&2
|
||||||
echo " Hint: render it — 'terragrunt apply' in tf/deployment/<env>/ceph/, then scripts/render-inventories.sh <env>." >&2
|
echo " Hint: render it — 'terragrunt apply' in tf/deployment/<partition>/<region>/ceph/, then scripts/render-inventories.sh <partition> <region>." >&2
|
||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
|
|
||||||
@@ -27,7 +27,7 @@ TEMPLATE="$CEPH_ENV_DIR/secrets.yml.tpl"
|
|||||||
|
|
||||||
if [ ! -f "$TEMPLATE" ]; then
|
if [ ! -f "$TEMPLATE" ]; then
|
||||||
echo "ansible-play.sh: secrets template not found: $TEMPLATE" >&2
|
echo "ansible-play.sh: secrets template not found: $TEMPLATE" >&2
|
||||||
echo " Hint: render it — 'terragrunt apply' in tf/deployment/<env>/ceph/, then scripts/render-inventories.sh <env>." >&2
|
echo " Hint: render it — 'terragrunt apply' in tf/deployment/<partition>/<region>/ceph/, then scripts/render-inventories.sh <partition> <region>." >&2
|
||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
|
|
||||||
|
|||||||
@@ -5,7 +5,7 @@
|
|||||||
#
|
#
|
||||||
# Target inventory + secrets template are located via CEPH_ENV (path to
|
# Target inventory + secrets template are located via CEPH_ENV (path to
|
||||||
# the TF-rendered inventory file — TF is the authoritative source of
|
# the TF-rendered inventory file — TF is the authoritative source of
|
||||||
# cluster identity, declared in tf/deployment/<env>/ceph/clusters.auto.tfvars).
|
# cluster identity, declared in tf/deployment/<partition>/<region>/ceph/clusters.auto.tfvars).
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
|
|
||||||
: "${CEPH_ENV:?CEPH_ENV must be set to the target cluster inventory.ini}"
|
: "${CEPH_ENV:?CEPH_ENV must be set to the target cluster inventory.ini}"
|
||||||
|
|||||||
@@ -14,15 +14,21 @@
|
|||||||
# reflects the current cluster spec). This script is read-only against state.
|
# reflects the current cluster spec). This script is read-only against state.
|
||||||
#
|
#
|
||||||
# Usage (from anywhere):
|
# Usage (from anywhere):
|
||||||
# ansible/ceph/scripts/render-inventories.sh [env] # env defaults to dev
|
# ansible/ceph/scripts/render-inventories.sh [partition] [region]
|
||||||
|
# # partition defaults to staging, region defaults to austin (sietch's home)
|
||||||
|
#
|
||||||
|
# The TF `render` output carries each cluster's own `dirname` (region-scoped,
|
||||||
|
# friendly-cluster leaf — e.g. `staging-austin/sietch`), so this wrapper does
|
||||||
|
# not derive the inventory layout; it only locates the partition/region stack.
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
|
|
||||||
ENVIRONMENT="${1:-dev}"
|
PARTITION="${1:-staging}"
|
||||||
|
REGION="${2:-austin}"
|
||||||
|
|
||||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||||
ANSIBLE_CEPH="$(cd "$SCRIPT_DIR/.." && pwd)" # ansible/ceph
|
ANSIBLE_CEPH="$(cd "$SCRIPT_DIR/.." && pwd)" # ansible/ceph
|
||||||
REPO_ROOT="$(cd "$ANSIBLE_CEPH/../.." && pwd)" # repo root of THIS checkout
|
REPO_ROOT="$(cd "$ANSIBLE_CEPH/../.." && pwd)" # repo root of THIS checkout
|
||||||
STACK_DIR="$REPO_ROOT/tf/deployment/${ENVIRONMENT}/ceph"
|
STACK_DIR="$REPO_ROOT/tf/deployment/${PARTITION}/${REGION}/ceph"
|
||||||
INV_ROOT="$ANSIBLE_CEPH/inventories"
|
INV_ROOT="$ANSIBLE_CEPH/inventories"
|
||||||
|
|
||||||
[ -d "$STACK_DIR" ] || {
|
[ -d "$STACK_DIR" ] || {
|
||||||
@@ -54,4 +60,4 @@ for cluster, spec in data.items():
|
|||||||
print(f"wrote {p}")
|
print(f"wrote {p}")
|
||||||
PY
|
PY
|
||||||
|
|
||||||
echo "render-inventories: done (${ENVIRONMENT})."
|
echo "render-inventories: done (${PARTITION}/${REGION})."
|
||||||
|
|||||||
@@ -45,9 +45,9 @@ truth and writes these **gitignored** files:
|
|||||||
|
|
||||||
| file | generated from |
|
| file | generated from |
|
||||||
|---|---|
|
|---|---|
|
||||||
| `inventories/<site>/hosts.yml` | `mgmt-hosts.yaml` (host names + public IPs) |
|
| `inventories/<region>/hosts.yml` | `mgmt-hosts.yaml` (host names + public IPs) |
|
||||||
| `inventories/<site>/host_vars/*.yml` | `mgmt-hosts.yaml` (NIC) + `fabric-addressing` (VLAN ids/addresses, subnet route) |
|
| `inventories/<region>/host_vars/*.yml` | `mgmt-hosts.yaml` (NIC) + `fabric-addressing` (VLAN ids/addresses, subnet route) |
|
||||||
| `inventories/<site>/group_vars/all/users.generated.yml` | `tf/shared/modules/identity` (`server`-mapped users) |
|
| `inventories/<region>/group_vars/all/users.generated.yml` | `tf/shared/modules/identity` (`server`-mapped users) |
|
||||||
|
|
||||||
Only `group_vars/all/main.yml` (static config) and `roles/**` are committed. To
|
Only `group_vars/all/main.yml` (static config) and `roles/**` are committed. To
|
||||||
change hosts, addresses, or users, edit the Terraform sources — never the
|
change hosts, addresses, or users, edit the Terraform sources — never the
|
||||||
@@ -82,7 +82,7 @@ shred -u /tmp/htz-fsn1-prov-key
|
|||||||
optional. The inventory connects as `ansible_user: root`.
|
optional. The inventory connects as `ansible_user: root`.
|
||||||
|
|
||||||
The whole render-key + run flow above is wrapped by `mise run mgmt:ansible`
|
The whole render-key + run flow above is wrapped by `mise run mgmt:ansible`
|
||||||
(`SITE` selects the inventory; defaults to `htz-fsn1`), which CI also runs on
|
(`REGION` selects the inventory; defaults to `htz-fsn1`), which CI also runs on
|
||||||
every **prod** apply — the `Ansible converge (mgmt hosts)` step of
|
every **prod** apply — the `Ansible converge (mgmt hosts)` step of
|
||||||
`.github/workflows/infra.yml`, right after the Terraform apply. It's idempotent
|
`.github/workflows/infra.yml`, right after the Terraform apply. It's idempotent
|
||||||
and reaches the hosts over their public IP, so it requires them to already be
|
and reaches the hosts over their public IP, so it requires them to already be
|
||||||
@@ -143,5 +143,5 @@ Nothing secret is committed. SSH public keys are public data and are TF-generate
|
|||||||
into `group_vars/all/users.generated.yml`. Runtime secrets are passed via `op read`:
|
into `group_vars/all/users.generated.yml`. Runtime secrets are passed via `op read`:
|
||||||
|
|
||||||
- Provisioning private key: `op://yucca_tf_prod/HTZ_FSN1_PROVISIONING_SSH_PRIVATE_KEY/password`
|
- Provisioning private key: `op://yucca_tf_prod/HTZ_FSN1_PROVISIONING_SSH_PRIVATE_KEY/password`
|
||||||
- NetBird mgmt setup key: `op://yucca_tf_prod/NETBIRD_YUCCA_PROD_<SITE>_MGMT_SETUP_KEY/password`
|
- NetBird mgmt setup key: `op://yucca_tf_prod/NETBIRD_YUCCA_PROD_<REGION>_MGMT_SETUP_KEY/password`
|
||||||
(the site's reusable `mgmt` key, `auto_groups=["mgmt"]`; minted by the prod netbird stack)
|
(the site's reusable `mgmt` key, `auto_groups=["mgmt"]`; minted by the prod netbird stack)
|
||||||
|
|||||||
@@ -10,5 +10,5 @@ timezone: UTC
|
|||||||
mgmt_domain: fsn.htz.futo.cloud
|
mgmt_domain: fsn.htz.futo.cloud
|
||||||
|
|
||||||
# NetBird "mgmt" setup key — NOT hardcoded; passed at run time by
|
# NetBird "mgmt" setup key — NOT hardcoded; passed at run time by
|
||||||
# `mise run mgmt:ansible` from op://yucca_tf_prod/NETBIRD_YUCCA_PROD_<SITE>_MGMT_SETUP_KEY.
|
# `mise run mgmt:ansible` from op://yucca_tf_prod/NETBIRD_YUCCA_PROD_<REGION>_MGMT_SETUP_KEY.
|
||||||
mgmt_netbird_setup_key: ""
|
mgmt_netbird_setup_key: ""
|
||||||
|
|||||||
@@ -5,5 +5,5 @@ mgmt_netbird_management_url: "https://api.netbird.io"
|
|||||||
# Setup key — NEVER hardcoded. Passed at run time (op:// ref, see README). The
|
# Setup key — NEVER hardcoded. Passed at run time (op:// ref, see README). The
|
||||||
# site's reusable "mgmt" key (auto_groups=["mgmt"]) so the node joins the mgmt
|
# site's reusable "mgmt" key (auto_groups=["mgmt"]) so the node joins the mgmt
|
||||||
# group and becomes a route peer for the site subnets (10.40.5.0/24, api, cluster
|
# group and becomes a route peer for the site subnets (10.40.5.0/24, api, cluster
|
||||||
# nets — the routed network is defined in TF: tf/deployment/prod/<site>/netbird).
|
# nets — the routed network is defined in TF: tf/deployment/prod/<region>/netbird).
|
||||||
mgmt_netbird_setup_key: ""
|
mgmt_netbird_setup_key: ""
|
||||||
|
|||||||
@@ -23,7 +23,7 @@
|
|||||||
- mgmt_netbird_setup_key | length > 0
|
- mgmt_netbird_setup_key | length > 0
|
||||||
fail_msg: >-
|
fail_msg: >-
|
||||||
mgmt_netbird_setup_key is empty. Pass it at run time (mgmt:ansible does this
|
mgmt_netbird_setup_key is empty. Pass it at run time (mgmt:ansible does this
|
||||||
from op://yucca_tf_prod/NETBIRD_YUCCA_PROD_<SITE>_MGMT_SETUP_KEY).
|
from op://yucca_tf_prod/NETBIRD_YUCCA_PROD_<REGION>_MGMT_SETUP_KEY).
|
||||||
|
|
||||||
# mgmt nodes route the site subnets to the overlay.
|
# mgmt nodes route the site subnets to the overlay.
|
||||||
- name: Enable IP forwarding (NetBird route peer)
|
- name: Enable IP forwarding (NetBird route peer)
|
||||||
|
|||||||
@@ -8,7 +8,7 @@
|
|||||||
#
|
#
|
||||||
# Canonical entrypoint is `mise run mgmt:ansible` (renders the inventory from
|
# Canonical entrypoint is `mise run mgmt:ansible` (renders the inventory from
|
||||||
# Terraform, then runs this). It passes the provisioning key + the NetBird
|
# Terraform, then runs this). It passes the provisioning key + the NetBird
|
||||||
# "mgmt" setup key (op://yucca_tf_prod/NETBIRD_YUCCA_PROD_<SITE>_MGMT_SETUP_KEY).
|
# "mgmt" setup key (op://yucca_tf_prod/NETBIRD_YUCCA_PROD_<REGION>_MGMT_SETUP_KEY).
|
||||||
#
|
#
|
||||||
# NOTE: the networkd role (25G VLAN sub-interfaces) is gated on
|
# NOTE: the networkd role (25G VLAN sub-interfaces) is gated on
|
||||||
# mgmt_networkd_enabled (default false) because the 25G fabric link is
|
# mgmt_networkd_enabled (default false) because the 25G fabric link is
|
||||||
|
|||||||
@@ -8,8 +8,8 @@ PATH = "{{config_root}}/.venv/bin:{{env.PATH}}"
|
|||||||
# Note: TALOS_ENV is intentionally NOT declared here. mise's [env] block
|
# Note: TALOS_ENV is intentionally NOT declared here. mise's [env] block
|
||||||
# overrides shell-exported values, which silently sends operators to the
|
# overrides shell-exported values, which silently sends operators to the
|
||||||
# wrong cluster. Operators must export TALOS_ENV once per shell session:
|
# wrong cluster. Operators must export TALOS_ENV once per shell session:
|
||||||
# export TALOS_ENV=inventories/sietch-talos.dev.austin.int/inventory.ini
|
# export TALOS_ENV=inventories/staging-austin/inventory.ini
|
||||||
# Inventory file is TF-rendered — `TF_STACK_DIR=tf/deployment/dev/talos mise run tf:apply`
|
# Inventory file is TF-rendered — `TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply`
|
||||||
# (from yucca root) if missing.
|
# (from yucca root) if missing.
|
||||||
|
|
||||||
[tasks.setup]
|
[tasks.setup]
|
||||||
@@ -35,7 +35,7 @@ yamllint --version
|
|||||||
|
|
||||||
echo ""
|
echo ""
|
||||||
echo "Environment ready. Run 'mise trust' if prompted."
|
echo "Environment ready. Run 'mise trust' if prompted."
|
||||||
echo "Next: 'TF_STACK_DIR=tf/deployment/dev/talos mise run tf:apply' from yucca root to render inventory."
|
echo "Next: 'TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply' from yucca root to render inventory."
|
||||||
"""
|
"""
|
||||||
|
|
||||||
[tasks.lint]
|
[tasks.lint]
|
||||||
@@ -64,7 +64,7 @@ run = """
|
|||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
# Syntax-check only parses YAML — any valid inventory works. Default to
|
# Syntax-check only parses YAML — any valid inventory works. Default to
|
||||||
# sietch-talos when TALOS_ENV isn't inline-prefixed; the parse is identical.
|
# sietch-talos when TALOS_ENV isn't inline-prefixed; the parse is identical.
|
||||||
TALOS_ENV="${TALOS_ENV:-inventories/sietch-talos.dev.austin.int/inventory.ini}"
|
TALOS_ENV="${TALOS_ENV:-inventories/staging-austin/inventory.ini}"
|
||||||
# Fall back to inventory.example.ini before tf:apply has rendered the runtime one.
|
# Fall back to inventory.example.ini before tf:apply has rendered the runtime one.
|
||||||
if [ ! -f "$TALOS_ENV" ]; then
|
if [ ! -f "$TALOS_ENV" ]; then
|
||||||
EXAMPLE="$(dirname "$TALOS_ENV")/inventory.example.ini"
|
EXAMPLE="$(dirname "$TALOS_ENV")/inventory.example.ini"
|
||||||
|
|||||||
@@ -8,7 +8,7 @@ hardware.
|
|||||||
This subtree is the **Ansible substrate**: it provisions the VLAN 50/51
|
This subtree is the **Ansible substrate**: it provisions the VLAN 50/51
|
||||||
bridges, the libvirt/KVM stack, and stages the Talos VMs — stopping when
|
bridges, the libvirt/KVM stack, and stages the Talos VMs — stopping when
|
||||||
the VMs are running and ready for `talosctl`. The Terraform half in
|
the VMs are running and ready for `talosctl`. The Terraform half in
|
||||||
[tf/deployment/dev/talos/](../../tf/deployment/dev/talos/) renders the
|
[tf/deployment/staging/austin/talos/](../../tf/deployment/staging/austin/talos/) renders the
|
||||||
Ansible inventory and drives cluster bring-up (machine config, bootstrap,
|
Ansible inventory and drives cluster bring-up (machine config, bootstrap,
|
||||||
kubeconfig).
|
kubeconfig).
|
||||||
|
|
||||||
@@ -31,7 +31,7 @@ cluster secrets out of Ansible's fact cache keeps the blast radius small.
|
|||||||
|
|
||||||
- `talosctl gen config / apply-config / bootstrap` — the Terraform
|
- `talosctl gen config / apply-config / bootstrap` — the Terraform
|
||||||
`siderolabs/talos` provider owns these
|
`siderolabs/talos` provider owns these
|
||||||
([tf/deployment/dev/talos/](../../tf/deployment/dev/talos/));
|
([tf/deployment/staging/austin/talos/](../../tf/deployment/staging/austin/talos/));
|
||||||
`docs/operator-handoff.md` keeps the manual sequence for recovery.
|
`docs/operator-handoff.md` keeps the manual sequence for recovery.
|
||||||
- Persistent storage for VMs (RBD-backed boot disks, ceph-csi). Boot
|
- Persistent storage for VMs (RBD-backed boot disks, ceph-csi). Boot
|
||||||
disks are local qcow2 today; RBD lands in a follow-up.
|
disks are local qcow2 today; RBD lands in a follow-up.
|
||||||
@@ -54,19 +54,19 @@ validation tool, selected explicitly with `-e profile=smoke`.
|
|||||||
|
|
||||||
## Operator workflow
|
## Operator workflow
|
||||||
|
|
||||||
Inventory is **TF-rendered** from `tf/deployment/dev/talos/`. Run TF first
|
Inventory is **TF-rendered** from `tf/deployment/staging/austin/talos/`. Run TF first
|
||||||
(from the repo root):
|
(from the repo root):
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
TF_STACK_DIR=tf/deployment/dev/talos mise run tf:init # first time only
|
TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:init # first time only
|
||||||
TF_STACK_DIR=tf/deployment/dev/talos mise run tf:apply # renders inventory + host_vars
|
TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply # renders inventory + host_vars
|
||||||
```
|
```
|
||||||
|
|
||||||
Then point at the rendered inventory and use the talos-subtree tasks:
|
Then point at the rendered inventory and use the talos-subtree tasks:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
cd ansible/talos/
|
cd ansible/talos/
|
||||||
export TALOS_ENV=inventories/sietch-talos.dev.austin.int/inventory.ini
|
export TALOS_ENV=inventories/staging-austin/inventory.ini
|
||||||
|
|
||||||
mise run setup # first time only: python venv + ansible collections
|
mise run setup # first time only: python venv + ansible collections
|
||||||
mise run lint # yamllint + ansible-lint + shellcheck
|
mise run lint # yamllint + ansible-lint + shellcheck
|
||||||
|
|||||||
@@ -17,7 +17,7 @@ Provisioning split across the monorepo:
|
|||||||
| Subtree | Owns |
|
| Subtree | Owns |
|
||||||
|---|---|
|
|---|---|
|
||||||
| `ansible/talos/` (this) | Hypervisor substrate: bridges, libvirt, image, VM definitions. Stops at "VMs ready"; Terraform owns bootstrap (the subtree README explains why the boundary sits here). |
|
| `ansible/talos/` (this) | Hypervisor substrate: bridges, libvirt, image, VM definitions. Stops at "VMs ready"; Terraform owns bootstrap (the subtree README explains why the boundary sits here). |
|
||||||
| `tf/shared/modules/talos-cluster/modules/inventory-renderer/` | Renders the Ansible inventory + (future) `secrets.yml.tpl` from `tf/deployment/dev/talos/clusters.auto.tfvars`. Parity with the ceph-cluster module. |
|
| `tf/shared/modules/talos-cluster/modules/inventory-renderer/` | Renders the Ansible inventory + (future) `secrets.yml.tpl` from `tf/deployment/staging/austin/talos/clusters.auto.tfvars`. Parity with the ceph-cluster module. |
|
||||||
| `tf/shared/modules/talos-cluster/modules/talos-bootstrap/` | `siderolabs/talos` provider — machine_secrets, configuration_apply per node, bootstrap, kubeconfig. Drives the talosctl sequence so operators don't run it by hand. |
|
| `tf/shared/modules/talos-cluster/modules/talos-bootstrap/` | `siderolabs/talos` provider — machine_secrets, configuration_apply per node, bootstrap, kubeconfig. Drives the talosctl sequence so operators don't run it by hand. |
|
||||||
|
|
||||||
## Physical layout
|
## Physical layout
|
||||||
@@ -77,7 +77,7 @@ tailnet/VPN subnet route, or being directly on the VLAN) before
|
|||||||
## Naming scheme
|
## Naming scheme
|
||||||
|
|
||||||
Monorepo FQDN: `<cluster>-<role>-<name>.<domain>` —
|
Monorepo FQDN: `<cluster>-<role>-<name>.<domain>` —
|
||||||
e.g. `sietch-talos-cp1.dev.austin.int.futo.cloud`. The
|
e.g. `sietch-talos-cp1.staging.austin.int.futo.cloud`. The
|
||||||
`<cluster>-<role>` prefix (`sietch-talos`) is the
|
`<cluster>-<role>` prefix (`sietch-talos`) is the
|
||||||
`talos_domain_prefix` group_vars value; `<name>` is the entry in
|
`talos_domain_prefix` group_vars value; `<name>` is the entry in
|
||||||
each host's `host_vars/*.yml` talos_vms list.
|
each host's `host_vars/*.yml` talos_vms list.
|
||||||
@@ -167,8 +167,8 @@ pulls). The TF bootstrap sets
|
|||||||
`machine.network.nameservers: [1.1.1.1, 8.8.8.8]` on Talos VMs for
|
`machine.network.nameservers: [1.1.1.1, 8.8.8.8]` on Talos VMs for
|
||||||
upstream image-registry resolution. Kubeconfig uses
|
upstream image-registry resolution. Kubeconfig uses
|
||||||
the CP VIP IP directly (`https://10.50.0.10:6443`). DNS records
|
the CP VIP IP directly (`https://10.50.0.10:6443`). DNS records
|
||||||
under `*.compute.dev.austin.int.futo.cloud` and
|
under `*.compute.staging.austin.int.futo.cloud` and
|
||||||
`*.services.dev.austin.int.futo.cloud` land with the DNS-layer
|
`*.services.staging.austin.int.futo.cloud` land with the DNS-layer
|
||||||
follow-up (LB → ingress → external-dns).
|
follow-up (LB → ingress → external-dns).
|
||||||
|
|
||||||
## Out of scope
|
## Out of scope
|
||||||
|
|||||||
@@ -22,12 +22,12 @@ Sietch hypervisors. The TF stack owns this flow;
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
cd <yucca>
|
cd <yucca>
|
||||||
TF_STACK_DIR=tf/deployment/dev/talos mise run tf:init # first time
|
TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:init # first time
|
||||||
TF_STACK_DIR=tf/deployment/dev/talos mise run tf:apply
|
TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply
|
||||||
```
|
```
|
||||||
|
|
||||||
What lands:
|
What lands:
|
||||||
- `inventories/sietch-talos.dev.austin.int/inventory.ini` and per-host
|
- `inventories/staging-austin/inventory.ini` and per-host
|
||||||
`host_vars/*.yml` (the `talos_vms` lists), both rendered from
|
`host_vars/*.yml` (the `talos_vms` lists), both rendered from
|
||||||
`clusters.auto.tfvars` — `nodes[]` is the single source of truth for
|
`clusters.auto.tfvars` — `nodes[]` is the single source of truth for
|
||||||
topology, so never hand-edit `host_vars`.
|
topology, so never hand-edit `host_vars`.
|
||||||
@@ -51,7 +51,7 @@ re-run the apply in step 3.
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
cd <yucca>/ansible/talos
|
cd <yucca>/ansible/talos
|
||||||
export TALOS_ENV=inventories/sietch-talos.dev.austin.int/inventory.ini
|
export TALOS_ENV=inventories/staging-austin/inventory.ini
|
||||||
|
|
||||||
# Optional first time: bootstrap python venv + ansible collections
|
# Optional first time: bootstrap python venv + ansible collections
|
||||||
mise run setup
|
mise run setup
|
||||||
@@ -106,7 +106,7 @@ the desktop app (or `OP_SERVICE_ACCOUNT_TOKEN` if set).
|
|||||||
ip route get 10.50.0.10
|
ip route get 10.50.0.10
|
||||||
|
|
||||||
cd <yucca>
|
cd <yucca>
|
||||||
TF_STACK_DIR=tf/deployment/dev/talos mise run tf:apply
|
TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply
|
||||||
# Watch for:
|
# Watch for:
|
||||||
# - module.cluster["sietch"].module.talos_bootstrap.talos_machine_configuration_apply.controlplane["cp1"]: Creating...
|
# - module.cluster["sietch"].module.talos_bootstrap.talos_machine_configuration_apply.controlplane["cp1"]: Creating...
|
||||||
# - ... .talos_machine_bootstrap.this: Creating...
|
# - ... .talos_machine_bootstrap.this: Creating...
|
||||||
@@ -124,10 +124,10 @@ mkdir -p ~/.kube ~/.talos
|
|||||||
# Use `output -json` alone — `-raw -json` are mutually exclusive and
|
# Use `output -json` alone — `-raw -json` are mutually exclusive and
|
||||||
# yield an empty file. op-run.sh resolves the S3 state creds.
|
# yield an empty file. op-run.sh resolves the S3 state creds.
|
||||||
tf/op-run.sh terragrunt \
|
tf/op-run.sh terragrunt \
|
||||||
--working-dir tf/deployment/dev/talos output -json kubeconfigs \
|
--working-dir tf/deployment/staging/austin/talos output -json kubeconfigs \
|
||||||
| jq -r .sietch > ~/.kube/sietch-talos.config
|
| jq -r .sietch > ~/.kube/sietch-talos.config
|
||||||
tf/op-run.sh terragrunt \
|
tf/op-run.sh terragrunt \
|
||||||
--working-dir tf/deployment/dev/talos output -json talosconfigs \
|
--working-dir tf/deployment/staging/austin/talos output -json talosconfigs \
|
||||||
| jq -r .sietch > ~/.talos/sietch-talos.config
|
| jq -r .sietch > ~/.talos/sietch-talos.config
|
||||||
|
|
||||||
export KUBECONFIG=~/.kube/sietch-talos.config
|
export KUBECONFIG=~/.kube/sietch-talos.config
|
||||||
@@ -171,7 +171,7 @@ talosctl reset --graceful --reboot \
|
|||||||
# deletes the rendered inventory.ini this play needs. (If you ran them in
|
# deletes the rendered inventory.ini this play needs. (If you ran them in
|
||||||
# the wrong order: cp inventory.example.ini inventory.ini and re-run.)
|
# the wrong order: cp inventory.example.ini inventory.ini and re-run.)
|
||||||
cd <yucca>/ansible/talos
|
cd <yucca>/ansible/talos
|
||||||
export TALOS_ENV=inventories/sietch-talos.dev.austin.int/inventory.ini
|
export TALOS_ENV=inventories/staging-austin/inventory.ini
|
||||||
mise run destroy-vms
|
mise run destroy-vms
|
||||||
|
|
||||||
# Destroy TF cluster state, LAST (kubeconfig/talosconfig outputs blanked,
|
# Destroy TF cluster state, LAST (kubeconfig/talosconfig outputs blanked,
|
||||||
@@ -180,7 +180,7 @@ mise run destroy-vms
|
|||||||
# wiped/destroyed nodes and aborts with "cluster health check failed".
|
# wiped/destroyed nodes and aborts with "cluster health check failed".
|
||||||
# (Destroying the talos resources is state-only — nothing is un-applied.)
|
# (Destroying the talos resources is state-only — nothing is un-applied.)
|
||||||
cd <yucca>
|
cd <yucca>
|
||||||
TF_STACK_DIR=tf/deployment/dev/talos mise run tf:destroy -- -refresh=false
|
TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:destroy -- -refresh=false
|
||||||
```
|
```
|
||||||
|
|
||||||
To redeploy from clean: repeat steps 1 → 2 → 3 → 4. Static IPs are
|
To redeploy from clean: repeat steps 1 → 2 → 3 → 4. Static IPs are
|
||||||
@@ -196,7 +196,7 @@ production topology:
|
|||||||
# 1. Restore the production profile in TF input — if you ran smoke, the
|
# 1. Restore the production profile in TF input — if you ran smoke, the
|
||||||
# tfvars still says `profile = "smoke"` and the TF re-apply would
|
# tfvars still says `profile = "smoke"` and the TF re-apply would
|
||||||
# bootstrap nothing new:
|
# bootstrap nothing new:
|
||||||
# tf/deployment/dev/talos/clusters.auto.tfvars → profile = "full"
|
# tf/deployment/staging/austin/talos/clusters.auto.tfvars → profile = "full"
|
||||||
|
|
||||||
# 2. Provision the remaining VMs on lawson + samara (cp2/worker2,
|
# 2. Provision the remaining VMs on lawson + samara (cp2/worker2,
|
||||||
# cp3/worker3). Existing cp1/worker1 are a no-op.
|
# cp3/worker3). Existing cp1/worker1 are a no-op.
|
||||||
@@ -206,7 +206,7 @@ mise run provision # profile=full is the default
|
|||||||
# 3. Re-apply TF for the 4 new node bootstraps (cp2/cp3 join etcd,
|
# 3. Re-apply TF for the 4 new node bootstraps (cp2/cp3 join etcd,
|
||||||
# worker2/worker3 join the cluster)
|
# worker2/worker3 join the cluster)
|
||||||
cd ../..
|
cd ../..
|
||||||
TF_STACK_DIR=tf/deployment/dev/talos mise run tf:apply
|
TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply
|
||||||
```
|
```
|
||||||
|
|
||||||
NOTE: going from 1 CP (smoke) to 3 CP (full) grows the etcd quorum.
|
NOTE: going from 1 CP (smoke) to 3 CP (full) grows the etcd quorum.
|
||||||
|
|||||||
@@ -33,7 +33,7 @@ set `profile = "smoke"` in `clusters.auto.tfvars` before `tf:apply`:
|
|||||||
mise run provision -- -e profile=smoke
|
mise run provision -- -e profile=smoke
|
||||||
|
|
||||||
# TF bootstrap — set profile = "smoke" in clusters.auto.tfvars first
|
# TF bootstrap — set profile = "smoke" in clusters.auto.tfvars first
|
||||||
TF_STACK_DIR=tf/deployment/dev/talos mise run tf:apply
|
TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply
|
||||||
```
|
```
|
||||||
|
|
||||||
Everything else (static `ip=` addressing, factory image, direct kernel
|
Everything else (static `ip=` addressing, factory image, direct kernel
|
||||||
@@ -50,7 +50,7 @@ laurel, so the limit costs nothing.
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
cd ansible/talos
|
cd ansible/talos
|
||||||
export TALOS_ENV=inventories/sietch-talos.dev.austin.int/inventory.ini
|
export TALOS_ENV=inventories/staging-austin/inventory.ini
|
||||||
|
|
||||||
# A. preflight only — read-only, zero state change
|
# A. preflight only — read-only, zero state change
|
||||||
scripts/ansible-play.sh preflight.yml --limit sietch-ceph-laurel
|
scripts/ansible-play.sh preflight.yml --limit sietch-ceph-laurel
|
||||||
|
|||||||
@@ -1,126 +0,0 @@
|
|||||||
---
|
|
||||||
# Cluster-wide variables for the sietch-talos cluster.
|
|
||||||
#
|
|
||||||
# Most operators editing this file want one of:
|
|
||||||
# - talos_version — bump Talos release
|
|
||||||
# - talos_vm_mac_prefix — change MAC OUI / locally-administered prefix
|
|
||||||
# - vlan_*_* — tweak network details
|
|
||||||
|
|
||||||
# ─── Cluster identity ───────────────────────────────────────────────────
|
|
||||||
# Short cluster name. Also drives talos_domain_prefix below (and the
|
|
||||||
# TF-side 1P secret prefixes). A second cluster would get its own
|
|
||||||
# inventory dir with a different talos_cluster_name.
|
|
||||||
talos_cluster_name: "sietch"
|
|
||||||
|
|
||||||
# Libvirt domain prefix for Talos VMs on this cluster. talos_vms role
|
|
||||||
# reads this to compose `<prefix>-<vm-name>` (e.g. sietch-talos-cp1).
|
|
||||||
# Aligns with the monorepo FQDN scheme `<cluster>-<role>-<name>.<domain>`.
|
|
||||||
talos_domain_prefix: "{{ talos_cluster_name }}-talos"
|
|
||||||
|
|
||||||
# Cluster FQDN suffix (where the K8s API + node FQDNs land). The TF
|
|
||||||
# bootstrap uses this to seed kubeconfig endpoints and Talos cert SANs.
|
|
||||||
talos_domain: "dev.austin.int.futo.cloud"
|
|
||||||
|
|
||||||
# ─── Talos version + image (Image Factory) ──────────────────────────────
|
|
||||||
# Pinned for reproducibility. Bump deliberately; refresh all three
|
|
||||||
# checksums (image + kernel + initramfs) with the version.
|
|
||||||
talos_version: "1.13.3"
|
|
||||||
|
|
||||||
# Factory schematic — extensions baked into the boot assets:
|
|
||||||
# siderolabs/{qemu-guest-agent, util-linux-tools}. ID is sha256 of the
|
|
||||||
# schematic; regenerate at https://factory.talos.dev and refresh checksums
|
|
||||||
# if the set changes.
|
|
||||||
talos_schematic_id: "a7bcadbc1b6d03c0e687be3a5d9789ef7113362a6a1a038653dfd16283a92b6b"
|
|
||||||
talos_factory_base: "https://factory.talos.dev/image/{{ talos_schematic_id }}/v{{ talos_version }}"
|
|
||||||
|
|
||||||
# Disk image — qcow2 overlays back onto this pre-seeded raw (partitions
|
|
||||||
# present, no install step). Boots to maintenance mode until TF applies config.
|
|
||||||
talos_image_url: "{{ talos_factory_base }}/metal-amd64.raw.zst"
|
|
||||||
talos_image_sha256: "c7515f1076db75f6ec4287b4b5663ea0d490363a0d49b04e61fe96bfe71b2cc6"
|
|
||||||
|
|
||||||
# Kernel + initramfs — staged and direct-booted by libvirt (not on-disk
|
|
||||||
# GRUB) to inject a per-VM `ip=` for static first-boot addressing.
|
|
||||||
# Upgrades: re-stage assets + reboot (the disk kernel is bypassed).
|
|
||||||
talos_kernel_url: "{{ talos_factory_base }}/kernel-amd64"
|
|
||||||
talos_initramfs_url: "{{ talos_factory_base }}/initramfs-amd64.xz"
|
|
||||||
talos_kernel_sha256: "c36d4aac36d081af8016ff9874f29115bf1493d23e556b99d355316afa0e98ff"
|
|
||||||
talos_initramfs_sha256: "d613bd10a0b7742e17705e0643ec62e19150522a018fb27855535c75918eabc7"
|
|
||||||
|
|
||||||
talos_image_dir: "/var/lib/libvirt/images"
|
|
||||||
|
|
||||||
# Filenames keyed on version + short schematic hash so a schematic change at
|
|
||||||
# the same version yields new filenames and re-stages, instead of serving
|
|
||||||
# the stale cached image. default('stock') when no schematic.
|
|
||||||
talos_image_tag: "v{{ talos_version }}-{{ (talos_schematic_id | default('stock', true))[:12] }}"
|
|
||||||
talos_base_image_path: "{{ talos_image_dir }}/talos-{{ talos_image_tag }}.raw"
|
|
||||||
talos_kernel_path: "{{ talos_image_dir }}/talos-vmlinuz-{{ talos_image_tag }}"
|
|
||||||
talos_initramfs_path: "{{ talos_image_dir }}/talos-initramfs-{{ talos_image_tag }}.xz"
|
|
||||||
|
|
||||||
# ─── Static addressing (libvirt kernel `ip=`) ────────────────────────────
|
|
||||||
# Per-VM static_ip is in host_vars; gateway/mask/interface are VLAN-50
|
|
||||||
# constants. Gateway hardcoded (first usable in 10.50.0.0/16) to avoid an
|
|
||||||
# ansible.utils dependency; TF derives the same via cidrhost().
|
|
||||||
talos_vms_network_gateway: "10.50.0.1"
|
|
||||||
talos_vms_network_mask: "255.255.0.0"
|
|
||||||
talos_vms_network_interface: "enp1s0"
|
|
||||||
|
|
||||||
# ─── VM MAC prefix ──────────────────────────────────────────────────────
|
|
||||||
# SideroLabs has no registered OUI (verified against IEEE registry +
|
|
||||||
# maclookup.app; Talos prescribes no convention). Choosing 52:54:00:50:XX:XX:
|
|
||||||
# - 52:54:00 — QEMU/KVM OUI; widely understood by network tools as
|
|
||||||
# virtual-machine traffic and matches libvirt defaults.
|
|
||||||
# - 50 — fixed distinguisher; literal hex "50" mnemonically
|
|
||||||
# matches VLAN 50, separating Talos VMs from any other
|
|
||||||
# QEMU VMs that might land on these hosts later.
|
|
||||||
# - last 2 bytes — computed in talos_vms role from a deterministic
|
|
||||||
# hash of vm_name so the same VM gets the same MAC
|
|
||||||
# across reprovisions (stable identity at the switch;
|
|
||||||
# workers' VLAN-51 DHCP leases stay pinned).
|
|
||||||
talos_vm_mac_prefix: "52:54:00:50"
|
|
||||||
|
|
||||||
# Secondary MAC prefix for the VLAN-51 Services NIC on worker VMs.
|
|
||||||
# Same scheme — 52:54:00 QEMU OUI, fourth octet "51" for VLAN-51
|
|
||||||
# mnemonic, last two bytes from the same per-vm-name sha1 hash as the
|
|
||||||
# primary NIC. Workers thus end up with two related but distinct MACs
|
|
||||||
# (e.g., worker3 → 52:54:00:50:5a:cd on VLAN 50, 52:54:00:51:5a:cd on
|
|
||||||
# VLAN 51), which is easy to correlate at switch / DHCP server level.
|
|
||||||
# CPs do not get a Services NIC; only used when item.role == 'worker'.
|
|
||||||
talos_vm_mac_prefix_services: "52:54:00:51"
|
|
||||||
|
|
||||||
# ─── VLANs ──────────────────────────────────────────────────────────────
|
|
||||||
# Both VLANs are tagged upstream at the switch and arrive on the host
|
|
||||||
# as tagged frames over bond0. Roles add bridge + VLAN sub-interfaces;
|
|
||||||
# they don't touch the physical bond.
|
|
||||||
vlan_compute_id: 50
|
|
||||||
vlan_compute_name: "compute"
|
|
||||||
vlan_compute_bridge: "br-vlan50"
|
|
||||||
vlan_compute_subnet: "10.50.0.0/16"
|
|
||||||
vlan_compute_dhcp_start: "10.50.0.16"
|
|
||||||
vlan_compute_dhcp_end: "10.50.4.255"
|
|
||||||
|
|
||||||
vlan_services_id: 51
|
|
||||||
vlan_services_name: "services"
|
|
||||||
vlan_services_bridge: "br-vlan51"
|
|
||||||
vlan_services_subnet: "10.51.0.0/16"
|
|
||||||
|
|
||||||
# ─── Talos cluster topology ─────────────────────────────────────────────
|
|
||||||
# Reserved (static) IPs sit OUTSIDE the DHCP range above so leases
|
|
||||||
# can't collide. VIP is the documented control-plane endpoint that
|
|
||||||
# operators put in their kubeconfig.
|
|
||||||
talos_cp_vip: "10.50.0.10"
|
|
||||||
|
|
||||||
# ─── libvirt host config ────────────────────────────────────────────────
|
|
||||||
# Storage pool that holds base image + per-VM overlays + (later)
|
|
||||||
# snapshots. Lives on the OS root by default; switch to RBD in a
|
|
||||||
# follow-up iteration.
|
|
||||||
libvirt_storage_pool_name: "talos"
|
|
||||||
libvirt_storage_pool_path: "{{ talos_image_dir }}"
|
|
||||||
|
|
||||||
# ─── Profile selector ───────────────────────────────────────────────────
|
|
||||||
# `full` = 3 CP + 3 workers (production, the default) — one CP + one
|
|
||||||
# worker per hypervisor, so each host is its own failure domain
|
|
||||||
# for both the control plane and the worker pool.
|
|
||||||
# `smoke` = 1 CP + 1 worker (both on laurel) — single-host validation for
|
|
||||||
# substrate changes; select explicitly with `-e profile=smoke`.
|
|
||||||
# Per-host VM lists are declared in host_vars and gated by this.
|
|
||||||
talos_profile: "full"
|
|
||||||
@@ -1,26 +0,0 @@
|
|||||||
; Example inventory for sietch-talos. The RUNTIME inventory.ini in this
|
|
||||||
; same directory is TF-rendered (tf/shared/modules/talos-cluster/
|
|
||||||
; modules/inventory-renderer/) and gitignored. This .example file is
|
|
||||||
; checked in as documentation of the expected shape.
|
|
||||||
;
|
|
||||||
; To render the runtime inventory:
|
|
||||||
; TF_STACK_DIR=tf/deployment/dev/talos mise run tf:apply
|
|
||||||
;
|
|
||||||
; Sietch hypervisors. The 3 nodes are co-located in Austin, run Ceph
|
|
||||||
; OSDs on bare metal, and host Talos VMs as guests.
|
|
||||||
;
|
|
||||||
; ansible_host = bond0 IP (10.10.10.0/24 management/storage VLAN, the
|
|
||||||
; existing Ceph cluster network — NOT the new Talos VLAN 50).
|
|
||||||
; If 10.10.10.0/24 isn't directly reachable from your workstation, put any
|
|
||||||
; needed SSH bastion / ProxyJump in your ~/.ssh/config.
|
|
||||||
|
|
||||||
[hypervisors]
|
|
||||||
sietch-ceph-laurel ansible_host=10.10.10.90
|
|
||||||
sietch-ceph-lawson ansible_host=10.10.10.91
|
|
||||||
sietch-ceph-samara ansible_host=10.10.10.92
|
|
||||||
|
|
||||||
[hypervisors:vars]
|
|
||||||
ansible_user=ansible-iac
|
|
||||||
ansible_ssh_private_key_file=~/.ssh/id_ed25519_sietch
|
|
||||||
ansible_ssh_common_args='-o ControlMaster=auto -o ControlPersist=60s'
|
|
||||||
ansible_python_interpreter=/usr/bin/python3
|
|
||||||
@@ -5,7 +5,7 @@ talos_vms_enabled: false
|
|||||||
# Libvirt domain + qcow2 + NVRAM file name prefix.
|
# Libvirt domain + qcow2 + NVRAM file name prefix.
|
||||||
# Default scopes to "sietch-talos" so VMs end up as sietch-talos-cp1,
|
# Default scopes to "sietch-talos" so VMs end up as sietch-talos-cp1,
|
||||||
# sietch-talos-worker1, etc. — aligns with the monorepo FQDN scheme
|
# sietch-talos-worker1, etc. — aligns with the monorepo FQDN scheme
|
||||||
# `<cluster>-<role>-<name>.<domain>` (e.g. sietch-talos-cp1.dev.austin.int.futo.cloud)
|
# `<cluster>-<role>-<name>.<domain>` (e.g. sietch-talos-cp1.staging.austin.int.futo.cloud)
|
||||||
# and matches the *_TALOS_* 1P secret-prefix convention. Read from
|
# and matches the *_TALOS_* 1P secret-prefix convention. Read from
|
||||||
# group_vars `talos_domain_prefix` when present so a future second
|
# group_vars `talos_domain_prefix` when present so a future second
|
||||||
# cluster can override without role edits.
|
# cluster can override without role edits.
|
||||||
|
|||||||
@@ -7,4 +7,4 @@ appVersion: "0.0.1"
|
|||||||
dependencies:
|
dependencies:
|
||||||
- name: yucca-common
|
- name: yucca-common
|
||||||
version: 0.1.0
|
version: 0.1.0
|
||||||
repository: "file://../yucca-common"
|
repository: "file://../../lib/yucca-common"
|
||||||
@@ -7,4 +7,4 @@ appVersion: "0.0.1"
|
|||||||
dependencies:
|
dependencies:
|
||||||
- name: yucca-common
|
- name: yucca-common
|
||||||
version: 0.1.0
|
version: 0.1.0
|
||||||
repository: "file://../yucca-common"
|
repository: "file://../../lib/yucca-common"
|
||||||
@@ -7,4 +7,4 @@ appVersion: "0.0.1"
|
|||||||
dependencies:
|
dependencies:
|
||||||
- name: yucca-common
|
- name: yucca-common
|
||||||
version: 0.1.0
|
version: 0.1.0
|
||||||
repository: "file://../yucca-common"
|
repository: "file://../../lib/yucca-common"
|
||||||
@@ -7,4 +7,4 @@ appVersion: "0.0.1"
|
|||||||
dependencies:
|
dependencies:
|
||||||
- name: yucca-common
|
- name: yucca-common
|
||||||
version: 0.1.0
|
version: 0.1.0
|
||||||
repository: "file://../yucca-common"
|
repository: "file://../../lib/yucca-common"
|
||||||
+1
-1
@@ -7,4 +7,4 @@ appVersion: "0.0.1"
|
|||||||
dependencies:
|
dependencies:
|
||||||
- name: yucca-common
|
- name: yucca-common
|
||||||
version: 0.1.0
|
version: 0.1.0
|
||||||
repository: "file://../yucca-common"
|
repository: "file://../../lib/yucca-common"
|
||||||
@@ -7,4 +7,4 @@ appVersion: "0.0.1"
|
|||||||
dependencies:
|
dependencies:
|
||||||
- name: yucca-common
|
- name: yucca-common
|
||||||
version: 0.1.0
|
version: 0.1.0
|
||||||
repository: "file://../yucca-common"
|
repository: "file://../../lib/yucca-common"
|
||||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user