* feat: introduce partition/region/ceph-cluster model across the stack
Formalize partition -> region -> {one k8s cluster, many ceph clusters} and
thread it through every layer plus a new yuctl ops CLI.
- tf: deployment/<partition>/<region>/<stack> layout; terragrunt path-parse +
state key yucca/<partition>/<region>/<stack>; per-region region.hcl (role,
site_id, datacenter, provider_code, domain); env->partition / site->region
renames (NetBird object names byte-identical); standardized per-stack
`discovery` output contract (secrets as op:// refs).
- k8s: clusters/<partition>/<region>/ (staging/austin, prod/htz-fsn1, dev/local);
role-based kustomize components (primary/secondary); hybrid cluster-settings
(TF-rendered identity + human fragment); dev-mirror folded into dev/local;
charts regrouped into charts/{apps,platform,lib,dev}.
- ci: infra.yml partition/region discovery matrix; partition-keyed path filters;
<partition>-<region> environment gates; image-versions path moves.
- ansible: inventories under <partition>-<region>/<cluster>.
- yuctl: Go/cobra CLI reading the discovery contract from TF state.
- Retire the sietch-talos libvirt VM cluster (dev@local is the k3d cluster);
ceph inventory_dirname -> <partition>-<region>/<cluster>.
Verified: mise k8s:validate green (3 clusters); yuctl go build/vet; tofu
validate pre-merge (all 9 stacks). Live-staging state migration NOT run.
* fix typo
* commit
mgmt — Hetzner FSN1 Management Hosts
Ansible automation that configures the two Hetzner management hosts after they
are reprovisioned to Debian 13 ("trixie"). Mirrors the conventions of the
sibling ansible/ceph/ tree (layout, ansible.cfg, role structure,
systemd-networkd templating).
| Host | Public IP | Role |
|---|---|---|
htz-fsn-mgmt-1 |
178.63.124.40 |
NetBird route peer (routes 10.40.5.0/24 et al.) |
htz-fsn-mgmt-2 |
178.63.124.41 |
NetBird route peer |
Both are AX41-NVMe, both join the NetBird mgmt group and route the site
subnets (the routed network is declared in TF, tf/deployment/prod/htz-fsn1/netbird).
The inventory is TF-generated at run time (see "Generated inventory").
What it does
site.yml applies these roles in order:
- baseline — apt cache + base packages, timezone UTC, hostname from
inventory,
/etc/hosts. - users — login users for the members of the identity registry's
server-mapped groups (nutgood,andy— both inserver_admins,sudo = ALL). Themgmt_userslist is TF-generated fromtf/shared/modules/identity(see "Generated inventory" below). - security — nftables firewall, SSH hardening (no password auth,
PermitRootLogin prohibit-password), unattended-upgrades. - networkd — the host L3 that lets these nodes forward the NetBird-routed
site subnets: the OOB 1G NIC on
10.40.5.0/24(the switch vme) + the 25G fabric VLAN sub-interfaces (public/private/api). Enabled by default (mgmt_networkd_enabled: true); only the OOB + fabric NICs are matched, so the primary public NIC is untouched. The OOB path is reliable; the 25G VLANs are written but stay carrier-down until that link is up — see the 25G caveat below. - netbird — install NetBird,
netbird upwith themgmtsetup key, enable IP forwarding (the mgmt nodes are the NetBird route peers for the site subnets; the routed network itself is declared in TF).
Generated inventory
The inventory is generated by Terraform at run time, not hand-written.
mise run mgmt:render-inventory (run automatically by mgmt:ansible) invokes
tf/render/ansible-mgmt, which derives everything from the single sources of
truth and writes these gitignored files:
| file | generated from |
|---|---|
inventories/<region>/hosts.yml |
mgmt-hosts.yaml (host names + public IPs) |
inventories/<region>/host_vars/*.yml |
mgmt-hosts.yaml (NIC) + fabric-addressing (VLAN ids/addresses, subnet route) |
inventories/<region>/group_vars/all/users.generated.yml |
tf/shared/modules/identity (server-mapped users) |
Only group_vars/all/main.yml (static config) and roles/** are committed. To
change hosts, addresses, or users, edit the Terraform sources — never the
generated files. The render uses only the local provider (no backend, no
secrets), so it runs anywhere.
Reprovision → Ansible flow
- Reprovision both hosts to Debian 13 via Hetzner robot auto-install.
- Post-reprovision, root is reachable over SSH with the TF-generated
provisioning key (stored in 1Password). Render it and run
site.yml.
# Render the provisioning private key from 1Password to a temp file
umask 077
op read --account team-futo \
"op://yucca_tf_prod/HTZ_FSN1_PROVISIONING_SSH_PRIVATE_KEY/password" \
> /tmp/htz-fsn1-prov-key
chmod 600 /tmp/htz-fsn1-prov-key
# Run the playbook (root, provisioning key, NetBird mgmt setup key from 1P)
ansible-playbook -i inventories/htz-fsn1 site.yml \
--private-key /tmp/htz-fsn1-prov-key \
--extra-vars "mgmt_netbird_setup_key=$(op read --account team-futo \
'op://yucca_tf_prod/NETBIRD_YUCCA_PROD_HTZ_FSN1_MGMT_SETUP_KEY/password')"
# Clean up
shred -u /tmp/htz-fsn1-prov-key
ansible.cfg sets inventory = inventories/htz-fsn1/hosts.yml, so -i is
optional. The inventory connects as ansible_user: root.
The whole render-key + run flow above is wrapped by mise run mgmt:ansible
(REGION selects the inventory; defaults to htz-fsn1), which CI also runs on
every prod apply — the Ansible converge (mgmt hosts) step of
.github/workflows/infra.yml, right after the Terraform apply. It's idempotent
and reaches the hosts over their public IP, so it requires them to already be
reprovisioned (provisioning key authorized).
The
--private-keyflow is preferred overansible_ssh_private_key_filein group_vars so the key never has to be persisted to a committed path — it is rendered to a temp file, used, and shredded.
Connection: public IP → NetBird
A freshly-reprovisioned host is only reachable over its public IP, so that's
the bootstrap address (mgmt_public_ip in host_vars). The first play in
site.yml probes the overlay from the control node (netbird status --json):
if the host is a connected peer, it switches ansible_host to the host's
NetBird IP for the rest of the run; otherwise it stays on the public IP.
So the first run provisions over the public IP and brings the host onto the
overlay (the netbird role); every subsequent run reconnects over NetBird
automatically. This needs the control node on the overlay too — the CI
prod-ansible job joins via the ci setup key; locally, be connected to NetBird.
For the very first run after a reinstall, pass -e mgmt_bootstrap=true to force
the public IP (skips any stale peer entry for the host).
Host L3 / routing & the 25G caveat
These nodes are the NetBird route peers for the site, so they need L3 paths to
the routed subnets in order to forward to them. The networkd role
(enabled by default) configures, per the TF-rendered host_vars:
| Path | Network | mgmt-1 | mgmt-2 | NIC |
|---|---|---|---|---|
| OOB management (switch vme) | 10.40.5.0/24 |
10.40.5.50 |
10.40.5.51 |
enp37s0 (1G) |
| VLAN 120 (cluster public) | 10.40.20.0/23 |
10.40.20.2 |
10.40.20.3 |
enp33s0f0np0 (25G) |
| VLAN 122 (cluster private) | 10.40.22.0/23 |
10.40.22.2 |
10.40.22.3 |
enp33s0f0np0 (25G) |
| VLAN 10 (api) | 10.40.10.0/24 |
10.40.10.2 |
10.40.10.3 |
enp33s0f0np0 (25G) |
The OOB path is 1G and reliable — that's what carries the switch traffic.
The 25G fabric link (Intel E810, "ice") is currently physically unreliable;
its VLAN sub-interfaces are written but stay carrier-down (RequiredForOnline=no,
so they never block boot) until the link is up — harmless until then.
After a reinstall, confirm both NIC names with ip link and correct
oob_nic / fabric_nic in tf/deployment/prod/htz-fsn1/mgmt-hosts.yaml if they
differ (predictable names can change). Only these NICs are matched; the primary
public NIC keeps Hetzner's DHCP default — this tree does not touch it.
Setup
ansible-galaxy collection install -r requirements.yml
ansible-playbook -i inventories/htz-fsn1 --syntax-check site.yml
Secrets
Nothing secret is committed. SSH public keys are public data and are TF-generated
into group_vars/all/users.generated.yml. Runtime secrets are passed via op read:
- Provisioning private key:
op://yucca_tf_prod/HTZ_FSN1_PROVISIONING_SSH_PRIVATE_KEY/password - NetBird mgmt setup key:
op://yucca_tf_prod/NETBIRD_YUCCA_PROD_<REGION>_MGMT_SETUP_KEY/password(the site's reusablemgmtkey,auto_groups=["mgmt"]; minted by the prod netbird stack)