Files
yucca/tf
Andy MolendaandClaude Opus 4.8 0357aac743 feat(ceph): prod spice cluster on htz-fsn1 (#256)
* feat(ceph): scaffold spice prod cluster (reprovision + inventory + TF)

Stand up spice (48x Hetzner SX295, prod/htz-fsn1): the reprovision_hetzner
ansible role (rescue -> reset -> installimage -> verify, base-OS-only, with
stale-mdraid pre-clean and resume markers), the prod-htz-fsn1/spice inventory
+ group_vars/host_vars, the prod ceph TF stack, and gen-spice-host-vars.py.
Adds a `mise reprovision` task (op run + tf/.env.prod), the spice SSH key
(yucca_tf_prod), and per-host roles-based [ceph_mon] filtering so large
clusters pin a fixed MON quorum instead of defaulting to all nodes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): create spice block.db + ssd-osd LVs at converge (NVMe-RAID)

Add an NVMe-RAID branch to ceph_deploy lvm-setup so spice's single vg0 (built
by installimage) gets its 14 block.db LVs + one ssd-osd LV created at converge,
before osds.yml. This replaces painbox's fragile installimage -x chroot
post-install with an idempotent, observable ansible step; installimage stays
base-OS-only. Narrows the old blanket "externally-managed LVM" skip so it only
fires when neither the sietch dual-SSD nor the NVMe-RAID shape applies, and
wires ceph_db_vg / ceph_ssd_osd_lv / ceph_ssd_osd_reserve_gib in the spice
group_vars.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): set spice bond NICs in group_vars, drop per-host placeholders

The 25G fabric NICs are uniform across all 48 SX295 (verified live on 36
nodes): bond members enp193s0f0/f1 (Intel E800/ice, PCI c1:00.0/.1), WAN
enp197s0 (igb, c5:00.0). Set bond_interfaces + oob_nic once in the spice
group_vars -- the networkd role reads bond_interfaces -- and remove the dead
per-host fabric_nic/oob_nic PLACEHOLDER lines from all 48 host_vars,
spice-hosts.yaml, and gen-spice-host-vars.py. No MAC-based naming needed: the
predictable names are PCI-derived and identical fleet-wide.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): LACP + VLAN sub-interfaces on the bond, lldpd fabric config

networkd role: emit 802.3ad LACP params (TransmitHashPolicy, LACPTransmitRate)
when bond_mode is 802.3ad, and support tagged VLAN sub-interfaces on the bond
via networkd_bond_vlans -- when set, bond0 carries no L3 and each VLAN carries
its own address. Opt-in: an empty list keeps the flat active-backup path
(sietch) byte-identical (verified by render). verify.yml gates the
bond0-IP/gateway asserts to the flat case and checks per-VLAN addresses
otherwise.

spice: bond0 becomes an 802.3ad LACP bond of enp193s0f0/f1 (MLAG to the QFX5240
leaves) carrying VLAN 120 (public 10.40.20.<idx>/23) + VLAN 122 (private
10.40.22.<idx>/23); default route stays on the 1G WAN until cutover.

baseline role: install + configure lldpd (portid ifname, cluster system
description, service enabled) when lldpd is in baseline_extra_packages, and add
lldpd/ethtool/tcpdump for spice. This also newly applies lldpd config to sietch
(already lists lldpd) on its next converge.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): set spice RGW pool to EC 16+4 host domain for the beta

Capacity-max profile (80% usable, m=4) for the wipeable beta behind michael.
Host failure domain accepts rack-loss risk in exchange for capacity, which is
fine for a beta and matches the single-rack API tier. Also lifts the RGW DRAFT
marker and starts the bulk EC data pool at pg_num 4096 (672 OSDs, 20-chunk pool)
instead of the 36-OSD-era default to avoid PG splitting during the first fill.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* docs(ceph): drop stale -x chroot reference from spice autosetup header

installimage.yml runs base-OS-only (installimage -a -c /autosetup, no -x). The
autosetup header still carried the painbox `-x /tmp/post-install.sh` line, which
implied a chroot post-install the role deliberately dropped. Point it at the
convergence steps (lvm-setup, baseline) that replaced it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): minimal -x chroot to seed root + ansible-iac keys on reprovision

installimage sometimes runs without our key present in the rescue (Robot armed
rescue without it, in-place installimage, manual rescue), so its "copy the rescue
authorized_keys" step leaves root keyless and the node unreachable for
convergence. Add a minimal `-x /tmp/post-install.sh` that does ONLY inert
user/key writes: seed root's authorized_keys with the iac key (baseline never
manages root, so it persists) and create an ansible-iac account with the key +
NOPASSWD sudo. No apt/LVM in the chroot -- that fragile step stays at post-boot
convergence, so set -e cannot spuriously disrupt installimage.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): converge iac access on existing nodes without reimaging

Nodes imaged before the reprovision -x chroot lack root's iac key and the
ansible-iac sudo account; reseeding them meant a reinstall. baseline gains an
iac-access task that ensures the same state idempotently -- root's authorized
key (non-exclusive) plus the ansible-iac user with key and NOPASSWD sudo -- so
the 36 already-installed nodes converge in place. Gated on the cluster iac key;
run standalone with --tags iac_access.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): manual rescue-first add-node path, Robot API optional

The Hetzner Robot rescue/reset step is both the unreliable part (nodes that arm
rescue but never boot it) and the highest blast radius (an errant /reset on a
live spice node in production). Gate it behind reprovision_use_robot (default
true, preserving reprovision.yml) and add add-node.yml, which sets it false: the
operator arms rescue by hand in the portal and Ansible runs only install, chroot
and verify on a node already in rescue. Robot key registration moves to
register_robot_keys.yml, imported from preflight only on the Robot path. With the
API out of the loop nothing here can flip a running node into rescue -- wait_rescue
refuses any node without the installimage ramdisk.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): relax stale-rescue guard on the manual add-node path

wait_rescue rejects a rescue older than reprovision_rescue_max_uptime (1200s) to
catch a stale boot -- valid only when Ansible triggered the reset. On the manual
path the operator rescues by hand, so uptime just measures wait-before-run and a
node legitimately in rescue for hours would be refused. Relax the guard to a day
on add-node.yml; the installimage-ramdisk check is the real in-rescue proof.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): skip the this-boot rescue guard on the manual path

philip has sat in rescue ~8 days, so any finite uptime ceiling rejects it. The
guard only means something when Ansible triggered the reset (uptime proves this
boot); on the manual add-node path a node may sit in rescue for weeks, so gate
the assert on reprovision_use_robot rather than inflating the timeout. The
installimage-ramdisk check remains the real in-rescue proof.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): networkd/ifupdown coexistence to protect the 1G WAN

The role assumed a full migration off ifupdown -- correct for sietch's flat bond,
but on spice it would disable ifupdown and leave the 1G WAN (a separate NIC that
ifupdown owns, carrying the default route) unconfigured on the next boot: the
painbox failure. Add networkd_replace_ifupdown (default true, sietch unchanged).
When false, networkd manages only the bond and VLANs; an Unmanaged=yes guard
fences the WAN NIC off, commit leaves networking.service enabled, the bond drops
to RequiredForOnline=no so a carrier-less fabric cannot stall boot, and verify
asserts the WAN kept its address and default route before anything commits. The
rollback script targets the right interface per mode. spice sets it false.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): make migrate-networkd usable pre-Ceph (gate health checks)

The ceph-health pre/post gate assumed a live cluster and a ceph_bootstrap group,
so the play could not run during spice bringup (coexistence activation before any
cephadm deploy). Gate the four ceph tasks on ceph_bootstrap being populated; on a
pre-Ceph cluster they skip and the networkd role runs on its own.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* chore(ceph): tunable serial + halt-on-failure for migrate-networkd

Fleet rollout of the coexistence activation wants batches, not one-at-a-time, and
must stop the instant a node fails (a dropped WAN shows up as unreachable) rather
than silently skip it. Template serial (networkd_serial, default 1 unchanged) and
set max_fail_percentage 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): install Intel ice DDP firmware so E810 NICs leave Safe Mode

The minimal Debian image ships no firmware-misc-nonfree, so the E810 DDP package
(intel/ice/ddp/ice.pkg) is absent and every NIC boots into Safe Mode -- whose
crippled classifier drops reserved-multicast control frames (LACP, LLDP), so the
25G fabric bond never aggregates despite a correctly configured switch. Add
firmware-misc-nonfree to the spice package set and a baseline nic-firmware task
that reboots an E810 host once to load the DDP when it is still in Safe Mode
(self-gating: no-op on non-ice or already-loaded hosts; refuses to reboot when the
DDP is absent, so it cannot loop). The WAN is a separate igb NIC, so the reboot
keeps the node reachable.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): correct ice DDP activation gating (pipefail + dpkg check)

The Safe Mode check used `set -o pipefail`, but `ethtool --show-fec` exits
non-zero on "not supported", which pipefail propagated and masked grep's match --
so the check always yielded "ok" and the activating reboot never fired. Drop
pipefail there, and gate the reboot on the dpkg DB (firmware-misc-nonfree
installed) rather than a file stat that can lag a large apt transaction. Verified
on one node end-to-end: DDP loads, Safe Mode clears, LACP converges with the
switch, both VLAN gateways ping.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): pin chrony time sources for Ceph time sync

chrony was listed in baseline_enable_services but never installed, so system.yml
would fail starting it -- and there was no time-source config at all, leaving sync
to the distro default pool. Ceph mon quorum is skew-sensitive, so make it explicit:
add baseline/chrony.yml (install + templated chrony.conf + enable) driven by
chrony_ntp_servers (default Debian pool, makestep for the initial correction), drop
chrony from baseline_enable_services so it is owned in one place, and point spice at
Hetzner NTP (ntp1/2/3.hetzner.de) -- low-latency from FSN1, consistent across all 47.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): populate spice OSD by-path map; exclude noreen until recovered

Inspected the SX295 disk layout: 14x 20TB SATA HDDs per node across 3 AHCI
controllers, by-path uniform on 46/47 nodes -- so ceph_hdd_osds (path_phy + db
vg0/db-slotN) and ceph_ssd_osds (vg0/ssd-osd) live in group_vars, not 47 host_vars.
spice-ceph-miguel has one disk on 46:00.0-ata-4 rather than 87:00.0-ata-4 and
overrides the map in its host_vars. Verified by rendering osd-spec.yml.j2: 658 HDD
OSD paths (47x14) + 47 NVMe ssd-osd, miguel's override resolving correctly. noreen
(boot-order casualty, held in triage) is commented out of clusters.auto.tfvars so
the rendered inventory + deploy target only the 47 live nodes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): bootstrap mon on the fabric public_network, not the WAN

cephadm bootstrap --mon-ip must fall inside public_network, but spice
passed bond_ip, the 1G WAN address used for the ansible/SSH connection
and cephadm host management. spice's public_network is the 25G fabric
(10.40.20.0/23), so the initial mon would fail to bind. Bootstrap now
uses ceph_public_ip (10.40.20.<host_index>, derived like the bond VLANs);
sietch is flat with no host_index and falls back to bond_ip, already in
its own public_network.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* ci(ceph): static validation gate for the ansible/ceph stack

CI only validated the Kubernetes/Flux surface; nothing parsed the ceph
playbooks, inventory or scripts, so a broken playbook or host_var first
executed during the post-merge prod apply against real nodes. This runs
the existing local validators (yamllint, ansible-lint, shellcheck,
ansible-playbook --syntax-check, py_compile) on ansible/ceph PRs with no
secrets and no connection to any host; a throwaway localhost inventory
satisfies --syntax-check since the real inventory is TF-generated.

Its own workflow, not a job in ci.yml, so the ansible/ceph/** path filter
is workflow-scoped and does not gate the unrelated k8s jobs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): make the ceph converge deliberate and cluster-parameterized

A merge touching ansible/ceph/** could open the staging partition and
auto-run the full baseline->tune->deploy->harden pipeline against the
live sietch cluster, and the converge was hardcoded to sietch/austin
(render-inventories.sh got no region, install-ssh-keys.sh got a literal
sietch, CEPH_ENV pointed at .../sietch/), so a prod converge could never
reach spice.

The ceph converge now runs only on workflow_dispatch with
run_ceph_ansible=true, pinned to the one matrix entry that owns the
chosen ceph_cluster (spice=prod/htz-fsn1, sietch=staging/austin), so a
push never reconverges a live cluster and one dispatch cannot converge
both. Region and cluster are threaded from the dispatch inputs / matrix
into the render, key install, and CEPH_ENV. The staging paths-filter is
scoped to ansible/ceph/inventories/staging-** so a prod ceph change no
longer opens the staging matrix. The ceph TF apply stays auto and
env-gated; the mgmt converge is unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): make the RGW firewall source scope configurable

The RGW/S3 port was accepted from any source unconditionally, while
every other Ceph service (mon, osd, dashboard, monitoring) is scoped to
ceph_firewall_trusted_networks. On spice that port is plain HTTP 7480, so
the S3 endpoint sat open to the public internet with no lever to close
it. A new ceph_firewall_rgw_any_source (default true, mirroring
ceph_firewall_ssh_any_source) keeps the open behavior by default but lets
production restrict RGW to the trusted networks, which already cover the
fabric and NetBird overlay.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* chore(ceph): fix stale OSD-map comment on the noreen host_var

noreen was generated before the by-path map moved into group_vars, so it
still carried the DEFERRED note while its 47 siblings point at the group
var. The host stays pre-staged for re-add once recovered; only the
comment was wrong.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): add the host-mgmt VLAN (124) to the spice bond

The FSN1-C1 fabric carries a third VLAN on the 25G bond,
FSN1-C1-HOST-MGMT (124, 10.40.24.0/24), alongside Ceph public (120) and
private (122). The leaf already trunks it to every server bond and
advertises 10.40.24.0/24 into the NetBird overlay, so it is the intended
in-band ansible/SSH reach once the WAN is retired, but the host side had
no matching sub-interface. Add bond0.124 at 10.40.24.<host_index>/24 with
no gateway, so the default route stays on the 1G WAN until cutover.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): apply networkd config deltas to an already-running daemon

The networkd role only started systemd-networkd, which is a no-op once it
is already running, so a re-run that added config (e.g. a new bond VLAN)
wrote the .netdev/.network files but never applied them, and verify then
failed asserting the sub-interface had no address. Add a networkctl reload
plus a per-VLAN settle wait after the start, so a re-run creates the added
sub-interfaces live without tearing down existing links; the WAN on
ifupdown is untouched regardless. Non-disruptive on the flat sietch path
(the settle wait is gated on networkd_bond_vlans).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): bind Ceph services on the fabric only, never the WAN

Ceph on spice must never listen on the 1G WAN. mon/mgr/OSD already bind
per public_network/cluster_network, but the beast RGW frontend and the
cephadm monitoring stack bound 0.0.0.0, and the mon-ip, cephadm host-add
address, and dashboard URLs advertised the WAN bond_ip. Add a per-cluster
policy: ceph_service_ip (the address Ceph advertises and binds for point
services; defaults to bond_ip, resolving to the fabric on spice via
ceph_public_ip) and ceph_bind_networks (the cephadm networks: bind
restriction for RGW and the monitoring daemons; defaults to
public_network). Route mon-ip, the host-add address, the dashboard
monitoring URLs, RGW, and prometheus/grafana/alertmanager/node-exporter/
ceph-exporter through them, and scope the RGW firewall to the trusted
networks. The RGW zonegroup hostnames and TLS SAN gain the fabric address
so signed requests to the fabric IP still validate. Flat clusters like
sietch have no ceph_public_ip, so ceph_service_ip falls back to bond_ip
and ceph_bind_networks to their single flat network; point services are
unchanged there.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): define the metrics-worker RGW user on spice

rgw.yml Step 14.5 creates a read-only RGW admin for the yucca-metrics-worker
service and references ceph_rgw_metrics_user_*, which only sietch defined,
so the RGW phase failed on spice with AnsibleUndefinedVariable. Add the
block; the keys are the TF-minted SPICE_METRICS_WORKER_* items, injected
via secrets.yml.tpl as vault_metrics_worker_*, so op inject resolves them.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): refuse to deploy against an un-rendered inventory

The ceph_deploy role gates every phase on the ceph_bootstrap/ceph_mon/
ceph_join groups that render-inventories.sh emits from TF state; run
against the hand-written stopgap inventory that lacks them, the role
errors mid-deploy or falls through to placing a MON on every host. A
pre-task assert now requires ceph_bootstrap to be exactly one host that
is also in ceph_mon, and ceph_mon to be a non-empty odd quorum, and
stops with a render hint otherwise.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): make the fabric-only bind restriction opt-in per cluster

ceph_bind_networks defaulted to public_network, which would have injected
a networks: bind restriction into every cluster - re-binding sietch's live
RGW and monitoring daemons from 0.0.0.0 to its flat network on the next
converge (a redeploy, and a break for any access path not on that subnet).
Default it empty instead: no networks: field is emitted and the monitoring
re-spec is skipped, so flat clusters bind every interface exactly as
cephadm ships them. spice opts in to the fabric public network in its
group_vars.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): make ceph_service_ip a group_var so hostvars can read it

ceph_service_ip was a ceph_deploy role default, but join/rgw/monitoring
read it as hostvars[<host>]['ceph_service_ip'], and role defaults are NOT
exposed through hostvars - the lookup raised "HostVarsVars object has no
attribute ceph_service_ip" and would abort the deploy on every cluster at
the join phase (a regression the fabric-only change introduced for both
spice and the live sietch). Define it in each cluster's group_vars instead
(group_vars do resolve through hostvars, verified per host): spice to the
fabric ceph_public_ip, sietch to bond_ip (unchanged from pre-fabric).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* chore(ceph): drop the unused ceph_cluster_ip var on spice

ceph_cluster_ip was defined for symmetry with ceph_public_ip but consumed
nowhere - OSD replication binds to the cluster_network CIDR, which cephadm
resolves to each node's bond0.122 address on its own. Remove the dead var
and correct the neighbouring comment (ceph_service_ip is a group_var now,
not a role default).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): serve spice RGW over TLS on 443 with a 10-year self-signed cert

Match sietch: flip spice RGW/beast from plain HTTP 7480 to HTTPS 443 with a
self-signed 10-year cert. rgw.yml generates /etc/ceph/rgw-ssl.{crt,key} and
cephadm distributes it to every RGW daemon via the service spec; the SANs
already cover s3.<domain> + the wildcard + each node's fabric IP
(ceph_service_ip). The firewall RGW port derives from ceph_rgw_port, so it
follows to 443, scoped to the trusted networks (RGW binds the fabric only,
never the WAN). This also makes the discovery rgw_s3_endpoint (hardcoded
https) correct for spice.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(dns): prod Cloudflare DNS stack for the spice RGW S3 endpoint

Add tf/deployment/prod/global/dns (mirroring staging/global/dns) so
s3.prod.fsn1.htz.futo.cloud and its wildcard resolve. Round-robin A across
all 47 spice ceph nodes' fabric public IPs (10.40.20.<host_index>, VLAN 120)
where RGW/beast binds; proxied=false (private RFC1918, reached over the
NetBird-advertised cls1_public 10.40.20.0/23). The wildcard serves S3
virtual-hosted buckets, and both names are in the self-signed TLS cert SANs.
noreen (host_index 40) excluded. CI auto-discovers the stack (applies at
order 0 under the prod-global environment); tf/.env.prod gains the token ref.

Next up: create op://yucca_tf_prod/CLOUDFLARE_API_TOKEN before the apply.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): preflight-assert the Ceph service IP is up before deploy

A node whose fabric VLAN sub-interface (spice bond0.120) did not come back
after a reboot would otherwise fail deep inside cephadm bootstrap/join. Add
a per-host pre_task that asserts ceph_service_ip is present in
ansible_all_ipv4_addresses, so a fabric-down node halts up front with a
clear message. Passes on a healthy node (verified on spice-ceph-adelia);
on sietch ceph_service_ip is bond_ip, the connection address, so it is
trivially satisfied.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* chore(ceph): honest ssd_model_pattern note + guard empty monitoring networks

ssd_model_pattern on spice was "SAMSUNG" with a "verify on node 1" note, but
the NVMe is SOLIDIGM/KIOXIA and the var is only read by ceph_destroy's
partition-6 SSD wipe (sietch's dual-SSD shape); spice is NVMe-RAID with no
partition 6 and a vg0 ssd-osd LV cleaned by the generic VG/PV removal, so
the pattern is a deliberate no-match. Reword to say so (it must stay defined
because cleanup.yml references it unconditionally). Also wrap the
monitoring-spec networks block in a length guard so an empty ceph_bind_networks
renders no dangling `networks:` key (defensive; the apply is already gated).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): lock down root/password SSH in baseline, not just harden

installimage ships PermitRootLogin yes + a root password, and all sshd
hardening lived in harden.yml (the last deploy stage), so a freshly imaged
node sat root-password-open on the public WAN for the whole campaign (the
11-node exposure the swarm audit found). Add an early baseline task, right
after iac-access authorizes the key on root, that deploys a 10-baseline-ssh
drop-in (PermitRootLogin prohibit-password + PasswordAuthentication no,
same values as security/50-hardening.conf so they never disagree), locks
the root password, and removes the interim remediation drop-in. A
Validate->Reload handler chain runs sshd -t before reloading, only on
change. Key-safe on both clusters (ansible connects by key), so nothing
can lock out.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): pin RGW metadata + mgr pools to the replicated size

The RGW system-metadata pools (.rgw.root, {zone}.rgw.{log,control,meta})
and .mgr inherited the cluster-spec default osd_pool_default_size 2 /
min_size 1, unlike the index/extra pools which are pinned. Losing two OSDs
behind a metadata PG would take out the RGW/mgr control plane, and
min_size 1 permits single-replica writes. Pin them to ceph_rgw_replicated_
size/_min_size (spice 3/2; sietch keeps 2/1 via its own vars), existence-
gated like the pg_num loop and idempotent (only sets when the value
differs).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): make the ops password hash idempotent

users.yml hashed ops_password with password_hash('sha512') and no salt, so
a fresh random salt was drawn every run: the hash never matched /etc/shadow
and the user module rewrote it reporting 'changed' on every converge (a
clean converge was never a true green signal, on both clusters). Derive a
stable salt from a one-way sha256 of the password so the hash is
deterministic and idempotent, while still reconciling an out-of-band
password change. The salt in /etc/shadow is public and leaks nothing.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): make live fabric/hosts/timezone reproducible from the repo

Three drift gaps where the running fleet did not match the committed IaC.
networkd_enabled was committed for sietch only, so spice's fabric VLANs
came up from an ad-hoc -e and a reprovisioned node would not bring them up
from the repo alone; commit it true for spice (networkd_replace_ifupdown
false still fences the 1G WAN). hosts.j2 mapped every node name to bond_ip
(the 1G WAN) rather than the fabric ceph_service_ip, so in-cluster name
resolution pointed off the fabric. And timezone: UTC was declared but never
applied, leaving nodes on the image default (Europe/Berlin); add a
community.general.timezone task to baseline/system.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): scope OSD/RGW readiness waits to the in-play host set

join.yml is --limit-aware (it intersects ansible_play_hosts), but the OSD
provisioning wait computed EXPECTED from all ceph_nodes and the RGW wait
blocked until running rgw >= all ceph_nodes, so a subset (--limit) deploy
joined the subset then deadlocked at both gates. Base both counts on
groups['ceph_nodes'] intersect ansible_play_hosts, matching join.yml. On a
full deploy the intersection is all nodes, so behavior is unchanged; the
rgw-spec placement stays pinned to all ceph_nodes (only the readiness gate
is play-scoped, so a --limit re-run cannot shrink the spec).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): scope the nftables ruleset to table inet filter

nftables.conf.j2 did a global `flush ruleset`, and the no-op guard diffed
the whole `nft list ruleset`. On a ceph node (cephadm runs podman) that
wiped the podman/docker/fail2ban tables on every apply, and made the
desired-vs-live diff never match (podman tables are live but absent in the
throwaway netns), so the firewall reloaded on every converge - each one
flushing podman again. Replace only table inet filter (add, delete,
re-add), compare only that table in the guard, and reload via ExecReload
(nft -f, no global flush) instead of restart (whose ExecStop flushes the
whole ruleset). The fabric/ceph rules themselves are unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): deploy MGR at a capped count, not one per host

placement.yml applied mgr to every active host (47 mgr daemons on spice:
1 active + 46 standby), which is wasteful and non-standard - MON already
uses a bounded ceph_mon set. Deploy count:{{ ceph_mgr_count }} (default 3:
1 active + 2 standby) instead; cephadm schedules them and caps at the host
count on small clusters like sietch.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): mark the RGW EC data pool as bulk

The EC data pool is pre-seeded to pg_num 4096 (sized for spice's 658 HDD
OSDs), but without the bulk flag the pg_autoscaler can walk it back toward 1
during the first fill, causing PG splitting under load. Set bulk=true so the
autoscaler targets a full-capacity pg_num and treats the pre-seed as a
floor. Idempotent: only set when not already true.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): harden the ceph_destroy SSD-model match

The SSD OSD PV-remove and partition-wipe loops used
grep -i "$ssd_model_pattern", which interpreted the pattern as a regex and,
on an empty pattern, matched every disk - the worst-case foot-gun in a
destroy path (it would target all disks). Use grep -iF (fixed string, no
regex) and skip the loop entirely when the pattern is empty. Kept -F
without -w, since -w would fail to match underscore-containing model
strings like Micron_5100_MTF... Shared role, so it hardens sietch too.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* chore(ceph): drop the dead s3_key_drift notice; reconcile harden header

The "S3 key drift notice" debug task was gated on s3_key_drift, a variable
nothing ever sets (the drift detection it stood in for was never built), so
the branch could never fire - remove it. And update the harden.yml header:
the security-critical sshd lockdown (root key-only, no password auth, locked
root) now runs early in baseline, not gated behind this last stage; harden
adds the firewall and the remaining sshd hardening on top.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): accurate change reporting in tuning roles (T4)

CRUSH tunables and per-device HDD/SSD sysfs writes ran a mutating
command every converge under changed_when:false, so a run reported ok
even when it changed state and --check misled register consumers.
Switch to get-then-set: read the current value first, only write when
it differs, and emit CHANGED so changed_when reflects reality. Same
settings applied; only change-detection becomes truthful. Also drops
two pre-existing em-dashes to keep the file plain ASCII.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* refactor(ceph): declarative lvg/lvol for OSD block.db setup

Replace hand-rolled pvcreate/vgcreate/lvcreate shell in ceph_deploy
lvm-setup with community.general.lvg/lvol (already pinned, and the
idiom provision_host/disks.yml already uses). Same VG/LV names, sizes,
fixed-vs-100%FREE split, per-shape loops, and device targets on both
shapes. Kept read-only asserts/verify/show tasks as-is.

Preserved the wipefs -af signature clear ahead of PV creation: lvg does
pvcreate -f but not --yes, so it will not wipe a stale foreign signature
on a reused partition. wipefs stays gated on the VG being absent so a
live PV is never touched. Kept a minimal shell to compute the spice
ssd-osd size (vg0 free - reserve) since lvol has no free-minus-N
primitive, plus an existence guard so re-runs don't recompute a bad
size or trigger an unforced lvol shrink.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* chore(ceph): ASCII-clean the lvm-setup header comment

Convert the pre-existing arrows and em-dashes in the header (left untouched
by the lvg/lvol refactor) to plain ASCII.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): never let lvol shrink a live block.db LV (data safety)

The lvg/lvol refactor let community.general.lvol reconcile an existing
fixed-size db-slot to its size, and lvol defaults to shrink:true - so a
db-slot that ever drifted LARGER would be TRUNCATED, corrupting a live
BlueStore block.db and taking out the OSD. The old shell skipped existing
LVs entirely, so it could never do this. Set module_defaults shrink:false
on both lvm-setup blocks: lvol still creates and may grow, but never
truncates an existing LV (it leaves a larger one alone). Also restore
opts:-Wy so (re)created LVs wipe stale signatures - the destroy->recreate
recovery-path freshness the explicit -Wy gave.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): RGW self-signed cert generation aborted on first converge

The openssl -subj and -addext args used backslash-newline continuations
inside their quoted strings, so shell line-folding kept the next line's
indentation in the value; countryName rendered as "DE    " (over the
2-char max) and openssl exited 1 under set -euo pipefail, leaving no S3
endpoint on the first spice deploy. Precompute the subject and SAN as
single-line facts (whitespace-controlled Jinja for the SAN) and pass
them quoted, so each openssl arg is one flat string.

See ansible/ceph/roles/ceph_deploy/tasks/rgw.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): green the ansible-lint CI gate on HEAD

Bare ansible-lint (what mise run lint and the required CI check run)
exited 2 on two deliberate patterns: run_once on the fleet-wide
placement assert, and an intentional no-pipefail shell in the ice DDP
Safe-Mode probe (pipefail there would mask grep's match). Waive both
with inline noqa so the reasoned patterns stay and the gate passes.

See ansible/ceph/deploy-ceph.yml and roles/baseline/tasks/nic-firmware.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): pin RGW pools to device-class CRUSH rules (SSD fast tier)

The replicated_ssd/replicated_hdd rules were created but bound to no pool,
so every replicated pool fell to the default class-agnostic rule and the
omap-heavy bucket index landed on HDD (~98.5% of weight) while spice's
47-OSD NVMe ssd-osd tier sat idle. Pin the index, RGW metadata, and mgr
pools to the SSD rule and the bulk non-EC pool to HDD; force the OSD
device class at creation (ssd-osd -> ssd, HDD -> hdd) so the rules bind
deterministically, and assert the placement in verify.yml. Pinning is
opt-in per cluster (empty default) so a live cluster is never re-homed
implicitly; the pin runs after RGW readiness because the .rgw.* system
pools are created lazily by the realm/zone setup and the daemon.

See ansible/ceph/roles/ceph_deploy/{templates/osd-spec.yml.j2,tasks/rgw.yml,tasks/verify.yml}.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): apply RGW pool PG/durability pins on the first converge

The pool list was captured once at the top of the RGW play, before any
pool was created, so every "pool in list" gate skipped on a fresh run 1:
the data/index/non-EC PG pre-sizing deferred to a second converge, and
the .rgw.* system pools (created lazily by the daemon) never got their
size/min_size pin at all, sitting at the inherited 2/1 single-replica
until someone ran the play twice. Re-query the pool list once the
explicit pools exist, and move the system-pool PG + size/min_size pins
below the RGW readiness wait where those pools are real. Parameterize the
bootstrap osd_pool_default_size per cluster (spice 3/2, role default keeps
sietch at 2/1) so an unpinned future pool is not born single-replica, and
assert final size/min_size in verify.yml so a miss fails the play loudly.

See ansible/ceph/roles/ceph_deploy/tasks/{rgw.yml,verify.yml}.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): real ok-to-stop gate before a live-node reimage

Reimaging a spice node zeroes the OS NVMe, and the 14 HDD OSD block.db
LVs plus the ssd-osd data LV all live on that NVMe, so a reimage destroys
every OSD on the node; the old guard only checked a hand-typed
i_know_its_dead flag and the prepare-os-disks comment wrongly claimed a
reimage keeps OSD data. Replace the stub with a mon-delegated check that
queries the node's OSD ids, refuses on HEALTH_ERR or a failed
ceph osd ok-to-stop, and sets noout on those OSDs for the reimage window.
The gate is default-on: auto mode enforces whenever a live cluster with
this node's OSDs is reachable and no-ops for the initial bootstrap;
strict fails closed if it cannot verify; permissive is the explicit
escape hatch.

See ansible/ceph/roles/reprovision_hetzner/tasks/ceph_safety.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): jumbo MTU 9000 on the VLAN 122 cluster/replication network

Jumbo pays off on the OSD replication path (a closed, homogeneous fabric)
but is a partial-blackhole risk on the RGW-facing public network, which
serves heterogeneous 1500-byte S3 clients and the ~1400 NetBird overlay,
so raise MTU on VLAN 122 (Ceph private) only and leave 120/124 at 1500. A
VLAN opts into jumbo with a per-entry mtu; the bond ceiling and both 25G
members auto-raise to the largest child MTU so a 1500 parent or slave
cannot silently cap the jumbo frames. verify.yml asserts the applied MTU
on the bond, members, and each VLAN (local, no switch dependency). Flat
clusters resolve to 1500 and emit no MTU, so sietch is unchanged.

Paired with a QFX-side change (leaf server-LAG + VLAN 122 IRB >= 9000) and
a ping -M do -s 8972 matrix before the cluster network is relied on.

See ansible/ceph/roles/networkd.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* chore(dns): commit prod DNS provider lock file

The prod DNS stack was the only DNS stack without a committed
.terraform.lock.hcl, so cloudflare ~> 5.0 would resolve unpinned at the
first prod apply of the stack fronting the S3 endpoint. Pin it to 5.21.0,
matching the staging DNS stack, with linux_amd64 + darwin_arm64 hashes.

See tf/deployment/prod/global/dns/.terraform.lock.hcl.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): opt-in alertmanager receiver so alerts reach a human

cephadm deploys alertmanager with no receiver, so every Prometheus alert
rule (OSD down, host down, PG degraded, near-full) fires into the default
null route and nobody is notified. Add ceph_alertmanager_webhook_urls: a
non-empty list routes all alerts to those webhooks via cephadm
user_data.default_webhook_urls. The monitoring re-spec now applies when
either fabric bind or alerting is configured (independent gates), and
verify.yml reports the delivery status, warning loudly when none is set.
Empty by default, so flat clusters (sietch) are unchanged; the spice
destination is a deferred operator decision (TODO in defaults).

See ansible/ceph/roles/ceph_deploy/tasks/monitoring.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(fabric): jumbo MTU on the leaf server bonds + VLAN 122 IRB

Completes the host-side VLAN 122 jumbo change on the QFX leaf: the server
LACP bonds (ae1..aeN) get the 9216 L2 jumbo MTU that the spine uplink ae0
already carries (802.3ad members inherit it, so no per-member mtu), and the
private/cluster IRB gateway (VLAN 122) gets an L3 family-inet MTU of 9000 to
match the Ceph hosts. Public (120) and host-mgmt (124) IRBs stay at 1500 by
design, so jumbo is confined to the closed OSD-replication path.

See tf/shared/modules/cluster-fabric.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ci): pin + SHA256-verify NetBird install, drop curl|sh

The netbird-connect action piped pkgs.netbird.io/install.sh straight into
sh on the runner that holds the prod-write 1Password token and overlay
access to the live nodes, so a compromised installer would run as root
there. Download a pinned release tarball (v0.74.4) from the GitHub release
and verify its SHA256 against an in-repo pin before unpacking; fail closed
on any mismatch. A version input plus a refresh comment keep the pin
maintainable.

See .github/actions/netbird-connect/action.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ci): close the prod apply gate (reviewed-plan binding + fail-closed reviewers)

The prod approval gate could fail open and the apply was not bound to the
reviewed diff. Manage the four gate Environments as code (new
_meta/github stack: github_repository_environment + required reviewers,
protected-branches-only, a validation that rejects a reviewerless gate),
and make the discover job fail closed by asserting every emitted
partition-region Environment already carries a required reviewer. Bind
apply to the plan: the plan job writes -out to an absolute path and
uploads it, the gated apply downloads that exact file and applies it with
no re-plan and no -auto-approve, so a stale plan fails closed. Scope a
bare workflow_dispatch to plan-only and a toggled dispatch to just the
partition it touches; add fail-fast and timeout-minutes across the jobs.

The _meta/github stack is bootstrap-applied out of band and needs real
reviewer IDs plus a GH_ENV_ADMIN_TOKEN secret before the gate passes.

See .github/workflows/infra.yml and tf/deployment/_meta/github.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): blast-radius control on the day-2 converge plays

The converge chain (baseline, tune-os, tune-hardware, tune-ceph, harden)
ran against all nodes at once with no canary and no stop-on-failure, the
inverse of the serial:1/max_fail:0 destructive plays. Batch each converge
play (serial, default 10%) with max_fail_percentage:0, and gate every
batch on cluster health via a shared pre/post ceph-health checkpoint that
halts the roll on HEALTH_ERR. The gate is a no-op until a cluster is live,
so the bootstrap ordering (baseline -> tune -> deploy) is unaffected;
deploy-ceph stays big-bang. Tune the batch size with ceph_converge_serial.

See ansible/ceph/tasks/health_gate.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): freeze load-bearing host packages instead of auto-upgrading

We run no unattended-upgrades: an apt upgrade that moved cephadm's
container runtime (podman + ecosystem) or the chrony time source that mon
quorum depends on out from under a live cluster is the uncoordinated
change we avoid. Hold those packages at their installed version (dpkg
selection) so apt upgrade skips them; upgrading is then deliberate and
health-gated. Diagnostic/ops tools are left unheld. Runs last in baseline
so nothing is held before it exists; tune via baseline_held_packages.

See ansible/ceph/roles/baseline/tasks/hold-packages.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): THP=madvise and stop the GRUB cmdline clobber

Transparent Huge Pages sat at the kernel default 'always', which bloats
BlueStore/tcmalloc RSS and drives allocation-stall latency across the OSD
fleet; set it to madvise via a oneshot unit ordered Before=ceph-osd.target
(mirroring ceph-cpu-governor.service) plus a live sysfs write, on by
default for every ceph cluster. Separately, the processor.max_cstate GRUB
task rewrote GRUB_CMDLINE_LINUX_DEFAULT wholesale (dropping existing
tokens, injecting quiet) - a latent footgun behind the default-off
governor flag; replace it with a /etc/default/grub.d drop-in that only
appends the cstate token.

See ansible/ceph/roles/hardware_tuning.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): host security sysctls + persistent journald

Add a host kernel/network hardening sysctl set (syncookies, ignore/deny
ICMP redirects + source routing, kptr/dmesg restrict, log martians) to the
existing os_tuning sysctl drop-in. rp_filter is set to LOOSE (2), not
strict (1): these nodes are multi-homed (1G WAN default route + 25G fabric
VLAN sub-interfaces), so strict reverse-path filtering would blackhole
asymmetric fabric traffic. Separately make journald persistent
(Storage=persistent, bounded SystemMaxUse/RuntimeMaxUse) so logs survive
the pipeline's own reboots on these headless nodes.

See ansible/ceph/roles/os_tuning.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): health-gated cephadm upgrade playbook + rollback runbook

No upgrade automation existed for the pinned 20.2.x tentacle train. Add a
deliberate, manually-invoked upgrade-ceph.yml: preflight-gates on HEALTH_OK
(with an explicit allow-WARN toggle), no in-progress upgrade, and all OSDs
up+in; records and prints the rollback image before starting; refuses to
run without an explicit target (ceph_upgrade_target_image/_version, no
default); drives ceph orch upgrade with a bounded poll and fails loudly on
a stall; asserts health + version convergence after. Not wired into
site.yml. Rollback procedure lives in the play header and
docs/runbooks/upgrade-ceph.md.

See ansible/ceph/upgrade-ceph.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): scheduled cluster-state backups with a systemd timer

Config/topology state was captured only manually and controller-local. Add
a ceph_backup role that installs a capture script + daily systemd timer on
the bootstrap node: each run tars fsid, config dump, mon/osd/crush maps
(raw + decoded), osd tree, orch ls/host ls, and the RGW realm/zonegroup/
zone into a root-only 0700 dir, pruned by retention. No secret keyrings are
dumped. An offsite target (rsync or s3://) is a var left empty for the
operator. Invoke via backup-ceph.yml (mise run backup-timer); not wired
into converge. Opt out with ceph_backup_enabled=false.

See ansible/ceph/roles/ceph_backup and docs/runbooks/backup-restore.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): pin mgr dashboard/prometheus bind to fabric; pin EC min_size

The mgr dashboard (8443) and prometheus module (9283) run inside the active
mgr and bound 0.0.0.0, closed only by nftables. Pin each mgr instance's own
localized server_addr (mgr/<module>/<mgr_id>/server_addr, the key the
modules read via get_localized_module_option) to that host's fabric IP,
discovering mgr ids at runtime so it is failover-safe; a global server_addr
would leave the dashboard unbindable after failover. Opt-in on
ceph_bind_networks; takes effect on the next mgr cycle (nftables holds the
WAN closed meanwhile). Separately pin the EC data pool min_size explicitly
to k+1 (derived from ceph_rgw_ec_k) so the write-availability floor is
documented and cannot drift, and record the single-site host-failure-domain
DR ceiling in docs/capacity-planning.md.

See ansible/ceph/roles/ceph_deploy/tasks/monitoring.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* chore(ceph): freeze Galaxy collection versions to a major

The collection requirements floated on >= lower bounds, so a new
community.general major could be pulled mid-campaign. Pin compatible-release
ranges to the currently-resolving majors (ansible.posix >=2,<3;
community.general >=12,<13) so a 47-node campaign cannot cross a major
between runs. No lockfile mechanism exists in the repo, so the ranges live
in requirements.yml.

See ansible/ceph/requirements.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* test(ceph): molecule scenario for the fabric-only firewall branch

The only molecule scenario tested the default open-firewall branch
(*_any_source: true), the opposite of spice's production posture, with no
idempotence check. Add a fabric-only scenario that flips the closed branch
(rgw/ssh any_source: false, trusted networks = fabric + NetBird) and
asserts RGW/SSH are NOT accepted from any source, are restricted to the
fabric/overlay, and that a second render is idempotent. Mirrors the default
scenario's render-and-grep idiom; the existing scenario is untouched.

See ansible/ceph/roles/security/molecule/fabric-only.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* ci(dns): drift-check the prod S3 roster against the ceph node list

The 47 S3 RGW A-records were hand-copied with no link to the ceph roster,
so add/replace/remove of a node silently blackholed the endpoint or dropped
capacity. A terragrunt dependency is not viable here (no dependency idiom in
the repo, and the ceph discovery output carries only bond_ip, never the
fabric IP), so add a stdlib-only, credential-free check that reconstructs
the expected fabric-IP roster from clusters.auto.tfvars (in-service names) +
spice-hosts.yaml (host_index) and asserts the apex and wildcard A-records
match exactly. A path-scoped CI gate fails the PR on drift before the DNS
apply; also runnable via mise run tf:check-dns-roster.

See tf/scripts/check-s3-dns-roster.py.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): inotify headroom for podman scale; podman-compat upgrade note

cephadm-on-podman runs many containers/systemd units per node and leans on
inotify; Debian's default fs.inotify.max_user_instances=128 is a known
bottleneck at that scale, so raise instances to 512 and watches to 524288
in os_tuning. Separately, document in the upgrade runbook that the podman
dpkg-hold must be lifted (unhold -> bump to a cephadm-supported version ->
re-hold) before a cross-major Ceph upgrade; patch-train upgrades need no
podman change.

See ansible/ceph/roles/os_tuning and docs/runbooks/upgrade-ceph.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): health-checked RGW ingress VIP (haproxy + keepalived)

S3 was bare round-robin DNS across 47 RGW A-records, so a dead node
blackholed ~1/47 of new connections for the TTL. Add an opt-in cephadm
ingress service: a keepalived floating VIP (spice 10.40.20.250) with
haproxy health-checking the RGW backends and dropping a dead one from
rotation. haproxy terminates TLS on the VIP with the existing self-signed
RGW cert (its SAN now carries the VIP) and re-encrypts to beast: L4
passthrough for RGW needs IngressSpec.use_tcp_mode_over_rgw, which is
absent on the pinned 20.2.2 (source-verified; applying it TypeErrors), so
terminate is the only health-checked mode available. Applies after RGW
readiness; asserts its preconditions (VIP set, ceph_bind_networks non-empty
so beast binds a per-node IP and does not collide with the VIP on :443,
ceph_rgw_ssl true). Opt-in per cluster; sietch unaffected.

See ansible/ceph/roles/ceph_deploy/templates/rgw-ingress-spec.yaml.j2.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(dns): point the prod S3 endpoint at the ingress VIP

Collapse the s3.prod.fsn1.htz.futo.cloud apex + wildcard from 47 round-robin
A-records to the single health-checked ingress VIP 10.40.20.250, and switch
the drift-check to assert both records equal that VIP (the ceph
ceph_rgw_ingress_vip is the source of truth). Apply ordering matters and is
noted in the tfvars: the ingress must be live before this DNS cutover, or
s3 resolves to a VIP nothing answers; roll back in reverse.

See tf/deployment/prod/global/dns/records.auto.tfvars.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* chore(ceph): fix stale staging copy-paste comments in the prod stack; fmt

The prod ceph stack carried the staging stack's comments verbatim: versions.tf
and secrets.tf described creating SIETCH_CEPH_* items in yucca_tf_staging via
OP_TF_YUCCA_STAGING_ENV_WRITE, and a variable example said (sietch, ...). This
stack creates SPICE_CEPH_* in yucca_tf_prod via OP_TF_YUCCA_PROD_ENV_WRITE;
correct the comments to match. Legitimate cross-refs (the mirrors-staging/talos
provenance, the partition-slug enumeration) are left as-is. Also tofu fmt the
drifted clusters.auto.tfvars (whitespace only).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* ci(infra): scope reviewer fail-closed gate to apply events; prettier-format workflows

* ci(infra): pin 1Password CLI version to avoid flaky latest resolution

* feat(dns): resolve prod S3 apex+wildcard to the node fabric-IP roster

* ci(infra): park the Environments approval gate (disable fail-closed step + stub _meta/github)

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 12:01:10 -07:00
..
2026-07-13 13:15:50 +00:00

yucca/tf

Terraform/OpenTofu authority for cluster identity, 1P secret items, and the inventory artifacts Ansible consumes. Multi-partition / multi-region via terragrunt.

The partition / region model

Everything is keyed on a single first-class topology:

partition → region → { exactly one K8s cluster, one-or-more Ceph clusters }

  • Partition = prod | staging | dev (formerly env).
  • Region = a site: htz-fsn1, austin, local (formerly site/datacenter). Plus a reserved global pseudo-region for partition-wide stacks (DNS, the account-wide NetBird layer) — global is not a physical site, so its role/FQDN metadata is null.
  • The human form is partition@region (prod@htz-fsn1); the canonical slug is partition-region (prod-htz-fsn1), which derives names on every surface.

Every stack sits at three path segments: deployment/<partition>/<region>/<stack>. The account-wide prod layer is prod/global/netbird (a stack under the global region), so there is no two-segment special case.

Layout

tf/
├── .env                              ← op:// references (committed; no literal secrets)
├── op-run.sh                         ← op-run wrapper used by the mise tf:* tasks
├── shared/
│   └── modules/
│       ├── ceph-cluster/             ← per-cluster ceph orchestration module
│       │   ├── main.tf, variables.tf, outputs.tf, rendering.tf
│       │   ├── wordlist.txt          ← 923 words for auto-picked hostnames
│       │   └── templates/
│       │       ├── inventory.ini.tftpl
│       │       ├── inventory-destroy.ini.tftpl
│       │       ├── inventory-provision-debian-live.ini.tftpl
│       │       └── secrets.yml.tpl.tftpl
│       └── talos-baremetal/         ← Talos on bare-metal nodes already in maintenance mode
│           ├── main.tf, variables.tf, outputs.tf
│           └── firewall.tf              ← Talos host ingress firewall (default-deny + allow-lists)
└── deployment/
    ├── terragrunt.hcl                ← root: state backend + partition/region/stack from path
    ├── staging/
    │   ├── austin/
    │   │   ├── region.hcl            ← role + FQDN parts for staging@austin
    │   │   ├── ceph/                 ← sietch ceph cluster (clusters.auto.tfvars)
    │   │   └── talos/                ← bare-metal Talos cluster (3× CP, Cilium CNI)
    │   └── global/
    │       ├── region.hcl            ← role=null (pseudo-region)
    │       ├── dns/                  ← Cloudflare records (futo.cloud)
    │       └── netbird/              ← NetBird Cloud access control (staging)
    ├── dev/
    │   ├── local/
    │   │   ├── region.hcl
    │   │   └── talos/                ← Talos VMs on the sietch hypervisors
    │   └── global/
    │       ├── region.hcl
    │       └── dns/
    └── prod/
        ├── htz-fsn1/
        │   ├── region.hcl           ← role + FQDN parts; site_id=40
        │   ├── mgmt-hosts.yaml      ← mgmt-host roster (region root; fabric + render read it)
        │   ├── fabric/              ← Junos switch fabric + NetBox + mgmt reprovision
        │   └── netbird/             ← htz-fsn1 site NetBird layer (routed htz-fsn1 network)
        └── global/
            ├── region.hcl
            ├── terragrunt.hcl       ← account-wide prod NetBird (two-segment region-root stack)
            └── netbird.tf

Add a region: create deployment/<partition>/<region>/region.hcl + stacks under it. Add a stack: a new sibling dir under a region (<region>/monitoring/ …). NetBird Cloud access control lives in staging/global/netbird/, and for prod is layered: prod/global/ (account-wide) above per-region prod/<region>/netbird/ (e.g. prod/htz-fsn1/netbird/). See "The netbird-env module" below.

The dns stack manages infrastructure names in the futo.cloud Cloudflare zone (today: the Sietch RGW S3 endpoint + virtual-hosted wildcard). Records are declarative in records.auto.tfvars; the API token resolves from op://yucca_tf_manual/CLOUDFLARE_API_TOKEN via tf/.env.

The talos stack is documented in ansible/talos/README.md and ansible/talos/docs/runbooks/cluster-bring-up.md (the TF + Ansible flow is interleaved — TF renders the inventory Ansible consumes, then bootstraps the VMs Ansible created).

Conventions

Partition, region, and stack are derived from the directory path

deployment/terragrunt.hcl parses the child's relative path: partition = segs[0], region = segs[1], stack = join(segs[2:]), and slug = "<partition>-<region>".

deployment/staging/austin/ceph     → partition=staging, region=austin,   stack=ceph
deployment/staging/global/dns      → partition=staging, region=global,   stack=dns
deployment/prod/htz-fsn1/fabric    → partition=prod,    region=htz-fsn1, stack=fabric
deployment/prod/htz-fsn1/netbird   → partition=prod,    region=htz-fsn1, stack=netbird
deployment/prod/global/netbird     → partition=prod,    region=global,   stack=netbird

The state backend key is derived from these: yucca/${partition}/${region}/${stack}/terraform.tfstate in the shared yucca-tf-state S3 bucket. (The legacy ceph/ project prefix is dropped — talos/dns/netbird/fabric all share the bucket now.) Every stack is three segments — the account-wide prod layer is prod/global/netbird (not a bare prod/global), so region.hcl at prod/global/ is found by find_in_parent_folders from the stack dir, exactly like staging/global. The n==2 branch in terragrunt.hcl is now dead and can be removed.

Per-region metadata + role (region.hcl)

Each region dir carries a deployment/<partition>/<region>/region.hcl holding role (primary | secondary), site_id, datacenter, provider_code, and domain. The root terragrunt finds it via find_in_parent_folders("region.hcl", "") (with a not-found guard) and merges its locals into every stack's inputs — so every stack in a region inherits the metadata without per-tfvars duplication. global pseudo-regions set role = null (and null FQDN parts). Each stack declares matching variable blocks with null defaults.

role encodes the product invariant: when a partition spans multiple regions, yucca-api + the database run only in the primary region; secondary regions run the storage-local subset. It is authoritative here in TF state (discovery.role) and consumed downstream (Flux role components, yuctl).

The op run --env-file=tf/.env -- pattern

tf/.env holds 1Password op:// references — not literal secrets:

export OP_SERVICE_ACCOUNT_TOKEN="op://yucca_tf_dev/yucca_futo_1pass_superuser_service_account/password"

Wrap every terragrunt invocation with op run --env-file=tf/.env -- (the mise tf:* tasks do this automatically). The op CLI resolves the op:// reference and injects the actual token as OP_SERVICE_ACCOUNT_TOKEN into the child process's environment. The 1P Terraform provider picks it up from the env var and authenticates.

The same pattern is used in immich-app/devtools and is the Futo-wide convention for TF secret injection.

Committed .env is safe because it's just pointers

Yucca's root .gitignore normally excludes .env files — we add an explicit !tf/.env exception. This file contains only op:// URIs; no secret ever transits the repo. It's a committed manifest of "which 1P items this TF depends on."

Stack override via TF_STACK_DIR

The default mise run tf:* tasks target tf/deployment/staging/austin/ceph. Point them at another stack via the TF_STACK_DIR env var:

TF_STACK_DIR=tf/deployment/staging/austin/ceph  mise run tf:plan
TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply

Running TF

One-shot (preferred for now)

mise run tf:init      # first time in a stack
mise run tf:plan      # dry run
mise run tf:apply     # render artifacts + (future) create 1P items

These wrap: op run --env-file=tf/.env -- terragrunt --working-dir <stack> <cmd>.

In CI (staging stacks)

.github/workflows/infra.yml runs the staging stacks from GitHub Actions:

  • Plan on every PR touching tf/**; apply on merge to main, gated behind the staging-infra Environment (required reviewers).
  • Applies staging/austin/talos (cluster + Flux + secrets) then staging/global/dns.
  • The Talos stack reaches the 10.10.10.0/24 nodes by joining the tailnet (tailscale/github-action, --accept-routes) — the cluster firewall already trusts the Tailscale CIDRs. The DNS stack is pure Cloudflare API, no tailnet.
  • Secrets come from the same op run --env-file=tf/.env path; CI just supplies the per-env 1P service-account token (the rest resolves from 1P). The token is injected as OP_SERVICE_ACCOUNT_TOKEN from the environment-specific secret — OP_TF_YUCCA_STAGING_ENV here (dev/prod workflows use OP_TF_YUCCA_DEV_ENV / OP_TF_YUCCA_PROD_ENV) — replacing a shared superuser SA with a scoped one.

Prerequisites (out-of-band): repo secret OP_TF_YUCCA_STAGING_ENV — a staging 1P service account mirroring the dev/prod ones, i.e. granted shared_tf, shared_tf_staging, yucca_tf (read) + yucca_tf_staging (read/write for the JWT-keypair item). Note: a copy of the dev SA token won't work — it can't read yucca_tf_staging. Also: TS_OAUTH_CLIENT_ID, TS_OAUTH_SECRET; a Tailscale subnet router advertising 10.10.10.0/24 with tag:project-yucca approved for it; and the staging-infra Environment with required reviewers.

State backend

Remote: shared yucca-tf-state S3 bucket at OVH Paris (https://s3.eu-west-par.io.cloud.ovh.net/). Key path: yucca/${partition}/${region}/${stack}/terraform.tfstate (region-root stacks: yucca/${partition}/${region}/...). All yucca stacks live under the yucca/ prefix in the shared bucket.

Credentials are AWS-compatible env vars (AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY), injected via op run --env-file=tf/.env from the TF_STATE_S3_* items in the yucca_tf vault. OVH-specific config (skip AWS validation, path-style URLs) is in deployment/terragrunt.hcl.

State locking is not enabled. OVH has no DynamoDB equivalent; OpenTofu's use_lockfile = true option would handle single-bucket locking but expects the lockfile object to already exist — terragrunt init against a fresh backend fails with 404 before it can create one. Enable it once the concern is concurrent applies (multiple operators working the same stack simultaneously). Single operator today → low risk.

Discovery outputs contract

Every stack emits a single non-sensitive top-level discovery output (plus discovery_schema_version). It is the machine-readable description of the topology that yuctl consumes — it reads the state object straight from S3 and parses .outputs.discovery.value, so no checkout, init, or provider is needed.

Secrets are always op:// references, never values. The in-stack sensitive kubeconfig/talosconfig outputs stay for in-stack use; discovery only carries the reference (op://<vault>/<title>/password).

Common envelope (every stack):

{
  schema_version, partition, region, slug, role, stack, stack_type,
  region_meta: { site_id, datacenter, provider_code, domain }
}

Per stack_type payload:

stack_type stacks payload key + fields
region-k8s */talos kubernetes: { cluster_name, api_endpoint, operator_endpoint, cp_node_ips, kubeconfig_ref, talosconfig_ref }
ceph */ceph ceph_clusters: { <name> => { cluster_name, fqdn, rgw_s3_endpoint, health_cred_ref, s3_admin_cred_refs, secret_item_titles, bootstrap_host } }
dns */dns dns: { provider, zone, record_fqdns, api_token_ref }
netbird */netbird, prod/global netbird: { name_prefix, vault, group_ids, policy_ids, network_ids, setup_key_item_titles }
fabric prod/htz-fsn1/fabric fabric: { site_id, kube_cidr, mgmt_cidr, cluster_cidrs }

The talos stacks persist their kube/talosconfig into 1Password (onepassword_item, mirroring the JWT-keypair / netbird-setup-key pattern) so the *_ref fields resolve — staging writes YUCCA_STAGING_{KUBE,TALOS}CONFIG; dev writes YUCCA_DEV_<CLUSTER>_{KUBE,TALOS}CONFIG.

Migrating an existing stack to the new layout

The dir moves + terragrunt rewrite change the state key (ceph/<env>/... → yucca/<partition>/<region>/...), so a live stack's remote state must be re-pointed — reversibly, no destroy/recreate. Per stack (single-operator window; no state locking):

  1. Pre-flight terragrunt plan on the OLD layout → confirm no-op; back up the old state object.
  2. Land the terragrunt.hcl rewrite + dir moves + renames (no apply yet).
  3. Server-side S3 copy old key → new yucca/... key (old object stays as a rollback anchor).
  4. Delete the stale generated backend.tf; terragrunt init in the new dir (decline the state-copy prompt; -reconfigure if needed).
  5. terragrunt plan → must be no-op (the gate; a non-empty plan = a rename/source-path mismatch, not a state problem — STOP + roll back).
  6. terragrunt output discovery → confirm the contract resolves.
  7. After all stacks are green, aws s3 rm the legacy ceph/* objects.

Order: staging first, then prod (global → region → fabric — prod/global before prod/htz-fsn1/netbird, which has a real terragrunt dependency on it). dev is dir-move only (no remote state). The NetBird plans must show no resource renames — the rendered names are byte-identical across the rename.

The ceph-cluster module

Declarative input in clusters.auto.tfvars:

clusters = {
  sietch = {
    domain            = "staging.austin.int.futo.cloud"
    partition         = "staging"
    region            = "austin"
    provider_code     = "int"
    role_in_hostname  = "ceph"
    ansible_ssh_user  = "ansible-iac"
    ansible_ssh_key   = "~/.ssh/id_ed25519_sietch"
    vault             = "yucca_tf_staging"
    provision_profile = "debian-live"   # null for Hetzner-installimage clusters
    hosts = [
      { name = "laurel", bond_ip = "10.10.10.90", bootstrap = true },
      { name = "lawson", bond_ip = "10.10.10.91" },
      { name = "samara", bond_ip = "10.10.10.92" },
    ]
  }
}

On apply, the module:

  1. Picks wordlist names for hosts[].name == null (stable across applies; seeded per-cluster; operator-declared names excluded from the pool to prevent collisions).
  2. Renders inventory.ini (normal ops), inventory-destroy.ini (explicit destroy flag), secrets.yml.tpl (op:// references to yucca_tf_staging/<CLUSTER>_CEPH_*_PASSWORD/password). Optionally renders inventory-provision.ini when provision_profile != null.
  3. (Not yet TF-managed) onepassword_item resources for cluster secrets are dormant — items are created via op CLI today and read by Ansible at play time. See ansible/ceph/docs/secrets.md for the re-enable plan.

The talos-baremetal module (staging/austin/talos)

Brings up Talos on bare-metal nodes already running in maintenance mode at known addresses — no Ansible, no hypervisors, no VLANs (the earlier VM-oriented talos-cluster module was removed unused). It dials each node's maintenance IP, applies machine config (which installs to disk + reboots), bootstraps one CP, then emits kube/talosconfig and gates on cluster health.

Declarative input in deployment/staging/austin/talos/clusters.auto.tfvars:

clusters = {
  yucca-staging = {
    talos_version      = "1.13.4"
    kubernetes_version = "v1.36.1"
    install_disk       = "/dev/sda"        # WIPED — the 240GB DELLBOSS; NVMe left raw
    cluster_vip        = "10.10.10.15"     # L2 VIP, etcd-elected across CPs
    gateway            = "10.10.10.1"
    subnet_cidr        = "10.10.10.0/24"
    cni                = "cilium"           # cni:none in Talos + Cilium via Helm
    disable_kube_proxy = true               # Cilium kube-proxy replacement (KubePrism)
    cilium_version     = "1.19.5"
    hubble             = true
    bond = { interfaces = ["eno1np0", "eno2np1"], mode = "active-backup" } # flip to 802.3ad after the switch is LACP'd
    nodes = [
      { name = "staging-cp1", address = "10.10.10.47" },
      { name = "staging-cp2", address = "10.10.10.242" },
      { name = "staging-cp3", address = "10.10.10.117" },
    ]
  }
}

Notes:

  • Static IP = maintenance IP. Each node's address is pinned as the static IP on bond0, so TF stays reachable across the install reboot.
  • bond comes up active-backup (no switch config needed). Migrate to 802.3ad later, node-by-node, after converting the switch ports to LACP port-channels — LACP needs both ends configured at once, so a big-bang flip drops connectivity until both sides agree.
  • Ingress firewall (firewall.tf): default-deny + per-service allow-lists scoped to the subnet (+ pod CIDR on kubelet). apid + apiserver also trust the Tailscale ranges (trust_tailscale). ⚠️ The host running tf apply must have a source IP inside an allowed range or apid (50000) is blocked and bootstrap hangs — add operator/jump subnets to trusted_cidrs.
  • CNI is installed in the same apply. With cni:none the nodes are NotReady until Cilium lands, so the module's health gate runs skip_kubernetes_checks; helm.tf installs Cilium, then a second (full) health gate enforces Ready.
  • One cluster per stack. The helm provider binds to a single cluster (one(...)); add more clusters in their own stack.

Run it (see "Running TF" below — needs 1Password unlocked + an on-LAN apply host):

TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:init
TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:plan
TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply   # WIPES /dev/sda, installs Talos

The netbird-env module (NetBird Cloud access control)

Manages one layer's NetBird Cloud footprint: groups, access policies, device auth (setup) keys, and routed networks. One NetBird Cloud account (api.netbird.io) backs everything — the module namespaces every object <name_prefix>_<key> (all underscores) so all envs/sites coexist.

Group model

Per env (and per prod site), the baseline groups are:

group who rendered (staging / prod htz-fsn1)
ci ephemeral CI runners yucca-staging-ci / yucca-prod-htz-fsn1-ci
mgmt management nodes (configured via Ansible; also the route peers) yucca-staging-mgmt / …
talos Talos cluster nodes yucca-staging-talos / …
k8s_operator in-cluster kubernetes operator yucca-staging-k8s-operator / …

Logical keys (the tfvars map keys, e.g. ci) stay lowercase; rendered NetBird names are UPPER_SNAKE (uppercased, hyphens → underscores). CI is per-env (ci, reaching only that env's groups) — no cross-env CI plane. k8s is split into talos (the nodes) and k8s_operator (the operator identity) so they can carry different policies.

Stacks & layering

stack partition / scope state key
deployment/staging/global/netbird staging (global region) yucca/staging/global/netbird/…
deployment/prod/global/netbird prod, account-wide (cross-region) yucca/prod/global/netbird/…
deployment/prod/htz-fsn1/netbird prod, region htz-fsn1 yucca/prod/htz-fsn1/netbird/…

Staging is single-layer. Prod is layered: a global layer (reserved for account-wide / cross-region groups + policies) above per-region layers. The global layer is empty today — each region owns its own resource group and the yucca → resources policy is module-generated per layer (see below). Region groups are region-scoped (yucca-prod-<region>-<role>) so a network router's peers are unambiguously that region's mgmt nodes. A region layer can still consume a global group via a terragrunt dependency on prod/global → the module's external_groups input, but none do today. The root terragrunt derives stack from the full sub-path, so prod/htz-fsn1/netbird gets its own state key. (Now that the fabric stack lives in its own prod/htz-fsn1/fabric/ sub-stack, the region root carries no terragrunt.hcl, so prod/htz-fsn1/netbird uses a normal find_in_parent_folders include — the old direct-root-include workaround is gone.) Rendered NetBird object names and 1P item titles are unchanged by the rename (YUCCA_STAGING_*, NETBIRD_YUCCA_PROD_HTZ_FSN1_*): the env→partition / site→region swap keeps the same string values.

The yucca → yucca-tags access model

yucca members reach everything tagged as a yucca tag (NetBird object names are lowercase-kebab — e.g. yucca-prod-htz-fsn1-mgmt; the 1Password setup-key item titles stay UPPER_SNAKE, since CI/ansible/talos read them by op:// string):

  • yucca — the existing users group (people). External (looked up by its actual name yucca); never managed here.
  • yucca tags — any group flagged resource = true in a layer's groups. This marks the group yucca-reachable; it applies to peer/node groups (so a yucca member can SSH yucca-prod-htz-fsn1-mgmt, the mgmt nodes) as well as routed-subnet tags (yucca-prod-htz-fsn1-resources, which the site's netbird_network_resources are tagged into). Today every group a layer owns is flagged.

For each layer that owns ≥1 yucca tag, the netbird-env module auto-generates a <prefix>-yucca-to-resources policy (bidirectional = false) whose destinations are all of that layer's flagged groups — so flagging a group grants yucca users access to it (its peers and any tagged resources) with no policy to edit. The source is var.yucca_users_group (default yucca, looked up by name; set null to opt a layer out). bidirectional = false means yucca users only initiate — this policy never makes a tagged group a source. The union of these per-layer policies is the account-wide "yucca reaches every yucca tag we create" guarantee.

The former single shared yucca_resource tag in prod/global (one account-wide yucca → yucca_resource policy, consumed by sites via a terragrunt dependency) was retired in favour of this per-layer model — no cross-stack group reference, and the destination list is derived, not hand-maintained.

Staging additionally grants its ci group access to the existing Liberty Park infra groups (where the staging nodes live today) — those are external groups resolved by name in staging/global/netbird/main.tf.

Declarative input (netbird.auto.tfvars)

Groups, setup keys, policies and networks reference groups by logical key, never opaque NetBird IDs:

# resource = true ⇒ yucca-reachable ("yucca tag"); here every group is flagged,
# so yucca users reach all of them (SSH the nodes + the routed subnets)
groups = { ci = { resource = true }, mgmt = { resource = true },
           talos = { resource = true }, k8s_operator = { resource = true },
           resources = { resource = true } }

setup_keys = {
  ci           = { type = "reusable", ephemeral = true, auto_groups = ["ci"] }
  mgmt         = { type = "reusable", auto_groups = ["mgmt"] }
  talos        = { type = "reusable", auto_groups = ["talos"] }
  k8s_operator = { type = "reusable", auto_groups = ["k8s_operator"] }
}

policies = {
  ci-to-all = {                       # CI reaches every node group in this env
    rules = [{ name = "ci-to-all", protocol = "all"
               sources = ["ci"], destinations = ["mgmt", "talos", "k8s_operator"] }]
  }
}

NetBird is default-deny — a peer gets only the access its groups' policies grant; an empty policies map means total isolation. The yucca → resource policy is not declared here: the module generates it from every group flagged resource = true (see the access model above).

Networks (prod htz-fsn1) — CIDRs propagated, not hardcoded

The htz-fsn1 site layer exposes a NetBird Network named htz-fsn1: the mgmt group are the routing peers, and each routed subnet is a netbird_network_resource. The CIDRs are derived from the same fabric-addressing module the fabric stack uses (re-instantiated in the layer's addressing.tf — a pure, stateless module, so no duplication and no cross-stack coupling). Every resource is tagged into the site's own resources group, so access is the module-generated yucca-prod-htz-fsn1-yucca-to-resources policy. The only per-site input is the site id (the CIDRs flow from it):

site_id = 40   # mirrors prod/htz-fsn1; feeds fabric-addressing → the routed CIDRs
               #   mgmt 10.40.5.0/24 · api 10.40.10.0/24
               #   cls1_public 10.40.20.0/23 · cls1_private 10.40.22.0/23

Setup-key plaintext → 1Password. Each setup key's secret key is written to the per-env vault (yucca_tf_<env>) as item NETBIRD_<UPPERCASED_NAMESPACED_NAME>_SETUP_KEY (onepassword_item, same "TF mints secrets into 1P" pattern as the JWT keypair). The namespaced title keeps multiple prod sites writing to the one yucca_tf_prod vault from colliding.

Auth. Two providers, both fed by op run --env-file=tf/.env[.prod]:

  • netbird — admin PAT from NB_PAT (op://shared_tf/NETBIRD_TF_PAT, shared across all envs; management_url defaults to NetBird Cloud).
  • onepassword — OP_SERVICE_ACCOUNT_TOKEN (same session), writes the keys.

Run it (pure cloud API — no tailnet, no node contact):

TF_STACK_DIR=tf/deployment/staging/global/netbird mise run tf:init   # then tf:plan / tf:apply
# prod — global layer first, then each region layer (uses the prod env file + SA):
OP_ENV_FILE=tf/.env.prod TF_STACK_DIR=tf/deployment/prod/global/netbird   mise run tf:apply
OP_ENV_FILE=tf/.env.prod TF_STACK_DIR=tf/deployment/prod/htz-fsn1/netbird mise run tf:apply

CI (.github/workflows/infra.yml) applies staging/global/netbird in the staging matrix, and the prod layers (prod/global then prod/htz-fsn1/netbird) as gated prod-infra jobs on the prod 1P SA / tf/.env.prod. Prod CI needs the OP_TF_YUCCA_PROD_ENV[_WRITE] repo secrets + a prod-infra Environment — see the workflow header.

CI connects over NetBird

CI reaches the staging 10.10.10.0/24 nodes over the NetBird overlay (this replaced the Tailscale subnet-router path). The .github/actions/netbird-connect composite action installs the client and runs netbird up with the ci setup key read from 1P (op://yucca_tf_staging/NETBIRD_YUCCA_STAGING_CI_SETUP_KEY); the runner joins as a ci peer and the existing staging route advertises the LAN. The apply job applies staging/global/netbird first (minting that key) before connecting, so a fresh bootstrap is self-contained. The prod fabric workflow (fabric.yml) still uses Tailscale — 10.40.5.0/24 isn't on NetBird yet.

Where secrets actually live

  • yucca_tf_dev (team-shared): live values consumed by Ansible at play time. Password items per cluster (ops, dashboard, grafana, S3 svc-user access + secret), SSH Key items per cluster (ansible-iac keypairs), and DR-capture items per cluster (RGW TLS cert + key, client.admin keyring — populated by mise run capture).
  • yucca_tf_dev_manual (team-shared): placeholders for human-fillable secrets (API tokens, OAuth client secrets). Not yet used by ceph-cluster.

Service accounts themselves are in yucca_tf_dev as two items:

SA Purpose Consumed by
yucca_futo_1pass_superuser_service_account Read + write all yucca_tf_* vaults TF (via tf/.env)
yucca_futo_1pass_service_account Read-only on yucca_tf and yucca_tf_dev Ansible runtime / CI

Both are shared with other Futo consumers (o11y, base Yucca infra). Rotation affects all of them — see ansible/ceph/docs/runbooks/rotate-sa-token.md for the coordination procedure.

Adding a new cluster

  1. Add an entry to clusters.auto.tfvars.
  2. Create the 1P items in yucca_tf_dev:
    • Password items: <CLUSTER>_CEPH_{OPS,DASHBOARD,GRAFANA}_PASSWORD, plus <CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_{ACCESS,SECRET}_KEY. Use op item create --generate-password for each.
    • SSH Key item: <CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY via op item create --category "SSH Key" --ssh-generate-key=ed25519.
  3. mise run tf:apply — renders inventory + secrets template.
  4. Create per-node host_vars/*.yml files in the new inventory dir (hardware topology; not TF-rendered yet).
  5. On operator workstation: scripts/install-ssh-keys.sh <cluster> to pull the private key from 1P.
  6. After first successful deploy: mise run capture to snapshot the RGW TLS material + admin keyring to 1P for DR.
  7. See ansible/ceph/docs/adding-a-cluster.md for the full walk-through.