* feat(ceph): scaffold spice prod cluster (reprovision + inventory + TF) Stand up spice (48x Hetzner SX295, prod/htz-fsn1): the reprovision_hetzner ansible role (rescue -> reset -> installimage -> verify, base-OS-only, with stale-mdraid pre-clean and resume markers), the prod-htz-fsn1/spice inventory + group_vars/host_vars, the prod ceph TF stack, and gen-spice-host-vars.py. Adds a `mise reprovision` task (op run + tf/.env.prod), the spice SSH key (yucca_tf_prod), and per-host roles-based [ceph_mon] filtering so large clusters pin a fixed MON quorum instead of defaulting to all nodes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): create spice block.db + ssd-osd LVs at converge (NVMe-RAID) Add an NVMe-RAID branch to ceph_deploy lvm-setup so spice's single vg0 (built by installimage) gets its 14 block.db LVs + one ssd-osd LV created at converge, before osds.yml. This replaces painbox's fragile installimage -x chroot post-install with an idempotent, observable ansible step; installimage stays base-OS-only. Narrows the old blanket "externally-managed LVM" skip so it only fires when neither the sietch dual-SSD nor the NVMe-RAID shape applies, and wires ceph_db_vg / ceph_ssd_osd_lv / ceph_ssd_osd_reserve_gib in the spice group_vars. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): set spice bond NICs in group_vars, drop per-host placeholders The 25G fabric NICs are uniform across all 48 SX295 (verified live on 36 nodes): bond members enp193s0f0/f1 (Intel E800/ice, PCI c1:00.0/.1), WAN enp197s0 (igb, c5:00.0). Set bond_interfaces + oob_nic once in the spice group_vars -- the networkd role reads bond_interfaces -- and remove the dead per-host fabric_nic/oob_nic PLACEHOLDER lines from all 48 host_vars, spice-hosts.yaml, and gen-spice-host-vars.py. No MAC-based naming needed: the predictable names are PCI-derived and identical fleet-wide. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): LACP + VLAN sub-interfaces on the bond, lldpd fabric config networkd role: emit 802.3ad LACP params (TransmitHashPolicy, LACPTransmitRate) when bond_mode is 802.3ad, and support tagged VLAN sub-interfaces on the bond via networkd_bond_vlans -- when set, bond0 carries no L3 and each VLAN carries its own address. Opt-in: an empty list keeps the flat active-backup path (sietch) byte-identical (verified by render). verify.yml gates the bond0-IP/gateway asserts to the flat case and checks per-VLAN addresses otherwise. spice: bond0 becomes an 802.3ad LACP bond of enp193s0f0/f1 (MLAG to the QFX5240 leaves) carrying VLAN 120 (public 10.40.20.<idx>/23) + VLAN 122 (private 10.40.22.<idx>/23); default route stays on the 1G WAN until cutover. baseline role: install + configure lldpd (portid ifname, cluster system description, service enabled) when lldpd is in baseline_extra_packages, and add lldpd/ethtool/tcpdump for spice. This also newly applies lldpd config to sietch (already lists lldpd) on its next converge. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): set spice RGW pool to EC 16+4 host domain for the beta Capacity-max profile (80% usable, m=4) for the wipeable beta behind michael. Host failure domain accepts rack-loss risk in exchange for capacity, which is fine for a beta and matches the single-rack API tier. Also lifts the RGW DRAFT marker and starts the bulk EC data pool at pg_num 4096 (672 OSDs, 20-chunk pool) instead of the 36-OSD-era default to avoid PG splitting during the first fill. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * docs(ceph): drop stale -x chroot reference from spice autosetup header installimage.yml runs base-OS-only (installimage -a -c /autosetup, no -x). The autosetup header still carried the painbox `-x /tmp/post-install.sh` line, which implied a chroot post-install the role deliberately dropped. Point it at the convergence steps (lvm-setup, baseline) that replaced it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): minimal -x chroot to seed root + ansible-iac keys on reprovision installimage sometimes runs without our key present in the rescue (Robot armed rescue without it, in-place installimage, manual rescue), so its "copy the rescue authorized_keys" step leaves root keyless and the node unreachable for convergence. Add a minimal `-x /tmp/post-install.sh` that does ONLY inert user/key writes: seed root's authorized_keys with the iac key (baseline never manages root, so it persists) and create an ansible-iac account with the key + NOPASSWD sudo. No apt/LVM in the chroot -- that fragile step stays at post-boot convergence, so set -e cannot spuriously disrupt installimage. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): converge iac access on existing nodes without reimaging Nodes imaged before the reprovision -x chroot lack root's iac key and the ansible-iac sudo account; reseeding them meant a reinstall. baseline gains an iac-access task that ensures the same state idempotently -- root's authorized key (non-exclusive) plus the ansible-iac user with key and NOPASSWD sudo -- so the 36 already-installed nodes converge in place. Gated on the cluster iac key; run standalone with --tags iac_access. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): manual rescue-first add-node path, Robot API optional The Hetzner Robot rescue/reset step is both the unreliable part (nodes that arm rescue but never boot it) and the highest blast radius (an errant /reset on a live spice node in production). Gate it behind reprovision_use_robot (default true, preserving reprovision.yml) and add add-node.yml, which sets it false: the operator arms rescue by hand in the portal and Ansible runs only install, chroot and verify on a node already in rescue. Robot key registration moves to register_robot_keys.yml, imported from preflight only on the Robot path. With the API out of the loop nothing here can flip a running node into rescue -- wait_rescue refuses any node without the installimage ramdisk. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): relax stale-rescue guard on the manual add-node path wait_rescue rejects a rescue older than reprovision_rescue_max_uptime (1200s) to catch a stale boot -- valid only when Ansible triggered the reset. On the manual path the operator rescues by hand, so uptime just measures wait-before-run and a node legitimately in rescue for hours would be refused. Relax the guard to a day on add-node.yml; the installimage-ramdisk check is the real in-rescue proof. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): skip the this-boot rescue guard on the manual path philip has sat in rescue ~8 days, so any finite uptime ceiling rejects it. The guard only means something when Ansible triggered the reset (uptime proves this boot); on the manual add-node path a node may sit in rescue for weeks, so gate the assert on reprovision_use_robot rather than inflating the timeout. The installimage-ramdisk check remains the real in-rescue proof. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): networkd/ifupdown coexistence to protect the 1G WAN The role assumed a full migration off ifupdown -- correct for sietch's flat bond, but on spice it would disable ifupdown and leave the 1G WAN (a separate NIC that ifupdown owns, carrying the default route) unconfigured on the next boot: the painbox failure. Add networkd_replace_ifupdown (default true, sietch unchanged). When false, networkd manages only the bond and VLANs; an Unmanaged=yes guard fences the WAN NIC off, commit leaves networking.service enabled, the bond drops to RequiredForOnline=no so a carrier-less fabric cannot stall boot, and verify asserts the WAN kept its address and default route before anything commits. The rollback script targets the right interface per mode. spice sets it false. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make migrate-networkd usable pre-Ceph (gate health checks) The ceph-health pre/post gate assumed a live cluster and a ceph_bootstrap group, so the play could not run during spice bringup (coexistence activation before any cephadm deploy). Gate the four ceph tasks on ceph_bootstrap being populated; on a pre-Ceph cluster they skip and the networkd role runs on its own. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): tunable serial + halt-on-failure for migrate-networkd Fleet rollout of the coexistence activation wants batches, not one-at-a-time, and must stop the instant a node fails (a dropped WAN shows up as unreachable) rather than silently skip it. Template serial (networkd_serial, default 1 unchanged) and set max_fail_percentage 0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): install Intel ice DDP firmware so E810 NICs leave Safe Mode The minimal Debian image ships no firmware-misc-nonfree, so the E810 DDP package (intel/ice/ddp/ice.pkg) is absent and every NIC boots into Safe Mode -- whose crippled classifier drops reserved-multicast control frames (LACP, LLDP), so the 25G fabric bond never aggregates despite a correctly configured switch. Add firmware-misc-nonfree to the spice package set and a baseline nic-firmware task that reboots an E810 host once to load the DDP when it is still in Safe Mode (self-gating: no-op on non-ice or already-loaded hosts; refuses to reboot when the DDP is absent, so it cannot loop). The WAN is a separate igb NIC, so the reboot keeps the node reachable. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): correct ice DDP activation gating (pipefail + dpkg check) The Safe Mode check used `set -o pipefail`, but `ethtool --show-fec` exits non-zero on "not supported", which pipefail propagated and masked grep's match -- so the check always yielded "ok" and the activating reboot never fired. Drop pipefail there, and gate the reboot on the dpkg DB (firmware-misc-nonfree installed) rather than a file stat that can lag a large apt transaction. Verified on one node end-to-end: DDP loads, Safe Mode clears, LACP converges with the switch, both VLAN gateways ping. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): pin chrony time sources for Ceph time sync chrony was listed in baseline_enable_services but never installed, so system.yml would fail starting it -- and there was no time-source config at all, leaving sync to the distro default pool. Ceph mon quorum is skew-sensitive, so make it explicit: add baseline/chrony.yml (install + templated chrony.conf + enable) driven by chrony_ntp_servers (default Debian pool, makestep for the initial correction), drop chrony from baseline_enable_services so it is owned in one place, and point spice at Hetzner NTP (ntp1/2/3.hetzner.de) -- low-latency from FSN1, consistent across all 47. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): populate spice OSD by-path map; exclude noreen until recovered Inspected the SX295 disk layout: 14x 20TB SATA HDDs per node across 3 AHCI controllers, by-path uniform on 46/47 nodes -- so ceph_hdd_osds (path_phy + db vg0/db-slotN) and ceph_ssd_osds (vg0/ssd-osd) live in group_vars, not 47 host_vars. spice-ceph-miguel has one disk on 46:00.0-ata-4 rather than 87:00.0-ata-4 and overrides the map in its host_vars. Verified by rendering osd-spec.yml.j2: 658 HDD OSD paths (47x14) + 47 NVMe ssd-osd, miguel's override resolving correctly. noreen (boot-order casualty, held in triage) is commented out of clusters.auto.tfvars so the rendered inventory + deploy target only the 47 live nodes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): bootstrap mon on the fabric public_network, not the WAN cephadm bootstrap --mon-ip must fall inside public_network, but spice passed bond_ip, the 1G WAN address used for the ansible/SSH connection and cephadm host management. spice's public_network is the 25G fabric (10.40.20.0/23), so the initial mon would fail to bind. Bootstrap now uses ceph_public_ip (10.40.20.<host_index>, derived like the bond VLANs); sietch is flat with no host_index and falls back to bond_ip, already in its own public_network. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * ci(ceph): static validation gate for the ansible/ceph stack CI only validated the Kubernetes/Flux surface; nothing parsed the ceph playbooks, inventory or scripts, so a broken playbook or host_var first executed during the post-merge prod apply against real nodes. This runs the existing local validators (yamllint, ansible-lint, shellcheck, ansible-playbook --syntax-check, py_compile) on ansible/ceph PRs with no secrets and no connection to any host; a throwaway localhost inventory satisfies --syntax-check since the real inventory is TF-generated. Its own workflow, not a job in ci.yml, so the ansible/ceph/** path filter is workflow-scoped and does not gate the unrelated k8s jobs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make the ceph converge deliberate and cluster-parameterized A merge touching ansible/ceph/** could open the staging partition and auto-run the full baseline->tune->deploy->harden pipeline against the live sietch cluster, and the converge was hardcoded to sietch/austin (render-inventories.sh got no region, install-ssh-keys.sh got a literal sietch, CEPH_ENV pointed at .../sietch/), so a prod converge could never reach spice. The ceph converge now runs only on workflow_dispatch with run_ceph_ansible=true, pinned to the one matrix entry that owns the chosen ceph_cluster (spice=prod/htz-fsn1, sietch=staging/austin), so a push never reconverges a live cluster and one dispatch cannot converge both. Region and cluster are threaded from the dispatch inputs / matrix into the render, key install, and CEPH_ENV. The staging paths-filter is scoped to ansible/ceph/inventories/staging-** so a prod ceph change no longer opens the staging matrix. The ceph TF apply stays auto and env-gated; the mgmt converge is unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make the RGW firewall source scope configurable The RGW/S3 port was accepted from any source unconditionally, while every other Ceph service (mon, osd, dashboard, monitoring) is scoped to ceph_firewall_trusted_networks. On spice that port is plain HTTP 7480, so the S3 endpoint sat open to the public internet with no lever to close it. A new ceph_firewall_rgw_any_source (default true, mirroring ceph_firewall_ssh_any_source) keeps the open behavior by default but lets production restrict RGW to the trusted networks, which already cover the fabric and NetBird overlay. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): fix stale OSD-map comment on the noreen host_var noreen was generated before the by-path map moved into group_vars, so it still carried the DEFERRED note while its 47 siblings point at the group var. The host stays pre-staged for re-add once recovered; only the comment was wrong. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): add the host-mgmt VLAN (124) to the spice bond The FSN1-C1 fabric carries a third VLAN on the 25G bond, FSN1-C1-HOST-MGMT (124, 10.40.24.0/24), alongside Ceph public (120) and private (122). The leaf already trunks it to every server bond and advertises 10.40.24.0/24 into the NetBird overlay, so it is the intended in-band ansible/SSH reach once the WAN is retired, but the host side had no matching sub-interface. Add bond0.124 at 10.40.24.<host_index>/24 with no gateway, so the default route stays on the 1G WAN until cutover. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): apply networkd config deltas to an already-running daemon The networkd role only started systemd-networkd, which is a no-op once it is already running, so a re-run that added config (e.g. a new bond VLAN) wrote the .netdev/.network files but never applied them, and verify then failed asserting the sub-interface had no address. Add a networkctl reload plus a per-VLAN settle wait after the start, so a re-run creates the added sub-interfaces live without tearing down existing links; the WAN on ifupdown is untouched regardless. Non-disruptive on the flat sietch path (the settle wait is gated on networkd_bond_vlans). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): bind Ceph services on the fabric only, never the WAN Ceph on spice must never listen on the 1G WAN. mon/mgr/OSD already bind per public_network/cluster_network, but the beast RGW frontend and the cephadm monitoring stack bound 0.0.0.0, and the mon-ip, cephadm host-add address, and dashboard URLs advertised the WAN bond_ip. Add a per-cluster policy: ceph_service_ip (the address Ceph advertises and binds for point services; defaults to bond_ip, resolving to the fabric on spice via ceph_public_ip) and ceph_bind_networks (the cephadm networks: bind restriction for RGW and the monitoring daemons; defaults to public_network). Route mon-ip, the host-add address, the dashboard monitoring URLs, RGW, and prometheus/grafana/alertmanager/node-exporter/ ceph-exporter through them, and scope the RGW firewall to the trusted networks. The RGW zonegroup hostnames and TLS SAN gain the fabric address so signed requests to the fabric IP still validate. Flat clusters like sietch have no ceph_public_ip, so ceph_service_ip falls back to bond_ip and ceph_bind_networks to their single flat network; point services are unchanged there. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): define the metrics-worker RGW user on spice rgw.yml Step 14.5 creates a read-only RGW admin for the yucca-metrics-worker service and references ceph_rgw_metrics_user_*, which only sietch defined, so the RGW phase failed on spice with AnsibleUndefinedVariable. Add the block; the keys are the TF-minted SPICE_METRICS_WORKER_* items, injected via secrets.yml.tpl as vault_metrics_worker_*, so op inject resolves them. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): refuse to deploy against an un-rendered inventory The ceph_deploy role gates every phase on the ceph_bootstrap/ceph_mon/ ceph_join groups that render-inventories.sh emits from TF state; run against the hand-written stopgap inventory that lacks them, the role errors mid-deploy or falls through to placing a MON on every host. A pre-task assert now requires ceph_bootstrap to be exactly one host that is also in ceph_mon, and ceph_mon to be a non-empty odd quorum, and stops with a render hint otherwise. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make the fabric-only bind restriction opt-in per cluster ceph_bind_networks defaulted to public_network, which would have injected a networks: bind restriction into every cluster - re-binding sietch's live RGW and monitoring daemons from 0.0.0.0 to its flat network on the next converge (a redeploy, and a break for any access path not on that subnet). Default it empty instead: no networks: field is emitted and the monitoring re-spec is skipped, so flat clusters bind every interface exactly as cephadm ships them. spice opts in to the fabric public network in its group_vars. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make ceph_service_ip a group_var so hostvars can read it ceph_service_ip was a ceph_deploy role default, but join/rgw/monitoring read it as hostvars[<host>]['ceph_service_ip'], and role defaults are NOT exposed through hostvars - the lookup raised "HostVarsVars object has no attribute ceph_service_ip" and would abort the deploy on every cluster at the join phase (a regression the fabric-only change introduced for both spice and the live sietch). Define it in each cluster's group_vars instead (group_vars do resolve through hostvars, verified per host): spice to the fabric ceph_public_ip, sietch to bond_ip (unchanged from pre-fabric). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): drop the unused ceph_cluster_ip var on spice ceph_cluster_ip was defined for symmetry with ceph_public_ip but consumed nowhere - OSD replication binds to the cluster_network CIDR, which cephadm resolves to each node's bond0.122 address on its own. Remove the dead var and correct the neighbouring comment (ceph_service_ip is a group_var now, not a role default). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): serve spice RGW over TLS on 443 with a 10-year self-signed cert Match sietch: flip spice RGW/beast from plain HTTP 7480 to HTTPS 443 with a self-signed 10-year cert. rgw.yml generates /etc/ceph/rgw-ssl.{crt,key} and cephadm distributes it to every RGW daemon via the service spec; the SANs already cover s3.<domain> + the wildcard + each node's fabric IP (ceph_service_ip). The firewall RGW port derives from ceph_rgw_port, so it follows to 443, scoped to the trusted networks (RGW binds the fabric only, never the WAN). This also makes the discovery rgw_s3_endpoint (hardcoded https) correct for spice. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(dns): prod Cloudflare DNS stack for the spice RGW S3 endpoint Add tf/deployment/prod/global/dns (mirroring staging/global/dns) so s3.prod.fsn1.htz.futo.cloud and its wildcard resolve. Round-robin A across all 47 spice ceph nodes' fabric public IPs (10.40.20.<host_index>, VLAN 120) where RGW/beast binds; proxied=false (private RFC1918, reached over the NetBird-advertised cls1_public 10.40.20.0/23). The wildcard serves S3 virtual-hosted buckets, and both names are in the self-signed TLS cert SANs. noreen (host_index 40) excluded. CI auto-discovers the stack (applies at order 0 under the prod-global environment); tf/.env.prod gains the token ref. Next up: create op://yucca_tf_prod/CLOUDFLARE_API_TOKEN before the apply. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): preflight-assert the Ceph service IP is up before deploy A node whose fabric VLAN sub-interface (spice bond0.120) did not come back after a reboot would otherwise fail deep inside cephadm bootstrap/join. Add a per-host pre_task that asserts ceph_service_ip is present in ansible_all_ipv4_addresses, so a fabric-down node halts up front with a clear message. Passes on a healthy node (verified on spice-ceph-adelia); on sietch ceph_service_ip is bond_ip, the connection address, so it is trivially satisfied. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): honest ssd_model_pattern note + guard empty monitoring networks ssd_model_pattern on spice was "SAMSUNG" with a "verify on node 1" note, but the NVMe is SOLIDIGM/KIOXIA and the var is only read by ceph_destroy's partition-6 SSD wipe (sietch's dual-SSD shape); spice is NVMe-RAID with no partition 6 and a vg0 ssd-osd LV cleaned by the generic VG/PV removal, so the pattern is a deliberate no-match. Reword to say so (it must stay defined because cleanup.yml references it unconditionally). Also wrap the monitoring-spec networks block in a length guard so an empty ceph_bind_networks renders no dangling `networks:` key (defensive; the apply is already gated). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): lock down root/password SSH in baseline, not just harden installimage ships PermitRootLogin yes + a root password, and all sshd hardening lived in harden.yml (the last deploy stage), so a freshly imaged node sat root-password-open on the public WAN for the whole campaign (the 11-node exposure the swarm audit found). Add an early baseline task, right after iac-access authorizes the key on root, that deploys a 10-baseline-ssh drop-in (PermitRootLogin prohibit-password + PasswordAuthentication no, same values as security/50-hardening.conf so they never disagree), locks the root password, and removes the interim remediation drop-in. A Validate->Reload handler chain runs sshd -t before reloading, only on change. Key-safe on both clusters (ansible connects by key), so nothing can lock out. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): pin RGW metadata + mgr pools to the replicated size The RGW system-metadata pools (.rgw.root, {zone}.rgw.{log,control,meta}) and .mgr inherited the cluster-spec default osd_pool_default_size 2 / min_size 1, unlike the index/extra pools which are pinned. Losing two OSDs behind a metadata PG would take out the RGW/mgr control plane, and min_size 1 permits single-replica writes. Pin them to ceph_rgw_replicated_ size/_min_size (spice 3/2; sietch keeps 2/1 via its own vars), existence- gated like the pg_num loop and idempotent (only sets when the value differs). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make the ops password hash idempotent users.yml hashed ops_password with password_hash('sha512') and no salt, so a fresh random salt was drawn every run: the hash never matched /etc/shadow and the user module rewrote it reporting 'changed' on every converge (a clean converge was never a true green signal, on both clusters). Derive a stable salt from a one-way sha256 of the password so the hash is deterministic and idempotent, while still reconciling an out-of-band password change. The salt in /etc/shadow is public and leaks nothing. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make live fabric/hosts/timezone reproducible from the repo Three drift gaps where the running fleet did not match the committed IaC. networkd_enabled was committed for sietch only, so spice's fabric VLANs came up from an ad-hoc -e and a reprovisioned node would not bring them up from the repo alone; commit it true for spice (networkd_replace_ifupdown false still fences the 1G WAN). hosts.j2 mapped every node name to bond_ip (the 1G WAN) rather than the fabric ceph_service_ip, so in-cluster name resolution pointed off the fabric. And timezone: UTC was declared but never applied, leaving nodes on the image default (Europe/Berlin); add a community.general.timezone task to baseline/system.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): scope OSD/RGW readiness waits to the in-play host set join.yml is --limit-aware (it intersects ansible_play_hosts), but the OSD provisioning wait computed EXPECTED from all ceph_nodes and the RGW wait blocked until running rgw >= all ceph_nodes, so a subset (--limit) deploy joined the subset then deadlocked at both gates. Base both counts on groups['ceph_nodes'] intersect ansible_play_hosts, matching join.yml. On a full deploy the intersection is all nodes, so behavior is unchanged; the rgw-spec placement stays pinned to all ceph_nodes (only the readiness gate is play-scoped, so a --limit re-run cannot shrink the spec). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): scope the nftables ruleset to table inet filter nftables.conf.j2 did a global `flush ruleset`, and the no-op guard diffed the whole `nft list ruleset`. On a ceph node (cephadm runs podman) that wiped the podman/docker/fail2ban tables on every apply, and made the desired-vs-live diff never match (podman tables are live but absent in the throwaway netns), so the firewall reloaded on every converge - each one flushing podman again. Replace only table inet filter (add, delete, re-add), compare only that table in the guard, and reload via ExecReload (nft -f, no global flush) instead of restart (whose ExecStop flushes the whole ruleset). The fabric/ceph rules themselves are unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): deploy MGR at a capped count, not one per host placement.yml applied mgr to every active host (47 mgr daemons on spice: 1 active + 46 standby), which is wasteful and non-standard - MON already uses a bounded ceph_mon set. Deploy count:{{ ceph_mgr_count }} (default 3: 1 active + 2 standby) instead; cephadm schedules them and caps at the host count on small clusters like sietch. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): mark the RGW EC data pool as bulk The EC data pool is pre-seeded to pg_num 4096 (sized for spice's 658 HDD OSDs), but without the bulk flag the pg_autoscaler can walk it back toward 1 during the first fill, causing PG splitting under load. Set bulk=true so the autoscaler targets a full-capacity pg_num and treats the pre-seed as a floor. Idempotent: only set when not already true. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): harden the ceph_destroy SSD-model match The SSD OSD PV-remove and partition-wipe loops used grep -i "$ssd_model_pattern", which interpreted the pattern as a regex and, on an empty pattern, matched every disk - the worst-case foot-gun in a destroy path (it would target all disks). Use grep -iF (fixed string, no regex) and skip the loop entirely when the pattern is empty. Kept -F without -w, since -w would fail to match underscore-containing model strings like Micron_5100_MTF... Shared role, so it hardens sietch too. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): drop the dead s3_key_drift notice; reconcile harden header The "S3 key drift notice" debug task was gated on s3_key_drift, a variable nothing ever sets (the drift detection it stood in for was never built), so the branch could never fire - remove it. And update the harden.yml header: the security-critical sshd lockdown (root key-only, no password auth, locked root) now runs early in baseline, not gated behind this last stage; harden adds the firewall and the remaining sshd hardening on top. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): accurate change reporting in tuning roles (T4) CRUSH tunables and per-device HDD/SSD sysfs writes ran a mutating command every converge under changed_when:false, so a run reported ok even when it changed state and --check misled register consumers. Switch to get-then-set: read the current value first, only write when it differs, and emit CHANGED so changed_when reflects reality. Same settings applied; only change-detection becomes truthful. Also drops two pre-existing em-dashes to keep the file plain ASCII. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * refactor(ceph): declarative lvg/lvol for OSD block.db setup Replace hand-rolled pvcreate/vgcreate/lvcreate shell in ceph_deploy lvm-setup with community.general.lvg/lvol (already pinned, and the idiom provision_host/disks.yml already uses). Same VG/LV names, sizes, fixed-vs-100%FREE split, per-shape loops, and device targets on both shapes. Kept read-only asserts/verify/show tasks as-is. Preserved the wipefs -af signature clear ahead of PV creation: lvg does pvcreate -f but not --yes, so it will not wipe a stale foreign signature on a reused partition. wipefs stays gated on the VG being absent so a live PV is never touched. Kept a minimal shell to compute the spice ssd-osd size (vg0 free - reserve) since lvol has no free-minus-N primitive, plus an existence guard so re-runs don't recompute a bad size or trigger an unforced lvol shrink. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): ASCII-clean the lvm-setup header comment Convert the pre-existing arrows and em-dashes in the header (left untouched by the lvg/lvol refactor) to plain ASCII. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): never let lvol shrink a live block.db LV (data safety) The lvg/lvol refactor let community.general.lvol reconcile an existing fixed-size db-slot to its size, and lvol defaults to shrink:true - so a db-slot that ever drifted LARGER would be TRUNCATED, corrupting a live BlueStore block.db and taking out the OSD. The old shell skipped existing LVs entirely, so it could never do this. Set module_defaults shrink:false on both lvm-setup blocks: lvol still creates and may grow, but never truncates an existing LV (it leaves a larger one alone). Also restore opts:-Wy so (re)created LVs wipe stale signatures - the destroy->recreate recovery-path freshness the explicit -Wy gave. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): RGW self-signed cert generation aborted on first converge The openssl -subj and -addext args used backslash-newline continuations inside their quoted strings, so shell line-folding kept the next line's indentation in the value; countryName rendered as "DE " (over the 2-char max) and openssl exited 1 under set -euo pipefail, leaving no S3 endpoint on the first spice deploy. Precompute the subject and SAN as single-line facts (whitespace-controlled Jinja for the SAN) and pass them quoted, so each openssl arg is one flat string. See ansible/ceph/roles/ceph_deploy/tasks/rgw.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): green the ansible-lint CI gate on HEAD Bare ansible-lint (what mise run lint and the required CI check run) exited 2 on two deliberate patterns: run_once on the fleet-wide placement assert, and an intentional no-pipefail shell in the ice DDP Safe-Mode probe (pipefail there would mask grep's match). Waive both with inline noqa so the reasoned patterns stay and the gate passes. See ansible/ceph/deploy-ceph.yml and roles/baseline/tasks/nic-firmware.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): pin RGW pools to device-class CRUSH rules (SSD fast tier) The replicated_ssd/replicated_hdd rules were created but bound to no pool, so every replicated pool fell to the default class-agnostic rule and the omap-heavy bucket index landed on HDD (~98.5% of weight) while spice's 47-OSD NVMe ssd-osd tier sat idle. Pin the index, RGW metadata, and mgr pools to the SSD rule and the bulk non-EC pool to HDD; force the OSD device class at creation (ssd-osd -> ssd, HDD -> hdd) so the rules bind deterministically, and assert the placement in verify.yml. Pinning is opt-in per cluster (empty default) so a live cluster is never re-homed implicitly; the pin runs after RGW readiness because the .rgw.* system pools are created lazily by the realm/zone setup and the daemon. See ansible/ceph/roles/ceph_deploy/{templates/osd-spec.yml.j2,tasks/rgw.yml,tasks/verify.yml}. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): apply RGW pool PG/durability pins on the first converge The pool list was captured once at the top of the RGW play, before any pool was created, so every "pool in list" gate skipped on a fresh run 1: the data/index/non-EC PG pre-sizing deferred to a second converge, and the .rgw.* system pools (created lazily by the daemon) never got their size/min_size pin at all, sitting at the inherited 2/1 single-replica until someone ran the play twice. Re-query the pool list once the explicit pools exist, and move the system-pool PG + size/min_size pins below the RGW readiness wait where those pools are real. Parameterize the bootstrap osd_pool_default_size per cluster (spice 3/2, role default keeps sietch at 2/1) so an unpinned future pool is not born single-replica, and assert final size/min_size in verify.yml so a miss fails the play loudly. See ansible/ceph/roles/ceph_deploy/tasks/{rgw.yml,verify.yml}. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): real ok-to-stop gate before a live-node reimage Reimaging a spice node zeroes the OS NVMe, and the 14 HDD OSD block.db LVs plus the ssd-osd data LV all live on that NVMe, so a reimage destroys every OSD on the node; the old guard only checked a hand-typed i_know_its_dead flag and the prepare-os-disks comment wrongly claimed a reimage keeps OSD data. Replace the stub with a mon-delegated check that queries the node's OSD ids, refuses on HEALTH_ERR or a failed ceph osd ok-to-stop, and sets noout on those OSDs for the reimage window. The gate is default-on: auto mode enforces whenever a live cluster with this node's OSDs is reachable and no-ops for the initial bootstrap; strict fails closed if it cannot verify; permissive is the explicit escape hatch. See ansible/ceph/roles/reprovision_hetzner/tasks/ceph_safety.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): jumbo MTU 9000 on the VLAN 122 cluster/replication network Jumbo pays off on the OSD replication path (a closed, homogeneous fabric) but is a partial-blackhole risk on the RGW-facing public network, which serves heterogeneous 1500-byte S3 clients and the ~1400 NetBird overlay, so raise MTU on VLAN 122 (Ceph private) only and leave 120/124 at 1500. A VLAN opts into jumbo with a per-entry mtu; the bond ceiling and both 25G members auto-raise to the largest child MTU so a 1500 parent or slave cannot silently cap the jumbo frames. verify.yml asserts the applied MTU on the bond, members, and each VLAN (local, no switch dependency). Flat clusters resolve to 1500 and emit no MTU, so sietch is unchanged. Paired with a QFX-side change (leaf server-LAG + VLAN 122 IRB >= 9000) and a ping -M do -s 8972 matrix before the cluster network is relied on. See ansible/ceph/roles/networkd. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(dns): commit prod DNS provider lock file The prod DNS stack was the only DNS stack without a committed .terraform.lock.hcl, so cloudflare ~> 5.0 would resolve unpinned at the first prod apply of the stack fronting the S3 endpoint. Pin it to 5.21.0, matching the staging DNS stack, with linux_amd64 + darwin_arm64 hashes. See tf/deployment/prod/global/dns/.terraform.lock.hcl. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): opt-in alertmanager receiver so alerts reach a human cephadm deploys alertmanager with no receiver, so every Prometheus alert rule (OSD down, host down, PG degraded, near-full) fires into the default null route and nobody is notified. Add ceph_alertmanager_webhook_urls: a non-empty list routes all alerts to those webhooks via cephadm user_data.default_webhook_urls. The monitoring re-spec now applies when either fabric bind or alerting is configured (independent gates), and verify.yml reports the delivery status, warning loudly when none is set. Empty by default, so flat clusters (sietch) are unchanged; the spice destination is a deferred operator decision (TODO in defaults). See ansible/ceph/roles/ceph_deploy/tasks/monitoring.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(fabric): jumbo MTU on the leaf server bonds + VLAN 122 IRB Completes the host-side VLAN 122 jumbo change on the QFX leaf: the server LACP bonds (ae1..aeN) get the 9216 L2 jumbo MTU that the spine uplink ae0 already carries (802.3ad members inherit it, so no per-member mtu), and the private/cluster IRB gateway (VLAN 122) gets an L3 family-inet MTU of 9000 to match the Ceph hosts. Public (120) and host-mgmt (124) IRBs stay at 1500 by design, so jumbo is confined to the closed OSD-replication path. See tf/shared/modules/cluster-fabric. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ci): pin + SHA256-verify NetBird install, drop curl|sh The netbird-connect action piped pkgs.netbird.io/install.sh straight into sh on the runner that holds the prod-write 1Password token and overlay access to the live nodes, so a compromised installer would run as root there. Download a pinned release tarball (v0.74.4) from the GitHub release and verify its SHA256 against an in-repo pin before unpacking; fail closed on any mismatch. A version input plus a refresh comment keep the pin maintainable. See .github/actions/netbird-connect/action.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ci): close the prod apply gate (reviewed-plan binding + fail-closed reviewers) The prod approval gate could fail open and the apply was not bound to the reviewed diff. Manage the four gate Environments as code (new _meta/github stack: github_repository_environment + required reviewers, protected-branches-only, a validation that rejects a reviewerless gate), and make the discover job fail closed by asserting every emitted partition-region Environment already carries a required reviewer. Bind apply to the plan: the plan job writes -out to an absolute path and uploads it, the gated apply downloads that exact file and applies it with no re-plan and no -auto-approve, so a stale plan fails closed. Scope a bare workflow_dispatch to plan-only and a toggled dispatch to just the partition it touches; add fail-fast and timeout-minutes across the jobs. The _meta/github stack is bootstrap-applied out of band and needs real reviewer IDs plus a GH_ENV_ADMIN_TOKEN secret before the gate passes. See .github/workflows/infra.yml and tf/deployment/_meta/github. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): blast-radius control on the day-2 converge plays The converge chain (baseline, tune-os, tune-hardware, tune-ceph, harden) ran against all nodes at once with no canary and no stop-on-failure, the inverse of the serial:1/max_fail:0 destructive plays. Batch each converge play (serial, default 10%) with max_fail_percentage:0, and gate every batch on cluster health via a shared pre/post ceph-health checkpoint that halts the roll on HEALTH_ERR. The gate is a no-op until a cluster is live, so the bootstrap ordering (baseline -> tune -> deploy) is unaffected; deploy-ceph stays big-bang. Tune the batch size with ceph_converge_serial. See ansible/ceph/tasks/health_gate.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): freeze load-bearing host packages instead of auto-upgrading We run no unattended-upgrades: an apt upgrade that moved cephadm's container runtime (podman + ecosystem) or the chrony time source that mon quorum depends on out from under a live cluster is the uncoordinated change we avoid. Hold those packages at their installed version (dpkg selection) so apt upgrade skips them; upgrading is then deliberate and health-gated. Diagnostic/ops tools are left unheld. Runs last in baseline so nothing is held before it exists; tune via baseline_held_packages. See ansible/ceph/roles/baseline/tasks/hold-packages.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): THP=madvise and stop the GRUB cmdline clobber Transparent Huge Pages sat at the kernel default 'always', which bloats BlueStore/tcmalloc RSS and drives allocation-stall latency across the OSD fleet; set it to madvise via a oneshot unit ordered Before=ceph-osd.target (mirroring ceph-cpu-governor.service) plus a live sysfs write, on by default for every ceph cluster. Separately, the processor.max_cstate GRUB task rewrote GRUB_CMDLINE_LINUX_DEFAULT wholesale (dropping existing tokens, injecting quiet) - a latent footgun behind the default-off governor flag; replace it with a /etc/default/grub.d drop-in that only appends the cstate token. See ansible/ceph/roles/hardware_tuning. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): host security sysctls + persistent journald Add a host kernel/network hardening sysctl set (syncookies, ignore/deny ICMP redirects + source routing, kptr/dmesg restrict, log martians) to the existing os_tuning sysctl drop-in. rp_filter is set to LOOSE (2), not strict (1): these nodes are multi-homed (1G WAN default route + 25G fabric VLAN sub-interfaces), so strict reverse-path filtering would blackhole asymmetric fabric traffic. Separately make journald persistent (Storage=persistent, bounded SystemMaxUse/RuntimeMaxUse) so logs survive the pipeline's own reboots on these headless nodes. See ansible/ceph/roles/os_tuning. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): health-gated cephadm upgrade playbook + rollback runbook No upgrade automation existed for the pinned 20.2.x tentacle train. Add a deliberate, manually-invoked upgrade-ceph.yml: preflight-gates on HEALTH_OK (with an explicit allow-WARN toggle), no in-progress upgrade, and all OSDs up+in; records and prints the rollback image before starting; refuses to run without an explicit target (ceph_upgrade_target_image/_version, no default); drives ceph orch upgrade with a bounded poll and fails loudly on a stall; asserts health + version convergence after. Not wired into site.yml. Rollback procedure lives in the play header and docs/runbooks/upgrade-ceph.md. See ansible/ceph/upgrade-ceph.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): scheduled cluster-state backups with a systemd timer Config/topology state was captured only manually and controller-local. Add a ceph_backup role that installs a capture script + daily systemd timer on the bootstrap node: each run tars fsid, config dump, mon/osd/crush maps (raw + decoded), osd tree, orch ls/host ls, and the RGW realm/zonegroup/ zone into a root-only 0700 dir, pruned by retention. No secret keyrings are dumped. An offsite target (rsync or s3://) is a var left empty for the operator. Invoke via backup-ceph.yml (mise run backup-timer); not wired into converge. Opt out with ceph_backup_enabled=false. See ansible/ceph/roles/ceph_backup and docs/runbooks/backup-restore.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): pin mgr dashboard/prometheus bind to fabric; pin EC min_size The mgr dashboard (8443) and prometheus module (9283) run inside the active mgr and bound 0.0.0.0, closed only by nftables. Pin each mgr instance's own localized server_addr (mgr/<module>/<mgr_id>/server_addr, the key the modules read via get_localized_module_option) to that host's fabric IP, discovering mgr ids at runtime so it is failover-safe; a global server_addr would leave the dashboard unbindable after failover. Opt-in on ceph_bind_networks; takes effect on the next mgr cycle (nftables holds the WAN closed meanwhile). Separately pin the EC data pool min_size explicitly to k+1 (derived from ceph_rgw_ec_k) so the write-availability floor is documented and cannot drift, and record the single-site host-failure-domain DR ceiling in docs/capacity-planning.md. See ansible/ceph/roles/ceph_deploy/tasks/monitoring.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): freeze Galaxy collection versions to a major The collection requirements floated on >= lower bounds, so a new community.general major could be pulled mid-campaign. Pin compatible-release ranges to the currently-resolving majors (ansible.posix >=2,<3; community.general >=12,<13) so a 47-node campaign cannot cross a major between runs. No lockfile mechanism exists in the repo, so the ranges live in requirements.yml. See ansible/ceph/requirements.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * test(ceph): molecule scenario for the fabric-only firewall branch The only molecule scenario tested the default open-firewall branch (*_any_source: true), the opposite of spice's production posture, with no idempotence check. Add a fabric-only scenario that flips the closed branch (rgw/ssh any_source: false, trusted networks = fabric + NetBird) and asserts RGW/SSH are NOT accepted from any source, are restricted to the fabric/overlay, and that a second render is idempotent. Mirrors the default scenario's render-and-grep idiom; the existing scenario is untouched. See ansible/ceph/roles/security/molecule/fabric-only. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * ci(dns): drift-check the prod S3 roster against the ceph node list The 47 S3 RGW A-records were hand-copied with no link to the ceph roster, so add/replace/remove of a node silently blackholed the endpoint or dropped capacity. A terragrunt dependency is not viable here (no dependency idiom in the repo, and the ceph discovery output carries only bond_ip, never the fabric IP), so add a stdlib-only, credential-free check that reconstructs the expected fabric-IP roster from clusters.auto.tfvars (in-service names) + spice-hosts.yaml (host_index) and asserts the apex and wildcard A-records match exactly. A path-scoped CI gate fails the PR on drift before the DNS apply; also runnable via mise run tf:check-dns-roster. See tf/scripts/check-s3-dns-roster.py. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): inotify headroom for podman scale; podman-compat upgrade note cephadm-on-podman runs many containers/systemd units per node and leans on inotify; Debian's default fs.inotify.max_user_instances=128 is a known bottleneck at that scale, so raise instances to 512 and watches to 524288 in os_tuning. Separately, document in the upgrade runbook that the podman dpkg-hold must be lifted (unhold -> bump to a cephadm-supported version -> re-hold) before a cross-major Ceph upgrade; patch-train upgrades need no podman change. See ansible/ceph/roles/os_tuning and docs/runbooks/upgrade-ceph.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): health-checked RGW ingress VIP (haproxy + keepalived) S3 was bare round-robin DNS across 47 RGW A-records, so a dead node blackholed ~1/47 of new connections for the TTL. Add an opt-in cephadm ingress service: a keepalived floating VIP (spice 10.40.20.250) with haproxy health-checking the RGW backends and dropping a dead one from rotation. haproxy terminates TLS on the VIP with the existing self-signed RGW cert (its SAN now carries the VIP) and re-encrypts to beast: L4 passthrough for RGW needs IngressSpec.use_tcp_mode_over_rgw, which is absent on the pinned 20.2.2 (source-verified; applying it TypeErrors), so terminate is the only health-checked mode available. Applies after RGW readiness; asserts its preconditions (VIP set, ceph_bind_networks non-empty so beast binds a per-node IP and does not collide with the VIP on :443, ceph_rgw_ssl true). Opt-in per cluster; sietch unaffected. See ansible/ceph/roles/ceph_deploy/templates/rgw-ingress-spec.yaml.j2. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(dns): point the prod S3 endpoint at the ingress VIP Collapse the s3.prod.fsn1.htz.futo.cloud apex + wildcard from 47 round-robin A-records to the single health-checked ingress VIP 10.40.20.250, and switch the drift-check to assert both records equal that VIP (the ceph ceph_rgw_ingress_vip is the source of truth). Apply ordering matters and is noted in the tfvars: the ingress must be live before this DNS cutover, or s3 resolves to a VIP nothing answers; roll back in reverse. See tf/deployment/prod/global/dns/records.auto.tfvars. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): fix stale staging copy-paste comments in the prod stack; fmt The prod ceph stack carried the staging stack's comments verbatim: versions.tf and secrets.tf described creating SIETCH_CEPH_* items in yucca_tf_staging via OP_TF_YUCCA_STAGING_ENV_WRITE, and a variable example said (sietch, ...). This stack creates SPICE_CEPH_* in yucca_tf_prod via OP_TF_YUCCA_PROD_ENV_WRITE; correct the comments to match. Legitimate cross-refs (the mirrors-staging/talos provenance, the partition-slug enumeration) are left as-is. Also tofu fmt the drifted clusters.auto.tfvars (whitespace only). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * ci(infra): scope reviewer fail-closed gate to apply events; prettier-format workflows * ci(infra): pin 1Password CLI version to avoid flaky latest resolution * feat(dns): resolve prod S3 apex+wildcard to the node fabric-IP roster * ci(infra): park the Environments approval gate (disable fail-closed step + stub _meta/github) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
yucca/tf
Terraform/OpenTofu authority for cluster identity, 1P secret items, and the inventory artifacts Ansible consumes. Multi-partition / multi-region via terragrunt.
The partition / region model
Everything is keyed on a single first-class topology:
partition → region → { exactly one K8s cluster, one-or-more Ceph clusters }
- Partition =
prod|staging|dev(formerlyenv). - Region = a site:
htz-fsn1,austin,local(formerlysite/datacenter). Plus a reservedglobalpseudo-region for partition-wide stacks (DNS, the account-wide NetBird layer) —globalis not a physical site, so itsrole/FQDN metadata is null. - The human form is
partition@region(prod@htz-fsn1); the canonical slug ispartition-region(prod-htz-fsn1), which derives names on every surface.
Every stack sits at three path segments:
deployment/<partition>/<region>/<stack>. The account-wide prod layer is
prod/global/netbird (a stack under the global region), so there is no
two-segment special case.
Layout
tf/
├── .env ← op:// references (committed; no literal secrets)
├── op-run.sh ← op-run wrapper used by the mise tf:* tasks
├── shared/
│ └── modules/
│ ├── ceph-cluster/ ← per-cluster ceph orchestration module
│ │ ├── main.tf, variables.tf, outputs.tf, rendering.tf
│ │ ├── wordlist.txt ← 923 words for auto-picked hostnames
│ │ └── templates/
│ │ ├── inventory.ini.tftpl
│ │ ├── inventory-destroy.ini.tftpl
│ │ ├── inventory-provision-debian-live.ini.tftpl
│ │ └── secrets.yml.tpl.tftpl
│ └── talos-baremetal/ ← Talos on bare-metal nodes already in maintenance mode
│ ├── main.tf, variables.tf, outputs.tf
│ └── firewall.tf ← Talos host ingress firewall (default-deny + allow-lists)
└── deployment/
├── terragrunt.hcl ← root: state backend + partition/region/stack from path
├── staging/
│ ├── austin/
│ │ ├── region.hcl ← role + FQDN parts for staging@austin
│ │ ├── ceph/ ← sietch ceph cluster (clusters.auto.tfvars)
│ │ └── talos/ ← bare-metal Talos cluster (3× CP, Cilium CNI)
│ └── global/
│ ├── region.hcl ← role=null (pseudo-region)
│ ├── dns/ ← Cloudflare records (futo.cloud)
│ └── netbird/ ← NetBird Cloud access control (staging)
├── dev/
│ ├── local/
│ │ ├── region.hcl
│ │ └── talos/ ← Talos VMs on the sietch hypervisors
│ └── global/
│ ├── region.hcl
│ └── dns/
└── prod/
├── htz-fsn1/
│ ├── region.hcl ← role + FQDN parts; site_id=40
│ ├── mgmt-hosts.yaml ← mgmt-host roster (region root; fabric + render read it)
│ ├── fabric/ ← Junos switch fabric + NetBox + mgmt reprovision
│ └── netbird/ ← htz-fsn1 site NetBird layer (routed htz-fsn1 network)
└── global/
├── region.hcl
├── terragrunt.hcl ← account-wide prod NetBird (two-segment region-root stack)
└── netbird.tf
Add a region: create deployment/<partition>/<region>/region.hcl + stacks under
it. Add a stack: a new sibling dir under a region (<region>/monitoring/ …).
NetBird Cloud access control lives in staging/global/netbird/, and for prod is
layered: prod/global/ (account-wide) above per-region
prod/<region>/netbird/ (e.g. prod/htz-fsn1/netbird/). See "The netbird-env
module" below.
The dns stack manages infrastructure names in the futo.cloud Cloudflare
zone (today: the Sietch RGW S3 endpoint + virtual-hosted wildcard).
Records are declarative in records.auto.tfvars; the API token resolves
from op://yucca_tf_manual/CLOUDFLARE_API_TOKEN via tf/.env.
The talos stack is documented in ansible/talos/README.md and
ansible/talos/docs/runbooks/cluster-bring-up.md (the TF + Ansible flow
is interleaved — TF renders the inventory Ansible consumes, then
bootstraps the VMs Ansible created).
Conventions
Partition, region, and stack are derived from the directory path
deployment/terragrunt.hcl parses the child's relative path:
partition = segs[0], region = segs[1], stack = join(segs[2:]), and
slug = "<partition>-<region>".
deployment/staging/austin/ceph → partition=staging, region=austin, stack=ceph
deployment/staging/global/dns → partition=staging, region=global, stack=dns
deployment/prod/htz-fsn1/fabric → partition=prod, region=htz-fsn1, stack=fabric
deployment/prod/htz-fsn1/netbird → partition=prod, region=htz-fsn1, stack=netbird
deployment/prod/global/netbird → partition=prod, region=global, stack=netbird
The state backend key is derived from these:
yucca/${partition}/${region}/${stack}/terraform.tfstate in the shared
yucca-tf-state S3 bucket. (The legacy ceph/ project prefix is dropped —
talos/dns/netbird/fabric all share the bucket now.) Every stack is three
segments — the account-wide prod layer is prod/global/netbird (not a bare
prod/global), so region.hcl at prod/global/ is found by
find_in_parent_folders from the stack dir, exactly like staging/global. The
n==2 branch in terragrunt.hcl is now dead and can be removed.
Per-region metadata + role (region.hcl)
Each region dir carries a deployment/<partition>/<region>/region.hcl holding
role (primary | secondary), site_id, datacenter, provider_code, and
domain. The root terragrunt finds it via
find_in_parent_folders("region.hcl", "") (with a not-found guard) and merges
its locals into every stack's inputs — so every stack in a region inherits the
metadata without per-tfvars duplication. global pseudo-regions set role = null
(and null FQDN parts). Each stack declares matching variable blocks with null
defaults.
role encodes the product invariant: when a partition spans multiple regions,
yucca-api + the database run only in the primary region; secondary regions
run the storage-local subset. It is authoritative here in TF state
(discovery.role) and consumed downstream (Flux role components, yuctl).
The op run --env-file=tf/.env -- pattern
tf/.env holds 1Password op:// references — not literal secrets:
export OP_SERVICE_ACCOUNT_TOKEN="op://yucca_tf_dev/yucca_futo_1pass_superuser_service_account/password"
Wrap every terragrunt invocation with op run --env-file=tf/.env -- (the
mise tf:* tasks do this automatically). The op CLI resolves the op://
reference and injects the actual token as OP_SERVICE_ACCOUNT_TOKEN into
the child process's environment. The 1P Terraform provider picks it up
from the env var and authenticates.
The same pattern is used in immich-app/devtools and is the Futo-wide
convention for TF secret injection.
Committed .env is safe because it's just pointers
Yucca's root .gitignore normally excludes .env files — we add an
explicit !tf/.env exception. This file contains only op:// URIs; no
secret ever transits the repo. It's a committed manifest of "which 1P
items this TF depends on."
Stack override via TF_STACK_DIR
The default mise run tf:* tasks target tf/deployment/staging/austin/ceph.
Point them at another stack via the TF_STACK_DIR env var:
TF_STACK_DIR=tf/deployment/staging/austin/ceph mise run tf:plan
TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply
Running TF
One-shot (preferred for now)
mise run tf:init # first time in a stack
mise run tf:plan # dry run
mise run tf:apply # render artifacts + (future) create 1P items
These wrap: op run --env-file=tf/.env -- terragrunt --working-dir <stack> <cmd>.
In CI (staging stacks)
.github/workflows/infra.yml runs the staging stacks from GitHub Actions:
- Plan on every PR touching
tf/**; apply on merge tomain, gated behind thestaging-infraEnvironment (required reviewers). - Applies
staging/austin/talos(cluster + Flux + secrets) thenstaging/global/dns. - The Talos stack reaches the
10.10.10.0/24nodes by joining the tailnet (tailscale/github-action,--accept-routes) — the cluster firewall already trusts the Tailscale CIDRs. The DNS stack is pure Cloudflare API, no tailnet. - Secrets come from the same
op run --env-file=tf/.envpath; CI just supplies the per-env 1P service-account token (the rest resolves from 1P). The token is injected asOP_SERVICE_ACCOUNT_TOKENfrom the environment-specific secret —OP_TF_YUCCA_STAGING_ENVhere (dev/prod workflows useOP_TF_YUCCA_DEV_ENV/OP_TF_YUCCA_PROD_ENV) — replacing a shared superuser SA with a scoped one.
Prerequisites (out-of-band): repo secret OP_TF_YUCCA_STAGING_ENV — a staging
1P service account mirroring the dev/prod ones, i.e. granted shared_tf,
shared_tf_staging, yucca_tf (read) + yucca_tf_staging (read/write for the
JWT-keypair item). Note: a copy of the dev SA token won't work — it can't read
yucca_tf_staging. Also: TS_OAUTH_CLIENT_ID, TS_OAUTH_SECRET; a Tailscale
subnet router advertising 10.10.10.0/24 with tag:project-yucca approved for
it; and the staging-infra Environment with required reviewers.
State backend
Remote: shared yucca-tf-state S3 bucket at OVH Paris
(https://s3.eu-west-par.io.cloud.ovh.net/). Key path:
yucca/${partition}/${region}/${stack}/terraform.tfstate (region-root stacks:
yucca/${partition}/${region}/...). All yucca stacks live under the yucca/
prefix in the shared bucket.
Credentials are AWS-compatible env vars (AWS_ACCESS_KEY_ID /
AWS_SECRET_ACCESS_KEY), injected via op run --env-file=tf/.env from
the TF_STATE_S3_* items in the yucca_tf vault. OVH-specific config
(skip AWS validation, path-style URLs) is in
deployment/terragrunt.hcl.
State locking is not enabled. OVH has no DynamoDB equivalent;
OpenTofu's use_lockfile = true option would handle single-bucket locking
but expects the lockfile object to already exist — terragrunt init
against a fresh backend fails with 404 before it can create one. Enable
it once the concern is concurrent applies (multiple operators working the
same stack simultaneously). Single operator today → low risk.
Discovery outputs contract
Every stack emits a single non-sensitive top-level discovery output (plus
discovery_schema_version). It is the machine-readable description of the
topology that yuctl consumes — it reads the state object straight from S3 and
parses .outputs.discovery.value, so no checkout, init, or provider is needed.
Secrets are always op:// references, never values. The in-stack sensitive
kubeconfig/talosconfig outputs stay for in-stack use; discovery only carries
the reference (op://<vault>/<title>/password).
Common envelope (every stack):
{
schema_version, partition, region, slug, role, stack, stack_type,
region_meta: { site_id, datacenter, provider_code, domain }
}
Per stack_type payload:
| stack_type | stacks | payload key + fields |
|---|---|---|
region-k8s |
*/talos |
kubernetes: { cluster_name, api_endpoint, operator_endpoint, cp_node_ips, kubeconfig_ref, talosconfig_ref } |
ceph |
*/ceph |
ceph_clusters: { <name> => { cluster_name, fqdn, rgw_s3_endpoint, health_cred_ref, s3_admin_cred_refs, secret_item_titles, bootstrap_host } } |
dns |
*/dns |
dns: { provider, zone, record_fqdns, api_token_ref } |
netbird |
*/netbird, prod/global |
netbird: { name_prefix, vault, group_ids, policy_ids, network_ids, setup_key_item_titles } |
fabric |
prod/htz-fsn1/fabric |
fabric: { site_id, kube_cidr, mgmt_cidr, cluster_cidrs } |
The talos stacks persist their kube/talosconfig into 1Password
(onepassword_item, mirroring the JWT-keypair / netbird-setup-key pattern) so
the *_ref fields resolve — staging writes YUCCA_STAGING_{KUBE,TALOS}CONFIG;
dev writes YUCCA_DEV_<CLUSTER>_{KUBE,TALOS}CONFIG.
Migrating an existing stack to the new layout
The dir moves + terragrunt rewrite change the state key (ceph/<env>/... →
yucca/<partition>/<region>/...), so a live stack's remote state must be
re-pointed — reversibly, no destroy/recreate. Per stack (single-operator window;
no state locking):
- Pre-flight
terragrunt planon the OLD layout → confirm no-op; back up the old state object. - Land the terragrunt.hcl rewrite + dir moves + renames (no apply yet).
- Server-side S3 copy old key → new
yucca/...key (old object stays as a rollback anchor). - Delete the stale generated
backend.tf;terragrunt initin the new dir (decline the state-copy prompt;-reconfigureif needed). terragrunt plan→ must be no-op (the gate; a non-empty plan = a rename/source-path mismatch, not a state problem — STOP + roll back).terragrunt output discovery→ confirm the contract resolves.- After all stacks are green,
aws s3 rmthe legacyceph/*objects.
Order: staging first, then prod (global → region → fabric — prod/global
before prod/htz-fsn1/netbird, which has a real terragrunt dependency on it).
dev is dir-move only (no remote state). The NetBird plans must show no
resource renames — the rendered names are byte-identical across the rename.
The ceph-cluster module
Declarative input in clusters.auto.tfvars:
clusters = {
sietch = {
domain = "staging.austin.int.futo.cloud"
partition = "staging"
region = "austin"
provider_code = "int"
role_in_hostname = "ceph"
ansible_ssh_user = "ansible-iac"
ansible_ssh_key = "~/.ssh/id_ed25519_sietch"
vault = "yucca_tf_staging"
provision_profile = "debian-live" # null for Hetzner-installimage clusters
hosts = [
{ name = "laurel", bond_ip = "10.10.10.90", bootstrap = true },
{ name = "lawson", bond_ip = "10.10.10.91" },
{ name = "samara", bond_ip = "10.10.10.92" },
]
}
}
On apply, the module:
- Picks wordlist names for
hosts[].name == null(stable across applies; seeded per-cluster; operator-declared names excluded from the pool to prevent collisions). - Renders
inventory.ini(normal ops),inventory-destroy.ini(explicit destroy flag),secrets.yml.tpl(op:// references toyucca_tf_staging/<CLUSTER>_CEPH_*_PASSWORD/password). Optionally rendersinventory-provision.iniwhenprovision_profile != null. - (Not yet TF-managed)
onepassword_itemresources for cluster secrets are dormant — items are created viaopCLI today and read by Ansible at play time. Seeansible/ceph/docs/secrets.mdfor the re-enable plan.
The talos-baremetal module (staging/austin/talos)
Brings up Talos on bare-metal nodes already running in maintenance mode at
known addresses — no Ansible, no hypervisors, no VLANs (the earlier VM-oriented
talos-cluster module was removed unused). It dials each node's maintenance IP, applies machine
config (which installs to disk + reboots), bootstraps one CP, then emits
kube/talosconfig and gates on cluster health.
Declarative input in deployment/staging/austin/talos/clusters.auto.tfvars:
clusters = {
yucca-staging = {
talos_version = "1.13.4"
kubernetes_version = "v1.36.1"
install_disk = "/dev/sda" # WIPED — the 240GB DELLBOSS; NVMe left raw
cluster_vip = "10.10.10.15" # L2 VIP, etcd-elected across CPs
gateway = "10.10.10.1"
subnet_cidr = "10.10.10.0/24"
cni = "cilium" # cni:none in Talos + Cilium via Helm
disable_kube_proxy = true # Cilium kube-proxy replacement (KubePrism)
cilium_version = "1.19.5"
hubble = true
bond = { interfaces = ["eno1np0", "eno2np1"], mode = "active-backup" } # flip to 802.3ad after the switch is LACP'd
nodes = [
{ name = "staging-cp1", address = "10.10.10.47" },
{ name = "staging-cp2", address = "10.10.10.242" },
{ name = "staging-cp3", address = "10.10.10.117" },
]
}
}
Notes:
- Static IP = maintenance IP. Each node's
addressis pinned as the static IP onbond0, so TF stays reachable across the install reboot. - bond comes up
active-backup(no switch config needed). Migrate to802.3adlater, node-by-node, after converting the switch ports to LACP port-channels — LACP needs both ends configured at once, so a big-bang flip drops connectivity until both sides agree. - Ingress firewall (
firewall.tf): default-deny + per-service allow-lists scoped to the subnet (+ pod CIDR on kubelet). apid + apiserver also trust the Tailscale ranges (trust_tailscale). ⚠️ The host runningtf applymust have a source IP inside an allowed range or apid (50000) is blocked and bootstrap hangs — add operator/jump subnets totrusted_cidrs. - CNI is installed in the same apply. With
cni:nonethe nodes are NotReady until Cilium lands, so the module's health gate runsskip_kubernetes_checks;helm.tfinstalls Cilium, then a second (full) health gate enforces Ready. - One cluster per stack. The helm provider binds to a single cluster
(
one(...)); add more clusters in their own stack.
Run it (see "Running TF" below — needs 1Password unlocked + an on-LAN apply host):
TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:init
TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:plan
TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply # WIPES /dev/sda, installs Talos
The netbird-env module (NetBird Cloud access control)
Manages one layer's NetBird Cloud footprint: groups,
access policies, device auth (setup) keys, and routed networks. One
NetBird Cloud account (api.netbird.io) backs everything — the module namespaces
every object <name_prefix>_<key> (all underscores) so all envs/sites coexist.
Group model
Per env (and per prod site), the baseline groups are:
| group | who | rendered (staging / prod htz-fsn1) |
|---|---|---|
ci |
ephemeral CI runners | yucca-staging-ci / yucca-prod-htz-fsn1-ci |
mgmt |
management nodes (configured via Ansible; also the route peers) | yucca-staging-mgmt / … |
talos |
Talos cluster nodes | yucca-staging-talos / … |
k8s_operator |
in-cluster kubernetes operator | yucca-staging-k8s-operator / … |
Logical keys (the tfvars map keys, e.g. ci) stay lowercase; rendered NetBird
names are UPPER_SNAKE (uppercased, hyphens → underscores). CI is per-env
(ci, reaching only that env's groups) — no cross-env CI plane. k8s is split
into talos (the nodes) and k8s_operator (the operator identity) so they can
carry different policies.
Stacks & layering
| stack | partition / scope | state key |
|---|---|---|
deployment/staging/global/netbird |
staging (global region) | yucca/staging/global/netbird/… |
deployment/prod/global/netbird |
prod, account-wide (cross-region) | yucca/prod/global/netbird/… |
deployment/prod/htz-fsn1/netbird |
prod, region htz-fsn1 | yucca/prod/htz-fsn1/netbird/… |
Staging is single-layer. Prod is layered: a global layer (reserved for
account-wide / cross-region groups + policies) above per-region layers. The
global layer is empty today — each region owns its own resource group and the
yucca → resources policy is module-generated per layer (see below). Region
groups are region-scoped (yucca-prod-<region>-<role>) so a network router's
peers are unambiguously that region's mgmt nodes. A region layer can still
consume a global group via a terragrunt dependency on prod/global → the
module's external_groups input, but none do today. The root terragrunt derives stack from the full sub-path,
so prod/htz-fsn1/netbird gets its own state key. (Now that the fabric stack
lives in its own prod/htz-fsn1/fabric/ sub-stack, the region root carries no
terragrunt.hcl, so prod/htz-fsn1/netbird uses a normal
find_in_parent_folders include — the old direct-root-include workaround is
gone.) Rendered NetBird object names and 1P item titles are unchanged by the
rename (YUCCA_STAGING_*, NETBIRD_YUCCA_PROD_HTZ_FSN1_*): the env→partition
/ site→region swap keeps the same string values.
The yucca → yucca-tags access model
yucca members reach everything tagged as a yucca tag (NetBird object names are
lowercase-kebab — e.g. yucca-prod-htz-fsn1-mgmt; the 1Password setup-key item
titles stay UPPER_SNAKE, since CI/ansible/talos read them by op:// string):
yucca— the existing users group (people). External (looked up by its actual nameyucca); never managed here.- yucca tags — any group flagged
resource = truein a layer'sgroups. This marks the group yucca-reachable; it applies to peer/node groups (so a yucca member can SSHyucca-prod-htz-fsn1-mgmt, the mgmt nodes) as well as routed-subnet tags (yucca-prod-htz-fsn1-resources, which the site'snetbird_network_resources are tagged into). Today every group a layer owns is flagged.
For each layer that owns ≥1 yucca tag, the netbird-env module auto-generates
a <prefix>-yucca-to-resources policy (bidirectional = false) whose
destinations are all of that layer's flagged groups — so flagging a group grants
yucca users access to it (its peers and any tagged resources) with no policy to
edit. The source is var.yucca_users_group (default yucca, looked up by name;
set null to opt a layer out). bidirectional = false means yucca users only
initiate — this policy never makes a tagged group a source. The union of these
per-layer policies is the account-wide "yucca reaches every yucca tag we create"
guarantee.
The former single shared
yucca_resourcetag inprod/global(one account-wideyucca → yucca_resourcepolicy, consumed by sites via a terragrunt dependency) was retired in favour of this per-layer model — no cross-stack group reference, and the destination list is derived, not hand-maintained.
Staging additionally grants its ci group access to the existing Liberty
Park infra groups (where the staging nodes live today) — those are external
groups resolved by name in staging/global/netbird/main.tf.
Declarative input (netbird.auto.tfvars)
Groups, setup keys, policies and networks reference groups by logical key, never opaque NetBird IDs:
# resource = true ⇒ yucca-reachable ("yucca tag"); here every group is flagged,
# so yucca users reach all of them (SSH the nodes + the routed subnets)
groups = { ci = { resource = true }, mgmt = { resource = true },
talos = { resource = true }, k8s_operator = { resource = true },
resources = { resource = true } }
setup_keys = {
ci = { type = "reusable", ephemeral = true, auto_groups = ["ci"] }
mgmt = { type = "reusable", auto_groups = ["mgmt"] }
talos = { type = "reusable", auto_groups = ["talos"] }
k8s_operator = { type = "reusable", auto_groups = ["k8s_operator"] }
}
policies = {
ci-to-all = { # CI reaches every node group in this env
rules = [{ name = "ci-to-all", protocol = "all"
sources = ["ci"], destinations = ["mgmt", "talos", "k8s_operator"] }]
}
}
NetBird is default-deny — a peer gets only the access its groups' policies
grant; an empty policies map means total isolation. The yucca → resource
policy is not declared here: the module generates it from every group flagged
resource = true (see the access model above).
Networks (prod htz-fsn1) — CIDRs propagated, not hardcoded
The htz-fsn1 site layer exposes a NetBird Network named htz-fsn1: the
mgmt group are the routing peers, and each routed subnet is a
netbird_network_resource. The CIDRs are derived from the same
fabric-addressing module the fabric stack uses (re-instantiated in the
layer's addressing.tf — a pure, stateless module, so no duplication and no
cross-stack coupling). Every resource is tagged into the site's own resources
group, so access is the module-generated yucca-prod-htz-fsn1-yucca-to-resources
policy. The only per-site input is the site id (the CIDRs flow from it):
site_id = 40 # mirrors prod/htz-fsn1; feeds fabric-addressing → the routed CIDRs
# mgmt 10.40.5.0/24 · api 10.40.10.0/24
# cls1_public 10.40.20.0/23 · cls1_private 10.40.22.0/23
Setup-key plaintext → 1Password. Each setup key's secret key is written to
the per-env vault (yucca_tf_<env>) as item
NETBIRD_<UPPERCASED_NAMESPACED_NAME>_SETUP_KEY (onepassword_item, same "TF
mints secrets into 1P" pattern as the JWT keypair). The namespaced title keeps
multiple prod sites writing to the one yucca_tf_prod vault from colliding.
Auth. Two providers, both fed by op run --env-file=tf/.env[.prod]:
netbird— admin PAT fromNB_PAT(op://shared_tf/NETBIRD_TF_PAT, shared across all envs;management_urldefaults to NetBird Cloud).onepassword—OP_SERVICE_ACCOUNT_TOKEN(same session), writes the keys.
Run it (pure cloud API — no tailnet, no node contact):
TF_STACK_DIR=tf/deployment/staging/global/netbird mise run tf:init # then tf:plan / tf:apply
# prod — global layer first, then each region layer (uses the prod env file + SA):
OP_ENV_FILE=tf/.env.prod TF_STACK_DIR=tf/deployment/prod/global/netbird mise run tf:apply
OP_ENV_FILE=tf/.env.prod TF_STACK_DIR=tf/deployment/prod/htz-fsn1/netbird mise run tf:apply
CI (.github/workflows/infra.yml) applies staging/global/netbird in the staging
matrix, and the prod layers (prod/global then prod/htz-fsn1/netbird) as gated
prod-infra jobs on the prod 1P SA / tf/.env.prod. Prod CI needs the
OP_TF_YUCCA_PROD_ENV[_WRITE] repo secrets + a prod-infra Environment — see
the workflow header.
CI connects over NetBird
CI reaches the staging 10.10.10.0/24 nodes over the NetBird overlay (this
replaced the Tailscale subnet-router path). The .github/actions/netbird-connect
composite action installs the client and runs netbird up with the ci setup
key read from 1P (op://yucca_tf_staging/NETBIRD_YUCCA_STAGING_CI_SETUP_KEY);
the runner joins as a ci peer and the existing staging route advertises the LAN.
The apply job applies staging/global/netbird first (minting that key) before
connecting, so a fresh bootstrap is self-contained. The prod fabric workflow
(fabric.yml) still uses Tailscale — 10.40.5.0/24 isn't on NetBird yet.
Where secrets actually live
yucca_tf_dev(team-shared): live values consumed by Ansible at play time. Password items per cluster (ops, dashboard, grafana, S3 svc-user access + secret), SSH Key items per cluster (ansible-iac keypairs), and DR-capture items per cluster (RGW TLS cert + key, client.admin keyring — populated bymise run capture).yucca_tf_dev_manual(team-shared): placeholders for human-fillable secrets (API tokens, OAuth client secrets). Not yet used by ceph-cluster.
Service accounts themselves are in yucca_tf_dev as two items:
| SA | Purpose | Consumed by |
|---|---|---|
yucca_futo_1pass_superuser_service_account |
Read + write all yucca_tf_* vaults |
TF (via tf/.env) |
yucca_futo_1pass_service_account |
Read-only on yucca_tf and yucca_tf_dev |
Ansible runtime / CI |
Both are shared with other Futo consumers (o11y, base Yucca infra).
Rotation affects all of them — see ansible/ceph/docs/runbooks/rotate-sa-token.md
for the coordination procedure.
Adding a new cluster
- Add an entry to
clusters.auto.tfvars. - Create the 1P items in
yucca_tf_dev:- Password items:
<CLUSTER>_CEPH_{OPS,DASHBOARD,GRAFANA}_PASSWORD, plus<CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_{ACCESS,SECRET}_KEY. Useop item create --generate-passwordfor each. - SSH Key item:
<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEYviaop item create --category "SSH Key" --ssh-generate-key=ed25519.
- Password items:
mise run tf:apply— renders inventory + secrets template.- Create per-node
host_vars/*.ymlfiles in the new inventory dir (hardware topology; not TF-rendered yet). - On operator workstation:
scripts/install-ssh-keys.sh <cluster>to pull the private key from 1P. - After first successful deploy:
mise run captureto snapshot the RGW TLS material + admin keyring to 1P for DR. - See
ansible/ceph/docs/adding-a-cluster.mdfor the full walk-through.
Related
- TF-first + op inject model:
ansible/ceph/docs/secrets.md - Immich devtools (upstream pattern): https://github.com/immich-app/devtools/tree/main/tf