mirror of
https://github.com/immich-app/yucca.git
synced 2026-09-30 13:33:00 +08:00
* feat(ceph): scaffold spice prod cluster (reprovision + inventory + TF) Stand up spice (48x Hetzner SX295, prod/htz-fsn1): the reprovision_hetzner ansible role (rescue -> reset -> installimage -> verify, base-OS-only, with stale-mdraid pre-clean and resume markers), the prod-htz-fsn1/spice inventory + group_vars/host_vars, the prod ceph TF stack, and gen-spice-host-vars.py. Adds a `mise reprovision` task (op run + tf/.env.prod), the spice SSH key (yucca_tf_prod), and per-host roles-based [ceph_mon] filtering so large clusters pin a fixed MON quorum instead of defaulting to all nodes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): create spice block.db + ssd-osd LVs at converge (NVMe-RAID) Add an NVMe-RAID branch to ceph_deploy lvm-setup so spice's single vg0 (built by installimage) gets its 14 block.db LVs + one ssd-osd LV created at converge, before osds.yml. This replaces painbox's fragile installimage -x chroot post-install with an idempotent, observable ansible step; installimage stays base-OS-only. Narrows the old blanket "externally-managed LVM" skip so it only fires when neither the sietch dual-SSD nor the NVMe-RAID shape applies, and wires ceph_db_vg / ceph_ssd_osd_lv / ceph_ssd_osd_reserve_gib in the spice group_vars. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): set spice bond NICs in group_vars, drop per-host placeholders The 25G fabric NICs are uniform across all 48 SX295 (verified live on 36 nodes): bond members enp193s0f0/f1 (Intel E800/ice, PCI c1:00.0/.1), WAN enp197s0 (igb, c5:00.0). Set bond_interfaces + oob_nic once in the spice group_vars -- the networkd role reads bond_interfaces -- and remove the dead per-host fabric_nic/oob_nic PLACEHOLDER lines from all 48 host_vars, spice-hosts.yaml, and gen-spice-host-vars.py. No MAC-based naming needed: the predictable names are PCI-derived and identical fleet-wide. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): LACP + VLAN sub-interfaces on the bond, lldpd fabric config networkd role: emit 802.3ad LACP params (TransmitHashPolicy, LACPTransmitRate) when bond_mode is 802.3ad, and support tagged VLAN sub-interfaces on the bond via networkd_bond_vlans -- when set, bond0 carries no L3 and each VLAN carries its own address. Opt-in: an empty list keeps the flat active-backup path (sietch) byte-identical (verified by render). verify.yml gates the bond0-IP/gateway asserts to the flat case and checks per-VLAN addresses otherwise. spice: bond0 becomes an 802.3ad LACP bond of enp193s0f0/f1 (MLAG to the QFX5240 leaves) carrying VLAN 120 (public 10.40.20.<idx>/23) + VLAN 122 (private 10.40.22.<idx>/23); default route stays on the 1G WAN until cutover. baseline role: install + configure lldpd (portid ifname, cluster system description, service enabled) when lldpd is in baseline_extra_packages, and add lldpd/ethtool/tcpdump for spice. This also newly applies lldpd config to sietch (already lists lldpd) on its next converge. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): set spice RGW pool to EC 16+4 host domain for the beta Capacity-max profile (80% usable, m=4) for the wipeable beta behind michael. Host failure domain accepts rack-loss risk in exchange for capacity, which is fine for a beta and matches the single-rack API tier. Also lifts the RGW DRAFT marker and starts the bulk EC data pool at pg_num 4096 (672 OSDs, 20-chunk pool) instead of the 36-OSD-era default to avoid PG splitting during the first fill. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * docs(ceph): drop stale -x chroot reference from spice autosetup header installimage.yml runs base-OS-only (installimage -a -c /autosetup, no -x). The autosetup header still carried the painbox `-x /tmp/post-install.sh` line, which implied a chroot post-install the role deliberately dropped. Point it at the convergence steps (lvm-setup, baseline) that replaced it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): minimal -x chroot to seed root + ansible-iac keys on reprovision installimage sometimes runs without our key present in the rescue (Robot armed rescue without it, in-place installimage, manual rescue), so its "copy the rescue authorized_keys" step leaves root keyless and the node unreachable for convergence. Add a minimal `-x /tmp/post-install.sh` that does ONLY inert user/key writes: seed root's authorized_keys with the iac key (baseline never manages root, so it persists) and create an ansible-iac account with the key + NOPASSWD sudo. No apt/LVM in the chroot -- that fragile step stays at post-boot convergence, so set -e cannot spuriously disrupt installimage. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): converge iac access on existing nodes without reimaging Nodes imaged before the reprovision -x chroot lack root's iac key and the ansible-iac sudo account; reseeding them meant a reinstall. baseline gains an iac-access task that ensures the same state idempotently -- root's authorized key (non-exclusive) plus the ansible-iac user with key and NOPASSWD sudo -- so the 36 already-installed nodes converge in place. Gated on the cluster iac key; run standalone with --tags iac_access. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): manual rescue-first add-node path, Robot API optional The Hetzner Robot rescue/reset step is both the unreliable part (nodes that arm rescue but never boot it) and the highest blast radius (an errant /reset on a live spice node in production). Gate it behind reprovision_use_robot (default true, preserving reprovision.yml) and add add-node.yml, which sets it false: the operator arms rescue by hand in the portal and Ansible runs only install, chroot and verify on a node already in rescue. Robot key registration moves to register_robot_keys.yml, imported from preflight only on the Robot path. With the API out of the loop nothing here can flip a running node into rescue -- wait_rescue refuses any node without the installimage ramdisk. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): relax stale-rescue guard on the manual add-node path wait_rescue rejects a rescue older than reprovision_rescue_max_uptime (1200s) to catch a stale boot -- valid only when Ansible triggered the reset. On the manual path the operator rescues by hand, so uptime just measures wait-before-run and a node legitimately in rescue for hours would be refused. Relax the guard to a day on add-node.yml; the installimage-ramdisk check is the real in-rescue proof. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): skip the this-boot rescue guard on the manual path philip has sat in rescue ~8 days, so any finite uptime ceiling rejects it. The guard only means something when Ansible triggered the reset (uptime proves this boot); on the manual add-node path a node may sit in rescue for weeks, so gate the assert on reprovision_use_robot rather than inflating the timeout. The installimage-ramdisk check remains the real in-rescue proof. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): networkd/ifupdown coexistence to protect the 1G WAN The role assumed a full migration off ifupdown -- correct for sietch's flat bond, but on spice it would disable ifupdown and leave the 1G WAN (a separate NIC that ifupdown owns, carrying the default route) unconfigured on the next boot: the painbox failure. Add networkd_replace_ifupdown (default true, sietch unchanged). When false, networkd manages only the bond and VLANs; an Unmanaged=yes guard fences the WAN NIC off, commit leaves networking.service enabled, the bond drops to RequiredForOnline=no so a carrier-less fabric cannot stall boot, and verify asserts the WAN kept its address and default route before anything commits. The rollback script targets the right interface per mode. spice sets it false. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make migrate-networkd usable pre-Ceph (gate health checks) The ceph-health pre/post gate assumed a live cluster and a ceph_bootstrap group, so the play could not run during spice bringup (coexistence activation before any cephadm deploy). Gate the four ceph tasks on ceph_bootstrap being populated; on a pre-Ceph cluster they skip and the networkd role runs on its own. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): tunable serial + halt-on-failure for migrate-networkd Fleet rollout of the coexistence activation wants batches, not one-at-a-time, and must stop the instant a node fails (a dropped WAN shows up as unreachable) rather than silently skip it. Template serial (networkd_serial, default 1 unchanged) and set max_fail_percentage 0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): install Intel ice DDP firmware so E810 NICs leave Safe Mode The minimal Debian image ships no firmware-misc-nonfree, so the E810 DDP package (intel/ice/ddp/ice.pkg) is absent and every NIC boots into Safe Mode -- whose crippled classifier drops reserved-multicast control frames (LACP, LLDP), so the 25G fabric bond never aggregates despite a correctly configured switch. Add firmware-misc-nonfree to the spice package set and a baseline nic-firmware task that reboots an E810 host once to load the DDP when it is still in Safe Mode (self-gating: no-op on non-ice or already-loaded hosts; refuses to reboot when the DDP is absent, so it cannot loop). The WAN is a separate igb NIC, so the reboot keeps the node reachable. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): correct ice DDP activation gating (pipefail + dpkg check) The Safe Mode check used `set -o pipefail`, but `ethtool --show-fec` exits non-zero on "not supported", which pipefail propagated and masked grep's match -- so the check always yielded "ok" and the activating reboot never fired. Drop pipefail there, and gate the reboot on the dpkg DB (firmware-misc-nonfree installed) rather than a file stat that can lag a large apt transaction. Verified on one node end-to-end: DDP loads, Safe Mode clears, LACP converges with the switch, both VLAN gateways ping. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): pin chrony time sources for Ceph time sync chrony was listed in baseline_enable_services but never installed, so system.yml would fail starting it -- and there was no time-source config at all, leaving sync to the distro default pool. Ceph mon quorum is skew-sensitive, so make it explicit: add baseline/chrony.yml (install + templated chrony.conf + enable) driven by chrony_ntp_servers (default Debian pool, makestep for the initial correction), drop chrony from baseline_enable_services so it is owned in one place, and point spice at Hetzner NTP (ntp1/2/3.hetzner.de) -- low-latency from FSN1, consistent across all 47. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): populate spice OSD by-path map; exclude noreen until recovered Inspected the SX295 disk layout: 14x 20TB SATA HDDs per node across 3 AHCI controllers, by-path uniform on 46/47 nodes -- so ceph_hdd_osds (path_phy + db vg0/db-slotN) and ceph_ssd_osds (vg0/ssd-osd) live in group_vars, not 47 host_vars. spice-ceph-miguel has one disk on 46:00.0-ata-4 rather than 87:00.0-ata-4 and overrides the map in its host_vars. Verified by rendering osd-spec.yml.j2: 658 HDD OSD paths (47x14) + 47 NVMe ssd-osd, miguel's override resolving correctly. noreen (boot-order casualty, held in triage) is commented out of clusters.auto.tfvars so the rendered inventory + deploy target only the 47 live nodes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): bootstrap mon on the fabric public_network, not the WAN cephadm bootstrap --mon-ip must fall inside public_network, but spice passed bond_ip, the 1G WAN address used for the ansible/SSH connection and cephadm host management. spice's public_network is the 25G fabric (10.40.20.0/23), so the initial mon would fail to bind. Bootstrap now uses ceph_public_ip (10.40.20.<host_index>, derived like the bond VLANs); sietch is flat with no host_index and falls back to bond_ip, already in its own public_network. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * ci(ceph): static validation gate for the ansible/ceph stack CI only validated the Kubernetes/Flux surface; nothing parsed the ceph playbooks, inventory or scripts, so a broken playbook or host_var first executed during the post-merge prod apply against real nodes. This runs the existing local validators (yamllint, ansible-lint, shellcheck, ansible-playbook --syntax-check, py_compile) on ansible/ceph PRs with no secrets and no connection to any host; a throwaway localhost inventory satisfies --syntax-check since the real inventory is TF-generated. Its own workflow, not a job in ci.yml, so the ansible/ceph/** path filter is workflow-scoped and does not gate the unrelated k8s jobs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make the ceph converge deliberate and cluster-parameterized A merge touching ansible/ceph/** could open the staging partition and auto-run the full baseline->tune->deploy->harden pipeline against the live sietch cluster, and the converge was hardcoded to sietch/austin (render-inventories.sh got no region, install-ssh-keys.sh got a literal sietch, CEPH_ENV pointed at .../sietch/), so a prod converge could never reach spice. The ceph converge now runs only on workflow_dispatch with run_ceph_ansible=true, pinned to the one matrix entry that owns the chosen ceph_cluster (spice=prod/htz-fsn1, sietch=staging/austin), so a push never reconverges a live cluster and one dispatch cannot converge both. Region and cluster are threaded from the dispatch inputs / matrix into the render, key install, and CEPH_ENV. The staging paths-filter is scoped to ansible/ceph/inventories/staging-** so a prod ceph change no longer opens the staging matrix. The ceph TF apply stays auto and env-gated; the mgmt converge is unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make the RGW firewall source scope configurable The RGW/S3 port was accepted from any source unconditionally, while every other Ceph service (mon, osd, dashboard, monitoring) is scoped to ceph_firewall_trusted_networks. On spice that port is plain HTTP 7480, so the S3 endpoint sat open to the public internet with no lever to close it. A new ceph_firewall_rgw_any_source (default true, mirroring ceph_firewall_ssh_any_source) keeps the open behavior by default but lets production restrict RGW to the trusted networks, which already cover the fabric and NetBird overlay. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): fix stale OSD-map comment on the noreen host_var noreen was generated before the by-path map moved into group_vars, so it still carried the DEFERRED note while its 47 siblings point at the group var. The host stays pre-staged for re-add once recovered; only the comment was wrong. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): add the host-mgmt VLAN (124) to the spice bond The FSN1-C1 fabric carries a third VLAN on the 25G bond, FSN1-C1-HOST-MGMT (124, 10.40.24.0/24), alongside Ceph public (120) and private (122). The leaf already trunks it to every server bond and advertises 10.40.24.0/24 into the NetBird overlay, so it is the intended in-band ansible/SSH reach once the WAN is retired, but the host side had no matching sub-interface. Add bond0.124 at 10.40.24.<host_index>/24 with no gateway, so the default route stays on the 1G WAN until cutover. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): apply networkd config deltas to an already-running daemon The networkd role only started systemd-networkd, which is a no-op once it is already running, so a re-run that added config (e.g. a new bond VLAN) wrote the .netdev/.network files but never applied them, and verify then failed asserting the sub-interface had no address. Add a networkctl reload plus a per-VLAN settle wait after the start, so a re-run creates the added sub-interfaces live without tearing down existing links; the WAN on ifupdown is untouched regardless. Non-disruptive on the flat sietch path (the settle wait is gated on networkd_bond_vlans). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): bind Ceph services on the fabric only, never the WAN Ceph on spice must never listen on the 1G WAN. mon/mgr/OSD already bind per public_network/cluster_network, but the beast RGW frontend and the cephadm monitoring stack bound 0.0.0.0, and the mon-ip, cephadm host-add address, and dashboard URLs advertised the WAN bond_ip. Add a per-cluster policy: ceph_service_ip (the address Ceph advertises and binds for point services; defaults to bond_ip, resolving to the fabric on spice via ceph_public_ip) and ceph_bind_networks (the cephadm networks: bind restriction for RGW and the monitoring daemons; defaults to public_network). Route mon-ip, the host-add address, the dashboard monitoring URLs, RGW, and prometheus/grafana/alertmanager/node-exporter/ ceph-exporter through them, and scope the RGW firewall to the trusted networks. The RGW zonegroup hostnames and TLS SAN gain the fabric address so signed requests to the fabric IP still validate. Flat clusters like sietch have no ceph_public_ip, so ceph_service_ip falls back to bond_ip and ceph_bind_networks to their single flat network; point services are unchanged there. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): define the metrics-worker RGW user on spice rgw.yml Step 14.5 creates a read-only RGW admin for the yucca-metrics-worker service and references ceph_rgw_metrics_user_*, which only sietch defined, so the RGW phase failed on spice with AnsibleUndefinedVariable. Add the block; the keys are the TF-minted SPICE_METRICS_WORKER_* items, injected via secrets.yml.tpl as vault_metrics_worker_*, so op inject resolves them. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): refuse to deploy against an un-rendered inventory The ceph_deploy role gates every phase on the ceph_bootstrap/ceph_mon/ ceph_join groups that render-inventories.sh emits from TF state; run against the hand-written stopgap inventory that lacks them, the role errors mid-deploy or falls through to placing a MON on every host. A pre-task assert now requires ceph_bootstrap to be exactly one host that is also in ceph_mon, and ceph_mon to be a non-empty odd quorum, and stops with a render hint otherwise. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make the fabric-only bind restriction opt-in per cluster ceph_bind_networks defaulted to public_network, which would have injected a networks: bind restriction into every cluster - re-binding sietch's live RGW and monitoring daemons from 0.0.0.0 to its flat network on the next converge (a redeploy, and a break for any access path not on that subnet). Default it empty instead: no networks: field is emitted and the monitoring re-spec is skipped, so flat clusters bind every interface exactly as cephadm ships them. spice opts in to the fabric public network in its group_vars. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make ceph_service_ip a group_var so hostvars can read it ceph_service_ip was a ceph_deploy role default, but join/rgw/monitoring read it as hostvars[<host>]['ceph_service_ip'], and role defaults are NOT exposed through hostvars - the lookup raised "HostVarsVars object has no attribute ceph_service_ip" and would abort the deploy on every cluster at the join phase (a regression the fabric-only change introduced for both spice and the live sietch). Define it in each cluster's group_vars instead (group_vars do resolve through hostvars, verified per host): spice to the fabric ceph_public_ip, sietch to bond_ip (unchanged from pre-fabric). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): drop the unused ceph_cluster_ip var on spice ceph_cluster_ip was defined for symmetry with ceph_public_ip but consumed nowhere - OSD replication binds to the cluster_network CIDR, which cephadm resolves to each node's bond0.122 address on its own. Remove the dead var and correct the neighbouring comment (ceph_service_ip is a group_var now, not a role default). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): serve spice RGW over TLS on 443 with a 10-year self-signed cert Match sietch: flip spice RGW/beast from plain HTTP 7480 to HTTPS 443 with a self-signed 10-year cert. rgw.yml generates /etc/ceph/rgw-ssl.{crt,key} and cephadm distributes it to every RGW daemon via the service spec; the SANs already cover s3.<domain> + the wildcard + each node's fabric IP (ceph_service_ip). The firewall RGW port derives from ceph_rgw_port, so it follows to 443, scoped to the trusted networks (RGW binds the fabric only, never the WAN). This also makes the discovery rgw_s3_endpoint (hardcoded https) correct for spice. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(dns): prod Cloudflare DNS stack for the spice RGW S3 endpoint Add tf/deployment/prod/global/dns (mirroring staging/global/dns) so s3.prod.fsn1.htz.futo.cloud and its wildcard resolve. Round-robin A across all 47 spice ceph nodes' fabric public IPs (10.40.20.<host_index>, VLAN 120) where RGW/beast binds; proxied=false (private RFC1918, reached over the NetBird-advertised cls1_public 10.40.20.0/23). The wildcard serves S3 virtual-hosted buckets, and both names are in the self-signed TLS cert SANs. noreen (host_index 40) excluded. CI auto-discovers the stack (applies at order 0 under the prod-global environment); tf/.env.prod gains the token ref. Next up: create op://yucca_tf_prod/CLOUDFLARE_API_TOKEN before the apply. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): preflight-assert the Ceph service IP is up before deploy A node whose fabric VLAN sub-interface (spice bond0.120) did not come back after a reboot would otherwise fail deep inside cephadm bootstrap/join. Add a per-host pre_task that asserts ceph_service_ip is present in ansible_all_ipv4_addresses, so a fabric-down node halts up front with a clear message. Passes on a healthy node (verified on spice-ceph-adelia); on sietch ceph_service_ip is bond_ip, the connection address, so it is trivially satisfied. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): honest ssd_model_pattern note + guard empty monitoring networks ssd_model_pattern on spice was "SAMSUNG" with a "verify on node 1" note, but the NVMe is SOLIDIGM/KIOXIA and the var is only read by ceph_destroy's partition-6 SSD wipe (sietch's dual-SSD shape); spice is NVMe-RAID with no partition 6 and a vg0 ssd-osd LV cleaned by the generic VG/PV removal, so the pattern is a deliberate no-match. Reword to say so (it must stay defined because cleanup.yml references it unconditionally). Also wrap the monitoring-spec networks block in a length guard so an empty ceph_bind_networks renders no dangling `networks:` key (defensive; the apply is already gated). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): lock down root/password SSH in baseline, not just harden installimage ships PermitRootLogin yes + a root password, and all sshd hardening lived in harden.yml (the last deploy stage), so a freshly imaged node sat root-password-open on the public WAN for the whole campaign (the 11-node exposure the swarm audit found). Add an early baseline task, right after iac-access authorizes the key on root, that deploys a 10-baseline-ssh drop-in (PermitRootLogin prohibit-password + PasswordAuthentication no, same values as security/50-hardening.conf so they never disagree), locks the root password, and removes the interim remediation drop-in. A Validate->Reload handler chain runs sshd -t before reloading, only on change. Key-safe on both clusters (ansible connects by key), so nothing can lock out. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): pin RGW metadata + mgr pools to the replicated size The RGW system-metadata pools (.rgw.root, {zone}.rgw.{log,control,meta}) and .mgr inherited the cluster-spec default osd_pool_default_size 2 / min_size 1, unlike the index/extra pools which are pinned. Losing two OSDs behind a metadata PG would take out the RGW/mgr control plane, and min_size 1 permits single-replica writes. Pin them to ceph_rgw_replicated_ size/_min_size (spice 3/2; sietch keeps 2/1 via its own vars), existence- gated like the pg_num loop and idempotent (only sets when the value differs). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make the ops password hash idempotent users.yml hashed ops_password with password_hash('sha512') and no salt, so a fresh random salt was drawn every run: the hash never matched /etc/shadow and the user module rewrote it reporting 'changed' on every converge (a clean converge was never a true green signal, on both clusters). Derive a stable salt from a one-way sha256 of the password so the hash is deterministic and idempotent, while still reconciling an out-of-band password change. The salt in /etc/shadow is public and leaks nothing. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make live fabric/hosts/timezone reproducible from the repo Three drift gaps where the running fleet did not match the committed IaC. networkd_enabled was committed for sietch only, so spice's fabric VLANs came up from an ad-hoc -e and a reprovisioned node would not bring them up from the repo alone; commit it true for spice (networkd_replace_ifupdown false still fences the 1G WAN). hosts.j2 mapped every node name to bond_ip (the 1G WAN) rather than the fabric ceph_service_ip, so in-cluster name resolution pointed off the fabric. And timezone: UTC was declared but never applied, leaving nodes on the image default (Europe/Berlin); add a community.general.timezone task to baseline/system.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): scope OSD/RGW readiness waits to the in-play host set join.yml is --limit-aware (it intersects ansible_play_hosts), but the OSD provisioning wait computed EXPECTED from all ceph_nodes and the RGW wait blocked until running rgw >= all ceph_nodes, so a subset (--limit) deploy joined the subset then deadlocked at both gates. Base both counts on groups['ceph_nodes'] intersect ansible_play_hosts, matching join.yml. On a full deploy the intersection is all nodes, so behavior is unchanged; the rgw-spec placement stays pinned to all ceph_nodes (only the readiness gate is play-scoped, so a --limit re-run cannot shrink the spec). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): scope the nftables ruleset to table inet filter nftables.conf.j2 did a global `flush ruleset`, and the no-op guard diffed the whole `nft list ruleset`. On a ceph node (cephadm runs podman) that wiped the podman/docker/fail2ban tables on every apply, and made the desired-vs-live diff never match (podman tables are live but absent in the throwaway netns), so the firewall reloaded on every converge - each one flushing podman again. Replace only table inet filter (add, delete, re-add), compare only that table in the guard, and reload via ExecReload (nft -f, no global flush) instead of restart (whose ExecStop flushes the whole ruleset). The fabric/ceph rules themselves are unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): deploy MGR at a capped count, not one per host placement.yml applied mgr to every active host (47 mgr daemons on spice: 1 active + 46 standby), which is wasteful and non-standard - MON already uses a bounded ceph_mon set. Deploy count:{{ ceph_mgr_count }} (default 3: 1 active + 2 standby) instead; cephadm schedules them and caps at the host count on small clusters like sietch. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): mark the RGW EC data pool as bulk The EC data pool is pre-seeded to pg_num 4096 (sized for spice's 658 HDD OSDs), but without the bulk flag the pg_autoscaler can walk it back toward 1 during the first fill, causing PG splitting under load. Set bulk=true so the autoscaler targets a full-capacity pg_num and treats the pre-seed as a floor. Idempotent: only set when not already true. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): harden the ceph_destroy SSD-model match The SSD OSD PV-remove and partition-wipe loops used grep -i "$ssd_model_pattern", which interpreted the pattern as a regex and, on an empty pattern, matched every disk - the worst-case foot-gun in a destroy path (it would target all disks). Use grep -iF (fixed string, no regex) and skip the loop entirely when the pattern is empty. Kept -F without -w, since -w would fail to match underscore-containing model strings like Micron_5100_MTF... Shared role, so it hardens sietch too. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): drop the dead s3_key_drift notice; reconcile harden header The "S3 key drift notice" debug task was gated on s3_key_drift, a variable nothing ever sets (the drift detection it stood in for was never built), so the branch could never fire - remove it. And update the harden.yml header: the security-critical sshd lockdown (root key-only, no password auth, locked root) now runs early in baseline, not gated behind this last stage; harden adds the firewall and the remaining sshd hardening on top. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): accurate change reporting in tuning roles (T4) CRUSH tunables and per-device HDD/SSD sysfs writes ran a mutating command every converge under changed_when:false, so a run reported ok even when it changed state and --check misled register consumers. Switch to get-then-set: read the current value first, only write when it differs, and emit CHANGED so changed_when reflects reality. Same settings applied; only change-detection becomes truthful. Also drops two pre-existing em-dashes to keep the file plain ASCII. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * refactor(ceph): declarative lvg/lvol for OSD block.db setup Replace hand-rolled pvcreate/vgcreate/lvcreate shell in ceph_deploy lvm-setup with community.general.lvg/lvol (already pinned, and the idiom provision_host/disks.yml already uses). Same VG/LV names, sizes, fixed-vs-100%FREE split, per-shape loops, and device targets on both shapes. Kept read-only asserts/verify/show tasks as-is. Preserved the wipefs -af signature clear ahead of PV creation: lvg does pvcreate -f but not --yes, so it will not wipe a stale foreign signature on a reused partition. wipefs stays gated on the VG being absent so a live PV is never touched. Kept a minimal shell to compute the spice ssd-osd size (vg0 free - reserve) since lvol has no free-minus-N primitive, plus an existence guard so re-runs don't recompute a bad size or trigger an unforced lvol shrink. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): ASCII-clean the lvm-setup header comment Convert the pre-existing arrows and em-dashes in the header (left untouched by the lvg/lvol refactor) to plain ASCII. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): never let lvol shrink a live block.db LV (data safety) The lvg/lvol refactor let community.general.lvol reconcile an existing fixed-size db-slot to its size, and lvol defaults to shrink:true - so a db-slot that ever drifted LARGER would be TRUNCATED, corrupting a live BlueStore block.db and taking out the OSD. The old shell skipped existing LVs entirely, so it could never do this. Set module_defaults shrink:false on both lvm-setup blocks: lvol still creates and may grow, but never truncates an existing LV (it leaves a larger one alone). Also restore opts:-Wy so (re)created LVs wipe stale signatures - the destroy->recreate recovery-path freshness the explicit -Wy gave. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): RGW self-signed cert generation aborted on first converge The openssl -subj and -addext args used backslash-newline continuations inside their quoted strings, so shell line-folding kept the next line's indentation in the value; countryName rendered as "DE " (over the 2-char max) and openssl exited 1 under set -euo pipefail, leaving no S3 endpoint on the first spice deploy. Precompute the subject and SAN as single-line facts (whitespace-controlled Jinja for the SAN) and pass them quoted, so each openssl arg is one flat string. See ansible/ceph/roles/ceph_deploy/tasks/rgw.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): green the ansible-lint CI gate on HEAD Bare ansible-lint (what mise run lint and the required CI check run) exited 2 on two deliberate patterns: run_once on the fleet-wide placement assert, and an intentional no-pipefail shell in the ice DDP Safe-Mode probe (pipefail there would mask grep's match). Waive both with inline noqa so the reasoned patterns stay and the gate passes. See ansible/ceph/deploy-ceph.yml and roles/baseline/tasks/nic-firmware.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): pin RGW pools to device-class CRUSH rules (SSD fast tier) The replicated_ssd/replicated_hdd rules were created but bound to no pool, so every replicated pool fell to the default class-agnostic rule and the omap-heavy bucket index landed on HDD (~98.5% of weight) while spice's 47-OSD NVMe ssd-osd tier sat idle. Pin the index, RGW metadata, and mgr pools to the SSD rule and the bulk non-EC pool to HDD; force the OSD device class at creation (ssd-osd -> ssd, HDD -> hdd) so the rules bind deterministically, and assert the placement in verify.yml. Pinning is opt-in per cluster (empty default) so a live cluster is never re-homed implicitly; the pin runs after RGW readiness because the .rgw.* system pools are created lazily by the realm/zone setup and the daemon. See ansible/ceph/roles/ceph_deploy/{templates/osd-spec.yml.j2,tasks/rgw.yml,tasks/verify.yml}. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): apply RGW pool PG/durability pins on the first converge The pool list was captured once at the top of the RGW play, before any pool was created, so every "pool in list" gate skipped on a fresh run 1: the data/index/non-EC PG pre-sizing deferred to a second converge, and the .rgw.* system pools (created lazily by the daemon) never got their size/min_size pin at all, sitting at the inherited 2/1 single-replica until someone ran the play twice. Re-query the pool list once the explicit pools exist, and move the system-pool PG + size/min_size pins below the RGW readiness wait where those pools are real. Parameterize the bootstrap osd_pool_default_size per cluster (spice 3/2, role default keeps sietch at 2/1) so an unpinned future pool is not born single-replica, and assert final size/min_size in verify.yml so a miss fails the play loudly. See ansible/ceph/roles/ceph_deploy/tasks/{rgw.yml,verify.yml}. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): real ok-to-stop gate before a live-node reimage Reimaging a spice node zeroes the OS NVMe, and the 14 HDD OSD block.db LVs plus the ssd-osd data LV all live on that NVMe, so a reimage destroys every OSD on the node; the old guard only checked a hand-typed i_know_its_dead flag and the prepare-os-disks comment wrongly claimed a reimage keeps OSD data. Replace the stub with a mon-delegated check that queries the node's OSD ids, refuses on HEALTH_ERR or a failed ceph osd ok-to-stop, and sets noout on those OSDs for the reimage window. The gate is default-on: auto mode enforces whenever a live cluster with this node's OSDs is reachable and no-ops for the initial bootstrap; strict fails closed if it cannot verify; permissive is the explicit escape hatch. See ansible/ceph/roles/reprovision_hetzner/tasks/ceph_safety.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): jumbo MTU 9000 on the VLAN 122 cluster/replication network Jumbo pays off on the OSD replication path (a closed, homogeneous fabric) but is a partial-blackhole risk on the RGW-facing public network, which serves heterogeneous 1500-byte S3 clients and the ~1400 NetBird overlay, so raise MTU on VLAN 122 (Ceph private) only and leave 120/124 at 1500. A VLAN opts into jumbo with a per-entry mtu; the bond ceiling and both 25G members auto-raise to the largest child MTU so a 1500 parent or slave cannot silently cap the jumbo frames. verify.yml asserts the applied MTU on the bond, members, and each VLAN (local, no switch dependency). Flat clusters resolve to 1500 and emit no MTU, so sietch is unchanged. Paired with a QFX-side change (leaf server-LAG + VLAN 122 IRB >= 9000) and a ping -M do -s 8972 matrix before the cluster network is relied on. See ansible/ceph/roles/networkd. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(dns): commit prod DNS provider lock file The prod DNS stack was the only DNS stack without a committed .terraform.lock.hcl, so cloudflare ~> 5.0 would resolve unpinned at the first prod apply of the stack fronting the S3 endpoint. Pin it to 5.21.0, matching the staging DNS stack, with linux_amd64 + darwin_arm64 hashes. See tf/deployment/prod/global/dns/.terraform.lock.hcl. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): opt-in alertmanager receiver so alerts reach a human cephadm deploys alertmanager with no receiver, so every Prometheus alert rule (OSD down, host down, PG degraded, near-full) fires into the default null route and nobody is notified. Add ceph_alertmanager_webhook_urls: a non-empty list routes all alerts to those webhooks via cephadm user_data.default_webhook_urls. The monitoring re-spec now applies when either fabric bind or alerting is configured (independent gates), and verify.yml reports the delivery status, warning loudly when none is set. Empty by default, so flat clusters (sietch) are unchanged; the spice destination is a deferred operator decision (TODO in defaults). See ansible/ceph/roles/ceph_deploy/tasks/monitoring.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(fabric): jumbo MTU on the leaf server bonds + VLAN 122 IRB Completes the host-side VLAN 122 jumbo change on the QFX leaf: the server LACP bonds (ae1..aeN) get the 9216 L2 jumbo MTU that the spine uplink ae0 already carries (802.3ad members inherit it, so no per-member mtu), and the private/cluster IRB gateway (VLAN 122) gets an L3 family-inet MTU of 9000 to match the Ceph hosts. Public (120) and host-mgmt (124) IRBs stay at 1500 by design, so jumbo is confined to the closed OSD-replication path. See tf/shared/modules/cluster-fabric. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ci): pin + SHA256-verify NetBird install, drop curl|sh The netbird-connect action piped pkgs.netbird.io/install.sh straight into sh on the runner that holds the prod-write 1Password token and overlay access to the live nodes, so a compromised installer would run as root there. Download a pinned release tarball (v0.74.4) from the GitHub release and verify its SHA256 against an in-repo pin before unpacking; fail closed on any mismatch. A version input plus a refresh comment keep the pin maintainable. See .github/actions/netbird-connect/action.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ci): close the prod apply gate (reviewed-plan binding + fail-closed reviewers) The prod approval gate could fail open and the apply was not bound to the reviewed diff. Manage the four gate Environments as code (new _meta/github stack: github_repository_environment + required reviewers, protected-branches-only, a validation that rejects a reviewerless gate), and make the discover job fail closed by asserting every emitted partition-region Environment already carries a required reviewer. Bind apply to the plan: the plan job writes -out to an absolute path and uploads it, the gated apply downloads that exact file and applies it with no re-plan and no -auto-approve, so a stale plan fails closed. Scope a bare workflow_dispatch to plan-only and a toggled dispatch to just the partition it touches; add fail-fast and timeout-minutes across the jobs. The _meta/github stack is bootstrap-applied out of band and needs real reviewer IDs plus a GH_ENV_ADMIN_TOKEN secret before the gate passes. See .github/workflows/infra.yml and tf/deployment/_meta/github. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): blast-radius control on the day-2 converge plays The converge chain (baseline, tune-os, tune-hardware, tune-ceph, harden) ran against all nodes at once with no canary and no stop-on-failure, the inverse of the serial:1/max_fail:0 destructive plays. Batch each converge play (serial, default 10%) with max_fail_percentage:0, and gate every batch on cluster health via a shared pre/post ceph-health checkpoint that halts the roll on HEALTH_ERR. The gate is a no-op until a cluster is live, so the bootstrap ordering (baseline -> tune -> deploy) is unaffected; deploy-ceph stays big-bang. Tune the batch size with ceph_converge_serial. See ansible/ceph/tasks/health_gate.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): freeze load-bearing host packages instead of auto-upgrading We run no unattended-upgrades: an apt upgrade that moved cephadm's container runtime (podman + ecosystem) or the chrony time source that mon quorum depends on out from under a live cluster is the uncoordinated change we avoid. Hold those packages at their installed version (dpkg selection) so apt upgrade skips them; upgrading is then deliberate and health-gated. Diagnostic/ops tools are left unheld. Runs last in baseline so nothing is held before it exists; tune via baseline_held_packages. See ansible/ceph/roles/baseline/tasks/hold-packages.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): THP=madvise and stop the GRUB cmdline clobber Transparent Huge Pages sat at the kernel default 'always', which bloats BlueStore/tcmalloc RSS and drives allocation-stall latency across the OSD fleet; set it to madvise via a oneshot unit ordered Before=ceph-osd.target (mirroring ceph-cpu-governor.service) plus a live sysfs write, on by default for every ceph cluster. Separately, the processor.max_cstate GRUB task rewrote GRUB_CMDLINE_LINUX_DEFAULT wholesale (dropping existing tokens, injecting quiet) - a latent footgun behind the default-off governor flag; replace it with a /etc/default/grub.d drop-in that only appends the cstate token. See ansible/ceph/roles/hardware_tuning. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): host security sysctls + persistent journald Add a host kernel/network hardening sysctl set (syncookies, ignore/deny ICMP redirects + source routing, kptr/dmesg restrict, log martians) to the existing os_tuning sysctl drop-in. rp_filter is set to LOOSE (2), not strict (1): these nodes are multi-homed (1G WAN default route + 25G fabric VLAN sub-interfaces), so strict reverse-path filtering would blackhole asymmetric fabric traffic. Separately make journald persistent (Storage=persistent, bounded SystemMaxUse/RuntimeMaxUse) so logs survive the pipeline's own reboots on these headless nodes. See ansible/ceph/roles/os_tuning. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): health-gated cephadm upgrade playbook + rollback runbook No upgrade automation existed for the pinned 20.2.x tentacle train. Add a deliberate, manually-invoked upgrade-ceph.yml: preflight-gates on HEALTH_OK (with an explicit allow-WARN toggle), no in-progress upgrade, and all OSDs up+in; records and prints the rollback image before starting; refuses to run without an explicit target (ceph_upgrade_target_image/_version, no default); drives ceph orch upgrade with a bounded poll and fails loudly on a stall; asserts health + version convergence after. Not wired into site.yml. Rollback procedure lives in the play header and docs/runbooks/upgrade-ceph.md. See ansible/ceph/upgrade-ceph.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): scheduled cluster-state backups with a systemd timer Config/topology state was captured only manually and controller-local. Add a ceph_backup role that installs a capture script + daily systemd timer on the bootstrap node: each run tars fsid, config dump, mon/osd/crush maps (raw + decoded), osd tree, orch ls/host ls, and the RGW realm/zonegroup/ zone into a root-only 0700 dir, pruned by retention. No secret keyrings are dumped. An offsite target (rsync or s3://) is a var left empty for the operator. Invoke via backup-ceph.yml (mise run backup-timer); not wired into converge. Opt out with ceph_backup_enabled=false. See ansible/ceph/roles/ceph_backup and docs/runbooks/backup-restore.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): pin mgr dashboard/prometheus bind to fabric; pin EC min_size The mgr dashboard (8443) and prometheus module (9283) run inside the active mgr and bound 0.0.0.0, closed only by nftables. Pin each mgr instance's own localized server_addr (mgr/<module>/<mgr_id>/server_addr, the key the modules read via get_localized_module_option) to that host's fabric IP, discovering mgr ids at runtime so it is failover-safe; a global server_addr would leave the dashboard unbindable after failover. Opt-in on ceph_bind_networks; takes effect on the next mgr cycle (nftables holds the WAN closed meanwhile). Separately pin the EC data pool min_size explicitly to k+1 (derived from ceph_rgw_ec_k) so the write-availability floor is documented and cannot drift, and record the single-site host-failure-domain DR ceiling in docs/capacity-planning.md. See ansible/ceph/roles/ceph_deploy/tasks/monitoring.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): freeze Galaxy collection versions to a major The collection requirements floated on >= lower bounds, so a new community.general major could be pulled mid-campaign. Pin compatible-release ranges to the currently-resolving majors (ansible.posix >=2,<3; community.general >=12,<13) so a 47-node campaign cannot cross a major between runs. No lockfile mechanism exists in the repo, so the ranges live in requirements.yml. See ansible/ceph/requirements.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * test(ceph): molecule scenario for the fabric-only firewall branch The only molecule scenario tested the default open-firewall branch (*_any_source: true), the opposite of spice's production posture, with no idempotence check. Add a fabric-only scenario that flips the closed branch (rgw/ssh any_source: false, trusted networks = fabric + NetBird) and asserts RGW/SSH are NOT accepted from any source, are restricted to the fabric/overlay, and that a second render is idempotent. Mirrors the default scenario's render-and-grep idiom; the existing scenario is untouched. See ansible/ceph/roles/security/molecule/fabric-only. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * ci(dns): drift-check the prod S3 roster against the ceph node list The 47 S3 RGW A-records were hand-copied with no link to the ceph roster, so add/replace/remove of a node silently blackholed the endpoint or dropped capacity. A terragrunt dependency is not viable here (no dependency idiom in the repo, and the ceph discovery output carries only bond_ip, never the fabric IP), so add a stdlib-only, credential-free check that reconstructs the expected fabric-IP roster from clusters.auto.tfvars (in-service names) + spice-hosts.yaml (host_index) and asserts the apex and wildcard A-records match exactly. A path-scoped CI gate fails the PR on drift before the DNS apply; also runnable via mise run tf:check-dns-roster. See tf/scripts/check-s3-dns-roster.py. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): inotify headroom for podman scale; podman-compat upgrade note cephadm-on-podman runs many containers/systemd units per node and leans on inotify; Debian's default fs.inotify.max_user_instances=128 is a known bottleneck at that scale, so raise instances to 512 and watches to 524288 in os_tuning. Separately, document in the upgrade runbook that the podman dpkg-hold must be lifted (unhold -> bump to a cephadm-supported version -> re-hold) before a cross-major Ceph upgrade; patch-train upgrades need no podman change. See ansible/ceph/roles/os_tuning and docs/runbooks/upgrade-ceph.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): health-checked RGW ingress VIP (haproxy + keepalived) S3 was bare round-robin DNS across 47 RGW A-records, so a dead node blackholed ~1/47 of new connections for the TTL. Add an opt-in cephadm ingress service: a keepalived floating VIP (spice 10.40.20.250) with haproxy health-checking the RGW backends and dropping a dead one from rotation. haproxy terminates TLS on the VIP with the existing self-signed RGW cert (its SAN now carries the VIP) and re-encrypts to beast: L4 passthrough for RGW needs IngressSpec.use_tcp_mode_over_rgw, which is absent on the pinned 20.2.2 (source-verified; applying it TypeErrors), so terminate is the only health-checked mode available. Applies after RGW readiness; asserts its preconditions (VIP set, ceph_bind_networks non-empty so beast binds a per-node IP and does not collide with the VIP on :443, ceph_rgw_ssl true). Opt-in per cluster; sietch unaffected. See ansible/ceph/roles/ceph_deploy/templates/rgw-ingress-spec.yaml.j2. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(dns): point the prod S3 endpoint at the ingress VIP Collapse the s3.prod.fsn1.htz.futo.cloud apex + wildcard from 47 round-robin A-records to the single health-checked ingress VIP 10.40.20.250, and switch the drift-check to assert both records equal that VIP (the ceph ceph_rgw_ingress_vip is the source of truth). Apply ordering matters and is noted in the tfvars: the ingress must be live before this DNS cutover, or s3 resolves to a VIP nothing answers; roll back in reverse. See tf/deployment/prod/global/dns/records.auto.tfvars. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): fix stale staging copy-paste comments in the prod stack; fmt The prod ceph stack carried the staging stack's comments verbatim: versions.tf and secrets.tf described creating SIETCH_CEPH_* items in yucca_tf_staging via OP_TF_YUCCA_STAGING_ENV_WRITE, and a variable example said (sietch, ...). This stack creates SPICE_CEPH_* in yucca_tf_prod via OP_TF_YUCCA_PROD_ENV_WRITE; correct the comments to match. Legitimate cross-refs (the mirrors-staging/talos provenance, the partition-slug enumeration) are left as-is. Also tofu fmt the drifted clusters.auto.tfvars (whitespace only). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * ci(infra): scope reviewer fail-closed gate to apply events; prettier-format workflows * ci(infra): pin 1Password CLI version to avoid flaky latest resolution * feat(dns): resolve prod S3 apex+wildcard to the node fabric-IP roster * ci(infra): park the Environments approval gate (disable fail-closed step + stub _meta/github) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
55 lines
2.2 KiB
YAML
55 lines
2.2 KiB
YAML
---
|
|
# ops user - human-interactive account.
|
|
# Convergeable: running this again fixes drift in password, sudo, shell.
|
|
|
|
- name: Ensure ops user exists
|
|
ansible.builtin.user:
|
|
name: "{{ baseline_ops_user }}"
|
|
uid: "{{ baseline_ops_uid }}"
|
|
shell: /bin/bash
|
|
groups: sudo
|
|
append: true
|
|
create_home: true
|
|
state: present
|
|
|
|
- name: Assert ops_password is defined and non-empty
|
|
ansible.builtin.assert:
|
|
that:
|
|
- ops_password is defined
|
|
- ops_password | length > 0
|
|
fail_msg: >-
|
|
ops_password is undefined or empty. Run the play via
|
|
scripts/ansible-play.sh (not bare ansible-playbook) so the cluster's
|
|
secrets.yml.tpl is resolved via op inject into --extra-vars. If that's
|
|
already what you did, verify vault_ops_password resolves:
|
|
`op read op://<vault>/<CLUSTER>_CEPH_OPS_PASSWORD/password`.
|
|
|
|
- name: Set ops user password
|
|
ansible.builtin.user:
|
|
name: "{{ baseline_ops_user }}"
|
|
# password_hash draws a RANDOM salt by default, so the hash differs every run
|
|
# and the user module rewrites /etc/shadow (reporting 'changed') on every
|
|
# converge. Derive a STABLE salt from the password (one-way sha256, so the
|
|
# public salt in /etc/shadow leaks nothing) to make the hash deterministic and
|
|
# the task idempotent, while still reconciling an out-of-band password change.
|
|
password: "{{ ops_password | password_hash('sha512', salt=(ops_password | hash('sha256'))[:16]) }}"
|
|
no_log: true
|
|
|
|
- name: Configure ops sudo
|
|
ansible.builtin.copy:
|
|
content: "{{ baseline_ops_user }} ALL=(ALL) {{ baseline_ops_sudo }}\n"
|
|
dest: "/etc/sudoers.d/{{ baseline_ops_user }}"
|
|
mode: '0440'
|
|
validate: "visudo -cf %s"
|
|
|
|
# Operator SSH access to the shared ops account, sourced from the identity
|
|
# registry (rendered into group_vars/all/operators.yml). exclusive: true makes
|
|
# the registry the single source of truth -- keys not in it are removed. Skipped
|
|
# when the list is empty so ops stays password-only rather than getting wiped.
|
|
- name: Authorize operator SSH keys on the ops account
|
|
ansible.posix.authorized_key:
|
|
user: "{{ baseline_ops_user }}"
|
|
key: "{{ ops_authorized_keys | join('\n') }}"
|
|
exclusive: true
|
|
when: ops_authorized_keys | length > 0
|