mirror of
https://github.com/immich-app/yucca.git
synced 2026-09-30 13:33:00 +08:00
feat(ceph): prod spice cluster on htz-fsn1 (#256)
* feat(ceph): scaffold spice prod cluster (reprovision + inventory + TF) Stand up spice (48x Hetzner SX295, prod/htz-fsn1): the reprovision_hetzner ansible role (rescue -> reset -> installimage -> verify, base-OS-only, with stale-mdraid pre-clean and resume markers), the prod-htz-fsn1/spice inventory + group_vars/host_vars, the prod ceph TF stack, and gen-spice-host-vars.py. Adds a `mise reprovision` task (op run + tf/.env.prod), the spice SSH key (yucca_tf_prod), and per-host roles-based [ceph_mon] filtering so large clusters pin a fixed MON quorum instead of defaulting to all nodes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): create spice block.db + ssd-osd LVs at converge (NVMe-RAID) Add an NVMe-RAID branch to ceph_deploy lvm-setup so spice's single vg0 (built by installimage) gets its 14 block.db LVs + one ssd-osd LV created at converge, before osds.yml. This replaces painbox's fragile installimage -x chroot post-install with an idempotent, observable ansible step; installimage stays base-OS-only. Narrows the old blanket "externally-managed LVM" skip so it only fires when neither the sietch dual-SSD nor the NVMe-RAID shape applies, and wires ceph_db_vg / ceph_ssd_osd_lv / ceph_ssd_osd_reserve_gib in the spice group_vars. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): set spice bond NICs in group_vars, drop per-host placeholders The 25G fabric NICs are uniform across all 48 SX295 (verified live on 36 nodes): bond members enp193s0f0/f1 (Intel E800/ice, PCI c1:00.0/.1), WAN enp197s0 (igb, c5:00.0). Set bond_interfaces + oob_nic once in the spice group_vars -- the networkd role reads bond_interfaces -- and remove the dead per-host fabric_nic/oob_nic PLACEHOLDER lines from all 48 host_vars, spice-hosts.yaml, and gen-spice-host-vars.py. No MAC-based naming needed: the predictable names are PCI-derived and identical fleet-wide. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): LACP + VLAN sub-interfaces on the bond, lldpd fabric config networkd role: emit 802.3ad LACP params (TransmitHashPolicy, LACPTransmitRate) when bond_mode is 802.3ad, and support tagged VLAN sub-interfaces on the bond via networkd_bond_vlans -- when set, bond0 carries no L3 and each VLAN carries its own address. Opt-in: an empty list keeps the flat active-backup path (sietch) byte-identical (verified by render). verify.yml gates the bond0-IP/gateway asserts to the flat case and checks per-VLAN addresses otherwise. spice: bond0 becomes an 802.3ad LACP bond of enp193s0f0/f1 (MLAG to the QFX5240 leaves) carrying VLAN 120 (public 10.40.20.<idx>/23) + VLAN 122 (private 10.40.22.<idx>/23); default route stays on the 1G WAN until cutover. baseline role: install + configure lldpd (portid ifname, cluster system description, service enabled) when lldpd is in baseline_extra_packages, and add lldpd/ethtool/tcpdump for spice. This also newly applies lldpd config to sietch (already lists lldpd) on its next converge. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): set spice RGW pool to EC 16+4 host domain for the beta Capacity-max profile (80% usable, m=4) for the wipeable beta behind michael. Host failure domain accepts rack-loss risk in exchange for capacity, which is fine for a beta and matches the single-rack API tier. Also lifts the RGW DRAFT marker and starts the bulk EC data pool at pg_num 4096 (672 OSDs, 20-chunk pool) instead of the 36-OSD-era default to avoid PG splitting during the first fill. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * docs(ceph): drop stale -x chroot reference from spice autosetup header installimage.yml runs base-OS-only (installimage -a -c /autosetup, no -x). The autosetup header still carried the painbox `-x /tmp/post-install.sh` line, which implied a chroot post-install the role deliberately dropped. Point it at the convergence steps (lvm-setup, baseline) that replaced it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): minimal -x chroot to seed root + ansible-iac keys on reprovision installimage sometimes runs without our key present in the rescue (Robot armed rescue without it, in-place installimage, manual rescue), so its "copy the rescue authorized_keys" step leaves root keyless and the node unreachable for convergence. Add a minimal `-x /tmp/post-install.sh` that does ONLY inert user/key writes: seed root's authorized_keys with the iac key (baseline never manages root, so it persists) and create an ansible-iac account with the key + NOPASSWD sudo. No apt/LVM in the chroot -- that fragile step stays at post-boot convergence, so set -e cannot spuriously disrupt installimage. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): converge iac access on existing nodes without reimaging Nodes imaged before the reprovision -x chroot lack root's iac key and the ansible-iac sudo account; reseeding them meant a reinstall. baseline gains an iac-access task that ensures the same state idempotently -- root's authorized key (non-exclusive) plus the ansible-iac user with key and NOPASSWD sudo -- so the 36 already-installed nodes converge in place. Gated on the cluster iac key; run standalone with --tags iac_access. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): manual rescue-first add-node path, Robot API optional The Hetzner Robot rescue/reset step is both the unreliable part (nodes that arm rescue but never boot it) and the highest blast radius (an errant /reset on a live spice node in production). Gate it behind reprovision_use_robot (default true, preserving reprovision.yml) and add add-node.yml, which sets it false: the operator arms rescue by hand in the portal and Ansible runs only install, chroot and verify on a node already in rescue. Robot key registration moves to register_robot_keys.yml, imported from preflight only on the Robot path. With the API out of the loop nothing here can flip a running node into rescue -- wait_rescue refuses any node without the installimage ramdisk. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): relax stale-rescue guard on the manual add-node path wait_rescue rejects a rescue older than reprovision_rescue_max_uptime (1200s) to catch a stale boot -- valid only when Ansible triggered the reset. On the manual path the operator rescues by hand, so uptime just measures wait-before-run and a node legitimately in rescue for hours would be refused. Relax the guard to a day on add-node.yml; the installimage-ramdisk check is the real in-rescue proof. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): skip the this-boot rescue guard on the manual path philip has sat in rescue ~8 days, so any finite uptime ceiling rejects it. The guard only means something when Ansible triggered the reset (uptime proves this boot); on the manual add-node path a node may sit in rescue for weeks, so gate the assert on reprovision_use_robot rather than inflating the timeout. The installimage-ramdisk check remains the real in-rescue proof. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): networkd/ifupdown coexistence to protect the 1G WAN The role assumed a full migration off ifupdown -- correct for sietch's flat bond, but on spice it would disable ifupdown and leave the 1G WAN (a separate NIC that ifupdown owns, carrying the default route) unconfigured on the next boot: the painbox failure. Add networkd_replace_ifupdown (default true, sietch unchanged). When false, networkd manages only the bond and VLANs; an Unmanaged=yes guard fences the WAN NIC off, commit leaves networking.service enabled, the bond drops to RequiredForOnline=no so a carrier-less fabric cannot stall boot, and verify asserts the WAN kept its address and default route before anything commits. The rollback script targets the right interface per mode. spice sets it false. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make migrate-networkd usable pre-Ceph (gate health checks) The ceph-health pre/post gate assumed a live cluster and a ceph_bootstrap group, so the play could not run during spice bringup (coexistence activation before any cephadm deploy). Gate the four ceph tasks on ceph_bootstrap being populated; on a pre-Ceph cluster they skip and the networkd role runs on its own. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): tunable serial + halt-on-failure for migrate-networkd Fleet rollout of the coexistence activation wants batches, not one-at-a-time, and must stop the instant a node fails (a dropped WAN shows up as unreachable) rather than silently skip it. Template serial (networkd_serial, default 1 unchanged) and set max_fail_percentage 0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): install Intel ice DDP firmware so E810 NICs leave Safe Mode The minimal Debian image ships no firmware-misc-nonfree, so the E810 DDP package (intel/ice/ddp/ice.pkg) is absent and every NIC boots into Safe Mode -- whose crippled classifier drops reserved-multicast control frames (LACP, LLDP), so the 25G fabric bond never aggregates despite a correctly configured switch. Add firmware-misc-nonfree to the spice package set and a baseline nic-firmware task that reboots an E810 host once to load the DDP when it is still in Safe Mode (self-gating: no-op on non-ice or already-loaded hosts; refuses to reboot when the DDP is absent, so it cannot loop). The WAN is a separate igb NIC, so the reboot keeps the node reachable. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): correct ice DDP activation gating (pipefail + dpkg check) The Safe Mode check used `set -o pipefail`, but `ethtool --show-fec` exits non-zero on "not supported", which pipefail propagated and masked grep's match -- so the check always yielded "ok" and the activating reboot never fired. Drop pipefail there, and gate the reboot on the dpkg DB (firmware-misc-nonfree installed) rather than a file stat that can lag a large apt transaction. Verified on one node end-to-end: DDP loads, Safe Mode clears, LACP converges with the switch, both VLAN gateways ping. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): pin chrony time sources for Ceph time sync chrony was listed in baseline_enable_services but never installed, so system.yml would fail starting it -- and there was no time-source config at all, leaving sync to the distro default pool. Ceph mon quorum is skew-sensitive, so make it explicit: add baseline/chrony.yml (install + templated chrony.conf + enable) driven by chrony_ntp_servers (default Debian pool, makestep for the initial correction), drop chrony from baseline_enable_services so it is owned in one place, and point spice at Hetzner NTP (ntp1/2/3.hetzner.de) -- low-latency from FSN1, consistent across all 47. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): populate spice OSD by-path map; exclude noreen until recovered Inspected the SX295 disk layout: 14x 20TB SATA HDDs per node across 3 AHCI controllers, by-path uniform on 46/47 nodes -- so ceph_hdd_osds (path_phy + db vg0/db-slotN) and ceph_ssd_osds (vg0/ssd-osd) live in group_vars, not 47 host_vars. spice-ceph-miguel has one disk on 46:00.0-ata-4 rather than 87:00.0-ata-4 and overrides the map in its host_vars. Verified by rendering osd-spec.yml.j2: 658 HDD OSD paths (47x14) + 47 NVMe ssd-osd, miguel's override resolving correctly. noreen (boot-order casualty, held in triage) is commented out of clusters.auto.tfvars so the rendered inventory + deploy target only the 47 live nodes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): bootstrap mon on the fabric public_network, not the WAN cephadm bootstrap --mon-ip must fall inside public_network, but spice passed bond_ip, the 1G WAN address used for the ansible/SSH connection and cephadm host management. spice's public_network is the 25G fabric (10.40.20.0/23), so the initial mon would fail to bind. Bootstrap now uses ceph_public_ip (10.40.20.<host_index>, derived like the bond VLANs); sietch is flat with no host_index and falls back to bond_ip, already in its own public_network. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * ci(ceph): static validation gate for the ansible/ceph stack CI only validated the Kubernetes/Flux surface; nothing parsed the ceph playbooks, inventory or scripts, so a broken playbook or host_var first executed during the post-merge prod apply against real nodes. This runs the existing local validators (yamllint, ansible-lint, shellcheck, ansible-playbook --syntax-check, py_compile) on ansible/ceph PRs with no secrets and no connection to any host; a throwaway localhost inventory satisfies --syntax-check since the real inventory is TF-generated. Its own workflow, not a job in ci.yml, so the ansible/ceph/** path filter is workflow-scoped and does not gate the unrelated k8s jobs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make the ceph converge deliberate and cluster-parameterized A merge touching ansible/ceph/** could open the staging partition and auto-run the full baseline->tune->deploy->harden pipeline against the live sietch cluster, and the converge was hardcoded to sietch/austin (render-inventories.sh got no region, install-ssh-keys.sh got a literal sietch, CEPH_ENV pointed at .../sietch/), so a prod converge could never reach spice. The ceph converge now runs only on workflow_dispatch with run_ceph_ansible=true, pinned to the one matrix entry that owns the chosen ceph_cluster (spice=prod/htz-fsn1, sietch=staging/austin), so a push never reconverges a live cluster and one dispatch cannot converge both. Region and cluster are threaded from the dispatch inputs / matrix into the render, key install, and CEPH_ENV. The staging paths-filter is scoped to ansible/ceph/inventories/staging-** so a prod ceph change no longer opens the staging matrix. The ceph TF apply stays auto and env-gated; the mgmt converge is unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make the RGW firewall source scope configurable The RGW/S3 port was accepted from any source unconditionally, while every other Ceph service (mon, osd, dashboard, monitoring) is scoped to ceph_firewall_trusted_networks. On spice that port is plain HTTP 7480, so the S3 endpoint sat open to the public internet with no lever to close it. A new ceph_firewall_rgw_any_source (default true, mirroring ceph_firewall_ssh_any_source) keeps the open behavior by default but lets production restrict RGW to the trusted networks, which already cover the fabric and NetBird overlay. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): fix stale OSD-map comment on the noreen host_var noreen was generated before the by-path map moved into group_vars, so it still carried the DEFERRED note while its 47 siblings point at the group var. The host stays pre-staged for re-add once recovered; only the comment was wrong. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): add the host-mgmt VLAN (124) to the spice bond The FSN1-C1 fabric carries a third VLAN on the 25G bond, FSN1-C1-HOST-MGMT (124, 10.40.24.0/24), alongside Ceph public (120) and private (122). The leaf already trunks it to every server bond and advertises 10.40.24.0/24 into the NetBird overlay, so it is the intended in-band ansible/SSH reach once the WAN is retired, but the host side had no matching sub-interface. Add bond0.124 at 10.40.24.<host_index>/24 with no gateway, so the default route stays on the 1G WAN until cutover. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): apply networkd config deltas to an already-running daemon The networkd role only started systemd-networkd, which is a no-op once it is already running, so a re-run that added config (e.g. a new bond VLAN) wrote the .netdev/.network files but never applied them, and verify then failed asserting the sub-interface had no address. Add a networkctl reload plus a per-VLAN settle wait after the start, so a re-run creates the added sub-interfaces live without tearing down existing links; the WAN on ifupdown is untouched regardless. Non-disruptive on the flat sietch path (the settle wait is gated on networkd_bond_vlans). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): bind Ceph services on the fabric only, never the WAN Ceph on spice must never listen on the 1G WAN. mon/mgr/OSD already bind per public_network/cluster_network, but the beast RGW frontend and the cephadm monitoring stack bound 0.0.0.0, and the mon-ip, cephadm host-add address, and dashboard URLs advertised the WAN bond_ip. Add a per-cluster policy: ceph_service_ip (the address Ceph advertises and binds for point services; defaults to bond_ip, resolving to the fabric on spice via ceph_public_ip) and ceph_bind_networks (the cephadm networks: bind restriction for RGW and the monitoring daemons; defaults to public_network). Route mon-ip, the host-add address, the dashboard monitoring URLs, RGW, and prometheus/grafana/alertmanager/node-exporter/ ceph-exporter through them, and scope the RGW firewall to the trusted networks. The RGW zonegroup hostnames and TLS SAN gain the fabric address so signed requests to the fabric IP still validate. Flat clusters like sietch have no ceph_public_ip, so ceph_service_ip falls back to bond_ip and ceph_bind_networks to their single flat network; point services are unchanged there. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): define the metrics-worker RGW user on spice rgw.yml Step 14.5 creates a read-only RGW admin for the yucca-metrics-worker service and references ceph_rgw_metrics_user_*, which only sietch defined, so the RGW phase failed on spice with AnsibleUndefinedVariable. Add the block; the keys are the TF-minted SPICE_METRICS_WORKER_* items, injected via secrets.yml.tpl as vault_metrics_worker_*, so op inject resolves them. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): refuse to deploy against an un-rendered inventory The ceph_deploy role gates every phase on the ceph_bootstrap/ceph_mon/ ceph_join groups that render-inventories.sh emits from TF state; run against the hand-written stopgap inventory that lacks them, the role errors mid-deploy or falls through to placing a MON on every host. A pre-task assert now requires ceph_bootstrap to be exactly one host that is also in ceph_mon, and ceph_mon to be a non-empty odd quorum, and stops with a render hint otherwise. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make the fabric-only bind restriction opt-in per cluster ceph_bind_networks defaulted to public_network, which would have injected a networks: bind restriction into every cluster - re-binding sietch's live RGW and monitoring daemons from 0.0.0.0 to its flat network on the next converge (a redeploy, and a break for any access path not on that subnet). Default it empty instead: no networks: field is emitted and the monitoring re-spec is skipped, so flat clusters bind every interface exactly as cephadm ships them. spice opts in to the fabric public network in its group_vars. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make ceph_service_ip a group_var so hostvars can read it ceph_service_ip was a ceph_deploy role default, but join/rgw/monitoring read it as hostvars[<host>]['ceph_service_ip'], and role defaults are NOT exposed through hostvars - the lookup raised "HostVarsVars object has no attribute ceph_service_ip" and would abort the deploy on every cluster at the join phase (a regression the fabric-only change introduced for both spice and the live sietch). Define it in each cluster's group_vars instead (group_vars do resolve through hostvars, verified per host): spice to the fabric ceph_public_ip, sietch to bond_ip (unchanged from pre-fabric). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): drop the unused ceph_cluster_ip var on spice ceph_cluster_ip was defined for symmetry with ceph_public_ip but consumed nowhere - OSD replication binds to the cluster_network CIDR, which cephadm resolves to each node's bond0.122 address on its own. Remove the dead var and correct the neighbouring comment (ceph_service_ip is a group_var now, not a role default). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): serve spice RGW over TLS on 443 with a 10-year self-signed cert Match sietch: flip spice RGW/beast from plain HTTP 7480 to HTTPS 443 with a self-signed 10-year cert. rgw.yml generates /etc/ceph/rgw-ssl.{crt,key} and cephadm distributes it to every RGW daemon via the service spec; the SANs already cover s3.<domain> + the wildcard + each node's fabric IP (ceph_service_ip). The firewall RGW port derives from ceph_rgw_port, so it follows to 443, scoped to the trusted networks (RGW binds the fabric only, never the WAN). This also makes the discovery rgw_s3_endpoint (hardcoded https) correct for spice. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(dns): prod Cloudflare DNS stack for the spice RGW S3 endpoint Add tf/deployment/prod/global/dns (mirroring staging/global/dns) so s3.prod.fsn1.htz.futo.cloud and its wildcard resolve. Round-robin A across all 47 spice ceph nodes' fabric public IPs (10.40.20.<host_index>, VLAN 120) where RGW/beast binds; proxied=false (private RFC1918, reached over the NetBird-advertised cls1_public 10.40.20.0/23). The wildcard serves S3 virtual-hosted buckets, and both names are in the self-signed TLS cert SANs. noreen (host_index 40) excluded. CI auto-discovers the stack (applies at order 0 under the prod-global environment); tf/.env.prod gains the token ref. Next up: create op://yucca_tf_prod/CLOUDFLARE_API_TOKEN before the apply. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): preflight-assert the Ceph service IP is up before deploy A node whose fabric VLAN sub-interface (spice bond0.120) did not come back after a reboot would otherwise fail deep inside cephadm bootstrap/join. Add a per-host pre_task that asserts ceph_service_ip is present in ansible_all_ipv4_addresses, so a fabric-down node halts up front with a clear message. Passes on a healthy node (verified on spice-ceph-adelia); on sietch ceph_service_ip is bond_ip, the connection address, so it is trivially satisfied. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): honest ssd_model_pattern note + guard empty monitoring networks ssd_model_pattern on spice was "SAMSUNG" with a "verify on node 1" note, but the NVMe is SOLIDIGM/KIOXIA and the var is only read by ceph_destroy's partition-6 SSD wipe (sietch's dual-SSD shape); spice is NVMe-RAID with no partition 6 and a vg0 ssd-osd LV cleaned by the generic VG/PV removal, so the pattern is a deliberate no-match. Reword to say so (it must stay defined because cleanup.yml references it unconditionally). Also wrap the monitoring-spec networks block in a length guard so an empty ceph_bind_networks renders no dangling `networks:` key (defensive; the apply is already gated). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): lock down root/password SSH in baseline, not just harden installimage ships PermitRootLogin yes + a root password, and all sshd hardening lived in harden.yml (the last deploy stage), so a freshly imaged node sat root-password-open on the public WAN for the whole campaign (the 11-node exposure the swarm audit found). Add an early baseline task, right after iac-access authorizes the key on root, that deploys a 10-baseline-ssh drop-in (PermitRootLogin prohibit-password + PasswordAuthentication no, same values as security/50-hardening.conf so they never disagree), locks the root password, and removes the interim remediation drop-in. A Validate->Reload handler chain runs sshd -t before reloading, only on change. Key-safe on both clusters (ansible connects by key), so nothing can lock out. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): pin RGW metadata + mgr pools to the replicated size The RGW system-metadata pools (.rgw.root, {zone}.rgw.{log,control,meta}) and .mgr inherited the cluster-spec default osd_pool_default_size 2 / min_size 1, unlike the index/extra pools which are pinned. Losing two OSDs behind a metadata PG would take out the RGW/mgr control plane, and min_size 1 permits single-replica writes. Pin them to ceph_rgw_replicated_ size/_min_size (spice 3/2; sietch keeps 2/1 via its own vars), existence- gated like the pg_num loop and idempotent (only sets when the value differs). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make the ops password hash idempotent users.yml hashed ops_password with password_hash('sha512') and no salt, so a fresh random salt was drawn every run: the hash never matched /etc/shadow and the user module rewrote it reporting 'changed' on every converge (a clean converge was never a true green signal, on both clusters). Derive a stable salt from a one-way sha256 of the password so the hash is deterministic and idempotent, while still reconciling an out-of-band password change. The salt in /etc/shadow is public and leaks nothing. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): make live fabric/hosts/timezone reproducible from the repo Three drift gaps where the running fleet did not match the committed IaC. networkd_enabled was committed for sietch only, so spice's fabric VLANs came up from an ad-hoc -e and a reprovisioned node would not bring them up from the repo alone; commit it true for spice (networkd_replace_ifupdown false still fences the 1G WAN). hosts.j2 mapped every node name to bond_ip (the 1G WAN) rather than the fabric ceph_service_ip, so in-cluster name resolution pointed off the fabric. And timezone: UTC was declared but never applied, leaving nodes on the image default (Europe/Berlin); add a community.general.timezone task to baseline/system.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): scope OSD/RGW readiness waits to the in-play host set join.yml is --limit-aware (it intersects ansible_play_hosts), but the OSD provisioning wait computed EXPECTED from all ceph_nodes and the RGW wait blocked until running rgw >= all ceph_nodes, so a subset (--limit) deploy joined the subset then deadlocked at both gates. Base both counts on groups['ceph_nodes'] intersect ansible_play_hosts, matching join.yml. On a full deploy the intersection is all nodes, so behavior is unchanged; the rgw-spec placement stays pinned to all ceph_nodes (only the readiness gate is play-scoped, so a --limit re-run cannot shrink the spec). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): scope the nftables ruleset to table inet filter nftables.conf.j2 did a global `flush ruleset`, and the no-op guard diffed the whole `nft list ruleset`. On a ceph node (cephadm runs podman) that wiped the podman/docker/fail2ban tables on every apply, and made the desired-vs-live diff never match (podman tables are live but absent in the throwaway netns), so the firewall reloaded on every converge - each one flushing podman again. Replace only table inet filter (add, delete, re-add), compare only that table in the guard, and reload via ExecReload (nft -f, no global flush) instead of restart (whose ExecStop flushes the whole ruleset). The fabric/ceph rules themselves are unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): deploy MGR at a capped count, not one per host placement.yml applied mgr to every active host (47 mgr daemons on spice: 1 active + 46 standby), which is wasteful and non-standard - MON already uses a bounded ceph_mon set. Deploy count:{{ ceph_mgr_count }} (default 3: 1 active + 2 standby) instead; cephadm schedules them and caps at the host count on small clusters like sietch. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): mark the RGW EC data pool as bulk The EC data pool is pre-seeded to pg_num 4096 (sized for spice's 658 HDD OSDs), but without the bulk flag the pg_autoscaler can walk it back toward 1 during the first fill, causing PG splitting under load. Set bulk=true so the autoscaler targets a full-capacity pg_num and treats the pre-seed as a floor. Idempotent: only set when not already true. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): harden the ceph_destroy SSD-model match The SSD OSD PV-remove and partition-wipe loops used grep -i "$ssd_model_pattern", which interpreted the pattern as a regex and, on an empty pattern, matched every disk - the worst-case foot-gun in a destroy path (it would target all disks). Use grep -iF (fixed string, no regex) and skip the loop entirely when the pattern is empty. Kept -F without -w, since -w would fail to match underscore-containing model strings like Micron_5100_MTF... Shared role, so it hardens sietch too. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): drop the dead s3_key_drift notice; reconcile harden header The "S3 key drift notice" debug task was gated on s3_key_drift, a variable nothing ever sets (the drift detection it stood in for was never built), so the branch could never fire - remove it. And update the harden.yml header: the security-critical sshd lockdown (root key-only, no password auth, locked root) now runs early in baseline, not gated behind this last stage; harden adds the firewall and the remaining sshd hardening on top. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): accurate change reporting in tuning roles (T4) CRUSH tunables and per-device HDD/SSD sysfs writes ran a mutating command every converge under changed_when:false, so a run reported ok even when it changed state and --check misled register consumers. Switch to get-then-set: read the current value first, only write when it differs, and emit CHANGED so changed_when reflects reality. Same settings applied; only change-detection becomes truthful. Also drops two pre-existing em-dashes to keep the file plain ASCII. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * refactor(ceph): declarative lvg/lvol for OSD block.db setup Replace hand-rolled pvcreate/vgcreate/lvcreate shell in ceph_deploy lvm-setup with community.general.lvg/lvol (already pinned, and the idiom provision_host/disks.yml already uses). Same VG/LV names, sizes, fixed-vs-100%FREE split, per-shape loops, and device targets on both shapes. Kept read-only asserts/verify/show tasks as-is. Preserved the wipefs -af signature clear ahead of PV creation: lvg does pvcreate -f but not --yes, so it will not wipe a stale foreign signature on a reused partition. wipefs stays gated on the VG being absent so a live PV is never touched. Kept a minimal shell to compute the spice ssd-osd size (vg0 free - reserve) since lvol has no free-minus-N primitive, plus an existence guard so re-runs don't recompute a bad size or trigger an unforced lvol shrink. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): ASCII-clean the lvm-setup header comment Convert the pre-existing arrows and em-dashes in the header (left untouched by the lvg/lvol refactor) to plain ASCII. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): never let lvol shrink a live block.db LV (data safety) The lvg/lvol refactor let community.general.lvol reconcile an existing fixed-size db-slot to its size, and lvol defaults to shrink:true - so a db-slot that ever drifted LARGER would be TRUNCATED, corrupting a live BlueStore block.db and taking out the OSD. The old shell skipped existing LVs entirely, so it could never do this. Set module_defaults shrink:false on both lvm-setup blocks: lvol still creates and may grow, but never truncates an existing LV (it leaves a larger one alone). Also restore opts:-Wy so (re)created LVs wipe stale signatures - the destroy->recreate recovery-path freshness the explicit -Wy gave. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): RGW self-signed cert generation aborted on first converge The openssl -subj and -addext args used backslash-newline continuations inside their quoted strings, so shell line-folding kept the next line's indentation in the value; countryName rendered as "DE " (over the 2-char max) and openssl exited 1 under set -euo pipefail, leaving no S3 endpoint on the first spice deploy. Precompute the subject and SAN as single-line facts (whitespace-controlled Jinja for the SAN) and pass them quoted, so each openssl arg is one flat string. See ansible/ceph/roles/ceph_deploy/tasks/rgw.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): green the ansible-lint CI gate on HEAD Bare ansible-lint (what mise run lint and the required CI check run) exited 2 on two deliberate patterns: run_once on the fleet-wide placement assert, and an intentional no-pipefail shell in the ice DDP Safe-Mode probe (pipefail there would mask grep's match). Waive both with inline noqa so the reasoned patterns stay and the gate passes. See ansible/ceph/deploy-ceph.yml and roles/baseline/tasks/nic-firmware.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): pin RGW pools to device-class CRUSH rules (SSD fast tier) The replicated_ssd/replicated_hdd rules were created but bound to no pool, so every replicated pool fell to the default class-agnostic rule and the omap-heavy bucket index landed on HDD (~98.5% of weight) while spice's 47-OSD NVMe ssd-osd tier sat idle. Pin the index, RGW metadata, and mgr pools to the SSD rule and the bulk non-EC pool to HDD; force the OSD device class at creation (ssd-osd -> ssd, HDD -> hdd) so the rules bind deterministically, and assert the placement in verify.yml. Pinning is opt-in per cluster (empty default) so a live cluster is never re-homed implicitly; the pin runs after RGW readiness because the .rgw.* system pools are created lazily by the realm/zone setup and the daemon. See ansible/ceph/roles/ceph_deploy/{templates/osd-spec.yml.j2,tasks/rgw.yml,tasks/verify.yml}. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): apply RGW pool PG/durability pins on the first converge The pool list was captured once at the top of the RGW play, before any pool was created, so every "pool in list" gate skipped on a fresh run 1: the data/index/non-EC PG pre-sizing deferred to a second converge, and the .rgw.* system pools (created lazily by the daemon) never got their size/min_size pin at all, sitting at the inherited 2/1 single-replica until someone ran the play twice. Re-query the pool list once the explicit pools exist, and move the system-pool PG + size/min_size pins below the RGW readiness wait where those pools are real. Parameterize the bootstrap osd_pool_default_size per cluster (spice 3/2, role default keeps sietch at 2/1) so an unpinned future pool is not born single-replica, and assert final size/min_size in verify.yml so a miss fails the play loudly. See ansible/ceph/roles/ceph_deploy/tasks/{rgw.yml,verify.yml}. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ceph): real ok-to-stop gate before a live-node reimage Reimaging a spice node zeroes the OS NVMe, and the 14 HDD OSD block.db LVs plus the ssd-osd data LV all live on that NVMe, so a reimage destroys every OSD on the node; the old guard only checked a hand-typed i_know_its_dead flag and the prepare-os-disks comment wrongly claimed a reimage keeps OSD data. Replace the stub with a mon-delegated check that queries the node's OSD ids, refuses on HEALTH_ERR or a failed ceph osd ok-to-stop, and sets noout on those OSDs for the reimage window. The gate is default-on: auto mode enforces whenever a live cluster with this node's OSDs is reachable and no-ops for the initial bootstrap; strict fails closed if it cannot verify; permissive is the explicit escape hatch. See ansible/ceph/roles/reprovision_hetzner/tasks/ceph_safety.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): jumbo MTU 9000 on the VLAN 122 cluster/replication network Jumbo pays off on the OSD replication path (a closed, homogeneous fabric) but is a partial-blackhole risk on the RGW-facing public network, which serves heterogeneous 1500-byte S3 clients and the ~1400 NetBird overlay, so raise MTU on VLAN 122 (Ceph private) only and leave 120/124 at 1500. A VLAN opts into jumbo with a per-entry mtu; the bond ceiling and both 25G members auto-raise to the largest child MTU so a 1500 parent or slave cannot silently cap the jumbo frames. verify.yml asserts the applied MTU on the bond, members, and each VLAN (local, no switch dependency). Flat clusters resolve to 1500 and emit no MTU, so sietch is unchanged. Paired with a QFX-side change (leaf server-LAG + VLAN 122 IRB >= 9000) and a ping -M do -s 8972 matrix before the cluster network is relied on. See ansible/ceph/roles/networkd. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(dns): commit prod DNS provider lock file The prod DNS stack was the only DNS stack without a committed .terraform.lock.hcl, so cloudflare ~> 5.0 would resolve unpinned at the first prod apply of the stack fronting the S3 endpoint. Pin it to 5.21.0, matching the staging DNS stack, with linux_amd64 + darwin_arm64 hashes. See tf/deployment/prod/global/dns/.terraform.lock.hcl. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): opt-in alertmanager receiver so alerts reach a human cephadm deploys alertmanager with no receiver, so every Prometheus alert rule (OSD down, host down, PG degraded, near-full) fires into the default null route and nobody is notified. Add ceph_alertmanager_webhook_urls: a non-empty list routes all alerts to those webhooks via cephadm user_data.default_webhook_urls. The monitoring re-spec now applies when either fabric bind or alerting is configured (independent gates), and verify.yml reports the delivery status, warning loudly when none is set. Empty by default, so flat clusters (sietch) are unchanged; the spice destination is a deferred operator decision (TODO in defaults). See ansible/ceph/roles/ceph_deploy/tasks/monitoring.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(fabric): jumbo MTU on the leaf server bonds + VLAN 122 IRB Completes the host-side VLAN 122 jumbo change on the QFX leaf: the server LACP bonds (ae1..aeN) get the 9216 L2 jumbo MTU that the spine uplink ae0 already carries (802.3ad members inherit it, so no per-member mtu), and the private/cluster IRB gateway (VLAN 122) gets an L3 family-inet MTU of 9000 to match the Ceph hosts. Public (120) and host-mgmt (124) IRBs stay at 1500 by design, so jumbo is confined to the closed OSD-replication path. See tf/shared/modules/cluster-fabric. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ci): pin + SHA256-verify NetBird install, drop curl|sh The netbird-connect action piped pkgs.netbird.io/install.sh straight into sh on the runner that holds the prod-write 1Password token and overlay access to the live nodes, so a compromised installer would run as root there. Download a pinned release tarball (v0.74.4) from the GitHub release and verify its SHA256 against an in-repo pin before unpacking; fail closed on any mismatch. A version input plus a refresh comment keep the pin maintainable. See .github/actions/netbird-connect/action.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * fix(ci): close the prod apply gate (reviewed-plan binding + fail-closed reviewers) The prod approval gate could fail open and the apply was not bound to the reviewed diff. Manage the four gate Environments as code (new _meta/github stack: github_repository_environment + required reviewers, protected-branches-only, a validation that rejects a reviewerless gate), and make the discover job fail closed by asserting every emitted partition-region Environment already carries a required reviewer. Bind apply to the plan: the plan job writes -out to an absolute path and uploads it, the gated apply downloads that exact file and applies it with no re-plan and no -auto-approve, so a stale plan fails closed. Scope a bare workflow_dispatch to plan-only and a toggled dispatch to just the partition it touches; add fail-fast and timeout-minutes across the jobs. The _meta/github stack is bootstrap-applied out of band and needs real reviewer IDs plus a GH_ENV_ADMIN_TOKEN secret before the gate passes. See .github/workflows/infra.yml and tf/deployment/_meta/github. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): blast-radius control on the day-2 converge plays The converge chain (baseline, tune-os, tune-hardware, tune-ceph, harden) ran against all nodes at once with no canary and no stop-on-failure, the inverse of the serial:1/max_fail:0 destructive plays. Batch each converge play (serial, default 10%) with max_fail_percentage:0, and gate every batch on cluster health via a shared pre/post ceph-health checkpoint that halts the roll on HEALTH_ERR. The gate is a no-op until a cluster is live, so the bootstrap ordering (baseline -> tune -> deploy) is unaffected; deploy-ceph stays big-bang. Tune the batch size with ceph_converge_serial. See ansible/ceph/tasks/health_gate.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): freeze load-bearing host packages instead of auto-upgrading We run no unattended-upgrades: an apt upgrade that moved cephadm's container runtime (podman + ecosystem) or the chrony time source that mon quorum depends on out from under a live cluster is the uncoordinated change we avoid. Hold those packages at their installed version (dpkg selection) so apt upgrade skips them; upgrading is then deliberate and health-gated. Diagnostic/ops tools are left unheld. Runs last in baseline so nothing is held before it exists; tune via baseline_held_packages. See ansible/ceph/roles/baseline/tasks/hold-packages.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): THP=madvise and stop the GRUB cmdline clobber Transparent Huge Pages sat at the kernel default 'always', which bloats BlueStore/tcmalloc RSS and drives allocation-stall latency across the OSD fleet; set it to madvise via a oneshot unit ordered Before=ceph-osd.target (mirroring ceph-cpu-governor.service) plus a live sysfs write, on by default for every ceph cluster. Separately, the processor.max_cstate GRUB task rewrote GRUB_CMDLINE_LINUX_DEFAULT wholesale (dropping existing tokens, injecting quiet) - a latent footgun behind the default-off governor flag; replace it with a /etc/default/grub.d drop-in that only appends the cstate token. See ansible/ceph/roles/hardware_tuning. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): host security sysctls + persistent journald Add a host kernel/network hardening sysctl set (syncookies, ignore/deny ICMP redirects + source routing, kptr/dmesg restrict, log martians) to the existing os_tuning sysctl drop-in. rp_filter is set to LOOSE (2), not strict (1): these nodes are multi-homed (1G WAN default route + 25G fabric VLAN sub-interfaces), so strict reverse-path filtering would blackhole asymmetric fabric traffic. Separately make journald persistent (Storage=persistent, bounded SystemMaxUse/RuntimeMaxUse) so logs survive the pipeline's own reboots on these headless nodes. See ansible/ceph/roles/os_tuning. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): health-gated cephadm upgrade playbook + rollback runbook No upgrade automation existed for the pinned 20.2.x tentacle train. Add a deliberate, manually-invoked upgrade-ceph.yml: preflight-gates on HEALTH_OK (with an explicit allow-WARN toggle), no in-progress upgrade, and all OSDs up+in; records and prints the rollback image before starting; refuses to run without an explicit target (ceph_upgrade_target_image/_version, no default); drives ceph orch upgrade with a bounded poll and fails loudly on a stall; asserts health + version convergence after. Not wired into site.yml. Rollback procedure lives in the play header and docs/runbooks/upgrade-ceph.md. See ansible/ceph/upgrade-ceph.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): scheduled cluster-state backups with a systemd timer Config/topology state was captured only manually and controller-local. Add a ceph_backup role that installs a capture script + daily systemd timer on the bootstrap node: each run tars fsid, config dump, mon/osd/crush maps (raw + decoded), osd tree, orch ls/host ls, and the RGW realm/zonegroup/ zone into a root-only 0700 dir, pruned by retention. No secret keyrings are dumped. An offsite target (rsync or s3://) is a var left empty for the operator. Invoke via backup-ceph.yml (mise run backup-timer); not wired into converge. Opt out with ceph_backup_enabled=false. See ansible/ceph/roles/ceph_backup and docs/runbooks/backup-restore.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): pin mgr dashboard/prometheus bind to fabric; pin EC min_size The mgr dashboard (8443) and prometheus module (9283) run inside the active mgr and bound 0.0.0.0, closed only by nftables. Pin each mgr instance's own localized server_addr (mgr/<module>/<mgr_id>/server_addr, the key the modules read via get_localized_module_option) to that host's fabric IP, discovering mgr ids at runtime so it is failover-safe; a global server_addr would leave the dashboard unbindable after failover. Opt-in on ceph_bind_networks; takes effect on the next mgr cycle (nftables holds the WAN closed meanwhile). Separately pin the EC data pool min_size explicitly to k+1 (derived from ceph_rgw_ec_k) so the write-availability floor is documented and cannot drift, and record the single-site host-failure-domain DR ceiling in docs/capacity-planning.md. See ansible/ceph/roles/ceph_deploy/tasks/monitoring.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): freeze Galaxy collection versions to a major The collection requirements floated on >= lower bounds, so a new community.general major could be pulled mid-campaign. Pin compatible-release ranges to the currently-resolving majors (ansible.posix >=2,<3; community.general >=12,<13) so a 47-node campaign cannot cross a major between runs. No lockfile mechanism exists in the repo, so the ranges live in requirements.yml. See ansible/ceph/requirements.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * test(ceph): molecule scenario for the fabric-only firewall branch The only molecule scenario tested the default open-firewall branch (*_any_source: true), the opposite of spice's production posture, with no idempotence check. Add a fabric-only scenario that flips the closed branch (rgw/ssh any_source: false, trusted networks = fabric + NetBird) and asserts RGW/SSH are NOT accepted from any source, are restricted to the fabric/overlay, and that a second render is idempotent. Mirrors the default scenario's render-and-grep idiom; the existing scenario is untouched. See ansible/ceph/roles/security/molecule/fabric-only. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * ci(dns): drift-check the prod S3 roster against the ceph node list The 47 S3 RGW A-records were hand-copied with no link to the ceph roster, so add/replace/remove of a node silently blackholed the endpoint or dropped capacity. A terragrunt dependency is not viable here (no dependency idiom in the repo, and the ceph discovery output carries only bond_ip, never the fabric IP), so add a stdlib-only, credential-free check that reconstructs the expected fabric-IP roster from clusters.auto.tfvars (in-service names) + spice-hosts.yaml (host_index) and asserts the apex and wildcard A-records match exactly. A path-scoped CI gate fails the PR on drift before the DNS apply; also runnable via mise run tf:check-dns-roster. See tf/scripts/check-s3-dns-roster.py. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): inotify headroom for podman scale; podman-compat upgrade note cephadm-on-podman runs many containers/systemd units per node and leans on inotify; Debian's default fs.inotify.max_user_instances=128 is a known bottleneck at that scale, so raise instances to 512 and watches to 524288 in os_tuning. Separately, document in the upgrade runbook that the podman dpkg-hold must be lifted (unhold -> bump to a cephadm-supported version -> re-hold) before a cross-major Ceph upgrade; patch-train upgrades need no podman change. See ansible/ceph/roles/os_tuning and docs/runbooks/upgrade-ceph.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(ceph): health-checked RGW ingress VIP (haproxy + keepalived) S3 was bare round-robin DNS across 47 RGW A-records, so a dead node blackholed ~1/47 of new connections for the TTL. Add an opt-in cephadm ingress service: a keepalived floating VIP (spice 10.40.20.250) with haproxy health-checking the RGW backends and dropping a dead one from rotation. haproxy terminates TLS on the VIP with the existing self-signed RGW cert (its SAN now carries the VIP) and re-encrypts to beast: L4 passthrough for RGW needs IngressSpec.use_tcp_mode_over_rgw, which is absent on the pinned 20.2.2 (source-verified; applying it TypeErrors), so terminate is the only health-checked mode available. Applies after RGW readiness; asserts its preconditions (VIP set, ceph_bind_networks non-empty so beast binds a per-node IP and does not collide with the VIP on :443, ceph_rgw_ssl true). Opt-in per cluster; sietch unaffected. See ansible/ceph/roles/ceph_deploy/templates/rgw-ingress-spec.yaml.j2. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * feat(dns): point the prod S3 endpoint at the ingress VIP Collapse the s3.prod.fsn1.htz.futo.cloud apex + wildcard from 47 round-robin A-records to the single health-checked ingress VIP 10.40.20.250, and switch the drift-check to assert both records equal that VIP (the ceph ceph_rgw_ingress_vip is the source of truth). Apply ordering matters and is noted in the tfvars: the ingress must be live before this DNS cutover, or s3 resolves to a VIP nothing answers; roll back in reverse. See tf/deployment/prod/global/dns/records.auto.tfvars. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * chore(ceph): fix stale staging copy-paste comments in the prod stack; fmt The prod ceph stack carried the staging stack's comments verbatim: versions.tf and secrets.tf described creating SIETCH_CEPH_* items in yucca_tf_staging via OP_TF_YUCCA_STAGING_ENV_WRITE, and a variable example said (sietch, ...). This stack creates SPICE_CEPH_* in yucca_tf_prod via OP_TF_YUCCA_PROD_ENV_WRITE; correct the comments to match. Legitimate cross-refs (the mirrors-staging/talos provenance, the partition-slug enumeration) are left as-is. Also tofu fmt the drifted clusters.auto.tfvars (whitespace only). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK * ci(infra): scope reviewer fail-closed gate to apply events; prettier-format workflows * ci(infra): pin 1Password CLI version to avoid flaky latest resolution * feat(dns): resolve prod S3 apex+wildcard to the node fabric-IP roster * ci(infra): park the Environments approval gate (disable fail-closed step + stub _meta/github) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
26b58feb15
commit
0357aac743
@@ -20,15 +20,62 @@ inputs:
|
||||
description: Optional peer hostname to register the runner as.
|
||||
required: false
|
||||
default: ''
|
||||
version:
|
||||
description: >-
|
||||
Pinned NetBird client version to install. MUST match the in-repo SHA256
|
||||
pins in the install step below (bump both together).
|
||||
required: false
|
||||
default: '0.74.4'
|
||||
|
||||
runs:
|
||||
using: composite
|
||||
steps:
|
||||
- name: Install NetBird client
|
||||
# Install a PINNED NetBird release and verify its SHA256 against an in-repo
|
||||
# pin before executing anything -- never `curl | sh`. This runner holds the
|
||||
# prod-write 1Password SA token and overlay access to the live nodes, so a
|
||||
# compromised pkgs.netbird.io/install.sh would run as root here.
|
||||
#
|
||||
# To bump: pick a version, then read the published checksums and copy the
|
||||
# `netbird_<ver>_linux_<arch>.tar.gz` lines into SHA256 below:
|
||||
# curl -fsSL \
|
||||
# https://github.com/netbirdio/netbird/releases/download/v<ver>/netbird_<ver>_checksums.txt
|
||||
- name: Install NetBird client (pinned + SHA256-verified)
|
||||
shell: bash
|
||||
env:
|
||||
NB_VERSION: ${{ inputs.version }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
curl -fsSL https://pkgs.netbird.io/install.sh | sh
|
||||
|
||||
# Pinned SHA256 of netbird_${NB_VERSION}_linux_<arch>.tar.gz.
|
||||
declare -A SHA256=(
|
||||
[amd64]=b854710409e7de79071642e2c328b181d4db7b9c90c298ee32b6e3b5d9ebd36f
|
||||
[arm64]=af6fc4909bcc68e55197eddbecd8a8ec1b02bc9b34cc28ef72e2858375e138d9
|
||||
)
|
||||
|
||||
case "$(uname -m)" in
|
||||
x86_64) arch=amd64 ;;
|
||||
aarch64) arch=arm64 ;;
|
||||
*) echo "unsupported CPU arch: $(uname -m)" >&2; exit 1 ;;
|
||||
esac
|
||||
want="${SHA256[$arch]:-}"
|
||||
[ -n "$want" ] || { echo "no pinned SHA256 for arch '$arch'" >&2; exit 1; }
|
||||
|
||||
tarball="netbird_${NB_VERSION}_linux_${arch}.tar.gz"
|
||||
url="https://github.com/netbirdio/netbird/releases/download/v${NB_VERSION}/${tarball}"
|
||||
|
||||
tmp="$(mktemp -d)"; trap 'rm -rf "$tmp"' EXIT
|
||||
curl -fsSL --retry 3 --proto '=https' --tlsv1.2 -o "$tmp/$tarball" "$url"
|
||||
|
||||
# Fail closed on any mismatch BEFORE the artifact is unpacked or run.
|
||||
echo "${want} ${tmp}/${tarball}" | sha256sum -c -
|
||||
|
||||
tar -xzf "$tmp/$tarball" -C "$tmp" netbird
|
||||
sudo install -m 0755 "$tmp/netbird" /usr/local/bin/netbird
|
||||
|
||||
# Register + start the daemon (what the upstream installer would do), so
|
||||
# the later `netbird up` has a service to talk to.
|
||||
sudo netbird service install 2>/dev/null || true
|
||||
sudo netbird service start 2>/dev/null || true
|
||||
netbird version
|
||||
|
||||
- name: Connect to NetBird
|
||||
|
||||
@@ -0,0 +1,91 @@
|
||||
name: ansible-ceph-validate
|
||||
|
||||
# Static validation gate for the ansible/ceph stack. ci.yml only validates the
|
||||
# Kubernetes/Flux surface (mise run k8s:validate); nothing there parses the ceph
|
||||
# playbooks, inventory or scripts. Without this gate a broken playbook, host_var
|
||||
# or script arg first executes during the post-merge PROD apply against real
|
||||
# nodes. This job runs the existing local validators (yamllint + ansible-lint +
|
||||
# shellcheck + syntax-check) with NO secrets and NO connection to any host.
|
||||
#
|
||||
# It lives in its own workflow (not a job in ci.yml) so the path filter below is
|
||||
# workflow-scoped: `on.<event>.paths` gates the whole file. Adding the same
|
||||
# filter inside ci.yml would gate every ci.yml job (k8s validate, tests) too.
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
paths:
|
||||
- 'ansible/ceph/**'
|
||||
- '.github/workflows/ansible-ceph-validate.yml'
|
||||
push:
|
||||
branches: [main]
|
||||
paths:
|
||||
- 'ansible/ceph/**'
|
||||
- '.github/workflows/ansible-ceph-validate.yml'
|
||||
|
||||
# Read-only: checkout + tool downloads only. No writes through the token, no
|
||||
# deploy credentials, no 1Password.
|
||||
permissions:
|
||||
contents: read
|
||||
|
||||
concurrency:
|
||||
group: ansible-ceph-validate-${{ github.ref }}
|
||||
cancel-in-progress: true
|
||||
|
||||
jobs:
|
||||
validate:
|
||||
name: Validate ansible/ceph (lint + syntax)
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 15
|
||||
defaults:
|
||||
run:
|
||||
working-directory: ansible/ceph
|
||||
steps:
|
||||
- uses: actions/checkout@1af3b93b6815bc44a9784bd300feb67ff0d1eeb3 # v6.0.0
|
||||
with:
|
||||
persist-credentials: false
|
||||
|
||||
- name: Setup Mise
|
||||
uses: immich-app/devtools/actions/use-mise@cd24790a7f5f6439ac32cc94f5523cb2de8bfa8c # use-mise-action-v1.1.0
|
||||
env:
|
||||
# mise downloads tools from GitHub releases; anonymous requests hit
|
||||
# API rate limits.
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
|
||||
# Trust the nested ansible/ceph/.mise.toml and install its pinned Python
|
||||
# (3.12). The use-mise action installs root tools from the repo root; the
|
||||
# ceph config is only picked up here, where cwd is ansible/ceph.
|
||||
- name: Trust and install mise tools
|
||||
run: |
|
||||
mise trust --all
|
||||
mise install
|
||||
|
||||
# Create the project .venv exactly as `mise run setup` does: pip install
|
||||
# requirements.txt (pins ansible-core, ansible-lint, yamllint) and the
|
||||
# ansible-galaxy collections (ansible.posix, community.general) that
|
||||
# ansible-lint / syntax-check need to resolve modules. PyPI + Galaxy only;
|
||||
# no secrets.
|
||||
- name: Install ansible tooling
|
||||
run: mise run setup
|
||||
|
||||
# yamllint + ansible-lint + shellcheck (scripts/*.sh). None need a live
|
||||
# inventory or secrets; shellcheck ships on ubuntu-latest runners.
|
||||
- name: Lint (yamllint + ansible-lint + shellcheck)
|
||||
run: mise run lint
|
||||
|
||||
# ansible-playbook --syntax-check needs *an* inventory, but the real
|
||||
# inventory.ini is TF-generated and gitignored, so it does not exist in
|
||||
# CI. A syntax-check only parses YAML/role structure, so any valid
|
||||
# inventory works: write a throwaway localhost inventory and point CEPH_ENV
|
||||
# at it (the ceph:check task honours an exported CEPH_ENV). This never
|
||||
# connects to a host.
|
||||
- name: Syntax-check playbooks
|
||||
env:
|
||||
CEPH_ENV: /tmp/ci-inventory.ini
|
||||
run: |
|
||||
printf '[all]\nlocalhost ansible_connection=local\n' > /tmp/ci-inventory.ini
|
||||
mise run check
|
||||
|
||||
# Byte-compile the helper scripts so a broken generator is caught here,
|
||||
# not mid-provision. (shellcheck above covers the *.sh scripts.)
|
||||
- name: Compile Python scripts
|
||||
run: mise exec -- python -m py_compile scripts/*.py
|
||||
+260
-53
@@ -28,13 +28,19 @@ name: Infra (Terraform)
|
||||
# advertises the node subnets the overlay-joining jobs reach; prod/global must
|
||||
# precede prod/htz-fsn1/netbird (a real terragrunt dependency).
|
||||
#
|
||||
# The two post-apply Ansible converges (ceph RGW users on the ceph stack; mgmt
|
||||
# hosts on the fabric stack) are gated independently of their TF apply: on push
|
||||
# they run only when their own surface changed (ceph: ansible/ceph/**; mgmt:
|
||||
# ansible/mgmt/** + tf/render/ansible-mgmt/** + .mise/tasks/mgmt/**), and on
|
||||
# workflow_dispatch only when the run_ceph_ansible / run_mgmt_ansible toggles are
|
||||
# set (both default off). A pure-Terraform change to either stack thus applies
|
||||
# without reconverging the nodes — and the ceph overlay join is skipped with it.
|
||||
# The two post-apply Ansible converges are gated independently of their TF apply.
|
||||
# They differ by blast radius:
|
||||
# - ceph (full baseline -> tune -> deploy -> harden against a LIVE bare-metal
|
||||
# cluster) is DELIBERATE: it runs ONLY on workflow_dispatch with
|
||||
# run_ceph_ansible=true, targeting the single cluster named by the
|
||||
# ceph_cluster input (spice = prod/htz-fsn1, sietch = staging/austin). A push
|
||||
# NEVER converges a ceph cluster, so merging a ceph change (prod- or
|
||||
# staging-scoped) can never auto-run an Ansible converge on a live cluster.
|
||||
# - mgmt (mgmt hosts on the fabric stack) still auto-runs on push when its own
|
||||
# surface changed (ansible/mgmt/** + tf/render/ansible-mgmt/** +
|
||||
# .mise/tasks/mgmt/**), and on workflow_dispatch when run_mgmt_ansible is set.
|
||||
# A pure-Terraform change to either stack thus applies without reconverging the
|
||||
# nodes, and the matching overlay join is skipped with it.
|
||||
#
|
||||
# The fabric (switch) apply is gated the same way: on push it applies only when the
|
||||
# fabric surface changed (tf/deployment/prod/htz-fsn1/fabric/**, the fabric shared
|
||||
@@ -45,13 +51,26 @@ name: Infra (Terraform)
|
||||
#
|
||||
# Environment gates rekey to <partition>-<region> (one per stack's region):
|
||||
# staging-austin, staging-global, prod-global, prod-htz-fsn1. Each matrix apply
|
||||
# entry references its own gate, so an unprovisioned Environment hangs the apply.
|
||||
# entry references its own gate. GitHub does NOT hang on an unprovisioned
|
||||
# Environment -- it auto-creates it on first reference with ZERO protection
|
||||
# rules, so a brand-new <partition>-<region> would apply with -auto-approve and
|
||||
# no required reviewer (fail OPEN). The intended fix has TWO parts, currently PARKED:
|
||||
# 1. The Environments are managed as code (tf/deployment/_meta/github, the
|
||||
# integrations/github provider: github_repository_environment + required
|
||||
# reviewers) so every gate exists WITH reviewers by construction.
|
||||
# 2. The `discover` job FAILS CLOSED (a "required reviewers" step) on apply-bearing
|
||||
# events, refusing to proceed unless each Environment carries a reviewer.
|
||||
# BOTH are DISABLED for now (the _meta/github stack is stubbed and the discover step
|
||||
# is `if: false`): enforcing them requires a one-time bootstrap (GH_ENV_ADMIN_TOKEN
|
||||
# secret + the four Environments created with reviewers) that would otherwise break
|
||||
# the first push-to-main. Until that blast radius is redesigned, apply gates behave
|
||||
# as GitHub's default (auto-create UNPROTECTED). Tracked for follow-up.
|
||||
#
|
||||
# Prerequisites (provisioned out-of-band):
|
||||
# - Repo secrets: OP_TF_YUCCA_STAGING_ENV (+ _WRITE); OP_TF_YUCCA_PROD_ENV
|
||||
# (read) + OP_TF_YUCCA_PROD_ENV_WRITE (netbird apply / write escalation source).
|
||||
# - GitHub Environments with required reviewers: staging-austin, staging-global,
|
||||
# prod-global, prod-htz-fsn1.
|
||||
# - GitHub Environments with required reviewers (staging-austin, staging-global,
|
||||
# prod-global, prod-htz-fsn1): DEFERRED -- the approval gate is parked (above).
|
||||
# - BOOTSTRAP — the netbird stacks applied ONCE out-of-band so the CI/mgmt setup
|
||||
# keys exist in 1P before anything tries to connect (CI can't mint them itself:
|
||||
# the apply that mints them is gated behind the plan that needs them). E.g.:
|
||||
@@ -80,9 +99,16 @@ on:
|
||||
workflow_dispatch:
|
||||
inputs:
|
||||
run_ceph_ansible:
|
||||
description: 'Run the Ceph Ansible converge (RGW users) after the ceph apply'
|
||||
description: 'Run the Ceph Ansible converge (full pipeline) against ceph_cluster'
|
||||
type: boolean
|
||||
default: false
|
||||
ceph_cluster:
|
||||
description: 'Ceph cluster the converge targets (only used when run_ceph_ansible is true)'
|
||||
type: choice
|
||||
options:
|
||||
- sietch
|
||||
- spice
|
||||
default: sietch
|
||||
run_mgmt_ansible:
|
||||
description: 'Run the mgmt Ansible converge after the fabric apply'
|
||||
type: boolean
|
||||
@@ -93,7 +119,10 @@ on:
|
||||
default: false
|
||||
|
||||
# Serialize: the OVH S3 backend has no state locking (single-operator model),
|
||||
# so never let two infra runs apply concurrently.
|
||||
# so never let two infra runs apply concurrently. This workflow-wide concurrency
|
||||
# group IS the single-operator assertion that stands in for a backend lock: at
|
||||
# most one infra run holds it at a time, and cancel-in-progress:false lets a
|
||||
# running apply finish rather than being interrupted mid-state-write.
|
||||
concurrency:
|
||||
group: ${{ github.workflow }}
|
||||
cancel-in-progress: false
|
||||
@@ -108,14 +137,15 @@ jobs:
|
||||
# Skip on fork PRs (no access to secrets / the overlay anyway).
|
||||
if: github.event_name != 'pull_request' || github.event.pull_request.head.repo.full_name == github.repository
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
outputs:
|
||||
staging: ${{ steps.filter.outputs.staging }}
|
||||
prod: ${{ steps.filter.outputs.prod }}
|
||||
dev: ${{ steps.filter.outputs.dev }}
|
||||
shared: ${{ steps.filter.outputs.shared }}
|
||||
# Ansible-converge gates: isolate the Ansible surfaces so a pure-Terraform
|
||||
# change to the ceph/fabric stack applies without reconverging the nodes.
|
||||
ansible_ceph: ${{ steps.filter.outputs.ansible_ceph }}
|
||||
# Mgmt Ansible-converge gate: isolate the mgmt Ansible surface so a
|
||||
# pure-Terraform change to the fabric stack applies without reconverging the
|
||||
# mgmt hosts. (The ceph converge is workflow_dispatch-only -- no push gate.)
|
||||
ansible_mgmt: ${{ steps.filter.outputs.ansible_mgmt }}
|
||||
# Fabric-apply gate: the fabric (switch) stack applies only when its own
|
||||
# surface changed — not on every prod change.
|
||||
@@ -130,10 +160,15 @@ jobs:
|
||||
filters: |
|
||||
staging:
|
||||
- 'tf/deployment/staging/**'
|
||||
- 'ansible/ceph/**'
|
||||
# Only staging's own ceph inventory activates the staging matrix.
|
||||
# A prod ceph change (or a shared ceph role/play) must NEVER open
|
||||
# the staging partition -- see DEFECT 3. The ceph Ansible converge
|
||||
# itself is workflow_dispatch-only regardless (no push converge).
|
||||
- 'ansible/ceph/inventories/staging-**'
|
||||
prod:
|
||||
- 'tf/deployment/prod/**'
|
||||
- 'ansible/mgmt/**'
|
||||
- 'ansible/ceph/inventories/prod-**'
|
||||
- 'tf/render/**'
|
||||
- 'tf/providers/**'
|
||||
- '.mise/tasks/infra/**'
|
||||
@@ -150,11 +185,11 @@ jobs:
|
||||
- '.mise/config.toml'
|
||||
- '.github/actions/netbird-connect/**'
|
||||
- '.github/workflows/infra.yml'
|
||||
# ── Ansible-converge gates (stack-step level, not partition level) ──
|
||||
# These don't open/close a partition's matrix; the apply job uses them
|
||||
# to decide whether to run the post-apply Ansible converge for its stack.
|
||||
ansible_ceph:
|
||||
- 'ansible/ceph/**'
|
||||
# -- Mgmt Ansible-converge gate (stack-step level, not partition) --
|
||||
# Does not open/close a partition's matrix; the apply job uses it to
|
||||
# decide whether to run the post-apply mgmt converge for the fabric
|
||||
# stack. (There is no ceph equivalent: the ceph converge is
|
||||
# workflow_dispatch-only, so it needs no push paths-filter.)
|
||||
ansible_mgmt:
|
||||
- 'ansible/mgmt/**'
|
||||
- 'tf/render/ansible-mgmt/**'
|
||||
@@ -184,6 +219,9 @@ jobs:
|
||||
|| needs.changes.outputs.shared == 'true'
|
||||
|| github.event_name == 'workflow_dispatch'
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
permissions:
|
||||
contents: read
|
||||
outputs:
|
||||
matrix: ${{ steps.gen.outputs.matrix }}
|
||||
has_stacks: ${{ steps.gen.outputs.has_stacks }}
|
||||
@@ -198,13 +236,39 @@ jobs:
|
||||
PROD: ${{ needs.changes.outputs.prod }}
|
||||
SHARED: ${{ needs.changes.outputs.shared }}
|
||||
DISPATCH: ${{ github.event_name == 'workflow_dispatch' }}
|
||||
# Dispatch scoping: a manual run only touches the partition(s) its
|
||||
# toggles actually act on, instead of forcing BOTH partitions active.
|
||||
DISPATCH_CEPH: ${{ inputs.run_ceph_ansible }}
|
||||
DISPATCH_MGMT: ${{ inputs.run_mgmt_ansible }}
|
||||
DISPATCH_FABRIC: ${{ inputs.run_fabric }}
|
||||
CEPH_CLUSTER: ${{ inputs.ceph_cluster }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
|
||||
# Active partitions: own filter OR a shared/manual force. dev is
|
||||
# local-only and never emitted (no remote state / no CI SA).
|
||||
# Active partitions: own filter OR a shared force. dev is local-only and
|
||||
# never emitted (no remote state / no CI SA).
|
||||
declare -A active=()
|
||||
if [ "$SHARED" = "true" ] || [ "$DISPATCH" = "true" ]; then
|
||||
if [ "$DISPATCH" = "true" ]; then
|
||||
# Manual run: scope to the partition each requested action touches.
|
||||
# - ceph converge -> the partition owning the chosen cluster
|
||||
# (spice = prod, sietch = staging)
|
||||
# - mgmt converge / fabric apply -> prod (htz-fsn1 only)
|
||||
# A BARE dispatch (no toggle set) is a plan-only preview: emit both
|
||||
# partitions so the operator sees every diff, but the apply job is
|
||||
# skipped entirely (see its `if:`), so nothing applies unreviewed.
|
||||
if [ "$DISPATCH_CEPH" = "true" ]; then
|
||||
case "$CEPH_CLUSTER" in
|
||||
spice) active[prod]=1 ;;
|
||||
sietch) active[staging]=1 ;;
|
||||
*) echo "unknown ceph_cluster: '$CEPH_CLUSTER'" >&2; exit 1 ;;
|
||||
esac
|
||||
fi
|
||||
[ "$DISPATCH_MGMT" = "true" ] && active[prod]=1
|
||||
[ "$DISPATCH_FABRIC" = "true" ] && active[prod]=1
|
||||
if [ "${#active[@]}" -eq 0 ]; then
|
||||
active[staging]=1; active[prod]=1 # bare dispatch = plan-only preview
|
||||
fi
|
||||
elif [ "$SHARED" = "true" ]; then
|
||||
active[staging]=1; active[prod]=1
|
||||
else
|
||||
[ "$STAGING" = "true" ] && active[staging]=1
|
||||
@@ -251,12 +315,59 @@ jobs:
|
||||
echo "has_stacks=$has" >> "$GITHUB_OUTPUT"
|
||||
echo "Discovered stacks:"; echo "$include" | jq -r '.[] | " [\(.order)] \(.partition)@\(.region)/\(.stack) (\(.dir))"'
|
||||
|
||||
# -- Approval-gate assertion: PARKED (disabled) ----------------------
|
||||
# This step asserted that every target Environment carries a required
|
||||
# reviewer and failed the run otherwise (fail-closed four-eyes on prod
|
||||
# applies). It is DISABLED pending a blast-radius redesign: as written it
|
||||
# blocks the FIRST push-to-main until BOTH (a) a GH_ENV_ADMIN_TOKEN secret
|
||||
# exists (the built-in GITHUB_TOKEN cannot read protection rules) AND (b) the
|
||||
# four Environments are created WITH reviewers via the parked
|
||||
# tf/deployment/_meta/github stack. Neither is bootstrapped yet, so enforcing
|
||||
# here would break every prod/staging apply on merge. While parked, an
|
||||
# unprovisioned Environment auto-creates UNPROTECTED (fail OPEN) -- the
|
||||
# pre-hardening status quo, tracked for follow-up. The full logic is kept
|
||||
# below for revival; re-enable once the gate can be provisioned without a
|
||||
# hard cutover. See tf/deployment/_meta/github/terragrunt.hcl.disabled.
|
||||
- name: Fail closed unless every target Environment has required reviewers
|
||||
# PARKED: re-enable by restoring this guard (blast-radius redesign pending):
|
||||
# steps.gen.outputs.has_stacks == 'true'
|
||||
# && (github.event_name == 'push' || github.event_name == 'workflow_dispatch')
|
||||
if: false
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GH_ENV_ADMIN_TOKEN || secrets.GITHUB_TOKEN }}
|
||||
REPO: ${{ github.repository }}
|
||||
MATRIX: ${{ steps.gen.outputs.matrix }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
envs=$(jq -r '.include[] | "\(.partition)-\(.region)"' <<<"$MATRIX" | sort -u)
|
||||
[ -n "$envs" ] || { echo "No environments to verify."; exit 0; }
|
||||
fail=0
|
||||
while IFS= read -r envname; do
|
||||
[ -n "$envname" ] || continue
|
||||
if ! resp=$(gh api "repos/$REPO/environments/$envname" 2>/dev/null); then
|
||||
echo "FAIL-CLOSED: environment '$envname' is missing or unreadable -> treated as unprotected." >&2
|
||||
fail=1; continue
|
||||
fi
|
||||
n=$(jq '[.protection_rules[]? | select(.type=="required_reviewers") | .reviewers[]?] | length' <<<"$resp")
|
||||
if [ "${n:-0}" -lt 1 ]; then
|
||||
echo "FAIL-CLOSED: environment '$envname' has no required reviewers (approval gate would fail open)." >&2
|
||||
fail=1
|
||||
else
|
||||
echo "OK: environment '$envname' has ${n} required reviewer(s)."
|
||||
fi
|
||||
done <<<"$envs"
|
||||
[ "$fail" -eq 0 ] || {
|
||||
echo "One or more target Environments lack a required reviewer; refusing to plan/apply (fail closed)." >&2
|
||||
exit 1
|
||||
}
|
||||
|
||||
# ── Plan every changed stack (parallel; read-only) ───────────────────────────
|
||||
plan:
|
||||
name: Plan ${{ matrix.partition }}@${{ matrix.region }}/${{ matrix.stack }}
|
||||
needs: [changes, discover]
|
||||
if: needs.discover.outputs.has_stacks == 'true'
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 20
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix: ${{ fromJSON(needs.discover.outputs.matrix) }}
|
||||
@@ -272,10 +383,22 @@ jobs:
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
||||
with:
|
||||
persist-credentials: false
|
||||
# Filesystem-safe artifact id for this stack (stack may nest '/').
|
||||
- id: planid
|
||||
name: Compute plan artifact id
|
||||
env:
|
||||
DIR: ${{ matrix.dir }}
|
||||
run: echo "id=$(printf '%s' "$DIR" | tr '/' '-')" >> "$GITHUB_OUTPUT"
|
||||
- name: Set up mise (go + opentofu + terragrunt)
|
||||
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
|
||||
- name: Install 1Password CLI
|
||||
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
|
||||
with:
|
||||
# Pin, don't resolve "latest": the version-check endpoint 1P hits to
|
||||
# resolve "latest" intermittently returns empty JSON (all parallel
|
||||
# plan/apply jobs fail at "Getting latest version number"). A pinned
|
||||
# version installs directly, skipping that call.
|
||||
version: '2.31.1'
|
||||
|
||||
# Guardrail: prod Flux must sync `main` once we're ON main. flux_git_ref
|
||||
# defaults to a feature branch during bring-up (deliberate, documented
|
||||
@@ -319,29 +442,63 @@ jobs:
|
||||
# Prod fabric goes through the mise task (builds the hetzner provider, renders
|
||||
# the NETCONF key; junos comes from the registry). Every other stack is a
|
||||
# plain registry-provider terragrunt plan.
|
||||
# Persist the diff to an ABSOLUTE path (terragrunt runs tofu inside a
|
||||
# .terragrunt-cache dir, so a relative -out would be unreachable). The gated
|
||||
# apply downloads this exact file and applies it -- the human approves the
|
||||
# DIFF, not a plan-job log. TG_PLAN threads the path into the mise task too.
|
||||
- name: Terragrunt plan (fabric)
|
||||
if: matrix.stack == 'fabric'
|
||||
env:
|
||||
TG_PLAN: ${{ github.workspace }}/tfplan.bin
|
||||
run: mise run infra:plan -- --non-interactive
|
||||
- name: Terragrunt plan
|
||||
if: matrix.stack != 'fabric'
|
||||
env:
|
||||
STACK_DIR: ${{ matrix.dir }}
|
||||
TG_PLAN: ${{ github.workspace }}/tfplan.bin
|
||||
run: >-
|
||||
tf/op-run.sh terragrunt
|
||||
--working-dir "tf/deployment/$STACK_DIR"
|
||||
--non-interactive plan
|
||||
--non-interactive plan -out="$TG_PLAN"
|
||||
|
||||
# Hand the saved plan to the apply job. NOTE: a tofu plan file can embed
|
||||
# sensitive values in the clear, so keep retention short and rely on the
|
||||
# repo's private-artifact scoping. if-no-files-found:error makes a missing
|
||||
# plan fail loudly rather than silently fall back to a fresh apply.
|
||||
- name: Upload reviewed plan
|
||||
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
||||
with:
|
||||
name: tfplan-${{ steps.planid.outputs.id }}
|
||||
path: ${{ github.workspace }}/tfplan.bin
|
||||
if-no-files-found: error
|
||||
retention-days: 1
|
||||
|
||||
# ── Apply (gated, ordered) ───────────────────────────────────────────────────
|
||||
apply:
|
||||
name: Apply ${{ matrix.partition }}@${{ matrix.region }}/${{ matrix.stack }} (gated)
|
||||
needs: [changes, discover, plan]
|
||||
# Only on merge to main (or manual dispatch) — never on PRs.
|
||||
if: (github.event_name == 'push' || github.event_name == 'workflow_dispatch') && needs.discover.outputs.has_stacks == 'true'
|
||||
# Only on merge to main (or a SCOPED manual dispatch) -- never on PRs. A bare
|
||||
# workflow_dispatch (no run_ceph_ansible / run_mgmt_ansible / run_fabric) is
|
||||
# plan-only: the apply job is skipped so a manual "just look" run can never
|
||||
# apply every emitted stack unreviewed.
|
||||
if: >-
|
||||
needs.discover.outputs.has_stacks == 'true'
|
||||
&& (
|
||||
github.event_name == 'push'
|
||||
|| (github.event_name == 'workflow_dispatch'
|
||||
&& (inputs.run_ceph_ansible || inputs.run_mgmt_ansible || inputs.run_fabric))
|
||||
)
|
||||
runs-on: ubuntu-latest
|
||||
# Ceiling for a worst-case entry (TF apply + up to 90-min ceph converge).
|
||||
# Finer per-phase bounds are enforced at step level (apply 30 / converge 90).
|
||||
timeout-minutes: 120
|
||||
strategy:
|
||||
fail-fast: false
|
||||
# Serialize so the sorted `order` is honored: account netbird → site netbird
|
||||
# → node-touching stacks, and prod/global before prod/htz-fsn1/netbird.
|
||||
# fail-fast so a failed lower-order apply (e.g. netbird) cancels the
|
||||
# remaining ordered node-touching applies instead of running against a
|
||||
# half-provisioned overlay.
|
||||
fail-fast: true
|
||||
# Serialize so the sorted `order` is honored: account netbird -> site netbird
|
||||
# -> node-touching stacks, and prod/global before prod/htz-fsn1/netbird.
|
||||
max-parallel: 1
|
||||
matrix: ${{ fromJSON(needs.discover.outputs.matrix) }}
|
||||
# Per-region approval gate.
|
||||
@@ -360,32 +517,67 @@ jobs:
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
||||
with:
|
||||
persist-credentials: false
|
||||
# Same filesystem-safe id the plan job used to name the artifact.
|
||||
- id: planid
|
||||
name: Compute plan artifact id
|
||||
env:
|
||||
DIR: ${{ matrix.dir }}
|
||||
run: echo "id=$(printf '%s' "$DIR" | tr '/' '-')" >> "$GITHUB_OUTPUT"
|
||||
- name: Set up mise (go + opentofu + terragrunt + ansible)
|
||||
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
|
||||
- name: Install 1Password CLI
|
||||
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
|
||||
with:
|
||||
# See the plan job's note: pin to skip the flaky "latest" resolution.
|
||||
version: '2.31.1'
|
||||
|
||||
# Decide whether each post-apply Ansible converge runs. On push, gate on the
|
||||
# paths-filter (did the Ansible surface change?); on manual dispatch, gate on
|
||||
# the run_*_ansible toggles (default off) since there's no diff to detect.
|
||||
# Pull the EXACT plan the reviewer approved, downloaded to the same absolute
|
||||
# path the plan job wrote. `terragrunt apply <planfile>` fails closed if the
|
||||
# saved plan is stale relative to current state -- so a drifted backend can
|
||||
# never silently apply a different diff than the one that was reviewed.
|
||||
- name: Download reviewed plan
|
||||
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
|
||||
with:
|
||||
name: tfplan-${{ steps.planid.outputs.id }}
|
||||
path: ${{ github.workspace }}
|
||||
|
||||
# Decide whether each post-apply Ansible converge runs.
|
||||
# - ceph: DELIBERATE and blast-radius-guarded. Runs ONLY on
|
||||
# workflow_dispatch with run_ceph_ansible=true, and ONLY for the ceph
|
||||
# matrix entry whose partition/region owns the selected ceph_cluster
|
||||
# (spice -> prod/htz-fsn1, sietch -> staging/austin). A push never sets
|
||||
# this gate, so merging a ceph change cannot converge a live cluster.
|
||||
# - mgmt: on push, gate on the paths-filter (did the mgmt surface change?);
|
||||
# on manual dispatch, gate on the run_mgmt_ansible toggle (default off).
|
||||
- name: Resolve Ansible-converge gates
|
||||
id: ansible_gate
|
||||
env:
|
||||
DISPATCH: ${{ github.event_name == 'workflow_dispatch' }}
|
||||
DISPATCH_CEPH: ${{ inputs.run_ceph_ansible }}
|
||||
DISPATCH_MGMT: ${{ inputs.run_mgmt_ansible }}
|
||||
CHANGED_CEPH: ${{ needs.changes.outputs.ansible_ceph }}
|
||||
CEPH_CLUSTER: ${{ inputs.ceph_cluster }}
|
||||
PARTITION: ${{ matrix.partition }}
|
||||
REGION: ${{ matrix.region }}
|
||||
CHANGED_MGMT: ${{ needs.changes.outputs.ansible_mgmt }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
if [ "$DISPATCH" = "true" ]; then
|
||||
ceph=$DISPATCH_CEPH; mgmt=$DISPATCH_MGMT
|
||||
else
|
||||
ceph=$CHANGED_CEPH; mgmt=$CHANGED_MGMT
|
||||
# ceph converge: dispatch-only, and pinned to the one cluster/entry.
|
||||
ceph=false
|
||||
if [ "$DISPATCH" = "true" ] && [ "$DISPATCH_CEPH" = "true" ]; then
|
||||
case "$CEPH_CLUSTER" in
|
||||
spice) want_partition=prod; want_region=htz-fsn1 ;;
|
||||
sietch) want_partition=staging; want_region=austin ;;
|
||||
*) echo "unknown ceph_cluster: '$CEPH_CLUSTER'" >&2; exit 1 ;;
|
||||
esac
|
||||
if [ "$PARTITION" = "$want_partition" ] && [ "$REGION" = "$want_region" ]; then
|
||||
ceph=true
|
||||
fi
|
||||
fi
|
||||
# mgmt converge: push -> paths-filter; dispatch -> toggle.
|
||||
if [ "$DISPATCH" = "true" ]; then mgmt=$DISPATCH_MGMT; else mgmt=$CHANGED_MGMT; fi
|
||||
echo "ceph=${ceph:-false}" >> "$GITHUB_OUTPUT"
|
||||
echo "mgmt=${mgmt:-false}" >> "$GITHUB_OUTPUT"
|
||||
echo "Ansible converge gates → ceph=${ceph:-false} mgmt=${mgmt:-false}"
|
||||
echo "Ansible converge gates -> ceph=${ceph:-false} (cluster=${CEPH_CLUSTER:-none}) mgmt=${mgmt:-false}"
|
||||
|
||||
# Decide whether the fabric (switch) apply runs. Same model as the Ansible
|
||||
# gates: on push, gate on the paths-filter (did the fabric surface change?);
|
||||
@@ -428,37 +620,50 @@ jobs:
|
||||
# Gated: only when the fabric surface changed (push) or run_fabric (dispatch).
|
||||
- name: Terragrunt apply (fabric)
|
||||
if: matrix.stack == 'fabric' && steps.fabric_gate.outputs.run == 'true'
|
||||
run: mise run infra:apply -- --non-interactive -auto-approve
|
||||
# Everything else: direct terragrunt apply with the partition's write SA.
|
||||
timeout-minutes: 30
|
||||
# TG_PLAN routes the reviewed plan into the mise task, which applies that
|
||||
# file (no fresh re-plan, no -auto-approve) and fails on a stale plan.
|
||||
env:
|
||||
TG_PLAN: ${{ github.workspace }}/tfplan.bin
|
||||
run: mise run infra:apply -- --non-interactive
|
||||
# Everything else: direct terragrunt apply of the reviewed plan file with
|
||||
# the partition's write SA (no -auto-approve -- the plan is pre-approved).
|
||||
- name: Terragrunt apply
|
||||
if: matrix.stack != 'fabric'
|
||||
timeout-minutes: 30
|
||||
env:
|
||||
STACK_DIR: ${{ matrix.dir }}
|
||||
TG_PLAN: ${{ github.workspace }}/tfplan.bin
|
||||
run: >-
|
||||
tf/op-run.sh terragrunt
|
||||
--working-dir "tf/deployment/$STACK_DIR"
|
||||
--non-interactive apply -auto-approve
|
||||
--non-interactive apply "$TG_PLAN"
|
||||
|
||||
# ── Ceph convergence (Ansible) — ceph stack, gated on the converge gate ──
|
||||
# -- Ceph convergence (Ansible) - ceph stack, workflow_dispatch-only --
|
||||
# The TF apply above only minted the RGW keys into 1P + the cluster Secret;
|
||||
# this creates the matching RGW users on the bare-metal cluster. Reuses the
|
||||
# NetBird overlay + 1Password session already established in this entry.
|
||||
# Skipped (along with the overlay join above) when ansible/ceph/** is
|
||||
# unchanged on push, or when run_ceph_ansible is off on manual dispatch.
|
||||
# this converges the selected bare-metal cluster. Reuses the NetBird overlay
|
||||
# + 1Password session already established in this entry. The gate fires ONLY
|
||||
# on workflow_dispatch with run_ceph_ansible=true, and only for the entry
|
||||
# that owns the chosen ceph_cluster (spice=prod/htz-fsn1, sietch=staging/
|
||||
# austin) -- a push never converges here. Cluster + region are threaded from
|
||||
# the dispatch inputs / matrix so this is not hardcoded to sietch/austin.
|
||||
- name: Render Ansible inventory from the ceph TF state
|
||||
if: matrix.stack == 'ceph' && steps.ansible_gate.outputs.ceph == 'true'
|
||||
env:
|
||||
PARTITION: ${{ matrix.partition }}
|
||||
run: ansible/ceph/scripts/render-inventories.sh "$PARTITION"
|
||||
REGION: ${{ matrix.region }}
|
||||
run: ansible/ceph/scripts/render-inventories.sh "$PARTITION" "$REGION"
|
||||
- name: Install the ansible-iac SSH key from 1Password
|
||||
if: matrix.stack == 'ceph' && steps.ansible_gate.outputs.ceph == 'true'
|
||||
# Pulls op://yucca_tf_<partition>/SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY to
|
||||
# ~/.ssh/id_ed25519_sietch — the path the rendered inventory references.
|
||||
# Pulls op://yucca_tf_<partition>/<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY to
|
||||
# ~/.ssh/id_ed25519_<cluster> -- the path the rendered inventory references
|
||||
# (spice -> id_ed25519_spice, sietch -> id_ed25519_sietch).
|
||||
env:
|
||||
PARTITION: ${{ matrix.partition }}
|
||||
CEPH_CLUSTER: ${{ inputs.ceph_cluster }}
|
||||
run: |
|
||||
mkdir -p ~/.ssh && chmod 700 ~/.ssh
|
||||
OP_VAULT="yucca_tf_$PARTITION" ansible/ceph/scripts/install-ssh-keys.sh sietch
|
||||
OP_VAULT="yucca_tf_$PARTITION" ansible/ceph/scripts/install-ssh-keys.sh "$CEPH_CLUSTER"
|
||||
- name: Provision the ceph Ansible toolchain (venv + collections)
|
||||
if: matrix.stack == 'ceph' && steps.ansible_gate.outputs.ceph == 'true'
|
||||
working-directory: ansible/ceph
|
||||
@@ -466,8 +671,9 @@ jobs:
|
||||
mise trust
|
||||
mise install
|
||||
mise run setup
|
||||
- name: Deploy Ceph (full pipeline — baseline → tune → deploy → harden)
|
||||
- name: Deploy Ceph (full pipeline -- baseline -> tune -> deploy -> harden)
|
||||
if: matrix.stack == 'ceph' && steps.ansible_gate.outputs.ceph == 'true'
|
||||
timeout-minutes: 90
|
||||
# Run from ansible/ceph so mise loads ansible/ceph/.mise.toml (where the
|
||||
# `deploy` task + CEPH_ENV-relative inventory paths live); the root config
|
||||
# has no `deploy` task.
|
||||
@@ -476,7 +682,7 @@ jobs:
|
||||
# over the NetBird overlay, so disable strict host-key checking for this run.
|
||||
env:
|
||||
ANSIBLE_HOST_KEY_CHECKING: 'false'
|
||||
CEPH_ENV: inventories/${{ matrix.partition }}-${{ matrix.region }}/sietch/inventory.ini
|
||||
CEPH_ENV: inventories/${{ matrix.partition }}-${{ matrix.region }}/${{ inputs.ceph_cluster }}/inventory.ini
|
||||
run: mise run deploy
|
||||
|
||||
# ── Mgmt convergence (Ansible) — fabric stack, gated on the converge gate ─
|
||||
@@ -488,4 +694,5 @@ jobs:
|
||||
# overlay join regardless (it touches the switch vme directly).
|
||||
- name: Ansible converge (mgmt hosts)
|
||||
if: matrix.stack == 'fabric' && steps.ansible_gate.outputs.mgmt == 'true'
|
||||
timeout-minutes: 30
|
||||
run: mise run mgmt:ansible
|
||||
|
||||
@@ -0,0 +1,54 @@
|
||||
name: s3-dns-roster-validate
|
||||
|
||||
# Static gate: the prod S3 RGW DNS roster (the A-records for
|
||||
# s3.prod.fsn1.htz.futo.cloud + the wildcard, in tf/deployment/prod/global/dns)
|
||||
# is a hand-maintained round-robin over every in-service spice ceph node's fabric
|
||||
# IP. Nothing links it to the ceph stack, so adding/replacing/removing a node can
|
||||
# silently blackhole the endpoint (a stale IP that routes nowhere) or drop capacity
|
||||
# (a live node left out of rotation). Without this gate the mismatch first surfaces
|
||||
# in production, after the DNS apply.
|
||||
#
|
||||
# tf/scripts/check-s3-dns-roster.py reconstructs the expected roster from the ceph
|
||||
# source of truth (clusters.auto.tfvars in-service set + spice-hosts.yaml
|
||||
# host_index -> fabric IP) and fails on any divergence. Credential-free and
|
||||
# stdlib-only: no remote state, no provider, no 1Password, no host connection.
|
||||
#
|
||||
# Its own workflow (not a job in infra.yml or ci.yml) so the path filter below is
|
||||
# workflow-scoped: it fires only when the DNS roster or a ceph roster input
|
||||
# changes, mirroring ansible-ceph-validate.yml.
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
paths: &paths
|
||||
- 'tf/deployment/prod/global/dns/records.auto.tfvars'
|
||||
- 'tf/deployment/prod/htz-fsn1/ceph/clusters.auto.tfvars'
|
||||
- 'tf/deployment/prod/htz-fsn1/ceph/spice-hosts.yaml'
|
||||
- 'tf/scripts/check-s3-dns-roster.py'
|
||||
- '.github/workflows/s3-dns-roster-validate.yml'
|
||||
push:
|
||||
branches: [main]
|
||||
paths: *paths
|
||||
|
||||
# Read-only: checkout only. No writes through the token, no deploy credentials,
|
||||
# no 1Password.
|
||||
permissions:
|
||||
contents: read
|
||||
|
||||
concurrency:
|
||||
group: s3-dns-roster-validate-${{ github.ref }}
|
||||
cancel-in-progress: true
|
||||
|
||||
jobs:
|
||||
validate:
|
||||
name: Check S3 DNS roster vs ceph node roster
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 5
|
||||
steps:
|
||||
- uses: actions/checkout@1af3b93b6815bc44a9784bd300feb67ff0d1eeb3 # v6.0.0
|
||||
with:
|
||||
persist-credentials: false
|
||||
|
||||
# Stdlib-only Python; ubuntu-latest ships python3. No mise/tool install
|
||||
# needed, so this gate stays fast and dependency-free.
|
||||
- name: Roster drift check
|
||||
run: tf/scripts/check-s3-dns-roster.py
|
||||
@@ -97,6 +97,10 @@ run = "tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/stagi
|
||||
description = "Format terraform + terragrunt files recursively"
|
||||
run = "tofu fmt -recursive tf/ && terragrunt hcl format --working-dir tf/"
|
||||
|
||||
[tasks."tf:check-dns-roster"]
|
||||
description = "Drift-check the prod S3 RGW DNS roster against the ceph node roster (credential-free, no state)"
|
||||
run = "tf/scripts/check-s3-dns-roster.py"
|
||||
|
||||
[env]
|
||||
NODE_ENV = "development"
|
||||
LOG_LEVEL = "debug"
|
||||
|
||||
+14
-1
@@ -34,10 +34,23 @@ SITE="${SITE:-htz-fsn1}"
|
||||
# provider rejects having both set ("service_account_token and account are set").
|
||||
unset OP_ACCOUNT
|
||||
STACK_DIR="tf/deployment/prod/$SITE/fabric"
|
||||
|
||||
# Bind the apply to the reviewed plan. When CI passes TG_PLAN pointing at the
|
||||
# saved plan artifact, apply THAT file (no fresh re-plan, no -auto-approve): tofu
|
||||
# fails closed if the saved plan is stale relative to current state. Strip any
|
||||
# -auto-approve the caller passed (meaningless with a plan file). Local/manual
|
||||
# runs leave TG_PLAN unset and keep the interactive/`-auto-approve` behaviour.
|
||||
APPLY_ARGS=("$@")
|
||||
if [ -n "${TG_PLAN:-}" ]; then
|
||||
[ -f "$TG_PLAN" ] || { echo "infra:apply: TG_PLAN set but plan file missing: $TG_PLAN" >&2; exit 1; }
|
||||
filtered=(); for a in "${APPLY_ARGS[@]}"; do [ "$a" = "-auto-approve" ] && continue; filtered+=("$a"); done
|
||||
APPLY_ARGS=("${filtered[@]}" "$TG_PLAN")
|
||||
fi
|
||||
|
||||
# jeremmfr/junos is concurrency-safe (per-resource CRUD); reads/refresh run in
|
||||
# parallel and commits serialize on the per-device config lock (handled by the
|
||||
# provider). The device NETCONF connection-limit is raised to 250.
|
||||
OP_ENV_FILE=tf/.env.prod "$ROOT/tf/op-run.sh" terragrunt --working-dir "$STACK_DIR" apply -parallelism=4 "$@"
|
||||
OP_ENV_FILE=tf/.env.prod "$ROOT/tf/op-run.sh" terragrunt --working-dir "$STACK_DIR" apply -parallelism=4 "${APPLY_ARGS[@]}"
|
||||
|
||||
# Confirm the dangling `commit confirmed` jeremmfr leaves on the last commit per
|
||||
# device: its per-resource "confirm" is `commit check`, which does NOT cancel the
|
||||
|
||||
@@ -33,5 +33,13 @@ SITE="${SITE:-htz-fsn1}"
|
||||
# With a service-account token in the env, drop OP_ACCOUNT — the onepassword
|
||||
# provider rejects having both set ("service_account_token and account are set").
|
||||
unset OP_ACCOUNT
|
||||
|
||||
# When CI sets TG_PLAN, persist the diff to that (absolute) path so the gated
|
||||
# apply can consume the EXACT reviewed plan (terragrunt runs tofu inside a
|
||||
# .terragrunt-cache dir, so a relative -out would be unreachable -- pass an
|
||||
# absolute path). Local `tf:plan` leaves TG_PLAN unset -> a plain read-only plan.
|
||||
PLAN_OUT=()
|
||||
[ -n "${TG_PLAN:-}" ] && PLAN_OUT=(-out="$TG_PLAN")
|
||||
|
||||
# jeremmfr/junos is concurrency-safe; the device NETCONF connection-limit is 250.
|
||||
OP_ENV_FILE=tf/.env.prod "$ROOT/tf/op-run.sh" terragrunt --working-dir "tf/deployment/prod/$SITE/fabric" plan -parallelism=4 "$@"
|
||||
OP_ENV_FILE=tf/.env.prod "$ROOT/tf/op-run.sh" terragrunt --working-dir "tf/deployment/prod/$SITE/fabric" plan -parallelism=4 "${PLAN_OUT[@]}" "$@"
|
||||
|
||||
@@ -155,6 +155,10 @@ run = "scripts/ansible-play.sh backup-config.yml"
|
||||
description = "Snapshot bootstrap secrets (RGW TLS, admin keyring) to 1P (DR belt)"
|
||||
run = "scripts/ansible-play.sh post-deploy-capture.yml"
|
||||
|
||||
[tasks."backup-timer"]
|
||||
description = "Install + enable the scheduled cluster-state backup timer on the bootstrap node"
|
||||
run = "scripts/ansible-play.sh backup-ceph.yml"
|
||||
|
||||
[tasks."rotate-certs"]
|
||||
description = "Rotate RGW TLS cert (regenerate self-signed, restart RGW daemons)"
|
||||
run = "scripts/ansible-play.sh rotate-certs.yml"
|
||||
@@ -170,3 +174,14 @@ run = "scripts/ansible-play.sh hardware-inventory.yml"
|
||||
[tasks."migrate-networkd"]
|
||||
description = "One-shot ifupdown→networkd migration (rolling, serial=1, noout-gated)"
|
||||
run = "scripts/ansible-play.sh migrate-networkd.yml"
|
||||
|
||||
[tasks.reprovision]
|
||||
description = "Hetzner installimage reprovision (rescue -> reset -> installimage -> verify). DESTRUCTIVE - canary first."
|
||||
run = """
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
# HETZNER_ROBOT_* come from tf/.env.prod via op run; the same op session also
|
||||
# satisfies ansible-play.sh's inner op inject of secrets.yml.tpl. Set CEPH_ENV first.
|
||||
OP_ACCOUNT=team-futo op run --env-file=../../tf/.env.prod -- \\
|
||||
scripts/ansible-play.sh reprovision.yml "$@"
|
||||
"""
|
||||
|
||||
@@ -326,6 +326,7 @@ scripts/ansible-play.sh deploy-ceph.yml --tags bootstrap
|
||||
| `drift` | `mise run drift` | Detect configuration drift (via ansible-play.sh) |
|
||||
| `deploy` | `mise run deploy` | Full deploy pipeline (via ansible-play.sh) |
|
||||
| `backup` | `mise run backup` | Export cluster config for DR (via ansible-play.sh) |
|
||||
| `backup-timer` | `mise run backup-timer` | Install the scheduled on-node cluster-state backup timer (bootstrap node) |
|
||||
| `capture` | `mise run capture` | Snapshot RGW TLS + admin keyring to 1P for disaster recovery |
|
||||
| `bench` | `mise run bench` | S3 benchmark (RGW round-trip) |
|
||||
| `bench-rados` | `mise run bench-rados` | RADOS bench (raw cluster I/O) |
|
||||
|
||||
@@ -0,0 +1,67 @@
|
||||
---
|
||||
# Manual, rescue-first "add a node" path: install + chroot + verify on a node the
|
||||
# operator has ALREADY put into the Hetzner rescue system by hand. This is the
|
||||
# default way to (re)image a spice node -- the Robot API (register keys / arm
|
||||
# rescue / hw reset) is deliberately OUT of the loop (reprovision_use_robot=false).
|
||||
#
|
||||
# WHY manual: the Robot rescue/reset step is the unreliable part (nodes that arm
|
||||
# rescue but never boot it) AND the highest blast radius (an errant /reset on a
|
||||
# LIVE spice node once the cluster is in production). Taking it out means nothing
|
||||
# Ansible does here can flip a running node into rescue: installimage exists only
|
||||
# in the rescue ramdisk, so wait_rescue hard-refuses any node a human did not
|
||||
# deliberately rescue. The install/chroot/verify automation -- the reliable part --
|
||||
# still runs untouched.
|
||||
#
|
||||
# For the rare remote "nuke a dead node" case, use reprovision.yml (Robot path).
|
||||
#
|
||||
# OPERATOR PRECONDITION (do this in the Hetzner Robot portal, per node):
|
||||
# 1. Enable the Linux rescue system, authorizing the spice-ansible-iac key
|
||||
# (and operator break-glass keys). 2. Trigger a hardware reset. 3. Wait until
|
||||
# the node is reachable in rescue over its WAN IP. THEN run this play.
|
||||
# Run it promptly after rescue boots (wait_rescue rejects a stale rescue whose
|
||||
# uptime exceeds reprovision_rescue_max_uptime; raise it with -e if needed).
|
||||
#
|
||||
# No secrets and no Robot creds are needed -- run with plain ansible-playbook:
|
||||
# Canary (one node, full wipe + verify):
|
||||
# ansible-playbook -i inventories/prod-htz-fsn1/spice/inventory.ini add-node.yml \
|
||||
# -e confirm_wipe=true --limit spice-ceph-philip
|
||||
# Fan-out (only nodes you have already rescued; low serial, canary-gated):
|
||||
# ansible-playbook -i inventories/prod-htz-fsn1/spice/inventory.ini add-node.yml \
|
||||
# -e confirm_wipe=true -e allow_fanout=true -e reprovision_serial=2 \
|
||||
# --limit 'rescue:!spice-ceph-philip'
|
||||
#
|
||||
# ALWAYS DESTRUCTIVE to the NVMe OS disks (installimage). Requires confirm_wipe=true.
|
||||
|
||||
- name: Add node (manual rescue -> installimage -> installed OS)
|
||||
hosts: ceph_nodes
|
||||
gather_facts: false
|
||||
become: false
|
||||
serial: "{{ reprovision_serial | default(1) }}"
|
||||
max_fail_percentage: 0 # any node fails -> halt the batch (no further wipes)
|
||||
vars:
|
||||
reprovision_use_robot: false # no Robot API: rescue is armed by hand, not here
|
||||
# With use_robot=false, wait_rescue skips the this-boot uptime guard entirely
|
||||
# (it only makes sense when Ansible triggered the reset). A node may sit in
|
||||
# rescue for weeks; the installimage-ramdisk check is the real in-rescue proof.
|
||||
pre_tasks:
|
||||
# run_once + through the role so defaults load and --limit can't skip it. With
|
||||
# use_robot=false this asserts keys + the fan-out gate only (no Robot calls).
|
||||
- name: Preflight (once) # noqa: run-once[task]
|
||||
ansible.builtin.include_role:
|
||||
name: reprovision_hetzner
|
||||
tasks_from: preflight
|
||||
run_once: true
|
||||
roles:
|
||||
- role: reprovision_hetzner
|
||||
|
||||
- name: Verify added nodes after reboot
|
||||
hosts: ceph_nodes
|
||||
gather_facts: false
|
||||
become: false
|
||||
vars:
|
||||
reprovision_use_robot: false
|
||||
tasks:
|
||||
- name: Verify installed OS marker
|
||||
ansible.builtin.include_role:
|
||||
name: reprovision_hetzner
|
||||
tasks_from: verify
|
||||
@@ -0,0 +1,29 @@
|
||||
---
|
||||
# backup-ceph.yml - install (and enable) SCHEDULED Ceph cluster-state backups.
|
||||
#
|
||||
# The deep review flagged that config/topology backups were manual, unscheduled,
|
||||
# and controller-local only. This play installs an on-node capture script plus a
|
||||
# systemd timer on the bootstrap node so recoverable cluster state (fsid, config
|
||||
# dump, monmap/osdmap/crushmap, osd tree, orch ls/host ls, RGW realm/zonegroup/
|
||||
# zone) is snapshot to a timestamped tarball daily under a root-only local dir,
|
||||
# pruned by retention, and ready to ship offsite once an operator wires a target.
|
||||
#
|
||||
# Deliberate/manual: run it to install the timer; the timer then runs unattended.
|
||||
# It complements (does not replace):
|
||||
# - mise run capture (post-deploy-capture.yml) -- secrets -> 1Password
|
||||
# - mise run backup (backup-config.yml) -- one-shot controller-local export
|
||||
#
|
||||
# Usage:
|
||||
# scripts/ansible-play.sh backup-ceph.yml
|
||||
# mise run backup-timer
|
||||
#
|
||||
# Opt a cluster out with `ceph_backup_enabled: false` in its group_vars.
|
||||
# Set an offsite target with `ceph_backup_offsite_dest` (empty by default).
|
||||
|
||||
- name: Install scheduled Ceph cluster-state backups
|
||||
hosts: ceph_bootstrap
|
||||
become: true
|
||||
gather_facts: false
|
||||
|
||||
roles:
|
||||
- ceph_backup
|
||||
@@ -16,6 +16,20 @@
|
||||
hosts: ceph_nodes
|
||||
become: true
|
||||
gather_facts: false
|
||||
# Blast-radius control: converge in small batches (serial) and halt the whole
|
||||
# roll if any node in a batch fails (max_fail_percentage:0) or the batch
|
||||
# degrades the cluster (the ceph-health gate). deploy-ceph stays big-bang (its
|
||||
# bootstrap needs all nodes at once); this governs the day-2 converge plays.
|
||||
serial: "{{ ceph_converge_serial | default('10%') }}"
|
||||
max_fail_percentage: 0
|
||||
|
||||
pre_tasks:
|
||||
- name: Ceph-health gate (pre-batch)
|
||||
ansible.builtin.import_tasks: tasks/health_gate.yml
|
||||
|
||||
roles:
|
||||
- baseline
|
||||
|
||||
post_tasks:
|
||||
- name: Ceph-health gate (post-batch)
|
||||
ansible.builtin.import_tasks: tasks/health_gate.yml
|
||||
|
||||
@@ -18,5 +18,46 @@
|
||||
become: true
|
||||
gather_facts: true
|
||||
|
||||
pre_tasks:
|
||||
# Fail fast if the inventory was not rendered from TF. The ceph_deploy role
|
||||
# gates every phase on the ceph_bootstrap/ceph_mon/ceph_join groups and indexes
|
||||
# groups['ceph_bootstrap'][0]; against a hand-written stopgap inventory that
|
||||
# lacks them the role would error mid-run or, worse, fall through to placing a
|
||||
# MON on every host. render-inventories.sh emits these groups from TF state
|
||||
# (bootstrap = exactly one host, mon = an odd quorum including the bootstrap
|
||||
# node) - refuse to deploy against anything else.
|
||||
- name: Assert the TF-rendered placement groups are present and sane # noqa: run-once[task]
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- (groups['ceph_bootstrap'] | default([])) | length == 1
|
||||
- (groups['ceph_mon'] | default([])) | length >= 1
|
||||
- (groups['ceph_mon'] | default([]) | length) is odd
|
||||
- ((groups['ceph_bootstrap'] | default([])) | intersect(groups['ceph_mon'] | default([])) | length) == 1
|
||||
fail_msg: >-
|
||||
Inventory is missing or has bad TF-rendered placement groups
|
||||
(need ceph_bootstrap = exactly 1 host that is also in ceph_mon, and
|
||||
ceph_mon = a non-empty ODD quorum). Render it first:
|
||||
scripts/render-inventories.sh <partition> <region>. Refusing to deploy.
|
||||
success_msg: >-
|
||||
Placement OK: bootstrap={{ groups['ceph_bootstrap'] | default([]) | length }},
|
||||
mon={{ groups['ceph_mon'] | default([]) | length }},
|
||||
join={{ groups['ceph_join'] | default([]) | length }}.
|
||||
run_once: true
|
||||
|
||||
# Per host: the address the mon/OSD/RGW bind (ceph_service_ip) must be live on
|
||||
# a local interface before bootstrap. For spice this is the fabric IP on
|
||||
# bond0.120; a node whose VLAN sub-interface did not come back after a reboot
|
||||
# would otherwise fail deep inside cephadm bootstrap/join. gather_facts is on,
|
||||
# so ansible_all_ipv4_addresses is populated.
|
||||
- name: Assert the Ceph service IP is live on this node (fabric/bond up)
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- ceph_service_ip in ansible_all_ipv4_addresses
|
||||
fail_msg: >-
|
||||
ceph_service_ip {{ ceph_service_ip }} is not on any local interface on
|
||||
{{ inventory_hostname }} - the Ceph bind network (spice: bond0.120) is
|
||||
down on this node. Bring the fabric up before deploying it.
|
||||
success_msg: "Ceph service IP {{ ceph_service_ip }} is live on {{ inventory_hostname }}."
|
||||
|
||||
roles:
|
||||
- ceph_deploy
|
||||
|
||||
@@ -62,6 +62,30 @@ but a full-node loss degrades a large share of PGs. For host-level
|
||||
failure domain you need **11+ nodes minimum**. Production (Yucca) will
|
||||
want this; dev can tolerate the weaker guarantee.
|
||||
|
||||
## Durability and DR ceiling
|
||||
|
||||
The spice (production, htz-fsn1) backup-of-record stores objects in an EC 16+4
|
||||
data pool with `failure_domain=host`. min_size is pinned to k+1 = 17 by the RGW
|
||||
role. What that pool tolerates:
|
||||
|
||||
- **Reads:** survive up to m = 4 simultaneous host (or OSD) losses. With 4
|
||||
chunks gone the remaining 16 = k chunks still reconstruct every object.
|
||||
- **Writes:** survive up to m - 1 = 3 simultaneous host losses. At the k+1
|
||||
floor a PG stays writable only while at least 17 shards are up; the 4th host
|
||||
loss drops a PG to k = 16 shards, which still serves reads but blocks writes
|
||||
until recovery restores a 17th shard. min_size is never set below k+1 --
|
||||
permitting writes at k shards would risk data loss if another shard were lost
|
||||
mid-recovery.
|
||||
|
||||
This is a **single-site** guarantee with **no rack diversity**: every node sits
|
||||
in one datacenter (FSN1) and the failure domain is host, not rack. A rack, row,
|
||||
PDU, or switch fault that takes down more than m hosts at once, or a whole-
|
||||
datacenter loss (power, network, fire, flood), exceeds this ceiling -- there is
|
||||
no second site and no off-region copy. The pool protects against disk and node
|
||||
failure, not site failure. Off-site DR (a second region, or an external copy of
|
||||
the backup-of-record) is out of scope for this cluster and would need a separate
|
||||
replication path.
|
||||
|
||||
## block.db sizing
|
||||
|
||||
Rule of thumb: block.db ~ 4% of OSD data size.
|
||||
|
||||
@@ -40,7 +40,24 @@ Backups are written to `backups/<timestamp>/` on the Ansible controller
|
||||
|
||||
### Backup schedule
|
||||
|
||||
No automated schedule is configured. Run manually:
|
||||
**Scheduled (automated):** `mise run backup-timer` (or
|
||||
`scripts/ansible-play.sh backup-ceph.yml`) installs an on-node capture
|
||||
script (`/usr/local/sbin/ceph-backup.sh`) plus a systemd timer on the
|
||||
bootstrap node. The timer runs daily (with up to 1h of jitter) and writes a
|
||||
timestamped tarball to `/var/backups/ceph/` (root-only, `0700`), pruning
|
||||
tarballs older than `ceph_backup_retention_days` (default 14). Each tarball
|
||||
holds the recoverable cluster state: `fsid`, `ceph config dump`, monmap,
|
||||
osdmap, crushmap (binary + decompiled), `ceph osd tree`, `ceph orch ls` /
|
||||
`host ls`, and the RGW realm/zonegroup/zone. It does **not** contain secret
|
||||
keyrings. Opt a cluster out with `ceph_backup_enabled: false` in its
|
||||
group_vars.
|
||||
|
||||
**Offsite:** left as an operator decision. Set `ceph_backup_offsite_dest`
|
||||
(empty by default) to an rsync target (`user@host:/path/`) or S3 URI
|
||||
(`s3://bucket/prefix`) and the script ships each tarball after capture.
|
||||
|
||||
**Manual (`mise run backup`):** the controller-local export below is still
|
||||
available for ad-hoc snapshots. Run it manually:
|
||||
|
||||
- Before any cluster topology change (add/remove node or OSD)
|
||||
- Before Ceph version upgrades
|
||||
|
||||
@@ -0,0 +1,129 @@
|
||||
# Runbook: Upgrade Ceph (cephadm-orchestrated) and Roll Back
|
||||
|
||||
**When:** moving the cluster to a new patch release on the pinned Tentacle
|
||||
20.2.x train (e.g. 20.2.1 -> 20.2.2), or rolling back a bad upgrade.
|
||||
|
||||
**Time estimate:** 20-60 min for a small cluster; scales with daemon/OSD count
|
||||
(cephadm rolls one daemon type at a time and waits for health between steps).
|
||||
|
||||
**Playbook:** `upgrade-ceph.yml` (+ `tasks/ceph_upgrade.yml`). It is a
|
||||
deliberate, manually-invoked day-2 op and is **not** in `site.yml` -- converge
|
||||
never triggers an upgrade.
|
||||
|
||||
---
|
||||
|
||||
## What the playbook guarantees
|
||||
|
||||
1. **Preflight gate** (refuses otherwise): cluster is `HEALTH_OK`, no upgrade
|
||||
already in progress, and every OSD is up + in.
|
||||
2. **Rollback anchor recorded**: it prints the image(s) the daemons are running
|
||||
*now* before touching anything -- copy this for step-4 rollback.
|
||||
3. **Explicit target required**: no default image/version is baked in.
|
||||
4. **Bounded poll**: watches `ceph orch upgrade status`; fails loudly if
|
||||
cephadm pauses/stalls the upgrade (it auto-pauses on a failed step).
|
||||
5. **Post-checks**: asserts `HEALTH_OK` and that all daemons converged on one
|
||||
version, then prints a summary.
|
||||
6. **Idempotent**: if the cluster already runs the target image and is
|
||||
version-converged, it is a clean no-op.
|
||||
|
||||
## Upgrade
|
||||
|
||||
```bash
|
||||
# Preferred: pin the exact image (digest-pinnable, and what rollback consumes).
|
||||
scripts/ansible-play.sh upgrade-ceph.yml \
|
||||
-e ceph_upgrade_target_image=quay.io/ceph/ceph:v20.2.2
|
||||
|
||||
# Or by version (cephadm resolves the image):
|
||||
scripts/ansible-play.sh upgrade-ceph.yml -e ceph_upgrade_version=20.2.2
|
||||
```
|
||||
|
||||
Optional toggles:
|
||||
|
||||
```bash
|
||||
# Proceed on a benign HEALTH_WARN (review `ceph health detail` first):
|
||||
-e ceph_upgrade_allow_health_warn=true
|
||||
|
||||
# Widen the poll bound (default 120 * 30s = 60 min):
|
||||
-e ceph_upgrade_poll_retries=240 -e ceph_upgrade_poll_delay=30
|
||||
```
|
||||
|
||||
Preview the preflight/plan without starting anything:
|
||||
|
||||
```bash
|
||||
scripts/ansible-play.sh upgrade-ceph.yml \
|
||||
-e ceph_upgrade_target_image=quay.io/ceph/ceph:v20.2.2 --check
|
||||
```
|
||||
|
||||
## Rollback
|
||||
|
||||
cephadm upgrades roll one daemon type at a time and auto-**pause** on the first
|
||||
failure rather than tearing the cluster down. There is no destructive cutover,
|
||||
so rollback = point the orchestrator back at the **prior** image and let it
|
||||
converge.
|
||||
|
||||
```bash
|
||||
# 0. Prior image: the play PRINTED it ("ROLLBACK ANCHOR ..."). Otherwise:
|
||||
ceph orch ps --format json | \
|
||||
python3 -c 'import sys,json;print(sorted({d["container_image_name"] for d in json.load(sys.stdin)}))'
|
||||
|
||||
# 1. Stop the in-flight upgrade (or `pause`/`resume` to hold and inspect):
|
||||
ceph orch upgrade stop
|
||||
|
||||
# 2. Re-target the prior image (only daemons ahead of it move):
|
||||
ceph orch upgrade start --image <PRIOR_IMAGE>
|
||||
|
||||
# 3. Watch it converge:
|
||||
ceph orch upgrade status
|
||||
ceph -W cephadm
|
||||
watch ceph versions
|
||||
|
||||
# 4. Health checks once status shows not in_progress:
|
||||
ceph health detail # expect HEALTH_OK
|
||||
ceph versions # expect ONE overall version
|
||||
ceph osd stat # all OSDs up + in
|
||||
```
|
||||
|
||||
Re-running `upgrade-ceph.yml` with the prior image as the target performs the
|
||||
same health-gated rollback with all the same checks.
|
||||
|
||||
## Podman compatibility (cross-major upgrades)
|
||||
|
||||
The baseline role `dpkg`-holds podman (and the rest of its ecosystem) at its
|
||||
installed version so a stray `apt upgrade` cannot move the container runtime out
|
||||
from under a live cluster. Patch-train upgrades (20.2.x -> 20.2.y) run on the same
|
||||
podman, so the hold needs no attention. But a **cross-major Ceph upgrade** may
|
||||
require a newer podman per the cephadm compatibility matrix, and the hold will
|
||||
block the podman bump. When that applies:
|
||||
|
||||
```bash
|
||||
# On each node (drive via a serial, health-gated converge, not all at once):
|
||||
apt-mark unhold podman
|
||||
apt-get install -y --only-upgrade podman # to a version the cephadm matrix allows
|
||||
apt-mark hold podman # re-freeze at the new version
|
||||
```
|
||||
|
||||
Do this BEFORE `ceph orch upgrade start`, verify `podman version` on every node,
|
||||
then upgrade Ceph. Check the cephadm podman support matrix for the target release
|
||||
first; too-new podman can also break cephadm, so bump to a *supported* version,
|
||||
not merely the latest. `baseline_held_packages` controls which packages are held.
|
||||
|
||||
## Gotchas
|
||||
|
||||
- **No cross-major downgrade.** Rollback is only safe *within* the pinned
|
||||
20.2.x train (patch level). Do not use this to cross 20.x -> 19.x -- Ceph does
|
||||
not support it.
|
||||
- **Paused mid-upgrade.** If the poll assert fails, the upgrade paused. Inspect
|
||||
`ceph orch upgrade status` (`message` field) and `ceph -W cephadm`, fix the
|
||||
offending daemon/host, then `ceph orch upgrade resume` -- or stop and roll
|
||||
back per above.
|
||||
- **Bootstrap host in `--limit`.** All cluster ops delegate to the bootstrap
|
||||
node; do not `--limit` it out of the run.
|
||||
- **Registry reachability.** Nodes must be able to pull the target image. A
|
||||
stalled pull shows up as a paused upgrade with a pull error in the message.
|
||||
|
||||
## References
|
||||
|
||||
- `upgrade-ceph.yml` -- the playbook; its header carries the same rollback
|
||||
procedure inline.
|
||||
- `docs/runbooks/replace-host.md`, `docs/runbooks/add-node.md` -- related day-2
|
||||
ops that also drive `ceph` from the bootstrap node.
|
||||
+23
-4
@@ -1,11 +1,15 @@
|
||||
---
|
||||
# Security hardening for Ceph nodes.
|
||||
# Deploys nftables firewall rules and SSH hardening.
|
||||
# Run AFTER deploy-ceph.yml — cephadm needs unrestricted access during deploy.
|
||||
# Deploys the nftables firewall + the full sshd hardening drop-in.
|
||||
# Runs AFTER deploy-ceph.yml - cephadm needs unrestricted access during deploy.
|
||||
# NB: the security-critical sshd lockdown (root key-only, no password auth, locked
|
||||
# root password) is applied EARLY in the baseline role, not here, so a node is
|
||||
# never left root-password-open on the WAN during the deploy campaign; this stage
|
||||
# adds the firewall and the remaining sshd hardening on top.
|
||||
#
|
||||
# WARNING: This locks down inbound traffic. Ensure ceph_firewall_trusted_networks
|
||||
# includes all subnets that need to reach Ceph services. SSH (port 22) is always
|
||||
# open from any source to prevent lockout.
|
||||
# includes all subnets that need to reach Ceph services. SSH (port 22) stays open
|
||||
# from any source (ceph_firewall_ssh_any_source) to prevent lockout.
|
||||
#
|
||||
# Usage:
|
||||
# scripts/ansible-play.sh harden.yml
|
||||
@@ -14,6 +18,21 @@
|
||||
hosts: ceph_nodes
|
||||
become: true
|
||||
gather_facts: false
|
||||
# Blast-radius control (see baseline.yml): batch the roll + halt on any node
|
||||
# failure or cluster degradation. Especially important here - a bad nftables
|
||||
# ruleset that fences the fabric shows up as an unreachable batch (halts on
|
||||
# max_fail_percentage) or a degraded cluster at the post-batch health gate,
|
||||
# before it reaches the rest of the fleet.
|
||||
serial: "{{ ceph_converge_serial | default('10%') }}"
|
||||
max_fail_percentage: 0
|
||||
|
||||
pre_tasks:
|
||||
- name: Ceph-health gate (pre-batch)
|
||||
ansible.builtin.import_tasks: tasks/health_gate.yml
|
||||
|
||||
roles:
|
||||
- security
|
||||
|
||||
post_tasks:
|
||||
- name: Ceph-health gate (post-batch)
|
||||
ansible.builtin.import_tasks: tasks/health_gate.yml
|
||||
|
||||
@@ -0,0 +1,310 @@
|
||||
---
|
||||
# Cluster-wide Ansible config for spice (prod / htz-fsn1, 48x SX295).
|
||||
# Committed. Hand-maintained. Adapted from the painbox (SX295/Hetzner) precedent
|
||||
# with the network reshaped onto the bonded-25G fabric VLANs.
|
||||
#
|
||||
# NOTE: sections marked DRAFT are ceph data-protection / RGW decisions that must
|
||||
# be reviewed before `mise run deploy` against 48 nodes. They do NOT affect the
|
||||
# OS-install milestone (installimage), only the later cephadm convergence.
|
||||
|
||||
# === Naming ===
|
||||
cluster_name: spice
|
||||
cluster_role: ceph
|
||||
cluster_domain: prod.fsn1.htz.futo.cloud
|
||||
|
||||
# === Network ===
|
||||
# Ceph public/cluster traffic runs on the bonded 25G fabric VLANs (NetBox
|
||||
# FSN1-C1-PUBLIC / FSN1-C1-PRIVATE). Per-host IPs are 10.40.20.<host_index> /
|
||||
# 10.40.22.<host_index> (host_index from spice-hosts.yaml; gateway .1 = leaf IRB).
|
||||
# A third fabric VLAN, FSN1-C1-HOST-MGMT (124, 10.40.24.0/24), is the in-band
|
||||
# host-management path: the leaf trunks it to every server bond and the fabric
|
||||
# advertises 10.40.24.0/24 into the NetBird overlay, so it is the intended
|
||||
# ansible/SSH reach once the WAN is retired.
|
||||
# The DEFAULT ROUTE stays on the 1G WAN for now (Hetzner gateway); ansible reaches
|
||||
# nodes over their WAN address until the fabric is cut over. Revisit if/when the
|
||||
# default route moves off the WAN (to Ceph Public or Host-Mgmt).
|
||||
public_network: 10.40.20.0/23 # VLAN 120 (Ceph Public)
|
||||
cluster_network: 10.40.22.0/23 # VLAN 122 (Ceph Private)
|
||||
# Per-host Ceph BIND address on the fabric public VLAN. Derived from host_index the
|
||||
# same way as networkd_bond_vlans below (a per-host host_var, so it resolves per
|
||||
# host at play time). Lives INSIDE public_network -- unlike bond_ip, the 1G WAN
|
||||
# address used only for the ansible/SSH connection (ansible_host). (No
|
||||
# ceph_cluster_ip var: OSD replication binds to the cluster_network CIDR, which
|
||||
# cephadm resolves to each node's bond0.122 address on its own.)
|
||||
ceph_public_ip: "10.40.20.{{ host_index }}" # bond0.120, inside public_network
|
||||
# The address Ceph point-services advertise + bind on (mon --mon-ip, cephadm
|
||||
# host-add, dashboard monitoring URLs) = the fabric public IP. Ceph never touches
|
||||
# the WAN. A group_var, NOT a role default, because join/rgw/monitoring read it via
|
||||
# hostvars[<host>], which does not expose role defaults; here it resolves per host
|
||||
# to 10.40.20.<host_index>. sietch is flat with no ceph_public_ip, so there it
|
||||
# falls back to bond_ip (already inside its public_network).
|
||||
ceph_service_ip: "{{ ceph_public_ip | default(bond_ip) }}"
|
||||
# Fabric-only bind opt-in: cephadm restricts the RGW + monitoring daemon binds to
|
||||
# these CIDR(s) so nothing Ceph listens on the 1G WAN (mon/mgr/OSD already bind
|
||||
# public_network/cluster_network). The role default is empty (flat clusters bind
|
||||
# every interface, unchanged); spice pins the fabric public network.
|
||||
ceph_bind_networks:
|
||||
- "{{ public_network }}"
|
||||
# Bond of the two 25G NICs, LACP to match the leaf ae<k> aggregate (cluster-fabric).
|
||||
bond_mode: "802.3ad"
|
||||
# COEXISTENCE: keep the 1G WAN (oob_nic) on ifupdown exactly as installimage set
|
||||
# it (static IP + default route), and let networkd manage ONLY the bond + VLANs.
|
||||
# The WAN is never migrated, so a networkd error can't cost us the reachability
|
||||
# link (the painbox failure). networkd's commit leaves networking.service enabled
|
||||
# and an Unmanaged=yes guard fences oob_nic off. Flip to true only if the default
|
||||
# route ever moves onto the fabric and the WAN is retired.
|
||||
networkd_replace_ifupdown: false
|
||||
# Enable the networkd role (defaults false, so it is a no-op unless set). spice's
|
||||
# 25G fabric bond + VLANs are networkd-managed, so a reprovisioned node must bring
|
||||
# them up from the repo alone; committed here so the live state is reproducible
|
||||
# (previously supplied ad-hoc via -e). Safe: networkd_replace_ifupdown:false fences
|
||||
# the 1G WAN. Roll one node at a time on any live reconverge.
|
||||
networkd_enabled: true
|
||||
# The two Intel E800 (ice) 25G ports, verified uniform across all 48 SX295 nodes
|
||||
# (PCI 0000:c1:00.0 / .1). Consumed by the networkd role as the bond0 members.
|
||||
bond_interfaces:
|
||||
- enp193s0f0
|
||||
- enp193s0f1
|
||||
# 1G WAN (igb, PCI 0000:c5:00.0); holds the default route until the fabric cutover.
|
||||
oob_nic: enp197s0
|
||||
fabric_public_vlan: 120
|
||||
fabric_private_vlan: 122
|
||||
fabric_mgmt_vlan: 124
|
||||
dns_server: 185.12.64.1 # Hetzner recursive (reachable via WAN default route)
|
||||
# Time sync: Hetzner's own NTP (low-latency from FSN1, consistent across all 47).
|
||||
# Ceph mon quorum is skew-sensitive; pin the source rather than the distro pool.
|
||||
chrony_ntp_servers:
|
||||
- ntp1.hetzner.de
|
||||
- ntp2.hetzner.de
|
||||
- ntp3.hetzner.de
|
||||
|
||||
# Tagged VLAN sub-interfaces on the 25G LACP bond (consumed by the networkd role).
|
||||
# Addresses derive from host_index; masks match each network (public/private /23,
|
||||
# host-mgmt /24). No Gateway here -- the default route stays on the 1G WAN until
|
||||
# fabric cutover.
|
||||
# Jumbo (MTU 9000) is enabled on VLAN 122 (Ceph PRIVATE / OSD replication) ONLY.
|
||||
# That is a closed, homogeneous, all-under-our-control path where jumbo's packet/
|
||||
# interrupt savings on replication+recovery actually pay off. VLAN 120 (Ceph public,
|
||||
# where RGW faces heterogeneous 1500 S3 clients + the ~1400 NetBird overlay) and
|
||||
# VLAN 124 (mgmt) stay at 1500: jumbo on a client-facing path is a partial-blackhole
|
||||
# risk (small ops fine, large PUT/GET hang) for near-zero gain. The bond ceiling and
|
||||
# both 25G members auto-raise to the max child MTU so the jumbo VLAN is not capped.
|
||||
# PAIRED REQUIREMENT: the leaf server-LAG + the VLAN 122 IRB must be >= 9000 on the
|
||||
# QFX side before the cluster network passes jumbo; validate with a bidirectional
|
||||
# `ping -M do -s 8972` matrix on 10.40.22.0/23 + green ceph health before relying on it.
|
||||
networkd_bond_vlans:
|
||||
- id: "{{ fabric_public_vlan }}" # 120 - Ceph public (10.40.20.0/23)
|
||||
address: "10.40.20.{{ host_index }}/23"
|
||||
- id: "{{ fabric_private_vlan }}" # 122 - Ceph private (10.40.22.0/23)
|
||||
address: "10.40.22.{{ host_index }}/23"
|
||||
mtu: 9000
|
||||
- id: "{{ fabric_mgmt_vlan }}" # 124 - Host mgmt (10.40.24.0/24)
|
||||
address: "10.40.24.{{ host_index }}/24"
|
||||
|
||||
# Cluster-specific packages, merged with baseline_diag_packages. Spice hardware:
|
||||
# Intel E800 (ice) 25G NICs on an LACP fabric; direct-attach SATA HDDs + NVMe (no
|
||||
# SAS expander / Mellanox, so none of sietch's mstflint/ledmon/sg3-utils here).
|
||||
baseline_extra_packages:
|
||||
- lldpd # LLDP neighbor/switch discovery on the 25G fabric
|
||||
- ethtool # NIC/bond link, driver, and 25G speed inspection
|
||||
- tcpdump # packet capture for fabric/network debugging
|
||||
# Intel E810 (ice) DDP package (intel/ice/ddp/ice.pkg). WITHOUT it the NIC boots
|
||||
# into Safe Mode, whose crippled classifier drops LACP/LLDP control frames -> the
|
||||
# 25G bond never aggregates. The minimal Debian image omits it; baseline installs
|
||||
# it and (nic-firmware.yml) reboots once to load the DDP. Needs the non-free-firmware
|
||||
# apt component (present in the Hetzner Debian base).
|
||||
- firmware-misc-nonfree
|
||||
timezone: UTC
|
||||
|
||||
# === Ceph ===
|
||||
ceph_release: tentacle
|
||||
ceph_repo_url: "https://download.ceph.com/debian-{{ ceph_release }}/"
|
||||
ceph_repo_key_url: "https://download.ceph.com/keys/release.asc"
|
||||
|
||||
# === OS / Auth ===
|
||||
# installimage boots root; the post-install script authorizes the cluster
|
||||
# ansible-iac key for root, and ansible connects as root (matches painbox).
|
||||
admin_user: root
|
||||
provision_iac_ssh_key_path: "~/.ssh/id_ed25519_spice"
|
||||
|
||||
# 1P vault for cluster secret lookups.
|
||||
cluster_secrets_vault: yucca_tf_prod
|
||||
|
||||
# === Secret aliases (op inject -> vault_* at play time) ===
|
||||
ops_password: "{{ vault_ops_password }}"
|
||||
ceph_dashboard_user: admin
|
||||
ceph_dashboard_password: "{{ vault_ceph_dashboard_password }}"
|
||||
ceph_grafana_admin_user: admin
|
||||
ceph_grafana_admin_password: "{{ vault_grafana_admin_password }}"
|
||||
ceph_rgw_s3_user_access_key: "{{ vault_s3_restic_access_key }}"
|
||||
ceph_rgw_s3_user_secret_key: "{{ vault_s3_restic_secret_key }}"
|
||||
|
||||
# === Reprovision (installimage; reprovision_hetzner role) ===
|
||||
# Most knobs live in roles/reprovision_hetzner/defaults; override the operator-
|
||||
# facing ones here. VERIFY image name + boot_mode on node 1 (mise run reprovision
|
||||
# -- --limit spice-ceph-adelia -e confirm_wipe=true -e reprovision_mode=inspect).
|
||||
reprovision_boot_mode: bios # SX295 painbox precedent = BIOS (verified)
|
||||
reprovision_expect_nvme: 2
|
||||
reprovision_expect_sata: 14
|
||||
# reprovision_image: override only if node-1 inspect shows a different image path/name.
|
||||
|
||||
# === Storage (SX295: 2x NVMe RAID-1 OS + 14x SATA HDD per node) ===
|
||||
# block.db LVs carved from the NVMe RAID-1 vg0 at converge by ceph_deploy's
|
||||
# lvm-setup (NVMe-RAID branch): 14x 128G db-slots + one ssd-osd LV from the vg0
|
||||
# remainder minus reserve. HDDs become OSDs with block.db on those LVs.
|
||||
# Consumed ONLY by ceph_destroy's partition-6 SSD-OSD wipe (sietch's dual-SSD
|
||||
# shape). spice is NVMe-RAID: the ssd-osd is a vg0 LV removed by the generic VG/PV
|
||||
# cleanup, and the NVMe carries no partition 6, so this pattern is a deliberate
|
||||
# no-match here (the SSD-partition wipe finds nothing). cleanup.yml references it
|
||||
# unconditionally, so it must stay defined; the value is a don't-care on spice.
|
||||
ssd_model_pattern: "__no_ssd_osd_partitions_on_nvme_raid__"
|
||||
ceph_db_vg: vg0 # single NVMe-RAID VG created by installimage
|
||||
ceph_db_lv_size: "128G"
|
||||
ceph_db_lvs_per_node: 14
|
||||
ceph_ssd_osd_lv: ssd-osd # NVMe-backed OSD LV (fast replicated pool)
|
||||
ceph_ssd_osd_reserve_gib: 512 # vg0 free left unallocated for headroom/expansion
|
||||
|
||||
# OSD device map (by-path, consumed by ceph_deploy/osd-spec.yml.j2, NVMe-RAID shape:
|
||||
# no sas_path_prefix -> data path is /dev/disk/by-path/<path_phy>, db is /dev/<db>).
|
||||
# The 14x 20TB SATA HDDs sit on 3 AHCI controllers; the by-path layout is IDENTICAL
|
||||
# across 46 of the 47 nodes (verified 2026-07-10), so it lives here as a group_var
|
||||
# instead of 47 host_vars. Exception: spice-ceph-miguel has one disk on
|
||||
# 46:00.0-ata-4 (not 87:00.0-ata-4) and overrides ceph_hdd_osds in its host_vars.
|
||||
# db-slotN and ssd-osd LVs are carved from vg0 at deploy by lvm-setup.
|
||||
ceph_hdd_osds:
|
||||
- { path_phy: pci-0000:45:00.0-ata-1, db: vg0/db-slot0 }
|
||||
- { path_phy: pci-0000:45:00.0-ata-2, db: vg0/db-slot1 }
|
||||
- { path_phy: pci-0000:45:00.0-ata-3, db: vg0/db-slot2 }
|
||||
- { path_phy: pci-0000:45:00.0-ata-4, db: vg0/db-slot3 }
|
||||
- { path_phy: pci-0000:45:00.0-ata-5, db: vg0/db-slot4 }
|
||||
- { path_phy: pci-0000:45:00.0-ata-6, db: vg0/db-slot5 }
|
||||
- { path_phy: pci-0000:45:00.0-ata-7, db: vg0/db-slot6 }
|
||||
- { path_phy: pci-0000:45:00.0-ata-8, db: vg0/db-slot7 }
|
||||
- { path_phy: pci-0000:46:00.0-ata-1, db: vg0/db-slot8 }
|
||||
- { path_phy: pci-0000:46:00.0-ata-2, db: vg0/db-slot9 }
|
||||
- { path_phy: pci-0000:87:00.0-ata-1, db: vg0/db-slot10 }
|
||||
- { path_phy: pci-0000:87:00.0-ata-2, db: vg0/db-slot11 }
|
||||
- { path_phy: pci-0000:87:00.0-ata-3, db: vg0/db-slot12 }
|
||||
- { path_phy: pci-0000:87:00.0-ata-4, db: vg0/db-slot13 }
|
||||
ceph_ssd_osds:
|
||||
- { lv: vg0/ssd-osd }
|
||||
|
||||
# === RGW ===
|
||||
# EC 16+4, HOST failure domain -- capacity-max profile (80% usable, m=4) for the
|
||||
# beta behind michael. Host domain trades rack-loss protection for capacity: a
|
||||
# destroyed rack means data loss, accepted because the beta is wipeable and the
|
||||
# API tier is already single-rack. Rationale + alternatives in Brain Box
|
||||
# spice-ec-profile-analysis.md. Hold m=4 (host domain is the only durability
|
||||
# left); do not widen past k+m=20.
|
||||
ceph_rgw_realm: spice
|
||||
ceph_rgw_zonegroup: eu-central-1
|
||||
ceph_rgw_zonegroup_api_name: eu-central-1
|
||||
ceph_rgw_zone: prod-z1
|
||||
ceph_rgw_ec_profile: ec-k16m4-host
|
||||
ceph_rgw_ec_k: 16
|
||||
ceph_rgw_ec_m: 4
|
||||
ceph_rgw_ec_failure_domain: host
|
||||
ceph_rgw_ec_device_class: hdd
|
||||
# Start the bulk EC data pool near its autoscaler target. The role default (128)
|
||||
# is sized for ~36 OSDs; spice has 672 HDD OSDs on a 20-chunk pool. pg_num 4096
|
||||
# => ~122 PG-shards/OSD ( pg_num * (k+m) / osds ), inside the recommended 100-250
|
||||
# per OSD for large BlueStore clusters (default target 100, docs recommend 200;
|
||||
# >500 hurts). 2048 would be only ~61/OSD (too low); 4096 is the right power of 2.
|
||||
# Starting low would PG-split during the first fill.
|
||||
ceph_pg_init_data: 4096
|
||||
ceph_rgw_data_pool: "{{ ceph_rgw_zone }}.rgw.buckets.data"
|
||||
ceph_rgw_index_pool: "{{ ceph_rgw_zone }}.rgw.buckets.index"
|
||||
ceph_rgw_extra_pool: "{{ ceph_rgw_zone }}.rgw.buckets.non-ec"
|
||||
ceph_rgw_dns_name: s3.{{ cluster_domain }}
|
||||
ceph_rgw_replicated_size: 3
|
||||
ceph_rgw_replicated_min_size: 2
|
||||
# Bootstrap default for any pool that does not pin its own size (footgun guard:
|
||||
# an ad-hoc or future pool on a prod cluster should not land at single-replica
|
||||
# 2/1). Applied via cluster-spec.yml.j2 at cephadm bootstrap; role default is
|
||||
# 2/1 so sietch is unchanged. Named RGW pools still pin their own size above.
|
||||
ceph_osd_pool_default_size: 3
|
||||
ceph_osd_pool_default_min_size: 2
|
||||
# --- Device-class CRUSH pool placement ---
|
||||
# crush-rules.yml creates replicated_ssd/replicated_hdd but binds nothing; every
|
||||
# replicated pool otherwise falls to the default class-agnostic rule, so the
|
||||
# omap-heavy bucket index lands on HDD (~98.5% of weight) while the 47-OSD NVMe
|
||||
# ssd-osd tier sits idle. Pin the latency-sensitive index/metadata/mgr pools to
|
||||
# the SSD tier; keep the bulk non-EC pool on HDD. The EC data pool is pinned to
|
||||
# hdd via ceph_rgw_ec_device_class (the EC profile), not here. OSD device classes
|
||||
# are forced at creation in osd-spec.yml.j2 (ssd-osd -> ssd, HDD -> hdd), and
|
||||
# verify.yml asserts every intended pool carries its rule and the ssd class has
|
||||
# live OSDs.
|
||||
ceph_rgw_ssd_crush_rule: replicated_ssd
|
||||
ceph_rgw_hdd_crush_rule: replicated_hdd
|
||||
ceph_rgw_ssd_pools:
|
||||
- "{{ ceph_rgw_index_pool }}"
|
||||
- "{{ ceph_rgw_zone }}.rgw.meta"
|
||||
- .rgw.root
|
||||
- "{{ ceph_rgw_zone }}.rgw.log"
|
||||
- "{{ ceph_rgw_zone }}.rgw.control"
|
||||
- .mgr
|
||||
ceph_rgw_hdd_pools:
|
||||
- "{{ ceph_rgw_extra_pool }}"
|
||||
ceph_rgw_count_per_host: 1
|
||||
ceph_rgw_port: 443
|
||||
# RGW/beast terminates TLS on 443 with a 10-year self-signed cert: rgw.yml
|
||||
# generates /etc/ceph/rgw-ssl.{crt,key} and cephadm distributes it to every RGW
|
||||
# daemon via the service spec. Clients reach RGW over the fabric / NetBird, so the
|
||||
# cert SANs cover the s3.<domain> name (+ wildcard for virtual-hosted buckets) and
|
||||
# each node's fabric IP (ceph_service_ip). Self-signed: clients trust-on-first-use
|
||||
# or pin the CA. Falkenstein/Saxony/DE matches the FSN1 site (cosmetic for a
|
||||
# self-signed cert; only the CN + SANs are validated).
|
||||
ceph_rgw_ssl: true
|
||||
ceph_rgw_ssl_cert_days: 3650 # 10 years
|
||||
ceph_rgw_ssl_cert_subject_c: DE
|
||||
ceph_rgw_ssl_cert_subject_st: Saxony
|
||||
ceph_rgw_ssl_cert_subject_l: Falkenstein
|
||||
ceph_rgw_ssl_cert_subject_o: FUTO
|
||||
ceph_rgw_ssl_cert_email: yucca@futo.org
|
||||
ceph_rgw_scheme: "{{ 'https' if ceph_rgw_ssl else 'http' }}"
|
||||
# --- RGW ingress: health-checked S3 VIP (haproxy + keepalived, TLS terminate) ---
|
||||
# A cephadm ingress service puts a keepalived floating VIP with haproxy health
|
||||
# checks in front of the 47 RGW daemons, in TLS TERMINATION mode (haproxy
|
||||
# `mode http`: haproxy terminates the client TLS on the VIP with the reused RGW
|
||||
# self-signed cert, then re-encrypts to each beast backend). Without it a dead RGW
|
||||
# node blackholes ~1/47 of S3 connections; with it the node is health-checked out
|
||||
# of the VIP. The VIP sits on the fabric public network (10.40.20.0/23, bond0.120);
|
||||
# NetBird already advertises 10.40.20.0/23 into the overlay, so the VIP is
|
||||
# reachable, and DNS points s3.<domain> at it (that DNS cutover is a separate TF
|
||||
# apply that MUST follow the ingress being live, not precede it). The VIP is added
|
||||
# to the RGW cert SAN (rgw.yml) so a client hitting it validates; virtual-hosted
|
||||
# buckets keep working because mode http forwards the Host header to the wildcard
|
||||
# cert. Terminate (not L4 passthrough) because passthrough needs a post-20.2.2
|
||||
# cephadm field (use_tcp_mode_over_rgw) absent on the pinned 20.2.2.
|
||||
#
|
||||
# Preconditions are satisfied here: ceph_bind_networks non-empty (VIP:443 vs beast
|
||||
# per-node-IP:443 are distinct sockets) and ceph_rgw_ssl true (cert exists).
|
||||
ceph_rgw_ingress_enabled: true
|
||||
ceph_rgw_ingress_vip: "10.40.20.250"
|
||||
ceph_rgw_ingress_vip_prefix: 23 # matches public_network 10.40.20.0/23
|
||||
ceph_rgw_s3_user_uid: svc-yucca-restic
|
||||
ceph_rgw_s3_user_display_name: "yucca/restic service account"
|
||||
# Metrics-worker RGW admin user (read-only): the yucca-metrics-worker service
|
||||
# reads bucket usage from RadosGW. Keys are TF-minted (SPICE_METRICS_WORKER_*)
|
||||
# and injected via secrets.yml.tpl as vault_metrics_worker_*; rgw.yml (Step 14.5)
|
||||
# creates the user with these caps. Without this block the RGW phase fails with
|
||||
# AnsibleUndefinedVariable. Mirror of the sietch definition.
|
||||
ceph_rgw_metrics_user_uid: metrics-worker
|
||||
ceph_rgw_metrics_user_display_name: "yucca/metrics-worker RGW admin (read-only)"
|
||||
ceph_rgw_metrics_user_access_key: "{{ vault_metrics_worker_access_key }}"
|
||||
ceph_rgw_metrics_user_secret_key: "{{ vault_metrics_worker_secret_key }}"
|
||||
ceph_rgw_metrics_user_caps: "buckets=read;usage=read;metadata=read;users=read"
|
||||
|
||||
# === Monitoring Stack ===
|
||||
ceph_prometheus_port: 9095
|
||||
ceph_grafana_port: 3000
|
||||
ceph_alertmanager_port: 9093
|
||||
# ceph_grafana_admin_user/password are defined once in the Secret aliases block above.
|
||||
|
||||
# === Security / firewall ===
|
||||
# Spice is fabric-only: every Ceph daemon binds the 25G fabric (mon/mgr/OSD via
|
||||
# public_network/cluster_network; RGW + monitoring via ceph_bind_networks) and
|
||||
# never listens on the 1G WAN. RGW is therefore scoped to the trusted networks
|
||||
# rather than left world-open (ceph_firewall_rgw_any_source defaults true for
|
||||
# flat/dev clusters; spice pins it closed).
|
||||
ceph_firewall_rgw_any_source: false
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-adelia
|
||||
bond_ip: 178.63.139.248
|
||||
hetzner_server_number: 3008187
|
||||
host_index: 4
|
||||
mon: true
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-alexus
|
||||
bond_ip: 178.63.139.254
|
||||
hetzner_server_number: 3008189
|
||||
host_index: 5
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-alyssa
|
||||
bond_ip: 178.63.139.228
|
||||
hetzner_server_number: 3008190
|
||||
host_index: 6
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-ardith
|
||||
bond_ip: 178.63.139.227
|
||||
hetzner_server_number: 3008191
|
||||
host_index: 7
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-athena
|
||||
bond_ip: 178.63.139.225
|
||||
hetzner_server_number: 3008192
|
||||
host_index: 8
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-bernie
|
||||
bond_ip: 178.63.139.226
|
||||
hetzner_server_number: 3008193
|
||||
host_index: 9
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-braden
|
||||
bond_ip: 178.63.139.240
|
||||
hetzner_server_number: 3008194
|
||||
host_index: 10
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-callie
|
||||
bond_ip: 178.63.139.243
|
||||
hetzner_server_number: 3008195
|
||||
host_index: 11
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-catina
|
||||
bond_ip: 178.63.139.244
|
||||
hetzner_server_number: 3008196
|
||||
host_index: 12
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-cletus
|
||||
bond_ip: 178.63.139.253
|
||||
hetzner_server_number: 3008197
|
||||
host_index: 13
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-curtis
|
||||
bond_ip: 178.63.139.252
|
||||
hetzner_server_number: 3008198
|
||||
host_index: 14
|
||||
mon: true
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-darion
|
||||
bond_ip: 178.63.139.251
|
||||
hetzner_server_number: 3008199
|
||||
host_index: 15
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-deidra
|
||||
bond_ip: 178.63.139.250
|
||||
hetzner_server_number: 3008200
|
||||
host_index: 16
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-deonte
|
||||
bond_ip: 178.63.139.249
|
||||
hetzner_server_number: 3008201
|
||||
host_index: 17
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-dorian
|
||||
bond_ip: 178.63.139.242
|
||||
hetzner_server_number: 3008202
|
||||
host_index: 18
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-edythe
|
||||
bond_ip: 178.63.139.247
|
||||
hetzner_server_number: 3008203
|
||||
host_index: 19
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-eloise
|
||||
bond_ip: 178.63.139.246
|
||||
hetzner_server_number: 3008204
|
||||
host_index: 20
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-evelyn
|
||||
bond_ip: 178.63.139.245
|
||||
hetzner_server_number: 3008205
|
||||
host_index: 21
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-gaylon
|
||||
bond_ip: 178.63.139.241
|
||||
hetzner_server_number: 3008206
|
||||
host_index: 22
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-graham
|
||||
bond_ip: 178.63.139.239
|
||||
hetzner_server_number: 3008207
|
||||
host_index: 23
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-hayley
|
||||
bond_ip: 178.63.139.230
|
||||
hetzner_server_number: 3012008
|
||||
host_index: 24
|
||||
mon: true
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-howell
|
||||
bond_ip: 178.63.139.222
|
||||
hetzner_server_number: 3012009
|
||||
host_index: 25
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-jacque
|
||||
bond_ip: 178.63.139.221
|
||||
hetzner_server_number: 3012010
|
||||
host_index: 26
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-javier
|
||||
bond_ip: 178.63.139.207
|
||||
hetzner_server_number: 3012011
|
||||
host_index: 27
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-jewell
|
||||
bond_ip: 178.63.139.210
|
||||
hetzner_server_number: 3012012
|
||||
host_index: 28
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-joseph
|
||||
bond_ip: 178.63.139.218
|
||||
hetzner_server_number: 3012013
|
||||
host_index: 29
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-kassie
|
||||
bond_ip: 178.63.139.208
|
||||
hetzner_server_number: 3012014
|
||||
host_index: 30
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-kelsea
|
||||
bond_ip: 178.63.139.217
|
||||
hetzner_server_number: 3012015
|
||||
host_index: 31
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-kurtis
|
||||
bond_ip: 178.63.139.234
|
||||
hetzner_server_number: 3012016
|
||||
host_index: 32
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-lemuel
|
||||
bond_ip: 178.63.139.214
|
||||
hetzner_server_number: 3012017
|
||||
host_index: 33
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-lizzie
|
||||
bond_ip: 178.63.139.213
|
||||
hetzner_server_number: 3014572
|
||||
host_index: 34
|
||||
mon: true
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-lucius
|
||||
bond_ip: 178.63.139.212
|
||||
hetzner_server_number: 3014573
|
||||
host_index: 35
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-marian
|
||||
bond_ip: 178.63.139.211
|
||||
hetzner_server_number: 3014574
|
||||
host_index: 36
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-mattie
|
||||
bond_ip: 178.63.139.237
|
||||
hetzner_server_number: 3014575
|
||||
host_index: 37
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,24 @@
|
||||
---
|
||||
hostname_short: spice-ceph-miguel
|
||||
bond_ip: 178.63.139.216
|
||||
hetzner_server_number: 3014576
|
||||
host_index: 38
|
||||
mon: false
|
||||
# OVERRIDE of the group_vars ceph_hdd_osds: miguel has one HDD on 46:00.0-ata-4
|
||||
# (not 87:00.0-ata-4 like the other 46 nodes) -- verified 2026-07-10. host_vars
|
||||
# fully replaces the group_var, so the whole 14-disk map is repeated here.
|
||||
ceph_hdd_osds:
|
||||
- { path_phy: pci-0000:45:00.0-ata-1, db: vg0/db-slot0 }
|
||||
- { path_phy: pci-0000:45:00.0-ata-2, db: vg0/db-slot1 }
|
||||
- { path_phy: pci-0000:45:00.0-ata-3, db: vg0/db-slot2 }
|
||||
- { path_phy: pci-0000:45:00.0-ata-4, db: vg0/db-slot3 }
|
||||
- { path_phy: pci-0000:45:00.0-ata-5, db: vg0/db-slot4 }
|
||||
- { path_phy: pci-0000:45:00.0-ata-6, db: vg0/db-slot5 }
|
||||
- { path_phy: pci-0000:45:00.0-ata-7, db: vg0/db-slot6 }
|
||||
- { path_phy: pci-0000:45:00.0-ata-8, db: vg0/db-slot7 }
|
||||
- { path_phy: pci-0000:46:00.0-ata-1, db: vg0/db-slot8 }
|
||||
- { path_phy: pci-0000:46:00.0-ata-2, db: vg0/db-slot9 }
|
||||
- { path_phy: pci-0000:46:00.0-ata-4, db: vg0/db-slot10 }
|
||||
- { path_phy: pci-0000:87:00.0-ata-1, db: vg0/db-slot11 }
|
||||
- { path_phy: pci-0000:87:00.0-ata-2, db: vg0/db-slot12 }
|
||||
- { path_phy: pci-0000:87:00.0-ata-3, db: vg0/db-slot13 }
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-murphy
|
||||
bond_ip: 178.63.139.232
|
||||
hetzner_server_number: 3014577
|
||||
host_index: 39
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,9 @@
|
||||
---
|
||||
hostname_short: spice-ceph-noreen
|
||||
bond_ip: 178.63.139.220
|
||||
hetzner_server_number: 3014578
|
||||
host_index: 40
|
||||
mon: false
|
||||
# ceph_hdd_osds: defined in group_vars/all/vars.yml (SX295 by-path map is
|
||||
# uniform across nodes). Override here ONLY if a node's by-path differs
|
||||
# (e.g. spice-ceph-miguel). Verify with: ls -l /dev/disk/by-path/pci-*-ata-*
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-philip
|
||||
bond_ip: 178.63.139.219
|
||||
hetzner_server_number: 3014579
|
||||
host_index: 41
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-raymon
|
||||
bond_ip: 178.63.139.223
|
||||
hetzner_server_number: 3014580
|
||||
host_index: 42
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-romona
|
||||
bond_ip: 178.63.139.236
|
||||
hetzner_server_number: 3014581
|
||||
host_index: 43
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-serena
|
||||
bond_ip: 178.63.139.229
|
||||
hetzner_server_number: 3014582
|
||||
host_index: 44
|
||||
mon: true
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-shanna
|
||||
bond_ip: 178.63.139.209
|
||||
hetzner_server_number: 3014583
|
||||
host_index: 45
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-shelby
|
||||
bond_ip: 178.63.139.224
|
||||
hetzner_server_number: 3014584
|
||||
host_index: 46
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-sommer
|
||||
bond_ip: 178.63.139.215
|
||||
hetzner_server_number: 3014585
|
||||
host_index: 47
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-sylvia
|
||||
bond_ip: 178.63.139.235
|
||||
hetzner_server_number: 3014586
|
||||
host_index: 48
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-theron
|
||||
bond_ip: 178.63.139.238
|
||||
hetzner_server_number: 3014587
|
||||
host_index: 49
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-trista
|
||||
bond_ip: 178.63.139.233
|
||||
hetzner_server_number: 3014588
|
||||
host_index: 50
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
hostname_short: spice-ceph-virgie
|
||||
bond_ip: 178.63.139.231
|
||||
hetzner_server_number: 3014589
|
||||
host_index: 51
|
||||
mon: false
|
||||
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
|
||||
@@ -7,6 +7,12 @@ cluster_role: ceph
|
||||
cluster_domain: staging.austin.int.futo.cloud
|
||||
public_network: 10.10.10.0/24
|
||||
cluster_network: 10.10.10.0/24
|
||||
# The address Ceph point-services advertise + bind on (mon --mon-ip, cephadm
|
||||
# host-add, dashboard monitoring URLs). Flat cluster: no ceph_public_ip, so this
|
||||
# resolves to bond_ip (10.10.10.9x) - identical to the pre-fabric behavior. A
|
||||
# group_var, NOT a role default, because it is read via hostvars[<host>] which
|
||||
# does not expose role defaults.
|
||||
ceph_service_ip: "{{ ceph_public_ip | default(bond_ip) }}"
|
||||
gateway: 10.10.10.1
|
||||
dns_server: 10.10.10.1
|
||||
bond_mode: active-backup
|
||||
|
||||
@@ -29,14 +29,19 @@
|
||||
hosts: ceph_nodes
|
||||
become: true
|
||||
gather_facts: false
|
||||
serial: 1
|
||||
serial: "{{ networkd_serial | default(1) }}"
|
||||
max_fail_percentage: 0 # any node fails (e.g. WAN drops -> unreachable) -> halt
|
||||
|
||||
pre_tasks:
|
||||
# Ceph-health gate applies only to a live cluster. Pre-Ceph clusters (spice
|
||||
# bringup: coexistence activation before any cephadm deploy) have no
|
||||
# ceph_bootstrap group and no ceph binary, so skip it automatically.
|
||||
- name: Preflight -- check Ceph health
|
||||
ansible.builtin.command: ceph health --format json
|
||||
register: ceph_health_pre
|
||||
changed_when: false
|
||||
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
|
||||
when: (groups['ceph_bootstrap'] | default([]) | length) > 0
|
||||
|
||||
- name: Assert cluster is healthy
|
||||
ansible.builtin.assert:
|
||||
@@ -46,6 +51,7 @@
|
||||
Ceph is {{ (ceph_health_pre.stdout | from_json).status }}.
|
||||
Fix cluster health before migrating.
|
||||
success_msg: "Ceph: {{ (ceph_health_pre.stdout | from_json).status }}"
|
||||
when: (groups['ceph_bootstrap'] | default([]) | length) > 0
|
||||
|
||||
roles:
|
||||
- role: networkd
|
||||
@@ -56,7 +62,9 @@
|
||||
register: ceph_health_post
|
||||
changed_when: false
|
||||
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
|
||||
when: (groups['ceph_bootstrap'] | default([]) | length) > 0
|
||||
|
||||
- name: Report Ceph health
|
||||
ansible.builtin.debug:
|
||||
msg: "Ceph: {{ (ceph_health_post.stdout | from_json).status }}"
|
||||
when: (groups['ceph_bootstrap'] | default([]) | length) > 0
|
||||
|
||||
@@ -0,0 +1,55 @@
|
||||
---
|
||||
# Reprovision Hetzner SX295 ceph nodes via the Robot API + installimage.
|
||||
# Ansible-native replacement for the out-of-band installimage step:
|
||||
# arm rescue (uri) -> hw reset (uri) -> wait for rescue -> [zap OSD HDDs] ->
|
||||
# installimage -> reboot -> verify installed OS. Then hand off to `mise run deploy`.
|
||||
#
|
||||
# ALWAYS DESTRUCTIVE to the NVMe OS disks. NEVER fold this into deploy/converge.
|
||||
# Launch ONLY via `mise run reprovision` (wraps op run --env-file=../../tf/.env.prod,
|
||||
# which materializes HETZNER_ROBOT_* and lets ansible-play.sh op-inject secrets).
|
||||
# Set CEPH_ENV first (e.g. inventories/prod-htz-fsn1/spice/inventory.ini).
|
||||
#
|
||||
# Inspect node 1 (arm rescue + reset + assert disks/boot-mode, NO wipe):
|
||||
# mise run reprovision -- --limit spice-ceph-adelia -e confirm_wipe=true -e reprovision_mode=inspect
|
||||
# Canary (one node, full wipe + verify; verifying it unlocks fan-out):
|
||||
# mise run reprovision -- --limit spice-ceph-adelia -e confirm_wipe=true
|
||||
# Fan-out (remaining nodes, batched; gated on the canary marker):
|
||||
# mise run reprovision -- --limit 'ceph_nodes:!spice-ceph-adelia' -e confirm_wipe=true -e reprovision_serial=5
|
||||
# Resume after a halt (skip already-provisioned nodes, converge only the rest):
|
||||
# mise run reprovision -- --limit ceph_nodes -e confirm_wipe=true -e allow_fanout=true \
|
||||
# -e reprovision_skip_if_provisioned=true -e reprovision_serial=5
|
||||
# Day-2 reimage of a node (also zap its OSD HDDs; the ceph-safety gate verifies
|
||||
# ok-to-stop against a live mon and refuses if the OSDs are not safe to lose):
|
||||
# mise run reprovision -- --limit <host> -e confirm_wipe=true -e reprovision_wipe_osd_disks=true \
|
||||
# -e reprovision_ceph_safety=strict
|
||||
|
||||
- name: Reprovision ceph nodes (rescue -> installimage -> installed OS)
|
||||
hosts: ceph_nodes
|
||||
gather_facts: false
|
||||
become: false
|
||||
serial: "{{ reprovision_serial | default(1) }}"
|
||||
max_fail_percentage: 0 # any node fails -> halt the batch (no further wipes)
|
||||
pre_tasks:
|
||||
# run_once + included through the role so role defaults load and --limit can't
|
||||
# skip it (a hosts: localhost play would be). Asserts creds/keys, enforces the
|
||||
# fan-out gate, registers rescue keys, publishes fingerprints to localhost facts.
|
||||
- name: Preflight + register robot keys (once) # noqa: run-once[task]
|
||||
ansible.builtin.include_role:
|
||||
name: reprovision_hetzner
|
||||
tasks_from: preflight
|
||||
run_once: true
|
||||
roles:
|
||||
- role: reprovision_hetzner
|
||||
|
||||
- name: Verify reprovisioned nodes after reboot
|
||||
hosts: ceph_nodes
|
||||
gather_facts: false
|
||||
become: false
|
||||
tasks:
|
||||
- name: Skip verify in inspect mode (no install happened)
|
||||
ansible.builtin.meta: end_play
|
||||
when: reprovision_mode | default('install') != 'install'
|
||||
- name: Verify installed OS marker
|
||||
ansible.builtin.include_role:
|
||||
name: reprovision_hetzner
|
||||
tasks_from: verify
|
||||
@@ -1,6 +1,10 @@
|
||||
---
|
||||
# Galaxy collections are pinned to the currently-resolved major and floored at
|
||||
# a known-good minor. Compatible-release ranges (no lockfile in this repo) keep
|
||||
# a fleet-wide Ceph campaign reproducible: a new major cannot be pulled in
|
||||
# mid-run and change module behavior between the first and last node.
|
||||
collections:
|
||||
- name: ansible.posix
|
||||
version: ">=2.0.0"
|
||||
version: ">=2.0.0,<3.0.0"
|
||||
- name: community.general
|
||||
version: ">=9.0.0"
|
||||
version: ">=12.0.0,<13.0.0"
|
||||
|
||||
@@ -2,6 +2,22 @@
|
||||
# Post-boot OS baseline — runs via ansible-iac after provisioning.
|
||||
# Configures everything that doesn't need to be in the chroot.
|
||||
|
||||
# --- iac automation access (convergence twin of the reprovision -x chroot) ---
|
||||
# Ensures root's iac key + the ansible-iac sudo account on already-installed nodes
|
||||
# without reimaging. baseline_iac_pubkey derives from the cluster's iac key; the
|
||||
# iac-access task is gated on provision_iac_ssh_key_path so clusters without one
|
||||
# skip. Override baseline_iac_pubkey directly to authorize a different key.
|
||||
baseline_iac_user: ansible-iac
|
||||
baseline_iac_pubkey: >-
|
||||
{{ lookup('file', (provision_iac_ssh_key_path | expanduser) ~ '.pub') }}
|
||||
|
||||
# --- Intel ice NIC firmware (DDP) ---
|
||||
# nic-firmware.yml reboots an E810 host once to load the DDP (exit Safe Mode) after
|
||||
# firmware-misc-nonfree is installed. Set false to install the package but skip the
|
||||
# activating reboot (an operator can reboot on their own schedule).
|
||||
baseline_nic_firmware_activate: true
|
||||
baseline_nic_firmware_reboot_timeout: 600
|
||||
|
||||
# --- ops user ---
|
||||
baseline_ops_user: ops
|
||||
baseline_ops_uid: 1001
|
||||
@@ -55,8 +71,28 @@ baseline_diag_packages:
|
||||
# ...) sets its own list or leaves it empty.
|
||||
baseline_extra_packages: []
|
||||
|
||||
# Load-bearing host packages frozen at their installed version via a dpkg hold, so
|
||||
# a stray `apt upgrade` cannot move them out from under a live cluster. We run NO
|
||||
# unattended-upgrades by design; upgrading these is a deliberate, health-gated act
|
||||
# (unhold -> upgrade -> re-hold, ideally through the serial converge). The default
|
||||
# covers cephadm's container runtime (podman + ecosystem) and the chrony time
|
||||
# source that mon quorum depends on. Diagnostic/ops tools are intentionally NOT
|
||||
# held (a version drift there is harmless). Extend per cluster, or set [] to
|
||||
# disable. For an exact-version pin (rather than freeze-at-installed) add an apt
|
||||
# preferences entry; the hold covers the common "do not let it drift" need.
|
||||
baseline_held_packages: "{{ baseline_podman_packages + ['chrony'] }}"
|
||||
|
||||
# --- Time sync (chrony) ---
|
||||
# Pinned NTP sources for chrony.conf. Default to the Debian pool as a safe,
|
||||
# explicit fallback; each cluster overrides with a consistent low-latency set
|
||||
# (e.g. spice -> Hetzner NTP). Owned by chrony.yml (install+config+enable), NOT
|
||||
# baseline_enable_services, so system.yml never starts a not-yet-installed unit.
|
||||
chrony_ntp_servers:
|
||||
- 0.debian.pool.ntp.org
|
||||
- 1.debian.pool.ntp.org
|
||||
- 2.debian.pool.ntp.org
|
||||
|
||||
# --- Services ---
|
||||
baseline_enable_services:
|
||||
- dbus
|
||||
- chrony
|
||||
- podman.socket
|
||||
|
||||
@@ -0,0 +1,23 @@
|
||||
---
|
||||
- name: Restart lldpd
|
||||
ansible.builtin.systemd:
|
||||
name: lldpd
|
||||
state: restarted
|
||||
|
||||
- name: Restart chrony
|
||||
ansible.builtin.systemd:
|
||||
name: chrony
|
||||
state: restarted
|
||||
|
||||
# sshd drop-in changes notify "Validate sshd config"; it runs sshd -t against the
|
||||
# real combined config and only then notifies "Reload sshd" (handlers fire in
|
||||
# listed order), so an invalid config halts the play before any reload.
|
||||
- name: Validate sshd config
|
||||
ansible.builtin.command: /usr/sbin/sshd -t
|
||||
changed_when: false
|
||||
notify: Reload sshd
|
||||
|
||||
- name: Reload sshd
|
||||
ansible.builtin.systemd:
|
||||
name: ssh
|
||||
state: reloaded
|
||||
@@ -0,0 +1,24 @@
|
||||
---
|
||||
# Time sync (chrony) for Ceph. Owns the package + config + service so the source
|
||||
# set is explicit and versioned, not left to the distro default pool. Ceph mon
|
||||
# quorum is skew-sensitive, so this is a hard prerequisite, not a nicety.
|
||||
|
||||
- name: Install chrony
|
||||
ansible.builtin.apt:
|
||||
name: chrony
|
||||
state: present
|
||||
|
||||
- name: Deploy chrony.conf (pinned NTP sources)
|
||||
ansible.builtin.template:
|
||||
src: chrony.conf.j2
|
||||
dest: /etc/chrony/chrony.conf
|
||||
owner: root
|
||||
group: root
|
||||
mode: '0644'
|
||||
notify: Restart chrony
|
||||
|
||||
- name: Enable and start chrony
|
||||
ansible.builtin.systemd:
|
||||
name: chrony
|
||||
enabled: true
|
||||
state: started
|
||||
@@ -0,0 +1,21 @@
|
||||
---
|
||||
# Freeze the load-bearing host packages at their installed version.
|
||||
#
|
||||
# We deliberately do NOT run unattended-upgrades: an apt upgrade that moved the
|
||||
# cephadm container runtime (podman + ecosystem) or the time source (chrony, which
|
||||
# mon quorum depends on) out from under a live cluster is exactly the kind of
|
||||
# uncoordinated change we avoid. Instead we HOLD what matters (dpkg selection) so
|
||||
# `apt upgrade` skips them; upgrading them becomes a deliberate, health-gated act
|
||||
# (unhold -> upgrade -> re-hold, ideally through the serial converge). Diagnostic
|
||||
# and ops tools are intentionally not held (a version drift there is harmless).
|
||||
#
|
||||
# Runs last in the baseline role so every held package is already installed.
|
||||
# Idempotent: dpkg_selections only reports changed when the selection differs.
|
||||
# Set baseline_held_packages: [] to disable holds entirely.
|
||||
|
||||
- name: Freeze load-bearing packages at their installed version (dpkg hold)
|
||||
ansible.builtin.dpkg_selections:
|
||||
name: "{{ item }}"
|
||||
selection: hold
|
||||
loop: "{{ baseline_held_packages }}"
|
||||
when: baseline_held_packages | length > 0
|
||||
@@ -0,0 +1,39 @@
|
||||
---
|
||||
# Idempotent convergence twin of the reprovision -x chroot (reprovision_hetzner/
|
||||
# templates/post-install.sh.j2). Seeds the SAME login access -- root's iac key and
|
||||
# the ansible-iac sudo account -- so nodes that were installed WITHOUT the chroot
|
||||
# (e.g. the first 36 imaged before it existed) converge to the same state without
|
||||
# reimaging. Fresh nodes already have it from the chroot; this just re-affirms.
|
||||
#
|
||||
# Runs as the connection user (spice: root). Gated on provision_iac_ssh_key_path so
|
||||
# clusters that do not define an iac key skip cleanly. Standalone:
|
||||
# scripts/ansible-play.sh <spice> baseline.yml --tags iac_access --limit <hosts>
|
||||
|
||||
- name: Ensure the ansible-iac automation user
|
||||
ansible.builtin.user:
|
||||
name: "{{ baseline_iac_user }}"
|
||||
shell: /bin/bash
|
||||
create_home: true
|
||||
state: present
|
||||
|
||||
- name: Passwordless sudo for ansible-iac
|
||||
ansible.builtin.copy:
|
||||
content: "{{ baseline_iac_user }} ALL=(ALL) NOPASSWD:ALL\n"
|
||||
dest: "/etc/sudoers.d/{{ baseline_iac_user }}"
|
||||
owner: root
|
||||
group: root
|
||||
mode: '0440'
|
||||
validate: "visudo -cf %s"
|
||||
|
||||
- name: Authorize the iac key on ansible-iac
|
||||
ansible.posix.authorized_key:
|
||||
user: "{{ baseline_iac_user }}"
|
||||
key: "{{ baseline_iac_pubkey }}"
|
||||
state: present
|
||||
|
||||
# Non-exclusive: root is the reachability guarantee and must never be pruned here.
|
||||
- name: Authorize the iac key on root
|
||||
ansible.posix.authorized_key:
|
||||
user: root
|
||||
key: "{{ baseline_iac_pubkey }}"
|
||||
state: present
|
||||
@@ -0,0 +1,20 @@
|
||||
---
|
||||
# Configure lldpd (installed via baseline_extra_packages on fabric clusters).
|
||||
# Advertises the node over LLDP so the leaf switches and NetBox can discover it
|
||||
# on the 25G fabric. Imported from main.yml only when 'lldpd' is in the package
|
||||
# list, so clusters without lldpd are unaffected.
|
||||
|
||||
- name: Deploy lldpd configuration
|
||||
ansible.builtin.template:
|
||||
src: lldpd.conf.j2
|
||||
dest: /etc/lldpd.d/10-ceph.conf
|
||||
owner: root
|
||||
group: root
|
||||
mode: '0644'
|
||||
notify: Restart lldpd
|
||||
|
||||
- name: Enable and start lldpd
|
||||
ansible.builtin.systemd:
|
||||
name: lldpd
|
||||
enabled: true
|
||||
state: started
|
||||
@@ -5,6 +5,18 @@
|
||||
# (created by provision_host during initial OS install).
|
||||
# Everything here is convergeable via normal deploy pipeline.
|
||||
|
||||
- name: Ensure iac automation access (root key + ansible-iac)
|
||||
ansible.builtin.import_tasks: iac-access.yml
|
||||
when: provision_iac_ssh_key_path is defined
|
||||
tags: [iac_access, users]
|
||||
|
||||
# Immediately after the iac key is authorized on root: close root password login
|
||||
# and password auth (do NOT wait for the harden stage). Prevents an installimage
|
||||
# node from sitting root-password-open on the WAN during the deploy campaign.
|
||||
- name: Harden SSH early (root key-only, no password auth, lock root)
|
||||
ansible.builtin.import_tasks: ssh-hardening.yml
|
||||
tags: [ssh, hardening, security]
|
||||
|
||||
- name: Configure ops user
|
||||
ansible.builtin.import_tasks: users.yml
|
||||
tags: [users]
|
||||
@@ -13,6 +25,28 @@
|
||||
ansible.builtin.import_tasks: packages.yml
|
||||
tags: [packages]
|
||||
|
||||
- name: Configure time sync (chrony) for Ceph
|
||||
ansible.builtin.import_tasks: chrony.yml
|
||||
tags: [chrony, time, system]
|
||||
|
||||
# After packages (firmware-misc-nonfree) are installed: load the Intel ice DDP so
|
||||
# E810 NICs leave Safe Mode. Self-gating (no-op on non-ice hosts / already-loaded).
|
||||
- name: Activate Intel ice NIC firmware (DDP / exit Safe Mode)
|
||||
ansible.builtin.import_tasks: nic-firmware.yml
|
||||
tags: [packages, nic_firmware, network]
|
||||
|
||||
- name: Configure lldpd for fabric discovery
|
||||
ansible.builtin.import_tasks: lldpd.yml
|
||||
when: "'lldpd' in (baseline_extra_packages | default([]))"
|
||||
tags: [lldpd, network]
|
||||
|
||||
- name: Configure /etc/hosts and services
|
||||
ansible.builtin.import_tasks: system.yml
|
||||
tags: [system]
|
||||
|
||||
# Last: freeze the load-bearing packages installed above (podman ecosystem,
|
||||
# chrony) so a stray apt upgrade cannot move them under a live cluster. Runs after
|
||||
# every install so nothing is held before it exists.
|
||||
- name: Freeze load-bearing package versions (no unattended upgrades)
|
||||
ansible.builtin.import_tasks: hold-packages.yml
|
||||
tags: [packages, hold]
|
||||
|
||||
@@ -0,0 +1,71 @@
|
||||
---
|
||||
# Intel E810 (ice) NICs need their DDP package (intel/ice/ddp/ice.pkg, provided by
|
||||
# firmware-misc-nonfree, installed via baseline_extra_packages). Without it the NIC
|
||||
# boots into Safe Mode: the packet classifier is crippled and drops reserved-multicast
|
||||
# control frames (LACP 01:80:c2:00:00:02, LLDP ..:0e), so the 25G fabric bond never
|
||||
# aggregates even though the switch is correctly configured. The DDP only loads when
|
||||
# ice (re)probes, so once the package is present a single reboot clears Safe Mode.
|
||||
#
|
||||
# Safe by design: only acts on hosts with an ice NIC; only reboots when the NIC is
|
||||
# actually in Safe Mode AND the DDP file is present (so the reboot will fix it, and
|
||||
# healthy nodes are never rebooted); refuses to reboot if the DDP is still missing
|
||||
# (would loop). The WAN is on a separate igb NIC, so the reboot keeps the node
|
||||
# reachable (coexistence). Idempotent: a converged node is a no-op.
|
||||
|
||||
- name: Find an ice-bound netdev (is this an E810 host?)
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
for d in /sys/bus/pci/drivers/ice/0000:*/net/*; do
|
||||
[ -e "$d" ] && { basename "$d"; break; }
|
||||
done
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: baseline_ice_nic
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
|
||||
- name: Activate the ice DDP (exit Safe Mode) when needed
|
||||
when: baseline_ice_nic.stdout | trim | length > 0
|
||||
block:
|
||||
- name: Check whether the ice NIC is in Safe Mode # noqa: risky-shell-pipe
|
||||
# In Safe Mode advanced ops are disabled: --show-fec fails "not supported".
|
||||
# NB: no `set -o pipefail` here -- ethtool exits non-zero on "not supported",
|
||||
# which pipefail would propagate and mask grep's match, always yielding "ok".
|
||||
ansible.builtin.shell: |
|
||||
if ethtool --show-fec {{ baseline_ice_nic.stdout | trim }} 2>&1 | grep -qi "not supported"; then
|
||||
echo safemode
|
||||
else
|
||||
echo ok
|
||||
fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: baseline_ice_mode
|
||||
changed_when: false
|
||||
|
||||
# Gate on the dpkg DB (deterministic), not a file stat: after a large apt
|
||||
# transaction the DDP file can lag briefly on disk, but dpkg-query reflects the
|
||||
# install immediately -- and if the package is installed, the DDP file is present.
|
||||
- name: Check firmware-misc-nonfree is installed (provides the DDP)
|
||||
ansible.builtin.command: dpkg-query -W -f=${Status} firmware-misc-nonfree
|
||||
register: baseline_ice_ddp_pkg
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
|
||||
- name: Warn - NIC in Safe Mode but firmware-misc-nonfree not installed (do not reboot; would loop)
|
||||
ansible.builtin.debug:
|
||||
msg: >-
|
||||
ice NIC {{ baseline_ice_nic.stdout | trim }} is in Safe Mode but
|
||||
firmware-misc-nonfree is not installed. NOT rebooting. Check that it is in
|
||||
baseline_extra_packages and that the non-free-firmware apt component is enabled.
|
||||
when:
|
||||
- "'safemode' in baseline_ice_mode.stdout"
|
||||
- "'install ok installed' not in baseline_ice_ddp_pkg.stdout"
|
||||
|
||||
- name: Reboot once to load the ice DDP and exit Safe Mode
|
||||
ansible.builtin.reboot:
|
||||
reboot_timeout: "{{ baseline_nic_firmware_reboot_timeout }}"
|
||||
msg: "Loading Intel ice DDP (exit Safe Mode) after firmware-misc-nonfree install"
|
||||
when:
|
||||
- baseline_nic_firmware_activate | bool
|
||||
- "'safemode' in baseline_ice_mode.stdout"
|
||||
- "'install ok installed' in baseline_ice_ddp_pkg.stdout"
|
||||
@@ -0,0 +1,42 @@
|
||||
---
|
||||
# Early SSH lockdown. Runs in baseline (right after iac-access authorizes the iac
|
||||
# key on root), NOT gated behind the harden stage -- so a freshly installimaged
|
||||
# node is never left with root PASSWORD login exposed on the public WAN before
|
||||
# harden runs (installimage ships PermitRootLogin yes + a root password). The
|
||||
# security role's 50-hardening.conf applies the FULL sshd hardening later; this is
|
||||
# the security-critical subset only (root key-only, no password auth, locked root
|
||||
# password), using the SAME values so the two drop-ins never disagree.
|
||||
#
|
||||
# Safe on both clusters and idempotent on live nodes: ansible reaches every node
|
||||
# by KEY (root on spice, ansible-iac on sietch) and iac-access has already
|
||||
# authorized the key on root, so disabling password auth and locking the root
|
||||
# password cannot lock anyone out. The Validate->Reload handler chain reloads
|
||||
# (never restarts) only on change, after sshd -t passes.
|
||||
|
||||
- name: Deploy the baseline SSH lockdown drop-in
|
||||
ansible.builtin.copy:
|
||||
dest: /etc/ssh/sshd_config.d/10-baseline-ssh.conf
|
||||
content: |
|
||||
# Managed by ansible (roles/baseline/tasks/ssh-hardening.yml). Early
|
||||
# root/password lockdown; roles/security 50-hardening.conf adds the rest.
|
||||
PermitRootLogin prohibit-password
|
||||
PasswordAuthentication no
|
||||
owner: root
|
||||
group: root
|
||||
mode: '0644'
|
||||
notify: Validate sshd config
|
||||
tags: [ssh, hardening, security]
|
||||
|
||||
- name: Lock the root password (root is key-only; ansible/cephadm use keys)
|
||||
ansible.builtin.user:
|
||||
name: root
|
||||
password_lock: true
|
||||
tags: [ssh, hardening, security]
|
||||
|
||||
# Supersede the one-off interim drop-in from the pre-baseline BLOCKER remediation.
|
||||
- name: Remove the interim root-password drop-in if present
|
||||
ansible.builtin.file:
|
||||
path: /etc/ssh/sshd_config.d/00-disable-root-password.conf
|
||||
state: absent
|
||||
notify: Validate sshd config
|
||||
tags: [ssh, hardening, security]
|
||||
@@ -17,3 +17,10 @@
|
||||
enabled: true
|
||||
state: started
|
||||
loop: "{{ baseline_enable_services }}"
|
||||
|
||||
# timezone: UTC is declared in group_vars but nothing set it, so nodes drifted to
|
||||
# the image default (Europe/Berlin) and logs render in CEST. Idempotent, reboot-free.
|
||||
- name: Set the system timezone
|
||||
community.general.timezone:
|
||||
name: "{{ timezone }}"
|
||||
when: timezone is defined
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
---
|
||||
# ops user — human-interactive account.
|
||||
# ops user - human-interactive account.
|
||||
# Convergeable: running this again fixes drift in password, sudo, shell.
|
||||
|
||||
- name: Ensure ops user exists
|
||||
@@ -27,7 +27,12 @@
|
||||
- name: Set ops user password
|
||||
ansible.builtin.user:
|
||||
name: "{{ baseline_ops_user }}"
|
||||
password: "{{ ops_password | password_hash('sha512') }}"
|
||||
# password_hash draws a RANDOM salt by default, so the hash differs every run
|
||||
# and the user module rewrites /etc/shadow (reporting 'changed') on every
|
||||
# converge. Derive a STABLE salt from the password (one-way sha256, so the
|
||||
# public salt in /etc/shadow leaks nothing) to make the hash deterministic and
|
||||
# the task idempotent, while still reconciling an out-of-band password change.
|
||||
password: "{{ ops_password | password_hash('sha512', salt=(ops_password | hash('sha256'))[:16]) }}"
|
||||
no_log: true
|
||||
|
||||
- name: Configure ops sudo
|
||||
|
||||
@@ -0,0 +1,15 @@
|
||||
# {{ ansible_managed }}
|
||||
# Time sync for Ceph. Mons will not form/keep quorum with clock skew, so pin a
|
||||
# consistent, low-latency source set (chrony_ntp_servers) rather than the distro
|
||||
# default pool. makestep allows a large one-off correction on freshly-imaged nodes
|
||||
# before the cluster forms; after that chrony only slews.
|
||||
|
||||
{% for s in chrony_ntp_servers %}
|
||||
server {{ s }} iburst
|
||||
{% endfor %}
|
||||
|
||||
driftfile /var/lib/chrony/chrony.drift
|
||||
makestep 1.0 3
|
||||
rtcsync
|
||||
logdir /var/log/chrony
|
||||
leapsectz right/UTC
|
||||
@@ -1,8 +1,10 @@
|
||||
127.0.0.1 localhost
|
||||
|
||||
# Ceph cluster nodes
|
||||
# Ceph cluster nodes. Map each name to ceph_service_ip -- the fabric IP on spice
|
||||
# (bond0.120), falling back to bond_ip on flat clusters like sietch -- so in-cluster
|
||||
# name resolution matches where Ceph binds, not the 1G WAN.
|
||||
{% for host in groups['ceph_nodes'] %}
|
||||
{{ hostvars[host]['bond_ip'] }} {{ hostvars[host]['hostname_short'] }}.{{ cluster_domain }} {{ hostvars[host]['hostname_short'] }}
|
||||
{{ hostvars[host]['ceph_service_ip'] }} {{ hostvars[host]['hostname_short'] }}.{{ cluster_domain }} {{ hostvars[host]['hostname_short'] }}
|
||||
{% endfor %}
|
||||
|
||||
::1 localhost ip6-localhost ip6-loopback
|
||||
|
||||
@@ -0,0 +1,7 @@
|
||||
# {{ ansible_managed }}
|
||||
# lldpd config for ceph fabric nodes (read from /etc/lldpd.d/ at daemon start).
|
||||
# Port ID as the interface name (readable in NetBox / on the switch) and a
|
||||
# cluster-tagged system description. lldpd advertises on all interfaces by
|
||||
# default; the leaf only cares about the bonded 25G members.
|
||||
configure lldp portidsubtype ifname
|
||||
configure system description "{{ cluster_name | default('ceph') }} {{ cluster_role | default('ceph') }} node"
|
||||
@@ -0,0 +1,32 @@
|
||||
---
|
||||
# Scheduled Ceph cluster-state backup defaults.
|
||||
#
|
||||
# The capture is read-only (dumps config/topology maps via the on-node admin
|
||||
# keyring) and writes to a root-only local dir, so it is universally safe. Set
|
||||
# ceph_backup_enabled: false in a cluster's group_vars to opt out.
|
||||
ceph_backup_enabled: true
|
||||
|
||||
# Local landing dir for tarballs. Root-only (0700): the capture set is
|
||||
# config/topology only, but keep it locked down regardless.
|
||||
ceph_backup_dir: /var/backups/ceph
|
||||
|
||||
# Prune tarballs older than N days on each run. 0 disables pruning.
|
||||
ceph_backup_retention_days: 14
|
||||
|
||||
# Where the capture script is installed on the bootstrap node.
|
||||
ceph_backup_script_path: /usr/local/sbin/ceph-backup.sh
|
||||
|
||||
# Offsite sync target - EMPTY BY DEFAULT (operator decision; do not hardcode a
|
||||
# destination). When set the script ships each tarball after capture:
|
||||
# rsync form: user@host:/path/ or rsync://host/module/path
|
||||
# S3 form: s3://bucket/prefix (requires the aws CLI on the bootstrap node)
|
||||
# TODO(operator): choose + wire an offsite destination for true DR.
|
||||
ceph_backup_offsite_dest: ""
|
||||
|
||||
# systemd timer schedule + jitter. Daily with an up-to-1h randomized delay so a
|
||||
# fleet of clusters does not stampede a shared offsite target at the same second.
|
||||
ceph_backup_oncalendar: "daily"
|
||||
ceph_backup_randomized_delay: "1h"
|
||||
|
||||
# Kick one immediate capture at install time so the operator sees it work.
|
||||
ceph_backup_run_now: true
|
||||
@@ -0,0 +1,4 @@
|
||||
---
|
||||
- name: Reload systemd
|
||||
ansible.builtin.systemd:
|
||||
daemon_reload: true
|
||||
@@ -0,0 +1,14 @@
|
||||
---
|
||||
# No dependencies - deploys a self-contained capture script + systemd timer on
|
||||
# the bootstrap node. Independent of the deploy pipeline; run deliberately via
|
||||
# backup-ceph.yml. Complements post-deploy-capture.yml (secrets -> 1P) and
|
||||
# backup-config.yml (one-shot controller-local export) with a SCHEDULED,
|
||||
# on-node config/topology backup.
|
||||
dependencies: []
|
||||
|
||||
galaxy_info:
|
||||
author: FUTO
|
||||
license: AGPL-3.0-only
|
||||
role_name: ceph_backup
|
||||
description: Scheduled capture of Ceph cluster config/topology state to local tarballs for DR
|
||||
min_ansible_version: "2.19"
|
||||
@@ -0,0 +1,80 @@
|
||||
---
|
||||
# Deploy the scheduled cluster-state backup: capture script + systemd
|
||||
# service/timer on the bootstrap node. Read-only capture to a root-only local
|
||||
# dir, so it is universally safe; gate with ceph_backup_enabled to opt out.
|
||||
|
||||
- name: Scheduled backup install
|
||||
when: ceph_backup_enabled | bool
|
||||
block:
|
||||
- name: Ensure the backup directory exists (root-only)
|
||||
ansible.builtin.file:
|
||||
path: "{{ ceph_backup_dir }}"
|
||||
state: directory
|
||||
owner: root
|
||||
group: root
|
||||
mode: '0700'
|
||||
|
||||
- name: Deploy the ceph-backup.sh capture script
|
||||
ansible.builtin.template:
|
||||
src: ceph-backup.sh.j2
|
||||
dest: "{{ ceph_backup_script_path }}"
|
||||
owner: root
|
||||
group: root
|
||||
mode: '0700'
|
||||
|
||||
- name: Deploy the ceph-backup.service unit
|
||||
ansible.builtin.template:
|
||||
src: ceph-backup.service.j2
|
||||
dest: /etc/systemd/system/ceph-backup.service
|
||||
owner: root
|
||||
group: root
|
||||
mode: '0644'
|
||||
notify: Reload systemd
|
||||
|
||||
- name: Deploy the ceph-backup.timer unit
|
||||
ansible.builtin.template:
|
||||
src: ceph-backup.timer.j2
|
||||
dest: /etc/systemd/system/ceph-backup.timer
|
||||
owner: root
|
||||
group: root
|
||||
mode: '0644'
|
||||
notify: Reload systemd
|
||||
|
||||
- name: Reload systemd now so the units are current before enabling
|
||||
ansible.builtin.meta: flush_handlers
|
||||
|
||||
- name: Enable and start the backup timer
|
||||
ansible.builtin.systemd:
|
||||
name: ceph-backup.timer
|
||||
enabled: true
|
||||
state: started
|
||||
|
||||
# Starting the oneshot service blocks until ExecStart returns, so this both
|
||||
# proves the script runs and leaves a first tarball to inspect.
|
||||
- name: Trigger one immediate backup to validate the install
|
||||
ansible.builtin.systemd:
|
||||
name: ceph-backup.service
|
||||
state: started
|
||||
when: ceph_backup_run_now | bool
|
||||
|
||||
- name: Find the most recent backup tarball
|
||||
ansible.builtin.find:
|
||||
paths: "{{ ceph_backup_dir }}"
|
||||
patterns: 'ceph-state-*.tar.gz'
|
||||
register: ceph_backup_tarballs
|
||||
when: ceph_backup_run_now | bool
|
||||
|
||||
- name: Report backup install status
|
||||
ansible.builtin.debug:
|
||||
msg: >-
|
||||
Scheduled backups installed on {{ inventory_hostname }}: timer
|
||||
{{ ceph_backup_oncalendar }} (jitter {{ ceph_backup_randomized_delay }}),
|
||||
dir {{ ceph_backup_dir }}, retention {{ ceph_backup_retention_days }}d,
|
||||
offsite {{ ('set: ' ~ ceph_backup_offsite_dest)
|
||||
if (ceph_backup_offsite_dest | length > 0)
|
||||
else 'DISABLED (empty - operator TODO)' }}.
|
||||
Latest tarball:
|
||||
{{ (ceph_backup_tarballs.files | map(attribute='path') | sort | last)
|
||||
if (ceph_backup_run_now | bool
|
||||
and (ceph_backup_tarballs.files | default([]) | length) > 0)
|
||||
else 'none yet (timer will produce one)' }}
|
||||
@@ -0,0 +1,13 @@
|
||||
# {{ ansible_managed }}
|
||||
# Oneshot capture of Ceph cluster config/topology state. Triggered by
|
||||
# ceph-backup.timer; can also be run on demand: systemctl start ceph-backup.service
|
||||
[Unit]
|
||||
Description=Ceph cluster-state backup (config/topology to {{ ceph_backup_dir }})
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
# Belt-and-suspenders: any file the script writes lands root-only.
|
||||
UMask=0077
|
||||
ExecStart={{ ceph_backup_script_path }}
|
||||
@@ -0,0 +1,132 @@
|
||||
#!/usr/bin/env bash
|
||||
# {{ ansible_managed }}
|
||||
# ceph-backup.sh - capture recoverable Ceph cluster state to a timestamped
|
||||
# tarball under a local backup dir, prune old tarballs, optionally sync offsite.
|
||||
#
|
||||
# Runs on the bootstrap node as root and uses the on-node admin keyring; it
|
||||
# embeds NO credentials. Installed + scheduled by ansible/ceph/backup-ceph.yml.
|
||||
#
|
||||
# Does NOT dump secret keyrings. The backup dir is root-only (0700) and the
|
||||
# capture set is config/topology only: fsid, ceph config dump, monmap, osdmap,
|
||||
# crushmap (+ decompiled), osd tree, orch ls/host ls, RGW realm/zonegroup/zone.
|
||||
set -euo pipefail
|
||||
|
||||
# --- Config (rendered from ansible vars; environment overrides win) ---
|
||||
BACKUP_DIR="${CEPH_BACKUP_DIR:-{{ ceph_backup_dir }}}"
|
||||
RETENTION_DAYS="${CEPH_BACKUP_RETENTION_DAYS:-{{ ceph_backup_retention_days }}}"
|
||||
# Offsite target: EMPTY by default. Operator decision - set to an rsync target
|
||||
# (user@host:/path or rsync://...) or an S3 URI (s3://bucket/prefix) to ship
|
||||
# offsite. Leaving it empty skips the offsite step entirely.
|
||||
OFFSITE_DEST="${CEPH_BACKUP_OFFSITE_DEST:-{{ ceph_backup_offsite_dest }}}"
|
||||
|
||||
TS="$(date -u +%Y%m%dT%H%M%SZ)"
|
||||
STAGE="$(mktemp -d "${TMPDIR:-/tmp}/ceph-backup.XXXXXX")"
|
||||
OUT="${STAGE}/ceph-state-${TS}"
|
||||
mkdir -p "${OUT}"
|
||||
trap 'rm -rf "${STAGE}"' EXIT
|
||||
|
||||
log() { echo "ceph-backup: $*" >&2; }
|
||||
|
||||
# cap <outfile> <cmd...>: run a capture command; keep going if a subsystem is
|
||||
# absent (e.g. a cluster with no RGW realm) so one gap does not abort the run.
|
||||
cap() {
|
||||
local out="$1"
|
||||
shift
|
||||
if "$@" >"${OUT}/${out}" 2>"${OUT}/${out}.err"; then
|
||||
rm -f "${OUT}/${out}.err"
|
||||
else
|
||||
log "WARN: capture failed: ${out} (kept ${out}.err)"
|
||||
fi
|
||||
}
|
||||
|
||||
# --- Capture set ---
|
||||
cap fsid.txt ceph fsid
|
||||
cap ceph-status.txt ceph status
|
||||
cap config-dump.json ceph config dump --format json
|
||||
cap config-dump.txt ceph config dump
|
||||
cap osd-tree.txt ceph osd tree
|
||||
cap osd-tree.json ceph osd tree --format json
|
||||
cap osd-dump.json ceph osd dump --format json
|
||||
cap mon-dump.json ceph mon dump --format json
|
||||
cap orch-services.yaml ceph orch ls --format yaml
|
||||
cap orch-hosts.yaml ceph orch host ls --format yaml
|
||||
|
||||
# Binary maps + human-readable decodes (decode tools are best-effort).
|
||||
if ceph mon getmap -o "${OUT}/monmap.bin" 2>"${OUT}/monmap.err"; then
|
||||
rm -f "${OUT}/monmap.err"
|
||||
if command -v monmaptool >/dev/null 2>&1; then
|
||||
monmaptool --print "${OUT}/monmap.bin" >"${OUT}/monmap.txt" 2>/dev/null || true
|
||||
fi
|
||||
else
|
||||
log "WARN: monmap capture failed (kept monmap.err)"
|
||||
fi
|
||||
|
||||
if ceph osd getmap -o "${OUT}/osdmap.bin" 2>"${OUT}/osdmap.err"; then
|
||||
rm -f "${OUT}/osdmap.err"
|
||||
if command -v osdmaptool >/dev/null 2>&1; then
|
||||
osdmaptool --print "${OUT}/osdmap.bin" >"${OUT}/osdmap.txt" 2>/dev/null || true
|
||||
fi
|
||||
else
|
||||
log "WARN: osdmap capture failed (kept osdmap.err)"
|
||||
fi
|
||||
|
||||
if ceph osd getcrushmap -o "${OUT}/crushmap.bin" 2>"${OUT}/crushmap.err"; then
|
||||
rm -f "${OUT}/crushmap.err"
|
||||
if command -v crushtool >/dev/null 2>&1; then
|
||||
crushtool -d "${OUT}/crushmap.bin" -o "${OUT}/crushmap.txt" 2>/dev/null || true
|
||||
fi
|
||||
else
|
||||
log "WARN: crushmap capture failed (kept crushmap.err)"
|
||||
fi
|
||||
|
||||
# RGW multisite config (absent on clusters with no realm - non-fatal).
|
||||
cap rgw-realm.json radosgw-admin realm get
|
||||
cap rgw-zonegroup.json radosgw-admin zonegroup get
|
||||
cap rgw-zone.json radosgw-admin zone get
|
||||
|
||||
# Provenance.
|
||||
{
|
||||
echo "captured_at=${TS}"
|
||||
echo "host=$(hostname -f 2>/dev/null || hostname)"
|
||||
echo "ceph_version=$(ceph --version 2>/dev/null || echo unknown)"
|
||||
} >"${OUT}/MANIFEST.txt"
|
||||
|
||||
# --- Package ---
|
||||
mkdir -p "${BACKUP_DIR}"
|
||||
chmod 0700 "${BACKUP_DIR}"
|
||||
TARBALL="${BACKUP_DIR}/ceph-state-${TS}.tar.gz"
|
||||
tar -C "${STAGE}" -czf "${TARBALL}" "ceph-state-${TS}"
|
||||
chmod 0600 "${TARBALL}"
|
||||
log "wrote ${TARBALL}"
|
||||
|
||||
# --- Prune ---
|
||||
if [ "${RETENTION_DAYS}" -gt 0 ] 2>/dev/null; then
|
||||
find "${BACKUP_DIR}" -maxdepth 1 -type f -name 'ceph-state-*.tar.gz' \
|
||||
-mtime "+${RETENTION_DAYS}" -print -delete \
|
||||
| while read -r p; do log "pruned ${p}"; done || true
|
||||
fi
|
||||
|
||||
# --- Optional offsite sync (skipped unless OFFSITE_DEST is set) ---
|
||||
if [ -n "${OFFSITE_DEST}" ]; then
|
||||
case "${OFFSITE_DEST}" in
|
||||
s3://*)
|
||||
if command -v aws >/dev/null 2>&1; then
|
||||
aws s3 cp "${TARBALL}" "${OFFSITE_DEST%/}/$(basename "${TARBALL}")" \
|
||||
&& log "synced to ${OFFSITE_DEST}"
|
||||
else
|
||||
log "WARN: OFFSITE_DEST is S3 but the aws CLI is not installed - skipping offsite sync"
|
||||
fi
|
||||
;;
|
||||
*)
|
||||
if command -v rsync >/dev/null 2>&1; then
|
||||
rsync -a "${TARBALL}" "${OFFSITE_DEST}" && log "synced to ${OFFSITE_DEST}"
|
||||
else
|
||||
log "WARN: rsync not installed - skipping offsite sync"
|
||||
fi
|
||||
;;
|
||||
esac
|
||||
else
|
||||
log "offsite sync disabled (CEPH_BACKUP_OFFSITE_DEST empty)"
|
||||
fi
|
||||
|
||||
log "done"
|
||||
@@ -0,0 +1,13 @@
|
||||
# {{ ansible_managed }}
|
||||
# Runs ceph-backup.service on a schedule. Persistent=true catches up a missed
|
||||
# run (host powered off at the scheduled time) on the next boot.
|
||||
[Unit]
|
||||
Description=Scheduled Ceph cluster-state backup
|
||||
|
||||
[Timer]
|
||||
OnCalendar={{ ceph_backup_oncalendar }}
|
||||
RandomizedDelaySec={{ ceph_backup_randomized_delay }}
|
||||
Persistent=true
|
||||
|
||||
[Install]
|
||||
WantedBy=timers.target
|
||||
@@ -2,7 +2,7 @@
|
||||
# Ceph release train. Upstream publishes per-release apt trees at
|
||||
# download.ceph.com/debian-<release>/, and each tree has subdirectories
|
||||
# per Debian codename. prerequisites.yml pins the codename to 'bookworm'
|
||||
# in the sources.list entry — when upgrading the base OS to Trixie,
|
||||
# in the sources.list entry - when upgrading the base OS to Trixie,
|
||||
# flip the codename there. This split keeps the Ceph release and the
|
||||
# Debian release independently versionable.
|
||||
ceph_release: tentacle
|
||||
@@ -16,7 +16,106 @@ cephadm_install_method: repo # 'repo' or 'curl'
|
||||
# splitting under load which tanks performance during the first fill.
|
||||
# These values are sized for ~36 OSDs. Scale proportionally for larger
|
||||
# clusters.
|
||||
ceph_pg_init_data: 128 # EC data pool — bulk of I/O
|
||||
ceph_pg_init_data: 128 # EC data pool - bulk of I/O
|
||||
ceph_pg_init_index: 16 # bucket index
|
||||
ceph_pg_init_non_ec: 16 # multipart uploads
|
||||
ceph_pg_init_meta: 8 # .rgw.root, .meta, .log, .control
|
||||
|
||||
# RGW EC data-pool min_size (write-availability floor). Ceph defaults an EC pool
|
||||
# to min_size = k+1, which is the correct floor: at k+1 up shards a PG stays
|
||||
# writable, so writes tolerate up to m-1 host losses while reads still tolerate m.
|
||||
# Pin it EXPLICITLY (get-then-set in rgw.yml) so the value is documented and
|
||||
# cannot drift. NEVER set below k+1: permitting writes at k shards risks data loss
|
||||
# if another shard is lost mid-recovery. Derived from ceph_rgw_ec_k so it tracks
|
||||
# whatever profile a cluster runs (spice 16+4 -> 17, sietch 8+3 -> 9).
|
||||
ceph_rgw_ec_min_size: "{{ (ceph_rgw_ec_k | int) + 1 }}"
|
||||
|
||||
# MGR daemon count (1 active + standbys). cephadm's default and upstream guidance
|
||||
# is a small fixed count; deploying one per host (e.g. 47 on spice) is wasteful and
|
||||
# non-standard. count:N lets cephadm place them (caps at the host count on small
|
||||
# clusters like sietch).
|
||||
ceph_mgr_count: 3
|
||||
|
||||
# --- Service bind/advertise policy (per-cluster; fabric-only on spice) ---
|
||||
# ceph_service_ip (the address Ceph point-services advertise + bind on: mon
|
||||
# --mon-ip, the cephadm host-add address, the dashboard monitoring URLs) is
|
||||
# defined in each cluster's group_vars, NOT here. join/rgw/monitoring read it
|
||||
# through hostvars[<host>], and role defaults are INVISIBLE via hostvars (a role
|
||||
# default raises "HostVarsVars object has no attribute ceph_service_ip" and aborts
|
||||
# the deploy), whereas group_vars resolve through hostvars per host. spice sets it
|
||||
# to the fabric ceph_public_ip; sietch (flat) to bond_ip. It is NOT the ansible/SSH
|
||||
# connection address (that stays bond_ip / ansible_host until the WAN is retired).
|
||||
#
|
||||
# CIDR(s) cephadm restricts daemon binds to, injected as the `networks:` field of
|
||||
# the RGW + monitoring service specs. mon/mgr/OSD already bind per public_network/
|
||||
# cluster_network; this closes the remaining daemons (beast RGW, prometheus,
|
||||
# grafana, alertmanager, node-exporter, ceph-exporter) that otherwise bind 0.0.0.0.
|
||||
# EMPTY by default: no `networks:` field is emitted and the monitoring re-spec is
|
||||
# skipped, so daemons bind every interface exactly as cephadm ships them - flat
|
||||
# clusters like sietch are untouched. A cluster that must be fabric-only (spice)
|
||||
# opts in by setting this to its fabric CIDR(s) in group_vars.
|
||||
ceph_bind_networks: []
|
||||
|
||||
# --- Device-class CRUSH pool placement (per-cluster; opt-in) ---
|
||||
# Pools listed here are pinned to the named device-class CRUSH rule after RGW is
|
||||
# up. EMPTY by default so a cluster's pools stay on whatever rule they already
|
||||
# carry - re-homing a live pool triggers a full rebalance under load, so this is
|
||||
# never applied implicitly. A cluster with a dedicated fast tier (spice's ssd-osd
|
||||
# NVMe) opts in via group_vars. The rule names must match crush-rules.yml.
|
||||
ceph_rgw_ssd_crush_rule: replicated_ssd
|
||||
ceph_rgw_hdd_crush_rule: replicated_hdd
|
||||
ceph_rgw_ssd_pools: []
|
||||
ceph_rgw_hdd_pools: []
|
||||
|
||||
# --- RGW ingress: health-checked S3 VIP (haproxy + keepalived; per-cluster; opt-in) ---
|
||||
# A cephadm `ingress` service that puts a keepalived-managed floating VIP with
|
||||
# haproxy health checks in front of RGW, in TLS TERMINATION mode (haproxy
|
||||
# `mode http`: haproxy terminates the client TLS on the VIP with the reused RGW
|
||||
# self-signed cert, then re-encrypts to the RGW backends). A dead RGW node is then
|
||||
# dropped from the VIP instead of blackholing ~1/N of S3 connections. DEFAULT OFF:
|
||||
# no ingress spec is rendered or applied, so flat/dev clusters (sietch) and any
|
||||
# cluster that has not opted in are a complete no-op.
|
||||
#
|
||||
# (Terminate, not L4 passthrough, because passthrough for an RGW backend needs
|
||||
# IngressSpec.use_tcp_mode_over_rgw, a post-20.2.2 upstream field absent on the
|
||||
# pinned 20.2.2 -- source-verified; an unknown spec key makes `orch apply`
|
||||
# TypeError. Terminate is the only health-checked ingress mode available here.)
|
||||
#
|
||||
# A cluster opts in by setting ceph_rgw_ingress_enabled: true AND
|
||||
# ceph_rgw_ingress_vip in its group_vars. Three hard preconditions, all asserted
|
||||
# at apply time:
|
||||
# 1. ceph_rgw_ingress_vip must be set (the VIP the clients hit).
|
||||
# 2. ceph_bind_networks must be non-empty - otherwise beast binds
|
||||
# 0.0.0.0:<ceph_rgw_port> and the co-located haproxy VIP bind on the same
|
||||
# port collides (EADDRINUSE). Restricted beast binds a per-node IP, the VIP
|
||||
# is a distinct socket, so they coexist.
|
||||
# 3. ceph_rgw_ssl must be true - haproxy terminates with the RGW self-signed
|
||||
# cert (rgw_ssl_cert_combined_pem), which only exists when RGW ssl is on.
|
||||
ceph_rgw_ingress_enabled: false
|
||||
# Floating VIP the S3 clients hit (bare IP; the prefix is added below). Empty by
|
||||
# default; a cluster that opts in sets this to an address on its RGW-facing
|
||||
# network (e.g. spice fabric public 10.40.20.250).
|
||||
ceph_rgw_ingress_vip: ""
|
||||
# CIDR prefix of the network the VIP lives on - MUST match the prefix of the
|
||||
# public/RGW-facing network (spice fabric public = /23) so keepalived binds the
|
||||
# VIP to the right interface.
|
||||
ceph_rgw_ingress_vip_prefix: 23
|
||||
# Port clients hit on the VIP. Must match ceph_rgw_port (the S3/beast port).
|
||||
ceph_rgw_ingress_frontend_port: 443
|
||||
# haproxy stats/monitor port (load-balancer status; not the S3 data path).
|
||||
ceph_rgw_ingress_monitor_port: 1967
|
||||
# Number of haproxy + keepalived instances cephadm spreads across nodes (VRRP
|
||||
# needs >= 2 for failover; 3 tolerates one node loss with quorum to spare).
|
||||
ceph_rgw_ingress_count: 3
|
||||
|
||||
# --- Alert delivery (per-cluster; opt-in) ---
|
||||
# cephadm auto-deploys alertmanager but with NO receiver, so the ~89 Prometheus
|
||||
# alert rules (OSD down, host down, PG degraded, near-full) fire into the default
|
||||
# null route and no human is notified. Set this to one or more webhook URLs (a
|
||||
# chat incoming-webhook, or a pager like Opsgenie/PagerDuty, or a FUTO endpoint)
|
||||
# and the alertmanager service spec routes all alerts there via cephadm's
|
||||
# user_data.default_webhook_urls. Values may be op:// refs resolved before the
|
||||
# play. EMPTY by default: no receiver is wired (cephadm default is unchanged), and
|
||||
# verify.yml warns that alert delivery is not configured. TODO(operator): set the
|
||||
# destination for spice.
|
||||
ceph_alertmanager_webhook_urls: []
|
||||
|
||||
@@ -32,7 +32,7 @@
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
cephadm bootstrap \
|
||||
--mon-ip {{ bond_ip }} \
|
||||
--mon-ip {{ ceph_service_ip }} \
|
||||
--initial-dashboard-user {{ ceph_dashboard_user }} \
|
||||
--initial-dashboard-password "$(cat /tmp/.ceph-dashboard-pw)" \
|
||||
--dashboard-password-noupdate \
|
||||
@@ -75,7 +75,7 @@
|
||||
# `--dashboard-password-noupdate` means it is never re-applied afterwards. If a
|
||||
# cluster was ever bootstrapped out-of-band (or before this var existed),
|
||||
# cephadm generated a *random* admin password and no playbook run would ever
|
||||
# correct it — the vaulted password stays aspirational. These tasks reconcile
|
||||
# correct it; the vaulted password stays aspirational. These tasks reconcile
|
||||
# the running cluster to `ceph_dashboard_password` on every converge, so the
|
||||
# vault is the single source of truth for both initial set and rotation.
|
||||
- name: Write dashboard password to temp file (reconcile)
|
||||
|
||||
@@ -38,7 +38,7 @@
|
||||
ansible.builtin.shell: |
|
||||
set -euo pipefail
|
||||
HOST="{{ hostvars[item]['hostname_short'] }}"
|
||||
ADDR="{{ hostvars[item]['bond_ip'] }}"
|
||||
ADDR="{{ hostvars[item]['ceph_service_ip'] }}"
|
||||
echo "Testing SSH to $HOST ($ADDR)..."
|
||||
ceph cephadm check-host "$HOST" "$ADDR" 2>&1 || {
|
||||
echo "WARN: ceph cephadm check-host failed, testing raw SSH..."
|
||||
@@ -75,10 +75,10 @@
|
||||
grep -q '"{{ hostvars[item]['hostname_short'] }}"'; then
|
||||
echo "Host {{ hostvars[item]['hostname_short'] }} already in cluster"
|
||||
else
|
||||
echo "Adding {{ hostvars[item]['hostname_short'] }} ({{ hostvars[item]['bond_ip'] }})"
|
||||
echo "Adding {{ hostvars[item]['hostname_short'] }} ({{ hostvars[item]['ceph_service_ip'] }})"
|
||||
ceph orch host add \
|
||||
{{ hostvars[item]['hostname_short'] }} \
|
||||
{{ hostvars[item]['bond_ip'] }} 2>&1
|
||||
{{ hostvars[item]['ceph_service_ip'] }} 2>&1
|
||||
fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
|
||||
@@ -2,33 +2,122 @@
|
||||
# Phase 4.5: Ensure Ceph block.db LVM is set up on each node (sietch-shape only)
|
||||
#
|
||||
# This task is sietch-shape-specific: it assumes dual SAS-attached SSDs each
|
||||
# with partition 5 → its own VG, with 6 db-slot LVs per VG (12 total). It
|
||||
# ensures PVs, VGs, and LVs exist. Idempotent — skips if already present.
|
||||
# with partition 5 -> its own VG, with 6 db-slot LVs per VG (12 total). It
|
||||
# ensures PVs, VGs, and LVs exist. Idempotent - skips if already present.
|
||||
# Provides a recovery path if the LVM was destroyed by `mise run destroy`.
|
||||
#
|
||||
# Painbox-shape (Hetzner SX295, NVMe RAID-1 → single vg0 with all LVs created
|
||||
# by installimage post-install) skips the whole block — its host_vars don't
|
||||
# Painbox-shape (Hetzner SX295, NVMe RAID-1 -> single vg0 with all LVs created
|
||||
# by installimage post-install) skips the whole block - its host_vars don't
|
||||
# define `sas_path_prefix` so the gate below short-circuits.
|
||||
#
|
||||
# Required host_vars (sietch-shape only):
|
||||
# sas_path_prefix — by-path prefix for SSD discovery (e.g., pci-0000:02:00.0-sas-exp...)
|
||||
# ssd1_phy, ssd2_phy — PHY slot numbers for the two SSDs
|
||||
# ceph_db_vg1, ceph_db_vg2 — VG names for SSD1/SSD2 partition 5
|
||||
# sas_path_prefix - by-path prefix for SSD discovery (e.g., pci-0000:02:00.0-sas-exp...)
|
||||
# ssd1_phy, ssd2_phy - PHY slot numbers for the two SSDs
|
||||
# ceph_db_vg1, ceph_db_vg2 - VG names for SSD1/SSD2 partition 5
|
||||
#
|
||||
# VG mapping: ceph_db_vg1 on SSD1 partition 5, ceph_db_vg2 on SSD2 partition 5
|
||||
# LV naming: db-slot0..5 on VG1, db-slot6..11 on VG2 (one per HDD OSD)
|
||||
|
||||
- name: Skip in-role LVM setup (no sas_path_prefix — externally-managed LVM)
|
||||
# NVMe-RAID shape (spice): a single vg0 on the NVMe RAID-1 already carries the
|
||||
# OS LVs from installimage. Create the block.db LVs (one per HDD OSD) + one
|
||||
# ssd-osd LV here at converge, before osds.yml. This replaces painbox's fragile
|
||||
# installimage -x chroot post-install with an idempotent, observable ansible
|
||||
# step. Gated on ceph_db_lvs_per_node (spice sets it; genuinely
|
||||
# externally-managed hosts do not).
|
||||
- name: Set up LVM (NVMe-RAID single-vg0 topology)
|
||||
when:
|
||||
- sas_path_prefix is undefined
|
||||
- ceph_db_lvs_per_node is defined
|
||||
# DATA SAFETY: never let lvol SHRINK an existing LV -- reconciling a drifted
|
||||
# db-slot down would truncate a live block.db and corrupt the OSD. shrink:false
|
||||
# means lvol still creates and may grow, but leaves a larger existing LV alone
|
||||
# (the old shell skipped existing LVs entirely). opts -Wy wipes stale signatures
|
||||
# on (re)created LVs, preserving the destroy->recreate recovery-path freshness.
|
||||
module_defaults:
|
||||
community.general.lvol:
|
||||
shrink: false
|
||||
opts: "-Wy"
|
||||
block:
|
||||
- name: Assert vg0 exists (created by installimage)
|
||||
ansible.builtin.command: vgs {{ ceph_db_vg }}
|
||||
changed_when: false
|
||||
|
||||
- name: Create block.db LVs on vg0 (fixed size, one per HDD OSD)
|
||||
community.general.lvol:
|
||||
vg: "{{ ceph_db_vg }}"
|
||||
lv: "db-slot{{ item }}"
|
||||
size: "{{ ceph_db_lv_size }}"
|
||||
state: present
|
||||
loop: "{{ range(0, ceph_db_lvs_per_node | int) | list }}"
|
||||
|
||||
# ssd-osd is sized as (vg0 free - reserve); lvol has no "free minus N GiB"
|
||||
# primitive, so compute the GiB with a read-only shell and feed it to lvol.
|
||||
# Gate the compute + create on the LV being absent: once ssd-osd exists, vg0
|
||||
# free == reserve, so recomputing would yield <= 0 and the original shell's
|
||||
# FATAL guard would (correctly) refuse -- and an unforced lvol resize would
|
||||
# fail. The existence check preserves that idempotency exactly.
|
||||
- name: Check whether the ssd-osd LV already exists on vg0
|
||||
ansible.builtin.command: lvs {{ ceph_db_vg }}/{{ ceph_ssd_osd_lv }}
|
||||
register: vg0_ssd_check
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
|
||||
- name: Compute the ssd-osd size (vg0 free minus reserve, in GiB)
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
free_g=$(vgs {{ ceph_db_vg }} --noheadings --nosuffix --units g -o vg_free | tr -d ' ' | cut -d. -f1)
|
||||
echo $(( free_g - {{ ceph_ssd_osd_reserve_gib }} ))
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: vg0_ssd_size
|
||||
changed_when: false
|
||||
when: vg0_ssd_check.rc != 0
|
||||
|
||||
- name: Fail when vg0 free is at or below the reserve
|
||||
ansible.builtin.fail:
|
||||
msg: >-
|
||||
FATAL: {{ ceph_db_vg }} free minus reserve
|
||||
{{ ceph_ssd_osd_reserve_gib }}G leaves no room for {{ ceph_ssd_osd_lv }}
|
||||
when:
|
||||
- vg0_ssd_check.rc != 0
|
||||
- (vg0_ssd_size.stdout | int) <= 0
|
||||
|
||||
- name: Create the ssd-osd LV from vg0 remainder minus reserve
|
||||
community.general.lvol:
|
||||
vg: "{{ ceph_db_vg }}"
|
||||
lv: "{{ ceph_ssd_osd_lv }}"
|
||||
size: "{{ vg0_ssd_size.stdout | int }}G"
|
||||
state: present
|
||||
when: vg0_ssd_check.rc != 0
|
||||
|
||||
- name: Verify vg0 LVM layout
|
||||
ansible.builtin.command: lvs {{ ceph_db_vg }} -o lv_name,lv_size --noheadings
|
||||
register: vg0_lvm_state
|
||||
changed_when: false
|
||||
|
||||
- name: Show vg0 LVM layout
|
||||
ansible.builtin.debug:
|
||||
msg: "{{ vg0_lvm_state.stdout_lines }}"
|
||||
|
||||
# Genuinely externally-managed LVM (neither sietch dual-SSD nor spice NVMe-RAID).
|
||||
- name: Skip in-role LVM setup (externally-managed LVM)
|
||||
ansible.builtin.debug:
|
||||
msg: >-
|
||||
LVM is externally managed on this host (e.g., Hetzner installimage
|
||||
post-install creates vg0 + db-slots on a RAID-1 NVMe). Skipping the
|
||||
sietch-shape dual-SSD-VG setup. If LVM is missing on this host,
|
||||
LVM is externally managed on this host (no sas_path_prefix and no
|
||||
ceph_db_lvs_per_node). Skipping in-role LVM setup. If LVM is missing,
|
||||
reprovision via the cluster's installimage flow.
|
||||
when: sas_path_prefix is undefined
|
||||
when:
|
||||
- sas_path_prefix is undefined
|
||||
- ceph_db_lvs_per_node is undefined
|
||||
|
||||
- name: Set up LVM (sietch-shape dual-SSD-VG topology)
|
||||
when: sas_path_prefix is defined
|
||||
# DATA SAFETY: shrink:false so lvol never truncates a live block.db LV on drift
|
||||
# (it may still create/grow); opts -Wy wipes stale signatures on (re)created LVs.
|
||||
module_defaults:
|
||||
community.general.lvol:
|
||||
shrink: false
|
||||
opts: "-Wy"
|
||||
block:
|
||||
- name: Resolve SSD1 device path
|
||||
ansible.builtin.command: >
|
||||
@@ -54,51 +143,64 @@
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
|
||||
- name: Create PV and VG on SSD1 partition 5
|
||||
ansible.builtin.shell: |
|
||||
PART="{{ ssd1_dev.stdout }}5"
|
||||
wipefs -af "$PART"
|
||||
pvcreate -f "$PART" && vgcreate {{ ceph_db_vg1 }} "$PART"
|
||||
# lvg runs "pvcreate -f" but not "--yes", so it will not clear a stale
|
||||
# foreign (ceph/ext) signature on a reused partition. Preserve the original
|
||||
# wipefs -af, gated exactly as before on the VG being absent, so we never
|
||||
# wipe a live PV. lvg then creates the PV+VG idempotently.
|
||||
- name: Wipe stale signatures on SSD1 partition 5
|
||||
ansible.builtin.command: wipefs -af {{ ssd1_dev.stdout }}5
|
||||
when: vg1_check.rc != 0
|
||||
changed_when: true
|
||||
|
||||
- name: Create PV and VG on SSD2 partition 5
|
||||
ansible.builtin.shell: |
|
||||
PART="{{ ssd2_dev.stdout }}5"
|
||||
wipefs -af "$PART"
|
||||
pvcreate -f "$PART" && vgcreate {{ ceph_db_vg2 }} "$PART"
|
||||
- name: Create PV and VG on SSD1 partition 5
|
||||
community.general.lvg:
|
||||
vg: "{{ ceph_db_vg1 }}"
|
||||
pvs: "{{ ssd1_dev.stdout }}5"
|
||||
state: present
|
||||
|
||||
- name: Wipe stale signatures on SSD2 partition 5
|
||||
ansible.builtin.command: wipefs -af {{ ssd2_dev.stdout }}5
|
||||
when: vg2_check.rc != 0
|
||||
changed_when: true
|
||||
|
||||
- name: Create block.db LVs on VG1 (slots 0-4 at 240G, slot 5 gets remainder)
|
||||
ansible.builtin.shell: |
|
||||
if lvs {{ ceph_db_vg1 }}/db-slot{{ item }} 2>/dev/null; then
|
||||
echo "db-slot{{ item }} already exists"
|
||||
else
|
||||
{% if item < 5 %}
|
||||
lvcreate --yes -Wy -L {{ ceph_db_lv_size }} -n db-slot{{ item }} {{ ceph_db_vg1 }}
|
||||
{% else %}
|
||||
lvcreate --yes -Wy -l 100%FREE -n db-slot{{ item }} {{ ceph_db_vg1 }}
|
||||
{% endif %}
|
||||
fi
|
||||
loop: [0, 1, 2, 3, 4, 5]
|
||||
register: vg1_lv_results
|
||||
changed_when: "'already exists' not in (vg1_lv_results.stdout | default(''))"
|
||||
- name: Create PV and VG on SSD2 partition 5
|
||||
community.general.lvg:
|
||||
vg: "{{ ceph_db_vg2 }}"
|
||||
pvs: "{{ ssd2_dev.stdout }}5"
|
||||
state: present
|
||||
|
||||
- name: Create block.db LVs on VG2 (slots 6-10 at 240G, slot 11 gets remainder)
|
||||
ansible.builtin.shell: |
|
||||
if lvs {{ ceph_db_vg2 }}/db-slot{{ item }} 2>/dev/null; then
|
||||
echo "db-slot{{ item }} already exists"
|
||||
else
|
||||
{% if item < 11 %}
|
||||
lvcreate --yes -Wy -L {{ ceph_db_lv_size }} -n db-slot{{ item }} {{ ceph_db_vg2 }}
|
||||
{% else %}
|
||||
lvcreate --yes -Wy -l 100%FREE -n db-slot{{ item }} {{ ceph_db_vg2 }}
|
||||
{% endif %}
|
||||
fi
|
||||
loop: [6, 7, 8, 9, 10, 11]
|
||||
register: vg2_lv_results
|
||||
changed_when: "'already exists' not in (vg2_lv_results.stdout | default(''))"
|
||||
# Fixed slots first, then the 100%FREE remainder, so the remainder LV only
|
||||
# claims the space left after the fixed LVs (matches the sequential order of
|
||||
# the original single-loop shell).
|
||||
- name: Create fixed-size block.db LVs on VG1 (slots 0-4 at 240G)
|
||||
community.general.lvol:
|
||||
vg: "{{ ceph_db_vg1 }}"
|
||||
lv: "db-slot{{ item }}"
|
||||
size: "{{ ceph_db_lv_size }}"
|
||||
state: present
|
||||
loop: [0, 1, 2, 3, 4]
|
||||
|
||||
- name: Create remainder block.db LV on VG1 (slot 5, 100%FREE)
|
||||
community.general.lvol:
|
||||
vg: "{{ ceph_db_vg1 }}"
|
||||
lv: db-slot5
|
||||
size: 100%FREE
|
||||
state: present
|
||||
|
||||
- name: Create fixed-size block.db LVs on VG2 (slots 6-10 at 240G)
|
||||
community.general.lvol:
|
||||
vg: "{{ ceph_db_vg2 }}"
|
||||
lv: "db-slot{{ item }}"
|
||||
size: "{{ ceph_db_lv_size }}"
|
||||
state: present
|
||||
loop: [6, 7, 8, 9, 10]
|
||||
|
||||
- name: Create remainder block.db LV on VG2 (slot 11, 100%FREE)
|
||||
community.general.lvol:
|
||||
vg: "{{ ceph_db_vg2 }}"
|
||||
lv: db-slot11
|
||||
size: 100%FREE
|
||||
state: present
|
||||
|
||||
# --- SSD OSD data LVs (partition 6) ---
|
||||
# part6 of each SSD is the SSD-class OSD data device. cephadm cannot deploy
|
||||
@@ -120,46 +222,50 @@
|
||||
failed_when: false
|
||||
when: ceph_ssd_vg2 is defined
|
||||
|
||||
- name: Create PV and VG on SSD1 partition 6 (SSD OSD)
|
||||
ansible.builtin.shell: |
|
||||
PART="{{ ssd1_dev.stdout }}6"
|
||||
wipefs -af "$PART"
|
||||
pvcreate -f "$PART" && vgcreate {{ ceph_ssd_vg1 }} "$PART"
|
||||
# Same wipefs-then-lvg safety as partition 5, gated on the SSD-OSD VG being
|
||||
# requested (ceph_ssd_vg1/2 defined) and absent.
|
||||
- name: Wipe stale signatures on SSD1 partition 6 (SSD OSD)
|
||||
ansible.builtin.command: wipefs -af {{ ssd1_dev.stdout }}6
|
||||
when:
|
||||
- ceph_ssd_vg1 is defined
|
||||
- ssd_vg1_check.rc | default(1) != 0
|
||||
changed_when: true
|
||||
|
||||
- name: Create PV and VG on SSD2 partition 6 (SSD OSD)
|
||||
ansible.builtin.shell: |
|
||||
PART="{{ ssd2_dev.stdout }}6"
|
||||
wipefs -af "$PART"
|
||||
pvcreate -f "$PART" && vgcreate {{ ceph_ssd_vg2 }} "$PART"
|
||||
- name: Create PV and VG on SSD1 partition 6 (SSD OSD)
|
||||
community.general.lvg:
|
||||
vg: "{{ ceph_ssd_vg1 }}"
|
||||
pvs: "{{ ssd1_dev.stdout }}6"
|
||||
state: present
|
||||
when: ceph_ssd_vg1 is defined
|
||||
|
||||
- name: Wipe stale signatures on SSD2 partition 6 (SSD OSD)
|
||||
ansible.builtin.command: wipefs -af {{ ssd2_dev.stdout }}6
|
||||
when:
|
||||
- ceph_ssd_vg2 is defined
|
||||
- ssd_vg2_check.rc | default(1) != 0
|
||||
changed_when: true
|
||||
|
||||
- name: Create PV and VG on SSD2 partition 6 (SSD OSD)
|
||||
community.general.lvg:
|
||||
vg: "{{ ceph_ssd_vg2 }}"
|
||||
pvs: "{{ ssd2_dev.stdout }}6"
|
||||
state: present
|
||||
when: ceph_ssd_vg2 is defined
|
||||
|
||||
- name: Create SSD OSD LV on VG1 (100%FREE)
|
||||
ansible.builtin.shell: |
|
||||
if lvs {{ ceph_ssd_vg1 }}/osd-ssd 2>/dev/null; then
|
||||
echo "osd-ssd already exists"
|
||||
else
|
||||
lvcreate --yes -Wy -l 100%FREE -n osd-ssd {{ ceph_ssd_vg1 }}
|
||||
fi
|
||||
register: ssd_lv1_result
|
||||
changed_when: "'already exists' not in (ssd_lv1_result.stdout | default(''))"
|
||||
community.general.lvol:
|
||||
vg: "{{ ceph_ssd_vg1 }}"
|
||||
lv: osd-ssd
|
||||
size: 100%FREE
|
||||
state: present
|
||||
when: ceph_ssd_vg1 is defined
|
||||
|
||||
- name: Create SSD OSD LV on VG2 (100%FREE)
|
||||
ansible.builtin.shell: |
|
||||
if lvs {{ ceph_ssd_vg2 }}/osd-ssd 2>/dev/null; then
|
||||
echo "osd-ssd already exists"
|
||||
else
|
||||
lvcreate --yes -Wy -l 100%FREE -n osd-ssd {{ ceph_ssd_vg2 }}
|
||||
fi
|
||||
register: ssd_lv2_result
|
||||
changed_when: "'already exists' not in (ssd_lv2_result.stdout | default(''))"
|
||||
community.general.lvol:
|
||||
vg: "{{ ceph_ssd_vg2 }}"
|
||||
lv: osd-ssd
|
||||
size: 100%FREE
|
||||
state: present
|
||||
when: ceph_ssd_vg2 is defined
|
||||
|
||||
- name: Verify LVM setup
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
---
|
||||
# Phase 5.8: Monitoring stack — dashboard integration
|
||||
# Phase 5.8: Monitoring stack - dashboard integration
|
||||
#
|
||||
# cephadm auto-deploys the full monitoring stack (node-exporter,
|
||||
# ceph-exporter, prometheus, alertmanager, grafana) during bootstrap.
|
||||
@@ -62,26 +62,129 @@
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
# --- Step 2.5: Re-spec the monitoring stack (fabric bind + alert delivery) ---
|
||||
# cephadm auto-deployed the stack binding all interfaces and with NO alertmanager
|
||||
# receiver. Re-apply the spec to (a) restrict the bind to ceph_bind_networks
|
||||
# (fabric-only on spice, so ceph never listens on the WAN) and (b) route alerts to
|
||||
# ceph_alertmanager_webhook_urls. Placement matches cephadm's defaults, so no
|
||||
# daemon relocates. Each piece is independently gated in the template, and the
|
||||
# whole re-spec is SKIPPED when BOTH are empty (flat clusters with no alerting, like
|
||||
# sietch) - the auto-deployed stack is then left exactly as cephadm shipped it.
|
||||
- name: Render monitoring service spec (fabric bind + alert receiver)
|
||||
ansible.builtin.template:
|
||||
src: monitoring-spec.yaml.j2
|
||||
dest: /etc/ceph/monitoring-spec.yaml
|
||||
owner: root
|
||||
group: root
|
||||
mode: '0644'
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- (ceph_bind_networks | default([]) | length) > 0
|
||||
or (ceph_alertmanager_webhook_urls | default([]) | length) > 0
|
||||
|
||||
- name: Apply monitoring service spec
|
||||
ansible.builtin.command: ceph orch apply -i /etc/ceph/monitoring-spec.yaml
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- (ceph_bind_networks | default([]) | length) > 0
|
||||
or (ceph_alertmanager_webhook_urls | default([]) | length) > 0
|
||||
changed_when: false
|
||||
|
||||
# --- Step 2.6: Pin the mgr dashboard + prometheus module bind to the fabric ---
|
||||
#
|
||||
# The dashboard (8443) and the prometheus mgr module (9283) run INSIDE the active
|
||||
# ceph-mgr, not as standalone cephadm daemons, so the monitoring re-spec above
|
||||
# does not reach them -- they bind :: (every interface, WAN included) and are
|
||||
# closed only by nftables. Pin each mgr instance's own server_addr to that
|
||||
# instance's fabric IP (ceph_service_ip) so whichever mgr is active listens on
|
||||
# the fabric, never the WAN.
|
||||
#
|
||||
# PER-INSTANCE, not global: the active mgr fails over across the mgr hosts and
|
||||
# each host has a DIFFERENT fabric IP (ceph_service_ip = 10.40.20.<index>). A
|
||||
# single global mgr/dashboard/server_addr would bind a foreign IP after failover
|
||||
# and the dashboard could not start. Ceph resolves the localized key
|
||||
# mgr/<module>/<mgr_id>/server_addr ahead of the global one (the module reads it
|
||||
# via get_localized_module_option), and each mgr reads ITS OWN value when its
|
||||
# module starts serving (on mgr start / failover), so per-instance binding is
|
||||
# inherently failover-safe. The mgr_id is the cephadm daemon name (<host>.<rand>,
|
||||
# e.g. spice-ceph-adelia.abcdef), NOT hostname_short -- it is discovered at
|
||||
# runtime from `ceph orch ps`, and the daemon's host is mapped back to that
|
||||
# host's ceph_service_ip.
|
||||
#
|
||||
# TIMING: the new bind takes effect the next time each mgr (re)starts its module
|
||||
# (mgr restart, failover, or upgrade); nftables keeps the WAN closed until then.
|
||||
# We deliberately do NOT force a live module reload (disable/enable) -- that would
|
||||
# drop the dashboard mid-converge for no durability gain, and the per-instance
|
||||
# config can never leave the dashboard unbindable the way a global one could.
|
||||
#
|
||||
# Opt-in: only when ceph_bind_networks is non-empty. sietch (empty) is a no-op
|
||||
# and keeps cephadm's bind-all default.
|
||||
- name: Pin mgr dashboard + prometheus module bind to the fabric (per mgr instance)
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
python3 - <<'PYEOF'
|
||||
import json, subprocess, sys
|
||||
|
||||
fabric_ip = {
|
||||
{% for h in groups['ceph_nodes'] %}
|
||||
"{{ hostvars[h]['hostname_short'] }}": "{{ hostvars[h]['ceph_service_ip'] }}",
|
||||
{% endfor %}
|
||||
}
|
||||
|
||||
mgrs = json.loads(subprocess.check_output(
|
||||
["ceph", "orch", "ps", "--daemon-type", "mgr", "--format", "json"]))
|
||||
cfg = json.loads(subprocess.check_output(
|
||||
["ceph", "config", "dump", "--format", "json"]))
|
||||
current = {(c["section"], c["name"]): c["value"] for c in cfg}
|
||||
|
||||
changed = False
|
||||
for d in mgrs:
|
||||
mgr_id = d["daemon_id"]
|
||||
host = d["hostname"]
|
||||
ip = fabric_ip.get(host)
|
||||
if not ip:
|
||||
sys.stderr.write(
|
||||
"WARNING: no fabric IP for mgr host %s, skipping\n" % host)
|
||||
continue
|
||||
for mod in ("dashboard", "prometheus"):
|
||||
key = "mgr/%s/%s/server_addr" % (mod, mgr_id)
|
||||
if current.get(("mgr", key)) != ip:
|
||||
subprocess.check_call(["ceph", "config", "set", "mgr", key, ip])
|
||||
print("CHANGED %s -> %s" % (key, ip))
|
||||
changed = True
|
||||
print("CHANGED" if changed else "ok")
|
||||
PYEOF
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- (ceph_bind_networks | default([]) | length) > 0
|
||||
register: mgr_module_bind
|
||||
changed_when: "'CHANGED' in (mgr_module_bind.stdout | default(''))"
|
||||
|
||||
# --- Step 3: Configure dashboard integrations ---
|
||||
# The dashboard runs on the active mgr and reaches these services on the fabric;
|
||||
# advertise ceph_service_ip (fabric on spice), not bond_ip (the 1G WAN), so the
|
||||
# URLs stay inside ceph_firewall_trusted_networks after harden.
|
||||
|
||||
- name: Configure dashboard Prometheus URL
|
||||
ansible.builtin.command: >
|
||||
ceph dashboard set-prometheus-api-host
|
||||
http://{{ hostvars[groups['ceph_bootstrap'][0]]['bond_ip'] }}:{{ ceph_prometheus_port }}
|
||||
http://{{ hostvars[groups['ceph_bootstrap'][0]]['ceph_service_ip'] }}:{{ ceph_prometheus_port }}
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Configure dashboard Alertmanager URL
|
||||
ansible.builtin.command: >
|
||||
ceph dashboard set-alertmanager-api-host
|
||||
http://{{ hostvars[groups['ceph_bootstrap'][0]]['bond_ip'] }}:{{ ceph_alertmanager_port }}
|
||||
http://{{ hostvars[groups['ceph_bootstrap'][0]]['ceph_service_ip'] }}:{{ ceph_alertmanager_port }}
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Configure dashboard Grafana URL
|
||||
ansible.builtin.command: >
|
||||
ceph dashboard set-grafana-api-url
|
||||
https://{{ hostvars[groups['ceph_bootstrap'][0]]['bond_ip'] }}:{{ ceph_grafana_port }}
|
||||
https://{{ hostvars[groups['ceph_bootstrap'][0]]['ceph_service_ip'] }}:{{ ceph_grafana_port }}
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
@@ -159,9 +262,9 @@
|
||||
ceph orch ls --service-type ceph-exporter
|
||||
echo ""
|
||||
echo "=== Endpoints ==="
|
||||
echo "Prometheus: http://{{ bond_ip }}:{{ ceph_prometheus_port }}"
|
||||
echo "Grafana: https://{{ bond_ip }}:{{ ceph_grafana_port }}"
|
||||
echo "Alertmanager: http://{{ bond_ip }}:{{ ceph_alertmanager_port }}"
|
||||
echo "Prometheus: http://{{ ceph_service_ip }}:{{ ceph_prometheus_port }}"
|
||||
echo "Grafana: https://{{ ceph_service_ip }}:{{ ceph_grafana_port }}"
|
||||
echo "Alertmanager: http://{{ ceph_service_ip }}:{{ ceph_alertmanager_port }}"
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: monitoring_status
|
||||
|
||||
@@ -52,17 +52,17 @@
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
EXPECTED={{
|
||||
groups['ceph_nodes']
|
||||
(groups['ceph_nodes'] | intersect(ansible_play_hosts))
|
||||
| map('extract', hostvars)
|
||||
| map(attribute='ceph_hdd_osds', default=[])
|
||||
| map('length') | sum
|
||||
+
|
||||
groups['ceph_nodes']
|
||||
(groups['ceph_nodes'] | intersect(ansible_play_hosts))
|
||||
| map('extract', hostvars)
|
||||
| map(attribute='ceph_ssd_osds', default=[])
|
||||
| map('length') | sum
|
||||
}}
|
||||
echo "Expecting $EXPECTED OSDs total across the cluster"
|
||||
echo "Expecting $EXPECTED OSDs across the nodes in this run"
|
||||
for i in $(seq 1 180); do
|
||||
ACTUAL=$(ceph osd stat --format json 2>/dev/null \
|
||||
| python3 -c "import sys, json; print(json.load(sys.stdin).get('num_osds', 0))" \
|
||||
|
||||
@@ -6,7 +6,8 @@
|
||||
# 2. Else if <=2 active hosts: 1 MON (bootstrap only, avoids 2-MON fragility)
|
||||
# 3. Else: MON on all active hosts
|
||||
#
|
||||
# MGR always deploys on all active hosts (standbys are harmless).
|
||||
# MGR deploys a capped count (ceph_mgr_count, default 3: 1 active + standbys),
|
||||
# not one per host -- cephadm schedules them and caps at the host count.
|
||||
|
||||
- name: Determine active cluster hosts
|
||||
ansible.builtin.shell: |
|
||||
@@ -45,10 +46,10 @@
|
||||
MON placement: {{ mon_hosts }}
|
||||
({{ 'explicit ceph_mon group'
|
||||
if groups['ceph_mon'] | default([]) | length > 0
|
||||
else ('single MON — avoids 2-mon quorum fragility'
|
||||
else ('single MON - avoids 2-mon quorum fragility'
|
||||
if active_hosts.stdout.split(',') | length <= 2
|
||||
else 'MON on all hosts') }}).
|
||||
MGR placement: {{ active_hosts.stdout }} (all hosts).
|
||||
MGR placement: count:{{ ceph_mgr_count }}.
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- mon_hosts is defined
|
||||
@@ -61,9 +62,9 @@
|
||||
- mon_hosts | default('') | length > 0
|
||||
changed_when: false
|
||||
|
||||
- name: Deploy MGR on all active hosts
|
||||
- name: Deploy MGR at a capped count (not one per host)
|
||||
ansible.builtin.command: >
|
||||
ceph orch apply mgr --placement="{{ active_hosts.stdout }}"
|
||||
ceph orch apply mgr --placement="count:{{ ceph_mgr_count }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- active_hosts.stdout | default('') | length > 0
|
||||
|
||||
@@ -46,6 +46,45 @@
|
||||
- create_data_pool is changed
|
||||
changed_when: true
|
||||
|
||||
- name: Mark the EC data pool as bulk (autoscaler sizing floor)
|
||||
# bulk=true tells the pg_autoscaler to target a pg_num sized for the pool's
|
||||
# expected full capacity, so it will not shrink the pre-seeded pg_num
|
||||
# (ceph_pg_init_data) back toward 1 during the first fill. Idempotent: only sets
|
||||
# when not already true.
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
[ "$(ceph osd pool get {{ ceph_rgw_data_pool }} bulk | awk '{print $2}')" = "true" ] || { ceph osd pool set {{ ceph_rgw_data_pool }} bulk true; echo CHANGED; }
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- (create_data_pool is changed) or (ceph_rgw_data_pool in pool_list.stdout)
|
||||
register: bulk_flag_result
|
||||
changed_when: "'CHANGED' in (bulk_flag_result.stdout | default(''))"
|
||||
|
||||
- name: Pin EC data pool min_size to k+1 ({{ ceph_rgw_ec_min_size }})
|
||||
# The EC pool otherwise inherits Ceph's default min_size = k+1 implicitly. Pin
|
||||
# it EXPLICITLY so the write-availability floor is documented and cannot drift:
|
||||
# k+1 keeps a PG writable while at least k+1 shards are up, so writes tolerate
|
||||
# up to m-1 host losses while reads still tolerate m. NEVER set below k+1 --
|
||||
# allowing writes at k shards risks data loss if another shard is lost during
|
||||
# recovery. Idempotent get-then-set; gated on the pool existing (fresh run:
|
||||
# create_data_pool changed; re-run: pool already in pool_list).
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
cur=$(ceph osd pool get {{ ceph_rgw_data_pool }} min_size | awk '{print $2}')
|
||||
if [ "$cur" != "{{ ceph_rgw_ec_min_size }}" ]; then
|
||||
ceph osd pool set {{ ceph_rgw_data_pool }} min_size {{ ceph_rgw_ec_min_size }}
|
||||
echo CHANGED
|
||||
fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- (create_data_pool is changed) or (ceph_rgw_data_pool in pool_list.stdout)
|
||||
register: ec_min_size_result
|
||||
changed_when: "'CHANGED' in (ec_min_size_result.stdout | default(''))"
|
||||
|
||||
# --- Step 4: Replicated index pool ---
|
||||
|
||||
- name: Create replicated index pool {{ ceph_rgw_index_pool }}
|
||||
@@ -117,6 +156,18 @@
|
||||
# spikes and throughput drops. Pre-sizing avoids this penalty during
|
||||
# the first fill. Idempotent: only increases pg_num, never decreases.
|
||||
|
||||
# The top-of-file pool_list snapshot predates the data/index/non-EC pools
|
||||
# created just above, so on a fresh run 1 it does NOT contain them and every
|
||||
# "in pool_list.stdout" gate below would skip - deferring all pre-sizing to a
|
||||
# second converge. Re-query now that the explicit pools exist so run 1 sizes
|
||||
# them. (The lazily-created .rgw.* system pools are handled later, after the
|
||||
# RGW daemon is up.)
|
||||
- name: Re-query pools after explicit RGW pools are created
|
||||
ansible.builtin.command: ceph osd pool ls
|
||||
register: pool_list
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Set initial PG count on data pool
|
||||
ansible.builtin.command: >
|
||||
ceph osd pool set {{ ceph_rgw_data_pool }}
|
||||
@@ -156,26 +207,10 @@
|
||||
- pg_extra_result.rc != 0
|
||||
- "'is not >= current' not in pg_extra_result.stderr | default('')"
|
||||
|
||||
- name: Set initial PG count on RGW metadata pools
|
||||
ansible.builtin.command: >
|
||||
ceph osd pool set {{ item }} pg_num {{ ceph_pg_init_meta }}
|
||||
loop:
|
||||
- .rgw.root
|
||||
- "{{ ceph_rgw_zone }}.rgw.log"
|
||||
- "{{ ceph_rgw_zone }}.rgw.control"
|
||||
- "{{ ceph_rgw_zone }}.rgw.meta"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
# Skip pools that haven't materialized yet — RGW creates them
|
||||
# lazily on first daemon startup. Mirrors the per-pool gate used
|
||||
# by the data/index/extra pool tasks above.
|
||||
- item in pool_list.stdout
|
||||
register: pg_meta_result
|
||||
changed_when: "'set' in pg_meta_result.stdout | default('')"
|
||||
failed_when:
|
||||
- pg_meta_result.rc is defined
|
||||
- pg_meta_result.rc != 0
|
||||
- "'is not >= current' not in pg_meta_result.stderr | default('')"
|
||||
# NB: the .rgw.* system-pool PG-count and size/min_size pins are NOT here - those
|
||||
# pools are created lazily by the realm/zone setup and the RGW daemon, so they do
|
||||
# not exist yet at this point in the play. They are pinned in Step 13.6 below,
|
||||
# after the RGW readiness wait, against a freshly re-queried pool list.
|
||||
|
||||
- name: Wait for PG peering to complete
|
||||
ansible.builtin.shell: |
|
||||
@@ -345,6 +380,7 @@
|
||||
{% for h in groups['ceph_nodes'] %}
|
||||
'{{ hostvars[h]['hostname_short'] }}',
|
||||
'{{ hostvars[h]['bond_ip'] }}',
|
||||
'{{ hostvars[h]['ceph_service_ip'] }}',
|
||||
{% endfor %}
|
||||
'{{ ceph_rgw_dns_name }}',
|
||||
])
|
||||
@@ -412,6 +448,28 @@
|
||||
# - a wildcard under the same name for virtual-hosted buckets
|
||||
# - per-node FQDNs so direct-host addressing also validates
|
||||
# - per-node bond IPs so IP-based S3 clients also validate
|
||||
# - the RGW ingress VIP (only when ceph_rgw_ingress_enabled) so a client hitting
|
||||
# the health-checked VIP, where haproxy terminates TLS with this cert, validates
|
||||
|
||||
# Build the -subj and SAN as single-line strings up front. Doing this inline in
|
||||
# the openssl command forced a backslash-newline INSIDE the quoted args, which
|
||||
# folds the next line's indentation into the value (e.g. C="DE ") and openssl
|
||||
# aborts on the over-length field. Precomputing here keeps the shell one flat line
|
||||
# per arg. The {%- ... %} whitespace-control strips the fold-inserted spaces so the
|
||||
# SAN concatenates with no embedded whitespace.
|
||||
- name: Build RGW self-signed cert subject and SAN
|
||||
ansible.builtin.set_fact:
|
||||
_rgw_cert_subj: >-
|
||||
/C={{ ceph_rgw_ssl_cert_subject_c }}/ST={{ ceph_rgw_ssl_cert_subject_st }}/L={{ ceph_rgw_ssl_cert_subject_l }}/O={{ ceph_rgw_ssl_cert_subject_o }}/CN={{ ceph_rgw_dns_name }}/emailAddress={{ ceph_rgw_ssl_cert_email }}
|
||||
_rgw_cert_san: >-
|
||||
subjectAltName=DNS:{{ ceph_rgw_dns_name }},DNS:*.{{ ceph_rgw_dns_name }}
|
||||
{%- for h in groups['ceph_nodes'] %},DNS:{{ hostvars[h]['hostname_short'] }}.{{ cluster_domain }}{% endfor %}
|
||||
{%- for h in groups['ceph_nodes'] %},IP:{{ hostvars[h]['bond_ip'] }}{% endfor %}
|
||||
{%- for h in groups['ceph_nodes'] %},IP:{{ hostvars[h]['ceph_service_ip'] }}{% endfor %}
|
||||
{%- if ceph_rgw_ingress_enabled | default(false) | bool and (ceph_rgw_ingress_vip | default('')) | length > 0 %},IP:{{ ceph_rgw_ingress_vip }}{% endif %}
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_rgw_ssl | default(false) | bool
|
||||
|
||||
- name: Generate RGW self-signed cert and key (10 year validity)
|
||||
ansible.builtin.shell: |
|
||||
@@ -421,17 +479,8 @@
|
||||
-keyout /etc/ceph/rgw-ssl.key \
|
||||
-out /etc/ceph/rgw-ssl.crt \
|
||||
-days {{ ceph_rgw_ssl_cert_days }} \
|
||||
-subj "/C={{ ceph_rgw_ssl_cert_subject_c }}\
|
||||
/ST={{ ceph_rgw_ssl_cert_subject_st }}\
|
||||
/L={{ ceph_rgw_ssl_cert_subject_l }}\
|
||||
/O={{ ceph_rgw_ssl_cert_subject_o }}\
|
||||
/CN={{ ceph_rgw_dns_name }}\
|
||||
/emailAddress={{ ceph_rgw_ssl_cert_email }}" \
|
||||
-addext "subjectAltName=\
|
||||
DNS:{{ ceph_rgw_dns_name }},\
|
||||
DNS:*.{{ ceph_rgw_dns_name }}\
|
||||
{% for h in groups['ceph_nodes'] %},DNS:{{ hostvars[h]['hostname_short'] }}.{{ cluster_domain }}{% endfor %}\
|
||||
{% for h in groups['ceph_nodes'] %},IP:{{ hostvars[h]['bond_ip'] }}{% endfor %}"
|
||||
-subj {{ _rgw_cert_subj | quote }} \
|
||||
-addext {{ _rgw_cert_san | quote }}
|
||||
args:
|
||||
executable: /bin/bash
|
||||
creates: /etc/ceph/rgw-ssl.crt
|
||||
@@ -519,12 +568,192 @@
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: rgw_running_count
|
||||
until: "(rgw_running_count.stdout | int) >= (groups['ceph_nodes'] | length)"
|
||||
# Wait for RGW on the nodes in THIS run (respects --limit), not all ceph_nodes,
|
||||
# so a subset deploy does not block forever waiting for daemons on hosts it never
|
||||
# joined. The rgw-spec placement stays pinned to all ceph_nodes (it must not
|
||||
# shrink on a --limit re-run); only the readiness gate is play-scoped.
|
||||
until: "(rgw_running_count.stdout | int) >= (groups['ceph_nodes'] | intersect(ansible_play_hosts) | length)"
|
||||
retries: 30
|
||||
delay: 10
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
# --- Step 13.7: RGW ingress (haproxy + keepalived) -- health-checked S3 VIP ---
|
||||
#
|
||||
# Opt-in floating VIP in front of RGW in TLS TERMINATION mode: haproxy `mode http`
|
||||
# terminates the client's TLS on the VIP with the RGW self-signed cert (reused,
|
||||
# combined cert+key -- see template), then re-encrypts to each RGW/beast backend
|
||||
# (default-server ssl / verify none). The keepalived VIP + haproxy per-backend
|
||||
# health check drop a dead RGW node from rotation instead of blackholing ~1/N of
|
||||
# S3 connections.
|
||||
#
|
||||
# Runs HERE, after the RGW readiness wait above, on purpose: the ingress spec's
|
||||
# backend_service references the RGW service, which must already exist and be
|
||||
# running before cephadm can wire the ingress to it. It also needs
|
||||
# rgw_ssl_cert_combined_pem, the fact built in Step 11.6.
|
||||
#
|
||||
# WHY TERMINATE, not passthrough: L4 passthrough (`mode tcp`) for an RGW backend
|
||||
# requires IngressSpec.use_tcp_mode_over_rgw, a post-20.2.2 upstream field absent
|
||||
# on the pinned 20.2.2 (source-verified: an unknown spec key makes `ceph orch
|
||||
# apply` TypeError). Terminate is the only health-checked ingress mode available
|
||||
# here. haproxy forwards the Host header in mode http, so virtual-hosted buckets
|
||||
# keep working against haproxy's wildcard cert.
|
||||
#
|
||||
# NO 443 COLLISION: haproxy terminates on VIP:{{ ceph_rgw_ingress_frontend_port }},
|
||||
# beast binds each node's fabric IP:{{ ceph_rgw_port }} -- distinct sockets. This
|
||||
# holds ONLY when the RGW spec restricts beast to a specific IP (ceph_bind_networks
|
||||
# non-empty); otherwise beast binds 0.0.0.0 and the co-located VIP bind collides.
|
||||
# Asserted below.
|
||||
#
|
||||
# APPLY ORDERING (cross-tool): the DNS cutover that points s3.<domain> at the VIP
|
||||
# is done separately in TF and MUST NOT precede this ingress being live, or S3
|
||||
# resolves to a VIP that nothing answers.
|
||||
#
|
||||
# Idempotent + declarative: `ceph orch apply -i` reconciles to the spec.
|
||||
|
||||
- name: Assert RGW ingress preconditions (VIP + restricted beast bind + RGW ssl)
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- (ceph_rgw_ingress_vip | default('')) | length > 0
|
||||
- ceph_bind_networks | default([]) | length > 0
|
||||
- ceph_rgw_ssl | default(false) | bool
|
||||
fail_msg: >-
|
||||
ceph_rgw_ingress_enabled is true but a precondition is unmet.
|
||||
Require: ceph_rgw_ingress_vip set (e.g. 10.40.20.250); ceph_bind_networks
|
||||
non-empty so beast binds a specific per-node IP (else beast binds
|
||||
0.0.0.0:{{ ceph_rgw_port }} and the haproxy VIP bind on the same port fails
|
||||
with EADDRINUSE); and ceph_rgw_ssl true so the self-signed cert exists for
|
||||
haproxy to terminate with (rgw_ssl_cert_combined_pem).
|
||||
quiet: true
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_rgw_ingress_enabled | default(false) | bool
|
||||
|
||||
- name: Render RGW ingress service spec
|
||||
ansible.builtin.template:
|
||||
src: rgw-ingress-spec.yaml.j2
|
||||
dest: /etc/ceph/rgw-ingress-spec.yaml
|
||||
owner: root
|
||||
group: root
|
||||
mode: '0600'
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_rgw_ingress_enabled | default(false) | bool
|
||||
- (ceph_rgw_ingress_vip | default('')) | length > 0
|
||||
no_log: true # spec embeds the RGW private key (combined PEM)
|
||||
|
||||
- name: Apply RGW ingress service spec
|
||||
ansible.builtin.command: ceph orch apply -i /etc/ceph/rgw-ingress-spec.yaml
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_rgw_ingress_enabled | default(false) | bool
|
||||
- (ceph_rgw_ingress_vip | default('')) | length > 0
|
||||
changed_when: false
|
||||
|
||||
# --- Step 13.6: Pin pools to device-class CRUSH rules ---
|
||||
#
|
||||
# crush-rules.yml creates replicated_ssd/replicated_hdd but binds no pool, so
|
||||
# every replicated pool falls to the default class-agnostic rule and the
|
||||
# omap-heavy bucket index lands on HDD while the ssd-osd NVMe tier sits idle.
|
||||
# Pin the latency-sensitive index/metadata/mgr pools to the SSD rule and the
|
||||
# bulk non-EC pool to HDD.
|
||||
#
|
||||
# Runs HERE (after RGW readiness) on purpose: .rgw.root and the prod-z1.rgw.*
|
||||
# metadata pools are created lazily by the realm/zone setup and the RGW daemon,
|
||||
# so a pool snapshot taken before the daemon is up (like the top-of-file one)
|
||||
# would miss them. Re-query fresh, and gate each pin on the pool existing now.
|
||||
# Opt-in per cluster via ceph_rgw_ssd_pools / ceph_rgw_hdd_pools (empty default
|
||||
# leaves a live cluster's pools where they are - re-homing forces a rebalance).
|
||||
|
||||
- name: Re-query pools after RGW is up (system pools are created lazily)
|
||||
ansible.builtin.command: ceph osd pool ls
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
register: pool_list_post
|
||||
changed_when: false
|
||||
|
||||
- name: Set initial PG count on RGW metadata pools
|
||||
ansible.builtin.command: >
|
||||
ceph osd pool set {{ item }} pg_num {{ ceph_pg_init_meta }}
|
||||
loop:
|
||||
- .rgw.root
|
||||
- "{{ ceph_rgw_zone }}.rgw.log"
|
||||
- "{{ ceph_rgw_zone }}.rgw.control"
|
||||
- "{{ ceph_rgw_zone }}.rgw.meta"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- item in (pool_list_post.stdout | default(''))
|
||||
register: pg_meta_result
|
||||
changed_when: "'set' in pg_meta_result.stdout | default('')"
|
||||
failed_when:
|
||||
- pg_meta_result.rc is defined
|
||||
- pg_meta_result.rc != 0
|
||||
- "'is not >= current' not in pg_meta_result.stderr | default('')"
|
||||
|
||||
- name: Pin replica size on the RGW metadata + mgr pools (durability)
|
||||
# These otherwise inherit the cluster-spec osd_pool_default_size / min_size,
|
||||
# unlike index/extra which are pinned. Losing 2 OSDs behind a metadata PG would
|
||||
# take out the RGW/mgr control plane, and min_size 1 permits single-replica
|
||||
# writes. Pin to the same replicated size (spice 3/2; sietch stays 2/1 via its
|
||||
# own var). Runs here, after RGW readiness, because these pools are created
|
||||
# lazily by the daemon; the pre-daemon snapshot missed them so run 1 skipped
|
||||
# the pin and they sat at the inherited 2/1 until a second converge. Idempotent:
|
||||
# only sets when the current value differs.
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
ch=0
|
||||
[ "$(ceph osd pool get {{ item }} size | awk '{print $2}')" = "{{ ceph_rgw_replicated_size }}" ] || { ceph osd pool set {{ item }} size {{ ceph_rgw_replicated_size }}; ch=1; }
|
||||
[ "$(ceph osd pool get {{ item }} min_size | awk '{print $2}')" = "{{ ceph_rgw_replicated_min_size }}" ] || { ceph osd pool set {{ item }} min_size {{ ceph_rgw_replicated_min_size }}; ch=1; }
|
||||
[ "$ch" = 1 ] && echo CHANGED || echo ok
|
||||
args:
|
||||
executable: /bin/bash
|
||||
loop:
|
||||
- .rgw.root
|
||||
- "{{ ceph_rgw_zone }}.rgw.log"
|
||||
- "{{ ceph_rgw_zone }}.rgw.control"
|
||||
- "{{ ceph_rgw_zone }}.rgw.meta"
|
||||
- .mgr
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- item in (pool_list_post.stdout | default(''))
|
||||
register: size_meta_result
|
||||
changed_when: "'CHANGED' in (size_meta_result.stdout | default(''))"
|
||||
|
||||
- name: Pin pools to the SSD device-class CRUSH rule
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
cur=$(ceph osd pool get {{ item | quote }} crush_rule | awk '{print $2}')
|
||||
if [ "$cur" != {{ ceph_rgw_ssd_crush_rule | quote }} ]; then
|
||||
ceph osd pool set {{ item | quote }} crush_rule {{ ceph_rgw_ssd_crush_rule | quote }}
|
||||
echo CHANGED
|
||||
fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
loop: "{{ ceph_rgw_ssd_pools }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_rgw_ssd_pools | length > 0
|
||||
- item in (pool_list_post.stdout | default(''))
|
||||
register: ssd_pin_result
|
||||
changed_when: "'CHANGED' in (ssd_pin_result.stdout | default(''))"
|
||||
|
||||
- name: Pin pools to the HDD device-class CRUSH rule
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
cur=$(ceph osd pool get {{ item | quote }} crush_rule | awk '{print $2}')
|
||||
if [ "$cur" != {{ ceph_rgw_hdd_crush_rule | quote }} ]; then
|
||||
ceph osd pool set {{ item | quote }} crush_rule {{ ceph_rgw_hdd_crush_rule | quote }}
|
||||
echo CHANGED
|
||||
fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
loop: "{{ ceph_rgw_hdd_pools }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_rgw_hdd_pools | length > 0
|
||||
- item in (pool_list_post.stdout | default(''))
|
||||
register: hdd_pin_result
|
||||
changed_when: "'CHANGED' in (hdd_pin_result.stdout | default(''))"
|
||||
|
||||
# --- Step 13.5: Ensure dashboard RGW user has admin caps ---
|
||||
#
|
||||
# cephadm auto-creates a "dashboard" RGW user with system=true but
|
||||
@@ -694,16 +923,6 @@
|
||||
no_log: true
|
||||
tags: [s3_keys]
|
||||
|
||||
- name: S3 key drift notice (dry flag — no action taken)
|
||||
ansible.builtin.debug:
|
||||
msg:
|
||||
- "S3 svc-user keys differ from 1P values."
|
||||
- "To align (rotates keys, requires Yucca-app re-config): re-run with -e rotate_s3_keys=true"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- s3_key_drift | default(false)
|
||||
- not (rotate_s3_keys | default(false) | bool)
|
||||
|
||||
- name: Display S3 endpoint (credentials live in 1P, not log output)
|
||||
ansible.builtin.debug:
|
||||
msg:
|
||||
@@ -711,7 +930,7 @@
|
||||
- "Access Key: op://{{ cluster_secrets_vault }}/{{ cluster_name | upper }}_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY/password"
|
||||
- "Secret Key: op://{{ cluster_secrets_vault }}/{{ cluster_name | upper }}_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY/password"
|
||||
- "Endpoint (DNS): {{ ceph_rgw_scheme }}://{{ ceph_rgw_dns_name }}:{{ ceph_rgw_port }}"
|
||||
- "Endpoint (direct): {{ ceph_rgw_scheme }}://{{ bond_ip }}:{{ ceph_rgw_port }}"
|
||||
- "Endpoint (direct): {{ ceph_rgw_scheme }}://{{ ceph_service_ip }}:{{ ceph_rgw_port }}"
|
||||
- "Region: {{ ceph_rgw_zonegroup_api_name }}"
|
||||
- "Virtual-hosted: {{ ceph_rgw_scheme }}://<bucket>.{{ ceph_rgw_dns_name }}:{{ ceph_rgw_port }}"
|
||||
- "Path-style: {{ ceph_rgw_scheme }}://{{ ceph_rgw_dns_name }}:{{ ceph_rgw_port }}/<bucket>"
|
||||
|
||||
@@ -84,3 +84,98 @@
|
||||
The running cluster's password does not match ceph_dashboard_password.
|
||||
Re-run the ceph_deploy role to reconcile, or investigate an out-of-band change.
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
|
||||
# --- Validation: device-class CRUSH pool placement ---
|
||||
# Guards the "rule created but no pool bound" failure mode surfaced by the deep
|
||||
# review: assert every pool we intended on a tier actually carries that rule, and
|
||||
# that the SSD rule bound to live OSDs (a rule whose device class has zero OSDs
|
||||
# selects nothing and leaves the pinned pool's PGs inactive). Only runs when the
|
||||
# cluster opts in to pinning (ceph_rgw_ssd_pools / _hdd_pools); skips absent pools.
|
||||
- name: Validate device-class pool placement
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
rc=0
|
||||
check() { # $1 pool, $2 want_rule
|
||||
ceph osd pool ls | grep -qxF "$1" || { echo "skip (absent): $1"; return; }
|
||||
got=$(ceph osd pool get "$1" crush_rule | awk '{print $2}')
|
||||
if [ "$got" = "$2" ]; then echo "ok: $1 -> $2"
|
||||
else echo "FAIL: $1 -> $got (want $2)"; rc=1; fi
|
||||
}
|
||||
{% for p in ceph_rgw_ssd_pools %}
|
||||
check {{ p | quote }} {{ ceph_rgw_ssd_crush_rule | quote }}
|
||||
{% endfor %}
|
||||
{% for p in ceph_rgw_hdd_pools %}
|
||||
check {{ p | quote }} {{ ceph_rgw_hdd_crush_rule | quote }}
|
||||
{% endfor %}
|
||||
n=$(ceph osd crush class ls-osd ssd 2>/dev/null | grep -c '^[0-9]' || true)
|
||||
if [ "$n" -gt 0 ]; then echo "ok: ssd class has $n OSDs"
|
||||
else echo "FAIL: ssd class has 0 OSDs (replicated_ssd would select nothing)"; rc=1; fi
|
||||
[ "$rc" = 0 ] && echo VALIDATION_OK || echo VALIDATION_FAIL
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- (ceph_rgw_ssd_pools | length + ceph_rgw_hdd_pools | length) > 0
|
||||
register: crush_placement_check
|
||||
changed_when: false
|
||||
failed_when: "'VALIDATION_FAIL' in (crush_placement_check.stdout | default(''))"
|
||||
|
||||
- name: Show device-class pool placement validation
|
||||
ansible.builtin.debug:
|
||||
msg: "{{ crush_placement_check.stdout_lines }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- crush_placement_check.stdout_lines is defined
|
||||
|
||||
# --- Alert delivery status ---
|
||||
# cephadm's alertmanager ships with no receiver, so a cluster with no configured
|
||||
# webhook routes every alert rule into a null route. Surface this loudly rather
|
||||
# than failing the play (the destination is a deliberate per-cluster decision).
|
||||
- name: Report alert-delivery status
|
||||
ansible.builtin.debug:
|
||||
msg: >-
|
||||
{{ ('ALERT DELIVERY: ' ~ (ceph_alertmanager_webhook_urls | length | string)
|
||||
~ ' webhook receiver(s) configured.')
|
||||
if (ceph_alertmanager_webhook_urls | default([]) | length) > 0
|
||||
else ('WARNING: no alert delivery configured - alertmanager has no receiver, so the '
|
||||
~ 'Prometheus alert rules fire into a null route and reach no human. '
|
||||
~ 'Set ceph_alertmanager_webhook_urls to a webhook/pager endpoint.') }}
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
|
||||
# --- Validation: RGW replicated-pool durability ---
|
||||
# Guards the run-1 durability regression surfaced by the deep review: the
|
||||
# .rgw.* system pools are created lazily by the RGW daemon, so a size/min_size
|
||||
# pin taken before the daemon is up skips them and they sit at the inherited
|
||||
# 2/1. Assert every RGW replicated pool that exists carries the cluster's target
|
||||
# size/min_size, so a miss fails the play instead of shipping single-replica.
|
||||
- name: Validate RGW replicated-pool durability
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
rc=0
|
||||
check() { # $1 pool
|
||||
ceph osd pool ls | grep -qxF "$1" || { echo "skip (absent): $1"; return; }
|
||||
sz=$(ceph osd pool get "$1" size | awk '{print $2}')
|
||||
mn=$(ceph osd pool get "$1" min_size | awk '{print $2}')
|
||||
if [ "$sz" = {{ ceph_rgw_replicated_size | quote }} ] && [ "$mn" = {{ ceph_rgw_replicated_min_size | quote }} ]; then
|
||||
echo "ok: $1 size=$sz min_size=$mn"
|
||||
else
|
||||
echo "FAIL: $1 size=$sz min_size=$mn (want {{ ceph_rgw_replicated_size }}/{{ ceph_rgw_replicated_min_size }})"; rc=1
|
||||
fi
|
||||
}
|
||||
{% for p in [ceph_rgw_index_pool, ceph_rgw_extra_pool, '.rgw.root', ceph_rgw_zone ~ '.rgw.log', ceph_rgw_zone ~ '.rgw.control', ceph_rgw_zone ~ '.rgw.meta', '.mgr'] %}
|
||||
check {{ p | quote }}
|
||||
{% endfor %}
|
||||
[ "$rc" = 0 ] && echo VALIDATION_OK || echo VALIDATION_FAIL
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
register: pool_durability_check
|
||||
changed_when: false
|
||||
failed_when: "'VALIDATION_FAIL' in (pool_durability_check.stdout | default(''))"
|
||||
|
||||
- name: Show RGW replicated-pool durability validation
|
||||
ansible.builtin.debug:
|
||||
msg: "{{ pool_durability_check.stdout_lines }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- pool_durability_check.stdout_lines is defined
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user