feat(ceph): prod spice cluster on htz-fsn1 (#256)

* feat(ceph): scaffold spice prod cluster (reprovision + inventory + TF)

Stand up spice (48x Hetzner SX295, prod/htz-fsn1): the reprovision_hetzner
ansible role (rescue -> reset -> installimage -> verify, base-OS-only, with
stale-mdraid pre-clean and resume markers), the prod-htz-fsn1/spice inventory
+ group_vars/host_vars, the prod ceph TF stack, and gen-spice-host-vars.py.
Adds a `mise reprovision` task (op run + tf/.env.prod), the spice SSH key
(yucca_tf_prod), and per-host roles-based [ceph_mon] filtering so large
clusters pin a fixed MON quorum instead of defaulting to all nodes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): create spice block.db + ssd-osd LVs at converge (NVMe-RAID)

Add an NVMe-RAID branch to ceph_deploy lvm-setup so spice's single vg0 (built
by installimage) gets its 14 block.db LVs + one ssd-osd LV created at converge,
before osds.yml. This replaces painbox's fragile installimage -x chroot
post-install with an idempotent, observable ansible step; installimage stays
base-OS-only. Narrows the old blanket "externally-managed LVM" skip so it only
fires when neither the sietch dual-SSD nor the NVMe-RAID shape applies, and
wires ceph_db_vg / ceph_ssd_osd_lv / ceph_ssd_osd_reserve_gib in the spice
group_vars.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): set spice bond NICs in group_vars, drop per-host placeholders

The 25G fabric NICs are uniform across all 48 SX295 (verified live on 36
nodes): bond members enp193s0f0/f1 (Intel E800/ice, PCI c1:00.0/.1), WAN
enp197s0 (igb, c5:00.0). Set bond_interfaces + oob_nic once in the spice
group_vars -- the networkd role reads bond_interfaces -- and remove the dead
per-host fabric_nic/oob_nic PLACEHOLDER lines from all 48 host_vars,
spice-hosts.yaml, and gen-spice-host-vars.py. No MAC-based naming needed: the
predictable names are PCI-derived and identical fleet-wide.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): LACP + VLAN sub-interfaces on the bond, lldpd fabric config

networkd role: emit 802.3ad LACP params (TransmitHashPolicy, LACPTransmitRate)
when bond_mode is 802.3ad, and support tagged VLAN sub-interfaces on the bond
via networkd_bond_vlans -- when set, bond0 carries no L3 and each VLAN carries
its own address. Opt-in: an empty list keeps the flat active-backup path
(sietch) byte-identical (verified by render). verify.yml gates the
bond0-IP/gateway asserts to the flat case and checks per-VLAN addresses
otherwise.

spice: bond0 becomes an 802.3ad LACP bond of enp193s0f0/f1 (MLAG to the QFX5240
leaves) carrying VLAN 120 (public 10.40.20.<idx>/23) + VLAN 122 (private
10.40.22.<idx>/23); default route stays on the 1G WAN until cutover.

baseline role: install + configure lldpd (portid ifname, cluster system
description, service enabled) when lldpd is in baseline_extra_packages, and add
lldpd/ethtool/tcpdump for spice. This also newly applies lldpd config to sietch
(already lists lldpd) on its next converge.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): set spice RGW pool to EC 16+4 host domain for the beta

Capacity-max profile (80% usable, m=4) for the wipeable beta behind michael.
Host failure domain accepts rack-loss risk in exchange for capacity, which is
fine for a beta and matches the single-rack API tier. Also lifts the RGW DRAFT
marker and starts the bulk EC data pool at pg_num 4096 (672 OSDs, 20-chunk pool)
instead of the 36-OSD-era default to avoid PG splitting during the first fill.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* docs(ceph): drop stale -x chroot reference from spice autosetup header

installimage.yml runs base-OS-only (installimage -a -c /autosetup, no -x). The
autosetup header still carried the painbox `-x /tmp/post-install.sh` line, which
implied a chroot post-install the role deliberately dropped. Point it at the
convergence steps (lvm-setup, baseline) that replaced it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): minimal -x chroot to seed root + ansible-iac keys on reprovision

installimage sometimes runs without our key present in the rescue (Robot armed
rescue without it, in-place installimage, manual rescue), so its "copy the rescue
authorized_keys" step leaves root keyless and the node unreachable for
convergence. Add a minimal `-x /tmp/post-install.sh` that does ONLY inert
user/key writes: seed root's authorized_keys with the iac key (baseline never
manages root, so it persists) and create an ansible-iac account with the key +
NOPASSWD sudo. No apt/LVM in the chroot -- that fragile step stays at post-boot
convergence, so set -e cannot spuriously disrupt installimage.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): converge iac access on existing nodes without reimaging

Nodes imaged before the reprovision -x chroot lack root's iac key and the
ansible-iac sudo account; reseeding them meant a reinstall. baseline gains an
iac-access task that ensures the same state idempotently -- root's authorized
key (non-exclusive) plus the ansible-iac user with key and NOPASSWD sudo -- so
the 36 already-installed nodes converge in place. Gated on the cluster iac key;
run standalone with --tags iac_access.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): manual rescue-first add-node path, Robot API optional

The Hetzner Robot rescue/reset step is both the unreliable part (nodes that arm
rescue but never boot it) and the highest blast radius (an errant /reset on a
live spice node in production). Gate it behind reprovision_use_robot (default
true, preserving reprovision.yml) and add add-node.yml, which sets it false: the
operator arms rescue by hand in the portal and Ansible runs only install, chroot
and verify on a node already in rescue. Robot key registration moves to
register_robot_keys.yml, imported from preflight only on the Robot path. With the
API out of the loop nothing here can flip a running node into rescue -- wait_rescue
refuses any node without the installimage ramdisk.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): relax stale-rescue guard on the manual add-node path

wait_rescue rejects a rescue older than reprovision_rescue_max_uptime (1200s) to
catch a stale boot -- valid only when Ansible triggered the reset. On the manual
path the operator rescues by hand, so uptime just measures wait-before-run and a
node legitimately in rescue for hours would be refused. Relax the guard to a day
on add-node.yml; the installimage-ramdisk check is the real in-rescue proof.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): skip the this-boot rescue guard on the manual path

philip has sat in rescue ~8 days, so any finite uptime ceiling rejects it. The
guard only means something when Ansible triggered the reset (uptime proves this
boot); on the manual add-node path a node may sit in rescue for weeks, so gate
the assert on reprovision_use_robot rather than inflating the timeout. The
installimage-ramdisk check remains the real in-rescue proof.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): networkd/ifupdown coexistence to protect the 1G WAN

The role assumed a full migration off ifupdown -- correct for sietch's flat bond,
but on spice it would disable ifupdown and leave the 1G WAN (a separate NIC that
ifupdown owns, carrying the default route) unconfigured on the next boot: the
painbox failure. Add networkd_replace_ifupdown (default true, sietch unchanged).
When false, networkd manages only the bond and VLANs; an Unmanaged=yes guard
fences the WAN NIC off, commit leaves networking.service enabled, the bond drops
to RequiredForOnline=no so a carrier-less fabric cannot stall boot, and verify
asserts the WAN kept its address and default route before anything commits. The
rollback script targets the right interface per mode. spice sets it false.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): make migrate-networkd usable pre-Ceph (gate health checks)

The ceph-health pre/post gate assumed a live cluster and a ceph_bootstrap group,
so the play could not run during spice bringup (coexistence activation before any
cephadm deploy). Gate the four ceph tasks on ceph_bootstrap being populated; on a
pre-Ceph cluster they skip and the networkd role runs on its own.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* chore(ceph): tunable serial + halt-on-failure for migrate-networkd

Fleet rollout of the coexistence activation wants batches, not one-at-a-time, and
must stop the instant a node fails (a dropped WAN shows up as unreachable) rather
than silently skip it. Template serial (networkd_serial, default 1 unchanged) and
set max_fail_percentage 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): install Intel ice DDP firmware so E810 NICs leave Safe Mode

The minimal Debian image ships no firmware-misc-nonfree, so the E810 DDP package
(intel/ice/ddp/ice.pkg) is absent and every NIC boots into Safe Mode -- whose
crippled classifier drops reserved-multicast control frames (LACP, LLDP), so the
25G fabric bond never aggregates despite a correctly configured switch. Add
firmware-misc-nonfree to the spice package set and a baseline nic-firmware task
that reboots an E810 host once to load the DDP when it is still in Safe Mode
(self-gating: no-op on non-ice or already-loaded hosts; refuses to reboot when the
DDP is absent, so it cannot loop). The WAN is a separate igb NIC, so the reboot
keeps the node reachable.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): correct ice DDP activation gating (pipefail + dpkg check)

The Safe Mode check used `set -o pipefail`, but `ethtool --show-fec` exits
non-zero on "not supported", which pipefail propagated and masked grep's match --
so the check always yielded "ok" and the activating reboot never fired. Drop
pipefail there, and gate the reboot on the dpkg DB (firmware-misc-nonfree
installed) rather than a file stat that can lag a large apt transaction. Verified
on one node end-to-end: DDP loads, Safe Mode clears, LACP converges with the
switch, both VLAN gateways ping.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): pin chrony time sources for Ceph time sync

chrony was listed in baseline_enable_services but never installed, so system.yml
would fail starting it -- and there was no time-source config at all, leaving sync
to the distro default pool. Ceph mon quorum is skew-sensitive, so make it explicit:
add baseline/chrony.yml (install + templated chrony.conf + enable) driven by
chrony_ntp_servers (default Debian pool, makestep for the initial correction), drop
chrony from baseline_enable_services so it is owned in one place, and point spice at
Hetzner NTP (ntp1/2/3.hetzner.de) -- low-latency from FSN1, consistent across all 47.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): populate spice OSD by-path map; exclude noreen until recovered

Inspected the SX295 disk layout: 14x 20TB SATA HDDs per node across 3 AHCI
controllers, by-path uniform on 46/47 nodes -- so ceph_hdd_osds (path_phy + db
vg0/db-slotN) and ceph_ssd_osds (vg0/ssd-osd) live in group_vars, not 47 host_vars.
spice-ceph-miguel has one disk on 46:00.0-ata-4 rather than 87:00.0-ata-4 and
overrides the map in its host_vars. Verified by rendering osd-spec.yml.j2: 658 HDD
OSD paths (47x14) + 47 NVMe ssd-osd, miguel's override resolving correctly. noreen
(boot-order casualty, held in triage) is commented out of clusters.auto.tfvars so
the rendered inventory + deploy target only the 47 live nodes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): bootstrap mon on the fabric public_network, not the WAN

cephadm bootstrap --mon-ip must fall inside public_network, but spice
passed bond_ip, the 1G WAN address used for the ansible/SSH connection
and cephadm host management. spice's public_network is the 25G fabric
(10.40.20.0/23), so the initial mon would fail to bind. Bootstrap now
uses ceph_public_ip (10.40.20.<host_index>, derived like the bond VLANs);
sietch is flat with no host_index and falls back to bond_ip, already in
its own public_network.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* ci(ceph): static validation gate for the ansible/ceph stack

CI only validated the Kubernetes/Flux surface; nothing parsed the ceph
playbooks, inventory or scripts, so a broken playbook or host_var first
executed during the post-merge prod apply against real nodes. This runs
the existing local validators (yamllint, ansible-lint, shellcheck,
ansible-playbook --syntax-check, py_compile) on ansible/ceph PRs with no
secrets and no connection to any host; a throwaway localhost inventory
satisfies --syntax-check since the real inventory is TF-generated.

Its own workflow, not a job in ci.yml, so the ansible/ceph/** path filter
is workflow-scoped and does not gate the unrelated k8s jobs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): make the ceph converge deliberate and cluster-parameterized

A merge touching ansible/ceph/** could open the staging partition and
auto-run the full baseline->tune->deploy->harden pipeline against the
live sietch cluster, and the converge was hardcoded to sietch/austin
(render-inventories.sh got no region, install-ssh-keys.sh got a literal
sietch, CEPH_ENV pointed at .../sietch/), so a prod converge could never
reach spice.

The ceph converge now runs only on workflow_dispatch with
run_ceph_ansible=true, pinned to the one matrix entry that owns the
chosen ceph_cluster (spice=prod/htz-fsn1, sietch=staging/austin), so a
push never reconverges a live cluster and one dispatch cannot converge
both. Region and cluster are threaded from the dispatch inputs / matrix
into the render, key install, and CEPH_ENV. The staging paths-filter is
scoped to ansible/ceph/inventories/staging-** so a prod ceph change no
longer opens the staging matrix. The ceph TF apply stays auto and
env-gated; the mgmt converge is unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): make the RGW firewall source scope configurable

The RGW/S3 port was accepted from any source unconditionally, while
every other Ceph service (mon, osd, dashboard, monitoring) is scoped to
ceph_firewall_trusted_networks. On spice that port is plain HTTP 7480, so
the S3 endpoint sat open to the public internet with no lever to close
it. A new ceph_firewall_rgw_any_source (default true, mirroring
ceph_firewall_ssh_any_source) keeps the open behavior by default but lets
production restrict RGW to the trusted networks, which already cover the
fabric and NetBird overlay.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* chore(ceph): fix stale OSD-map comment on the noreen host_var

noreen was generated before the by-path map moved into group_vars, so it
still carried the DEFERRED note while its 47 siblings point at the group
var. The host stays pre-staged for re-add once recovered; only the
comment was wrong.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): add the host-mgmt VLAN (124) to the spice bond

The FSN1-C1 fabric carries a third VLAN on the 25G bond,
FSN1-C1-HOST-MGMT (124, 10.40.24.0/24), alongside Ceph public (120) and
private (122). The leaf already trunks it to every server bond and
advertises 10.40.24.0/24 into the NetBird overlay, so it is the intended
in-band ansible/SSH reach once the WAN is retired, but the host side had
no matching sub-interface. Add bond0.124 at 10.40.24.<host_index>/24 with
no gateway, so the default route stays on the 1G WAN until cutover.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): apply networkd config deltas to an already-running daemon

The networkd role only started systemd-networkd, which is a no-op once it
is already running, so a re-run that added config (e.g. a new bond VLAN)
wrote the .netdev/.network files but never applied them, and verify then
failed asserting the sub-interface had no address. Add a networkctl reload
plus a per-VLAN settle wait after the start, so a re-run creates the added
sub-interfaces live without tearing down existing links; the WAN on
ifupdown is untouched regardless. Non-disruptive on the flat sietch path
(the settle wait is gated on networkd_bond_vlans).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): bind Ceph services on the fabric only, never the WAN

Ceph on spice must never listen on the 1G WAN. mon/mgr/OSD already bind
per public_network/cluster_network, but the beast RGW frontend and the
cephadm monitoring stack bound 0.0.0.0, and the mon-ip, cephadm host-add
address, and dashboard URLs advertised the WAN bond_ip. Add a per-cluster
policy: ceph_service_ip (the address Ceph advertises and binds for point
services; defaults to bond_ip, resolving to the fabric on spice via
ceph_public_ip) and ceph_bind_networks (the cephadm networks: bind
restriction for RGW and the monitoring daemons; defaults to
public_network). Route mon-ip, the host-add address, the dashboard
monitoring URLs, RGW, and prometheus/grafana/alertmanager/node-exporter/
ceph-exporter through them, and scope the RGW firewall to the trusted
networks. The RGW zonegroup hostnames and TLS SAN gain the fabric address
so signed requests to the fabric IP still validate. Flat clusters like
sietch have no ceph_public_ip, so ceph_service_ip falls back to bond_ip
and ceph_bind_networks to their single flat network; point services are
unchanged there.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): define the metrics-worker RGW user on spice

rgw.yml Step 14.5 creates a read-only RGW admin for the yucca-metrics-worker
service and references ceph_rgw_metrics_user_*, which only sietch defined,
so the RGW phase failed on spice with AnsibleUndefinedVariable. Add the
block; the keys are the TF-minted SPICE_METRICS_WORKER_* items, injected
via secrets.yml.tpl as vault_metrics_worker_*, so op inject resolves them.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): refuse to deploy against an un-rendered inventory

The ceph_deploy role gates every phase on the ceph_bootstrap/ceph_mon/
ceph_join groups that render-inventories.sh emits from TF state; run
against the hand-written stopgap inventory that lacks them, the role
errors mid-deploy or falls through to placing a MON on every host. A
pre-task assert now requires ceph_bootstrap to be exactly one host that
is also in ceph_mon, and ceph_mon to be a non-empty odd quorum, and
stops with a render hint otherwise.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): make the fabric-only bind restriction opt-in per cluster

ceph_bind_networks defaulted to public_network, which would have injected
a networks: bind restriction into every cluster - re-binding sietch's live
RGW and monitoring daemons from 0.0.0.0 to its flat network on the next
converge (a redeploy, and a break for any access path not on that subnet).
Default it empty instead: no networks: field is emitted and the monitoring
re-spec is skipped, so flat clusters bind every interface exactly as
cephadm ships them. spice opts in to the fabric public network in its
group_vars.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): make ceph_service_ip a group_var so hostvars can read it

ceph_service_ip was a ceph_deploy role default, but join/rgw/monitoring
read it as hostvars[<host>]['ceph_service_ip'], and role defaults are NOT
exposed through hostvars - the lookup raised "HostVarsVars object has no
attribute ceph_service_ip" and would abort the deploy on every cluster at
the join phase (a regression the fabric-only change introduced for both
spice and the live sietch). Define it in each cluster's group_vars instead
(group_vars do resolve through hostvars, verified per host): spice to the
fabric ceph_public_ip, sietch to bond_ip (unchanged from pre-fabric).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* chore(ceph): drop the unused ceph_cluster_ip var on spice

ceph_cluster_ip was defined for symmetry with ceph_public_ip but consumed
nowhere - OSD replication binds to the cluster_network CIDR, which cephadm
resolves to each node's bond0.122 address on its own. Remove the dead var
and correct the neighbouring comment (ceph_service_ip is a group_var now,
not a role default).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): serve spice RGW over TLS on 443 with a 10-year self-signed cert

Match sietch: flip spice RGW/beast from plain HTTP 7480 to HTTPS 443 with a
self-signed 10-year cert. rgw.yml generates /etc/ceph/rgw-ssl.{crt,key} and
cephadm distributes it to every RGW daemon via the service spec; the SANs
already cover s3.<domain> + the wildcard + each node's fabric IP
(ceph_service_ip). The firewall RGW port derives from ceph_rgw_port, so it
follows to 443, scoped to the trusted networks (RGW binds the fabric only,
never the WAN). This also makes the discovery rgw_s3_endpoint (hardcoded
https) correct for spice.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(dns): prod Cloudflare DNS stack for the spice RGW S3 endpoint

Add tf/deployment/prod/global/dns (mirroring staging/global/dns) so
s3.prod.fsn1.htz.futo.cloud and its wildcard resolve. Round-robin A across
all 47 spice ceph nodes' fabric public IPs (10.40.20.<host_index>, VLAN 120)
where RGW/beast binds; proxied=false (private RFC1918, reached over the
NetBird-advertised cls1_public 10.40.20.0/23). The wildcard serves S3
virtual-hosted buckets, and both names are in the self-signed TLS cert SANs.
noreen (host_index 40) excluded. CI auto-discovers the stack (applies at
order 0 under the prod-global environment); tf/.env.prod gains the token ref.

Next up: create op://yucca_tf_prod/CLOUDFLARE_API_TOKEN before the apply.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): preflight-assert the Ceph service IP is up before deploy

A node whose fabric VLAN sub-interface (spice bond0.120) did not come back
after a reboot would otherwise fail deep inside cephadm bootstrap/join. Add
a per-host pre_task that asserts ceph_service_ip is present in
ansible_all_ipv4_addresses, so a fabric-down node halts up front with a
clear message. Passes on a healthy node (verified on spice-ceph-adelia);
on sietch ceph_service_ip is bond_ip, the connection address, so it is
trivially satisfied.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* chore(ceph): honest ssd_model_pattern note + guard empty monitoring networks

ssd_model_pattern on spice was "SAMSUNG" with a "verify on node 1" note, but
the NVMe is SOLIDIGM/KIOXIA and the var is only read by ceph_destroy's
partition-6 SSD wipe (sietch's dual-SSD shape); spice is NVMe-RAID with no
partition 6 and a vg0 ssd-osd LV cleaned by the generic VG/PV removal, so
the pattern is a deliberate no-match. Reword to say so (it must stay defined
because cleanup.yml references it unconditionally). Also wrap the
monitoring-spec networks block in a length guard so an empty ceph_bind_networks
renders no dangling `networks:` key (defensive; the apply is already gated).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): lock down root/password SSH in baseline, not just harden

installimage ships PermitRootLogin yes + a root password, and all sshd
hardening lived in harden.yml (the last deploy stage), so a freshly imaged
node sat root-password-open on the public WAN for the whole campaign (the
11-node exposure the swarm audit found). Add an early baseline task, right
after iac-access authorizes the key on root, that deploys a 10-baseline-ssh
drop-in (PermitRootLogin prohibit-password + PasswordAuthentication no,
same values as security/50-hardening.conf so they never disagree), locks
the root password, and removes the interim remediation drop-in. A
Validate->Reload handler chain runs sshd -t before reloading, only on
change. Key-safe on both clusters (ansible connects by key), so nothing
can lock out.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): pin RGW metadata + mgr pools to the replicated size

The RGW system-metadata pools (.rgw.root, {zone}.rgw.{log,control,meta})
and .mgr inherited the cluster-spec default osd_pool_default_size 2 /
min_size 1, unlike the index/extra pools which are pinned. Losing two OSDs
behind a metadata PG would take out the RGW/mgr control plane, and
min_size 1 permits single-replica writes. Pin them to ceph_rgw_replicated_
size/_min_size (spice 3/2; sietch keeps 2/1 via its own vars), existence-
gated like the pg_num loop and idempotent (only sets when the value
differs).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): make the ops password hash idempotent

users.yml hashed ops_password with password_hash('sha512') and no salt, so
a fresh random salt was drawn every run: the hash never matched /etc/shadow
and the user module rewrote it reporting 'changed' on every converge (a
clean converge was never a true green signal, on both clusters). Derive a
stable salt from a one-way sha256 of the password so the hash is
deterministic and idempotent, while still reconciling an out-of-band
password change. The salt in /etc/shadow is public and leaks nothing.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): make live fabric/hosts/timezone reproducible from the repo

Three drift gaps where the running fleet did not match the committed IaC.
networkd_enabled was committed for sietch only, so spice's fabric VLANs
came up from an ad-hoc -e and a reprovisioned node would not bring them up
from the repo alone; commit it true for spice (networkd_replace_ifupdown
false still fences the 1G WAN). hosts.j2 mapped every node name to bond_ip
(the 1G WAN) rather than the fabric ceph_service_ip, so in-cluster name
resolution pointed off the fabric. And timezone: UTC was declared but never
applied, leaving nodes on the image default (Europe/Berlin); add a
community.general.timezone task to baseline/system.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): scope OSD/RGW readiness waits to the in-play host set

join.yml is --limit-aware (it intersects ansible_play_hosts), but the OSD
provisioning wait computed EXPECTED from all ceph_nodes and the RGW wait
blocked until running rgw >= all ceph_nodes, so a subset (--limit) deploy
joined the subset then deadlocked at both gates. Base both counts on
groups['ceph_nodes'] intersect ansible_play_hosts, matching join.yml. On a
full deploy the intersection is all nodes, so behavior is unchanged; the
rgw-spec placement stays pinned to all ceph_nodes (only the readiness gate
is play-scoped, so a --limit re-run cannot shrink the spec).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): scope the nftables ruleset to table inet filter

nftables.conf.j2 did a global `flush ruleset`, and the no-op guard diffed
the whole `nft list ruleset`. On a ceph node (cephadm runs podman) that
wiped the podman/docker/fail2ban tables on every apply, and made the
desired-vs-live diff never match (podman tables are live but absent in the
throwaway netns), so the firewall reloaded on every converge - each one
flushing podman again. Replace only table inet filter (add, delete,
re-add), compare only that table in the guard, and reload via ExecReload
(nft -f, no global flush) instead of restart (whose ExecStop flushes the
whole ruleset). The fabric/ceph rules themselves are unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): deploy MGR at a capped count, not one per host

placement.yml applied mgr to every active host (47 mgr daemons on spice:
1 active + 46 standby), which is wasteful and non-standard - MON already
uses a bounded ceph_mon set. Deploy count:{{ ceph_mgr_count }} (default 3:
1 active + 2 standby) instead; cephadm schedules them and caps at the host
count on small clusters like sietch.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): mark the RGW EC data pool as bulk

The EC data pool is pre-seeded to pg_num 4096 (sized for spice's 658 HDD
OSDs), but without the bulk flag the pg_autoscaler can walk it back toward 1
during the first fill, causing PG splitting under load. Set bulk=true so the
autoscaler targets a full-capacity pg_num and treats the pre-seed as a
floor. Idempotent: only set when not already true.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): harden the ceph_destroy SSD-model match

The SSD OSD PV-remove and partition-wipe loops used
grep -i "$ssd_model_pattern", which interpreted the pattern as a regex and,
on an empty pattern, matched every disk - the worst-case foot-gun in a
destroy path (it would target all disks). Use grep -iF (fixed string, no
regex) and skip the loop entirely when the pattern is empty. Kept -F
without -w, since -w would fail to match underscore-containing model
strings like Micron_5100_MTF... Shared role, so it hardens sietch too.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* chore(ceph): drop the dead s3_key_drift notice; reconcile harden header

The "S3 key drift notice" debug task was gated on s3_key_drift, a variable
nothing ever sets (the drift detection it stood in for was never built), so
the branch could never fire - remove it. And update the harden.yml header:
the security-critical sshd lockdown (root key-only, no password auth, locked
root) now runs early in baseline, not gated behind this last stage; harden
adds the firewall and the remaining sshd hardening on top.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): accurate change reporting in tuning roles (T4)

CRUSH tunables and per-device HDD/SSD sysfs writes ran a mutating
command every converge under changed_when:false, so a run reported ok
even when it changed state and --check misled register consumers.
Switch to get-then-set: read the current value first, only write when
it differs, and emit CHANGED so changed_when reflects reality. Same
settings applied; only change-detection becomes truthful. Also drops
two pre-existing em-dashes to keep the file plain ASCII.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* refactor(ceph): declarative lvg/lvol for OSD block.db setup

Replace hand-rolled pvcreate/vgcreate/lvcreate shell in ceph_deploy
lvm-setup with community.general.lvg/lvol (already pinned, and the
idiom provision_host/disks.yml already uses). Same VG/LV names, sizes,
fixed-vs-100%FREE split, per-shape loops, and device targets on both
shapes. Kept read-only asserts/verify/show tasks as-is.

Preserved the wipefs -af signature clear ahead of PV creation: lvg does
pvcreate -f but not --yes, so it will not wipe a stale foreign signature
on a reused partition. wipefs stays gated on the VG being absent so a
live PV is never touched. Kept a minimal shell to compute the spice
ssd-osd size (vg0 free - reserve) since lvol has no free-minus-N
primitive, plus an existence guard so re-runs don't recompute a bad
size or trigger an unforced lvol shrink.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* chore(ceph): ASCII-clean the lvm-setup header comment

Convert the pre-existing arrows and em-dashes in the header (left untouched
by the lvg/lvol refactor) to plain ASCII.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): never let lvol shrink a live block.db LV (data safety)

The lvg/lvol refactor let community.general.lvol reconcile an existing
fixed-size db-slot to its size, and lvol defaults to shrink:true - so a
db-slot that ever drifted LARGER would be TRUNCATED, corrupting a live
BlueStore block.db and taking out the OSD. The old shell skipped existing
LVs entirely, so it could never do this. Set module_defaults shrink:false
on both lvm-setup blocks: lvol still creates and may grow, but never
truncates an existing LV (it leaves a larger one alone). Also restore
opts:-Wy so (re)created LVs wipe stale signatures - the destroy->recreate
recovery-path freshness the explicit -Wy gave.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): RGW self-signed cert generation aborted on first converge

The openssl -subj and -addext args used backslash-newline continuations
inside their quoted strings, so shell line-folding kept the next line's
indentation in the value; countryName rendered as "DE    " (over the
2-char max) and openssl exited 1 under set -euo pipefail, leaving no S3
endpoint on the first spice deploy. Precompute the subject and SAN as
single-line facts (whitespace-controlled Jinja for the SAN) and pass
them quoted, so each openssl arg is one flat string.

See ansible/ceph/roles/ceph_deploy/tasks/rgw.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): green the ansible-lint CI gate on HEAD

Bare ansible-lint (what mise run lint and the required CI check run)
exited 2 on two deliberate patterns: run_once on the fleet-wide
placement assert, and an intentional no-pipefail shell in the ice DDP
Safe-Mode probe (pipefail there would mask grep's match). Waive both
with inline noqa so the reasoned patterns stay and the gate passes.

See ansible/ceph/deploy-ceph.yml and roles/baseline/tasks/nic-firmware.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): pin RGW pools to device-class CRUSH rules (SSD fast tier)

The replicated_ssd/replicated_hdd rules were created but bound to no pool,
so every replicated pool fell to the default class-agnostic rule and the
omap-heavy bucket index landed on HDD (~98.5% of weight) while spice's
47-OSD NVMe ssd-osd tier sat idle. Pin the index, RGW metadata, and mgr
pools to the SSD rule and the bulk non-EC pool to HDD; force the OSD
device class at creation (ssd-osd -> ssd, HDD -> hdd) so the rules bind
deterministically, and assert the placement in verify.yml. Pinning is
opt-in per cluster (empty default) so a live cluster is never re-homed
implicitly; the pin runs after RGW readiness because the .rgw.* system
pools are created lazily by the realm/zone setup and the daemon.

See ansible/ceph/roles/ceph_deploy/{templates/osd-spec.yml.j2,tasks/rgw.yml,tasks/verify.yml}.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): apply RGW pool PG/durability pins on the first converge

The pool list was captured once at the top of the RGW play, before any
pool was created, so every "pool in list" gate skipped on a fresh run 1:
the data/index/non-EC PG pre-sizing deferred to a second converge, and
the .rgw.* system pools (created lazily by the daemon) never got their
size/min_size pin at all, sitting at the inherited 2/1 single-replica
until someone ran the play twice. Re-query the pool list once the
explicit pools exist, and move the system-pool PG + size/min_size pins
below the RGW readiness wait where those pools are real. Parameterize the
bootstrap osd_pool_default_size per cluster (spice 3/2, role default keeps
sietch at 2/1) so an unpinned future pool is not born single-replica, and
assert final size/min_size in verify.yml so a miss fails the play loudly.

See ansible/ceph/roles/ceph_deploy/tasks/{rgw.yml,verify.yml}.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ceph): real ok-to-stop gate before a live-node reimage

Reimaging a spice node zeroes the OS NVMe, and the 14 HDD OSD block.db
LVs plus the ssd-osd data LV all live on that NVMe, so a reimage destroys
every OSD on the node; the old guard only checked a hand-typed
i_know_its_dead flag and the prepare-os-disks comment wrongly claimed a
reimage keeps OSD data. Replace the stub with a mon-delegated check that
queries the node's OSD ids, refuses on HEALTH_ERR or a failed
ceph osd ok-to-stop, and sets noout on those OSDs for the reimage window.
The gate is default-on: auto mode enforces whenever a live cluster with
this node's OSDs is reachable and no-ops for the initial bootstrap;
strict fails closed if it cannot verify; permissive is the explicit
escape hatch.

See ansible/ceph/roles/reprovision_hetzner/tasks/ceph_safety.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): jumbo MTU 9000 on the VLAN 122 cluster/replication network

Jumbo pays off on the OSD replication path (a closed, homogeneous fabric)
but is a partial-blackhole risk on the RGW-facing public network, which
serves heterogeneous 1500-byte S3 clients and the ~1400 NetBird overlay,
so raise MTU on VLAN 122 (Ceph private) only and leave 120/124 at 1500. A
VLAN opts into jumbo with a per-entry mtu; the bond ceiling and both 25G
members auto-raise to the largest child MTU so a 1500 parent or slave
cannot silently cap the jumbo frames. verify.yml asserts the applied MTU
on the bond, members, and each VLAN (local, no switch dependency). Flat
clusters resolve to 1500 and emit no MTU, so sietch is unchanged.

Paired with a QFX-side change (leaf server-LAG + VLAN 122 IRB >= 9000) and
a ping -M do -s 8972 matrix before the cluster network is relied on.

See ansible/ceph/roles/networkd.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* chore(dns): commit prod DNS provider lock file

The prod DNS stack was the only DNS stack without a committed
.terraform.lock.hcl, so cloudflare ~> 5.0 would resolve unpinned at the
first prod apply of the stack fronting the S3 endpoint. Pin it to 5.21.0,
matching the staging DNS stack, with linux_amd64 + darwin_arm64 hashes.

See tf/deployment/prod/global/dns/.terraform.lock.hcl.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): opt-in alertmanager receiver so alerts reach a human

cephadm deploys alertmanager with no receiver, so every Prometheus alert
rule (OSD down, host down, PG degraded, near-full) fires into the default
null route and nobody is notified. Add ceph_alertmanager_webhook_urls: a
non-empty list routes all alerts to those webhooks via cephadm
user_data.default_webhook_urls. The monitoring re-spec now applies when
either fabric bind or alerting is configured (independent gates), and
verify.yml reports the delivery status, warning loudly when none is set.
Empty by default, so flat clusters (sietch) are unchanged; the spice
destination is a deferred operator decision (TODO in defaults).

See ansible/ceph/roles/ceph_deploy/tasks/monitoring.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(fabric): jumbo MTU on the leaf server bonds + VLAN 122 IRB

Completes the host-side VLAN 122 jumbo change on the QFX leaf: the server
LACP bonds (ae1..aeN) get the 9216 L2 jumbo MTU that the spine uplink ae0
already carries (802.3ad members inherit it, so no per-member mtu), and the
private/cluster IRB gateway (VLAN 122) gets an L3 family-inet MTU of 9000 to
match the Ceph hosts. Public (120) and host-mgmt (124) IRBs stay at 1500 by
design, so jumbo is confined to the closed OSD-replication path.

See tf/shared/modules/cluster-fabric.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ci): pin + SHA256-verify NetBird install, drop curl|sh

The netbird-connect action piped pkgs.netbird.io/install.sh straight into
sh on the runner that holds the prod-write 1Password token and overlay
access to the live nodes, so a compromised installer would run as root
there. Download a pinned release tarball (v0.74.4) from the GitHub release
and verify its SHA256 against an in-repo pin before unpacking; fail closed
on any mismatch. A version input plus a refresh comment keep the pin
maintainable.

See .github/actions/netbird-connect/action.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* fix(ci): close the prod apply gate (reviewed-plan binding + fail-closed reviewers)

The prod approval gate could fail open and the apply was not bound to the
reviewed diff. Manage the four gate Environments as code (new
_meta/github stack: github_repository_environment + required reviewers,
protected-branches-only, a validation that rejects a reviewerless gate),
and make the discover job fail closed by asserting every emitted
partition-region Environment already carries a required reviewer. Bind
apply to the plan: the plan job writes -out to an absolute path and
uploads it, the gated apply downloads that exact file and applies it with
no re-plan and no -auto-approve, so a stale plan fails closed. Scope a
bare workflow_dispatch to plan-only and a toggled dispatch to just the
partition it touches; add fail-fast and timeout-minutes across the jobs.

The _meta/github stack is bootstrap-applied out of band and needs real
reviewer IDs plus a GH_ENV_ADMIN_TOKEN secret before the gate passes.

See .github/workflows/infra.yml and tf/deployment/_meta/github.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): blast-radius control on the day-2 converge plays

The converge chain (baseline, tune-os, tune-hardware, tune-ceph, harden)
ran against all nodes at once with no canary and no stop-on-failure, the
inverse of the serial:1/max_fail:0 destructive plays. Batch each converge
play (serial, default 10%) with max_fail_percentage:0, and gate every
batch on cluster health via a shared pre/post ceph-health checkpoint that
halts the roll on HEALTH_ERR. The gate is a no-op until a cluster is live,
so the bootstrap ordering (baseline -> tune -> deploy) is unaffected;
deploy-ceph stays big-bang. Tune the batch size with ceph_converge_serial.

See ansible/ceph/tasks/health_gate.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): freeze load-bearing host packages instead of auto-upgrading

We run no unattended-upgrades: an apt upgrade that moved cephadm's
container runtime (podman + ecosystem) or the chrony time source that mon
quorum depends on out from under a live cluster is the uncoordinated
change we avoid. Hold those packages at their installed version (dpkg
selection) so apt upgrade skips them; upgrading is then deliberate and
health-gated. Diagnostic/ops tools are left unheld. Runs last in baseline
so nothing is held before it exists; tune via baseline_held_packages.

See ansible/ceph/roles/baseline/tasks/hold-packages.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): THP=madvise and stop the GRUB cmdline clobber

Transparent Huge Pages sat at the kernel default 'always', which bloats
BlueStore/tcmalloc RSS and drives allocation-stall latency across the OSD
fleet; set it to madvise via a oneshot unit ordered Before=ceph-osd.target
(mirroring ceph-cpu-governor.service) plus a live sysfs write, on by
default for every ceph cluster. Separately, the processor.max_cstate GRUB
task rewrote GRUB_CMDLINE_LINUX_DEFAULT wholesale (dropping existing
tokens, injecting quiet) - a latent footgun behind the default-off
governor flag; replace it with a /etc/default/grub.d drop-in that only
appends the cstate token.

See ansible/ceph/roles/hardware_tuning.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): host security sysctls + persistent journald

Add a host kernel/network hardening sysctl set (syncookies, ignore/deny
ICMP redirects + source routing, kptr/dmesg restrict, log martians) to the
existing os_tuning sysctl drop-in. rp_filter is set to LOOSE (2), not
strict (1): these nodes are multi-homed (1G WAN default route + 25G fabric
VLAN sub-interfaces), so strict reverse-path filtering would blackhole
asymmetric fabric traffic. Separately make journald persistent
(Storage=persistent, bounded SystemMaxUse/RuntimeMaxUse) so logs survive
the pipeline's own reboots on these headless nodes.

See ansible/ceph/roles/os_tuning.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): health-gated cephadm upgrade playbook + rollback runbook

No upgrade automation existed for the pinned 20.2.x tentacle train. Add a
deliberate, manually-invoked upgrade-ceph.yml: preflight-gates on HEALTH_OK
(with an explicit allow-WARN toggle), no in-progress upgrade, and all OSDs
up+in; records and prints the rollback image before starting; refuses to
run without an explicit target (ceph_upgrade_target_image/_version, no
default); drives ceph orch upgrade with a bounded poll and fails loudly on
a stall; asserts health + version convergence after. Not wired into
site.yml. Rollback procedure lives in the play header and
docs/runbooks/upgrade-ceph.md.

See ansible/ceph/upgrade-ceph.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): scheduled cluster-state backups with a systemd timer

Config/topology state was captured only manually and controller-local. Add
a ceph_backup role that installs a capture script + daily systemd timer on
the bootstrap node: each run tars fsid, config dump, mon/osd/crush maps
(raw + decoded), osd tree, orch ls/host ls, and the RGW realm/zonegroup/
zone into a root-only 0700 dir, pruned by retention. No secret keyrings are
dumped. An offsite target (rsync or s3://) is a var left empty for the
operator. Invoke via backup-ceph.yml (mise run backup-timer); not wired
into converge. Opt out with ceph_backup_enabled=false.

See ansible/ceph/roles/ceph_backup and docs/runbooks/backup-restore.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): pin mgr dashboard/prometheus bind to fabric; pin EC min_size

The mgr dashboard (8443) and prometheus module (9283) run inside the active
mgr and bound 0.0.0.0, closed only by nftables. Pin each mgr instance's own
localized server_addr (mgr/<module>/<mgr_id>/server_addr, the key the
modules read via get_localized_module_option) to that host's fabric IP,
discovering mgr ids at runtime so it is failover-safe; a global server_addr
would leave the dashboard unbindable after failover. Opt-in on
ceph_bind_networks; takes effect on the next mgr cycle (nftables holds the
WAN closed meanwhile). Separately pin the EC data pool min_size explicitly
to k+1 (derived from ceph_rgw_ec_k) so the write-availability floor is
documented and cannot drift, and record the single-site host-failure-domain
DR ceiling in docs/capacity-planning.md.

See ansible/ceph/roles/ceph_deploy/tasks/monitoring.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* chore(ceph): freeze Galaxy collection versions to a major

The collection requirements floated on >= lower bounds, so a new
community.general major could be pulled mid-campaign. Pin compatible-release
ranges to the currently-resolving majors (ansible.posix >=2,<3;
community.general >=12,<13) so a 47-node campaign cannot cross a major
between runs. No lockfile mechanism exists in the repo, so the ranges live
in requirements.yml.

See ansible/ceph/requirements.yml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* test(ceph): molecule scenario for the fabric-only firewall branch

The only molecule scenario tested the default open-firewall branch
(*_any_source: true), the opposite of spice's production posture, with no
idempotence check. Add a fabric-only scenario that flips the closed branch
(rgw/ssh any_source: false, trusted networks = fabric + NetBird) and
asserts RGW/SSH are NOT accepted from any source, are restricted to the
fabric/overlay, and that a second render is idempotent. Mirrors the default
scenario's render-and-grep idiom; the existing scenario is untouched.

See ansible/ceph/roles/security/molecule/fabric-only.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* ci(dns): drift-check the prod S3 roster against the ceph node list

The 47 S3 RGW A-records were hand-copied with no link to the ceph roster,
so add/replace/remove of a node silently blackholed the endpoint or dropped
capacity. A terragrunt dependency is not viable here (no dependency idiom in
the repo, and the ceph discovery output carries only bond_ip, never the
fabric IP), so add a stdlib-only, credential-free check that reconstructs
the expected fabric-IP roster from clusters.auto.tfvars (in-service names) +
spice-hosts.yaml (host_index) and asserts the apex and wildcard A-records
match exactly. A path-scoped CI gate fails the PR on drift before the DNS
apply; also runnable via mise run tf:check-dns-roster.

See tf/scripts/check-s3-dns-roster.py.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): inotify headroom for podman scale; podman-compat upgrade note

cephadm-on-podman runs many containers/systemd units per node and leans on
inotify; Debian's default fs.inotify.max_user_instances=128 is a known
bottleneck at that scale, so raise instances to 512 and watches to 524288
in os_tuning. Separately, document in the upgrade runbook that the podman
dpkg-hold must be lifted (unhold -> bump to a cephadm-supported version ->
re-hold) before a cross-major Ceph upgrade; patch-train upgrades need no
podman change.

See ansible/ceph/roles/os_tuning and docs/runbooks/upgrade-ceph.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(ceph): health-checked RGW ingress VIP (haproxy + keepalived)

S3 was bare round-robin DNS across 47 RGW A-records, so a dead node
blackholed ~1/47 of new connections for the TTL. Add an opt-in cephadm
ingress service: a keepalived floating VIP (spice 10.40.20.250) with
haproxy health-checking the RGW backends and dropping a dead one from
rotation. haproxy terminates TLS on the VIP with the existing self-signed
RGW cert (its SAN now carries the VIP) and re-encrypts to beast: L4
passthrough for RGW needs IngressSpec.use_tcp_mode_over_rgw, which is
absent on the pinned 20.2.2 (source-verified; applying it TypeErrors), so
terminate is the only health-checked mode available. Applies after RGW
readiness; asserts its preconditions (VIP set, ceph_bind_networks non-empty
so beast binds a per-node IP and does not collide with the VIP on :443,
ceph_rgw_ssl true). Opt-in per cluster; sietch unaffected.

See ansible/ceph/roles/ceph_deploy/templates/rgw-ingress-spec.yaml.j2.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* feat(dns): point the prod S3 endpoint at the ingress VIP

Collapse the s3.prod.fsn1.htz.futo.cloud apex + wildcard from 47 round-robin
A-records to the single health-checked ingress VIP 10.40.20.250, and switch
the drift-check to assert both records equal that VIP (the ceph
ceph_rgw_ingress_vip is the source of truth). Apply ordering matters and is
noted in the tfvars: the ingress must be live before this DNS cutover, or
s3 resolves to a VIP nothing answers; roll back in reverse.

See tf/deployment/prod/global/dns/records.auto.tfvars.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* chore(ceph): fix stale staging copy-paste comments in the prod stack; fmt

The prod ceph stack carried the staging stack's comments verbatim: versions.tf
and secrets.tf described creating SIETCH_CEPH_* items in yucca_tf_staging via
OP_TF_YUCCA_STAGING_ENV_WRITE, and a variable example said (sietch, ...). This
stack creates SPICE_CEPH_* in yucca_tf_prod via OP_TF_YUCCA_PROD_ENV_WRITE;
correct the comments to match. Legitimate cross-refs (the mirrors-staging/talos
provenance, the partition-slug enumeration) are left as-is. Also tofu fmt the
drifted clusters.auto.tfvars (whitespace only).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK

* ci(infra): scope reviewer fail-closed gate to apply events; prettier-format workflows

* ci(infra): pin 1Password CLI version to avoid flaky latest resolution

* feat(dns): resolve prod S3 apex+wildcard to the node fabric-IP roster

* ci(infra): park the Environments approval gate (disable fail-closed step + stub _meta/github)

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Andy Molenda
2026-07-14 12:01:10 -07:00
committed by GitHub
co-authored by Claude Opus 4.8
parent 26b58feb15
commit 0357aac743
188 changed files with 6823 additions and 290 deletions
+49 -2
View File
@@ -20,15 +20,62 @@ inputs:
description: Optional peer hostname to register the runner as.
required: false
default: ''
version:
description: >-
Pinned NetBird client version to install. MUST match the in-repo SHA256
pins in the install step below (bump both together).
required: false
default: '0.74.4'
runs:
using: composite
steps:
- name: Install NetBird client
# Install a PINNED NetBird release and verify its SHA256 against an in-repo
# pin before executing anything -- never `curl | sh`. This runner holds the
# prod-write 1Password SA token and overlay access to the live nodes, so a
# compromised pkgs.netbird.io/install.sh would run as root here.
#
# To bump: pick a version, then read the published checksums and copy the
# `netbird_<ver>_linux_<arch>.tar.gz` lines into SHA256 below:
# curl -fsSL \
# https://github.com/netbirdio/netbird/releases/download/v<ver>/netbird_<ver>_checksums.txt
- name: Install NetBird client (pinned + SHA256-verified)
shell: bash
env:
NB_VERSION: ${{ inputs.version }}
run: |
set -euo pipefail
curl -fsSL https://pkgs.netbird.io/install.sh | sh
# Pinned SHA256 of netbird_${NB_VERSION}_linux_<arch>.tar.gz.
declare -A SHA256=(
[amd64]=b854710409e7de79071642e2c328b181d4db7b9c90c298ee32b6e3b5d9ebd36f
[arm64]=af6fc4909bcc68e55197eddbecd8a8ec1b02bc9b34cc28ef72e2858375e138d9
)
case "$(uname -m)" in
x86_64) arch=amd64 ;;
aarch64) arch=arm64 ;;
*) echo "unsupported CPU arch: $(uname -m)" >&2; exit 1 ;;
esac
want="${SHA256[$arch]:-}"
[ -n "$want" ] || { echo "no pinned SHA256 for arch '$arch'" >&2; exit 1; }
tarball="netbird_${NB_VERSION}_linux_${arch}.tar.gz"
url="https://github.com/netbirdio/netbird/releases/download/v${NB_VERSION}/${tarball}"
tmp="$(mktemp -d)"; trap 'rm -rf "$tmp"' EXIT
curl -fsSL --retry 3 --proto '=https' --tlsv1.2 -o "$tmp/$tarball" "$url"
# Fail closed on any mismatch BEFORE the artifact is unpacked or run.
echo "${want} ${tmp}/${tarball}" | sha256sum -c -
tar -xzf "$tmp/$tarball" -C "$tmp" netbird
sudo install -m 0755 "$tmp/netbird" /usr/local/bin/netbird
# Register + start the daemon (what the upstream installer would do), so
# the later `netbird up` has a service to talk to.
sudo netbird service install 2>/dev/null || true
sudo netbird service start 2>/dev/null || true
netbird version
- name: Connect to NetBird
@@ -0,0 +1,91 @@
name: ansible-ceph-validate
# Static validation gate for the ansible/ceph stack. ci.yml only validates the
# Kubernetes/Flux surface (mise run k8s:validate); nothing there parses the ceph
# playbooks, inventory or scripts. Without this gate a broken playbook, host_var
# or script arg first executes during the post-merge PROD apply against real
# nodes. This job runs the existing local validators (yamllint + ansible-lint +
# shellcheck + syntax-check) with NO secrets and NO connection to any host.
#
# It lives in its own workflow (not a job in ci.yml) so the path filter below is
# workflow-scoped: `on.<event>.paths` gates the whole file. Adding the same
# filter inside ci.yml would gate every ci.yml job (k8s validate, tests) too.
on:
pull_request:
paths:
- 'ansible/ceph/**'
- '.github/workflows/ansible-ceph-validate.yml'
push:
branches: [main]
paths:
- 'ansible/ceph/**'
- '.github/workflows/ansible-ceph-validate.yml'
# Read-only: checkout + tool downloads only. No writes through the token, no
# deploy credentials, no 1Password.
permissions:
contents: read
concurrency:
group: ansible-ceph-validate-${{ github.ref }}
cancel-in-progress: true
jobs:
validate:
name: Validate ansible/ceph (lint + syntax)
runs-on: ubuntu-latest
timeout-minutes: 15
defaults:
run:
working-directory: ansible/ceph
steps:
- uses: actions/checkout@1af3b93b6815bc44a9784bd300feb67ff0d1eeb3 # v6.0.0
with:
persist-credentials: false
- name: Setup Mise
uses: immich-app/devtools/actions/use-mise@cd24790a7f5f6439ac32cc94f5523cb2de8bfa8c # use-mise-action-v1.1.0
env:
# mise downloads tools from GitHub releases; anonymous requests hit
# API rate limits.
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
# Trust the nested ansible/ceph/.mise.toml and install its pinned Python
# (3.12). The use-mise action installs root tools from the repo root; the
# ceph config is only picked up here, where cwd is ansible/ceph.
- name: Trust and install mise tools
run: |
mise trust --all
mise install
# Create the project .venv exactly as `mise run setup` does: pip install
# requirements.txt (pins ansible-core, ansible-lint, yamllint) and the
# ansible-galaxy collections (ansible.posix, community.general) that
# ansible-lint / syntax-check need to resolve modules. PyPI + Galaxy only;
# no secrets.
- name: Install ansible tooling
run: mise run setup
# yamllint + ansible-lint + shellcheck (scripts/*.sh). None need a live
# inventory or secrets; shellcheck ships on ubuntu-latest runners.
- name: Lint (yamllint + ansible-lint + shellcheck)
run: mise run lint
# ansible-playbook --syntax-check needs *an* inventory, but the real
# inventory.ini is TF-generated and gitignored, so it does not exist in
# CI. A syntax-check only parses YAML/role structure, so any valid
# inventory works: write a throwaway localhost inventory and point CEPH_ENV
# at it (the ceph:check task honours an exported CEPH_ENV). This never
# connects to a host.
- name: Syntax-check playbooks
env:
CEPH_ENV: /tmp/ci-inventory.ini
run: |
printf '[all]\nlocalhost ansible_connection=local\n' > /tmp/ci-inventory.ini
mise run check
# Byte-compile the helper scripts so a broken generator is caught here,
# not mid-provision. (shellcheck above covers the *.sh scripts.)
- name: Compile Python scripts
run: mise exec -- python -m py_compile scripts/*.py
+260 -53
View File
@@ -28,13 +28,19 @@ name: Infra (Terraform)
# advertises the node subnets the overlay-joining jobs reach; prod/global must
# precede prod/htz-fsn1/netbird (a real terragrunt dependency).
#
# The two post-apply Ansible converges (ceph RGW users on the ceph stack; mgmt
# hosts on the fabric stack) are gated independently of their TF apply: on push
# they run only when their own surface changed (ceph: ansible/ceph/**; mgmt:
# ansible/mgmt/** + tf/render/ansible-mgmt/** + .mise/tasks/mgmt/**), and on
# workflow_dispatch only when the run_ceph_ansible / run_mgmt_ansible toggles are
# set (both default off). A pure-Terraform change to either stack thus applies
# without reconverging the nodes — and the ceph overlay join is skipped with it.
# The two post-apply Ansible converges are gated independently of their TF apply.
# They differ by blast radius:
# - ceph (full baseline -> tune -> deploy -> harden against a LIVE bare-metal
# cluster) is DELIBERATE: it runs ONLY on workflow_dispatch with
# run_ceph_ansible=true, targeting the single cluster named by the
# ceph_cluster input (spice = prod/htz-fsn1, sietch = staging/austin). A push
# NEVER converges a ceph cluster, so merging a ceph change (prod- or
# staging-scoped) can never auto-run an Ansible converge on a live cluster.
# - mgmt (mgmt hosts on the fabric stack) still auto-runs on push when its own
# surface changed (ansible/mgmt/** + tf/render/ansible-mgmt/** +
# .mise/tasks/mgmt/**), and on workflow_dispatch when run_mgmt_ansible is set.
# A pure-Terraform change to either stack thus applies without reconverging the
# nodes, and the matching overlay join is skipped with it.
#
# The fabric (switch) apply is gated the same way: on push it applies only when the
# fabric surface changed (tf/deployment/prod/htz-fsn1/fabric/**, the fabric shared
@@ -45,13 +51,26 @@ name: Infra (Terraform)
#
# Environment gates rekey to <partition>-<region> (one per stack's region):
# staging-austin, staging-global, prod-global, prod-htz-fsn1. Each matrix apply
# entry references its own gate, so an unprovisioned Environment hangs the apply.
# entry references its own gate. GitHub does NOT hang on an unprovisioned
# Environment -- it auto-creates it on first reference with ZERO protection
# rules, so a brand-new <partition>-<region> would apply with -auto-approve and
# no required reviewer (fail OPEN). The intended fix has TWO parts, currently PARKED:
# 1. The Environments are managed as code (tf/deployment/_meta/github, the
# integrations/github provider: github_repository_environment + required
# reviewers) so every gate exists WITH reviewers by construction.
# 2. The `discover` job FAILS CLOSED (a "required reviewers" step) on apply-bearing
# events, refusing to proceed unless each Environment carries a reviewer.
# BOTH are DISABLED for now (the _meta/github stack is stubbed and the discover step
# is `if: false`): enforcing them requires a one-time bootstrap (GH_ENV_ADMIN_TOKEN
# secret + the four Environments created with reviewers) that would otherwise break
# the first push-to-main. Until that blast radius is redesigned, apply gates behave
# as GitHub's default (auto-create UNPROTECTED). Tracked for follow-up.
#
# Prerequisites (provisioned out-of-band):
# - Repo secrets: OP_TF_YUCCA_STAGING_ENV (+ _WRITE); OP_TF_YUCCA_PROD_ENV
# (read) + OP_TF_YUCCA_PROD_ENV_WRITE (netbird apply / write escalation source).
# - GitHub Environments with required reviewers: staging-austin, staging-global,
# prod-global, prod-htz-fsn1.
# - GitHub Environments with required reviewers (staging-austin, staging-global,
# prod-global, prod-htz-fsn1): DEFERRED -- the approval gate is parked (above).
# - BOOTSTRAP — the netbird stacks applied ONCE out-of-band so the CI/mgmt setup
# keys exist in 1P before anything tries to connect (CI can't mint them itself:
# the apply that mints them is gated behind the plan that needs them). E.g.:
@@ -80,9 +99,16 @@ on:
workflow_dispatch:
inputs:
run_ceph_ansible:
description: 'Run the Ceph Ansible converge (RGW users) after the ceph apply'
description: 'Run the Ceph Ansible converge (full pipeline) against ceph_cluster'
type: boolean
default: false
ceph_cluster:
description: 'Ceph cluster the converge targets (only used when run_ceph_ansible is true)'
type: choice
options:
- sietch
- spice
default: sietch
run_mgmt_ansible:
description: 'Run the mgmt Ansible converge after the fabric apply'
type: boolean
@@ -93,7 +119,10 @@ on:
default: false
# Serialize: the OVH S3 backend has no state locking (single-operator model),
# so never let two infra runs apply concurrently.
# so never let two infra runs apply concurrently. This workflow-wide concurrency
# group IS the single-operator assertion that stands in for a backend lock: at
# most one infra run holds it at a time, and cancel-in-progress:false lets a
# running apply finish rather than being interrupted mid-state-write.
concurrency:
group: ${{ github.workflow }}
cancel-in-progress: false
@@ -108,14 +137,15 @@ jobs:
# Skip on fork PRs (no access to secrets / the overlay anyway).
if: github.event_name != 'pull_request' || github.event.pull_request.head.repo.full_name == github.repository
runs-on: ubuntu-latest
timeout-minutes: 10
outputs:
staging: ${{ steps.filter.outputs.staging }}
prod: ${{ steps.filter.outputs.prod }}
dev: ${{ steps.filter.outputs.dev }}
shared: ${{ steps.filter.outputs.shared }}
# Ansible-converge gates: isolate the Ansible surfaces so a pure-Terraform
# change to the ceph/fabric stack applies without reconverging the nodes.
ansible_ceph: ${{ steps.filter.outputs.ansible_ceph }}
# Mgmt Ansible-converge gate: isolate the mgmt Ansible surface so a
# pure-Terraform change to the fabric stack applies without reconverging the
# mgmt hosts. (The ceph converge is workflow_dispatch-only -- no push gate.)
ansible_mgmt: ${{ steps.filter.outputs.ansible_mgmt }}
# Fabric-apply gate: the fabric (switch) stack applies only when its own
# surface changed — not on every prod change.
@@ -130,10 +160,15 @@ jobs:
filters: |
staging:
- 'tf/deployment/staging/**'
- 'ansible/ceph/**'
# Only staging's own ceph inventory activates the staging matrix.
# A prod ceph change (or a shared ceph role/play) must NEVER open
# the staging partition -- see DEFECT 3. The ceph Ansible converge
# itself is workflow_dispatch-only regardless (no push converge).
- 'ansible/ceph/inventories/staging-**'
prod:
- 'tf/deployment/prod/**'
- 'ansible/mgmt/**'
- 'ansible/ceph/inventories/prod-**'
- 'tf/render/**'
- 'tf/providers/**'
- '.mise/tasks/infra/**'
@@ -150,11 +185,11 @@ jobs:
- '.mise/config.toml'
- '.github/actions/netbird-connect/**'
- '.github/workflows/infra.yml'
# ── Ansible-converge gates (stack-step level, not partition level) ──
# These don't open/close a partition's matrix; the apply job uses them
# to decide whether to run the post-apply Ansible converge for its stack.
ansible_ceph:
- 'ansible/ceph/**'
# -- Mgmt Ansible-converge gate (stack-step level, not partition) --
# Does not open/close a partition's matrix; the apply job uses it to
# decide whether to run the post-apply mgmt converge for the fabric
# stack. (There is no ceph equivalent: the ceph converge is
# workflow_dispatch-only, so it needs no push paths-filter.)
ansible_mgmt:
- 'ansible/mgmt/**'
- 'tf/render/ansible-mgmt/**'
@@ -184,6 +219,9 @@ jobs:
|| needs.changes.outputs.shared == 'true'
|| github.event_name == 'workflow_dispatch'
runs-on: ubuntu-latest
timeout-minutes: 10
permissions:
contents: read
outputs:
matrix: ${{ steps.gen.outputs.matrix }}
has_stacks: ${{ steps.gen.outputs.has_stacks }}
@@ -198,13 +236,39 @@ jobs:
PROD: ${{ needs.changes.outputs.prod }}
SHARED: ${{ needs.changes.outputs.shared }}
DISPATCH: ${{ github.event_name == 'workflow_dispatch' }}
# Dispatch scoping: a manual run only touches the partition(s) its
# toggles actually act on, instead of forcing BOTH partitions active.
DISPATCH_CEPH: ${{ inputs.run_ceph_ansible }}
DISPATCH_MGMT: ${{ inputs.run_mgmt_ansible }}
DISPATCH_FABRIC: ${{ inputs.run_fabric }}
CEPH_CLUSTER: ${{ inputs.ceph_cluster }}
run: |
set -euo pipefail
# Active partitions: own filter OR a shared/manual force. dev is
# local-only and never emitted (no remote state / no CI SA).
# Active partitions: own filter OR a shared force. dev is local-only and
# never emitted (no remote state / no CI SA).
declare -A active=()
if [ "$SHARED" = "true" ] || [ "$DISPATCH" = "true" ]; then
if [ "$DISPATCH" = "true" ]; then
# Manual run: scope to the partition each requested action touches.
# - ceph converge -> the partition owning the chosen cluster
# (spice = prod, sietch = staging)
# - mgmt converge / fabric apply -> prod (htz-fsn1 only)
# A BARE dispatch (no toggle set) is a plan-only preview: emit both
# partitions so the operator sees every diff, but the apply job is
# skipped entirely (see its `if:`), so nothing applies unreviewed.
if [ "$DISPATCH_CEPH" = "true" ]; then
case "$CEPH_CLUSTER" in
spice) active[prod]=1 ;;
sietch) active[staging]=1 ;;
*) echo "unknown ceph_cluster: '$CEPH_CLUSTER'" >&2; exit 1 ;;
esac
fi
[ "$DISPATCH_MGMT" = "true" ] && active[prod]=1
[ "$DISPATCH_FABRIC" = "true" ] && active[prod]=1
if [ "${#active[@]}" -eq 0 ]; then
active[staging]=1; active[prod]=1 # bare dispatch = plan-only preview
fi
elif [ "$SHARED" = "true" ]; then
active[staging]=1; active[prod]=1
else
[ "$STAGING" = "true" ] && active[staging]=1
@@ -251,12 +315,59 @@ jobs:
echo "has_stacks=$has" >> "$GITHUB_OUTPUT"
echo "Discovered stacks:"; echo "$include" | jq -r '.[] | " [\(.order)] \(.partition)@\(.region)/\(.stack) (\(.dir))"'
# -- Approval-gate assertion: PARKED (disabled) ----------------------
# This step asserted that every target Environment carries a required
# reviewer and failed the run otherwise (fail-closed four-eyes on prod
# applies). It is DISABLED pending a blast-radius redesign: as written it
# blocks the FIRST push-to-main until BOTH (a) a GH_ENV_ADMIN_TOKEN secret
# exists (the built-in GITHUB_TOKEN cannot read protection rules) AND (b) the
# four Environments are created WITH reviewers via the parked
# tf/deployment/_meta/github stack. Neither is bootstrapped yet, so enforcing
# here would break every prod/staging apply on merge. While parked, an
# unprovisioned Environment auto-creates UNPROTECTED (fail OPEN) -- the
# pre-hardening status quo, tracked for follow-up. The full logic is kept
# below for revival; re-enable once the gate can be provisioned without a
# hard cutover. See tf/deployment/_meta/github/terragrunt.hcl.disabled.
- name: Fail closed unless every target Environment has required reviewers
# PARKED: re-enable by restoring this guard (blast-radius redesign pending):
# steps.gen.outputs.has_stacks == 'true'
# && (github.event_name == 'push' || github.event_name == 'workflow_dispatch')
if: false
env:
GH_TOKEN: ${{ secrets.GH_ENV_ADMIN_TOKEN || secrets.GITHUB_TOKEN }}
REPO: ${{ github.repository }}
MATRIX: ${{ steps.gen.outputs.matrix }}
run: |
set -euo pipefail
envs=$(jq -r '.include[] | "\(.partition)-\(.region)"' <<<"$MATRIX" | sort -u)
[ -n "$envs" ] || { echo "No environments to verify."; exit 0; }
fail=0
while IFS= read -r envname; do
[ -n "$envname" ] || continue
if ! resp=$(gh api "repos/$REPO/environments/$envname" 2>/dev/null); then
echo "FAIL-CLOSED: environment '$envname' is missing or unreadable -> treated as unprotected." >&2
fail=1; continue
fi
n=$(jq '[.protection_rules[]? | select(.type=="required_reviewers") | .reviewers[]?] | length' <<<"$resp")
if [ "${n:-0}" -lt 1 ]; then
echo "FAIL-CLOSED: environment '$envname' has no required reviewers (approval gate would fail open)." >&2
fail=1
else
echo "OK: environment '$envname' has ${n} required reviewer(s)."
fi
done <<<"$envs"
[ "$fail" -eq 0 ] || {
echo "One or more target Environments lack a required reviewer; refusing to plan/apply (fail closed)." >&2
exit 1
}
# ── Plan every changed stack (parallel; read-only) ───────────────────────────
plan:
name: Plan ${{ matrix.partition }}@${{ matrix.region }}/${{ matrix.stack }}
needs: [changes, discover]
if: needs.discover.outputs.has_stacks == 'true'
runs-on: ubuntu-latest
timeout-minutes: 20
strategy:
fail-fast: false
matrix: ${{ fromJSON(needs.discover.outputs.matrix) }}
@@ -272,10 +383,22 @@ jobs:
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
# Filesystem-safe artifact id for this stack (stack may nest '/').
- id: planid
name: Compute plan artifact id
env:
DIR: ${{ matrix.dir }}
run: echo "id=$(printf '%s' "$DIR" | tr '/' '-')" >> "$GITHUB_OUTPUT"
- name: Set up mise (go + opentofu + terragrunt)
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
- name: Install 1Password CLI
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
with:
# Pin, don't resolve "latest": the version-check endpoint 1P hits to
# resolve "latest" intermittently returns empty JSON (all parallel
# plan/apply jobs fail at "Getting latest version number"). A pinned
# version installs directly, skipping that call.
version: '2.31.1'
# Guardrail: prod Flux must sync `main` once we're ON main. flux_git_ref
# defaults to a feature branch during bring-up (deliberate, documented
@@ -319,29 +442,63 @@ jobs:
# Prod fabric goes through the mise task (builds the hetzner provider, renders
# the NETCONF key; junos comes from the registry). Every other stack is a
# plain registry-provider terragrunt plan.
# Persist the diff to an ABSOLUTE path (terragrunt runs tofu inside a
# .terragrunt-cache dir, so a relative -out would be unreachable). The gated
# apply downloads this exact file and applies it -- the human approves the
# DIFF, not a plan-job log. TG_PLAN threads the path into the mise task too.
- name: Terragrunt plan (fabric)
if: matrix.stack == 'fabric'
env:
TG_PLAN: ${{ github.workspace }}/tfplan.bin
run: mise run infra:plan -- --non-interactive
- name: Terragrunt plan
if: matrix.stack != 'fabric'
env:
STACK_DIR: ${{ matrix.dir }}
TG_PLAN: ${{ github.workspace }}/tfplan.bin
run: >-
tf/op-run.sh terragrunt
--working-dir "tf/deployment/$STACK_DIR"
--non-interactive plan
--non-interactive plan -out="$TG_PLAN"
# Hand the saved plan to the apply job. NOTE: a tofu plan file can embed
# sensitive values in the clear, so keep retention short and rely on the
# repo's private-artifact scoping. if-no-files-found:error makes a missing
# plan fail loudly rather than silently fall back to a fresh apply.
- name: Upload reviewed plan
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: tfplan-${{ steps.planid.outputs.id }}
path: ${{ github.workspace }}/tfplan.bin
if-no-files-found: error
retention-days: 1
# ── Apply (gated, ordered) ───────────────────────────────────────────────────
apply:
name: Apply ${{ matrix.partition }}@${{ matrix.region }}/${{ matrix.stack }} (gated)
needs: [changes, discover, plan]
# Only on merge to main (or manual dispatch) — never on PRs.
if: (github.event_name == 'push' || github.event_name == 'workflow_dispatch') && needs.discover.outputs.has_stacks == 'true'
# Only on merge to main (or a SCOPED manual dispatch) -- never on PRs. A bare
# workflow_dispatch (no run_ceph_ansible / run_mgmt_ansible / run_fabric) is
# plan-only: the apply job is skipped so a manual "just look" run can never
# apply every emitted stack unreviewed.
if: >-
needs.discover.outputs.has_stacks == 'true'
&& (
github.event_name == 'push'
|| (github.event_name == 'workflow_dispatch'
&& (inputs.run_ceph_ansible || inputs.run_mgmt_ansible || inputs.run_fabric))
)
runs-on: ubuntu-latest
# Ceiling for a worst-case entry (TF apply + up to 90-min ceph converge).
# Finer per-phase bounds are enforced at step level (apply 30 / converge 90).
timeout-minutes: 120
strategy:
fail-fast: false
# Serialize so the sorted `order` is honored: account netbird → site netbird
# → node-touching stacks, and prod/global before prod/htz-fsn1/netbird.
# fail-fast so a failed lower-order apply (e.g. netbird) cancels the
# remaining ordered node-touching applies instead of running against a
# half-provisioned overlay.
fail-fast: true
# Serialize so the sorted `order` is honored: account netbird -> site netbird
# -> node-touching stacks, and prod/global before prod/htz-fsn1/netbird.
max-parallel: 1
matrix: ${{ fromJSON(needs.discover.outputs.matrix) }}
# Per-region approval gate.
@@ -360,32 +517,67 @@ jobs:
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
# Same filesystem-safe id the plan job used to name the artifact.
- id: planid
name: Compute plan artifact id
env:
DIR: ${{ matrix.dir }}
run: echo "id=$(printf '%s' "$DIR" | tr '/' '-')" >> "$GITHUB_OUTPUT"
- name: Set up mise (go + opentofu + terragrunt + ansible)
uses: jdx/mise-action@e6a8b3978addb5a52f2b4cd9d91eafa7f0ab959d # v4.2.0
- name: Install 1Password CLI
uses: 1password/install-cli-action@a5215d3a7f75c1629216c465ea9ab3ab399c4b71 # v4.0.0
with:
# See the plan job's note: pin to skip the flaky "latest" resolution.
version: '2.31.1'
# Decide whether each post-apply Ansible converge runs. On push, gate on the
# paths-filter (did the Ansible surface change?); on manual dispatch, gate on
# the run_*_ansible toggles (default off) since there's no diff to detect.
# Pull the EXACT plan the reviewer approved, downloaded to the same absolute
# path the plan job wrote. `terragrunt apply <planfile>` fails closed if the
# saved plan is stale relative to current state -- so a drifted backend can
# never silently apply a different diff than the one that was reviewed.
- name: Download reviewed plan
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
with:
name: tfplan-${{ steps.planid.outputs.id }}
path: ${{ github.workspace }}
# Decide whether each post-apply Ansible converge runs.
# - ceph: DELIBERATE and blast-radius-guarded. Runs ONLY on
# workflow_dispatch with run_ceph_ansible=true, and ONLY for the ceph
# matrix entry whose partition/region owns the selected ceph_cluster
# (spice -> prod/htz-fsn1, sietch -> staging/austin). A push never sets
# this gate, so merging a ceph change cannot converge a live cluster.
# - mgmt: on push, gate on the paths-filter (did the mgmt surface change?);
# on manual dispatch, gate on the run_mgmt_ansible toggle (default off).
- name: Resolve Ansible-converge gates
id: ansible_gate
env:
DISPATCH: ${{ github.event_name == 'workflow_dispatch' }}
DISPATCH_CEPH: ${{ inputs.run_ceph_ansible }}
DISPATCH_MGMT: ${{ inputs.run_mgmt_ansible }}
CHANGED_CEPH: ${{ needs.changes.outputs.ansible_ceph }}
CEPH_CLUSTER: ${{ inputs.ceph_cluster }}
PARTITION: ${{ matrix.partition }}
REGION: ${{ matrix.region }}
CHANGED_MGMT: ${{ needs.changes.outputs.ansible_mgmt }}
run: |
set -euo pipefail
if [ "$DISPATCH" = "true" ]; then
ceph=$DISPATCH_CEPH; mgmt=$DISPATCH_MGMT
else
ceph=$CHANGED_CEPH; mgmt=$CHANGED_MGMT
# ceph converge: dispatch-only, and pinned to the one cluster/entry.
ceph=false
if [ "$DISPATCH" = "true" ] && [ "$DISPATCH_CEPH" = "true" ]; then
case "$CEPH_CLUSTER" in
spice) want_partition=prod; want_region=htz-fsn1 ;;
sietch) want_partition=staging; want_region=austin ;;
*) echo "unknown ceph_cluster: '$CEPH_CLUSTER'" >&2; exit 1 ;;
esac
if [ "$PARTITION" = "$want_partition" ] && [ "$REGION" = "$want_region" ]; then
ceph=true
fi
fi
# mgmt converge: push -> paths-filter; dispatch -> toggle.
if [ "$DISPATCH" = "true" ]; then mgmt=$DISPATCH_MGMT; else mgmt=$CHANGED_MGMT; fi
echo "ceph=${ceph:-false}" >> "$GITHUB_OUTPUT"
echo "mgmt=${mgmt:-false}" >> "$GITHUB_OUTPUT"
echo "Ansible converge gates → ceph=${ceph:-false} mgmt=${mgmt:-false}"
echo "Ansible converge gates -> ceph=${ceph:-false} (cluster=${CEPH_CLUSTER:-none}) mgmt=${mgmt:-false}"
# Decide whether the fabric (switch) apply runs. Same model as the Ansible
# gates: on push, gate on the paths-filter (did the fabric surface change?);
@@ -428,37 +620,50 @@ jobs:
# Gated: only when the fabric surface changed (push) or run_fabric (dispatch).
- name: Terragrunt apply (fabric)
if: matrix.stack == 'fabric' && steps.fabric_gate.outputs.run == 'true'
run: mise run infra:apply -- --non-interactive -auto-approve
# Everything else: direct terragrunt apply with the partition's write SA.
timeout-minutes: 30
# TG_PLAN routes the reviewed plan into the mise task, which applies that
# file (no fresh re-plan, no -auto-approve) and fails on a stale plan.
env:
TG_PLAN: ${{ github.workspace }}/tfplan.bin
run: mise run infra:apply -- --non-interactive
# Everything else: direct terragrunt apply of the reviewed plan file with
# the partition's write SA (no -auto-approve -- the plan is pre-approved).
- name: Terragrunt apply
if: matrix.stack != 'fabric'
timeout-minutes: 30
env:
STACK_DIR: ${{ matrix.dir }}
TG_PLAN: ${{ github.workspace }}/tfplan.bin
run: >-
tf/op-run.sh terragrunt
--working-dir "tf/deployment/$STACK_DIR"
--non-interactive apply -auto-approve
--non-interactive apply "$TG_PLAN"
# ── Ceph convergence (Ansible) — ceph stack, gated on the converge gate ──
# -- Ceph convergence (Ansible) - ceph stack, workflow_dispatch-only --
# The TF apply above only minted the RGW keys into 1P + the cluster Secret;
# this creates the matching RGW users on the bare-metal cluster. Reuses the
# NetBird overlay + 1Password session already established in this entry.
# Skipped (along with the overlay join above) when ansible/ceph/** is
# unchanged on push, or when run_ceph_ansible is off on manual dispatch.
# this converges the selected bare-metal cluster. Reuses the NetBird overlay
# + 1Password session already established in this entry. The gate fires ONLY
# on workflow_dispatch with run_ceph_ansible=true, and only for the entry
# that owns the chosen ceph_cluster (spice=prod/htz-fsn1, sietch=staging/
# austin) -- a push never converges here. Cluster + region are threaded from
# the dispatch inputs / matrix so this is not hardcoded to sietch/austin.
- name: Render Ansible inventory from the ceph TF state
if: matrix.stack == 'ceph' && steps.ansible_gate.outputs.ceph == 'true'
env:
PARTITION: ${{ matrix.partition }}
run: ansible/ceph/scripts/render-inventories.sh "$PARTITION"
REGION: ${{ matrix.region }}
run: ansible/ceph/scripts/render-inventories.sh "$PARTITION" "$REGION"
- name: Install the ansible-iac SSH key from 1Password
if: matrix.stack == 'ceph' && steps.ansible_gate.outputs.ceph == 'true'
# Pulls op://yucca_tf_<partition>/SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY to
# ~/.ssh/id_ed25519_sietch — the path the rendered inventory references.
# Pulls op://yucca_tf_<partition>/<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY to
# ~/.ssh/id_ed25519_<cluster> -- the path the rendered inventory references
# (spice -> id_ed25519_spice, sietch -> id_ed25519_sietch).
env:
PARTITION: ${{ matrix.partition }}
CEPH_CLUSTER: ${{ inputs.ceph_cluster }}
run: |
mkdir -p ~/.ssh && chmod 700 ~/.ssh
OP_VAULT="yucca_tf_$PARTITION" ansible/ceph/scripts/install-ssh-keys.sh sietch
OP_VAULT="yucca_tf_$PARTITION" ansible/ceph/scripts/install-ssh-keys.sh "$CEPH_CLUSTER"
- name: Provision the ceph Ansible toolchain (venv + collections)
if: matrix.stack == 'ceph' && steps.ansible_gate.outputs.ceph == 'true'
working-directory: ansible/ceph
@@ -466,8 +671,9 @@ jobs:
mise trust
mise install
mise run setup
- name: Deploy Ceph (full pipeline — baseline → tune → deploy → harden)
- name: Deploy Ceph (full pipeline -- baseline -> tune -> deploy -> harden)
if: matrix.stack == 'ceph' && steps.ansible_gate.outputs.ceph == 'true'
timeout-minutes: 90
# Run from ansible/ceph so mise loads ansible/ceph/.mise.toml (where the
# `deploy` task + CEPH_ENV-relative inventory paths live); the root config
# has no `deploy` task.
@@ -476,7 +682,7 @@ jobs:
# over the NetBird overlay, so disable strict host-key checking for this run.
env:
ANSIBLE_HOST_KEY_CHECKING: 'false'
CEPH_ENV: inventories/${{ matrix.partition }}-${{ matrix.region }}/sietch/inventory.ini
CEPH_ENV: inventories/${{ matrix.partition }}-${{ matrix.region }}/${{ inputs.ceph_cluster }}/inventory.ini
run: mise run deploy
# ── Mgmt convergence (Ansible) — fabric stack, gated on the converge gate ─
@@ -488,4 +694,5 @@ jobs:
# overlay join regardless (it touches the switch vme directly).
- name: Ansible converge (mgmt hosts)
if: matrix.stack == 'fabric' && steps.ansible_gate.outputs.mgmt == 'true'
timeout-minutes: 30
run: mise run mgmt:ansible
@@ -0,0 +1,54 @@
name: s3-dns-roster-validate
# Static gate: the prod S3 RGW DNS roster (the A-records for
# s3.prod.fsn1.htz.futo.cloud + the wildcard, in tf/deployment/prod/global/dns)
# is a hand-maintained round-robin over every in-service spice ceph node's fabric
# IP. Nothing links it to the ceph stack, so adding/replacing/removing a node can
# silently blackhole the endpoint (a stale IP that routes nowhere) or drop capacity
# (a live node left out of rotation). Without this gate the mismatch first surfaces
# in production, after the DNS apply.
#
# tf/scripts/check-s3-dns-roster.py reconstructs the expected roster from the ceph
# source of truth (clusters.auto.tfvars in-service set + spice-hosts.yaml
# host_index -> fabric IP) and fails on any divergence. Credential-free and
# stdlib-only: no remote state, no provider, no 1Password, no host connection.
#
# Its own workflow (not a job in infra.yml or ci.yml) so the path filter below is
# workflow-scoped: it fires only when the DNS roster or a ceph roster input
# changes, mirroring ansible-ceph-validate.yml.
on:
pull_request:
paths: &paths
- 'tf/deployment/prod/global/dns/records.auto.tfvars'
- 'tf/deployment/prod/htz-fsn1/ceph/clusters.auto.tfvars'
- 'tf/deployment/prod/htz-fsn1/ceph/spice-hosts.yaml'
- 'tf/scripts/check-s3-dns-roster.py'
- '.github/workflows/s3-dns-roster-validate.yml'
push:
branches: [main]
paths: *paths
# Read-only: checkout only. No writes through the token, no deploy credentials,
# no 1Password.
permissions:
contents: read
concurrency:
group: s3-dns-roster-validate-${{ github.ref }}
cancel-in-progress: true
jobs:
validate:
name: Check S3 DNS roster vs ceph node roster
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- uses: actions/checkout@1af3b93b6815bc44a9784bd300feb67ff0d1eeb3 # v6.0.0
with:
persist-credentials: false
# Stdlib-only Python; ubuntu-latest ships python3. No mise/tool install
# needed, so this gate stays fast and dependency-free.
- name: Roster drift check
run: tf/scripts/check-s3-dns-roster.py
+4
View File
@@ -97,6 +97,10 @@ run = "tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/stagi
description = "Format terraform + terragrunt files recursively"
run = "tofu fmt -recursive tf/ && terragrunt hcl format --working-dir tf/"
[tasks."tf:check-dns-roster"]
description = "Drift-check the prod S3 RGW DNS roster against the ceph node roster (credential-free, no state)"
run = "tf/scripts/check-s3-dns-roster.py"
[env]
NODE_ENV = "development"
LOG_LEVEL = "debug"
+14 -1
View File
@@ -34,10 +34,23 @@ SITE="${SITE:-htz-fsn1}"
# provider rejects having both set ("service_account_token and account are set").
unset OP_ACCOUNT
STACK_DIR="tf/deployment/prod/$SITE/fabric"
# Bind the apply to the reviewed plan. When CI passes TG_PLAN pointing at the
# saved plan artifact, apply THAT file (no fresh re-plan, no -auto-approve): tofu
# fails closed if the saved plan is stale relative to current state. Strip any
# -auto-approve the caller passed (meaningless with a plan file). Local/manual
# runs leave TG_PLAN unset and keep the interactive/`-auto-approve` behaviour.
APPLY_ARGS=("$@")
if [ -n "${TG_PLAN:-}" ]; then
[ -f "$TG_PLAN" ] || { echo "infra:apply: TG_PLAN set but plan file missing: $TG_PLAN" >&2; exit 1; }
filtered=(); for a in "${APPLY_ARGS[@]}"; do [ "$a" = "-auto-approve" ] && continue; filtered+=("$a"); done
APPLY_ARGS=("${filtered[@]}" "$TG_PLAN")
fi
# jeremmfr/junos is concurrency-safe (per-resource CRUD); reads/refresh run in
# parallel and commits serialize on the per-device config lock (handled by the
# provider). The device NETCONF connection-limit is raised to 250.
OP_ENV_FILE=tf/.env.prod "$ROOT/tf/op-run.sh" terragrunt --working-dir "$STACK_DIR" apply -parallelism=4 "$@"
OP_ENV_FILE=tf/.env.prod "$ROOT/tf/op-run.sh" terragrunt --working-dir "$STACK_DIR" apply -parallelism=4 "${APPLY_ARGS[@]}"
# Confirm the dangling `commit confirmed` jeremmfr leaves on the last commit per
# device: its per-resource "confirm" is `commit check`, which does NOT cancel the
+9 -1
View File
@@ -33,5 +33,13 @@ SITE="${SITE:-htz-fsn1}"
# With a service-account token in the env, drop OP_ACCOUNT — the onepassword
# provider rejects having both set ("service_account_token and account are set").
unset OP_ACCOUNT
# When CI sets TG_PLAN, persist the diff to that (absolute) path so the gated
# apply can consume the EXACT reviewed plan (terragrunt runs tofu inside a
# .terragrunt-cache dir, so a relative -out would be unreachable -- pass an
# absolute path). Local `tf:plan` leaves TG_PLAN unset -> a plain read-only plan.
PLAN_OUT=()
[ -n "${TG_PLAN:-}" ] && PLAN_OUT=(-out="$TG_PLAN")
# jeremmfr/junos is concurrency-safe; the device NETCONF connection-limit is 250.
OP_ENV_FILE=tf/.env.prod "$ROOT/tf/op-run.sh" terragrunt --working-dir "tf/deployment/prod/$SITE/fabric" plan -parallelism=4 "$@"
OP_ENV_FILE=tf/.env.prod "$ROOT/tf/op-run.sh" terragrunt --working-dir "tf/deployment/prod/$SITE/fabric" plan -parallelism=4 "${PLAN_OUT[@]}" "$@"
+15
View File
@@ -155,6 +155,10 @@ run = "scripts/ansible-play.sh backup-config.yml"
description = "Snapshot bootstrap secrets (RGW TLS, admin keyring) to 1P (DR belt)"
run = "scripts/ansible-play.sh post-deploy-capture.yml"
[tasks."backup-timer"]
description = "Install + enable the scheduled cluster-state backup timer on the bootstrap node"
run = "scripts/ansible-play.sh backup-ceph.yml"
[tasks."rotate-certs"]
description = "Rotate RGW TLS cert (regenerate self-signed, restart RGW daemons)"
run = "scripts/ansible-play.sh rotate-certs.yml"
@@ -170,3 +174,14 @@ run = "scripts/ansible-play.sh hardware-inventory.yml"
[tasks."migrate-networkd"]
description = "One-shot ifupdown→networkd migration (rolling, serial=1, noout-gated)"
run = "scripts/ansible-play.sh migrate-networkd.yml"
[tasks.reprovision]
description = "Hetzner installimage reprovision (rescue -> reset -> installimage -> verify). DESTRUCTIVE - canary first."
run = """
#!/usr/bin/env bash
set -euo pipefail
# HETZNER_ROBOT_* come from tf/.env.prod via op run; the same op session also
# satisfies ansible-play.sh's inner op inject of secrets.yml.tpl. Set CEPH_ENV first.
OP_ACCOUNT=team-futo op run --env-file=../../tf/.env.prod -- \\
scripts/ansible-play.sh reprovision.yml "$@"
"""
+1
View File
@@ -326,6 +326,7 @@ scripts/ansible-play.sh deploy-ceph.yml --tags bootstrap
| `drift` | `mise run drift` | Detect configuration drift (via ansible-play.sh) |
| `deploy` | `mise run deploy` | Full deploy pipeline (via ansible-play.sh) |
| `backup` | `mise run backup` | Export cluster config for DR (via ansible-play.sh) |
| `backup-timer` | `mise run backup-timer` | Install the scheduled on-node cluster-state backup timer (bootstrap node) |
| `capture` | `mise run capture` | Snapshot RGW TLS + admin keyring to 1P for disaster recovery |
| `bench` | `mise run bench` | S3 benchmark (RGW round-trip) |
| `bench-rados` | `mise run bench-rados` | RADOS bench (raw cluster I/O) |
+67
View File
@@ -0,0 +1,67 @@
---
# Manual, rescue-first "add a node" path: install + chroot + verify on a node the
# operator has ALREADY put into the Hetzner rescue system by hand. This is the
# default way to (re)image a spice node -- the Robot API (register keys / arm
# rescue / hw reset) is deliberately OUT of the loop (reprovision_use_robot=false).
#
# WHY manual: the Robot rescue/reset step is the unreliable part (nodes that arm
# rescue but never boot it) AND the highest blast radius (an errant /reset on a
# LIVE spice node once the cluster is in production). Taking it out means nothing
# Ansible does here can flip a running node into rescue: installimage exists only
# in the rescue ramdisk, so wait_rescue hard-refuses any node a human did not
# deliberately rescue. The install/chroot/verify automation -- the reliable part --
# still runs untouched.
#
# For the rare remote "nuke a dead node" case, use reprovision.yml (Robot path).
#
# OPERATOR PRECONDITION (do this in the Hetzner Robot portal, per node):
# 1. Enable the Linux rescue system, authorizing the spice-ansible-iac key
# (and operator break-glass keys). 2. Trigger a hardware reset. 3. Wait until
# the node is reachable in rescue over its WAN IP. THEN run this play.
# Run it promptly after rescue boots (wait_rescue rejects a stale rescue whose
# uptime exceeds reprovision_rescue_max_uptime; raise it with -e if needed).
#
# No secrets and no Robot creds are needed -- run with plain ansible-playbook:
# Canary (one node, full wipe + verify):
# ansible-playbook -i inventories/prod-htz-fsn1/spice/inventory.ini add-node.yml \
# -e confirm_wipe=true --limit spice-ceph-philip
# Fan-out (only nodes you have already rescued; low serial, canary-gated):
# ansible-playbook -i inventories/prod-htz-fsn1/spice/inventory.ini add-node.yml \
# -e confirm_wipe=true -e allow_fanout=true -e reprovision_serial=2 \
# --limit 'rescue:!spice-ceph-philip'
#
# ALWAYS DESTRUCTIVE to the NVMe OS disks (installimage). Requires confirm_wipe=true.
- name: Add node (manual rescue -> installimage -> installed OS)
hosts: ceph_nodes
gather_facts: false
become: false
serial: "{{ reprovision_serial | default(1) }}"
max_fail_percentage: 0 # any node fails -> halt the batch (no further wipes)
vars:
reprovision_use_robot: false # no Robot API: rescue is armed by hand, not here
# With use_robot=false, wait_rescue skips the this-boot uptime guard entirely
# (it only makes sense when Ansible triggered the reset). A node may sit in
# rescue for weeks; the installimage-ramdisk check is the real in-rescue proof.
pre_tasks:
# run_once + through the role so defaults load and --limit can't skip it. With
# use_robot=false this asserts keys + the fan-out gate only (no Robot calls).
- name: Preflight (once) # noqa: run-once[task]
ansible.builtin.include_role:
name: reprovision_hetzner
tasks_from: preflight
run_once: true
roles:
- role: reprovision_hetzner
- name: Verify added nodes after reboot
hosts: ceph_nodes
gather_facts: false
become: false
vars:
reprovision_use_robot: false
tasks:
- name: Verify installed OS marker
ansible.builtin.include_role:
name: reprovision_hetzner
tasks_from: verify
+29
View File
@@ -0,0 +1,29 @@
---
# backup-ceph.yml - install (and enable) SCHEDULED Ceph cluster-state backups.
#
# The deep review flagged that config/topology backups were manual, unscheduled,
# and controller-local only. This play installs an on-node capture script plus a
# systemd timer on the bootstrap node so recoverable cluster state (fsid, config
# dump, monmap/osdmap/crushmap, osd tree, orch ls/host ls, RGW realm/zonegroup/
# zone) is snapshot to a timestamped tarball daily under a root-only local dir,
# pruned by retention, and ready to ship offsite once an operator wires a target.
#
# Deliberate/manual: run it to install the timer; the timer then runs unattended.
# It complements (does not replace):
# - mise run capture (post-deploy-capture.yml) -- secrets -> 1Password
# - mise run backup (backup-config.yml) -- one-shot controller-local export
#
# Usage:
# scripts/ansible-play.sh backup-ceph.yml
# mise run backup-timer
#
# Opt a cluster out with `ceph_backup_enabled: false` in its group_vars.
# Set an offsite target with `ceph_backup_offsite_dest` (empty by default).
- name: Install scheduled Ceph cluster-state backups
hosts: ceph_bootstrap
become: true
gather_facts: false
roles:
- ceph_backup
+14
View File
@@ -16,6 +16,20 @@
hosts: ceph_nodes
become: true
gather_facts: false
# Blast-radius control: converge in small batches (serial) and halt the whole
# roll if any node in a batch fails (max_fail_percentage:0) or the batch
# degrades the cluster (the ceph-health gate). deploy-ceph stays big-bang (its
# bootstrap needs all nodes at once); this governs the day-2 converge plays.
serial: "{{ ceph_converge_serial | default('10%') }}"
max_fail_percentage: 0
pre_tasks:
- name: Ceph-health gate (pre-batch)
ansible.builtin.import_tasks: tasks/health_gate.yml
roles:
- baseline
post_tasks:
- name: Ceph-health gate (post-batch)
ansible.builtin.import_tasks: tasks/health_gate.yml
+41
View File
@@ -18,5 +18,46 @@
become: true
gather_facts: true
pre_tasks:
# Fail fast if the inventory was not rendered from TF. The ceph_deploy role
# gates every phase on the ceph_bootstrap/ceph_mon/ceph_join groups and indexes
# groups['ceph_bootstrap'][0]; against a hand-written stopgap inventory that
# lacks them the role would error mid-run or, worse, fall through to placing a
# MON on every host. render-inventories.sh emits these groups from TF state
# (bootstrap = exactly one host, mon = an odd quorum including the bootstrap
# node) - refuse to deploy against anything else.
- name: Assert the TF-rendered placement groups are present and sane # noqa: run-once[task]
ansible.builtin.assert:
that:
- (groups['ceph_bootstrap'] | default([])) | length == 1
- (groups['ceph_mon'] | default([])) | length >= 1
- (groups['ceph_mon'] | default([]) | length) is odd
- ((groups['ceph_bootstrap'] | default([])) | intersect(groups['ceph_mon'] | default([])) | length) == 1
fail_msg: >-
Inventory is missing or has bad TF-rendered placement groups
(need ceph_bootstrap = exactly 1 host that is also in ceph_mon, and
ceph_mon = a non-empty ODD quorum). Render it first:
scripts/render-inventories.sh <partition> <region>. Refusing to deploy.
success_msg: >-
Placement OK: bootstrap={{ groups['ceph_bootstrap'] | default([]) | length }},
mon={{ groups['ceph_mon'] | default([]) | length }},
join={{ groups['ceph_join'] | default([]) | length }}.
run_once: true
# Per host: the address the mon/OSD/RGW bind (ceph_service_ip) must be live on
# a local interface before bootstrap. For spice this is the fabric IP on
# bond0.120; a node whose VLAN sub-interface did not come back after a reboot
# would otherwise fail deep inside cephadm bootstrap/join. gather_facts is on,
# so ansible_all_ipv4_addresses is populated.
- name: Assert the Ceph service IP is live on this node (fabric/bond up)
ansible.builtin.assert:
that:
- ceph_service_ip in ansible_all_ipv4_addresses
fail_msg: >-
ceph_service_ip {{ ceph_service_ip }} is not on any local interface on
{{ inventory_hostname }} - the Ceph bind network (spice: bond0.120) is
down on this node. Bring the fabric up before deploying it.
success_msg: "Ceph service IP {{ ceph_service_ip }} is live on {{ inventory_hostname }}."
roles:
- ceph_deploy
+24
View File
@@ -62,6 +62,30 @@ but a full-node loss degrades a large share of PGs. For host-level
failure domain you need **11+ nodes minimum**. Production (Yucca) will
want this; dev can tolerate the weaker guarantee.
## Durability and DR ceiling
The spice (production, htz-fsn1) backup-of-record stores objects in an EC 16+4
data pool with `failure_domain=host`. min_size is pinned to k+1 = 17 by the RGW
role. What that pool tolerates:
- **Reads:** survive up to m = 4 simultaneous host (or OSD) losses. With 4
chunks gone the remaining 16 = k chunks still reconstruct every object.
- **Writes:** survive up to m - 1 = 3 simultaneous host losses. At the k+1
floor a PG stays writable only while at least 17 shards are up; the 4th host
loss drops a PG to k = 16 shards, which still serves reads but blocks writes
until recovery restores a 17th shard. min_size is never set below k+1 --
permitting writes at k shards would risk data loss if another shard were lost
mid-recovery.
This is a **single-site** guarantee with **no rack diversity**: every node sits
in one datacenter (FSN1) and the failure domain is host, not rack. A rack, row,
PDU, or switch fault that takes down more than m hosts at once, or a whole-
datacenter loss (power, network, fire, flood), exceeds this ceiling -- there is
no second site and no off-region copy. The pool protects against disk and node
failure, not site failure. Off-site DR (a second region, or an external copy of
the backup-of-record) is out of scope for this cluster and would need a separate
replication path.
## block.db sizing
Rule of thumb: block.db ~ 4% of OSD data size.
+18 -1
View File
@@ -40,7 +40,24 @@ Backups are written to `backups/<timestamp>/` on the Ansible controller
### Backup schedule
No automated schedule is configured. Run manually:
**Scheduled (automated):** `mise run backup-timer` (or
`scripts/ansible-play.sh backup-ceph.yml`) installs an on-node capture
script (`/usr/local/sbin/ceph-backup.sh`) plus a systemd timer on the
bootstrap node. The timer runs daily (with up to 1h of jitter) and writes a
timestamped tarball to `/var/backups/ceph/` (root-only, `0700`), pruning
tarballs older than `ceph_backup_retention_days` (default 14). Each tarball
holds the recoverable cluster state: `fsid`, `ceph config dump`, monmap,
osdmap, crushmap (binary + decompiled), `ceph osd tree`, `ceph orch ls` /
`host ls`, and the RGW realm/zonegroup/zone. It does **not** contain secret
keyrings. Opt a cluster out with `ceph_backup_enabled: false` in its
group_vars.
**Offsite:** left as an operator decision. Set `ceph_backup_offsite_dest`
(empty by default) to an rsync target (`user@host:/path/`) or S3 URI
(`s3://bucket/prefix`) and the script ships each tarball after capture.
**Manual (`mise run backup`):** the controller-local export below is still
available for ad-hoc snapshots. Run it manually:
- Before any cluster topology change (add/remove node or OSD)
- Before Ceph version upgrades
+129
View File
@@ -0,0 +1,129 @@
# Runbook: Upgrade Ceph (cephadm-orchestrated) and Roll Back
**When:** moving the cluster to a new patch release on the pinned Tentacle
20.2.x train (e.g. 20.2.1 -> 20.2.2), or rolling back a bad upgrade.
**Time estimate:** 20-60 min for a small cluster; scales with daemon/OSD count
(cephadm rolls one daemon type at a time and waits for health between steps).
**Playbook:** `upgrade-ceph.yml` (+ `tasks/ceph_upgrade.yml`). It is a
deliberate, manually-invoked day-2 op and is **not** in `site.yml` -- converge
never triggers an upgrade.
---
## What the playbook guarantees
1. **Preflight gate** (refuses otherwise): cluster is `HEALTH_OK`, no upgrade
already in progress, and every OSD is up + in.
2. **Rollback anchor recorded**: it prints the image(s) the daemons are running
*now* before touching anything -- copy this for step-4 rollback.
3. **Explicit target required**: no default image/version is baked in.
4. **Bounded poll**: watches `ceph orch upgrade status`; fails loudly if
cephadm pauses/stalls the upgrade (it auto-pauses on a failed step).
5. **Post-checks**: asserts `HEALTH_OK` and that all daemons converged on one
version, then prints a summary.
6. **Idempotent**: if the cluster already runs the target image and is
version-converged, it is a clean no-op.
## Upgrade
```bash
# Preferred: pin the exact image (digest-pinnable, and what rollback consumes).
scripts/ansible-play.sh upgrade-ceph.yml \
-e ceph_upgrade_target_image=quay.io/ceph/ceph:v20.2.2
# Or by version (cephadm resolves the image):
scripts/ansible-play.sh upgrade-ceph.yml -e ceph_upgrade_version=20.2.2
```
Optional toggles:
```bash
# Proceed on a benign HEALTH_WARN (review `ceph health detail` first):
-e ceph_upgrade_allow_health_warn=true
# Widen the poll bound (default 120 * 30s = 60 min):
-e ceph_upgrade_poll_retries=240 -e ceph_upgrade_poll_delay=30
```
Preview the preflight/plan without starting anything:
```bash
scripts/ansible-play.sh upgrade-ceph.yml \
-e ceph_upgrade_target_image=quay.io/ceph/ceph:v20.2.2 --check
```
## Rollback
cephadm upgrades roll one daemon type at a time and auto-**pause** on the first
failure rather than tearing the cluster down. There is no destructive cutover,
so rollback = point the orchestrator back at the **prior** image and let it
converge.
```bash
# 0. Prior image: the play PRINTED it ("ROLLBACK ANCHOR ..."). Otherwise:
ceph orch ps --format json | \
python3 -c 'import sys,json;print(sorted({d["container_image_name"] for d in json.load(sys.stdin)}))'
# 1. Stop the in-flight upgrade (or `pause`/`resume` to hold and inspect):
ceph orch upgrade stop
# 2. Re-target the prior image (only daemons ahead of it move):
ceph orch upgrade start --image <PRIOR_IMAGE>
# 3. Watch it converge:
ceph orch upgrade status
ceph -W cephadm
watch ceph versions
# 4. Health checks once status shows not in_progress:
ceph health detail # expect HEALTH_OK
ceph versions # expect ONE overall version
ceph osd stat # all OSDs up + in
```
Re-running `upgrade-ceph.yml` with the prior image as the target performs the
same health-gated rollback with all the same checks.
## Podman compatibility (cross-major upgrades)
The baseline role `dpkg`-holds podman (and the rest of its ecosystem) at its
installed version so a stray `apt upgrade` cannot move the container runtime out
from under a live cluster. Patch-train upgrades (20.2.x -> 20.2.y) run on the same
podman, so the hold needs no attention. But a **cross-major Ceph upgrade** may
require a newer podman per the cephadm compatibility matrix, and the hold will
block the podman bump. When that applies:
```bash
# On each node (drive via a serial, health-gated converge, not all at once):
apt-mark unhold podman
apt-get install -y --only-upgrade podman # to a version the cephadm matrix allows
apt-mark hold podman # re-freeze at the new version
```
Do this BEFORE `ceph orch upgrade start`, verify `podman version` on every node,
then upgrade Ceph. Check the cephadm podman support matrix for the target release
first; too-new podman can also break cephadm, so bump to a *supported* version,
not merely the latest. `baseline_held_packages` controls which packages are held.
## Gotchas
- **No cross-major downgrade.** Rollback is only safe *within* the pinned
20.2.x train (patch level). Do not use this to cross 20.x -> 19.x -- Ceph does
not support it.
- **Paused mid-upgrade.** If the poll assert fails, the upgrade paused. Inspect
`ceph orch upgrade status` (`message` field) and `ceph -W cephadm`, fix the
offending daemon/host, then `ceph orch upgrade resume` -- or stop and roll
back per above.
- **Bootstrap host in `--limit`.** All cluster ops delegate to the bootstrap
node; do not `--limit` it out of the run.
- **Registry reachability.** Nodes must be able to pull the target image. A
stalled pull shows up as a paused upgrade with a pull error in the message.
## References
- `upgrade-ceph.yml` -- the playbook; its header carries the same rollback
procedure inline.
- `docs/runbooks/replace-host.md`, `docs/runbooks/add-node.md` -- related day-2
ops that also drive `ceph` from the bootstrap node.
+23 -4
View File
@@ -1,11 +1,15 @@
---
# Security hardening for Ceph nodes.
# Deploys nftables firewall rules and SSH hardening.
# Run AFTER deploy-ceph.yml — cephadm needs unrestricted access during deploy.
# Deploys the nftables firewall + the full sshd hardening drop-in.
# Runs AFTER deploy-ceph.yml - cephadm needs unrestricted access during deploy.
# NB: the security-critical sshd lockdown (root key-only, no password auth, locked
# root password) is applied EARLY in the baseline role, not here, so a node is
# never left root-password-open on the WAN during the deploy campaign; this stage
# adds the firewall and the remaining sshd hardening on top.
#
# WARNING: This locks down inbound traffic. Ensure ceph_firewall_trusted_networks
# includes all subnets that need to reach Ceph services. SSH (port 22) is always
# open from any source to prevent lockout.
# includes all subnets that need to reach Ceph services. SSH (port 22) stays open
# from any source (ceph_firewall_ssh_any_source) to prevent lockout.
#
# Usage:
# scripts/ansible-play.sh harden.yml
@@ -14,6 +18,21 @@
hosts: ceph_nodes
become: true
gather_facts: false
# Blast-radius control (see baseline.yml): batch the roll + halt on any node
# failure or cluster degradation. Especially important here - a bad nftables
# ruleset that fences the fabric shows up as an unreachable batch (halts on
# max_fail_percentage) or a degraded cluster at the post-batch health gate,
# before it reaches the rest of the fleet.
serial: "{{ ceph_converge_serial | default('10%') }}"
max_fail_percentage: 0
pre_tasks:
- name: Ceph-health gate (pre-batch)
ansible.builtin.import_tasks: tasks/health_gate.yml
roles:
- security
post_tasks:
- name: Ceph-health gate (post-batch)
ansible.builtin.import_tasks: tasks/health_gate.yml
@@ -0,0 +1,310 @@
---
# Cluster-wide Ansible config for spice (prod / htz-fsn1, 48x SX295).
# Committed. Hand-maintained. Adapted from the painbox (SX295/Hetzner) precedent
# with the network reshaped onto the bonded-25G fabric VLANs.
#
# NOTE: sections marked DRAFT are ceph data-protection / RGW decisions that must
# be reviewed before `mise run deploy` against 48 nodes. They do NOT affect the
# OS-install milestone (installimage), only the later cephadm convergence.
# === Naming ===
cluster_name: spice
cluster_role: ceph
cluster_domain: prod.fsn1.htz.futo.cloud
# === Network ===
# Ceph public/cluster traffic runs on the bonded 25G fabric VLANs (NetBox
# FSN1-C1-PUBLIC / FSN1-C1-PRIVATE). Per-host IPs are 10.40.20.<host_index> /
# 10.40.22.<host_index> (host_index from spice-hosts.yaml; gateway .1 = leaf IRB).
# A third fabric VLAN, FSN1-C1-HOST-MGMT (124, 10.40.24.0/24), is the in-band
# host-management path: the leaf trunks it to every server bond and the fabric
# advertises 10.40.24.0/24 into the NetBird overlay, so it is the intended
# ansible/SSH reach once the WAN is retired.
# The DEFAULT ROUTE stays on the 1G WAN for now (Hetzner gateway); ansible reaches
# nodes over their WAN address until the fabric is cut over. Revisit if/when the
# default route moves off the WAN (to Ceph Public or Host-Mgmt).
public_network: 10.40.20.0/23 # VLAN 120 (Ceph Public)
cluster_network: 10.40.22.0/23 # VLAN 122 (Ceph Private)
# Per-host Ceph BIND address on the fabric public VLAN. Derived from host_index the
# same way as networkd_bond_vlans below (a per-host host_var, so it resolves per
# host at play time). Lives INSIDE public_network -- unlike bond_ip, the 1G WAN
# address used only for the ansible/SSH connection (ansible_host). (No
# ceph_cluster_ip var: OSD replication binds to the cluster_network CIDR, which
# cephadm resolves to each node's bond0.122 address on its own.)
ceph_public_ip: "10.40.20.{{ host_index }}" # bond0.120, inside public_network
# The address Ceph point-services advertise + bind on (mon --mon-ip, cephadm
# host-add, dashboard monitoring URLs) = the fabric public IP. Ceph never touches
# the WAN. A group_var, NOT a role default, because join/rgw/monitoring read it via
# hostvars[<host>], which does not expose role defaults; here it resolves per host
# to 10.40.20.<host_index>. sietch is flat with no ceph_public_ip, so there it
# falls back to bond_ip (already inside its public_network).
ceph_service_ip: "{{ ceph_public_ip | default(bond_ip) }}"
# Fabric-only bind opt-in: cephadm restricts the RGW + monitoring daemon binds to
# these CIDR(s) so nothing Ceph listens on the 1G WAN (mon/mgr/OSD already bind
# public_network/cluster_network). The role default is empty (flat clusters bind
# every interface, unchanged); spice pins the fabric public network.
ceph_bind_networks:
- "{{ public_network }}"
# Bond of the two 25G NICs, LACP to match the leaf ae<k> aggregate (cluster-fabric).
bond_mode: "802.3ad"
# COEXISTENCE: keep the 1G WAN (oob_nic) on ifupdown exactly as installimage set
# it (static IP + default route), and let networkd manage ONLY the bond + VLANs.
# The WAN is never migrated, so a networkd error can't cost us the reachability
# link (the painbox failure). networkd's commit leaves networking.service enabled
# and an Unmanaged=yes guard fences oob_nic off. Flip to true only if the default
# route ever moves onto the fabric and the WAN is retired.
networkd_replace_ifupdown: false
# Enable the networkd role (defaults false, so it is a no-op unless set). spice's
# 25G fabric bond + VLANs are networkd-managed, so a reprovisioned node must bring
# them up from the repo alone; committed here so the live state is reproducible
# (previously supplied ad-hoc via -e). Safe: networkd_replace_ifupdown:false fences
# the 1G WAN. Roll one node at a time on any live reconverge.
networkd_enabled: true
# The two Intel E800 (ice) 25G ports, verified uniform across all 48 SX295 nodes
# (PCI 0000:c1:00.0 / .1). Consumed by the networkd role as the bond0 members.
bond_interfaces:
- enp193s0f0
- enp193s0f1
# 1G WAN (igb, PCI 0000:c5:00.0); holds the default route until the fabric cutover.
oob_nic: enp197s0
fabric_public_vlan: 120
fabric_private_vlan: 122
fabric_mgmt_vlan: 124
dns_server: 185.12.64.1 # Hetzner recursive (reachable via WAN default route)
# Time sync: Hetzner's own NTP (low-latency from FSN1, consistent across all 47).
# Ceph mon quorum is skew-sensitive; pin the source rather than the distro pool.
chrony_ntp_servers:
- ntp1.hetzner.de
- ntp2.hetzner.de
- ntp3.hetzner.de
# Tagged VLAN sub-interfaces on the 25G LACP bond (consumed by the networkd role).
# Addresses derive from host_index; masks match each network (public/private /23,
# host-mgmt /24). No Gateway here -- the default route stays on the 1G WAN until
# fabric cutover.
# Jumbo (MTU 9000) is enabled on VLAN 122 (Ceph PRIVATE / OSD replication) ONLY.
# That is a closed, homogeneous, all-under-our-control path where jumbo's packet/
# interrupt savings on replication+recovery actually pay off. VLAN 120 (Ceph public,
# where RGW faces heterogeneous 1500 S3 clients + the ~1400 NetBird overlay) and
# VLAN 124 (mgmt) stay at 1500: jumbo on a client-facing path is a partial-blackhole
# risk (small ops fine, large PUT/GET hang) for near-zero gain. The bond ceiling and
# both 25G members auto-raise to the max child MTU so the jumbo VLAN is not capped.
# PAIRED REQUIREMENT: the leaf server-LAG + the VLAN 122 IRB must be >= 9000 on the
# QFX side before the cluster network passes jumbo; validate with a bidirectional
# `ping -M do -s 8972` matrix on 10.40.22.0/23 + green ceph health before relying on it.
networkd_bond_vlans:
- id: "{{ fabric_public_vlan }}" # 120 - Ceph public (10.40.20.0/23)
address: "10.40.20.{{ host_index }}/23"
- id: "{{ fabric_private_vlan }}" # 122 - Ceph private (10.40.22.0/23)
address: "10.40.22.{{ host_index }}/23"
mtu: 9000
- id: "{{ fabric_mgmt_vlan }}" # 124 - Host mgmt (10.40.24.0/24)
address: "10.40.24.{{ host_index }}/24"
# Cluster-specific packages, merged with baseline_diag_packages. Spice hardware:
# Intel E800 (ice) 25G NICs on an LACP fabric; direct-attach SATA HDDs + NVMe (no
# SAS expander / Mellanox, so none of sietch's mstflint/ledmon/sg3-utils here).
baseline_extra_packages:
- lldpd # LLDP neighbor/switch discovery on the 25G fabric
- ethtool # NIC/bond link, driver, and 25G speed inspection
- tcpdump # packet capture for fabric/network debugging
# Intel E810 (ice) DDP package (intel/ice/ddp/ice.pkg). WITHOUT it the NIC boots
# into Safe Mode, whose crippled classifier drops LACP/LLDP control frames -> the
# 25G bond never aggregates. The minimal Debian image omits it; baseline installs
# it and (nic-firmware.yml) reboots once to load the DDP. Needs the non-free-firmware
# apt component (present in the Hetzner Debian base).
- firmware-misc-nonfree
timezone: UTC
# === Ceph ===
ceph_release: tentacle
ceph_repo_url: "https://download.ceph.com/debian-{{ ceph_release }}/"
ceph_repo_key_url: "https://download.ceph.com/keys/release.asc"
# === OS / Auth ===
# installimage boots root; the post-install script authorizes the cluster
# ansible-iac key for root, and ansible connects as root (matches painbox).
admin_user: root
provision_iac_ssh_key_path: "~/.ssh/id_ed25519_spice"
# 1P vault for cluster secret lookups.
cluster_secrets_vault: yucca_tf_prod
# === Secret aliases (op inject -> vault_* at play time) ===
ops_password: "{{ vault_ops_password }}"
ceph_dashboard_user: admin
ceph_dashboard_password: "{{ vault_ceph_dashboard_password }}"
ceph_grafana_admin_user: admin
ceph_grafana_admin_password: "{{ vault_grafana_admin_password }}"
ceph_rgw_s3_user_access_key: "{{ vault_s3_restic_access_key }}"
ceph_rgw_s3_user_secret_key: "{{ vault_s3_restic_secret_key }}"
# === Reprovision (installimage; reprovision_hetzner role) ===
# Most knobs live in roles/reprovision_hetzner/defaults; override the operator-
# facing ones here. VERIFY image name + boot_mode on node 1 (mise run reprovision
# -- --limit spice-ceph-adelia -e confirm_wipe=true -e reprovision_mode=inspect).
reprovision_boot_mode: bios # SX295 painbox precedent = BIOS (verified)
reprovision_expect_nvme: 2
reprovision_expect_sata: 14
# reprovision_image: override only if node-1 inspect shows a different image path/name.
# === Storage (SX295: 2x NVMe RAID-1 OS + 14x SATA HDD per node) ===
# block.db LVs carved from the NVMe RAID-1 vg0 at converge by ceph_deploy's
# lvm-setup (NVMe-RAID branch): 14x 128G db-slots + one ssd-osd LV from the vg0
# remainder minus reserve. HDDs become OSDs with block.db on those LVs.
# Consumed ONLY by ceph_destroy's partition-6 SSD-OSD wipe (sietch's dual-SSD
# shape). spice is NVMe-RAID: the ssd-osd is a vg0 LV removed by the generic VG/PV
# cleanup, and the NVMe carries no partition 6, so this pattern is a deliberate
# no-match here (the SSD-partition wipe finds nothing). cleanup.yml references it
# unconditionally, so it must stay defined; the value is a don't-care on spice.
ssd_model_pattern: "__no_ssd_osd_partitions_on_nvme_raid__"
ceph_db_vg: vg0 # single NVMe-RAID VG created by installimage
ceph_db_lv_size: "128G"
ceph_db_lvs_per_node: 14
ceph_ssd_osd_lv: ssd-osd # NVMe-backed OSD LV (fast replicated pool)
ceph_ssd_osd_reserve_gib: 512 # vg0 free left unallocated for headroom/expansion
# OSD device map (by-path, consumed by ceph_deploy/osd-spec.yml.j2, NVMe-RAID shape:
# no sas_path_prefix -> data path is /dev/disk/by-path/<path_phy>, db is /dev/<db>).
# The 14x 20TB SATA HDDs sit on 3 AHCI controllers; the by-path layout is IDENTICAL
# across 46 of the 47 nodes (verified 2026-07-10), so it lives here as a group_var
# instead of 47 host_vars. Exception: spice-ceph-miguel has one disk on
# 46:00.0-ata-4 (not 87:00.0-ata-4) and overrides ceph_hdd_osds in its host_vars.
# db-slotN and ssd-osd LVs are carved from vg0 at deploy by lvm-setup.
ceph_hdd_osds:
- { path_phy: pci-0000:45:00.0-ata-1, db: vg0/db-slot0 }
- { path_phy: pci-0000:45:00.0-ata-2, db: vg0/db-slot1 }
- { path_phy: pci-0000:45:00.0-ata-3, db: vg0/db-slot2 }
- { path_phy: pci-0000:45:00.0-ata-4, db: vg0/db-slot3 }
- { path_phy: pci-0000:45:00.0-ata-5, db: vg0/db-slot4 }
- { path_phy: pci-0000:45:00.0-ata-6, db: vg0/db-slot5 }
- { path_phy: pci-0000:45:00.0-ata-7, db: vg0/db-slot6 }
- { path_phy: pci-0000:45:00.0-ata-8, db: vg0/db-slot7 }
- { path_phy: pci-0000:46:00.0-ata-1, db: vg0/db-slot8 }
- { path_phy: pci-0000:46:00.0-ata-2, db: vg0/db-slot9 }
- { path_phy: pci-0000:87:00.0-ata-1, db: vg0/db-slot10 }
- { path_phy: pci-0000:87:00.0-ata-2, db: vg0/db-slot11 }
- { path_phy: pci-0000:87:00.0-ata-3, db: vg0/db-slot12 }
- { path_phy: pci-0000:87:00.0-ata-4, db: vg0/db-slot13 }
ceph_ssd_osds:
- { lv: vg0/ssd-osd }
# === RGW ===
# EC 16+4, HOST failure domain -- capacity-max profile (80% usable, m=4) for the
# beta behind michael. Host domain trades rack-loss protection for capacity: a
# destroyed rack means data loss, accepted because the beta is wipeable and the
# API tier is already single-rack. Rationale + alternatives in Brain Box
# spice-ec-profile-analysis.md. Hold m=4 (host domain is the only durability
# left); do not widen past k+m=20.
ceph_rgw_realm: spice
ceph_rgw_zonegroup: eu-central-1
ceph_rgw_zonegroup_api_name: eu-central-1
ceph_rgw_zone: prod-z1
ceph_rgw_ec_profile: ec-k16m4-host
ceph_rgw_ec_k: 16
ceph_rgw_ec_m: 4
ceph_rgw_ec_failure_domain: host
ceph_rgw_ec_device_class: hdd
# Start the bulk EC data pool near its autoscaler target. The role default (128)
# is sized for ~36 OSDs; spice has 672 HDD OSDs on a 20-chunk pool. pg_num 4096
# => ~122 PG-shards/OSD ( pg_num * (k+m) / osds ), inside the recommended 100-250
# per OSD for large BlueStore clusters (default target 100, docs recommend 200;
# >500 hurts). 2048 would be only ~61/OSD (too low); 4096 is the right power of 2.
# Starting low would PG-split during the first fill.
ceph_pg_init_data: 4096
ceph_rgw_data_pool: "{{ ceph_rgw_zone }}.rgw.buckets.data"
ceph_rgw_index_pool: "{{ ceph_rgw_zone }}.rgw.buckets.index"
ceph_rgw_extra_pool: "{{ ceph_rgw_zone }}.rgw.buckets.non-ec"
ceph_rgw_dns_name: s3.{{ cluster_domain }}
ceph_rgw_replicated_size: 3
ceph_rgw_replicated_min_size: 2
# Bootstrap default for any pool that does not pin its own size (footgun guard:
# an ad-hoc or future pool on a prod cluster should not land at single-replica
# 2/1). Applied via cluster-spec.yml.j2 at cephadm bootstrap; role default is
# 2/1 so sietch is unchanged. Named RGW pools still pin their own size above.
ceph_osd_pool_default_size: 3
ceph_osd_pool_default_min_size: 2
# --- Device-class CRUSH pool placement ---
# crush-rules.yml creates replicated_ssd/replicated_hdd but binds nothing; every
# replicated pool otherwise falls to the default class-agnostic rule, so the
# omap-heavy bucket index lands on HDD (~98.5% of weight) while the 47-OSD NVMe
# ssd-osd tier sits idle. Pin the latency-sensitive index/metadata/mgr pools to
# the SSD tier; keep the bulk non-EC pool on HDD. The EC data pool is pinned to
# hdd via ceph_rgw_ec_device_class (the EC profile), not here. OSD device classes
# are forced at creation in osd-spec.yml.j2 (ssd-osd -> ssd, HDD -> hdd), and
# verify.yml asserts every intended pool carries its rule and the ssd class has
# live OSDs.
ceph_rgw_ssd_crush_rule: replicated_ssd
ceph_rgw_hdd_crush_rule: replicated_hdd
ceph_rgw_ssd_pools:
- "{{ ceph_rgw_index_pool }}"
- "{{ ceph_rgw_zone }}.rgw.meta"
- .rgw.root
- "{{ ceph_rgw_zone }}.rgw.log"
- "{{ ceph_rgw_zone }}.rgw.control"
- .mgr
ceph_rgw_hdd_pools:
- "{{ ceph_rgw_extra_pool }}"
ceph_rgw_count_per_host: 1
ceph_rgw_port: 443
# RGW/beast terminates TLS on 443 with a 10-year self-signed cert: rgw.yml
# generates /etc/ceph/rgw-ssl.{crt,key} and cephadm distributes it to every RGW
# daemon via the service spec. Clients reach RGW over the fabric / NetBird, so the
# cert SANs cover the s3.<domain> name (+ wildcard for virtual-hosted buckets) and
# each node's fabric IP (ceph_service_ip). Self-signed: clients trust-on-first-use
# or pin the CA. Falkenstein/Saxony/DE matches the FSN1 site (cosmetic for a
# self-signed cert; only the CN + SANs are validated).
ceph_rgw_ssl: true
ceph_rgw_ssl_cert_days: 3650 # 10 years
ceph_rgw_ssl_cert_subject_c: DE
ceph_rgw_ssl_cert_subject_st: Saxony
ceph_rgw_ssl_cert_subject_l: Falkenstein
ceph_rgw_ssl_cert_subject_o: FUTO
ceph_rgw_ssl_cert_email: yucca@futo.org
ceph_rgw_scheme: "{{ 'https' if ceph_rgw_ssl else 'http' }}"
# --- RGW ingress: health-checked S3 VIP (haproxy + keepalived, TLS terminate) ---
# A cephadm ingress service puts a keepalived floating VIP with haproxy health
# checks in front of the 47 RGW daemons, in TLS TERMINATION mode (haproxy
# `mode http`: haproxy terminates the client TLS on the VIP with the reused RGW
# self-signed cert, then re-encrypts to each beast backend). Without it a dead RGW
# node blackholes ~1/47 of S3 connections; with it the node is health-checked out
# of the VIP. The VIP sits on the fabric public network (10.40.20.0/23, bond0.120);
# NetBird already advertises 10.40.20.0/23 into the overlay, so the VIP is
# reachable, and DNS points s3.<domain> at it (that DNS cutover is a separate TF
# apply that MUST follow the ingress being live, not precede it). The VIP is added
# to the RGW cert SAN (rgw.yml) so a client hitting it validates; virtual-hosted
# buckets keep working because mode http forwards the Host header to the wildcard
# cert. Terminate (not L4 passthrough) because passthrough needs a post-20.2.2
# cephadm field (use_tcp_mode_over_rgw) absent on the pinned 20.2.2.
#
# Preconditions are satisfied here: ceph_bind_networks non-empty (VIP:443 vs beast
# per-node-IP:443 are distinct sockets) and ceph_rgw_ssl true (cert exists).
ceph_rgw_ingress_enabled: true
ceph_rgw_ingress_vip: "10.40.20.250"
ceph_rgw_ingress_vip_prefix: 23 # matches public_network 10.40.20.0/23
ceph_rgw_s3_user_uid: svc-yucca-restic
ceph_rgw_s3_user_display_name: "yucca/restic service account"
# Metrics-worker RGW admin user (read-only): the yucca-metrics-worker service
# reads bucket usage from RadosGW. Keys are TF-minted (SPICE_METRICS_WORKER_*)
# and injected via secrets.yml.tpl as vault_metrics_worker_*; rgw.yml (Step 14.5)
# creates the user with these caps. Without this block the RGW phase fails with
# AnsibleUndefinedVariable. Mirror of the sietch definition.
ceph_rgw_metrics_user_uid: metrics-worker
ceph_rgw_metrics_user_display_name: "yucca/metrics-worker RGW admin (read-only)"
ceph_rgw_metrics_user_access_key: "{{ vault_metrics_worker_access_key }}"
ceph_rgw_metrics_user_secret_key: "{{ vault_metrics_worker_secret_key }}"
ceph_rgw_metrics_user_caps: "buckets=read;usage=read;metadata=read;users=read"
# === Monitoring Stack ===
ceph_prometheus_port: 9095
ceph_grafana_port: 3000
ceph_alertmanager_port: 9093
# ceph_grafana_admin_user/password are defined once in the Secret aliases block above.
# === Security / firewall ===
# Spice is fabric-only: every Ceph daemon binds the 25G fabric (mon/mgr/OSD via
# public_network/cluster_network; RGW + monitoring via ceph_bind_networks) and
# never listens on the 1G WAN. RGW is therefore scoped to the trusted networks
# rather than left world-open (ceph_firewall_rgw_any_source defaults true for
# flat/dev clusters; spice pins it closed).
ceph_firewall_rgw_any_source: false
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-adelia
bond_ip: 178.63.139.248
hetzner_server_number: 3008187
host_index: 4
mon: true
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-alexus
bond_ip: 178.63.139.254
hetzner_server_number: 3008189
host_index: 5
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-alyssa
bond_ip: 178.63.139.228
hetzner_server_number: 3008190
host_index: 6
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-ardith
bond_ip: 178.63.139.227
hetzner_server_number: 3008191
host_index: 7
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-athena
bond_ip: 178.63.139.225
hetzner_server_number: 3008192
host_index: 8
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-bernie
bond_ip: 178.63.139.226
hetzner_server_number: 3008193
host_index: 9
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-braden
bond_ip: 178.63.139.240
hetzner_server_number: 3008194
host_index: 10
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-callie
bond_ip: 178.63.139.243
hetzner_server_number: 3008195
host_index: 11
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-catina
bond_ip: 178.63.139.244
hetzner_server_number: 3008196
host_index: 12
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-cletus
bond_ip: 178.63.139.253
hetzner_server_number: 3008197
host_index: 13
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-curtis
bond_ip: 178.63.139.252
hetzner_server_number: 3008198
host_index: 14
mon: true
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-darion
bond_ip: 178.63.139.251
hetzner_server_number: 3008199
host_index: 15
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-deidra
bond_ip: 178.63.139.250
hetzner_server_number: 3008200
host_index: 16
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-deonte
bond_ip: 178.63.139.249
hetzner_server_number: 3008201
host_index: 17
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-dorian
bond_ip: 178.63.139.242
hetzner_server_number: 3008202
host_index: 18
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-edythe
bond_ip: 178.63.139.247
hetzner_server_number: 3008203
host_index: 19
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-eloise
bond_ip: 178.63.139.246
hetzner_server_number: 3008204
host_index: 20
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-evelyn
bond_ip: 178.63.139.245
hetzner_server_number: 3008205
host_index: 21
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-gaylon
bond_ip: 178.63.139.241
hetzner_server_number: 3008206
host_index: 22
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-graham
bond_ip: 178.63.139.239
hetzner_server_number: 3008207
host_index: 23
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-hayley
bond_ip: 178.63.139.230
hetzner_server_number: 3012008
host_index: 24
mon: true
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-howell
bond_ip: 178.63.139.222
hetzner_server_number: 3012009
host_index: 25
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-jacque
bond_ip: 178.63.139.221
hetzner_server_number: 3012010
host_index: 26
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-javier
bond_ip: 178.63.139.207
hetzner_server_number: 3012011
host_index: 27
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-jewell
bond_ip: 178.63.139.210
hetzner_server_number: 3012012
host_index: 28
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-joseph
bond_ip: 178.63.139.218
hetzner_server_number: 3012013
host_index: 29
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-kassie
bond_ip: 178.63.139.208
hetzner_server_number: 3012014
host_index: 30
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-kelsea
bond_ip: 178.63.139.217
hetzner_server_number: 3012015
host_index: 31
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-kurtis
bond_ip: 178.63.139.234
hetzner_server_number: 3012016
host_index: 32
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-lemuel
bond_ip: 178.63.139.214
hetzner_server_number: 3012017
host_index: 33
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-lizzie
bond_ip: 178.63.139.213
hetzner_server_number: 3014572
host_index: 34
mon: true
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-lucius
bond_ip: 178.63.139.212
hetzner_server_number: 3014573
host_index: 35
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-marian
bond_ip: 178.63.139.211
hetzner_server_number: 3014574
host_index: 36
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-mattie
bond_ip: 178.63.139.237
hetzner_server_number: 3014575
host_index: 37
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,24 @@
---
hostname_short: spice-ceph-miguel
bond_ip: 178.63.139.216
hetzner_server_number: 3014576
host_index: 38
mon: false
# OVERRIDE of the group_vars ceph_hdd_osds: miguel has one HDD on 46:00.0-ata-4
# (not 87:00.0-ata-4 like the other 46 nodes) -- verified 2026-07-10. host_vars
# fully replaces the group_var, so the whole 14-disk map is repeated here.
ceph_hdd_osds:
- { path_phy: pci-0000:45:00.0-ata-1, db: vg0/db-slot0 }
- { path_phy: pci-0000:45:00.0-ata-2, db: vg0/db-slot1 }
- { path_phy: pci-0000:45:00.0-ata-3, db: vg0/db-slot2 }
- { path_phy: pci-0000:45:00.0-ata-4, db: vg0/db-slot3 }
- { path_phy: pci-0000:45:00.0-ata-5, db: vg0/db-slot4 }
- { path_phy: pci-0000:45:00.0-ata-6, db: vg0/db-slot5 }
- { path_phy: pci-0000:45:00.0-ata-7, db: vg0/db-slot6 }
- { path_phy: pci-0000:45:00.0-ata-8, db: vg0/db-slot7 }
- { path_phy: pci-0000:46:00.0-ata-1, db: vg0/db-slot8 }
- { path_phy: pci-0000:46:00.0-ata-2, db: vg0/db-slot9 }
- { path_phy: pci-0000:46:00.0-ata-4, db: vg0/db-slot10 }
- { path_phy: pci-0000:87:00.0-ata-1, db: vg0/db-slot11 }
- { path_phy: pci-0000:87:00.0-ata-2, db: vg0/db-slot12 }
- { path_phy: pci-0000:87:00.0-ata-3, db: vg0/db-slot13 }
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-murphy
bond_ip: 178.63.139.232
hetzner_server_number: 3014577
host_index: 39
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,9 @@
---
hostname_short: spice-ceph-noreen
bond_ip: 178.63.139.220
hetzner_server_number: 3014578
host_index: 40
mon: false
# ceph_hdd_osds: defined in group_vars/all/vars.yml (SX295 by-path map is
# uniform across nodes). Override here ONLY if a node's by-path differs
# (e.g. spice-ceph-miguel). Verify with: ls -l /dev/disk/by-path/pci-*-ata-*
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-philip
bond_ip: 178.63.139.219
hetzner_server_number: 3014579
host_index: 41
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-raymon
bond_ip: 178.63.139.223
hetzner_server_number: 3014580
host_index: 42
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-romona
bond_ip: 178.63.139.236
hetzner_server_number: 3014581
host_index: 43
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-serena
bond_ip: 178.63.139.229
hetzner_server_number: 3014582
host_index: 44
mon: true
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-shanna
bond_ip: 178.63.139.209
hetzner_server_number: 3014583
host_index: 45
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-shelby
bond_ip: 178.63.139.224
hetzner_server_number: 3014584
host_index: 46
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-sommer
bond_ip: 178.63.139.215
hetzner_server_number: 3014585
host_index: 47
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-sylvia
bond_ip: 178.63.139.235
hetzner_server_number: 3014586
host_index: 48
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-theron
bond_ip: 178.63.139.238
hetzner_server_number: 3014587
host_index: 49
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-trista
bond_ip: 178.63.139.233
hetzner_server_number: 3014588
host_index: 50
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -0,0 +1,7 @@
---
hostname_short: spice-ceph-virgie
bond_ip: 178.63.139.231
hetzner_server_number: 3014589
host_index: 51
mon: false
# ceph_hdd_osds: DEFERRED until node-1 inspection (uniform SX295 by-path map)
@@ -7,6 +7,12 @@ cluster_role: ceph
cluster_domain: staging.austin.int.futo.cloud
public_network: 10.10.10.0/24
cluster_network: 10.10.10.0/24
# The address Ceph point-services advertise + bind on (mon --mon-ip, cephadm
# host-add, dashboard monitoring URLs). Flat cluster: no ceph_public_ip, so this
# resolves to bond_ip (10.10.10.9x) - identical to the pre-fabric behavior. A
# group_var, NOT a role default, because it is read via hostvars[<host>] which
# does not expose role defaults.
ceph_service_ip: "{{ ceph_public_ip | default(bond_ip) }}"
gateway: 10.10.10.1
dns_server: 10.10.10.1
bond_mode: active-backup
+9 -1
View File
@@ -29,14 +29,19 @@
hosts: ceph_nodes
become: true
gather_facts: false
serial: 1
serial: "{{ networkd_serial | default(1) }}"
max_fail_percentage: 0 # any node fails (e.g. WAN drops -> unreachable) -> halt
pre_tasks:
# Ceph-health gate applies only to a live cluster. Pre-Ceph clusters (spice
# bringup: coexistence activation before any cephadm deploy) have no
# ceph_bootstrap group and no ceph binary, so skip it automatically.
- name: Preflight -- check Ceph health
ansible.builtin.command: ceph health --format json
register: ceph_health_pre
changed_when: false
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
when: (groups['ceph_bootstrap'] | default([]) | length) > 0
- name: Assert cluster is healthy
ansible.builtin.assert:
@@ -46,6 +51,7 @@
Ceph is {{ (ceph_health_pre.stdout | from_json).status }}.
Fix cluster health before migrating.
success_msg: "Ceph: {{ (ceph_health_pre.stdout | from_json).status }}"
when: (groups['ceph_bootstrap'] | default([]) | length) > 0
roles:
- role: networkd
@@ -56,7 +62,9 @@
register: ceph_health_post
changed_when: false
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
when: (groups['ceph_bootstrap'] | default([]) | length) > 0
- name: Report Ceph health
ansible.builtin.debug:
msg: "Ceph: {{ (ceph_health_post.stdout | from_json).status }}"
when: (groups['ceph_bootstrap'] | default([]) | length) > 0
+55
View File
@@ -0,0 +1,55 @@
---
# Reprovision Hetzner SX295 ceph nodes via the Robot API + installimage.
# Ansible-native replacement for the out-of-band installimage step:
# arm rescue (uri) -> hw reset (uri) -> wait for rescue -> [zap OSD HDDs] ->
# installimage -> reboot -> verify installed OS. Then hand off to `mise run deploy`.
#
# ALWAYS DESTRUCTIVE to the NVMe OS disks. NEVER fold this into deploy/converge.
# Launch ONLY via `mise run reprovision` (wraps op run --env-file=../../tf/.env.prod,
# which materializes HETZNER_ROBOT_* and lets ansible-play.sh op-inject secrets).
# Set CEPH_ENV first (e.g. inventories/prod-htz-fsn1/spice/inventory.ini).
#
# Inspect node 1 (arm rescue + reset + assert disks/boot-mode, NO wipe):
# mise run reprovision -- --limit spice-ceph-adelia -e confirm_wipe=true -e reprovision_mode=inspect
# Canary (one node, full wipe + verify; verifying it unlocks fan-out):
# mise run reprovision -- --limit spice-ceph-adelia -e confirm_wipe=true
# Fan-out (remaining nodes, batched; gated on the canary marker):
# mise run reprovision -- --limit 'ceph_nodes:!spice-ceph-adelia' -e confirm_wipe=true -e reprovision_serial=5
# Resume after a halt (skip already-provisioned nodes, converge only the rest):
# mise run reprovision -- --limit ceph_nodes -e confirm_wipe=true -e allow_fanout=true \
# -e reprovision_skip_if_provisioned=true -e reprovision_serial=5
# Day-2 reimage of a node (also zap its OSD HDDs; the ceph-safety gate verifies
# ok-to-stop against a live mon and refuses if the OSDs are not safe to lose):
# mise run reprovision -- --limit <host> -e confirm_wipe=true -e reprovision_wipe_osd_disks=true \
# -e reprovision_ceph_safety=strict
- name: Reprovision ceph nodes (rescue -> installimage -> installed OS)
hosts: ceph_nodes
gather_facts: false
become: false
serial: "{{ reprovision_serial | default(1) }}"
max_fail_percentage: 0 # any node fails -> halt the batch (no further wipes)
pre_tasks:
# run_once + included through the role so role defaults load and --limit can't
# skip it (a hosts: localhost play would be). Asserts creds/keys, enforces the
# fan-out gate, registers rescue keys, publishes fingerprints to localhost facts.
- name: Preflight + register robot keys (once) # noqa: run-once[task]
ansible.builtin.include_role:
name: reprovision_hetzner
tasks_from: preflight
run_once: true
roles:
- role: reprovision_hetzner
- name: Verify reprovisioned nodes after reboot
hosts: ceph_nodes
gather_facts: false
become: false
tasks:
- name: Skip verify in inspect mode (no install happened)
ansible.builtin.meta: end_play
when: reprovision_mode | default('install') != 'install'
- name: Verify installed OS marker
ansible.builtin.include_role:
name: reprovision_hetzner
tasks_from: verify
+6 -2
View File
@@ -1,6 +1,10 @@
---
# Galaxy collections are pinned to the currently-resolved major and floored at
# a known-good minor. Compatible-release ranges (no lockfile in this repo) keep
# a fleet-wide Ceph campaign reproducible: a new major cannot be pulled in
# mid-run and change module behavior between the first and last node.
collections:
- name: ansible.posix
version: ">=2.0.0"
version: ">=2.0.0,<3.0.0"
- name: community.general
version: ">=9.0.0"
version: ">=12.0.0,<13.0.0"
+37 -1
View File
@@ -2,6 +2,22 @@
# Post-boot OS baseline — runs via ansible-iac after provisioning.
# Configures everything that doesn't need to be in the chroot.
# --- iac automation access (convergence twin of the reprovision -x chroot) ---
# Ensures root's iac key + the ansible-iac sudo account on already-installed nodes
# without reimaging. baseline_iac_pubkey derives from the cluster's iac key; the
# iac-access task is gated on provision_iac_ssh_key_path so clusters without one
# skip. Override baseline_iac_pubkey directly to authorize a different key.
baseline_iac_user: ansible-iac
baseline_iac_pubkey: >-
{{ lookup('file', (provision_iac_ssh_key_path | expanduser) ~ '.pub') }}
# --- Intel ice NIC firmware (DDP) ---
# nic-firmware.yml reboots an E810 host once to load the DDP (exit Safe Mode) after
# firmware-misc-nonfree is installed. Set false to install the package but skip the
# activating reboot (an operator can reboot on their own schedule).
baseline_nic_firmware_activate: true
baseline_nic_firmware_reboot_timeout: 600
# --- ops user ---
baseline_ops_user: ops
baseline_ops_uid: 1001
@@ -55,8 +71,28 @@ baseline_diag_packages:
# ...) sets its own list or leaves it empty.
baseline_extra_packages: []
# Load-bearing host packages frozen at their installed version via a dpkg hold, so
# a stray `apt upgrade` cannot move them out from under a live cluster. We run NO
# unattended-upgrades by design; upgrading these is a deliberate, health-gated act
# (unhold -> upgrade -> re-hold, ideally through the serial converge). The default
# covers cephadm's container runtime (podman + ecosystem) and the chrony time
# source that mon quorum depends on. Diagnostic/ops tools are intentionally NOT
# held (a version drift there is harmless). Extend per cluster, or set [] to
# disable. For an exact-version pin (rather than freeze-at-installed) add an apt
# preferences entry; the hold covers the common "do not let it drift" need.
baseline_held_packages: "{{ baseline_podman_packages + ['chrony'] }}"
# --- Time sync (chrony) ---
# Pinned NTP sources for chrony.conf. Default to the Debian pool as a safe,
# explicit fallback; each cluster overrides with a consistent low-latency set
# (e.g. spice -> Hetzner NTP). Owned by chrony.yml (install+config+enable), NOT
# baseline_enable_services, so system.yml never starts a not-yet-installed unit.
chrony_ntp_servers:
- 0.debian.pool.ntp.org
- 1.debian.pool.ntp.org
- 2.debian.pool.ntp.org
# --- Services ---
baseline_enable_services:
- dbus
- chrony
- podman.socket
@@ -0,0 +1,23 @@
---
- name: Restart lldpd
ansible.builtin.systemd:
name: lldpd
state: restarted
- name: Restart chrony
ansible.builtin.systemd:
name: chrony
state: restarted
# sshd drop-in changes notify "Validate sshd config"; it runs sshd -t against the
# real combined config and only then notifies "Reload sshd" (handlers fire in
# listed order), so an invalid config halts the play before any reload.
- name: Validate sshd config
ansible.builtin.command: /usr/sbin/sshd -t
changed_when: false
notify: Reload sshd
- name: Reload sshd
ansible.builtin.systemd:
name: ssh
state: reloaded
@@ -0,0 +1,24 @@
---
# Time sync (chrony) for Ceph. Owns the package + config + service so the source
# set is explicit and versioned, not left to the distro default pool. Ceph mon
# quorum is skew-sensitive, so this is a hard prerequisite, not a nicety.
- name: Install chrony
ansible.builtin.apt:
name: chrony
state: present
- name: Deploy chrony.conf (pinned NTP sources)
ansible.builtin.template:
src: chrony.conf.j2
dest: /etc/chrony/chrony.conf
owner: root
group: root
mode: '0644'
notify: Restart chrony
- name: Enable and start chrony
ansible.builtin.systemd:
name: chrony
enabled: true
state: started
@@ -0,0 +1,21 @@
---
# Freeze the load-bearing host packages at their installed version.
#
# We deliberately do NOT run unattended-upgrades: an apt upgrade that moved the
# cephadm container runtime (podman + ecosystem) or the time source (chrony, which
# mon quorum depends on) out from under a live cluster is exactly the kind of
# uncoordinated change we avoid. Instead we HOLD what matters (dpkg selection) so
# `apt upgrade` skips them; upgrading them becomes a deliberate, health-gated act
# (unhold -> upgrade -> re-hold, ideally through the serial converge). Diagnostic
# and ops tools are intentionally not held (a version drift there is harmless).
#
# Runs last in the baseline role so every held package is already installed.
# Idempotent: dpkg_selections only reports changed when the selection differs.
# Set baseline_held_packages: [] to disable holds entirely.
- name: Freeze load-bearing packages at their installed version (dpkg hold)
ansible.builtin.dpkg_selections:
name: "{{ item }}"
selection: hold
loop: "{{ baseline_held_packages }}"
when: baseline_held_packages | length > 0
@@ -0,0 +1,39 @@
---
# Idempotent convergence twin of the reprovision -x chroot (reprovision_hetzner/
# templates/post-install.sh.j2). Seeds the SAME login access -- root's iac key and
# the ansible-iac sudo account -- so nodes that were installed WITHOUT the chroot
# (e.g. the first 36 imaged before it existed) converge to the same state without
# reimaging. Fresh nodes already have it from the chroot; this just re-affirms.
#
# Runs as the connection user (spice: root). Gated on provision_iac_ssh_key_path so
# clusters that do not define an iac key skip cleanly. Standalone:
# scripts/ansible-play.sh <spice> baseline.yml --tags iac_access --limit <hosts>
- name: Ensure the ansible-iac automation user
ansible.builtin.user:
name: "{{ baseline_iac_user }}"
shell: /bin/bash
create_home: true
state: present
- name: Passwordless sudo for ansible-iac
ansible.builtin.copy:
content: "{{ baseline_iac_user }} ALL=(ALL) NOPASSWD:ALL\n"
dest: "/etc/sudoers.d/{{ baseline_iac_user }}"
owner: root
group: root
mode: '0440'
validate: "visudo -cf %s"
- name: Authorize the iac key on ansible-iac
ansible.posix.authorized_key:
user: "{{ baseline_iac_user }}"
key: "{{ baseline_iac_pubkey }}"
state: present
# Non-exclusive: root is the reachability guarantee and must never be pruned here.
- name: Authorize the iac key on root
ansible.posix.authorized_key:
user: root
key: "{{ baseline_iac_pubkey }}"
state: present
@@ -0,0 +1,20 @@
---
# Configure lldpd (installed via baseline_extra_packages on fabric clusters).
# Advertises the node over LLDP so the leaf switches and NetBox can discover it
# on the 25G fabric. Imported from main.yml only when 'lldpd' is in the package
# list, so clusters without lldpd are unaffected.
- name: Deploy lldpd configuration
ansible.builtin.template:
src: lldpd.conf.j2
dest: /etc/lldpd.d/10-ceph.conf
owner: root
group: root
mode: '0644'
notify: Restart lldpd
- name: Enable and start lldpd
ansible.builtin.systemd:
name: lldpd
enabled: true
state: started
@@ -5,6 +5,18 @@
# (created by provision_host during initial OS install).
# Everything here is convergeable via normal deploy pipeline.
- name: Ensure iac automation access (root key + ansible-iac)
ansible.builtin.import_tasks: iac-access.yml
when: provision_iac_ssh_key_path is defined
tags: [iac_access, users]
# Immediately after the iac key is authorized on root: close root password login
# and password auth (do NOT wait for the harden stage). Prevents an installimage
# node from sitting root-password-open on the WAN during the deploy campaign.
- name: Harden SSH early (root key-only, no password auth, lock root)
ansible.builtin.import_tasks: ssh-hardening.yml
tags: [ssh, hardening, security]
- name: Configure ops user
ansible.builtin.import_tasks: users.yml
tags: [users]
@@ -13,6 +25,28 @@
ansible.builtin.import_tasks: packages.yml
tags: [packages]
- name: Configure time sync (chrony) for Ceph
ansible.builtin.import_tasks: chrony.yml
tags: [chrony, time, system]
# After packages (firmware-misc-nonfree) are installed: load the Intel ice DDP so
# E810 NICs leave Safe Mode. Self-gating (no-op on non-ice hosts / already-loaded).
- name: Activate Intel ice NIC firmware (DDP / exit Safe Mode)
ansible.builtin.import_tasks: nic-firmware.yml
tags: [packages, nic_firmware, network]
- name: Configure lldpd for fabric discovery
ansible.builtin.import_tasks: lldpd.yml
when: "'lldpd' in (baseline_extra_packages | default([]))"
tags: [lldpd, network]
- name: Configure /etc/hosts and services
ansible.builtin.import_tasks: system.yml
tags: [system]
# Last: freeze the load-bearing packages installed above (podman ecosystem,
# chrony) so a stray apt upgrade cannot move them under a live cluster. Runs after
# every install so nothing is held before it exists.
- name: Freeze load-bearing package versions (no unattended upgrades)
ansible.builtin.import_tasks: hold-packages.yml
tags: [packages, hold]
@@ -0,0 +1,71 @@
---
# Intel E810 (ice) NICs need their DDP package (intel/ice/ddp/ice.pkg, provided by
# firmware-misc-nonfree, installed via baseline_extra_packages). Without it the NIC
# boots into Safe Mode: the packet classifier is crippled and drops reserved-multicast
# control frames (LACP 01:80:c2:00:00:02, LLDP ..:0e), so the 25G fabric bond never
# aggregates even though the switch is correctly configured. The DDP only loads when
# ice (re)probes, so once the package is present a single reboot clears Safe Mode.
#
# Safe by design: only acts on hosts with an ice NIC; only reboots when the NIC is
# actually in Safe Mode AND the DDP file is present (so the reboot will fix it, and
# healthy nodes are never rebooted); refuses to reboot if the DDP is still missing
# (would loop). The WAN is on a separate igb NIC, so the reboot keeps the node
# reachable (coexistence). Idempotent: a converged node is a no-op.
- name: Find an ice-bound netdev (is this an E810 host?)
ansible.builtin.shell: |
set -o pipefail
for d in /sys/bus/pci/drivers/ice/0000:*/net/*; do
[ -e "$d" ] && { basename "$d"; break; }
done
args:
executable: /bin/bash
register: baseline_ice_nic
changed_when: false
failed_when: false
- name: Activate the ice DDP (exit Safe Mode) when needed
when: baseline_ice_nic.stdout | trim | length > 0
block:
- name: Check whether the ice NIC is in Safe Mode # noqa: risky-shell-pipe
# In Safe Mode advanced ops are disabled: --show-fec fails "not supported".
# NB: no `set -o pipefail` here -- ethtool exits non-zero on "not supported",
# which pipefail would propagate and mask grep's match, always yielding "ok".
ansible.builtin.shell: |
if ethtool --show-fec {{ baseline_ice_nic.stdout | trim }} 2>&1 | grep -qi "not supported"; then
echo safemode
else
echo ok
fi
args:
executable: /bin/bash
register: baseline_ice_mode
changed_when: false
# Gate on the dpkg DB (deterministic), not a file stat: after a large apt
# transaction the DDP file can lag briefly on disk, but dpkg-query reflects the
# install immediately -- and if the package is installed, the DDP file is present.
- name: Check firmware-misc-nonfree is installed (provides the DDP)
ansible.builtin.command: dpkg-query -W -f=${Status} firmware-misc-nonfree
register: baseline_ice_ddp_pkg
changed_when: false
failed_when: false
- name: Warn - NIC in Safe Mode but firmware-misc-nonfree not installed (do not reboot; would loop)
ansible.builtin.debug:
msg: >-
ice NIC {{ baseline_ice_nic.stdout | trim }} is in Safe Mode but
firmware-misc-nonfree is not installed. NOT rebooting. Check that it is in
baseline_extra_packages and that the non-free-firmware apt component is enabled.
when:
- "'safemode' in baseline_ice_mode.stdout"
- "'install ok installed' not in baseline_ice_ddp_pkg.stdout"
- name: Reboot once to load the ice DDP and exit Safe Mode
ansible.builtin.reboot:
reboot_timeout: "{{ baseline_nic_firmware_reboot_timeout }}"
msg: "Loading Intel ice DDP (exit Safe Mode) after firmware-misc-nonfree install"
when:
- baseline_nic_firmware_activate | bool
- "'safemode' in baseline_ice_mode.stdout"
- "'install ok installed' in baseline_ice_ddp_pkg.stdout"
@@ -0,0 +1,42 @@
---
# Early SSH lockdown. Runs in baseline (right after iac-access authorizes the iac
# key on root), NOT gated behind the harden stage -- so a freshly installimaged
# node is never left with root PASSWORD login exposed on the public WAN before
# harden runs (installimage ships PermitRootLogin yes + a root password). The
# security role's 50-hardening.conf applies the FULL sshd hardening later; this is
# the security-critical subset only (root key-only, no password auth, locked root
# password), using the SAME values so the two drop-ins never disagree.
#
# Safe on both clusters and idempotent on live nodes: ansible reaches every node
# by KEY (root on spice, ansible-iac on sietch) and iac-access has already
# authorized the key on root, so disabling password auth and locking the root
# password cannot lock anyone out. The Validate->Reload handler chain reloads
# (never restarts) only on change, after sshd -t passes.
- name: Deploy the baseline SSH lockdown drop-in
ansible.builtin.copy:
dest: /etc/ssh/sshd_config.d/10-baseline-ssh.conf
content: |
# Managed by ansible (roles/baseline/tasks/ssh-hardening.yml). Early
# root/password lockdown; roles/security 50-hardening.conf adds the rest.
PermitRootLogin prohibit-password
PasswordAuthentication no
owner: root
group: root
mode: '0644'
notify: Validate sshd config
tags: [ssh, hardening, security]
- name: Lock the root password (root is key-only; ansible/cephadm use keys)
ansible.builtin.user:
name: root
password_lock: true
tags: [ssh, hardening, security]
# Supersede the one-off interim drop-in from the pre-baseline BLOCKER remediation.
- name: Remove the interim root-password drop-in if present
ansible.builtin.file:
path: /etc/ssh/sshd_config.d/00-disable-root-password.conf
state: absent
notify: Validate sshd config
tags: [ssh, hardening, security]
@@ -17,3 +17,10 @@
enabled: true
state: started
loop: "{{ baseline_enable_services }}"
# timezone: UTC is declared in group_vars but nothing set it, so nodes drifted to
# the image default (Europe/Berlin) and logs render in CEST. Idempotent, reboot-free.
- name: Set the system timezone
community.general.timezone:
name: "{{ timezone }}"
when: timezone is defined
+7 -2
View File
@@ -1,5 +1,5 @@
---
# ops user — human-interactive account.
# ops user - human-interactive account.
# Convergeable: running this again fixes drift in password, sudo, shell.
- name: Ensure ops user exists
@@ -27,7 +27,12 @@
- name: Set ops user password
ansible.builtin.user:
name: "{{ baseline_ops_user }}"
password: "{{ ops_password | password_hash('sha512') }}"
# password_hash draws a RANDOM salt by default, so the hash differs every run
# and the user module rewrites /etc/shadow (reporting 'changed') on every
# converge. Derive a STABLE salt from the password (one-way sha256, so the
# public salt in /etc/shadow leaks nothing) to make the hash deterministic and
# the task idempotent, while still reconciling an out-of-band password change.
password: "{{ ops_password | password_hash('sha512', salt=(ops_password | hash('sha256'))[:16]) }}"
no_log: true
- name: Configure ops sudo
@@ -0,0 +1,15 @@
# {{ ansible_managed }}
# Time sync for Ceph. Mons will not form/keep quorum with clock skew, so pin a
# consistent, low-latency source set (chrony_ntp_servers) rather than the distro
# default pool. makestep allows a large one-off correction on freshly-imaged nodes
# before the cluster forms; after that chrony only slews.
{% for s in chrony_ntp_servers %}
server {{ s }} iburst
{% endfor %}
driftfile /var/lib/chrony/chrony.drift
makestep 1.0 3
rtcsync
logdir /var/log/chrony
leapsectz right/UTC
@@ -1,8 +1,10 @@
127.0.0.1 localhost
# Ceph cluster nodes
# Ceph cluster nodes. Map each name to ceph_service_ip -- the fabric IP on spice
# (bond0.120), falling back to bond_ip on flat clusters like sietch -- so in-cluster
# name resolution matches where Ceph binds, not the 1G WAN.
{% for host in groups['ceph_nodes'] %}
{{ hostvars[host]['bond_ip'] }} {{ hostvars[host]['hostname_short'] }}.{{ cluster_domain }} {{ hostvars[host]['hostname_short'] }}
{{ hostvars[host]['ceph_service_ip'] }} {{ hostvars[host]['hostname_short'] }}.{{ cluster_domain }} {{ hostvars[host]['hostname_short'] }}
{% endfor %}
::1 localhost ip6-localhost ip6-loopback
@@ -0,0 +1,7 @@
# {{ ansible_managed }}
# lldpd config for ceph fabric nodes (read from /etc/lldpd.d/ at daemon start).
# Port ID as the interface name (readable in NetBox / on the switch) and a
# cluster-tagged system description. lldpd advertises on all interfaces by
# default; the leaf only cares about the bonded 25G members.
configure lldp portidsubtype ifname
configure system description "{{ cluster_name | default('ceph') }} {{ cluster_role | default('ceph') }} node"
@@ -0,0 +1,32 @@
---
# Scheduled Ceph cluster-state backup defaults.
#
# The capture is read-only (dumps config/topology maps via the on-node admin
# keyring) and writes to a root-only local dir, so it is universally safe. Set
# ceph_backup_enabled: false in a cluster's group_vars to opt out.
ceph_backup_enabled: true
# Local landing dir for tarballs. Root-only (0700): the capture set is
# config/topology only, but keep it locked down regardless.
ceph_backup_dir: /var/backups/ceph
# Prune tarballs older than N days on each run. 0 disables pruning.
ceph_backup_retention_days: 14
# Where the capture script is installed on the bootstrap node.
ceph_backup_script_path: /usr/local/sbin/ceph-backup.sh
# Offsite sync target - EMPTY BY DEFAULT (operator decision; do not hardcode a
# destination). When set the script ships each tarball after capture:
# rsync form: user@host:/path/ or rsync://host/module/path
# S3 form: s3://bucket/prefix (requires the aws CLI on the bootstrap node)
# TODO(operator): choose + wire an offsite destination for true DR.
ceph_backup_offsite_dest: ""
# systemd timer schedule + jitter. Daily with an up-to-1h randomized delay so a
# fleet of clusters does not stampede a shared offsite target at the same second.
ceph_backup_oncalendar: "daily"
ceph_backup_randomized_delay: "1h"
# Kick one immediate capture at install time so the operator sees it work.
ceph_backup_run_now: true
@@ -0,0 +1,4 @@
---
- name: Reload systemd
ansible.builtin.systemd:
daemon_reload: true
@@ -0,0 +1,14 @@
---
# No dependencies - deploys a self-contained capture script + systemd timer on
# the bootstrap node. Independent of the deploy pipeline; run deliberately via
# backup-ceph.yml. Complements post-deploy-capture.yml (secrets -> 1P) and
# backup-config.yml (one-shot controller-local export) with a SCHEDULED,
# on-node config/topology backup.
dependencies: []
galaxy_info:
author: FUTO
license: AGPL-3.0-only
role_name: ceph_backup
description: Scheduled capture of Ceph cluster config/topology state to local tarballs for DR
min_ansible_version: "2.19"
@@ -0,0 +1,80 @@
---
# Deploy the scheduled cluster-state backup: capture script + systemd
# service/timer on the bootstrap node. Read-only capture to a root-only local
# dir, so it is universally safe; gate with ceph_backup_enabled to opt out.
- name: Scheduled backup install
when: ceph_backup_enabled | bool
block:
- name: Ensure the backup directory exists (root-only)
ansible.builtin.file:
path: "{{ ceph_backup_dir }}"
state: directory
owner: root
group: root
mode: '0700'
- name: Deploy the ceph-backup.sh capture script
ansible.builtin.template:
src: ceph-backup.sh.j2
dest: "{{ ceph_backup_script_path }}"
owner: root
group: root
mode: '0700'
- name: Deploy the ceph-backup.service unit
ansible.builtin.template:
src: ceph-backup.service.j2
dest: /etc/systemd/system/ceph-backup.service
owner: root
group: root
mode: '0644'
notify: Reload systemd
- name: Deploy the ceph-backup.timer unit
ansible.builtin.template:
src: ceph-backup.timer.j2
dest: /etc/systemd/system/ceph-backup.timer
owner: root
group: root
mode: '0644'
notify: Reload systemd
- name: Reload systemd now so the units are current before enabling
ansible.builtin.meta: flush_handlers
- name: Enable and start the backup timer
ansible.builtin.systemd:
name: ceph-backup.timer
enabled: true
state: started
# Starting the oneshot service blocks until ExecStart returns, so this both
# proves the script runs and leaves a first tarball to inspect.
- name: Trigger one immediate backup to validate the install
ansible.builtin.systemd:
name: ceph-backup.service
state: started
when: ceph_backup_run_now | bool
- name: Find the most recent backup tarball
ansible.builtin.find:
paths: "{{ ceph_backup_dir }}"
patterns: 'ceph-state-*.tar.gz'
register: ceph_backup_tarballs
when: ceph_backup_run_now | bool
- name: Report backup install status
ansible.builtin.debug:
msg: >-
Scheduled backups installed on {{ inventory_hostname }}: timer
{{ ceph_backup_oncalendar }} (jitter {{ ceph_backup_randomized_delay }}),
dir {{ ceph_backup_dir }}, retention {{ ceph_backup_retention_days }}d,
offsite {{ ('set: ' ~ ceph_backup_offsite_dest)
if (ceph_backup_offsite_dest | length > 0)
else 'DISABLED (empty - operator TODO)' }}.
Latest tarball:
{{ (ceph_backup_tarballs.files | map(attribute='path') | sort | last)
if (ceph_backup_run_now | bool
and (ceph_backup_tarballs.files | default([]) | length) > 0)
else 'none yet (timer will produce one)' }}
@@ -0,0 +1,13 @@
# {{ ansible_managed }}
# Oneshot capture of Ceph cluster config/topology state. Triggered by
# ceph-backup.timer; can also be run on demand: systemctl start ceph-backup.service
[Unit]
Description=Ceph cluster-state backup (config/topology to {{ ceph_backup_dir }})
After=network-online.target
Wants=network-online.target
[Service]
Type=oneshot
# Belt-and-suspenders: any file the script writes lands root-only.
UMask=0077
ExecStart={{ ceph_backup_script_path }}
@@ -0,0 +1,132 @@
#!/usr/bin/env bash
# {{ ansible_managed }}
# ceph-backup.sh - capture recoverable Ceph cluster state to a timestamped
# tarball under a local backup dir, prune old tarballs, optionally sync offsite.
#
# Runs on the bootstrap node as root and uses the on-node admin keyring; it
# embeds NO credentials. Installed + scheduled by ansible/ceph/backup-ceph.yml.
#
# Does NOT dump secret keyrings. The backup dir is root-only (0700) and the
# capture set is config/topology only: fsid, ceph config dump, monmap, osdmap,
# crushmap (+ decompiled), osd tree, orch ls/host ls, RGW realm/zonegroup/zone.
set -euo pipefail
# --- Config (rendered from ansible vars; environment overrides win) ---
BACKUP_DIR="${CEPH_BACKUP_DIR:-{{ ceph_backup_dir }}}"
RETENTION_DAYS="${CEPH_BACKUP_RETENTION_DAYS:-{{ ceph_backup_retention_days }}}"
# Offsite target: EMPTY by default. Operator decision - set to an rsync target
# (user@host:/path or rsync://...) or an S3 URI (s3://bucket/prefix) to ship
# offsite. Leaving it empty skips the offsite step entirely.
OFFSITE_DEST="${CEPH_BACKUP_OFFSITE_DEST:-{{ ceph_backup_offsite_dest }}}"
TS="$(date -u +%Y%m%dT%H%M%SZ)"
STAGE="$(mktemp -d "${TMPDIR:-/tmp}/ceph-backup.XXXXXX")"
OUT="${STAGE}/ceph-state-${TS}"
mkdir -p "${OUT}"
trap 'rm -rf "${STAGE}"' EXIT
log() { echo "ceph-backup: $*" >&2; }
# cap <outfile> <cmd...>: run a capture command; keep going if a subsystem is
# absent (e.g. a cluster with no RGW realm) so one gap does not abort the run.
cap() {
local out="$1"
shift
if "$@" >"${OUT}/${out}" 2>"${OUT}/${out}.err"; then
rm -f "${OUT}/${out}.err"
else
log "WARN: capture failed: ${out} (kept ${out}.err)"
fi
}
# --- Capture set ---
cap fsid.txt ceph fsid
cap ceph-status.txt ceph status
cap config-dump.json ceph config dump --format json
cap config-dump.txt ceph config dump
cap osd-tree.txt ceph osd tree
cap osd-tree.json ceph osd tree --format json
cap osd-dump.json ceph osd dump --format json
cap mon-dump.json ceph mon dump --format json
cap orch-services.yaml ceph orch ls --format yaml
cap orch-hosts.yaml ceph orch host ls --format yaml
# Binary maps + human-readable decodes (decode tools are best-effort).
if ceph mon getmap -o "${OUT}/monmap.bin" 2>"${OUT}/monmap.err"; then
rm -f "${OUT}/monmap.err"
if command -v monmaptool >/dev/null 2>&1; then
monmaptool --print "${OUT}/monmap.bin" >"${OUT}/monmap.txt" 2>/dev/null || true
fi
else
log "WARN: monmap capture failed (kept monmap.err)"
fi
if ceph osd getmap -o "${OUT}/osdmap.bin" 2>"${OUT}/osdmap.err"; then
rm -f "${OUT}/osdmap.err"
if command -v osdmaptool >/dev/null 2>&1; then
osdmaptool --print "${OUT}/osdmap.bin" >"${OUT}/osdmap.txt" 2>/dev/null || true
fi
else
log "WARN: osdmap capture failed (kept osdmap.err)"
fi
if ceph osd getcrushmap -o "${OUT}/crushmap.bin" 2>"${OUT}/crushmap.err"; then
rm -f "${OUT}/crushmap.err"
if command -v crushtool >/dev/null 2>&1; then
crushtool -d "${OUT}/crushmap.bin" -o "${OUT}/crushmap.txt" 2>/dev/null || true
fi
else
log "WARN: crushmap capture failed (kept crushmap.err)"
fi
# RGW multisite config (absent on clusters with no realm - non-fatal).
cap rgw-realm.json radosgw-admin realm get
cap rgw-zonegroup.json radosgw-admin zonegroup get
cap rgw-zone.json radosgw-admin zone get
# Provenance.
{
echo "captured_at=${TS}"
echo "host=$(hostname -f 2>/dev/null || hostname)"
echo "ceph_version=$(ceph --version 2>/dev/null || echo unknown)"
} >"${OUT}/MANIFEST.txt"
# --- Package ---
mkdir -p "${BACKUP_DIR}"
chmod 0700 "${BACKUP_DIR}"
TARBALL="${BACKUP_DIR}/ceph-state-${TS}.tar.gz"
tar -C "${STAGE}" -czf "${TARBALL}" "ceph-state-${TS}"
chmod 0600 "${TARBALL}"
log "wrote ${TARBALL}"
# --- Prune ---
if [ "${RETENTION_DAYS}" -gt 0 ] 2>/dev/null; then
find "${BACKUP_DIR}" -maxdepth 1 -type f -name 'ceph-state-*.tar.gz' \
-mtime "+${RETENTION_DAYS}" -print -delete \
| while read -r p; do log "pruned ${p}"; done || true
fi
# --- Optional offsite sync (skipped unless OFFSITE_DEST is set) ---
if [ -n "${OFFSITE_DEST}" ]; then
case "${OFFSITE_DEST}" in
s3://*)
if command -v aws >/dev/null 2>&1; then
aws s3 cp "${TARBALL}" "${OFFSITE_DEST%/}/$(basename "${TARBALL}")" \
&& log "synced to ${OFFSITE_DEST}"
else
log "WARN: OFFSITE_DEST is S3 but the aws CLI is not installed - skipping offsite sync"
fi
;;
*)
if command -v rsync >/dev/null 2>&1; then
rsync -a "${TARBALL}" "${OFFSITE_DEST}" && log "synced to ${OFFSITE_DEST}"
else
log "WARN: rsync not installed - skipping offsite sync"
fi
;;
esac
else
log "offsite sync disabled (CEPH_BACKUP_OFFSITE_DEST empty)"
fi
log "done"
@@ -0,0 +1,13 @@
# {{ ansible_managed }}
# Runs ceph-backup.service on a schedule. Persistent=true catches up a missed
# run (host powered off at the scheduled time) on the next boot.
[Unit]
Description=Scheduled Ceph cluster-state backup
[Timer]
OnCalendar={{ ceph_backup_oncalendar }}
RandomizedDelaySec={{ ceph_backup_randomized_delay }}
Persistent=true
[Install]
WantedBy=timers.target
@@ -2,7 +2,7 @@
# Ceph release train. Upstream publishes per-release apt trees at
# download.ceph.com/debian-<release>/, and each tree has subdirectories
# per Debian codename. prerequisites.yml pins the codename to 'bookworm'
# in the sources.list entry — when upgrading the base OS to Trixie,
# in the sources.list entry - when upgrading the base OS to Trixie,
# flip the codename there. This split keeps the Ceph release and the
# Debian release independently versionable.
ceph_release: tentacle
@@ -16,7 +16,106 @@ cephadm_install_method: repo # 'repo' or 'curl'
# splitting under load which tanks performance during the first fill.
# These values are sized for ~36 OSDs. Scale proportionally for larger
# clusters.
ceph_pg_init_data: 128 # EC data pool — bulk of I/O
ceph_pg_init_data: 128 # EC data pool - bulk of I/O
ceph_pg_init_index: 16 # bucket index
ceph_pg_init_non_ec: 16 # multipart uploads
ceph_pg_init_meta: 8 # .rgw.root, .meta, .log, .control
# RGW EC data-pool min_size (write-availability floor). Ceph defaults an EC pool
# to min_size = k+1, which is the correct floor: at k+1 up shards a PG stays
# writable, so writes tolerate up to m-1 host losses while reads still tolerate m.
# Pin it EXPLICITLY (get-then-set in rgw.yml) so the value is documented and
# cannot drift. NEVER set below k+1: permitting writes at k shards risks data loss
# if another shard is lost mid-recovery. Derived from ceph_rgw_ec_k so it tracks
# whatever profile a cluster runs (spice 16+4 -> 17, sietch 8+3 -> 9).
ceph_rgw_ec_min_size: "{{ (ceph_rgw_ec_k | int) + 1 }}"
# MGR daemon count (1 active + standbys). cephadm's default and upstream guidance
# is a small fixed count; deploying one per host (e.g. 47 on spice) is wasteful and
# non-standard. count:N lets cephadm place them (caps at the host count on small
# clusters like sietch).
ceph_mgr_count: 3
# --- Service bind/advertise policy (per-cluster; fabric-only on spice) ---
# ceph_service_ip (the address Ceph point-services advertise + bind on: mon
# --mon-ip, the cephadm host-add address, the dashboard monitoring URLs) is
# defined in each cluster's group_vars, NOT here. join/rgw/monitoring read it
# through hostvars[<host>], and role defaults are INVISIBLE via hostvars (a role
# default raises "HostVarsVars object has no attribute ceph_service_ip" and aborts
# the deploy), whereas group_vars resolve through hostvars per host. spice sets it
# to the fabric ceph_public_ip; sietch (flat) to bond_ip. It is NOT the ansible/SSH
# connection address (that stays bond_ip / ansible_host until the WAN is retired).
#
# CIDR(s) cephadm restricts daemon binds to, injected as the `networks:` field of
# the RGW + monitoring service specs. mon/mgr/OSD already bind per public_network/
# cluster_network; this closes the remaining daemons (beast RGW, prometheus,
# grafana, alertmanager, node-exporter, ceph-exporter) that otherwise bind 0.0.0.0.
# EMPTY by default: no `networks:` field is emitted and the monitoring re-spec is
# skipped, so daemons bind every interface exactly as cephadm ships them - flat
# clusters like sietch are untouched. A cluster that must be fabric-only (spice)
# opts in by setting this to its fabric CIDR(s) in group_vars.
ceph_bind_networks: []
# --- Device-class CRUSH pool placement (per-cluster; opt-in) ---
# Pools listed here are pinned to the named device-class CRUSH rule after RGW is
# up. EMPTY by default so a cluster's pools stay on whatever rule they already
# carry - re-homing a live pool triggers a full rebalance under load, so this is
# never applied implicitly. A cluster with a dedicated fast tier (spice's ssd-osd
# NVMe) opts in via group_vars. The rule names must match crush-rules.yml.
ceph_rgw_ssd_crush_rule: replicated_ssd
ceph_rgw_hdd_crush_rule: replicated_hdd
ceph_rgw_ssd_pools: []
ceph_rgw_hdd_pools: []
# --- RGW ingress: health-checked S3 VIP (haproxy + keepalived; per-cluster; opt-in) ---
# A cephadm `ingress` service that puts a keepalived-managed floating VIP with
# haproxy health checks in front of RGW, in TLS TERMINATION mode (haproxy
# `mode http`: haproxy terminates the client TLS on the VIP with the reused RGW
# self-signed cert, then re-encrypts to the RGW backends). A dead RGW node is then
# dropped from the VIP instead of blackholing ~1/N of S3 connections. DEFAULT OFF:
# no ingress spec is rendered or applied, so flat/dev clusters (sietch) and any
# cluster that has not opted in are a complete no-op.
#
# (Terminate, not L4 passthrough, because passthrough for an RGW backend needs
# IngressSpec.use_tcp_mode_over_rgw, a post-20.2.2 upstream field absent on the
# pinned 20.2.2 -- source-verified; an unknown spec key makes `orch apply`
# TypeError. Terminate is the only health-checked ingress mode available here.)
#
# A cluster opts in by setting ceph_rgw_ingress_enabled: true AND
# ceph_rgw_ingress_vip in its group_vars. Three hard preconditions, all asserted
# at apply time:
# 1. ceph_rgw_ingress_vip must be set (the VIP the clients hit).
# 2. ceph_bind_networks must be non-empty - otherwise beast binds
# 0.0.0.0:<ceph_rgw_port> and the co-located haproxy VIP bind on the same
# port collides (EADDRINUSE). Restricted beast binds a per-node IP, the VIP
# is a distinct socket, so they coexist.
# 3. ceph_rgw_ssl must be true - haproxy terminates with the RGW self-signed
# cert (rgw_ssl_cert_combined_pem), which only exists when RGW ssl is on.
ceph_rgw_ingress_enabled: false
# Floating VIP the S3 clients hit (bare IP; the prefix is added below). Empty by
# default; a cluster that opts in sets this to an address on its RGW-facing
# network (e.g. spice fabric public 10.40.20.250).
ceph_rgw_ingress_vip: ""
# CIDR prefix of the network the VIP lives on - MUST match the prefix of the
# public/RGW-facing network (spice fabric public = /23) so keepalived binds the
# VIP to the right interface.
ceph_rgw_ingress_vip_prefix: 23
# Port clients hit on the VIP. Must match ceph_rgw_port (the S3/beast port).
ceph_rgw_ingress_frontend_port: 443
# haproxy stats/monitor port (load-balancer status; not the S3 data path).
ceph_rgw_ingress_monitor_port: 1967
# Number of haproxy + keepalived instances cephadm spreads across nodes (VRRP
# needs >= 2 for failover; 3 tolerates one node loss with quorum to spare).
ceph_rgw_ingress_count: 3
# --- Alert delivery (per-cluster; opt-in) ---
# cephadm auto-deploys alertmanager but with NO receiver, so the ~89 Prometheus
# alert rules (OSD down, host down, PG degraded, near-full) fire into the default
# null route and no human is notified. Set this to one or more webhook URLs (a
# chat incoming-webhook, or a pager like Opsgenie/PagerDuty, or a FUTO endpoint)
# and the alertmanager service spec routes all alerts there via cephadm's
# user_data.default_webhook_urls. Values may be op:// refs resolved before the
# play. EMPTY by default: no receiver is wired (cephadm default is unchanged), and
# verify.yml warns that alert delivery is not configured. TODO(operator): set the
# destination for spice.
ceph_alertmanager_webhook_urls: []
@@ -32,7 +32,7 @@
ansible.builtin.shell: |
set -o pipefail
cephadm bootstrap \
--mon-ip {{ bond_ip }} \
--mon-ip {{ ceph_service_ip }} \
--initial-dashboard-user {{ ceph_dashboard_user }} \
--initial-dashboard-password "$(cat /tmp/.ceph-dashboard-pw)" \
--dashboard-password-noupdate \
@@ -75,7 +75,7 @@
# `--dashboard-password-noupdate` means it is never re-applied afterwards. If a
# cluster was ever bootstrapped out-of-band (or before this var existed),
# cephadm generated a *random* admin password and no playbook run would ever
# correct it — the vaulted password stays aspirational. These tasks reconcile
# correct it; the vaulted password stays aspirational. These tasks reconcile
# the running cluster to `ceph_dashboard_password` on every converge, so the
# vault is the single source of truth for both initial set and rotation.
- name: Write dashboard password to temp file (reconcile)
@@ -38,7 +38,7 @@
ansible.builtin.shell: |
set -euo pipefail
HOST="{{ hostvars[item]['hostname_short'] }}"
ADDR="{{ hostvars[item]['bond_ip'] }}"
ADDR="{{ hostvars[item]['ceph_service_ip'] }}"
echo "Testing SSH to $HOST ($ADDR)..."
ceph cephadm check-host "$HOST" "$ADDR" 2>&1 || {
echo "WARN: ceph cephadm check-host failed, testing raw SSH..."
@@ -75,10 +75,10 @@
grep -q '"{{ hostvars[item]['hostname_short'] }}"'; then
echo "Host {{ hostvars[item]['hostname_short'] }} already in cluster"
else
echo "Adding {{ hostvars[item]['hostname_short'] }} ({{ hostvars[item]['bond_ip'] }})"
echo "Adding {{ hostvars[item]['hostname_short'] }} ({{ hostvars[item]['ceph_service_ip'] }})"
ceph orch host add \
{{ hostvars[item]['hostname_short'] }} \
{{ hostvars[item]['bond_ip'] }} 2>&1
{{ hostvars[item]['ceph_service_ip'] }} 2>&1
fi
args:
executable: /bin/bash
@@ -2,33 +2,122 @@
# Phase 4.5: Ensure Ceph block.db LVM is set up on each node (sietch-shape only)
#
# This task is sietch-shape-specific: it assumes dual SAS-attached SSDs each
# with partition 5 → its own VG, with 6 db-slot LVs per VG (12 total). It
# ensures PVs, VGs, and LVs exist. Idempotent — skips if already present.
# with partition 5 -> its own VG, with 6 db-slot LVs per VG (12 total). It
# ensures PVs, VGs, and LVs exist. Idempotent - skips if already present.
# Provides a recovery path if the LVM was destroyed by `mise run destroy`.
#
# Painbox-shape (Hetzner SX295, NVMe RAID-1 → single vg0 with all LVs created
# by installimage post-install) skips the whole block — its host_vars don't
# Painbox-shape (Hetzner SX295, NVMe RAID-1 -> single vg0 with all LVs created
# by installimage post-install) skips the whole block - its host_vars don't
# define `sas_path_prefix` so the gate below short-circuits.
#
# Required host_vars (sietch-shape only):
# sas_path_prefix — by-path prefix for SSD discovery (e.g., pci-0000:02:00.0-sas-exp...)
# ssd1_phy, ssd2_phy — PHY slot numbers for the two SSDs
# ceph_db_vg1, ceph_db_vg2 — VG names for SSD1/SSD2 partition 5
# sas_path_prefix - by-path prefix for SSD discovery (e.g., pci-0000:02:00.0-sas-exp...)
# ssd1_phy, ssd2_phy - PHY slot numbers for the two SSDs
# ceph_db_vg1, ceph_db_vg2 - VG names for SSD1/SSD2 partition 5
#
# VG mapping: ceph_db_vg1 on SSD1 partition 5, ceph_db_vg2 on SSD2 partition 5
# LV naming: db-slot0..5 on VG1, db-slot6..11 on VG2 (one per HDD OSD)
- name: Skip in-role LVM setup (no sas_path_prefix — externally-managed LVM)
# NVMe-RAID shape (spice): a single vg0 on the NVMe RAID-1 already carries the
# OS LVs from installimage. Create the block.db LVs (one per HDD OSD) + one
# ssd-osd LV here at converge, before osds.yml. This replaces painbox's fragile
# installimage -x chroot post-install with an idempotent, observable ansible
# step. Gated on ceph_db_lvs_per_node (spice sets it; genuinely
# externally-managed hosts do not).
- name: Set up LVM (NVMe-RAID single-vg0 topology)
when:
- sas_path_prefix is undefined
- ceph_db_lvs_per_node is defined
# DATA SAFETY: never let lvol SHRINK an existing LV -- reconciling a drifted
# db-slot down would truncate a live block.db and corrupt the OSD. shrink:false
# means lvol still creates and may grow, but leaves a larger existing LV alone
# (the old shell skipped existing LVs entirely). opts -Wy wipes stale signatures
# on (re)created LVs, preserving the destroy->recreate recovery-path freshness.
module_defaults:
community.general.lvol:
shrink: false
opts: "-Wy"
block:
- name: Assert vg0 exists (created by installimage)
ansible.builtin.command: vgs {{ ceph_db_vg }}
changed_when: false
- name: Create block.db LVs on vg0 (fixed size, one per HDD OSD)
community.general.lvol:
vg: "{{ ceph_db_vg }}"
lv: "db-slot{{ item }}"
size: "{{ ceph_db_lv_size }}"
state: present
loop: "{{ range(0, ceph_db_lvs_per_node | int) | list }}"
# ssd-osd is sized as (vg0 free - reserve); lvol has no "free minus N GiB"
# primitive, so compute the GiB with a read-only shell and feed it to lvol.
# Gate the compute + create on the LV being absent: once ssd-osd exists, vg0
# free == reserve, so recomputing would yield <= 0 and the original shell's
# FATAL guard would (correctly) refuse -- and an unforced lvol resize would
# fail. The existence check preserves that idempotency exactly.
- name: Check whether the ssd-osd LV already exists on vg0
ansible.builtin.command: lvs {{ ceph_db_vg }}/{{ ceph_ssd_osd_lv }}
register: vg0_ssd_check
changed_when: false
failed_when: false
- name: Compute the ssd-osd size (vg0 free minus reserve, in GiB)
ansible.builtin.shell: |
set -o pipefail
free_g=$(vgs {{ ceph_db_vg }} --noheadings --nosuffix --units g -o vg_free | tr -d ' ' | cut -d. -f1)
echo $(( free_g - {{ ceph_ssd_osd_reserve_gib }} ))
args:
executable: /bin/bash
register: vg0_ssd_size
changed_when: false
when: vg0_ssd_check.rc != 0
- name: Fail when vg0 free is at or below the reserve
ansible.builtin.fail:
msg: >-
FATAL: {{ ceph_db_vg }} free minus reserve
{{ ceph_ssd_osd_reserve_gib }}G leaves no room for {{ ceph_ssd_osd_lv }}
when:
- vg0_ssd_check.rc != 0
- (vg0_ssd_size.stdout | int) <= 0
- name: Create the ssd-osd LV from vg0 remainder minus reserve
community.general.lvol:
vg: "{{ ceph_db_vg }}"
lv: "{{ ceph_ssd_osd_lv }}"
size: "{{ vg0_ssd_size.stdout | int }}G"
state: present
when: vg0_ssd_check.rc != 0
- name: Verify vg0 LVM layout
ansible.builtin.command: lvs {{ ceph_db_vg }} -o lv_name,lv_size --noheadings
register: vg0_lvm_state
changed_when: false
- name: Show vg0 LVM layout
ansible.builtin.debug:
msg: "{{ vg0_lvm_state.stdout_lines }}"
# Genuinely externally-managed LVM (neither sietch dual-SSD nor spice NVMe-RAID).
- name: Skip in-role LVM setup (externally-managed LVM)
ansible.builtin.debug:
msg: >-
LVM is externally managed on this host (e.g., Hetzner installimage
post-install creates vg0 + db-slots on a RAID-1 NVMe). Skipping the
sietch-shape dual-SSD-VG setup. If LVM is missing on this host,
LVM is externally managed on this host (no sas_path_prefix and no
ceph_db_lvs_per_node). Skipping in-role LVM setup. If LVM is missing,
reprovision via the cluster's installimage flow.
when: sas_path_prefix is undefined
when:
- sas_path_prefix is undefined
- ceph_db_lvs_per_node is undefined
- name: Set up LVM (sietch-shape dual-SSD-VG topology)
when: sas_path_prefix is defined
# DATA SAFETY: shrink:false so lvol never truncates a live block.db LV on drift
# (it may still create/grow); opts -Wy wipes stale signatures on (re)created LVs.
module_defaults:
community.general.lvol:
shrink: false
opts: "-Wy"
block:
- name: Resolve SSD1 device path
ansible.builtin.command: >
@@ -54,51 +143,64 @@
changed_when: false
failed_when: false
- name: Create PV and VG on SSD1 partition 5
ansible.builtin.shell: |
PART="{{ ssd1_dev.stdout }}5"
wipefs -af "$PART"
pvcreate -f "$PART" && vgcreate {{ ceph_db_vg1 }} "$PART"
# lvg runs "pvcreate -f" but not "--yes", so it will not clear a stale
# foreign (ceph/ext) signature on a reused partition. Preserve the original
# wipefs -af, gated exactly as before on the VG being absent, so we never
# wipe a live PV. lvg then creates the PV+VG idempotently.
- name: Wipe stale signatures on SSD1 partition 5
ansible.builtin.command: wipefs -af {{ ssd1_dev.stdout }}5
when: vg1_check.rc != 0
changed_when: true
- name: Create PV and VG on SSD2 partition 5
ansible.builtin.shell: |
PART="{{ ssd2_dev.stdout }}5"
wipefs -af "$PART"
pvcreate -f "$PART" && vgcreate {{ ceph_db_vg2 }} "$PART"
- name: Create PV and VG on SSD1 partition 5
community.general.lvg:
vg: "{{ ceph_db_vg1 }}"
pvs: "{{ ssd1_dev.stdout }}5"
state: present
- name: Wipe stale signatures on SSD2 partition 5
ansible.builtin.command: wipefs -af {{ ssd2_dev.stdout }}5
when: vg2_check.rc != 0
changed_when: true
- name: Create block.db LVs on VG1 (slots 0-4 at 240G, slot 5 gets remainder)
ansible.builtin.shell: |
if lvs {{ ceph_db_vg1 }}/db-slot{{ item }} 2>/dev/null; then
echo "db-slot{{ item }} already exists"
else
{% if item < 5 %}
lvcreate --yes -Wy -L {{ ceph_db_lv_size }} -n db-slot{{ item }} {{ ceph_db_vg1 }}
{% else %}
lvcreate --yes -Wy -l 100%FREE -n db-slot{{ item }} {{ ceph_db_vg1 }}
{% endif %}
fi
loop: [0, 1, 2, 3, 4, 5]
register: vg1_lv_results
changed_when: "'already exists' not in (vg1_lv_results.stdout | default(''))"
- name: Create PV and VG on SSD2 partition 5
community.general.lvg:
vg: "{{ ceph_db_vg2 }}"
pvs: "{{ ssd2_dev.stdout }}5"
state: present
- name: Create block.db LVs on VG2 (slots 6-10 at 240G, slot 11 gets remainder)
ansible.builtin.shell: |
if lvs {{ ceph_db_vg2 }}/db-slot{{ item }} 2>/dev/null; then
echo "db-slot{{ item }} already exists"
else
{% if item < 11 %}
lvcreate --yes -Wy -L {{ ceph_db_lv_size }} -n db-slot{{ item }} {{ ceph_db_vg2 }}
{% else %}
lvcreate --yes -Wy -l 100%FREE -n db-slot{{ item }} {{ ceph_db_vg2 }}
{% endif %}
fi
loop: [6, 7, 8, 9, 10, 11]
register: vg2_lv_results
changed_when: "'already exists' not in (vg2_lv_results.stdout | default(''))"
# Fixed slots first, then the 100%FREE remainder, so the remainder LV only
# claims the space left after the fixed LVs (matches the sequential order of
# the original single-loop shell).
- name: Create fixed-size block.db LVs on VG1 (slots 0-4 at 240G)
community.general.lvol:
vg: "{{ ceph_db_vg1 }}"
lv: "db-slot{{ item }}"
size: "{{ ceph_db_lv_size }}"
state: present
loop: [0, 1, 2, 3, 4]
- name: Create remainder block.db LV on VG1 (slot 5, 100%FREE)
community.general.lvol:
vg: "{{ ceph_db_vg1 }}"
lv: db-slot5
size: 100%FREE
state: present
- name: Create fixed-size block.db LVs on VG2 (slots 6-10 at 240G)
community.general.lvol:
vg: "{{ ceph_db_vg2 }}"
lv: "db-slot{{ item }}"
size: "{{ ceph_db_lv_size }}"
state: present
loop: [6, 7, 8, 9, 10]
- name: Create remainder block.db LV on VG2 (slot 11, 100%FREE)
community.general.lvol:
vg: "{{ ceph_db_vg2 }}"
lv: db-slot11
size: 100%FREE
state: present
# --- SSD OSD data LVs (partition 6) ---
# part6 of each SSD is the SSD-class OSD data device. cephadm cannot deploy
@@ -120,46 +222,50 @@
failed_when: false
when: ceph_ssd_vg2 is defined
- name: Create PV and VG on SSD1 partition 6 (SSD OSD)
ansible.builtin.shell: |
PART="{{ ssd1_dev.stdout }}6"
wipefs -af "$PART"
pvcreate -f "$PART" && vgcreate {{ ceph_ssd_vg1 }} "$PART"
# Same wipefs-then-lvg safety as partition 5, gated on the SSD-OSD VG being
# requested (ceph_ssd_vg1/2 defined) and absent.
- name: Wipe stale signatures on SSD1 partition 6 (SSD OSD)
ansible.builtin.command: wipefs -af {{ ssd1_dev.stdout }}6
when:
- ceph_ssd_vg1 is defined
- ssd_vg1_check.rc | default(1) != 0
changed_when: true
- name: Create PV and VG on SSD2 partition 6 (SSD OSD)
ansible.builtin.shell: |
PART="{{ ssd2_dev.stdout }}6"
wipefs -af "$PART"
pvcreate -f "$PART" && vgcreate {{ ceph_ssd_vg2 }} "$PART"
- name: Create PV and VG on SSD1 partition 6 (SSD OSD)
community.general.lvg:
vg: "{{ ceph_ssd_vg1 }}"
pvs: "{{ ssd1_dev.stdout }}6"
state: present
when: ceph_ssd_vg1 is defined
- name: Wipe stale signatures on SSD2 partition 6 (SSD OSD)
ansible.builtin.command: wipefs -af {{ ssd2_dev.stdout }}6
when:
- ceph_ssd_vg2 is defined
- ssd_vg2_check.rc | default(1) != 0
changed_when: true
- name: Create PV and VG on SSD2 partition 6 (SSD OSD)
community.general.lvg:
vg: "{{ ceph_ssd_vg2 }}"
pvs: "{{ ssd2_dev.stdout }}6"
state: present
when: ceph_ssd_vg2 is defined
- name: Create SSD OSD LV on VG1 (100%FREE)
ansible.builtin.shell: |
if lvs {{ ceph_ssd_vg1 }}/osd-ssd 2>/dev/null; then
echo "osd-ssd already exists"
else
lvcreate --yes -Wy -l 100%FREE -n osd-ssd {{ ceph_ssd_vg1 }}
fi
register: ssd_lv1_result
changed_when: "'already exists' not in (ssd_lv1_result.stdout | default(''))"
community.general.lvol:
vg: "{{ ceph_ssd_vg1 }}"
lv: osd-ssd
size: 100%FREE
state: present
when: ceph_ssd_vg1 is defined
- name: Create SSD OSD LV on VG2 (100%FREE)
ansible.builtin.shell: |
if lvs {{ ceph_ssd_vg2 }}/osd-ssd 2>/dev/null; then
echo "osd-ssd already exists"
else
lvcreate --yes -Wy -l 100%FREE -n osd-ssd {{ ceph_ssd_vg2 }}
fi
register: ssd_lv2_result
changed_when: "'already exists' not in (ssd_lv2_result.stdout | default(''))"
community.general.lvol:
vg: "{{ ceph_ssd_vg2 }}"
lv: osd-ssd
size: 100%FREE
state: present
when: ceph_ssd_vg2 is defined
- name: Verify LVM setup
@@ -1,5 +1,5 @@
---
# Phase 5.8: Monitoring stack — dashboard integration
# Phase 5.8: Monitoring stack - dashboard integration
#
# cephadm auto-deploys the full monitoring stack (node-exporter,
# ceph-exporter, prometheus, alertmanager, grafana) during bootstrap.
@@ -62,26 +62,129 @@
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
# --- Step 2.5: Re-spec the monitoring stack (fabric bind + alert delivery) ---
# cephadm auto-deployed the stack binding all interfaces and with NO alertmanager
# receiver. Re-apply the spec to (a) restrict the bind to ceph_bind_networks
# (fabric-only on spice, so ceph never listens on the WAN) and (b) route alerts to
# ceph_alertmanager_webhook_urls. Placement matches cephadm's defaults, so no
# daemon relocates. Each piece is independently gated in the template, and the
# whole re-spec is SKIPPED when BOTH are empty (flat clusters with no alerting, like
# sietch) - the auto-deployed stack is then left exactly as cephadm shipped it.
- name: Render monitoring service spec (fabric bind + alert receiver)
ansible.builtin.template:
src: monitoring-spec.yaml.j2
dest: /etc/ceph/monitoring-spec.yaml
owner: root
group: root
mode: '0644'
when:
- inventory_hostname in groups['ceph_bootstrap']
- (ceph_bind_networks | default([]) | length) > 0
or (ceph_alertmanager_webhook_urls | default([]) | length) > 0
- name: Apply monitoring service spec
ansible.builtin.command: ceph orch apply -i /etc/ceph/monitoring-spec.yaml
when:
- inventory_hostname in groups['ceph_bootstrap']
- (ceph_bind_networks | default([]) | length) > 0
or (ceph_alertmanager_webhook_urls | default([]) | length) > 0
changed_when: false
# --- Step 2.6: Pin the mgr dashboard + prometheus module bind to the fabric ---
#
# The dashboard (8443) and the prometheus mgr module (9283) run INSIDE the active
# ceph-mgr, not as standalone cephadm daemons, so the monitoring re-spec above
# does not reach them -- they bind :: (every interface, WAN included) and are
# closed only by nftables. Pin each mgr instance's own server_addr to that
# instance's fabric IP (ceph_service_ip) so whichever mgr is active listens on
# the fabric, never the WAN.
#
# PER-INSTANCE, not global: the active mgr fails over across the mgr hosts and
# each host has a DIFFERENT fabric IP (ceph_service_ip = 10.40.20.<index>). A
# single global mgr/dashboard/server_addr would bind a foreign IP after failover
# and the dashboard could not start. Ceph resolves the localized key
# mgr/<module>/<mgr_id>/server_addr ahead of the global one (the module reads it
# via get_localized_module_option), and each mgr reads ITS OWN value when its
# module starts serving (on mgr start / failover), so per-instance binding is
# inherently failover-safe. The mgr_id is the cephadm daemon name (<host>.<rand>,
# e.g. spice-ceph-adelia.abcdef), NOT hostname_short -- it is discovered at
# runtime from `ceph orch ps`, and the daemon's host is mapped back to that
# host's ceph_service_ip.
#
# TIMING: the new bind takes effect the next time each mgr (re)starts its module
# (mgr restart, failover, or upgrade); nftables keeps the WAN closed until then.
# We deliberately do NOT force a live module reload (disable/enable) -- that would
# drop the dashboard mid-converge for no durability gain, and the per-instance
# config can never leave the dashboard unbindable the way a global one could.
#
# Opt-in: only when ceph_bind_networks is non-empty. sietch (empty) is a no-op
# and keeps cephadm's bind-all default.
- name: Pin mgr dashboard + prometheus module bind to the fabric (per mgr instance)
ansible.builtin.shell: |
set -o pipefail
python3 - <<'PYEOF'
import json, subprocess, sys
fabric_ip = {
{% for h in groups['ceph_nodes'] %}
"{{ hostvars[h]['hostname_short'] }}": "{{ hostvars[h]['ceph_service_ip'] }}",
{% endfor %}
}
mgrs = json.loads(subprocess.check_output(
["ceph", "orch", "ps", "--daemon-type", "mgr", "--format", "json"]))
cfg = json.loads(subprocess.check_output(
["ceph", "config", "dump", "--format", "json"]))
current = {(c["section"], c["name"]): c["value"] for c in cfg}
changed = False
for d in mgrs:
mgr_id = d["daemon_id"]
host = d["hostname"]
ip = fabric_ip.get(host)
if not ip:
sys.stderr.write(
"WARNING: no fabric IP for mgr host %s, skipping\n" % host)
continue
for mod in ("dashboard", "prometheus"):
key = "mgr/%s/%s/server_addr" % (mod, mgr_id)
if current.get(("mgr", key)) != ip:
subprocess.check_call(["ceph", "config", "set", "mgr", key, ip])
print("CHANGED %s -> %s" % (key, ip))
changed = True
print("CHANGED" if changed else "ok")
PYEOF
args:
executable: /bin/bash
when:
- inventory_hostname in groups['ceph_bootstrap']
- (ceph_bind_networks | default([]) | length) > 0
register: mgr_module_bind
changed_when: "'CHANGED' in (mgr_module_bind.stdout | default(''))"
# --- Step 3: Configure dashboard integrations ---
# The dashboard runs on the active mgr and reaches these services on the fabric;
# advertise ceph_service_ip (fabric on spice), not bond_ip (the 1G WAN), so the
# URLs stay inside ceph_firewall_trusted_networks after harden.
- name: Configure dashboard Prometheus URL
ansible.builtin.command: >
ceph dashboard set-prometheus-api-host
http://{{ hostvars[groups['ceph_bootstrap'][0]]['bond_ip'] }}:{{ ceph_prometheus_port }}
http://{{ hostvars[groups['ceph_bootstrap'][0]]['ceph_service_ip'] }}:{{ ceph_prometheus_port }}
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Configure dashboard Alertmanager URL
ansible.builtin.command: >
ceph dashboard set-alertmanager-api-host
http://{{ hostvars[groups['ceph_bootstrap'][0]]['bond_ip'] }}:{{ ceph_alertmanager_port }}
http://{{ hostvars[groups['ceph_bootstrap'][0]]['ceph_service_ip'] }}:{{ ceph_alertmanager_port }}
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Configure dashboard Grafana URL
ansible.builtin.command: >
ceph dashboard set-grafana-api-url
https://{{ hostvars[groups['ceph_bootstrap'][0]]['bond_ip'] }}:{{ ceph_grafana_port }}
https://{{ hostvars[groups['ceph_bootstrap'][0]]['ceph_service_ip'] }}:{{ ceph_grafana_port }}
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
@@ -159,9 +262,9 @@
ceph orch ls --service-type ceph-exporter
echo ""
echo "=== Endpoints ==="
echo "Prometheus: http://{{ bond_ip }}:{{ ceph_prometheus_port }}"
echo "Grafana: https://{{ bond_ip }}:{{ ceph_grafana_port }}"
echo "Alertmanager: http://{{ bond_ip }}:{{ ceph_alertmanager_port }}"
echo "Prometheus: http://{{ ceph_service_ip }}:{{ ceph_prometheus_port }}"
echo "Grafana: https://{{ ceph_service_ip }}:{{ ceph_grafana_port }}"
echo "Alertmanager: http://{{ ceph_service_ip }}:{{ ceph_alertmanager_port }}"
args:
executable: /bin/bash
register: monitoring_status
@@ -52,17 +52,17 @@
ansible.builtin.shell: |
set -o pipefail
EXPECTED={{
groups['ceph_nodes']
(groups['ceph_nodes'] | intersect(ansible_play_hosts))
| map('extract', hostvars)
| map(attribute='ceph_hdd_osds', default=[])
| map('length') | sum
+
groups['ceph_nodes']
(groups['ceph_nodes'] | intersect(ansible_play_hosts))
| map('extract', hostvars)
| map(attribute='ceph_ssd_osds', default=[])
| map('length') | sum
}}
echo "Expecting $EXPECTED OSDs total across the cluster"
echo "Expecting $EXPECTED OSDs across the nodes in this run"
for i in $(seq 1 180); do
ACTUAL=$(ceph osd stat --format json 2>/dev/null \
| python3 -c "import sys, json; print(json.load(sys.stdin).get('num_osds', 0))" \
@@ -6,7 +6,8 @@
# 2. Else if <=2 active hosts: 1 MON (bootstrap only, avoids 2-MON fragility)
# 3. Else: MON on all active hosts
#
# MGR always deploys on all active hosts (standbys are harmless).
# MGR deploys a capped count (ceph_mgr_count, default 3: 1 active + standbys),
# not one per host -- cephadm schedules them and caps at the host count.
- name: Determine active cluster hosts
ansible.builtin.shell: |
@@ -45,10 +46,10 @@
MON placement: {{ mon_hosts }}
({{ 'explicit ceph_mon group'
if groups['ceph_mon'] | default([]) | length > 0
else ('single MON — avoids 2-mon quorum fragility'
else ('single MON - avoids 2-mon quorum fragility'
if active_hosts.stdout.split(',') | length <= 2
else 'MON on all hosts') }}).
MGR placement: {{ active_hosts.stdout }} (all hosts).
MGR placement: count:{{ ceph_mgr_count }}.
when:
- inventory_hostname in groups['ceph_bootstrap']
- mon_hosts is defined
@@ -61,9 +62,9 @@
- mon_hosts | default('') | length > 0
changed_when: false
- name: Deploy MGR on all active hosts
- name: Deploy MGR at a capped count (not one per host)
ansible.builtin.command: >
ceph orch apply mgr --placement="{{ active_hosts.stdout }}"
ceph orch apply mgr --placement="count:{{ ceph_mgr_count }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- active_hosts.stdout | default('') | length > 0
+262 -43
View File
@@ -46,6 +46,45 @@
- create_data_pool is changed
changed_when: true
- name: Mark the EC data pool as bulk (autoscaler sizing floor)
# bulk=true tells the pg_autoscaler to target a pg_num sized for the pool's
# expected full capacity, so it will not shrink the pre-seeded pg_num
# (ceph_pg_init_data) back toward 1 during the first fill. Idempotent: only sets
# when not already true.
ansible.builtin.shell: |
set -o pipefail
[ "$(ceph osd pool get {{ ceph_rgw_data_pool }} bulk | awk '{print $2}')" = "true" ] || { ceph osd pool set {{ ceph_rgw_data_pool }} bulk true; echo CHANGED; }
args:
executable: /bin/bash
when:
- inventory_hostname in groups['ceph_bootstrap']
- (create_data_pool is changed) or (ceph_rgw_data_pool in pool_list.stdout)
register: bulk_flag_result
changed_when: "'CHANGED' in (bulk_flag_result.stdout | default(''))"
- name: Pin EC data pool min_size to k+1 ({{ ceph_rgw_ec_min_size }})
# The EC pool otherwise inherits Ceph's default min_size = k+1 implicitly. Pin
# it EXPLICITLY so the write-availability floor is documented and cannot drift:
# k+1 keeps a PG writable while at least k+1 shards are up, so writes tolerate
# up to m-1 host losses while reads still tolerate m. NEVER set below k+1 --
# allowing writes at k shards risks data loss if another shard is lost during
# recovery. Idempotent get-then-set; gated on the pool existing (fresh run:
# create_data_pool changed; re-run: pool already in pool_list).
ansible.builtin.shell: |
set -o pipefail
cur=$(ceph osd pool get {{ ceph_rgw_data_pool }} min_size | awk '{print $2}')
if [ "$cur" != "{{ ceph_rgw_ec_min_size }}" ]; then
ceph osd pool set {{ ceph_rgw_data_pool }} min_size {{ ceph_rgw_ec_min_size }}
echo CHANGED
fi
args:
executable: /bin/bash
when:
- inventory_hostname in groups['ceph_bootstrap']
- (create_data_pool is changed) or (ceph_rgw_data_pool in pool_list.stdout)
register: ec_min_size_result
changed_when: "'CHANGED' in (ec_min_size_result.stdout | default(''))"
# --- Step 4: Replicated index pool ---
- name: Create replicated index pool {{ ceph_rgw_index_pool }}
@@ -117,6 +156,18 @@
# spikes and throughput drops. Pre-sizing avoids this penalty during
# the first fill. Idempotent: only increases pg_num, never decreases.
# The top-of-file pool_list snapshot predates the data/index/non-EC pools
# created just above, so on a fresh run 1 it does NOT contain them and every
# "in pool_list.stdout" gate below would skip - deferring all pre-sizing to a
# second converge. Re-query now that the explicit pools exist so run 1 sizes
# them. (The lazily-created .rgw.* system pools are handled later, after the
# RGW daemon is up.)
- name: Re-query pools after explicit RGW pools are created
ansible.builtin.command: ceph osd pool ls
register: pool_list
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Set initial PG count on data pool
ansible.builtin.command: >
ceph osd pool set {{ ceph_rgw_data_pool }}
@@ -156,26 +207,10 @@
- pg_extra_result.rc != 0
- "'is not >= current' not in pg_extra_result.stderr | default('')"
- name: Set initial PG count on RGW metadata pools
ansible.builtin.command: >
ceph osd pool set {{ item }} pg_num {{ ceph_pg_init_meta }}
loop:
- .rgw.root
- "{{ ceph_rgw_zone }}.rgw.log"
- "{{ ceph_rgw_zone }}.rgw.control"
- "{{ ceph_rgw_zone }}.rgw.meta"
when:
- inventory_hostname in groups['ceph_bootstrap']
# Skip pools that haven't materialized yet — RGW creates them
# lazily on first daemon startup. Mirrors the per-pool gate used
# by the data/index/extra pool tasks above.
- item in pool_list.stdout
register: pg_meta_result
changed_when: "'set' in pg_meta_result.stdout | default('')"
failed_when:
- pg_meta_result.rc is defined
- pg_meta_result.rc != 0
- "'is not >= current' not in pg_meta_result.stderr | default('')"
# NB: the .rgw.* system-pool PG-count and size/min_size pins are NOT here - those
# pools are created lazily by the realm/zone setup and the RGW daemon, so they do
# not exist yet at this point in the play. They are pinned in Step 13.6 below,
# after the RGW readiness wait, against a freshly re-queried pool list.
- name: Wait for PG peering to complete
ansible.builtin.shell: |
@@ -345,6 +380,7 @@
{% for h in groups['ceph_nodes'] %}
'{{ hostvars[h]['hostname_short'] }}',
'{{ hostvars[h]['bond_ip'] }}',
'{{ hostvars[h]['ceph_service_ip'] }}',
{% endfor %}
'{{ ceph_rgw_dns_name }}',
])
@@ -412,6 +448,28 @@
# - a wildcard under the same name for virtual-hosted buckets
# - per-node FQDNs so direct-host addressing also validates
# - per-node bond IPs so IP-based S3 clients also validate
# - the RGW ingress VIP (only when ceph_rgw_ingress_enabled) so a client hitting
# the health-checked VIP, where haproxy terminates TLS with this cert, validates
# Build the -subj and SAN as single-line strings up front. Doing this inline in
# the openssl command forced a backslash-newline INSIDE the quoted args, which
# folds the next line's indentation into the value (e.g. C="DE ") and openssl
# aborts on the over-length field. Precomputing here keeps the shell one flat line
# per arg. The {%- ... %} whitespace-control strips the fold-inserted spaces so the
# SAN concatenates with no embedded whitespace.
- name: Build RGW self-signed cert subject and SAN
ansible.builtin.set_fact:
_rgw_cert_subj: >-
/C={{ ceph_rgw_ssl_cert_subject_c }}/ST={{ ceph_rgw_ssl_cert_subject_st }}/L={{ ceph_rgw_ssl_cert_subject_l }}/O={{ ceph_rgw_ssl_cert_subject_o }}/CN={{ ceph_rgw_dns_name }}/emailAddress={{ ceph_rgw_ssl_cert_email }}
_rgw_cert_san: >-
subjectAltName=DNS:{{ ceph_rgw_dns_name }},DNS:*.{{ ceph_rgw_dns_name }}
{%- for h in groups['ceph_nodes'] %},DNS:{{ hostvars[h]['hostname_short'] }}.{{ cluster_domain }}{% endfor %}
{%- for h in groups['ceph_nodes'] %},IP:{{ hostvars[h]['bond_ip'] }}{% endfor %}
{%- for h in groups['ceph_nodes'] %},IP:{{ hostvars[h]['ceph_service_ip'] }}{% endfor %}
{%- if ceph_rgw_ingress_enabled | default(false) | bool and (ceph_rgw_ingress_vip | default('')) | length > 0 %},IP:{{ ceph_rgw_ingress_vip }}{% endif %}
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_rgw_ssl | default(false) | bool
- name: Generate RGW self-signed cert and key (10 year validity)
ansible.builtin.shell: |
@@ -421,17 +479,8 @@
-keyout /etc/ceph/rgw-ssl.key \
-out /etc/ceph/rgw-ssl.crt \
-days {{ ceph_rgw_ssl_cert_days }} \
-subj "/C={{ ceph_rgw_ssl_cert_subject_c }}\
/ST={{ ceph_rgw_ssl_cert_subject_st }}\
/L={{ ceph_rgw_ssl_cert_subject_l }}\
/O={{ ceph_rgw_ssl_cert_subject_o }}\
/CN={{ ceph_rgw_dns_name }}\
/emailAddress={{ ceph_rgw_ssl_cert_email }}" \
-addext "subjectAltName=\
DNS:{{ ceph_rgw_dns_name }},\
DNS:*.{{ ceph_rgw_dns_name }}\
{% for h in groups['ceph_nodes'] %},DNS:{{ hostvars[h]['hostname_short'] }}.{{ cluster_domain }}{% endfor %}\
{% for h in groups['ceph_nodes'] %},IP:{{ hostvars[h]['bond_ip'] }}{% endfor %}"
-subj {{ _rgw_cert_subj | quote }} \
-addext {{ _rgw_cert_san | quote }}
args:
executable: /bin/bash
creates: /etc/ceph/rgw-ssl.crt
@@ -519,12 +568,192 @@
args:
executable: /bin/bash
register: rgw_running_count
until: "(rgw_running_count.stdout | int) >= (groups['ceph_nodes'] | length)"
# Wait for RGW on the nodes in THIS run (respects --limit), not all ceph_nodes,
# so a subset deploy does not block forever waiting for daemons on hosts it never
# joined. The rgw-spec placement stays pinned to all ceph_nodes (it must not
# shrink on a --limit re-run); only the readiness gate is play-scoped.
until: "(rgw_running_count.stdout | int) >= (groups['ceph_nodes'] | intersect(ansible_play_hosts) | length)"
retries: 30
delay: 10
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
# --- Step 13.7: RGW ingress (haproxy + keepalived) -- health-checked S3 VIP ---
#
# Opt-in floating VIP in front of RGW in TLS TERMINATION mode: haproxy `mode http`
# terminates the client's TLS on the VIP with the RGW self-signed cert (reused,
# combined cert+key -- see template), then re-encrypts to each RGW/beast backend
# (default-server ssl / verify none). The keepalived VIP + haproxy per-backend
# health check drop a dead RGW node from rotation instead of blackholing ~1/N of
# S3 connections.
#
# Runs HERE, after the RGW readiness wait above, on purpose: the ingress spec's
# backend_service references the RGW service, which must already exist and be
# running before cephadm can wire the ingress to it. It also needs
# rgw_ssl_cert_combined_pem, the fact built in Step 11.6.
#
# WHY TERMINATE, not passthrough: L4 passthrough (`mode tcp`) for an RGW backend
# requires IngressSpec.use_tcp_mode_over_rgw, a post-20.2.2 upstream field absent
# on the pinned 20.2.2 (source-verified: an unknown spec key makes `ceph orch
# apply` TypeError). Terminate is the only health-checked ingress mode available
# here. haproxy forwards the Host header in mode http, so virtual-hosted buckets
# keep working against haproxy's wildcard cert.
#
# NO 443 COLLISION: haproxy terminates on VIP:{{ ceph_rgw_ingress_frontend_port }},
# beast binds each node's fabric IP:{{ ceph_rgw_port }} -- distinct sockets. This
# holds ONLY when the RGW spec restricts beast to a specific IP (ceph_bind_networks
# non-empty); otherwise beast binds 0.0.0.0 and the co-located VIP bind collides.
# Asserted below.
#
# APPLY ORDERING (cross-tool): the DNS cutover that points s3.<domain> at the VIP
# is done separately in TF and MUST NOT precede this ingress being live, or S3
# resolves to a VIP that nothing answers.
#
# Idempotent + declarative: `ceph orch apply -i` reconciles to the spec.
- name: Assert RGW ingress preconditions (VIP + restricted beast bind + RGW ssl)
ansible.builtin.assert:
that:
- (ceph_rgw_ingress_vip | default('')) | length > 0
- ceph_bind_networks | default([]) | length > 0
- ceph_rgw_ssl | default(false) | bool
fail_msg: >-
ceph_rgw_ingress_enabled is true but a precondition is unmet.
Require: ceph_rgw_ingress_vip set (e.g. 10.40.20.250); ceph_bind_networks
non-empty so beast binds a specific per-node IP (else beast binds
0.0.0.0:{{ ceph_rgw_port }} and the haproxy VIP bind on the same port fails
with EADDRINUSE); and ceph_rgw_ssl true so the self-signed cert exists for
haproxy to terminate with (rgw_ssl_cert_combined_pem).
quiet: true
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_rgw_ingress_enabled | default(false) | bool
- name: Render RGW ingress service spec
ansible.builtin.template:
src: rgw-ingress-spec.yaml.j2
dest: /etc/ceph/rgw-ingress-spec.yaml
owner: root
group: root
mode: '0600'
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_rgw_ingress_enabled | default(false) | bool
- (ceph_rgw_ingress_vip | default('')) | length > 0
no_log: true # spec embeds the RGW private key (combined PEM)
- name: Apply RGW ingress service spec
ansible.builtin.command: ceph orch apply -i /etc/ceph/rgw-ingress-spec.yaml
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_rgw_ingress_enabled | default(false) | bool
- (ceph_rgw_ingress_vip | default('')) | length > 0
changed_when: false
# --- Step 13.6: Pin pools to device-class CRUSH rules ---
#
# crush-rules.yml creates replicated_ssd/replicated_hdd but binds no pool, so
# every replicated pool falls to the default class-agnostic rule and the
# omap-heavy bucket index lands on HDD while the ssd-osd NVMe tier sits idle.
# Pin the latency-sensitive index/metadata/mgr pools to the SSD rule and the
# bulk non-EC pool to HDD.
#
# Runs HERE (after RGW readiness) on purpose: .rgw.root and the prod-z1.rgw.*
# metadata pools are created lazily by the realm/zone setup and the RGW daemon,
# so a pool snapshot taken before the daemon is up (like the top-of-file one)
# would miss them. Re-query fresh, and gate each pin on the pool existing now.
# Opt-in per cluster via ceph_rgw_ssd_pools / ceph_rgw_hdd_pools (empty default
# leaves a live cluster's pools where they are - re-homing forces a rebalance).
- name: Re-query pools after RGW is up (system pools are created lazily)
ansible.builtin.command: ceph osd pool ls
when: inventory_hostname in groups['ceph_bootstrap']
register: pool_list_post
changed_when: false
- name: Set initial PG count on RGW metadata pools
ansible.builtin.command: >
ceph osd pool set {{ item }} pg_num {{ ceph_pg_init_meta }}
loop:
- .rgw.root
- "{{ ceph_rgw_zone }}.rgw.log"
- "{{ ceph_rgw_zone }}.rgw.control"
- "{{ ceph_rgw_zone }}.rgw.meta"
when:
- inventory_hostname in groups['ceph_bootstrap']
- item in (pool_list_post.stdout | default(''))
register: pg_meta_result
changed_when: "'set' in pg_meta_result.stdout | default('')"
failed_when:
- pg_meta_result.rc is defined
- pg_meta_result.rc != 0
- "'is not >= current' not in pg_meta_result.stderr | default('')"
- name: Pin replica size on the RGW metadata + mgr pools (durability)
# These otherwise inherit the cluster-spec osd_pool_default_size / min_size,
# unlike index/extra which are pinned. Losing 2 OSDs behind a metadata PG would
# take out the RGW/mgr control plane, and min_size 1 permits single-replica
# writes. Pin to the same replicated size (spice 3/2; sietch stays 2/1 via its
# own var). Runs here, after RGW readiness, because these pools are created
# lazily by the daemon; the pre-daemon snapshot missed them so run 1 skipped
# the pin and they sat at the inherited 2/1 until a second converge. Idempotent:
# only sets when the current value differs.
ansible.builtin.shell: |
set -o pipefail
ch=0
[ "$(ceph osd pool get {{ item }} size | awk '{print $2}')" = "{{ ceph_rgw_replicated_size }}" ] || { ceph osd pool set {{ item }} size {{ ceph_rgw_replicated_size }}; ch=1; }
[ "$(ceph osd pool get {{ item }} min_size | awk '{print $2}')" = "{{ ceph_rgw_replicated_min_size }}" ] || { ceph osd pool set {{ item }} min_size {{ ceph_rgw_replicated_min_size }}; ch=1; }
[ "$ch" = 1 ] && echo CHANGED || echo ok
args:
executable: /bin/bash
loop:
- .rgw.root
- "{{ ceph_rgw_zone }}.rgw.log"
- "{{ ceph_rgw_zone }}.rgw.control"
- "{{ ceph_rgw_zone }}.rgw.meta"
- .mgr
when:
- inventory_hostname in groups['ceph_bootstrap']
- item in (pool_list_post.stdout | default(''))
register: size_meta_result
changed_when: "'CHANGED' in (size_meta_result.stdout | default(''))"
- name: Pin pools to the SSD device-class CRUSH rule
ansible.builtin.shell: |
set -o pipefail
cur=$(ceph osd pool get {{ item | quote }} crush_rule | awk '{print $2}')
if [ "$cur" != {{ ceph_rgw_ssd_crush_rule | quote }} ]; then
ceph osd pool set {{ item | quote }} crush_rule {{ ceph_rgw_ssd_crush_rule | quote }}
echo CHANGED
fi
args:
executable: /bin/bash
loop: "{{ ceph_rgw_ssd_pools }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_rgw_ssd_pools | length > 0
- item in (pool_list_post.stdout | default(''))
register: ssd_pin_result
changed_when: "'CHANGED' in (ssd_pin_result.stdout | default(''))"
- name: Pin pools to the HDD device-class CRUSH rule
ansible.builtin.shell: |
set -o pipefail
cur=$(ceph osd pool get {{ item | quote }} crush_rule | awk '{print $2}')
if [ "$cur" != {{ ceph_rgw_hdd_crush_rule | quote }} ]; then
ceph osd pool set {{ item | quote }} crush_rule {{ ceph_rgw_hdd_crush_rule | quote }}
echo CHANGED
fi
args:
executable: /bin/bash
loop: "{{ ceph_rgw_hdd_pools }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_rgw_hdd_pools | length > 0
- item in (pool_list_post.stdout | default(''))
register: hdd_pin_result
changed_when: "'CHANGED' in (hdd_pin_result.stdout | default(''))"
# --- Step 13.5: Ensure dashboard RGW user has admin caps ---
#
# cephadm auto-creates a "dashboard" RGW user with system=true but
@@ -694,16 +923,6 @@
no_log: true
tags: [s3_keys]
- name: S3 key drift notice (dry flag — no action taken)
ansible.builtin.debug:
msg:
- "S3 svc-user keys differ from 1P values."
- "To align (rotates keys, requires Yucca-app re-config): re-run with -e rotate_s3_keys=true"
when:
- inventory_hostname in groups['ceph_bootstrap']
- s3_key_drift | default(false)
- not (rotate_s3_keys | default(false) | bool)
- name: Display S3 endpoint (credentials live in 1P, not log output)
ansible.builtin.debug:
msg:
@@ -711,7 +930,7 @@
- "Access Key: op://{{ cluster_secrets_vault }}/{{ cluster_name | upper }}_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY/password"
- "Secret Key: op://{{ cluster_secrets_vault }}/{{ cluster_name | upper }}_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY/password"
- "Endpoint (DNS): {{ ceph_rgw_scheme }}://{{ ceph_rgw_dns_name }}:{{ ceph_rgw_port }}"
- "Endpoint (direct): {{ ceph_rgw_scheme }}://{{ bond_ip }}:{{ ceph_rgw_port }}"
- "Endpoint (direct): {{ ceph_rgw_scheme }}://{{ ceph_service_ip }}:{{ ceph_rgw_port }}"
- "Region: {{ ceph_rgw_zonegroup_api_name }}"
- "Virtual-hosted: {{ ceph_rgw_scheme }}://<bucket>.{{ ceph_rgw_dns_name }}:{{ ceph_rgw_port }}"
- "Path-style: {{ ceph_rgw_scheme }}://{{ ceph_rgw_dns_name }}:{{ ceph_rgw_port }}/<bucket>"
@@ -84,3 +84,98 @@
The running cluster's password does not match ceph_dashboard_password.
Re-run the ceph_deploy role to reconcile, or investigate an out-of-band change.
when: inventory_hostname in groups['ceph_bootstrap']
# --- Validation: device-class CRUSH pool placement ---
# Guards the "rule created but no pool bound" failure mode surfaced by the deep
# review: assert every pool we intended on a tier actually carries that rule, and
# that the SSD rule bound to live OSDs (a rule whose device class has zero OSDs
# selects nothing and leaves the pinned pool's PGs inactive). Only runs when the
# cluster opts in to pinning (ceph_rgw_ssd_pools / _hdd_pools); skips absent pools.
- name: Validate device-class pool placement
ansible.builtin.shell: |
set -o pipefail
rc=0
check() { # $1 pool, $2 want_rule
ceph osd pool ls | grep -qxF "$1" || { echo "skip (absent): $1"; return; }
got=$(ceph osd pool get "$1" crush_rule | awk '{print $2}')
if [ "$got" = "$2" ]; then echo "ok: $1 -> $2"
else echo "FAIL: $1 -> $got (want $2)"; rc=1; fi
}
{% for p in ceph_rgw_ssd_pools %}
check {{ p | quote }} {{ ceph_rgw_ssd_crush_rule | quote }}
{% endfor %}
{% for p in ceph_rgw_hdd_pools %}
check {{ p | quote }} {{ ceph_rgw_hdd_crush_rule | quote }}
{% endfor %}
n=$(ceph osd crush class ls-osd ssd 2>/dev/null | grep -c '^[0-9]' || true)
if [ "$n" -gt 0 ]; then echo "ok: ssd class has $n OSDs"
else echo "FAIL: ssd class has 0 OSDs (replicated_ssd would select nothing)"; rc=1; fi
[ "$rc" = 0 ] && echo VALIDATION_OK || echo VALIDATION_FAIL
args:
executable: /bin/bash
when:
- inventory_hostname in groups['ceph_bootstrap']
- (ceph_rgw_ssd_pools | length + ceph_rgw_hdd_pools | length) > 0
register: crush_placement_check
changed_when: false
failed_when: "'VALIDATION_FAIL' in (crush_placement_check.stdout | default(''))"
- name: Show device-class pool placement validation
ansible.builtin.debug:
msg: "{{ crush_placement_check.stdout_lines }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- crush_placement_check.stdout_lines is defined
# --- Alert delivery status ---
# cephadm's alertmanager ships with no receiver, so a cluster with no configured
# webhook routes every alert rule into a null route. Surface this loudly rather
# than failing the play (the destination is a deliberate per-cluster decision).
- name: Report alert-delivery status
ansible.builtin.debug:
msg: >-
{{ ('ALERT DELIVERY: ' ~ (ceph_alertmanager_webhook_urls | length | string)
~ ' webhook receiver(s) configured.')
if (ceph_alertmanager_webhook_urls | default([]) | length) > 0
else ('WARNING: no alert delivery configured - alertmanager has no receiver, so the '
~ 'Prometheus alert rules fire into a null route and reach no human. '
~ 'Set ceph_alertmanager_webhook_urls to a webhook/pager endpoint.') }}
when: inventory_hostname in groups['ceph_bootstrap']
# --- Validation: RGW replicated-pool durability ---
# Guards the run-1 durability regression surfaced by the deep review: the
# .rgw.* system pools are created lazily by the RGW daemon, so a size/min_size
# pin taken before the daemon is up skips them and they sit at the inherited
# 2/1. Assert every RGW replicated pool that exists carries the cluster's target
# size/min_size, so a miss fails the play instead of shipping single-replica.
- name: Validate RGW replicated-pool durability
ansible.builtin.shell: |
set -o pipefail
rc=0
check() { # $1 pool
ceph osd pool ls | grep -qxF "$1" || { echo "skip (absent): $1"; return; }
sz=$(ceph osd pool get "$1" size | awk '{print $2}')
mn=$(ceph osd pool get "$1" min_size | awk '{print $2}')
if [ "$sz" = {{ ceph_rgw_replicated_size | quote }} ] && [ "$mn" = {{ ceph_rgw_replicated_min_size | quote }} ]; then
echo "ok: $1 size=$sz min_size=$mn"
else
echo "FAIL: $1 size=$sz min_size=$mn (want {{ ceph_rgw_replicated_size }}/{{ ceph_rgw_replicated_min_size }})"; rc=1
fi
}
{% for p in [ceph_rgw_index_pool, ceph_rgw_extra_pool, '.rgw.root', ceph_rgw_zone ~ '.rgw.log', ceph_rgw_zone ~ '.rgw.control', ceph_rgw_zone ~ '.rgw.meta', '.mgr'] %}
check {{ p | quote }}
{% endfor %}
[ "$rc" = 0 ] && echo VALIDATION_OK || echo VALIDATION_FAIL
args:
executable: /bin/bash
when: inventory_hostname in groups['ceph_bootstrap']
register: pool_durability_check
changed_when: false
failed_when: "'VALIDATION_FAIL' in (pool_durability_check.stdout | default(''))"
- name: Show RGW replicated-pool durability validation
ansible.builtin.debug:
msg: "{{ pool_durability_check.stdout_lines }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- pool_durability_check.stdout_lines is defined

Some files were not shown because too many files have changed in this diff Show More