* feat(ceph): scaffold spice prod cluster (reprovision + inventory + TF)
Stand up spice (48x Hetzner SX295, prod/htz-fsn1): the reprovision_hetzner
ansible role (rescue -> reset -> installimage -> verify, base-OS-only, with
stale-mdraid pre-clean and resume markers), the prod-htz-fsn1/spice inventory
+ group_vars/host_vars, the prod ceph TF stack, and gen-spice-host-vars.py.
Adds a `mise reprovision` task (op run + tf/.env.prod), the spice SSH key
(yucca_tf_prod), and per-host roles-based [ceph_mon] filtering so large
clusters pin a fixed MON quorum instead of defaulting to all nodes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): create spice block.db + ssd-osd LVs at converge (NVMe-RAID)
Add an NVMe-RAID branch to ceph_deploy lvm-setup so spice's single vg0 (built
by installimage) gets its 14 block.db LVs + one ssd-osd LV created at converge,
before osds.yml. This replaces painbox's fragile installimage -x chroot
post-install with an idempotent, observable ansible step; installimage stays
base-OS-only. Narrows the old blanket "externally-managed LVM" skip so it only
fires when neither the sietch dual-SSD nor the NVMe-RAID shape applies, and
wires ceph_db_vg / ceph_ssd_osd_lv / ceph_ssd_osd_reserve_gib in the spice
group_vars.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): set spice bond NICs in group_vars, drop per-host placeholders
The 25G fabric NICs are uniform across all 48 SX295 (verified live on 36
nodes): bond members enp193s0f0/f1 (Intel E800/ice, PCI c1:00.0/.1), WAN
enp197s0 (igb, c5:00.0). Set bond_interfaces + oob_nic once in the spice
group_vars -- the networkd role reads bond_interfaces -- and remove the dead
per-host fabric_nic/oob_nic PLACEHOLDER lines from all 48 host_vars,
spice-hosts.yaml, and gen-spice-host-vars.py. No MAC-based naming needed: the
predictable names are PCI-derived and identical fleet-wide.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): LACP + VLAN sub-interfaces on the bond, lldpd fabric config
networkd role: emit 802.3ad LACP params (TransmitHashPolicy, LACPTransmitRate)
when bond_mode is 802.3ad, and support tagged VLAN sub-interfaces on the bond
via networkd_bond_vlans -- when set, bond0 carries no L3 and each VLAN carries
its own address. Opt-in: an empty list keeps the flat active-backup path
(sietch) byte-identical (verified by render). verify.yml gates the
bond0-IP/gateway asserts to the flat case and checks per-VLAN addresses
otherwise.
spice: bond0 becomes an 802.3ad LACP bond of enp193s0f0/f1 (MLAG to the QFX5240
leaves) carrying VLAN 120 (public 10.40.20.<idx>/23) + VLAN 122 (private
10.40.22.<idx>/23); default route stays on the 1G WAN until cutover.
baseline role: install + configure lldpd (portid ifname, cluster system
description, service enabled) when lldpd is in baseline_extra_packages, and add
lldpd/ethtool/tcpdump for spice. This also newly applies lldpd config to sietch
(already lists lldpd) on its next converge.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): set spice RGW pool to EC 16+4 host domain for the beta
Capacity-max profile (80% usable, m=4) for the wipeable beta behind michael.
Host failure domain accepts rack-loss risk in exchange for capacity, which is
fine for a beta and matches the single-rack API tier. Also lifts the RGW DRAFT
marker and starts the bulk EC data pool at pg_num 4096 (672 OSDs, 20-chunk pool)
instead of the 36-OSD-era default to avoid PG splitting during the first fill.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* docs(ceph): drop stale -x chroot reference from spice autosetup header
installimage.yml runs base-OS-only (installimage -a -c /autosetup, no -x). The
autosetup header still carried the painbox `-x /tmp/post-install.sh` line, which
implied a chroot post-install the role deliberately dropped. Point it at the
convergence steps (lvm-setup, baseline) that replaced it.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): minimal -x chroot to seed root + ansible-iac keys on reprovision
installimage sometimes runs without our key present in the rescue (Robot armed
rescue without it, in-place installimage, manual rescue), so its "copy the rescue
authorized_keys" step leaves root keyless and the node unreachable for
convergence. Add a minimal `-x /tmp/post-install.sh` that does ONLY inert
user/key writes: seed root's authorized_keys with the iac key (baseline never
manages root, so it persists) and create an ansible-iac account with the key +
NOPASSWD sudo. No apt/LVM in the chroot -- that fragile step stays at post-boot
convergence, so set -e cannot spuriously disrupt installimage.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): converge iac access on existing nodes without reimaging
Nodes imaged before the reprovision -x chroot lack root's iac key and the
ansible-iac sudo account; reseeding them meant a reinstall. baseline gains an
iac-access task that ensures the same state idempotently -- root's authorized
key (non-exclusive) plus the ansible-iac user with key and NOPASSWD sudo -- so
the 36 already-installed nodes converge in place. Gated on the cluster iac key;
run standalone with --tags iac_access.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): manual rescue-first add-node path, Robot API optional
The Hetzner Robot rescue/reset step is both the unreliable part (nodes that arm
rescue but never boot it) and the highest blast radius (an errant /reset on a
live spice node in production). Gate it behind reprovision_use_robot (default
true, preserving reprovision.yml) and add add-node.yml, which sets it false: the
operator arms rescue by hand in the portal and Ansible runs only install, chroot
and verify on a node already in rescue. Robot key registration moves to
register_robot_keys.yml, imported from preflight only on the Robot path. With the
API out of the loop nothing here can flip a running node into rescue -- wait_rescue
refuses any node without the installimage ramdisk.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): relax stale-rescue guard on the manual add-node path
wait_rescue rejects a rescue older than reprovision_rescue_max_uptime (1200s) to
catch a stale boot -- valid only when Ansible triggered the reset. On the manual
path the operator rescues by hand, so uptime just measures wait-before-run and a
node legitimately in rescue for hours would be refused. Relax the guard to a day
on add-node.yml; the installimage-ramdisk check is the real in-rescue proof.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): skip the this-boot rescue guard on the manual path
philip has sat in rescue ~8 days, so any finite uptime ceiling rejects it. The
guard only means something when Ansible triggered the reset (uptime proves this
boot); on the manual add-node path a node may sit in rescue for weeks, so gate
the assert on reprovision_use_robot rather than inflating the timeout. The
installimage-ramdisk check remains the real in-rescue proof.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): networkd/ifupdown coexistence to protect the 1G WAN
The role assumed a full migration off ifupdown -- correct for sietch's flat bond,
but on spice it would disable ifupdown and leave the 1G WAN (a separate NIC that
ifupdown owns, carrying the default route) unconfigured on the next boot: the
painbox failure. Add networkd_replace_ifupdown (default true, sietch unchanged).
When false, networkd manages only the bond and VLANs; an Unmanaged=yes guard
fences the WAN NIC off, commit leaves networking.service enabled, the bond drops
to RequiredForOnline=no so a carrier-less fabric cannot stall boot, and verify
asserts the WAN kept its address and default route before anything commits. The
rollback script targets the right interface per mode. spice sets it false.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): make migrate-networkd usable pre-Ceph (gate health checks)
The ceph-health pre/post gate assumed a live cluster and a ceph_bootstrap group,
so the play could not run during spice bringup (coexistence activation before any
cephadm deploy). Gate the four ceph tasks on ceph_bootstrap being populated; on a
pre-Ceph cluster they skip and the networkd role runs on its own.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* chore(ceph): tunable serial + halt-on-failure for migrate-networkd
Fleet rollout of the coexistence activation wants batches, not one-at-a-time, and
must stop the instant a node fails (a dropped WAN shows up as unreachable) rather
than silently skip it. Template serial (networkd_serial, default 1 unchanged) and
set max_fail_percentage 0.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): install Intel ice DDP firmware so E810 NICs leave Safe Mode
The minimal Debian image ships no firmware-misc-nonfree, so the E810 DDP package
(intel/ice/ddp/ice.pkg) is absent and every NIC boots into Safe Mode -- whose
crippled classifier drops reserved-multicast control frames (LACP, LLDP), so the
25G fabric bond never aggregates despite a correctly configured switch. Add
firmware-misc-nonfree to the spice package set and a baseline nic-firmware task
that reboots an E810 host once to load the DDP when it is still in Safe Mode
(self-gating: no-op on non-ice or already-loaded hosts; refuses to reboot when the
DDP is absent, so it cannot loop). The WAN is a separate igb NIC, so the reboot
keeps the node reachable.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): correct ice DDP activation gating (pipefail + dpkg check)
The Safe Mode check used `set -o pipefail`, but `ethtool --show-fec` exits
non-zero on "not supported", which pipefail propagated and masked grep's match --
so the check always yielded "ok" and the activating reboot never fired. Drop
pipefail there, and gate the reboot on the dpkg DB (firmware-misc-nonfree
installed) rather than a file stat that can lag a large apt transaction. Verified
on one node end-to-end: DDP loads, Safe Mode clears, LACP converges with the
switch, both VLAN gateways ping.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): pin chrony time sources for Ceph time sync
chrony was listed in baseline_enable_services but never installed, so system.yml
would fail starting it -- and there was no time-source config at all, leaving sync
to the distro default pool. Ceph mon quorum is skew-sensitive, so make it explicit:
add baseline/chrony.yml (install + templated chrony.conf + enable) driven by
chrony_ntp_servers (default Debian pool, makestep for the initial correction), drop
chrony from baseline_enable_services so it is owned in one place, and point spice at
Hetzner NTP (ntp1/2/3.hetzner.de) -- low-latency from FSN1, consistent across all 47.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): populate spice OSD by-path map; exclude noreen until recovered
Inspected the SX295 disk layout: 14x 20TB SATA HDDs per node across 3 AHCI
controllers, by-path uniform on 46/47 nodes -- so ceph_hdd_osds (path_phy + db
vg0/db-slotN) and ceph_ssd_osds (vg0/ssd-osd) live in group_vars, not 47 host_vars.
spice-ceph-miguel has one disk on 46:00.0-ata-4 rather than 87:00.0-ata-4 and
overrides the map in its host_vars. Verified by rendering osd-spec.yml.j2: 658 HDD
OSD paths (47x14) + 47 NVMe ssd-osd, miguel's override resolving correctly. noreen
(boot-order casualty, held in triage) is commented out of clusters.auto.tfvars so
the rendered inventory + deploy target only the 47 live nodes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): bootstrap mon on the fabric public_network, not the WAN
cephadm bootstrap --mon-ip must fall inside public_network, but spice
passed bond_ip, the 1G WAN address used for the ansible/SSH connection
and cephadm host management. spice's public_network is the 25G fabric
(10.40.20.0/23), so the initial mon would fail to bind. Bootstrap now
uses ceph_public_ip (10.40.20.<host_index>, derived like the bond VLANs);
sietch is flat with no host_index and falls back to bond_ip, already in
its own public_network.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* ci(ceph): static validation gate for the ansible/ceph stack
CI only validated the Kubernetes/Flux surface; nothing parsed the ceph
playbooks, inventory or scripts, so a broken playbook or host_var first
executed during the post-merge prod apply against real nodes. This runs
the existing local validators (yamllint, ansible-lint, shellcheck,
ansible-playbook --syntax-check, py_compile) on ansible/ceph PRs with no
secrets and no connection to any host; a throwaway localhost inventory
satisfies --syntax-check since the real inventory is TF-generated.
Its own workflow, not a job in ci.yml, so the ansible/ceph/** path filter
is workflow-scoped and does not gate the unrelated k8s jobs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): make the ceph converge deliberate and cluster-parameterized
A merge touching ansible/ceph/** could open the staging partition and
auto-run the full baseline->tune->deploy->harden pipeline against the
live sietch cluster, and the converge was hardcoded to sietch/austin
(render-inventories.sh got no region, install-ssh-keys.sh got a literal
sietch, CEPH_ENV pointed at .../sietch/), so a prod converge could never
reach spice.
The ceph converge now runs only on workflow_dispatch with
run_ceph_ansible=true, pinned to the one matrix entry that owns the
chosen ceph_cluster (spice=prod/htz-fsn1, sietch=staging/austin), so a
push never reconverges a live cluster and one dispatch cannot converge
both. Region and cluster are threaded from the dispatch inputs / matrix
into the render, key install, and CEPH_ENV. The staging paths-filter is
scoped to ansible/ceph/inventories/staging-** so a prod ceph change no
longer opens the staging matrix. The ceph TF apply stays auto and
env-gated; the mgmt converge is unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): make the RGW firewall source scope configurable
The RGW/S3 port was accepted from any source unconditionally, while
every other Ceph service (mon, osd, dashboard, monitoring) is scoped to
ceph_firewall_trusted_networks. On spice that port is plain HTTP 7480, so
the S3 endpoint sat open to the public internet with no lever to close
it. A new ceph_firewall_rgw_any_source (default true, mirroring
ceph_firewall_ssh_any_source) keeps the open behavior by default but lets
production restrict RGW to the trusted networks, which already cover the
fabric and NetBird overlay.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* chore(ceph): fix stale OSD-map comment on the noreen host_var
noreen was generated before the by-path map moved into group_vars, so it
still carried the DEFERRED note while its 47 siblings point at the group
var. The host stays pre-staged for re-add once recovered; only the
comment was wrong.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): add the host-mgmt VLAN (124) to the spice bond
The FSN1-C1 fabric carries a third VLAN on the 25G bond,
FSN1-C1-HOST-MGMT (124, 10.40.24.0/24), alongside Ceph public (120) and
private (122). The leaf already trunks it to every server bond and
advertises 10.40.24.0/24 into the NetBird overlay, so it is the intended
in-band ansible/SSH reach once the WAN is retired, but the host side had
no matching sub-interface. Add bond0.124 at 10.40.24.<host_index>/24 with
no gateway, so the default route stays on the 1G WAN until cutover.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): apply networkd config deltas to an already-running daemon
The networkd role only started systemd-networkd, which is a no-op once it
is already running, so a re-run that added config (e.g. a new bond VLAN)
wrote the .netdev/.network files but never applied them, and verify then
failed asserting the sub-interface had no address. Add a networkctl reload
plus a per-VLAN settle wait after the start, so a re-run creates the added
sub-interfaces live without tearing down existing links; the WAN on
ifupdown is untouched regardless. Non-disruptive on the flat sietch path
(the settle wait is gated on networkd_bond_vlans).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): bind Ceph services on the fabric only, never the WAN
Ceph on spice must never listen on the 1G WAN. mon/mgr/OSD already bind
per public_network/cluster_network, but the beast RGW frontend and the
cephadm monitoring stack bound 0.0.0.0, and the mon-ip, cephadm host-add
address, and dashboard URLs advertised the WAN bond_ip. Add a per-cluster
policy: ceph_service_ip (the address Ceph advertises and binds for point
services; defaults to bond_ip, resolving to the fabric on spice via
ceph_public_ip) and ceph_bind_networks (the cephadm networks: bind
restriction for RGW and the monitoring daemons; defaults to
public_network). Route mon-ip, the host-add address, the dashboard
monitoring URLs, RGW, and prometheus/grafana/alertmanager/node-exporter/
ceph-exporter through them, and scope the RGW firewall to the trusted
networks. The RGW zonegroup hostnames and TLS SAN gain the fabric address
so signed requests to the fabric IP still validate. Flat clusters like
sietch have no ceph_public_ip, so ceph_service_ip falls back to bond_ip
and ceph_bind_networks to their single flat network; point services are
unchanged there.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): define the metrics-worker RGW user on spice
rgw.yml Step 14.5 creates a read-only RGW admin for the yucca-metrics-worker
service and references ceph_rgw_metrics_user_*, which only sietch defined,
so the RGW phase failed on spice with AnsibleUndefinedVariable. Add the
block; the keys are the TF-minted SPICE_METRICS_WORKER_* items, injected
via secrets.yml.tpl as vault_metrics_worker_*, so op inject resolves them.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): refuse to deploy against an un-rendered inventory
The ceph_deploy role gates every phase on the ceph_bootstrap/ceph_mon/
ceph_join groups that render-inventories.sh emits from TF state; run
against the hand-written stopgap inventory that lacks them, the role
errors mid-deploy or falls through to placing a MON on every host. A
pre-task assert now requires ceph_bootstrap to be exactly one host that
is also in ceph_mon, and ceph_mon to be a non-empty odd quorum, and
stops with a render hint otherwise.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): make the fabric-only bind restriction opt-in per cluster
ceph_bind_networks defaulted to public_network, which would have injected
a networks: bind restriction into every cluster - re-binding sietch's live
RGW and monitoring daemons from 0.0.0.0 to its flat network on the next
converge (a redeploy, and a break for any access path not on that subnet).
Default it empty instead: no networks: field is emitted and the monitoring
re-spec is skipped, so flat clusters bind every interface exactly as
cephadm ships them. spice opts in to the fabric public network in its
group_vars.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): make ceph_service_ip a group_var so hostvars can read it
ceph_service_ip was a ceph_deploy role default, but join/rgw/monitoring
read it as hostvars[<host>]['ceph_service_ip'], and role defaults are NOT
exposed through hostvars - the lookup raised "HostVarsVars object has no
attribute ceph_service_ip" and would abort the deploy on every cluster at
the join phase (a regression the fabric-only change introduced for both
spice and the live sietch). Define it in each cluster's group_vars instead
(group_vars do resolve through hostvars, verified per host): spice to the
fabric ceph_public_ip, sietch to bond_ip (unchanged from pre-fabric).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* chore(ceph): drop the unused ceph_cluster_ip var on spice
ceph_cluster_ip was defined for symmetry with ceph_public_ip but consumed
nowhere - OSD replication binds to the cluster_network CIDR, which cephadm
resolves to each node's bond0.122 address on its own. Remove the dead var
and correct the neighbouring comment (ceph_service_ip is a group_var now,
not a role default).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): serve spice RGW over TLS on 443 with a 10-year self-signed cert
Match sietch: flip spice RGW/beast from plain HTTP 7480 to HTTPS 443 with a
self-signed 10-year cert. rgw.yml generates /etc/ceph/rgw-ssl.{crt,key} and
cephadm distributes it to every RGW daemon via the service spec; the SANs
already cover s3.<domain> + the wildcard + each node's fabric IP
(ceph_service_ip). The firewall RGW port derives from ceph_rgw_port, so it
follows to 443, scoped to the trusted networks (RGW binds the fabric only,
never the WAN). This also makes the discovery rgw_s3_endpoint (hardcoded
https) correct for spice.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(dns): prod Cloudflare DNS stack for the spice RGW S3 endpoint
Add tf/deployment/prod/global/dns (mirroring staging/global/dns) so
s3.prod.fsn1.htz.futo.cloud and its wildcard resolve. Round-robin A across
all 47 spice ceph nodes' fabric public IPs (10.40.20.<host_index>, VLAN 120)
where RGW/beast binds; proxied=false (private RFC1918, reached over the
NetBird-advertised cls1_public 10.40.20.0/23). The wildcard serves S3
virtual-hosted buckets, and both names are in the self-signed TLS cert SANs.
noreen (host_index 40) excluded. CI auto-discovers the stack (applies at
order 0 under the prod-global environment); tf/.env.prod gains the token ref.
Next up: create op://yucca_tf_prod/CLOUDFLARE_API_TOKEN before the apply.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): preflight-assert the Ceph service IP is up before deploy
A node whose fabric VLAN sub-interface (spice bond0.120) did not come back
after a reboot would otherwise fail deep inside cephadm bootstrap/join. Add
a per-host pre_task that asserts ceph_service_ip is present in
ansible_all_ipv4_addresses, so a fabric-down node halts up front with a
clear message. Passes on a healthy node (verified on spice-ceph-adelia);
on sietch ceph_service_ip is bond_ip, the connection address, so it is
trivially satisfied.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* chore(ceph): honest ssd_model_pattern note + guard empty monitoring networks
ssd_model_pattern on spice was "SAMSUNG" with a "verify on node 1" note, but
the NVMe is SOLIDIGM/KIOXIA and the var is only read by ceph_destroy's
partition-6 SSD wipe (sietch's dual-SSD shape); spice is NVMe-RAID with no
partition 6 and a vg0 ssd-osd LV cleaned by the generic VG/PV removal, so
the pattern is a deliberate no-match. Reword to say so (it must stay defined
because cleanup.yml references it unconditionally). Also wrap the
monitoring-spec networks block in a length guard so an empty ceph_bind_networks
renders no dangling `networks:` key (defensive; the apply is already gated).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): lock down root/password SSH in baseline, not just harden
installimage ships PermitRootLogin yes + a root password, and all sshd
hardening lived in harden.yml (the last deploy stage), so a freshly imaged
node sat root-password-open on the public WAN for the whole campaign (the
11-node exposure the swarm audit found). Add an early baseline task, right
after iac-access authorizes the key on root, that deploys a 10-baseline-ssh
drop-in (PermitRootLogin prohibit-password + PasswordAuthentication no,
same values as security/50-hardening.conf so they never disagree), locks
the root password, and removes the interim remediation drop-in. A
Validate->Reload handler chain runs sshd -t before reloading, only on
change. Key-safe on both clusters (ansible connects by key), so nothing
can lock out.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): pin RGW metadata + mgr pools to the replicated size
The RGW system-metadata pools (.rgw.root, {zone}.rgw.{log,control,meta})
and .mgr inherited the cluster-spec default osd_pool_default_size 2 /
min_size 1, unlike the index/extra pools which are pinned. Losing two OSDs
behind a metadata PG would take out the RGW/mgr control plane, and
min_size 1 permits single-replica writes. Pin them to ceph_rgw_replicated_
size/_min_size (spice 3/2; sietch keeps 2/1 via its own vars), existence-
gated like the pg_num loop and idempotent (only sets when the value
differs).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): make the ops password hash idempotent
users.yml hashed ops_password with password_hash('sha512') and no salt, so
a fresh random salt was drawn every run: the hash never matched /etc/shadow
and the user module rewrote it reporting 'changed' on every converge (a
clean converge was never a true green signal, on both clusters). Derive a
stable salt from a one-way sha256 of the password so the hash is
deterministic and idempotent, while still reconciling an out-of-band
password change. The salt in /etc/shadow is public and leaks nothing.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): make live fabric/hosts/timezone reproducible from the repo
Three drift gaps where the running fleet did not match the committed IaC.
networkd_enabled was committed for sietch only, so spice's fabric VLANs
came up from an ad-hoc -e and a reprovisioned node would not bring them up
from the repo alone; commit it true for spice (networkd_replace_ifupdown
false still fences the 1G WAN). hosts.j2 mapped every node name to bond_ip
(the 1G WAN) rather than the fabric ceph_service_ip, so in-cluster name
resolution pointed off the fabric. And timezone: UTC was declared but never
applied, leaving nodes on the image default (Europe/Berlin); add a
community.general.timezone task to baseline/system.yml.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): scope OSD/RGW readiness waits to the in-play host set
join.yml is --limit-aware (it intersects ansible_play_hosts), but the OSD
provisioning wait computed EXPECTED from all ceph_nodes and the RGW wait
blocked until running rgw >= all ceph_nodes, so a subset (--limit) deploy
joined the subset then deadlocked at both gates. Base both counts on
groups['ceph_nodes'] intersect ansible_play_hosts, matching join.yml. On a
full deploy the intersection is all nodes, so behavior is unchanged; the
rgw-spec placement stays pinned to all ceph_nodes (only the readiness gate
is play-scoped, so a --limit re-run cannot shrink the spec).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): scope the nftables ruleset to table inet filter
nftables.conf.j2 did a global `flush ruleset`, and the no-op guard diffed
the whole `nft list ruleset`. On a ceph node (cephadm runs podman) that
wiped the podman/docker/fail2ban tables on every apply, and made the
desired-vs-live diff never match (podman tables are live but absent in the
throwaway netns), so the firewall reloaded on every converge - each one
flushing podman again. Replace only table inet filter (add, delete,
re-add), compare only that table in the guard, and reload via ExecReload
(nft -f, no global flush) instead of restart (whose ExecStop flushes the
whole ruleset). The fabric/ceph rules themselves are unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): deploy MGR at a capped count, not one per host
placement.yml applied mgr to every active host (47 mgr daemons on spice:
1 active + 46 standby), which is wasteful and non-standard - MON already
uses a bounded ceph_mon set. Deploy count:{{ ceph_mgr_count }} (default 3:
1 active + 2 standby) instead; cephadm schedules them and caps at the host
count on small clusters like sietch.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): mark the RGW EC data pool as bulk
The EC data pool is pre-seeded to pg_num 4096 (sized for spice's 658 HDD
OSDs), but without the bulk flag the pg_autoscaler can walk it back toward 1
during the first fill, causing PG splitting under load. Set bulk=true so the
autoscaler targets a full-capacity pg_num and treats the pre-seed as a
floor. Idempotent: only set when not already true.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): harden the ceph_destroy SSD-model match
The SSD OSD PV-remove and partition-wipe loops used
grep -i "$ssd_model_pattern", which interpreted the pattern as a regex and,
on an empty pattern, matched every disk - the worst-case foot-gun in a
destroy path (it would target all disks). Use grep -iF (fixed string, no
regex) and skip the loop entirely when the pattern is empty. Kept -F
without -w, since -w would fail to match underscore-containing model
strings like Micron_5100_MTF... Shared role, so it hardens sietch too.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* chore(ceph): drop the dead s3_key_drift notice; reconcile harden header
The "S3 key drift notice" debug task was gated on s3_key_drift, a variable
nothing ever sets (the drift detection it stood in for was never built), so
the branch could never fire - remove it. And update the harden.yml header:
the security-critical sshd lockdown (root key-only, no password auth, locked
root) now runs early in baseline, not gated behind this last stage; harden
adds the firewall and the remaining sshd hardening on top.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): accurate change reporting in tuning roles (T4)
CRUSH tunables and per-device HDD/SSD sysfs writes ran a mutating
command every converge under changed_when:false, so a run reported ok
even when it changed state and --check misled register consumers.
Switch to get-then-set: read the current value first, only write when
it differs, and emit CHANGED so changed_when reflects reality. Same
settings applied; only change-detection becomes truthful. Also drops
two pre-existing em-dashes to keep the file plain ASCII.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* refactor(ceph): declarative lvg/lvol for OSD block.db setup
Replace hand-rolled pvcreate/vgcreate/lvcreate shell in ceph_deploy
lvm-setup with community.general.lvg/lvol (already pinned, and the
idiom provision_host/disks.yml already uses). Same VG/LV names, sizes,
fixed-vs-100%FREE split, per-shape loops, and device targets on both
shapes. Kept read-only asserts/verify/show tasks as-is.
Preserved the wipefs -af signature clear ahead of PV creation: lvg does
pvcreate -f but not --yes, so it will not wipe a stale foreign signature
on a reused partition. wipefs stays gated on the VG being absent so a
live PV is never touched. Kept a minimal shell to compute the spice
ssd-osd size (vg0 free - reserve) since lvol has no free-minus-N
primitive, plus an existence guard so re-runs don't recompute a bad
size or trigger an unforced lvol shrink.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* chore(ceph): ASCII-clean the lvm-setup header comment
Convert the pre-existing arrows and em-dashes in the header (left untouched
by the lvg/lvol refactor) to plain ASCII.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): never let lvol shrink a live block.db LV (data safety)
The lvg/lvol refactor let community.general.lvol reconcile an existing
fixed-size db-slot to its size, and lvol defaults to shrink:true - so a
db-slot that ever drifted LARGER would be TRUNCATED, corrupting a live
BlueStore block.db and taking out the OSD. The old shell skipped existing
LVs entirely, so it could never do this. Set module_defaults shrink:false
on both lvm-setup blocks: lvol still creates and may grow, but never
truncates an existing LV (it leaves a larger one alone). Also restore
opts:-Wy so (re)created LVs wipe stale signatures - the destroy->recreate
recovery-path freshness the explicit -Wy gave.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): RGW self-signed cert generation aborted on first converge
The openssl -subj and -addext args used backslash-newline continuations
inside their quoted strings, so shell line-folding kept the next line's
indentation in the value; countryName rendered as "DE " (over the
2-char max) and openssl exited 1 under set -euo pipefail, leaving no S3
endpoint on the first spice deploy. Precompute the subject and SAN as
single-line facts (whitespace-controlled Jinja for the SAN) and pass
them quoted, so each openssl arg is one flat string.
See ansible/ceph/roles/ceph_deploy/tasks/rgw.yml.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): green the ansible-lint CI gate on HEAD
Bare ansible-lint (what mise run lint and the required CI check run)
exited 2 on two deliberate patterns: run_once on the fleet-wide
placement assert, and an intentional no-pipefail shell in the ice DDP
Safe-Mode probe (pipefail there would mask grep's match). Waive both
with inline noqa so the reasoned patterns stay and the gate passes.
See ansible/ceph/deploy-ceph.yml and roles/baseline/tasks/nic-firmware.yml.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): pin RGW pools to device-class CRUSH rules (SSD fast tier)
The replicated_ssd/replicated_hdd rules were created but bound to no pool,
so every replicated pool fell to the default class-agnostic rule and the
omap-heavy bucket index landed on HDD (~98.5% of weight) while spice's
47-OSD NVMe ssd-osd tier sat idle. Pin the index, RGW metadata, and mgr
pools to the SSD rule and the bulk non-EC pool to HDD; force the OSD
device class at creation (ssd-osd -> ssd, HDD -> hdd) so the rules bind
deterministically, and assert the placement in verify.yml. Pinning is
opt-in per cluster (empty default) so a live cluster is never re-homed
implicitly; the pin runs after RGW readiness because the .rgw.* system
pools are created lazily by the realm/zone setup and the daemon.
See ansible/ceph/roles/ceph_deploy/{templates/osd-spec.yml.j2,tasks/rgw.yml,tasks/verify.yml}.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): apply RGW pool PG/durability pins on the first converge
The pool list was captured once at the top of the RGW play, before any
pool was created, so every "pool in list" gate skipped on a fresh run 1:
the data/index/non-EC PG pre-sizing deferred to a second converge, and
the .rgw.* system pools (created lazily by the daemon) never got their
size/min_size pin at all, sitting at the inherited 2/1 single-replica
until someone ran the play twice. Re-query the pool list once the
explicit pools exist, and move the system-pool PG + size/min_size pins
below the RGW readiness wait where those pools are real. Parameterize the
bootstrap osd_pool_default_size per cluster (spice 3/2, role default keeps
sietch at 2/1) so an unpinned future pool is not born single-replica, and
assert final size/min_size in verify.yml so a miss fails the play loudly.
See ansible/ceph/roles/ceph_deploy/tasks/{rgw.yml,verify.yml}.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ceph): real ok-to-stop gate before a live-node reimage
Reimaging a spice node zeroes the OS NVMe, and the 14 HDD OSD block.db
LVs plus the ssd-osd data LV all live on that NVMe, so a reimage destroys
every OSD on the node; the old guard only checked a hand-typed
i_know_its_dead flag and the prepare-os-disks comment wrongly claimed a
reimage keeps OSD data. Replace the stub with a mon-delegated check that
queries the node's OSD ids, refuses on HEALTH_ERR or a failed
ceph osd ok-to-stop, and sets noout on those OSDs for the reimage window.
The gate is default-on: auto mode enforces whenever a live cluster with
this node's OSDs is reachable and no-ops for the initial bootstrap;
strict fails closed if it cannot verify; permissive is the explicit
escape hatch.
See ansible/ceph/roles/reprovision_hetzner/tasks/ceph_safety.yml.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): jumbo MTU 9000 on the VLAN 122 cluster/replication network
Jumbo pays off on the OSD replication path (a closed, homogeneous fabric)
but is a partial-blackhole risk on the RGW-facing public network, which
serves heterogeneous 1500-byte S3 clients and the ~1400 NetBird overlay,
so raise MTU on VLAN 122 (Ceph private) only and leave 120/124 at 1500. A
VLAN opts into jumbo with a per-entry mtu; the bond ceiling and both 25G
members auto-raise to the largest child MTU so a 1500 parent or slave
cannot silently cap the jumbo frames. verify.yml asserts the applied MTU
on the bond, members, and each VLAN (local, no switch dependency). Flat
clusters resolve to 1500 and emit no MTU, so sietch is unchanged.
Paired with a QFX-side change (leaf server-LAG + VLAN 122 IRB >= 9000) and
a ping -M do -s 8972 matrix before the cluster network is relied on.
See ansible/ceph/roles/networkd.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* chore(dns): commit prod DNS provider lock file
The prod DNS stack was the only DNS stack without a committed
.terraform.lock.hcl, so cloudflare ~> 5.0 would resolve unpinned at the
first prod apply of the stack fronting the S3 endpoint. Pin it to 5.21.0,
matching the staging DNS stack, with linux_amd64 + darwin_arm64 hashes.
See tf/deployment/prod/global/dns/.terraform.lock.hcl.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): opt-in alertmanager receiver so alerts reach a human
cephadm deploys alertmanager with no receiver, so every Prometheus alert
rule (OSD down, host down, PG degraded, near-full) fires into the default
null route and nobody is notified. Add ceph_alertmanager_webhook_urls: a
non-empty list routes all alerts to those webhooks via cephadm
user_data.default_webhook_urls. The monitoring re-spec now applies when
either fabric bind or alerting is configured (independent gates), and
verify.yml reports the delivery status, warning loudly when none is set.
Empty by default, so flat clusters (sietch) are unchanged; the spice
destination is a deferred operator decision (TODO in defaults).
See ansible/ceph/roles/ceph_deploy/tasks/monitoring.yml.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(fabric): jumbo MTU on the leaf server bonds + VLAN 122 IRB
Completes the host-side VLAN 122 jumbo change on the QFX leaf: the server
LACP bonds (ae1..aeN) get the 9216 L2 jumbo MTU that the spine uplink ae0
already carries (802.3ad members inherit it, so no per-member mtu), and the
private/cluster IRB gateway (VLAN 122) gets an L3 family-inet MTU of 9000 to
match the Ceph hosts. Public (120) and host-mgmt (124) IRBs stay at 1500 by
design, so jumbo is confined to the closed OSD-replication path.
See tf/shared/modules/cluster-fabric.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ci): pin + SHA256-verify NetBird install, drop curl|sh
The netbird-connect action piped pkgs.netbird.io/install.sh straight into
sh on the runner that holds the prod-write 1Password token and overlay
access to the live nodes, so a compromised installer would run as root
there. Download a pinned release tarball (v0.74.4) from the GitHub release
and verify its SHA256 against an in-repo pin before unpacking; fail closed
on any mismatch. A version input plus a refresh comment keep the pin
maintainable.
See .github/actions/netbird-connect/action.yml.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* fix(ci): close the prod apply gate (reviewed-plan binding + fail-closed reviewers)
The prod approval gate could fail open and the apply was not bound to the
reviewed diff. Manage the four gate Environments as code (new
_meta/github stack: github_repository_environment + required reviewers,
protected-branches-only, a validation that rejects a reviewerless gate),
and make the discover job fail closed by asserting every emitted
partition-region Environment already carries a required reviewer. Bind
apply to the plan: the plan job writes -out to an absolute path and
uploads it, the gated apply downloads that exact file and applies it with
no re-plan and no -auto-approve, so a stale plan fails closed. Scope a
bare workflow_dispatch to plan-only and a toggled dispatch to just the
partition it touches; add fail-fast and timeout-minutes across the jobs.
The _meta/github stack is bootstrap-applied out of band and needs real
reviewer IDs plus a GH_ENV_ADMIN_TOKEN secret before the gate passes.
See .github/workflows/infra.yml and tf/deployment/_meta/github.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): blast-radius control on the day-2 converge plays
The converge chain (baseline, tune-os, tune-hardware, tune-ceph, harden)
ran against all nodes at once with no canary and no stop-on-failure, the
inverse of the serial:1/max_fail:0 destructive plays. Batch each converge
play (serial, default 10%) with max_fail_percentage:0, and gate every
batch on cluster health via a shared pre/post ceph-health checkpoint that
halts the roll on HEALTH_ERR. The gate is a no-op until a cluster is live,
so the bootstrap ordering (baseline -> tune -> deploy) is unaffected;
deploy-ceph stays big-bang. Tune the batch size with ceph_converge_serial.
See ansible/ceph/tasks/health_gate.yml.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): freeze load-bearing host packages instead of auto-upgrading
We run no unattended-upgrades: an apt upgrade that moved cephadm's
container runtime (podman + ecosystem) or the chrony time source that mon
quorum depends on out from under a live cluster is the uncoordinated
change we avoid. Hold those packages at their installed version (dpkg
selection) so apt upgrade skips them; upgrading is then deliberate and
health-gated. Diagnostic/ops tools are left unheld. Runs last in baseline
so nothing is held before it exists; tune via baseline_held_packages.
See ansible/ceph/roles/baseline/tasks/hold-packages.yml.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): THP=madvise and stop the GRUB cmdline clobber
Transparent Huge Pages sat at the kernel default 'always', which bloats
BlueStore/tcmalloc RSS and drives allocation-stall latency across the OSD
fleet; set it to madvise via a oneshot unit ordered Before=ceph-osd.target
(mirroring ceph-cpu-governor.service) plus a live sysfs write, on by
default for every ceph cluster. Separately, the processor.max_cstate GRUB
task rewrote GRUB_CMDLINE_LINUX_DEFAULT wholesale (dropping existing
tokens, injecting quiet) - a latent footgun behind the default-off
governor flag; replace it with a /etc/default/grub.d drop-in that only
appends the cstate token.
See ansible/ceph/roles/hardware_tuning.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): host security sysctls + persistent journald
Add a host kernel/network hardening sysctl set (syncookies, ignore/deny
ICMP redirects + source routing, kptr/dmesg restrict, log martians) to the
existing os_tuning sysctl drop-in. rp_filter is set to LOOSE (2), not
strict (1): these nodes are multi-homed (1G WAN default route + 25G fabric
VLAN sub-interfaces), so strict reverse-path filtering would blackhole
asymmetric fabric traffic. Separately make journald persistent
(Storage=persistent, bounded SystemMaxUse/RuntimeMaxUse) so logs survive
the pipeline's own reboots on these headless nodes.
See ansible/ceph/roles/os_tuning.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): health-gated cephadm upgrade playbook + rollback runbook
No upgrade automation existed for the pinned 20.2.x tentacle train. Add a
deliberate, manually-invoked upgrade-ceph.yml: preflight-gates on HEALTH_OK
(with an explicit allow-WARN toggle), no in-progress upgrade, and all OSDs
up+in; records and prints the rollback image before starting; refuses to
run without an explicit target (ceph_upgrade_target_image/_version, no
default); drives ceph orch upgrade with a bounded poll and fails loudly on
a stall; asserts health + version convergence after. Not wired into
site.yml. Rollback procedure lives in the play header and
docs/runbooks/upgrade-ceph.md.
See ansible/ceph/upgrade-ceph.yml.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): scheduled cluster-state backups with a systemd timer
Config/topology state was captured only manually and controller-local. Add
a ceph_backup role that installs a capture script + daily systemd timer on
the bootstrap node: each run tars fsid, config dump, mon/osd/crush maps
(raw + decoded), osd tree, orch ls/host ls, and the RGW realm/zonegroup/
zone into a root-only 0700 dir, pruned by retention. No secret keyrings are
dumped. An offsite target (rsync or s3://) is a var left empty for the
operator. Invoke via backup-ceph.yml (mise run backup-timer); not wired
into converge. Opt out with ceph_backup_enabled=false.
See ansible/ceph/roles/ceph_backup and docs/runbooks/backup-restore.md.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): pin mgr dashboard/prometheus bind to fabric; pin EC min_size
The mgr dashboard (8443) and prometheus module (9283) run inside the active
mgr and bound 0.0.0.0, closed only by nftables. Pin each mgr instance's own
localized server_addr (mgr/<module>/<mgr_id>/server_addr, the key the
modules read via get_localized_module_option) to that host's fabric IP,
discovering mgr ids at runtime so it is failover-safe; a global server_addr
would leave the dashboard unbindable after failover. Opt-in on
ceph_bind_networks; takes effect on the next mgr cycle (nftables holds the
WAN closed meanwhile). Separately pin the EC data pool min_size explicitly
to k+1 (derived from ceph_rgw_ec_k) so the write-availability floor is
documented and cannot drift, and record the single-site host-failure-domain
DR ceiling in docs/capacity-planning.md.
See ansible/ceph/roles/ceph_deploy/tasks/monitoring.yml.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* chore(ceph): freeze Galaxy collection versions to a major
The collection requirements floated on >= lower bounds, so a new
community.general major could be pulled mid-campaign. Pin compatible-release
ranges to the currently-resolving majors (ansible.posix >=2,<3;
community.general >=12,<13) so a 47-node campaign cannot cross a major
between runs. No lockfile mechanism exists in the repo, so the ranges live
in requirements.yml.
See ansible/ceph/requirements.yml.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* test(ceph): molecule scenario for the fabric-only firewall branch
The only molecule scenario tested the default open-firewall branch
(*_any_source: true), the opposite of spice's production posture, with no
idempotence check. Add a fabric-only scenario that flips the closed branch
(rgw/ssh any_source: false, trusted networks = fabric + NetBird) and
asserts RGW/SSH are NOT accepted from any source, are restricted to the
fabric/overlay, and that a second render is idempotent. Mirrors the default
scenario's render-and-grep idiom; the existing scenario is untouched.
See ansible/ceph/roles/security/molecule/fabric-only.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* ci(dns): drift-check the prod S3 roster against the ceph node list
The 47 S3 RGW A-records were hand-copied with no link to the ceph roster,
so add/replace/remove of a node silently blackholed the endpoint or dropped
capacity. A terragrunt dependency is not viable here (no dependency idiom in
the repo, and the ceph discovery output carries only bond_ip, never the
fabric IP), so add a stdlib-only, credential-free check that reconstructs
the expected fabric-IP roster from clusters.auto.tfvars (in-service names) +
spice-hosts.yaml (host_index) and asserts the apex and wildcard A-records
match exactly. A path-scoped CI gate fails the PR on drift before the DNS
apply; also runnable via mise run tf:check-dns-roster.
See tf/scripts/check-s3-dns-roster.py.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): inotify headroom for podman scale; podman-compat upgrade note
cephadm-on-podman runs many containers/systemd units per node and leans on
inotify; Debian's default fs.inotify.max_user_instances=128 is a known
bottleneck at that scale, so raise instances to 512 and watches to 524288
in os_tuning. Separately, document in the upgrade runbook that the podman
dpkg-hold must be lifted (unhold -> bump to a cephadm-supported version ->
re-hold) before a cross-major Ceph upgrade; patch-train upgrades need no
podman change.
See ansible/ceph/roles/os_tuning and docs/runbooks/upgrade-ceph.md.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(ceph): health-checked RGW ingress VIP (haproxy + keepalived)
S3 was bare round-robin DNS across 47 RGW A-records, so a dead node
blackholed ~1/47 of new connections for the TTL. Add an opt-in cephadm
ingress service: a keepalived floating VIP (spice 10.40.20.250) with
haproxy health-checking the RGW backends and dropping a dead one from
rotation. haproxy terminates TLS on the VIP with the existing self-signed
RGW cert (its SAN now carries the VIP) and re-encrypts to beast: L4
passthrough for RGW needs IngressSpec.use_tcp_mode_over_rgw, which is
absent on the pinned 20.2.2 (source-verified; applying it TypeErrors), so
terminate is the only health-checked mode available. Applies after RGW
readiness; asserts its preconditions (VIP set, ceph_bind_networks non-empty
so beast binds a per-node IP and does not collide with the VIP on :443,
ceph_rgw_ssl true). Opt-in per cluster; sietch unaffected.
See ansible/ceph/roles/ceph_deploy/templates/rgw-ingress-spec.yaml.j2.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* feat(dns): point the prod S3 endpoint at the ingress VIP
Collapse the s3.prod.fsn1.htz.futo.cloud apex + wildcard from 47 round-robin
A-records to the single health-checked ingress VIP 10.40.20.250, and switch
the drift-check to assert both records equal that VIP (the ceph
ceph_rgw_ingress_vip is the source of truth). Apply ordering matters and is
noted in the tfvars: the ingress must be live before this DNS cutover, or
s3 resolves to a VIP nothing answers; roll back in reverse.
See tf/deployment/prod/global/dns/records.auto.tfvars.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* chore(ceph): fix stale staging copy-paste comments in the prod stack; fmt
The prod ceph stack carried the staging stack's comments verbatim: versions.tf
and secrets.tf described creating SIETCH_CEPH_* items in yucca_tf_staging via
OP_TF_YUCCA_STAGING_ENV_WRITE, and a variable example said (sietch, ...). This
stack creates SPICE_CEPH_* in yucca_tf_prod via OP_TF_YUCCA_PROD_ENV_WRITE;
correct the comments to match. Legitimate cross-refs (the mirrors-staging/talos
provenance, the partition-slug enumeration) are left as-is. Also tofu fmt the
drifted clusters.auto.tfvars (whitespace only).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016FvsRx453wDkwzrEwcGgbK
* ci(infra): scope reviewer fail-closed gate to apply events; prettier-format workflows
* ci(infra): pin 1Password CLI version to avoid flaky latest resolution
* feat(dns): resolve prod S3 apex+wildcard to the node fabric-IP roster
* ci(infra): park the Environments approval gate (disable fail-closed step + stub _meta/github)
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* chore(netbird): move to kebab naming
Render NetBird object names (groups, setup keys, policies, networks,
network-resources) as lowercase-kebab instead of UPPER_SNAKE, e.g.
YUCCA_PROD_HTZ_FSN1_MGMT → yucca-prod-htz-fsn1-mgmt. The 1Password setup-key
item titles stay UPPER_SNAKE (decoupled) so CI/ansible/talos op:// consumers
keep resolving.
Pin the futo-org/netbird provider to 1.0.2, which fixes the group
resources TF→API decode so a resource-tag group (htz-fsn1 `resources`) can be
renamed in place — no name pin needed.
* update locks
* chore(naming): naming names
* chore(netbird): move to kebab naming
Render NetBird object names (groups, setup keys, policies, networks,
network-resources) as lowercase-kebab instead of UPPER_SNAKE, e.g.
YUCCA_PROD_HTZ_FSN1_MGMT → yucca-prod-htz-fsn1-mgmt. The 1Password setup-key
item titles stay UPPER_SNAKE (decoupled) so CI/ansible/talos op:// consumers
keep resolving.
Pin the futo-org/netbird provider to 1.0.2, which fixes the group
resources TF→API decode so a resource-tag group (htz-fsn1 `resources`) can be
renamed in place — no name pin needed.
* update locks
* docs(ceph): inline ADR rationale and drop the ADR set
Fold each linked ADR's rationale into the prose it supported, then remove
the ADR files, the README index row, and the stray code-comment reference --
no ADR trace remains.
True up the docs to the partition/region/ceph-cluster layout (#222) in the
same pass: state keys (yucca/<partition>/<region>/<stack>), inventory paths
(<partition>-<region>/<cluster>), stack dirs, and the recover runbook's S3
state paths. Reframe architecture's env section around partitions/regions
with sietch as staging/austin.
* docs(ceph): editorial pass to align docs with current code and CI/CD
Rewrite the ceph docs against the actual code rather than the pre-refactor
state:
- partition/region/ceph-cluster layout throughout: state keys
yucca/<partition>/<region>/<stack>, inventories <partition>-<region>/<cluster>,
the real clusters.auto.tfvars schema (partition/region, not environment/datacenter)
- sietch reframed as staging/austin with secrets in yucca_tf_staging; vault
hierarchy flipped from dev-primary to staging-primary
- live CI/CD (.github/workflows/infra.yml): per-partition read/write service
accounts delivered as GitHub secrets (OP_TF_YUCCA_<ENV>_ENV[_WRITE]),
plan/apply gating, NetBird overlay (Tailscale retired)
- correct the CEPH_ENV guidance (export works; deliberately kept out of mise
[env]) across README, CONTRIBUTING, scripts.md, adding-a-cluster
- drop the obsolete "read SA token from a 1P item" dance from the runbooks
- fix stale vaults, paths, examples, and the inventory-provision.ini name
* docs(ceph): transliterate docs to plain ASCII
Replace non-ASCII punctuation and box-drawing with ASCII equivalents across
the ceph docs: em/en dashes to --/-, middot separators to commas, arrows to
->, directory-tree box-drawing to |-- / `--, section sign to "section", x for
the multiply glyph, and >= / <= / ~ for the math glyphs. No content changes.
* docs(ceph): fix broken rotate-ssh-key link in scripts.md
The "Related" link pointed at runbooks/rotate-ssh-key.md, which does not
exist (a pre-existing dangling link). SSH-key rotation lives in
rotate-secrets.md; point at its "Rotating SSH keys" section.
* docs(ceph): style polish from per-doc review
Tighten verbal texture flagged by a per-doc style pass; no structural changes.
- correctness: ansible-play.sh runs ansible-playbook, it does not exec, so the
trap fires from the still-alive wrapper -- fix the "exec"/SIGKILL claims in
scripts.md, architecture.md, secrets.md
- cut recurring tics: "DR belt"/"belt-and-suspenders" -> "disaster recovery",
"a lost laptop is a non-event", "by design", "system mesh"
- unstuff long dash/semicolon sentences in architecture (vault-password
history, provision/baseline split), secrets (SSH-key paragraph), patterns
- recover-bad-tofu-apply: move the dormant-1P-items aside into one note,
consolidate the repeated caveats
- misc: naming grammar fix + drop trivia, hardware "would"/"blindly",
rotate-secrets "after confidence", drop a dead snippet line, fix the 16+
numeric hedge
* docs(ceph): make add-node and recover runbooks CI-aware
Now that infra.yml applies the stacks and runs the full ceph convergence on
merge, refresh the two runbooks the pipeline changed:
- add-node: lead with the manual-vs-CI split. The TF + host_vars change is a
PR; the only operator-only step is the physical provisioning (live-image
boot + provision.yml), which CI can't do; baseline/tune/join/harden run in
CI on merge. Keep the by-hand convergence as a documented fallback.
- recover-bad-tofu-apply: note that applies now run in CI with the partition
write SA, so the bad apply is usually a failed CI run; CI does not self-heal,
recovery is operator-run locally.
* docs(tf,talos): finish ADR purge into tf, fix sietch vault, mark talos dormant
- complete the ADR removal that stopped at ansible/ceph: drop the dangling
ADR-009/010 references from tf/README.md (link + related line) and the ceph
module / stack code comments, so no ADR trace remains repo-wide
- tf/README: the sietch cluster example uses yucca_tf_staging (was yucca_tf_dev)
- ansible/talos: add a "second-class, not actively used" status banner to the
README and architecture doc so readers don't treat the converged/libvirt
Talos docs as live
Left untouched: the Tailscale / SA-token / CI sections of tf/README -- those
are mid-migration in another owner's lane (bye-tailscale is in flight; fabric
still rides Tailscale by design).