Commit Graph
92 Commits
Author SHA1 Message Date
renovate[bot] 9d9a9cf765 chore(deps): update dependency ansible-lint to v26.8.0 (#482)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-09-15 12:35:32 +01:00
renovate[bot] 794f285e61 chore(deps): update dependency boto3 to v1.43.91 (#237)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-09-15 08:52:14 +01:00
renovate[bot] c982d82543 chore(deps): update dependency python to v3.14.7 (#278)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-09-15 02:01:54 +01:00
renovate[bot] b5dd78d1be chore(deps): update dependency ansible-core to v2.21.4 (#236)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-09-15 00:23:00 +01:00
Andy Molenda db5d9cecb8 chore(monk): pin 0.0.393 on spice and sietch (#656)
* chore(monk): pin 0.0.393 on spice and sietch

* chore(monk): request review of the 0.0.393 pin
2026-09-11 10:09:30 -07:00
Antoine Lecompte 3e213f0891 fix(o11y): show traffic rate in bits per second on all dashboards (#639) 2026-09-10 10:50:05 -04:00
Andy Molenda f0a5c49765 feat(ci): let the ceph converge target one playbook and host (#620)
* feat(ci): let the ceph converge target one playbook and host

The dispatch could only run the full provisioning pipeline, which never touched monk.yml at all, so a pin bump or a single-host canary had no CI path and ran from a workstation against prod. The playbook is an allowlisted choice mapped to existing mise tasks, and the limit is free text so it reaches the shell through the environment rather than a template expansion. harden and site had no mise task despite being the two playbooks operators run targeted most often. Directory-local invocation stays correct when #612 turns on monorepo mode.

* fix(ci): forward arguments through the ceph deploy task

mise appends task arguments to a single-command task but not to a script block, so a --limit passed to deploy was silently dropped and every host in the inventory converged. That made the new ceph_limit input a guard that reads as applied while doing nothing, on the option the dispatch defaults to. Verified against a stub: each play now receives the limit, and the no-argument case is unchanged.

* fix(ci): stop a ceph converge from applying prod terraform

run_ceph_ansible activates the partition so the converge can run, but the Terragrunt apply step was gated only on the stack not being fabric, so dispatching a converge also applied every prod stack with the write service account into environments that carry no reviewers. The inventory is rendered from TF state rather than from an apply, so the converge needs nothing applied first, and force_apply remains the explicit way to apply without a code change. Also fails the step when a limit leaves a play with no hosts, which ansible reports as success, and corrects the header and input descriptions that described the old coupling and oversold what a limit does.

* fix(ci): default the ceph converge to a read-only check and fail closed

ceph_playbook defaulted to deploy, so setting run_ceph_ansible and nothing else ran the six-play provisioning pipeline against a live cluster; status is the safe default and the destructive choices are now explicit. The no-hosts guard keys on ansible wording and fails open if that changes, so a recap check fails closed alongside it. Records that skipping the apply trades an implicit reconcile for a staleness window: state means the last apply, so a clusters.auto.tfvars change needs force_apply.

* docs(ci): record the verified monorepo task behavior

Monorepo mode is live after #612, so the claim about it is no longer a prediction: ansible/ceph is a config_root, the tasks are addressable as //ansible/ceph:<task>, and the bare name still resolves and still forwards arguments from inside the root. Verified both invocation forms against the merged change.

* style(ci): drop trailing whitespace left by the ascii pass

* fix(ci): only fail the converge when no play matched anything

The skip check fired on any play matching no hosts, so a site run narrowed to a non-mon host converged its other eight plays and was then reported as having converged nothing. A skipped play is now a warning and the recap check is the sole failure condition, which still catches the case that motivated the guard: monk alone under an off-group limit, where no play runs at all.
2026-09-08 12:04:17 -07:00
Devin Buhl b03e7d1c93 ci: use official mise action and lock tools for linux/macos x64/arm64 (#624)
Signed-off-by: Devin Buhl <devin@buhl.casa>
2026-09-03 09:36:36 -04:00
Andy Molenda d6df0764cd chore(ceph): bump the monk pin to the build with completion counters (#618)
Moves both clusters to `0.0.371`, the first build that counts scrub completions and tracks omap apart from data bytes (#617). Digest resolved from the registry index for the tag Deploy run 371 published.

Spice runs five mon instances and sietch three, all currently on `0.0.355`; the fleet digest matches that pin exactly today, so this is the whole delta.

The sietch comment named the previous tag and the reason for it, which stopped being true the moment the pin moved, so it now names what `0.0.371` adds instead. The rest of that file's diff is the ASCII normalization the repo asks for in files being touched.

Nothing consumes the new metrics yet: the pace and time-to-clear panels land once the converge has run and the counters have a predecessor snapshot to diff against.

- `main` <!-- branch-stack -->
  - \#618 :point\_left:
2026-09-02 05:15:52 -07:00
Andy Molenda c713549419 feat(ceph): enable monk on the spice mons and restore its scrape job (#615)
Flips monk on for the spice mons and restores the `ceph-scrub` scrape job, which #599 held back so that targets and exporter would land together.

The adelia canary collected cleanly against the live cluster: the interval read resolved spice's tuned 28d deep target rather than the 7d cold-start fallback, 4257 PGs classified with zero parse errors, on-host collection took 2.1s, and mgr CPU sat at 50.6% of one core, inside the 45.7-52.3% idle band measured before monk existed.

All five instances are scraped rather than one. Every instance returns the same cluster-wide data, so the redundancy costs about 3% of one mgr core and buys availability that matches the quorum's own fault model; a host rebuild gets the exporter back on the next converge because the role targets the TF-rendered `ceph_mon` group. Queries dedup with `max by (cluster, pool_id)`, applied identically to both sides of any ratio.

- `main` <!-- branch-stack -->
  - \#615 :point\_left:
2026-09-01 08:07:45 -07:00
Andy Molenda 55f37fe7d5 feat(ceph): pin monk on spice, gated behind a canary (#599)
Two parts.

**Pin and gate.** monk is pinned on spice to the digest sietch has been running since this morning, and left off. The fluent-bit precedent in the same file is the reason: `site.yml` carries `monk.yml`, so an enabled flag lands the exporter on all five mons at the next converge rather than on the canary host. The canary runs `harden.yml` first, because the `:9284` accept rule ships in `roles/security` and `roles/scrub_exporter` asserts it is live before staging anything; without that ordering the canary aborts on the assert, after `run_once` has already minted `client.scrub-exporter`.

**Scrape targets follow the exporter.** The `ceph-scrub` job shipped in #595 with five spice targets while monk runs on none of them. All five alert rules are `noDataState: OK`, so no targets is silence, but five unreachable targets synthesizes `up=0` and holds `MonkDown` firing per instance. The job returns with the enablement flip so targets and exporter land together.

Measured on adelia, which hosts the active mgr: `pg ls` is 11.5 MB over 4222 PGs and costs the mgr 0.78 cpu-s per call, from 20 saturated queries against a paired idle window. Five mons at the 2m default therefore cost it about 3% of one core, so the default stands.

- `main` <!-- branch-stack -->
  - \#599 :point\_left:
2026-08-31 13:47:25 -07:00
Andy Molenda b7629172be feat(ceph): enable monk on sietch (#591) 2026-08-31 07:33:16 -07:00
Andy Molenda 6eca8f094b fix(ceph): converge-safe pulls, live-ruleset gate, and container hardening for monk (#593)
* fix(ceph): converge-safe pulls, live-ruleset gate, caps reconcile, container hardening

* docs(ceph): monk rotation, upgrade, and retirement coverage

* fix(ceph): let the caps task own caps, split cgroups, guard check mode

* fix(ceph): pair Delegate with cgroups=split and probe collect_success
2026-08-31 06:25:54 -07:00
Andy Molenda be10b15a12 feat(ceph): deploy monk onto the mon hosts (#588)
* feat(ceph): deploy monk onto the mon hosts

* fix(ceph): target the ceph_mon group and no_log the keyring tasks

* chore(ceph): drop the monk interval vars, monk reads the cluster now

* fix(ceph): restart monk when the pull lands a new digest
2026-08-28 21:58:53 +00:00
Andy Molenda 30431b8d01 fix(ceph): drop wt0 interface alerts before they reach alertmanager (#539) 2026-08-25 10:06:11 -07:00
Antoine Lecompte 8392f4f379 feat(ceph): move user creation inside terraform (#541) 2026-08-25 06:32:50 -07:00
Andy Molenda 30eb44e6cb feat(ceph): log firewall drops over nflog and ship them to o11y (#518)
* feat(ceph): log firewall drops over nflog and ship them to o11y

* docs(ceph): note the drop-event source in the log-shipping headers
2026-08-21 13:34:58 -07:00
Andy Molenda 1cdf388cc9 feat(ceph): persist the HDD write-cache disable and guard it to rotational media (#531) 2026-08-21 13:23:28 -07:00
Antoine Lecompte bbcfa7c295 fix(cnpg): fix backup alerts (#522) 2026-08-21 13:17:57 +00:00
Antoine Lecompte 045491d437 feat(cnpg): make better (#496) 2026-08-19 15:27:09 -04:00
Andy Molenda d5972ef980 fix(ceph): deliver recovery progress and repeat chronic alerts every 12h (#487) 2026-08-19 15:59:18 +00:00
Andy Molenda fc0b712a24 fix(ceph): inhibit CephPGsUnclean while backfill is moving (#484) 2026-08-19 07:59:37 -07:00
Andy Molenda 5eb353a498 feat(ceph): install hdparm on every ceph node (#479) 2026-08-19 07:59:16 -07:00
Andy Molenda 9091c1fc9d fix(ceph): group alertmanager notifications by pool (#463) 2026-08-18 06:19:18 -07:00
Andy Molenda 48d09d79af feat(ceph): rebuild a replaced OSD onto its intended db slot (#462)
* fix(ceph): let read-only probes run in check mode

* feat(ceph): add replace-osd playbook for a swapped HDD

* docs(ceph): correct the spice disk replacement procedure

* docs(ceph): hetzner hot-swaps SX295 drives on a live host

* feat(ceph): let replace-osd keep the old osd id

* docs(ceph): reserve the osd id before the replace-osd rebuild

* docs(ceph): spell out what orch osd rm --replace does and does not do

* fix(ceph): verify the rebuilt OSD by its own report and survive a queued stop

* docs(ceph): replace-osd when the cluster is not active+clean

* fix(ceph): let the queued action fire before judging the rebuilt OSD

* fix(ceph): refuse an explicit db slot the scan did not find free
2026-08-18 05:53:35 -07:00
renovate[bot] f6762d31bc chore(deps): update dependency ansible-lint to v26.6.0 (#250)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-08-06 15:18:35 +01:00
renovate[bot] 7acb2bd6f3 chore(deps): update dependency ansible-core to v2.20.7 [security] (#458)
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
2026-08-06 15:11:39 +01:00
Andy Molenda 5d60aae2fd fix(ceph): repair the monitoring path from spec apply to Zulip message (#408) 2026-08-03 13:05:04 +00:00
Andy Molenda 5e047f0215 feat(ceph): move the ceph config model into terraform, deep scrub interval to global (#410)
Claude-Session: https://claude.ai/code/session_01UXpK3sqBGrsJQFwP6MeXxd
2026-08-03 06:04:47 -07:00
Andy Molenda 8a2feb5128 docs(ceph): runbook for triaging a host with a dead root filesystem (#406) 2026-08-03 13:04:29 +00:00
Andy Molenda cbf9114b28 fix(ceph): skip the dpkg unhold on nodes where the package is absent (#403) 2026-07-31 15:43:47 +00:00
Andy Molenda 7cfb06494b fix(ceph): deliver alertmanager alerts to the configured receiver (#404) 2026-07-31 15:43:29 +00:00
Andy Molenda 957afec6e3 feat(ceph): ship spice host logs to o11y over netbird (#378)
* feat(ceph): ship spice host logs to o11y over netbird

* refactor(ceph): use fluent-bit instead of vector for host logs

* docs(ceph): tighten fluent-bit role comments

* fix(ceph): compress host logs to o11y with zstd

* fix(ceph): pin fluent-bit 5.0.9 and converge log shipping in site.yml

* refactor(ceph): resolve o11y over mesh DNS instead of a hosts entry

* fix(ceph): run log shipping last and tolerate isolated node failures

* fix(ceph): check log shipping preconditions before mutating the node

* fix(ceph): bind fluent-bit metrics to the fabric and open the port

* test(ceph): add a molecule scenario for the fluent-bit config

* fix(ceph): verify the o11y cert hostname, not just the chain

* fix(ceph): wait for the bind address so the metrics endpoint survives a reboot

* docs(ceph): plain-technical pass over the fluent-bit comments

* docs(ceph): cite the source for the journald field names

* chore(ceph): hold host log shipping off until the canary passes
2026-07-30 10:56:40 -07:00
Andy Molenda 289a2ee03b fix(ceph): widen spice deep scrub interval and unblock scrub scheduling (#387) 2026-07-30 05:53:44 -07:00
Andy Molenda a54c4f59af docs(ceph): treat the spice 1G WAN default route as permanent (#388)
* docs(ceph): drop the resolved incident from the spice alert comment

* docs(ceph): treat the spice 1G WAN default route as permanent
2026-07-29 12:07:35 -07:00
Andy Molenda ce575b1113 feat(ceph): make cluster config a layered, mask-aware model (#377) 2026-07-29 10:17:07 -07:00
Andy Molenda c606d25c8e fix(ceph): only commit the rgw period when something changed (#372) 2026-07-29 07:59:28 -07:00
Andy Molenda 65d75e2d28 fix(ceph): expect the declared MON quorum in drift detection (#374) 2026-07-29 07:55:24 -07:00
Andy Molenda ee7fd68800 fix(ceph): compare drift against declared config, not role defaults (#373) 2026-07-29 07:55:11 -07:00
Andy Molenda 3104a35d78 fix(ceph): take osd specs out of the cephadm serve loop (#364)
* docs(ceph): cover spice alongside sietch in the ceph docs

* fix(ceph): take osd specs out of the cephadm serve loop

* fix(ceph): keep spice on per-host osd specs, collapse is greenfield-only

* docs(ceph): correct the committed-inventory claim for both clusters
2026-07-29 12:28:37 +00:00
Andy Molenda 4ef8a1886c feat(ceph): route spice alertmanager to a real receiver (#355) 2026-07-27 11:55:07 -07:00
Andy Molenda efe72b89a4 docs(ceph): cluster-profiles reference and per-cluster runbook parameterization (#354)
* docs(ceph): add a cluster-profiles reference for sietch and spice

* docs(ceph): parameterize the cert-rotation and backup runbooks per cluster
2026-07-27 11:35:51 -07:00
Andy Molenda a499e41327 docs(ceph): spice disk-replacement runbook and db-slot pairing note (#353)
* docs(ceph): add the spice branch of the disk-replacement runbook

* docs(ceph): record that the ceph_hdd_osds db pairing is not binding
2026-07-27 11:25:25 -07:00
Andy Molenda 866d69f712 fix(ceph): pin SATA link power to max_performance on OSD nodes (#347)
* fix(ceph): pin SATA link power to max_performance on OSD nodes

* feat(ceph): drift-check the SATA link power management policy
2026-07-27 10:20:34 -07:00
Andy Molenda 03d9bdd497 feat(ceph): enable telemetry phone-home for transparency (#316) 2026-07-24 06:10:30 -07:00
Antoine Lecompte 9a7b72d89c fix(mgmt): vlan interface names exceeded IFNAMSIZ (#305) 2026-07-22 19:17:06 +00:00
Andy Molenda 526bbfcca5 feat(ceph): passwordless ops sudo on sietch and spice (#299)
* feat(ceph): passwordless ops sudo on sietch

* feat(ceph): passwordless ops sudo on spice
2026-07-22 10:57:22 -07:00
Andy Molenda b2d49cce46 feat(ceph): netbird ssh server for spice (sso-gated human access) (#290) 2026-07-22 06:57:36 -07:00
Andy Molenda 2881be69a4 feat(ceph): netbird enrollment role for the spice nodes (#287) 2026-07-22 06:08:44 -07:00
Andy Molenda 4782ef8d69 feat(ceph): trust the netbird overlay interface in nftables (#285) 2026-07-21 15:20:34 -07:00