Commit Graph
505 Commits
Author SHA1 Message Date
Antoine Lecompte dcf3065fee fix(michael): answer GET of a missing blob or config with 404, not 500 (#650) 2026-09-11 08:49:06 -04:00
Antoine Lecompte 3e213f0891 fix(o11y): show traffic rate in bits per second on all dashboards (#639) 2026-09-10 10:50:05 -04:00
immich-push-o-matic[bot] edfe1c9aaa chore(main): release 0.40.1 (#638) v0.40.1 2026-09-10 13:40:26 +01:00
Paul Makles 1e4e1b0612 fix: zod 4.4.3 - explicit defaults for coerce.boolean (#637) 2026-09-10 13:34:00 +01:00
immich-push-o-matic[bot] 58b1c0418b chore(main): release 0.40.0 (#635) v0.40.0 2026-09-10 11:15:31 +01:00
Paul Makles eed66845c8 feat: use zod for validation (#636) 2026-09-10 11:13:02 +01:00
Paul Makles 87e0066ff2 fix: more types that should be integers (#634)
Signed-off-by: izzy <me@insrt.uk>
2026-09-09 12:33:01 +01:00
immich-push-o-matic[bot] 59f2eaf84d chore(main): release 0.39.0 (#609) v0.39.0 2026-09-09 07:05:55 -04:00
Paul Makles 7a04d36482 fix: explicitly mark keepLast as an integer (#632) 2026-09-09 10:45:45 +00:00
Paul Makles 037745c682 feat: use restic proxy in orchestrator when available (#631) 2026-09-09 11:02:19 +01:00
Andy Molenda cadcba2389 fix(ceph): let recovery outrank scrubbing on spice (#630) 2026-09-08 12:29:30 -07:00
Andy Molenda f0a5c49765 feat(ci): let the ceph converge target one playbook and host (#620)
* feat(ci): let the ceph converge target one playbook and host

The dispatch could only run the full provisioning pipeline, which never touched monk.yml at all, so a pin bump or a single-host canary had no CI path and ran from a workstation against prod. The playbook is an allowlisted choice mapped to existing mise tasks, and the limit is free text so it reaches the shell through the environment rather than a template expansion. harden and site had no mise task despite being the two playbooks operators run targeted most often. Directory-local invocation stays correct when #612 turns on monorepo mode.

* fix(ci): forward arguments through the ceph deploy task

mise appends task arguments to a single-command task but not to a script block, so a --limit passed to deploy was silently dropped and every host in the inventory converged. That made the new ceph_limit input a guard that reads as applied while doing nothing, on the option the dispatch defaults to. Verified against a stub: each play now receives the limit, and the no-argument case is unchanged.

* fix(ci): stop a ceph converge from applying prod terraform

run_ceph_ansible activates the partition so the converge can run, but the Terragrunt apply step was gated only on the stack not being fabric, so dispatching a converge also applied every prod stack with the write service account into environments that carry no reviewers. The inventory is rendered from TF state rather than from an apply, so the converge needs nothing applied first, and force_apply remains the explicit way to apply without a code change. Also fails the step when a limit leaves a play with no hosts, which ansible reports as success, and corrects the header and input descriptions that described the old coupling and oversold what a limit does.

* fix(ci): default the ceph converge to a read-only check and fail closed

ceph_playbook defaulted to deploy, so setting run_ceph_ansible and nothing else ran the six-play provisioning pipeline against a live cluster; status is the safe default and the destructive choices are now explicit. The no-hosts guard keys on ansible wording and fails open if that changes, so a recap check fails closed alongside it. Records that skipping the apply trades an implicit reconcile for a staleness window: state means the last apply, so a clusters.auto.tfvars change needs force_apply.

* docs(ci): record the verified monorepo task behavior

Monorepo mode is live after #612, so the claim about it is no longer a prediction: ansible/ceph is a config_root, the tasks are addressable as //ansible/ceph:<task>, and the bare name still resolves and still forwards arguments from inside the root. Verified both invocation forms against the merged change.

* style(ci): drop trailing whitespace left by the ascii pass

* fix(ci): only fail the converge when no play matched anything

The skip check fired on any play matching no hosts, so a site run narrowed to a non-mon host converged its other eight plays and was then reported as having converged nothing. A skipped play is now a warning and the recap check is the sole failure condition, which still catches the case that motivated the guard: monk alone under an off-group limit, where no play runs at all.
2026-09-08 12:04:17 -07:00
Devin Buhl b03e7d1c93 ci: use official mise action and lock tools for linux/macos x64/arm64 (#624)
Signed-off-by: Devin Buhl <devin@buhl.casa>
2026-09-03 09:36:36 -04:00
Paul Makles 312062e33c feat: standalone app login & work around Zitadel race condition (#556) 2026-09-03 13:35:35 +01:00
Paul Makles f31f9be024 feat: destructive actions - delete repos & disable write-only (#607) 2026-09-03 13:02:52 +01:00
Paul Makles 76d48c26b4 feat: restic proxy (#611)
* feat(restic proxy): initial commit

* feat(restic proxy): working impl.

* fix(restic proxy): use repoId as grant key

* fix(restic proxy): handle http errors for meta

* fix(restic proxy): close http response body

* fix(restic proxy): use path not encoded path

* fix(restic proxy): unset RawPath

* docs(restic proxy): notes for later work

* fix(restic proxy): timeout on meta requests

* fix(restic proxy): concurrent minting

* fix(restic proxy): don't nil error

* fix(restic proxy): robust claims/exp check

* docs(restic proxy): document next work

* fix(restic proxy): respect log level

* chore(restic proxy): set default log level to info

* fix(restic proxy): permit temporary failures

* refactor(restic proxy): remove exit dead code

* chore: mise config

* feat(restic proxy): configurable well known URL

* chore(restic proxy): use real Exp

* refactor(restic proxy): use username as repoId

* fix(restic proxy): check repositoryId is set

* test(restic proxy): generate tests

* feat(restic proxy): publish Dockerfile

* feat(standalone app): include restic-proxy binary

* refactor(restic proxy): allow configuring api & meta URLs
test: generate accompanying tests

* chore: set config_roots

* test: drop comments from generated tests

* docs(restic proxy): add to USER_MANUAL

* fix(restic proxy): remint grant if auth token changes

* chore(restic proxy): update mise.toml

* refactor(restic proxy): differentiate session token w. 'token'
feat(restic proxy): denial cache
test(restic proxy): update generated tests for deny cache

* chore: remove `restic-proxy:check` from `check` dependencies (covered by
lint)
2026-09-02 16:37:31 +01:00
Paul Makles f8497b8033 feat: enable mise mono repo (#612)
* feat: enable mise mono repo

* fix: set config_roots

* ci: update mise action for software/publish

* chore: generate mise.lock

* ci: only update `ci.yml` with mise 2026.8.6 (no lock)

* ci: prefer `mise run`
2026-09-02 16:37:31 +01:00
Antoine Lecompte 85c936abc6 feat(bot): add dashboard (#622) 2026-09-02 14:34:46 +00:00
Andy Molenda d6df0764cd chore(ceph): bump the monk pin to the build with completion counters (#618)
Moves both clusters to `0.0.371`, the first build that counts scrub completions and tracks omap apart from data bytes (#617). Digest resolved from the registry index for the tag Deploy run 371 published.

Spice runs five mon instances and sietch three, all currently on `0.0.355`; the fleet digest matches that pin exactly today, so this is the whole delta.

The sietch comment named the previous tag and the reason for it, which stopped being true the moment the pin moved, so it now names what `0.0.371` adds instead. The rest of that file's diff is the ASCII normalization the repo asks for in files being touched.

Nothing consumes the new metrics yet: the pace and time-to-clear panels land once the converge has run and the counters have a predecessor snapshot to diff against.

- `main` <!-- branch-stack -->
  - \#618 :point\_left:
2026-09-02 05:15:52 -07:00
Andy Molenda b9d60c667f feat(monk): count scrub completions and track omap bytes (#617)
Backlog size and age answer how far behind the cluster is; nothing answered how fast it drains. Ceph exports no per-PG completion counter, and the old dashboard approximated pace from OSD scrub read-byte counters, which mixed raw bytes with logical ones and counted a re-scrub of the same PG as progress.

monk now diffs each PG's stamp against the previous refresh and counts the PGs and bytes that actually completed a scrub. A PG missing from either side is skipped rather than counted, so pool deletion and PG splits do not register as completions, and the first refresh after start has no predecessor so the counters simply begin at zero.

These are counters, so they dedup differently from the gauges: rate first, then `max by (cluster, pool_id)`. Taking `max` across instances before `rate` would read an instance restart as a jump, because the five instances start at different times and each counts independently. The README carries that rule and the time-to-clear query.

Also tracks omap apart from data bytes. `stat_sum.num_bytes` excludes omap, so `prod-z1.rgw.buckets.index` reports zero bytes while holding 164 GB across 524M keys on spice: every one of its PGs could go overdue without moving byte-weighted coverage. Existing metrics keep their meaning and `ceph_scrub_pool_bytes` still matches `ceph_pool_stored`; the omap figures are additive.

- `main` <!-- branch-stack -->
  - \#617 :point\_left:
2026-09-01 14:33:22 -07:00
Andy Molenda 078c8f756c feat(o11y): put the scrub work-outstanding row on measured metrics (#616)
Replaces the inferred work-outstanding row with monk's measured metrics.

The old row derived coverage from `ceph_osd_scrub_*_read_bytes` against a `$deep_interval_days` offset. That estimate needed 28 days of stored metrics before it meant anything, counted re-scrubs of the same PG as progress, and divided read bytes by `ceph_osd_stat_bytes_used`. It also depended on two hardcoded interval variables that could drift from what the scrub scheduler actually targets.

Coverage, overdue bytes, overdue PGs and the age distribution now come from per-PG stamps. Every query dedups with `max by (cluster, pool_id)` because all five monk instances report the same cluster-wide data, and both sides of the coverage ratio are stored bytes, so nothing mixes logical with raw. The two interval variables are gone: panels judge against `ceph_scrub_target_interval_seconds`, read live from the cluster.

`Oldest deep scrub vs target` is stated as a ratio rather than days on purpose, so a single threshold is correct on both spice at 28 days and sietch at 7.

Validation: all ten expressions parse against o11y, `ceph_pool_metadata` has no duplicate `(cluster, pool_id)` so the `group_left(name)` join cannot fail, and running the identical join shape with `ceph_pool_stored` returns `prod-z1.rgw.buckets.data = 8252.03 TiB`, matching monk's own `ceph_scrub_pool_bytes` to the same figure. Panel values cannot be checked until the scrape job reaches prod at the next release promotion, since prod Flux is tag-pinned.

- `main` <!-- branch-stack -->
  - \#616 :point\_left:
2026-09-01 14:29:07 -07:00
Andy Molenda c713549419 feat(ceph): enable monk on the spice mons and restore its scrape job (#615)
Flips monk on for the spice mons and restores the `ceph-scrub` scrape job, which #599 held back so that targets and exporter would land together.

The adelia canary collected cleanly against the live cluster: the interval read resolved spice's tuned 28d deep target rather than the 7d cold-start fallback, 4257 PGs classified with zero parse errors, on-host collection took 2.1s, and mgr CPU sat at 50.6% of one core, inside the 45.7-52.3% idle band measured before monk existed.

All five instances are scraped rather than one. Every instance returns the same cluster-wide data, so the redundancy costs about 3% of one mgr core and buys availability that matches the quorum's own fault model; a host rebuild gets the exporter back on the next converge because the role targets the TF-rendered `ceph_mon` group. Queries dedup with `max by (cluster, pool_id)`, applied identically to both sides of any ratio.

- `main` <!-- branch-stack -->
  - \#615 :point\_left:
2026-09-01 08:07:45 -07:00
Antoine Lecompte 56a1b18061 feat(cnpg): enable quorum commit on the real database clusters (#613) 2026-09-01 09:57:18 -04:00
Antoine Lecompte ec696ed762 fix(flux): make the database health check actually apply (#614) 2026-09-01 09:57:06 -04:00
Antoine Lecompte 2044b5095e fix(flux): gate the database kustomization on a serving primary (#610) 2026-09-01 09:26:14 -04:00
Antoine Lecompte f1640326b9 feat(infra): activate the futo-backups-bot prod creds (#601) 2026-09-01 09:00:41 -04:00
immich-push-o-matic[bot] cf3795a196 chore(main): release 0.38.0 (#538) v0.38.0 2026-09-01 07:00:32 -04:00
Paul Makles a08e47ff8a fix: align migrations to actual schema & correct migrations task (#608) 2026-09-01 11:41:17 +01:00
Antoine Lecompte 7750b71620 fix(futo-backups-bot): ignore interactions from other guilds (#602) 2026-08-31 15:42:57 -07:00
Andy Molenda 55f37fe7d5 feat(ceph): pin monk on spice, gated behind a canary (#599)
Two parts.

**Pin and gate.** monk is pinned on spice to the digest sietch has been running since this morning, and left off. The fluent-bit precedent in the same file is the reason: `site.yml` carries `monk.yml`, so an enabled flag lands the exporter on all five mons at the next converge rather than on the canary host. The canary runs `harden.yml` first, because the `:9284` accept rule ships in `roles/security` and `roles/scrub_exporter` asserts it is live before staging anything; without that ordering the canary aborts on the assert, after `run_once` has already minted `client.scrub-exporter`.

**Scrape targets follow the exporter.** The `ceph-scrub` job shipped in #595 with five spice targets while monk runs on none of them. All five alert rules are `noDataState: OK`, so no targets is silence, but five unreachable targets synthesizes `up=0` and holds `MonkDown` firing per instance. The job returns with the enablement flip so targets and exporter land together.

Measured on adelia, which hosts the active mgr: `pg ls` is 11.5 MB over 4222 PGs and costs the mgr 0.78 cpu-s per call, from 20 saturated queries against a paired idle window. Five mons at the 2m default therefore cost it about 3% of one core, so the default stands.

- `main` <!-- branch-stack -->
  - \#599 :point\_left:
2026-08-31 13:47:25 -07:00
Antoine Lecompte bb0cd74bee feat(columbo): give the agent fixed fleet-health probes and a bigger tool budget (#600) 2026-08-31 16:10:02 -04:00
Andy Molenda a1f5454e3e ci: publish release images on the release event and gate PR image builds (#594)
Two independent parts.

**Release delivery.** `v` tags publish on the release event, in a new Release Images workflow, rather than from whichever Deploy run survives the concurrency queue. Deploy now queues FIFO instead of discarding superseded runs, and ships only build tags. Release Images waits out that commit's Deploy, then retags the matching `sha-` image per app, or builds from the tag tree when none exists; it is re-runnable against any tag.

**PR image gating.** build-check runs on every PR and aggregates into one `Result` check, so image builds become requirable in branch protection. The per-app job names are dynamic and cannot be.
2026-08-31 12:13:18 -07:00
Antoine Lecompte 8e9cd64c16 feat(columbo): retry model calls with per-attempt deadlines (#598) 2026-08-31 11:28:12 -07:00
Andy Molenda b7629172be feat(ceph): enable monk on sietch (#591) 2026-08-31 07:33:16 -07:00
Andy Molenda 30c38f7f8e feat(o11y): scrape monk on the spice mons and alert on its health (#595)
* feat(o11y): scrape monk on the spice mons and alert on its health

* fix(o11y): cover dead monk instances and fit staleness inside snapshot expiry

* fix(o11y): name the rule that actually persists past snapshot expiry

* fix(o11y): fire collect-failing before the staleness rule expires
2026-08-31 06:41:33 -07:00
Andy Molenda 16d9c700ce fix(monk): harden failure paths, exec handling, and the metric contract (#592)
* fix(monk): harden failure paths, exec handling, and the metric contract

* fix(monk): log collection state transitions, not every failing refresh

* fix(monk): correct exec cancellation, waitdelay scope, and log levels
2026-08-31 06:26:57 -07:00
Andy Molenda 6eca8f094b fix(ceph): converge-safe pulls, live-ruleset gate, and container hardening for monk (#593)
* fix(ceph): converge-safe pulls, live-ruleset gate, caps reconcile, container hardening

* docs(ceph): monk rotation, upgrade, and retirement coverage

* fix(ceph): let the caps task own caps, split cgroups, guard check mode

* fix(ceph): pair Delegate with cgroups=split and probe collect_success
2026-08-31 06:25:54 -07:00
Antoine Lecompte ce32472f89 fix(michael): log the resolved route plus op and blob_type on access lines (#597) 2026-08-31 08:38:17 -04:00
Antoine Lecompte 500107b5ce feat(columbo): triage-refusal staff notes + per-user telemetry catalog (#596) 2026-08-31 08:38:05 -04:00
Andy Molenda be10b15a12 feat(ceph): deploy monk onto the mon hosts (#588)
* feat(ceph): deploy monk onto the mon hosts

* fix(ceph): target the ceph_mon group and no_log the keyring tasks

* chore(ceph): drop the monk interval vars, monk reads the cluster now

* fix(ceph): restart monk when the pull lands a new digest
2026-08-28 21:58:53 +00:00
Andy Molenda 919398b0c9 feat(monk): add monk, a measured ceph scrub-backlog exporter (#587)
* feat(monk): add monk, a measured ceph scrub-backlog exporter

* fix(monk): harden collection failure paths

* docs(monk): defer the deploy role reference to its follow-up PR

* fix(monk): emit the age distribution as a real histogram and document operational contracts

* feat(monk): read overdue targets from the cluster instead of flags

* ci(monk): register the monk image build filter

* chore(monk): cut comments the types already carry
2026-08-28 14:55:11 -07:00
Andy Molenda 768a259c9d ci: build changed app images on pull requests (#589)
* ci: build changed app images on pull requests

* ci: cover shared and root image inputs in build-check

* ci: treat .dockerignore as an image input

* ci: dockerignore governs the go image contexts too

* ci: apps register their build filter in the PR that adds them
2026-08-28 14:20:25 -07:00
Andy Molenda e2a7f657a8 test(yucca-admin-api): expect discordLink in the single-user response (#590) 2026-08-28 13:11:29 -07:00
Antoine Lecompte b3aa57dc5d feat(columbo): per-investigation audit log, tool-error recovery, pipe-aware log scoping (#586) 2026-08-28 13:57:28 -04:00
Antoine Lecompte 88267dc1f7 feat(columbo): add columbo (#585) 2026-08-28 13:26:11 -04:00
Antoine Lecompte 7988ac3487 fix(futo-backups-bot): fall back to a generic agent name when the profile read is forbidden (#584) 2026-08-27 12:51:41 -07:00
Antoine Lecompte 001da241f7 chore(tf): consume the freshdesk-stack webhook credentials in the talos stacks (#582) 2026-08-27 15:33:58 -04:00
Antoine Lecompte 73006bd23b chore(tf): bump slop-place/freshdesk to 0.1.2 for the automation-action passthrough fix (#581) 2026-08-27 15:21:10 -04:00
Antoine Lecompte fa9cfd3891 refactor(tf): move freshdesk ownership into partition-level global/freshdesk stacks (#580) 2026-08-27 14:35:04 -04:00
Antoine Lecompte 64371fbaa9 feat(infra): wire the freshdesk webhook route, network policy and secrets (#578) 2026-08-27 13:09:21 -04:00