Commit Graph
115 Commits
Author SHA1 Message Date
Andy Molenda 8d5bb532e3 feat(ceph): add staging/ceph terraform stack (#165) 2026-06-25 07:47:11 -07:00
Andy Molenda 0794082771 fix(ceph): parameterize secrets vault for non-dev clusters (#163) 2026-06-25 06:53:23 -07:00
immich-push-o-matic[bot] cd9ac63b2b chore: release main (#161) 2026-06-25 12:48:49 +00:00
Paul Makles 1d94bebcdc fix: missing GatewayEvent export (#162) 2026-06-25 12:46:23 +00:00
immich-push-o-matic[bot] 96119643ea chore: release main (#141) v0.4.1 2026-06-25 12:25:00 +00:00
Paul Makles 4956b5105b fix: package imports for orchestrator api (#160) 2026-06-25 12:04:47 +00:00
Andy MolendaandClaude Opus 4.8 1cf742b0f4 fix(ceph): converge RGW S3 user to exactly the 1P key on rotation (#159)
Step 15 previously only added the canonical 1P key on drift, leaving any
cephadm-generated random key in place as a second valid credential, and its
drift check only inspected keys[0]. When rotate_s3_keys is set, converge to
exactly the canonical key: add it if missing, then retire every non-canonical
key. The canonical key is added before any stray is removed, so the user is
never left without a working key. Still gated behind rotate_s3_keys because
rotating a live S3 key is disruptive to current consumers.

Verified on sietch (already reconciled): with rotate_s3_keys=true the run
reports canonical_present=True stray_count=0 and makes no changes.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 09:58:52 -07:00
Andy MolendaandClaude Opus 4.8 7eaa9e2993 fix(ceph): open cephadm service discovery port 8765 in firewall (#156)
The security role allowlisted every monitoring port except 8765, the cephadm
service discovery endpoint on the mgr. Prometheus fetches its scrape target
list from there via http_sd, so with the port dropped it discovered zero
targets and Grafana showed no data.

Add ceph_firewall_service_discovery_port (8765) to the role defaults and the
nftables template. Applied to sietch: the SD endpoint now returns 200 from the
Prometheus host, Prometheus has 7 active targets up, and ceph metrics flow
again.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 09:16:36 -07:00
Andy MolendaandClaude Opus 4.8 27ffe9d1e7 fix(ceph): enforce Grafana admin login password from the vault (#157)
cephadm seeds the Grafana admin login password only at first deploy, and
nothing reconciled it, so it drifted from the vault. Direct login at :3000
failed, and the dashboard Grafana API calls (which authenticate as the same
admin user) were also affected.

Add an idempotent task that resets the Grafana admin password to
ceph_grafana_admin_password on every converge via grafana cli
reset-admin-password, which writes the sqlite DB on the host volume so it
persists across restarts. Also rename the existing task to make clear it sets
the dashboard Grafana API password, not the admin login.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 09:15:31 -07:00
Andy MolendaandClaude Opus 4.8 94e7b993fd fix(ceph): enforce dashboard admin password on every converge (#154)
The dashboard password was only set during the one-time cephadm bootstrap
(guarded by `not ceph_conf.stat.exists`), with no idempotent reconcile, so a
cluster bootstrapped out of band kept a random admin password that no later
playbook run would correct.

Add an enforce task that sets the password from ceph_dashboard_password on
every run, plus a verify-phase assertion that logs into the dashboard with the
vaulted credential and fails the play on drift.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 08:15:50 -07:00
Andy MolendaandClaude Opus 4.8 83dab28c91 fix(tf/ceph): stop writing rendered inventories via local_file (#155)
local_file stored the destination path in shared remote state, derived from
get_repo_root(). git worktrees each resolve get_repo_root() to their own root,
so state became bound to whichever worktree last applied. Any apply from a
different checkout then force-replaced every rendered file and rebound state.

The module now emits inventory and secrets content as a render output instead
of local_file resources. ansible/ceph/scripts/render-inventories.sh reads that
output and writes the files using a path derived from its own location, so no
checkout-specific path ever enters state.

Also exclude painbox from the managed clusters (in active use by Zack); its
spec moves to clusters.example.tfvars and its 1Password items are left alone.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 08:15:46 -07:00
Antoine Lecompte 03e2802c06 fix(staging): correct vmauth url (#153) 2026-06-24 10:21:38 -04:00
Antoine Lecompte 1736b7cd3a fix(staging): set web port (#152) 2026-06-24 14:08:17 +00:00
Antoine Lecompte 8f8be89592 fix(staging): oidc unfun (#151) 2026-06-24 14:01:28 +00:00
Antoine Lecompte b03327f7c7 fix(staging): missing quotes (#150) 2026-06-24 13:48:09 +00:00
Antoine Lecompte a92d39f4ec fix(staging): something (#149) 2026-06-24 09:41:40 -04:00
Antoine Lecompte 7d5af4559f feat(staging): wire in oidc (#148) 2026-06-24 09:22:29 -04:00
Antoine Lecompte df65da6f92 fix(staging): a lot (#147) 2026-06-24 13:10:57 +00:00
Antoine Lecompte ff995bc59c fix(staging): typo (#146) 2026-06-24 12:43:06 +00:00
Antoine Lecompte fad10fdae9 feat(ci): add kubeconfig / talosconfig to op (#145) 2026-06-24 12:35:40 +00:00
Antoine Lecompte 02dab5bcc3 fix(ci): fix secret not having write access (#144) 2026-06-24 08:22:32 -04:00
Antoine Lecompte a4e15123be fix(ci): op wrong secret field (#143)
* fix(ci): op wrong secret field
2026-06-24 07:54:39 -04:00
Antoine Lecompte a6eeb8bb67 fix(ci): adjust secret name (#142) 2026-06-24 07:43:23 -04:00
Antoine Lecompte f700b18cd6 feat: staging (#135)
* feat: staging

* more stuff

* pin actions

* adjust

* adjust
2026-06-23 18:21:12 +00:00
Andy Molenda 6463c87a32 chore: update ansible/ceph/inventories/sietch[...]/group_vars/all/var… (#140)
chore: update ansible/ceph/inventories/sietch[...]/group_vars/all/vars.yml

Updated Sietch-Ceph bond_interfaces to match the new 25G network link names:
- eno1 -> eno1np0
- eno2 -> eno2np1
2026-06-23 08:17:20 -07:00
Paul Makles 3e9f844c8e chore: general clean up (#134) 2026-06-22 13:37:21 +00:00
immich-push-o-matic[bot] cc09a64eb6 chore: release main (#133) v0.4.0 2026-06-19 10:48:05 +00:00
renovate[bot] 73463cd42f chore(deps): update dependency restic to v0.19.0 (#104) 2026-06-19 10:34:22 +00:00
Paul Makles 65b585a435 feat: use .well-known/yucca.json to discover backend (#126) 2026-06-18 18:21:09 +01:00
Paul Makles 27d1ac10c2 feat: telemetry & start up error reporting / robustness (#132) 2026-06-18 16:19:11 +00:00
immich-push-o-matic[bot] 359971ab68 chore: release main (#131) v0.3.1 2026-06-17 14:05:03 +00:00
Paul Makles 64b5fa7897 fix: add package metadata for provenance (#130) 2026-06-17 14:03:51 +00:00
immich-push-o-matic[bot]andizzy a6aa7f26c4 chore: release main (#128)
Co-authored-by: izzy <me@insrt.uk>
v0.3.0
2026-06-17 13:47:04 +00:00
Paul Makles 18c5e6e1f0 ci: combine release please components with root (#129) 2026-06-17 13:32:54 +00:00
Paul Makles faa288087d feat: metrics worker (radosgw ingest) (#123) 2026-06-17 14:16:28 +01:00
immich-push-o-matic[bot] db6415c6fa chore: release main (#96) backups-orchestrator-ui-v0.2.0 backups-orchestrator-api-v0.2.0 backups-api-client-v0.2.0 2026-06-17 12:15:33 +00:00
Paul Makles 10c0a3b39c feat: reconfigure primary repository backend (#118) 2026-06-15 13:39:19 +00:00
Paul Makles fcecc8dc20 fix: include python3 in Nix flake, needed for Tilt (#122) 2026-06-15 13:39:16 +00:00
Andy Molenda b36f63f3f8 feat(dns): Cloudflare-managed DNS for futo.cloud (#121)
s3.dev.austin.int.futo.cloud and its virtual-hosted wildcard now resolve
publicly, round-robin across the three Sietch ceph nodes. Records are
declarative in tf/deployment/dev/dns; the Cloudflare token resolves from
1Password via tf/.env. The cluster already expected these names, so
cephadm needed no changes.

See tf/README.md and ansible/ceph/docs/s3-integration.md.
2026-06-12 09:44:52 -07:00
Andy Molenda 30fdf6da88 feat(talos): hyper-converged Talos Kubernetes on the Sietch Ceph hosts (#120)
Run a Talos K8s cluster as libvirt VMs on the existing 3-node Ceph
cluster, using its idle CPU/memory headroom instead of new hardware.
Ansible provisions the hypervisor substrate and VMs; Terraform renders
the inventory and bootstraps the cluster.

See ansible/talos/README.md and docs/runbooks/cluster-bring-up.md.
2026-06-12 13:42:39 +00:00
Antoine Lecompte 070e22a7bb feat: local k8s (#85)
* impl. local kube

* add support for op injected oidc secrets

* ci: set least-privilege workflow token permissions
2026-06-12 13:17:37 +00:00
immich-tofu[bot] 179ae89605 chore: modify CODE_OF_CONDUCT.md 2026-06-03 15:33:58 +00:00
immich-tofu[bot] a67248c41e chore: modify .github/FUNDING.yml 2026-06-02 21:57:52 +00:00
immich-tofu[bot] 1d7f6dbc34 chore: modify SECURITY.md 2026-06-02 21:57:50 +00:00
renovate[bot] b7d2515750 chore(deps): update dependency @sveltejs/kit to v2.57.1 [security] (#52) 2026-06-02 15:49:04 +00:00
renovate[bot] 183ee00c7b chore(deps): update dependency vite to v7.3.2 [security] (#67) 2026-06-02 15:43:35 +00:00
renovate[bot] 35d3cd34ef chore(deps): update dependency @opentelemetry/sdk-node to ^0.217.0 [security] (#88) 2026-06-02 15:29:35 +00:00
renovate[bot] 6aaccb6b93 chore(deps): update dependency svelte to v5.55.7 [security] (#53) 2026-06-02 15:28:43 +00:00
renovate[bot] 95564ee818 chore(deps): pin dependencies (#82) 2026-06-02 16:05:30 +01:00
Paul Makles 6bce03d888 refactor: another general clean-up & tests (#111)
Signed-off-by: izzy <me@insrt.uk>
2026-06-02 15:54:12 +01:00