Commit Graph
14 Commits
Author SHA1 Message Date
Antoine Lecompte 963364a034 feat(prod): prod (#247) 2026-07-13 13:15:50 +00:00
Antoine Lecompte 75b7808993 chore(netbird): move to kebab naming (#242)
* chore(netbird): move to kebab naming

Render NetBird object names (groups, setup keys, policies, networks,
network-resources) as lowercase-kebab instead of UPPER_SNAKE, e.g.
YUCCA_PROD_HTZ_FSN1_MGMT → yucca-prod-htz-fsn1-mgmt. The 1Password setup-key
item titles stay UPPER_SNAKE (decoupled) so CI/ansible/talos op:// consumers
keep resolving.

Pin the futo-org/netbird provider to 1.0.2, which fixes the group
resources TF→API decode so a resource-tag group (htz-fsn1 `resources`) can be
renamed in place — no name pin needed.

* update locks
2026-06-30 16:20:10 +00:00
Antoine Lecompte 012db81704 fix(netbird): policies (#239) 2026-06-30 10:31:31 -04:00
Antoine Lecompte 2ec0e668ad chore(netbird): switch provider (#238)
* chore(netbird): change provider, normalize

* commit
2026-06-30 10:09:42 -04:00
Andy Molenda 8821fad205 docs(ceph): realign with partition/region model and CI/CD, retire ADRs
* docs(ceph): inline ADR rationale and drop the ADR set

Fold each linked ADR's rationale into the prose it supported, then remove
the ADR files, the README index row, and the stray code-comment reference --
no ADR trace remains.

True up the docs to the partition/region/ceph-cluster layout (#222) in the
same pass: state keys (yucca/<partition>/<region>/<stack>), inventory paths
(<partition>-<region>/<cluster>), stack dirs, and the recover runbook's S3
state paths. Reframe architecture's env section around partitions/regions
with sietch as staging/austin.

* docs(ceph): editorial pass to align docs with current code and CI/CD

Rewrite the ceph docs against the actual code rather than the pre-refactor
state:

- partition/region/ceph-cluster layout throughout: state keys
  yucca/<partition>/<region>/<stack>, inventories <partition>-<region>/<cluster>,
  the real clusters.auto.tfvars schema (partition/region, not environment/datacenter)
- sietch reframed as staging/austin with secrets in yucca_tf_staging; vault
  hierarchy flipped from dev-primary to staging-primary
- live CI/CD (.github/workflows/infra.yml): per-partition read/write service
  accounts delivered as GitHub secrets (OP_TF_YUCCA_<ENV>_ENV[_WRITE]),
  plan/apply gating, NetBird overlay (Tailscale retired)
- correct the CEPH_ENV guidance (export works; deliberately kept out of mise
  [env]) across README, CONTRIBUTING, scripts.md, adding-a-cluster
- drop the obsolete "read SA token from a 1P item" dance from the runbooks
- fix stale vaults, paths, examples, and the inventory-provision.ini name

* docs(ceph): transliterate docs to plain ASCII

Replace non-ASCII punctuation and box-drawing with ASCII equivalents across
the ceph docs: em/en dashes to --/-, middot separators to commas, arrows to
->, directory-tree box-drawing to |-- / `--, section sign to "section", x for
the multiply glyph, and >= / <= / ~ for the math glyphs. No content changes.

* docs(ceph): fix broken rotate-ssh-key link in scripts.md

The "Related" link pointed at runbooks/rotate-ssh-key.md, which does not
exist (a pre-existing dangling link). SSH-key rotation lives in
rotate-secrets.md; point at its "Rotating SSH keys" section.

* docs(ceph): style polish from per-doc review

Tighten verbal texture flagged by a per-doc style pass; no structural changes.

- correctness: ansible-play.sh runs ansible-playbook, it does not exec, so the
  trap fires from the still-alive wrapper -- fix the "exec"/SIGKILL claims in
  scripts.md, architecture.md, secrets.md
- cut recurring tics: "DR belt"/"belt-and-suspenders" -> "disaster recovery",
  "a lost laptop is a non-event", "by design", "system mesh"
- unstuff long dash/semicolon sentences in architecture (vault-password
  history, provision/baseline split), secrets (SSH-key paragraph), patterns
- recover-bad-tofu-apply: move the dormant-1P-items aside into one note,
  consolidate the repeated caveats
- misc: naming grammar fix + drop trivia, hardware "would"/"blindly",
  rotate-secrets "after confidence", drop a dead snippet line, fix the 16+
  numeric hedge

* docs(ceph): make add-node and recover runbooks CI-aware

Now that infra.yml applies the stacks and runs the full ceph convergence on
merge, refresh the two runbooks the pipeline changed:

- add-node: lead with the manual-vs-CI split. The TF + host_vars change is a
  PR; the only operator-only step is the physical provisioning (live-image
  boot + provision.yml), which CI can't do; baseline/tune/join/harden run in
  CI on merge. Keep the by-hand convergence as a documented fallback.
- recover-bad-tofu-apply: note that applies now run in CI with the partition
  write SA, so the bad apply is usually a failed CI run; CI does not self-heal,
  recovery is operator-run locally.

* docs(tf,talos): finish ADR purge into tf, fix sietch vault, mark talos dormant

- complete the ADR removal that stopped at ansible/ceph: drop the dangling
  ADR-009/010 references from tf/README.md (link + related line) and the ceph
  module / stack code comments, so no ADR trace remains repo-wide
- tf/README: the sietch cluster example uses yucca_tf_staging (was yucca_tf_dev)
- ansible/talos: add a "second-class, not actively used" status banner to the
  README and architecture doc so readers don't treat the converged/libvirt
  Talos docs as live

Left untouched: the Tailscale / SA-token / CI sections of tf/README -- those
are mid-migration in another owner's lane (bye-tailscale is in flight; fabric
still rides Tailscale by design).
2026-06-29 13:04:15 -07:00
Antoine Lecompte c6985d902c feat(all): introduce partition/region/ceph-cluster model across the stack (#222)
* feat: introduce partition/region/ceph-cluster model across the stack

Formalize partition -> region -> {one k8s cluster, many ceph clusters} and
thread it through every layer plus a new yuctl ops CLI.

- tf: deployment/<partition>/<region>/<stack> layout; terragrunt path-parse +
  state key yucca/<partition>/<region>/<stack>; per-region region.hcl (role,
  site_id, datacenter, provider_code, domain); env->partition / site->region
  renames (NetBird object names byte-identical); standardized per-stack
  `discovery` output contract (secrets as op:// refs).
- k8s: clusters/<partition>/<region>/ (staging/austin, prod/htz-fsn1, dev/local);
  role-based kustomize components (primary/secondary); hybrid cluster-settings
  (TF-rendered identity + human fragment); dev-mirror folded into dev/local;
  charts regrouped into charts/{apps,platform,lib,dev}.
- ci: infra.yml partition/region discovery matrix; partition-keyed path filters;
  <partition>-<region> environment gates; image-versions path moves.
- ansible: inventories under <partition>-<region>/<cluster>.
- yuctl: Go/cobra CLI reading the discovery contract from TF state.
- Retire the sietch-talos libvirt VM cluster (dev@local is the k3d cluster);
  ceph inventory_dirname -> <partition>-<region>/<cluster>.

Verified: mise k8s:validate green (3 clusters); yuctl go build/vet; tofu
validate pre-merge (all 9 stacks). Live-staging state migration NOT run.

* fix typo

* commit
2026-06-29 08:40:29 -04:00
Antoine Lecompte 948b15cf0b fix(netbird): some issues (#212) 2026-06-26 18:16:41 +00:00
Antoine Lecompte 031974a32b feat(netbird): wire this sweetie up (#209)
* feat(netbird): wire this sweetie up

* remove dev

* fix

* more features
2026-06-26 18:03:02 +00:00
Andy Molenda 96dfa44c84 docs(ceph): repoint example refs from dev to staging (#187) 2026-06-25 15:06:46 -07:00
Antoine Lecompte a6eeb8bb67 fix(ci): adjust secret name (#142) 2026-06-24 07:43:23 -04:00
Antoine Lecompte f700b18cd6 feat: staging (#135)
* feat: staging

* more stuff

* pin actions

* adjust

* adjust
2026-06-23 18:21:12 +00:00
Andy Molenda b36f63f3f8 feat(dns): Cloudflare-managed DNS for futo.cloud (#121)
s3.dev.austin.int.futo.cloud and its virtual-hosted wildcard now resolve
publicly, round-robin across the three Sietch ceph nodes. Records are
declarative in tf/deployment/dev/dns; the Cloudflare token resolves from
1Password via tf/.env. The cluster already expected these names, so
cephadm needed no changes.

See tf/README.md and ansible/ceph/docs/s3-integration.md.
2026-06-12 09:44:52 -07:00
Andy Molenda 30fdf6da88 feat(talos): hyper-converged Talos Kubernetes on the Sietch Ceph hosts (#120)
Run a Talos K8s cluster as libvirt VMs on the existing 3-node Ceph
cluster, using its idle CPU/memory headroom instead of new hardware.
Ansible provisions the hypervisor substrate and VMs; Terraform renders
the inventory and bootstraps the cluster.

See ansible/talos/README.md and docs/runbooks/cluster-bring-up.md.
2026-06-12 13:42:39 +00:00
Andy Molenda 63087f6850 feat(ceph): import yucca-ceph ansible + terraform infrastructure (#86)
* feat(ceph): import yucca-ceph ansible + terraform infrastructure

Imports the yucca-ceph Ansible tree into ansible/ceph/ and adds the
Terraform stack at tf/ that drives it. Cuts over from ansible-vault
to the hybrid secrets architecture (TF as inventory authority, 1P
as secrets store, op-inject at deploy time) in one atomic move.
Source: internal yucca-ceph working tree; fresh subtree-style
import, history not preserved. Andy continues operating sietch +
painbox post-merge; yucca-team hosts the code and reviews changes.

What it adds:
  - sietch (3-node Austin, production Ceph S3 backend, untouched
    by this PR)
  - painbox (single-node Hetzner SX295 in Helsinki) as a second
    deployable cluster
  - Future clusters land by appending to clusters.auto.tfvars in
    the matching environment stack (tf/deployment/<env>/ceph/) —
    no per-cluster TF code required

How it works (full map: ansible/ceph/docs/architecture.md):
  - tf/shared/modules/ceph-cluster renders inventory.ini variants
    + secrets.yml.tpl per cluster from clusters.auto.tfvars
  - secrets.yml.tpl carries op:// refs; `op inject -f` resolves
    them at deploy time from the matching yucca_tf_<env> vault
  - State in OVH yucca-tf-state bucket (key ceph/<env>/<stack>/)
  - 11 ADRs capture the decisions: ansible/ceph/docs/adr/

Out of scope (intentional):
  - LUKS keys not yet in 1P (deferred until hybrid is stable)
  - tf/shared/modules/ceph-cluster/secrets.tf.disabled is dormant;
    today's 1P items via `op item create` per
    ansible/ceph/docs/adding-a-cluster.md
  - Talos K8s on sietch is a separate workstream

Atomicity + rollback: TF-rendered inventory + secrets-template
files are gitignored (TF generates them) and ansible-vault removal
is coupled to the op-inject path. Splitting this PR lands in a
non-bootable state — merge as one unit. The merge itself is
reversible via `git revert` until the post-merge `tf:apply` runs;
after apply, full rollback needs state restore or `tofu state mv`
(land + validate before applying).

Dev-env impact: adds opentofu + terragrunt to yucca root mise tools
plus a self-contained ansible/ceph/.mise.toml. No new commands or
prereqs for immich-side contributors who don't touch ceph or run
tf:* tasks.

Verification:
  - `mise run lint` (from ansible/ceph/): 130 files, 0 warnings
  - `mise run check`: 19 playbooks parse clean
  - `mise run tf:plan`: succeeds; 7 expected file path-rename
    replacements (3 painbox + 4 sietch). State drift from import,
    no cluster-side change.
  - painbox deployed 2026-04-26 on the new code path: Bookworm +
    Ceph Tentacle, 15 OSDs (14 HDD + 1 SSD) up + in, mon/mgr/rgw
    running. HEALTH_WARN is expected on a single-node cluster.

Post-merge: from the yucca root, `mise run tf:apply` flips the
bucket state to the new monorepo paths (the 7 renames above).

* fix(ceph): exempt ansible/ and tf/ subtrees from root prettier

The imported infrastructure subtrees enforce their own format
conventions (yamllint + ansible-lint inside ansible/ceph/; tofu fmt
inside tf/). Prettier on ansible YAML reflows long Jinja2 expressions
and shell command blocks in unwanted ways, so root prettier checks
are skipped for both subtrees.

Also reformat root README.md table column alignment to match prettier
conventions (only the imported subtrees are exempt; yucca-side files
including the root README still follow root prettier rules).

* fix(ceph): clean up secrets tmpfile after ansible-playbook exits

`ansible-play.sh` rendered the resolved secrets file via `op inject`
into a `mktemp` tmpfile, set up a `trap 'rm -f "$TMPFILE"' EXIT INT
TERM`, then `exec`'d ansible-playbook. The `exec` replaced the bash
shell entirely, so the EXIT trap never fired — every play left a
plaintext-secrets file in /tmp.

In practice this was masked because /tmp is tmpfs (RAM only on this
operator's setup), so files evaporate on reboot. But within an
operator session, files accumulated linearly with each playbook
invocation. Recent count on the import-PR session: 38 files.

Drop the `exec`. With `set -euo pipefail` already on, bash:

  - propagates ansible-playbook's exit code (set -e)
  - fires the EXIT trap before exiting (always)
  - cleans up the tmpfile on success, failure, or signal

Verified: `CEPH_ENV=... scripts/ansible-play.sh status.yml
--syntax-check` creates and removes the tmpfile within the same
invocation — /tmp is clean before and after.

`scripts/preflight.sh` uses the same trap pattern but does not
`exec`, so its tmpfile cleanup was already correct (and the suffix
differs: `-secrets-test.yml` vs `-secrets.yml`, confirming
ansible-play.sh as the sole offender).
2026-05-18 06:17:56 -07:00