Files
yucca/ansible/talos/README.md
Andy Molenda 8821fad205 docs(ceph): realign with partition/region model and CI/CD, retire ADRs
* docs(ceph): inline ADR rationale and drop the ADR set

Fold each linked ADR's rationale into the prose it supported, then remove
the ADR files, the README index row, and the stray code-comment reference --
no ADR trace remains.

True up the docs to the partition/region/ceph-cluster layout (#222) in the
same pass: state keys (yucca/<partition>/<region>/<stack>), inventory paths
(<partition>-<region>/<cluster>), stack dirs, and the recover runbook's S3
state paths. Reframe architecture's env section around partitions/regions
with sietch as staging/austin.

* docs(ceph): editorial pass to align docs with current code and CI/CD

Rewrite the ceph docs against the actual code rather than the pre-refactor
state:

- partition/region/ceph-cluster layout throughout: state keys
  yucca/<partition>/<region>/<stack>, inventories <partition>-<region>/<cluster>,
  the real clusters.auto.tfvars schema (partition/region, not environment/datacenter)
- sietch reframed as staging/austin with secrets in yucca_tf_staging; vault
  hierarchy flipped from dev-primary to staging-primary
- live CI/CD (.github/workflows/infra.yml): per-partition read/write service
  accounts delivered as GitHub secrets (OP_TF_YUCCA_<ENV>_ENV[_WRITE]),
  plan/apply gating, NetBird overlay (Tailscale retired)
- correct the CEPH_ENV guidance (export works; deliberately kept out of mise
  [env]) across README, CONTRIBUTING, scripts.md, adding-a-cluster
- drop the obsolete "read SA token from a 1P item" dance from the runbooks
- fix stale vaults, paths, examples, and the inventory-provision.ini name

* docs(ceph): transliterate docs to plain ASCII

Replace non-ASCII punctuation and box-drawing with ASCII equivalents across
the ceph docs: em/en dashes to --/-, middot separators to commas, arrows to
->, directory-tree box-drawing to |-- / `--, section sign to "section", x for
the multiply glyph, and >= / <= / ~ for the math glyphs. No content changes.

* docs(ceph): fix broken rotate-ssh-key link in scripts.md

The "Related" link pointed at runbooks/rotate-ssh-key.md, which does not
exist (a pre-existing dangling link). SSH-key rotation lives in
rotate-secrets.md; point at its "Rotating SSH keys" section.

* docs(ceph): style polish from per-doc review

Tighten verbal texture flagged by a per-doc style pass; no structural changes.

- correctness: ansible-play.sh runs ansible-playbook, it does not exec, so the
  trap fires from the still-alive wrapper -- fix the "exec"/SIGKILL claims in
  scripts.md, architecture.md, secrets.md
- cut recurring tics: "DR belt"/"belt-and-suspenders" -> "disaster recovery",
  "a lost laptop is a non-event", "by design", "system mesh"
- unstuff long dash/semicolon sentences in architecture (vault-password
  history, provision/baseline split), secrets (SSH-key paragraph), patterns
- recover-bad-tofu-apply: move the dormant-1P-items aside into one note,
  consolidate the repeated caveats
- misc: naming grammar fix + drop trivia, hardware "would"/"blindly",
  rotate-secrets "after confidence", drop a dead snippet line, fix the 16+
  numeric hedge

* docs(ceph): make add-node and recover runbooks CI-aware

Now that infra.yml applies the stacks and runs the full ceph convergence on
merge, refresh the two runbooks the pipeline changed:

- add-node: lead with the manual-vs-CI split. The TF + host_vars change is a
  PR; the only operator-only step is the physical provisioning (live-image
  boot + provision.yml), which CI can't do; baseline/tune/join/harden run in
  CI on merge. Keep the by-hand convergence as a documented fallback.
- recover-bad-tofu-apply: note that applies now run in CI with the partition
  write SA, so the bad apply is usually a failed CI run; CI does not self-heal,
  recovery is operator-run locally.

* docs(tf,talos): finish ADR purge into tf, fix sietch vault, mark talos dormant

- complete the ADR removal that stopped at ansible/ceph: drop the dangling
  ADR-009/010 references from tf/README.md (link + related line) and the ceph
  module / stack code comments, so no ADR trace remains repo-wide
- tf/README: the sietch cluster example uses yucca_tf_staging (was yucca_tf_dev)
- ansible/talos: add a "second-class, not actively used" status banner to the
  README and architecture doc so readers don't treat the converged/libvirt
  Talos docs as live

Left untouched: the Tailscale / SA-token / CI sections of tf/README -- those
are mid-migration in another owner's lane (bye-tailscale is in flight; fabric
still rides Tailscale by design).
2026-06-29 13:04:15 -07:00

4.7 KiB

yucca/ansible/talos

Status: second-class, not actively used. This converged/libvirt Talos substrate is maintained at low priority and is not part of the actively-run deployment. Treat the docs in this subtree as informational -- they may lag the live stacks; validate against current code before relying on any step.

Hyper-converged Talos K8s on the 3-node Sietch Ceph cluster. Those hosts have idle CPU + memory headroom (Ceph spends its budget on disk and network I/O), so the Talos VMs run there instead of on dedicated K8s hardware.

This subtree is the Ansible substrate: it provisions the VLAN 50/51 bridges, the libvirt/KVM stack, and stages the Talos VMs — stopping when the VMs are running and ready for talosctl. The Terraform half in tf/deployment/staging/austin/talos/ renders the Ansible inventory and drives cluster bring-up (machine config, bootstrap, kubeconfig).

That boundary is deliberate: Ansible stays declarative and idempotent, and the one-shot cluster bootstrap (etcd init, cluster CA) lives in Terraform. Mixing a one-shot bootstrap into idempotent plays is awkward, and keeping cluster secrets out of Ansible's fact cache keeps the blast radius small.

Scope (in)

  • VLAN 50 (Compute, 10.50.0.0/16) + VLAN 51 (Services, 10.51.0.0/16) L2 bridges on each hypervisor's bond0.
  • Idempotent install of qemu-kvm + libvirt-daemon + ovmf + virtinst.
  • br_netfilter bypass so Talos VM↔VM traffic isn't dropped by Ceph's existing inet filter forward DROP.
  • Fetch + checksum-verify + stage the official Talos metal-amd64 image.
  • Define and start Talos VMs per profile.

Scope (out)

  • talosctl gen config / apply-config / bootstrap — the Terraform siderolabs/talos provider owns these (tf/deployment/staging/austin/talos/); docs/operator-handoff.md keeps the manual sequence for recovery.
  • Persistent storage for VMs (RBD-backed boot disks, ceph-csi). Boot disks are local qcow2 today; RBD lands in a follow-up.
  • Switch/VLAN tagging — assumed already done upstream (VLAN 50/51 tagged on every hypervisor's bond0 uplink).
  • GitOps / Flux / workloads — land after the storage follow-up.

Hardware

Host Role Full VMs
sietch-ceph-laurel Ceph + hypervisor sietch-talos-cp1, worker1
sietch-ceph-lawson Ceph + hypervisor sietch-talos-cp2, worker2
sietch-ceph-samara Ceph + hypervisor sietch-talos-cp3, worker3

Full profile (production, the default everywhere) = 3 CP + 3 workers, one CP + one worker per hypervisor (each host is its own failure domain). Smoke profile = cp1 + worker1 on laurel — a single-host validation tool, selected explicitly with -e profile=smoke.

Operator workflow

Inventory is TF-rendered from tf/deployment/staging/austin/talos/. Run TF first (from the repo root):

TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:init    # first time only
TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply   # renders inventory + host_vars

Then point at the rendered inventory and use the talos-subtree tasks:

cd ansible/talos/
export TALOS_ENV=inventories/staging-austin/inventory.ini

mise run setup       # first time only: python venv + ansible collections
mise run lint        # yamllint + ansible-lint + shellcheck
mise run check       # ansible-playbook --syntax-check all playbooks
mise run preflight   # SSH + Ceph health + bond0 + /dev/kvm
mise run prepare-hypervisors  # idempotent infra (bridges + libvirt + image)
mise run provision   # production full profile (default: 3 CP + 3 workers)
# single-host validation instead: mise run provision -- -e profile=smoke

After provision: the Terraform stack takes over cluster bring-up. See docs/runbooks/cluster-bring-up.md for the end-to-end flow, and docs/operator-handoff.md for driving talosctl by hand when recovering.

Safety posture

  • All roles refuse to do anything unless their *_enabled flag is set (defaults false). The prepare-hypervisors playbook flips them.
  • preflight is import_playbook'd by prepare + provision — Ceph HEALTH_OK enforcement is automatic.
  • Network changes are config-only and applied via networkctl reload, never systemctl restart systemd-networkd (preserves Ceph storage VLAN).
  • mise run destroy-vms removes VMs + overlays only; bridges, libvirt, and the base image stay intact for fast redeploy.

See docs/runbooks/smoke-plan.md for single-host validation (smoke profile, bounded first pass on a fresh hypervisor, rollback paths).