* docs(ceph): inline ADR rationale and drop the ADR set Fold each linked ADR's rationale into the prose it supported, then remove the ADR files, the README index row, and the stray code-comment reference -- no ADR trace remains. True up the docs to the partition/region/ceph-cluster layout (#222) in the same pass: state keys (yucca/<partition>/<region>/<stack>), inventory paths (<partition>-<region>/<cluster>), stack dirs, and the recover runbook's S3 state paths. Reframe architecture's env section around partitions/regions with sietch as staging/austin. * docs(ceph): editorial pass to align docs with current code and CI/CD Rewrite the ceph docs against the actual code rather than the pre-refactor state: - partition/region/ceph-cluster layout throughout: state keys yucca/<partition>/<region>/<stack>, inventories <partition>-<region>/<cluster>, the real clusters.auto.tfvars schema (partition/region, not environment/datacenter) - sietch reframed as staging/austin with secrets in yucca_tf_staging; vault hierarchy flipped from dev-primary to staging-primary - live CI/CD (.github/workflows/infra.yml): per-partition read/write service accounts delivered as GitHub secrets (OP_TF_YUCCA_<ENV>_ENV[_WRITE]), plan/apply gating, NetBird overlay (Tailscale retired) - correct the CEPH_ENV guidance (export works; deliberately kept out of mise [env]) across README, CONTRIBUTING, scripts.md, adding-a-cluster - drop the obsolete "read SA token from a 1P item" dance from the runbooks - fix stale vaults, paths, examples, and the inventory-provision.ini name * docs(ceph): transliterate docs to plain ASCII Replace non-ASCII punctuation and box-drawing with ASCII equivalents across the ceph docs: em/en dashes to --/-, middot separators to commas, arrows to ->, directory-tree box-drawing to |-- / `--, section sign to "section", x for the multiply glyph, and >= / <= / ~ for the math glyphs. No content changes. * docs(ceph): fix broken rotate-ssh-key link in scripts.md The "Related" link pointed at runbooks/rotate-ssh-key.md, which does not exist (a pre-existing dangling link). SSH-key rotation lives in rotate-secrets.md; point at its "Rotating SSH keys" section. * docs(ceph): style polish from per-doc review Tighten verbal texture flagged by a per-doc style pass; no structural changes. - correctness: ansible-play.sh runs ansible-playbook, it does not exec, so the trap fires from the still-alive wrapper -- fix the "exec"/SIGKILL claims in scripts.md, architecture.md, secrets.md - cut recurring tics: "DR belt"/"belt-and-suspenders" -> "disaster recovery", "a lost laptop is a non-event", "by design", "system mesh" - unstuff long dash/semicolon sentences in architecture (vault-password history, provision/baseline split), secrets (SSH-key paragraph), patterns - recover-bad-tofu-apply: move the dormant-1P-items aside into one note, consolidate the repeated caveats - misc: naming grammar fix + drop trivia, hardware "would"/"blindly", rotate-secrets "after confidence", drop a dead snippet line, fix the 16+ numeric hedge * docs(ceph): make add-node and recover runbooks CI-aware Now that infra.yml applies the stacks and runs the full ceph convergence on merge, refresh the two runbooks the pipeline changed: - add-node: lead with the manual-vs-CI split. The TF + host_vars change is a PR; the only operator-only step is the physical provisioning (live-image boot + provision.yml), which CI can't do; baseline/tune/join/harden run in CI on merge. Keep the by-hand convergence as a documented fallback. - recover-bad-tofu-apply: note that applies now run in CI with the partition write SA, so the bad apply is usually a failed CI run; CI does not self-heal, recovery is operator-run locally. * docs(tf,talos): finish ADR purge into tf, fix sietch vault, mark talos dormant - complete the ADR removal that stopped at ansible/ceph: drop the dangling ADR-009/010 references from tf/README.md (link + related line) and the ceph module / stack code comments, so no ADR trace remains repo-wide - tf/README: the sietch cluster example uses yucca_tf_staging (was yucca_tf_dev) - ansible/talos: add a "second-class, not actively used" status banner to the README and architecture doc so readers don't treat the converged/libvirt Talos docs as live Left untouched: the Tailscale / SA-token / CI sections of tf/README -- those are mid-migration in another owner's lane (bye-tailscale is in flight; fabric still rides Tailscale by design).
4.7 KiB
yucca/ansible/talos
Status: second-class, not actively used. This converged/libvirt Talos substrate is maintained at low priority and is not part of the actively-run deployment. Treat the docs in this subtree as informational -- they may lag the live stacks; validate against current code before relying on any step.
Hyper-converged Talos K8s on the 3-node Sietch Ceph cluster. Those hosts have idle CPU + memory headroom (Ceph spends its budget on disk and network I/O), so the Talos VMs run there instead of on dedicated K8s hardware.
This subtree is the Ansible substrate: it provisions the VLAN 50/51
bridges, the libvirt/KVM stack, and stages the Talos VMs — stopping when
the VMs are running and ready for talosctl. The Terraform half in
tf/deployment/staging/austin/talos/ renders the
Ansible inventory and drives cluster bring-up (machine config, bootstrap,
kubeconfig).
That boundary is deliberate: Ansible stays declarative and idempotent, and the one-shot cluster bootstrap (etcd init, cluster CA) lives in Terraform. Mixing a one-shot bootstrap into idempotent plays is awkward, and keeping cluster secrets out of Ansible's fact cache keeps the blast radius small.
Scope (in)
- VLAN 50 (Compute,
10.50.0.0/16) + VLAN 51 (Services,10.51.0.0/16) L2 bridges on each hypervisor'sbond0. - Idempotent install of qemu-kvm + libvirt-daemon + ovmf + virtinst.
br_netfilterbypass so Talos VM↔VM traffic isn't dropped by Ceph's existinginet filterforward DROP.- Fetch + checksum-verify + stage the official Talos
metal-amd64image. - Define and start Talos VMs per profile.
Scope (out)
talosctl gen config / apply-config / bootstrap— the Terraformsiderolabs/talosprovider owns these (tf/deployment/staging/austin/talos/);docs/operator-handoff.mdkeeps the manual sequence for recovery.- Persistent storage for VMs (RBD-backed boot disks, ceph-csi). Boot disks are local qcow2 today; RBD lands in a follow-up.
- Switch/VLAN tagging — assumed already done upstream (VLAN 50/51 tagged on every hypervisor's bond0 uplink).
- GitOps / Flux / workloads — land after the storage follow-up.
Hardware
| Host | Role | Full VMs |
|---|---|---|
| sietch-ceph-laurel | Ceph + hypervisor | sietch-talos-cp1, worker1 |
| sietch-ceph-lawson | Ceph + hypervisor | sietch-talos-cp2, worker2 |
| sietch-ceph-samara | Ceph + hypervisor | sietch-talos-cp3, worker3 |
Full profile (production, the default everywhere) = 3 CP + 3 workers,
one CP + one worker per hypervisor (each host is its own failure
domain). Smoke profile = cp1 + worker1 on laurel — a single-host
validation tool, selected explicitly with -e profile=smoke.
Operator workflow
Inventory is TF-rendered from tf/deployment/staging/austin/talos/. Run TF first
(from the repo root):
TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:init # first time only
TF_STACK_DIR=tf/deployment/staging/austin/talos mise run tf:apply # renders inventory + host_vars
Then point at the rendered inventory and use the talos-subtree tasks:
cd ansible/talos/
export TALOS_ENV=inventories/staging-austin/inventory.ini
mise run setup # first time only: python venv + ansible collections
mise run lint # yamllint + ansible-lint + shellcheck
mise run check # ansible-playbook --syntax-check all playbooks
mise run preflight # SSH + Ceph health + bond0 + /dev/kvm
mise run prepare-hypervisors # idempotent infra (bridges + libvirt + image)
mise run provision # production full profile (default: 3 CP + 3 workers)
# single-host validation instead: mise run provision -- -e profile=smoke
After provision: the Terraform stack takes over cluster bring-up. See
docs/runbooks/cluster-bring-up.md for the end-to-end flow, and
docs/operator-handoff.md for driving talosctl by hand when recovering.
Safety posture
- All roles refuse to do anything unless their
*_enabledflag is set (defaults false). The prepare-hypervisors playbook flips them. - preflight is
import_playbook'd by prepare + provision — Ceph HEALTH_OK enforcement is automatic. - Network changes are config-only and applied via
networkctl reload, neversystemctl restart systemd-networkd(preserves Ceph storage VLAN). mise run destroy-vmsremoves VMs + overlays only; bridges, libvirt, and the base image stay intact for fast redeploy.
See docs/runbooks/smoke-plan.md for single-host validation (smoke
profile, bounded first pass on a fresh hypervisor, rollback paths).