Files
yucca/ansible/ceph/README.md
T
Andy Molenda 8821fad205 docs(ceph): realign with partition/region model and CI/CD, retire ADRs
* docs(ceph): inline ADR rationale and drop the ADR set

Fold each linked ADR's rationale into the prose it supported, then remove
the ADR files, the README index row, and the stray code-comment reference --
no ADR trace remains.

True up the docs to the partition/region/ceph-cluster layout (#222) in the
same pass: state keys (yucca/<partition>/<region>/<stack>), inventory paths
(<partition>-<region>/<cluster>), stack dirs, and the recover runbook's S3
state paths. Reframe architecture's env section around partitions/regions
with sietch as staging/austin.

* docs(ceph): editorial pass to align docs with current code and CI/CD

Rewrite the ceph docs against the actual code rather than the pre-refactor
state:

- partition/region/ceph-cluster layout throughout: state keys
  yucca/<partition>/<region>/<stack>, inventories <partition>-<region>/<cluster>,
  the real clusters.auto.tfvars schema (partition/region, not environment/datacenter)
- sietch reframed as staging/austin with secrets in yucca_tf_staging; vault
  hierarchy flipped from dev-primary to staging-primary
- live CI/CD (.github/workflows/infra.yml): per-partition read/write service
  accounts delivered as GitHub secrets (OP_TF_YUCCA_<ENV>_ENV[_WRITE]),
  plan/apply gating, NetBird overlay (Tailscale retired)
- correct the CEPH_ENV guidance (export works; deliberately kept out of mise
  [env]) across README, CONTRIBUTING, scripts.md, adding-a-cluster
- drop the obsolete "read SA token from a 1P item" dance from the runbooks
- fix stale vaults, paths, examples, and the inventory-provision.ini name

* docs(ceph): transliterate docs to plain ASCII

Replace non-ASCII punctuation and box-drawing with ASCII equivalents across
the ceph docs: em/en dashes to --/-, middot separators to commas, arrows to
->, directory-tree box-drawing to |-- / `--, section sign to "section", x for
the multiply glyph, and >= / <= / ~ for the math glyphs. No content changes.

* docs(ceph): fix broken rotate-ssh-key link in scripts.md

The "Related" link pointed at runbooks/rotate-ssh-key.md, which does not
exist (a pre-existing dangling link). SSH-key rotation lives in
rotate-secrets.md; point at its "Rotating SSH keys" section.

* docs(ceph): style polish from per-doc review

Tighten verbal texture flagged by a per-doc style pass; no structural changes.

- correctness: ansible-play.sh runs ansible-playbook, it does not exec, so the
  trap fires from the still-alive wrapper -- fix the "exec"/SIGKILL claims in
  scripts.md, architecture.md, secrets.md
- cut recurring tics: "DR belt"/"belt-and-suspenders" -> "disaster recovery",
  "a lost laptop is a non-event", "by design", "system mesh"
- unstuff long dash/semicolon sentences in architecture (vault-password
  history, provision/baseline split), secrets (SSH-key paragraph), patterns
- recover-bad-tofu-apply: move the dormant-1P-items aside into one note,
  consolidate the repeated caveats
- misc: naming grammar fix + drop trivia, hardware "would"/"blindly",
  rotate-secrets "after confidence", drop a dead snippet line, fix the 16+
  numeric hedge

* docs(ceph): make add-node and recover runbooks CI-aware

Now that infra.yml applies the stacks and runs the full ceph convergence on
merge, refresh the two runbooks the pipeline changed:

- add-node: lead with the manual-vs-CI split. The TF + host_vars change is a
  PR; the only operator-only step is the physical provisioning (live-image
  boot + provision.yml), which CI can't do; baseline/tune/join/harden run in
  CI on merge. Keep the by-hand convergence as a documented fallback.
- recover-bad-tofu-apply: note that applies now run in CI with the partition
  write SA, so the bad apply is usually a failed CI run; CI does not self-heal,
  recovery is operator-run locally.

* docs(tf,talos): finish ADR purge into tf, fix sietch vault, mark talos dormant

- complete the ADR removal that stopped at ansible/ceph: drop the dangling
  ADR-009/010 references from tf/README.md (link + related line) and the ceph
  module / stack code comments, so no ADR trace remains repo-wide
- tf/README: the sietch cluster example uses yucca_tf_staging (was yucca_tf_dev)
- ansible/talos: add a "second-class, not actively used" status banner to the
  README and architecture doc so readers don't treat the converged/libvirt
  Talos docs as live

Left untouched: the Tailscale / SA-token / CI sections of tf/README -- those
are mid-migration in another owner's lane (bye-tailscale is in flight; fabric
still rides Tailscale by design).
2026-06-29 13:04:15 -07:00

7.1 KiB

ceph -- Ceph Tentacle on Bare Metal

Ansible automation for provisioning, deploying, tuning, hardening, and operating Ceph Tentacle (v20) clusters on bare-metal hardware via cephadm. Lives in the yucca monorepo at ansible/ceph/; secrets and inventory scaffolding are provisioned from yucca/tf/ (see ../../tf/).

Cluster Domain Location Hardware Nodes
sietch staging.austin.int.futo.cloud Austin DC Dell R730xd 3

Clusters are declared in yucca/tf/deployment/staging/austin/ceph/clusters.auto.tfvars; tofu apply renders inventories/<partition>-<region>/<cluster>/inventory.ini and secrets.yml.tpl per cluster. The CEPH_ENV variable selects the active cluster for any mise run or direct ansible invocation.

Architecture

graph TB
    subgraph "Controller (your workstation)"
        A[mise + ansible + 1Password CLI]
    end

    subgraph "Austin DC -- 10.10.10.0/24"
        direction TB
        L[laurel<br/>MON+MGR+OSD+RGW]
        W[lawson<br/>MON+MGR+OSD+RGW]
        S[samara<br/>MON+MGR+OSD+RGW]
    end

    A -->|SSH| L
    A -->|SSH| W
    A -->|SSH| S

See docs/architecture.md for role dependencies, data flow, and design rationale.

Quick Start

# 1. Render cluster inventories + secrets templates (once, from yucca/tf/)
(cd ../../tf/deployment/staging/austin/ceph && tofu init && tofu apply)

# 2. Set up the ansible side
mise trust && mise run setup          # bootstrap dev environment

# 3. Run mise tasks against the target cluster. Export CEPH_ENV once per
#    shell (or prefix it inline). It is deliberately NOT in mise's [env]
#    block -- that would override your shell value and silently target the
#    wrong cluster; see docs/scripts.md "Setting CEPH_ENV".
export CEPH_ENV=inventories/staging-austin/sietch/inventory.ini
mise run preflight       # TF artifacts + 1P + SSH + connectivity
mise run status          # read-only cluster health check
mise run drift           # configuration drift detection
mise run deploy          # full pipeline (idempotent)

See CONTRIBUTING.md for the full development workflow.

Roles

Order Role Description
1 provision_host Bare-metal Debian 12 install via debootstrap (Austin only)
2 baseline Post-boot OS baseline: ops user, packages, /etc/hosts, services
3 os_tuning Kernel sysctl, TCP buffers, optional centralized logging
4 hardware_tuning I/O scheduler, readahead, udev rules, optional CPU governor
5 ceph_deploy cephadm bootstrap, join, placement, OSDs, RGW, monitoring
6 ceph_tuning Recovery throttling, scrub window, CRUSH, telemetry, audit
7 security nftables firewall, SSH hardening
8 ceph_destroy Complete cluster teardown (safety-gated)
9 s3_bench Parallel S3 benchmark against local RGW

Playbooks

Playbook Description
site.yml Full pipeline: baseline + tune + deploy + tune + harden
provision.yml Bare-metal provisioning (Austin, -i provision inventory)
baseline.yml OS baseline (users, packages, hosts)
tune-os.yml Kernel/sysctl tuning
tune-hardware.yml Disk I/O tuning
deploy-ceph.yml Ceph cluster deployment (tags: prerequisites, bootstrap, join, placement, lvm, osds, crush, rgw, monitoring, verify)
tune-ceph.yml Post-deploy Ceph config tuning
harden.yml Firewall + SSH hardening
destroy-ceph.yml Cluster teardown (destroy inventory, requires flags)
status.yml Quick health check (read-only)
drift.yml Configuration drift detection
bench.yml S3 benchmark (RGW round-trip)
rados-bench.yml RADOS bench (raw cluster I/O, bypasses RGW)
backup-config.yml Export cluster config for DR
post-deploy-capture.yml Snapshot RGW TLS + admin keyring to 1P for disaster recovery
rotate-certs.yml RGW TLS certificate rotation
rotate-ssh-key.yml Distribute current ansible-iac pubkey from 1P to nodes
migrate-networkd.yml One-shot networkd/bridge migration (rolling, noout-gated)
hardware-inventory.yml Hardware facts to JSON

mise Tasks

Task Description
setup Bootstrap dev environment (venv, deps, collections)
lint yamllint + ansible-lint + shellcheck (no 1P required)
check Syntax-check all playbooks (no 1P required)
test Molecule tests
preflight TF artifacts + 1P session + SSH + connectivity
status Cluster health check
drift Configuration drift detection
deploy Full pipeline
destroy Cluster teardown (interactive)
backup Export cluster config for DR
capture Snapshot RGW TLS + admin keyring to 1P for disaster recovery
bench S3 benchmark (RGW round-trip)
bench-rados RADOS bench (raw cluster I/O)
rotate-certs RGW TLS certificate rotation
rotate-ssh-key Distribute current ansible-iac pubkey from 1P to nodes
hardware-inventory Capture hardware facts to JSON
migrate-networkd One-shot networkd/bridge migration (rolling, noout-gated)

Inventory + secrets-template scaffolding live in yucca/tf/ -- run tofu apply in tf/deployment/staging/austin/ceph/ to (re-)render inventories/<partition>-<region>/<cluster>/inventory.ini and secrets.yml.tpl.

Documentation

Document Audience
CONTRIBUTING.md Developers -- setup, workflow, conventions
docs/architecture.md Developers -- role graph, data flow, design
docs/patterns.md Developers -- coding idioms, anti-patterns
docs/adding-a-role.md Developers -- role skeleton, conventions
docs/adding-a-cluster.md Developers -- inventory setup, secrets
docs/secrets.md Developers/ops -- 1Password integration
docs/naming.md Everyone -- hostname and inventory naming
docs/hardware.md Ops/procurement -- node specs and hardware shapes
docs/s3-integration.md App developers -- endpoints, boto3, certs
docs/security-model.md InfoSec -- encryption, users, firewall
docs/capacity-planning.md Managers -- costs, formulas, growth
docs/troubleshooting.md SRE/on-call -- symptom/diagnosis/fix
docs/runbooks/ Ops -- add/replace node, replace disk, rotate certs/secrets/SSH/SA token, remote hands, bad-tofu-apply recovery, backup/restore

Known Limitations

  • Single-network topology: public = cluster network on sietch.
  • Self-signed TLS: RGW clients need --no-verify-ssl. Production needs real certs.
  • ops user is password-only: no SSH keys installed; password sourced from 1P. Intended as an interactive console or recovery account, not for automation.
  • DNS not managed: s3.<domain> and *.s3.<domain> records must exist externally.