mirror of
https://github.com/immich-app/yucca.git
synced 2026-09-30 21:37:50 +08:00
* docs(ceph): inline ADR rationale and drop the ADR set Fold each linked ADR's rationale into the prose it supported, then remove the ADR files, the README index row, and the stray code-comment reference -- no ADR trace remains. True up the docs to the partition/region/ceph-cluster layout (#222) in the same pass: state keys (yucca/<partition>/<region>/<stack>), inventory paths (<partition>-<region>/<cluster>), stack dirs, and the recover runbook's S3 state paths. Reframe architecture's env section around partitions/regions with sietch as staging/austin. * docs(ceph): editorial pass to align docs with current code and CI/CD Rewrite the ceph docs against the actual code rather than the pre-refactor state: - partition/region/ceph-cluster layout throughout: state keys yucca/<partition>/<region>/<stack>, inventories <partition>-<region>/<cluster>, the real clusters.auto.tfvars schema (partition/region, not environment/datacenter) - sietch reframed as staging/austin with secrets in yucca_tf_staging; vault hierarchy flipped from dev-primary to staging-primary - live CI/CD (.github/workflows/infra.yml): per-partition read/write service accounts delivered as GitHub secrets (OP_TF_YUCCA_<ENV>_ENV[_WRITE]), plan/apply gating, NetBird overlay (Tailscale retired) - correct the CEPH_ENV guidance (export works; deliberately kept out of mise [env]) across README, CONTRIBUTING, scripts.md, adding-a-cluster - drop the obsolete "read SA token from a 1P item" dance from the runbooks - fix stale vaults, paths, examples, and the inventory-provision.ini name * docs(ceph): transliterate docs to plain ASCII Replace non-ASCII punctuation and box-drawing with ASCII equivalents across the ceph docs: em/en dashes to --/-, middot separators to commas, arrows to ->, directory-tree box-drawing to |-- / `--, section sign to "section", x for the multiply glyph, and >= / <= / ~ for the math glyphs. No content changes. * docs(ceph): fix broken rotate-ssh-key link in scripts.md The "Related" link pointed at runbooks/rotate-ssh-key.md, which does not exist (a pre-existing dangling link). SSH-key rotation lives in rotate-secrets.md; point at its "Rotating SSH keys" section. * docs(ceph): style polish from per-doc review Tighten verbal texture flagged by a per-doc style pass; no structural changes. - correctness: ansible-play.sh runs ansible-playbook, it does not exec, so the trap fires from the still-alive wrapper -- fix the "exec"/SIGKILL claims in scripts.md, architecture.md, secrets.md - cut recurring tics: "DR belt"/"belt-and-suspenders" -> "disaster recovery", "a lost laptop is a non-event", "by design", "system mesh" - unstuff long dash/semicolon sentences in architecture (vault-password history, provision/baseline split), secrets (SSH-key paragraph), patterns - recover-bad-tofu-apply: move the dormant-1P-items aside into one note, consolidate the repeated caveats - misc: naming grammar fix + drop trivia, hardware "would"/"blindly", rotate-secrets "after confidence", drop a dead snippet line, fix the 16+ numeric hedge * docs(ceph): make add-node and recover runbooks CI-aware Now that infra.yml applies the stacks and runs the full ceph convergence on merge, refresh the two runbooks the pipeline changed: - add-node: lead with the manual-vs-CI split. The TF + host_vars change is a PR; the only operator-only step is the physical provisioning (live-image boot + provision.yml), which CI can't do; baseline/tune/join/harden run in CI on merge. Keep the by-hand convergence as a documented fallback. - recover-bad-tofu-apply: note that applies now run in CI with the partition write SA, so the bad apply is usually a failed CI run; CI does not self-heal, recovery is operator-run locally. * docs(tf,talos): finish ADR purge into tf, fix sietch vault, mark talos dormant - complete the ADR removal that stopped at ansible/ceph: drop the dangling ADR-009/010 references from tf/README.md (link + related line) and the ceph module / stack code comments, so no ADR trace remains repo-wide - tf/README: the sietch cluster example uses yucca_tf_staging (was yucca_tf_dev) - ansible/talos: add a "second-class, not actively used" status banner to the README and architecture doc so readers don't treat the converged/libvirt Talos docs as live Left untouched: the Tailscale / SA-token / CI sections of tf/README -- those are mid-migration in another owner's lane (bye-tailscale is in flight; fabric still rides Tailscale by design).
875 lines
40 KiB
Markdown
875 lines
40 KiB
Markdown
# Architecture
|
|
|
|
How the Ceph automation in `yucca/ansible/ceph/` is shaped, what each tool
|
|
owns, and how the four tools (Terraform, 1Password, Ansible, mise) -- plus the
|
|
op CLI that resolves secrets between them -- hand off work to each other.
|
|
|
|
This is a structural reference. For step-by-step usage see
|
|
[CONTRIBUTING.md](../CONTRIBUTING.md); for narrower topics see the
|
|
specialized docs under `docs/`.
|
|
|
|
---
|
|
|
|
## 1. System context
|
|
|
|
```mermaid
|
|
flowchart LR
|
|
OP([Operator workstation<br/>mise, op CLI, ansible, tofu])
|
|
YUCCA[/Yucca monorepo<br/>tf/ + ansible/ceph/ + kubernetes//]
|
|
ONEP[("1Password org<br/>yucca_tf, yucca_tf_staging, ...")]
|
|
S3[("OVH S3<br/>yucca-tf-state bucket")]
|
|
SIETCH["Sietch, Austin DC<br/>3x Dell R730xd"]
|
|
|
|
OP -->|edits| YUCCA
|
|
OP -->|reads/writes secrets| ONEP
|
|
OP -->|TF state I/O| S3
|
|
OP -->|SSH ansible-iac| SIETCH
|
|
```
|
|
|
|
The yucca monorepo is the single source of truth for cluster identity and
|
|
configuration. Operators run mise tasks on their workstation; secrets stay
|
|
in 1Password (never on disk); TF state lives in OVH S3; Ansible drives
|
|
configuration over SSH against bare-metal Ceph nodes.
|
|
|
|
External dependencies are minimal and explicit:
|
|
|
|
- **1Password org** -- organization-scoped vaults shared with other Futo infra
|
|
(Immich, o11y). Authoritative store for live secret values.
|
|
- **OVH S3** -- `yucca-tf-state` bucket at `s3.eu-west-par.io.cloud.ovh.net`.
|
|
Keyed by `yucca/<partition>/<region>/<stack>/terraform.tfstate` so multiple
|
|
stacks share the bucket without collision.
|
|
- **Hardware** -- Austin colo for sietch (Dell R730xd x 3, single 10G bond).
|
|
Detail in [hardware.md](hardware.md).
|
|
|
|
---
|
|
|
|
## 2. Partitions and regions
|
|
|
|
Partition (dev / staging / prod) and region are first-class concerns: every
|
|
tool in the mesh derives both from the same source -- directory layout
|
|
(`tf/deployment/<partition>/<region>/<stack>`) -- so isolation is structural,
|
|
not flag-driven.
|
|
|
|
| Layer | staging / austin (today) | dev / local (planned) | prod / htz-fsn1 (planned) |
|
|
|--------------|---------------------------------------------------------|------------------------------------------------|------------------------------------------------|
|
|
| TF stack dir | `tf/deployment/staging/austin/ceph/` | `tf/deployment/dev/local/ceph/` | `tf/deployment/prod/htz-fsn1/ceph/` |
|
|
| TF state key | `yucca/staging/austin/ceph/terraform.tfstate` | `yucca/dev/local/ceph/terraform.tfstate` | `yucca/prod/htz-fsn1/ceph/terraform.tfstate` |
|
|
| 1P vaults | `yucca_tf_staging`, `yucca_tf_staging_manual` | `yucca_tf_dev`, `yucca_tf_dev_manual` | `yucca_tf` (live), `yucca_tf_prod_manual` |
|
|
| Ansible inv | `inventories/staging-austin/<cluster>/` | `inventories/dev-local/<cluster>/` | `inventories/prod-htz-fsn1/<cluster>/` |
|
|
| mise default | `CEPH_ENV=...staging-austin/sietch/inventory.ini` | overridden via env at invocation | overridden via env at invocation |
|
|
|
|
Today the only deployed cluster is sietch (staging / austin). Adding another
|
|
region or partition is purely additive: create the matching
|
|
`tf/deployment/<partition>/<region>/ceph/` directory, populate
|
|
`clusters.auto.tfvars`, and the same module + Ansible roles + mise tasks work
|
|
unchanged. The state backend key path, 1P vault selection, and inventory
|
|
directory naming all derive from the partition + region segments.
|
|
|
|
`TF_STACK_DIR` is the operator-side override for `mise run tf:*` tasks; it
|
|
defaults to `tf/deployment/staging/austin/ceph` and points at any sibling stack directory.
|
|
`CEPH_ENV` is the matching override for Ansible -- points at the rendered
|
|
`inventory.ini` for the cluster you intend to operate on.
|
|
|
|
---
|
|
|
|
## 3. The tool mesh
|
|
|
|
```mermaid
|
|
flowchart TB
|
|
subgraph ws["Operator workstation"]
|
|
direction LR
|
|
MISE([mise<br/>orchestration])
|
|
TF[Terraform / Tofu<br/>via Terragrunt]
|
|
OP[op CLI]
|
|
ANS[Ansible]
|
|
WRAP[scripts/<br/>ansible-play.sh<br/>install-ssh-keys.sh]
|
|
end
|
|
|
|
ONEP[("1Password<br/>yucca_tf_*")]
|
|
S3[("OVH S3<br/>tfstate")]
|
|
REPO[/"Yucca repo<br/>inventories/<partition>-<region>/<cluster>/<br/>(host_vars committed,<br/>TF outputs gitignored)"/]
|
|
NODES[Ceph nodes]
|
|
|
|
MISE -->|tf:*| TF
|
|
MISE -->|deploy / status / drift| WRAP
|
|
MISE -->|capture| WRAP
|
|
|
|
TF -->|reads SA token via op run --env-file| OP
|
|
TF -->|reads/writes state| S3
|
|
TF -->|renders| REPO
|
|
|
|
WRAP -->|reads inventory + secrets.yml.tpl| REPO
|
|
WRAP -->|op inject / op read| OP
|
|
WRAP -->|runs| ANS
|
|
ANS -->|SSH ansible-iac| NODES
|
|
|
|
OP <-->|item CRUD| ONEP
|
|
```
|
|
|
|
### Who owns what
|
|
|
|
| Tool | Owns | Reads from |
|
|
|---------------|--------------------------------------------------------------------------------|---------------------------------------------|
|
|
| **Terraform** | Cluster identity, host names, rendered Ansible artifacts, TF state | `clusters.auto.tfvars`, 1P (via op CLI) |
|
|
| **1Password** | Live secret values, SSH keypairs, service-account tokens | nothing -- authoritative store |
|
|
| **op CLI** | Auth and resolution: env-injection, file-template injection, single-value read | 1P (session or SA token) |
|
|
| **Ansible** | Convergence: applying configuration to nodes | Rendered inventory + op-injected tmpfile |
|
|
| **mise** | Task discovery, toolchain pinning, env defaults | `.mise.toml`, `tf/.env` |
|
|
|
|
### Handoff points (the edges of the mesh)
|
|
|
|
1. **TF -> repo** -- `terragrunt apply` renders `inventory.ini`,
|
|
`inventory-destroy.ini`, optional `inventory-provision.ini`,
|
|
and `secrets.yml.tpl` into `inventories/<partition>-<region>/<cluster>/`. These files are
|
|
gitignored -- the source of truth is `clusters.auto.tfvars` + the
|
|
ceph-cluster module.
|
|
2. **TF <-> op CLI** -- TF runs are wrapped with `op run --env-file=tf/.env`,
|
|
which resolves `op://...` references in `tf/.env` and injects them as
|
|
`OP_SERVICE_ACCOUNT_TOKEN`, `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`
|
|
for the child process. The `tf/.env` file is committed (it contains only
|
|
pointers, never literal secrets).
|
|
3. **Ansible <-> op CLI** -- `scripts/ansible-play.sh` reads the cluster's
|
|
`secrets.yml.tpl`, runs `op inject -f` to resolve `op://` references into
|
|
a `mktemp`'d tmpfile (chmod 600, trap-cleaned), then runs
|
|
`ansible-playbook --extra-vars @<tmpfile>`. The tmpfile lives only for
|
|
the duration of the play.
|
|
4. **mise -> wrappers** -- `mise run deploy` invokes
|
|
`scripts/ansible-play.sh deploy.yml ...`; `mise run tf:*` invokes
|
|
`tf/op-run.sh terragrunt --working-dir <stack> <cmd>` (op-run.sh is a thin
|
|
`op run --env-file=tf/.env --` wrapper). mise tasks never call
|
|
`ansible-playbook` directly.
|
|
|
|
---
|
|
|
|
## 4. Terraform (authority)
|
|
|
|
### Layout
|
|
|
|
```
|
|
tf/
|
|
|-- .env op:// references (committed; no literal secrets)
|
|
|-- shared/modules/ceph-cluster/ per-cluster orchestration module
|
|
| |-- main.tf, variables.tf, outputs.tf, rendering.tf
|
|
| |-- wordlist.txt 923 words for auto-picked hostnames
|
|
| `-- templates/ inventory + secrets.yml.tpl templates
|
|
`-- deployment/
|
|
|-- terragrunt.hcl root: state backend, partition/region/stack derived from path
|
|
`-- staging/austin/ceph/
|
|
|-- terragrunt.hcl includes root, sets ansible_project_root
|
|
|-- main.tf, variables.tf, versions.tf
|
|
`-- clusters.auto.tfvars declarative cluster list
|
|
```
|
|
|
|
### Cluster identity is declared, not derived
|
|
|
|
The `clusters` map in `clusters.auto.tfvars` is the source of truth.
|
|
Each top-level key becomes a cluster:
|
|
|
|
```hcl
|
|
sietch = {
|
|
domain = "staging.austin.int.futo.cloud"
|
|
partition = "staging"
|
|
region = "austin"
|
|
provider_code = "int"
|
|
role_in_hostname = "ceph"
|
|
ansible_ssh_user = "ansible-iac"
|
|
ansible_ssh_key = "~/.ssh/id_ed25519_sietch"
|
|
vault = "yucca_tf_staging"
|
|
provision_profile = "debian-live"
|
|
hosts = [
|
|
{ name = "laurel", bond_ip = "10.10.10.90", bootstrap = true },
|
|
{ name = "lawson", bond_ip = "10.10.10.91" },
|
|
{ name = "samara", bond_ip = "10.10.10.92" },
|
|
]
|
|
}
|
|
```
|
|
|
|
The module computes everything else: hostname (`<cluster>-<role>-<name>`),
|
|
FQDN (`<hostname>.<domain>`), 1P item names
|
|
(`<CLUSTER>_CEPH_<ROLE>_PASSWORD`), inventory directory path
|
|
(`inventories/<partition>-<region>/<cluster>/`).
|
|
|
|
### Auto-naming via wordlist
|
|
|
|
For hosts where `name = null`, the module picks a stable name from a
|
|
923-word pool using `random_shuffle` seeded by `(cluster_name, name_seed)`.
|
|
Operator-declared names are excluded from the pool to prevent collisions
|
|
within a cluster. Adding hosts at the tail is safe -- existing positions
|
|
keep their names across applies.
|
|
|
|
A host declared with no `name` demonstrates this: TF auto-picks a stable
|
|
word (e.g. `evelyn`) -> hostname `<cluster>-ceph-evelyn`.
|
|
|
|
### Rendered artifacts (gitignored)
|
|
|
|
Per `tf/shared/modules/ceph-cluster/rendering.tf`, the module writes four
|
|
files into `inventories/<partition>-<region>/<cluster>/`:
|
|
|
|
| File | Purpose |
|
|
|--------------------------------------------|------------------------------------------------------------------------------------------|
|
|
| `inventory.ini` | Normal-ops inventory: `ansible-iac` user + cluster SSH key |
|
|
| `inventory-destroy.ini` | Destroy-mode inventory (same credentials; separate file as a speed bump) |
|
|
| `inventory-provision.ini` | Provisioning inventory (only when `provision_profile != null`; uses live-image creds; the profile names the template, not the output) |
|
|
| `secrets.yml.tpl` | `vault_*: op://<vault>/<CLUSTER>_CEPH_*/password` pointers, consumed by `op inject -f` |
|
|
|
|
All four are in `ansible/ceph/.gitignore`. Re-render with `mise run tf:apply`.
|
|
|
|
### State backend
|
|
|
|
S3 backend in `tf/deployment/terragrunt.hcl`:
|
|
|
|
- Bucket: `yucca-tf-state` (shared with o11y and other Futo stacks)
|
|
- Region: `eu-west-par` (OVH Paris)
|
|
- Endpoint: `https://s3.eu-west-par.io.cloud.ovh.net/`
|
|
- Key: `yucca/${partition}/${region}/${stack}/terraform.tfstate` -- derived
|
|
from the child stack's path under `deployment/`
|
|
- Skip AWS-specific validation; use path-style URLs (OVH compatibility)
|
|
|
|
State locking is **not enabled today**. OVH has no DynamoDB equivalent.
|
|
OpenTofu's `use_lockfile = true` would work but expects the lockfile object
|
|
to already exist -- fresh-backend init fails with 404 before it can create
|
|
one. Single-operator workflow today; revisit when concurrent applies become
|
|
likely. See `deployment/terragrunt.hcl` for the inline rationale.
|
|
|
|
### What TF does not yet manage
|
|
|
|
`onepassword_item` resources are **dormant** (`tf/.../secrets.tf.disabled`).
|
|
1P items are created today via the `op` CLI (operator runs `op item create`
|
|
once per cluster). The gate to re-enabling them is the dedicated
|
|
`sietch-ceph` service account that lets us split write authority from the
|
|
org-wide superuser SA.
|
|
|
|
---
|
|
|
|
## 5. 1Password (live values)
|
|
|
|
### Vault hierarchy
|
|
|
|
| Vault | Purpose | Who reads it | Who writes it |
|
|
|---------------------------|--------------------------------------------------------|------------------------------------|-------------------------------------|
|
|
| `yucca_tf` | Cross-partition shared (TF state S3 creds) | TF (via `tf/op-run.sh`) | Operator (manual) |
|
|
| `yucca_tf_staging` | staging live values (sietch today) | Ansible runtime (op inject) | Superuser SA (TF) + operator (op CLI) |
|
|
| `yucca_tf_staging_manual` | staging human-fillable placeholders (API tokens, OAuth)| Ansible runtime | Operator (manual) |
|
|
| `yucca_tf_dev(_manual)`, `yucca_tf`, `yucca_tf_prod_manual` | dev + prod analogues, same shape | per partition | per partition |
|
|
|
|
Each partition has its own live + `_manual` vault pair; sietch runs in
|
|
staging, so its items live in `yucca_tf_staging`. The `_manual` vaults exist
|
|
for items that can't be auto-generated (third-party API tokens, OAuth client
|
|
secrets) -- they're populated by humans, not by TF.
|
|
|
|
### Service accounts
|
|
|
|
Each partition has a **read** and a **write** 1Password service account, scoped
|
|
to that partition's vaults. CI consumes them as GitHub repo secrets, injected
|
|
as `OP_SERVICE_ACCOUNT_TOKEN` per job -- the read token for `plan`, the write
|
|
token for `apply`:
|
|
|
|
| Partition | Read SA secret | Write SA secret |
|
|
|-----------|---------------------------|---------------------------------|
|
|
| staging | `OP_TF_YUCCA_STAGING_ENV` | `OP_TF_YUCCA_STAGING_ENV_WRITE` |
|
|
| prod | `OP_TF_YUCCA_PROD_ENV` | `OP_TF_YUCCA_PROD_ENV_WRITE` |
|
|
| dev | local-only -- no CI service account | -- |
|
|
|
|
Locally, operators authenticate with their own 1Password desktop session
|
|
(Futo membership) rather than a service-account token. The split -- read for
|
|
plan, write for apply -- keeps drift-detection runs from holding write
|
|
authority.
|
|
|
|
Rotation procedure: [docs/runbooks/rotate-sa-token.md](runbooks/rotate-sa-token.md).
|
|
|
|
### Item categories and naming
|
|
|
|
Per cluster, the following items live in the cluster's `vault` (currently
|
|
`yucca_tf_staging` for sietch):
|
|
|
|
| Category | Item title pattern | Field consumed |
|
|
|------------|---------------------------------------------------------|----------------------|
|
|
| Password | `<CLUSTER>_CEPH_OPS_PASSWORD` | `password` |
|
|
| Password | `<CLUSTER>_CEPH_DASHBOARD_PASSWORD` | `password` |
|
|
| Password | `<CLUSTER>_CEPH_GRAFANA_PASSWORD` | `password` |
|
|
| Password | `<CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY` | `password` |
|
|
| Password | `<CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY` | `password` |
|
|
| SSH Key | `<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY` | `private_key` / `public_key` |
|
|
| Document | `<CLUSTER>_CEPH_RGW_TLS_CERT`, `..._RGW_TLS_KEY` | file content |
|
|
| Document | `<CLUSTER>_CEPH_CLIENT_ADMIN_KEYRING` | file content |
|
|
|
|
The `<CLUSTER>_CEPH_*` prefix is hardcoded in
|
|
`tf/shared/modules/ceph-cluster/main.tf` (`secret_prefix = "${upper(var.cluster_name)}_CEPH"`)
|
|
so every Ceph-project item across all clusters is grep-discoverable as
|
|
`*_CEPH_*` regardless of the role segment in node hostnames.
|
|
|
|
Full item-by-item catalog: [docs/secrets.md](secrets.md).
|
|
|
|
### Three op-CLI patterns
|
|
|
|
The op CLI is invoked in three distinct ways across the codebase. Each
|
|
serves a different shape of secret consumption:
|
|
|
|
1. **`op run --env-file=tf/.env -- <cmd>`** -- env-var injection.
|
|
Resolves `op://` references in a dotenv file and injects the resolved
|
|
values as env vars into the child process. Used for TF (SA token) and
|
|
the S3 backend (AWS creds). Wrapped by all `mise run tf:*` tasks.
|
|
2. **`op inject -f -i <tpl> -o <out>`** -- file-template resolution.
|
|
Reads a file containing inline `op://` references, resolves each, writes
|
|
to the output path. Used by `scripts/ansible-play.sh` to render
|
|
`secrets.yml.tpl` -> tmpfile, and by the Hetzner installimage flow to
|
|
render `post-install.sh.tpl` -> `post-install.sh`.
|
|
3. **`op read "op://<vault>/<item>/<field>"`** -- single-value read.
|
|
Used by `scripts/install-ssh-keys.sh`, `rotate-ssh-key.yml`,
|
|
`post-deploy-capture.yml`. Returns one value to stdout for one specific
|
|
field; fails closed if missing.
|
|
|
|
No custom password-script (no `vault-password.sh`); no
|
|
`ansible-vault`-encrypted file in git. Lint and syntax-check tasks don't
|
|
invoke op at all -- they don't need secrets, so "1P unavailable" never
|
|
silently degrades them. This replaced an earlier `vault-password.sh` +
|
|
`ansible-vault` setup. That setup fell back to a dummy password when 1P was
|
|
unavailable, which masked real auth failures until a downstream task blew up.
|
|
The current flow fails closed instead.
|
|
|
|
---
|
|
|
|
## 6. Ansible (consumer)
|
|
|
|
### Role dependency graph
|
|
|
|
Provisioning is a separate concern (`provision.yml`, sietch only). The main
|
|
pipeline (`site.yml`) runs everything else in this order:
|
|
|
|
```mermaid
|
|
flowchart TB
|
|
PROV["provision_host<br/><i>separate playbook, live image</i>"]
|
|
BASE["baseline<br/><i>users, packages, /etc/hosts</i>"]
|
|
OST["os_tuning<br/><i>sysctl, TCP buffers</i>"]
|
|
HWT["hardware_tuning<br/><i>I/O scheduler, readahead</i>"]
|
|
DEPLOY["ceph_deploy<br/><i>bootstrap, join, OSDs, RGW,<br/>crush rules, monitoring</i>"]
|
|
CTUNE["ceph_tuning<br/><i>recovery throttling, scrub<br/>window, telemetry, audit</i>"]
|
|
SEC["security<br/><i>nftables, SSH hardening</i>"]
|
|
|
|
PROV -.->|reboot into installed OS| BASE
|
|
BASE --> OST
|
|
BASE --> HWT
|
|
OST --> DEPLOY
|
|
HWT --> DEPLOY
|
|
DEPLOY --> CTUNE
|
|
CTUNE --> SEC
|
|
|
|
classDef separate stroke-dasharray: 4 4
|
|
class PROV separate
|
|
```
|
|
|
|
`site.yml` starts at `baseline` -- `provision_host` runs only on first
|
|
install via `provision.yml`. (The roles above are imported as per-role
|
|
playbooks: `baseline.yml`, `tune-os.yml`, `tune-hardware.yml`,
|
|
`deploy-ceph.yml`, `tune-ceph.yml`, `harden.yml`.)
|
|
|
|
The split is deliberate. `provision_host` does the minimum inside the
|
|
live-image chroot -- just the `ansible-iac` user, so Ansible can connect after
|
|
reboot -- because chroot work is fragile. The ops user, packages, and
|
|
`/etc/hosts` move to the convergeable `baseline` role, which re-runs against a
|
|
live node to fix drift without reprovisioning.
|
|
|
|
The OS is installed with `debootstrap` from the live image rather than a
|
|
preseed/autoinstall. The disk layout (mdraid-1 across two SSDs, partitions
|
|
reserved for ceph block.db and SSD OSDs) needs scripted partitioning and
|
|
pre-flight hardware validation that preseed's `partman` recipes can't express.
|
|
|
|
### Why this order matters
|
|
|
|
1. **baseline before tuning** -- cephadm needs podman, dbus, chrony.
|
|
The baseline role installs these and enables the services. Running
|
|
tuning on a node without podman would leave cephadm unable to bootstrap.
|
|
2. **tuning before deploy** -- OSD daemons inherit kernel parameters
|
|
active at startup. Applying sysctl (`vm.min_free_kbytes`, `fs.aio-max-nr`)
|
|
and I/O scheduler (`mq-deadline` for HDD, `none` for SSD) before bootstrap
|
|
means daemons launch with correct limits from the first second.
|
|
3. **ceph_tuning after deploy** -- these settings use `ceph config set`
|
|
which requires a running cluster. Recovery throttling, scrub windows, and
|
|
PG autoscaler targets cannot be applied until MONs are up.
|
|
4. **security last** -- nftables drops all traffic not explicitly allowed.
|
|
Running it before ceph_deploy would block cephadm's inter-node SSH,
|
|
container image pulls, and MON/OSD port negotiation. Once the cluster is
|
|
healthy, the firewall locks it down.
|
|
|
|
### ceph_deploy internal pipeline
|
|
|
|
`roles/ceph_deploy/tasks/main.yml` orchestrates ten phases:
|
|
|
|
```mermaid
|
|
flowchart TB
|
|
P1["Phase 1, prerequisites.yml<br/><i>Ceph repo, cephadm, ceph-common</i>"]
|
|
P2["Phase 2, bootstrap.yml<br/><i>cephadm bootstrap on first node</i>"]
|
|
P3["Phase 3, join.yml<br/><i>ceph orch host add for remaining nodes</i>"]
|
|
P4["Phase 4, placement.yml<br/><i>MON/MGR placement calculation</i>"]
|
|
P45["Phase 4.5, lvm-setup.yml<br/><i>ensure block.db VGs/LVs exist (sietch-shape only;<br/>NVMe-RAID shape skips -- LVM owned by installimage post-install)</i>"]
|
|
P5["Phase 5, osds.yml<br/><i>render osd-spec.yml.j2 -> ceph orch apply osd<br/>(cephadm provisions LUKS + LVM internally)</i>"]
|
|
P55["Phase 5.5, crush-rules.yml<br/><i>replicated_hdd / replicated_ssd rules</i>"]
|
|
P575["Phase 5.75, rgw.yml<br/><i>EC pools, realm/zone, TLS, S3 user</i>"]
|
|
P58["Phase 5.8, monitoring.yml<br/><i>dashboard URL integration, Grafana creds</i>"]
|
|
P6["Phase 6, verify.yml<br/><i>cluster health report</i>"]
|
|
|
|
P1 --> P2 --> P3 --> P4 --> P45 --> P5 --> P55 --> P575 --> P58 --> P6
|
|
```
|
|
|
|
Tag-driven re-runs are first-class:
|
|
`scripts/ansible-play.sh deploy-ceph.yml --tags rgw,monitoring` re-runs
|
|
just those phases.
|
|
|
|
### Inventory layout
|
|
|
|
```
|
|
inventories/
|
|
staging-austin/sietch/ Austin staging cluster
|
|
inventory.ini TF-generated, gitignored
|
|
inventory-destroy.ini TF-generated, gitignored
|
|
inventory-provision.ini TF-generated, gitignored
|
|
secrets.yml.tpl TF-generated, gitignored
|
|
group_vars/all/vars.yml cluster-wide variables (committed)
|
|
host_vars/ per-node hardware topology (committed)
|
|
sietch-ceph-laurel.yml bond_ip, SAS path prefix, OSD maps
|
|
sietch-ceph-lawson.yml
|
|
sietch-ceph-samara.yml
|
|
installimage/ Hetzner installimage assets (sietch n/a)
|
|
```
|
|
|
|
A Hetzner NVMe-RAID cluster would follow the same layout, adding an
|
|
`installimage/` directory with a `post-install.sh.tpl` (op-injected) that
|
|
owns LVM setup. No such cluster is deployed today -- sietch is the only live
|
|
cluster -- but the module and roles already support the shape.
|
|
|
|
`host_vars/*.yml` is committed because per-node hardware facts (bond_ip, SAS
|
|
expander paths, SSD PHY positions, HDD-to-block.db mappings) are stable
|
|
inventory truth -- not operator preference. The `.local.yml` suffix is
|
|
gitignored as an escape hatch for operator-local overrides.
|
|
|
|
### Variable precedence
|
|
|
|
```mermaid
|
|
flowchart TB
|
|
D["<b>role defaults</b><br/>roles/*/defaults/main.yml<br/><i>lowest priority</i>"]
|
|
G["<b>group_vars</b><br/>inventories/<cluster>/group_vars/all/vars.yml"]
|
|
H["<b>host_vars</b><br/>inventories/<cluster>/host_vars/<host>.yml"]
|
|
T["<b>extra-vars @tmpfile</b><br/>scripts/ansible-play.sh<br/><i>op-injected secrets</i>"]
|
|
E["<b>extra-vars -e X=Y</b><br/>-e confirm_wipe=true<br/><i>highest priority</i>"]
|
|
|
|
D --> G --> H --> T --> E
|
|
```
|
|
|
|
- **Role defaults** define every tunable with a safe value
|
|
(`ceph_firewall_ssh_any_source: true`, `ceph_cpu_governor_enabled: false`).
|
|
- **group_vars/all/vars.yml** sets cluster-wide values: network topology,
|
|
Ceph release, RGW config, monitoring ports, plus the `vault_*` ->
|
|
consumable-name aliases (`ops_password: "{{ vault_ops_password }}"`).
|
|
- **host_vars** provides per-node physical topology.
|
|
- **extra-vars from @tmpfile** carries op-injected `vault_ops_password`,
|
|
`vault_ceph_dashboard_password`, `vault_grafana_admin_password`,
|
|
`vault_s3_restic_access_key`, `vault_s3_restic_secret_key`.
|
|
- **extra-vars via `-e`** carries safety gates: `confirm_wipe=true`,
|
|
`provision_skip_reboot=true`, `yes_destroy_ceph=true`.
|
|
|
|
### ansible.cfg stays generic
|
|
|
|
`ansible.cfg` contains zero site-specific values. No default inventory, no
|
|
ProxyJump, no hardcoded key paths. Site-specifics live exclusively in
|
|
`clusters.auto.tfvars` (which TF renders into the inventory) or in the
|
|
inventory's `group_vars`. The same `ansible.cfg` and the same roles work
|
|
unchanged across Austin, Hetzner, or any future cluster -- only the cluster
|
|
entry in `clusters.auto.tfvars` differs.
|
|
|
|
---
|
|
|
|
## 7. mise (orchestration surface)
|
|
|
|
### Why mise
|
|
|
|
- **Toolchain pinning** -- `.mise.toml` declares the exact versions of
|
|
`python`, `tofu`, `terragrunt`, `op`. New operators get a working
|
|
environment with `mise trust && mise run setup`.
|
|
- **Task discovery** -- `mise tasks` lists every operation; tasks are
|
|
shell-script-shaped, kept in `.mise.toml`, and committed.
|
|
- **Devtools parity** -- matches the conventions in `immich-app/devtools`
|
|
(where the `op run --env-file=tf/.env --` pattern originated).
|
|
|
|
### Task taxonomy
|
|
|
|
| Group | Tasks |
|
|
|----------------|--------------------------------------------------------------------------|
|
|
| Bootstrap | `setup` |
|
|
| Verify | `lint`, `check`, `test`, `preflight` |
|
|
| Read-only ops | `status`, `drift` |
|
|
| State change | `deploy`, `destroy`, `capture`, `backup` |
|
|
| Rotation | `rotate-certs`, `rotate-ssh-key` |
|
|
| Inventory | `hardware-inventory`, `migrate-networkd` |
|
|
| Benchmarks | `bench`, `bench-rados` |
|
|
|
|
The ceph ops tasks above live in `ansible/ceph/.mise.toml`. The `tf:*` tasks
|
|
(`tf:init`, `tf:plan`, `tf:apply`, `tf:destroy`, `tf:fmt`) live in the
|
|
yucca-root `.mise/config.toml` and run from the repo root -- they wrap
|
|
terragrunt for any stack, not just ceph.
|
|
|
|
### How mise wraps the underlying CLIs
|
|
|
|
- `mise run tf:*` -> `tf/op-run.sh terragrunt --working-dir ${TF_STACK_DIR} <cmd>` (op-run.sh = `op run --env-file=tf/.env --`)
|
|
- `mise run deploy` -> `scripts/ansible-play.sh deploy-ceph.yml ...` (per phase)
|
|
- `mise run status` -> `scripts/ansible-play.sh status.yml`
|
|
- `mise run capture` -> `scripts/ansible-play.sh post-deploy-capture.yml`
|
|
|
|
mise never invokes `ansible-playbook` or `terragrunt` directly. The wrappers
|
|
own secrets injection and pre-flight checks; mise owns task discovery and
|
|
env defaults.
|
|
|
|
### Env defaults
|
|
|
|
`CEPH_ENV` is deliberately **not** declared in `[env]` -- mise's `[env]` block
|
|
overrides shell-exported values, which would silently send an operator to the
|
|
wrong cluster. Instead each ceph ops task falls back to sietch only when
|
|
`CEPH_ENV` is unset:
|
|
|
|
```bash
|
|
# the default baked into each task
|
|
CEPH_ENV="${CEPH_ENV:-inventories/staging-austin/sietch/inventory.ini}"
|
|
|
|
# operate on another cluster by exporting once per shell, or inline:
|
|
export CEPH_ENV=inventories/staging-austin/sietch/inventory.ini
|
|
CEPH_ENV=inventories/<partition>-<region>/<cluster>/inventory.ini mise run status
|
|
```
|
|
|
|
`TF_STACK_DIR` works the same way for the root `tf:*` tasks -- it defaults to
|
|
`tf/deployment/staging/austin/ceph` and is overridden per-invocation:
|
|
|
|
```bash
|
|
TF_STACK_DIR=tf/deployment/<partition>/<region>/ceph mise run tf:plan
|
|
```
|
|
|
|
---
|
|
|
|
## 8. Wrapper scripts (the glue layer)
|
|
|
|
Three scripts under `ansible/ceph/scripts/` sit between mise and the
|
|
underlying CLIs. They exist to keep secrets out of `argv`, fail closed
|
|
when 1P is unreachable, and give better error messages than the raw tools.
|
|
|
|
| Script | Purpose |
|
|
|-------------------------|----------------------------------------------------------------------------------|
|
|
| `ansible-play.sh` | Render secrets via `op inject -f` to a `mktemp`'d file (chmod 600, trap-cleaned), then run `ansible-playbook --extra-vars @<tmpfile>` |
|
|
| `install-ssh-keys.sh` | Idempotent `op read` -> `~/.ssh/id_ed25519_<cluster>` installer; refuses overwrite on fingerprint mismatch |
|
|
| `preflight.sh` | Verifies TF artifacts present, 1P session live, SSH reachable, Python on targets -- surfaced via `mise run preflight` |
|
|
|
|
Per-script reference (synopsis, args, env, exit codes, examples):
|
|
[docs/scripts.md](scripts.md).
|
|
|
|
---
|
|
|
|
## 9. Data flow: concrete operations
|
|
|
|
### 9.1 `mise run tf:apply` -- render artifacts
|
|
|
|
```mermaid
|
|
sequenceDiagram
|
|
actor OP as Operator
|
|
participant MISE as mise
|
|
participant OPCLI as op CLI
|
|
participant ONEP as 1Password
|
|
participant TG as terragrunt / tofu
|
|
participant S3 as OVH S3
|
|
participant REPO as Repo (inventories/)
|
|
|
|
OP->>MISE: mise run tf:apply
|
|
MISE->>OPCLI: tf/op-run.sh -- ...<br/>(op run --env-file=tf/.env)
|
|
OPCLI->>ONEP: resolve op:// references
|
|
ONEP-->>OPCLI: SA token + AWS keys
|
|
OPCLI->>TG: exec child process<br/>with env vars injected
|
|
TG->>S3: read tfstate<br/>(yucca/<partition>/<region>/<stack>/terraform.tfstate)
|
|
S3-->>TG: current state
|
|
TG->>TG: plan + apply
|
|
TG->>S3: write updated tfstate
|
|
TG->>REPO: render inventory.ini,<br/>secrets.yml.tpl, ...
|
|
```
|
|
|
|
### 9.2 `mise run deploy` -- full Ceph deploy
|
|
|
|
```mermaid
|
|
sequenceDiagram
|
|
actor OP as Operator
|
|
participant MISE as mise
|
|
participant WRAP as ansible-play.sh
|
|
participant OPCLI as op CLI
|
|
participant ONEP as 1Password
|
|
participant TMP as /tmp/<id>-secrets.yml
|
|
participant ANS as ansible-playbook
|
|
participant NODES as Ceph nodes
|
|
|
|
OP->>MISE: mise run deploy
|
|
MISE->>WRAP: ansible-play.sh deploy-ceph.yml
|
|
WRAP->>OPCLI: op account get
|
|
OPCLI-->>WRAP: session OK
|
|
WRAP->>TMP: mktemp + chmod 600 + trap rm
|
|
WRAP->>OPCLI: op inject -f -i secrets.yml.tpl -o TMP
|
|
OPCLI->>ONEP: resolve op://yucca_tf_staging/SIETCH_CEPH_*/password
|
|
ONEP-->>OPCLI: secret values
|
|
OPCLI->>TMP: write resolved YAML
|
|
WRAP->>ANS: run --extra-vars @TMP
|
|
loop phases 1..6
|
|
ANS->>NODES: SSH ansible-iac@<bond_ip><br/>via id_ed25519_<cluster>
|
|
end
|
|
Note over WRAP,TMP: tmpfile rm'd on exit (trap)
|
|
```
|
|
|
|
### 9.3 `mise run capture` -- DR snapshot
|
|
|
|
```mermaid
|
|
sequenceDiagram
|
|
actor OP as Operator
|
|
participant MISE as mise
|
|
participant ANS as ansible-playbook<br/>(post-deploy-capture.yml)
|
|
participant BOOT as Bootstrap node
|
|
participant LOCAL as localhost (delegated)
|
|
participant OPCLI as op CLI
|
|
participant ONEP as 1Password
|
|
|
|
OP->>MISE: mise run capture
|
|
MISE->>ANS: ansible-play.sh post-deploy-capture.yml
|
|
ANS->>BOOT: SSH read /etc/ceph/rgw-ssl.crt
|
|
ANS->>BOOT: SSH read /etc/ceph/rgw-ssl.key
|
|
ANS->>BOOT: SSH read /etc/ceph/ceph.client.admin.keyring
|
|
BOOT-->>ANS: file contents
|
|
loop for each artifact
|
|
ANS->>LOCAL: delegate_to: localhost
|
|
LOCAL->>OPCLI: op item edit/create<br/><CLUSTER>_CEPH_<ITEM>
|
|
OPCLI->>ONEP: upsert Document item<br/>in yucca_tf_staging
|
|
end
|
|
Note over ONEP: Now holds RGW_TLS_CERT,<br/>RGW_TLS_KEY, CLIENT_ADMIN_KEYRING
|
|
```
|
|
|
|
### 9.4 `scripts/install-ssh-keys.sh` -- fresh workstation
|
|
|
|
```mermaid
|
|
sequenceDiagram
|
|
actor OP as Operator (new ws)
|
|
participant SCRIPT as install-ssh-keys.sh
|
|
participant OPCLI as op CLI
|
|
participant ONEP as 1Password
|
|
participant SSH as ~/.ssh/
|
|
|
|
OP->>SCRIPT: install-ssh-keys.sh sietch
|
|
SCRIPT->>OPCLI: op read .../public_key
|
|
OPCLI->>ONEP: SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY
|
|
ONEP-->>OPCLI: public_key
|
|
OPCLI-->>SCRIPT: pubkey content
|
|
SCRIPT->>SSH: compare with id_ed25519_sietch (if exists)
|
|
alt fingerprint match
|
|
SCRIPT-->>OP: skip (already present)
|
|
else fingerprint mismatch
|
|
SCRIPT-->>OP: refuse (operator must mv aside)
|
|
else file missing
|
|
SCRIPT->>OPCLI: op read .../private_key
|
|
OPCLI-->>SCRIPT: private_key
|
|
SCRIPT->>SSH: write id_ed25519_sietch (0600)<br/>+ .pub (0644)<br/>(umask 077)
|
|
end
|
|
```
|
|
|
|
---
|
|
|
|
## 10. OSD lifecycle
|
|
|
|
### Phase flow
|
|
|
|
```mermaid
|
|
flowchart TB
|
|
SIETCH["Sietch prep:<br/>provision_host/disks.yml partitions SSDs<br/>then ceph_deploy/lvm-setup.yml<br/><i>creates VG + db-slot LVs on each SSD's partition 5</i>"]
|
|
NVMERAID["NVMe-RAID prep:<br/>installimage post-install.sh<br/><i>NVMe RAID-1 -> vg0 -> db-slot0..13 + ssd-osd LVs</i>"]
|
|
SPEC["ceph_deploy/osds.yml renders<br/>templates/osd-spec.yml.j2 -> /etc/ceph/osd-spec.yml<br/><i>one document per host; paths from host_vars</i>"]
|
|
APPLY["ceph orch apply osd -i /etc/ceph/osd-spec.yml<br/><i>cephadm: discover disks, LUKS-format, LVM, deploy daemons</i>"]
|
|
POLL["Wait for cephadm to provision<br/><i>poll num_osds until expected count reached</i>"]
|
|
UP["Wait for OSDs up<br/><i>poll num_up_osds == num_osds</i>"]
|
|
UNSET["Defensive: ceph osd unset noin<br/><i>idempotent -- clears stale flag from prior runs</i>"]
|
|
REWEIGHT["Safety net: fix any reweight=0 OSDs"]
|
|
|
|
SIETCH --> SPEC
|
|
NVMERAID --> SPEC
|
|
SPEC --> APPLY --> POLL --> UP --> UNSET --> REWEIGHT
|
|
```
|
|
|
|
### Service-spec model, not per-disk loops
|
|
|
|
Earlier versions of this role iterated `cephadm ceph-volume lvm create`
|
|
per disk and composed `/dev/disk/by-path/...` paths from host_vars
|
|
(`sas_path_prefix` + `path_phy`). That assumed sietch's SAS expander
|
|
topology and broke on the NVMe-RAID shape's PCI-ATA disks plus LV-backed
|
|
SSD OSD.
|
|
|
|
The current flow renders a cephadm OSD service spec from per-host data
|
|
and applies it via `ceph orch apply osd -i`. Cephadm handles device
|
|
path resolution, LUKS encryption (`encrypted: true`), LVM provisioning,
|
|
and daemon deployment. The role is hardware-shape-agnostic -- the only
|
|
shape-aware logic is the template's Jinja conditional. Device paths are
|
|
listed explicitly rather than filtered by `rotational`, so cephadm never
|
|
auto-discovers and claims an OS or block.db partition.
|
|
|
|
### Hardware-shape independence in the template
|
|
|
|
`templates/osd-spec.yml.j2` renders one document per host (sietch nodes
|
|
have unique SAS prefixes per chassis, so a shared spec doesn't work)
|
|
with two shape branches:
|
|
|
|
- **Sietch** (`sas_path_prefix` defined): data path =
|
|
`/dev/disk/by-path/{{ sas_path_prefix }}-{{ path_phy }}-lun-0`; SSD
|
|
OSD = partition on the SAS-attached SSD via `path_phy + partition`.
|
|
- **NVMe-RAID shape** (`sas_path_prefix` undefined): data path =
|
|
`/dev/disk/by-path/{{ path_phy }}` (operator authors the full PCI-ATA
|
|
identifier in host_vars); SSD OSD = LV via the `lv` field
|
|
(`/dev/{{ lv }}`).
|
|
|
|
`db_devices.paths` is always `/dev/{{ db }}` -- both shapes use LVs for
|
|
block.db, no composition needed.
|
|
|
|
### Idempotency
|
|
|
|
`ceph orch apply osd` is idempotent -- re-applying the same spec is a
|
|
no-op when deployed OSDs match. New disks (populating an empty bay
|
|
later, future expansion) are picked up automatically on the next apply.
|
|
Existing OSDs are not destroyed by a spec apply -- removal requires
|
|
explicit `ceph orch osd rm`.
|
|
|
|
### Defensive noin handling
|
|
|
|
The spec-based flow doesn't need the `noin` flag (cephadm rolls out
|
|
OSDs gracefully one at a time). The role's tail still includes a
|
|
`ceph osd unset noin` task as a defensive cleanup -- stale `noin` flags
|
|
from a prior failed run of the older imperative flow can leave the
|
|
cluster degraded; the unconditional unset clears that safely (no-op
|
|
when already unset).
|
|
|
|
### Reweight-zero safety net
|
|
|
|
`osds.yml` ends with a task that fixes any OSD stuck at `reweight=0` by
|
|
running `ceph osd reweight <id> 1.0`. Rare with the spec-based flow but
|
|
kept as a backstop against an OSD coming up while `noin` was set
|
|
externally.
|
|
|
|
---
|
|
|
|
## 11. Monitoring
|
|
|
|
### What cephadm auto-deploys
|
|
|
|
cephadm's bootstrap automatically deploys:
|
|
- **node-exporter** on every node
|
|
- **ceph-exporter** on every node
|
|
- **prometheus** (single instance, cephadm-managed)
|
|
- **alertmanager** (single instance)
|
|
- **grafana** (single instance, with pre-built Ceph dashboards)
|
|
- **89 Prometheus alert rules** across 16 groups
|
|
|
|
### What we configure
|
|
|
|
`roles/ceph_deploy/tasks/monitoring.yml` handles only integration:
|
|
|
|
1. Enable the `prometheus` MGR module (if not already enabled)
|
|
2. Wait for all five monitoring service types to report `running > 0`
|
|
3. Set dashboard integration URLs for Prometheus, Alertmanager, Grafana
|
|
(using the bootstrap node's `bond_ip`)
|
|
4. Set Grafana admin credentials from the op-injected
|
|
`vault_grafana_admin_password`
|
|
5. Disable Grafana SSL cert verification in dashboard (self-signed cert)
|
|
|
|
`roles/ceph_tuning/tasks/main.yml` verifies the alert rule count and warns
|
|
if fewer than 10 rule groups are loaded (expects 16+).
|
|
|
|
---
|
|
|
|
## 12. Provision host internals
|
|
|
|
### Ten-phase flow
|
|
|
|
```mermaid
|
|
flowchart TB
|
|
P1["detect.yml<br/><i>live image + UEFI assertions, SSD discovery</i>"]
|
|
P2["prerequisites_live.yml<br/><i>apt setup on live image, install debootstrap/mdadm/lvm2</i>"]
|
|
P3["disks.yml<br/><i>partition, mdraid, LVM, mount at /mnt</i>"]
|
|
P4["install.yml<br/><i>debootstrap Bookworm into /mnt</i>"]
|
|
P5["configure.yml<br/><i>hostname, hosts, network, fstab, mdadm templates</i>"]
|
|
P6["chroot_packages.yml<br/><i>bind mounts, apt install, machine-id, SSH keys</i>"]
|
|
P7["admin_user.yml<br/><i>ansible-iac (key-only) inside chroot;<br/>ops user is created post-boot by the baseline role</i>"]
|
|
P8["bootloader.yml<br/><i>initramfs, grub-install, efibootmgr</i>"]
|
|
P9["finalize.yml<br/><i>marker, ESP mirror, unmount, reboot</i>"]
|
|
P10["unmount.yml<br/><i>reverse-order cleanup (shared with rescue)</i>"]
|
|
|
|
P1 --> P2 --> P3 --> P4 --> P5 --> P6 --> P7 --> P8 --> P9 --> P10
|
|
```
|
|
|
|
An optional `wipe-osds.yml` runs right after `disks.yml` when
|
|
`provision_wipe_osd_disks=true`, zapping prior OSD signatures off the data
|
|
disks before install -- used when rebuilding a node that was previously a Ceph
|
|
member.
|
|
|
|
### Marker-driven resume gate
|
|
|
|
After `disks.yml` runs, `main.yml` checks for
|
|
`/mnt/etc/ceph-provisioned.json`. If present and the hostname matches, all
|
|
chroot phases (4-8) plus the marker/ESP block are skipped. The role goes
|
|
straight to unmount + reboot.
|
|
|
|
This prevents:
|
|
- Re-binding bind mounts that are already in place
|
|
- Re-rotating SSH host keys (would break known_hosts)
|
|
- Re-hashing the ops password with a fresh salt
|
|
- Re-running grub-install for no reason
|
|
- Overwriting the marker with a stale `provisioned_at` timestamp
|
|
|
|
The marker filename (`ceph-provisioned.json`) is project-scoped, not
|
|
cluster-scoped -- every Ceph cluster (sietch, future) writes the
|
|
same filename. The marker's *contents* identify which cluster + host the
|
|
machine belongs to.
|
|
|
|
### Block/rescue cleanup
|
|
|
|
The entire provisioning sequence (phases 2-9) runs inside a `block/rescue`.
|
|
If any phase fails, the rescue block includes `unmount.yml` which tears
|
|
down chroot bind mounts and the /mnt hierarchy in reverse order, then
|
|
re-raises the failure. This ensures the next run starts from a clean mount
|
|
state.
|
|
|
|
---
|
|
|
|
## 13. CI/CD and roadmap
|
|
|
|
### Live today
|
|
|
|
- **CI / GitHub Actions** -- `.github/workflows/infra.yml` applies the
|
|
Terragrunt stacks from CI. A `discover` job scans
|
|
`tf/deployment/<partition>/<region>/<stack>/` into a
|
|
`{partition, region, stack}` matrix, so adding a stack needs no workflow
|
|
edit. `plan` runs with each partition's read SA; `apply` runs with the
|
|
write SA, gated behind a per-region GitHub Environment with required
|
|
reviewers (`staging-austin`, `staging-global`, `prod-global`,
|
|
`prod-htz-fsn1`). Apply order is global (NetBird + DNS) -> site NetBird ->
|
|
node-touching stacks (ceph / talos / fabric). Connectivity to the
|
|
bare-metal nodes is over the NetBird overlay.
|
|
- **Talos K8s as a sibling stack** -- `tf/deployment/<partition>/<region>/talos/`
|
|
shares the terragrunt root config and S3 backend, with its own state key
|
|
(`yucca/<partition>/<region>/talos/terraform.tfstate`). The staging/austin
|
|
talos stack is in the tree.
|
|
|
|
### Roadmap
|
|
|
|
- **TF-managed `onepassword_item` resources** -- re-enable the dormant
|
|
resources in `secrets.tf.disabled` once the dedicated ceph service account
|
|
lands, so 1P items are TF-owned rather than created by hand.
|
|
- **OSD LUKS keys in 1P** -- store dm-crypt keys for DR. Deferred until the
|
|
hybrid is stable.
|
|
|
|
---
|
|
|
|
## See also
|
|
|
|
| Topic | Doc |
|
|
|------------------------------------|--------------------------------------------------------------------------------|
|
|
| TF/Terragrunt detail | [`tf/README.md`](../../../tf/README.md) |
|
|
| Wrapper script reference | [docs/scripts.md](scripts.md) |
|
|
| Secrets catalog + rotation | [docs/secrets.md](secrets.md) |
|
|
| Trust boundaries + encryption | [docs/security-model.md](security-model.md) |
|
|
| Hardware specs + network topology | [docs/hardware.md](hardware.md) |
|
|
| Coding idioms and anti-patterns | [docs/patterns.md](patterns.md) |
|
|
| Adding a new cluster (walkthrough) | [docs/adding-a-cluster.md](adding-a-cluster.md) |
|