mirror of
https://github.com/immich-app/yucca.git
synced 2026-09-30 21:37:50 +08:00
feat(ceph): import yucca-ceph ansible + terraform infrastructure (#86)
* feat(ceph): import yucca-ceph ansible + terraform infrastructure
Imports the yucca-ceph Ansible tree into ansible/ceph/ and adds the
Terraform stack at tf/ that drives it. Cuts over from ansible-vault
to the hybrid secrets architecture (TF as inventory authority, 1P
as secrets store, op-inject at deploy time) in one atomic move.
Source: internal yucca-ceph working tree; fresh subtree-style
import, history not preserved. Andy continues operating sietch +
painbox post-merge; yucca-team hosts the code and reviews changes.
What it adds:
- sietch (3-node Austin, production Ceph S3 backend, untouched
by this PR)
- painbox (single-node Hetzner SX295 in Helsinki) as a second
deployable cluster
- Future clusters land by appending to clusters.auto.tfvars in
the matching environment stack (tf/deployment/<env>/ceph/) —
no per-cluster TF code required
How it works (full map: ansible/ceph/docs/architecture.md):
- tf/shared/modules/ceph-cluster renders inventory.ini variants
+ secrets.yml.tpl per cluster from clusters.auto.tfvars
- secrets.yml.tpl carries op:// refs; `op inject -f` resolves
them at deploy time from the matching yucca_tf_<env> vault
- State in OVH yucca-tf-state bucket (key ceph/<env>/<stack>/)
- 11 ADRs capture the decisions: ansible/ceph/docs/adr/
Out of scope (intentional):
- LUKS keys not yet in 1P (deferred until hybrid is stable)
- tf/shared/modules/ceph-cluster/secrets.tf.disabled is dormant;
today's 1P items via `op item create` per
ansible/ceph/docs/adding-a-cluster.md
- Talos K8s on sietch is a separate workstream
Atomicity + rollback: TF-rendered inventory + secrets-template
files are gitignored (TF generates them) and ansible-vault removal
is coupled to the op-inject path. Splitting this PR lands in a
non-bootable state — merge as one unit. The merge itself is
reversible via `git revert` until the post-merge `tf:apply` runs;
after apply, full rollback needs state restore or `tofu state mv`
(land + validate before applying).
Dev-env impact: adds opentofu + terragrunt to yucca root mise tools
plus a self-contained ansible/ceph/.mise.toml. No new commands or
prereqs for immich-side contributors who don't touch ceph or run
tf:* tasks.
Verification:
- `mise run lint` (from ansible/ceph/): 130 files, 0 warnings
- `mise run check`: 19 playbooks parse clean
- `mise run tf:plan`: succeeds; 7 expected file path-rename
replacements (3 painbox + 4 sietch). State drift from import,
no cluster-side change.
- painbox deployed 2026-04-26 on the new code path: Bookworm +
Ceph Tentacle, 15 OSDs (14 HDD + 1 SSD) up + in, mon/mgr/rgw
running. HEALTH_WARN is expected on a single-node cluster.
Post-merge: from the yucca root, `mise run tf:apply` flips the
bucket state to the new monorepo paths (the 7 renames above).
* fix(ceph): exempt ansible/ and tf/ subtrees from root prettier
The imported infrastructure subtrees enforce their own format
conventions (yamllint + ansible-lint inside ansible/ceph/; tofu fmt
inside tf/). Prettier on ansible YAML reflows long Jinja2 expressions
and shell command blocks in unwanted ways, so root prettier checks
are skipped for both subtrees.
Also reformat root README.md table column alignment to match prettier
conventions (only the imported subtrees are exempt; yucca-side files
including the root README still follow root prettier rules).
* fix(ceph): clean up secrets tmpfile after ansible-playbook exits
`ansible-play.sh` rendered the resolved secrets file via `op inject`
into a `mktemp` tmpfile, set up a `trap 'rm -f "$TMPFILE"' EXIT INT
TERM`, then `exec`'d ansible-playbook. The `exec` replaced the bash
shell entirely, so the EXIT trap never fired — every play left a
plaintext-secrets file in /tmp.
In practice this was masked because /tmp is tmpfs (RAM only on this
operator's setup), so files evaporate on reboot. But within an
operator session, files accumulated linearly with each playbook
invocation. Recent count on the import-PR session: 38 files.
Drop the `exec`. With `set -euo pipefail` already on, bash:
- propagates ansible-playbook's exit code (set -e)
- fires the EXIT trap before exiting (always)
- cleans up the tmpfile on success, failure, or signal
Verified: `CEPH_ENV=... scripts/ansible-play.sh status.yml
--syntax-check` creates and removes the tmpfile within the same
invocation — /tmp is clean before and after.
`scripts/preflight.sh` uses the same trap pattern but does not
`exec`, so its tmpfile cleanup was already correct (and the suffix
differs: `-secrets-test.yml` vs `-secrets.yml`, confirming
ansible-play.sh as the sole offender).
This commit is contained in:
+202
@@ -0,0 +1,202 @@
|
||||
# yucca/tf
|
||||
|
||||
Terraform/OpenTofu authority for cluster identity, 1P secret items, and the
|
||||
inventory artifacts Ansible consumes. Multi-env via terragrunt.
|
||||
|
||||
## Layout
|
||||
|
||||
```
|
||||
tf/
|
||||
├── .env ← op:// references (committed; no literal secrets)
|
||||
├── shared/
|
||||
│ └── modules/
|
||||
│ └── ceph-cluster/ ← per-cluster orchestration module
|
||||
│ ├── main.tf, variables.tf, outputs.tf, rendering.tf
|
||||
│ ├── wordlist.txt ← 923 words for auto-picked hostnames
|
||||
│ └── templates/
|
||||
│ ├── inventory.ini.tftpl
|
||||
│ ├── inventory-destroy.ini.tftpl
|
||||
│ ├── inventory-provision-debian-live.ini.tftpl
|
||||
│ └── secrets.yml.tpl.tftpl
|
||||
└── deployment/
|
||||
├── terragrunt.hcl ← root: state backend, env/stack derived from path
|
||||
└── dev/
|
||||
└── ceph/
|
||||
├── terragrunt.hcl ← include root + stack-level inputs
|
||||
├── versions.tf, variables.tf, main.tf
|
||||
├── clusters.auto.tfvars ← declarative cluster list (edit here to add/modify)
|
||||
└── .terraform.lock.hcl
|
||||
```
|
||||
|
||||
Future envs land as siblings: `deployment/staging/ceph/`, `deployment/prod/ceph/`.
|
||||
Future stacks land as siblings of `ceph/` within an env:
|
||||
`deployment/dev/talos/`, `deployment/dev/monitoring/`, etc.
|
||||
|
||||
## Conventions
|
||||
|
||||
### Env and stack are derived from the directory path
|
||||
|
||||
`deployment/terragrunt.hcl` parses the child's relative path to extract
|
||||
`env` and `stack`:
|
||||
|
||||
```
|
||||
deployment/dev/ceph → env = dev, stack = ceph
|
||||
deployment/staging/ceph → env = staging, stack = ceph
|
||||
deployment/prod/talos → env = prod, stack = talos
|
||||
```
|
||||
|
||||
The state backend key is derived from these values:
|
||||
`ceph/${env}/${stack}/terraform.tfstate` in the shared `yucca-tf-state` S3
|
||||
bucket. Project-scoped so ceph state doesn't collide with o11y or future
|
||||
stacks in the same bucket.
|
||||
|
||||
### The `op run --env-file=tf/.env --` pattern
|
||||
|
||||
`tf/.env` holds 1Password `op://` references — **not literal secrets**:
|
||||
|
||||
```sh
|
||||
export OP_SERVICE_ACCOUNT_TOKEN="op://yucca_tf_dev/yucca_futo_1pass_superuser_service_account/password"
|
||||
```
|
||||
|
||||
Wrap every terragrunt invocation with `op run --env-file=tf/.env --` (the
|
||||
mise `tf:*` tasks do this automatically). The op CLI resolves the `op://`
|
||||
reference and injects the actual token as `OP_SERVICE_ACCOUNT_TOKEN` into
|
||||
the child process's environment. The 1P Terraform provider picks it up
|
||||
from the env var and authenticates.
|
||||
|
||||
The same pattern is used in `immich-app/devtools` and is the Futo-wide
|
||||
convention for TF secret injection.
|
||||
|
||||
### Committed `.env` is safe because it's just pointers
|
||||
|
||||
Yucca's root `.gitignore` normally excludes `.env` files — we add an
|
||||
explicit `!tf/.env` exception. This file contains only `op://` URIs; no
|
||||
secret ever transits the repo. It's a committed manifest of "which 1P
|
||||
items this TF depends on."
|
||||
|
||||
### Stack override via `TF_STACK_DIR`
|
||||
|
||||
The default `mise run tf:*` tasks target `tf/deployment/dev/ceph`. Point
|
||||
them at another stack via the `TF_STACK_DIR` env var:
|
||||
|
||||
```bash
|
||||
TF_STACK_DIR=tf/deployment/staging/ceph mise run tf:plan
|
||||
TF_STACK_DIR=tf/deployment/dev/talos mise run tf:apply
|
||||
```
|
||||
|
||||
## Running TF
|
||||
|
||||
### One-shot (preferred for now)
|
||||
|
||||
```bash
|
||||
mise run tf:init # first time in a stack
|
||||
mise run tf:plan # dry run
|
||||
mise run tf:apply # render artifacts + (future) create 1P items
|
||||
```
|
||||
|
||||
These wrap: `op run --env-file=tf/.env -- terragrunt --working-dir <stack> <cmd>`.
|
||||
|
||||
### State backend
|
||||
|
||||
Remote: shared `yucca-tf-state` S3 bucket at OVH Paris
|
||||
(`https://s3.eu-west-par.io.cloud.ovh.net/`). Key path:
|
||||
`ceph/${env}/${stack}/terraform.tfstate` — project-scoped so ceph state
|
||||
doesn't collide with o11y or future stacks in the same bucket.
|
||||
|
||||
Credentials are AWS-compatible env vars (`AWS_ACCESS_KEY_ID` /
|
||||
`AWS_SECRET_ACCESS_KEY`), injected via `op run --env-file=tf/.env` from
|
||||
the `TF_STATE_S3_*` items in the `yucca_tf` vault. OVH-specific config
|
||||
(skip AWS validation, path-style URLs) is in
|
||||
`deployment/terragrunt.hcl`.
|
||||
|
||||
**State locking is not enabled.** OVH has no DynamoDB equivalent;
|
||||
OpenTofu's `use_lockfile = true` option would handle single-bucket locking
|
||||
but expects the lockfile object to already exist — `terragrunt init`
|
||||
against a fresh backend fails with 404 before it can create one. Enable
|
||||
it once the concern is concurrent applies (multiple operators working the
|
||||
same stack simultaneously). Single operator today → low risk.
|
||||
|
||||
## The ceph-cluster module
|
||||
|
||||
Declarative input in `clusters.auto.tfvars`:
|
||||
|
||||
```hcl
|
||||
clusters = {
|
||||
sietch = {
|
||||
domain = "dev.austin.int.futo.cloud"
|
||||
environment = "dev"
|
||||
datacenter = "austin"
|
||||
provider_code = "int"
|
||||
role_in_hostname = "ceph"
|
||||
ansible_ssh_user = "ansible-iac"
|
||||
ansible_ssh_key = "~/.ssh/id_ed25519_sietch"
|
||||
vault = "yucca_tf_dev"
|
||||
provision_profile = "debian-live" # null for Hetzner-installimage clusters
|
||||
hosts = [
|
||||
{ name = "laurel", bond_ip = "10.10.10.90", bootstrap = true },
|
||||
{ name = "lawson", bond_ip = "10.10.10.91" },
|
||||
{ name = "samara", bond_ip = "10.10.10.92" },
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
On apply, the module:
|
||||
|
||||
1. Picks wordlist names for `hosts[].name == null` (stable across applies;
|
||||
seeded per-cluster; operator-declared names excluded from the pool to
|
||||
prevent collisions).
|
||||
2. Renders `inventory.ini` (normal ops), `inventory-destroy.ini` (explicit
|
||||
destroy flag), `secrets.yml.tpl` (op:// references to
|
||||
`yucca_tf_dev/<CLUSTER>_CEPH_*_PASSWORD/password`). Optionally renders
|
||||
`inventory-provision.ini` when `provision_profile != null`.
|
||||
3. **(Not yet TF-managed)** `onepassword_item` resources for cluster
|
||||
secrets are dormant — items are created via `op` CLI today and read
|
||||
by Ansible at play time. See [ADR-009](../ansible/ceph/docs/adr/009-tf-first-op-inject-over-vault-password-sh.md)
|
||||
for the re-enable plan.
|
||||
|
||||
## Where secrets actually live
|
||||
|
||||
- **`yucca_tf_dev`** (team-shared): live values consumed by Ansible at
|
||||
play time. Password items per cluster (ops, dashboard, grafana, S3
|
||||
svc-user access + secret), SSH Key items per cluster (ansible-iac
|
||||
keypairs), and DR-capture items per cluster (RGW TLS cert + key,
|
||||
client.admin keyring — populated by `mise run capture`).
|
||||
- **`yucca_tf_dev_manual`** (team-shared): placeholders for
|
||||
human-fillable secrets (API tokens, OAuth client secrets). Not yet used
|
||||
by ceph-cluster.
|
||||
|
||||
Service accounts themselves are in `yucca_tf_dev` as two items:
|
||||
|
||||
| SA | Purpose | Consumed by |
|
||||
|---|---|---|
|
||||
| `yucca_futo_1pass_superuser_service_account` | Read + write all `yucca_tf_*` vaults | TF (via `tf/.env`) |
|
||||
| `yucca_futo_1pass_service_account` | Read-only on `yucca_tf` and `yucca_tf_dev` | Ansible runtime / CI |
|
||||
|
||||
Both are shared with other Futo consumers (o11y, base Yucca infra).
|
||||
Rotation affects all of them — see `ansible/ceph/docs/runbooks/rotate-sa-token.md`
|
||||
for the coordination procedure.
|
||||
|
||||
## Adding a new cluster
|
||||
|
||||
1. Add an entry to `clusters.auto.tfvars`.
|
||||
2. Create the 1P items in `yucca_tf_dev`:
|
||||
- Password items: `<CLUSTER>_CEPH_{OPS,DASHBOARD,GRAFANA}_PASSWORD`, plus
|
||||
`<CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_{ACCESS,SECRET}_KEY`. Use
|
||||
`op item create --generate-password` for each.
|
||||
- SSH Key item: `<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY` via
|
||||
`op item create --category "SSH Key" --ssh-generate-key=ed25519`.
|
||||
3. `mise run tf:apply` — renders inventory + secrets template.
|
||||
4. Create per-node `host_vars/*.yml` files in the new inventory dir (hardware
|
||||
topology; not TF-rendered yet).
|
||||
5. On operator workstation: `scripts/install-ssh-keys.sh <cluster>` to pull
|
||||
the private key from 1P.
|
||||
6. After first successful deploy: `mise run capture` to snapshot the RGW
|
||||
TLS material + admin keyring to 1P for DR.
|
||||
7. See [`ansible/ceph/docs/adding-a-cluster.md`](../ansible/ceph/docs/adding-a-cluster.md)
|
||||
for the full walk-through.
|
||||
|
||||
## Related
|
||||
|
||||
- ADR-009 (TF-first + op inject): `ansible/ceph/docs/adr/009-tf-first-op-inject-over-vault-password-sh.md`
|
||||
- Immich devtools (upstream pattern): <https://github.com/immich-app/devtools/tree/main/tf>
|
||||
Reference in New Issue
Block a user