* feat(ceph): import yucca-ceph ansible + terraform infrastructure
Imports the yucca-ceph Ansible tree into ansible/ceph/ and adds the
Terraform stack at tf/ that drives it. Cuts over from ansible-vault
to the hybrid secrets architecture (TF as inventory authority, 1P
as secrets store, op-inject at deploy time) in one atomic move.
Source: internal yucca-ceph working tree; fresh subtree-style
import, history not preserved. Andy continues operating sietch +
painbox post-merge; yucca-team hosts the code and reviews changes.
What it adds:
- sietch (3-node Austin, production Ceph S3 backend, untouched
by this PR)
- painbox (single-node Hetzner SX295 in Helsinki) as a second
deployable cluster
- Future clusters land by appending to clusters.auto.tfvars in
the matching environment stack (tf/deployment/<env>/ceph/) —
no per-cluster TF code required
How it works (full map: ansible/ceph/docs/architecture.md):
- tf/shared/modules/ceph-cluster renders inventory.ini variants
+ secrets.yml.tpl per cluster from clusters.auto.tfvars
- secrets.yml.tpl carries op:// refs; `op inject -f` resolves
them at deploy time from the matching yucca_tf_<env> vault
- State in OVH yucca-tf-state bucket (key ceph/<env>/<stack>/)
- 11 ADRs capture the decisions: ansible/ceph/docs/adr/
Out of scope (intentional):
- LUKS keys not yet in 1P (deferred until hybrid is stable)
- tf/shared/modules/ceph-cluster/secrets.tf.disabled is dormant;
today's 1P items via `op item create` per
ansible/ceph/docs/adding-a-cluster.md
- Talos K8s on sietch is a separate workstream
Atomicity + rollback: TF-rendered inventory + secrets-template
files are gitignored (TF generates them) and ansible-vault removal
is coupled to the op-inject path. Splitting this PR lands in a
non-bootable state — merge as one unit. The merge itself is
reversible via `git revert` until the post-merge `tf:apply` runs;
after apply, full rollback needs state restore or `tofu state mv`
(land + validate before applying).
Dev-env impact: adds opentofu + terragrunt to yucca root mise tools
plus a self-contained ansible/ceph/.mise.toml. No new commands or
prereqs for immich-side contributors who don't touch ceph or run
tf:* tasks.
Verification:
- `mise run lint` (from ansible/ceph/): 130 files, 0 warnings
- `mise run check`: 19 playbooks parse clean
- `mise run tf:plan`: succeeds; 7 expected file path-rename
replacements (3 painbox + 4 sietch). State drift from import,
no cluster-side change.
- painbox deployed 2026-04-26 on the new code path: Bookworm +
Ceph Tentacle, 15 OSDs (14 HDD + 1 SSD) up + in, mon/mgr/rgw
running. HEALTH_WARN is expected on a single-node cluster.
Post-merge: from the yucca root, `mise run tf:apply` flips the
bucket state to the new monorepo paths (the 7 renames above).
* fix(ceph): exempt ansible/ and tf/ subtrees from root prettier
The imported infrastructure subtrees enforce their own format
conventions (yamllint + ansible-lint inside ansible/ceph/; tofu fmt
inside tf/). Prettier on ansible YAML reflows long Jinja2 expressions
and shell command blocks in unwanted ways, so root prettier checks
are skipped for both subtrees.
Also reformat root README.md table column alignment to match prettier
conventions (only the imported subtrees are exempt; yucca-side files
including the root README still follow root prettier rules).
* fix(ceph): clean up secrets tmpfile after ansible-playbook exits
`ansible-play.sh` rendered the resolved secrets file via `op inject`
into a `mktemp` tmpfile, set up a `trap 'rm -f "$TMPFILE"' EXIT INT
TERM`, then `exec`'d ansible-playbook. The `exec` replaced the bash
shell entirely, so the EXIT trap never fired — every play left a
plaintext-secrets file in /tmp.
In practice this was masked because /tmp is tmpfs (RAM only on this
operator's setup), so files evaporate on reboot. But within an
operator session, files accumulated linearly with each playbook
invocation. Recent count on the import-PR session: 38 files.
Drop the `exec`. With `set -euo pipefail` already on, bash:
- propagates ansible-playbook's exit code (set -e)
- fires the EXIT trap before exiting (always)
- cleans up the tmpfile on success, failure, or signal
Verified: `CEPH_ENV=... scripts/ansible-play.sh status.yml
--syntax-check` creates and removes the tmpfile within the same
invocation — /tmp is clean before and after.
`scripts/preflight.sh` uses the same trap pattern but does not
`exec`, so its tmpfile cleanup was already correct (and the suffix
differs: `-secrets-test.yml` vs `-secrets.yml`, confirming
ansible-play.sh as the sole offender).
7.5 KiB
ADR-011: Cephadm OSD Service Specs over Imperative Per-Disk Loops
Status
Accepted (2026-04-26). Refines ADR-002
(explicit OSD-to-disk mapping) — preserves its anti-auto-discovery spirit
while changing where the explicit mapping is expressed (cephadm spec
applied via ceph orch apply, not an Ansible per-disk loop).
Context
roles/ceph_deploy/tasks/osds.yml originally created OSDs by iterating
each declared HDD/SSD in host_vars and invoking cephadm ceph-volume lvm create --dmcrypt --data <disk> --block.db <lv> per disk. The disk path
was composed inside the role:
DISK="/dev/disk/by-path/{{ sas_path_prefix }}-{{ item.path_phy }}-lun-0"
This worked for sietch (Dell R730xd, SAS expander, dual-SSD-VG topology) because every assumption baked into the composition was sietch-shape:
sas_path_prefixexists per hostpath_phyis a slot identifier (phy0,phy1, ...) appended to the prefix-lun-0suffix is the SAS LUN convention- SSD OSDs are partitions on the boot SSD (
-part6)
When the painbox (Hetzner SX295) cluster came online for the first integration test of this codebase, every one of those assumptions broke:
- Painbox is SATA + NVMe, no SAS expander →
sas_path_prefixundefined, role failed at template-resolution time. - SATA by-path strings are
pci-XXXX:XX:XX.X-ata-Ndirectly (no-lun-Nsuffix) — operator authors the full path identifier aspath_phy. - Painbox SSD OSD is an LV on the NVMe RAID-1 VG (
vg0/ssd-osd), not a partition — different shape entirely from sietch's SSD OSDs.
The fix could have been per-shape branching inside osds.yml's shell loop
(if/else for path composition; if/else for partition vs LV), but every
phase of the role beyond OSDs (RGW, monitoring, etc.) was already migrating
toward cephadm's declarative service-spec pattern (rgw-spec.yaml.j2 was
already in place). The OSD path was the obvious next migration.
Decision
OSD creation is now declarative via a cephadm OSD service spec rendered
from per-host data, applied via ceph orch apply osd -i /etc/ceph/osd-spec.yml.
templates/osd-spec.yml.j2renders one document per host (HDD spec- optional SSD spec). Per-host because sietch nodes have unique SAS
prefixes per chassis — a shared spec with
data_devices.pathsrequires identical paths across allplacement.hosts.
- optional SSD spec). Per-host because sietch nodes have unique SAS
prefixes per chassis — a shared spec with
- Hardware-shape branching lives in the template's Jinja conditional,
not in role logic:
sas_path_prefix is defined→ sietch-shape path compositionsas_path_prefix is undefined→ painbox-shape, usepath_phydirectly (full PCI-ATA identifier from host_vars)- SSD OSD:
lv is defined→ painbox-style LV path; else sietch-style partition-on-SAS-disk
tasks/osds.ymlis now thin: render template →ceph orch apply→ wait for provisioning → wait for OSDs up → defensiveceph osd unset noin→ reweight-zero safety net. ~150 LOC down from ~230.- Encryption is set in the spec (
encrypted: true), not via a per-disk--dmcryptflag. Cephadm provisions LUKS internally and stores the dm-crypt key in the MON config-key store as before. - Idempotency is provided by cephadm — re-applying the same spec is
a no-op. New disks (e.g. populating an empty bay) are picked up
automatically. Existing OSDs are not destroyed by a spec change;
removal requires
ceph orch osd rm.
Out of scope (for this ADR)
- Extending the spec pattern to MON/MGR/RGW/monitoring placement
—
rgw-spec.yaml.j2already does this for RGW; the rest of the role still uses imperativeceph orch apply mon --placement=...(and similar for MGR/monitoring). Refactoring those is a separate body of work — see "Option C" in the architecture session log; planned as a dedicated follow-up PR. - TF-rendered service specs. TF already has all the cluster identity it needs to render every cephadm spec deterministically. Moving spec rendering from Ansible (Jinja in role templates) to TF (Jinja in module templates) would shift the boundary further toward "TF declares, Ansible applies." Same Option-C follow-up.
- Auto-discovery via
data_devices.rotational: 1filter. Cephadm supports filter-based device selection. We chose explicitpathsto preserve ADR-002's "no auto-discovery surprises" stance — empty bays on sietch and the SSD OSD on painbox can't be expressed with filters alone.
Consequences
- Positive: hardware-shape independence. Painbox (SATA + NVMe-RAID
- LV-backed SSD OSD) and sietch (SAS expander + dual-SSD-VG) deploy through the same role with the same task file. Future clusters with yet other shapes (Hetzner AX-line, Equinix bare-metal, etc.) need only their own host_vars; the role doesn't change.
- Positive:
osds.ymlshrinks ~35%; the deleted code was the most fragile part (per-disk shell loops with embedded Jinja path composition). - Positive: ADR-002's explicit-mapping spirit is preserved. Paths are still listed explicitly in the spec — cephadm doesn't auto-discover. Empty bays stay empty; partitions stay reserved for OS/block.db.
- Positive: Aligns with cephadm's intended deployment model. All modern cephadm operators use service specs; the imperative-loop pattern is legacy.
- Neutral: The spec is rendered to
/etc/ceph/osd-spec.ymlon the bootstrap node and then applied. The file is overwritten on each re-render — operators inspecting the deployed state should query cephadm (ceph orch ls --service-type osd --export) rather than reading the on-disk spec, which may have been updated since the last apply. - Negative: Cephadm's spec apply is async — the role polls until the expected OSD count is reached, with a 15-minute timeout. On a large cluster (hundreds of OSDs), this could exceed the timeout. Not a concern at current scale (15 OSDs painbox, 30 OSDs sietch).
- Negative: Debugging "why isn't this disk becoming an OSD?" is
harder than with the per-disk loop, where the failing disk had its
own log line. With cephadm, you check
ceph orch ls,ceph orch ps, andceph cephadm osd activate <host> --dry-runon the bootstrap node.
Migration that landed with this ADR
templates/osd-spec.yml.j2written (replaces stub that existed but was never wired in).tasks/osds.ymlrewritten to render + apply + poll + safety net.tasks/lvm-setup.ymlgated onsas_path_prefix is defined— sietch still needs the role's defensive VG/LV recovery path; painbox skips because installimage's post-install owns LVM lifecycle.- Validated end-to-end against a freshly-reprovisioned painbox: 15/15
OSDs up + in, all encrypted (15 dm-crypt keys in MON store), correct
per-host service IDs (
osd.painbox-ceph-evelyn-hdd,osd.painbox-ceph-evelyn-ssd). - Sietch validation deferred — sietch is currently deployed and healthy; the spec apply against existing OSDs is idempotent (no-op when paths match), but a real test on sietch is a separate operational step planned for the next sietch maintenance window.
References
- ADR-002 — explicit OSD-to-disk mapping (refined, not superseded)
- ADR-009 — TF-first secrets (companion architectural shift)
roles/ceph_deploy/templates/osd-spec.yml.j2— the templateroles/ceph_deploy/tasks/osds.yml— the thin applier- Cephadm OSD Service docs