Files
yucca/ansible/ceph/backup-config.yml
T
Andy Molenda 63087f6850 feat(ceph): import yucca-ceph ansible + terraform infrastructure (#86)
* feat(ceph): import yucca-ceph ansible + terraform infrastructure

Imports the yucca-ceph Ansible tree into ansible/ceph/ and adds the
Terraform stack at tf/ that drives it. Cuts over from ansible-vault
to the hybrid secrets architecture (TF as inventory authority, 1P
as secrets store, op-inject at deploy time) in one atomic move.
Source: internal yucca-ceph working tree; fresh subtree-style
import, history not preserved. Andy continues operating sietch +
painbox post-merge; yucca-team hosts the code and reviews changes.

What it adds:
  - sietch (3-node Austin, production Ceph S3 backend, untouched
    by this PR)
  - painbox (single-node Hetzner SX295 in Helsinki) as a second
    deployable cluster
  - Future clusters land by appending to clusters.auto.tfvars in
    the matching environment stack (tf/deployment/<env>/ceph/) —
    no per-cluster TF code required

How it works (full map: ansible/ceph/docs/architecture.md):
  - tf/shared/modules/ceph-cluster renders inventory.ini variants
    + secrets.yml.tpl per cluster from clusters.auto.tfvars
  - secrets.yml.tpl carries op:// refs; `op inject -f` resolves
    them at deploy time from the matching yucca_tf_<env> vault
  - State in OVH yucca-tf-state bucket (key ceph/<env>/<stack>/)
  - 11 ADRs capture the decisions: ansible/ceph/docs/adr/

Out of scope (intentional):
  - LUKS keys not yet in 1P (deferred until hybrid is stable)
  - tf/shared/modules/ceph-cluster/secrets.tf.disabled is dormant;
    today's 1P items via `op item create` per
    ansible/ceph/docs/adding-a-cluster.md
  - Talos K8s on sietch is a separate workstream

Atomicity + rollback: TF-rendered inventory + secrets-template
files are gitignored (TF generates them) and ansible-vault removal
is coupled to the op-inject path. Splitting this PR lands in a
non-bootable state — merge as one unit. The merge itself is
reversible via `git revert` until the post-merge `tf:apply` runs;
after apply, full rollback needs state restore or `tofu state mv`
(land + validate before applying).

Dev-env impact: adds opentofu + terragrunt to yucca root mise tools
plus a self-contained ansible/ceph/.mise.toml. No new commands or
prereqs for immich-side contributors who don't touch ceph or run
tf:* tasks.

Verification:
  - `mise run lint` (from ansible/ceph/): 130 files, 0 warnings
  - `mise run check`: 19 playbooks parse clean
  - `mise run tf:plan`: succeeds; 7 expected file path-rename
    replacements (3 painbox + 4 sietch). State drift from import,
    no cluster-side change.
  - painbox deployed 2026-04-26 on the new code path: Bookworm +
    Ceph Tentacle, 15 OSDs (14 HDD + 1 SSD) up + in, mon/mgr/rgw
    running. HEALTH_WARN is expected on a single-node cluster.

Post-merge: from the yucca root, `mise run tf:apply` flips the
bucket state to the new monorepo paths (the 7 renames above).

* fix(ceph): exempt ansible/ and tf/ subtrees from root prettier

The imported infrastructure subtrees enforce their own format
conventions (yamllint + ansible-lint inside ansible/ceph/; tofu fmt
inside tf/). Prettier on ansible YAML reflows long Jinja2 expressions
and shell command blocks in unwanted ways, so root prettier checks
are skipped for both subtrees.

Also reformat root README.md table column alignment to match prettier
conventions (only the imported subtrees are exempt; yucca-side files
including the root README still follow root prettier rules).

* fix(ceph): clean up secrets tmpfile after ansible-playbook exits

`ansible-play.sh` rendered the resolved secrets file via `op inject`
into a `mktemp` tmpfile, set up a `trap 'rm -f "$TMPFILE"' EXIT INT
TERM`, then `exec`'d ansible-playbook. The `exec` replaced the bash
shell entirely, so the EXIT trap never fired — every play left a
plaintext-secrets file in /tmp.

In practice this was masked because /tmp is tmpfs (RAM only on this
operator's setup), so files evaporate on reboot. But within an
operator session, files accumulated linearly with each playbook
invocation. Recent count on the import-PR session: 38 files.

Drop the `exec`. With `set -euo pipefail` already on, bash:

  - propagates ansible-playbook's exit code (set -e)
  - fires the EXIT trap before exiting (always)
  - cleans up the tmpfile on success, failure, or signal

Verified: `CEPH_ENV=... scripts/ansible-play.sh status.yml
--syntax-check` creates and removes the tmpfile within the same
invocation — /tmp is clean before and after.

`scripts/preflight.sh` uses the same trap pattern but does not
`exec`, so its tmpfile cleanup was already correct (and the suffix
differs: `-secrets-test.yml` vs `-secrets.yml`, confirming
ansible-play.sh as the sole offender).
2026-05-18 06:17:56 -07:00

160 lines
5.3 KiB
YAML

---
# Export Ceph cluster configuration for disaster recovery.
# Captures: ceph.conf, admin keyring, CRUSH map, RGW realm config,
# monitor map, and OSD map to the controller's backups/ directory.
#
# Usage:
# scripts/ansible-play.sh backup-config.yml
#
# Outputs: backups/<timestamp>/ on the controller (gitignored)
# Restore: see docs/runbooks/recovery.md
- name: Backup Ceph cluster configuration
hosts: ceph_nodes
become: true
gather_facts: false
vars:
backup_timestamp: "{{ lookup('pipe', 'date +%Y%m%dT%H%M%S') }}"
backup_local_dir: "{{ playbook_dir }}/backups/{{ backup_timestamp }}"
tasks:
- name: Create local backup directory
ansible.builtin.file:
path: "{{ backup_local_dir }}"
state: directory
mode: '0700'
delegate_to: localhost
run_once: true # noqa: run-once[task]
- name: Export cluster configuration
when: inventory_hostname in groups['ceph_bootstrap']
block:
- name: Export ceph.conf
ansible.builtin.command: ceph config generate-minimal-conf
register: ceph_conf_export
changed_when: false
- name: Export admin keyring
ansible.builtin.command: ceph auth get client.admin
register: admin_keyring_export
changed_when: false
- name: Export CRUSH map (decompiled)
ansible.builtin.shell: |
set -o pipefail
ceph osd getcrushmap -o /tmp/crushmap.bin 2>/dev/null
crushtool -d /tmp/crushmap.bin -o /tmp/crushmap.txt
cat /tmp/crushmap.txt
rm -f /tmp/crushmap.bin /tmp/crushmap.txt
args:
executable: /bin/bash
register: crush_export
changed_when: false
- name: Export OSD map summary
ansible.builtin.command: ceph osd dump --format json
register: osd_dump_export
changed_when: false
- name: Export monitor map
ansible.builtin.command: ceph mon dump --format json
register: mon_dump_export
changed_when: false
- name: Export cluster config dump
ansible.builtin.command: ceph config dump --format json
register: config_dump_export
changed_when: false
- name: Export RGW realm configuration
ansible.builtin.shell: |
set -o pipefail
echo '{"realm":' && radosgw-admin realm get 2>/dev/null
echo ',"zonegroup":' && radosgw-admin zonegroup get 2>/dev/null
echo ',"zone":' && radosgw-admin zone get 2>/dev/null
echo '}'
args:
executable: /bin/bash
register: rgw_realm_export
changed_when: false
failed_when: false
- name: Export service specs
ansible.builtin.command: ceph orch ls --format yaml
register: orch_services_export
changed_when: false
- name: Export host list
ansible.builtin.command: ceph orch host ls --format yaml
register: orch_hosts_export
changed_when: false
- name: Write ceph.conf backup
ansible.builtin.copy:
content: "{{ ceph_conf_export.stdout }}\n"
dest: "{{ backup_local_dir }}/ceph.conf"
mode: '0600'
delegate_to: localhost
- name: Write admin keyring backup
ansible.builtin.copy:
content: "{{ admin_keyring_export.stdout }}\n"
dest: "{{ backup_local_dir }}/ceph.client.admin.keyring"
mode: '0600'
delegate_to: localhost
- name: Write CRUSH map backup
ansible.builtin.copy:
content: "{{ crush_export.stdout }}\n"
dest: "{{ backup_local_dir }}/crushmap.txt"
mode: '0600'
delegate_to: localhost
- name: Write OSD dump backup
ansible.builtin.copy:
content: "{{ osd_dump_export.stdout }}\n"
dest: "{{ backup_local_dir }}/osd-dump.json"
mode: '0600'
delegate_to: localhost
- name: Write monitor dump backup
ansible.builtin.copy:
content: "{{ mon_dump_export.stdout }}\n"
dest: "{{ backup_local_dir }}/mon-dump.json"
mode: '0600'
delegate_to: localhost
- name: Write config dump backup
ansible.builtin.copy:
content: "{{ config_dump_export.stdout }}\n"
dest: "{{ backup_local_dir }}/config-dump.json"
mode: '0600'
delegate_to: localhost
- name: Write RGW realm backup
ansible.builtin.copy:
content: "{{ rgw_realm_export.stdout }}\n"
dest: "{{ backup_local_dir }}/rgw-realm.json"
mode: '0600'
delegate_to: localhost
- name: Write service specs backup
ansible.builtin.copy:
content: "{{ orch_services_export.stdout }}\n"
dest: "{{ backup_local_dir }}/orch-services.yaml"
mode: '0600'
delegate_to: localhost
- name: Write host list backup
ansible.builtin.copy:
content: "{{ orch_hosts_export.stdout }}\n"
dest: "{{ backup_local_dir }}/orch-hosts.yaml"
mode: '0600'
delegate_to: localhost
- name: Report backup location
ansible.builtin.debug:
msg: "Backup written to {{ backup_local_dir }}/"
run_once: true # noqa: run-once[task]