Files
yucca/ansible/ceph/docs/patterns.md
T
Andy Molenda 63087f6850 feat(ceph): import yucca-ceph ansible + terraform infrastructure (#86)
* feat(ceph): import yucca-ceph ansible + terraform infrastructure

Imports the yucca-ceph Ansible tree into ansible/ceph/ and adds the
Terraform stack at tf/ that drives it. Cuts over from ansible-vault
to the hybrid secrets architecture (TF as inventory authority, 1P
as secrets store, op-inject at deploy time) in one atomic move.
Source: internal yucca-ceph working tree; fresh subtree-style
import, history not preserved. Andy continues operating sietch +
painbox post-merge; yucca-team hosts the code and reviews changes.

What it adds:
  - sietch (3-node Austin, production Ceph S3 backend, untouched
    by this PR)
  - painbox (single-node Hetzner SX295 in Helsinki) as a second
    deployable cluster
  - Future clusters land by appending to clusters.auto.tfvars in
    the matching environment stack (tf/deployment/<env>/ceph/) —
    no per-cluster TF code required

How it works (full map: ansible/ceph/docs/architecture.md):
  - tf/shared/modules/ceph-cluster renders inventory.ini variants
    + secrets.yml.tpl per cluster from clusters.auto.tfvars
  - secrets.yml.tpl carries op:// refs; `op inject -f` resolves
    them at deploy time from the matching yucca_tf_<env> vault
  - State in OVH yucca-tf-state bucket (key ceph/<env>/<stack>/)
  - 11 ADRs capture the decisions: ansible/ceph/docs/adr/

Out of scope (intentional):
  - LUKS keys not yet in 1P (deferred until hybrid is stable)
  - tf/shared/modules/ceph-cluster/secrets.tf.disabled is dormant;
    today's 1P items via `op item create` per
    ansible/ceph/docs/adding-a-cluster.md
  - Talos K8s on sietch is a separate workstream

Atomicity + rollback: TF-rendered inventory + secrets-template
files are gitignored (TF generates them) and ansible-vault removal
is coupled to the op-inject path. Splitting this PR lands in a
non-bootable state — merge as one unit. The merge itself is
reversible via `git revert` until the post-merge `tf:apply` runs;
after apply, full rollback needs state restore or `tofu state mv`
(land + validate before applying).

Dev-env impact: adds opentofu + terragrunt to yucca root mise tools
plus a self-contained ansible/ceph/.mise.toml. No new commands or
prereqs for immich-side contributors who don't touch ceph or run
tf:* tasks.

Verification:
  - `mise run lint` (from ansible/ceph/): 130 files, 0 warnings
  - `mise run check`: 19 playbooks parse clean
  - `mise run tf:plan`: succeeds; 7 expected file path-rename
    replacements (3 painbox + 4 sietch). State drift from import,
    no cluster-side change.
  - painbox deployed 2026-04-26 on the new code path: Bookworm +
    Ceph Tentacle, 15 OSDs (14 HDD + 1 SSD) up + in, mon/mgr/rgw
    running. HEALTH_WARN is expected on a single-node cluster.

Post-merge: from the yucca root, `mise run tf:apply` flips the
bucket state to the new monorepo paths (the 7 renames above).

* fix(ceph): exempt ansible/ and tf/ subtrees from root prettier

The imported infrastructure subtrees enforce their own format
conventions (yamllint + ansible-lint inside ansible/ceph/; tofu fmt
inside tf/). Prettier on ansible YAML reflows long Jinja2 expressions
and shell command blocks in unwanted ways, so root prettier checks
are skipped for both subtrees.

Also reformat root README.md table column alignment to match prettier
conventions (only the imported subtrees are exempt; yucca-side files
including the root README still follow root prettier rules).

* fix(ceph): clean up secrets tmpfile after ansible-playbook exits

`ansible-play.sh` rendered the resolved secrets file via `op inject`
into a `mktemp` tmpfile, set up a `trap 'rm -f "$TMPFILE"' EXIT INT
TERM`, then `exec`'d ansible-playbook. The `exec` replaced the bash
shell entirely, so the EXIT trap never fired — every play left a
plaintext-secrets file in /tmp.

In practice this was masked because /tmp is tmpfs (RAM only on this
operator's setup), so files evaporate on reboot. But within an
operator session, files accumulated linearly with each playbook
invocation. Recent count on the import-PR session: 38 files.

Drop the `exec`. With `set -euo pipefail` already on, bash:

  - propagates ansible-playbook's exit code (set -e)
  - fires the EXIT trap before exiting (always)
  - cleans up the tmpfile on success, failure, or signal

Verified: `CEPH_ENV=... scripts/ansible-play.sh status.yml
--syntax-check` creates and removes the tmpfile within the same
invocation — /tmp is clean before and after.

`scripts/preflight.sh` uses the same trap pattern but does not
`exec`, so its tmpfile cleanup was already correct (and the suffix
differs: `-secrets-test.yml` vs `-secrets.yml`, confirming
ansible-play.sh as the sole offender).
2026-05-18 06:17:56 -07:00

15 KiB

Patterns

Project-specific Ansible idioms. This doc skips generic Ansible hygiene (FQCN, set -o pipefail, etc. — those are table stakes) and focuses on patterns that are non-obvious or specific to how this Ceph automation is built.

For how patterns wire into the wrapper + secrets flow, see scripts.md. For role structure and the pre-submit checklist, see adding-a-role.md.


Check-then-set against the cluster

Problem: ceph config set always reports changed. Running it unconditionally makes every play dirty and obscures real drift.

Pattern: read the current value, compare to the expected value, apply only if different. The comparison is the key — raw ceph config set with changed_when: true is a lie.

Canonical example — roles/ceph_tuning/tasks/main.yml:

- name: Read current OSD config values
  ansible.builtin.shell: |
    set -o pipefail
    echo "recovery_max_active=$(ceph config get osd osd_recovery_max_active)"
    # ...
  register: current_osd_config
  changed_when: false

- name: Apply Ceph config values
  ansible.builtin.command: >
    ceph config set {{ item.section }} {{ item.key }} {{ item.value }}
  loop:
    - { section: osd, key: osd_recovery_max_active, value: "{{ ... }}",
        current: "{{ osd_cfg.recovery_max_active | default('') | float }}" }
  when: item.current | float != item.value | float
  changed_when: true

Other instances: rgw.yml (zonegroup hostnames), crush-rules.yml (rule existence), monitoring.yml (module enable check).

Float comparison gotcha

ceph config get osd osd_recovery_sleep_hdd returns 0.100000, but the Ansible variable is 0.1. String comparison fails; integer comparison truncates. Always cast both sides to | float before comparing. Affects any Ceph config value returned with trailing zeros (osd_recovery_sleep_hdd, osd_deep_scrub_interval, etc.).


Marker-driven idempotency

Problem: provisioning is destructive and multi-phase. A crash during phase 6 must not re-wipe disks on the next run. But the role still needs to handle a completely fresh node.

Pattern: write a JSON marker at the end of provisioning. On subsequent runs, check for the marker and skip completed phases.

The marker filename (/etc/ceph-provisioned.json) is project-scoped — every Ceph cluster writes the same filename. The marker's contents identify which cluster + host the machine belongs to (hostname, fqdn, bond_ip, cluster_name, SSD serials, provisioned_at timestamp).

Template: roles/provision_host/templates/ceph-provisioned.json.j2.

Resume gate — roles/provision_host/tasks/main.yml:

- name: Check if provisioning marker is present in chroot
  ansible.builtin.stat:
    path: "{{ provision_mnt }}{{ provision_marker_path }}"
  register: marker_stat

- name: Set provisioning_done fact
  ansible.builtin.set_fact:
    provisioning_done: "{{ marker_stat.stat.exists | bool }}"

Then every chroot phase is gated on when: not provisioning_done.

disks.yml additionally handles the "md array already assembled but nothing mounted" case — mounts, reads the marker, validates the hostname matches inventory_hostname, and either resumes (marker matches) or unmounts and re-wipes (marker missing or wrong host).

Use when: any multi-step destructive workflow where partial completion must be resumable. The marker must contain enough identity information to distinguish "this node's previous run" from "a different node's leftover state."


Block/rescue cleanup

Problem: if provisioning fails mid-chroot (e.g., debootstrap network error), bind mounts at /mnt/dev, /mnt/proc, /mnt/sys remain active. The next run fails because it can't cleanly remount.

Pattern: wrap the phase sequence in block/rescue. The rescue includes a shared unmount.yml that tears down mounts in reverse order, then re-raises the failure.

- name: Provisioning phases
  block:
    - ansible.builtin.import_tasks: prerequisites_live.yml
    - ansible.builtin.import_tasks: disks.yml
    # ... phases 4-8 ...
    - ansible.builtin.import_tasks: finalize.yml
  rescue:
    - name: Unmount /mnt hierarchy on failure (best-effort cleanup)
      ansible.builtin.include_tasks: unmount.yml
    - name: Re-raise failure
      ansible.builtin.fail:
        msg: "Provisioning phase failed for {{ inventory_hostname }}."

The unmount tasks use failed_when: false — if a path isn't mounted, we just want to keep going.


Conditional features

Problem: not every cluster needs iSCSI, NFS, CPU governor tuning, or centralized logging. These features should be zero-overhead when disabled.

Pattern: gate on <feature>_enabled | bool with defaults of false.

Current feature flags:

Feature Toggle Default Consumer
CPU governor ceph_cpu_governor_enabled false hardware_tuning
Centralized logging ceph_logging_enabled false os_tuning
iSCSI firewall ceph_firewall_iscsi_enabled false security
NFS firewall ceph_firewall_nfs_enabled false security
Firewall overall ceph_firewall_enabled true security
RGW TLS ceph_rgw_ssl false ceph_deploy/rgw
Audit logging ceph_audit_enabled true ceph_tuning
SSH open to all sources (dev) ceph_firewall_ssh_any_source true (dev) security
Weekly fstrim timer (SSDs) ceph_enable_fstrim_timer true hardware_tuning
LVM device filter ceph_lvm_filter_enabled false hardware_tuning

The | bool filter is mandatory. Ansible may pass booleans as strings from inventory or extra-vars; without | bool, the string "false" is truthy.

In Jinja templates (e.g. nftables.conf.j2):

{% if ceph_firewall_iscsi_enabled | bool %}
        ip saddr {{ net }} tcp dport {{ ceph_firewall_iscsi_port }} accept
{% endif %}

Drift detection pattern

drift.yml is a read-only play that compares expected state against live cluster. Three steps:

  1. Load all role defaults via vars_files — gives drift detection access to expected values without depending on any role's execution:

    vars_files:
      - roles/baseline/defaults/main.yml
      - roles/os_tuning/defaults/main.yml
      # ...
    
  2. Accumulate results into a list via set_fact:

    drift_results: >-
      {{ drift_results + [{
        'category': 'sysctl',
        'item': item.item.key,
        'expected': item.item.expected | string,
        'actual': item.stdout | trim,
        'match': (item.stdout | trim) == (item.item.expected | string)
      }] }}
    
  3. Generate a formatted report using Jinja in set_fact.

Categories checked: sysctl values, HDD/SSD I/O schedulers, nftables policy, SSH PasswordAuthentication, ops sudo config, OSD status, MON quorum, RGW daemon count, cluster health, Ceph config values.

Use when: building read-only comparison plays. The pattern generalizes to any "expected vs actual" audit.


CEPH_ENV as inventory selector

Wrappers and downstream scripts derive the cluster's paths from the CEPH_ENV environment variable, which points at the TF-rendered inventory file. Cluster identity is authoritative in clusters.auto.tfvars; CEPH_ENV is the runtime pointer.

CEPH_ENV = inventories/sietch-ceph.dev.austin.int/inventory.ini
           |
           dirname ->  inventories/sietch-ceph.dev.austin.int
                       |
                       + "/secrets.yml.tpl"  -> op inject input

scripts/ansible-play.sh derives the secrets template path as $(dirname $CEPH_ENV)/secrets.yml.tpl and fails closed if either file is missing. See scripts.md for the full contract.

Destroy task in .mise.toml extracts the domain for the safety gate:

CLUSTER_ID=$(basename "$CEPH_ENV_DIR")        # sietch-ceph.dev.austin.int
DOMAIN=${CLUSTER_ID#*-ceph.}.futo.cloud       # dev.austin.int.futo.cloud

Placement group logic

MON placement strategy varies by cluster size. A 2-node cluster should NOT run 2 MONs (no quorum majority possible). A 3+ node cluster should run MON on all nodes.

roles/ceph_deploy/tasks/placement.yml:

- name: Calculate MON placement
  ansible.builtin.set_fact:
    mon_hosts: >-
      {%- if groups['ceph_mon'] | default([]) | length > 0 -%}
      {{ groups['ceph_mon'] | map('extract', hostvars, 'hostname_short') | join(',') -}}
      {%- elif active_hosts.stdout.split(',') | length <= 2 -%}
      {{ hostvars[groups['ceph_bootstrap'][0]]['hostname_short'] -}}
      {%- else -%}
      {{ active_hosts.stdout -}}
      {%- endif -%}

Three branches: (1) explicit ceph_mon inventory group → use those; (2) ≤ 2 active hosts → single MON (bootstrap only, avoids 2-MON quorum fragility); (3) 3+ active hosts → MON on all active hosts.

MGR always deploys on all active hosts — standby MGRs are harmless and provide fast failover.


Shell + changed_when discipline

Two rules worth stating explicitly because they're the most common lint-clean failures:

  1. changed_when: false for read-only commands (checks, queries, status).

  2. changed_when: true only when guarded by when: — the task only runs when something needs to change. Unguarded changed_when: true reports "changed" on every run; ansible-lint catches this.

  3. Output-based changed_when for shell tasks that may or may not change state:

    changed_when: "'CHANGED' in zonegroup_hostnames.stdout"
    
  4. failed_when with fallthrough for commands where a specific error is expected and acceptable:

    failed_when:
      - pg_data_result.rc != 0
      - "'is not >= current' not in pg_data_result.stderr | default('')"
    

Every shell/command task in this codebase sets one of these — no bare command: without a changed_when. Lint enforces it.

no_log: true on every task handling passwords, keys, or credentials. Ansible output is committed to ansible.log and displayed to operators — secrets must never land there.


Cephadm service specs over imperative loops

Problem: the role's first instinct is "iterate every disk / daemon / service in Ansible and run cephadm per item." This couples the role tightly to per-host hardware shape (path composition, partition layout, LV vs disk topology) and breaks on any new cluster shape — the original sietch-shape osds.yml failed on painbox's PCI-ATA + LV-backed SSD OSD topology because sas_path_prefix was hardcoded into the path composition.

Pattern: for any cephadm-managed surface (OSDs, RGW, MON/MGR placement, monitoring), render a declarative service spec and apply it via ceph orch apply -i <spec>.yaml. Cephadm handles per-disk discovery, daemon lifecycle, encryption, LVM, etc. internally. The role becomes a thin renderer + applier; hardware shape moves into the template's Jinja conditional, not the role logic.

Examples in this codebase:

  • templates/rgw-spec.yaml.j2 + tasks/rgw.yml's ceph orch apply -i task — RGW daemon placement spec.
  • templates/osd-spec.yml.j2 + tasks/osds.yml's ceph orch apply osd -i task — OSD service spec; per-host documents in a multi-doc YAML; Jinja conditional handles sietch-shape vs painbox-shape path composition.

Template skeleton:

{% for host in groups['ceph_nodes'] %}
{% set h = hostvars[host] %}
---
service_type: <kind>
service_id: {{ h.hostname_short }}-<role>
placement:
  hosts: [{{ h.hostname_short }}]
spec:
  <kind-specific fields>
{% endfor %}

Apply pattern in tasks/*.yml:

- name: Render cephadm service spec
  ansible.builtin.template:
    src: <kind>-spec.yml.j2
    dest: /etc/ceph/<kind>-spec.yml
    mode: '0644'
  delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
  run_once: true

- name: Apply cephadm service spec
  ansible.builtin.command: ceph orch apply <kind> -i /etc/ceph/<kind>-spec.yml
  delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
  run_once: true
  changed_when: ...

Use when: the surface you're managing is a cephadm-orch-supported service type (host, mon, mgr, osd, rgw, mds, nfs, prometheus, grafana, alertmanager, node-exporter, ceph-exporter, etc.). Don't use for surfaces cephadm doesn't manage declaratively (CRUSH rules, pools, ceph config tunables, RGW realm/zone setup, S3 user creation) — those still need imperative ceph / radosgw-admin calls.

Trade-off vs imperative loops: debugging "why isn't this disk becoming an OSD?" is harder — there's no per-disk log line. Check ceph orch ls / ceph orch ps / ceph cephadm osd activate <host> --dry-run instead. Worth the trade-off because the role becomes hardware-shape-agnostic.

See ADR-011 for the full decision record on the OSD-path migration.


Anti-patterns

changed_when: true without a when guard

# BAD — reports changed on every run even when idempotent
- ansible.builtin.command: ceph config set osd foo bar
  changed_when: true

# GOOD — only runs when needed, so changed_when: true is accurate
- ansible.builtin.command: ceph config set osd foo bar
  when: current_foo != 'bar'
  changed_when: true

Shell without pipefail

# BAD — if `ceph osd dump` fails, grep runs on empty input and task succeeds
- ansible.builtin.shell: ceph osd dump | grep noin

# GOOD — pipefail propagates the ceph failure
- ansible.builtin.shell: |
    set -o pipefail
    ceph osd dump | grep noin
  args:
    executable: /bin/bash

Hardcoded site-specific values in roles

# BAD — in a role's tasks/main.yml
- ansible.builtin.command: ceph config set osd osd_recovery_max_active 1

# GOOD — value comes from defaults, overridable via group_vars
- ansible.builtin.command: >
    ceph config set osd osd_recovery_max_active
    {{ ceph_osd_recovery_max_active }}

Roles use defaults/main.yml for all tunables. Site-specific values live in inventories/<cluster>/group_vars/all/vars.yml.

Using ansible_play_batch / ansible_play_hosts for placement specs

# BAD — --limit shrinks play_batch, cephadm removes daemons from omitted hosts
- ansible.builtin.command:
    ceph orch apply rgw --placement="{{ ansible_play_batch | join(',') }}"

cephadm is declarative: applying a smaller placement list REMOVES daemons from hosts not in the list. Always use groups['ceph_nodes'] (the full inventory group) for placement specs, never ansible_play_batch or ansible_play_hosts. The comment at roles/ceph_deploy/tasks/rgw.yml:466 explains the failure mode in detail.

Running plays with bare ansible-playbook

Every playbook in this project consumes op-injected secrets. Running ansible-playbook foo.yml directly skips the wrapper, op inject never runs, and vault_* variables are empty — tasks that need them fail with confusing errors. Always scripts/ansible-play.sh <playbook.yml>. See scripts.md for the full contract.