feat(ceph): import yucca-ceph ansible + terraform infrastructure (#86)

* feat(ceph): import yucca-ceph ansible + terraform infrastructure

Imports the yucca-ceph Ansible tree into ansible/ceph/ and adds the
Terraform stack at tf/ that drives it. Cuts over from ansible-vault
to the hybrid secrets architecture (TF as inventory authority, 1P
as secrets store, op-inject at deploy time) in one atomic move.
Source: internal yucca-ceph working tree; fresh subtree-style
import, history not preserved. Andy continues operating sietch +
painbox post-merge; yucca-team hosts the code and reviews changes.

What it adds:
  - sietch (3-node Austin, production Ceph S3 backend, untouched
    by this PR)
  - painbox (single-node Hetzner SX295 in Helsinki) as a second
    deployable cluster
  - Future clusters land by appending to clusters.auto.tfvars in
    the matching environment stack (tf/deployment/<env>/ceph/) —
    no per-cluster TF code required

How it works (full map: ansible/ceph/docs/architecture.md):
  - tf/shared/modules/ceph-cluster renders inventory.ini variants
    + secrets.yml.tpl per cluster from clusters.auto.tfvars
  - secrets.yml.tpl carries op:// refs; `op inject -f` resolves
    them at deploy time from the matching yucca_tf_<env> vault
  - State in OVH yucca-tf-state bucket (key ceph/<env>/<stack>/)
  - 11 ADRs capture the decisions: ansible/ceph/docs/adr/

Out of scope (intentional):
  - LUKS keys not yet in 1P (deferred until hybrid is stable)
  - tf/shared/modules/ceph-cluster/secrets.tf.disabled is dormant;
    today's 1P items via `op item create` per
    ansible/ceph/docs/adding-a-cluster.md
  - Talos K8s on sietch is a separate workstream

Atomicity + rollback: TF-rendered inventory + secrets-template
files are gitignored (TF generates them) and ansible-vault removal
is coupled to the op-inject path. Splitting this PR lands in a
non-bootable state — merge as one unit. The merge itself is
reversible via `git revert` until the post-merge `tf:apply` runs;
after apply, full rollback needs state restore or `tofu state mv`
(land + validate before applying).

Dev-env impact: adds opentofu + terragrunt to yucca root mise tools
plus a self-contained ansible/ceph/.mise.toml. No new commands or
prereqs for immich-side contributors who don't touch ceph or run
tf:* tasks.

Verification:
  - `mise run lint` (from ansible/ceph/): 130 files, 0 warnings
  - `mise run check`: 19 playbooks parse clean
  - `mise run tf:plan`: succeeds; 7 expected file path-rename
    replacements (3 painbox + 4 sietch). State drift from import,
    no cluster-side change.
  - painbox deployed 2026-04-26 on the new code path: Bookworm +
    Ceph Tentacle, 15 OSDs (14 HDD + 1 SSD) up + in, mon/mgr/rgw
    running. HEALTH_WARN is expected on a single-node cluster.

Post-merge: from the yucca root, `mise run tf:apply` flips the
bucket state to the new monorepo paths (the 7 renames above).

* fix(ceph): exempt ansible/ and tf/ subtrees from root prettier

The imported infrastructure subtrees enforce their own format
conventions (yamllint + ansible-lint inside ansible/ceph/; tofu fmt
inside tf/). Prettier on ansible YAML reflows long Jinja2 expressions
and shell command blocks in unwanted ways, so root prettier checks
are skipped for both subtrees.

Also reformat root README.md table column alignment to match prettier
conventions (only the imported subtrees are exempt; yucca-side files
including the root README still follow root prettier rules).

* fix(ceph): clean up secrets tmpfile after ansible-playbook exits

`ansible-play.sh` rendered the resolved secrets file via `op inject`
into a `mktemp` tmpfile, set up a `trap 'rm -f "$TMPFILE"' EXIT INT
TERM`, then `exec`'d ansible-playbook. The `exec` replaced the bash
shell entirely, so the EXIT trap never fired — every play left a
plaintext-secrets file in /tmp.

In practice this was masked because /tmp is tmpfs (RAM only on this
operator's setup), so files evaporate on reboot. But within an
operator session, files accumulated linearly with each playbook
invocation. Recent count on the import-PR session: 38 files.

Drop the `exec`. With `set -euo pipefail` already on, bash:

  - propagates ansible-playbook's exit code (set -e)
  - fires the EXIT trap before exiting (always)
  - cleans up the tmpfile on success, failure, or signal

Verified: `CEPH_ENV=... scripts/ansible-play.sh status.yml
--syntax-check` creates and removes the tmpfile within the same
invocation — /tmp is clean before and after.

`scripts/preflight.sh` uses the same trap pattern but does not
`exec`, so its tmpfile cleanup was already correct (and the suffix
differs: `-secrets-test.yml` vs `-secrets.yml`, confirming
ansible-play.sh as the sole offender).
This commit is contained in:
Andy Molenda
2026-05-18 06:17:56 -07:00
committed by GitHub
parent e703b08b83
commit 63087f6850
192 changed files with 17080 additions and 2 deletions
+25
View File
@@ -4,7 +4,32 @@ node_modules
.env.local
.env
# Exception: tf/.env is committed — contains only op:// references to 1P items,
# no literal secrets. Resolved at runtime by `op run --env-file=tf/.env -- ...`.
!tf/.env
mise.local.toml
dist/
packages/michael/michael
# OpenTofu / Terraform
# .terraform.lock.hcl IS committed for reproducibility
**/.terraform/
**/terraform.tfstate
**/terraform.tfstate.*
**/*.tfvars.local
# terragrunt-generated backend.tf contains absolute per-operator paths
**/backend.tf
# Ansible runtime artifacts
ansible/*/ansible.log
ansible/*/ansible.*.log
ansible/*/.ansible/
ansible/*/.ansible_facts_cache/
ansible/*/.venv/
# Lens / decision-support artifacts are controller-local notes, not repo code.
# Per-project: kept outside the repo (e.g., ~/Projects/immich/yucca-ceph-import/analysis/
# on operator workstation, but not tracked).
analysis/
+29
View File
@@ -7,6 +7,10 @@ restic = "0.18.0"
gh = "2.25.0"
"github:git-town/git-town" = "22.4.0"
# Infrastructure tooling (added for tf/ and ansible/ subtrees)
opentofu = "1.11.5"
terragrunt = "0.99.4"
[tasks.dev]
description = "Start all services in development mode"
depends = ["install:deps", "common:build", "docker:start"]
@@ -44,6 +48,31 @@ run = [{ tasks = ["lint", "*:check"] }, { task = "format" }, { task = "test" }]
description = "Run all possible code fixes"
run = [{ task = "*:fix" }, { task = "web:lingui" }]
# ─── Infrastructure (tf + ansible) ──────────────────────────────────────────
# All tf:* tasks wrap terragrunt with `op run --env-file=tf/.env` so the
# OP_SERVICE_ACCOUNT_TOKEN is injected from 1Password at invocation time.
# No literal secrets in the .env file — just op:// references.
[tasks."tf:init"]
description = "Terragrunt init for a given stack (default: deployment/dev/ceph)"
run = "op run --env-file=tf/.env -- terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/dev/ceph} init"
[tasks."tf:plan"]
description = "Terragrunt plan for a given stack"
run = "op run --env-file=tf/.env -- terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/dev/ceph} plan"
[tasks."tf:apply"]
description = "Terragrunt apply for a given stack"
run = "op run --env-file=tf/.env -- terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/dev/ceph} apply"
[tasks."tf:destroy"]
description = "Terragrunt destroy for a given stack (use with care)"
run = "op run --env-file=tf/.env -- terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/dev/ceph} destroy"
[tasks."tf:fmt"]
description = "Format terraform + terragrunt files recursively"
run = "tofu fmt -recursive tf/ && terragrunt hcl format --working-dir tf/"
[env]
NODE_ENV = "development"
LOG_LEVEL = "debug"
+7
View File
@@ -23,3 +23,10 @@ yarn.lock
**/locales
**/fetch-client.ts
**/openapi-specs.json
# Infrastructure subtrees own their own format conventions (yamllint +
# ansible-lint inside ansible/ceph/; tofu fmt inside tf/). Prettier on
# ansible YAML is opinionated in ways that conflict with Jinja2 + ansible
# task structure, so these subtrees are exempt from root prettier checks.
ansible/
tf/
+24 -1
View File
@@ -1,4 +1,11 @@
## Development Guide
# Yucca
Application code lives under `packages/`. Infrastructure that operates Yucca
(Ceph storage backend, future Talos K8s, deployment/state managed via
Terraform+1Password) lives at the top level in `ansible/`, `tf/`, and
`kubernetes/`.
## Development Guide (application)
Ensure you have prerequisites installed:
@@ -20,3 +27,19 @@ mise test:integration # integration tests
mise test:e2e # e2e tests
mise test:e2e:web # e2e web tests
```
## Infrastructure
| Start here | Path | Purpose |
| ---------------------------------------------------- | --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- |
| [`ansible/ceph/README.md`](./ansible/ceph/README.md) | `ansible/ceph/` | Ansible automation for Ceph clusters (sietch, painbox). Deploys + operates via cephadm on bare-metal and Hetzner. |
| [`tf/README.md`](./tf/README.md) | `tf/` | Terraform/OpenTofu authority for cluster identity, 1P secret items, rendered Ansible inventories. Terragrunt multi-env (`deployment/<env>/<stack>/`). |
| (coming in follow-up) | `kubernetes/` | Flux GitOps surface for the (future) Talos K8s cluster. |
**Secrets are managed via the `yucca_tf_*` 1Password vaults.** Runtime reads use a
read-only service account; TF writes use a superuser service account. See
`ansible/ceph/docs/secrets.md` and `tf/README.md` for the full model.
`mise run tf:init / tf:plan / tf:apply` wraps terragrunt via
`op run --env-file=tf/.env --` so the superuser token is injected from 1P at
invocation time.
+11
View File
@@ -0,0 +1,11 @@
---
# All roles share the ceph_ variable prefix for consistency across the
# cluster stack. The var-naming[no-role-prefix] rule expects each role
# to use its own prefix, which doesn't fit our single-product layout.
skip_list:
- var-naming[no-role-prefix]
# Molecule verify playbooks use relative paths to test template rendering
# outside of role context. This is intentional and expected.
exclude_paths:
- roles/*/molecule/
+29
View File
@@ -0,0 +1,29 @@
root = true
[*]
end_of_line = lf
insert_final_newline = true
trim_trailing_whitespace = true
charset = utf-8
[*.{yml,yaml}]
indent_style = space
indent_size = 2
[*.{j2,jinja2}]
indent_style = space
indent_size = 2
[*.py]
indent_style = space
indent_size = 4
[*.{sh,bash}]
indent_style = space
indent_size = 2
[Makefile]
indent_style = tab
[*.md]
trim_trailing_whitespace = false
+40
View File
@@ -0,0 +1,40 @@
# Ansible
ansible.log
ansible.*.log
*.retry
.ansible/
.ansible_facts_cache/
# Transient outputs — benchmark results, hardware inventory snapshots, backups, exports
bench/
hardware/
backups/
exports/
# TF-generated artifacts — do not commit; re-render with `tofu apply` in tf/deployment/<env>/ceph/
inventories/*/inventory.ini
inventories/*/inventory-provision.ini
inventories/*/inventory-destroy.ini
inventories/*/secrets.yml.tpl
# host_vars IS committed (per-node hardware facts — stable, part of the inventory
# source of truth). Operator-local overrides use host_vars/*.local.yml if needed.
inventories/*/host_vars/*.local.yml
# Rendered installimage scripts (source of truth is the .tpl file).
inventories/*/installimage/post-install.sh
# Python virtualenv (mise setup) and bytecode
.venv/
__pycache__/
*.pyc
# Ansible navigator artifacts
.artifacts/
ansible-navigator.log
# OS / editor
.DS_Store
*.swp
*.swo
*~
+173
View File
@@ -0,0 +1,173 @@
[tools]
python = "3.12"
[env]
# Project-local virtualenv for all Python tooling
VIRTUAL_ENV = "{{config_root}}/.venv"
PATH = "{{config_root}}/.venv/bin:{{env.PATH}}"
# Note: CEPH_ENV is intentionally NOT declared here. mise's [env] block
# overrides shell-exported values, which silently sends operators to the
# wrong cluster. Operators must export CEPH_ENV once per shell session:
# export CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory.ini
# export CEPH_ENV=inventories/painbox-ceph.dev.hel.htz/inventory.ini
# Inventory files are TF-generated — `mise run tf:apply` (from yucca root) if missing.
[tasks.setup]
description = "Bootstrap development environment"
run = """
#!/usr/bin/env bash
set -euo pipefail
echo "=== Creating virtualenv ==="
python -m venv .venv
source .venv/bin/activate
echo "=== Installing Python dependencies ==="
pip install -q -r requirements.txt
echo "=== Installing Ansible collections ==="
ansible-galaxy collection install -r requirements.yml
echo "=== Verifying ==="
ansible --version | head -1
ansible-lint --version | head -1
yamllint --version
echo ""
echo "Environment ready. Run 'mise trust' if prompted."
echo "Next: run 'mise run tf:apply' (from the yucca root) to render inventories."
"""
[tasks.lint]
description = "Run all linters (yamllint + ansible-lint + shellcheck)"
run = """
#!/usr/bin/env bash
set -euo pipefail
echo "=== yamllint ==="
yamllint roles/ *.yml
echo ""
echo "=== ansible-lint ==="
ansible-lint
echo ""
echo "=== shellcheck ==="
if command -v shellcheck &>/dev/null; then
shellcheck scripts/*.sh && echo "All scripts pass"
else
echo "shellcheck not installed — skipping (install: pacman -S shellcheck)"
fi
"""
[tasks.check]
description = "Syntax-check all playbooks (no 1P access required)"
run = """
#!/usr/bin/env bash
set -euo pipefail
# Syntax-check only parses YAML — any valid inventory works. Default to
# sietch when CEPH_ENV isn't inline-prefixed; the parse is identical.
CEPH_ENV="${CEPH_ENV:-inventories/sietch-ceph.dev.austin.int/inventory.ini}"
for pb in *.yml; do
case "$pb" in
requirements.yml|ansible-navigator.yml) continue ;;
*)
echo "Checking $pb..."
ansible-playbook -i "$CEPH_ENV" --syntax-check "$pb" 2>&1 | tail -1
;;
esac
done
"""
[tasks.test]
description = "Run Molecule tests for roles with test scenarios"
run = """
#!/usr/bin/env bash
set -euo pipefail
for role in roles/*/molecule; do
ROLE_DIR=$(dirname "$role")
ROLE_NAME=$(basename "$ROLE_DIR")
echo "=== Testing $ROLE_NAME ==="
(cd "$ROLE_DIR" && molecule verify)
done
"""
[tasks.preflight]
description = "Pre-flight checks (TF artifacts, 1P session, SSH, connectivity)"
run = "scripts/preflight.sh"
[tasks.status]
description = "Quick cluster health check (read-only)"
run = "scripts/ansible-play.sh status.yml"
[tasks.drift]
description = "Detect configuration drift from expected state"
run = "scripts/ansible-play.sh drift.yml"
[tasks.deploy]
description = "Full deploy: baseline → tune → deploy → tune ceph → harden"
run = """
#!/usr/bin/env bash
set -euo pipefail
# Rotate ansible.log before a full deploy
if [ -f ansible.log ] && [ "$(wc -c < ansible.log)" -gt 1048576 ]; then
mv ansible.log "ansible.$(date +%Y%m%dT%H%M%S).log"
echo "Rotated ansible.log (>1MB)"
fi
echo "=== Baseline ===" && scripts/ansible-play.sh baseline.yml
echo "=== OS tuning ===" && scripts/ansible-play.sh tune-os.yml
echo "=== Hardware tuning ===" && scripts/ansible-play.sh tune-hardware.yml
echo "=== Deploy Ceph ===" && scripts/ansible-play.sh deploy-ceph.yml
echo "=== Ceph tuning ===" && scripts/ansible-play.sh tune-ceph.yml
echo "=== Harden ===" && scripts/ansible-play.sh harden.yml
"""
[tasks.destroy]
description = "Destroy Ceph cluster (requires confirmation)"
run = """
#!/usr/bin/env bash
set -euo pipefail
CEPH_ENV_DIR=$(dirname "$CEPH_ENV")
CLUSTER_ID=$(basename "$CEPH_ENV_DIR") # e.g., sietch-ceph.dev.austin.int
DOMAIN=${CLUSTER_ID#*-ceph.}.futo.cloud # e.g., dev.austin.int.futo.cloud
DESTROY_INV="$CEPH_ENV_DIR/inventory-destroy.ini"
echo "Usage: scripts/ansible-play.sh destroy-ceph.yml \\"
echo " -e yes_destroy_ceph=true -e destroy_target_domain=$DOMAIN"
echo " (with CEPH_ENV=$DESTROY_INV if destroying as root)"
echo ""
read -rp "Run destroy now? [y/N] " confirm
if [[ "$confirm" =~ ^[Yy]$ ]]; then
CEPH_ENV="$DESTROY_INV" scripts/ansible-play.sh destroy-ceph.yml \
-e yes_destroy_ceph=true \
-e destroy_target_domain="$DOMAIN"
fi
"""
[tasks.bench]
description = "S3 benchmark (parallel PUT/GET against local RGW)"
run = "scripts/ansible-play.sh bench.yml"
[tasks.bench-rados]
description = "RADOS bench (raw cluster I/O, bypasses RGW)"
run = "scripts/ansible-play.sh rados-bench.yml"
[tasks.backup]
description = "Export Ceph cluster config for disaster recovery"
run = "scripts/ansible-play.sh backup-config.yml"
[tasks.capture]
description = "Snapshot bootstrap secrets (RGW TLS, admin keyring) to 1P (DR belt)"
run = "scripts/ansible-play.sh post-deploy-capture.yml"
[tasks."rotate-certs"]
description = "Rotate RGW TLS cert (regenerate self-signed, restart RGW daemons)"
run = "scripts/ansible-play.sh rotate-certs.yml"
[tasks."rotate-ssh-key"]
description = "Distribute current ansible-iac pubkey from 1P to nodes (forward-only)"
run = "scripts/ansible-play.sh rotate-ssh-key.yml"
[tasks."hardware-inventory"]
description = "Snapshot per-node hardware facts to JSON (informational only)"
run = "scripts/ansible-play.sh hardware-inventory.yml"
[tasks."migrate-networkd"]
description = "One-shot ifupdown→networkd migration (rolling, serial=1, noout-gated)"
run = "scripts/ansible-play.sh migrate-networkd.yml"
+12
View File
@@ -0,0 +1,12 @@
---
extends: relaxed
rules:
line-length:
max: 260
allow-non-breakable-inline-mappings: true
comments:
min-spaces-from-content: 1
octal-values:
forbid-implicit-octal: true
forbid-explicit-octal: true
+435
View File
@@ -0,0 +1,435 @@
# Contributing
Developer guide for the ceph Ansible project in the `yucca` monorepo. Read
this before making changes.
## Prerequisites
| Tool | Version | Install |
|------|---------|---------|
| [mise](https://mise.jdx.dev/) | latest | `curl https://mise.jdx.dev/install.sh \| sh` |
| Python | 3.12 (managed by mise) | Automatic via `.mise.toml` |
| [1Password CLI](https://developer.1password.com/docs/cli) | v2 | `pacman -S 1password-cli` / `brew install 1password-cli` |
| [shellcheck](https://www.shellcheck.net/) | latest | `pacman -S shellcheck` |
| SSH config | Access to cluster nodes | See below |
### 1Password access
You need read access to the **`yucca_tf_dev`** 1Password vault (Futo team
membership grants this). The `scripts/ansible-play.sh` wrapper uses `op
inject` to resolve secrets at playbook time — desktop session unlock or
`OP_SERVICE_ACCOUNT_TOKEN` satisfies auth. No ansible-vault password to
manage.
### SSH setup
The `ansible-iac` SSH keys live in `yucca_tf_dev` as items
`SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY` and `PAINBOX_CEPH_ANSIBLE_IAC_SSH_KEY`
(see [ADR-010](docs/adr/010-ssh-keys-in-1password.md) for rationale).
**First-time workstation setup:**
```bash
cd yucca/ansible/ceph
scripts/install-ssh-keys.sh # op read → ~/.ssh/id_ed25519_{sietch,painbox}
```
The script is idempotent and refuses to overwrite an existing key whose
fingerprint doesn't match 1P. Keys land as `~/.ssh/id_ed25519_sietch` +
`~/.ssh/id_ed25519_painbox` (private + `.pub` both 0600/0644).
**Jump hosts / proxies** belong in your personal `~/.ssh/config`, not in
this repo.
**Recommended:** run the 1Password desktop app's SSH Agent globally
(`Host * IdentityAgent ~/.1password/agent.sock` in `~/.ssh/config`). Keys
are served from 1P; private material never leaves the app. The
inventory's explicit `ansible_ssh_private_key_file` still works as a
fallback.
## First-time setup
```bash
# Clone the monorepo and navigate to this subproject
git clone <yucca-monorepo-url> && cd yucca/ansible/ceph
# Trust the mise config (one-time per checkout)
mise trust
# Bootstrap the dev environment (creates .venv, installs Python deps + Ansible collections)
mise run setup
# Render cluster inventories + secrets templates (run once, or after any
# change to tf/deployment/dev/ceph/clusters.auto.tfvars)
cd ../../tf/deployment/dev/ceph && tofu init && tofu apply && cd -
```
This runs:
1. Creates `.venv/` with Python 3.12
2. `pip install -r requirements.txt` (ansible-core, ansible-lint, yamllint, molecule, boto3)
3. `ansible-galaxy collection install -r requirements.yml` (ansible.posix, community.general)
Verify:
```bash
ansible --version | head -1 # ansible-core 2.20.x
ansible-lint --version # 26.x
yamllint --version # 1.38.x
```
## Selecting a cluster
All tooling uses the `CEPH_ENV` variable to select the target cluster.
It points to an **inventory file** (not a directory):
```bash
# Inline prefix — required for `mise run` invocations:
CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory.ini mise run preflight
CEPH_ENV=inventories/painbox-ceph.dev.hel.htz/inventory.ini mise run status
```
**`export CEPH_ENV=...` does NOT work with `mise run`.** mise's `[env]`
block strips shell-exported vars when launching tasks; the wrapper exits
with `CEPH_ENV must be set` even though your shell clearly has it set.
The inline-prefix form passes the var directly into mise's invocation
env where it's preserved. See [docs/scripts.md "Setting CEPH_ENV"](docs/scripts.md)
for the full explanation.
For multiple commands against the same cluster, set a local (non-exported)
shell variable and inline-prefix each invocation:
```bash
CE=inventories/painbox-ceph.dev.hel.htz/inventory.ini
CEPH_ENV=$CE mise run preflight
CEPH_ENV=$CE mise run status
CEPH_ENV=$CE mise run deploy
```
Calling scripts directly (e.g., `scripts/preflight.sh`) DOES respect
shell `export` — it's only the `mise run` path that filters the env.
Cluster identity is declared in `tf/deployment/dev/ceph/clusters.auto.tfvars`
(keyed by short cluster name). TF renders the directory name, inventory
file, and secrets template from that entry. `CEPH_ENV` is just a pointer
to the rendered inventory file; wrappers like `scripts/ansible-play.sh`
and the `destroy` mise task extract the cluster name from its path for
convenience:
```
CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory.ini
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
inventory dir = sietch-ceph.dev.austin.int (rendered by TF)
cluster name = sietch (map key in clusters.auto.tfvars)
domain = dev.austin.int.futo.cloud (domain field in tfvars)
```
Running `mise run tf:apply` regenerates `inventories/<cluster>/inventory.ini`
and `secrets.yml.tpl` any time the tfvars entry changes.
## Development workflow
### 1. Edit
Make your changes in `roles/`, playbooks, or inventory files.
### 2. Lint
```bash
mise run lint
```
Runs three linters in sequence:
- **yamllint** -- relaxed profile, 260-char line limit (`.yamllint`)
- **ansible-lint** -- skips `var-naming[no-role-prefix]` because all roles share
the `ceph_` prefix (`.ansible-lint`)
- **shellcheck** -- all scripts in `scripts/`
### 3. Syntax-check
```bash
mise run check
```
Runs `ansible-playbook --syntax-check` against every playbook using the active
`CEPH_ENV` inventory.
### 4. Re-render inventory (TF)
```bash
mise run tf:apply
```
Re-renders `inventory.ini` and `secrets.yml.tpl` for every cluster declared
in `tf/deployment/dev/ceph/clusters.auto.tfvars`. Required after any cluster-
spec edit.
### 5. Dry-run against real nodes
```bash
scripts/ansible-play.sh baseline.yml --check --diff
```
`--check` simulates changes without applying them. `--diff` shows what would
change. Safe to run against production. Use the wrapper (not bare
`ansible-playbook`) so secrets resolve via `op inject`.
### 6. Test with Molecule
```bash
mise run test
```
Roles with `molecule/` directories get template-rendering verification.
Molecule doesn't converge (requires real hardware) but verifies Jinja2
templates render without errors.
## Code conventions
### FQCN everywhere
Always use fully-qualified collection names for modules:
```yaml
# Good
- name: Install packages
ansible.builtin.apt:
name: htop
state: present
# Bad
- name: Install packages
apt:
name: htop
```
### changed_when is required
Every `shell` and `command` task must declare `changed_when`:
```yaml
- name: Check OSD count
ansible.builtin.shell: |
set -o pipefail
ceph osd stat --format json | python3 -c "..."
args:
executable: /bin/bash
register: osd_count
changed_when: false # read-only command
- name: Create OSD
ansible.builtin.shell: ...
changed_when: "'Created osd' in result.stdout"
```
### Shell task rules
- Always set `args.executable: /bin/bash`
- Always start with `set -o pipefail` (or `set -euo pipefail` for multi-line)
- Use `>` folded style for single-line commands, `|` literal style for
multi-line:
```yaml
# Single logical command
- name: Check disk
ansible.builtin.shell: >
set -euo pipefail;
pvs /dev/sda 2>/dev/null | grep -q ceph
# Multi-line script
- name: Wait for OSDs
ansible.builtin.shell: |
set -o pipefail
ceph osd stat --format json
args:
executable: /bin/bash
```
### Secrets handling
- Use `no_log: true` on any task that handles passwords, keys, or tokens
- Secrets are provisioned in 1Password (see [docs/secrets.md](docs/secrets.md))
and consumed at playbook time via `scripts/ansible-play.sh`, which runs
`op inject` on the cluster's `secrets.yml.tpl` and passes resolved values
as `--extra-vars`.
- In ansible code, reference secrets as regular variables via the existing
alias pattern in each cluster's `group_vars/all/vars.yml`:
```yaml
# group_vars/all/vars.yml (plaintext, committed)
ops_password: "{{ vault_ops_password }}"
```
`vault_ops_password` is populated from `op inject` — no ansible-vault, no
encrypted file in git.
- `secrets.yml.tpl` is TF-generated and gitignored; don't edit it by hand.
Add new secrets via `tf/shared/modules/ceph-cluster/main.tf` (the
`local.secrets` map), then `tofu apply` to re-render.
### Handler pattern
Define handlers in `roles/<role>/handlers/main.yml`:
```yaml
---
- name: Restart sshd
ansible.builtin.systemd:
name: ssh
state: restarted
```
Trigger with `notify`:
```yaml
- name: Update SSH config
ansible.builtin.copy:
src: sshd_config
dest: /etc/ssh/sshd_config
notify: Restart sshd
```
### Variable naming
- All cluster-level variables use the `ceph_` prefix
- Role-specific internal variables use `<role_name>_` prefix (e.g.,
`baseline_ops_user`, `baseline_podman_packages`)
- snake_case for everything
- Document every variable in `defaults/main.yml` with a comment explaining
what it does and what valid values look like
### Role names
- snake_case: `ceph_deploy`, `os_tuning`, `hardware_tuning`
- Not: `ceph-deploy`, `CephDeploy`, `cephDeploy`
### Tags
Use tags on `import_tasks` in the role's `main.yml` to allow selective runs:
```yaml
- name: Phase 1 - Prerequisites
ansible.builtin.import_tasks: prerequisites.yml
tags: [prerequisites]
- name: Phase 2 - Bootstrap cluster
ansible.builtin.import_tasks: bootstrap.yml
tags: [bootstrap]
```
Run a specific phase:
```bash
scripts/ansible-play.sh deploy-ceph.yml --tags bootstrap
```
## mise tasks reference
| Task | Command | Description |
|------|---------|-------------|
| `setup` | `mise run setup` | Bootstrap dev environment (venv, pip, galaxy) |
| `lint` | `mise run lint` | yamllint + ansible-lint + shellcheck (no 1P required) |
| `check` | `mise run check` | Syntax-check all playbooks (no 1P required) |
| `test` | `mise run test` | Molecule verify for roles with test scenarios |
| `preflight` | `mise run preflight` | TF artifacts + 1P session + SSH + connectivity |
| `status` | `mise run status` | Read-only cluster health check (via ansible-play.sh) |
| `drift` | `mise run drift` | Detect configuration drift (via ansible-play.sh) |
| `deploy` | `mise run deploy` | Full deploy pipeline (via ansible-play.sh) |
| `backup` | `mise run backup` | Export cluster config for DR (via ansible-play.sh) |
| `capture` | `mise run capture` | Snapshot RGW TLS + admin keyring to 1P (DR belt) |
| `bench-rados` | `mise run bench-rados` | RADOS bench (raw cluster I/O) |
| `destroy` | `mise run destroy` | Destroy cluster (interactive confirmation) |
Inventory rendering and secret-item management are TF responsibilities
— `tofu apply` in `tf/deployment/dev/ceph/` renders `inventory.ini` and
`secrets.yml.tpl` for every cluster declared in `clusters.auto.tfvars`.
Cluster secrets live in `yucca_tf_dev` (see [docs/secrets.md](docs/secrets.md)).
## File organization
```
yucca/
├── tf/ # Terraform state + secrets + rendering (authoritative)
│ ├── shared/modules/ceph-cluster/ # Module: per-cluster orchestration + rendering
│ └── deployment/dev/ceph/ # Cluster declarations + tofu apply target
└── ansible/ceph/ # This directory
├── *.yml # Top-level playbooks (site.yml, deploy-ceph.yml, etc.)
├── inventories/
│ └── <cluster>-ceph.<env>.<dc>.<provider>/
│ ├── inventory.ini # TF-rendered (gitignored)
│ ├── secrets.yml.tpl # TF-rendered, consumed by op inject (gitignored)
│ ├── group_vars/all/
│ │ └── vars.yml # Cluster-wide variables (plaintext, committed)
│ └── host_vars/
│ ├── <hostname>.yml # Per-node hardware config (committed)
│ └── <hostname>.local.yml # Per-operator overrides (gitignored)
├── roles/
│ └── <role_name>/
│ ├── defaults/main.yml # Default variables (documented)
│ ├── meta/main.yml # Role metadata + dependencies
│ ├── tasks/main.yml # Entry point (imports sub-task files)
│ ├── handlers/main.yml # Service restart handlers
│ ├── templates/*.j2 # Jinja2 templates
│ └── molecule/default/ # Test scenario (optional)
├── scripts/
│ ├── ansible-play.sh # Wrapper: op inject + ansible-playbook
│ ├── install-ssh-keys.sh # Pull ansible-iac private keys from 1P → ~/.ssh/
│ └── preflight.sh # Pre-deploy checks
├── docs/ # Operational documentation
│ ├── runbooks/ # Step-by-step operational procedures
│ └── archive/ # Historical deployment notes
├── .mise.toml # Task runner + Python version + CEPH_ENV
├── ansible.cfg # Ansible config (SSH, caching, output)
├── requirements.txt # Python dependencies
└── requirements.yml # Ansible Galaxy collections
```
### What goes where
- **`roles/`** -- Reusable automation. Each role handles one concern (baseline
setup, Ceph deploy, OS tuning, security). Roles never reference a specific
cluster -- they use variables from inventory.
- **`inventories/`** -- Cluster-specific data. IPs, hardware mappings, network
config, secrets. Each cluster is fully self-contained in its directory.
- **`scripts/`** -- Controller-side tooling. Things that run on your workstation,
not on target nodes. Secret management, name generation, validation.
- **Top-level `*.yml`** -- Playbooks that wire roles to hosts. Each playbook is
a thin wrapper: set `hosts`, `become`, and import a role.
## Commit messages
Use [conventional commits](https://www.conventionalcommits.org/) for future CI
compatibility:
```
feat(ceph_deploy): add RGW virtual-hosted bucket support
fix(os_tuning): correct TCP buffer sizes for 10GbE
chore(deps): bump ansible-core to 2.20.4
docs(runbooks): add disk replacement procedure
refactor(baseline): split packages into podman and diagnostics
```
Types: `feat`, `fix`, `chore`, `docs`, `refactor`, `test`, `ci`
Scope: role name, `scripts`, `inventory`, `deps`, or omit for repo-wide changes.
## Playbook execution order
The full deploy pipeline (`mise run deploy` or `site.yml`) runs in this order:
```
1. baseline.yml -- ops user, packages, /etc/hosts, services
2. tune-os.yml -- sysctl, TCP buffers, ulimits
3. tune-hardware.yml -- I/O scheduler, readahead, queue depth
4. deploy-ceph.yml -- cephadm bootstrap, join, placement, OSDs, RGW
5. tune-ceph.yml -- recovery throttles, scrub window, telemetry
6. harden.yml -- nftables firewall, SSH hardening
```
Provisioning (`provision.yml`) is a separate concern that runs before this
pipeline on bare-metal live images.
## Editor config
The `.editorconfig` enforces:
- YAML/Jinja2: 2-space indent
- Python: 4-space indent
- Shell: 2-space indent
- All files: UTF-8, LF line endings, trailing whitespace trimmed
+150
View File
@@ -0,0 +1,150 @@
# ceph — Ceph Tentacle on Bare Metal
Ansible automation for provisioning, deploying, tuning, hardening, and
operating Ceph Tentacle (v20) clusters on bare-metal hardware via cephadm.
Lives in the `yucca` monorepo at `ansible/ceph/`; secrets and inventory
scaffolding are provisioned from `yucca/tf/` (see `../../tf/`).
| Cluster | Domain | Location | Hardware | Nodes |
|---------|--------|----------|----------|-------|
| **sietch** | `dev.austin.int.futo.cloud` | Austin DC | Dell R730xd | 3 |
| **painbox** | `dev.hel.htz.futo.cloud` | Hetzner Helsinki | SX295 | 1 |
Clusters are declared in `yucca/tf/deployment/dev/ceph/clusters.auto.tfvars`;
`tofu apply` renders `inventories/<cluster>/inventory.ini` and
`secrets.yml.tpl` per cluster. The `CEPH_ENV` variable selects the active
cluster for any `mise run` or direct ansible invocation.
## Architecture
```mermaid
graph TB
subgraph "Controller (your workstation)"
A[mise + ansible + 1Password CLI]
end
subgraph "Austin DC -- 10.10.10.0/24"
direction TB
L[laurel<br/>MON+MGR+OSD+RGW]
W[lawson<br/>MON+MGR+OSD+RGW]
S[samara<br/>MON+MGR+OSD+RGW]
end
subgraph "Hetzner Helsinki"
P[painbox-ceph-evelyn<br/>MON+MGR+OSD+RGW]
end
A -->|SSH| L
A -->|SSH| W
A -->|SSH| S
A -->|SSH| P
```
See [docs/architecture.md](docs/architecture.md) for role dependencies,
data flow, and design rationale.
## Quick Start
```bash
# 1. Render cluster inventories + secrets templates (once, from yucca/tf/)
(cd ../../tf/deployment/dev/ceph && tofu init && tofu apply)
# 2. Set up the ansible side
mise trust && mise run setup # bootstrap dev environment
# 3. Run mise tasks against the target cluster. CEPH_ENV must be set
# inline (NOT via `export`) — see docs/scripts.md "Setting CEPH_ENV"
# for why mise's [env] block strips shell exports.
CE=inventories/sietch-ceph.dev.austin.int/inventory.ini
CEPH_ENV=$CE mise run preflight # TF artifacts + 1P + SSH + connectivity
CEPH_ENV=$CE mise run status # read-only cluster health check
CEPH_ENV=$CE mise run drift # configuration drift detection
CEPH_ENV=$CE mise run deploy # full pipeline (idempotent)
```
See [CONTRIBUTING.md](CONTRIBUTING.md) for the full development workflow.
## Roles
| Order | Role | Description |
|-------|------|-------------|
| 1 | `provision_host` | Bare-metal Debian 12 install via debootstrap (Austin only) |
| 2 | `baseline` | Post-boot OS baseline: ops user, packages, /etc/hosts, services |
| 3 | `os_tuning` | Kernel sysctl, TCP buffers, optional centralized logging |
| 4 | `hardware_tuning` | I/O scheduler, readahead, udev rules, optional CPU governor |
| 5 | `ceph_deploy` | cephadm bootstrap, join, placement, OSDs, RGW, monitoring |
| 6 | `ceph_tuning` | Recovery throttling, scrub window, CRUSH, telemetry, audit |
| 7 | `security` | nftables firewall, SSH hardening |
| 8 | `ceph_destroy` | Complete cluster teardown (safety-gated) |
| 9 | `s3_bench` | Parallel S3 benchmark against local RGW |
## Playbooks
| Playbook | Description |
|----------|-------------|
| `site.yml` | Full pipeline: baseline + tune + deploy + tune + harden |
| `provision.yml` | Bare-metal provisioning (Austin, `-i` provision inventory) |
| `baseline.yml` | OS baseline (users, packages, hosts) |
| `tune-os.yml` | Kernel/sysctl tuning |
| `tune-hardware.yml` | Disk I/O tuning |
| `deploy-ceph.yml` | Ceph cluster deployment (tags: prerequisites, bootstrap, join, placement, lvm, osds, crush, rgw, monitoring, verify) |
| `tune-ceph.yml` | Post-deploy Ceph config tuning |
| `harden.yml` | Firewall + SSH hardening |
| `destroy-ceph.yml` | Cluster teardown (destroy inventory, requires flags) |
| `status.yml` | Quick health check (read-only) |
| `drift.yml` | Configuration drift detection |
| `bench.yml` | S3 benchmark (RGW round-trip) |
| `rados-bench.yml` | RADOS bench (raw cluster I/O, bypasses RGW) |
| `backup-config.yml` | Export cluster config for DR |
| `post-deploy-capture.yml` | Snapshot RGW TLS + admin keyring to 1P (DR belt) |
| `rotate-certs.yml` | RGW TLS certificate rotation |
| `rotate-ssh-key.yml` | Distribute current ansible-iac pubkey from 1P to nodes |
| `migrate-networkd.yml` | One-shot networkd/bridge migration (rolling, noout-gated) |
| `hardware-inventory.yml` | Hardware facts to JSON |
## mise Tasks
| Task | Description |
|------|-------------|
| `setup` | Bootstrap dev environment (venv, deps, collections) |
| `lint` | yamllint + ansible-lint + shellcheck (no 1P required) |
| `check` | Syntax-check all playbooks (no 1P required) |
| `test` | Molecule tests |
| `preflight` | TF artifacts + 1P session + SSH + connectivity |
| `status` | Cluster health check |
| `drift` | Configuration drift detection |
| `deploy` | Full pipeline |
| `destroy` | Cluster teardown (interactive) |
| `backup` | Export cluster config for DR |
| `capture` | Snapshot RGW TLS + admin keyring to 1P (DR belt) |
| `bench-rados` | RADOS bench (raw cluster I/O) |
Inventory scaffolding + secret-item provisioning live in `yucca/tf/` — run
`tofu apply` in `tf/deployment/dev/ceph/` to (re-)render
`inventories/<cluster>/inventory.ini` and `secrets.yml.tpl`.
## Documentation
| Document | Audience |
|----------|----------|
| [CONTRIBUTING.md](CONTRIBUTING.md) | Developers -- setup, workflow, conventions |
| [docs/architecture.md](docs/architecture.md) | Developers -- role graph, data flow, design |
| [docs/patterns.md](docs/patterns.md) | Developers -- coding idioms, anti-patterns |
| [docs/adding-a-role.md](docs/adding-a-role.md) | Developers -- role skeleton, conventions |
| [docs/adding-a-cluster.md](docs/adding-a-cluster.md) | Developers -- inventory setup, secrets |
| [docs/secrets.md](docs/secrets.md) | Developers/ops -- 1Password integration |
| [docs/naming.md](docs/naming.md) | Everyone -- hostname and inventory naming |
| [docs/hardware.md](docs/hardware.md) | Ops/procurement -- R730xd vs SX295 specs |
| [docs/s3-integration.md](docs/s3-integration.md) | App developers -- endpoints, boto3, certs |
| [docs/security-model.md](docs/security-model.md) | InfoSec -- encryption, users, firewall |
| [docs/capacity-planning.md](docs/capacity-planning.md) | Managers -- costs, formulas, growth |
| [docs/troubleshooting.md](docs/troubleshooting.md) | SRE/on-call -- symptom/diagnosis/fix |
| [docs/runbooks/](docs/runbooks/) | Ops -- add/replace node, replace disk, rotate certs/secrets/SSH/SA token, remote hands, painbox reprovision, bad-tofu-apply recovery, backup/restore |
| [docs/adr/](docs/adr/) | Everyone -- architecture decision records |
## Known Limitations
- **Single-network topology**: public = cluster network on both clusters.
- **Self-signed TLS**: RGW clients need `--no-verify-ssl`. Production needs real certs.
- **`ops` user is password-only**: no SSH keys installed; password sourced from 1P. Intended as an interactive console / recovery account, not for automation.
- **DNS not managed**: `s3.<domain>` and `*.s3.<domain>` records must exist externally.
+21
View File
@@ -0,0 +1,21 @@
---
ansible-navigator:
mode: stdout
playbook-artifact:
enable: true
save-as: .artifacts/{playbook_name}-{time_stamp}.json
ansible:
config:
path: ansible.cfg
logging:
level: warning
file: ansible-navigator.log
color:
enable: true
osc4: true
execution-environment:
enabled: false
+30
View File
@@ -0,0 +1,30 @@
[defaults]
# No default inventory — use -i or CEPH_ENV. Run plays via scripts/ansible-play.sh
# which injects secrets (from secrets.yml.tpl via op inject) as --extra-vars.
log_path = ansible.log
# Performance
forks = 20
gathering = smart
fact_caching = jsonfile
fact_caching_connection = .ansible_facts_cache
fact_caching_timeout = 3600
# Output
stdout_callback = default
result_format = yaml
callbacks_enabled = ansible.posix.timer, ansible.posix.profile_tasks
force_color = True
diff_always = True
deprecation_warnings = False
retry_files_enabled = False
display_skipped_hosts = False
# Security
host_key_checking = False
timeout = 30
[ssh_connection]
# No ProxyJump here — set per-inventory via ansible_ssh_common_args
ssh_args = -o ControlMaster=auto -o ControlPersist=60s -o StrictHostKeyChecking=no
pipelining = True
+159
View File
@@ -0,0 +1,159 @@
---
# Export Ceph cluster configuration for disaster recovery.
# Captures: ceph.conf, admin keyring, CRUSH map, RGW realm config,
# monitor map, and OSD map to the controller's backups/ directory.
#
# Usage:
# scripts/ansible-play.sh backup-config.yml
#
# Outputs: backups/<timestamp>/ on the controller (gitignored)
# Restore: see docs/runbooks/recovery.md
- name: Backup Ceph cluster configuration
hosts: ceph_nodes
become: true
gather_facts: false
vars:
backup_timestamp: "{{ lookup('pipe', 'date +%Y%m%dT%H%M%S') }}"
backup_local_dir: "{{ playbook_dir }}/backups/{{ backup_timestamp }}"
tasks:
- name: Create local backup directory
ansible.builtin.file:
path: "{{ backup_local_dir }}"
state: directory
mode: '0700'
delegate_to: localhost
run_once: true # noqa: run-once[task]
- name: Export cluster configuration
when: inventory_hostname in groups['ceph_bootstrap']
block:
- name: Export ceph.conf
ansible.builtin.command: ceph config generate-minimal-conf
register: ceph_conf_export
changed_when: false
- name: Export admin keyring
ansible.builtin.command: ceph auth get client.admin
register: admin_keyring_export
changed_when: false
- name: Export CRUSH map (decompiled)
ansible.builtin.shell: |
set -o pipefail
ceph osd getcrushmap -o /tmp/crushmap.bin 2>/dev/null
crushtool -d /tmp/crushmap.bin -o /tmp/crushmap.txt
cat /tmp/crushmap.txt
rm -f /tmp/crushmap.bin /tmp/crushmap.txt
args:
executable: /bin/bash
register: crush_export
changed_when: false
- name: Export OSD map summary
ansible.builtin.command: ceph osd dump --format json
register: osd_dump_export
changed_when: false
- name: Export monitor map
ansible.builtin.command: ceph mon dump --format json
register: mon_dump_export
changed_when: false
- name: Export cluster config dump
ansible.builtin.command: ceph config dump --format json
register: config_dump_export
changed_when: false
- name: Export RGW realm configuration
ansible.builtin.shell: |
set -o pipefail
echo '{"realm":' && radosgw-admin realm get 2>/dev/null
echo ',"zonegroup":' && radosgw-admin zonegroup get 2>/dev/null
echo ',"zone":' && radosgw-admin zone get 2>/dev/null
echo '}'
args:
executable: /bin/bash
register: rgw_realm_export
changed_when: false
failed_when: false
- name: Export service specs
ansible.builtin.command: ceph orch ls --format yaml
register: orch_services_export
changed_when: false
- name: Export host list
ansible.builtin.command: ceph orch host ls --format yaml
register: orch_hosts_export
changed_when: false
- name: Write ceph.conf backup
ansible.builtin.copy:
content: "{{ ceph_conf_export.stdout }}\n"
dest: "{{ backup_local_dir }}/ceph.conf"
mode: '0600'
delegate_to: localhost
- name: Write admin keyring backup
ansible.builtin.copy:
content: "{{ admin_keyring_export.stdout }}\n"
dest: "{{ backup_local_dir }}/ceph.client.admin.keyring"
mode: '0600'
delegate_to: localhost
- name: Write CRUSH map backup
ansible.builtin.copy:
content: "{{ crush_export.stdout }}\n"
dest: "{{ backup_local_dir }}/crushmap.txt"
mode: '0600'
delegate_to: localhost
- name: Write OSD dump backup
ansible.builtin.copy:
content: "{{ osd_dump_export.stdout }}\n"
dest: "{{ backup_local_dir }}/osd-dump.json"
mode: '0600'
delegate_to: localhost
- name: Write monitor dump backup
ansible.builtin.copy:
content: "{{ mon_dump_export.stdout }}\n"
dest: "{{ backup_local_dir }}/mon-dump.json"
mode: '0600'
delegate_to: localhost
- name: Write config dump backup
ansible.builtin.copy:
content: "{{ config_dump_export.stdout }}\n"
dest: "{{ backup_local_dir }}/config-dump.json"
mode: '0600'
delegate_to: localhost
- name: Write RGW realm backup
ansible.builtin.copy:
content: "{{ rgw_realm_export.stdout }}\n"
dest: "{{ backup_local_dir }}/rgw-realm.json"
mode: '0600'
delegate_to: localhost
- name: Write service specs backup
ansible.builtin.copy:
content: "{{ orch_services_export.stdout }}\n"
dest: "{{ backup_local_dir }}/orch-services.yaml"
mode: '0600'
delegate_to: localhost
- name: Write host list backup
ansible.builtin.copy:
content: "{{ orch_hosts_export.stdout }}\n"
dest: "{{ backup_local_dir }}/orch-hosts.yaml"
mode: '0600'
delegate_to: localhost
- name: Report backup location
ansible.builtin.debug:
msg: "Backup written to {{ backup_local_dir }}/"
run_once: true # noqa: run-once[task]
+21
View File
@@ -0,0 +1,21 @@
---
# Post-boot OS baseline — ops user, packages, /etc/hosts, services.
# Run after provisioning, before ceph_deploy.
# Fully convergeable — re-running fixes any drift in user config,
# packages, or host entries.
#
# Usage:
# scripts/ansible-play.sh baseline.yml
#
# Tags:
# --tags users ops user only
# --tags packages podman + diagnostics only
# --tags system /etc/hosts + services only
- name: OS baseline configuration
hosts: ceph_nodes
become: true
gather_facts: false
roles:
- baseline
+21
View File
@@ -0,0 +1,21 @@
---
# S3 benchmark — runs in parallel on all Ceph nodes against local RGW.
#
# Default: 1000 objects x 16 MiB x 3 nodes = ~48 GiB PUT workload
#
# Usage:
# scripts/ansible-play.sh bench.yml # default PUT
# scripts/ansible-play.sh bench.yml -e s3bench_ops=get # GET (after PUT)
# scripts/ansible-play.sh bench.yml -e s3bench_ops=mixed # 70/20/10 mix
# scripts/ansible-play.sh bench.yml -e s3bench_num_objects=10000 # ~480 GiB fill
# scripts/ansible-play.sh bench.yml -e s3bench_ops=delete # cleanup
#
# Results: bench/ directory (gitignored)
- name: S3 benchmark
hosts: ceph_nodes
become: true
gather_facts: false
roles:
- s3_bench
+22
View File
@@ -0,0 +1,22 @@
---
# Deploy Ceph Tentacle cluster (cluster chosen via CEPH_ENV)
#
# Prerequisites:
# 1. All nodes provisioned and running Debian 12 with bond0 networking
# 2. SSH access via inventory.ini credentials (rendered by tofu apply)
# 3. Storage partitioned with Ceph block.db LVs and SSD OSD partitions
#
# Usage:
# scripts/ansible-play.sh deploy-ceph.yml
#
# To run a specific phase only:
# scripts/ansible-play.sh deploy-ceph.yml --tags bootstrap
# scripts/ansible-play.sh deploy-ceph.yml --tags osds
- name: Deploy Ceph Tentacle cluster
hosts: ceph_nodes
become: true
gather_facts: true
roles:
- ceph_deploy
+26
View File
@@ -0,0 +1,26 @@
---
# Destroy Ceph Tentacle cluster — COMPLETE TEARDOWN
#
# This removes ALL pools, OSDs, daemons, and cluster state.
# All data on Ceph-managed devices will be permanently lost.
#
# Guardrails:
# 1. Must pass yes_destroy_ceph=true or it refuses to run
# 2. Must pass destroy_target_domain matching cluster_domain
# 3. Pauses for manual confirmation before destructive steps
#
# Usage:
# scripts/ansible-play.sh destroy-ceph.yml \
# -e "yes_destroy_ceph=true destroy_target_domain=dev.austin.int.futo.cloud"
#
# To run a specific phase only:
# scripts/ansible-play.sh destroy-ceph.yml -e "yes_destroy_ceph=true destroy_target_domain=dev.austin.int.futo.cloud" --tags purge
# scripts/ansible-play.sh destroy-ceph.yml -e "yes_destroy_ceph=true destroy_target_domain=dev.austin.int.futo.cloud" --tags cleanup
- name: Destroy Ceph Tentacle cluster
hosts: ceph_nodes
become: true
gather_facts: true
roles:
- ceph_destroy
+378
View File
@@ -0,0 +1,378 @@
# Adding a cluster
Clusters are declared in `tf/deployment/<env>/ceph/clusters.auto.tfvars`.
Every cluster-scoped concern — inventory file, hostname, 1P item names, SSH
key path, secrets template — is derived from that one entry. Most of what
this walkthrough describes is editing that file and running
`mise run tf:apply`; the rest is creating the 1P items TF expects to read
at playbook time.
For the broader architecture see [docs/architecture.md](architecture.md); for
the per-item secrets catalog see [docs/secrets.md](secrets.md); for the
naming rules see [docs/naming.md](naming.md).
## What TF does vs. what you do
| TF (`mise run tf:apply`) | You (one-time per cluster) |
|-------------------------------------------------------------------------|---------------------------------------------------------------------|
| Renders `inventory.ini`, `inventory-destroy.ini`, `secrets.yml.tpl` | Create `group_vars/all/vars.yml` (cluster-wide Ansible config) |
| Renders `inventory-provision-<profile>.ini` when `provision_profile` set | Create one `host_vars/<hostname>.yml` per node (hardware topology) |
| Picks auto-names from the wordlist for hosts where `name = null` | Create 1P items: passwords, SSH keypair |
| Computes hostnames, FQDNs, 1P item titles, inventory directory path | Run `scripts/install-ssh-keys.sh <cluster>` on your workstation |
| (Future) Creates `onepassword_item` resources for passwords | Run `mise run preflight` + `mise run deploy` |
Every operator doing a cluster add follows the same steps — nothing in
this walkthrough is machine- or operator-specific.
## Inventory directory naming
```
inventories/<cluster>-<role>.<env>.<datacenter>.<provider>/
```
The `<role>` segment comes from `role_in_hostname` in the TFvars (defaults
to `ceph`). Every Ceph-project inventory grep-matches `*-ceph.*` regardless
of datacenter or environment.
Existing examples:
- `sietch-ceph.dev.austin.int/` — Austin DC, internal network, dev
- `painbox-ceph.dev.hel.htz/` — Hetzner Helsinki, dev
Future environments land as siblings: `*-ceph.staging.<dc>.<provider>/`,
`*-ceph.prod.<dc>.<provider>/`.
## Step-by-step
### 1. Choose a cluster name
The engineer adding the cluster picks the name. Conventions and constraints
live in [docs/naming.md](naming.md#cluster-naming). Quick summary:
- **Convention:** Dune-themed (existing: `sietch`, `painbox`). Not enforced.
- **Constraints:** lowercase, short (6–10 chars ideal), no dashes or dots,
unique within the `yucca_tf_*` item namespace, not already a key in
`clusters.auto.tfvars`.
- **Cost of renaming later:** expensive (touches hostnames, 1P items,
cephadm identity, SSH keys, DNS). Pick deliberately.
Host names within a cluster can be operator-declared in the TFvars or
auto-picked from the 923-word wordlist — see [docs/naming.md](naming.md#host-naming).
### 2. Declare the cluster in TF
Edit `tf/deployment/<env>/ceph/clusters.auto.tfvars` and add an entry.
Working example for a hypothetical `mesa` cluster at Hetzner Falkenstein:
```hcl
clusters = {
sietch = { ... }
painbox = { ... }
mesa = {
domain = "dev.fsn.htz.futo.cloud"
environment = "dev"
datacenter = "fsn"
provider_code = "htz"
role_in_hostname = "ceph"
ansible_ssh_user = "ansible-iac" # Hetzner installimage boots as root;
# baseline creates ansible-iac before first deploy
ansible_ssh_key = "~/.ssh/id_ed25519_mesa"
vault = "yucca_tf_dev" # or yucca_tf_staging / yucca_tf_prod
provision_profile = null # Hetzner installimage; no debian-live provisioning
hosts = [
{ bond_ip = "<public-ip>", bootstrap = true }, # name auto-picked from wordlist
]
}
}
```
Notes:
- `vault` declares which 1Password vault TF rendering will write into the
`secrets.yml.tpl`. `yucca_tf_dev` for dev clusters, `yucca_tf_staging`
or `yucca_tf` (prod) for their respective environments.
- `provision_profile = "debian-live"` enables bare-metal provisioning via
`provision.yml` (rendered `inventory-provision.ini`). Leave null for
Hetzner installimage workflows — the post-install script uses its own
path (`inventories/<cluster>/installimage/post-install.sh.tpl`).
- Host `name = null` (omitted) → TF picks a stable wordlist name seeded
per-cluster. Auto-picks don't change on subsequent applies.
### 3. Render the inventory + secrets template
```bash
mise run tf:apply
```
Or, for a non-default stack:
```bash
TF_STACK_DIR=tf/deployment/<env>/ceph mise run tf:apply
```
This creates (per the module's `rendering.tf`):
- `ansible/ceph/inventories/<cluster>-ceph.<env>.<dc>.<provider>/inventory.ini`
- `.../inventory-destroy.ini`
- `.../secrets.yml.tpl`
- `.../inventory-provision-<profile>.ini` (only when `provision_profile` is set)
All of these are gitignored — re-run `mise run tf:apply` after any
`clusters.auto.tfvars` change.
### 4. Create `group_vars/all/vars.yml`
Hand-maintained, committed. Copy the closer existing analogue as a starting
point:
- **Bare-metal cluster:** copy from `sietch-ceph.dev.austin.int/group_vars/all/vars.yml`
- **Hetzner/single-NIC cluster:** copy from `painbox-ceph.dev.hel.htz/group_vars/all/vars.yml`
```bash
cp inventories/painbox-ceph.dev.hel.htz/group_vars/all/vars.yml \
inventories/mesa-ceph.dev.fsn.htz/group_vars/all/vars.yml
```
Edit every value. Required shape:
```yaml
---
# === Naming ===
cluster_name: mesa
cluster_role: ceph
cluster_domain: dev.fsn.htz.futo.cloud
# === Network ===
public_network: <subnet or public /32>
cluster_network: <same as public for single-network topology>
# Bonds / gateway / DNS — omit or customize per hardware
# === Ceph ===
ceph_release: tentacle
ceph_repo_url: "https://download.ceph.com/debian-{{ ceph_release }}/"
ceph_repo_key_url: "https://download.ceph.com/keys/release.asc"
# === OS Provisioning ===
admin_user: ansible-iac
timezone: UTC
# Used by provision.yml's post-reboot SSH verification to read the marker.
provision_iac_ssh_key_path: "~/.ssh/id_ed25519_mesa"
# === Secret aliases (populated by op inject at playbook time) ===
# These map TF-rendered vault_* names into the role-facing names the
# playbooks consume. Add one alias per secret declared in the module's
# secrets map (tf/shared/modules/ceph-cluster/main.tf).
ops_password: "{{ vault_ops_password }}"
ceph_dashboard_user: admin
ceph_dashboard_password: "{{ vault_ceph_dashboard_password }}"
ceph_grafana_admin_user: admin
ceph_grafana_admin_password: "{{ vault_grafana_admin_password }}"
ceph_rgw_s3_user_access_key: "{{ vault_s3_restic_access_key }}"
ceph_rgw_s3_user_secret_key: "{{ vault_s3_restic_secret_key }}"
# === RGW ===
ceph_rgw_realm: <cluster-name>
ceph_rgw_zonegroup: <zonegroup>
ceph_rgw_zone: <zone>
# === Storage ===
ssd_model_pattern: "Micron_5100" # match your SSD model
# ... (see the cluster you copied from for full hardware config)
```
### 5. Create `host_vars/<hostname>.yml` per node
Host files are committed (per-cluster hardware topology is stable inventory
truth — not operator preference). Use `<cluster>/host_vars/example.yml` as
a template.
```bash
CLUSTER_DIR=inventories/mesa-ceph.dev.fsn.htz
# Look up the hostname TF picked (or declared) — visible in the rendered inventory.ini
TF_OUTPUT=$(cat "$CLUSTER_DIR/inventory.ini")
# Create one host_vars file per hostname_short shown in the [ceph_nodes] section
cp "$CLUSTER_DIR/host_vars/example.yml" "$CLUSTER_DIR/host_vars/<hostname_short>.yml"
```
Edit with node-specific hardware facts: `bond_ip`, SAS expander path
prefix, SSD PHY positions, HDD-to-block.db-LV mappings. See
[docs/hardware.md](hardware.md) for the shape.
Operator-local overrides (e.g., testing a workaround on one node) can go
in `<hostname_short>.local.yml` — that suffix is gitignored.
### 6. Create 1Password items
For the target vault declared in the cluster's TFvars entry:
```bash
VAULT=yucca_tf_dev # match the vault field in clusters.auto.tfvars
CLUSTER=MESA # uppercase cluster_name
# Password items — 3 ending in _PASSWORD
for role in OPS DASHBOARD GRAFANA; do
op item create --vault "$VAULT" --category password \
--title "${CLUSTER}_CEPH_${role}_PASSWORD" \
--generate-password='letters,digits,32'
done
# S3 service-user keys — 2 items; names already end in _KEY
for suffix in S3_SVC_YUCCA_RESTIC_ACCESS_KEY S3_SVC_YUCCA_RESTIC_SECRET_KEY; do
op item create --vault "$VAULT" --category password \
--title "${CLUSTER}_CEPH_${suffix}" \
--generate-password='letters,digits,32'
done
# SSH Key item — one keypair per cluster. op CLI GENERATES the key inside 1P;
# we never create the private key on disk first (see ADR-010).
op item create --vault "$VAULT" \
--category "SSH Key" \
--title "${CLUSTER}_CEPH_ANSIBLE_IAC_SSH_KEY" \
--ssh-generate-key=ed25519
```
Verify the TF-rendered template resolves:
```bash
CEPH_ENV=inventories/mesa-ceph.dev.fsn.htz/inventory.ini
op inject -f -i "$(dirname $CEPH_ENV)/secrets.yml.tpl" -o /tmp/test-secrets.yml
head -5 /tmp/test-secrets.yml && rm /tmp/test-secrets.yml
```
Disaster-recovery items (`<CLUSTER>_CEPH_RGW_TLS_CERT`, `_RGW_TLS_KEY`,
`_CLIENT_ADMIN_KEYRING`) are **not** created here — they're populated by
`mise run capture` after the first successful deploy. Skipping that step
is the most common gotcha.
### 7. Install the SSH keypair on your workstation
```bash
scripts/install-ssh-keys.sh mesa
```
The wrapper reads `private_key` and `public_key` from
`${CLUSTER}_CEPH_ANSIBLE_IAC_SSH_KEY` and writes `~/.ssh/id_ed25519_mesa`
(0600) + `.pub` (0644). Idempotent — re-running is safe. Every operator
who will run plays against this cluster runs this command once on their
workstation (or any time they wipe `~/.ssh/`).
The `ansible_ssh_key` path in `clusters.auto.tfvars` must match what
`install-ssh-keys.sh` writes. If you chose a non-default filename,
update both together (or update the mapping in `install-ssh-keys.sh`).
See [docs/scripts.md](scripts.md#install-ssh-keyssh) for the script
reference and [ADR-010](adr/010-ssh-keys-in-1password.md) for the
rationale.
### 8. Preflight
```bash
CEPH_ENV=inventories/mesa-ceph.dev.fsn.htz/inventory.ini mise run preflight
```
(Inline-prefix form — `export CEPH_ENV=...` then `mise run preflight`
does NOT work; mise's `[env]` block strips shell exports. See
[docs/scripts.md "Setting CEPH_ENV"](scripts.md).)
Verifies: TF artifacts present, 1P session live, `op inject` resolves
the template, SSH reachable, Python 3 on targets.
### 9. Deploy
For Hetzner installimage clusters, run the installimage flow first
(out-of-band; see [runbooks/painbox-reprovision.md](runbooks/painbox-reprovision.md)
for the pattern). For Austin bare-metal clusters, run `provision.yml`
first (boot into the live image, then `scripts/ansible-play.sh
provision.yml -e confirm_wipe=true` with
`CEPH_ENV=.../inventory-provision.ini`). Then:
```bash
mise run deploy
```
Every task invocation goes through `scripts/ansible-play.sh`, which
`op inject`s the secrets template into a short-lived tmpfile and passes
it as `--extra-vars @<tmpfile>`.
### 10. Capture the DR belt-and-suspenders items
After the first successful deploy:
```bash
mise run capture
```
This reads `/etc/ceph/rgw-ssl.crt`, `/etc/ceph/rgw-ssl.key`, and
`/etc/ceph/ceph.client.admin.keyring` from the bootstrap node and upserts
them as Document items in the cluster's vault
(`<CLUSTER>_CEPH_RGW_TLS_CERT`, `_RGW_TLS_KEY`, `_CLIENT_ADMIN_KEYRING`). Safe
to re-run — updates in place on content drift.
## How `CEPH_ENV` works
`CEPH_ENV` points to the inventory file, not the directory:
```bash
# Correct
CEPH_ENV=inventories/mesa-ceph.dev.fsn.htz/inventory.ini
# Wrong — directory mode loads every .ini including destroy inventory
CEPH_ENV=inventories/mesa-ceph.dev.fsn.htz/
```
Default is set in `.mise.toml` (`sietch` in dev). Override per-command:
```bash
CEPH_ENV=inventories/painbox-ceph.dev.hel.htz/inventory.ini mise run status
```
`scripts/ansible-play.sh` derives the secrets template path from `CEPH_ENV`
(same directory, `secrets.yml.tpl`).
## What lives where in the repo
| Committed | Gitignored (TF-rendered or operator-local) |
|------------------------------------------------------|--------------------------------------------|
| `clusters.auto.tfvars` | `inventories/*/inventory.ini` |
| `inventories/<cluster>/group_vars/all/vars.yml` | `inventories/*/inventory-provision.ini` |
| `inventories/<cluster>/host_vars/<hostname>.yml` | `inventories/*/inventory-destroy.ini` |
| `inventories/<cluster>/installimage/*.tpl` | `inventories/*/secrets.yml.tpl` |
| | `inventories/*/host_vars/*.local.yml` |
| | `inventories/*/installimage/post-install.sh` |
No secrets are ever committed. The `.tpl` file contains `op://` references
only; `op inject` resolves them at play time into a 0600 tmpfile that's
trap-cleaned on exit.
## Common gotchas
- **Forgot step 10 (`mise run capture`)** — DR items are missing in 1P.
Running `capture` after the fact works; it just needs the bootstrap
node's filesystem intact.
- **`ansible_ssh_key` path mismatch** between `clusters.auto.tfvars` and
`install-ssh-keys.sh` — key installed under a different name than what
the inventory expects. Keep them aligned.
- **Fingerprint mismatch on `install-ssh-keys.sh`** — happens after a
key rotation if you haven't moved the old key aside. Follow the
`mv ~/.ssh/id_ed25519_<cluster>{,.$(date +%Y%m%d).bak}` path in the
wrapper's error message.
- **`host_vars/` out of date after `tofu apply` re-picks a wordlist
name** — auto-names are stable across applies, but if you add hosts
at positions other than the tail, shuffled names may shift. Add new
hosts at the end of the `hosts = [...]` list to keep existing
hostnames stable.
- **Painbox-style Hetzner clusters and `provision_profile`** — leave
`provision_profile = null` so TF doesn't render the debian-live
inventory. The Hetzner installimage flow has its own post-install
script under `installimage/`, not Ansible-driven.
## See also
- [architecture.md §4 (Terraform)](architecture.md) — what TF owns
- [secrets.md](secrets.md) — per-item catalog + vault selection
- [naming.md](naming.md) — cluster + host naming conventions
- [hardware.md](hardware.md) — `host_vars/` shape, network topology
- [scripts.md](scripts.md) — wrapper reference
- [ADR-009](adr/009-tf-first-op-inject-over-vault-password-sh.md) — why TF is authoritative
- [ADR-010](adr/010-ssh-keys-in-1password.md) — why SSH keys live in 1P
- [runbooks/painbox-reprovision.md](runbooks/painbox-reprovision.md) — Hetzner-specific reprovisioning pattern
+116
View File
@@ -0,0 +1,116 @@
# Adding a role
How to create a new Ansible role in this project. The heavy lifting is in
the existing roles — this doc is the skeleton + wiring + pre-submit
checklist; copy an exemplar for idioms.
For code-level patterns (idempotency, `changed_when`, handlers, secrets
handling, shell conventions), see [patterns.md](patterns.md). For how roles
compose into the overall pipeline, see
[architecture.md §6](architecture.md).
## Skeleton
```
roles/<role_name>/
├── defaults/main.yml required — every variable with a default + comment
├── meta/main.yml required — author, license, Ansible version
├── tasks/main.yml required — imports sub-task files, tagged
├── handlers/main.yml if the role restarts/reloads services
├── templates/*.j2 Jinja2 templates
└── molecule/default/ optional — test scenario
```
Role names use `snake_case` (`ceph_deploy`, `os_tuning`).
**Cluster-level variables** use the `ceph_` prefix (overridable in
`group_vars/all/vars.yml`). **Role-internal variables** use the
`<role_name>_` prefix. The `.ansible-lint` config skips
`var-naming[no-role-prefix]` because all roles share the `ceph_` prefix
for cluster-level settings — this is intentional.
## Exemplars to copy from
| Copy this when you're writing... | Role |
|---|---|
| A phased deployment with tags | `ceph_deploy` |
| Modular sub-task files with header comments | `baseline` |
| Kernel/sysctl values with units documented | `os_tuning` |
| Per-device-class settings + feature toggles | `hardware_tuning` |
| Templated firewall config with opt-in features | `security` |
| Ceph CLI shell tasks with `changed_when` patterns | `ceph_tuning` |
Every existing role has a header comment block at the top of
`defaults/main.yml` explaining its scope — open one and mirror the shape.
## Wiring into the pipeline
### 1. Playbook wrapper
Create a top-level playbook (e.g. `my-feature.yml`) that imports the role:
```yaml
---
# Brief description + usage tags.
- name: My feature
hosts: ceph_nodes
become: true
roles:
- my_role_name
```
### 2. `site.yml` (if part of full deploy)
Insert in the correct dependency position in `site.yml`:
```yaml
- import_playbook: my-feature.yml
```
Order matters — see [architecture.md §6 "Why this order matters"](architecture.md).
### 3. `mise run deploy` (if part of full deploy)
Add to the `deploy` task in yucca-root `.mise/config.toml` (or
`ansible/ceph/.mise.toml` if the task lives there). Always invoke via
the wrapper so secrets resolve:
```toml
echo "=== My feature ===" && scripts/ansible-play.sh my-feature.yml
```
### 4. Optional: standalone mise task
For playbooks useful to run independently (like `bench`, `drift`,
`status`):
```toml
[tasks.my-feature]
description = "One-line description"
run = "scripts/ansible-play.sh my-feature.yml"
```
See [scripts.md](scripts.md) for the wrapper reference.
## Molecule test (optional)
Not every role needs one — the `ceph_deploy` role is the only one with
a full scenario today. If you're writing something non-trivial, copy
`roles/ceph_deploy/molecule/default/` as a starting point.
## Pre-submit checklist
Before opening a PR, verify:
- [ ] `mise run lint` passes clean (yamllint + ansible-lint + shellcheck)
- [ ] `mise run check` passes syntax-check with the role's playbook
- [ ] Every variable in `defaults/main.yml` has a comment explaining what it does
- [ ] All modules use FQCN (`ansible.builtin.apt`, not `apt`)
- [ ] All `shell`/`command` tasks have `changed_when`
- [ ] All `shell` tasks set `args.executable: /bin/bash` and include `set -o pipefail`
- [ ] Secret-handling tasks have `no_log: true`
- [ ] `meta/main.yml` has author, license, description, `min_ansible_version: "2.19"`
- [ ] Tags on `import_tasks` in `main.yml`
- [ ] Playbook wrapper exists at `ansible/ceph/` root
- [ ] Role is added to `site.yml` in the correct position (if part of full deploy)
- [ ] Anti-patterns in [patterns.md §Anti-patterns](patterns.md) not violated
@@ -0,0 +1,49 @@
# ADR-001: cephadm + Ansible Wrapper Roles over ceph-ansible
## Status
Accepted
## Context
The official `ceph-ansible` project targets Ceph Quincy and older releases. Ceph
Tentacle (the release this cluster runs) is a cephadm-native release where
`ceph-ansible` is deprecated and no longer tested. Additionally, `ceph-ansible`
is a large, opinionated framework that owns the entire node lifecycle -- from
package installation to OSD creation -- making it difficult to compose with our
own provisioning pipeline (debootstrap, baseline role, security hardening).
We needed a deployment approach that:
- Works with Ceph Tentacle (cephadm-native release).
- Gives us explicit control over each phase (bootstrap, join, OSD creation,
RGW, monitoring) so we can debug and re-run individual stages.
- Stays composable with our existing Ansible roles for OS provisioning,
baseline configuration, and security.
## Decision
We use **cephadm directly**, wrapped in thin Ansible task files inside the
`ceph_deploy` role. Each deployment phase is a separate task file
(`bootstrap.yml`, `join.yml`, `osds.yml`, `rgw.yml`, etc.) that calls
`cephadm` and `ceph` CLI commands with idempotency guards.
The Ansible layer handles orchestration (ordering, host targeting, variable
interpolation, idempotency checks) while cephadm handles container management,
daemon lifecycle, and config distribution. Comments in `bootstrap.yml` note
that cephadm automatically distributes `ceph.conf`, admin keyrings, and SSH
keys during `ceph orch host add` -- we lean on that rather than reimplementing
distribution logic.
## Consequences
- **Positive:** Full compatibility with Ceph Tentacle and future releases. Each
phase is independently re-runnable. Task files are small and auditable. No
dependency on an external Ansible Galaxy role with its own release cycle.
- **Positive:** Operators can `--tags osds` to re-run just OSD creation after a
failure, or `--tags rgw` to redeploy the gateway layer independently.
- **Negative:** We own the idempotency logic (e.g., `pvs | grep ceph` checks,
`stat` on `/etc/ceph/ceph.conf`). Upstream ceph-ansible handled this
automatically, but it also hid failures behind abstraction layers.
- **Negative:** New Ceph features require manual task-file additions rather than
a Galaxy role version bump.
@@ -0,0 +1,47 @@
# ADR-002: Explicit OSD-to-Disk Mapping in host_vars
## Status
Accepted
## Context
Ceph can auto-discover available drives via `ceph orch apply osd --all-available-devices`.
This is convenient but dangerous in a mixed-media cluster with complex disk
layouts. Our Austin nodes have:
- 12 front-bay HDDs assigned to OSD data, each needing a specific block.db LV
on one of two rear-bay SSDs.
- 2 rear-bay SSDs that are multi-purpose: partition 1-4 for OS (mdraid+LVM),
partition 5 for block.db LVs, partition 6 for SSD OSDs.
- Device paths that use `/dev/disk/by-path/` with SAS expander PHY addresses
for slot stability (replacing a drive in the same bay keeps the same path).
Auto-discovery cannot express the HDD-to-SSD-db mapping. It also risks
claiming OS partitions or db partitions as OSD data. A mismap during automated
discovery would silently create OSDs without block.db acceleration, or worse,
destroy the OS volume.
## Decision
Every OSD is explicitly listed in each node's `host_vars` file. HDD OSDs
specify both the PHY slot (`path_phy`) and the exact block.db LV (`db`). SSD
OSDs specify the PHY slot and partition number. The `osds.yml` task file
iterates these lists with `loop:`, resolving each entry to a
`/dev/disk/by-path/` stable path.
Each entry is idempotent: the task checks `pvs` for an existing Ceph PV on the
resolved device and skips creation if one exists. Empty bays (missing block
device) are also skipped gracefully.
## Consequences
- **Positive:** Every OSD-to-disk-to-db mapping is version-controlled and
auditable. Drive replacements are tracked with serial number comments in
host_vars. No risk of accidental OSD creation on OS or db partitions.
- **Positive:** Slot-stable `by-path` addressing means a replacement drive in
the same bay inherits the same OSD definition -- no host_vars edit needed.
- **Negative:** Adding a new node requires writing a complete host_vars file
with all PHY mappings. This is manual but happens rarely (new hardware).
- **Negative:** Cannot scale to hundreds of heterogeneous nodes without
templating. Acceptable for our 3-node (scaling to ~10) cluster size.
@@ -0,0 +1,51 @@
# ADR-003: Baseline Role Split from Provisioning
## Status
Accepted
## Context
The `provision_host` role runs inside a chroot on a live image. It must create
the minimum viable user (`ansible-iac`) so Ansible can connect after first
reboot. Originally, the ops user, diagnostic packages, `/etc/hosts`, and
service enablement were also done in chroot.
This caused problems:
- Chroot operations are fragile -- bind mounts for `/dev`, `/proc`, `/sys`
must be set up and torn down correctly. More chroot work means more failure
surface during provisioning.
- The ops user password comes from 1Password (resolved at play time via
`op inject` — see ADR-009). Embedding it in the chroot phase means
provisioning depends on the operator's 1P session being live during
install, which complicates the live-image environment.
- Post-provision drift (stale `/etc/hosts`, missing packages, changed
passwords) required re-provisioning from the live image to fix. There was
no way to converge config on a running system.
## Decision
The `provision_host` role now creates only `ansible-iac` (uid 1000, locked
password, key-only SSH, NOPASSWD sudo) inside chroot. Everything else moved
to a separate **baseline** role that runs post-boot via the normal Ansible
pipeline:
- `users.yml` -- ops user with op-injected password, sudo config
- `packages.yml` -- podman ecosystem, diagnostic tools
- `system.yml` -- `/etc/hosts` template, service enablement, timezone
The baseline role is convergeable: re-running it on a live system corrects
drift without re-provisioning.
## Consequences
- **Positive:** Provisioning is faster and less fragile -- fewer chroot
operations, no vault dependency during OS install.
- **Positive:** Password changes, package additions, and `/etc/hosts` updates
are applied by re-running `baseline.yml` against running nodes. No reboot
or live-image cycle needed.
- **Positive:** Clear responsibility boundary: provision_host owns
"bare metal to bootable OS", baseline owns "bootable OS to operational node".
- **Negative:** Two-step initial setup (provision, then baseline) instead of
one. Mitigated by the `site.yml` playbook which chains both.
@@ -0,0 +1,50 @@
# ADR-004: Multi-Inventory over Multi-Repo
## Status
Accepted. The core decision (one repo, per-cluster inventory dirs) stands.
Refined by ADR-009 — inventory files are now TF-rendered rather than
hand-authored, and the `vault-password.sh`/`secrets-init.sh` mechanism
referenced below has been replaced by `op inject` + `scripts/ansible-play.sh`.
## Context
We operate multiple Ceph clusters across different sites and purposes:
- `sietch-ceph.dev.austin.int` -- 3-node dev cluster in Austin datacenter
- `painbox-ceph.dev.hel.htz` -- single-node dev cluster at Hetzner Helsinki
Each cluster has different hardware, networking, SSH access, and credentials.
The common approaches are: (a) one repository per cluster, or (b) one
repository with per-cluster inventory directories.
Separate repos cause role drift -- a fix to `ceph_deploy` in one repo must be
cherry-picked to every other repo. Shared roles via Git submodules or Galaxy
add dependency management overhead. The clusters share the same roles, playbooks,
and scripts; only the inventory data (host lists, IPs, credentials, host_vars)
differs.
## Decision
One repository with a top-level `inventories/` directory containing one
subdirectory per cluster. Each cluster directory has its own `inventory.ini`,
`group_vars/`, and `host_vars/`. The `ansible.cfg` has no default inventory --
the active cluster is selected via `-i inventories/<cluster>/inventory.ini` or
the `CEPH_ENV` environment variable.
Scripts like `vault-password.sh` and `secrets-init.sh` derive the cluster
identity from `CEPH_ENV` (parsing the inventory path to extract the cluster
ID), so they work against any cluster without hardcoded names.
## Consequences
- **Positive:** Role and playbook changes apply to all clusters immediately.
No cherry-picking or submodule syncing.
- **Positive:** Cluster-specific config (IPs, OSD maps, SSH keys, vault
passwords) is cleanly isolated in per-cluster inventory directories.
- **Positive:** Scripts auto-detect cluster context from `CEPH_ENV`, so the
same tooling works for Austin and Hetzner without modification.
- **Negative:** A broken role change affects all clusters. Mitigated by testing
on painbox (Hetzner) before applying to sietch (Austin production path).
- **Negative:** Repository grows with each cluster. Acceptable at our scale
(inventory data is small).
@@ -0,0 +1,58 @@
# ADR-005: 1Password CLI over HashiCorp Vault
## Status
Superseded by [ADR-009](./009-tf-first-op-inject-over-vault-password-sh.md).
The core decision (1Password over HashiCorp Vault) stands — the team already
runs on 1P for credentials, and a self-hosted Vault server wasn't justified
for this scale. What's changed is the *mechanism*: the `vault-password.sh`
script + `ansible-vault` + `vault.yml` pipeline described below has been
retired in favor of TF-provisioned 1P items + `op inject` at playbook time.
See ADR-009 for the current implementation. This document is preserved for
historical context on the Vault-vs-1P choice.
## Context
The Ansible playbooks need secrets: vault passwords for encrypted vars files,
dashboard credentials, ops user passwords, and S3 keys. The standard options
are:
- `ansible-vault` with a password file or `--ask-vault-pass` -- simple but
the vault password itself needs to live somewhere (plaintext file, env var,
or manual entry every run).
- HashiCorp Vault -- powerful but requires its own infrastructure (server,
unsealing, token management, HA). Overkill for a small team managing a few
clusters.
- 1Password -- already used by the team for credential management. Has a CLI
(`op`) with service account tokens for CI and desktop app integration for
developer workstations.
## Decision
Secrets are managed in 1Password. The `vault-password.sh` script is the
Ansible `vault_password_file`. It fetches the vault password from a 1Password
item, trying four auth methods in order:
1. Service account token (`OP_SERVICE_ACCOUNT_TOKEN`) for CI/headless.
2. Desktop app integration (polkit/YubiKey) for developer workstations.
3. Interactive prompt as TTY fallback.
4. Dummy password for lint/syntax-check (no TTY, no `op`).
A companion `secrets-init.sh` script bootstraps all 1Password items for a new
cluster (vault password, dashboard login, grafana login, ops user) with
auto-generated passwords and multi-URL entries for browser autofill. Items
are stored in a dedicated "Yucca" vault and named by cluster FQDN.
## Consequences
- **Positive:** No additional infrastructure to manage. 1Password is already
the team's credential store -- secrets live alongside other org credentials.
- **Positive:** YubiKey/biometric unlock on workstations means no plaintext
password files on disk. Service account tokens provide headless CI access.
- **Positive:** `secrets-init.sh` makes cluster credential bootstrapping
repeatable -- new clusters get consistent 1Password items automatically.
- **Negative:** Hard dependency on 1Password CLI (`op`). Mitigated by the
interactive and dummy fallbacks in `vault-password.sh`.
- **Negative:** 1Password is a SaaS dependency. Acceptable trade-off vs.
self-hosting HashiCorp Vault for a small operations team.
@@ -0,0 +1,48 @@
# ADR-006: Self-Signed TLS for RGW
## Status
Accepted
## Context
The RadosGW (S3) frontend needs TLS. The options are:
- **Let's Encrypt** -- requires public DNS and HTTP-01 or DNS-01 challenge
validation. Our clusters sit on private networks (10.10.10.0/24) with no
public DNS records and no inbound internet access. ACME is not viable
without a DNS provider API and split-horizon DNS.
- **Private CA** -- proper chain of trust, but requires CA infrastructure
(key ceremony, CRL/OCSP, distribution of the CA cert to every client).
Significant operational overhead for an internal dev/staging cluster.
- **Self-signed certs** -- simple to generate, no external dependencies. S3
clients (aws-cli, boto3, rclone) all support disabling cert verification
or trusting a custom cert.
## Decision
RGW uses a self-signed certificate generated by `openssl` during deployment.
The cert is created on the bootstrap node with a 10-year validity period and
includes SANs for:
- The canonical DNS name and a wildcard under it (for virtual-hosted buckets).
- Per-node FQDNs and bond IPs (so direct-host and IP-based access validates).
The combined PEM (cert + key) is embedded in the cephadm RGW service spec.
Cephadm distributes it to every RGW daemon container. The dashboard's RGW API
SSL verification is disabled to accommodate the self-signed cert.
Rotation is manual: delete the cert files on the bootstrap node and re-run the
role. The `creates:` guard makes the generation idempotent.
## Consequences
- **Positive:** Zero external dependencies. Works on air-gapped and private
networks. No DNS provider API, no ACME client, no CA infrastructure.
- **Positive:** 10-year validity avoids renewal automation for a dev cluster.
SANs cover all access patterns (DNS, hostname, IP, virtual-hosted buckets).
- **Negative:** S3 clients must either disable TLS verification or import the
self-signed cert. Documented in s3-integration.md.
- **Negative:** Not suitable for production clusters serving untrusted clients.
A future production ADR will revisit this with a private CA or ACME via
DNS-01 challenge.
@@ -0,0 +1,51 @@
# ADR-007: nftables over iptables
## Status
Accepted
## Context
Ceph nodes expose many services: MON (3300, 6789), OSD (6800-7568), MGR/dashboard
(8443), RGW/S3 (443), Prometheus (9095), Grafana (3000), Alertmanager (9093),
node-exporter (9100), plus optional iSCSI and NFS ports. Without a firewall,
all of these are reachable from any source on the network.
The two main Linux firewall frameworks are:
- **iptables/ip6tables** -- legacy, being replaced upstream. Uses separate
tables for IPv4/IPv6. Debian 12 still ships it but marks it deprecated.
- **nftables** -- the successor. Single framework for IPv4/IPv6/ARP. Native
in the kernel since 3.13. Debian 12's default backend for `iptables` is
already `nft`.
## Decision
We use **nftables** with a Jinja2-templated ruleset (`nftables.conf.j2`)
managed by the security role. The template generates a complete `inet filter`
table with:
- Default-drop input policy.
- Established/related connection tracking.
- SSH open to all (or restricted to trusted networks, controlled by a variable).
- RGW/S3 port open to all (public-facing service).
- All other Ceph services restricted to `ceph_firewall_trusted_networks`.
- Rate-limited logging of dropped packets for diagnostics.
- Optional iSCSI and NFS blocks gated by boolean variables.
The ruleset is rendered from inventory variables (port numbers, trusted
networks, feature flags), making it consistent across nodes and clusters
without manual rule management.
## Consequences
- **Positive:** Single `inet` table covers both IPv4 and IPv6. No dual-stack
rule duplication.
- **Positive:** Template-driven rules are version-controlled, auditable, and
consistent across all nodes. Adding a new service means adding one variable
and one template block.
- **Positive:** `flush ruleset` at the top ensures convergence -- re-applying
the template replaces the entire ruleset atomically.
- **Negative:** Operators familiar only with `iptables` syntax need to learn
nftables. Mitigated by the template being well-commented and the Debian 12
ecosystem defaulting to nft.
@@ -0,0 +1,65 @@
# ADR-008: debootstrap from Live Image over Preseed/Autoinstall
## Status
Accepted
## Context
Austin nodes are bare-metal Dell servers that need Debian 12 installed from
scratch. The standard approaches are:
- **Preseed/Autoinstall** -- unattended Debian installer driven by a preseed
file. Requires PXE boot infrastructure or a custom ISO. The installer is a
black box: partition layout is expressed in preseed's declarative syntax,
which cannot handle our complex disk layout (mdraid-1 across two SSDs,
5 partitions per SSD for ESP/boot/swap/root/ceph-db, plus SSD OSD
partitions).
- **debootstrap from a live image** -- boot a Debian live USB, run
`debootstrap` to install the base OS into a prepared mount point. Full
scripting control over partitioning, mdraid, LVM, and chroot configuration.
Our disk layout requires:
1. Detecting exactly 2 SSDs by model pattern, validating they are non-rotational
and above a minimum size.
2. GPT partitioning with 5-6 partitions per SSD (ESP, boot, swap, root LVM,
ceph block.db, ceph SSD OSD).
3. mdraid-1 mirrors across matching partitions on both SSDs.
4. LVM on the root mdraid for flexible volume management.
Preseed cannot express step 1 (hardware validation with abort-on-failure) or
step 2 (partition 5-6 reserved for Ceph with specific sizes). Custom
partitioning in preseed uses `partman` recipes, which are notoriously fragile
and poorly documented for non-standard layouts.
## Decision
We boot nodes from a Debian 12 live image (USB stick via iDRAC virtual media),
then run the `provision_host` role via Ansible. The role:
1. Validates the live-image environment (UEFI, correct SSDs, sizes, non-rotational).
2. Partitions, creates mdraid arrays, sets up LVM -- all via shell commands
with full error handling and idempotency guards.
3. Runs `debootstrap --arch amd64 bookworm /mnt` to install the base OS.
4. Templates hostname, hosts, network, fstab, mdadm.conf inside the chroot.
5. Installs packages, creates the `ansible-iac` user, installs GRUB.
6. Writes a provisioning marker, mirrors the ESP, and reboots.
The entire process has block/rescue error handling: on failure, bind mounts
are cleaned up so the next run starts clean. A marker file makes completed
provisioning idempotent -- re-running skips directly to unmount and reboot.
## Consequences
- **Positive:** Full scripting control over disk layout. Complex partitioning,
mdraid, and multi-purpose SSD partitions are straightforward shell commands,
not preseed recipes.
- **Positive:** Hardware validation before any destructive operation. The role
aborts if SSDs don't match expectations (wrong model, too small, rotational).
- **Positive:** Idempotent with marker-driven resume. Partially failed runs
can be safely retried.
- **Negative:** Requires a live-image boot mechanism (iDRAC virtual media or
physical USB). Cannot provision purely over the network without PXE.
- **Negative:** More Ansible code to maintain than a preseed file. Justified
by the disk layout complexity that preseed cannot express.
@@ -0,0 +1,112 @@
# ADR-009: TF-First Authority + `op inject` over ansible-vault
## Status
Accepted (2026-04-22). Supersedes the implementation portions of [ADR-005](./005-1password-over-hashicorp-vault.md);
the "1Password over HashiCorp Vault" decision itself stands.
## Context
ADR-005 landed `vault-password.sh` (4-tier auth fallback) + `ansible-vault`
+ committed-encrypted `vault.yml` + `secrets-sync.sh` as the secrets
mechanism. Over time this accumulated friction:
1. **Silent-failure footgun**: `vault-password.sh` fell back to a dummy
password when 1Password was unavailable and no TTY was present. This
masked real auth failures — plays ran with wrong secrets until a task
that used them downstream blew up with a confusing error.
2. **Two source-of-truth problem**: secrets lived both in 1Password (master)
and in `vault.yml` (cache). `secrets-sync.sh` reconciled them but the
reconciliation step was easy to forget.
3. **Cluster identity authority**: `assign-names.py`, `vault-password.sh`,
`secrets-init.sh` all derived cluster identity from the inventory
directory path — operator-authored. Naming drift between scripts was a
recurring gotcha.
4. **Move to yucca monorepo**: migrating `sietch-ceph-dev-austin-int-futo-cloud`
into the `immich-apps/yucca` monorepo made the existing ad-hoc bash
scripting look out of place next to the Immich devtools pattern already
in use for other Futo infra (TF + 1P, `op run --env-file`, terragrunt).
5. **Painbox reprovision** surfaced that the Ansible layer's implicit
cluster-identity (host names, inventory paths) was hard to change — TF
as the authority makes identity a declarative input.
## Decision
**TF is the authority for cluster identity and 1P-item lifecycle; Ansible
is a consumer.** Specifically:
1. **Cluster + host identity** declared in `tf/deployment/<env>/<stack>/clusters.auto.tfvars`.
TF renders `inventory.ini`, `inventory-destroy.ini`, `inventory-provision.ini`,
and `secrets.yml.tpl` per cluster via `templatefile()` + `local_file`
resources.
2. **1P items** live in the `yucca_tf_*` team-shared vaults. Item names
follow `<CLUSTER>_CEPH_<ROLE>_*` (SHOUTY_SNAKE_CASE, `CEPH` hardcoded —
project-scoped, not hostname-role-scoped). TF-managed `onepassword_item`
resources are currently **dormant** in `secrets.tf.disabled` — items are
created today via the `op` CLI (operator runs `op item create` once per
cluster; see `docs/adding-a-cluster.md` §6). Re-enable once the
dedicated ceph-scoped service account replaces the org-wide superuser SA.
3. **Ansible consumption** via `scripts/ansible-play.sh`, which:
- Verifies `op account get` succeeds (fails closed).
- Renders the cluster's `secrets.yml.tpl` (op:// references → real
values) via `op inject -f` into a `0600` tmpfile cleaned by `trap`.
- Execs `ansible-playbook --extra-vars @<tmpfile>`.
4. **Service-account split**: superuser SA writes items (TF); read-only SA
reads items (Ansible runtime + CI). Both SAs live as 1P items in
`yucca_tf_dev`; `tf/.env` holds their `op://` references — resolved at
invocation time via `op run --env-file=tf/.env -- terragrunt ...`.
5. **Terragrunt multi-env** layout (`tf/deployment/<env>/<stack>/` with root
`terragrunt.hcl`). State backend is S3 against the shared `yucca-tf-state`
bucket at OVH Paris, keyed `ceph/${env}/${stack}/terraform.tfstate` —
project-scoped so ceph state doesn't collide with o11y or future stacks.
State locking (`use_lockfile`) deferred — OVH has no DynamoDB equivalent,
and OpenTofu's lockfile-object mode fails on fresh backends with a 404
before it can create one. Single-operator workflow today; revisit when
concurrent applies become likely.
## Consequences
- **Positive:** No ansible-vault. No `vault.yml`. No `vault-password.sh`.
No `secrets-sync.sh`. No `secrets-init.sh`. No dummy-password fallback.
No reconciliation step. Rotation = edit 1P item (or `terragrunt taint`),
next play run picks it up.
- **Positive:** Cluster identity is declarative. Adding a cluster or host is
an HCL edit + `tofu apply`. TF renders everything Ansible needs to
consume.
- **Positive:** CI lint and syntax-check don't require 1P access — ansible
parsing doesn't hit secrets until play-execution.
- **Positive:** Aligns with Immich devtools conventions — same `op run
--env-file=` pattern, same TF+1P module shape, same vault naming
structure (yucca_tf_*).
- **Negative:** `op` CLI must be available on every control node; no
offline/airgap fallback. Operationally acceptable.
- **Negative:** TF state now gates Ansible runs — rendered files must exist
before `ansible-playbook` has inputs. Mitigated by running `tofu apply`
as the one-time bootstrap step (gitignored outputs re-render on demand).
- **Negative:** Two service-account tokens to manage (read-only + superuser).
Shared with o11y and other Futo consumers via `yucca_tf_dev` — rotation
coordination required (see `docs/runbooks/rotate-sa-token.md`).
## Future direction
The TF-renders-Ansible-consumes pattern this ADR establishes for inventory
and secrets extends naturally to cephadm service specs. [ADR-011](./011-cephadm-osd-service-specs.md)
takes the first step (OSDs) — `osd-spec.yml.j2` is currently rendered by
Ansible from per-host data. A logical next step ("Option C" in the
architecture session log) is to move spec rendering into TF itself, so
TF emits `inventory.ini` + `secrets.yml.tpl` + the full set of cephadm
service specs (host registration, MON/MGR placement, OSD, RGW, monitoring)
as a coherent set of artifacts. The Ansible role becomes a thin applier:
`ceph orch apply -i <each-spec>`. Tracked as a follow-up PR after the
import lands.
## References
- `tf/shared/modules/ceph-cluster/` — the module that renders inventories
and (when re-enabled) provisions 1P items.
- `tf/deployment/dev/ceph/` — the dev-env declaration.
- `scripts/ansible-play.sh` — the Ansible-consumer wrapper.
- `docs/secrets.md` — end-user documentation of the model.
- [ADR-011](./011-cephadm-osd-service-specs.md) — first cephadm-spec
refactor (OSD path).
- [Immich devtools](https://github.com/immich-app/devtools) `tf/shared/modules/secrets/` — upstream pattern we copied.
@@ -0,0 +1,126 @@
# ADR-010: SSH Keys in 1Password (Native Category, Forward-Only Rotation)
## Status
Accepted (2026-04-22). Complements [ADR-009](./009-tf-first-op-inject-over-vault-password-sh.md)
(TF-first secrets) by applying the same "1P is the storage layer" model to
SSH keys.
## Context
Before this ADR, the `ansible-iac` SSH keypair for each cluster lived only on
operator workstations — `~/.ssh/id_ed25519_sietch-ceph` and
`~/.ssh/id_ed25519_ceph-painbox-lab-hel-htz`. Implications:
1. **No durable storage** — a laptop loss meant the private key was
gone; re-access required provisioning a new key and manually
distributing the pubkey to every host.
2. **No clean onboarding path** — adding a new operator required
out-of-band key sharing (insecure) or cutting them a separate key
pair (possible but undocumented).
3. **Remote-hands operators** had no scripted path to get cluster SSH
access beyond an out-of-band share of the private key.
4. **Inconsistent with the rest of the secrets model** — passwords live
in `yucca_tf_dev`, but SSH keys lived only on disk. Operators had to
remember two different security boundaries.
1Password has native `SSH Key` category support: items store the private
key, automatically derive the public key, and integrate with 1P SSH Agent
on workstations (already configured globally on the primary operator's
machine — `Host * IdentityAgent ~/.1password/agent.sock` in
`~/.ssh/config`).
## Decision
SSH keys join the hybrid secrets model:
1. **New keypairs are generated natively in 1P** via
`op item create --category "SSH Key" --ssh-generate-key=ed25519`. Private
key never touches operator disk during generation.
2. **Items live in `yucca_tf_dev`** with title
`<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY` (SHOUTY_SNAKE_CASE, same scheme
as password items).
3. **Public keys are consumed by Ansible at play time** via
`op read "op://yucca_tf_dev/<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY/public_key"`
in a new `rotate-ssh-key.yml` playbook, which ensures the current
pubkey is present in `ansible-iac@<host>:~/.ssh/authorized_keys`
(additive — doesn't remove others).
4. **Private keys reach operator workstations** via
`scripts/install-ssh-keys.sh` (idempotent `op read` → `~/.ssh/`,
refuses to overwrite on fingerprint mismatch).
5. **Rotation is forward-only**: new keys are generated, distributed,
verified; old keys are cleaned up manually (SSH in, delete from
authorized_keys, delete disk files) after confidence.
## Out of scope (for this ADR)
- **TF-managed SSH key resources.** Attempted earlier but op CLI's
import path for SSH Keys produces malformed items (accepts
private_key on create, strips the field on retrieval). Verified
against op `2.34.0`. Generation via `--ssh-generate-key` works
correctly. Future ADR may move generation into TF via
`onepassword_item` with `tls_private_key` — but that path puts the
private key into TF state, which is a step backwards for this
particular category.
- **1Password SSH Agent as the primary transport.** Operator workstations
may already have it globally configured (recommended), but ansible
inventory still references explicit `ansible_ssh_private_key_file`
paths. A future change could
drop the path entirely and rely on `IdentityAgent` routing — defer
until we've run under the new keys long enough to trust the agent
path.
- **CI integration.** No CI pipeline runs Ansible yet. When that
lands, CI would authenticate with the read-only SA, `install-ssh-keys.sh`
pulls keys into a CI-scoped `~/.ssh/`. Separate PR.
## Consequences
- **Positive:** keys are durable (laptop loss is a non-event), rotatable
(regenerate in 1P via `op item edit --ssh-generate-key`), auditable
(1P logs reads), and scoped (read-only SA can pull them, superuser SA
can rotate them).
- **Positive:** remote-hands onboarding is a documented `scripts/install-ssh-keys.sh`
invocation — no insecure key-over-chat.
- **Positive:** consistent with ADR-009. Operators have one mental
model: `op://<vault>/<ITEM>/<field>`, whether it's a password or a
key.
- **Negative:** op CLI can't import existing SSH keys reliably into the
`SSH Key` category (bug in 2.34.0). Rotation is the only way in —
existing keys are retired, not migrated. Operationally fine; makes
"put my personal bastion key in 1P" harder.
- **Negative:** ansible `rotate-ssh-key.yml` depends on `op` running on
the controller (already required by `scripts/ansible-play.sh`). Not
a new dependency, but does mean the key-rotation playbook can't run
offline.
- **Neutral:** private keys exist on operator disks during the
transition between `install-ssh-keys.sh` and adoption of 1P SSH
Agent as the sole transport. Files are `0600` and same security
properties as the pre-ADR state.
## Migration that landed with this ADR
- New items `SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY` and
`PAINBOX_CEPH_ANSIBLE_IAC_SSH_KEY` generated in `yucca_tf_dev`.
- `clusters.auto.tfvars` `ansible_ssh_key` and
`inventories/<cluster>/group_vars/all/vars.yml` `provision_iac_ssh_key_path`
updated to `~/.ssh/id_ed25519_sietch` and `~/.ssh/id_ed25519_painbox`.
- Rendered `inventory.ini` (all variants) and `secrets.yml.tpl` reflect
the new key paths after `tofu apply`.
- `scripts/install-ssh-keys.sh` added for operator-side key install.
- `rotate-ssh-key.yml` playbook added for pubkey distribution.
- End-to-end rotation verified against sietch: `install-ssh-keys.sh`
pulled the new keypair to the operator workstation,
`rotate-ssh-key.yml` distributed the public key to all three nodes'
`ansible-iac:~/.ssh/authorized_keys` (exclusive: false — additive),
and a subsequent `ansible -m ping` against the cluster succeeded
using the new key. Old per-operator keypairs
(`id_ed25519_sietch-ceph`, `id_ed25519_ceph-painbox-lab-hel-htz`) are
retired — delete from disk at operator convenience. Painbox's old
key is irrelevant post-reprovision.
## References
- [ADR-009](./009-tf-first-op-inject-over-vault-password-sh.md) — companion, TF-first secrets
- `scripts/install-ssh-keys.sh` — the operator-side tool
- `rotate-ssh-key.yml` — the host-side distribution playbook
- [1Password SSH Agent docs](https://developer.1password.com/docs/ssh/agent/)
@@ -0,0 +1,146 @@
# ADR-011: Cephadm OSD Service Specs over Imperative Per-Disk Loops
## Status
Accepted (2026-04-26). Refines [ADR-002](./002-explicit-osd-mapping.md)
(explicit OSD-to-disk mapping) — preserves its anti-auto-discovery spirit
while changing *where* the explicit mapping is expressed (cephadm spec
applied via `ceph orch apply`, not an Ansible per-disk loop).
## Context
`roles/ceph_deploy/tasks/osds.yml` originally created OSDs by iterating
each declared HDD/SSD in host_vars and invoking `cephadm ceph-volume lvm
create --dmcrypt --data <disk> --block.db <lv>` per disk. The disk path
was composed inside the role:
```jinja2
DISK="/dev/disk/by-path/{{ sas_path_prefix }}-{{ item.path_phy }}-lun-0"
```
This worked for sietch (Dell R730xd, SAS expander, dual-SSD-VG topology)
because every assumption baked into the composition was sietch-shape:
- `sas_path_prefix` exists per host
- `path_phy` is a slot identifier (`phy0`, `phy1`, ...) appended to the prefix
- `-lun-0` suffix is the SAS LUN convention
- SSD OSDs are partitions on the boot SSD (`-part6`)
When the painbox (Hetzner SX295) cluster came online for the first integration
test of this codebase, every one of those assumptions broke:
1. Painbox is SATA + NVMe, no SAS expander → `sas_path_prefix` undefined,
role failed at template-resolution time.
2. SATA by-path strings are `pci-XXXX:XX:XX.X-ata-N` directly (no `-lun-N`
suffix) — operator authors the full path identifier as `path_phy`.
3. Painbox SSD OSD is an LV on the NVMe RAID-1 VG (`vg0/ssd-osd`), not a
partition — different shape entirely from sietch's SSD OSDs.
The fix could have been per-shape branching inside `osds.yml`'s shell loop
(if/else for path composition; if/else for partition vs LV), but every
phase of the role beyond OSDs (RGW, monitoring, etc.) was already migrating
toward cephadm's declarative service-spec pattern (`rgw-spec.yaml.j2` was
already in place). The OSD path was the obvious next migration.
## Decision
OSD creation is now declarative via a cephadm OSD service spec rendered
from per-host data, applied via `ceph orch apply osd -i /etc/ceph/osd-spec.yml`.
1. **`templates/osd-spec.yml.j2`** renders one document per host (HDD spec
+ optional SSD spec). Per-host because sietch nodes have unique SAS
prefixes per chassis — a shared spec with `data_devices.paths` requires
identical paths across all `placement.hosts`.
2. **Hardware-shape branching** lives in the template's Jinja conditional,
not in role logic:
- `sas_path_prefix is defined` → sietch-shape path composition
- `sas_path_prefix is undefined` → painbox-shape, use `path_phy`
directly (full PCI-ATA identifier from host_vars)
- SSD OSD: `lv is defined` → painbox-style LV path; else sietch-style
partition-on-SAS-disk
3. **`tasks/osds.yml`** is now thin: render template → `ceph orch apply` →
wait for provisioning → wait for OSDs up → defensive `ceph osd unset
noin` → reweight-zero safety net. ~150 LOC down from ~230.
4. **Encryption** is set in the spec (`encrypted: true`), not via a
per-disk `--dmcrypt` flag. Cephadm provisions LUKS internally and
stores the dm-crypt key in the MON config-key store as before.
5. **Idempotency** is provided by cephadm — re-applying the same spec is
a no-op. New disks (e.g. populating an empty bay) are picked up
automatically. Existing OSDs are not destroyed by a spec change;
removal requires `ceph orch osd rm`.
## Out of scope (for this ADR)
- **Extending the spec pattern to MON/MGR/RGW/monitoring placement**
— `rgw-spec.yaml.j2` already does this for RGW; the rest of the role
still uses imperative `ceph orch apply mon --placement=...` (and
similar for MGR/monitoring). Refactoring those is a separate body of
work — see "Option C" in the architecture session log; planned as a
dedicated follow-up PR.
- **TF-rendered service specs.** TF already has all the cluster identity
it needs to render every cephadm spec deterministically. Moving spec
rendering from Ansible (Jinja in role templates) to TF (Jinja in
module templates) would shift the boundary further toward
"TF declares, Ansible applies." Same Option-C follow-up.
- **Auto-discovery via `data_devices.rotational: 1` filter.** Cephadm
supports filter-based device selection. We chose explicit `paths`
to preserve ADR-002's "no auto-discovery surprises" stance — empty
bays on sietch and the SSD OSD on painbox can't be expressed with
filters alone.
## Consequences
- **Positive:** hardware-shape independence. Painbox (SATA + NVMe-RAID
+ LV-backed SSD OSD) and sietch (SAS expander + dual-SSD-VG) deploy
through the same role with the same task file. Future clusters with
yet other shapes (Hetzner AX-line, Equinix bare-metal, etc.) need
only their own host_vars; the role doesn't change.
- **Positive:** `osds.yml` shrinks ~35%; the deleted code was the most
fragile part (per-disk shell loops with embedded Jinja path
composition).
- **Positive:** ADR-002's explicit-mapping spirit is preserved. Paths
are still listed explicitly in the spec — cephadm doesn't auto-discover.
Empty bays stay empty; partitions stay reserved for OS/block.db.
- **Positive:** Aligns with cephadm's intended deployment model. All
modern cephadm operators use service specs; the imperative-loop
pattern is legacy.
- **Neutral:** The spec is rendered to `/etc/ceph/osd-spec.yml` on the
bootstrap node and then applied. The file is overwritten on each
re-render — operators inspecting the deployed state should query
cephadm (`ceph orch ls --service-type osd --export`) rather than
reading the on-disk spec, which may have been updated since the last
apply.
- **Negative:** Cephadm's spec apply is async — the role polls until
the expected OSD count is reached, with a 15-minute timeout. On a
large cluster (hundreds of OSDs), this could exceed the timeout. Not
a concern at current scale (15 OSDs painbox, 30 OSDs sietch).
- **Negative:** Debugging "why isn't this disk becoming an OSD?" is
harder than with the per-disk loop, where the failing disk had its
own log line. With cephadm, you check `ceph orch ls`,
`ceph orch ps`, and `ceph cephadm osd activate <host> --dry-run` on
the bootstrap node.
## Migration that landed with this ADR
- `templates/osd-spec.yml.j2` written (replaces stub that existed but
was never wired in).
- `tasks/osds.yml` rewritten to render + apply + poll + safety net.
- `tasks/lvm-setup.yml` gated on `sas_path_prefix is defined` — sietch
still needs the role's defensive VG/LV recovery path; painbox skips
because installimage's post-install owns LVM lifecycle.
- Validated end-to-end against a freshly-reprovisioned painbox: 15/15
OSDs up + in, all encrypted (15 dm-crypt keys in MON store), correct
per-host service IDs (`osd.painbox-ceph-evelyn-hdd`,
`osd.painbox-ceph-evelyn-ssd`).
- Sietch validation deferred — sietch is currently deployed and
healthy; the spec apply against existing OSDs is idempotent (no-op
when paths match), but a real test on sietch is a separate operational
step planned for the next sietch maintenance window.
## References
- [ADR-002](./002-explicit-osd-mapping.md) — explicit OSD-to-disk mapping (refined, not superseded)
- [ADR-009](./009-tf-first-op-inject-over-vault-password-sh.md) — TF-first secrets (companion architectural shift)
- `roles/ceph_deploy/templates/osd-spec.yml.j2` — the template
- `roles/ceph_deploy/tasks/osds.yml` — the thin applier
- [Cephadm OSD Service docs](https://docs.ceph.com/en/latest/cephadm/services/osd/)
+825
View File
@@ -0,0 +1,825 @@
# Architecture
How the Ceph automation in `yucca/ansible/ceph/` is shaped, what each tool
owns, and how the four tools (Terraform, 1Password, Ansible, mise) hand off
work to each other.
This is a structural reference. For step-by-step usage see
[CONTRIBUTING.md](../CONTRIBUTING.md); for narrower topics see the
specialized docs under `docs/`.
---
## 1. System context
```mermaid
flowchart LR
OP([Operator workstation<br/>mise · op CLI · ansible · tofu])
YUCCA[/Yucca monorepo<br/>tf/ + ansible/ceph/ + kubernetes//]
ONEP[("1Password org<br/>yucca_tf · yucca_tf_dev · ...")]
S3[("OVH S3<br/>yucca-tf-state bucket")]
SIETCH["Sietch · Austin DC<br/>3× Dell R730xd"]
PAINBOX["Painbox · Hetzner Helsinki<br/>1× SX295"]
OP -->|edits| YUCCA
OP -->|reads/writes secrets| ONEP
OP -->|TF state I/O| S3
OP -->|SSH ansible-iac| SIETCH
OP -->|SSH ansible-iac| PAINBOX
```
The yucca monorepo is the single source of truth for cluster identity and
configuration. Operators run mise tasks on their workstation; secrets stay
in 1Password (never on disk); TF state lives in OVH S3; Ansible drives
configuration over SSH against bare-metal Ceph nodes.
External dependencies are minimal and explicit:
- **1Password org** — organization-scoped vaults shared with other Futo infra
(Immich, o11y). Authoritative store for live secret values.
- **OVH S3** — `yucca-tf-state` bucket at `s3.eu-west-par.io.cloud.ovh.net`.
Project-keyed (`ceph/<env>/<stack>/terraform.tfstate`) so multiple stacks
share the bucket without collision.
- **Hardware** — Austin colo for sietch (Dell R730xd × 3, single 10G bond);
Hetzner Helsinki for painbox (SX295 × 1). Detail in [hardware.md](hardware.md).
---
## 2. Environments
Environments are a first-class concern: every tool in the mesh derives
its environment from the same source — directory layout — so dev / staging
/ prod isolation is structural, not flag-driven.
| Layer | dev (today) | staging (planned) | prod (planned) |
|--------------|----------------------------------------------------------------|---------------------------------------------------------|------------------------------------------------------|
| TF stack dir | `tf/deployment/dev/ceph/` | `tf/deployment/staging/ceph/` | `tf/deployment/prod/ceph/` |
| TF state key | `ceph/dev/ceph/terraform.tfstate` | `ceph/staging/ceph/terraform.tfstate` | `ceph/prod/ceph/terraform.tfstate` |
| 1P vaults | `yucca_tf_dev` (live) · `yucca_tf_dev_manual` (human-fillable) | `yucca_tf_staging` · `yucca_tf_staging_manual` | `yucca_tf` (live) · `yucca_tf_prod_manual` |
| Ansible inv | `inventories/<cluster>-ceph.dev.<dc>.<provider>/` | `inventories/<cluster>-ceph.staging.<dc>.<provider>/` | `inventories/<cluster>-ceph.prod.<dc>.<provider>/` |
| mise default | `CEPH_ENV=...sietch-ceph.dev.austin.int/inventory.ini` | overridden via env at invocation | overridden via env at invocation |
Today the only deployed environment is dev (sietch + painbox). Adding
staging/prod is purely additive: create the matching `tf/deployment/<env>/ceph/`
directory, populate `clusters.auto.tfvars`, and the same module + Ansible
roles + mise tasks work unchanged. The state backend key path, 1P vault
selection, and inventory directory naming all derive from the env segment.
`TF_STACK_DIR` is the operator-side override for `mise run tf:*` tasks; it
defaults to `tf/deployment/dev/ceph` and points at any sibling stack directory.
`CEPH_ENV` is the matching override for Ansible — points at the rendered
`inventory.ini` for the cluster you intend to operate on.
---
## 3. The four-tool mesh
```mermaid
flowchart TB
subgraph ws["Operator workstation"]
direction LR
MISE([mise<br/>orchestration])
TF[Terraform / Tofu<br/>via Terragrunt]
OP[op CLI]
ANS[Ansible]
WRAP[scripts/<br/>ansible-play.sh<br/>install-ssh-keys.sh]
end
ONEP[("1Password<br/>yucca_tf_*")]
S3[("OVH S3<br/>tfstate")]
REPO[/"Yucca repo<br/>inventories/<cluster>/<br/>(host_vars committed,<br/>TF outputs gitignored)"/]
NODES[Ceph nodes]
MISE -->|tf:*| TF
MISE -->|deploy / status / drift| WRAP
MISE -->|capture| WRAP
TF -->|reads SA token via op run --env-file| OP
TF -->|reads/writes state| S3
TF -->|renders| REPO
WRAP -->|reads inventory + secrets.yml.tpl| REPO
WRAP -->|op inject / op read| OP
WRAP -->|exec| ANS
ANS -->|SSH ansible-iac| NODES
OP <-->|item CRUD| ONEP
```
### Who owns what
| Tool | Owns | Reads from |
|---------------|--------------------------------------------------------------------------------|---------------------------------------------|
| **Terraform** | Cluster identity, host names, rendered Ansible artifacts, TF state | `clusters.auto.tfvars` · 1P (via op CLI) |
| **1Password** | Live secret values, SSH keypairs, service-account tokens | nothing — authoritative store |
| **op CLI** | Auth and resolution: env-injection, file-template injection, single-value read | 1P (session or SA token) |
| **Ansible** | Convergence: applying configuration to nodes | Rendered inventory + op-injected tmpfile |
| **mise** | Task discovery, toolchain pinning, env defaults | `.mise.toml`, `tf/.env` |
### Handoff points (the edges of the mesh)
1. **TF → repo** — `terragrunt apply` renders `inventory.ini`,
`inventory-destroy.ini`, optional `inventory-provision-<profile>.ini`,
and `secrets.yml.tpl` into `inventories/<cluster>/`. These files are
gitignored — the source of truth is `clusters.auto.tfvars` + the
ceph-cluster module.
2. **TF ↔ op CLI** — TF runs are wrapped with `op run --env-file=tf/.env`,
which resolves `op://...` references in `tf/.env` and injects them as
`OP_SERVICE_ACCOUNT_TOKEN`, `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`
for the child process. The `tf/.env` file is committed (it contains only
pointers, never literal secrets).
3. **Ansible ↔ op CLI** — `scripts/ansible-play.sh` reads the cluster's
`secrets.yml.tpl`, runs `op inject -f` to resolve `op://` references into
a `mktemp`'d tmpfile (chmod 600, trap-cleaned), then execs
`ansible-playbook --extra-vars @<tmpfile>`. The tmpfile lives only for
the duration of the play.
4. **mise → wrappers** — `mise run deploy` invokes
`scripts/ansible-play.sh deploy.yml ...`; `mise run tf:*` invokes
`op run --env-file=tf/.env -- terragrunt --working-dir <stack> <cmd>`.
mise tasks never call `ansible-playbook` directly.
---
## 4. Terraform (authority)
### Layout
```
tf/
├── .env op:// references (committed; no literal secrets)
├── shared/modules/ceph-cluster/ per-cluster orchestration module
│ ├── main.tf · variables.tf · outputs.tf · rendering.tf
│ ├── wordlist.txt 923 words for auto-picked hostnames
│ └── templates/ inventory + secrets.yml.tpl templates
└── deployment/
├── terragrunt.hcl root: state backend, env/stack derived from path
└── dev/ceph/
├── terragrunt.hcl includes root, sets ansible_project_root
├── main.tf · variables.tf · versions.tf
└── clusters.auto.tfvars declarative cluster list
```
### Cluster identity is declared, not derived
The `clusters` map in `clusters.auto.tfvars` is the source of truth.
Each top-level key becomes a cluster:
```hcl
sietch = {
domain = "dev.austin.int.futo.cloud"
environment = "dev"
datacenter = "austin"
provider_code = "int"
role_in_hostname = "ceph"
ansible_ssh_user = "ansible-iac"
ansible_ssh_key = "~/.ssh/id_ed25519_sietch"
vault = "yucca_tf_dev"
provision_profile = "debian-live"
hosts = [
{ name = "laurel", bond_ip = "10.10.10.90", bootstrap = true },
{ name = "lawson", bond_ip = "10.10.10.91" },
{ name = "samara", bond_ip = "10.10.10.92" },
]
}
```
The module computes everything else: hostname (`<cluster>-<role>-<name>`),
FQDN (`<hostname>.<domain>`), 1P item names
(`<CLUSTER>_CEPH_<ROLE>_PASSWORD`), inventory directory path
(`inventories/<cluster>-<role>.<env>.<dc>.<provider>/`).
### Auto-naming via wordlist
For hosts where `name = null`, the module picks a stable name from a
923-word pool using `random_shuffle` seeded by `(cluster_name, name_seed)`.
Operator-declared names are excluded from the pool to prevent collisions
within a cluster. Adding hosts at the tail is safe — existing positions
keep their names across applies.
Painbox today demonstrates this: its single host has no `name`, so TF
auto-picked `evelyn` → hostname `painbox-ceph-evelyn`.
### Rendered artifacts (gitignored)
Per `tf/shared/modules/ceph-cluster/rendering.tf`, the module writes four
files into `inventories/<cluster>/`:
| File | Purpose |
|--------------------------------------------|------------------------------------------------------------------------------------------|
| `inventory.ini` | Normal-ops inventory: `ansible-iac` user + cluster SSH key |
| `inventory-destroy.ini` | Destroy-mode inventory (same credentials; separate file as a speed bump) |
| `inventory-provision-<profile>.ini` | Provisioning inventory (only when `provision_profile != null`; uses live-image creds) |
| `secrets.yml.tpl` | `vault_*: op://<vault>/<CLUSTER>_CEPH_*/password` pointers, consumed by `op inject -f` |
All four are in `ansible/ceph/.gitignore`. Re-render with `mise run tf:apply`.
### State backend
S3 backend in `tf/deployment/terragrunt.hcl`:
- Bucket: `yucca-tf-state` (shared with o11y and other Futo stacks)
- Region: `eu-west-par` (OVH Paris)
- Endpoint: `https://s3.eu-west-par.io.cloud.ovh.net/`
- Key: `ceph/${env}/${stack}/terraform.tfstate` — derived from
the child stack's path under `deployment/`
- Skip AWS-specific validation; use path-style URLs (OVH compatibility)
State locking is **not enabled today**. OVH has no DynamoDB equivalent.
OpenTofu's `use_lockfile = true` would work but expects the lockfile object
to already exist — fresh-backend init fails with 404 before it can create
one. Single-operator workflow today; revisit when concurrent applies become
likely. See `deployment/terragrunt.hcl` for the inline rationale.
### What TF does not yet manage
`onepassword_item` resources are **dormant** (`tf/.../secrets.tf.disabled`).
1P items are created today via the `op` CLI (operator runs `op item create`
once per cluster). Re-enabling them is tracked in
[ADR-009](adr/009-tf-first-op-inject-over-vault-password-sh.md) — the gate
is the dedicated `sietch-ceph` service account that lets us split write
authority from the org-wide superuser SA.
---
## 5. 1Password (live values)
### Vault hierarchy
| Vault | Purpose | Who reads it | Who writes it |
|---------------------------|--------------------------------------------------------|------------------------------------|-------------------------------------|
| `yucca_tf` | Cross-env shared (TF state S3 creds) | TF (via op run --env-file) | Operator (manual) |
| `yucca_tf_dev` | dev environment live values | Ansible runtime (op inject) | Superuser SA (TF) + operator (op CLI) |
| `yucca_tf_dev_manual` | dev human-fillable placeholders (API tokens, OAuth) | Ansible runtime | Operator (manual) |
| `yucca_tf_staging(_manual)` · `yucca_tf_prod_manual` | (planned) staging/prod analogues | (future) | (future) |
The `_manual` vaults exist for items that can't be auto-generated (third-party
API tokens, OAuth client secrets) — they're populated by humans, not by TF.
### Service accounts
Two service accounts in `yucca_tf_dev`, both shared org-wide:
| SA | Scope | Consumed by |
|-----------------------------------------------|---------------------------------------------|-------------------------------------|
| `yucca_futo_1pass_superuser_service_account` | Read + write all `yucca_tf_*` vaults | TF (via `tf/.env`) + interactive op CLI |
| `yucca_futo_1pass_service_account` | Read-only on `yucca_tf` and `yucca_tf_dev` | Ansible runtime / future CI |
Rotation procedure: [docs/runbooks/rotate-sa-token.md](runbooks/rotate-sa-token.md).
### Item categories and naming
Per cluster, the following items live in the cluster's `vault` (currently
`yucca_tf_dev` for both sietch and painbox):
| Category | Item title pattern | Field consumed |
|------------|---------------------------------------------------------|----------------------|
| Password | `<CLUSTER>_CEPH_OPS_PASSWORD` | `password` |
| Password | `<CLUSTER>_CEPH_DASHBOARD_PASSWORD` | `password` |
| Password | `<CLUSTER>_CEPH_GRAFANA_PASSWORD` | `password` |
| Password | `<CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY` | `password` |
| Password | `<CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY` | `password` |
| SSH Key | `<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY` | `private_key` / `public_key` |
| Document | `<CLUSTER>_CEPH_RGW_TLS_CERT` · `..._RGW_TLS_KEY` | file content |
| Document | `<CLUSTER>_CEPH_CLIENT_ADMIN_KEYRING` | file content |
The `<CLUSTER>_CEPH_*` prefix is hardcoded in
`tf/shared/modules/ceph-cluster/main.tf` (`secret_prefix = "${upper(var.cluster_name)}_CEPH"`)
so every Ceph-project item across all clusters is grep-discoverable as
`*_CEPH_*` regardless of the role segment in node hostnames.
Full item-by-item catalog: [docs/secrets.md](secrets.md).
### Three op-CLI patterns
The op CLI is invoked in three distinct ways across the codebase. Each
serves a different shape of secret consumption:
1. **`op run --env-file=tf/.env -- <cmd>`** — env-var injection.
Resolves `op://` references in a dotenv file and injects the resolved
values as env vars into the child process. Used for TF (SA token) and
the S3 backend (AWS creds). Wrapped by all `mise run tf:*` tasks.
2. **`op inject -f -i <tpl> -o <out>`** — file-template resolution.
Reads a file containing inline `op://` references, resolves each, writes
to the output path. Used by `scripts/ansible-play.sh` to render
`secrets.yml.tpl` → tmpfile, and by the Hetzner installimage flow to
render `post-install.sh.tpl` → `post-install.sh`.
3. **`op read "op://<vault>/<item>/<field>"`** — single-value read.
Used by `scripts/install-ssh-keys.sh`, `rotate-ssh-key.yml`,
`post-deploy-capture.yml`. Returns one value to stdout for one specific
field; fails closed if missing.
No custom password-script (no `vault-password.sh`); no
`ansible-vault`-encrypted file in git. Lint and syntax-check tasks don't
invoke op at all — they don't need secrets, so "1P unavailable" never
silently degrades them. See [ADR-009](adr/009-tf-first-op-inject-over-vault-password-sh.md).
---
## 6. Ansible (consumer)
### Role dependency graph
Provisioning is a separate concern (`provision.yml`, sietch only). The main
pipeline (`site.yml`) runs everything else in this order:
```mermaid
flowchart TB
PROV["provision_host<br/><i>separate playbook, live image</i>"]
BASE["baseline<br/><i>users, packages, /etc/hosts</i>"]
OST["os_tuning<br/><i>sysctl, TCP buffers</i>"]
HWT["hardware_tuning<br/><i>I/O scheduler, readahead</i>"]
DEPLOY["ceph_deploy<br/><i>bootstrap, join, OSDs, RGW,<br/>crush rules, monitoring</i>"]
CTUNE["ceph_tuning<br/><i>recovery throttling, scrub<br/>window, telemetry, audit</i>"]
SEC["security<br/><i>nftables, SSH hardening</i>"]
PROV -.->|reboot into installed OS| BASE
BASE --> OST
BASE --> HWT
OST --> DEPLOY
HWT --> DEPLOY
DEPLOY --> CTUNE
CTUNE --> SEC
classDef separate stroke-dasharray: 4 4
class PROV separate
```
`site.yml` starts at `baseline` — `provision_host` runs only on first
install via `provision.yml`.
### Why this order matters
1. **baseline before tuning** — cephadm needs podman, dbus, chrony.
The baseline role installs these and enables the services. Running
tuning on a node without podman would leave cephadm unable to bootstrap.
2. **tuning before deploy** — OSD daemons inherit kernel parameters
active at startup. Applying sysctl (`vm.min_free_kbytes`, `fs.aio-max-nr`)
and I/O scheduler (`mq-deadline` for HDD, `none` for SSD) before bootstrap
means daemons launch with correct limits from the first second.
3. **ceph_tuning after deploy** — these settings use `ceph config set`
which requires a running cluster. Recovery throttling, scrub windows, and
PG autoscaler targets cannot be applied until MONs are up.
4. **security last** — nftables drops all traffic not explicitly allowed.
Running it before ceph_deploy would block cephadm's inter-node SSH,
container image pulls, and MON/OSD port negotiation. Once the cluster is
healthy, the firewall locks it down.
### ceph_deploy internal pipeline
`roles/ceph_deploy/tasks/main.yml` orchestrates ten phases:
```mermaid
flowchart TB
P1["Phase 1 · prerequisites.yml<br/><i>Ceph repo, cephadm, ceph-common</i>"]
P2["Phase 2 · bootstrap.yml<br/><i>cephadm bootstrap on first node</i>"]
P3["Phase 3 · join.yml<br/><i>ceph orch host add for remaining nodes</i>"]
P4["Phase 4 · placement.yml<br/><i>MON/MGR placement calculation</i>"]
P45["Phase 4.5 · lvm-setup.yml<br/><i>ensure block.db VGs/LVs exist (sietch-shape only;<br/>painbox skips — LVM owned by installimage post-install)</i>"]
P5["Phase 5 · osds.yml<br/><i>render osd-spec.yml.j2 → ceph orch apply osd<br/>(cephadm provisions LUKS + LVM internally)</i>"]
P55["Phase 5.5 · crush-rules.yml<br/><i>replicated_hdd / replicated_ssd rules</i>"]
P575["Phase 5.75 · rgw.yml<br/><i>EC pools, realm/zone, TLS, S3 user</i>"]
P58["Phase 5.8 · monitoring.yml<br/><i>dashboard URL integration, Grafana creds</i>"]
P6["Phase 6 · verify.yml<br/><i>cluster health report</i>"]
P1 --> P2 --> P3 --> P4 --> P45 --> P5 --> P55 --> P575 --> P58 --> P6
```
Tag-driven re-runs are first-class:
`scripts/ansible-play.sh deploy-ceph.yml --tags rgw,monitoring` re-runs
just those phases.
### Inventory layout
```
inventories/
sietch-ceph.dev.austin.int/ Austin dev cluster
inventory.ini TF-generated, gitignored
inventory-destroy.ini TF-generated, gitignored
inventory-provision.ini TF-generated, gitignored
secrets.yml.tpl TF-generated, gitignored
group_vars/all/vars.yml cluster-wide variables (committed)
host_vars/ per-node hardware topology (committed)
sietch-ceph-laurel.yml bond_ip, SAS path prefix, OSD maps
sietch-ceph-lawson.yml
sietch-ceph-samara.yml
installimage/ Hetzner installimage assets (sietch n/a)
painbox-ceph.dev.hel.htz/ Hetzner dev cluster
inventory.ini TF-generated, gitignored
inventory-destroy.ini TF-generated, gitignored
secrets.yml.tpl TF-generated, gitignored
group_vars/all/vars.yml committed
host_vars/painbox-ceph-evelyn.yml committed
installimage/ post-install.sh.tpl (op-injected)
```
`host_vars/*.yml` is committed because per-node hardware facts (bond_ip, SAS
expander paths, SSD PHY positions, HDD-to-block.db mappings) are stable
inventory truth — not operator preference. The `.local.yml` suffix is
gitignored as an escape hatch for operator-local overrides.
### Variable precedence
```mermaid
flowchart TB
D["<b>role defaults</b><br/>roles/*/defaults/main.yml<br/><i>lowest priority</i>"]
G["<b>group_vars</b><br/>inventories/&lt;cluster&gt;/group_vars/all/vars.yml"]
H["<b>host_vars</b><br/>inventories/&lt;cluster&gt;/host_vars/&lt;host&gt;.yml"]
T["<b>extra-vars @tmpfile</b><br/>scripts/ansible-play.sh<br/><i>op-injected secrets</i>"]
E["<b>extra-vars -e X=Y</b><br/>-e confirm_wipe=true<br/><i>highest priority</i>"]
D --> G --> H --> T --> E
```
- **Role defaults** define every tunable with a safe value
(`ceph_firewall_ssh_any_source: true`, `ceph_cpu_governor_enabled: false`).
- **group_vars/all/vars.yml** sets cluster-wide values: network topology,
Ceph release, RGW config, monitoring ports, plus the `vault_*` →
consumable-name aliases (`ops_password: "{{ vault_ops_password }}"`).
- **host_vars** provides per-node physical topology.
- **extra-vars from @tmpfile** carries op-injected `vault_ops_password`,
`vault_ceph_dashboard_password`, `vault_grafana_admin_password`,
`vault_s3_restic_access_key`, `vault_s3_restic_secret_key`.
- **extra-vars via `-e`** carries safety gates: `confirm_wipe=true`,
`provision_skip_reboot=true`, `yes_destroy_ceph=true`.
### ansible.cfg stays generic
`ansible.cfg` contains zero site-specific values. No default inventory, no
ProxyJump, no hardcoded key paths. Site-specifics live exclusively in
`clusters.auto.tfvars` (which TF renders into the inventory) or in the
inventory's `group_vars`. The same `ansible.cfg` and the same roles work
unchanged across Austin, Hetzner, or any future cluster — only the cluster
entry in `clusters.auto.tfvars` differs.
---
## 7. mise (orchestration surface)
### Why mise
- **Toolchain pinning** — `.mise.toml` declares the exact versions of
`python`, `tofu`, `terragrunt`, `op`. New operators get a working
environment with `mise trust && mise run setup`.
- **Task discovery** — `mise tasks` lists every operation; tasks are
shell-script-shaped, kept in `.mise.toml`, and committed.
- **Devtools parity** — matches the conventions in `immich-app/devtools`
(where the `op run --env-file=tf/.env --` pattern originated).
### Task taxonomy
| Group | Tasks |
|----------------|--------------------------------------------------------------------------|
| Bootstrap | `setup` |
| Verify | `lint` · `check` · `test` · `preflight` |
| TF (wrapped) | `tf:init` · `tf:plan` · `tf:apply` · `tf:destroy` |
| Read-only ops | `status` · `drift` |
| State change | `deploy` · `destroy` · `capture` · `backup` |
| Benchmarks | `bench` · `bench-rados` |
### How mise wraps the underlying CLIs
- `mise run tf:*` → `op run --env-file=tf/.env -- terragrunt --working-dir <stack> <cmd>`
- `mise run deploy` → `scripts/ansible-play.sh deploy-ceph.yml ...` (per phase)
- `mise run status` → `scripts/ansible-play.sh status.yml`
- `mise run capture` → `scripts/ansible-play.sh post-deploy-capture.yml`
mise never invokes `ansible-playbook` or `terragrunt` directly. The wrappers
own secrets injection and pre-flight checks; mise owns task discovery and
env defaults.
### Env defaults
```toml
[env]
CEPH_ENV = "inventories/sietch-ceph.dev.austin.int/inventory.ini"
```
`TF_STACK_DIR` defaults to `tf/deployment/dev/ceph` inside each `tf:*` task.
Both are overridable per-invocation:
```bash
TF_STACK_DIR=tf/deployment/staging/ceph mise run tf:plan
CEPH_ENV=inventories/painbox-ceph.dev.hel.htz/inventory.ini mise run status
```
---
## 8. Wrapper scripts (the glue layer)
Three scripts under `ansible/ceph/scripts/` sit between mise and the
underlying CLIs. They exist to keep secrets out of `argv`, fail closed
when 1P is unreachable, and give better error messages than the raw tools.
| Script | Purpose |
|-------------------------|----------------------------------------------------------------------------------|
| `ansible-play.sh` | Render secrets via `op inject -f` to a `mktemp`'d file (chmod 600, trap-cleaned), then exec `ansible-playbook --extra-vars @<tmpfile>` |
| `install-ssh-keys.sh` | Idempotent `op read` → `~/.ssh/id_ed25519_<cluster>` installer; refuses overwrite on fingerprint mismatch |
| `preflight.sh` | Verifies TF artifacts present, 1P session live, SSH reachable, Python on targets — surfaced via `mise run preflight` |
Per-script reference (synopsis, args, env, exit codes, examples):
[docs/scripts.md](scripts.md).
---
## 9. Data flow: concrete operations
### 9.1 `mise run tf:apply` — render artifacts
```mermaid
sequenceDiagram
actor OP as Operator
participant MISE as mise
participant OPCLI as op CLI
participant ONEP as 1Password
participant TG as terragrunt / tofu
participant S3 as OVH S3
participant REPO as Repo (inventories/)
OP->>MISE: mise run tf:apply
MISE->>OPCLI: op run --env-file=tf/.env -- ...
OPCLI->>ONEP: resolve op:// references
ONEP-->>OPCLI: SA token + AWS keys
OPCLI->>TG: exec child process<br/>with env vars injected
TG->>S3: read tfstate<br/>(ceph/<env>/<stack>/terraform.tfstate)
S3-->>TG: current state
TG->>TG: plan + apply
TG->>S3: write updated tfstate
TG->>REPO: render inventory.ini,<br/>secrets.yml.tpl, ...
```
### 9.2 `mise run deploy` — full Ceph deploy
```mermaid
sequenceDiagram
actor OP as Operator
participant MISE as mise
participant WRAP as ansible-play.sh
participant OPCLI as op CLI
participant ONEP as 1Password
participant TMP as /tmp/<id>-secrets.yml
participant ANS as ansible-playbook
participant NODES as Ceph nodes
OP->>MISE: mise run deploy
MISE->>WRAP: ansible-play.sh deploy-ceph.yml
WRAP->>OPCLI: op account get
OPCLI-->>WRAP: session OK
WRAP->>TMP: mktemp + chmod 600 + trap rm
WRAP->>OPCLI: op inject -f -i secrets.yml.tpl -o TMP
OPCLI->>ONEP: resolve op://yucca_tf_dev/SIETCH_CEPH_*/password
ONEP-->>OPCLI: secret values
OPCLI->>TMP: write resolved YAML
WRAP->>ANS: exec --extra-vars @TMP
loop phases 1..6
ANS->>NODES: SSH ansible-iac@<bond_ip><br/>via id_ed25519_<cluster>
end
Note over WRAP,TMP: tmpfile rm'd on exit (trap)
```
### 9.3 `mise run capture` — DR snapshot
```mermaid
sequenceDiagram
actor OP as Operator
participant MISE as mise
participant ANS as ansible-playbook<br/>(post-deploy-capture.yml)
participant BOOT as Bootstrap node
participant LOCAL as localhost (delegated)
participant OPCLI as op CLI
participant ONEP as 1Password
OP->>MISE: mise run capture
MISE->>ANS: ansible-play.sh post-deploy-capture.yml
ANS->>BOOT: SSH read /etc/ceph/rgw-ssl.crt
ANS->>BOOT: SSH read /etc/ceph/rgw-ssl.key
ANS->>BOOT: SSH read /etc/ceph/ceph.client.admin.keyring
BOOT-->>ANS: file contents
loop for each artifact
ANS->>LOCAL: delegate_to: localhost
LOCAL->>OPCLI: op item edit/create<br/><CLUSTER>_CEPH_<ITEM>
OPCLI->>ONEP: upsert Document item<br/>in yucca_tf_dev
end
Note over ONEP: Now holds RGW_TLS_CERT,<br/>RGW_TLS_KEY, CLIENT_ADMIN_KEYRING
```
### 9.4 `scripts/install-ssh-keys.sh` — fresh workstation
```mermaid
sequenceDiagram
actor OP as Operator (new ws)
participant SCRIPT as install-ssh-keys.sh
participant OPCLI as op CLI
participant ONEP as 1Password
participant SSH as ~/.ssh/
OP->>SCRIPT: install-ssh-keys.sh sietch
SCRIPT->>OPCLI: op read .../public_key
OPCLI->>ONEP: SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY
ONEP-->>OPCLI: public_key
OPCLI-->>SCRIPT: pubkey content
SCRIPT->>SSH: compare with id_ed25519_sietch (if exists)
alt fingerprint match
SCRIPT-->>OP: skip (already present)
else fingerprint mismatch
SCRIPT-->>OP: refuse (operator must mv aside)
else file missing
SCRIPT->>OPCLI: op read .../private_key
OPCLI-->>SCRIPT: private_key
SCRIPT->>SSH: write id_ed25519_sietch (0600)<br/>+ .pub (0644)<br/>(umask 077)
end
```
---
## 10. OSD lifecycle
### Phase flow
```mermaid
flowchart TB
SIETCH["Sietch prep:<br/>provision_host/disks.yml partitions SSDs<br/>then ceph_deploy/lvm-setup.yml<br/><i>creates VG + db-slot LVs on each SSD's partition 5</i>"]
PAINBOX["Painbox prep:<br/>installimage post-install.sh<br/><i>NVMe RAID-1 → vg0 → db-slot0..13 + ssd-osd LVs</i>"]
SPEC["ceph_deploy/osds.yml renders<br/>templates/osd-spec.yml.j2 → /etc/ceph/osd-spec.yml<br/><i>one document per host; paths from host_vars</i>"]
APPLY["ceph orch apply osd -i /etc/ceph/osd-spec.yml<br/><i>cephadm: discover disks, LUKS-format, LVM, deploy daemons</i>"]
POLL["Wait for cephadm to provision<br/><i>poll num_osds until expected count reached</i>"]
UP["Wait for OSDs up<br/><i>poll num_up_osds == num_osds</i>"]
UNSET["Defensive: ceph osd unset noin<br/><i>idempotent — clears stale flag from prior runs</i>"]
REWEIGHT["Safety net: fix any reweight=0 OSDs"]
SIETCH --> SPEC
PAINBOX --> SPEC
SPEC --> APPLY --> POLL --> UP --> UNSET --> REWEIGHT
```
### Service-spec model, not per-disk loops
Earlier versions of this role iterated `cephadm ceph-volume lvm create`
per disk and composed `/dev/disk/by-path/...` paths from host_vars
(`sas_path_prefix` + `path_phy`). That assumed sietch's SAS expander
topology and broke on painbox's PCI-ATA disks plus LV-backed SSD OSD.
The current flow renders a cephadm OSD service spec from per-host data
and applies it via `ceph orch apply osd -i`. Cephadm handles device
path resolution, LUKS encryption (`encrypted: true`), LVM provisioning,
and daemon deployment. The role is hardware-shape-agnostic — the only
shape-aware logic is the template's Jinja conditional. See
[ADR-011](adr/011-cephadm-osd-service-specs.md) for the decision record.
### Hardware-shape independence in the template
`templates/osd-spec.yml.j2` renders one document per host (sietch nodes
have unique SAS prefixes per chassis, so a shared spec doesn't work)
with two shape branches:
- **Sietch** (`sas_path_prefix` defined): data path =
`/dev/disk/by-path/{{ sas_path_prefix }}-{{ path_phy }}-lun-0`; SSD
OSD = partition on the SAS-attached SSD via `path_phy + partition`.
- **Painbox** (`sas_path_prefix` undefined): data path =
`/dev/disk/by-path/{{ path_phy }}` (operator authors the full PCI-ATA
identifier in host_vars); SSD OSD = LV via the `lv` field
(`/dev/{{ lv }}`).
`db_devices.paths` is always `/dev/{{ db }}` — both shapes use LVs for
block.db, no composition needed.
### Idempotency
`ceph orch apply osd` is idempotent — re-applying the same spec is a
no-op when deployed OSDs match. New disks (populating an empty bay
later, future expansion) are picked up automatically on the next apply.
Existing OSDs are not destroyed by a spec apply — removal requires
explicit `ceph orch osd rm`.
### Defensive noin handling
The spec-based flow doesn't need the `noin` flag (cephadm rolls out
OSDs gracefully one at a time). The role's tail still includes a
`ceph osd unset noin` task as a defensive cleanup — stale `noin` flags
from a prior failed run of the older imperative flow can leave the
cluster degraded; the unconditional unset clears that safely (no-op
when already unset).
### Reweight-zero safety net
`osds.yml` ends with a task that fixes any OSD stuck at `reweight=0` by
running `ceph osd reweight <id> 1.0`. Rare with the spec-based flow but
kept as a backstop against an OSD coming up while `noin` was set
externally.
---
## 11. Monitoring
### What cephadm auto-deploys
cephadm's bootstrap automatically deploys:
- **node-exporter** on every node
- **ceph-exporter** on every node
- **prometheus** (single instance, cephadm-managed)
- **alertmanager** (single instance)
- **grafana** (single instance, with pre-built Ceph dashboards)
- **89 Prometheus alert rules** across 16 groups
### What we configure
`roles/ceph_deploy/tasks/monitoring.yml` handles only integration:
1. Enable the `prometheus` MGR module (if not already enabled)
2. Wait for all five monitoring service types to report `running > 0`
3. Set dashboard integration URLs for Prometheus, Alertmanager, Grafana
(using the bootstrap node's `bond_ip`)
4. Set Grafana admin credentials from the op-injected
`vault_grafana_admin_password`
5. Disable Grafana SSL cert verification in dashboard (self-signed cert)
`roles/ceph_tuning/tasks/main.yml` verifies the alert rule count and warns
if fewer than 10 rule groups are loaded (expects 16+).
---
## 12. Provision host internals
### Ten-phase flow
```mermaid
flowchart TB
P1["detect.yml<br/><i>live image + UEFI assertions, SSD discovery</i>"]
P2["prerequisites_live.yml<br/><i>apt setup on live image, install debootstrap/mdadm/lvm2</i>"]
P3["disks.yml<br/><i>partition, mdraid, LVM, mount at /mnt</i>"]
P4["install.yml<br/><i>debootstrap Bookworm into /mnt</i>"]
P5["configure.yml<br/><i>hostname, hosts, network, fstab, mdadm templates</i>"]
P6["chroot_packages.yml<br/><i>bind mounts, apt install, machine-id, SSH keys</i>"]
P7["admin_user.yml<br/><i>ansible-iac (key-only) inside chroot;<br/>ops user is created post-boot by baseline (ADR-003)</i>"]
P8["bootloader.yml<br/><i>initramfs, grub-install, efibootmgr</i>"]
P9["finalize.yml<br/><i>marker, ESP mirror, unmount, reboot</i>"]
P10["unmount.yml<br/><i>reverse-order cleanup (shared with rescue)</i>"]
P1 --> P2 --> P3 --> P4 --> P5 --> P6 --> P7 --> P8 --> P9 --> P10
```
### Marker-driven resume gate
After `disks.yml` runs, `main.yml` checks for
`/mnt/etc/ceph-provisioned.json`. If present and the hostname matches, all
chroot phases (4-8) plus the marker/ESP block are skipped. The role goes
straight to unmount + reboot.
This prevents:
- Re-binding bind mounts that are already in place
- Re-rotating SSH host keys (would break known_hosts)
- Re-hashing the ops password with a fresh salt
- Re-running grub-install for no reason
- Overwriting the marker with a stale `provisioned_at` timestamp
The marker filename (`ceph-provisioned.json`) is project-scoped, not
cluster-scoped — every Ceph cluster (sietch, painbox, future) writes the
same filename. The marker's *contents* identify which cluster + host the
machine belongs to.
### Block/rescue cleanup
The entire provisioning sequence (phases 2-9) runs inside a `block/rescue`.
If any phase fails, the rescue block includes `unmount.yml` which tears
down chroot bind mounts and the /mnt hierarchy in reverse order, then
re-raises the failure. This ensures the next run starts from a clean mount
state.
---
## 13. Future
- **CI / GitHub Actions** — the SA split (superuser write vs read-only
consume) already enables it. Read-only SA in CI runs `mise run lint`,
`mise run check`, `mise run test`, `mise run preflight` against every PR.
Superuser SA only runs `mise run tf:plan` (never `apply`) to detect drift.
- **Talos K8s as a sibling stack** — `tf/deployment/<env>/talos/` would
share the same terragrunt root config and S3 backend, with its own state
key (`ceph/<env>/talos/terraform.tfstate`). Deployment plan lives
outside this repo until Phase A begins; a per-stack README lands
alongside the code when it's implemented.
- **TF-managed `onepassword_item` resources** — re-enable the dormant
resources in `secrets.tf.disabled` once the dedicated ceph service
account lands. Tracked in [ADR-009](adr/009-tf-first-op-inject-over-vault-password-sh.md).
- **OSD LUKS keys in 1P** — store dm-crypt keys for DR. Deferred until
the hybrid is stable.
---
## See also
| Topic | Doc |
|------------------------------------|--------------------------------------------------------------------------------|
| TF/Terragrunt detail | [`tf/README.md`](../../../tf/README.md) |
| Wrapper script reference | [docs/scripts.md](scripts.md) |
| Secrets catalog + rotation | [docs/secrets.md](secrets.md) |
| Trust boundaries + encryption | [docs/security-model.md](security-model.md) |
| Hardware specs + network topology | [docs/hardware.md](hardware.md) |
| Coding idioms and anti-patterns | [docs/patterns.md](patterns.md) |
| Adding a new cluster (walkthrough) | [docs/adding-a-cluster.md](adding-a-cluster.md) |
| TF-first + op inject decision | [ADR-009](adr/009-tf-first-op-inject-over-vault-password-sh.md) |
| SSH keys in 1P decision | [ADR-010](adr/010-ssh-keys-in-1password.md) |
| Cephadm OSD service spec decision | [ADR-011](adr/011-cephadm-osd-service-specs.md) |
| Why explicit OSD-to-disk mapping | [ADR-002](adr/002-explicit-osd-mapping.md) |
| Why baseline is split from provision | [ADR-003](adr/003-baseline-split-from-provision.md) |
| Why debootstrap (not preseed) for sietch | [ADR-008](adr/008-debootstrap-over-preseed.md) |
+79
View File
@@ -0,0 +1,79 @@
# Capacity Planning
Audience: Managers, procurement, budget planning. For hardware specs and
disk layouts see [hardware.md](hardware.md); this doc is the sizing math.
## Current deployments
| Cluster | Nodes | HDD × size | Raw HDD | EC-usable (8+3) | 70%-full target |
|---|---|---|---|---|---|
| sietch (Austin) | 3 × Dell R730xd | 30 × 6 TB | ~164 TiB | ~119 TiB | ~83 TiB |
| painbox (Hetzner Helsinki) | 1 × SX295 | 14 × 22 TB | ~280 TiB | ~204 TiB | ~143 TiB |
| **Combined** | | | **~444 TiB** | **~323 TiB** | **~226 TiB** |
Each cluster also contributes ~1 TiB of SSD OSD on the boot SSDs (minor;
ignored in the math above).
Sietch has ~6 empty bays across its three nodes (block.db LVs pre-created),
worth +36 TB raw (~33 TiB) by populating them. No LVM or network changes
needed — cheapest expansion path.
## Sizing formulas
```
EC-usable = raw × (k / (k + m)) = raw × 8/11 = raw × 0.727
Operational target = EC-usable × 0.70 (keep cluster below 70% full)
Raw needed = target_data / 0.727 / 0.70
```
`backfillfull` triggers at 85% full (stops recovery); `full` triggers at
95% (stops writes). 70% is the conservative operational ceiling —
substantial headroom for failures, rebalancing, and growth.
Ceph also consumes small amounts for index pools, RGW metadata, and the
non-EC multipart-upload pool. All negligible relative to data (< 1% each
at steady state).
## Worked example
Target: 50 TiB of application data.
```
raw_needed = 50 / 0.727 / 0.70 = 98 TiB raw HDD
HDDs = 98 TiB / 5.45 TiB = 18 drives (at 6 TB each)
Nodes = 18 / 12 bays = 2 nodes minimum (populated)
```
For 22 TB Hetzner drives the drive count is much lower (~5 drives) but
you still need enough failure domains for EC — see below.
## When to add drives vs. nodes
| | Add drives | Add a node |
|---|---|---|
| When | Empty bays exist with pre-created block.db LVs | All bays populated; need more IOPS, network, or failure domains |
| Cost | ~$30–50 per 6 TB HDD (used) | ~$1,000–1,500 per fully-populated node (used) |
| Adds | +3.96 TiB EC-usable per drive | +47 TiB EC-usable per fully-populated R730xd |
| Impact | Backfill only | Backfill + CRUSH reshape + monitoring/SSH/cephadm host onboarding |
## Failure domain ceiling
EC 8+3 needs 11 failure domains. Austin currently uses
`failure_domain=osd` (spreads across 30 OSDs across 3 nodes) — works today
but a full-node loss degrades a large share of PGs. For host-level
failure domain you need **11+ nodes minimum**. Production (Yucca) will
want this; dev can tolerate the weaker guarantee.
## block.db sizing
Rule of thumb: block.db ≈ 4% of OSD data size.
- **Austin (6 TB HDDs):** 240 GiB LV per HDD matches the rule; 6 LVs
consume 1,440 GiB of each SSD's partition 5 (see hardware.md).
- **Hetzner (22 TB HDDs):** 128 GiB LV is undersized against the 4% rule
(would want 256–512 GiB). Acceptable for dev/benchmark use; production
deployments with 22 TB HDDs should target larger block.db.
If block.db fills up, BlueStore spills metadata to the HDD data partition
— OSD keeps working but small-object operations slow down. Fix: grow the
block.db LV or reduce metadata density.
+162
View File
@@ -0,0 +1,162 @@
# Hardware Reference
Audience: Ops, procurement, capacity planning. Per-node hardware facts that
differ across clusters live in `ansible/ceph/inventories/<cluster>/host_vars/`
(bond_ip, SAS path prefix, SSD PHY positions, HDD-to-block.db mappings) —
those files are committed and are authoritative for physical topology.
For where this fits in the tool mesh, see [architecture.md](architecture.md).
## Network topology
Both clusters use a **single-network** design today — public and cluster
traffic share one subnet. Production will separate them; see
[security-model.md](security-model.md) for the threat-model implications.
| Cluster | Subnet | Connection |
|----------|-------------------|------------------------------------------------------------------------|
| sietch | `10.10.10.0/24` | 2× 10GbE bonded active-backup (eno1 + eno2) per node; private switch |
| painbox | public /32 | Single 1GbE, direct SSH (no bond, no ProxyJump) |
Per-node connection IPs (`bond_ip`) are declared in
`tf/deployment/dev/ceph/clusters.auto.tfvars` and rendered by TF into the
cluster's `inventory.ini`. They're also mirrored into `host_vars/` for use
by roles that need the IP as a variable (e.g., cephadm public-network
resolution, dashboard URL construction).
## sietch -- Dell R730xd (Austin)
| Component | Spec |
|-----------|------|
| Chassis | Dell R730xd 12-bay LFF + 2 rear 2.5" bays |
| CPU | 2x Intel Xeon E5-2697A v4 (64 vCPUs) |
| RAM | 128 GB DDR4 |
| Boot SSDs | 2x Micron 5100 3.8TB (rear bays 12/13, mdraid-1) |
| HDD OSDs | 8-12x SAS 6TB (HGST HUS726060AL4210 / Seagate ST6000NKCLAR6000) |
| SSD OSDs | Partition 6 on each boot SSD (colocated, no separate block.db) |
| Block.db | 6x 240GB LVs per SSD (partition 5, LVM VG) |
| HBA | Broadcom/LSI SAS3008 IT mode (mpt3sas, no RAID) |
| Network | 2x 10GbE bonded active-backup (eno1 + eno2) |
| Boot | UEFI, dual ESP (one per SSD, rsync-mirrored) |
| OS | Debian 12 Bookworm (debootstrap provisioned) |
| Provisioning | Live image → `provision.yml` → debootstrap |
### SSD partition layout (per SSD)
| Partition | Size | Use |
|-----------|------|-----|
| 1 | 512 MB | ESP (FAT32, UEFI boot) |
| 2 | 1 GB | /boot (mdraid-1, ext4, metadata 1.0) |
| 3 | 80 GB | / (mdraid-1, ext4) |
| 4 | 8 GB | swap (mdraid-1) |
| 5 | ~1.4 TB | Ceph block.db LVs (LVM VG) |
| 6 | ~2 TB | SSD OSD data (ceph-volume) |
## painbox -- Hetzner SX295 (Helsinki)
| Component | Spec |
|-----------|------|
| Chassis | Hetzner SX295 storage server |
| CPU | AMD EPYC 7502P (32C/64T) |
| RAM | 128 GB DDR4 ECC |
| Boot NVMe | 2x Samsung 7.68TB (installimage RAID-1, vg0) |
| HDD OSDs | 14x Seagate Exos X22 22TB SATA |
| SSD OSD | 1x ~4.4TB LV on vg0 (NVMe remainder) |
| Block.db | 14x 128GB LVs on vg0 |
| SATA | 3 onboard controllers (8+2+4 ports = 14 total) |
| Network | Single 1GbE, direct SSH (no bond, no ProxyJump) |
| Boot | BIOS (Hetzner standard) |
| OS | Debian 12 Bookworm (Hetzner installimage) |
| Provisioning | Rescue mode → `installimage/autosetup` + `post-install.sh` |
### NVMe layout (vg0 on md1)
| LV | Size | Use |
|----|------|-----|
| swap | 32 GB | Swap |
| root | 100 GB | / |
| var | 200 GB | /var |
| varlog | 50 GB | /var/log |
| db-slot0..13 | 14x 128 GB | Block.db per HDD OSD |
| ssd-osd | ~4.4 TB | SSD OSD data |
| (reserve) | ~512 GB | Future expansion |
> **Why Bookworm and not Trixie:** upstream Ceph Tentacle's Debian
> repository at `download.ceph.com/debian-tentacle/dists/` publishes only
> for `bookworm`, `jammy`, and `noble`. Trixie is not yet supported. The
> autosetup `IMAGE` line MUST select a Bookworm tarball until upstream
> ships Trixie packages.
## Comparison
| | sietch (per node) | painbox (single node) |
|---|---|---|
| HDD OSD count | 8-12 | 14 |
| SSD OSD count | 2 | 1 |
| block.db per HDD | 240 GB | 128 GB |
| Total raw HDD | 48-72 TB | 308 TB |
| Device path format | `/dev/disk/by-path/sas-exp*-phy*-lun-0` | `/dev/disk/by-path/pci-*-ata-*` |
| EC profile | 8+3 (failure domain: OSD) | 8+3 (failure domain: OSD) |
| Replicated pool size | 2 (min_size 1) | 2 (min_size 1) |
## host_vars schema by hardware shape
The two clusters have fundamentally different storage topologies, which
shows up in their `host_vars/<host>.yml` schemas. When adding a new cluster,
operators must pick the schema matching the hardware — not just copy from
either existing cluster blindly.
### sietch-shape (SAS expander + dual-SSD-VG)
```yaml
hostname_short: <cluster>-ceph-<name>
bond_ip: 10.10.X.Y
sas_path_prefix: "pci-XXXX:XX:XX.X-sas-exp0xXXXX..." # REQUIRED — disambiguates SAS topology
ssd1_phy: 12 # PHY slot of boot SSD #1
ssd2_phy: 13 # PHY slot of boot SSD #2
ceph_db_vg1: ceph-db-rear12 # VG on SSD1 partition 5
ceph_db_vg2: ceph-db-rear13 # VG on SSD2 partition 5
ceph_db_lvs_per_ssd: 6 # 6 db-slot LVs per VG → 12 total
ceph_hdd_osds:
- { path_phy: phy0, db: ceph-db-rear12/db-slot0 }
- ...
ceph_ssd_osds: # SSD OSD = partition 6 of each boot SSD
- { path_phy: phy12, partition: 6 }
- { path_phy: phy13, partition: 6 }
```
The role composes full disk paths as
`/dev/disk/by-path/<sas_path_prefix>-<path_phy>-lun-0` (HDDs) or with
`-part<N>` suffix (SSD partitions). `roles/ceph_deploy/tasks/lvm-setup.yml`
runs this shape's VG/LV recovery path.
### painbox-shape (NVMe-RAID + single VG + LV-backed SSD OSD)
```yaml
hostname_short: <cluster>-ceph-<name>
bond_ip: <public IP>
ceph_db_vg: vg0 # single VG on NVMe RAID-1 (no sas_path_prefix)
ceph_db_lvs_per_node: 14 # all db-slot LVs on the one VG
ceph_hdd_osds:
- { path_phy: pci-XXXX:XX:XX.X-ata-1, db: vg0/db-slot0 } # full PCI-ATA path
- ...
ceph_ssd_osds:
- { lv: vg0/ssd-osd } # SSD OSD = LV on same VG
```
The role uses `path_phy` directly as the by-path identifier (no composition
needed — operator authors the full string). For SSD OSDs, `lv` field is
used (`/dev/<lv>`) instead of `path_phy + partition`. `lvm-setup.yml` is
**skipped** on this shape (gated `when: sas_path_prefix is defined`) —
LVM is owned by the Hetzner installimage post-install script.
### Decision rule for new clusters
- **Has a SAS expander** (PERC HBA, mpt3sas, etc.) and **dedicated boot SSDs
partitioned for both block.db and OS** → sietch-shape.
- **Single VG covering boot + block.db + SSD OSD** (typical for
hosting-provider servers with NVMe RAID-1) → painbox-shape.
- **Other shapes** (e.g., dedicated NVMe block.db drives) require either
a new shape branch in `roles/ceph_deploy/tasks/osds.yml`'s template or
a fresh decision — see [ADR-011](adr/011-cephadm-osd-service-specs.md)
for how cephadm OSD service specs handle hardware-shape independence.
+243
View File
@@ -0,0 +1,243 @@
# Naming
Every name derivable in this project — hostname, inventory directory, 1P
item title, SSH key filename — traces back to one entry per cluster in
`tf/deployment/<env>/ceph/clusters.auto.tfvars`. TF's `ceph-cluster` module
assembles the rest.
For how naming fits the broader system see
[architecture.md §4 (Terraform)](architecture.md); for the per-item 1P
catalog see [secrets.md](secrets.md); for the SSH-key lifecycle see
[ADR-010](adr/010-ssh-keys-in-1password.md).
## The three name layers
Three parallel naming surfaces derive from the same tfvars entry, but
**only the hostname carries the `role_in_hostname` segment**. Inventory
directory and 1P item prefix are both hardcoded to `ceph` / `CEPH` in the
ceph-cluster module — they're project-scoped, not role-scoped.
| Surface | Pattern | Role source |
|------------------------|---------------------------------------------------|---------------------------------|
| Hostname (short + FQDN)| `<cluster>-<role>-<name>[.<env>.<dc>.<provider>.futo.cloud]` | `role_in_hostname` tfvar |
| Inventory directory | `inventories/<cluster>-ceph.<env>.<dc>.<provider>/` | always `ceph` (module-hardcoded) |
| 1P item prefix | `<CLUSTER>_CEPH_*` | always `CEPH` (module-hardcoded) |
This split is deliberate. A future cluster where every node is a dedicated
OSD might set `role_in_hostname = "osd"` (yielding hostnames like
`mesa-osd-willow`) but its inventory dir and 1P items would still grep-match
`*-ceph.*` and `*_CEPH_*` alongside every other Ceph-project cluster.
### Hostname segments
| Component | Example | Source (per-cluster tfvars field) |
|-----------|---------------------------------|--------------------------------------|
| cluster | `sietch`, `painbox` | top-level map key |
| role | `ceph` (small clusters), `osd` | `role_in_hostname` (default `ceph`) |
| name | `laurel`, `evelyn` | `hosts[].name`, or TF-picked |
| env | `dev`, `staging`, `prod` | `environment` |
| dc | `austin`, `hel`, `fsn` | `datacenter` |
| provider | `int`, `htz` | `provider_code` |
Current `role_in_hostname` values: both sietch and painbox use `ceph`
(mixed-role, all-nodes-are-everything). Dedicated-role hostnames (`osd`,
`mon`) are supported but not used today.
## Cluster naming
**Who chooses:** the engineer adding the cluster picks the name at the
moment they add the entry to `clusters.auto.tfvars`. No automation — it's
a deliberate, one-time decision.
**When:** before any `mise run tf:apply`, any 1P item creation, any cluster
bootstrap. Renaming after deployment is expensive (see "Cost of renaming"
below).
**Convention (unenforced):** Dune-themed. Existing clusters are `sietch`
(an underground Fremen community) and `painbox` (the Bene Gesserit
gom-jabbar test apparatus). Dune candidates not yet used, and that fit
the hard constraints below: `arrakis`, `caladan`, `giedi`, `ixian`,
`kwisatz`, `muaddib`, `fremen`, `chani`, `leto`, `jessica`. Nothing in
the code enforces Dune specifically — mixed themes or theme breaks are
acceptable when they communicate intent better (e.g., a cluster named
after its datacenter for a production tier).
**Hard constraints:**
- **lowercase alphanumeric** — becomes the HCL map key, the inventory
directory segment (`<name>-ceph.<env>...`), the hostname prefix, and
(uppercased) the 1P item prefix (`<NAME>_CEPH_*`).
- **Short** — appears in every hostname and every 1P item title. Aim for
6–10 characters; 15 is the realistic ceiling.
- **Unique within the `yucca_tf_*` 1P item namespace** — other Futo
consumers (o11y, future stacks) write items to the same vault set.
Before committing, check:
```bash
op item list --vault yucca_tf_dev --format=json \
| jq -r '.[] | .title' | grep -i "^<PROPOSED_NAME>_"
```
Must return empty.
- **Not already a `clusters.auto.tfvars` key** — TF enforces this with
a plan-time error.
- **No dashes, dots, or underscores in the cluster name itself** — those
are segment separators in inventory directories and hostnames. A cluster
name `my-cluster` would produce `my-cluster-ceph-laurel` which parses
ambiguously. Use `mycluster` instead.
**Soft guidance:**
- Memorable — operators will say it out loud in incidents.
- Distinct from existing clusters' first 3 letters (grep-friendly in logs).
- Doesn't encode environment or datacenter — those live in separate
segments. The cluster name is project identity, not location.
### Cost of renaming after deployment
A rename touches all of:
1. `clusters.auto.tfvars` map key
2. Inventory directory name
3. Every hostname (short + FQDN) and every SSH `known_hosts` entry for every operator
4. 1P item titles (`<OLD>_CEPH_*` → `<NEW>_CEPH_*`) including the SSH Key item
5. cephadm cluster identity (requires cluster rebuild in the common case)
6. `ansible_ssh_key` path in the tfvars (`~/.ssh/id_ed25519_<cluster>`) and
the mapping in `scripts/install-ssh-keys.sh`
7. Any DNS records and external systems that reference the hostnames
Expect hours-to-days of work, cluster downtime, and coordination with
every consumer of the cluster's S3/dashboards/etc. Painbox's rename
(from `painbox-osd-5c3cac.lab.*` to `painbox-ceph-evelyn.dev.*`) was
cheap only because painbox is idle — a running production Ceph cluster
makes this a multi-week project.
**Pick once. Pick deliberately.**
## Host naming
Each host entry in `hosts = [...]` can either declare a name or omit it
to let TF pick one from the wordlist. Both paths are first-class;
different clusters use different paths based on operator preference.
### Operator-declared
Declare the name explicitly in the tfvars. Used when the operator has a
specific name in mind — typically because it's been spoken during
planning and the team already uses it.
```hcl
hosts = [
{ name = "laurel", bond_ip = "10.10.10.90", bootstrap = true },
{ name = "lawson", bond_ip = "10.10.10.91" },
{ name = "samara", bond_ip = "10.10.10.92" },
]
```
Sietch uses this path.
Names must be unique **within a cluster**, not globally. A future
`mesa-ceph-laurel` can coexist with `sietch-ceph-laurel` — the FQDN
disambiguates.
### Auto-picked from wordlist
Omit `name` (leave the field absent) and the module picks from
`tf/shared/modules/ceph-cluster/wordlist.txt` (923 words) via
`random_shuffle`, seeded by `cluster_name` + `name_seed`. Picks are
stable across subsequent applies.
```hcl
hosts = [
{ bond_ip = "157.180.105.198", bootstrap = true }, # TF picks
]
```
Painbox uses this path — `hosts[0]` has no name, and TF picked `evelyn`
on first apply, yielding `painbox-ceph-evelyn`.
Operator-declared names are excluded from the available pool to prevent
collisions within the cluster.
### Stability rules (either path, or mixed)
- **Add new hosts at the tail of the list.** Auto-picked names are
positional — `hosts[0]` gets `random_shuffle.result[0]`, `hosts[1]`
gets `result[1]`, and so on. Inserting a new entry at position 0 would
shift every subsequent host's result-index. Always append.
- **Never bump `name_seed` once a cluster has deployed hosts.** Bumping
re-rolls every auto-picked name in the cluster, which cascades into
certs, SSH `known_hosts`, 1P items, DNS, cephadm identity.
- **Mixing paths has a non-obvious side effect.** Auto-picked names are
drawn from `available_words = wordlist − explicit_names`. Adding or
removing an *operator-declared* host changes `explicit_names`, which
changes the shuffle input length, which re-permutes the result. An
auto-picked host at `hosts[2]` could get a different name even if
nothing about its own entry changed. The safe patterns:
1. All-operator-declared within a cluster (sietch's model), or
2. All-auto-picked within a cluster (painbox's model).
Mixed works for initial setup but complicates later add/remove.
- **Converting between paths after deploy** (e.g., adding `name = "evelyn"`
to a host that previously auto-picked `evelyn`) **does not preserve the
name** despite appearing to match — it rewrites `available_words` and
re-rolls every other auto-pick. Only do this if you're prepared to pin
every auto-named host in the same apply, or accept the cluster-wide
rename.
## Inventory directory naming
TF renders inventory directories as:
```
inventories/<cluster>-ceph.<env>.<dc>.<provider>/
```
The `-ceph` segment is hardcoded in the ceph-cluster module regardless of
`role_in_hostname`. Keeps all Ceph-project inventory paths grep-matchable
as `*-ceph.*` — even a hypothetical cluster with `role_in_hostname = "osd"`
(hostnames `mesa-osd-*`) still renders `mesa-ceph.prod.fsn.htz/`.
Defined in `tf/deployment/<env>/ceph/main.tf` (`local.inventory_dirs`).
## 1Password item naming
Items are titled `<CLUSTER>_CEPH_<specifier>` (SHOUTY_SNAKE_CASE). The
`<CLUSTER>_CEPH_*` prefix is hardcoded in
`tf/shared/modules/ceph-cluster/main.tf` (`local.secret_prefix`) — same
project-scoping rationale as the inventory directory.
Per cluster, the expected item set:
| Category | Title | Source of values |
|----------|--------------------------------------------------------|-----------------------------------------------------|
| Password | `<CLUSTER>_CEPH_OPS_PASSWORD` | `op item create --generate-password` at setup |
| Password | `<CLUSTER>_CEPH_DASHBOARD_PASSWORD` | same |
| Password | `<CLUSTER>_CEPH_GRAFANA_PASSWORD` | same |
| Password | `<CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY` | same |
| Password | `<CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY` | same |
| SSH Key | `<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY` | `op item create --category "SSH Key" --ssh-generate-key=ed25519` |
| Document | `<CLUSTER>_CEPH_RGW_TLS_CERT` | populated by `mise run capture` post-deploy |
| Document | `<CLUSTER>_CEPH_RGW_TLS_KEY` | same |
| Document | `<CLUSTER>_CEPH_CLIENT_ADMIN_KEYRING` | same |
Passwords + SSH Key are created at cluster-add time (step 6 of
[adding-a-cluster.md](adding-a-cluster.md)). DR Documents are upserted
automatically on the first `mise run capture` after deploy (step 10).
For the full consumption flow (which Ansible variable each item maps to,
which role reads it) see [secrets.md](secrets.md).
## Workstation SSH key filenames
The operator-side private key path is derived from the cluster name:
```
~/.ssh/id_ed25519_<cluster>
```
Examples: `~/.ssh/id_ed25519_sietch`, `~/.ssh/id_ed25519_painbox`. Per
[ADR-010](adr/010-ssh-keys-in-1password.md), the keypair lives in
`<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY` in 1P and is installed via
`scripts/install-ssh-keys.sh <cluster>`.
The `ansible_ssh_key` field in the cluster's tfvars entry must match this
path. If you choose a non-default filename (unusual), update both together
and also adjust the mapping in `scripts/install-ssh-keys.sh`.
+441
View File
@@ -0,0 +1,441 @@
# Patterns
Project-specific Ansible idioms. This doc skips generic Ansible hygiene
(FQCN, `set -o pipefail`, etc. — those are table stakes) and focuses on
patterns that are non-obvious or specific to how this Ceph automation is
built.
For how patterns wire into the wrapper + secrets flow, see
[scripts.md](scripts.md). For role structure and the pre-submit checklist,
see [adding-a-role.md](adding-a-role.md).
---
## Check-then-set against the cluster
**Problem:** `ceph config set` always reports `changed`. Running it
unconditionally makes every play dirty and obscures real drift.
**Pattern:** read the current value, compare to the expected value, apply
only if different. The comparison is the key — raw `ceph config set` with
`changed_when: true` is a lie.
Canonical example — `roles/ceph_tuning/tasks/main.yml`:
```yaml
- name: Read current OSD config values
ansible.builtin.shell: |
set -o pipefail
echo "recovery_max_active=$(ceph config get osd osd_recovery_max_active)"
# ...
register: current_osd_config
changed_when: false
- name: Apply Ceph config values
ansible.builtin.command: >
ceph config set {{ item.section }} {{ item.key }} {{ item.value }}
loop:
- { section: osd, key: osd_recovery_max_active, value: "{{ ... }}",
current: "{{ osd_cfg.recovery_max_active | default('') | float }}" }
when: item.current | float != item.value | float
changed_when: true
```
Other instances: `rgw.yml` (zonegroup hostnames), `crush-rules.yml`
(rule existence), `monitoring.yml` (module enable check).
### Float comparison gotcha
`ceph config get osd osd_recovery_sleep_hdd` returns `0.100000`, but the
Ansible variable is `0.1`. String comparison fails; integer comparison
truncates. Always cast both sides to `| float` before comparing. Affects
any Ceph config value returned with trailing zeros
(`osd_recovery_sleep_hdd`, `osd_deep_scrub_interval`, etc.).
---
## Marker-driven idempotency
**Problem:** provisioning is destructive and multi-phase. A crash during
phase 6 must not re-wipe disks on the next run. But the role still needs
to handle a completely fresh node.
**Pattern:** write a JSON marker at the end of provisioning. On
subsequent runs, check for the marker and skip completed phases.
The marker filename (`/etc/ceph-provisioned.json`) is project-scoped —
every Ceph cluster writes the same filename. The marker's *contents*
identify which cluster + host the machine belongs to (hostname, fqdn,
bond_ip, cluster_name, SSD serials, provisioned_at timestamp).
**Template:** `roles/provision_host/templates/ceph-provisioned.json.j2`.
**Resume gate** — `roles/provision_host/tasks/main.yml`:
```yaml
- name: Check if provisioning marker is present in chroot
ansible.builtin.stat:
path: "{{ provision_mnt }}{{ provision_marker_path }}"
register: marker_stat
- name: Set provisioning_done fact
ansible.builtin.set_fact:
provisioning_done: "{{ marker_stat.stat.exists | bool }}"
```
Then every chroot phase is gated on `when: not provisioning_done`.
`disks.yml` additionally handles the "md array already assembled but
nothing mounted" case — mounts, reads the marker, validates the hostname
matches `inventory_hostname`, and either resumes (marker matches) or
unmounts and re-wipes (marker missing or wrong host).
**Use when:** any multi-step destructive workflow where partial
completion must be resumable. The marker must contain enough identity
information to distinguish "this node's previous run" from "a different
node's leftover state."
---
## Block/rescue cleanup
**Problem:** if provisioning fails mid-chroot (e.g., debootstrap network
error), bind mounts at `/mnt/dev`, `/mnt/proc`, `/mnt/sys` remain active.
The next run fails because it can't cleanly remount.
**Pattern:** wrap the phase sequence in `block/rescue`. The rescue
includes a shared `unmount.yml` that tears down mounts in reverse order,
then re-raises the failure.
```yaml
- name: Provisioning phases
block:
- ansible.builtin.import_tasks: prerequisites_live.yml
- ansible.builtin.import_tasks: disks.yml
# ... phases 4-8 ...
- ansible.builtin.import_tasks: finalize.yml
rescue:
- name: Unmount /mnt hierarchy on failure (best-effort cleanup)
ansible.builtin.include_tasks: unmount.yml
- name: Re-raise failure
ansible.builtin.fail:
msg: "Provisioning phase failed for {{ inventory_hostname }}."
```
The unmount tasks use `failed_when: false` — if a path isn't mounted, we
just want to keep going.
---
## Conditional features
**Problem:** not every cluster needs iSCSI, NFS, CPU governor tuning, or
centralized logging. These features should be zero-overhead when
disabled.
**Pattern:** gate on `<feature>_enabled | bool` with defaults of `false`.
Current feature flags:
| Feature | Toggle | Default | Consumer |
|---|---|---|---|
| CPU governor | `ceph_cpu_governor_enabled` | false | `hardware_tuning` |
| Centralized logging | `ceph_logging_enabled` | false | `os_tuning` |
| iSCSI firewall | `ceph_firewall_iscsi_enabled` | false | `security` |
| NFS firewall | `ceph_firewall_nfs_enabled` | false | `security` |
| Firewall overall | `ceph_firewall_enabled` | true | `security` |
| RGW TLS | `ceph_rgw_ssl` | false | `ceph_deploy/rgw` |
| Audit logging | `ceph_audit_enabled` | true | `ceph_tuning` |
| SSH open to all sources (dev) | `ceph_firewall_ssh_any_source` | true (dev) | `security` |
| Weekly fstrim timer (SSDs) | `ceph_enable_fstrim_timer` | true | `hardware_tuning` |
| LVM device filter | `ceph_lvm_filter_enabled` | false | `hardware_tuning` |
The `| bool` filter is mandatory. Ansible may pass booleans as strings
from inventory or extra-vars; without `| bool`, the string `"false"` is
truthy.
**In Jinja templates** (e.g. `nftables.conf.j2`):
```jinja2
{% if ceph_firewall_iscsi_enabled | bool %}
ip saddr {{ net }} tcp dport {{ ceph_firewall_iscsi_port }} accept
{% endif %}
```
---
## Drift detection pattern
`drift.yml` is a read-only play that compares expected state against
live cluster. Three steps:
1. **Load all role defaults** via `vars_files` — gives drift detection
access to expected values without depending on any role's execution:
```yaml
vars_files:
- roles/baseline/defaults/main.yml
- roles/os_tuning/defaults/main.yml
# ...
```
2. **Accumulate results** into a list via `set_fact`:
```yaml
drift_results: >-
{{ drift_results + [{
'category': 'sysctl',
'item': item.item.key,
'expected': item.item.expected | string,
'actual': item.stdout | trim,
'match': (item.stdout | trim) == (item.item.expected | string)
}] }}
```
3. **Generate a formatted report** using Jinja in `set_fact`.
Categories checked: sysctl values, HDD/SSD I/O schedulers, nftables
policy, SSH `PasswordAuthentication`, ops sudo config, OSD status, MON
quorum, RGW daemon count, cluster health, Ceph config values.
**Use when:** building read-only comparison plays. The pattern generalizes
to any "expected vs actual" audit.
---
## CEPH_ENV as inventory selector
Wrappers and downstream scripts derive the cluster's paths from the
`CEPH_ENV` environment variable, which points at the TF-rendered
inventory file. Cluster identity is authoritative in
`clusters.auto.tfvars`; `CEPH_ENV` is the runtime pointer.
```
CEPH_ENV = inventories/sietch-ceph.dev.austin.int/inventory.ini
|
dirname -> inventories/sietch-ceph.dev.austin.int
|
+ "/secrets.yml.tpl" -> op inject input
```
**`scripts/ansible-play.sh`** derives the secrets template path as
`$(dirname $CEPH_ENV)/secrets.yml.tpl` and fails closed if either file
is missing. See [scripts.md](scripts.md) for the full contract.
**Destroy task** in `.mise.toml` extracts the domain for the safety gate:
```bash
CLUSTER_ID=$(basename "$CEPH_ENV_DIR") # sietch-ceph.dev.austin.int
DOMAIN=${CLUSTER_ID#*-ceph.}.futo.cloud # dev.austin.int.futo.cloud
```
---
## Placement group logic
MON placement strategy varies by cluster size. A 2-node cluster should
NOT run 2 MONs (no quorum majority possible). A 3+ node cluster should
run MON on all nodes.
`roles/ceph_deploy/tasks/placement.yml`:
```yaml
- name: Calculate MON placement
ansible.builtin.set_fact:
mon_hosts: >-
{%- if groups['ceph_mon'] | default([]) | length > 0 -%}
{{ groups['ceph_mon'] | map('extract', hostvars, 'hostname_short') | join(',') -}}
{%- elif active_hosts.stdout.split(',') | length <= 2 -%}
{{ hostvars[groups['ceph_bootstrap'][0]]['hostname_short'] -}}
{%- else -%}
{{ active_hosts.stdout -}}
{%- endif -%}
```
Three branches: (1) explicit `ceph_mon` inventory group → use those;
(2) ≤ 2 active hosts → single MON (bootstrap only, avoids 2-MON
quorum fragility); (3) 3+ active hosts → MON on all active hosts.
MGR always deploys on all active hosts — standby MGRs are harmless and
provide fast failover.
---
## Shell + `changed_when` discipline
Two rules worth stating explicitly because they're the most common
lint-clean failures:
1. **`changed_when: false`** for read-only commands (checks, queries,
status).
2. **`changed_when: true`** only when guarded by `when:` — the task only
runs when something needs to change. Unguarded `changed_when: true`
reports "changed" on every run; ansible-lint catches this.
3. **Output-based `changed_when`** for shell tasks that may or may not
change state:
```yaml
changed_when: "'CHANGED' in zonegroup_hostnames.stdout"
```
4. **`failed_when` with fallthrough** for commands where a specific
error is expected and acceptable:
```yaml
failed_when:
- pg_data_result.rc != 0
- "'is not >= current' not in pg_data_result.stderr | default('')"
```
Every `shell`/`command` task in this codebase sets one of these — no bare
`command:` without a `changed_when`. Lint enforces it.
`no_log: true` on every task handling passwords, keys, or credentials.
Ansible output is committed to `ansible.log` and displayed to operators
— secrets must never land there.
---
## Cephadm service specs over imperative loops
**Problem:** the role's first instinct is "iterate every disk / daemon /
service in Ansible and run `cephadm` per item." This couples the role
tightly to per-host hardware shape (path composition, partition layout,
LV vs disk topology) and breaks on any new cluster shape — the original
sietch-shape `osds.yml` failed on painbox's PCI-ATA + LV-backed SSD OSD
topology because `sas_path_prefix` was hardcoded into the path
composition.
**Pattern:** for any cephadm-managed surface (OSDs, RGW, MON/MGR
placement, monitoring), render a **declarative service spec** and apply
it via `ceph orch apply -i <spec>.yaml`. Cephadm handles per-disk
discovery, daemon lifecycle, encryption, LVM, etc. internally. The role
becomes a thin renderer + applier; hardware shape moves into the
template's Jinja conditional, not the role logic.
**Examples in this codebase:**
- `templates/rgw-spec.yaml.j2` + `tasks/rgw.yml`'s `ceph orch apply -i`
task — RGW daemon placement spec.
- `templates/osd-spec.yml.j2` + `tasks/osds.yml`'s
`ceph orch apply osd -i` task — OSD service spec; per-host documents
in a multi-doc YAML; Jinja conditional handles sietch-shape vs
painbox-shape path composition.
**Template skeleton:**
```jinja2
{% for host in groups['ceph_nodes'] %}
{% set h = hostvars[host] %}
---
service_type: <kind>
service_id: {{ h.hostname_short }}-<role>
placement:
hosts: [{{ h.hostname_short }}]
spec:
<kind-specific fields>
{% endfor %}
```
**Apply pattern in tasks/*.yml:**
```yaml
- name: Render cephadm service spec
ansible.builtin.template:
src: <kind>-spec.yml.j2
dest: /etc/ceph/<kind>-spec.yml
mode: '0644'
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
run_once: true
- name: Apply cephadm service spec
ansible.builtin.command: ceph orch apply <kind> -i /etc/ceph/<kind>-spec.yml
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
run_once: true
changed_when: ...
```
**Use when:** the surface you're managing is a cephadm-orch-supported
service type (`host`, `mon`, `mgr`, `osd`, `rgw`, `mds`, `nfs`,
`prometheus`, `grafana`, `alertmanager`, `node-exporter`,
`ceph-exporter`, etc.). Don't use for surfaces cephadm doesn't manage
declaratively (CRUSH rules, pools, ceph config tunables, RGW realm/zone
setup, S3 user creation) — those still need imperative `ceph` /
`radosgw-admin` calls.
**Trade-off vs imperative loops:** debugging "why isn't this disk
becoming an OSD?" is harder — there's no per-disk log line. Check
`ceph orch ls` / `ceph orch ps` / `ceph cephadm osd activate <host>
--dry-run` instead. Worth the trade-off because the role becomes
hardware-shape-agnostic.
See [ADR-011](adr/011-cephadm-osd-service-specs.md) for the full
decision record on the OSD-path migration.
---
## Anti-patterns
### `changed_when: true` without a `when` guard
```yaml
# BAD — reports changed on every run even when idempotent
- ansible.builtin.command: ceph config set osd foo bar
changed_when: true
# GOOD — only runs when needed, so changed_when: true is accurate
- ansible.builtin.command: ceph config set osd foo bar
when: current_foo != 'bar'
changed_when: true
```
### Shell without pipefail
```yaml
# BAD — if `ceph osd dump` fails, grep runs on empty input and task succeeds
- ansible.builtin.shell: ceph osd dump | grep noin
# GOOD — pipefail propagates the ceph failure
- ansible.builtin.shell: |
set -o pipefail
ceph osd dump | grep noin
args:
executable: /bin/bash
```
### Hardcoded site-specific values in roles
```yaml
# BAD — in a role's tasks/main.yml
- ansible.builtin.command: ceph config set osd osd_recovery_max_active 1
# GOOD — value comes from defaults, overridable via group_vars
- ansible.builtin.command: >
ceph config set osd osd_recovery_max_active
{{ ceph_osd_recovery_max_active }}
```
Roles use `defaults/main.yml` for all tunables. Site-specific values live
in `inventories/<cluster>/group_vars/all/vars.yml`.
### Using `ansible_play_batch` / `ansible_play_hosts` for placement specs
```yaml
# BAD — --limit shrinks play_batch, cephadm removes daemons from omitted hosts
- ansible.builtin.command:
ceph orch apply rgw --placement="{{ ansible_play_batch | join(',') }}"
```
cephadm is declarative: applying a smaller placement list REMOVES daemons
from hosts not in the list. Always use `groups['ceph_nodes']` (the full
inventory group) for placement specs, never `ansible_play_batch` or
`ansible_play_hosts`. The comment at
`roles/ceph_deploy/tasks/rgw.yml:466` explains the failure mode in detail.
### Running plays with bare `ansible-playbook`
Every playbook in this project consumes op-injected secrets. Running
`ansible-playbook foo.yml` directly skips the wrapper, `op inject` never
runs, and `vault_*` variables are empty — tasks that need them fail with
confusing errors. Always `scripts/ansible-play.sh <playbook.yml>`. See
[scripts.md](scripts.md) for the full contract.
+208
View File
@@ -0,0 +1,208 @@
# Runbook: Add a Node to an Existing Cluster
**When:** Expanding cluster capacity or replacing a failed chassis.
**Time estimate:** 30-60 minutes (Austin physical), 15-30 minutes (Hetzner).
**Prerequisites:**
- Cluster is healthy (`ceph health` returns `HEALTH_OK` or understood warnings)
- `op` session live (desktop unlocked or `OP_SERVICE_ACCOUNT_TOKEN` set)
- SSH access to existing cluster nodes working
---
## 1. Declare the new host in TF
Pick an unused name, or omit `name` to let TF auto-pick from the
923-word list seeded per-cluster (see [docs/naming.md](../naming.md)).
Edit `tf/deployment/dev/ceph/clusters.auto.tfvars` and append to the
target cluster's `hosts` list:
```hcl
sietch = {
# ...
hosts = [
{ name = "laurel", bond_ip = "10.10.10.90", bootstrap = true },
{ name = "lawson", bond_ip = "10.10.10.91" },
{ name = "samara", bond_ip = "10.10.10.92" },
{ name = "maxton", bond_ip = "10.10.10.93" }, # NEW
]
}
```
Apply:
```bash
mise run tf:apply
```
This re-renders `inventory.ini` with the new host in `[ceph_nodes]`,
`[ceph_mon]`, and `[ceph_join]`. Bootstrap assignment doesn't change —
still pinned to the first/explicitly-declared bootstrap host. Add new
hosts at the **tail** of the list so existing auto-picked names keep
their positions.
## 2. Create host_vars
```bash
cd inventories/sietch-ceph.dev.austin.int
cp host_vars/example.yml host_vars/sietch-ceph-<name>.yml
```
Edit the new file. Every field is node-specific and must match the physical
hardware:
| Field | How to find it |
|---|---|
| `hostname_short` | The full `sietch-ceph-<name>` from step 1 |
| `bond_ip` | Next available IP in 10.10.10.0/24. Austin convention: yucca-N = 10.10.10.9N |
| `sas_path_prefix` | SSH into node, run `ls /dev/disk/by-path/ \| grep sas` |
| `ssd1_phy` / `ssd2_phy` | Identify SSD PHY positions from `lsscsi -t` output |
| `ceph_db_vg1/2` | Name the VGs by slot, e.g., `ceph-db-rear12`, `ceph-db-rear13` |
| `ceph_hdd_osds` | Map each populated HDD bay to its PHY and corresponding db-slot LV |
| `ceph_ssd_osds` | SSD partition 6 on each SSD (no separate block.db) |
## 3. Confirm inventory regenerated correctly
After `tofu apply` in step 1, inspect the rendered inventory:
```bash
cat inventories/sietch-ceph.dev.austin.int/inventory.ini
```
The new host should appear in:
- `[ceph_nodes]` — all cluster members
- `[ceph_mon]` — MON/MGR placement
- `[ceph_join]` — everything except the bootstrap host
Bootstrap host is unchanged. TF never moves an existing bootstrap assignment.
## 4. Provision the OS
### Austin (physical servers)
1. Boot the server to the Debian 12 live image via iDRAC virtual console
2. Verify the live image is reachable on the node's bond IP
3. Run provisioning:
```bash
CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory-provision.ini \
scripts/ansible-play.sh provision.yml \
-e confirm_wipe=true \
--limit sietch-ceph-<name>
```
4. Wait for reboot and verify SSH access as `ansible-iac`:
```bash
ssh -i ~/.ssh/id_ed25519_sietch ansible-iac@sietch-ceph-<name> hostname -f
```
Expected output: `sietch-ceph-<name>.dev.austin.int.futo.cloud`
### Hetzner (remote servers)
1. Boot into rescue mode via Hetzner Robot panel
2. SSH into rescue system as root
3. Run installimage for Debian 12 Bookworm
4. Run the post-install script to configure networking and partitioning
5. Reboot into installed OS
6. Verify SSH access
## 5. Run baseline
Installs podman, diagnostic tools, creates the ops user, renders /etc/hosts,
enables dbus/chrony/podman.socket:
```bash
scripts/ansible-play.sh baseline.yml --limit sietch-ceph-<name>
```
## 6. Apply tuning
OS-level sysctl, hardware I/O schedulers, CPU governor:
```bash
scripts/ansible-play.sh tune-os.yml --limit sietch-ceph-<name>
scripts/ansible-play.sh tune-hardware.yml --limit sietch-ceph-<name>
```
## 7. Join node to Ceph cluster
This runs all deploy phases. For an existing cluster, bootstrap is skipped
(ceph.conf already exists on the bootstrap node). The node gets joined,
LVM is set up, OSDs are created, and services are placed:
```bash
scripts/ansible-play.sh deploy-ceph.yml \
--limit sietch-ceph-<name>,sietch-ceph-laurel
```
**Important:** You must include the bootstrap node (`sietch-ceph-laurel`)
in `--limit` because join, placement, OSD activation, and RGW service spec
updates all run from the bootstrap node.
## 8. Apply Ceph tuning and security hardening
```bash
scripts/ansible-play.sh tune-ceph.yml --limit sietch-ceph-laurel
scripts/ansible-play.sh harden.yml --limit sietch-ceph-<name>
```
## 9. Verify
### Check the node appears in the cluster
```bash
ssh ansible-iac@sietch-ceph-laurel
ceph orch host ls
```
Expected: new hostname listed with status empty (= online).
### Check OSDs are up
```bash
ceph osd tree
```
Expected: new node appears as a host bucket with its OSDs in `up` state.
### Check overall health
```bash
ceph status
```
Expected: `HEALTH_OK` or `HEALTH_WARN` with only backfill-related warnings
(which clear as data rebalances).
### Check RGW is running on new node
```bash
ceph orch ls --service-type rgw
```
Expected: running count incremented by 1.
### Run drift detection
```bash
mise run drift
```
Expected: no drift on the new node.
## Rollback
If the node needs to be removed:
```bash
# From bootstrap node
ceph orch host drain sietch-ceph-<name> --force
# Wait for daemons to migrate (~5 min)
ceph orch host rm sietch-ceph-<name> --force
```
Then remove the node from `inventory.ini` and delete its `host_vars` file.
@@ -0,0 +1,178 @@
# Runbook: Backup and Restore
**When:** Before major operations (upgrades, topology changes), on a
regular schedule, or during disaster recovery.
**Time estimate:** Backup: 2 minutes. Restore: depends on scenario.
---
## Backup
### Run a backup
```bash
mise run backup
```
Or directly:
```bash
scripts/ansible-play.sh backup-config.yml
```
### What gets captured
Backups are written to `backups/<timestamp>/` on the Ansible controller
(gitignored). Each backup contains:
| File | Contents |
|---|---|
| `ceph.conf` | Minimal cluster config (fsid, mon_host, auth settings) |
| `ceph.client.admin.keyring` | Admin authentication keyring |
| `crushmap.txt` | Decompiled CRUSH map (host/OSD topology and rules) |
| `osd-dump.json` | Full OSD map: pool definitions, PG counts, flags, weights |
| `mon-dump.json` | Monitor map: mon addresses, quorum members |
| `config-dump.json` | All runtime config overrides (`ceph config dump`) |
| `rgw-realm.json` | RGW realm, zonegroup, and zone configuration |
| `orch-services.yaml` | cephadm service specs (RGW, monitoring, crash, etc.) |
| `orch-hosts.yaml` | Cluster host list with addresses and labels |
### Backup schedule
No automated schedule is configured. Run manually:
- Before any cluster topology change (add/remove node or OSD)
- Before Ceph version upgrades
- Before CRUSH map modifications
- Weekly during active development
### Verify a backup
```bash
ls -la backups/$(ls -t backups/ | head -1)/
```
Check that all files are present and non-empty. The admin keyring and
ceph.conf are the most critical -- without them, cluster access is lost.
---
## Restore Scenarios
### Scenario 1: Lost ceph.conf / admin keyring on a single node
**Cause:** Accidental deletion, failed re-provision.
**Fix:** cephadm automatically distributes ceph.conf and the admin keyring
to managed hosts. Force redistribution:
```bash
# From bootstrap node
ceph cephadm config-check enable
ceph orch host rescan <hostname>
```
Or manually copy from the backup:
```bash
scp backups/<timestamp>/ceph.conf ansible-iac@<host>:/etc/ceph/ceph.conf
scp backups/<timestamp>/ceph.client.admin.keyring ansible-iac@<host>:/etc/ceph/ceph.client.admin.keyring
```
### Scenario 2: Lost admin keyring on ALL nodes
**Cause:** Full cluster purge without backup, or corruption.
**Fix:** Restore the keyring from the backup to the bootstrap node:
```bash
scp backups/<timestamp>/ceph.client.admin.keyring \
ansible-iac@sietch-ceph-laurel:/etc/ceph/
ssh ansible-iac@sietch-ceph-laurel
sudo chmod 600 /etc/ceph/ceph.client.admin.keyring
sudo chown ceph:ceph /etc/ceph/ceph.client.admin.keyring
```
Verify access is restored:
```bash
ceph status
```
### Scenario 3: CRUSH map corruption
**Cause:** Bad CRUSH rule edit, accidental tunables change.
**Fix:** Restore the CRUSH map from backup:
```bash
# Compile the decompiled map
crushtool -c backups/<timestamp>/crushmap.txt -o /tmp/crushmap.bin
# Inject it
ceph osd setcrushmap -i /tmp/crushmap.bin
```
**WARNING:** This overwrites the entire CRUSH topology. Any OSDs added
since the backup was taken will not be in the restored map.
### Scenario 4: RGW realm/zone misconfiguration
**Cause:** Bad radosgw-admin command, zone placement errors.
**Fix:** Use the backup as a reference to reconstruct:
```bash
cat backups/<timestamp>/rgw-realm.json | python3 -m json.tool
```
Then re-apply zone placement targets, zonegroup hostnames, etc. using
`radosgw-admin zone set` / `radosgw-admin zonegroup set` with the JSON
from the backup piped in.
### Scenario 5: Full cluster rebuild (total loss)
**Cause:** All nodes destroyed, starting from scratch.
1. Re-provision all nodes (see add-node runbook)
2. Run the full deploy pipeline:
```bash
mise run deploy
```
3. Restore configuration from backup:
```bash
# After bootstrap, apply saved CRUSH map
crushtool -c backups/<timestamp>/crushmap.txt -o /tmp/crushmap.bin
ceph osd setcrushmap -i /tmp/crushmap.bin
# Re-apply runtime config overrides
# Review config-dump.json and apply relevant settings
cat backups/<timestamp>/config-dump.json | python3 -c "
import sys, json
for item in json.load(sys.stdin):
section = item.get('section', 'global')
name = item.get('name', '')
value = item.get('value', '')
if name and section != 'mds':
print(f'ceph config set {section} {name} {value}')
"
```
**Note:** Object data (S3 objects, bucket contents) is stored on the OSDs
and cannot be restored from this config backup. This backup only preserves
cluster metadata and configuration. For data protection, rely on Ceph's
built-in replication (size=2+) and erasure coding.
### Scenario 6: Restore service specs after cephadm reset
```bash
ceph orch apply -i backups/<timestamp>/orch-services.yaml
```
This re-deploys RGW, monitoring, and crash daemons with the saved
placement and configuration.
@@ -0,0 +1,170 @@
# Runbook: Reprovision Painbox
**Status:** Validated 2026-04-26. Executed end-to-end during the yucca
monorepo import PR — painbox now runs as `painbox-ceph-evelyn` on
Bookworm with Ceph Tentacle deployed (15 OSDs up + in, all LUKS-encrypted).
This runbook captures the procedure for future reprovisions (disaster
recovery, OS upgrade, hardware refresh) and remains the canonical
Hetzner-installimage flow for any future SX-class clusters.
**When:** any future painbox reprovision — DR scenario, kernel/OS upgrade
that requires a fresh install, or hardware change that invalidates the
existing layout.
**Time estimate:** 45-60 minutes including installimage wait + reboot
+ post-deploy verification.
---
## Pre-flight
Canonical painbox identity (a future reprovision keeps this — a DR scenario
restores the same name on the same hardware):
| Item | Value |
|----------------|----------------------------------------------------------------|
| Inventory dir | `inventories/painbox-ceph.dev.hel.htz/` |
| Short hostname | `painbox-ceph-evelyn` |
| FQDN | `painbox-ceph-evelyn.dev.hel.htz.futo.cloud` |
| Public IP | `157.180.105.198` |
| OS image | Debian 12 Bookworm (Hetzner installimage tarball) |
| SSH key path | `~/.ssh/id_ed25519_painbox` (per [ADR-010](../adr/010-ssh-keys-in-1password.md)) |
1P items (`PAINBOX_CEPH_*`) already exist in `yucca_tf_dev`, including
the `PAINBOX_CEPH_ANSIBLE_IAC_SSH_KEY` SSH Key item.
> **Historical note:** the first reprovision under this identity ran
> 2026-04-26 as a rename from `painbox-osd-5c3cac.lab.hel.htz.futo.cloud`
> + old SSH key `~/.ssh/id_ed25519_ceph-painbox-lab-hel-htz` to the
> values above. Subsequent reprovisions (DR, OS refresh, etc.) restore
> the same identity — no rename involved.
## Steps
### 1. Install the ansible-iac SSH key on the operator workstation
Per [ADR-010](../adr/010-ssh-keys-in-1password.md), the keypair lives in
1P and is pulled to the workstation idempotently:
```bash
scripts/install-ssh-keys.sh painbox
# Writes ~/.ssh/id_ed25519_painbox (0600) + .pub (0644)
```
If the key was already installed from a previous session, this is a
no-op (fingerprint compare + skip). See [docs/scripts.md](../scripts.md)
for the wrapper's behavior.
### 2. Boot into Hetzner rescue
From Hetzner Robot panel: **Activate rescue system** → **Reboot**. SSH in
as root to the rescue IP.
> **Hetzner rotates the rescue root password on every activation.** The
> password shown in the Robot panel after activation is single-use for
> that rescue session — the next activation gets a fresh one. The 1Password
> entry `painbox-ceph-evelyn.dev.hel.htz.futo.cloud` (vault: `Yucca`) holds
> a *cached* copy from a prior activation; **update it from the Robot
> panel each time** before SSH'ing in, or the cached password fails auth.
> Rescue host keys also rotate per activation, so use `-o
> StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null` (or
> `ssh-keygen -R <ip>` first) when connecting.
### 3. Render + upload installimage scripts
`post-install.sh` is TEMPLATED — source of truth is `post-install.sh.tpl`,
which embeds an `op://` reference to the current
`PAINBOX_CEPH_ANSIBLE_IAC_SSH_KEY` pubkey. Render locally so rotations
propagate at reprovision time without edits:
```bash
cd ansible/ceph/inventories/painbox-ceph.dev.hel.htz/installimage
op inject -f -i post-install.sh.tpl -o /tmp/post-install.sh.rendered
```
Upload both files to the rescue system:
```bash
scp /tmp/post-install.sh.rendered root@<rescue-ip>:/tmp/post-install.sh
scp autosetup root@<rescue-ip>:/autosetup
```
### 4. Run installimage
```bash
# On the rescue system
chmod +x /tmp/post-install.sh
installimage -a -c /autosetup -x /tmp/post-install.sh
```
(`autosetup` targets `painbox-ceph-evelyn.dev.hel.htz.futo.cloud` as
HOSTNAME statically. If you ever rotate the hostname, edit autosetup
directly.)
### 5. Reboot into installed OS
```bash
reboot
```
SSH in with the new keypair:
```bash
ssh -i ~/.ssh/id_ed25519_painbox root@157.180.105.198 hostname -f
# Expect: painbox-ceph-evelyn.dev.hel.htz.futo.cloud
```
### 6. Update `yucca_tf_dev` state (if items were renamed)
The `PAINBOX_CEPH_*` items in `yucca_tf_dev` are already correctly named —
no action needed. Disaster-recovery items (`PAINBOX_CEPH_RGW_TLS_CERT`,
`_RGW_TLS_KEY`, `_CLIENT_ADMIN_KEYRING`) are upserted by `mise run capture` after
a successful deploy (step 8).
### 7. Run ansible deploy
```bash
cd ~/yucca/ansible/ceph
export CEPH_ENV=inventories/painbox-ceph.dev.hel.htz/inventory.ini
mise run preflight # should pass
mise run deploy # baseline → tune → deploy → tune → harden
```
### 8. Verify
```bash
ssh ansible-iac@painbox-ceph-evelyn.dev.hel.htz.futo.cloud sudo ceph -s
# HEALTH_OK (single-node cluster)
```
## Rollback
If reprovisioning fails at installimage or post-install, Hetzner rescue
is still accessible:
1. Activate rescue, SSH in.
2. Re-run installimage with the OLD scripts (preserve them in
`docs/archive/painbox-installimage-2026-04-xx/` before overwriting).
3. Old hostname + old SSH key restore the prior-state box.
4. No cluster data is on this box (painbox never held Ceph data),
so nothing to restore.
## Post-reprovision cleanup
- Delete the old SSH key file `~/.ssh/id_ed25519_ceph-painbox-lab-hel-htz*`
from the operator workstation once verification is complete.
- If any external DNS records pointed at `painbox-osd-5c3cac.lab.*`,
update them to `painbox-ceph-evelyn.dev.*`.
(No edit to `post-install.sh` is required — the committed source of
truth is `post-install.sh.tpl`, which embeds a live `op://` reference
to the current `PAINBOX_CEPH_ANSIBLE_IAC_SSH_KEY/public_key`.
Re-rendering via `op inject -f` always picks up whatever is current in
1P at the moment of reprovision.)
## References
- `docs/runbooks/add-node.md` §Hetzner — the general installimage flow
- `docs/adr/010-ssh-keys-in-1password.md` — SSH keypair lifecycle
- `docs/scripts.md` — `install-ssh-keys.sh` reference
- `inventories/painbox-ceph.dev.hel.htz/installimage/` — scripts
@@ -0,0 +1,164 @@
# Runbook: Recover from a Bad `tofu apply`
**When:** state corruption, drift, rendered files don't match cluster spec,
or TF destroyed an item/file that shouldn't have been touched.
**Time estimate:** 5-30 minutes depending on severity.
---
## Common failure modes and their fixes
### Missing or unreadable state file
```
tofu init
# Error: failed to load state: ...
```
**Cause:** S3 credentials unresolved, bucket object missing, or the backend
config differs from what the state was written with.
**Fix:**
```bash
# Verify S3 credentials resolve via op run
op run --env-file=tf/.env -- env | grep AWS_
# Expect AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY populated
# Verify the state object exists in the bucket
op run --env-file=tf/.env -- \
aws --endpoint-url=https://s3.eu-west-par.io.cloud.ovh.net/ \
s3 ls s3://yucca-tf-state/ceph/dev/ceph/
# If present: re-init should pick it up
mise run tf:init
# If the object is missing: state was never created or was deleted. You
# can recover by re-applying — TF recreates local_file resources
# (idempotent, same content; no 1P items are harmed because the module's
# onepassword_item resources are currently dormant — see ADR-009 §2).
mise run tf:apply
```
### Rendered file on disk doesn't match tfvars
```bash
cat ansible/ceph/inventories/sietch-ceph.dev.austin.int/inventory.ini
# says ansible_user=root but tfvars says ansible-iac
```
**Cause:** someone hand-edited the rendered file; TF state shows it
unchanged; operator is surprised.
**Fix:**
```bash
mise run tf:plan # TF shows drift
mise run tf:apply # TF overwrites with correct content
```
All TF-rendered files are gitignored — the single source of truth is
`clusters.auto.tfvars`. Hand-edits are ephemeral.
### Wrong vault referenced in rendered secrets.yml.tpl
**Cause:** `vault` field in clusters.auto.tfvars typo'd or set to a
vault you don't have access to.
**Fix:** Fix the tfvars entry, `mise run tf:apply`. `scripts/ansible-play.sh`
will fail loudly on the next run (op inject exits non-zero on unresolvable
references — per ADR-009 fail-closed principle).
### Wordlist auto-pick renamed a deployed host
**Cause:** `name_seed` bumped, or a new auto-named host was prepended such
that existing auto-name indices shifted.
**Symptom:**
```bash
mise run tf:plan
# Plan shows: ~inventory.ini content with hostname change from painbox-ceph-evelyn to painbox-ceph-<other>
```
**Do NOT apply** — renaming a deployed host cascades into SSH known_hosts,
cephadm host registration, certs, 1P item names, DNS.
**Fix:** pin the existing name by adding `name = "evelyn"` to the host
entry in `clusters.auto.tfvars`, then apply. The pinned name takes
precedence over the shuffle output.
### Item disappeared from yucca_tf_dev
**Cause:** an operator, another consumer of `yucca_tf_dev`, or a TF
delete-on-destroy run removed an item the Ansible side depends on.
**Symptom:** `op inject -f -i secrets.yml.tpl` fails with "item not found".
**Fix:**
```bash
# Re-create the item manually via superuser SA
SU_TOKEN=$(op read "op://yucca_tf_dev/yucca_futo_1pass_superuser_service_account/password")
OP_SERVICE_ACCOUNT_TOKEN="$SU_TOKEN" op item create \
--vault yucca_tf_dev \
--category password \
--title SIETCH_CEPH_OPS_PASSWORD \
--generate-password='letters,digits,32'
unset SU_TOKEN
```
Then either (a) run `mise run deploy` to push the new password to the
cluster (baseline role converges ops user pw), or (b) if the cluster
already has the old value and you want to keep it, retrieve from a
backup and set the item via `op item edit password=...`.
### TF state object corrupted or lost
The state lives in S3 (`yucca-tf-state` bucket, key
`ceph/${env}/${stack}/terraform.tfstate`). Recovery options in order of
preference:
**Option A — roll back via S3 versioning.** The bucket has versioning
enabled; list prior versions and restore the last-known-good:
```bash
op run --env-file=tf/.env -- \
aws --endpoint-url=https://s3.eu-west-par.io.cloud.ovh.net/ \
s3api list-object-versions \
--bucket yucca-tf-state \
--prefix ceph/dev/ceph/terraform.tfstate
# Identify the VersionId of a good snapshot, then:
op run --env-file=tf/.env -- \
aws --endpoint-url=https://s3.eu-west-par.io.cloud.ovh.net/ \
s3api copy-object \
--bucket yucca-tf-state \
--copy-source 'yucca-tf-state/ceph/dev/ceph/terraform.tfstate?versionId=<VID>' \
--key ceph/dev/ceph/terraform.tfstate
```
**Option B — re-apply from clean state.** Delete the state object and
re-init + re-apply. TF recreates the `local_file` resources (idempotent,
same content). Safe today because `onepassword_item` resources are
dormant (see ADR-009 §2).
```bash
op run --env-file=tf/.env -- \
aws --endpoint-url=https://s3.eu-west-par.io.cloud.ovh.net/ \
s3 rm s3://yucca-tf-state/ceph/dev/ceph/terraform.tfstate
mise run tf:init
mise run tf:apply
```
Once ADR-009 §2 ships and 1P items are TF-owned, Option B becomes
destructive — recovery will require `terraform import` of each item
against the new state. Document that path when re-enabling.
## References
- ADR-009 — fail-closed principle
- `tf/README.md` §"State backend" — bucket, endpoint, key path rationale
- `architecture.md` §4 — what TF owns vs renders
@@ -0,0 +1,80 @@
# Runbook: Remote Hands Access
**When:** a remote-hands operator (on-site at the datacenter) needs to
run playbooks against the cluster.
**Time estimate:** 10 minutes of setup per operator, one-time. 5 minutes
per remote-hands session after that.
---
## The access model
Remote-hands operators don't touch 1P desktop sessions. They authenticate
via a scoped service-account token, set as `OP_SERVICE_ACCOUNT_TOKEN` in
their shell environment. `scripts/ansible-play.sh` picks it up automatically
and uses it to resolve `secrets.yml.tpl`.
Which SA: the **read-only** `yucca_futo_1pass_service_account` in
`yucca_tf_dev`. Remote hands shouldn't have write access to 1P items —
least-privilege principle.
## One-time setup (per operator)
1. **Grant the operator access to the Immich 1P group** — the group's
1P admin does this. Gives them read on `yucca_tf_dev` (enough to read
the SA token and SSH key items).
2. **Operator retrieves the SA token**:
```bash
op read "op://yucca_tf_dev/yucca_futo_1pass_service_account/password"
```
3. **Operator sets it in their shell profile** (`~/.bashrc` or equivalent):
```bash
export OP_SERVICE_ACCOUNT_TOKEN="ops_eyJ...." # 862-char token
```
4. **Operator clones yucca**, installs mise tooling, runs:
```bash
cd ~/yucca/ansible/ceph
mise trust && mise run setup
```
5. **Operator installs the ansible-iac SSH keys** from 1P:
```bash
scripts/install-ssh-keys.sh
```
Lands `~/.ssh/id_ed25519_sietch` and `~/.ssh/id_ed25519_painbox` (0600).
6. **Verify**:
```bash
mise run preflight # should pass; no desktop 1P required
```
## Per-session workflow
```bash
cd ~/yucca/ansible/ceph
export CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory.ini
mise run status # read-only smoke test
mise run deploy # or any other task
```
No auth prompts — the SA token in env gets picked up by `op inject` inside
`scripts/ansible-play.sh`.
## What remote hands CAN'T do
The read-only SA token can't:
- Modify 1P items (e.g., rotate secrets).
- Run `tf:apply` (needs superuser SA for writes).
Actions requiring writes must route through the primary operator or
come with explicit superuser-token provisioning.
## Rotation
If a remote-hands operator leaves the project, rotate the read-only SA
token (see `rotate-sa-token.md`). Their environment still holds the old
token but it's invalidated in 1P; the next command fails loudly.
## References
- `tf/README.md` §"Where secrets actually live" — SA scopes
- `docs/secrets.md` — full secrets model
+246
View File
@@ -0,0 +1,246 @@
# Runbook: Replace a Failed HDD
**When:** SMART failure, unresponsive OSD, or predictive disk replacement.
**Time estimate:** 15-30 minutes (plus backfill time, which varies by data volume).
**Prerequisites:**
- Physical or remote hands access for the disk swap
- Cluster has enough free capacity to absorb the missing OSD during backfill
---
## 1. Identify the failed disk
### Check cluster health
```bash
ceph health detail
```
Look for messages like:
- `OSD_DOWN` -- `1 osds down`
- `DEVICE_HEALTH_TOOMANY` -- `1 device(s) expected to fail soon`
- `PG_DEGRADED` -- degraded placement groups
### Find the OSD ID and host
```bash
ceph osd tree
```
Note the OSD ID (e.g., `osd.7`) and the host it belongs to.
### Check SMART data
SSH to the host and find the physical device:
```bash
# Find the device backing the OSD
ceph-volume lvm list | grep -A5 "osd.7"
# Or find by path
ceph device ls | grep osd.7
# Check SMART
smartctl -a /dev/sdX
```
Look for: `Reallocated_Sector_Ct`, `Current_Pending_Sector`,
`SMART overall-health self-assessment test result: FAILED`.
## 2. Mark OSD out (start draining)
```bash
ceph osd out osd.7
```
This begins backfilling data away from the OSD. Monitor progress:
```bash
ceph -s
# or watch:
ceph -w
```
Wait until backfill completes and all PGs are `active+clean`:
```bash
ceph pg stat
```
Expected output includes `active+clean` for all PGs, zero `remapped` or
`backfilling`.
**Do not proceed until backfill is complete.** Pulling a disk during
backfill risks data loss if another disk fails simultaneously.
## 3. Stop and purge the OSD
```bash
# Stop the OSD daemon
ceph orch daemon stop osd.7
# Purge OSD from cluster (removes from CRUSH, auth keys, etc.)
ceph osd purge osd.7 --yes-i-really-mean-it
```
Verify removal:
```bash
ceph osd tree
```
The OSD should no longer appear.
## 4. Close LUKS and clean up on the host
SSH to the host where the OSD lived:
```bash
ssh ansible-iac@sietch-ceph-<host>
sudo -i
```
### Close the LUKS mapping
```bash
# List dm-crypt mappings to find the right one
dmsetup ls --target crypt
# Close it (name will be something like ceph-<uuid>-...-block-dmcrypt)
cryptsetup close <mapping-name>
```
### Remove LVM artifacts
```bash
# Find and remove orphaned PV on the disk
pvs | grep /dev/sdX
pvremove -f /dev/sdX
```
### Wipe device signatures
```bash
wipefs -af /dev/sdX
dd if=/dev/zero of=/dev/sdX bs=1M count=10
```
## 5. Physical disk swap
### Austin (on-site)
1. Identify the bay number from the SAS PHY mapping in host_vars
2. The disk bay maps to a by-path device: `/dev/disk/by-path/<sas_path_prefix>-phy<N>-lun-0`
3. Turn on the drive bay LED if available via iDRAC / `ledctl`
4. Coordinate with the remote hands team for the physical swap
5. Hot-swap the drive -- these are SAS/SATA hot-plug bays
6. Verify the new disk appears:
```bash
lsblk
ls /dev/disk/by-path/ | grep phy<N>
```
### Hetzner (remote)
Requires a support ticket to Hetzner for disk replacement. Schedule
maintenance window.
## 6. Re-create the OSD
The new disk must get an encrypted OSD with a block.db LV on the SSD,
matching the original configuration.
### Verify the block.db LV still exists
```bash
lvs | grep db-slot<N>
```
If the LV was destroyed, re-run LVM setup first:
```bash
scripts/ansible-play.sh deploy-ceph.yml \
--tags lvm --limit sietch-ceph-<host>,sietch-ceph-laurel
```
### Create the OSD manually
```bash
# Ensure bootstrap-osd keyring is present
ceph auth get client.bootstrap-osd -o /var/lib/ceph/bootstrap-osd/ceph.keyring
# Create encrypted OSD with block.db
DISK="/dev/disk/by-path/<sas_path_prefix>-phy<N>-lun-0"
DB_LV="<vg-name>/db-slot<N>"
cephadm ceph-volume \
--keyring /var/lib/ceph/bootstrap-osd/ceph.keyring \
lvm create --dmcrypt --no-systemd \
--data "$DISK" --block.db "$DB_LV"
```
### Activate the OSD
From the bootstrap node:
```bash
ceph cephadm osd activate <hostname>
```
### Or re-run via Ansible
Alternatively, re-run the OSD creation phase through Ansible (idempotent --
skips existing OSDs):
```bash
scripts/ansible-play.sh deploy-ceph.yml \
--tags osds --limit sietch-ceph-<host>,sietch-ceph-laurel
```
## 7. Verify
### Check the new OSD is up
```bash
ceph osd tree
```
Expected: new OSD appears under the correct host with status `up` and
weight > 0.
### Check reweight
```bash
# If reweight is 0, fix it
ceph osd tree | grep "osd.<new-id>"
# If needed:
ceph osd reweight <new-id> 1.0
```
### Monitor backfill to the new OSD
```bash
ceph -s
```
Wait for `active+clean` on all PGs.
### Verify dmcrypt
```bash
ceph config-key dump | grep dm-crypt | wc -l
```
Should show one more key than before the replacement.
### Verify SMART on new disk
```bash
smartctl -a /dev/sdX
```
Confirm the new disk has zero errors.
+107
View File
@@ -0,0 +1,107 @@
# Runbook: Replace a Host (Hardware Swap, Preserving Name)
**When:** chassis failure, motherboard swap, or preventative hardware
refresh. You want the new box to assume the old host's identity — same
hostname, same IP, same Ceph OSD identities.
**Time estimate:** 1-2 hours (bare-metal) / 30 min (Hetzner).
---
## Preserving name = preserving trust
Keeping the old hostname avoids:
- Re-issuing the ansible-iac SSH key to a new hostname
- Rotating the host's Ceph auth keys
- Updating monitoring dashboards / alert rules with new labels
- Updating any external systems that refer to the hostname
The TF inventory declaration stays unchanged — cluster name, host name,
bond IP all match. Only the physical box changes.
## Steps
### 1. Mark OSDs out, wait for backfill
```bash
# SSH to bootstrap node
sudo ceph osd out $(sudo ceph osd ls-tree <hostname-being-replaced>)
sudo ceph -w # watch for HEALTH_OK or acceptable degraded state
```
Depending on cluster size, full backfill can take hours. For time-critical
replacements, skip this step and accept degraded state during rebuild —
but confirm you have headroom.
### 2. Power down old host, rack new one in same slot
Preserve the physical network cabling and IPMI address. The new box gets
the same bond IP via the same DHCP reservation / static config.
### 3. Re-run provisioning against the new host
**Bare-metal (sietch)**:
```bash
# Boot the new box from the Debian 12 live image (same procedure as first
# provision — see docs/runbooks/add-node.md for iDRAC steps).
CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory-provision.ini \
scripts/ansible-play.sh provision.yml \
-e confirm_wipe=true \
--limit sietch-ceph-<name>
```
**Hetzner (painbox)**: reboot into rescue, run installimage + post-install
scripts from `inventories/painbox-ceph.dev.hel.htz/installimage/`.
### 4. Clear the old host's Ceph state
The cluster still has the old host's OSDs, CRUSH entries, and cephadm
host record. Clean those up BEFORE the new box rejoins:
```bash
# On bootstrap node:
for osd in $(sudo ceph osd ls-tree <hostname>); do
sudo ceph osd purge $osd --yes-i-really-mean-it
done
sudo ceph osd crush remove <hostname>
sudo ceph orch host rm <hostname> --force
```
### 5. Re-run baseline + deploy for that host
```bash
scripts/ansible-play.sh baseline.yml --limit <hostname>
scripts/ansible-play.sh deploy-ceph.yml \
--limit <hostname>,<bootstrap-hostname>
```
Note bootstrap host must be in `--limit` — join/placement/OSD activation
all run from bootstrap.
### 6. Verify
```bash
# All OSDs for the replaced host are up+in
sudo ceph osd tree | grep <hostname>
# Cluster is rebalancing or healthy
sudo ceph -s
```
## Gotchas
- **Hardware topology changed**: new chassis may have different PCI paths
for SAS controllers / NVMe. Update `host_vars/<hostname>.yml`
(`sas_path_prefix`, `ceph_hdd_osds`) before step 5.
- **Different serial numbers**: acceptable — `by-path` is used for OSD
slot identity, not `by-id`.
- **Backfill thundering herd**: if you skipped step 1, the new OSDs enter
the cluster and backfill aggressively. Throttle with
`ceph_osd_recovery_max_active` in vars.yml if you see client-IO impact.
## References
- `docs/runbooks/add-node.md` — related but for net-new hosts
- `docs/runbooks/replace-disk.md` — single-disk replacement within a host
+121
View File
@@ -0,0 +1,121 @@
# Runbook: Rotate RGW TLS Certificates
**When:** Certificate approaching expiry, compromised key material, or SAN
changes (new nodes added, DNS name changed).
**Time estimate:** 5 minutes. Brief RGW restart causes ~10s S3 downtime.
**Prerequisites:**
- `op` session live (desktop unlocked or `OP_SERVICE_ACCOUNT_TOKEN` set)
- Cluster is healthy
---
## 1. Check current certificate expiry
```bash
ssh ansible-iac@sietch-ceph-laurel
sudo openssl x509 -in /etc/ceph/rgw-ssl.crt -noout -subject -dates -ext subjectAltName
```
Output shows:
```
subject=C = US, ST = Texas, L = Austin, O = FUTO, CN = s3.dev.austin.int.futo.cloud
notBefore=...
notAfter=...
X509v3 Subject Alternative Name:
DNS:s3.dev.austin.int.futo.cloud, DNS:*.s3.dev.austin.int.futo.cloud, ...
```
## 2. Run the rotation playbook
```bash
scripts/ansible-play.sh rotate-certs.yml
```
### What this does
1. **Deletes** `/etc/ceph/rgw-ssl.crt` and `/etc/ceph/rgw-ssl.key` on the
bootstrap node
2. **Re-runs the RGW role** (`roles/ceph_deploy/tasks/rgw.yml`) which:
- Generates a new 4096-bit RSA self-signed cert with 10-year validity
- SANs include: the canonical DNS name, wildcard for virtual-hosted
buckets, per-node FQDNs, and per-node bond IPs
- Renders the RGW service spec with the new cert embedded
- Applies the service spec via `ceph orch apply` -- cephadm distributes
the cert to all RGW daemon containers
3. **Restarts all RGW daemons** via `ceph orch restart rgw` to load the
new cert
4. **Displays** the new certificate subject, dates, and SANs
### Impact
- RGW daemons restart sequentially. S3 requests will fail for ~10 seconds
during the restart window.
- Clients using the old self-signed cert in their trust store will need the
new cert. Export with:
```bash
ssh ansible-iac@sietch-ceph-laurel sudo cat /etc/ceph/rgw-ssl.crt
```
## 3. Verify after rotation
### Check new cert details
The playbook prints this, but to verify manually:
```bash
ssh ansible-iac@sietch-ceph-laurel
sudo openssl x509 -in /etc/ceph/rgw-ssl.crt -noout -subject -dates -ext subjectAltName
```
### Test RGW endpoint
```bash
# From a node in the cluster (self-signed cert)
curl -k https://s3.dev.austin.int.futo.cloud:443/
```
Expected: XML response with `ListAllMyBucketsResult` or `AccessDenied`
(both mean RGW is serving TLS correctly).
### Test direct node access
```bash
curl -k https://10.10.10.90:443/
```
### Check RGW daemons are running
```bash
ceph orch ls --service-type rgw
```
Expected: running count matches the number of ceph_nodes (currently 3).
### Check dashboard can reach RGW
Open `https://<bootstrap-ip>:8443` and navigate to Object Gateway.
The page should load without 500 errors (dashboard has RGW API SSL
verification disabled for self-signed certs).
## Certificate configuration
The cert parameters are controlled by these variables in
`inventories/sietch-ceph.dev.austin.int/group_vars/all/vars.yml`:
| Variable | Default | Purpose |
|---|---|---|
| `ceph_rgw_ssl` | `true` | Enable TLS on RGW frontend |
| `ceph_rgw_ssl_cert_days` | `3650` | Validity period (10 years) |
| `ceph_rgw_ssl_cert_subject_c` | `US` | Country |
| `ceph_rgw_ssl_cert_subject_st` | `Texas` | State |
| `ceph_rgw_ssl_cert_subject_l` | `Austin` | Locality |
| `ceph_rgw_ssl_cert_subject_o` | `FUTO` | Organization |
| `ceph_rgw_ssl_cert_email` | `yucca@futo.org` | Contact email |
| `ceph_rgw_dns_name` | `s3.dev.austin.int.futo.cloud` | CN and primary SAN |
SANs are auto-generated from inventory: per-node FQDNs and bond IPs are
included so direct-host access validates.
@@ -0,0 +1,70 @@
# Runbook: Rotate 1Password Service Account Token
**When:** token leak, team personnel change, or scheduled rotation.
**Time estimate:** 15-30 minutes including cross-team coordination.
**Blast radius:** **cross-project**. The SA tokens in `yucca_tf_dev` are
shared with o11y and other Futo consumers. Rotating breaks every consumer
until they pick up the new token.
---
## Prerequisites
- Coordination with other consumers of the SA (at minimum: o11y team,
whoever else uses `yucca_tf_*`). Ask in Discord before rotating.
- 1Password access to create a replacement SA token.
- List of repos/CI pipelines that consume the token — so you know who to
notify to pick up the rotation.
## Which SA to rotate
Two SAs live in `yucca_tf_dev`:
| SA | Scope |
|---|---|
| `yucca_futo_1pass_superuser_service_account` | read+write all yucca_tf_* vaults |
| `yucca_futo_1pass_service_account` | read-only on yucca_tf + yucca_tf_dev |
Sietch-ceph consumes both (TF uses superuser for writes, Ansible uses
read-only for runtime). Read-only rotation has lower blast radius.
## Steps
1. **Announce in Discord** to Immich maintainers + o11y team: "Rotating
`yucca_futo_1pass_<type>_service_account` at HH:MM. Expect a brief
window where CI pipelines fail — update your vars after."
2. **Create a new SA token in 1Password Admin UI** with matching scope.
Give it a dated suffix (e.g. `-2026-04-22`) so old and new coexist
briefly.
3. **Update the SA's `password` field** in the existing 1P item (so
consumers reading via `op://yucca_tf_dev/<name>/password` pick up the
new value with no code change).
4. **Test from this repo**:
```bash
mise run tf:plan # should succeed with new superuser token
scripts/ansible-play.sh status.yml # should succeed with new read-only
```
5. **Notify consumers** the rotation is done — they restart any daemons /
re-pull tokens as needed.
6. **Revoke the old SA token** in 1Password Admin UI once all consumers
confirm green.
## Recovery: rotation broke something
- If the new token doesn't work: check the new SA has the same vault
scopes as the old. Empty `op vault list` output under the new token =
scope misconfigured.
- If a consumer is still using the old token: the item's `password` field
is the single source of truth; consumers re-reading pick up the new
value. If a consumer caches the token in its own env (CI secret,
systemd EnvironmentFile), update that manually.
- Roll back: edit the 1P item's `password` field back to the old token
value (if you kept a copy). The old SA is still valid until
explicitly revoked.
## References
- `tf/.env` consumes `op://yucca_tf_dev/yucca_futo_1pass_superuser_service_account/password`
- `tf/README.md` §"Where secrets actually live" — describes the SA scopes
@@ -0,0 +1,201 @@
# Runbook: Rotate Passwords and Keys
**When:** scheduled rotation, suspected compromise, or personnel change.
**Time estimate:** 5 minutes per password (plus any service-restart windows).
**Prerequisites:**
- `op` CLI authenticated (desktop session unlocked or `OP_SERVICE_ACCOUNT_TOKEN` set)
- Inventory and secrets template rendered (`tofu apply` has been run at least once)
---
## Model
1Password is the sole source of truth. Rotation = edit the item's `password`
field, done — no re-encrypt, no commit, no sync step. The next `mise run
deploy` picks up the new value via `op inject`.
## Secrets managed by this cluster
| Secret | 1Password item | Vault | Where it's applied |
|---------------------------|----------------------------------------|----------------|--------------------------------------------|
| `ops` user password | `<CLUSTER>_CEPH_OPS_PASSWORD` | `yucca_tf_dev` | Linux login on all nodes |
| Ceph dashboard password | `<CLUSTER>_CEPH_DASHBOARD_PASSWORD` | `yucca_tf_dev` | `https://<node>:8443` admin login |
| Grafana admin password | `<CLUSTER>_CEPH_GRAFANA_PASSWORD` | `yucca_tf_dev` | `https://<node>:3000` admin login |
| S3 service-user access | `<CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY` | `yucca_tf_dev` | RGW S3 user `yucca-restic` access key |
| S3 service-user secret | `<CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY` | `yucca_tf_dev` | RGW S3 user `yucca-restic` secret key |
Replace `<CLUSTER>` with `SIETCH` or `PAINBOX`. The active vault name is
declared per-cluster in the `vault` field of the cluster's entry in
`tf/deployment/dev/ceph/clusters.auto.tfvars` — `yucca_tf_dev` for dev,
future `yucca_tf_staging` / `yucca_tf` for staging/prod.
## 1. Rotate in 1Password
Either edit via the desktop app, or from CLI:
```bash
# Generate + set a new password in one shot
op item edit SIETCH_CEPH_DASHBOARD_PASSWORD --vault yucca_tf_dev --generate-password='letters,digits,32'
```
Or set an explicit value:
```bash
op item edit SIETCH_CEPH_DASHBOARD_PASSWORD --vault yucca_tf_dev password=<literal-new-value>
```
Verify the new value resolves:
```bash
op read "op://yucca_tf_dev/SIETCH_CEPH_DASHBOARD_PASSWORD/password" | wc -c
```
## 2. Apply to the cluster
### ops user password
The baseline role converges the ops user password on every run:
```bash
scripts/ansible-play.sh baseline.yml --tags users
```
The old password stops working immediately.
### Dashboard password
Not re-applied by normal deploys. Set directly on the bootstrap MON:
```bash
ssh ansible-iac@sietch-ceph-laurel \
sudo ceph dashboard ac-user-set-password admin \
"$(op read 'op://yucca_tf_dev/SIETCH_CEPH_DASHBOARD_PASSWORD/password')"
```
Expect `User admin updated`.
### Grafana admin password
Re-run the monitoring tag:
```bash
scripts/ansible-play.sh deploy-ceph.yml --tags monitoring
```
## 3. Verify
| Check | Command / action |
|---|---|
| ops login | `ssh ops@sietch-ceph-laurel` — new password prompts and works |
| Dashboard | browse `https://<bootstrap-ip>:8443`, log in as `admin` with new password |
| Grafana | browse `https://<bootstrap-ip>:3000`, log in as `admin` with new password |
## 4. Clean up
No cleanup needed on the secrets side — old values are gone from 1Password
the moment you save the new ones. No vault.yml re-encryption, no git commit
required for rotation.
---
## Rotating SSH keys
SSH keys for `ansible-iac` live in `yucca_tf_dev` as category `SSH Key`
items titled `<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY`. Rotation is
forward-only — generate a new key in 1P, distribute the new pubkey,
retire the old key after confidence.
### Steps (sietch example; same pattern for any cluster)
1. **Generate the new keypair natively in 1P**. The current item must be
replaced (1P doesn't support multiple key versions per item).
Snapshot the old item first for rollback:
```bash
# Get the new superuser SA token
SU_TOKEN=$(op read "op://yucca_tf_dev/yucca_futo_1pass_superuser_service_account/password")
# Save the old item as a dated rollback copy (private_key is derivable
# via op item get; use --reveal if you need it on disk)
OP_SERVICE_ACCOUNT_TOKEN="$SU_TOKEN" op item edit SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY \
--vault yucca_tf_dev \
--title "SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY_RETIRED_$(date +%Y%m%d)"
# Create the replacement with the canonical title
OP_SERVICE_ACCOUNT_TOKEN="$SU_TOKEN" op item create \
--vault yucca_tf_dev --category "SSH Key" \
--title SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY \
--ssh-generate-key=ed25519
unset SU_TOKEN
```
2. **Install the new private key on your workstation**
(`scripts/install-ssh-keys.sh` refuses to overwrite mismatched
fingerprints — move the old file aside first):
```bash
mv ~/.ssh/id_ed25519_sietch ~/.ssh/id_ed25519_sietch.retired-$(date +%Y%m%d)
mv ~/.ssh/id_ed25519_sietch.pub ~/.ssh/id_ed25519_sietch.pub.retired-$(date +%Y%m%d)
scripts/install-ssh-keys.sh sietch
```
3. **Distribute the new pubkey to every node** (non-destructive;
old key stays in authorized_keys until step 5):
```bash
scripts/ansible-play.sh rotate-ssh-key.yml
```
4. **Test** SSH with the new key:
```bash
ssh -i ~/.ssh/id_ed25519_sietch ansible-iac@sietch-ceph-laurel hostname -f
```
5. **Remove the old key from authorized_keys** once confident (manual
— no playbook for this yet). On each node:
```bash
ssh ansible-iac@sietch-ceph-laurel \
"sed -i '/RETIRED-KEY-COMMENT/d' ~/.ssh/authorized_keys"
```
Match by key comment (e.g., the email/hostname in the pubkey's
trailing field).
6. **Delete the retired 1P item** (optional; keep ~30 days for
audit/rollback):
```bash
SU_TOKEN=$(op read "op://yucca_tf_dev/yucca_futo_1pass_superuser_service_account/password")
OP_SERVICE_ACCOUNT_TOKEN="$SU_TOKEN" op item delete \
--vault yucca_tf_dev "SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY_RETIRED_<date>"
unset SU_TOKEN
```
### Provisioning a fresh host uses the current key
`roles/provision_host/tasks/admin_user.yml` reads
`{{ provision_iac_ssh_key_path }}.pub` (e.g., `~/.ssh/id_ed25519_sietch.pub`
on operator disk) and installs it as the bootstrap `authorized_keys`
during Debian live-image provisioning. Make sure `install-ssh-keys.sh`
has run on any workstation that'll drive provisioning.
---
## Future: rotating TF-provisioned secrets
Once the sietch-ceph service account lands and `secrets.tf.disabled` is
re-enabled, rotations become a `terraform taint` + `apply`:
```bash
cd tf/deployment/dev/ceph
terragrunt taint 'module.cluster["sietch"].onepassword_item.secret["dashboard"]'
terragrunt apply
```
This regenerates the password in 1Password; the apply step to the live
cluster (dashboard / grafana commands above) is still required.
+268
View File
@@ -0,0 +1,268 @@
# S3 Integration Guide
Audience: Application developers (Yucca, Immich, Restic, internal tooling).
## Endpoints
The cluster runs Ceph RGW (RADOS Gateway) on every node behind a self-signed
wildcard TLS certificate on **port 443**.
| Style | URL |
|---|---|
| Path-style | `https://s3.dev.austin.int.futo.cloud/<bucket>/<key>` |
| Virtual-hosted | `https://<bucket>.s3.dev.austin.int.futo.cloud/<key>` |
| Direct (per-node) | `https://10.10.10.90:443`, `https://10.10.10.91:443`, `https://10.10.10.92:443` |
Region: **us-east-1**
Path-style is recommended for simplicity. Virtual-hosted requires wildcard DNS
(see DNS section below).
## Getting credentials
### Option A: 1Password (preferred)
S3 credentials for the `svc-yucca-restic` service account are stored in
1Password after initial deployment. Ask the infrastructure team for access to
the "Ceph S3" vault entry.
### Option B: radosgw-admin (infra operators only)
SSH to the bootstrap node (sietch-ceph-laurel) and run:
```bash
radosgw-admin user info --uid=svc-yucca-restic
```
The `keys[0].access_key` and `keys[0].secret_key` fields contain the
credentials.
To create a new service account:
```bash
radosgw-admin user create \
--uid=svc-myapp \
--display-name='myapp service account' \
--max-buckets=100
```
## Self-signed certificate handling
The cluster uses a self-signed wildcard certificate. Every client must either
trust the CA or disable TLS verification.
### Trusting the cert (recommended for production workloads)
Copy the cert from the bootstrap node:
```bash
scp ansible-iac@10.10.10.90:/etc/ceph/rgw-ssl.crt ./rgw-ssl.crt
```
Then pass it to your client (examples below).
### Disabling verification (quick testing only)
Pass `--no-verify-ssl` (AWS CLI) or `verify=False` (boto3). Fine for
benchmarking, not for production.
## AWS CLI configuration
### ~/.aws/credentials
```ini
[sietch]
aws_access_key_id = YOUR_ACCESS_KEY
aws_secret_access_key = YOUR_SECRET_KEY
```
### ~/.aws/config
```ini
[profile sietch]
region = us-east-1
endpoint_url = https://s3.dev.austin.int.futo.cloud
s3 =
signature_version = s3v4
addressing_style = path
```
### Basic operations
```bash
# List buckets
aws --profile sietch --no-verify-ssl s3 ls
# Create a bucket
aws --profile sietch --no-verify-ssl s3 mb s3://my-bucket
# Upload a file
aws --profile sietch --no-verify-ssl s3 cp ./file.txt s3://my-bucket/file.txt
# List objects
aws --profile sietch --no-verify-ssl s3 ls s3://my-bucket/
# Download
aws --profile sietch --no-verify-ssl s3 cp s3://my-bucket/file.txt ./downloaded.txt
# Using the CA bundle instead of --no-verify-ssl
aws --profile sietch --ca-bundle ./rgw-ssl.crt s3 ls
```
## boto3 (Python)
```python
import boto3
import botocore
import urllib3
# Suppress InsecureRequestWarning when verify=False
urllib3.disable_warnings(urllib3.exceptions.InsecureRequestWarning)
s3 = boto3.client(
"s3",
endpoint_url="https://s3.dev.austin.int.futo.cloud",
aws_access_key_id="YOUR_ACCESS_KEY",
aws_secret_access_key="YOUR_SECRET_KEY",
region_name="us-east-1",
verify=False, # or path to rgw-ssl.crt
config=botocore.config.Config(
signature_version="s3v4",
s3={"addressing_style": "path"},
),
)
# Create a bucket
s3.create_bucket(Bucket="my-bucket")
# Upload
s3.put_object(Bucket="my-bucket", Key="hello.txt", Body=b"hello world")
# Download
obj = s3.get_object(Bucket="my-bucket", Key="hello.txt")
data = obj["Body"].read()
# List objects
response = s3.list_objects_v2(Bucket="my-bucket")
for item in response.get("Contents", []):
print(item["Key"], item["Size"])
```
To use the CA bundle instead of disabling verification:
```python
s3 = boto3.client(
"s3",
endpoint_url="https://s3.dev.austin.int.futo.cloud",
aws_access_key_id="YOUR_ACCESS_KEY",
aws_secret_access_key="YOUR_SECRET_KEY",
region_name="us-east-1",
verify="/path/to/rgw-ssl.crt",
config=botocore.config.Config(
signature_version="s3v4",
s3={"addressing_style": "path"},
),
)
```
## Restic
```bash
export AWS_ACCESS_KEY_ID="YOUR_ACCESS_KEY"
export AWS_SECRET_ACCESS_KEY="YOUR_SECRET_KEY"
export RESTIC_REPOSITORY="s3:https://s3.dev.austin.int.futo.cloud/restic-backups"
# Init (first time)
restic init --option s3.region=us-east-1
# Backup
restic backup /data --option s3.region=us-east-1
```
Note: Restic uses the Go AWS SDK. For self-signed certs, set
`AWS_CA_BUNDLE=/path/to/rgw-ssl.crt` or use the system trust store.
## Bucket creation
Buckets are created via any S3 client. The `svc-yucca-restic` service account
has a limit of 100 buckets (configurable via `radosgw-admin user modify
--max-buckets`).
```bash
# AWS CLI
aws --profile sietch --no-verify-ssl s3 mb s3://my-new-bucket
# boto3
s3.create_bucket(Bucket="my-new-bucket")
```
Bucket data lands in the EC data pool (`dev-z1.rgw.buckets.data`). Index
metadata goes to a separate replicated pool. No pool-level configuration is
needed from the application side.
## DNS setup for virtual-hosted buckets
Virtual-hosted bucket addressing (e.g., `https://my-bucket.s3.dev.austin.int.futo.cloud`)
requires two DNS records:
```
s3.dev.austin.int.futo.cloud. A 10.10.10.90
s3.dev.austin.int.futo.cloud. A 10.10.10.91
s3.dev.austin.int.futo.cloud. A 10.10.10.92
*.s3.dev.austin.int.futo.cloud. A 10.10.10.90
*.s3.dev.austin.int.futo.cloud. A 10.10.10.91
*.s3.dev.austin.int.futo.cloud. A 10.10.10.92
```
Round-robin A records across all three nodes. Without these records,
use path-style addressing or connect directly via IP.
Note: this repo does not manage DNS. Coordinate with the network/DNS team.
## Performance characteristics
| Property | Value |
|---|---|
| Data pool | Erasure coded, k=8 m=3, failure domain=OSD |
| Index pool | Replicated, size=2, min_size=1 |
| Storage media | HDD-backed (HGST 6 TB SAS drives) |
| Block.db | SSD-backed (Micron 5100 3.8 TB) |
| Nodes | 3 (Dell R730xd) |
| RGW daemons | 1 per node |
| TLS | Self-signed wildcard, 10-year validity |
| Network | 10 GbE bonded active-backup (no LACP) |
### What to expect
- **Throughput**: HDD-bound for large objects. A single node can sustain
roughly 500-800 MiB/s aggregate reads from its HDDs. With 3 nodes and
EC 8+3, expect 300-600 MiB/s aggregate for large sequential workloads
depending on concurrency and object size.
- **Latency**: Higher than SSD or cloud S3. Small object PUTs (< 1 MiB) will
see 10-50 ms latency due to HDD seeks. Use concurrency to amortize.
- **IOPS**: Low single-drive IOPS (100-200 per HDD). Use larger objects
(16+ MiB) to maximize throughput.
- **EC overhead**: Usable capacity is raw * k/(k+m) = raw * 8/11 = ~72.7%
of raw HDD capacity.
- **Single network**: Public and cluster traffic share the same 10 GbE bond.
Recovery/rebalancing events will compete with client I/O.
### Benchmarking
The cluster includes a benchmark tool at `roles/s3_bench/files/s3bench.py`:
```bash
python3 s3bench.py \
--endpoint https://10.10.10.90:443 \
--access-key KEY \
--secret-key SECRET \
--bucket s3bench \
--num-objects 100 \
--object-size-mb 16 \
--concurrency 8 \
--ops put \
--log /tmp/bench.jsonl
```
Supports `put`, `get`, `delete`, and `mixed` (70/20/10 split) operations.
Results include throughput (MiB/s), IOPS, and p50/p95/p99 latencies.
+292
View File
@@ -0,0 +1,292 @@
# Wrapper Scripts
The scripts under `ansible/ceph/scripts/` sit between `mise` tasks and the
underlying CLIs (`ansible-playbook`, `op`, `ssh`). They exist to:
- Keep secrets **out of argv** — `op inject` writes to a file; the file is
passed as `--extra-vars @<path>`, never expanded inline.
- **Fail closed** — if 1Password is unreachable or the inventory is missing,
the wrappers exit before invoking ansible. The raw tools' error messages
degrade silently in these cases.
- Offer **better error messages** — "inventory not found at `<path>` — has
`tofu apply` been run?" beats "No inventory was parsed".
For the architectural role these scripts play, see
[architecture.md §8](architecture.md).
## Setting `CEPH_ENV`
Two of the three wrappers (`ansible-play.sh`, `preflight.sh`) require
`CEPH_ENV` to point at the target cluster's `inventory.ini`. **Always set
it inline, never via `export`:**
```bash
# Correct — inline prefix, applies to one mise/script invocation
CEPH_ENV=inventories/painbox-ceph.dev.hel.htz/inventory.ini mise run preflight
# WRONG — mise's [env] machinery silently strips shell-exported vars
# when launching tasks; CEPH_ENV reaches an empty environment and the
# wrapper exits with "CEPH_ENV must be set". Confusing because your shell
# clearly has it set.
export CEPH_ENV=inventories/painbox-ceph.dev.hel.htz/inventory.ini
mise run preflight # fails
```
The `ansible/ceph/.mise.toml` `[env]` block intentionally does NOT declare
a `CEPH_ENV` default — silent default-cluster behavior is more dangerous
than requiring an explicit choice. If you call the scripts directly
(bypassing `mise run`), shell `export` works normally because mise isn't
in the path.
## Quick reference
| Script | What it does | Called by |
|------------------------|------------------------------------------------------------------------------------|---------------------------------------------------|
| `ansible-play.sh` | Render `secrets.yml.tpl` via `op inject` to a tmpfile, then exec `ansible-playbook` | Every `mise run` task that runs a playbook |
| `install-ssh-keys.sh` | Pull per-cluster ansible-iac SSH keys from 1P into `~/.ssh/` | Operator (once per workstation / after rotation) |
| `preflight.sh` | Verify TF artifacts, 1P session, SSH reachability, Python on targets | `mise run preflight` |
---
## `ansible-play.sh`
Wrapper around `ansible-playbook` that resolves TF-rendered secrets via
`op inject` and passes them as an ephemeral extra-vars file.
### Synopsis
```
CEPH_ENV=inventories/<cluster>/inventory.ini \
scripts/ansible-play.sh <playbook.yml> [ansible-playbook args...]
```
### What it does
1. Validates `$CEPH_ENV` is set and the inventory file exists.
2. Derives the secrets template path: `$(dirname $CEPH_ENV)/secrets.yml.tpl`.
3. Verifies `op account get` succeeds (1P desktop session or
`OP_SERVICE_ACCOUNT_TOKEN`). Fails fast if unavailable.
4. `mktemp`s a tmpfile, `chmod 600`, registers a `trap` to delete it on
`EXIT INT TERM` (including operator Ctrl-C or SIGKILL'd parents).
5. Runs `op inject -f -i <template> -o <tmpfile>`. Fails if any `op://`
reference can't be resolved.
6. `exec`s `ansible-playbook -i $CEPH_ENV --extra-vars @<tmpfile> <args>`.
The `exec` means the wrapper process is replaced — the trap still fires via
the shell's EXIT handler on the child's termination.
### Environment
| Variable | Required | Purpose |
|-----------------------------|----------|-----------------------------------------------------------------|
| `CEPH_ENV` | yes | Path to the target cluster's `inventory.ini` |
| `OP_SERVICE_ACCOUNT_TOKEN` | no | CI headless auth. Falls back to `op` desktop session if unset. |
### Arguments
Everything after the playbook name is passed to `ansible-playbook` verbatim.
Common patterns:
```bash
scripts/ansible-play.sh baseline.yml --check --diff
scripts/ansible-play.sh deploy-ceph.yml --tags rgw,monitoring
scripts/ansible-play.sh destroy-ceph.yml -e yes_destroy_ceph=true -e destroy_target_domain=dev.austin.int.futo.cloud
```
The destroy playbook requires both safety gates:
- `yes_destroy_ceph=true` — explicit confirmation (otherwise the play refuses to run)
- `destroy_target_domain=<cluster domain>` — must match the inventory's `cluster_domain`. Mismatch aborts the play, guarding against running destroy with the wrong `CEPH_ENV`.
The `mise run destroy` task (in `.mise.toml`) builds these arguments automatically from `CEPH_ENV` and adds an interactive `[y/N]` prompt — prefer it over invoking the wrapper directly.
### Exit codes
| Code | Meaning |
|------|-------------------------------------------------------------------|
| 0 | ansible-playbook completed successfully |
| 1 | Inventory file missing or secrets template missing |
| 2 | 1Password session unavailable (`op account get` failed) |
| 3 | `op inject` failed (bad `op://` reference, missing item, etc.) |
| ≥4 | ansible-playbook's own exit code (unreachable hosts, failed tasks) |
### Examples
```bash
# Standard deploy
CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory.ini \
scripts/ansible-play.sh deploy-ceph.yml
# Dry-run a role via tags
CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory.ini \
scripts/ansible-play.sh site.yml --check --diff --tags baseline
# CI / headless (SA token from env)
OP_SERVICE_ACCOUNT_TOKEN="$(...)" \
CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory.ini \
scripts/ansible-play.sh status.yml
```
### Related
- [architecture.md §9.2](architecture.md) — deploy data flow
- [secrets.md](secrets.md) — what's in `secrets.yml.tpl`
---
## `install-ssh-keys.sh`
Idempotent installer that pulls per-cluster `ansible-iac` SSH keypairs from
1Password into `~/.ssh/`. For new workstations or after a key rotation.
### Synopsis
```
scripts/install-ssh-keys.sh [cluster...]
```
If no cluster arguments are given, installs keys for every known cluster
(`sietch`, `painbox`).
### What it does
For each target cluster:
1. Resolves the 1P item name: `<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY` in
`yucca_tf_dev`.
2. Resolves the target filename: `~/.ssh/id_ed25519_<cluster>`.
3. If the key already exists on disk:
- Compares fingerprints (local vs 1P).
- **Match** → skip (idempotent re-run).
- **Mismatch** → refuse to overwrite. Prints the `mv` command the
operator should run manually. Exit 2.
4. If the key is missing:
- `umask 077`.
- Writes `private_key` → `~/.ssh/id_ed25519_<cluster>` (`chmod 600`).
- Writes `public_key` → `~/.ssh/id_ed25519_<cluster>.pub` (`chmod 644`).
- Prints the installed key's fingerprint.
### Environment
| Variable | Required | Purpose |
|-----------------------------|----------|---------------------------------------------------------|
| `OP_SERVICE_ACCOUNT_TOKEN` | no | CI headless auth. Read-only scope on `yucca_tf_dev` is enough. |
### Arguments
Zero or more cluster short names (`sietch`, `painbox`). With no args, all
known clusters are installed.
### Exit codes
| Code | Meaning |
|------|------------------------------------------------------------------|
| 0 | All requested keys installed (or already present and matching) |
| 1 | Unknown cluster name |
| 2 | Fingerprint mismatch — refused to overwrite existing key on disk |
### Examples
```bash
# Install all known keys
scripts/install-ssh-keys.sh
# Just one cluster
scripts/install-ssh-keys.sh sietch
# Re-run after a rotation (operator already moved the old key aside manually)
mv ~/.ssh/id_ed25519_sietch ~/.ssh/id_ed25519_sietch.20260423.bak
mv ~/.ssh/id_ed25519_sietch.pub ~/.ssh/id_ed25519_sietch.pub.20260423.bak
scripts/install-ssh-keys.sh sietch
```
### Related
- [runbooks/rotate-ssh-key.md](runbooks/rotate-ssh-key.md) — the full rotation flow
- [ADR-010](adr/010-ssh-keys-in-1password.md) — why SSH keys live in 1P
---
## `preflight.sh`
Read-only smoke test — verifies the controller environment is ready to run
destructive playbooks against the target cluster. Surfaced via
`mise run preflight`.
### Synopsis
```
CEPH_ENV=inventories/<cluster>/inventory.ini scripts/preflight.sh
```
### What it checks
**Controller:**
- `scripts/ansible-play.sh` is executable
- Inventory file exists (TF-rendered)
- Secrets template exists (TF-rendered)
- SSH private + public keys exist on disk
- `ansible` and `op` CLIs installed
- 1Password session live (`op account get`)
**Secrets:**
- `op inject` resolves the cluster's `secrets.yml.tpl` successfully
- Resolved file has a non-empty `vault_ops_password` line (sanity check
that op isn't silently substituting empty strings)
**Target connectivity** (one check per host in `ceph_nodes`):
- SSH reachability via `ansible -m ping`
- Python 3 available via `ansible -m raw`
### Environment
| Variable | Required | Purpose |
|-----------------------------|----------|-------------------------------------------------------|
| `CEPH_ENV` | yes | Path to the target cluster's `inventory.ini` |
| `OP_SERVICE_ACCOUNT_TOKEN` | no | CI headless auth. Falls back to `op` desktop session. |
### Exit codes
| Code | Meaning |
|------|------------------------------------------------------------------------------|
| 0 | All checks passed |
| 1 | One or more checks failed — summary printed, unsafe to proceed |
Warnings (non-blocking) are reported in the summary but don't affect exit.
### Examples
```bash
# Via mise (recommended)
mise run preflight
# Direct, against painbox
CEPH_ENV=inventories/painbox-ceph.dev.hel.htz/inventory.ini \
scripts/preflight.sh
```
### Related
- [architecture.md §8](architecture.md) — where the wrappers fit in the system mesh
---
## Adding a new wrapper
Follow these conventions:
1. **Fail closed** — `set -euo pipefail` at the top. Any uncaught error
aborts the script.
2. **Validate inputs early** — check required env vars (`: "${CEPH_ENV:?...}"`),
then check that referenced files exist, before doing any real work.
3. **Never interpolate secrets into argv** — write them to a `mktemp`'d
file (`chmod 600`) and pass the path. Always `trap 'rm -f "$TMP"' EXIT INT TERM`.
4. **Distinct exit codes** — the caller (mise task or another script) should
be able to tell "inventory missing" from "1P unreachable" from "ansible
failed" without parsing stderr.
5. **Idempotent where plausible** — re-running the script should not make
things worse. Prefer skip-if-already-correct over unconditional overwrite.
6. **Cross-reference from [architecture.md §8](architecture.md)** and add a
section to this file with the same shape as the ones above.
+143
View File
@@ -0,0 +1,143 @@
# Secrets
This project uses a **TF-provisions, op-injects, ansible-consumes** model. No
`ansible-vault`, no encrypted `vault.yml` in git, no custom password-file script.
For how secrets fit into the broader architecture, see
[architecture.md §5 (1Password)](architecture.md).
```mermaid
flowchart TB
ONEP[("1Password<br/>yucca_tf · yucca_tf_dev · ...<br/><i>source of truth</i>")]
TF[Terraform / Tofu<br/>tf/deployment/dev/ceph/]
REPO[/"inventories/&lt;cluster&gt;/<br/>inventory.ini (TF-gen, gitignored)<br/>secrets.yml.tpl (TF-gen, gitignored)"/]
WRAP[scripts/ansible-play.sh<br/><i>mktemp + op inject → exec ansible-playbook --extra-vars @tmp</i>]
ANS[ansible-playbook]
TF -.->|reads via op run --env-file| ONEP
TF -->|renders| REPO
WRAP -->|reads template| REPO
WRAP -.->|op inject -f<br/>at play time| ONEP
WRAP --> ANS
```
## What lives where
| Vault | Purpose | Who writes |
|------------------------------------|-------------------------------------------------------------------------------------------|-----------------------------------------------------------------|
| `yucca_tf` (team-shared) | Cross-env shared state: TF state S3 credentials | Operator (manual) |
| `yucca_tf_dev` (team-shared) | Live values for the `dev` environment — `<CLUSTER>_CEPH_*` items | Superuser service account (TF) + operator via `op` CLI |
| `yucca_tf_dev_manual` (team-shared)| Human-fillable dev placeholders (3rd-party API tokens, OAuth client secrets) — not yet used by ceph-cluster | Operator (manual) |
Future environments land as siblings: `yucca_tf_staging(_manual)`,
`yucca_tf_prod_manual`. The vault a given cluster reads from is declared
per-cluster in `tf/deployment/<env>/ceph/clusters.auto.tfvars` (field
`vault`). TF derives item paths from that field at render time; changing
it + `tofu apply` re-renders `secrets.yml.tpl` with the new vault path.
## Item naming
Format: `<CLUSTER>_CEPH_<ROLE>_PASSWORD`
**Password items** (category `Password`, consumed via `op inject` at
playbook time):
| Item | Field | Consumed as |
|---|---|---|
| `SIETCH_CEPH_OPS_PASSWORD` | `password` | `vault_ops_password` |
| `SIETCH_CEPH_DASHBOARD_PASSWORD` | `password` | `vault_ceph_dashboard_password` |
| `SIETCH_CEPH_GRAFANA_PASSWORD` | `password` | `vault_grafana_admin_password` |
| `PAINBOX_CEPH_OPS_PASSWORD` | `password` | `vault_ops_password` |
| `PAINBOX_CEPH_DASHBOARD_PASSWORD` | `password` | `vault_ceph_dashboard_password` |
| `PAINBOX_CEPH_GRAFANA_PASSWORD` | `password` | `vault_grafana_admin_password` |
**SSH Key items** (category `SSH Key`, consumed via
`scripts/install-ssh-keys.sh` on operator workstations and
`rotate-ssh-key.yml` on cluster nodes — see [ADR-010](adr/010-ssh-keys-in-1password.md)):
| Item | Field | Consumed as |
|---|---|---|
| `SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY` | `private_key` | `~/.ssh/id_ed25519_sietch` on operator workstation |
| `SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY` | `public_key` | sietch nodes' `ansible-iac@:~/.ssh/authorized_keys` |
| `PAINBOX_CEPH_ANSIBLE_IAC_SSH_KEY` | `private_key` | `~/.ssh/id_ed25519_painbox` on operator workstation |
| `PAINBOX_CEPH_ANSIBLE_IAC_SSH_KEY` | `public_key` | painbox nodes' `ansible-iac@:~/.ssh/authorized_keys` |
**S3 service-user items** (predetermined keys passed to
`radosgw-admin user create --access-key=X --secret-key=Y` at deploy
time — Yucca app / restic client can be pre-configured with matching
credentials without waiting for post-bootstrap capture):
| Item | Field | Consumed as |
|---|---|---|
| `SIETCH_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY` | `password` | `vault_s3_restic_access_key` → `ceph_rgw_s3_user_access_key` |
| `SIETCH_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY` | `password` | `vault_s3_restic_secret_key` → `ceph_rgw_s3_user_secret_key` |
| `PAINBOX_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY` | `password` | same, painbox |
| `PAINBOX_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY` | `password` | same, painbox |
**Disaster-recovery items** (populated by `mise run capture` after
deploy — stored in 1P for recovery if the bootstrap node's filesystem
is lost):
| Item | Field | Source |
|---|---|---|
| `<CLUSTER>_CEPH_RGW_TLS_CERT` | `password` (concealed) | `/etc/ceph/rgw-ssl.crt` on bootstrap |
| `<CLUSTER>_CEPH_RGW_TLS_KEY` | `password` (concealed) | `/etc/ceph/rgw-ssl.key` on bootstrap |
| `<CLUSTER>_CEPH_CLIENT_ADMIN_KEYRING` | `password` (concealed) | `/etc/ceph/ceph.client.admin.keyring` on bootstrap |
Items get created on first `mise run capture`; subsequent runs update
in place on content drift.
Item names are derived in `tf/shared/modules/ceph-cluster/main.tf`
(`local.secret_prefix`). Hardcoded `CEPH` (not `role_in_hostname`) so every
Ceph-project secret grep-matches `*_CEPH_*` regardless of whether hostnames
use `ceph`, `osd`, or `mon` as the role segment.
## Runtime flow
The op CLI is invoked in three distinct patterns across this project:
| Pattern | Used by | What it does |
|----------------------------|---------------------------------------------------------------------------|--------------------------------------------------------|
| `op run --env-file=tf/.env --` | Every `mise run tf:*` task | Resolves `op://` references in a dotenv file, injects as env vars into child process (TF: SA token + AWS creds) |
| `op inject -f -i tpl -o out` | `scripts/ansible-play.sh`, Hetzner installimage post-install rendering | Resolves all `op://` references in a file template, writes resolved file |
| `op read "op://..."` | `scripts/install-ssh-keys.sh`, `rotate-ssh-key.yml`, `post-deploy-capture.yml` | Reads a single field from a single item to stdout |
### The `ansible-play.sh` flow
`scripts/ansible-play.sh` wraps every `ansible-playbook` invocation:
1. Verifies `op account get` succeeds — fails closed if 1Password is locked.
2. `mktemp` a `0600` tmpfile with `trap` cleanup on `EXIT` / `INT` / `TERM`.
3. `op inject -f -i <cluster>/secrets.yml.tpl -o $tmpfile` — resolves every
`op://` reference. Exit non-zero if any reference can't be resolved.
4. `exec ansible-playbook --extra-vars @$tmpfile ...`.
Ansible task references stay unchanged — `vault_ops_password` etc. are
regular variables populated from extra-vars (highest precedence).
Full wrapper reference: [scripts.md](scripts.md).
## CI / headless
Set `OP_SERVICE_ACCOUNT_TOKEN` (from a service-account item) in the CI
environment. `op inject` uses it automatically. The cutover PR must
demonstrate green `lint` and `check` with **zero** op credentials — those
tasks do not require secrets.
## Rotating secrets
See [docs/runbooks/rotate-secrets.md](runbooks/rotate-secrets.md).
## Adding a new secret
1. Add the item name to `local.secrets` in
`tf/shared/modules/ceph-cluster/main.tf`.
2. Add the matching line to the secrets template
(`templates/secrets.yml.tpl.tftpl`).
3. Add the ansible variable alias in each cluster's
`inventories/<cluster>/group_vars/all/vars.yml`.
4. Create the item manually in the target vault (or let TF do it post
service account): `op item create --vault <vault> --category password
--title NEW_SECRET_NAME --generate-password=letters,digits,32`.
5. `tofu apply` — re-renders templates with the new reference.
6. Playbooks consuming the variable now have it available.
+236
View File
@@ -0,0 +1,236 @@
# Security Model
Audience: InfoSec, compliance review, security audits.
For how these controls fit the broader tool mesh, see
[architecture.md](architecture.md). For the per-item secrets catalog see
[secrets.md](secrets.md).
## Encryption at rest
All OSD volumes (HDD and SSD) are created with `--dmcrypt`, which wraps each
OSD's BlueStore volume in a dm-crypt/LUKS layer.
| Property | Detail |
|---|---|
| Encryption layer | dm-crypt (LUKS) via cephadm OSD service spec (`encrypted: true`); see [ADR-011](adr/011-cephadm-osd-service-specs.md) |
| Scope | Every OSD data volume and every OSD block.db volume |
| Key storage | MON config-key database (`config-key dump` shows dm-crypt entries) |
| Key distribution | MONs hand keys to OSD daemons at startup via the Ceph auth subsystem |
| Algorithm | AES-256-XTS (dm-crypt default) |
The encryption keys never leave the MON quorum. If a drive is removed from
the chassis, the data is unreadable without access to the MON key store.
### Verification
```bash
# Count dm-crypt keys in MON store
ceph config-key dump 2>/dev/null | grep -c dm-crypt
# Confirm OSDs are on dm-crypt devices
lsblk --output NAME,TYPE,MOUNTPOINT | grep crypt
```
## Encryption in transit
### RGW / S3 (client-facing)
| Property | Detail |
|---|---|
| Protocol | HTTPS (TLS 1.2+) on port 443 |
| Certificate | Self-signed RSA 4096-bit, 10-year validity |
| CN | `s3.dev.austin.int.futo.cloud` |
| SANs | `s3.dev.austin.int.futo.cloud`, `*.s3.dev.austin.int.futo.cloud`, per-node FQDNs, per-node bond IPs |
| Issuer | Self-signed (O=FUTO, L=Austin, ST=Texas, C=US) |
| Cert location | `/etc/ceph/rgw-ssl.crt` + `/etc/ceph/rgw-ssl.key` on bootstrap node |
| Distribution | cephadm distributes combined PEM to all RGW daemon containers |
Clients must either trust the self-signed cert via `--ca-bundle` / `verify=`
or disable TLS verification (`--no-verify-ssl`).
### Intra-cluster (MON/OSD/MGR)
Ceph messenger v2 (`msgr2`) is used for all intra-cluster communication.
The `cephx` authentication protocol provides mutual authentication between
daemons. Wire encryption (`ms_client_mode`, `ms_cluster_mode`) is available
but not explicitly forced in this deployment -- the default Tentacle
configuration uses `crc` mode (authenticated but not encrypted on the wire).
### Dashboard
Ceph dashboard runs on port 8443 (HTTPS) with its own self-signed cert,
restricted to trusted networks only via firewall rules.
## User model
Three user accounts exist on every node:
| User | UID | Authentication | Sudo | Purpose |
|---|---|---|---|---|
| `ansible-iac` | 1000 | SSH key only (password locked) | `NOPASSWD:ALL` | Ansible automation. No interactive use. |
| `ops` | 1001 | Password (from 1Password via `op inject`) | `ALL` (password required) | Human interactive access. Operator SSH keys distributed out-of-band post-deploy. |
| `root` | 0 | SSH key (cephadm inter-node) | n/a | Required by cephadm orchestrator for SSH between nodes. `PermitRootLogin prohibit-password`. |
| `ceph` | system | nologin shell | none | Ceph daemon processes. No login capability. |
### SSH hardening
| Setting | Value |
|---|---|
| `PermitRootLogin` | `prohibit-password` (key-only, required by cephadm) |
| `MaxAuthTries` | 3 |
| `AllowUsers` | `ansible-iac ops root` |
| `PasswordAuthentication` | Not explicitly disabled (ops user uses password + key) |
### Key management
- **ansible-iac key**: The keypair is stored in 1Password as an SSH Key item
(`<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY` in `yucca_tf_dev`). Operator
workstations install it via `scripts/install-ssh-keys.sh` →
`~/.ssh/id_ed25519_<cluster>`. Rotation is forward-only (additive) via
`rotate-ssh-key.yml` — new pubkey distributed to `authorized_keys` with
`exclusive: false`, old keys pruned out-of-band. See
[ADR-010](adr/010-ssh-keys-in-1password.md).
- **ops user keys**: Distributed out-of-band after initial deployment.
No keys are provisioned at deploy time.
- **cephadm key**: Generated during `cephadm bootstrap`, distributed to all
nodes during the join phase. Used for orchestrator SSH between nodes.
## Firewall rules (nftables)
Every node runs nftables with a default-drop input policy. The ruleset is
templated from `roles/security/templates/nftables.conf.j2`.
### Open to all sources
| Port | Service |
|---|---|
| 22/tcp | SSH (configurable: `ceph_firewall_ssh_any_source` defaults to true for dev) |
| 443/tcp | RGW / S3 endpoint |
| ICMP | Ping and MTU discovery |
### Restricted to trusted networks only
Trusted networks: `10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16`,
`100.64.0.0/10` (Tailscale CGNAT).
| Port(s) | Service |
|---|---|
| 3300, 6789/tcp | Ceph MON |
| 6800-7300/tcp | Ceph OSD |
| 8443/tcp | Ceph Dashboard |
| 9095/tcp | Prometheus |
| 3000/tcp | Grafana |
| 9093/tcp | Alertmanager |
| 9100/tcp | Node Exporter |
| 9283/tcp | MGR Exporter |
| 9926/tcp | Ceph Exporter |
### Default policy
```
chain input { policy drop; } # all unmatched traffic is dropped
chain forward { policy drop; } # no forwarding
chain output { policy accept; } # outbound unrestricted
```
Dropped packets are logged at rate 5/minute with prefix `nftables-drop: ` for
forensic review.
### Optional gateways (disabled by default)
- iSCSI (port 3260 + API 5000) -- `ceph_firewall_iscsi_enabled: false`
- NFS-Ganesha (port 2049 + mgmt 12049) -- `ceph_firewall_nfs_enabled: false`
## Secrets management
### 1Password + op inject
All deployment secrets live in 1Password in the `yucca_tf_dev` vault (dev
environment; staging/prod land as sibling `yucca_tf_staging` /
`yucca_tf_prod` vaults). Items are named `<CLUSTER>_CEPH_<ROLE>_*` — see
[docs/secrets.md](secrets.md) for the full catalog.
At playbook time, `scripts/ansible-play.sh` invokes `op inject` on the
TF-rendered `secrets.yml.tpl`, writes the resolved values to a `0600`
tmpfile, and passes it as `--extra-vars @tmpfile`. Tmpfile is `trap`-cleaned
on `EXIT`/`INT`/`TERM`. No at-rest encrypted file in git.
| Secret | Variable |
|----------------------------------|-----------------------------------|
| ops user password | `vault_ops_password` |
| Ceph dashboard admin password | `vault_ceph_dashboard_password` |
| Grafana admin password | `vault_grafana_admin_password` |
| S3 service-user access key | `vault_s3_restic_access_key` |
| S3 service-user secret key | `vault_s3_restic_secret_key` |
### Trust boundaries
Two service accounts separate write authority from runtime consumption:
| Service account | Scope | Used by |
|-----------------------------------------------|--------------------------------------------|----------------------------------------|
| `yucca_futo_1pass_superuser_service_account` | Read + write all `yucca_tf_*` vaults | TF (`tf/.env`) + interactive `op` CLI |
| `yucca_futo_1pass_service_account` | Read-only on `yucca_tf` and `yucca_tf_dev` | Ansible runtime / future CI |
The ansible-play.sh wrapper runs with whichever session is active on the
operator workstation (typically desktop unlock → superuser SA via 1Password
desktop). CI will use the read-only SA via `OP_SERVICE_ACCOUNT_TOKEN`.
Rotation procedure: [runbooks/rotate-sa-token.md](runbooks/rotate-sa-token.md).
### What is NOT in 1Password
- dm-crypt OSD encryption keys — stored in the MON config-key database.
Deferred to 1P as a future belt-and-suspenders item (see yucca memory
`project_luks_keys_in_1pass.md`).
- cephadm bootstrap SSH key — generated at bootstrap, distributed by
cephadm's orchestrator. Not needed outside the cluster.
- ops user SSH keys — distributed out-of-band after initial deployment.
See "ops user keys" in Key management above.
## Audit logging
| Control | Detail |
|---|---|
| Ceph audit log | Enabled (`ceph config set global log_to_cluster_level audit`) |
| Ceph manager log | `mgr/cephadm/log_to_cluster true` |
| View audit log | `ceph log last -W audit` |
| Firewall logging | Dropped packets logged at 5/min with `nftables-drop:` prefix |
| Prometheus alerts | 16+ alert rule groups (89 built-in rules) covering OSD, MON, PG, pool, MDS, hardware, network |
The audit channel records all `ceph` admin commands executed against the
cluster, including the authenticated user, timestamp, and command arguments.
## Telemetry
Telemetry phone-home is explicitly disabled:
```
ceph_telemetry_enabled: false
```
No cluster metadata, performance data, or crash reports are sent to upstream
Ceph. All data stays within the cluster boundary.
## Network topology
| Network | CIDR | Purpose |
|---|---|---|
| Public/Cluster | 10.10.10.0/24 | Combined public + cluster traffic (single-network topology) |
| iDRAC/BMC | 10.10.11.0/24 | Out-of-band management (separate VLAN) |
Nodes are on a private network -- the 10.10.10.0/24 subnet is not directly
routable from the public internet.
## Known gaps and mitigations
| Gap | Risk | Mitigation |
|---|---|---|
| Self-signed TLS cert | Clients must disable verification or trust the CA manually. MITM possible if cert is not pinned. | Cert has 10-year validity with specific SANs. Production (Yucca) will use cert-manager + real CA behind haproxy. |
| Single network (public = cluster) | Cluster rebalancing traffic is visible on the client network. A compromised client could sniff inter-OSD traffic. | Trusted network firewall restricts cluster ports to RFC1918 + Tailscale. Production will separate public and cluster networks. |
| `msgr2` wire encryption not forced | Intra-cluster traffic is authenticated (cephx) but not encrypted on the wire by default. | All traffic stays within 10.10.10.0/24 on a private switch. Can be enabled via `ceph config set global ms_cluster_mode secure` if needed. |
| SSH open to all sources (dev default) | SSH is reachable from any IP that can route to the nodes. | On a private network (not internet-exposed). `MaxAuthTries=3`, `AllowUsers` whitelist. Production should set `ceph_firewall_ssh_any_source: false`. |
| `ops` password auth | Password-based SSH is not disabled. | Password is vault-encrypted, rotated via Ansible. Interactive use only -- automation uses key-only `ansible-iac`. |
| Operator SSH keys distributed out-of-band | No automated key lifecycle for `ops` user. | Acceptable for dev. Production should use centralized key management (e.g., Teleport, Vault SSH CA). |
| RGW S3 credentials static | No automatic rotation of S3 access/secret keys. | Keys stored in 1Password with access control. `radosgw-admin key create/rm` available for manual rotation. |
+684
View File
@@ -0,0 +1,684 @@
# Troubleshooting Guide
Decision-tree format: symptom, diagnosis, fix.
---
## Cluster Health
### HEALTH_WARN
#### Symptom: `HEALTH_WARN: N osds down`
**Diagnose:**
```bash
ceph osd tree | grep down
ceph health detail
```
**Common causes and fixes:**
1. **OSD daemon crashed** -- check logs:
```bash
ceph crash ls-new
ceph crash info <crash-id>
# Or on the host:
journalctl -u ceph-osd@<id> --since "1 hour ago"
```
Restart the daemon:
```bash
ceph orch daemon restart osd.<id>
```
2. **Host unreachable** -- the node itself is down:
```bash
ssh ansible-iac@sietch-ceph-<name> hostname
# If unreachable, check iDRAC / physical console
```
3. **Disk failure** -- see the [replace-disk runbook](runbooks/replace-disk.md).
4. **OSD out of disk space** -- check nearfull/full ratios:
```bash
ceph osd df
```
---
#### Symptom: `HEALTH_WARN: N pgs not active+clean`
**Diagnose:**
```bash
ceph pg stat
ceph pg dump_stuck
```
**Common causes and fixes:**
1. **Backfill in progress** -- normal after adding/removing OSDs. Monitor:
```bash
ceph -s
```
Wait for completion. No action needed.
2. **Stale PGs** -- PGs stuck in `stale` state:
```bash
ceph pg dump_stuck stale
```
Usually indicates the hosting OSD is down. Fix the OSD first.
3. **Inactive PGs** -- PGs in `creating` or `peering`:
```bash
ceph pg dump_stuck inactive
```
If stuck for >15 minutes, check MON logs. May need `ceph pg force-create-pg <pgid>`.
---
#### Symptom: `HEALTH_WARN: clock skew detected`
**Diagnose:**
```bash
ceph time-sync-status
# On each node:
chronyc tracking
```
**Fix:** Restart chrony on the affected node:
```bash
sudo systemctl restart chrony
```
If persistent, check NTP sources:
```bash
chronyc sources -v
```
---
#### Symptom: `HEALTH_WARN: N daemons have recently crashed`
**Diagnose:**
```bash
ceph crash ls-new
ceph crash info <crash-id>
```
**Fix:** Review the crash, then archive it:
```bash
ceph crash archive <crash-id>
# Or archive all:
ceph crash archive-all
```
If crashes are recurring, investigate the daemon logs on the host.
---
#### Symptom: `HEALTH_WARN: N pool(s) have no replicas configured`
**Diagnose:**
```bash
ceph osd pool ls detail | grep 'size 1'
```
**Fix:** Set appropriate replication:
```bash
ceph osd pool set <pool> size 2
ceph osd pool set <pool> min_size 1
```
---
### HEALTH_ERR
#### Symptom: `HEALTH_ERR: N pgs are stuck inactive`
**Diagnose:**
```bash
ceph pg dump_stuck inactive
ceph osd tree
```
**Fix:** This is critical -- data may be inaccessible.
1. Check if the hosting OSDs are down. Bring them up first.
2. If OSDs are permanently lost and data cannot be recovered:
```bash
# DANGER: marks missing PGs as complete with potential data loss
ceph pg force-recovery <pgid>
# Last resort:
ceph osd force-create-pg <pgid> --yes-i-really-mean-it
```
---
#### Symptom: `HEALTH_ERR: N scrub errors`
**Diagnose:**
```bash
ceph health detail
# Find the affected PGs
ceph pg dump | grep inconsistent
# Deep scrub the PG
ceph pg deep-scrub <pgid>
```
**Fix:**
```bash
ceph pg repair <pgid>
```
If repair fails, the underlying disk may have bit rot. Check SMART data on
the hosting OSDs.
---
#### Symptom: `HEALTH_ERR: full osds`
**Diagnose:**
```bash
ceph osd df
ceph df
```
**Fix:** This is an emergency. The cluster stops accepting writes.
1. Delete unnecessary data or pools if possible
2. Temporarily raise the full ratio:
```bash
ceph osd set-full-ratio 0.97
```
3. Add more OSDs (see [add-node runbook](runbooks/add-node.md))
4. Set the ratio back after capacity is restored:
```bash
ceph osd set-full-ratio 0.95
```
---
## OSD Issues
### Symptom: OSD down and won't start
**Diagnose:**
```bash
# Check daemon status
ceph orch ps --daemon-type osd | grep <host>
# Check container logs
ssh ansible-iac@<host>
sudo podman logs ceph-<fsid>-osd.<id>
sudo journalctl -u ceph-<fsid>@osd.<id>
```
**Common causes:**
1. **LUKS key missing** -- dmcrypt key not in MON store:
```bash
ceph config-key dump | grep dm-crypt | grep <osd-id>
```
If missing, the OSD cannot be unlocked. Rebuild it (see replace-disk
runbook).
2. **Corrupt BlueStore DB** -- look for `fsck` errors in the OSD log.
May need `ceph-bluestore-tool repair`.
3. **Block device disappeared** -- check the SAS path:
```bash
ls /dev/disk/by-path/ | grep phy<N>
```
If missing, the disk or cable has failed.
---
### Symptom: Slow ops / blocked requests
**Diagnose:**
```bash
ceph daemon osd.<id> dump_ops_in_flight
ceph daemon osd.<id> perf dump | grep -i slow
```
**Common causes:**
1. **Disk latency** -- check I/O wait:
```bash
iostat -xz 5 3
```
Look for `%util > 90%` or `await > 100ms` on HDD devices.
2. **Network issues** -- check for packet loss:
```bash
ping -c 100 <other-node-ip>
ethtool -S eno1 | grep -i error
```
3. **Recovery throttling too aggressive** -- reduce recovery impact:
```bash
ceph config set osd osd_recovery_max_active 1
ceph config set osd osd_max_backfills 1
ceph config set osd osd_recovery_sleep_hdd 0.1
```
---
## RGW (S3) Issues
### Symptom: S3 requests return 403 Forbidden / SignatureDoesNotMatch
**Diagnose:**
```bash
# Check if the Host header hostname is in the zonegroup
radosgw-admin zonegroup get --rgw-zonegroup=us-east-1 | python3 -c "
import sys, json
zg = json.load(sys.stdin)
print('Hostnames:', zg.get('hostnames', []))
print('API name:', zg.get('api_name'))
"
```
**Common causes:**
1. **Missing hostname in zonegroup** -- the Host header used by the client
is not in the zonegroup's hostnames list. S3 signature verification
includes the Host header, so mismatches cause 403.
Fix:
```bash
# Re-run the RGW role to add all node hostnames/IPs
scripts/ansible-play.sh deploy-ceph.yml --tags rgw \
--limit sietch-ceph-laurel
```
2. **Wrong access/secret key** -- verify credentials:
```bash
radosgw-admin user info --uid=svc-yucca-restic
```
3. **Clock skew** -- S3 signatures are time-sensitive. Check client and
server clocks are within 15 minutes.
---
### Symptom: S3 requests return 500 Internal Server Error
**Diagnose:**
```bash
# Check RGW daemon logs
ceph log last 50 --channel=cluster | grep rgw
# Check if RGW daemons are running
ceph orch ls --service-type rgw
ceph orch ps --daemon-type rgw
```
**Common causes:**
1. **RGW daemons down** -- restart:
```bash
ceph orch restart rgw
```
2. **Pool issues** -- check that RGW pools exist and are healthy:
```bash
ceph osd pool ls | grep rgw
ceph pg stat
```
3. **TLS cert expired or corrupted** -- see
[rotate-certs runbook](runbooks/rotate-certs.md).
---
### Symptom: Dashboard Object Gateway page shows 500 or is empty
**Diagnose:**
```bash
# Check dashboard RGW API SSL verification setting
ceph dashboard get-rgw-api-ssl-verify
# Check dashboard RGW user has caps
radosgw-admin user info --uid=dashboard | python3 -c "
import sys, json
u = json.load(sys.stdin)
print('Caps:', u.get('caps', []))
print('System:', u.get('system'))
"
```
**Fix:**
1. Disable SSL verification (self-signed certs):
```bash
ceph dashboard set-rgw-api-ssl-verify false
```
2. Add admin caps to dashboard user:
```bash
radosgw-admin caps add --uid=dashboard \
--caps='buckets=*;users=*;usage=*;metadata=*;zone=*'
```
3. Sync dashboard credentials:
```bash
ceph dashboard set-rgw-credentials
```
---
## Dashboard Issues
### Symptom: Dashboard unreachable at https://<ip>:8443
**Diagnose:**
```bash
# Check MGR daemon status
ceph orch ps --daemon-type mgr
# Check which MGR is active
ceph mgr stat
# Check dashboard module is enabled
ceph mgr module ls --format json | python3 -c "
import sys, json
d = json.load(sys.stdin)
print('dashboard' in d.get('enabled_modules', []))
"
```
**Fix:**
1. **MGR daemon down** -- restart:
```bash
ceph orch restart mgr
```
2. **Dashboard module disabled**:
```bash
ceph mgr module enable dashboard
```
3. **Firewall blocking port 8443** -- check nftables:
```bash
ssh ansible-iac@<host> sudo nft list ruleset | grep 8443
```
If the port is not open, re-run security hardening:
```bash
scripts/ansible-play.sh harden.yml
```
4. **Wrong IP/port** -- check dashboard URL:
```bash
ceph mgr services
```
---
## Deploy Issues
### Symptom: `mise run deploy` fails — `no available installation candidate for cephadm=20.2.*`
**Cause:** apt cache is stale. `prerequisites.yml` adds the Ceph Tentacle
repo (`download.ceph.com/debian-tentacle`) and then runs `apt update`.
On a freshly-installed OS (Hetzner installimage runs apt internally),
the cache is "fresh" enough that an `apt update` with `cache_valid_time`
set will skip the refresh — apt never sees the Ceph repo's Packages file
and only knows about Debian's older `cephadm 16.2.x`.
**Diagnose:** SSH to the target node and run:
```bash
cat /etc/apt/sources.list.d/ceph.list # confirm the repo file exists
apt-cache policy cephadm # if download.ceph.com is missing, cache is stale
apt-get update # manual refresh
apt-cache policy cephadm # should now show 20.2.x candidate
```
**Fix:** `tasks/prerequisites.yml` was patched to drop `cache_valid_time`
on the apt-update task; refresh is unconditional after the repo is
added. If you're seeing this on a node that ran the OLD prerequisites
task (cached deploy state), `apt update` manually then re-run deploy.
---
### Symptom: `mise run deploy` fails — `'ceph_rgw_dns_name' is undefined`
**Cause:** the per-cluster `group_vars/all/vars.yml` is missing the
`ceph_rgw_dns_name` declaration. RGW zonegroup creation needs it for
the `--endpoints` and master zonegroup hostname.
**Fix:** add to `inventories/<cluster>/group_vars/all/vars.yml`:
```yaml
ceph_rgw_dns_name: s3.{{ cluster_domain }}
```
This derives the DNS name from `cluster_domain` (e.g.
`s3.dev.hel.htz.futo.cloud`). Sietch defines this explicitly; painbox
was missing it before its first deploy. Future clusters should include
it from the start — see [docs/adding-a-cluster.md](adding-a-cluster.md)
group_vars template.
---
### Symptom: `HEALTH_WARN: OSDMAP_FLAGS: noin flag(s) set` after deploy
**Cause:** the `noin` flag was set by an earlier failed run of the
older imperative OSD-creation flow and never unset. The current
spec-based flow doesn't set `noin` (cephadm rolls out OSDs gracefully)
and includes a defensive unset task at the tail of `osds.yml`, but the
flag can persist if the deploy never reached that tail (e.g., a failure
in an earlier phase).
**Fix:** clear it manually, or just re-run `mise run deploy` — the
defensive task at the end of `tasks/osds.yml` unsets `noin`
unconditionally (idempotent no-op when already unset):
```bash
ssh -i ~/.ssh/id_ed25519_<cluster> root@<bootstrap-ip> 'ceph osd unset noin'
```
Validate:
```bash
ceph osd dump | grep -E "^flags"
# Should NOT contain 'noin'. Default healthy: sortbitwise,recovery_deletes,purged_snapdirs,pglog_hardlimit
```
---
## Provisioning Issues
### Symptom: provision.yml fails with "REFUSING TO RUN"
**Cause:** Missing safety flag.
**Fix:**
```bash
CEPH_ENV=<inventory> scripts/ansible-play.sh provision.yml \
-e confirm_wipe=true
```
---
### Symptom: Provisioning fails at debootstrap / chroot phase
**Diagnose:** Check which task failed in the Ansible output. The rescue
block automatically unmounts `/mnt`, so it's safe to re-run.
**Common causes:**
1. **apt sources unreachable from live image** -- check network
connectivity from the live image. DNS resolution and internet access
are required for debootstrap.
2. **Disk detection failed** -- SSD not found at expected path:
```bash
ls /dev/disk/by-path/ | grep sas
lsblk
```
3. **Previous partial provision** -- the role is idempotent. If the
provisioning marker exists at `/mnt/etc/ceph-provisioned.json`, all
chroot phases are skipped. To force re-provision, boot into the live
image and re-run.
---
### Symptom: Post-reboot SSH fails after provisioning
**Diagnose:**
```bash
# Try with verbose SSH
ssh -vvv -i ~/.ssh/id_ed25519_sietch ansible-iac@sietch-ceph-<name>
```
**Common causes:**
1. **Node still booting** -- R730xd POST takes 60-90 seconds. Wait and
retry.
2. **SSH host key changed** -- fresh provision generates new host keys:
```bash
ssh-keygen -R sietch-ceph-<name>
```
3. **Network not up** -- bond interface may not have configured. Check
via iDRAC virtual console.
4. **Wrong IP** -- verify `bond_ip` in host_vars matches the actual
network config.
---
## SSH Connectivity
### Symptom: Cannot SSH to cluster nodes from controller
**Diagnose:**
```bash
# Test SSH to a node directly
ssh ansible-iac@10.10.10.90 hostname
# If using a jump host, verify it's reachable (check your ~/.ssh/config)
ssh <jump-host> hostname
```
**Common causes:**
1. **SSH config issue** -- if nodes are behind a jump host, verify your
`~/.ssh/config` has the correct ProxyJump or ProxyCommand settings.
This is personal config, not managed by the repo.
2. **Wrong SSH key** -- inventory uses `~/.ssh/id_ed25519_sietch`:
```bash
ls -la ~/.ssh/id_ed25519_sietch*
```
3. **sntrup761 kex hang** -- cephadm's asyncssh does not support
post-quantum key exchange. The baseline role deploys
`/etc/ssh/sshd_config.d/no-sntrup.conf` to disable it. If missing:
```bash
scripts/ansible-play.sh deploy-ceph.yml \
--tags prerequisites --limit sietch-ceph-<name>,sietch-ceph-laurel
```
4. **nftables blocking SSH** -- verify port 22 is allowed:
```bash
# From the node (via iDRAC console if SSH is blocked)
nft list ruleset | grep 22
```
---
## Quick Health Check Commands
```bash
# Overall status
ceph status
# OSD health
ceph osd tree
ceph osd df
# PG health
ceph pg stat
ceph pg dump_stuck
# Services
ceph orch ls
ceph orch ps
# Recent crashes
ceph crash ls-new
# Drift from expected config
mise run drift
# Cluster capacity
ceph df
```
+342
View File
@@ -0,0 +1,342 @@
---
# Drift detection — compares expected Ansible state against live cluster.
# Read-only. No changes. Reports mismatches.
#
# Usage:
# scripts/ansible-play.sh drift.yml
# mise run drift
- name: Drift detection
hosts: ceph_nodes
become: true
gather_facts: false
vars:
drift_results: []
vars_files:
- roles/baseline/defaults/main.yml
- roles/os_tuning/defaults/main.yml
- roles/hardware_tuning/defaults/main.yml
- roles/ceph_tuning/defaults/main.yml
- roles/security/defaults/main.yml
tasks:
# --- Collect all checks ---
# sysctl
- name: Check sysctl values
ansible.builtin.shell: |
set -o pipefail
sysctl -n {{ item.key }} 2>/dev/null | tr -d ' '
args:
executable: /bin/bash
loop:
- key: vm.swappiness
expected: "{{ ceph_sysctl_vm_swappiness }}"
- key: vm.min_free_kbytes
expected: "{{ ceph_sysctl_vm_min_free_kbytes }}"
- key: vm.zone_reclaim_mode
expected: "{{ ceph_sysctl_vm_zone_reclaim_mode }}"
- key: fs.aio-max-nr
expected: "{{ ceph_sysctl_fs_aio_max_nr }}"
- key: kernel.pid_max
expected: "{{ ceph_sysctl_kernel_pid_max }}"
register: sysctl_checks
changed_when: false
loop_control:
label: "{{ item.key }}"
- name: Record sysctl drift
ansible.builtin.set_fact:
drift_results: >-
{{ drift_results + [{
'category': 'sysctl',
'item': item.item.key,
'expected': item.item.expected | string,
'actual': item.stdout | trim,
'match': (item.stdout | trim) == (item.item.expected | string)
}] }}
loop: "{{ sysctl_checks.results }}"
loop_control:
label: "{{ item.item.key }}"
# I/O scheduler
- name: Check HDD I/O scheduler
ansible.builtin.shell: |
set -o pipefail
for dev in $(lsblk -dnpo NAME,ROTA,TYPE \
| awk '$2==1 && $3=="disk" {print $1}' | head -1); do
BDEV=$(basename "$dev")
cat /sys/block/$BDEV/queue/scheduler \
| grep -oP '\[\K[^\]]+'
done
args:
executable: /bin/bash
register: hdd_sched
changed_when: false
- name: Record HDD scheduler drift
ansible.builtin.set_fact:
drift_results: >-
{{ drift_results + [{
'category': 'hardware',
'item': 'HDD I/O scheduler',
'expected': ceph_hdd_scheduler,
'actual': hdd_sched.stdout | trim | default('none'),
'match': (hdd_sched.stdout | trim | default('none'))
== ceph_hdd_scheduler
}] }}
- name: Check SSD I/O scheduler
ansible.builtin.shell: |
set -o pipefail
for dev in $(lsblk -dnpo NAME,ROTA,TYPE \
| awk '$2==0 && $3=="disk" {print $1}' | head -1); do
BDEV=$(basename "$dev")
cat /sys/block/$BDEV/queue/scheduler \
| grep -oP '\[\K[^\]]+'
done
args:
executable: /bin/bash
register: ssd_sched
changed_when: false
- name: Record SSD scheduler drift
ansible.builtin.set_fact:
drift_results: >-
{{ drift_results + [{
'category': 'hardware',
'item': 'SSD I/O scheduler',
'expected': ceph_ssd_scheduler,
'actual': ssd_sched.stdout | trim | default('none'),
'match': (ssd_sched.stdout | trim | default('none'))
== ceph_ssd_scheduler
}] }}
# nftables
- name: Check nftables policy
ansible.builtin.shell: |
set -o pipefail
nft list chain inet filter input 2>/dev/null \
| grep -c 'policy drop' || echo 0
args:
executable: /bin/bash
register: nft_policy
changed_when: false
- name: Record firewall drift
ansible.builtin.set_fact:
drift_results: >-
{{ drift_results + [{
'category': 'security',
'item': 'nftables policy drop',
'expected': 'enabled',
'actual': 'enabled' if (nft_policy.stdout | trim | int > 0)
else 'missing',
'match': (nft_policy.stdout | trim | int > 0)
}] }}
# SSH hardening
- name: Check PasswordAuthentication
ansible.builtin.shell: |
set -o pipefail
sshd -T 2>/dev/null | grep -i passwordauthentication \
| awk '{print $2}'
args:
executable: /bin/bash
register: ssh_pwauth
changed_when: false
- name: Record SSH drift
ansible.builtin.set_fact:
drift_results: >-
{{ drift_results + [{
'category': 'security',
'item': 'SSH PasswordAuthentication',
'expected': 'no',
'actual': ssh_pwauth.stdout | trim,
'match': (ssh_pwauth.stdout | trim) == 'no'
}] }}
# sudo config
- name: Check ops sudo config
ansible.builtin.shell: |
grep -c 'NOPASSWD' /etc/sudoers.d/ops 2>/dev/null || echo 0
args:
executable: /bin/bash
register: ops_sudo
changed_when: false
- name: Record sudo drift
ansible.builtin.set_fact:
drift_results: >-
{{ drift_results + [{
'category': 'security',
'item': 'ops sudo requires password',
'expected': 'yes',
'actual': 'no (NOPASSWD)' if (ops_sudo.stdout | trim | int > 0)
else 'yes',
'match': (ops_sudo.stdout | trim | int == 0)
}] }}
# --- Ceph cluster checks (bootstrap only) ---
- name: Check OSD status
ansible.builtin.shell: |
set -o pipefail
ceph osd stat --format json 2>/dev/null | python3 -c "
import sys, json
d = json.load(sys.stdin)
print(d.get('num_osds', 0), d.get('num_up_osds', 0))
"
args:
executable: /bin/bash
register: osd_stat
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Record OSD drift
ansible.builtin.set_fact:
drift_results: >-
{{ drift_results + [{
'category': 'ceph',
'item': 'OSDs up/total',
'expected': (osd_stat.stdout.split()[0]) ~ '/'
~ (osd_stat.stdout.split()[0]),
'actual': (osd_stat.stdout.split()[1]) ~ '/'
~ (osd_stat.stdout.split()[0]),
'match': osd_stat.stdout.split()[0]
== osd_stat.stdout.split()[1]
}] }}
when: inventory_hostname in groups['ceph_bootstrap']
- name: Check MON quorum
ansible.builtin.shell: |
set -o pipefail
ceph mon stat --format json 2>/dev/null | python3 -c "
import sys, json
print(json.load(sys.stdin).get('num_mons', 0))
"
args:
executable: /bin/bash
register: mon_count
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Record MON drift
ansible.builtin.set_fact:
drift_results: >-
{{ drift_results + [{
'category': 'ceph',
'item': 'MON count',
'expected': groups['ceph_nodes'] | length | string,
'actual': mon_count.stdout | trim,
'match': (mon_count.stdout | trim)
== (groups['ceph_nodes'] | length | string)
}] }}
when: inventory_hostname in groups['ceph_bootstrap']
- name: Check RGW daemons
ansible.builtin.shell: |
set -o pipefail
ceph orch ls --service-type rgw --format json 2>/dev/null \
| python3 -c "
import sys, json
svcs = json.load(sys.stdin)
print(sum(s.get('status',{}).get('running',0) for s in svcs))
"
args:
executable: /bin/bash
register: rgw_count
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Record RGW drift
ansible.builtin.set_fact:
drift_results: >-
{{ drift_results + [{
'category': 'ceph',
'item': 'RGW daemons',
'expected': groups['ceph_nodes'] | length | string,
'actual': rgw_count.stdout | trim,
'match': (rgw_count.stdout | trim)
== (groups['ceph_nodes'] | length | string)
}] }}
when: inventory_hostname in groups['ceph_bootstrap']
- name: Check cluster health
ansible.builtin.command: ceph health --format json
register: health_check
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Record health drift
ansible.builtin.set_fact:
drift_results: >-
{{ drift_results + [{
'category': 'ceph',
'item': 'cluster health',
'expected': 'HEALTH_OK',
'actual': (health_check.stdout | from_json).status,
'match': (health_check.stdout | from_json).status
== 'HEALTH_OK'
}] }}
when: inventory_hostname in groups['ceph_bootstrap']
# Ceph config values
- name: Check Ceph config values
ansible.builtin.command: "ceph config get osd {{ item.key }}"
loop:
- key: osd_recovery_max_active
expected: "{{ ceph_osd_recovery_max_active }}"
- key: osd_max_backfills
expected: "{{ ceph_osd_max_backfills }}"
- key: osd_scrub_begin_hour
expected: "{{ ceph_osd_scrub_begin_hour }}"
- key: osd_scrub_end_hour
expected: "{{ ceph_osd_scrub_end_hour }}"
register: ceph_config_checks
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
loop_control:
label: "{{ item.key }}"
- name: Record Ceph config drift
ansible.builtin.set_fact:
drift_results: >-
{{ drift_results + [{
'category': 'ceph-config',
'item': item.item.key,
'expected': item.item.expected | string,
'actual': item.stdout | trim,
'match': (item.stdout | trim) == (item.item.expected | string)
}] }}
loop: "{{ ceph_config_checks.results | default([]) }}"
when: inventory_hostname in groups['ceph_bootstrap']
loop_control:
label: "{{ item.item.key | default('skipped') }}"
# --- Report ---
- name: Build drift report
ansible.builtin.set_fact:
drift_report: |
=== Drift Report: {{ inventory_hostname }} @ {{ lookup('pipe', 'date -u +%Y-%m-%dT%H:%M:%SZ') }} ===
{% for r in drift_results %}
{% if r.match %}
{{ '%-14s' | format(r.category) }} {{ '%-30s' | format(r.item) }} expected: {{ '%-12s' | format(r.expected) }} actual: {{ '%-12s' | format(r.actual) }} OK
{% else %}
{{ '%-14s' | format(r.category) }} {{ '%-30s' | format(r.item) }} expected: {{ '%-12s' | format(r.expected) }} actual: {{ '%-12s' | format(r.actual) }} !! DRIFT
{% endif %}
{% endfor %}
{% set drifts = drift_results | selectattr('match', 'equalto', false) | list %}
{% if drifts | length == 0 %}
No drift detected. Cluster matches expected state.
{% else %}
{{ drifts | length }} drift(s) detected. Run 'mise run deploy' to converge.
{% endif %}
- name: Display drift report
ansible.builtin.debug:
msg: "{{ drift_report.split('\n') }}"
+19
View File
@@ -0,0 +1,19 @@
---
# Security hardening for Ceph nodes.
# Deploys nftables firewall rules and SSH hardening.
# Run AFTER deploy-ceph.yml — cephadm needs unrestricted access during deploy.
#
# WARNING: This locks down inbound traffic. Ensure ceph_firewall_trusted_networks
# includes all subnets that need to reach Ceph services. SSH (port 22) is always
# open from any source to prevent lockout.
#
# Usage:
# scripts/ansible-play.sh harden.yml
- name: Security hardening for Ceph nodes
hosts: ceph_nodes
become: true
gather_facts: false
roles:
- security
+85
View File
@@ -0,0 +1,85 @@
---
- name: Hardware inventory
hosts: ceph_nodes
gather_facts: true
become: true
tasks:
- name: Gather dmidecode memory info
ansible.builtin.command: dmidecode -t memory
register: dmidecode_memory
changed_when: false
- name: Gather lsblk disk info
ansible.builtin.command: lsblk -d -J -o NAME,SIZE,MODEL,ROTA,TRAN,SERIAL
register: lsblk_output
changed_when: false
- name: Gather network interface info
ansible.builtin.command: ip -j link show
register: ip_link_output
changed_when: false
- name: Parse system info
ansible.builtin.set_fact:
hw_info:
hostname: "{{ ansible_hostname }}"
ip: "{{ ansible_host }}"
system:
manufacturer: "{{ ansible_system_vendor }}"
product: "{{ ansible_product_name }}"
serial: "{{ ansible_product_serial }}"
cpu:
model: "{{ ansible_processor[2] }}"
sockets: "{{ ansible_processor_count }}"
cores_per_socket: "{{ ansible_processor_cores }}"
threads_per_core: "{{ ansible_processor_threads_per_core }}"
total_vcpus: "{{ ansible_processor_vcpus }}"
memory:
total_mb: "{{ ansible_memtotal_mb }}"
dimms: "{{ dmidecode_memory.stdout }}"
storage:
disks: "{{ (lsblk_output.stdout | from_json).blockdevices }}"
network:
interfaces: "{{ ip_link_output.stdout | from_json }}"
os:
distribution: "{{ ansible_distribution }}"
version: "{{ ansible_distribution_version }}"
kernel: "{{ ansible_kernel }}"
- name: Ensure hardware/ output directory exists
delegate_to: localhost
become: false
ansible.builtin.file:
path: "{{ playbook_dir }}/hardware"
state: directory
mode: '0755'
run_once: true # noqa: run-once[task]
- name: Write per-host JSON
delegate_to: localhost
become: false
ansible.builtin.copy:
content: "{{ hw_info | to_nice_json }}"
dest: "{{ playbook_dir }}/hardware/{{ inventory_hostname }}.json"
mode: '0644'
- name: Stash for aggregation
ansible.builtin.set_fact:
_collected: "{{ hw_info }}"
- name: Combine inventory into single file
hosts: localhost
gather_facts: false
tasks:
- name: Build combined inventory
ansible.builtin.set_fact:
combined: >-
{{ combined | default({}) |
combine({item: hostvars[item]['_collected']}) }}
loop: "{{ groups['ceph_nodes'] }}"
- name: Write combined JSON
ansible.builtin.copy:
content: "{{ combined | to_nice_json }}"
dest: "{{ playbook_dir }}/hardware/combined.json"
mode: '0644'
@@ -0,0 +1,71 @@
---
# === Naming ===
cluster_name: painbox
cluster_role: ceph
# === Network ===
cluster_domain: dev.hel.htz.futo.cloud
public_network: 157.180.105.192/26
cluster_network: 157.180.105.192/26
gateway: 157.180.105.193
dns_server: 185.12.64.1
timezone: UTC
# No bond — single NIC, direct SSH (no ProxyJump)
# === Ceph ===
ceph_release: tentacle
ceph_repo_url: "https://download.ceph.com/debian-{{ ceph_release }}/"
ceph_repo_key_url: "https://download.ceph.com/keys/release.asc"
# === OS / Auth ===
admin_user: root
# 1P vault for cluster secret lookups (e.g., rotate-ssh-key.yml reads pubkey from here).
cluster_secrets_vault: yucca_tf_dev
# Secret aliases — vault_* vars populated by scripts/ansible-play.sh via op inject
ops_password: "{{ vault_ops_password }}"
ceph_dashboard_user: admin
ceph_dashboard_password: "{{ vault_ceph_dashboard_password }}"
# S3 svc-user (yucca-restic consumer) — TF+1P-predetermined keys passed to
# `radosgw-admin user create --access-key=... --secret-key=...` so the Yucca
# app can be pre-configured with matching credentials.
ceph_rgw_s3_user_access_key: "{{ vault_s3_restic_access_key }}"
ceph_rgw_s3_user_secret_key: "{{ vault_s3_restic_secret_key }}"
# === Storage ===
ssd_model_pattern: "SAMSUNG"
ceph_db_lv_size: "128G"
ceph_db_lvs_per_node: 14
# === RGW (Object Gateway) ===
ceph_rgw_realm: painbox
ceph_rgw_zonegroup: eu-central-1
ceph_rgw_zonegroup_api_name: eu-central-1
ceph_rgw_zone: dev-z1
ceph_rgw_ec_profile: ec-k8m3-osd
ceph_rgw_ec_k: 8
ceph_rgw_ec_m: 3
ceph_rgw_ec_failure_domain: osd
ceph_rgw_ec_device_class: hdd
ceph_rgw_data_pool: "{{ ceph_rgw_zone }}.rgw.buckets.data"
ceph_rgw_index_pool: "{{ ceph_rgw_zone }}.rgw.buckets.index"
ceph_rgw_extra_pool: "{{ ceph_rgw_zone }}.rgw.buckets.non-ec"
ceph_rgw_dns_name: s3.{{ cluster_domain }}
ceph_rgw_replicated_size: 2
ceph_rgw_replicated_min_size: 1
ceph_rgw_count_per_host: 1
ceph_rgw_port: 7480
ceph_rgw_ssl: false
ceph_rgw_scheme: "{{ 'https' if ceph_rgw_ssl else 'http' }}"
ceph_rgw_s3_user_uid: svc-yucca-restic
ceph_rgw_s3_user_display_name: "yucca/restic service account"
# === Monitoring Stack ===
ceph_prometheus_port: 9095
ceph_grafana_port: 3000
ceph_alertmanager_port: 9093
ceph_grafana_admin_user: admin
ceph_grafana_admin_password: "{{ vault_grafana_admin_password }}"
@@ -0,0 +1,35 @@
---
# Hetzner SX295 OSD node template — painbox cluster
#
# Copy to host_vars/<hostname>.yml (full inventory_hostname).
# Filename MUST match inventory_hostname for Ansible auto-load.
# Values are node-specific — SATA controller addresses differ per server.
#
# Hardware: EPYC 7502P / 2x 7.68TB NVMe (RAID-1) / 14x 22TB SATA
# NVMe pair is installimage RAID-1, partitioned as:
# md0 = /boot, md1 = vg0 (root, var, swap, block.db LVs, SSD OSD)
# SATA drives are left raw for Ceph HDD OSDs.
hostname_short: painbox-ceph-EXAMPLE
bond_ip: 157.180.105.X
# Block.db VG — on NVMe RAID-1 vg0 (created by installimage post-install)
ceph_db_vg: vg0
# HDD OSD mappings: SATA by-path -> block.db LV
# Find SATA controllers with: ls /dev/disk/by-path/ | grep ata
# Each SATA port maps to one db-slot LV, assigned positionally.
#
# Typical SX295 SATA topology (3 controllers, 14 ports total):
# 0000:45:00.0 (8 ports): ata-1..8
# 0000:46:00.0 (2 ports): ata-1..2
# 0000:87:00.0 (4 ports): ata-1..4
ceph_hdd_osds:
- path_phy: pci-0000:45:00.0-ata-1
db: vg0/db-slot0
- path_phy: pci-0000:45:00.0-ata-2
db: vg0/db-slot1
# ... one entry per HDD (14 total for full SX295)
# SSD OSD — remaining NVMe space after block.db LVs + reserve
ceph_ssd_osds:
- lv: vg0/ssd-osd
@@ -0,0 +1,53 @@
---
# Painbox host_vars — the single Hetzner SX295 node.
# Reprovisioned under this identity 2026-04-26 (was painbox-osd-5c3cac).
# Reprovision procedure: docs/runbooks/painbox-reprovision.md.
hostname_short: painbox-ceph-evelyn
bond_ip: 157.180.105.198
# Block.db VG — single NVMe RAID-1 VG created by installimage post-install
ceph_db_vg: vg0
# HDD OSD mappings: SATA path -> block.db LV
# by-path is slot-stable: replacing a drive in the same bay keeps the same path.
# All 14 SATA ports populated with Seagate Exos X22 22TB (ST22000NM001E).
#
# SATA controller topology:
# 0000:45:00.0 (8 ports): ata-1..8
# 0000:46:00.0 (2 ports): ata-1..2
# 0000:87:00.0 (4 ports): ata-1..4
# Total: 14 ports -> 14 HDDs -> 14 db-slots (db-slot0..13)
ceph_hdd_osds:
- path_phy: pci-0000:45:00.0-ata-1 # ZX299K3G
db: vg0/db-slot0
- path_phy: pci-0000:45:00.0-ata-2 # ZX297VJE
db: vg0/db-slot1
- path_phy: pci-0000:45:00.0-ata-3 # ZX2992CN
db: vg0/db-slot2
- path_phy: pci-0000:45:00.0-ata-4 # ZX295NSB
db: vg0/db-slot3
- path_phy: pci-0000:45:00.0-ata-5 # ZX29DE1R
db: vg0/db-slot4
- path_phy: pci-0000:45:00.0-ata-6 # ZX299AGR
db: vg0/db-slot5
- path_phy: pci-0000:45:00.0-ata-7 # ZX298TEQ
db: vg0/db-slot6
- path_phy: pci-0000:45:00.0-ata-8 # ZX29DE34
db: vg0/db-slot7
- path_phy: pci-0000:46:00.0-ata-1 # ZX294A8Z
db: vg0/db-slot8
- path_phy: pci-0000:46:00.0-ata-2 # ZX299EHE
db: vg0/db-slot9
- path_phy: pci-0000:87:00.0-ata-1 # ZX29DE21
db: vg0/db-slot10
- path_phy: pci-0000:87:00.0-ata-2 # ZX294PJ3
db: vg0/db-slot11
- path_phy: pci-0000:87:00.0-ata-3 # ZX29DDD1
db: vg0/db-slot12
- path_phy: pci-0000:87:00.0-ata-4 # ZX297AZ8
db: vg0/db-slot13
# SSD OSD — remaining NVMe RAID-1 space after block.db LVs + reserve
ceph_ssd_osds:
- lv: vg0/ssd-osd
@@ -0,0 +1,32 @@
## Hetzner installimage autosetup for SX295 Ceph node
## Cluster: painbox-ceph.dev.hel.htz.futo.cloud (reprovision target)
## Hardware: EPYC 7502P / 2x 7.68TB NVMe / 14x 22TB SATA
##
## Usage (from rescue):
## cat > /autosetup < autosetup
## installimage -a -c /autosetup -x /tmp/post-install.sh
##
## IMPORTANT: Only NVMe drives are listed. SATA drives are untouched.
## Ceph: Tentacle (v20) — EOL 2027-11-18
## BOOT MODE: BIOS (verified 2026-03-27). If UEFI, add before /boot:
## PART /boot/efi esp 256M
DRIVE1 /dev/nvme0n1
DRIVE2 /dev/nvme1n1
SWRAID 1
SWRAIDLEVEL 1
BOOTLOADER grub
HOSTNAME painbox-ceph-evelyn.dev.hel.htz.futo.cloud
PART /boot ext4 1G
PART lvm vg0 all
LV vg0 swap swap swap 32G
LV vg0 root / ext4 100G
LV vg0 var /var ext4 200G
LV vg0 varlog /var/log ext4 50G
IMAGE /root/.oldroot/nfs/install/../images/Debian-bookworm-latest-amd64-base.tar.zst
@@ -0,0 +1,90 @@
#!/bin/bash
# Post-install script for SX295 Ceph node (painbox)
# Runs in chroot after Hetzner installimage completes.
#
# Sets up: LVM for block.db + SSD OSD, SSH keys, cephadm prereqs
# Does NOT: touch SATA drives, deploy Ceph (that's Ansible's job)
#
# RENDER BEFORE USE:
# op inject -f -i post-install.sh.tpl -o /tmp/post-install.sh.rendered
# Then upload the rendered file to the rescue system. See
# docs/runbooks/painbox-reprovision.md.
#
# The rendered script authorizes only the ansible-iac pubkey fetched live
# from 1P — rotations propagate at reprovision time without a template edit.
set -euo pipefail
set -x
echo "=== post-install: starting ==="
# --- SSH authorized keys ---
mkdir -p /root/.ssh
chmod 700 /root/.ssh
cat >> /root/.ssh/authorized_keys << 'SSHKEYS'
op://yucca_tf_dev/PAINBOX_CEPH_ANSIBLE_IAC_SSH_KEY/public_key
SSHKEYS
chmod 600 /root/.ssh/authorized_keys
# --- System packages ---
apt-get update -qq
apt-get install -y -qq \
lvm2 \
python3 \
python3-pip \
curl \
jq \
smartmontools \
chrony \
podman \
bzip2
# --- NVMe RAID-1 LVM for Ceph ---
# After installimage, the NVMe RAID-1 has:
# md0 = /boot (1G)
# md1 = LVM PV for vg0 (rest of NVMe)
# vg0 = swap(32G) + root(100G) + var(200G) + varlog(50G)
#
# Remaining space in vg0: ~7.28 TiB - 382 GiB = ~5.53 TiB free
# Allocate: 14 x 128 GiB block.db LVs + leave 0.5 TiB reserve + rest for SSD OSD
echo "=== post-install: creating block.db LVs ==="
if ! vgs vg0 &>/dev/null; then
echo "FATAL: vg0 not found. Check installimage partition config."
exit 1
fi
for i in $(seq 0 13); do
lvcreate --yes -Wy -L 128G -n db-slot${i} vg0
echo " created vg0/db-slot${i} (128 GiB)"
done
# SSD OSD: allocate all remaining space minus 512 GiB reserve
VG_FREE_GIB=$(vgs vg0 --noheadings --nosuffix --units g -o vg_free | tr -d ' ' | cut -d. -f1)
RESERVE_GIB=512
OSD_GIB=$((VG_FREE_GIB - RESERVE_GIB))
if [ "$OSD_GIB" -gt 0 ]; then
lvcreate --yes -Wy -L "${OSD_GIB}G" -n ssd-osd vg0
echo " created vg0/ssd-osd (${OSD_GIB} GiB, ${RESERVE_GIB} GiB reserved)"
else
echo " WARNING: not enough space for SSD OSD LV"
fi
echo "=== post-install: LVM layout ==="
lvs vg0 -o lv_name,lv_size --noheadings
# --- Timezone ---
ln -sf /usr/share/zoneinfo/UTC /etc/localtime
# --- Sysctl tuning for Ceph ---
cat > /etc/sysctl.d/90-ceph.conf << 'SYSCTL'
# Ceph OSD tuning
vm.min_free_kbytes = 1048576
vm.swappiness = 10
net.core.rmem_max = 67108864
net.core.wmem_max = 67108864
SYSCTL
echo "=== post-install: complete ==="
@@ -0,0 +1,111 @@
---
# === Naming ===
cluster_name: sietch
cluster_role: ceph
# === Network ===
cluster_domain: dev.austin.int.futo.cloud
public_network: 10.10.10.0/24
cluster_network: 10.10.10.0/24
gateway: 10.10.10.1
dns_server: 10.10.10.1
bond_mode: active-backup
bond_interfaces:
- eno1
- eno2
networkd_enabled: true
# === Ceph ===
ceph_release: tentacle
ceph_repo_url: "https://download.ceph.com/debian-{{ ceph_release }}/"
ceph_repo_key_url: "https://download.ceph.com/keys/release.asc"
# === OS Provisioning ===
admin_user: ansible-iac
timezone: UTC
# Deploy keypair path on the controller. Both {path} and {path}.pub must
# exist before running provision.yml (the preflight play in provision.yml
# enforces this). ansible-iac's authorized_keys on every node is populated
# by a file lookup at provision time — rotating the key on the controller
# automatically propagates on the next re-provision.
provision_iac_ssh_key_path: "~/.ssh/id_ed25519_sietch"
# 1P vault for cluster secret lookups (e.g., rotate-ssh-key.yml reads pubkey from here).
cluster_secrets_vault: yucca_tf_dev
# Secret aliases — vault_* vars populated by scripts/ansible-play.sh via op inject
#
# Note: the deploy public key is NOT a secret — it's public data and lives
# at {{ provision_iac_ssh_key_path }}.pub on the controller. Reading it via
# file lookup avoids drift between vault and disk.
ops_password: "{{ vault_ops_password }}"
ceph_dashboard_user: admin
ceph_dashboard_password: "{{ vault_ceph_dashboard_password }}"
# S3 svc-user (yucca-restic consumer) — TF+1P-predetermined keys passed to
# `radosgw-admin user create --access-key=... --secret-key=...` so the Yucca
# app can be pre-configured with matching credentials. Rotation path documented
# in docs/runbooks/rotate-secrets.md.
ceph_rgw_s3_user_access_key: "{{ vault_s3_restic_access_key }}"
ceph_rgw_s3_user_secret_key: "{{ vault_s3_restic_secret_key }}"
# === Storage ===
ssd_model_pattern: "Micron_5100"
os_partitions:
esp_size: "512M"
boot_size: "1G"
root_size: "80G"
swap_size: "8G"
ceph_db_size: "1440G"
ceph_db_lv_size: "240G"
ceph_db_lvs_per_ssd: 6
# === RGW (Object Gateway) ===
ceph_rgw_realm: sietch
ceph_rgw_zonegroup: us-east-1
ceph_rgw_zonegroup_api_name: us-east-1
ceph_rgw_zone: dev-z1
ceph_rgw_ec_profile: ec-k8m3-osd
ceph_rgw_ec_k: 8
ceph_rgw_ec_m: 3
ceph_rgw_ec_failure_domain: osd
ceph_rgw_ec_device_class: hdd
ceph_rgw_data_pool: "{{ ceph_rgw_zone }}.rgw.buckets.data"
ceph_rgw_index_pool: "{{ ceph_rgw_zone }}.rgw.buckets.index"
ceph_rgw_extra_pool: "{{ ceph_rgw_zone }}.rgw.buckets.non-ec"
ceph_rgw_replicated_size: 2
ceph_rgw_replicated_min_size: 1
ceph_rgw_port: 443
ceph_rgw_s3_user_uid: svc-yucca-restic
ceph_rgw_s3_user_display_name: "yucca/restic service account"
# --- RGW DNS + TLS ---
# Virtual-hosted S3 support: setting rgw_dns_name tells RGW to strip this
# suffix from the Host header and treat the remainder as the bucket name.
# Requires matching DNS: both s3.<domain> and *.s3.<domain> should resolve
# to the cluster nodes (round-robin A, or VIP/LB in prod).
ceph_rgw_dns_name: s3.dev.austin.int.futo.cloud
# Self-signed wildcard cert handed to cephadm via service spec.
# cephadm distributes to all RGW daemons. To rotate, delete
# /etc/ceph/rgw-ssl.{crt,key} on the bootstrap node and re-run the role.
ceph_rgw_ssl: true
ceph_rgw_ssl_cert_days: 3650 # 10 years
ceph_rgw_ssl_cert_subject_c: US
ceph_rgw_ssl_cert_subject_st: Texas
ceph_rgw_ssl_cert_subject_l: Austin
ceph_rgw_ssl_cert_subject_o: FUTO
ceph_rgw_ssl_cert_email: yucca@futo.org
# Computed: scheme used for endpoints, debug output, and zonegroup/zone URLs
ceph_rgw_scheme: "{{ 'https' if ceph_rgw_ssl else 'http' }}"
# === Monitoring Stack ===
ceph_prometheus_port: 9095
ceph_grafana_port: 3000
ceph_alertmanager_port: 9093
ceph_grafana_admin_user: admin
ceph_grafana_admin_password: "{{ vault_grafana_admin_password }}"
@@ -0,0 +1,36 @@
---
# Copy to host_vars/sietch-ceph-<name>.yml (full inventory_hostname).
# The <name> segment is either operator-declared in clusters.auto.tfvars
# or TF-auto-picked from wordlist.txt — see docs/naming.md.
# Filename MUST match inventory_hostname for Ansible auto-load.
# Values are node-specific — hardware paths differ per chassis.
hostname_short: sietch-ceph-EXAMPLE
bond_ip: 10.0.0.1
# SAS expander base path (unique per chassis)
# Find with: ls /dev/disk/by-path/ | grep sas
sas_path_prefix: "pci-0000:02:00.0-sas-exp0x500056b3XXXXXXXX"
# SSD PHY positions in the SAS topology
ssd1_phy: 12
ssd2_phy: 13
# LVM volume group names for block.db (on SSD partition 5)
ceph_db_vg1: ceph-db-ssd1
ceph_db_vg2: ceph-db-ssd2
# HDD OSD mappings: SAS PHY slot -> block.db LV
# 6 HDDs per SSD, each gets a dedicated 240G block.db LV
ceph_hdd_osds:
- path_phy: phy0
db: ceph-db-ssd1/db-slot0
- path_phy: phy1
db: ceph-db-ssd1/db-slot1
# ... one entry per HDD
# SSD OSD partitions (partition 6 on each SSD, no separate block.db)
ceph_ssd_osds:
- path_phy: phy12
partition: 6
- path_phy: phy13
partition: 6
@@ -0,0 +1,52 @@
---
hostname_short: sietch-ceph-laurel
bond_ip: 10.10.10.90
# SAS expander base path (unique per chassis — different backplane address per node)
sas_path_prefix: "pci-0000:02:00.0-sas-exp0x500056b3fcf498ff"
# SSD PHY positions (rear bays)
ssd1_phy: 12 # serial 17321A07BA4A
ssd2_phy: 13 # serial 17321A07CFEE
# LVM VGs on SSD partition 5
ceph_db_vg1: ceph-db-rear12 # VG on SSD1 (phy12) partition 5
ceph_db_vg2: ceph-db-rear13 # VG on SSD2 (phy13) partition 5
# HDD OSD mappings: PHY slot -> block.db LV
# PHY 0-5 -> SSD1 (ceph-db-rear12/db-slot0..5)
# PHY 6-11 -> SSD2 (ceph-db-rear13/db-slot6..11)
# by-path is slot-stable: replacing a drive in the same bay keeps the same path
# 12 HDDs — all front bays populated
ceph_hdd_osds:
- path_phy: phy0
db: ceph-db-rear12/db-slot0
- path_phy: phy1
db: ceph-db-rear12/db-slot1
- path_phy: phy2
db: ceph-db-rear12/db-slot2
- path_phy: phy3
db: ceph-db-rear12/db-slot3 # serial Z4D09B99 (ST6000NKCLAR6000, added 2026-04-10)
- path_phy: phy4
db: ceph-db-rear12/db-slot4
- path_phy: phy5
db: ceph-db-rear12/db-slot5 # serial Z4D0G7VC (ST6000NKCLAR6000, added 2026-04-10)
- path_phy: phy6
db: ceph-db-rear13/db-slot6
- path_phy: phy7
db: ceph-db-rear13/db-slot7 # serial Z4D0G7SC (ST6000NKCLAR6000, added 2026-04-10)
- path_phy: phy8
db: ceph-db-rear13/db-slot8
- path_phy: phy9
db: ceph-db-rear13/db-slot9
- path_phy: phy10
db: ceph-db-rear13/db-slot10 # serial Z4D0G7XK (ST6000NKCLAR6000, added 2026-04-10)
- path_phy: phy11
db: ceph-db-rear13/db-slot11
# SSD OSD partitions (partition 6 on each rear SSD, no separate block.db)
ceph_ssd_osds:
- path_phy: phy12
partition: 6
- path_phy: phy13
partition: 6
@@ -0,0 +1,52 @@
---
hostname_short: sietch-ceph-lawson
bond_ip: 10.10.10.91
# SAS expander base path (unique per chassis — different backplane address per node)
sas_path_prefix: "pci-0000:02:00.0-sas-exp0x500056b35be6b4ff"
# SSD PHY positions (rear bays)
ssd1_phy: 12 # serial 17321A07CF91
ssd2_phy: 13 # serial 17251B44C4D8
# LVM VGs on SSD partition 5
ceph_db_vg1: ceph-db-rear12 # VG on SSD1 (phy12) partition 5
ceph_db_vg2: ceph-db-rear13 # VG on SSD2 (phy13) partition 5
# HDD OSD mappings: PHY slot -> block.db LV
# PHY 0-5 -> SSD1 (ceph-db-rear12/db-slot0..5)
# PHY 6-11 -> SSD2 (ceph-db-rear13/db-slot6..11)
# by-path is slot-stable: replacing a drive in the same bay keeps the same path
# 12 HDDs — all front bays populated
ceph_hdd_osds:
- path_phy: phy0
db: ceph-db-rear12/db-slot0
- path_phy: phy1
db: ceph-db-rear12/db-slot1
- path_phy: phy2
db: ceph-db-rear12/db-slot2
- path_phy: phy3
db: ceph-db-rear12/db-slot3
- path_phy: phy4
db: ceph-db-rear12/db-slot4
- path_phy: phy5
db: ceph-db-rear12/db-slot5
- path_phy: phy6
db: ceph-db-rear13/db-slot6
- path_phy: phy7
db: ceph-db-rear13/db-slot7 # serial Z4D0GKJX (ST6000NKCLAR6000, added 2026-04-10)
- path_phy: phy8
db: ceph-db-rear13/db-slot8
- path_phy: phy9
db: ceph-db-rear13/db-slot9
- path_phy: phy10
db: ceph-db-rear13/db-slot10
- path_phy: phy11
db: ceph-db-rear13/db-slot11
# SSD OSD partitions (partition 6 on each rear SSD, no separate block.db)
ceph_ssd_osds:
- path_phy: phy12
partition: 6
- path_phy: phy13
partition: 6
@@ -0,0 +1,52 @@
---
hostname_short: sietch-ceph-samara
bond_ip: 10.10.10.92
# SAS expander base path (unique per chassis — different backplane address per node)
sas_path_prefix: "pci-0000:02:00.0-sas-exp0x500056b3393ba1ff"
# SSD PHY positions (rear bays)
ssd1_phy: 12 # serial 17321A07D1D3
ssd2_phy: 13 # serial 17251B44DF51
# LVM VGs on SSD partition 5
ceph_db_vg1: ceph-db-rear12 # VG on SSD1 (phy12) partition 5
ceph_db_vg2: ceph-db-rear13 # VG on SSD2 (phy13) partition 5
# HDD OSD mappings: PHY slot -> block.db LV
# PHY 0-5 -> SSD1 (ceph-db-rear12/db-slot0..5)
# PHY 6-11 -> SSD2 (ceph-db-rear13/db-slot6..11)
# by-path is slot-stable: replacing a drive in the same bay keeps the same path
# 12 HDDs — all front bays populated
ceph_hdd_osds:
- path_phy: phy0
db: ceph-db-rear12/db-slot0
- path_phy: phy1
db: ceph-db-rear12/db-slot1
- path_phy: phy2
db: ceph-db-rear12/db-slot2
- path_phy: phy3
db: ceph-db-rear12/db-slot3
- path_phy: phy4
db: ceph-db-rear12/db-slot4 # serial Z4D0GKB2 (replacement, added 2026-04-11)
- path_phy: phy5
db: ceph-db-rear12/db-slot5
- path_phy: phy6
db: ceph-db-rear13/db-slot6
- path_phy: phy7
db: ceph-db-rear13/db-slot7
- path_phy: phy8
db: ceph-db-rear13/db-slot8
- path_phy: phy9
db: ceph-db-rear13/db-slot9
- path_phy: phy10
db: ceph-db-rear13/db-slot10
- path_phy: phy11
db: ceph-db-rear13/db-slot11
# SSD OSD partitions (partition 6 on each rear SSD, no separate block.db)
ceph_ssd_osds:
- path_phy: phy12
partition: 6
- path_phy: phy13
partition: 6
+62
View File
@@ -0,0 +1,62 @@
---
# Migrate ifupdown -> systemd-networkd on sietch Ceph nodes.
#
# Live migration — no reboot, no Ceph disruption:
# 1. Deploys networkd configs
# 2. Starts networkd (adopts existing bond0)
# 3. Verifies bond0 IP, members, gateway
# 4. Commits: disables ifupdown, enables networkd for boot
#
# If verification fails, the play stops before committing.
# networkd can be stopped and ifupdown still owns next boot.
#
# Safety:
# - serial: 1 -- one node at a time
# - /etc/network/interfaces preserved on disk
# - Rollback: /usr/local/sbin/rollback-networkd.sh
# - iDRAC at 10.10.11.9{0,1,2} for emergency recovery
#
# Usage (samara first):
# scripts/ansible-play.sh migrate-networkd.yml --limit sietch-ceph-samara
#
# Usage (remaining nodes):
# scripts/ansible-play.sh migrate-networkd.yml --limit 'ceph_nodes:!sietch-ceph-samara'
#
# Dry run:
# scripts/ansible-play.sh migrate-networkd.yml --limit sietch-ceph-samara --check --diff
- name: Migrate to systemd-networkd
hosts: ceph_nodes
become: true
gather_facts: false
serial: 1
pre_tasks:
- name: Preflight -- check Ceph health
ansible.builtin.command: ceph health --format json
register: ceph_health_pre
changed_when: false
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
- name: Assert cluster is healthy
ansible.builtin.assert:
that:
- (ceph_health_pre.stdout | from_json).status in ['HEALTH_OK', 'HEALTH_WARN']
fail_msg: >-
Ceph is {{ (ceph_health_pre.stdout | from_json).status }}.
Fix cluster health before migrating.
success_msg: "Ceph: {{ (ceph_health_pre.stdout | from_json).status }}"
roles:
- role: networkd
post_tasks:
- name: Confirm Ceph health after migration
ansible.builtin.command: ceph health --format json
register: ceph_health_post
changed_when: false
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
- name: Report Ceph health
ansible.builtin.debug:
msg: "Ceph: {{ (ceph_health_post.stdout | from_json).status }}"
+96
View File
@@ -0,0 +1,96 @@
---
# post-deploy-capture.yml — snapshot bootstrap-side secrets into 1Password.
#
# Captures:
# /etc/ceph/rgw-ssl.crt -> <CLUSTER>_CEPH_RGW_TLS_CERT
# /etc/ceph/rgw-ssl.key -> <CLUSTER>_CEPH_RGW_TLS_KEY
# /etc/ceph/ceph.client.admin.keyring -> <CLUSTER>_CEPH_CLIENT_ADMIN_KEYRING
#
# Idempotent: creates item if missing, updates if content drifted.
#
# Belt-and-suspenders only. The cluster continues to run from on-node
# copies; the 1P copies exist for disaster recovery (e.g., laurel dies +
# filesystem loss + you want to restore the RGW cert to a new bootstrap
# node without re-trusting it in every S3 client).
#
# Usage:
# scripts/ansible-play.sh post-deploy-capture.yml
#
# Requires: superuser SA read access from controller (already required by
# mise run tf:* tasks), op CLI on controller, 1P session live.
- name: Capture bootstrap-side secrets to 1Password
hosts: ceph_bootstrap
become: true
gather_facts: false
strategy: linear
vars:
capture_items:
- file: /etc/ceph/rgw-ssl.crt
title_suffix: RGW_TLS_CERT
- file: /etc/ceph/rgw-ssl.key
title_suffix: RGW_TLS_KEY
- file: /etc/ceph/ceph.client.admin.keyring
title_suffix: CLIENT_ADMIN_KEYRING
tasks:
- name: Read bootstrap files
ansible.builtin.slurp:
src: "{{ item.file }}"
loop: "{{ capture_items }}"
register: slurped
no_log: true
- name: Fetch superuser SA token from 1P (localhost) # noqa: run-once[task]
ansible.builtin.command:
argv:
- op
- read
- "op://yucca_tf_dev/yucca_futo_1pass_superuser_service_account/password"
register: su_token
delegate_to: localhost
become: false # run as operator on controller — root has no op session
run_once: true
changed_when: false
check_mode: false
no_log: true
# exec 0<&- closes stdin. op CLI treats non-TTY piped stdin as a JSON
# item template; without this redirect op fails with "invalid JSON in
# piped input" because ansible pipes an empty stdin to the shell.
- name: Upsert each captured file to 1Password # noqa: run-once[task]
ansible.builtin.shell: |
set -euo pipefail
exec 0<&-
TITLE="{{ cluster_name | upper }}_CEPH_{{ item.item.title_suffix }}"
VALUE="$CONTENT"
if op item get "$TITLE" --vault yucca_tf_dev >/dev/null 2>&1; then
op item edit "$TITLE" --vault yucca_tf_dev "password=$VALUE" >/dev/null
echo "updated: $TITLE"
else
op item create --vault yucca_tf_dev --category password --title "$TITLE" "password=$VALUE" >/dev/null
echo "created: $TITLE"
fi
args:
executable: /bin/bash
environment:
OP_SERVICE_ACCOUNT_TOKEN: "{{ su_token.stdout | trim }}"
CONTENT: "{{ item.content | b64decode }}"
loop: "{{ slurped.results }}"
loop_control:
label: "{{ item.item.title_suffix }}"
delegate_to: localhost
become: false # run as operator on controller — op session is operator-owned
run_once: true
register: capture_result
changed_when: "'created:' in capture_result.stdout or 'updated:' in capture_result.stdout"
no_log: true # content is cert/key/keyring material — don't echo via ansible logs
- name: Capture summary # noqa: run-once[task]
ansible.builtin.debug:
msg: "{{ item.stdout }}"
loop: "{{ capture_result.results }}"
loop_control:
label: "{{ item.item.item.title_suffix }}"
run_once: true
+208
View File
@@ -0,0 +1,208 @@
---
# Provision sietch-ceph nodes from the Debian 12 live image.
#
# Fresh-install workflow (no cloning):
# 1. Boot all target nodes to Debian 12 live image (user/live credentials)
# 2. Ensure the live image is reachable on its reserved IP (= bond_ip)
# 3. CEPH_ENV=inventories/<cluster>/inventory-provision.ini \
# scripts/ansible-play.sh provision.yml -e confirm_wipe=true
#
# Flags:
# confirm_wipe=true required — destroys all data on SSDs
# provision_create_iac_keypair=true opt-in: preflight auto-generates
# the ansible-iac deploy keypair if
# it doesn't already exist on disk
# provision_skip_reboot=true opt-in: skip the final reboot and
# the post-reboot verification play
# (canary inspection of the chroot)
# --limit <host> provision one node at a time
#
# Post-reboot: a separate verification play waits on the installed OS
# (reached via the production sietch-ceph-<name> SSH alias / ansible-iac
# user with key) and confirms the provisioning marker is in place.
# --- Preflight: controller-side prerequisites -------------------------
#
# Runs BEFORE the target hosts are touched. Verifies the ansible-iac
# deploy keypair exists on the controller (it gets installed into every
# node's ~ansible-iac/.ssh/authorized_keys via file lookup during the
# main play). Keypair generation is strictly opt-in via
# -e provision_create_iac_keypair=true — default behavior is to fail
# fast with a clear message if the keypair is missing.
- name: Preflight — verify controller prerequisites
hosts: localhost
gather_facts: false
connection: local
tasks:
- name: Stat ansible-iac private key
ansible.builtin.stat:
path: "{{ provision_iac_ssh_key_path | expanduser }}"
register: iac_privkey_stat
- name: Stat ansible-iac public key
ansible.builtin.stat:
path: "{{ (provision_iac_ssh_key_path | expanduser) ~ '.pub' }}"
register: iac_pubkey_stat
- name: Generate ansible-iac keypair (only when opt-in flag is set)
ansible.builtin.command:
cmd: >-
ssh-keygen -t ed25519 -N ''
-C "ansible-iac@{{ cluster_name }}-{{ cluster_role }}.{{ cluster_domain }}"
-f {{ provision_iac_ssh_key_path | expanduser }}
args:
creates: "{{ provision_iac_ssh_key_path | expanduser }}"
when:
- not iac_privkey_stat.stat.exists
- provision_create_iac_keypair | default(false) | bool
- name: Fail if ansible-iac keypair missing and create flag not set
ansible.builtin.fail:
msg: |
ansible-iac deploy keypair is missing on the controller:
{{ provision_iac_ssh_key_path | expanduser }}
{{ provision_iac_ssh_key_path | expanduser }}.pub
This keypair is required to provision the cluster — it is
installed into ansible-iac's authorized_keys on every node
and used by ansible_ssh_private_key_file in inventory.ini.
To auto-generate it on this run, re-invoke with:
-e provision_create_iac_keypair=true
Or generate it yourself first:
ssh-keygen -t ed25519 -N '' \
-C "ansible-iac@{{ cluster_name }}-{{ cluster_role }}.{{ cluster_domain }}" \
-f {{ provision_iac_ssh_key_path }}
when:
- not iac_privkey_stat.stat.exists
- not (provision_create_iac_keypair | default(false) | bool)
- name: Re-stat public key after potential creation
ansible.builtin.stat:
path: "{{ (provision_iac_ssh_key_path | expanduser) ~ '.pub' }}"
register: iac_pubkey_final_stat
- name: Assert ansible-iac public key exists after preflight
ansible.builtin.assert:
that:
- iac_pubkey_final_stat.stat.exists
fail_msg: >-
ansible-iac public key still missing at
{{ (provision_iac_ssh_key_path | expanduser) ~ '.pub' }}
after preflight completed. Investigate ssh-keygen task above.
- name: Report ansible-iac keypair ready
ansible.builtin.debug:
msg: >-
ansible-iac keypair present at
{{ provision_iac_ssh_key_path | expanduser }}(.pub) —
preflight OK.
- name: Provision sietch-ceph nodes (live image → installed OS)
hosts: provision_targets
gather_facts: false
serial: 1
become: true
max_fail_percentage: 0
roles:
- role: provision_host
- name: Verify provisioned nodes after reboot
hosts: provision_targets
gather_facts: false
become: false
tasks:
# Verification runs on localhost via delegate_to + the production
# sietch-ceph-<name> SSH alias. That avoids having to juggle live
# vs installed credentials inside a single play.
- name: Skip verification play when reboot was intentionally skipped
ansible.builtin.meta: end_play
when: provision_skip_reboot | default(false) | bool
# Give the box a head start before we start hammering ssh. R730xd
# POST + GRUB + initramfs + systemd brings the bond up in ~60-90s.
- name: Initial reboot wait
delegate_to: localhost
ansible.builtin.pause:
seconds: "{{ provision_reboot_wait_delay | default(60) }}"
# The ssh probe is the actual gate — it honors ~/.ssh/config
# (including any ProxyJump or ProxyCommand entries) so it works
# from the controller regardless of whether the target subnet is
# directly routable. wait_for/tcp does NOT honor ssh_config, so
# we rely on retries here instead.
#
# We force user + identity explicitly with -l / -i / IdentitiesOnly
# because ~/.ssh/config for the host alias lands as the 'ops' human
# user (which has NO authorized_keys on a fresh provision — operator
# keys are distributed out-of-band). The verify play is an automation
# reachability test, so it uses ansible-iac (the same identity
# ansible_ssh_private_key_file points at in inventory.ini).
#
# UserKnownHostsFile=/dev/null + StrictHostKeyChecking=no bypass
# known_hosts entirely: fresh provisioning generates new host keys,
# and on reprovisioning a stale known_hosts entry would otherwise
# trigger "REMOTE HOST IDENTIFICATION HAS CHANGED" and abort. The
# real identity check happens via the marker assertion below, not
# via the SSH host key cache. LogLevel=ERROR suppresses the
# "Warning: Permanently added..." noise.
- name: Probe installed OS via production SSH alias (as ansible-iac)
delegate_to: localhost
ansible.builtin.command:
cmd: >-
ssh -o BatchMode=yes -o ConnectTimeout=10
-o StrictHostKeyChecking=no
-o UserKnownHostsFile=/dev/null
-o IdentitiesOnly=yes
-o LogLevel=ERROR
-l ansible-iac
-i {{ provision_iac_ssh_key_path | expanduser }}
{{ hostname_short }}
"cat {{ provision_marker_path | default('/etc/ceph-provisioned.json') }}"
register: marker_read
changed_when: false
retries: "{{ provision_verify_retries | default(30) }}"
delay: "{{ provision_verify_delay | default(15) }}"
until: marker_read.rc == 0
- name: Parse provisioning marker
ansible.builtin.set_fact:
marker: "{{ marker_read.stdout | from_json }}"
- name: Verify marker contents match inventory expectations
ansible.builtin.assert:
that:
- marker.hostname == hostname_short
- marker.bond_ip == bond_ip
- marker.cluster_domain == cluster_domain
- marker.admin_user == admin_user
fail_msg: "Marker contents do not match inventory: {{ marker }}"
success_msg: "{{ hostname_short }} provisioned OK at {{ marker.provisioned_at }}"
- name: Gather installed OS hostname (as ansible-iac)
delegate_to: localhost
ansible.builtin.command:
cmd: >-
ssh -o BatchMode=yes -o ConnectTimeout=10
-o StrictHostKeyChecking=no
-o UserKnownHostsFile=/dev/null
-o IdentitiesOnly=yes
-o LogLevel=ERROR
-l ansible-iac
-i {{ provision_iac_ssh_key_path | expanduser }}
{{ hostname_short }} hostname -f
register: installed_hostname
changed_when: false
- name: Display post-provision state
ansible.builtin.debug:
msg:
- "hostname -f: {{ installed_hostname.stdout | trim }}"
- "marker.hostname: {{ marker.hostname }}"
- "marker.fqdn: {{ marker.fqdn }}"
- "marker.bond_ip: {{ marker.bond_ip }}"
- "marker.provisioned: {{ marker.provisioned_at }}"
+103
View File
@@ -0,0 +1,103 @@
---
# RADOS benchmark — tests raw cluster I/O bypassing RGW/S3.
# Useful for isolating storage performance from S3 overhead.
#
# Usage:
# scripts/ansible-play.sh rados-bench.yml # 30s write, replicated
# scripts/ansible-play.sh rados-bench.yml -e rados_ec=true # EC 8+3 (matches RGW data pool)
# scripts/ansible-play.sh rados-bench.yml -e rados_mode=seq # sequential read
# scripts/ansible-play.sh rados-bench.yml -e rados_seconds=60
# scripts/ansible-play.sh rados-bench.yml -e rados_cleanup=false # keep data for read test
#
# Modes: write, seq (sequential read), rand (random read)
# Read modes require a prior write run with rados_cleanup=false.
- name: RADOS benchmark
hosts: ceph_nodes
become: true
gather_facts: false
vars:
rados_pool: "rados-bench"
rados_mode: write
rados_seconds: 30
rados_threads: 16
rados_cleanup: true
rados_ec: false
rados_ec_profile: "ec-k8m3-osd" # reuse the RGW EC profile (8+3)
tasks:
- name: Create EC bench pool if missing
ansible.builtin.shell: |
set -o pipefail
if ceph osd pool ls | grep -q '^{{ rados_pool }}$'; then
echo "EXISTS"
else
ceph osd pool create {{ rados_pool }} 64 erasure {{ rados_ec_profile }}
ceph osd pool set {{ rados_pool }} allow_ec_overwrites true
echo "CREATED"
fi
args:
executable: /bin/bash
when:
- inventory_hostname in groups['ceph_bootstrap']
- rados_ec | bool
register: ec_pool_create
changed_when: "'CREATED' in ec_pool_create.stdout | default('')"
- name: Create replicated bench pool if missing
ansible.builtin.shell: |
set -o pipefail
if ceph osd pool ls | grep -q '^{{ rados_pool }}$'; then
echo "EXISTS"
else
ceph osd pool create {{ rados_pool }} 32 replicated
ceph osd pool set {{ rados_pool }} size 2
ceph osd pool set {{ rados_pool }} min_size 1
echo "CREATED"
fi
args:
executable: /bin/bash
when:
- inventory_hostname in groups['ceph_bootstrap']
- not (rados_ec | bool)
register: rep_pool_create
changed_when: "'CREATED' in rep_pool_create.stdout | default('')"
- name: Show benchmark parameters
ansible.builtin.debug:
msg:
- "Node: {{ inventory_hostname }}"
- "Pool: {{ rados_pool }}"
- "Mode: {{ rados_mode }}"
- "Duration: {{ rados_seconds }}s"
- "Threads: {{ rados_threads }}"
- name: Run rados bench
ansible.builtin.command: >
rados bench -p {{ rados_pool }}
{{ rados_seconds }} {{ rados_mode }}
-t {{ rados_threads }}
--run-name {{ inventory_hostname }}
{{ '--no-cleanup' if not (rados_cleanup | bool) else '' }}
register: rados_result
changed_when: true
timeout: "{{ (rados_seconds | int * 3) + 60 }}"
- name: Display results
ansible.builtin.debug:
msg: "{{ rados_result.stdout_lines[-20:] }}"
- name: Clean up bench pool
ansible.builtin.shell: |
ceph config set mon mon_allow_pool_delete true
ceph osd pool delete {{ rados_pool }} {{ rados_pool }} \
--yes-i-really-really-mean-it
ceph config set mon mon_allow_pool_delete false
args:
executable: /bin/bash
when:
- inventory_hostname in groups['ceph_bootstrap']
- rados_cleanup | bool
- rados_mode == 'write'
changed_when: true
+6
View File
@@ -0,0 +1,6 @@
ansible-core==2.20.4
ansible-lint==26.4.0
ansible-navigator>=25.0
yamllint==1.38.0
molecule>=25.0
boto3==1.42.87
+6
View File
@@ -0,0 +1,6 @@
---
collections:
- name: ansible.posix
version: ">=2.0.0"
- name: community.general
version: ">=9.0.0"
@@ -0,0 +1,49 @@
---
# Post-boot OS baseline — runs via ansible-iac after provisioning.
# Configures everything that doesn't need to be in the chroot.
# --- ops user ---
baseline_ops_user: ops
baseline_ops_uid: 1001
baseline_ops_sudo: "ALL" # "ALL" = password required; "NOPASSWD:ALL" for dev
# --- /etc/hosts ---
# Rendered from the same cluster variables used by ceph_deploy/hosts.j2.
# Keeps host entries converged even if ceph_deploy hasn't run yet.
baseline_manage_hosts: true
# --- Packages ---
# Podman ecosystem (needed by cephadm)
baseline_podman_packages:
- podman
- catatonit
- dbus-user-session
- fuse-overlayfs
- slirp4netns
- uidmap
# Diagnostic and ops toolset
baseline_diag_packages:
- smartmontools
- ipmitool
- lsscsi
- pciutils
- lshw
- ethtool
- jq
- htop
- btop
- tmux
- bat
- ncdu
- nload
- sysstat
- iotop
- bsdmainutils
- python3-boto3
# --- Services ---
baseline_enable_services:
- dbus
- chrony
- podman.socket
+11
View File
@@ -0,0 +1,11 @@
---
# Runs after provision_host, before ceph_deploy.
# Requires ansible-iac SSH access (set up during provisioning).
dependencies: []
galaxy_info:
author: FUTO
license: AGPL-3.0-only
role_name: baseline
description: Post-boot OS baseline — ops user, packages, hosts, services
min_ansible_version: "2.19"
@@ -0,0 +1,18 @@
---
# Post-boot OS baseline.
# Runs via ansible-iac on the installed OS (not in chroot).
# Assumes: ansible-iac user exists with key-only SSH + NOPASSWD sudo
# (created by provision_host during initial OS install).
# Everything here is convergeable via normal deploy pipeline.
- name: Configure ops user
ansible.builtin.import_tasks: users.yml
tags: [users]
- name: Configure system packages
ansible.builtin.import_tasks: packages.yml
tags: [packages]
- name: Configure /etc/hosts and services
ansible.builtin.import_tasks: system.yml
tags: [system]
@@ -0,0 +1,18 @@
---
# Install podman ecosystem + diagnostic toolset.
# Convergeable: re-running installs anything missing.
- name: Update apt cache
ansible.builtin.apt:
update_cache: true
cache_valid_time: 3600
- name: Install podman ecosystem
ansible.builtin.apt:
name: "{{ baseline_podman_packages }}"
state: present
- name: Install diagnostic and ops packages
ansible.builtin.apt:
name: "{{ baseline_diag_packages }}"
state: present
@@ -0,0 +1,19 @@
---
# /etc/hosts, service enablement, timezone.
# Convergeable: re-running corrects drift in hosts file and services.
- name: Render /etc/hosts with cluster node entries
ansible.builtin.template:
src: hosts.j2
dest: /etc/hosts
owner: root
group: root
mode: '0644'
when: baseline_manage_hosts | bool
- name: Enable and start required services
ansible.builtin.systemd:
name: "{{ item }}"
enabled: true
state: started
loop: "{{ baseline_enable_services }}"
@@ -0,0 +1,38 @@
---
# ops user — human-interactive account.
# Convergeable: running this again fixes drift in password, sudo, shell.
- name: Ensure ops user exists
ansible.builtin.user:
name: "{{ baseline_ops_user }}"
uid: "{{ baseline_ops_uid }}"
shell: /bin/bash
groups: sudo
append: true
create_home: true
state: present
- name: Assert ops_password is defined and non-empty
ansible.builtin.assert:
that:
- ops_password is defined
- ops_password | length > 0
fail_msg: >-
ops_password is undefined or empty. Run the play via
scripts/ansible-play.sh (not bare ansible-playbook) so the cluster's
secrets.yml.tpl is resolved via op inject into --extra-vars. If that's
already what you did, verify vault_ops_password resolves:
`op read op://<vault>/<CLUSTER>_CEPH_OPS_PASSWORD/password`.
- name: Set ops user password
ansible.builtin.user:
name: "{{ baseline_ops_user }}"
password: "{{ ops_password | password_hash('sha512') }}"
no_log: true
- name: Configure ops sudo
ansible.builtin.copy:
content: "{{ baseline_ops_user }} ALL=(ALL) {{ baseline_ops_sudo }}\n"
dest: "/etc/sudoers.d/{{ baseline_ops_user }}"
mode: '0440'
validate: "visudo -cf %s"
@@ -0,0 +1,10 @@
127.0.0.1 localhost
# Ceph cluster nodes
{% for host in groups['ceph_nodes'] %}
{{ hostvars[host]['bond_ip'] }} {{ hostvars[host]['hostname_short'] }}.{{ cluster_domain }} {{ hostvars[host]['hostname_short'] }}
{% endfor %}
::1 localhost ip6-localhost ip6-loopback
ff02::1 ip6-allnodes
ff02::2 ip6-allrouters
@@ -0,0 +1,22 @@
---
# Ceph release train. Upstream publishes per-release apt trees at
# download.ceph.com/debian-<release>/, and each tree has subdirectories
# per Debian codename. prerequisites.yml pins the codename to 'bookworm'
# in the sources.list entry — when upgrading the base OS to Trixie,
# flip the codename there. This split keeps the Ceph release and the
# Debian release independently versionable.
ceph_release: tentacle
ceph_release_version: "20.2.*" # pin to patch range; set "20.2.1" to pin exactly
ceph_repo_url: "https://download.ceph.com/debian-{{ ceph_release }}/"
ceph_repo_key_url: "https://download.ceph.com/keys/release.asc"
cephadm_install_method: repo # 'repo' or 'curl'
# Initial PG counts per pool. The autoscaler is left on and will adjust
# over time, but starting at pg_num=1 (the ceph default) causes PG
# splitting under load which tanks performance during the first fill.
# These values are sized for ~36 OSDs. Scale proportionally for larger
# clusters.
ceph_pg_init_data: 128 # EC data pool — bulk of I/O
ceph_pg_init_index: 16 # bucket index
ceph_pg_init_non_ec: 16 # multipart uploads
ceph_pg_init_meta: 8 # .rgw.root, .meta, .log, .control
@@ -0,0 +1,14 @@
---
- name: Update apt cache
ansible.builtin.apt:
update_cache: true
- name: Restart chronyd
ansible.builtin.systemd:
name: chronyd
state: restarted
- name: Restart sshd
ansible.builtin.systemd:
name: ssh
state: restarted
@@ -0,0 +1,18 @@
---
# Advisory: os_tuning and hardware_tuning should run before this role.
# Not declared as hard dependencies because the tuning playbooks are
# separate operational steps, not automatic prerequisites.
#
# Execution order:
# 1. provision_host (bare metal → Debian)
# 2. os_tuning (sysctl, ulimits)
# 3. hardware_tuning (I/O scheduler, readahead)
# 4. ceph_deploy (this role)
dependencies: []
galaxy_info:
author: FUTO
license: AGPL-3.0-only
role_name: ceph_deploy
description: Deploy Ceph Tentacle cluster via cephadm
min_ansible_version: "2.19"
@@ -0,0 +1,50 @@
---
dependency:
name: galaxy
options:
requirements-file: ${MOLECULE_PROJECT_DIRECTORY}/../../requirements.yml
driver:
name: default
platforms:
- name: molecule-ceph-test
image: debian:bookworm
pre_build_image: true
provisioner:
name: ansible
inventory:
hosts:
all:
hosts:
molecule-ceph-test:
hostname_short: molecule-ceph-test
bond_ip: 10.10.10.99
sas_path_prefix: "pci-0000:02:00.0-sas-exp0x500056b300000000"
ssd1_phy: 12
ssd2_phy: 13
ceph_db_vg1: ceph-db-rear12
ceph_db_vg2: ceph-db-rear13
ceph_hdd_osds:
- path_phy: phy0
db: ceph-db-rear12/db-slot0
ceph_ssd_osds:
- path_phy: phy12
partition: 6
children:
ceph_nodes:
hosts:
molecule-ceph-test:
ceph_bootstrap:
hosts:
molecule-ceph-test:
ceph_join:
hosts: {}
config_options:
defaults:
gathering: implicit
# Only verify templates render — don't converge (needs real hardware)
verifier:
name: ansible
@@ -0,0 +1,55 @@
---
# Molecule verify — validates that templates render without errors.
# Does NOT require a running Ceph cluster.
- name: Verify ceph_deploy template rendering
hosts: all
gather_facts: false
vars:
cluster_domain: dev.austin.int.futo.cloud
public_network: 10.10.10.0/24
cluster_network: 10.10.10.0/24
ceph_rgw_realm: sietch
ceph_rgw_zone: dev-z1
ceph_rgw_port: 443
ceph_rgw_ssl: true
rgw_ssl_cert_combined_pem: "MOCK_CERT_CONTENT"
ceph_rgw_dns_name: s3.dev.austin.int.futo.cloud
tasks:
- name: Render hosts.j2
ansible.builtin.template:
src: "../../templates/hosts.j2"
dest: /tmp/molecule-hosts.txt
mode: '0644'
- name: Verify hosts.j2 has no 127.0.1.1
ansible.builtin.command: grep -c '127.0.1.1' /tmp/molecule-hosts.txt
register: hosts_check
failed_when: hosts_check.rc == 0
changed_when: false
- name: Render cluster-spec.yml.j2
ansible.builtin.template:
src: "../../templates/cluster-spec.yml.j2"
dest: /tmp/molecule-cluster-spec.yml
mode: '0644'
- name: Render rgw-spec.yaml.j2
ansible.builtin.template:
src: "../../templates/rgw-spec.yaml.j2"
dest: /tmp/molecule-rgw-spec.yaml
mode: '0644'
- name: Verify rgw-spec uses hostname_short
ansible.builtin.shell: |
set -o pipefail
grep 'molecule-ceph-test' /tmp/molecule-rgw-spec.yaml
args:
executable: /bin/bash
changed_when: false
- name: Report success
ansible.builtin.debug:
msg: "All templates rendered and validated successfully"
@@ -0,0 +1,70 @@
---
# Phase 2: Bootstrap cluster (first node only)
- name: Check if Ceph cluster already exists
ansible.builtin.stat:
path: /etc/ceph/ceph.conf
register: ceph_conf
when: inventory_hostname in groups['ceph_bootstrap']
- name: Write initial ceph config
ansible.builtin.template:
src: cluster-spec.yml.j2
dest: /tmp/ceph-initial.conf
mode: '0644'
when:
- inventory_hostname in groups['ceph_bootstrap']
- not (ceph_conf.stat.exists | default(false))
- name: Write dashboard password to temp file
ansible.builtin.copy:
content: "{{ ceph_dashboard_password }}"
dest: /tmp/.ceph-dashboard-pw
owner: root
group: root
mode: '0600'
when:
- inventory_hostname in groups['ceph_bootstrap']
- not (ceph_conf.stat.exists | default(false))
no_log: true
- name: Bootstrap Ceph cluster
ansible.builtin.shell: |
set -o pipefail
cephadm bootstrap \
--mon-ip {{ bond_ip }} \
--initial-dashboard-user {{ ceph_dashboard_user }} \
--initial-dashboard-password "$(cat /tmp/.ceph-dashboard-pw)" \
--dashboard-password-noupdate \
--allow-fqdn-hostname \
--config /tmp/ceph-initial.conf \
2>&1 | tee /var/log/ceph-bootstrap.log
RC=$?
rm -f /tmp/.ceph-dashboard-pw
exit $RC
args:
executable: /bin/bash
register: bootstrap_result
when:
- inventory_hostname in groups['ceph_bootstrap']
- not (ceph_conf.stat.exists | default(false))
changed_when: bootstrap_result.rc | default(1) == 0
no_log: true
- name: Clean up dashboard password file on skip
ansible.builtin.file:
path: /tmp/.ceph-dashboard-pw
state: absent
when: inventory_hostname in groups['ceph_bootstrap']
timeout: 600
- name: Show bootstrap result
ansible.builtin.debug:
msg: "{{ bootstrap_result.stdout_lines[-15:] | default(['Already bootstrapped']) }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- bootstrap_result is not skipped
# cephadm automatically distributes ceph.conf, admin keyring, and
# its SSH key to managed hosts during `ceph orch host add`. No
# manual capture or distribution needed.
@@ -0,0 +1,36 @@
---
# Phase 5.5: Create CRUSH device-class rules
# Ensures pools can target HDD-only or SSD-only OSDs
- name: Check existing CRUSH rules
ansible.builtin.command: ceph osd crush rule ls
register: crush_rules
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Create replicated_hdd CRUSH rule
ansible.builtin.command: >
ceph osd crush rule create-replicated replicated_hdd default host hdd
when:
- inventory_hostname in groups['ceph_bootstrap']
- "'replicated_hdd' not in crush_rules.stdout"
changed_when: true
- name: Create replicated_ssd CRUSH rule
ansible.builtin.command: >
ceph osd crush rule create-replicated replicated_ssd default host ssd
when:
- inventory_hostname in groups['ceph_bootstrap']
- "'replicated_ssd' not in crush_rules.stdout"
changed_when: true
- name: Show CRUSH rules
ansible.builtin.command: ceph osd crush rule ls
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
register: final_rules
- name: Display CRUSH rules
ansible.builtin.debug:
msg: "{{ final_rules.stdout_lines }}"
when: inventory_hostname in groups['ceph_bootstrap']
@@ -0,0 +1,106 @@
---
# Phase 3: Add nodes to cluster (from bootstrap node)
#
# Only attempts to add hosts that are in the current play (--limit aware).
# This prevents failures when nodes are defined in inventory
# but not yet available.
- name: Build list of joinable hosts
ansible.builtin.set_fact:
joinable_hosts: >-
{{ groups['ceph_join'] | default([])
| intersect(ansible_play_hosts)
| list }}
when: inventory_hostname in groups['ceph_bootstrap']
- name: Verify cephadm SSH to join nodes
ansible.builtin.shell: |
set -euo pipefail
HOST="{{ hostvars[item]['hostname_short'] }}"
ADDR="{{ hostvars[item]['bond_ip'] }}"
echo "Testing SSH to $HOST ($ADDR)..."
ceph cephadm check-host "$HOST" "$ADDR" 2>&1 || {
echo "WARN: ceph cephadm check-host failed, testing raw SSH..."
cephadm shell -- ssh -o StrictHostKeyChecking=no -o ConnectTimeout=10 \
"root@${ADDR}" "cephadm check-host --expect-hostname $HOST" 2>&1
}
args:
executable: /bin/bash
loop: "{{ joinable_hosts | default([]) }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- joinable_hosts | default([]) | length > 0
loop_control:
label: "{{ hostvars[item]['hostname_short'] }}"
changed_when: false
failed_when: false
register: ssh_check_results
- name: Show SSH check results
ansible.builtin.debug:
msg: "{{ item.stdout_lines | default([]) }}"
loop: "{{ ssh_check_results.results | default([]) }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- item.stdout is defined
loop_control:
label: "{{ item.item | default('unknown') }}"
- name: Add remaining nodes to Ceph cluster
ansible.builtin.shell: |
set -o pipefail
# Check if host already added
if ceph orch host ls --format json |
grep -q '"{{ hostvars[item]['hostname_short'] }}"'; then
echo "Host {{ hostvars[item]['hostname_short'] }} already in cluster"
else
echo "Adding {{ hostvars[item]['hostname_short'] }} ({{ hostvars[item]['bond_ip'] }})"
ceph orch host add \
{{ hostvars[item]['hostname_short'] }} \
{{ hostvars[item]['bond_ip'] }} 2>&1
fi
args:
executable: /bin/bash
loop: "{{ joinable_hosts | default([]) }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- joinable_hosts | default([]) | length > 0
register: host_add_results
changed_when: "'Adding' in (host_add_results.stdout | default(''))"
- name: Show host add results
ansible.builtin.debug:
msg: "{{ item.stdout }}"
loop: "{{ host_add_results.results | default([]) }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- item.stdout is defined
loop_control:
label: "{{ item.item | default('unknown') }}"
# cephadm automatically syncs ceph.conf and admin keyring to managed
# hosts after `ceph orch host add`. No manual distribution needed.
- name: Wait for new hosts to be online
ansible.builtin.shell: |
set -o pipefail
for i in $(seq 1 30); do
online=$(ceph orch host ls --format json | python3 -c "
import sys, json
hosts = json.load(sys.stdin)
print(len([h for h in hosts if h.get('status','') == '']))
" 2>/dev/null || echo 0)
expected={{ (joinable_hosts | default([]) | length) + 1 }}
echo "Hosts online: $online/$expected"
if [ "$online" -ge "$expected" ]; then
exit 0
fi
sleep 10
done
echo "WARNING: Not all hosts online yet"
args:
executable: /bin/bash
when:
- inventory_hostname in groups['ceph_bootstrap']
- joinable_hosts | default([]) | length > 0
changed_when: false
@@ -0,0 +1,110 @@
---
# Phase 4.5: Ensure Ceph block.db LVM is set up on each node (sietch-shape only)
#
# This task is sietch-shape-specific: it assumes dual SAS-attached SSDs each
# with partition 5 → its own VG, with 6 db-slot LVs per VG (12 total). It
# ensures PVs, VGs, and LVs exist. Idempotent — skips if already present.
# Provides a recovery path if the LVM was destroyed by `mise run destroy`.
#
# Painbox-shape (Hetzner SX295, NVMe RAID-1 → single vg0 with all LVs created
# by installimage post-install) skips the whole block — its host_vars don't
# define `sas_path_prefix` so the gate below short-circuits.
#
# Required host_vars (sietch-shape only):
# sas_path_prefix — by-path prefix for SSD discovery (e.g., pci-0000:02:00.0-sas-exp...)
# ssd1_phy, ssd2_phy — PHY slot numbers for the two SSDs
# ceph_db_vg1, ceph_db_vg2 — VG names for SSD1/SSD2 partition 5
#
# VG mapping: ceph_db_vg1 on SSD1 partition 5, ceph_db_vg2 on SSD2 partition 5
# LV naming: db-slot0..5 on VG1, db-slot6..11 on VG2 (one per HDD OSD)
- name: Skip in-role LVM setup (no sas_path_prefix — externally-managed LVM)
ansible.builtin.debug:
msg: >-
LVM is externally managed on this host (e.g., Hetzner installimage
post-install creates vg0 + db-slots on a RAID-1 NVMe). Skipping the
sietch-shape dual-SSD-VG setup. If LVM is missing on this host,
reprovision via the cluster's installimage flow.
when: sas_path_prefix is undefined
- name: Set up LVM (sietch-shape dual-SSD-VG topology)
when: sas_path_prefix is defined
block:
- name: Resolve SSD1 device path
ansible.builtin.command: >
readlink -f /dev/disk/by-path/{{ sas_path_prefix }}-phy{{ ssd1_phy }}-lun-0
register: ssd1_dev
changed_when: false
- name: Resolve SSD2 device path
ansible.builtin.command: >
readlink -f /dev/disk/by-path/{{ sas_path_prefix }}-phy{{ ssd2_phy }}-lun-0
register: ssd2_dev
changed_when: false
- name: Check if VG1 exists
ansible.builtin.shell: vgs {{ ceph_db_vg1 }} 2>/dev/null
register: vg1_check
changed_when: false
failed_when: false
- name: Check if VG2 exists
ansible.builtin.shell: vgs {{ ceph_db_vg2 }} 2>/dev/null
register: vg2_check
changed_when: false
failed_when: false
- name: Create PV and VG on SSD1 partition 5
ansible.builtin.shell: |
PART="{{ ssd1_dev.stdout }}5"
wipefs -af "$PART"
pvcreate -f "$PART" && vgcreate {{ ceph_db_vg1 }} "$PART"
when: vg1_check.rc != 0
changed_when: true
- name: Create PV and VG on SSD2 partition 5
ansible.builtin.shell: |
PART="{{ ssd2_dev.stdout }}5"
wipefs -af "$PART"
pvcreate -f "$PART" && vgcreate {{ ceph_db_vg2 }} "$PART"
when: vg2_check.rc != 0
changed_when: true
- name: Create block.db LVs on VG1 (slots 0-4 at 240G, slot 5 gets remainder)
ansible.builtin.shell: |
if lvs {{ ceph_db_vg1 }}/db-slot{{ item }} 2>/dev/null; then
echo "db-slot{{ item }} already exists"
else
{% if item < 5 %}
lvcreate --yes -Wy -L {{ ceph_db_lv_size }} -n db-slot{{ item }} {{ ceph_db_vg1 }}
{% else %}
lvcreate --yes -Wy -l 100%FREE -n db-slot{{ item }} {{ ceph_db_vg1 }}
{% endif %}
fi
loop: [0, 1, 2, 3, 4, 5]
register: vg1_lv_results
changed_when: "'already exists' not in (vg1_lv_results.stdout | default(''))"
- name: Create block.db LVs on VG2 (slots 6-10 at 240G, slot 11 gets remainder)
ansible.builtin.shell: |
if lvs {{ ceph_db_vg2 }}/db-slot{{ item }} 2>/dev/null; then
echo "db-slot{{ item }} already exists"
else
{% if item < 11 %}
lvcreate --yes -Wy -L {{ ceph_db_lv_size }} -n db-slot{{ item }} {{ ceph_db_vg2 }}
{% else %}
lvcreate --yes -Wy -l 100%FREE -n db-slot{{ item }} {{ ceph_db_vg2 }}
{% endif %}
fi
loop: [6, 7, 8, 9, 10, 11]
register: vg2_lv_results
changed_when: "'already exists' not in (vg2_lv_results.stdout | default(''))"
- name: Verify LVM setup
ansible.builtin.command: lvs -o lv_name,vg_name,lv_size --noheadings
register: lvm_state
changed_when: false
- name: Show LVM state
ansible.builtin.debug:
msg: "{{ lvm_state.stdout_lines }}"
@@ -0,0 +1,43 @@
---
# Ceph Tentacle deployment via cephadm
# Phases: prerequisites -> bootstrap -> join nodes -> placement -> OSDs -> verify
- name: Phase 1 - Prerequisites
ansible.builtin.import_tasks: prerequisites.yml
tags: [prerequisites]
- name: Phase 2 - Bootstrap cluster
ansible.builtin.import_tasks: bootstrap.yml
tags: [bootstrap]
- name: Phase 3 - Join nodes
ansible.builtin.import_tasks: join.yml
tags: [join]
- name: Phase 4 - MON/MGR placement
ansible.builtin.import_tasks: placement.yml
tags: [placement]
- name: Phase 4.5 - Ensure block.db LVM exists
ansible.builtin.import_tasks: lvm-setup.yml
tags: [lvm, osds]
- name: Phase 5 - Create OSDs
ansible.builtin.import_tasks: osds.yml
tags: [osds]
- name: Phase 5.5 - CRUSH device-class rules
ansible.builtin.import_tasks: crush-rules.yml
tags: [crush, osds]
- name: Phase 5.75 - RadosGW (S3 Object Gateway)
ansible.builtin.import_tasks: rgw.yml
tags: [rgw]
- name: Phase 5.8 - Monitoring stack
ansible.builtin.import_tasks: monitoring.yml
tags: [monitoring]
- name: Phase 6 - Verify cluster
ansible.builtin.import_tasks: verify.yml
tags: [verify]
@@ -0,0 +1,138 @@
---
# Phase 5.8: Monitoring stack — dashboard integration
#
# cephadm auto-deploys the full monitoring stack (node-exporter,
# ceph-exporter, prometheus, alertmanager, grafana) during bootstrap.
# This task file only:
# 1. Ensures the prometheus mgr module is enabled
# 2. Waits for all monitoring daemons to come up
# 3. Configures dashboard integration URLs and credentials
# --- Step 1: Enable prometheus mgr module ---
- name: Check prometheus mgr module status
ansible.builtin.shell: |
set -o pipefail
ceph mgr module ls --format json 2>/dev/null | python3 -c "
import sys, json
data = json.load(sys.stdin)
enabled = data.get('enabled_modules', [])
print('enabled' if 'prometheus' in enabled else 'disabled')
"
args:
executable: /bin/bash
register: prometheus_module
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Enable prometheus mgr module
ansible.builtin.command: ceph mgr module enable prometheus
when:
- inventory_hostname in groups['ceph_bootstrap']
- "'disabled' in prometheus_module.stdout"
changed_when: true
# --- Step 2: Wait for cephadm to deploy monitoring ---
- name: Wait for monitoring daemons to start
ansible.builtin.shell: |
python3 -c "
import json, subprocess
services = ['node-exporter', 'ceph-exporter', 'prometheus',
'alertmanager', 'grafana']
for svc in services:
out = subprocess.check_output(
['ceph', 'orch', 'ls', '--service-type', svc,
'--format', 'json'],
stderr=subprocess.DEVNULL)
data = json.loads(out)
running = sum(
s.get('status', {}).get('running', 0) for s in data)
if running == 0:
print(f'waiting:{svc}')
exit(1)
print('all_running')
" 2>/dev/null
args:
executable: /bin/bash
register: monitoring_wait
until: "'all_running' in monitoring_wait.stdout"
retries: 30
delay: 10
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
# --- Step 3: Configure dashboard integrations ---
- name: Configure dashboard Prometheus URL
ansible.builtin.command: >
ceph dashboard set-prometheus-api-host
http://{{ hostvars[groups['ceph_bootstrap'][0]]['bond_ip'] }}:{{ ceph_prometheus_port }}
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Configure dashboard Alertmanager URL
ansible.builtin.command: >
ceph dashboard set-alertmanager-api-host
http://{{ hostvars[groups['ceph_bootstrap'][0]]['bond_ip'] }}:{{ ceph_alertmanager_port }}
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Configure dashboard Grafana URL
ansible.builtin.command: >
ceph dashboard set-grafana-api-url
https://{{ hostvars[groups['ceph_bootstrap'][0]]['bond_ip'] }}:{{ ceph_grafana_port }}
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Disable Grafana SSL cert verification in dashboard
ansible.builtin.command: >
ceph dashboard set-grafana-api-ssl-verify false
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
# --- Step 4: Set Grafana admin credentials ---
- name: Set Grafana admin user
ansible.builtin.command: >
ceph dashboard set-grafana-api-username {{ ceph_grafana_admin_user }}
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Set Grafana admin password
ansible.builtin.shell: |
set -o pipefail
echo '{{ ceph_grafana_admin_password }}' \
| ceph dashboard set-grafana-api-password -i -
args:
executable: /bin/bash
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
no_log: true
# --- Step 5: Display monitoring endpoints ---
- name: Show monitoring status
ansible.builtin.shell: |
set -o pipefail
echo "=== Monitoring Services ==="
ceph orch ls --service-type prometheus
ceph orch ls --service-type grafana
ceph orch ls --service-type alertmanager
ceph orch ls --service-type node-exporter
ceph orch ls --service-type ceph-exporter
echo ""
echo "=== Endpoints ==="
echo "Prometheus: http://{{ bond_ip }}:{{ ceph_prometheus_port }}"
echo "Grafana: https://{{ bond_ip }}:{{ ceph_grafana_port }}"
echo "Alertmanager: http://{{ bond_ip }}:{{ ceph_alertmanager_port }}"
args:
executable: /bin/bash
register: monitoring_status
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Display monitoring status
ansible.builtin.debug:
msg: "{{ monitoring_status.stdout_lines }}"
when: inventory_hostname in groups['ceph_bootstrap']
@@ -0,0 +1,183 @@
---
# Phase 5: Create OSDs via cephadm orch service spec
#
# Renders a multi-document OSD service spec (one document per host, see
# templates/osd-spec.yml.j2) and applies via `ceph orch apply osd`. Cephadm
# handles per-disk LUKS provisioning, LVM creation, daemon deployment, and
# CRUSH placement internally — the role no longer iterates disks.
#
# Hardware-shape independence: the template's Jinja conditionals handle
# both sietch (SAS expander + dual-SSD-VG) and painbox (PCI-ATA + single
# NVMe-RAID VG + LV-backed SSD OSD) without per-shape branching here.
#
# Required host_vars (per host, in inventories/<cluster>/host_vars/<host>.yml):
# ceph_hdd_osds — list of {path_phy, db} mappings per HDD
# ceph_ssd_osds — list of {path_phy + partition} (sietch) or {lv} (painbox)
# sas_path_prefix — REQUIRED for sietch; OMITTED for painbox-shape hosts
#
# Idempotent: re-applying the same spec is a no-op when OSDs match. New
# disks (e.g. populating an empty bay later) are picked up automatically.
- name: Render OSD service spec on bootstrap node
ansible.builtin.template:
src: osd-spec.yml.j2
dest: /etc/ceph/osd-spec.yml
mode: '0644'
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
run_once: true
register: osd_spec_render
- name: Show rendered OSD spec (for verification)
ansible.builtin.command: cat /etc/ceph/osd-spec.yml
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
run_once: true
changed_when: false
register: osd_spec_content
- name: Display rendered OSD spec
ansible.builtin.debug:
msg: "{{ osd_spec_content.stdout_lines }}"
run_once: true
- name: Apply OSD service spec via cephadm orch
ansible.builtin.command: ceph orch apply osd -i /etc/ceph/osd-spec.yml
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
run_once: true
register: orch_apply
changed_when: >-
'Scheduled' in (orch_apply.stdout | default(''))
or 'created' in (orch_apply.stdout | default('') | lower)
- name: Wait for cephadm to provision all expected OSDs
ansible.builtin.shell: |
set -o pipefail
EXPECTED={{
groups['ceph_nodes']
| map('extract', hostvars)
| map(attribute='ceph_hdd_osds', default=[])
| map('length') | sum
+
groups['ceph_nodes']
| map('extract', hostvars)
| map(attribute='ceph_ssd_osds', default=[])
| map('length') | sum
}}
echo "Expecting $EXPECTED OSDs total across the cluster"
for i in $(seq 1 90); do
ACTUAL=$(ceph osd stat --format json 2>/dev/null \
| python3 -c "import sys, json; print(json.load(sys.stdin).get('num_osds', 0))" \
2>/dev/null || echo 0)
echo " attempt $i: $ACTUAL/$EXPECTED OSDs created"
if [ "$ACTUAL" -ge "$EXPECTED" ]; then
echo "All expected OSDs provisioned"
exit 0
fi
sleep 10
done
echo "WARNING: only $ACTUAL of $EXPECTED OSDs provisioned after 15 minutes"
exit 1
args:
executable: /bin/bash
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
run_once: true
changed_when: false
timeout: 960
- name: Wait for OSDs to come up
ansible.builtin.shell: |
set -o pipefail
for i in $(seq 1 90); do
up=$(ceph osd stat --format json 2>/dev/null |
python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('num_up_osds',0))" \
2>/dev/null || echo 0)
total=$(ceph osd stat --format json 2>/dev/null |
python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('num_osds',0))" \
2>/dev/null || echo 0)
echo "OSDs: $up/$total up"
if [ "$up" -eq "$total" ] && [ "$total" -gt 0 ]; then
echo "All OSDs up"; exit 0
fi
sleep 10
done
echo "WARNING: Not all OSDs up after 15 minutes"
exit 1
args:
executable: /bin/bash
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
run_once: true
changed_when: false
timeout: 960
# Safety net: if any OSD comes up with reweight=0 (rare with orch apply,
# but possible if a noin flag was set externally before this run), fix it.
- name: Fix any OSDs stuck at reweight 0
ansible.builtin.shell: |
set -o pipefail
FIXED=0
for osd_id in $(ceph osd tree --format json 2>/dev/null | python3 -c "
import sys, json
tree = json.load(sys.stdin)
for node in tree.get('nodes', []):
if node.get('type') == 'osd' and node.get('status') == 'up' \
and node.get('reweight', 1) == 0:
print(node['id'])
"); do
echo "Reweighting osd.$osd_id from 0 to 1.0"
ceph osd reweight "$osd_id" 1.0
FIXED=$((FIXED + 1))
done
if [ "$FIXED" -eq 0 ]; then
echo "All up OSDs have proper reweight"
else
echo "Fixed $FIXED OSDs with reweight=0"
fi
args:
executable: /bin/bash
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
run_once: true
register: reweight_fix
changed_when: "'Fixed' in reweight_fix.stdout"
- name: Show OSD tree
ansible.builtin.command: ceph osd tree
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
run_once: true
changed_when: false
register: osd_tree
- name: Display OSD tree
ansible.builtin.debug:
msg: "{{ osd_tree.stdout_lines }}"
run_once: true
- name: Verify dmcrypt keys stored in MONs
ansible.builtin.shell: |
set -o pipefail
ceph config-key dump 2>/dev/null | grep -c dm-crypt
args:
executable: /bin/bash
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
run_once: true
changed_when: false
register: dmcrypt_keys
- name: Show dmcrypt key count
ansible.builtin.debug:
msg: >-
{{ dmcrypt_keys.stdout }} dm-crypt keys stored
in MON config-key database
run_once: true
# Defensive: clear the noin flag if it's set. The current spec-apply flow
# doesn't set noin (cephadm orch handles backfill rollout gracefully), but
# a prior failed run of this role's older imperative OSD-creation flow may
# have set it and not unset it. Unset is idempotent — no-op when already off.
- name: Defensive — clear noin flag if set
ansible.builtin.command: ceph osd unset noin
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
run_once: true
register: noin_unset
changed_when: "'noin is unset' in (noin_unset.stdout | default(''))"
failed_when: false
@@ -0,0 +1,91 @@
---
# Phase 4: Configure MON and MGR placement
#
# MON placement strategy:
# 1. If ceph_mon group is defined and non-empty: use those hosts (explicit)
# 2. Else if <=2 active hosts: 1 MON (bootstrap only, avoids 2-MON fragility)
# 3. Else: MON on all active hosts
#
# MGR always deploys on all active hosts (standbys are harmless).
- name: Determine active cluster hosts
ansible.builtin.shell: |
set -o pipefail
ceph orch host ls --format json | python3 -c "
import sys, json
hosts = json.load(sys.stdin)
print(','.join(h['hostname'] for h in hosts if h.get('status','') == ''))
"
args:
executable: /bin/bash
register: active_hosts
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Calculate MON placement
ansible.builtin.set_fact:
mon_hosts: >-
{%- if groups['ceph_mon'] | default([]) | length > 0 -%}
{{ groups['ceph_mon']
| map('extract', hostvars, 'hostname_short')
| join(',') -}}
{%- elif active_hosts.stdout.split(',') | length <= 2 -%}
{{ hostvars[groups['ceph_bootstrap'][0]]['hostname_short'] -}}
{%- else -%}
{{ active_hosts.stdout -}}
{%- endif -%}
when:
- inventory_hostname in groups['ceph_bootstrap']
- active_hosts.stdout | default('') | length > 0
- name: Show placement decision
ansible.builtin.debug:
msg: >-
Active hosts: {{ active_hosts.stdout }}.
MON placement: {{ mon_hosts }}
({{ 'explicit ceph_mon group'
if groups['ceph_mon'] | default([]) | length > 0
else ('single MON — avoids 2-mon quorum fragility'
if active_hosts.stdout.split(',') | length <= 2
else 'MON on all hosts') }}).
MGR placement: {{ active_hosts.stdout }} (all hosts).
when:
- inventory_hostname in groups['ceph_bootstrap']
- mon_hosts is defined
- name: Deploy MON on calculated hosts
ansible.builtin.command: >
ceph orch apply mon --placement="{{ mon_hosts }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- mon_hosts | default('') | length > 0
changed_when: false
- name: Deploy MGR on all active hosts
ansible.builtin.command: >
ceph orch apply mgr --placement="{{ active_hosts.stdout }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- active_hosts.stdout | default('') | length > 0
changed_when: false
- name: Wait for MONs to be ready
ansible.builtin.shell: |
set -o pipefail
for i in $(seq 1 30); do
mon_count=$(ceph mon stat --format json 2>/dev/null |
python3 -c "import sys,json; print(json.load(sys.stdin).get('num_mons',0))" \
2>/dev/null || echo 0)
expected={{ mon_hosts.split(',') | length }}
if [ "$mon_count" -ge "$expected" ]; then
echo "All MONs ready ($mon_count/$expected)"
exit 0
fi
echo "Waiting for MONs... ($mon_count/$expected)"
sleep 10
done
echo "WARNING: Not all MONs ready, continuing"
args:
executable: /bin/bash
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
@@ -0,0 +1,56 @@
---
# Phase 1: Prerequisites (all nodes)
# /etc/hosts is managed by the baseline role.
- name: Ensure keyrings directory exists
ansible.builtin.file:
path: /etc/apt/keyrings
state: directory
mode: '0755'
- name: Add Ceph release GPG key
ansible.builtin.shell: |
set -o pipefail
curl -fsSL {{ ceph_repo_key_url }} | gpg --dearmor -o /etc/apt/keyrings/ceph.gpg
args:
creates: /etc/apt/keyrings/ceph.gpg
executable: /bin/bash
- name: Add Ceph Tentacle repository
ansible.builtin.copy:
content: "deb [signed-by=/etc/apt/keyrings/ceph.gpg] {{ ceph_repo_url }} bookworm main\n"
dest: /etc/apt/sources.list.d/ceph.list
mode: '0644'
- name: Update apt cache
# Always refresh — no cache_valid_time. Hetzner installimage runs apt
# during install, leaving the cache "fresh" (~minutes old) but without the
# Ceph repo's Packages file. cache_valid_time=3600 here would skip refresh
# and apt install would only see Debian's older cephadm (16.x) → version
# pin fails. Always-refresh is the safe default for any new repo addition.
ansible.builtin.apt:
update_cache: true
# Podman ecosystem, diagnostic tools, and service enablement are
# handled by the baseline role. This task installs only Ceph-specific
# packages that baseline doesn't cover.
- name: Install cephadm and Ceph-specific dependencies
ansible.builtin.apt:
name:
- "cephadm={{ ceph_release_version }}"
- "ceph-common={{ ceph_release_version }}"
- python3-asyncssh
state: present
# dbus, chrony, and podman.socket are enabled by the baseline role.
- name: Disable sntrup761 post-quantum kex (hangs with asyncssh/cephadm)
ansible.builtin.copy:
# yamllint disable-line rule:line-length
content: >-
KexAlgorithms curve25519-sha256,curve25519-sha256@libssh.org,ecdh-sha2-nistp256,ecdh-sha2-nistp384,ecdh-sha2-nistp521,diffie-hellman-group-exchange-sha256,diffie-hellman-group16-sha512,diffie-hellman-group18-sha512,diffie-hellman-group14-sha256
dest: /etc/ssh/sshd_config.d/no-sntrup.conf
mode: '0644'
notify: Restart sshd
@@ -0,0 +1,658 @@
---
# Phase 5.75: RadosGW (S3 Object Gateway)
# Deploys RGW with EC data pool, realm/zonegroup/zone, and S3 service user
# --- Step 1: EC profile ---
- name: Check existing EC profiles
ansible.builtin.command: ceph osd erasure-code-profile ls
register: ec_profiles
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Create EC profile {{ ceph_rgw_ec_profile }}
ansible.builtin.command: >
ceph osd erasure-code-profile set {{ ceph_rgw_ec_profile }}
k={{ ceph_rgw_ec_k }} m={{ ceph_rgw_ec_m }}
crush-failure-domain={{ ceph_rgw_ec_failure_domain }}
crush-device-class={{ ceph_rgw_ec_device_class }}
when:
- inventory_hostname in groups['ceph_bootstrap']
- "ceph_rgw_ec_profile not in ec_profiles.stdout"
changed_when: true
# --- Step 2-3: EC data pool ---
- name: Check existing pools
ansible.builtin.command: ceph osd pool ls
register: pool_list
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Create EC data pool {{ ceph_rgw_data_pool }}
ansible.builtin.command: >
ceph osd pool create {{ ceph_rgw_data_pool }} erasure {{ ceph_rgw_ec_profile }}
register: create_data_pool
when:
- inventory_hostname in groups['ceph_bootstrap']
- "ceph_rgw_data_pool not in pool_list.stdout"
changed_when: true
- name: Enable EC overwrites on data pool
ansible.builtin.command: >
ceph osd pool set {{ ceph_rgw_data_pool }} allow_ec_overwrites true
when:
- inventory_hostname in groups['ceph_bootstrap']
- create_data_pool is changed
changed_when: true
# --- Step 4: Replicated index pool ---
- name: Create replicated index pool {{ ceph_rgw_index_pool }}
ansible.builtin.command: >
ceph osd pool create {{ ceph_rgw_index_pool }} replicated
when:
- inventory_hostname in groups['ceph_bootstrap']
- "ceph_rgw_index_pool not in pool_list.stdout"
changed_when: true
- name: Set index pool replication parameters
ansible.builtin.shell: |
set -e
ceph osd pool set {{ ceph_rgw_index_pool }} size {{ ceph_rgw_replicated_size }}
ceph osd pool set {{ ceph_rgw_index_pool }} min_size {{ ceph_rgw_replicated_min_size }}
args:
executable: /bin/bash
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
# --- Step 5: Replicated non-EC pool (multipart uploads) ---
- name: Create replicated non-EC pool {{ ceph_rgw_extra_pool }}
ansible.builtin.command: >
ceph osd pool create {{ ceph_rgw_extra_pool }} replicated
when:
- inventory_hostname in groups['ceph_bootstrap']
- "ceph_rgw_extra_pool not in pool_list.stdout"
changed_when: true
- name: Set non-EC pool replication parameters
ansible.builtin.shell: |
set -e
ceph osd pool set {{ ceph_rgw_extra_pool }} size {{ ceph_rgw_replicated_size }}
ceph osd pool set {{ ceph_rgw_extra_pool }} min_size {{ ceph_rgw_replicated_min_size }}
args:
executable: /bin/bash
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
# --- Step 6: Pool application tags ---
- name: Check pool application tags
ansible.builtin.command: "ceph osd pool application get {{ item }}"
loop:
- "{{ ceph_rgw_data_pool }}"
- "{{ ceph_rgw_index_pool }}"
- "{{ ceph_rgw_extra_pool }}"
register: pool_apps
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
failed_when: false
- name: Enable rgw application on pools
ansible.builtin.command: "ceph osd pool application enable {{ item.item }} rgw"
loop: "{{ pool_apps.results }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- item is not skipped
- "'rgw' not in (item.stdout | default(''))"
changed_when: true
loop_control:
label: "{{ item.item | default('skipped') }}"
# --- Step 6.5: Initial PG counts ---
#
# Pools are created with pg_num=1 (ceph default). The autoscaler will
# eventually grow them, but splitting PGs under load causes latency
# spikes and throughput drops. Pre-sizing avoids this penalty during
# the first fill. Idempotent: only increases pg_num, never decreases.
- name: Set initial PG count on data pool
ansible.builtin.command: >
ceph osd pool set {{ ceph_rgw_data_pool }}
pg_num {{ ceph_pg_init_data }}
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_rgw_data_pool in pool_list.stdout
register: pg_data_result
changed_when: "'set' in pg_data_result.stdout | default('')"
failed_when:
- pg_data_result.rc != 0
- "'is not >= current' not in pg_data_result.stderr | default('')"
- name: Set initial PG count on index pool
ansible.builtin.command: >
ceph osd pool set {{ ceph_rgw_index_pool }}
pg_num {{ ceph_pg_init_index }}
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_rgw_index_pool in pool_list.stdout
register: pg_index_result
changed_when: "'set' in pg_index_result.stdout | default('')"
failed_when:
- pg_index_result.rc != 0
- "'is not >= current' not in pg_index_result.stderr | default('')"
- name: Set initial PG count on non-EC pool
ansible.builtin.command: >
ceph osd pool set {{ ceph_rgw_extra_pool }}
pg_num {{ ceph_pg_init_non_ec }}
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_rgw_extra_pool in pool_list.stdout
register: pg_extra_result
changed_when: "'set' in pg_extra_result.stdout | default('')"
failed_when:
- pg_extra_result.rc != 0
- "'is not >= current' not in pg_extra_result.stderr | default('')"
- name: Set initial PG count on RGW metadata pools
ansible.builtin.command: >
ceph osd pool set {{ item }} pg_num {{ ceph_pg_init_meta }}
loop:
- .rgw.root
- "{{ ceph_rgw_zone }}.rgw.log"
- "{{ ceph_rgw_zone }}.rgw.control"
- "{{ ceph_rgw_zone }}.rgw.meta"
when:
- inventory_hostname in groups['ceph_bootstrap']
# Skip pools that haven't materialized yet — RGW creates them
# lazily on first daemon startup. Mirrors the per-pool gate used
# by the data/index/extra pool tasks above.
- item in pool_list.stdout
register: pg_meta_result
changed_when: "'set' in pg_meta_result.stdout | default('')"
failed_when:
- pg_meta_result.rc is defined
- pg_meta_result.rc != 0
- "'is not >= current' not in pg_meta_result.stderr | default('')"
- name: Wait for PG peering to complete
ansible.builtin.shell: |
set -o pipefail
for i in $(seq 1 60); do
STATE=$(ceph pg stat --format json 2>/dev/null | python3 -c "
import sys, json
d = json.load(sys.stdin)
s = d.get('pg_summary', {}).get('num_pg_by_state', [])
non_clean = sum(x['num'] for x in s if 'active+clean' not in x['name'])
print(non_clean)
" 2>/dev/null || echo 999)
if [ "$STATE" = "0" ]; then
echo "All PGs active+clean"
exit 0
fi
echo "Waiting for PG peering... ($STATE PGs not clean)"
sleep 5
done
echo "WARNING: PGs still peering after 5 minutes"
args:
executable: /bin/bash
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
timeout: 330
# --- Step 7: Realm ---
- name: Check existing realms
ansible.builtin.command: radosgw-admin realm list
register: realm_list
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
failed_when: false
- name: Create realm {{ ceph_rgw_realm }}
ansible.builtin.command: >
radosgw-admin realm create --rgw-realm={{ ceph_rgw_realm }} --default
when:
- inventory_hostname in groups['ceph_bootstrap']
- "ceph_rgw_realm not in (realm_list.stdout | default(''))"
changed_when: true
- name: Ensure realm is default
ansible.builtin.command: >
radosgw-admin realm default --rgw-realm={{ ceph_rgw_realm }}
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
# --- Step 8: Zonegroup ---
- name: Check existing zonegroups
ansible.builtin.command: radosgw-admin zonegroup list
register: zonegroup_list
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
failed_when: false
- name: Create zonegroup {{ ceph_rgw_zonegroup }}
ansible.builtin.command: >
radosgw-admin zonegroup create
--rgw-realm={{ ceph_rgw_realm }}
--rgw-zonegroup={{ ceph_rgw_zonegroup }}
--endpoints={{ ceph_rgw_scheme }}://{{ ceph_rgw_dns_name }}:{{ ceph_rgw_port }}
--master --default
when:
- inventory_hostname in groups['ceph_bootstrap']
- "ceph_rgw_zonegroup not in (zonegroup_list.stdout | default(''))"
changed_when: true
- name: Set zonegroup api_name to {{ ceph_rgw_zonegroup_api_name }}
ansible.builtin.shell: |
set -e
python3 -c "
import json, subprocess, sys
zg = json.loads(subprocess.check_output(
['radosgw-admin', 'zonegroup', 'get',
'--rgw-zonegroup', '{{ ceph_rgw_zonegroup }}']))
if zg.get('api_name') == '{{ ceph_rgw_zonegroup_api_name }}':
sys.exit(0)
zg['api_name'] = '{{ ceph_rgw_zonegroup_api_name }}'
subprocess.run(
['radosgw-admin', 'zonegroup', 'set',
'--rgw-zonegroup', '{{ ceph_rgw_zonegroup }}'],
input=json.dumps(zg).encode(), check=True)
print('CHANGED')
"
args:
executable: /bin/bash
register: set_api_name
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: "'CHANGED' in (set_api_name.stdout | default(''))"
# --- Step 9: Zone ---
- name: Check existing zones
ansible.builtin.command: radosgw-admin zone list
register: zone_list
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
failed_when: false
- name: Create zone {{ ceph_rgw_zone }}
ansible.builtin.command: >
radosgw-admin zone create
--rgw-realm={{ ceph_rgw_realm }}
--rgw-zonegroup={{ ceph_rgw_zonegroup }}
--rgw-zone={{ ceph_rgw_zone }}
--endpoints={{ ceph_rgw_scheme }}://{{ ceph_rgw_dns_name }}:{{ ceph_rgw_port }}
--master --default
when:
- inventory_hostname in groups['ceph_bootstrap']
- "ceph_rgw_zone not in (zone_list.stdout | default(''))"
changed_when: true
# --- Step 10: Zone placement targets ---
- name: Configure zone placement targets for EC pools
ansible.builtin.shell: |
set -e
python3 -c "
import json, subprocess, sys
zone = json.loads(subprocess.check_output(
['radosgw-admin', 'zone', 'get', '--rgw-zone', '{{ ceph_rgw_zone }}']))
changed = False
for pt in zone.get('placement_targets', []):
if pt['name'] == 'default-placement':
if (pt.get('data_pool') != '{{ ceph_rgw_data_pool }}' or
pt.get('index_pool') != '{{ ceph_rgw_index_pool }}' or
pt.get('data_extra_pool') != '{{ ceph_rgw_extra_pool }}'):
pt['data_pool'] = '{{ ceph_rgw_data_pool }}'
pt['index_pool'] = '{{ ceph_rgw_index_pool }}'
pt['data_extra_pool'] = '{{ ceph_rgw_extra_pool }}'
changed = True
if not changed:
sys.exit(0)
subprocess.run(
['radosgw-admin', 'zone', 'set', '--rgw-zone', '{{ ceph_rgw_zone }}'],
input=json.dumps(zone).encode(), check=True)
print('CHANGED')
"
args:
executable: /bin/bash
register: zone_placement
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: "'CHANGED' in (zone_placement.stdout | default(''))"
# --- Step 10.5: Zonegroup hostnames ---
#
# The dashboard connects to RGW using the node's hostname or IP in the
# Host header. RGW validates the Host header during S3 signature
# verification — if the hostname isn't in the zonegroup's hostnames
# list, signature verification fails with SignatureDoesNotMatch (403).
# Adding all node hostnames and IPs ensures both dashboard and direct
# client access work regardless of which address is used.
- name: Add cluster hostnames and IPs to zonegroup
ansible.builtin.shell: |
set -o pipefail
python3 -c "
import json, subprocess
zg = json.loads(subprocess.check_output(
['radosgw-admin', 'zonegroup', 'get',
'--rgw-zonegroup', '{{ ceph_rgw_zonegroup }}']))
hostnames = set(zg.get('hostnames', []))
needed = set([
{% for h in groups['ceph_nodes'] %}
'{{ hostvars[h]['hostname_short'] }}',
'{{ hostvars[h]['bond_ip'] }}',
{% endfor %}
'{{ ceph_rgw_dns_name }}',
])
if needed.issubset(hostnames):
print('OK')
else:
hostnames.update(needed)
zg['hostnames'] = sorted(hostnames)
subprocess.run(
['radosgw-admin', 'zonegroup', 'set',
'--rgw-zonegroup', '{{ ceph_rgw_zonegroup }}'],
input=json.dumps(zg).encode(), check=True,
capture_output=True)
print('CHANGED: ' + ','.join(sorted(needed)))
"
args:
executable: /bin/bash
when: inventory_hostname in groups['ceph_bootstrap']
register: zonegroup_hostnames
changed_when: "'CHANGED' in zonegroup_hostnames.stdout"
# --- Step 11: Commit period ---
- name: Commit the period
ansible.builtin.command: >
radosgw-admin period update --commit --rgw-realm={{ ceph_rgw_realm }}
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
# --- Step 11.5: rgw_dns_name for virtual-hosted bucket addressing ---
#
# With this set, a request to bucket.{{ ceph_rgw_dns_name }} is
# recognized as the bucket "bucket" rather than a literal bucket whose
# name is the full FQDN. Clients can then use either path-style
# ({{ ceph_rgw_scheme }}://{{ ceph_rgw_dns_name }}/bucket/key) or
# virtual-hosted style ({{ ceph_rgw_scheme }}://bucket.{{ ceph_rgw_dns_name }}/key).
- name: Get current rgw_dns_name
ansible.builtin.command: ceph config get client.rgw rgw_dns_name
register: current_rgw_dns_name
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
failed_when: false
- name: Set rgw_dns_name for virtual-hosted bucket support
ansible.builtin.command: >
ceph config set client.rgw rgw_dns_name {{ ceph_rgw_dns_name }}
when:
- inventory_hostname in groups['ceph_bootstrap']
- "(current_rgw_dns_name.stdout | default('')) | trim != ceph_rgw_dns_name"
changed_when: true
# --- Step 11.6: Self-signed TLS cert for RGW frontend ---
#
# Generated on the bootstrap node via openssl, handed to cephadm via the
# service spec template below. cephadm distributes the combined PEM to
# every RGW daemon container on apply.
#
# Idempotent via `creates:` — the openssl run only fires if the cert file
# is absent. To rotate: delete /etc/ceph/rgw-ssl.{crt,key} on the bootstrap
# node and re-run the role.
#
# SANs include:
# - the canonical DNS name (CN)
# - a wildcard under the same name for virtual-hosted buckets
# - per-node FQDNs so direct-host addressing also validates
# - per-node bond IPs so IP-based S3 clients also validate
- name: Generate RGW self-signed cert and key (10 year validity)
ansible.builtin.shell: |
set -euo pipefail
umask 077
openssl req -x509 -newkey rsa:4096 -nodes \
-keyout /etc/ceph/rgw-ssl.key \
-out /etc/ceph/rgw-ssl.crt \
-days {{ ceph_rgw_ssl_cert_days }} \
-subj "/C={{ ceph_rgw_ssl_cert_subject_c }}\
/ST={{ ceph_rgw_ssl_cert_subject_st }}\
/L={{ ceph_rgw_ssl_cert_subject_l }}\
/O={{ ceph_rgw_ssl_cert_subject_o }}\
/CN={{ ceph_rgw_dns_name }}\
/emailAddress={{ ceph_rgw_ssl_cert_email }}" \
-addext "subjectAltName=\
DNS:{{ ceph_rgw_dns_name }},\
DNS:*.{{ ceph_rgw_dns_name }}\
{% for h in groups['ceph_nodes'] %},DNS:{{ hostvars[h]['hostname_short'] }}.{{ cluster_domain }}{% endfor %}\
{% for h in groups['ceph_nodes'] %},IP:{{ hostvars[h]['bond_ip'] }}{% endfor %}"
args:
executable: /bin/bash
creates: /etc/ceph/rgw-ssl.crt
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_rgw_ssl | default(false) | bool
changed_when: true
- name: Read combined cert + key for service spec
ansible.builtin.shell: cat /etc/ceph/rgw-ssl.crt /etc/ceph/rgw-ssl.key
args:
executable: /bin/bash
register: rgw_ssl_combined
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_rgw_ssl | default(false) | bool
changed_when: false
no_log: true
- name: Set combined PEM fact
ansible.builtin.set_fact:
rgw_ssl_cert_combined_pem: "{{ rgw_ssl_combined.stdout }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_rgw_ssl | default(false) | bool
no_log: true
# --- Step 12: Deploy RGW daemons via cephadm service spec ---
#
# We render a YAML service spec rather than using `ceph orch apply rgw
# <realm> <zone> --placement=... --port=...` because the command-line
# form doesn't accept cert content. The YAML spec supports both
# `ssl: true` and inline `rgw_frontend_ssl_certificate` — cephadm
# distributes the cert to every RGW container.
#
# Placement is pinned to groups['ceph_nodes'] (not ansible_play_batch)
# so a --limit re-run doesn't accidentally shrink the spec and
# de-deploy daemons on omitted hosts (cephadm is declarative: applying
# a smaller placement REMOVES daemons).
- name: Render RGW service spec
ansible.builtin.template:
src: rgw-spec.yaml.j2
dest: /etc/ceph/rgw-spec.yaml
owner: root
group: root
mode: '0600'
when: inventory_hostname in groups['ceph_bootstrap']
no_log: true
- name: Apply RGW service spec
ansible.builtin.command: ceph orch apply -i /etc/ceph/rgw-spec.yaml
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
# --- Step 12.5: Dashboard RGW SSL verify ---
#
# With self-signed certs the dashboard's RGW client rejects the cert
# and returns 500 on the Object Gateway page. Disable verification so
# the dashboard can talk to the local RGW endpoints.
- name: Disable dashboard RGW API SSL verification (self-signed cert)
ansible.builtin.command: ceph dashboard set-rgw-api-ssl-verify false
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_rgw_ssl | default(false) | bool
changed_when: false
- name: Sync dashboard RGW credentials
ansible.builtin.command: ceph dashboard set-rgw-credentials
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
# --- Step 13: Wait for RGW daemons ---
- name: Wait for RGW daemons to start
ansible.builtin.shell: |
set -eo pipefail
ceph orch ls --service-type rgw --format json | python3 -c "
import sys, json
svcs = json.load(sys.stdin)
running = sum(s.get('status', {}).get('running', 0) for s in svcs)
print(running)
"
args:
executable: /bin/bash
register: rgw_running_count
until: "(rgw_running_count.stdout | int) >= (groups['ceph_nodes'] | length)"
retries: 30
delay: 10
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
# --- Step 13.5: Ensure dashboard RGW user has admin caps ---
#
# cephadm auto-creates a "dashboard" RGW user with system=true but
# zero caps. Without admin caps the dashboard's Object Gateway page
# returns 403 SignatureDoesNotMatch on every admin API call.
- name: Ensure dashboard RGW user has admin caps
ansible.builtin.shell: |
set -o pipefail
CAPS=$(radosgw-admin user info --uid=dashboard --format json 2>/dev/null \
| python3 -c "import sys,json; print(len(json.load(sys.stdin).get('caps',[])))")
if [ "$CAPS" = "0" ]; then
radosgw-admin caps add --uid=dashboard \
--caps='buckets=*;users=*;usage=*;metadata=*;zone=*' >/dev/null 2>&1
echo "CHANGED"
else
echo "OK"
fi
args:
executable: /bin/bash
when: inventory_hostname in groups['ceph_bootstrap']
register: dashboard_caps
changed_when: "'CHANGED' in dashboard_caps.stdout"
# --- Step 14: Create S3 user ---
- name: Check if S3 user exists
ansible.builtin.command: "radosgw-admin user info --uid={{ ceph_rgw_s3_user_uid }}"
register: s3_user_check
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
failed_when: false
- name: Create S3 service user with predetermined keys — {{ ceph_rgw_s3_user_uid }}
ansible.builtin.command: >
radosgw-admin user create
--uid={{ ceph_rgw_s3_user_uid }}
--display-name='{{ ceph_rgw_s3_user_display_name }}'
--access-key='{{ ceph_rgw_s3_user_access_key }}'
--secret-key='{{ ceph_rgw_s3_user_secret_key }}'
--max-buckets=100
register: s3_user_create
when:
- inventory_hostname in groups['ceph_bootstrap']
- s3_user_check.rc != 0
changed_when: true
no_log: true
# --- Step 15: Align existing S3 user keys with 1P (idempotent rotation) ---
#
# For clusters that predate the TF+1P-predetermined-keys pattern, the user
# exists but with cephadm-generated random keys. This task detects that
# drift and realigns — destructive for the Yucca-app side, so gated by an
# explicit flag.
- name: Get current S3 user info
ansible.builtin.command: "radosgw-admin user info --uid={{ ceph_rgw_s3_user_uid }}"
register: s3_user_info
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
no_log: true
- name: Detect S3 key drift (1P vs cluster)
ansible.builtin.set_fact:
s3_key_drift: >-
{{
(s3_user_info.stdout | from_json)['keys'][0]['access_key']
!= ceph_rgw_s3_user_access_key
}}
when: inventory_hostname in groups['ceph_bootstrap']
no_log: true
- name: Align S3 user keys with 1P values (only when rotate_s3_keys=true)
ansible.builtin.command: >
radosgw-admin key create
--uid={{ ceph_rgw_s3_user_uid }}
--access-key='{{ ceph_rgw_s3_user_access_key }}'
--secret-key='{{ ceph_rgw_s3_user_secret_key }}'
when:
- inventory_hostname in groups['ceph_bootstrap']
- s3_key_drift | default(false)
- rotate_s3_keys | default(false) | bool
changed_when: true
no_log: true
- name: S3 key drift notice (dry flag — no action taken)
ansible.builtin.debug:
msg:
- "S3 svc-user keys differ from 1P values."
- "To align (rotates keys, requires Yucca-app re-config): re-run with -e rotate_s3_keys=true"
when:
- inventory_hostname in groups['ceph_bootstrap']
- s3_key_drift | default(false)
- not (rotate_s3_keys | default(false) | bool)
- name: Display S3 endpoint (credentials live in 1P, not log output)
ansible.builtin.debug:
msg:
- "=== S3 endpoint for {{ ceph_rgw_s3_user_uid }} ==="
- "Access Key: op://{{ cluster_secrets_vault }}/{{ cluster_name | upper }}_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY/password"
- "Secret Key: op://{{ cluster_secrets_vault }}/{{ cluster_name | upper }}_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY/password"
- "Endpoint (DNS): {{ ceph_rgw_scheme }}://{{ ceph_rgw_dns_name }}:{{ ceph_rgw_port }}"
- "Endpoint (direct): {{ ceph_rgw_scheme }}://{{ bond_ip }}:{{ ceph_rgw_port }}"
- "Region: {{ ceph_rgw_zonegroup_api_name }}"
- "Virtual-hosted: {{ ceph_rgw_scheme }}://<bucket>.{{ ceph_rgw_dns_name }}:{{ ceph_rgw_port }}"
- "Path-style: {{ ceph_rgw_scheme }}://{{ ceph_rgw_dns_name }}:{{ ceph_rgw_port }}/<bucket>"
when: inventory_hostname in groups['ceph_bootstrap']
no_log: false
# --- Step 16: Show RGW status ---
- name: Show RGW status
ansible.builtin.shell: |
echo "=== RGW Services ==="
ceph orch ls --service-type rgw
echo ""
echo "=== Realm ==="
radosgw-admin realm list
echo ""
echo "=== Zonegroup ==="
radosgw-admin zonegroup get --rgw-zonegroup={{ ceph_rgw_zonegroup }}
args:
executable: /bin/bash
register: rgw_final_status
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Display RGW status
ansible.builtin.debug:
msg: "{{ rgw_final_status.stdout_lines }}"
when: inventory_hostname in groups['ceph_bootstrap']
@@ -0,0 +1,44 @@
---
# Phase 6: Verify cluster health
- name: Show cluster status
ansible.builtin.shell: |
set -o pipefail
echo "=== Cluster Status ==="
ceph status
echo ""
echo "=== OSD Tree ==="
ceph osd tree
echo ""
echo "=== Cluster Capacity ==="
ceph df
echo ""
echo "=== RGW Services ==="
ceph orch ls --service-type rgw 2>/dev/null || echo "No RGW services deployed"
echo ""
echo "=== RGW Endpoints ==="
radosgw-admin zonegroup get --rgw-zonegroup={{ ceph_rgw_zonegroup }} 2>/dev/null | python3 -c "
import sys, json
try:
zg = json.load(sys.stdin)
print(f\"Region (api_name): {zg.get('api_name', 'N/A')}\")
print(f\"Endpoints: {zg.get('endpoints', [])}\")
except: print('RGW not configured')
" || echo "RGW not configured"
echo ""
echo "=== Monitoring Services ==="
ceph orch ls --service-type prometheus 2>/dev/null || echo "No monitoring deployed"
ceph orch ls --service-type grafana 2>/dev/null || true
ceph orch ls --service-type alertmanager 2>/dev/null || true
ceph orch ls --service-type node-exporter 2>/dev/null || true
ceph orch ls --service-type ceph-exporter 2>/dev/null || true
args:
executable: /bin/bash
when: inventory_hostname in groups['ceph_bootstrap']
register: cluster_status
changed_when: false
- name: Display final cluster state
ansible.builtin.debug:
msg: "{{ cluster_status.stdout_lines }}"
when: inventory_hostname in groups['ceph_bootstrap']
@@ -0,0 +1,23 @@
[global]
# Initial cluster config applied during cephadm bootstrap
osd_pool_default_size = 2
osd_pool_default_min_size = 1
public_network = {{ public_network }}
cluster_network = {{ cluster_network }}
# Recovery/backfill throttling (small cluster, shared network)
osd_recovery_max_active = 1
osd_recovery_sleep_hdd = 0.1
osd_max_backfills = 1
# PG autoscaler target
mon_target_pg_per_osd = 100
[osd]
# Let BlueStore use the entire block.db LV (pre-created at 240G)
bluestore_block_db_size = 0
# Scrub window — 2-6 AM, deep scrub weekly
osd_scrub_begin_hour = 2
osd_scrub_end_hour = 6
osd_deep_scrub_interval = 604800
@@ -0,0 +1,62 @@
# OSD service spec — rendered by ceph_deploy role
# Applied via: ceph orch apply osd -i /etc/ceph/osd-spec.yml
#
# One document per host because per-host disk paths differ:
# - sietch nodes: each has a unique sas_path_prefix (different SAS expander
# address per chassis), so explicit data_devices.paths can't be shared.
# - painbox: single host, single spec.
#
# Hardware-shape independence: when sas_path_prefix is defined the host is
# sietch-shape (SAS expander, dual-SSD-VG block.db); otherwise the host is
# painbox-shape (PCI-ATA disks, single NVMe-RAID VG, SSD OSD as an LV).
#
# Idempotent: cephadm no-ops when the spec matches what's deployed. New
# disks (e.g. populating an empty bay) are picked up automatically on apply.
#
# Do not edit by hand. Regenerate by re-running the ceph_deploy role.
{% for host in groups['ceph_nodes'] %}
{% set h = hostvars[host] %}
{% if h.ceph_hdd_osds | default([]) | length > 0 %}
---
service_type: osd
service_id: {{ h.hostname_short }}-hdd
placement:
hosts:
- {{ h.hostname_short }}
spec:
data_devices:
paths:
{% for hdd in h.ceph_hdd_osds %}
{% if h.sas_path_prefix is defined %}
- /dev/disk/by-path/{{ h.sas_path_prefix }}-{{ hdd.path_phy }}-lun-0
{% else %}
- /dev/disk/by-path/{{ hdd.path_phy }}
{% endif %}
{% endfor %}
db_devices:
paths:
{% for hdd in h.ceph_hdd_osds %}
- /dev/{{ hdd.db }}
{% endfor %}
encrypted: true
{% endif %}
{% if h.ceph_ssd_osds | default([]) | length > 0 %}
---
service_type: osd
service_id: {{ h.hostname_short }}-ssd
placement:
hosts:
- {{ h.hostname_short }}
spec:
data_devices:
paths:
{% for ssd in h.ceph_ssd_osds %}
{% if ssd.lv is defined %}
- /dev/{{ ssd.lv }}
{% else %}
- /dev/disk/by-path/{{ h.sas_path_prefix }}-{{ ssd.path_phy }}-lun-0-part{{ ssd.partition }}
{% endif %}
{% endfor %}
encrypted: true
{% endif %}
{% endfor %}
@@ -0,0 +1,21 @@
# RGW service spec — rendered by ceph-deploy role
# Applied via: ceph orch apply -i /etc/ceph/rgw-spec.yaml
#
# Do not edit by hand. Regenerate by re-running the ceph-deploy role.
service_type: rgw
service_id: {{ ceph_rgw_realm }}.{{ ceph_rgw_zone }}
placement:
hosts:
{% for host in groups['ceph_nodes'] %}
- {{ hostvars[host]['hostname_short'] }}
{% endfor %}
spec:
rgw_realm: {{ ceph_rgw_realm }}
rgw_zone: {{ ceph_rgw_zone }}
rgw_frontend_port: {{ ceph_rgw_port }}
rgw_frontend_type: beast
{% if ceph_rgw_ssl | default(false) %}
ssl: true
rgw_frontend_ssl_certificate: |
{{ rgw_ssl_cert_combined_pem | trim | indent(4, True) }}
{% endif %}
@@ -0,0 +1,210 @@
---
# Phase 5: Clean up LVM, device signatures, and leftover Ceph state on all nodes
#
# Order matters:
# 1. Close dm-crypt/LUKS mappings (devices are locked until closed)
# 2. Remove device-mapper entries
# 3. Remove LVM volume groups and physical volumes
# 4. Wipe device signatures (HDD and SSD OSD partitions)
# 5. Clean stale LVM metadata files
# 6. Reset failed systemd units
# 7. Remove Ceph directories
# --- Step 1: Close dm-crypt LUKS mappings ---
# ceph-volume --dmcrypt creates LUKS volumes on OSD devices. These persist after
# cephadm rm-cluster and lock the underlying block devices. Must close before wipefs.
- name: Close dm-crypt LUKS mappings
ansible.builtin.shell: |
set -o pipefail
CLOSED=0
TARGETS=$(dmsetup ls --target crypt 2>/dev/null | grep -v 'No devices found' || true)
for dm in $(echo "$TARGETS" | awk 'NF {print $1}'); do
echo "Closing LUKS: $dm"
cryptsetup close "$dm" 2>&1 || dmsetup remove -f "$dm" 2>&1 || true
CLOSED=$((CLOSED+1))
done
if [ $CLOSED -eq 0 ]; then echo "No LUKS mappings to close"; fi
args:
executable: /bin/bash
register: luks_cleanup
changed_when: "'Closing LUKS' in luks_cleanup.stdout"
# --- Step 2: Remove leftover device mapper entries ---
- name: Remove Ceph device mapper entries
ansible.builtin.shell: |
set -o pipefail
REMOVED=0
for dm in $(dmsetup ls 2>/dev/null | grep -i ceph | awk '{print $1}'); do
echo "Removing DM: $dm"
dmsetup remove -f "$dm" 2>&1 || true
REMOVED=$((REMOVED+1))
done
if [ $REMOVED -eq 0 ]; then echo "No Ceph DM entries found"; fi
args:
executable: /bin/bash
register: dm_cleanup
changed_when: "'Removing DM' in dm_cleanup.stdout"
# --- Step 3: Remove Ceph LVM volume groups ---
# Catches both ceph-db VGs (ceph-db-rear12, ceph-db-rear13) and ceph-volume
# created VGs (ceph-<uuid> from dmcrypt OSDs).
- name: Remove Ceph LVM volume groups
ansible.builtin.shell: |
set -o pipefail
REMOVED=0
for vg in $(vgs --noheadings -o vg_name 2>/dev/null | grep -i ceph | tr -d ' '); do
echo "Removing VG: $vg"
vgremove -f "$vg" 2>&1 || true
REMOVED=$((REMOVED+1))
done
if [ $REMOVED -eq 0 ]; then echo "No Ceph VGs found"; fi
args:
executable: /bin/bash
register: vg_cleanup
changed_when: "'Removing VG' in vg_cleanup.stdout"
# --- Step 4: Remove LVM physical volumes ---
# After VG removal, PVs are orphaned (no VG name in pvs output). Must catch BOTH
# PVs that still reference a ceph VG AND orphaned PVs on HDD/SSD OSD devices.
- name: Remove Ceph LVM physical volumes (named VGs)
ansible.builtin.shell: |
set -o pipefail
REMOVED=0
for pv in $(pvs --noheadings -o pv_name,vg_name 2>/dev/null | grep -i ceph | awk '{print $1}'); do
echo "Removing PV: $pv"
pvremove -f "$pv" 2>&1 || true
REMOVED=$((REMOVED+1))
done
if [ $REMOVED -eq 0 ]; then echo "No named Ceph PVs found"; fi
args:
executable: /bin/bash
register: pv_cleanup_named
changed_when: "'Removing PV' in pv_cleanup_named.stdout"
- name: Remove orphaned LVM physical volumes on HDD devices
ansible.builtin.shell: |
set -o pipefail
REMOVED=0
for dev in $(lsblk -dnpo NAME,ROTA,TYPE | awk '$2==1 && $3=="disk" {print $1}'); do
if pvs "$dev" --noheadings 2>/dev/null | grep -q "$dev"; then
echo "Removing orphaned PV: $dev"
pvremove -f "$dev" 2>&1 || true
REMOVED=$((REMOVED+1))
fi
done
if [ $REMOVED -eq 0 ]; then echo "No orphaned HDD PVs found"; fi
args:
executable: /bin/bash
register: pv_cleanup_hdd
changed_when: "'Removing orphaned PV' in pv_cleanup_hdd.stdout"
- name: Remove orphaned LVM physical volumes on SSD OSD partitions
ansible.builtin.shell: |
set -o pipefail
REMOVED=0
for ssd in $(lsblk -dnpo NAME,MODEL | grep -i "{{ ssd_model_pattern }}" | awk '{print $1}'); do
if [[ "$ssd" == *nvme* ]]; then PART="${ssd}p6"; else PART="${ssd}6"; fi
if [ -b "$PART" ] && pvs "$PART" --noheadings 2>/dev/null | grep -q "$PART"; then
echo "Removing orphaned PV: $PART"
pvremove -f "$PART" 2>&1 || true
REMOVED=$((REMOVED+1))
fi
done
if [ $REMOVED -eq 0 ]; then echo "No orphaned SSD OSD PVs found"; fi
args:
executable: /bin/bash
register: pv_cleanup_ssd
changed_when: "'Removing orphaned PV' in pv_cleanup_ssd.stdout"
# --- Step 5: Wipe device signatures ---
# Wipe ALL rotational (HDD) devices that have any filesystem/LVM/LUKS signatures.
# Every HDD in these servers is a Ceph OSD — no ambiguity about which to wipe.
# Only wipes devices that actually have signatures (idempotent on clean disks).
- name: Wipe HDD device signatures
ansible.builtin.shell: |
set -o pipefail
ZAPPED=0
for dev in $(lsblk -dnpo NAME,ROTA,TYPE | awk '$2==1 && $3=="disk" {print $1}'); do
SIGS=$(wipefs "$dev" 2>/dev/null | tail -n +2)
if [ -n "$SIGS" ]; then
echo "Wiping $dev"
wipefs -af "$dev" 2>&1 || true
dd if=/dev/zero of="$dev" bs=1M count=10 2>/dev/null || true
ZAPPED=$((ZAPPED+1))
fi
done
if [ $ZAPPED -eq 0 ]; then echo "No HDD devices to wipe"; fi
args:
executable: /bin/bash
register: hdd_zap
changed_when: "'Wiping' in hdd_zap.stdout"
- name: Wipe SSD OSD partition signatures (partition 6)
ansible.builtin.shell: |
set -o pipefail
ZAPPED=0
for ssd in $(lsblk -dnpo NAME,MODEL | grep -i "{{ ssd_model_pattern }}" | awk '{print $1}'); do
if [[ "$ssd" == *nvme* ]]; then PART="${ssd}p6"; else PART="${ssd}6"; fi
if [ -b "$PART" ]; then
SIGS=$(wipefs "$PART" 2>/dev/null | tail -n +2)
if [ -n "$SIGS" ]; then
echo "Wiping SSD OSD partition $PART"
wipefs -af "$PART" 2>&1 || true
dd if=/dev/zero of="$PART" bs=1M count=10 2>/dev/null || true
ZAPPED=$((ZAPPED+1))
fi
fi
done
if [ $ZAPPED -eq 0 ]; then echo "No SSD OSD partitions to wipe"; fi
args:
executable: /bin/bash
register: ssd_zap
changed_when: "'Wiping' in ssd_zap.stdout"
# --- Step 6: Clean stale LVM metadata ---
# ceph-volume operations generate LVM archive/backup files that accumulate across
# cluster lifecycles. Hundreds of ceph-* files can build up in /etc/lvm/.
- name: Remove stale Ceph LVM archive and backup files
ansible.builtin.shell: |
REMOVED=0
for dir in /etc/lvm/archive /etc/lvm/backup; do
for f in "$dir"/ceph-*; do
[ -f "$f" ] || continue
rm -f "$f"
REMOVED=$((REMOVED+1))
done
done
echo "Removed $REMOVED stale LVM metadata files"
register: lvm_meta_cleanup
changed_when: "lvm_meta_cleanup.stdout is not search('Removed 0')"
# --- Step 7: Reset failed Ceph systemd units ---
- name: Reset failed Ceph systemd units
ansible.builtin.shell: |
set -o pipefail
RESET=0
for unit in $(systemctl list-units --all --failed --no-legend 2>/dev/null | grep -i ceph | awk '{print $1}'); do
echo "Resetting failed unit: $unit"
systemctl reset-failed "$unit" 2>&1 || true
RESET=$((RESET+1))
done
if [ $RESET -eq 0 ]; then echo "No failed Ceph units to reset"; fi
args:
executable: /bin/bash
register: systemd_cleanup
changed_when: "'Resetting failed unit' in systemd_cleanup.stdout"
# --- Step 8: Remove Ceph directories ---
- name: Clean up leftover Ceph directories
ansible.builtin.file:
path: "{{ item }}"
state: absent
loop:
- /etc/ceph
- /var/lib/ceph
- /var/log/ceph
- /tmp/ceph-initial.conf
- name: Summary
ansible.builtin.debug:
msg: "Ceph cluster teardown complete on {{ hostname_short }}. Devices are clean and ready for reuse."
@@ -0,0 +1,47 @@
---
# Phase 3: Remove non-bootstrap hosts from cluster (from bootstrap node)
- name: Get cluster hosts
ansible.builtin.shell: |
set -o pipefail
timeout 10 ceph orch host ls --format json 2>/dev/null | python3 -c "
import sys, json
hosts = json.load(sys.stdin)
for h in hosts:
print(h['hostname'])
" 2>/dev/null || true
args:
executable: /bin/bash
register: cluster_hosts
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Remove non-bootstrap hosts from cluster
ansible.builtin.shell: |
HOST="{{ item }}"
BOOTSTRAP="{{ hostvars[groups['ceph_bootstrap'][0]]['hostname_short'] }}"
if [ "$HOST" = "$BOOTSTRAP" ]; then
echo "Skipping bootstrap host $HOST (removed during purge)"
else
echo "Removing host $HOST from cluster"
ceph orch host drain "$HOST" --force 2>&1 || true
# Wait briefly for daemons to be removed
sleep 10
ceph orch host rm "$HOST" --force 2>&1 || true
fi
loop: "{{ cluster_hosts.stdout_lines | default([]) }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- cluster_hosts.stdout_lines | default([]) | length > 0
register: host_rm_results
changed_when: "'Removing host' in (item.stdout | default(''))"
- name: Show host removal results
ansible.builtin.debug:
msg: "{{ item.stdout }}"
loop: "{{ host_rm_results.results | default([]) }}"
loop_control:
label: "{{ item.item | default('unknown') }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- item.stdout is defined
@@ -0,0 +1,27 @@
---
# Ceph cluster teardown — reverse order of deploy
# Phases: preflight → pools → osds → hosts → purge → cleanup
- name: Phase 0 - Preflight safety checks
ansible.builtin.import_tasks: preflight.yml
tags: [preflight, pools, osds, hosts, purge, cleanup]
- name: Phase 1 - Remove all pools
ansible.builtin.import_tasks: pools.yml
tags: [pools]
- name: Phase 2 - Remove all OSDs
ansible.builtin.import_tasks: osds.yml
tags: [osds]
- name: Phase 3 - Remove non-bootstrap hosts
ansible.builtin.import_tasks: hosts.yml
tags: [hosts]
- name: Phase 4 - Purge cluster from all nodes
ansible.builtin.import_tasks: purge.yml
tags: [purge]
- name: Phase 5 - Clean up devices and LVM
ansible.builtin.import_tasks: cleanup.yml
tags: [cleanup]
@@ -0,0 +1,105 @@
---
# Phase 2: Remove all OSDs (from bootstrap node)
# For each OSD: mark out → stop daemon → purge (removes from CRUSH, auth, etc.)
- name: Get all OSD IDs
ansible.builtin.shell: |
timeout 10 ceph osd ls --format json 2>/dev/null || echo "[]"
register: osd_list
when: inventory_hostname in groups['ceph_bootstrap']
changed_when: false
- name: Parse OSD IDs
ansible.builtin.set_fact:
ceph_osd_ids: "{{ osd_list.stdout | from_json }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- osd_list.stdout is defined
- name: Show OSDs to be removed
ansible.builtin.debug:
msg: "OSDs to remove: {{ ceph_osd_ids | default([]) }}"
when: inventory_hostname in groups['ceph_bootstrap']
- name: Set noout flag to prevent rebalancing during teardown
ansible.builtin.command: ceph osd set noout
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_osd_ids | default([]) | length > 0
changed_when: true
- name: Mark all OSDs out
ansible.builtin.shell: |
ceph osd out osd.{{ item }} 2>&1 || true
loop: "{{ ceph_osd_ids | default([]) }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_osd_ids | default([]) | length > 0
changed_when: true
- name: Stop all OSD daemons via orchestrator
ansible.builtin.shell: |
ceph orch daemon stop osd.{{ item }} 2>&1 || true
loop: "{{ ceph_osd_ids | default([]) }}"
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_osd_ids | default([]) | length > 0
changed_when: true
- name: Wait for OSD daemons to stop
ansible.builtin.shell: |
set -o pipefail
for i in $(seq 1 30); do
up=$(ceph osd stat --format json 2>/dev/null | python3 -c "import sys,json; print(json.load(sys.stdin).get('num_up_osds',0))" 2>/dev/null || echo 0)
if [ "$up" -eq 0 ]; then
echo "All OSD daemons stopped"
exit 0
fi
echo "Waiting for OSDs to stop... ($up still up)"
sleep 5
done
echo "WARNING: Some OSDs still up, proceeding anyway"
args:
executable: /bin/bash
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_osd_ids | default([]) | length > 0
changed_when: false
- name: Remove all OSDs via orchestrator
ansible.builtin.shell: |
ceph orch osd rm {{ ceph_osd_ids | join(' ') }} --zap --force 2>&1 || true
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_osd_ids | default([]) | length > 0
changed_when: true
- name: Wait for OSD removal to complete
ansible.builtin.shell: |
set -o pipefail
for i in $(seq 1 60); do
remaining=$(ceph osd ls --format json 2>/dev/null | python3 -c "import sys,json; print(len(json.load(sys.stdin)))" 2>/dev/null || echo 0)
if [ "$remaining" -eq 0 ]; then
echo "All OSDs removed"
exit 0
fi
echo "Waiting for OSD removal... ($remaining remaining)"
sleep 10
done
echo "WARNING: OSD removal timed out, force-purging remaining"
args:
executable: /bin/bash
when:
- inventory_hostname in groups['ceph_bootstrap']
- ceph_osd_ids | default([]) | length > 0
changed_when: false
- name: Force-purge any remaining OSDs
ansible.builtin.shell: |
for osd_id in $(ceph osd ls 2>/dev/null); do
echo "Force-purging osd.${osd_id}"
ceph osd purge osd.${osd_id} --yes-i-really-mean-it 2>&1 || true
done
when: inventory_hostname in groups['ceph_bootstrap']
register: force_purge
changed_when: "'Force-purging' in force_purge.stdout"

Some files were not shown because too many files have changed in this diff Show More