mirror of
https://github.com/immich-app/yucca.git
synced 2026-09-30 13:33:00 +08:00
feat(ceph): import yucca-ceph ansible + terraform infrastructure (#86)
* feat(ceph): import yucca-ceph ansible + terraform infrastructure
Imports the yucca-ceph Ansible tree into ansible/ceph/ and adds the
Terraform stack at tf/ that drives it. Cuts over from ansible-vault
to the hybrid secrets architecture (TF as inventory authority, 1P
as secrets store, op-inject at deploy time) in one atomic move.
Source: internal yucca-ceph working tree; fresh subtree-style
import, history not preserved. Andy continues operating sietch +
painbox post-merge; yucca-team hosts the code and reviews changes.
What it adds:
- sietch (3-node Austin, production Ceph S3 backend, untouched
by this PR)
- painbox (single-node Hetzner SX295 in Helsinki) as a second
deployable cluster
- Future clusters land by appending to clusters.auto.tfvars in
the matching environment stack (tf/deployment/<env>/ceph/) —
no per-cluster TF code required
How it works (full map: ansible/ceph/docs/architecture.md):
- tf/shared/modules/ceph-cluster renders inventory.ini variants
+ secrets.yml.tpl per cluster from clusters.auto.tfvars
- secrets.yml.tpl carries op:// refs; `op inject -f` resolves
them at deploy time from the matching yucca_tf_<env> vault
- State in OVH yucca-tf-state bucket (key ceph/<env>/<stack>/)
- 11 ADRs capture the decisions: ansible/ceph/docs/adr/
Out of scope (intentional):
- LUKS keys not yet in 1P (deferred until hybrid is stable)
- tf/shared/modules/ceph-cluster/secrets.tf.disabled is dormant;
today's 1P items via `op item create` per
ansible/ceph/docs/adding-a-cluster.md
- Talos K8s on sietch is a separate workstream
Atomicity + rollback: TF-rendered inventory + secrets-template
files are gitignored (TF generates them) and ansible-vault removal
is coupled to the op-inject path. Splitting this PR lands in a
non-bootable state — merge as one unit. The merge itself is
reversible via `git revert` until the post-merge `tf:apply` runs;
after apply, full rollback needs state restore or `tofu state mv`
(land + validate before applying).
Dev-env impact: adds opentofu + terragrunt to yucca root mise tools
plus a self-contained ansible/ceph/.mise.toml. No new commands or
prereqs for immich-side contributors who don't touch ceph or run
tf:* tasks.
Verification:
- `mise run lint` (from ansible/ceph/): 130 files, 0 warnings
- `mise run check`: 19 playbooks parse clean
- `mise run tf:plan`: succeeds; 7 expected file path-rename
replacements (3 painbox + 4 sietch). State drift from import,
no cluster-side change.
- painbox deployed 2026-04-26 on the new code path: Bookworm +
Ceph Tentacle, 15 OSDs (14 HDD + 1 SSD) up + in, mon/mgr/rgw
running. HEALTH_WARN is expected on a single-node cluster.
Post-merge: from the yucca root, `mise run tf:apply` flips the
bucket state to the new monorepo paths (the 7 renames above).
* fix(ceph): exempt ansible/ and tf/ subtrees from root prettier
The imported infrastructure subtrees enforce their own format
conventions (yamllint + ansible-lint inside ansible/ceph/; tofu fmt
inside tf/). Prettier on ansible YAML reflows long Jinja2 expressions
and shell command blocks in unwanted ways, so root prettier checks
are skipped for both subtrees.
Also reformat root README.md table column alignment to match prettier
conventions (only the imported subtrees are exempt; yucca-side files
including the root README still follow root prettier rules).
* fix(ceph): clean up secrets tmpfile after ansible-playbook exits
`ansible-play.sh` rendered the resolved secrets file via `op inject`
into a `mktemp` tmpfile, set up a `trap 'rm -f "$TMPFILE"' EXIT INT
TERM`, then `exec`'d ansible-playbook. The `exec` replaced the bash
shell entirely, so the EXIT trap never fired — every play left a
plaintext-secrets file in /tmp.
In practice this was masked because /tmp is tmpfs (RAM only on this
operator's setup), so files evaporate on reboot. But within an
operator session, files accumulated linearly with each playbook
invocation. Recent count on the import-PR session: 38 files.
Drop the `exec`. With `set -euo pipefail` already on, bash:
- propagates ansible-playbook's exit code (set -e)
- fires the EXIT trap before exiting (always)
- cleans up the tmpfile on success, failure, or signal
Verified: `CEPH_ENV=... scripts/ansible-play.sh status.yml
--syntax-check` creates and removes the tmpfile within the same
invocation — /tmp is clean before and after.
`scripts/preflight.sh` uses the same trap pattern but does not
`exec`, so its tmpfile cleanup was already correct (and the suffix
differs: `-secrets-test.yml` vs `-secrets.yml`, confirming
ansible-play.sh as the sole offender).
This commit is contained in:
+25
@@ -4,7 +4,32 @@ node_modules
|
||||
.env.local
|
||||
.env
|
||||
|
||||
# Exception: tf/.env is committed — contains only op:// references to 1P items,
|
||||
# no literal secrets. Resolved at runtime by `op run --env-file=tf/.env -- ...`.
|
||||
!tf/.env
|
||||
|
||||
mise.local.toml
|
||||
|
||||
dist/
|
||||
packages/michael/michael
|
||||
|
||||
# OpenTofu / Terraform
|
||||
# .terraform.lock.hcl IS committed for reproducibility
|
||||
**/.terraform/
|
||||
**/terraform.tfstate
|
||||
**/terraform.tfstate.*
|
||||
**/*.tfvars.local
|
||||
# terragrunt-generated backend.tf contains absolute per-operator paths
|
||||
**/backend.tf
|
||||
|
||||
# Ansible runtime artifacts
|
||||
ansible/*/ansible.log
|
||||
ansible/*/ansible.*.log
|
||||
ansible/*/.ansible/
|
||||
ansible/*/.ansible_facts_cache/
|
||||
ansible/*/.venv/
|
||||
|
||||
# Lens / decision-support artifacts are controller-local notes, not repo code.
|
||||
# Per-project: kept outside the repo (e.g., ~/Projects/immich/yucca-ceph-import/analysis/
|
||||
# on operator workstation, but not tracked).
|
||||
analysis/
|
||||
|
||||
@@ -7,6 +7,10 @@ restic = "0.18.0"
|
||||
gh = "2.25.0"
|
||||
"github:git-town/git-town" = "22.4.0"
|
||||
|
||||
# Infrastructure tooling (added for tf/ and ansible/ subtrees)
|
||||
opentofu = "1.11.5"
|
||||
terragrunt = "0.99.4"
|
||||
|
||||
[tasks.dev]
|
||||
description = "Start all services in development mode"
|
||||
depends = ["install:deps", "common:build", "docker:start"]
|
||||
@@ -44,6 +48,31 @@ run = [{ tasks = ["lint", "*:check"] }, { task = "format" }, { task = "test" }]
|
||||
description = "Run all possible code fixes"
|
||||
run = [{ task = "*:fix" }, { task = "web:lingui" }]
|
||||
|
||||
# ─── Infrastructure (tf + ansible) ──────────────────────────────────────────
|
||||
# All tf:* tasks wrap terragrunt with `op run --env-file=tf/.env` so the
|
||||
# OP_SERVICE_ACCOUNT_TOKEN is injected from 1Password at invocation time.
|
||||
# No literal secrets in the .env file — just op:// references.
|
||||
|
||||
[tasks."tf:init"]
|
||||
description = "Terragrunt init for a given stack (default: deployment/dev/ceph)"
|
||||
run = "op run --env-file=tf/.env -- terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/dev/ceph} init"
|
||||
|
||||
[tasks."tf:plan"]
|
||||
description = "Terragrunt plan for a given stack"
|
||||
run = "op run --env-file=tf/.env -- terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/dev/ceph} plan"
|
||||
|
||||
[tasks."tf:apply"]
|
||||
description = "Terragrunt apply for a given stack"
|
||||
run = "op run --env-file=tf/.env -- terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/dev/ceph} apply"
|
||||
|
||||
[tasks."tf:destroy"]
|
||||
description = "Terragrunt destroy for a given stack (use with care)"
|
||||
run = "op run --env-file=tf/.env -- terragrunt --working-dir ${TF_STACK_DIR:-tf/deployment/dev/ceph} destroy"
|
||||
|
||||
[tasks."tf:fmt"]
|
||||
description = "Format terraform + terragrunt files recursively"
|
||||
run = "tofu fmt -recursive tf/ && terragrunt hcl format --working-dir tf/"
|
||||
|
||||
[env]
|
||||
NODE_ENV = "development"
|
||||
LOG_LEVEL = "debug"
|
||||
|
||||
@@ -23,3 +23,10 @@ yarn.lock
|
||||
**/locales
|
||||
**/fetch-client.ts
|
||||
**/openapi-specs.json
|
||||
|
||||
# Infrastructure subtrees own their own format conventions (yamllint +
|
||||
# ansible-lint inside ansible/ceph/; tofu fmt inside tf/). Prettier on
|
||||
# ansible YAML is opinionated in ways that conflict with Jinja2 + ansible
|
||||
# task structure, so these subtrees are exempt from root prettier checks.
|
||||
ansible/
|
||||
tf/
|
||||
|
||||
@@ -1,4 +1,11 @@
|
||||
## Development Guide
|
||||
# Yucca
|
||||
|
||||
Application code lives under `packages/`. Infrastructure that operates Yucca
|
||||
(Ceph storage backend, future Talos K8s, deployment/state managed via
|
||||
Terraform+1Password) lives at the top level in `ansible/`, `tf/`, and
|
||||
`kubernetes/`.
|
||||
|
||||
## Development Guide (application)
|
||||
|
||||
Ensure you have prerequisites installed:
|
||||
|
||||
@@ -20,3 +27,19 @@ mise test:integration # integration tests
|
||||
mise test:e2e # e2e tests
|
||||
mise test:e2e:web # e2e web tests
|
||||
```
|
||||
|
||||
## Infrastructure
|
||||
|
||||
| Start here | Path | Purpose |
|
||||
| ---------------------------------------------------- | --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| [`ansible/ceph/README.md`](./ansible/ceph/README.md) | `ansible/ceph/` | Ansible automation for Ceph clusters (sietch, painbox). Deploys + operates via cephadm on bare-metal and Hetzner. |
|
||||
| [`tf/README.md`](./tf/README.md) | `tf/` | Terraform/OpenTofu authority for cluster identity, 1P secret items, rendered Ansible inventories. Terragrunt multi-env (`deployment/<env>/<stack>/`). |
|
||||
| (coming in follow-up) | `kubernetes/` | Flux GitOps surface for the (future) Talos K8s cluster. |
|
||||
|
||||
**Secrets are managed via the `yucca_tf_*` 1Password vaults.** Runtime reads use a
|
||||
read-only service account; TF writes use a superuser service account. See
|
||||
`ansible/ceph/docs/secrets.md` and `tf/README.md` for the full model.
|
||||
|
||||
`mise run tf:init / tf:plan / tf:apply` wraps terragrunt via
|
||||
`op run --env-file=tf/.env --` so the superuser token is injected from 1P at
|
||||
invocation time.
|
||||
|
||||
@@ -0,0 +1,11 @@
|
||||
---
|
||||
# All roles share the ceph_ variable prefix for consistency across the
|
||||
# cluster stack. The var-naming[no-role-prefix] rule expects each role
|
||||
# to use its own prefix, which doesn't fit our single-product layout.
|
||||
skip_list:
|
||||
- var-naming[no-role-prefix]
|
||||
|
||||
# Molecule verify playbooks use relative paths to test template rendering
|
||||
# outside of role context. This is intentional and expected.
|
||||
exclude_paths:
|
||||
- roles/*/molecule/
|
||||
@@ -0,0 +1,29 @@
|
||||
root = true
|
||||
|
||||
[*]
|
||||
end_of_line = lf
|
||||
insert_final_newline = true
|
||||
trim_trailing_whitespace = true
|
||||
charset = utf-8
|
||||
|
||||
[*.{yml,yaml}]
|
||||
indent_style = space
|
||||
indent_size = 2
|
||||
|
||||
[*.{j2,jinja2}]
|
||||
indent_style = space
|
||||
indent_size = 2
|
||||
|
||||
[*.py]
|
||||
indent_style = space
|
||||
indent_size = 4
|
||||
|
||||
[*.{sh,bash}]
|
||||
indent_style = space
|
||||
indent_size = 2
|
||||
|
||||
[Makefile]
|
||||
indent_style = tab
|
||||
|
||||
[*.md]
|
||||
trim_trailing_whitespace = false
|
||||
@@ -0,0 +1,40 @@
|
||||
# Ansible
|
||||
ansible.log
|
||||
ansible.*.log
|
||||
*.retry
|
||||
.ansible/
|
||||
.ansible_facts_cache/
|
||||
|
||||
# Transient outputs — benchmark results, hardware inventory snapshots, backups, exports
|
||||
bench/
|
||||
hardware/
|
||||
backups/
|
||||
exports/
|
||||
|
||||
# TF-generated artifacts — do not commit; re-render with `tofu apply` in tf/deployment/<env>/ceph/
|
||||
inventories/*/inventory.ini
|
||||
inventories/*/inventory-provision.ini
|
||||
inventories/*/inventory-destroy.ini
|
||||
inventories/*/secrets.yml.tpl
|
||||
|
||||
# host_vars IS committed (per-node hardware facts — stable, part of the inventory
|
||||
# source of truth). Operator-local overrides use host_vars/*.local.yml if needed.
|
||||
inventories/*/host_vars/*.local.yml
|
||||
|
||||
# Rendered installimage scripts (source of truth is the .tpl file).
|
||||
inventories/*/installimage/post-install.sh
|
||||
|
||||
# Python virtualenv (mise setup) and bytecode
|
||||
.venv/
|
||||
__pycache__/
|
||||
*.pyc
|
||||
|
||||
# Ansible navigator artifacts
|
||||
.artifacts/
|
||||
ansible-navigator.log
|
||||
|
||||
# OS / editor
|
||||
.DS_Store
|
||||
*.swp
|
||||
*.swo
|
||||
*~
|
||||
@@ -0,0 +1,173 @@
|
||||
[tools]
|
||||
python = "3.12"
|
||||
|
||||
[env]
|
||||
# Project-local virtualenv for all Python tooling
|
||||
VIRTUAL_ENV = "{{config_root}}/.venv"
|
||||
PATH = "{{config_root}}/.venv/bin:{{env.PATH}}"
|
||||
# Note: CEPH_ENV is intentionally NOT declared here. mise's [env] block
|
||||
# overrides shell-exported values, which silently sends operators to the
|
||||
# wrong cluster. Operators must export CEPH_ENV once per shell session:
|
||||
# export CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory.ini
|
||||
# export CEPH_ENV=inventories/painbox-ceph.dev.hel.htz/inventory.ini
|
||||
# Inventory files are TF-generated — `mise run tf:apply` (from yucca root) if missing.
|
||||
|
||||
[tasks.setup]
|
||||
description = "Bootstrap development environment"
|
||||
run = """
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
echo "=== Creating virtualenv ==="
|
||||
python -m venv .venv
|
||||
source .venv/bin/activate
|
||||
|
||||
echo "=== Installing Python dependencies ==="
|
||||
pip install -q -r requirements.txt
|
||||
|
||||
echo "=== Installing Ansible collections ==="
|
||||
ansible-galaxy collection install -r requirements.yml
|
||||
|
||||
echo "=== Verifying ==="
|
||||
ansible --version | head -1
|
||||
ansible-lint --version | head -1
|
||||
yamllint --version
|
||||
|
||||
echo ""
|
||||
echo "Environment ready. Run 'mise trust' if prompted."
|
||||
echo "Next: run 'mise run tf:apply' (from the yucca root) to render inventories."
|
||||
"""
|
||||
|
||||
[tasks.lint]
|
||||
description = "Run all linters (yamllint + ansible-lint + shellcheck)"
|
||||
run = """
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
echo "=== yamllint ==="
|
||||
yamllint roles/ *.yml
|
||||
echo ""
|
||||
echo "=== ansible-lint ==="
|
||||
ansible-lint
|
||||
echo ""
|
||||
echo "=== shellcheck ==="
|
||||
if command -v shellcheck &>/dev/null; then
|
||||
shellcheck scripts/*.sh && echo "All scripts pass"
|
||||
else
|
||||
echo "shellcheck not installed — skipping (install: pacman -S shellcheck)"
|
||||
fi
|
||||
"""
|
||||
|
||||
[tasks.check]
|
||||
description = "Syntax-check all playbooks (no 1P access required)"
|
||||
run = """
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
# Syntax-check only parses YAML — any valid inventory works. Default to
|
||||
# sietch when CEPH_ENV isn't inline-prefixed; the parse is identical.
|
||||
CEPH_ENV="${CEPH_ENV:-inventories/sietch-ceph.dev.austin.int/inventory.ini}"
|
||||
for pb in *.yml; do
|
||||
case "$pb" in
|
||||
requirements.yml|ansible-navigator.yml) continue ;;
|
||||
*)
|
||||
echo "Checking $pb..."
|
||||
ansible-playbook -i "$CEPH_ENV" --syntax-check "$pb" 2>&1 | tail -1
|
||||
;;
|
||||
esac
|
||||
done
|
||||
"""
|
||||
|
||||
[tasks.test]
|
||||
description = "Run Molecule tests for roles with test scenarios"
|
||||
run = """
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
for role in roles/*/molecule; do
|
||||
ROLE_DIR=$(dirname "$role")
|
||||
ROLE_NAME=$(basename "$ROLE_DIR")
|
||||
echo "=== Testing $ROLE_NAME ==="
|
||||
(cd "$ROLE_DIR" && molecule verify)
|
||||
done
|
||||
"""
|
||||
|
||||
[tasks.preflight]
|
||||
description = "Pre-flight checks (TF artifacts, 1P session, SSH, connectivity)"
|
||||
run = "scripts/preflight.sh"
|
||||
|
||||
[tasks.status]
|
||||
description = "Quick cluster health check (read-only)"
|
||||
run = "scripts/ansible-play.sh status.yml"
|
||||
|
||||
[tasks.drift]
|
||||
description = "Detect configuration drift from expected state"
|
||||
run = "scripts/ansible-play.sh drift.yml"
|
||||
|
||||
[tasks.deploy]
|
||||
description = "Full deploy: baseline → tune → deploy → tune ceph → harden"
|
||||
run = """
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
# Rotate ansible.log before a full deploy
|
||||
if [ -f ansible.log ] && [ "$(wc -c < ansible.log)" -gt 1048576 ]; then
|
||||
mv ansible.log "ansible.$(date +%Y%m%dT%H%M%S).log"
|
||||
echo "Rotated ansible.log (>1MB)"
|
||||
fi
|
||||
echo "=== Baseline ===" && scripts/ansible-play.sh baseline.yml
|
||||
echo "=== OS tuning ===" && scripts/ansible-play.sh tune-os.yml
|
||||
echo "=== Hardware tuning ===" && scripts/ansible-play.sh tune-hardware.yml
|
||||
echo "=== Deploy Ceph ===" && scripts/ansible-play.sh deploy-ceph.yml
|
||||
echo "=== Ceph tuning ===" && scripts/ansible-play.sh tune-ceph.yml
|
||||
echo "=== Harden ===" && scripts/ansible-play.sh harden.yml
|
||||
"""
|
||||
|
||||
[tasks.destroy]
|
||||
description = "Destroy Ceph cluster (requires confirmation)"
|
||||
run = """
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
CEPH_ENV_DIR=$(dirname "$CEPH_ENV")
|
||||
CLUSTER_ID=$(basename "$CEPH_ENV_DIR") # e.g., sietch-ceph.dev.austin.int
|
||||
DOMAIN=${CLUSTER_ID#*-ceph.}.futo.cloud # e.g., dev.austin.int.futo.cloud
|
||||
DESTROY_INV="$CEPH_ENV_DIR/inventory-destroy.ini"
|
||||
echo "Usage: scripts/ansible-play.sh destroy-ceph.yml \\"
|
||||
echo " -e yes_destroy_ceph=true -e destroy_target_domain=$DOMAIN"
|
||||
echo " (with CEPH_ENV=$DESTROY_INV if destroying as root)"
|
||||
echo ""
|
||||
read -rp "Run destroy now? [y/N] " confirm
|
||||
if [[ "$confirm" =~ ^[Yy]$ ]]; then
|
||||
CEPH_ENV="$DESTROY_INV" scripts/ansible-play.sh destroy-ceph.yml \
|
||||
-e yes_destroy_ceph=true \
|
||||
-e destroy_target_domain="$DOMAIN"
|
||||
fi
|
||||
"""
|
||||
|
||||
[tasks.bench]
|
||||
description = "S3 benchmark (parallel PUT/GET against local RGW)"
|
||||
run = "scripts/ansible-play.sh bench.yml"
|
||||
|
||||
[tasks.bench-rados]
|
||||
description = "RADOS bench (raw cluster I/O, bypasses RGW)"
|
||||
run = "scripts/ansible-play.sh rados-bench.yml"
|
||||
|
||||
[tasks.backup]
|
||||
description = "Export Ceph cluster config for disaster recovery"
|
||||
run = "scripts/ansible-play.sh backup-config.yml"
|
||||
|
||||
[tasks.capture]
|
||||
description = "Snapshot bootstrap secrets (RGW TLS, admin keyring) to 1P (DR belt)"
|
||||
run = "scripts/ansible-play.sh post-deploy-capture.yml"
|
||||
|
||||
[tasks."rotate-certs"]
|
||||
description = "Rotate RGW TLS cert (regenerate self-signed, restart RGW daemons)"
|
||||
run = "scripts/ansible-play.sh rotate-certs.yml"
|
||||
|
||||
[tasks."rotate-ssh-key"]
|
||||
description = "Distribute current ansible-iac pubkey from 1P to nodes (forward-only)"
|
||||
run = "scripts/ansible-play.sh rotate-ssh-key.yml"
|
||||
|
||||
[tasks."hardware-inventory"]
|
||||
description = "Snapshot per-node hardware facts to JSON (informational only)"
|
||||
run = "scripts/ansible-play.sh hardware-inventory.yml"
|
||||
|
||||
[tasks."migrate-networkd"]
|
||||
description = "One-shot ifupdown→networkd migration (rolling, serial=1, noout-gated)"
|
||||
run = "scripts/ansible-play.sh migrate-networkd.yml"
|
||||
@@ -0,0 +1,12 @@
|
||||
---
|
||||
extends: relaxed
|
||||
|
||||
rules:
|
||||
line-length:
|
||||
max: 260
|
||||
allow-non-breakable-inline-mappings: true
|
||||
comments:
|
||||
min-spaces-from-content: 1
|
||||
octal-values:
|
||||
forbid-implicit-octal: true
|
||||
forbid-explicit-octal: true
|
||||
@@ -0,0 +1,435 @@
|
||||
# Contributing
|
||||
|
||||
Developer guide for the ceph Ansible project in the `yucca` monorepo. Read
|
||||
this before making changes.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
| Tool | Version | Install |
|
||||
|------|---------|---------|
|
||||
| [mise](https://mise.jdx.dev/) | latest | `curl https://mise.jdx.dev/install.sh \| sh` |
|
||||
| Python | 3.12 (managed by mise) | Automatic via `.mise.toml` |
|
||||
| [1Password CLI](https://developer.1password.com/docs/cli) | v2 | `pacman -S 1password-cli` / `brew install 1password-cli` |
|
||||
| [shellcheck](https://www.shellcheck.net/) | latest | `pacman -S shellcheck` |
|
||||
| SSH config | Access to cluster nodes | See below |
|
||||
|
||||
### 1Password access
|
||||
|
||||
You need read access to the **`yucca_tf_dev`** 1Password vault (Futo team
|
||||
membership grants this). The `scripts/ansible-play.sh` wrapper uses `op
|
||||
inject` to resolve secrets at playbook time — desktop session unlock or
|
||||
`OP_SERVICE_ACCOUNT_TOKEN` satisfies auth. No ansible-vault password to
|
||||
manage.
|
||||
|
||||
### SSH setup
|
||||
|
||||
The `ansible-iac` SSH keys live in `yucca_tf_dev` as items
|
||||
`SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY` and `PAINBOX_CEPH_ANSIBLE_IAC_SSH_KEY`
|
||||
(see [ADR-010](docs/adr/010-ssh-keys-in-1password.md) for rationale).
|
||||
|
||||
**First-time workstation setup:**
|
||||
|
||||
```bash
|
||||
cd yucca/ansible/ceph
|
||||
scripts/install-ssh-keys.sh # op read → ~/.ssh/id_ed25519_{sietch,painbox}
|
||||
```
|
||||
|
||||
The script is idempotent and refuses to overwrite an existing key whose
|
||||
fingerprint doesn't match 1P. Keys land as `~/.ssh/id_ed25519_sietch` +
|
||||
`~/.ssh/id_ed25519_painbox` (private + `.pub` both 0600/0644).
|
||||
|
||||
**Jump hosts / proxies** belong in your personal `~/.ssh/config`, not in
|
||||
this repo.
|
||||
|
||||
**Recommended:** run the 1Password desktop app's SSH Agent globally
|
||||
(`Host * IdentityAgent ~/.1password/agent.sock` in `~/.ssh/config`). Keys
|
||||
are served from 1P; private material never leaves the app. The
|
||||
inventory's explicit `ansible_ssh_private_key_file` still works as a
|
||||
fallback.
|
||||
|
||||
## First-time setup
|
||||
|
||||
```bash
|
||||
# Clone the monorepo and navigate to this subproject
|
||||
git clone <yucca-monorepo-url> && cd yucca/ansible/ceph
|
||||
|
||||
# Trust the mise config (one-time per checkout)
|
||||
mise trust
|
||||
# Bootstrap the dev environment (creates .venv, installs Python deps + Ansible collections)
|
||||
mise run setup
|
||||
|
||||
# Render cluster inventories + secrets templates (run once, or after any
|
||||
# change to tf/deployment/dev/ceph/clusters.auto.tfvars)
|
||||
cd ../../tf/deployment/dev/ceph && tofu init && tofu apply && cd -
|
||||
```
|
||||
|
||||
This runs:
|
||||
1. Creates `.venv/` with Python 3.12
|
||||
2. `pip install -r requirements.txt` (ansible-core, ansible-lint, yamllint, molecule, boto3)
|
||||
3. `ansible-galaxy collection install -r requirements.yml` (ansible.posix, community.general)
|
||||
|
||||
Verify:
|
||||
|
||||
```bash
|
||||
ansible --version | head -1 # ansible-core 2.20.x
|
||||
ansible-lint --version # 26.x
|
||||
yamllint --version # 1.38.x
|
||||
```
|
||||
|
||||
## Selecting a cluster
|
||||
|
||||
All tooling uses the `CEPH_ENV` variable to select the target cluster.
|
||||
It points to an **inventory file** (not a directory):
|
||||
|
||||
```bash
|
||||
# Inline prefix — required for `mise run` invocations:
|
||||
CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory.ini mise run preflight
|
||||
CEPH_ENV=inventories/painbox-ceph.dev.hel.htz/inventory.ini mise run status
|
||||
```
|
||||
|
||||
**`export CEPH_ENV=...` does NOT work with `mise run`.** mise's `[env]`
|
||||
block strips shell-exported vars when launching tasks; the wrapper exits
|
||||
with `CEPH_ENV must be set` even though your shell clearly has it set.
|
||||
The inline-prefix form passes the var directly into mise's invocation
|
||||
env where it's preserved. See [docs/scripts.md "Setting CEPH_ENV"](docs/scripts.md)
|
||||
for the full explanation.
|
||||
|
||||
For multiple commands against the same cluster, set a local (non-exported)
|
||||
shell variable and inline-prefix each invocation:
|
||||
|
||||
```bash
|
||||
CE=inventories/painbox-ceph.dev.hel.htz/inventory.ini
|
||||
CEPH_ENV=$CE mise run preflight
|
||||
CEPH_ENV=$CE mise run status
|
||||
CEPH_ENV=$CE mise run deploy
|
||||
```
|
||||
|
||||
Calling scripts directly (e.g., `scripts/preflight.sh`) DOES respect
|
||||
shell `export` — it's only the `mise run` path that filters the env.
|
||||
|
||||
Cluster identity is declared in `tf/deployment/dev/ceph/clusters.auto.tfvars`
|
||||
(keyed by short cluster name). TF renders the directory name, inventory
|
||||
file, and secrets template from that entry. `CEPH_ENV` is just a pointer
|
||||
to the rendered inventory file; wrappers like `scripts/ansible-play.sh`
|
||||
and the `destroy` mise task extract the cluster name from its path for
|
||||
convenience:
|
||||
|
||||
```
|
||||
CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory.ini
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
inventory dir = sietch-ceph.dev.austin.int (rendered by TF)
|
||||
cluster name = sietch (map key in clusters.auto.tfvars)
|
||||
domain = dev.austin.int.futo.cloud (domain field in tfvars)
|
||||
```
|
||||
|
||||
Running `mise run tf:apply` regenerates `inventories/<cluster>/inventory.ini`
|
||||
and `secrets.yml.tpl` any time the tfvars entry changes.
|
||||
|
||||
## Development workflow
|
||||
|
||||
### 1. Edit
|
||||
|
||||
Make your changes in `roles/`, playbooks, or inventory files.
|
||||
|
||||
### 2. Lint
|
||||
|
||||
```bash
|
||||
mise run lint
|
||||
```
|
||||
|
||||
Runs three linters in sequence:
|
||||
- **yamllint** -- relaxed profile, 260-char line limit (`.yamllint`)
|
||||
- **ansible-lint** -- skips `var-naming[no-role-prefix]` because all roles share
|
||||
the `ceph_` prefix (`.ansible-lint`)
|
||||
- **shellcheck** -- all scripts in `scripts/`
|
||||
|
||||
### 3. Syntax-check
|
||||
|
||||
```bash
|
||||
mise run check
|
||||
```
|
||||
|
||||
Runs `ansible-playbook --syntax-check` against every playbook using the active
|
||||
`CEPH_ENV` inventory.
|
||||
|
||||
### 4. Re-render inventory (TF)
|
||||
|
||||
```bash
|
||||
mise run tf:apply
|
||||
```
|
||||
|
||||
Re-renders `inventory.ini` and `secrets.yml.tpl` for every cluster declared
|
||||
in `tf/deployment/dev/ceph/clusters.auto.tfvars`. Required after any cluster-
|
||||
spec edit.
|
||||
|
||||
### 5. Dry-run against real nodes
|
||||
|
||||
```bash
|
||||
scripts/ansible-play.sh baseline.yml --check --diff
|
||||
```
|
||||
|
||||
`--check` simulates changes without applying them. `--diff` shows what would
|
||||
change. Safe to run against production. Use the wrapper (not bare
|
||||
`ansible-playbook`) so secrets resolve via `op inject`.
|
||||
|
||||
### 6. Test with Molecule
|
||||
|
||||
```bash
|
||||
mise run test
|
||||
```
|
||||
|
||||
Roles with `molecule/` directories get template-rendering verification.
|
||||
Molecule doesn't converge (requires real hardware) but verifies Jinja2
|
||||
templates render without errors.
|
||||
|
||||
## Code conventions
|
||||
|
||||
### FQCN everywhere
|
||||
|
||||
Always use fully-qualified collection names for modules:
|
||||
|
||||
```yaml
|
||||
# Good
|
||||
- name: Install packages
|
||||
ansible.builtin.apt:
|
||||
name: htop
|
||||
state: present
|
||||
|
||||
# Bad
|
||||
- name: Install packages
|
||||
apt:
|
||||
name: htop
|
||||
```
|
||||
|
||||
### changed_when is required
|
||||
|
||||
Every `shell` and `command` task must declare `changed_when`:
|
||||
|
||||
```yaml
|
||||
- name: Check OSD count
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
ceph osd stat --format json | python3 -c "..."
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: osd_count
|
||||
changed_when: false # read-only command
|
||||
|
||||
- name: Create OSD
|
||||
ansible.builtin.shell: ...
|
||||
changed_when: "'Created osd' in result.stdout"
|
||||
```
|
||||
|
||||
### Shell task rules
|
||||
|
||||
- Always set `args.executable: /bin/bash`
|
||||
- Always start with `set -o pipefail` (or `set -euo pipefail` for multi-line)
|
||||
- Use `>` folded style for single-line commands, `|` literal style for
|
||||
multi-line:
|
||||
|
||||
```yaml
|
||||
# Single logical command
|
||||
- name: Check disk
|
||||
ansible.builtin.shell: >
|
||||
set -euo pipefail;
|
||||
pvs /dev/sda 2>/dev/null | grep -q ceph
|
||||
|
||||
# Multi-line script
|
||||
- name: Wait for OSDs
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
ceph osd stat --format json
|
||||
args:
|
||||
executable: /bin/bash
|
||||
```
|
||||
|
||||
### Secrets handling
|
||||
|
||||
- Use `no_log: true` on any task that handles passwords, keys, or tokens
|
||||
- Secrets are provisioned in 1Password (see [docs/secrets.md](docs/secrets.md))
|
||||
and consumed at playbook time via `scripts/ansible-play.sh`, which runs
|
||||
`op inject` on the cluster's `secrets.yml.tpl` and passes resolved values
|
||||
as `--extra-vars`.
|
||||
- In ansible code, reference secrets as regular variables via the existing
|
||||
alias pattern in each cluster's `group_vars/all/vars.yml`:
|
||||
|
||||
```yaml
|
||||
# group_vars/all/vars.yml (plaintext, committed)
|
||||
ops_password: "{{ vault_ops_password }}"
|
||||
```
|
||||
|
||||
`vault_ops_password` is populated from `op inject` — no ansible-vault, no
|
||||
encrypted file in git.
|
||||
- `secrets.yml.tpl` is TF-generated and gitignored; don't edit it by hand.
|
||||
Add new secrets via `tf/shared/modules/ceph-cluster/main.tf` (the
|
||||
`local.secrets` map), then `tofu apply` to re-render.
|
||||
|
||||
### Handler pattern
|
||||
|
||||
Define handlers in `roles/<role>/handlers/main.yml`:
|
||||
|
||||
```yaml
|
||||
---
|
||||
- name: Restart sshd
|
||||
ansible.builtin.systemd:
|
||||
name: ssh
|
||||
state: restarted
|
||||
```
|
||||
|
||||
Trigger with `notify`:
|
||||
|
||||
```yaml
|
||||
- name: Update SSH config
|
||||
ansible.builtin.copy:
|
||||
src: sshd_config
|
||||
dest: /etc/ssh/sshd_config
|
||||
notify: Restart sshd
|
||||
```
|
||||
|
||||
### Variable naming
|
||||
|
||||
- All cluster-level variables use the `ceph_` prefix
|
||||
- Role-specific internal variables use `<role_name>_` prefix (e.g.,
|
||||
`baseline_ops_user`, `baseline_podman_packages`)
|
||||
- snake_case for everything
|
||||
- Document every variable in `defaults/main.yml` with a comment explaining
|
||||
what it does and what valid values look like
|
||||
|
||||
### Role names
|
||||
|
||||
- snake_case: `ceph_deploy`, `os_tuning`, `hardware_tuning`
|
||||
- Not: `ceph-deploy`, `CephDeploy`, `cephDeploy`
|
||||
|
||||
### Tags
|
||||
|
||||
Use tags on `import_tasks` in the role's `main.yml` to allow selective runs:
|
||||
|
||||
```yaml
|
||||
- name: Phase 1 - Prerequisites
|
||||
ansible.builtin.import_tasks: prerequisites.yml
|
||||
tags: [prerequisites]
|
||||
|
||||
- name: Phase 2 - Bootstrap cluster
|
||||
ansible.builtin.import_tasks: bootstrap.yml
|
||||
tags: [bootstrap]
|
||||
```
|
||||
|
||||
Run a specific phase:
|
||||
|
||||
```bash
|
||||
scripts/ansible-play.sh deploy-ceph.yml --tags bootstrap
|
||||
```
|
||||
|
||||
## mise tasks reference
|
||||
|
||||
| Task | Command | Description |
|
||||
|------|---------|-------------|
|
||||
| `setup` | `mise run setup` | Bootstrap dev environment (venv, pip, galaxy) |
|
||||
| `lint` | `mise run lint` | yamllint + ansible-lint + shellcheck (no 1P required) |
|
||||
| `check` | `mise run check` | Syntax-check all playbooks (no 1P required) |
|
||||
| `test` | `mise run test` | Molecule verify for roles with test scenarios |
|
||||
| `preflight` | `mise run preflight` | TF artifacts + 1P session + SSH + connectivity |
|
||||
| `status` | `mise run status` | Read-only cluster health check (via ansible-play.sh) |
|
||||
| `drift` | `mise run drift` | Detect configuration drift (via ansible-play.sh) |
|
||||
| `deploy` | `mise run deploy` | Full deploy pipeline (via ansible-play.sh) |
|
||||
| `backup` | `mise run backup` | Export cluster config for DR (via ansible-play.sh) |
|
||||
| `capture` | `mise run capture` | Snapshot RGW TLS + admin keyring to 1P (DR belt) |
|
||||
| `bench-rados` | `mise run bench-rados` | RADOS bench (raw cluster I/O) |
|
||||
| `destroy` | `mise run destroy` | Destroy cluster (interactive confirmation) |
|
||||
|
||||
Inventory rendering and secret-item management are TF responsibilities
|
||||
— `tofu apply` in `tf/deployment/dev/ceph/` renders `inventory.ini` and
|
||||
`secrets.yml.tpl` for every cluster declared in `clusters.auto.tfvars`.
|
||||
Cluster secrets live in `yucca_tf_dev` (see [docs/secrets.md](docs/secrets.md)).
|
||||
|
||||
## File organization
|
||||
|
||||
```
|
||||
yucca/
|
||||
├── tf/ # Terraform state + secrets + rendering (authoritative)
|
||||
│ ├── shared/modules/ceph-cluster/ # Module: per-cluster orchestration + rendering
|
||||
│ └── deployment/dev/ceph/ # Cluster declarations + tofu apply target
|
||||
└── ansible/ceph/ # This directory
|
||||
├── *.yml # Top-level playbooks (site.yml, deploy-ceph.yml, etc.)
|
||||
├── inventories/
|
||||
│ └── <cluster>-ceph.<env>.<dc>.<provider>/
|
||||
│ ├── inventory.ini # TF-rendered (gitignored)
|
||||
│ ├── secrets.yml.tpl # TF-rendered, consumed by op inject (gitignored)
|
||||
│ ├── group_vars/all/
|
||||
│ │ └── vars.yml # Cluster-wide variables (plaintext, committed)
|
||||
│ └── host_vars/
|
||||
│ ├── <hostname>.yml # Per-node hardware config (committed)
|
||||
│ └── <hostname>.local.yml # Per-operator overrides (gitignored)
|
||||
├── roles/
|
||||
│ └── <role_name>/
|
||||
│ ├── defaults/main.yml # Default variables (documented)
|
||||
│ ├── meta/main.yml # Role metadata + dependencies
|
||||
│ ├── tasks/main.yml # Entry point (imports sub-task files)
|
||||
│ ├── handlers/main.yml # Service restart handlers
|
||||
│ ├── templates/*.j2 # Jinja2 templates
|
||||
│ └── molecule/default/ # Test scenario (optional)
|
||||
├── scripts/
|
||||
│ ├── ansible-play.sh # Wrapper: op inject + ansible-playbook
|
||||
│ ├── install-ssh-keys.sh # Pull ansible-iac private keys from 1P → ~/.ssh/
|
||||
│ └── preflight.sh # Pre-deploy checks
|
||||
├── docs/ # Operational documentation
|
||||
│ ├── runbooks/ # Step-by-step operational procedures
|
||||
│ └── archive/ # Historical deployment notes
|
||||
├── .mise.toml # Task runner + Python version + CEPH_ENV
|
||||
├── ansible.cfg # Ansible config (SSH, caching, output)
|
||||
├── requirements.txt # Python dependencies
|
||||
└── requirements.yml # Ansible Galaxy collections
|
||||
```
|
||||
|
||||
### What goes where
|
||||
|
||||
- **`roles/`** -- Reusable automation. Each role handles one concern (baseline
|
||||
setup, Ceph deploy, OS tuning, security). Roles never reference a specific
|
||||
cluster -- they use variables from inventory.
|
||||
- **`inventories/`** -- Cluster-specific data. IPs, hardware mappings, network
|
||||
config, secrets. Each cluster is fully self-contained in its directory.
|
||||
- **`scripts/`** -- Controller-side tooling. Things that run on your workstation,
|
||||
not on target nodes. Secret management, name generation, validation.
|
||||
- **Top-level `*.yml`** -- Playbooks that wire roles to hosts. Each playbook is
|
||||
a thin wrapper: set `hosts`, `become`, and import a role.
|
||||
|
||||
## Commit messages
|
||||
|
||||
Use [conventional commits](https://www.conventionalcommits.org/) for future CI
|
||||
compatibility:
|
||||
|
||||
```
|
||||
feat(ceph_deploy): add RGW virtual-hosted bucket support
|
||||
fix(os_tuning): correct TCP buffer sizes for 10GbE
|
||||
chore(deps): bump ansible-core to 2.20.4
|
||||
docs(runbooks): add disk replacement procedure
|
||||
refactor(baseline): split packages into podman and diagnostics
|
||||
```
|
||||
|
||||
Types: `feat`, `fix`, `chore`, `docs`, `refactor`, `test`, `ci`
|
||||
|
||||
Scope: role name, `scripts`, `inventory`, `deps`, or omit for repo-wide changes.
|
||||
|
||||
## Playbook execution order
|
||||
|
||||
The full deploy pipeline (`mise run deploy` or `site.yml`) runs in this order:
|
||||
|
||||
```
|
||||
1. baseline.yml -- ops user, packages, /etc/hosts, services
|
||||
2. tune-os.yml -- sysctl, TCP buffers, ulimits
|
||||
3. tune-hardware.yml -- I/O scheduler, readahead, queue depth
|
||||
4. deploy-ceph.yml -- cephadm bootstrap, join, placement, OSDs, RGW
|
||||
5. tune-ceph.yml -- recovery throttles, scrub window, telemetry
|
||||
6. harden.yml -- nftables firewall, SSH hardening
|
||||
```
|
||||
|
||||
Provisioning (`provision.yml`) is a separate concern that runs before this
|
||||
pipeline on bare-metal live images.
|
||||
|
||||
## Editor config
|
||||
|
||||
The `.editorconfig` enforces:
|
||||
- YAML/Jinja2: 2-space indent
|
||||
- Python: 4-space indent
|
||||
- Shell: 2-space indent
|
||||
- All files: UTF-8, LF line endings, trailing whitespace trimmed
|
||||
@@ -0,0 +1,150 @@
|
||||
# ceph — Ceph Tentacle on Bare Metal
|
||||
|
||||
Ansible automation for provisioning, deploying, tuning, hardening, and
|
||||
operating Ceph Tentacle (v20) clusters on bare-metal hardware via cephadm.
|
||||
Lives in the `yucca` monorepo at `ansible/ceph/`; secrets and inventory
|
||||
scaffolding are provisioned from `yucca/tf/` (see `../../tf/`).
|
||||
|
||||
| Cluster | Domain | Location | Hardware | Nodes |
|
||||
|---------|--------|----------|----------|-------|
|
||||
| **sietch** | `dev.austin.int.futo.cloud` | Austin DC | Dell R730xd | 3 |
|
||||
| **painbox** | `dev.hel.htz.futo.cloud` | Hetzner Helsinki | SX295 | 1 |
|
||||
|
||||
Clusters are declared in `yucca/tf/deployment/dev/ceph/clusters.auto.tfvars`;
|
||||
`tofu apply` renders `inventories/<cluster>/inventory.ini` and
|
||||
`secrets.yml.tpl` per cluster. The `CEPH_ENV` variable selects the active
|
||||
cluster for any `mise run` or direct ansible invocation.
|
||||
|
||||
## Architecture
|
||||
|
||||
```mermaid
|
||||
graph TB
|
||||
subgraph "Controller (your workstation)"
|
||||
A[mise + ansible + 1Password CLI]
|
||||
end
|
||||
|
||||
subgraph "Austin DC -- 10.10.10.0/24"
|
||||
direction TB
|
||||
L[laurel<br/>MON+MGR+OSD+RGW]
|
||||
W[lawson<br/>MON+MGR+OSD+RGW]
|
||||
S[samara<br/>MON+MGR+OSD+RGW]
|
||||
end
|
||||
|
||||
subgraph "Hetzner Helsinki"
|
||||
P[painbox-ceph-evelyn<br/>MON+MGR+OSD+RGW]
|
||||
end
|
||||
|
||||
A -->|SSH| L
|
||||
A -->|SSH| W
|
||||
A -->|SSH| S
|
||||
A -->|SSH| P
|
||||
```
|
||||
|
||||
See [docs/architecture.md](docs/architecture.md) for role dependencies,
|
||||
data flow, and design rationale.
|
||||
|
||||
## Quick Start
|
||||
|
||||
```bash
|
||||
# 1. Render cluster inventories + secrets templates (once, from yucca/tf/)
|
||||
(cd ../../tf/deployment/dev/ceph && tofu init && tofu apply)
|
||||
|
||||
# 2. Set up the ansible side
|
||||
mise trust && mise run setup # bootstrap dev environment
|
||||
|
||||
# 3. Run mise tasks against the target cluster. CEPH_ENV must be set
|
||||
# inline (NOT via `export`) — see docs/scripts.md "Setting CEPH_ENV"
|
||||
# for why mise's [env] block strips shell exports.
|
||||
CE=inventories/sietch-ceph.dev.austin.int/inventory.ini
|
||||
CEPH_ENV=$CE mise run preflight # TF artifacts + 1P + SSH + connectivity
|
||||
CEPH_ENV=$CE mise run status # read-only cluster health check
|
||||
CEPH_ENV=$CE mise run drift # configuration drift detection
|
||||
CEPH_ENV=$CE mise run deploy # full pipeline (idempotent)
|
||||
```
|
||||
|
||||
See [CONTRIBUTING.md](CONTRIBUTING.md) for the full development workflow.
|
||||
|
||||
## Roles
|
||||
|
||||
| Order | Role | Description |
|
||||
|-------|------|-------------|
|
||||
| 1 | `provision_host` | Bare-metal Debian 12 install via debootstrap (Austin only) |
|
||||
| 2 | `baseline` | Post-boot OS baseline: ops user, packages, /etc/hosts, services |
|
||||
| 3 | `os_tuning` | Kernel sysctl, TCP buffers, optional centralized logging |
|
||||
| 4 | `hardware_tuning` | I/O scheduler, readahead, udev rules, optional CPU governor |
|
||||
| 5 | `ceph_deploy` | cephadm bootstrap, join, placement, OSDs, RGW, monitoring |
|
||||
| 6 | `ceph_tuning` | Recovery throttling, scrub window, CRUSH, telemetry, audit |
|
||||
| 7 | `security` | nftables firewall, SSH hardening |
|
||||
| 8 | `ceph_destroy` | Complete cluster teardown (safety-gated) |
|
||||
| 9 | `s3_bench` | Parallel S3 benchmark against local RGW |
|
||||
|
||||
## Playbooks
|
||||
|
||||
| Playbook | Description |
|
||||
|----------|-------------|
|
||||
| `site.yml` | Full pipeline: baseline + tune + deploy + tune + harden |
|
||||
| `provision.yml` | Bare-metal provisioning (Austin, `-i` provision inventory) |
|
||||
| `baseline.yml` | OS baseline (users, packages, hosts) |
|
||||
| `tune-os.yml` | Kernel/sysctl tuning |
|
||||
| `tune-hardware.yml` | Disk I/O tuning |
|
||||
| `deploy-ceph.yml` | Ceph cluster deployment (tags: prerequisites, bootstrap, join, placement, lvm, osds, crush, rgw, monitoring, verify) |
|
||||
| `tune-ceph.yml` | Post-deploy Ceph config tuning |
|
||||
| `harden.yml` | Firewall + SSH hardening |
|
||||
| `destroy-ceph.yml` | Cluster teardown (destroy inventory, requires flags) |
|
||||
| `status.yml` | Quick health check (read-only) |
|
||||
| `drift.yml` | Configuration drift detection |
|
||||
| `bench.yml` | S3 benchmark (RGW round-trip) |
|
||||
| `rados-bench.yml` | RADOS bench (raw cluster I/O, bypasses RGW) |
|
||||
| `backup-config.yml` | Export cluster config for DR |
|
||||
| `post-deploy-capture.yml` | Snapshot RGW TLS + admin keyring to 1P (DR belt) |
|
||||
| `rotate-certs.yml` | RGW TLS certificate rotation |
|
||||
| `rotate-ssh-key.yml` | Distribute current ansible-iac pubkey from 1P to nodes |
|
||||
| `migrate-networkd.yml` | One-shot networkd/bridge migration (rolling, noout-gated) |
|
||||
| `hardware-inventory.yml` | Hardware facts to JSON |
|
||||
|
||||
## mise Tasks
|
||||
|
||||
| Task | Description |
|
||||
|------|-------------|
|
||||
| `setup` | Bootstrap dev environment (venv, deps, collections) |
|
||||
| `lint` | yamllint + ansible-lint + shellcheck (no 1P required) |
|
||||
| `check` | Syntax-check all playbooks (no 1P required) |
|
||||
| `test` | Molecule tests |
|
||||
| `preflight` | TF artifacts + 1P session + SSH + connectivity |
|
||||
| `status` | Cluster health check |
|
||||
| `drift` | Configuration drift detection |
|
||||
| `deploy` | Full pipeline |
|
||||
| `destroy` | Cluster teardown (interactive) |
|
||||
| `backup` | Export cluster config for DR |
|
||||
| `capture` | Snapshot RGW TLS + admin keyring to 1P (DR belt) |
|
||||
| `bench-rados` | RADOS bench (raw cluster I/O) |
|
||||
|
||||
Inventory scaffolding + secret-item provisioning live in `yucca/tf/` — run
|
||||
`tofu apply` in `tf/deployment/dev/ceph/` to (re-)render
|
||||
`inventories/<cluster>/inventory.ini` and `secrets.yml.tpl`.
|
||||
|
||||
## Documentation
|
||||
|
||||
| Document | Audience |
|
||||
|----------|----------|
|
||||
| [CONTRIBUTING.md](CONTRIBUTING.md) | Developers -- setup, workflow, conventions |
|
||||
| [docs/architecture.md](docs/architecture.md) | Developers -- role graph, data flow, design |
|
||||
| [docs/patterns.md](docs/patterns.md) | Developers -- coding idioms, anti-patterns |
|
||||
| [docs/adding-a-role.md](docs/adding-a-role.md) | Developers -- role skeleton, conventions |
|
||||
| [docs/adding-a-cluster.md](docs/adding-a-cluster.md) | Developers -- inventory setup, secrets |
|
||||
| [docs/secrets.md](docs/secrets.md) | Developers/ops -- 1Password integration |
|
||||
| [docs/naming.md](docs/naming.md) | Everyone -- hostname and inventory naming |
|
||||
| [docs/hardware.md](docs/hardware.md) | Ops/procurement -- R730xd vs SX295 specs |
|
||||
| [docs/s3-integration.md](docs/s3-integration.md) | App developers -- endpoints, boto3, certs |
|
||||
| [docs/security-model.md](docs/security-model.md) | InfoSec -- encryption, users, firewall |
|
||||
| [docs/capacity-planning.md](docs/capacity-planning.md) | Managers -- costs, formulas, growth |
|
||||
| [docs/troubleshooting.md](docs/troubleshooting.md) | SRE/on-call -- symptom/diagnosis/fix |
|
||||
| [docs/runbooks/](docs/runbooks/) | Ops -- add/replace node, replace disk, rotate certs/secrets/SSH/SA token, remote hands, painbox reprovision, bad-tofu-apply recovery, backup/restore |
|
||||
| [docs/adr/](docs/adr/) | Everyone -- architecture decision records |
|
||||
|
||||
## Known Limitations
|
||||
|
||||
- **Single-network topology**: public = cluster network on both clusters.
|
||||
- **Self-signed TLS**: RGW clients need `--no-verify-ssl`. Production needs real certs.
|
||||
- **`ops` user is password-only**: no SSH keys installed; password sourced from 1P. Intended as an interactive console / recovery account, not for automation.
|
||||
- **DNS not managed**: `s3.<domain>` and `*.s3.<domain>` records must exist externally.
|
||||
@@ -0,0 +1,21 @@
|
||||
---
|
||||
ansible-navigator:
|
||||
mode: stdout
|
||||
playbook-artifact:
|
||||
enable: true
|
||||
save-as: .artifacts/{playbook_name}-{time_stamp}.json
|
||||
|
||||
ansible:
|
||||
config:
|
||||
path: ansible.cfg
|
||||
|
||||
logging:
|
||||
level: warning
|
||||
file: ansible-navigator.log
|
||||
|
||||
color:
|
||||
enable: true
|
||||
osc4: true
|
||||
|
||||
execution-environment:
|
||||
enabled: false
|
||||
@@ -0,0 +1,30 @@
|
||||
[defaults]
|
||||
# No default inventory — use -i or CEPH_ENV. Run plays via scripts/ansible-play.sh
|
||||
# which injects secrets (from secrets.yml.tpl via op inject) as --extra-vars.
|
||||
log_path = ansible.log
|
||||
|
||||
# Performance
|
||||
forks = 20
|
||||
gathering = smart
|
||||
fact_caching = jsonfile
|
||||
fact_caching_connection = .ansible_facts_cache
|
||||
fact_caching_timeout = 3600
|
||||
|
||||
# Output
|
||||
stdout_callback = default
|
||||
result_format = yaml
|
||||
callbacks_enabled = ansible.posix.timer, ansible.posix.profile_tasks
|
||||
force_color = True
|
||||
diff_always = True
|
||||
deprecation_warnings = False
|
||||
retry_files_enabled = False
|
||||
display_skipped_hosts = False
|
||||
|
||||
# Security
|
||||
host_key_checking = False
|
||||
timeout = 30
|
||||
|
||||
[ssh_connection]
|
||||
# No ProxyJump here — set per-inventory via ansible_ssh_common_args
|
||||
ssh_args = -o ControlMaster=auto -o ControlPersist=60s -o StrictHostKeyChecking=no
|
||||
pipelining = True
|
||||
@@ -0,0 +1,159 @@
|
||||
---
|
||||
# Export Ceph cluster configuration for disaster recovery.
|
||||
# Captures: ceph.conf, admin keyring, CRUSH map, RGW realm config,
|
||||
# monitor map, and OSD map to the controller's backups/ directory.
|
||||
#
|
||||
# Usage:
|
||||
# scripts/ansible-play.sh backup-config.yml
|
||||
#
|
||||
# Outputs: backups/<timestamp>/ on the controller (gitignored)
|
||||
# Restore: see docs/runbooks/recovery.md
|
||||
|
||||
- name: Backup Ceph cluster configuration
|
||||
hosts: ceph_nodes
|
||||
become: true
|
||||
gather_facts: false
|
||||
|
||||
vars:
|
||||
backup_timestamp: "{{ lookup('pipe', 'date +%Y%m%dT%H%M%S') }}"
|
||||
backup_local_dir: "{{ playbook_dir }}/backups/{{ backup_timestamp }}"
|
||||
|
||||
tasks:
|
||||
- name: Create local backup directory
|
||||
ansible.builtin.file:
|
||||
path: "{{ backup_local_dir }}"
|
||||
state: directory
|
||||
mode: '0700'
|
||||
delegate_to: localhost
|
||||
run_once: true # noqa: run-once[task]
|
||||
|
||||
- name: Export cluster configuration
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
block:
|
||||
- name: Export ceph.conf
|
||||
ansible.builtin.command: ceph config generate-minimal-conf
|
||||
register: ceph_conf_export
|
||||
changed_when: false
|
||||
|
||||
- name: Export admin keyring
|
||||
ansible.builtin.command: ceph auth get client.admin
|
||||
register: admin_keyring_export
|
||||
changed_when: false
|
||||
|
||||
- name: Export CRUSH map (decompiled)
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
ceph osd getcrushmap -o /tmp/crushmap.bin 2>/dev/null
|
||||
crushtool -d /tmp/crushmap.bin -o /tmp/crushmap.txt
|
||||
cat /tmp/crushmap.txt
|
||||
rm -f /tmp/crushmap.bin /tmp/crushmap.txt
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: crush_export
|
||||
changed_when: false
|
||||
|
||||
- name: Export OSD map summary
|
||||
ansible.builtin.command: ceph osd dump --format json
|
||||
register: osd_dump_export
|
||||
changed_when: false
|
||||
|
||||
- name: Export monitor map
|
||||
ansible.builtin.command: ceph mon dump --format json
|
||||
register: mon_dump_export
|
||||
changed_when: false
|
||||
|
||||
- name: Export cluster config dump
|
||||
ansible.builtin.command: ceph config dump --format json
|
||||
register: config_dump_export
|
||||
changed_when: false
|
||||
|
||||
- name: Export RGW realm configuration
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
echo '{"realm":' && radosgw-admin realm get 2>/dev/null
|
||||
echo ',"zonegroup":' && radosgw-admin zonegroup get 2>/dev/null
|
||||
echo ',"zone":' && radosgw-admin zone get 2>/dev/null
|
||||
echo '}'
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: rgw_realm_export
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
|
||||
- name: Export service specs
|
||||
ansible.builtin.command: ceph orch ls --format yaml
|
||||
register: orch_services_export
|
||||
changed_when: false
|
||||
|
||||
- name: Export host list
|
||||
ansible.builtin.command: ceph orch host ls --format yaml
|
||||
register: orch_hosts_export
|
||||
changed_when: false
|
||||
|
||||
- name: Write ceph.conf backup
|
||||
ansible.builtin.copy:
|
||||
content: "{{ ceph_conf_export.stdout }}\n"
|
||||
dest: "{{ backup_local_dir }}/ceph.conf"
|
||||
mode: '0600'
|
||||
delegate_to: localhost
|
||||
|
||||
- name: Write admin keyring backup
|
||||
ansible.builtin.copy:
|
||||
content: "{{ admin_keyring_export.stdout }}\n"
|
||||
dest: "{{ backup_local_dir }}/ceph.client.admin.keyring"
|
||||
mode: '0600'
|
||||
delegate_to: localhost
|
||||
|
||||
- name: Write CRUSH map backup
|
||||
ansible.builtin.copy:
|
||||
content: "{{ crush_export.stdout }}\n"
|
||||
dest: "{{ backup_local_dir }}/crushmap.txt"
|
||||
mode: '0600'
|
||||
delegate_to: localhost
|
||||
|
||||
- name: Write OSD dump backup
|
||||
ansible.builtin.copy:
|
||||
content: "{{ osd_dump_export.stdout }}\n"
|
||||
dest: "{{ backup_local_dir }}/osd-dump.json"
|
||||
mode: '0600'
|
||||
delegate_to: localhost
|
||||
|
||||
- name: Write monitor dump backup
|
||||
ansible.builtin.copy:
|
||||
content: "{{ mon_dump_export.stdout }}\n"
|
||||
dest: "{{ backup_local_dir }}/mon-dump.json"
|
||||
mode: '0600'
|
||||
delegate_to: localhost
|
||||
|
||||
- name: Write config dump backup
|
||||
ansible.builtin.copy:
|
||||
content: "{{ config_dump_export.stdout }}\n"
|
||||
dest: "{{ backup_local_dir }}/config-dump.json"
|
||||
mode: '0600'
|
||||
delegate_to: localhost
|
||||
|
||||
- name: Write RGW realm backup
|
||||
ansible.builtin.copy:
|
||||
content: "{{ rgw_realm_export.stdout }}\n"
|
||||
dest: "{{ backup_local_dir }}/rgw-realm.json"
|
||||
mode: '0600'
|
||||
delegate_to: localhost
|
||||
|
||||
- name: Write service specs backup
|
||||
ansible.builtin.copy:
|
||||
content: "{{ orch_services_export.stdout }}\n"
|
||||
dest: "{{ backup_local_dir }}/orch-services.yaml"
|
||||
mode: '0600'
|
||||
delegate_to: localhost
|
||||
|
||||
- name: Write host list backup
|
||||
ansible.builtin.copy:
|
||||
content: "{{ orch_hosts_export.stdout }}\n"
|
||||
dest: "{{ backup_local_dir }}/orch-hosts.yaml"
|
||||
mode: '0600'
|
||||
delegate_to: localhost
|
||||
|
||||
- name: Report backup location
|
||||
ansible.builtin.debug:
|
||||
msg: "Backup written to {{ backup_local_dir }}/"
|
||||
run_once: true # noqa: run-once[task]
|
||||
@@ -0,0 +1,21 @@
|
||||
---
|
||||
# Post-boot OS baseline — ops user, packages, /etc/hosts, services.
|
||||
# Run after provisioning, before ceph_deploy.
|
||||
# Fully convergeable — re-running fixes any drift in user config,
|
||||
# packages, or host entries.
|
||||
#
|
||||
# Usage:
|
||||
# scripts/ansible-play.sh baseline.yml
|
||||
#
|
||||
# Tags:
|
||||
# --tags users ops user only
|
||||
# --tags packages podman + diagnostics only
|
||||
# --tags system /etc/hosts + services only
|
||||
|
||||
- name: OS baseline configuration
|
||||
hosts: ceph_nodes
|
||||
become: true
|
||||
gather_facts: false
|
||||
|
||||
roles:
|
||||
- baseline
|
||||
@@ -0,0 +1,21 @@
|
||||
---
|
||||
# S3 benchmark — runs in parallel on all Ceph nodes against local RGW.
|
||||
#
|
||||
# Default: 1000 objects x 16 MiB x 3 nodes = ~48 GiB PUT workload
|
||||
#
|
||||
# Usage:
|
||||
# scripts/ansible-play.sh bench.yml # default PUT
|
||||
# scripts/ansible-play.sh bench.yml -e s3bench_ops=get # GET (after PUT)
|
||||
# scripts/ansible-play.sh bench.yml -e s3bench_ops=mixed # 70/20/10 mix
|
||||
# scripts/ansible-play.sh bench.yml -e s3bench_num_objects=10000 # ~480 GiB fill
|
||||
# scripts/ansible-play.sh bench.yml -e s3bench_ops=delete # cleanup
|
||||
#
|
||||
# Results: bench/ directory (gitignored)
|
||||
|
||||
- name: S3 benchmark
|
||||
hosts: ceph_nodes
|
||||
become: true
|
||||
gather_facts: false
|
||||
|
||||
roles:
|
||||
- s3_bench
|
||||
@@ -0,0 +1,22 @@
|
||||
---
|
||||
# Deploy Ceph Tentacle cluster (cluster chosen via CEPH_ENV)
|
||||
#
|
||||
# Prerequisites:
|
||||
# 1. All nodes provisioned and running Debian 12 with bond0 networking
|
||||
# 2. SSH access via inventory.ini credentials (rendered by tofu apply)
|
||||
# 3. Storage partitioned with Ceph block.db LVs and SSD OSD partitions
|
||||
#
|
||||
# Usage:
|
||||
# scripts/ansible-play.sh deploy-ceph.yml
|
||||
#
|
||||
# To run a specific phase only:
|
||||
# scripts/ansible-play.sh deploy-ceph.yml --tags bootstrap
|
||||
# scripts/ansible-play.sh deploy-ceph.yml --tags osds
|
||||
|
||||
- name: Deploy Ceph Tentacle cluster
|
||||
hosts: ceph_nodes
|
||||
become: true
|
||||
gather_facts: true
|
||||
|
||||
roles:
|
||||
- ceph_deploy
|
||||
@@ -0,0 +1,26 @@
|
||||
---
|
||||
# Destroy Ceph Tentacle cluster — COMPLETE TEARDOWN
|
||||
#
|
||||
# This removes ALL pools, OSDs, daemons, and cluster state.
|
||||
# All data on Ceph-managed devices will be permanently lost.
|
||||
#
|
||||
# Guardrails:
|
||||
# 1. Must pass yes_destroy_ceph=true or it refuses to run
|
||||
# 2. Must pass destroy_target_domain matching cluster_domain
|
||||
# 3. Pauses for manual confirmation before destructive steps
|
||||
#
|
||||
# Usage:
|
||||
# scripts/ansible-play.sh destroy-ceph.yml \
|
||||
# -e "yes_destroy_ceph=true destroy_target_domain=dev.austin.int.futo.cloud"
|
||||
#
|
||||
# To run a specific phase only:
|
||||
# scripts/ansible-play.sh destroy-ceph.yml -e "yes_destroy_ceph=true destroy_target_domain=dev.austin.int.futo.cloud" --tags purge
|
||||
# scripts/ansible-play.sh destroy-ceph.yml -e "yes_destroy_ceph=true destroy_target_domain=dev.austin.int.futo.cloud" --tags cleanup
|
||||
|
||||
- name: Destroy Ceph Tentacle cluster
|
||||
hosts: ceph_nodes
|
||||
become: true
|
||||
gather_facts: true
|
||||
|
||||
roles:
|
||||
- ceph_destroy
|
||||
@@ -0,0 +1,378 @@
|
||||
# Adding a cluster
|
||||
|
||||
Clusters are declared in `tf/deployment/<env>/ceph/clusters.auto.tfvars`.
|
||||
Every cluster-scoped concern — inventory file, hostname, 1P item names, SSH
|
||||
key path, secrets template — is derived from that one entry. Most of what
|
||||
this walkthrough describes is editing that file and running
|
||||
`mise run tf:apply`; the rest is creating the 1P items TF expects to read
|
||||
at playbook time.
|
||||
|
||||
For the broader architecture see [docs/architecture.md](architecture.md); for
|
||||
the per-item secrets catalog see [docs/secrets.md](secrets.md); for the
|
||||
naming rules see [docs/naming.md](naming.md).
|
||||
|
||||
## What TF does vs. what you do
|
||||
|
||||
| TF (`mise run tf:apply`) | You (one-time per cluster) |
|
||||
|-------------------------------------------------------------------------|---------------------------------------------------------------------|
|
||||
| Renders `inventory.ini`, `inventory-destroy.ini`, `secrets.yml.tpl` | Create `group_vars/all/vars.yml` (cluster-wide Ansible config) |
|
||||
| Renders `inventory-provision-<profile>.ini` when `provision_profile` set | Create one `host_vars/<hostname>.yml` per node (hardware topology) |
|
||||
| Picks auto-names from the wordlist for hosts where `name = null` | Create 1P items: passwords, SSH keypair |
|
||||
| Computes hostnames, FQDNs, 1P item titles, inventory directory path | Run `scripts/install-ssh-keys.sh <cluster>` on your workstation |
|
||||
| (Future) Creates `onepassword_item` resources for passwords | Run `mise run preflight` + `mise run deploy` |
|
||||
|
||||
Every operator doing a cluster add follows the same steps — nothing in
|
||||
this walkthrough is machine- or operator-specific.
|
||||
|
||||
## Inventory directory naming
|
||||
|
||||
```
|
||||
inventories/<cluster>-<role>.<env>.<datacenter>.<provider>/
|
||||
```
|
||||
|
||||
The `<role>` segment comes from `role_in_hostname` in the TFvars (defaults
|
||||
to `ceph`). Every Ceph-project inventory grep-matches `*-ceph.*` regardless
|
||||
of datacenter or environment.
|
||||
|
||||
Existing examples:
|
||||
- `sietch-ceph.dev.austin.int/` — Austin DC, internal network, dev
|
||||
- `painbox-ceph.dev.hel.htz/` — Hetzner Helsinki, dev
|
||||
|
||||
Future environments land as siblings: `*-ceph.staging.<dc>.<provider>/`,
|
||||
`*-ceph.prod.<dc>.<provider>/`.
|
||||
|
||||
## Step-by-step
|
||||
|
||||
### 1. Choose a cluster name
|
||||
|
||||
The engineer adding the cluster picks the name. Conventions and constraints
|
||||
live in [docs/naming.md](naming.md#cluster-naming). Quick summary:
|
||||
|
||||
- **Convention:** Dune-themed (existing: `sietch`, `painbox`). Not enforced.
|
||||
- **Constraints:** lowercase, short (6–10 chars ideal), no dashes or dots,
|
||||
unique within the `yucca_tf_*` item namespace, not already a key in
|
||||
`clusters.auto.tfvars`.
|
||||
- **Cost of renaming later:** expensive (touches hostnames, 1P items,
|
||||
cephadm identity, SSH keys, DNS). Pick deliberately.
|
||||
|
||||
Host names within a cluster can be operator-declared in the TFvars or
|
||||
auto-picked from the 923-word wordlist — see [docs/naming.md](naming.md#host-naming).
|
||||
|
||||
### 2. Declare the cluster in TF
|
||||
|
||||
Edit `tf/deployment/<env>/ceph/clusters.auto.tfvars` and add an entry.
|
||||
Working example for a hypothetical `mesa` cluster at Hetzner Falkenstein:
|
||||
|
||||
```hcl
|
||||
clusters = {
|
||||
sietch = { ... }
|
||||
painbox = { ... }
|
||||
|
||||
mesa = {
|
||||
domain = "dev.fsn.htz.futo.cloud"
|
||||
environment = "dev"
|
||||
datacenter = "fsn"
|
||||
provider_code = "htz"
|
||||
role_in_hostname = "ceph"
|
||||
ansible_ssh_user = "ansible-iac" # Hetzner installimage boots as root;
|
||||
# baseline creates ansible-iac before first deploy
|
||||
ansible_ssh_key = "~/.ssh/id_ed25519_mesa"
|
||||
vault = "yucca_tf_dev" # or yucca_tf_staging / yucca_tf_prod
|
||||
provision_profile = null # Hetzner installimage; no debian-live provisioning
|
||||
hosts = [
|
||||
{ bond_ip = "<public-ip>", bootstrap = true }, # name auto-picked from wordlist
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Notes:
|
||||
|
||||
- `vault` declares which 1Password vault TF rendering will write into the
|
||||
`secrets.yml.tpl`. `yucca_tf_dev` for dev clusters, `yucca_tf_staging`
|
||||
or `yucca_tf` (prod) for their respective environments.
|
||||
- `provision_profile = "debian-live"` enables bare-metal provisioning via
|
||||
`provision.yml` (rendered `inventory-provision.ini`). Leave null for
|
||||
Hetzner installimage workflows — the post-install script uses its own
|
||||
path (`inventories/<cluster>/installimage/post-install.sh.tpl`).
|
||||
- Host `name = null` (omitted) → TF picks a stable wordlist name seeded
|
||||
per-cluster. Auto-picks don't change on subsequent applies.
|
||||
|
||||
### 3. Render the inventory + secrets template
|
||||
|
||||
```bash
|
||||
mise run tf:apply
|
||||
```
|
||||
|
||||
Or, for a non-default stack:
|
||||
|
||||
```bash
|
||||
TF_STACK_DIR=tf/deployment/<env>/ceph mise run tf:apply
|
||||
```
|
||||
|
||||
This creates (per the module's `rendering.tf`):
|
||||
|
||||
- `ansible/ceph/inventories/<cluster>-ceph.<env>.<dc>.<provider>/inventory.ini`
|
||||
- `.../inventory-destroy.ini`
|
||||
- `.../secrets.yml.tpl`
|
||||
- `.../inventory-provision-<profile>.ini` (only when `provision_profile` is set)
|
||||
|
||||
All of these are gitignored — re-run `mise run tf:apply` after any
|
||||
`clusters.auto.tfvars` change.
|
||||
|
||||
### 4. Create `group_vars/all/vars.yml`
|
||||
|
||||
Hand-maintained, committed. Copy the closer existing analogue as a starting
|
||||
point:
|
||||
|
||||
- **Bare-metal cluster:** copy from `sietch-ceph.dev.austin.int/group_vars/all/vars.yml`
|
||||
- **Hetzner/single-NIC cluster:** copy from `painbox-ceph.dev.hel.htz/group_vars/all/vars.yml`
|
||||
|
||||
```bash
|
||||
cp inventories/painbox-ceph.dev.hel.htz/group_vars/all/vars.yml \
|
||||
inventories/mesa-ceph.dev.fsn.htz/group_vars/all/vars.yml
|
||||
```
|
||||
|
||||
Edit every value. Required shape:
|
||||
|
||||
```yaml
|
||||
---
|
||||
# === Naming ===
|
||||
cluster_name: mesa
|
||||
cluster_role: ceph
|
||||
cluster_domain: dev.fsn.htz.futo.cloud
|
||||
|
||||
# === Network ===
|
||||
public_network: <subnet or public /32>
|
||||
cluster_network: <same as public for single-network topology>
|
||||
# Bonds / gateway / DNS — omit or customize per hardware
|
||||
|
||||
# === Ceph ===
|
||||
ceph_release: tentacle
|
||||
ceph_repo_url: "https://download.ceph.com/debian-{{ ceph_release }}/"
|
||||
ceph_repo_key_url: "https://download.ceph.com/keys/release.asc"
|
||||
|
||||
# === OS Provisioning ===
|
||||
admin_user: ansible-iac
|
||||
timezone: UTC
|
||||
# Used by provision.yml's post-reboot SSH verification to read the marker.
|
||||
provision_iac_ssh_key_path: "~/.ssh/id_ed25519_mesa"
|
||||
|
||||
# === Secret aliases (populated by op inject at playbook time) ===
|
||||
# These map TF-rendered vault_* names into the role-facing names the
|
||||
# playbooks consume. Add one alias per secret declared in the module's
|
||||
# secrets map (tf/shared/modules/ceph-cluster/main.tf).
|
||||
ops_password: "{{ vault_ops_password }}"
|
||||
ceph_dashboard_user: admin
|
||||
ceph_dashboard_password: "{{ vault_ceph_dashboard_password }}"
|
||||
ceph_grafana_admin_user: admin
|
||||
ceph_grafana_admin_password: "{{ vault_grafana_admin_password }}"
|
||||
ceph_rgw_s3_user_access_key: "{{ vault_s3_restic_access_key }}"
|
||||
ceph_rgw_s3_user_secret_key: "{{ vault_s3_restic_secret_key }}"
|
||||
|
||||
# === RGW ===
|
||||
ceph_rgw_realm: <cluster-name>
|
||||
ceph_rgw_zonegroup: <zonegroup>
|
||||
ceph_rgw_zone: <zone>
|
||||
|
||||
# === Storage ===
|
||||
ssd_model_pattern: "Micron_5100" # match your SSD model
|
||||
# ... (see the cluster you copied from for full hardware config)
|
||||
```
|
||||
|
||||
### 5. Create `host_vars/<hostname>.yml` per node
|
||||
|
||||
Host files are committed (per-cluster hardware topology is stable inventory
|
||||
truth — not operator preference). Use `<cluster>/host_vars/example.yml` as
|
||||
a template.
|
||||
|
||||
```bash
|
||||
CLUSTER_DIR=inventories/mesa-ceph.dev.fsn.htz
|
||||
# Look up the hostname TF picked (or declared) — visible in the rendered inventory.ini
|
||||
TF_OUTPUT=$(cat "$CLUSTER_DIR/inventory.ini")
|
||||
# Create one host_vars file per hostname_short shown in the [ceph_nodes] section
|
||||
cp "$CLUSTER_DIR/host_vars/example.yml" "$CLUSTER_DIR/host_vars/<hostname_short>.yml"
|
||||
```
|
||||
|
||||
Edit with node-specific hardware facts: `bond_ip`, SAS expander path
|
||||
prefix, SSD PHY positions, HDD-to-block.db-LV mappings. See
|
||||
[docs/hardware.md](hardware.md) for the shape.
|
||||
|
||||
Operator-local overrides (e.g., testing a workaround on one node) can go
|
||||
in `<hostname_short>.local.yml` — that suffix is gitignored.
|
||||
|
||||
### 6. Create 1Password items
|
||||
|
||||
For the target vault declared in the cluster's TFvars entry:
|
||||
|
||||
```bash
|
||||
VAULT=yucca_tf_dev # match the vault field in clusters.auto.tfvars
|
||||
CLUSTER=MESA # uppercase cluster_name
|
||||
|
||||
# Password items — 3 ending in _PASSWORD
|
||||
for role in OPS DASHBOARD GRAFANA; do
|
||||
op item create --vault "$VAULT" --category password \
|
||||
--title "${CLUSTER}_CEPH_${role}_PASSWORD" \
|
||||
--generate-password='letters,digits,32'
|
||||
done
|
||||
|
||||
# S3 service-user keys — 2 items; names already end in _KEY
|
||||
for suffix in S3_SVC_YUCCA_RESTIC_ACCESS_KEY S3_SVC_YUCCA_RESTIC_SECRET_KEY; do
|
||||
op item create --vault "$VAULT" --category password \
|
||||
--title "${CLUSTER}_CEPH_${suffix}" \
|
||||
--generate-password='letters,digits,32'
|
||||
done
|
||||
|
||||
# SSH Key item — one keypair per cluster. op CLI GENERATES the key inside 1P;
|
||||
# we never create the private key on disk first (see ADR-010).
|
||||
op item create --vault "$VAULT" \
|
||||
--category "SSH Key" \
|
||||
--title "${CLUSTER}_CEPH_ANSIBLE_IAC_SSH_KEY" \
|
||||
--ssh-generate-key=ed25519
|
||||
```
|
||||
|
||||
Verify the TF-rendered template resolves:
|
||||
|
||||
```bash
|
||||
CEPH_ENV=inventories/mesa-ceph.dev.fsn.htz/inventory.ini
|
||||
op inject -f -i "$(dirname $CEPH_ENV)/secrets.yml.tpl" -o /tmp/test-secrets.yml
|
||||
head -5 /tmp/test-secrets.yml && rm /tmp/test-secrets.yml
|
||||
```
|
||||
|
||||
Disaster-recovery items (`<CLUSTER>_CEPH_RGW_TLS_CERT`, `_RGW_TLS_KEY`,
|
||||
`_CLIENT_ADMIN_KEYRING`) are **not** created here — they're populated by
|
||||
`mise run capture` after the first successful deploy. Skipping that step
|
||||
is the most common gotcha.
|
||||
|
||||
### 7. Install the SSH keypair on your workstation
|
||||
|
||||
```bash
|
||||
scripts/install-ssh-keys.sh mesa
|
||||
```
|
||||
|
||||
The wrapper reads `private_key` and `public_key` from
|
||||
`${CLUSTER}_CEPH_ANSIBLE_IAC_SSH_KEY` and writes `~/.ssh/id_ed25519_mesa`
|
||||
(0600) + `.pub` (0644). Idempotent — re-running is safe. Every operator
|
||||
who will run plays against this cluster runs this command once on their
|
||||
workstation (or any time they wipe `~/.ssh/`).
|
||||
|
||||
The `ansible_ssh_key` path in `clusters.auto.tfvars` must match what
|
||||
`install-ssh-keys.sh` writes. If you chose a non-default filename,
|
||||
update both together (or update the mapping in `install-ssh-keys.sh`).
|
||||
|
||||
See [docs/scripts.md](scripts.md#install-ssh-keyssh) for the script
|
||||
reference and [ADR-010](adr/010-ssh-keys-in-1password.md) for the
|
||||
rationale.
|
||||
|
||||
### 8. Preflight
|
||||
|
||||
```bash
|
||||
CEPH_ENV=inventories/mesa-ceph.dev.fsn.htz/inventory.ini mise run preflight
|
||||
```
|
||||
|
||||
(Inline-prefix form — `export CEPH_ENV=...` then `mise run preflight`
|
||||
does NOT work; mise's `[env]` block strips shell exports. See
|
||||
[docs/scripts.md "Setting CEPH_ENV"](scripts.md).)
|
||||
|
||||
Verifies: TF artifacts present, 1P session live, `op inject` resolves
|
||||
the template, SSH reachable, Python 3 on targets.
|
||||
|
||||
### 9. Deploy
|
||||
|
||||
For Hetzner installimage clusters, run the installimage flow first
|
||||
(out-of-band; see [runbooks/painbox-reprovision.md](runbooks/painbox-reprovision.md)
|
||||
for the pattern). For Austin bare-metal clusters, run `provision.yml`
|
||||
first (boot into the live image, then `scripts/ansible-play.sh
|
||||
provision.yml -e confirm_wipe=true` with
|
||||
`CEPH_ENV=.../inventory-provision.ini`). Then:
|
||||
|
||||
```bash
|
||||
mise run deploy
|
||||
```
|
||||
|
||||
Every task invocation goes through `scripts/ansible-play.sh`, which
|
||||
`op inject`s the secrets template into a short-lived tmpfile and passes
|
||||
it as `--extra-vars @<tmpfile>`.
|
||||
|
||||
### 10. Capture the DR belt-and-suspenders items
|
||||
|
||||
After the first successful deploy:
|
||||
|
||||
```bash
|
||||
mise run capture
|
||||
```
|
||||
|
||||
This reads `/etc/ceph/rgw-ssl.crt`, `/etc/ceph/rgw-ssl.key`, and
|
||||
`/etc/ceph/ceph.client.admin.keyring` from the bootstrap node and upserts
|
||||
them as Document items in the cluster's vault
|
||||
(`<CLUSTER>_CEPH_RGW_TLS_CERT`, `_RGW_TLS_KEY`, `_CLIENT_ADMIN_KEYRING`). Safe
|
||||
to re-run — updates in place on content drift.
|
||||
|
||||
## How `CEPH_ENV` works
|
||||
|
||||
`CEPH_ENV` points to the inventory file, not the directory:
|
||||
|
||||
```bash
|
||||
# Correct
|
||||
CEPH_ENV=inventories/mesa-ceph.dev.fsn.htz/inventory.ini
|
||||
|
||||
# Wrong — directory mode loads every .ini including destroy inventory
|
||||
CEPH_ENV=inventories/mesa-ceph.dev.fsn.htz/
|
||||
```
|
||||
|
||||
Default is set in `.mise.toml` (`sietch` in dev). Override per-command:
|
||||
|
||||
```bash
|
||||
CEPH_ENV=inventories/painbox-ceph.dev.hel.htz/inventory.ini mise run status
|
||||
```
|
||||
|
||||
`scripts/ansible-play.sh` derives the secrets template path from `CEPH_ENV`
|
||||
(same directory, `secrets.yml.tpl`).
|
||||
|
||||
## What lives where in the repo
|
||||
|
||||
| Committed | Gitignored (TF-rendered or operator-local) |
|
||||
|------------------------------------------------------|--------------------------------------------|
|
||||
| `clusters.auto.tfvars` | `inventories/*/inventory.ini` |
|
||||
| `inventories/<cluster>/group_vars/all/vars.yml` | `inventories/*/inventory-provision.ini` |
|
||||
| `inventories/<cluster>/host_vars/<hostname>.yml` | `inventories/*/inventory-destroy.ini` |
|
||||
| `inventories/<cluster>/installimage/*.tpl` | `inventories/*/secrets.yml.tpl` |
|
||||
| | `inventories/*/host_vars/*.local.yml` |
|
||||
| | `inventories/*/installimage/post-install.sh` |
|
||||
|
||||
No secrets are ever committed. The `.tpl` file contains `op://` references
|
||||
only; `op inject` resolves them at play time into a 0600 tmpfile that's
|
||||
trap-cleaned on exit.
|
||||
|
||||
## Common gotchas
|
||||
|
||||
- **Forgot step 10 (`mise run capture`)** — DR items are missing in 1P.
|
||||
Running `capture` after the fact works; it just needs the bootstrap
|
||||
node's filesystem intact.
|
||||
- **`ansible_ssh_key` path mismatch** between `clusters.auto.tfvars` and
|
||||
`install-ssh-keys.sh` — key installed under a different name than what
|
||||
the inventory expects. Keep them aligned.
|
||||
- **Fingerprint mismatch on `install-ssh-keys.sh`** — happens after a
|
||||
key rotation if you haven't moved the old key aside. Follow the
|
||||
`mv ~/.ssh/id_ed25519_<cluster>{,.$(date +%Y%m%d).bak}` path in the
|
||||
wrapper's error message.
|
||||
- **`host_vars/` out of date after `tofu apply` re-picks a wordlist
|
||||
name** — auto-names are stable across applies, but if you add hosts
|
||||
at positions other than the tail, shuffled names may shift. Add new
|
||||
hosts at the end of the `hosts = [...]` list to keep existing
|
||||
hostnames stable.
|
||||
- **Painbox-style Hetzner clusters and `provision_profile`** — leave
|
||||
`provision_profile = null` so TF doesn't render the debian-live
|
||||
inventory. The Hetzner installimage flow has its own post-install
|
||||
script under `installimage/`, not Ansible-driven.
|
||||
|
||||
## See also
|
||||
|
||||
- [architecture.md §4 (Terraform)](architecture.md) — what TF owns
|
||||
- [secrets.md](secrets.md) — per-item catalog + vault selection
|
||||
- [naming.md](naming.md) — cluster + host naming conventions
|
||||
- [hardware.md](hardware.md) — `host_vars/` shape, network topology
|
||||
- [scripts.md](scripts.md) — wrapper reference
|
||||
- [ADR-009](adr/009-tf-first-op-inject-over-vault-password-sh.md) — why TF is authoritative
|
||||
- [ADR-010](adr/010-ssh-keys-in-1password.md) — why SSH keys live in 1P
|
||||
- [runbooks/painbox-reprovision.md](runbooks/painbox-reprovision.md) — Hetzner-specific reprovisioning pattern
|
||||
@@ -0,0 +1,116 @@
|
||||
# Adding a role
|
||||
|
||||
How to create a new Ansible role in this project. The heavy lifting is in
|
||||
the existing roles — this doc is the skeleton + wiring + pre-submit
|
||||
checklist; copy an exemplar for idioms.
|
||||
|
||||
For code-level patterns (idempotency, `changed_when`, handlers, secrets
|
||||
handling, shell conventions), see [patterns.md](patterns.md). For how roles
|
||||
compose into the overall pipeline, see
|
||||
[architecture.md §6](architecture.md).
|
||||
|
||||
## Skeleton
|
||||
|
||||
```
|
||||
roles/<role_name>/
|
||||
├── defaults/main.yml required — every variable with a default + comment
|
||||
├── meta/main.yml required — author, license, Ansible version
|
||||
├── tasks/main.yml required — imports sub-task files, tagged
|
||||
├── handlers/main.yml if the role restarts/reloads services
|
||||
├── templates/*.j2 Jinja2 templates
|
||||
└── molecule/default/ optional — test scenario
|
||||
```
|
||||
|
||||
Role names use `snake_case` (`ceph_deploy`, `os_tuning`).
|
||||
|
||||
**Cluster-level variables** use the `ceph_` prefix (overridable in
|
||||
`group_vars/all/vars.yml`). **Role-internal variables** use the
|
||||
`<role_name>_` prefix. The `.ansible-lint` config skips
|
||||
`var-naming[no-role-prefix]` because all roles share the `ceph_` prefix
|
||||
for cluster-level settings — this is intentional.
|
||||
|
||||
## Exemplars to copy from
|
||||
|
||||
| Copy this when you're writing... | Role |
|
||||
|---|---|
|
||||
| A phased deployment with tags | `ceph_deploy` |
|
||||
| Modular sub-task files with header comments | `baseline` |
|
||||
| Kernel/sysctl values with units documented | `os_tuning` |
|
||||
| Per-device-class settings + feature toggles | `hardware_tuning` |
|
||||
| Templated firewall config with opt-in features | `security` |
|
||||
| Ceph CLI shell tasks with `changed_when` patterns | `ceph_tuning` |
|
||||
|
||||
Every existing role has a header comment block at the top of
|
||||
`defaults/main.yml` explaining its scope — open one and mirror the shape.
|
||||
|
||||
## Wiring into the pipeline
|
||||
|
||||
### 1. Playbook wrapper
|
||||
|
||||
Create a top-level playbook (e.g. `my-feature.yml`) that imports the role:
|
||||
|
||||
```yaml
|
||||
---
|
||||
# Brief description + usage tags.
|
||||
- name: My feature
|
||||
hosts: ceph_nodes
|
||||
become: true
|
||||
roles:
|
||||
- my_role_name
|
||||
```
|
||||
|
||||
### 2. `site.yml` (if part of full deploy)
|
||||
|
||||
Insert in the correct dependency position in `site.yml`:
|
||||
|
||||
```yaml
|
||||
- import_playbook: my-feature.yml
|
||||
```
|
||||
|
||||
Order matters — see [architecture.md §6 "Why this order matters"](architecture.md).
|
||||
|
||||
### 3. `mise run deploy` (if part of full deploy)
|
||||
|
||||
Add to the `deploy` task in yucca-root `.mise/config.toml` (or
|
||||
`ansible/ceph/.mise.toml` if the task lives there). Always invoke via
|
||||
the wrapper so secrets resolve:
|
||||
|
||||
```toml
|
||||
echo "=== My feature ===" && scripts/ansible-play.sh my-feature.yml
|
||||
```
|
||||
|
||||
### 4. Optional: standalone mise task
|
||||
|
||||
For playbooks useful to run independently (like `bench`, `drift`,
|
||||
`status`):
|
||||
|
||||
```toml
|
||||
[tasks.my-feature]
|
||||
description = "One-line description"
|
||||
run = "scripts/ansible-play.sh my-feature.yml"
|
||||
```
|
||||
|
||||
See [scripts.md](scripts.md) for the wrapper reference.
|
||||
|
||||
## Molecule test (optional)
|
||||
|
||||
Not every role needs one — the `ceph_deploy` role is the only one with
|
||||
a full scenario today. If you're writing something non-trivial, copy
|
||||
`roles/ceph_deploy/molecule/default/` as a starting point.
|
||||
|
||||
## Pre-submit checklist
|
||||
|
||||
Before opening a PR, verify:
|
||||
|
||||
- [ ] `mise run lint` passes clean (yamllint + ansible-lint + shellcheck)
|
||||
- [ ] `mise run check` passes syntax-check with the role's playbook
|
||||
- [ ] Every variable in `defaults/main.yml` has a comment explaining what it does
|
||||
- [ ] All modules use FQCN (`ansible.builtin.apt`, not `apt`)
|
||||
- [ ] All `shell`/`command` tasks have `changed_when`
|
||||
- [ ] All `shell` tasks set `args.executable: /bin/bash` and include `set -o pipefail`
|
||||
- [ ] Secret-handling tasks have `no_log: true`
|
||||
- [ ] `meta/main.yml` has author, license, description, `min_ansible_version: "2.19"`
|
||||
- [ ] Tags on `import_tasks` in `main.yml`
|
||||
- [ ] Playbook wrapper exists at `ansible/ceph/` root
|
||||
- [ ] Role is added to `site.yml` in the correct position (if part of full deploy)
|
||||
- [ ] Anti-patterns in [patterns.md §Anti-patterns](patterns.md) not violated
|
||||
@@ -0,0 +1,49 @@
|
||||
# ADR-001: cephadm + Ansible Wrapper Roles over ceph-ansible
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The official `ceph-ansible` project targets Ceph Quincy and older releases. Ceph
|
||||
Tentacle (the release this cluster runs) is a cephadm-native release where
|
||||
`ceph-ansible` is deprecated and no longer tested. Additionally, `ceph-ansible`
|
||||
is a large, opinionated framework that owns the entire node lifecycle -- from
|
||||
package installation to OSD creation -- making it difficult to compose with our
|
||||
own provisioning pipeline (debootstrap, baseline role, security hardening).
|
||||
|
||||
We needed a deployment approach that:
|
||||
|
||||
- Works with Ceph Tentacle (cephadm-native release).
|
||||
- Gives us explicit control over each phase (bootstrap, join, OSD creation,
|
||||
RGW, monitoring) so we can debug and re-run individual stages.
|
||||
- Stays composable with our existing Ansible roles for OS provisioning,
|
||||
baseline configuration, and security.
|
||||
|
||||
## Decision
|
||||
|
||||
We use **cephadm directly**, wrapped in thin Ansible task files inside the
|
||||
`ceph_deploy` role. Each deployment phase is a separate task file
|
||||
(`bootstrap.yml`, `join.yml`, `osds.yml`, `rgw.yml`, etc.) that calls
|
||||
`cephadm` and `ceph` CLI commands with idempotency guards.
|
||||
|
||||
The Ansible layer handles orchestration (ordering, host targeting, variable
|
||||
interpolation, idempotency checks) while cephadm handles container management,
|
||||
daemon lifecycle, and config distribution. Comments in `bootstrap.yml` note
|
||||
that cephadm automatically distributes `ceph.conf`, admin keyrings, and SSH
|
||||
keys during `ceph orch host add` -- we lean on that rather than reimplementing
|
||||
distribution logic.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Positive:** Full compatibility with Ceph Tentacle and future releases. Each
|
||||
phase is independently re-runnable. Task files are small and auditable. No
|
||||
dependency on an external Ansible Galaxy role with its own release cycle.
|
||||
- **Positive:** Operators can `--tags osds` to re-run just OSD creation after a
|
||||
failure, or `--tags rgw` to redeploy the gateway layer independently.
|
||||
- **Negative:** We own the idempotency logic (e.g., `pvs | grep ceph` checks,
|
||||
`stat` on `/etc/ceph/ceph.conf`). Upstream ceph-ansible handled this
|
||||
automatically, but it also hid failures behind abstraction layers.
|
||||
- **Negative:** New Ceph features require manual task-file additions rather than
|
||||
a Galaxy role version bump.
|
||||
@@ -0,0 +1,47 @@
|
||||
# ADR-002: Explicit OSD-to-Disk Mapping in host_vars
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
Ceph can auto-discover available drives via `ceph orch apply osd --all-available-devices`.
|
||||
This is convenient but dangerous in a mixed-media cluster with complex disk
|
||||
layouts. Our Austin nodes have:
|
||||
|
||||
- 12 front-bay HDDs assigned to OSD data, each needing a specific block.db LV
|
||||
on one of two rear-bay SSDs.
|
||||
- 2 rear-bay SSDs that are multi-purpose: partition 1-4 for OS (mdraid+LVM),
|
||||
partition 5 for block.db LVs, partition 6 for SSD OSDs.
|
||||
- Device paths that use `/dev/disk/by-path/` with SAS expander PHY addresses
|
||||
for slot stability (replacing a drive in the same bay keeps the same path).
|
||||
|
||||
Auto-discovery cannot express the HDD-to-SSD-db mapping. It also risks
|
||||
claiming OS partitions or db partitions as OSD data. A mismap during automated
|
||||
discovery would silently create OSDs without block.db acceleration, or worse,
|
||||
destroy the OS volume.
|
||||
|
||||
## Decision
|
||||
|
||||
Every OSD is explicitly listed in each node's `host_vars` file. HDD OSDs
|
||||
specify both the PHY slot (`path_phy`) and the exact block.db LV (`db`). SSD
|
||||
OSDs specify the PHY slot and partition number. The `osds.yml` task file
|
||||
iterates these lists with `loop:`, resolving each entry to a
|
||||
`/dev/disk/by-path/` stable path.
|
||||
|
||||
Each entry is idempotent: the task checks `pvs` for an existing Ceph PV on the
|
||||
resolved device and skips creation if one exists. Empty bays (missing block
|
||||
device) are also skipped gracefully.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Positive:** Every OSD-to-disk-to-db mapping is version-controlled and
|
||||
auditable. Drive replacements are tracked with serial number comments in
|
||||
host_vars. No risk of accidental OSD creation on OS or db partitions.
|
||||
- **Positive:** Slot-stable `by-path` addressing means a replacement drive in
|
||||
the same bay inherits the same OSD definition -- no host_vars edit needed.
|
||||
- **Negative:** Adding a new node requires writing a complete host_vars file
|
||||
with all PHY mappings. This is manual but happens rarely (new hardware).
|
||||
- **Negative:** Cannot scale to hundreds of heterogeneous nodes without
|
||||
templating. Acceptable for our 3-node (scaling to ~10) cluster size.
|
||||
@@ -0,0 +1,51 @@
|
||||
# ADR-003: Baseline Role Split from Provisioning
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The `provision_host` role runs inside a chroot on a live image. It must create
|
||||
the minimum viable user (`ansible-iac`) so Ansible can connect after first
|
||||
reboot. Originally, the ops user, diagnostic packages, `/etc/hosts`, and
|
||||
service enablement were also done in chroot.
|
||||
|
||||
This caused problems:
|
||||
|
||||
- Chroot operations are fragile -- bind mounts for `/dev`, `/proc`, `/sys`
|
||||
must be set up and torn down correctly. More chroot work means more failure
|
||||
surface during provisioning.
|
||||
- The ops user password comes from 1Password (resolved at play time via
|
||||
`op inject` — see ADR-009). Embedding it in the chroot phase means
|
||||
provisioning depends on the operator's 1P session being live during
|
||||
install, which complicates the live-image environment.
|
||||
- Post-provision drift (stale `/etc/hosts`, missing packages, changed
|
||||
passwords) required re-provisioning from the live image to fix. There was
|
||||
no way to converge config on a running system.
|
||||
|
||||
## Decision
|
||||
|
||||
The `provision_host` role now creates only `ansible-iac` (uid 1000, locked
|
||||
password, key-only SSH, NOPASSWD sudo) inside chroot. Everything else moved
|
||||
to a separate **baseline** role that runs post-boot via the normal Ansible
|
||||
pipeline:
|
||||
|
||||
- `users.yml` -- ops user with op-injected password, sudo config
|
||||
- `packages.yml` -- podman ecosystem, diagnostic tools
|
||||
- `system.yml` -- `/etc/hosts` template, service enablement, timezone
|
||||
|
||||
The baseline role is convergeable: re-running it on a live system corrects
|
||||
drift without re-provisioning.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Positive:** Provisioning is faster and less fragile -- fewer chroot
|
||||
operations, no vault dependency during OS install.
|
||||
- **Positive:** Password changes, package additions, and `/etc/hosts` updates
|
||||
are applied by re-running `baseline.yml` against running nodes. No reboot
|
||||
or live-image cycle needed.
|
||||
- **Positive:** Clear responsibility boundary: provision_host owns
|
||||
"bare metal to bootable OS", baseline owns "bootable OS to operational node".
|
||||
- **Negative:** Two-step initial setup (provision, then baseline) instead of
|
||||
one. Mitigated by the `site.yml` playbook which chains both.
|
||||
@@ -0,0 +1,50 @@
|
||||
# ADR-004: Multi-Inventory over Multi-Repo
|
||||
|
||||
## Status
|
||||
|
||||
Accepted. The core decision (one repo, per-cluster inventory dirs) stands.
|
||||
Refined by ADR-009 — inventory files are now TF-rendered rather than
|
||||
hand-authored, and the `vault-password.sh`/`secrets-init.sh` mechanism
|
||||
referenced below has been replaced by `op inject` + `scripts/ansible-play.sh`.
|
||||
|
||||
## Context
|
||||
|
||||
We operate multiple Ceph clusters across different sites and purposes:
|
||||
|
||||
- `sietch-ceph.dev.austin.int` -- 3-node dev cluster in Austin datacenter
|
||||
- `painbox-ceph.dev.hel.htz` -- single-node dev cluster at Hetzner Helsinki
|
||||
|
||||
Each cluster has different hardware, networking, SSH access, and credentials.
|
||||
The common approaches are: (a) one repository per cluster, or (b) one
|
||||
repository with per-cluster inventory directories.
|
||||
|
||||
Separate repos cause role drift -- a fix to `ceph_deploy` in one repo must be
|
||||
cherry-picked to every other repo. Shared roles via Git submodules or Galaxy
|
||||
add dependency management overhead. The clusters share the same roles, playbooks,
|
||||
and scripts; only the inventory data (host lists, IPs, credentials, host_vars)
|
||||
differs.
|
||||
|
||||
## Decision
|
||||
|
||||
One repository with a top-level `inventories/` directory containing one
|
||||
subdirectory per cluster. Each cluster directory has its own `inventory.ini`,
|
||||
`group_vars/`, and `host_vars/`. The `ansible.cfg` has no default inventory --
|
||||
the active cluster is selected via `-i inventories/<cluster>/inventory.ini` or
|
||||
the `CEPH_ENV` environment variable.
|
||||
|
||||
Scripts like `vault-password.sh` and `secrets-init.sh` derive the cluster
|
||||
identity from `CEPH_ENV` (parsing the inventory path to extract the cluster
|
||||
ID), so they work against any cluster without hardcoded names.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Positive:** Role and playbook changes apply to all clusters immediately.
|
||||
No cherry-picking or submodule syncing.
|
||||
- **Positive:** Cluster-specific config (IPs, OSD maps, SSH keys, vault
|
||||
passwords) is cleanly isolated in per-cluster inventory directories.
|
||||
- **Positive:** Scripts auto-detect cluster context from `CEPH_ENV`, so the
|
||||
same tooling works for Austin and Hetzner without modification.
|
||||
- **Negative:** A broken role change affects all clusters. Mitigated by testing
|
||||
on painbox (Hetzner) before applying to sietch (Austin production path).
|
||||
- **Negative:** Repository grows with each cluster. Acceptable at our scale
|
||||
(inventory data is small).
|
||||
@@ -0,0 +1,58 @@
|
||||
# ADR-005: 1Password CLI over HashiCorp Vault
|
||||
|
||||
## Status
|
||||
|
||||
Superseded by [ADR-009](./009-tf-first-op-inject-over-vault-password-sh.md).
|
||||
|
||||
The core decision (1Password over HashiCorp Vault) stands — the team already
|
||||
runs on 1P for credentials, and a self-hosted Vault server wasn't justified
|
||||
for this scale. What's changed is the *mechanism*: the `vault-password.sh`
|
||||
script + `ansible-vault` + `vault.yml` pipeline described below has been
|
||||
retired in favor of TF-provisioned 1P items + `op inject` at playbook time.
|
||||
See ADR-009 for the current implementation. This document is preserved for
|
||||
historical context on the Vault-vs-1P choice.
|
||||
|
||||
## Context
|
||||
|
||||
The Ansible playbooks need secrets: vault passwords for encrypted vars files,
|
||||
dashboard credentials, ops user passwords, and S3 keys. The standard options
|
||||
are:
|
||||
|
||||
- `ansible-vault` with a password file or `--ask-vault-pass` -- simple but
|
||||
the vault password itself needs to live somewhere (plaintext file, env var,
|
||||
or manual entry every run).
|
||||
- HashiCorp Vault -- powerful but requires its own infrastructure (server,
|
||||
unsealing, token management, HA). Overkill for a small team managing a few
|
||||
clusters.
|
||||
- 1Password -- already used by the team for credential management. Has a CLI
|
||||
(`op`) with service account tokens for CI and desktop app integration for
|
||||
developer workstations.
|
||||
|
||||
## Decision
|
||||
|
||||
Secrets are managed in 1Password. The `vault-password.sh` script is the
|
||||
Ansible `vault_password_file`. It fetches the vault password from a 1Password
|
||||
item, trying four auth methods in order:
|
||||
|
||||
1. Service account token (`OP_SERVICE_ACCOUNT_TOKEN`) for CI/headless.
|
||||
2. Desktop app integration (polkit/YubiKey) for developer workstations.
|
||||
3. Interactive prompt as TTY fallback.
|
||||
4. Dummy password for lint/syntax-check (no TTY, no `op`).
|
||||
|
||||
A companion `secrets-init.sh` script bootstraps all 1Password items for a new
|
||||
cluster (vault password, dashboard login, grafana login, ops user) with
|
||||
auto-generated passwords and multi-URL entries for browser autofill. Items
|
||||
are stored in a dedicated "Yucca" vault and named by cluster FQDN.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Positive:** No additional infrastructure to manage. 1Password is already
|
||||
the team's credential store -- secrets live alongside other org credentials.
|
||||
- **Positive:** YubiKey/biometric unlock on workstations means no plaintext
|
||||
password files on disk. Service account tokens provide headless CI access.
|
||||
- **Positive:** `secrets-init.sh` makes cluster credential bootstrapping
|
||||
repeatable -- new clusters get consistent 1Password items automatically.
|
||||
- **Negative:** Hard dependency on 1Password CLI (`op`). Mitigated by the
|
||||
interactive and dummy fallbacks in `vault-password.sh`.
|
||||
- **Negative:** 1Password is a SaaS dependency. Acceptable trade-off vs.
|
||||
self-hosting HashiCorp Vault for a small operations team.
|
||||
@@ -0,0 +1,48 @@
|
||||
# ADR-006: Self-Signed TLS for RGW
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The RadosGW (S3) frontend needs TLS. The options are:
|
||||
|
||||
- **Let's Encrypt** -- requires public DNS and HTTP-01 or DNS-01 challenge
|
||||
validation. Our clusters sit on private networks (10.10.10.0/24) with no
|
||||
public DNS records and no inbound internet access. ACME is not viable
|
||||
without a DNS provider API and split-horizon DNS.
|
||||
- **Private CA** -- proper chain of trust, but requires CA infrastructure
|
||||
(key ceremony, CRL/OCSP, distribution of the CA cert to every client).
|
||||
Significant operational overhead for an internal dev/staging cluster.
|
||||
- **Self-signed certs** -- simple to generate, no external dependencies. S3
|
||||
clients (aws-cli, boto3, rclone) all support disabling cert verification
|
||||
or trusting a custom cert.
|
||||
|
||||
## Decision
|
||||
|
||||
RGW uses a self-signed certificate generated by `openssl` during deployment.
|
||||
The cert is created on the bootstrap node with a 10-year validity period and
|
||||
includes SANs for:
|
||||
|
||||
- The canonical DNS name and a wildcard under it (for virtual-hosted buckets).
|
||||
- Per-node FQDNs and bond IPs (so direct-host and IP-based access validates).
|
||||
|
||||
The combined PEM (cert + key) is embedded in the cephadm RGW service spec.
|
||||
Cephadm distributes it to every RGW daemon container. The dashboard's RGW API
|
||||
SSL verification is disabled to accommodate the self-signed cert.
|
||||
|
||||
Rotation is manual: delete the cert files on the bootstrap node and re-run the
|
||||
role. The `creates:` guard makes the generation idempotent.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Positive:** Zero external dependencies. Works on air-gapped and private
|
||||
networks. No DNS provider API, no ACME client, no CA infrastructure.
|
||||
- **Positive:** 10-year validity avoids renewal automation for a dev cluster.
|
||||
SANs cover all access patterns (DNS, hostname, IP, virtual-hosted buckets).
|
||||
- **Negative:** S3 clients must either disable TLS verification or import the
|
||||
self-signed cert. Documented in s3-integration.md.
|
||||
- **Negative:** Not suitable for production clusters serving untrusted clients.
|
||||
A future production ADR will revisit this with a private CA or ACME via
|
||||
DNS-01 challenge.
|
||||
@@ -0,0 +1,51 @@
|
||||
# ADR-007: nftables over iptables
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
Ceph nodes expose many services: MON (3300, 6789), OSD (6800-7568), MGR/dashboard
|
||||
(8443), RGW/S3 (443), Prometheus (9095), Grafana (3000), Alertmanager (9093),
|
||||
node-exporter (9100), plus optional iSCSI and NFS ports. Without a firewall,
|
||||
all of these are reachable from any source on the network.
|
||||
|
||||
The two main Linux firewall frameworks are:
|
||||
|
||||
- **iptables/ip6tables** -- legacy, being replaced upstream. Uses separate
|
||||
tables for IPv4/IPv6. Debian 12 still ships it but marks it deprecated.
|
||||
- **nftables** -- the successor. Single framework for IPv4/IPv6/ARP. Native
|
||||
in the kernel since 3.13. Debian 12's default backend for `iptables` is
|
||||
already `nft`.
|
||||
|
||||
## Decision
|
||||
|
||||
We use **nftables** with a Jinja2-templated ruleset (`nftables.conf.j2`)
|
||||
managed by the security role. The template generates a complete `inet filter`
|
||||
table with:
|
||||
|
||||
- Default-drop input policy.
|
||||
- Established/related connection tracking.
|
||||
- SSH open to all (or restricted to trusted networks, controlled by a variable).
|
||||
- RGW/S3 port open to all (public-facing service).
|
||||
- All other Ceph services restricted to `ceph_firewall_trusted_networks`.
|
||||
- Rate-limited logging of dropped packets for diagnostics.
|
||||
- Optional iSCSI and NFS blocks gated by boolean variables.
|
||||
|
||||
The ruleset is rendered from inventory variables (port numbers, trusted
|
||||
networks, feature flags), making it consistent across nodes and clusters
|
||||
without manual rule management.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Positive:** Single `inet` table covers both IPv4 and IPv6. No dual-stack
|
||||
rule duplication.
|
||||
- **Positive:** Template-driven rules are version-controlled, auditable, and
|
||||
consistent across all nodes. Adding a new service means adding one variable
|
||||
and one template block.
|
||||
- **Positive:** `flush ruleset` at the top ensures convergence -- re-applying
|
||||
the template replaces the entire ruleset atomically.
|
||||
- **Negative:** Operators familiar only with `iptables` syntax need to learn
|
||||
nftables. Mitigated by the template being well-commented and the Debian 12
|
||||
ecosystem defaulting to nft.
|
||||
@@ -0,0 +1,65 @@
|
||||
# ADR-008: debootstrap from Live Image over Preseed/Autoinstall
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
Austin nodes are bare-metal Dell servers that need Debian 12 installed from
|
||||
scratch. The standard approaches are:
|
||||
|
||||
- **Preseed/Autoinstall** -- unattended Debian installer driven by a preseed
|
||||
file. Requires PXE boot infrastructure or a custom ISO. The installer is a
|
||||
black box: partition layout is expressed in preseed's declarative syntax,
|
||||
which cannot handle our complex disk layout (mdraid-1 across two SSDs,
|
||||
5 partitions per SSD for ESP/boot/swap/root/ceph-db, plus SSD OSD
|
||||
partitions).
|
||||
- **debootstrap from a live image** -- boot a Debian live USB, run
|
||||
`debootstrap` to install the base OS into a prepared mount point. Full
|
||||
scripting control over partitioning, mdraid, LVM, and chroot configuration.
|
||||
|
||||
Our disk layout requires:
|
||||
|
||||
1. Detecting exactly 2 SSDs by model pattern, validating they are non-rotational
|
||||
and above a minimum size.
|
||||
2. GPT partitioning with 5-6 partitions per SSD (ESP, boot, swap, root LVM,
|
||||
ceph block.db, ceph SSD OSD).
|
||||
3. mdraid-1 mirrors across matching partitions on both SSDs.
|
||||
4. LVM on the root mdraid for flexible volume management.
|
||||
|
||||
Preseed cannot express step 1 (hardware validation with abort-on-failure) or
|
||||
step 2 (partition 5-6 reserved for Ceph with specific sizes). Custom
|
||||
partitioning in preseed uses `partman` recipes, which are notoriously fragile
|
||||
and poorly documented for non-standard layouts.
|
||||
|
||||
## Decision
|
||||
|
||||
We boot nodes from a Debian 12 live image (USB stick via iDRAC virtual media),
|
||||
then run the `provision_host` role via Ansible. The role:
|
||||
|
||||
1. Validates the live-image environment (UEFI, correct SSDs, sizes, non-rotational).
|
||||
2. Partitions, creates mdraid arrays, sets up LVM -- all via shell commands
|
||||
with full error handling and idempotency guards.
|
||||
3. Runs `debootstrap --arch amd64 bookworm /mnt` to install the base OS.
|
||||
4. Templates hostname, hosts, network, fstab, mdadm.conf inside the chroot.
|
||||
5. Installs packages, creates the `ansible-iac` user, installs GRUB.
|
||||
6. Writes a provisioning marker, mirrors the ESP, and reboots.
|
||||
|
||||
The entire process has block/rescue error handling: on failure, bind mounts
|
||||
are cleaned up so the next run starts clean. A marker file makes completed
|
||||
provisioning idempotent -- re-running skips directly to unmount and reboot.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Positive:** Full scripting control over disk layout. Complex partitioning,
|
||||
mdraid, and multi-purpose SSD partitions are straightforward shell commands,
|
||||
not preseed recipes.
|
||||
- **Positive:** Hardware validation before any destructive operation. The role
|
||||
aborts if SSDs don't match expectations (wrong model, too small, rotational).
|
||||
- **Positive:** Idempotent with marker-driven resume. Partially failed runs
|
||||
can be safely retried.
|
||||
- **Negative:** Requires a live-image boot mechanism (iDRAC virtual media or
|
||||
physical USB). Cannot provision purely over the network without PXE.
|
||||
- **Negative:** More Ansible code to maintain than a preseed file. Justified
|
||||
by the disk layout complexity that preseed cannot express.
|
||||
@@ -0,0 +1,112 @@
|
||||
# ADR-009: TF-First Authority + `op inject` over ansible-vault
|
||||
|
||||
## Status
|
||||
|
||||
Accepted (2026-04-22). Supersedes the implementation portions of [ADR-005](./005-1password-over-hashicorp-vault.md);
|
||||
the "1Password over HashiCorp Vault" decision itself stands.
|
||||
|
||||
## Context
|
||||
|
||||
ADR-005 landed `vault-password.sh` (4-tier auth fallback) + `ansible-vault`
|
||||
+ committed-encrypted `vault.yml` + `secrets-sync.sh` as the secrets
|
||||
mechanism. Over time this accumulated friction:
|
||||
|
||||
1. **Silent-failure footgun**: `vault-password.sh` fell back to a dummy
|
||||
password when 1Password was unavailable and no TTY was present. This
|
||||
masked real auth failures — plays ran with wrong secrets until a task
|
||||
that used them downstream blew up with a confusing error.
|
||||
2. **Two source-of-truth problem**: secrets lived both in 1Password (master)
|
||||
and in `vault.yml` (cache). `secrets-sync.sh` reconciled them but the
|
||||
reconciliation step was easy to forget.
|
||||
3. **Cluster identity authority**: `assign-names.py`, `vault-password.sh`,
|
||||
`secrets-init.sh` all derived cluster identity from the inventory
|
||||
directory path — operator-authored. Naming drift between scripts was a
|
||||
recurring gotcha.
|
||||
4. **Move to yucca monorepo**: migrating `sietch-ceph-dev-austin-int-futo-cloud`
|
||||
into the `immich-apps/yucca` monorepo made the existing ad-hoc bash
|
||||
scripting look out of place next to the Immich devtools pattern already
|
||||
in use for other Futo infra (TF + 1P, `op run --env-file`, terragrunt).
|
||||
5. **Painbox reprovision** surfaced that the Ansible layer's implicit
|
||||
cluster-identity (host names, inventory paths) was hard to change — TF
|
||||
as the authority makes identity a declarative input.
|
||||
|
||||
## Decision
|
||||
|
||||
**TF is the authority for cluster identity and 1P-item lifecycle; Ansible
|
||||
is a consumer.** Specifically:
|
||||
|
||||
1. **Cluster + host identity** declared in `tf/deployment/<env>/<stack>/clusters.auto.tfvars`.
|
||||
TF renders `inventory.ini`, `inventory-destroy.ini`, `inventory-provision.ini`,
|
||||
and `secrets.yml.tpl` per cluster via `templatefile()` + `local_file`
|
||||
resources.
|
||||
2. **1P items** live in the `yucca_tf_*` team-shared vaults. Item names
|
||||
follow `<CLUSTER>_CEPH_<ROLE>_*` (SHOUTY_SNAKE_CASE, `CEPH` hardcoded —
|
||||
project-scoped, not hostname-role-scoped). TF-managed `onepassword_item`
|
||||
resources are currently **dormant** in `secrets.tf.disabled` — items are
|
||||
created today via the `op` CLI (operator runs `op item create` once per
|
||||
cluster; see `docs/adding-a-cluster.md` §6). Re-enable once the
|
||||
dedicated ceph-scoped service account replaces the org-wide superuser SA.
|
||||
3. **Ansible consumption** via `scripts/ansible-play.sh`, which:
|
||||
- Verifies `op account get` succeeds (fails closed).
|
||||
- Renders the cluster's `secrets.yml.tpl` (op:// references → real
|
||||
values) via `op inject -f` into a `0600` tmpfile cleaned by `trap`.
|
||||
- Execs `ansible-playbook --extra-vars @<tmpfile>`.
|
||||
4. **Service-account split**: superuser SA writes items (TF); read-only SA
|
||||
reads items (Ansible runtime + CI). Both SAs live as 1P items in
|
||||
`yucca_tf_dev`; `tf/.env` holds their `op://` references — resolved at
|
||||
invocation time via `op run --env-file=tf/.env -- terragrunt ...`.
|
||||
5. **Terragrunt multi-env** layout (`tf/deployment/<env>/<stack>/` with root
|
||||
`terragrunt.hcl`). State backend is S3 against the shared `yucca-tf-state`
|
||||
bucket at OVH Paris, keyed `ceph/${env}/${stack}/terraform.tfstate` —
|
||||
project-scoped so ceph state doesn't collide with o11y or future stacks.
|
||||
State locking (`use_lockfile`) deferred — OVH has no DynamoDB equivalent,
|
||||
and OpenTofu's lockfile-object mode fails on fresh backends with a 404
|
||||
before it can create one. Single-operator workflow today; revisit when
|
||||
concurrent applies become likely.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Positive:** No ansible-vault. No `vault.yml`. No `vault-password.sh`.
|
||||
No `secrets-sync.sh`. No `secrets-init.sh`. No dummy-password fallback.
|
||||
No reconciliation step. Rotation = edit 1P item (or `terragrunt taint`),
|
||||
next play run picks it up.
|
||||
- **Positive:** Cluster identity is declarative. Adding a cluster or host is
|
||||
an HCL edit + `tofu apply`. TF renders everything Ansible needs to
|
||||
consume.
|
||||
- **Positive:** CI lint and syntax-check don't require 1P access — ansible
|
||||
parsing doesn't hit secrets until play-execution.
|
||||
- **Positive:** Aligns with Immich devtools conventions — same `op run
|
||||
--env-file=` pattern, same TF+1P module shape, same vault naming
|
||||
structure (yucca_tf_*).
|
||||
- **Negative:** `op` CLI must be available on every control node; no
|
||||
offline/airgap fallback. Operationally acceptable.
|
||||
- **Negative:** TF state now gates Ansible runs — rendered files must exist
|
||||
before `ansible-playbook` has inputs. Mitigated by running `tofu apply`
|
||||
as the one-time bootstrap step (gitignored outputs re-render on demand).
|
||||
- **Negative:** Two service-account tokens to manage (read-only + superuser).
|
||||
Shared with o11y and other Futo consumers via `yucca_tf_dev` — rotation
|
||||
coordination required (see `docs/runbooks/rotate-sa-token.md`).
|
||||
|
||||
## Future direction
|
||||
|
||||
The TF-renders-Ansible-consumes pattern this ADR establishes for inventory
|
||||
and secrets extends naturally to cephadm service specs. [ADR-011](./011-cephadm-osd-service-specs.md)
|
||||
takes the first step (OSDs) — `osd-spec.yml.j2` is currently rendered by
|
||||
Ansible from per-host data. A logical next step ("Option C" in the
|
||||
architecture session log) is to move spec rendering into TF itself, so
|
||||
TF emits `inventory.ini` + `secrets.yml.tpl` + the full set of cephadm
|
||||
service specs (host registration, MON/MGR placement, OSD, RGW, monitoring)
|
||||
as a coherent set of artifacts. The Ansible role becomes a thin applier:
|
||||
`ceph orch apply -i <each-spec>`. Tracked as a follow-up PR after the
|
||||
import lands.
|
||||
|
||||
## References
|
||||
|
||||
- `tf/shared/modules/ceph-cluster/` — the module that renders inventories
|
||||
and (when re-enabled) provisions 1P items.
|
||||
- `tf/deployment/dev/ceph/` — the dev-env declaration.
|
||||
- `scripts/ansible-play.sh` — the Ansible-consumer wrapper.
|
||||
- `docs/secrets.md` — end-user documentation of the model.
|
||||
- [ADR-011](./011-cephadm-osd-service-specs.md) — first cephadm-spec
|
||||
refactor (OSD path).
|
||||
- [Immich devtools](https://github.com/immich-app/devtools) `tf/shared/modules/secrets/` — upstream pattern we copied.
|
||||
@@ -0,0 +1,126 @@
|
||||
# ADR-010: SSH Keys in 1Password (Native Category, Forward-Only Rotation)
|
||||
|
||||
## Status
|
||||
|
||||
Accepted (2026-04-22). Complements [ADR-009](./009-tf-first-op-inject-over-vault-password-sh.md)
|
||||
(TF-first secrets) by applying the same "1P is the storage layer" model to
|
||||
SSH keys.
|
||||
|
||||
## Context
|
||||
|
||||
Before this ADR, the `ansible-iac` SSH keypair for each cluster lived only on
|
||||
operator workstations — `~/.ssh/id_ed25519_sietch-ceph` and
|
||||
`~/.ssh/id_ed25519_ceph-painbox-lab-hel-htz`. Implications:
|
||||
|
||||
1. **No durable storage** — a laptop loss meant the private key was
|
||||
gone; re-access required provisioning a new key and manually
|
||||
distributing the pubkey to every host.
|
||||
2. **No clean onboarding path** — adding a new operator required
|
||||
out-of-band key sharing (insecure) or cutting them a separate key
|
||||
pair (possible but undocumented).
|
||||
3. **Remote-hands operators** had no scripted path to get cluster SSH
|
||||
access beyond an out-of-band share of the private key.
|
||||
4. **Inconsistent with the rest of the secrets model** — passwords live
|
||||
in `yucca_tf_dev`, but SSH keys lived only on disk. Operators had to
|
||||
remember two different security boundaries.
|
||||
|
||||
1Password has native `SSH Key` category support: items store the private
|
||||
key, automatically derive the public key, and integrate with 1P SSH Agent
|
||||
on workstations (already configured globally on the primary operator's
|
||||
machine — `Host * IdentityAgent ~/.1password/agent.sock` in
|
||||
`~/.ssh/config`).
|
||||
|
||||
## Decision
|
||||
|
||||
SSH keys join the hybrid secrets model:
|
||||
|
||||
1. **New keypairs are generated natively in 1P** via
|
||||
`op item create --category "SSH Key" --ssh-generate-key=ed25519`. Private
|
||||
key never touches operator disk during generation.
|
||||
2. **Items live in `yucca_tf_dev`** with title
|
||||
`<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY` (SHOUTY_SNAKE_CASE, same scheme
|
||||
as password items).
|
||||
3. **Public keys are consumed by Ansible at play time** via
|
||||
`op read "op://yucca_tf_dev/<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY/public_key"`
|
||||
in a new `rotate-ssh-key.yml` playbook, which ensures the current
|
||||
pubkey is present in `ansible-iac@<host>:~/.ssh/authorized_keys`
|
||||
(additive — doesn't remove others).
|
||||
4. **Private keys reach operator workstations** via
|
||||
`scripts/install-ssh-keys.sh` (idempotent `op read` → `~/.ssh/`,
|
||||
refuses to overwrite on fingerprint mismatch).
|
||||
5. **Rotation is forward-only**: new keys are generated, distributed,
|
||||
verified; old keys are cleaned up manually (SSH in, delete from
|
||||
authorized_keys, delete disk files) after confidence.
|
||||
|
||||
## Out of scope (for this ADR)
|
||||
|
||||
- **TF-managed SSH key resources.** Attempted earlier but op CLI's
|
||||
import path for SSH Keys produces malformed items (accepts
|
||||
private_key on create, strips the field on retrieval). Verified
|
||||
against op `2.34.0`. Generation via `--ssh-generate-key` works
|
||||
correctly. Future ADR may move generation into TF via
|
||||
`onepassword_item` with `tls_private_key` — but that path puts the
|
||||
private key into TF state, which is a step backwards for this
|
||||
particular category.
|
||||
- **1Password SSH Agent as the primary transport.** Operator workstations
|
||||
may already have it globally configured (recommended), but ansible
|
||||
inventory still references explicit `ansible_ssh_private_key_file`
|
||||
paths. A future change could
|
||||
drop the path entirely and rely on `IdentityAgent` routing — defer
|
||||
until we've run under the new keys long enough to trust the agent
|
||||
path.
|
||||
- **CI integration.** No CI pipeline runs Ansible yet. When that
|
||||
lands, CI would authenticate with the read-only SA, `install-ssh-keys.sh`
|
||||
pulls keys into a CI-scoped `~/.ssh/`. Separate PR.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Positive:** keys are durable (laptop loss is a non-event), rotatable
|
||||
(regenerate in 1P via `op item edit --ssh-generate-key`), auditable
|
||||
(1P logs reads), and scoped (read-only SA can pull them, superuser SA
|
||||
can rotate them).
|
||||
- **Positive:** remote-hands onboarding is a documented `scripts/install-ssh-keys.sh`
|
||||
invocation — no insecure key-over-chat.
|
||||
- **Positive:** consistent with ADR-009. Operators have one mental
|
||||
model: `op://<vault>/<ITEM>/<field>`, whether it's a password or a
|
||||
key.
|
||||
- **Negative:** op CLI can't import existing SSH keys reliably into the
|
||||
`SSH Key` category (bug in 2.34.0). Rotation is the only way in —
|
||||
existing keys are retired, not migrated. Operationally fine; makes
|
||||
"put my personal bastion key in 1P" harder.
|
||||
- **Negative:** ansible `rotate-ssh-key.yml` depends on `op` running on
|
||||
the controller (already required by `scripts/ansible-play.sh`). Not
|
||||
a new dependency, but does mean the key-rotation playbook can't run
|
||||
offline.
|
||||
- **Neutral:** private keys exist on operator disks during the
|
||||
transition between `install-ssh-keys.sh` and adoption of 1P SSH
|
||||
Agent as the sole transport. Files are `0600` and same security
|
||||
properties as the pre-ADR state.
|
||||
|
||||
## Migration that landed with this ADR
|
||||
|
||||
- New items `SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY` and
|
||||
`PAINBOX_CEPH_ANSIBLE_IAC_SSH_KEY` generated in `yucca_tf_dev`.
|
||||
- `clusters.auto.tfvars` `ansible_ssh_key` and
|
||||
`inventories/<cluster>/group_vars/all/vars.yml` `provision_iac_ssh_key_path`
|
||||
updated to `~/.ssh/id_ed25519_sietch` and `~/.ssh/id_ed25519_painbox`.
|
||||
- Rendered `inventory.ini` (all variants) and `secrets.yml.tpl` reflect
|
||||
the new key paths after `tofu apply`.
|
||||
- `scripts/install-ssh-keys.sh` added for operator-side key install.
|
||||
- `rotate-ssh-key.yml` playbook added for pubkey distribution.
|
||||
- End-to-end rotation verified against sietch: `install-ssh-keys.sh`
|
||||
pulled the new keypair to the operator workstation,
|
||||
`rotate-ssh-key.yml` distributed the public key to all three nodes'
|
||||
`ansible-iac:~/.ssh/authorized_keys` (exclusive: false — additive),
|
||||
and a subsequent `ansible -m ping` against the cluster succeeded
|
||||
using the new key. Old per-operator keypairs
|
||||
(`id_ed25519_sietch-ceph`, `id_ed25519_ceph-painbox-lab-hel-htz`) are
|
||||
retired — delete from disk at operator convenience. Painbox's old
|
||||
key is irrelevant post-reprovision.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR-009](./009-tf-first-op-inject-over-vault-password-sh.md) — companion, TF-first secrets
|
||||
- `scripts/install-ssh-keys.sh` — the operator-side tool
|
||||
- `rotate-ssh-key.yml` — the host-side distribution playbook
|
||||
- [1Password SSH Agent docs](https://developer.1password.com/docs/ssh/agent/)
|
||||
@@ -0,0 +1,146 @@
|
||||
# ADR-011: Cephadm OSD Service Specs over Imperative Per-Disk Loops
|
||||
|
||||
## Status
|
||||
|
||||
Accepted (2026-04-26). Refines [ADR-002](./002-explicit-osd-mapping.md)
|
||||
(explicit OSD-to-disk mapping) — preserves its anti-auto-discovery spirit
|
||||
while changing *where* the explicit mapping is expressed (cephadm spec
|
||||
applied via `ceph orch apply`, not an Ansible per-disk loop).
|
||||
|
||||
## Context
|
||||
|
||||
`roles/ceph_deploy/tasks/osds.yml` originally created OSDs by iterating
|
||||
each declared HDD/SSD in host_vars and invoking `cephadm ceph-volume lvm
|
||||
create --dmcrypt --data <disk> --block.db <lv>` per disk. The disk path
|
||||
was composed inside the role:
|
||||
|
||||
```jinja2
|
||||
DISK="/dev/disk/by-path/{{ sas_path_prefix }}-{{ item.path_phy }}-lun-0"
|
||||
```
|
||||
|
||||
This worked for sietch (Dell R730xd, SAS expander, dual-SSD-VG topology)
|
||||
because every assumption baked into the composition was sietch-shape:
|
||||
|
||||
- `sas_path_prefix` exists per host
|
||||
- `path_phy` is a slot identifier (`phy0`, `phy1`, ...) appended to the prefix
|
||||
- `-lun-0` suffix is the SAS LUN convention
|
||||
- SSD OSDs are partitions on the boot SSD (`-part6`)
|
||||
|
||||
When the painbox (Hetzner SX295) cluster came online for the first integration
|
||||
test of this codebase, every one of those assumptions broke:
|
||||
|
||||
1. Painbox is SATA + NVMe, no SAS expander → `sas_path_prefix` undefined,
|
||||
role failed at template-resolution time.
|
||||
2. SATA by-path strings are `pci-XXXX:XX:XX.X-ata-N` directly (no `-lun-N`
|
||||
suffix) — operator authors the full path identifier as `path_phy`.
|
||||
3. Painbox SSD OSD is an LV on the NVMe RAID-1 VG (`vg0/ssd-osd`), not a
|
||||
partition — different shape entirely from sietch's SSD OSDs.
|
||||
|
||||
The fix could have been per-shape branching inside `osds.yml`'s shell loop
|
||||
(if/else for path composition; if/else for partition vs LV), but every
|
||||
phase of the role beyond OSDs (RGW, monitoring, etc.) was already migrating
|
||||
toward cephadm's declarative service-spec pattern (`rgw-spec.yaml.j2` was
|
||||
already in place). The OSD path was the obvious next migration.
|
||||
|
||||
## Decision
|
||||
|
||||
OSD creation is now declarative via a cephadm OSD service spec rendered
|
||||
from per-host data, applied via `ceph orch apply osd -i /etc/ceph/osd-spec.yml`.
|
||||
|
||||
1. **`templates/osd-spec.yml.j2`** renders one document per host (HDD spec
|
||||
+ optional SSD spec). Per-host because sietch nodes have unique SAS
|
||||
prefixes per chassis — a shared spec with `data_devices.paths` requires
|
||||
identical paths across all `placement.hosts`.
|
||||
2. **Hardware-shape branching** lives in the template's Jinja conditional,
|
||||
not in role logic:
|
||||
- `sas_path_prefix is defined` → sietch-shape path composition
|
||||
- `sas_path_prefix is undefined` → painbox-shape, use `path_phy`
|
||||
directly (full PCI-ATA identifier from host_vars)
|
||||
- SSD OSD: `lv is defined` → painbox-style LV path; else sietch-style
|
||||
partition-on-SAS-disk
|
||||
3. **`tasks/osds.yml`** is now thin: render template → `ceph orch apply` →
|
||||
wait for provisioning → wait for OSDs up → defensive `ceph osd unset
|
||||
noin` → reweight-zero safety net. ~150 LOC down from ~230.
|
||||
4. **Encryption** is set in the spec (`encrypted: true`), not via a
|
||||
per-disk `--dmcrypt` flag. Cephadm provisions LUKS internally and
|
||||
stores the dm-crypt key in the MON config-key store as before.
|
||||
5. **Idempotency** is provided by cephadm — re-applying the same spec is
|
||||
a no-op. New disks (e.g. populating an empty bay) are picked up
|
||||
automatically. Existing OSDs are not destroyed by a spec change;
|
||||
removal requires `ceph orch osd rm`.
|
||||
|
||||
## Out of scope (for this ADR)
|
||||
|
||||
- **Extending the spec pattern to MON/MGR/RGW/monitoring placement**
|
||||
— `rgw-spec.yaml.j2` already does this for RGW; the rest of the role
|
||||
still uses imperative `ceph orch apply mon --placement=...` (and
|
||||
similar for MGR/monitoring). Refactoring those is a separate body of
|
||||
work — see "Option C" in the architecture session log; planned as a
|
||||
dedicated follow-up PR.
|
||||
- **TF-rendered service specs.** TF already has all the cluster identity
|
||||
it needs to render every cephadm spec deterministically. Moving spec
|
||||
rendering from Ansible (Jinja in role templates) to TF (Jinja in
|
||||
module templates) would shift the boundary further toward
|
||||
"TF declares, Ansible applies." Same Option-C follow-up.
|
||||
- **Auto-discovery via `data_devices.rotational: 1` filter.** Cephadm
|
||||
supports filter-based device selection. We chose explicit `paths`
|
||||
to preserve ADR-002's "no auto-discovery surprises" stance — empty
|
||||
bays on sietch and the SSD OSD on painbox can't be expressed with
|
||||
filters alone.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Positive:** hardware-shape independence. Painbox (SATA + NVMe-RAID
|
||||
+ LV-backed SSD OSD) and sietch (SAS expander + dual-SSD-VG) deploy
|
||||
through the same role with the same task file. Future clusters with
|
||||
yet other shapes (Hetzner AX-line, Equinix bare-metal, etc.) need
|
||||
only their own host_vars; the role doesn't change.
|
||||
- **Positive:** `osds.yml` shrinks ~35%; the deleted code was the most
|
||||
fragile part (per-disk shell loops with embedded Jinja path
|
||||
composition).
|
||||
- **Positive:** ADR-002's explicit-mapping spirit is preserved. Paths
|
||||
are still listed explicitly in the spec — cephadm doesn't auto-discover.
|
||||
Empty bays stay empty; partitions stay reserved for OS/block.db.
|
||||
- **Positive:** Aligns with cephadm's intended deployment model. All
|
||||
modern cephadm operators use service specs; the imperative-loop
|
||||
pattern is legacy.
|
||||
- **Neutral:** The spec is rendered to `/etc/ceph/osd-spec.yml` on the
|
||||
bootstrap node and then applied. The file is overwritten on each
|
||||
re-render — operators inspecting the deployed state should query
|
||||
cephadm (`ceph orch ls --service-type osd --export`) rather than
|
||||
reading the on-disk spec, which may have been updated since the last
|
||||
apply.
|
||||
- **Negative:** Cephadm's spec apply is async — the role polls until
|
||||
the expected OSD count is reached, with a 15-minute timeout. On a
|
||||
large cluster (hundreds of OSDs), this could exceed the timeout. Not
|
||||
a concern at current scale (15 OSDs painbox, 30 OSDs sietch).
|
||||
- **Negative:** Debugging "why isn't this disk becoming an OSD?" is
|
||||
harder than with the per-disk loop, where the failing disk had its
|
||||
own log line. With cephadm, you check `ceph orch ls`,
|
||||
`ceph orch ps`, and `ceph cephadm osd activate <host> --dry-run` on
|
||||
the bootstrap node.
|
||||
|
||||
## Migration that landed with this ADR
|
||||
|
||||
- `templates/osd-spec.yml.j2` written (replaces stub that existed but
|
||||
was never wired in).
|
||||
- `tasks/osds.yml` rewritten to render + apply + poll + safety net.
|
||||
- `tasks/lvm-setup.yml` gated on `sas_path_prefix is defined` — sietch
|
||||
still needs the role's defensive VG/LV recovery path; painbox skips
|
||||
because installimage's post-install owns LVM lifecycle.
|
||||
- Validated end-to-end against a freshly-reprovisioned painbox: 15/15
|
||||
OSDs up + in, all encrypted (15 dm-crypt keys in MON store), correct
|
||||
per-host service IDs (`osd.painbox-ceph-evelyn-hdd`,
|
||||
`osd.painbox-ceph-evelyn-ssd`).
|
||||
- Sietch validation deferred — sietch is currently deployed and
|
||||
healthy; the spec apply against existing OSDs is idempotent (no-op
|
||||
when paths match), but a real test on sietch is a separate operational
|
||||
step planned for the next sietch maintenance window.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR-002](./002-explicit-osd-mapping.md) — explicit OSD-to-disk mapping (refined, not superseded)
|
||||
- [ADR-009](./009-tf-first-op-inject-over-vault-password-sh.md) — TF-first secrets (companion architectural shift)
|
||||
- `roles/ceph_deploy/templates/osd-spec.yml.j2` — the template
|
||||
- `roles/ceph_deploy/tasks/osds.yml` — the thin applier
|
||||
- [Cephadm OSD Service docs](https://docs.ceph.com/en/latest/cephadm/services/osd/)
|
||||
@@ -0,0 +1,825 @@
|
||||
# Architecture
|
||||
|
||||
How the Ceph automation in `yucca/ansible/ceph/` is shaped, what each tool
|
||||
owns, and how the four tools (Terraform, 1Password, Ansible, mise) hand off
|
||||
work to each other.
|
||||
|
||||
This is a structural reference. For step-by-step usage see
|
||||
[CONTRIBUTING.md](../CONTRIBUTING.md); for narrower topics see the
|
||||
specialized docs under `docs/`.
|
||||
|
||||
---
|
||||
|
||||
## 1. System context
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
OP([Operator workstation<br/>mise · op CLI · ansible · tofu])
|
||||
YUCCA[/Yucca monorepo<br/>tf/ + ansible/ceph/ + kubernetes//]
|
||||
ONEP[("1Password org<br/>yucca_tf · yucca_tf_dev · ...")]
|
||||
S3[("OVH S3<br/>yucca-tf-state bucket")]
|
||||
SIETCH["Sietch · Austin DC<br/>3× Dell R730xd"]
|
||||
PAINBOX["Painbox · Hetzner Helsinki<br/>1× SX295"]
|
||||
|
||||
OP -->|edits| YUCCA
|
||||
OP -->|reads/writes secrets| ONEP
|
||||
OP -->|TF state I/O| S3
|
||||
OP -->|SSH ansible-iac| SIETCH
|
||||
OP -->|SSH ansible-iac| PAINBOX
|
||||
```
|
||||
|
||||
The yucca monorepo is the single source of truth for cluster identity and
|
||||
configuration. Operators run mise tasks on their workstation; secrets stay
|
||||
in 1Password (never on disk); TF state lives in OVH S3; Ansible drives
|
||||
configuration over SSH against bare-metal Ceph nodes.
|
||||
|
||||
External dependencies are minimal and explicit:
|
||||
|
||||
- **1Password org** — organization-scoped vaults shared with other Futo infra
|
||||
(Immich, o11y). Authoritative store for live secret values.
|
||||
- **OVH S3** — `yucca-tf-state` bucket at `s3.eu-west-par.io.cloud.ovh.net`.
|
||||
Project-keyed (`ceph/<env>/<stack>/terraform.tfstate`) so multiple stacks
|
||||
share the bucket without collision.
|
||||
- **Hardware** — Austin colo for sietch (Dell R730xd × 3, single 10G bond);
|
||||
Hetzner Helsinki for painbox (SX295 × 1). Detail in [hardware.md](hardware.md).
|
||||
|
||||
---
|
||||
|
||||
## 2. Environments
|
||||
|
||||
Environments are a first-class concern: every tool in the mesh derives
|
||||
its environment from the same source — directory layout — so dev / staging
|
||||
/ prod isolation is structural, not flag-driven.
|
||||
|
||||
| Layer | dev (today) | staging (planned) | prod (planned) |
|
||||
|--------------|----------------------------------------------------------------|---------------------------------------------------------|------------------------------------------------------|
|
||||
| TF stack dir | `tf/deployment/dev/ceph/` | `tf/deployment/staging/ceph/` | `tf/deployment/prod/ceph/` |
|
||||
| TF state key | `ceph/dev/ceph/terraform.tfstate` | `ceph/staging/ceph/terraform.tfstate` | `ceph/prod/ceph/terraform.tfstate` |
|
||||
| 1P vaults | `yucca_tf_dev` (live) · `yucca_tf_dev_manual` (human-fillable) | `yucca_tf_staging` · `yucca_tf_staging_manual` | `yucca_tf` (live) · `yucca_tf_prod_manual` |
|
||||
| Ansible inv | `inventories/<cluster>-ceph.dev.<dc>.<provider>/` | `inventories/<cluster>-ceph.staging.<dc>.<provider>/` | `inventories/<cluster>-ceph.prod.<dc>.<provider>/` |
|
||||
| mise default | `CEPH_ENV=...sietch-ceph.dev.austin.int/inventory.ini` | overridden via env at invocation | overridden via env at invocation |
|
||||
|
||||
Today the only deployed environment is dev (sietch + painbox). Adding
|
||||
staging/prod is purely additive: create the matching `tf/deployment/<env>/ceph/`
|
||||
directory, populate `clusters.auto.tfvars`, and the same module + Ansible
|
||||
roles + mise tasks work unchanged. The state backend key path, 1P vault
|
||||
selection, and inventory directory naming all derive from the env segment.
|
||||
|
||||
`TF_STACK_DIR` is the operator-side override for `mise run tf:*` tasks; it
|
||||
defaults to `tf/deployment/dev/ceph` and points at any sibling stack directory.
|
||||
`CEPH_ENV` is the matching override for Ansible — points at the rendered
|
||||
`inventory.ini` for the cluster you intend to operate on.
|
||||
|
||||
---
|
||||
|
||||
## 3. The four-tool mesh
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph ws["Operator workstation"]
|
||||
direction LR
|
||||
MISE([mise<br/>orchestration])
|
||||
TF[Terraform / Tofu<br/>via Terragrunt]
|
||||
OP[op CLI]
|
||||
ANS[Ansible]
|
||||
WRAP[scripts/<br/>ansible-play.sh<br/>install-ssh-keys.sh]
|
||||
end
|
||||
|
||||
ONEP[("1Password<br/>yucca_tf_*")]
|
||||
S3[("OVH S3<br/>tfstate")]
|
||||
REPO[/"Yucca repo<br/>inventories/<cluster>/<br/>(host_vars committed,<br/>TF outputs gitignored)"/]
|
||||
NODES[Ceph nodes]
|
||||
|
||||
MISE -->|tf:*| TF
|
||||
MISE -->|deploy / status / drift| WRAP
|
||||
MISE -->|capture| WRAP
|
||||
|
||||
TF -->|reads SA token via op run --env-file| OP
|
||||
TF -->|reads/writes state| S3
|
||||
TF -->|renders| REPO
|
||||
|
||||
WRAP -->|reads inventory + secrets.yml.tpl| REPO
|
||||
WRAP -->|op inject / op read| OP
|
||||
WRAP -->|exec| ANS
|
||||
ANS -->|SSH ansible-iac| NODES
|
||||
|
||||
OP <-->|item CRUD| ONEP
|
||||
```
|
||||
|
||||
### Who owns what
|
||||
|
||||
| Tool | Owns | Reads from |
|
||||
|---------------|--------------------------------------------------------------------------------|---------------------------------------------|
|
||||
| **Terraform** | Cluster identity, host names, rendered Ansible artifacts, TF state | `clusters.auto.tfvars` · 1P (via op CLI) |
|
||||
| **1Password** | Live secret values, SSH keypairs, service-account tokens | nothing — authoritative store |
|
||||
| **op CLI** | Auth and resolution: env-injection, file-template injection, single-value read | 1P (session or SA token) |
|
||||
| **Ansible** | Convergence: applying configuration to nodes | Rendered inventory + op-injected tmpfile |
|
||||
| **mise** | Task discovery, toolchain pinning, env defaults | `.mise.toml`, `tf/.env` |
|
||||
|
||||
### Handoff points (the edges of the mesh)
|
||||
|
||||
1. **TF → repo** — `terragrunt apply` renders `inventory.ini`,
|
||||
`inventory-destroy.ini`, optional `inventory-provision-<profile>.ini`,
|
||||
and `secrets.yml.tpl` into `inventories/<cluster>/`. These files are
|
||||
gitignored — the source of truth is `clusters.auto.tfvars` + the
|
||||
ceph-cluster module.
|
||||
2. **TF ↔ op CLI** — TF runs are wrapped with `op run --env-file=tf/.env`,
|
||||
which resolves `op://...` references in `tf/.env` and injects them as
|
||||
`OP_SERVICE_ACCOUNT_TOKEN`, `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`
|
||||
for the child process. The `tf/.env` file is committed (it contains only
|
||||
pointers, never literal secrets).
|
||||
3. **Ansible ↔ op CLI** — `scripts/ansible-play.sh` reads the cluster's
|
||||
`secrets.yml.tpl`, runs `op inject -f` to resolve `op://` references into
|
||||
a `mktemp`'d tmpfile (chmod 600, trap-cleaned), then execs
|
||||
`ansible-playbook --extra-vars @<tmpfile>`. The tmpfile lives only for
|
||||
the duration of the play.
|
||||
4. **mise → wrappers** — `mise run deploy` invokes
|
||||
`scripts/ansible-play.sh deploy.yml ...`; `mise run tf:*` invokes
|
||||
`op run --env-file=tf/.env -- terragrunt --working-dir <stack> <cmd>`.
|
||||
mise tasks never call `ansible-playbook` directly.
|
||||
|
||||
---
|
||||
|
||||
## 4. Terraform (authority)
|
||||
|
||||
### Layout
|
||||
|
||||
```
|
||||
tf/
|
||||
├── .env op:// references (committed; no literal secrets)
|
||||
├── shared/modules/ceph-cluster/ per-cluster orchestration module
|
||||
│ ├── main.tf · variables.tf · outputs.tf · rendering.tf
|
||||
│ ├── wordlist.txt 923 words for auto-picked hostnames
|
||||
│ └── templates/ inventory + secrets.yml.tpl templates
|
||||
└── deployment/
|
||||
├── terragrunt.hcl root: state backend, env/stack derived from path
|
||||
└── dev/ceph/
|
||||
├── terragrunt.hcl includes root, sets ansible_project_root
|
||||
├── main.tf · variables.tf · versions.tf
|
||||
└── clusters.auto.tfvars declarative cluster list
|
||||
```
|
||||
|
||||
### Cluster identity is declared, not derived
|
||||
|
||||
The `clusters` map in `clusters.auto.tfvars` is the source of truth.
|
||||
Each top-level key becomes a cluster:
|
||||
|
||||
```hcl
|
||||
sietch = {
|
||||
domain = "dev.austin.int.futo.cloud"
|
||||
environment = "dev"
|
||||
datacenter = "austin"
|
||||
provider_code = "int"
|
||||
role_in_hostname = "ceph"
|
||||
ansible_ssh_user = "ansible-iac"
|
||||
ansible_ssh_key = "~/.ssh/id_ed25519_sietch"
|
||||
vault = "yucca_tf_dev"
|
||||
provision_profile = "debian-live"
|
||||
hosts = [
|
||||
{ name = "laurel", bond_ip = "10.10.10.90", bootstrap = true },
|
||||
{ name = "lawson", bond_ip = "10.10.10.91" },
|
||||
{ name = "samara", bond_ip = "10.10.10.92" },
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
The module computes everything else: hostname (`<cluster>-<role>-<name>`),
|
||||
FQDN (`<hostname>.<domain>`), 1P item names
|
||||
(`<CLUSTER>_CEPH_<ROLE>_PASSWORD`), inventory directory path
|
||||
(`inventories/<cluster>-<role>.<env>.<dc>.<provider>/`).
|
||||
|
||||
### Auto-naming via wordlist
|
||||
|
||||
For hosts where `name = null`, the module picks a stable name from a
|
||||
923-word pool using `random_shuffle` seeded by `(cluster_name, name_seed)`.
|
||||
Operator-declared names are excluded from the pool to prevent collisions
|
||||
within a cluster. Adding hosts at the tail is safe — existing positions
|
||||
keep their names across applies.
|
||||
|
||||
Painbox today demonstrates this: its single host has no `name`, so TF
|
||||
auto-picked `evelyn` → hostname `painbox-ceph-evelyn`.
|
||||
|
||||
### Rendered artifacts (gitignored)
|
||||
|
||||
Per `tf/shared/modules/ceph-cluster/rendering.tf`, the module writes four
|
||||
files into `inventories/<cluster>/`:
|
||||
|
||||
| File | Purpose |
|
||||
|--------------------------------------------|------------------------------------------------------------------------------------------|
|
||||
| `inventory.ini` | Normal-ops inventory: `ansible-iac` user + cluster SSH key |
|
||||
| `inventory-destroy.ini` | Destroy-mode inventory (same credentials; separate file as a speed bump) |
|
||||
| `inventory-provision-<profile>.ini` | Provisioning inventory (only when `provision_profile != null`; uses live-image creds) |
|
||||
| `secrets.yml.tpl` | `vault_*: op://<vault>/<CLUSTER>_CEPH_*/password` pointers, consumed by `op inject -f` |
|
||||
|
||||
All four are in `ansible/ceph/.gitignore`. Re-render with `mise run tf:apply`.
|
||||
|
||||
### State backend
|
||||
|
||||
S3 backend in `tf/deployment/terragrunt.hcl`:
|
||||
|
||||
- Bucket: `yucca-tf-state` (shared with o11y and other Futo stacks)
|
||||
- Region: `eu-west-par` (OVH Paris)
|
||||
- Endpoint: `https://s3.eu-west-par.io.cloud.ovh.net/`
|
||||
- Key: `ceph/${env}/${stack}/terraform.tfstate` — derived from
|
||||
the child stack's path under `deployment/`
|
||||
- Skip AWS-specific validation; use path-style URLs (OVH compatibility)
|
||||
|
||||
State locking is **not enabled today**. OVH has no DynamoDB equivalent.
|
||||
OpenTofu's `use_lockfile = true` would work but expects the lockfile object
|
||||
to already exist — fresh-backend init fails with 404 before it can create
|
||||
one. Single-operator workflow today; revisit when concurrent applies become
|
||||
likely. See `deployment/terragrunt.hcl` for the inline rationale.
|
||||
|
||||
### What TF does not yet manage
|
||||
|
||||
`onepassword_item` resources are **dormant** (`tf/.../secrets.tf.disabled`).
|
||||
1P items are created today via the `op` CLI (operator runs `op item create`
|
||||
once per cluster). Re-enabling them is tracked in
|
||||
[ADR-009](adr/009-tf-first-op-inject-over-vault-password-sh.md) — the gate
|
||||
is the dedicated `sietch-ceph` service account that lets us split write
|
||||
authority from the org-wide superuser SA.
|
||||
|
||||
---
|
||||
|
||||
## 5. 1Password (live values)
|
||||
|
||||
### Vault hierarchy
|
||||
|
||||
| Vault | Purpose | Who reads it | Who writes it |
|
||||
|---------------------------|--------------------------------------------------------|------------------------------------|-------------------------------------|
|
||||
| `yucca_tf` | Cross-env shared (TF state S3 creds) | TF (via op run --env-file) | Operator (manual) |
|
||||
| `yucca_tf_dev` | dev environment live values | Ansible runtime (op inject) | Superuser SA (TF) + operator (op CLI) |
|
||||
| `yucca_tf_dev_manual` | dev human-fillable placeholders (API tokens, OAuth) | Ansible runtime | Operator (manual) |
|
||||
| `yucca_tf_staging(_manual)` · `yucca_tf_prod_manual` | (planned) staging/prod analogues | (future) | (future) |
|
||||
|
||||
The `_manual` vaults exist for items that can't be auto-generated (third-party
|
||||
API tokens, OAuth client secrets) — they're populated by humans, not by TF.
|
||||
|
||||
### Service accounts
|
||||
|
||||
Two service accounts in `yucca_tf_dev`, both shared org-wide:
|
||||
|
||||
| SA | Scope | Consumed by |
|
||||
|-----------------------------------------------|---------------------------------------------|-------------------------------------|
|
||||
| `yucca_futo_1pass_superuser_service_account` | Read + write all `yucca_tf_*` vaults | TF (via `tf/.env`) + interactive op CLI |
|
||||
| `yucca_futo_1pass_service_account` | Read-only on `yucca_tf` and `yucca_tf_dev` | Ansible runtime / future CI |
|
||||
|
||||
Rotation procedure: [docs/runbooks/rotate-sa-token.md](runbooks/rotate-sa-token.md).
|
||||
|
||||
### Item categories and naming
|
||||
|
||||
Per cluster, the following items live in the cluster's `vault` (currently
|
||||
`yucca_tf_dev` for both sietch and painbox):
|
||||
|
||||
| Category | Item title pattern | Field consumed |
|
||||
|------------|---------------------------------------------------------|----------------------|
|
||||
| Password | `<CLUSTER>_CEPH_OPS_PASSWORD` | `password` |
|
||||
| Password | `<CLUSTER>_CEPH_DASHBOARD_PASSWORD` | `password` |
|
||||
| Password | `<CLUSTER>_CEPH_GRAFANA_PASSWORD` | `password` |
|
||||
| Password | `<CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY` | `password` |
|
||||
| Password | `<CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY` | `password` |
|
||||
| SSH Key | `<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY` | `private_key` / `public_key` |
|
||||
| Document | `<CLUSTER>_CEPH_RGW_TLS_CERT` · `..._RGW_TLS_KEY` | file content |
|
||||
| Document | `<CLUSTER>_CEPH_CLIENT_ADMIN_KEYRING` | file content |
|
||||
|
||||
The `<CLUSTER>_CEPH_*` prefix is hardcoded in
|
||||
`tf/shared/modules/ceph-cluster/main.tf` (`secret_prefix = "${upper(var.cluster_name)}_CEPH"`)
|
||||
so every Ceph-project item across all clusters is grep-discoverable as
|
||||
`*_CEPH_*` regardless of the role segment in node hostnames.
|
||||
|
||||
Full item-by-item catalog: [docs/secrets.md](secrets.md).
|
||||
|
||||
### Three op-CLI patterns
|
||||
|
||||
The op CLI is invoked in three distinct ways across the codebase. Each
|
||||
serves a different shape of secret consumption:
|
||||
|
||||
1. **`op run --env-file=tf/.env -- <cmd>`** — env-var injection.
|
||||
Resolves `op://` references in a dotenv file and injects the resolved
|
||||
values as env vars into the child process. Used for TF (SA token) and
|
||||
the S3 backend (AWS creds). Wrapped by all `mise run tf:*` tasks.
|
||||
2. **`op inject -f -i <tpl> -o <out>`** — file-template resolution.
|
||||
Reads a file containing inline `op://` references, resolves each, writes
|
||||
to the output path. Used by `scripts/ansible-play.sh` to render
|
||||
`secrets.yml.tpl` → tmpfile, and by the Hetzner installimage flow to
|
||||
render `post-install.sh.tpl` → `post-install.sh`.
|
||||
3. **`op read "op://<vault>/<item>/<field>"`** — single-value read.
|
||||
Used by `scripts/install-ssh-keys.sh`, `rotate-ssh-key.yml`,
|
||||
`post-deploy-capture.yml`. Returns one value to stdout for one specific
|
||||
field; fails closed if missing.
|
||||
|
||||
No custom password-script (no `vault-password.sh`); no
|
||||
`ansible-vault`-encrypted file in git. Lint and syntax-check tasks don't
|
||||
invoke op at all — they don't need secrets, so "1P unavailable" never
|
||||
silently degrades them. See [ADR-009](adr/009-tf-first-op-inject-over-vault-password-sh.md).
|
||||
|
||||
---
|
||||
|
||||
## 6. Ansible (consumer)
|
||||
|
||||
### Role dependency graph
|
||||
|
||||
Provisioning is a separate concern (`provision.yml`, sietch only). The main
|
||||
pipeline (`site.yml`) runs everything else in this order:
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
PROV["provision_host<br/><i>separate playbook, live image</i>"]
|
||||
BASE["baseline<br/><i>users, packages, /etc/hosts</i>"]
|
||||
OST["os_tuning<br/><i>sysctl, TCP buffers</i>"]
|
||||
HWT["hardware_tuning<br/><i>I/O scheduler, readahead</i>"]
|
||||
DEPLOY["ceph_deploy<br/><i>bootstrap, join, OSDs, RGW,<br/>crush rules, monitoring</i>"]
|
||||
CTUNE["ceph_tuning<br/><i>recovery throttling, scrub<br/>window, telemetry, audit</i>"]
|
||||
SEC["security<br/><i>nftables, SSH hardening</i>"]
|
||||
|
||||
PROV -.->|reboot into installed OS| BASE
|
||||
BASE --> OST
|
||||
BASE --> HWT
|
||||
OST --> DEPLOY
|
||||
HWT --> DEPLOY
|
||||
DEPLOY --> CTUNE
|
||||
CTUNE --> SEC
|
||||
|
||||
classDef separate stroke-dasharray: 4 4
|
||||
class PROV separate
|
||||
```
|
||||
|
||||
`site.yml` starts at `baseline` — `provision_host` runs only on first
|
||||
install via `provision.yml`.
|
||||
|
||||
### Why this order matters
|
||||
|
||||
1. **baseline before tuning** — cephadm needs podman, dbus, chrony.
|
||||
The baseline role installs these and enables the services. Running
|
||||
tuning on a node without podman would leave cephadm unable to bootstrap.
|
||||
2. **tuning before deploy** — OSD daemons inherit kernel parameters
|
||||
active at startup. Applying sysctl (`vm.min_free_kbytes`, `fs.aio-max-nr`)
|
||||
and I/O scheduler (`mq-deadline` for HDD, `none` for SSD) before bootstrap
|
||||
means daemons launch with correct limits from the first second.
|
||||
3. **ceph_tuning after deploy** — these settings use `ceph config set`
|
||||
which requires a running cluster. Recovery throttling, scrub windows, and
|
||||
PG autoscaler targets cannot be applied until MONs are up.
|
||||
4. **security last** — nftables drops all traffic not explicitly allowed.
|
||||
Running it before ceph_deploy would block cephadm's inter-node SSH,
|
||||
container image pulls, and MON/OSD port negotiation. Once the cluster is
|
||||
healthy, the firewall locks it down.
|
||||
|
||||
### ceph_deploy internal pipeline
|
||||
|
||||
`roles/ceph_deploy/tasks/main.yml` orchestrates ten phases:
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
P1["Phase 1 · prerequisites.yml<br/><i>Ceph repo, cephadm, ceph-common</i>"]
|
||||
P2["Phase 2 · bootstrap.yml<br/><i>cephadm bootstrap on first node</i>"]
|
||||
P3["Phase 3 · join.yml<br/><i>ceph orch host add for remaining nodes</i>"]
|
||||
P4["Phase 4 · placement.yml<br/><i>MON/MGR placement calculation</i>"]
|
||||
P45["Phase 4.5 · lvm-setup.yml<br/><i>ensure block.db VGs/LVs exist (sietch-shape only;<br/>painbox skips — LVM owned by installimage post-install)</i>"]
|
||||
P5["Phase 5 · osds.yml<br/><i>render osd-spec.yml.j2 → ceph orch apply osd<br/>(cephadm provisions LUKS + LVM internally)</i>"]
|
||||
P55["Phase 5.5 · crush-rules.yml<br/><i>replicated_hdd / replicated_ssd rules</i>"]
|
||||
P575["Phase 5.75 · rgw.yml<br/><i>EC pools, realm/zone, TLS, S3 user</i>"]
|
||||
P58["Phase 5.8 · monitoring.yml<br/><i>dashboard URL integration, Grafana creds</i>"]
|
||||
P6["Phase 6 · verify.yml<br/><i>cluster health report</i>"]
|
||||
|
||||
P1 --> P2 --> P3 --> P4 --> P45 --> P5 --> P55 --> P575 --> P58 --> P6
|
||||
```
|
||||
|
||||
Tag-driven re-runs are first-class:
|
||||
`scripts/ansible-play.sh deploy-ceph.yml --tags rgw,monitoring` re-runs
|
||||
just those phases.
|
||||
|
||||
### Inventory layout
|
||||
|
||||
```
|
||||
inventories/
|
||||
sietch-ceph.dev.austin.int/ Austin dev cluster
|
||||
inventory.ini TF-generated, gitignored
|
||||
inventory-destroy.ini TF-generated, gitignored
|
||||
inventory-provision.ini TF-generated, gitignored
|
||||
secrets.yml.tpl TF-generated, gitignored
|
||||
group_vars/all/vars.yml cluster-wide variables (committed)
|
||||
host_vars/ per-node hardware topology (committed)
|
||||
sietch-ceph-laurel.yml bond_ip, SAS path prefix, OSD maps
|
||||
sietch-ceph-lawson.yml
|
||||
sietch-ceph-samara.yml
|
||||
installimage/ Hetzner installimage assets (sietch n/a)
|
||||
|
||||
painbox-ceph.dev.hel.htz/ Hetzner dev cluster
|
||||
inventory.ini TF-generated, gitignored
|
||||
inventory-destroy.ini TF-generated, gitignored
|
||||
secrets.yml.tpl TF-generated, gitignored
|
||||
group_vars/all/vars.yml committed
|
||||
host_vars/painbox-ceph-evelyn.yml committed
|
||||
installimage/ post-install.sh.tpl (op-injected)
|
||||
```
|
||||
|
||||
`host_vars/*.yml` is committed because per-node hardware facts (bond_ip, SAS
|
||||
expander paths, SSD PHY positions, HDD-to-block.db mappings) are stable
|
||||
inventory truth — not operator preference. The `.local.yml` suffix is
|
||||
gitignored as an escape hatch for operator-local overrides.
|
||||
|
||||
### Variable precedence
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
D["<b>role defaults</b><br/>roles/*/defaults/main.yml<br/><i>lowest priority</i>"]
|
||||
G["<b>group_vars</b><br/>inventories/<cluster>/group_vars/all/vars.yml"]
|
||||
H["<b>host_vars</b><br/>inventories/<cluster>/host_vars/<host>.yml"]
|
||||
T["<b>extra-vars @tmpfile</b><br/>scripts/ansible-play.sh<br/><i>op-injected secrets</i>"]
|
||||
E["<b>extra-vars -e X=Y</b><br/>-e confirm_wipe=true<br/><i>highest priority</i>"]
|
||||
|
||||
D --> G --> H --> T --> E
|
||||
```
|
||||
|
||||
- **Role defaults** define every tunable with a safe value
|
||||
(`ceph_firewall_ssh_any_source: true`, `ceph_cpu_governor_enabled: false`).
|
||||
- **group_vars/all/vars.yml** sets cluster-wide values: network topology,
|
||||
Ceph release, RGW config, monitoring ports, plus the `vault_*` →
|
||||
consumable-name aliases (`ops_password: "{{ vault_ops_password }}"`).
|
||||
- **host_vars** provides per-node physical topology.
|
||||
- **extra-vars from @tmpfile** carries op-injected `vault_ops_password`,
|
||||
`vault_ceph_dashboard_password`, `vault_grafana_admin_password`,
|
||||
`vault_s3_restic_access_key`, `vault_s3_restic_secret_key`.
|
||||
- **extra-vars via `-e`** carries safety gates: `confirm_wipe=true`,
|
||||
`provision_skip_reboot=true`, `yes_destroy_ceph=true`.
|
||||
|
||||
### ansible.cfg stays generic
|
||||
|
||||
`ansible.cfg` contains zero site-specific values. No default inventory, no
|
||||
ProxyJump, no hardcoded key paths. Site-specifics live exclusively in
|
||||
`clusters.auto.tfvars` (which TF renders into the inventory) or in the
|
||||
inventory's `group_vars`. The same `ansible.cfg` and the same roles work
|
||||
unchanged across Austin, Hetzner, or any future cluster — only the cluster
|
||||
entry in `clusters.auto.tfvars` differs.
|
||||
|
||||
---
|
||||
|
||||
## 7. mise (orchestration surface)
|
||||
|
||||
### Why mise
|
||||
|
||||
- **Toolchain pinning** — `.mise.toml` declares the exact versions of
|
||||
`python`, `tofu`, `terragrunt`, `op`. New operators get a working
|
||||
environment with `mise trust && mise run setup`.
|
||||
- **Task discovery** — `mise tasks` lists every operation; tasks are
|
||||
shell-script-shaped, kept in `.mise.toml`, and committed.
|
||||
- **Devtools parity** — matches the conventions in `immich-app/devtools`
|
||||
(where the `op run --env-file=tf/.env --` pattern originated).
|
||||
|
||||
### Task taxonomy
|
||||
|
||||
| Group | Tasks |
|
||||
|----------------|--------------------------------------------------------------------------|
|
||||
| Bootstrap | `setup` |
|
||||
| Verify | `lint` · `check` · `test` · `preflight` |
|
||||
| TF (wrapped) | `tf:init` · `tf:plan` · `tf:apply` · `tf:destroy` |
|
||||
| Read-only ops | `status` · `drift` |
|
||||
| State change | `deploy` · `destroy` · `capture` · `backup` |
|
||||
| Benchmarks | `bench` · `bench-rados` |
|
||||
|
||||
### How mise wraps the underlying CLIs
|
||||
|
||||
- `mise run tf:*` → `op run --env-file=tf/.env -- terragrunt --working-dir <stack> <cmd>`
|
||||
- `mise run deploy` → `scripts/ansible-play.sh deploy-ceph.yml ...` (per phase)
|
||||
- `mise run status` → `scripts/ansible-play.sh status.yml`
|
||||
- `mise run capture` → `scripts/ansible-play.sh post-deploy-capture.yml`
|
||||
|
||||
mise never invokes `ansible-playbook` or `terragrunt` directly. The wrappers
|
||||
own secrets injection and pre-flight checks; mise owns task discovery and
|
||||
env defaults.
|
||||
|
||||
### Env defaults
|
||||
|
||||
```toml
|
||||
[env]
|
||||
CEPH_ENV = "inventories/sietch-ceph.dev.austin.int/inventory.ini"
|
||||
```
|
||||
|
||||
`TF_STACK_DIR` defaults to `tf/deployment/dev/ceph` inside each `tf:*` task.
|
||||
Both are overridable per-invocation:
|
||||
|
||||
```bash
|
||||
TF_STACK_DIR=tf/deployment/staging/ceph mise run tf:plan
|
||||
CEPH_ENV=inventories/painbox-ceph.dev.hel.htz/inventory.ini mise run status
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 8. Wrapper scripts (the glue layer)
|
||||
|
||||
Three scripts under `ansible/ceph/scripts/` sit between mise and the
|
||||
underlying CLIs. They exist to keep secrets out of `argv`, fail closed
|
||||
when 1P is unreachable, and give better error messages than the raw tools.
|
||||
|
||||
| Script | Purpose |
|
||||
|-------------------------|----------------------------------------------------------------------------------|
|
||||
| `ansible-play.sh` | Render secrets via `op inject -f` to a `mktemp`'d file (chmod 600, trap-cleaned), then exec `ansible-playbook --extra-vars @<tmpfile>` |
|
||||
| `install-ssh-keys.sh` | Idempotent `op read` → `~/.ssh/id_ed25519_<cluster>` installer; refuses overwrite on fingerprint mismatch |
|
||||
| `preflight.sh` | Verifies TF artifacts present, 1P session live, SSH reachable, Python on targets — surfaced via `mise run preflight` |
|
||||
|
||||
Per-script reference (synopsis, args, env, exit codes, examples):
|
||||
[docs/scripts.md](scripts.md).
|
||||
|
||||
---
|
||||
|
||||
## 9. Data flow: concrete operations
|
||||
|
||||
### 9.1 `mise run tf:apply` — render artifacts
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
actor OP as Operator
|
||||
participant MISE as mise
|
||||
participant OPCLI as op CLI
|
||||
participant ONEP as 1Password
|
||||
participant TG as terragrunt / tofu
|
||||
participant S3 as OVH S3
|
||||
participant REPO as Repo (inventories/)
|
||||
|
||||
OP->>MISE: mise run tf:apply
|
||||
MISE->>OPCLI: op run --env-file=tf/.env -- ...
|
||||
OPCLI->>ONEP: resolve op:// references
|
||||
ONEP-->>OPCLI: SA token + AWS keys
|
||||
OPCLI->>TG: exec child process<br/>with env vars injected
|
||||
TG->>S3: read tfstate<br/>(ceph/<env>/<stack>/terraform.tfstate)
|
||||
S3-->>TG: current state
|
||||
TG->>TG: plan + apply
|
||||
TG->>S3: write updated tfstate
|
||||
TG->>REPO: render inventory.ini,<br/>secrets.yml.tpl, ...
|
||||
```
|
||||
|
||||
### 9.2 `mise run deploy` — full Ceph deploy
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
actor OP as Operator
|
||||
participant MISE as mise
|
||||
participant WRAP as ansible-play.sh
|
||||
participant OPCLI as op CLI
|
||||
participant ONEP as 1Password
|
||||
participant TMP as /tmp/<id>-secrets.yml
|
||||
participant ANS as ansible-playbook
|
||||
participant NODES as Ceph nodes
|
||||
|
||||
OP->>MISE: mise run deploy
|
||||
MISE->>WRAP: ansible-play.sh deploy-ceph.yml
|
||||
WRAP->>OPCLI: op account get
|
||||
OPCLI-->>WRAP: session OK
|
||||
WRAP->>TMP: mktemp + chmod 600 + trap rm
|
||||
WRAP->>OPCLI: op inject -f -i secrets.yml.tpl -o TMP
|
||||
OPCLI->>ONEP: resolve op://yucca_tf_dev/SIETCH_CEPH_*/password
|
||||
ONEP-->>OPCLI: secret values
|
||||
OPCLI->>TMP: write resolved YAML
|
||||
WRAP->>ANS: exec --extra-vars @TMP
|
||||
loop phases 1..6
|
||||
ANS->>NODES: SSH ansible-iac@<bond_ip><br/>via id_ed25519_<cluster>
|
||||
end
|
||||
Note over WRAP,TMP: tmpfile rm'd on exit (trap)
|
||||
```
|
||||
|
||||
### 9.3 `mise run capture` — DR snapshot
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
actor OP as Operator
|
||||
participant MISE as mise
|
||||
participant ANS as ansible-playbook<br/>(post-deploy-capture.yml)
|
||||
participant BOOT as Bootstrap node
|
||||
participant LOCAL as localhost (delegated)
|
||||
participant OPCLI as op CLI
|
||||
participant ONEP as 1Password
|
||||
|
||||
OP->>MISE: mise run capture
|
||||
MISE->>ANS: ansible-play.sh post-deploy-capture.yml
|
||||
ANS->>BOOT: SSH read /etc/ceph/rgw-ssl.crt
|
||||
ANS->>BOOT: SSH read /etc/ceph/rgw-ssl.key
|
||||
ANS->>BOOT: SSH read /etc/ceph/ceph.client.admin.keyring
|
||||
BOOT-->>ANS: file contents
|
||||
loop for each artifact
|
||||
ANS->>LOCAL: delegate_to: localhost
|
||||
LOCAL->>OPCLI: op item edit/create<br/><CLUSTER>_CEPH_<ITEM>
|
||||
OPCLI->>ONEP: upsert Document item<br/>in yucca_tf_dev
|
||||
end
|
||||
Note over ONEP: Now holds RGW_TLS_CERT,<br/>RGW_TLS_KEY, CLIENT_ADMIN_KEYRING
|
||||
```
|
||||
|
||||
### 9.4 `scripts/install-ssh-keys.sh` — fresh workstation
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
actor OP as Operator (new ws)
|
||||
participant SCRIPT as install-ssh-keys.sh
|
||||
participant OPCLI as op CLI
|
||||
participant ONEP as 1Password
|
||||
participant SSH as ~/.ssh/
|
||||
|
||||
OP->>SCRIPT: install-ssh-keys.sh sietch
|
||||
SCRIPT->>OPCLI: op read .../public_key
|
||||
OPCLI->>ONEP: SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY
|
||||
ONEP-->>OPCLI: public_key
|
||||
OPCLI-->>SCRIPT: pubkey content
|
||||
SCRIPT->>SSH: compare with id_ed25519_sietch (if exists)
|
||||
alt fingerprint match
|
||||
SCRIPT-->>OP: skip (already present)
|
||||
else fingerprint mismatch
|
||||
SCRIPT-->>OP: refuse (operator must mv aside)
|
||||
else file missing
|
||||
SCRIPT->>OPCLI: op read .../private_key
|
||||
OPCLI-->>SCRIPT: private_key
|
||||
SCRIPT->>SSH: write id_ed25519_sietch (0600)<br/>+ .pub (0644)<br/>(umask 077)
|
||||
end
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 10. OSD lifecycle
|
||||
|
||||
### Phase flow
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
SIETCH["Sietch prep:<br/>provision_host/disks.yml partitions SSDs<br/>then ceph_deploy/lvm-setup.yml<br/><i>creates VG + db-slot LVs on each SSD's partition 5</i>"]
|
||||
PAINBOX["Painbox prep:<br/>installimage post-install.sh<br/><i>NVMe RAID-1 → vg0 → db-slot0..13 + ssd-osd LVs</i>"]
|
||||
SPEC["ceph_deploy/osds.yml renders<br/>templates/osd-spec.yml.j2 → /etc/ceph/osd-spec.yml<br/><i>one document per host; paths from host_vars</i>"]
|
||||
APPLY["ceph orch apply osd -i /etc/ceph/osd-spec.yml<br/><i>cephadm: discover disks, LUKS-format, LVM, deploy daemons</i>"]
|
||||
POLL["Wait for cephadm to provision<br/><i>poll num_osds until expected count reached</i>"]
|
||||
UP["Wait for OSDs up<br/><i>poll num_up_osds == num_osds</i>"]
|
||||
UNSET["Defensive: ceph osd unset noin<br/><i>idempotent — clears stale flag from prior runs</i>"]
|
||||
REWEIGHT["Safety net: fix any reweight=0 OSDs"]
|
||||
|
||||
SIETCH --> SPEC
|
||||
PAINBOX --> SPEC
|
||||
SPEC --> APPLY --> POLL --> UP --> UNSET --> REWEIGHT
|
||||
```
|
||||
|
||||
### Service-spec model, not per-disk loops
|
||||
|
||||
Earlier versions of this role iterated `cephadm ceph-volume lvm create`
|
||||
per disk and composed `/dev/disk/by-path/...` paths from host_vars
|
||||
(`sas_path_prefix` + `path_phy`). That assumed sietch's SAS expander
|
||||
topology and broke on painbox's PCI-ATA disks plus LV-backed SSD OSD.
|
||||
|
||||
The current flow renders a cephadm OSD service spec from per-host data
|
||||
and applies it via `ceph orch apply osd -i`. Cephadm handles device
|
||||
path resolution, LUKS encryption (`encrypted: true`), LVM provisioning,
|
||||
and daemon deployment. The role is hardware-shape-agnostic — the only
|
||||
shape-aware logic is the template's Jinja conditional. See
|
||||
[ADR-011](adr/011-cephadm-osd-service-specs.md) for the decision record.
|
||||
|
||||
### Hardware-shape independence in the template
|
||||
|
||||
`templates/osd-spec.yml.j2` renders one document per host (sietch nodes
|
||||
have unique SAS prefixes per chassis, so a shared spec doesn't work)
|
||||
with two shape branches:
|
||||
|
||||
- **Sietch** (`sas_path_prefix` defined): data path =
|
||||
`/dev/disk/by-path/{{ sas_path_prefix }}-{{ path_phy }}-lun-0`; SSD
|
||||
OSD = partition on the SAS-attached SSD via `path_phy + partition`.
|
||||
- **Painbox** (`sas_path_prefix` undefined): data path =
|
||||
`/dev/disk/by-path/{{ path_phy }}` (operator authors the full PCI-ATA
|
||||
identifier in host_vars); SSD OSD = LV via the `lv` field
|
||||
(`/dev/{{ lv }}`).
|
||||
|
||||
`db_devices.paths` is always `/dev/{{ db }}` — both shapes use LVs for
|
||||
block.db, no composition needed.
|
||||
|
||||
### Idempotency
|
||||
|
||||
`ceph orch apply osd` is idempotent — re-applying the same spec is a
|
||||
no-op when deployed OSDs match. New disks (populating an empty bay
|
||||
later, future expansion) are picked up automatically on the next apply.
|
||||
Existing OSDs are not destroyed by a spec apply — removal requires
|
||||
explicit `ceph orch osd rm`.
|
||||
|
||||
### Defensive noin handling
|
||||
|
||||
The spec-based flow doesn't need the `noin` flag (cephadm rolls out
|
||||
OSDs gracefully one at a time). The role's tail still includes a
|
||||
`ceph osd unset noin` task as a defensive cleanup — stale `noin` flags
|
||||
from a prior failed run of the older imperative flow can leave the
|
||||
cluster degraded; the unconditional unset clears that safely (no-op
|
||||
when already unset).
|
||||
|
||||
### Reweight-zero safety net
|
||||
|
||||
`osds.yml` ends with a task that fixes any OSD stuck at `reweight=0` by
|
||||
running `ceph osd reweight <id> 1.0`. Rare with the spec-based flow but
|
||||
kept as a backstop against an OSD coming up while `noin` was set
|
||||
externally.
|
||||
|
||||
---
|
||||
|
||||
## 11. Monitoring
|
||||
|
||||
### What cephadm auto-deploys
|
||||
|
||||
cephadm's bootstrap automatically deploys:
|
||||
- **node-exporter** on every node
|
||||
- **ceph-exporter** on every node
|
||||
- **prometheus** (single instance, cephadm-managed)
|
||||
- **alertmanager** (single instance)
|
||||
- **grafana** (single instance, with pre-built Ceph dashboards)
|
||||
- **89 Prometheus alert rules** across 16 groups
|
||||
|
||||
### What we configure
|
||||
|
||||
`roles/ceph_deploy/tasks/monitoring.yml` handles only integration:
|
||||
|
||||
1. Enable the `prometheus` MGR module (if not already enabled)
|
||||
2. Wait for all five monitoring service types to report `running > 0`
|
||||
3. Set dashboard integration URLs for Prometheus, Alertmanager, Grafana
|
||||
(using the bootstrap node's `bond_ip`)
|
||||
4. Set Grafana admin credentials from the op-injected
|
||||
`vault_grafana_admin_password`
|
||||
5. Disable Grafana SSL cert verification in dashboard (self-signed cert)
|
||||
|
||||
`roles/ceph_tuning/tasks/main.yml` verifies the alert rule count and warns
|
||||
if fewer than 10 rule groups are loaded (expects 16+).
|
||||
|
||||
---
|
||||
|
||||
## 12. Provision host internals
|
||||
|
||||
### Ten-phase flow
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
P1["detect.yml<br/><i>live image + UEFI assertions, SSD discovery</i>"]
|
||||
P2["prerequisites_live.yml<br/><i>apt setup on live image, install debootstrap/mdadm/lvm2</i>"]
|
||||
P3["disks.yml<br/><i>partition, mdraid, LVM, mount at /mnt</i>"]
|
||||
P4["install.yml<br/><i>debootstrap Bookworm into /mnt</i>"]
|
||||
P5["configure.yml<br/><i>hostname, hosts, network, fstab, mdadm templates</i>"]
|
||||
P6["chroot_packages.yml<br/><i>bind mounts, apt install, machine-id, SSH keys</i>"]
|
||||
P7["admin_user.yml<br/><i>ansible-iac (key-only) inside chroot;<br/>ops user is created post-boot by baseline (ADR-003)</i>"]
|
||||
P8["bootloader.yml<br/><i>initramfs, grub-install, efibootmgr</i>"]
|
||||
P9["finalize.yml<br/><i>marker, ESP mirror, unmount, reboot</i>"]
|
||||
P10["unmount.yml<br/><i>reverse-order cleanup (shared with rescue)</i>"]
|
||||
|
||||
P1 --> P2 --> P3 --> P4 --> P5 --> P6 --> P7 --> P8 --> P9 --> P10
|
||||
```
|
||||
|
||||
### Marker-driven resume gate
|
||||
|
||||
After `disks.yml` runs, `main.yml` checks for
|
||||
`/mnt/etc/ceph-provisioned.json`. If present and the hostname matches, all
|
||||
chroot phases (4-8) plus the marker/ESP block are skipped. The role goes
|
||||
straight to unmount + reboot.
|
||||
|
||||
This prevents:
|
||||
- Re-binding bind mounts that are already in place
|
||||
- Re-rotating SSH host keys (would break known_hosts)
|
||||
- Re-hashing the ops password with a fresh salt
|
||||
- Re-running grub-install for no reason
|
||||
- Overwriting the marker with a stale `provisioned_at` timestamp
|
||||
|
||||
The marker filename (`ceph-provisioned.json`) is project-scoped, not
|
||||
cluster-scoped — every Ceph cluster (sietch, painbox, future) writes the
|
||||
same filename. The marker's *contents* identify which cluster + host the
|
||||
machine belongs to.
|
||||
|
||||
### Block/rescue cleanup
|
||||
|
||||
The entire provisioning sequence (phases 2-9) runs inside a `block/rescue`.
|
||||
If any phase fails, the rescue block includes `unmount.yml` which tears
|
||||
down chroot bind mounts and the /mnt hierarchy in reverse order, then
|
||||
re-raises the failure. This ensures the next run starts from a clean mount
|
||||
state.
|
||||
|
||||
---
|
||||
|
||||
## 13. Future
|
||||
|
||||
- **CI / GitHub Actions** — the SA split (superuser write vs read-only
|
||||
consume) already enables it. Read-only SA in CI runs `mise run lint`,
|
||||
`mise run check`, `mise run test`, `mise run preflight` against every PR.
|
||||
Superuser SA only runs `mise run tf:plan` (never `apply`) to detect drift.
|
||||
- **Talos K8s as a sibling stack** — `tf/deployment/<env>/talos/` would
|
||||
share the same terragrunt root config and S3 backend, with its own state
|
||||
key (`ceph/<env>/talos/terraform.tfstate`). Deployment plan lives
|
||||
outside this repo until Phase A begins; a per-stack README lands
|
||||
alongside the code when it's implemented.
|
||||
- **TF-managed `onepassword_item` resources** — re-enable the dormant
|
||||
resources in `secrets.tf.disabled` once the dedicated ceph service
|
||||
account lands. Tracked in [ADR-009](adr/009-tf-first-op-inject-over-vault-password-sh.md).
|
||||
- **OSD LUKS keys in 1P** — store dm-crypt keys for DR. Deferred until
|
||||
the hybrid is stable.
|
||||
|
||||
---
|
||||
|
||||
## See also
|
||||
|
||||
| Topic | Doc |
|
||||
|------------------------------------|--------------------------------------------------------------------------------|
|
||||
| TF/Terragrunt detail | [`tf/README.md`](../../../tf/README.md) |
|
||||
| Wrapper script reference | [docs/scripts.md](scripts.md) |
|
||||
| Secrets catalog + rotation | [docs/secrets.md](secrets.md) |
|
||||
| Trust boundaries + encryption | [docs/security-model.md](security-model.md) |
|
||||
| Hardware specs + network topology | [docs/hardware.md](hardware.md) |
|
||||
| Coding idioms and anti-patterns | [docs/patterns.md](patterns.md) |
|
||||
| Adding a new cluster (walkthrough) | [docs/adding-a-cluster.md](adding-a-cluster.md) |
|
||||
| TF-first + op inject decision | [ADR-009](adr/009-tf-first-op-inject-over-vault-password-sh.md) |
|
||||
| SSH keys in 1P decision | [ADR-010](adr/010-ssh-keys-in-1password.md) |
|
||||
| Cephadm OSD service spec decision | [ADR-011](adr/011-cephadm-osd-service-specs.md) |
|
||||
| Why explicit OSD-to-disk mapping | [ADR-002](adr/002-explicit-osd-mapping.md) |
|
||||
| Why baseline is split from provision | [ADR-003](adr/003-baseline-split-from-provision.md) |
|
||||
| Why debootstrap (not preseed) for sietch | [ADR-008](adr/008-debootstrap-over-preseed.md) |
|
||||
@@ -0,0 +1,79 @@
|
||||
# Capacity Planning
|
||||
|
||||
Audience: Managers, procurement, budget planning. For hardware specs and
|
||||
disk layouts see [hardware.md](hardware.md); this doc is the sizing math.
|
||||
|
||||
## Current deployments
|
||||
|
||||
| Cluster | Nodes | HDD × size | Raw HDD | EC-usable (8+3) | 70%-full target |
|
||||
|---|---|---|---|---|---|
|
||||
| sietch (Austin) | 3 × Dell R730xd | 30 × 6 TB | ~164 TiB | ~119 TiB | ~83 TiB |
|
||||
| painbox (Hetzner Helsinki) | 1 × SX295 | 14 × 22 TB | ~280 TiB | ~204 TiB | ~143 TiB |
|
||||
| **Combined** | | | **~444 TiB** | **~323 TiB** | **~226 TiB** |
|
||||
|
||||
Each cluster also contributes ~1 TiB of SSD OSD on the boot SSDs (minor;
|
||||
ignored in the math above).
|
||||
|
||||
Sietch has ~6 empty bays across its three nodes (block.db LVs pre-created),
|
||||
worth +36 TB raw (~33 TiB) by populating them. No LVM or network changes
|
||||
needed — cheapest expansion path.
|
||||
|
||||
## Sizing formulas
|
||||
|
||||
```
|
||||
EC-usable = raw × (k / (k + m)) = raw × 8/11 = raw × 0.727
|
||||
Operational target = EC-usable × 0.70 (keep cluster below 70% full)
|
||||
Raw needed = target_data / 0.727 / 0.70
|
||||
```
|
||||
|
||||
`backfillfull` triggers at 85% full (stops recovery); `full` triggers at
|
||||
95% (stops writes). 70% is the conservative operational ceiling —
|
||||
substantial headroom for failures, rebalancing, and growth.
|
||||
|
||||
Ceph also consumes small amounts for index pools, RGW metadata, and the
|
||||
non-EC multipart-upload pool. All negligible relative to data (< 1% each
|
||||
at steady state).
|
||||
|
||||
## Worked example
|
||||
|
||||
Target: 50 TiB of application data.
|
||||
|
||||
```
|
||||
raw_needed = 50 / 0.727 / 0.70 = 98 TiB raw HDD
|
||||
HDDs = 98 TiB / 5.45 TiB = 18 drives (at 6 TB each)
|
||||
Nodes = 18 / 12 bays = 2 nodes minimum (populated)
|
||||
```
|
||||
|
||||
For 22 TB Hetzner drives the drive count is much lower (~5 drives) but
|
||||
you still need enough failure domains for EC — see below.
|
||||
|
||||
## When to add drives vs. nodes
|
||||
|
||||
| | Add drives | Add a node |
|
||||
|---|---|---|
|
||||
| When | Empty bays exist with pre-created block.db LVs | All bays populated; need more IOPS, network, or failure domains |
|
||||
| Cost | ~$30–50 per 6 TB HDD (used) | ~$1,000–1,500 per fully-populated node (used) |
|
||||
| Adds | +3.96 TiB EC-usable per drive | +47 TiB EC-usable per fully-populated R730xd |
|
||||
| Impact | Backfill only | Backfill + CRUSH reshape + monitoring/SSH/cephadm host onboarding |
|
||||
|
||||
## Failure domain ceiling
|
||||
|
||||
EC 8+3 needs 11 failure domains. Austin currently uses
|
||||
`failure_domain=osd` (spreads across 30 OSDs across 3 nodes) — works today
|
||||
but a full-node loss degrades a large share of PGs. For host-level
|
||||
failure domain you need **11+ nodes minimum**. Production (Yucca) will
|
||||
want this; dev can tolerate the weaker guarantee.
|
||||
|
||||
## block.db sizing
|
||||
|
||||
Rule of thumb: block.db ≈ 4% of OSD data size.
|
||||
|
||||
- **Austin (6 TB HDDs):** 240 GiB LV per HDD matches the rule; 6 LVs
|
||||
consume 1,440 GiB of each SSD's partition 5 (see hardware.md).
|
||||
- **Hetzner (22 TB HDDs):** 128 GiB LV is undersized against the 4% rule
|
||||
(would want 256–512 GiB). Acceptable for dev/benchmark use; production
|
||||
deployments with 22 TB HDDs should target larger block.db.
|
||||
|
||||
If block.db fills up, BlueStore spills metadata to the HDD data partition
|
||||
— OSD keeps working but small-object operations slow down. Fix: grow the
|
||||
block.db LV or reduce metadata density.
|
||||
@@ -0,0 +1,162 @@
|
||||
# Hardware Reference
|
||||
|
||||
Audience: Ops, procurement, capacity planning. Per-node hardware facts that
|
||||
differ across clusters live in `ansible/ceph/inventories/<cluster>/host_vars/`
|
||||
(bond_ip, SAS path prefix, SSD PHY positions, HDD-to-block.db mappings) —
|
||||
those files are committed and are authoritative for physical topology.
|
||||
|
||||
For where this fits in the tool mesh, see [architecture.md](architecture.md).
|
||||
|
||||
## Network topology
|
||||
|
||||
Both clusters use a **single-network** design today — public and cluster
|
||||
traffic share one subnet. Production will separate them; see
|
||||
[security-model.md](security-model.md) for the threat-model implications.
|
||||
|
||||
| Cluster | Subnet | Connection |
|
||||
|----------|-------------------|------------------------------------------------------------------------|
|
||||
| sietch | `10.10.10.0/24` | 2× 10GbE bonded active-backup (eno1 + eno2) per node; private switch |
|
||||
| painbox | public /32 | Single 1GbE, direct SSH (no bond, no ProxyJump) |
|
||||
|
||||
Per-node connection IPs (`bond_ip`) are declared in
|
||||
`tf/deployment/dev/ceph/clusters.auto.tfvars` and rendered by TF into the
|
||||
cluster's `inventory.ini`. They're also mirrored into `host_vars/` for use
|
||||
by roles that need the IP as a variable (e.g., cephadm public-network
|
||||
resolution, dashboard URL construction).
|
||||
|
||||
## sietch -- Dell R730xd (Austin)
|
||||
|
||||
| Component | Spec |
|
||||
|-----------|------|
|
||||
| Chassis | Dell R730xd 12-bay LFF + 2 rear 2.5" bays |
|
||||
| CPU | 2x Intel Xeon E5-2697A v4 (64 vCPUs) |
|
||||
| RAM | 128 GB DDR4 |
|
||||
| Boot SSDs | 2x Micron 5100 3.8TB (rear bays 12/13, mdraid-1) |
|
||||
| HDD OSDs | 8-12x SAS 6TB (HGST HUS726060AL4210 / Seagate ST6000NKCLAR6000) |
|
||||
| SSD OSDs | Partition 6 on each boot SSD (colocated, no separate block.db) |
|
||||
| Block.db | 6x 240GB LVs per SSD (partition 5, LVM VG) |
|
||||
| HBA | Broadcom/LSI SAS3008 IT mode (mpt3sas, no RAID) |
|
||||
| Network | 2x 10GbE bonded active-backup (eno1 + eno2) |
|
||||
| Boot | UEFI, dual ESP (one per SSD, rsync-mirrored) |
|
||||
| OS | Debian 12 Bookworm (debootstrap provisioned) |
|
||||
| Provisioning | Live image → `provision.yml` → debootstrap |
|
||||
|
||||
### SSD partition layout (per SSD)
|
||||
|
||||
| Partition | Size | Use |
|
||||
|-----------|------|-----|
|
||||
| 1 | 512 MB | ESP (FAT32, UEFI boot) |
|
||||
| 2 | 1 GB | /boot (mdraid-1, ext4, metadata 1.0) |
|
||||
| 3 | 80 GB | / (mdraid-1, ext4) |
|
||||
| 4 | 8 GB | swap (mdraid-1) |
|
||||
| 5 | ~1.4 TB | Ceph block.db LVs (LVM VG) |
|
||||
| 6 | ~2 TB | SSD OSD data (ceph-volume) |
|
||||
|
||||
## painbox -- Hetzner SX295 (Helsinki)
|
||||
|
||||
| Component | Spec |
|
||||
|-----------|------|
|
||||
| Chassis | Hetzner SX295 storage server |
|
||||
| CPU | AMD EPYC 7502P (32C/64T) |
|
||||
| RAM | 128 GB DDR4 ECC |
|
||||
| Boot NVMe | 2x Samsung 7.68TB (installimage RAID-1, vg0) |
|
||||
| HDD OSDs | 14x Seagate Exos X22 22TB SATA |
|
||||
| SSD OSD | 1x ~4.4TB LV on vg0 (NVMe remainder) |
|
||||
| Block.db | 14x 128GB LVs on vg0 |
|
||||
| SATA | 3 onboard controllers (8+2+4 ports = 14 total) |
|
||||
| Network | Single 1GbE, direct SSH (no bond, no ProxyJump) |
|
||||
| Boot | BIOS (Hetzner standard) |
|
||||
| OS | Debian 12 Bookworm (Hetzner installimage) |
|
||||
| Provisioning | Rescue mode → `installimage/autosetup` + `post-install.sh` |
|
||||
|
||||
### NVMe layout (vg0 on md1)
|
||||
|
||||
| LV | Size | Use |
|
||||
|----|------|-----|
|
||||
| swap | 32 GB | Swap |
|
||||
| root | 100 GB | / |
|
||||
| var | 200 GB | /var |
|
||||
| varlog | 50 GB | /var/log |
|
||||
| db-slot0..13 | 14x 128 GB | Block.db per HDD OSD |
|
||||
| ssd-osd | ~4.4 TB | SSD OSD data |
|
||||
| (reserve) | ~512 GB | Future expansion |
|
||||
|
||||
> **Why Bookworm and not Trixie:** upstream Ceph Tentacle's Debian
|
||||
> repository at `download.ceph.com/debian-tentacle/dists/` publishes only
|
||||
> for `bookworm`, `jammy`, and `noble`. Trixie is not yet supported. The
|
||||
> autosetup `IMAGE` line MUST select a Bookworm tarball until upstream
|
||||
> ships Trixie packages.
|
||||
|
||||
## Comparison
|
||||
|
||||
| | sietch (per node) | painbox (single node) |
|
||||
|---|---|---|
|
||||
| HDD OSD count | 8-12 | 14 |
|
||||
| SSD OSD count | 2 | 1 |
|
||||
| block.db per HDD | 240 GB | 128 GB |
|
||||
| Total raw HDD | 48-72 TB | 308 TB |
|
||||
| Device path format | `/dev/disk/by-path/sas-exp*-phy*-lun-0` | `/dev/disk/by-path/pci-*-ata-*` |
|
||||
| EC profile | 8+3 (failure domain: OSD) | 8+3 (failure domain: OSD) |
|
||||
| Replicated pool size | 2 (min_size 1) | 2 (min_size 1) |
|
||||
|
||||
## host_vars schema by hardware shape
|
||||
|
||||
The two clusters have fundamentally different storage topologies, which
|
||||
shows up in their `host_vars/<host>.yml` schemas. When adding a new cluster,
|
||||
operators must pick the schema matching the hardware — not just copy from
|
||||
either existing cluster blindly.
|
||||
|
||||
### sietch-shape (SAS expander + dual-SSD-VG)
|
||||
|
||||
```yaml
|
||||
hostname_short: <cluster>-ceph-<name>
|
||||
bond_ip: 10.10.X.Y
|
||||
sas_path_prefix: "pci-XXXX:XX:XX.X-sas-exp0xXXXX..." # REQUIRED — disambiguates SAS topology
|
||||
ssd1_phy: 12 # PHY slot of boot SSD #1
|
||||
ssd2_phy: 13 # PHY slot of boot SSD #2
|
||||
ceph_db_vg1: ceph-db-rear12 # VG on SSD1 partition 5
|
||||
ceph_db_vg2: ceph-db-rear13 # VG on SSD2 partition 5
|
||||
ceph_db_lvs_per_ssd: 6 # 6 db-slot LVs per VG → 12 total
|
||||
ceph_hdd_osds:
|
||||
- { path_phy: phy0, db: ceph-db-rear12/db-slot0 }
|
||||
- ...
|
||||
ceph_ssd_osds: # SSD OSD = partition 6 of each boot SSD
|
||||
- { path_phy: phy12, partition: 6 }
|
||||
- { path_phy: phy13, partition: 6 }
|
||||
```
|
||||
|
||||
The role composes full disk paths as
|
||||
`/dev/disk/by-path/<sas_path_prefix>-<path_phy>-lun-0` (HDDs) or with
|
||||
`-part<N>` suffix (SSD partitions). `roles/ceph_deploy/tasks/lvm-setup.yml`
|
||||
runs this shape's VG/LV recovery path.
|
||||
|
||||
### painbox-shape (NVMe-RAID + single VG + LV-backed SSD OSD)
|
||||
|
||||
```yaml
|
||||
hostname_short: <cluster>-ceph-<name>
|
||||
bond_ip: <public IP>
|
||||
ceph_db_vg: vg0 # single VG on NVMe RAID-1 (no sas_path_prefix)
|
||||
ceph_db_lvs_per_node: 14 # all db-slot LVs on the one VG
|
||||
ceph_hdd_osds:
|
||||
- { path_phy: pci-XXXX:XX:XX.X-ata-1, db: vg0/db-slot0 } # full PCI-ATA path
|
||||
- ...
|
||||
ceph_ssd_osds:
|
||||
- { lv: vg0/ssd-osd } # SSD OSD = LV on same VG
|
||||
```
|
||||
|
||||
The role uses `path_phy` directly as the by-path identifier (no composition
|
||||
needed — operator authors the full string). For SSD OSDs, `lv` field is
|
||||
used (`/dev/<lv>`) instead of `path_phy + partition`. `lvm-setup.yml` is
|
||||
**skipped** on this shape (gated `when: sas_path_prefix is defined`) —
|
||||
LVM is owned by the Hetzner installimage post-install script.
|
||||
|
||||
### Decision rule for new clusters
|
||||
|
||||
- **Has a SAS expander** (PERC HBA, mpt3sas, etc.) and **dedicated boot SSDs
|
||||
partitioned for both block.db and OS** → sietch-shape.
|
||||
- **Single VG covering boot + block.db + SSD OSD** (typical for
|
||||
hosting-provider servers with NVMe RAID-1) → painbox-shape.
|
||||
- **Other shapes** (e.g., dedicated NVMe block.db drives) require either
|
||||
a new shape branch in `roles/ceph_deploy/tasks/osds.yml`'s template or
|
||||
a fresh decision — see [ADR-011](adr/011-cephadm-osd-service-specs.md)
|
||||
for how cephadm OSD service specs handle hardware-shape independence.
|
||||
@@ -0,0 +1,243 @@
|
||||
# Naming
|
||||
|
||||
Every name derivable in this project — hostname, inventory directory, 1P
|
||||
item title, SSH key filename — traces back to one entry per cluster in
|
||||
`tf/deployment/<env>/ceph/clusters.auto.tfvars`. TF's `ceph-cluster` module
|
||||
assembles the rest.
|
||||
|
||||
For how naming fits the broader system see
|
||||
[architecture.md §4 (Terraform)](architecture.md); for the per-item 1P
|
||||
catalog see [secrets.md](secrets.md); for the SSH-key lifecycle see
|
||||
[ADR-010](adr/010-ssh-keys-in-1password.md).
|
||||
|
||||
## The three name layers
|
||||
|
||||
Three parallel naming surfaces derive from the same tfvars entry, but
|
||||
**only the hostname carries the `role_in_hostname` segment**. Inventory
|
||||
directory and 1P item prefix are both hardcoded to `ceph` / `CEPH` in the
|
||||
ceph-cluster module — they're project-scoped, not role-scoped.
|
||||
|
||||
| Surface | Pattern | Role source |
|
||||
|------------------------|---------------------------------------------------|---------------------------------|
|
||||
| Hostname (short + FQDN)| `<cluster>-<role>-<name>[.<env>.<dc>.<provider>.futo.cloud]` | `role_in_hostname` tfvar |
|
||||
| Inventory directory | `inventories/<cluster>-ceph.<env>.<dc>.<provider>/` | always `ceph` (module-hardcoded) |
|
||||
| 1P item prefix | `<CLUSTER>_CEPH_*` | always `CEPH` (module-hardcoded) |
|
||||
|
||||
This split is deliberate. A future cluster where every node is a dedicated
|
||||
OSD might set `role_in_hostname = "osd"` (yielding hostnames like
|
||||
`mesa-osd-willow`) but its inventory dir and 1P items would still grep-match
|
||||
`*-ceph.*` and `*_CEPH_*` alongside every other Ceph-project cluster.
|
||||
|
||||
### Hostname segments
|
||||
|
||||
| Component | Example | Source (per-cluster tfvars field) |
|
||||
|-----------|---------------------------------|--------------------------------------|
|
||||
| cluster | `sietch`, `painbox` | top-level map key |
|
||||
| role | `ceph` (small clusters), `osd` | `role_in_hostname` (default `ceph`) |
|
||||
| name | `laurel`, `evelyn` | `hosts[].name`, or TF-picked |
|
||||
| env | `dev`, `staging`, `prod` | `environment` |
|
||||
| dc | `austin`, `hel`, `fsn` | `datacenter` |
|
||||
| provider | `int`, `htz` | `provider_code` |
|
||||
|
||||
Current `role_in_hostname` values: both sietch and painbox use `ceph`
|
||||
(mixed-role, all-nodes-are-everything). Dedicated-role hostnames (`osd`,
|
||||
`mon`) are supported but not used today.
|
||||
|
||||
## Cluster naming
|
||||
|
||||
**Who chooses:** the engineer adding the cluster picks the name at the
|
||||
moment they add the entry to `clusters.auto.tfvars`. No automation — it's
|
||||
a deliberate, one-time decision.
|
||||
|
||||
**When:** before any `mise run tf:apply`, any 1P item creation, any cluster
|
||||
bootstrap. Renaming after deployment is expensive (see "Cost of renaming"
|
||||
below).
|
||||
|
||||
**Convention (unenforced):** Dune-themed. Existing clusters are `sietch`
|
||||
(an underground Fremen community) and `painbox` (the Bene Gesserit
|
||||
gom-jabbar test apparatus). Dune candidates not yet used, and that fit
|
||||
the hard constraints below: `arrakis`, `caladan`, `giedi`, `ixian`,
|
||||
`kwisatz`, `muaddib`, `fremen`, `chani`, `leto`, `jessica`. Nothing in
|
||||
the code enforces Dune specifically — mixed themes or theme breaks are
|
||||
acceptable when they communicate intent better (e.g., a cluster named
|
||||
after its datacenter for a production tier).
|
||||
|
||||
**Hard constraints:**
|
||||
|
||||
- **lowercase alphanumeric** — becomes the HCL map key, the inventory
|
||||
directory segment (`<name>-ceph.<env>...`), the hostname prefix, and
|
||||
(uppercased) the 1P item prefix (`<NAME>_CEPH_*`).
|
||||
- **Short** — appears in every hostname and every 1P item title. Aim for
|
||||
6–10 characters; 15 is the realistic ceiling.
|
||||
- **Unique within the `yucca_tf_*` 1P item namespace** — other Futo
|
||||
consumers (o11y, future stacks) write items to the same vault set.
|
||||
Before committing, check:
|
||||
```bash
|
||||
op item list --vault yucca_tf_dev --format=json \
|
||||
| jq -r '.[] | .title' | grep -i "^<PROPOSED_NAME>_"
|
||||
```
|
||||
Must return empty.
|
||||
- **Not already a `clusters.auto.tfvars` key** — TF enforces this with
|
||||
a plan-time error.
|
||||
- **No dashes, dots, or underscores in the cluster name itself** — those
|
||||
are segment separators in inventory directories and hostnames. A cluster
|
||||
name `my-cluster` would produce `my-cluster-ceph-laurel` which parses
|
||||
ambiguously. Use `mycluster` instead.
|
||||
|
||||
**Soft guidance:**
|
||||
|
||||
- Memorable — operators will say it out loud in incidents.
|
||||
- Distinct from existing clusters' first 3 letters (grep-friendly in logs).
|
||||
- Doesn't encode environment or datacenter — those live in separate
|
||||
segments. The cluster name is project identity, not location.
|
||||
|
||||
### Cost of renaming after deployment
|
||||
|
||||
A rename touches all of:
|
||||
|
||||
1. `clusters.auto.tfvars` map key
|
||||
2. Inventory directory name
|
||||
3. Every hostname (short + FQDN) and every SSH `known_hosts` entry for every operator
|
||||
4. 1P item titles (`<OLD>_CEPH_*` → `<NEW>_CEPH_*`) including the SSH Key item
|
||||
5. cephadm cluster identity (requires cluster rebuild in the common case)
|
||||
6. `ansible_ssh_key` path in the tfvars (`~/.ssh/id_ed25519_<cluster>`) and
|
||||
the mapping in `scripts/install-ssh-keys.sh`
|
||||
7. Any DNS records and external systems that reference the hostnames
|
||||
|
||||
Expect hours-to-days of work, cluster downtime, and coordination with
|
||||
every consumer of the cluster's S3/dashboards/etc. Painbox's rename
|
||||
(from `painbox-osd-5c3cac.lab.*` to `painbox-ceph-evelyn.dev.*`) was
|
||||
cheap only because painbox is idle — a running production Ceph cluster
|
||||
makes this a multi-week project.
|
||||
|
||||
**Pick once. Pick deliberately.**
|
||||
|
||||
## Host naming
|
||||
|
||||
Each host entry in `hosts = [...]` can either declare a name or omit it
|
||||
to let TF pick one from the wordlist. Both paths are first-class;
|
||||
different clusters use different paths based on operator preference.
|
||||
|
||||
### Operator-declared
|
||||
|
||||
Declare the name explicitly in the tfvars. Used when the operator has a
|
||||
specific name in mind — typically because it's been spoken during
|
||||
planning and the team already uses it.
|
||||
|
||||
```hcl
|
||||
hosts = [
|
||||
{ name = "laurel", bond_ip = "10.10.10.90", bootstrap = true },
|
||||
{ name = "lawson", bond_ip = "10.10.10.91" },
|
||||
{ name = "samara", bond_ip = "10.10.10.92" },
|
||||
]
|
||||
```
|
||||
|
||||
Sietch uses this path.
|
||||
|
||||
Names must be unique **within a cluster**, not globally. A future
|
||||
`mesa-ceph-laurel` can coexist with `sietch-ceph-laurel` — the FQDN
|
||||
disambiguates.
|
||||
|
||||
### Auto-picked from wordlist
|
||||
|
||||
Omit `name` (leave the field absent) and the module picks from
|
||||
`tf/shared/modules/ceph-cluster/wordlist.txt` (923 words) via
|
||||
`random_shuffle`, seeded by `cluster_name` + `name_seed`. Picks are
|
||||
stable across subsequent applies.
|
||||
|
||||
```hcl
|
||||
hosts = [
|
||||
{ bond_ip = "157.180.105.198", bootstrap = true }, # TF picks
|
||||
]
|
||||
```
|
||||
|
||||
Painbox uses this path — `hosts[0]` has no name, and TF picked `evelyn`
|
||||
on first apply, yielding `painbox-ceph-evelyn`.
|
||||
|
||||
Operator-declared names are excluded from the available pool to prevent
|
||||
collisions within the cluster.
|
||||
|
||||
### Stability rules (either path, or mixed)
|
||||
|
||||
- **Add new hosts at the tail of the list.** Auto-picked names are
|
||||
positional — `hosts[0]` gets `random_shuffle.result[0]`, `hosts[1]`
|
||||
gets `result[1]`, and so on. Inserting a new entry at position 0 would
|
||||
shift every subsequent host's result-index. Always append.
|
||||
- **Never bump `name_seed` once a cluster has deployed hosts.** Bumping
|
||||
re-rolls every auto-picked name in the cluster, which cascades into
|
||||
certs, SSH `known_hosts`, 1P items, DNS, cephadm identity.
|
||||
- **Mixing paths has a non-obvious side effect.** Auto-picked names are
|
||||
drawn from `available_words = wordlist − explicit_names`. Adding or
|
||||
removing an *operator-declared* host changes `explicit_names`, which
|
||||
changes the shuffle input length, which re-permutes the result. An
|
||||
auto-picked host at `hosts[2]` could get a different name even if
|
||||
nothing about its own entry changed. The safe patterns:
|
||||
1. All-operator-declared within a cluster (sietch's model), or
|
||||
2. All-auto-picked within a cluster (painbox's model).
|
||||
Mixed works for initial setup but complicates later add/remove.
|
||||
- **Converting between paths after deploy** (e.g., adding `name = "evelyn"`
|
||||
to a host that previously auto-picked `evelyn`) **does not preserve the
|
||||
name** despite appearing to match — it rewrites `available_words` and
|
||||
re-rolls every other auto-pick. Only do this if you're prepared to pin
|
||||
every auto-named host in the same apply, or accept the cluster-wide
|
||||
rename.
|
||||
|
||||
## Inventory directory naming
|
||||
|
||||
TF renders inventory directories as:
|
||||
|
||||
```
|
||||
inventories/<cluster>-ceph.<env>.<dc>.<provider>/
|
||||
```
|
||||
|
||||
The `-ceph` segment is hardcoded in the ceph-cluster module regardless of
|
||||
`role_in_hostname`. Keeps all Ceph-project inventory paths grep-matchable
|
||||
as `*-ceph.*` — even a hypothetical cluster with `role_in_hostname = "osd"`
|
||||
(hostnames `mesa-osd-*`) still renders `mesa-ceph.prod.fsn.htz/`.
|
||||
|
||||
Defined in `tf/deployment/<env>/ceph/main.tf` (`local.inventory_dirs`).
|
||||
|
||||
## 1Password item naming
|
||||
|
||||
Items are titled `<CLUSTER>_CEPH_<specifier>` (SHOUTY_SNAKE_CASE). The
|
||||
`<CLUSTER>_CEPH_*` prefix is hardcoded in
|
||||
`tf/shared/modules/ceph-cluster/main.tf` (`local.secret_prefix`) — same
|
||||
project-scoping rationale as the inventory directory.
|
||||
|
||||
Per cluster, the expected item set:
|
||||
|
||||
| Category | Title | Source of values |
|
||||
|----------|--------------------------------------------------------|-----------------------------------------------------|
|
||||
| Password | `<CLUSTER>_CEPH_OPS_PASSWORD` | `op item create --generate-password` at setup |
|
||||
| Password | `<CLUSTER>_CEPH_DASHBOARD_PASSWORD` | same |
|
||||
| Password | `<CLUSTER>_CEPH_GRAFANA_PASSWORD` | same |
|
||||
| Password | `<CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY` | same |
|
||||
| Password | `<CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY` | same |
|
||||
| SSH Key | `<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY` | `op item create --category "SSH Key" --ssh-generate-key=ed25519` |
|
||||
| Document | `<CLUSTER>_CEPH_RGW_TLS_CERT` | populated by `mise run capture` post-deploy |
|
||||
| Document | `<CLUSTER>_CEPH_RGW_TLS_KEY` | same |
|
||||
| Document | `<CLUSTER>_CEPH_CLIENT_ADMIN_KEYRING` | same |
|
||||
|
||||
Passwords + SSH Key are created at cluster-add time (step 6 of
|
||||
[adding-a-cluster.md](adding-a-cluster.md)). DR Documents are upserted
|
||||
automatically on the first `mise run capture` after deploy (step 10).
|
||||
|
||||
For the full consumption flow (which Ansible variable each item maps to,
|
||||
which role reads it) see [secrets.md](secrets.md).
|
||||
|
||||
## Workstation SSH key filenames
|
||||
|
||||
The operator-side private key path is derived from the cluster name:
|
||||
|
||||
```
|
||||
~/.ssh/id_ed25519_<cluster>
|
||||
```
|
||||
|
||||
Examples: `~/.ssh/id_ed25519_sietch`, `~/.ssh/id_ed25519_painbox`. Per
|
||||
[ADR-010](adr/010-ssh-keys-in-1password.md), the keypair lives in
|
||||
`<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY` in 1P and is installed via
|
||||
`scripts/install-ssh-keys.sh <cluster>`.
|
||||
|
||||
The `ansible_ssh_key` field in the cluster's tfvars entry must match this
|
||||
path. If you choose a non-default filename (unusual), update both together
|
||||
and also adjust the mapping in `scripts/install-ssh-keys.sh`.
|
||||
@@ -0,0 +1,441 @@
|
||||
# Patterns
|
||||
|
||||
Project-specific Ansible idioms. This doc skips generic Ansible hygiene
|
||||
(FQCN, `set -o pipefail`, etc. — those are table stakes) and focuses on
|
||||
patterns that are non-obvious or specific to how this Ceph automation is
|
||||
built.
|
||||
|
||||
For how patterns wire into the wrapper + secrets flow, see
|
||||
[scripts.md](scripts.md). For role structure and the pre-submit checklist,
|
||||
see [adding-a-role.md](adding-a-role.md).
|
||||
|
||||
---
|
||||
|
||||
## Check-then-set against the cluster
|
||||
|
||||
**Problem:** `ceph config set` always reports `changed`. Running it
|
||||
unconditionally makes every play dirty and obscures real drift.
|
||||
|
||||
**Pattern:** read the current value, compare to the expected value, apply
|
||||
only if different. The comparison is the key — raw `ceph config set` with
|
||||
`changed_when: true` is a lie.
|
||||
|
||||
Canonical example — `roles/ceph_tuning/tasks/main.yml`:
|
||||
|
||||
```yaml
|
||||
- name: Read current OSD config values
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
echo "recovery_max_active=$(ceph config get osd osd_recovery_max_active)"
|
||||
# ...
|
||||
register: current_osd_config
|
||||
changed_when: false
|
||||
|
||||
- name: Apply Ceph config values
|
||||
ansible.builtin.command: >
|
||||
ceph config set {{ item.section }} {{ item.key }} {{ item.value }}
|
||||
loop:
|
||||
- { section: osd, key: osd_recovery_max_active, value: "{{ ... }}",
|
||||
current: "{{ osd_cfg.recovery_max_active | default('') | float }}" }
|
||||
when: item.current | float != item.value | float
|
||||
changed_when: true
|
||||
```
|
||||
|
||||
Other instances: `rgw.yml` (zonegroup hostnames), `crush-rules.yml`
|
||||
(rule existence), `monitoring.yml` (module enable check).
|
||||
|
||||
### Float comparison gotcha
|
||||
|
||||
`ceph config get osd osd_recovery_sleep_hdd` returns `0.100000`, but the
|
||||
Ansible variable is `0.1`. String comparison fails; integer comparison
|
||||
truncates. Always cast both sides to `| float` before comparing. Affects
|
||||
any Ceph config value returned with trailing zeros
|
||||
(`osd_recovery_sleep_hdd`, `osd_deep_scrub_interval`, etc.).
|
||||
|
||||
---
|
||||
|
||||
## Marker-driven idempotency
|
||||
|
||||
**Problem:** provisioning is destructive and multi-phase. A crash during
|
||||
phase 6 must not re-wipe disks on the next run. But the role still needs
|
||||
to handle a completely fresh node.
|
||||
|
||||
**Pattern:** write a JSON marker at the end of provisioning. On
|
||||
subsequent runs, check for the marker and skip completed phases.
|
||||
|
||||
The marker filename (`/etc/ceph-provisioned.json`) is project-scoped —
|
||||
every Ceph cluster writes the same filename. The marker's *contents*
|
||||
identify which cluster + host the machine belongs to (hostname, fqdn,
|
||||
bond_ip, cluster_name, SSD serials, provisioned_at timestamp).
|
||||
|
||||
**Template:** `roles/provision_host/templates/ceph-provisioned.json.j2`.
|
||||
|
||||
**Resume gate** — `roles/provision_host/tasks/main.yml`:
|
||||
|
||||
```yaml
|
||||
- name: Check if provisioning marker is present in chroot
|
||||
ansible.builtin.stat:
|
||||
path: "{{ provision_mnt }}{{ provision_marker_path }}"
|
||||
register: marker_stat
|
||||
|
||||
- name: Set provisioning_done fact
|
||||
ansible.builtin.set_fact:
|
||||
provisioning_done: "{{ marker_stat.stat.exists | bool }}"
|
||||
```
|
||||
|
||||
Then every chroot phase is gated on `when: not provisioning_done`.
|
||||
|
||||
`disks.yml` additionally handles the "md array already assembled but
|
||||
nothing mounted" case — mounts, reads the marker, validates the hostname
|
||||
matches `inventory_hostname`, and either resumes (marker matches) or
|
||||
unmounts and re-wipes (marker missing or wrong host).
|
||||
|
||||
**Use when:** any multi-step destructive workflow where partial
|
||||
completion must be resumable. The marker must contain enough identity
|
||||
information to distinguish "this node's previous run" from "a different
|
||||
node's leftover state."
|
||||
|
||||
---
|
||||
|
||||
## Block/rescue cleanup
|
||||
|
||||
**Problem:** if provisioning fails mid-chroot (e.g., debootstrap network
|
||||
error), bind mounts at `/mnt/dev`, `/mnt/proc`, `/mnt/sys` remain active.
|
||||
The next run fails because it can't cleanly remount.
|
||||
|
||||
**Pattern:** wrap the phase sequence in `block/rescue`. The rescue
|
||||
includes a shared `unmount.yml` that tears down mounts in reverse order,
|
||||
then re-raises the failure.
|
||||
|
||||
```yaml
|
||||
- name: Provisioning phases
|
||||
block:
|
||||
- ansible.builtin.import_tasks: prerequisites_live.yml
|
||||
- ansible.builtin.import_tasks: disks.yml
|
||||
# ... phases 4-8 ...
|
||||
- ansible.builtin.import_tasks: finalize.yml
|
||||
rescue:
|
||||
- name: Unmount /mnt hierarchy on failure (best-effort cleanup)
|
||||
ansible.builtin.include_tasks: unmount.yml
|
||||
- name: Re-raise failure
|
||||
ansible.builtin.fail:
|
||||
msg: "Provisioning phase failed for {{ inventory_hostname }}."
|
||||
```
|
||||
|
||||
The unmount tasks use `failed_when: false` — if a path isn't mounted, we
|
||||
just want to keep going.
|
||||
|
||||
---
|
||||
|
||||
## Conditional features
|
||||
|
||||
**Problem:** not every cluster needs iSCSI, NFS, CPU governor tuning, or
|
||||
centralized logging. These features should be zero-overhead when
|
||||
disabled.
|
||||
|
||||
**Pattern:** gate on `<feature>_enabled | bool` with defaults of `false`.
|
||||
|
||||
Current feature flags:
|
||||
|
||||
| Feature | Toggle | Default | Consumer |
|
||||
|---|---|---|---|
|
||||
| CPU governor | `ceph_cpu_governor_enabled` | false | `hardware_tuning` |
|
||||
| Centralized logging | `ceph_logging_enabled` | false | `os_tuning` |
|
||||
| iSCSI firewall | `ceph_firewall_iscsi_enabled` | false | `security` |
|
||||
| NFS firewall | `ceph_firewall_nfs_enabled` | false | `security` |
|
||||
| Firewall overall | `ceph_firewall_enabled` | true | `security` |
|
||||
| RGW TLS | `ceph_rgw_ssl` | false | `ceph_deploy/rgw` |
|
||||
| Audit logging | `ceph_audit_enabled` | true | `ceph_tuning` |
|
||||
| SSH open to all sources (dev) | `ceph_firewall_ssh_any_source` | true (dev) | `security` |
|
||||
| Weekly fstrim timer (SSDs) | `ceph_enable_fstrim_timer` | true | `hardware_tuning` |
|
||||
| LVM device filter | `ceph_lvm_filter_enabled` | false | `hardware_tuning` |
|
||||
|
||||
The `| bool` filter is mandatory. Ansible may pass booleans as strings
|
||||
from inventory or extra-vars; without `| bool`, the string `"false"` is
|
||||
truthy.
|
||||
|
||||
**In Jinja templates** (e.g. `nftables.conf.j2`):
|
||||
|
||||
```jinja2
|
||||
{% if ceph_firewall_iscsi_enabled | bool %}
|
||||
ip saddr {{ net }} tcp dport {{ ceph_firewall_iscsi_port }} accept
|
||||
{% endif %}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Drift detection pattern
|
||||
|
||||
`drift.yml` is a read-only play that compares expected state against
|
||||
live cluster. Three steps:
|
||||
|
||||
1. **Load all role defaults** via `vars_files` — gives drift detection
|
||||
access to expected values without depending on any role's execution:
|
||||
|
||||
```yaml
|
||||
vars_files:
|
||||
- roles/baseline/defaults/main.yml
|
||||
- roles/os_tuning/defaults/main.yml
|
||||
# ...
|
||||
```
|
||||
|
||||
2. **Accumulate results** into a list via `set_fact`:
|
||||
|
||||
```yaml
|
||||
drift_results: >-
|
||||
{{ drift_results + [{
|
||||
'category': 'sysctl',
|
||||
'item': item.item.key,
|
||||
'expected': item.item.expected | string,
|
||||
'actual': item.stdout | trim,
|
||||
'match': (item.stdout | trim) == (item.item.expected | string)
|
||||
}] }}
|
||||
```
|
||||
|
||||
3. **Generate a formatted report** using Jinja in `set_fact`.
|
||||
|
||||
Categories checked: sysctl values, HDD/SSD I/O schedulers, nftables
|
||||
policy, SSH `PasswordAuthentication`, ops sudo config, OSD status, MON
|
||||
quorum, RGW daemon count, cluster health, Ceph config values.
|
||||
|
||||
**Use when:** building read-only comparison plays. The pattern generalizes
|
||||
to any "expected vs actual" audit.
|
||||
|
||||
---
|
||||
|
||||
## CEPH_ENV as inventory selector
|
||||
|
||||
Wrappers and downstream scripts derive the cluster's paths from the
|
||||
`CEPH_ENV` environment variable, which points at the TF-rendered
|
||||
inventory file. Cluster identity is authoritative in
|
||||
`clusters.auto.tfvars`; `CEPH_ENV` is the runtime pointer.
|
||||
|
||||
```
|
||||
CEPH_ENV = inventories/sietch-ceph.dev.austin.int/inventory.ini
|
||||
|
|
||||
dirname -> inventories/sietch-ceph.dev.austin.int
|
||||
|
|
||||
+ "/secrets.yml.tpl" -> op inject input
|
||||
```
|
||||
|
||||
**`scripts/ansible-play.sh`** derives the secrets template path as
|
||||
`$(dirname $CEPH_ENV)/secrets.yml.tpl` and fails closed if either file
|
||||
is missing. See [scripts.md](scripts.md) for the full contract.
|
||||
|
||||
**Destroy task** in `.mise.toml` extracts the domain for the safety gate:
|
||||
|
||||
```bash
|
||||
CLUSTER_ID=$(basename "$CEPH_ENV_DIR") # sietch-ceph.dev.austin.int
|
||||
DOMAIN=${CLUSTER_ID#*-ceph.}.futo.cloud # dev.austin.int.futo.cloud
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Placement group logic
|
||||
|
||||
MON placement strategy varies by cluster size. A 2-node cluster should
|
||||
NOT run 2 MONs (no quorum majority possible). A 3+ node cluster should
|
||||
run MON on all nodes.
|
||||
|
||||
`roles/ceph_deploy/tasks/placement.yml`:
|
||||
|
||||
```yaml
|
||||
- name: Calculate MON placement
|
||||
ansible.builtin.set_fact:
|
||||
mon_hosts: >-
|
||||
{%- if groups['ceph_mon'] | default([]) | length > 0 -%}
|
||||
{{ groups['ceph_mon'] | map('extract', hostvars, 'hostname_short') | join(',') -}}
|
||||
{%- elif active_hosts.stdout.split(',') | length <= 2 -%}
|
||||
{{ hostvars[groups['ceph_bootstrap'][0]]['hostname_short'] -}}
|
||||
{%- else -%}
|
||||
{{ active_hosts.stdout -}}
|
||||
{%- endif -%}
|
||||
```
|
||||
|
||||
Three branches: (1) explicit `ceph_mon` inventory group → use those;
|
||||
(2) ≤ 2 active hosts → single MON (bootstrap only, avoids 2-MON
|
||||
quorum fragility); (3) 3+ active hosts → MON on all active hosts.
|
||||
|
||||
MGR always deploys on all active hosts — standby MGRs are harmless and
|
||||
provide fast failover.
|
||||
|
||||
---
|
||||
|
||||
## Shell + `changed_when` discipline
|
||||
|
||||
Two rules worth stating explicitly because they're the most common
|
||||
lint-clean failures:
|
||||
|
||||
1. **`changed_when: false`** for read-only commands (checks, queries,
|
||||
status).
|
||||
2. **`changed_when: true`** only when guarded by `when:` — the task only
|
||||
runs when something needs to change. Unguarded `changed_when: true`
|
||||
reports "changed" on every run; ansible-lint catches this.
|
||||
3. **Output-based `changed_when`** for shell tasks that may or may not
|
||||
change state:
|
||||
|
||||
```yaml
|
||||
changed_when: "'CHANGED' in zonegroup_hostnames.stdout"
|
||||
```
|
||||
|
||||
4. **`failed_when` with fallthrough** for commands where a specific
|
||||
error is expected and acceptable:
|
||||
|
||||
```yaml
|
||||
failed_when:
|
||||
- pg_data_result.rc != 0
|
||||
- "'is not >= current' not in pg_data_result.stderr | default('')"
|
||||
```
|
||||
|
||||
Every `shell`/`command` task in this codebase sets one of these — no bare
|
||||
`command:` without a `changed_when`. Lint enforces it.
|
||||
|
||||
`no_log: true` on every task handling passwords, keys, or credentials.
|
||||
Ansible output is committed to `ansible.log` and displayed to operators
|
||||
— secrets must never land there.
|
||||
|
||||
---
|
||||
|
||||
## Cephadm service specs over imperative loops
|
||||
|
||||
**Problem:** the role's first instinct is "iterate every disk / daemon /
|
||||
service in Ansible and run `cephadm` per item." This couples the role
|
||||
tightly to per-host hardware shape (path composition, partition layout,
|
||||
LV vs disk topology) and breaks on any new cluster shape — the original
|
||||
sietch-shape `osds.yml` failed on painbox's PCI-ATA + LV-backed SSD OSD
|
||||
topology because `sas_path_prefix` was hardcoded into the path
|
||||
composition.
|
||||
|
||||
**Pattern:** for any cephadm-managed surface (OSDs, RGW, MON/MGR
|
||||
placement, monitoring), render a **declarative service spec** and apply
|
||||
it via `ceph orch apply -i <spec>.yaml`. Cephadm handles per-disk
|
||||
discovery, daemon lifecycle, encryption, LVM, etc. internally. The role
|
||||
becomes a thin renderer + applier; hardware shape moves into the
|
||||
template's Jinja conditional, not the role logic.
|
||||
|
||||
**Examples in this codebase:**
|
||||
|
||||
- `templates/rgw-spec.yaml.j2` + `tasks/rgw.yml`'s `ceph orch apply -i`
|
||||
task — RGW daemon placement spec.
|
||||
- `templates/osd-spec.yml.j2` + `tasks/osds.yml`'s
|
||||
`ceph orch apply osd -i` task — OSD service spec; per-host documents
|
||||
in a multi-doc YAML; Jinja conditional handles sietch-shape vs
|
||||
painbox-shape path composition.
|
||||
|
||||
**Template skeleton:**
|
||||
|
||||
```jinja2
|
||||
{% for host in groups['ceph_nodes'] %}
|
||||
{% set h = hostvars[host] %}
|
||||
---
|
||||
service_type: <kind>
|
||||
service_id: {{ h.hostname_short }}-<role>
|
||||
placement:
|
||||
hosts: [{{ h.hostname_short }}]
|
||||
spec:
|
||||
<kind-specific fields>
|
||||
{% endfor %}
|
||||
```
|
||||
|
||||
**Apply pattern in tasks/*.yml:**
|
||||
|
||||
```yaml
|
||||
- name: Render cephadm service spec
|
||||
ansible.builtin.template:
|
||||
src: <kind>-spec.yml.j2
|
||||
dest: /etc/ceph/<kind>-spec.yml
|
||||
mode: '0644'
|
||||
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
|
||||
run_once: true
|
||||
|
||||
- name: Apply cephadm service spec
|
||||
ansible.builtin.command: ceph orch apply <kind> -i /etc/ceph/<kind>-spec.yml
|
||||
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
|
||||
run_once: true
|
||||
changed_when: ...
|
||||
```
|
||||
|
||||
**Use when:** the surface you're managing is a cephadm-orch-supported
|
||||
service type (`host`, `mon`, `mgr`, `osd`, `rgw`, `mds`, `nfs`,
|
||||
`prometheus`, `grafana`, `alertmanager`, `node-exporter`,
|
||||
`ceph-exporter`, etc.). Don't use for surfaces cephadm doesn't manage
|
||||
declaratively (CRUSH rules, pools, ceph config tunables, RGW realm/zone
|
||||
setup, S3 user creation) — those still need imperative `ceph` /
|
||||
`radosgw-admin` calls.
|
||||
|
||||
**Trade-off vs imperative loops:** debugging "why isn't this disk
|
||||
becoming an OSD?" is harder — there's no per-disk log line. Check
|
||||
`ceph orch ls` / `ceph orch ps` / `ceph cephadm osd activate <host>
|
||||
--dry-run` instead. Worth the trade-off because the role becomes
|
||||
hardware-shape-agnostic.
|
||||
|
||||
See [ADR-011](adr/011-cephadm-osd-service-specs.md) for the full
|
||||
decision record on the OSD-path migration.
|
||||
|
||||
---
|
||||
|
||||
## Anti-patterns
|
||||
|
||||
### `changed_when: true` without a `when` guard
|
||||
|
||||
```yaml
|
||||
# BAD — reports changed on every run even when idempotent
|
||||
- ansible.builtin.command: ceph config set osd foo bar
|
||||
changed_when: true
|
||||
|
||||
# GOOD — only runs when needed, so changed_when: true is accurate
|
||||
- ansible.builtin.command: ceph config set osd foo bar
|
||||
when: current_foo != 'bar'
|
||||
changed_when: true
|
||||
```
|
||||
|
||||
### Shell without pipefail
|
||||
|
||||
```yaml
|
||||
# BAD — if `ceph osd dump` fails, grep runs on empty input and task succeeds
|
||||
- ansible.builtin.shell: ceph osd dump | grep noin
|
||||
|
||||
# GOOD — pipefail propagates the ceph failure
|
||||
- ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
ceph osd dump | grep noin
|
||||
args:
|
||||
executable: /bin/bash
|
||||
```
|
||||
|
||||
### Hardcoded site-specific values in roles
|
||||
|
||||
```yaml
|
||||
# BAD — in a role's tasks/main.yml
|
||||
- ansible.builtin.command: ceph config set osd osd_recovery_max_active 1
|
||||
|
||||
# GOOD — value comes from defaults, overridable via group_vars
|
||||
- ansible.builtin.command: >
|
||||
ceph config set osd osd_recovery_max_active
|
||||
{{ ceph_osd_recovery_max_active }}
|
||||
```
|
||||
|
||||
Roles use `defaults/main.yml` for all tunables. Site-specific values live
|
||||
in `inventories/<cluster>/group_vars/all/vars.yml`.
|
||||
|
||||
### Using `ansible_play_batch` / `ansible_play_hosts` for placement specs
|
||||
|
||||
```yaml
|
||||
# BAD — --limit shrinks play_batch, cephadm removes daemons from omitted hosts
|
||||
- ansible.builtin.command:
|
||||
ceph orch apply rgw --placement="{{ ansible_play_batch | join(',') }}"
|
||||
```
|
||||
|
||||
cephadm is declarative: applying a smaller placement list REMOVES daemons
|
||||
from hosts not in the list. Always use `groups['ceph_nodes']` (the full
|
||||
inventory group) for placement specs, never `ansible_play_batch` or
|
||||
`ansible_play_hosts`. The comment at
|
||||
`roles/ceph_deploy/tasks/rgw.yml:466` explains the failure mode in detail.
|
||||
|
||||
### Running plays with bare `ansible-playbook`
|
||||
|
||||
Every playbook in this project consumes op-injected secrets. Running
|
||||
`ansible-playbook foo.yml` directly skips the wrapper, `op inject` never
|
||||
runs, and `vault_*` variables are empty — tasks that need them fail with
|
||||
confusing errors. Always `scripts/ansible-play.sh <playbook.yml>`. See
|
||||
[scripts.md](scripts.md) for the full contract.
|
||||
@@ -0,0 +1,208 @@
|
||||
# Runbook: Add a Node to an Existing Cluster
|
||||
|
||||
**When:** Expanding cluster capacity or replacing a failed chassis.
|
||||
|
||||
**Time estimate:** 30-60 minutes (Austin physical), 15-30 minutes (Hetzner).
|
||||
|
||||
**Prerequisites:**
|
||||
- Cluster is healthy (`ceph health` returns `HEALTH_OK` or understood warnings)
|
||||
- `op` session live (desktop unlocked or `OP_SERVICE_ACCOUNT_TOKEN` set)
|
||||
- SSH access to existing cluster nodes working
|
||||
|
||||
---
|
||||
|
||||
## 1. Declare the new host in TF
|
||||
|
||||
Pick an unused name, or omit `name` to let TF auto-pick from the
|
||||
923-word list seeded per-cluster (see [docs/naming.md](../naming.md)).
|
||||
Edit `tf/deployment/dev/ceph/clusters.auto.tfvars` and append to the
|
||||
target cluster's `hosts` list:
|
||||
|
||||
```hcl
|
||||
sietch = {
|
||||
# ...
|
||||
hosts = [
|
||||
{ name = "laurel", bond_ip = "10.10.10.90", bootstrap = true },
|
||||
{ name = "lawson", bond_ip = "10.10.10.91" },
|
||||
{ name = "samara", bond_ip = "10.10.10.92" },
|
||||
{ name = "maxton", bond_ip = "10.10.10.93" }, # NEW
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Apply:
|
||||
|
||||
```bash
|
||||
mise run tf:apply
|
||||
```
|
||||
|
||||
This re-renders `inventory.ini` with the new host in `[ceph_nodes]`,
|
||||
`[ceph_mon]`, and `[ceph_join]`. Bootstrap assignment doesn't change —
|
||||
still pinned to the first/explicitly-declared bootstrap host. Add new
|
||||
hosts at the **tail** of the list so existing auto-picked names keep
|
||||
their positions.
|
||||
|
||||
## 2. Create host_vars
|
||||
|
||||
```bash
|
||||
cd inventories/sietch-ceph.dev.austin.int
|
||||
cp host_vars/example.yml host_vars/sietch-ceph-<name>.yml
|
||||
```
|
||||
|
||||
Edit the new file. Every field is node-specific and must match the physical
|
||||
hardware:
|
||||
|
||||
| Field | How to find it |
|
||||
|---|---|
|
||||
| `hostname_short` | The full `sietch-ceph-<name>` from step 1 |
|
||||
| `bond_ip` | Next available IP in 10.10.10.0/24. Austin convention: yucca-N = 10.10.10.9N |
|
||||
| `sas_path_prefix` | SSH into node, run `ls /dev/disk/by-path/ \| grep sas` |
|
||||
| `ssd1_phy` / `ssd2_phy` | Identify SSD PHY positions from `lsscsi -t` output |
|
||||
| `ceph_db_vg1/2` | Name the VGs by slot, e.g., `ceph-db-rear12`, `ceph-db-rear13` |
|
||||
| `ceph_hdd_osds` | Map each populated HDD bay to its PHY and corresponding db-slot LV |
|
||||
| `ceph_ssd_osds` | SSD partition 6 on each SSD (no separate block.db) |
|
||||
|
||||
## 3. Confirm inventory regenerated correctly
|
||||
|
||||
After `tofu apply` in step 1, inspect the rendered inventory:
|
||||
|
||||
```bash
|
||||
cat inventories/sietch-ceph.dev.austin.int/inventory.ini
|
||||
```
|
||||
|
||||
The new host should appear in:
|
||||
- `[ceph_nodes]` — all cluster members
|
||||
- `[ceph_mon]` — MON/MGR placement
|
||||
- `[ceph_join]` — everything except the bootstrap host
|
||||
|
||||
Bootstrap host is unchanged. TF never moves an existing bootstrap assignment.
|
||||
|
||||
## 4. Provision the OS
|
||||
|
||||
### Austin (physical servers)
|
||||
|
||||
1. Boot the server to the Debian 12 live image via iDRAC virtual console
|
||||
2. Verify the live image is reachable on the node's bond IP
|
||||
3. Run provisioning:
|
||||
|
||||
```bash
|
||||
CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory-provision.ini \
|
||||
scripts/ansible-play.sh provision.yml \
|
||||
-e confirm_wipe=true \
|
||||
--limit sietch-ceph-<name>
|
||||
```
|
||||
|
||||
4. Wait for reboot and verify SSH access as `ansible-iac`:
|
||||
|
||||
```bash
|
||||
ssh -i ~/.ssh/id_ed25519_sietch ansible-iac@sietch-ceph-<name> hostname -f
|
||||
```
|
||||
|
||||
Expected output: `sietch-ceph-<name>.dev.austin.int.futo.cloud`
|
||||
|
||||
### Hetzner (remote servers)
|
||||
|
||||
1. Boot into rescue mode via Hetzner Robot panel
|
||||
2. SSH into rescue system as root
|
||||
3. Run installimage for Debian 12 Bookworm
|
||||
4. Run the post-install script to configure networking and partitioning
|
||||
5. Reboot into installed OS
|
||||
6. Verify SSH access
|
||||
|
||||
## 5. Run baseline
|
||||
|
||||
Installs podman, diagnostic tools, creates the ops user, renders /etc/hosts,
|
||||
enables dbus/chrony/podman.socket:
|
||||
|
||||
```bash
|
||||
scripts/ansible-play.sh baseline.yml --limit sietch-ceph-<name>
|
||||
```
|
||||
|
||||
## 6. Apply tuning
|
||||
|
||||
OS-level sysctl, hardware I/O schedulers, CPU governor:
|
||||
|
||||
```bash
|
||||
scripts/ansible-play.sh tune-os.yml --limit sietch-ceph-<name>
|
||||
scripts/ansible-play.sh tune-hardware.yml --limit sietch-ceph-<name>
|
||||
```
|
||||
|
||||
## 7. Join node to Ceph cluster
|
||||
|
||||
This runs all deploy phases. For an existing cluster, bootstrap is skipped
|
||||
(ceph.conf already exists on the bootstrap node). The node gets joined,
|
||||
LVM is set up, OSDs are created, and services are placed:
|
||||
|
||||
```bash
|
||||
scripts/ansible-play.sh deploy-ceph.yml \
|
||||
--limit sietch-ceph-<name>,sietch-ceph-laurel
|
||||
```
|
||||
|
||||
**Important:** You must include the bootstrap node (`sietch-ceph-laurel`)
|
||||
in `--limit` because join, placement, OSD activation, and RGW service spec
|
||||
updates all run from the bootstrap node.
|
||||
|
||||
## 8. Apply Ceph tuning and security hardening
|
||||
|
||||
```bash
|
||||
scripts/ansible-play.sh tune-ceph.yml --limit sietch-ceph-laurel
|
||||
scripts/ansible-play.sh harden.yml --limit sietch-ceph-<name>
|
||||
```
|
||||
|
||||
## 9. Verify
|
||||
|
||||
### Check the node appears in the cluster
|
||||
|
||||
```bash
|
||||
ssh ansible-iac@sietch-ceph-laurel
|
||||
|
||||
ceph orch host ls
|
||||
```
|
||||
|
||||
Expected: new hostname listed with status empty (= online).
|
||||
|
||||
### Check OSDs are up
|
||||
|
||||
```bash
|
||||
ceph osd tree
|
||||
```
|
||||
|
||||
Expected: new node appears as a host bucket with its OSDs in `up` state.
|
||||
|
||||
### Check overall health
|
||||
|
||||
```bash
|
||||
ceph status
|
||||
```
|
||||
|
||||
Expected: `HEALTH_OK` or `HEALTH_WARN` with only backfill-related warnings
|
||||
(which clear as data rebalances).
|
||||
|
||||
### Check RGW is running on new node
|
||||
|
||||
```bash
|
||||
ceph orch ls --service-type rgw
|
||||
```
|
||||
|
||||
Expected: running count incremented by 1.
|
||||
|
||||
### Run drift detection
|
||||
|
||||
```bash
|
||||
mise run drift
|
||||
```
|
||||
|
||||
Expected: no drift on the new node.
|
||||
|
||||
## Rollback
|
||||
|
||||
If the node needs to be removed:
|
||||
|
||||
```bash
|
||||
# From bootstrap node
|
||||
ceph orch host drain sietch-ceph-<name> --force
|
||||
# Wait for daemons to migrate (~5 min)
|
||||
ceph orch host rm sietch-ceph-<name> --force
|
||||
```
|
||||
|
||||
Then remove the node from `inventory.ini` and delete its `host_vars` file.
|
||||
@@ -0,0 +1,178 @@
|
||||
# Runbook: Backup and Restore
|
||||
|
||||
**When:** Before major operations (upgrades, topology changes), on a
|
||||
regular schedule, or during disaster recovery.
|
||||
|
||||
**Time estimate:** Backup: 2 minutes. Restore: depends on scenario.
|
||||
|
||||
---
|
||||
|
||||
## Backup
|
||||
|
||||
### Run a backup
|
||||
|
||||
```bash
|
||||
mise run backup
|
||||
```
|
||||
|
||||
Or directly:
|
||||
|
||||
```bash
|
||||
scripts/ansible-play.sh backup-config.yml
|
||||
```
|
||||
|
||||
### What gets captured
|
||||
|
||||
Backups are written to `backups/<timestamp>/` on the Ansible controller
|
||||
(gitignored). Each backup contains:
|
||||
|
||||
| File | Contents |
|
||||
|---|---|
|
||||
| `ceph.conf` | Minimal cluster config (fsid, mon_host, auth settings) |
|
||||
| `ceph.client.admin.keyring` | Admin authentication keyring |
|
||||
| `crushmap.txt` | Decompiled CRUSH map (host/OSD topology and rules) |
|
||||
| `osd-dump.json` | Full OSD map: pool definitions, PG counts, flags, weights |
|
||||
| `mon-dump.json` | Monitor map: mon addresses, quorum members |
|
||||
| `config-dump.json` | All runtime config overrides (`ceph config dump`) |
|
||||
| `rgw-realm.json` | RGW realm, zonegroup, and zone configuration |
|
||||
| `orch-services.yaml` | cephadm service specs (RGW, monitoring, crash, etc.) |
|
||||
| `orch-hosts.yaml` | Cluster host list with addresses and labels |
|
||||
|
||||
### Backup schedule
|
||||
|
||||
No automated schedule is configured. Run manually:
|
||||
|
||||
- Before any cluster topology change (add/remove node or OSD)
|
||||
- Before Ceph version upgrades
|
||||
- Before CRUSH map modifications
|
||||
- Weekly during active development
|
||||
|
||||
### Verify a backup
|
||||
|
||||
```bash
|
||||
ls -la backups/$(ls -t backups/ | head -1)/
|
||||
```
|
||||
|
||||
Check that all files are present and non-empty. The admin keyring and
|
||||
ceph.conf are the most critical -- without them, cluster access is lost.
|
||||
|
||||
---
|
||||
|
||||
## Restore Scenarios
|
||||
|
||||
### Scenario 1: Lost ceph.conf / admin keyring on a single node
|
||||
|
||||
**Cause:** Accidental deletion, failed re-provision.
|
||||
|
||||
**Fix:** cephadm automatically distributes ceph.conf and the admin keyring
|
||||
to managed hosts. Force redistribution:
|
||||
|
||||
```bash
|
||||
# From bootstrap node
|
||||
ceph cephadm config-check enable
|
||||
ceph orch host rescan <hostname>
|
||||
```
|
||||
|
||||
Or manually copy from the backup:
|
||||
|
||||
```bash
|
||||
scp backups/<timestamp>/ceph.conf ansible-iac@<host>:/etc/ceph/ceph.conf
|
||||
scp backups/<timestamp>/ceph.client.admin.keyring ansible-iac@<host>:/etc/ceph/ceph.client.admin.keyring
|
||||
```
|
||||
|
||||
### Scenario 2: Lost admin keyring on ALL nodes
|
||||
|
||||
**Cause:** Full cluster purge without backup, or corruption.
|
||||
|
||||
**Fix:** Restore the keyring from the backup to the bootstrap node:
|
||||
|
||||
```bash
|
||||
scp backups/<timestamp>/ceph.client.admin.keyring \
|
||||
ansible-iac@sietch-ceph-laurel:/etc/ceph/
|
||||
|
||||
ssh ansible-iac@sietch-ceph-laurel
|
||||
sudo chmod 600 /etc/ceph/ceph.client.admin.keyring
|
||||
sudo chown ceph:ceph /etc/ceph/ceph.client.admin.keyring
|
||||
```
|
||||
|
||||
Verify access is restored:
|
||||
|
||||
```bash
|
||||
ceph status
|
||||
```
|
||||
|
||||
### Scenario 3: CRUSH map corruption
|
||||
|
||||
**Cause:** Bad CRUSH rule edit, accidental tunables change.
|
||||
|
||||
**Fix:** Restore the CRUSH map from backup:
|
||||
|
||||
```bash
|
||||
# Compile the decompiled map
|
||||
crushtool -c backups/<timestamp>/crushmap.txt -o /tmp/crushmap.bin
|
||||
|
||||
# Inject it
|
||||
ceph osd setcrushmap -i /tmp/crushmap.bin
|
||||
```
|
||||
|
||||
**WARNING:** This overwrites the entire CRUSH topology. Any OSDs added
|
||||
since the backup was taken will not be in the restored map.
|
||||
|
||||
### Scenario 4: RGW realm/zone misconfiguration
|
||||
|
||||
**Cause:** Bad radosgw-admin command, zone placement errors.
|
||||
|
||||
**Fix:** Use the backup as a reference to reconstruct:
|
||||
|
||||
```bash
|
||||
cat backups/<timestamp>/rgw-realm.json | python3 -m json.tool
|
||||
```
|
||||
|
||||
Then re-apply zone placement targets, zonegroup hostnames, etc. using
|
||||
`radosgw-admin zone set` / `radosgw-admin zonegroup set` with the JSON
|
||||
from the backup piped in.
|
||||
|
||||
### Scenario 5: Full cluster rebuild (total loss)
|
||||
|
||||
**Cause:** All nodes destroyed, starting from scratch.
|
||||
|
||||
1. Re-provision all nodes (see add-node runbook)
|
||||
2. Run the full deploy pipeline:
|
||||
|
||||
```bash
|
||||
mise run deploy
|
||||
```
|
||||
|
||||
3. Restore configuration from backup:
|
||||
|
||||
```bash
|
||||
# After bootstrap, apply saved CRUSH map
|
||||
crushtool -c backups/<timestamp>/crushmap.txt -o /tmp/crushmap.bin
|
||||
ceph osd setcrushmap -i /tmp/crushmap.bin
|
||||
|
||||
# Re-apply runtime config overrides
|
||||
# Review config-dump.json and apply relevant settings
|
||||
cat backups/<timestamp>/config-dump.json | python3 -c "
|
||||
import sys, json
|
||||
for item in json.load(sys.stdin):
|
||||
section = item.get('section', 'global')
|
||||
name = item.get('name', '')
|
||||
value = item.get('value', '')
|
||||
if name and section != 'mds':
|
||||
print(f'ceph config set {section} {name} {value}')
|
||||
"
|
||||
```
|
||||
|
||||
**Note:** Object data (S3 objects, bucket contents) is stored on the OSDs
|
||||
and cannot be restored from this config backup. This backup only preserves
|
||||
cluster metadata and configuration. For data protection, rely on Ceph's
|
||||
built-in replication (size=2+) and erasure coding.
|
||||
|
||||
### Scenario 6: Restore service specs after cephadm reset
|
||||
|
||||
```bash
|
||||
ceph orch apply -i backups/<timestamp>/orch-services.yaml
|
||||
```
|
||||
|
||||
This re-deploys RGW, monitoring, and crash daemons with the saved
|
||||
placement and configuration.
|
||||
@@ -0,0 +1,170 @@
|
||||
# Runbook: Reprovision Painbox
|
||||
|
||||
**Status:** Validated 2026-04-26. Executed end-to-end during the yucca
|
||||
monorepo import PR — painbox now runs as `painbox-ceph-evelyn` on
|
||||
Bookworm with Ceph Tentacle deployed (15 OSDs up + in, all LUKS-encrypted).
|
||||
This runbook captures the procedure for future reprovisions (disaster
|
||||
recovery, OS upgrade, hardware refresh) and remains the canonical
|
||||
Hetzner-installimage flow for any future SX-class clusters.
|
||||
|
||||
**When:** any future painbox reprovision — DR scenario, kernel/OS upgrade
|
||||
that requires a fresh install, or hardware change that invalidates the
|
||||
existing layout.
|
||||
|
||||
**Time estimate:** 45-60 minutes including installimage wait + reboot
|
||||
+ post-deploy verification.
|
||||
|
||||
---
|
||||
|
||||
## Pre-flight
|
||||
|
||||
Canonical painbox identity (a future reprovision keeps this — a DR scenario
|
||||
restores the same name on the same hardware):
|
||||
|
||||
| Item | Value |
|
||||
|----------------|----------------------------------------------------------------|
|
||||
| Inventory dir | `inventories/painbox-ceph.dev.hel.htz/` |
|
||||
| Short hostname | `painbox-ceph-evelyn` |
|
||||
| FQDN | `painbox-ceph-evelyn.dev.hel.htz.futo.cloud` |
|
||||
| Public IP | `157.180.105.198` |
|
||||
| OS image | Debian 12 Bookworm (Hetzner installimage tarball) |
|
||||
| SSH key path | `~/.ssh/id_ed25519_painbox` (per [ADR-010](../adr/010-ssh-keys-in-1password.md)) |
|
||||
|
||||
1P items (`PAINBOX_CEPH_*`) already exist in `yucca_tf_dev`, including
|
||||
the `PAINBOX_CEPH_ANSIBLE_IAC_SSH_KEY` SSH Key item.
|
||||
|
||||
> **Historical note:** the first reprovision under this identity ran
|
||||
> 2026-04-26 as a rename from `painbox-osd-5c3cac.lab.hel.htz.futo.cloud`
|
||||
> + old SSH key `~/.ssh/id_ed25519_ceph-painbox-lab-hel-htz` to the
|
||||
> values above. Subsequent reprovisions (DR, OS refresh, etc.) restore
|
||||
> the same identity — no rename involved.
|
||||
|
||||
## Steps
|
||||
|
||||
### 1. Install the ansible-iac SSH key on the operator workstation
|
||||
|
||||
Per [ADR-010](../adr/010-ssh-keys-in-1password.md), the keypair lives in
|
||||
1P and is pulled to the workstation idempotently:
|
||||
|
||||
```bash
|
||||
scripts/install-ssh-keys.sh painbox
|
||||
# Writes ~/.ssh/id_ed25519_painbox (0600) + .pub (0644)
|
||||
```
|
||||
|
||||
If the key was already installed from a previous session, this is a
|
||||
no-op (fingerprint compare + skip). See [docs/scripts.md](../scripts.md)
|
||||
for the wrapper's behavior.
|
||||
|
||||
### 2. Boot into Hetzner rescue
|
||||
|
||||
From Hetzner Robot panel: **Activate rescue system** → **Reboot**. SSH in
|
||||
as root to the rescue IP.
|
||||
|
||||
> **Hetzner rotates the rescue root password on every activation.** The
|
||||
> password shown in the Robot panel after activation is single-use for
|
||||
> that rescue session — the next activation gets a fresh one. The 1Password
|
||||
> entry `painbox-ceph-evelyn.dev.hel.htz.futo.cloud` (vault: `Yucca`) holds
|
||||
> a *cached* copy from a prior activation; **update it from the Robot
|
||||
> panel each time** before SSH'ing in, or the cached password fails auth.
|
||||
> Rescue host keys also rotate per activation, so use `-o
|
||||
> StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null` (or
|
||||
> `ssh-keygen -R <ip>` first) when connecting.
|
||||
|
||||
### 3. Render + upload installimage scripts
|
||||
|
||||
`post-install.sh` is TEMPLATED — source of truth is `post-install.sh.tpl`,
|
||||
which embeds an `op://` reference to the current
|
||||
`PAINBOX_CEPH_ANSIBLE_IAC_SSH_KEY` pubkey. Render locally so rotations
|
||||
propagate at reprovision time without edits:
|
||||
|
||||
```bash
|
||||
cd ansible/ceph/inventories/painbox-ceph.dev.hel.htz/installimage
|
||||
op inject -f -i post-install.sh.tpl -o /tmp/post-install.sh.rendered
|
||||
```
|
||||
|
||||
Upload both files to the rescue system:
|
||||
|
||||
```bash
|
||||
scp /tmp/post-install.sh.rendered root@<rescue-ip>:/tmp/post-install.sh
|
||||
scp autosetup root@<rescue-ip>:/autosetup
|
||||
```
|
||||
|
||||
### 4. Run installimage
|
||||
|
||||
```bash
|
||||
# On the rescue system
|
||||
chmod +x /tmp/post-install.sh
|
||||
installimage -a -c /autosetup -x /tmp/post-install.sh
|
||||
```
|
||||
|
||||
(`autosetup` targets `painbox-ceph-evelyn.dev.hel.htz.futo.cloud` as
|
||||
HOSTNAME statically. If you ever rotate the hostname, edit autosetup
|
||||
directly.)
|
||||
|
||||
### 5. Reboot into installed OS
|
||||
|
||||
```bash
|
||||
reboot
|
||||
```
|
||||
|
||||
SSH in with the new keypair:
|
||||
|
||||
```bash
|
||||
ssh -i ~/.ssh/id_ed25519_painbox root@157.180.105.198 hostname -f
|
||||
# Expect: painbox-ceph-evelyn.dev.hel.htz.futo.cloud
|
||||
```
|
||||
|
||||
### 6. Update `yucca_tf_dev` state (if items were renamed)
|
||||
|
||||
The `PAINBOX_CEPH_*` items in `yucca_tf_dev` are already correctly named —
|
||||
no action needed. Disaster-recovery items (`PAINBOX_CEPH_RGW_TLS_CERT`,
|
||||
`_RGW_TLS_KEY`, `_CLIENT_ADMIN_KEYRING`) are upserted by `mise run capture` after
|
||||
a successful deploy (step 8).
|
||||
|
||||
### 7. Run ansible deploy
|
||||
|
||||
```bash
|
||||
cd ~/yucca/ansible/ceph
|
||||
export CEPH_ENV=inventories/painbox-ceph.dev.hel.htz/inventory.ini
|
||||
mise run preflight # should pass
|
||||
mise run deploy # baseline → tune → deploy → tune → harden
|
||||
```
|
||||
|
||||
### 8. Verify
|
||||
|
||||
```bash
|
||||
ssh ansible-iac@painbox-ceph-evelyn.dev.hel.htz.futo.cloud sudo ceph -s
|
||||
# HEALTH_OK (single-node cluster)
|
||||
```
|
||||
|
||||
## Rollback
|
||||
|
||||
If reprovisioning fails at installimage or post-install, Hetzner rescue
|
||||
is still accessible:
|
||||
|
||||
1. Activate rescue, SSH in.
|
||||
2. Re-run installimage with the OLD scripts (preserve them in
|
||||
`docs/archive/painbox-installimage-2026-04-xx/` before overwriting).
|
||||
3. Old hostname + old SSH key restore the prior-state box.
|
||||
4. No cluster data is on this box (painbox never held Ceph data),
|
||||
so nothing to restore.
|
||||
|
||||
## Post-reprovision cleanup
|
||||
|
||||
- Delete the old SSH key file `~/.ssh/id_ed25519_ceph-painbox-lab-hel-htz*`
|
||||
from the operator workstation once verification is complete.
|
||||
- If any external DNS records pointed at `painbox-osd-5c3cac.lab.*`,
|
||||
update them to `painbox-ceph-evelyn.dev.*`.
|
||||
|
||||
(No edit to `post-install.sh` is required — the committed source of
|
||||
truth is `post-install.sh.tpl`, which embeds a live `op://` reference
|
||||
to the current `PAINBOX_CEPH_ANSIBLE_IAC_SSH_KEY/public_key`.
|
||||
Re-rendering via `op inject -f` always picks up whatever is current in
|
||||
1P at the moment of reprovision.)
|
||||
|
||||
## References
|
||||
|
||||
- `docs/runbooks/add-node.md` §Hetzner — the general installimage flow
|
||||
- `docs/adr/010-ssh-keys-in-1password.md` — SSH keypair lifecycle
|
||||
- `docs/scripts.md` — `install-ssh-keys.sh` reference
|
||||
- `inventories/painbox-ceph.dev.hel.htz/installimage/` — scripts
|
||||
@@ -0,0 +1,164 @@
|
||||
# Runbook: Recover from a Bad `tofu apply`
|
||||
|
||||
**When:** state corruption, drift, rendered files don't match cluster spec,
|
||||
or TF destroyed an item/file that shouldn't have been touched.
|
||||
|
||||
**Time estimate:** 5-30 minutes depending on severity.
|
||||
|
||||
---
|
||||
|
||||
## Common failure modes and their fixes
|
||||
|
||||
### Missing or unreadable state file
|
||||
|
||||
```
|
||||
tofu init
|
||||
# Error: failed to load state: ...
|
||||
```
|
||||
|
||||
**Cause:** S3 credentials unresolved, bucket object missing, or the backend
|
||||
config differs from what the state was written with.
|
||||
|
||||
**Fix:**
|
||||
|
||||
```bash
|
||||
# Verify S3 credentials resolve via op run
|
||||
op run --env-file=tf/.env -- env | grep AWS_
|
||||
# Expect AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY populated
|
||||
|
||||
# Verify the state object exists in the bucket
|
||||
op run --env-file=tf/.env -- \
|
||||
aws --endpoint-url=https://s3.eu-west-par.io.cloud.ovh.net/ \
|
||||
s3 ls s3://yucca-tf-state/ceph/dev/ceph/
|
||||
|
||||
# If present: re-init should pick it up
|
||||
mise run tf:init
|
||||
|
||||
# If the object is missing: state was never created or was deleted. You
|
||||
# can recover by re-applying — TF recreates local_file resources
|
||||
# (idempotent, same content; no 1P items are harmed because the module's
|
||||
# onepassword_item resources are currently dormant — see ADR-009 §2).
|
||||
mise run tf:apply
|
||||
```
|
||||
|
||||
### Rendered file on disk doesn't match tfvars
|
||||
|
||||
```bash
|
||||
cat ansible/ceph/inventories/sietch-ceph.dev.austin.int/inventory.ini
|
||||
# says ansible_user=root but tfvars says ansible-iac
|
||||
```
|
||||
|
||||
**Cause:** someone hand-edited the rendered file; TF state shows it
|
||||
unchanged; operator is surprised.
|
||||
|
||||
**Fix:**
|
||||
|
||||
```bash
|
||||
mise run tf:plan # TF shows drift
|
||||
mise run tf:apply # TF overwrites with correct content
|
||||
```
|
||||
|
||||
All TF-rendered files are gitignored — the single source of truth is
|
||||
`clusters.auto.tfvars`. Hand-edits are ephemeral.
|
||||
|
||||
### Wrong vault referenced in rendered secrets.yml.tpl
|
||||
|
||||
**Cause:** `vault` field in clusters.auto.tfvars typo'd or set to a
|
||||
vault you don't have access to.
|
||||
|
||||
**Fix:** Fix the tfvars entry, `mise run tf:apply`. `scripts/ansible-play.sh`
|
||||
will fail loudly on the next run (op inject exits non-zero on unresolvable
|
||||
references — per ADR-009 fail-closed principle).
|
||||
|
||||
### Wordlist auto-pick renamed a deployed host
|
||||
|
||||
**Cause:** `name_seed` bumped, or a new auto-named host was prepended such
|
||||
that existing auto-name indices shifted.
|
||||
|
||||
**Symptom:**
|
||||
|
||||
```bash
|
||||
mise run tf:plan
|
||||
# Plan shows: ~inventory.ini content with hostname change from painbox-ceph-evelyn to painbox-ceph-<other>
|
||||
```
|
||||
|
||||
**Do NOT apply** — renaming a deployed host cascades into SSH known_hosts,
|
||||
cephadm host registration, certs, 1P item names, DNS.
|
||||
|
||||
**Fix:** pin the existing name by adding `name = "evelyn"` to the host
|
||||
entry in `clusters.auto.tfvars`, then apply. The pinned name takes
|
||||
precedence over the shuffle output.
|
||||
|
||||
### Item disappeared from yucca_tf_dev
|
||||
|
||||
**Cause:** an operator, another consumer of `yucca_tf_dev`, or a TF
|
||||
delete-on-destroy run removed an item the Ansible side depends on.
|
||||
|
||||
**Symptom:** `op inject -f -i secrets.yml.tpl` fails with "item not found".
|
||||
|
||||
**Fix:**
|
||||
|
||||
```bash
|
||||
# Re-create the item manually via superuser SA
|
||||
SU_TOKEN=$(op read "op://yucca_tf_dev/yucca_futo_1pass_superuser_service_account/password")
|
||||
OP_SERVICE_ACCOUNT_TOKEN="$SU_TOKEN" op item create \
|
||||
--vault yucca_tf_dev \
|
||||
--category password \
|
||||
--title SIETCH_CEPH_OPS_PASSWORD \
|
||||
--generate-password='letters,digits,32'
|
||||
unset SU_TOKEN
|
||||
```
|
||||
|
||||
Then either (a) run `mise run deploy` to push the new password to the
|
||||
cluster (baseline role converges ops user pw), or (b) if the cluster
|
||||
already has the old value and you want to keep it, retrieve from a
|
||||
backup and set the item via `op item edit password=...`.
|
||||
|
||||
### TF state object corrupted or lost
|
||||
|
||||
The state lives in S3 (`yucca-tf-state` bucket, key
|
||||
`ceph/${env}/${stack}/terraform.tfstate`). Recovery options in order of
|
||||
preference:
|
||||
|
||||
**Option A — roll back via S3 versioning.** The bucket has versioning
|
||||
enabled; list prior versions and restore the last-known-good:
|
||||
|
||||
```bash
|
||||
op run --env-file=tf/.env -- \
|
||||
aws --endpoint-url=https://s3.eu-west-par.io.cloud.ovh.net/ \
|
||||
s3api list-object-versions \
|
||||
--bucket yucca-tf-state \
|
||||
--prefix ceph/dev/ceph/terraform.tfstate
|
||||
|
||||
# Identify the VersionId of a good snapshot, then:
|
||||
op run --env-file=tf/.env -- \
|
||||
aws --endpoint-url=https://s3.eu-west-par.io.cloud.ovh.net/ \
|
||||
s3api copy-object \
|
||||
--bucket yucca-tf-state \
|
||||
--copy-source 'yucca-tf-state/ceph/dev/ceph/terraform.tfstate?versionId=<VID>' \
|
||||
--key ceph/dev/ceph/terraform.tfstate
|
||||
```
|
||||
|
||||
**Option B — re-apply from clean state.** Delete the state object and
|
||||
re-init + re-apply. TF recreates the `local_file` resources (idempotent,
|
||||
same content). Safe today because `onepassword_item` resources are
|
||||
dormant (see ADR-009 §2).
|
||||
|
||||
```bash
|
||||
op run --env-file=tf/.env -- \
|
||||
aws --endpoint-url=https://s3.eu-west-par.io.cloud.ovh.net/ \
|
||||
s3 rm s3://yucca-tf-state/ceph/dev/ceph/terraform.tfstate
|
||||
|
||||
mise run tf:init
|
||||
mise run tf:apply
|
||||
```
|
||||
|
||||
Once ADR-009 §2 ships and 1P items are TF-owned, Option B becomes
|
||||
destructive — recovery will require `terraform import` of each item
|
||||
against the new state. Document that path when re-enabling.
|
||||
|
||||
## References
|
||||
|
||||
- ADR-009 — fail-closed principle
|
||||
- `tf/README.md` §"State backend" — bucket, endpoint, key path rationale
|
||||
- `architecture.md` §4 — what TF owns vs renders
|
||||
@@ -0,0 +1,80 @@
|
||||
# Runbook: Remote Hands Access
|
||||
|
||||
**When:** a remote-hands operator (on-site at the datacenter) needs to
|
||||
run playbooks against the cluster.
|
||||
|
||||
**Time estimate:** 10 minutes of setup per operator, one-time. 5 minutes
|
||||
per remote-hands session after that.
|
||||
|
||||
---
|
||||
|
||||
## The access model
|
||||
|
||||
Remote-hands operators don't touch 1P desktop sessions. They authenticate
|
||||
via a scoped service-account token, set as `OP_SERVICE_ACCOUNT_TOKEN` in
|
||||
their shell environment. `scripts/ansible-play.sh` picks it up automatically
|
||||
and uses it to resolve `secrets.yml.tpl`.
|
||||
|
||||
Which SA: the **read-only** `yucca_futo_1pass_service_account` in
|
||||
`yucca_tf_dev`. Remote hands shouldn't have write access to 1P items —
|
||||
least-privilege principle.
|
||||
|
||||
## One-time setup (per operator)
|
||||
|
||||
1. **Grant the operator access to the Immich 1P group** — the group's
|
||||
1P admin does this. Gives them read on `yucca_tf_dev` (enough to read
|
||||
the SA token and SSH key items).
|
||||
2. **Operator retrieves the SA token**:
|
||||
```bash
|
||||
op read "op://yucca_tf_dev/yucca_futo_1pass_service_account/password"
|
||||
```
|
||||
3. **Operator sets it in their shell profile** (`~/.bashrc` or equivalent):
|
||||
```bash
|
||||
export OP_SERVICE_ACCOUNT_TOKEN="ops_eyJ...." # 862-char token
|
||||
```
|
||||
4. **Operator clones yucca**, installs mise tooling, runs:
|
||||
```bash
|
||||
cd ~/yucca/ansible/ceph
|
||||
mise trust && mise run setup
|
||||
```
|
||||
5. **Operator installs the ansible-iac SSH keys** from 1P:
|
||||
```bash
|
||||
scripts/install-ssh-keys.sh
|
||||
```
|
||||
Lands `~/.ssh/id_ed25519_sietch` and `~/.ssh/id_ed25519_painbox` (0600).
|
||||
6. **Verify**:
|
||||
```bash
|
||||
mise run preflight # should pass; no desktop 1P required
|
||||
```
|
||||
|
||||
## Per-session workflow
|
||||
|
||||
```bash
|
||||
cd ~/yucca/ansible/ceph
|
||||
export CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory.ini
|
||||
mise run status # read-only smoke test
|
||||
mise run deploy # or any other task
|
||||
```
|
||||
|
||||
No auth prompts — the SA token in env gets picked up by `op inject` inside
|
||||
`scripts/ansible-play.sh`.
|
||||
|
||||
## What remote hands CAN'T do
|
||||
|
||||
The read-only SA token can't:
|
||||
- Modify 1P items (e.g., rotate secrets).
|
||||
- Run `tf:apply` (needs superuser SA for writes).
|
||||
|
||||
Actions requiring writes must route through the primary operator or
|
||||
come with explicit superuser-token provisioning.
|
||||
|
||||
## Rotation
|
||||
|
||||
If a remote-hands operator leaves the project, rotate the read-only SA
|
||||
token (see `rotate-sa-token.md`). Their environment still holds the old
|
||||
token but it's invalidated in 1P; the next command fails loudly.
|
||||
|
||||
## References
|
||||
|
||||
- `tf/README.md` §"Where secrets actually live" — SA scopes
|
||||
- `docs/secrets.md` — full secrets model
|
||||
@@ -0,0 +1,246 @@
|
||||
# Runbook: Replace a Failed HDD
|
||||
|
||||
**When:** SMART failure, unresponsive OSD, or predictive disk replacement.
|
||||
|
||||
**Time estimate:** 15-30 minutes (plus backfill time, which varies by data volume).
|
||||
|
||||
**Prerequisites:**
|
||||
- Physical or remote hands access for the disk swap
|
||||
- Cluster has enough free capacity to absorb the missing OSD during backfill
|
||||
|
||||
---
|
||||
|
||||
## 1. Identify the failed disk
|
||||
|
||||
### Check cluster health
|
||||
|
||||
```bash
|
||||
ceph health detail
|
||||
```
|
||||
|
||||
Look for messages like:
|
||||
- `OSD_DOWN` -- `1 osds down`
|
||||
- `DEVICE_HEALTH_TOOMANY` -- `1 device(s) expected to fail soon`
|
||||
- `PG_DEGRADED` -- degraded placement groups
|
||||
|
||||
### Find the OSD ID and host
|
||||
|
||||
```bash
|
||||
ceph osd tree
|
||||
```
|
||||
|
||||
Note the OSD ID (e.g., `osd.7`) and the host it belongs to.
|
||||
|
||||
### Check SMART data
|
||||
|
||||
SSH to the host and find the physical device:
|
||||
|
||||
```bash
|
||||
# Find the device backing the OSD
|
||||
ceph-volume lvm list | grep -A5 "osd.7"
|
||||
|
||||
# Or find by path
|
||||
ceph device ls | grep osd.7
|
||||
|
||||
# Check SMART
|
||||
smartctl -a /dev/sdX
|
||||
```
|
||||
|
||||
Look for: `Reallocated_Sector_Ct`, `Current_Pending_Sector`,
|
||||
`SMART overall-health self-assessment test result: FAILED`.
|
||||
|
||||
## 2. Mark OSD out (start draining)
|
||||
|
||||
```bash
|
||||
ceph osd out osd.7
|
||||
```
|
||||
|
||||
This begins backfilling data away from the OSD. Monitor progress:
|
||||
|
||||
```bash
|
||||
ceph -s
|
||||
# or watch:
|
||||
ceph -w
|
||||
```
|
||||
|
||||
Wait until backfill completes and all PGs are `active+clean`:
|
||||
|
||||
```bash
|
||||
ceph pg stat
|
||||
```
|
||||
|
||||
Expected output includes `active+clean` for all PGs, zero `remapped` or
|
||||
`backfilling`.
|
||||
|
||||
**Do not proceed until backfill is complete.** Pulling a disk during
|
||||
backfill risks data loss if another disk fails simultaneously.
|
||||
|
||||
## 3. Stop and purge the OSD
|
||||
|
||||
```bash
|
||||
# Stop the OSD daemon
|
||||
ceph orch daemon stop osd.7
|
||||
|
||||
# Purge OSD from cluster (removes from CRUSH, auth keys, etc.)
|
||||
ceph osd purge osd.7 --yes-i-really-mean-it
|
||||
```
|
||||
|
||||
Verify removal:
|
||||
|
||||
```bash
|
||||
ceph osd tree
|
||||
```
|
||||
|
||||
The OSD should no longer appear.
|
||||
|
||||
## 4. Close LUKS and clean up on the host
|
||||
|
||||
SSH to the host where the OSD lived:
|
||||
|
||||
```bash
|
||||
ssh ansible-iac@sietch-ceph-<host>
|
||||
sudo -i
|
||||
```
|
||||
|
||||
### Close the LUKS mapping
|
||||
|
||||
```bash
|
||||
# List dm-crypt mappings to find the right one
|
||||
dmsetup ls --target crypt
|
||||
|
||||
# Close it (name will be something like ceph-<uuid>-...-block-dmcrypt)
|
||||
cryptsetup close <mapping-name>
|
||||
```
|
||||
|
||||
### Remove LVM artifacts
|
||||
|
||||
```bash
|
||||
# Find and remove orphaned PV on the disk
|
||||
pvs | grep /dev/sdX
|
||||
pvremove -f /dev/sdX
|
||||
```
|
||||
|
||||
### Wipe device signatures
|
||||
|
||||
```bash
|
||||
wipefs -af /dev/sdX
|
||||
dd if=/dev/zero of=/dev/sdX bs=1M count=10
|
||||
```
|
||||
|
||||
## 5. Physical disk swap
|
||||
|
||||
### Austin (on-site)
|
||||
|
||||
1. Identify the bay number from the SAS PHY mapping in host_vars
|
||||
2. The disk bay maps to a by-path device: `/dev/disk/by-path/<sas_path_prefix>-phy<N>-lun-0`
|
||||
3. Turn on the drive bay LED if available via iDRAC / `ledctl`
|
||||
4. Coordinate with the remote hands team for the physical swap
|
||||
5. Hot-swap the drive -- these are SAS/SATA hot-plug bays
|
||||
6. Verify the new disk appears:
|
||||
|
||||
```bash
|
||||
lsblk
|
||||
ls /dev/disk/by-path/ | grep phy<N>
|
||||
```
|
||||
|
||||
### Hetzner (remote)
|
||||
|
||||
Requires a support ticket to Hetzner for disk replacement. Schedule
|
||||
maintenance window.
|
||||
|
||||
## 6. Re-create the OSD
|
||||
|
||||
The new disk must get an encrypted OSD with a block.db LV on the SSD,
|
||||
matching the original configuration.
|
||||
|
||||
### Verify the block.db LV still exists
|
||||
|
||||
```bash
|
||||
lvs | grep db-slot<N>
|
||||
```
|
||||
|
||||
If the LV was destroyed, re-run LVM setup first:
|
||||
|
||||
```bash
|
||||
scripts/ansible-play.sh deploy-ceph.yml \
|
||||
--tags lvm --limit sietch-ceph-<host>,sietch-ceph-laurel
|
||||
```
|
||||
|
||||
### Create the OSD manually
|
||||
|
||||
```bash
|
||||
# Ensure bootstrap-osd keyring is present
|
||||
ceph auth get client.bootstrap-osd -o /var/lib/ceph/bootstrap-osd/ceph.keyring
|
||||
|
||||
# Create encrypted OSD with block.db
|
||||
DISK="/dev/disk/by-path/<sas_path_prefix>-phy<N>-lun-0"
|
||||
DB_LV="<vg-name>/db-slot<N>"
|
||||
|
||||
cephadm ceph-volume \
|
||||
--keyring /var/lib/ceph/bootstrap-osd/ceph.keyring \
|
||||
lvm create --dmcrypt --no-systemd \
|
||||
--data "$DISK" --block.db "$DB_LV"
|
||||
```
|
||||
|
||||
### Activate the OSD
|
||||
|
||||
From the bootstrap node:
|
||||
|
||||
```bash
|
||||
ceph cephadm osd activate <hostname>
|
||||
```
|
||||
|
||||
### Or re-run via Ansible
|
||||
|
||||
Alternatively, re-run the OSD creation phase through Ansible (idempotent --
|
||||
skips existing OSDs):
|
||||
|
||||
```bash
|
||||
scripts/ansible-play.sh deploy-ceph.yml \
|
||||
--tags osds --limit sietch-ceph-<host>,sietch-ceph-laurel
|
||||
```
|
||||
|
||||
## 7. Verify
|
||||
|
||||
### Check the new OSD is up
|
||||
|
||||
```bash
|
||||
ceph osd tree
|
||||
```
|
||||
|
||||
Expected: new OSD appears under the correct host with status `up` and
|
||||
weight > 0.
|
||||
|
||||
### Check reweight
|
||||
|
||||
```bash
|
||||
# If reweight is 0, fix it
|
||||
ceph osd tree | grep "osd.<new-id>"
|
||||
|
||||
# If needed:
|
||||
ceph osd reweight <new-id> 1.0
|
||||
```
|
||||
|
||||
### Monitor backfill to the new OSD
|
||||
|
||||
```bash
|
||||
ceph -s
|
||||
```
|
||||
|
||||
Wait for `active+clean` on all PGs.
|
||||
|
||||
### Verify dmcrypt
|
||||
|
||||
```bash
|
||||
ceph config-key dump | grep dm-crypt | wc -l
|
||||
```
|
||||
|
||||
Should show one more key than before the replacement.
|
||||
|
||||
### Verify SMART on new disk
|
||||
|
||||
```bash
|
||||
smartctl -a /dev/sdX
|
||||
```
|
||||
|
||||
Confirm the new disk has zero errors.
|
||||
@@ -0,0 +1,107 @@
|
||||
# Runbook: Replace a Host (Hardware Swap, Preserving Name)
|
||||
|
||||
**When:** chassis failure, motherboard swap, or preventative hardware
|
||||
refresh. You want the new box to assume the old host's identity — same
|
||||
hostname, same IP, same Ceph OSD identities.
|
||||
|
||||
**Time estimate:** 1-2 hours (bare-metal) / 30 min (Hetzner).
|
||||
|
||||
---
|
||||
|
||||
## Preserving name = preserving trust
|
||||
|
||||
Keeping the old hostname avoids:
|
||||
- Re-issuing the ansible-iac SSH key to a new hostname
|
||||
- Rotating the host's Ceph auth keys
|
||||
- Updating monitoring dashboards / alert rules with new labels
|
||||
- Updating any external systems that refer to the hostname
|
||||
|
||||
The TF inventory declaration stays unchanged — cluster name, host name,
|
||||
bond IP all match. Only the physical box changes.
|
||||
|
||||
## Steps
|
||||
|
||||
### 1. Mark OSDs out, wait for backfill
|
||||
|
||||
```bash
|
||||
# SSH to bootstrap node
|
||||
sudo ceph osd out $(sudo ceph osd ls-tree <hostname-being-replaced>)
|
||||
sudo ceph -w # watch for HEALTH_OK or acceptable degraded state
|
||||
```
|
||||
|
||||
Depending on cluster size, full backfill can take hours. For time-critical
|
||||
replacements, skip this step and accept degraded state during rebuild —
|
||||
but confirm you have headroom.
|
||||
|
||||
### 2. Power down old host, rack new one in same slot
|
||||
|
||||
Preserve the physical network cabling and IPMI address. The new box gets
|
||||
the same bond IP via the same DHCP reservation / static config.
|
||||
|
||||
### 3. Re-run provisioning against the new host
|
||||
|
||||
**Bare-metal (sietch)**:
|
||||
|
||||
```bash
|
||||
# Boot the new box from the Debian 12 live image (same procedure as first
|
||||
# provision — see docs/runbooks/add-node.md for iDRAC steps).
|
||||
|
||||
CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory-provision.ini \
|
||||
scripts/ansible-play.sh provision.yml \
|
||||
-e confirm_wipe=true \
|
||||
--limit sietch-ceph-<name>
|
||||
```
|
||||
|
||||
**Hetzner (painbox)**: reboot into rescue, run installimage + post-install
|
||||
scripts from `inventories/painbox-ceph.dev.hel.htz/installimage/`.
|
||||
|
||||
### 4. Clear the old host's Ceph state
|
||||
|
||||
The cluster still has the old host's OSDs, CRUSH entries, and cephadm
|
||||
host record. Clean those up BEFORE the new box rejoins:
|
||||
|
||||
```bash
|
||||
# On bootstrap node:
|
||||
for osd in $(sudo ceph osd ls-tree <hostname>); do
|
||||
sudo ceph osd purge $osd --yes-i-really-mean-it
|
||||
done
|
||||
sudo ceph osd crush remove <hostname>
|
||||
sudo ceph orch host rm <hostname> --force
|
||||
```
|
||||
|
||||
### 5. Re-run baseline + deploy for that host
|
||||
|
||||
```bash
|
||||
scripts/ansible-play.sh baseline.yml --limit <hostname>
|
||||
scripts/ansible-play.sh deploy-ceph.yml \
|
||||
--limit <hostname>,<bootstrap-hostname>
|
||||
```
|
||||
|
||||
Note bootstrap host must be in `--limit` — join/placement/OSD activation
|
||||
all run from bootstrap.
|
||||
|
||||
### 6. Verify
|
||||
|
||||
```bash
|
||||
# All OSDs for the replaced host are up+in
|
||||
sudo ceph osd tree | grep <hostname>
|
||||
|
||||
# Cluster is rebalancing or healthy
|
||||
sudo ceph -s
|
||||
```
|
||||
|
||||
## Gotchas
|
||||
|
||||
- **Hardware topology changed**: new chassis may have different PCI paths
|
||||
for SAS controllers / NVMe. Update `host_vars/<hostname>.yml`
|
||||
(`sas_path_prefix`, `ceph_hdd_osds`) before step 5.
|
||||
- **Different serial numbers**: acceptable — `by-path` is used for OSD
|
||||
slot identity, not `by-id`.
|
||||
- **Backfill thundering herd**: if you skipped step 1, the new OSDs enter
|
||||
the cluster and backfill aggressively. Throttle with
|
||||
`ceph_osd_recovery_max_active` in vars.yml if you see client-IO impact.
|
||||
|
||||
## References
|
||||
|
||||
- `docs/runbooks/add-node.md` — related but for net-new hosts
|
||||
- `docs/runbooks/replace-disk.md` — single-disk replacement within a host
|
||||
@@ -0,0 +1,121 @@
|
||||
# Runbook: Rotate RGW TLS Certificates
|
||||
|
||||
**When:** Certificate approaching expiry, compromised key material, or SAN
|
||||
changes (new nodes added, DNS name changed).
|
||||
|
||||
**Time estimate:** 5 minutes. Brief RGW restart causes ~10s S3 downtime.
|
||||
|
||||
**Prerequisites:**
|
||||
- `op` session live (desktop unlocked or `OP_SERVICE_ACCOUNT_TOKEN` set)
|
||||
- Cluster is healthy
|
||||
|
||||
---
|
||||
|
||||
## 1. Check current certificate expiry
|
||||
|
||||
```bash
|
||||
ssh ansible-iac@sietch-ceph-laurel
|
||||
sudo openssl x509 -in /etc/ceph/rgw-ssl.crt -noout -subject -dates -ext subjectAltName
|
||||
```
|
||||
|
||||
Output shows:
|
||||
|
||||
```
|
||||
subject=C = US, ST = Texas, L = Austin, O = FUTO, CN = s3.dev.austin.int.futo.cloud
|
||||
notBefore=...
|
||||
notAfter=...
|
||||
X509v3 Subject Alternative Name:
|
||||
DNS:s3.dev.austin.int.futo.cloud, DNS:*.s3.dev.austin.int.futo.cloud, ...
|
||||
```
|
||||
|
||||
## 2. Run the rotation playbook
|
||||
|
||||
```bash
|
||||
scripts/ansible-play.sh rotate-certs.yml
|
||||
```
|
||||
|
||||
### What this does
|
||||
|
||||
1. **Deletes** `/etc/ceph/rgw-ssl.crt` and `/etc/ceph/rgw-ssl.key` on the
|
||||
bootstrap node
|
||||
2. **Re-runs the RGW role** (`roles/ceph_deploy/tasks/rgw.yml`) which:
|
||||
- Generates a new 4096-bit RSA self-signed cert with 10-year validity
|
||||
- SANs include: the canonical DNS name, wildcard for virtual-hosted
|
||||
buckets, per-node FQDNs, and per-node bond IPs
|
||||
- Renders the RGW service spec with the new cert embedded
|
||||
- Applies the service spec via `ceph orch apply` -- cephadm distributes
|
||||
the cert to all RGW daemon containers
|
||||
3. **Restarts all RGW daemons** via `ceph orch restart rgw` to load the
|
||||
new cert
|
||||
4. **Displays** the new certificate subject, dates, and SANs
|
||||
|
||||
### Impact
|
||||
|
||||
- RGW daemons restart sequentially. S3 requests will fail for ~10 seconds
|
||||
during the restart window.
|
||||
- Clients using the old self-signed cert in their trust store will need the
|
||||
new cert. Export with:
|
||||
|
||||
```bash
|
||||
ssh ansible-iac@sietch-ceph-laurel sudo cat /etc/ceph/rgw-ssl.crt
|
||||
```
|
||||
|
||||
## 3. Verify after rotation
|
||||
|
||||
### Check new cert details
|
||||
|
||||
The playbook prints this, but to verify manually:
|
||||
|
||||
```bash
|
||||
ssh ansible-iac@sietch-ceph-laurel
|
||||
sudo openssl x509 -in /etc/ceph/rgw-ssl.crt -noout -subject -dates -ext subjectAltName
|
||||
```
|
||||
|
||||
### Test RGW endpoint
|
||||
|
||||
```bash
|
||||
# From a node in the cluster (self-signed cert)
|
||||
curl -k https://s3.dev.austin.int.futo.cloud:443/
|
||||
```
|
||||
|
||||
Expected: XML response with `ListAllMyBucketsResult` or `AccessDenied`
|
||||
(both mean RGW is serving TLS correctly).
|
||||
|
||||
### Test direct node access
|
||||
|
||||
```bash
|
||||
curl -k https://10.10.10.90:443/
|
||||
```
|
||||
|
||||
### Check RGW daemons are running
|
||||
|
||||
```bash
|
||||
ceph orch ls --service-type rgw
|
||||
```
|
||||
|
||||
Expected: running count matches the number of ceph_nodes (currently 3).
|
||||
|
||||
### Check dashboard can reach RGW
|
||||
|
||||
Open `https://<bootstrap-ip>:8443` and navigate to Object Gateway.
|
||||
The page should load without 500 errors (dashboard has RGW API SSL
|
||||
verification disabled for self-signed certs).
|
||||
|
||||
## Certificate configuration
|
||||
|
||||
The cert parameters are controlled by these variables in
|
||||
`inventories/sietch-ceph.dev.austin.int/group_vars/all/vars.yml`:
|
||||
|
||||
| Variable | Default | Purpose |
|
||||
|---|---|---|
|
||||
| `ceph_rgw_ssl` | `true` | Enable TLS on RGW frontend |
|
||||
| `ceph_rgw_ssl_cert_days` | `3650` | Validity period (10 years) |
|
||||
| `ceph_rgw_ssl_cert_subject_c` | `US` | Country |
|
||||
| `ceph_rgw_ssl_cert_subject_st` | `Texas` | State |
|
||||
| `ceph_rgw_ssl_cert_subject_l` | `Austin` | Locality |
|
||||
| `ceph_rgw_ssl_cert_subject_o` | `FUTO` | Organization |
|
||||
| `ceph_rgw_ssl_cert_email` | `yucca@futo.org` | Contact email |
|
||||
| `ceph_rgw_dns_name` | `s3.dev.austin.int.futo.cloud` | CN and primary SAN |
|
||||
|
||||
SANs are auto-generated from inventory: per-node FQDNs and bond IPs are
|
||||
included so direct-host access validates.
|
||||
@@ -0,0 +1,70 @@
|
||||
# Runbook: Rotate 1Password Service Account Token
|
||||
|
||||
**When:** token leak, team personnel change, or scheduled rotation.
|
||||
|
||||
**Time estimate:** 15-30 minutes including cross-team coordination.
|
||||
|
||||
**Blast radius:** **cross-project**. The SA tokens in `yucca_tf_dev` are
|
||||
shared with o11y and other Futo consumers. Rotating breaks every consumer
|
||||
until they pick up the new token.
|
||||
|
||||
---
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Coordination with other consumers of the SA (at minimum: o11y team,
|
||||
whoever else uses `yucca_tf_*`). Ask in Discord before rotating.
|
||||
- 1Password access to create a replacement SA token.
|
||||
- List of repos/CI pipelines that consume the token — so you know who to
|
||||
notify to pick up the rotation.
|
||||
|
||||
## Which SA to rotate
|
||||
|
||||
Two SAs live in `yucca_tf_dev`:
|
||||
|
||||
| SA | Scope |
|
||||
|---|---|
|
||||
| `yucca_futo_1pass_superuser_service_account` | read+write all yucca_tf_* vaults |
|
||||
| `yucca_futo_1pass_service_account` | read-only on yucca_tf + yucca_tf_dev |
|
||||
|
||||
Sietch-ceph consumes both (TF uses superuser for writes, Ansible uses
|
||||
read-only for runtime). Read-only rotation has lower blast radius.
|
||||
|
||||
## Steps
|
||||
|
||||
1. **Announce in Discord** to Immich maintainers + o11y team: "Rotating
|
||||
`yucca_futo_1pass_<type>_service_account` at HH:MM. Expect a brief
|
||||
window where CI pipelines fail — update your vars after."
|
||||
2. **Create a new SA token in 1Password Admin UI** with matching scope.
|
||||
Give it a dated suffix (e.g. `-2026-04-22`) so old and new coexist
|
||||
briefly.
|
||||
3. **Update the SA's `password` field** in the existing 1P item (so
|
||||
consumers reading via `op://yucca_tf_dev/<name>/password` pick up the
|
||||
new value with no code change).
|
||||
4. **Test from this repo**:
|
||||
```bash
|
||||
mise run tf:plan # should succeed with new superuser token
|
||||
scripts/ansible-play.sh status.yml # should succeed with new read-only
|
||||
```
|
||||
5. **Notify consumers** the rotation is done — they restart any daemons /
|
||||
re-pull tokens as needed.
|
||||
6. **Revoke the old SA token** in 1Password Admin UI once all consumers
|
||||
confirm green.
|
||||
|
||||
## Recovery: rotation broke something
|
||||
|
||||
- If the new token doesn't work: check the new SA has the same vault
|
||||
scopes as the old. Empty `op vault list` output under the new token =
|
||||
scope misconfigured.
|
||||
- If a consumer is still using the old token: the item's `password` field
|
||||
is the single source of truth; consumers re-reading pick up the new
|
||||
value. If a consumer caches the token in its own env (CI secret,
|
||||
systemd EnvironmentFile), update that manually.
|
||||
- Roll back: edit the 1P item's `password` field back to the old token
|
||||
value (if you kept a copy). The old SA is still valid until
|
||||
explicitly revoked.
|
||||
|
||||
## References
|
||||
|
||||
- `tf/.env` consumes `op://yucca_tf_dev/yucca_futo_1pass_superuser_service_account/password`
|
||||
- `tf/README.md` §"Where secrets actually live" — describes the SA scopes
|
||||
@@ -0,0 +1,201 @@
|
||||
# Runbook: Rotate Passwords and Keys
|
||||
|
||||
**When:** scheduled rotation, suspected compromise, or personnel change.
|
||||
|
||||
**Time estimate:** 5 minutes per password (plus any service-restart windows).
|
||||
|
||||
**Prerequisites:**
|
||||
- `op` CLI authenticated (desktop session unlocked or `OP_SERVICE_ACCOUNT_TOKEN` set)
|
||||
- Inventory and secrets template rendered (`tofu apply` has been run at least once)
|
||||
|
||||
---
|
||||
|
||||
## Model
|
||||
|
||||
1Password is the sole source of truth. Rotation = edit the item's `password`
|
||||
field, done — no re-encrypt, no commit, no sync step. The next `mise run
|
||||
deploy` picks up the new value via `op inject`.
|
||||
|
||||
## Secrets managed by this cluster
|
||||
|
||||
| Secret | 1Password item | Vault | Where it's applied |
|
||||
|---------------------------|----------------------------------------|----------------|--------------------------------------------|
|
||||
| `ops` user password | `<CLUSTER>_CEPH_OPS_PASSWORD` | `yucca_tf_dev` | Linux login on all nodes |
|
||||
| Ceph dashboard password | `<CLUSTER>_CEPH_DASHBOARD_PASSWORD` | `yucca_tf_dev` | `https://<node>:8443` admin login |
|
||||
| Grafana admin password | `<CLUSTER>_CEPH_GRAFANA_PASSWORD` | `yucca_tf_dev` | `https://<node>:3000` admin login |
|
||||
| S3 service-user access | `<CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY` | `yucca_tf_dev` | RGW S3 user `yucca-restic` access key |
|
||||
| S3 service-user secret | `<CLUSTER>_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY` | `yucca_tf_dev` | RGW S3 user `yucca-restic` secret key |
|
||||
|
||||
Replace `<CLUSTER>` with `SIETCH` or `PAINBOX`. The active vault name is
|
||||
declared per-cluster in the `vault` field of the cluster's entry in
|
||||
`tf/deployment/dev/ceph/clusters.auto.tfvars` — `yucca_tf_dev` for dev,
|
||||
future `yucca_tf_staging` / `yucca_tf` for staging/prod.
|
||||
|
||||
## 1. Rotate in 1Password
|
||||
|
||||
Either edit via the desktop app, or from CLI:
|
||||
|
||||
```bash
|
||||
# Generate + set a new password in one shot
|
||||
op item edit SIETCH_CEPH_DASHBOARD_PASSWORD --vault yucca_tf_dev --generate-password='letters,digits,32'
|
||||
```
|
||||
|
||||
Or set an explicit value:
|
||||
|
||||
```bash
|
||||
op item edit SIETCH_CEPH_DASHBOARD_PASSWORD --vault yucca_tf_dev password=<literal-new-value>
|
||||
```
|
||||
|
||||
Verify the new value resolves:
|
||||
|
||||
```bash
|
||||
op read "op://yucca_tf_dev/SIETCH_CEPH_DASHBOARD_PASSWORD/password" | wc -c
|
||||
```
|
||||
|
||||
## 2. Apply to the cluster
|
||||
|
||||
### ops user password
|
||||
|
||||
The baseline role converges the ops user password on every run:
|
||||
|
||||
```bash
|
||||
scripts/ansible-play.sh baseline.yml --tags users
|
||||
```
|
||||
|
||||
The old password stops working immediately.
|
||||
|
||||
### Dashboard password
|
||||
|
||||
Not re-applied by normal deploys. Set directly on the bootstrap MON:
|
||||
|
||||
```bash
|
||||
ssh ansible-iac@sietch-ceph-laurel \
|
||||
sudo ceph dashboard ac-user-set-password admin \
|
||||
"$(op read 'op://yucca_tf_dev/SIETCH_CEPH_DASHBOARD_PASSWORD/password')"
|
||||
```
|
||||
|
||||
Expect `User admin updated`.
|
||||
|
||||
### Grafana admin password
|
||||
|
||||
Re-run the monitoring tag:
|
||||
|
||||
```bash
|
||||
scripts/ansible-play.sh deploy-ceph.yml --tags monitoring
|
||||
```
|
||||
|
||||
## 3. Verify
|
||||
|
||||
| Check | Command / action |
|
||||
|---|---|
|
||||
| ops login | `ssh ops@sietch-ceph-laurel` — new password prompts and works |
|
||||
| Dashboard | browse `https://<bootstrap-ip>:8443`, log in as `admin` with new password |
|
||||
| Grafana | browse `https://<bootstrap-ip>:3000`, log in as `admin` with new password |
|
||||
|
||||
## 4. Clean up
|
||||
|
||||
No cleanup needed on the secrets side — old values are gone from 1Password
|
||||
the moment you save the new ones. No vault.yml re-encryption, no git commit
|
||||
required for rotation.
|
||||
|
||||
---
|
||||
|
||||
## Rotating SSH keys
|
||||
|
||||
SSH keys for `ansible-iac` live in `yucca_tf_dev` as category `SSH Key`
|
||||
items titled `<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY`. Rotation is
|
||||
forward-only — generate a new key in 1P, distribute the new pubkey,
|
||||
retire the old key after confidence.
|
||||
|
||||
### Steps (sietch example; same pattern for any cluster)
|
||||
|
||||
1. **Generate the new keypair natively in 1P**. The current item must be
|
||||
replaced (1P doesn't support multiple key versions per item).
|
||||
Snapshot the old item first for rollback:
|
||||
|
||||
```bash
|
||||
# Get the new superuser SA token
|
||||
SU_TOKEN=$(op read "op://yucca_tf_dev/yucca_futo_1pass_superuser_service_account/password")
|
||||
|
||||
# Save the old item as a dated rollback copy (private_key is derivable
|
||||
# via op item get; use --reveal if you need it on disk)
|
||||
OP_SERVICE_ACCOUNT_TOKEN="$SU_TOKEN" op item edit SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY \
|
||||
--vault yucca_tf_dev \
|
||||
--title "SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY_RETIRED_$(date +%Y%m%d)"
|
||||
|
||||
# Create the replacement with the canonical title
|
||||
OP_SERVICE_ACCOUNT_TOKEN="$SU_TOKEN" op item create \
|
||||
--vault yucca_tf_dev --category "SSH Key" \
|
||||
--title SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY \
|
||||
--ssh-generate-key=ed25519
|
||||
|
||||
unset SU_TOKEN
|
||||
```
|
||||
|
||||
2. **Install the new private key on your workstation**
|
||||
(`scripts/install-ssh-keys.sh` refuses to overwrite mismatched
|
||||
fingerprints — move the old file aside first):
|
||||
|
||||
```bash
|
||||
mv ~/.ssh/id_ed25519_sietch ~/.ssh/id_ed25519_sietch.retired-$(date +%Y%m%d)
|
||||
mv ~/.ssh/id_ed25519_sietch.pub ~/.ssh/id_ed25519_sietch.pub.retired-$(date +%Y%m%d)
|
||||
scripts/install-ssh-keys.sh sietch
|
||||
```
|
||||
|
||||
3. **Distribute the new pubkey to every node** (non-destructive;
|
||||
old key stays in authorized_keys until step 5):
|
||||
|
||||
```bash
|
||||
scripts/ansible-play.sh rotate-ssh-key.yml
|
||||
```
|
||||
|
||||
4. **Test** SSH with the new key:
|
||||
|
||||
```bash
|
||||
ssh -i ~/.ssh/id_ed25519_sietch ansible-iac@sietch-ceph-laurel hostname -f
|
||||
```
|
||||
|
||||
5. **Remove the old key from authorized_keys** once confident (manual
|
||||
— no playbook for this yet). On each node:
|
||||
|
||||
```bash
|
||||
ssh ansible-iac@sietch-ceph-laurel \
|
||||
"sed -i '/RETIRED-KEY-COMMENT/d' ~/.ssh/authorized_keys"
|
||||
```
|
||||
|
||||
Match by key comment (e.g., the email/hostname in the pubkey's
|
||||
trailing field).
|
||||
|
||||
6. **Delete the retired 1P item** (optional; keep ~30 days for
|
||||
audit/rollback):
|
||||
|
||||
```bash
|
||||
SU_TOKEN=$(op read "op://yucca_tf_dev/yucca_futo_1pass_superuser_service_account/password")
|
||||
OP_SERVICE_ACCOUNT_TOKEN="$SU_TOKEN" op item delete \
|
||||
--vault yucca_tf_dev "SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY_RETIRED_<date>"
|
||||
unset SU_TOKEN
|
||||
```
|
||||
|
||||
### Provisioning a fresh host uses the current key
|
||||
|
||||
`roles/provision_host/tasks/admin_user.yml` reads
|
||||
`{{ provision_iac_ssh_key_path }}.pub` (e.g., `~/.ssh/id_ed25519_sietch.pub`
|
||||
on operator disk) and installs it as the bootstrap `authorized_keys`
|
||||
during Debian live-image provisioning. Make sure `install-ssh-keys.sh`
|
||||
has run on any workstation that'll drive provisioning.
|
||||
|
||||
---
|
||||
|
||||
## Future: rotating TF-provisioned secrets
|
||||
|
||||
Once the sietch-ceph service account lands and `secrets.tf.disabled` is
|
||||
re-enabled, rotations become a `terraform taint` + `apply`:
|
||||
|
||||
```bash
|
||||
cd tf/deployment/dev/ceph
|
||||
terragrunt taint 'module.cluster["sietch"].onepassword_item.secret["dashboard"]'
|
||||
terragrunt apply
|
||||
```
|
||||
|
||||
This regenerates the password in 1Password; the apply step to the live
|
||||
cluster (dashboard / grafana commands above) is still required.
|
||||
@@ -0,0 +1,268 @@
|
||||
# S3 Integration Guide
|
||||
|
||||
Audience: Application developers (Yucca, Immich, Restic, internal tooling).
|
||||
|
||||
## Endpoints
|
||||
|
||||
The cluster runs Ceph RGW (RADOS Gateway) on every node behind a self-signed
|
||||
wildcard TLS certificate on **port 443**.
|
||||
|
||||
| Style | URL |
|
||||
|---|---|
|
||||
| Path-style | `https://s3.dev.austin.int.futo.cloud/<bucket>/<key>` |
|
||||
| Virtual-hosted | `https://<bucket>.s3.dev.austin.int.futo.cloud/<key>` |
|
||||
| Direct (per-node) | `https://10.10.10.90:443`, `https://10.10.10.91:443`, `https://10.10.10.92:443` |
|
||||
|
||||
Region: **us-east-1**
|
||||
|
||||
Path-style is recommended for simplicity. Virtual-hosted requires wildcard DNS
|
||||
(see DNS section below).
|
||||
|
||||
## Getting credentials
|
||||
|
||||
### Option A: 1Password (preferred)
|
||||
|
||||
S3 credentials for the `svc-yucca-restic` service account are stored in
|
||||
1Password after initial deployment. Ask the infrastructure team for access to
|
||||
the "Ceph S3" vault entry.
|
||||
|
||||
### Option B: radosgw-admin (infra operators only)
|
||||
|
||||
SSH to the bootstrap node (sietch-ceph-laurel) and run:
|
||||
|
||||
```bash
|
||||
radosgw-admin user info --uid=svc-yucca-restic
|
||||
```
|
||||
|
||||
The `keys[0].access_key` and `keys[0].secret_key` fields contain the
|
||||
credentials.
|
||||
|
||||
To create a new service account:
|
||||
|
||||
```bash
|
||||
radosgw-admin user create \
|
||||
--uid=svc-myapp \
|
||||
--display-name='myapp service account' \
|
||||
--max-buckets=100
|
||||
```
|
||||
|
||||
## Self-signed certificate handling
|
||||
|
||||
The cluster uses a self-signed wildcard certificate. Every client must either
|
||||
trust the CA or disable TLS verification.
|
||||
|
||||
### Trusting the cert (recommended for production workloads)
|
||||
|
||||
Copy the cert from the bootstrap node:
|
||||
|
||||
```bash
|
||||
scp ansible-iac@10.10.10.90:/etc/ceph/rgw-ssl.crt ./rgw-ssl.crt
|
||||
```
|
||||
|
||||
Then pass it to your client (examples below).
|
||||
|
||||
### Disabling verification (quick testing only)
|
||||
|
||||
Pass `--no-verify-ssl` (AWS CLI) or `verify=False` (boto3). Fine for
|
||||
benchmarking, not for production.
|
||||
|
||||
## AWS CLI configuration
|
||||
|
||||
### ~/.aws/credentials
|
||||
|
||||
```ini
|
||||
[sietch]
|
||||
aws_access_key_id = YOUR_ACCESS_KEY
|
||||
aws_secret_access_key = YOUR_SECRET_KEY
|
||||
```
|
||||
|
||||
### ~/.aws/config
|
||||
|
||||
```ini
|
||||
[profile sietch]
|
||||
region = us-east-1
|
||||
endpoint_url = https://s3.dev.austin.int.futo.cloud
|
||||
s3 =
|
||||
signature_version = s3v4
|
||||
addressing_style = path
|
||||
```
|
||||
|
||||
### Basic operations
|
||||
|
||||
```bash
|
||||
# List buckets
|
||||
aws --profile sietch --no-verify-ssl s3 ls
|
||||
|
||||
# Create a bucket
|
||||
aws --profile sietch --no-verify-ssl s3 mb s3://my-bucket
|
||||
|
||||
# Upload a file
|
||||
aws --profile sietch --no-verify-ssl s3 cp ./file.txt s3://my-bucket/file.txt
|
||||
|
||||
# List objects
|
||||
aws --profile sietch --no-verify-ssl s3 ls s3://my-bucket/
|
||||
|
||||
# Download
|
||||
aws --profile sietch --no-verify-ssl s3 cp s3://my-bucket/file.txt ./downloaded.txt
|
||||
|
||||
# Using the CA bundle instead of --no-verify-ssl
|
||||
aws --profile sietch --ca-bundle ./rgw-ssl.crt s3 ls
|
||||
```
|
||||
|
||||
## boto3 (Python)
|
||||
|
||||
```python
|
||||
import boto3
|
||||
import botocore
|
||||
import urllib3
|
||||
|
||||
# Suppress InsecureRequestWarning when verify=False
|
||||
urllib3.disable_warnings(urllib3.exceptions.InsecureRequestWarning)
|
||||
|
||||
s3 = boto3.client(
|
||||
"s3",
|
||||
endpoint_url="https://s3.dev.austin.int.futo.cloud",
|
||||
aws_access_key_id="YOUR_ACCESS_KEY",
|
||||
aws_secret_access_key="YOUR_SECRET_KEY",
|
||||
region_name="us-east-1",
|
||||
verify=False, # or path to rgw-ssl.crt
|
||||
config=botocore.config.Config(
|
||||
signature_version="s3v4",
|
||||
s3={"addressing_style": "path"},
|
||||
),
|
||||
)
|
||||
|
||||
# Create a bucket
|
||||
s3.create_bucket(Bucket="my-bucket")
|
||||
|
||||
# Upload
|
||||
s3.put_object(Bucket="my-bucket", Key="hello.txt", Body=b"hello world")
|
||||
|
||||
# Download
|
||||
obj = s3.get_object(Bucket="my-bucket", Key="hello.txt")
|
||||
data = obj["Body"].read()
|
||||
|
||||
# List objects
|
||||
response = s3.list_objects_v2(Bucket="my-bucket")
|
||||
for item in response.get("Contents", []):
|
||||
print(item["Key"], item["Size"])
|
||||
```
|
||||
|
||||
To use the CA bundle instead of disabling verification:
|
||||
|
||||
```python
|
||||
s3 = boto3.client(
|
||||
"s3",
|
||||
endpoint_url="https://s3.dev.austin.int.futo.cloud",
|
||||
aws_access_key_id="YOUR_ACCESS_KEY",
|
||||
aws_secret_access_key="YOUR_SECRET_KEY",
|
||||
region_name="us-east-1",
|
||||
verify="/path/to/rgw-ssl.crt",
|
||||
config=botocore.config.Config(
|
||||
signature_version="s3v4",
|
||||
s3={"addressing_style": "path"},
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
## Restic
|
||||
|
||||
```bash
|
||||
export AWS_ACCESS_KEY_ID="YOUR_ACCESS_KEY"
|
||||
export AWS_SECRET_ACCESS_KEY="YOUR_SECRET_KEY"
|
||||
export RESTIC_REPOSITORY="s3:https://s3.dev.austin.int.futo.cloud/restic-backups"
|
||||
|
||||
# Init (first time)
|
||||
restic init --option s3.region=us-east-1
|
||||
|
||||
# Backup
|
||||
restic backup /data --option s3.region=us-east-1
|
||||
```
|
||||
|
||||
Note: Restic uses the Go AWS SDK. For self-signed certs, set
|
||||
`AWS_CA_BUNDLE=/path/to/rgw-ssl.crt` or use the system trust store.
|
||||
|
||||
## Bucket creation
|
||||
|
||||
Buckets are created via any S3 client. The `svc-yucca-restic` service account
|
||||
has a limit of 100 buckets (configurable via `radosgw-admin user modify
|
||||
--max-buckets`).
|
||||
|
||||
```bash
|
||||
# AWS CLI
|
||||
aws --profile sietch --no-verify-ssl s3 mb s3://my-new-bucket
|
||||
|
||||
# boto3
|
||||
s3.create_bucket(Bucket="my-new-bucket")
|
||||
```
|
||||
|
||||
Bucket data lands in the EC data pool (`dev-z1.rgw.buckets.data`). Index
|
||||
metadata goes to a separate replicated pool. No pool-level configuration is
|
||||
needed from the application side.
|
||||
|
||||
## DNS setup for virtual-hosted buckets
|
||||
|
||||
Virtual-hosted bucket addressing (e.g., `https://my-bucket.s3.dev.austin.int.futo.cloud`)
|
||||
requires two DNS records:
|
||||
|
||||
```
|
||||
s3.dev.austin.int.futo.cloud. A 10.10.10.90
|
||||
s3.dev.austin.int.futo.cloud. A 10.10.10.91
|
||||
s3.dev.austin.int.futo.cloud. A 10.10.10.92
|
||||
*.s3.dev.austin.int.futo.cloud. A 10.10.10.90
|
||||
*.s3.dev.austin.int.futo.cloud. A 10.10.10.91
|
||||
*.s3.dev.austin.int.futo.cloud. A 10.10.10.92
|
||||
```
|
||||
|
||||
Round-robin A records across all three nodes. Without these records,
|
||||
use path-style addressing or connect directly via IP.
|
||||
|
||||
Note: this repo does not manage DNS. Coordinate with the network/DNS team.
|
||||
|
||||
## Performance characteristics
|
||||
|
||||
| Property | Value |
|
||||
|---|---|
|
||||
| Data pool | Erasure coded, k=8 m=3, failure domain=OSD |
|
||||
| Index pool | Replicated, size=2, min_size=1 |
|
||||
| Storage media | HDD-backed (HGST 6 TB SAS drives) |
|
||||
| Block.db | SSD-backed (Micron 5100 3.8 TB) |
|
||||
| Nodes | 3 (Dell R730xd) |
|
||||
| RGW daemons | 1 per node |
|
||||
| TLS | Self-signed wildcard, 10-year validity |
|
||||
| Network | 10 GbE bonded active-backup (no LACP) |
|
||||
|
||||
### What to expect
|
||||
|
||||
- **Throughput**: HDD-bound for large objects. A single node can sustain
|
||||
roughly 500-800 MiB/s aggregate reads from its HDDs. With 3 nodes and
|
||||
EC 8+3, expect 300-600 MiB/s aggregate for large sequential workloads
|
||||
depending on concurrency and object size.
|
||||
- **Latency**: Higher than SSD or cloud S3. Small object PUTs (< 1 MiB) will
|
||||
see 10-50 ms latency due to HDD seeks. Use concurrency to amortize.
|
||||
- **IOPS**: Low single-drive IOPS (100-200 per HDD). Use larger objects
|
||||
(16+ MiB) to maximize throughput.
|
||||
- **EC overhead**: Usable capacity is raw * k/(k+m) = raw * 8/11 = ~72.7%
|
||||
of raw HDD capacity.
|
||||
- **Single network**: Public and cluster traffic share the same 10 GbE bond.
|
||||
Recovery/rebalancing events will compete with client I/O.
|
||||
|
||||
### Benchmarking
|
||||
|
||||
The cluster includes a benchmark tool at `roles/s3_bench/files/s3bench.py`:
|
||||
|
||||
```bash
|
||||
python3 s3bench.py \
|
||||
--endpoint https://10.10.10.90:443 \
|
||||
--access-key KEY \
|
||||
--secret-key SECRET \
|
||||
--bucket s3bench \
|
||||
--num-objects 100 \
|
||||
--object-size-mb 16 \
|
||||
--concurrency 8 \
|
||||
--ops put \
|
||||
--log /tmp/bench.jsonl
|
||||
```
|
||||
|
||||
Supports `put`, `get`, `delete`, and `mixed` (70/20/10 split) operations.
|
||||
Results include throughput (MiB/s), IOPS, and p50/p95/p99 latencies.
|
||||
@@ -0,0 +1,292 @@
|
||||
# Wrapper Scripts
|
||||
|
||||
The scripts under `ansible/ceph/scripts/` sit between `mise` tasks and the
|
||||
underlying CLIs (`ansible-playbook`, `op`, `ssh`). They exist to:
|
||||
|
||||
- Keep secrets **out of argv** — `op inject` writes to a file; the file is
|
||||
passed as `--extra-vars @<path>`, never expanded inline.
|
||||
- **Fail closed** — if 1Password is unreachable or the inventory is missing,
|
||||
the wrappers exit before invoking ansible. The raw tools' error messages
|
||||
degrade silently in these cases.
|
||||
- Offer **better error messages** — "inventory not found at `<path>` — has
|
||||
`tofu apply` been run?" beats "No inventory was parsed".
|
||||
|
||||
For the architectural role these scripts play, see
|
||||
[architecture.md §8](architecture.md).
|
||||
|
||||
## Setting `CEPH_ENV`
|
||||
|
||||
Two of the three wrappers (`ansible-play.sh`, `preflight.sh`) require
|
||||
`CEPH_ENV` to point at the target cluster's `inventory.ini`. **Always set
|
||||
it inline, never via `export`:**
|
||||
|
||||
```bash
|
||||
# Correct — inline prefix, applies to one mise/script invocation
|
||||
CEPH_ENV=inventories/painbox-ceph.dev.hel.htz/inventory.ini mise run preflight
|
||||
|
||||
# WRONG — mise's [env] machinery silently strips shell-exported vars
|
||||
# when launching tasks; CEPH_ENV reaches an empty environment and the
|
||||
# wrapper exits with "CEPH_ENV must be set". Confusing because your shell
|
||||
# clearly has it set.
|
||||
export CEPH_ENV=inventories/painbox-ceph.dev.hel.htz/inventory.ini
|
||||
mise run preflight # fails
|
||||
```
|
||||
|
||||
The `ansible/ceph/.mise.toml` `[env]` block intentionally does NOT declare
|
||||
a `CEPH_ENV` default — silent default-cluster behavior is more dangerous
|
||||
than requiring an explicit choice. If you call the scripts directly
|
||||
(bypassing `mise run`), shell `export` works normally because mise isn't
|
||||
in the path.
|
||||
|
||||
## Quick reference
|
||||
|
||||
| Script | What it does | Called by |
|
||||
|------------------------|------------------------------------------------------------------------------------|---------------------------------------------------|
|
||||
| `ansible-play.sh` | Render `secrets.yml.tpl` via `op inject` to a tmpfile, then exec `ansible-playbook` | Every `mise run` task that runs a playbook |
|
||||
| `install-ssh-keys.sh` | Pull per-cluster ansible-iac SSH keys from 1P into `~/.ssh/` | Operator (once per workstation / after rotation) |
|
||||
| `preflight.sh` | Verify TF artifacts, 1P session, SSH reachability, Python on targets | `mise run preflight` |
|
||||
|
||||
---
|
||||
|
||||
## `ansible-play.sh`
|
||||
|
||||
Wrapper around `ansible-playbook` that resolves TF-rendered secrets via
|
||||
`op inject` and passes them as an ephemeral extra-vars file.
|
||||
|
||||
### Synopsis
|
||||
|
||||
```
|
||||
CEPH_ENV=inventories/<cluster>/inventory.ini \
|
||||
scripts/ansible-play.sh <playbook.yml> [ansible-playbook args...]
|
||||
```
|
||||
|
||||
### What it does
|
||||
|
||||
1. Validates `$CEPH_ENV` is set and the inventory file exists.
|
||||
2. Derives the secrets template path: `$(dirname $CEPH_ENV)/secrets.yml.tpl`.
|
||||
3. Verifies `op account get` succeeds (1P desktop session or
|
||||
`OP_SERVICE_ACCOUNT_TOKEN`). Fails fast if unavailable.
|
||||
4. `mktemp`s a tmpfile, `chmod 600`, registers a `trap` to delete it on
|
||||
`EXIT INT TERM` (including operator Ctrl-C or SIGKILL'd parents).
|
||||
5. Runs `op inject -f -i <template> -o <tmpfile>`. Fails if any `op://`
|
||||
reference can't be resolved.
|
||||
6. `exec`s `ansible-playbook -i $CEPH_ENV --extra-vars @<tmpfile> <args>`.
|
||||
|
||||
The `exec` means the wrapper process is replaced — the trap still fires via
|
||||
the shell's EXIT handler on the child's termination.
|
||||
|
||||
### Environment
|
||||
|
||||
| Variable | Required | Purpose |
|
||||
|-----------------------------|----------|-----------------------------------------------------------------|
|
||||
| `CEPH_ENV` | yes | Path to the target cluster's `inventory.ini` |
|
||||
| `OP_SERVICE_ACCOUNT_TOKEN` | no | CI headless auth. Falls back to `op` desktop session if unset. |
|
||||
|
||||
### Arguments
|
||||
|
||||
Everything after the playbook name is passed to `ansible-playbook` verbatim.
|
||||
Common patterns:
|
||||
|
||||
```bash
|
||||
scripts/ansible-play.sh baseline.yml --check --diff
|
||||
scripts/ansible-play.sh deploy-ceph.yml --tags rgw,monitoring
|
||||
scripts/ansible-play.sh destroy-ceph.yml -e yes_destroy_ceph=true -e destroy_target_domain=dev.austin.int.futo.cloud
|
||||
```
|
||||
|
||||
The destroy playbook requires both safety gates:
|
||||
- `yes_destroy_ceph=true` — explicit confirmation (otherwise the play refuses to run)
|
||||
- `destroy_target_domain=<cluster domain>` — must match the inventory's `cluster_domain`. Mismatch aborts the play, guarding against running destroy with the wrong `CEPH_ENV`.
|
||||
|
||||
The `mise run destroy` task (in `.mise.toml`) builds these arguments automatically from `CEPH_ENV` and adds an interactive `[y/N]` prompt — prefer it over invoking the wrapper directly.
|
||||
|
||||
### Exit codes
|
||||
|
||||
| Code | Meaning |
|
||||
|------|-------------------------------------------------------------------|
|
||||
| 0 | ansible-playbook completed successfully |
|
||||
| 1 | Inventory file missing or secrets template missing |
|
||||
| 2 | 1Password session unavailable (`op account get` failed) |
|
||||
| 3 | `op inject` failed (bad `op://` reference, missing item, etc.) |
|
||||
| ≥4 | ansible-playbook's own exit code (unreachable hosts, failed tasks) |
|
||||
|
||||
### Examples
|
||||
|
||||
```bash
|
||||
# Standard deploy
|
||||
CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory.ini \
|
||||
scripts/ansible-play.sh deploy-ceph.yml
|
||||
|
||||
# Dry-run a role via tags
|
||||
CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory.ini \
|
||||
scripts/ansible-play.sh site.yml --check --diff --tags baseline
|
||||
|
||||
# CI / headless (SA token from env)
|
||||
OP_SERVICE_ACCOUNT_TOKEN="$(...)" \
|
||||
CEPH_ENV=inventories/sietch-ceph.dev.austin.int/inventory.ini \
|
||||
scripts/ansible-play.sh status.yml
|
||||
```
|
||||
|
||||
### Related
|
||||
|
||||
- [architecture.md §9.2](architecture.md) — deploy data flow
|
||||
- [secrets.md](secrets.md) — what's in `secrets.yml.tpl`
|
||||
|
||||
---
|
||||
|
||||
## `install-ssh-keys.sh`
|
||||
|
||||
Idempotent installer that pulls per-cluster `ansible-iac` SSH keypairs from
|
||||
1Password into `~/.ssh/`. For new workstations or after a key rotation.
|
||||
|
||||
### Synopsis
|
||||
|
||||
```
|
||||
scripts/install-ssh-keys.sh [cluster...]
|
||||
```
|
||||
|
||||
If no cluster arguments are given, installs keys for every known cluster
|
||||
(`sietch`, `painbox`).
|
||||
|
||||
### What it does
|
||||
|
||||
For each target cluster:
|
||||
|
||||
1. Resolves the 1P item name: `<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY` in
|
||||
`yucca_tf_dev`.
|
||||
2. Resolves the target filename: `~/.ssh/id_ed25519_<cluster>`.
|
||||
3. If the key already exists on disk:
|
||||
- Compares fingerprints (local vs 1P).
|
||||
- **Match** → skip (idempotent re-run).
|
||||
- **Mismatch** → refuse to overwrite. Prints the `mv` command the
|
||||
operator should run manually. Exit 2.
|
||||
4. If the key is missing:
|
||||
- `umask 077`.
|
||||
- Writes `private_key` → `~/.ssh/id_ed25519_<cluster>` (`chmod 600`).
|
||||
- Writes `public_key` → `~/.ssh/id_ed25519_<cluster>.pub` (`chmod 644`).
|
||||
- Prints the installed key's fingerprint.
|
||||
|
||||
### Environment
|
||||
|
||||
| Variable | Required | Purpose |
|
||||
|-----------------------------|----------|---------------------------------------------------------|
|
||||
| `OP_SERVICE_ACCOUNT_TOKEN` | no | CI headless auth. Read-only scope on `yucca_tf_dev` is enough. |
|
||||
|
||||
### Arguments
|
||||
|
||||
Zero or more cluster short names (`sietch`, `painbox`). With no args, all
|
||||
known clusters are installed.
|
||||
|
||||
### Exit codes
|
||||
|
||||
| Code | Meaning |
|
||||
|------|------------------------------------------------------------------|
|
||||
| 0 | All requested keys installed (or already present and matching) |
|
||||
| 1 | Unknown cluster name |
|
||||
| 2 | Fingerprint mismatch — refused to overwrite existing key on disk |
|
||||
|
||||
### Examples
|
||||
|
||||
```bash
|
||||
# Install all known keys
|
||||
scripts/install-ssh-keys.sh
|
||||
|
||||
# Just one cluster
|
||||
scripts/install-ssh-keys.sh sietch
|
||||
|
||||
# Re-run after a rotation (operator already moved the old key aside manually)
|
||||
mv ~/.ssh/id_ed25519_sietch ~/.ssh/id_ed25519_sietch.20260423.bak
|
||||
mv ~/.ssh/id_ed25519_sietch.pub ~/.ssh/id_ed25519_sietch.pub.20260423.bak
|
||||
scripts/install-ssh-keys.sh sietch
|
||||
```
|
||||
|
||||
### Related
|
||||
|
||||
- [runbooks/rotate-ssh-key.md](runbooks/rotate-ssh-key.md) — the full rotation flow
|
||||
- [ADR-010](adr/010-ssh-keys-in-1password.md) — why SSH keys live in 1P
|
||||
|
||||
---
|
||||
|
||||
## `preflight.sh`
|
||||
|
||||
Read-only smoke test — verifies the controller environment is ready to run
|
||||
destructive playbooks against the target cluster. Surfaced via
|
||||
`mise run preflight`.
|
||||
|
||||
### Synopsis
|
||||
|
||||
```
|
||||
CEPH_ENV=inventories/<cluster>/inventory.ini scripts/preflight.sh
|
||||
```
|
||||
|
||||
### What it checks
|
||||
|
||||
**Controller:**
|
||||
|
||||
- `scripts/ansible-play.sh` is executable
|
||||
- Inventory file exists (TF-rendered)
|
||||
- Secrets template exists (TF-rendered)
|
||||
- SSH private + public keys exist on disk
|
||||
- `ansible` and `op` CLIs installed
|
||||
- 1Password session live (`op account get`)
|
||||
|
||||
**Secrets:**
|
||||
|
||||
- `op inject` resolves the cluster's `secrets.yml.tpl` successfully
|
||||
- Resolved file has a non-empty `vault_ops_password` line (sanity check
|
||||
that op isn't silently substituting empty strings)
|
||||
|
||||
**Target connectivity** (one check per host in `ceph_nodes`):
|
||||
|
||||
- SSH reachability via `ansible -m ping`
|
||||
- Python 3 available via `ansible -m raw`
|
||||
|
||||
### Environment
|
||||
|
||||
| Variable | Required | Purpose |
|
||||
|-----------------------------|----------|-------------------------------------------------------|
|
||||
| `CEPH_ENV` | yes | Path to the target cluster's `inventory.ini` |
|
||||
| `OP_SERVICE_ACCOUNT_TOKEN` | no | CI headless auth. Falls back to `op` desktop session. |
|
||||
|
||||
### Exit codes
|
||||
|
||||
| Code | Meaning |
|
||||
|------|------------------------------------------------------------------------------|
|
||||
| 0 | All checks passed |
|
||||
| 1 | One or more checks failed — summary printed, unsafe to proceed |
|
||||
|
||||
Warnings (non-blocking) are reported in the summary but don't affect exit.
|
||||
|
||||
### Examples
|
||||
|
||||
```bash
|
||||
# Via mise (recommended)
|
||||
mise run preflight
|
||||
|
||||
# Direct, against painbox
|
||||
CEPH_ENV=inventories/painbox-ceph.dev.hel.htz/inventory.ini \
|
||||
scripts/preflight.sh
|
||||
```
|
||||
|
||||
### Related
|
||||
|
||||
- [architecture.md §8](architecture.md) — where the wrappers fit in the system mesh
|
||||
|
||||
---
|
||||
|
||||
## Adding a new wrapper
|
||||
|
||||
Follow these conventions:
|
||||
|
||||
1. **Fail closed** — `set -euo pipefail` at the top. Any uncaught error
|
||||
aborts the script.
|
||||
2. **Validate inputs early** — check required env vars (`: "${CEPH_ENV:?...}"`),
|
||||
then check that referenced files exist, before doing any real work.
|
||||
3. **Never interpolate secrets into argv** — write them to a `mktemp`'d
|
||||
file (`chmod 600`) and pass the path. Always `trap 'rm -f "$TMP"' EXIT INT TERM`.
|
||||
4. **Distinct exit codes** — the caller (mise task or another script) should
|
||||
be able to tell "inventory missing" from "1P unreachable" from "ansible
|
||||
failed" without parsing stderr.
|
||||
5. **Idempotent where plausible** — re-running the script should not make
|
||||
things worse. Prefer skip-if-already-correct over unconditional overwrite.
|
||||
6. **Cross-reference from [architecture.md §8](architecture.md)** and add a
|
||||
section to this file with the same shape as the ones above.
|
||||
@@ -0,0 +1,143 @@
|
||||
# Secrets
|
||||
|
||||
This project uses a **TF-provisions, op-injects, ansible-consumes** model. No
|
||||
`ansible-vault`, no encrypted `vault.yml` in git, no custom password-file script.
|
||||
|
||||
For how secrets fit into the broader architecture, see
|
||||
[architecture.md §5 (1Password)](architecture.md).
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
ONEP[("1Password<br/>yucca_tf · yucca_tf_dev · ...<br/><i>source of truth</i>")]
|
||||
TF[Terraform / Tofu<br/>tf/deployment/dev/ceph/]
|
||||
REPO[/"inventories/<cluster>/<br/>inventory.ini (TF-gen, gitignored)<br/>secrets.yml.tpl (TF-gen, gitignored)"/]
|
||||
WRAP[scripts/ansible-play.sh<br/><i>mktemp + op inject → exec ansible-playbook --extra-vars @tmp</i>]
|
||||
ANS[ansible-playbook]
|
||||
|
||||
TF -.->|reads via op run --env-file| ONEP
|
||||
TF -->|renders| REPO
|
||||
WRAP -->|reads template| REPO
|
||||
WRAP -.->|op inject -f<br/>at play time| ONEP
|
||||
WRAP --> ANS
|
||||
```
|
||||
|
||||
## What lives where
|
||||
|
||||
| Vault | Purpose | Who writes |
|
||||
|------------------------------------|-------------------------------------------------------------------------------------------|-----------------------------------------------------------------|
|
||||
| `yucca_tf` (team-shared) | Cross-env shared state: TF state S3 credentials | Operator (manual) |
|
||||
| `yucca_tf_dev` (team-shared) | Live values for the `dev` environment — `<CLUSTER>_CEPH_*` items | Superuser service account (TF) + operator via `op` CLI |
|
||||
| `yucca_tf_dev_manual` (team-shared)| Human-fillable dev placeholders (3rd-party API tokens, OAuth client secrets) — not yet used by ceph-cluster | Operator (manual) |
|
||||
|
||||
Future environments land as siblings: `yucca_tf_staging(_manual)`,
|
||||
`yucca_tf_prod_manual`. The vault a given cluster reads from is declared
|
||||
per-cluster in `tf/deployment/<env>/ceph/clusters.auto.tfvars` (field
|
||||
`vault`). TF derives item paths from that field at render time; changing
|
||||
it + `tofu apply` re-renders `secrets.yml.tpl` with the new vault path.
|
||||
|
||||
## Item naming
|
||||
|
||||
Format: `<CLUSTER>_CEPH_<ROLE>_PASSWORD`
|
||||
|
||||
**Password items** (category `Password`, consumed via `op inject` at
|
||||
playbook time):
|
||||
|
||||
| Item | Field | Consumed as |
|
||||
|---|---|---|
|
||||
| `SIETCH_CEPH_OPS_PASSWORD` | `password` | `vault_ops_password` |
|
||||
| `SIETCH_CEPH_DASHBOARD_PASSWORD` | `password` | `vault_ceph_dashboard_password` |
|
||||
| `SIETCH_CEPH_GRAFANA_PASSWORD` | `password` | `vault_grafana_admin_password` |
|
||||
| `PAINBOX_CEPH_OPS_PASSWORD` | `password` | `vault_ops_password` |
|
||||
| `PAINBOX_CEPH_DASHBOARD_PASSWORD` | `password` | `vault_ceph_dashboard_password` |
|
||||
| `PAINBOX_CEPH_GRAFANA_PASSWORD` | `password` | `vault_grafana_admin_password` |
|
||||
|
||||
**SSH Key items** (category `SSH Key`, consumed via
|
||||
`scripts/install-ssh-keys.sh` on operator workstations and
|
||||
`rotate-ssh-key.yml` on cluster nodes — see [ADR-010](adr/010-ssh-keys-in-1password.md)):
|
||||
|
||||
| Item | Field | Consumed as |
|
||||
|---|---|---|
|
||||
| `SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY` | `private_key` | `~/.ssh/id_ed25519_sietch` on operator workstation |
|
||||
| `SIETCH_CEPH_ANSIBLE_IAC_SSH_KEY` | `public_key` | sietch nodes' `ansible-iac@:~/.ssh/authorized_keys` |
|
||||
| `PAINBOX_CEPH_ANSIBLE_IAC_SSH_KEY` | `private_key` | `~/.ssh/id_ed25519_painbox` on operator workstation |
|
||||
| `PAINBOX_CEPH_ANSIBLE_IAC_SSH_KEY` | `public_key` | painbox nodes' `ansible-iac@:~/.ssh/authorized_keys` |
|
||||
|
||||
**S3 service-user items** (predetermined keys passed to
|
||||
`radosgw-admin user create --access-key=X --secret-key=Y` at deploy
|
||||
time — Yucca app / restic client can be pre-configured with matching
|
||||
credentials without waiting for post-bootstrap capture):
|
||||
|
||||
| Item | Field | Consumed as |
|
||||
|---|---|---|
|
||||
| `SIETCH_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY` | `password` | `vault_s3_restic_access_key` → `ceph_rgw_s3_user_access_key` |
|
||||
| `SIETCH_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY` | `password` | `vault_s3_restic_secret_key` → `ceph_rgw_s3_user_secret_key` |
|
||||
| `PAINBOX_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY` | `password` | same, painbox |
|
||||
| `PAINBOX_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY` | `password` | same, painbox |
|
||||
|
||||
**Disaster-recovery items** (populated by `mise run capture` after
|
||||
deploy — stored in 1P for recovery if the bootstrap node's filesystem
|
||||
is lost):
|
||||
|
||||
| Item | Field | Source |
|
||||
|---|---|---|
|
||||
| `<CLUSTER>_CEPH_RGW_TLS_CERT` | `password` (concealed) | `/etc/ceph/rgw-ssl.crt` on bootstrap |
|
||||
| `<CLUSTER>_CEPH_RGW_TLS_KEY` | `password` (concealed) | `/etc/ceph/rgw-ssl.key` on bootstrap |
|
||||
| `<CLUSTER>_CEPH_CLIENT_ADMIN_KEYRING` | `password` (concealed) | `/etc/ceph/ceph.client.admin.keyring` on bootstrap |
|
||||
|
||||
Items get created on first `mise run capture`; subsequent runs update
|
||||
in place on content drift.
|
||||
|
||||
Item names are derived in `tf/shared/modules/ceph-cluster/main.tf`
|
||||
(`local.secret_prefix`). Hardcoded `CEPH` (not `role_in_hostname`) so every
|
||||
Ceph-project secret grep-matches `*_CEPH_*` regardless of whether hostnames
|
||||
use `ceph`, `osd`, or `mon` as the role segment.
|
||||
|
||||
## Runtime flow
|
||||
|
||||
The op CLI is invoked in three distinct patterns across this project:
|
||||
|
||||
| Pattern | Used by | What it does |
|
||||
|----------------------------|---------------------------------------------------------------------------|--------------------------------------------------------|
|
||||
| `op run --env-file=tf/.env --` | Every `mise run tf:*` task | Resolves `op://` references in a dotenv file, injects as env vars into child process (TF: SA token + AWS creds) |
|
||||
| `op inject -f -i tpl -o out` | `scripts/ansible-play.sh`, Hetzner installimage post-install rendering | Resolves all `op://` references in a file template, writes resolved file |
|
||||
| `op read "op://..."` | `scripts/install-ssh-keys.sh`, `rotate-ssh-key.yml`, `post-deploy-capture.yml` | Reads a single field from a single item to stdout |
|
||||
|
||||
### The `ansible-play.sh` flow
|
||||
|
||||
`scripts/ansible-play.sh` wraps every `ansible-playbook` invocation:
|
||||
|
||||
1. Verifies `op account get` succeeds — fails closed if 1Password is locked.
|
||||
2. `mktemp` a `0600` tmpfile with `trap` cleanup on `EXIT` / `INT` / `TERM`.
|
||||
3. `op inject -f -i <cluster>/secrets.yml.tpl -o $tmpfile` — resolves every
|
||||
`op://` reference. Exit non-zero if any reference can't be resolved.
|
||||
4. `exec ansible-playbook --extra-vars @$tmpfile ...`.
|
||||
|
||||
Ansible task references stay unchanged — `vault_ops_password` etc. are
|
||||
regular variables populated from extra-vars (highest precedence).
|
||||
|
||||
Full wrapper reference: [scripts.md](scripts.md).
|
||||
|
||||
## CI / headless
|
||||
|
||||
Set `OP_SERVICE_ACCOUNT_TOKEN` (from a service-account item) in the CI
|
||||
environment. `op inject` uses it automatically. The cutover PR must
|
||||
demonstrate green `lint` and `check` with **zero** op credentials — those
|
||||
tasks do not require secrets.
|
||||
|
||||
## Rotating secrets
|
||||
|
||||
See [docs/runbooks/rotate-secrets.md](runbooks/rotate-secrets.md).
|
||||
|
||||
## Adding a new secret
|
||||
|
||||
1. Add the item name to `local.secrets` in
|
||||
`tf/shared/modules/ceph-cluster/main.tf`.
|
||||
2. Add the matching line to the secrets template
|
||||
(`templates/secrets.yml.tpl.tftpl`).
|
||||
3. Add the ansible variable alias in each cluster's
|
||||
`inventories/<cluster>/group_vars/all/vars.yml`.
|
||||
4. Create the item manually in the target vault (or let TF do it post
|
||||
service account): `op item create --vault <vault> --category password
|
||||
--title NEW_SECRET_NAME --generate-password=letters,digits,32`.
|
||||
5. `tofu apply` — re-renders templates with the new reference.
|
||||
6. Playbooks consuming the variable now have it available.
|
||||
@@ -0,0 +1,236 @@
|
||||
# Security Model
|
||||
|
||||
Audience: InfoSec, compliance review, security audits.
|
||||
|
||||
For how these controls fit the broader tool mesh, see
|
||||
[architecture.md](architecture.md). For the per-item secrets catalog see
|
||||
[secrets.md](secrets.md).
|
||||
|
||||
## Encryption at rest
|
||||
|
||||
All OSD volumes (HDD and SSD) are created with `--dmcrypt`, which wraps each
|
||||
OSD's BlueStore volume in a dm-crypt/LUKS layer.
|
||||
|
||||
| Property | Detail |
|
||||
|---|---|
|
||||
| Encryption layer | dm-crypt (LUKS) via cephadm OSD service spec (`encrypted: true`); see [ADR-011](adr/011-cephadm-osd-service-specs.md) |
|
||||
| Scope | Every OSD data volume and every OSD block.db volume |
|
||||
| Key storage | MON config-key database (`config-key dump` shows dm-crypt entries) |
|
||||
| Key distribution | MONs hand keys to OSD daemons at startup via the Ceph auth subsystem |
|
||||
| Algorithm | AES-256-XTS (dm-crypt default) |
|
||||
|
||||
The encryption keys never leave the MON quorum. If a drive is removed from
|
||||
the chassis, the data is unreadable without access to the MON key store.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
# Count dm-crypt keys in MON store
|
||||
ceph config-key dump 2>/dev/null | grep -c dm-crypt
|
||||
|
||||
# Confirm OSDs are on dm-crypt devices
|
||||
lsblk --output NAME,TYPE,MOUNTPOINT | grep crypt
|
||||
```
|
||||
|
||||
## Encryption in transit
|
||||
|
||||
### RGW / S3 (client-facing)
|
||||
|
||||
| Property | Detail |
|
||||
|---|---|
|
||||
| Protocol | HTTPS (TLS 1.2+) on port 443 |
|
||||
| Certificate | Self-signed RSA 4096-bit, 10-year validity |
|
||||
| CN | `s3.dev.austin.int.futo.cloud` |
|
||||
| SANs | `s3.dev.austin.int.futo.cloud`, `*.s3.dev.austin.int.futo.cloud`, per-node FQDNs, per-node bond IPs |
|
||||
| Issuer | Self-signed (O=FUTO, L=Austin, ST=Texas, C=US) |
|
||||
| Cert location | `/etc/ceph/rgw-ssl.crt` + `/etc/ceph/rgw-ssl.key` on bootstrap node |
|
||||
| Distribution | cephadm distributes combined PEM to all RGW daemon containers |
|
||||
|
||||
Clients must either trust the self-signed cert via `--ca-bundle` / `verify=`
|
||||
or disable TLS verification (`--no-verify-ssl`).
|
||||
|
||||
### Intra-cluster (MON/OSD/MGR)
|
||||
|
||||
Ceph messenger v2 (`msgr2`) is used for all intra-cluster communication.
|
||||
The `cephx` authentication protocol provides mutual authentication between
|
||||
daemons. Wire encryption (`ms_client_mode`, `ms_cluster_mode`) is available
|
||||
but not explicitly forced in this deployment -- the default Tentacle
|
||||
configuration uses `crc` mode (authenticated but not encrypted on the wire).
|
||||
|
||||
### Dashboard
|
||||
|
||||
Ceph dashboard runs on port 8443 (HTTPS) with its own self-signed cert,
|
||||
restricted to trusted networks only via firewall rules.
|
||||
|
||||
## User model
|
||||
|
||||
Three user accounts exist on every node:
|
||||
|
||||
| User | UID | Authentication | Sudo | Purpose |
|
||||
|---|---|---|---|---|
|
||||
| `ansible-iac` | 1000 | SSH key only (password locked) | `NOPASSWD:ALL` | Ansible automation. No interactive use. |
|
||||
| `ops` | 1001 | Password (from 1Password via `op inject`) | `ALL` (password required) | Human interactive access. Operator SSH keys distributed out-of-band post-deploy. |
|
||||
| `root` | 0 | SSH key (cephadm inter-node) | n/a | Required by cephadm orchestrator for SSH between nodes. `PermitRootLogin prohibit-password`. |
|
||||
| `ceph` | system | nologin shell | none | Ceph daemon processes. No login capability. |
|
||||
|
||||
### SSH hardening
|
||||
|
||||
| Setting | Value |
|
||||
|---|---|
|
||||
| `PermitRootLogin` | `prohibit-password` (key-only, required by cephadm) |
|
||||
| `MaxAuthTries` | 3 |
|
||||
| `AllowUsers` | `ansible-iac ops root` |
|
||||
| `PasswordAuthentication` | Not explicitly disabled (ops user uses password + key) |
|
||||
|
||||
### Key management
|
||||
|
||||
- **ansible-iac key**: The keypair is stored in 1Password as an SSH Key item
|
||||
(`<CLUSTER>_CEPH_ANSIBLE_IAC_SSH_KEY` in `yucca_tf_dev`). Operator
|
||||
workstations install it via `scripts/install-ssh-keys.sh` →
|
||||
`~/.ssh/id_ed25519_<cluster>`. Rotation is forward-only (additive) via
|
||||
`rotate-ssh-key.yml` — new pubkey distributed to `authorized_keys` with
|
||||
`exclusive: false`, old keys pruned out-of-band. See
|
||||
[ADR-010](adr/010-ssh-keys-in-1password.md).
|
||||
- **ops user keys**: Distributed out-of-band after initial deployment.
|
||||
No keys are provisioned at deploy time.
|
||||
- **cephadm key**: Generated during `cephadm bootstrap`, distributed to all
|
||||
nodes during the join phase. Used for orchestrator SSH between nodes.
|
||||
|
||||
## Firewall rules (nftables)
|
||||
|
||||
Every node runs nftables with a default-drop input policy. The ruleset is
|
||||
templated from `roles/security/templates/nftables.conf.j2`.
|
||||
|
||||
### Open to all sources
|
||||
|
||||
| Port | Service |
|
||||
|---|---|
|
||||
| 22/tcp | SSH (configurable: `ceph_firewall_ssh_any_source` defaults to true for dev) |
|
||||
| 443/tcp | RGW / S3 endpoint |
|
||||
| ICMP | Ping and MTU discovery |
|
||||
|
||||
### Restricted to trusted networks only
|
||||
|
||||
Trusted networks: `10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16`,
|
||||
`100.64.0.0/10` (Tailscale CGNAT).
|
||||
|
||||
| Port(s) | Service |
|
||||
|---|---|
|
||||
| 3300, 6789/tcp | Ceph MON |
|
||||
| 6800-7300/tcp | Ceph OSD |
|
||||
| 8443/tcp | Ceph Dashboard |
|
||||
| 9095/tcp | Prometheus |
|
||||
| 3000/tcp | Grafana |
|
||||
| 9093/tcp | Alertmanager |
|
||||
| 9100/tcp | Node Exporter |
|
||||
| 9283/tcp | MGR Exporter |
|
||||
| 9926/tcp | Ceph Exporter |
|
||||
|
||||
### Default policy
|
||||
|
||||
```
|
||||
chain input { policy drop; } # all unmatched traffic is dropped
|
||||
chain forward { policy drop; } # no forwarding
|
||||
chain output { policy accept; } # outbound unrestricted
|
||||
```
|
||||
|
||||
Dropped packets are logged at rate 5/minute with prefix `nftables-drop: ` for
|
||||
forensic review.
|
||||
|
||||
### Optional gateways (disabled by default)
|
||||
|
||||
- iSCSI (port 3260 + API 5000) -- `ceph_firewall_iscsi_enabled: false`
|
||||
- NFS-Ganesha (port 2049 + mgmt 12049) -- `ceph_firewall_nfs_enabled: false`
|
||||
|
||||
## Secrets management
|
||||
|
||||
### 1Password + op inject
|
||||
|
||||
All deployment secrets live in 1Password in the `yucca_tf_dev` vault (dev
|
||||
environment; staging/prod land as sibling `yucca_tf_staging` /
|
||||
`yucca_tf_prod` vaults). Items are named `<CLUSTER>_CEPH_<ROLE>_*` — see
|
||||
[docs/secrets.md](secrets.md) for the full catalog.
|
||||
|
||||
At playbook time, `scripts/ansible-play.sh` invokes `op inject` on the
|
||||
TF-rendered `secrets.yml.tpl`, writes the resolved values to a `0600`
|
||||
tmpfile, and passes it as `--extra-vars @tmpfile`. Tmpfile is `trap`-cleaned
|
||||
on `EXIT`/`INT`/`TERM`. No at-rest encrypted file in git.
|
||||
|
||||
| Secret | Variable |
|
||||
|----------------------------------|-----------------------------------|
|
||||
| ops user password | `vault_ops_password` |
|
||||
| Ceph dashboard admin password | `vault_ceph_dashboard_password` |
|
||||
| Grafana admin password | `vault_grafana_admin_password` |
|
||||
| S3 service-user access key | `vault_s3_restic_access_key` |
|
||||
| S3 service-user secret key | `vault_s3_restic_secret_key` |
|
||||
|
||||
### Trust boundaries
|
||||
|
||||
Two service accounts separate write authority from runtime consumption:
|
||||
|
||||
| Service account | Scope | Used by |
|
||||
|-----------------------------------------------|--------------------------------------------|----------------------------------------|
|
||||
| `yucca_futo_1pass_superuser_service_account` | Read + write all `yucca_tf_*` vaults | TF (`tf/.env`) + interactive `op` CLI |
|
||||
| `yucca_futo_1pass_service_account` | Read-only on `yucca_tf` and `yucca_tf_dev` | Ansible runtime / future CI |
|
||||
|
||||
The ansible-play.sh wrapper runs with whichever session is active on the
|
||||
operator workstation (typically desktop unlock → superuser SA via 1Password
|
||||
desktop). CI will use the read-only SA via `OP_SERVICE_ACCOUNT_TOKEN`.
|
||||
|
||||
Rotation procedure: [runbooks/rotate-sa-token.md](runbooks/rotate-sa-token.md).
|
||||
|
||||
### What is NOT in 1Password
|
||||
|
||||
- dm-crypt OSD encryption keys — stored in the MON config-key database.
|
||||
Deferred to 1P as a future belt-and-suspenders item (see yucca memory
|
||||
`project_luks_keys_in_1pass.md`).
|
||||
- cephadm bootstrap SSH key — generated at bootstrap, distributed by
|
||||
cephadm's orchestrator. Not needed outside the cluster.
|
||||
- ops user SSH keys — distributed out-of-band after initial deployment.
|
||||
See "ops user keys" in Key management above.
|
||||
|
||||
## Audit logging
|
||||
|
||||
| Control | Detail |
|
||||
|---|---|
|
||||
| Ceph audit log | Enabled (`ceph config set global log_to_cluster_level audit`) |
|
||||
| Ceph manager log | `mgr/cephadm/log_to_cluster true` |
|
||||
| View audit log | `ceph log last -W audit` |
|
||||
| Firewall logging | Dropped packets logged at 5/min with `nftables-drop:` prefix |
|
||||
| Prometheus alerts | 16+ alert rule groups (89 built-in rules) covering OSD, MON, PG, pool, MDS, hardware, network |
|
||||
|
||||
The audit channel records all `ceph` admin commands executed against the
|
||||
cluster, including the authenticated user, timestamp, and command arguments.
|
||||
|
||||
## Telemetry
|
||||
|
||||
Telemetry phone-home is explicitly disabled:
|
||||
|
||||
```
|
||||
ceph_telemetry_enabled: false
|
||||
```
|
||||
|
||||
No cluster metadata, performance data, or crash reports are sent to upstream
|
||||
Ceph. All data stays within the cluster boundary.
|
||||
|
||||
## Network topology
|
||||
|
||||
| Network | CIDR | Purpose |
|
||||
|---|---|---|
|
||||
| Public/Cluster | 10.10.10.0/24 | Combined public + cluster traffic (single-network topology) |
|
||||
| iDRAC/BMC | 10.10.11.0/24 | Out-of-band management (separate VLAN) |
|
||||
|
||||
Nodes are on a private network -- the 10.10.10.0/24 subnet is not directly
|
||||
routable from the public internet.
|
||||
|
||||
## Known gaps and mitigations
|
||||
|
||||
| Gap | Risk | Mitigation |
|
||||
|---|---|---|
|
||||
| Self-signed TLS cert | Clients must disable verification or trust the CA manually. MITM possible if cert is not pinned. | Cert has 10-year validity with specific SANs. Production (Yucca) will use cert-manager + real CA behind haproxy. |
|
||||
| Single network (public = cluster) | Cluster rebalancing traffic is visible on the client network. A compromised client could sniff inter-OSD traffic. | Trusted network firewall restricts cluster ports to RFC1918 + Tailscale. Production will separate public and cluster networks. |
|
||||
| `msgr2` wire encryption not forced | Intra-cluster traffic is authenticated (cephx) but not encrypted on the wire by default. | All traffic stays within 10.10.10.0/24 on a private switch. Can be enabled via `ceph config set global ms_cluster_mode secure` if needed. |
|
||||
| SSH open to all sources (dev default) | SSH is reachable from any IP that can route to the nodes. | On a private network (not internet-exposed). `MaxAuthTries=3`, `AllowUsers` whitelist. Production should set `ceph_firewall_ssh_any_source: false`. |
|
||||
| `ops` password auth | Password-based SSH is not disabled. | Password is vault-encrypted, rotated via Ansible. Interactive use only -- automation uses key-only `ansible-iac`. |
|
||||
| Operator SSH keys distributed out-of-band | No automated key lifecycle for `ops` user. | Acceptable for dev. Production should use centralized key management (e.g., Teleport, Vault SSH CA). |
|
||||
| RGW S3 credentials static | No automatic rotation of S3 access/secret keys. | Keys stored in 1Password with access control. `radosgw-admin key create/rm` available for manual rotation. |
|
||||
@@ -0,0 +1,684 @@
|
||||
# Troubleshooting Guide
|
||||
|
||||
Decision-tree format: symptom, diagnosis, fix.
|
||||
|
||||
---
|
||||
|
||||
## Cluster Health
|
||||
|
||||
### HEALTH_WARN
|
||||
|
||||
#### Symptom: `HEALTH_WARN: N osds down`
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
ceph osd tree | grep down
|
||||
ceph health detail
|
||||
```
|
||||
|
||||
**Common causes and fixes:**
|
||||
|
||||
1. **OSD daemon crashed** -- check logs:
|
||||
|
||||
```bash
|
||||
ceph crash ls-new
|
||||
ceph crash info <crash-id>
|
||||
# Or on the host:
|
||||
journalctl -u ceph-osd@<id> --since "1 hour ago"
|
||||
```
|
||||
|
||||
Restart the daemon:
|
||||
|
||||
```bash
|
||||
ceph orch daemon restart osd.<id>
|
||||
```
|
||||
|
||||
2. **Host unreachable** -- the node itself is down:
|
||||
|
||||
```bash
|
||||
ssh ansible-iac@sietch-ceph-<name> hostname
|
||||
# If unreachable, check iDRAC / physical console
|
||||
```
|
||||
|
||||
3. **Disk failure** -- see the [replace-disk runbook](runbooks/replace-disk.md).
|
||||
|
||||
4. **OSD out of disk space** -- check nearfull/full ratios:
|
||||
|
||||
```bash
|
||||
ceph osd df
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
#### Symptom: `HEALTH_WARN: N pgs not active+clean`
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
ceph pg stat
|
||||
ceph pg dump_stuck
|
||||
```
|
||||
|
||||
**Common causes and fixes:**
|
||||
|
||||
1. **Backfill in progress** -- normal after adding/removing OSDs. Monitor:
|
||||
|
||||
```bash
|
||||
ceph -s
|
||||
```
|
||||
|
||||
Wait for completion. No action needed.
|
||||
|
||||
2. **Stale PGs** -- PGs stuck in `stale` state:
|
||||
|
||||
```bash
|
||||
ceph pg dump_stuck stale
|
||||
```
|
||||
|
||||
Usually indicates the hosting OSD is down. Fix the OSD first.
|
||||
|
||||
3. **Inactive PGs** -- PGs in `creating` or `peering`:
|
||||
|
||||
```bash
|
||||
ceph pg dump_stuck inactive
|
||||
```
|
||||
|
||||
If stuck for >15 minutes, check MON logs. May need `ceph pg force-create-pg <pgid>`.
|
||||
|
||||
---
|
||||
|
||||
#### Symptom: `HEALTH_WARN: clock skew detected`
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
ceph time-sync-status
|
||||
# On each node:
|
||||
chronyc tracking
|
||||
```
|
||||
|
||||
**Fix:** Restart chrony on the affected node:
|
||||
|
||||
```bash
|
||||
sudo systemctl restart chrony
|
||||
```
|
||||
|
||||
If persistent, check NTP sources:
|
||||
|
||||
```bash
|
||||
chronyc sources -v
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
#### Symptom: `HEALTH_WARN: N daemons have recently crashed`
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
ceph crash ls-new
|
||||
ceph crash info <crash-id>
|
||||
```
|
||||
|
||||
**Fix:** Review the crash, then archive it:
|
||||
|
||||
```bash
|
||||
ceph crash archive <crash-id>
|
||||
# Or archive all:
|
||||
ceph crash archive-all
|
||||
```
|
||||
|
||||
If crashes are recurring, investigate the daemon logs on the host.
|
||||
|
||||
---
|
||||
|
||||
#### Symptom: `HEALTH_WARN: N pool(s) have no replicas configured`
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
ceph osd pool ls detail | grep 'size 1'
|
||||
```
|
||||
|
||||
**Fix:** Set appropriate replication:
|
||||
|
||||
```bash
|
||||
ceph osd pool set <pool> size 2
|
||||
ceph osd pool set <pool> min_size 1
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### HEALTH_ERR
|
||||
|
||||
#### Symptom: `HEALTH_ERR: N pgs are stuck inactive`
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
ceph pg dump_stuck inactive
|
||||
ceph osd tree
|
||||
```
|
||||
|
||||
**Fix:** This is critical -- data may be inaccessible.
|
||||
|
||||
1. Check if the hosting OSDs are down. Bring them up first.
|
||||
2. If OSDs are permanently lost and data cannot be recovered:
|
||||
|
||||
```bash
|
||||
# DANGER: marks missing PGs as complete with potential data loss
|
||||
ceph pg force-recovery <pgid>
|
||||
# Last resort:
|
||||
ceph osd force-create-pg <pgid> --yes-i-really-mean-it
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
#### Symptom: `HEALTH_ERR: N scrub errors`
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
ceph health detail
|
||||
# Find the affected PGs
|
||||
ceph pg dump | grep inconsistent
|
||||
# Deep scrub the PG
|
||||
ceph pg deep-scrub <pgid>
|
||||
```
|
||||
|
||||
**Fix:**
|
||||
|
||||
```bash
|
||||
ceph pg repair <pgid>
|
||||
```
|
||||
|
||||
If repair fails, the underlying disk may have bit rot. Check SMART data on
|
||||
the hosting OSDs.
|
||||
|
||||
---
|
||||
|
||||
#### Symptom: `HEALTH_ERR: full osds`
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
ceph osd df
|
||||
ceph df
|
||||
```
|
||||
|
||||
**Fix:** This is an emergency. The cluster stops accepting writes.
|
||||
|
||||
1. Delete unnecessary data or pools if possible
|
||||
2. Temporarily raise the full ratio:
|
||||
|
||||
```bash
|
||||
ceph osd set-full-ratio 0.97
|
||||
```
|
||||
|
||||
3. Add more OSDs (see [add-node runbook](runbooks/add-node.md))
|
||||
4. Set the ratio back after capacity is restored:
|
||||
|
||||
```bash
|
||||
ceph osd set-full-ratio 0.95
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## OSD Issues
|
||||
|
||||
### Symptom: OSD down and won't start
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
# Check daemon status
|
||||
ceph orch ps --daemon-type osd | grep <host>
|
||||
|
||||
# Check container logs
|
||||
ssh ansible-iac@<host>
|
||||
sudo podman logs ceph-<fsid>-osd.<id>
|
||||
sudo journalctl -u ceph-<fsid>@osd.<id>
|
||||
```
|
||||
|
||||
**Common causes:**
|
||||
|
||||
1. **LUKS key missing** -- dmcrypt key not in MON store:
|
||||
|
||||
```bash
|
||||
ceph config-key dump | grep dm-crypt | grep <osd-id>
|
||||
```
|
||||
|
||||
If missing, the OSD cannot be unlocked. Rebuild it (see replace-disk
|
||||
runbook).
|
||||
|
||||
2. **Corrupt BlueStore DB** -- look for `fsck` errors in the OSD log.
|
||||
May need `ceph-bluestore-tool repair`.
|
||||
|
||||
3. **Block device disappeared** -- check the SAS path:
|
||||
|
||||
```bash
|
||||
ls /dev/disk/by-path/ | grep phy<N>
|
||||
```
|
||||
|
||||
If missing, the disk or cable has failed.
|
||||
|
||||
---
|
||||
|
||||
### Symptom: Slow ops / blocked requests
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
ceph daemon osd.<id> dump_ops_in_flight
|
||||
ceph daemon osd.<id> perf dump | grep -i slow
|
||||
```
|
||||
|
||||
**Common causes:**
|
||||
|
||||
1. **Disk latency** -- check I/O wait:
|
||||
|
||||
```bash
|
||||
iostat -xz 5 3
|
||||
```
|
||||
|
||||
Look for `%util > 90%` or `await > 100ms` on HDD devices.
|
||||
|
||||
2. **Network issues** -- check for packet loss:
|
||||
|
||||
```bash
|
||||
ping -c 100 <other-node-ip>
|
||||
ethtool -S eno1 | grep -i error
|
||||
```
|
||||
|
||||
3. **Recovery throttling too aggressive** -- reduce recovery impact:
|
||||
|
||||
```bash
|
||||
ceph config set osd osd_recovery_max_active 1
|
||||
ceph config set osd osd_max_backfills 1
|
||||
ceph config set osd osd_recovery_sleep_hdd 0.1
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## RGW (S3) Issues
|
||||
|
||||
### Symptom: S3 requests return 403 Forbidden / SignatureDoesNotMatch
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
# Check if the Host header hostname is in the zonegroup
|
||||
radosgw-admin zonegroup get --rgw-zonegroup=us-east-1 | python3 -c "
|
||||
import sys, json
|
||||
zg = json.load(sys.stdin)
|
||||
print('Hostnames:', zg.get('hostnames', []))
|
||||
print('API name:', zg.get('api_name'))
|
||||
"
|
||||
```
|
||||
|
||||
**Common causes:**
|
||||
|
||||
1. **Missing hostname in zonegroup** -- the Host header used by the client
|
||||
is not in the zonegroup's hostnames list. S3 signature verification
|
||||
includes the Host header, so mismatches cause 403.
|
||||
|
||||
Fix:
|
||||
|
||||
```bash
|
||||
# Re-run the RGW role to add all node hostnames/IPs
|
||||
scripts/ansible-play.sh deploy-ceph.yml --tags rgw \
|
||||
--limit sietch-ceph-laurel
|
||||
```
|
||||
|
||||
2. **Wrong access/secret key** -- verify credentials:
|
||||
|
||||
```bash
|
||||
radosgw-admin user info --uid=svc-yucca-restic
|
||||
```
|
||||
|
||||
3. **Clock skew** -- S3 signatures are time-sensitive. Check client and
|
||||
server clocks are within 15 minutes.
|
||||
|
||||
---
|
||||
|
||||
### Symptom: S3 requests return 500 Internal Server Error
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
# Check RGW daemon logs
|
||||
ceph log last 50 --channel=cluster | grep rgw
|
||||
|
||||
# Check if RGW daemons are running
|
||||
ceph orch ls --service-type rgw
|
||||
ceph orch ps --daemon-type rgw
|
||||
```
|
||||
|
||||
**Common causes:**
|
||||
|
||||
1. **RGW daemons down** -- restart:
|
||||
|
||||
```bash
|
||||
ceph orch restart rgw
|
||||
```
|
||||
|
||||
2. **Pool issues** -- check that RGW pools exist and are healthy:
|
||||
|
||||
```bash
|
||||
ceph osd pool ls | grep rgw
|
||||
ceph pg stat
|
||||
```
|
||||
|
||||
3. **TLS cert expired or corrupted** -- see
|
||||
[rotate-certs runbook](runbooks/rotate-certs.md).
|
||||
|
||||
---
|
||||
|
||||
### Symptom: Dashboard Object Gateway page shows 500 or is empty
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
# Check dashboard RGW API SSL verification setting
|
||||
ceph dashboard get-rgw-api-ssl-verify
|
||||
|
||||
# Check dashboard RGW user has caps
|
||||
radosgw-admin user info --uid=dashboard | python3 -c "
|
||||
import sys, json
|
||||
u = json.load(sys.stdin)
|
||||
print('Caps:', u.get('caps', []))
|
||||
print('System:', u.get('system'))
|
||||
"
|
||||
```
|
||||
|
||||
**Fix:**
|
||||
|
||||
1. Disable SSL verification (self-signed certs):
|
||||
|
||||
```bash
|
||||
ceph dashboard set-rgw-api-ssl-verify false
|
||||
```
|
||||
|
||||
2. Add admin caps to dashboard user:
|
||||
|
||||
```bash
|
||||
radosgw-admin caps add --uid=dashboard \
|
||||
--caps='buckets=*;users=*;usage=*;metadata=*;zone=*'
|
||||
```
|
||||
|
||||
3. Sync dashboard credentials:
|
||||
|
||||
```bash
|
||||
ceph dashboard set-rgw-credentials
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Dashboard Issues
|
||||
|
||||
### Symptom: Dashboard unreachable at https://<ip>:8443
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
# Check MGR daemon status
|
||||
ceph orch ps --daemon-type mgr
|
||||
|
||||
# Check which MGR is active
|
||||
ceph mgr stat
|
||||
|
||||
# Check dashboard module is enabled
|
||||
ceph mgr module ls --format json | python3 -c "
|
||||
import sys, json
|
||||
d = json.load(sys.stdin)
|
||||
print('dashboard' in d.get('enabled_modules', []))
|
||||
"
|
||||
```
|
||||
|
||||
**Fix:**
|
||||
|
||||
1. **MGR daemon down** -- restart:
|
||||
|
||||
```bash
|
||||
ceph orch restart mgr
|
||||
```
|
||||
|
||||
2. **Dashboard module disabled**:
|
||||
|
||||
```bash
|
||||
ceph mgr module enable dashboard
|
||||
```
|
||||
|
||||
3. **Firewall blocking port 8443** -- check nftables:
|
||||
|
||||
```bash
|
||||
ssh ansible-iac@<host> sudo nft list ruleset | grep 8443
|
||||
```
|
||||
|
||||
If the port is not open, re-run security hardening:
|
||||
|
||||
```bash
|
||||
scripts/ansible-play.sh harden.yml
|
||||
```
|
||||
|
||||
4. **Wrong IP/port** -- check dashboard URL:
|
||||
|
||||
```bash
|
||||
ceph mgr services
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Deploy Issues
|
||||
|
||||
### Symptom: `mise run deploy` fails — `no available installation candidate for cephadm=20.2.*`
|
||||
|
||||
**Cause:** apt cache is stale. `prerequisites.yml` adds the Ceph Tentacle
|
||||
repo (`download.ceph.com/debian-tentacle`) and then runs `apt update`.
|
||||
On a freshly-installed OS (Hetzner installimage runs apt internally),
|
||||
the cache is "fresh" enough that an `apt update` with `cache_valid_time`
|
||||
set will skip the refresh — apt never sees the Ceph repo's Packages file
|
||||
and only knows about Debian's older `cephadm 16.2.x`.
|
||||
|
||||
**Diagnose:** SSH to the target node and run:
|
||||
|
||||
```bash
|
||||
cat /etc/apt/sources.list.d/ceph.list # confirm the repo file exists
|
||||
apt-cache policy cephadm # if download.ceph.com is missing, cache is stale
|
||||
apt-get update # manual refresh
|
||||
apt-cache policy cephadm # should now show 20.2.x candidate
|
||||
```
|
||||
|
||||
**Fix:** `tasks/prerequisites.yml` was patched to drop `cache_valid_time`
|
||||
on the apt-update task; refresh is unconditional after the repo is
|
||||
added. If you're seeing this on a node that ran the OLD prerequisites
|
||||
task (cached deploy state), `apt update` manually then re-run deploy.
|
||||
|
||||
---
|
||||
|
||||
### Symptom: `mise run deploy` fails — `'ceph_rgw_dns_name' is undefined`
|
||||
|
||||
**Cause:** the per-cluster `group_vars/all/vars.yml` is missing the
|
||||
`ceph_rgw_dns_name` declaration. RGW zonegroup creation needs it for
|
||||
the `--endpoints` and master zonegroup hostname.
|
||||
|
||||
**Fix:** add to `inventories/<cluster>/group_vars/all/vars.yml`:
|
||||
|
||||
```yaml
|
||||
ceph_rgw_dns_name: s3.{{ cluster_domain }}
|
||||
```
|
||||
|
||||
This derives the DNS name from `cluster_domain` (e.g.
|
||||
`s3.dev.hel.htz.futo.cloud`). Sietch defines this explicitly; painbox
|
||||
was missing it before its first deploy. Future clusters should include
|
||||
it from the start — see [docs/adding-a-cluster.md](adding-a-cluster.md)
|
||||
group_vars template.
|
||||
|
||||
---
|
||||
|
||||
### Symptom: `HEALTH_WARN: OSDMAP_FLAGS: noin flag(s) set` after deploy
|
||||
|
||||
**Cause:** the `noin` flag was set by an earlier failed run of the
|
||||
older imperative OSD-creation flow and never unset. The current
|
||||
spec-based flow doesn't set `noin` (cephadm rolls out OSDs gracefully)
|
||||
and includes a defensive unset task at the tail of `osds.yml`, but the
|
||||
flag can persist if the deploy never reached that tail (e.g., a failure
|
||||
in an earlier phase).
|
||||
|
||||
**Fix:** clear it manually, or just re-run `mise run deploy` — the
|
||||
defensive task at the end of `tasks/osds.yml` unsets `noin`
|
||||
unconditionally (idempotent no-op when already unset):
|
||||
|
||||
```bash
|
||||
ssh -i ~/.ssh/id_ed25519_<cluster> root@<bootstrap-ip> 'ceph osd unset noin'
|
||||
```
|
||||
|
||||
Validate:
|
||||
|
||||
```bash
|
||||
ceph osd dump | grep -E "^flags"
|
||||
# Should NOT contain 'noin'. Default healthy: sortbitwise,recovery_deletes,purged_snapdirs,pglog_hardlimit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Provisioning Issues
|
||||
|
||||
### Symptom: provision.yml fails with "REFUSING TO RUN"
|
||||
|
||||
**Cause:** Missing safety flag.
|
||||
|
||||
**Fix:**
|
||||
|
||||
```bash
|
||||
CEPH_ENV=<inventory> scripts/ansible-play.sh provision.yml \
|
||||
-e confirm_wipe=true
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Symptom: Provisioning fails at debootstrap / chroot phase
|
||||
|
||||
**Diagnose:** Check which task failed in the Ansible output. The rescue
|
||||
block automatically unmounts `/mnt`, so it's safe to re-run.
|
||||
|
||||
**Common causes:**
|
||||
|
||||
1. **apt sources unreachable from live image** -- check network
|
||||
connectivity from the live image. DNS resolution and internet access
|
||||
are required for debootstrap.
|
||||
|
||||
2. **Disk detection failed** -- SSD not found at expected path:
|
||||
|
||||
```bash
|
||||
ls /dev/disk/by-path/ | grep sas
|
||||
lsblk
|
||||
```
|
||||
|
||||
3. **Previous partial provision** -- the role is idempotent. If the
|
||||
provisioning marker exists at `/mnt/etc/ceph-provisioned.json`, all
|
||||
chroot phases are skipped. To force re-provision, boot into the live
|
||||
image and re-run.
|
||||
|
||||
---
|
||||
|
||||
### Symptom: Post-reboot SSH fails after provisioning
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
# Try with verbose SSH
|
||||
ssh -vvv -i ~/.ssh/id_ed25519_sietch ansible-iac@sietch-ceph-<name>
|
||||
```
|
||||
|
||||
**Common causes:**
|
||||
|
||||
1. **Node still booting** -- R730xd POST takes 60-90 seconds. Wait and
|
||||
retry.
|
||||
|
||||
2. **SSH host key changed** -- fresh provision generates new host keys:
|
||||
|
||||
```bash
|
||||
ssh-keygen -R sietch-ceph-<name>
|
||||
```
|
||||
|
||||
3. **Network not up** -- bond interface may not have configured. Check
|
||||
via iDRAC virtual console.
|
||||
|
||||
4. **Wrong IP** -- verify `bond_ip` in host_vars matches the actual
|
||||
network config.
|
||||
|
||||
---
|
||||
|
||||
## SSH Connectivity
|
||||
|
||||
### Symptom: Cannot SSH to cluster nodes from controller
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
# Test SSH to a node directly
|
||||
ssh ansible-iac@10.10.10.90 hostname
|
||||
|
||||
# If using a jump host, verify it's reachable (check your ~/.ssh/config)
|
||||
ssh <jump-host> hostname
|
||||
```
|
||||
|
||||
**Common causes:**
|
||||
|
||||
1. **SSH config issue** -- if nodes are behind a jump host, verify your
|
||||
`~/.ssh/config` has the correct ProxyJump or ProxyCommand settings.
|
||||
This is personal config, not managed by the repo.
|
||||
|
||||
2. **Wrong SSH key** -- inventory uses `~/.ssh/id_ed25519_sietch`:
|
||||
|
||||
```bash
|
||||
ls -la ~/.ssh/id_ed25519_sietch*
|
||||
```
|
||||
|
||||
3. **sntrup761 kex hang** -- cephadm's asyncssh does not support
|
||||
post-quantum key exchange. The baseline role deploys
|
||||
`/etc/ssh/sshd_config.d/no-sntrup.conf` to disable it. If missing:
|
||||
|
||||
```bash
|
||||
scripts/ansible-play.sh deploy-ceph.yml \
|
||||
--tags prerequisites --limit sietch-ceph-<name>,sietch-ceph-laurel
|
||||
```
|
||||
|
||||
4. **nftables blocking SSH** -- verify port 22 is allowed:
|
||||
|
||||
```bash
|
||||
# From the node (via iDRAC console if SSH is blocked)
|
||||
nft list ruleset | grep 22
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Quick Health Check Commands
|
||||
|
||||
```bash
|
||||
# Overall status
|
||||
ceph status
|
||||
|
||||
# OSD health
|
||||
ceph osd tree
|
||||
ceph osd df
|
||||
|
||||
# PG health
|
||||
ceph pg stat
|
||||
ceph pg dump_stuck
|
||||
|
||||
# Services
|
||||
ceph orch ls
|
||||
ceph orch ps
|
||||
|
||||
# Recent crashes
|
||||
ceph crash ls-new
|
||||
|
||||
# Drift from expected config
|
||||
mise run drift
|
||||
|
||||
# Cluster capacity
|
||||
ceph df
|
||||
```
|
||||
@@ -0,0 +1,342 @@
|
||||
---
|
||||
# Drift detection — compares expected Ansible state against live cluster.
|
||||
# Read-only. No changes. Reports mismatches.
|
||||
#
|
||||
# Usage:
|
||||
# scripts/ansible-play.sh drift.yml
|
||||
# mise run drift
|
||||
|
||||
- name: Drift detection
|
||||
hosts: ceph_nodes
|
||||
become: true
|
||||
gather_facts: false
|
||||
|
||||
vars:
|
||||
drift_results: []
|
||||
vars_files:
|
||||
- roles/baseline/defaults/main.yml
|
||||
- roles/os_tuning/defaults/main.yml
|
||||
- roles/hardware_tuning/defaults/main.yml
|
||||
- roles/ceph_tuning/defaults/main.yml
|
||||
- roles/security/defaults/main.yml
|
||||
|
||||
tasks:
|
||||
# --- Collect all checks ---
|
||||
|
||||
# sysctl
|
||||
- name: Check sysctl values
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
sysctl -n {{ item.key }} 2>/dev/null | tr -d ' '
|
||||
args:
|
||||
executable: /bin/bash
|
||||
loop:
|
||||
- key: vm.swappiness
|
||||
expected: "{{ ceph_sysctl_vm_swappiness }}"
|
||||
- key: vm.min_free_kbytes
|
||||
expected: "{{ ceph_sysctl_vm_min_free_kbytes }}"
|
||||
- key: vm.zone_reclaim_mode
|
||||
expected: "{{ ceph_sysctl_vm_zone_reclaim_mode }}"
|
||||
- key: fs.aio-max-nr
|
||||
expected: "{{ ceph_sysctl_fs_aio_max_nr }}"
|
||||
- key: kernel.pid_max
|
||||
expected: "{{ ceph_sysctl_kernel_pid_max }}"
|
||||
register: sysctl_checks
|
||||
changed_when: false
|
||||
loop_control:
|
||||
label: "{{ item.key }}"
|
||||
|
||||
- name: Record sysctl drift
|
||||
ansible.builtin.set_fact:
|
||||
drift_results: >-
|
||||
{{ drift_results + [{
|
||||
'category': 'sysctl',
|
||||
'item': item.item.key,
|
||||
'expected': item.item.expected | string,
|
||||
'actual': item.stdout | trim,
|
||||
'match': (item.stdout | trim) == (item.item.expected | string)
|
||||
}] }}
|
||||
loop: "{{ sysctl_checks.results }}"
|
||||
loop_control:
|
||||
label: "{{ item.item.key }}"
|
||||
|
||||
# I/O scheduler
|
||||
- name: Check HDD I/O scheduler
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
for dev in $(lsblk -dnpo NAME,ROTA,TYPE \
|
||||
| awk '$2==1 && $3=="disk" {print $1}' | head -1); do
|
||||
BDEV=$(basename "$dev")
|
||||
cat /sys/block/$BDEV/queue/scheduler \
|
||||
| grep -oP '\[\K[^\]]+'
|
||||
done
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: hdd_sched
|
||||
changed_when: false
|
||||
|
||||
- name: Record HDD scheduler drift
|
||||
ansible.builtin.set_fact:
|
||||
drift_results: >-
|
||||
{{ drift_results + [{
|
||||
'category': 'hardware',
|
||||
'item': 'HDD I/O scheduler',
|
||||
'expected': ceph_hdd_scheduler,
|
||||
'actual': hdd_sched.stdout | trim | default('none'),
|
||||
'match': (hdd_sched.stdout | trim | default('none'))
|
||||
== ceph_hdd_scheduler
|
||||
}] }}
|
||||
|
||||
- name: Check SSD I/O scheduler
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
for dev in $(lsblk -dnpo NAME,ROTA,TYPE \
|
||||
| awk '$2==0 && $3=="disk" {print $1}' | head -1); do
|
||||
BDEV=$(basename "$dev")
|
||||
cat /sys/block/$BDEV/queue/scheduler \
|
||||
| grep -oP '\[\K[^\]]+'
|
||||
done
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: ssd_sched
|
||||
changed_when: false
|
||||
|
||||
- name: Record SSD scheduler drift
|
||||
ansible.builtin.set_fact:
|
||||
drift_results: >-
|
||||
{{ drift_results + [{
|
||||
'category': 'hardware',
|
||||
'item': 'SSD I/O scheduler',
|
||||
'expected': ceph_ssd_scheduler,
|
||||
'actual': ssd_sched.stdout | trim | default('none'),
|
||||
'match': (ssd_sched.stdout | trim | default('none'))
|
||||
== ceph_ssd_scheduler
|
||||
}] }}
|
||||
|
||||
# nftables
|
||||
- name: Check nftables policy
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
nft list chain inet filter input 2>/dev/null \
|
||||
| grep -c 'policy drop' || echo 0
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: nft_policy
|
||||
changed_when: false
|
||||
|
||||
- name: Record firewall drift
|
||||
ansible.builtin.set_fact:
|
||||
drift_results: >-
|
||||
{{ drift_results + [{
|
||||
'category': 'security',
|
||||
'item': 'nftables policy drop',
|
||||
'expected': 'enabled',
|
||||
'actual': 'enabled' if (nft_policy.stdout | trim | int > 0)
|
||||
else 'missing',
|
||||
'match': (nft_policy.stdout | trim | int > 0)
|
||||
}] }}
|
||||
|
||||
# SSH hardening
|
||||
- name: Check PasswordAuthentication
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
sshd -T 2>/dev/null | grep -i passwordauthentication \
|
||||
| awk '{print $2}'
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: ssh_pwauth
|
||||
changed_when: false
|
||||
|
||||
- name: Record SSH drift
|
||||
ansible.builtin.set_fact:
|
||||
drift_results: >-
|
||||
{{ drift_results + [{
|
||||
'category': 'security',
|
||||
'item': 'SSH PasswordAuthentication',
|
||||
'expected': 'no',
|
||||
'actual': ssh_pwauth.stdout | trim,
|
||||
'match': (ssh_pwauth.stdout | trim) == 'no'
|
||||
}] }}
|
||||
|
||||
# sudo config
|
||||
- name: Check ops sudo config
|
||||
ansible.builtin.shell: |
|
||||
grep -c 'NOPASSWD' /etc/sudoers.d/ops 2>/dev/null || echo 0
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: ops_sudo
|
||||
changed_when: false
|
||||
|
||||
- name: Record sudo drift
|
||||
ansible.builtin.set_fact:
|
||||
drift_results: >-
|
||||
{{ drift_results + [{
|
||||
'category': 'security',
|
||||
'item': 'ops sudo requires password',
|
||||
'expected': 'yes',
|
||||
'actual': 'no (NOPASSWD)' if (ops_sudo.stdout | trim | int > 0)
|
||||
else 'yes',
|
||||
'match': (ops_sudo.stdout | trim | int == 0)
|
||||
}] }}
|
||||
|
||||
# --- Ceph cluster checks (bootstrap only) ---
|
||||
|
||||
- name: Check OSD status
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
ceph osd stat --format json 2>/dev/null | python3 -c "
|
||||
import sys, json
|
||||
d = json.load(sys.stdin)
|
||||
print(d.get('num_osds', 0), d.get('num_up_osds', 0))
|
||||
"
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: osd_stat
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Record OSD drift
|
||||
ansible.builtin.set_fact:
|
||||
drift_results: >-
|
||||
{{ drift_results + [{
|
||||
'category': 'ceph',
|
||||
'item': 'OSDs up/total',
|
||||
'expected': (osd_stat.stdout.split()[0]) ~ '/'
|
||||
~ (osd_stat.stdout.split()[0]),
|
||||
'actual': (osd_stat.stdout.split()[1]) ~ '/'
|
||||
~ (osd_stat.stdout.split()[0]),
|
||||
'match': osd_stat.stdout.split()[0]
|
||||
== osd_stat.stdout.split()[1]
|
||||
}] }}
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
|
||||
- name: Check MON quorum
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
ceph mon stat --format json 2>/dev/null | python3 -c "
|
||||
import sys, json
|
||||
print(json.load(sys.stdin).get('num_mons', 0))
|
||||
"
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: mon_count
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Record MON drift
|
||||
ansible.builtin.set_fact:
|
||||
drift_results: >-
|
||||
{{ drift_results + [{
|
||||
'category': 'ceph',
|
||||
'item': 'MON count',
|
||||
'expected': groups['ceph_nodes'] | length | string,
|
||||
'actual': mon_count.stdout | trim,
|
||||
'match': (mon_count.stdout | trim)
|
||||
== (groups['ceph_nodes'] | length | string)
|
||||
}] }}
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
|
||||
- name: Check RGW daemons
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
ceph orch ls --service-type rgw --format json 2>/dev/null \
|
||||
| python3 -c "
|
||||
import sys, json
|
||||
svcs = json.load(sys.stdin)
|
||||
print(sum(s.get('status',{}).get('running',0) for s in svcs))
|
||||
"
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: rgw_count
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Record RGW drift
|
||||
ansible.builtin.set_fact:
|
||||
drift_results: >-
|
||||
{{ drift_results + [{
|
||||
'category': 'ceph',
|
||||
'item': 'RGW daemons',
|
||||
'expected': groups['ceph_nodes'] | length | string,
|
||||
'actual': rgw_count.stdout | trim,
|
||||
'match': (rgw_count.stdout | trim)
|
||||
== (groups['ceph_nodes'] | length | string)
|
||||
}] }}
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
|
||||
- name: Check cluster health
|
||||
ansible.builtin.command: ceph health --format json
|
||||
register: health_check
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Record health drift
|
||||
ansible.builtin.set_fact:
|
||||
drift_results: >-
|
||||
{{ drift_results + [{
|
||||
'category': 'ceph',
|
||||
'item': 'cluster health',
|
||||
'expected': 'HEALTH_OK',
|
||||
'actual': (health_check.stdout | from_json).status,
|
||||
'match': (health_check.stdout | from_json).status
|
||||
== 'HEALTH_OK'
|
||||
}] }}
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
|
||||
# Ceph config values
|
||||
- name: Check Ceph config values
|
||||
ansible.builtin.command: "ceph config get osd {{ item.key }}"
|
||||
loop:
|
||||
- key: osd_recovery_max_active
|
||||
expected: "{{ ceph_osd_recovery_max_active }}"
|
||||
- key: osd_max_backfills
|
||||
expected: "{{ ceph_osd_max_backfills }}"
|
||||
- key: osd_scrub_begin_hour
|
||||
expected: "{{ ceph_osd_scrub_begin_hour }}"
|
||||
- key: osd_scrub_end_hour
|
||||
expected: "{{ ceph_osd_scrub_end_hour }}"
|
||||
register: ceph_config_checks
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
loop_control:
|
||||
label: "{{ item.key }}"
|
||||
|
||||
- name: Record Ceph config drift
|
||||
ansible.builtin.set_fact:
|
||||
drift_results: >-
|
||||
{{ drift_results + [{
|
||||
'category': 'ceph-config',
|
||||
'item': item.item.key,
|
||||
'expected': item.item.expected | string,
|
||||
'actual': item.stdout | trim,
|
||||
'match': (item.stdout | trim) == (item.item.expected | string)
|
||||
}] }}
|
||||
loop: "{{ ceph_config_checks.results | default([]) }}"
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
loop_control:
|
||||
label: "{{ item.item.key | default('skipped') }}"
|
||||
|
||||
# --- Report ---
|
||||
|
||||
- name: Build drift report
|
||||
ansible.builtin.set_fact:
|
||||
drift_report: |
|
||||
=== Drift Report: {{ inventory_hostname }} @ {{ lookup('pipe', 'date -u +%Y-%m-%dT%H:%M:%SZ') }} ===
|
||||
{% for r in drift_results %}
|
||||
{% if r.match %}
|
||||
{{ '%-14s' | format(r.category) }} {{ '%-30s' | format(r.item) }} expected: {{ '%-12s' | format(r.expected) }} actual: {{ '%-12s' | format(r.actual) }} OK
|
||||
{% else %}
|
||||
{{ '%-14s' | format(r.category) }} {{ '%-30s' | format(r.item) }} expected: {{ '%-12s' | format(r.expected) }} actual: {{ '%-12s' | format(r.actual) }} !! DRIFT
|
||||
{% endif %}
|
||||
{% endfor %}
|
||||
|
||||
{% set drifts = drift_results | selectattr('match', 'equalto', false) | list %}
|
||||
{% if drifts | length == 0 %}
|
||||
No drift detected. Cluster matches expected state.
|
||||
{% else %}
|
||||
{{ drifts | length }} drift(s) detected. Run 'mise run deploy' to converge.
|
||||
{% endif %}
|
||||
|
||||
- name: Display drift report
|
||||
ansible.builtin.debug:
|
||||
msg: "{{ drift_report.split('\n') }}"
|
||||
@@ -0,0 +1,19 @@
|
||||
---
|
||||
# Security hardening for Ceph nodes.
|
||||
# Deploys nftables firewall rules and SSH hardening.
|
||||
# Run AFTER deploy-ceph.yml — cephadm needs unrestricted access during deploy.
|
||||
#
|
||||
# WARNING: This locks down inbound traffic. Ensure ceph_firewall_trusted_networks
|
||||
# includes all subnets that need to reach Ceph services. SSH (port 22) is always
|
||||
# open from any source to prevent lockout.
|
||||
#
|
||||
# Usage:
|
||||
# scripts/ansible-play.sh harden.yml
|
||||
|
||||
- name: Security hardening for Ceph nodes
|
||||
hosts: ceph_nodes
|
||||
become: true
|
||||
gather_facts: false
|
||||
|
||||
roles:
|
||||
- security
|
||||
@@ -0,0 +1,85 @@
|
||||
---
|
||||
- name: Hardware inventory
|
||||
hosts: ceph_nodes
|
||||
gather_facts: true
|
||||
become: true
|
||||
tasks:
|
||||
- name: Gather dmidecode memory info
|
||||
ansible.builtin.command: dmidecode -t memory
|
||||
register: dmidecode_memory
|
||||
changed_when: false
|
||||
|
||||
- name: Gather lsblk disk info
|
||||
ansible.builtin.command: lsblk -d -J -o NAME,SIZE,MODEL,ROTA,TRAN,SERIAL
|
||||
register: lsblk_output
|
||||
changed_when: false
|
||||
|
||||
- name: Gather network interface info
|
||||
ansible.builtin.command: ip -j link show
|
||||
register: ip_link_output
|
||||
changed_when: false
|
||||
|
||||
- name: Parse system info
|
||||
ansible.builtin.set_fact:
|
||||
hw_info:
|
||||
hostname: "{{ ansible_hostname }}"
|
||||
ip: "{{ ansible_host }}"
|
||||
system:
|
||||
manufacturer: "{{ ansible_system_vendor }}"
|
||||
product: "{{ ansible_product_name }}"
|
||||
serial: "{{ ansible_product_serial }}"
|
||||
cpu:
|
||||
model: "{{ ansible_processor[2] }}"
|
||||
sockets: "{{ ansible_processor_count }}"
|
||||
cores_per_socket: "{{ ansible_processor_cores }}"
|
||||
threads_per_core: "{{ ansible_processor_threads_per_core }}"
|
||||
total_vcpus: "{{ ansible_processor_vcpus }}"
|
||||
memory:
|
||||
total_mb: "{{ ansible_memtotal_mb }}"
|
||||
dimms: "{{ dmidecode_memory.stdout }}"
|
||||
storage:
|
||||
disks: "{{ (lsblk_output.stdout | from_json).blockdevices }}"
|
||||
network:
|
||||
interfaces: "{{ ip_link_output.stdout | from_json }}"
|
||||
os:
|
||||
distribution: "{{ ansible_distribution }}"
|
||||
version: "{{ ansible_distribution_version }}"
|
||||
kernel: "{{ ansible_kernel }}"
|
||||
|
||||
- name: Ensure hardware/ output directory exists
|
||||
delegate_to: localhost
|
||||
become: false
|
||||
ansible.builtin.file:
|
||||
path: "{{ playbook_dir }}/hardware"
|
||||
state: directory
|
||||
mode: '0755'
|
||||
run_once: true # noqa: run-once[task]
|
||||
|
||||
- name: Write per-host JSON
|
||||
delegate_to: localhost
|
||||
become: false
|
||||
ansible.builtin.copy:
|
||||
content: "{{ hw_info | to_nice_json }}"
|
||||
dest: "{{ playbook_dir }}/hardware/{{ inventory_hostname }}.json"
|
||||
mode: '0644'
|
||||
|
||||
- name: Stash for aggregation
|
||||
ansible.builtin.set_fact:
|
||||
_collected: "{{ hw_info }}"
|
||||
|
||||
- name: Combine inventory into single file
|
||||
hosts: localhost
|
||||
gather_facts: false
|
||||
tasks:
|
||||
- name: Build combined inventory
|
||||
ansible.builtin.set_fact:
|
||||
combined: >-
|
||||
{{ combined | default({}) |
|
||||
combine({item: hostvars[item]['_collected']}) }}
|
||||
loop: "{{ groups['ceph_nodes'] }}"
|
||||
|
||||
- name: Write combined JSON
|
||||
ansible.builtin.copy:
|
||||
content: "{{ combined | to_nice_json }}"
|
||||
dest: "{{ playbook_dir }}/hardware/combined.json"
|
||||
mode: '0644'
|
||||
@@ -0,0 +1,71 @@
|
||||
---
|
||||
# === Naming ===
|
||||
cluster_name: painbox
|
||||
cluster_role: ceph
|
||||
|
||||
# === Network ===
|
||||
cluster_domain: dev.hel.htz.futo.cloud
|
||||
public_network: 157.180.105.192/26
|
||||
cluster_network: 157.180.105.192/26
|
||||
gateway: 157.180.105.193
|
||||
dns_server: 185.12.64.1
|
||||
timezone: UTC
|
||||
|
||||
# No bond — single NIC, direct SSH (no ProxyJump)
|
||||
|
||||
# === Ceph ===
|
||||
ceph_release: tentacle
|
||||
ceph_repo_url: "https://download.ceph.com/debian-{{ ceph_release }}/"
|
||||
ceph_repo_key_url: "https://download.ceph.com/keys/release.asc"
|
||||
|
||||
# === OS / Auth ===
|
||||
admin_user: root
|
||||
|
||||
# 1P vault for cluster secret lookups (e.g., rotate-ssh-key.yml reads pubkey from here).
|
||||
cluster_secrets_vault: yucca_tf_dev
|
||||
|
||||
# Secret aliases — vault_* vars populated by scripts/ansible-play.sh via op inject
|
||||
ops_password: "{{ vault_ops_password }}"
|
||||
ceph_dashboard_user: admin
|
||||
ceph_dashboard_password: "{{ vault_ceph_dashboard_password }}"
|
||||
|
||||
# S3 svc-user (yucca-restic consumer) — TF+1P-predetermined keys passed to
|
||||
# `radosgw-admin user create --access-key=... --secret-key=...` so the Yucca
|
||||
# app can be pre-configured with matching credentials.
|
||||
ceph_rgw_s3_user_access_key: "{{ vault_s3_restic_access_key }}"
|
||||
ceph_rgw_s3_user_secret_key: "{{ vault_s3_restic_secret_key }}"
|
||||
|
||||
# === Storage ===
|
||||
ssd_model_pattern: "SAMSUNG"
|
||||
ceph_db_lv_size: "128G"
|
||||
ceph_db_lvs_per_node: 14
|
||||
|
||||
# === RGW (Object Gateway) ===
|
||||
ceph_rgw_realm: painbox
|
||||
ceph_rgw_zonegroup: eu-central-1
|
||||
ceph_rgw_zonegroup_api_name: eu-central-1
|
||||
ceph_rgw_zone: dev-z1
|
||||
ceph_rgw_ec_profile: ec-k8m3-osd
|
||||
ceph_rgw_ec_k: 8
|
||||
ceph_rgw_ec_m: 3
|
||||
ceph_rgw_ec_failure_domain: osd
|
||||
ceph_rgw_ec_device_class: hdd
|
||||
ceph_rgw_data_pool: "{{ ceph_rgw_zone }}.rgw.buckets.data"
|
||||
ceph_rgw_index_pool: "{{ ceph_rgw_zone }}.rgw.buckets.index"
|
||||
ceph_rgw_extra_pool: "{{ ceph_rgw_zone }}.rgw.buckets.non-ec"
|
||||
ceph_rgw_dns_name: s3.{{ cluster_domain }}
|
||||
ceph_rgw_replicated_size: 2
|
||||
ceph_rgw_replicated_min_size: 1
|
||||
ceph_rgw_count_per_host: 1
|
||||
ceph_rgw_port: 7480
|
||||
ceph_rgw_ssl: false
|
||||
ceph_rgw_scheme: "{{ 'https' if ceph_rgw_ssl else 'http' }}"
|
||||
ceph_rgw_s3_user_uid: svc-yucca-restic
|
||||
ceph_rgw_s3_user_display_name: "yucca/restic service account"
|
||||
|
||||
# === Monitoring Stack ===
|
||||
ceph_prometheus_port: 9095
|
||||
ceph_grafana_port: 3000
|
||||
ceph_alertmanager_port: 9093
|
||||
ceph_grafana_admin_user: admin
|
||||
ceph_grafana_admin_password: "{{ vault_grafana_admin_password }}"
|
||||
@@ -0,0 +1,35 @@
|
||||
---
|
||||
# Hetzner SX295 OSD node template — painbox cluster
|
||||
#
|
||||
# Copy to host_vars/<hostname>.yml (full inventory_hostname).
|
||||
# Filename MUST match inventory_hostname for Ansible auto-load.
|
||||
# Values are node-specific — SATA controller addresses differ per server.
|
||||
#
|
||||
# Hardware: EPYC 7502P / 2x 7.68TB NVMe (RAID-1) / 14x 22TB SATA
|
||||
# NVMe pair is installimage RAID-1, partitioned as:
|
||||
# md0 = /boot, md1 = vg0 (root, var, swap, block.db LVs, SSD OSD)
|
||||
# SATA drives are left raw for Ceph HDD OSDs.
|
||||
hostname_short: painbox-ceph-EXAMPLE
|
||||
bond_ip: 157.180.105.X
|
||||
|
||||
# Block.db VG — on NVMe RAID-1 vg0 (created by installimage post-install)
|
||||
ceph_db_vg: vg0
|
||||
|
||||
# HDD OSD mappings: SATA by-path -> block.db LV
|
||||
# Find SATA controllers with: ls /dev/disk/by-path/ | grep ata
|
||||
# Each SATA port maps to one db-slot LV, assigned positionally.
|
||||
#
|
||||
# Typical SX295 SATA topology (3 controllers, 14 ports total):
|
||||
# 0000:45:00.0 (8 ports): ata-1..8
|
||||
# 0000:46:00.0 (2 ports): ata-1..2
|
||||
# 0000:87:00.0 (4 ports): ata-1..4
|
||||
ceph_hdd_osds:
|
||||
- path_phy: pci-0000:45:00.0-ata-1
|
||||
db: vg0/db-slot0
|
||||
- path_phy: pci-0000:45:00.0-ata-2
|
||||
db: vg0/db-slot1
|
||||
# ... one entry per HDD (14 total for full SX295)
|
||||
|
||||
# SSD OSD — remaining NVMe space after block.db LVs + reserve
|
||||
ceph_ssd_osds:
|
||||
- lv: vg0/ssd-osd
|
||||
@@ -0,0 +1,53 @@
|
||||
---
|
||||
# Painbox host_vars — the single Hetzner SX295 node.
|
||||
# Reprovisioned under this identity 2026-04-26 (was painbox-osd-5c3cac).
|
||||
# Reprovision procedure: docs/runbooks/painbox-reprovision.md.
|
||||
|
||||
hostname_short: painbox-ceph-evelyn
|
||||
bond_ip: 157.180.105.198
|
||||
|
||||
# Block.db VG — single NVMe RAID-1 VG created by installimage post-install
|
||||
ceph_db_vg: vg0
|
||||
|
||||
# HDD OSD mappings: SATA path -> block.db LV
|
||||
# by-path is slot-stable: replacing a drive in the same bay keeps the same path.
|
||||
# All 14 SATA ports populated with Seagate Exos X22 22TB (ST22000NM001E).
|
||||
#
|
||||
# SATA controller topology:
|
||||
# 0000:45:00.0 (8 ports): ata-1..8
|
||||
# 0000:46:00.0 (2 ports): ata-1..2
|
||||
# 0000:87:00.0 (4 ports): ata-1..4
|
||||
# Total: 14 ports -> 14 HDDs -> 14 db-slots (db-slot0..13)
|
||||
ceph_hdd_osds:
|
||||
- path_phy: pci-0000:45:00.0-ata-1 # ZX299K3G
|
||||
db: vg0/db-slot0
|
||||
- path_phy: pci-0000:45:00.0-ata-2 # ZX297VJE
|
||||
db: vg0/db-slot1
|
||||
- path_phy: pci-0000:45:00.0-ata-3 # ZX2992CN
|
||||
db: vg0/db-slot2
|
||||
- path_phy: pci-0000:45:00.0-ata-4 # ZX295NSB
|
||||
db: vg0/db-slot3
|
||||
- path_phy: pci-0000:45:00.0-ata-5 # ZX29DE1R
|
||||
db: vg0/db-slot4
|
||||
- path_phy: pci-0000:45:00.0-ata-6 # ZX299AGR
|
||||
db: vg0/db-slot5
|
||||
- path_phy: pci-0000:45:00.0-ata-7 # ZX298TEQ
|
||||
db: vg0/db-slot6
|
||||
- path_phy: pci-0000:45:00.0-ata-8 # ZX29DE34
|
||||
db: vg0/db-slot7
|
||||
- path_phy: pci-0000:46:00.0-ata-1 # ZX294A8Z
|
||||
db: vg0/db-slot8
|
||||
- path_phy: pci-0000:46:00.0-ata-2 # ZX299EHE
|
||||
db: vg0/db-slot9
|
||||
- path_phy: pci-0000:87:00.0-ata-1 # ZX29DE21
|
||||
db: vg0/db-slot10
|
||||
- path_phy: pci-0000:87:00.0-ata-2 # ZX294PJ3
|
||||
db: vg0/db-slot11
|
||||
- path_phy: pci-0000:87:00.0-ata-3 # ZX29DDD1
|
||||
db: vg0/db-slot12
|
||||
- path_phy: pci-0000:87:00.0-ata-4 # ZX297AZ8
|
||||
db: vg0/db-slot13
|
||||
|
||||
# SSD OSD — remaining NVMe RAID-1 space after block.db LVs + reserve
|
||||
ceph_ssd_osds:
|
||||
- lv: vg0/ssd-osd
|
||||
@@ -0,0 +1,32 @@
|
||||
## Hetzner installimage autosetup for SX295 Ceph node
|
||||
## Cluster: painbox-ceph.dev.hel.htz.futo.cloud (reprovision target)
|
||||
## Hardware: EPYC 7502P / 2x 7.68TB NVMe / 14x 22TB SATA
|
||||
##
|
||||
## Usage (from rescue):
|
||||
## cat > /autosetup < autosetup
|
||||
## installimage -a -c /autosetup -x /tmp/post-install.sh
|
||||
##
|
||||
## IMPORTANT: Only NVMe drives are listed. SATA drives are untouched.
|
||||
## Ceph: Tentacle (v20) — EOL 2027-11-18
|
||||
## BOOT MODE: BIOS (verified 2026-03-27). If UEFI, add before /boot:
|
||||
## PART /boot/efi esp 256M
|
||||
|
||||
DRIVE1 /dev/nvme0n1
|
||||
DRIVE2 /dev/nvme1n1
|
||||
|
||||
SWRAID 1
|
||||
SWRAIDLEVEL 1
|
||||
|
||||
BOOTLOADER grub
|
||||
|
||||
HOSTNAME painbox-ceph-evelyn.dev.hel.htz.futo.cloud
|
||||
|
||||
PART /boot ext4 1G
|
||||
PART lvm vg0 all
|
||||
|
||||
LV vg0 swap swap swap 32G
|
||||
LV vg0 root / ext4 100G
|
||||
LV vg0 var /var ext4 200G
|
||||
LV vg0 varlog /var/log ext4 50G
|
||||
|
||||
IMAGE /root/.oldroot/nfs/install/../images/Debian-bookworm-latest-amd64-base.tar.zst
|
||||
@@ -0,0 +1,90 @@
|
||||
#!/bin/bash
|
||||
# Post-install script for SX295 Ceph node (painbox)
|
||||
# Runs in chroot after Hetzner installimage completes.
|
||||
#
|
||||
# Sets up: LVM for block.db + SSD OSD, SSH keys, cephadm prereqs
|
||||
# Does NOT: touch SATA drives, deploy Ceph (that's Ansible's job)
|
||||
#
|
||||
# RENDER BEFORE USE:
|
||||
# op inject -f -i post-install.sh.tpl -o /tmp/post-install.sh.rendered
|
||||
# Then upload the rendered file to the rescue system. See
|
||||
# docs/runbooks/painbox-reprovision.md.
|
||||
#
|
||||
# The rendered script authorizes only the ansible-iac pubkey fetched live
|
||||
# from 1P — rotations propagate at reprovision time without a template edit.
|
||||
|
||||
set -euo pipefail
|
||||
set -x
|
||||
|
||||
echo "=== post-install: starting ==="
|
||||
|
||||
# --- SSH authorized keys ---
|
||||
mkdir -p /root/.ssh
|
||||
chmod 700 /root/.ssh
|
||||
cat >> /root/.ssh/authorized_keys << 'SSHKEYS'
|
||||
op://yucca_tf_dev/PAINBOX_CEPH_ANSIBLE_IAC_SSH_KEY/public_key
|
||||
SSHKEYS
|
||||
chmod 600 /root/.ssh/authorized_keys
|
||||
|
||||
# --- System packages ---
|
||||
apt-get update -qq
|
||||
apt-get install -y -qq \
|
||||
lvm2 \
|
||||
python3 \
|
||||
python3-pip \
|
||||
curl \
|
||||
jq \
|
||||
smartmontools \
|
||||
chrony \
|
||||
podman \
|
||||
bzip2
|
||||
|
||||
# --- NVMe RAID-1 LVM for Ceph ---
|
||||
# After installimage, the NVMe RAID-1 has:
|
||||
# md0 = /boot (1G)
|
||||
# md1 = LVM PV for vg0 (rest of NVMe)
|
||||
# vg0 = swap(32G) + root(100G) + var(200G) + varlog(50G)
|
||||
#
|
||||
# Remaining space in vg0: ~7.28 TiB - 382 GiB = ~5.53 TiB free
|
||||
# Allocate: 14 x 128 GiB block.db LVs + leave 0.5 TiB reserve + rest for SSD OSD
|
||||
|
||||
echo "=== post-install: creating block.db LVs ==="
|
||||
|
||||
if ! vgs vg0 &>/dev/null; then
|
||||
echo "FATAL: vg0 not found. Check installimage partition config."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
for i in $(seq 0 13); do
|
||||
lvcreate --yes -Wy -L 128G -n db-slot${i} vg0
|
||||
echo " created vg0/db-slot${i} (128 GiB)"
|
||||
done
|
||||
|
||||
# SSD OSD: allocate all remaining space minus 512 GiB reserve
|
||||
VG_FREE_GIB=$(vgs vg0 --noheadings --nosuffix --units g -o vg_free | tr -d ' ' | cut -d. -f1)
|
||||
RESERVE_GIB=512
|
||||
OSD_GIB=$((VG_FREE_GIB - RESERVE_GIB))
|
||||
|
||||
if [ "$OSD_GIB" -gt 0 ]; then
|
||||
lvcreate --yes -Wy -L "${OSD_GIB}G" -n ssd-osd vg0
|
||||
echo " created vg0/ssd-osd (${OSD_GIB} GiB, ${RESERVE_GIB} GiB reserved)"
|
||||
else
|
||||
echo " WARNING: not enough space for SSD OSD LV"
|
||||
fi
|
||||
|
||||
echo "=== post-install: LVM layout ==="
|
||||
lvs vg0 -o lv_name,lv_size --noheadings
|
||||
|
||||
# --- Timezone ---
|
||||
ln -sf /usr/share/zoneinfo/UTC /etc/localtime
|
||||
|
||||
# --- Sysctl tuning for Ceph ---
|
||||
cat > /etc/sysctl.d/90-ceph.conf << 'SYSCTL'
|
||||
# Ceph OSD tuning
|
||||
vm.min_free_kbytes = 1048576
|
||||
vm.swappiness = 10
|
||||
net.core.rmem_max = 67108864
|
||||
net.core.wmem_max = 67108864
|
||||
SYSCTL
|
||||
|
||||
echo "=== post-install: complete ==="
|
||||
@@ -0,0 +1,111 @@
|
||||
---
|
||||
# === Naming ===
|
||||
cluster_name: sietch
|
||||
cluster_role: ceph
|
||||
|
||||
# === Network ===
|
||||
cluster_domain: dev.austin.int.futo.cloud
|
||||
public_network: 10.10.10.0/24
|
||||
cluster_network: 10.10.10.0/24
|
||||
gateway: 10.10.10.1
|
||||
dns_server: 10.10.10.1
|
||||
bond_mode: active-backup
|
||||
bond_interfaces:
|
||||
- eno1
|
||||
- eno2
|
||||
networkd_enabled: true
|
||||
|
||||
# === Ceph ===
|
||||
ceph_release: tentacle
|
||||
ceph_repo_url: "https://download.ceph.com/debian-{{ ceph_release }}/"
|
||||
ceph_repo_key_url: "https://download.ceph.com/keys/release.asc"
|
||||
|
||||
# === OS Provisioning ===
|
||||
admin_user: ansible-iac
|
||||
timezone: UTC
|
||||
|
||||
# Deploy keypair path on the controller. Both {path} and {path}.pub must
|
||||
# exist before running provision.yml (the preflight play in provision.yml
|
||||
# enforces this). ansible-iac's authorized_keys on every node is populated
|
||||
# by a file lookup at provision time — rotating the key on the controller
|
||||
# automatically propagates on the next re-provision.
|
||||
provision_iac_ssh_key_path: "~/.ssh/id_ed25519_sietch"
|
||||
|
||||
# 1P vault for cluster secret lookups (e.g., rotate-ssh-key.yml reads pubkey from here).
|
||||
cluster_secrets_vault: yucca_tf_dev
|
||||
|
||||
# Secret aliases — vault_* vars populated by scripts/ansible-play.sh via op inject
|
||||
#
|
||||
# Note: the deploy public key is NOT a secret — it's public data and lives
|
||||
# at {{ provision_iac_ssh_key_path }}.pub on the controller. Reading it via
|
||||
# file lookup avoids drift between vault and disk.
|
||||
ops_password: "{{ vault_ops_password }}"
|
||||
ceph_dashboard_user: admin
|
||||
ceph_dashboard_password: "{{ vault_ceph_dashboard_password }}"
|
||||
|
||||
# S3 svc-user (yucca-restic consumer) — TF+1P-predetermined keys passed to
|
||||
# `radosgw-admin user create --access-key=... --secret-key=...` so the Yucca
|
||||
# app can be pre-configured with matching credentials. Rotation path documented
|
||||
# in docs/runbooks/rotate-secrets.md.
|
||||
ceph_rgw_s3_user_access_key: "{{ vault_s3_restic_access_key }}"
|
||||
ceph_rgw_s3_user_secret_key: "{{ vault_s3_restic_secret_key }}"
|
||||
|
||||
# === Storage ===
|
||||
ssd_model_pattern: "Micron_5100"
|
||||
|
||||
os_partitions:
|
||||
esp_size: "512M"
|
||||
boot_size: "1G"
|
||||
root_size: "80G"
|
||||
swap_size: "8G"
|
||||
ceph_db_size: "1440G"
|
||||
|
||||
ceph_db_lv_size: "240G"
|
||||
ceph_db_lvs_per_ssd: 6
|
||||
|
||||
# === RGW (Object Gateway) ===
|
||||
ceph_rgw_realm: sietch
|
||||
ceph_rgw_zonegroup: us-east-1
|
||||
ceph_rgw_zonegroup_api_name: us-east-1
|
||||
ceph_rgw_zone: dev-z1
|
||||
ceph_rgw_ec_profile: ec-k8m3-osd
|
||||
ceph_rgw_ec_k: 8
|
||||
ceph_rgw_ec_m: 3
|
||||
ceph_rgw_ec_failure_domain: osd
|
||||
ceph_rgw_ec_device_class: hdd
|
||||
ceph_rgw_data_pool: "{{ ceph_rgw_zone }}.rgw.buckets.data"
|
||||
ceph_rgw_index_pool: "{{ ceph_rgw_zone }}.rgw.buckets.index"
|
||||
ceph_rgw_extra_pool: "{{ ceph_rgw_zone }}.rgw.buckets.non-ec"
|
||||
ceph_rgw_replicated_size: 2
|
||||
ceph_rgw_replicated_min_size: 1
|
||||
ceph_rgw_port: 443
|
||||
ceph_rgw_s3_user_uid: svc-yucca-restic
|
||||
ceph_rgw_s3_user_display_name: "yucca/restic service account"
|
||||
|
||||
# --- RGW DNS + TLS ---
|
||||
# Virtual-hosted S3 support: setting rgw_dns_name tells RGW to strip this
|
||||
# suffix from the Host header and treat the remainder as the bucket name.
|
||||
# Requires matching DNS: both s3.<domain> and *.s3.<domain> should resolve
|
||||
# to the cluster nodes (round-robin A, or VIP/LB in prod).
|
||||
ceph_rgw_dns_name: s3.dev.austin.int.futo.cloud
|
||||
|
||||
# Self-signed wildcard cert handed to cephadm via service spec.
|
||||
# cephadm distributes to all RGW daemons. To rotate, delete
|
||||
# /etc/ceph/rgw-ssl.{crt,key} on the bootstrap node and re-run the role.
|
||||
ceph_rgw_ssl: true
|
||||
ceph_rgw_ssl_cert_days: 3650 # 10 years
|
||||
ceph_rgw_ssl_cert_subject_c: US
|
||||
ceph_rgw_ssl_cert_subject_st: Texas
|
||||
ceph_rgw_ssl_cert_subject_l: Austin
|
||||
ceph_rgw_ssl_cert_subject_o: FUTO
|
||||
ceph_rgw_ssl_cert_email: yucca@futo.org
|
||||
|
||||
# Computed: scheme used for endpoints, debug output, and zonegroup/zone URLs
|
||||
ceph_rgw_scheme: "{{ 'https' if ceph_rgw_ssl else 'http' }}"
|
||||
|
||||
# === Monitoring Stack ===
|
||||
ceph_prometheus_port: 9095
|
||||
ceph_grafana_port: 3000
|
||||
ceph_alertmanager_port: 9093
|
||||
ceph_grafana_admin_user: admin
|
||||
ceph_grafana_admin_password: "{{ vault_grafana_admin_password }}"
|
||||
@@ -0,0 +1,36 @@
|
||||
---
|
||||
# Copy to host_vars/sietch-ceph-<name>.yml (full inventory_hostname).
|
||||
# The <name> segment is either operator-declared in clusters.auto.tfvars
|
||||
# or TF-auto-picked from wordlist.txt — see docs/naming.md.
|
||||
# Filename MUST match inventory_hostname for Ansible auto-load.
|
||||
# Values are node-specific — hardware paths differ per chassis.
|
||||
hostname_short: sietch-ceph-EXAMPLE
|
||||
bond_ip: 10.0.0.1
|
||||
|
||||
# SAS expander base path (unique per chassis)
|
||||
# Find with: ls /dev/disk/by-path/ | grep sas
|
||||
sas_path_prefix: "pci-0000:02:00.0-sas-exp0x500056b3XXXXXXXX"
|
||||
|
||||
# SSD PHY positions in the SAS topology
|
||||
ssd1_phy: 12
|
||||
ssd2_phy: 13
|
||||
|
||||
# LVM volume group names for block.db (on SSD partition 5)
|
||||
ceph_db_vg1: ceph-db-ssd1
|
||||
ceph_db_vg2: ceph-db-ssd2
|
||||
|
||||
# HDD OSD mappings: SAS PHY slot -> block.db LV
|
||||
# 6 HDDs per SSD, each gets a dedicated 240G block.db LV
|
||||
ceph_hdd_osds:
|
||||
- path_phy: phy0
|
||||
db: ceph-db-ssd1/db-slot0
|
||||
- path_phy: phy1
|
||||
db: ceph-db-ssd1/db-slot1
|
||||
# ... one entry per HDD
|
||||
|
||||
# SSD OSD partitions (partition 6 on each SSD, no separate block.db)
|
||||
ceph_ssd_osds:
|
||||
- path_phy: phy12
|
||||
partition: 6
|
||||
- path_phy: phy13
|
||||
partition: 6
|
||||
@@ -0,0 +1,52 @@
|
||||
---
|
||||
hostname_short: sietch-ceph-laurel
|
||||
bond_ip: 10.10.10.90
|
||||
|
||||
# SAS expander base path (unique per chassis — different backplane address per node)
|
||||
sas_path_prefix: "pci-0000:02:00.0-sas-exp0x500056b3fcf498ff"
|
||||
|
||||
# SSD PHY positions (rear bays)
|
||||
ssd1_phy: 12 # serial 17321A07BA4A
|
||||
ssd2_phy: 13 # serial 17321A07CFEE
|
||||
|
||||
# LVM VGs on SSD partition 5
|
||||
ceph_db_vg1: ceph-db-rear12 # VG on SSD1 (phy12) partition 5
|
||||
ceph_db_vg2: ceph-db-rear13 # VG on SSD2 (phy13) partition 5
|
||||
|
||||
# HDD OSD mappings: PHY slot -> block.db LV
|
||||
# PHY 0-5 -> SSD1 (ceph-db-rear12/db-slot0..5)
|
||||
# PHY 6-11 -> SSD2 (ceph-db-rear13/db-slot6..11)
|
||||
# by-path is slot-stable: replacing a drive in the same bay keeps the same path
|
||||
# 12 HDDs — all front bays populated
|
||||
ceph_hdd_osds:
|
||||
- path_phy: phy0
|
||||
db: ceph-db-rear12/db-slot0
|
||||
- path_phy: phy1
|
||||
db: ceph-db-rear12/db-slot1
|
||||
- path_phy: phy2
|
||||
db: ceph-db-rear12/db-slot2
|
||||
- path_phy: phy3
|
||||
db: ceph-db-rear12/db-slot3 # serial Z4D09B99 (ST6000NKCLAR6000, added 2026-04-10)
|
||||
- path_phy: phy4
|
||||
db: ceph-db-rear12/db-slot4
|
||||
- path_phy: phy5
|
||||
db: ceph-db-rear12/db-slot5 # serial Z4D0G7VC (ST6000NKCLAR6000, added 2026-04-10)
|
||||
- path_phy: phy6
|
||||
db: ceph-db-rear13/db-slot6
|
||||
- path_phy: phy7
|
||||
db: ceph-db-rear13/db-slot7 # serial Z4D0G7SC (ST6000NKCLAR6000, added 2026-04-10)
|
||||
- path_phy: phy8
|
||||
db: ceph-db-rear13/db-slot8
|
||||
- path_phy: phy9
|
||||
db: ceph-db-rear13/db-slot9
|
||||
- path_phy: phy10
|
||||
db: ceph-db-rear13/db-slot10 # serial Z4D0G7XK (ST6000NKCLAR6000, added 2026-04-10)
|
||||
- path_phy: phy11
|
||||
db: ceph-db-rear13/db-slot11
|
||||
|
||||
# SSD OSD partitions (partition 6 on each rear SSD, no separate block.db)
|
||||
ceph_ssd_osds:
|
||||
- path_phy: phy12
|
||||
partition: 6
|
||||
- path_phy: phy13
|
||||
partition: 6
|
||||
@@ -0,0 +1,52 @@
|
||||
---
|
||||
hostname_short: sietch-ceph-lawson
|
||||
bond_ip: 10.10.10.91
|
||||
|
||||
# SAS expander base path (unique per chassis — different backplane address per node)
|
||||
sas_path_prefix: "pci-0000:02:00.0-sas-exp0x500056b35be6b4ff"
|
||||
|
||||
# SSD PHY positions (rear bays)
|
||||
ssd1_phy: 12 # serial 17321A07CF91
|
||||
ssd2_phy: 13 # serial 17251B44C4D8
|
||||
|
||||
# LVM VGs on SSD partition 5
|
||||
ceph_db_vg1: ceph-db-rear12 # VG on SSD1 (phy12) partition 5
|
||||
ceph_db_vg2: ceph-db-rear13 # VG on SSD2 (phy13) partition 5
|
||||
|
||||
# HDD OSD mappings: PHY slot -> block.db LV
|
||||
# PHY 0-5 -> SSD1 (ceph-db-rear12/db-slot0..5)
|
||||
# PHY 6-11 -> SSD2 (ceph-db-rear13/db-slot6..11)
|
||||
# by-path is slot-stable: replacing a drive in the same bay keeps the same path
|
||||
# 12 HDDs — all front bays populated
|
||||
ceph_hdd_osds:
|
||||
- path_phy: phy0
|
||||
db: ceph-db-rear12/db-slot0
|
||||
- path_phy: phy1
|
||||
db: ceph-db-rear12/db-slot1
|
||||
- path_phy: phy2
|
||||
db: ceph-db-rear12/db-slot2
|
||||
- path_phy: phy3
|
||||
db: ceph-db-rear12/db-slot3
|
||||
- path_phy: phy4
|
||||
db: ceph-db-rear12/db-slot4
|
||||
- path_phy: phy5
|
||||
db: ceph-db-rear12/db-slot5
|
||||
- path_phy: phy6
|
||||
db: ceph-db-rear13/db-slot6
|
||||
- path_phy: phy7
|
||||
db: ceph-db-rear13/db-slot7 # serial Z4D0GKJX (ST6000NKCLAR6000, added 2026-04-10)
|
||||
- path_phy: phy8
|
||||
db: ceph-db-rear13/db-slot8
|
||||
- path_phy: phy9
|
||||
db: ceph-db-rear13/db-slot9
|
||||
- path_phy: phy10
|
||||
db: ceph-db-rear13/db-slot10
|
||||
- path_phy: phy11
|
||||
db: ceph-db-rear13/db-slot11
|
||||
|
||||
# SSD OSD partitions (partition 6 on each rear SSD, no separate block.db)
|
||||
ceph_ssd_osds:
|
||||
- path_phy: phy12
|
||||
partition: 6
|
||||
- path_phy: phy13
|
||||
partition: 6
|
||||
@@ -0,0 +1,52 @@
|
||||
---
|
||||
hostname_short: sietch-ceph-samara
|
||||
bond_ip: 10.10.10.92
|
||||
|
||||
# SAS expander base path (unique per chassis — different backplane address per node)
|
||||
sas_path_prefix: "pci-0000:02:00.0-sas-exp0x500056b3393ba1ff"
|
||||
|
||||
# SSD PHY positions (rear bays)
|
||||
ssd1_phy: 12 # serial 17321A07D1D3
|
||||
ssd2_phy: 13 # serial 17251B44DF51
|
||||
|
||||
# LVM VGs on SSD partition 5
|
||||
ceph_db_vg1: ceph-db-rear12 # VG on SSD1 (phy12) partition 5
|
||||
ceph_db_vg2: ceph-db-rear13 # VG on SSD2 (phy13) partition 5
|
||||
|
||||
# HDD OSD mappings: PHY slot -> block.db LV
|
||||
# PHY 0-5 -> SSD1 (ceph-db-rear12/db-slot0..5)
|
||||
# PHY 6-11 -> SSD2 (ceph-db-rear13/db-slot6..11)
|
||||
# by-path is slot-stable: replacing a drive in the same bay keeps the same path
|
||||
# 12 HDDs — all front bays populated
|
||||
ceph_hdd_osds:
|
||||
- path_phy: phy0
|
||||
db: ceph-db-rear12/db-slot0
|
||||
- path_phy: phy1
|
||||
db: ceph-db-rear12/db-slot1
|
||||
- path_phy: phy2
|
||||
db: ceph-db-rear12/db-slot2
|
||||
- path_phy: phy3
|
||||
db: ceph-db-rear12/db-slot3
|
||||
- path_phy: phy4
|
||||
db: ceph-db-rear12/db-slot4 # serial Z4D0GKB2 (replacement, added 2026-04-11)
|
||||
- path_phy: phy5
|
||||
db: ceph-db-rear12/db-slot5
|
||||
- path_phy: phy6
|
||||
db: ceph-db-rear13/db-slot6
|
||||
- path_phy: phy7
|
||||
db: ceph-db-rear13/db-slot7
|
||||
- path_phy: phy8
|
||||
db: ceph-db-rear13/db-slot8
|
||||
- path_phy: phy9
|
||||
db: ceph-db-rear13/db-slot9
|
||||
- path_phy: phy10
|
||||
db: ceph-db-rear13/db-slot10
|
||||
- path_phy: phy11
|
||||
db: ceph-db-rear13/db-slot11
|
||||
|
||||
# SSD OSD partitions (partition 6 on each rear SSD, no separate block.db)
|
||||
ceph_ssd_osds:
|
||||
- path_phy: phy12
|
||||
partition: 6
|
||||
- path_phy: phy13
|
||||
partition: 6
|
||||
@@ -0,0 +1,62 @@
|
||||
---
|
||||
# Migrate ifupdown -> systemd-networkd on sietch Ceph nodes.
|
||||
#
|
||||
# Live migration — no reboot, no Ceph disruption:
|
||||
# 1. Deploys networkd configs
|
||||
# 2. Starts networkd (adopts existing bond0)
|
||||
# 3. Verifies bond0 IP, members, gateway
|
||||
# 4. Commits: disables ifupdown, enables networkd for boot
|
||||
#
|
||||
# If verification fails, the play stops before committing.
|
||||
# networkd can be stopped and ifupdown still owns next boot.
|
||||
#
|
||||
# Safety:
|
||||
# - serial: 1 -- one node at a time
|
||||
# - /etc/network/interfaces preserved on disk
|
||||
# - Rollback: /usr/local/sbin/rollback-networkd.sh
|
||||
# - iDRAC at 10.10.11.9{0,1,2} for emergency recovery
|
||||
#
|
||||
# Usage (samara first):
|
||||
# scripts/ansible-play.sh migrate-networkd.yml --limit sietch-ceph-samara
|
||||
#
|
||||
# Usage (remaining nodes):
|
||||
# scripts/ansible-play.sh migrate-networkd.yml --limit 'ceph_nodes:!sietch-ceph-samara'
|
||||
#
|
||||
# Dry run:
|
||||
# scripts/ansible-play.sh migrate-networkd.yml --limit sietch-ceph-samara --check --diff
|
||||
|
||||
- name: Migrate to systemd-networkd
|
||||
hosts: ceph_nodes
|
||||
become: true
|
||||
gather_facts: false
|
||||
serial: 1
|
||||
|
||||
pre_tasks:
|
||||
- name: Preflight -- check Ceph health
|
||||
ansible.builtin.command: ceph health --format json
|
||||
register: ceph_health_pre
|
||||
changed_when: false
|
||||
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
|
||||
|
||||
- name: Assert cluster is healthy
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- (ceph_health_pre.stdout | from_json).status in ['HEALTH_OK', 'HEALTH_WARN']
|
||||
fail_msg: >-
|
||||
Ceph is {{ (ceph_health_pre.stdout | from_json).status }}.
|
||||
Fix cluster health before migrating.
|
||||
success_msg: "Ceph: {{ (ceph_health_pre.stdout | from_json).status }}"
|
||||
|
||||
roles:
|
||||
- role: networkd
|
||||
|
||||
post_tasks:
|
||||
- name: Confirm Ceph health after migration
|
||||
ansible.builtin.command: ceph health --format json
|
||||
register: ceph_health_post
|
||||
changed_when: false
|
||||
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
|
||||
|
||||
- name: Report Ceph health
|
||||
ansible.builtin.debug:
|
||||
msg: "Ceph: {{ (ceph_health_post.stdout | from_json).status }}"
|
||||
@@ -0,0 +1,96 @@
|
||||
---
|
||||
# post-deploy-capture.yml — snapshot bootstrap-side secrets into 1Password.
|
||||
#
|
||||
# Captures:
|
||||
# /etc/ceph/rgw-ssl.crt -> <CLUSTER>_CEPH_RGW_TLS_CERT
|
||||
# /etc/ceph/rgw-ssl.key -> <CLUSTER>_CEPH_RGW_TLS_KEY
|
||||
# /etc/ceph/ceph.client.admin.keyring -> <CLUSTER>_CEPH_CLIENT_ADMIN_KEYRING
|
||||
#
|
||||
# Idempotent: creates item if missing, updates if content drifted.
|
||||
#
|
||||
# Belt-and-suspenders only. The cluster continues to run from on-node
|
||||
# copies; the 1P copies exist for disaster recovery (e.g., laurel dies +
|
||||
# filesystem loss + you want to restore the RGW cert to a new bootstrap
|
||||
# node without re-trusting it in every S3 client).
|
||||
#
|
||||
# Usage:
|
||||
# scripts/ansible-play.sh post-deploy-capture.yml
|
||||
#
|
||||
# Requires: superuser SA read access from controller (already required by
|
||||
# mise run tf:* tasks), op CLI on controller, 1P session live.
|
||||
|
||||
- name: Capture bootstrap-side secrets to 1Password
|
||||
hosts: ceph_bootstrap
|
||||
become: true
|
||||
gather_facts: false
|
||||
strategy: linear
|
||||
|
||||
vars:
|
||||
capture_items:
|
||||
- file: /etc/ceph/rgw-ssl.crt
|
||||
title_suffix: RGW_TLS_CERT
|
||||
- file: /etc/ceph/rgw-ssl.key
|
||||
title_suffix: RGW_TLS_KEY
|
||||
- file: /etc/ceph/ceph.client.admin.keyring
|
||||
title_suffix: CLIENT_ADMIN_KEYRING
|
||||
|
||||
tasks:
|
||||
- name: Read bootstrap files
|
||||
ansible.builtin.slurp:
|
||||
src: "{{ item.file }}"
|
||||
loop: "{{ capture_items }}"
|
||||
register: slurped
|
||||
no_log: true
|
||||
|
||||
- name: Fetch superuser SA token from 1P (localhost) # noqa: run-once[task]
|
||||
ansible.builtin.command:
|
||||
argv:
|
||||
- op
|
||||
- read
|
||||
- "op://yucca_tf_dev/yucca_futo_1pass_superuser_service_account/password"
|
||||
register: su_token
|
||||
delegate_to: localhost
|
||||
become: false # run as operator on controller — root has no op session
|
||||
run_once: true
|
||||
changed_when: false
|
||||
check_mode: false
|
||||
no_log: true
|
||||
|
||||
# exec 0<&- closes stdin. op CLI treats non-TTY piped stdin as a JSON
|
||||
# item template; without this redirect op fails with "invalid JSON in
|
||||
# piped input" because ansible pipes an empty stdin to the shell.
|
||||
- name: Upsert each captured file to 1Password # noqa: run-once[task]
|
||||
ansible.builtin.shell: |
|
||||
set -euo pipefail
|
||||
exec 0<&-
|
||||
TITLE="{{ cluster_name | upper }}_CEPH_{{ item.item.title_suffix }}"
|
||||
VALUE="$CONTENT"
|
||||
if op item get "$TITLE" --vault yucca_tf_dev >/dev/null 2>&1; then
|
||||
op item edit "$TITLE" --vault yucca_tf_dev "password=$VALUE" >/dev/null
|
||||
echo "updated: $TITLE"
|
||||
else
|
||||
op item create --vault yucca_tf_dev --category password --title "$TITLE" "password=$VALUE" >/dev/null
|
||||
echo "created: $TITLE"
|
||||
fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
environment:
|
||||
OP_SERVICE_ACCOUNT_TOKEN: "{{ su_token.stdout | trim }}"
|
||||
CONTENT: "{{ item.content | b64decode }}"
|
||||
loop: "{{ slurped.results }}"
|
||||
loop_control:
|
||||
label: "{{ item.item.title_suffix }}"
|
||||
delegate_to: localhost
|
||||
become: false # run as operator on controller — op session is operator-owned
|
||||
run_once: true
|
||||
register: capture_result
|
||||
changed_when: "'created:' in capture_result.stdout or 'updated:' in capture_result.stdout"
|
||||
no_log: true # content is cert/key/keyring material — don't echo via ansible logs
|
||||
|
||||
- name: Capture summary # noqa: run-once[task]
|
||||
ansible.builtin.debug:
|
||||
msg: "{{ item.stdout }}"
|
||||
loop: "{{ capture_result.results }}"
|
||||
loop_control:
|
||||
label: "{{ item.item.item.title_suffix }}"
|
||||
run_once: true
|
||||
@@ -0,0 +1,208 @@
|
||||
---
|
||||
# Provision sietch-ceph nodes from the Debian 12 live image.
|
||||
#
|
||||
# Fresh-install workflow (no cloning):
|
||||
# 1. Boot all target nodes to Debian 12 live image (user/live credentials)
|
||||
# 2. Ensure the live image is reachable on its reserved IP (= bond_ip)
|
||||
# 3. CEPH_ENV=inventories/<cluster>/inventory-provision.ini \
|
||||
# scripts/ansible-play.sh provision.yml -e confirm_wipe=true
|
||||
#
|
||||
# Flags:
|
||||
# confirm_wipe=true required — destroys all data on SSDs
|
||||
# provision_create_iac_keypair=true opt-in: preflight auto-generates
|
||||
# the ansible-iac deploy keypair if
|
||||
# it doesn't already exist on disk
|
||||
# provision_skip_reboot=true opt-in: skip the final reboot and
|
||||
# the post-reboot verification play
|
||||
# (canary inspection of the chroot)
|
||||
# --limit <host> provision one node at a time
|
||||
#
|
||||
# Post-reboot: a separate verification play waits on the installed OS
|
||||
# (reached via the production sietch-ceph-<name> SSH alias / ansible-iac
|
||||
# user with key) and confirms the provisioning marker is in place.
|
||||
|
||||
# --- Preflight: controller-side prerequisites -------------------------
|
||||
#
|
||||
# Runs BEFORE the target hosts are touched. Verifies the ansible-iac
|
||||
# deploy keypair exists on the controller (it gets installed into every
|
||||
# node's ~ansible-iac/.ssh/authorized_keys via file lookup during the
|
||||
# main play). Keypair generation is strictly opt-in via
|
||||
# -e provision_create_iac_keypair=true — default behavior is to fail
|
||||
# fast with a clear message if the keypair is missing.
|
||||
|
||||
- name: Preflight — verify controller prerequisites
|
||||
hosts: localhost
|
||||
gather_facts: false
|
||||
connection: local
|
||||
tasks:
|
||||
- name: Stat ansible-iac private key
|
||||
ansible.builtin.stat:
|
||||
path: "{{ provision_iac_ssh_key_path | expanduser }}"
|
||||
register: iac_privkey_stat
|
||||
|
||||
- name: Stat ansible-iac public key
|
||||
ansible.builtin.stat:
|
||||
path: "{{ (provision_iac_ssh_key_path | expanduser) ~ '.pub' }}"
|
||||
register: iac_pubkey_stat
|
||||
|
||||
- name: Generate ansible-iac keypair (only when opt-in flag is set)
|
||||
ansible.builtin.command:
|
||||
cmd: >-
|
||||
ssh-keygen -t ed25519 -N ''
|
||||
-C "ansible-iac@{{ cluster_name }}-{{ cluster_role }}.{{ cluster_domain }}"
|
||||
-f {{ provision_iac_ssh_key_path | expanduser }}
|
||||
args:
|
||||
creates: "{{ provision_iac_ssh_key_path | expanduser }}"
|
||||
when:
|
||||
- not iac_privkey_stat.stat.exists
|
||||
- provision_create_iac_keypair | default(false) | bool
|
||||
|
||||
- name: Fail if ansible-iac keypair missing and create flag not set
|
||||
ansible.builtin.fail:
|
||||
msg: |
|
||||
ansible-iac deploy keypair is missing on the controller:
|
||||
{{ provision_iac_ssh_key_path | expanduser }}
|
||||
{{ provision_iac_ssh_key_path | expanduser }}.pub
|
||||
|
||||
This keypair is required to provision the cluster — it is
|
||||
installed into ansible-iac's authorized_keys on every node
|
||||
and used by ansible_ssh_private_key_file in inventory.ini.
|
||||
|
||||
To auto-generate it on this run, re-invoke with:
|
||||
-e provision_create_iac_keypair=true
|
||||
|
||||
Or generate it yourself first:
|
||||
ssh-keygen -t ed25519 -N '' \
|
||||
-C "ansible-iac@{{ cluster_name }}-{{ cluster_role }}.{{ cluster_domain }}" \
|
||||
-f {{ provision_iac_ssh_key_path }}
|
||||
when:
|
||||
- not iac_privkey_stat.stat.exists
|
||||
- not (provision_create_iac_keypair | default(false) | bool)
|
||||
|
||||
- name: Re-stat public key after potential creation
|
||||
ansible.builtin.stat:
|
||||
path: "{{ (provision_iac_ssh_key_path | expanduser) ~ '.pub' }}"
|
||||
register: iac_pubkey_final_stat
|
||||
|
||||
- name: Assert ansible-iac public key exists after preflight
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- iac_pubkey_final_stat.stat.exists
|
||||
fail_msg: >-
|
||||
ansible-iac public key still missing at
|
||||
{{ (provision_iac_ssh_key_path | expanduser) ~ '.pub' }}
|
||||
after preflight completed. Investigate ssh-keygen task above.
|
||||
|
||||
- name: Report ansible-iac keypair ready
|
||||
ansible.builtin.debug:
|
||||
msg: >-
|
||||
ansible-iac keypair present at
|
||||
{{ provision_iac_ssh_key_path | expanduser }}(.pub) —
|
||||
preflight OK.
|
||||
|
||||
- name: Provision sietch-ceph nodes (live image → installed OS)
|
||||
hosts: provision_targets
|
||||
gather_facts: false
|
||||
serial: 1
|
||||
become: true
|
||||
max_fail_percentage: 0
|
||||
|
||||
roles:
|
||||
- role: provision_host
|
||||
|
||||
- name: Verify provisioned nodes after reboot
|
||||
hosts: provision_targets
|
||||
gather_facts: false
|
||||
become: false
|
||||
tasks:
|
||||
# Verification runs on localhost via delegate_to + the production
|
||||
# sietch-ceph-<name> SSH alias. That avoids having to juggle live
|
||||
# vs installed credentials inside a single play.
|
||||
|
||||
- name: Skip verification play when reboot was intentionally skipped
|
||||
ansible.builtin.meta: end_play
|
||||
when: provision_skip_reboot | default(false) | bool
|
||||
|
||||
# Give the box a head start before we start hammering ssh. R730xd
|
||||
# POST + GRUB + initramfs + systemd brings the bond up in ~60-90s.
|
||||
- name: Initial reboot wait
|
||||
delegate_to: localhost
|
||||
ansible.builtin.pause:
|
||||
seconds: "{{ provision_reboot_wait_delay | default(60) }}"
|
||||
|
||||
# The ssh probe is the actual gate — it honors ~/.ssh/config
|
||||
# (including any ProxyJump or ProxyCommand entries) so it works
|
||||
# from the controller regardless of whether the target subnet is
|
||||
# directly routable. wait_for/tcp does NOT honor ssh_config, so
|
||||
# we rely on retries here instead.
|
||||
#
|
||||
# We force user + identity explicitly with -l / -i / IdentitiesOnly
|
||||
# because ~/.ssh/config for the host alias lands as the 'ops' human
|
||||
# user (which has NO authorized_keys on a fresh provision — operator
|
||||
# keys are distributed out-of-band). The verify play is an automation
|
||||
# reachability test, so it uses ansible-iac (the same identity
|
||||
# ansible_ssh_private_key_file points at in inventory.ini).
|
||||
#
|
||||
# UserKnownHostsFile=/dev/null + StrictHostKeyChecking=no bypass
|
||||
# known_hosts entirely: fresh provisioning generates new host keys,
|
||||
# and on reprovisioning a stale known_hosts entry would otherwise
|
||||
# trigger "REMOTE HOST IDENTIFICATION HAS CHANGED" and abort. The
|
||||
# real identity check happens via the marker assertion below, not
|
||||
# via the SSH host key cache. LogLevel=ERROR suppresses the
|
||||
# "Warning: Permanently added..." noise.
|
||||
- name: Probe installed OS via production SSH alias (as ansible-iac)
|
||||
delegate_to: localhost
|
||||
ansible.builtin.command:
|
||||
cmd: >-
|
||||
ssh -o BatchMode=yes -o ConnectTimeout=10
|
||||
-o StrictHostKeyChecking=no
|
||||
-o UserKnownHostsFile=/dev/null
|
||||
-o IdentitiesOnly=yes
|
||||
-o LogLevel=ERROR
|
||||
-l ansible-iac
|
||||
-i {{ provision_iac_ssh_key_path | expanduser }}
|
||||
{{ hostname_short }}
|
||||
"cat {{ provision_marker_path | default('/etc/ceph-provisioned.json') }}"
|
||||
register: marker_read
|
||||
changed_when: false
|
||||
retries: "{{ provision_verify_retries | default(30) }}"
|
||||
delay: "{{ provision_verify_delay | default(15) }}"
|
||||
until: marker_read.rc == 0
|
||||
|
||||
- name: Parse provisioning marker
|
||||
ansible.builtin.set_fact:
|
||||
marker: "{{ marker_read.stdout | from_json }}"
|
||||
|
||||
- name: Verify marker contents match inventory expectations
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- marker.hostname == hostname_short
|
||||
- marker.bond_ip == bond_ip
|
||||
- marker.cluster_domain == cluster_domain
|
||||
- marker.admin_user == admin_user
|
||||
fail_msg: "Marker contents do not match inventory: {{ marker }}"
|
||||
success_msg: "{{ hostname_short }} provisioned OK at {{ marker.provisioned_at }}"
|
||||
|
||||
- name: Gather installed OS hostname (as ansible-iac)
|
||||
delegate_to: localhost
|
||||
ansible.builtin.command:
|
||||
cmd: >-
|
||||
ssh -o BatchMode=yes -o ConnectTimeout=10
|
||||
-o StrictHostKeyChecking=no
|
||||
-o UserKnownHostsFile=/dev/null
|
||||
-o IdentitiesOnly=yes
|
||||
-o LogLevel=ERROR
|
||||
-l ansible-iac
|
||||
-i {{ provision_iac_ssh_key_path | expanduser }}
|
||||
{{ hostname_short }} hostname -f
|
||||
register: installed_hostname
|
||||
changed_when: false
|
||||
|
||||
- name: Display post-provision state
|
||||
ansible.builtin.debug:
|
||||
msg:
|
||||
- "hostname -f: {{ installed_hostname.stdout | trim }}"
|
||||
- "marker.hostname: {{ marker.hostname }}"
|
||||
- "marker.fqdn: {{ marker.fqdn }}"
|
||||
- "marker.bond_ip: {{ marker.bond_ip }}"
|
||||
- "marker.provisioned: {{ marker.provisioned_at }}"
|
||||
@@ -0,0 +1,103 @@
|
||||
---
|
||||
# RADOS benchmark — tests raw cluster I/O bypassing RGW/S3.
|
||||
# Useful for isolating storage performance from S3 overhead.
|
||||
#
|
||||
# Usage:
|
||||
# scripts/ansible-play.sh rados-bench.yml # 30s write, replicated
|
||||
# scripts/ansible-play.sh rados-bench.yml -e rados_ec=true # EC 8+3 (matches RGW data pool)
|
||||
# scripts/ansible-play.sh rados-bench.yml -e rados_mode=seq # sequential read
|
||||
# scripts/ansible-play.sh rados-bench.yml -e rados_seconds=60
|
||||
# scripts/ansible-play.sh rados-bench.yml -e rados_cleanup=false # keep data for read test
|
||||
#
|
||||
# Modes: write, seq (sequential read), rand (random read)
|
||||
# Read modes require a prior write run with rados_cleanup=false.
|
||||
|
||||
- name: RADOS benchmark
|
||||
hosts: ceph_nodes
|
||||
become: true
|
||||
gather_facts: false
|
||||
|
||||
vars:
|
||||
rados_pool: "rados-bench"
|
||||
rados_mode: write
|
||||
rados_seconds: 30
|
||||
rados_threads: 16
|
||||
rados_cleanup: true
|
||||
rados_ec: false
|
||||
rados_ec_profile: "ec-k8m3-osd" # reuse the RGW EC profile (8+3)
|
||||
|
||||
tasks:
|
||||
- name: Create EC bench pool if missing
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
if ceph osd pool ls | grep -q '^{{ rados_pool }}$'; then
|
||||
echo "EXISTS"
|
||||
else
|
||||
ceph osd pool create {{ rados_pool }} 64 erasure {{ rados_ec_profile }}
|
||||
ceph osd pool set {{ rados_pool }} allow_ec_overwrites true
|
||||
echo "CREATED"
|
||||
fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- rados_ec | bool
|
||||
register: ec_pool_create
|
||||
changed_when: "'CREATED' in ec_pool_create.stdout | default('')"
|
||||
|
||||
- name: Create replicated bench pool if missing
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
if ceph osd pool ls | grep -q '^{{ rados_pool }}$'; then
|
||||
echo "EXISTS"
|
||||
else
|
||||
ceph osd pool create {{ rados_pool }} 32 replicated
|
||||
ceph osd pool set {{ rados_pool }} size 2
|
||||
ceph osd pool set {{ rados_pool }} min_size 1
|
||||
echo "CREATED"
|
||||
fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- not (rados_ec | bool)
|
||||
register: rep_pool_create
|
||||
changed_when: "'CREATED' in rep_pool_create.stdout | default('')"
|
||||
|
||||
- name: Show benchmark parameters
|
||||
ansible.builtin.debug:
|
||||
msg:
|
||||
- "Node: {{ inventory_hostname }}"
|
||||
- "Pool: {{ rados_pool }}"
|
||||
- "Mode: {{ rados_mode }}"
|
||||
- "Duration: {{ rados_seconds }}s"
|
||||
- "Threads: {{ rados_threads }}"
|
||||
|
||||
- name: Run rados bench
|
||||
ansible.builtin.command: >
|
||||
rados bench -p {{ rados_pool }}
|
||||
{{ rados_seconds }} {{ rados_mode }}
|
||||
-t {{ rados_threads }}
|
||||
--run-name {{ inventory_hostname }}
|
||||
{{ '--no-cleanup' if not (rados_cleanup | bool) else '' }}
|
||||
register: rados_result
|
||||
changed_when: true
|
||||
timeout: "{{ (rados_seconds | int * 3) + 60 }}"
|
||||
|
||||
- name: Display results
|
||||
ansible.builtin.debug:
|
||||
msg: "{{ rados_result.stdout_lines[-20:] }}"
|
||||
|
||||
- name: Clean up bench pool
|
||||
ansible.builtin.shell: |
|
||||
ceph config set mon mon_allow_pool_delete true
|
||||
ceph osd pool delete {{ rados_pool }} {{ rados_pool }} \
|
||||
--yes-i-really-really-mean-it
|
||||
ceph config set mon mon_allow_pool_delete false
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- rados_cleanup | bool
|
||||
- rados_mode == 'write'
|
||||
changed_when: true
|
||||
@@ -0,0 +1,6 @@
|
||||
ansible-core==2.20.4
|
||||
ansible-lint==26.4.0
|
||||
ansible-navigator>=25.0
|
||||
yamllint==1.38.0
|
||||
molecule>=25.0
|
||||
boto3==1.42.87
|
||||
@@ -0,0 +1,6 @@
|
||||
---
|
||||
collections:
|
||||
- name: ansible.posix
|
||||
version: ">=2.0.0"
|
||||
- name: community.general
|
||||
version: ">=9.0.0"
|
||||
@@ -0,0 +1,49 @@
|
||||
---
|
||||
# Post-boot OS baseline — runs via ansible-iac after provisioning.
|
||||
# Configures everything that doesn't need to be in the chroot.
|
||||
|
||||
# --- ops user ---
|
||||
baseline_ops_user: ops
|
||||
baseline_ops_uid: 1001
|
||||
baseline_ops_sudo: "ALL" # "ALL" = password required; "NOPASSWD:ALL" for dev
|
||||
|
||||
# --- /etc/hosts ---
|
||||
# Rendered from the same cluster variables used by ceph_deploy/hosts.j2.
|
||||
# Keeps host entries converged even if ceph_deploy hasn't run yet.
|
||||
baseline_manage_hosts: true
|
||||
|
||||
# --- Packages ---
|
||||
# Podman ecosystem (needed by cephadm)
|
||||
baseline_podman_packages:
|
||||
- podman
|
||||
- catatonit
|
||||
- dbus-user-session
|
||||
- fuse-overlayfs
|
||||
- slirp4netns
|
||||
- uidmap
|
||||
|
||||
# Diagnostic and ops toolset
|
||||
baseline_diag_packages:
|
||||
- smartmontools
|
||||
- ipmitool
|
||||
- lsscsi
|
||||
- pciutils
|
||||
- lshw
|
||||
- ethtool
|
||||
- jq
|
||||
- htop
|
||||
- btop
|
||||
- tmux
|
||||
- bat
|
||||
- ncdu
|
||||
- nload
|
||||
- sysstat
|
||||
- iotop
|
||||
- bsdmainutils
|
||||
- python3-boto3
|
||||
|
||||
# --- Services ---
|
||||
baseline_enable_services:
|
||||
- dbus
|
||||
- chrony
|
||||
- podman.socket
|
||||
@@ -0,0 +1,11 @@
|
||||
---
|
||||
# Runs after provision_host, before ceph_deploy.
|
||||
# Requires ansible-iac SSH access (set up during provisioning).
|
||||
dependencies: []
|
||||
|
||||
galaxy_info:
|
||||
author: FUTO
|
||||
license: AGPL-3.0-only
|
||||
role_name: baseline
|
||||
description: Post-boot OS baseline — ops user, packages, hosts, services
|
||||
min_ansible_version: "2.19"
|
||||
@@ -0,0 +1,18 @@
|
||||
---
|
||||
# Post-boot OS baseline.
|
||||
# Runs via ansible-iac on the installed OS (not in chroot).
|
||||
# Assumes: ansible-iac user exists with key-only SSH + NOPASSWD sudo
|
||||
# (created by provision_host during initial OS install).
|
||||
# Everything here is convergeable via normal deploy pipeline.
|
||||
|
||||
- name: Configure ops user
|
||||
ansible.builtin.import_tasks: users.yml
|
||||
tags: [users]
|
||||
|
||||
- name: Configure system packages
|
||||
ansible.builtin.import_tasks: packages.yml
|
||||
tags: [packages]
|
||||
|
||||
- name: Configure /etc/hosts and services
|
||||
ansible.builtin.import_tasks: system.yml
|
||||
tags: [system]
|
||||
@@ -0,0 +1,18 @@
|
||||
---
|
||||
# Install podman ecosystem + diagnostic toolset.
|
||||
# Convergeable: re-running installs anything missing.
|
||||
|
||||
- name: Update apt cache
|
||||
ansible.builtin.apt:
|
||||
update_cache: true
|
||||
cache_valid_time: 3600
|
||||
|
||||
- name: Install podman ecosystem
|
||||
ansible.builtin.apt:
|
||||
name: "{{ baseline_podman_packages }}"
|
||||
state: present
|
||||
|
||||
- name: Install diagnostic and ops packages
|
||||
ansible.builtin.apt:
|
||||
name: "{{ baseline_diag_packages }}"
|
||||
state: present
|
||||
@@ -0,0 +1,19 @@
|
||||
---
|
||||
# /etc/hosts, service enablement, timezone.
|
||||
# Convergeable: re-running corrects drift in hosts file and services.
|
||||
|
||||
- name: Render /etc/hosts with cluster node entries
|
||||
ansible.builtin.template:
|
||||
src: hosts.j2
|
||||
dest: /etc/hosts
|
||||
owner: root
|
||||
group: root
|
||||
mode: '0644'
|
||||
when: baseline_manage_hosts | bool
|
||||
|
||||
- name: Enable and start required services
|
||||
ansible.builtin.systemd:
|
||||
name: "{{ item }}"
|
||||
enabled: true
|
||||
state: started
|
||||
loop: "{{ baseline_enable_services }}"
|
||||
@@ -0,0 +1,38 @@
|
||||
---
|
||||
# ops user — human-interactive account.
|
||||
# Convergeable: running this again fixes drift in password, sudo, shell.
|
||||
|
||||
- name: Ensure ops user exists
|
||||
ansible.builtin.user:
|
||||
name: "{{ baseline_ops_user }}"
|
||||
uid: "{{ baseline_ops_uid }}"
|
||||
shell: /bin/bash
|
||||
groups: sudo
|
||||
append: true
|
||||
create_home: true
|
||||
state: present
|
||||
|
||||
- name: Assert ops_password is defined and non-empty
|
||||
ansible.builtin.assert:
|
||||
that:
|
||||
- ops_password is defined
|
||||
- ops_password | length > 0
|
||||
fail_msg: >-
|
||||
ops_password is undefined or empty. Run the play via
|
||||
scripts/ansible-play.sh (not bare ansible-playbook) so the cluster's
|
||||
secrets.yml.tpl is resolved via op inject into --extra-vars. If that's
|
||||
already what you did, verify vault_ops_password resolves:
|
||||
`op read op://<vault>/<CLUSTER>_CEPH_OPS_PASSWORD/password`.
|
||||
|
||||
- name: Set ops user password
|
||||
ansible.builtin.user:
|
||||
name: "{{ baseline_ops_user }}"
|
||||
password: "{{ ops_password | password_hash('sha512') }}"
|
||||
no_log: true
|
||||
|
||||
- name: Configure ops sudo
|
||||
ansible.builtin.copy:
|
||||
content: "{{ baseline_ops_user }} ALL=(ALL) {{ baseline_ops_sudo }}\n"
|
||||
dest: "/etc/sudoers.d/{{ baseline_ops_user }}"
|
||||
mode: '0440'
|
||||
validate: "visudo -cf %s"
|
||||
@@ -0,0 +1,10 @@
|
||||
127.0.0.1 localhost
|
||||
|
||||
# Ceph cluster nodes
|
||||
{% for host in groups['ceph_nodes'] %}
|
||||
{{ hostvars[host]['bond_ip'] }} {{ hostvars[host]['hostname_short'] }}.{{ cluster_domain }} {{ hostvars[host]['hostname_short'] }}
|
||||
{% endfor %}
|
||||
|
||||
::1 localhost ip6-localhost ip6-loopback
|
||||
ff02::1 ip6-allnodes
|
||||
ff02::2 ip6-allrouters
|
||||
@@ -0,0 +1,22 @@
|
||||
---
|
||||
# Ceph release train. Upstream publishes per-release apt trees at
|
||||
# download.ceph.com/debian-<release>/, and each tree has subdirectories
|
||||
# per Debian codename. prerequisites.yml pins the codename to 'bookworm'
|
||||
# in the sources.list entry — when upgrading the base OS to Trixie,
|
||||
# flip the codename there. This split keeps the Ceph release and the
|
||||
# Debian release independently versionable.
|
||||
ceph_release: tentacle
|
||||
ceph_release_version: "20.2.*" # pin to patch range; set "20.2.1" to pin exactly
|
||||
ceph_repo_url: "https://download.ceph.com/debian-{{ ceph_release }}/"
|
||||
ceph_repo_key_url: "https://download.ceph.com/keys/release.asc"
|
||||
cephadm_install_method: repo # 'repo' or 'curl'
|
||||
|
||||
# Initial PG counts per pool. The autoscaler is left on and will adjust
|
||||
# over time, but starting at pg_num=1 (the ceph default) causes PG
|
||||
# splitting under load which tanks performance during the first fill.
|
||||
# These values are sized for ~36 OSDs. Scale proportionally for larger
|
||||
# clusters.
|
||||
ceph_pg_init_data: 128 # EC data pool — bulk of I/O
|
||||
ceph_pg_init_index: 16 # bucket index
|
||||
ceph_pg_init_non_ec: 16 # multipart uploads
|
||||
ceph_pg_init_meta: 8 # .rgw.root, .meta, .log, .control
|
||||
@@ -0,0 +1,14 @@
|
||||
---
|
||||
- name: Update apt cache
|
||||
ansible.builtin.apt:
|
||||
update_cache: true
|
||||
|
||||
- name: Restart chronyd
|
||||
ansible.builtin.systemd:
|
||||
name: chronyd
|
||||
state: restarted
|
||||
|
||||
- name: Restart sshd
|
||||
ansible.builtin.systemd:
|
||||
name: ssh
|
||||
state: restarted
|
||||
@@ -0,0 +1,18 @@
|
||||
---
|
||||
# Advisory: os_tuning and hardware_tuning should run before this role.
|
||||
# Not declared as hard dependencies because the tuning playbooks are
|
||||
# separate operational steps, not automatic prerequisites.
|
||||
#
|
||||
# Execution order:
|
||||
# 1. provision_host (bare metal → Debian)
|
||||
# 2. os_tuning (sysctl, ulimits)
|
||||
# 3. hardware_tuning (I/O scheduler, readahead)
|
||||
# 4. ceph_deploy (this role)
|
||||
dependencies: []
|
||||
|
||||
galaxy_info:
|
||||
author: FUTO
|
||||
license: AGPL-3.0-only
|
||||
role_name: ceph_deploy
|
||||
description: Deploy Ceph Tentacle cluster via cephadm
|
||||
min_ansible_version: "2.19"
|
||||
@@ -0,0 +1,50 @@
|
||||
---
|
||||
dependency:
|
||||
name: galaxy
|
||||
options:
|
||||
requirements-file: ${MOLECULE_PROJECT_DIRECTORY}/../../requirements.yml
|
||||
|
||||
driver:
|
||||
name: default
|
||||
|
||||
platforms:
|
||||
- name: molecule-ceph-test
|
||||
image: debian:bookworm
|
||||
pre_build_image: true
|
||||
|
||||
provisioner:
|
||||
name: ansible
|
||||
inventory:
|
||||
hosts:
|
||||
all:
|
||||
hosts:
|
||||
molecule-ceph-test:
|
||||
hostname_short: molecule-ceph-test
|
||||
bond_ip: 10.10.10.99
|
||||
sas_path_prefix: "pci-0000:02:00.0-sas-exp0x500056b300000000"
|
||||
ssd1_phy: 12
|
||||
ssd2_phy: 13
|
||||
ceph_db_vg1: ceph-db-rear12
|
||||
ceph_db_vg2: ceph-db-rear13
|
||||
ceph_hdd_osds:
|
||||
- path_phy: phy0
|
||||
db: ceph-db-rear12/db-slot0
|
||||
ceph_ssd_osds:
|
||||
- path_phy: phy12
|
||||
partition: 6
|
||||
children:
|
||||
ceph_nodes:
|
||||
hosts:
|
||||
molecule-ceph-test:
|
||||
ceph_bootstrap:
|
||||
hosts:
|
||||
molecule-ceph-test:
|
||||
ceph_join:
|
||||
hosts: {}
|
||||
config_options:
|
||||
defaults:
|
||||
gathering: implicit
|
||||
|
||||
# Only verify templates render — don't converge (needs real hardware)
|
||||
verifier:
|
||||
name: ansible
|
||||
@@ -0,0 +1,55 @@
|
||||
---
|
||||
# Molecule verify — validates that templates render without errors.
|
||||
# Does NOT require a running Ceph cluster.
|
||||
|
||||
- name: Verify ceph_deploy template rendering
|
||||
hosts: all
|
||||
gather_facts: false
|
||||
|
||||
vars:
|
||||
cluster_domain: dev.austin.int.futo.cloud
|
||||
public_network: 10.10.10.0/24
|
||||
cluster_network: 10.10.10.0/24
|
||||
ceph_rgw_realm: sietch
|
||||
ceph_rgw_zone: dev-z1
|
||||
ceph_rgw_port: 443
|
||||
ceph_rgw_ssl: true
|
||||
rgw_ssl_cert_combined_pem: "MOCK_CERT_CONTENT"
|
||||
ceph_rgw_dns_name: s3.dev.austin.int.futo.cloud
|
||||
|
||||
tasks:
|
||||
- name: Render hosts.j2
|
||||
ansible.builtin.template:
|
||||
src: "../../templates/hosts.j2"
|
||||
dest: /tmp/molecule-hosts.txt
|
||||
mode: '0644'
|
||||
|
||||
- name: Verify hosts.j2 has no 127.0.1.1
|
||||
ansible.builtin.command: grep -c '127.0.1.1' /tmp/molecule-hosts.txt
|
||||
register: hosts_check
|
||||
failed_when: hosts_check.rc == 0
|
||||
changed_when: false
|
||||
|
||||
- name: Render cluster-spec.yml.j2
|
||||
ansible.builtin.template:
|
||||
src: "../../templates/cluster-spec.yml.j2"
|
||||
dest: /tmp/molecule-cluster-spec.yml
|
||||
mode: '0644'
|
||||
|
||||
- name: Render rgw-spec.yaml.j2
|
||||
ansible.builtin.template:
|
||||
src: "../../templates/rgw-spec.yaml.j2"
|
||||
dest: /tmp/molecule-rgw-spec.yaml
|
||||
mode: '0644'
|
||||
|
||||
- name: Verify rgw-spec uses hostname_short
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
grep 'molecule-ceph-test' /tmp/molecule-rgw-spec.yaml
|
||||
args:
|
||||
executable: /bin/bash
|
||||
changed_when: false
|
||||
|
||||
- name: Report success
|
||||
ansible.builtin.debug:
|
||||
msg: "All templates rendered and validated successfully"
|
||||
@@ -0,0 +1,70 @@
|
||||
---
|
||||
# Phase 2: Bootstrap cluster (first node only)
|
||||
|
||||
- name: Check if Ceph cluster already exists
|
||||
ansible.builtin.stat:
|
||||
path: /etc/ceph/ceph.conf
|
||||
register: ceph_conf
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
|
||||
- name: Write initial ceph config
|
||||
ansible.builtin.template:
|
||||
src: cluster-spec.yml.j2
|
||||
dest: /tmp/ceph-initial.conf
|
||||
mode: '0644'
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- not (ceph_conf.stat.exists | default(false))
|
||||
|
||||
- name: Write dashboard password to temp file
|
||||
ansible.builtin.copy:
|
||||
content: "{{ ceph_dashboard_password }}"
|
||||
dest: /tmp/.ceph-dashboard-pw
|
||||
owner: root
|
||||
group: root
|
||||
mode: '0600'
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- not (ceph_conf.stat.exists | default(false))
|
||||
no_log: true
|
||||
|
||||
- name: Bootstrap Ceph cluster
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
cephadm bootstrap \
|
||||
--mon-ip {{ bond_ip }} \
|
||||
--initial-dashboard-user {{ ceph_dashboard_user }} \
|
||||
--initial-dashboard-password "$(cat /tmp/.ceph-dashboard-pw)" \
|
||||
--dashboard-password-noupdate \
|
||||
--allow-fqdn-hostname \
|
||||
--config /tmp/ceph-initial.conf \
|
||||
2>&1 | tee /var/log/ceph-bootstrap.log
|
||||
RC=$?
|
||||
rm -f /tmp/.ceph-dashboard-pw
|
||||
exit $RC
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: bootstrap_result
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- not (ceph_conf.stat.exists | default(false))
|
||||
changed_when: bootstrap_result.rc | default(1) == 0
|
||||
no_log: true
|
||||
|
||||
- name: Clean up dashboard password file on skip
|
||||
ansible.builtin.file:
|
||||
path: /tmp/.ceph-dashboard-pw
|
||||
state: absent
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
timeout: 600
|
||||
|
||||
- name: Show bootstrap result
|
||||
ansible.builtin.debug:
|
||||
msg: "{{ bootstrap_result.stdout_lines[-15:] | default(['Already bootstrapped']) }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- bootstrap_result is not skipped
|
||||
|
||||
# cephadm automatically distributes ceph.conf, admin keyring, and
|
||||
# its SSH key to managed hosts during `ceph orch host add`. No
|
||||
# manual capture or distribution needed.
|
||||
@@ -0,0 +1,36 @@
|
||||
---
|
||||
# Phase 5.5: Create CRUSH device-class rules
|
||||
# Ensures pools can target HDD-only or SSD-only OSDs
|
||||
|
||||
- name: Check existing CRUSH rules
|
||||
ansible.builtin.command: ceph osd crush rule ls
|
||||
register: crush_rules
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Create replicated_hdd CRUSH rule
|
||||
ansible.builtin.command: >
|
||||
ceph osd crush rule create-replicated replicated_hdd default host hdd
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- "'replicated_hdd' not in crush_rules.stdout"
|
||||
changed_when: true
|
||||
|
||||
- name: Create replicated_ssd CRUSH rule
|
||||
ansible.builtin.command: >
|
||||
ceph osd crush rule create-replicated replicated_ssd default host ssd
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- "'replicated_ssd' not in crush_rules.stdout"
|
||||
changed_when: true
|
||||
|
||||
- name: Show CRUSH rules
|
||||
ansible.builtin.command: ceph osd crush rule ls
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
register: final_rules
|
||||
|
||||
- name: Display CRUSH rules
|
||||
ansible.builtin.debug:
|
||||
msg: "{{ final_rules.stdout_lines }}"
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
@@ -0,0 +1,106 @@
|
||||
---
|
||||
# Phase 3: Add nodes to cluster (from bootstrap node)
|
||||
#
|
||||
# Only attempts to add hosts that are in the current play (--limit aware).
|
||||
# This prevents failures when nodes are defined in inventory
|
||||
# but not yet available.
|
||||
|
||||
- name: Build list of joinable hosts
|
||||
ansible.builtin.set_fact:
|
||||
joinable_hosts: >-
|
||||
{{ groups['ceph_join'] | default([])
|
||||
| intersect(ansible_play_hosts)
|
||||
| list }}
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
|
||||
- name: Verify cephadm SSH to join nodes
|
||||
ansible.builtin.shell: |
|
||||
set -euo pipefail
|
||||
HOST="{{ hostvars[item]['hostname_short'] }}"
|
||||
ADDR="{{ hostvars[item]['bond_ip'] }}"
|
||||
echo "Testing SSH to $HOST ($ADDR)..."
|
||||
ceph cephadm check-host "$HOST" "$ADDR" 2>&1 || {
|
||||
echo "WARN: ceph cephadm check-host failed, testing raw SSH..."
|
||||
cephadm shell -- ssh -o StrictHostKeyChecking=no -o ConnectTimeout=10 \
|
||||
"root@${ADDR}" "cephadm check-host --expect-hostname $HOST" 2>&1
|
||||
}
|
||||
args:
|
||||
executable: /bin/bash
|
||||
loop: "{{ joinable_hosts | default([]) }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- joinable_hosts | default([]) | length > 0
|
||||
loop_control:
|
||||
label: "{{ hostvars[item]['hostname_short'] }}"
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
register: ssh_check_results
|
||||
|
||||
- name: Show SSH check results
|
||||
ansible.builtin.debug:
|
||||
msg: "{{ item.stdout_lines | default([]) }}"
|
||||
loop: "{{ ssh_check_results.results | default([]) }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- item.stdout is defined
|
||||
loop_control:
|
||||
label: "{{ item.item | default('unknown') }}"
|
||||
|
||||
- name: Add remaining nodes to Ceph cluster
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
# Check if host already added
|
||||
if ceph orch host ls --format json |
|
||||
grep -q '"{{ hostvars[item]['hostname_short'] }}"'; then
|
||||
echo "Host {{ hostvars[item]['hostname_short'] }} already in cluster"
|
||||
else
|
||||
echo "Adding {{ hostvars[item]['hostname_short'] }} ({{ hostvars[item]['bond_ip'] }})"
|
||||
ceph orch host add \
|
||||
{{ hostvars[item]['hostname_short'] }} \
|
||||
{{ hostvars[item]['bond_ip'] }} 2>&1
|
||||
fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
loop: "{{ joinable_hosts | default([]) }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- joinable_hosts | default([]) | length > 0
|
||||
register: host_add_results
|
||||
changed_when: "'Adding' in (host_add_results.stdout | default(''))"
|
||||
|
||||
- name: Show host add results
|
||||
ansible.builtin.debug:
|
||||
msg: "{{ item.stdout }}"
|
||||
loop: "{{ host_add_results.results | default([]) }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- item.stdout is defined
|
||||
loop_control:
|
||||
label: "{{ item.item | default('unknown') }}"
|
||||
|
||||
# cephadm automatically syncs ceph.conf and admin keyring to managed
|
||||
# hosts after `ceph orch host add`. No manual distribution needed.
|
||||
|
||||
- name: Wait for new hosts to be online
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
for i in $(seq 1 30); do
|
||||
online=$(ceph orch host ls --format json | python3 -c "
|
||||
import sys, json
|
||||
hosts = json.load(sys.stdin)
|
||||
print(len([h for h in hosts if h.get('status','') == '']))
|
||||
" 2>/dev/null || echo 0)
|
||||
expected={{ (joinable_hosts | default([]) | length) + 1 }}
|
||||
echo "Hosts online: $online/$expected"
|
||||
if [ "$online" -ge "$expected" ]; then
|
||||
exit 0
|
||||
fi
|
||||
sleep 10
|
||||
done
|
||||
echo "WARNING: Not all hosts online yet"
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- joinable_hosts | default([]) | length > 0
|
||||
changed_when: false
|
||||
@@ -0,0 +1,110 @@
|
||||
---
|
||||
# Phase 4.5: Ensure Ceph block.db LVM is set up on each node (sietch-shape only)
|
||||
#
|
||||
# This task is sietch-shape-specific: it assumes dual SAS-attached SSDs each
|
||||
# with partition 5 → its own VG, with 6 db-slot LVs per VG (12 total). It
|
||||
# ensures PVs, VGs, and LVs exist. Idempotent — skips if already present.
|
||||
# Provides a recovery path if the LVM was destroyed by `mise run destroy`.
|
||||
#
|
||||
# Painbox-shape (Hetzner SX295, NVMe RAID-1 → single vg0 with all LVs created
|
||||
# by installimage post-install) skips the whole block — its host_vars don't
|
||||
# define `sas_path_prefix` so the gate below short-circuits.
|
||||
#
|
||||
# Required host_vars (sietch-shape only):
|
||||
# sas_path_prefix — by-path prefix for SSD discovery (e.g., pci-0000:02:00.0-sas-exp...)
|
||||
# ssd1_phy, ssd2_phy — PHY slot numbers for the two SSDs
|
||||
# ceph_db_vg1, ceph_db_vg2 — VG names for SSD1/SSD2 partition 5
|
||||
#
|
||||
# VG mapping: ceph_db_vg1 on SSD1 partition 5, ceph_db_vg2 on SSD2 partition 5
|
||||
# LV naming: db-slot0..5 on VG1, db-slot6..11 on VG2 (one per HDD OSD)
|
||||
|
||||
- name: Skip in-role LVM setup (no sas_path_prefix — externally-managed LVM)
|
||||
ansible.builtin.debug:
|
||||
msg: >-
|
||||
LVM is externally managed on this host (e.g., Hetzner installimage
|
||||
post-install creates vg0 + db-slots on a RAID-1 NVMe). Skipping the
|
||||
sietch-shape dual-SSD-VG setup. If LVM is missing on this host,
|
||||
reprovision via the cluster's installimage flow.
|
||||
when: sas_path_prefix is undefined
|
||||
|
||||
- name: Set up LVM (sietch-shape dual-SSD-VG topology)
|
||||
when: sas_path_prefix is defined
|
||||
block:
|
||||
- name: Resolve SSD1 device path
|
||||
ansible.builtin.command: >
|
||||
readlink -f /dev/disk/by-path/{{ sas_path_prefix }}-phy{{ ssd1_phy }}-lun-0
|
||||
register: ssd1_dev
|
||||
changed_when: false
|
||||
|
||||
- name: Resolve SSD2 device path
|
||||
ansible.builtin.command: >
|
||||
readlink -f /dev/disk/by-path/{{ sas_path_prefix }}-phy{{ ssd2_phy }}-lun-0
|
||||
register: ssd2_dev
|
||||
changed_when: false
|
||||
|
||||
- name: Check if VG1 exists
|
||||
ansible.builtin.shell: vgs {{ ceph_db_vg1 }} 2>/dev/null
|
||||
register: vg1_check
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
|
||||
- name: Check if VG2 exists
|
||||
ansible.builtin.shell: vgs {{ ceph_db_vg2 }} 2>/dev/null
|
||||
register: vg2_check
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
|
||||
- name: Create PV and VG on SSD1 partition 5
|
||||
ansible.builtin.shell: |
|
||||
PART="{{ ssd1_dev.stdout }}5"
|
||||
wipefs -af "$PART"
|
||||
pvcreate -f "$PART" && vgcreate {{ ceph_db_vg1 }} "$PART"
|
||||
when: vg1_check.rc != 0
|
||||
changed_when: true
|
||||
|
||||
- name: Create PV and VG on SSD2 partition 5
|
||||
ansible.builtin.shell: |
|
||||
PART="{{ ssd2_dev.stdout }}5"
|
||||
wipefs -af "$PART"
|
||||
pvcreate -f "$PART" && vgcreate {{ ceph_db_vg2 }} "$PART"
|
||||
when: vg2_check.rc != 0
|
||||
changed_when: true
|
||||
|
||||
- name: Create block.db LVs on VG1 (slots 0-4 at 240G, slot 5 gets remainder)
|
||||
ansible.builtin.shell: |
|
||||
if lvs {{ ceph_db_vg1 }}/db-slot{{ item }} 2>/dev/null; then
|
||||
echo "db-slot{{ item }} already exists"
|
||||
else
|
||||
{% if item < 5 %}
|
||||
lvcreate --yes -Wy -L {{ ceph_db_lv_size }} -n db-slot{{ item }} {{ ceph_db_vg1 }}
|
||||
{% else %}
|
||||
lvcreate --yes -Wy -l 100%FREE -n db-slot{{ item }} {{ ceph_db_vg1 }}
|
||||
{% endif %}
|
||||
fi
|
||||
loop: [0, 1, 2, 3, 4, 5]
|
||||
register: vg1_lv_results
|
||||
changed_when: "'already exists' not in (vg1_lv_results.stdout | default(''))"
|
||||
|
||||
- name: Create block.db LVs on VG2 (slots 6-10 at 240G, slot 11 gets remainder)
|
||||
ansible.builtin.shell: |
|
||||
if lvs {{ ceph_db_vg2 }}/db-slot{{ item }} 2>/dev/null; then
|
||||
echo "db-slot{{ item }} already exists"
|
||||
else
|
||||
{% if item < 11 %}
|
||||
lvcreate --yes -Wy -L {{ ceph_db_lv_size }} -n db-slot{{ item }} {{ ceph_db_vg2 }}
|
||||
{% else %}
|
||||
lvcreate --yes -Wy -l 100%FREE -n db-slot{{ item }} {{ ceph_db_vg2 }}
|
||||
{% endif %}
|
||||
fi
|
||||
loop: [6, 7, 8, 9, 10, 11]
|
||||
register: vg2_lv_results
|
||||
changed_when: "'already exists' not in (vg2_lv_results.stdout | default(''))"
|
||||
|
||||
- name: Verify LVM setup
|
||||
ansible.builtin.command: lvs -o lv_name,vg_name,lv_size --noheadings
|
||||
register: lvm_state
|
||||
changed_when: false
|
||||
|
||||
- name: Show LVM state
|
||||
ansible.builtin.debug:
|
||||
msg: "{{ lvm_state.stdout_lines }}"
|
||||
@@ -0,0 +1,43 @@
|
||||
---
|
||||
# Ceph Tentacle deployment via cephadm
|
||||
# Phases: prerequisites -> bootstrap -> join nodes -> placement -> OSDs -> verify
|
||||
|
||||
- name: Phase 1 - Prerequisites
|
||||
ansible.builtin.import_tasks: prerequisites.yml
|
||||
tags: [prerequisites]
|
||||
|
||||
- name: Phase 2 - Bootstrap cluster
|
||||
ansible.builtin.import_tasks: bootstrap.yml
|
||||
tags: [bootstrap]
|
||||
|
||||
- name: Phase 3 - Join nodes
|
||||
ansible.builtin.import_tasks: join.yml
|
||||
tags: [join]
|
||||
|
||||
- name: Phase 4 - MON/MGR placement
|
||||
ansible.builtin.import_tasks: placement.yml
|
||||
tags: [placement]
|
||||
|
||||
- name: Phase 4.5 - Ensure block.db LVM exists
|
||||
ansible.builtin.import_tasks: lvm-setup.yml
|
||||
tags: [lvm, osds]
|
||||
|
||||
- name: Phase 5 - Create OSDs
|
||||
ansible.builtin.import_tasks: osds.yml
|
||||
tags: [osds]
|
||||
|
||||
- name: Phase 5.5 - CRUSH device-class rules
|
||||
ansible.builtin.import_tasks: crush-rules.yml
|
||||
tags: [crush, osds]
|
||||
|
||||
- name: Phase 5.75 - RadosGW (S3 Object Gateway)
|
||||
ansible.builtin.import_tasks: rgw.yml
|
||||
tags: [rgw]
|
||||
|
||||
- name: Phase 5.8 - Monitoring stack
|
||||
ansible.builtin.import_tasks: monitoring.yml
|
||||
tags: [monitoring]
|
||||
|
||||
- name: Phase 6 - Verify cluster
|
||||
ansible.builtin.import_tasks: verify.yml
|
||||
tags: [verify]
|
||||
@@ -0,0 +1,138 @@
|
||||
---
|
||||
# Phase 5.8: Monitoring stack — dashboard integration
|
||||
#
|
||||
# cephadm auto-deploys the full monitoring stack (node-exporter,
|
||||
# ceph-exporter, prometheus, alertmanager, grafana) during bootstrap.
|
||||
# This task file only:
|
||||
# 1. Ensures the prometheus mgr module is enabled
|
||||
# 2. Waits for all monitoring daemons to come up
|
||||
# 3. Configures dashboard integration URLs and credentials
|
||||
|
||||
# --- Step 1: Enable prometheus mgr module ---
|
||||
|
||||
- name: Check prometheus mgr module status
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
ceph mgr module ls --format json 2>/dev/null | python3 -c "
|
||||
import sys, json
|
||||
data = json.load(sys.stdin)
|
||||
enabled = data.get('enabled_modules', [])
|
||||
print('enabled' if 'prometheus' in enabled else 'disabled')
|
||||
"
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: prometheus_module
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Enable prometheus mgr module
|
||||
ansible.builtin.command: ceph mgr module enable prometheus
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- "'disabled' in prometheus_module.stdout"
|
||||
changed_when: true
|
||||
|
||||
# --- Step 2: Wait for cephadm to deploy monitoring ---
|
||||
|
||||
- name: Wait for monitoring daemons to start
|
||||
ansible.builtin.shell: |
|
||||
python3 -c "
|
||||
import json, subprocess
|
||||
services = ['node-exporter', 'ceph-exporter', 'prometheus',
|
||||
'alertmanager', 'grafana']
|
||||
for svc in services:
|
||||
out = subprocess.check_output(
|
||||
['ceph', 'orch', 'ls', '--service-type', svc,
|
||||
'--format', 'json'],
|
||||
stderr=subprocess.DEVNULL)
|
||||
data = json.loads(out)
|
||||
running = sum(
|
||||
s.get('status', {}).get('running', 0) for s in data)
|
||||
if running == 0:
|
||||
print(f'waiting:{svc}')
|
||||
exit(1)
|
||||
print('all_running')
|
||||
" 2>/dev/null
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: monitoring_wait
|
||||
until: "'all_running' in monitoring_wait.stdout"
|
||||
retries: 30
|
||||
delay: 10
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
# --- Step 3: Configure dashboard integrations ---
|
||||
|
||||
- name: Configure dashboard Prometheus URL
|
||||
ansible.builtin.command: >
|
||||
ceph dashboard set-prometheus-api-host
|
||||
http://{{ hostvars[groups['ceph_bootstrap'][0]]['bond_ip'] }}:{{ ceph_prometheus_port }}
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Configure dashboard Alertmanager URL
|
||||
ansible.builtin.command: >
|
||||
ceph dashboard set-alertmanager-api-host
|
||||
http://{{ hostvars[groups['ceph_bootstrap'][0]]['bond_ip'] }}:{{ ceph_alertmanager_port }}
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Configure dashboard Grafana URL
|
||||
ansible.builtin.command: >
|
||||
ceph dashboard set-grafana-api-url
|
||||
https://{{ hostvars[groups['ceph_bootstrap'][0]]['bond_ip'] }}:{{ ceph_grafana_port }}
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Disable Grafana SSL cert verification in dashboard
|
||||
ansible.builtin.command: >
|
||||
ceph dashboard set-grafana-api-ssl-verify false
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
# --- Step 4: Set Grafana admin credentials ---
|
||||
|
||||
- name: Set Grafana admin user
|
||||
ansible.builtin.command: >
|
||||
ceph dashboard set-grafana-api-username {{ ceph_grafana_admin_user }}
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Set Grafana admin password
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
echo '{{ ceph_grafana_admin_password }}' \
|
||||
| ceph dashboard set-grafana-api-password -i -
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
no_log: true
|
||||
|
||||
# --- Step 5: Display monitoring endpoints ---
|
||||
|
||||
- name: Show monitoring status
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
echo "=== Monitoring Services ==="
|
||||
ceph orch ls --service-type prometheus
|
||||
ceph orch ls --service-type grafana
|
||||
ceph orch ls --service-type alertmanager
|
||||
ceph orch ls --service-type node-exporter
|
||||
ceph orch ls --service-type ceph-exporter
|
||||
echo ""
|
||||
echo "=== Endpoints ==="
|
||||
echo "Prometheus: http://{{ bond_ip }}:{{ ceph_prometheus_port }}"
|
||||
echo "Grafana: https://{{ bond_ip }}:{{ ceph_grafana_port }}"
|
||||
echo "Alertmanager: http://{{ bond_ip }}:{{ ceph_alertmanager_port }}"
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: monitoring_status
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Display monitoring status
|
||||
ansible.builtin.debug:
|
||||
msg: "{{ monitoring_status.stdout_lines }}"
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
@@ -0,0 +1,183 @@
|
||||
---
|
||||
# Phase 5: Create OSDs via cephadm orch service spec
|
||||
#
|
||||
# Renders a multi-document OSD service spec (one document per host, see
|
||||
# templates/osd-spec.yml.j2) and applies via `ceph orch apply osd`. Cephadm
|
||||
# handles per-disk LUKS provisioning, LVM creation, daemon deployment, and
|
||||
# CRUSH placement internally — the role no longer iterates disks.
|
||||
#
|
||||
# Hardware-shape independence: the template's Jinja conditionals handle
|
||||
# both sietch (SAS expander + dual-SSD-VG) and painbox (PCI-ATA + single
|
||||
# NVMe-RAID VG + LV-backed SSD OSD) without per-shape branching here.
|
||||
#
|
||||
# Required host_vars (per host, in inventories/<cluster>/host_vars/<host>.yml):
|
||||
# ceph_hdd_osds — list of {path_phy, db} mappings per HDD
|
||||
# ceph_ssd_osds — list of {path_phy + partition} (sietch) or {lv} (painbox)
|
||||
# sas_path_prefix — REQUIRED for sietch; OMITTED for painbox-shape hosts
|
||||
#
|
||||
# Idempotent: re-applying the same spec is a no-op when OSDs match. New
|
||||
# disks (e.g. populating an empty bay later) are picked up automatically.
|
||||
|
||||
- name: Render OSD service spec on bootstrap node
|
||||
ansible.builtin.template:
|
||||
src: osd-spec.yml.j2
|
||||
dest: /etc/ceph/osd-spec.yml
|
||||
mode: '0644'
|
||||
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
|
||||
run_once: true
|
||||
register: osd_spec_render
|
||||
|
||||
- name: Show rendered OSD spec (for verification)
|
||||
ansible.builtin.command: cat /etc/ceph/osd-spec.yml
|
||||
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
|
||||
run_once: true
|
||||
changed_when: false
|
||||
register: osd_spec_content
|
||||
|
||||
- name: Display rendered OSD spec
|
||||
ansible.builtin.debug:
|
||||
msg: "{{ osd_spec_content.stdout_lines }}"
|
||||
run_once: true
|
||||
|
||||
- name: Apply OSD service spec via cephadm orch
|
||||
ansible.builtin.command: ceph orch apply osd -i /etc/ceph/osd-spec.yml
|
||||
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
|
||||
run_once: true
|
||||
register: orch_apply
|
||||
changed_when: >-
|
||||
'Scheduled' in (orch_apply.stdout | default(''))
|
||||
or 'created' in (orch_apply.stdout | default('') | lower)
|
||||
|
||||
- name: Wait for cephadm to provision all expected OSDs
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
EXPECTED={{
|
||||
groups['ceph_nodes']
|
||||
| map('extract', hostvars)
|
||||
| map(attribute='ceph_hdd_osds', default=[])
|
||||
| map('length') | sum
|
||||
+
|
||||
groups['ceph_nodes']
|
||||
| map('extract', hostvars)
|
||||
| map(attribute='ceph_ssd_osds', default=[])
|
||||
| map('length') | sum
|
||||
}}
|
||||
echo "Expecting $EXPECTED OSDs total across the cluster"
|
||||
for i in $(seq 1 90); do
|
||||
ACTUAL=$(ceph osd stat --format json 2>/dev/null \
|
||||
| python3 -c "import sys, json; print(json.load(sys.stdin).get('num_osds', 0))" \
|
||||
2>/dev/null || echo 0)
|
||||
echo " attempt $i: $ACTUAL/$EXPECTED OSDs created"
|
||||
if [ "$ACTUAL" -ge "$EXPECTED" ]; then
|
||||
echo "All expected OSDs provisioned"
|
||||
exit 0
|
||||
fi
|
||||
sleep 10
|
||||
done
|
||||
echo "WARNING: only $ACTUAL of $EXPECTED OSDs provisioned after 15 minutes"
|
||||
exit 1
|
||||
args:
|
||||
executable: /bin/bash
|
||||
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
|
||||
run_once: true
|
||||
changed_when: false
|
||||
timeout: 960
|
||||
|
||||
- name: Wait for OSDs to come up
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
for i in $(seq 1 90); do
|
||||
up=$(ceph osd stat --format json 2>/dev/null |
|
||||
python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('num_up_osds',0))" \
|
||||
2>/dev/null || echo 0)
|
||||
total=$(ceph osd stat --format json 2>/dev/null |
|
||||
python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('num_osds',0))" \
|
||||
2>/dev/null || echo 0)
|
||||
echo "OSDs: $up/$total up"
|
||||
if [ "$up" -eq "$total" ] && [ "$total" -gt 0 ]; then
|
||||
echo "All OSDs up"; exit 0
|
||||
fi
|
||||
sleep 10
|
||||
done
|
||||
echo "WARNING: Not all OSDs up after 15 minutes"
|
||||
exit 1
|
||||
args:
|
||||
executable: /bin/bash
|
||||
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
|
||||
run_once: true
|
||||
changed_when: false
|
||||
timeout: 960
|
||||
|
||||
# Safety net: if any OSD comes up with reweight=0 (rare with orch apply,
|
||||
# but possible if a noin flag was set externally before this run), fix it.
|
||||
|
||||
- name: Fix any OSDs stuck at reweight 0
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
FIXED=0
|
||||
for osd_id in $(ceph osd tree --format json 2>/dev/null | python3 -c "
|
||||
import sys, json
|
||||
tree = json.load(sys.stdin)
|
||||
for node in tree.get('nodes', []):
|
||||
if node.get('type') == 'osd' and node.get('status') == 'up' \
|
||||
and node.get('reweight', 1) == 0:
|
||||
print(node['id'])
|
||||
"); do
|
||||
echo "Reweighting osd.$osd_id from 0 to 1.0"
|
||||
ceph osd reweight "$osd_id" 1.0
|
||||
FIXED=$((FIXED + 1))
|
||||
done
|
||||
if [ "$FIXED" -eq 0 ]; then
|
||||
echo "All up OSDs have proper reweight"
|
||||
else
|
||||
echo "Fixed $FIXED OSDs with reweight=0"
|
||||
fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
|
||||
run_once: true
|
||||
register: reweight_fix
|
||||
changed_when: "'Fixed' in reweight_fix.stdout"
|
||||
|
||||
- name: Show OSD tree
|
||||
ansible.builtin.command: ceph osd tree
|
||||
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
|
||||
run_once: true
|
||||
changed_when: false
|
||||
register: osd_tree
|
||||
|
||||
- name: Display OSD tree
|
||||
ansible.builtin.debug:
|
||||
msg: "{{ osd_tree.stdout_lines }}"
|
||||
run_once: true
|
||||
|
||||
- name: Verify dmcrypt keys stored in MONs
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
ceph config-key dump 2>/dev/null | grep -c dm-crypt
|
||||
args:
|
||||
executable: /bin/bash
|
||||
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
|
||||
run_once: true
|
||||
changed_when: false
|
||||
register: dmcrypt_keys
|
||||
|
||||
- name: Show dmcrypt key count
|
||||
ansible.builtin.debug:
|
||||
msg: >-
|
||||
{{ dmcrypt_keys.stdout }} dm-crypt keys stored
|
||||
in MON config-key database
|
||||
run_once: true
|
||||
|
||||
# Defensive: clear the noin flag if it's set. The current spec-apply flow
|
||||
# doesn't set noin (cephadm orch handles backfill rollout gracefully), but
|
||||
# a prior failed run of this role's older imperative OSD-creation flow may
|
||||
# have set it and not unset it. Unset is idempotent — no-op when already off.
|
||||
|
||||
- name: Defensive — clear noin flag if set
|
||||
ansible.builtin.command: ceph osd unset noin
|
||||
delegate_to: "{{ groups['ceph_bootstrap'][0] }}"
|
||||
run_once: true
|
||||
register: noin_unset
|
||||
changed_when: "'noin is unset' in (noin_unset.stdout | default(''))"
|
||||
failed_when: false
|
||||
@@ -0,0 +1,91 @@
|
||||
---
|
||||
# Phase 4: Configure MON and MGR placement
|
||||
#
|
||||
# MON placement strategy:
|
||||
# 1. If ceph_mon group is defined and non-empty: use those hosts (explicit)
|
||||
# 2. Else if <=2 active hosts: 1 MON (bootstrap only, avoids 2-MON fragility)
|
||||
# 3. Else: MON on all active hosts
|
||||
#
|
||||
# MGR always deploys on all active hosts (standbys are harmless).
|
||||
|
||||
- name: Determine active cluster hosts
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
ceph orch host ls --format json | python3 -c "
|
||||
import sys, json
|
||||
hosts = json.load(sys.stdin)
|
||||
print(','.join(h['hostname'] for h in hosts if h.get('status','') == ''))
|
||||
"
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: active_hosts
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Calculate MON placement
|
||||
ansible.builtin.set_fact:
|
||||
mon_hosts: >-
|
||||
{%- if groups['ceph_mon'] | default([]) | length > 0 -%}
|
||||
{{ groups['ceph_mon']
|
||||
| map('extract', hostvars, 'hostname_short')
|
||||
| join(',') -}}
|
||||
{%- elif active_hosts.stdout.split(',') | length <= 2 -%}
|
||||
{{ hostvars[groups['ceph_bootstrap'][0]]['hostname_short'] -}}
|
||||
{%- else -%}
|
||||
{{ active_hosts.stdout -}}
|
||||
{%- endif -%}
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- active_hosts.stdout | default('') | length > 0
|
||||
|
||||
- name: Show placement decision
|
||||
ansible.builtin.debug:
|
||||
msg: >-
|
||||
Active hosts: {{ active_hosts.stdout }}.
|
||||
MON placement: {{ mon_hosts }}
|
||||
({{ 'explicit ceph_mon group'
|
||||
if groups['ceph_mon'] | default([]) | length > 0
|
||||
else ('single MON — avoids 2-mon quorum fragility'
|
||||
if active_hosts.stdout.split(',') | length <= 2
|
||||
else 'MON on all hosts') }}).
|
||||
MGR placement: {{ active_hosts.stdout }} (all hosts).
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- mon_hosts is defined
|
||||
|
||||
- name: Deploy MON on calculated hosts
|
||||
ansible.builtin.command: >
|
||||
ceph orch apply mon --placement="{{ mon_hosts }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- mon_hosts | default('') | length > 0
|
||||
changed_when: false
|
||||
|
||||
- name: Deploy MGR on all active hosts
|
||||
ansible.builtin.command: >
|
||||
ceph orch apply mgr --placement="{{ active_hosts.stdout }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- active_hosts.stdout | default('') | length > 0
|
||||
changed_when: false
|
||||
|
||||
- name: Wait for MONs to be ready
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
for i in $(seq 1 30); do
|
||||
mon_count=$(ceph mon stat --format json 2>/dev/null |
|
||||
python3 -c "import sys,json; print(json.load(sys.stdin).get('num_mons',0))" \
|
||||
2>/dev/null || echo 0)
|
||||
expected={{ mon_hosts.split(',') | length }}
|
||||
if [ "$mon_count" -ge "$expected" ]; then
|
||||
echo "All MONs ready ($mon_count/$expected)"
|
||||
exit 0
|
||||
fi
|
||||
echo "Waiting for MONs... ($mon_count/$expected)"
|
||||
sleep 10
|
||||
done
|
||||
echo "WARNING: Not all MONs ready, continuing"
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
@@ -0,0 +1,56 @@
|
||||
---
|
||||
# Phase 1: Prerequisites (all nodes)
|
||||
|
||||
# /etc/hosts is managed by the baseline role.
|
||||
|
||||
- name: Ensure keyrings directory exists
|
||||
ansible.builtin.file:
|
||||
path: /etc/apt/keyrings
|
||||
state: directory
|
||||
mode: '0755'
|
||||
|
||||
- name: Add Ceph release GPG key
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
curl -fsSL {{ ceph_repo_key_url }} | gpg --dearmor -o /etc/apt/keyrings/ceph.gpg
|
||||
args:
|
||||
creates: /etc/apt/keyrings/ceph.gpg
|
||||
executable: /bin/bash
|
||||
|
||||
- name: Add Ceph Tentacle repository
|
||||
ansible.builtin.copy:
|
||||
content: "deb [signed-by=/etc/apt/keyrings/ceph.gpg] {{ ceph_repo_url }} bookworm main\n"
|
||||
dest: /etc/apt/sources.list.d/ceph.list
|
||||
mode: '0644'
|
||||
|
||||
- name: Update apt cache
|
||||
# Always refresh — no cache_valid_time. Hetzner installimage runs apt
|
||||
# during install, leaving the cache "fresh" (~minutes old) but without the
|
||||
# Ceph repo's Packages file. cache_valid_time=3600 here would skip refresh
|
||||
# and apt install would only see Debian's older cephadm (16.x) → version
|
||||
# pin fails. Always-refresh is the safe default for any new repo addition.
|
||||
ansible.builtin.apt:
|
||||
update_cache: true
|
||||
|
||||
# Podman ecosystem, diagnostic tools, and service enablement are
|
||||
# handled by the baseline role. This task installs only Ceph-specific
|
||||
# packages that baseline doesn't cover.
|
||||
|
||||
- name: Install cephadm and Ceph-specific dependencies
|
||||
ansible.builtin.apt:
|
||||
name:
|
||||
- "cephadm={{ ceph_release_version }}"
|
||||
- "ceph-common={{ ceph_release_version }}"
|
||||
- python3-asyncssh
|
||||
state: present
|
||||
|
||||
# dbus, chrony, and podman.socket are enabled by the baseline role.
|
||||
|
||||
- name: Disable sntrup761 post-quantum kex (hangs with asyncssh/cephadm)
|
||||
ansible.builtin.copy:
|
||||
# yamllint disable-line rule:line-length
|
||||
content: >-
|
||||
KexAlgorithms curve25519-sha256,curve25519-sha256@libssh.org,ecdh-sha2-nistp256,ecdh-sha2-nistp384,ecdh-sha2-nistp521,diffie-hellman-group-exchange-sha256,diffie-hellman-group16-sha512,diffie-hellman-group18-sha512,diffie-hellman-group14-sha256
|
||||
dest: /etc/ssh/sshd_config.d/no-sntrup.conf
|
||||
mode: '0644'
|
||||
notify: Restart sshd
|
||||
@@ -0,0 +1,658 @@
|
||||
---
|
||||
# Phase 5.75: RadosGW (S3 Object Gateway)
|
||||
# Deploys RGW with EC data pool, realm/zonegroup/zone, and S3 service user
|
||||
|
||||
# --- Step 1: EC profile ---
|
||||
|
||||
- name: Check existing EC profiles
|
||||
ansible.builtin.command: ceph osd erasure-code-profile ls
|
||||
register: ec_profiles
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Create EC profile {{ ceph_rgw_ec_profile }}
|
||||
ansible.builtin.command: >
|
||||
ceph osd erasure-code-profile set {{ ceph_rgw_ec_profile }}
|
||||
k={{ ceph_rgw_ec_k }} m={{ ceph_rgw_ec_m }}
|
||||
crush-failure-domain={{ ceph_rgw_ec_failure_domain }}
|
||||
crush-device-class={{ ceph_rgw_ec_device_class }}
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- "ceph_rgw_ec_profile not in ec_profiles.stdout"
|
||||
changed_when: true
|
||||
|
||||
# --- Step 2-3: EC data pool ---
|
||||
|
||||
- name: Check existing pools
|
||||
ansible.builtin.command: ceph osd pool ls
|
||||
register: pool_list
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Create EC data pool {{ ceph_rgw_data_pool }}
|
||||
ansible.builtin.command: >
|
||||
ceph osd pool create {{ ceph_rgw_data_pool }} erasure {{ ceph_rgw_ec_profile }}
|
||||
register: create_data_pool
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- "ceph_rgw_data_pool not in pool_list.stdout"
|
||||
changed_when: true
|
||||
|
||||
- name: Enable EC overwrites on data pool
|
||||
ansible.builtin.command: >
|
||||
ceph osd pool set {{ ceph_rgw_data_pool }} allow_ec_overwrites true
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- create_data_pool is changed
|
||||
changed_when: true
|
||||
|
||||
# --- Step 4: Replicated index pool ---
|
||||
|
||||
- name: Create replicated index pool {{ ceph_rgw_index_pool }}
|
||||
ansible.builtin.command: >
|
||||
ceph osd pool create {{ ceph_rgw_index_pool }} replicated
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- "ceph_rgw_index_pool not in pool_list.stdout"
|
||||
changed_when: true
|
||||
|
||||
- name: Set index pool replication parameters
|
||||
ansible.builtin.shell: |
|
||||
set -e
|
||||
ceph osd pool set {{ ceph_rgw_index_pool }} size {{ ceph_rgw_replicated_size }}
|
||||
ceph osd pool set {{ ceph_rgw_index_pool }} min_size {{ ceph_rgw_replicated_min_size }}
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
# --- Step 5: Replicated non-EC pool (multipart uploads) ---
|
||||
|
||||
- name: Create replicated non-EC pool {{ ceph_rgw_extra_pool }}
|
||||
ansible.builtin.command: >
|
||||
ceph osd pool create {{ ceph_rgw_extra_pool }} replicated
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- "ceph_rgw_extra_pool not in pool_list.stdout"
|
||||
changed_when: true
|
||||
|
||||
- name: Set non-EC pool replication parameters
|
||||
ansible.builtin.shell: |
|
||||
set -e
|
||||
ceph osd pool set {{ ceph_rgw_extra_pool }} size {{ ceph_rgw_replicated_size }}
|
||||
ceph osd pool set {{ ceph_rgw_extra_pool }} min_size {{ ceph_rgw_replicated_min_size }}
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
# --- Step 6: Pool application tags ---
|
||||
|
||||
- name: Check pool application tags
|
||||
ansible.builtin.command: "ceph osd pool application get {{ item }}"
|
||||
loop:
|
||||
- "{{ ceph_rgw_data_pool }}"
|
||||
- "{{ ceph_rgw_index_pool }}"
|
||||
- "{{ ceph_rgw_extra_pool }}"
|
||||
register: pool_apps
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
|
||||
- name: Enable rgw application on pools
|
||||
ansible.builtin.command: "ceph osd pool application enable {{ item.item }} rgw"
|
||||
loop: "{{ pool_apps.results }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- item is not skipped
|
||||
- "'rgw' not in (item.stdout | default(''))"
|
||||
changed_when: true
|
||||
loop_control:
|
||||
label: "{{ item.item | default('skipped') }}"
|
||||
|
||||
# --- Step 6.5: Initial PG counts ---
|
||||
#
|
||||
# Pools are created with pg_num=1 (ceph default). The autoscaler will
|
||||
# eventually grow them, but splitting PGs under load causes latency
|
||||
# spikes and throughput drops. Pre-sizing avoids this penalty during
|
||||
# the first fill. Idempotent: only increases pg_num, never decreases.
|
||||
|
||||
- name: Set initial PG count on data pool
|
||||
ansible.builtin.command: >
|
||||
ceph osd pool set {{ ceph_rgw_data_pool }}
|
||||
pg_num {{ ceph_pg_init_data }}
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_rgw_data_pool in pool_list.stdout
|
||||
register: pg_data_result
|
||||
changed_when: "'set' in pg_data_result.stdout | default('')"
|
||||
failed_when:
|
||||
- pg_data_result.rc != 0
|
||||
- "'is not >= current' not in pg_data_result.stderr | default('')"
|
||||
|
||||
- name: Set initial PG count on index pool
|
||||
ansible.builtin.command: >
|
||||
ceph osd pool set {{ ceph_rgw_index_pool }}
|
||||
pg_num {{ ceph_pg_init_index }}
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_rgw_index_pool in pool_list.stdout
|
||||
register: pg_index_result
|
||||
changed_when: "'set' in pg_index_result.stdout | default('')"
|
||||
failed_when:
|
||||
- pg_index_result.rc != 0
|
||||
- "'is not >= current' not in pg_index_result.stderr | default('')"
|
||||
|
||||
- name: Set initial PG count on non-EC pool
|
||||
ansible.builtin.command: >
|
||||
ceph osd pool set {{ ceph_rgw_extra_pool }}
|
||||
pg_num {{ ceph_pg_init_non_ec }}
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_rgw_extra_pool in pool_list.stdout
|
||||
register: pg_extra_result
|
||||
changed_when: "'set' in pg_extra_result.stdout | default('')"
|
||||
failed_when:
|
||||
- pg_extra_result.rc != 0
|
||||
- "'is not >= current' not in pg_extra_result.stderr | default('')"
|
||||
|
||||
- name: Set initial PG count on RGW metadata pools
|
||||
ansible.builtin.command: >
|
||||
ceph osd pool set {{ item }} pg_num {{ ceph_pg_init_meta }}
|
||||
loop:
|
||||
- .rgw.root
|
||||
- "{{ ceph_rgw_zone }}.rgw.log"
|
||||
- "{{ ceph_rgw_zone }}.rgw.control"
|
||||
- "{{ ceph_rgw_zone }}.rgw.meta"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
# Skip pools that haven't materialized yet — RGW creates them
|
||||
# lazily on first daemon startup. Mirrors the per-pool gate used
|
||||
# by the data/index/extra pool tasks above.
|
||||
- item in pool_list.stdout
|
||||
register: pg_meta_result
|
||||
changed_when: "'set' in pg_meta_result.stdout | default('')"
|
||||
failed_when:
|
||||
- pg_meta_result.rc is defined
|
||||
- pg_meta_result.rc != 0
|
||||
- "'is not >= current' not in pg_meta_result.stderr | default('')"
|
||||
|
||||
- name: Wait for PG peering to complete
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
for i in $(seq 1 60); do
|
||||
STATE=$(ceph pg stat --format json 2>/dev/null | python3 -c "
|
||||
import sys, json
|
||||
d = json.load(sys.stdin)
|
||||
s = d.get('pg_summary', {}).get('num_pg_by_state', [])
|
||||
non_clean = sum(x['num'] for x in s if 'active+clean' not in x['name'])
|
||||
print(non_clean)
|
||||
" 2>/dev/null || echo 999)
|
||||
if [ "$STATE" = "0" ]; then
|
||||
echo "All PGs active+clean"
|
||||
exit 0
|
||||
fi
|
||||
echo "Waiting for PG peering... ($STATE PGs not clean)"
|
||||
sleep 5
|
||||
done
|
||||
echo "WARNING: PGs still peering after 5 minutes"
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
timeout: 330
|
||||
|
||||
# --- Step 7: Realm ---
|
||||
|
||||
- name: Check existing realms
|
||||
ansible.builtin.command: radosgw-admin realm list
|
||||
register: realm_list
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
|
||||
- name: Create realm {{ ceph_rgw_realm }}
|
||||
ansible.builtin.command: >
|
||||
radosgw-admin realm create --rgw-realm={{ ceph_rgw_realm }} --default
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- "ceph_rgw_realm not in (realm_list.stdout | default(''))"
|
||||
changed_when: true
|
||||
|
||||
- name: Ensure realm is default
|
||||
ansible.builtin.command: >
|
||||
radosgw-admin realm default --rgw-realm={{ ceph_rgw_realm }}
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
# --- Step 8: Zonegroup ---
|
||||
|
||||
- name: Check existing zonegroups
|
||||
ansible.builtin.command: radosgw-admin zonegroup list
|
||||
register: zonegroup_list
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
|
||||
- name: Create zonegroup {{ ceph_rgw_zonegroup }}
|
||||
ansible.builtin.command: >
|
||||
radosgw-admin zonegroup create
|
||||
--rgw-realm={{ ceph_rgw_realm }}
|
||||
--rgw-zonegroup={{ ceph_rgw_zonegroup }}
|
||||
--endpoints={{ ceph_rgw_scheme }}://{{ ceph_rgw_dns_name }}:{{ ceph_rgw_port }}
|
||||
--master --default
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- "ceph_rgw_zonegroup not in (zonegroup_list.stdout | default(''))"
|
||||
changed_when: true
|
||||
|
||||
- name: Set zonegroup api_name to {{ ceph_rgw_zonegroup_api_name }}
|
||||
ansible.builtin.shell: |
|
||||
set -e
|
||||
python3 -c "
|
||||
import json, subprocess, sys
|
||||
zg = json.loads(subprocess.check_output(
|
||||
['radosgw-admin', 'zonegroup', 'get',
|
||||
'--rgw-zonegroup', '{{ ceph_rgw_zonegroup }}']))
|
||||
if zg.get('api_name') == '{{ ceph_rgw_zonegroup_api_name }}':
|
||||
sys.exit(0)
|
||||
zg['api_name'] = '{{ ceph_rgw_zonegroup_api_name }}'
|
||||
subprocess.run(
|
||||
['radosgw-admin', 'zonegroup', 'set',
|
||||
'--rgw-zonegroup', '{{ ceph_rgw_zonegroup }}'],
|
||||
input=json.dumps(zg).encode(), check=True)
|
||||
print('CHANGED')
|
||||
"
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: set_api_name
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: "'CHANGED' in (set_api_name.stdout | default(''))"
|
||||
|
||||
# --- Step 9: Zone ---
|
||||
|
||||
- name: Check existing zones
|
||||
ansible.builtin.command: radosgw-admin zone list
|
||||
register: zone_list
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
|
||||
- name: Create zone {{ ceph_rgw_zone }}
|
||||
ansible.builtin.command: >
|
||||
radosgw-admin zone create
|
||||
--rgw-realm={{ ceph_rgw_realm }}
|
||||
--rgw-zonegroup={{ ceph_rgw_zonegroup }}
|
||||
--rgw-zone={{ ceph_rgw_zone }}
|
||||
--endpoints={{ ceph_rgw_scheme }}://{{ ceph_rgw_dns_name }}:{{ ceph_rgw_port }}
|
||||
--master --default
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- "ceph_rgw_zone not in (zone_list.stdout | default(''))"
|
||||
changed_when: true
|
||||
|
||||
# --- Step 10: Zone placement targets ---
|
||||
|
||||
- name: Configure zone placement targets for EC pools
|
||||
ansible.builtin.shell: |
|
||||
set -e
|
||||
python3 -c "
|
||||
import json, subprocess, sys
|
||||
zone = json.loads(subprocess.check_output(
|
||||
['radosgw-admin', 'zone', 'get', '--rgw-zone', '{{ ceph_rgw_zone }}']))
|
||||
changed = False
|
||||
for pt in zone.get('placement_targets', []):
|
||||
if pt['name'] == 'default-placement':
|
||||
if (pt.get('data_pool') != '{{ ceph_rgw_data_pool }}' or
|
||||
pt.get('index_pool') != '{{ ceph_rgw_index_pool }}' or
|
||||
pt.get('data_extra_pool') != '{{ ceph_rgw_extra_pool }}'):
|
||||
pt['data_pool'] = '{{ ceph_rgw_data_pool }}'
|
||||
pt['index_pool'] = '{{ ceph_rgw_index_pool }}'
|
||||
pt['data_extra_pool'] = '{{ ceph_rgw_extra_pool }}'
|
||||
changed = True
|
||||
if not changed:
|
||||
sys.exit(0)
|
||||
subprocess.run(
|
||||
['radosgw-admin', 'zone', 'set', '--rgw-zone', '{{ ceph_rgw_zone }}'],
|
||||
input=json.dumps(zone).encode(), check=True)
|
||||
print('CHANGED')
|
||||
"
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: zone_placement
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: "'CHANGED' in (zone_placement.stdout | default(''))"
|
||||
|
||||
# --- Step 10.5: Zonegroup hostnames ---
|
||||
#
|
||||
# The dashboard connects to RGW using the node's hostname or IP in the
|
||||
# Host header. RGW validates the Host header during S3 signature
|
||||
# verification — if the hostname isn't in the zonegroup's hostnames
|
||||
# list, signature verification fails with SignatureDoesNotMatch (403).
|
||||
# Adding all node hostnames and IPs ensures both dashboard and direct
|
||||
# client access work regardless of which address is used.
|
||||
|
||||
- name: Add cluster hostnames and IPs to zonegroup
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
python3 -c "
|
||||
import json, subprocess
|
||||
zg = json.loads(subprocess.check_output(
|
||||
['radosgw-admin', 'zonegroup', 'get',
|
||||
'--rgw-zonegroup', '{{ ceph_rgw_zonegroup }}']))
|
||||
hostnames = set(zg.get('hostnames', []))
|
||||
needed = set([
|
||||
{% for h in groups['ceph_nodes'] %}
|
||||
'{{ hostvars[h]['hostname_short'] }}',
|
||||
'{{ hostvars[h]['bond_ip'] }}',
|
||||
{% endfor %}
|
||||
'{{ ceph_rgw_dns_name }}',
|
||||
])
|
||||
if needed.issubset(hostnames):
|
||||
print('OK')
|
||||
else:
|
||||
hostnames.update(needed)
|
||||
zg['hostnames'] = sorted(hostnames)
|
||||
subprocess.run(
|
||||
['radosgw-admin', 'zonegroup', 'set',
|
||||
'--rgw-zonegroup', '{{ ceph_rgw_zonegroup }}'],
|
||||
input=json.dumps(zg).encode(), check=True,
|
||||
capture_output=True)
|
||||
print('CHANGED: ' + ','.join(sorted(needed)))
|
||||
"
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
register: zonegroup_hostnames
|
||||
changed_when: "'CHANGED' in zonegroup_hostnames.stdout"
|
||||
|
||||
# --- Step 11: Commit period ---
|
||||
|
||||
- name: Commit the period
|
||||
ansible.builtin.command: >
|
||||
radosgw-admin period update --commit --rgw-realm={{ ceph_rgw_realm }}
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
# --- Step 11.5: rgw_dns_name for virtual-hosted bucket addressing ---
|
||||
#
|
||||
# With this set, a request to bucket.{{ ceph_rgw_dns_name }} is
|
||||
# recognized as the bucket "bucket" rather than a literal bucket whose
|
||||
# name is the full FQDN. Clients can then use either path-style
|
||||
# ({{ ceph_rgw_scheme }}://{{ ceph_rgw_dns_name }}/bucket/key) or
|
||||
# virtual-hosted style ({{ ceph_rgw_scheme }}://bucket.{{ ceph_rgw_dns_name }}/key).
|
||||
|
||||
- name: Get current rgw_dns_name
|
||||
ansible.builtin.command: ceph config get client.rgw rgw_dns_name
|
||||
register: current_rgw_dns_name
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
|
||||
- name: Set rgw_dns_name for virtual-hosted bucket support
|
||||
ansible.builtin.command: >
|
||||
ceph config set client.rgw rgw_dns_name {{ ceph_rgw_dns_name }}
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- "(current_rgw_dns_name.stdout | default('')) | trim != ceph_rgw_dns_name"
|
||||
changed_when: true
|
||||
|
||||
# --- Step 11.6: Self-signed TLS cert for RGW frontend ---
|
||||
#
|
||||
# Generated on the bootstrap node via openssl, handed to cephadm via the
|
||||
# service spec template below. cephadm distributes the combined PEM to
|
||||
# every RGW daemon container on apply.
|
||||
#
|
||||
# Idempotent via `creates:` — the openssl run only fires if the cert file
|
||||
# is absent. To rotate: delete /etc/ceph/rgw-ssl.{crt,key} on the bootstrap
|
||||
# node and re-run the role.
|
||||
#
|
||||
# SANs include:
|
||||
# - the canonical DNS name (CN)
|
||||
# - a wildcard under the same name for virtual-hosted buckets
|
||||
# - per-node FQDNs so direct-host addressing also validates
|
||||
# - per-node bond IPs so IP-based S3 clients also validate
|
||||
|
||||
- name: Generate RGW self-signed cert and key (10 year validity)
|
||||
ansible.builtin.shell: |
|
||||
set -euo pipefail
|
||||
umask 077
|
||||
openssl req -x509 -newkey rsa:4096 -nodes \
|
||||
-keyout /etc/ceph/rgw-ssl.key \
|
||||
-out /etc/ceph/rgw-ssl.crt \
|
||||
-days {{ ceph_rgw_ssl_cert_days }} \
|
||||
-subj "/C={{ ceph_rgw_ssl_cert_subject_c }}\
|
||||
/ST={{ ceph_rgw_ssl_cert_subject_st }}\
|
||||
/L={{ ceph_rgw_ssl_cert_subject_l }}\
|
||||
/O={{ ceph_rgw_ssl_cert_subject_o }}\
|
||||
/CN={{ ceph_rgw_dns_name }}\
|
||||
/emailAddress={{ ceph_rgw_ssl_cert_email }}" \
|
||||
-addext "subjectAltName=\
|
||||
DNS:{{ ceph_rgw_dns_name }},\
|
||||
DNS:*.{{ ceph_rgw_dns_name }}\
|
||||
{% for h in groups['ceph_nodes'] %},DNS:{{ hostvars[h]['hostname_short'] }}.{{ cluster_domain }}{% endfor %}\
|
||||
{% for h in groups['ceph_nodes'] %},IP:{{ hostvars[h]['bond_ip'] }}{% endfor %}"
|
||||
args:
|
||||
executable: /bin/bash
|
||||
creates: /etc/ceph/rgw-ssl.crt
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_rgw_ssl | default(false) | bool
|
||||
changed_when: true
|
||||
|
||||
- name: Read combined cert + key for service spec
|
||||
ansible.builtin.shell: cat /etc/ceph/rgw-ssl.crt /etc/ceph/rgw-ssl.key
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: rgw_ssl_combined
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_rgw_ssl | default(false) | bool
|
||||
changed_when: false
|
||||
no_log: true
|
||||
|
||||
- name: Set combined PEM fact
|
||||
ansible.builtin.set_fact:
|
||||
rgw_ssl_cert_combined_pem: "{{ rgw_ssl_combined.stdout }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_rgw_ssl | default(false) | bool
|
||||
no_log: true
|
||||
|
||||
# --- Step 12: Deploy RGW daemons via cephadm service spec ---
|
||||
#
|
||||
# We render a YAML service spec rather than using `ceph orch apply rgw
|
||||
# <realm> <zone> --placement=... --port=...` because the command-line
|
||||
# form doesn't accept cert content. The YAML spec supports both
|
||||
# `ssl: true` and inline `rgw_frontend_ssl_certificate` — cephadm
|
||||
# distributes the cert to every RGW container.
|
||||
#
|
||||
# Placement is pinned to groups['ceph_nodes'] (not ansible_play_batch)
|
||||
# so a --limit re-run doesn't accidentally shrink the spec and
|
||||
# de-deploy daemons on omitted hosts (cephadm is declarative: applying
|
||||
# a smaller placement REMOVES daemons).
|
||||
|
||||
- name: Render RGW service spec
|
||||
ansible.builtin.template:
|
||||
src: rgw-spec.yaml.j2
|
||||
dest: /etc/ceph/rgw-spec.yaml
|
||||
owner: root
|
||||
group: root
|
||||
mode: '0600'
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
no_log: true
|
||||
|
||||
- name: Apply RGW service spec
|
||||
ansible.builtin.command: ceph orch apply -i /etc/ceph/rgw-spec.yaml
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
# --- Step 12.5: Dashboard RGW SSL verify ---
|
||||
#
|
||||
# With self-signed certs the dashboard's RGW client rejects the cert
|
||||
# and returns 500 on the Object Gateway page. Disable verification so
|
||||
# the dashboard can talk to the local RGW endpoints.
|
||||
|
||||
- name: Disable dashboard RGW API SSL verification (self-signed cert)
|
||||
ansible.builtin.command: ceph dashboard set-rgw-api-ssl-verify false
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_rgw_ssl | default(false) | bool
|
||||
changed_when: false
|
||||
|
||||
- name: Sync dashboard RGW credentials
|
||||
ansible.builtin.command: ceph dashboard set-rgw-credentials
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
# --- Step 13: Wait for RGW daemons ---
|
||||
|
||||
- name: Wait for RGW daemons to start
|
||||
ansible.builtin.shell: |
|
||||
set -eo pipefail
|
||||
ceph orch ls --service-type rgw --format json | python3 -c "
|
||||
import sys, json
|
||||
svcs = json.load(sys.stdin)
|
||||
running = sum(s.get('status', {}).get('running', 0) for s in svcs)
|
||||
print(running)
|
||||
"
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: rgw_running_count
|
||||
until: "(rgw_running_count.stdout | int) >= (groups['ceph_nodes'] | length)"
|
||||
retries: 30
|
||||
delay: 10
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
# --- Step 13.5: Ensure dashboard RGW user has admin caps ---
|
||||
#
|
||||
# cephadm auto-creates a "dashboard" RGW user with system=true but
|
||||
# zero caps. Without admin caps the dashboard's Object Gateway page
|
||||
# returns 403 SignatureDoesNotMatch on every admin API call.
|
||||
|
||||
- name: Ensure dashboard RGW user has admin caps
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
CAPS=$(radosgw-admin user info --uid=dashboard --format json 2>/dev/null \
|
||||
| python3 -c "import sys,json; print(len(json.load(sys.stdin).get('caps',[])))")
|
||||
if [ "$CAPS" = "0" ]; then
|
||||
radosgw-admin caps add --uid=dashboard \
|
||||
--caps='buckets=*;users=*;usage=*;metadata=*;zone=*' >/dev/null 2>&1
|
||||
echo "CHANGED"
|
||||
else
|
||||
echo "OK"
|
||||
fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
register: dashboard_caps
|
||||
changed_when: "'CHANGED' in dashboard_caps.stdout"
|
||||
|
||||
# --- Step 14: Create S3 user ---
|
||||
|
||||
- name: Check if S3 user exists
|
||||
ansible.builtin.command: "radosgw-admin user info --uid={{ ceph_rgw_s3_user_uid }}"
|
||||
register: s3_user_check
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
failed_when: false
|
||||
|
||||
- name: Create S3 service user with predetermined keys — {{ ceph_rgw_s3_user_uid }}
|
||||
ansible.builtin.command: >
|
||||
radosgw-admin user create
|
||||
--uid={{ ceph_rgw_s3_user_uid }}
|
||||
--display-name='{{ ceph_rgw_s3_user_display_name }}'
|
||||
--access-key='{{ ceph_rgw_s3_user_access_key }}'
|
||||
--secret-key='{{ ceph_rgw_s3_user_secret_key }}'
|
||||
--max-buckets=100
|
||||
register: s3_user_create
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- s3_user_check.rc != 0
|
||||
changed_when: true
|
||||
no_log: true
|
||||
|
||||
# --- Step 15: Align existing S3 user keys with 1P (idempotent rotation) ---
|
||||
#
|
||||
# For clusters that predate the TF+1P-predetermined-keys pattern, the user
|
||||
# exists but with cephadm-generated random keys. This task detects that
|
||||
# drift and realigns — destructive for the Yucca-app side, so gated by an
|
||||
# explicit flag.
|
||||
|
||||
- name: Get current S3 user info
|
||||
ansible.builtin.command: "radosgw-admin user info --uid={{ ceph_rgw_s3_user_uid }}"
|
||||
register: s3_user_info
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
no_log: true
|
||||
|
||||
- name: Detect S3 key drift (1P vs cluster)
|
||||
ansible.builtin.set_fact:
|
||||
s3_key_drift: >-
|
||||
{{
|
||||
(s3_user_info.stdout | from_json)['keys'][0]['access_key']
|
||||
!= ceph_rgw_s3_user_access_key
|
||||
}}
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
no_log: true
|
||||
|
||||
- name: Align S3 user keys with 1P values (only when rotate_s3_keys=true)
|
||||
ansible.builtin.command: >
|
||||
radosgw-admin key create
|
||||
--uid={{ ceph_rgw_s3_user_uid }}
|
||||
--access-key='{{ ceph_rgw_s3_user_access_key }}'
|
||||
--secret-key='{{ ceph_rgw_s3_user_secret_key }}'
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- s3_key_drift | default(false)
|
||||
- rotate_s3_keys | default(false) | bool
|
||||
changed_when: true
|
||||
no_log: true
|
||||
|
||||
- name: S3 key drift notice (dry flag — no action taken)
|
||||
ansible.builtin.debug:
|
||||
msg:
|
||||
- "S3 svc-user keys differ from 1P values."
|
||||
- "To align (rotates keys, requires Yucca-app re-config): re-run with -e rotate_s3_keys=true"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- s3_key_drift | default(false)
|
||||
- not (rotate_s3_keys | default(false) | bool)
|
||||
|
||||
- name: Display S3 endpoint (credentials live in 1P, not log output)
|
||||
ansible.builtin.debug:
|
||||
msg:
|
||||
- "=== S3 endpoint for {{ ceph_rgw_s3_user_uid }} ==="
|
||||
- "Access Key: op://{{ cluster_secrets_vault }}/{{ cluster_name | upper }}_CEPH_S3_SVC_YUCCA_RESTIC_ACCESS_KEY/password"
|
||||
- "Secret Key: op://{{ cluster_secrets_vault }}/{{ cluster_name | upper }}_CEPH_S3_SVC_YUCCA_RESTIC_SECRET_KEY/password"
|
||||
- "Endpoint (DNS): {{ ceph_rgw_scheme }}://{{ ceph_rgw_dns_name }}:{{ ceph_rgw_port }}"
|
||||
- "Endpoint (direct): {{ ceph_rgw_scheme }}://{{ bond_ip }}:{{ ceph_rgw_port }}"
|
||||
- "Region: {{ ceph_rgw_zonegroup_api_name }}"
|
||||
- "Virtual-hosted: {{ ceph_rgw_scheme }}://<bucket>.{{ ceph_rgw_dns_name }}:{{ ceph_rgw_port }}"
|
||||
- "Path-style: {{ ceph_rgw_scheme }}://{{ ceph_rgw_dns_name }}:{{ ceph_rgw_port }}/<bucket>"
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
no_log: false
|
||||
|
||||
# --- Step 16: Show RGW status ---
|
||||
|
||||
- name: Show RGW status
|
||||
ansible.builtin.shell: |
|
||||
echo "=== RGW Services ==="
|
||||
ceph orch ls --service-type rgw
|
||||
echo ""
|
||||
echo "=== Realm ==="
|
||||
radosgw-admin realm list
|
||||
echo ""
|
||||
echo "=== Zonegroup ==="
|
||||
radosgw-admin zonegroup get --rgw-zonegroup={{ ceph_rgw_zonegroup }}
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: rgw_final_status
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Display RGW status
|
||||
ansible.builtin.debug:
|
||||
msg: "{{ rgw_final_status.stdout_lines }}"
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
@@ -0,0 +1,44 @@
|
||||
---
|
||||
# Phase 6: Verify cluster health
|
||||
|
||||
- name: Show cluster status
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
echo "=== Cluster Status ==="
|
||||
ceph status
|
||||
echo ""
|
||||
echo "=== OSD Tree ==="
|
||||
ceph osd tree
|
||||
echo ""
|
||||
echo "=== Cluster Capacity ==="
|
||||
ceph df
|
||||
echo ""
|
||||
echo "=== RGW Services ==="
|
||||
ceph orch ls --service-type rgw 2>/dev/null || echo "No RGW services deployed"
|
||||
echo ""
|
||||
echo "=== RGW Endpoints ==="
|
||||
radosgw-admin zonegroup get --rgw-zonegroup={{ ceph_rgw_zonegroup }} 2>/dev/null | python3 -c "
|
||||
import sys, json
|
||||
try:
|
||||
zg = json.load(sys.stdin)
|
||||
print(f\"Region (api_name): {zg.get('api_name', 'N/A')}\")
|
||||
print(f\"Endpoints: {zg.get('endpoints', [])}\")
|
||||
except: print('RGW not configured')
|
||||
" || echo "RGW not configured"
|
||||
echo ""
|
||||
echo "=== Monitoring Services ==="
|
||||
ceph orch ls --service-type prometheus 2>/dev/null || echo "No monitoring deployed"
|
||||
ceph orch ls --service-type grafana 2>/dev/null || true
|
||||
ceph orch ls --service-type alertmanager 2>/dev/null || true
|
||||
ceph orch ls --service-type node-exporter 2>/dev/null || true
|
||||
ceph orch ls --service-type ceph-exporter 2>/dev/null || true
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
register: cluster_status
|
||||
changed_when: false
|
||||
|
||||
- name: Display final cluster state
|
||||
ansible.builtin.debug:
|
||||
msg: "{{ cluster_status.stdout_lines }}"
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
@@ -0,0 +1,23 @@
|
||||
[global]
|
||||
# Initial cluster config applied during cephadm bootstrap
|
||||
osd_pool_default_size = 2
|
||||
osd_pool_default_min_size = 1
|
||||
public_network = {{ public_network }}
|
||||
cluster_network = {{ cluster_network }}
|
||||
|
||||
# Recovery/backfill throttling (small cluster, shared network)
|
||||
osd_recovery_max_active = 1
|
||||
osd_recovery_sleep_hdd = 0.1
|
||||
osd_max_backfills = 1
|
||||
|
||||
# PG autoscaler target
|
||||
mon_target_pg_per_osd = 100
|
||||
|
||||
[osd]
|
||||
# Let BlueStore use the entire block.db LV (pre-created at 240G)
|
||||
bluestore_block_db_size = 0
|
||||
|
||||
# Scrub window — 2-6 AM, deep scrub weekly
|
||||
osd_scrub_begin_hour = 2
|
||||
osd_scrub_end_hour = 6
|
||||
osd_deep_scrub_interval = 604800
|
||||
@@ -0,0 +1,62 @@
|
||||
# OSD service spec — rendered by ceph_deploy role
|
||||
# Applied via: ceph orch apply osd -i /etc/ceph/osd-spec.yml
|
||||
#
|
||||
# One document per host because per-host disk paths differ:
|
||||
# - sietch nodes: each has a unique sas_path_prefix (different SAS expander
|
||||
# address per chassis), so explicit data_devices.paths can't be shared.
|
||||
# - painbox: single host, single spec.
|
||||
#
|
||||
# Hardware-shape independence: when sas_path_prefix is defined the host is
|
||||
# sietch-shape (SAS expander, dual-SSD-VG block.db); otherwise the host is
|
||||
# painbox-shape (PCI-ATA disks, single NVMe-RAID VG, SSD OSD as an LV).
|
||||
#
|
||||
# Idempotent: cephadm no-ops when the spec matches what's deployed. New
|
||||
# disks (e.g. populating an empty bay) are picked up automatically on apply.
|
||||
#
|
||||
# Do not edit by hand. Regenerate by re-running the ceph_deploy role.
|
||||
{% for host in groups['ceph_nodes'] %}
|
||||
{% set h = hostvars[host] %}
|
||||
{% if h.ceph_hdd_osds | default([]) | length > 0 %}
|
||||
---
|
||||
service_type: osd
|
||||
service_id: {{ h.hostname_short }}-hdd
|
||||
placement:
|
||||
hosts:
|
||||
- {{ h.hostname_short }}
|
||||
spec:
|
||||
data_devices:
|
||||
paths:
|
||||
{% for hdd in h.ceph_hdd_osds %}
|
||||
{% if h.sas_path_prefix is defined %}
|
||||
- /dev/disk/by-path/{{ h.sas_path_prefix }}-{{ hdd.path_phy }}-lun-0
|
||||
{% else %}
|
||||
- /dev/disk/by-path/{{ hdd.path_phy }}
|
||||
{% endif %}
|
||||
{% endfor %}
|
||||
db_devices:
|
||||
paths:
|
||||
{% for hdd in h.ceph_hdd_osds %}
|
||||
- /dev/{{ hdd.db }}
|
||||
{% endfor %}
|
||||
encrypted: true
|
||||
{% endif %}
|
||||
{% if h.ceph_ssd_osds | default([]) | length > 0 %}
|
||||
---
|
||||
service_type: osd
|
||||
service_id: {{ h.hostname_short }}-ssd
|
||||
placement:
|
||||
hosts:
|
||||
- {{ h.hostname_short }}
|
||||
spec:
|
||||
data_devices:
|
||||
paths:
|
||||
{% for ssd in h.ceph_ssd_osds %}
|
||||
{% if ssd.lv is defined %}
|
||||
- /dev/{{ ssd.lv }}
|
||||
{% else %}
|
||||
- /dev/disk/by-path/{{ h.sas_path_prefix }}-{{ ssd.path_phy }}-lun-0-part{{ ssd.partition }}
|
||||
{% endif %}
|
||||
{% endfor %}
|
||||
encrypted: true
|
||||
{% endif %}
|
||||
{% endfor %}
|
||||
@@ -0,0 +1,21 @@
|
||||
# RGW service spec — rendered by ceph-deploy role
|
||||
# Applied via: ceph orch apply -i /etc/ceph/rgw-spec.yaml
|
||||
#
|
||||
# Do not edit by hand. Regenerate by re-running the ceph-deploy role.
|
||||
service_type: rgw
|
||||
service_id: {{ ceph_rgw_realm }}.{{ ceph_rgw_zone }}
|
||||
placement:
|
||||
hosts:
|
||||
{% for host in groups['ceph_nodes'] %}
|
||||
- {{ hostvars[host]['hostname_short'] }}
|
||||
{% endfor %}
|
||||
spec:
|
||||
rgw_realm: {{ ceph_rgw_realm }}
|
||||
rgw_zone: {{ ceph_rgw_zone }}
|
||||
rgw_frontend_port: {{ ceph_rgw_port }}
|
||||
rgw_frontend_type: beast
|
||||
{% if ceph_rgw_ssl | default(false) %}
|
||||
ssl: true
|
||||
rgw_frontend_ssl_certificate: |
|
||||
{{ rgw_ssl_cert_combined_pem | trim | indent(4, True) }}
|
||||
{% endif %}
|
||||
@@ -0,0 +1,210 @@
|
||||
---
|
||||
# Phase 5: Clean up LVM, device signatures, and leftover Ceph state on all nodes
|
||||
#
|
||||
# Order matters:
|
||||
# 1. Close dm-crypt/LUKS mappings (devices are locked until closed)
|
||||
# 2. Remove device-mapper entries
|
||||
# 3. Remove LVM volume groups and physical volumes
|
||||
# 4. Wipe device signatures (HDD and SSD OSD partitions)
|
||||
# 5. Clean stale LVM metadata files
|
||||
# 6. Reset failed systemd units
|
||||
# 7. Remove Ceph directories
|
||||
|
||||
# --- Step 1: Close dm-crypt LUKS mappings ---
|
||||
# ceph-volume --dmcrypt creates LUKS volumes on OSD devices. These persist after
|
||||
# cephadm rm-cluster and lock the underlying block devices. Must close before wipefs.
|
||||
- name: Close dm-crypt LUKS mappings
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
CLOSED=0
|
||||
TARGETS=$(dmsetup ls --target crypt 2>/dev/null | grep -v 'No devices found' || true)
|
||||
for dm in $(echo "$TARGETS" | awk 'NF {print $1}'); do
|
||||
echo "Closing LUKS: $dm"
|
||||
cryptsetup close "$dm" 2>&1 || dmsetup remove -f "$dm" 2>&1 || true
|
||||
CLOSED=$((CLOSED+1))
|
||||
done
|
||||
if [ $CLOSED -eq 0 ]; then echo "No LUKS mappings to close"; fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: luks_cleanup
|
||||
changed_when: "'Closing LUKS' in luks_cleanup.stdout"
|
||||
|
||||
# --- Step 2: Remove leftover device mapper entries ---
|
||||
- name: Remove Ceph device mapper entries
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
REMOVED=0
|
||||
for dm in $(dmsetup ls 2>/dev/null | grep -i ceph | awk '{print $1}'); do
|
||||
echo "Removing DM: $dm"
|
||||
dmsetup remove -f "$dm" 2>&1 || true
|
||||
REMOVED=$((REMOVED+1))
|
||||
done
|
||||
if [ $REMOVED -eq 0 ]; then echo "No Ceph DM entries found"; fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: dm_cleanup
|
||||
changed_when: "'Removing DM' in dm_cleanup.stdout"
|
||||
|
||||
# --- Step 3: Remove Ceph LVM volume groups ---
|
||||
# Catches both ceph-db VGs (ceph-db-rear12, ceph-db-rear13) and ceph-volume
|
||||
# created VGs (ceph-<uuid> from dmcrypt OSDs).
|
||||
- name: Remove Ceph LVM volume groups
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
REMOVED=0
|
||||
for vg in $(vgs --noheadings -o vg_name 2>/dev/null | grep -i ceph | tr -d ' '); do
|
||||
echo "Removing VG: $vg"
|
||||
vgremove -f "$vg" 2>&1 || true
|
||||
REMOVED=$((REMOVED+1))
|
||||
done
|
||||
if [ $REMOVED -eq 0 ]; then echo "No Ceph VGs found"; fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: vg_cleanup
|
||||
changed_when: "'Removing VG' in vg_cleanup.stdout"
|
||||
|
||||
# --- Step 4: Remove LVM physical volumes ---
|
||||
# After VG removal, PVs are orphaned (no VG name in pvs output). Must catch BOTH
|
||||
# PVs that still reference a ceph VG AND orphaned PVs on HDD/SSD OSD devices.
|
||||
- name: Remove Ceph LVM physical volumes (named VGs)
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
REMOVED=0
|
||||
for pv in $(pvs --noheadings -o pv_name,vg_name 2>/dev/null | grep -i ceph | awk '{print $1}'); do
|
||||
echo "Removing PV: $pv"
|
||||
pvremove -f "$pv" 2>&1 || true
|
||||
REMOVED=$((REMOVED+1))
|
||||
done
|
||||
if [ $REMOVED -eq 0 ]; then echo "No named Ceph PVs found"; fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: pv_cleanup_named
|
||||
changed_when: "'Removing PV' in pv_cleanup_named.stdout"
|
||||
|
||||
- name: Remove orphaned LVM physical volumes on HDD devices
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
REMOVED=0
|
||||
for dev in $(lsblk -dnpo NAME,ROTA,TYPE | awk '$2==1 && $3=="disk" {print $1}'); do
|
||||
if pvs "$dev" --noheadings 2>/dev/null | grep -q "$dev"; then
|
||||
echo "Removing orphaned PV: $dev"
|
||||
pvremove -f "$dev" 2>&1 || true
|
||||
REMOVED=$((REMOVED+1))
|
||||
fi
|
||||
done
|
||||
if [ $REMOVED -eq 0 ]; then echo "No orphaned HDD PVs found"; fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: pv_cleanup_hdd
|
||||
changed_when: "'Removing orphaned PV' in pv_cleanup_hdd.stdout"
|
||||
|
||||
- name: Remove orphaned LVM physical volumes on SSD OSD partitions
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
REMOVED=0
|
||||
for ssd in $(lsblk -dnpo NAME,MODEL | grep -i "{{ ssd_model_pattern }}" | awk '{print $1}'); do
|
||||
if [[ "$ssd" == *nvme* ]]; then PART="${ssd}p6"; else PART="${ssd}6"; fi
|
||||
if [ -b "$PART" ] && pvs "$PART" --noheadings 2>/dev/null | grep -q "$PART"; then
|
||||
echo "Removing orphaned PV: $PART"
|
||||
pvremove -f "$PART" 2>&1 || true
|
||||
REMOVED=$((REMOVED+1))
|
||||
fi
|
||||
done
|
||||
if [ $REMOVED -eq 0 ]; then echo "No orphaned SSD OSD PVs found"; fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: pv_cleanup_ssd
|
||||
changed_when: "'Removing orphaned PV' in pv_cleanup_ssd.stdout"
|
||||
|
||||
# --- Step 5: Wipe device signatures ---
|
||||
# Wipe ALL rotational (HDD) devices that have any filesystem/LVM/LUKS signatures.
|
||||
# Every HDD in these servers is a Ceph OSD — no ambiguity about which to wipe.
|
||||
# Only wipes devices that actually have signatures (idempotent on clean disks).
|
||||
- name: Wipe HDD device signatures
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
ZAPPED=0
|
||||
for dev in $(lsblk -dnpo NAME,ROTA,TYPE | awk '$2==1 && $3=="disk" {print $1}'); do
|
||||
SIGS=$(wipefs "$dev" 2>/dev/null | tail -n +2)
|
||||
if [ -n "$SIGS" ]; then
|
||||
echo "Wiping $dev"
|
||||
wipefs -af "$dev" 2>&1 || true
|
||||
dd if=/dev/zero of="$dev" bs=1M count=10 2>/dev/null || true
|
||||
ZAPPED=$((ZAPPED+1))
|
||||
fi
|
||||
done
|
||||
if [ $ZAPPED -eq 0 ]; then echo "No HDD devices to wipe"; fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: hdd_zap
|
||||
changed_when: "'Wiping' in hdd_zap.stdout"
|
||||
|
||||
- name: Wipe SSD OSD partition signatures (partition 6)
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
ZAPPED=0
|
||||
for ssd in $(lsblk -dnpo NAME,MODEL | grep -i "{{ ssd_model_pattern }}" | awk '{print $1}'); do
|
||||
if [[ "$ssd" == *nvme* ]]; then PART="${ssd}p6"; else PART="${ssd}6"; fi
|
||||
if [ -b "$PART" ]; then
|
||||
SIGS=$(wipefs "$PART" 2>/dev/null | tail -n +2)
|
||||
if [ -n "$SIGS" ]; then
|
||||
echo "Wiping SSD OSD partition $PART"
|
||||
wipefs -af "$PART" 2>&1 || true
|
||||
dd if=/dev/zero of="$PART" bs=1M count=10 2>/dev/null || true
|
||||
ZAPPED=$((ZAPPED+1))
|
||||
fi
|
||||
fi
|
||||
done
|
||||
if [ $ZAPPED -eq 0 ]; then echo "No SSD OSD partitions to wipe"; fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: ssd_zap
|
||||
changed_when: "'Wiping' in ssd_zap.stdout"
|
||||
|
||||
# --- Step 6: Clean stale LVM metadata ---
|
||||
# ceph-volume operations generate LVM archive/backup files that accumulate across
|
||||
# cluster lifecycles. Hundreds of ceph-* files can build up in /etc/lvm/.
|
||||
- name: Remove stale Ceph LVM archive and backup files
|
||||
ansible.builtin.shell: |
|
||||
REMOVED=0
|
||||
for dir in /etc/lvm/archive /etc/lvm/backup; do
|
||||
for f in "$dir"/ceph-*; do
|
||||
[ -f "$f" ] || continue
|
||||
rm -f "$f"
|
||||
REMOVED=$((REMOVED+1))
|
||||
done
|
||||
done
|
||||
echo "Removed $REMOVED stale LVM metadata files"
|
||||
register: lvm_meta_cleanup
|
||||
changed_when: "lvm_meta_cleanup.stdout is not search('Removed 0')"
|
||||
|
||||
# --- Step 7: Reset failed Ceph systemd units ---
|
||||
- name: Reset failed Ceph systemd units
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
RESET=0
|
||||
for unit in $(systemctl list-units --all --failed --no-legend 2>/dev/null | grep -i ceph | awk '{print $1}'); do
|
||||
echo "Resetting failed unit: $unit"
|
||||
systemctl reset-failed "$unit" 2>&1 || true
|
||||
RESET=$((RESET+1))
|
||||
done
|
||||
if [ $RESET -eq 0 ]; then echo "No failed Ceph units to reset"; fi
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: systemd_cleanup
|
||||
changed_when: "'Resetting failed unit' in systemd_cleanup.stdout"
|
||||
|
||||
# --- Step 8: Remove Ceph directories ---
|
||||
- name: Clean up leftover Ceph directories
|
||||
ansible.builtin.file:
|
||||
path: "{{ item }}"
|
||||
state: absent
|
||||
loop:
|
||||
- /etc/ceph
|
||||
- /var/lib/ceph
|
||||
- /var/log/ceph
|
||||
- /tmp/ceph-initial.conf
|
||||
|
||||
- name: Summary
|
||||
ansible.builtin.debug:
|
||||
msg: "Ceph cluster teardown complete on {{ hostname_short }}. Devices are clean and ready for reuse."
|
||||
@@ -0,0 +1,47 @@
|
||||
---
|
||||
# Phase 3: Remove non-bootstrap hosts from cluster (from bootstrap node)
|
||||
|
||||
- name: Get cluster hosts
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
timeout 10 ceph orch host ls --format json 2>/dev/null | python3 -c "
|
||||
import sys, json
|
||||
hosts = json.load(sys.stdin)
|
||||
for h in hosts:
|
||||
print(h['hostname'])
|
||||
" 2>/dev/null || true
|
||||
args:
|
||||
executable: /bin/bash
|
||||
register: cluster_hosts
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Remove non-bootstrap hosts from cluster
|
||||
ansible.builtin.shell: |
|
||||
HOST="{{ item }}"
|
||||
BOOTSTRAP="{{ hostvars[groups['ceph_bootstrap'][0]]['hostname_short'] }}"
|
||||
if [ "$HOST" = "$BOOTSTRAP" ]; then
|
||||
echo "Skipping bootstrap host $HOST (removed during purge)"
|
||||
else
|
||||
echo "Removing host $HOST from cluster"
|
||||
ceph orch host drain "$HOST" --force 2>&1 || true
|
||||
# Wait briefly for daemons to be removed
|
||||
sleep 10
|
||||
ceph orch host rm "$HOST" --force 2>&1 || true
|
||||
fi
|
||||
loop: "{{ cluster_hosts.stdout_lines | default([]) }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- cluster_hosts.stdout_lines | default([]) | length > 0
|
||||
register: host_rm_results
|
||||
changed_when: "'Removing host' in (item.stdout | default(''))"
|
||||
|
||||
- name: Show host removal results
|
||||
ansible.builtin.debug:
|
||||
msg: "{{ item.stdout }}"
|
||||
loop: "{{ host_rm_results.results | default([]) }}"
|
||||
loop_control:
|
||||
label: "{{ item.item | default('unknown') }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- item.stdout is defined
|
||||
@@ -0,0 +1,27 @@
|
||||
---
|
||||
# Ceph cluster teardown — reverse order of deploy
|
||||
# Phases: preflight → pools → osds → hosts → purge → cleanup
|
||||
|
||||
- name: Phase 0 - Preflight safety checks
|
||||
ansible.builtin.import_tasks: preflight.yml
|
||||
tags: [preflight, pools, osds, hosts, purge, cleanup]
|
||||
|
||||
- name: Phase 1 - Remove all pools
|
||||
ansible.builtin.import_tasks: pools.yml
|
||||
tags: [pools]
|
||||
|
||||
- name: Phase 2 - Remove all OSDs
|
||||
ansible.builtin.import_tasks: osds.yml
|
||||
tags: [osds]
|
||||
|
||||
- name: Phase 3 - Remove non-bootstrap hosts
|
||||
ansible.builtin.import_tasks: hosts.yml
|
||||
tags: [hosts]
|
||||
|
||||
- name: Phase 4 - Purge cluster from all nodes
|
||||
ansible.builtin.import_tasks: purge.yml
|
||||
tags: [purge]
|
||||
|
||||
- name: Phase 5 - Clean up devices and LVM
|
||||
ansible.builtin.import_tasks: cleanup.yml
|
||||
tags: [cleanup]
|
||||
@@ -0,0 +1,105 @@
|
||||
---
|
||||
# Phase 2: Remove all OSDs (from bootstrap node)
|
||||
# For each OSD: mark out → stop daemon → purge (removes from CRUSH, auth, etc.)
|
||||
|
||||
- name: Get all OSD IDs
|
||||
ansible.builtin.shell: |
|
||||
timeout 10 ceph osd ls --format json 2>/dev/null || echo "[]"
|
||||
register: osd_list
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
changed_when: false
|
||||
|
||||
- name: Parse OSD IDs
|
||||
ansible.builtin.set_fact:
|
||||
ceph_osd_ids: "{{ osd_list.stdout | from_json }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- osd_list.stdout is defined
|
||||
|
||||
- name: Show OSDs to be removed
|
||||
ansible.builtin.debug:
|
||||
msg: "OSDs to remove: {{ ceph_osd_ids | default([]) }}"
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
|
||||
- name: Set noout flag to prevent rebalancing during teardown
|
||||
ansible.builtin.command: ceph osd set noout
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_osd_ids | default([]) | length > 0
|
||||
changed_when: true
|
||||
|
||||
- name: Mark all OSDs out
|
||||
ansible.builtin.shell: |
|
||||
ceph osd out osd.{{ item }} 2>&1 || true
|
||||
loop: "{{ ceph_osd_ids | default([]) }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_osd_ids | default([]) | length > 0
|
||||
changed_when: true
|
||||
|
||||
- name: Stop all OSD daemons via orchestrator
|
||||
ansible.builtin.shell: |
|
||||
ceph orch daemon stop osd.{{ item }} 2>&1 || true
|
||||
loop: "{{ ceph_osd_ids | default([]) }}"
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_osd_ids | default([]) | length > 0
|
||||
changed_when: true
|
||||
|
||||
- name: Wait for OSD daemons to stop
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
for i in $(seq 1 30); do
|
||||
up=$(ceph osd stat --format json 2>/dev/null | python3 -c "import sys,json; print(json.load(sys.stdin).get('num_up_osds',0))" 2>/dev/null || echo 0)
|
||||
if [ "$up" -eq 0 ]; then
|
||||
echo "All OSD daemons stopped"
|
||||
exit 0
|
||||
fi
|
||||
echo "Waiting for OSDs to stop... ($up still up)"
|
||||
sleep 5
|
||||
done
|
||||
echo "WARNING: Some OSDs still up, proceeding anyway"
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_osd_ids | default([]) | length > 0
|
||||
changed_when: false
|
||||
|
||||
- name: Remove all OSDs via orchestrator
|
||||
ansible.builtin.shell: |
|
||||
ceph orch osd rm {{ ceph_osd_ids | join(' ') }} --zap --force 2>&1 || true
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_osd_ids | default([]) | length > 0
|
||||
changed_when: true
|
||||
|
||||
- name: Wait for OSD removal to complete
|
||||
ansible.builtin.shell: |
|
||||
set -o pipefail
|
||||
for i in $(seq 1 60); do
|
||||
remaining=$(ceph osd ls --format json 2>/dev/null | python3 -c "import sys,json; print(len(json.load(sys.stdin)))" 2>/dev/null || echo 0)
|
||||
if [ "$remaining" -eq 0 ]; then
|
||||
echo "All OSDs removed"
|
||||
exit 0
|
||||
fi
|
||||
echo "Waiting for OSD removal... ($remaining remaining)"
|
||||
sleep 10
|
||||
done
|
||||
echo "WARNING: OSD removal timed out, force-purging remaining"
|
||||
args:
|
||||
executable: /bin/bash
|
||||
when:
|
||||
- inventory_hostname in groups['ceph_bootstrap']
|
||||
- ceph_osd_ids | default([]) | length > 0
|
||||
changed_when: false
|
||||
|
||||
- name: Force-purge any remaining OSDs
|
||||
ansible.builtin.shell: |
|
||||
for osd_id in $(ceph osd ls 2>/dev/null); do
|
||||
echo "Force-purging osd.${osd_id}"
|
||||
ceph osd purge osd.${osd_id} --yes-i-really-mean-it 2>&1 || true
|
||||
done
|
||||
when: inventory_hostname in groups['ceph_bootstrap']
|
||||
register: force_purge
|
||||
changed_when: "'Force-purging' in force_purge.stdout"
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user