Files
yucca-o11y/docs/01-bootstrap-guide.md
Devin Buhl 76e4e69f0d feat(deployment): guard servers from destroy and stage production talos reboots (#269)
* feat(deployment): guard servers from destroy and stage production talos reboots

Add prevent_destroy to the OVH control-plane instances, the bare-metal
workers, and the Talos machine secrets so a plan that would replace or
delete one fails instead of applying.

Add a talos apply_mode variable. Production applies with
staged_if_needing_reboot: a machine-config change that needs a reboot is
written to the node but stays inactive until an operator reboots it, one
node at a time. Staging keeps auto so the reboot path is exercised there
first. The provider has no try-style auto-revert mode, and the per-node
applies are unordered, so this is the only lever that stops a bad config
from rebooting every control plane at once.

Document both behaviours in the bootstrap guide.

Signed-off-by: Devin Buhl <devin@buhl.casa>

* chore(rootly): refresh provider lock hashes

Signed-off-by: Devin Buhl <devin@buhl.casa>

---------

Signed-off-by: Devin Buhl <devin@buhl.casa>
2026-09-02 09:08:17 -04:00

129 lines
8.0 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Bootstrap guide
How to stand up an environment from nothing. The cluster is built by Terragrunt in module order, then handed to Flux. Run everything through `mise` so tool versions and 1Password credential injection are consistent. See the [README](../README.md) for the repository layout — where each module and manifest lives, and where state and secrets come from.
## One-time prep (per environment)
1. Create the environment's 1Password vault (`o11y_tf_staging` or `o11y_tf_prod`).
2. Set up `.private/openstack/<env>/openrc.sh` for the environment's OVH Public Cloud project.
3. Upload the Talos **OpenStack** image (control planes only) to the environment's regions:
```bash
mise run //deployment/modules/ovh/account:image:download && mise run //deployment/modules/ovh/account:image:upload
```
This image carries the `qemu-guest-agent` + `netbird` schematic.
4. Workers need **no download or upload** — they are OVH BYOI and pull the bare-metal raw straight from the Talos Factory at order time. The worker schematic must stay **netbird-only**: `qemu-guest-agent` on bare metal blocks on a virtio port that never appears and reboot-loops the node.
5. For production: delete the apex DNS records via the OVH dashboard before applying.
## Apply order
Set the environment in the shell first:
```bash
export ENVIRONMENT=staging
export TF_VAR_env=staging
```
1. **OVH** — cloud project, vRack, private network, CP instances, workers, IPLB, DNS. First-time runs are slow (CP instances ~5 min each, bare-metal orders 20–45 min each, IPLB ~10 min).
```bash
mise run //deployment/modules/ovh/account:apply
```
2. **NetBird** — the per-environment mesh objects (all named `o11y-<env>-*`): groups and setup keys for the Talos nodes and the in-cluster routing peers, the vRack network route, the mesh-gateway VIP resource and DNS zone, the pod-egress network, and the access policies (`yucca → resource`, `yucca → talos`, `yucca → gateway`, `talos → bootstrap opc`). The Talos module consumes the node setup key and mesh zone from here, so apply NetBird first.
```bash
mise run //deployment/modules/netbird/cluster:apply
```
3. **NetBox** — registers the environment's ranges (vRack, gateway ServiceCIDR, pod-egress CIDR) in IPAM, from the same values the other modules allocate.
```bash
mise run //deployment/modules/netbox/cluster:apply
```
4. **Talos (bootstrap)** — initial bring-up over public IPs, because the NetBird extension isn't running yet.
```bash
TF_VAR_use_public_endpoints=true mise run //deployment/modules/talos/cluster:apply
```
5. **Verify** the cluster is up and operator-side NetBird routing works. Pull the configs (see [Cluster access](#cluster-access)) and hit the APIs over the NetBird network — use the **direct** kubeconfig context during bring-up, since the default context targets the mesh gateway, which only exists once Flux has reconciled:
```bash
mise run //deployment/modules/talos/cluster:kubeconfig && mise run //deployment/modules/talos/cluster:talosconfig
kubectl --kubeconfig .private/$ENVIRONMENT/kubeconfig --context o11y-$ENVIRONMENT-direct get nodes -o wide
talosctl --talosconfig .private/$ENVIRONMENT/talosconfig -n 10.150.200.10 get members
```
6. **Talos (steady state)** — drop the public-endpoints override now that NetBird routes work; the host firewall closes the public NIC (everything except `:30443` on workers).
```bash
unset TF_VAR_use_public_endpoints
mise run //deployment/modules/talos/cluster:apply
```
7. **Kubernetes/Helm** — install CoreDNS (Terraform-seeded — Flux needs cluster DNS from its first reconcile; Talos's copy is disabled), the Flux Operator + Instance, the `bootstrap-settings` ConfigMap, and the bootstrap secrets (cert-manager OVH DNS credentials, the 1Password Connect token for external-secrets). After this, Flux owns cluster state.
```bash
mise run //deployment/modules/kubernetes/helm:apply
```
Flux then reconciles from `kubernetes/clusters/<env>/apps.yaml`, fanning out to the per-app Kustomizations in dependency order.
## Cluster access
How to get `kubectl` / `talosctl` access to an **existing** cluster (no bootstrap required).
**Prerequisites:**
* **NetBird** — the cluster APIs are reachable only over the NetBird network, so your host must be running the NetBird client (`netbird up`) and joined to the FUTO NetBird account, which places your peer in the `yucca` group. The access policy then distributes the route to the cluster's vRack CIDR, so `kubectl`/`talosctl` can reach the nodes' private IPs.
* **1Password CLI (`op`)** — installed and signed in to the `team-futo.1password.com` account. `mise run //deployment/modules/talos/cluster:kubeconfig` fetches the configs through the `//:tg` terragrunt wrapper, which wraps `op run` to inject the Terraform state credentials; without an authenticated `op` it can't read state.
**Fetch the configs.** Two tasks pull `kubeconfig` and `talosconfig` for the environment (run whichever you need):
```bash
export ENVIRONMENT=staging # or production
export TF_VAR_env=$ENVIRONMENT
mise run //deployment/modules/talos/cluster:kubeconfig # for kubectl
mise run //deployment/modules/talos/cluster:talosconfig # for talosctl
```
Each writes to `.private/$ENVIRONMENT/` (mode 600) from the Talos module's Terraform outputs. The kubeconfig is TF-authored with two contexts:
* **`o11y-<env>`** (default) — the HA endpoint `kube.<mesh-zone>:6443`, fronted by the Envoy mesh gateway (TLS passthrough to every apiserver). Survives any single CP being down and never hairpins through a NetBird routing peer.
* **`o11y-<env>-direct`** — a control-plane private IP, for bootstrap/DR before the mesh gateway exists: `kubectl --context o11y-<env>-direct`. Every CP IP is an apiserver cert SAN, so TLS validates on both paths. (The floating VIP `.5` is in-cluster-only — it doesn't ARP across DCs.)
**Point your tools at them:**
```bash
export KUBECONFIG=$PWD/.private/$ENVIRONMENT/kubeconfig
export TALOSCONFIG=$PWD/.private/$ENVIRONMENT/talosconfig
kubectl get nodes
talosctl -n 10.150.200.10 health # any CP private IP; the talosconfig also lists the workers
```
## Common operations
| Task | Command |
| --- | --- |
| Plan/apply one module | `mise run //deployment/modules/<m>:plan` / `mise run //deployment/modules/<m>:apply` (or `mise run :plan` from inside the module) |
| Plan/apply all in dep order | `mise run //deployment:plan` / `mise run //deployment:apply` |
| Re-init backends | `mise run //deployment:init` |
| Any other terragrunt subcommand | `mise run //:tg <args>` (runs in the current directory, e.g. `state rm`, `import`, `force-unlock`) |
| Format HCL / Terraform | `mise run //deployment:fmt` |
| Lint docs | `mise run //:md:lint` |
| Fetch kubeconfig / talosconfig | `mise run //deployment/modules/talos/cluster:kubeconfig` / `mise run //deployment/modules/talos/cluster:talosconfig` (see [Cluster access](#cluster-access)) |
## Apply guardrails
* **Production never reboots unattended.** The production Talos module applies with `staged_if_needing_reboot`: a machine-config change that needs a reboot is written to the node but stays inactive until the node reboots (the plan shows `resolved_apply_mode = "staged"` for those nodes). Roll them yourself, one at a time, checking `talosctl health` between nodes and doing the elected NetBird routing peer last (rebooting it drops the vRack route until another peer is elected). Staging applies with `auto`, so it reboots on its own and proves the change first.
* **Servers cannot be destroyed by an apply.** The OVH control-plane instances, the bare-metal workers, and the Talos machine secrets carry `prevent_destroy`; a plan that would replace or delete one fails instead. Tearing one down deliberately means removing that lifecycle rule in the same change.
## Tooling
* **OpenTofu** + **Terragrunt** for IaC; **mise** drives tool versions and task wrappers. Version pins live in `.mise/config.toml` and each module's lock file.
* **1Password CLI** (`op run --env-file deployment/.env`) injects API credentials at invocation time.