Files
yucca-o11y/docs/01-bootstrap-guide.md
T
2026-06-08 13:01:28 -04:00

5.8 KiB
Raw Blame History

Bootstrap guide

How to stand up an environment from nothing. The cluster is built by Terragrunt in module order, then handed to Flux. Run everything through mise so tool versions and 1Password credential injection are consistent. See the README for the repository layout — where each module and manifest lives, and where state and secrets come from.

One-time prep (per environment)

  1. Create the environment's 1Password vault (o11y_tf_staging or o11y_tf_prod).

  2. Set up .private/openstack/<env>/openrc.sh for the environment's OVH Public Cloud project.

  3. Upload the Talos OpenStack image (control planes only) to the environment's regions:

    mise run talos:dl:cp && mise run talos:ul:cp
    

    This image carries the qemu-guest-agent + tailscale schematic.

  4. Workers need no download or upload — they are OVH BYOI and pull the bare-metal raw straight from the Talos Factory at order time. The worker schematic must stay tailscale-only: qemu-guest-agent on bare metal blocks on a virtio port that never appears and reboot-loops the node.

  5. For production: delete the apex DNS records via the OVH dashboard before applying.

Apply order

Set the environment in the shell first:

export ENVIRONMENT=staging
export TF_VAR_env=staging
  1. OVH — cloud project, vRack, private network, CP instances, workers, IPLB, DNS. First-time runs are slow (CP instances ~5 min each, bare-metal orders 20–45 min each, IPLB ~10 min).

    mise run tg run --working-dir deployment/modules/ovh/account apply
    
  2. Tailscale — tailnet-global ACL (only needs one run across all environments).

    mise run tg run --working-dir deployment/modules/tailscale/account apply
    
  3. Talos (bootstrap) — initial bring-up over public IPs, because the Tailscale extension isn't running yet.

    TF_VAR_use_public_endpoints=true mise run tg run --working-dir deployment/modules/talos/cluster apply
    
  4. Verify the cluster is up and operator-side Tailscale routing works. Pull the configs (see Cluster access) and hit the APIs over the tailnet:

    mise run talos:kubeconfig && mise run talos:talosconfig
    kubectl --kubeconfig .private/$ENVIRONMENT/kubeconfig get nodes -o wide
    talosctl --talosconfig .private/$ENVIRONMENT/talosconfig -n 10.150.200.10 get members
    
  5. Talos (steady state) — drop the public-endpoints override now that Tailscale routes work; the host firewall closes the public NIC (everything except :30443 on workers).

    unset TF_VAR_use_public_endpoints
    mise run tg run --working-dir deployment/modules/talos/cluster apply
    
  6. Kubernetes/Helm — install the Flux Operator + Instance and create bootstrap secrets (cert-manager, OVH DNS credentials, external-secrets 1Password token). After this, Flux owns cluster state.

    mise run tg run --working-dir deployment/modules/kubernetes/helm apply
    

Flux then reconciles from kubernetes/clusters/<env>/apps.yaml, fanning out to the per-app Kustomizations in dependency order.

Cluster access

How to get kubectl / talosctl access to an existing cluster (no bootstrap required).

Prerequisites:

  • Tailscale — the cluster APIs are reachable only over the tailnet, so your host needs Tailscale running with subnet-route consumption enabled: tailscale set --accept-routes on Linux, or the "Use Tailscale subnets" toggle in the macOS app.
  • 1Password CLI (op) — installed and signed in to the team-futo.1password.com account. mise run talos:config fetches the configs through mise run tg, which wraps op run to inject the Terraform state credentials; without an authenticated op it can't read state.

Fetch the configs. Two tasks pull kubeconfig and talosconfig for the environment (run whichever you need):

export ENVIRONMENT=staging        # or production
export TF_VAR_env=$ENVIRONMENT
mise run talos:kubeconfig         # for kubectl
mise run talos:talosconfig        # for talosctl

Each writes to .private/$ENVIRONMENT/ (mode 600) from the Talos module's Terraform outputs. talos:kubeconfig also repoints the kubeconfig server: from the floating VIP (10.150.200.5) to a control-plane private IP (10.150.200.10) — the VIP doesn't ARP reliably across DCs over Tailscale, and every CP IP is in the apiserver cert SANs so TLS still validates.

A highly-available operator API endpoint is TBD. kubectl is pinned to a single control-plane IP, so if that CP is down you currently repoint to another by hand (any CP IP works — they're all cert SANs). The floating VIP is HA inside the cluster (kubelet and in-cluster clients use it) but doesn't ARP across DCs over Tailscale, so there's no HA endpoint for operators yet.

Point your tools at them:

export KUBECONFIG=$PWD/.private/$ENVIRONMENT/kubeconfig
export TALOSCONFIG=$PWD/.private/$ENVIRONMENT/talosconfig

kubectl get nodes
talosctl -n 10.150.200.10 health   # any CP private IP; the talosconfig also lists the workers

Common operations

Task Command
Plan/apply one module mise run tg run --working-dir deployment/modules/<m> {plan,apply}
Plan/apply all in dep order mise run tf:{plan,apply}
Re-init backends mise run tf:init
Format HCL / Terraform mise run tg:fmt / mise run tf:fmt
Lint docs mise run md:lint
Fetch kubeconfig / talosconfig mise run talos:kubeconfig / mise run talos:talosconfig (see Cluster access)

Tooling

  • OpenTofu + Terragrunt for IaC; mise drives tool versions and task wrappers. Version pins live in .mise/config.toml and each module's lock file.
  • 1Password CLI (op run --env-file deployment/.env) injects API credentials at invocation time.