NAME

Rex::Rancher - Rancher Kubernetes (RKE2/K3s) deployment automation for Rex

VERSION

version 0.002

SYNOPSIS

use Rex -feature => ['1.4'];
use Rex::Rancher;

# Deploy RKE2 control plane (no GPU)
task "deploy_server", sub {
  rancher_deploy_server(
    distribution    => 'rke2',
    hostname        => 'cp-01',
    domain          => 'k8s.example.com',
    token           => 'my-secret',
    tls_san         => 'k8s.example.com',
    kubeconfig_file => "$ENV{HOME}/.kube/mycluster.yaml",
  );
};

# Deploy RKE2 control plane with GPU support
task "deploy_gpu_server", sub {
  rancher_deploy_server(
    distribution    => 'rke2',
    gpu             => 1,    # requires Rex::GPU installed
    reboot          => 1,    # reboot after driver install (first deploy)
    hostname        => 'gpu-cp-01',
    domain          => 'k8s.example.com',
    token           => 'my-secret',
    tls_san         => 'gpu-cp-01.k8s.example.com',
    kubeconfig_file => "$ENV{HOME}/.kube/gpu-cluster.yaml",
  );
};

# Deploy K3s worker with GPU support
task "deploy_gpu_worker", sub {
  rancher_deploy_agent(
    distribution => 'k3s',
    gpu          => 1,    # requires Rex::GPU installed
    hostname     => 'gpu-01',
    domain       => 'k8s.example.com',
    server       => 'https://10.0.0.1:6443',
    token        => 'K10...',
  );
};

# Deploy a single-node cluster (control plane + workloads on same node)
task "deploy_single_node", sub {
  rancher_deploy_server(
    distribution    => 'rke2',
    token           => 'my-secret',
    tls_san         => '10.0.0.1',
    kubeconfig_file => "$ENV{HOME}/.kube/single.yaml",
  );
  # Remove control-plane taint so workloads can be scheduled
  untaint_node(kubeconfig => "$ENV{HOME}/.kube/single.yaml");
};

DESCRIPTION

Rex::Rancher provides complete, zero-touch Kubernetes cluster deployment for Rancher distributions (RKE2 and K3s) using the Rex orchestration framework. It handles everything from raw Linux node preparation through to a running CNI and GPU device plugin.

GPU support is optional. Pass gpu => 1 and install Rex::GPU separately. Rex::Rancher works identically for non-GPU nodes. Clusters that hand the GPU to the NVIDIA GPU Operator pass gpu_setup => 0 and/or gpu_device_plugin => 0 and need no Rex::GPU.

When deploying a GPU server node, the full pipeline runs automatically:

1. Node preparation — base packages (on Debian/Ubuntu after stopping the automatic apt services), hostname, timezone, locale, NTP, swap off, kernel modules (br_netfilter, overlay), sysctl for Kubernetes networking.
2. GPU setup (gpu => 1, unless gpu_setup => 0) — NVIDIA driver via DKMS, optional reboot, Container Toolkit, CDI specs, containerd runtime config. Handled by Rex::GPU.
3. Cluster bring-up — write config (with cilium, the distribution's own CNI switched off), run RKE2 or K3s install script, wait for kubeconfig file on the remote host, then, with kubeconfig_file, fetch and save it locally and wait for API server readiness via Kubernetes::REST.
4. Cilium CNI (skipped with cilium => 0, which leaves the distribution's own CNI in place) — Cilium CLI installed on the remote host, Cilium installed, upgraded or left alone with distribution-appropriate Helm values (kube-proxy replacement on both).
5. NVIDIA device plugin (gpu => 1 + kubeconfig_file, unless gpu_device_plugin => 0) — DaemonSet applied via the Kubernetes API, wait for nvidia.com/gpu capacity on the node. No kubectl required anywhere.

All Kubernetes API operations (steps 3 and 5) run locally on the machine executing Rex using Kubernetes::REST and IO::K8s. No kubectl binary is needed on the remote host.

This distribution supports hosts without an SFTP subsystem (common on Hetzner dedicated servers). Use set connection => "LibSSH" and install Rex::LibSSH.

For fine-grained control, use the individual modules directly:

Rex::Rancher::Node — Node preparation
Rex::Rancher::Server — Control plane installation and config retrieval
Rex::Rancher::Agent — Worker node installation
Rex::Rancher::Cilium — Cilium CNI installation and upgrade
Rex::Rancher::K8s — Kubernetes API operations (device plugin, readiness, untaint)

GPU hardware support

With gpu => 1 (and gpu_setup not switched off), driver choice and hardware checks are made by Rex::GPU's gpu_setup; Rex::Rancher passes only the distribution and reboot. Newer Rex::GPU versions behave as follows:

  • Blackwell (RTX 50xx, RTX PRO, B200/GB200, B300): the open kernel driver on Ubuntu; on Debian 12 and 13 the driver comes from NVIDIA's CUDA repository. Any other release or architecture makes gpu_setup die.

  • Maxwell, Pascal, Volta (e.g. P100, V100, GTX 9xx/10xx, GT 1030, GeForce MX): pinned to the proprietary 580 driver branch.

  • Which GPUs get a driver is decided by generation, not by name: every Maxwell-or-newer GPU counts, consumer and laptop cards (GeForce MX, GT 1030, GTX 9xx, laptop RTX) included. Earlier Rex::GPU versions skipped these, so a re-deploy with gpu => 1 on such a node now installs the driver and, with reboot, reboots it; afterwards the node reports nvidia.com/gpu.

  • Kepler and older (e.g. GT 710, GTX 7xx, Tesla K80/K40/K20): skipped with a warning, no driver is installed, and a newer GPU on the same host is still set up. A node whose only NVIDIA GPUs are Kepler no longer dies: the deploy carries on without a GPU, and step 6 deploys the device plugin, waits about two minutes for nvidia.com/gpu and ends with a warning — pass gpu_device_plugin => 0 (or leave out gpu) for such a node.

  • VMs with a passed-through GPU next to an emulated console are detected as GPU hosts and get the full GPU pipeline, including the driver reboot — set reboot accordingly.

  • NVIDIA vGPU guests (e.g. Azure NVadsA10 v5, AWS G6f; told apart from a passed-through card by PCI subsystem ID) need NVIDIA's licensed vGPU guest driver, which Rex::GPU does not install. If it already works (nvidia-smi -L lists the GPU and libcuda.so.1 is in the linker cache) the deploy goes on as usual; otherwise gpu_setup dies naming the vGPU type, also when a non-vGPU GPU sits on the same host.

  • HGX B200/B300: the driver is installed as for any Blackwell, then a warning notes that CUDA needs NVIDIA Fabric Manager, the NVLink Subnet Manager (nvlsm), OFED/MOFED and kernel 5.17 or newer, none of which Rex::GPU sets up (it only checks whether Fabric Manager is running). GB200/GB300 NVL72 trays get an info line that multi-node NVLink needs nvidia-imex. Log output only; the deploy is unchanged.

  • Several compute GPUs: the driver must satisfy all of them (Ada + V100 gives 580, Ada + B200 gives the open driver). If no driver fits (V100 + B200) and none is already installed, or no package source serves the required driver for the OS (e.g. Debian 14 with Blackwell or V100), gpu_setup dies.

Such a die comes before any driver package is installed. At most pciutils has been installed for detection (when lspci was missing), unless the missing package source only shows once the package index is refreshed (e.g. Ubuntu, where no fitting driver package is found): then the package sources have already been prepared (apt-get update, repositories added or enabled). The die is not caught: "rancher_deploy_server" and "rancher_deploy_agent" abort with it. It does, however, come after "prepare_node" in Rex::Rancher::Node has already run, so base packages, hostname, timezone, locale, swap, kernel modules and sysctl are already changed; no Kubernetes distribution has been installed yet. See Rex::GPU for the details of detection and driver selection.

rancher_deploy_server(%opts)

Full control plane deployment in a single call: prepare the node, optionally set up GPU support, install the Kubernetes distribution, wait for the API, install Cilium CNI, and deploy the NVIDIA device plugin.

When gpu => 1 is passed and Rex::GPU is installed, GPU detection and driver installation are performed automatically as step 2 before the cluster is brought up. After Cilium is running, the NVIDIA device plugin DaemonSet is deployed via the local Kubernetes API (no kubectl required on the remote host) and the function waits for nvidia.com/gpu resources to appear on the node.

The full pipeline for a GPU server deployment:

1. prepare_node — base packages, hostname, timezone, locale, NTP, swap off, kernel modules, sysctl
2. gpu_setup (only with gpu => 1, unless gpu_setup => 0) — driver + toolkit + CDI + containerd config
3. install_server — write config, run installer, wait for the service to be active, then for the kubeconfig file
4. Fetch kubeconfig locally, patch 127.0.0.1 to the real server address, save to kubeconfig_file, wait for API with "wait_for_api" in Rex::Rancher::K8s (skipped without kubeconfig_file)
5. install_cilium (skipped with cilium => 0) — install Cilium CLI on remote, then install, upgrade or leave Cilium alone according to its Helm release (read through the saved kubeconfig once the API answered; without one, plain cilium install)
6. deploy_nvidia_device_plugin (only with gpu => 1 and kubeconfig_file, unless gpu_device_plugin => 0)

If the API does not answer through the saved kubeconfig within "wait_for_api" in Rex::Rancher::K8s's five minutes, the deploy dies naming the address it tried: the distribution is installed and running, but Cilium and the device plugin are not, and the node stays NotReady until a re-run (which reuses the token and picks up from there). The usual causes are the address (tls_san, kubeconfig_server), a firewall between this machine and port 6443, or a SAN missing from the certificate. Without kubeconfig_file nothing is awaited and Cilium is installed through the remote host alone.

Options:

distribution

Kubernetes distribution to install. rke2 (default) or k3s.

gpu

If true, detect GPUs and run the full GPU setup pipeline via Rex::GPU before installing the Kubernetes distribution, and deploy the NVIDIA device plugin once the API answers. Requires Rex::GPU to be installed unless gpu_setup => 0. Default: 0; without it no GPU step runs and gpu_setup/gpu_device_plugin are ignored. Driver selection depends on the GPU generation, and some hardware/OS combinations make the deploy die instead — see "GPU hardware support".

gpu_setup

With gpu => 1: whether step 2 runs Rex::GPU's gpu_setup (driver, container toolkit, CDI, containerd config). Default: 1. Pass 0 when the NVIDIA GPU Operator (driver.enabled, toolkit.enabled) or the host image provides these; Rex::GPU is then not loaded and need not be installed. On rke2, gpu_setup => 0 also turns on nvidia_runtime_path.

nvidia_runtime_path

Passed to "install_server" in Rex::Rancher::Server (and "install_agent" in Rex::Rancher::Agent): write a PATH to /etc/default/rke2-server (rke2-agent) before the first start, so rke2 finds a host-installed nvidia-container-runtime (DGX OS, a preinstalled toolkit in /usr/bin); skipped when there is none on the host. Default: on for gpu => 1, gpu_setup => 0, off otherwise. No effect on k3s.

gpu_device_plugin

With gpu => 1: whether step 6 deploys the NVIDIA device plugin DaemonSet. Default: 1. Pass 0 when the GPU Operator runs its own device plugin (devicePlugin.enabled) — two plugins would both advertise nvidia.com/gpu. With gpu_setup => 0 and this left on, the driver and the nvidia runtime must already be on the host, or the plugin finds no GPU and the deploy ends with a warning.

reboot

If true, reboot the host after GPU driver installation and wait for it to come back before proceeding. Only meaningful when Rex::GPU's gpu_setup runs (gpu => 1 without gpu_setup => 0); otherwise it is ignored with a warning. Required on first deploy when nouveau was previously loaded. Default: 0.

hostname

Short hostname to set on the node (optional). If omitted, the existing hostname is left unchanged.

domain

Domain suffix for the FQDN (optional). Used together with hostname to set /etc/hosts. If hostname is given without domain, hostname is still set but no hosts entry is written.

timezone

Timezone string, e.g. Europe/Berlin. Default: UTC.

locale

System locale, e.g. de_DE.UTF-8. Default: en_US.UTF-8. See "prepare_node" in Rex::Rancher::Node.

ntp

Install and start chrony. Default: 1; pass 0 to leave time sync to the host (e.g. a VM with hypervisor time sync).

server

URL of an existing server to join as an additional control plane node (HA), passed to "install_server" in Rex::Rancher::Server: https://SERVER:9345 for RKE2, https://SERVER:6443 for K3s. Omit it for the first server.

token

Shared cluster secret used for node joining. If omitted, the token of an already-installed server on the host is reused (a re-run never rotates it); only a fresh server gets a generated one. See "install_server" in Rex::Rancher::Server.

tls_san

Additional TLS Subject Alternative Names for the API server certificate. Accepts a string (single SAN or comma-separated list) or an arrayref. The first SAN is used as the server address when patching the kubeconfig (see kubeconfig_file below), and on K3s with Cilium as the address Cilium reaches the API server at (see k8s_service_host), so it must be reachable from every node.

kubeconfig_file

Local file path where the cluster kubeconfig is saved after the server is running. Required for the NVIDIA device plugin step to work. Optional — if omitted no local kubeconfig is saved and device plugin deployment is skipped even when gpu => 1.

RKE2 and K3s write https://127.0.0.1 into the kubeconfig. The first tls_san entry (or kubeconfig_server if provided) is substituted for 127.0.0.1 so the saved file connects to the real server address. If no address can be derived (neither tls_san nor kubeconfig_server given), or the derived address is itself loopback (127.0.0.1, localhost, ::1), the file is still saved but a warning is logged: it will keep pointing at https://127.0.0.1 and cannot reach the cluster from the operator's machine.

If the kubeconfig cannot be fetched from the host or written to kubeconfig_file, the deploy dies naming the cause, like the API timeout above: the server is installed, Cilium and the device plugin are not, and a re-run picks up from there.

kubeconfig_server

Explicit server address to use when patching the kubeconfig. Overrides the tls_san-based default, so the kubeconfig can point at an address that is not the first tls_san entry (e.g. the node's advertised host while tls_san lists every control-plane address). Only the kubeconfig is affected; the address must still be a name in the API server certificate (the node's own IPs and hostname, or a tls_san entry).

version

Pinned distribution version (INSTALL_RKE2_VERSION / INSTALL_K3S_VERSION). Default: latest stable. When given, the installed version is verified and a mismatch dies. See "install_server" in Rex::Rancher::Server.

install_method

script (default, curl | sh) or artifact (checksum-verified release artifact for the node's architecture, downloaded on the host; requires version). See "install_server" in Rex::Rancher::Server.

node_name

Kubernetes node name (node-name in config.yaml). Default: the hostname.

disable

Packaged components to switch off (disable in config.yaml). Default: rke2-ingress-nginx, rke2-traefik and rke2-traefik-crd on rke2, traefik and servicelb on k3s; a given list replaces the default. With gateway_api the rke2 default also holds rke2-gateway-api-crd (see gateway_api below). See "install_server" in Rex::Rancher::Server.

node_labels

Node labels to apply, as an arrayref of key=value strings.

registries

Private registry mirror configuration hashref, written to registries.yaml. See "install_server" in Rex::Rancher::Server for the structure.

cilium

Whether Cilium is the cluster's CNI. Default: 1: the distribution's own CNI and kube-proxy are switched off in config.yaml (RKE2: cni: none and disable-kube-proxy: true; K3s: flannel-backend: none, disable-network-policy: true, disable-kube-proxy: true and cluster-cidr: 10.42.0.0/16) and Cilium is installed in step 5 with kube-proxy replacement. K3s carries the configuration kubernetes-ocp verified live, but has not been run live through Rex::Rancher, which defaults to an older Cilium (see Rex::Rancher::Cilium). Set to 0 and Rex::Rancher does nothing CNI-related: the distribution's built-in CNI comes up (Canal for RKE2, Flannel for K3s) and the pipeline skips "install_cilium" in Rex::Rancher::Cilium entirely. Passing gateway_api, cilium_version, cilium_cli_version, cilium_helm_values or k8s_service_host together with cilium => 0 dies before the node is touched.

cilium_version, cilium_cli_version, cilium_helm_values

Passed to "install_cilium" in Rex::Rancher::Cilium as version, cli_version and helm_values.

k8s_service_host

K3s only: the control plane address Cilium reaches the API server at from every node, passed to "install_cilium" in Rex::Rancher::Cilium. Default: the first tls_san. Without either, a K3s deploy with Cilium dies before the node is touched: K3s agents serve the API on 127.0.0.1:6444, so no localhost address works on every node. Passing it on RKE2 dies, also before the node is touched.

gateway_api, gateway_api_version, gateway_api_channel

Passed to "install_cilium" in Rex::Rancher::Cilium unchanged. gateway_api needs kubeconfig_file; invalid Cilium options die before the node is touched.

With gateway_api, Cilium's CRDs have to be the only ones: RKE2 v1.37+ would otherwise install its own rke2-gateway-api-crd chart over them (older RKE2 ignores the name). Without a disable of yours it is added to the default list; a disable of yours that lacks it is used as given, with a warning. On a running cluster the change to config.yaml takes effect only when rke2-server restarts; until then "install_cilium" in Rex::Rancher::Cilium dies while RKE2's release exists.

rancher_deploy_agent(%opts)

Full worker node deployment: prepare the node, optionally set up GPU support, install the Kubernetes agent, and join the existing cluster.

The pipeline is shorter than "rancher_deploy_server" — there is no Cilium installation or kubeconfig retrieval. GPU host setup via gpu => 1 (and gpu_setup, reboot) works identically to the server case; there is no device plugin step, so gpu_device_plugin has no effect here.

Options:

server

URL of the server to join. For RKE2: https://SERVER_IP:9345. For K3s: https://SERVER_IP:6443. Required.

token

Node join token. Obtain from the server with "get_token" in Rex::Rancher::Server. Required.

node_name

Override the node name registered in Kubernetes (optional).

distribution, version, install_method, node_labels, registries, nvidia_runtime_path

As for "rancher_deploy_server"; passed to "install_agent" in Rex::Rancher::Agent.

hostname, domain, timezone, locale, ntp

As for "rancher_deploy_server"; passed to "prepare_node" in Rex::Rancher::Node.

gpu, gpu_setup, reboot

As for "rancher_deploy_server" (step 2).

A missing server or token dies before the host is touched. As on the server, an SFTP-less host needs the LibSSH connection backend; without it the deploy dies before the first step with a hint to Rex::LibSSH.

The server-only options have no effect on an agent and are ignored: tls_san, disable, kubeconfig_file, kubeconfig_server, cilium, cilium_version, cilium_cli_version, cilium_helm_values, gateway_api, gateway_api_version, gateway_api_channel, k8s_service_host and gpu_device_plugin. Whether a K3s agent runs Flannel and kube-proxy or leaves both to Cilium follows the server's config.yaml.

rancher_scan_known_hosts($host, %opts)

Pre-seed the local known_hosts with $host's SSH host key by running ssh-keyscan on the machine executing Rex (not on the target). Returns true if a key was added, false/undef otherwise.

This is a pre-connect helper: Rex::LibSSH >= 0.004 verifies the server host key against known_hosts (a CWE-322 fix; earlier versions never checked). A freshly-installed host — the Hetzner dedicated servers this distribution targets — has no known_hosts entry, so the very first verified connect dies with host key is not in known_hosts and strict_hostkeycheck is on. Scanning the key in beforehand fixes that while keeping host-key verification on, which is why this is preferred over disabling the check.

Because Rex opens the connection before the task body runs, call this from a before hook so it executes ahead of connect (see the before 'ALL' block in eg/hetzner-gpu.Rexfile). It shells out locally and never uses Rex's run, since there is no connection yet:

before 'ALL' => sub {
  my ($server) = @_;
  rancher_scan_known_hosts($server);
};

It is idempotent (an already-trusted host is left untouched) and degrades to a warning — never a hard failure — when ssh-keyscan is absent or the host is unreachable; the subsequent verified connect then surfaces the real error.

Options:

known_hosts

Path to the known_hosts file to update. Defaults to $HOME/.ssh/known_hosts.

SEE ALSO

Rex, Rex::LibSSH, Rex::GPU, Rex::Rancher::K8s, Kubernetes::REST, IO::K8s

SUPPORT

Issues

Please report bugs and feature requests on GitHub at https://github.com/Getty/rex-rancher/issues.

CONTRIBUTING

Contributions are welcome! Please fork the repository and submit a pull request.

AUTHOR

Torsten Raudssus <getty@cpan.org>

COPYRIGHT AND LICENSE

This software is copyright (c) 2026 by Torsten Raudssus <torsten@raudssus.de> https://raudssus.de/.

This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.