NAME

Rex::Rancher::Server - Rancher Kubernetes server (control plane) installation

VERSION

version 0.002

SYNOPSIS

use Rex::Rancher::Server;

# Install RKE2 server (default)
install_server(
  token   => 'my-cluster-secret',
  tls_san => ['lb.example.com'],
);

# Install K3s server
install_server(
  distribution => 'k3s',
  token        => 'my-cluster-secret',
  tls_san      => ['lb.example.com'],
);

# Join additional control plane node (HA setup)
install_server(
  distribution => 'rke2',
  token        => 'my-cluster-secret',
  server       => 'https://first-server:9345',
);

# Retrieve kubeconfig and join token from a running server
my $kubeconfig = get_kubeconfig('rke2');
my $token      = get_token('rke2');

# Update registry mirrors on an already-running node
update_registries(
  distribution => 'rke2',
  registries   => {
    mirrors => { 'docker.io' => { endpoint => ['http://cache:5000'] } },
  },
);

DESCRIPTION

Rex::Rancher::Server handles control plane installation for both RKE2 and K3s Kubernetes distributions. It provides a unified interface for installing, configuring, and managing server nodes.

RKE2 installation

By default the official install script at https://get.rke2.io is fetched and run via curl -sfL … | sh -; with install_method => 'artifact' the checksum-verified release tarball is installed instead (see "install_server"). The service is started with --no-block to avoid systemd's 90-second activation timeout (RKE2's first start pulls many container images), then polled with systemctl is-active for up to 10 minutes; a failed or never-active service dies with its journal tail. After that the function waits until the kubeconfig file appears at /etc/rancher/rke2/rke2.yaml; API readiness is confirmed separately by the caller using "wait_for_api" in Rex::Rancher::K8s.

K3s installation

The official install script at https://get.k3s.io is used (piped, or run against the checksum-verified binary with install_method => 'artifact'), with K3S_URL set when joining an existing server. The script runs with INSTALL_K3S_SKIP_START: instead of its own blocking restart, k3s.service is restarted with --no-block, then the same systemctl is-active wait (at most 10 minutes, journal tail on failure) as for RKE2 follows, so a joining server that cannot reach the first one dies instead of hanging the deploy. Then the kubeconfig wait. The token is read from config.yaml and never passed on the command line. Traefik and ServiceLB are disabled by default (disable in config.yaml, see "install_server") to leave room for Cilium and external load balancers.

Config layout

Both distributions use /etc/rancher/<dist>/config.yaml with the same key names (token, tls-san, node-name, node-label, disable, cni, etc.). When cilium => 1 (the default), the distribution's own CNI is switched off so that Cilium is the only one: on RKE2 cni: none and disable-kube-proxy: true, on K3s flannel-backend: none, disable-network-policy: true, disable-kube-proxy: true and cluster-cidr: 10.42.0.0/16; on both, Cilium's kube-proxy replacement takes over. See "install_server"'s cilium.

Registry mirrors are written to registries.yaml in the same directory. Both files are 0600 root:root: config.yaml holds the join token, registries.yaml may hold registry credentials.

FUNCTIONS

install_server(%opts)

Write the cluster configuration file, optionally write registries.yaml, install the distribution, start the service, wait until systemctl is-active reports it active, and then wait until the kubeconfig file is written to disk by the server process.

Returns 1 on success. Dies if installation fails, the distribution is unknown, the installed version differs from a pinned version, or the service does not become active within 10 minutes. A service that ends up failed or never gets active makes the die message carry the last 50 lines of its journal (journalctl -u SERVICE -n 50 --no-pager).

Options:

distribution

rke2 (default) or k3s. rke2 is the verified distribution. The k3s path carries the configuration kubernetes-ocp verified live (see "cilium"), but has not itself been run live through Rex::Rancher.

token

Shared secret used for node joining. If omitted, the token the server is already sealed with (/var/lib/rancher/rke2/server/token, K3s: /var/lib/rancher/k3s/server/token) is reused, so re-running install_server on a live control plane never rotates its token. Only on a fresh server (no such file) is a new one generated (up to 48 random alphanumeric characters, never fewer than 32). A passed token always wins.

The token is written to config.yaml only; it is never put on the installer command line or into its environment, where ps would show it. config.yaml is written 0600 root:root, including when it already exists.

server

URL of an existing server node to join. Used for multi-server HA setups (omit for the first/only server). For RKE2 the port is 9345; for K3s it is 6443.

tls_san

Additional TLS Subject Alternative Names for the API server certificate, as an arrayref or a comma-separated string. Include the load balancer address, public IP, or DNS name so that kubeconfig clients can connect.

version

Pinned version string, e.g. v1.30.4+rke2r1 for RKE2 or v1.30.4+k3s1 for K3s, handed to the installer as INSTALL_RKE2_VERSION / INSTALL_K3S_VERSION. If omitted, the latest stable release is installed.

When given, the version the installed binary reports (rke2 --version / k3s --version) is compared with it after the installer ran, and a mismatch dies (on RKE2 before the service is started; the K3s install script has already restarted it). This catches a pinned install or upgrade that failed while an older binary is still on the host.

install_method

How the distribution gets onto the host. script (default) pipes the official install script into sh (curl -sfL https://get.rke2.io | sh -, K3s: https://get.k3s.io), exactly as without this option.

artifact pre-downloads the release artifact for the node's own architecture (uname -m on the host: amd64 or arm64, anything else dies) from the GitHub release, verifies it against the release's official sha256sum-ARCH.txt and dies loudly on a mismatch, then runs the install script against the local file: RKE2 via INSTALL_RKE2_ARTIFACT_PATH (tarball rke2.linux-ARCH.tar.gz), K3s by installing the binary to /usr/local/bin/k3s and running the script with INSTALL_K3S_SKIP_DOWNLOAD=binary. Downloads run on the host with curl (no SFTP, nothing is uploaded) into /tmp/rke2-artifacts / /tmp/k3s-artifacts, which are emptied first. Requires version; dies without one.

On RPM-based hosts (Rocky, RHEL) RKE2's install script uses the tarball instead of its RPM method when given an artifact path, so no rke2-selinux package is installed, and a host that already carries RKE2 from RPMs is refused by the script ("existing RKE2 RPMs").

node_name

Kubernetes node name, written as node-name to config.yaml. If omitted, the system hostname is used.

disable

Packaged components to switch off, as an arrayref or a comma-separated string, written as disable to config.yaml. The names are distribution-specific. Default: ['rke2-ingress-nginx', 'rke2-traefik', 'rke2-traefik-crd'] on rke2 (no bundled ingress controller: RKE2 ships the Traefik charts since v1.30.3, opt-in, and deploys Traefik by default on new clusters since v1.36; a name the installed RKE2 does not ship is ignored), ['traefik', 'servicelb'] on k3s. A given list replaces the default rather than extending it; [] disables nothing. Independent of cilium.

# keep the default and also drop metrics-server
disable => [qw( rke2-ingress-nginx rke2-traefik rke2-traefik-crd
                rke2-metrics-server )],
node_labels

Node labels applied at join time, as an arrayref of key=value strings.

registries

Private registry mirror configuration. Written to registries.yaml in the distribution config directory, 0600 root:root (it may hold registry passwords). Structure:

{
  mirrors => {
    'docker.io' => { endpoint => ['http://registry.internal:5000'] },
  },
  configs => {
    'registry.internal:5000' => {
      auth => { username => 'user', password => 'pass' },
    },
  },
}
cilium

If true (default: 1), switch off the distribution's own CNI in the server config so that Cilium is the only one. Set to 0 to keep the distribution's default CNI (Canal on RKE2, Flannel on K3s).

On rke2, cni: none and disable-kube-proxy: true are written, preparing the node for Cilium with full kube-proxy replacement.

On k3s, flannel-backend: none, disable-network-policy: true, disable-kube-proxy: true and cluster-cidr: 10.42.0.0/16 are written: Flannel, k3s's embedded network policy controller and kube-proxy are switched off, and Cilium takes over all three with kube-proxy replacement. cluster-cidr is k3s's own default, written out because Cilium's cluster-pool IPAM is given the same range (see Rex::Rancher::Cilium). These are server settings that k3s agents take from the server; an additional server joining with server gets the same keys, as k3s requires them to match across servers. The same keys and Cilium values were verified live in kubernetes-ocp (k3s v1.36.4+k3s1, Cilium 1.20.0, Gateway API v1.6.1); the k3s path through Rex::Rancher has not been run live, and differs in the Cilium version it defaults to (see Rex::Rancher::Cilium).

nvidia_runtime_path

If true, and nvidia-container-runtime is on the host's PATH, write PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin to /etc/default/rke2-server before the installer runs. The rke2 unit sets no PATH, and rke2 looks for the NVIDIA runtime only when the service starts; without it a host- or vendor-installed toolkit (/usr/bin, e.g. DGX OS) is not wired into containerd. Other lines of the file are kept, an existing PATH= line is replaced. If the file changed while the service is already running, a warning asks for a restart; nothing is restarted. The GPU Operator's toolkit (/usr/local/nvidia/toolkit) is found by rke2 without this. No effect on k3s. Default: 0; "rancher_deploy_server" in Rex::Rancher turns it on for gpu => 1, gpu_setup => 0.

install_server(
  distribution => 'rke2',
  token        => 'my-cluster-secret',
  tls_san      => ['loadbalancer.example.com'],
  node_labels  => ['role=control-plane'],
  version      => 'v1.30.4+rke2r1',
  node_name    => 'cp-01',
);

# Checksum-verified release artifact instead of curl | sh
install_server(
  version        => 'v1.30.4+rke2r1',
  install_method => 'artifact',
);

update_registries(%opts)

Update registries.yaml on an already-running node and restart the distribution service to pick up the new registry mirror configuration.

Use this to add or change registry mirrors after the cluster is up — for example, after deploying an in-cluster registry that you want every node to use as a pull-through cache.

Required options:

registries

Registry mirror hashref (same structure as install_server's registries option).

Optional options:

distribution

rke2 (default) or k3s. Controls which service is restarted.

update_registries(
  distribution => 'rke2',
  registries   => {
    mirrors => {
      'docker.io'         => { endpoint => ['http://registry.internal:5000'] },
      'registry.internal' => { endpoint => ['http://registry.internal:5000'] },
    },
  },
);

get_kubeconfig($distribution)

Read the kubeconfig file from the remote server and return its content as a string. The file is read directly via cat over SSH; no SFTP is used.

$distribution defaults to rke2.

Note: RKE2 and K3s both write https://127.0.0.1 as the server address. The caller is responsible for substituting the real server address before saving the kubeconfig for external use. "rancher_deploy_server" in Rex::Rancher performs this substitution automatically.

Dies if the file cannot be read.

get_token($distribution)

Read the node join token from the server and return it as a string (trailing newline stripped).

$distribution defaults to rke2.

The token is stored at:

RKE2: /var/lib/rancher/rke2/server/node-token
K3s: /var/lib/rancher/k3s/server/node-token

Dies if the file cannot be read (e.g. server not yet started).

SEE ALSO

Rex::Rancher, Rex::Rancher::Node, Rex::Rancher::Agent, Rex::Rancher::Cilium, Rex::Rancher::K8s, Rex

SUPPORT

Issues

Please report bugs and feature requests on GitHub at https://github.com/Getty/rex-rancher/issues.

CONTRIBUTING

Contributions are welcome! Please fork the repository and submit a pull request.

AUTHOR

Torsten Raudssus <getty@cpan.org>

COPYRIGHT AND LICENSE

This software is copyright (c) 2026 by Torsten Raudssus <torsten@raudssus.de> https://raudssus.de/.

This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.