NAME
Rex::Rancher::Cilium - Cilium CNI installation for Rancher Kubernetes distributions
VERSION
version 0.002
SYNOPSIS
use Rex::Rancher::Cilium;
use JSON::MaybeXS; # JSON()->true below
# Install Cilium on an RKE2 cluster (defaults to version 1.17.0)
install_cilium(
distribution => 'rke2',
);
# Install, upgrade or leave alone -- decided from the Helm release,
# read through the local kubeconfig; with Gateway API and extra values
install_cilium(
distribution => 'rke2',
kubeconfig => "$ENV{HOME}/.kube/mycluster.yaml",
version => '1.17.0',
gateway_api => 1,
gateway_api_version => 'v1.2.0',
helm_values => { hubble => { relay => { enabled => JSON()->true } } },
);
# Install Cilium on a K3s cluster with explicit version
install_cilium(
distribution => 'k3s',
k8s_service_host => '10.0.0.1', # the control plane, not localhost
version => '1.17.0',
cli_version => 'v0.16.23',
);
# Upgrade an existing Cilium installation
upgrade_cilium(
distribution => 'rke2',
version => '1.17.0',
);
DESCRIPTION
Rex::Rancher::Cilium provides Cilium CNI installation and upgrade for Rancher Kubernetes distributions (RKE2 and K3s). The Cilium CLI runs on the remote server host via SSH; the Helm release state and the Gateway API CRDs are read and written from the local machine via Kubernetes::REST when a local kubeconfig is given.
Prerequisites
The server must already be running with its own CNI switched off in config.yaml, so that Cilium is the only one: on RKE2 cni: none and disable-kube-proxy: true, on K3s flannel-backend: none, disable-network-policy: true, disable-kube-proxy: true and cluster-cidr: 10.42.0.0/16; on both Cilium takes over kube-proxy's role. "install_server" in Rex::Rancher::Server sets these options by default when cilium => 1.
Helm values
Distribution-specific Helm values are written to /tmp/cilium-values-<dist>.yaml:
- RKE2
-
kubeProxyReplacement: true,k8sServiceHost: 127.0.0.1,k8sServicePort: "6443",cni.exclusive: false,operator.replicas: 1,ipam.mode: kubernetes. - K3s
-
kubeProxyReplacement: true,k8sServiceHost:k8s_service_host,k8sServicePort: "6443",cni.exclusive: true,operator.replicas: 1,ipam.mode: cluster-poolwithipam.operator.clusterPoolIPv4PodCIDRList: [10.42.0.0/16](K3s'scluster-cidr).
Both distributions share the same CNI binary/config paths (/opt/cni/bin, /etc/cni/net.d). gateway_api adds gatewayAPI.enabled: true, and helm_values is merged over all of it. --set kubeProxyReplacement=true is also passed on the command line and wins over any value file.
Default versions
The module ships with pinned defaults for reproducibility: Cilium 1.17.0 and Cilium CLI v0.16.23. Override with the version and cli_version options. The Gateway API version has no default; see gateway_api_version.
FUNCTIONS
install_cilium(%opts)
Install Cilium CNI on a Rancher Kubernetes cluster (RKE2 or K3s), or bring an existing installation to the requested version and values.
The Cilium CLI binary is downloaded from GitHub to /usr/local/bin/cilium on the remote host (skipped if the correct version is already present). Distribution-appropriate Helm values, merged with helm_values, are written to /tmp/cilium-values-<dist>.yaml and handed to the CLI.
With kubeconfig (a local kubeconfig path), the existing Helm release is read from the cluster via Kubernetes::REST before anything is installed, and the outcome depends on its state:
no release:
cilium install.deployedat the requested version with every requested value already in effect: nothing is done.deployedat another version, or with a requested value differing:cilium upgrade.failedwith an earlier revision stilldeployed(a failed upgrade):cilium upgradeagain.failedwith nodeployedrevision (a failed first install),pending-install,uninstallingoruninstalled: the stale release is removed (cilium uninstall, then its Helm release Secrets) and Cilium is installed fresh. No working Cilium exists in these states, so nothing running is taken down.an upgrade that would change
ipam.mode(the release sets one explicitly and the requested values differ): dies beforecilium upgrade, naming both modes. Cilium cannot switch IPAM mode under running pods; redeploy the cluster, or pass the deployed mode inhelm_values.pending-upgradeorpending-rollback: dies. An earlier run was interrupted or another is still running, and the deployed revision still carries the pod network; the message names the Secret to delete once no other deploy is running.
After a fresh install the cilium DaemonSet must exist, or it dies: the CLI has been seen to exit 0 without creating anything.
Without kubeconfig the release state cannot be read, and the previous behaviour applies: cilium install runs, and its "cannot re-use a name" error (the release already exists) counts as success, with a warning that version and values were not reconciled. Use "upgrade_cilium" to change an existing installation in that case.
On both distributions kubeProxyReplacement=true is passed to enable Cilium's eBPF-based kube-proxy replacement, so the server config must have switched off the distribution's CNI and kube-proxy: on RKE2 cni: none and disable-kube-proxy: true, on K3s flannel-backend: none, disable-network-policy: true, disable-kube-proxy: true and cluster-cidr: 10.42.0.0/16 ("install_server" in Rex::Rancher::Server writes these). Cilium then needs an API server address that works before any Service does, on every node: on RKE2 that is 127.0.0.1:6443, where servers and agents alike serve it. K3s agents serve it on 127.0.0.1:6444 instead, so on K3s Cilium is given the control plane's own address, k8s_service_host, and dies without one.
rke2 is the verified distribution. The k3s values are those kubernetes-ocp verified live (k3s v1.36.4+k3s1, Cilium 1.20.0, Gateway API v1.6.1 standard); Rex::Rancher's k3s path has not been run live itself and defaults to Cilium 1.17.0.
Options:
distribution-
rke2(default) ork3s. version-
Cilium version to install, e.g.
1.17.0. Default:1.17.0. cli_version-
Cilium CLI version to download, e.g.
v0.16.23. Default:v0.16.23. k8s_service_host-
K3s only, and required there: the control plane address Cilium reaches the API server at on port 6443 from every node (
k8sServiceHost), e.g. the server's IP or a name in its certificate. A loopback address dies, because K3s agents serve the API on127.0.0.1:6444, not 6443.helm_values->{k8sServiceHost}may take its place. "rancher_deploy_server" in Rex::Rancher passes the firsttls_san. Passing it on RKE2 dies: RKE2 uses127.0.0.1, which works on every node there. api_server-
Kubernetes API server URL, passed as
--api-serverto the CLI. Optional; the CLI uses the kubeconfig's server address if omitted. kubeconfig-
Local path to the cluster kubeconfig (as saved by "rancher_deploy_server" in Rex::Rancher). Enables the release-state handling above and is required for
gateway_api. Optional. helm_values-
Hashref of additional Helm values, deep-merged over the defaults (hashes merge key by key, anything else replaces the default). Optional.
gateway_api-
If true, apply the Gateway API CRDs before Cilium and set
gatewayAPI.enabled: true. Default: off. Requireskubeconfigandgateway_api_version. The CRDs are fetched from the kubernetes-sigs/gateway-api GitHub release on the machine running Rex and applied through Kubernetes::REST (nokubectl). They are skipped when the cluster already carries that bundle version and channel; when they are applied to a cluster with a runningcilium-operator, the operator is restarted so it picks up the new CRDs.RKE2 v1.37+ ships the same CRDs as its own chart,
rke2-gateway-api-crd, which would overwrite them. "rancher_deploy_server" in Rex::Rancher disables it for you; calling this directly, put it ininstall_server'sdisable. While that chart's Helm release exists this dies before applying anything, naming the way out: disable the chart and restartrke2-server(Helm uninstalls it but keeps itsgateway.networking.k8s.ioCRDs, so Gateways and routes survive), or use RKE2's CRDs viahelm_valueswithoutgateway_api. gateway_api_version-
Gateway API release to apply, e.g.
v1.2.0. It must match what the Ciliumversionsupports; there is no default because the two are version-locked. gateway_api_channel-
experimental(default) orstandard. What Cilium needs depends on both versions. Cilium up to 1.16 requiresTLSRoutev1alpha2, which only the experimental channel carries. Cilium 1.17 to 1.19 require standard-channel CRDs only and handleTLSRoutev1alpha2 when it is there; before Gateway API v1.5 the standard channel has noTLSRoute, sostandardcosts TLS passthrough. Gateway API v1.5 movedTLSRoute(as v1) into the standard channel, v1.6 alsoTCPRouteandUDPRoute; Cilium 1.20 requiresTLSRoutev1 andBackendTLSPolicyv1, which the v1.5+ standard channel carries. The default staysexperimentalso the default Cilium keepsTLSRoute.Gateway API v1.5+ ships an admission policy that refuses experimental CRDs on top of standard ones: a cluster that started on
standardcannot move toexperimental.
upgrade_cilium(%opts)
Upgrade an existing Cilium installation to a new version using cilium upgrade, unconditionally. The Cilium CLI is updated first if needed. The same Helm values generation logic as "install_cilium" is used, and gateway_api applies the CRDs the same way (restarting a running cilium-operator when they were applied).
Options are the same as "install_cilium".
upgrade_cilium(
distribution => 'rke2',
version => '1.17.0',
);
SEE ALSO
Rex::Rancher, Rex::Rancher::Server, Rex, https://docs.cilium.io/
SUPPORT
Issues
Please report bugs and feature requests on GitHub at https://github.com/Getty/rex-rancher/issues.
CONTRIBUTING
Contributions are welcome! Please fork the repository and submit a pull request.
AUTHOR
Torsten Raudssus <getty@cpan.org>
COPYRIGHT AND LICENSE
This software is copyright (c) 2026 by Torsten Raudssus <torsten@raudssus.de> https://raudssus.de/.
This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.