NAME
Rex::GPU::NVIDIA - NVIDIA GPU driver and container toolkit management
VERSION
version 0.002
SYNOPSIS
use Rex::GPU::NVIDIA;
# Step 1: Install driver (with reboot on first deploy)
install_driver(reboot => 1);
# Step 2: Install NVIDIA Container Toolkit
install_container_toolkit();
# Step 3: Generate CDI specs for the device plugin
generate_cdi_specs();
# Step 4: Configure containerd for Kubernetes
configure_containerd('rke2'); # 'rke2', 'k3s', or 'containerd'
# Verify the current installation status
my $ok = verify_nvidia();
DESCRIPTION
Rex::GPU::NVIDIA manages the full NVIDIA software stack needed to run GPU-accelerated workloads in Kubernetes: driver installation, the Container Toolkit, CDI spec generation, and containerd runtime configuration.
Each step is OS-aware and handles Debian/Ubuntu, RHEL/Rocky/CentOS, and openSUSE Leap without further configuration.
Driver installation
Drivers are installed via DKMS, so they survive kernel upgrades without needing reinstallation. The nouveau open-source driver is blacklisted and the initramfs is regenerated to prevent it from loading at boot.
On Debian, whichever of contrib, non-free and non-free-firmware is missing is added to each Debian archive entry, in both source formats: active deb lines of /etc/apt/sources.list (appended after the last component; non-free-firmware alone, as the Debian 12 installer writes it, does not count as non-free), and the Components: field of stanzas in the deb822 format (/etc/apt/sources.list.d/*.sources, e.g. debian.sources on Debian 13 and Debian cloud images). An entry is a Debian archive when it is of type deb (not deb-src, not commented out or Enabled: no), its components include main, and either Signed-By / [signed-by=...] names only debian-archive-* keyrings under /usr/share/keyrings (whatever the URI, so a mirror of your own signed with Debian's key counts), or there is no signed-by and every URI is a *.debian.org host, Hetzner's mirror.hetzner.com/debian/ (or .de) mirror or the cloud images' mirror+file:/etc/apt/mirrors/debian*.list. Third-party sources and unknown mirrors are left untouched, and a file with nothing to add is not rewritten; if no Debian archive entry is recognised in either format, a warning is logged. To recognise a mirror of your own, override "is_debian_archive_uri" in Rex::GPU::NVIDIA::Setup::Debian in a subclass.
The install is done by Rex::GPU::NVIDIA::Setup::Debian, Rex::GPU::NVIDIA::Setup::Ubuntu, Rex::GPU::NVIDIA::Setup::RHEL and Rex::GPU::NVIDIA::Setup::SUSE (experimental classes, see Rex::GPU::NVIDIA::Setup); the steps and commands are the ones described here. A subclass of your own replaces them through the setup option of "install_driver" or set gpu_nvidia_setup -- see "WRITING YOUR OWN SETUP" in Rex::GPU::NVIDIA::Setup.
The package choice also depends on the GPU generation, read from the PCI device IDs of the GPUs passed as gpus (or gpu) to "install_driver" ("gpu_setup" in Rex::GPU passes every compute GPU it detected); one driver has to fit them all. Without that option every host gets the default selection below.
Turing, Ampere, Ada, Hopper and unknown IDs: the default per-distro selection.
Blackwell (B200/GB200/B300, GeForce RTX 50xx, RTX PRO Blackwell, GB10), on any CPU architecture: it has no proprietary kernel module. Ubuntu selects the
-server-openvariant. Debian 12/13 installs the open-module set from NVIDIA's CUDA repository instead ofnon-free.Maxwell, Pascal, Volta (e.g. V100, P100, GeForce GT 1030, GTX 980): only the proprietary module of the 580 branch supports them. Ubuntu pins
nvidia-driver-580-server, RHEL pins branch 580 (module stream or versionlock) withkmod-nvidia-latest-dkms, and openSUSE usesnvidia-driver-G06-kmp-meta.Kepler or older: no supported branch is installed and "install_driver" dies before changing the host. "gpu_setup" in Rex::GPU never passes one: detection skips it with a warning.
NVIDIA vGPU guest (
vgpu => 1): it needs NVIDIA's licensed vGPU guest driver, which is not installed here; unless a working driver is already there, "install_driver" dies before changing the host.
See the gpus option of "install_driver" for the exact packages.
On Ubuntu, the newest available nvidia-driver-NNN-server package is auto-detected and installed by default. It is looked up after apt-get update; if the refreshed index lists none, install_driver dies before installing a driver package.
On RHEL/Rocky/AlmaLinux/CentOS Stream, the NVIDIA CUDA repository is added and the open-kernel DKMS variant is used by default. For RHEL 10+ the module streams approach is not available; kmod-nvidia-open-dkms is installed directly. The CUDA repository URL is architecture-aware: aarch64 hosts use the sbsa tree (repos/rhelN/sbsa/), x86_64 hosts the x86_64 tree.
On openSUSE Leap, a kmp-meta package is used (by default nvidia-open-driver-G06-signed-kmp-meta for Leap 15.x, nvidia-open-driver-G07-signed-kmp-meta for Leap 16.x) to ensure the kernel module and userspace libraries are always at the same version. Stale OSS non-free packages are removed before installation and locked afterwards to prevent nvidia-smi from reporting a Driver/library version mismatch.
Container Toolkit
nvidia-container-toolkit is installed from the official NVIDIA GitHub package repository (https://nvidia.github.io/libnvidia-container/).
CDI specs
Container Device Interface specifications let the Kubernetes device plugin enumerate GPU resources without requiring privileged container access. When a managed CDI source — the nvidia-cdi-refresh systemd unit shipped by modern nvidia-container-toolkit — already keeps /run/cdi/nvidia.yaml fresh, that source is left to own CDI; otherwise a static spec is written to /etc/cdi/nvidia.yaml by nvidia-ctk cdi generate. Only one of the two default scan dirs (/etc/cdi, /run/cdi) is populated, so nvidia.com/gpu is never defined twice.
Containerd configuration
For RKE2 and K3s, the NVIDIA runtime is registered additively and version-aware, without clobbering the config that the distribution generates: a no-op when RKE2/K3s already wired the runtime natively, a config-v3.toml.d/ drop-in on modern (containerd 2.x / config v3) hosts, or a base-extending ({{ template "base" . }}) config.toml.tmpl on legacy (containerd 1.x / config v2) hosts. A stale full-config config.toml.tmpl left by the 0.001 release is detected by its exact bare-clobber signature and removed so the distribution regenerates its native config (the operator must restart the service or reboot for that to take effect). For standalone containerd, nvidia-ctk runtime configure is used.
Supported distributions:
Debian 11 (bullseye), 12 (bookworm), 13 (trixie)
Ubuntu 22.04 (jammy), 24.04 (noble)
RHEL / Rocky Linux / AlmaLinux 8, 9, 10 — CentOS Stream 9, 10
The verified target set is the RKE2 Linux family above. openSUSE Leap / SLES is unverified and unsupported — SUSE is not a deploy target for the GPU-on-Rancher pipeline. The Rex::GPU::NVIDIA::Setup::SUSE path exists but is not exercised; do not treat a SUSE run as evidence.
Tested on Hetzner dedicated servers with NVIDIA RTX 4000 SFF Ada Generation.
install_driver
Install NVIDIA GPU drivers appropriate for the detected OS using DKMS. Blacklists the nouveau driver and rebuilds the initramfs so the blacklist takes effect on next boot.
After installation (and after reboot, if reboot => 1), calls "verify_nvidia_driver" to confirm the kernel module loaded correctly. Not the full "verify_nvidia": the container toolkit is installed after the driver, so its check could only warn here.
Dies if the detected OS is not supported.
If a working NVIDIA driver is already loaded and functional (nvidia-smi -L lists a GPU and libcuda.so.1 is in the linker cache, ldconfig -p) — for example on a host provisioned via the NVIDIA CUDA package repository, or on a re-run — install_driver logs this and returns immediately without installing anything, and without blacklisting nouveau or rebooting. A host where nvidia-smi works but libcuda.so.1 is missing gets a warning and the driver install (see "already_installed" in Rex::GPU::NVIDIA::Setup). This keeps the call idempotent and stops the per-distro package selection from installing a second, version-conflicting (or lower) driver over the one already present.
Options:
reboot-
If true, the host is rebooted immediately after driver installation. The function waits up to 5 minutes for the host to come back (polling every 5 seconds via SSH reconnect; the host counts as back once
echo okrun over the new connection printsok), then continues with verification. Default:0.Rebooting is required on the first deployment when the
nouveauopen-source driver was previously loaded, because nouveau must be unloaded before the NVIDIA kernel module can bind to the device. gpus-
Optional arrayref of the GPUs this driver install is for. "gpu_setup" in Rex::GPU passes every CUDA-capable NVIDIA GPU it detected here, in the shape "detect" in Rex::GPU::Detect returns. Only
device_id,nameand the vGPU keys (vgpu,vgpu_type,subsystem_id) are read, so a caller that finds the GPUs itself -- withoutlspci-- passes just the first two; a GPU withoutvgpuis not a vGPU:install_driver(gpu => { device_id => '2b85', name => 'NVIDIA GeForce RTX 5090' });NVIDIA vGPU guest (karr #24;
vgpu => 1, see "NVIDIA vGPU guests" in Rex::GPU::Detect): such a device needs NVIDIA's licensed vGPU guest (GRID) driver, which none of the package sources Rex::GPU installs from carries, and the opennvidia.koof the datacenter packages refuses an Ampere-or-newer vGPU. The already-installed check below runs first: a guest whose vGPU driver already works (nvidia-smi -Llists the GPU andlibcuda.so.1is in the linker cache) goes on as usual, and "gpu_setup" in Rex::GPU then installs the container toolkit, CDI and containerd for it. Without a working driverinstall_driverdies after that probe and before anything on the host is changed:NVIDIA vGPU guest (type NVIDIA A10-2Q, 10de:2236 sub 14b9): install the licensed NVIDIA vGPU guest driver, then run again. No driver package was installed and no package source was addedThe same when a vGPU is passed together with a GPU that is not one (a passed-through card next to it): one NVIDIA kernel module drives every GPU of the host, so the vGPU guest driver and the driver Rex::GPU would install exclude each other; the message names both. With a working driver that mixed host also goes on as usual.
device_idis four hex digits, without0x(sysfsdevicereads0x2b85) and without a newline; any other defined value dies before the host is touched. See "gpus" in Rex::GPU::NVIDIA::Setup.install_driveritself never runslspcior installspciutils, with or withoutgpus. One driver has to drive them all: the driver is chosen for the intersection of their requirements (Rex::GPU::NVIDIA::Requirement: kernel module and driver-branch range, keyed on thedevice_id), andinstall_driverdies before anything on the host is changed when the GPUs cannot share one driver -- e.g. a V100 (proprietary module, 580 or older) next to a B200 (open module only) -- naming the GPUs on each side. gpu-
A single GPU hashref:
gpu => $gisgpus => [ $g ]. Kept for callers from beforegpus; passing both dies. setup-
Experimental. A Rex::GPU::NVIDIA::Setup class name or object to install with, instead of the class for the OS. Chosen in this order: this option, then
set gpu_nvidia_setup => ...in the Rexfile, then "setup_class_for_os" (see "setup_for"). A class name is loaded from@INC-- Rex puts thelib/directory next to the Rexfile there -- unless the package is already defined, e.g. in the Rexfile itself. An object is used as it is, except that one without GPUs of its own getsgpus(see "adopt" in Rex::GPU::NVIDIA::Setup). A module that is missing or does not compile, or a class that is not a Setup, dies before anything touches the host. One object serves one host: build it inside the task (it caches that host's facts); a second install with it dies. With a class of your own, an OS without a built-in class is no longer refused. How to write one: "WRITING YOUR OWN SETUP" in Rex::GPU::NVIDIA::Setup. requirement-
Experimental. An extra constraint on the driver: a hashref with any of
kernel_module(open,proprietary,either),min_branch,max_branch, or a Rex::GPU::NVIDIA::Requirement object. It is intersected with whatgpusneed, never replaces it:{ kernel_module => 'open' }moves an Ada to the open driver, but on a V100 (proprietary only) makesinstall_driverdie before anything on the host is changed, and a Kepler is refused whatever it says. An unknown key or a bad value dies before the host is touched. See "extra_requirement" in Rex::GPU::NVIDIA::Setup. nvswitches-
Optional arrayref of the host's NVSwitch chips, what "detect" in Rex::GPU::Detect returns under
nvswitch; "gpu_setup" in Rex::GPU passes it when it found one. Non-empty means an HGX baseboard whose GPUs need NVIDIA Fabric Manager:the driver source must provide one -- a source that does not (Debian
non-free, openSUSE's GFX repository) is rejected, and with none leftinstall_driverdies before anything on the host is changed;on apt its package must have an installation candidate after the index refresh, checked before the driver is installed;
after the driver packages are verified it is installed at exactly the installed driver's upstream version, checked with
dpkg-query/rpm -q(dies otherwise; the driver stays installed), andnvidia-fabricmanager.serviceis enabled;with
rebootthe enabled unit starts on boot; without it,install_driverstarts it aftermodprobe nvidia. Either way it then checkssystemctl is-activeand only warns if the unit is not running (e.g. nouveau still holds the GPUs until the reboot).
The package: Ubuntu
nvidia-fabricmanager-NNNof the chosen-serverbranch; Debian 12/13 and RHEL/Rocky/Almanvidia-fabricmanagerfrom NVIDIA's CUDA repository.If a working driver is already installed (e.g. an HGX host provisioned before Rex::GPU installed Fabric Manager), the driver is left alone and "retrofit_fabric_manager" in Rex::GPU::NVIDIA::Setup decides, host-read-only first:
a Fabric Manager package is already installed: nothing is changed; if its version is not the loaded driver's (
nvidia-smi --query-gpu=driver_version) it warns;none is: after
apt-get update(apt; dnf refreshes expired metadata itself) it asks the host's current package sources --apt-cache madison/dnf list --showduplicates-- for the same package name a fresh install would use, at exactly the loaded driver's version (on apt a simulated install must also remove nothing). Offered: installed, checked withdpkg-query/rpm -q, the unit enabled and started; a failure there dies, the driver untouched. Not offered, the version unreadable, or no package name known (openSUSE): it warns with the reason and the version needed, and installs nothing. No package source is ever added for this.
Then, as after an install, it warns if the unit is not active. Omitted or empty (the default, and every caller that finds its GPUs without
lspci): no Fabric Manager, nothing changes.HGX B200/B300 (karr #56): their NVSwitches are not PCI devices on the host, so there are no
nvswitches; they are recognised by the GPUs' device IDs instead (B2002901/2909, B3003182; "nvlink_platforms" in Rex::GPU::NVIDIA::Setup, fromgpus, so also for a caller without Rex::GPU::Detect), and everything above applies to them as to an NVSwitch host -- the driver source must provide Fabric Manager, it is installed at exactly the driver's version and its unit enabled and started. On top of that, after Fabric Manager:nvlsm(the NVLink Subnet Manager; no service of its own,nvidia-fabricmanager.servicestarts it),infiniband-diagsandlibibumad3(RHEL family:libibumad), plus on Ubuntulinux-modules-extraof the running kernel, unversioned -- the newest the repository has, as NVIDIA's own gpu-driver-container installs them -- throughapt-get/dnfand checked withdpkg -l/rpm -q(dies if one is missing; the driver and Fabric Manager stay installed);nvlsmcomes from NVIDIA's CUDA repository. On Debian 12/13 and the RHEL family that is where the driver came from. Ubuntu's driver comes from Ubuntu's archive, which has nonvlsm: the CUDA repository is added there, only on these hosts, after the driver and Fabric Manager are installed, with an apt pin that lets nothing butnvlsmcome from it (see "prepare_nvlink_fabric_source" in Rex::GPU::NVIDIA::Setup::Ubuntu);ib_umadis loaded (modprobe) and listed in/etc/modules-load.d/ib_umad.conf-- Fabric Manager's start script refuses to run without it;a running kernel older than 5.17 gets a warning (not on the RHEL family, whose 5.14 kernel NVIDIA supports for these boards); nothing stops;
after Fabric Manager is started (or the reboot),
nvidia-smi -qmust showFabric State: Completed,Status: Successfor every GPU, read up to 12 times 10 seconds apart while the unit is active ("check_nvlink_fabric" in Rex::GPU::NVIDIA::Setup). If not, one loud warning with what it read and where to look;install_driverdoes not die and "verify_nvidia" is not affected.
Where Rex::GPU knows no
nvlsmsource -- Ubuntu other than 22.04/24.04 or not amd64, RHEL before 9, Debian other than 12/13 -- and on openSUSE (no Fabric Manager source),install_driverdies before anything on the host is changed. With an already-installed driver the missing ones of those packages are installed from the host's current package sources only (no repository is added, as for Fabric Manager above; a package still missing only warns),ib_umadis loaded, the unit started if anything was installed or loaded, and the Fabric State is checked the same way. GB200/GB300 NVL72 compute trays get an info line instead: multi-node NVLink needsnvidia-imexand its configuration, which Rex::GPU does not set up.
Omit gpus and gpu (or pass undef) to keep the GPU-agnostic package selection.
Each distro's Rex::GPU::NVIDIA::Setup class has an ordered list of driver sources; the first that fits the requirement is installed, and if none fits, install_driver dies before anything is changed, listing every source and why it was rejected. What the requirement picks:
No constraint (Turing, Ampere, Ada, Hopper and every GPU the table does not know; no GPU): the first source -- Ubuntu the newest
-server, Debiannon-freenvidia-driver, RHELopen-dkms, openSUSE openG06/G07.Blackwell (open kernel module only, branch 570 or newer; GB10 and Blackwell Ultra 580 or newer): Ubuntu the newest
-server-open. On Debian no Debian-packaged driver fits (bookworm ships 535, trixie 550), so the driver comes from NVIDIA's CUDA apt repository instead ofnon-free: thecuda-keyringpackage fordebian12ordebian13(x86_64for amd64,sbsafor arm64), then the compute-only open-module setnvidia-driver-cuda+nvidia-kernel-open-dkms;non-freeis not enabled on that path. On any other Debian release or architecture it dies. RHEL and openSUSE: their default open driver.Maxwell, Pascal, Volta (
1340-1DF6, e.g. Tesla M60, P100, P40, V100): the proprietary driver of the 580 branch, their last. On Ubuntunvidia-driver-580-server; if apt has no candidate for it,install_driverdies before installing and does not fall back to another branch. On RHEL/Rocky/Alma 8 and 9 module streamnvidia-driver:580-dkms; on RHEL 10python3-dnf-plugin-versionlockand adnf versionlockon*nvidia*580*; both then installkmod-nvidia-latest-dkms+nvidia-driver+nvidia-driver-cuda, and verify the kmod and a 580nvidia-driver. On openSUSE Leap 15 and 16nvidia-driver-G06-kmp-meta, verified withrpm -q. On Debian 11/12/13 thenon-freedriver (470, 535, 550); on a Debian release without a knownnon-freebranch it dies.Kepler or older (device ID below
1340, e.g. Tesla K80/K40), as any one of the GPUs: the newest driver that supports it is the end-of-life 470 branch.install_driverdies on every distro before anything on the host is changed. A host whose driver was installed by hand (nvidia-smi -Llists the GPU andlibcuda.so.1is in the linker cache) passes the already-installed check above instead. "gpu_setup" in Rex::GPU passes only compute GPUs, and detection never counts a Kepler as compute, at any PCI class (karr #55): a Kepler display card and a class-0302Tesla K80/K40/K20 are skipped there with a warning and never reachinstall_driver. The refusal here is for a caller that passes one directly.
A mixed host gets what the combination needs: an Ada next to a B200 gets the open driver (Ubuntu -server-open), an Ada next to a V100 the proprietary 580 one (Ubuntu nvidia-driver-580-server).
install_driver(); # install only, load module without reboot
install_driver(reboot => 1); # install, reboot, verify
install_driver(gpus => [ grep { $_->{compute} } @{ $gpus->{nvidia} } ]);
install_driver(gpu => $gpus->{nvidia}[0]); # one GPU, the older form
install_driver(gpus => \@compute, setup => 'My::GPU::Setup');
install_driver(gpus => \@compute, requirement => { min_branch => 580 });
setup_class_for_os
my $class = Rex::GPU::NVIDIA->setup_class_for_os;
Experimental. The Rex::GPU::NVIDIA::Setup class "install_driver" uses on this host: Rex::GPU::NVIDIA::Setup::Ubuntu on Ubuntu, Rex::GPU::NVIDIA::Setup::Debian on every other Debian-family host, Rex::GPU::NVIDIA::Setup::RHEL on the RHEL family, Rex::GPU::NVIDIA::Setup::SUSE on openSUSE, and undef elsewhere ("install_driver" then dies). Asked only when neither the setup option nor set gpu_nvidia_setup chose a class (see "setup_for").
The RHEL family is what "is_redhat" in Rex::Commands::Gather accepts, plus the names Rex reports for Rocky Linux, AlmaLinux and CentOS Stream when lsb_release is installed and is_redhat does not know them: Rocky, RockyLinux, AlmaLinux, CentOSStream. "install_container_toolkit" recognises the same names. Reads nothing from the host.
setup_for
my $setup = Rex::GPU::NVIDIA->setup_for(
gpus => \@gpus,
setup => 'My::GPU::Setup', # optional
extra_requirement => { ... }, # optional
nvswitches => \@nvswitches, # optional
);
Experimental. The Rex::GPU::NVIDIA::Setup object "install_driver" runs, chosen in this order:
- 1.
setup-- a class name or an object ("install_driver"'ssetupoption); - 3. "setup_class_for_os".
A class is loaded ("custom_setup") and built with gpus, extra_requirement and nvswitches; an object gets them through "adopt" in Rex::GPU::NVIDIA::Setup. Returns undef only when nothing chose a class and the OS has none. Reads nothing from the host.
custom_setup
my $class_or_object = Rex::GPU::NVIDIA->custom_setup($setup_option);
Experimental. The setup the user chose -- $setup_option if defined, else set gpu_nvidia_setup -- validated, or undef for "choose by OS". A class name is loaded with "use_module" in Module::Runtime unless the package is already defined (e.g. written into the Rexfile itself), from @INC, where Rex puts the lib/ directory next to the Rexfile and in the current directory. Croaks, before anything touches the host, if the name is not a package name, the module cannot be found or does not compile, or the class or object is not a Rex::GPU::NVIDIA::Setup.
install_container_toolkit
Install the NVIDIA Container Toolkit (nvidia-container-toolkit package) from the official NVIDIA package repository at https://nvidia.github.io/libnvidia-container/.
Already installed toolkit. If nvidia-ctk --version runs and the package manager lists nvidia-container-toolkit as installed (dpkg -l ii/hi, or rpm -q) -- a re-run, or an image that ships the toolkit, such as NVIDIA DGX OS -- install_container_toolkit logs this and returns without touching the repository, the key or the package. The trade-off: a re-run no longer upgrades an installed toolkit to the repository's newest version; upgrade it with the package manager (apt-get install --only-upgrade nvidia-container-toolkit, dnf upgrade nvidia-container-toolkit, zypper update nvidia-container-toolkit). An nvidia-ctk that the package manager does not know about (e.g. unpacked by the GPU Operator) does not count: the package is installed, because "configure_containerd" relies on the packaged /usr/bin/nvidia-container-runtime.
Otherwise the repository GPG key is imported and the package repository is registered before installing. On Debian/Ubuntu the apt timers (unattended-upgrades, apt-daily, apt-daily-upgrade) are stopped first and every apt-get -- the curl/gnupg helpers included -- waits up to 120 seconds for the dpkg lock, as in "install_driver"; the key is downloaded and dearmored to a temporary file that replaces /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg only when it is non-empty, so a re-run refreshes a rotated key, and a failed download or dearmor dies. Then the signed APT source list is written. On RHEL the .repo file is downloaded with curl -f to a temporary file that replaces /etc/yum.repos.d/nvidia-container-toolkit.repo only when it contains the [nvidia-container-toolkit] section; a failed download (e.g. an HTTP error) or a file without that section dies and leaves an existing .repo as it was. On openSUSE Leap the base repository URL is added directly (zypper cannot parse RPM .repo files directly) and refreshed, replacing any existing nvidia-container-toolkit entry; a failed zypper addrepo or refresh (e.g. an HTTP error or an unresolvable host, which only the refresh reveals) dies before zypper install, and a repository that fails its refresh is removed again ("add_repo" in Rex::GPU::NVIDIA::Setup::SUSE). Every zypper call waits up to 120 seconds for the zypp lock ("zypper" in Rex::GPU::NVIDIA::Setup::SUSE).
The package is installed with apt-get/dnf/zypper directly, never through "pkg" in Rex::Commands::Pkg, and the result is checked with dpkg -l (ii) or rpm -q on every distro, openSUSE included.
Dies if the OS is not supported or if installation fails.
configure_containerd
configure_containerd($runtime);
Configure the containerd runtime to use the NVIDIA container runtime. The nvidia-container-runtime binary must already be installed ("install_container_toolkit" provides it); if it is not present this function returns immediately without error.
An unknown $runtime makes it die before any command runs on the host, naming the valid values -- with or without nvidia-container-runtime. none is not one of them here; that is "gpu_setup" in Rex::GPU's switch for not calling this function at all.
$runtime selects how containerd is configured:
rke2ork3s(default:rke2)-
Registers the NVIDIA runtime
additively, without replacing the base containerd config that RKE2/K3s generate. The mechanism is chosen from the effective, generatedconfig.tomlunder/var/lib/rancher/{rke2,k3s}/agent/etc/containerd/:Already wired — if that
config.tomlalready contains annvidiaruntime block (modern RKE2/K3s auto-detectnvidia-container-runtimeonPATHand wire it themselves), this is a no-op: nothing is written and the native config is left untouched.Modern (containerd 2.x / config v3) — writes an additive drop-in at
config-v3.toml.d/99-nvidia.toml(RKE2/K3s already importconfig-v3.toml.d/*.toml). The base config —SystemdCgroup, the pinned sandbox image, snapshotter options and the registrycerts.dpath — is preserved.Legacy (containerd 1.x / config v2) — writes a
config.toml.tmplthat begins with{{ template "base" . }}and only adds the nvidia runtime, so the rendered base config is preserved.
Both RKE2 and K3s use the same logic (only the
/var/lib/rancher/*base directory differs). If no generatedconfig.tomlexists yet and no config version marker is found, the modern v3 drop-in is written and a warning is logged.Healing an earlier clobber. A host set up by the 0.001 release carries a stale full-config
config.toml.tmpl— the bareimports = [...]+version = 2template (no{{ template "base" . }}) that replaced the distribution's base config. RKE2/K3s render that stale template, so the generatedconfig.tomlalready shows annvidiaruntime and the Already wired check above would no-op and leave the clobber (missingSystemdCgroup/ pinned sandbox /certs.d) in place. Before that check,configure_containerdtherefore removes thatconfig.toml.tmplonly when it matches the exact bare-clobber signature (never the base-extending template above, never a user's own customconfig.toml.tmpl, neverconfig-v3.toml.d/), then writes the additive v3 drop-in. Removing the template lets the distribution regenerate its native config, but the liveconfig.tomlstays clobbered until then: this function does not restart RKE2/K3s (that would bounce the node's containerd and its workloads); it logs a warning that the operator must restart the service or reboot the node for the native config to regenerate. containerd-
Calls
nvidia-ctk runtime configure --runtime=containerdand restarts thecontainerdsystemd service. Suitable for standalone (non-Rancher) containerd installations.
verify_nvidia
Verify the current NVIDIA installation by checking three things:
- 1.
nvidiakernel module is loaded (lsmod | grep nvidia) - 2.
nvidia-smi -Lreports at least one GPU - 3.
nvidia-ctkbinary is available (Container Toolkit present)
Returns 1 if all checks pass, 0 if any check fails. A warning is logged for each failure; the function does not die. A partial installation (e.g. driver installed but host not yet rebooted) emits a summary warning noting that features may not work until reboot.
verify_nvidia_driver
The driver half of "verify_nvidia", what "install_driver" runs after an install:
- 1.
nvidiakernel module is loaded (lsmod | grep nvidia) - 2.
nvidia-smi -Lreports at least one GPU - 3.
libcuda.so.1is in the linker cache ("libcuda_command" in Rex::GPU::NVIDIA::Setup) -- what "already_installed" in Rex::GPU::NVIDIA::Setup requires, so a failure here means the next run installs again
Does not look for the container toolkit. Returns 1 if all checks pass, 0 otherwise; logs a warning per failure and never dies. Not exported: call it as Rex::GPU::NVIDIA::verify_nvidia_driver().
generate_cdi_specs
Generate CDI (Container Device Interface) specifications for the detected NVIDIA GPUs so the Kubernetes NVIDIA device plugin can enumerate GPU resources without requiring a privileged container.
If a managed CDI source already owns the runtime scan dir, this function does not also write a static /etc/cdi/nvidia.yaml. Modern nvidia-container-toolkit ships nvidia-cdi-refresh.path/.service, which regenerate /run/cdi/nvidia.yaml and keep it fresh across driver updates. Because /etc/cdi and /run/cdi are both default CDI scan dirs, a second static copy would define the same device kind (nvidia.com/gpu) twice — a duplicate-device load error in CDI consumers — and would drift against the refreshed copy over driver updates. In that case the managed generator is triggered once (so /run/cdi is populated immediately) and left to own CDI.
Otherwise — no managed refresh unit and no existing /run/cdi/nvidia.yaml — output is written to /etc/cdi/nvidia.yaml via nvidia-ctk cdi generate, and the /etc/cdi/ directory is created if it does not exist. Only one of the two scan dirs is ever populated.
The managed-source check keys on the nvidia-cdi-refresh.path systemd unit state (installed/armed) rather than only on the presence of /run/cdi/nvidia.yaml, because /run is tmpfs and is empty right after the first-deploy reboot even though the refresh unit owns CDI from then on.
This step must be run after "install_container_toolkit" (which provides nvidia-ctk and the refresh unit) and, on first deploy, after the reboot that activates the NVIDIA kernel module (so the tool can enumerate physical devices).
MIG. The spec reflects the MIG layout at the time it is generated (MIG instances are included by default). Neither the static /etc/cdi/nvidia.yaml nor nvidia-cdi-refresh regenerates it when MIG mode or instances are reconfigured. After changing MIG, run systemctl restart nvidia-cdi-refresh.service, or on a host without that unit nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml. The MIG strategy Kubernetes exposes (single / mixed) is configured in the NVIDIA device plugin or GPU Operator, not here.
SEE ALSO
Rex::GPU, Rex::GPU::Detect, https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/
SUPPORT
Issues
Please report bugs and feature requests on GitHub at https://github.com/Getty/rex-gpu/issues.
CONTRIBUTING
Contributions are welcome! Please fork the repository and submit a pull request.
AUTHOR
Torsten Raudssus <getty@cpan.org>
COPYRIGHT AND LICENSE
This software is copyright (c) 2026 by Torsten Raudssus <torsten@raudssus.de> https://raudssus.de/.
This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.