NAME
Rex::GPU::Detect - GPU hardware detection via PCI class codes
VERSION
version 0.002
SYNOPSIS
use Rex::GPU::Detect;
my $gpus = detect();
if (@{ $gpus->{nvidia} }) {
for my $gpu (@{ $gpus->{nvidia} }) {
printf "NVIDIA %s (class %s, compute: %s)\n",
$gpu->{name}, $gpu->{pci_class}, $gpu->{compute} ? 'yes' : 'no';
}
}
DESCRIPTION
Rex::GPU::Detect detects GPU hardware on a remote host by parsing lspci -nn output and matching PCI vendor and class codes.
Detection approach
PCI class codes 0300 (VGA compatible controller) and 0302 (3D controller) identify display/GPU hardware. The module filters lspci -nn output for these class codes, then classifies devices by vendor ID:
10de— NVIDIA1002— AMD / ATI
Virtual GPU filtering
Display devices with vendor IDs 1af4 (virtio), 1b36 (QEMU/QXL), 15ad (VMware), or 80ee (VirtualBox) are skipped line by line; they need no host driver. Skipping one does not end the scan: on a VM with a passed-through GPU (vfio-pci) the emulated console display and the real card appear side by side, and the real card is still detected. A VM whose display devices are all virtual returns empty arrays, as before. The vendor checks run first, so a 10de/1002 line is never treated as virtual — this also means an NVIDIA vGPU guest device (vendor 10de) is detected like a passed-through card; lspci -nn cannot tell the two apart. The subsystem ID can, see "NVIDIA vGPU guests".
NVIDIA vGPU guests
A VM on an NVIDIA vGPU (a slice of a physical GPU: Azure NVadsA10 v5, AWS G6f, a vGPU on VMware or KVM) sees a PCI device with the physical GPU's vendor and device ID, the same [10de:XXXX] line in lspci -nn as the card itself. What differs is the subsystem ID: NVIDIA gives every vGPU type its own, and publishes the pairs in its open GPU kernel modules. "detect" reads them with lspci -vmmnn -d 10de: (run only when an NVIDIA GPU was found) and looks each GPU's device ID, subsystem vendor and subsystem ID up in Rex::GPU::NVIDIA::VGPU (1135 pairs, Turing to Blackwell Ultra, from NVIDIA's open-gpu-kernel-modules 615.71.09):
a known pair with subsystem vendor
10de:vgpu => 1andvgpu_type(e.g.GRID A100X-1-5C,NVIDIA A10-2Q);anything else -- a physical or passed-through card (none of the 616 physical device/subsystem pairs NVIDIA's 615.71.09 README lists is in the vGPU table), a vGPU type newer than the table, a slot
lspci -vmmnndid not list:vgpu => 0, detection as before.
compute stays what the generation says: a vGPU of a compute GPU is compute. A vGPU guest needs NVIDIA's licensed vGPU guest driver, not the datacenter driver "install_driver" in Rex::GPU::NVIDIA installs (whose open kernel module refuses an Ampere-or-newer vGPU); install_driver therefore dies for one before it changes the host, unless a working driver is already there.
NVIDIA compute classification
NVIDIA GPUs are further classified as compute-capable. Only compute-capable GPUs trigger driver installation in Rex::GPU. Every GPU that can be used for AI counts -- GeForce MX, GT and GTX 9xx with 2 GB of memory included -- as long as a current driver branch supports it: the criterion is the GPU generation, read from the PCI device ID, not the marketing name. The rules, first match wins:
A Kepler-or-older device ID (below
1340, see below) — not compute, whatever the PCI class. This is checked first (karr #55), so a Kepler Tesla (K80, K40, K20), which enumerates as class0302, is skipped with the Kepler warning like a Kepler display card, and a newer GPU on the same host is still installed.PCI class
0302(3D controller) — compute/datacenter. Datacenter GPUs such as the A100, H100, and RTX 4000 Ada typically enumerate as class0302. A class-0302device whose ID no table row covers, or that has no ID, is compute by its class alone.The PCI device ID's generation, from the generations table of Rex::GPU::NVIDIA::Requirement (its compute flag), whatever the PCI class and whatever name
lspciprints. The table covers every ID from0000to2FFFand the Blackwell Ultra IDs:Maxwell, Pascal, Volta (
1340-1DF6; GeForce GTX 750 Ti/9xx/10xx, GT 1030, MX110-MX350, Tesla M/P/V100, ...): compute. They get the proprietary 580-branch driver.Turing, Ampere, Ada, Hopper (
1DF7-28FF; MX450/MX550/MX570, GTX 16xx, RTX 20xx-40xx, T4, A100, L40, H100, ...): compute.Blackwell and Blackwell Ultra (
2900-2FFF,3182,31C2/31C3; GeForce RTX 50xx desktop and laptop, RTX PRO Blackwell, B200/GB200/B300/GB300, the GB1010de:2e12of NVIDIA DGX Spark): compute.Kepler or older (below
1340; GeForce GT 710/730, GTX 6xx/7xx, Quadro K4000, Tesla K20/K40/K80, ...): not compute (checked before the class rule, see above). Their last driver branch is 470, which the current distributions no longer package, so the GPU is skipped with a warning ("... is Kepler or older silicon: it needs driver branch 470 or older, which current distributions no longer package -- skipped, no driver installed") instead of making the driver installation die. A Kepler card next to a newer GPU does not stop the newer one's installation.
Many GPUs enumerate as a VGA controller (class
0300), and on a host whosepci.idspredates the siliconlspciprints onlyDeviceas the name; the device ID is present regardless.Name rules, only for an ID no table row covers (
3000and up, except the Blackwell Ultra IDs -- silicon newer than the table): RTX, GTX 10xx/16xx, GTX 9xx/745/750, GT 1xxx, GeForce MX1xx-5xx, TITAN X/Xp/V/RTX, Quadro M/P/T/GP/GV, Tesla M/P/V/T4 and datacenter short codes (A100, H100, L40, ...). Each names only Maxwell-or-later products.
Unrecognised NVIDIA GPU models default to compute => 0 and emit a warning. AMD GPU compute is always 0; AMD driver support is not yet implemented.
Each detected NVIDIA GPU also carries its raw device_id (the [10de:XXXX] field, or undef if lspci printed none). Rex::GPU passes every compute GPU hashref through to "install_driver" in Rex::GPU::NVIDIA, which chooses the driver from the device IDs through Rex::GPU::NVIDIA::Requirement: the open kernel module for Blackwell-architecture silicon (B200/GB200/B300, GeForce RTX 50xx, RTX PRO Blackwell, GB10), which has no proprietary one, the proprietary 580 branch for a pre-Turing GPU (Maxwell/Pascal/Volta, e.g. the V100), and a refusal for a Kepler-or-older one passed to it directly.
FUNCTIONS
detect
Detect GPU hardware on the remote host: parses lspci -nn output filtered to PCI display-class devices (class codes 03xx).
If lspci is on the remote PATH (command -v lspci), nothing is installed. Otherwise pciutils is installed first: through "pkg" in Rex::Commands::Pkg (unless is_installed says it already is), or, on Rocky Linux, AlmaLinux and CentOS Stream under the names Rex reports when lsb_release is installed (Rocky, RockyLinux, AlmaLinux, CentOSStream; Rex::Pkg cannot handle those), with dnf install -y pciutils checked by rpm -q pciutils. Dies if that check fails or lspci is still not found afterwards -- before lspci runs, so a host without it never reports "no GPU".
Returns a hashref with three array refs, always all present: nvidia and amd (one hashref per detected GPU) and nvswitch (one hashref per detected NVSwitch, see below):
{
nvidia => [
{
name => "AD104GL [RTX 4000 SFF Ada Generation]",
vendor => "nvidia",
pci_class => "0302", # "0300" = VGA controller, "0302" = 3D controller
compute => 1, # 1 if CUDA-capable, 0 otherwise
device_id => "27b0", # [10de:XXXX]; undef if lspci printed no vendor:device pair
subsystem_vendor_id => "10de", # from lspci -vmmnn; undef if not found
subsystem_id => "16fa", # likewise
vgpu => 0, # 1 for an NVIDIA vGPU guest device, see below
# vgpu_type => "NVIDIA A10-2Q", # only when vgpu is 1
}
],
amd => [
{
name => "Navi 31 [Radeon RX 7900 XTX]",
vendor => "amd",
pci_class => "0300",
compute => 0, # AMD compute support not yet implemented
}
],
nvswitch => [
{
name => "GH100 [H100 NVSwitch]",
vendor => "nvidia",
pci_class => "0680",
device_id => "22a3",
}
],
}
nvswitch lists the NVSwitch chips of an HGX baseboard (NVIDIA 10de devices of PCI class 0680, "Bridge"), found by a second, read-only lspci -nn -d 10de: that runs only when an NVIDIA GPU was found; it is [] otherwise. A device counts only if its ID is a known NVSwitch (1ac2 HGX-2, 1af1 HGX A100, 22a3 HGX H100/H200) or lspci names it ... NVSwitch; another NVIDIA bridge device is logged and skipped. An NVSwitch host needs NVIDIA Fabric Manager, which "gpu_setup" in Rex::GPU installs with the driver. HGX B200/B300 NVSwitches are not detected: they are not PCI devices on the host (NVIDIA's Fabric Manager guide), so nvswitch stays [] there. Their NVLink fabric is recognised by the GPU device IDs instead, when the driver is installed ("nvlink_platforms" in Rex::GPU::NVIDIA::Setup).
When an NVIDIA GPU was found, a third read-only command, lspci -vmmnn -d 10de:, reads each NVIDIA device's subsystem IDs by PCI slot. Every nvidia element gets subsystem_vendor_id and subsystem_id (undef if that output has no record for its slot), and vgpu: 1 if the pair of device ID and subsystem ID is an NVIDIA vGPU type, 0 otherwise; with 1 also vgpu_type, NVIDIA's name for the type. See "NVIDIA vGPU guests". compute is not affected by it. A host without an NVIDIA GPU runs neither this command nor the NVSwitch one.
If no supported GPU is found, or if the only display devices are virtual, all three arrays are empty ([]) -- nvswitch too, since it is only probed when an NVIDIA GPU was found. A virtual display next to a real NVIDIA/AMD card (GPU passthrough, cloud GPU VM) is skipped on its own line and the real card is still reported.
open_kernel_module_required
Rex::GPU::Detect::open_kernel_module_required($device_id);
Given an NVIDIA PCI device ID (the XXXX in [10de:XXXX], lowercase or uppercase), returns true if that device is known to have no proprietary kernel module at all — NVIDIA's open GPU kernel modules are the only option: every Blackwell-architecture part, on any CPU architecture. True for an ID in the Blackwell device-ID ranges taken from NVIDIA's open-gpu-kernel-modules supported-GPU table (2900-2FFF: B200, GB200, GeForce RTX 50xx, RTX PRO Blackwell, GB10; plus B300 3182 and GB300 31C2/31C3). Returns false for undef, a malformed ID, and every ID outside those ranges — Turing/Ampere/Ada/Hopper parts and any future generation keep the default proprietary -server selection. This function only answers the driver-variant question; which GPUs are compute-capable is "NVIDIA compute classification".
A wrapper: true exactly when "for_device_id" in Rex::GPU::NVIDIA::Requirement gives kernel_module open. The device-ID ranges live in that class's generations table only, so the driver installer carries no second hardcoded device list.
Not in @EXPORT — this is a Rex::GPU::NVIDIA-internal lookup, not a Rexfile-facing command.
legacy_driver_requirement
my $legacy = Rex::GPU::Detect::legacy_driver_requirement($device_id);
# { generation => 'Maxwell/Pascal/Volta', max_branch => 580 } or undef
Given an NVIDIA PCI device ID (the XXXX in [10de:XXXX], any case), returns a hashref for a pre-Turing GPU that current NVIDIA drivers no longer support: generation (a label for messages) and max_branch, the newest driver branch that still does. These GPUs work only with NVIDIA's proprietary kernel module; the open module does not support them.
1340-1DF6— Maxwell, Pascal and Volta (Tesla M60/M40, P100, P40, P4, V100, V100S, TITAN V, GeForce 9xx/10xx, ...):max_branch580.below
1340— Kepler (0FC6-12BA, Tesla K80/K40) and older (Fermi and earlier):max_branch470.
Returns undef for undef, a malformed ID, and every ID from 1DF7 up (Turing and every later or unknown generation), which keep the default driver selection. The ranges are taken from the legacy sections of NVIDIA's supportedchips README (driver 615.71.09). This only chooses the driver; whether a GPU is compute-capable is "NVIDIA compute classification" (Maxwell/Pascal/Volta: yes, Kepler or older: no).
A wrapper over "for_device_id" in Rex::GPU::NVIDIA::Requirement: a hashref of its generation and max_branch when the requirement has a max_branch, undef otherwise.
Not in @EXPORT — a Rex::GPU::NVIDIA-internal lookup.
SEE ALSO
Rex::GPU, Rex::GPU::NVIDIA, https://pci-ids.ucw.cz/ (PCI ID database)
SUPPORT
Issues
Please report bugs and feature requests on GitHub at https://github.com/Getty/rex-gpu/issues.
CONTRIBUTING
Contributions are welcome! Please fork the repository and submit a pull request.
AUTHOR
Torsten Raudssus <getty@cpan.org>
COPYRIGHT AND LICENSE
This software is copyright (c) 2026 by Torsten Raudssus <torsten@raudssus.de> https://raudssus.de/.
This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.