The same set of changes described here have also been applied in the NVIDIA driver packages published by NVIDIA as part of the CUDA repositories. The changes will land in the 620 branch releases.

The same set of changes will land at the same time in the official driver installation guide.

DNF 5 reboot suggestion

The kernel module and the user space libraries must always be the same version, so after a driver update you really want to reboot. With DNF 4 on RHEL/EL 8 to 10, the packages have been adding themselves to the needs-restarting DNF plugin configuration for a while. DNF 5 did not have an equivalent until recently, but luckily my merge request to DNF 5 got merged without issues so we now have a configurable list of packages that can trigger the notification.

The snippets reside in /usr/share/dnf5/suggest-reboot.d/ and are available in Fedora (and eventually RHEL 11); after an update to the driver, “dnf5 needs-restarting” and the reboot hint at the end of the transaction tell you that a reboot is needed, like for a kernel update. This is also available in the 580 branch driver packages, the LTS driver that still supports older GPUs.

NVIDIA drivers on a system without an NVIDIA GPU

The NVIDIA driver packages can now be installed on any system, whether or not it has an NVIDIA GPU. That includes OS images that end up on very different hardware, and machines that only get an NVIDIA GPU later in their life, for example a laptop that gets plugged into a Thunderbolt dock or eGPU enclosure with an NVIDIA card inside.

Until now the packages assumed that if you installed them, you had the hardware. On a system without an NVIDIA GPU you would get services that fail at every boot, an autostart entry that complains at every login, and so on. None of it was harmful, but it made the driver a bad fit for generic images. Everything that needs the GPU now stays “dormant” and waits for the GPU to appear in the system.

Most of the changes are variations on the work done for Anatase (https://anatase.org/). Thanks to Antheas Kapenekakis for figuring it all out.

Kernel modules

The kernel modules are built as usual, through akmods (akmod-nvidia) or DKMS (dkms-nvidia), whether or not a GPU is present. They are not included in the initrd, and they are only loaded when the kernel finds a matching PCI device through its modalias. On a system without an NVIDIA GPU they just sit on disk. When a GPU appears, even after boot, udev loads the driver.

On top of the open kernel modules, nvidia-kmod and dkms-nvidia now carry two extra patches from Anatase:

  • Tunneled PCIe link speed. A GPU attached over Thunderbolt or USB4 sits behind a tunneled PCIe link, and the link often does not train to the best speed the tunnel supports. With this patch, whenever such a GPU is in use (at startup, and every time it wakes up from runtime power management), the driver stops the hardware from changing the link speed on its own at both ends of the link. It then asks the kernel’s PCI core to train the link once at the maximum speed available, and gives control back to the hardware when the GPU goes idle again. It is only enabled for Thunderbolt-attached GPUs and on kernels that export pcie_set_target_speed(), so internal GPUs are not affected. This is what makes the “plug in a dock with an NVIDIA GPU” case perform properly and not just work.
  • CEC over DisplayPort to HDMI adapters. This one is not about GPU detection, but it is a nice addition: it exposes the DisplayPort AUX channel to the kernel’s DRM helpers, so HDMI-CEC works through DP-to-HDMI adapters that support it. For example, you can control the Steam interface with a remote or turn on the TV once your HTPC set powers on.

I’m trying to push both patches internally in NVIDIA, let’s see if we can get them in.

Auotostart of programs

Previously the packages shipped systemd presets that enabled nvidia-persistenced and nvidia-powerd. An enabled service starts at every boot, GPU or not. The presets are gone. Both packages now ship a udev rule that pulls in the service when an NVIDIA display controller (VGA or 3D class) is added to the system:

ACTION=="add", SUBSYSTEM=="pci", ATTR{vendor}=="0x10de", ATTR{class}=="0x030000|0x030200", TAG+="systemd", ENV{SYSTEMD_WANTS}+="nvidia-persistenced.service"

On a system without an NVIDIA GPU the services never start. On a system with one they start at boot as before. When a GPU is hot-plugged, they start the moment it appears. There is nothing to enable or disable by hand.

nvidia-settings installs an autostart entry that restores the saved settings at login with “nvidia-settings --load-config-only“. Without an NVIDIA GPU this just fails at every login. The entry now checks for the device first:

Exec=sh -c "[ -e /dev/nvidia0 ] && exec /usr/bin/nvidia-settings --load-config-only"

No GPU means no device node, so that covers nvidia-settings as well.

CUDA in unprivileged containers

CUDA also needs the /dev/nvidia-uvm and /dev/nvidia-uvm-tools device nodes. They are normally created on demand: the NVIDIA libraries call nvidia-modprobe, a setuid helper that loads the nvidia-uvm module and creates the nodes.

Unprivileged containers (rootless Podman, Toolbx, Distrobox) can not call setuid programs on the system, so CUDA fails even if the rest of the GPU is passed through correctly. This is the standard mechanism for most of the non open source libraries that are part of the driver to make sure the correct device nodes are created upon execution.

nvidia-kmod-common now does this in advance with a udev rule. When an NVIDIA GPU is added, it runs “nvidia-modprobe -u -c 0“, which loads nvidia-uvm and creates both nodes. The nodes are then already there to pass into containers.