AlmaLinux 9 VMs hang on reboot after updates

For about three months now, we’ve been having problems with AlmaLinux 9 VMs not rebooting reliably after installing Updates.

Unfortunately, it does not happen every time and not all VMs have had this behavior. And it does not happen when we reboot a VM without installing updates. That makes it difficult to debug.

It’s always VMs running on OpenStack based clouds, although different clouds from different providers. It’s been happening for VMs I run privately that have been created from the official AlmaLinux 9 cloud images as well as for VMs that we are running at work that have been created from private internal images.

The behavior is always the same:

  • We run the reboot command.
  • The VM will initiate the shutdown / reboot procedure
  • but will hang itself during reboot.
  • Sending an ACPI Shut down will finally shut down the VM all the way.

Wehn VM is hung, we cannot connect to the VM via SSH and when openting the TTY, no input is registered. Pressing cltr+alt+delet, which usually initiates a reboot, does nothing.

If you have any ideas for how we could try to reproduce the error, or how to debug it without it being reliable, please let us know.

Here is the journal of an affected VM: Sep 10 03:15:26 myserver.example.com automated-maintenance[1778813]: + rebootS - Pastebin.com

Hello

I reviewed the logs.
It appears the VM is hanging due to a failure to free graphics memory.
The cause seems to be a bug in the qxl driver.

I believe you can reproduce the error by using the qxl driver in QEMU-KVM.
Also, as a workaround, using the VirtIO driver seemed to resolve the issue.

Reference site

thank you


Sep 10 03:15:26 myserver.example.com systemd[1]: Stopping Record System Boot/Shutdown in UTMP..

Sep 10 03:34:34 myserver.example.com kernel:  fb_pan_display+0x87/0x110
Sep 10 03:34:34 myserver.example.com kernel:  bit_update_start+0x1a/0x40
Sep 10 03:34:34 myserver.example.com kernel:  fbcon_switch+0x330/0x4b0

Sep 10 03:36:37 myserver.example.com kernel:  drm_modeset_lock+0x4a/0x110 [drm]
Sep 10 03:36:37 myserver.example.com kernel:  drm_atomic_get_plane_state+0x7a/0x170 [drm]
Sep 10 03:36:37 myserver.example.com kernel:  drm_client_modeset_commit_atomic+0xaa/0x230 [drm]
Sep 10 03:36:37 myserver.example.com kernel:  drm_client_modeset_commit_locked+0x56/0x160 [drm]

Sep 10 03:50:33 myserver.example.com kernel: [TTM] Buffer eviction failed
Sep 10 03:50:33 myserver.example.com kernel: qxl 0000:00:01.0: object_init failed for (3149824, 0x00000001)
Sep 10 03:50:33 myserver.example.com kernel: [drm:qxl_alloc_bo_reserved [qxl]] *ERROR* failed to allocate VRAM BO

Hi. Thank you so much for the links. This has helped a lot in explaining what’s going on and allowed us to look into work arounds.

Do I understand correctly, that the issue is with the guest OS’ kernel (the kernel running inside the VM) and not the carrier’s kernel?

Also, do you happen to know if this is a known issue? Should I report it in the hopes that AlmaLinux will include a fix, or is that hopeless?

Hello

The error logs are being output by the guest OS kernel, and the issue is surfacing in the guest OS’s qxl kernel driver.

However, I believe the root cause is highly likely to be a bug in the interaction with the host-side QEMU.

The qxl issue is a known problem.

It seems to be discussed in the linked Proxmox forum as well.

The qxl driver appears to be particularly unstable when combined with newer kernels.

Also, as a contributor, I don’t know whether AlmaLinux will address this.

If you wish to report the bug,
Please try reporting the bug here:

The linked thread on the Proxmox Support Forum makes a few mentions. It took me a bit to find and verify them all. So in order to spare any who come after me the same work, here is what I have found:

This is a bug in the kernel running inside the VMs.

  • It was initially introduced with commit 5a838e5d5825 «drm/qxl: simplify qxl_fence_wait».
  • Later the issues we’re having have been discovered and described in Bug#1054514.
  • This bug was resolved in commit 07ed11afb68d «Revert “drm/qxl: simplify qxl_fence_wait”» by reverting the first commit.
  • This, however, caused new issues, described in this E-Mail thread here
  • As a result of that, the initial commit was then re-applied in commit 3628e0383dd3 «Reapply “drm/qxl: simplify qxl_fence_wait”»
    • When doing so, the reality of re-introducing a known bug, was willingly accepted.
    • During this discussion, several kernel developers expressed the opinion, that qxl should no longer be used, most notably David Airlie from RedHat the discussion spawned from the above commit.

I can’t recommend anyone use qxl hw over virtio-gpu hw in their VMs, since virtio-gpu is actually hw designed for virt.

The initial bug linked above contains a useful script to reproduce the problems.

P.s. I would have linked all the commits, but the forum software won’t let me put more than two links in my post, so…

Hello,

Thank you for your confirmation.
In that case, I believe a temporary workaround would be to disable the qxl module on the VM side.

When I tried the following, the qxl module was not loaded.

※ Please verify this thoroughly yourself if you attempt it.

[root@alma9 modprobe.d]# echo "install qxl /bin/true" > /etc/modprobe.d/disable-qxl.conf
[root@alma9 modprobe.d]# modprobe -v qxl
insmod /lib/modules/5.14.0-570.46.1.el9_6.x86_64/kernel/drivers/gpu/drm/ttm/ttm.ko.xz
insmod /lib/modules/5.14.0-570.46.1.el9_6.x86_64/kernel/drivers/gpu/drm/drm_ttm_helper.ko.xz
install /bin/true

I created a Jira issue.