Skip to main content
Crusoe Support Help Center home page
Crusoe

VM Hangs With Kernel Soft Lockup in the NFS Write Path on Shared Storage

Rasul Imanov
Rasul Imanov
Updated

Introduction

A VM with a Crusoe Shared Disk (NFS) mounted may become unresponsive with a continuously repeating kernel soft lockup. One CPU becomes pinned in the kernel's NFS write path, and any process that subsequently writes to the affected file enters uninterruptible I/O ("D" state) and cannot be killed or timed out.

On Slurm clusters, this typically surfaces as slurmd becoming unresponsive and the controller marking the node down, with in-flight jobs lost.

This condition has been observed on versions of the VAST NFS client driver (vastnfs) earlier than 4.5.7. In our investigation of affected VMs, kernel logs show a writeback worker stuck in the page-group locking code of the kernel's NFS write path (nfs_lock_and_join_requests): the same CPU spins there indefinitely — in one observed case for over 31 hours — and the condition does not self-resolve until the VM is reset. The exact code-level root cause has not been conclusively identified. In practice, upgrading the driver to 4.5.7 or later has resolved the issue: affected nodes have had no recurrence after upgrading.

The evidence points to the NFS client driver rather than the shared disk itself, the storage backend, or your workload.

Prerequisites

  • Crusoe Shared Disk (NFS) Mounted on One or More VMs
  • VAST NFS Client Driver (vastnfs DKMS Module) Installed, as Included in Curated Images Such as ubuntu24.04-nvidia-slurm
  • Root or sudo Access on the Affected VM
  • Maintenance Window With the NFS Volume Unmounted and No Jobs Running (Driver Upgrade Only)

Symptoms

  • The kernel log shows repeating soft lockup messages with a trace through the NFS write path, for example a CPU stuck in nfs_page_set_headlock or nfs_lock_and_join_requests.
  • Worker processes are stuck in uninterruptible I/O ("D" state) writing to the NFS mount, and cannot be killed with kill -9.
  • On Slurm clusters, slurmd stops responding, the controller marks the node down, and jobs on the node are marked failed.
  • The Modules linked in: line of the lockup trace tags the NFS modules with (OE) — or run cat /sys/module/nfs/taint, which returns OE — confirming the out-of-tree vastnfs driver is in use rather than the in-tree client.

Instructions

Step 1: Confirm the Driver Version

Run the following on the affected VM:

 
vastnfs-ctl status

This issue has only been observed on versions earlier than 4.5.7. If your version is 4.5.7 or later and the VM shows similar symptoms, you are seeing a different issue — capture the diagnostics in Step 3 and contact Crusoe Support.

Step 2: Check for the Lockup Signature

Review the kernel log:

 
sudo dmesg | grep -i "soft lockup"
sudo journalctl -k --since "-24h" | grep -A 20 "soft lockup"

If the VM no longer accepts SSH, use the serial console instead — note that a local user password must already be set on the VM for console login, so it is worth setting one before you need it.

A trace referencing nfs_page_set_headlock or nfs_lock_and_join_requests, with processes in "D" state, matches this issue. Both conditions should match — the driver version and the lockup signature. A driver version match alone is not a diagnosis.

Step 3: Capture Diagnostics Before Resetting

⚠️ Warning: Resetting the VM destroys the evidence. Capture the diagnostics below before you reset — without them the root cause cannot be confirmed after the fact.

Save the kernel journal and the process state:

 
sudo journalctl -k --since "-48h" > ~/kernel-log.txt
ps -eo pid,tgid,ppid,stat,wchan:32,cmd | awk '$4 ~ /D/' > ~/dstate.txt

Optionally, dump the kernel stacks of all blocked tasks into the journal before capturing it — this records exactly where each stuck process is waiting:

 
echo w | sudo tee /proc/sysrq-trigger

Copy both files off the VM (for example with scp) before resetting, and attach them to your support ticket. This allows Crusoe Support to confirm the root cause rather than infer it.

Step 4: Reset the VM

Processes in uninterruptible I/O cannot be killed, so a reset is the only way to clear the condition. Reset the VM from the Crusoe Cloud console, or with the CLI:

 
crusoe compute vms reset <vm-name>

Step 5: Upgrade the Driver

A curated image containing the updated driver is in progress. Until it ships, install the updated driver manually. The current published package is 4.5.8.

⚠️ Known issue with 4.5.8: on NFSv3 mounts, vastnfs 4.5.8 can return "Permission denied" when opening an existing file owned by another user with a create flag (for example > or >> redirection, or lock files). Workaround: run stat on the file or ls -l on its directory first. If multiple users write to the same files on your shared storage, install 4.5.7 instead — it resolves the lockup and predates this regression — or contact Crusoe Support.

⚠️ Warning: This upgrade must be applied per node and needs a quiet maintenance window — no jobs running and the NFS volume unmounted. Reloading the driver under active I/O can hang the node again.

On each node, in order:

  1. Drain the node (Slurm):
 
scontrol update nodename=<node> state=drain reason="vastnfs driver upgrade"
  1. Once jobs have drained, unmount the NFS volume.
  2. Download and install the package (apt pulls in dkms and kernel headers if missing):
 
wget https://raw.githubusercontent.com/crusoecloud/crusoe-nfs-support/main/debs/vastnfs-dkms_4.5.8-vastdata_all.deb
sudo apt install ./vastnfs-dkms_4.5.8-vastdata_all.deb
  1. Refresh the initramfs and reload the driver:
 
sudo update-initramfs -u -k $(uname -r)
sudo vastnfs-ctl reload
  1. Remount the NFS volume and verify the new version is active:
 
vastnfs-ctl status
  1. Return the node to service:
 
scontrol update nodename=<node> state=resume

Packages are published in the crusoe-nfs-support repository.

Resolution

After upgrading to vastnfs 4.5.7 or later, this lockup has not recurred in our experience: nodes stay available under the same workloads that previously hung them. If a node on 4.5.7 or later becomes unresponsive with NFS-related symptoms, treat it as a new issue — capture diagnostics and open a support ticket rather than assuming this defect.

Before upgrading, note the known 4.5.8 issue described in Step 5.

On Slurm clusters, it is also worth confirming requeue behavior so a future node failure does not silently kill jobs:

 
scontrol show config | grep Requeue

Check whether jobs are submitted with --no-requeue, or via srun, which does not requeue. Requeue restarts the job from the beginning (or its last checkpoint), instead of losing it.

ℹ️ Note: An updated curated image with the updated driver is planned. This article will be updated when it is available.

Example

A Slurm training job runs on a node writing checkpoints and logs to a shared NFS mount. Partway through the run the node stops responding to scontrol, and the controller marks it down. The job is lost.

On the serial console, dmesg shows a soft lockup repeating every few seconds with nfs_page_set_headlock in the trace, and ps shows worker processes in "D" state on files on the NFS mount. kill -9 does nothing. vastnfs-ctl status reports version 4.5.1.

The team captures the journal and the D-state process list, resets the VM, and upgrades the driver during their next maintenance window. The lockup does not recur after the upgrade.

Related Articles

Additional Resources

Related to

Was this article helpful?

0 out of 0 found this helpful

Still need help?

Our support team is ready to assist you with any questions.

Have more questions? Submit a request

Recently Viewed

Comments

0 comments

Article is closed for comments.