Frequently Asked Questions (FAQ)
-
FAQ: Persistent Disk Offering
Promoted
-
FAQ: Common dataset errors in Serverless Fine-Tuning
-
FAQ: Local NVMe Drive Monitoring and Diagnostics
-
FAQ: How Do I Identify the Physical Hardware Behind My GPU VM?
-
FAQ: How Node Pool Configuration Changes Affect Existing Nodes
-
FAQ: Storage Quotas on VAST Shared Volumes
Solutions
-
Getting Started with SLURM on Crusoe Cloud
Promoted
-
Slurm jobs fail with "ImportError: libGL.so.1" after node restarts
-
Slurm login node restarts with OOMKilled (exit code 137)
-
API key fails with "401 Authentication failed" after user removal
-
Move Ethernet NIC Interrupts Off Your Training CPUs in Crusoe GPU VMs
-
GPU Operator pods stuck on B200/B300 CMK nodes due to MOFED race condition
How-To
-
How-To Capture NVIDIA Bug Report via Command Center
Promoted
-
How To Enable GPU Direct Storage (GDS) on Crusoe GPU Instances
Promoted
-
How To Download Large Models from AWS S3 to Local NVMe Using rclone
Promoted
-
How-To Setup GB200 NVL72 Rack on CMK Cluster and Run NCCL Performance Validation
Promoted
-
How-To Validate Infiniband Performance with NCCL All Reduce Test
Promoted
-
How-To Setup Serial-Console to Access Your VM
Promoted