Multi-GPU tensor-parallel vLLM on AMD Radeon RX 7900 XT / XTX / GRE (7900XT, 7900XTX, RX7900XT, gfx1100, RDNA3, ROCm): root cause and fix for the RCCL hostcall / PCIe atomics (AtomicOps) crash "NCCL error: unhandled cuda error" / "operation cannot be performed in the present state", Proxmox VFIO passthrough; LLM inference benchmarks on 13 machines
benchmark amd hip proxmox rccl multi-gpu vfio rocm radeon gpu-passthrough tensor-parallelism vllm 7900xtx llm-inference gfx1100 h100 rdna3 pcie-atomics 7900xt rx7900xt
-
Updated
Sep 17, 2026 - Python