Mixed-precision acceleration patch for MiniMax H3 inference on NVIDIA V100
-
Updated
Sep 6, 2026 - Python
Mixed-precision acceleration patch for MiniMax H3 inference on NVIDIA V100
Fork of llama.cpp Nvida Volta V100 (tensor parallelism 4 x v100 GPU)
V100-optimized mixed precision for Qwen Image in ComfyUI — 1.43x faster sampling on NVIDIA Tesla V100 32GB (Volta/sm_70).
VirtualV LLM benchmark testsuite: local hardware-aware model evaluation (GGUF/1Cat-vLLM), governance standard, and results
To associate your repository with the nvidia-v100 topic, visit your repo's landing page and select "manage topics."