Popular repositories Loading
-
-
legacy-gpu-llm-notes
legacy-gpu-llm-notes PublicMeasured llama.cpp and quantization results on Tesla V100 (sm_70), Pascal GTX 1070, and RTX 4070 - hardware most projects don't test on
Python
-
volta-bonsai
volta-bonsai PublicTernary-Bonsai-2-27B on a Tesla V100 (sm_70): ternary kernels and vision both work on Volta
Shell
-
volta-gpt-oss
volta-gpt-oss Publicgpt-oss-20b MXFP4 runs on a Tesla V100 (sm_70) at 143 t/s, despite MXFP4 being documented as requiring compute capability 9.0+
-
volta-ik-llama
volta-ik-llama Publicik_llama.cpp builds clean and runs correctly on a Tesla V100 (sm_70) with no source changes
-
volta-dual-card
volta-dual-card PublicMixed sm_70 + sm_89 layer split in one llama.cpp job: works, cheap, no VRAM leak
If the problem persists, check the GitHub status page or contact support.