Hand-written NVFP4 W4A16 CUDA kernels for Volta
-
Updated
Aug 19, 2026 - Python
Hand-written NVFP4 W4A16 CUDA kernels for Volta
Benchmarks and notes for running modern LLMs with vLLM on 8x Tesla V100-32GB in 2026.
Run Qwen3.6-27B on four Tesla V100s at 366 tok/s using hand-written NVFP4 CUDA kernels and chain-MTP speculation.
llama.cpp optimized for NVIDIA Tesla V100 (Volta, SM70): 2–6 GPU tensor parallelism, Qwen3.8-27B Q8_0/Q4 with DFlash2 speculative decoding, up to 512K context, multimodal (image/PDF/video) and concurrent serving.
Qwen3.8-27B in native NVFP4/FP8 on 2x PCIe Tesla V100-32GB (SM70): the PCIe runbook for v100-skinny + 1Cat-vLLM, with the 3 fixes that make it work without NVLink. 61-74 tok/s decode, MTP speculative decoding, OpenAI-compatible.
Qwen3.8-Flash-Next (125B MoE, NVFP4) on 4x V100-SXM2-32GB and DeepSeek-V4.1-flash on 8x V100 — a Volta port of SGLang for agentic coding.
Multi-GPU acceleration for MiniMax H3 video generation on NVIDIA V100 (sm_70). Ulysses sequence parallelism as a drop-in ComfyUI custom node — ~19 min to ~7 min on 8x V100.
Reproducible llama.cpp kernel and runtime optimization lab for dual NVIDIA Tesla V100 GPUs (SM70)
Runbook + benchmarks: Qwen3.8-Flash-Next-ABLITERATED NVFP4 on 4× Tesla V100-32GB (reflashed SXM2→PCIe, 2+2 NVLink + PLX). 1Cat-vLLM 1.5.0, TP4 — 262,144-token context validated, 46 tok/s decode, 122 tok/s aggregate at 4 concurrent streams. Full E0–E17 optimization log with measured evidence.
SGLang fork for IBM POWER9 (ppc64le): Tesla V100 sm70, CUDA 12.4, Granite LLM inference. Triton attention, float16, OpenAI-compatible API.
DP4A FlashAttention-2 and GEMM for Volta GPUs. 46 TOP/s INT8 on CMP 100-210 where tensor cores are firmware-disabled. The .superl8 format loads weights at memory speed.
V100 (sm_70) tuning kit for Ternary Bonsai 2 27B: q8_0-KV flash-attention direct read, D256 Split-D prefill kernel, PTQ1_0 planar mat-vec, MTP graft + draft micro-batch fix, ubatch presets. Prefill 175 -> 854 t/s, VRAM -1.9 GiB, PPL bit-identical (2-chunk). 中英双语 README (README.md / README.zh-CN.md).
Tesla V100 32GB (sm_70) running Qwen3.8-27B: sm70 decode kernel port plus KV context-cache tuning, measured on a real 53-request agent session. Decode 42.3-89.4 tok/s, TTFT 0.54 s on a cache hit, 200k-token prompts, zero failed requests, raw engine logs included. Published by an AI on the machine owner's behalf. 中文版:README.zh-CN.md
Serve Qwen3.5-397B-A17B (AWQ) on 8x Tesla V100-SXM2-32GB (DGX-1, TP8) for agentic coding & ops — a downstream fork of 1Cat-vLLM.
VastLLM: a production-oriented FastLLM fork for native C++ inference, V100/SM70, long context, and Qwen3.8/3.6/3.5 series; upstream: ztxz16/fastllm
Wire the 1CatAI Split-D D256 FlashAttention kernel (fishlikeX/sm70-attn, MIT) into NInfer on Tesla V100 sm_70: +36-41% prefill, TTFT -3min, decode unchanged. Measured data + integration guide. Published by the user with AI assistance.
PyTorch 2.12 fork for IBM POWER9/POWER10 (ppc64le) with CUDA 12.4 and Triton. Tesla V100 sm70, GPU training and LLM inference.
Measured llama.cpp and quantization results on Tesla V100 (sm_70), Pascal GTX 1070, and RTX 4070 - hardware most projects don't test on
To associate your repository with the sm70 topic, visit your repo's landing page and select "manage topics."