cmp-170hx
Here are 8 public repositories matching this topic...
GLM-5.3-Flash (320B MoE) served with vLLM on 8x CMP 170HX (SM80, 64 GB, PCIe Gen2 x4): pipeline-parallel 8, NVFP4 weights, DFlash2 speculative decoding, 1M context. Patches, launch config and benchmarks.
-
Updated
Aug 31, 2026 - Python
Running large LLMs on pre-Ampere NVIDIA hardware — Tesla V100 (sm_70), RTX 2080 Ti (sm_75), CMP 170HX. Measured benchmarks, vLLM forks, and the hardware side: NVLink on SXM2 carrier boards, driver traps, cooling, used-kit acceptance.
-
Updated
Aug 19, 2026
Capacity-aware single-GPU SGLang benchmarks with MTP A/B, core/extended context matrices, raw telemetry, model/KV memory capture, and generated Markdown/PDF reports.
-
Updated
Aug 1, 2026 - Python
vLLM for NVIDIA CMP 170HX (Ampere sm_80): fast, repeatable model serving on one card or several
-
Updated
Oct 1, 2026 - Python
64 GB of HBM2e on a dead mining card: running Qwen3-Next-80B on an NVIDIA CMP 170HX with vLLM. Configs, gotchas, and measured benchmarks.
-
Updated
Sep 22, 2026 - Python
Add this topic to your repo
To associate your repository with the cmp-170hx topic, visit your repo's landing page and select "manage topics."