PyTorch kernel-regression library for spectrum cartography.
-
Updated
Sep 15, 2026 - Python
PyTorch kernel-regression library for spectrum cartography.
Tesla V100 32GB (sm_70) running Qwen3.8-27B: sm70 decode kernel port plus KV context-cache tuning, measured on a real 53-request agent session. Decode 42.3-89.4 tok/s, TTFT 0.54 s on a cache hit, 200k-token prompts, zero failed requests, raw engine logs included. Published by an AI on the machine owner's behalf. 中文版:README.zh-CN.md
Wire the 1CatAI Split-D D256 FlashAttention kernel (fishlikeX/sm70-attn, MIT) into NInfer on Tesla V100 sm_70: +36-41% prefill, TTFT -3min, decode unchanged. Measured data + integration guide. Published by the user with AI assistance.
Data-driven INT8 KV-cache quantization: SQNR-guided granularity, near-lossless PPL (+0.2%), 0.50× memory, fused Triton int8 decode attention at 1.56× — gap to 2× fully explained.
To associate your repository with the attention-kernel topic, visit your repo's landing page and select "manage topics."