HIP/ROCm fork optimized for AMD RDNA2 (gfx1030) with PrismML Q1_0_G128 1-bit quant support, RotorQuant, TurboQuant, EAGLE3 and P-EAGLE speculative decoding, and full Wave32 kernel optimizations.
-
Updated
Jun 23, 2026 - C++
HIP/ROCm fork optimized for AMD RDNA2 (gfx1030) with PrismML Q1_0_G128 1-bit quant support, RotorQuant, TurboQuant, EAGLE3 and P-EAGLE speculative decoding, and full Wave32 kernel optimizations.
AMD ROCm (gfx1030) inference fork with RotorQuant/TurboQuant KV compression, PHANTOM-X zero-copy draft speculation, EAGLE3 speculative decoding, 12 RDNA2 crash fixes, and PrismML Bonsai Q1_0_G128 1-bit GGUF support.
Experimental open-source compatibility research for running DLSS Neural Rendering workloads on AMD RDNA2/gfx1030. Includes ISA translation, host emulation, proof-oriented validation, and bounded hardware testing. Early development: no end-to-end game frame yet.
Exact FlashAttention for AMD RDNA2 (gfx1030 / RX 6800 XT), written in HIP. Drop-in replacement for flash-attn: 19.8 TFLOP/s forward, decode at 99% of HBM bandwidth.
Custom SPIR-V kernel factory for PHANTOM speculative decoding — LLVM IR to GPU (SPIR-V/HIP) and CPU (native x86) cross-target compilation, RDNA2/gfx1030 optimized pre-compiled kernels, dynamic kernel swapping, zero-JIT inference pipeline
ComfyUI + ZLUDA on Windows for AMD RDNA2 desktop GPUs (RX 6950/6900/6800 XT, RX 6700/6750 XT) — fixes HIP/ZLUDA version mismatches, torch DLL loading, the mem-efficient attention backend that resets the display driver, and installs rocBLAS kernels on gfx1031.
27B Qwen3.8-class uncensored ternary model on a 16GB GPU: 256K context + MTP + vision | Qwen3.8 系 27B 无审查三值量化 · 256K 上下文 · 16GB 显卡实测 | RX 6900 XT (gfx1030)
To associate your repository with the gfx1030 topic, visit your repo's landing page and select "manage topics."