gemv
Here are 12 public repositories matching this topic...
PCCX™ specification, documentation, and ecosystem coordination hub for open AI accelerator IP.
-
Updated
May 31, 2026 - SystemVerilog
Matilda is a library to repeatedly multiply a constant matrix with a variable vector
-
Updated
Jan 26, 2026 - C++
High-performance FP32 GEMV for AMD RDNA 3 (gfx1100), with reproducible performance and numerical analysis.
-
Updated
Aug 31, 2026 - C++
PCCX™ v002 IP-core package — board- and model-agnostic reusable RTL for LLM, Vision, Voice, and common subsystems.
-
Updated
Jun 3, 2026 - SystemVerilog
Measure and visualize why LLM inference is slow: bottleneck analysis, model dissection, KV-cache, GEMM/GEMV, quantization, and memory-bound decoding.
-
Updated
Apr 30, 2026 - HTML
Hand-written CUDA kernels for batch-1 LLM decode: FP16 GEMV and online softmax at the measured memory-bandwidth limit, INT4/INT8 dequant GEMV turning compression into throughput. Nsight-profiled, correctness-gated.
-
Updated
Jul 6, 2026 - Cuda
🧮 CereMath is a library of Machine Learning kernels for the Wafer Scale Engine (WSE) from Cerebras. Made as a BSc thesis project
-
Updated
Apr 15, 2026 - Zig
InfiniTensor 训练营作业:OpenCL Q8_0 量化 GEMV 优化(RTX 4090 D 实测)
-
Updated
Aug 11, 2026 - C++
Benchmark workbench for MAX (Mojo) LLM decode kernels vs llama.cpp / cuBLAS / FlashInfer and the memory roofline on consumer NVIDIA GPUs (sm_86/sm_89). A public record, not a competing kernel library.
-
Updated
Sep 3, 2026 - Python
Add this topic to your repo
To associate your repository with the gemv topic, visit your repo's landing page and select "manage topics."