[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
-
Updated
Jan 17, 2026 - Cuda
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
LLM algorithm practice lab with theory, solutions, and test cases.《大模型算法与系统教程》面向大模型入门到进阶的算法实战教程,覆盖原理讲解、答案解析、测试用例与 CUDA/Triton 实战。
🌱 A tiny, readable LLM serving engine with vLLM/SGLang-style features.
Deterministic intermediate representation for AI agents — compile, verify, execute, and replay structured intent.
A lightweight Bun + Express template that connects to the Testune AI API and streams chat responses in real time using Server-Sent Events (SSE)
AI gateway and observability infrastructure for production LLM systems
Add a description, image, and links to the llm-infra topic page so that developers can more easily learn about it.
To associate your repository with the llm-infra topic, visit your repo's landing page and select "manage topics."