Skip to content
@kvcache-ai

kvcache.ai

KVCache.AI is a joint research project between MADSys and top industry collaborators, focusing on efficient LLM serving.

Pinned Loading

  1. Mooncake Mooncake Public

    Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

    C++ 5.9k 985

  2. ktransformers ktransformers Public

    A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

    Python 18.7k 1.5k

  3. TrEnv-X TrEnv-X Public

    Go 95 8

Repositories

Showing 10 of 15 repositories
  • sglang Public Forked from sgl-project/sglang

    SGLang is a fast serving framework for large language models and vision language models.

    kvcache-ai/sglang's past year of commit activity
    Python 14 Apache-2.0 7,338 0 12 Updated Jul 20, 2026
  • Mooncake Public

    Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

    kvcache-ai/Mooncake's past year of commit activity
    C++ 5,927 Apache-2.0 985 217 (1 issue needs help) 259 Updated Jul 20, 2026
  • ktransformers Public

    A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

    kvcache-ai/ktransformers's past year of commit activity
    Python 18,740 Apache-2.0 1,463 451 (1 issue needs help) 13 Updated Jul 20, 2026
  • kvcache-blog Public
    kvcache-ai/kvcache-blog's past year of commit activity
    JavaScript 18 MIT 13 0 1 Updated Jul 20, 2026
  • vllm Public Forked from vllm-project/vllm

    A high-throughput and memory-efficient inference and serving engine for LLMs

    kvcache-ai/vllm's past year of commit activity
    Python 16 Apache-2.0 19,875 0 0 Updated Jul 1, 2026
  • Model-Optimizer Public Forked from NVIDIA/Model-Optimizer

    A unified library of SOTA model optimization techniques like quantization, pruning, distillation, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

    kvcache-ai/Model-Optimizer's past year of commit activity
    Python 1 Apache-2.0 509 0 0 Updated Jun 17, 2026
  • accelerate Public Forked from huggingface/accelerate

    🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support

    kvcache-ai/accelerate's past year of commit activity
    Python 2 Apache-2.0 1,445 0 1 Updated May 9, 2026
  • transformers Public Forked from huggingface/transformers

    🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

    kvcache-ai/transformers's past year of commit activity
    Python 1 Apache-2.0 34,682 0 1 Updated May 9, 2026
  • sglang_awq Public Forked from sgl-project/sglang

    SGLang is a fast serving framework for large language models and vision language models.

    kvcache-ai/sglang_awq's past year of commit activity
    Python 2 Apache-2.0 7,342 0 0 Updated Apr 22, 2026
  • evalscope Public Forked from modelscope/evalscope

    A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

    kvcache-ai/evalscope's past year of commit activity
    Python 0 Apache-2.0 429 0 0 Updated Apr 13, 2026

Top languages

Loading…

Most used topics

Loading…