Skip to content
#

eviction

Here are 34 public repositories matching this topic...

Custom CUDA kernels for KV-cache eviction + INT8 quantized paged attention (vLLM-oriented), PyTorch C++ extensions with Python API: eviction @16k blocks 1707.75us host -> 562us fused (~3x); INT8 attention p50 1.76-14.18ms; ~49.9% KV memory vs FP16; 26/26 pytest passing on T1000.

  • Updated Jul 14, 2026
  • Python

Analytical benchmark for sliding window attention KV cache management: quality vs window size tradeoffs, SWA vs eviction comparison, prefix sharing interaction, and operational window recommendations across four attention distributions

  • Updated Jul 26, 2026
  • Python

Improve this page

Add a description, image, and links to the eviction topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the eviction topic, visit your repo's landing page and select "manage topics."

Learn more