FastAPI proxy that strips volatile fields from OpenClaw requests to dramatically improve llama-server KV cache hit rates (~22× faster prompt eval)
-
Updated
Feb 23, 2026 - Python
FastAPI proxy that strips volatile fields from OpenClaw requests to dramatically improve llama-server KV cache hit rates (~22× faster prompt eval)
Privacy-first , fast Windows dictation with local Whisper on CPU, NVIDIA CUDA, or Vulkan GPUs, plus optional Groq, OpenAI, > and Gemini cloud services.
To associate your repository with the amd-vulkan topic, visit your repo's landing page and select "manage topics."