Skip to content
View davetha's full-sized avatar

Block or report davetha

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Popular repositories Loading

  1. aiter-cdna2 aiter-cdna2 Public

    Run AMD AITER's hand-written ASM kernels on CDNA2 / gfx90a (MI210, MI250). 242 of 1,422 kernels translated, with the tests and benchmarks to prove which actually run.

    Python 10

  2. mi210-llm-stack mi210-llm-stack Public

    AMD MI210 (gfx90a/CDNA2) LLM inference optimization — TurboQuant, KIVI, per-layer KV types, FlashAttention, MoE expert caching

    Shell 9

  3. r9700-lru-expert-cache r9700-lru-expert-cache Public

    Device-side LRU expert cache + kernel-count patches: Qwen3.8-Flash-Next decode on 2x AMD R9700 (gfx1201), ROCm 10, tcclaviger vLLM fork

    Python 4 1

  4. cve-af-alg-block cve-af-alg-block Public

    Shell 3

  5. vllm-int8-moe-rocm vllm-int8-moe-rocm Public

    vLLM refuses INT8 MoE on every AMD GPU because of a CUDA-only check. One-line fix + benchmark harness. On MI210: 3.20s TTFT vs 5.07s for AWQ-Int4.

    Python 3

  6. mi210-vllm mi210-vllm Public

    Pinned, gate-verified vLLM deployment stack for AMD MI210 (gfx90a/CDNA2)

    Python 3