Lynn GitHub 镜像仓 · Primary repository: https://github.com/MerkyorLynn/Lynn · Downloads: https://download.merkyorlynn.com/download.html
-
Updated
Aug 22, 2026 - TypeScript
Lynn GitHub 镜像仓 · Primary repository: https://github.com/MerkyorLynn/Lynn · Downloads: https://download.merkyorlynn.com/download.html
Benchmark of local LLMs (Qwen 3.6, Gemma 4) vs ChatGPT and Gemini on coding tasks. Apple M5, 32GB. Methodology, raw outputs, judge prompts, scores, and charts.
Pokedex benchmark for LLMs
Saotri Bench — coding benchmark for evaluating LLM agents on multi-phase programming tasks with hidden requirements.
面向 AI IDE / 编码 Agent 的编程能力基准测试框架,内置防背题三防线:公私用例分离、泛化差距检测、参数化生成
Raw logs of Claude Code running on local Qwen3.5-27B (llama.cpp). Builds a Python todo app with 50 tests. Real-world performance data: 30 min, cache thrashing, 38 t/s generation.
31 local and hosted LLMs tested with one difficult falling-block browser game prom31 local and hosted LLMs tested with one difficult falling-block browser game prompt. Runnable outputs and hands-on findings.pt. Runnable outputs and hands-on findings.
Add a description, image, and links to the coding-benchmark topic page so that developers can more easily learn about it.
To associate your repository with the coding-benchmark topic, visit your repo's landing page and select "manage topics."