You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.
A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/Linux with full GPU capability
Source-only offline deployment, testing, and operations toolkit for DeepSeek-V4-Flash-0731 on 4x/8x NVIDIA A100 GPUs, pinned to a reviewed community vLLM R1 stack.
Complete GB10 (DGX Spark, sm_121) block-FP8 Triton kernel configs for DeepSeek-V4-Flash-0731 at TP=2, including the DSpark main_proj shape missing from every earlier run