Create tokcalc.json - #1670
Create tokcalc.json#1670stevecrates489-commits wants to merge 1 commit into
Conversation
|
Hi @stevecrates489-commits 👋 — thanks for your interest in the MCP Registry! It looks like this PR is trying to publish an MCP server by adding or editing server files in this repository (under Servers are published with the To publish your server, follow the Publishing Quickstart. In short: # Build the publisher CLI
make publisher
# Authenticate (e.g. via GitHub) and publish your server.json
./bin/mcp-publisher login github
./bin/mcp-publisher publishNote: If you believe this was closed in error, please leave a comment and a maintainer will take a look. 🙇 This is an automated message. |
Summary of changes
Adds
tokcalc— an open-source LLM serving capacity planner MCP server with 6 read-only tools for AI agents (Cursor, Claude Desktop, Cline).Motivation and Context
LLM engineers and platform teams need to answer deployment questions that simple "tokens per second" calculators cannot: model fit in VRAM, KV-cache pressure at different context lengths, multi-GPU topology recommendations, continuous batching multipliers, prompt caching economics, and build-vs-buy break-even analysis.
Currently, AI agents helping with LLM infrastructure planning have no tool to call for deterministic capacity calculations. They hallucinate GPU specs, invent throughput numbers, and miss critical constraints like KV-cache memory limits at long context.
tokcalcsolves this by exposing a bounded, deterministic, read-only calculation engine as an MCP server.The server covers 35 models (Llama 3/4, Qwen 2/3, DeepSeek V3/R1, Mixtral, Gemma, Phi, etc.), 30 GPUs (H100/H200/B200, A100, RTX 4090/5090, AMD MI300X, Intel Gaudi 3, TPU, Groq LPU, Apple Silicon), and 16 quantization formats (FP16/FP8, GGUF Q2_K–Q8_0, GPTQ, AWQ, EXL2, NVFP4).
How Has This Been Tested?
npx -y @tokcalc/mcp-server, tested with prompts like "List the GPUs tokcalc supports" → agent calledlist_gpusand received 30 GPU records. Also tested "Recommend a topology for Llama 3.3 70B at 32K context" → agent calledlist_models→recommend_topology→ returned feasible topology options.@tokcalc/mcp-server@0.1.2, installable vianpx.formatInputSchema()with$refStrategy: "none"and strip the top-level$schemameta-key, ensuring strict MCP client compatibility.Breaking Changes
None. This is a new server with no prior versions in the registry.
Types of changes
Checklist
Additional context
Architecture: The MCP server shares its calculation engine (
src/lib/token-calc.ts) with the tokcalc web app (Next.js). The engine is pure TypeScript with no network access — all calculations are deterministic and offline. The MCP server bundles this engine into a single self-contained file (~0.66 MB, 207 modules) usingbun build --target=node.Tool design principles (per Perplexity MCP research):
formatInputSchema()for strict MCP client compatibilityserver.json is included at the repository root with:
name:dev.tokcalc/capacity-plannerpackages: npm@tokcalc/mcp-serverv0.1.2, stdio transportrepository: GitHub (stevecrates489-commits/tokcalc)npm:
@tokcalc/mcp-server@0.1.2—npx -y @tokcalc/mcp-serverworks out of the boxLicense: Apache 2.0 (same as the main tokcalc project)
Live demo: https://tokcalc.vercel.app