Stop overpaying to run your agents. Kalibr routes every request to lower-cost model and tool paths without degrading performance.
-
Updated
Jun 3, 2026 - Python
Stop overpaying to run your agents. Kalibr routes every request to lower-cost model and tool paths without degrading performance.
Measures dollars per correct outcome on LLMs. Contamination-resistant: benchmark questions are generated fresh at runtime.
What LLM inference actually costs. 874 verified price rows across 24 providers, 378 GPU rental rates, 395 cited throughput datapoints, and a self-hosting break-even model. Every number carries a source URL, a retrieval date, and passed a mechanical evidence gate. Dated immutable snapshots. CC BY 4.0.
Caveman prompting, measured. A two-channel evaluation protocol scoring what input and output compression actually cost LLMs in dollars, accuracy, and surface-text fidelity across seven models and five benchmarks.
CachePin
Drop-in OpenAI-compatible proxy that routes each request to the best model by cost, quality, or latency — with a pluggable scoring engine and a built-in eval harness. Self-hosted.
Tamper-evident, stranger-verifiable receipts for LLM cost-savings — anyone can recompute your caching/routing savings math offline, no trust in your dashboard required. Pure stdlib, zero-dependency.
AI Infrastructure Portfolio: agent orchestration, model routing, inference cost analysis — working code, real measurements.
Inference cost allocation for autonomous AI agent collaborations — Shapley-fair splitting, congestion pricing, token metering. Part of the Agent Trust Stack.
Cost-aware LLM gateway: semantic cache + difficulty-based model routing that cuts spend (~66%) and latency vs always using the frontier model, measured against a baseline and gated in CI. Fully offline.
Budget-aware degradation for autonomous AI agents — map a spend balance to survival tiers, enforce per-call, hourly and daily LLM inference budgets, and automatically downgrade the model when funds run low.
Fit more signal into fewer tokens. Prompt and RAG context compression with a fidelity evaluator: 48% fewer tokens at 100% fact recall.
Locational Cost of Intelligence: a location-adjusted cost function with QoS chance constraints, with code and data schemas for an Intelligence Price Deflator. Headline empirical result withdrawn pending reconstruction, see STATUS.md.
Inference cost allocation for autonomous AI agent collaborations — Shapley-fair splitting, congestion pricing, token metering. Part of the Agent Trust Stack.
Bittensor subnet 92. A decentralised efficiency layer for autonomous agents, producing a Verified Agent Runtime scored on verified work per dollar.
NeoSmith Maestro — frontier-level coding accuracy from a bouquet of small models at ~1/30th the cost. Technical report + reproducible evidence (LiveCodeBench v6 92.2%, SWE-bench Pro 74.2/88.9).
Add a description, image, and links to the inference-cost topic page so that developers can more easily learn about it.
To associate your repository with the inference-cost topic, visit your repo's landing page and select "manage topics."