See what changed in your agent application, using summaries instead of raw traces.
Did token usage rise because you made more requests, or because each request used more tokens? Which tracked contributors changed? Is the comparison missing usage data? fleetdiff reads small summary files and answers locally. Start with one application; combine compatible exports when you add workers or separately operated systems. Keep your existing trace backend. No account, upload, or model API key is needed.
git clone https://github.com/llm-measurement/fleetdiff.git
cd fleetdiff
sh examples/investigate.shNeeds Git and Go 1.25 or 1.26 with a current security patch, on Linux or macOS. No Docker. The command builds current source and prints a single-application report from synthetic collector exports:
Reported tokens: 200 -> 600
Model attempts: 2 -> 3
Tokens per attempt: 100.00 -> 200.00
Attempt-count contribution: +150.00 tokens
Tokens-per-attempt contribution: +250.00 tokens
This is an arithmetic split, not proof of cause. Try
sh examples/investigate.sh --missing-usage: one missing usage field makes the
explanation cannot_determine, not a guessed saving. Add --two-stacks to run
the same investigation over disjoint gateway and direct-call exports.
See the example and its limits.
Model attempts are observed model spans, including failures and retries, not
unique user requests. Existing JSON fields keep their requests names.
investigate is included in release v0.2.0 and later.
No Go compiler is needed to use a release binary. Download the
v0.2.0 archive
for Linux or macOS, on AMD64 or ARM64. Follow the
download and verification instructions
before extracting it. Then run ./fleetdiff --version or investigate your exports:
./fleetdiff investigate --before ./before --after ./after --expected appDid token usage jump after an application change? Send LiteLLM traces to the
collector recipe,
then use fleetdiff investigate to compare its before-and-after summary exports.
See whether the increase came from more requests or more recorded tokens per
request, with missing usage shown. Keep your existing tracing backend.
The recipe documents the tested LiteLLM version and an optional callback that preserves missing provider usage for supported non-streaming responses. Streaming usage provenance remains unknown: LiteLLM can supply estimates when provider counts are absent. The comparison is not invoice reconciliation or proof of savings.
Run sh examples/demo.sh for the existing two-operator walkthrough.
Watch the one-minute two-operator walkthrough.
A research supervisor delegates to an internal-document specialist your team runs and a research specialist a partner runs. This is the supervisor-and-specialists pattern described by Anthropic and LangChain. After a deployment:
1. Did total reported usage fall?
Our reported tokens: 800 -> 360 (down 440)
Fleet reported tokens: 1200 -> 1800 (up 600)
2. Did model activity rise while runs stayed flat?
Model requests: 6 -> 10
Observed root-agent runs: 2 -> 2
Requests missing token usage: 2 -> 2
Your system reported less. The fleet reported more. Reported token usage shifted toward the partner and increased overall. Whether the answers got better is outside what fleetdiff measures. The data is synthetic, not provider traffic or evidence of savings.
| Question | In the demo |
|---|---|
| More requests or more tokens per request? | Single-app report: +150 and +250 tokens respectively; refused if usage is incomplete |
| Which tracked contributors changed? | Prompt-weight delta bounds and shares of recorded sketch weight, not a guaranteed top-k ranking |
| Can it identify a runaway session? | cannot_determine: current exports do not attribute model tokens to sessions |
| Did total reported usage fall, or just move? | Ours 800 to 360 tokens; fleet 1,200 to 1,800 |
| Did model activity rise while runs stayed flat? | 6 to 10 model requests; 2 root-agent runs in both windows |
| Did the workload touch more documents? | About 1 to 4 distinct MCP resources; a shared one counts once |
| Did particular tool-error signatures increase? | One rises from 1 to 4; one appears; one is unchanged |
| Is part of the fleet missing? | A missing export is refused, or reported as explicitly partial |
Operator A Operator B
traces -> collector traces -> collector
| |
summary file summary file counters + keyed sketches,
\ / no raw values
'------> fleetdiff compare <-' runs locally, read-only
|
fleet report (text or JSON)
Each operator keeps its own trace backend. The OpenTelemetry collector or llm-sketchkit exports a summary per time window. Replayed summaries do not double-count. Shared identities count once with compatible hashing settings. Operators must avoid counting the same requests twice: fleetdiff cannot deduplicate requests observed by different operators. It will not call a comparison complete when an expected operator is missing.
With Docker running, sh examples/demo.sh --live starts two pinned, released
collectors, sends the same synthetic traffic, and compares their real exports.
It checks every answer above, plus propagated trace context and privacy scans
for raw fields. See the two-operator walkthrough.
From the checkout, build the command if you skipped the demo:
go build -o bin/fleetdiff ./cmd/fleetdiff- Each operator enables summary export and agrees on time windows, hashing settings, and who observes which requests.
- Put one window's files in
before/and the later window's files inafter/. - Investigate, naming every expected producer, even if one export is missing:
bin/fleetdiff investigate --before ./before --after ./after --expected appAdd --format json for automation. fleetdiff only reads files; it never modifies
inputs or makes network requests. Sharing an export still requires
authorization. For multiple disjoint producers, use --expected team,partner.
The lower-level compare command remains unchanged. See the
investigation contract and, for separate operators, the
two-operator trial checklist.
- Exact: request, token, and agent-run counts, for observed spans.
- Estimated: distinct users, sessions, and resources, with a nominal error.
- Bounded: changes in prompt and tool-error signatures, with lower and upper bounds.
- Missing stays missing: requests without token usage are counted as missing, never as zero.
Reported tokens are not an invoice or a measure of useful work, and a
before-and-after difference does not prove its cause. More activity does not
prove retries, over-delegation, or better answers. Agent runs count only
invoke_agent spans with no parent, so an agent under an HTTP request or
workflow span is not counted. See the
agent and MCP caveats.
Hashes are keyed and pseudonymous, not anonymous; treat exports and reports as
sensitive.
- Adversarial and fuzz tests on malformed inputs, with bounded file, byte, and directory-entry limits.
- Weekly dependency and security scans.
- Release binaries for Linux and macOS on AMD64 and ARM64, with SBOM and provenance, tested on native runners.
- Resource measurements from one machine, with the commands to reproduce them.
Run the checks yourself with go test -race ./... and go vet ./....
- FAQ: inputs, accuracy, privacy, and troubleshooting
- Comparison contract
- Investigation questions and JSON API
- Operations: installation verification, offline use, and upgrades
- Resource measurements: sizing on one machine
- Security policy · Changelog · Releasing
The 0.2.x release line provides local, read-only investigate and compare commands.
Apache-2.0. Code authors: Vijay and Codex.
