LLM observability and evaluation. Argus traces calls to OpenAI, Anthropic and Gemini, recording latency, tokens, cost and a few quality signals per request. Events go through Kafka into ClickHouse and are shown in Grafana. An eval pipeline scrapes documentation, builds a benchmark with an LLM judge and replays it against target models.
FastAPI, Kafka, ClickHouse, Grafana, Docker. Deploys to Cloud Run from GitHub Actions.
app + Argus SDK (provider wrapper -> tracer -> emitter)
| HTTP
v
collector (FastAPI) --> Kafka --> consumer --> ClickHouse <-- eval runner
| (scrape, generate,
v judge, replay)
Grafana, /v1/* API
Apps send events to the collector over HTTP by default, so they don't need a
Kafka client. The collector publishes to Kafka and the consumer batches
inserts into ClickHouse. ARGUS_EMITTER selects http (default), kafka
(publish directly) or noop. Emitting is best effort and never raises in the
host app.
No API keys needed; the demo seeds synthetic data.
docker compose up -d --build
docker compose run --rm api python -m scripts.init_clickhouse
docker compose run --rm api python -m scripts.seed_demo --hours 72 --rate 40Grafana is at http://localhost:3000 (anonymous access on, admin/admin). The collector API docs are at http://localhost:8000/docs.
from argus.providers.openai_provider import OpenAIProvider
llm = OpenAIProvider(default_model="gpt-4o-mini")
response = llm.complete("Summarize the theory of relativity in two sentences.")
print(response.text)Each call records latency, token usage, estimated cost, refusal detection and
JSON validity. AnthropicProvider and GeminiProvider work the same way.
Pass prompt_version="v2" to compare prompt revisions.
from argus.eval.scraper import scrape_sync
from argus.eval.benchmark import build_benchmark, save_benchmark
from argus.eval.runner import run_benchmark
from argus.providers.anthropic_provider import AnthropicProvider
judge = AnthropicProvider(default_model="claude-sonnet-4")
docs = scrape_sync(["https://en.wikipedia.org/wiki/Observability"])
items = build_benchmark(judge, docs, per_doc=5)
save_benchmark(items, "benchmarks/observability.jsonl")
target = AnthropicProvider(default_model="claude-haiku-4")
summary = run_benchmark(target, items, judge_provider=judge, benchmark="observability")
print(summary.pass_rate, summary.per_category)Eval results are stored in ClickHouse next to the traces.
| Method | Endpoint | |
|---|---|---|
| POST | /v1/traces |
ingest a TraceEvent |
| POST | /v1/evals |
ingest an EvalEvent |
| GET | /v1/metrics/cost?hours=24 |
cost by provider and model |
| GET | /v1/metrics/latency?hours=24 |
p50, p95, p99 latency |
| GET | /v1/metrics/volume?hours=24 |
request volume and errors |
| GET | /v1/metrics/quality?hours=24 |
refusal, error and JSON validity rates |
| GET | /v1/drift?metric=cost_usd |
drift analysis |
| GET | /healthz |
health check |
argus/analytics/drift.py compares a recent window to a baseline with the
Population Stability Index and a z-score on the mean. PSI below 0.10 is
stable, 0.10 to 0.25 moderate, above 0.25 significant. The demo data adds
latency and cost drift to OpenAI so the dashboard has something to show.
config/pricing.yaml has placeholder prices for local development and may not
match current provider pricing. Models are matched exactly, then by longest
prefix. Unknown models get zero cost and a warning.
argus/ analytics, collector, eval, ingest, providers,
config, emitter, events, pricing, quality, tracer
scripts/ init_clickhouse.py, seed_demo.py
grafana/ dashboards and provisioning
deploy/ Cloud Run script and service config
pip install -e ".[dev]" # core + pytest, ruff
pytest
pip install -e ".[all]" # every provider and service dependencyCI runs ruff and the tests, builds the image, and deploys the collector to
Cloud Run on pushes to main. To deploy by hand:
export GCP_PROJECT_ID=...
export GCP_REGION=us-central1
export AR_REPO=argus
./deploy/deploy.shKafka and ClickHouse run separately; point the collector at them with
ARGUS_KAFKA_BOOTSTRAP and ARGUS_CH_HOST.
MIT