Skip to content

Repository files navigation

Argus

LLM observability and evaluation. Argus traces calls to OpenAI, Anthropic and Gemini, recording latency, tokens, cost and a few quality signals per request. Events go through Kafka into ClickHouse and are shown in Grafana. An eval pipeline scrapes documentation, builds a benchmark with an LLM judge and replays it against target models.

FastAPI, Kafka, ClickHouse, Grafana, Docker. Deploys to Cloud Run from GitHub Actions.

Architecture

app + Argus SDK (provider wrapper -> tracer -> emitter)
        | HTTP
        v
collector (FastAPI) --> Kafka --> consumer --> ClickHouse <-- eval runner
                                                   |          (scrape, generate,
                                                   v           judge, replay)
                                          Grafana, /v1/* API

Apps send events to the collector over HTTP by default, so they don't need a Kafka client. The collector publishes to Kafka and the consumer batches inserts into ClickHouse. ARGUS_EMITTER selects http (default), kafka (publish directly) or noop. Emitting is best effort and never raises in the host app.

Quick start

No API keys needed; the demo seeds synthetic data.

docker compose up -d --build
docker compose run --rm api python -m scripts.init_clickhouse
docker compose run --rm api python -m scripts.seed_demo --hours 72 --rate 40

Grafana is at http://localhost:3000 (anonymous access on, admin/admin). The collector API docs are at http://localhost:8000/docs.

Tracing calls

from argus.providers.openai_provider import OpenAIProvider

llm = OpenAIProvider(default_model="gpt-4o-mini")
response = llm.complete("Summarize the theory of relativity in two sentences.")
print(response.text)

Each call records latency, token usage, estimated cost, refusal detection and JSON validity. AnthropicProvider and GeminiProvider work the same way. Pass prompt_version="v2" to compare prompt revisions.

Evals

from argus.eval.scraper import scrape_sync
from argus.eval.benchmark import build_benchmark, save_benchmark
from argus.eval.runner import run_benchmark
from argus.providers.anthropic_provider import AnthropicProvider

judge = AnthropicProvider(default_model="claude-sonnet-4")
docs = scrape_sync(["https://en.wikipedia.org/wiki/Observability"])
items = build_benchmark(judge, docs, per_doc=5)
save_benchmark(items, "benchmarks/observability.jsonl")

target = AnthropicProvider(default_model="claude-haiku-4")
summary = run_benchmark(target, items, judge_provider=judge, benchmark="observability")
print(summary.pass_rate, summary.per_category)

Eval results are stored in ClickHouse next to the traces.

API

Method Endpoint
POST /v1/traces ingest a TraceEvent
POST /v1/evals ingest an EvalEvent
GET /v1/metrics/cost?hours=24 cost by provider and model
GET /v1/metrics/latency?hours=24 p50, p95, p99 latency
GET /v1/metrics/volume?hours=24 request volume and errors
GET /v1/metrics/quality?hours=24 refusal, error and JSON validity rates
GET /v1/drift?metric=cost_usd drift analysis
GET /healthz health check

Drift

argus/analytics/drift.py compares a recent window to a baseline with the Population Stability Index and a z-score on the mean. PSI below 0.10 is stable, 0.10 to 0.25 moderate, above 0.25 significant. The demo data adds latency and cost drift to OpenAI so the dashboard has something to show.

Pricing

config/pricing.yaml has placeholder prices for local development and may not match current provider pricing. Models are matched exactly, then by longest prefix. Unknown models get zero cost and a warning.

Layout

argus/        analytics, collector, eval, ingest, providers,
              config, emitter, events, pricing, quality, tracer
scripts/      init_clickhouse.py, seed_demo.py
grafana/      dashboards and provisioning
deploy/       Cloud Run script and service config

Development

pip install -e ".[dev]"   # core + pytest, ruff
pytest
pip install -e ".[all]"   # every provider and service dependency

Deployment

CI runs ruff and the tests, builds the image, and deploys the collector to Cloud Run on pushes to main. To deploy by hand:

export GCP_PROJECT_ID=...
export GCP_REGION=us-central1
export AR_REPO=argus
./deploy/deploy.sh

Kafka and ClickHouse run separately; point the collector at them with ARGUS_KAFKA_BOOTSTRAP and ARGUS_CH_HOST.

License

MIT

About

Kafka-to-ClickHouse telemetry pipeline for LLM request latency, cost, and evaluation events.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages