tracking performance of rustc-generated binaries over time
-
Updated
Dec 2, 2022 - Rust
tracking performance of rustc-generated binaries over time
Detects when a deploy silently ships a noindex, drops a schema block, or rewrites your canonical. Baselines a page's SEO contract, diffs every change with Critical/Warning/Info severity, and points fixes to claude-seo sub-skills. A Claude Code skill. Open source.
Agentic eval framework for Laravel: golden datasets, LLM-as-judge, adversarial harness, regression detection
Fast, lightweight Rust benchmarking library. Dev mode: profile like Criterion. Prod mode: monitor with <10ns overhead. Features: async benchmarks, regression alerts, health metrics, p95/p99 latency tracking. Thread-safe, minimal deps. Benchmark any code block or function with one macro.
Statistical regression detection and CI quality gates for AI agents
Nextflow plugin for cross-run benchmarking. Records per-process metrics for every run into a local or shared history store, compares runs against their own history, and flags regressions with likely causes, over-provisioned resources, and retries. Includes a zero-dependency dashboard, an MCP server for AI assistants, and opt-in AI narration.
Detects semantic regressions across PRs that git's merge conflict detection misses
Unified benchmarking and profiling framework for the JAX scientific ML ecosystem. Timing, GPU/energy monitoring, FLOPS counting, roofline analysis, statistical testing, regression detection, and CI integration.
A new package that helps developers integration-test AI and LLM applications by validating structured outputs. It takes a user's test scenario or prompt as input, sends it to an LLM, and uses pattern
Linux/C++ CLI for benchmarking, perf-counter capture, and regression detection with configurable thresholds and statistical validation.
AI-powered benchmark regression analyzer with CLI, REST API, and interactive Dashboard. Compares cloud performance runs via OpenSearch + LLM, with automated regression detection.
LLM eval harness (aibench): 21 seeded, reproducible graded tasks with CUSUM drift detection for silent model degradation — plus a Karpathy-style council that auto-elects OpenRouter advisors.
Online evaluation, regression detection, and statistical A/B experimentation for LLM systems in production
SQL service observability demo: SLA monitoring, regression detection, root-cause analysis, and capacity forecasting.
Catch silent agent failures in CI. Pytest plugin with regression detection for LLM agents.
This program enables users to microbenchmark their Java code to detect regression in performance.
Compare two AI Procurement Pulse snapshots, flag regressions beyond threshold, render GitHub issue body + vendor-outreach email template citing the specific field diff. Productizes the quarterly Pulse regression signal. Drop in as GH Action in the Pulse engine repo.
Agent and LLM evaluation harness: golden datasets, multi-scorer execution, regression detection across model versions, cost-quality leaderboards, and CI gates for model promotion.
Proof, not guesses, for agent-caused regressions.
Framework-agnostic ledger for quality signals: append-only history, cross-run diff and ratchet gating on regressions
To associate your repository with the regression-detection topic, visit your repo's landing page and select "manage topics."