Group-safe research prototype for prioritising potentially inconsistent nucleus annotations for expert review.
-
Updated
Sep 1, 2026 - Python
Group-safe research prototype for prioritising potentially inconsistent nucleus annotations for expert review.
Emotion architecture from Reddit comments: rater behavior, semantic clusters, and contradiction mapping in GoEmotions.
Statistical validity checks for human-graded AI evaluations
Code, logs and evaluation artifacts for "Shortcut Learning in a Public Grape Disease Dataset: Annotation Granularity as a Modulator, Not a Cause" — inconsistent annotation granularity induces a quantifiable shortcut, established by counterfactual retraining and a pre-registered negative result.
Modular visual evaluation and computer-vision framework for image QA, structured annotation, reliability analysis, and human pose data.
Production IAA engine — Krippendorff's Alpha, SBERT semantic agreement, Fleiss' Kappa, Shannon entropy, 3-layer collusion detection. Ran behind RawEval's 9-annotator workbench.
Reliability, rogue-rater, drift and leakage diagnostics for labelled data — on the raw GoEmotions ratings, 27 of 28 emotions fall below the 0.667 agreement floor.
To associate your repository with the annotation-quality topic, visit your repo's landing page and select "manage topics."