Evaluating Weakly Supervised Object Localization Methods Right (CVPR 2020)
-
Updated
Sep 20, 2022 - Python
Evaluating Weakly Supervised Object Localization Methods Right (CVPR 2020)
Leakage-audit and prevention framework for Alzheimer's disease MRI deep learning benchmarks. Validated across Kaggle, OASIS-1, and ADNI.
Capability Schema Spec defines a shared semantic language for world model evaluation. Standardize capability definition, observation, and verification across models and benchmarks. Not a benchmark—a shared language. Define • Observe • Verify
面向基础教育作业反馈的人机协同批改框架:概念设计、风险治理与可复现评测协议。 A research-oriented prototype and evaluation protocol for human-in-the-loop assignment grading in K-12 education.
A Unified Safety-Centric Benchmark for Quadruped Robot Locomotion and Fall Recovery
Reference implementation of the Capability Schema Specification. Proves that world model capabilities can be defined, observed, and verified in practice — with real checkpoints, real simulators, and real scores. Define • Observe • Verify • Deliver
The Recall Ceiling of LLM Recommendation Reranking (CIKM 2026) — code, processed datasets, and every result JSON cited in the paper
A standardized streaming evaluation protocol and open-source harness for continuous sign language recognition.
Barkley Labs | The systems lab for individual intelligence. Evidence and intelligence protocols for decisions averages cannot make: Barkley AI (behavior), IREP Protocol (human evaluation), ACTA Music (creative provenance). barkleylabs.ai
Add a description, image, and links to the evaluation-protocol topic page so that developers can more easily learn about it.
To associate your repository with the evaluation-protocol topic, visit your repo's landing page and select "manage topics."