Skip to content

Latest commit

 

History

1,009 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

VERA-MH 2.0

CI

VERA-MH (Validation of Ethical and Responsible AI in Mental Health) is a framework for evaluating how AI systems respond in mental health conversations that raise safety concerns. The result is a standardized score researchers, developers, and clinicians can use to compare systems before, during, and after deployment.

VERA-MH 2.0 is a multi-turn, multi-target evaluator.

  • Multi-turn—an LLM role-plays a clinically developed persona in a full conversation with the chatbot under test (30 turns by default), so the evaluation covers how the chatbot handles risk as it unfolds, not a single prompt and reply. An LLM judge then rates the whole transcript against a clinical rubric.
  • Multi-target—each target is a self-contained bundle of personas, prompts, and a clinical rubric, selected with --target. Two clinical rubrics ship by default: SI (suicidal ideation) and HFO (harm from others).
  • Bring your own rubric—anyone can add a target by dropping a folder with a manifest.json into data/; vera picks it up by name with no code changes. See docs/targets.md for the bundle layout and docs/rubric.md for the rubric format and the current compatibility requirements (for now, scoring expects the same five dimensions the shipped rubrics use).

2.0 also adds:

  • The vera CLI—a single entry point (vera.py) with generate, judge, score, and pipeline commands, replacing the separate scripts.
  • Config files—any run can be described by a JSON config and replayed exactly; each command prints its fully resolved config when it starts.

2.0 introduces breaking changes; see the CHANGELOG. The project is a continuous work in progress, and feedback is welcome under the Code of Conduct.

Papers: Reliability and Validity of VERA-MH · Concept paper · Announcement

Quick start

pip install uv                 # if you don't have it
uv sync
cp .env.example .env           # add ANTHROPIC_API_KEY, OPENAI_API_KEY, ... as needed

Run a smoke test (generate, judge, and score on 3 personas):

uv run python vera.py pipeline \
  -c gpt-4o \
  -u claude-sonnet-4-5-20250929:1 \
  -j gpt-5.4:1 \
  --judge-params reasoning_effort=low \
  --target SI \
  --sample 3
  • -c is the chatbot under test.
  • -u is the user model that role-plays the personas (model:repeats).
  • -j is the judge (model:instances). GPT 5.4 with reasoning_effort=low is the recommended judge.

Drop --sample to run every persona. Results land under output/: transcripts in conversations/, judge ratings in evaluations/j_*/results.csv, and scores in evaluations/j_*/scores/.

Each stage can also be run on its own (vera generate, vera judge, vera score), and --help on any command lists its flags. The full reference, including config files and resuming interrupted runs, is in docs/cli.md.

To evaluate your own chatbot or API rather than a built-in model, see docs/evaluating.md.

Getting a reliable VERA-MH score

For a score comparable to the published VERA-MH numbers, use the recommended profile checked in at configs/recommended-SI.json: all 100 SI personas, 30 turns, two user models (GPT 5.2 and Claude Opus 4.5), and GPT 5.4 as judge. Fill in the model under test and run it:

jq '.generation.chatbot.name = "<model-under-test>"' configs/recommended-SI.json \
  | uv run python vera.py pipeline --config -

This produces one evaluation per user model. The headline score is the pooled result across both:

uv run python scripts/pool_vera_scores.py <evaluation-folder-A> <evaluation-folder-B>

vera pipeline prints both evaluation folders when it finishes. How the score is computed and what the output files contain is covered in docs/scoring.md.

Documentation

Doc What's in it
docs/cli.md Full vera CLI reference: every command and flag, config files, --into, model parameters, output layout
docs/targets.md Targets, personas, and prompts: what ships, how they're structured, how to customize them
docs/evaluating.md Connecting your own LLM, agent, or API as the chatbot under test
docs/judge.md, docs/rubric.md How the rubric-driven judge works; adding a compatible rubric
docs/scoring.md The VERA-MH score formula, score outputs, pooling, cross-model comparison, improvement reports
docs/architecture.md Target architecture and design invariants
docs/legacy-scripts.md The pre-2.0 scripts (deprecated)

Contributing

Development conventions, the architecture map, testing policy, and git workflow are in AGENTS.md. Claude Code users get slash commands (/test, /format, /create-pr, ...) described in CLAUDE.md.

uv run pytest -m "not live"    # CI-safe test suite, no API keys needed
pre-commit install             # optional: format and lint on commit

Legacy scripts

The pre-2.0 entry points—generate.py, judge.py, run_pipeline.py, judge/score.py, and scripts/run_recommended_vera_pipeline.sh—still work, but they are deprecated and will be removed soon. Their flags differ from the vera CLI (for example -c means --max-concurrent in generate.py). Use vera for new work; if you still depend on a script, see docs/legacy-scripts.md for its options and the vera equivalent.

License

MIT with conditions: "software" is replaced by "materials" to better describe the project. See LICENSE.

About

VERA-MH official repository

Resources

Security policy

Stars

58 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages