VERA-MH (Validation of Ethical and Responsible AI in Mental Health) is a framework for evaluating how AI systems respond in mental health conversations that raise safety concerns. The result is a standardized score researchers, developers, and clinicians can use to compare systems before, during, and after deployment.
VERA-MH 2.0 is a multi-turn, multi-target evaluator.
- Multi-turn—an LLM role-plays a clinically developed persona in a full conversation with the chatbot under test (30 turns by default), so the evaluation covers how the chatbot handles risk as it unfolds, not a single prompt and reply. An LLM judge then rates the whole transcript against a clinical rubric.
- Multi-target—each target is a self-contained bundle of personas, prompts, and a clinical rubric, selected with
--target. Two clinical rubrics ship by default:SI(suicidal ideation) andHFO(harm from others). - Bring your own rubric—anyone can add a target by dropping a folder with a
manifest.jsonintodata/;verapicks it up by name with no code changes. See docs/targets.md for the bundle layout and docs/rubric.md for the rubric format and the current compatibility requirements (for now, scoring expects the same five dimensions the shipped rubrics use).
2.0 also adds:
- The
veraCLI—a single entry point (vera.py) withgenerate,judge,score, andpipelinecommands, replacing the separate scripts. - Config files—any run can be described by a JSON config and replayed exactly; each command prints its fully resolved config when it starts.
2.0 introduces breaking changes; see the CHANGELOG. The project is a continuous work in progress, and feedback is welcome under the Code of Conduct.
Papers: Reliability and Validity of VERA-MH · Concept paper · Announcement
pip install uv # if you don't have it
uv sync
cp .env.example .env # add ANTHROPIC_API_KEY, OPENAI_API_KEY, ... as neededRun a smoke test (generate, judge, and score on 3 personas):
uv run python vera.py pipeline \
-c gpt-4o \
-u claude-sonnet-4-5-20250929:1 \
-j gpt-5.4:1 \
--judge-params reasoning_effort=low \
--target SI \
--sample 3-cis the chatbot under test.-uis the user model that role-plays the personas (model:repeats).-jis the judge (model:instances). GPT 5.4 withreasoning_effort=lowis the recommended judge.
Drop --sample to run every persona. Results land under output/: transcripts in conversations/, judge ratings in evaluations/j_*/results.csv, and scores in evaluations/j_*/scores/.
Each stage can also be run on its own (vera generate, vera judge, vera score), and --help on any command lists its flags. The full reference, including config files and resuming interrupted runs, is in docs/cli.md.
To evaluate your own chatbot or API rather than a built-in model, see docs/evaluating.md.
For a score comparable to the published VERA-MH numbers, use the recommended profile checked in at configs/recommended-SI.json: all 100 SI personas, 30 turns, two user models (GPT 5.2 and Claude Opus 4.5), and GPT 5.4 as judge. Fill in the model under test and run it:
jq '.generation.chatbot.name = "<model-under-test>"' configs/recommended-SI.json \
| uv run python vera.py pipeline --config -This produces one evaluation per user model. The headline score is the pooled result across both:
uv run python scripts/pool_vera_scores.py <evaluation-folder-A> <evaluation-folder-B>vera pipeline prints both evaluation folders when it finishes. How the score is computed and what the output files contain is covered in docs/scoring.md.
| Doc | What's in it |
|---|---|
| docs/cli.md | Full vera CLI reference: every command and flag, config files, --into, model parameters, output layout |
| docs/targets.md | Targets, personas, and prompts: what ships, how they're structured, how to customize them |
| docs/evaluating.md | Connecting your own LLM, agent, or API as the chatbot under test |
| docs/judge.md, docs/rubric.md | How the rubric-driven judge works; adding a compatible rubric |
| docs/scoring.md | The VERA-MH score formula, score outputs, pooling, cross-model comparison, improvement reports |
| docs/architecture.md | Target architecture and design invariants |
| docs/legacy-scripts.md | The pre-2.0 scripts (deprecated) |
Development conventions, the architecture map, testing policy, and git workflow are in AGENTS.md. Claude Code users get slash commands (/test, /format, /create-pr, ...) described in CLAUDE.md.
uv run pytest -m "not live" # CI-safe test suite, no API keys needed
pre-commit install # optional: format and lint on commitThe pre-2.0 entry points—generate.py, judge.py, run_pipeline.py, judge/score.py, and scripts/run_recommended_vera_pipeline.sh—still work, but they are deprecated and will be removed soon. Their flags differ from the vera CLI (for example -c means --max-concurrent in generate.py). Use vera for new work; if you still depend on a script, see docs/legacy-scripts.md for its options and the vera equivalent.
MIT with conditions: "software" is replaced by "materials" to better describe the project. See LICENSE.