Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 22 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,28 @@ linkml-validate -s schema.yaml data.yaml \

---

## Agent skill

Install the [linkml-reference-validator workflow skill](skills/linkml-reference-validator/SKILL.md) with the
[skills CLI](https://skills.sh/) (requires Node.js). Preview available skills first:

```bash
npx skills add linkml/linkml-reference-validator --list
npx skills add linkml/linkml-reference-validator --skill linkml-reference-validator
```

Installation defaults to the current project. Use `-a codex` or
`-a claude-code` to select an agent; add `-g` for your user-wide skills directory:

```bash
npx skills add linkml/linkml-reference-validator --skill linkml-reference-validator -a codex -g
```

The skill helps agents configure reference QC, interpret failures and coverage,
and set up repository hooks and CI using bundled reference guides. Install the
runtime separately as described above; installing the skill does not activate
hooks or configure CI, credentials, or data sources.

## Why Use This Tool?

Scientific data often includes claims supported by quotes from publications. But how do you know the quotes are accurate?
Expand Down
67 changes: 67 additions & 0 deletions skills/linkml-reference-validator/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,67 @@
---
name: linkml-reference-validator
description: Set up, configure, and troubleshoot LinkML Reference Validator (LRV) as deterministic reference QC in repository hooks and CI. Use when interpreting quote or title failures, diagnosing unchecked evidence and cache coverage, or configuring schema extraction, sources, and matching policy.
---

# Work with reference QC

LRV performs deterministic quote and title checks against retrieved source
content. Run those checks through repository automation. The agent's role is to
integrate the checker, explain findings, and make evidence-based corrections.
Maintainers define required coverage and exception policy; curators decide
whether the evidence supports the scientific claim. A matching quote settles
neither its relevance nor the truth of the claim.

A passing automated check needs no agent reenactment. Apply scientific review
when curating or reviewing evidence, rather than rereading every source on each
QC run.

## Establish the validation context

Read the project's guidance, validation recipe or wrapper, schema, explicit
config, dependency lock, and existing hook/CI output. Identify the affected
files, target class, source cache, and whether this check is blocking or
advisory. Use the existing project command: a wrapper may encode policy absent
from a bare CLI invocation. Reproduce a finding on the affected file with the
same inputs before changing data; avoid repeated whole-corpus retrieval.

- For a new gate or hook/CI changes, read [Guardrail setup](references/guardrails.md).
- For extraction, retrieval, cache, or matching changes, read
[Configuration and targeted diagnosis](references/configuration.md).

Installing this skill supplies instructions; it does not activate a hook,
install LRV, or establish a required CI check.

## Interpret the result before repairing it

| Finding | Agent's next step |
| --- | --- |
| Quote matched | Report source matching as passed. During evidence review, assess the attached claim, population, and direction of effect separately. |
| Quote did not match | Compare the exact input with the retrieved content. Distinguish paraphrase, wrong citation, normalization, and missing full text. |
| Title mismatch | Check identifier and source metadata together; changing the title to fit the wrong paper hides the real error. |
| Source unavailable or prefix skipped | Report evidence as unchecked. Diagnose retrieval or source configuration; do not label the quote fabricated. |
| Zero comparisons or unexpectedly low coverage | Inspect schema annotations, target class, missing evidence, file selection, and skip policy. Exit 0 does not establish coverage. |
| Crash or invocation error | Repair the execution/configuration problem before drawing conclusions about the evidence. |

An abstract-only record cannot establish that a quotation is absent from the
full paper. Keep matched, failed, skipped, and unavailable evidence distinct;
read coverage diagnostics alongside the exit status. Advisory output from a
hook is not completion of the repository's required checks.

## Correct the cause and close the loop

When fixing evidence, inspect the source and the intended claim together.
Transcribe a supported quote accurately; never invent wording or swap citations
solely to obtain a pass. Preserve the source cache as evidence rather than
editing it to match the submitted quote. Review any automated repair suggestion
against the paper and schema before applying it.

For configuration work, explain which records become checked or unchecked and
demonstrate the intended behavior with representative passing and failing
examples. Preserve established policy during routine repairs: adding a skipped
prefix or lowering retrieval severity changes the QC contract.

Rerun the affected check after corrections and the required repository checks
before delivery. Report the file/field, reference ID, diagnosis, correction,
command and outcome, source coverage, and any evidence still unchecked. Surface
unresolved scientific interpretation or policy choices to the human curator.
99 changes: 99 additions & 0 deletions skills/linkml-reference-validator/references/configuration.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,99 @@
# Configure and diagnose reference checks

## Confirm the effective inputs

Use the project's wrapper and locked runtime. The commands below illustrate
targeted investigation with the nested CLI; substitute real paths and IDs.
LRV's `--config` selects a config file; `-c` means **cache directory**, unlike
LTV. An explicit config avoids depending on the invocation directory's
`.linkml-reference-validator.yaml` or `.yml` autodiscovery.

The config accepts a `validation` envelope, a legacy `reference_validation`
envelope, or flat validation fields. Prefer one unambiguous representation:

```yaml
validation:
cache_dir: references_cache
unknown_prefix_severity: ERROR
skip_prefixes: []
```

Inspect CLI overrides and project wrappers as well as the file. Check a known
failure to confirm configuration is being consumed; a misplaced envelope can
leave defaults in effect. Keep custom source definitions reproducible in the
project rather than relying on a developer's home-directory configuration.

## Map the schema to the evidence

Plausible field names alone do not establish schema-aware coverage. Excerpt
slots need `implements` or `slot_uri` identifying `oa:exact`; reference slots
need `dcterms:references`. Legacy `linkml:excerpt` and
`linkml:authoritative_reference` are also supported. Check the schema's actual
nesting and target class, including nested reference objects. Confirm a known
evidence item contributes a comparison and a bad quote fails.

For plain text, `validate text-file` uses regex extraction: inspect `--help`
for text/reference capture groups and test the extracted pairs. A regex that
matches nothing does not validate the document.

## Diagnose retrieval before matching

```bash
uv run --locked linkml-reference-validator lookup PMID:16888623 \
--format json --cache-dir references_cache --config conf/reference-validator.yaml
```

Lookup can return success when only some supplied identifiers resolve. Inspect
each record and its `content_type` (for example, `abstract_only` or full text),
not just the status. Ordinary validation does not read the private research
library cache. Access to a paper during research does not establish that the
shared validation cache contains that paper's body.

Unknown-prefix/fetch failures, explicitly skipped prefixes, and quote
mismatches need different remedies. `skip_prefixes` is case-insensitive and
bypasses checking those references. `unknown_prefix_severity` changes how
retrieval failures are reported; it does not supply missing evidence. When
enabling a new source, populate its cache and assess newly exposed title and
snippet failures before changing the gate.

`file:` references can support local sources; use absolute paths or configure
`reference_base_dir` so a temporary hook file or changed working directory does
not resolve a different source.

## Diagnose matching and configuration policy

For one quote, arguments are **text first, reference second**:

```bash
uv run --locked linkml-reference-validator validate text \
"MUC1 oncoprotein blocks nuclear targeting" PMID:16888623 \
--cache-dir references_cache --config conf/reference-validator.yaml
```

Add `--title` with the supplied title to check metadata too. Quote fragments
separated by ellipses and bracketed editorial material have normalization
rules; review the source before treating every substring failure as paraphrase.

`literal_bracket_patterns` preserves bracket contents matching configured
regexes. This matters when a source itself contains bracketed abbreviations or
statistics. Choose patterns using actual source/quote pairs and test both
literal text and editorial glosses. A broad pattern changes which editorial
material is treated as a verbatim quotation. `min_excerpt_length` can require
more substantive excerpts, but length alone does not establish evidential
strength. Explain the corpus impact of either policy change.

## Use repair as a reviewed suggestion

When a repair is appropriate, preview with `repair text "QUOTE" REFERENCE_ID`
or `repair data ... --dry-run`. Data repair uses a simpler extractor than
schema-aware validation and expects scalar reference IDs in common evidence
fields; it does not support every nested layout. For
`reference: {id: PMID:...}`, use the extracted ID with text repair rather than
rewriting the dataset to suit the repair command.

Review suggestions for fidelity to the source and the intended claim. When
applying a data repair, `--no-dry-run --output repaired.yaml` preserves the input;
otherwise the CLI overwrites it with a backup. Revalidate the result with the
same schema, config, and source content. Consult the
[CLI reference](https://linkml.io/linkml-reference-validator/reference/cli/)
for source-specific options supported by the installed release.
100 changes: 100 additions & 0 deletions skills/linkml-reference-validator/references/guardrails.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,100 @@
# Put LRV in repository guardrails

Read this when adding or changing automated validation. Routine evidence
correction should use the project's existing commands.

## Establish one repository contract

Inspect existing task runners, dependency locks, schema annotations, wrappers,
hooks, and CI before adding another entry point. Record the selected files,
target class, explicit config, cache location, and which findings block each
stage. Keep these decisions in version control so humans and agents run the
same checks.

If LRV is absent from a uv project, add it with `uv add linkml-reference-validator`
and commit the dependency and lockfile changes as part of setup. Skill installation
alone does not supply the runtime. Check the locked release's subcommand help when
integrating or upgrading: repository wrappers can target an older release than
upstream documentation.

For a project with `schema.yaml`, an evidence-bearing class `Statement`, and
`conf/reference-validator.yaml`, a validation command is:

```bash
uv run --locked linkml-reference-validator validate data data.yaml \
--schema schema.yaml --target-class Statement \
--config conf/reference-validator.yaml --cache-dir references_cache
```

Adapt those paths/class to the actual project and wrap the command in its task
runner. Use the same wrapper from CI and hooks. Include structural LinkML
validation separately: source matching is not schema validation. Check that
file selection and schema extraction actually produce comparisons.

## Separate retrieval from the frequent check

Matching is deterministic for fixed inputs, config, tool version, and source
content. Live retrieval and changing caches affect what can be compared.
Define how references enter the cache, how source provenance is retained, and
how unavailable content is reported. Prepare newly cited references before
expecting cached checks to cover them.

Keep expensive enrichment and cache normalization out of every edit. Where
supported, `--no-full-text` disables full-text fetching; it is **not an offline
switch** and new references may still be fetched and cached. Test both a cache
hit and a cache miss before describing a hook as offline or non-mutating.
Do not silently treat unavailable content as verified evidence.

## Add a hook appropriate to the editing stage

Use a command hook that invokes the deterministic checker. Configure the event
using the installed Claude Code version's
[hook protocol](https://code.claude.com/docs/en/hooks), preserving existing hooks.

- A `PreToolUse` edit hook must validate the **proposed content** in a temporary
file, faithfully implementing the tool's edit semantics, including repeated
replacements. Checking the old file cannot reject a bad proposed edit.
- Resolve paths against the checkout containing the edited file. A hook's
script or launch directory can belong to the main checkout while the agent
edits a worktree. Preserve the real schema/config/cache context for temporary
files, including relative reference paths.
- Map a blocking validation failure to hook exit 2 with an actionable stderr
diagnostic. Do not merely forward CLI exit 1 and assume it blocks. Bound
subprocess runtime so the hook can report failure before its own timeout.
- A `PostToolUse` hook can report on the saved result; it cannot undo the write.
A tool-specific hook also misses edits by other tools or humans. CI provides
the common check for all changes.

Decide explicitly whether incomplete retrieval should interrupt editing or be
reported for resolution before merge. Preserve the project's chosen distinction
between advisory feedback and required validation.

## Wire CI and verify the integration

Install locked dependencies, prepare the intended cache, and call the repository
recipe. Cover changes to data, schema, validation config, wrappers, dependencies,
and cache policy. Schema/config changes can affect the whole corpus even when
no data file changed. Handle deleted files and an intentionally empty selection
explicitly; accidental selection of zero files must not look like successful QC.

Exercise a matching quote, a mismatch, a wrong title, a missing source, and an
input with no extractable evidence. Verify the actual wrapper/hook/CI outcomes,
not just the standalone CLI. Confirm that failures remain failures through
shell pipelines and that advisory results stay visible. Inspect cache changes
and runtime on representative files. Retain logs with comparison coverage and
the source/tool/config versions needed to reproduce findings.

## Example: dismech's two stages

At [dismech's inspected revision](https://github.com/monarch-initiative/dismech/tree/6bd2810f2f896fd9aa05f8223f7f50caebe457b3),
the [pre-edit hook](https://github.com/monarch-initiative/dismech/blob/6bd2810f2f896fd9aa05f8223f7f50caebe457b3/.claude/hooks/validate_disorder_hook.py)
finds the edited file's worktree and calls `just validate-pre-edit` on candidate
content. The [recipes](https://github.com/monarch-initiative/dismech/blob/6bd2810f2f896fd9aa05f8223f7f50caebe457b3/project.justfile)
block schema and term failures at this stage, but report reference findings as
advisory with `--no-full-text`. Broader validation runs before merge.

The [reference wrapper](https://github.com/monarch-initiative/dismech/blob/6bd2810f2f896fd9aa05f8223f7f50caebe457b3/scripts/run_reference_validator.sh)
implements project-specific warning handling and appends a separate, advisory
snippet audit. Recipes explicitly select `conf/reference_validator_config.yaml`;
reading only the root dotfile would describe the wrong effective policy.
These are design examples, not a universal configuration to copy wholesale.
Loading