Skip to content

On Linux, linkml-reference-validator 0.1.7 finds none of the 107 committed doi_ caches: #690 renamed them to a casing it cannot read #1091

Description

@realmarcin

Context

Found while reviewing #1089. That PR only adds a .gitignore rule (gh pr view 1089: .gitignore +5 -0) and neither causes nor fixes this. Among the untracked leftovers it accounts for is references_cache/doi_10.1007_s11270-011-0818-5.md, whose frontmatter reads reference_id: DOI:10.1007/s11270-011-0818-5 (the same frontmatter #697 quoted). The validator wrote that id in upper case, yet the file carries a lower-case doi_ name. Following that mismatch leads to the problem below.

Summary

In linkml-reference-validator 0.1.7 (the version in uv.lock:2799-2800), fetch() normalises doi: to DOI: before it reads the disk cache. So the file it looks for is DOI_<slug>.md, not doi_<slug>.md:

  • etl/reference_fetcher.py:86: normalized_reference_id = self.normalize_reference_id(reference_id)
  • :94: self._load_from_disk(normalized_reference_id), then :358 get_cache_path(...), then :223-225 DOI_…md
  • :360-365: .exists() check, with a .txt fallback of the same casing

#690 quoted get_cache_path (:204-225) without the normalisation at :86 that feeds it. It concluded the validator looks for doi_, and PR #692 renamed the 133 DOI_* files to doi_*. The opposite was true. Of the doi: references cited at the pre-#692 tree, the validator's fetcher found 118 of 118 cached ones by exact name. On origin/main (8505a56) it finds 0 of 107.

PR #129 (merged 2026-06-10) had already stated the right convention: "The validator uppercases the prefix (doi: → DOI:) before get_cache_path, so the committed DOI_*.md abstract caches are already validator-visible." #129's case-sensitive APFS run covered only PMID:10049867, so its DOI sentence came from reading the source. The reproduction below confirms it.

#697 recorded the normalisation correctly ("normalises a reference id to DOI: and builds its cache path from that"). It then concluded that a DOI_* file is "invisible" on Linux and that the loop "never converges". For the validator, DOI_ is the only name it can see, and the next run reads its own DOI_ file. The normaliser added for #697 (PR #702) makes things worse on Linux: renaming that file to doi_ turns the next run's cache hit into a miss.

macOS hides all of this, because default APFS is case-insensitive.

Reproduction (no network)

All commands were run from a git archive origin/main copy, with the worktree venv for #1089 (linkml-reference-validator 0.1.7) and PYTHONPATH=<copy>/src. communitymech.paths.REFERENCES_CACHE was confirmed to resolve to the copy.

1. What the fetcher reads, and why reference_prefix_map cannot change it

from linkml_reference_validator.models import ReferenceValidationConfig
from linkml_reference_validator.etl.reference_fetcher import ReferenceFetcher
ref = "doi:10.1007/s11270-011-0818-5"
for pm in ({}, {"DOI": "doi"}, {"doi": "doi"}):
    f = ReferenceFetcher(ReferenceValidationConfig(cache_dir="references_cache", reference_prefix_map=pm))
    n = f.normalize_reference_id(ref)
    print(pm, n, f.get_cache_path(ref).name, f.get_cache_path(n).name, f._normalized_prefix_map())
{} DOI:10.1007/s11270-011-0818-5 doi_10.1007_s11270-011-0818-5.md DOI_10.1007_s11270-011-0818-5.md {}
{'DOI': 'doi'} DOI:10.1007/s11270-011-0818-5 doi_10.1007_s11270-011-0818-5.md DOI_10.1007_s11270-011-0818-5.md {'DOI': 'DOI'}
{'doi': 'doi'} DOI:10.1007/s11270-011-0818-5 doi_10.1007_s11270-011-0818-5.md DOI_10.1007_s11270-011-0818-5.md {'DOI': 'DOI'}

get_cache_path(raw) returns doi_, but fetch() never passes it the raw id; it passes the normalised one, which gives DOI_. Prefix-map values are upper-cased as well (_normalized_prefix_map, :197-202, via _normalize_prefix, :186-190), so no config value produces doi_. conf/reference_validator.yaml sets no map today. The plugin passes the record's reference string through unchanged (plugins/reference_validation_plugin.py:461-464).

2. Corpus count on origin/main

I used the regex from tests/test_reference_cache_names_are_case_exact.py:53 over kb/communities, data/isolates and kb/taxa, and exact-name membership in os.listdir(references_cache):

cache prefixes: PMID 616, doi 224, DOI 0
distinct doi: refs: 107 (1031 `reference: doi:` occurrences), cited by 71 records
  with an exact doi_ cache (.md or .txt):     107   (99 .md, 8 .txt-only)
  whose DOI_ name (what fetch() reads) exists:  0

The same count at the parent of PR #692's squash-merge commit (80b882cb^1, which had DOI_=133 and doi_=79) gives 128 distinct doi: refs, 118 cached, and 118 found by the fetcher's exact name. Of those 118, 117 resolved only by case, which matches #690's 117. The 118th, doi:10.1039/C3EE42189A, had both DOI_…md and doi_…txt committed.

3. End to end on a case-sensitive filesystem

I made a case-sensitive APFS image to stand in for Linux:

hdiutil create -size 20m -fs 'Case-sensitive APFS' -volname cs cs.dmg
hdiutil attach -nobrowse -mountpoint "$MNT" cs.dmg
mkdir -p "$MNT"/{conf,kb/communities,references_cache,src/communitymech}
cp conf/reference_validator.yaml "$MNT/conf/"; cp -R src/communitymech/schema "$MNT/src/communitymech/"
cp kb/communities/SynCom_Pseudomonas_Rahnella_Artemisia_Phytoremediation.yaml "$MNT/kb/communities/"
cp references_cache/doi_10.1016_j.indcrop.2025.122518.md "$MNT/references_cache/"

repro_doi_case.py runs every EvidenceItem in that record through ReferenceValidationPlugin._validate_excerpt, with the config loaded from conf/reference_validator.yaml. It blocks sockets and replaces ReferenceSourceRegistry with a stub, so a cache miss is counted but cannot reach Crossref:

"""usage: repro_doi_case.py ROOT miss|write  -- no network possible."""
import os, socket, sys
from pathlib import Path
import yaml

def _blocked(*a, **k): raise RuntimeError("network blocked")
socket.socket.connect = socket.create_connection = socket.getaddrinfo = _blocked

root, mode = Path(sys.argv[1]), sys.argv[2]
os.chdir(root)
from linkml_reference_validator.cli.shared import load_validation_config
from linkml_reference_validator.etl import reference_fetcher as rf
from linkml_reference_validator.models import ReferenceContent
from linkml_reference_validator.plugins.reference_validation_plugin import ReferenceValidationPlugin
from linkml_runtime.utils.schemaview import SchemaView

misses = []
class Stub:  # stands in for every real source (Crossref, PubMed, ...)
    def fetch(self, identifier, config):
        return ReferenceContent(reference_id=f"DOI:{identifier}", title="STUB",
                                content="stub", content_type="stub") if mode == "write" else None
def get_source(ref):
    misses.append(ref)
    return Stub
rf.ReferenceSourceRegistry = type("Registry", (), {"get_source": staticmethod(get_source)})

plugin = ReferenceValidationPlugin(config=load_validation_config(Path("conf/reference_validator.yaml")))
plugin.schema_view = SchemaView("src/communitymech/schema/communitymech.yaml")
doc = yaml.safe_load(Path("kb/communities/SynCom_Pseudomonas_Rahnella_Artemisia_Phytoremediation.yaml").read_text())
def walk(o):
    if isinstance(o, dict):
        if "reference" in o and "snippet" in o: yield o
        for v in o.values(): yield from walk(v)
    elif isinstance(o, list):
        for v in o: yield from walk(v)
items = list(walk(doc))
res = [r for i in items for r in plugin._validate_excerpt(i["snippet"], i["reference"], None, "")]
print(f"items={len(items)} failures={len(res)} misses_reaching_source={len(misses)} "
      f"messages={sorted({str(r.message)[:60] for r in res})}")
print("cache:", sorted(os.listdir("references_cache")))

Arms 6 and 7 drive the real CLI in-process (linkml_reference_validator.cli.main, validate data <record> -s src/communitymech/schema/communitymech.yaml --config conf/reference_validator.yaml). Arm 6 uses the same stubbed registry. Arm 7 keeps the real DOISource but patches its requests.get to raise requests.exceptions.ConnectionError, which is what an offline host raises. Sockets are blocked as well.

Observed:

Arm Setup Result
1 checkout on default (case-insensitive) APFS items=5 failures=0 misses_reaching_source=0
2 case-sensitive volume, committed doi_…md items=5 failures=5 misses_reaching_source=5, Could not fetch reference: doi:10.1016/j.indcrop.2025.122518
3 control: same volume, same bytes renamed DOI_…md items=5 failures=0 misses_reaching_source=0
4 arm 2, but the stub returns content (write), then scripts/normalize_cache_names.py --cache-dir $MNT/references_cache fetch writes DOI_…md beside doi_…md; the normaliser prints its #706 conflict banner and [conflict] DOI_10.1016_j.indcrop.2025.122518.md: doi_10.1016_j.indcrop.2025.122518.md already exists and differs, then exits 1
5 fresh DOI, empty cache: write, miss, normaliser, miss misses 1, then 0 (it reads its own DOI_ file), then [renamed] DOI_… -> doi_… (exit 0), then 5 misses again
6 arm 2 vs arm 3 files, real validate data CLI, sources stubbed to return nothing doi_: five [WARN] Could not fetch reference, exit 1. DOI_ control: All validations passed!, exit 0
7 arm 2 vs arm 3 files, real CLI and real DOISource, requests.get raises ConnectionError doi_: uncaught requests.exceptions.ConnectionError from sources/doi.py:104 (via supporting_text_validator.py:176 → reference_fetcher.py:108) aborts the run on the first item, exit 1. DOI_ control: 0 requests, exit 0

The failures in arms 4 and 5 read Text part not found because the stub's content is a placeholder; those arms test file names, not snippets. On Linux, skip hdiutil and run the script in the checkout itself.

Expected vs actual

  • Expected: a doi: evidence item whose source is committed under references_cache/ validates from the committed file on every platform, with no network access.
  • Actual: on a case-sensitive filesystem, all 107 cached doi: references (71 records) miss. What happens next depends on the network:
    • Network reachable: each miss goes to sources/doi.py:99 (https://api.crossref.org/works/{doi}), falling back to DataCite (:85). The snippet is then checked against whatever that returns, not the committed text. The fetch writes DOI_*.md beside the committed doi_*.md, and normalize_cache_names.py reports a [conflict]. The recipes deliberately discard the normaliser's exit code (justfile:163-166, :193-199, :397-400; kept that way by test: gate the cache-name conflict, and say so where it happens (#706) #727 for The cache normaliser's exit code is discarded, so a conflict cannot fail anything #706).
    • Network unreachable: 0.1.7's DOI source does not catch requests exceptions (sources/doi.py:104). There is no try in sources/doi.py, etl/reference_fetcher.py or validation/supporting_text_validator.py. So the first doi: miss raises ConnectionError and aborts validation of the whole file (arm 7). Every later evidence item in that file, PMID-cited ones included, goes unchecked.
    • Both sources answer with nothing: each item yields Could not fetch reference at unknown_prefix_severity (WARNING, conf/reference_validator.yaml:4; supporting_text_validator.py:176-184), so the snippet is never checked. The CLI still exits 1, because it exits 1 on any result, WARNING included (cli/validate.py:300-317; arm 6).
    • A DOI with no committed cache: the normaliser renames the one file the validator can read (arm 5).

Impact, stated honestly

  • Not a CI failure today. No workflow invokes validate-references, validate-references-all, qc-references, repair-references, linkml-reference-validator or the normaliser (grep -rn over .github/ finds nothing). All 10 jobs run on ubuntu-latest. The one CI-run test that drives the upstream fetcher (tests/test_reference_validator_actually_validates.py) uses PMID:28287150 only, and PMID: normalisation is a no-op. The e2e-marked test_batch_snippet_fixer_validation.py test is deselected by addopts = "-m 'not e2e'", and its record cites PMIDs only.

  • It affects anyone who runs reference validation on Linux, and any future Linux job that does. On a case-sensitive filesystem, just validate-references on a record that cites a cached doi: exits 1: it re-fetches, or aborts when offline. The same command passes on macOS.

  • The repo's own gate enforces the unreadable casing. tests/test_reference_cache_names_are_case_exact.py passes on origin/main (6 passed). Renaming one cache file to DOI_ turns two of its tests red: test_no_reference_resolves_only_by_filename_case and test_every_cache_filename_uses_the_canonical_prefix_casing. The mutation was confirmed applied, and the restore was green on the next run. Both tests take their casing from CANONICAL_CACHE_PREFIXES = {"doi": "doi", ...} in src/communitymech/paths.py:52-56, which the normaliser imports too.

  • The inverted premise is stated as fact in at least these places:

    • src/communitymech/paths.py:42-51
    • tests/test_reference_cache_names_are_case_exact.py:13-18 and :199
    • scripts/normalize_cache_names.py:3-9
    • tests/test_cache_names_are_normalised_after_fetching.py:3-9
    • scripts/cache_fulltext.py:200-211
    • tests/test_cultivation_source_studied_the_community.py:263-275
    • justfile:158-162

    NEXT_TASKS.md:1043 records the opposite, correct fact: DOI_<doi>.md is "the cache filename the reference validator already reads".

Suggested fix

  1. Do not simply rename back to DOI_. The repo's own readers build doi_ exactly: scripts/annotate_reference_errors.py:47, cache_supplements.py:207,245, cultivation_spans.py:74 and taxon_absent_from_source.py:96,115. Only two handle either spelling: cache_fulltext.py:217-221 tries both, and evidence_snippet_audit.py:190 is case-insensitive. Under 0.1.7, no single committed name serves both the validator and the scripts on a case-sensitive filesystem.
  2. Fix the read side. The upstream fix is partial, and the bump is not drop-in. Keep a line break in metadata, and one cache file per DOI linkml/linkml-reference-validator#87 (merged 2026-09-21 as eb944441; the DOI change starts in its first commit, c61044d4) makes _cache_path resolve a DOI to an existing cache file that differs only in case. It does this with _existing_case_variant, which is built from real directory entries rather than exists(), so it does not reintroduce A cache-resolution fallback that could never match made a snippet gate report different results on Linux and macOS #694's platform split. The fix is present in tags v0.3.0rc1–rc3 and absent from v0.2.1, the latest non-prerelease. Two things I ran from rc3 source on the case-sensitive volume:
    • Staleness. rc3 treats every frontmatter entry without an extractor_version: stamp as stale on the validation read path (_is_stale_cache_entry) and re-fetches it whenever a source is reachable. None of the 384 frontmatter .md files in references_cache/ carries that stamp. With today's config, a bump would send every such reference, PMID and DOI alike, back to the network on every platform. Setting trust_cached_entries: true (default false) in conf/reference_validator.yaml switches that off. Observed: the committed doi_10.1016_j.indcrop.2025.122518.md was not loaded with the default setting, but loaded with trust_cached_entries or with an extractor_version: 1 stamp.
    • .txt gap. Even with that setting, the case-variant lookup covers only the .md name. The .txt fallback is derived from the unresolved DOI_ path, so the 8 cited doi: references cached only as doi_*.txt still miss on a case-sensitive filesystem (e.g. doi:10.1186/s13568-015-0124-5, doi:10.1128/mSystems.00352-19). Control: the same bytes named DOI_*.txt load.
    • Options:
  3. Once the read side is fixed, the post-fetch normaliser is harmless. Until then it is actively harmful on Linux (arm 5). In the meantime, consider disabling its rename step or limiting it to case-insensitive filesystems.
  4. Correct the statements listed above, starting with CANONICAL_CACHE_PREFIXES, which is where the casing is actually enforced. Add a regression test that checks the name the fetcher actually reads, f.get_cache_path(f.normalize_reference_id(ref)).name plus its .txt sibling, by exact os.listdir membership. That test is red today (0/107) on both platforms. Give it a control arm (The mutation-testing protocol is undocumented, and three of its failure modes produce false greens #696/A mutation check shipped able to pass whatever the mutation did (#696's own failure mode) #709).

Not verified

  • Not run on real Linux or ext4. I used a case-sensitive APFS disk image (confirmed: a and A coexist; diskutil reports Case-sensitive APFS).
  • No network call was made. What Crossref or DataCite return for these DOIs, and whether snippets would match it, is unknown. The offline arm simulates an unreachable host by raising requests.exceptions.ConnectionError from requests.get; I did not cut a real link or test a DNS failure.
  • The CLI was driven in-process through linkml_reference_validator.cli.main with patched sources, not through uv run or just.
  • rc3 was run from its source tarball (fetched with a read-only gh api call) against the 0.1.7 venv's dependencies, calling only get_cache_path and _load_from_disk. I did not run rc3's full fetch or validate path, or install it with its own pins.
  • The 1031 occurrence count comes from the test's regex, so it may include reference: slots the plugin never validates. The distinct-reference (107) and record (71) counts are the meaningful ones.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions