You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Found while reviewing #1089. That PR only adds a .gitignore rule (gh pr view 1089: .gitignore +5 -0) and neither causes nor fixes this. Among the untracked leftovers it accounts for is references_cache/doi_10.1007_s11270-011-0818-5.md, whose frontmatter reads reference_id: DOI:10.1007/s11270-011-0818-5 (the same frontmatter #697 quoted). The validator wrote that id in upper case, yet the file carries a lower-case doi_ name. Following that mismatch leads to the problem below.
Summary
In linkml-reference-validator 0.1.7 (the version in uv.lock:2799-2800), fetch() normalises doi: to DOI:before it reads the disk cache. So the file it looks for is DOI_<slug>.md, not doi_<slug>.md:
:94: self._load_from_disk(normalized_reference_id), then :358get_cache_path(...), then :223-225DOI_…md
:360-365: .exists() check, with a .txt fallback of the same casing
#690 quoted get_cache_path (:204-225) without the normalisation at :86 that feeds it. It concluded the validator looks for doi_, and PR #692 renamed the 133 DOI_* files to doi_*. The opposite was true. Of the doi: references cited at the pre-#692 tree, the validator's fetcher found 118 of 118 cached ones by exact name. On origin/main (8505a56) it finds 0 of 107.
PR #129 (merged 2026-06-10) had already stated the right convention: "The validator uppercases the prefix (doi: → DOI:) beforeget_cache_path, so the committed DOI_*.md abstract caches are already validator-visible." #129's case-sensitive APFS run covered only PMID:10049867, so its DOI sentence came from reading the source. The reproduction below confirms it.
#697 recorded the normalisation correctly ("normalises a reference id to DOI: and builds its cache path from that"). It then concluded that a DOI_* file is "invisible" on Linux and that the loop "never converges". For the validator, DOI_ is the only name it can see, and the next run reads its own DOI_ file. The normaliser added for #697 (PR #702) makes things worse on Linux: renaming that file to doi_ turns the next run's cache hit into a miss.
macOS hides all of this, because default APFS is case-insensitive.
Reproduction (no network)
All commands were run from a git archive origin/main copy, with the worktree venv for #1089 (linkml-reference-validator 0.1.7) and PYTHONPATH=<copy>/src. communitymech.paths.REFERENCES_CACHE was confirmed to resolve to the copy.
1. What the fetcher reads, and why reference_prefix_map cannot change it
get_cache_path(raw) returns doi_, but fetch() never passes it the raw id; it passes the normalised one, which gives DOI_. Prefix-map values are upper-cased as well (_normalized_prefix_map, :197-202, via _normalize_prefix, :186-190), so no config value produces doi_. conf/reference_validator.yaml sets no map today. The plugin passes the record's reference string through unchanged (plugins/reference_validation_plugin.py:461-464).
2. Corpus count on origin/main
I used the regex from tests/test_reference_cache_names_are_case_exact.py:53 over kb/communities, data/isolates and kb/taxa, and exact-name membership in os.listdir(references_cache):
cache prefixes: PMID 616, doi 224, DOI 0
distinct doi: refs: 107 (1031 `reference: doi:` occurrences), cited by 71 records
with an exact doi_ cache (.md or .txt): 107 (99 .md, 8 .txt-only)
whose DOI_ name (what fetch() reads) exists: 0
The same count at the parent of PR #692's squash-merge commit (80b882cb^1, which had DOI_=133 and doi_=79) gives 128 distinct doi: refs, 118 cached, and 118 found by the fetcher's exact name. Of those 118, 117 resolved only by case, which matches #690's 117. The 118th, doi:10.1039/C3EE42189A, had both DOI_…md and doi_…txt committed.
3. End to end on a case-sensitive filesystem
I made a case-sensitive APFS image to stand in for Linux:
repro_doi_case.py runs every EvidenceItem in that record through ReferenceValidationPlugin._validate_excerpt, with the config loaded from conf/reference_validator.yaml. It blocks sockets and replaces ReferenceSourceRegistry with a stub, so a cache miss is counted but cannot reach Crossref:
Arms 6 and 7 drive the real CLI in-process (linkml_reference_validator.cli.main, validate data <record> -s src/communitymech/schema/communitymech.yaml --config conf/reference_validator.yaml). Arm 6 uses the same stubbed registry. Arm 7 keeps the real DOISource but patches its requests.get to raise requests.exceptions.ConnectionError, which is what an offline host raises. Sockets are blocked as well.
Observed:
Arm
Setup
Result
1
checkout on default (case-insensitive) APFS
items=5 failures=0 misses_reaching_source=0
2
case-sensitive volume, committed doi_…md
items=5 failures=5 misses_reaching_source=5, Could not fetch reference: doi:10.1016/j.indcrop.2025.122518
3
control: same volume, same bytes renamed DOI_…md
items=5 failures=0 misses_reaching_source=0
4
arm 2, but the stub returns content (write), then scripts/normalize_cache_names.py --cache-dir $MNT/references_cache
fetch writes DOI_…mdbesidedoi_…md; the normaliser prints its #706 conflict banner and [conflict] DOI_10.1016_j.indcrop.2025.122518.md: doi_10.1016_j.indcrop.2025.122518.md already exists and differs, then exits 1
5
fresh DOI, empty cache: write, miss, normaliser, miss
misses 1, then 0 (it reads its own DOI_ file), then [renamed] DOI_… -> doi_… (exit 0), then 5 misses again
6
arm 2 vs arm 3 files, real validate data CLI, sources stubbed to return nothing
doi_: five [WARN] Could not fetch reference, exit 1. DOI_ control: All validations passed!, exit 0
7
arm 2 vs arm 3 files, real CLI and real DOISource, requests.get raises ConnectionError
doi_: uncaught requests.exceptions.ConnectionError from sources/doi.py:104 (via supporting_text_validator.py:176 → reference_fetcher.py:108) aborts the run on the first item, exit 1. DOI_ control: 0 requests, exit 0
The failures in arms 4 and 5 read Text part not found because the stub's content is a placeholder; those arms test file names, not snippets. On Linux, skip hdiutil and run the script in the checkout itself.
Expected vs actual
Expected: a doi: evidence item whose source is committed under references_cache/ validates from the committed file on every platform, with no network access.
Actual: on a case-sensitive filesystem, all 107 cached doi: references (71 records) miss. What happens next depends on the network:
Network unreachable: 0.1.7's DOI source does not catch requests exceptions (sources/doi.py:104). There is no try in sources/doi.py, etl/reference_fetcher.py or validation/supporting_text_validator.py. So the first doi: miss raises ConnectionError and aborts validation of the whole file (arm 7). Every later evidence item in that file, PMID-cited ones included, goes unchecked.
Both sources answer with nothing: each item yields Could not fetch reference at unknown_prefix_severity (WARNING, conf/reference_validator.yaml:4; supporting_text_validator.py:176-184), so the snippet is never checked. The CLI still exits 1, because it exits 1 on any result, WARNING included (cli/validate.py:300-317; arm 6).
A DOI with no committed cache: the normaliser renames the one file the validator can read (arm 5).
Impact, stated honestly
Not a CI failure today. No workflow invokes validate-references, validate-references-all, qc-references, repair-references, linkml-reference-validator or the normaliser (grep -rn over .github/ finds nothing). All 10 jobs run on ubuntu-latest. The one CI-run test that drives the upstream fetcher (tests/test_reference_validator_actually_validates.py) uses PMID:28287150 only, and PMID: normalisation is a no-op. The e2e-marked test_batch_snippet_fixer_validation.py test is deselected by addopts = "-m 'not e2e'", and its record cites PMIDs only.
It affects anyone who runs reference validation on Linux, and any future Linux job that does. On a case-sensitive filesystem, just validate-references on a record that cites a cached doi: exits 1: it re-fetches, or aborts when offline. The same command passes on macOS.
The repo's own gate enforces the unreadable casing. tests/test_reference_cache_names_are_case_exact.py passes on origin/main (6 passed). Renaming one cache file to DOI_ turns two of its tests red: test_no_reference_resolves_only_by_filename_case and test_every_cache_filename_uses_the_canonical_prefix_casing. The mutation was confirmed applied, and the restore was green on the next run. Both tests take their casing from CANONICAL_CACHE_PREFIXES = {"doi": "doi", ...} in src/communitymech/paths.py:52-56, which the normaliser imports too.
The inverted premise is stated as fact in at least these places:
src/communitymech/paths.py:42-51
tests/test_reference_cache_names_are_case_exact.py:13-18 and :199
NEXT_TASKS.md:1043 records the opposite, correct fact: DOI_<doi>.md is "the cache filename the reference validator already reads".
Suggested fix
Do not simply rename back to DOI_. The repo's own readers build doi_ exactly: scripts/annotate_reference_errors.py:47, cache_supplements.py:207,245, cultivation_spans.py:74 and taxon_absent_from_source.py:96,115. Only two handle either spelling: cache_fulltext.py:217-221 tries both, and evidence_snippet_audit.py:190 is case-insensitive. Under 0.1.7, no single committed name serves both the validator and the scripts on a case-sensitive filesystem.
Staleness. rc3 treats every frontmatter entry without an extractor_version: stamp as stale on the validation read path (_is_stale_cache_entry) and re-fetches it whenever a source is reachable. None of the 384 frontmatter .md files in references_cache/ carries that stamp. With today's config, a bump would send every such reference, PMID and DOI alike, back to the network on every platform. Setting trust_cached_entries: true (default false) in conf/reference_validator.yaml switches that off. Observed: the committed doi_10.1016_j.indcrop.2025.122518.md was not loaded with the default setting, but loaded with trust_cached_entries or with an extractor_version: 1 stamp.
.txt gap. Even with that setting, the case-variant lookup covers only the .md name. The .txt fallback is derived from the unresolved DOI_ path, so the 8 cited doi: references cached only as doi_*.txt still miss on a case-sensitive filesystem (e.g. doi:10.1186/s13568-015-0124-5, doi:10.1128/mSystems.00352-19). Control: the same bytes named DOI_*.txt load.
Options:
Bump to 0.3.x once it is stable, together with trust_cached_entries: true and a fix for the .txt case (upstream or local). pyproject.toml:17 only says >=0.1.0; the lock holds 0.1.7. Canary first: re-run arms 2 and 3 with both an .md and a .txt-only DOI.
Once the read side is fixed, the post-fetch normaliser is harmless. Until then it is actively harmful on Linux (arm 5). In the meantime, consider disabling its rename step or limiting it to case-insensitive filesystems.
Not run on real Linux or ext4. I used a case-sensitive APFS disk image (confirmed: a and A coexist; diskutil reports Case-sensitive APFS).
No network call was made. What Crossref or DataCite return for these DOIs, and whether snippets would match it, is unknown. The offline arm simulates an unreachable host by raising requests.exceptions.ConnectionError from requests.get; I did not cut a real link or test a DNS failure.
The CLI was driven in-process through linkml_reference_validator.cli.main with patched sources, not through uv run or just.
rc3 was run from its source tarball (fetched with a read-only gh api call) against the 0.1.7 venv's dependencies, calling only get_cache_path and _load_from_disk. I did not run rc3's full fetch or validate path, or install it with its own pins.
The 1031 occurrence count comes from the test's regex, so it may include reference: slots the plugin never validates. The distinct-reference (107) and record (71) counts are the meaningful ones.
Context
Found while reviewing #1089. That PR only adds a
.gitignorerule (gh pr view 1089:.gitignore +5 -0) and neither causes nor fixes this. Among the untracked leftovers it accounts for isreferences_cache/doi_10.1007_s11270-011-0818-5.md, whose frontmatter readsreference_id: DOI:10.1007/s11270-011-0818-5(the same frontmatter #697 quoted). The validator wrote that id in upper case, yet the file carries a lower-casedoi_name. Following that mismatch leads to the problem below.Summary
In
linkml-reference-validator0.1.7 (the version inuv.lock:2799-2800),fetch()normalisesdoi:toDOI:before it reads the disk cache. So the file it looks for isDOI_<slug>.md, notdoi_<slug>.md:etl/reference_fetcher.py:86:normalized_reference_id = self.normalize_reference_id(reference_id):94:self._load_from_disk(normalized_reference_id), then:358get_cache_path(...), then:223-225DOI_…md:360-365:.exists()check, with a.txtfallback of the same casing#690 quoted
get_cache_path(:204-225) without the normalisation at:86that feeds it. It concluded the validator looks fordoi_, and PR #692 renamed the 133DOI_*files todoi_*. The opposite was true. Of thedoi:references cited at the pre-#692 tree, the validator's fetcher found 118 of 118 cached ones by exact name. Onorigin/main(8505a56) it finds 0 of 107.PR #129 (merged 2026-06-10) had already stated the right convention: "The validator uppercases the prefix (
doi:→DOI:) beforeget_cache_path, so the committedDOI_*.mdabstract caches are already validator-visible." #129's case-sensitive APFS run covered onlyPMID:10049867, so its DOI sentence came from reading the source. The reproduction below confirms it.#697 recorded the normalisation correctly ("normalises a reference id to
DOI:and builds its cache path from that"). It then concluded that aDOI_*file is "invisible" on Linux and that the loop "never converges". For the validator,DOI_is the only name it can see, and the next run reads its ownDOI_file. The normaliser added for #697 (PR #702) makes things worse on Linux: renaming that file todoi_turns the next run's cache hit into a miss.macOS hides all of this, because default APFS is case-insensitive.
Reproduction (no network)
All commands were run from a
git archive origin/maincopy, with the worktree venv for #1089 (linkml-reference-validator 0.1.7) andPYTHONPATH=<copy>/src.communitymech.paths.REFERENCES_CACHEwas confirmed to resolve to the copy.1. What the fetcher reads, and why
reference_prefix_mapcannot change itget_cache_path(raw)returnsdoi_, butfetch()never passes it the raw id; it passes the normalised one, which givesDOI_. Prefix-map values are upper-cased as well (_normalized_prefix_map,:197-202, via_normalize_prefix,:186-190), so no config value producesdoi_.conf/reference_validator.yamlsets no map today. The plugin passes the record's reference string through unchanged (plugins/reference_validation_plugin.py:461-464).2. Corpus count on
origin/mainI used the regex from
tests/test_reference_cache_names_are_case_exact.py:53overkb/communities,data/isolatesandkb/taxa, and exact-name membership inos.listdir(references_cache):The same count at the parent of PR #692's squash-merge commit (
80b882cb^1, which hadDOI_=133 anddoi_=79) gives 128 distinctdoi:refs, 118 cached, and 118 found by the fetcher's exact name. Of those 118, 117 resolved only by case, which matches #690's 117. The 118th,doi:10.1039/C3EE42189A, had bothDOI_…mdanddoi_…txtcommitted.3. End to end on a case-sensitive filesystem
I made a case-sensitive APFS image to stand in for Linux:
repro_doi_case.pyruns everyEvidenceItemin that record throughReferenceValidationPlugin._validate_excerpt, with the config loaded fromconf/reference_validator.yaml. It blocks sockets and replacesReferenceSourceRegistrywith a stub, so a cache miss is counted but cannot reach Crossref:Arms 6 and 7 drive the real CLI in-process (
linkml_reference_validator.cli.main,validate data <record> -s src/communitymech/schema/communitymech.yaml --config conf/reference_validator.yaml). Arm 6 uses the same stubbed registry. Arm 7 keeps the realDOISourcebut patches itsrequests.getto raiserequests.exceptions.ConnectionError, which is what an offline host raises. Sockets are blocked as well.Observed:
items=5 failures=0 misses_reaching_source=0doi_…mditems=5 failures=5 misses_reaching_source=5,Could not fetch reference: doi:10.1016/j.indcrop.2025.122518DOI_…mditems=5 failures=0 misses_reaching_source=0write), thenscripts/normalize_cache_names.py --cache-dir $MNT/references_cacheDOI_…mdbesidedoi_…md; the normaliser prints its #706 conflict banner and[conflict] DOI_10.1016_j.indcrop.2025.122518.md: doi_10.1016_j.indcrop.2025.122518.md already exists and differs, then exits 1write,miss, normaliser,missDOI_file), then[renamed] DOI_… -> doi_…(exit 0), then 5 misses againvalidate dataCLI, sources stubbed to return nothingdoi_: five[WARN] Could not fetch reference, exit 1.DOI_control:All validations passed!, exit 0DOISource,requests.getraisesConnectionErrordoi_: uncaughtrequests.exceptions.ConnectionErrorfromsources/doi.py:104(viasupporting_text_validator.py:176→reference_fetcher.py:108) aborts the run on the first item, exit 1.DOI_control: 0 requests, exit 0The failures in arms 4 and 5 read
Text part not foundbecause the stub's content is a placeholder; those arms test file names, not snippets. On Linux, skiphdiutiland run the script in the checkout itself.Expected vs actual
doi:evidence item whose source is committed underreferences_cache/validates from the committed file on every platform, with no network access.doi:references (71 records) miss. What happens next depends on the network:sources/doi.py:99(https://api.crossref.org/works/{doi}), falling back to DataCite (:85). The snippet is then checked against whatever that returns, not the committed text. The fetch writesDOI_*.mdbeside the committeddoi_*.md, andnormalize_cache_names.pyreports a[conflict]. The recipes deliberately discard the normaliser's exit code (justfile:163-166,:193-199,:397-400; kept that way by test: gate the cache-name conflict, and say so where it happens (#706) #727 for The cache normaliser's exit code is discarded, so a conflict cannot fail anything #706).requestsexceptions (sources/doi.py:104). There is notryinsources/doi.py,etl/reference_fetcher.pyorvalidation/supporting_text_validator.py. So the firstdoi:miss raisesConnectionErrorand aborts validation of the whole file (arm 7). Every later evidence item in that file, PMID-cited ones included, goes unchecked.Could not fetch referenceatunknown_prefix_severity(WARNING,conf/reference_validator.yaml:4;supporting_text_validator.py:176-184), so the snippet is never checked. The CLI still exits 1, because it exits 1 on any result, WARNING included (cli/validate.py:300-317; arm 6).Impact, stated honestly
Not a CI failure today. No workflow invokes
validate-references,validate-references-all,qc-references,repair-references,linkml-reference-validatoror the normaliser (grep -rnover.github/finds nothing). All 10 jobs run onubuntu-latest. The one CI-run test that drives the upstream fetcher (tests/test_reference_validator_actually_validates.py) usesPMID:28287150only, andPMID:normalisation is a no-op. Thee2e-markedtest_batch_snippet_fixer_validation.pytest is deselected byaddopts = "-m 'not e2e'", and its record cites PMIDs only.It affects anyone who runs reference validation on Linux, and any future Linux job that does. On a case-sensitive filesystem,
just validate-referenceson a record that cites a cacheddoi:exits 1: it re-fetches, or aborts when offline. The same command passes on macOS.The repo's own gate enforces the unreadable casing.
tests/test_reference_cache_names_are_case_exact.pypasses onorigin/main(6 passed). Renaming one cache file toDOI_turns two of its tests red:test_no_reference_resolves_only_by_filename_caseandtest_every_cache_filename_uses_the_canonical_prefix_casing. The mutation was confirmed applied, and the restore was green on the next run. Both tests take their casing fromCANONICAL_CACHE_PREFIXES = {"doi": "doi", ...}insrc/communitymech/paths.py:52-56, which the normaliser imports too.The inverted premise is stated as fact in at least these places:
src/communitymech/paths.py:42-51tests/test_reference_cache_names_are_case_exact.py:13-18and:199scripts/normalize_cache_names.py:3-9tests/test_cache_names_are_normalised_after_fetching.py:3-9scripts/cache_fulltext.py:200-211tests/test_cultivation_source_studied_the_community.py:263-275justfile:158-162NEXT_TASKS.md:1043records the opposite, correct fact:DOI_<doi>.mdis "the cache filename the reference validator already reads".Suggested fix
DOI_. The repo's own readers builddoi_exactly:scripts/annotate_reference_errors.py:47,cache_supplements.py:207,245,cultivation_spans.py:74andtaxon_absent_from_source.py:96,115. Only two handle either spelling:cache_fulltext.py:217-221tries both, andevidence_snippet_audit.py:190is case-insensitive. Under 0.1.7, no single committed name serves both the validator and the scripts on a case-sensitive filesystem.eb944441; the DOI change starts in its first commit,c61044d4) makes_cache_pathresolve a DOI to an existing cache file that differs only in case. It does this with_existing_case_variant, which is built from real directory entries rather thanexists(), so it does not reintroduce A cache-resolution fallback that could never match made a snippet gate report different results on Linux and macOS #694's platform split. The fix is present in tagsv0.3.0rc1–rc3and absent fromv0.2.1, the latest non-prerelease. Two things I ran from rc3 source on the case-sensitive volume:extractor_version:stamp as stale on the validation read path (_is_stale_cache_entry) and re-fetches it whenever a source is reachable. None of the 384 frontmatter.mdfiles inreferences_cache/carries that stamp. With today's config, a bump would send every such reference, PMID and DOI alike, back to the network on every platform. Settingtrust_cached_entries: true(defaultfalse) inconf/reference_validator.yamlswitches that off. Observed: the committeddoi_10.1016_j.indcrop.2025.122518.mdwas not loaded with the default setting, but loaded withtrust_cached_entriesor with anextractor_version: 1stamp..txtgap. Even with that setting, the case-variant lookup covers only the.mdname. The.txtfallback is derived from the unresolvedDOI_path, so the 8 citeddoi:references cached only asdoi_*.txtstill miss on a case-sensitive filesystem (e.g.doi:10.1186/s13568-015-0124-5,doi:10.1128/mSystems.00352-19). Control: the same bytes namedDOI_*.txtload.trust_cached_entries: trueand a fix for the.txtcase (upstream or local).pyproject.toml:17only says>=0.1.0; the lock holds 0.1.7. Canary first: re-run arms 2 and 3 with both an.mdand a.txt-only DOI.doi_spelling for both.mdand.txt. Bare NCT references are cached under one path and read from another, so every run re-fetches them linkml/linkml-reference-validator#69 cites a downstream project that wrapsget_cache_paththe same way.CANONICAL_CACHE_PREFIXES, which is where the casing is actually enforced. Add a regression test that checks the name the fetcher actually reads,f.get_cache_path(f.normalize_reference_id(ref)).nameplus its.txtsibling, by exactos.listdirmembership. That test is red today (0/107) on both platforms. Give it a control arm (The mutation-testing protocol is undocumented, and three of its failure modes produce false greens #696/A mutation check shipped able to pass whatever the mutation did (#696's own failure mode) #709).Not verified
aandAcoexist;diskutilreportsCase-sensitive APFS).requests.exceptions.ConnectionErrorfromrequests.get; I did not cut a real link or test a DNS failure.linkml_reference_validator.cli.mainwith patched sources, not throughuv runorjust.gh apicall) against the 0.1.7 venv's dependencies, calling onlyget_cache_pathand_load_from_disk. I did not run rc3's full fetch or validate path, or install it with its own pins.reference:slots the plugin never validates. The distinct-reference (107) and record (71) counts are the meaningful ones.