Refreshing a cached PubMed record can silently delete the abstract it already held. For old records whose abstract PubMed stores in <OtherAbstract> rather than <Abstract>, the extractor finds nothing, writes content_type: unavailable, and overwrites the cache file with an empty ## Content section. Any snippet validated against that text then fails with "No content available for reference".
PMID:5697815 is a worked example. Its cache entry held a 967-character abstract; after a refresh the file went from 1787 to 850 bytes with the abstract gone. PubMed does serve the text — it is just in a different element:
$ curl -s "https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi?db=pubmed&id=5697815&retmode=xml" \
| grep -oE "<(Abstract|OtherAbstract)[^>]*>"
<OtherAbstract Type="PIP" Language="eng">
<AbstractText>
There is no <Abstract> element at all, so this line in _parse_abstract returns None:
https://github.com/linkml/linkml-reference-validator/blob/main/src/linkml_reference_validator/etl/sources/pmid.py#L233
abstract = soup.find("Abstract")
if not abstract:
return None
PMID:6680124 fails the same way (1695 → 650 bytes, a 1175-character PIP abstract dropped).
Suggested fix
Fall back to <OtherAbstract> when <Abstract> is absent. Type is one of PIP, KIE, NASA, or AIDS — all genuine abstracts contributed by other indexing programs, mostly on pre-1990 records. Prefer Language="eng", since OtherAbstract is also where PubMed puts translated abstracts, and a French translation should not become the text that English snippets are matched against.
Worth noting this is a data-loss bug rather than a coverage gap: the entry is not merely left un-enriched, the text it already had is deleted. That puts it in the same family as #85, which stopped a refresh demoting a full-text entry — the same protection does not cover abstract_only, because an abstract-only entry returning no abstract looks like an honest no-op.
How I hit it
Running validate over one dismech KB entry refreshed 148 cache files. Four went abstract_only → unavailable; two of those (the ones above) were this bug. The other two are correct — PMID:4869291 and PMID:37769103 genuinely have no abstract in PubMed, and the previous cache had stored the MEDLINE citation header as though it were content, so unavailable is the more honest label there.
Across the dismech cache, 15 files currently hold an abstract that lives in OtherAbstract (11 PIP, 3 AIDS, 2 NASA) and would be scrubbed on their next refresh.
Refreshing a cached PubMed record can silently delete the abstract it already held. For old records whose abstract PubMed stores in
<OtherAbstract>rather than<Abstract>, the extractor finds nothing, writescontent_type: unavailable, and overwrites the cache file with an empty## Contentsection. Any snippet validated against that text then fails with "No content available for reference".PMID:5697815is a worked example. Its cache entry held a 967-character abstract; after a refresh the file went from 1787 to 850 bytes with the abstract gone. PubMed does serve the text — it is just in a different element:There is no
<Abstract>element at all, so this line in_parse_abstractreturnsNone:https://github.com/linkml/linkml-reference-validator/blob/main/src/linkml_reference_validator/etl/sources/pmid.py#L233
PMID:6680124fails the same way (1695 → 650 bytes, a 1175-character PIP abstract dropped).Suggested fix
Fall back to
<OtherAbstract>when<Abstract>is absent.Typeis one ofPIP,KIE,NASA, orAIDS— all genuine abstracts contributed by other indexing programs, mostly on pre-1990 records. PreferLanguage="eng", sinceOtherAbstractis also where PubMed puts translated abstracts, and a French translation should not become the text that English snippets are matched against.Worth noting this is a data-loss bug rather than a coverage gap: the entry is not merely left un-enriched, the text it already had is deleted. That puts it in the same family as #85, which stopped a refresh demoting a full-text entry — the same protection does not cover
abstract_only, because an abstract-only entry returning no abstract looks like an honest no-op.How I hit it
Running
validateover one dismech KB entry refreshed 148 cache files. Four wentabstract_only→unavailable; two of those (the ones above) were this bug. The other two are correct —PMID:4869291andPMID:37769103genuinely have no abstract in PubMed, and the previous cache had stored the MEDLINE citation header as though it were content, sounavailableis the more honest label there.Across the dismech cache, 15 files currently hold an abstract that lives in
OtherAbstract(11PIP, 3AIDS, 2NASA) and would be scrubbed on their next refresh.