Skip to content

Refresh deletes a cached abstract when PubMed stores it in <OtherAbstract> (PIP/NASA/KIE/AIDS) #88

Description

@cmungall

Refreshing a cached PubMed record can silently delete the abstract it already held. For old records whose abstract PubMed stores in <OtherAbstract> rather than <Abstract>, the extractor finds nothing, writes content_type: unavailable, and overwrites the cache file with an empty ## Content section. Any snippet validated against that text then fails with "No content available for reference".

PMID:5697815 is a worked example. Its cache entry held a 967-character abstract; after a refresh the file went from 1787 to 850 bytes with the abstract gone. PubMed does serve the text — it is just in a different element:

$ curl -s "https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi?db=pubmed&id=5697815&retmode=xml" \
    | grep -oE "<(Abstract|OtherAbstract)[^>]*>"
<OtherAbstract Type="PIP" Language="eng">
<AbstractText>

There is no <Abstract> element at all, so this line in _parse_abstract returns None:

https://github.com/linkml/linkml-reference-validator/blob/main/src/linkml_reference_validator/etl/sources/pmid.py#L233

abstract = soup.find("Abstract")

if not abstract:
    return None

PMID:6680124 fails the same way (1695 → 650 bytes, a 1175-character PIP abstract dropped).

Suggested fix

Fall back to <OtherAbstract> when <Abstract> is absent. Type is one of PIP, KIE, NASA, or AIDS — all genuine abstracts contributed by other indexing programs, mostly on pre-1990 records. Prefer Language="eng", since OtherAbstract is also where PubMed puts translated abstracts, and a French translation should not become the text that English snippets are matched against.

Worth noting this is a data-loss bug rather than a coverage gap: the entry is not merely left un-enriched, the text it already had is deleted. That puts it in the same family as #85, which stopped a refresh demoting a full-text entry — the same protection does not cover abstract_only, because an abstract-only entry returning no abstract looks like an honest no-op.

How I hit it

Running validate over one dismech KB entry refreshed 148 cache files. Four went abstract_only → unavailable; two of those (the ones above) were this bug. The other two are correct — PMID:4869291 and PMID:37769103 genuinely have no abstract in PubMed, and the previous cache had stored the MEDLINE citation header as though it were content, so unavailable is the more honest label there.

Across the dismech cache, 15 files currently hold an abstract that lives in OtherAbstract (11 PIP, 3 AIDS, 2 NASA) and would be scrubbed on their next refresh.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions