Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
118 changes: 113 additions & 5 deletions docs/troubleshooting.md
Original file line number Diff line number Diff line change
Expand Up @@ -791,6 +791,60 @@ Run through this checklist when encountering issues:
- [How to Repair Validation Errors](how-to/repair-validation-errors.md) - Fixing common issues
- [GitHub Issues](https://github.com/linkml/linkml-reference-validator/issues) - Report bugs

### Which full text is fetched, and which is not

Two rules govern what reaches the public reference cache. Both follow from the
contract the Zotero provider established: material you may lawfully *read* is
not automatically material a project may *redistribute*, and a checked-in cache
is redistribution.

**A landing page is not full text.** A provider returns a location only for a
file — OpenAlex's `pdf_url`, Unpaywall's `url_for_pdf`. A record whose only
location is an article *page* yields nothing, and the record stays
`abstract_only`. Downloading such a page is scraping, and hosts refuse it:
PMC answers a rate-limited request with a reCAPTCHA interstitial served on an
HTTP 200, which no status check downstream can distinguish from article text.

**Bronze open access is not redistributable.** `oa_status: bronze` means the
publisher has made an article free to read on its own site under no open
licence. Bronze locations are marked with a non-open `access_type` and are
skipped by ordinary validation, the same route private-library material takes.
`gold`, `diamond` and `hybrid` are treated as open; `green` only when the
location states a licence, since that status names a repository rather than a
permission. An unrecognised or missing status is **not** assumed open: an
unknown licence is not a grant.

This governs files found by asking an OA *index* (OpenAlex, Unpaywall). It does
not govern the `pmc` and `epmc_preprint` providers, which ask an archive's own
API for a document it serves for machine retrieval — a stronger warrant than an
index's summary of a third-party host. The same green deposit can therefore be
declined from OpenAlex and admitted from PMC.

**A declined reference is retried when the extractor version moves, not on every
run.** Declining is a decision, not a finding, so it must not be recorded as
`full_text_attempted` — that flag means a clean run concluded no full text
exists, and it stops the provider chain running again. But leaving nothing
recorded is its own problem: a decline never clears the way a transient failure
does, so every run would re-walk the whole chain for every bronze or
page-only reference. On one real corpus that is tens of thousands of entries and
hours of rate-limit sleep per run, aimed at the hosts whose rate limiting causes
the interstitial in the first place. So the decline is recorded as
`full_text_declined: <reason>` in the cache entry, and re-examined when the entry
is next re-fetched.

If you previously relied on landing-page or bronze full text, those references
now resolve as `abstract_only`. Excerpts quoted from that text will no longer
verify. Existing cache entries are not rewritten — see the refresh note below.

**`cache enrich` is unaffected in where it writes.** That command has always
passed `private=True`, so every file it fetches already went to the private
research cache (`~/.cache/linkml-reference-validator/private` by default),
whatever its licence. What changes on that path is the frontmatter: an enriched
bronze entry now records `full_text_access_type: publisher_free`. Its
`full_text_url` is still recorded — a publisher link is a stable public address,
and it is the *text* that licensing keeps out of the public cache, not the
address.

### JATS tables and XML cache refresh

JATS/PMC XML extraction appends pipe-delimited tables after the existing body
Expand Down Expand Up @@ -828,11 +882,65 @@ XML remains available with the existing stale-cache warning and is not rewritten
it may still lack table rows. A later process retries the refresh. Existing
stale HTML rejection remains unchanged.

A successful source refresh that returns only an abstract can replace the old
full text if no provider supplies a body. This is existing refresh behavior;
the stale fallback applies when the source returns no record, not when it
returns an abstract-only record. Keep a backup if retaining older full text is
necessary.
**A cache-wide refresh adds one line to every enriched entry.** Both index
providers now set `access_type` where they previously left it unset, so a
refreshed public entry gains a `full_text_access_type: open` line. It means the
same as its absence, but it does show up as a one-line diff on every such entry
— worth knowing before reviewing that commit.

**`cache reference` exits 1 on a preserved entry, every run.** A preserved entry
is never written, so it is never stamped, so the next run reaches the same
decision and fails again. A script that caches a list of references and gates on
the exit status will keep failing on those until the source serves full text
again, or `--force` accepts the abstract-only refresh in its place. That is the guard working as
intended, but it is the kind of thing that gets diagnosed twice if it is not
written down.

A successful source refresh that returns only an abstract **no longer replaces**
cached full text. Full-text retrieval fails transiently and silently — a
rate-limited PMC request is answered with a reCAPTCHA interstitial carried on an
HTTP 200 — so an abstract-only refresh is not evidence that the article has no
full text. When a refresh loses full text the cached entry is kept, a warning
names it, and the entry stays stale so a later run tries again.

Losing it means coming back with *no* full text — an `abstract_only`,
`unavailable` or `summary` record where the cache holds `full_text_*`. Sizes are
not compared **to decide a refusal**. An earlier version of this guard also refused a refresh whose text
was a fraction of the cached length, to catch a PDF whose text layer is a
publisher cover sheet; it caught that, and wrongly refused four kinds of genuine
improvement — a scraped page replaced by a clean XML body, by a clean HTML body,
a plain-text API body replaced by XML, and a re-extraction that merely trimmed a
trailing section. A shorter extraction is usually a better one, and a length
comparison cannot tell those apart.

Wrongly refusing is the worse error: a refused entry is never written, so it is
never stamped, so every later run re-fetches and re-refuses it. Judging whether
text *is* an article belongs in the acceptance layer, where a wrong answer costs
one skipped fetch instead of a cache that can never migrate.

Size is still *reported*, which has none of those properties. When a refresh
keeps full text but returns under a fifth as much article text as the cache
held, the entry is written as usual and a warning names both figures:

```
Refresh of PMID:9177246 replaced the cached full_text_html entry (18,464
characters of article text) with a much shorter one: 602 characters of
full_text_pdf. Written as usual, since a shorter extraction is often a cleaner
one — but check it if quoted excerpts stop verifying.
```

Usually that is a cleaner extraction and there is nothing to do. It is the
thread to pull when a quoted excerpt stops verifying: re-read the cached entry,
and if the text is a publisher cover sheet rather than the article, re-fetch when
the source will serve the real thing. The lengths quoted are of the article text,
with the abstract both records carry subtracted.

A refresh that *finds* full text still rewrites the entry. `force_refresh`
(`--force`) overrides the refusal, but not the notice: it logs a warning naming
the cached entry it is about to replace and its length, because it is the remedy
this tool recommends and following that advice should not quietly discard an
article body. This is narrower than the stale fallback, which applies
only when the source returns no record at all.

This pass targets JATS `table-wrap` content and searches the whole document;
tables and notes inside embedded `sub-article` or `response` elements are
Expand Down
35 changes: 24 additions & 11 deletions src/linkml_reference_validator/cli/cache.py
Original file line number Diff line number Diff line change
Expand Up @@ -150,19 +150,24 @@ def reference_command(

outcome = fetcher.fetch_with_provenance(reference_id, force_refresh=force)

# A reference that could not be re-fetched falls back to an out-of-date cache
# entry, which is the right answer for validation but not here. What this
# command promises is that the cache holds a current entry afterwards - not
# that it downloaded one, since an entry the current extractor already wrote
# needs no download. A stale entry leaves that promise unmet, so reporting
# success would take a script that gates on the exit status green through an
# outage. The reason is deliberately left open: the source may be unreachable,
# or no source may handle this identifier at all.
# What this command promises is that the cache holds a current entry
# afterwards - not that it downloaded one, since an entry the current
# extractor already wrote needs no download. Two paths leave that promise
# unmet, and the message has to be true of both: the reference could not be
# re-fetched and an out-of-date entry was served instead, or it was
# re-fetched and the result was refused for holding less full text than the
# entry already cached. Either way the write was skipped, so what is
# reported is that, rather than a guess at which path ran. Reporting success
# would take a script that gates on the exit status green through an outage.
if outcome.served_stale:
typer.echo(
f"Failed to cache {reference_id}: it could not be re-fetched, so an "
"out-of-date cache entry was served. The cache still holds no current "
"entry for it.",
f"Failed to cache {reference_id}: it could not be re-fetched, or the "
"refresh came back with no full text where the cache holds some. "
"Either way no entry was written, so the cache still holds no current "
"entry for it. Re-run when the source serves full text again. "
"(--force replaces cached text with whatever a refresh returns, so it "
"resolves the second case and not the first: with the source "
"unreachable there is nothing to put in its place.)",
err=True,
)
raise typer.Exit(1)
Expand Down Expand Up @@ -303,6 +308,14 @@ def enrich_command(
typer.echo(f"{reference.reference_id}\tnot_found\t-")
continue

if location.declined:
# Reported as its own outcome, not as a find and not as an absence.
# A declined location carries no url and no text, so counting it
# would inflate `Found:` by exactly the references the provider
# refused -- and this command's whole output is an inventory.
typer.echo(f"{reference.reference_id}\tdeclined\t{location.declined}")
continue

found += 1
source = f"{location.provider or provider}:{location.source_item_id or '-'}"
if dry_run:
Expand Down
130 changes: 130 additions & 0 deletions src/linkml_reference_validator/etl/fulltext/base.py
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,136 @@

logger = logging.getLogger(__name__)

#: ``oa_status`` values that are openly licensed on their own.
#:
#: ``bronze`` is deliberately absent. It means the publisher has made the article
#: free to read on its own site under no open licence: readable, not
#: redistributable. Treating "I may read this" as "this project may republish
#: this" is the distinction PR #61 drew for Zotero private libraries, and it
#: applies identically to a bronze PDF.
#:
#: ``green`` is absent too, and for a subtler reason: it describes *where* a copy
#: lives -- a repository -- not the terms it lives there under. A green PMC
#: author manuscript is free to read under a funder policy whose redistribution
#: terms vary by publisher, which is the same shape of claim as bronze with a
#: different host. It is admitted by :func:`access_type_for_oa_status` only when
#: the location states a licence.
#:
#: **Scope.** This governs a file found by asking an OA *index* -- OpenAlex or
#: Unpaywall -- where the index's own status is all that is known about the
#: terms. It deliberately does not govern ``pmc`` or ``epmc_preprint``, which
#: return ``oa_status="green"`` with no ``access_type`` and so continue to write
#: to the public cache. Those providers ask the archive's own API for a document
#: it serves for machine retrieval, which is a stronger warrant than an index's
#: summary of a third-party host, and routing them through this rule would
#: decline most PMID full text for want of a licence field their API does not
#: return. The asymmetry is deliberate: the same green deposit can be declined
#: from OpenAlex and admitted from PMC.
SELF_EVIDENTLY_OPEN_OA_STATUSES = frozenset({"gold", "diamond", "hybrid"})

#: ``oa_status`` values that are open *if* the location states a licence. See
#: the note on ``green`` above.
LICENCE_DEPENDENT_OA_STATUSES = frozenset({"green"})

#: Licence values outside the ``cc-`` family that grant redistribution.
OPEN_LICENCES = frozenset({"cc0", "public-domain", "mit"})


def states_a_licence(licence: Optional[str]) -> bool:
"""Report whether a location's ``license`` field grants redistribution.

An **allowlist**, matching how this module treats ``oa_status``: a value it
has not heard of is an unknown licence, and an unknown licence is not a
grant. A truthiness test would not do, because neither API uses ``None`` as
its only way of saying "no licence statement". OpenAlex's vocabulary
includes ``other-oa`` (5.9M works) and ``publisher-specific-oa``, and
Unpaywall documents ``implied-oa`` for a copy it believes free with no
licence statement found. All three are truthy strings naming the *absence*
of a licence -- and a bare funder-policy repository deposit, which is the
case the ``green`` rule was written about, is exactly where they appear.

Every Creative Commons variant qualifies. ``nc`` restricts commercial use
and ``nd`` restricts derivatives; neither restricts holding a verbatim copy,
which is all a cache does.

Examples:
>>> states_a_licence("cc-by")
True
>>> states_a_licence("cc-by-nc-nd")
True
>>> states_a_licence("public-domain")
True

The sentinels that mean "no licence statement" do not:

>>> states_a_licence("other-oa")
False
>>> states_a_licence("implied-oa")
False
>>> states_a_licence(None)
False
"""
value = (licence or "").strip().lower()
return value.startswith("cc-") or value in OPEN_LICENCES


#: ``access_type`` for a location that is free to read but not openly licensed.
#: Any non-``open`` value is skipped by ``_enrich_with_full_text``; naming it
#: distinctly keeps the reason legible in a log line.
PUBLISHER_FREE_ACCESS = "publisher_free"


def access_type_for_oa_status(
oa_status: Optional[str], licence: Optional[str] = None
) -> str:
"""Map an ``oa_status`` (and licence, where it decides) to an ``access_type``.

Unrecognised and missing statuses are **not** treated as open. A status this
version has not heard of is an unknown licence, and the safe reading of an
unknown licence is that it does not grant redistribution.

``green`` is decided by the licence rather than the status, because the
status names a repository rather than a permission: a deposit that states
its licence is open, a bare one is not. "States its licence" means
:func:`states_a_licence`, not merely a non-empty field -- see there for why
the difference matters.

Known miss: a work's ``oa_status`` is its *best* status across locations, so
a work marked ``bronze`` may still carry an openly-licensed repository copy
in a location this code never inspects. Conservative rather than wrong --
some redistributable full text is skipped -- and widening it means walking
every location instead of the best one.

Examples:
>>> access_type_for_oa_status("gold")
'open'
>>> access_type_for_oa_status("bronze")
'publisher_free'
>>> access_type_for_oa_status(None)
'publisher_free'
>>> access_type_for_oa_status("something-new")
'publisher_free'

A repository deposit is open when it states its terms, and not when it
merely states its address:

>>> access_type_for_oa_status("green", licence="cc-by")
'open'
>>> access_type_for_oa_status("green")
'publisher_free'

A licence does not rescue a status that is not licence-dependent:

>>> access_type_for_oa_status("bronze", licence="cc-by")
'publisher_free'
"""
status = (oa_status or "").strip().lower()
if status in SELF_EVIDENTLY_OPEN_OA_STATUSES:
return "open"
if status in LICENCE_DEPENDENT_OA_STATUSES and states_a_licence(licence):
return "open"
return PUBLISHER_FREE_ACCESS


class FullTextProvider(ABC):
"""Abstract base class for full-text providers."""
Expand Down
Loading
Loading