Skip to content

Repository files navigation

EDP-1 — Exposure Data Provenance

DOI

Status: Draft / RFC. A schema for the thing every machine-readable evidence standard leaves out: how well a study's exposure was actually measured.

HL7 FHIR's Evidence and EvidenceVariable describe a result beautifully — population, exposure, comparator, outcome, statistics, certainty. Nothing in that stack describes the instrument. A nutrition cohort reporting HR 0.74, 95% CI 0.59–0.93, certainty moderate on food-frequency-questionnaire data gets recorded faithfully and with complete silence about an instrument that correlates ~0.3 with true intake. The certainty rating is about the statistics. Nothing grades the input.

That is FDP-1's weakest-link rule one layer up, and EDP-1 is the missing field.

HL7 FHIR Evidence / EvidenceVariable   ← exists: describes the result
  └─ EDP-1 exposure provenance         ← this: how well the exposure was measured
       └─ FDP-1 value provenance       ← exists: how well the food value was known

One rule runs through all three: a claim's grade is the weakest of its layers. A large, well-analysed cohort whose exposure came from an FFQ over label-derived composition values is Grade D, and the JSON says so.

Contents

path what
EDP-1-exposure-provenance.md the spec — 7 exposure fields, 7 claim fields, 3-layer weakest link, FHIR binding
examples/fibre-two-instruments.json one claim declared twice, FFQ vs weighed record; plus what a real extraction looks like
validator/validate_edp.py reference validator, dependency-free. Also --source-audit
harness/ PMC Open Access → EDP-1 coverage measurement
harness/SLICE-SELECTION-FINDING.md why the corpus query must not mention what the corpus measures

Prior art, stated first

Most of the field list is inherited, and §1.1 of the spec says so: STROBE-nut (2016) specified nearly all of it as a reporting checklist; ONE (2019) made that checklist an OWL ontology; Nutritools/DIET@NET curates the validation coefficients; NutriGrade, NUQUEST, ROBINS-E and RoB-NObs grade exposure measurement by expert judgment; FoodData Central DOIs its releases.

What is new is narrower: binding those facts to a published effect estimate as required, machine-checkable fields; the composition-database release version on the study record; OPEN/NONE absence semantics a validator can enforce (OWL is open-world and cannot express them); and a weakest-link minimum instead of a compensatory sum. This repo claims descent on the vocabulary and novelty only on the binding.

The two gaps it closes

1. Measurement provenance (§2.4). attenuation records the instrument's agreement with a reference method — the coefficient, which statistic it is, what it was validated against, in which population, and the citation. These coefficients exist in instrument registries (Nutritools/DIET@NET, NCI), but those are indexed by instrument: nothing resolves a published cohort paper to the validation record of the instrument it used. EDP-1 carries it on the study and makes instrument_ref the join key. Rules with teeth: a coefficient without a citation is non-conforming, and instrument: ffq-validated with attenuation: OPEN regrades to ffq-unvalidated (Grade D). "Validated" is a claim with a citation attached or it is not a claim.

2. Specification space (§3.2). Patel/Burford/Ioannidis fit 8,192 specifications per exposure and found a third of variables flipped sign depending only on covariate choice. A single hazard ratio is one draw from a distribution the paper doesn't report. specification declares how many models were fit, how many were reported, how covariates were chosen, and the vibration across the space. A model selected from an undeclared many, reported as one, with no vibration analysis, is Grade D by construction.

Run it

python validator/validate_edp.py examples/fibre-two-instruments.json

cd harness
python test_offline.py                 # offline assertions, before any network call
python extract_edp.py --search --limit 300
python extract_edp.py --extract --report
python ../validator/validate_edp.py out/declarations.json --source-audit
python extract_edp.py --audit 25       # hand-validation worklist
python score_audit.py                  # precision/recall/corrected% + Wilson CI

python validate_against_qu2024.py      # score the classifier vs a PUBLISHED gold standard
python grade_citations.py --scan DIR   # grade the papers a model's parameters rest on

Set NCBI_EMAIL before any networked command (NCBI asks for a contact address), and NCBI_API_KEY to go roughly 3x faster. Fetches are cached on disk and resumable, so re-runs cost NCBI nothing.

Pilot result (n=113, 2026-07-26)

A 113-paper pilot over PMC Open Access nutrition cohorts (MeSH Diet[Majr] + cohort + cardiometabolic). Floors, not estimates — see caveats below.

disclosed %
named a dietary instrument 40.7%
named which instrument (items/version) 8.0%
reported a validation coefficient 6.2%
named a food composition database 11.5%
…with a version or access date fewer than 5% (2/113)
reported an effect estimate + CI 74.3%
…with an explicit contrast 25.7%
stated how covariates were chosen 18.6%
declared how many specifications were fit 0.9%

Evidence grade across the slice: 0% reach A, B, or C. 8.8% D, 91.2% ungradeable.

The row to stare at is database_version. A paper that names FoodData Central or NDSR but never says which release has nutrient values that are not reproducible even in principle.

Report this as a replication, not as a point estimate. 1.8% is two papers; its 95% Wilson interval is [0.5%, 6.2%], which will not support a decimal. The defensible claim is fewer than 5% name a release — and the evidence that carries weight is that the figure landed at 1.6% and 1.8% in two independently drawn corpora selected on different criteria (see SLICE-SELECTION-FINDING.md). The replication is the finding; the decimal is noise.

Read before quoting any number

  1. Floors, not estimates. Regex under-detects, and it under-detects in the direction that flatters the hypothesis. That is the first thing a reviewer will attack. Run --audit → score_audit.py and report corrected% = floor% × (precision/recall) with n and a CI. Report the floor alongside it and say which is which. Hand-code at least 50 papers — the worklist ships unfilled, and a smaller sample yields agreement, not a defensible correction.
  2. Attach a confidence interval to every rate, and never quote a decimal a handful of papers cannot support. Several rows here rest on single-digit counts.
  3. Disclosure, never practice. "X% of papers report the database version" — never "X% of studies used a versioned database." Absence means the paper did not say, not that the researchers did not do it.
  4. The corpus is biased and you must say so. PMC Open Access skews toward newer papers and OA-mandated funders. If anything it over-represents good disclosure, which makes a low number more striking, not less.
  5. Never select the slice on a field the slice measures. Requiring an instrument phrase in the query drives "named an instrument" from 40.7% to 96.7% and drags four other fields up with it. Details and the controlled comparison are in SLICE-SELECTION-FINDING.md.
  6. If you add an LLM extraction pass, validate it the same way. Publishing unvalidated machine extractions in a paper about provenance would be self-refuting.

Where this goes

The FHIR binding (§8) is a complex extension on EvidenceVariable, not a new resource — extensions are the sanctioned mechanism and it is a far smaller ask. Its sibling is an extension on NutritionProduct.nutrient, whose entire nutrient element today is item (which nutrient) and amount (how much) and nothing else. Propose them together: EDP-1 grades the study's exposure, FDP-1 grades the number that exposure was computed from, and §3.1 connects them.

The instrument and attenuation.statistic enums are candidate SEVCO terms rather than a private ValueSet — SEVCO already codes study design, risk of bias, and statistics, so these belong in the same code system rather than in a locally minted one.

The FHIR binding is being prepared as a ballot comment; see §8 of the spec.

License

Apache-2.0 — see LICENSE — with a patent non-assertion covenant. Cite via CITATION.cff.

No dependencies: Python ≥3.11, standard library only.

About

EDP-1: Exposure Data Provenance Declaration — a minimal, RFC-style spec for declaring how well a study's exposure was actually measured. Draft/RFC. Companion to FDP-1.

Topics

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages