Skip to content

Fast path: pdf-inspector for text-based PDFs, Marker only when needed - #8

Closed
odfalik wants to merge 1 commit into
mainfrom
fast-convert-pdf-inspector
Closed

odfalik wants to merge 1 commit into
mainfrom
fast-convert-pdf-inspector

Conversation

@odfalik

@odfalik odfalik commented Aug 1, 2026

Copy link
Copy Markdown
Contributor

Adds a CPU-only fast conversion path so agents (and laptops) can ingest papers without loading Marker's ML models.

What it does

  • fast_convert.py: classifies the PDF with pdf-inspector (pure Rust, no models). Text-based/mixed PDFs with clean encodings convert to markdown in ~100ms; figures are dumped with PyMuPDF into the usual images/ folder.
  • convert_pdf_auto() routes: fast path first, automatic fallback to Marker for scanned/image-based PDFs, encoding issues, or any failure. process_paper now calls it, so behavior is unchanged for PDFs that need OCR.
  • Output layout, metadata and hashes are identical; metadata gains engine and pdf_type.

Why
On the opendataloader benchmark pdf-inspector scores 0.875 overall / 0.814 TEDS on tables, ahead of pymupdf4llm and markitdown, at ~35x the speed. Verified locally: 'Attention Is All You Need' classified text_based (confidence 1.0), converted in 111ms, 3 figures extracted, on an 8-CPU box with no GPU.

Note: pdf-inspector does not extract figures itself — that's why PyMuPDF handles images on this path.

Routes text-based PDFs through pdf-inspector (pure Rust, ~100ms, CPU-only)
with PyMuPDF figure extraction, falling back to Marker for scanned PDFs or
broken font encodings.
@odfalik

odfalik commented Aug 1, 2026

Copy link
Copy Markdown
Contributor Author

Ran the head-to-head against the existing library before merging this — results argue for re-scoping the PR.

Compared 5 papers already in R2 (their paper.md was produced by Marker on GPU) against pdf-inspector on the same source PDFs, on an 8-CPU box:

Paper Pages pdf-inspector Old chars New chars Old table rows New table rows Old math New math
CLAM_2004.09666 35 0.13s 91,789 88,586 0 9 42 0
CONCH_arxiv 57 0.32s 145,688 129,908 292 255 22 0
CellViT 23 0.28s 159,147 110,591 141 114 458 0
DNABERT-2 23 0.09s 110,733 78,701 317 222 58 0
GigaTIME_Cell_2025 35 1.93s 106,214 99,435 32 16 68 0

Speed is 100-1000x better, but: no LaTeX at all (equation fragments bleed into prose in two-column papers), fewer table rows on 4/5 papers with complex tables merged, and PyMuPDF figure extraction is unreliable (155 fragments vs 9 real figures on CellViT, 0 on DNABERT-2 whose figures are vector).

Proposal: don't make this the default path for process_paper. Keep Marker canonical, expose pdf-inspector as an explicit quick_convert + use classify_pdf as an OCR router. I'll push that revision unless someone prefers the current shape.

@odfalik odfalik closed this Aug 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant