Conversation
Routes text-based PDFs through pdf-inspector (pure Rust, ~100ms, CPU-only) with PyMuPDF figure extraction, falling back to Marker for scanned PDFs or broken font encodings.
|
Ran the head-to-head against the existing library before merging this — results argue for re-scoping the PR. Compared 5 papers already in R2 (their
Speed is 100-1000x better, but: no LaTeX at all (equation fragments bleed into prose in two-column papers), fewer table rows on 4/5 papers with complex tables merged, and PyMuPDF figure extraction is unreliable (155 fragments vs 9 real figures on CellViT, 0 on DNABERT-2 whose figures are vector). Proposal: don't make this the default path for |
Adds a CPU-only fast conversion path so agents (and laptops) can ingest papers without loading Marker's ML models.
What it does
fast_convert.py: classifies the PDF with pdf-inspector (pure Rust, no models). Text-based/mixed PDFs with clean encodings convert to markdown in ~100ms; figures are dumped with PyMuPDF into the usualimages/folder.convert_pdf_auto()routes: fast path first, automatic fallback to Marker for scanned/image-based PDFs, encoding issues, or any failure.process_papernow calls it, so behavior is unchanged for PDFs that need OCR.engineandpdf_type.Why
On the opendataloader benchmark pdf-inspector scores 0.875 overall / 0.814 TEDS on tables, ahead of pymupdf4llm and markitdown, at ~35x the speed. Verified locally: 'Attention Is All You Need' classified text_based (confidence 1.0), converted in 111ms, 3 figures extracted, on an 8-CPU box with no GPU.
Note: pdf-inspector does not extract figures itself — that's why PyMuPDF handles images on this path.