Skip to content

feat(documents): PDF to Word (DOCX) with OCR fallback - #156

Merged
slaveofcode merged 2 commits into
developfrom
feat/documents-pdf-to-docx
Aug 3, 2026
Merged

feat(documents): PDF to Word (DOCX) with OCR fallback#156
slaveofcode merged 2 commits into
developfrom
feat/documents-pdf-to-docx

Conversation

@slaveofcode

Copy link
Copy Markdown
Owner

What

New PDF to Word (DOCX) converter — turns a PDF into an editable .docx entirely client-side, with on-device OCR for scanned pages. Nothing is uploaded.

How

  • pdf.js extracts positioned text per page → a pure reconstruction lib rebuilds real paragraphs and headings (line grouping by baseline, paragraph splitting by vertical gap, heading detection by font size).
  • Pages with little/no selectable text auto-fall back to on-device OCR (PaddleOCR, reused from the OCR tool). A Force OCR toggle runs OCR on every page (scanned docs).
  • The docx library (MIT) generates the Word file in-browser; per-page progress; page breaks between pages.

Honest limits (surfaced in the UI)

PDF stores positioned glyphs, not paragraphs — so the output is editable, reflowable text, not a pixel-perfect copy. Exact layout, tables and multi-column pages are reconstructed heuristically and may need cleanup. Faithful table/column reconstruction needs a server-side/ML engine, which would break the client-side promise.

Details

  • src/tools/documents/pdf-docx.lib.ts — pure groupLines/paragraphsFromLines/reconstruct/textDensity; 8 unit tests
  • src/islands/documents/PdfToDocx.tsx — pdf.js extraction + OCR fallback + docx generation; full worker/page cleanup
  • New dep: docx 9.7.1 (MIT, audit-clean). Registered pdf-to-docx (beta); EN + ID SEO + OG
  • 712 tests pass · 0 lint errors · tsc --noEmit clean for touched files · build green

🤖 Generated with Claude Code

slaveofcode and others added 2 commits August 3, 2026 17:24
Promote to production: Code Scratchpad lite-VS Code editing
Convert a PDF to an editable .docx entirely client-side:
- Extract positioned text per page with pdf.js, then reconstruct real
  paragraphs and headings (heuristic line/paragraph grouping + font-size
  heading detection) so the Word output is editable and reflowable.
- Pages with little/no selectable text (scanned/image PDFs) fall back to
  on-device OCR (PaddleOCR, reused from the OCR tool); a 'Force OCR' toggle
  runs OCR on every page.
- Generate the .docx in-browser with the 'docx' library (MIT). Per-page
  progress; page breaks between pages. Nothing is uploaded.

Honest about limits: PDF stores positioned glyphs, not paragraphs, so exact
layout / tables / multi-column pages are reconstructed heuristically and may
need cleanup — surfaced in the UI. Pure reconstruction lib with 8 unit tests.
Bilingual (EN + ID) UI, SEO and OG.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ubfx4XocHcECaL8twp9zsr
@cloudflare-workers-and-pages

Copy link
Copy Markdown

Deploying with  Cloudflare Workers  Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

Status Name Latest Commit Updated (UTC)
✅ Deployment successful!
View logs
goodwebtools de986d2 Aug 03 2026, 10:59 AM

@slaveofcode
slaveofcode merged commit 9e7d983 into develop Aug 3, 2026
2 checks passed
@slaveofcode
slaveofcode deleted the feat/documents-pdf-to-docx branch August 3, 2026 11:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant