Skip to content

About

No description, website, or topics provided.

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Doc Stitch

Doc Stitch

Document Merger — combine Word (.docx) and PDF documents into one file, from the command line or in your browser, with nothing uploaded anywhere.

Use it in your browser →


npx doc-stitch intro.docx body.docx appendix.docx -o report.docx
npx doc-stitch "scans/*.pdf" -o all.pdf --bookmarks per-document

Why another one

Both formats have the same problem, for different reasons.

For Word, the JavaScript ecosystem has one library, last published in 2017, plus four forks that exist mainly because the original produces files Word refuses to open.

For PDF, the standard library works well but merging with it drops your bookmarks entirely — a long-standing, much-reported gap. Merge ten chapters and the table of contents simply vanishes.

And every hosted alternative asks you to upload the document first, which is a hard no for the things people actually merge: resumes, contracts, legal drafts, medical records.

So doc-stitch aims at three things:

  1. The output opens, and works. Every merge is re-parsed from its own bytes and structurally validated before you get it. Word rejects a broken .docx outright; a PDF with a broken outline opens fine and just fails to navigate, which is worse — nobody finds out until later. Both are checked.
  2. Nothing is uploaded. The web app runs the engines in a Web Worker in your tab. Open the Network tab and merge — you will see no requests. It works offline.
  3. You are told what breaks. --dry-run reports what each document contains and exactly what the merge will do to it, before anything is written.

Install

npm install -g doc-stitch

Or run it without installing:

npx doc-stitch a.docx b.docx -o merged.docx

Usage

doc-stitch <input...> -o <output> [options]

  Inputs may be .docx or .pdf. They are identified by content, not by extension.

WORD OPTIONS
      --break <mode>        page | continuous | none        (default: page)
      --styles <policy>     rename | first-wins             (default: rename)
      --flatten-sections    Drop per-document page setup.
      --no-update-fields    Do not ask Word to refresh a table of contents.

PDF OPTIONS
      --bookmarks <mode>    preserve | per-document | none  (default: preserve)
      --no-forms            Do not carry interactive form fields across.
      --no-page-labels      Do not carry page labels (i, ii, iii …) across.

GENERAL
  -o, --output <file>       Where to write the merged document.
      --dry-run             Report what would happen; write nothing.
      --json                Machine-readable output.
  -q, --quiet               Only print errors.

Exit codes: 0 clean, 1 failed, 2 merged with fidelity warnings — so you can gate CI on a lossless merge.

Mixed batches

Word and PDF cannot be combined into a single file without a full layout engine to render one into the other, so a mixed batch produces one file per format:

$ doc-stitch report.docx appendix.pdf -o combined.docx
✓ Merged 1 document into combined.docx (18 KB)
✓ Merged 1 document into combined.pdf (94 KB)

See what a merge will do first

$ doc-stitch front.pdf body.pdf form.pdf --dry-run

Merge order
  PDF documents (3)
   1. front.pdf
     2 pages, Letter, portrait, 2 bookmarks
     preserved: 2 bookmarks, 1 page label range
   2. body.pdf
     3 pages, Letter, portrait, 3 bookmarks
     preserved: 3 bookmarks, 2 internal links, 1 page label range
   3. form.pdf
     1 page, Letter, portrait, 2 form fields
     preserved: 2 form fields

Warnings
  form.pdf
    • This document has no bookmarks of its own, so it contributes none to the
      merged file. Use per-document bookmarks to add one pointing at where it starts.

As a library

import { merge } from "@doc-stitch/docx";   // synchronous
import { merge as mergePdf } from "@doc-stitch/pdf";  // async
import { groupByKind } from "@doc-stitch/shared";

Neither engine uses Node built-ins, so the same code runs in the browser. The web app imports them directly and loads each on demand — merging Word documents never downloads the PDF engine.

What survives a merge

Honest accounting. "Preserved" means tested — see packages/*/test.

Word documents

Feature Status Notes
Text, paragraphs, tables ✅ Preserved
Styles and headings ✅ Preserved Collisions resolved; see below
Lists and numbering ✅ Preserved Each document keeps its own counters
Headers, footers, page setup ✅ Preserved Each document keeps its own section
Images, hyperlinks, bookmarks ✅ Preserved Bookmarks renamed on collision
Tracked changes, content controls ✅ Preserved Not flattened; ids shifted
Table of contents ⚠️ Refreshed Word rebuilds page numbers on open
Footnotes / endnotes / comments ❌ Dropped Markers removed, body text kept. Warned.
Charts, SmartArt, OLE ⚠️ Copied as-is Carried across but not reconciled

PDFs

Feature Status Notes
Pages and their content ✅ Preserved
Bookmarks / outlines ✅ Preserved Nesting, open/closed state and /Count
Internal links ✅ Preserved Retargeted to the merged page
Form fields ✅ Preserved Renamed on name collision
Page labels ✅ Preserved i, ii, iii front matter stays roman
Metadata ✅ First document's Title, author, dates
Annotations ✅ Preserved Copied with their page
Attachments, JavaScript ❌ Dropped Warned
Screen-reader tagging ❌ Dropped Warned — the merged file is less accessible
Encrypted PDFs ❌ Refused Remove the password first; no attempt is made to break it

Nothing in the ❌ rows is dropped silently — each produces a warning naming the document.

Style collisions (Word)

  • --styles rename (default) imports the second definition as Heading12 and repoints that document's content at it. Both looks survive.
  • --styles first-wins makes later documents adopt the first document's definition.

Word's structural defaults (Normal, DefaultParagraphFont, TableNormal, NoList) are always first-wins regardless of policy. Renaming them accomplishes nothing: a paragraph with no explicit style resolves to Normal by definition, so a renamed copy would sit unused while the text kept the base's formatting anyway.

Bookmarks (PDF)

  • --bookmarks preserve (default) concatenates each document's tree, retargeted to its new pages. The output looks like the inputs did.
  • --bookmarks per-document additionally wraps each document's bookmarks under a top-level entry named after the file. This is the only way to get navigation when the sources have no bookmarks of their own.

Development

npm install
npm test          # 104 tests, including validation gates on every output
npm run build     # shared + docx + pdf + cli
npm run dev       # web app on http://localhost:5173

Generate the sample documents used by the CLI and web demos:

node scripts/make-fixtures.ts

Deploying the web app

.github/workflows/deploy.yml publishes packages/web to GitHub Pages on every push to main.

It needs Pages enabled once, by hand: Settings → Pages → Source: GitHub Actions. This cannot be automated from the workflow — actions/configure-pages offers an enablement input, but creating a Pages site requires admin rights that GITHUB_TOKEN does not carry, so it fails with Resource not accessible by integration regardless of what permissions: grants. Until Pages is enabled the deploy job fails at that step; the CI workflow is unaffected.

vite.config.ts sets base: "./", so the build works from a project subpath without extra configuration.

Brand assets

brand/logo-source.png is the master artwork (2000×2000). The sizes the site actually serves live in packages/web/public/ and are committed, so no image tooling is needed to build.

There are two marks, deliberately:

Mark Used for Why
Full logo (logo-192/512.png) Header, home-screen icon, social preview Drawn at 96px or larger, where the wordmark reads
Simplified mark (favicon.svg → favicon-16/32/48/96.png) Browser tab The full logo's three elements and wide margins turn to mud at 16px

The simplified mark keeps the idea and the exact sampled palette, reduced to two sheets and a stitched seam. Regenerate after changing either source:

powershell -ExecutionPolicy Bypass -File scripts/make-logo-assets.ps1     # from brand/logo-source.png
powershell -ExecutionPolicy Bypass -File scripts/make-favicon-assets.ps1  # from favicon.svg, needs LibreOffice

To review the tab icon at the size it is actually drawn:

powershell -ExecutionPolicy Bypass -File scripts/preview-favicon.ps1

The brand teal is #0097B2. It is used for the logo and theme-color only — at 3.2:1 on the page background it is below the contrast floor for text, so links, buttons and the wordmark subtitle use --accent (#007A91 light, #4FC9E0 dark), which is the same hue at 4.7:1 or better in both themes.

Fixtures are authored in code (packages/*/test/helpers/fixtures.ts) rather than committed as binaries, so every adversarial case — colliding styles, continued lists, landscape sections, nested bookmarks, duplicate field names, roman-numeral front matter — is readable in the diff.

How it works

Word. A .docx is a ZIP of XML. Concatenating two word/document.xml bodies takes ten minutes and produces a broken file; the work is reconciling everything the body references:

Module Problem it solves
docx/merge/styles.ts Same style id, different definition
docx/merge/numbering.ts numId is document-scoped, so lists continue the wrong counter
docx/merge/rels.ts rId3 means something different in every part
docx/merge/sections.ts Page setup lives in sectPr; without a break, doc B inherits doc A's
docx/merge/ids.ts Bookmarks and revisions each number from zero

PDF. A PDF is a graph of numbered objects. Pages copy cleanly; everything that points at a page does not:

Module Problem it solves
pdf/dest.ts Destinations name page objects, whose numbers change on merge
pdf/outlines.ts Outlines are a linked tree with a /Count that must match visible children
pdf/forms.ts /AcroForm /Fields is document-level, so copied widgets stop being fillable
pdf/labels.ts An unlabelled document silently inherits the previous one's numbering

The destination work happens before pages are copied, because pdf-lib's copier follows references — an unhandled link pointing at another page drags a duplicate copy of that page into the output.

Each engine's validate.ts re-opens its own result and refuses to emit a broken file.

Verified against independent readers

The unit tests check structure with the same libraries that wrote it, which cannot catch a shared misunderstanding of a format. So both engines are also checked by an implementation that shares no code with them:

  • Word — LibreOffice converts the merged file and the page count and orientation are asserted.
  • PDF — Mozilla's pdf.js reads back the bookmarks, page labels and form fields (packages/pdf/test/independent-reader.test.ts). The same test demonstrates that a stock copyPages merge returns null for the outline where ours returns the entries.

Both run in CI.

Security

Every input is a file somebody else made, and both engines are written on that assumption. XXE, billion-laughs expansion, zip bombs and zip-slip part names are each defended against and locked in by tests in packages/docx/test/security.test.ts. See SECURITY.md for the threat model, the known limits, and how to report a vulnerability.

License

MIT

About

No description, website, or topics provided.

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages