Skip to content

fix(convert): escape text that read writes - #217

Merged
willkg merged 9 commits into
mainfrom
fix/escape-text
Sep 28, 2026
Merged

willkg merged 9 commits into
mainfrom
fix/escape-text

Conversation

@willkg

@willkg willkg commented Sep 28, 2026

Copy link
Copy Markdown
Member

Fixes #203. Plan: _plans/056_escape-text-on-read.md, including an "Amended during implementation" section.

read and export wrote a text node's characters straight into the Markdown, so page text that spells Markdown syntax became that syntax on the next publish. #203 filed the block half (<p>1. not a list</p> publishes as a list). The same gap existed inline: literal *not emphasis* published as <em>, <b>x</b> as real markup, &copy; as ©, and a line [x]: y vanished as a link definition. This PR escapes both, the way any Markdown writer does.

What changes

  • Inline escaping (escape.go's escapeText) runs on each text node as it renders. It escapes a character only where the character could take effect, so snake_case, a * b, about ~5 min, AT&T and a < b stay readable. Every rule and every escape was probed against goldmark; the tables are in the plan.
  • Line starts (escapeLineStarts) are handled in a pass over each paragraph's rendered text, split at hard breaks. It covers list markers, ATX headings, >, setext underlines, thematic breaks, fences and table delimiter rows. A heading's trailing # (a closing sequence) is escaped too.
  • Link text also escapes ]. escapeLinkText becomes the full escaper, so a page title or display name like *Foo* no longer renders as emphasis. user-find prints MentionMarkdown, so its line still matches read's by construction, but it changes for names containing these characters.
  • Bare URLs and email addresses in plain text are escaped (https\://…, www\.…, a\@b.com) so they don't turn into links on publish. This is the one visibly ugly escape; the docs say why and how to opt into a link.
  • Raw storage written inline (an inline macro, a list in a pipe-table cell, a passed-through <ac:link>) escapes its text as well, because goldmark parses text between inline HTML tags as Markdown. A newline there is written as &#10;. This part is a plan amendment: the plan assumed raw storage never needs escaping, which is true only for raw blocks.
  • Anchors ignore escapes. linkindex removes them before building a heading's slug, so links to an escaped heading keep working.

The rule that took the most work

An escaping decision must not change when publishing re-splits the same text, or the Markdown isn't a fixed point. Several things split text differently on the next read:

  • renderMark moving edge whitespace out of a mark;
  • coalesceSplitMarks merging two runs of one mark;
  • a <span> or <u> wrapper dropping away.

So whitespace at a node's edge counts as unknown, adjacent text nodes are merged, and wrappers are flattened before escaping. One goldmark quirk remains: \~~~b~~ counts as a run of three however the first tilde is escaped. So a tilde at a node's end is written &#126;.

Testing

  • Unit tests: one row per rule, each also checking that the escaped form publishes back as the original text. There's also a table of every shape the code review found, which checks what publishing actually sends, not only that the Markdown is stable.
  • Goldens: a new storage2md/escaping case (also run through the round-trip fixed-point check) and a new regression/escaped-heading-anchor case. No existing golden changed.
  • Table property test: the generator now writes Markdown-significant text. Against main's converter it fails at seed 0; with this PR all 3000 seeds pass, as did two 60s fuzz runs (~195k execs each).
  • Real pages: exports of three live SRE pages are byte-identical to main's.
  • make check passes.

Left for later

read wrote a text node's characters straight into the Markdown, so text
that spells inline syntax became that syntax on the next publish:
literal "*not emphasis*" published as <em>, "<b>x</b>" as real markup,
"&copy;" as the character, "[x]: y" vanished as a link definition.

escapeText (escape.go) escapes each text node as it is rendered, and
only where a character could take effect, so snake_case, "a * b",
"about ~5 min", AT&T and "a < b" stay readable. A node's edges border
whatever the next node renders, so every rule escapes there. Each rule
and each escape was measured against goldmark (_plans/056, D2).

Link text escapes a "]" too, tracked by a depth counter on the
renderer. escapeLinkText becomes escapeText with that rule on, so a
page title, an anchor or a mention's display name holding "*Foo*" no
longer renders as emphasis. inlineTextForLink's plain-text-only split
is gone: text is escaped at its source, before markup surrounds it.

TestStorageToMarkdownOutputParsesAsMarkdown's bracket checks now remove
escapes first, and allow a literal "]" outside link text.

Part of #203.
Text that starts a line with a block marker opened that block on the
next publish: "<p>1. not a list</p>" became a list, "# x" a heading,
"> 90 days" a blockquote, and "---" after a hard break turned the line
above into a setext heading.

escapeLineStarts runs over each paragraph's rendered text, split at
hard breaks, and escapes a marker at the start of each line: ordered
and bulleted list markers, ATX headings, blockquotes, setext
underlines, thematic breaks, fences and table delimiter rows. It runs
on rendered Markdown because only the assembled paragraph knows where
its lines start, which is safe because nothing markfluence emits
inline begins with a block marker (_plans/056, D3/D4). It applies to
paragraphs, list-item text, loose text, and raw cells' loose runs.

A heading's trailing " #" was read as a closing sequence and dropped,
so "## Item #" published as "Item"; its first "#" is now escaped.

Fixes #203.
GFM autolinks a bare URL, a www. address and an email address, so plain
text holding one published as a link. Storage holds a URL as plain text
only when someone chose that -- the editor turns a typed URL into a
link -- so the escape keeps it text: "https\://", "www\.", "a\@b.com".
It reads worse in an exported file, which docs will say (_plans/056,
D7).

An "@" at a text node's start is left alone, the one exception to
escaping at an edge: a link whose text starts with "@" is how a mention
is written (#91).
Letting the table property test write Markdown-significant text found
five ways the escaping so far could change from one read to the next,
or miss text entirely:

- Whitespace at a text node's edge now counts as unknown, like the edge
  itself. renderMark moves a mark's edge whitespace outside it and
  publishing stores it in the next node, so trusting it escaped "\" in
  one read and not the next -- and a backslash before a space the mark
  gave away went on to escape the closing delimiter.
- Adjacent text nodes are merged before rendering. coalesceSplitMarks
  leaves two when it merges two runs of one mark, and the split moved
  an edge into the middle of what publishing writes back as one node.
- A thematic break must repeat one character, as CommonMark requires.
  The looser pattern matched "**---**", bold markup markfluence emits.
- A heading that is nothing but "#" is an empty heading with a closing
  sequence, so its "#" is escaped too.
- Raw storage written inline -- an inline macro, a list in a pipe-table
  cell, an <ac:link> passed through -- has its text parsed as Markdown
  between the tags, so "_x_" in a status macro's title published as
  <em>. serializeInline escapes those text nodes; a raw block needs
  nothing, since an HTML block is not parsed.

Part of #203.
The generator's words now include text that spells Markdown -- block
markers, emphasis, tildes, backticks, brackets, a tag, an entity, a
backslash, a URL and an address -- which it avoided until #203. Against
the converter from before #203 it fails at seed 0.

Two older gaps stay out of its way: a code span never holds a backtick,
since read writes one with a single backtick whatever it holds, and a
mark's text never starts or ends with punctuation (#216).
docs/markdown-file.md says why an exported file holds backslash
escapes, including the https\:// form a plain-text URL gets and how to
turn it into a link. CLAUDE.md describes escape.go's two layers and the
rule the property test holds them to: an escaping decision must not
change when publishing re-splits the same text. The plan records what
changed during implementation.
- Escapes in a heading reached its anchor. linkindex builds a heading's
  slug from its Markdown line, so "## Setup \[beta]" sent links to
  "#Setup-\[beta]", which Confluence never creates. extractHeadings now
  removes backslash escapes, since the anchor comes from the published
  text.
- Transparent wrappers (a coloured <span>, <u>, <sup>) are flattened
  into the surrounding run before escaping. Text they split was escaped
  in pieces: an address, URL or entity split by one was missed, and
  "<u>a_</u>b" escaped differently from the "a_b" read back next time.
- A tilde at a node's end could meet a del mark's "~~", and goldmark
  counts "\~~~b~~" as a run of three however the first is escaped, so
  the strikethrough was lost. It is now written as "&#126;".
- Raw storage written inline escapes "]" inside link text, and writes a
  newline as "&#10;": a real one let goldmark start a block at the next
  line, or read the element as an HTML block whose escapes then
  published. renderRawBlock's one-line hasLooseText form is inline at
  the top level and is escaped there too.
- ordinary prose skips the per-character pass: "." and ":" are looked
  for only as "www." and "://", and neighbours are computed only by the
  rules that need them.
@willkg
willkg merged commit e7ac876 into main Sep 28, 2026
1 check passed
@willkg
willkg deleted the fix/escape-text branch September 28, 2026 21:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

read: paragraph text that looks like a Markdown block becomes structure on publish

1 participant