fix(convert): escape text that read writes - #217
Merged
Merged
Conversation
read wrote a text node's characters straight into the Markdown, so text that spells inline syntax became that syntax on the next publish: literal "*not emphasis*" published as <em>, "<b>x</b>" as real markup, "©" as the character, "[x]: y" vanished as a link definition. escapeText (escape.go) escapes each text node as it is rendered, and only where a character could take effect, so snake_case, "a * b", "about ~5 min", AT&T and "a < b" stay readable. A node's edges border whatever the next node renders, so every rule escapes there. Each rule and each escape was measured against goldmark (_plans/056, D2). Link text escapes a "]" too, tracked by a depth counter on the renderer. escapeLinkText becomes escapeText with that rule on, so a page title, an anchor or a mention's display name holding "*Foo*" no longer renders as emphasis. inlineTextForLink's plain-text-only split is gone: text is escaped at its source, before markup surrounds it. TestStorageToMarkdownOutputParsesAsMarkdown's bracket checks now remove escapes first, and allow a literal "]" outside link text. Part of #203.
Text that starts a line with a block marker opened that block on the next publish: "<p>1. not a list</p>" became a list, "# x" a heading, "> 90 days" a blockquote, and "---" after a hard break turned the line above into a setext heading. escapeLineStarts runs over each paragraph's rendered text, split at hard breaks, and escapes a marker at the start of each line: ordered and bulleted list markers, ATX headings, blockquotes, setext underlines, thematic breaks, fences and table delimiter rows. It runs on rendered Markdown because only the assembled paragraph knows where its lines start, which is safe because nothing markfluence emits inline begins with a block marker (_plans/056, D3/D4). It applies to paragraphs, list-item text, loose text, and raw cells' loose runs. A heading's trailing " #" was read as a closing sequence and dropped, so "## Item #" published as "Item"; its first "#" is now escaped. Fixes #203.
GFM autolinks a bare URL, a www. address and an email address, so plain text holding one published as a link. Storage holds a URL as plain text only when someone chose that -- the editor turns a typed URL into a link -- so the escape keeps it text: "https\://", "www\.", "a\@b.com". It reads worse in an exported file, which docs will say (_plans/056, D7). An "@" at a text node's start is left alone, the one exception to escaping at an edge: a link whose text starts with "@" is how a mention is written (#91).
Letting the table property test write Markdown-significant text found five ways the escaping so far could change from one read to the next, or miss text entirely: - Whitespace at a text node's edge now counts as unknown, like the edge itself. renderMark moves a mark's edge whitespace outside it and publishing stores it in the next node, so trusting it escaped "\" in one read and not the next -- and a backslash before a space the mark gave away went on to escape the closing delimiter. - Adjacent text nodes are merged before rendering. coalesceSplitMarks leaves two when it merges two runs of one mark, and the split moved an edge into the middle of what publishing writes back as one node. - A thematic break must repeat one character, as CommonMark requires. The looser pattern matched "**---**", bold markup markfluence emits. - A heading that is nothing but "#" is an empty heading with a closing sequence, so its "#" is escaped too. - Raw storage written inline -- an inline macro, a list in a pipe-table cell, an <ac:link> passed through -- has its text parsed as Markdown between the tags, so "_x_" in a status macro's title published as <em>. serializeInline escapes those text nodes; a raw block needs nothing, since an HTML block is not parsed. Part of #203.
The generator's words now include text that spells Markdown -- block markers, emphasis, tildes, backticks, brackets, a tag, an entity, a backslash, a URL and an address -- which it avoided until #203. Against the converter from before #203 it fails at seed 0. Two older gaps stay out of its way: a code span never holds a backtick, since read writes one with a single backtick whatever it holds, and a mark's text never starts or ends with punctuation (#216).
docs/markdown-file.md says why an exported file holds backslash escapes, including the https\:// form a plain-text URL gets and how to turn it into a link. CLAUDE.md describes escape.go's two layers and the rule the property test holds them to: an escaping decision must not change when publishing re-splits the same text. The plan records what changed during implementation.
- Escapes in a heading reached its anchor. linkindex builds a heading's slug from its Markdown line, so "## Setup \[beta]" sent links to "#Setup-\[beta]", which Confluence never creates. extractHeadings now removes backslash escapes, since the anchor comes from the published text. - Transparent wrappers (a coloured <span>, <u>, <sup>) are flattened into the surrounding run before escaping. Text they split was escaped in pieces: an address, URL or entity split by one was missed, and "<u>a_</u>b" escaped differently from the "a_b" read back next time. - A tilde at a node's end could meet a del mark's "~~", and goldmark counts "\~~~b~~" as a run of three however the first is escaped, so the strikethrough was lost. It is now written as "~". - Raw storage written inline escapes "]" inside link text, and writes a newline as " ": a real one let goldmark start a block at the next line, or read the element as an HTML block whose escapes then published. renderRawBlock's one-line hasLooseText form is inline at the top level and is escaped there too. - ordinary prose skips the per-character pass: "." and ":" are looked for only as "www." and "://", and neighbours are computed only by the rules that need them.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #203. Plan:
_plans/056_escape-text-on-read.md, including an "Amended during implementation" section.readandexportwrote a text node's characters straight into the Markdown, so page text that spells Markdown syntax became that syntax on the next publish. #203 filed the block half (<p>1. not a list</p>publishes as a list). The same gap existed inline: literal*not emphasis*published as<em>,<b>x</b>as real markup,©as ©, and a line[x]: yvanished as a link definition. This PR escapes both, the way any Markdown writer does.What changes
escape.go'sescapeText) runs on each text node as it renders. It escapes a character only where the character could take effect, sosnake_case,a * b,about ~5 min,AT&Tanda < bstay readable. Every rule and every escape was probed against goldmark; the tables are in the plan.escapeLineStarts) are handled in a pass over each paragraph's rendered text, split at hard breaks. It covers list markers, ATX headings,>, setext underlines, thematic breaks, fences and table delimiter rows. A heading's trailing#(a closing sequence) is escaped too.].escapeLinkTextbecomes the full escaper, so a page title or display name like*Foo*no longer renders as emphasis.user-findprintsMentionMarkdown, so its line still matchesread's by construction, but it changes for names containing these characters.https\://…,www\.…,a\@b.com) so they don't turn into links on publish. This is the one visibly ugly escape; the docs say why and how to opt into a link.<ac:link>) escapes its text as well, because goldmark parses text between inline HTML tags as Markdown. A newline there is written as . This part is a plan amendment: the plan assumed raw storage never needs escaping, which is true only for raw blocks.linkindexremoves them before building a heading's slug, so links to an escaped heading keep working.The rule that took the most work
An escaping decision must not change when publishing re-splits the same text, or the Markdown isn't a fixed point. Several things split text differently on the next read:
renderMarkmoving edge whitespace out of a mark;coalesceSplitMarksmerging two runs of one mark;<span>or<u>wrapper dropping away.So whitespace at a node's edge counts as unknown, adjacent text nodes are merged, and wrappers are flattened before escaping. One goldmark quirk remains:
\~~~b~~counts as a run of three however the first tilde is escaped. So a tilde at a node's end is written~.Testing
storage2md/escapingcase (also run through the round-trip fixed-point check) and a newregression/escaped-heading-anchorcase. No existing golden changed.main's converter it fails at seed 0; with this PR all 3000 seeds pass, as did two 60s fuzz runs (~195k execs each).main's.make checkpasses.Left for later
a**(b)**c) reads back as literal asterisks. Escaping can't fix it. The generator avoids the shape.