diff --git a/CLAUDE.md b/CLAUDE.md index 40f249f..804c3a6 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -79,7 +79,7 @@ Module `github.com/mozilla/markfluence` (`go 1.25`). `main.go` is a shim to `cmd - `internal/pagetree` — `Walk`, the traversal of pages *and folders* under a node, plus `WalkSpace` (the same traversal seeded from a space's root pages, via `client.ListSpaceRootPages`) and `AllDepths`. Both go through one `walker`, so the depth rule and the visited guard exist in a single copy. It is a package rather than command-local because listing a subtree and exporting one (#59) need the identical walk, and its rules must not exist in two copies: siblings arrive from two requests (`/child/page`, `/child/folder`) and are **merged by `extensions.position`**, or the output loses the order Confluence displays; a folder **counts as a level** like a page, which is only reasonable because folders are reported rather than silently traversed; and the walk descends folders even when only pages matter, since a folder may hold the only pages in a subtree. `nodeURL` uses `SiteURL()` — a v1 child row carries `webui` but no `base`. A visited set guards the unbounded case. - `internal/pageref` — `Resolve`, the single page-argument resolver: a numeric id, a Confluence page **or folder** URL (`pagePathRE` matches both `/pages/` and `/folder/`, since `children` takes a folder and a folder URL is what a browser hands you — the id is all it returns, so a command that can only use a page reports its own not-found), or a `.md` file that **declares** a `page_id` — in its own frontmatter or in a `pages:` entry for it (#139), resolved through `pagemeta` after discovering the root from the file's own directory (stat'd first, so `123.md` is a file). Every command taking a page uses it, which is why the manifest lookup lives here rather than at the seven call sites: without it the page argument meant one thing to `update` and another to `page-info`/`read`/`children`/`export`/`attachment-*`, so a file `update` could publish could not be named to any of them. A **disagreement** between the two locations is fatal here — it is the question being asked — while a **malformed** `markfluence.yaml` is not: a project file this resolver never consults must not make `page-info 123` fail, and the commands that bound reads by the root report it themselves. `message.go` also owns the wording for the two ways a *frontmatter* `page_id` is wrong — `NotFoundMessage` (caller supplies the remedy, which differs per command) and `NotNumericMessage` — because `create` and `update` report both and `check` reports the non-numeric one, and a reader should recognize the same problem across all of them. They return strings, not errors: `create` wraps the text in its typed `pageIDFailure` (which also carries the `--json` fields), the others want a plain error. Anything checking a `page_id` before a request uses `IsDigits`, since the API answers a non-numeric id with a 400 whose body says nothing useful. - `internal/client` — `ConfluenceClient` over `net/http` with basic auth. Built from a `Config` (site URL, cloud ID, username, token) via `New`; it carries **two bases**: `BaseURL()` is where requests go (the gateway when a cloud ID is set) and `SiteURL()` is always the site. Anything a reader sees uses `SiteURL()` — printed page URLs and, critically, the `baseURL` handed to `convert.MdToConfluence`, since rewritten links are published *into* the page. Pages are Confluence **v2**; attachment writes and the user lookup are **v1** (`/wiki/rest/api/...`). A **folder** — the Cloud content type that can parent a page — has its own v2 route, `GetFolderOrNil` against `/wiki/api/v2/folders/{id}`, because every v2 *page* route answers a folder id with 404; enumerating children, if it is ever added, must be v1, since v2 cannot list inside a folder at all and its page-children route silently omits folders ([docs/confluence/folders.md](docs/confluence/folders.md)). Page **status** — the title lozenge — is v1 only and lives in `state.go` (`PageState`/`AvailableStates`/`SetPageState`, plus `StateVocabulary`, which keeps the space's statuses and the caller's own custom ones in separate fields because only the first are valid for a file); v2 carries no state field on a page in any form and there is no expansion that adds one, so a lozenge is one extra request per page, always. `AvailableStates` must be asked about the page the status is going on — its answer varies by caller *and* page, and it needs edit permission on that page. A **move** is `MovePage` (`move.go`), always the v1 `PUT /content/{id}/move/{position}/{targetId}` with `append` (under a page or folder) or `after` (after the last top-level page, the only way to the top of a space): the v1 route leaves the page version alone where a v2 `parentId` change bumps it, and v2 silently ignores a null `parentId`, so it cannot reach the top at all ([docs/confluence/api.md](docs/confluence/api.md#moving-a-page)). There is deliberately **no `ClearPageState`**: the `DELETE` route exists, but nothing can reach it until a clearing spelling does, and an unused write method is a loaded gun. Typed `HTTPError`, per-attempt context timeouts, centralized retry/backoff in `send`. `HTTPError.Error()` appends a **hint** for the three auth failures whose status misleads, matched on the response *body* rather than deduced from the status and always **appended** to it, never replacing it. The one that matters: **a rejected credential is a 404 on every v2 route**, so a revoked token used to make `read` answer `page ... not found` about a page that exists. `RejectedCredential` tells it apart by the fact that every genuine v2 404 *names* what it could not find and the auth one does not, `notFound` gates the three `…OrNil` helpers on it so they stop reading it as "absent", and `jsonout.CodeFor` checks it before the status switch so `--json` reports `AUTH` rather than `NOT_FOUND`. **Two error types on the request path, and one predicate for them**: an `*HTTPError` once a response has a status, an unexported `requestError` when there is none (a transport failure, a request that would not build, a body that would not decode), and `FromRequest` answers whether an error is either. That is what lets a caller tell a server failure from a local one — `jsonout.CodeOr(err, fallback)` is the whole point of it, since `CodeFor` alone reports every non-`HTTPError` as `NETWORK` and so turns `no title given` into a network problem (#133). The rule is deliberately scoped to the request: `DownloadAttachment` writing to the caller's writer, `uploadAttachment` opening the caller's file, and `Resolve` reading the environment stay untyped, because tagging them would misreport an unreadable file as a network failure. The wrapper carries no message of its own, so `Error()` is the inner text verbatim and nothing a reader sees changed. A 403 that is not one of the two measured credential phrasings gets no hint, because that is what a genuine permission denial looks like ([docs/confluence/api.md](docs/confluence/api.md#scopes)). **Retry rules**: 429 for any method; 502/503/504 for idempotent methods; **any other 5xx only when the response carries `Retry-After`** — that is how a 500 becomes retryable, and it is why `parseRetryAfter` reports the header's *presence* apart from its delay (`Retry-After: 0` means "retry now", not "no header"). The exponential delay is jittered, a server-supplied `Retry-After` never is. Decisions go to a package-level hook (`SetRetryLogger`, set once in `root.go` beside `ui.SetDebug`) and fire whichever way they went, because `internal/client` prints nothing and a silent twelve-minute retry storm is otherwise indistinguishable from a hang. **A versioned PUT is not as idempotent as its method**: `SetContentProperty` retry-once on top (recovers a lost create-POST response) and `UpdatePage`'s `updateLanded` both exist for the same reason — a write whose response was lost gets re-sent, and the re-sent version is refused. `updateLanded` requires version *and* title *and* body to match what was sent, since a concurrent edit could have produced the version alone and claiming success over someone else's content is worse than a false failure ([docs/confluence/api.md](docs/confluence/api.md)). `SyncAttachments` (skip/update by a SHA-256 recorded in the attachment's comment, alongside the source path so `read` recovers image paths exactly; only the current comment form is parsed — an attachment stamped by a markfluence predating a comment-format change reads as unmanaged and is re-uploaded once, the same as any hand-uploaded file — except that a *recorded path disagreeing with the local source* is an update even when the checksum matches, so a mangled path repairs itself instead of surviving every later publish; a comment with no source recorded at all is not a disagreement. Every text part of the upload form must go through `writeTextField`, never `multipart.Writer.WriteField`, which emits no charset and gets decoded as Latin-1), `_links.next` pagination. Every attachment file is read through `LocalAttachment.Open`, twice — once for `planAttachments`' checksum, once by `uploadAttachment`, which reads the file whole and takes the comment's checksum from *those* bytes, since the file may have changed in between and a comment misdescribing its content reads as up to date on the next publish. **Four pagination schemes, and picking the wrong one truncates silently.** v1 *child/attachment* collections page through the generic `listV1` helper by `start`/`limit` offset, never `_links.next` (absent when the results fit one page, so it cannot terminate a loop); `ListAttachments`, `ListChildPages`, and `ListChildFolders` all go through it. v2 collections page through `listV2` by the cursor in `_links.next` (whose loop is `walkV2`, shared rather than copied so a counting caller can stream — a second implementation of v2 paging is how one of them comes to terminate on a short page), which is a `/wiki`-prefixed absolute path `resolveNext` handles unchanged; `ListContentProperties` and `SearchPagesByTitle` share it. **`/wiki/rest/api/search` is neither**: it ignores `start` outright, its `next` is context-relative so it needs the `/wiki` prefix `resolveNext` does not add, a short page does *not* mean the end, and `totalSize` can be nonzero against an empty `results` — so `searchCQL` terminates only on a missing `next` and nothing may branch on `totalSize` ([docs/confluence/search.md](docs/confluence/search.md)). `searchCQLBounded` adds a row bound under it (`SearchCQL` is that call with no bound, which is why `find` is unaffected): it asks for `max+1` and reports the surplus as `more`, since `totalSize` cannot supply a count. **`/wiki/rest/api/space` is a fifth, and the one that punishes the obvious choice**: it pages by `start`/`limit` offset exactly as the child collections do, and **a short page is not the end** — asked for 250 from `start=0` it answered 200, and `start=200` then answered 250 more, against 525 spaces. `listV1` stops on that short page, so `WalkSpaceOperations` has its own loop terminating on an **empty** page, advancing by rows *returned* rather than by the limit asked for, bounded by `maxSpacePages` since an empty page is the only end signal offset paging has here. The first version of the probe that found this trusted the short page and reported 200 spaces with total confidence ([users.md](docs/confluence/users.md)). **`/wiki/rest/api/search/user` is a fourth**, and the one that looks most like an existing scheme while not being it: it pages by `start`/`limit` offset exactly as `listV1` does, so `listV1` is the obvious home for it and is a trap — the route **caps a page at 100 rows while echoing back whatever limit was requested** (101, 250 and 500 all answer 100), and `listV1` asks for `v1PageSize = 250` and reads a short page as the end of the collection, so it would truncate every result set past 100 with no error at all. `SearchUsers` lives in its own `users.go` with `userPageSize = 100` and the measurement beside it for that reason, and `TestPageCapDoesNotTruncate` is the regression. It also carries `maxUserPages`, `searchCQLBounded`'s guard for the same hazard reached a different way: a short page is the *only* end signal offset paging here has, so a server that clamped `start` — or ignored it the way `/wiki/rest/api/search` ignores it outright — would return a full page forever and an unbounded walk would collect rows until it ran out of memory. Its `totalSize` is a *third* kind of wrong: not absent like v1's and not an estimate like `/search`'s, but the row count of the page just fetched, so `limit=3` answers 3 and `limit=500` answers 100 against 304 real matches. `user.go` holds the two identity routes (`CurrentUser`, `UserInfo` — both `read:confluence-user`, both seeing a deactivated account the directory cannot) and `WalkSpaceOperations`; `space.go` holds `GetSpace` (one v1 request answering identity, the caller's own operations, description, labels and the homepage *with its title*), `SpaceStateSettings` (space-admin only, so a 403 that is not a rejected credential is `(nil, nil)` rather than an error) and `WalkSpacePages`. Both space routes decode the space `id` as a `json.Number`: v1 reports it as a **number** where every v2 route reports a string, and `homepage.id` in the same response is a string. Full text goes through `SearchText`/`SearchRawCQL`, which return the cleaned `SearchMatch` the way `FindByTitle` returns `TitleMatch` — and **every field of a match comes from the row's `content` object**, because the row-level `title` is HTML-escaped *and* wrapped in `@@@hl@@@` markers where `content.title` is neither. The `excerpt` exists only at row level, so `cleanExcerpt` strips those markers, unescapes once, and collapses to one line — in the client, so the human and `--json` paths cannot disagree about it. `excerpt=highlight` is passed explicitly and **re-attached when following the cursor** (the `next` link carries `cql` and `limit` but not `excerpt`, and `doJSON` appends params with a bare `?`); an unrecognized value there yields an empty excerpt with a 200, so a rename by Atlassian degrades to no excerpts rather than an error. A row with no `content` object is skipped and **counted** — `type = space` answers with hundreds of them, and a silent skip would report a successful empty result. A bare v1 child row already carries `webui`, `status`, and `extensions.position`, so child listing needs no `expand`. `DownloadAttachment` goes through `send` (inheriting retry/backoff) against `_links.download`; **never** add a `CheckRedirect` that forwards headers — it would leak site credentials to Atlassian's media host, which neither needs nor wants them. `credentials.go` holds the credentials file's own functions — `CredentialsPath`, `ReadCredentials` (the rules `Resolve` uses, minus the permission warning, for a caller about to rewrite the file; one read returning a `CredentialsFile` of values, the lines a rewrite drops, and the mode), `WriteCredentials` (temp-file-and-rename at `0600`, writing through a symbolic link, quoting a value exactly when `readDotenv` would not read it back unchanged, then reading the result back and comparing), `DisplayPath`, and `FetchCloudID`, the unauthenticated `tenant_info` request, which bypasses `send` because `send` always sets basic auth. The setting names (`URLVar`/`UsernameVar`/`TokenVar`/`CloudIDVar`), `CredentialsDoc` and `LooseMode` are exported so `credentials-init` copies none of them. `HTTPError.ScopeMismatch`/`SiteRejectedAuth`, beside `RejectedCredential`, are the shapes `hint` matches, exported for `credentials-init`. `config.go` holds `Resolve` and the env-file reader (`loadDotenv`, which warns, over `readDotenv`, which does not), plus the **permission warning** (#136): a *regular* credentials file or `--env-file` reachable by anyone but its owner (`mode.Perm()&0o077`; a pipe from `--env-file <(pass show …)` reports 0440 and no chmod can fix it) *and* containing `CONFLUENCE_TOKEN` earns a warning naming the file, its mode, and the `chmod`. Both halves matter — a file holding only the URL and username leaks nothing, and a warning that fires on a file with no secret in it is how one becomes something people scroll past. It stats rather than lstats (a link's own `0777` would cry wolf over a `0600` target), lives in `loadDotenv` because that is the one function both the credentials file and `--env-file` pass through, and reaches the reader through `SetSecurityWarner` for the same reason `SetRetryLogger` exists — wired to `cmd/root.go`'s `reportSecurityWarning`, which prints it (human mode) *and* records it via `jsonout.AddWarning`, since stderr under `--json` is a schema-validated document with no room for a stray line. A group/world-*writable* file with no token in it is knowingly **not** covered: the same-source rule means a URL rewritten there can no longer be paired with a token from somewhere else, and the cloud ID follows the URL. Why each of these is shaped this way, with the evidence: [docs/confluence/api.md](docs/confluence/api.md) and [attachments.md](docs/confluence/attachments.md). -- `internal/convert` — the converter (the crux). `MdToConfluence(md *frontmatter.MarkdownFile, root *project.Root, index *linkindex.Index, baseURL, spaceKey string) (*ConfluencePage, error)`. `root` bounds which images and parent references may be read (S1/S2) and is what an image's recorded `Source` is relative to; `index` is the tree-wide link/anchor index for `root` (`internal/linkindex.Build`), built once and shared across every file converted under it rather than rebuilt per conversion — both are discovered/built by the caller (`internal/project`/`internal/linkindex`), which is why this package stays client-free. It parses with goldmark (GFM) and renders through a custom `storageRenderer` registered at priority 100 (below the default HTML=1000 and table=500 renderers) that emits Confluence storage format. `shield.go` renames raw `ac:`/`ri:` tags to colon-free sentinels around the goldmark step so pasted storage passes through; `callouts.go` is an AST transformer + blockquote renderer for GitHub alerts; `mention.go` owns the user-mention mapping in both directions (#91): a mention is 80% of all `` usage, and it converts to `[@Display Name](https://home.atlassian.com/people/{accountId})`. Three things decide its shape, each measured rather than reasoned. **The URL is Atlassian Home, not the site** — Confluence's own renderer still emits `{site}/wiki/people/{id}`, which no longer resolves usefully in a browser, so a mention in Markdown names *no site*, needs nothing from configuration, and is therefore recognisable by `check` with no client at all. **Matching is on the path, ignoring host and query**, because several spellings of one target circulate (the Home URL, the modal's `?cloudId=` copy, the `/o/{orgId}` redirect, both Confluence forms, root-relative) and none of `cloudId`/`ref`/the org segment identifies the person — only the id does, and `ri:user` stores nothing else. **The `@` on the link text is the marker**, load-bearing rather than decoration: the URL cannot tell "mention this person" from "link to their profile", so without it anyone writing the second would silently get the first. The account id is **not** pattern-validated (two shapes are live on one instance, so a pattern tight enough for one rejects the other) and `ri:local-id` is never emitted (a mention carrying only the id resolves to the same person, verified via ADF). `MentionMarkdown` is the shared builder for a mention's whole Markdown line, exported because `user-find` (#143) prints exactly it and a second copy there would be a second place to get the `@` marker and name escaping right. `ConfluencePage.Mentions` reports the ids the *forward* direction emitted so the caller can warn about one that names nobody — the `Attachments` arrangement, and necessary because Confluence accepts any id and renders `@Unlicensed user` rather than failing, and the profile URL 200s either way. An unresolvable mention still renders as a link, `[@Unlicensed user](…)` — that wording mirrors Confluence because the only ids reaching it are the ones the page labels that way: a **deactivated account resolves normally** and keeps its name (measured across every mention on a real page — 18 of them, six departed, all 200, returning e.g. `Mark Reid (Deactivated)`), so a departed colleague never takes that branch. Name resolution is `pagedoc.UserCache`, a per-run cross-page cache, and the tri-state is the part to preserve: `client.LookupUser` separates a name from `ErrNoSuchUser` from an unaskable question, `StorageOptions.UserNames` carries that as name / `""` / absent, and only a **confirmed** absence renders the placeholder. Flattening those would write a fabricated name over a real one the moment a VPN dropped mid-export, across a whole tree, into a file that then looks authoritative — which is also why the cache remembers a 404 but not a timeout (one is an answer, the other is not) and why `MentionWarnings` warns only about a confirmed absence; `aclink.go` is the *inverse* direction's one element with enough shape to need its own file — ``, which the editor writes for every internal link and `MdToConfluence` never emits, so nothing in the regression suite covers it. One rule decides its whole mapping: **convert when the Markdown republishes to a link resolving to the same target, pass the storage through when it would not** — so a page link and a space link convert, while a mention and a space link convert, and an attachment link (only images are uploaded, so a relative href would be dead) and an unresolvable target stay raw, which the shield republishes byte-identical. A page target is a **title, never an id**, so `PageLinkTargets` reports what needs resolving and `StorageOptions.PageLinks` carries the answers back. An `ac:anchor` is **percent-encoded** where `confluenceSlug` output is not: decode it before matching a heading, leave it encoded inside a URL. A same-page anchor recovers its heading from the document rather than inverting the slug, which is impossible — `confluenceSlug` turns both a space and a hyphen into `-`. The survey the mapping rests on, and the `xml.HTMLAutoClose` trap that made `` crash the parser outright (#88), are in [docs/confluence/links-and-anchors.md](docs/confluence/links-and-anchors.md); Plain text used as a Markdown link's text goes through `escapeLinkText` (`\`, `[`, `]`), applied to the *raw* sources only — a page title, a space key, an anchor, a display name — and via `inlineTextForLink` to a body whose every descendant is a text node. Never to already-rendered output: an `ac:link-body` holding markup has been converted to Markdown already, and escaping it yields a literal `\*\*bold\*\*`. Both directions are tested, because a fix at either extreme passes one and fails the other. `attachname.go` owns the source-path→attachment-name mapping, which is now the path's **base name** and nothing else (#59/`_plans/029`): the name is the attachment's identity, so an encoded path moved the name every time the file moved and orphaned the old attachment, and the path is recorded in the comment anyway. The mapping is therefore lossy, and what the bijection used to buy is an explicit refusal — two assets in one document whose base names agree return a typed `NameCollisionError` from `MdToConfluence`, which is a *failure* and not a `Broken` entry, since nothing blocks a publish on `Broken`. `check` catches that error and reports it as `Broken` anyway, because there it is a document defect like a dead link rather than a converter failure. A stored name is never interpreted in the other direction either: `sourceFor` reads the recorded path or uses the name verbatim. What names Confluence accepts is in [docs/confluence/attachments.md](docs/confluence/attachments.md); `destination.go` owns the **other** codec, destination↔path (`decodeDestination`/`encodeDestination`), shared by images *and* doc links — a Markdown destination is a URL, so decode inbound (**before** `withinRoot`, or an encoded `..%2F` slips the clamp) and encode outbound in `storage_to_md.go` (or `export` emits Markdown that no longer parses, and `sourceFor`'s absolute-path refusal is undone by the next read); an undecodable destination is a literal `%` in a filename, not an error; the reasoning is in [docs/confluence/links-and-anchors.md](docs/confluence/links-and-anchors.md); `images.go` (resolution stays page-relative like GitHub; the documentation root — cwd — bounds what may be published, and an image above it is `IMAGE BROKEN`), `links.go` (GitHub/Confluence slugs, doc-link + anchor rewriting against `internal/linkindex`'s tree-wide index; `resolveDocKey` resolves a destination to the index's root-relative key and reports `escapes` — a purely lexical check on the *query* side, since the index itself needs no clamp: an escaping key can never be in it, built by walking downward from root). A doc-link target is one of four severities, #42: missing entirely or escaping root is **Broken** (`LINK BROKEN: … (not found|outside the documentation root)`) and replaces the whole `` element — tags and visible text alike — with that literal message, matching `images.go`'s precedent for a missing image (`renderLink` needs a small per-node flag, `linkBrokenText`, since goldmark still invokes a container node's renderer on the matching leaving call regardless of `WalkSkipChildren` on entering, and there is no `` to write in the broken case); existing on disk with no `page_id` yet is unchanged — a **warning**, the normal state of an unpublished tree; a `#fragment` matching no heading on an otherwise-resolving target also **warns**, gated on `linkindex.Index.FileExists` so a missing/escaping target isn't double-reported. `tables.go` (the `` tag, stamped with `data-layout="align-start"` so tables auto-size and left-align — this must stay if column widths are ever emitted, or a `` silently induces a layout; plus cells: an AST transformer consumes a leading `` comment in a cell and `renderTableCell` emits it as `data-highlight-colour`; `storage_to_md_table.go`'s `cellTexts` reverses this, reading `data-highlight-colour` back into a `bg:` marker (`cellBGNames`, the reverse of `tables.go`'s swatch map — a hex outside the 21 swatches round-trips as the literal hex, and where two names share a hex the British spelling wins, matching Confluence's own `-colour`), and a column's GFM alignment becomes a `

` wrapper **inside** the cell — never the `align` attribute the GFM renderer would emit, which is the one form Confluence discards. Only center and right are emitted: Confluence has no explicit left, so `:---` publishes bare and `read` recovers it as `---`. Rows still fall through to the GFM renderer. **The way back writes a pipe table only when Markdown expresses the whole table** (#55, `_plans/055`), and otherwise the table as raw storage through `renderRawBlock`, with `

`, a span of 1, cell children that are paragraphs, lists or inline markup, a `text-align` only of a known value, and every non-empty cell in a column agreeing on alignment. **What a line's alignment is was measured, not assumed** ([storage-format.md](docs/confluence/storage-format.md#a-paragraphs-alignment-and-its-cells)): only `center` and `right` do anything, and a paragraph saying one overrides its cell's; `left` and `justify` do nothing, and `start` and `end` are stripped by Confluence (a cell's too), so a paragraph saying any of those four is aligned by its cell like one saying nothing — `end` is ADF's *name* for right, not a storage value that means it. Dropping those four on the way back is what keeps the Markdown a fixed point, since `start`/`end` written back would be stripped and read differently — since alignment is per-paragraph in Confluence and per-column in GFM, a column that disagrees could not be written without republishing cells with an alignment they lacked; this replaced a majority vote that did exactly that. Some markup is **ignored** rather than honoured because the browser editor writes it on every table it saves (measured on a markfluence table saved in the browser, [storage-format.md](docs/confluence/storage-format.md#what-the-editor-writes-on-a-table)): `ac:local-id` and the bare `local-id` it puts on every paragraph, `data-table-width` whatever its value, `data-table-display-mode="default"`, `data-layout="align-start"`, and — on an `align-start` table only — a `` of pixel widths, which the editor measures and writes on any save. Honouring them would turn every markfluence table someone edits in Confluence into HTML on the next `read`. The known cost: a hand resize writes the same `data-table-width` and ``, so it is lost when a read-back table republishes. A `` on any other layout keeps the table raw, since `default`/`full-width` are visible layouts markfluence never writes. In a raw cell, `rawCellBlocks` groups loose inline content into one paragraph, keeps an empty paragraph as `

` and one that renders to nothing (a `
`, a `

`; `renderRawBlock` writes any element holding loose text whole (`hasLooseText`), since one line per child would add whitespace to it. The bare `local-id` is in `droppedAttrs` beside `ac:local-id`. A list in an aligned column keeps a table raw: the pipe cell would publish it inside the aligned `

`. So does a list holding a `|` anywhere (unescaped it splits the row; escaped, the backslash survives into an `href`), a paragraph of nothing but `
`s, and an empty paragraph at either end of a cell. `renderCellLines` puts no `
` beside a list, which ends its line by being a block. In `cellAlign`, loose inline content has the cell's alignment (measured too) and an empty paragraph has none to disagree with. A multi-line cell is one `

` per line — Enter in the editor starts a new `

`, it does not insert a `
` — so `renderCellLines` in `storage_to_md.go` joins sibling `

` children with a literal `
` rather than nothing: a GFM table row is exactly one physical line, so a real newline isn't an option, and the same substitution catches a bare mid-line `
` (Shift+Enter) that would otherwise render as the two-space hard break valid in ordinary block content but not inside a table row. A `

`/`` as content containers so each cell's body stays Markdown between blank lines (a Markdown paragraph has no alignment, so `normalizeCellAlign` drops every declaration that does nothing and moves a cell's shared alignment onto the cell as `style="text-align: center|right;"`; an aligned `

` stays storage only when a cell's lines disagree — storage would also keep an image in it as ``, which is never uploaded to a new page). `tableShape` decides, over an **allowlist**, so an attribute nobody has seen keeps a table raw rather than being dropped: one header row of `

` and nothing but `` below it (GFM has no headerless table, and promoting the first row turned ``s into ``s on the next publish), every row as wide as the header, no `
` tag, stamped with `data-layout="align-start"` so tables auto-size and left-align — this must stay if column widths are ever emitted, or a `` silently induces a layout; plus cells: an AST transformer consumes a leading `` comment in a cell and `renderTableCell` emits it as `data-highlight-colour`; `storage_to_md_table.go`'s `cellTexts` reverses this, reading `data-highlight-colour` back into a `bg:` marker (`cellBGNames`, the reverse of `tables.go`'s swatch map — a hex outside the 21 swatches round-trips as the literal hex, and where two names share a hex the British spelling wins, matching Confluence's own `-colour`), and a column's GFM alignment becomes a `

` wrapper **inside** the cell — never the `align` attribute the GFM renderer would emit, which is the one form Confluence discards. Only center and right are emitted: Confluence has no explicit left, so `:---` publishes bare and `read` recovers it as `---`. Rows still fall through to the GFM renderer. **The way back writes a pipe table only when Markdown expresses the whole table** (#55, `_plans/055`), and otherwise the table as raw storage through `renderRawBlock`, with `

`, a span of 1, cell children that are paragraphs, lists or inline markup, a `text-align` only of a known value, and every non-empty cell in a column agreeing on alignment. **What a line's alignment is was measured, not assumed** ([storage-format.md](docs/confluence/storage-format.md#a-paragraphs-alignment-and-its-cells)): only `center` and `right` do anything, and a paragraph saying one overrides its cell's; `left` and `justify` do nothing, and `start` and `end` are stripped by Confluence (a cell's too), so a paragraph saying any of those four is aligned by its cell like one saying nothing — `end` is ADF's *name* for right, not a storage value that means it. Dropping those four on the way back is what keeps the Markdown a fixed point, since `start`/`end` written back would be stripped and read differently — since alignment is per-paragraph in Confluence and per-column in GFM, a column that disagrees could not be written without republishing cells with an alignment they lacked; this replaced a majority vote that did exactly that. Some markup is **ignored** rather than honoured because the browser editor writes it on every table it saves (measured on a markfluence table saved in the browser, [storage-format.md](docs/confluence/storage-format.md#what-the-editor-writes-on-a-table)): `ac:local-id` and the bare `local-id` it puts on every paragraph, `data-table-width` whatever its value, `data-table-display-mode="default"`, `data-layout="align-start"`, and — on an `align-start` table only — a `` of pixel widths, which the editor measures and writes on any save. Honouring them would turn every markfluence table someone edits in Confluence into HTML on the next `read`. The known cost: a hand resize writes the same `data-table-width` and ``, so it is lost when a read-back table republishes. A `` on any other layout keeps the table raw, since `default`/`full-width` are visible layouts markfluence never writes. In a raw cell, `rawCellBlocks` groups loose inline content into one paragraph, keeps an empty paragraph as `

` and one that renders to nothing (a `
`, a `

`; `renderRawBlock` writes any element holding loose text whole (`hasLooseText`), since one line per child would add whitespace to it. The bare `local-id` is in `droppedAttrs` beside `ac:local-id`. A list in an aligned column keeps a table raw: the pipe cell would publish it inside the aligned `

`. So does a list holding a `|` anywhere (unescaped it splits the row; escaped, the backslash survives into an `href`), a paragraph of nothing but `
`s, and an empty paragraph at either end of a cell. `renderCellLines` puts no `
` beside a list, which ends its line by being a block. In `cellAlign`, loose inline content has the cell's alignment (measured too) and an empty paragraph has none to disagree with. A multi-line cell is one `

` per line — Enter in the editor starts a new `

`, it does not insert a `
` — so `renderCellLines` in `storage_to_md.go` joins sibling `

` children with a literal `
` rather than nothing: a GFM table row is exactly one physical line, so a real newline isn't an option, and the same substitution catches a bare mid-line `
` (Shift+Enter) that would otherwise render as the two-space hard break valid in ordinary block content but not inside a table row. A `

`/`` as content containers so each cell's body stays Markdown between blank lines (a Markdown paragraph has no alignment, so `normalizeCellAlign` drops every declaration that does nothing and moves a cell's shared alignment onto the cell as `style="text-align: center|right;"`; an aligned `

` stays storage only when a cell's lines disagree — storage would also keep an image in it as ``, which is never uploaded to a new page). `tableShape` decides, over an **allowlist**, so an attribute nobody has seen keeps a table raw rather than being dropped: one header row of `

` and nothing but `` below it (GFM has no headerless table, and promoting the first row turned ``s into ``s on the next publish), every row as wide as the header, no `
Pipe cell
*x* and a|b
+

Raw cell: *x*

+

Layout cell: *x*

+

1. not a list

+

# not a heading, and a line after a break:
# not a heading either

+

> 90 days

+

- not a list in a callout

+

1. not a list in a raw cell

> loose text in a raw cell
+

+ not a list in a layout cell

+

Item #

+

Plain text, not links: https://example.com, www.example.com and ops@example.com.

+

A status macro: _x_ *y* stays raw.

+
A list in a pipe cell
  • _x_ and [y]
+

A backslash before a mark's moved space: a \ x, and two runs of one mark: \ down.

+

---

+

#

diff --git a/internal/convert/testdata/storage2md/escaping/output.md b/internal/convert/testdata/storage2md/escaping/output.md new file mode 100644 index 0000000..d26ab49 --- /dev/null +++ b/internal/convert/testdata/storage2md/escaping/output.md @@ -0,0 +1,99 @@ +Emphasis: \*not emphasis\*, 2\*3\*4, a \_b\_ c, and a * b * c stays. + +Strikethrough: a ~b\~ c, a \~\~b\~\~ c, and about ~5 min stays. + +Code: \`not code\`. + +HTML: \not bold\, \, \not storage\, and a < b stays. + +Entities: \© and \©, and AT&T stays. + +Brackets: \[x], \[x]\[y], and x] stays. + +Backslashes: a\\\*b, and C:\path stays. + +A link: [Q1 \*Draft\*\]](https://example.com), and one with markup: [**bold** \_text\_](https://example.com). + +- In a list item: \*x\* and \[y] + +> [!NOTE] +> In a callout: \*x\* + +| Pipe cell | +| --- | +| \*x\* and a\|b | + + + + + + + +
+ +Raw cell: \*x\* + +
+ + + + + +Layout cell: \*x\* + + + + + +1\. not a list + +\# not a heading, and a line after a break: +\# not a heading either + +\> 90 days + +> [!WARNING] +> \- not a list in a callout + + + + + + + + +
+ +1\. not a list in a raw cell + + + +\> loose text in a raw cell + +
+ + + + + +\+ not a list in a layout cell + + + + + +## Item \# + +Plain text, not links: https\://example.com, www\.example.com and ops\@example.com. + +A status macro: \_x\_ \*y\* stays raw. + +| A list in a pipe cell | +| --- | +|
  • \_x\_ and \[y]
| + +A backslash before a mark's moved space: a \\ *x*, and two runs of one mark: *\ down*. + +**---** + +### \# diff --git a/internal/linkindex/linkindex.go b/internal/linkindex/linkindex.go index 1fc1e25..b463c5f 100644 --- a/internal/linkindex/linkindex.go +++ b/internal/linkindex/linkindex.go @@ -184,9 +184,16 @@ func ConfluenceSlug(heading string) string { return whitespaceRunRE.ReplaceAllString(strings.TrimSpace(heading), "-") } +// backslashEscapeRE matches a CommonMark backslash escape: a backslash before +// an ASCII punctuation character, which the group captures. +var backslashEscapeRE = regexp.MustCompile("\\\\([!-/:-@\\[-`{-~])") + // extractHeadings returns the text of each ATX heading in a // frontmatter-stripped body, skipping fenced code blocks so "#" lines inside -// samples aren't headings. +// samples aren't headings. Backslash escapes are removed, since Confluence +// builds a heading's anchor from its published text, where they are gone: +// read writes "## Setup \[beta]" for a heading whose anchor is "Setup-[beta]" +// (#203). func extractHeadings(body string) []string { var headings []string inCode := false @@ -210,7 +217,7 @@ func extractHeadings(body string) []string { continue } if text := strings.TrimSpace(rest); text != "" { - headings = append(headings, text) + headings = append(headings, backslashEscapeRE.ReplaceAllString(text, "$1")) } } return headings diff --git a/internal/linkindex/slug_test.go b/internal/linkindex/slug_test.go index fc5e78f..1ea4e21 100644 --- a/internal/linkindex/slug_test.go +++ b/internal/linkindex/slug_test.go @@ -67,6 +67,21 @@ func TestExtractHeadings(t *testing.T) { } } +// TestExtractHeadingsRemovesEscapes: an anchor is built from a heading's +// published text, which has no backslash escapes, and read writes them (#203). +func TestExtractHeadingsRemovesEscapes(t *testing.T) { + got := extractHeadings("## Setup \\[beta]\n## Item \\#\n## C:\\path\n## a \\\\ b\n") + want := []string{"Setup [beta]", "Item #", "C:\\path", "a \\ b"} + if len(got) != len(want) { + t.Fatalf("got %q, want %q", got, want) + } + for i := range want { + if got[i] != want[i] { + t.Errorf("heading %d = %q, want %q", i, got[i], want[i]) + } + } +} + func TestExtractHeadingsUnterminatedFenceSkipsToEnd(t *testing.T) { // A fence that's never closed must not leave later real headings exposed // by some off-by-one toggle; everything after the open fence is "in code."