From e8540bdd8bd9669b4abc4ea4bc28efbdacb4a3b2 Mon Sep 17 00:00:00 2001 From: Oskar Hagberg Date: Sat, 12 Sep 2026 18:45:44 +0200 Subject: [PATCH] =?UTF-8?q?feat(hosting):=20S30=20=E2=80=94=20export=20art?= =?UTF-8?q?efact=20HTML=20(GUI=20download=20+=20MCP=20read-back)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Artefactor could take HTML in but never give it back. The client had no download affordance, and an agent on the connector could create/update an artefact but never *read* one — so it could neither derive a new artefact from an existing one nor safely update one in place after losing the original from context. The latter is a correctness problem: `update_artefact` replaces the HTML and leaves every per-user data blob untouched, and the backend treats blobs as opaque (AD8), so it cannot migrate them. Export endpoint — `GET /api/artefacts/:ref/download` (slug or id). Resolution and access reuse `resolveViewableArtefact`, so effective-tier resolution through a collection root (AH20) and the `AccessPolicy` cell (AH18) are inherited rather than re-implemented. The body is `PayloadStore.get()` **verbatim** — never the injected render: no S13 bootstrap, no S12 shell — so download → edit → re-upload round-trips. Unknown ref, not-viewable, and archived are a flat 404, with no owner carve-out for archived (AH7 keeps it inert; restore → download → re-archive is one click). New pure filename helper in `src/shared/` emits both `filename` and the RFC 5987 `filename*`. MCP read-back — `get_artefact_html` and `get_artefact_data`, both owner-scoped via `loadOwnActiveArtefact`. The data tool returns the caller's OWN entry only (AD2/AD4), verbatim through `getOwnDataEntry`, with no server-side summarising or key/type digest: passing bytes through is transport, describing their structure would be the backend interpreting the blob. Each hard-errors above its context cap (~1 MB HTML, 256 KB blob) naming the real size and pointing at the GUI download — no truncation, since truncated HTML can't be edited and truncated JSON can't be parsed. Declared data schema — an artefact may state its own data shape in an inert ` +``` + +This is a **payload convention**, not a domain rule. The backend **may forward** the block (the +snapshot read returns it as parsed JSON, or `null` when it is absent, malformed, or not a JSON +object — never an error) but **never interprets or enforces it**: a blob is never validated +against a declared schema, because enforcing a schema would make the backend interpret the blob +and collapse AD8. An unknown script `type` is ignored by browsers, so the block is inert, and +because it lives in the trusted HTML it travels with an export automatically (see the export +read path in `artefact-hosting.md`). + +It is best-effort by nature — a *second* representation of a truth that actually lives in the +JS, with nothing enforcing agreement — so a stale declaration produces a **confident** wrong +migration, which is worse than inference (inference fails visibly). The authoring rule is +therefore: schema and code are written in the same breath; a reader trusts the schema for +orientation and **verifies it against the HTML before any shape-changing write**. The block also +says nothing about the population, so the "tolerate shapes you never saw" rule above stands +regardless. + +**Two version notions, reconciled.** The schema's `version` and AD9's `authoredAgainstVersion` +answer different questions and must not be conflated: + +| | Says | Set by | +|--|------|--------| +| `authoredAgainstVersion` (AD9) | this blob was written against payload hash X | the **backend**, on every write | +| schema `version` | this HTML expects shape v2 | the **authoring agent**, by hand | + +They are complementary: the pin is *mechanical* and per-entry (is this user's data stale?), the +declared version is *semantic* and per-payload (what shape does this HTML expect?). Keep both. ## Artefact runtime contract @@ -159,6 +215,9 @@ read-only (AD5). - **Cross-user viewing is a host feature**, not an artefact capability: a host user-picker loads another author's data **read-only** by re-seeding the artefact. The artefact never knows whose data it holds. +- **A snapshot may be read, never interpreted (S30).** The connector can return the caller's + own blob verbatim and forward the artefact's declared schema block, but the backend neither + summarises the blob nor validates it against that schema — AD8 is unchanged. ## Amendment (post-v0.2) — payload version pin @@ -185,7 +244,16 @@ host can detect the mismatch. **Use.** - **OSS:** even with a single mutable payload, the host can tell whether a viewer's saved data **predates the current payload** (pin ≠ current hash) — a sharper form of the `dataAuthorCount` - breaking-change signal already exposed to the MCP connector. + breaking-change signal already exposed to the MCP connector. **S30 reserves this field + without depending on it**: the snapshot read already returns the pair + (`currentPayloadVersion`, `authoredAgainstVersion`), with the pin `null` until this slice + lands. That is sound precisely because AD9 is advisory and gates nothing — when S19 ships, + the pin populates with no change to the tool's shape and no rewrite of the doctrine written + against it (`null` ⇒ unknown, treat as possibly stale; ≠ current ⇒ written against older + HTML, migration owed; = current ⇒ matches what is deployed). + + Do not conflate this pin with the **declared schema's** `version` (see "Declared data schema" + above): this one is mechanical and backend-set, that one semantic and author-set. - **Superset (history / rollback):** before serving an old or rolled-back version, compare the seeded entry's pin to the target version to decide whether the data is compatible (seed it, or warn / seed cautiously). This is the **pin-for-compatibility** decision — there is deliberately diff --git a/docs/specs/ddd/artefact-hosting.md b/docs/specs/ddd/artefact-hosting.md index b65e2fa..48a968d 100644 --- a/docs/specs/ddd/artefact-hosting.md +++ b/docs/specs/ddd/artefact-hosting.md @@ -134,6 +134,17 @@ States are the product of `visibility × status`. Allowed transitions: concern, not a security boundary. - Because artefacts are served **same-origin**, their JS can call the backend store (see `artefact-data.md`) carrying the viewer's session — this is how forms persist data. +- **Export (S30).** Alongside the render there is a second read path: the **export**, which + returns the **stored payload** — never the injected render. It is governed by the same + matrix as the render (AH7/AH8/AH9, on the *effective* tier per AH20): unknown handle, + not-viewable, and **archived** all surface as a flat 404, with **no owner carve-out** for an + archived artefact (restore → export → re-archive is the escape hatch; AH7 keeps archived + inert). The distinction from the render is the point: the render injects the localStorage + bootstrap and wraps the artefact in the host shell, whereas the export is byte-identical to + what was uploaded, so export → edit → re-upload round-trips through the same create/edit + commands and yields the same artefact. A payload convention carried *inside* the HTML — such + as the declared data schema in `artefact-data.md` — therefore travels with the export for + free, with no extra plumbing. ## Relationship to Artefact Data diff --git a/docs/specs/fdd/slice-dag.md b/docs/specs/fdd/slice-dag.md index ce50ab6..2ecea11 100644 --- a/docs/specs/fdd/slice-dag.md +++ b/docs/specs/fdd/slice-dag.md @@ -27,6 +27,12 @@ S1 Identity (BetterAuth — email+password for dev; Google OAuth added later) │ └────────────► S18 MCP connector (remote MCP server + OAuth via BetterAuth `mcp` plugin; wraps the Hosting commands as tools — needs S2, S3, S4, S5, S7, S10) + │ + └──► S30 Export artefact HTML (GUI download + MCP read-back + tools — needs S2, S4, S6, S11, S18) + ┊ + ┊ (optional sharpener, NOT a dependency) + ┄┄┄ S19 data version pin (AD9) ~~S8 Issue / revoke API key~~ and ~~S9 API push ingestion~~ are **dropped** — the pinned better-auth has no api-key plugin and a raw token API was deemed unnecessary; programmatic @@ -785,6 +791,75 @@ Deps: **S28.** (CL1 relaxed; CL12/CL13/CL14.) not re-attach evicted artefacts. - **Boundary:** **OSS**. No schema change (contributors ride `collection_access`). +### S30 — Export artefact HTML (GUI download + MCP read-back tools) +Deps: **S2, S4, S6, S11, S18.** (AH7/AH8/AH9 for the export read path; AD2/AD4/AD8 for the +data snapshot.) Artefactor could take HTML in but never give it back: the client had no +download affordance, and an agent on the connector could `create`/`update` an artefact but +never *read* one — so it could neither derive a new artefact from an existing one (A) nor +safely update one in place after losing the original from context (B). (B) is a correctness +problem: `update_artefact` replaces the HTML and leaves every per-user data blob untouched, +and the backend treats blobs as opaque (AD8), so it cannot migrate them. + +- **Domain** — `extractDeclaredSchema(html)`: a pure, best-effort lift of the artefact's + **declared data schema** block (below) from its trusted HTML. Absent/malformed/non-object + → `null`, never an error. A payload *convention*, never enforced. +- **Shared** — `artefactFilename(title, fallback)` + `attachmentDisposition(filename)`: pure + download-naming helpers (slugify, bound, ASCII-fold, RFC 5987), unit-tested on their own. +- **BFF** — `GET /api/artefacts/:ref/download` (`:ref` = slug **or** id). Resolution and + access reuse `resolveViewableArtefact`, so effective-tier resolution through a collection + root (AH20) and the `AccessPolicy` cell (AH18) are **inherited, not re-implemented**. Body + = `PayloadStore.get(payloadRef)` **verbatim** — no S13 bootstrap, no S12 host shell, so the + download round-trips (download → edit → re-upload / `update_artefact` yields the same + artefact). `Content-Type: text/html; charset=UTF-8`; `Content-Disposition: attachment` with + both `filename` and `filename*`; `Content-Length` = `payloadBytes`. Unknown ref / + not-viewable / **archived** → flat 404. +- **MCP** — `get_artefact_html { id }` → `{ id, title, kind, html, dataAuthorCount }`; + `get_artefact_data { id }` → `{ id, blob, bytes, updatedAt, dataAuthorCount, schema, + currentPayloadVersion, authoredAgainstVersion }` — the caller's **own** entry only, verbatim + via `getOwnDataEntry`, with no server-side summarising or key/type digest (that would be the + backend interpreting the blob). Both owner-scoped via `loadOwnActiveArtefact`. Each + hard-errors above its cap (`MAX_MCP_HTML_BYTES` ≈ 1 MB, `MAX_MCP_BLOB_BYTES` ≈ 256 KB), + naming the real size and pointing at the GUI download — **no truncation**: truncated HTML is + unusable for editing and truncated JSON unparseable, and either invites the agent to act on + a fragment as though it were whole. +- **Doctrine** — `skills/artefactor/SKILL.md` + `PERSISTENCE_CONTRACT_SUMMARY`: **migrate + forward** now leads the breaking-change guidance (read old key → transform → write new; + idempotent, at load before first render, old key kept one generation, never `clear()`), the + snapshot is one blob and not the population, and the schema block joins the template + + checklist. Bumping the key without migrating is named for what it is — silent data loss. +- **Client** — "Download HTML" in `MoreMenu` (reaching `ArtefactRow` + `ArtefactCard`), owned + artefacts only. A plain anchor: session-cookie auth needs no fetch/blob dance. +- **Acceptance:** owner downloads own artefact by id at any visibility, byte-identical to what + was uploaded; viewer downloads a shared artefact by slug; anonymous downloads a `public` one; + private-to-a-non-owner, **archived (including for the owner)**, and unknown ref all 404; both + `Content-Disposition` forms present and `Content-Length` = `payloadBytes`; no bootstrap or + shell in the body; filename helper covers ASCII / non-ASCII (åäö) / symbol-only (falls back + to slug-or-id) / 300-char (bounded) / always `.html`. `get_artefact_html` returns the exact + stored HTML with `dataAuthorCount`; `get_artefact_data` returns the caller's own blob + verbatim, `blob: null` when they have none, never another author's; over-cap on either → + error naming the actual size; non-owner / unknown / archived / out-of-scope → not found; + `schema` is parsed JSON when present and `null` when absent, malformed, or not valid JSON — + **never** an error; a blob is never validated against a declared schema (AD8 holds); an + artefact with a declared schema survives export → re-upload intact; `currentPayloadVersion` + equals `payloadHash` and `authoredAgainstVersion` is `null` while S19 is unbuilt (both + fields' presence and shape asserted). +- **Archived stays inert (AH7)** — no owner carve-out. Restore → download → re-archive is one + click, which is not worth an exception in AH7 for an escape hatch. +- **On S19/AD9 — reserve, don't depend.** `get_artefact_data` returns the version-pin **pair** + but S19 is **not** a dependency edge, and the missing edge is deliberate, not an oversight: + AD9 is *advisory by spec* and never gates a read or write, so nothing here is incorrect while + the pin is `null`; and S19 also carries the unrelated AH15 `PayloadRetentionPolicy` port in + the edit command, which read-back has no business pulling in. `currentPayloadVersion` is free + today (`payloadHash` is already on the aggregate). When S19 lands, the pin populates with + **no tool-shape change and no doctrine rewrite** — the rule "pin present and ≠ current ⇒ that + user's data predates this payload" is written now and becomes true then. +- **Out of scope:** the download affordance for "shared with you" (`GalleryCard`/`GalleryRow`) + and the `/a/:slug` shell toolbar (the endpoint already honours the matrix — widening is + client-only); baking a data snapshot into the downloaded file; any data **write** tool (S31); + a per-artefact "allow download" toggle (new field + invariant + migration, and defeated by + view-source anyway). +- **Boundary:** **OSS**. No schema change. + ## Build order Topological: **S0 → S1 → S2 → {S3, S4, S5, S7, S10, S11}**, **S5 → {S6, S14, S16}**, @@ -804,4 +879,7 @@ dependency of the EE **Postgres persistence** context. **S25** (collections) dep hosting core + sharing (**S2/S5/S6/S10/S16**); **S26** (collection lifecycle) on **S25 + S15**; **S27** (bookmarks) on **S25**; **S28** (viewer-facing shared collections) on **S25** and **S29** (contributors + evict-on-cascade) on **S28**. All five are OSS feature slices -(context `ddd/artefact-collections.md`). +(context `ddd/artefact-collections.md`). **S30** (export HTML) depends on the hosting read +path + the data store + the connector (**S2/S4/S6/S11/S18**); **S19** is an *optional +sharpener* of S30's staleness signal, **not** a dependency edge (see the slice's +reserve-don't-depend note). diff --git a/skills/artefactor/SKILL.md b/skills/artefactor/SKILL.md index 7375e4f..40c212d 100644 --- a/skills/artefactor/SKILL.md +++ b/skills/artefactor/SKILL.md @@ -31,12 +31,27 @@ the user** (OAuth), so everything you create is owned by them. Tools: - **`list_artefacts`** / **`get_artefact`** — find what the user already has (use these before creating a duplicate; update in place when iterating). `get_artefact` returns **`dataAuthorCount`** — how many users have saved data in this artefact. +- **`get_artefact_html`** `{ id }` — the artefact's **stored HTML**, exactly as served. Use it + to derive a new artefact from an existing one, or to re-read an artefact before updating it + when you no longer have the source (`update_artefact` replaces the HTML wholesale, so you + need the current document to change it safely). +- **`get_artefact_data`** `{ id }` — **your own** saved data for the artefact, verbatim, plus + the shape the artefact **declares** for itself. Read it before any change to the data shape. + Returns `blob` (your entry, `null` if you have none), `bytes`, `updatedAt`, `schema` (the + declared block below, or `null`), `dataAuthorCount`, and the version pin pair + `currentPayloadVersion` / `authoredAgainstVersion`. - **`set_visibility`** / **`archive_artefact`** / **`restore_artefact`** — manage sharing and lifecycle. - **`get_authoring_guide`** — returns this guide. If you're working through the connector without this skill loaded (e.g. in Claude design), call it before writing artefact HTML to get the persistence contract, template, and checklist below. +Both read-back tools **refuse** a result too large for a tool call (roughly 1 MB of HTML, +256 KB of data) rather than truncating it — truncated HTML can't be edited and truncated JSON +can't be parsed. When that happens, the user can get the file from the Artefactor web app: +**"Download HTML"** in an artefact's ⋯ menu returns the stored document byte-for-byte, so +download → edit → re-upload round-trips. + **Typical flow:** write the HTML following the persistence rules below → `create_artefact` → share via `set_visibility` (or by passing `visibility`). When the user says "update the X artefact", prefer `list_artefacts`/`get_artefact` + `update_artefact` over creating a new one. @@ -79,17 +94,75 @@ keeps it **opaque** — the backend never reads or rewrites it. You shape and se won't (and can't) migrate them; the data is opaque to it. So if your new HTML expects a **different data shape** than the old one, returning users' saved data may be misread. -Before a shape-changing update, check `dataAuthorCount`. If it's `> 0` and the change is -breaking, do one of: - -- **Bump the storage-key version** in the HTML (`my-artefact-v1` → `my-artefact-v2`). Old data - is simply ignored and the artefact starts fresh — the localStorage-native migration (rule 2). +Before a shape-changing update: check `dataAuthorCount`, and call **`get_artefact_data`** to +see the shape actually saved. Then pick one of three, in this order: + +- **Migrate forward (preferred).** Ship migration code *in the new HTML*: read the old key, + transform it, write the new one. This is the only place a migration **can** live — a + server-side migration would mean parsing the blob, which the backend never does. Each user's + data migrates on their next visit. Write it so it is: + - **idempotent** — safe to run on every load, including when it already ran; + - **run at load, before first render**, so nothing renders against the old shape; + - **non-destructive** — leave the old key in place for one generation, and never `clear()`. +- **Bump the storage-key version** (`my-artefact-v1` → `my-artefact-v2`) **without** migrating. + Old data is simply ignored — which means every user **silently loses** what they saved. + Only do this when you know the saved data is disposable, and say so to the user. - **Publish a new artefact** (`create_artefact`) — a clean "v2" with its own id, link, and data — when you want to keep the old one intact for existing users. Non-breaking edits (copy, styling, bug fixes, additive fields your code already tolerates) are safe to `update_artefact` in place. +**The snapshot is one blob, not the population.** `get_artefact_data` returns the data of the +user you are acting for; `dataAuthorCount` tells you how many *other* people also hold data, +and you never see theirs. Their blobs may sit on older key versions, be partial, or have been +written by HTML two revisions back. So write the migration to tolerate shapes you never saw — +absent, partial, or of an unknown version — and never assume the one blob you read is +representative. + +**The pin is the per-user staleness signal.** `dataAuthorCount` says *some* users hold data; +comparing `authoredAgainstVersion` with `currentPayloadVersion` says whether *this* entry +predates the live payload: + +- `null` ⇒ unknown — treat as possibly stale; +- ≠ `currentPayloadVersion` ⇒ definitely written against older HTML, migration owed; +- = `currentPayloadVersion` ⇒ matches what is currently deployed. + +(The pin is backend-set and mechanical — "written against payload X". The `version` in the +declared schema below is author-set and semantic — "this HTML expects shape v2". They answer +different questions; keep both.) + +### Declaring your data shape + +Declare the shape your artefact saves, in an inert block in the HTML, so the shape is knowable +without reading all the code and without anyone's data being populated: + +```html + +``` + +The browser ignores an unknown script `type`, so the block does nothing at runtime, and it +lives in the HTML — so it travels with a download/re-upload automatically. `get_artefact_data` +returns it as `schema`. + +- **The `example` is the centrepiece.** Someone facing an empty blob needs a populated + exemplar, not a type list. Use **invented values only** — never real user data: the block + ships inside an artefact that may be downloaded or made `public`. One or two items, a couple + of KB at most. +- **`version` is what makes migrate-forward checkable** — compare the deployed artefact's + declared version with the one you are about to publish to know whether a migration is owed. +- **Write the schema and the code in the same breath.** Nothing enforces that they agree, and a + stale schema produces a *confident* wrong migration — worse than inference, which at least + fails visibly. So: trust the schema for orientation, and **verify it against the HTML before + any shape-changing write**. It also says nothing about the population — users on older + versions are still out there, so the "tolerate shapes you never saw" rule stands regardless. + ## Persisting data (localStorage) When Artefactor serves your artefact, it **replaces `window.localStorage`** with a shim backed @@ -151,15 +224,40 @@ manage who it belongs to. ### Recommended template ```html + hi`; + + it("returns the declared schema as parsed JSON", () => { + const html = block( + '{"key":"habit-tracker-v2","version":2,"example":{"habits":[]}}', + ); + expect(extractDeclaredSchema(html)).toEqual({ + key: "habit-tracker-v2", + version: 2, + example: { habits: [] }, + }); + }); + + it("tolerates attribute order, extra attributes, and whitespace", () => { + const html = + ''; + expect(extractDeclaredSchema(html)).toEqual({ key: "k" }); + }); + + it("accepts single-quoted and unquoted type attributes", () => { + expect( + extractDeclaredSchema( + "", + ), + ).toEqual({ key: "k" }); + expect( + extractDeclaredSchema( + '', + ), + ).toEqual({ key: "k" }); + }); + + it("returns null when there is no declaration", () => { + expect(extractDeclaredSchema("nothing here")).toBeNull(); + }); + + it("returns null when the block is not valid JSON", () => { + expect(extractDeclaredSchema(block("{ key: 'habit', }"))).toBeNull(); + }); + + it("returns null when the block is empty", () => { + expect(extractDeclaredSchema(block(" "))).toBeNull(); + }); + + it("returns null when the declaration is not a JSON object", () => { + expect(extractDeclaredSchema(block("[1,2,3]"))).toBeNull(); + expect(extractDeclaredSchema(block('"just a string"'))).toBeNull(); + expect(extractDeclaredSchema(block("null"))).toBeNull(); + }); + + it("ignores other script types, including a near-miss type", () => { + expect( + extractDeclaredSchema(''), + ).toBeNull(); + expect( + extractDeclaredSchema( + '', + ), + ).toBeNull(); + }); + + it("takes the first declaration when an artefact carries more than one", () => { + const html = block('{"key":"first"}') + block('{"key":"second"}'); + expect(extractDeclaredSchema(html)).toEqual({ key: "first" }); + }); +}); diff --git a/src/domain/data/declared-schema.ts b/src/domain/data/declared-schema.ts new file mode 100644 index 0000000..d617301 --- /dev/null +++ b/src/domain/data/declared-schema.ts @@ -0,0 +1,54 @@ +// S30 — the **declared data schema**: an authoring convention by which an +// artefact states its own data shape inside its HTML, so the shape is knowable +// without inferring it from the code and without a populated blob. +// +// +// +// The browser ignores an unknown script `type`, so the block is inert, and it +// lives in the trusted HTML — so it travels with an export automatically +// (download → edit → re-upload keeps it, no extra plumbing). +// +// This is a **payload convention**, not a domain rule: the backend may forward +// the declaration but never validates a blob against it (that would make the +// backend interpret the blob and collapse AD8 opacity). Extraction is therefore +// best-effort — absent, malformed, or non-object → `null`, never an error, and +// the caller falls back to reading the HTML itself. It is also only a *second* +// representation of a truth that lives in the JS, so a reader trusts it for +// orientation and verifies against the HTML before any shape-changing write. + +export const DECLARED_SCHEMA_TYPE = "application/artefactor-schema+json"; + +// The opening `"); + const body = await (await download(a.id, owner)).text(); + expect(body).toBe(""); + expect(body).not.toContain("data/me"); + expect(body).not.toContain(" { + const a = await create(owner); + const res = await download(a.id, owner); + expect(res.headers.get("content-type")).toBe("text/html; charset=UTF-8"); + const disposition = res.headers.get("content-disposition")!; + expect(disposition).toContain('filename="tradgard-report.html"'); + expect(disposition).toContain("filename*=UTF-8''tr%C3%A4dg%C3%A5rd-report.html"); + }); + + it("sets Content-Length to the artefact's payload size", async () => { + const a = await create(owner); + const res = await download(a.id, owner); + expect(res.headers.get("content-length")).toBe(String(a.payloadBytes)); + }); + + it("lets a signed-in viewer download an artefact shared with them by slug", async () => { + const a = await create(owner); + const shared = await share(owner, a.id, "authenticated"); + const res = await download(shared.publicSlug!, other); + expect(res.status).toBe(200); + expect(await res.text()).toBe(HTML); + }); + + it("lets an anonymous caller download a public artefact by slug", async () => { + const a = await create(owner); + const shared = await share(owner, a.id, "public"); + const res = await download(shared.publicSlug!, null); + expect(res.status).toBe(200); + expect(await res.text()).toBe(HTML); + }); + + it("404s a private artefact for a non-owner", async () => { + const a = await create(owner); + expect((await download(a.id, other)).status).toBe(404); + expect((await download(a.id, null)).status).toBe(404); + }); + + it("404s an unshared artefact's id for an anonymous caller", async () => { + const a = await create(owner); + await share(owner, a.id, "authenticated"); + expect((await download(a.id, null)).status).toBe(404); + }); + + it("404s an archived artefact, including for its owner (AH7)", async () => { + const a = await create(owner); + const shared = await share(owner, a.id, "public"); + await app.request(`/api/artefacts/${a.id}/archive`, { + method: "POST", + headers: { cookie: owner }, + }); + expect((await download(a.id, owner)).status).toBe(404); + expect((await download(shared.publicSlug!, null)).status).toBe(404); + + // Restore → download → re-archive is the owner's escape hatch (no carve-out). + await app.request(`/api/artefacts/${a.id}/restore`, { + method: "POST", + headers: { cookie: owner }, + }); + expect((await download(a.id, owner)).status).toBe(200); + }); + + it("round-trips a declared schema block through export -> re-upload", async () => { + const declared = + '

garden

"; + const a = await create(owner, "Garden", declared); + + // Download, then re-upload the downloaded bytes as a new artefact — the + // path a user takes to edit an artefact outside the app. + const downloaded = await (await download(a.id, owner)).text(); + const form = new FormData(); + form.set("title", "Garden v2"); + form.set("kind", "interactive-doc"); + form.set("payload", new File([downloaded], "garden.html")); + const reuploaded = (await ( + await app.request("/api/artefacts", { method: "POST", body: form, headers: { cookie: owner } }) + ).json()) as ArtefactSummary; + + const again = await (await download(reuploaded.id, owner)).text(); + expect(again).toBe(declared); + // The convention survives because it lives in the trusted HTML — no plumbing. + expect(extractDeclaredSchema(again)).toEqual({ + key: "garden-v1", + version: 1, + example: { beds: [{ id: "b1" }] }, + }); + }); + + it("404s an unknown ref", async () => { + expect((await download("no-such-artefact", owner)).status).toBe(404); + }); + + it("names the file after the ref when the title has no usable characters", async () => { + const a = await create(owner, "★ ※ ★"); + const res = await download(a.id, owner); + expect(res.headers.get("content-disposition")).toContain( + `filename*=UTF-8''${encodeURIComponent(a.id)}.html`, + ); + }); +}); diff --git a/src/server/mcp/authoring-guide.ts b/src/server/mcp/authoring-guide.ts index 120d40a..68bbb83 100644 --- a/src/server/mcp/authoring-guide.ts +++ b/src/server/mcp/authoring-guide.ts @@ -24,14 +24,23 @@ export const PERSISTENCE_CONTRACT_SUMMARY = `Artefactor hosts self-contained HTM When you AUTHOR an artefact's HTML, follow this persistence contract so saved data survives: 1. Persist ONLY through the standard localStorage API. Not IndexedDB, cookies, sessionStorage, or your own fetch/network code — only localStorage is backed by Artefactor's store. (sessionStorage looks similar but is NOT persisted.) -2. Keep your whole state as ONE JSON object under ONE versioned key (e.g. "my-artefact-v1"). Versioning the key lets you migrate later: bump it and old data is ignored. +2. Keep your whole state as ONE JSON object under ONE versioned key (e.g. "my-artefact-v1"). Versioning the key is what makes a later shape change tractable — but note that bumping it *without* a migration means old data is ignored, i.e. every user silently loses what they saved (see "Breaking data-shape changes" below). 3. Wrap every getItem/setItem in try/catch and degrade to in-memory. A write can fail and must never break the artefact: the data may be loaded read-only (writes rejected), over the 5 MB budget (QuotaExceededError), or unavailable (opened as a bare file). You never detect or control read-only mode yourself — just tolerate a failed write. 4. Stay under 5 MB total. Don't stuff big base64 images/files into saved state — keep them in the HTML or reference them by URL. 5. Debounce frequent writes (~400 ms, e.g. typing in a textarea). Don't depend on storage events or cross-tab sync. +6. Declare the shape you save, in an inert block the browser ignores, so it is knowable without reading all your code and without anyone's data being populated: + + The example is the centrepiece — a populated exemplar, using INVENTED values only (never real user data: the block ships inside a downloadable, possibly public artefact). Write the schema and the code in the same breath; nothing enforces that they agree. Reads are synchronous and instant (the server seeds saved data before your script runs, so getItem returns the saved value on first paint). Writes save automatically (debounced, flushed when the page is hidden/closed). Opened as a plain file, native localStorage is used — same code still works. -Publishing: use create_artefact to publish, update_artefact to iterate in place (HTML is a full replacement, not a patch), set_visibility to share. Before a breaking data-shape change to an artefact that already has saved data (get_artefact / update_artefact report dataAuthorCount), bump the storage-key version or publish a new artefact — never silently change the shape in place. +Publishing: use create_artefact to publish, update_artefact to iterate in place (HTML is a full replacement, not a patch), set_visibility to share. Read back with get_artefact_html (the stored HTML, to derive a new artefact or to re-read one before updating it) and get_artefact_data (your own saved blob, verbatim, plus the declared schema, dataAuthorCount, and the currentPayloadVersion / authoredAgainstVersion pin). Both refuse an over-cap result rather than truncating — the user can then get the file from the Artefactor web app's "Download HTML" menu item. + +Breaking data-shape changes: check dataAuthorCount, call get_artefact_data, then prefer to MIGRATE FORWARD — ship migration code in the new HTML that reads the old key, transforms it, and writes the new one. Idempotent, run at load before first render, old key left in place for one generation, never clear(). A migration can only live in the artefact: the backend treats blobs as opaque and cannot migrate them. Bumping the storage-key version without migrating silently discards every user's saved data — only do that for disposable data, and say so. Publishing a separate v2 keeps the old artefact intact for existing users. +The snapshot you read is ONE user's blob, not the population: others may sit on older key versions, be partial, or have been written by HTML two revisions back, so a migration must tolerate shapes you never saw. Staleness per user: authoredAgainstVersion null = unknown, treat as possibly stale; ≠ currentPayloadVersion = written against older HTML, migration owed; = current = matches what is deployed. (That pin is backend-set and mechanical; the declared schema's own version is author-set and semantic — keep both.) Trust the declared schema for orientation, but verify it against the HTML before any shape-changing write. IMAGES — there are two ways an artefact reaches Artefactor, and embedded raster images decide which: - A) Publish via the connector (create_artefact / update_artefact): you send the HTML directly through the tool call. This works ONLY for artefacts with NO embedded raster images — i.e. no PNG/JPEG photos or screenshots, whether as base64 "data:" URIs or binary. You cannot reliably emit base64 image bytes through a tool argument, so pushing an image-bearing artefact this way will TRUNCATE/CORRUPT it. Anything authored as text is fine: HTML/CSS, inline SVG, CSS-drawn graphics, charts, diagrams. Most artefacts qualify (forms, prototypes, slide decks, interactive docs) — so path A is the common case. diff --git a/src/server/mcp/tools.test.ts b/src/server/mcp/tools.test.ts index f575fb5..3e04a9d 100644 --- a/src/server/mcp/tools.test.ts +++ b/src/server/mcp/tools.test.ts @@ -2,6 +2,7 @@ import { beforeEach, describe, expect, it } from "vitest"; import { Client } from "@modelcontextprotocol/sdk/client/index.js"; import { InMemoryTransport } from "@modelcontextprotocol/sdk/inMemory.js"; import { buildMcpServer, type McpToolDeps } from "./server"; +import { MAX_MCP_BLOB_BYTES, MAX_MCP_HTML_BYTES } from "./tools"; import { InMemoryArtefactRepository } from "../../domain/artefact/in-memory-artefact-repository"; import { InMemoryCollectionRepository } from "../../domain/collection/in-memory-collection-repository"; import { SINGLETON_SCOPE } from "../../domain/artefact/tenant-scope"; @@ -77,6 +78,8 @@ describe("MCP artefact tools (S18)", () => { "archive_artefact", "create_artefact", "get_artefact", + "get_artefact_data", + "get_artefact_html", "get_authoring_guide", "list_artefacts", "restore_artefact", @@ -234,4 +237,203 @@ describe("MCP artefact tools (S18)", () => { const r = await call(client, "get_artefact", { id: "does-not-exist" }); expect(r.isError).toBe(true); }); + // ---------------------------------------------------------------- S30 ---- + // Read-back: an agent can pull an artefact's HTML and its own data snapshot, + // so it can derive a new artefact from an existing one (A) or update one in + // place after losing the original from context (B) — seeing, for (B), the + // data shape a breaking HTML change might orphan. + + it("get_artefact_html returns the exact stored HTML with dataAuthorCount", async () => { + const client = await clientFor("u1"); + const html = + "Rakna"; + const a = json( + await call(client, "create_artefact", { title: "Counter", kind: "form", html }), + ); + await deps.dataRepo.save( + upsertDataEntry({ id: "d1", artefactId: a.id, authorId: "u2", blob: "{}" }), + ); + + const r = json(await call(client, "get_artefact_html", { id: a.id })); + expect(r).toEqual({ + id: a.id, + title: "Counter", + kind: "form", + html, + dataAuthorCount: 1, + }); + }); + + it("get_artefact_html refuses an over-cap payload, naming the real size", async () => { + const client = await clientFor("u1"); + const html = "" + "a".repeat(MAX_MCP_HTML_BYTES) + ""; + const a = json( + await call(client, "create_artefact", { title: "Huge", kind: "other", html }), + ); + + const r = await call(client, "get_artefact_html", { id: a.id }); + expect(r.isError).toBe(true); + // Names the actual size and points at the GUI download — never truncates, + // since truncated HTML is unusable for editing. + expect(r.content[0]!.text).toContain( + String(new TextEncoder().encode(html).byteLength), + ); + expect(r.content[0]!.text).toMatch(/download/i); + }); + + it("get_artefact_data returns the caller's own blob verbatim", async () => { + const client = await clientFor("u1"); + const a = json( + await call(client, "create_artefact", { title: "Tracker", kind: "form", html: "t" }), + ); + const blob = JSON.stringify({ "habit-tracker-v2": '{"habits":[]}' }); + await deps.dataRepo.save( + upsertDataEntry({ + id: "d1", + artefactId: a.id, + authorId: "u1", + blob, + now: new Date("2026-09-12T10:00:00.000Z"), + }), + ); + + const r = json(await call(client, "get_artefact_data", { id: a.id })); + expect(r.blob).toBe(blob); + expect(r.bytes).toBe(new TextEncoder().encode(blob).byteLength); + expect(r.updatedAt).toBe("2026-09-12T10:00:00.000Z"); + expect(r.dataAuthorCount).toBe(1); + }); + + it("get_artefact_data returns blob: null when the caller has no entry, and never another author's", async () => { + const client = await clientFor("u1"); + const a = json( + await call(client, "create_artefact", { title: "Tracker", kind: "form", html: "t" }), + ); + await deps.dataRepo.save( + upsertDataEntry({ id: "d1", artefactId: a.id, authorId: "u2", blob: '{"theirs":1}' }), + ); + + const r = json(await call(client, "get_artefact_data", { id: a.id })); + expect(r.blob).toBeNull(); + expect(r.bytes).toBe(0); + expect(r.updatedAt).toBeNull(); + // The snapshot is one blob, not the population — the count says others exist. + expect(r.dataAuthorCount).toBe(1); + expect(JSON.stringify(r)).not.toContain("theirs"); + }); + + it("get_artefact_data refuses an over-cap blob, naming the real size", async () => { + const client = await clientFor("u1"); + const a = json( + await call(client, "create_artefact", { title: "Fat", kind: "form", html: "f" }), + ); + const blob = JSON.stringify({ big: "a".repeat(MAX_MCP_BLOB_BYTES) }); + await deps.dataRepo.save( + upsertDataEntry({ id: "d1", artefactId: a.id, authorId: "u1", blob }), + ); + + const r = await call(client, "get_artefact_data", { id: a.id }); + expect(r.isError).toBe(true); + expect(r.content[0]!.text).toContain( + String(new TextEncoder().encode(blob).byteLength), + ); + expect(r.content[0]!.text).toMatch(/download/i); + }); + + it("get_artefact_data returns the artefact's declared schema as parsed JSON", async () => { + const client = await clientFor("u1"); + const a = json( + await call(client, "create_artefact", { + title: "Habits", + kind: "form", + html: + '

habits

", + }), + ); + + const r = json(await call(client, "get_artefact_data", { id: a.id })); + // The shape arrives without pulling the whole HTML into context, and it + // answers the empty-blob case a blob-only read cannot. + expect(r.schema).toEqual({ + key: "habit-tracker-v2", + version: 2, + example: { habits: [{ id: "h1" }] }, + }); + }); + + it("get_artefact_data returns schema: null when it is absent or unparseable", async () => { + const client = await clientFor("u1"); + const none = json( + await call(client, "create_artefact", { title: "Plain", kind: "form", html: "p" }), + ); + expect(json(await call(client, "get_artefact_data", { id: none.id })).schema).toBeNull(); + + const broken = json( + await call(client, "create_artefact", { + title: "Broken", + kind: "form", + html: '', + }), + ); + const r = await call(client, "get_artefact_data", { id: broken.id }); + // Malformed is never an error — the agent falls back to reading the HTML. + expect(r.isError).toBeFalsy(); + expect(json(r).schema).toBeNull(); + }); + + it("never validates a blob against the declared schema (AD8 opacity)", async () => { + const client = await clientFor("u1"); + const a = json( + await call(client, "create_artefact", { + title: "Habits", + kind: "form", + html: + '', + }), + ); + // A blob that contradicts the declaration entirely. + const blob = '{"something-else":"[1,2,3]"}'; + await deps.dataRepo.save( + upsertDataEntry({ id: "d1", artefactId: a.id, authorId: "u1", blob }), + ); + + const r = await call(client, "get_artefact_data", { id: a.id }); + expect(r.isError).toBeFalsy(); + expect(json(r).blob).toBe(blob); + }); + + it("get_artefact_data reports the payload version pin (AD9 reserved)", async () => { + const client = await clientFor("u1"); + const a = json( + await call(client, "create_artefact", { title: "Pinned", kind: "form", html: "p" }), + ); + const stored = (await deps.repo.findById(a.id, SINGLETON_SCOPE))!; + + const r = json(await call(client, "get_artefact_data", { id: a.id })); + expect(r.currentPayloadVersion).toBe(stored.payloadHash); + // Reserved, not yet populated: S19 (ALI-269) sets the pin on every write. + // The field's presence and shape are asserted now so S19 needs no tool change. + expect(r).toHaveProperty("authoredAgainstVersion"); + expect(r.authoredAgainstVersion).toBeNull(); + }); + + it("both read-back tools are owner-scoped: unknown, another user's, and archived all -> not found", async () => { + const u1 = await clientFor("u1"); + const a = json( + await call(u1, "create_artefact", { title: "Mine", kind: "other", html: "m" }), + ); + + const u2 = await clientFor("u2"); + expect((await call(u2, "get_artefact_html", { id: a.id })).isError).toBe(true); + expect((await call(u2, "get_artefact_data", { id: a.id })).isError).toBe(true); + + expect((await call(u1, "get_artefact_html", { id: "nope" })).isError).toBe(true); + expect((await call(u1, "get_artefact_data", { id: "nope" })).isError).toBe(true); + + await call(u1, "archive_artefact", { id: a.id }); + expect((await call(u1, "get_artefact_html", { id: a.id })).isError).toBe(true); + expect((await call(u1, "get_artefact_data", { id: a.id })).isError).toBe(true); + }); }); diff --git a/src/server/mcp/tools.ts b/src/server/mcp/tools.ts index 2d6a2bb..d970e52 100644 --- a/src/server/mcp/tools.ts +++ b/src/server/mcp/tools.ts @@ -22,6 +22,8 @@ import { restoreArtefactCommand, } from "../artefacts/lifecycle.command"; import { loadOwnActiveArtefact } from "../artefacts/get-own-artefact"; +import { getOwnDataEntry } from "../data/own-data.command"; +import { extractDeclaredSchema } from "../../domain/data/declared-schema"; import { toArtefactSummary } from "../routes/artefacts"; import { loadAuthoringGuide } from "./authoring-guide"; import { env } from "../env"; @@ -41,6 +43,21 @@ export interface McpToolDeps { dataRepo: DataRepository; } +// S30 — caps on what a read-back tool may return. An MCP result lands in the +// model's context, but payloads are capped at 100 MB (AH2) and blobs at 5 MB +// (AD8) — orders of magnitude more than a context can take. Each tool therefore +// hard-errors above its cap, naming the real byte size and pointing at the GUI +// download, rather than truncating: truncated HTML is unusable for editing and +// truncated JSON is unparseable, and either invites the model to act on a +// fragment as though it were whole. (Same shape of refusal as create_artefact's +// base64-raster-image rule.) +export const MAX_MCP_HTML_BYTES = 1024 * 1024; // 1 MB +export const MAX_MCP_BLOB_BYTES = 256 * 1024; // 256 KB + +// A read-back result that is too big to put in the model's context. Carried as +// an error so it reaches the model as an `isError` result it can act on. +class ResultTooLarge extends Error {} + // An MCP tool result wrapping a JSON value as text content. function ok(value: unknown) { return { @@ -82,7 +99,11 @@ async function run(body: () => Promise) { try { return ok(await body()); } catch (err) { - if (err instanceof ArtefactNotFound || err instanceof InvariantViolation) { + if ( + err instanceof ArtefactNotFound || + err instanceof InvariantViolation || + err instanceof ResultTooLarge + ) { return fail(err.message); } throw err; @@ -235,6 +256,100 @@ export function registerArtefactTools( }), ); + // S30 — read-back. Until now an agent could write artefacts but never read one, + // which blocks two things: (A) deriving a new artefact from an existing one, + // and (B) updating one in place after losing the original from context. (B) is + // a correctness problem, not a convenience: `update_artefact` replaces the HTML + // and leaves every per-user data blob untouched, and the backend treats blobs + // as opaque (AD8), so it cannot migrate them — only the author's own HTML can. + // Both tools are owner-scoped via `loadOwnActiveArtefact`, like every other + // tool here: unknown / not yours / archived / out-of-scope → not found. + + server.registerTool( + "get_artefact_html", + { + title: "Get artefact HTML", + description: + "Return the stored HTML of one of your active artefacts, exactly as served. Use it to derive a new artefact from an existing one, or to re-read an artefact you are about to update after losing the original from context (update_artefact replaces the HTML wholesale, so you need the current source to change it safely). Also returns dataAuthorCount — if it is > 0, call get_artefact_data before any change to the data shape. Refuses an artefact whose HTML is too large for a tool result; download that one from the Artefactor web app instead.", + inputSchema: { id: z.string().min(1).describe("The artefact id.") }, + }, + async ({ id }) => + run(async () => { + const a = await loadOwnActiveArtefact(repo, { id, ownerId: userId, scope }); + if (a.payloadBytes > MAX_MCP_HTML_BYTES) { + throw new ResultTooLarge( + `This artefact's HTML is ${a.payloadBytes} bytes, over the ${MAX_MCP_HTML_BYTES}-byte limit for a tool result. It is not truncated, because partial HTML cannot be edited safely — download the file from the Artefactor web app ("Download HTML" on the artefact) and work from that instead.`, + ); + } + const html = new TextDecoder().decode( + await payloadStore.get(a.payloadRef), + ); + return { + id: a.id, + title: a.title, + kind: a.kind, + html, + dataAuthorCount: (await dataRepo.listAuthorsByArtefact(a.id)).length, + }; + }), + ); + + server.registerTool( + "get_artefact_data", + { + title: "Get artefact data snapshot", + description: + "Return YOUR OWN saved data for one of your active artefacts, verbatim, plus the data shape the artefact declares for itself. Read this before an HTML change that alters the shape the artefact reads from localStorage: `schema` is the artefact's own declaration (trust it for orientation, verify it against the HTML before acting — nothing enforces that they agree), `blob` is your own entry (null if you have none), and `dataAuthorCount` says how many users hold data in total. The snapshot is ONE user's blob, not the population: other users' entries may sit on older key versions, be partial, or have been written by HTML two revisions back, so any migration you ship must tolerate shapes you never saw. Compare authoredAgainstVersion with currentPayloadVersion to judge staleness (null = unknown, treat as possibly stale). Refuses a blob too large for a tool result.", + inputSchema: { id: z.string().min(1).describe("The artefact id.") }, + }, + async ({ id }) => + run(async () => { + const a = await loadOwnActiveArtefact(repo, { id, ownerId: userId, scope }); + // The caller's OWN entry only (AD2/AD4), through the same command the + // BFF uses — returned verbatim, never summarised (AD8 opacity). + const entry = await getOwnDataEntry( + { ref: a.id, authorId: userId, scope }, + { artefactRepo: repo, collectionRepo: deps.collectionRepo, dataRepo }, + ); + const bytes = entry + ? new TextEncoder().encode(entry.blob).byteLength + : 0; + if (bytes > MAX_MCP_BLOB_BYTES) { + throw new ResultTooLarge( + `Your saved data for this artefact is ${bytes} bytes, over the ${MAX_MCP_BLOB_BYTES}-byte limit for a tool result. It is not truncated, because partial JSON cannot be parsed — inspect it from the artefact itself, or download the artefact from the Artefactor web app.`, + ); + } + // The declared schema is lifted from the trusted HTML as a string, the + // same class of operation as locating to inject the bootstrap. + // The server forwards it and NEVER validates a blob against it. + const schema = extractDeclaredSchema( + new TextDecoder().decode(await payloadStore.get(a.payloadRef)), + ); + return { + id: a.id, + blob: entry?.blob ?? null, + bytes, + // Returned so a later write can be pinned against it (S31/ALI-268); + // without it an agent's write blindly overwrites whatever the user has + // saved since. + updatedAt: entry?.updatedAt.toISOString() ?? null, + dataAuthorCount: (await dataRepo.listAuthorsByArtefact(a.id)).length, + schema, + // Two version notions, deliberately kept apart (AD9): the *mechanical* + // pin below is the payload hash the backend stamps on a write, while + // the schema's own `version` is *semantic* and author-declared. + currentPayloadVersion: a.payloadHash, + // Reserved by S30, populated by S19 (ALI-269) — which is an optional + // sharpener of this staleness signal, not a dependency: AD9 is advisory + // and never gates a read or write, so `null` is correct until then. + // The doctrine is written now and becomes exact then: null ⇒ unknown, + // treat as possibly stale; ≠ current ⇒ written against older HTML, + // migration owed; = current ⇒ matches what is deployed. + authoredAgainstVersion: null as string | null, + }; + }), + ); + server.registerTool( "set_visibility", { diff --git a/src/server/routes/download.ts b/src/server/routes/download.ts new file mode 100644 index 0000000..fe47310 --- /dev/null +++ b/src/server/routes/download.ts @@ -0,0 +1,79 @@ +import { Hono } from "hono"; +import { ArtefactNotFound } from "../../domain/artefact/errors"; +import type { PayloadStore } from "../../domain/artefact/ports"; +import type { AccessPolicy } from "../../domain/artefact/access"; +import type { ArtefactRepository } from "../../domain/artefact/artefact-repository"; +import type { CollectionRepository } from "../../domain/collection/collection-repository"; +import { resolveViewableArtefact } from "../data/own-data.command"; +import type { AuthEnv } from "../middleware/auth"; +import type { TenantScopeResolver } from "../middleware/tenant-scope"; +import { + artefactFilename, + attachmentDisposition, +} from "../../shared/artefact-filename"; + +export interface DownloadRoutesDeps { + artefactRepo: ArtefactRepository; + // AH20 — the access decision needs the effective tier (collection tree root). + collectionRepo: CollectionRepository; + payloadStore: PayloadStore; + // S22 (AH17) — resolves the request's tenant scope for the id-fallback resolve. + resolveScope: TenantScopeResolver; + // S22 (AH18) — decides the `authenticated` tier for the slug-resolved artefact. + accessPolicy?: AccessPolicy; +} + +// S30 — export an artefact's HTML. Mounted at `/api/artefacts/:ref/download`, +// where `:ref` is the artefact's slug or its id (the id form is what the owner's +// dashboard uses for a never-shared artefact). Resolution and the access matrix +// are the *same* ones the data reads use (`resolveViewableArtefact`), so +// effective-tier resolution through a collection root (AH20) and the +// `AccessPolicy` cell (AH18) are inherited, not re-implemented: unknown ref, +// not-viewable, and archived all surface as a flat 404 (AH7/AH8). +// +// The body is the **stored** payload, verbatim — never the injected render: no +// S13 localStorage bootstrap and no S12 host shell. That is what makes the +// download round-trip (download → edit → re-upload / `update_artefact` yields +// the same artefact), and it keeps the export honest about what is hosted. +export function createDownloadRoutes(deps: DownloadRoutesDeps) { + const r = new Hono(); + + r.get("/:ref/download", async (c) => { + const ref = c.req.param("ref"); + try { + const artefact = await resolveViewableArtefact( + deps, + ref, + c.get("user")?.id ?? null, + await deps.resolveScope(c), + ); + const payload = await deps.payloadStore.get(artefact.payloadRef); + // Name the file after the title; a title with no letters or digits at all + // falls back to the handle the caller already has (slug, else id). + const filename = artefactFilename( + artefact.title, + artefact.publicSlug ?? artefact.id, + ); + // A raw Response (rather than `c.body`) so the stored bytes go out + // untouched — the payload is an opaque byte array, not a Hono body type. + // The cast only narrows the store's `ArrayBufferLike` backing to the + // `ArrayBuffer` a BodyInit is typed for; the bytes are never copied. + return new Response(payload as Uint8Array, { + status: 200, + headers: { + "Content-Type": "text/html; charset=UTF-8", + "Content-Disposition": attachmentDisposition(filename), + // The bytes actually being sent, which for a healthy store is the + // aggregate's `payloadBytes` — taking it from the body means a + // store/row disagreement can never produce a malformed response. + "Content-Length": String(payload.byteLength), + }, + }); + } catch (err) { + if (err instanceof ArtefactNotFound) return c.notFound(); + throw err; + } + }); + + return r; +} diff --git a/src/server/routes/index.ts b/src/server/routes/index.ts index 01f2bfc..8ab0bc5 100644 --- a/src/server/routes/index.ts +++ b/src/server/routes/index.ts @@ -22,6 +22,7 @@ import { createBookmarkRoutes } from "./bookmarks"; import { listEffectivelyShared } from "../collections/shared.query"; import { listSharedCollections } from "../collections/shared-collections.query"; import { createDataRoutes } from "./data"; +import { createDownloadRoutes } from "./download"; import { createViewRoutes } from "./views"; import { createUserRoutes } from "./users"; import type { @@ -206,6 +207,20 @@ export function createApiRoutes( }), ); + // S30 — Export: the artefact's stored HTML as a download, addressed by its + // slug or id. Access follows the same matrix as the data reads (AH20/AH18); + // anonymous callers may download a `public` artefact by slug. + api.route( + "/artefacts", + createDownloadRoutes({ + artefactRepo: artefactRepository, + collectionRepo: collectionRepository, + payloadStore, + resolveScope, + accessPolicy, + }), + ); + // S21 — Artefact Views: the "viewed by" list for an artefact, addressed by its // slug or id. A view itself is recorded on the serving path (see routes/serve.ts). api.route( diff --git a/src/shared/artefact-filename.test.ts b/src/shared/artefact-filename.test.ts new file mode 100644 index 0000000..16041db --- /dev/null +++ b/src/shared/artefact-filename.test.ts @@ -0,0 +1,65 @@ +import { describe, expect, it } from "vitest"; +import { artefactFilename, attachmentDisposition } from "./artefact-filename"; + +// S30 — the pure download-filename helper. The export endpoint turns an +// artefact's (title, fallback ref) into the name a browser saves the file as. +// Pure and framework-free, so it is unit-tested on its own. +describe("artefactFilename (S30)", () => { + it("slugifies an ASCII title", () => { + expect(artefactFilename("My Quarterly Report", "abc123")).toBe( + "my-quarterly-report.html", + ); + }); + + it("keeps non-ASCII letters in the name", () => { + expect(artefactFilename("Mötesprotokoll Öst", "abc123")).toBe( + "mötesprotokoll-öst.html", + ); + }); + + it("falls back to the ref when the title has no usable characters", () => { + expect(artefactFilename("★ ※ ★", "abc123")).toBe("abc123.html"); + }); + + it("falls back to the ref when the title is blank", () => { + expect(artefactFilename(" ", "abc123")).toBe("abc123.html"); + }); + + it("bounds a very long title", () => { + const name = artefactFilename("a".repeat(300), "abc123"); + expect(name.length).toBeLessThanOrEqual(100); + expect(name.endsWith(".html")).toBe(true); + }); + + it("collapses separators and strips path characters", () => { + expect(artefactFilename("../etc/passwd — v2", "abc123")).toBe( + "etc-passwd-v2.html", + ); + }); + + it("always ends in .html", () => { + expect(artefactFilename("index.html", "abc123")).toBe("index-html.html"); + }); +}); + +describe("attachmentDisposition (S30)", () => { + it("carries an ASCII filename and the RFC 5987 filename*", () => { + const header = attachmentDisposition("mötesprotokoll-öst.html"); + expect(header).toBe( + "attachment; filename=\"motesprotokoll-ost.html\"; filename*=UTF-8''m%C3%B6tesprotokoll-%C3%B6st.html", + ); + }); + + it("leaves an already-ASCII name identical in both forms", () => { + expect(attachmentDisposition("report.html")).toBe( + "attachment; filename=\"report.html\"; filename*=UTF-8''report.html", + ); + }); + + it("replaces characters that survive neither folding nor quoting", () => { + // A name that folds to nothing ASCII still needs a usable plain filename. + expect(attachmentDisposition("日本語.html")).toContain( + 'filename="artefact.html"', + ); + }); +}); diff --git a/src/shared/artefact-filename.ts b/src/shared/artefact-filename.ts new file mode 100644 index 0000000..ca53a8b --- /dev/null +++ b/src/shared/artefact-filename.ts @@ -0,0 +1,48 @@ +// S30 — download filenames. Turning an artefact's title into the name a browser +// saves it under is pure string work with no domain authority, so it lives here +// (shared between the BFF that sets the header and any client that shows it) +// rather than in a route handler. + +// Base name budget, leaving room for the ".html" suffix inside a 100-char name. +const MAX_BASE_LENGTH = 95; + +// The name used when a title folds away to nothing AND no usable ref is given. +const GENERIC_NAME = "artefact"; + +// Slugify a title: lowercase, keep letters/digits (in any script), collapse +// everything else to single hyphens. Path separators and dots go with the rest, +// so a title can never steer the saved file anywhere (`../etc/passwd` → +// `etc-passwd`). +function slugify(value: string): string { + return value + .toLowerCase() + .replace(/[^\p{L}\p{N}]+/gu, "-") + .replace(/^-+|-+$/g, ""); +} + +// The filename for an artefact download. `fallback` is the artefact's ref (slug +// or id) — used when the title carries no letters or digits at all (a +// symbol-only or blank title). Always ends in `.html`. +export function artefactFilename(title: string, fallback: string): string { + const base = slugify(title) || slugify(fallback) || GENERIC_NAME; + return `${base.slice(0, MAX_BASE_LENGTH).replace(/-+$/, "")}.html`; +} + +// Fold a name to ASCII for the legacy `filename` parameter: decompose accented +// letters and drop the combining marks, then drop anything still non-ASCII. +function foldToAscii(name: string): string { + return name + .normalize("NFD") + .replace(/\p{M}+/gu, "") + .replace(/[^\x20-\x7E]+/g, ""); +} + +// The `Content-Disposition` value for an artefact download. Carries both forms: +// the quoted ASCII `filename` every client understands, and the RFC 5987 +// `filename*` that preserves non-ASCII titles in clients that read it. +export function attachmentDisposition(filename: string): string { + // Fold the *base* name — the ".html" suffix is re-applied by the helper, so + // folding the whole filename would turn "report.html" into "report-html". + const ascii = artefactFilename(foldToAscii(filename.replace(/\.html$/i, "")), ""); + return `attachment; filename="${ascii}"; filename*=UTF-8''${encodeURIComponent(filename)}`; +}