Skip to content

fix(cursor): filter private transcript fields - #685

Closed
xnne-bot wants to merge 1 commit into
NevaMind-AI:mainfrom
XnneHangLab-Mirror:fix/cursor-transcript-private-fields
Closed

xnne-bot wants to merge 1 commit into
NevaMind-AI:mainfrom
XnneHangLab-Mirror:fix/cursor-transcript-private-fields

Conversation

@xnne-bot

@xnne-bot xnne-bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

📝 Pull Request Summary

Strip Cursor Agent's runtime prompt envelope from prepared transcripts while preserving the user's actual prompt, conversation/tool content, and other tags.


✅ What does this PR do?

  • Unwraps the real Cursor Agent JSONL user-prompt shape before writing prepared memory and skill inputs:
    • optional leading <timestamp>…</timestamp>
    • enclosing <user_query>…</user_query>
  • Retains the enclosed prompt verbatim and preserves all other content blocks and fields.
  • Leaves assistant records untouched, and only unwraps an exact top-level runtime envelope; user text that merely contains a tag remains unchanged.
  • Keeps raw Cursor session files unchanged; sanitization affects only generated bridging inputs.
  • Adds direct sanitizer and prepare-path tests, including source-file immutability and cursor advancement.

🤔 Why is this change needed?

Historical Cursor Agent JSONL sessions put runtime time metadata and a transport wrapper around every user prompt. These fields are not authored user content and introduce unnecessary scheduling/runtime noise into memory and skill mining.

The implementation and output comparison use an actual historical Cursor Agent session, not a synthetic fixture. That session is copied to a temporary input root and processed with main and this branch; the raw source is never modified.


🔍 Type of Change

Please check what applies:

  • Bug fix
  • New feature
  • Documentation update
  • Refactor / cleanup
  • Other (please explain)

✅ PR Quality Checklist

  • PR title follows an allowed format (for example feat:, fix:, docs:, memory base:)
  • Changes are limited in scope and easy to review
  • Documentation updated where applicable (not applicable: internal transcript-preparation behavior)
  • No breaking changes (or clearly documented)
  • Related issues or discussions linked (not applicable: no separate issue for this narrowly scoped fix)

Validation:

  • uv run python -m pytest tests/test_host_sessions.py tests/test_cursor_self_sessions.py — 58 passed
  • uv run ruff check src/memu/hosts/cursor/sessions.py tests/test_host_sessions.py
  • uv run ruff format --check src/memu/hosts/cursor/sessions.py tests/test_host_sessions.py
  • uv run mypy src/memu/hosts/cursor/sessions.py tests/test_host_sessions.py
  • Prepared a copied historical Cursor session against main and this branch; both emitted the same rows, with the PR removing only the runtime wrapper bytes.

📌 Optional

  • Screenshots or examples added (if applicable)
  • Edge cases considered
    • timestamps are optional in the wrapper
    • user text containing a non-envelope tag is unchanged
    • assistant text is unchanged
    • other content blocks are preserved
    • raw source transcripts remain unchanged
  • Follow-up tasks mentioned

@MrXnneHang
MrXnneHang force-pushed the fix/cursor-transcript-private-fields branch from bea9e93 to ea38068 Compare September 3, 2026 06:57
@MrXnneHang
MrXnneHang force-pushed the fix/cursor-transcript-private-fields branch from ea38068 to acc45eb Compare September 3, 2026 07:23
@xnne-bot

xnne-bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor Author

Actual historical Cursor Agent JSONL comparison, not a synthetic fixture:

I copied one 55-record historical session into a temporary Cursor input root and ran prepare_transcripts once against main and once against this PR. Both runs prepared one session, preserved the source JSONL, and emitted the same row counts and ordering.

Output Before (main) After (this PR) Reduction Rows
1.jsonl 44,566 B 44,481 B 85 B (0.2%) 38 → 38
1_full.jsonl 47,363 B 47,278 B 85 B (0.2%) 54 → 54

The 85-byte reduction is one real user runtime envelope:

<timestamp>Thursday, Aug 6, 2026, 6:01 PM (UTC+9)</timestamp>
<user_query>
…actual prompt…
</user_query>

After preparation, the generated transcript contains only …actual prompt…. The sanitizer does not remove arbitrary tags: it unwraps only an exact top-level user prompt envelope, leaves assistant records untouched, and preserves all other fields and content blocks.

@xnne-bot

xnne-bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor Author

实际历史数据核查后的结论:Cursor 的情况与 Claude Code / Codex / OpenClaw / Hermes 不同。

当前 adapter 的兼容边界是 Cursor 导出的 agent-transcripts/*.jsonl,而不是 Cursor 的内部 store.db。我扫描了本机历史 JSONL:对话记录已经是很干净的投影,只有 role、message.content 以及 text/tool block 所需字段;没有发现 requestId、providerOptions、model、usage、reasoning signature 等可由 delete-only sanitizer 删除的私有/runtime JSON 字段。

这些字段确实存在于 ~/.cursor/chats/.../store.db 的内部消息树中,但该 SQLite store 不是稳定的跨版本兼容边界。为它建立 extract-only projection 会将 adapter 绑定到 Cursor 内部 schema,并带来版本字段新增、移动或改名时丢失真实对话/工具内容的风险;这不符合当前 host adapter 保持兼容性的目标。

本 PR 最终能在真实 JSONL 中处理的只有 user text 外层的:

<timestamp>…</timestamp>
<user_query>…</user_query>

我以一个真实的 55-record 历史 session 分别在 main 和本 PR 上实际执行 prepare_transcripts:行数、顺序和有效内容均不变,1.jsonl 与 1_full.jsonl 各只减少 85 B(约 0.2%)。这不足以证明引入 Cursor 专用 sanitizer 的长期维护成本是合理的。

因此准备关闭本 PR,不合入当前实现。后续原则是:继续以 agent-transcripts/*.jsonl 作为 Cursor 的兼容边界,并保持 delete-only / preserve-unknown;只有在真实导出 JSONL 中确认出现具体的 private/runtime 字段时,才针对该字段新增有样本和回归测试的窄范围过滤。

@xnne-bot

xnne-bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor Author

基于上述真实历史 Cursor JSONL 的核查结论关闭此 PR:当前导出 transcript 已足够干净,尚无值得维护的 delete-only 私有字段过滤目标。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants