Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
27 changes: 11 additions & 16 deletions docs/reasoning-distillation-design.md
Original file line number Diff line number Diff line change
@@ -1,12 +1,13 @@
# 内化推理蒸馏:开发设计与验收规格

> **2026-09-29 产品契约(1.0.54):** 默认关闭;开启后,已完成轮次的可编辑思考一次交给配置的小模型整理,直接采用替换正文并事务回写。单段使用正文,多段一次返回轻量编号/正文,不再默认执行 claim 提取、coverage 证明或第二次 judge。
> 第一次和后续完成轮次均由应用作用域调度;调度与整理耗时分开统计。真实验收记录模型调用、解析与数据库写回,不能用前台先返回代替速度验收。
> 保留有效条件、数值、未决事项、当前方案执行状态和实际失败/回滚;删除重复和没有后续价值的自我纠错。结构由内容决定,不固定类别或条数。
> 整理后的说明文字用中文;路径、命令、符号、代码、URL、配置键、版本号和数值保持原样。此要求适用于当前单次整理路径,不要求恢复旧版 propose/judge 调用。
> 显式 `small_model` 优先;不可用时保留原文,不回退主模型。辅助调用 low、零自动重试。历史 native-wire 实验路径的 small/agent/primary 枚举不是当前默认完成轮次的调用链。
> 成功写回发送正常更新事件,TUI、重新打开的历史和后续请求读取同一正文;已经发出的请求保留原快照。原文和 metadata 留作关闭功能后的回放;签名/加密、过期或取消内容不写回。
> 当前实现、真实测量和验收见 [单次整理验收记录](reasoning-denoise-acceptance.md)。下方旧版设计保留为历史背景,与此契约冲突的逐槽位双调用、同步首轮等描述已被取代。
> **2026-10-08 产品契约(逐步重写):** 默认关闭;开启后,每段思考一结束(`reasoning-end`)就交给配置的小模型单独整理,与本步剩余输出和工具执行并行;多段思考各自一次调用、并发执行、互不牵连。
> 每次向模型发请求前先过“发送屏障”:等待本会话未完成的整理,至多 `OPENCODE_REASONING_DISTILLATION_SETTLE_MS`(默认 3000 ms),之后把剩余任务封存并中断。封存的任务永不写回,因此每段思考要么在第一次回传前完成替换,要么保持原文,不会因事后改写破坏已发请求的提示缓存前缀。写回按单段校验(未封存、会话未回退、持久化内容未变),后续消息不再导致整批作废。
> 整理调用走 `@opencode-ai/llm` 引擎(opencode 宿主经 native 适配器;引擎不支持的 SDK 包才回退 AI SDK):无工具、temperature 0、零自动重试,`none`/`low` 按协议翻译。预算按会话计 token(缺失用量按提示+输出估算,不再因此暂停),仅连续失败才暂停。
> 整理后的说明文字默认中文(`language` 可选 en);路径、命令、符号、代码、URL、配置键、版本号和数值保持原样。全文无信息时以非空占位“(无有效推理)”替换,保证 `reasoning_content` 等字段仍然存在。
> 哪些 metadata 把思考绑定到原文(签名、加密、存储引用、明文镜像)由引擎 `ReasoningCarrier.classify` 判定;签名/加密/未知载体不改写。原文和原 metadata 留作关闭功能后的回放。
> 设计与验收见 [逐步重写交付记录](reasoning-rewrite-engine-2026-10-08.md)。下方 1.0.54 契约与更早的旧版设计保留为历史背景,与此契约冲突处已被取代。
>
> **2026-09-29 产品契约(1.0.54,已被上方取代):** 默认关闭;开启后,已完成轮次的可编辑思考一次交给配置的小模型整理,直接采用替换正文并事务回写。单段使用正文,多段一次返回轻量编号/正文,不再默认执行 claim 提取、coverage 证明或第二次 judge。

```json
{
Expand All @@ -22,7 +23,7 @@ The original source is retained as host-owned provenance, excluded from provider
Background work belongs to the running application scope; shutdown cancels unfinished work without changing history.

- 日期:2026-09-22,第 4 版。
- 状态:提案。产品实现未开始;文档检查通过不代表产品、模型质量或上游兼容性通过验收。
- 状态:历史提案(第 4 版)。现行实现以文首产品契约为准;本节以下内容仅作背景。
- 调研基线:`cedcfb3647e50d9a25633b65e4f3e599adc7fb6c`;`origin/dev` = `f3f4e50a0164c3b11b3b0126b2cd9c1ebe25dc37`。实施前重新核对开发分支。
- 本次范围:完善设计,不修改产品代码、用户配置或原始历史,不部署、不创建 Issue 或提交。

Expand Down Expand Up @@ -253,8 +254,7 @@ type SupportResult =

type ReasoningSlotShape = "interleaved-field" | "unsigned-reasoning" | "downgraded-text"
type SlotEligibility =
| { allowed: true; capabilityFingerprint: string }
| { allowed: false; protection: "P1" | "P2" | "P3" | "P4" | "P5" }
{ allowed: true; capabilityFingerprint: string } | { allowed: false; protection: "P1" | "P2" | "P3" | "P4" | "P5" }
type WireReasoningMapping = {
refs: readonly SourceRef[]
shape: ReasoningSlotShape
Expand Down Expand Up @@ -298,12 +298,7 @@ type ModelProjection = {
text: string // Rendered only from validated claims and preserved source spans; 输出一律中文(§5.4.1),技术标识符逐字保留
}
type AuditViolationKind =
| "fabricated"
| "concealed"
| "simulated_execution"
| "unbacked_completion"
| "evidence_swap"
| "unverifiable"
"fabricated" | "concealed" | "simulated_execution" | "unbacked_completion" | "evidence_swap" | "unverifiable"
type AuditRecord = {
subject: "source-agent" | "distiller"
findings: readonly {
Expand Down
168 changes: 168 additions & 0 deletions docs/reasoning-rewrite-engine-2026-10-08.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,168 @@
# Reasoning rewrite: engine-level, per-step, rewrite-before-first-resend

Branch: `refactor/reasoning-rewrite-engine`. Baseline: `main` at `1b69adcdcf`.

## Why

The 1.0.54+ reasoning distillation ("重塑思考") organizes a completed turn's reasoning once, after the turn ends, and
adopts the result in the background. Review found:

- **Cache cost without matching benefit.** A step's reasoning is first resent by the next step of the same turn. Turn-end
rewriting therefore changes text that earlier requests already sent as input, invalidating the provider prompt cache
from that point for the next turn, while within-turn resends (where interleaved-thinking providers actually consume
reasoning) still carry the original.
- **Replacement is lost under normal use.** If the user continues before the background job commits, the next request
carries the original and adoption is rejected (`reasoning turn was retried or continued`); each turn is tried once.
- **Two organizer implementations.** The opencode host calls AI SDK `generateText` with an `openaiCompatible`-only effort
option; the core runner uses the `@opencode-ai/llm` engine. Multi-slot requests are one serial JSON call that is
rejected wholesale on any invalid item.
- **SDK metadata sniffing in core.** Editability is decided by key names from provider SDK metadata, duplicated in three
places, apart from the protocol code that defines those keys.
- **Dead native-wire projector.** The legacy propose/judge/claims/gates pipeline and the native-runtime hook have no
production caller. The core runner also converts the whole history on the foreground path and lowers the full main
request for the canonical path without using either result.
- **Empty replacement drops the field.** In the core runner an all-noise replacement removes the reasoning part, so the
provider request loses `reasoning_content` for that assistant message.

## Scope

- Reasoning distillation in `packages/core` (organizer, adoption, scheduling, replay, editability), the opencode session
host (`prompt.ts`, `processor.ts`, `llm.ts`, `reasoning-adoption.ts`), and an engine-owned reasoning carrier classifier
in `packages/llm`.
- No persisted-format change: `distillation` provenance, `providerMetadata` and the `reasoningDistillation` config keep
their shapes. The legacy `compatibility` config entry is still accepted and ignored.
- No public HTTP/event shape change. No DB migration.

## Approach

1. **Remove dead code.** Delete the native-wire projector (core runner `distillCycle`, claims/gates/plan/audit/cache/
projection/review/slot/budget modules, opencode propose/judge/interleaved helpers, native-runtime hook) and their tests;
drop the unused foreground conversion and `llm.prepare` from the core runner canonical path.
2. **One engine organizer.** `organizeReasoning` takes single-slot prompts only and runs slots concurrently; each slot
succeeds or fails independently. A shared engine transport builds an `LLMRequest` (no tools, temperature 0, no
retries, `none`/`low` effort translated per protocol) and calls `LLMClient.generate`. The opencode host uses it through
its native request adapter for engine-supported packages and keeps AI SDK only as an explicit fallback transport.
Budget: per-session token ceiling counted from reported usage, or estimated from prompt and output when the provider
omits usage; no per-session call cap; consecutive transport failures still pause the session.
3. **Rewrite before first resend.** A job starts as soon as a reasoning part ends (`reasoning-end`), overlapping the rest of
the step and its tool execution. Before each provider request a session `settle` waits for in-flight jobs up to
`OPENCODE_REASONING_DISTILLATION_SETTLE_MS` (default 3000), then seals and interrupts the remainder. A sealed job can
never adopt, so a part is rewritten before its first resend or never; adoption is validated per part (session not
reverted, part unchanged, not sealed) instead of rejecting the whole turn on continuation. All-noise replacements keep
an empty reasoning part so the provider field stays present.
4. **Engine-owned carrier classification.** `@opencode-ai/llm` exports `ReasoningCarrier.classify(metadata)`, declaring for
each provider namespace which keys are opaque (signatures, encrypted payloads, item references) and which are
plaintext mirrors. Core editability uses it; the duplicated key sniffers are removed.

## Acceptance

- Workspace typecheck and lint (no ratchet increase).
- Focused tests: core organizer/adoption/scheduler/canonical/runner, llm carrier classifier, opencode llm/prompt/session
reasoning tests, TUI adoption test; DAG-core gate.
- New regression coverage: per-slot partial success; settle adopts finished jobs before the next request and seals
unfinished ones (a sealed job never adopts); adoption after continuation succeeds when the part is unchanged; core
runner keeps `reasoning_content: ""` for an all-noise replacement; carrier classification per namespace.
- Live: an isolated run with a configured DeepSeek/QWEN/GLM small model shows the second step of a tool-using turn sending
the replaced reasoning, recorded with allowlisted timings.
- Savings comparison uses a script-calibrated estimate so a Chinese rewrite of English reasoning is judged by real cost.

## Results (2026-10-08)

Implementation note: steps 1 and 2 landed together in core because the organizer, adoption and runner modules were
rewritten in place; writing an intermediate multi-slot version only to delete it would have been wasted churn.

| Check | Result |
| ------------------------------ | ----------------------------------------------------------------------- |
| Workspace typecheck | 31/31 tasks |
| Lint | 0 errors, 4,738 warnings; ratchet lowered 4,850 → 4,750 (CI margin ~10) |
| `packages/core` full suite | 1,459 pass, 0 fail |
| `packages/llm` full suite | 325 pass, 30 skip, 0 fail |
| `packages/opencode` full suite | 5,241 pass, 33 skip, 1 todo, 0 fail |
| `packages/tui` full suite | 282 pass, 1 skip, 0 fail |
| DAG-core gate | all critical behavior and coverage floors passed |

Suites were rerun on the touched packages after the lint, estimator and review fixes below (core full suite; opencode
prompt, llm, session, processor, compaction and quality files: 337 pass, 0 fail).

Net change: about 11.7k lines removed (dead native-wire projector, propose/judge/claims/gates modules and their tests),
about 2k added including new tests.

### Lint

Every warning on a line this change adds or edits is fixed, including redundant non-null and `as` assertions removed
by oxlint's safe autofix in touched files. Remaining warnings in touched files predate this change.

### Token estimation for savings

Live runs first showed a Chinese organization of English reasoning rejected as `no-savings`. Measured on the configured
relay, the same reasoning was 114 DeepSeek tokens in English (3.05 chars/token) and 106 in its Chinese rewrite
(1.87 chars/token); GLM: 126 vs 124. The conservative reserve estimator (CJK = 1 token/char, other = 4 chars/token)
reported 87 vs 103 and flipped the sign. The savings check now uses `Token.estimateComparable` (CJK 0.55 token/char,
other 3 chars/token, calibrated on those tokenizers); reserves and limits keep the conservative estimator. The default
Chinese organizer language is kept.

### Self-review fixes

- Finished jobs that adopted nothing are dropped immediately, so sessions that never send again retain no work.
- The opencode host keeps only adopted part identities until the next barrier and re-reads the parts there.

### Review fixes

- The opencode host keeps its scheduler and budget in `InstanceState`: disposing a directory closes the instance scope
and cancels pending organizer calls (a mutation test with the application scope fails). The job owns an
`AbortController` that is aborted on interruption, so cancellation reaches the provider request on every path.
- Adoption also checks a durable fence in both runtimes: it is refused once the session has an assistant message newer
than the part's message, so a send attempt persisted by another process cannot be followed by a late rewrite. The
in-process barrier remains the primary mechanism; the fence covers what it cannot see.
- The opencode loop treats its persisted assistant message as the claim and re-reads the previous assistant
message's settled reasoning after it, before building the request. Only that message can still adopt before the
claim, so a rewrite another process commits between this loop's history read and its claim is carried by the
request, and one after the claim is refused. The fence is a direct query for later assistant rows instead of a paged
message scan inside the adoption transaction. A test commits a foreign rewrite in the claim's own transaction with a
SQLite trigger; without the re-read the request carries the original.
- The core runner persists its attempt with the first stream event, after building the request, so a concurrent sender
in another process can still send the original once. That costs one prompt-cache miss, not history consistency;
closing it would mean starting the step before reading history, which changes the runner's overflow and compaction
paths.
- Deleting a session cancels its pending jobs in both runtimes and releases its budget entry; a mutation test without
the cancellation times out.
- The organizer's output cap (24,576) is lowered to the small model's declared output limit on the engine and the
AI-SDK fallback, so providers that enforce their cap do not reject every organizer call.

### Smoke (built host binary, isolated)

`bun run build --single --skip-install` with the pinned Node 24.21.0 produced
`dist/opencode-darwin-arm64/bin/opencode` (`0.0.0-refactor/reasoning-rewrite-engine-202610081437`); the build's own
`--version` smoke passed. The cross-platform native-package install step was skipped: it fails locally on a frozen
lockfile, and this change does not touch dependencies. Installing over `/usr/local/bin/opencode` requires `sudo`, which
was not available, so **no installed-binary test is claimed**.

Each run used temporary config/data/state/cache/home, a synthetic project, one long-lived `serve` with two
`run --attach` turns, the configured `local-proxy-compatible` provider through existing credential references, main
model `deepseek`, and a local relay recording only role sequences and reasoning prefixes.

- **Feature on, organizer `deepseek` (`none`), default 3 s window**: step-1 reasoning (790 chars) organized in 503 ms
while the tool ran, adopted in 5 ms, and the **step-2 request carried the rewrite**. The final step's English reasoning
(215 chars) became 87 Chinese chars in 546 ms and the **next turn's request carried it**. 0 server errors.
- **Feature off**: no organizer calls; both requests resent the original reasoning; 0 server errors.

Earlier exploratory runs with `glm-flash` (no `none` variant, `low` effort) took 0.9–4.2 s per organizer call and, in
one run, made a step wait for the organizer inside the window. With a `none`-capable organizer the calls finish well
inside tool execution time.

### DEBUG regression and barrier visibility

A DEBUG-level live matrix on the installed binary (DeepSeek `none` and glm-flash `low` organizers at the default window,
a zero window, feature off) showed the intended behavior in every case: finished rewrites adopted before the next
request; a glm-flash call that missed the window and a call still running at a zero window were sealed and the original
stayed in every later request; no organizer calls when off; no server errors. Sealing was only inferable from missing
outcome lines, so each barrier now logs `reasoning rewrite sealed` (info) when it seals work and
`reasoning rewrite barrier` (debug) otherwise, with job, adopted and sealed counts and the wait time.

### Observations

- In the smoke run the organizer judged the step-1 reasoning entirely noise. It restated the request, listed three
error kinds and deliberated about formatting; the list reappeared in the visible answer. This is model judgment, not a
mechanical failure, but it shows that the organizer can drop short plans.
- A short-lived `opencode run` process ends before the final step's job; the barrier semantics need a long-lived process
(TUI, server), as intended.
Loading
Loading