Repository navigation
Long conversations shed old exchanges and recover from context refusals - #973
Conversation
WaylandYang
left a comment
There was a problem hiding this comment.
Thank you. The shape follows the decision on #964: the recogniser sits beside the completion ceiling and neither is mistaken for the other, trimming takes whole exchanges and counts characters, the retry is once per turn across tool calling, the RAG fallback and the final answer, tools are not run twice, and nothing stored changes. One thing blocks it and four are small.
1. Blocking: a large previous tool exchange evicts the answer just given.
Window::size counts the previous turn's tool exchange against the budget and keeps_exchange ties it to the last answer. One get_document result can be 24,000 characters, so two of them exceed the 32,000 default. trim then moves start past everything, and a follow-up such as "make it shorter" is sent with no history at all, on any model. Your test the_previous_tool_exchange_costs_budget_and_leaves_no_orphan_results asserts exactly that outcome.
Make the exchange droppable on its own and drop it first, before its question and answer. apply then has to remove it from the middle rather than as a prefix. Please add a test: after an evidence-heavy turn, the follow-up still carries the previous question and answer.
2. Do not retry when nothing can be dropped. With an empty history and an oversized question, recover returns true and sends the identical request again. Return false when trim does not move start, and pin one request in even_an_oversized_current_question_is_never_cut.
3. Give the learned window a floor. It only ever lowers and lasts until settings change or a restart, so one message stating a small number starves history. Ignore a stated window below 1024 tokens, as stated_completion_ceiling does, and never let the budget fall under 2,000 characters. Also use half the stated tokens as the character budget rather than the tokens themselves: a Chinese character is about one token, so min(32000, tokens) overshoots on a small window.
4. Comments in Rust are Chinese here, why-comments as you write them elsewhere. This covers chat_context.rs, the header of chat_context_tests.rs, the block in llm_util.rs, the field in state.rs and the new ones in lib.rs. context_window wants a comment like its sibling completion_ceiling.
5. Rebase onto dev and redo the two record edits. #974 cut every status line to one line and made the README an index that repeats only the status word.
- 0042's status line becomes: Implemented (#548) · revised 2026-09-26 (#937) · revised 2026-09-27 (#973)
- Your three paragraphs become a body section before Not done, with one dated sentence in decision 2 where the overturned claim stands, pointing to it.
- Drop the README hunk; the status word does not change.
- Leave Status history alone.
Smaller, as you see fit:
- failure() is shared, so extraction and governance now word an oversized chunk as a conversation that is too long and lose the HTTP status. A neutral sentence that keeps the status would serve both.
- The cfg(test) answer() in chat_finalization.rs is kept only for old tests; call the new one from them.
- mod context_tests belongs in the alphabetical path block of chat_empty_reply_tests.rs.
- Dropping the leading orphan answer of a 20-message window is right; please say so in the description, since it changes what long conversations send.
- Wordings from other vendors are worth adding only where you can confirm the real message. SiliconFlow matters most, since the benches run on it.
Thanks again; this is close.
Signed-off-by: dada-yan <BinjunYann@gmail.com>
…fusal only when history shrank, and floor the learned window Signed-off-by: dada-yan <BinjunYann@gmail.com>
…its decision 2 Signed-off-by: dada-yan <BinjunYann@gmail.com>
72fe3b2 to
6d499e5
Compare
|
Thank you. All five are in, and the smaller ones too.
Smaller: On vendor wordings: I confirmed DeepSeek's today with a 1.2M-token request. It is a 400 with only the generic code Verification, in containers on a Linux host (Rust 1.98.1, PostgreSQL 16):
|
Signed-off-by: dada-yan <BinjunYann@gmail.com> # Conflicts: # crates/utopia-server/src/llm_util.rs
WaylandYang
left a comment
There was a problem hiding this comment.
Thank you for turning all of it around, and for confirming the wording against a real DeepSeek refusal. I re-read the new apply and ran it: the boundary is found by count through the same functions that build rig's list, so a current question that repeats an earlier one is safe, several tool rounds survive in order, the exchange leaves or stays whole, and repeated calls give the same list. With dev merged in (#977, #978, #979) the server suite passes 509 tests, utopia-llm 42 and the web 205.
Two notes for a follow-up, neither blocking:
- apply trusts that rig's list holds the whole prior history. If it ever held fewer messages, the current question would be swallowed silently. A check that falls back to the unmodified list would make that impossible.
- A refusal caused by our own max_tokens on a small-window model is now worded as a conversation that is too long. It failed before too, so only the wording is off.
Landing it. #964 closes with it. Thanks again.
WaylandYang
left a comment
There was a problem hiding this comment.
Maintainer review: sound, landing.
A conversation with several large answers could exceed the model's context window on every new question. The refusal was treated as unsupported tools, so the compatibility retry and RAG fallback sent the same oversized history again.
This implements the decision in #964:
utopia-llm, separately from tool incompatibility and completion ceilings: thecontext_length_exceededcode first, then the agreed wording. Confirmed against DeepSeek on 2026-09-27 with a 1.2M-token request: 400 with only the generic codeinvalid_request_error, and the sentence "This model's maximum context length is 1048576 tokens. However, you requested 1200006 tokens (…)". It is recognized by wording and the window is read from it.context_too_long, with English and Chinese wording suggesting a new conversation.The history budget is half the stated window, counted in characters, and never below 2,000 or above 32,000; a stated window below 1,024 tokens is ignored, as
stated_completion_ceilingignores small numbers. This is a heuristic, not exact token accounting. A question or current evidence that is itself too large stays intact and can still fail. Limits are remembered in process, not persisted across restarts.Why the exchange goes first: one
get_documentresult can be 24,000 characters, so two of them exceed the default budget. Tied to the last answer, dropping them meant dropping every exchange, and a follow-up such as "make it shorter" went out with no history at all. The exchange is evidence the last answer already digested; the answers are what the person read.The dated revision in 0042 (one-line status, a sentence in decision 2 and a section before Not done) and the chat design document describe these boundaries.
Verification on dev
0fbcd38, with Rust 1.98.1 and PostgreSQL 16 with pgvector, in containers on a Linux host:an_evidence_heavy_turn_still_leaves_its_question_and_answer_for_the_follow_upfails there (the follow-up carries no history),even_an_oversized_current_question_is_never_cutfails (the same request is sent a second time), andoverflow_trims_once_and_the_next_turn_remembers_the_windowfails (the budget is the token count, not half of it). (logs 480–482)utopia-llm42;api::211;cargo fmt --all --checkandcargo clippy --locked --workspace --all-targets -- -D warningsclean;UTOPIA_TEST_REQUIRE_PDFTOTEXT=1 cargo test --locked --workspacewithUTOPIA_TEST_REQUIRE_DB=1: 1,240 passed, 0 failed, 5 pre-existing ignored;cargo build --locked --workspace; Node 20 / pnpm 10.2.1 frozen install, 197 frontend tests and the production build (logs 483–493).utopia-llm42, the 16 context tests,api::211, fmt and clippy clean.Fixes #964.