Skip to content

Long conversations shed old exchanges and recover from context refusals - #973

Merged
WaylandYang merged 5 commits into
deeplethe:devfrom
Maya-Kid:fix/chat-context-budget
Sep 27, 2026
Merged

WaylandYang merged 5 commits into
deeplethe:devfrom
Maya-Kid:fix/chat-context-budget

Conversation

@Maya-Kid

@Maya-Kid Maya-Kid commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

A conversation with several large answers could exceed the model's context window on every new question. The refusal was treated as unsupported tools, so the compatibility retry and RAG fallback sent the same oversized history again.

This implements the decision in #964:

  • Recognize a context-window refusal in utopia-llm, separately from tool incompatibility and completion ceilings: the context_length_exceeded code first, then the agreed wording. Confirmed against DeepSeek on 2026-09-27 with a 1.2M-token request: 400 with only the generic code invalid_request_error, and the sentence "This model's maximum context length is 1048576 tokens. However, you requested 1200006 tokens (…)". It is recognized by wording and the window is read from it.
  • Keep the 20-message window and add a 32,000-character history budget. The previous turn's tool exchange is the first thing dropped, then the oldest whole exchanges, including failed questions and stopped answers. An orphan answer at the start of the 20-message window is not sent, which changes what long conversations send. Stored conversation messages never change.
  • On a context refusal, halve the remaining history and retry once for the entire user turn. Keep the system prompt, current question and current tool results intact; do not execute tools again. Tool calling, RAG fallback and the reserved final answer share that window and recovery allowance. When nothing can be dropped, the request is not sent again.
  • Remember a stated context window on the client's clones and across later turns in the same workspace. Changing chat settings resets it. A second refusal returns context_too_long, with English and Chinese wording suggesting a new conversation.

The history budget is half the stated window, counted in characters, and never below 2,000 or above 32,000; a stated window below 1,024 tokens is ignored, as stated_completion_ceiling ignores small numbers. This is a heuristic, not exact token accounting. A question or current evidence that is itself too large stays intact and can still fail. Limits are remembered in process, not persisted across restarts.

Why the exchange goes first: one get_document result can be 24,000 characters, so two of them exceed the default budget. Tied to the last answer, dropping them meant dropping every exchange, and a follow-up such as "make it shorter" went out with no history at all. The exchange is evidence the last answer already digested; the answers are what the person read.

The dated revision in 0042 (one-line status, a sentence in decision 2 and a section before Not done) and the chat design document describe these boundaries.

Verification on dev 0fbcd38, with Rust 1.98.1 and PostgreSQL 16 with pgvector, in containers on a Linux host:

  • Control: the previous head plus the new handler tests. an_evidence_heavy_turn_still_leaves_its_question_and_answer_for_the_follow_up fails there (the follow-up carries no history), even_an_oversized_current_question_is_never_cut fails (the same request is sent a second time), and overflow_trims_once_and_the_next_turn_remembers_the_window fails (the budget is the token count, not half of it). (logs 480–482)
  • This branch: utopia-llm 42; api:: 211; cargo fmt --all --check and cargo clippy --locked --workspace --all-targets -- -D warnings clean; UTOPIA_TEST_REQUIRE_PDFTOTEXT=1 cargo test --locked --workspace with UTOPIA_TEST_REQUIRE_DB=1: 1,240 passed, 0 failed, 5 pre-existing ignored; cargo build --locked --workspace; Node 20 / pnpm 10.2.1 frozen install, 197 frontend tests and the production build (logs 483–493).
  • Also on a Mac with the same toolchain and a scratch PostgreSQL 16: utopia-llm 42, the 16 context tests, api:: 211, fmt and clippy clean.

Fixes #964.

@WaylandYang WaylandYang left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you. The shape follows the decision on #964: the recogniser sits beside the completion ceiling and neither is mistaken for the other, trimming takes whole exchanges and counts characters, the retry is once per turn across tool calling, the RAG fallback and the final answer, tools are not run twice, and nothing stored changes. One thing blocks it and four are small.

1. Blocking: a large previous tool exchange evicts the answer just given.
Window::size counts the previous turn's tool exchange against the budget and keeps_exchange ties it to the last answer. One get_document result can be 24,000 characters, so two of them exceed the 32,000 default. trim then moves start past everything, and a follow-up such as "make it shorter" is sent with no history at all, on any model. Your test the_previous_tool_exchange_costs_budget_and_leaves_no_orphan_results asserts exactly that outcome.
Make the exchange droppable on its own and drop it first, before its question and answer. apply then has to remove it from the middle rather than as a prefix. Please add a test: after an evidence-heavy turn, the follow-up still carries the previous question and answer.

2. Do not retry when nothing can be dropped. With an empty history and an oversized question, recover returns true and sends the identical request again. Return false when trim does not move start, and pin one request in even_an_oversized_current_question_is_never_cut.

3. Give the learned window a floor. It only ever lowers and lasts until settings change or a restart, so one message stating a small number starves history. Ignore a stated window below 1024 tokens, as stated_completion_ceiling does, and never let the budget fall under 2,000 characters. Also use half the stated tokens as the character budget rather than the tokens themselves: a Chinese character is about one token, so min(32000, tokens) overshoots on a small window.

4. Comments in Rust are Chinese here, why-comments as you write them elsewhere. This covers chat_context.rs, the header of chat_context_tests.rs, the block in llm_util.rs, the field in state.rs and the new ones in lib.rs. context_window wants a comment like its sibling completion_ceiling.

5. Rebase onto dev and redo the two record edits. #974 cut every status line to one line and made the README an index that repeats only the status word.

  • 0042's status line becomes: Implemented (#548) · revised 2026-09-26 (#937) · revised 2026-09-27 (#973)
  • Your three paragraphs become a body section before Not done, with one dated sentence in decision 2 where the overturned claim stands, pointing to it.
  • Drop the README hunk; the status word does not change.
  • Leave Status history alone.

Smaller, as you see fit:

  • failure() is shared, so extraction and governance now word an oversized chunk as a conversation that is too long and lose the HTTP status. A neutral sentence that keeps the status would serve both.
  • The cfg(test) answer() in chat_finalization.rs is kept only for old tests; call the new one from them.
  • mod context_tests belongs in the alphabetical path block of chat_empty_reply_tests.rs.
  • Dropping the leading orphan answer of a 20-message window is right; please say so in the description, since it changes what long conversations send.
  • Wordings from other vendors are worth adding only where you can confirm the real message. SiliconFlow matters most, since the benches run on it.

Thanks again; this is close.

Signed-off-by: dada-yan <BinjunYann@gmail.com>
…fusal only when history shrank, and floor the learned window

Signed-off-by: dada-yan <BinjunYann@gmail.com>
…its decision 2

Signed-off-by: dada-yan <BinjunYann@gmail.com>
@Maya-Kid
Maya-Kid force-pushed the fix/chat-context-budget branch from 72fe3b2 to 6d499e5 Compare September 27, 2026 09:47
@Maya-Kid

Copy link
Copy Markdown
Contributor Author

Thank you. All five are in, and the smaller ones too.

  1. The exchange goes first, on its own. Window has an exchange_kept flag independent of start: over budget, the previous turn's tool exchange is dropped before any question or answer, then the oldest exchanges. apply regenerates the kept history and swaps it in for rig's copy, so the exchange is removed from the middle rather than as a prefix. New handler test an_evidence_heavy_turn_still_leaves_its_question_and_answer_for_the_follow_up: after two 24,000-character get_document results, "make it shorter" is sent with the previous question and answer and no tool messages; a 100-character exchange is still sent. The unit test that asserted the old outcome now asserts the new one.
  2. No retry when nothing can be dropped. recover compares (start, exchange kept) before and after trimming and returns false when nothing moved. even_an_oversized_current_question_is_never_cut now pins one request.
  3. A floor, and half the tokens. stated_context_window ignores a number below 1,024 tokens, and history_char_budget is half the stated tokens as characters, clamped to 2,000–32,000. The handler test states 4,096 tokens and checks that the next turn carries question 5 but not question 4 under the 2,048-character budget.
  4. Comments are Chinese why-comments in chat_context.rs, the test header, llm_util.rs, state.rs and the new items in lib.rs; context_window has one beside completion_ceiling.
  5. Rebased onto dev 0fbcd38. 0042's status line is Implemented (#548) · revised 2026-09-26 (#937) · revised 2026-09-27 (#973), the three paragraphs are a section before Not done, decision 2 has the dated sentence pointing at it, the README hunk is gone and Status history is untouched.

Smaller: ContextTooLong keeps the status and reads "LLM request failed (400 Bad Request): the prompt exceeds the model's context window: …", so extraction and governance are not told a conversation is too long; the cfg(test) answer is gone and the old test calls answer_with_context; context_tests sits in the alphabetical block; the description says that the leading orphan answer of a 20-message window is dropped.

On vendor wordings: I confirmed DeepSeek's today with a 1.2M-token request. It is a 400 with only the generic code invalid_request_error, and the message "This model's maximum context length is 1048576 tokens. However, you requested 1200006 tokens (1200005 in the messages, 1 in the completion). Please reduce the length of the messages or completion." It is recognized by wording and the window is read from it; the code check does not fire for it, which is why both stay. SiliconFlow's docs give only the shape of its error body ({code, message, data}), not the sentence it uses for an oversized prompt, so I could not confirm it and added nothing; if you have that sentence, it is one entry in context_wording.

Verification, in containers on a Linux host (Rust 1.98.1, PostgreSQL 16):

  • Control, the previous head plus the new handler tests: the three that pin the changes fail there and the other four pass. an_evidence_heavy_turn_still_leaves_its_question_and_answer_for_the_follow_up fails because the follow-up carries no previous question; even_an_oversized_current_question_is_never_cut fails because the same request is sent a second time; overflow_trims_once_and_the_next_turn_remembers_the_window fails because the next turn is given 4,096 characters, not 2,048.
  • This branch: utopia-llm 42, api:: 211, fmt and clippy clean, the workspace 1,240 passed / 0 failed / 5 ignored with a real PostgreSQL and Poppler, the build, and the web install, 197 tests and build.

Signed-off-by: dada-yan <BinjunYann@gmail.com>

# Conflicts:
#	crates/utopia-server/src/llm_util.rs

@WaylandYang WaylandYang left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for turning all of it around, and for confirming the wording against a real DeepSeek refusal. I re-read the new apply and ran it: the boundary is found by count through the same functions that build rig's list, so a current question that repeats an earlier one is safe, several tool rounds survive in order, the exchange leaves or stays whole, and repeated calls give the same list. With dev merged in (#977, #978, #979) the server suite passes 509 tests, utopia-llm 42 and the web 205.

Two notes for a follow-up, neither blocking:

  • apply trusts that rig's list holds the whole prior history. If it ever held fewer messages, the current question would be swallowed silently. A check that falls back to the unmodified list would make that impossible.
  • A refusal caused by our own max_tokens on a small-window model is now worded as a conversation that is too long. It failed before too, so only the wording is off.

Landing it. #964 closes with it. Thanks again.

@WaylandYang WaylandYang left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maintainer review: sound, landing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

A conversation longer than the model's context window fails every turn as "the endpoint refused the request"

2 participants