Skip to content

Add opt-in context-window truncation to the chat endpoint - #296

Closed
stikves wants to merge 1 commit into
apple:mainfrom
stikves:sukru/context-truncation
Closed

stikves wants to merge 1 commit into
apple:mainfrom
stikves:sukru/context-truncation

Conversation

@stikves

@stikves stikves commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds an opt-in truncation field to /v1/chat/completions so an over-length conversation can be gracefully fit to the model's context window instead of being rejected outright. The default is unchanged (reject with 400), so this is purely additive and eval-safe.

API

One field, three modes:

truncation Behavior
"off" (default, also "disabled"/"none") Reject an over-length prompt with 400 — unchanged behavior
"auto" Drop oldest whole messages (keeping the system message and the newest turn) until the prompt fits maxContextLength
positive integer N (e.g. 500) As "auto", but target a budget of min(N, maxContextLength) tokens
POST /v1/chat/completions
{ "messages": [ ... ], "truncation": "auto" }   // or "off", or 500

Invalid values (non-positive integer, unknown string) return 400.

Design

  • Truncation runs at the message boundary, before the chat template is applied, so it requires no engine or KV-cache changes and cannot affect attention correctness.
  • ContextTruncation.apply is a pure policy — the prompt renderer is injected as a closure — so it is unit-tested in isolation with no tokenizer or server state.
  • The system message is pinned and the newest turn is always kept; oldest droppable messages go first. A keep-last token trim is the final safety net so an opt-in request never errors, even if the pinned set alone overflows.
  • Both the streaming and non-streaming chat paths go through one shared resolvePromptTokens helper.
  • Scope: chat only. /v1/completions (the loglikelihood/eval path, which resets per prompt) is deliberately untouched — truncating there would corrupt scoring.

This is the mechanical, lossless-at-the-boundary layer of context-window management. It mirrors the truncate_prompt_tokens request extension found in other local serving engines and the stateless "truncate to fit" convention, without changing the default.

Testing

  • ContextTruncationTests — every mode (off fits/overflows, auto no-op/drops-oldest, tokensAt budget cap, safety-net trim).
  • ServerAPITruncationTests — decode of the union field (absent → off, string forms, positive integer, rejection of non-positive/unknown).
  • Full CoreAILMCommonTests target green (89 tests); swift build --product llm-server clean; swift format clean.

Add a `truncation` field to /v1/chat/completions requests:
- "off" (default): reject an over-length prompt with 400 (unchanged behavior)
- "auto": drop oldest whole messages, keeping the system message and the newest
  turn, until the prompt fits maxContextLength
- positive integer N: as auto, targeting a budget of min(N, maxContextLength)

Truncation runs at the message boundary before the chat template is applied, so
it needs no engine or KV-cache changes. A keep-last token trim is the final
safety net so an opt-in request never errors. /v1/completions is unchanged.

ContextTruncation.apply is a pure policy (render injected as a closure) with
unit tests for the decode and every mode; the two chat paths share a
resolvePromptTokens helper.
@stikves stikves self-assigned this Sep 27, 2026
@stikves stikves closed this Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant