Skip to content

fix(openai): use word count, not char count, when truncating completion prompts - #53

Merged
tylermenezes merged 1 commit into
mainfrom
detail/bug-fix/fix-openai-use-word-count-not-char-count-when-trun-024870
Sep 18, 2026
Merged

tylermenezes merged 1 commit into
mainfrom
detail/bug-fix/fix-openai-use-word-count-not-char-count-when-trun-024870

Conversation

@detail-app

@detail-app detail-app Bot commented Sep 18, 2026

Copy link
Copy Markdown
Contributor

Detail bug report: View on Detail

Bug

textToCompletionPrompt (src/openai/format.ts) truncates oversized standup text so the final prompt fits the model's 2049-token budget. Its while loop computed the Array.slice keep-count from truncatedText.length — the character count — instead of the word count (truncatedText.split(' ').length).

Because the character count always exceeds the word-array length, slice(0, hugeIndex) returned the array unchanged, the rebuilt string was byte-identical, the token count never decreased, and the while spun forever. Any standup text exceeding ~2049 tokens sent the loop into a synchronous infinite loop that blocked the single Node process running every automation cron — no crash, no throw, no Fly restart, just a total cron blackout (no Slack reminders, email, auto-scoring, standup import, etc.) until manual restart, recurring daily.

The only unguarded caller is aiTrain (the daily 06:20 cron); the scoring path (getProjectStandupScore) short-circuits at ≤1000 chars before reaching the truncator, so it never tripped the bug.

Fix

Compute the slice end from the word array's own length — the unit being sliced — not the character count:

while (truncatedTextTokenLength > MODEL_MAX_TOKENS) {
  const words = truncatedText.split(' ');
  truncatedText = words
    .slice(0, words.length - Math.ceil((truncatedTextTokenLength - MODEL_MAX_TOKENS) / 2))
    .join(' ');
  truncatedTextTokenLength = encode(truncatedText).length;
}

This is the root-cause fix: while the loop body runs, the overflow term is >= 1, so at least one word is dropped each iteration and the word count strictly decreases — the loop now converges in at most words.length iterations for any input shape (prose, code, stack traces, no-space, multi-space).

I did not add the optional caller-side length filter on aiTrain's findMany; the report flags it as defense-in-depth that would also drop valid token-safe standups, and the truncator fix alone fully removes the bug.

Testing

Added src/openai/format.test.ts (offline, using only the repo's pinned gpt-tokenizer) covering:

  • Infinite-loop regression: oversized input (~4200 tokens) returns within a 5s wall-clock budget and the result fits the token budget. A regression to the char-count bug would hang the process before this assertion could be reached.
  • No over-truncation: under-limit input is passed through verbatim.
  • Prompt structure: 4 messages in system/user/assistant/user order with the correct system and assistant priming content.
  • Classification question per model type: the final user message matches BINARY_CLASSIFICATION_PROMPTS for both Vague and Workload.
  • Rating → completion matrix: getTrainingExample produces the correct yes/no assistant completion for each (modelType, rating) cell (Vague/1→yes, Vague/3→no, Workload/3→yes, Workload/1→no).

Verification:

  • Typecheck (tsc --skipLibCheck --noEmit) and the new unit tests both pass; tests pass under both tsx and ts-node.
  • Pre-existing suites (tests/testSlackReporting.ts, src/automation/tasks/syncAlumniInteractions.test.ts) still pass — no regression.
  • Lint could not be run: the repo's pinned @codeday/eslint-typescript-config pulls an @typescript-eslint/parser incompatible with typescript@5.2.2, which emits a DeprecationError: 'originalKeywordKind' parse failure at line 0 of every .ts file (including untouched files). This is a pre-existing toolchain issue, independent of this change; typecheck is the meaningful gate.

Automatic Fixes PRs can be configured here.

@detail-app
detail-app Bot requested a review from tylermenezes September 18, 2026 02:53
@tylermenezes
tylermenezes merged commit 7f5e346 into main Sep 18, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant