Repository navigation
fix(openai): use word count, not char count, when truncating completion prompts - #53
Merged
tylermenezes merged 1 commit intoSep 18, 2026
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Detail bug report: View on Detail
Bug
textToCompletionPrompt(src/openai/format.ts) truncates oversized standup text so the final prompt fits the model's 2049-token budget. Itswhileloop computed theArray.slicekeep-count fromtruncatedText.length— the character count — instead of the word count (truncatedText.split(' ').length).Because the character count always exceeds the word-array length,
slice(0, hugeIndex)returned the array unchanged, the rebuilt string was byte-identical, the token count never decreased, and thewhilespun forever. Any standup text exceeding ~2049 tokens sent the loop into a synchronous infinite loop that blocked the single Node process running every automation cron — no crash, no throw, no Fly restart, just a total cron blackout (no Slack reminders, email, auto-scoring, standup import, etc.) until manual restart, recurring daily.The only unguarded caller is
aiTrain(the daily 06:20 cron); the scoring path (getProjectStandupScore) short-circuits at ≤1000 chars before reaching the truncator, so it never tripped the bug.Fix
Compute the slice end from the word array's own length — the unit being sliced — not the character count:
This is the root-cause fix: while the loop body runs, the overflow term is
>= 1, so at least one word is dropped each iteration and the word count strictly decreases — the loop now converges in at mostwords.lengthiterations for any input shape (prose, code, stack traces, no-space, multi-space).I did not add the optional caller-side length filter on
aiTrain'sfindMany; the report flags it as defense-in-depth that would also drop valid token-safe standups, and the truncator fix alone fully removes the bug.Testing
Added
src/openai/format.test.ts(offline, using only the repo's pinnedgpt-tokenizer) covering:system/user/assistant/userorder with the correct system and assistant priming content.BINARY_CLASSIFICATION_PROMPTSfor bothVagueandWorkload.getTrainingExampleproduces the correctyes/noassistant completion for each(modelType, rating)cell (Vague/1→yes, Vague/3→no, Workload/3→yes, Workload/1→no).Verification:
tsc --skipLibCheck --noEmit) and the new unit tests both pass; tests pass under bothtsxandts-node.tests/testSlackReporting.ts,src/automation/tasks/syncAlumniInteractions.test.ts) still pass — no regression.@codeday/eslint-typescript-configpulls an@typescript-eslint/parserincompatible withtypescript@5.2.2, which emits aDeprecationError: 'originalKeywordKind'parse failure at line 0 of every.tsfile (including untouched files). This is a pre-existing toolchain issue, independent of this change; typecheck is the meaningful gate.Automatic Fixes PRs can be configured here.