fix: keep active request with prompt-guided tool result - #30
Conversation
Tiny Sweeper reviewThis pull request fixes the `anchor_user_request_after_tool_result` function to keep the active user request and tool result in a single user turn, preventing models from restarting tool discovery. The prior logic error (attempting to mutate a tool result as a user message) is corrected, and non-text content blocks are preserved. The change is safe to merge. All prior findings are resolved. _Memory was unavailable (model: cortex: v1/recall: timed out after 10s), so this review ran without it._ State: Ready for maintainer review Review snapshot
Completeness: Complete What changedThe change modifies `anchor_user_request_after_tool_result` in `crates/tinyinference-llm/src/prompt_tools/mod.rs` to prepend a text block containing the active user request to the terminal user-role tool result message instead of appending a separate user message. It preserves all existing content blocks. The test `tool_continuation_anchors_the_latest_request_after_search_results` is updated to expect the merged message, and a new test `tool_continuation_preserves_non_text_result_blocks` is added to verify non-text block preservation. Features
Tests
FindingsNo active actionable findings. Resolved this pass
Before mergeNone. How this fits togetherflowchart LR
n0["anchor_user_request_after_tool_result<br/>changed"]:::changed
n1["ensure_resolvable_user_turn<br/>changed"]:::changed
n2["...s_the_latest_request_after_search_results<br/>changed"]:::changed
n3["...thout_a_real_user_request_stays_unchanged<br/>changed"]:::changed
n4["system"]:::impacted
n5["user"]:::impacted
n6["coalesce_tool_results"]:::impacted
n7["recover_tool_calls"]:::impacted
n8["new"]:::impacted
n1 -->|calls| n5
n2 -->|calls| n0
n2 -->|tests| n0
n2 -->|calls| n4
n2 -->|tests| n4
n2 -->|calls| n5
n2 -->|tests| n5
n2 -->|calls| n6
n2 -->|tests| n6
n3 -->|calls| n4
n3 -->|tests| n4
n3 -->|calls| n6
n3 -->|tests| n6
n6 -->|calls| n8
n7 -->|calls| n8
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Agent review detailscritique
security
tests
commits
description
e2e
Evidence and run details
|
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (2)
Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review. 📝 WalkthroughWalkthroughThe anchoring function now adds the active request and continuation instructions to the existing terminal tool-result user turn. Tests check the unchanged message count, the continuation instruction, and preservation of an existing JSON content block. ChangesTool-result request anchoring
Priority: ⬇️ Low Estimated code review effort: 2 (Simple) | ~10 minutes Change: Bug fix Merge Risk: ⚪ Minimal · up to The active request remains with the tool result without accumulating duplicate instructions across translations. No actionable merge risk remains beyond normal checks. Architecture SummaryArchitecture risk: 🟡 Medium · up to The change affects 1 system. Changed systems: Architecture concerns Review detailsSystems and components
Before / after behavior
Reliability and maintainability
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
A rabbit reads the tool-result turn, Comment |
There was a problem hiding this comment.
Requesting changes: 2 lane(s) blocking, worst finding is critical.
Fix or reply to the findings below and push. The next review clears this automatically once they are gone — you should not need to dismiss anything by hand.
$0.0042 · 64,228 in / 11,735 out · 7,251 cached (11%) · ladder/vectors, gpt-5.6-luna, deepseek/deepseek-v4-flash · 238 embedded
critique: $0.0009 · 16,869 in / 2,586 out · 0 cached (0%) · gpt-5.6-luna, deepseek/deepseek-v4-flash
security: $0.0007 · 28,732 in / 1,963 out · 1,875 cached (7%) · gpt-5.6-luna
tests: $0.0012 · 12,660 in / 1,382 out · 3,072 cached (24%) · deepseek/deepseek-v4-flash
description: $0.0008 · 3,762 in / 3,310 out · 2,304 cached (61%) · deepseek/deepseek-v4-flash
There was a problem hiding this comment.
The previously-blocking findings are resolved. Clearing the changes request.
$0.0050 · 97,748 in / 10,792 out · 22,426 cached (23%) · ladder/vectors, gpt-5.6-luna, deepseek/deepseek-v4-flash · 358 embedded
critique: $0.0008 · 30,912 in / 1,459 out · 2,119 cached (7%) · gpt-5.6-luna
security: $0.0007 · 30,424 in / 820 out · 1,875 cached (6%) · gpt-5.6-luna
tests: $0.0022 · 26,720 in / 5,383 out · 16,128 cached (60%) · deepseek/deepseek-v4-flash
description: $0.0006 · 4,577 in / 1,858 out · 2,304 cached (50%) · deepseek/deepseek-v4-flash
Summary
Context
OpenHuman's live Gmail check exposed a gap in #29: the separate trailing request made DeepSeek search for the same Gmail action repeatedly. This change repairs that prompt shape. It does not claim to resolve the separate managed DeepSeek stale-greeting behavior; in a captured live run the backend received the correct email request and still returned the earlier greeting. That provider/routing issue remains under investigation.
Verification
cargo test --manifest-path Cargo.toml -p tinyinference-llm prompt_tools::test --lib(18 passed)cargo fmt --manifest-path Cargo.toml --all --checkSummary by CodeRabbit