Skip to content

fix: keep active request with prompt-guided tool result - #30

Merged
senamakel merged 2 commits into
tinyhumansai:mainfrom
senamakel:transcript-hi-e2e-live
Sep 25, 2026
Merged

senamakel merged 2 commits into
tinyhumansai:mainfrom
senamakel:transcript-hi-e2e-live

Conversation

@senamakel

@senamakel senamakel commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Keep a text-dialect tool result and the active user request in the same outgoing user message.
  • Avoid creating a separate trailing user turn that makes models restart tool discovery.
  • Update prompt-tool tests for the combined continuation, idempotency, and preservation of non-text result blocks.

Context

OpenHuman's live Gmail check exposed a gap in #29: the separate trailing request made DeepSeek search for the same Gmail action repeatedly. This change repairs that prompt shape. It does not claim to resolve the separate managed DeepSeek stale-greeting behavior; in a captured live run the backend received the correct email request and still returned the earlier greeting. That provider/routing issue remains under investigation.

Verification

  • cargo test --manifest-path Cargo.toml -p tinyinference-llm prompt_tools::test --lib (18 passed)
  • cargo fmt --manifest-path Cargo.toml --all --check
  • TinyAgents agent-loop E2E test in dependent PR exercises greeting → tool_search → deferred Gmail tool → answer.

Summary by CodeRabbit

  • Improvements
    • When continuing after a tool result, the latest request is now included in the existing conversation turn along with instructions to avoid repeating prior work.
    • Existing content in that turn is preserved.

@tinysweeper

tinysweeper Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

This pull request fixes the `anchor_user_request_after_tool_result` function to keep the active user request and tool result in a single user turn, preventing models from restarting tool discovery. The prior logic error (attempting to mutate a tool result as a user message) is corrected, and non-text content blocks are preserved. The change is safe to merge. All prior findings are resolved. _Memory was unavailable (model: cortex: v1/recall: timed out after 10s), so this review ran without it._

State: Ready for maintainer review
Priority: none
Reviewed head: 6aff95c54935
Updated: 1790339378 (Unix time)

Review snapshot

Change surface Files Review signal Count
Production 1 Active findings 0
Tests 1 Noted findings 0
Documentation 0 Resolved findings 12
Configuration 0 Pending checks/questions 0

Completeness: Complete
Test assessment: Test coverage is assessed from changed tests and lane evidence; execution is not claimed without trusted check data.

What changed

The change modifies `anchor_user_request_after_tool_result` in `crates/tinyinference-llm/src/prompt_tools/mod.rs` to prepend a text block containing the active user request to the terminal user-role tool result message instead of appending a separate user message. It preserves all existing content blocks. The test `tool_continuation_anchors_the_latest_request_after_search_results` is updated to expect the merged message, and a new test `tool_continuation_preserves_non_text_result_blocks` is added to verify non-text block preservation.

Features

  • Modified — anchor_user_request_after_tool_result: Keeps the active user request and tool result in a single user turn to prevent models from restarting tool discovery, while preserving non-text content blocks. This fixes a no-op behavior from the previous commit. (crates/tinyinference-llm/src/prompt_tools/mod.rs#pub fn anchor_user_request_after_tool_result(messages: &[Message]) -> Vec<Messag)

Tests

  • modification — Updated test `tool_continuation_anchors_the_latest_request_after_search_results` to expect the last message to contain both the user request and tool result, with a 'Do not repeat' instruction, and the message count unchanged.: Test validates the corrected merging behavior. (crates/tinyinference-llm/src/prompt_tools/test.rs#fn tool_continuation_anchors_the_latest_request_after_search_results() {)

Findings

No active actionable findings.

Resolved this pass

  • Mutate the terminal tool result instead of matching it as a user
  • Preserve non-text tool-result content
  • Modify the last user message instead of the tool result
  • Mutate the terminal tool result instead of matching it as a user
  • medium — Preserve non-text tool-result content
  • Modify the last user message instead of the tool result
  • Mutate the terminal tool result instead of matching it as a user
  • Preserve non-text tool-result content
  • Modify the last user message instead of the tool result
  • Mutate the terminal tool result instead of matching it as a user
  • Preserve non-text tool-result content
  • Modify the last user message instead of the tool result

Before merge

None.

How this fits together

flowchart LR
  n0["anchor_user_request_after_tool_result<br/>changed"]:::changed
  n1["ensure_resolvable_user_turn<br/>changed"]:::changed
  n2["...s_the_latest_request_after_search_results<br/>changed"]:::changed
  n3["...thout_a_real_user_request_stays_unchanged<br/>changed"]:::changed
  n4["system"]:::impacted
  n5["user"]:::impacted
  n6["coalesce_tool_results"]:::impacted
  n7["recover_tool_calls"]:::impacted
  n8["new"]:::impacted
  n1 -->|calls| n5
  n2 -->|calls| n0
  n2 -->|tests| n0
  n2 -->|calls| n4
  n2 -->|tests| n4
  n2 -->|calls| n5
  n2 -->|tests| n5
  n2 -->|calls| n6
  n2 -->|tests| n6
  n3 -->|calls| n4
  n3 -->|tests| n4
  n3 -->|calls| n6
  n3 -->|tests| n6
  n6 -->|calls| n8
  n7 -->|calls| n8
  classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
  classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
  classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
  classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading
Agent review details

critique

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The revision keeps the continuation in a single user turn and preserves existing non-text blocks, fixing the previously reported continuation-shape issues. The change looks safe to merge. _Memory was unavailable (model: cortex: v1/recall: timed out after 10s), so this review ran without it._

security

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The continuation anchoring change now keeps the latest request and completed tool result in one user turn while preserving all original content blocks. The previously reported issues are fixed, and the change is safe to merge. (1 earlier finding(s) still open) _Memory was unavailable (model: cortex: v1/recall: timed out after 10s), so this review ran without it._

tests

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: _Memory was unavailable (model: cortex: v1/recall: timed out after 10s), so this review ran without it._

commits

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: Nothing sensitive found in what this pull request commits.

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: Fix the prompt-guided tool continuation to keep the active request in the same user message as the tool result, preventing models from restarting tool discovery. All prior findings are resolved by this revision. _Memory was unavailable (model: cortex: v1/recall: timed out after 10s), so this review ran without it._

e2e

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: No end-to-end harness in this repository: no e2e test files and no e2e workflow.
Evidence and run details
  • Models: ladder/vectors, gpt-5.6-luna, deepseek/deepseek-v4-flash
  • Spend: $0.005004
  • Tokens: 97748 input · 10792 output · 22426 cached · 358 embedding
Head State Pass summary
71d3d2d350fb changes requested 3 active finding(s), 0 resolved finding(s) (at 1790339023)
6aff95c54935 ready for maintainer review 0 active finding(s), 12 resolved finding(s) (at 1790339378)

tinysweeper 0.1.0

@coderabbitai

coderabbitai Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: fba03287-85f4-4950-b15c-c7ef895d9b4a

📥 Commits

Reviewing files that changed from the base of the PR and between 6ed7dae and 6aff95c.

📒 Files selected for processing (2)
  • crates/tinyinference-llm/src/prompt_tools/mod.rs
  • crates/tinyinference-llm/src/prompt_tools/test.rs

Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

The anchoring function now adds the active request and continuation instructions to the existing terminal tool-result user turn. Tests check the unchanged message count, the continuation instruction, and preservation of an existing JSON content block.

Changes

Tool-result request anchoring

Layer / File(s) Summary
Update anchoring behavior
crates/tinyinference-llm/src/prompt_tools/mod.rs, crates/tinyinference-llm/src/prompt_tools/test.rs
The function prepends continuation instructions and the active request to the terminal user turn. Tests verify the message count remains unchanged and existing JSON content is preserved.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Bug fix

Merge Risk: ⚪ Minimal · up to 6aff9

The active request remains with the tool result without accumulating duplicate instructions across translations. No actionable merge risk remains beyond normal checks.

Architecture Summary

Architecture risk: 🟡 Medium · up to 6aff9

The change affects 1 system.

Changed systems: crates

Architecture concerns
No architecture-level concerns identified.

Review details

Systems and components

  • observed — crates (service) was modified; 2 changed files map to changed impact.

Before / after behavior

  • observed — Modified behavior in crates/tinyinference-llm/src/prompt_tools/mod.rs: The documentation now specifies that the active request and terminal tool result share one user turn, rather than describing the request as a separate message after the result.
  • observed — Modified behavior in crates/tinyinference-llm/src/prompt_tools/mod.rs: Instead of appending a new user message containing the latest request, the function prepends continuation instructions and the request text to the existing terminal user turn. This keeps the result and request in one turn and preserves the turn’s existing content blocks.
  • observed — Modified behavior in crates/tinyinference-llm/src/prompt_tools/test.rs: The test now expects anchoring to keep the message count unchanged and checks the final message, replacing the previous expectation that anchoring appended a message and checked the preceding one.
  • observed — Modified behavior in crates/tinyinference-llm/src/prompt_tools/test.rs: The test adds an assertion that the anchored final message contains “Do not repeat.”

Reliability and maintainability

  • inferred — Risk-relevant change factors for crates: blast_radius_2; direct_dependents_2
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: keeping the active request with the prompt-guided tool result.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 5 functions across 2 files.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

A rabbit reads the tool-result turn,
Then tucks the request where both belong.
“Do not repeat,” the prompt now says,
JSON stays safe along its way.
One turn holds the work to come,
And I hop off, the review done.

Comment @coderabbitai help to get the list of available commands.

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes: 2 lane(s) blocking, worst finding is critical.

Fix or reply to the findings below and push. The next review clears this automatically once they are gone — you should not need to dismiss anything by hand.

             $0.0042 · 64,228 in / 11,735 out · 7,251 cached (11%) · ladder/vectors, gpt-5.6-luna, deepseek/deepseek-v4-flash · 238 embedded
critique:    $0.0009 · 16,869 in / 2,586 out  · 0 cached (0%)      · gpt-5.6-luna, deepseek/deepseek-v4-flash
security:    $0.0007 · 28,732 in / 1,963 out  · 1,875 cached (7%)  · gpt-5.6-luna
tests:       $0.0012 · 12,660 in / 1,382 out  · 3,072 cached (24%) · deepseek/deepseek-v4-flash
description: $0.0008 · 3,762 in  / 3,310 out  · 2,304 cached (61%) · deepseek/deepseek-v4-flash

Comment thread crates/tinyinference-llm/src/prompt_tools/mod.rs
Comment thread crates/tinyinference-llm/src/prompt_tools/mod.rs Outdated

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The previously-blocking findings are resolved. Clearing the changes request.

             $0.0050 · 97,748 in / 10,792 out · 22,426 cached (23%) · ladder/vectors, gpt-5.6-luna, deepseek/deepseek-v4-flash · 358 embedded
critique:    $0.0008 · 30,912 in / 1,459 out  · 2,119 cached (7%)   · gpt-5.6-luna
security:    $0.0007 · 30,424 in / 820 out    · 1,875 cached (6%)   · gpt-5.6-luna
tests:       $0.0022 · 26,720 in / 5,383 out  · 16,128 cached (60%) · deepseek/deepseek-v4-flash
description: $0.0006 · 4,577 in  / 1,858 out  · 2,304 cached (50%)  · deepseek/deepseek-v4-flash

@senamakel
senamakel merged commit 8498950 into tinyhumansai:main Sep 25, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant