Repository navigation
fix(a2a): inactivity_action events run to completion on client disconnect - #76
Conversation
…nect Rails fires-and-forgets inactivity_action events from a Sidekiq job that completes in ~80ms. The processor always detects "client disconnected" and was cancelling the LLM call before any message reached the lead. Add ignore_disconnect=True path to run_unless_client_disconnects so system-initiated events complete regardless of HTTP connection lifetime. Interactive bot conversations are unaffected (CRM-236 behaviour preserved).
Reviewer's guide (collapsed on small PRs)Reviewer's GuideThe PR prevents inactivity_action processing from being cancelled when Rails’ short-lived Sidekiq request disconnects: the disconnect helper now supports an opt-in bypass, and the A2A route enables it only for inactivity_action events, preserving cancellation for interactive conversations. Sequence diagram for inactivity action completion after client disconnectsequenceDiagram
participant Rails as Rails Sidekiq
participant A2A as A2A route
participant Helper as run_unless_client_disconnects
participant Agent as run_agent
participant Lead as WhatsApp lead
Rails->>A2A: handle_message_send(metadata)
A2A->>A2A: metadata.get(evoai_crm_event)
A2A->>Helper: run_unless_client_disconnects(ignore_disconnect=True)
A2A--xRails: HTTP connection disconnects
Helper->>Agent: await run_agent()
Agent->>Lead: Send inactivity action message
Agent-->>Helper: final_response
Helper-->>A2A: result
Flow diagram for event-specific disconnect handlingflowchart TD
A[handle_message_send] --> B{evoai_crm_event == inactivity_action}
B -->|Yes| C["run_unless_client_disconnects(ignore_disconnect=True)"]
B -->|No| D["run_unless_client_disconnects(ignore_disconnect=False)"]
C --> E[run_agent runs to completion]
D --> F{Client disconnects first}
F -->|Yes| G[Cancel agent execution]
F -->|No| E
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
There was a problem hiding this comment.
Hey - I've found 1 issue
Prompt for AI Agents
Please address the comments from this code review:
## Individual Comments
### Comment 1
<location path="src/utils/client_disconnect.py" line_range="131" />
<code_context>
task = asyncio.ensure_future(coro)
- if not cancel_on_disconnect_enabled():
+ if not cancel_on_disconnect_enabled() or ignore_disconnect:
return await task
</code_context>
<issue_to_address>
**nitpick:** The module-level documentation says the utility stops an agent once nobody is waiting for its answer, but the new `ignore_disconnect` path intentionally continues running after the caller disconnects. The documentation is now misleading for the supported inactivity-action behavior.
**Triggers:** When maintainers rely on the module docstring to understand the utility's behavior.
**Suggested fix:** Update the module-level docstring to document the intentional fire-and-forget exception for system-initiated work.
</issue_to_address>Sourcery assessment
Needs a human reviewer. If this exception is wrong, an inactivity-triggered agent run will continue after the client disconnects and may write records or perform downstream actions that the caller would previously have prevented by hanging up. Reverting stops future runs from being detached, but any side effects from runs already completed would need to be repaired or undone separately.
…get callers conversation_opened, conversation_resolved and webwidget_triggered are sent from AgentBotListener via HttpRequestJob (Sidekiq) — the job disconnects as soon as the HTTP handshake completes, not when the LLM finishes. Without this exemption CRM-236 would silently cancel those agent runs too. Live bot conversation events (message_created, message_updated) are excluded: their caller is the bot runtime, which holds the connection until the reply is sent — disconnect there is a real timeout that should cancel the run.
…t exception Sourcery nitpick on PR evolution-foundation#76: the module docstring said the utility always stops an agent once nobody is waiting for the answer, but ignore_disconnect now lets system-initiated work (e.g. inactivity_action) run to completion after the caller disconnects. Document the exception. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Root cause
Rails fires
inactivity_actionevents from a Sidekiq job that completes in ~80ms. Because the HTTP connection drops as soon as Sidekiq finishes dispatching,run_unless_client_disconnectsalways detected "client disconnected" and cancelled the LLM call before any message was sent to the lead. The execution record was already committed in Rails (rule marked done), but no WhatsApp message ever reached the lead.Fix
src/utils/client_disconnect.py— addsignore_disconnect: bool = Falsetorun_unless_client_disconnects. WhenTrue, the cancel-on-disconnect watcher is bypassed entirely and the coroutine runs to completion.src/api/a2a_routes.py— detectsmetadata.evoai_crm_event == "inactivity_action"at the call site and passesignore_disconnect=Truefor those events.Behaviour
inactivity_action(fire-and-forget)Related
Summary by Sourcery
Ensure fire-and-forget system-initiated agent events run to completion without changing disconnect handling for interactive conversations.
Bug Fixes:
Enhancements: