Repository navigation
fix: speed up offline notification on disconnect (PMS #261905) - #783
Johnson-zs wants to merge 1 commit into
Conversation
1. Root cause: Offline notification had ~6s delay (3s ping threshold + 3s cache timer). preprocessOfflineStatus called connect() on every invocation without disconnect, leaking Qt connections. Also sendToRemote blocked on sync RPC during known offline state. 2. Fix: Move cacheTimer connect to initConnect for one-time setup, reduce cache interval to 500ms, lower ping threshold from 3 to 2, propagate setTimeout to m_trans_timeout, add early offline check in sendToRemote 3. Impact: Offline notification reaches frontend in ~2.5s instead of ~6s. No connection accumulation. Transfer exits immediately when offline known. zrpc timeout now configurable via setTimeout API Log: Speed up offline notification when network disconnects Influence: 1. Test file transfer with network disconnect (silent, no FIN/RST) 2. Test offline notification timing and repeated offline/online cycles 3. Verify normal file transfer unaffected fix: 加快断网时离线通知速度 1. 根因:离线通知全链路延迟约 6 秒(Ping 检测 3 秒 + 缓存定时器 3 秒)。 preprocessOfflineStatus 每次调用执行 connect 但从不 disconnect,连接 数线性累积。sendToRemote 在已知离线时仍阻塞在同步 RPC。 2. 方案:将 cacheTimer 连接移至 initConnect 一次性建立,缓存间隔降至 500ms,Ping 失败阈值从 3 降为 2,setTimeout 同时设置 m_trans_timeout, sendToRemote 入口增加离线早退检查。 3. 影响:离线通知约 2.5 秒到达前端(原约 6 秒)。无连接累积。传输路径 在已知离线时立即退出。zrpc 超时可通过现有 setTimeout API 配置。 Log: 加快网络断开时离线通知速度 Influence: 1. 测试断网场景下文件投送(静默断开,无 FIN/RST) 2. 测试离线通知时延及反复断网/恢复循环 3. 验证正常文件投送流程不受影响 PMS: BUG-261905
There was a problem hiding this comment.
Sorry @Johnson-zs, you have reached your weekly rate limit of 500000 diff characters.
Please try again later or upgrade to continue using Sourcery
|
[APPROVALNOTIFIER] This PR is NOT APPROVED This pull-request has been approved by: Johnson-zs The full list of commands accepted by this bot can be found here. DetailsNeeds approval from an approver in each of these files:Approvers can indicate their approval by writing |
Reviewer's GuideSpeeds up and de-duplicates offline notifications and avoids unnecessary blocking RPC calls by fixing the offline status cache timer wiring, tuning ping and timer thresholds, and propagating RPC timeouts from the client to the underlying TCP connection. Sequence diagram for faster offline notificationsequenceDiagram
participant Ping as SendRpcWork
participant IPC as SendIpcService
participant Timer as cacheTimer
participant Frontend as Frontend
Ping->>Ping: handlePing
alt 2 consecutive ping failures
Ping->>IPC: preprocessOfflineStatus
IPC->>Timer: startOfflineTimer
Timer-->>IPC: timeout after 500ms
IPC->>Frontend: Frontend.notifySendStatus
end
Flow diagram for skipping offline transfer RPCflowchart TD
Start[sendToRemote] --> Device{_device_not_enough?}
Device -->|yes| Reject[return false]
Device -->|no| Offline{_offlined?}
Offline -->|yes| Reject
Offline -->|no| RPC[Synchronous remote transfer RPC]
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
Root Cause Analysis
Offline notification on network disconnect had ~6s total delay: Ping detection required 3 consecutive failures (~3s at 1s interval) before declaring offline (
sendrpcservice.cpp:128), plus a fixed 3s cache timer inpreprocessOfflineStatus(sendipcservice.cpp:314). The core defect was thatpreprocessOfflineStatuscalledconnect(&_cacheTimer, ...)on every invocation without ever disconnecting (sendipcservice.cpp:315), causing Qt signal-slot connections to accumulate linearly — each timer timeout triggered multiple duplicate lambda executions. Additionally,sendToRemote(transferjob.cpp:748) held_send_mutexwhile blocking on a synchronous RPC whose recv timeout was hardcoded at 3000ms (tcpconnection.h:107), andsetTimeoutinrpcchannel.cpp:89only set the connect timeout (m_max_timeout) without propagating to the actual recv/send timeout (m_trans_timeout).Fix Approach
Five targeted changes: (1) moved the
connect(&_cacheTimer, ...)frompreprocessOfflineStatustoinitConnect()for one-time setup, eliminating connection accumulation — the lambda capture changed from[this, appName]to[this]since the body iterates all keys; (2) reduced cache timer interval from 3000ms to 500ms; (3) lowered ping failure threshold from> 2(3 failures) to> 1(2 failures); (4) addedsetTransTimeout()toTcpConnectionand modifiedTcpClient::setTimeoutto propagate the timeout tom_trans_timeoutvia the connection; (5) added an earlyif (_offlined) return false;check at the start ofsendToRemoteto skip the blocking RPC when offline is already known.Change Safety Assessment
Code Safety
preprocessOfflineStatusconnect accumulation was introduced in commitfcf512fa("Fix remote offline unstable") as an implementation oversight, not intentional design — moving toinitConnect()does not revert the original fix intent.0c85b579("no prompt when networks disconnected", PMS 228107) modifiedcurstatusat line 309 — our fix does not touch this line, no regression risk.tcpconnection.h,tcpclient.h) are additive: a new public methodsetTransTimeout()and a null-guarded propagation in the existingsetTimeoutsetter — no signature changes, no existing caller affected.Business Impact Scope
sendToRemoteexits immediately instead of blocking up to 3s on a sync RPC, preventing UI freezes.setTimeoutAPI, benefiting all RPC callers that set a controller timeout.Verification Suggestion
根因分析
断网时离线通知全链路延迟约 6 秒:Ping 检测需连续失败 3 次(1 秒间隔,约 3 秒)才判离线(
sendrpcservice.cpp:128),加上preprocessOfflineStatus固定 3 秒缓存定时器(sendipcservice.cpp:314)。核心缺陷在于preprocessOfflineStatus每次调用都执行connect(&_cacheTimer, ...)但从不 disconnect(sendipcservice.cpp:315),导致 Qt 信号槽连接数随调用线性累积——每次定时器超时触发多个重复 lambda 执行。此外,sendToRemote(transferjob.cpp:748)持_send_mutex阻塞在同步 RPC,recv 超时硬编码 3000ms(tcpconnection.h:107),而rpcchannel.cpp:89的setTimeout只设连接超时m_max_timeout,未联动 recv/send 实际使用的m_trans_timeout。修复方案
五处针对性改动:(1) 将
connect(&_cacheTimer, ...)从preprocessOfflineStatus移至initConnect()一次性建立,消除连接累积——lambda 捕获从[this, appName]改为[this](函数体遍历所有 key,不需要 appName);(2) 缓存定时器间隔从 3000ms 降至 500ms;(3) Ping 失败阈值从> 2(3 次)降至> 1(2 次);(4) 为TcpConnection新增setTransTimeout()方法,修改TcpClient::setTimeout将超时传播至m_trans_timeout;(5) 在sendToRemote入口增加if (_offlined) return false;早退检查,已知离线时跳过阻塞式 RPC。改动安全评估
代码安全评估
preprocessOfflineStatus的 connect 累积问题由 commitfcf512fa("Fix remote offline unstable")引入,属实现疏漏而非设计意图——移至initConnect()不改变原修复目标。0c85b579("no prompt when networks disconnected",PMS 228107)修改了第 309 行curstatus——本修复不涉及该行,无回归风险。tcpconnection.h、tcpclient.h)均为新增:一个公有方法setTransTimeout()和setTimeout中带空指针保护的传播——无签名变更,不影响现有调用者。业务影响范围
sendToRemote立即退出,不再阻塞最多 3 秒在同步 RPC 上,避免 UI 卡死。setTimeoutAPI 配置,惠及所有设置 controller 超时的 RPC 调用方。验证建议
PMS: BUG-261905
Summary by Sourcery
Improve disconnect handling so remote offline notifications arrive faster and known-offline transfers return promptly.
Bug Fixes:
Enhancements: