Conversation
Rename the running executable aside instead of handing replacement to a detached script, and exit only after the new daemon has started. A foreground daemon reports the release and leaves the restart to its owner. Document BSK_AUTO_UPDATE=off. Co-authored-by: Sthreal <2066885216@qq.com>
The outgoing daemon releases its lock, IPC endpoint and port but stays alive until a daemon started from the new executable answers. If that daemon exits or is not ready within 20 seconds, it puts the previous executable back and serves again on the same port. - Keep the previous executable on every platform until the update is confirmed, and run the new one once before relying on it. - Record each attempt's stage, result, error and recovery in update-state.json, show it in `bsk doctor`, and retry a failed version after 6 hours. - Serialize update attempts with an update lock. - Report new releases instead of installing them when bsk cannot write next to its executable. - `bsk update` installs before stopping the daemon and rolls back if the restarted daemon does not become ready.
It now exits only after it has seen the replacement serve, so it can still be running when the replacement publishes daemon.json.
`bsk update` stopped any running daemon and started a background one in its place. For a `--foreground` daemon that took it away from its terminal or supervisor, and inside a host that forbids Job breakaway it left no daemon at all. The daemon now records in daemon.json whether its owner manages it. `bsk update` installs the release but leaves such a daemon running and says to restart it there; the CLI hint and `bsk doctor` give the same steps. A background daemon is restarted on the port it served.
- `bsk update` restarts the daemon as a process it owns. One that is not ready in time is stopped and reaped before the previous version is restarted, so the rollback no longer meets it holding the daemon lock. - A handover or restart succeeds only when a daemon of the release's version serves the original port. `--version` must print that version. - A failed handover records a serving daemon only once the resumed one has published daemon.json, and appends the error if it cannot serve again. - A daemon keeps the executable path it started from, so it can update again after a rollback on Linux, where `current_exe` then names a deleted file. - The update lock sits next to the installed executable, so updates from different bsk homes are serialized through installation and rollback.
After a failed handover the previous daemon served again with its original configuration, so a daemon started with `--port 0` was given a new port and browsers could no longer reconnect. It now resumes on the port it served before the handover.
…ot restart - The daemon owns its update in one slot that its blocking steps fill themselves. Shutdown aborts the async task but waits for those steps, then hands the update over or undoes it, so a daemon that goes idle while the release downloads or is checked no longer leaves the new executable installed without a way back. A download that finishes after shutdown began installs nothing. - `bsk update` checks that it can start an independent daemon before stopping a background one. Inside a Windows Job that forbids breakaway it leaves that daemon running (`left_running`) instead of stopping a service it could not start again. - The installation takes the update lock and keeps it until it is confirmed or rolled back; the flaky lock-free concurrency test is replaced by one that takes turns through the lock. - The Windows lifecycle host gives each suite 480 seconds, keeps its output after a timeout, and stops daemons the suite started from test copies.
The lifecycle host stopped every process started from the system TEMP after a suite timed out, which could hit unrelated programs. Each suite now gets its own TEMP, and only processes running from it are stopped. Copies of bsk are removed from it afterwards so the artifact keeps logs. Compile the `locked` test helper only where its tests run, which removes a Windows dead-code warning.
|
Some of the documentation in this PR comes from PR#344, authored by @Sthreal, but the previous Windows upgrade solution design had issues that were difficult to fully resolve, so it was completely refactored. Thanks for the previous PR. @Sthreal Could you review the approach in this latest PR and see if there are any issues? |
|
补充这次 PR 的完整 Windows 本地验证结果,覆盖首次测试及两次更新后的复测。最后验证的提交是 结论:最新提交的相关回归为 413 passed / 0 failed;Windows Clippy 相对 base 没有新增告警。 不过,之前较广范围的 Windows 运行时测试仍有失败,不能据此宣称整个 workspace 在 Windows 上全绿;具体范围和限制列在下面。 测试环境与版本
测试期间发现的问题及最终状态
针对作者提出的 Windows Clippy 验证 在最新 PR 提交和实际 base 两边均退出 101,严格检查阶段报出相同的原有告警。为了避免早期失败遮住后续 integration-test targets,又在两边运行: 没有添加
因此可以确认
更新行为、路径、权限及进程边界覆盖 以下附加验证主要完成于
早期特殊 TEMP 重复运行曾出现启动等待预算超时;另一次成功交接套件耗时约 314.7 秒,超过旧脚本 180 秒上限。后续三轮相关回归未再出现该失败,脚本套件预算已调整为 480 秒;这证明本机这些运行通过,不代表所有机器负载下都不会达到超时上限。 最新提交 通过本提交未修改的 Worker 脚本,在独立宿主和每套件专用 TEMP/TMP 下运行:
合计 413 passed / 0 failed / 4 ignored / 1 filtered,五个套件的退出码全部为 0。交接套件耗时约 186 秒,未触发 480 秒上限。执行后没有测试 exe 或测试 daemon 残留,原有用户 daemon 保持运行。 忽略项是测试子进程入口;主动过滤的是 较广范围的 Windows 测试结果与归因边界 三轮提交、测试范围不同,以下计数不叠加,也不把最新的定向回归当作之前全部失败用例已经修复:
前两轮大部分失败属于 这些失败类别与前后两轮一致,已核对相关报错测试文件和 IPC 实现文件在 PR 相对对应 base 的 diff 中未改。但是没有在 base 上逐一 A/B 重跑全部运行时失败,所以不将“文件未改”当成它们全部属于基线问题的证明。Clippy 则已在最新提交和实际 base 上完整做了对照,两者需要区分。 部分 实现评价与适用范围 更新方案整体上比早期版本更清楚:原生 rename/restore 去掉了脱离运行的 cmd 更新助手; 本地测试期间发现的上述问题均已有对应修复和复测证据,从这次 Windows 验证范围看,没有发现需要继续阻止合并的新增问题。仓库原有 Windows lint 和较广范围运行时测试失败可继续独立跟进。 覆盖范围仍是同一台 Windows 11 x64 NTFS 机器上的多种 shell、Job、路径和权限条件,未实测 Windows 10、Windows Server、ARM64、非 NTFS/SMB、真实不同标准用户登录令牌或企业 EDR 策略。更新测试使用本地构建并等长替换版本字符串的可执行文件,以及受控的失败/慢自检测试程序;未验证正式签名 release。WebSocket 覆盖的是协议握手与重连,没有声称完成真实 Chrome/Edge 扩展 UI 端到端测试。 |
Summary
Fixes #336. A failed auto-update no longer leaves the browser disconnected.
daemon.json; if it cannot serve again, that error is appended. Browsers reconnect across the brief gap either way.bsk updateowns the daemon it restarts. It waits for that process to serve the new version on the original port, and stops and reaps it if it is not ready in time, before restoring and restarting the previous version. (The sharedbsk daemon startpath deliberately leaves a slow child running, since another client may reuse it; the update path does not use it any more.)current_exethen names<path> (deleted)..bsk.update.locknext to it), so updates from different bsk homes are serialized through installation, confirmation and rollback.--versionbefore any daemon is asked to use it.update-state.jsonin the bsk home: source, versions, stage (download / install / handover / restart), result, full error, recovery (unchanged,restoredwith whether a daemon is serving, orrestore_failedwith the manual step), and when the daemon retries. A replacement that fails to start leaves its own error for the outgoing daemon to record.bsk doctorshows this record, daemon/installed version skew, and files kept from earlier updates.cmdhelper. The runningbsk.exeis renamed aside and the new one takes its place, retrying briefly while a scanner holds the file.bsk updateinstalls and self-checks before stopping the daemon, and restores the previous executable and restarts the daemon from it if the restarted daemon does not become ready.bskstarted in the background installs updates. A--foregrounddaemon logs a new version once (WARN) and leaves the upgrade to its owner. It recordshost_managedindaemon.json, andbsk updatethen installs the release but leaves that daemon running ("daemon": "left_to_host"), instead of stopping it and starting a background daemon in its place, which would take it from its supervisor or, inside a host that forbids Job breakaway, leave no daemon. The CLI hint andbsk doctorsay: runbsk update, then restart the daemon in its terminal or supervisor. A background daemon is restarted on the port it served. If bsk cannot write next to its executable, it reports the release and the CLI hint points to the installer. Update attempts are serialized by an update lock; a failed version is retried after 6 hours, or immediately withbsk update.BSK_AUTO_UPDATE=off(credit: fix(update): fall back to foreground daemon restart #345).Scope
These changes apply on all platforms, not only Windows: the confirmed handover and rollback, the update record and doctor check, the foreground and unwritable-directory policies, and the
bsk updateordering. Windows additionally replaces the helper script with the in-place rename.Windows support follows the Rust
x86_64-pc-windows-msvctarget: Windows 10 / Windows Server 2016 or later. The rename of a running executable is verified on NTFS (CI). On a file system that refuses it, the first rename fails and the installation is left unchanged; that failure is recorded like any other.Updates that start from 0.3.1 or earlier still run the updater of that version; this behavior applies from the release after this one.
Test plan
cargo test --workspace --locked(macOS): all suites passcargo clippy --workspace --all-targets --locked -- -D warnings,cargo fmt --all -- --checktests/auto_update_handover.rs, cross-platform, with a real detached daemon and a local release server:succeededwith the new pid; kept files are removed--version): the previous executable is restored, the same process serves again on the same port, and the record has the error andrestoredinstall, daemon never stoppedbsk updatewhose new daemon takes the lock and hangs: it is stopped, the previous version serves the original port again, recordrestored/ servingdaemon_serving: falsewith both errorsbsk updatewith a--foregrounddaemon: release installed, the owner's daemon keeps running with the same pidbsk updatewith a background daemon: restarted from the new release on the port it servedFollow-ups (not in this PR)
bsk daemon install).--foregrounddaemon into a new version.