Skip to content

fix(update): keep Windows daemons serving across self-update - #346

Open
iuyo5678 wants to merge 10 commits into
mainfrom
fix/336-windows-in-place-update
Open

iuyo5678 wants to merge 10 commits into
mainfrom
fix/336-windows-in-place-update

Conversation

@iuyo5678

@iuyo5678 iuyo5678 commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Fixes #336. A failed auto-update no longer leaves the browser disconnected.

  • Handover is confirmed, not assumed. The outgoing daemon spawns a daemon from the new executable, releases its lock, IPC endpoint and port, but stays alive. It exits only once a daemon of the release's version serves the original port; a daemon answering with another version or port does not count and is named in the error. If the new daemon exits or is not ready within 20 seconds, it is stopped, and the outgoing daemon puts the previous executable back, takes the lock again and serves on the same port. The record says a daemon is serving only after the resumed one has published daemon.json; if it cannot serve again, that error is appended. Browsers reconnect across the brief gap either way.
  • bsk update owns the daemon it restarts. It waits for that process to serve the new version on the original port, and stops and reaps it if it is not ready in time, before restoring and restarting the previous version. (The shared bsk daemon start path deliberately leaves a slow child running, since another client may reuse it; the update path does not use it any more.)
  • The daemon keeps the executable path it started from, so it can update again after a rollback on Linux, where current_exe then names <path> (deleted).
  • The update lock guards the installed executable (.bsk.update.lock next to it), so updates from different bsk homes are serialized through installation, confirmation and rollback.
  • The previous executable is kept until the update is confirmed, on every platform: renamed aside on Windows (a running image can be renamed but not deleted), hard-linked (or copied) on Unix. The new executable must print exactly the release's version for --version before any daemon is asked to use it.
  • Every attempt is recorded in update-state.json in the bsk home: source, versions, stage (download / install / handover / restart), result, full error, recovery (unchanged, restored with whether a daemon is serving, or restore_failed with the manual step), and when the daemon retries. A replacement that fails to start leaves its own error for the outgoing daemon to record. bsk doctor shows this record, daemon/installed version skew, and files kept from earlier updates.
  • Windows no longer uses the detached cmd helper. The running bsk.exe is renamed aside and the new one takes its place, retrying briefly while a scanner holds the file.
  • bsk update installs and self-checks before stopping the daemon, and restores the previous executable and restarts the daemon from it if the restarted daemon does not become ready.
  • Only a daemon bsk started in the background installs updates. A --foreground daemon logs a new version once (WARN) and leaves the upgrade to its owner. It records host_managed in daemon.json, and bsk update then installs the release but leaves that daemon running ("daemon": "left_to_host"), instead of stopping it and starting a background daemon in its place, which would take it from its supervisor or, inside a host that forbids Job breakaway, leave no daemon. The CLI hint and bsk doctor say: run bsk update, then restart the daemon in its terminal or supervisor. A background daemon is restarted on the port it served. If bsk cannot write next to its executable, it reports the release and the CLI hint points to the installer. Update attempts are serialized by an update lock; a failed version is retried after 6 hours, or immediately with bsk update.
  • Documents BSK_AUTO_UPDATE=off (credit: fix(update): fall back to foreground daemon restart #345).

Scope

These changes apply on all platforms, not only Windows: the confirmed handover and rollback, the update record and doctor check, the foreground and unwritable-directory policies, and the bsk update ordering. Windows additionally replaces the helper script with the in-place rename.

Windows support follows the Rust x86_64-pc-windows-msvc target: Windows 10 / Windows Server 2016 or later. The rename of a running executable is verified on NTFS (CI). On a file system that refuses it, the first rename fails and the installation is left unchanged; that failure is recorded like any other.

Updates that start from 0.3.1 or earlier still run the updater of that version; this behavior applies from the release after this one.

Test plan

  • cargo test --workspace --locked (macOS): all suites pass
  • cargo clippy --workspace --all-targets --locked -- -D warnings, cargo fmt --all -- --check
  • New tests/auto_update_handover.rs, cross-platform, with a real detached daemon and a local release server:
    • successful handover: the old daemon exits only after the new one serves on the same port; the record says succeeded with the new pid; kept files are removed
    • a release whose daemon fails to start (compiled stand-in that passes --version): the previous executable is restored, the same process serves again on the same port, and the record has the error and restored
    • a release that cannot run, or reports another version than its manifest: rejected at install, daemon never stopped
    • bsk update whose new daemon takes the lock and hangs: it is stopped, the previous version serves the original port again, record restored / serving
    • a failed handover whose port is taken meanwhile: the resumed daemon cannot bind, record daemon_serving: false with both errors
    • fixtures are the bsk under test with its version string replaced by a newer one of the same length (re-signed ad hoc on macOS)
    • bsk update with a --foreground daemon: release installed, the owner's daemon keeps running with the same pid
    • bsk update with a background daemon: restarted from the new release on the port it served
  • Handover unit tests: a replacement that exits early is reported with its own error; one that never becomes ready is stopped
  • Windows unit tests: failed swap puts the original back; failed swap and restore reports the manual step; restore keeps an executable at the path; restore of a running image; concurrent replacements leave one intact executable
  • Update lock admits one attempt at a time and is shared by bsk homes that update the same executable; backoff, record, and hint unit tests; doctor check for each outcome
  • A daemon answering with another version or port is not accepted as the replacement (unit tests)
  • Linux: a process whose executable was replaced and rolled back installs again from the path it captured at startup
  • Linux and Windows CI
  • Windows CI, including the handover suite in the independent (non-Job) host

Follow-ups (not in this PR)

  • Optional per-user scheduled task for a persistent daemon (bsk daemon install).
  • An exit-code contract so supervisors can restart a --foreground daemon into a new version.
  • Authenticode signing of release binaries.

TencentXiaowei and others added 10 commits September 27, 2026 00:03
Rename the running executable aside instead of handing replacement to a
detached script, and exit only after the new daemon has started. A
foreground daemon reports the release and leaves the restart to its owner.

Document BSK_AUTO_UPDATE=off.

Co-authored-by: Sthreal <2066885216@qq.com>
The outgoing daemon releases its lock, IPC endpoint and port but stays
alive until a daemon started from the new executable answers. If that
daemon exits or is not ready within 20 seconds, it puts the previous
executable back and serves again on the same port.

- Keep the previous executable on every platform until the update is
  confirmed, and run the new one once before relying on it.
- Record each attempt's stage, result, error and recovery in
  update-state.json, show it in `bsk doctor`, and retry a failed
  version after 6 hours.
- Serialize update attempts with an update lock.
- Report new releases instead of installing them when bsk cannot write
  next to its executable.
- `bsk update` installs before stopping the daemon and rolls back if the
  restarted daemon does not become ready.
It now exits only after it has seen the replacement serve, so it can
still be running when the replacement publishes daemon.json.
`bsk update` stopped any running daemon and started a background one in
its place. For a `--foreground` daemon that took it away from its
terminal or supervisor, and inside a host that forbids Job breakaway it
left no daemon at all.

The daemon now records in daemon.json whether its owner manages it.
`bsk update` installs the release but leaves such a daemon running and
says to restart it there; the CLI hint and `bsk doctor` give the same
steps. A background daemon is restarted on the port it served.
- `bsk update` restarts the daemon as a process it owns. One that is not
  ready in time is stopped and reaped before the previous version is
  restarted, so the rollback no longer meets it holding the daemon lock.
- A handover or restart succeeds only when a daemon of the release's
  version serves the original port. `--version` must print that version.
- A failed handover records a serving daemon only once the resumed one has
  published daemon.json, and appends the error if it cannot serve again.
- A daemon keeps the executable path it started from, so it can update
  again after a rollback on Linux, where `current_exe` then names a
  deleted file.
- The update lock sits next to the installed executable, so updates from
  different bsk homes are serialized through installation and rollback.
After a failed handover the previous daemon served again with its
original configuration, so a daemon started with `--port 0` was given a
new port and browsers could no longer reconnect. It now resumes on the
port it served before the handover.
…ot restart

- The daemon owns its update in one slot that its blocking steps fill
  themselves. Shutdown aborts the async task but waits for those steps,
  then hands the update over or undoes it, so a daemon that goes idle
  while the release downloads or is checked no longer leaves the new
  executable installed without a way back. A download that finishes after
  shutdown began installs nothing.
- `bsk update` checks that it can start an independent daemon before
  stopping a background one. Inside a Windows Job that forbids breakaway
  it leaves that daemon running (`left_running`) instead of stopping a
  service it could not start again.
- The installation takes the update lock and keeps it until it is
  confirmed or rolled back; the flaky lock-free concurrency test is
  replaced by one that takes turns through the lock.
- The Windows lifecycle host gives each suite 480 seconds, keeps its output
  after a timeout, and stops daemons the suite started from test copies.
The lifecycle host stopped every process started from the system TEMP
after a suite timed out, which could hit unrelated programs. Each suite
now gets its own TEMP, and only processes running from it are stopped.
Copies of bsk are removed from it afterwards so the artifact keeps logs.

Compile the `locked` test helper only where its tests run, which removes
a Windows dead-code warning.
@iuyo5678

Copy link
Copy Markdown
Collaborator Author

Some of the documentation in this PR comes from PR#344, authored by @Sthreal, but the previous Windows upgrade solution design had issues that were difficult to fully resolve, so it was completely refactored. Thanks for the previous PR.

@Sthreal Could you review the approach in this latest PR and see if there are any issues?

@iuyo5678

Copy link
Copy Markdown
Collaborator Author

补充这次 PR 的完整 Windows 本地验证结果,覆盖首次测试及两次更新后的复测。最后验证的提交是 d691cc172642cc5ce6c24882a8cf5c33b2079eee,测试日期为 2026-09-27。

结论:最新提交的相关回归为 413 passed / 0 failed;Windows Clippy 相对 base 没有新增告警。 不过,之前较广范围的 Windows 运行时测试仍有失败,不能据此宣称整个 workspace 在 Windows 上全绿;具体范围和限制列在下面。

测试环境与版本

  • Windows 11 Pro x64,build 26200;C:、E: 均为 NTFS。
  • Rust 1.97.1 / Clippy 0.1.97,原生 x86_64-pc-windows-msvc 工具链,debug 构建。
  • Windows PowerShell 5.1.26100.9444、PowerShell 7.6.5、cmd.exe。
  • 三轮测试提交分别为 c7c3693、8c79b19f、d691cc1。前两轮包含较广范围测试及额外压力/边界测试;最后一轮针对实际新增改动,重新检查完整 Clippy targets,并通过新的 Windows Worker 脚本运行相关真实回归。
  • 测试使用私有 BSK_HOME、本地发布服务和独立端口;需要独立进程的测试从经检查不在 Job 内的宿主运行,受限及嵌套 Job 场景则显式创建真实 Win32 Job 对象。

测试期间发现的问题及最终状态

  1. daemon 在下载或安装期间 idle 退出,更新结果失去事务归属:已修复。

    在 c7c3693 上复现两次:归档下载延迟 8 秒,daemon 使用短 idle timeout;新版本能返回正确的 --version,但启动 daemon 会退出 3。daemon idle 退出后,阻塞任务仍安装新版本,异步任务已经取消,安装结果未被交接状态持有;旧文件没有恢复,记录停在 install/in_progress,下一次启动执行坏版本。

    在 8c79b19f 上分别针对“下载尚未完成”和“已经安装、进入慢自检”各复测两次:前者不安装,记录 failed/unchanged;后者恢复旧文件,记录 failed/restored、daemon_serving: false。两种场景之后均可正常启动旧版本。最新 d691cc1 的交接套件也再次通过这两个回归用例。

  2. 受限 Job 内手动更新可能停止旧 daemon,却无法重新启动:已修复。

    早期版本复现两次:旧后台 daemon 在 Job 外正常服务,更新命令位于禁止 breakaway 的 Job 中;安装和自检成功后停止旧 daemon,新版本及回滚后的重启均报 OS error 5,最终没有服务中的 daemon。

    修复后,重启前会检查独立启动能力。真实 Job 测试覆盖禁止 breakaway、允许 breakaway、silent breakaway、外层禁止/内层允许、两层均允许,共五种组合。禁止脱离时返回 left_running,旧 PID、端口和 IPC 服务保持正常;允许脱离时完成同端口重启。最新提交对应的受限 Job 集成回归也通过。

  3. 原始并发文件替换单元测试在 Windows 上不稳定:已修复。

    原测试直接并发调用不带生产更新锁的底层替换函数,在 20 次独立运行中失败 5 次;整个 Windows 更新单元测试组也在 5 次重复中失败 2 次。实际两个 CLI 并发更新的控制测试一直能正确做到一个成功、另一个报锁占用,因此当时没有把单元测试失败解释为生产 CLI 会损坏安装。

    改为通过生产更新锁执行后,新测试 concurrent_updates_of_one_executable_take_turns_and_leave_it_intact 独立运行 50/50 通过;真实并发 CLI 更新也通过,最终文件完整。Windows 文件替换单元测试另重复五轮,每轮 9 passed / 1 ignored,全部通过。上述压力测试是在 8c79b19f 上完成;最新提交未修改该实现,并再次通过相关 library 测试。

  4. 测试脚本超时清理可能误杀 TEMP 中的无关进程:已修复。

    在 8c79b19f 上,使用私有 TEMP 和无关哨兵进程,确认全局 TEMP 前缀加创建时间的筛选会误杀无关进程。

    在 d691cc1 上,直接运行未修改的 Worker 脚本,分别用 PowerShell 5.1 和 7 验证:本套件专用 TEMP 中的脱离子进程在超时后被回收;共享 TEMP、相似前缀目录、其他套件目录中的无关哨兵全部保留。每个套件的 TEMP/TMP 正确隔离;正常退出、非零退出、超时后继续执行后续套件、stdout/stderr 和诊断文件保留、测试 exe 副本清理都符合预期。测试路径包含中文、空格、& 和方括号。

  5. locked() 缺少 Windows 条件编译产生 dead-code 告警:已修复。

    最新提交加上对应 cfg 后,该告警不再出现。下面是完整的原生 Windows Clippy 对照结果。

针对作者提出的 Windows Clippy 验证

在最新 PR 提交和实际 base 8ca7911778c30a73ff0149d53d267023c31961bc 上分别运行:

cargo clippy --workspace --all-targets --locked -- -D warnings

两边均退出 101,严格检查阶段报出相同的原有告警。为了避免早期失败遮住后续 integration-test targets,又在两边运行:

cargo clippy --workspace --all-targets --locked --message-format=json

没有添加 -A 或关闭任何 lint;两边都完整执行完成并退出 0。按文件、位置、lint 和消息去重比较,得到完全相同的 5 条告警,PR 新增为 0:

  • crates/bsk-cli/src/daemon/file_transfer.rs:417:unused variable path。
  • crates/bsk-cli/src/daemon/file_transfer.rs:426:unused variable file。
  • crates/bsk-cli/src/skill_install/harness.rs:339:needless-return。
  • crates/bsk-cli/src/daemon/ipc.rs:793:handle_cancel_with_registry_only dead-code。
  • crates/bsk-cli/tests/tools_m9_ipc.rs:40:unused import support::wait_for_abort_registered。

因此可以确认 locked() 的修复有效,PR 新增及改动的代码没有产生额外 Windows Clippy 告警;严格 Clippy 仍受上述仓库基线告警影响。本轮使用原生 Windows MSVC 工具链并复用已有依赖缓存,没有遇到 aws-lc-sys 编译失败;未验证 macOS 上的交叉编译链,也未强制冷重建该依赖。

cargo fmt --all --check 通过。

更新行为、路径、权限及进程边界覆盖

以下附加验证主要完成于 8c79b19f,也是这次汇总的一部分;最新提交只继续修改了 locked() 的 cfg 和测试清理脚本,没有改动对应更新主流程。

  • 手动更新、自动更新、失败交接后的回滚:本地 WebSocket 协议客户端均能在原端口重新握手。成功更新切换至新版本;失败回滚恢复旧版本及原 daemon 的服务。
  • 前台 host-managed daemon:更新安装成功,但保留原 PID 和服务,返回 left_to_host;受限 Job 内同样通过。
  • --no-restart-daemon:安装更新后保留原 PID、状态仍可用。
  • PowerShell 5.1、PowerShell 7、cmd:中文、空格、%PATH%、! & ( )、单引号等路径场景均通过。
  • 长路径:首次 371 字符、后续 390 字符的扩展长度可执行文件路径更新均通过。调用进程时显式指定可执行文件,避免将测试驱动自身的路径搜索限制误判为更新器问题。
  • C: 与 E: NTFS 安装目录均通过;这不代表测试了跨卷原子 rename,更新暂存文件与目标仍在同一目录。
  • SHA-256 不匹配、错误版本、无效 PE:更新失败并保留或恢复原可执行文件。
  • 真实不允许 delete sharing 的文件句柄:持续占用时安全失败,原文件保留;临时占用测试等待 staged 文件完整写入后继续持有 1 秒,再释放句柄,重试成功。
  • 只读文件属性:更新成功;目录 ACL 拒绝创建文件:下载前即报告 not_writable,下载次数为 0,原文件保持不变。
  • 核心交接、Windows 更新、启动和更新策略套件在 8c79b19f 上连续三轮通过,其中两轮使用含中文、空格和特殊符号的 TEMP/TMP。交接每轮 10/10,Windows 更新每轮 4/4,启动每轮 9/9(同一默认端口用例过滤),策略每轮 1/1。

早期特殊 TEMP 重复运行曾出现启动等待预算超时;另一次成功交接套件耗时约 314.7 秒,超过旧脚本 180 秒上限。后续三轮相关回归未再出现该失败,脚本套件预算已调整为 480 秒;这证明本机这些运行通过,不代表所有机器负载下都不会达到超时上限。

最新提交 d691cc1 的真实回归结果

通过本提交未修改的 Worker 脚本,在独立宿主和每套件专用 TEMP/TMP 下运行:

  • CLI library:389 passed / 2 ignored。
  • Windows daemon startup:9 passed / 1 ignored / 1 filtered。
  • Windows update:4 passed / 1 ignored。
  • Auto-update policy:1 passed。
  • Auto-update handover:10 passed。

合计 413 passed / 0 failed / 4 ignored / 1 filtered,五个套件的退出码全部为 0。交接套件耗时约 186 秒,未触发 480 秒上限。执行后没有测试 exe 或测试 daemon 残留,原有用户 daemon 保持运行。

忽略项是测试子进程入口;主动过滤的是 automatic_start_returns_eof_and_reuses_the_same_daemon,因为本机已有 daemon 占用默认端口 52800,验证期间没有停止它。

较广范围的 Windows 测试结果与归因边界

三轮提交、测试范围不同,以下计数不叠加,也不把最新的定向回归当作之前全部失败用例已经修复:

  • c7c3693:运行已编译的 43 个测试二进制,656 passed / 89 failed / 3 ignored / 1 filtered。
  • 8c79b19f:再次运行 43 个测试二进制,668 passed / 88 failed / 4 ignored / 1 filtered。
  • d691cc1:本轮按两处实际改动进行 Clippy 全 targets 检查及上述五个相关运行时套件,结果为 413 passed / 0 failed / 4 ignored / 1 filtered,没有重新运行全部 43 个测试二进制。

前两轮大部分失败属于 .sock 路径被用于 Windows named pipe 的 IPC 测试;第二轮有 87 个此类失败,另 1 个是 authorization_file_contention_does_not_block_socket_messages 返回 401、预期 503。首次运行另见一个远程服务文件锁/访问失败。

这些失败类别与前后两轮一致,已核对相关报错测试文件和 IPC 实现文件在 PR 相对对应 base 的 diff 中未改。但是没有在 base 上逐一 A/B 重跑全部运行时失败,所以不将“文件未改”当成它们全部属于基线问题的证明。Clippy 则已在最新提交和实际 base 上完整做了对照,两者需要区分。

部分 cfg(unix) 测试二进制在 Windows 下只有 0 个测试,未将它们计作已验证的 Unix 行为;未运行 doctests。

实现评价与适用范围

更新方案整体上比早期版本更清楚:原生 rename/restore 去掉了脱离运行的 cmd 更新助手;Installed 持有更新锁直到确认或回滚;阻塞步骤直接保存事务状态,daemon 收尾时能够处理取消后的安装结果;只有确认新版本在预期端口服务才完成交接;手动更新会先验证 Job breakaway 能力再决定是否停止旧 daemon。最后的测试脚本改动也把清理范围明确限制到了每个套件自己的目录。

本地测试期间发现的上述问题均已有对应修复和复测证据,从这次 Windows 验证范围看,没有发现需要继续阻止合并的新增问题。仓库原有 Windows lint 和较广范围运行时测试失败可继续独立跟进。

覆盖范围仍是同一台 Windows 11 x64 NTFS 机器上的多种 shell、Job、路径和权限条件,未实测 Windows 10、Windows Server、ARM64、非 NTFS/SMB、真实不同标准用户登录令牌或企业 EDR 策略。更新测试使用本地构建并等长替换版本字符串的可执行文件,以及受控的失败/慢自检测试程序;未验证正式签名 release。WebSocket 覆盖的是协议握手与重连,没有声称完成真实 Chrome/Edge 扩展 UI 端到端测试。

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] Windows auto-update replaces the bsk binary but never restarts the daemon — and leaves no failure log

2 participants