Skip to content

composefs: Hold /boot open while a deployment is staged - #6

Closed
cgwalters-bot wants to merge 3 commits into
mainfrom
bot/finalize-staged-hold
Closed

cgwalters-bot wants to merge 3 commits into
mainfrom
bot/finalize-staged-hold

Conversation

@cgwalters-bot

@cgwalters-bot cgwalters-bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

On composefs installs /boot is usually the ESP, automounted by systemd-gpt-auto-generator with TimeoutIdleSec=2min. When that idle expire overlaps with a staged deployment, finalization breaks in one of two ways, and both end with the machine booting the old deployment:

  • If the expire starts after the shutdown transaction is queued but before bootc-finalize-staged.service's ExecStop looks at /boot (lsblk does statfs() on every mountpoint), they deadlock. The lookup blocks in autofs_expire_wait until systemd unmounts boot.mount, but boot.mount is ordered to stop after finalize (via local-fs.target). Finalize is killed after 5+5 minutes.
  • If the expire completes before shutdown, systemd won't trigger the automount again during shutdown, so on grub finalize fails right away with Failed to open /boot: Host is down (os error 112).

This is behind most of the plan-44-shadow-fixup failures in CI: test-44 stages, then waits about 120-130s before rebooting, which lands right on the expire. In the 50 most recent failed CI runs, plan-44 was the most common failing plan (24 legs). 13 of those were the "guest reboot timeout" (the deadlock, on systemd-boot, grub and grub-cc legs). The other 11 were expected exactly one testbootcgroup in /etc/group, got: [] on composefs grub legs, which is the EHOSTDOWN case: the old deployment booted, so the test found no group. Analyses with journals and CI links: deadlock, EHOSTDOWN.

ostree hit the same problem and fixed it in ostreedev/ostree@f3db79e7 ("finalize-staged: Ensure /boot automount doesn't expire", ostreedev/ostree#2543) with ostree-finalize-staged-hold.service. This ports that design. bootc-finalize-staged.service now Wants=/After= a new bootc-finalize-staged-hold.service, which runs bootc composefs-finalize-staged --hold in the root mount namespace (ExecStart=+). That keeps an fd for /boot open until it's stopped, which at shutdown happens after finalization. autofs never expires a busy mount, so nothing can race with shutdown. The hold only opens /boot and is dispatched before storage is loaded, so it doesn't depend on sysroot setup or end up in a private mount namespace, where autofs wouldn't see it.

Relation to bootc-dev#2488: that PR adds RequiresMountsFor=/boot to the finalize unit. Without a hold, the implied Requires=boot.mount makes the idle expire stop the finalize unit while the system is still running, so any reboot more than about 2 minutes after staging hangs. Its CI shows this: plan-44 hit a reboot timeout on 3 legs in the first attempt and 2 in the retry (centos-10 composefs systemd-boot, plus one fedora grub leg), each with finalize stopped 5-10s before the reboot (CI analysis). ostree has RequiresMountsFor=/boot only together with its hold unit, and that is what would make it safe here too. This PR deliberately keeps the finalize unit at RequiresMountsFor=/sysroot: the stop ordering against boot.mount already comes from local-fs.target, and without Requires=boot.mount a failed hold leaves us no worse off than today.

test-44 now also asserts that staging on composefs starts the hold unit.

Testing

All on a 64-core RHEL 10 runner with bcvk 0.19.0 and tmt 1.78.0, using CentOS Stream 10 composefs images built with just build (BOOTC_variant=composefs, systemd-boot, xfs, BLS, unsealed), where /boot is the ESP as a systemd-1 autofs with timeout=120, as in CI.

The key test reboots right inside the race: stage a trivial derived image the way test-44 does, then schedule systemctl reboot 0.25-1.0s before systemd's next /boot expire check (after /boot has been idle for more than 120s). Without this change that hung for about 15 minutes and booted the old deployment every time (4/4 at those offsets, 4/10 across a wider -1.0..0s sweep). With it, nothing hung: before the rebase 0/7 at those offsets and 0/13 across the sweep, and on the rebased commit 3/3 more at -0.25, -0.5 and -1.0s (reboot about 133s after staging). Every run came up in the new deployment, with bootc-finalize-staged-hold.service active and no boot.mount expire in the previous boot's journal. Letting a staged system sit idle for 240s left finalize active and /boot mounted. On a Fedora 43 composefs grub image, rebooting after /boot had expired booted the old deployment on main (4/4, the EHOSTDOWN case), and 3/3 runs with the hold that waited 300s before rebooting booted the new one.

plan-44-shadow-fixup via cargo xtask run-tmt with CI's flags passed 3/3 before the rebase and 2/2 after, and just validate and just unit-tests pass on the rebased commit. The tmt runs only show the plan and the new assertion work; the devspace is fast enough that switch→reboot stays under 120s there, so they don't measure the flake rate.

Caveats

  • The hold covers /boot. A system with the ESP automounted at /efi and /boot on the root filesystem isn't covered, and lsblk in finalize still statfs()es every mountpoint, so another expiring automount could in principle still hang it. Whether finalize needs lsblk at all is worth a separate look.
  • A deployment staged by an older bootc that doesn't start the hold unit is still exposed to the race until that reboot.
  • The commits have no Signed-off-by; add one before merging if you want it.

Generated-by: https://github.com/cgwalters/#llms


Review draft in cgwalters-forge, not upstream yet. This section is removed when the PR is opened upstream.

  • Upstream: bootc-dev/bootc, base main
  • Board item: PVTI_lAHOAQ_SPs4Bj2Gizg8P2EI

To review:

  • Approve to open it upstream, ready for review.
  • Comment /draft, then approve, to open it upstream as a draft (/ready undoes that).
  • Close to drop it.
  • Edit the title and description freely: they become the upstream PR's. Review comments are addressed with fixup commits and a reply here.

@cgwalters

Copy link
Copy Markdown

Looks fine go ahead and push a PR to proper upstream

@cgwalters-bot

Copy link
Copy Markdown
Collaborator Author

Thanks! Promotion only goes ahead on an approval or a comment line that is exactly /promote (tied to the current head), so it doesn't act on free-form text. Reply /promote here and the next poll opens the upstream PR.

Generated-by: https://github.com/cgwalters/#llms

@cgwalters

Copy link
Copy Markdown

Massively trim down the commit message it should basically just be like:

Copy over the logic from ostreedev/ostree@f3db79e7
("finalize-staged: Ensure /boot automount doesn't expire") to composefs as well. In order to properly handle shutdown we need ensure we have a file descriptor open.

Then link to the bug

}

/// The mount that must stay busy while a deployment is staged.
const HOLD_PATH: &str = "/boot";

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pretty sure we have another const for this somewhere

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Now using the existing crate::install::BOOT ("boot"), opening /{BOOT} like the other /boot users do; HOLD_PATH is gone.

Generated-by: https://github.com/cgwalters/#llms

alexlarsson and others added 3 commits September 24, 2026 13:06
This really is part of what an individual uki "boots".

Signed-off-by: Alexander Larsson <alexl@redhat.com>
Otherwise there is no difference in bootc status after a rollback.

Signed-off-by: Alexander Larsson <alexl@redhat.com>
On composefs installs /boot is usually the ESP, automounted by
systemd-gpt-auto-generator with a 2 minute idle timeout. If that
expire begins after the shutdown transaction is queued, but before
bootc-finalize-staged's ExecStop looks up /boot (lsblk statfs()es every
mountpoint), the two deadlock: the lookup blocks in
autofs_expire_wait until systemd unmounts boot.mount, and boot.mount is
ordered to stop after finalize via local-fs.target. After 5+5 minutes
finalize is killed, and the system boots the old deployment. That's
the "guest reboot timeout" flake seen mostly in plan-44, where the gap
between `bootc switch` and the reboot is about 130s in CI, on both
systemd-boot and grub legs.

If the expire instead completes before shutdown, systemd won't remount
/boot, so finalize fails immediately with "Failed to open /boot: Host
is down" on grub and the old deployment boots. That's the "expected
exactly one testbootcgroup" plan-44 failure on composefs grub legs.

See the analyses in
https://gist.github.com/cgwalters-bot/37b1789c26fd0af22f6879d93aa0a4bd
and
https://gist.github.com/cgwalters-bot/793b7485f34dd1586127982e40351e04

ostree hit the same problem and fixed it in ostreedev/ostree@f3db79e7
("finalize-staged: Ensure /boot automount doesn't expire") with
ostree-finalize-staged-hold.service, which keeps a file descriptor
for /boot open from the root mount namespace for as long as a
deployment is staged. Port that: bootc-finalize-staged.service now
pulls in bootc-finalize-staged-hold.service, which runs
`bootc composefs-finalize-staged --hold`. It's ordered before the
finalize unit, so at shutdown it's only stopped after finalization is
done. A busy /boot never expires, so it can't race with shutdown.
It would also make adding Requires=boot.mount to the finalize unit
(as bootc-dev#2488 does) safe, since without the hold an idle
expire stops the finalize unit while the system is still running.

Deviations from ostree: the hold only opens /boot rather than loading
the whole sysroot, since that's all that's needed and it keeps the
hold independent of storage setup (which may enter a private mount
namespace, where autofs wouldn't see the fd). For the same reason the
hold unit only has RequiresMountsFor=/boot, not /sysroot. And the
finalize unit keeps RequiresMountsFor=/sysroot without /boot: the stop
ordering against boot.mount already comes via local-fs.target, and
leaving out Requires=boot.mount means that if the hold somehow fails,
we're no worse off than before.

test-44 now checks that staging on composefs starts the hold unit.
Tested manually by rebooting inside the race window: that hung 4/4
times on main and 0/7 with this change, and plan-44 passed 3/3.

Generated-by: AI
@cgwalters-bot
cgwalters-bot force-pushed the bot/finalize-staged-hold branch from 6e8438f to 013018a Compare September 24, 2026 17:21
@cgwalters-bot

Copy link
Copy Markdown
Collaborator Author

Opened upstream as bootc-dev#2496. Closing this review draft.

@cgwalters-bot

Copy link
Copy Markdown
Collaborator Author

Both changes (trimmed commit message, reusing crate::install::BOOT) are applied in 31df035 on bootc-dev#2496. just validate and just unit-tests pass.

Generated-by: https://github.com/cgwalters/#llms

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants