Add patched-JDK rollout e2e (tier 18) and fix two provisioner bugs - #140
Merged
Bruno Borges (brunoborges) merged 2 commits intoSep 30, 2026
Merged
Conversation
Add e2e tier 18. It provisions a Temurin 21 digest, runs a workload, patches the NodeProfile digest, and checks that the node re-advertises the new version. It also checks that the running pod keeps its JDK and that a restart moves the same app image onto the patched JDK. Tier 18 runs in the e2e workflow. Fix two provisioner bugs this exposed: - kubectl --request-timeout disables kubectl's in-cluster fallback, so every provisioner pod crashed at the ownership fence (localhost:8080). The provisioner now writes an explicit kubeconfig from the projected service-account token when it runs in-cluster. - The retired-root sweep matched paths in /proc/mounts, which keep the pre-rename name, so a root in use by a running pod was deleted once its grace expired. References are now matched by device and the bind mount's mountinfo root field, which follows the rename. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Take one snapshot of mount roots per sweep, reading each mount namespace once. A namespace now counts as seen only after one of its PIDs gives a readable mount table, so a PID that exits mid-scan can no longer hide a namespace that other PIDs still hold. The scan also no longer repeats the full /proc walk for each retired root. Resolve device numbers with GNU or BSD stat so the provisioner unit tests also run on macOS, and add a regression test for the exited-PID race. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Bruno Borges (brunoborges)
deleted the
brunoborges-jdk-patch-rollout-e2e
branch
September 30, 2026 21:15
3 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
No automated test covered the main JDK CVE workflow: change a NodeProfile JDK digest, have the node pick up the patched JDK, and restart workloads onto it. Only a provisioner unit test and a manual benchmark touched this path. This PR adds e2e tier 18 for that flow. Writing it turned up two provisioner bugs, and this PR fixes both.
New e2e tier 18 (
integration-tests/e2e/tier18-jdk-patch.sh).brewlet-sourcecarry the new digest;kubectl rollout restartmoves the same app image onto 21.0.12, after which an idle sweep reclaims the retired root.run.sh,reset.sh,integration-tests/AGENTS.md, and the e2e workflow matrix.Fix 1: provisioner could not reach the API server in-cluster
The recent
--request-timeout=10sflags inverify_node_ownershipstop kubectl from falling back to in-cluster config, so it triedlocalhost:8080. Every provisioner pod crashed withownership-fence-failed, and existing tier 14 failed the same way. When running in a pod withoutKUBECONFIG, the provisioner now writes an explicit kubeconfig that points at the projected service-account token and CA files. It reads the token file directly, so token rotation still works.Fix 2: the sweep deleted retired JDK roots still in use by running pods
reclaim_retired_rootslooked for the retired path in/proc/mounts. Overlaylowerdir=and mount sources keep the path from sandbox start, which after rotation is the new active root's name. The live-mount check therefore never matched, and a running pod's JDK was deleted once the 1h grace expired, for example at the next patch. The check now resolves the root's device and filesystem-relative path. It then scans every mount namespace'smountinfofor a bind mount whose root field lies under that path; the kernel renders this field from the live dentry, so it follows the rename. It still fails safe: if no mount table can be read, the root is kept. The unit tests used an unrealistic mount mock and have been rewritten to model the real layout, including a path-prefix sibling case.Validation
integration-tests/e2e/run.sh --reset --tier 18on kind v0.30.0 (arm64, single node): 19 passed, 0 failed. Before fix 2 the "sweep kept the root a running pod still uses" check failed as predicted.integration-tests/e2e/run.sh --reset --tier 14on the same cluster: 22 passed, 0 failed. Before fix 1 it failed at the ownership fence.bash provisioner/entrypoint_test.shpasses as a non-root user in agolang:1.26-bookwormcontainer. It was run from a copy on the container filesystem because macOS bind mounts break the permission-based test cases.bash -non all changed scripts. shellcheck was not run because it isn't installed locally.Compatibility and operations
--request-timeoutchange could not start. This restores provisioning without any RBAC change.Checklist