Skip to content

Add patched-JDK rollout e2e (tier 18) and fix two provisioner bugs - #140

Merged
Bruno Borges (brunoborges) merged 2 commits into
mainfrom
brunoborges-jdk-patch-rollout-e2e
Sep 30, 2026
Merged

Bruno Borges (brunoborges) merged 2 commits into
mainfrom
brunoborges-jdk-patch-rollout-e2e

Conversation

@brunoborges

Copy link
Copy Markdown
Member

Summary

No automated test covered the main JDK CVE workflow: change a NodeProfile JDK digest, have the node pick up the patched JDK, and restart workloads onto it. Only a provisioner unit test and a manual benchmark touched this path. This PR adds e2e tier 18 for that flow. Writing it turned up two provisioner bugs, and this PR fixes both.

New e2e tier 18 (integration-tests/e2e/tier18-jdk-patch.sh)

  • Provisions Temurin 21.0.11 by multi-arch digest and runs a Brewlet workload on it.
  • Patches the NodeProfile to the 21.0.12 digest, then asserts that:
    • the node re-advertises 21.0.12 for the new profile generation;
    • the DaemonSet env and .brewlet-source carry the new digest;
    • the previous root is retired, not deleted.
  • Asserts the running pod keeps its container and still reports 21.0.11.
  • Asserts a retired-root sweep keeps the root that pod still uses.
  • Asserts kubectl rollout restart moves the same app image onto 21.0.12, after which an idle sweep reclaims the retired root.
  • Wired into run.sh, reset.sh, integration-tests/AGENTS.md, and the e2e workflow matrix.

Fix 1: provisioner could not reach the API server in-cluster
The recent --request-timeout=10s flags in verify_node_ownership stop kubectl from falling back to in-cluster config, so it tried localhost:8080. Every provisioner pod crashed with ownership-fence-failed, and existing tier 14 failed the same way. When running in a pod without KUBECONFIG, the provisioner now writes an explicit kubeconfig that points at the projected service-account token and CA files. It reads the token file directly, so token rotation still works.

Fix 2: the sweep deleted retired JDK roots still in use by running pods
reclaim_retired_roots looked for the retired path in /proc/mounts. Overlay lowerdir= and mount sources keep the path from sandbox start, which after rotation is the new active root's name. The live-mount check therefore never matched, and a running pod's JDK was deleted once the 1h grace expired, for example at the next patch. The check now resolves the root's device and filesystem-relative path. It then scans every mount namespace's mountinfo for a bind mount whose root field lies under that path; the kernel renders this field from the live dentry, so it follows the rename. It still fails safe: if no mount table can be read, the root is kept. The unit tests used an unrealistic mount mock and have been rewritten to model the real layout, including a path-prefix sibling case.

Validation

  • integration-tests/e2e/run.sh --reset --tier 18 on kind v0.30.0 (arm64, single node): 19 passed, 0 failed. Before fix 2 the "sweep kept the root a running pod still uses" check failed as predicted.
  • integration-tests/e2e/run.sh --reset --tier 14 on the same cluster: 22 passed, 0 failed. Before fix 1 it failed at the ownership fence.
  • bash provisioner/entrypoint_test.sh passes as a non-root user in a golang:1.26-bookworm container. It was run from a copy on the container filesystem because macOS bind mounts break the permission-based test cases.
  • bash -n on all changed scripts. shellcheck was not run because it isn't installed locally.

Compatibility and operations

  • Provisioner pods on released images with the --request-timeout change could not start. This restores provisioning without any RBAC change.
  • The retired-root sweep now keeps roots that running pods still use, instead of deleting them after the grace period. Disk is reclaimed on a later pass once those pods stop.
  • CI adds a "Tier 18: Patched JDK rollout" e2e job, which runs on schedule or manual dispatch only.

Checklist

  • Tests cover changed behavior.
  • Documentation is updated when needed.
  • No credentials, proprietary data, or unrelated generated files are included.

Add e2e tier 18. It provisions a Temurin 21 digest, runs a workload,
patches the NodeProfile digest, and checks that the node re-advertises the
new version. It also checks that the running pod keeps its JDK and that a
restart moves the same app image onto the patched JDK. Tier 18 runs in the
e2e workflow.

Fix two provisioner bugs this exposed:
- kubectl --request-timeout disables kubectl's in-cluster fallback, so
  every provisioner pod crashed at the ownership fence (localhost:8080).
  The provisioner now writes an explicit kubeconfig from the projected
  service-account token when it runs in-cluster.
- The retired-root sweep matched paths in /proc/mounts, which keep the
  pre-rename name, so a root in use by a running pod was deleted once its
  grace expired. References are now matched by device and the bind
  mount's mountinfo root field, which follows the rename.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Take one snapshot of mount roots per sweep, reading each mount namespace
once. A namespace now counts as seen only after one of its PIDs gives a
readable mount table, so a PID that exits mid-scan can no longer hide a
namespace that other PIDs still hold. The scan also no longer repeats the
full /proc walk for each retired root.

Resolve device numbers with GNU or BSD stat so the provisioner unit tests
also run on macOS, and add a regression test for the exited-PID race.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@brunoborges
Bruno Borges (brunoborges) merged commit 3c1d894 into main Sep 30, 2026
14 checks passed
@brunoborges
Bruno Borges (brunoborges) deleted the brunoborges-jdk-patch-rollout-e2e branch September 30, 2026 21:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant