You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Part of #93. Existing coverage proves that the controller creates an HPA (integration-tests/e2e/tier4-k8s.sh), that serving Brewlet Deployments can be scaled manually (tier9-japp-serving.sh), and that runtime metrics can be scraped (tier15-metrics-incluster.sh). It does not prove the complete CPU-driven autoscaling loop or that the Brewlet controller preserves live HPA decisions.
Remediation Phase 9 is distinct from the existing E2E Tier 9.
Proposed solution
Add a dedicated E2E scenario in a disposable cluster with a pinned metrics-server deployment, a provisioned Brewlet runtime, and a small controllable CPU workload created through JavaApplication. Reuse an existing suitable fixture, or add a minimal fixture with an explicit load switch. Configure CPU requests, an achievable utilization target, bounded min/max replicas, and documented test stabilization behavior.
Drive load through a reproducible client and observe the chain: actual CPU usage -> Kubernetes resource metrics API -> HPA recommendation -> Deployment replica changes -> ready Brewlet Pods. Remove load and observe HPA-driven scale-down. Do not use kubectl scale, patch the Deployment replica count, or supply synthetic metrics as substitutes for the autoscaler.
Acceptance criteria:
Preflight confirms usable CPU resource metrics for the workload and sufficient disposable-cluster capacity. Missing metrics or prerequisites fail the required job with diagnostics rather than becoming a skip/pass.
Under controlled sustained load, real observed CPU utilization causes the HPA to scale above the initial replica count, within configured bounds; additional Pods run through Brewlet and become Ready.
With load removed, actual CPU utilization falls and the HPA returns to the configured minimum after its stabilization period, using bounded condition-based waits rather than a fixed sleep as the assertion.
Across multiple reconciliation cycles, the Brewlet controller does not reset the HPA-selected replica count to a stale JavaApplication.spec.replicas value. Service endpoints and JavaApplication readiness/status reflect the serving workload correctly.
Verify serving through the Service during scaling, record transition/request outcomes, and establish eventual correct endpoints/readiness after both scale-up and scale-down. Avoid claiming stronger availability guarantees than the workload's declared behavior.
Load generation and teardown are bounded; failures stop fixture load and leave no invocation-owned workers or cluster resources behind. Capture HPA conditions/events, resource metrics, desired/current/ready replicas, JavaApplication status, and relevant controller/kubelet/workload logs.
Add a pinned, reproducible invocation and dedicated daily/manual E2E coverage. Mandatory assertions cannot silently skip. Preserve the existing workflow's scheduled/manual-only trigger policy.
Record at least two consecutive successful runs on fresh disposable clusters and document prerequisites, stabilization windows, diagnostics, and coverage in the runbook and relevant user docs.
Coordinate with #13 on harness determinism rather than duplicating a broad E2E rewrite. Follow integration-tests/AGENTS.md for fixture/harness changes and use the GitHub Pages workflow as the documentation source of truth. Keep this milestone independently runnable from Phase 9A to make failures attributable.
Compatibility and operational impact
Use only explicitly addressed disposable nodes; do not install metrics-server or generate load on a developer's shared/default cluster. Keep test resource requirements modest and explicit. CPU resource metrics come from Kubernetes/metrics-server; Brewlet's runtime exporter is not a replacement for that metrics path.
Preserve intended controller/HPA ownership and existing API behavior. Fix demonstrated conflicts rather than bypassing the JavaApplication controller. Test-only stabilization tuning must be explicit and must not silently change production defaults. Any metrics-server certificate exception needed for local kind must remain fixture-scoped and documented.
Out of scope: node/cluster-autoscaler integration, custom/external-metrics adapters, CPU HPA scale-to-zero, production capacity sizing, and Phase 10 performance benchmarks.
Pre-submission checks
I searched existing issues and roadmap items for this request.
Problem
Part of #93. Existing coverage proves that the controller creates an HPA (
integration-tests/e2e/tier4-k8s.sh), that serving Brewlet Deployments can be scaled manually (tier9-japp-serving.sh), and that runtime metrics can be scraped (tier15-metrics-incluster.sh). It does not prove the complete CPU-driven autoscaling loop or that the Brewlet controller preserves live HPA decisions.Remediation Phase 9 is distinct from the existing E2E Tier 9.
Proposed solution
Add a dedicated E2E scenario in a disposable cluster with a pinned metrics-server deployment, a provisioned Brewlet runtime, and a small controllable CPU workload created through
JavaApplication. Reuse an existing suitable fixture, or add a minimal fixture with an explicit load switch. Configure CPU requests, an achievable utilization target, bounded min/max replicas, and documented test stabilization behavior.Drive load through a reproducible client and observe the chain: actual CPU usage -> Kubernetes resource metrics API -> HPA recommendation -> Deployment replica changes -> ready Brewlet Pods. Remove load and observe HPA-driven scale-down. Do not use
kubectl scale, patch the Deployment replica count, or supply synthetic metrics as substitutes for the autoscaler.Acceptance criteria:
JavaApplication.spec.replicasvalue. Service endpoints and JavaApplication readiness/status reflect the serving workload correctly.Coordinate with #13 on harness determinism rather than duplicating a broad E2E rewrite. Follow
integration-tests/AGENTS.mdfor fixture/harness changes and use the GitHub Pages workflow as the documentation source of truth. Keep this milestone independently runnable from Phase 9A to make failures attributable.Compatibility and operational impact
Use only explicitly addressed disposable nodes; do not install metrics-server or generate load on a developer's shared/default cluster. Keep test resource requirements modest and explicit. CPU resource metrics come from Kubernetes/metrics-server; Brewlet's runtime exporter is not a replacement for that metrics path.
Preserve intended controller/HPA ownership and existing API behavior. Fix demonstrated conflicts rather than bypassing the JavaApplication controller. Test-only stabilization tuning must be explicit and must not silently change production defaults. Any metrics-server certificate exception needed for local kind must remain fixture-scoped and documented.
Out of scope: node/cluster-autoscaler integration, custom/external-metrics adapters, CPU HPA scale-to-zero, production capacity sizing, and Phase 10 performance benchmarks.
Pre-submission checks