fix(worker): Only publish occupancy between first task and drain - #796
Merged
enochtangg merged 2 commits intoSep 25, 2026
Merged
enochtangg merged 2 commits into
enochtangg merged 2 commits into
Conversation
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Want reviews to match your repository better? Bugbot Learning can learn team-specific rules from PR activity. A team admin can enable Learning in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit ef07cd0. Configure here.
evanh
approved these changes
Sep 25, 2026
enochtangg
deleted the
enochtang/stream-1843-make-taskworker-occupancy-robust-to-pod-churn
branch
September 25, 2026 17:35
3 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Pods that are not handling traffic still publish occupancy, which pulls the KEDA pool average down during rollouts and blocks scale-up. This limits publishing to the window where a pod's slots actually reflect pool load.
Change
A pod now waits until it accepts its first task before it publishes occupancy. A push worker can only receive a task after it is SERVING and the broker is routing to it. So publishing starts once the pod is actually doing work, not when its first child finishes warming up.
A pod stops publishing occupancy as soon as it starts draining. Instead of leaving the last value in place, the worker removes the Prometheus series, so Prometheus marks it stale and it drops out of the average. Only the metrics thread writes the gauge, and it removes the series on every flush once draining has started.
Draining starts as soon as the worker sees the file
/tmp/taskworker-draining. The pod's preStop hook creates this file when the pod is deleted, which currently is set to 15 seconds before SIGTERM arrives. That change is in a separate ops PR.A new metric,
taskworker.worker.occupancy.withheld, counts the times a pod skips publishing, tagged with the reason (no_task_yetorstopped). This makes the new behaviour easy to check during a rollout.