perf: queue tasks lock-free and size workers to GOMAXPROCS - #1
Merged
Merged
Conversation
Short tasks saw a long scheduling tail under CPU saturation. The worker target grew whenever a sampled task took 250 µs of wall time, which saturation alone produces, so the pool kept adding runnable workers that only lengthened the Go scheduler's run queues. And every submission took the queue mutex, which workers also held to take tasks, so under load producers queued on it. - Replace the mutex-guarded rings with a lock-free MPMC FIFO: a chain of bounded rings that doubles when full and shrinks back once a burst has drained. Batch submissions claim consecutive slots with one CAS, and submitting takes no lock while enough workers are running. - Base the running-worker target on GOMAXPROCS instead of twice that. A monitor, active only while tasks wait, adds workers when running ones block, when Ps sit idle (scheduler counts from runtime/metrics on Go 1.26+, the monitor's own wake-up lateness otherwise), or when tasks arrive more than twice as fast as they start, and throttles submitters past 65,536 queued tasks. WithConcurrency remains the hard cap on live workers. - Workers take one task at a time and claim small batches only when they contend on the queue head. Parked workers are still woken LIFO, so WithMaxIdle stack reuse and WithMaxJobs behave as before. - Queue and Task[T] share one implementation; the exported API is unchanged. - Make the stall reproduction's release deterministic and add a variant where the fast tasks queue behind blocking tasks that have not started. - README: document the new design and refresh the Reuse/NoReuse benchmarks.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Short tasks saw a long scheduling tail under CPU saturation. The worker target grew whenever a sampled task took 250 µs of wall time, which saturation alone produces, so the pool kept adding runnable workers that only lengthened the Go scheduler's run queues. And every submission took the queue mutex, which workers also held to take tasks, so under load producers queued on it.