Skip to content

perf: queue tasks lock-free and size workers to GOMAXPROCS - #1

Merged
limpo1989 merged 1 commit into
masterfrom
perf/lock-free-dispatch
Sep 25, 2026
Merged

limpo1989 merged 1 commit into
masterfrom
perf/lock-free-dispatch

Conversation

@limpo1989

Copy link
Copy Markdown
Owner

Short tasks saw a long scheduling tail under CPU saturation. The worker target grew whenever a sampled task took 250 µs of wall time, which saturation alone produces, so the pool kept adding runnable workers that only lengthened the Go scheduler's run queues. And every submission took the queue mutex, which workers also held to take tasks, so under load producers queued on it.

  • Replace the mutex-guarded rings with a lock-free MPMC FIFO: a chain of bounded rings that doubles when full and shrinks back once a burst has drained. Batch submissions claim consecutive slots with one CAS, and submitting takes no lock while enough workers are running.
  • Base the running-worker target on GOMAXPROCS instead of twice that. A monitor, active only while tasks wait, adds workers when running ones block, when Ps sit idle (scheduler counts from runtime/metrics on Go 1.26+, the monitor's own wake-up lateness otherwise), or when tasks arrive more than twice as fast as they start, and throttles submitters past 65,536 queued tasks. WithConcurrency remains the hard cap on live workers.
  • Workers take one task at a time and claim small batches only when they contend on the queue head. Parked workers are still woken LIFO, so WithMaxIdle stack reuse and WithMaxJobs behave as before.
  • Queue and Task[T] share one implementation; the exported API is unchanged.
  • Make the stall reproduction's release deterministic and add a variant where the fast tasks queue behind blocking tasks that have not started.
  • README: document the new design and refresh the Reuse/NoReuse benchmarks.

Short tasks saw a long scheduling tail under CPU saturation. The worker
target grew whenever a sampled task took 250 µs of wall time, which
saturation alone produces, so the pool kept adding runnable workers that
only lengthened the Go scheduler's run queues. And every submission took
the queue mutex, which workers also held to take tasks, so under load
producers queued on it.

- Replace the mutex-guarded rings with a lock-free MPMC FIFO: a chain of
  bounded rings that doubles when full and shrinks back once a burst has
  drained. Batch submissions claim consecutive slots with one CAS, and
  submitting takes no lock while enough workers are running.
- Base the running-worker target on GOMAXPROCS instead of twice that. A
  monitor, active only while tasks wait, adds workers when running ones
  block, when Ps sit idle (scheduler counts from runtime/metrics on Go
  1.26+, the monitor's own wake-up lateness otherwise), or when tasks
  arrive more than twice as fast as they start, and throttles submitters
  past 65,536 queued tasks. WithConcurrency remains the hard cap on live
  workers.
- Workers take one task at a time and claim small batches only when they
  contend on the queue head. Parked workers are still woken LIFO, so
  WithMaxIdle stack reuse and WithMaxJobs behave as before.
- Queue and Task[T] share one implementation; the exported API is
  unchanged.
- Make the stall reproduction's release deterministic and add a variant
  where the fast tasks queue behind blocking tasks that have not started.
- README: document the new design and refresh the Reuse/NoReuse
  benchmarks.
@limpo1989
limpo1989 merged commit 2045c49 into master Sep 25, 2026
20 checks passed
@limpo1989
limpo1989 deleted the perf/lock-free-dispatch branch September 25, 2026 09:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant