Skip to content

fix(lcb-service): use the fork start method explicitly - #433

Open
liayan wants to merge 2 commits into
mlcommons:mainfrom
liayan:fix/lcb-service-fork-start-method
Open

fix(lcb-service): use the fork start method explicitly#433
liayan wants to merge 2 commits into
mlcommons:mainfrom
liayan:fix/lcb-service-fork-start-method

Conversation

@liayan

@liayan liayan commented Jul 29, 2026

Copy link
Copy Markdown

Python 3.14 changed the default multiprocessing start method on Linux from fork to forkserver. The grading pipeline (pool workers forking a per-problem mp.Process + mp.Manager) only works with fork: under forkserver the grading children die at startup, every result comes back as an error, and execute_code_single_suppressed_errors turns that into all-failed tests, so the service sits at 0/N forever.

Pin the fork context for the executor, the per-problem Process and its Manager. Also raise if every subprocess reported an execution error -- that means the judge is broken, not that all samples failed -- and log those errors at error level instead of warning.

Seen on a python 3.14 lcb-service image: 0/349 after 3.5h, one defunct child per pool worker. Same inputs with fork forced: done in 6 min. The repo pins 3.12 so CI won't hit this, but shipped images have.

What does this PR do?

Type of change

  • Bug fix
  • New feature
  • Documentation update
  • Refactor/cleanup

Related issues

Testing

  • Tests added/updated
  • All tests pass locally
  • Manual testing completed

Checklist

  • Code follows project style
  • Pre-commit hooks pass
  • Documentation updated (if needed)

@liayan
liayan requested a review from a team July 29, 2026 19:32
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@github-actions

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

@codecov-commenter

codecov-commenter commented Jul 29, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
⚠️ Please upload report for BASE (main@1df1bb2). Learn more about missing BASE report.

Additional details and impacted files
@@           Coverage Diff           @@
##             main     #433   +/-   ##
=======================================
  Coverage        ?   81.70%           
=======================================
  Files           ?      146           
  Lines           ?    19343           
  Branches        ?        0           
=======================================
  Hits            ?    15804           
  Misses          ?     3539           
  Partials        ?        0           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@liayan
liayan force-pushed the fix/lcb-service-fork-start-method branch from e9aa1db to 4ff0df4 Compare July 29, 2026 19:42
@liayan

liayan commented Aug 4, 2026

Copy link
Copy Markdown
Author

Verified on the Python 3.14 lcb-service image (forkserver default): grading works with the fork pin; dead grading children classify as -6, and the all-infra-errors check raises, so the original failure is still caught; an all-timeout batch scores 0 without raising — that case would have raised on commit one fix.

@liayan
liayan force-pushed the fix/lcb-service-fork-start-method branch 2 times, most recently from 47115a5 to 9c6739e Compare August 4, 2026 14:26
@nvzhihanj
nvzhihanj requested a review from hvagadia August 4, 2026 18:33
liayan added 2 commits August 4, 2026 17:46
Python 3.14 changed the default multiprocessing start method on Linux
from fork to forkserver. The grading pipeline (pool workers forking a
per-problem mp.Process + mp.Manager) only works with fork: under
forkserver the grading children die at startup, every result comes back
as an error, and execute_code_single_suppressed_errors turns that into
all-failed tests, so the service sits at 0/N forever.

Pin the fork context for the executor, the per-problem Process and its
Manager. Also raise if every subprocess reported an execution error --
that means the judge is broken, not that all samples failed -- and log
those errors at error level instead of warning.

Seen on a python 3.14 lcb-service image: 0/349 after 3.5h, one defunct
child per pool worker. Same inputs with fork forced: done in 6 min.
The repo pins 3.12 so CI won't hit this, but shipped images have.
Timeouts were counted as execution errors, so a small batch where every
submission loops forever would trip the guard and raise instead of
scoring 0. Split the empty-buffer case in run_code_subprocess: child
still alive at the deadline -> timeout (-1, submission's fault), child
exited without reporting -> new GradingChildDied (-6, judge's fault).
The guard now only counts -5/-6, so the forkserver startup deaths still
raise and all-timeout batches score normally.

Also log the multiprocessing start method at service init; that would
have made the original 0/N a one-line diagnosis.
@liayan
liayan force-pushed the fix/lcb-service-fork-start-method branch from 9c6739e to 3dfa3ca Compare August 4, 2026 23:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants