Skip to content

Build an anchor to hold a netCDF file open in each batch Dask worker - #747

Open
rhaegar325 wants to merge 3 commits into
mainfrom
add_anchor_while_open_netcdf_to_avoid_netcdf_lock_issue
Open

rhaegar325 wants to merge 3 commits into
mainfrom
add_anchor_while_open_netcdf_to_avoid_netcdf_lock_issue

Conversation

@rhaegar325

@rhaegar325 rhaegar325 commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Fix #740

Problem

Batch jobs intermittently fail with RuntimeError: NetCDF: Not a valid ID,
raised from netCDF4.Dataset.__init__ (_get_format / _get_vars) after an
xarray file-cache miss. A plain rerun usually succeeds. Seen on direct
3D ocean variables with source partitioning (thkcello, wo, thetao, vo,
wmo, so) across piControl, esm-piControl and esm-hist.

netCDF-C keeps open files in a process-global table with no lock, and frees the
whole table when the open-file count reaches zero (del_from_NCList →
free_NCList in libdispatch/nclistmgr.c; unchanged on main). In a worker,
the compute thread opens files while xarray's CachingFileManager.__del__
closes them from another thread without the backend lock. If that close takes
the count to zero during an nc_open, the new ncid stops resolving.

In an isolated netCDF4 test (two threads opening/closing, nothing else open),
10/10 runs crashed; holding one extra file open made 10/10 clean.

Change

  • executors/nc_anchor.py: NetCDFAnchorPlugin, a WorkerPlugin that holds
    a tiny read-only netCDF file in the worker's local directory, so the count
    never reaches zero.
  • Worker template: register it right after the client starts, before any file
    is opened, and log netCDF anchor held on N/N workers.

A plugin rather than a preload: the scheduler re-applies it to workers the
nanny restarts, and no config is needed.

This is a workaround for the unsynchronised close, not a fix for it; a lost
update on the non-atomic counter could in principle still hit zero.

Testing

  • Plugin setup/teardown, including an existing anchor file.
  • LocalCluster(processes=True): every worker holds the anchor, and still
    does after client.restart().
  • Template: registration sits between dd.Client( and the first CMORiser.
  • Full unit suite: 2256 passed.
  • Not reproduced end to end: the race never fired in Dask-level stress tests
    (~46k slices) with or without the anchor, so production failure rates are
    the real check. The same anchor has been running via
    DASK_DISTRIBUTED__WORKER__PRELOAD on 1pctCO2, abrupt-4xCO2 and esm-flat10.

@codecov

codecov Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 79.8%. Comparing base (a78c3fc) to head (5ebf12f).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@           Coverage Diff           @@
##            main    #747     +/-   ##
=======================================
+ Coverage   79.7%   79.8%   +0.1%     
=======================================
  Files         41      42      +1     
  Lines       9173    9199     +26     
  Branches    1710    1713      +3     
=======================================
+ Hits        7315    7345     +30     
+ Misses      1524    1522      -2     
+ Partials     334     332      -2     
Flag Coverage Δ
unit 79.8% <100.0%> (+0.1%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@rhaegar325
rhaegar325 requested a review from rbeucher October 9, 2026 05:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Random NetCDF: Not a valid ID failures during CMORisation

1 participant