Prepare OM4 snapshot dataset generation - #857
Conversation
alxmrs
left a comment
There was a problem hiding this comment.
Self-review, noticed a new source for freshwater fluxes.
|
🤖 Self-review follow-up is in 2568fe0. Changes made:
I intentionally left wfo optional pending the discussion with Will about a matching freshwater-flux source for the averaged run. Snapshot input continues to preserve wfo when present. |
|
@fomo-bot Will you review this PR? |
|
Reviewed current head Findings: no blocking code issues. The snapshot/averaged source selection, OM4 variable contract, final output validation, Zarr v2 write path, and Slurm harness changes look coherent against the current diff. The remaining Validation:
Non-blocking note: the PR body’s validation command omits the data package test extra; in this fresh checkout, I’d be comfortable with this after it is marked ready for review and the normal review gate is satisfied. Keep the post-merge real-source smoke test before any public bucket writes. Debug info
|
Summary
DATA_VARIANT=snapshotsmode to the Torch preprocessing and normalization harnessesom4_5daily_snapshots.zarrand deriveom4_<resolution>_snapshotsoutput directorieswfoforcing while excluding undeclared raw diagnosticsv2026-08four-resolution workflowThe source and temporal semantics follow the discussion in #450, especially:
Validation
uv run --project data pytest data/tests -q— 31 passed, 1 skippedbash -n scripts/slurm_preprocess_om4.sbatch scripts/slurm_make_norm_om4.sbatch— passedDeployment after review
No preprocessing jobs or public-bucket writes were performed for this PR. After approval/merge, the intended sequence is:
s3://m2lines-pubs/Samudra/v2026-08/om4_<resolution>_snapshots/