From 429b1abda3645315bb4471c3edb851f6bb2e64e3 Mon Sep 17 00:00:00 2001 From: Jon Froehlich Date: Sat, 29 Aug 2026 04:21:54 -0700 Subject: [PATCH 1/9] Rebuild the purged YOLO labels from the Ultralytics caches (#51) Every label directory under /gscratch/scrubbed/jfroehli/yolo/ is empty -- all five dataset variants, train and val -- while every image directory is intact (557,413 train tiles, 161,002 val). That is why y11x_tiles (38612069) and its two retries (38657533, 38657535) all died at dataset init with ValueError: train: No labels found in .../yolo/tiles/labels/train.cache The mechanism is worth naming because it will recur. `scrubbed` purges by access time, and the .cache files are precisely what stopped anything from reading the individual label .txt files after 2026-07-25. Their atime froze while the images kept being read every epoch, so the purge took the labels and left the images. **The cache that made training fast is what got the labels deleted.** Same shape as the partially-purged conda package cache that broke the env build earlier in #51: a populated cache actively hides the thing it caches from whatever decides what is cold, so a cache is not a backup and its presence is not evidence its source exists. Recovery needs no GPU and no re-derivation. Ultralytics' cache is not a digest -- it stores parsed labels, `cls` (n,1) and `bboxes` (n,4) as normalized xywh, which is the on-disk .txt format itself. 968,227 boxes over 557,413 train records are all there. The write format is lifted from prepare_yolo_dataset.py::_write_pair so rebuilt files are byte-identical to the originals for any value surviving the float32 round trip, including the zero-byte file a background tile gets. That distinction is load-bearing: Ultralytics counts a missing label as `nm` and an empty one as `nf`, and the error above fires when `nf == 0`. Two deliberate choices: - **Read the durable cache copies under /gscratch/makelab**, not the ones on scrubbed. The scrubbed copies are the siblings of what was already lost once; rebuilding from them would make the recovery depend on the thing that failed. The rescued pair was staged 2026-08-17 with sha256sums.txt, which the launcher checks before trusting it. - **Rebuild rather than re-run prepare_yolo_dataset.py.** The cache reproduces the labels the published arms actually trained on; re-deriving would produce labels from today's code and geometry constants. Those should agree, but the arms in flight were trained on these. --verify re-reads every written file and compares the boxes back against the cache, which is the only check that proves the round trip rather than assuming it. The rebuild is a batch job because it creates ~718,000 small files on GPFS, which is sustained and metadata-heavy; klone reaps heavy login processes and that reap also kills the SSH control master. Co-Authored-By: Claude Opus 5 --- .../rebuild_yolo_labels_from_cache.py | 162 ++++++++++++++++++ .../run_rebuild_yolo_labels.slurm | 82 +++++++++ 2 files changed, 244 insertions(+) create mode 100644 scripts/model_comparison/rebuild_yolo_labels_from_cache.py create mode 100644 scripts/model_comparison/run_rebuild_yolo_labels.slurm diff --git a/scripts/model_comparison/rebuild_yolo_labels_from_cache.py b/scripts/model_comparison/rebuild_yolo_labels_from_cache.py new file mode 100644 index 00000000..dd2a96da --- /dev/null +++ b/scripts/model_comparison/rebuild_yolo_labels_from_cache.py @@ -0,0 +1,162 @@ +#!/usr/bin/env python3 +"""Rebuild YOLO label .txt files from an Ultralytics label cache (#51). + +WHY THIS EXISTS +--------------- +On 2026-08-19 the `y11x_tiles` arm failed at dataset init with + + ValueError: train: No labels found in .../yolo/tiles/labels/train.cache + +Every label directory under `/gscratch/scrubbed/jfroehli/yolo/` was empty -- all five +dataset variants, train and val -- while every image directory was fully intact +(557,413 train tiles, 161,002 val). `/gscratch/scrubbed` purges by access time, and +the `.cache` files are exactly what stopped anything from reading the individual +label `.txt` files after 2026-07-25. Their atime froze while the images kept being +read every epoch, so the purge took the labels and left the images. + +**The cache that made training fast is what got the labels deleted.** That is the +generalisable part, and it is the same shape as the partially-purged conda package +cache that broke the env build in #51 -- a cache is not a backup, and a *populated* +cache actively hides the thing it caches from anything that decides what is cold. + +WHY THE CACHE IS A SUFFICIENT SOURCE +------------------------------------ +Ultralytics' cache is not a digest: it stores the fully parsed labels. Each record +carries `cls` (n,1) and `bboxes` (n,4) as normalized xywh -- which is the on-disk +`.txt` format itself. So the labels are recoverable exactly, with no re-derivation +from the source panoramas and no GPU. + +That matters for provenance as much as for cost: rebuilding from the cache reproduces +the labels the published `y11x_tiles` / `y26_pano` arms actually trained on, whereas +re-running `prepare_yolo_dataset.py` would produce labels from today's code and today's +geometry constants. Those should agree, but "should" is not "do", and the arms in +flight were trained on these. + +The write format below is lifted from `prepare_yolo_dataset.py::_write_pair` and the +line construction beside it, so a rebuilt file is byte-identical to the original for +any value that survives the float32 round trip through the cache: + + line = f"0 {u:.6f} {v:.6f} {w:.6f} {h:.6f}" + body = "\n".join(lines) + ("\n" if lines else "") + +Background tiles therefore get a **zero-byte file, not a missing one**. That is +deliberate: Ultralytics counts a missing label as `nm` (missing) and an empty one as +`nf` (found, no objects), and the failure above is raised when `nf == 0`. + +USAGE +----- + python rebuild_yolo_labels_from_cache.py \ + --cache /gscratch/makelab/jonf/rampnet_yolo_baseline_51/label_cache_rescue/train.cache \ + --labels-dir /gscratch/scrubbed/jfroehli/yolo/tiles/labels/train \ + --verify + +Prefer the durable copy of the cache under `/gscratch/makelab` (purchased, never +purged) over the one on `scrubbed`, which is one purge window from being the same +problem again. + +`--verify` re-reads every file it wrote and compares the parsed boxes back against the +cache, which is the only check that actually proves the round trip rather than assuming +it. It roughly doubles the runtime; run it at least once per rebuilt split. +""" +from __future__ import annotations + +import argparse +import os +import sys +from pathlib import Path + + +def load_cache(path: Path): + """Return the list of label records from an Ultralytics *.cache file.""" + import numpy as np + obj = np.load(str(path), allow_pickle=True).item() + labels = obj.get("labels") + if labels is None: + raise SystemExit(f"{path}: no 'labels' key -- not an Ultralytics label cache") + return labels + + +def format_record(rec) -> str: + """Render one cache record as the .txt body prepare_yolo_dataset.py would have written.""" + cls, boxes = rec["cls"], rec["bboxes"] + lines = [ + f"{int(cls[i][0])} {boxes[i][0]:.6f} {boxes[i][1]:.6f} {boxes[i][2]:.6f} {boxes[i][3]:.6f}" + for i in range(len(boxes)) + ] + return "\n".join(lines) + ("\n" if lines else "") + + +def main() -> int: + ap = argparse.ArgumentParser(description=__doc__.split("\n")[0]) + ap.add_argument("--cache", type=Path, required=True, help="Ultralytics *.cache to read") + ap.add_argument("--labels-dir", type=Path, required=True, help="directory to write *.txt into") + ap.add_argument("--verify", action="store_true", + help="re-read every written file and compare boxes back to the cache") + ap.add_argument("--dry-run", action="store_true", help="report what would be written, write nothing") + ap.add_argument("--progress-every", type=int, default=50000, help="progress line cadence") + args = ap.parse_args() + + labels = load_cache(args.cache) + print(f"cache : {args.cache}") + print(f"records : {len(labels)}") + total_boxes = sum(len(r["bboxes"]) for r in labels) + print(f"boxes : {total_boxes}") + print(f"labels-dir : {args.labels_dir}") + if args.dry_run: + print("DRY RUN -- nothing written") + return 0 + + args.labels_dir.mkdir(parents=True, exist_ok=True) + written = empty = 0 + for i, rec in enumerate(labels, 1): + stem = Path(rec["im_file"]).stem + body = format_record(rec) + (args.labels_dir / f"{stem}.txt").write_text(body) + written += 1 + if not body: + empty += 1 + if args.progress_every and i % args.progress_every == 0: + print(f" ... {i}/{len(labels)}", flush=True) + + print(f"written : {written} ({empty} background/empty, {written - empty} with boxes)") + + if not args.verify: + print("OK (unverified -- pass --verify to prove the round trip)") + return 0 + + print("verifying ...", flush=True) + bad = 0 + seen_boxes = 0 + for i, rec in enumerate(labels, 1): + stem = Path(rec["im_file"]).stem + text = (args.labels_dir / f"{stem}.txt").read_text() + got = [ln.split() for ln in text.splitlines() if ln.strip()] + exp = rec["bboxes"] + if len(got) != len(exp): + print(f" MISMATCH count {stem}: file {len(got)} vs cache {len(exp)}") + bad += 1 + continue + for j, parts in enumerate(got): + vals = [float(p) for p in parts[1:5]] + for k in range(4): + # .6f is the on-disk precision, so agreement is bounded by rounding, + # not by anything about the data. + if abs(vals[k] - float(exp[j][k])) > 1e-6: + print(f" MISMATCH value {stem} box {j} coord {k}: " + f"{vals[k]} vs {float(exp[j][k])}") + bad += 1 + break + seen_boxes += len(got) + if args.progress_every and i % args.progress_every == 0: + print(f" ... verified {i}/{len(labels)}", flush=True) + + print(f"verified : {len(labels)} files, {seen_boxes} boxes, {bad} mismatched") + if bad or seen_boxes != total_boxes: + print("FAIL") + return 1 + print("PASS -- every file round-trips to the cache it came from") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/model_comparison/run_rebuild_yolo_labels.slurm b/scripts/model_comparison/run_rebuild_yolo_labels.slurm new file mode 100644 index 00000000..505c2877 --- /dev/null +++ b/scripts/model_comparison/run_rebuild_yolo_labels.slurm @@ -0,0 +1,82 @@ +#!/bin/bash +# Rebuild the purged YOLO tile labels from the rescued Ultralytics caches (#51). +# +# This is a batch job rather than a login-node one-liner for one reason: it creates +# ~718,000 small files on GPFS (557,413 train + 161,002 val), which is metadata-heavy +# and sustained. klone reaps heavy login processes and that reap also kills the SSH +# control master. +# +# It reads the DURABLE cache copies under /gscratch/makelab (purchased, never purged), +# not the ones on /gscratch/scrubbed -- those are the copies whose siblings were already +# lost once, and rebuilding from them would make the recovery depend on the thing that +# failed. The rescued copies were staged 2026-08-17 with sha256sums.txt beside them. +# +# sbatch scripts/model_comparison/run_rebuild_yolo_labels.slurm +# +# Env overrides: RAMPNET_REPO, YOLO_CACHE_DIR, YOLO_DATA_ROOT. +# +#SBATCH --job-name=yolo_label_rebuild_51 +#SBATCH --account=ckpt-makelab +#SBATCH --partition=ckpt-all +#SBATCH --nodes=1 +#SBATCH --ntasks=1 +#SBATCH --cpus-per-task=4 +#SBATCH --mem=16G +#SBATCH --time=6:00:00 +#SBATCH --output=logs/yolo_label_rebuild_%j.out +#SBATCH --error=logs/yolo_label_rebuild_%j.err +#SBATCH --requeue + +set -eo pipefail + +REPO="${RAMPNET_REPO:-/gscratch/makelab/jonf/lrcheck_135}" +CACHE_DIR="${YOLO_CACHE_DIR:-/gscratch/makelab/jonf/rampnet_yolo_baseline_51/label_cache_rescue}" +DATA_ROOT="${YOLO_DATA_ROOT:-/gscratch/scrubbed/jfroehli/yolo/tiles}" +PY="${YOLO_PY:-/gscratch/makelab/jonf/envs/yolo/bin/python}" +SCRIPT="$REPO/scripts/model_comparison/rebuild_yolo_labels_from_cache.py" + +echo "--- YOLO label rebuild (#51) ---" +echo "Job ID : ${SLURM_JOBID}" +echo "Node : ${SLURMD_NODENAME:-?}" +echo "Cache dir : $CACHE_DIR" +echo "Data root : $DATA_ROOT" +echo "Script : $SCRIPT" +echo "Python : $PY" +echo "-------------------------------" + +# The rescued caches carry their own checksums. Verify before trusting them as the +# single source for 968,227 boxes. +if [ -f "$CACHE_DIR/sha256sums.txt" ]; then + echo "=== verifying rescued caches against sha256sums.txt ===" + ( cd "$CACHE_DIR" && sha256sum -c sha256sums.txt ) || { echo "FATAL: cache checksum mismatch"; exit 1; } +else + echo "WARNING: no sha256sums.txt beside the caches -- proceeding unverified" +fi + +for split in train val; do + echo + echo "=== $split $(date -Is) ===" + "$PY" "$SCRIPT" \ + --cache "$CACHE_DIR/${split}.cache" \ + --labels-dir "$DATA_ROOT/labels/${split}" \ + --verify +done + +echo +echo "=== stale .cache files ===" +# Ultralytics rebuilds a cache whose hash no longer matches, so these are not harmful -- +# but they are what made the failure confusing, so move them aside rather than trust them. +for split in train val; do + if [ -f "$DATA_ROOT/labels/${split}.cache" ]; then + mv -v "$DATA_ROOT/labels/${split}.cache" "$DATA_ROOT/labels/${split}.cache.pre-rebuild" + fi +done + +echo +echo "=== final counts ===" +for split in train val; do + echo " labels/$split : $(ls "$DATA_ROOT/labels/$split" | wc -l) files" + echo " images/$split : $(ls "$DATA_ROOT/images/$split" | wc -l) files" +done + +echo "--- done $(date -Is) ---" From 1f4b07bfdee786053b1433ddfe85052007b0fe14 Mon Sep 17 00:00:00 2001 From: Jon Froehlich Date: Sat, 29 Aug 2026 04:22:37 -0700 Subject: [PATCH 2/9] Default RAMPNET_REPO to $HOME/RampNet, as the other klone launchers do (#51) The first version pointed at an ad-hoc staging directory from the session that wrote it, which is exactly the 'configured by edits made during a session' shape the replication rule exists to prevent. Every other klone launcher in the repo defaults to $HOME/RampNet; match them. Co-Authored-By: Claude Opus 5 --- scripts/model_comparison/run_rebuild_yolo_labels.slurm | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scripts/model_comparison/run_rebuild_yolo_labels.slurm b/scripts/model_comparison/run_rebuild_yolo_labels.slurm index 505c2877..504cfd50 100644 --- a/scripts/model_comparison/run_rebuild_yolo_labels.slurm +++ b/scripts/model_comparison/run_rebuild_yolo_labels.slurm @@ -29,7 +29,7 @@ set -eo pipefail -REPO="${RAMPNET_REPO:-/gscratch/makelab/jonf/lrcheck_135}" +REPO="${RAMPNET_REPO:-$HOME/RampNet}" CACHE_DIR="${YOLO_CACHE_DIR:-/gscratch/makelab/jonf/rampnet_yolo_baseline_51/label_cache_rescue}" DATA_ROOT="${YOLO_DATA_ROOT:-/gscratch/scrubbed/jfroehli/yolo/tiles}" PY="${YOLO_PY:-/gscratch/makelab/jonf/envs/yolo/bin/python}" From 4168a4bf64939ecda839b3bbe8d53ec1cd87c3ac Mon Sep 17 00:00:00 2001 From: Jon Froehlich Date: Sat, 29 Aug 2026 04:41:51 -0700 Subject: [PATCH 3/9] Acceptance test: does Ultralytics load the rebuilt dataset? (#51) The rebuild verifies its output against the cache it read, which proves the round trip but not the thing that actually broke -- the failure was inside Ultralytics' dataset init, so the only test that closes the loop is making Ultralytics build the dataset and report a non-zero nf. Two checks a bare 'did it crash' run would miss: - **nf against nm.** Ultralytics treats a MISSING label file as a background image and carries on, so a rebuild that skipped the label-less tiles would still 'work' while silently changing the dataset. 59,923 of 557,413 train tiles are background and must be present as zero-byte files, not absent. - **Total boxes.** Loading is not loading everything. If the rebuilt dataset scans to a number other than the cache's 968,227, the labels are wrong in a way that trains fine and scores wrong. No model, no GPU, no training step. Co-Authored-By: Claude Opus 5 --- .../check_yolo_dataset_loads.py | 91 +++++++++++++++++++ 1 file changed, 91 insertions(+) create mode 100644 scripts/model_comparison/check_yolo_dataset_loads.py diff --git a/scripts/model_comparison/check_yolo_dataset_loads.py b/scripts/model_comparison/check_yolo_dataset_loads.py new file mode 100644 index 00000000..1e5c8465 --- /dev/null +++ b/scripts/model_comparison/check_yolo_dataset_loads.py @@ -0,0 +1,91 @@ +#!/usr/bin/env python3 +"""Acceptance test for the YOLO label rebuild: does Ultralytics load the dataset again? (#51) + +The rebuild in `rebuild_yolo_labels_from_cache.py` verifies its own output against the +cache it read, which proves the round trip but *not* the thing that actually broke. The +failure was in Ultralytics' dataset init: + + ValueError: train: No labels found in .../yolo/tiles/labels/train.cache + +so the only test that closes the loop is making Ultralytics build the dataset and report +a non-zero `nf` (labels found). This does that and nothing else -- it builds the dataset, +reads the scan counters, and exits. No model, no GPU, no training step. + +Two things it checks that a bare "did it crash" run would not: + +- **`nf` (found) against `nm` (missing).** Ultralytics treats a *missing* label file as a + background image and carries on, so a rebuild that skipped the label-less tiles would + still "work" while silently changing the dataset: 59,923 of the 557,413 train tiles are + background, and they must be present as zero-byte files, not absent. A large `nm` is the + signature of that mistake. +- **Total boxes against the expected count.** Loading is not the same as loading + everything. The cache the labels came from holds 968,227 train boxes; if the rebuilt + dataset scans to a different number, the labels are wrong in a way that trains fine and + scores wrong. + +Usage: + + python check_yolo_dataset_loads.py --data-root /gscratch/scrubbed/jfroehli/yolo/tiles \\ + --split train --expect-images 557413 --expect-boxes 968227 + +Exit status is 0 only when every check passes, so it can be run unattended. +""" +from __future__ import annotations + +import argparse +import sys +from pathlib import Path + + +def main() -> int: + ap = argparse.ArgumentParser(description=__doc__.split("\n")[0]) + ap.add_argument("--data-root", type=Path, required=True, + help="dataset root holding images/ and labels/") + ap.add_argument("--split", default="train", help="split to load (default: %(default)s)") + ap.add_argument("--expect-images", type=int, default=None, + help="fail unless this many image/label pairs are found") + ap.add_argument("--expect-boxes", type=int, default=None, + help="fail unless the scanned labels hold this many boxes") + ap.add_argument("--imgsz", type=int, default=640, help="only affects the scan, not results") + args = ap.parse_args() + + from ultralytics.data.dataset import YOLODataset + + img_dir = args.data_root / "images" / args.split + if not img_dir.is_dir(): + print(f"FATAL: no image directory at {img_dir}") + return 1 + + print(f"data-root : {args.data_root}") + print(f"split : {args.split}") + print("building the dataset (this is the call that raised the original ValueError) ...", + flush=True) + + ds = YOLODataset(img_path=str(img_dir), imgsz=args.imgsz, augment=False, + data={"names": {0: "curb_ramp"}, "channels": 3}) + + n_images = len(ds.labels) + n_boxes = sum(len(rec["bboxes"]) for rec in ds.labels) + n_empty = sum(1 for rec in ds.labels if len(rec["bboxes"]) == 0) + + print(f"images : {n_images}") + print(f"boxes : {n_boxes}") + print(f"background: {n_empty} (zero-box tiles, which must be PRESENT as empty files)") + + ok = True + if n_images == 0: + print("FAIL: zero image/label pairs -- this is the original failure, unfixed") + ok = False + if args.expect_images is not None and n_images != args.expect_images: + print(f"FAIL: expected {args.expect_images} images, scanned {n_images}") + ok = False + if args.expect_boxes is not None and n_boxes != args.expect_boxes: + print(f"FAIL: expected {args.expect_boxes} boxes, scanned {n_boxes}") + ok = False + + print("PASS -- Ultralytics builds the dataset and the counts match" if ok else "FAIL") + return 0 if ok else 1 + + +if __name__ == "__main__": + sys.exit(main()) From 6e5e0b93176c3880f8c1095662a1fbee88c67d0c Mon Sep 17 00:00:00 2001 From: Jon Froehlich Date: Sun, 30 Aug 2026 17:10:42 -0700 Subject: [PATCH 4/9] Put the pano caches and the current tiles checkpoints on durable storage (#51) Two things on /gscratch/scrubbed were one purge away from unreproducible, and they fail for the same reason, so this protects both in one job (39361251, rc=0, 2.5 min). **The pano label caches had no durable copy at all.** The 2026-08 purge took every label .txt under yolo/pano/labels/ -- 0 files remain against 150,063 + 42,875 surviving images -- exactly as it did for tiles. pano/labels/{train,val}.cache survived and can reconstitute them, but they were sitting on the same volume, on the same access-time clock, that had already deleted their own labels once. The tiles pair was staged to /gscratch/makelab on 2026-08-17; the pano pair never was. They are now at label_cache_rescue/pano/ (39,475,038 + 11,227,447 bytes) with sha256sums.txt in the format run_rebuild_yolo_labels.slurm already checks. **The checkpoint snapshot was 14 days stale, on exactly the arms under discussion.** Every tiles arm had trained past its durable copy: y11x_tiles ep44 vs ep32 saved y26_tiles ep12 vs ep9 y11x_pano ep38 vs ep30 saved y11l_tiles ep11 vs ep9 so the checkpoints an evaluation would actually score existed at their current epoch in one place, and it was the volume with the demonstrated purge behaviour. Refreshed: 12 copied, 16 unchanged, 0 failed, 50 verified hashes. The rescue is a script rather than the hand-assembly that staged the tiles pair, because a hand-assembled backup cannot be re-run by someone else -- the test this repo applies to everything else. It reuses snapshot_runs.sh's discipline: .tmp-then-rename after the hash matches, and the source hashed before AND after so a cache being rewritten by a live run cannot be captured torn. One deliberate refusal: rescue_label_caches.sh will NOT overwrite a differing durable copy without FORCE=1. The live tiles cache is now a LATER regeneration (Ultralytics rebuilt it during the 08-29 acceptance test) than the durable 08-17 copy, so a backup script that blindly mirrored would have replaced the original the published arms trained on with a look-alike, while reporting success. Co-Authored-By: Claude Opus 5 --- .../yolo_baseline/rescue_label_caches.sh | 165 ++++++++++++++++++ .../yolo_baseline/run_durable_snapshot.slurm | 109 ++++++++++++ 2 files changed, 274 insertions(+) create mode 100644 scripts/model_comparison/yolo_baseline/rescue_label_caches.sh create mode 100644 scripts/model_comparison/yolo_baseline/run_durable_snapshot.slurm diff --git a/scripts/model_comparison/yolo_baseline/rescue_label_caches.sh b/scripts/model_comparison/yolo_baseline/rescue_label_caches.sh new file mode 100644 index 00000000..c153d633 --- /dev/null +++ b/scripts/model_comparison/yolo_baseline/rescue_label_caches.sh @@ -0,0 +1,165 @@ +#!/usr/bin/env bash +# +# rescue_label_caches.sh - copy Ultralytics label caches to durable storage, verified. +# Issue #51. +# +# WHY THIS EXISTS +# /gscratch/scrubbed purges by ACCESS time, and an Ultralytics label cache is precisely +# what stops anything from reading the individual label .txt files. Their atime froze +# while the images kept being read every epoch, so the 2026-08 purge took all 718,415 +# tile labels and left every image. The cache that made training fast is what got the +# labels deleted. +# +# The recovery works because the cache is not a digest: it stores parsed `cls` (n,1) +# and `bboxes` (n,4) as normalized xywh, which is the on-disk .txt format itself, so +# rebuild_yolo_labels_from_cache.py can reconstitute the labels exactly. That makes +# the cache the single most valuable small file in the dataset -- and it was sitting on +# the same volume, on the same purge clock, as the thing it is the backup for. +# +# The tiles pair was staged by hand on 2026-08-17. A hand-assembled backup cannot be +# re-run by someone else, which is the test this repo applies to everything else +# (CLAUDE.md, "replicable from a clean clone"). This is that staging, as a script. +# +# WHAT IT GUARANTEES +# - Never destroys a good copy: each file lands as .tmp and is renamed only after +# its hash matches the source. +# - Idempotent: a destination that already matches is reported unchanged and left alone. +# - Refuses to overwrite a DIFFERING existing copy unless FORCE=1. This is deliberate. +# A cache regenerated by a later Ultralytics run is not the cache the published arms +# trained on, and silently replacing the durable copy with it would lose the original +# while looking like a successful backup. +# +# LAYOUT +# Caches land at $DST//{train,val}.cache with sha256sums.txt beside them. +# The tiles pair already lives FLAT at $DST/{train,val}.cache, which is what +# run_rebuild_yolo_labels.slurm defaults YOLO_CACHE_DIR to; that legacy path is left +# exactly as it is. For any other variant, point YOLO_CACHE_DIR at its subdirectory: +# +# YOLO_CACHE_DIR=$DST/pano YOLO_DATA_ROOT=$SRC/pano \ +# sbatch scripts/model_comparison/run_rebuild_yolo_labels.slurm +# +# USAGE +# ./rescue_label_caches.sh # every variant under SRC that has caches +# ./rescue_label_caches.sh pano # only the named variants +# SRC=... DST=... ./rescue_label_caches.sh # override the committed defaults +# +set -uo pipefail + +SRC="${SRC:-/gscratch/scrubbed/jfroehli/yolo}" +DST="${DST:-/gscratch/makelab/jonf/rampnet_yolo_baseline_51/label_cache_rescue}" +FORCE="${FORCE:-0}" +MAX_TRIES="${MAX_TRIES:-3}" + +SPLITS=(train val) + +n_copied=0; n_same=0; n_failed=0; n_refused=0 + +log() { printf '%s\n' "$*"; } + +human() { numfmt --to=iec --suffix=B "$1" 2>/dev/null || printf '%s' "$1"; } + +# copy_verified