Skip to content

perf(metal): remove redundant host IDWT output transfers - #106

Closed
jcwal1516 wants to merge 1 commit into
perf/audit-00-metal-host-copy-baselinesfrom
perf/audit-10-metal-host-copy-removal
Closed

jcwal1516 wants to merge 1 commit into
perf/audit-00-metal-host-copy-baselinesfrom
perf/audit-10-metal-host-copy-removal

Conversation

@jcwal1516

Copy link
Copy Markdown
Member

The synchronous 5/3 and 9/7 host-slice IDWT callbacks now allocate only the required device output size, without uploading the caller's overwritten output. After the existing command wait, a checked private helper copies completed buffer bytes directly into the required host prefix, avoiding a temporary readback vector and preserving oversized destination tails.

Depends on baseline draft #105 (perf/audit-00-metal-host-copy-baselines, d868e4b). Candidate 33826b7 is independent of the compact-grid/fused-kernel and JPEG scratch optimizations. Element/byte size overflow, bounds, typed alignment, CPU visibility, empty ranges and zero-sized ABI checks are preserved through existing audited access boundaries. No new raw pointer access, public API, shader, route, dependency, or asynchronous semantics change. Owned-result callers retain their vector helpers.

Validation: the transfer regression first failed on the baseline (268 uploaded bytes, one 268-byte vector versus required zero), then passed. Four focused callback tests and two copy-helper tests passed, covering both wavelets/normalizations, odd and degenerate origins, destination capacities/sentinel tails, rejected huge/empty geometry, offsets, alignment, bounds, private storage and empty/ZST cases. cargo run --locked -p xtask -- release-metal --mode quick passed at 33826b7, including production/non-library Clippy, seven runtime/integration groups, inventory and all 21 required ignored hardware tests. Formatting and diff checks passed. Optional external Aperio input remains unavailable.

Identical committed harness, Apple M4 Pro, pinned Rust 1.96, release-bench, serial measurements: 50 calibrated samples after 3-second warmup for the copy/full-host diagnostics; four Criterion resident/readback controls with 50 samples, 3-second warmup and 10-second requested measurement.

Route Median time change 95% median-change interval
CPU-visible copy, 509×383 −50.17% [−50.37%, −50.05%]
CPU-visible copy, 512×512 −50.31% [−50.51%, −50.24%]
Full host-slice 5/3, 509×383 +0.02% [−0.72%, +0.74%]
Full host-slice 5/3, 512×512 +0.17% [−0.20%, +0.41%]
Full host-slice 9/7, 509×383 −1.00% [−1.46%, −0.48%]
Full host-slice 9/7, 512×512 +1.16% [+0.77%, +1.55%]

All routes pass the predeclared 2% regression gate. The four resident/readback controls range from 0.32% faster to 0.13% slower, with mean-change upper intervals at most +0.86%. This is an allocation/copy reduction, not a general full-decode speedup: the isolated CPU-visible copy is approximately 2.01× faster but excludes GPU execution/wait. Full host decode removes three temporary readback vectors and 1,024,780 or 1,376,256 bytes each of overwritten-output upload and vector readback per decode. Coefficient uploads, the final required host copy, and waits remain.

Raw samples, intervals and comparison CSVs are retained outside the checkout. Remaining merge gate: successful manual exact-head GitHub Metal validation workflow. CUDA validation is deferred until the end at the user's request; no CUDA result is claimed. External routing/adoption corpora were not run.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant