Repository navigation
Conversation
This was referenced Sep 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The existing JPEG Metal full-decode benchmark uses the queued submission API. This adds a baseline for the synchronous
Decoder::decode_to_device_with_sessionpath that task09 will optimize.The new matrix retains the decoder and backend session across calls, validates Metal residency and exact native pixels before timing, and returns a fresh surface on each iteration. It covers 420 with/without restart markers, 422, and 444 in Gray/RGB/RGBA; a whole-image scaled control; a 444 region-scaled control; and a real four-distinct-image batch using retained decoders and a reusable caller-owned output buffer. Partial-image controls are explicitly labeled as global-runtime calls because their public API does not accept a backend session.
A separate manual diagnostic counts actual shared/private Metal buffer allocations before download. The runtime is initialized outside those counts, and pixels are checked afterward. A pure regression test records the existing fail-closed pool behavior after an owner panic. No production ownership, codec behavior, API, dependencies, or routing change is included.
Known baseline gap: the initially proposed generated 420 restart region-scaled control differs from native at 545 of 6912 output bytes, maximum difference 39. Full-image and whole-image scaled probes pass. The evidence retains that failure; this PR does not fix or validate that partial 420 route, and uses 444 as the exact-parity region-scaled control. No tolerance was widened.
Validation: all 15 smoke probes passed exact native pixel comparison; scoped tests/bench Clippy, formatting, diff review, the allocation diagnostic, and the panic regression passed. The committed baseline at 070ef29 completed all 15 rows with 50 samples, 3-second warmup, and 10-second requested measurement in release-bench on Apple M4 Pro. The allocation diagnostic confirms 2 private + 5 shared buffers per Gray call and 3 private + 5 shared per RGB/RGBA call across the four sampling fixtures. Counts exclude runtime setup and download but include the returned shared output. No acceleration claim is made by this prerequisite. CUDA validation remains deferred until the end at the user's request. Exact-head manual GitHub Metal validation remains a premerge gate.