[HLSL] Add LinAlg MatVec interpretation and bias coverage#8672
Draft
JoeCitizen wants to merge 17 commits into
Draft
[HLSL] Add LinAlg MatVec interpretation and bias coverage#8672JoeCitizen wants to merge 17 commits into
JoeCitizen wants to merge 17 commits into
Conversation
Use the shared MatrixUse parameter for the OuterProduct result and set it to Accumulator, matching proposal 0035 and the public dx::linalg API. Add a host-side invariant to prevent the legacy A-use declaration from returning. Assisted-by: GitHub Copilot Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Create SRV buffers without UAV flags and transition them for both pixel and non-pixel shader access. Use a direct resource-initialization list so the graphics-only pixel state is legal. Assisted-by: GitHub Copilot Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add typed F16, F32, I32, and U32 matrix data with safe byte encoding, rectangular row/column-major storage mapping, and explicit exact, permitted-result, or excluded comparison policy. Cover offsets and padded strides with independent host goldens, and migrate the existing CopyConvert tests onto the oracle. Assisted-by: GitHub Copilot Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
This was referenced Jul 24, 2026
Handle packed row or column byte-count overflow before using the result, and include raw F32 bits in exact mismatch diagnostics. Cover adjacent float bit patterns in the host oracle test. Assisted-by: GitHub Copilot Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add ABI-checked wrappers for the six D3D12 Linear Algebra capability query categories and explicit applicability classification. Gate the rectangular F32 CopyConvert case using concrete supported wave sizes. Assisted-by: GitHub Copilot Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Compile capability-gated CopyConvert coverage at the exact wave size whose MatrixConstruction support was queried. Keep mandatory baseline cases on the existing ranged WaveSize attribute. Assisted-by: GitHub Copilot Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
JoeCitizen
force-pushed
the
linalg-hlk-matvec-semantics
branch
from
July 25, 2026 01:03
e1b8cfc to
30575b4
Compare
added 2 commits
July 25, 2026 14:10
Validate multiplication support flags per operation, exhaustively check the preview D3D12 ABI mirrors, and preserve query-backed optional skips in HLK mode. Assisted-by: GitHub Copilot Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add rectangular Length/GetCoordinate/GetElement coverage and the specified Get/Set out-of-bounds behaviour. Capture thread-local matrix records without UAV races and gate optional F32 cases at the exact queried wave size. Assisted-by: GitHub Copilot Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
JoeCitizen
force-pushed
the
linalg-hlk-matvec-semantics
branch
from
July 25, 2026 02:27
30575b4 to
a2deb88
Compare
Seed OOB Get outputs with non-zero sentinels and require every lane in the selected wave to execute and write the specified zero result. Assisted-by: GitHub Copilot Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add bounded raw descriptor-table bindings and independent whole-buffer oracles for LinAlg descriptor operations. Cover non-zero offsets, padded strides, row/column-major transfer, descriptor bounds, and capability-gated atomic accumulation. Assisted-by: GitHub Copilot Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
JoeCitizen
force-pushed
the
linalg-hlk-matvec-semantics
branch
from
July 25, 2026 02:55
a2deb88 to
40f613f
Compare
added 2 commits
July 25, 2026 15:21
Reject invalid raw-buffer views, conflicting shader-visible resource heaps, and ambiguous root-parameter bindings before ShaderOp execution. Assisted-by: GitHub Copilot Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add race-free Wave and ThreadGroup group-shared transfer coverage for row/column-major layouts, non-zero offsets, padded strides, and exact whole-buffer guards. Add capability-gated Wave atomic accumulation with coordinate-derived values, while keeping cross-component conversion out of scope pending runtime conformance. Assisted-by: GitHub Copilot Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
JoeCitizen
force-pushed
the
linalg-hlk-matvec-semantics
branch
from
July 25, 2026 03:23
40f613f to
3fbc17a
Compare
Extend each group-shared backing array by four typed sentinel elements so transfer and accumulation tests verify writes do not overrun the matrix extent. Assisted-by: GitHub Copilot Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add mixed F16/F32 CopyConvert cases and verify that conversion leaves the source matrix unchanged. Cover exact integer widening, RTNE plus saturating float narrowing, and capability-gated FP8 encoding and round-trip semantics with independent host oracles. Assisted-by: GitHub Copilot Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
JoeCitizen
force-pushed
the
linalg-hlk-matvec-semantics
branch
from
July 25, 2026 03:51
3fbc17a to
59708aa
Compare
added 2 commits
July 25, 2026 18:37
Feed host-derived packed FP8 bytes through an SRV for decode so the F16 result cannot false-pass through a folded shader encode/decode chain. Assisted-by: GitHub Copilot Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Refactor MatVec execution tests around independent matrix, vector, bias, and output resources with host-derived exact expectations. Add required interpreted input tuples, non-uniform layout coverage, unsigned output, and independent bias validation behind the runtime ThreadVectorMatrixMultiply capability query. The mandatory native F32-to-SInt8 case remains active and exposes the current preview WARP conversion defect. Assisted-by: GitHub Copilot Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
JoeCitizen
force-pushed
the
linalg-hlk-matvec-semantics
branch
from
July 25, 2026 06:40
59708aa to
f989f89
Compare
Separate native F32 inputs from hand-derived SInt8 values so MatVec exercises RTNE saturation, and use high-bit UInt8 lanes to distinguish unsigned packed interpretation. Assisted-by: GitHub Copilot Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The four existing MatVec baselines are migrated to the shared harness without changing their intended inputs or results.
Validation
ExecHLSLTeststarget1.721.2-preview: 10 passed and the ColumnMajor case skipped because WARP does not advertiseTRANSPOSE<8 x float>input with SInt8 interpretation, unsigned output withisOutputSigned=false, and a distinct bias SRVHlslExecTestUtils.cppandLinAlgTests.cppagainst the preview D3D12 headersgit diff --checkpassNo physical GPU or packaged-HLK execution is claimed.
Preview WARP defect
Public preview WARP advertises the mandatory F32-vector x SInt8-matrix to SInt32 tuple as supported, but the strengthened test receives
[-372, 372, -130, 262]instead of the independently derived[3336, -3336, -2, -254]. The physical inputs include half-way values and values beyond the SInt8 range, so the case now directly exercises RTNE ties and positive/negative saturation.The emitted DXIL correctly carries a native
<8 x float>vector and SInt8 conversion-target interpretation. The current WARP implementation treats the float words as packed I8 bytes instead of applying the required RTNE saturating F32-to-SInt8 conversion. The mandatory test remains active; skipping or weakening it would hide a Tier-1 conformance defect.Stack
This draft is stacked on PR #8671, which is stacked on PR #8670, PR #8669, PR #8668, PR #8667, PR #8666, PR #8665, and PR #8662. Until those ancestors land, this diff contains their commits as well. The implementation commit is
f989f89a9; review correction247f38d10strengthens the F32-to-SInt8 and UInt8 interpretation oracles.This remains a draft for named human review. The reviewer should verify the capability-query tuple, native-vector interpretation immediate, RTNE/saturation hand derivation, high-bit UInt8 packing, independent bias resource, CPU oracle, and preview-WARP defect analysis before requesting maintainer review.
Refs #7841
Refs #8559
Refs #8560
Refs #8650
Refs #8653
Assisted-by: GitHub Copilot