Skip to content

[HLSL] Add LinAlg MatVec interpretation and bias coverage#8672

Draft
JoeCitizen wants to merge 17 commits into
microsoft:mainfrom
JoeCitizen:linalg-hlk-matvec-semantics
Draft

[HLSL] Add LinAlg MatVec interpretation and bias coverage#8672
JoeCitizen wants to merge 17 commits into
microsoft:mainfrom
JoeCitizen:linalg-hlk-matvec-semantics

Conversation

@JoeCitizen

@JoeCitizen JoeCitizen commented Jul 24, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • refactor the MatVec execution harness to use independent matrix, vector, bias, and output resources with exact typed host encoding and CPU-derived dot-product expectations
  • add exact ThreadVectorMatrixMultiply capability queries and cover non-uniform F16, capability-gated ColumnMajor, required packed SInt8 and UInt8 tuples, the required native F32-to-SInt8 interpretation path, optional unsigned U32 output, and independent bias data
  • correct packed 8-bit element sizing and validate LSB-first vector packing with zero padding to the 32-bit storage boundary
  • separate native F32 source values from their hand-derived SInt8 interpretations, covering RTNE ties and positive/negative saturation, and use high-bit UInt8 lanes that distinguish unsigned from signed packed interpretation

The four existing MatVec baselines are migrated to the shared harness without changing their intended inputs or results.

Validation

  • built the Release ExecHLSLTests target
  • ran all 11 MatVec methods on the compatible local WARP and Agility SDK 1.721.2-preview: 10 passed and the ColumnMajor case skipped because WARP does not advertise TRANSPOSE
  • after propagating the correction through the complete stack, all 66 LinAlg methods reported 61 passed, 0 failed, and 5 authoritative capability-backed skips
  • compiled the capability-gated ColumnMajor shader independently
  • inspected emitted DXIL for packed SInt8/UInt8 inputs, native <8 x float> input with SInt8 interpretation, unsigned output with isOutputSigned=false, and a distinct bias SRV
  • compiled HlslExecTestUtils.cpp and LinAlgTests.cpp against the preview D3D12 headers
  • clang-format 17.0.1 and git diff --check pass

No physical GPU or packaged-HLK execution is claimed.

Preview WARP defect

Public preview WARP advertises the mandatory F32-vector x SInt8-matrix to SInt32 tuple as supported, but the strengthened test receives [-372, 372, -130, 262] instead of the independently derived [3336, -3336, -2, -254]. The physical inputs include half-way values and values beyond the SInt8 range, so the case now directly exercises RTNE ties and positive/negative saturation.

The emitted DXIL correctly carries a native <8 x float> vector and SInt8 conversion-target interpretation. The current WARP implementation treats the float words as packed I8 bytes instead of applying the required RTNE saturating F32-to-SInt8 conversion. The mandatory test remains active; skipping or weakening it would hide a Tier-1 conformance defect.

Stack

This draft is stacked on PR #8671, which is stacked on PR #8670, PR #8669, PR #8668, PR #8667, PR #8666, PR #8665, and PR #8662. Until those ancestors land, this diff contains their commits as well. The implementation commit is f989f89a9; review correction 247f38d10 strengthens the F32-to-SInt8 and UInt8 interpretation oracles.

This remains a draft for named human review. The reviewer should verify the capability-query tuple, native-vector interpretation immediate, RTNE/saturation hand derivation, high-bit UInt8 packing, independent bias resource, CPU oracle, and preview-WARP defect analysis before requesting maintainer review.

Refs #7841
Refs #8559
Refs #8560
Refs #8650
Refs #8653

Assisted-by: GitHub Copilot

Jack Elliott and others added 3 commits July 23, 2026 14:52
Use the shared MatrixUse parameter for the OuterProduct result and set it to Accumulator, matching proposal 0035 and the public dx::linalg API. Add a host-side invariant to prevent the legacy A-use declaration from returning.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Create SRV buffers without UAV flags and transition them for both pixel and non-pixel shader access. Use a direct resource-initialization list so the graphics-only pixel state is legal.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add typed F16, F32, I32, and U32 matrix data with safe byte encoding, rectangular row/column-major storage mapping, and explicit exact, permitted-result, or excluded comparison policy. Cover offsets and padded strides with independent host goldens, and migrate the existing CopyConvert tests onto the oracle.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Jack Elliott and others added 3 commits July 25, 2026 12:59
Handle packed row or column byte-count overflow before using the result, and include raw F32 bits in exact mismatch diagnostics. Cover adjacent float bit patterns in the host oracle test.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add ABI-checked wrappers for the six D3D12 Linear Algebra capability
query categories and explicit applicability classification. Gate the
rectangular F32 CopyConvert case using concrete supported wave sizes.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Compile capability-gated CopyConvert coverage at the exact wave size whose MatrixConstruction support was queried. Keep mandatory baseline cases on the existing ranged WaveSize attribute.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
@JoeCitizen
JoeCitizen force-pushed the linalg-hlk-matvec-semantics branch from e1b8cfc to 30575b4 Compare July 25, 2026 01:03
Jack Elliott added 2 commits July 25, 2026 14:10
Validate multiplication support flags per operation, exhaustively check the preview D3D12 ABI mirrors, and preserve query-backed optional skips in HLK mode.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add rectangular Length/GetCoordinate/GetElement coverage and the specified Get/Set out-of-bounds behaviour. Capture thread-local matrix records without UAV races and gate optional F32 cases at the exact queried wave size.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
@JoeCitizen
JoeCitizen force-pushed the linalg-hlk-matvec-semantics branch from 30575b4 to a2deb88 Compare July 25, 2026 02:27
Jack Elliott and others added 2 commits July 25, 2026 14:51
Seed OOB Get outputs with non-zero sentinels and require every lane in the selected wave to execute and write the specified zero result.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add bounded raw descriptor-table bindings and independent whole-buffer
oracles for LinAlg descriptor operations. Cover non-zero offsets, padded
strides, row/column-major transfer, descriptor bounds, and capability-gated
atomic accumulation.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
@JoeCitizen
JoeCitizen force-pushed the linalg-hlk-matvec-semantics branch from a2deb88 to 40f613f Compare July 25, 2026 02:55
Jack Elliott added 2 commits July 25, 2026 15:21
Reject invalid raw-buffer views, conflicting shader-visible resource heaps, and ambiguous root-parameter bindings before ShaderOp execution.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add race-free Wave and ThreadGroup group-shared transfer coverage for row/column-major layouts, non-zero offsets, padded strides, and exact whole-buffer guards. Add capability-gated Wave atomic accumulation with coordinate-derived values, while keeping cross-component conversion out of scope pending runtime conformance.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
@JoeCitizen
JoeCitizen force-pushed the linalg-hlk-matvec-semantics branch from 40f613f to 3fbc17a Compare July 25, 2026 03:23
Jack Elliott and others added 2 commits July 25, 2026 15:48
Extend each group-shared backing array by four typed sentinel elements so transfer and accumulation tests verify writes do not overrun the matrix extent.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Add mixed F16/F32 CopyConvert cases and verify that conversion leaves the source matrix unchanged.

Cover exact integer widening, RTNE plus saturating float narrowing, and capability-gated FP8 encoding and round-trip semantics with independent host oracles.

Assisted-by: GitHub Copilot
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
@JoeCitizen
JoeCitizen force-pushed the linalg-hlk-matvec-semantics branch from 3fbc17a to 59708aa Compare July 25, 2026 03:51
Jack Elliott added 2 commits July 25, 2026 18:37
Feed host-derived packed FP8 bytes through an SRV for decode so the F16 result cannot false-pass through a folded shader encode/decode chain.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Refactor MatVec execution tests around independent matrix, vector, bias, and output resources with host-derived exact expectations.

Add required interpreted input tuples, non-uniform layout coverage, unsigned output, and independent bias validation behind the runtime ThreadVectorMatrixMultiply capability query. The mandatory native F32-to-SInt8 case remains active and exposes the current preview WARP conversion defect.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
@JoeCitizen
JoeCitizen force-pushed the linalg-hlk-matvec-semantics branch from 59708aa to f989f89 Compare July 25, 2026 06:40
Separate native F32 inputs from hand-derived SInt8 values so MatVec exercises RTNE saturation, and use high-bit UInt8 lanes to distinguish unsigned packed interpretation.

Assisted-by: GitHub Copilot

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 83725f5d-8e98-4c1d-91ee-ad47629e007b
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: New

Development

Successfully merging this pull request may close these issues.

1 participant