SOFIE_quantization Full Library v0 - #53
Open
shaunlee8 wants to merge 31 commits into
Open
Conversation
…sion, FP8 chaining, vectorized elementwise
sanjibansg
requested changes
Sep 6, 2026
sanjibansg
left a comment
Member
There was a problem hiding this comment.
Ran a first review pass, couple of comments.
| #ifndef SOFIE_RQUANTIZATION_ALPAKA_DENSE_LINEAR_COMMON | ||
| #define SOFIE_RQUANTIZATION_ALPAKA_DENSE_LINEAR_COMMON | ||
|
|
||
| // cuBLASLt scaffolding both dense-linear precisions sit on: owned handles, layouts, heuristic |
Member
There was a problem hiding this comment.
can we instead interface this to sofieBLAS with guards to prevent anything else other than cuda? this way later when we support other architectures we just remove the guard and keeps the quantized SOFIE portable as well. This would need corresponding changes in sofieBLAS to support the CUDA primitives we have here.
Member
There was a problem hiding this comment.
I would keep the quantized models in a dedicated sub-dir.
Member
There was a problem hiding this comment.
Does this include any I/O operations different from what we already have?
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
An end-to-end quantized inference pipeline for SOFIE code generation. Support for ONNX quantized graphs, to self-contained generated headers, to low-precision execution. Structured as an installable codegen pass (
SOFIE_quantization) over the existing core.Frontend:
QuantizeLinear/DequantizeLinearboundary pairsQuantcustom opsData Types:
Operators:
Execution:
Pipeline:
RQuantization_*Pass.cxx),RQuantization_*.cxxfind regions where everything between a dequantize and a quantize can run directly on the codes (raw stored numbers, int or fp bit patterns). Regions grow through value-preserving ops, absorb activations, and detect boundaries between neighboring regions so chained layers requantize in place. Node elimination rules target leftover quantized ops.RQuantization_Lowering.hxx) where operand quantization metadata lives inQuantizedOperandContract(INT8, FP8, etc.) andQuantizationGridwhich maps codes to values.