Methods, component sensitivity, calibration, and deployment
Manuscript · Method catalog · Full bibliography · Citation
Baha Rababah, Yuzhang Shang, Carson K. Leung, Cuneyt G. Akcora, and Mubarak Shah
University of Manitoba and University of Central Florida.
This repository accompanies the survey and organizes its literature around one question: How does each method control quantization error? It connects the methods to the transformer components they affect, the calibration information they use, and the numerical formats and kernels required for deployment. The catalog contains 52 method entries from Section 3 and an index of all 108 references in the manuscript.
The four method families are compensation, rotation, salience, and optimization. The survey adopts this taxonomy from Zhao et al. 104 and assigns hybrid methods by their primary error-control mechanism. This repository preserves those assignments and records important secondary mechanisms in each entry.
Figure 1 from the manuscript, page 4. Weights, runtime activations, and cached keys and values are separate quantization targets; their memory costs and error paths differ.
| Edition | Contents |
|---|---|
| 6 September 2026 | Companion prepared from the supplied 35-page manuscript and main LaTeX source. Includes 52 method entries, 108 references, eight extracted figures, structured metadata, and source notes. |
| Section | Coverage |
|---|---|
| Background | Error definitions, quantization scope, number systems, granularity, and inference phases. |
| Taxonomy | Four error-control families and rules for reading hybrid methods. |
| Compensation | Hessian-aware reconstruction, low-rank reconstruction, and vector codebooks. |
| Rotation | Incoherence, fixed and learned rotations, structured transforms, and specialized settings. |
| Salience | Activation-aware scaling, selective preservation, mixed precision, and binary regimes. |
| Optimization | Rounding, clipping, block reconstruction, smoothing, and data-free balancing. |
| Component sensitivity | Attention, MLP projections, KV cache, normalization, residual paths, and output logits. |
| Calibration | Range fitting, distribution matching, reconstruction scope, activation statistics, and data selection. |
| Research directions | Formats and kernels, extreme precision, emerging architectures, inference workloads, and reliability. |
| Evaluation | Evaluation resources and the information needed to interpret reported results. |
| Libraries and implementations | Author repositories and systems resources cited in the survey. |
| Related surveys | Background and neighboring surveys from the bibliography. |
| Repository files | Documents, figures, structured indexes, and validation tools. |
| Contributing | Entry requirements and evidence rules. |
| Citation | Manuscript citation without unconfirmed publication metadata. |
| License and source status | Author decisions still required for publication and reuse. |
Post-training quantization converts a pretrained model to lower-precision representations after training. The survey treats the resulting approximation error as the common object across methods. For a linear layer with row-stacked inputs,
The shorthand W4A8 denotes 4-bit weights and 8-bit activations. It does not specify the KV-cache precision, quantization groups, codebook storage, retained high-precision components, or compute format of every operation. Those details are needed to compare methods with the same nominal precision.
| Design dimension | Choices discussed in the survey | Why the choice matters |
|---|---|---|
| Quantized tensors | Weights; weights and activations; cached keys and values. | Static weights can be processed offline. Activations and cached states depend on the input and generation process. |
| Quantization levels | Uniformly spaced values; non-uniform reconstruction values and codebooks. | The representation determines how limited levels fit the tensor distribution and how values are decoded. |
| Granularity | Per-tensor, per-channel, and per-group parameters. | Local scales can fit heterogeneous values more closely, with additional metadata and scaling work. |
| Precision assignment | Uniform precision; structured mixed precision; selected higher-precision entries or channels. | Selective precision changes the storage budget and may require separate execution paths. |
| Inference phase | Prefill and autoregressive decoding. | Prefill processes the prompt and builds the cache. Decode repeatedly uses cached states while producing new tokens. |
| Storage and execution | Packed values, scales, zero-points, indices, codebooks, correction matrices, and online transforms. | A compact representation must be evaluated together with the operations needed to execute it. |
Uniform quantization shares a fixed step between adjacent reconstruction values. Non-uniform quantization uses unevenly spaced reconstruction values; vector quantization extends the representation to groups of weights. The survey’s distinction between structured and unstructured access is also important: channel- or group-level selections can use regular layouts, while arbitrary preserved entries require location information and a suitable execution path. See Section 2.
| Family | Primary question | Mechanisms | Entries |
|---|---|---|---|
| Compensation | How can the quantization error be corrected? | Sequential curvature-aware updates, compact residual terms, and vector-codebook reconstruction. | 11 |
| Rotation | How can transformed coordinates make quantization easier? | Incoherence processing, orthogonal rotations, structured transforms, and related distribution reshaping. | 22 |
| Salience | Which parts need protection from precision loss? | Activation-aware scaling, preserved columns or entries, structured bit allocation, and binary masks. | 8 |
| Optimization | Which quantizer choices best preserve the selected output? | Learned rounding, clipping, scales, block reconstruction, and equivalent transformations. | 11 |
These families overlap. SEPTQ combines compensation and selective preservation; QuIP# combines rotations and codebooks; ROSAQ combines rotation and salience; several optimization methods use low-rank scaling. The family label records the survey’s organizing choice. The mechanism and deployment columns provide the additional information needed to interpret it. See Section 3.
Figure 2, page 7, is preserved as supplied. The method catalog follows the full Section 3 discussion: the original figure omits KurTail, ButterflyQuant, and SINQ and repeats PeRQ.
Each method name links to a primary paper record. The numbered reference links to its complete citation in this repository. Years and venues are taken from the supplied bibliography and may differ from the first arXiv posting date. “Author code” is shown where an author-linked project was identified; its absence means no such link was included in this preparation. The deployment column summarizes costs and reporting considerations discussed in the survey, not results from an implementation audit.
Compensation methods preserve the behavior of a projection through error correction or reconstruction. The survey separates sequential curvature-aware updates, low-rank reconstruction, and vector-codebook reconstruction. For a low-rank correction, the effective matrix is
Figure 3, page 7: Hessian-aware minimization, low-rank error reconstruction, and vector-codebook reconstruction.
The reconstruction signal comes from the outputs produced on calibration inputs. GPTQ uses the quadratic structure induced by
| Method and source | Error-control mechanism | Setting and deployment considerations |
|---|---|---|
| GPTQ · 20 arXiv, 2022 Author code |
Preserves linear projection outputs on calibration activations. It uses a second-order reconstruction objective, sequential rounding, inverse-Hessian error updates, lazy block updates, and a Cholesky reformulation. | Weight-only; 3-bit and 4-bit examples. Calibration requires activation statistics and curvature operations. Low-bit storage needs a compatible inference kernel. |
| SEPTQ · 48 KDD, 2025 |
Combines a static importance mask with sequential Hessian-guided compensation. Selected important weights retain their original values; the other weights use the low-bit approximation. | Weight-only; selective higher precision. Count preserved values and mask/index storage in the effective bit budget. Selective storage can introduce a separate high-precision path. |
LQER and QERA use an additive correction. LRQ is included in this subsection by the source, but its learned low-rank object is a weight-scaling matrix. The entries retain this distinction.
| Method and source | Error-control mechanism | Setting and deployment considerations |
|---|---|---|
| LQER · 101 ICML, 2024 Author code |
Approximates activation-scaled quantization error with a low-rank factorization. The deployed output combines the low-bit projection with a compact high-precision correction. | Weight and activation; W4A8 example. Include both correction matrices and their two matrix multiplications in memory and latency measurements. |
| QERA · 102 ICLR, 2025 Author code |
Derives an analytical low-rank correction from the calibration-output discrepancy. Its objective accounts for the inputs that multiply the weight error. | Weight-only error reconstruction. The residual correction is an extra deployed term. Rank selection controls its storage and compute cost. |
| ASER · 107 AAAI, 2025 |
Combines activation smoothing with low-rank reconstruction of weight-and-activation quantization error. Smoothing reduces outlier pressure before the correction recovers part of the output behavior. | Weight and activation; W4A8 and W4A6 examples. The correction rank, activation precision, and per-channel configuration must accompany any reported result. |
| LRQ · 40 NAACL, 2025 Author code |
Learns low-rank weight-scaling matrices through block reconstruction. The low-rank parameterization allows more flexible scaling than independent channel scales with fewer parameters than element-wise scaling. | Weight and activation; W8A8 and W4A8. Its low-rank object is a scaling parameterization. It is not the additive residual term used by LQER and QERA. |
These methods replace groups of weights with codebook indices. Their storage includes the codebooks and any residual or outlier representation; their execution includes decoding or lookup work.
| Method and source | Error-control mechanism | Setting and deployment considerations |
|---|---|---|
| AQLM · 18 ICML, 2024 Author code |
Reconstructs each weight vector as the sum of entries from several learned codebooks. Calibration preserves projection and transformer-block outputs, and codebook parameters are tuned jointly. | Weight-only; approximately 2–3 bits per parameter. Account for codebooks, indices, lookup work, and any adaptation stage. Nominal index precision alone does not describe total storage. |
| GPTVQ · 83 arXiv, 2024 Author code |
Splits weights into groups with small vector codebooks and index tables. The survey describes a storage-oriented implementation that decodes indices to a native compute type for matrix multiplication. | Weight-only; mobile CPU setting. Benefits depend on memory traffic and decoding cost on the target processor. Mobile CPU results are not GPU throughput measurements. |
| CRVQ · 93 TACL, 2025 |
Assigns a base codebook to all groups and extra codebooks to important channels. A Hessian-related score ranks channels before they are reordered into critical groups. | Weight-only; sub-2-bit settings. The average rate depends on how many groups receive extension codebooks. Include the channel allocation and all codebook storage. |
| VPTQ · 51 EMNLP, 2024 Author code |
Uses second-order information for vector reconstruction, a residual codebook for the error left by the first codebook, and a separate configuration for difficult vectors. | Weight-only; 2-bit examples. Report residual and outlier configurations with decoding throughput. The paper-specific comparison does not establish a hardware-independent ranking. |
| RSAVQ · 94 NeurIPS, 2025 |
Uses Fisher-information-based sensitivity to shape vector quantization error and allocate more precision to sensitive channels. Both error direction and unequal channel sensitivity enter the design. | Weight-only; 2-bit example. Sensitivity estimation and heterogeneous allocation add calibration and representation costs. Record the complete bit budget. |
The central comparison within this family is the object being reconstructed. A good entry-wise weight fit, an activation-weighted projection fit, and a full-block fit are different objectives. The survey’s reported numerical examples retain their original model and precision settings so they are not read as a common leaderboard.
Rotation methods change the coordinates in which quantization is performed. For an orthogonal matrix
Figure 4, page 12, is reproduced without alteration.
QuIP and QuIP# connect transformed coordinates to low-bit reconstruction. The source presents guarantees for this setting; these should be kept separate from claims about complete-model task degradation.
| Method and source | Error-control mechanism | Setting and deployment considerations |
|---|---|---|
| QuIP · 6 NeurIPS, 2023 Author code |
Applies randomized orthogonal preprocessing to weights and the curvature information used for adaptive rounding. The survey presents it as a low-bit method with theoretical guarantees. | Weight-only. Separate theoretical reconstruction guarantees from downstream task guarantees. Record the entire preprocessing and inference path. |
| QuIP# · 81 ICML, 2024 Author code |
Combines randomized Hadamard incoherence processing, lattice codebooks based on E8, and a brief fine-tuning stage. The structured transform reduces preprocessing cost. | Weight-only; 2-bit examples. Include codebook decoding and the reported fine-tuning stage when comparing calibration budgets or training-free claims. |
The survey distinguishes data-independent transforms from outlier-aware constructions. A training-free transform can still use activation statistics or require online operations.
| Method and source | Error-control mechanism | Setting and deployment considerations |
|---|---|---|
| QuaRot · 5 NeurIPS, 2024 Author code |
Uses randomized Hadamard rotations across the residual stream and selected attention/MLP paths. Two rotations are fused into weights; query/key and down-projection transforms are applied online. | Weights, activations, and KV cache; 4-bit pipeline. The whole pipeline is not free of online transforms. Report prefill and decode separately and include the cache configuration. |
| DuQuant · 45 NeurIPS, 2024 Author code |
Uses outlier-aware block rotations, a zigzag permutation to balance blocks, and a smoothing rotation. The construction uses distribution information without learned dense rotation parameters. | Weights and activations. Training-free construction still has preprocessing and calibration requirements. Permutation and block layout must fit the implementation. |
These methods estimate transforms from calibration signals. Their objectives include output preservation and distributional proxies; the optimization cost and the deployed transform are separate concerns.
| Method and source | Error-control mechanism | Setting and deployment considerations |
|---|---|---|
| SpinQuant · 53 ICLR, 2025 Author code |
Learns fusible residual and attention-path rotations on the orthogonal manifold. It retains the fixed online Hadamard transforms used for the other QuaRot paths. | Weights, activations, and KV cache. Dense rotation parameters and backpropagation increase calibration cost. Learned fusible rotations do not remove all online work. |
| OSTQuant · 30 ICLR, 2025 Author code |
Jointly learns orthogonal rotations and diagonal scaling. Its quantization-space utilization analysis motivates fitting the transformed distributions to the quantizer. | Weights, activations, and KV cache. Record the learned transformations, loss, and full target configuration. Scaling and rotation are distinct transformation families. |
| KurTail · 2 arXiv, 2025 |
Minimizes the kurtosis of rotated activations layer by layer. The survey uses it to illustrate learning a distributional proxy with reduced memory demands. | Weights, activations, and KV cache in Table 3. Its proxy uses activations. It should not be grouped with parameter-only rotation methods merely because the objective is a moment statistic. |
| DartQuant · 73 NeurIPS, 2025 Author code |
Uses distribution-aware calibration constraints with QR-based orthogonalization. The survey presents it as a lower-cost way to obtain quantization-friendly rotations. | Weights and activations. Compare its calibration cost under the same layer scope and sample budget as other learned rotations. |
Structure constrains how a transform is represented and executed. Data-free construction constrains which information is used to obtain it. These are independent design choices.
| Method and source | Error-control mechanism | Setting and deployment considerations |
|---|---|---|
| ButterflyQuant · 91 arXiv, 2025 |
Parameterizes an orthogonal butterfly transform with continuous Givens angles and a uniformity regularizer. The structure reduces the transform parameter count and computation. | Ultra-low-bit activations; 2-bit example. Structured transforms can still run online. Report runtime transform cost separately from the reduction in calibration parameters. |
| ParoQuant · 44 ICLR, 2026 Author code |
Uses independent pairwise Givens rotations together with channel-wise scaling. The survey connects the method to quantization errors in long reasoning generations. | Reasoning models; the cited 2.4% result is weight-only. Measure the cost of pairwise transforms in the deployed setting. The public abstract places the cited reasoning improvement in a weight-only setting and also discusses weight-activation results. |
| OptRot · 23 Machine Learning for Systems, 2025 |
Learns rotations from a weight-only fourth-power objective. The source separately describes OptRot+ as a data-dependent variant with activation-covariance information. | Weight-only; OptRot+ adds activation information. Keep the data-free and activation-informed variants separate in comparisons. A data-free objective can still require optimization. |
This source subsection groups closed-form rotations with methods that add scaling, bias correction, or more general invertible maps. The heading does not imply that every listed method uses a non-orthogonal transform.
| Method and source | Error-control mechanism | Setting and deployment considerations |
|---|---|---|
| SingleQuant · 90 arXiv, 2025 |
Constructs closed-form Givens alignment and uniformity rotations. The source uses it to discuss smoothing large outliers without iterative manifold optimization. | Weights and activations; W4A4. Its transforms remain orthogonal. The broader subsection title does not mean that this particular method uses a general affine map. |
| SmoothRot · 12 IEEE SMC, 2025 Author code |
Combines SmoothQuant-style channel scaling with a fixed Hadamard rotation. Scaling first moderates large activation ranges; rotation redistributes the remaining concentration. | Weights and activations. Distinguish fused scales from online Hadamard operations. A measured small overhead is not evidence that no transform is executed. |
| BASE-Q · 29 arXiv, 2025 Author code |
Adds bias correction and asymmetric scaling to a fixed rotation. The design addresses channel-mean offsets and asymmetric ranges through explicit additional operations. | Weights and activations. Record which correction parameters are absorbed offline and which operations remain in the deployed kernel. |
| FlatQuant · 78 ICML, 2025 Author code |
Learns per-linear-layer invertible transformations with a Kronecker parameterization. Its implementation fuses the online transform with quantization. | Weights and activations; cache flag differs across source tables. The transform is more general than an orthogonal rotation. Kernel fusion reduces overhead but does not imply zero extra arithmetic. |
Cache placement, positional operations, block scales, and selective precision impose additional design constraints. The entries retain their specific target settings.
| Method and source | Error-control mechanism | Setting and deployment considerations |
|---|---|---|
| RotateKV · 76 IJCAI, 2025 Author code |
Uses outlier-aware adaptive Hadamard rotations, channel reordering, grouped-head handling before rotary position encoding, and protection for attention-sink tokens. | KV cache; 2-bit examples. Report key/value layouts, protected tokens, rotation placement, and all retained high-precision cache storage. |
| KVLinC · 71 arXiv, 2025 |
Combines a Hadamard transform on values with small linear corrections for key-quantization error. Rotation and correction address different parts of attention. | KV cache. The correction adapters and online transform must be included in memory and latency accounting. |
| Block Rotation Quantization · 74 ICML, 2026 |
Matches the rotation block to the microscaling block used by the numeric format. The source presents this as a response to conflicts between global rotations and shared block scales. | Weights and activations; MXFP4. Preserve the format, group size, scale format, and kernel assumptions. A generic INT4 result does not establish MXFP4 performance. |
| PeRQ · 70 arXiv, 2026 |
Balances pre-rotation L1 mass across blocks with a permutation before block Hadamard transforms. The source relates the analysis to the empirical rotation-and-permutation design of DuQuant. | Weights and activations; block rotations. Use one catalog entry despite its repeated appearance in the original taxonomy figure. The primary arXiv record confirms the PeRQ name. |
| ROSAQ · 99 arXiv, 2025 |
Uses a closed-form principal-component rotation to align sensitive directions with selected components, then retains those directions in higher precision. | Weight quantization with salient FP16 directions. This is a rotation-and-salience hybrid. Include the protected directions in the effective precision and runtime cost. |
| Method and source | Error-control mechanism | Setting and deployment considerations |
|---|---|---|
| SpinOut · 65 IEEE Access, 2026 |
Injects artificial outliers during rotation training in a selected subset of sensitive layers. Layer scores and performance criteria determine where to apply the intervention. | Weights, activations, and KV cache. Separate layer search, rotation training, and calibration data requirements from inference cost. |
| ReSpinQuant · 35 ICML, 2026 |
Refolds learned layer-wise rotations into weights and approximates the resulting basis mismatch with a subspace residual rotation. The objective is to retain layer-wise accuracy at low runtime cost. | Weights and activations; W4A4 and W3A3. The residual approximation is part of the method. Record its accuracy and overhead alongside the fused transformations. |
FrameQuant is discussed beside this family but is explicitly excluded from rotation-based methods in the manuscript: it uses overcomplete fusion frames. It is available through FrameQuant · 1 and its author implementation. It is not included in the 22 rotation entries.
Salience methods allocate protection according to the effect of a perturbation on projection outputs or downstream behavior. The source considers activation magnitude, weight magnitude, curvature information, output reconstruction, and outlier statistics as importance signals. Protection may take the form of a scale transformation, higher-precision values, or a structured bit allocation. See Section 3.3.
Figure 5, page 16: activation-aware scaling, selective higher-precision preservation, and salience-weighted bit allocation.
| Method and source | Error-control mechanism | Setting and deployment considerations |
|---|---|---|
| LLM.int8() · 13 NeurIPS, 2022 Author code |
Uses vector-wise 8-bit computation for ordinary features and a 16-bit decomposition for the outlier feature dimensions. Activation outliers determine which dimensions receive special treatment. | 8-bit matrix multiplication with an FP16 outlier path. The effective execution is mixed precision. Include outlier extraction and the high-precision multiplication. |
| AWQ · 47 MLSys, 2024 Author code |
Uses activation statistics to identify important weight channels and searches equivalent channel scales that preserve projection outputs. The full-precision product is unchanged before quantization. | Weight-only; low-bit dense representation. Salient channels are protected through scaling. The survey explicitly states that the final quantized matrix does not keep separate FP16 outlier weights. |
The shape of the protected subset matters. A stored column and a collection of arbitrary entries impose different metadata and memory-access requirements.
| Method and source | Error-control mechanism | Setting and deployment considerations |
|---|---|---|
| OWQ · 36 AAAI, 2024 Author code |
Identifies weak columns that are particularly sensitive to quantization and preserves them at higher precision. The paper also includes Weak Column Tuning as a limited adaptation step. | Weight-only; selected high-precision columns. Include the preserved columns and any tuning stage. Structured preservation differs from arbitrary sparse weight exceptions. |
| SpQR · 15 ICLR, 2024 Author code |
Separates weights with unusually large quantization error into a sparse higher-precision representation. The dense remainder is quantized with group-wise scales. | Weight-only; approximately 3–4 bits per parameter. The sparse component needs values, indices, and decoding support. Include these costs when reporting compression. |
Average bit-width depends on the precision allocation, scales, masks, and residual representation. Binary or sub-2-bit labels must be read with those details.
| Method and source | Error-control mechanism | Setting and deployment considerations |
|---|---|---|
| SliM-LLM · 33 ICML, 2025 Author code |
Assigns bit-widths to groups according to salience structure and calibrates quantizers with element-level salience weights. The allocation avoids arbitrary sparse exceptions. | Group-level mixed-precision weights. Report the distribution of group precisions and metadata. SliM-LLM is distinct from the joint-compression method SLiM in reference [58]. |
| PB-LLM · 100 ICLR, 2024 Author code |
Protects salient weights at higher precision and binarizes the remaining weights. Its PTQ variant uses Hessian-guided reconstruction; the source also discusses a training-aware variant. | Partially binarized weights with a protected subset. Partial binarization is not uniform one-bit storage. State the protected fraction and whether the PTQ or training-aware variant is used. |
| BiLLM · 32 ICML, 2024 Author code |
Separates salient and non-salient weights, applies binary residual approximation to the salient part, and uses distribution-guided splitting for the remaining weights. | Binary-weight PTQ with salience-dependent treatment. Count all binary residual components, scales, and partition information. A one-bit label alone does not describe the full representation. |
| PTQ1.61 · 105 ACL, 2025 Author code |
Uses an activation-derived one-dimensional mask to select salient channels for 4-bit quantization. Other channels are binarized with block-wise scale optimization and preprocessing. | Sub-2-bit weights; salient channels at 4 bits. The reported average rate depends on the mask and allocation. Preserve the distinction between structured channel selection and element-wise masks. |
Optimization methods treat rounding, clipping, scales, or transformations as quantities to estimate during quantization. The source includes calibration-driven learning and parameter-only balancing in this family. It also includes foundational pre-LLM work where that work establishes a principle used by later LLM methods. See Section 3.4.
Figure 6, page 18: rounding and clipping, block-wise differentiable calibration, and equivalent transformations.
The quantizer chooses among discrete reconstruction values, but the calibration objective can measure error after a complete projection or block. The nearest weight value need not be the choice with the smallest output discrepancy.
| Method and source | Error-control mechanism | Setting and deployment considerations |
|---|---|---|
| AdaRound · 60 ICML, 2020 |
Learns up-or-down rounding decisions by minimizing local output reconstruction error on unlabeled calibration data. It motivates choosing grid assignments through the computation they affect. | Foundational PTQ; predates LLM-scale methods. Treat it as a foundational method. Inclusion in an LLM survey is not evidence of an original LLM-scale evaluation. |
| FlexRound · 39 ICML, 2023 Author code |
Uses element-wise division to learn weight positions relative to the quantization grid together with a common grid scale. Calibration reconstructs blocks. | Weights and activations; W8A8 example. Element-level calibration parameters increase optimization cost. The deployed representation should be described separately from calibration variables. |
| SignRound · 8 EMNLP Findings, 2024 Author code |
Optimizes rounding and clipping using signed gradient updates. The survey describes a short calibration procedure whose learned decisions do not add inference-time parameters. | Low-bit weight quantization. Compare calibration budgets and objectives explicitly. The original SignRound and SignRoundV2 have distinct bibliography entries. |
| SignRoundV2 · 7 arXiv, 2025 Author code |
Extends signed-gradient rounding with adaptive precision allocation and stabilization steps, including loss filtering and scale search. It addresses settings where one bit-width for every layer is too restrictive. | Extremely low-bit weights; adaptive mixed precision. Account for mixed-precision assignments when comparing to uniform precision. The linked AutoRound project serves both paper records. |
| TesseraQ · 43 arXiv, 2024 Author code |
Uses progressive adaptive rounding during block reconstruction. Some continuous rounding variables become fixed binary choices while other variables and dequantization scales remain optimized. | Ultra-low-bit PTQ; 2-bit weight-only example. Specify the complete pipeline and reconstruction scope. Table 4 includes backend stages beyond the central rounding contribution. |
| MPPQ · 87 IJCAI, 2025 |
Combines layer- and block-level supervision with magnitude and directional agreement. It also uses low-rank scaling parameters and a short search to initialize clipping. | Weights and activations; W4A4 example. Report the loss, initialization search, and calibration settings. Its low-rank scaling should not be described as an additive deployed residual. |
A block-level objective includes interactions through attention, gating, normalization, and residual addition. The source distinguishes the original vision-oriented BRECQ study from the LLM-specific OmniQuant method.
| Method and source | Error-control mechanism | Setting and deployment considerations |
|---|---|---|
| BRECQ · 42 ICLR, 2021 Author code |
Uses a full block as the reconstruction unit so calibration can account for interacting errors inside residual and nonlinear computation. It is a methodological predecessor of later LLM block calibration. | Foundational block reconstruction; originally CNNs. The original study is not an LLM benchmark. Its role in the survey is the calibration principle. |
| OmniQuant · 72 ICLR, 2024 Author code |
Freezes pretrained weights and learns clipping thresholds and equivalent transformation parameters through block reconstruction. Calibration variables are incorporated into stored weights and scales. | Weight-only and weight-activation configurations. Separate calibration-time learning from additional inference modules. State the precise weight, activation, and cache configuration. |
The shared idea is to alter tensor distributions while preserving the full-precision product. SmoothQuant chooses channel scales from statistics; LRQuant learns calibration parameters; SINQ balances weights without calibration inputs.
| Method and source | Error-control mechanism | Setting and deployment considerations |
|---|---|---|
| SmoothQuant · 89 ICML, 2023 Author code |
Uses activation and weight statistics to select diagonal scales. The transformation reduces activation ranges while moving part of the range into weights without changing the full-precision product. | Weight and activation; W8A8. The channel-scale rule is statistics-based. The source places it with optimization methods, but it does not require learned rounding or a dense learned rotation. |
| LRQuant · 106 ACL, 2024 Author code |
Learns smoothing and quantization parameters with magnitude and directional output agreement. It initializes scales by logarithmic activation equalization and also describes last-block test-time adaptation. | Weights and activations; W4A4. Report whether test-time adaptation is enabled. Test-set adaptation and ordinary fixed-quantizer evaluation are different protocols. |
| SINQ · 59 ICML, 2026 Author code |
Balances row and column standard deviations using a dampened Sinkhorn-Knopp-style procedure in log space. Weight statistics drive the scale computation without calibration inputs. | Calibration-free low-precision weights. Optimization-based placement does not imply use of external calibration text. Record which scale operations can be absorbed into adjacent computation. |
The survey defines component sensitivity through the degradation caused by quantizing a particular tensor or module while the remaining computation is held fixed or calibrated. This is a location-dependent question: weight perturbations, attention-score perturbations, reused cache errors, and output-logit changes do not enter the computation in the same way. See Section 4.
Figure 7, page 22. The numbered sites distinguish outliers after normalization, query/key score changes, reused cache error, MLP inputs and outputs, and vocabulary-logit distortion.
| Component | Sensitivity described in the survey | Implication for quantization |
|---|---|---|
| Query and key projections | Their errors perturb attention scores and can change how attention is distributed over earlier tokens. | Inspect attention behavior and the activations entering the projections; local weight error alone does not characterize the effect. |
| Value and output projections | Value errors change retrieved content. The output projection returns that content to the residual stream. | Measure projection and block outputs, and account for propagation into later blocks. |
| MLP up, gate, and down projections | Expansion inputs can contain large channel outliers. Down-projection errors enter the residual stream. | Large weight matrices offer storage savings, while lower-bit activations require careful clipping, smoothing, or reconstruction. |
| Key cache and value cache | Cached errors persist and are reused. Keys affect scores; values affect retrieved content. | Use explicit key/value settings and evaluate across context lengths. The source discusses channel-aware keys and token-aware values. |
| Normalization and residual paths | These operations influence the inputs and states used by many large projections. | Their small parameter count does not make them harmless targets. Record retained precision and fused transformations. |
| Embeddings, language-model head, and logits | The output path interacts directly with vocabulary probabilities. | Record exceptions in the final projection and evaluate generation behavior under the actual output configuration. |
The cache asymmetry is developed in KIVI · 52.
Calibration estimates the information and parameters needed by a quantizer for a fixed pretrained model. The survey connects this step to range estimation, channel importance, reconstruction error, and transformation selection. The full-precision model can supply layer or block targets on unlabeled inputs. See Section 5.
| Strategy | Information and objective | Source examples or discussion | Main issue to report |
|---|---|---|---|
| Range-based fitting | Observed minima and maxima; optional percentile clipping. | Section 5.1. | Granularity, clipping threshold, and treatment of persistent outlier channels. |
| Distributional fitting | Histograms and a distribution-discrepancy criterion. | Section 5.2. | Histogram construction, objective, and whether preserving tensor values predicts the relevant functional behavior. |
| Layer-wise reconstruction | Projection outputs on shared calibration inputs. | AdaRound 60; GPTQ 20. | Reconstruction scope, data, curvature estimation, and the treatment of earlier quantized layers. |
| Block-wise reconstruction | Outputs of composed attention, MLP, nonlinear, and residual computation. | BRECQ 42; OmniQuant 72; SignRound 8; TesseraQ 43. | Optimization budget, trainable quantizer parameters, reconstruction objective, and validation protocol. |
| Activation-aware selection | Channel statistics that determine scaling or weight salience. | SmoothQuant 89; AWQ 47; LRQuant 106. | Representativeness of activation ranges, outliers, and deployment prompts. |
| Granularity-aware fitting | Parameter sharing at tensor, channel, group, or block scope. | Table 6. | Additional metadata and compatibility with the target kernel. |
| Calibration-data selection | Examples chosen to cover the intended prompt structure and sequence lengths. | Section 5.5. | Data source, sample count, token length, selection procedure, and separation from evaluation. |
| Parameter-only fitting | Pretrained weights and analytical or optimization-based statistics. | EasyQuant 79; AdpQ 24; SINQ 59; OptRot 23. | Assumptions that substitute for observed deployment activations. |
For a linear layer, reconstruction can compare
In calibration, short generic text can fail to represent long prompts, dialogue, code, mathematics, or reasoning-heavy use cases. Calibration data should be described separately from the held-out data used to evaluate the quantized model. A larger sample count does not by itself demonstrate that the relevant activation patterns were covered. See Sections 5.5–5.6.
| Paper or resource | Year in survey | Role in the survey |
|---|---|---|
| EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs · 79 | 2023 | Parameter-only range optimization with outlier treatment. |
| AdpQ: A Zero-shot Calibration Free Adaptive Post Training Quantization Method for LLMs · 24 | 2024 | Calibration-free adaptive treatment of salient and non-salient weights. |
| Self-calibration for Language Model Quantization and Pruning · 88 | 2025 | Self-generated calibration inputs for quantization and pruning. |
| Zero-shot Quantization: A Comprehensive Survey · 34 | 2025 | Background survey for checking zero-shot and data-free terminology. |
Section 6 organizes open questions around the interaction between the quantization algorithm, numeric format, kernel, and hardware. Its five clusters also include architecture-specific sensitivity, long-horizon inference, and reliability. The following sections preserve that organization; they summarize the supplied manuscript’s research agenda and do not claim an exhaustive current-state survey. See Section 6.
Figure 8, page 27. The format, algorithm, kernel, and hardware need compatible choices; reliability and evaluation apply across the stack.
The source discusses native FP4 formats, block-level scales, the interaction between rotations and block layouts, and the gap between storage reduction and realized execution speed. Its examples include MXFP4 and NVFP4, format-aware quantization, block rotations, and fused low-bit kernels. The research question is which quantizer and transformation fit the deployed numerical representation and execution path.
| Paper or resource | Year in survey | Role in the survey |
|---|---|---|
| Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization · 17 | 2026 | Microscaling FP4 quantization and format-aware reconstruction. |
| Pretraining Large Language Models with NVFP4 · 64 | 2025 | NVFP4 numerical-format background cited by the survey. |
| Block Rotation is All You Need for MXFP4 Quantization · 74 | 2026 | Block rotations matched to MXFP4 quantization. |
| Pushing the Limits of Block Rotations in Post-Training Quantization · 70 | 2026 | Permutation before block rotations and analysis of block structure. |
| MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models · 21 | 2025 | Low-bit weight execution and kernel-level inference support. |
The source identifies difficult cases below four bits, especially joint low-bit weights and activations. Its proposed directions include codebooks, incoherence processing, selective bit allocation, ternarization, recovery, and joint compression with sparsity or low-rank structure. The theoretical question concerns the relation between local reconstruction error and complete-model behavior; the source’s existing guarantees and its end-to-end claims need to be distinguished.
| Paper or resource | Year in survey | Role in the survey |
|---|---|---|
| PT2-LLM: Post-Training Ternarization for Large Language Models · 95 | 2026 | Post-training ternarization. |
| VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models · 51 | 2024 | Extreme low-bit vector quantization. |
| Ptq1. 61: Push the real limit of extremely low-bit post-training quantization methods for large language models · 105 | 2025 | Structured sub-2-bit weight quantization. |
| Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization · 4 | 2025 | Layer-wise quantization-error propagation. |
| Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs · 28 | 2026 | Joint quantization and sparsification. |
| SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression · 58 | 2025 | Joint quantization, sparsity, and low-rank approximation. |
SliM-LLM 33 and SLiM 58 are different papers. The first concerns salience-driven mixed precision; the second concerns joint weight compression with sparsity and low-rank approximation. Their similarly spelled names should not be merged into one entry.
Mixture-of-experts models add routing decisions and uneven expert coverage during calibration. Recurrent and state-space architectures reuse states across sequence positions. Multimodal models require attention to modality-specific statistics, while diffusion language models introduce denoising-step-dependent activations. The source treats these as distinct settings whose sensitive components and calibration requirements need dedicated study.
The workload section concerns long contexts, long reasoning generations, test-time scaling, and repeated model calls in agentic tasks. We ask whether the quantized model preserves behavior across the complete generation or interaction process. Cache reuse, generation length, tool arguments, retained constraints, and recovery after errors require evaluation beyond short static prompts.
We argue that accuracy and perplexity can miss changes in model behavior. We include safety, fairness, factual recall, explanation quality, the interaction between quantization and adaptation, and reproducible cost measurement. The research agenda is to specify which behaviors are preserved and to test them under a documented quantization and deployment configuration.
The survey draws evidence from studies with different models, quantized tensors, calibration procedures, precision exceptions, and hardware.
The following reporting fields turn the distinctions in Sections 2–6 into an experiment record.
| Field | Information to record |
|---|---|
| Model | Exact checkpoint, size, model family, instruction or reasoning variant, and tokenizer. |
| Quantized scope | Weights, activations, keys, values, and all retained higher-precision components. |
| Numerical representation | Stored bit-widths, number format, grouping, scales, zero-points, codebooks, masks, residual terms, and effective storage. |
| Calibration | Dataset or information source, sample count, sequence length, selection rule, reconstruction scope, optimized parameters, and time. |
| Adaptation | Fine-tuning, codebook tuning, adapters, weak-column tuning, or test-time updates enabled in the reported pipeline. |
| Quality | Perplexity or task accuracy together with workload-specific reasoning, long-context, agentic, factuality, safety, or fidelity measurements. |
| Generation | Prompt lengths, output lengths, decoding settings, and cache policy used for the quality measurements. |
| System configuration | Named device, backend, numerical kernel, batch size, memory use, and whether transforms or corrections execute online. |
| Runtime | Separate prefill and decode measurements, throughput or latency, and the exact baseline. |
| Energy and reporting | Energy per token where measured, measurement scope, and enough configuration information to reproduce the comparison. |
WikiText-2, C4, and Penn Treebank occur in the manuscript’s perplexity examples. A result on one corpus does not establish preservation of long-context, reasoning, or safety behavior. Likewise, a nominal weight precision does not establish a complete storage budget or an on-device speedup. See Sections 6.4–6.5.
The following systems and toolkits are cited in the survey.
| Resource | Project | Role in the source |
|---|---|---|
| Marlin 21 | IST-DASLab/marlin | Kernel-level support for low-bit weight inference. |
| TensorRT-LLM 63 | NVIDIA/TensorRT-LLM | Inference framework discussed in the kernel-support section. |
| ExLlama 82 | turboderp/exllama | Quantized-weight inference implementation cited by the manuscript. |
| bitsandbytes / LLM.int8() 13 | bitsandbytes-foundation/bitsandbytes | Implementation associated with mixed-precision 8-bit matrix multiplication. |
| LLMC 27 | ModelTC/LightCompress | Compression toolkit and benchmarking resource. |
| Entry field | Required information |
|---|---|
| Identity | Exact paper title, authors, year or version, and a primary publication link. |
| Placement | Primary error-control family, subgroup, and any important secondary mechanism. |
| Scope | Quantized tensors, precision settings, grouping, and retained high-precision terms. |
| Calibration | Data requirements, reconstruction scope, optimized parameters, and any recovery training. |
| Execution | Online transforms, codebook or sparse lookup, auxiliary terms, and reported hardware evidence. |
| Code | An author-linked repository; distinguish the paper implementation from third-party reimplementations. |
| Evidence | The paper section or table supporting a numerical result, plus its model, metric, and baseline. |
New papers added after this edition will be marked as repository additions until they are incorporated into the manuscript.
@misc{rababah2026posttrainingquantization,
title = {Post-Training Quantization for Large Language Models: A Survey},
author = {Baha Rababah and Yuzhang Shang and Carson K. Leung and Cuneyt G. Akcora and Mubarak Shah},
year = {2026},
note = {Manuscript}
}This repository is licensed under the Creative Commons Attribution-ShareAlike 4.0 International license (CC BY-SA 4.0).
You may share and adapt the material for any purpose, provided that appropriate credit is given and derivative works are distributed under the same license.
Third-party papers, code, and assets remain subject to their own licenses and terms. The PDF’s existing notices remain unchanged, and this repository does not make any claim regarding journal acceptance or publication status.







