AI Architecture & Systems

Low-Precision and Quantized Kernels

2.13.1Quantize/Dequantize and Scale Computation#

2.13.2Weight-Only GEMM and Weight–Activation GEMM#

2.13.3INT8/INT4 and FP8/FP4 GEMM#

2.13.4Per-Tensor, Per-Channel, and Block Scaling#

2.13.5Microscaling Formats and Scale Layout#

2.13.6Online Quantization, Dequantization Fusion, and Numerical Validation#