AI Architecture & Systems
1
Microarchitecture
2
Kernel
3
Compiler
4
Architecture
5
System
6
Algorithms
Kernel
/
§2.13
Low-Precision and Quantized Kernels
2.13.1
Quantize/Dequantize and Scale Computation
#
2.13.2
Weight-Only GEMM and Weight–Activation GEMM
#
2.13.3
INT8/INT4 and FP8/FP4 GEMM
#
2.13.4
Per-Tensor, Per-Channel, and Block Scaling
#
2.13.5
Microscaling Formats and Scale Layout
#
2.13.6
Online Quantization, Dequantization Fusion, and Numerical Validation
#
← Previous
2.12 MoE Kernels
→ Next
2.14 Communication Kernels: Device-Initiated Transfer, Data Movement, and Compute Fusion