AI Architecture & Systems
1
Microarchitecture
2
Kernel
3
Compiler
4
Architecture
5
System
6
Algorithms
Compiler
/
§3.6
Loop Transformation and Tensorization
3.6.1
Tiling, Interchange, Unrolling, and Vectorization
#
3.6.2
Loop Fusion, Fission, and Software Pipelining
#
3.6.3
Affine Analysis and Polyhedral Optimization
#
3.6.4
Parallelization and Reduction Transformation
#
3.6.5
Tensorization and Matrix-Instruction Matching
#
3.6.6
Mapping Loop Schedules onto Systolic/SIMT Hardware
#
← Previous
3.5 Graph-Level Optimization
→ Next
3.7 Layout and Data-Movement Optimization