AI Architecture & Systems
1
Microarchitecture
2
Kernel
3
Compiler
4
Architecture
5
System
6
Algorithms
System
/
§5.25
System Support for Sparse Models and Long Context
5.25.1
Weight Sparsity and Sparse Weight Storage
#
5.25.2
Indexing, Routing, and Cache Access in Sparse Attention
#
5.25.3
Dynamic Token Selection and Batch Organization
#
5.25.4
KV Pruning/Compression and Memory Management
#
5.25.5
Load Balancing and Communication Under Dynamic Sparsity
#
5.25.6
Sparse Compute Gains Versus Preprocessing and Metadata Overhead
#
← Previous
5.24 Deploying and Tuning Quantized Models
→ Next
5.26 System Implementation of Speculative Decoding