AI Architecture & Systems
1
Microarchitecture
2
Kernel
3
Compiler
4
Architecture
5
System
6
Algorithms
System
/
§5.18
Batching and Online Scheduling
5.18.1
Static Batching, Dynamic Batching, and Continuous Batching
#
5.18.2
Iteration-Level Scheduling and Token Budgets
#
5.18.3
Chunked Prefill and Decode Priority
#
5.18.4
Preemption, Recompute, and Swap
#
5.18.5
CPU–GPU Overlap and Low-Overhead Schedulers
#
5.18.6
Fairness, Priority, and Mixing Short and Long Requests
#
5.18.7
Load-Aware Batch Size and Adaptive Scheduling
#
← Previous
5.17 The Full Lifecycle of an Inference Request
→ Next
5.19 KV Cache Allocation, Addressing, and Lifecycle