AI Architecture & Systems

System Metrics, Queueing, and Capacity Planning

5.3.1TTFT, TPOT/ITL, and End-to-End Latency#

5.3.2Request Throughput, Token Throughput, and Goodput#

5.3.3Latency Percentiles, SLOs, and Tail Latency#

5.3.4Arrival Process, Queueing, and Little's Law#

5.3.5MFU, HFU, Device Utilization, and Effective Work#

5.3.6Tokens/Dollar, Tokens/Joule, and Total Cost of Ownership#

5.3.7Capacity, Concurrency, Batch Size, and Load Inflection Points#