AI Architecture & Systems

System Implementation of Speculative Decoding

5.26.1Draft Model, Target Model, and Verification Worker#

5.26.2Same-Device, Separate-Device, and Heterogeneous Draft/Target Deployment#

5.26.3Acceptance-Aware and Load-Aware Draft Length#

5.26.4Tree/Block Verification and Batch Capacity#

5.26.5Draft/Target KV Cache Management and Rollback#

5.26.6Trade-offs Among Online Requests, Throughput, and Single-Request Latency#

5.26.7Draft Model Training, Updating, and Service Integration#