AI Architecture & Systems

The Execution Chain of Device Runtime, Driver, and Firmware

SourcesNVIDIA, AMD, Neuron, Ascend, Tenstorrent and Groq layer mappings.

5.2.1Model Loading, Executable Artifacts, and Device Resource Initialization#

5.2.2Context, Command Queue, Doorbell, Event, and Interrupt#

5.2.3Device Memory Allocation, Virtual Memory, and Host–Device Transfer#

5.2.4The Boundary Between the CUDA Runtime/Driver API and the GPU Kernel Module#

5.2.5The Boundary Among HIP/ROCr-HSA, AQL Queue, and amdgpu/KFD#

5.2.6Neuron libnrt/NEFF, AscendCL, and Device-Specific Execution#

5.2.7TT-Metalium Host/Device Programs and User-Space/Kernel-Space Drivers#

5.2.8Distinguishing a Thin Static Executor from a Dynamic Kernel-Launch Runtime#