AI Architecture & Systems
1
Microarchitecture
2
Kernel
3
Compiler
4
Architecture
5
System
6
Algorithms
Architecture
/
§4.5
NPU: Heterogeneous Compute Engines and Shared Local Memory
Sources
Survey §3 and Table 2; the final two entries broaden the case pool using the corpus.
4.5.1
Google TPU: MXU, Vector/Scalar, and Generation-Specific Units
#
4.5.2
AWS Trainium/Inferentia: NeuronCore and Explicit Scratchpad
#
4.5.3
Huawei Ascend: Da Vinci, Cube/Vector/Scalar, and CANN
#
4.5.4
Intel Gaudi: Matrix Engine, TPC, and Network Integration
#
4.5.5
Microsoft Maia: Matrix/Vector and Cloud Deployment
#
4.5.6
Qualcomm Cloud AI: Matrix, Vector, and Scalar Engines
#
4.5.7
Cambricon MLU: Multi-Core Neural Processor and BANG C
#
4.5.8
Corpus Extensions:
Alibaba T-Head
,
Kunlunxin
, and
Furiosa
#
4.5.9
Corpus Extensions:
Sophgo
,
Vastai
,
Tecorigin
, and
Stream Computing
#
← Previous
4.4 GPU: General-Purpose Parallel Execution and Specialized Matrix Computation
→ Next
4.6 Spatial Dataflow I: PE Arrays and Distributed Local Memory