6. AlgorithmsModels, Training Methods, and Generation Algorithms
Retained model and algorithm curriculum. The hardware survey motivates workload diversity and adaptability, but is not the source for the detailed algorithms or model-family histories below.
6.1Deep Learning and Model Computation Fundamentals66.2The Full Lifecycle of a Foundation Model66.3Tokenization and Input Representation66.4Transformer Structure and Its Evolution66.5The Basic Mechanism of Attention66.6MHA, MQA, GQA, and MLA76.7Position Encoding and Context Extension66.8Sparse and Compressed Attention86.9Linear Attention, SSMs, and Hybrid Models66.10Mixture of Experts76.11Pretraining Objectives and Optimization Algorithms66.12Data, Training Recipes, and Model Scaling66.13SFT and Parameter-Efficient Fine-Tuning66.14Alignment and Preference Learning66.15Reasoning and RL Post-Training76.16Autoregressive Inference and Sampling66.17Test-Time Compute and Inference Strategies66.18Speculative Decoding106.19Parallel, Blockwise, and Diffusion Language Generation76.20Quantization: Algorithms and Numerical Methods86.21Sparsification, Pruning, and Model Compression76.22Multimodal Foundation Models76.23Diffusion, Flow Matching, and Visual Generation86.24Retrieval, Tools, and Agent Algorithms76.25AI Workloads Beyond LLMs76.26Communication-Efficient Learning and Distributed Optimization76.27Model Evolution: From Sequence Models to Foundation Models66.28Model Family: Llama66.29Model Family: Mistral and Mixtral56.30Model Family: DeepSeek76.31Model Family: Qwen66.32Model Family: Kimi66.33Model Family: GLM56.34Other Model Families and Reading Public Reports66.35Model Evaluation and Cross-Layer Trade-offs8