About the survey

Rapidly growing AI workloads are driving large investments in AI datacenters. This survey classifies industrial AI accelerators into four architectural categories and compares their compute and memory organizations. It examines how node-, rack-, and pod-scale scale-up interconnects support collective communication, and it traces architectural evolution across accelerator generations, showing how increases in arithmetic throughput are accompanied by changes in data delivery, execution coordination, communication, and deployment infrastructure. Together, these analyses highlight key challenges in future AI datacenter design, from workload flexibility, data movement, and communication efficiency to power delivery, cooling, and rapid model evolution, and they show how the trade-off between generality and specialization extends from individual accelerators to datacenter-scale systems. The chart below is an interactive version of the survey's scaling figure: frontier-model parameter counts against per-accelerator dense FP16/BF16 compute, DRAM bandwidth and capacity, and scale-up interconnect bandwidth, extended with accelerator generations documented in the survey's open research corpus.

Model scaling versus accelerator scaling

Each series is normalized to its first observation in the survey's Figure 2 and plotted on a log scale. Legend percentages are endpoint annualized growth rates over the points currently shown. Hover a point for its value and source; click it for details. Scroll or drag to zoom, and use the filters to choose metrics, vendors, product status, and data sets.

Metric
Vendor
Status
Data set
Hollow markers: preliminary or forthcoming products.
Click a point to see the product, its recorded metrics, notes, and sources.

Methodology

Show the visible points as a table
Show the reference list

Sources are numbered in order of first use. Entries used by the points currently shown are highlighted; the rest are dimmed.