Model Evaluation and Cross-Layer Trade-offs
6.35.1Perplexity, Accuracy, and Task Quality#
6.35.2MMLU, GSM8K/MATH, HumanEval, and SWE-bench#
6.35.3Long-Context, Multimodal, and Agent Benchmarks#
6.35.4Pass@k, Success Rate, and Inference Compute Budget#
6.35.5Data Contamination, Judge Bias, and Evaluation Reproducibility#
6.35.6The Quality–Latency–Throughput–Memory–Energy–Cost Pareto Frontier#
6.35.7Distinguishing Algorithmic Improvement, Kernel Speedup, and End-to-End Gain#
6.35.8Model Architecture Change and the Adaptability Range of Deployed Hardware#