AI Architecture & Systems

Alignment and Preference Learning

6.14.1Reward Model, Preference Data, and Pairwise Ranking#

6.14.2RLHF and RLAIF#

6.14.3PPO, KL Regularization, and Reference Policy#

6.14.4DPO, IPO, KTO, and Direct Preference Optimization#

6.14.5Rejection Sampling and Best-of-N Data Selection#

6.14.6Safety Alignment, Refusal, and Preference Generalization#