This AI Research from Google DeepMind Explores the Performance Gap between Online and Offline Methods for AI Alignment
RLHF is the standard approach for aligning LLMs. However, recent advances in offline alignment methods, such as direct preference optimization (DPO) and its variants, challenge the necessity of on-policy sampling in RLHF. Offline methods, which […]
