On the Role of Preference Variance in Preference Optimization

Update: 2025-10-20

Description

This academic paper investigates the concept of Preference Variance (PVar) as a metric for improving the efficiency of Direct Preference Optimization (DPO), a method for aligning large language models (LLMs) with human feedback. The authors establish a theoretical foundation demonstrating that the magnitude of the DPO training gradient is bounded by the PVar of a given prompt, meaning prompts with low PVar contribute minimally to learning. Experimentally, the paper validates that training LLMs using subsets of data identified as having high PVar leads to faster convergence and superior performance compared to using randomly selected data or the entire dataset. Ultimately, the research suggests that strategically selecting high-PVar prompts can drastically reduce the cost of human annotation while maintaining or even improving the final quality of LLM alignment.

Comments

In Channel

Demystifying the Mechanisms Behind Emergent Exploration in Goal-conditioned RL

2025-10-2214:33

Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior

2025-10-2219:04

A Definition of AGI

2025-10-2216:28

Provably Learning from Language Feedback

2025-10-2119:55

In-Context Learning for Pure Exploration

2025-10-2116:30

On the Role of Preference Variance in Preference Optimization

2025-10-2014:42

Training LLM Agents to Empower Humans

2025-10-2013:38

Richard Sutton Declares LLMs a Dead End

2025-10-2013:20

Demystifying Reinforcement Learning in Agentic Reasoning

2025-10-1915:21

Emergent coordination in multi-agent language models

2025-10-1913:57

Learning-to-measure: in-context active feature acquisition

2025-10-1916:02

Andrej Karpathy's insights: AGI, Intelligence, and Evolution

2025-10-1916:11

Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data

2025-10-1812:48

Representation-Based Exploration for Language Models: From Test-Time to Post-Training

2025-10-1817:02

The attacker moves second: stronger adaptive attacks bypass defenses against LLM jail- Breaks and prompt injections

2025-10-1816:08

When can in-context learning generalize out of task distribution?

2025-10-1619:44

The Art of Scaling Reinforcement Learning Compute for LLMs

2025-10-1613:41

A small number of samples can poison LLMs of any size

2025-10-1613:58

Dual Goal Representations

2025-10-1417:11

Welcome to the Era of Experience

2025-10-1416:42

00:00

On the Role of Preference Variance in Preference Optimization

#box-pro-ellipsis-176110552605522{-webkit-line-clamp:2;}On the Role of Preference Variance in Preference Optimization

On the Role of Preference Variance in Preference Optimization

Enoch H. Kang

On the Role of Preference Variance in Preference Optimization