Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

ENPO uses adversarial policy exploration to remove iterative NLHF's exponential dependence on the inverse KL-regularization parameter, and consistently beats the evaluated RLHF and NLHF baselines on Llama-3-8B-Instruct

This work studies online iterative Nash learning from human feedback (NLHF) and identifies exploration as a key obstacle: standard iterative NLHF can incur an exponential dependence on the inverse KL-regularization parameter, showing that implicit exploration through policy updates can be insufficient; it proposes ENPO, which combines SFT-type regularization with adversarial policy exploration to eliminate this exponential dependence without minimax oracles or explicit preference-model estimation, further introduces BENPO, which uses additional oracles to achieve an O(log T) regret bound, and develops DENPO, a practical variant of ENPO for fine-tuning LLMs; with Llama-3-8B-Instruct, DENPO shows consistent improvements over the evaluated RLHF and NLHF baselines across multiple benchmarks.