Public articles linked to the same research event.
arXiv This work studies online iterative Nash learning from human feedback (NLHF) and identifies exploration as a key obstacle: standard iterative NLHF can incur an exponential dependence on the inverse KL-regularization parameter, showing that implicit exploration through policy updates can be insufficient; it proposes ENPO, which combines SFT-type regularization with adversarial policy exploration to eliminate this exponential dependence without minimax oracles or explicit preference-model estimation, further introduces BENPO, which uses additional oracles to achieve an O(log T) regret bound, and develops DENPO, a practical variant of ENPO for fine-tuning LLMs; with Llama-3-8B-Instruct, DENPO shows consistent improvements over the evaluated RLHF and NLHF baselines across multiple benchmarks.
This work studies online iterative Nash learning from human feedback (NLHF) and identifies exploration as a key obstacle: standard iterative NLHF can incur an exponential dependence on the inverse KL-regularization parameter, showing that implicit exploration through policy updates can be insufficient; it proposes ENPO, which combines SFT-type regularization with adversarial policy exploration to eliminate this exponential dependence without minimax oracles or explicit preference-model estimation, further introduces BENPO, which uses additional oracles to achieve an O(log T) regret bound, and develops DENPO, a practical variant of ENPO for fine-tuning LLMs; with Llama-3-8B-Instruct, DENPO shows consistent improvements over the evaluated RLHF and NLHF baselines across multiple benchmarks.
This work studies online iterative Nash learning from human feedback (NLHF) and identifies exploration as a key obstacle: standard iterative NLHF can incur an exponential dependence on the inverse KL-regularization parameter, showing that implicit exploration through policy updates can be insufficient; it proposes ENPO, which combines SFT-type regularization with adversarial policy exploration to eliminate this exponential dependence without minimax oracles or explicit preference-model estimation, further introduces BENPO, which uses additional oracles to achieve an O(log T) regret bound, and develops DENPO, a practical variant of ENPO for fine-tuning LLMs; with Llama-3-8B-Instruct, DENPO shows consistent improvements over the evaluated RLHF and NLHF baselines across multiple benchmarks.
This work studies online iterative Nash learning from human feedback (NLHF) and identifies exploration as a key obstacle: standard iterative NLHF can incur an exponential dependence on the inverse KL-regularization parameter, showing that implicit exploration through policy updates can be insufficient; it proposes ENPO, which combines SFT-type regularization with adversarial policy exploration to eliminate this exponential dependence without minimax oracles or explicit preference-model estimation, further introduces BENPO, which uses additional oracles to achieve an O(log T) regret bound, and develops DENPO, a practical variant of ENPO for fine-tuning LLMs; with Llama-3-8B-Instruct, DENPO shows consistent improvements over the evaluated RLHF and NLHF baselines across multiple benchmarks.