Skip to main content
Back to timeline
arXivSource publication:

Zeroth-Order Preference Alignment via Comparison Oracles: The ComPO Method

Synopsis

This paper proposes and analyzes Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method based on comparison oracles that, instead of directly optimizing a differentiable preference loss, perturbs the current policy and judges whether each perturbation raises the likelihood of preferred responses and lowers that of dispreferred ones to extract an update direction, thereby exploiting low-margin "noisy" preference pairs; the authors establish a convergence guarantee for the basic offline scheme under smoothness, gradient sparsity, and oracle compatibility, introduce an online version that uses unlabeled policy generations for reverse-KL control, and prove a performance guarantee for a basic constrained scheme under local coverage and in-distribution pairwise reward

Source-provided article image: A Zeroth-Order Paradigm for LLM Preference Alignment
Figure 1 ·

Figure 1: (Left) Percentage of nonzero entries in the final gradient as the gradient-entry threshold λ g \lambda_{g} varies. (Middle) Peak GPU memory used by ComPO for the three model families. (Right) Perturbed output-layer size and wall-clock time for completing 600 600 perturbations on 30 30 NVIDIA A40 GPUs.

arXiv

Interpretation

Introduces ComPO, which treats low-margin preference pairs as comparison signals rather than direct optimization targets, with a practical offline implementation using output-layer perturbations and entry-wise thresholding to refine an aligned policy. Unlike optimizing a fixed DPO-style margin loss, the method takes one-bit comparison signals per perturbation and aggregates them into a normalized update direction, so filtered-out low-margin pairs can still contribute alignment information. The paper provides algorithm descriptions (Algorithm 1 and Algorithm 2) and benchmark results on Mistral-7B, Llama-3-8B, and other models; for example, DPO_clean+ComPO raises AlpacaEval 2 LC from 9.41 to 11.66 on Mistral-7B-Base.

Establishes a best-iterate convergence guarantee for the basic offline scheme and a feasibility-preserving, coverage-based performance bound for the basic online scheme. The offline result shows comparison-query complexity depends only logarithmically on dimension under smoothness, gradient sparsity, and oracle compatibility; the online result links the performance gap to in-distribution pairwise reward error under an exact reverse-KL constraint and local coverage. Theorems 3.2 and 3.4 give explicit probability and error bounds, with proofs in Appendix B; these are theoretical guarantees whose assumptions (such as oracle compatibility with a latent objective and local coverage) are explicitly stated as separate conditions in the main text.

The online extension retains the offline comparison direction and uses only unlabeled current-policy generations to adjust the step size, with empirical evaluation of damping and replay. Unlike acquiring new preference labels online or adding an explicit exploration bonus, online samples only change the step size, not the comparison oracle or the preference labels. Table 10 shows that adding damping improves AlpacaEval 2 LC by 1.23 points for Qwen3-4B-Base and 2.07 points for Gemma-3-4B-it, with replay further improving all three reported metrics; the authors note the practical scheme is a heuristic approximation not covered by Theorem 3.4.

Pair-level likelihood diagnostics show the comparison-oracle update leaves preferred-response log-likelihood nondecreasing and dispreferred-response log-likelihood nonincreasing, consistent with mitigating likelihood displacement. This diagnostic directly inspects the direction of likelihood change for each noisy pair during training, rather than relying only on benchmark win rates. Table 2 reports three independent trials on Llama-3-Instruct-8B and Gemma-2-9B-it; for example, Llama-3-Instruct-8B at γ=1 moves from (-46.761, -47.410) to (-46.728, -47.520); the authors frame this as an in-training sanity check rather than a population-level performance guarantee.

Perspective

The work targets scenarios that align large language models using pairwise preference data, especially low-margin pairs where the reference policy assigns similar likelihoods to preferred and dispreferred responses; the practical offline scheme is limited to sparse entry updates in the output layer (lm_head), with multi-layer perturbation examined only in an ablation; the online scheme targets settings where unlabeled prompts can generate current-policy samples and one wishes to control deviation via a reverse-KL proxy. For practitioners seeking low-memory, task-specific refinement of an existing aligned checkpoint, this framework offers a reusable, modular path.

A careful reader would still watch: how the low-margin selection threshold δ_margin relates to richer criteria such as the CHES score, and whether conclusions are stable under different selection standards; whether the empirical benefits of damping and replay in the practical online scheme, whose length-normalized statistic is not the sequence-level reverse KL, hold on larger models and more tasks; the Arena-Hard drop for Mistral-7B-Instruct is associated with response-length differences, though the text notes this alone does not establish causality; additionally, this is a full-text load with figures presented as text, so verifying curve details would require returning to the original.

Sources