Kuaishou and Nanjing University propose BPS: adding reverse preference pairs for boundary failures so the same response is rejected where wrong and chosen where right
Synopsis
Kuaishou Technology and Nanjing University introduce Bidirectional Preference Synthesis (BPS): for each validated boundary failure, BPS keeps the conventional forward pair under the original prompt and adds a reverse pair under a synthesized achieved prompt, so the same response is rejected where it is wrong and chosen where it is right; on Qwen3-4B-Instruct-2507, BPS preserves original-side pairwise ranking while raising achieved-side ranking accuracy on held-out crossed anchors from 6.8% to 62.3%, without changing the DPO objective, training a reward model, or requiring online sampling.
Figure 1: Forward-only DPO data rejects τ fail \tau_{\text{fail}} only under P orig P_{\text{orig}} . BPS adds an achieved-prompt reverse pair, so the same response pair teaches a prompt-conditioned preference boundary.
arXivInterpretation
The paper reframes boundary failures as a preference-data construction problem: a rejected response is not necessarily bad in itself, but violates the given prompt while coherently satisfying a nearby intent or constraint setting, so forward-only supervision can conflate prompt-specific rejection with response-intrinsic undesirability. Prior correction-based offline preference pipelines treat failures only as rejected responses under the original prompt; the paper makes explicit that this supervision is incomplete for boundary failures and gives an actionable construction view. The paper illustrates the structure with instruction-following tasks carrying verifiable atomic constraints and notes the same structure appears in tool arguments and multilingual multi-turn settings; it is framed as a data-construction problem rather than a new loss.
For each validated boundary failure, BPS keeps the conventional forward pair under the original prompt and adds a reverse pair under a synthesized achieved prompt, forming a crossed anchor in which the same response is rejected where it is wrong and chosen where it is right. Unlike mechanical constraint reversal, BPS starts from the student's actual failed response, synthesizes a standalone readable achieved prompt, and keeps the reverse label only after eligibility and validation; unlike RPO/IOPO-style input-side construction, BPS ties each reverse pair to an observed boundary failure. The pipeline uses outcome-summary eligibility, achieved-prompt synthesis, and cycle-consistency validation; 282 of 316 pre-filtered reverse candidates pass validation (89.2%), and the main instantiation yields 1,443 forward pairs plus 282 reverse pairs (1,725 total, reverse density 0.195).
On held-out crossed anchors, BPS preserves original-side pairwise ranking while substantially improving achieved-side ranking: achieved-side pairwise ranking accuracy rises from 6.8% to 62.3% on the Gemini probe and from 9.3% to 61.6% on the Kimi-K2.6 cross-teacher probe. Forward-DPO rarely assigns the intended achieved-side ordering, whereas BPS shifts the reward geometry without sacrificing the original-side boundary; budget-matched controls show duplicating 282 forward pairs leaves 6.8% unchanged and mechanically reversed constraint pairs reach only 32.9%, indicating semantically matched achieved prompts are key. Reported on 146 Gemini pipeline-held-out crossed anchors using token-normalized DPO implicit-reward margins and the fraction with the intended sign; across three training seeds BPS stays above 63% achieved-side accuracy while Forward-DPO stays near 8%; a blind human audit of 200 held-out anchors finds 99.0% agreement with the reverse preference direction and 0.5% leakage.
In downstream evaluations, the clearest separation from Forward-DPO appears in multilingual multi-turn instruction following, with consistent capability-retention patterns on agentic/tool-use and code checks. Forward-DPO improves some shallow instruction-following metrics but falls below Base on most FollowBench metrics and later Multi-IF turns; BPS scores above Base on all listed IFBench and IFEval variants, attains the strongest listed FollowBench scores, and scores above both comparison models at every Multi-IF turn, with the BPS–Fwd gap growing with turn depth (0.72, 3.37, 4.34). The main comparison uses the same number of optimizer updates at step 100; a three-seed check gives IFEval averages of 87.79±0.23 for Forward-DPO versus 88.83±0.59 for BPS; in the function-calling instantiation BPS shows a consistent directional advantage across full-call accuracy, achieved-side preference ranking, achieved-prompt exact match, and joint boundary success.
Perspective
The result targets offline preference optimization pipelines using standard DPO: BPS changes only the preference data, introduces no new loss, trains no reward model, and requires no online sampling, so it can be embedded in existing offline pipelines. The method applies where verifiable constraints or structured action spaces exist; the paper rebuilds the full pipeline in both instruction following and function calling and notes each domain benefits from validation criteria matching its output structure (instruction following checks constraint satisfaction and leakage, while function calling also checks parseability, schema validity, function identity, and argument values). Achieved intent is represented as a natural-language prompt so standard preference-optimization pipelines can consume the data; in formal action spaces, schemas or constraint lists may also be stored alongside the achieved prompt for inspection. Construction cost shifts from online reinforcement learning to offline data construction: the main construction run uses about 2,500 teacher calls / 5M tokens, while a main DPO run takes about 0.2 hours (1.6 GPU-hours) and the total fine-tuning budget is about 15 GPU-hours, with teacher inference excluded from the GPU-hour estimate. The authors note these steps are independent of DPO training and can be batched, cached, or replaced with cheaper verifier models in larger runs.
Although achieved-side ranking accuracy rises from 6.8% to 62.3%, the absolute BPS achieved margin remains small, so the authors treat this as a clear shift in reward geometry rather than complete preference reversal. The Gemini held-out crossed-anchor probe is treated as a reward-geometry diagnostic rather than teacher-independent generalization, which is why a Kimi-K2.6 cross-teacher probe is added. Reverse density involves a trade-off: 0.10 gives a smaller correction, whereas 0.30 reduces IFBench, and the validation-determined density lies between them. The function-calling instantiation uses different boundary data and checkpoint selection, so its numbers are not mixed with the main IFBench-trained comparisons. Most validated anchors are format/structure or lexical-constraint mismatches, with few content-intent and language-mismatch samples, so the failure-type breakdown is described as descriptive rather than a claim about the full distribution of all possible boundary failures. In addition, main results use a fixed step-100 comparison rather than validation-selected best checkpoints, and an earlier Forward-DPO checkpoint exhibits a different narrow-performance/capability-retention trade-off.
