BPO replaces single-response rewards with counterfactual-view behavior packs, lifting Qwen2.5-VL-7B temporal-hard accuracy by 7.8 points
Related research and updatesSynopsis
The work proposes Behavior Pack Optimization (BPO), which replaces the single-response reward in video MLLM post-training with a jointly scored behavior pack of outputs across counterfactual views chosen by question type, requiring stability under irrelevant interventions, sensitivity when key evidence is removed, and abstention when no evidence remains, using the original-view response as an anchor-relative advantage; on TempCompass, MVBench, and NExT-QA, BPO improves Qwen2.5-VL-7B-Instruct macro accuracy by 4.7 pp, the temporal-hard subset by 7.8 pp, and abstention F1 by 20.0 pp over a budget-matched vanilla GRPO baseline from the same SFT checkpoint, transferring to Video-MME, LongVideoBench, and LLaVA-Video-7B.
Figure 2: Overview of Behavior Pack Optimization. Stage A installs the structured EST interface with SFT. Stage B routes each question to a question-typed view set, forms a behavior pack across counterfactual views, scores view-specific and cross-view contract rewards, and updates the policy with anchor-relative advantages.
arXivInterpretation
The paper observes that video MLLMs keep climbing video question answering benchmarks, yet shuffling frames, masking the segment supporting the answer, or occluding the target object barely changes their predictions, indicating accuracy rests on appearance and language priors rather than the temporal evidence the question asks for. It traces this to the unit of post-training: rewards are computed on a single response to the original clip, so the model is never asked to behave consistently across views. The diagnosis rests on the stated near-invariance of predictions under three counterfactual interventions (frame shuffling, evidence-segment masking, target occlusion), as reported at the abstract level.
It proposes Behavior Pack Optimization (BPO), replacing the single response with a behavior pack of outputs across counterfactual views chosen by question type, scored jointly, with a pack reward asking for stability when the intervention is irrelevant, sensitivity when key evidence is removed, and abstention when no evidence remains. It shifts the post-training objective from single-response correctness to cross-view behavioral consistency and folds abstention into the reward structure. The method description comes from the abstract; concrete view-set construction and reward form are not given.
To keep the objective stable at small pack sizes, BPO uses an anchor-relative advantage: the response on the original view serves as a per-prompt reference instead of a group mean over mixed views. It substitutes a per-prompt original-view anchor for the GRPO-style group-mean baseline to accommodate variance from mixing views within a pack. The abstract states the motivation and mechanism but reports neither pack-size values nor a variance analysis.
On TempCompass, MVBench, and NExT-QA, BPO improves Qwen2.5-VL-7B-Instruct macro accuracy by 4.7 pp, the temporal-hard subset by 7.8 pp, and abstention F1 by 20.0 pp against a budget-matched vanilla GRPO baseline from the same SFT checkpoint; gains transfer to Video-MME, LongVideoBench, and LLaVA-Video-7B, and ablations confirm they follow the view sets rather than the rollout count. It provides multi-benchmark, cross-model, ablation-backed comparisons under budget matching and a shared starting point, attributing gains to view-set design. The abstract reports specific percentage-point deltas, three evaluation benchmarks, two transfer benchmarks, one additional model, and an ablation separating view sets from rollout count.
Perspective
The work targets the post-training stage of video multimodal large language models, suited to video question answering settings where the model should answer from temporal evidence and abstain when evidence is absent; its gains are reported on Qwen2.5-VL-7B-Instruct and LLaVA-Video-7B, on TempCompass, MVBench, and NExT-QA plus the transfer benchmarks Video-MME and LongVideoBench. The method depends on counterfactual view sets chosen by question type, so a prerequisite is the ability to construct the corresponding intervention views for a given question type. The abstract positions the pack-level perspective as a starting point for the video MLLM and multimodal post-training community to carry forward toward evidence-grounded video reasoning.
The abstract does not give the concrete construction of view sets, pack-size values, the reward function form, or training cost, nor per-benchmark breakdowns or statistical uncertainty; the ablation is stated as gains following the view sets rather than the rollout count, and its exact comparison setup needs confirmation in the body. The abstention F1 improvement of 20.0 pp is large, so its evaluation protocol and negative-sample composition are worth checking in the body. These are open questions to keep in mind for reproduction and interpretation.
