Public articles linked to the same research event.
arXiv The work proposes Behavior Pack Optimization (BPO), which replaces the single-response reward in video MLLM post-training with a jointly scored behavior pack of outputs across counterfactual views chosen by question type, requiring stability under irrelevant interventions, sensitivity when key evidence is removed, and abstention when no evidence remains, using the original-view response as an anchor-relative advantage; on TempCompass, MVBench, and NExT-QA, BPO improves Qwen2.5-VL-7B-Instruct macro accuracy by 4.7 pp, the temporal-hard subset by 7.8 pp, and abstention F1 by 20.0 pp over a budget-matched vanilla GRPO baseline from the same SFT checkpoint, transferring to Video-MME, LongVideoBench, and LLaVA-Video-7B.
The work proposes Behavior Pack Optimization (BPO), which replaces the single-response reward in video MLLM post-training with a jointly scored behavior pack of outputs across counterfactual views chosen by question type, requiring stability under irrelevant interventions, sensitivity when key evidence is removed, and abstention when no evidence remains, using the original-view response as an anchor-relative advantage; on TempCompass, MVBench, and NExT-QA, BPO improves Qwen2.5-VL-7B-Instruct macro accuracy by 4.7 pp, the temporal-hard subset by 7.8 pp, and abstention F1 by 20.0 pp over a budget-matched vanilla GRPO baseline from the same SFT checkpoint, transferring to Video-MME, LongVideoBench, and LLaVA-Video-7B.
The work proposes Behavior Pack Optimization (BPO), which replaces the single-response reward in video MLLM post-training with a jointly scored behavior pack of outputs across counterfactual views chosen by question type, requiring stability under irrelevant interventions, sensitivity when key evidence is removed, and abstention when no evidence remains, using the original-view response as an anchor-relative advantage; on TempCompass, MVBench, and NExT-QA, BPO improves Qwen2.5-VL-7B-Instruct macro accuracy by 4.7 pp, the temporal-hard subset by 7.8 pp, and abstention F1 by 20.0 pp over a budget-matched vanilla GRPO baseline from the same SFT checkpoint, transferring to Video-MME, LongVideoBench, and LLaVA-Video-7B.
The work proposes Behavior Pack Optimization (BPO), which replaces the single-response reward in video MLLM post-training with a jointly scored behavior pack of outputs across counterfactual views chosen by question type, requiring stability under irrelevant interventions, sensitivity when key evidence is removed, and abstention when no evidence remains, using the original-view response as an anchor-relative advantage; on TempCompass, MVBench, and NExT-QA, BPO improves Qwen2.5-VL-7B-Instruct macro accuracy by 4.7 pp, the temporal-hard subset by 7.8 pp, and abstention F1 by 20.0 pp over a budget-matched vanilla GRPO baseline from the same SFT checkpoint, transferring to Video-MME, LongVideoBench, and LLaVA-Video-7B.