Skip to main content
Back to timeline
arXivSource publication:

BAC pairs identity-linked boxes with LoRA tuning to reach 93.20% person-centric QA at 8B scale

Synopsis

Using LSMDC v2 movie clips, the work builds an automatic character-identification and identity-linked spatial grounding pipeline, manually verifies a benchmark of 750 clips and 3,000 person-centric questions, and compares five identity-grounding representations; visual face boxes combined with textual coordinates prove most consistent across scales, and LoRA fine-tuning of Qwen yields BAC, with BAC-8B reaching 93.20% overall person-centric QA accuracy.

Source-provided article image: Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering
Figure 1 ·

Figure 1: Example of identity-aware captioning and person-centric QA using character identities grounded across sampled video frames.

arXiv

Interpretation

The paper formulates and systematically compares five ways of representing already known character identities to a general-purpose Video-MLLM: face textual coordinates (FTC), face visual boxes (FVB), estimated person visual boxes (PVB), face boxes plus face coordinates (FVB+FTC), and person boxes plus person coordinates (PVB+PTC). Earlier identity-aware description work either replaces SOMEONE tags after generation or requires the model to infer the reference-to-occurrence correspondence itself; here explicit identity-linked spatial cues are supplied whenever reliable detections and tracks exist, and the model must propagate identity across frames where cues are absent. On a manually verified benchmark of 750 clips and 3,000 questions, with Qwen3 at roughly 2B, 4B, and 8B in a same-family comparison, paired McNemar tests show FTC is significantly worse than every visually grounded variant at all scales.

Explicit visual grounding substantially improves person-centric QA: overall accuracy rises from 53.87% to 68.90% at 2B, 62.57% to 79.73% at 4B, and 59.77% to 88.53% at 8B. The result separates the spatial question of who is where from the semantic question of who does what, indicating that part of the identity-association bottleneck can be addressed through input representation rather than architecture change. The same clips, questions, prompts, and identity annotations are used across models within each comparison, with generation configuration held fixed; FVB+FTC gives the most consistent captioning and QA performance across scales, particularly at 2B.

Grounding effectiveness depends on target size: at 2B, PVB+PTC reaches 91.8% versus 80.0% for FVB+FTC on small faces, while FVB+FTC reaches 92.8% and 96.8% on medium and large faces versus 91.3% and 91.1% for PVB+PTC. This yields an actionable representation rule: use estimated person boxes to compensate when the face is very small, and prefer face boxes once the face is sufficiently visible, since expanding to the full person can introduce less precise spatial information. Results are binned by face area with 85 small-face, 1,865 medium-face, and 124 large-face questions, supported by qualitative cases at 2,222 px2 with 52 px height and 391,897 px2 with 760.5 px height.

LoRA fine-tuning of Qwen on about 32K identity-aware captioned clips produces BAC, which is supervised only on captioning yet generalizes to person-centric QA: 68.90% to 80.33% at 2B, 79.73% to 90.13% at 4B, and 87.70% to 93.20% at 8B, all statistically significant. BAC-8B's 93.20% ranks behind only GPT-5.6 Sol (97.07% with thinking, 96.30% without) among the frontier models evaluated here, and significantly exceeds Qwen3-VL-235B-A22B at 86.10%, Gemini 2.5 Flash at 82.63%, and Claude Sonnet 4.6 at 76.10%; BAC-4B's 90.13% also surpasses all evaluated 8-9B base models. LoRA adapters are applied only to the language-model projection layers while the vision encoder and multimodal projector stay frozen; after a hyperparameter sweep, the checkpoint with the lowest validation loss is used, with stable training and validation convergence at all three scales.

Perspective

The results apply to the controlled setting of movie narrative video: identities come from predefined cast lists and movie clips, the benchmark comprises 750 clips and 3,000 person-centric questions, and supervision is about 32K identity-aware captioned clips. The intended audience is video understanding researchers and engineering teams who need to link character identity to appearance, action, location, and interaction, especially those who prefer input representation plus lightweight adaptation over architectural change at 2B to 8B scale. In the ethical discussion the authors state the work aims to advance character-centric video understanding rather than person identification in unconstrained real-world environments, and encourage future applications to consider consent, privacy protections, and restrictions on biometric identity information.

Several values appear as placeholders in the body text, including significance-test p-values, the pixel thresholds for face-area bins, the LoRA learning rate, rank, scaling factor, and dropout, and the trainable parameter counts per scale; these specifics are not given numerically in the text read here and would need the appendices and tables. The correlation coefficients between face area and correctness are likewise presented as ranges. In addition, BAC is fine-tuned only on captioning, and the text offers no mechanism-level analysis of what representational change drives the QA gains; position questions (87.54% for BAC-8B) and interaction questions (91.28%) trail appearance and object questions, leaving spatial and relational reasoning as relatively harder directions.

Sources