Public articles linked to the same research event.
arXiv The authors introduce Query-Aware Streaming Latent Reasoning (QASLR), a post-training framework that selects a bounded frame set, restores its temporal order, and processes it clip by clip through a vision-language backbone, where persistent think tokens accumulate cross-clip evidence and an embed token reads out a fixed-dimensional representation after each update; training combines final contrastive learning, step-wise contrastive supervision, and final-embedding self-distillation, raising HourVideo retrieval Hit@1 from 54.2 to 70.7 and from 57.4 to 72.8 on 2B and 8B Qwen3-VL-Embedding, with gains on the evaluated moment-retrieval and video-QA tasks and transfer of the streaming head to the same-family RzenEmbed 7B.
The authors introduce Query-Aware Streaming Latent Reasoning (QASLR), a post-training framework that selects a bounded frame set, restores its temporal order, and processes it clip by clip through a vision-language backbone, where persistent think tokens accumulate cross-clip evidence and an embed token reads out a fixed-dimensional representation after each update; training combines final contrastive learning, step-wise contrastive supervision, and final-embedding self-distillation, raising HourVideo retrieval Hit@1 from 54.2 to 70.7 and from 57.4 to 72.8 on 2B and 8B Qwen3-VL-Embedding, with gains on the evaluated moment-retrieval and video-QA tasks and transfer of the streaming head to the same-family RzenEmbed 7B.
The authors introduce Query-Aware Streaming Latent Reasoning (QASLR), a post-training framework that selects a bounded frame set, restores its temporal order, and processes it clip by clip through a vision-language backbone, where persistent think tokens accumulate cross-clip evidence and an embed token reads out a fixed-dimensional representation after each update; training combines final contrastive learning, step-wise contrastive supervision, and final-embedding self-distillation, raising HourVideo retrieval Hit@1 from 54.2 to 70.7 and from 57.4 to 72.8 on 2B and 8B Qwen3-VL-Embedding, with gains on the evaluated moment-retrieval and video-QA tasks and transfer of the streaming head to the same-family RzenEmbed 7B.
The authors introduce Query-Aware Streaming Latent Reasoning (QASLR), a post-training framework that selects a bounded frame set, restores its temporal order, and processes it clip by clip through a vision-language backbone, where persistent think tokens accumulate cross-clip evidence and an embed token reads out a fixed-dimensional representation after each update; training combines final contrastive learning, step-wise contrastive supervision, and final-embedding self-distillation, raising HourVideo retrieval Hit@1 from 54.2 to 70.7 and from 57.4 to 72.8 on 2B and 8B Qwen3-VL-Embedding, with gains on the evaluated moment-retrieval and video-QA tasks and transfer of the streaming head to the same-family RzenEmbed 7B.