arXiv The work proposes Recursive Harness Distillation (RHD): a strong agent (GPT-6 Astra) first interacts with a frozen vision-language-action model and distills effective intervention strategies into a reusable playbook, then recursively revises that playbook using the light agent's (GPT-5.6 Luna) execution feedback, improving manipulation without updating any model parameters—real-world success rises from 37.3% to 64.0%, the light agent with the playbook reaches 66.7% on SimplerEnv Bridge versus 41.7% for GR00T alone, and the same playbook also lifts the strong agent to 79.2%.
The work proposes Recursive Harness Distillation (RHD): a strong agent (GPT-6 Astra) first interacts with a frozen vision-language-action model and distills effective intervention strategies into a reusable playbook, then recursively revises that playbook using the light agent's (GPT-5.6 Luna) execution feedback, improving manipulation without updating any model parameters—real-world success rises from 37.3% to 64.0%, the light agent with the playbook reaches 66.7% on SimplerEnv Bridge versus 41.7% for GR00T alone, and the same playbook also lifts the strong agent to 79.2%.
The work proposes Recursive Harness Distillation (RHD): a strong agent (GPT-6 Astra) first interacts with a frozen vision-language-action model and distills effective intervention strategies into a reusable playbook, then recursively revises that playbook using the light agent's (GPT-5.6 Luna) execution feedback, improving manipulation without updating any model parameters—real-world success rises from 37.3% to 64.0%, the light agent with the playbook reaches 66.7% on SimplerEnv Bridge versus 41.7% for GR00T alone, and the same playbook also lifts the strong agent to 79.2%.
The work proposes Recursive Harness Distillation (RHD): a strong agent (GPT-6 Astra) first interacts with a frozen vision-language-action model and distills effective intervention strategies into a reusable playbook, then recursively revises that playbook using the light agent's (GPT-5.6 Luna) execution feedback, improving manipulation without updating any model parameters—real-world success rises from 37.3% to 64.0%, the light agent with the playbook reaches 66.7% on SimplerEnv Bridge versus 41.7% for GR00T alone, and the same playbook also lifts the strong agent to 79.2%.
arXiv Using a matched ladder of 11 sparse MoE decoders (1.1B–44B total, 71M–2.4B active non-embedding parameters) with identical data mixture and optimization, the work fits separate text and multimodal scaling laws for encoder-free and encoder-based MLLMs and finds that removing the visual encoder shifts compute-optimal multimodal allocation toward larger models (model allocation exponent rising from about 0.464 to about 0.570) while leaving text nearly unchanged; encoder-free models underperform at small scales but their multimodal loss falls faster with compute, with extrapolation predicting an efficiency crossover on the order of 10^22 FLOPs (point estimate about 3.
Using a matched ladder of 11 sparse MoE decoders (1.1B–44B total, 71M–2.4B active non-embedding parameters) with identical data mixture and optimization, the work fits separate text and multimodal scaling laws for encoder-free and encoder-based MLLMs and finds that removing the visual encoder shifts compute-optimal multimodal allocation toward larger models (model allocation exponent rising from about 0.464 to about 0.570) while leaving text nearly unchanged; encoder-free models underperform at small scales but their multimodal loss falls faster with compute, with extrapolation predicting an efficiency crossover on the order of 10^22 FLOPs (point estimate about 3.
Using a matched ladder of 11 sparse MoE decoders (1.1B–44B total, 71M–2.4B active non-embedding parameters) with identical data mixture and optimization, the work fits separate text and multimodal scaling laws for encoder-free and encoder-based MLLMs and finds that removing the visual encoder shifts compute-optimal multimodal allocation toward larger models (model allocation exponent rising from about 0.464 to about 0.570) while leaving text nearly unchanged; encoder-free models underperform at small scales but their multimodal loss falls faster with compute, with extrapolation predicting an efficiency crossover on the order of 10^22 FLOPs (point estimate about 3.
Using a matched ladder of 11 sparse MoE decoders (1.1B–44B total, 71M–2.4B active non-embedding parameters) with identical data mixture and optimization, the work fits separate text and multimodal scaling laws for encoder-free and encoder-based MLLMs and finds that removing the visual encoder shifts compute-optimal multimodal allocation toward larger models (model allocation exponent rising from about 0.464 to about 0.570) while leaving text nearly unchanged; encoder-free models underperform at small scales but their multimodal loss falls faster with compute, with extrapolation predicting an efficiency crossover on the order of 10^22 FLOPs (point estimate about 3.
arXiv The work introduces G²PTQ, a post-training quantization framework that combines first- and second-order information under a globally supervised, block-wise objective: it recomputes gradient and Hessian estimates before quantizing each Transformer block to avoid staleness and uses trust-region scaling to bound the exact first-order compensation step, achieving lower KL divergence and higher downstream QA accuracy than baselines such as GPTQ, GuidedQuant, and GPTAQ across 13 dense and 2 MoE models from 0.6B to 125B parameters at 2/3/4-bit weight and 4-bit weight-activation settings.
The work introduces G²PTQ, a post-training quantization framework that combines first- and second-order information under a globally supervised, block-wise objective: it recomputes gradient and Hessian estimates before quantizing each Transformer block to avoid staleness and uses trust-region scaling to bound the exact first-order compensation step, achieving lower KL divergence and higher downstream QA accuracy than baselines such as GPTQ, GuidedQuant, and GPTAQ across 13 dense and 2 MoE models from 0.6B to 125B parameters at 2/3/4-bit weight and 4-bit weight-activation settings.
The work introduces G²PTQ, a post-training quantization framework that combines first- and second-order information under a globally supervised, block-wise objective: it recomputes gradient and Hessian estimates before quantizing each Transformer block to avoid staleness and uses trust-region scaling to bound the exact first-order compensation step, achieving lower KL divergence and higher downstream QA accuracy than baselines such as GPTQ, GuidedQuant, and GPTAQ across 13 dense and 2 MoE models from 0.6B to 125B parameters at 2/3/4-bit weight and 4-bit weight-activation settings.
The work introduces G²PTQ, a post-training quantization framework that combines first- and second-order information under a globally supervised, block-wise objective: it recomputes gradient and Hessian estimates before quantizing each Transformer block to avoid staleness and uses trust-region scaling to bound the exact first-order compensation step, achieving lower KL divergence and higher downstream QA accuracy than baselines such as GPTQ, GuidedQuant, and GPTAQ across 13 dense and 2 MoE models from 0.6B to 125B parameters at 2/3/4-bit weight and 4-bit weight-activation settings.
arXiv The work introduces CoWindow Attention (CoWA), in which all KV heads share near-diagonal and prefix-sink windows while complementary long-range windows partition the remaining history across heads, making full causal coverage a collective property of the head ensemble; in a window-matched 8K ablation CoWA reaches 89.73% versus 89.97% for FullAttn while duplicated long-range windows perform substantially worse, a 128K tensor-parallel operator benchmark reduces training forward and backward latency by 7.4x and 8.6x and decoding latency by 3.0x, and scaling-law training from 0.6B to 14B tracks FullAttn perplexity while cutting total training FLOPs by 28.5% in the 32K long-context stage.
The work introduces CoWindow Attention (CoWA), in which all KV heads share near-diagonal and prefix-sink windows while complementary long-range windows partition the remaining history across heads, making full causal coverage a collective property of the head ensemble; in a window-matched 8K ablation CoWA reaches 89.73% versus 89.97% for FullAttn while duplicated long-range windows perform substantially worse, a 128K tensor-parallel operator benchmark reduces training forward and backward latency by 7.4x and 8.6x and decoding latency by 3.0x, and scaling-law training from 0.6B to 14B tracks FullAttn perplexity while cutting total training FLOPs by 28.5% in the 32K long-context stage.
The work introduces CoWindow Attention (CoWA), in which all KV heads share near-diagonal and prefix-sink windows while complementary long-range windows partition the remaining history across heads, making full causal coverage a collective property of the head ensemble; in a window-matched 8K ablation CoWA reaches 89.73% versus 89.97% for FullAttn while duplicated long-range windows perform substantially worse, a 128K tensor-parallel operator benchmark reduces training forward and backward latency by 7.4x and 8.6x and decoding latency by 3.0x, and scaling-law training from 0.6B to 14B tracks FullAttn perplexity while cutting total training FLOPs by 28.5% in the 32K long-context stage.
The work introduces CoWindow Attention (CoWA), in which all KV heads share near-diagonal and prefix-sink windows while complementary long-range windows partition the remaining history across heads, making full causal coverage a collective property of the head ensemble; in a window-matched 8K ablation CoWA reaches 89.73% versus 89.97% for FullAttn while duplicated long-range windows perform substantially worse, a 128K tensor-parallel operator benchmark reduces training forward and backward latency by 7.4x and 8.6x and decoding latency by 3.0x, and scaling-law training from 0.6B to 14B tracks FullAttn perplexity while cutting total training FLOPs by 28.5% in the 32K long-context stage.
arXiv The study builds a curated resource of 1,813 matched American English–British English variant pairs and introduces DiAlign, a training-free method for estimating regional alignment, then uses them to audit six pretraining corpora, 21 post-training datasets, nine tokenizers, ten model checkpoints and generations under two prompting conditions, finding that American English is systematically favored across data exposure, representation and generation, and that British-English prompting only partially shifts this default.
The study builds a curated resource of 1,813 matched American English–British English variant pairs and introduces DiAlign, a training-free method for estimating regional alignment, then uses them to audit six pretraining corpora, 21 post-training datasets, nine tokenizers, ten model checkpoints and generations under two prompting conditions, finding that American English is systematically favored across data exposure, representation and generation, and that British-English prompting only partially shifts this default.
The study builds a curated resource of 1,813 matched American English–British English variant pairs and introduces DiAlign, a training-free method for estimating regional alignment, then uses them to audit six pretraining corpora, 21 post-training datasets, nine tokenizers, ten model checkpoints and generations under two prompting conditions, finding that American English is systematically favored across data exposure, representation and generation, and that British-English prompting only partially shifts this default.
The study builds a curated resource of 1,813 matched American English–British English variant pairs and introduces DiAlign, a training-free method for estimating regional alignment, then uses them to audit six pretraining corpora, 21 post-training datasets, nine tokenizers, ten model checkpoints and generations under two prompting conditions, finding that American English is systematically favored across data exposure, representation and generation, and that British-English prompting only partially shifts this default.
arXiv The work presents PReCache, a training-free KV cache sharing framework with two designs: PreLRShared precomputes a compact agent-specific low-rank cache for every agent when each segment of the shared context is first processed, removing later agents' repeated prefill over the accumulated trajectory, while ReBaseShared further reconstructs the shared base cache from adapter-free hidden states to reduce dependence on the previous agent, with lazy prefill for single-stream inference and double batching for concurrent serving; across LLaMA-3.
The work presents PReCache, a training-free KV cache sharing framework with two designs: PreLRShared precomputes a compact agent-specific low-rank cache for every agent when each segment of the shared context is first processed, removing later agents' repeated prefill over the accumulated trajectory, while ReBaseShared further reconstructs the shared base cache from adapter-free hidden states to reduce dependence on the previous agent, with lazy prefill for single-stream inference and double batching for concurrent serving; across LLaMA-3.
The work presents PReCache, a training-free KV cache sharing framework with two designs: PreLRShared precomputes a compact agent-specific low-rank cache for every agent when each segment of the shared context is first processed, removing later agents' repeated prefill over the accumulated trajectory, while ReBaseShared further reconstructs the shared base cache from adapter-free hidden states to reduce dependence on the previous agent, with lazy prefill for single-stream inference and double batching for concurrent serving; across LLaMA-3.
The work presents PReCache, a training-free KV cache sharing framework with two designs: PreLRShared precomputes a compact agent-specific low-rank cache for every agent when each segment of the shared context is first processed, removing later agents' repeated prefill over the accumulated trajectory, while ReBaseShared further reconstructs the shared base cache from adapter-free hidden states to reduce dependence on the previous agent, with lazy prefill for single-stream inference and double batching for concurrent serving; across LLaMA-3.
arXiv The work introduces UMM-Reflection, which cold-starts a unified multimodal model (BAGEL) with supervised fine-tuning on reflection trajectories and then applies reinforcement learning with a single trajectory-level advantage that updates both the reflection text and the flow-based image revisions, raising GenEval by 12.05 points over SFT and the conditional repair rate from 20.59% to 64.94%, with gains transferring to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none used in training.
The work introduces UMM-Reflection, which cold-starts a unified multimodal model (BAGEL) with supervised fine-tuning on reflection trajectories and then applies reinforcement learning with a single trajectory-level advantage that updates both the reflection text and the flow-based image revisions, raising GenEval by 12.05 points over SFT and the conditional repair rate from 20.59% to 64.94%, with gains transferring to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none used in training.
The work introduces UMM-Reflection, which cold-starts a unified multimodal model (BAGEL) with supervised fine-tuning on reflection trajectories and then applies reinforcement learning with a single trajectory-level advantage that updates both the reflection text and the flow-based image revisions, raising GenEval by 12.05 points over SFT and the conditional repair rate from 20.59% to 64.94%, with gains transferring to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none used in training.
The work introduces UMM-Reflection, which cold-starts a unified multimodal model (BAGEL) with supervised fine-tuning on reflection trajectories and then applies reinforcement learning with a single trajectory-level advantage that updates both the reflection text and the flow-based image revisions, raising GenEval by 12.05 points over SFT and the conditional repair rate from 20.59% to 64.94%, with gains transferring to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none used in training.
arXiv Through counterfactual interventions at intermediate reasoning states across two model families and multiple scales, the work finds that self-refinement largely consolidates probability mass onto solutions already reachable from the current state rather than making new ones reachable, distinguishes execution bottlenecks from knowledge bottlenecks, and introduces FlyBy, a selective querying framework that trains 4B and 8B models to reason first, diagnose what remains unresolved, and query stronger models at knowledge bottlenecks; on 1,158 hard problems across six benchmarks FlyBy-4B reaches 45.96% pass@8, surpassing Qwen3-14B's 41.64% at 2.7x lower serving cost, and FlyBy-8B raises pass@8 to 51.81%.
Through counterfactual interventions at intermediate reasoning states across two model families and multiple scales, the work finds that self-refinement largely consolidates probability mass onto solutions already reachable from the current state rather than making new ones reachable, distinguishes execution bottlenecks from knowledge bottlenecks, and introduces FlyBy, a selective querying framework that trains 4B and 8B models to reason first, diagnose what remains unresolved, and query stronger models at knowledge bottlenecks; on 1,158 hard problems across six benchmarks FlyBy-4B reaches 45.96% pass@8, surpassing Qwen3-14B's 41.64% at 2.7x lower serving cost, and FlyBy-8B raises pass@8 to 51.81%.
Through counterfactual interventions at intermediate reasoning states across two model families and multiple scales, the work finds that self-refinement largely consolidates probability mass onto solutions already reachable from the current state rather than making new ones reachable, distinguishes execution bottlenecks from knowledge bottlenecks, and introduces FlyBy, a selective querying framework that trains 4B and 8B models to reason first, diagnose what remains unresolved, and query stronger models at knowledge bottlenecks; on 1,158 hard problems across six benchmarks FlyBy-4B reaches 45.96% pass@8, surpassing Qwen3-14B's 41.64% at 2.7x lower serving cost, and FlyBy-8B raises pass@8 to 51.81%.
Through counterfactual interventions at intermediate reasoning states across two model families and multiple scales, the work finds that self-refinement largely consolidates probability mass onto solutions already reachable from the current state rather than making new ones reachable, distinguishes execution bottlenecks from knowledge bottlenecks, and introduces FlyBy, a selective querying framework that trains 4B and 8B models to reason first, diagnose what remains unresolved, and query stronger models at knowledge bottlenecks; on 1,158 hard problems across six benchmarks FlyBy-4B reaches 45.96% pass@8, surpassing Qwen3-14B's 41.64% at 2.7x lower serving cost, and FlyBy-8B raises pass@8 to 51.81%.
arXiv InfiniHand is an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video, trained in two progressive stages on roughly 5,000 hours of aggregated egocentric data; it reduces ARCTIC PA-p by 21.4% versus ViDiHand (9.82 to 7.72 mm), lowers EgoDex W-MPJPE from 78.41 to 28.77 mm versus Dyn-HaMR, and runs at 11.19 FPS, more than twice the throughput of HaWoR.
InfiniHand is an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video, trained in two progressive stages on roughly 5,000 hours of aggregated egocentric data; it reduces ARCTIC PA-p by 21.4% versus ViDiHand (9.82 to 7.72 mm), lowers EgoDex W-MPJPE from 78.41 to 28.77 mm versus Dyn-HaMR, and runs at 11.19 FPS, more than twice the throughput of HaWoR.
InfiniHand is an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video, trained in two progressive stages on roughly 5,000 hours of aggregated egocentric data; it reduces ARCTIC PA-p by 21.4% versus ViDiHand (9.82 to 7.72 mm), lowers EgoDex W-MPJPE from 78.41 to 28.77 mm versus Dyn-HaMR, and runs at 11.19 FPS, more than twice the throughput of HaWoR.
InfiniHand is an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video, trained in two progressive stages on roughly 5,000 hours of aggregated egocentric data; it reduces ARCTIC PA-p by 21.4% versus ViDiHand (9.82 to 7.72 mm), lowers EgoDex W-MPJPE from 78.41 to 28.77 mm versus Dyn-HaMR, and runs at 11.19 FPS, more than twice the throughput of HaWoR.
arXiv REALM is a coarse-to-fine framework for audio-driven reactive listening that fuses listener motion history with speaker audio through a Reactive Gated Speaker–Listener Fusion module using a delay-centered attention prior and adaptive gating, then predicts a base motion trajectory with a coarse decoder and augments it with audio-conditioned stochastic residuals in the expression subspace, improving over the evaluated baselines on multiple motion-quality metrics on ViCo and L2L and deploying to an Ameca humanoid robot with a perceptual user study.
REALM is a coarse-to-fine framework for audio-driven reactive listening that fuses listener motion history with speaker audio through a Reactive Gated Speaker–Listener Fusion module using a delay-centered attention prior and adaptive gating, then predicts a base motion trajectory with a coarse decoder and augments it with audio-conditioned stochastic residuals in the expression subspace, improving over the evaluated baselines on multiple motion-quality metrics on ViCo and L2L and deploying to an Ameca humanoid robot with a perceptual user study.
REALM is a coarse-to-fine framework for audio-driven reactive listening that fuses listener motion history with speaker audio through a Reactive Gated Speaker–Listener Fusion module using a delay-centered attention prior and adaptive gating, then predicts a base motion trajectory with a coarse decoder and augments it with audio-conditioned stochastic residuals in the expression subspace, improving over the evaluated baselines on multiple motion-quality metrics on ViCo and L2L and deploying to an Ameca humanoid robot with a perceptual user study.
REALM is a coarse-to-fine framework for audio-driven reactive listening that fuses listener motion history with speaker audio through a Reactive Gated Speaker–Listener Fusion module using a delay-centered attention prior and adaptive gating, then predicts a base motion trajectory with a coarse decoder and augments it with audio-conditioned stochastic residuals in the expression subspace, improving over the evaluated baselines on multiple motion-quality metrics on ViCo and L2L and deploying to an Ameca humanoid robot with a perceptual user study.
arXiv The work introduces WhiteMatter, which uses a learned router to mix past-token representations from any depth into shared key-value (KV) channels so every layer can draw on cross-layer information, together with cyclic Gauss–Seidel iteration to speed up training and prefill; under matched training-token budgets, half-cache WhiteMatter improves perplexity and average downstream scores over matched standard Transformers at two scales up to about 1.3B parameters, while the full-cache version performs comparably to a deeper standard Transformer.
The work introduces WhiteMatter, which uses a learned router to mix past-token representations from any depth into shared key-value (KV) channels so every layer can draw on cross-layer information, together with cyclic Gauss–Seidel iteration to speed up training and prefill; under matched training-token budgets, half-cache WhiteMatter improves perplexity and average downstream scores over matched standard Transformers at two scales up to about 1.3B parameters, while the full-cache version performs comparably to a deeper standard Transformer.
The work introduces WhiteMatter, which uses a learned router to mix past-token representations from any depth into shared key-value (KV) channels so every layer can draw on cross-layer information, together with cyclic Gauss–Seidel iteration to speed up training and prefill; under matched training-token budgets, half-cache WhiteMatter improves perplexity and average downstream scores over matched standard Transformers at two scales up to about 1.3B parameters, while the full-cache version performs comparably to a deeper standard Transformer.
The work introduces WhiteMatter, which uses a learned router to mix past-token representations from any depth into shared key-value (KV) channels so every layer can draw on cross-layer information, together with cyclic Gauss–Seidel iteration to speed up training and prefill; under matched training-token budgets, half-cache WhiteMatter improves perplexity and average downstream scores over matched standard Transformers at two scales up to about 1.3B parameters, while the full-cache version performs comparably to a deeper standard Transformer.
arXiv SentZero is a sentence-centric vision-language pretraining framework for chest X-ray: it uses an LLM to extract phrases from radiology reports and map them to concise topic-presence sentences that expand positive-pair diversity, adds an auxiliary loss that attracts the highest-similarity image patches toward clinical sentences recurring across studies to mitigate false negatives, and applies sentence-conditioned residual modulation to visual embeddings; after pretraining on MIMIC-CXR it outperforms prior multi-task zero-shot methods on most metrics for zero-shot multi-label classification on Open-I, ChestXray14, PadChest, ChestXDet10 and CheXpert and for zero-shot grounding on ChestXDet10 and MS-CXR.
SentZero is a sentence-centric vision-language pretraining framework for chest X-ray: it uses an LLM to extract phrases from radiology reports and map them to concise topic-presence sentences that expand positive-pair diversity, adds an auxiliary loss that attracts the highest-similarity image patches toward clinical sentences recurring across studies to mitigate false negatives, and applies sentence-conditioned residual modulation to visual embeddings; after pretraining on MIMIC-CXR it outperforms prior multi-task zero-shot methods on most metrics for zero-shot multi-label classification on Open-I, ChestXray14, PadChest, ChestXDet10 and CheXpert and for zero-shot grounding on ChestXDet10 and MS-CXR.
SentZero is a sentence-centric vision-language pretraining framework for chest X-ray: it uses an LLM to extract phrases from radiology reports and map them to concise topic-presence sentences that expand positive-pair diversity, adds an auxiliary loss that attracts the highest-similarity image patches toward clinical sentences recurring across studies to mitigate false negatives, and applies sentence-conditioned residual modulation to visual embeddings; after pretraining on MIMIC-CXR it outperforms prior multi-task zero-shot methods on most metrics for zero-shot multi-label classification on Open-I, ChestXray14, PadChest, ChestXDet10 and CheXpert and for zero-shot grounding on ChestXDet10 and MS-CXR.
SentZero is a sentence-centric vision-language pretraining framework for chest X-ray: it uses an LLM to extract phrases from radiology reports and map them to concise topic-presence sentences that expand positive-pair diversity, adds an auxiliary loss that attracts the highest-similarity image patches toward clinical sentences recurring across studies to mitigate false negatives, and applies sentence-conditioned residual modulation to visual embeddings; after pretraining on MIMIC-CXR it outperforms prior multi-task zero-shot methods on most metrics for zero-shot multi-label classification on Open-I, ChestXray14, PadChest, ChestXDet10 and CheXpert and for zero-shot grounding on ChestXDet10 and MS-CXR.
arXiv The work proposes a human-annotation-free synthetic supervision pipeline: it deeply rewrites 3,515 public documents via entity renaming and numeric perturbation, has LLMs generate questions and rubrics that require reasoning over the document, and admits only samples that genuinely depend on the document via a with-document versus without-document rubric gap check, yielding about 9,625 training samples; SFT raises Qwen3.6-35B-A3B on CL-bench from 13.7% to 22.8%, a subsequent rubric-reward RL stage reaches 24.6%, comparable to Qwen3.8-2.4T at 23.9%, with broad transfer to long-context understanding, instruction following, and reasoning, while code generation and knowledge stay mostly flat.
The work proposes a human-annotation-free synthetic supervision pipeline: it deeply rewrites 3,515 public documents via entity renaming and numeric perturbation, has LLMs generate questions and rubrics that require reasoning over the document, and admits only samples that genuinely depend on the document via a with-document versus without-document rubric gap check, yielding about 9,625 training samples; SFT raises Qwen3.6-35B-A3B on CL-bench from 13.7% to 22.8%, a subsequent rubric-reward RL stage reaches 24.6%, comparable to Qwen3.8-2.4T at 23.9%, with broad transfer to long-context understanding, instruction following, and reasoning, while code generation and knowledge stay mostly flat.
The work proposes a human-annotation-free synthetic supervision pipeline: it deeply rewrites 3,515 public documents via entity renaming and numeric perturbation, has LLMs generate questions and rubrics that require reasoning over the document, and admits only samples that genuinely depend on the document via a with-document versus without-document rubric gap check, yielding about 9,625 training samples; SFT raises Qwen3.6-35B-A3B on CL-bench from 13.7% to 22.8%, a subsequent rubric-reward RL stage reaches 24.6%, comparable to Qwen3.8-2.4T at 23.9%, with broad transfer to long-context understanding, instruction following, and reasoning, while code generation and knowledge stay mostly flat.
The work proposes a human-annotation-free synthetic supervision pipeline: it deeply rewrites 3,515 public documents via entity renaming and numeric perturbation, has LLMs generate questions and rubrics that require reasoning over the document, and admits only samples that genuinely depend on the document via a with-document versus without-document rubric gap check, yielding about 9,625 training samples; SFT raises Qwen3.6-35B-A3B on CL-bench from 13.7% to 22.8%, a subsequent rubric-reward RL stage reaches 24.6%, comparable to Qwen3.8-2.4T at 23.9%, with broad transfer to long-context understanding, instruction following, and reasoning, while code generation and knowledge stay mostly flat.
arXiv VGGT-Diff is a geometry-routed multi-view diffusion framework that uses a frozen VGGT-Ω to transform source-view features together with 3D points and confidence into query-aligned conditions via a confidence-aware Visual Geometry Router, injects them into a pretrained Wan2.1-I2V-14B video diffusion model, and regularizes predicted-clean residuals along reliable 3D tracks with Point-Track Residual Consistency; trained on only 980 DL3DV scenes, it ranks first in PSNR, LPIPS and DreamSim and second in SSIM on 6,188 DL3DV-Benchmark targets, and first in LPIPS and DreamSim with second-place PSNR on zero-shot Mip-NeRF 360.
VGGT-Diff is a geometry-routed multi-view diffusion framework that uses a frozen VGGT-Ω to transform source-view features together with 3D points and confidence into query-aligned conditions via a confidence-aware Visual Geometry Router, injects them into a pretrained Wan2.1-I2V-14B video diffusion model, and regularizes predicted-clean residuals along reliable 3D tracks with Point-Track Residual Consistency; trained on only 980 DL3DV scenes, it ranks first in PSNR, LPIPS and DreamSim and second in SSIM on 6,188 DL3DV-Benchmark targets, and first in LPIPS and DreamSim with second-place PSNR on zero-shot Mip-NeRF 360.
VGGT-Diff is a geometry-routed multi-view diffusion framework that uses a frozen VGGT-Ω to transform source-view features together with 3D points and confidence into query-aligned conditions via a confidence-aware Visual Geometry Router, injects them into a pretrained Wan2.1-I2V-14B video diffusion model, and regularizes predicted-clean residuals along reliable 3D tracks with Point-Track Residual Consistency; trained on only 980 DL3DV scenes, it ranks first in PSNR, LPIPS and DreamSim and second in SSIM on 6,188 DL3DV-Benchmark targets, and first in LPIPS and DreamSim with second-place PSNR on zero-shot Mip-NeRF 360.
VGGT-Diff is a geometry-routed multi-view diffusion framework that uses a frozen VGGT-Ω to transform source-view features together with 3D points and confidence into query-aligned conditions via a confidence-aware Visual Geometry Router, injects them into a pretrained Wan2.1-I2V-14B video diffusion model, and regularizes predicted-clean residuals along reliable 3D tracks with Point-Track Residual Consistency; trained on only 980 DL3DV scenes, it ranks first in PSNR, LPIPS and DreamSim and second in SSIM on 6,188 DL3DV-Benchmark targets, and first in LPIPS and DreamSim with second-place PSNR on zero-shot Mip-NeRF 360.
arXiv QwenGyre is an end-to-end framework for extreme-long-horizon (xLong) online agent RL that elastically reallocates GPUs between rollout and training without interrupting live harness executions, while its trajectory processor reconstructs branching executions into shared-prefix trajectory trees, scores partial progress after timeouts, and deduplicates sampled trajectories by role priority; scaled to Qwen 3.8 2.4T with 700K tokens per rollout, it raises NL2RepoBench pass rate from 52.5% to 58.5% in 48 steps and delivers up to 1.85x and 1.78x end-to-end speedups over Colocate and Async.
QwenGyre is an end-to-end framework for extreme-long-horizon (xLong) online agent RL that elastically reallocates GPUs between rollout and training without interrupting live harness executions, while its trajectory processor reconstructs branching executions into shared-prefix trajectory trees, scores partial progress after timeouts, and deduplicates sampled trajectories by role priority; scaled to Qwen 3.8 2.4T with 700K tokens per rollout, it raises NL2RepoBench pass rate from 52.5% to 58.5% in 48 steps and delivers up to 1.85x and 1.78x end-to-end speedups over Colocate and Async.
QwenGyre is an end-to-end framework for extreme-long-horizon (xLong) online agent RL that elastically reallocates GPUs between rollout and training without interrupting live harness executions, while its trajectory processor reconstructs branching executions into shared-prefix trajectory trees, scores partial progress after timeouts, and deduplicates sampled trajectories by role priority; scaled to Qwen 3.8 2.4T with 700K tokens per rollout, it raises NL2RepoBench pass rate from 52.5% to 58.5% in 48 steps and delivers up to 1.85x and 1.78x end-to-end speedups over Colocate and Async.
QwenGyre is an end-to-end framework for extreme-long-horizon (xLong) online agent RL that elastically reallocates GPUs between rollout and training without interrupting live harness executions, while its trajectory processor reconstructs branching executions into shared-prefix trajectory trees, scores partial progress after timeouts, and deduplicates sampled trajectories by role priority; scaled to Qwen 3.8 2.4T with 700K tokens per rollout, it raises NL2RepoBench pass rate from 52.5% to 58.5% in 48 steps and delivers up to 1.85x and 1.78x end-to-end speedups over Colocate and Async.
Google AI 与 Gemini 产品博客 Google describes its workhorse model Gemini 3.8 Flash as delivering significant improvements over 3.7 Flash in software engineering, agentic tasks, and multistep reasoning in specialized domains, and showcases four builder projects: live path mapping of satellites, space stations, and orbital rockets; turning Seigaiha waves into a moving ink painting; a T. rex skeleton built with a four-phase prompt and accuracy checks; and an interactive automatic transmission model with 10 camera views and four display modes.
Google describes its workhorse model Gemini 3.8 Flash as delivering significant improvements over 3.7 Flash in software engineering, agentic tasks, and multistep reasoning in specialized domains, and showcases four builder projects: live path mapping of satellites, space stations, and orbital rockets; turning Seigaiha waves into a moving ink painting; a T. rex skeleton built with a four-phase prompt and accuracy checks; and an interactive automatic transmission model with 10 camera views and four display modes.
Google describes its workhorse model Gemini 3.8 Flash as delivering significant improvements over 3.7 Flash in software engineering, agentic tasks, and multistep reasoning in specialized domains, and showcases four builder projects: live path mapping of satellites, space stations, and orbital rockets; turning Seigaiha waves into a moving ink painting; a T. rex skeleton built with a four-phase prompt and accuracy checks; and an interactive automatic transmission model with 10 camera views and four display modes.
Google describes its workhorse model Gemini 3.8 Flash as delivering significant improvements over 3.7 Flash in software engineering, agentic tasks, and multistep reasoning in specialized domains, and showcases four builder projects: live path mapping of satellites, space stations, and orbital rockets; turning Seigaiha waves into a moving ink painting; a T. rex skeleton built with a four-phase prompt and accuracy checks; and an interactive automatic transmission model with 10 camera views and four display modes.
arXiv The work builds SciGen-Verify, a benchmark for explainable verification of scientific image generation (1,350 samples: 406 instruction following, 287 multidisciplinary reasoning, 657 world knowledge, with a three-tier cascading protocol over binary judgement, explanation, and editing instruction), and develops SciGen-Verifier on Qwen3-VL-8B-Instruct trained by cold-start supervised fine-tuning (48,311 instances) followed by a curriculum-based two-stage reinforcement learning pipeline (rubric-guided process reward, then outcome reward), reaching 86.28% overall judgement accuracy, 81.79% explanation accuracy, and 57.42% editing-instruction accuracy on the benchmark while also serving as an online critic for iterative image rectification.
The work builds SciGen-Verify, a benchmark for explainable verification of scientific image generation (1,350 samples: 406 instruction following, 287 multidisciplinary reasoning, 657 world knowledge, with a three-tier cascading protocol over binary judgement, explanation, and editing instruction), and develops SciGen-Verifier on Qwen3-VL-8B-Instruct trained by cold-start supervised fine-tuning (48,311 instances) followed by a curriculum-based two-stage reinforcement learning pipeline (rubric-guided process reward, then outcome reward), reaching 86.28% overall judgement accuracy, 81.79% explanation accuracy, and 57.42% editing-instruction accuracy on the benchmark while also serving as an online critic for iterative image rectification.
The work builds SciGen-Verify, a benchmark for explainable verification of scientific image generation (1,350 samples: 406 instruction following, 287 multidisciplinary reasoning, 657 world knowledge, with a three-tier cascading protocol over binary judgement, explanation, and editing instruction), and develops SciGen-Verifier on Qwen3-VL-8B-Instruct trained by cold-start supervised fine-tuning (48,311 instances) followed by a curriculum-based two-stage reinforcement learning pipeline (rubric-guided process reward, then outcome reward), reaching 86.28% overall judgement accuracy, 81.79% explanation accuracy, and 57.42% editing-instruction accuracy on the benchmark while also serving as an online critic for iterative image rectification.
The work builds SciGen-Verify, a benchmark for explainable verification of scientific image generation (1,350 samples: 406 instruction following, 287 multidisciplinary reasoning, 657 world knowledge, with a three-tier cascading protocol over binary judgement, explanation, and editing instruction), and develops SciGen-Verifier on Qwen3-VL-8B-Instruct trained by cold-start supervised fine-tuning (48,311 instances) followed by a curriculum-based two-stage reinforcement learning pipeline (rubric-guided process reward, then outcome reward), reaching 86.28% overall judgement accuracy, 81.79% explanation accuracy, and 57.42% editing-instruction accuracy on the benchmark while also serving as an online critic for iterative image rectification.
arXiv The work proposes FactorEngram, which replaces monolithic per-pattern n-gram embeddings with sparsity-regularized coefficients over a dictionary shared across patterns and reuses that same dictionary for basis-level contextual gating, improving language modeling, downstream tasks, and long-context retrieval on 340M- and 1B-parameter Transformer backbones.
The work proposes FactorEngram, which replaces monolithic per-pattern n-gram embeddings with sparsity-regularized coefficients over a dictionary shared across patterns and reuses that same dictionary for basis-level contextual gating, improving language modeling, downstream tasks, and long-context retrieval on 340M- and 1B-parameter Transformer backbones.
The work proposes FactorEngram, which replaces monolithic per-pattern n-gram embeddings with sparsity-regularized coefficients over a dictionary shared across patterns and reuses that same dictionary for basis-level contextual gating, improving language modeling, downstream tasks, and long-context retrieval on 340M- and 1B-parameter Transformer backbones.
The work proposes FactorEngram, which replaces monolithic per-pattern n-gram embeddings with sparsity-regularized coefficients over a dictionary shared across patterns and reuses that same dictionary for basis-level contextual gating, improving language modeling, downstream tasks, and long-context retrieval on 340M- and 1B-parameter Transformer backbones.
arXiv This work revisits whether reverse KL divergence is necessary for on-policy distillation (OPD), proposes BinaryOPD, which keeps only the update direction as a binary reward and matches or slightly exceeds OPD across multiple student-teacher pairs on math and code, further finds that only the direction of a small subset of high-disagreement tokens is critical, and builds on this to propose C-MOPD, in which every sample is supervised by all teachers and which consistently outperforms MOPD on math and code benchmarks.
This work revisits whether reverse KL divergence is necessary for on-policy distillation (OPD), proposes BinaryOPD, which keeps only the update direction as a binary reward and matches or slightly exceeds OPD across multiple student-teacher pairs on math and code, further finds that only the direction of a small subset of high-disagreement tokens is critical, and builds on this to propose C-MOPD, in which every sample is supervised by all teachers and which consistently outperforms MOPD on math and code benchmarks.
This work revisits whether reverse KL divergence is necessary for on-policy distillation (OPD), proposes BinaryOPD, which keeps only the update direction as a binary reward and matches or slightly exceeds OPD across multiple student-teacher pairs on math and code, further finds that only the direction of a small subset of high-disagreement tokens is critical, and builds on this to propose C-MOPD, in which every sample is supervised by all teachers and which consistently outperforms MOPD on math and code benchmarks.
This work revisits whether reverse KL divergence is necessary for on-policy distillation (OPD), proposes BinaryOPD, which keeps only the update direction as a binary reward and matches or slightly exceeds OPD across multiple student-teacher pairs on math and code, further finds that only the direction of a small subset of high-disagreement tokens is critical, and builds on this to propose C-MOPD, in which every sample is supervised by all teachers and which consistently outperforms MOPD on math and code benchmarks.
arXiv The work introduces WaveFront Decoding (WFD), a training-free self-speculative decoding framework for looped language models that exploits two properties of these architectures, namely that intermediate recurrence outputs provide effective draft predictions and that weight sharing lets token states at different positions and recurrence depths be processed in one batched recurrent-block call, organizing these mixed-depth states into a diagonal wavefront so that drafting and verification proceed concurrently within the same recurrent calls while rejected drafts are corrected using full-depth predictions; across six Spec-Bench task categories it achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, rising to 4.81x on Huginn-3.
The work introduces WaveFront Decoding (WFD), a training-free self-speculative decoding framework for looped language models that exploits two properties of these architectures, namely that intermediate recurrence outputs provide effective draft predictions and that weight sharing lets token states at different positions and recurrence depths be processed in one batched recurrent-block call, organizing these mixed-depth states into a diagonal wavefront so that drafting and verification proceed concurrently within the same recurrent calls while rejected drafts are corrected using full-depth predictions; across six Spec-Bench task categories it achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, rising to 4.81x on Huginn-3.
The work introduces WaveFront Decoding (WFD), a training-free self-speculative decoding framework for looped language models that exploits two properties of these architectures, namely that intermediate recurrence outputs provide effective draft predictions and that weight sharing lets token states at different positions and recurrence depths be processed in one batched recurrent-block call, organizing these mixed-depth states into a diagonal wavefront so that drafting and verification proceed concurrently within the same recurrent calls while rejected drafts are corrected using full-depth predictions; across six Spec-Bench task categories it achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, rising to 4.81x on Huginn-3.
The work introduces WaveFront Decoding (WFD), a training-free self-speculative decoding framework for looped language models that exploits two properties of these architectures, namely that intermediate recurrence outputs provide effective draft predictions and that weight sharing lets token states at different positions and recurrence depths be processed in one batched recurrent-block call, organizing these mixed-depth states into a diagonal wavefront so that drafting and verification proceed concurrently within the same recurrent calls while rejected drafts are corrected using full-depth predictions; across six Spec-Bench task categories it achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, rising to 4.81x on Huginn-3.