arXiv The work introduces PlaylistEval, an agentic framework that builds video-language judge benchmarks without human annotation by automatically generating questions whose evidence is scattered across two distant segments of ~100-hour playlists and whose wrong answers differ only visually, yielding PlaylistBench with 630 pairs across seven domains; on a stratified subset of 152 pairs it agrees with human judgments 93.0% of the time (IAA 0.781), while evaluating 17 omnimodal and multimodal models from eight families shows frontier judges reach only 75.4% pairwise accuracy, open-source judge models perform far behind, retrieval improves accuracy by up to 10.5 points yet the best retriever finds both relevant segments in the top 10 only 37.
The work introduces PlaylistEval, an agentic framework that builds video-language judge benchmarks without human annotation by automatically generating questions whose evidence is scattered across two distant segments of ~100-hour playlists and whose wrong answers differ only visually, yielding PlaylistBench with 630 pairs across seven domains; on a stratified subset of 152 pairs it agrees with human judgments 93.0% of the time (IAA 0.781), while evaluating 17 omnimodal and multimodal models from eight families shows frontier judges reach only 75.4% pairwise accuracy, open-source judge models perform far behind, retrieval improves accuracy by up to 10.5 points yet the best retriever finds both relevant segments in the top 10 only 37.
The work introduces PlaylistEval, an agentic framework that builds video-language judge benchmarks without human annotation by automatically generating questions whose evidence is scattered across two distant segments of ~100-hour playlists and whose wrong answers differ only visually, yielding PlaylistBench with 630 pairs across seven domains; on a stratified subset of 152 pairs it agrees with human judgments 93.0% of the time (IAA 0.781), while evaluating 17 omnimodal and multimodal models from eight families shows frontier judges reach only 75.4% pairwise accuracy, open-source judge models perform far behind, retrieval improves accuracy by up to 10.5 points yet the best retriever finds both relevant segments in the top 10 only 37.
The work introduces PlaylistEval, an agentic framework that builds video-language judge benchmarks without human annotation by automatically generating questions whose evidence is scattered across two distant segments of ~100-hour playlists and whose wrong answers differ only visually, yielding PlaylistBench with 630 pairs across seven domains; on a stratified subset of 152 pairs it agrees with human judgments 93.0% of the time (IAA 0.781), while evaluating 17 omnimodal and multimodal models from eight families shows frontier judges reach only 75.4% pairwise accuracy, open-source judge models perform far behind, retrieval improves accuracy by up to 10.5 points yet the best retriever finds both relevant segments in the top 10 only 37.
arXiv AutoRef has a coding agent iteratively rewrite harness code while keeping both the image generator and the reasoning model frozen, separating the tasks whose feedback informs proposals from those used to select candidates and continuing the search from a beam of top-ranked harnesses; the resulting AutoRef-Harness raises open-weight FLUX.2 [klein] 4B from 5.72 to 7.37 (out of 10) on the held-out four-reference MultiBanana test split, matching or exceeding proprietary models such as Nano Banana Pro and GPT-Image-1.5, and transfers without re-optimization across reference counts, benchmarks and evaluators, generators, and reasoning models.
AutoRef has a coding agent iteratively rewrite harness code while keeping both the image generator and the reasoning model frozen, separating the tasks whose feedback informs proposals from those used to select candidates and continuing the search from a beam of top-ranked harnesses; the resulting AutoRef-Harness raises open-weight FLUX.2 [klein] 4B from 5.72 to 7.37 (out of 10) on the held-out four-reference MultiBanana test split, matching or exceeding proprietary models such as Nano Banana Pro and GPT-Image-1.5, and transfers without re-optimization across reference counts, benchmarks and evaluators, generators, and reasoning models.
AutoRef has a coding agent iteratively rewrite harness code while keeping both the image generator and the reasoning model frozen, separating the tasks whose feedback informs proposals from those used to select candidates and continuing the search from a beam of top-ranked harnesses; the resulting AutoRef-Harness raises open-weight FLUX.2 [klein] 4B from 5.72 to 7.37 (out of 10) on the held-out four-reference MultiBanana test split, matching or exceeding proprietary models such as Nano Banana Pro and GPT-Image-1.5, and transfers without re-optimization across reference counts, benchmarks and evaluators, generators, and reasoning models.
AutoRef has a coding agent iteratively rewrite harness code while keeping both the image generator and the reasoning model frozen, separating the tasks whose feedback informs proposals from those used to select candidates and continuing the search from a beam of top-ranked harnesses; the resulting AutoRef-Harness raises open-weight FLUX.2 [klein] 4B from 5.72 to 7.37 (out of 10) on the held-out four-reference MultiBanana test split, matching or exceeding proprietary models such as Nano Banana Pro and GPT-Image-1.5, and transfers without re-optimization across reference counts, benchmarks and evaluators, generators, and reasoning models.
arXiv The authors present ALICE, a foundation model trained once on a synthetic distribution family that turns mutual information estimation into in-context estimation of rectified-flow velocity fields for unseen distributions, matching per-distribution neural estimators on the Beyond Normal benchmark with at least twofold lower error at a 1k-sample budget and zero-shot reproducing and extending dedicated estimators' findings on biology, genetics, and neuroscience data the model never saw.
The authors present ALICE, a foundation model trained once on a synthetic distribution family that turns mutual information estimation into in-context estimation of rectified-flow velocity fields for unseen distributions, matching per-distribution neural estimators on the Beyond Normal benchmark with at least twofold lower error at a 1k-sample budget and zero-shot reproducing and extending dedicated estimators' findings on biology, genetics, and neuroscience data the model never saw.
The authors present ALICE, a foundation model trained once on a synthetic distribution family that turns mutual information estimation into in-context estimation of rectified-flow velocity fields for unseen distributions, matching per-distribution neural estimators on the Beyond Normal benchmark with at least twofold lower error at a 1k-sample budget and zero-shot reproducing and extending dedicated estimators' findings on biology, genetics, and neuroscience data the model never saw.
The authors present ALICE, a foundation model trained once on a synthetic distribution family that turns mutual information estimation into in-context estimation of rectified-flow velocity fields for unseen distributions, matching per-distribution neural estimators on the Beyond Normal benchmark with at least twofold lower error at a 1k-sample budget and zero-shot reproducing and extending dedicated estimators' findings on biology, genetics, and neuroscience data the model never saw.
arXiv Organized around how contextual evidence connects to execution, this survey sorts the robot in-context learning (ICL) literature into four families — context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution — compares their transfer assumptions and the roles of training, correspondence, and memory across manipulation and navigation, and proposes evaluation practices that separate responsiveness to teaching, physical transfer, and benefits from retained experience, together with an agenda linking compositional task acquisition, faithful transfer, and physical recursive self-improvement.
Organized around how contextual evidence connects to execution, this survey sorts the robot in-context learning (ICL) literature into four families — context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution — compares their transfer assumptions and the roles of training, correspondence, and memory across manipulation and navigation, and proposes evaluation practices that separate responsiveness to teaching, physical transfer, and benefits from retained experience, together with an agenda linking compositional task acquisition, faithful transfer, and physical recursive self-improvement.
Organized around how contextual evidence connects to execution, this survey sorts the robot in-context learning (ICL) literature into four families — context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution — compares their transfer assumptions and the roles of training, correspondence, and memory across manipulation and navigation, and proposes evaluation practices that separate responsiveness to teaching, physical transfer, and benefits from retained experience, together with an agenda linking compositional task acquisition, faithful transfer, and physical recursive self-improvement.
Organized around how contextual evidence connects to execution, this survey sorts the robot in-context learning (ICL) literature into four families — context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution — compares their transfer assumptions and the roles of training, correspondence, and memory across manipulation and navigation, and proposes evaluation practices that separate responsiveness to teaching, physical transfer, and benefits from retained experience, together with an agenda linking compositional task acquisition, faithful transfer, and physical recursive self-improvement.
arXiv The work formulates attack and defense for tool-using language-model agents as partially observed state control (SEAD), derives from it an execution-feedback tree-search attacker (DART) and a pre-execution defender that can run read-only state investigations (SAGE), and reports across 187 malicious tasks and four target models that DART exceeds baselines by 18.8–35.9 percentage points under both semantic and executable criteria, while SAGE preserves 95.79% of benign trajectories, intercepts 92.73% of harmful paths on recorded trajectories, and reduces DART's executable attack success from 48.0% to 4.0% online.
The work formulates attack and defense for tool-using language-model agents as partially observed state control (SEAD), derives from it an execution-feedback tree-search attacker (DART) and a pre-execution defender that can run read-only state investigations (SAGE), and reports across 187 malicious tasks and four target models that DART exceeds baselines by 18.8–35.9 percentage points under both semantic and executable criteria, while SAGE preserves 95.79% of benign trajectories, intercepts 92.73% of harmful paths on recorded trajectories, and reduces DART's executable attack success from 48.0% to 4.0% online.
The work formulates attack and defense for tool-using language-model agents as partially observed state control (SEAD), derives from it an execution-feedback tree-search attacker (DART) and a pre-execution defender that can run read-only state investigations (SAGE), and reports across 187 malicious tasks and four target models that DART exceeds baselines by 18.8–35.9 percentage points under both semantic and executable criteria, while SAGE preserves 95.79% of benign trajectories, intercepts 92.73% of harmful paths on recorded trajectories, and reduces DART's executable attack success from 48.0% to 4.0% online.
The work formulates attack and defense for tool-using language-model agents as partially observed state control (SEAD), derives from it an execution-feedback tree-search attacker (DART) and a pre-execution defender that can run read-only state investigations (SAGE), and reports across 187 malicious tasks and four target models that DART exceeds baselines by 18.8–35.9 percentage points under both semantic and executable criteria, while SAGE preserves 95.79% of benign trajectories, intercepts 92.73% of harmful paths on recorded trajectories, and reduces DART's executable attack success from 48.0% to 4.0% online.
arXiv By keeping the original diffusion or flow-matching objective and adding an adversarial loss to the predicted output at non-high-noise timesteps, this study jointly improves distribution fidelity, coverage, prompt alignment, and perceptual quality on the DeCo and PixelGen pixel backbones, attributing the gain to restored missing natural-image high-frequency power, while the same procedure yields no comparable gain in the tested latent diffusion configurations.
By keeping the original diffusion or flow-matching objective and adding an adversarial loss to the predicted output at non-high-noise timesteps, this study jointly improves distribution fidelity, coverage, prompt alignment, and perceptual quality on the DeCo and PixelGen pixel backbones, attributing the gain to restored missing natural-image high-frequency power, while the same procedure yields no comparable gain in the tested latent diffusion configurations.
By keeping the original diffusion or flow-matching objective and adding an adversarial loss to the predicted output at non-high-noise timesteps, this study jointly improves distribution fidelity, coverage, prompt alignment, and perceptual quality on the DeCo and PixelGen pixel backbones, attributing the gain to restored missing natural-image high-frequency power, while the same procedure yields no comparable gain in the tested latent diffusion configurations.
By keeping the original diffusion or flow-matching objective and adding an adversarial loss to the predicted output at non-high-noise timesteps, this study jointly improves distribution fidelity, coverage, prompt alignment, and perceptual quality on the DeCo and PixelGen pixel backbones, attributing the gain to restored missing natural-image high-frequency power, while the same procedure yields no comparable gain in the tested latent diffusion configurations.
arXiv The work introduces Context Language Models (CLMs), which mirror a model's live context as a freely editable file that the model can rewrite with Bash, shifting context management from external harness control to an intrinsic model capability; applied zero-shot to existing models such as Qwen3.6-27B and GPT5.6-Sol, CLMs achieve 11.4% higher accuracy with 21.5% fewer prefix-reuse FLOPs than the strongest baseline on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater end-to-end speedup at the same compute on a 24-hour multi-repository agent-swarm task, and can further learn context-management strategies through natural-language instructions, textual evolution, and online reinforcement learning.
The work introduces Context Language Models (CLMs), which mirror a model's live context as a freely editable file that the model can rewrite with Bash, shifting context management from external harness control to an intrinsic model capability; applied zero-shot to existing models such as Qwen3.6-27B and GPT5.6-Sol, CLMs achieve 11.4% higher accuracy with 21.5% fewer prefix-reuse FLOPs than the strongest baseline on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater end-to-end speedup at the same compute on a 24-hour multi-repository agent-swarm task, and can further learn context-management strategies through natural-language instructions, textual evolution, and online reinforcement learning.
The work introduces Context Language Models (CLMs), which mirror a model's live context as a freely editable file that the model can rewrite with Bash, shifting context management from external harness control to an intrinsic model capability; applied zero-shot to existing models such as Qwen3.6-27B and GPT5.6-Sol, CLMs achieve 11.4% higher accuracy with 21.5% fewer prefix-reuse FLOPs than the strongest baseline on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater end-to-end speedup at the same compute on a 24-hour multi-repository agent-swarm task, and can further learn context-management strategies through natural-language instructions, textual evolution, and online reinforcement learning.
The work introduces Context Language Models (CLMs), which mirror a model's live context as a freely editable file that the model can rewrite with Bash, shifting context management from external harness control to an intrinsic model capability; applied zero-shot to existing models such as Qwen3.6-27B and GPT5.6-Sol, CLMs achieve 11.4% higher accuracy with 21.5% fewer prefix-reuse FLOPs than the strongest baseline on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater end-to-end speedup at the same compute on a 24-hour multi-repository agent-swarm task, and can further learn context-management strategies through natural-language instructions, textual evolution, and online reinforcement learning.
arXiv PreviewDiff introduces a training-free test-time search method that decodes partial previews at selected denoising checkpoints, asks a multimodal judge to score and critique them, and uses the resulting natural-language feedback to branch over semantic prompt edits and locally re-noised latent continuations that are scored and selectively rolled forward, consistently outperforming budget-matched Best-of-N and scalar-search baselines on image and video generation benchmarks.
PreviewDiff introduces a training-free test-time search method that decodes partial previews at selected denoising checkpoints, asks a multimodal judge to score and critique them, and uses the resulting natural-language feedback to branch over semantic prompt edits and locally re-noised latent continuations that are scored and selectively rolled forward, consistently outperforming budget-matched Best-of-N and scalar-search baselines on image and video generation benchmarks.
PreviewDiff introduces a training-free test-time search method that decodes partial previews at selected denoising checkpoints, asks a multimodal judge to score and critique them, and uses the resulting natural-language feedback to branch over semantic prompt edits and locally re-noised latent continuations that are scored and selectively rolled forward, consistently outperforming budget-matched Best-of-N and scalar-search baselines on image and video generation benchmarks.
PreviewDiff introduces a training-free test-time search method that decodes partial previews at selected denoising checkpoints, asks a multimodal judge to score and critique them, and uses the resulting natural-language feedback to branch over semantic prompt edits and locally re-noised latent continuations that are scored and selectively rolled forward, consistently outperforming budget-matched Best-of-N and scalar-search baselines on image and video generation benchmarks.
arXiv The work proposes EmoRES, a training-free method that discovers that an emotion steering vector decomposes into a shared component moving speech away from neutral expression and a residual component directing generation toward the requested emotion, and controls the two independently; on IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the frozen IndexTTS-2 and CosyVoice2 backbones, improving rank correlation by 26.13 and 12.97 percentage points and emotion hit rate by 12.95 and 6.92 points, with human evaluation showing up to 35.0% relative improvement in listeners correctly identifying the dominant requested emotion and up to 63.8% naturalness preference in pairwise comparisons.
The work proposes EmoRES, a training-free method that discovers that an emotion steering vector decomposes into a shared component moving speech away from neutral expression and a residual component directing generation toward the requested emotion, and controls the two independently; on IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the frozen IndexTTS-2 and CosyVoice2 backbones, improving rank correlation by 26.13 and 12.97 percentage points and emotion hit rate by 12.95 and 6.92 points, with human evaluation showing up to 35.0% relative improvement in listeners correctly identifying the dominant requested emotion and up to 63.8% naturalness preference in pairwise comparisons.
The work proposes EmoRES, a training-free method that discovers that an emotion steering vector decomposes into a shared component moving speech away from neutral expression and a residual component directing generation toward the requested emotion, and controls the two independently; on IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the frozen IndexTTS-2 and CosyVoice2 backbones, improving rank correlation by 26.13 and 12.97 percentage points and emotion hit rate by 12.95 and 6.92 points, with human evaluation showing up to 35.0% relative improvement in listeners correctly identifying the dominant requested emotion and up to 63.8% naturalness preference in pairwise comparisons.
The work proposes EmoRES, a training-free method that discovers that an emotion steering vector decomposes into a shared component moving speech away from neutral expression and a residual component directing generation toward the requested emotion, and controls the two independently; on IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the frozen IndexTTS-2 and CosyVoice2 backbones, improving rank correlation by 26.13 and 12.97 percentage points and emotion hit rate by 12.95 and 6.92 points, with human evaluation showing up to 35.0% relative improvement in listeners correctly identifying the dominant requested emotion and up to 63.8% naturalness preference in pairwise comparisons.
arXiv Through controlled pretraining experiments, this study systematically examines when recurrence helps, where it should be applied, and how it should be conditioned in looped language models (LoopLMs), finding that extra loops beyond the training horizon improve reasoning (e.g., a reasoning score rising from 28.52 to 31.49) while degrading knowledge (from 62.80 to 52.11), that non-recurrent output layers improve robustness to under-unrolling, and that the proposed history-state injection (especially channel-wise) combined with timestep conditioning better preserves knowledge under extended unrolling and improves robustness across inference budgets.
Through controlled pretraining experiments, this study systematically examines when recurrence helps, where it should be applied, and how it should be conditioned in looped language models (LoopLMs), finding that extra loops beyond the training horizon improve reasoning (e.g., a reasoning score rising from 28.52 to 31.49) while degrading knowledge (from 62.80 to 52.11), that non-recurrent output layers improve robustness to under-unrolling, and that the proposed history-state injection (especially channel-wise) combined with timestep conditioning better preserves knowledge under extended unrolling and improves robustness across inference budgets.
Through controlled pretraining experiments, this study systematically examines when recurrence helps, where it should be applied, and how it should be conditioned in looped language models (LoopLMs), finding that extra loops beyond the training horizon improve reasoning (e.g., a reasoning score rising from 28.52 to 31.49) while degrading knowledge (from 62.80 to 52.11), that non-recurrent output layers improve robustness to under-unrolling, and that the proposed history-state injection (especially channel-wise) combined with timestep conditioning better preserves knowledge under extended unrolling and improves robustness across inference budgets.
Through controlled pretraining experiments, this study systematically examines when recurrence helps, where it should be applied, and how it should be conditioned in looped language models (LoopLMs), finding that extra loops beyond the training horizon improve reasoning (e.g., a reasoning score rising from 28.52 to 31.49) while degrading knowledge (from 62.80 to 52.11), that non-recurrent output layers improve robustness to under-unrolling, and that the proposed history-state injection (especially channel-wise) combined with timestep conditioning better preserves knowledge under extended unrolling and improves robustness across inference budgets.
arXiv The work builds HybridCUA-8K, a data pipeline of 5,000 GUI-only, CLI-only, and interleaved GUI–CLI trajectories plus 3,000 verified RLVR tasks, and a two-stage framework of supervised fine-tuning followed by online reinforcement learning with CLI-aware rewards; the resulting HybridCUA-9B reaches 53.6% accuracy on OSWorld (14.8 points above its Qwen3.5-9B base, with average steps falling from 31.6 to 14.0) and 36.0% on WindowsAgentArena (a 4.0-point gain).
The work builds HybridCUA-8K, a data pipeline of 5,000 GUI-only, CLI-only, and interleaved GUI–CLI trajectories plus 3,000 verified RLVR tasks, and a two-stage framework of supervised fine-tuning followed by online reinforcement learning with CLI-aware rewards; the resulting HybridCUA-9B reaches 53.6% accuracy on OSWorld (14.8 points above its Qwen3.5-9B base, with average steps falling from 31.6 to 14.0) and 36.0% on WindowsAgentArena (a 4.0-point gain).
The work builds HybridCUA-8K, a data pipeline of 5,000 GUI-only, CLI-only, and interleaved GUI–CLI trajectories plus 3,000 verified RLVR tasks, and a two-stage framework of supervised fine-tuning followed by online reinforcement learning with CLI-aware rewards; the resulting HybridCUA-9B reaches 53.6% accuracy on OSWorld (14.8 points above its Qwen3.5-9B base, with average steps falling from 31.6 to 14.0) and 36.0% on WindowsAgentArena (a 4.0-point gain).
The work builds HybridCUA-8K, a data pipeline of 5,000 GUI-only, CLI-only, and interleaved GUI–CLI trajectories plus 3,000 verified RLVR tasks, and a two-stage framework of supervised fine-tuning followed by online reinforcement learning with CLI-aware rewards; the resulting HybridCUA-9B reaches 53.6% accuracy on OSWorld (14.8 points above its Qwen3.5-9B base, with average steps falling from 31.6 to 14.0) and 36.0% on WindowsAgentArena (a 4.0-point gain).
arXiv CrossBFM freezes the backward map of a Behavior Foundation Model pretrained on the Unitree G1 and, using the frame-level cross-embodiment correspondence supplied by retargeted data, distills that latent space onto three humanoids (Inhouse M3, Booster T1, Fourier N1) with a unified encoder that has no robot-specific parameters, training the encoder in under one GPU-hour and each tracker in about 10 more GPU-hours; all three prompting modes transfer (motion tracking, goal reaching between poses, reward optimization), the latent-conditioned policy loses only 0.
CrossBFM freezes the backward map of a Behavior Foundation Model pretrained on the Unitree G1 and, using the frame-level cross-embodiment correspondence supplied by retargeted data, distills that latent space onto three humanoids (Inhouse M3, Booster T1, Fourier N1) with a unified encoder that has no robot-specific parameters, training the encoder in under one GPU-hour and each tracker in about 10 more GPU-hours; all three prompting modes transfer (motion tracking, goal reaching between poses, reward optimization), the latent-conditioned policy loses only 0.
CrossBFM freezes the backward map of a Behavior Foundation Model pretrained on the Unitree G1 and, using the frame-level cross-embodiment correspondence supplied by retargeted data, distills that latent space onto three humanoids (Inhouse M3, Booster T1, Fourier N1) with a unified encoder that has no robot-specific parameters, training the encoder in under one GPU-hour and each tracker in about 10 more GPU-hours; all three prompting modes transfer (motion tracking, goal reaching between poses, reward optimization), the latent-conditioned policy loses only 0.
CrossBFM freezes the backward map of a Behavior Foundation Model pretrained on the Unitree G1 and, using the frame-level cross-embodiment correspondence supplied by retargeted data, distills that latent space onto three humanoids (Inhouse M3, Booster T1, Fourier N1) with a unified encoder that has no robot-specific parameters, training the encoder in under one GPU-hour and each tracker in about 10 more GPU-hours; all three prompting modes transfer (motion tracking, goal reaching between poses, reward optimization), the latent-conditioned policy loses only 0.
arXiv The work adds a content axis to agent proactivity—horizontal proactivity pursues unstated information the current context already identifies, vertical proactivity pursues needs only earlier evidence reveals—scores runs against a need graph recovered mechanically from multi-hop benchmarks' own decompositions with no model judge, and proposes Q&D (questioner and drafter), which forks a recorded run at one step, samples eight candidate questions, continues each to the end, and prefers the question whose continuation retrieves more required evidence, so no reward model or judge is needed; on held-out splits of MuSiQue, StrategyQA and 2WikiMultiHopQA, at equal retrieval spend the trained 8B questioner recovers more required evidence than the same model prompted (89.5% versus 78.
The work adds a content axis to agent proactivity—horizontal proactivity pursues unstated information the current context already identifies, vertical proactivity pursues needs only earlier evidence reveals—scores runs against a need graph recovered mechanically from multi-hop benchmarks' own decompositions with no model judge, and proposes Q&D (questioner and drafter), which forks a recorded run at one step, samples eight candidate questions, continues each to the end, and prefers the question whose continuation retrieves more required evidence, so no reward model or judge is needed; on held-out splits of MuSiQue, StrategyQA and 2WikiMultiHopQA, at equal retrieval spend the trained 8B questioner recovers more required evidence than the same model prompted (89.5% versus 78.
The work adds a content axis to agent proactivity—horizontal proactivity pursues unstated information the current context already identifies, vertical proactivity pursues needs only earlier evidence reveals—scores runs against a need graph recovered mechanically from multi-hop benchmarks' own decompositions with no model judge, and proposes Q&D (questioner and drafter), which forks a recorded run at one step, samples eight candidate questions, continues each to the end, and prefers the question whose continuation retrieves more required evidence, so no reward model or judge is needed; on held-out splits of MuSiQue, StrategyQA and 2WikiMultiHopQA, at equal retrieval spend the trained 8B questioner recovers more required evidence than the same model prompted (89.5% versus 78.
The work adds a content axis to agent proactivity—horizontal proactivity pursues unstated information the current context already identifies, vertical proactivity pursues needs only earlier evidence reveals—scores runs against a need graph recovered mechanically from multi-hop benchmarks' own decompositions with no model judge, and proposes Q&D (questioner and drafter), which forks a recorded run at one step, samples eight candidate questions, continues each to the end, and prefers the question whose continuation retrieves more required evidence, so no reward model or judge is needed; on held-out splits of MuSiQue, StrategyQA and 2WikiMultiHopQA, at equal retrieval spend the trained 8B questioner recovers more required evidence than the same model prompted (89.5% versus 78.
arXiv The work proposes EVO-WAM, a framework that enables complete autoregressive rollouts without external execution feedback via state prediction and anchored multi-frame context, then selects task-completing prefixes with a vision-language model and verifies video-action consistency with an inverse dynamics model for iterative self-training; on seven unseen RoboTwin 2.0 tasks it raises average success from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, and on three real-world long-horizon composite tasks it raises Cosmos3 from 20.0% to 76.7%.
The work proposes EVO-WAM, a framework that enables complete autoregressive rollouts without external execution feedback via state prediction and anchored multi-frame context, then selects task-completing prefixes with a vision-language model and verifies video-action consistency with an inverse dynamics model for iterative self-training; on seven unseen RoboTwin 2.0 tasks it raises average success from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, and on three real-world long-horizon composite tasks it raises Cosmos3 from 20.0% to 76.7%.
The work proposes EVO-WAM, a framework that enables complete autoregressive rollouts without external execution feedback via state prediction and anchored multi-frame context, then selects task-completing prefixes with a vision-language model and verifies video-action consistency with an inverse dynamics model for iterative self-training; on seven unseen RoboTwin 2.0 tasks it raises average success from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, and on three real-world long-horizon composite tasks it raises Cosmos3 from 20.0% to 76.7%.
The work proposes EVO-WAM, a framework that enables complete autoregressive rollouts without external execution feedback via state prediction and anchored multi-frame context, then selects task-completing prefixes with a vision-language model and verifies video-action consistency with an inverse dynamics model for iterative self-training; on seven unseen RoboTwin 2.0 tasks it raises average success from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, and on three real-world long-horizon composite tasks it raises Cosmos3 from 20.0% to 76.7%.
arXiv The work shows that isotropic Gaussian regularization (SIGReg) can make Euclidean latent planning cost rank feasible outcomes differently from the task cost, and introduces AnisoWM with ΛReg, which learns a diagonal covariance target under fixed-trace and anisotropy constraints while leaving the prediction objective and Euclidean planner unchanged, improving planning success over LeWM in all four visual control environments (TwoRoom, Reacher, PushT, Cube) and yielding better agreement between latent costs and task outcomes.
The work shows that isotropic Gaussian regularization (SIGReg) can make Euclidean latent planning cost rank feasible outcomes differently from the task cost, and introduces AnisoWM with ΛReg, which learns a diagonal covariance target under fixed-trace and anisotropy constraints while leaving the prediction objective and Euclidean planner unchanged, improving planning success over LeWM in all four visual control environments (TwoRoom, Reacher, PushT, Cube) and yielding better agreement between latent costs and task outcomes.
The work shows that isotropic Gaussian regularization (SIGReg) can make Euclidean latent planning cost rank feasible outcomes differently from the task cost, and introduces AnisoWM with ΛReg, which learns a diagonal covariance target under fixed-trace and anisotropy constraints while leaving the prediction objective and Euclidean planner unchanged, improving planning success over LeWM in all four visual control environments (TwoRoom, Reacher, PushT, Cube) and yielding better agreement between latent costs and task outcomes.
The work shows that isotropic Gaussian regularization (SIGReg) can make Euclidean latent planning cost rank feasible outcomes differently from the task cost, and introduces AnisoWM with ΛReg, which learns a diagonal covariance target under fixed-trace and anisotropy constraints while leaving the prediction objective and Euclidean planner unchanged, improving planning success over LeWM in all four visual control environments (TwoRoom, Reacher, PushT, Cube) and yielding better agreement between latent costs and task outcomes.
arXiv The work proposes ReImaGin, a training-free multimodal agent framework in which a multimodal LLM calls a free-form generate_image tool backed by an instruction-following image generation model inside its reasoning loop; across six tasks—depth perception, puzzle completion, occlusion counting, collision prediction, multi-view spatial reasoning, and path tracing—it consistently outperforms text-only reasoning and the fixed specialist vision-tool baseline Visual Sketchpad, with gains of up to 25% on path tracing and about 40% relative on occlusion counting, and it further shows that effective visual reasoning strategies can be discovered automatically via prompt optimization.
The work proposes ReImaGin, a training-free multimodal agent framework in which a multimodal LLM calls a free-form generate_image tool backed by an instruction-following image generation model inside its reasoning loop; across six tasks—depth perception, puzzle completion, occlusion counting, collision prediction, multi-view spatial reasoning, and path tracing—it consistently outperforms text-only reasoning and the fixed specialist vision-tool baseline Visual Sketchpad, with gains of up to 25% on path tracing and about 40% relative on occlusion counting, and it further shows that effective visual reasoning strategies can be discovered automatically via prompt optimization.
The work proposes ReImaGin, a training-free multimodal agent framework in which a multimodal LLM calls a free-form generate_image tool backed by an instruction-following image generation model inside its reasoning loop; across six tasks—depth perception, puzzle completion, occlusion counting, collision prediction, multi-view spatial reasoning, and path tracing—it consistently outperforms text-only reasoning and the fixed specialist vision-tool baseline Visual Sketchpad, with gains of up to 25% on path tracing and about 40% relative on occlusion counting, and it further shows that effective visual reasoning strategies can be discovered automatically via prompt optimization.
The work proposes ReImaGin, a training-free multimodal agent framework in which a multimodal LLM calls a free-form generate_image tool backed by an instruction-following image generation model inside its reasoning loop; across six tasks—depth perception, puzzle completion, occlusion counting, collision prediction, multi-view spatial reasoning, and path tracing—it consistently outperforms text-only reasoning and the fixed specialist vision-tool baseline Visual Sketchpad, with gains of up to 25% on path tracing and about 40% relative on occlusion counting, and it further shows that effective visual reasoning strategies can be discovered automatically via prompt optimization.
arXiv Under matched backbone, data, and training budget, the study compares explicit and latent world action models and finds that latent models match explicit ones in distribution (96.85 vs 97.75) yet fall behind on all three generalization axes—environmental perturbation, data efficiency, and task generalization—while leaving future video tokens at pure noise and running a single forward pass recovers most of the gap, motivating Simple-WAM, which leads explicit models in generalization at latent-level inference speed.
Under matched backbone, data, and training budget, the study compares explicit and latent world action models and finds that latent models match explicit ones in distribution (96.85 vs 97.75) yet fall behind on all three generalization axes—environmental perturbation, data efficiency, and task generalization—while leaving future video tokens at pure noise and running a single forward pass recovers most of the gap, motivating Simple-WAM, which leads explicit models in generalization at latent-level inference speed.
Under matched backbone, data, and training budget, the study compares explicit and latent world action models and finds that latent models match explicit ones in distribution (96.85 vs 97.75) yet fall behind on all three generalization axes—environmental perturbation, data efficiency, and task generalization—while leaving future video tokens at pure noise and running a single forward pass recovers most of the gap, motivating Simple-WAM, which leads explicit models in generalization at latent-level inference speed.
Under matched backbone, data, and training budget, the study compares explicit and latent world action models and finds that latent models match explicit ones in distribution (96.85 vs 97.75) yet fall behind on all three generalization axes—environmental perturbation, data efficiency, and task generalization—while leaving future video tokens at pure noise and running a single forward pass recovers most of the gap, motivating Simple-WAM, which leads explicit models in generalization at latent-level inference speed.
arXiv The work derives a token- and layer-wise decomposition of how learning from one token changes another prediction, splitting it into a direct force-alignment channel (CH1) and a diffuse channel mediated by shared readout geometry (CH2), with a forward-computable approximation using only logits, hidden states, and the readout; following this interaction over time, the authors report that positive interaction supports data selection, negative interaction splits into concentrated collision and accumulated erosion with distinct controls, and that declining task-conditioned readout transmission during long-horizon training tracks declining future learnability, with readout restoration partially recovering plasticity.
The work derives a token- and layer-wise decomposition of how learning from one token changes another prediction, splitting it into a direct force-alignment channel (CH1) and a diffuse channel mediated by shared readout geometry (CH2), with a forward-computable approximation using only logits, hidden states, and the readout; following this interaction over time, the authors report that positive interaction supports data selection, negative interaction splits into concentrated collision and accumulated erosion with distinct controls, and that declining task-conditioned readout transmission during long-horizon training tracks declining future learnability, with readout restoration partially recovering plasticity.
The work derives a token- and layer-wise decomposition of how learning from one token changes another prediction, splitting it into a direct force-alignment channel (CH1) and a diffuse channel mediated by shared readout geometry (CH2), with a forward-computable approximation using only logits, hidden states, and the readout; following this interaction over time, the authors report that positive interaction supports data selection, negative interaction splits into concentrated collision and accumulated erosion with distinct controls, and that declining task-conditioned readout transmission during long-horizon training tracks declining future learnability, with readout restoration partially recovering plasticity.
The work derives a token- and layer-wise decomposition of how learning from one token changes another prediction, splitting it into a direct force-alignment channel (CH1) and a diffuse channel mediated by shared readout geometry (CH2), with a forward-computable approximation using only logits, hidden states, and the readout; following this interaction over time, the authors report that positive interaction supports data selection, negative interaction splits into concentrated collision and accumulated erosion with distinct controls, and that declining task-conditioned readout transmission during long-horizon training tracks declining future learnability, with readout restoration partially recovering plasticity.
arXiv Using sparse crosscoders and a proposed swap readout, the study reads how each student checkpoint uses every feature before and after on-policy distillation (OPD) across three OPD settings, finding that OPD neither creates the student's own features nor passes on the teacher's, but slightly reweights features the student already shares with the teacher (over 98% of frequently used features change firing rate by less than 20%), with the largest changes concentrated on decision tokens such as Wait, Hmm, and So; the SFT warm-up on the teacher's rollouts that commonly precedes OPD also adds no features, reweighting shared features partly along OPD's direction and partly beyond it, and imposing this reweighting on a directly distilled student's features without changing its weights brings its a
Using sparse crosscoders and a proposed swap readout, the study reads how each student checkpoint uses every feature before and after on-policy distillation (OPD) across three OPD settings, finding that OPD neither creates the student's own features nor passes on the teacher's, but slightly reweights features the student already shares with the teacher (over 98% of frequently used features change firing rate by less than 20%), with the largest changes concentrated on decision tokens such as Wait, Hmm, and So; the SFT warm-up on the teacher's rollouts that commonly precedes OPD also adds no features, reweighting shared features partly along OPD's direction and partly beyond it, and imposing this reweighting on a directly distilled student's features without changing its weights brings its a
Using sparse crosscoders and a proposed swap readout, the study reads how each student checkpoint uses every feature before and after on-policy distillation (OPD) across three OPD settings, finding that OPD neither creates the student's own features nor passes on the teacher's, but slightly reweights features the student already shares with the teacher (over 98% of frequently used features change firing rate by less than 20%), with the largest changes concentrated on decision tokens such as Wait, Hmm, and So; the SFT warm-up on the teacher's rollouts that commonly precedes OPD also adds no features, reweighting shared features partly along OPD's direction and partly beyond it, and imposing this reweighting on a directly distilled student's features without changing its weights brings its a
Using sparse crosscoders and a proposed swap readout, the study reads how each student checkpoint uses every feature before and after on-policy distillation (OPD) across three OPD settings, finding that OPD neither creates the student's own features nor passes on the teacher's, but slightly reweights features the student already shares with the teacher (over 98% of frequently used features change firing rate by less than 20%), with the largest changes concentrated on decision tokens such as Wait, Hmm, and So; the SFT warm-up on the teacher's rollouts that commonly precedes OPD also adds no features, reweighting shared features partly along OPD's direction and partly beyond it, and imposing this reweighting on a directly distilled student's features without changing its weights brings its a
arXiv The work introduces Chinese-Jev: a unified data pipeline converts heterogeneous Chinese annotations into probability targets over candidate options, a lightweight encoder-only model is first pre-trained on 10 million general Chinese decisions and then fine-tuned separately for medicine, law, and finance, and CJ-Bench is released with 307,900 held-out decisions; after pre-training it reaches 69.20% accuracy on general tasks, a 1.24% relative gain over the closed-source Jev, cuts expected calibration error from 11.45% to 3.78% at 14 ms latency, improves medical accuracy over Jev by 4.0%, reaches 92% of Jev's average accuracy across the three specialized domains at 15 ms average latency, and an INT8-quantized model runs about 1 second per decision in a mobile browser.
The work introduces Chinese-Jev: a unified data pipeline converts heterogeneous Chinese annotations into probability targets over candidate options, a lightweight encoder-only model is first pre-trained on 10 million general Chinese decisions and then fine-tuned separately for medicine, law, and finance, and CJ-Bench is released with 307,900 held-out decisions; after pre-training it reaches 69.20% accuracy on general tasks, a 1.24% relative gain over the closed-source Jev, cuts expected calibration error from 11.45% to 3.78% at 14 ms latency, improves medical accuracy over Jev by 4.0%, reaches 92% of Jev's average accuracy across the three specialized domains at 15 ms average latency, and an INT8-quantized model runs about 1 second per decision in a mobile browser.
The work introduces Chinese-Jev: a unified data pipeline converts heterogeneous Chinese annotations into probability targets over candidate options, a lightweight encoder-only model is first pre-trained on 10 million general Chinese decisions and then fine-tuned separately for medicine, law, and finance, and CJ-Bench is released with 307,900 held-out decisions; after pre-training it reaches 69.20% accuracy on general tasks, a 1.24% relative gain over the closed-source Jev, cuts expected calibration error from 11.45% to 3.78% at 14 ms latency, improves medical accuracy over Jev by 4.0%, reaches 92% of Jev's average accuracy across the three specialized domains at 15 ms average latency, and an INT8-quantized model runs about 1 second per decision in a mobile browser.
The work introduces Chinese-Jev: a unified data pipeline converts heterogeneous Chinese annotations into probability targets over candidate options, a lightweight encoder-only model is first pre-trained on 10 million general Chinese decisions and then fine-tuned separately for medicine, law, and finance, and CJ-Bench is released with 307,900 held-out decisions; after pre-training it reaches 69.20% accuracy on general tasks, a 1.24% relative gain over the closed-source Jev, cuts expected calibration error from 11.45% to 3.78% at 14 ms latency, improves medical accuracy over Jev by 4.0%, reaches 92% of Jev's average accuracy across the three specialized domains at 15 ms average latency, and an INT8-quantized model runs about 1 second per decision in a mobile browser.