arXiv The work adds a content axis to agent proactivity—horizontal proactivity pursues unstated information the current context already identifies, vertical proactivity pursues needs only earlier evidence reveals—scores runs against a need graph recovered mechanically from multi-hop benchmarks' own decompositions with no model judge, and proposes Q&D (questioner and drafter), which forks a recorded run at one step, samples eight candidate questions, continues each to the end, and prefers the question whose continuation retrieves more required evidence, so no reward model or judge is needed; on held-out splits of MuSiQue, StrategyQA and 2WikiMultiHopQA, at equal retrieval spend the trained 8B questioner recovers more required evidence than the same model prompted (89.5% versus 78.
The work adds a content axis to agent proactivity—horizontal proactivity pursues unstated information the current context already identifies, vertical proactivity pursues needs only earlier evidence reveals—scores runs against a need graph recovered mechanically from multi-hop benchmarks' own decompositions with no model judge, and proposes Q&D (questioner and drafter), which forks a recorded run at one step, samples eight candidate questions, continues each to the end, and prefers the question whose continuation retrieves more required evidence, so no reward model or judge is needed; on held-out splits of MuSiQue, StrategyQA and 2WikiMultiHopQA, at equal retrieval spend the trained 8B questioner recovers more required evidence than the same model prompted (89.5% versus 78.
The work adds a content axis to agent proactivity—horizontal proactivity pursues unstated information the current context already identifies, vertical proactivity pursues needs only earlier evidence reveals—scores runs against a need graph recovered mechanically from multi-hop benchmarks' own decompositions with no model judge, and proposes Q&D (questioner and drafter), which forks a recorded run at one step, samples eight candidate questions, continues each to the end, and prefers the question whose continuation retrieves more required evidence, so no reward model or judge is needed; on held-out splits of MuSiQue, StrategyQA and 2WikiMultiHopQA, at equal retrieval spend the trained 8B questioner recovers more required evidence than the same model prompted (89.5% versus 78.
The work adds a content axis to agent proactivity—horizontal proactivity pursues unstated information the current context already identifies, vertical proactivity pursues needs only earlier evidence reveals—scores runs against a need graph recovered mechanically from multi-hop benchmarks' own decompositions with no model judge, and proposes Q&D (questioner and drafter), which forks a recorded run at one step, samples eight candidate questions, continues each to the end, and prefers the question whose continuation retrieves more required evidence, so no reward model or judge is needed; on held-out splits of MuSiQue, StrategyQA and 2WikiMultiHopQA, at equal retrieval spend the trained 8B questioner recovers more required evidence than the same model prompted (89.5% versus 78.
arXiv The work proposes EVO-WAM, a framework that enables complete autoregressive rollouts without external execution feedback via state prediction and anchored multi-frame context, then selects task-completing prefixes with a vision-language model and verifies video-action consistency with an inverse dynamics model for iterative self-training; on seven unseen RoboTwin 2.0 tasks it raises average success from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, and on three real-world long-horizon composite tasks it raises Cosmos3 from 20.0% to 76.7%.
The work proposes EVO-WAM, a framework that enables complete autoregressive rollouts without external execution feedback via state prediction and anchored multi-frame context, then selects task-completing prefixes with a vision-language model and verifies video-action consistency with an inverse dynamics model for iterative self-training; on seven unseen RoboTwin 2.0 tasks it raises average success from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, and on three real-world long-horizon composite tasks it raises Cosmos3 from 20.0% to 76.7%.
The work proposes EVO-WAM, a framework that enables complete autoregressive rollouts without external execution feedback via state prediction and anchored multi-frame context, then selects task-completing prefixes with a vision-language model and verifies video-action consistency with an inverse dynamics model for iterative self-training; on seven unseen RoboTwin 2.0 tasks it raises average success from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, and on three real-world long-horizon composite tasks it raises Cosmos3 from 20.0% to 76.7%.
The work proposes EVO-WAM, a framework that enables complete autoregressive rollouts without external execution feedback via state prediction and anchored multi-frame context, then selects task-completing prefixes with a vision-language model and verifies video-action consistency with an inverse dynamics model for iterative self-training; on seven unseen RoboTwin 2.0 tasks it raises average success from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, and on three real-world long-horizon composite tasks it raises Cosmos3 from 20.0% to 76.7%.
arXiv The work shows that isotropic Gaussian regularization (SIGReg) can make Euclidean latent planning cost rank feasible outcomes differently from the task cost, and introduces AnisoWM with ΛReg, which learns a diagonal covariance target under fixed-trace and anisotropy constraints while leaving the prediction objective and Euclidean planner unchanged, improving planning success over LeWM in all four visual control environments (TwoRoom, Reacher, PushT, Cube) and yielding better agreement between latent costs and task outcomes.
The work shows that isotropic Gaussian regularization (SIGReg) can make Euclidean latent planning cost rank feasible outcomes differently from the task cost, and introduces AnisoWM with ΛReg, which learns a diagonal covariance target under fixed-trace and anisotropy constraints while leaving the prediction objective and Euclidean planner unchanged, improving planning success over LeWM in all four visual control environments (TwoRoom, Reacher, PushT, Cube) and yielding better agreement between latent costs and task outcomes.
The work shows that isotropic Gaussian regularization (SIGReg) can make Euclidean latent planning cost rank feasible outcomes differently from the task cost, and introduces AnisoWM with ΛReg, which learns a diagonal covariance target under fixed-trace and anisotropy constraints while leaving the prediction objective and Euclidean planner unchanged, improving planning success over LeWM in all four visual control environments (TwoRoom, Reacher, PushT, Cube) and yielding better agreement between latent costs and task outcomes.
The work shows that isotropic Gaussian regularization (SIGReg) can make Euclidean latent planning cost rank feasible outcomes differently from the task cost, and introduces AnisoWM with ΛReg, which learns a diagonal covariance target under fixed-trace and anisotropy constraints while leaving the prediction objective and Euclidean planner unchanged, improving planning success over LeWM in all four visual control environments (TwoRoom, Reacher, PushT, Cube) and yielding better agreement between latent costs and task outcomes.
arXiv The work proposes ReImaGin, a training-free multimodal agent framework in which a multimodal LLM calls a free-form generate_image tool backed by an instruction-following image generation model inside its reasoning loop; across six tasks—depth perception, puzzle completion, occlusion counting, collision prediction, multi-view spatial reasoning, and path tracing—it consistently outperforms text-only reasoning and the fixed specialist vision-tool baseline Visual Sketchpad, with gains of up to 25% on path tracing and about 40% relative on occlusion counting, and it further shows that effective visual reasoning strategies can be discovered automatically via prompt optimization.
The work proposes ReImaGin, a training-free multimodal agent framework in which a multimodal LLM calls a free-form generate_image tool backed by an instruction-following image generation model inside its reasoning loop; across six tasks—depth perception, puzzle completion, occlusion counting, collision prediction, multi-view spatial reasoning, and path tracing—it consistently outperforms text-only reasoning and the fixed specialist vision-tool baseline Visual Sketchpad, with gains of up to 25% on path tracing and about 40% relative on occlusion counting, and it further shows that effective visual reasoning strategies can be discovered automatically via prompt optimization.
The work proposes ReImaGin, a training-free multimodal agent framework in which a multimodal LLM calls a free-form generate_image tool backed by an instruction-following image generation model inside its reasoning loop; across six tasks—depth perception, puzzle completion, occlusion counting, collision prediction, multi-view spatial reasoning, and path tracing—it consistently outperforms text-only reasoning and the fixed specialist vision-tool baseline Visual Sketchpad, with gains of up to 25% on path tracing and about 40% relative on occlusion counting, and it further shows that effective visual reasoning strategies can be discovered automatically via prompt optimization.
The work proposes ReImaGin, a training-free multimodal agent framework in which a multimodal LLM calls a free-form generate_image tool backed by an instruction-following image generation model inside its reasoning loop; across six tasks—depth perception, puzzle completion, occlusion counting, collision prediction, multi-view spatial reasoning, and path tracing—it consistently outperforms text-only reasoning and the fixed specialist vision-tool baseline Visual Sketchpad, with gains of up to 25% on path tracing and about 40% relative on occlusion counting, and it further shows that effective visual reasoning strategies can be discovered automatically via prompt optimization.
arXiv Under matched backbone, data, and training budget, the study compares explicit and latent world action models and finds that latent models match explicit ones in distribution (96.85 vs 97.75) yet fall behind on all three generalization axes—environmental perturbation, data efficiency, and task generalization—while leaving future video tokens at pure noise and running a single forward pass recovers most of the gap, motivating Simple-WAM, which leads explicit models in generalization at latent-level inference speed.
Under matched backbone, data, and training budget, the study compares explicit and latent world action models and finds that latent models match explicit ones in distribution (96.85 vs 97.75) yet fall behind on all three generalization axes—environmental perturbation, data efficiency, and task generalization—while leaving future video tokens at pure noise and running a single forward pass recovers most of the gap, motivating Simple-WAM, which leads explicit models in generalization at latent-level inference speed.
Under matched backbone, data, and training budget, the study compares explicit and latent world action models and finds that latent models match explicit ones in distribution (96.85 vs 97.75) yet fall behind on all three generalization axes—environmental perturbation, data efficiency, and task generalization—while leaving future video tokens at pure noise and running a single forward pass recovers most of the gap, motivating Simple-WAM, which leads explicit models in generalization at latent-level inference speed.
Under matched backbone, data, and training budget, the study compares explicit and latent world action models and finds that latent models match explicit ones in distribution (96.85 vs 97.75) yet fall behind on all three generalization axes—environmental perturbation, data efficiency, and task generalization—while leaving future video tokens at pure noise and running a single forward pass recovers most of the gap, motivating Simple-WAM, which leads explicit models in generalization at latent-level inference speed.
arXiv The work derives a token- and layer-wise decomposition of how learning from one token changes another prediction, splitting it into a direct force-alignment channel (CH1) and a diffuse channel mediated by shared readout geometry (CH2), with a forward-computable approximation using only logits, hidden states, and the readout; following this interaction over time, the authors report that positive interaction supports data selection, negative interaction splits into concentrated collision and accumulated erosion with distinct controls, and that declining task-conditioned readout transmission during long-horizon training tracks declining future learnability, with readout restoration partially recovering plasticity.
The work derives a token- and layer-wise decomposition of how learning from one token changes another prediction, splitting it into a direct force-alignment channel (CH1) and a diffuse channel mediated by shared readout geometry (CH2), with a forward-computable approximation using only logits, hidden states, and the readout; following this interaction over time, the authors report that positive interaction supports data selection, negative interaction splits into concentrated collision and accumulated erosion with distinct controls, and that declining task-conditioned readout transmission during long-horizon training tracks declining future learnability, with readout restoration partially recovering plasticity.
The work derives a token- and layer-wise decomposition of how learning from one token changes another prediction, splitting it into a direct force-alignment channel (CH1) and a diffuse channel mediated by shared readout geometry (CH2), with a forward-computable approximation using only logits, hidden states, and the readout; following this interaction over time, the authors report that positive interaction supports data selection, negative interaction splits into concentrated collision and accumulated erosion with distinct controls, and that declining task-conditioned readout transmission during long-horizon training tracks declining future learnability, with readout restoration partially recovering plasticity.
The work derives a token- and layer-wise decomposition of how learning from one token changes another prediction, splitting it into a direct force-alignment channel (CH1) and a diffuse channel mediated by shared readout geometry (CH2), with a forward-computable approximation using only logits, hidden states, and the readout; following this interaction over time, the authors report that positive interaction supports data selection, negative interaction splits into concentrated collision and accumulated erosion with distinct controls, and that declining task-conditioned readout transmission during long-horizon training tracks declining future learnability, with readout restoration partially recovering plasticity.
arXiv Using sparse crosscoders and a proposed swap readout, the study reads how each student checkpoint uses every feature before and after on-policy distillation (OPD) across three OPD settings, finding that OPD neither creates the student's own features nor passes on the teacher's, but slightly reweights features the student already shares with the teacher (over 98% of frequently used features change firing rate by less than 20%), with the largest changes concentrated on decision tokens such as Wait, Hmm, and So; the SFT warm-up on the teacher's rollouts that commonly precedes OPD also adds no features, reweighting shared features partly along OPD's direction and partly beyond it, and imposing this reweighting on a directly distilled student's features without changing its weights brings its a
Using sparse crosscoders and a proposed swap readout, the study reads how each student checkpoint uses every feature before and after on-policy distillation (OPD) across three OPD settings, finding that OPD neither creates the student's own features nor passes on the teacher's, but slightly reweights features the student already shares with the teacher (over 98% of frequently used features change firing rate by less than 20%), with the largest changes concentrated on decision tokens such as Wait, Hmm, and So; the SFT warm-up on the teacher's rollouts that commonly precedes OPD also adds no features, reweighting shared features partly along OPD's direction and partly beyond it, and imposing this reweighting on a directly distilled student's features without changing its weights brings its a
Using sparse crosscoders and a proposed swap readout, the study reads how each student checkpoint uses every feature before and after on-policy distillation (OPD) across three OPD settings, finding that OPD neither creates the student's own features nor passes on the teacher's, but slightly reweights features the student already shares with the teacher (over 98% of frequently used features change firing rate by less than 20%), with the largest changes concentrated on decision tokens such as Wait, Hmm, and So; the SFT warm-up on the teacher's rollouts that commonly precedes OPD also adds no features, reweighting shared features partly along OPD's direction and partly beyond it, and imposing this reweighting on a directly distilled student's features without changing its weights brings its a
Using sparse crosscoders and a proposed swap readout, the study reads how each student checkpoint uses every feature before and after on-policy distillation (OPD) across three OPD settings, finding that OPD neither creates the student's own features nor passes on the teacher's, but slightly reweights features the student already shares with the teacher (over 98% of frequently used features change firing rate by less than 20%), with the largest changes concentrated on decision tokens such as Wait, Hmm, and So; the SFT warm-up on the teacher's rollouts that commonly precedes OPD also adds no features, reweighting shared features partly along OPD's direction and partly beyond it, and imposing this reweighting on a directly distilled student's features without changing its weights brings its a
arXiv The work introduces Chinese-Jev: a unified data pipeline converts heterogeneous Chinese annotations into probability targets over candidate options, a lightweight encoder-only model is first pre-trained on 10 million general Chinese decisions and then fine-tuned separately for medicine, law, and finance, and CJ-Bench is released with 307,900 held-out decisions; after pre-training it reaches 69.20% accuracy on general tasks, a 1.24% relative gain over the closed-source Jev, cuts expected calibration error from 11.45% to 3.78% at 14 ms latency, improves medical accuracy over Jev by 4.0%, reaches 92% of Jev's average accuracy across the three specialized domains at 15 ms average latency, and an INT8-quantized model runs about 1 second per decision in a mobile browser.
The work introduces Chinese-Jev: a unified data pipeline converts heterogeneous Chinese annotations into probability targets over candidate options, a lightweight encoder-only model is first pre-trained on 10 million general Chinese decisions and then fine-tuned separately for medicine, law, and finance, and CJ-Bench is released with 307,900 held-out decisions; after pre-training it reaches 69.20% accuracy on general tasks, a 1.24% relative gain over the closed-source Jev, cuts expected calibration error from 11.45% to 3.78% at 14 ms latency, improves medical accuracy over Jev by 4.0%, reaches 92% of Jev's average accuracy across the three specialized domains at 15 ms average latency, and an INT8-quantized model runs about 1 second per decision in a mobile browser.
The work introduces Chinese-Jev: a unified data pipeline converts heterogeneous Chinese annotations into probability targets over candidate options, a lightweight encoder-only model is first pre-trained on 10 million general Chinese decisions and then fine-tuned separately for medicine, law, and finance, and CJ-Bench is released with 307,900 held-out decisions; after pre-training it reaches 69.20% accuracy on general tasks, a 1.24% relative gain over the closed-source Jev, cuts expected calibration error from 11.45% to 3.78% at 14 ms latency, improves medical accuracy over Jev by 4.0%, reaches 92% of Jev's average accuracy across the three specialized domains at 15 ms average latency, and an INT8-quantized model runs about 1 second per decision in a mobile browser.
The work introduces Chinese-Jev: a unified data pipeline converts heterogeneous Chinese annotations into probability targets over candidate options, a lightweight encoder-only model is first pre-trained on 10 million general Chinese decisions and then fine-tuned separately for medicine, law, and finance, and CJ-Bench is released with 307,900 held-out decisions; after pre-training it reaches 69.20% accuracy on general tasks, a 1.24% relative gain over the closed-source Jev, cuts expected calibration error from 11.45% to 3.78% at 14 ms latency, improves medical accuracy over Jev by 4.0%, reaches 92% of Jev's average accuracy across the three specialized domains at 15 ms average latency, and an INT8-quantized model runs about 1 second per decision in a mobile browser.
arXiv The work proposes Temperature-Grouped Reinforcement Learning (TGRL), which partitions each prompt's rollout group into low- and high-temperature subsets, estimates prompt-level exploration gain from their reward contrast, and allocates that gain as token-level credit using Jensen–Shannon divergence between the temperature-scaled next-token distributions induced by the same logits; across 11 benchmarks, TGRL improves the six-benchmark math average by 1.6% at 32B, raises CodeForces rating by 196.7 points and LiveCodeBench Pass@16 by 4.4%, and improves ALFWorld/WebShop success rates by 6.3%/4.9%, reaching equivalent accuracy up to 36% faster than strong RLVR baselines without expanding the rollout budget.
The work proposes Temperature-Grouped Reinforcement Learning (TGRL), which partitions each prompt's rollout group into low- and high-temperature subsets, estimates prompt-level exploration gain from their reward contrast, and allocates that gain as token-level credit using Jensen–Shannon divergence between the temperature-scaled next-token distributions induced by the same logits; across 11 benchmarks, TGRL improves the six-benchmark math average by 1.6% at 32B, raises CodeForces rating by 196.7 points and LiveCodeBench Pass@16 by 4.4%, and improves ALFWorld/WebShop success rates by 6.3%/4.9%, reaching equivalent accuracy up to 36% faster than strong RLVR baselines without expanding the rollout budget.
The work proposes Temperature-Grouped Reinforcement Learning (TGRL), which partitions each prompt's rollout group into low- and high-temperature subsets, estimates prompt-level exploration gain from their reward contrast, and allocates that gain as token-level credit using Jensen–Shannon divergence between the temperature-scaled next-token distributions induced by the same logits; across 11 benchmarks, TGRL improves the six-benchmark math average by 1.6% at 32B, raises CodeForces rating by 196.7 points and LiveCodeBench Pass@16 by 4.4%, and improves ALFWorld/WebShop success rates by 6.3%/4.9%, reaching equivalent accuracy up to 36% faster than strong RLVR baselines without expanding the rollout budget.
The work proposes Temperature-Grouped Reinforcement Learning (TGRL), which partitions each prompt's rollout group into low- and high-temperature subsets, estimates prompt-level exploration gain from their reward contrast, and allocates that gain as token-level credit using Jensen–Shannon divergence between the temperature-scaled next-token distributions induced by the same logits; across 11 benchmarks, TGRL improves the six-benchmark math average by 1.6% at 32B, raises CodeForces rating by 196.7 points and LiveCodeBench Pass@16 by 4.4%, and improves ALFWorld/WebShop success rates by 6.3%/4.9%, reaching equivalent accuracy up to 36% faster than strong RLVR baselines without expanding the rollout budget.
arXiv The work proposes a preference-guided adaptation framework for open-vocabulary semantic segmentation that converts the segmentation differences different prompt templates produce on the same image (prompt disagreement) into binary preference supervision, adapting models with Region-Localized Preference Optimization (RLPO) plus consistency regularization in a streaming single-step setting, achieving consistent gains on the five domain groups of the MESS benchmark across SAN and CAT-Seg backbones (ViT-B/16 and ViT-L/14) without any pixel-level annotation and remaining effective under noisy preferences.
The work proposes a preference-guided adaptation framework for open-vocabulary semantic segmentation that converts the segmentation differences different prompt templates produce on the same image (prompt disagreement) into binary preference supervision, adapting models with Region-Localized Preference Optimization (RLPO) plus consistency regularization in a streaming single-step setting, achieving consistent gains on the five domain groups of the MESS benchmark across SAN and CAT-Seg backbones (ViT-B/16 and ViT-L/14) without any pixel-level annotation and remaining effective under noisy preferences.
The work proposes a preference-guided adaptation framework for open-vocabulary semantic segmentation that converts the segmentation differences different prompt templates produce on the same image (prompt disagreement) into binary preference supervision, adapting models with Region-Localized Preference Optimization (RLPO) plus consistency regularization in a streaming single-step setting, achieving consistent gains on the five domain groups of the MESS benchmark across SAN and CAT-Seg backbones (ViT-B/16 and ViT-L/14) without any pixel-level annotation and remaining effective under noisy preferences.
The work proposes a preference-guided adaptation framework for open-vocabulary semantic segmentation that converts the segmentation differences different prompt templates produce on the same image (prompt disagreement) into binary preference supervision, adapting models with Region-Localized Preference Optimization (RLPO) plus consistency regularization in a streaming single-step setting, achieving consistent gains on the five domain groups of the MESS benchmark across SAN and CAT-Seg backbones (ViT-B/16 and ViT-L/14) without any pixel-level annotation and remaining effective under noisy preferences.
arXiv The work proposes GRAFT, a framework that in Reinforcement Learning with Verifiable Rewards (RLVR) replaces a receiver's all-fail rollout groups with peer groups containing both successful and unsuccessful responses, controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping; across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT improves both models over GRPO at the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points, with 1.8 points on average preserved when reusing stored peer trajectories.
The work proposes GRAFT, a framework that in Reinforcement Learning with Verifiable Rewards (RLVR) replaces a receiver's all-fail rollout groups with peer groups containing both successful and unsuccessful responses, controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping; across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT improves both models over GRPO at the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points, with 1.8 points on average preserved when reusing stored peer trajectories.
The work proposes GRAFT, a framework that in Reinforcement Learning with Verifiable Rewards (RLVR) replaces a receiver's all-fail rollout groups with peer groups containing both successful and unsuccessful responses, controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping; across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT improves both models over GRPO at the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points, with 1.8 points on average preserved when reusing stored peer trajectories.
The work proposes GRAFT, a framework that in Reinforcement Learning with Verifiable Rewards (RLVR) replaces a receiver's all-fail rollout groups with peer groups containing both successful and unsuccessful responses, controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping; across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT improves both models over GRPO at the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points, with 1.8 points on average preserved when reusing stored peer trajectories.
arXiv LongLive-Plug introduces a once-for-all distillation framework that distills single-pass classifier-free guidance, four-step sampling, and long-context error correction into reusable LoRAs on a base model, enabling training-free plug-and-play deployment to 54 downstream video models across three backbone families, with results on SCOPE and Wan2.2-Fun-5B-Control that beat naive four-step sampling and approach per-target distillation.
LongLive-Plug introduces a once-for-all distillation framework that distills single-pass classifier-free guidance, four-step sampling, and long-context error correction into reusable LoRAs on a base model, enabling training-free plug-and-play deployment to 54 downstream video models across three backbone families, with results on SCOPE and Wan2.2-Fun-5B-Control that beat naive four-step sampling and approach per-target distillation.
LongLive-Plug introduces a once-for-all distillation framework that distills single-pass classifier-free guidance, four-step sampling, and long-context error correction into reusable LoRAs on a base model, enabling training-free plug-and-play deployment to 54 downstream video models across three backbone families, with results on SCOPE and Wan2.2-Fun-5B-Control that beat naive four-step sampling and approach per-target distillation.
LongLive-Plug introduces a once-for-all distillation framework that distills single-pass classifier-free guidance, four-step sampling, and long-context error correction into reusable LoRAs on a base model, enabling training-free plug-and-play deployment to 54 downstream video models across three backbone families, with results on SCOPE and Wan2.2-Fun-5B-Control that beat naive four-step sampling and approach per-target distillation.
arXiv Porcedda built Sys1Cal-v1, a synthetic dataset of 365 True/False items derived from 92 probability problems whose proposition probabilities are known by construction, and used total variation distance and distributional overlap (soft accuracy) to evaluate Jev's Noul, Choice, and Score primitives plus the open-source baseline SemIf, finding that Jev's Choice probabilities are systematically distorted and that inverting a latent third truth value of uncertainty estimated from Score expectations raises median Choice soft accuracy from 0.771 to 0.978.
Porcedda built Sys1Cal-v1, a synthetic dataset of 365 True/False items derived from 92 probability problems whose proposition probabilities are known by construction, and used total variation distance and distributional overlap (soft accuracy) to evaluate Jev's Noul, Choice, and Score primitives plus the open-source baseline SemIf, finding that Jev's Choice probabilities are systematically distorted and that inverting a latent third truth value of uncertainty estimated from Score expectations raises median Choice soft accuracy from 0.771 to 0.978.
Porcedda built Sys1Cal-v1, a synthetic dataset of 365 True/False items derived from 92 probability problems whose proposition probabilities are known by construction, and used total variation distance and distributional overlap (soft accuracy) to evaluate Jev's Noul, Choice, and Score primitives plus the open-source baseline SemIf, finding that Jev's Choice probabilities are systematically distorted and that inverting a latent third truth value of uncertainty estimated from Score expectations raises median Choice soft accuracy from 0.771 to 0.978.
Porcedda built Sys1Cal-v1, a synthetic dataset of 365 True/False items derived from 92 probability problems whose proposition probabilities are known by construction, and used total variation distance and distributional overlap (soft accuracy) to evaluate Jev's Noul, Choice, and Score primitives plus the open-source baseline SemIf, finding that Jev's Choice probabilities are systematically distorted and that inverting a latent third truth value of uncertainty estimated from Score expectations raises median Choice soft accuracy from 0.771 to 0.978.
arXiv Using controlled paired image-to-image (I2I) generation and image-to-text (I2T) understanding tasks on the unified multimodal model BAGEL, the authors find that a recipe of I2I training that updates shared parameters followed by I2T finetuning improves understanding with gains that grow as I2I data increases, build OmniTaskonomy, a taxonomy spanning 19 I2I tasks and 25 understanding capabilities, produce a selective task-dependent generation-to-understanding transfer map, and find that stronger gradient alignment between generation and understanding is associated with larger downstream transfer gains.
Using controlled paired image-to-image (I2I) generation and image-to-text (I2T) understanding tasks on the unified multimodal model BAGEL, the authors find that a recipe of I2I training that updates shared parameters followed by I2T finetuning improves understanding with gains that grow as I2I data increases, build OmniTaskonomy, a taxonomy spanning 19 I2I tasks and 25 understanding capabilities, produce a selective task-dependent generation-to-understanding transfer map, and find that stronger gradient alignment between generation and understanding is associated with larger downstream transfer gains.
Using controlled paired image-to-image (I2I) generation and image-to-text (I2T) understanding tasks on the unified multimodal model BAGEL, the authors find that a recipe of I2I training that updates shared parameters followed by I2T finetuning improves understanding with gains that grow as I2I data increases, build OmniTaskonomy, a taxonomy spanning 19 I2I tasks and 25 understanding capabilities, produce a selective task-dependent generation-to-understanding transfer map, and find that stronger gradient alignment between generation and understanding is associated with larger downstream transfer gains.
Using controlled paired image-to-image (I2I) generation and image-to-text (I2T) understanding tasks on the unified multimodal model BAGEL, the authors find that a recipe of I2I training that updates shared parameters followed by I2T finetuning improves understanding with gains that grow as I2I data increases, build OmniTaskonomy, a taxonomy spanning 19 I2I tasks and 25 understanding capabilities, produce a selective task-dependent generation-to-understanding transfer map, and find that stronger gradient alignment between generation and understanding is associated with larger downstream transfer gains.
arXiv The work develops a bias–variance theory for low-rank bilinear scoring in dense retrieval, proving that dual projections have lower risk exactly when squared directional signal exceeds the estimation cost of their extra degrees of freedom, and uses it to build the cross-fitted CARS selector, which reaches 90.1% mean geometry-selection accuracy and reduces held-out regret by 49–96% relative to the better fixed geometry across multiple datasets and embedding models.
The work develops a bias–variance theory for low-rank bilinear scoring in dense retrieval, proving that dual projections have lower risk exactly when squared directional signal exceeds the estimation cost of their extra degrees of freedom, and uses it to build the cross-fitted CARS selector, which reaches 90.1% mean geometry-selection accuracy and reduces held-out regret by 49–96% relative to the better fixed geometry across multiple datasets and embedding models.
The work develops a bias–variance theory for low-rank bilinear scoring in dense retrieval, proving that dual projections have lower risk exactly when squared directional signal exceeds the estimation cost of their extra degrees of freedom, and uses it to build the cross-fitted CARS selector, which reaches 90.1% mean geometry-selection accuracy and reduces held-out regret by 49–96% relative to the better fixed geometry across multiple datasets and embedding models.
The work develops a bias–variance theory for low-rank bilinear scoring in dense retrieval, proving that dual projections have lower risk exactly when squared directional signal exceeds the estimation cost of their extra degrees of freedom, and uses it to build the cross-fitted CARS selector, which reaches 90.1% mean geometry-selection accuracy and reduces held-out regret by 49–96% relative to the better fixed geometry across multiple datasets and embedding models.
arXiv Raven is an open-source multi-agent ecosystem that treats each executable model–harness pair as a composable unit of intelligence: a Host Agent decomposes goals into a directed acyclic graph and matches subtasks to registered specialists, HarnessBank-style evolution, EverOS memory, and Skill Forge reuse experience across tasks, and the paper gives sufficient conditions under which composition expands reliable task coverage under a shared resource budget, reporting gains over Claude Code, Hermes Agent, OpenClaw and other baselines on the MAOB planning benchmark, four native specialist domains, frozen-backbone harness evolution, and fixed skill-library reuse.
Raven is an open-source multi-agent ecosystem that treats each executable model–harness pair as a composable unit of intelligence: a Host Agent decomposes goals into a directed acyclic graph and matches subtasks to registered specialists, HarnessBank-style evolution, EverOS memory, and Skill Forge reuse experience across tasks, and the paper gives sufficient conditions under which composition expands reliable task coverage under a shared resource budget, reporting gains over Claude Code, Hermes Agent, OpenClaw and other baselines on the MAOB planning benchmark, four native specialist domains, frozen-backbone harness evolution, and fixed skill-library reuse.
Raven is an open-source multi-agent ecosystem that treats each executable model–harness pair as a composable unit of intelligence: a Host Agent decomposes goals into a directed acyclic graph and matches subtasks to registered specialists, HarnessBank-style evolution, EverOS memory, and Skill Forge reuse experience across tasks, and the paper gives sufficient conditions under which composition expands reliable task coverage under a shared resource budget, reporting gains over Claude Code, Hermes Agent, OpenClaw and other baselines on the MAOB planning benchmark, four native specialist domains, frozen-backbone harness evolution, and fixed skill-library reuse.
Raven is an open-source multi-agent ecosystem that treats each executable model–harness pair as a composable unit of intelligence: a Host Agent decomposes goals into a directed acyclic graph and matches subtasks to registered specialists, HarnessBank-style evolution, EverOS memory, and Skill Forge reuse experience across tasks, and the paper gives sufficient conditions under which composition expands reliable task coverage under a shared resource budget, reporting gains over Claude Code, Hermes Agent, OpenClaw and other baselines on the MAOB planning benchmark, four native specialist domains, frozen-backbone harness evolution, and fixed skill-library reuse.
arXiv The work names and systematically characterizes phase sensitivity in models with chunked KV-cache compression: the same information is easier or harder to retrieve depending on its phase relative to compression-window boundaries, with long-context retrieval accuracy in the DeepSeek-V4 family differing by up to 40.2 percentage points across phases at 128K context; the authors reproduce the effect by pretraining 23 compressed models from scratch (Qwen3-0.6B backbone, 100B tokens), use KV-head and layer mean-replacement knockouts to reveal phase specialization, and prove in an idealized induction model that gradient flow drives compression gates toward static preferences, showing that average benchmark scores can conceal periodic positional failures.
The work names and systematically characterizes phase sensitivity in models with chunked KV-cache compression: the same information is easier or harder to retrieve depending on its phase relative to compression-window boundaries, with long-context retrieval accuracy in the DeepSeek-V4 family differing by up to 40.2 percentage points across phases at 128K context; the authors reproduce the effect by pretraining 23 compressed models from scratch (Qwen3-0.6B backbone, 100B tokens), use KV-head and layer mean-replacement knockouts to reveal phase specialization, and prove in an idealized induction model that gradient flow drives compression gates toward static preferences, showing that average benchmark scores can conceal periodic positional failures.
The work names and systematically characterizes phase sensitivity in models with chunked KV-cache compression: the same information is easier or harder to retrieve depending on its phase relative to compression-window boundaries, with long-context retrieval accuracy in the DeepSeek-V4 family differing by up to 40.2 percentage points across phases at 128K context; the authors reproduce the effect by pretraining 23 compressed models from scratch (Qwen3-0.6B backbone, 100B tokens), use KV-head and layer mean-replacement knockouts to reveal phase specialization, and prove in an idealized induction model that gradient flow drives compression gates toward static preferences, showing that average benchmark scores can conceal periodic positional failures.
The work names and systematically characterizes phase sensitivity in models with chunked KV-cache compression: the same information is easier or harder to retrieve depending on its phase relative to compression-window boundaries, with long-context retrieval accuracy in the DeepSeek-V4 family differing by up to 40.2 percentage points across phases at 128K context; the authors reproduce the effect by pretraining 23 compressed models from scratch (Qwen3-0.6B backbone, 100B tokens), use KV-head and layer mean-replacement knockouts to reveal phase specialization, and prove in an idealized induction model that gradient flow drives compression gates toward static preferences, showing that average benchmark scores can conceal periodic positional failures.
arXiv The work proposes Marathoner, an autonomous agentic model trained through a post-training pipeline of ultra-long-horizon task synthesis, rejection sampling finetuning, and reinforcement learning, achieving consistent and substantial improvements over the base Qwen3.5-9B across 5 ultra-long-horizon benchmarks and surpassing strong proprietary models on some, with analysis showing it can execute for 10+ hours and make 1000+ tool calls on highly challenging tasks.
The work proposes Marathoner, an autonomous agentic model trained through a post-training pipeline of ultra-long-horizon task synthesis, rejection sampling finetuning, and reinforcement learning, achieving consistent and substantial improvements over the base Qwen3.5-9B across 5 ultra-long-horizon benchmarks and surpassing strong proprietary models on some, with analysis showing it can execute for 10+ hours and make 1000+ tool calls on highly challenging tasks.
The work proposes Marathoner, an autonomous agentic model trained through a post-training pipeline of ultra-long-horizon task synthesis, rejection sampling finetuning, and reinforcement learning, achieving consistent and substantial improvements over the base Qwen3.5-9B across 5 ultra-long-horizon benchmarks and surpassing strong proprietary models on some, with analysis showing it can execute for 10+ hours and make 1000+ tool calls on highly challenging tasks.
The work proposes Marathoner, an autonomous agentic model trained through a post-training pipeline of ultra-long-horizon task synthesis, rejection sampling finetuning, and reinforcement learning, achieving consistent and substantial improvements over the base Qwen3.5-9B across 5 ultra-long-horizon benchmarks and surpassing strong proprietary models on some, with analysis showing it can execute for 10+ hours and make 1000+ tool calls on highly challenging tasks.
arXiv Across about 4.7 million natural language autoencoder (NLA) explanations on four models and four auditing datasets, the study compares thirteen position-scoring signals computable in a single forward pass against a ranker trained only on chat structure, finding that chat structure beats the best individual signal in twelve of fourteen model-dataset cells, that explaining just 5% of positions retains nearly all of the success rate from explaining every position on three of four datasets, and that pretrained verbalizers recover words a fine-tuned model learned to conceal without additional verbalizer training.
Across about 4.7 million natural language autoencoder (NLA) explanations on four models and four auditing datasets, the study compares thirteen position-scoring signals computable in a single forward pass against a ranker trained only on chat structure, finding that chat structure beats the best individual signal in twelve of fourteen model-dataset cells, that explaining just 5% of positions retains nearly all of the success rate from explaining every position on three of four datasets, and that pretrained verbalizers recover words a fine-tuned model learned to conceal without additional verbalizer training.
Across about 4.7 million natural language autoencoder (NLA) explanations on four models and four auditing datasets, the study compares thirteen position-scoring signals computable in a single forward pass against a ranker trained only on chat structure, finding that chat structure beats the best individual signal in twelve of fourteen model-dataset cells, that explaining just 5% of positions retains nearly all of the success rate from explaining every position on three of four datasets, and that pretrained verbalizers recover words a fine-tuned model learned to conceal without additional verbalizer training.
Across about 4.7 million natural language autoencoder (NLA) explanations on four models and four auditing datasets, the study compares thirteen position-scoring signals computable in a single forward pass against a ranker trained only on chat structure, finding that chat structure beats the best individual signal in twelve of fourteen model-dataset cells, that explaining just 5% of positions retains nearly all of the success rate from explaining every position on three of four datasets, and that pretrained verbalizers recover words a fine-tuned model learned to conceal without additional verbalizer training.
arXiv The work introduces AutoDataBench, which turns "can an agent write a training task that a data pipeline would accept sample by sample" into an evaluation target: given an original benchmark task and a record of the target model attempting it, an author agent must write a new task for the same suite, and the score multiplies a gate, whether the target model's pass rate lands in a band, and coverage of a hidden rubric of behavioural modes; across three benchmarks of executable agent tasks, no agent evaluated scores above 20 out of 100 at the default 45-minute budget, and giving the strongest agent four times as long raises the score substantially while the cost of one usable task stays almost unchanged.
The work introduces AutoDataBench, which turns "can an agent write a training task that a data pipeline would accept sample by sample" into an evaluation target: given an original benchmark task and a record of the target model attempting it, an author agent must write a new task for the same suite, and the score multiplies a gate, whether the target model's pass rate lands in a band, and coverage of a hidden rubric of behavioural modes; across three benchmarks of executable agent tasks, no agent evaluated scores above 20 out of 100 at the default 45-minute budget, and giving the strongest agent four times as long raises the score substantially while the cost of one usable task stays almost unchanged.
The work introduces AutoDataBench, which turns "can an agent write a training task that a data pipeline would accept sample by sample" into an evaluation target: given an original benchmark task and a record of the target model attempting it, an author agent must write a new task for the same suite, and the score multiplies a gate, whether the target model's pass rate lands in a band, and coverage of a hidden rubric of behavioural modes; across three benchmarks of executable agent tasks, no agent evaluated scores above 20 out of 100 at the default 45-minute budget, and giving the strongest agent four times as long raises the score substantially while the cost of one usable task stays almost unchanged.
The work introduces AutoDataBench, which turns "can an agent write a training task that a data pipeline would accept sample by sample" into an evaluation target: given an original benchmark task and a record of the target model attempting it, an author agent must write a new task for the same suite, and the score multiplies a gate, whether the target model's pass rate lands in a band, and coverage of a hidden rubric of behavioural modes; across three benchmarks of executable agent tasks, no agent evaluated scores above 20 out of 100 at the default 45-minute budget, and giving the strongest agent four times as long raises the score substantially while the cost of one usable task stays almost unchanged.