arXiv The authors introduce Box2-Bench, which holds the model and task fixed while varying workflow availability and reliability across five matched conditions (none, good, bad, partial, mixed), finding that frontier models such as Gemini 3.7 Flash, DeepSeek V4 Flash, and GLM 5.2 benefit from good workflows in 8 of 12 model-task pairs but are hurt by bad workflows in all 12 (drops of 3.3 to 43.3 points), and that mixed workflows underperform their matched partial workflows in 10 of 12 pairs; training two open-weight models on bad workflows only, counterfactual SFT cuts the bad-workflow penalty from 20.0 to 6.7 points on AIME and from 9.8 to 0.6 points on WebShop while weakening use of held-out good workflows, and outcome-based RL restores AIME utilization from -3.3 to +6.
The authors introduce Box2-Bench, which holds the model and task fixed while varying workflow availability and reliability across five matched conditions (none, good, bad, partial, mixed), finding that frontier models such as Gemini 3.7 Flash, DeepSeek V4 Flash, and GLM 5.2 benefit from good workflows in 8 of 12 model-task pairs but are hurt by bad workflows in all 12 (drops of 3.3 to 43.3 points), and that mixed workflows underperform their matched partial workflows in 10 of 12 pairs; training two open-weight models on bad workflows only, counterfactual SFT cuts the bad-workflow penalty from 20.0 to 6.7 points on AIME and from 9.8 to 0.6 points on WebShop while weakening use of held-out good workflows, and outcome-based RL restores AIME utilization from -3.3 to +6.
The authors introduce Box2-Bench, which holds the model and task fixed while varying workflow availability and reliability across five matched conditions (none, good, bad, partial, mixed), finding that frontier models such as Gemini 3.7 Flash, DeepSeek V4 Flash, and GLM 5.2 benefit from good workflows in 8 of 12 model-task pairs but are hurt by bad workflows in all 12 (drops of 3.3 to 43.3 points), and that mixed workflows underperform their matched partial workflows in 10 of 12 pairs; training two open-weight models on bad workflows only, counterfactual SFT cuts the bad-workflow penalty from 20.0 to 6.7 points on AIME and from 9.8 to 0.6 points on WebShop while weakening use of held-out good workflows, and outcome-based RL restores AIME utilization from -3.3 to +6.
The authors introduce Box2-Bench, which holds the model and task fixed while varying workflow availability and reliability across five matched conditions (none, good, bad, partial, mixed), finding that frontier models such as Gemini 3.7 Flash, DeepSeek V4 Flash, and GLM 5.2 benefit from good workflows in 8 of 12 model-task pairs but are hurt by bad workflows in all 12 (drops of 3.3 to 43.3 points), and that mixed workflows underperform their matched partial workflows in 10 of 12 pairs; training two open-weight models on bad workflows only, counterfactual SFT cuts the bad-workflow penalty from 20.0 to 6.7 points on AIME and from 9.8 to 0.6 points on WebShop while weakening use of held-out good workflows, and outcome-based RL restores AIME utilization from -3.3 to +6.
arXiv The work introduces the Agent Error Dataset (AED), comprising 50,228 error–diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families and 23 policy models, and a five-stage Agentic Error-to-Training (AET) pipeline that links natural failures to diagnoses, proposed corrections and optional replay evidence; across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1% (a 32.7 percentage-point gain), full-diagnosis fine-tuning on a separately frozen diagnosis release raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6% (averaged over three seeds on a 943-case holdout), and in a single-seed actor-training comparison action-only repair training scores 6.
The work introduces the Agent Error Dataset (AED), comprising 50,228 error–diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families and 23 policy models, and a five-stage Agentic Error-to-Training (AET) pipeline that links natural failures to diagnoses, proposed corrections and optional replay evidence; across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1% (a 32.7 percentage-point gain), full-diagnosis fine-tuning on a separately frozen diagnosis release raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6% (averaged over three seeds on a 943-case holdout), and in a single-seed actor-training comparison action-only repair training scores 6.
The work introduces the Agent Error Dataset (AED), comprising 50,228 error–diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families and 23 policy models, and a five-stage Agentic Error-to-Training (AET) pipeline that links natural failures to diagnoses, proposed corrections and optional replay evidence; across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1% (a 32.7 percentage-point gain), full-diagnosis fine-tuning on a separately frozen diagnosis release raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6% (averaged over three seeds on a 943-case holdout), and in a single-seed actor-training comparison action-only repair training scores 6.
The work introduces the Agent Error Dataset (AED), comprising 50,228 error–diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families and 23 policy models, and a five-stage Agentic Error-to-Training (AET) pipeline that links natural failures to diagnoses, proposed corrections and optional replay evidence; across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1% (a 32.7 percentage-point gain), full-diagnosis fine-tuning on a separately frozen diagnosis release raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6% (averaged over three seeds on a 943-case holdout), and in a single-seed actor-training comparison action-only repair training scores 6.
arXiv The work introduces EvolvingNav, which builds a persistence–relocation belief from timestamped 3D object histories and closes the loop with an event-driven predict–observe–replan filter so an agent can infer where a target is when targets may move before the query or during navigation, and it releases EvoWorld-Bench with 54 scenes and 803,680 tasks, reporting improved navigation success and search efficiency over the evaluated baselines in simulation and real-robot experiments.
The work introduces EvolvingNav, which builds a persistence–relocation belief from timestamped 3D object histories and closes the loop with an event-driven predict–observe–replan filter so an agent can infer where a target is when targets may move before the query or during navigation, and it releases EvoWorld-Bench with 54 scenes and 803,680 tasks, reporting improved navigation success and search efficiency over the evaluated baselines in simulation and real-robot experiments.
The work introduces EvolvingNav, which builds a persistence–relocation belief from timestamped 3D object histories and closes the loop with an event-driven predict–observe–replan filter so an agent can infer where a target is when targets may move before the query or during navigation, and it releases EvoWorld-Bench with 54 scenes and 803,680 tasks, reporting improved navigation success and search efficiency over the evaluated baselines in simulation and real-robot experiments.
The work introduces EvolvingNav, which builds a persistence–relocation belief from timestamped 3D object histories and closes the loop with an event-driven predict–observe–replan filter so an agent can infer where a target is when targets may move before the query or during navigation, and it releases EvoWorld-Bench with 54 scenes and 803,680 tasks, reporting improved navigation success and search efficiency over the evaluated baselines in simulation and real-robot experiments.
arXiv LANTERN trains a linear classifier on Qwen3-32B activations to rank 50 million candidate relations among the 10,000 most-referenced OEIS sequences, then applies staged filtering, hypothesis generation, executable verification and analytical checking to produce 62 verified relations without an existing OEIS cross-reference; a content screen retained 13, nine of them informative or insightful, four not found in the OEIS or a targeted literature search, with the end-to-end process taking under 8 hours.
LANTERN trains a linear classifier on Qwen3-32B activations to rank 50 million candidate relations among the 10,000 most-referenced OEIS sequences, then applies staged filtering, hypothesis generation, executable verification and analytical checking to produce 62 verified relations without an existing OEIS cross-reference; a content screen retained 13, nine of them informative or insightful, four not found in the OEIS or a targeted literature search, with the end-to-end process taking under 8 hours.
LANTERN trains a linear classifier on Qwen3-32B activations to rank 50 million candidate relations among the 10,000 most-referenced OEIS sequences, then applies staged filtering, hypothesis generation, executable verification and analytical checking to produce 62 verified relations without an existing OEIS cross-reference; a content screen retained 13, nine of them informative or insightful, four not found in the OEIS or a targeted literature search, with the end-to-end process taking under 8 hours.
LANTERN trains a linear classifier on Qwen3-32B activations to rank 50 million candidate relations among the 10,000 most-referenced OEIS sequences, then applies staged filtering, hypothesis generation, executable verification and analytical checking to produce 62 verified relations without an existing OEIS cross-reference; a content screen retained 13, nine of them informative or insightful, four not found in the OEIS or a targeted literature search, with the end-to-end process taking under 8 hours.
arXiv Peking University and Singapore University of Technology and Design propose DC-SAE, a decoupled compact semantic autoencoder that pairs a frozen semantic encoder for a generation-friendly compact latent with a trainable pixel encoder for low-level detail, achieving 29.79 PSNR and 3.37 gFID on ImageNet at 32x spatial compression, improving over DC-AE by 13.5% in PSNR and 54.9% in gFID with a 4.41x speedup, while a B-parameter DiT reaches 0.84 on GenEval and 86.007 on DPG-Bench.
Peking University and Singapore University of Technology and Design propose DC-SAE, a decoupled compact semantic autoencoder that pairs a frozen semantic encoder for a generation-friendly compact latent with a trainable pixel encoder for low-level detail, achieving 29.79 PSNR and 3.37 gFID on ImageNet at 32x spatial compression, improving over DC-AE by 13.5% in PSNR and 54.9% in gFID with a 4.41x speedup, while a B-parameter DiT reaches 0.84 on GenEval and 86.007 on DPG-Bench.
Peking University and Singapore University of Technology and Design propose DC-SAE, a decoupled compact semantic autoencoder that pairs a frozen semantic encoder for a generation-friendly compact latent with a trainable pixel encoder for low-level detail, achieving 29.79 PSNR and 3.37 gFID on ImageNet at 32x spatial compression, improving over DC-AE by 13.5% in PSNR and 54.9% in gFID with a 4.41x speedup, while a B-parameter DiT reaches 0.84 on GenEval and 86.007 on DPG-Bench.
Peking University and Singapore University of Technology and Design propose DC-SAE, a decoupled compact semantic autoencoder that pairs a frozen semantic encoder for a generation-friendly compact latent with a trainable pixel encoder for low-level detail, achieving 29.79 PSNR and 3.37 gFID on ImageNet at 32x spatial compression, improving over DC-AE by 13.5% in PSNR and 54.9% in gFID with a 4.41x speedup, while a B-parameter DiT reaches 0.84 on GenEval and 86.007 on DPG-Bench.
arXiv The work proposes RIDE (RL-Induced Direction Extrapolation): on student-generated trajectories it computes the layerwise hidden-state residual between an RL-trained teacher and its pre-RL checkpoint, then regresses the student's hidden states toward targets displaced beyond the teacher along that residual; across four base/RL-teacher pairs (R1-Distill-1.5B, Qwen3-4B, Llama-3.2-3B, Phi-4-mini) RIDE's mean Avg@16 on AIME24, AIME25 and AIMO approaches or exceeds the RL-trained teacher on every pair, the only compared method whose mean does so, and it consistently outperforms output-space extrapolation.
The work proposes RIDE (RL-Induced Direction Extrapolation): on student-generated trajectories it computes the layerwise hidden-state residual between an RL-trained teacher and its pre-RL checkpoint, then regresses the student's hidden states toward targets displaced beyond the teacher along that residual; across four base/RL-teacher pairs (R1-Distill-1.5B, Qwen3-4B, Llama-3.2-3B, Phi-4-mini) RIDE's mean Avg@16 on AIME24, AIME25 and AIMO approaches or exceeds the RL-trained teacher on every pair, the only compared method whose mean does so, and it consistently outperforms output-space extrapolation.
The work proposes RIDE (RL-Induced Direction Extrapolation): on student-generated trajectories it computes the layerwise hidden-state residual between an RL-trained teacher and its pre-RL checkpoint, then regresses the student's hidden states toward targets displaced beyond the teacher along that residual; across four base/RL-teacher pairs (R1-Distill-1.5B, Qwen3-4B, Llama-3.2-3B, Phi-4-mini) RIDE's mean Avg@16 on AIME24, AIME25 and AIMO approaches or exceeds the RL-trained teacher on every pair, the only compared method whose mean does so, and it consistently outperforms output-space extrapolation.
The work proposes RIDE (RL-Induced Direction Extrapolation): on student-generated trajectories it computes the layerwise hidden-state residual between an RL-trained teacher and its pre-RL checkpoint, then regresses the student's hidden states toward targets displaced beyond the teacher along that residual; across four base/RL-teacher pairs (R1-Distill-1.5B, Qwen3-4B, Llama-3.2-3B, Phi-4-mini) RIDE's mean Avg@16 on AIME24, AIME25 and AIMO approaches or exceeds the RL-trained teacher on every pair, the only compared method whose mean does so, and it consistently outperforms output-space extrapolation.
arXiv The authors introduce CUA-SWE, a benchmark, interactive development environment, and evaluation pipeline spanning 105 tasks across Web, Game, DevOps, and Mobile, in which an agent edits code, runs commands, operates the running software, and inspects screenshots within a single task, with task-specific deterministic tests deciding the outcome; comparing code-only and hybrid CUA conditions across frontier models, GPT-6-Astra reaches the highest four-domain mean of 59.9%, and every frontier model in the four-domain comparison achieves higher aggregate success with hybrid access, with gains concentrated on tasks whose requirements must be recovered from application materials such as drawings, service contracts, or graphical reference cards.
The authors introduce CUA-SWE, a benchmark, interactive development environment, and evaluation pipeline spanning 105 tasks across Web, Game, DevOps, and Mobile, in which an agent edits code, runs commands, operates the running software, and inspects screenshots within a single task, with task-specific deterministic tests deciding the outcome; comparing code-only and hybrid CUA conditions across frontier models, GPT-6-Astra reaches the highest four-domain mean of 59.9%, and every frontier model in the four-domain comparison achieves higher aggregate success with hybrid access, with gains concentrated on tasks whose requirements must be recovered from application materials such as drawings, service contracts, or graphical reference cards.
The authors introduce CUA-SWE, a benchmark, interactive development environment, and evaluation pipeline spanning 105 tasks across Web, Game, DevOps, and Mobile, in which an agent edits code, runs commands, operates the running software, and inspects screenshots within a single task, with task-specific deterministic tests deciding the outcome; comparing code-only and hybrid CUA conditions across frontier models, GPT-6-Astra reaches the highest four-domain mean of 59.9%, and every frontier model in the four-domain comparison achieves higher aggregate success with hybrid access, with gains concentrated on tasks whose requirements must be recovered from application materials such as drawings, service contracts, or graphical reference cards.
The authors introduce CUA-SWE, a benchmark, interactive development environment, and evaluation pipeline spanning 105 tasks across Web, Game, DevOps, and Mobile, in which an agent edits code, runs commands, operates the running software, and inspects screenshots within a single task, with task-specific deterministic tests deciding the outcome; comparing code-only and hybrid CUA conditions across frontier models, GPT-6-Astra reaches the highest four-domain mean of 59.9%, and every frontier model in the four-domain comparison achieves higher aggregate success with hybrid access, with gains concentrated on tasks whose requirements must be recovered from application materials such as drawings, service contracts, or graphical reference cards.
arXiv By comparing each context's intermediate residual states with its own final state and an empirical bank of final states from other contexts, the study finds across Gemma-2B, Gemma-7B, Qwen2.5-1.5B, Qwen2.5-7B, Mistral-7B, and Llama-3-8B that the own endpoint becomes preferable to the average alternative at the earliest measured layer while many individual endpoints remain closer, that competitor sets generally shrink with depth yet show both entries and exits, and that directional alignment can improve while Euclidean distance to the final state changes little; the authors add a high-dimensional model separating norm, alignment, and endpoint geometry and prove that a straight path toward the own endpoint cannot introduce new competitors.
By comparing each context's intermediate residual states with its own final state and an empirical bank of final states from other contexts, the study finds across Gemma-2B, Gemma-7B, Qwen2.5-1.5B, Qwen2.5-7B, Mistral-7B, and Llama-3-8B that the own endpoint becomes preferable to the average alternative at the earliest measured layer while many individual endpoints remain closer, that competitor sets generally shrink with depth yet show both entries and exits, and that directional alignment can improve while Euclidean distance to the final state changes little; the authors add a high-dimensional model separating norm, alignment, and endpoint geometry and prove that a straight path toward the own endpoint cannot introduce new competitors.
By comparing each context's intermediate residual states with its own final state and an empirical bank of final states from other contexts, the study finds across Gemma-2B, Gemma-7B, Qwen2.5-1.5B, Qwen2.5-7B, Mistral-7B, and Llama-3-8B that the own endpoint becomes preferable to the average alternative at the earliest measured layer while many individual endpoints remain closer, that competitor sets generally shrink with depth yet show both entries and exits, and that directional alignment can improve while Euclidean distance to the final state changes little; the authors add a high-dimensional model separating norm, alignment, and endpoint geometry and prove that a straight path toward the own endpoint cannot introduce new competitors.
By comparing each context's intermediate residual states with its own final state and an empirical bank of final states from other contexts, the study finds across Gemma-2B, Gemma-7B, Qwen2.5-1.5B, Qwen2.5-7B, Mistral-7B, and Llama-3-8B that the own endpoint becomes preferable to the average alternative at the earliest measured layer while many individual endpoints remain closer, that competitor sets generally shrink with depth yet show both entries and exits, and that directional alignment can improve while Euclidean distance to the final state changes little; the authors add a high-dimensional model separating norm, alignment, and endpoint geometry and prove that a straight path toward the own endpoint cannot introduce new competitors.
arXiv CoEvoWhen introduces a policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from a VLM's agentic reasoning trajectories into a reusable skill without updating model parameters, improving ultra-long video temporal grounding accuracy and reducing inference visual token cost across five benchmarks and three VLMs, with the evolved skill transferring to general long-video QA without additional task-specific evolution.
CoEvoWhen introduces a policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from a VLM's agentic reasoning trajectories into a reusable skill without updating model parameters, improving ultra-long video temporal grounding accuracy and reducing inference visual token cost across five benchmarks and three VLMs, with the evolved skill transferring to general long-video QA without additional task-specific evolution.
CoEvoWhen introduces a policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from a VLM's agentic reasoning trajectories into a reusable skill without updating model parameters, improving ultra-long video temporal grounding accuracy and reducing inference visual token cost across five benchmarks and three VLMs, with the evolved skill transferring to general long-video QA without additional task-specific evolution.
CoEvoWhen introduces a policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from a VLM's agentic reasoning trajectories into a reusable skill without updating model parameters, improving ultra-long video temporal grounding accuracy and reducing inference visual token cost across five benchmarks and three VLMs, with the evolved skill transferring to general long-video QA without additional task-specific evolution.
arXiv The work introduces AdviSD, in which a small trainable advisor (Qwen3-8B) steers frozen frontier executors (Gemini 3.7 Flash, Claude Sonnet 4.6) with natural-language advice by pairing outcome-based GRPO with selective feedback-conditioned self-distillation: reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of that difference to choose which decisions to supervise, reaching 4.2–6.4 percentage points above advisor-GRPO on BFCL-v3 and 3.9–5.1 score points above it on EnvScaler while beating matched-count random selection and a no-gate variant.
The work introduces AdviSD, in which a small trainable advisor (Qwen3-8B) steers frozen frontier executors (Gemini 3.7 Flash, Claude Sonnet 4.6) with natural-language advice by pairing outcome-based GRPO with selective feedback-conditioned self-distillation: reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of that difference to choose which decisions to supervise, reaching 4.2–6.4 percentage points above advisor-GRPO on BFCL-v3 and 3.9–5.1 score points above it on EnvScaler while beating matched-count random selection and a no-gate variant.
The work introduces AdviSD, in which a small trainable advisor (Qwen3-8B) steers frozen frontier executors (Gemini 3.7 Flash, Claude Sonnet 4.6) with natural-language advice by pairing outcome-based GRPO with selective feedback-conditioned self-distillation: reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of that difference to choose which decisions to supervise, reaching 4.2–6.4 percentage points above advisor-GRPO on BFCL-v3 and 3.9–5.1 score points above it on EnvScaler while beating matched-count random selection and a no-gate variant.
The work introduces AdviSD, in which a small trainable advisor (Qwen3-8B) steers frozen frontier executors (Gemini 3.7 Flash, Claude Sonnet 4.6) with natural-language advice by pairing outcome-based GRPO with selective feedback-conditioned self-distillation: reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of that difference to choose which decisions to supervise, reaching 4.2–6.4 percentage points above advisor-GRPO on BFCL-v3 and 3.9–5.1 score points above it on EnvScaler while beating matched-count random selection and a no-gate variant.
arXiv Analyzing JEV 1.13 and three open KEV models (0.8B/4B/9B), the study finds that on ANLI JEV reaches 74.95% accuracy yet assigns 38.8% of predictions and 51.3% of errors to Neutral, and across 36 ordinal datasets final decisions cover only 67–76% of the effective gold support (versus 87–102% on four nominal tasks); through candidate-order randomization, same-item scale refinement from K=2 to K=14 (utilization falling to 26–75% at K=14), and targeted BA-LoRA post-training (gold-relative utilization rising from roughly 47% to 86% on eight supervised scales), it shows this ordinal scale-utilization bias is a learned and modifiable decision-stage compression of the candidate space.
Analyzing JEV 1.13 and three open KEV models (0.8B/4B/9B), the study finds that on ANLI JEV reaches 74.95% accuracy yet assigns 38.8% of predictions and 51.3% of errors to Neutral, and across 36 ordinal datasets final decisions cover only 67–76% of the effective gold support (versus 87–102% on four nominal tasks); through candidate-order randomization, same-item scale refinement from K=2 to K=14 (utilization falling to 26–75% at K=14), and targeted BA-LoRA post-training (gold-relative utilization rising from roughly 47% to 86% on eight supervised scales), it shows this ordinal scale-utilization bias is a learned and modifiable decision-stage compression of the candidate space.
Analyzing JEV 1.13 and three open KEV models (0.8B/4B/9B), the study finds that on ANLI JEV reaches 74.95% accuracy yet assigns 38.8% of predictions and 51.3% of errors to Neutral, and across 36 ordinal datasets final decisions cover only 67–76% of the effective gold support (versus 87–102% on four nominal tasks); through candidate-order randomization, same-item scale refinement from K=2 to K=14 (utilization falling to 26–75% at K=14), and targeted BA-LoRA post-training (gold-relative utilization rising from roughly 47% to 86% on eight supervised scales), it shows this ordinal scale-utilization bias is a learned and modifiable decision-stage compression of the candidate space.
Analyzing JEV 1.13 and three open KEV models (0.8B/4B/9B), the study finds that on ANLI JEV reaches 74.95% accuracy yet assigns 38.8% of predictions and 51.3% of errors to Neutral, and across 36 ordinal datasets final decisions cover only 67–76% of the effective gold support (versus 87–102% on four nominal tasks); through candidate-order randomization, same-item scale refinement from K=2 to K=14 (utilization falling to 26–75% at K=14), and targeted BA-LoRA post-training (gold-relative utilization rising from roughly 47% to 86% on eight supervised scales), it shows this ordinal scale-utilization bias is a learned and modifiable decision-stage compression of the candidate space.
arXiv SpatialCORE is a GRPO-based post-training framework that weights the geometric matching quality of each predicted bounding box by the model's own coordinate-token uncertainty in generating that grounding, and uses an answer gate to tie grounding optimization to final-answer correctness, so the model learns to reason from task-relevant objects that are both accurately and confidently localized; on the held-out OmniSpatial test set SpatialCORE-8B reaches the highest weighted-average accuracy among open-source and specialized spatial-reasoning models and leads zero-shot on the unseen SpatiaLab, while ablations show that removing confidence weighting drops accuracy from 56.33% to 52.33%.
SpatialCORE is a GRPO-based post-training framework that weights the geometric matching quality of each predicted bounding box by the model's own coordinate-token uncertainty in generating that grounding, and uses an answer gate to tie grounding optimization to final-answer correctness, so the model learns to reason from task-relevant objects that are both accurately and confidently localized; on the held-out OmniSpatial test set SpatialCORE-8B reaches the highest weighted-average accuracy among open-source and specialized spatial-reasoning models and leads zero-shot on the unseen SpatiaLab, while ablations show that removing confidence weighting drops accuracy from 56.33% to 52.33%.
SpatialCORE is a GRPO-based post-training framework that weights the geometric matching quality of each predicted bounding box by the model's own coordinate-token uncertainty in generating that grounding, and uses an answer gate to tie grounding optimization to final-answer correctness, so the model learns to reason from task-relevant objects that are both accurately and confidently localized; on the held-out OmniSpatial test set SpatialCORE-8B reaches the highest weighted-average accuracy among open-source and specialized spatial-reasoning models and leads zero-shot on the unseen SpatiaLab, while ablations show that removing confidence weighting drops accuracy from 56.33% to 52.33%.
SpatialCORE is a GRPO-based post-training framework that weights the geometric matching quality of each predicted bounding box by the model's own coordinate-token uncertainty in generating that grounding, and uses an answer gate to tie grounding optimization to final-answer correctness, so the model learns to reason from task-relevant objects that are both accurately and confidently localized; on the held-out OmniSpatial test set SpatialCORE-8B reaches the highest weighted-average accuracy among open-source and specialized spatial-reasoning models and leads zero-shot on the unseen SpatiaLab, while ablations show that removing confidence weighting drops accuracy from 56.33% to 52.33%.
arXiv The work proposes BiasReducer, a lightweight framework that edits only the linear reward head without retraining the reward model: an SAE-style encoder with semantic supervision learns internal representations for predefined attributes such as length, confidence, and sycophancy; the framework then determines the direction and amount by which to adjust the reward head for each attribute and stores these as an edit bank; for a new dataset it ranks attributes by their influence on reward scores and applies the matching edits. Across five reward models, BiasReducer-M improves RM-Bench-Hard, Arena-StyleConflict, and JudgeBiasBench by 8.3, 18.0, and 6.
The work proposes BiasReducer, a lightweight framework that edits only the linear reward head without retraining the reward model: an SAE-style encoder with semantic supervision learns internal representations for predefined attributes such as length, confidence, and sycophancy; the framework then determines the direction and amount by which to adjust the reward head for each attribute and stores these as an edit bank; for a new dataset it ranks attributes by their influence on reward scores and applies the matching edits. Across five reward models, BiasReducer-M improves RM-Bench-Hard, Arena-StyleConflict, and JudgeBiasBench by 8.3, 18.0, and 6.
The work proposes BiasReducer, a lightweight framework that edits only the linear reward head without retraining the reward model: an SAE-style encoder with semantic supervision learns internal representations for predefined attributes such as length, confidence, and sycophancy; the framework then determines the direction and amount by which to adjust the reward head for each attribute and stores these as an edit bank; for a new dataset it ranks attributes by their influence on reward scores and applies the matching edits. Across five reward models, BiasReducer-M improves RM-Bench-Hard, Arena-StyleConflict, and JudgeBiasBench by 8.3, 18.0, and 6.
The work proposes BiasReducer, a lightweight framework that edits only the linear reward head without retraining the reward model: an SAE-style encoder with semantic supervision learns internal representations for predefined attributes such as length, confidence, and sycophancy; the framework then determines the direction and amount by which to adjust the reward head for each attribute and stores these as an edit bank; for a new dataset it ranks attributes by their influence on reward scores and applies the matching edits. Across five reward models, BiasReducer-M improves RM-Bench-Hard, Arena-StyleConflict, and JudgeBiasBench by 8.3, 18.0, and 6.
arXiv The work introduces meta-skills — principles specifying when support is needed, what resources to provide, and how the Target should use them — which a Builder learns from the Target's execution feedback on a development set while both models' weights stay fixed, then freezes into a skill bank to construct harnesses for unseen tasks; across Harness-Bench and NewtonBench, full-bank meta-skills raise macro-average performance by 8.95 percentage points over no-skill construction and 12.02 points over delivering the same bank directly to the Target.
The work introduces meta-skills — principles specifying when support is needed, what resources to provide, and how the Target should use them — which a Builder learns from the Target's execution feedback on a development set while both models' weights stay fixed, then freezes into a skill bank to construct harnesses for unseen tasks; across Harness-Bench and NewtonBench, full-bank meta-skills raise macro-average performance by 8.95 percentage points over no-skill construction and 12.02 points over delivering the same bank directly to the Target.
The work introduces meta-skills — principles specifying when support is needed, what resources to provide, and how the Target should use them — which a Builder learns from the Target's execution feedback on a development set while both models' weights stay fixed, then freezes into a skill bank to construct harnesses for unseen tasks; across Harness-Bench and NewtonBench, full-bank meta-skills raise macro-average performance by 8.95 percentage points over no-skill construction and 12.02 points over delivering the same bank directly to the Target.
The work introduces meta-skills — principles specifying when support is needed, what resources to provide, and how the Target should use them — which a Builder learns from the Target's execution feedback on a development set while both models' weights stay fixed, then freezes into a skill bank to construct harnesses for unseen tasks; across Harness-Bench and NewtonBench, full-bank meta-skills raise macro-average performance by 8.95 percentage points over no-skill construction and 12.02 points over delivering the same bank directly to the Target.
arXiv LEAP splits an hour-scale recording into fixed-duration blocks, runs a lightweight localization pass per block to score short candidate windows, then re-encodes only the top-ranked windows in a single bounded answer pass, so the answer input and peak context stay independent of recording duration; across several AVQA benchmarks it improves over the Qwen3-Omni-30B-A3B baseline by 4.5–16.8% and transfers to MiniCPM-o 4.5, surpassing its published results by 3.1–13.0%.
LEAP splits an hour-scale recording into fixed-duration blocks, runs a lightweight localization pass per block to score short candidate windows, then re-encodes only the top-ranked windows in a single bounded answer pass, so the answer input and peak context stay independent of recording duration; across several AVQA benchmarks it improves over the Qwen3-Omni-30B-A3B baseline by 4.5–16.8% and transfers to MiniCPM-o 4.5, surpassing its published results by 3.1–13.0%.
LEAP splits an hour-scale recording into fixed-duration blocks, runs a lightweight localization pass per block to score short candidate windows, then re-encodes only the top-ranked windows in a single bounded answer pass, so the answer input and peak context stay independent of recording duration; across several AVQA benchmarks it improves over the Qwen3-Omni-30B-A3B baseline by 4.5–16.8% and transfers to MiniCPM-o 4.5, surpassing its published results by 3.1–13.0%.
LEAP splits an hour-scale recording into fixed-duration blocks, runs a lightweight localization pass per block to score short candidate windows, then re-encodes only the top-ranked windows in a single bounded answer pass, so the answer input and peak context stay independent of recording duration; across several AVQA benchmarks it improves over the Qwen3-Omni-30B-A3B baseline by 4.5–16.8% and transfers to MiniCPM-o 4.5, surpassing its published results by 3.1–13.0%.
arXiv The work presents RoboCoach, a world-model-guided active coaching framework whose Route-Imagine-Diagnose-Improve (RIDI) loop executes reusable skill experts in closed loop inside CoachWorld, a shared action-conditioned world model, uses a progress judge to record the first subtask that fails to complete, and thereby selects which subtask demonstrations to acquire and which expert adapters to update; across two simulation suites and two real-robot platforms, imagined and deployed success correlate with Spearman rho = 0.840 over 22 task-policy pairs, and with only 150 additional subtask demonstrations per platform success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX, while the coached experts reach 35.
The work presents RoboCoach, a world-model-guided active coaching framework whose Route-Imagine-Diagnose-Improve (RIDI) loop executes reusable skill experts in closed loop inside CoachWorld, a shared action-conditioned world model, uses a progress judge to record the first subtask that fails to complete, and thereby selects which subtask demonstrations to acquire and which expert adapters to update; across two simulation suites and two real-robot platforms, imagined and deployed success correlate with Spearman rho = 0.840 over 22 task-policy pairs, and with only 150 additional subtask demonstrations per platform success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX, while the coached experts reach 35.
The work presents RoboCoach, a world-model-guided active coaching framework whose Route-Imagine-Diagnose-Improve (RIDI) loop executes reusable skill experts in closed loop inside CoachWorld, a shared action-conditioned world model, uses a progress judge to record the first subtask that fails to complete, and thereby selects which subtask demonstrations to acquire and which expert adapters to update; across two simulation suites and two real-robot platforms, imagined and deployed success correlate with Spearman rho = 0.840 over 22 task-policy pairs, and with only 150 additional subtask demonstrations per platform success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX, while the coached experts reach 35.
The work presents RoboCoach, a world-model-guided active coaching framework whose Route-Imagine-Diagnose-Improve (RIDI) loop executes reusable skill experts in closed loop inside CoachWorld, a shared action-conditioned world model, uses a progress judge to record the first subtask that fails to complete, and thereby selects which subtask demonstrations to acquire and which expert adapters to update; across two simulation suites and two real-robot platforms, imagined and deployed success correlate with Spearman rho = 0.840 over 22 task-policy pairs, and with only 150 additional subtask demonstrations per platform success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX, while the coached experts reach 35.
arXiv Holding all other settings fixed and varying only the hidden current date injected into the system prompt across every day of 2024, this study of 9 recent LLMs and 6 datasets finds accuracy swings of up to 6% on multiple-choice QA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation, with model rankings reordering, an effect larger than batch size or numerical precision and one that chain-of-thought prompting amplifies rather than reduces.
Holding all other settings fixed and varying only the hidden current date injected into the system prompt across every day of 2024, this study of 9 recent LLMs and 6 datasets finds accuracy swings of up to 6% on multiple-choice QA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation, with model rankings reordering, an effect larger than batch size or numerical precision and one that chain-of-thought prompting amplifies rather than reduces.
Holding all other settings fixed and varying only the hidden current date injected into the system prompt across every day of 2024, this study of 9 recent LLMs and 6 datasets finds accuracy swings of up to 6% on multiple-choice QA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation, with model rankings reordering, an effect larger than batch size or numerical precision and one that chain-of-thought prompting amplifies rather than reduces.
Holding all other settings fixed and varying only the hidden current date injected into the system prompt across every day of 2024, this study of 9 recent LLMs and 6 datasets finds accuracy swings of up to 6% on multiple-choice QA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation, with model rankings reordering, an effect larger than batch size or numerical precision and one that chain-of-thought prompting amplifies rather than reduces.
arXiv Across two no-stakes cooperative games (Leader Election and Exclusion) and the GPQA-Diamond reasoning benchmark, with nine to twenty-five agents drawn from up to five open-weight model families, the study finds that when agents know each other's model family the interaction graph partitions into factions along that label, that the split follows the visible label rather than the underlying architecture (it persists under shuffled labels and neutral color tags and disappears when labels are removed), and that in strictly cooperative tasks labeled groups spend on average 30% more rounds and 55% more tokens to reach a decision while success drops from 96% to 81%.
Across two no-stakes cooperative games (Leader Election and Exclusion) and the GPQA-Diamond reasoning benchmark, with nine to twenty-five agents drawn from up to five open-weight model families, the study finds that when agents know each other's model family the interaction graph partitions into factions along that label, that the split follows the visible label rather than the underlying architecture (it persists under shuffled labels and neutral color tags and disappears when labels are removed), and that in strictly cooperative tasks labeled groups spend on average 30% more rounds and 55% more tokens to reach a decision while success drops from 96% to 81%.
Across two no-stakes cooperative games (Leader Election and Exclusion) and the GPQA-Diamond reasoning benchmark, with nine to twenty-five agents drawn from up to five open-weight model families, the study finds that when agents know each other's model family the interaction graph partitions into factions along that label, that the split follows the visible label rather than the underlying architecture (it persists under shuffled labels and neutral color tags and disappears when labels are removed), and that in strictly cooperative tasks labeled groups spend on average 30% more rounds and 55% more tokens to reach a decision while success drops from 96% to 81%.
Across two no-stakes cooperative games (Leader Election and Exclusion) and the GPQA-Diamond reasoning benchmark, with nine to twenty-five agents drawn from up to five open-weight model families, the study finds that when agents know each other's model family the interaction graph partitions into factions along that label, that the split follows the visible label rather than the underlying architecture (it persists under shuffled labels and neutral color tags and disappears when labels are removed), and that in strictly cooperative tasks labeled groups spend on average 30% more rounds and 55% more tokens to reach a decision while success drops from 96% to 81%.
arXiv The work introduces Loop Scaling Laws, the first scaling law that jointly models recurrence and MoE sparsity alongside model size and training data, using a bounded, sparsity-conditional recurrence mapping to characterize the effective-parameter gain from looping and its saturation; it predicts held-out loss more accurately than linear and power-law mappings, guides the choice of recurrence and expert count under compute and memory budgets, and at trillion-token scale lets a 0.3B-active/1.3B-total looped MoE match a roughly 2x larger non-looped MoE on reasoning benchmarks.
The work introduces Loop Scaling Laws, the first scaling law that jointly models recurrence and MoE sparsity alongside model size and training data, using a bounded, sparsity-conditional recurrence mapping to characterize the effective-parameter gain from looping and its saturation; it predicts held-out loss more accurately than linear and power-law mappings, guides the choice of recurrence and expert count under compute and memory budgets, and at trillion-token scale lets a 0.3B-active/1.3B-total looped MoE match a roughly 2x larger non-looped MoE on reasoning benchmarks.
The work introduces Loop Scaling Laws, the first scaling law that jointly models recurrence and MoE sparsity alongside model size and training data, using a bounded, sparsity-conditional recurrence mapping to characterize the effective-parameter gain from looping and its saturation; it predicts held-out loss more accurately than linear and power-law mappings, guides the choice of recurrence and expert count under compute and memory budgets, and at trillion-token scale lets a 0.3B-active/1.3B-total looped MoE match a roughly 2x larger non-looped MoE on reasoning benchmarks.
The work introduces Loop Scaling Laws, the first scaling law that jointly models recurrence and MoE sparsity alongside model size and training data, using a bounded, sparsity-conditional recurrence mapping to characterize the effective-parameter gain from looping and its saturation; it predicts held-out loss more accurately than linear and power-law mappings, guides the choice of recurrence and expert count under compute and memory budgets, and at trillion-token scale lets a 0.3B-active/1.3B-total looped MoE match a roughly 2x larger non-looped MoE on reasoning benchmarks.
arXiv SkillGym proposes an automatic pipeline that crawls skills from the internet and keeps those whose workflows run reproducibly offline, uses a builder-reviewer agent system to construct skill-critical tasks across four reasoning structures each with a reference solution and an executable verifier, then collects successful trajectories from three teacher models and four agent harnesses for supervised finetuning, improving six models from 2B to 122B parameters on 22 of 24 comparisons across four skill-use benchmarks, with the 9B model outperforming Qwen3.5-397B-A17B on the SkillGym test set and SkillEval and raising the rate of reading the relevant skill from 28% to 96%.
SkillGym proposes an automatic pipeline that crawls skills from the internet and keeps those whose workflows run reproducibly offline, uses a builder-reviewer agent system to construct skill-critical tasks across four reasoning structures each with a reference solution and an executable verifier, then collects successful trajectories from three teacher models and four agent harnesses for supervised finetuning, improving six models from 2B to 122B parameters on 22 of 24 comparisons across four skill-use benchmarks, with the 9B model outperforming Qwen3.5-397B-A17B on the SkillGym test set and SkillEval and raising the rate of reading the relevant skill from 28% to 96%.
SkillGym proposes an automatic pipeline that crawls skills from the internet and keeps those whose workflows run reproducibly offline, uses a builder-reviewer agent system to construct skill-critical tasks across four reasoning structures each with a reference solution and an executable verifier, then collects successful trajectories from three teacher models and four agent harnesses for supervised finetuning, improving six models from 2B to 122B parameters on 22 of 24 comparisons across four skill-use benchmarks, with the 9B model outperforming Qwen3.5-397B-A17B on the SkillGym test set and SkillEval and raising the rate of reading the relevant skill from 28% to 96%.
SkillGym proposes an automatic pipeline that crawls skills from the internet and keeps those whose workflows run reproducibly offline, uses a builder-reviewer agent system to construct skill-critical tasks across four reasoning structures each with a reference solution and an executable verifier, then collects successful trajectories from three teacher models and four agent harnesses for supervised finetuning, improving six models from 2B to 122B parameters on 22 of 24 comparisons across four skill-use benchmarks, with the 9B model outperforming Qwen3.5-397B-A17B on the SkillGym test set and SkillEval and raising the rate of reading the relevant skill from 28% to 96%.