Skip to main content

Search

“All disciplines” · 1364 results

Page 5 · showing 20
arXiv

SUSTech and CityU propose Box2-Bench: frontier models use good workflows but lose up to 43.3 points under bad ones

The authors introduce Box2-Bench, which holds the model and task fixed while varying workflow availability and reliability across five matched conditions (none, good, bad, partial, mixed), finding that frontier models such as Gemini 3.7 Flash, DeepSeek V4 Flash, and GLM 5.2 benefit from good workflows in 8 of 12 model-task pairs but are hurt by bad workflows in all 12 (drops of 3.3 to 43.3 points), and that mixed workflows underperform their matched partial workflows in 10 of 12 pairs; training two open-weight models on bad workflows only, counterfactual SFT cuts the bad-workflow penalty from 20.0 to 6.7 points on AIME and from 9.8 to 0.6 points on WebShop while weakening use of held-out good workflows, and outcome-based RL restores AIME utilization from -3.3 to +6.
arXiv

AED turns 50,228 failed agent rollouts into error–diagnosis pairs: first-proposal corrections lift verifier pass rates from 18.4% to 51.1%, and full-diagnosis fine-tuning raises Qwen3-8B's exact-step agreement from 47.2% to 63.6%

The work introduces the Agent Error Dataset (AED), comprising 50,228 error–diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families and 23 policy models, and a five-stage Agentic Error-to-Training (AET) pipeline that links natural failures to diagnoses, proposed corrections and optional replay evidence; across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1% (a 32.7 percentage-point gain), full-diagnosis fine-tuning on a separately frozen diagnosis release raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6% (averaged over three seeds on a 943-case holdout), and in a single-seed actor-training comparison action-only repair training scores 6.
arXiv

EvolvingNav predicts where targets go with a time-indexed 4D belief, lifting first-inspection success from 45.33% to 61.32% on EvoWorld-Bench

The work introduces EvolvingNav, which builds a persistence–relocation belief from timestamped 3D object histories and closes the loop with an event-driven predict–observe–replan filter so an agent can infer where a target is when targets may move before the query or during navigation, and it releases EvoWorld-Bench with 54 scenes and 803,680 tasks, reporting improved navigation success and search efficiency over the evaluated baselines in simulation and real-robot experiments.
arXiv

LANTERN used Qwen3-32B hidden states to rank 50 million OEIS sequence pairs and returned 62 verified relations in under 8 hours, four of them absent from OEIS and a targeted literature search

LANTERN trains a linear classifier on Qwen3-32B activations to rank 50 million candidate relations among the 10,000 most-referenced OEIS sequences, then applies staged filtering, hypothesis generation, executable verification and analytical checking to produce 62 verified relations without an existing OEIS cross-reference; a content screen retained 13, nine of them informative or insightful, four not found in the OEIS or a targeted literature search, with the end-to-end process taking under 8 hours.
arXiv

DC-SAE splits image tokenization into semantic and pixel branches, hitting 32x compression with 3.37 gFID and 29.79 PSNR on ImageNet and a 4.41x speedup over DC-AE

Peking University and Singapore University of Technology and Design propose DC-SAE, a decoupled compact semantic autoencoder that pairs a frozen semantic encoder for a generation-friendly compact latent with a trainable pixel encoder for low-level detail, achieving 29.79 PSNR and 3.37 gFID on ImageNet at 32x spatial compression, improving over DC-AE by 13.5% in PSNR and 54.9% in gFID with a 4.41x speedup, while a B-parameter DiT reaches 0.84 on GenEval and 86.007 on DPG-Bench.
arXiv

RIDE extrapolates RL teachers' hidden-state residuals and matches or beats the teacher on average across four base/teacher pairs

The work proposes RIDE (RL-Induced Direction Extrapolation): on student-generated trajectories it computes the layerwise hidden-state residual between an RL-trained teacher and its pre-RL checkpoint, then regresses the student's hidden states toward targets displaced beyond the teacher along that residual; across four base/RL-teacher pairs (R1-Distill-1.5B, Qwen3-4B, Llama-3.2-3B, Phi-4-mini) RIDE's mean Avg@16 on AIME24, AIME25 and AIMO approaches or exceeds the RL-trained teacher on every pair, the only compared method whose mean does so, and it consistently outperforms output-space extrapolation.
arXiv

CUA-SWE tested 105 executable tasks and found that giving agents the running interface lifted GPT-6-Astra's four-domain mean from 11.3% to 59.9%

The authors introduce CUA-SWE, a benchmark, interactive development environment, and evaluation pipeline spanning 105 tasks across Web, Game, DevOps, and Mobile, in which an agent edits code, runs commands, operates the running software, and inspects screenshots within a single task, with task-specific deterministic tests deciding the outcome; comparing code-only and hybrid CUA conditions across frontier models, GPT-6-Astra reaches the highest four-domain mean of 59.9%, and every frontier model in the four-domain comparison achieves higher aggregate success with hybrid access, with gains concentrated on tasks whose requirements must be recovered from application materials such as drawings, service contracts, or graphical reference cards.
arXiv

Across six pretrained language models, residual streams favor their own endpoint early while competitor sets shrink with changing membership

By comparing each context's intermediate residual states with its own final state and an empirical bank of final states from other contexts, the study finds across Gemma-2B, Gemma-7B, Qwen2.5-1.5B, Qwen2.5-7B, Mistral-7B, and Llama-3-8B that the own endpoint becomes preferable to the average alternative at the earliest measured layer while many individual endpoints remain closer, that competitor sets generally shrink with depth yet show both entries and exits, and that directional alignment can improve while Euclidean distance to the final state changes little; the authors add a high-dimensional model separating norm, alignment, and endpoint geometry and prove that a straight path toward the own endpoint cannot introduce new competitors.
arXiv

CoEvoWhen lets a frozen VLM coevolve policies and tools for ultra-long video, raising ExtremeWhenBench mIoU by 74.9% while cutting visual tokens by 11.4%

CoEvoWhen introduces a policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from a VLM's agentic reasoning trajectories into a reusable skill without updating model parameters, improving ultra-long video temporal grounding accuracy and reducing inference visual token cost across five benchmarks and three VLMs, with the evolved skill transferring to general long-video QA without additional task-specific evolution.
arXiv

AdviSD trains a Qwen3-8B advisor to learn selectively from feedback corrections, beating advisor-GRPO by 4.2–6.4 points on BFCL-v3

The work introduces AdviSD, in which a small trainable advisor (Qwen3-8B) steers frozen frontier executors (Gemini 3.7 Flash, Claude Sonnet 4.6) with natural-language advice by pairing outcome-based GRPO with selective feedback-conditioned self-distillation: reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of that difference to choose which decisions to supervise, reaching 4.2–6.4 percentage points above advisor-GRPO on BFCL-v3 and 3.9–5.1 score points above it on EnvScaler while beating matched-count random selection and a no-gate variant.
arXiv

JEV-like direct-decision models use only 67–76% of the effective ordinal label space across 36 datasets, and BA-LoRA post-training lifts utilization from about 47% to 86%

Analyzing JEV 1.13 and three open KEV models (0.8B/4B/9B), the study finds that on ANLI JEV reaches 74.95% accuracy yet assigns 38.8% of predictions and 51.3% of errors to Neutral, and across 36 ordinal datasets final decisions cover only 67–76% of the effective gold support (versus 87–102% on four nominal tasks); through candidate-order randomization, same-item scale refinement from K=2 to K=14 (utilization falling to 26–75% at K=14), and targeted BA-LoRA post-training (gold-relative utilization rising from roughly 47% to 86% on eight supervised scales), it shows this ordinal scale-utilization bias is a learned and modifiable decision-stage compression of the candidate space.
arXiv

SpatialCORE folds both box accuracy and box confidence into the reward, letting an 8B model top same-scale open-source and specialized spatial-reasoning models on OmniSpatial and SpatiaLab

SpatialCORE is a GRPO-based post-training framework that weights the geometric matching quality of each predicted bounding box by the model's own coordinate-token uncertainty in generating that grounding, and uses an answer gate to tie grounding optimization to final-answer correctness, so the model learns to reason from task-relevant objects that are both accurately and confidently localized; on the held-out OmniSpatial test set SpatialCORE-8B reaches the highest weighted-average accuracy among open-source and specialized spatial-reasoning models and leads zero-shot on the unseen SpatiaLab, while ablations show that removing confidence weighting drops accuracy from 56.33% to 52.33%.
arXiv

BiasReducer edits only the reward head, lifting five reward models by 8.3/18.0/6.9 points on average across three bias benchmarks

The work proposes BiasReducer, a lightweight framework that edits only the linear reward head without retraining the reward model: an SAE-style encoder with semantic supervision learns internal representations for predefined attributes such as length, confidence, and sycophancy; the framework then determines the direction and amount by which to adjust the reward head for each attribute and stores these as an edit bank; for a new dataset it ranks attributes by their influence on reward scores and applies the matching edits. Across five reward models, BiasReducer-M improves RM-Bench-Hard, Arena-StyleConflict, and JudgeBiasBench by 8.3, 18.0, and 6.
arXiv

Teaching a Builder to design execution environments: meta-skills lift macro-average scores by 8.95 points with frozen model weights

The work introduces meta-skills — principles specifying when support is needed, what resources to provide, and how the Target should use them — which a Builder learns from the Target's execution feedback on a development set while both models' weights stay fixed, then freezes into a skill bank to construct harnesses for unseen tasks; across Harness-Bench and NewtonBench, full-bank meta-skills raise macro-average performance by 8.95 percentage points over no-skill construction and 12.02 points over delivering the same bank directly to the Target.
arXiv

LEAP keeps hour-scale audio-video QA out of one giant context: 4.5–16.8% over the Qwen3-Omni-30B baseline and transfer to MiniCPM-o 4.5

LEAP splits an hour-scale recording into fixed-duration blocks, runs a lightweight localization pass per block to score short candidate windows, then re-encodes only the top-ranked windows in a single bounded answer pass, so the answer input and peak context stay independent of recording duration; across several AVQA benchmarks it improves over the Qwen3-Omni-30B-A3B baseline by 4.5–16.8% and transfers to MiniCPM-o 4.5, surpassing its published results by 3.1–13.0%.
arXiv

RoboCoach turns imagined failures into targeted demos: 150 extra subtask demonstrations lift Franka success from 13.3% to 75.0%

The work presents RoboCoach, a world-model-guided active coaching framework whose Route-Imagine-Diagnose-Improve (RIDI) loop executes reusable skill experts in closed loop inside CoachWorld, a shared action-conditioned world model, uses a progress judge to record the first subtask that fails to complete, and thereby selects which subtask demonstrations to acquire and which expert adapters to update; across two simulation suites and two real-robot platforms, imagined and deployed success correlate with Spearman rho = 0.840 over 22 task-policy pairs, and with only 150 additional subtask demonstrations per platform success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX, while the coached experts reach 35.
arXiv

A hidden current date in system prompts shifts nine models' scores by up to 14% and reshuffles leaderboards

Holding all other settings fixed and varying only the hidden current date injected into the system prompt across every day of 2024, this study of 9 recent LLMs and 6 datasets finds accuracy swings of up to 6% on multiple-choice QA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation, with model rankings reordering, an effect larger than batch size or numerical precision and one that chain-of-thought prompting amplifies rather than reduces.
arXiv

Labeling agents with their model family split multi-agent groups along the label, raising rounds and tokens and cutting success from 96% to 81%

Across two no-stakes cooperative games (Leader Election and Exclusion) and the GPQA-Diamond reasoning benchmark, with nine to twenty-five agents drawn from up to five open-weight model families, the study finds that when agents know each other's model family the interaction graph partitions into factions along that label, that the split follows the visible label rather than the underlying architecture (it persists under shuffled labels and neutral color tags and disappears when labels are removed), and that in strictly cooperative tasks labeled groups spend on average 30% more rounds and 55% more tokens to reach a decision while success drops from 96% to 81%.
arXiv

First joint scaling law for looped mixture of experts: sparsity raises and delays the saturation of recurrence's effective-parameter gain

The work introduces Loop Scaling Laws, the first scaling law that jointly models recurrence and MoE sparsity alongside model size and training data, using a bounded, sparsity-conditional recurrence mapping to characterize the effective-parameter gain from looping and its saturation; it predicts held-out loss more accurately than linear and power-law mappings, guides the choice of recurrence and expert count under compute and memory budgets, and at trillion-token scale lets a 0.3B-active/1.3B-total looped MoE match a roughly 2x larger non-looped MoE on reasoning benchmarks.
arXiv

SkillGym turns 184k community skills into verifiable training environments, and its 9B fine-tuned model beats a 397B untrained model on two benchmarks

SkillGym proposes an automatic pipeline that crawls skills from the internet and keeps those whose workflows run reproducibly offline, uses a builder-reviewer agent system to construct skill-critical tasks across four reasoning structures each with a reference solution and an executable verifier, then collects successful trajectories from three teacher models and four agent harnesses for supervised finetuning, improving six models from 2B to 122B parameters on 22 of 24 comparisons across four skill-use benchmarks, with the 9B model outperforming Qwen3.5-397B-A17B on the SkillGym test set and SkillEval and raising the rate of reading the relevant skill from 28% to 96%.