AI Core
927 items
AED turns 50,228 failed agent rollouts into error–diagnosis pairs: first-proposal corrections lift verifier pass rates from 18.4% to 51.1%, and full-diagnosis fine-tuning raises Qwen3-8B's exact-step agreement from 47.2% to 63.6%
The work introduces the Agent Error Dataset (AED), comprising 50,228 error–diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families and 23 policy models, and a five-stage Agentic Error-to-Training (AET) pipeline that links natural failures to diagnoses, proposed corrections and optional replay evidence; across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1% (a 32.7 percentage-point gain), full-diagnosis fine-tuning on a separately frozen diagnosis release raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6% (averaged over three seeds on a 943-case holdout), and in a single-seed actor-training comparison action-only repair training scores 6.
SeLMRoute turns query requirements into a probabilistic state via 16 interpretable semantic probes, reaching 72.08% average accuracy on LLMRouterBench
SeLMRoute splits LLM routing into three stages: a decision model first answers 16 interpretable semantic questions about each query and keeps the probability distributions as a 40-dimensional semantic state, a lightweight supervised regressor then predicts each candidate model's performance, and deployment objectives are applied last; on LLMRouterBench (15 datasets, 20 candidate models, 11,481 queries) it reaches 72.08%±0.45 average accuracy and 72.64% under grouped five-fold out-of-fold evaluation, versus 69.23% for the strongest fixed candidate, and in a 13-model performance-cost setting it improves performance in all five grouped splits with a mean PerfGain of 2.66%.
A2Z GameSpec-Bench tests coding agents on 100 long-form GDDs: near-99% verifiable rate, but top overall GDD Fidelity only 77.0
The work introduces A2Z GameSpec-Bench, a benchmark of 100 long-form game design documents (50 Small and 50 Big, averaging 14,085 and 26,297 tokens) that turns each GDD into a fixed Dependency-Aware Contract and evaluates coding agents' end-to-end game development faithfulness through source-code inspection, scenario-based replay, and adaptive playtesting, finding that compilable and runnable outputs do not imply faithful implementation (e.g., Claude-Fable-5.1 reaches a 98.7% verifiable rate but an overall GDD Fidelity of 77.0), that dependency context raises coverage of recorded playtest violations from 71.1% to 80.2%, and that requirement-specific feedback improves overall GDD Fidelity by 10.9% relative to self-revision after two rounds on Big GDDs.
LANTERN used Qwen3-32B hidden states to rank 50 million OEIS sequence pairs and returned 62 verified relations in under 8 hours, four of them absent from OEIS and a targeted literature search
LANTERN trains a linear classifier on Qwen3-32B activations to rank 50 million candidate relations among the 10,000 most-referenced OEIS sequences, then applies staged filtering, hypothesis generation, executable verification and analytical checking to produce 62 verified relations without an existing OEIS cross-reference; a content screen retained 13, nine of them informative or insightful, four not found in the OEIS or a targeted literature search, with the end-to-end process taking under 8 hours.
LSD cuts the length-scaling tax on already-solved queries from 19.0% to -3.7% while matching or beating RL Pass@1
The work defines the length-scaling tax (LST) as the excess response length that RL post-training adds on already-solved queries without a matching accuracy gain, and proposes Length Self-Distillation (LSD), which routes solved prompt groups to on-policy distillation while keeping the original RLVR objective for unsolved groups, using an exponential moving average of the same policy lineage as teacher with no external model; LSD reduces LST from 19.0% to -3.7% on single-turn reasoning and from 31.4% to 13.7% on multi-turn agentic tasks while matching or improving average Pass@1 over RL.
ReaLVR supervises latent visual reasoning with visual evidence, reaching a 63.7% five-task average on Qwen2.5-VL-7B and scaling latent visual reasoning to 235B
The work first analyzes latent-token behavior in latent visual reasoning (LVR) and identifies a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer, which the authors attribute to the lack of explicit supervision during GRPO training; it then proposes ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories, using correct-versus-wrong answer contrast to decide where stronger supervision is needed and relevant-versus-mismatched visual evidence to specify what to preserve, consistently outperforming evaluated LVR baselines across three model families, reaching a 63.7% five-task average on Qwen2.5-VL-7B, and being the first to scale visual reasoning in latent space up to 235B parameters.
RoboCoach turns imagined failures into targeted demos: 150 extra subtask demonstrations lift Franka success from 13.3% to 75.0%
The work presents RoboCoach, a world-model-guided active coaching framework whose Route-Imagine-Diagnose-Improve (RIDI) loop executes reusable skill experts in closed loop inside CoachWorld, a shared action-conditioned world model, uses a progress judge to record the first subtask that fails to complete, and thereby selects which subtask demonstrations to acquire and which expert adapters to update; across two simulation suites and two real-robot platforms, imagined and deployed success correlate with Spearman rho = 0.840 over 22 task-policy pairs, and with only 150 additional subtask demonstrations per platform success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX, while the coached experts reach 35.
SpatialCORE folds both box accuracy and box confidence into the reward, letting an 8B model top same-scale open-source and specialized spatial-reasoning models on OmniSpatial and SpatiaLab
SpatialCORE is a GRPO-based post-training framework that weights the geometric matching quality of each predicted bounding box by the model's own coordinate-token uncertainty in generating that grounding, and uses an answer gate to tie grounding optimization to final-answer correctness, so the model learns to reason from task-relevant objects that are both accurately and confidently localized; on the held-out OmniSpatial test set SpatialCORE-8B reaches the highest weighted-average accuracy among open-source and specialized spatial-reasoning models and leads zero-shot on the unseen SpatiaLab, while ablations show that removing confidence weighting drops accuracy from 56.33% to 52.33%.
Merriam-Webster adds agentic and vibe coding as Nature reports Google DeepMind's SynthIDBio watermarking AI-generated proteins
Nature reports that Merriam-Webster added AI and computing terms such as agentic, AGI, natural language processing, prompt engineering and vibe coding, plus technical terms including biohack, direct air capture, geoengineering and uncanny valley, with lexicographer Peter Sokolowski noting the advanced knowledge these entries require and cognitive scientist Lera Boroditsky linking rapid technological change to language change; the same evidence bundle carries a Nature paper introducing SynthIDBio, which adapts SynthID-text's tournament sampling into ProteinMPNN to watermark protein sequences and fine-tunes AlphaFold 3's diffusion module to watermark structures, with in vitro validation showing no significant population-level difference in binding affinity for SARS-CoV-2 RBD, VEGF-A and PD-L
Page 11 · showing 10