Skip to main content

AI Core

927 items

  1. arXiv

    GGSD turns five buttons into playable skills via 1v1 self-play: humans clear Maze and CubePush on Ant, Franka and G1 with no extra training

    The work presents Game-Guided Skill Discovery (GGSD), in which a hierarchical agent self-plays 1v1 competitive games against a pool of past checkpoints, a high-level policy selects among 5 discrete skills (6 for G1) and a skill-conditioned low-level policy outputs motor actions, with a mutual-information reward separating skill semantics; after training a human can replace the high-level policy and use the same discrete skills, composing them on Ant, a Franka arm and a Unitree G1 to solve unseen Maze and CubePush tasks, with human success rates of at least 84%.
  2. arXiv

    AED turns 50,228 failed agent rollouts into error–diagnosis pairs: first-proposal corrections lift verifier pass rates from 18.4% to 51.1%, and full-diagnosis fine-tuning raises Qwen3-8B's exact-step agreement from 47.2% to 63.6%

    The work introduces the Agent Error Dataset (AED), comprising 50,228 error–diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families and 23 policy models, and a five-stage Agentic Error-to-Training (AET) pipeline that links natural failures to diagnoses, proposed corrections and optional replay evidence; across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1% (a 32.7 percentage-point gain), full-diagnosis fine-tuning on a separately frozen diagnosis release raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6% (averaged over three seeds on a 943-case holdout), and in a single-seed actor-training comparison action-only repair training scores 6.
  3. arXiv

    SeLMRoute turns query requirements into a probabilistic state via 16 interpretable semantic probes, reaching 72.08% average accuracy on LLMRouterBench

    SeLMRoute splits LLM routing into three stages: a decision model first answers 16 interpretable semantic questions about each query and keeps the probability distributions as a 40-dimensional semantic state, a lightweight supervised regressor then predicts each candidate model's performance, and deployment objectives are applied last; on LLMRouterBench (15 datasets, 20 candidate models, 11,481 queries) it reaches 72.08%±0.45 average accuracy and 72.64% under grouped five-fold out-of-fold evaluation, versus 69.23% for the strongest fixed candidate, and in a 13-model performance-cost setting it improves performance in all five grouped splits with a mean PerfGain of 2.66%.
  4. arXiv

    A2Z GameSpec-Bench tests coding agents on 100 long-form GDDs: near-99% verifiable rate, but top overall GDD Fidelity only 77.0

    The work introduces A2Z GameSpec-Bench, a benchmark of 100 long-form game design documents (50 Small and 50 Big, averaging 14,085 and 26,297 tokens) that turns each GDD into a fixed Dependency-Aware Contract and evaluates coding agents' end-to-end game development faithfulness through source-code inspection, scenario-based replay, and adaptive playtesting, finding that compilable and runnable outputs do not imply faithful implementation (e.g., Claude-Fable-5.1 reaches a 98.7% verifiable rate but an overall GDD Fidelity of 77.0), that dependency context raises coverage of recorded playtest violations from 71.1% to 80.2%, and that requirement-specific feedback improves overall GDD Fidelity by 10.9% relative to self-revision after two rounds on Big GDDs.
  5. arXiv

    LANTERN used Qwen3-32B hidden states to rank 50 million OEIS sequence pairs and returned 62 verified relations in under 8 hours, four of them absent from OEIS and a targeted literature search

    LANTERN trains a linear classifier on Qwen3-32B activations to rank 50 million candidate relations among the 10,000 most-referenced OEIS sequences, then applies staged filtering, hypothesis generation, executable verification and analytical checking to produce 62 verified relations without an existing OEIS cross-reference; a content screen retained 13, nine of them informative or insightful, four not found in the OEIS or a targeted literature search, with the end-to-end process taking under 8 hours.
  6. arXiv

    LSD cuts the length-scaling tax on already-solved queries from 19.0% to -3.7% while matching or beating RL Pass@1

    The work defines the length-scaling tax (LST) as the excess response length that RL post-training adds on already-solved queries without a matching accuracy gain, and proposes Length Self-Distillation (LSD), which routes solved prompt groups to on-policy distillation while keeping the original RLVR objective for unsolved groups, using an exponential moving average of the same policy lineage as teacher with no external model; LSD reduces LST from 19.0% to -3.7% on single-turn reasoning and from 31.4% to 13.7% on multi-turn agentic tasks while matching or improving average Pass@1 over RL.
  7. arXiv

    ReaLVR supervises latent visual reasoning with visual evidence, reaching a 63.7% five-task average on Qwen2.5-VL-7B and scaling latent visual reasoning to 235B

    The work first analyzes latent-token behavior in latent visual reasoning (LVR) and identifies a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer, which the authors attribute to the lack of explicit supervision during GRPO training; it then proposes ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories, using correct-versus-wrong answer contrast to decide where stronger supervision is needed and relevant-versus-mismatched visual evidence to specify what to preserve, consistently outperforming evaluated LVR baselines across three model families, reaching a 63.7% five-task average on Qwen2.5-VL-7B, and being the first to scale visual reasoning in latent space up to 235B parameters.
  8. arXiv

    RoboCoach turns imagined failures into targeted demos: 150 extra subtask demonstrations lift Franka success from 13.3% to 75.0%

    The work presents RoboCoach, a world-model-guided active coaching framework whose Route-Imagine-Diagnose-Improve (RIDI) loop executes reusable skill experts in closed loop inside CoachWorld, a shared action-conditioned world model, uses a progress judge to record the first subtask that fails to complete, and thereby selects which subtask demonstrations to acquire and which expert adapters to update; across two simulation suites and two real-robot platforms, imagined and deployed success correlate with Spearman rho = 0.840 over 22 task-policy pairs, and with only 150 additional subtask demonstrations per platform success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX, while the coached experts reach 35.
  9. arXiv

    SpatialCORE folds both box accuracy and box confidence into the reward, letting an 8B model top same-scale open-source and specialized spatial-reasoning models on OmniSpatial and SpatiaLab

    SpatialCORE is a GRPO-based post-training framework that weights the geometric matching quality of each predicted bounding box by the model's own coordinate-token uncertainty in generating that grounding, and uses an answer gate to tie grounding optimization to final-answer correctness, so the model learns to reason from task-relevant objects that are both accurately and confidently localized; on the held-out OmniSpatial test set SpatialCORE-8B reaches the highest weighted-average accuracy among open-source and specialized spatial-reasoning models and leads zero-shot on the unseen SpatiaLab, while ablations show that removing confidence weighting drops accuracy from 56.33% to 52.33%.
  10. Nature News

    Merriam-Webster adds agentic and vibe coding as Nature reports Google DeepMind's SynthIDBio watermarking AI-generated proteins

    Nature reports that Merriam-Webster added AI and computing terms such as agentic, AGI, natural language processing, prompt engineering and vibe coding, plus technical terms including biohack, direct air capture, geoengineering and uncanny valley, with lexicographer Peter Sokolowski noting the advanced knowledge these entries require and cognitive scientist Lera Boroditsky linking rapid technological change to language change; the same evidence bundle carries a Nature paper introducing SynthIDBio, which adapts SynthID-text's tournament sampling into ProteinMPNN to watermark protein sequences and fine-tunes AlphaFold 3's diffusion module to watermark structures, with in vitro validation showing no significant population-level difference in binding affinity for SARS-CoV-2 RBD, VEGF-A and PD-L

Page 11 · showing 10