Skip to main content

AI Core

906 items

  1. arXiv

    EVOKE ranks the same actions under alternative goals at fixed states, pushing pretrained world knowledge into decisions: ALFWorld unseen success rises from 60.4% to 91.8%

    EVOKE is a post-training method that holds the environment state and interaction history fixed while introducing alternative goals for the same candidate actions and training on contrastive rankings, forcing the policy to elicit action-consequence knowledge already present from pretraining; across ALFWorld, WebShop, and search-based QA on three backbones it achieves the best average, raising ALFWorld unseen-game success from 60.4% for πboot to 91.8% while using fewer actions.
  2. arXiv

    SkillSeek matches an LLM-mediated retrieval loop on 89 SkillsBench tasks with BM25 plus a small reranker, cutting per-trial spend from USD 51.30 to USD 27.54

    The work presents SkillSeek, an open-source two-stage skill retriever (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP), and across a grid of two skill pools and two backbones on the 89-task SkillsBench benchmark observes that plain bm25 alone records a pass rate at or above Liu et al.'s LLM-mediated loop on three of four settings, with a small cross-encoder covering the remaining difference on the fourth (34K pool with Qwen3.5) at the same 0.442, while total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the USD 27.41 no-skill baseline), a pattern the authors attribute to a first-stage recall ceiling.
  3. arXiv

    Endless Exam scores nine models on 14 families of math construction problems: GPT-6 Astra reaches 143.16 with tools, yet none of 30 published frontiers is surpassed

    The authors introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems with 69 evaluation instances, where each submitted object is automatically checked for validity and given an uncapped relative quality score (100 marks average reference parity); across nine models and seventeen configurations, tool-free overall scores span 7.55 to 91.90, and with code and web access GPT-6 Astra, GPT-5.6 Luna and Claude Opus 5.5 reach 143.16, 119.35 and 180.06, while none of the 30 published-frontier references is surpassed.
  4. arXiv

    UniEvo-VL lets a multimodal model teach itself with its own critique, lifting GenEval from 0.747 to 0.808

    UniEvo-VL introduces an on-policy self-distillation (OPSD) recipe in which one multimodal model acts as both a student seeing only the vanilla prompt and a teacher conditioned on a critique-derived revised prompt, minimizing per-state divergence between their denoising distributions along the student's own trajectories; on Qwen-Image-2512 it raises GenEval from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53, while stronger external critics such as GPT5.6-Luna indicate a higher self-evolving ceiling and text-rendering gains remain uneven.
  5. arXiv

    ThinkV2V makes MLLMs think before editing: a 5B model tops both complex and standard video-editing benchmarks

    The work proposes ThinkV2V, a reasoning-driven video-editing framework that explicitly activates multimodal large language model thinking before visual generation: it uses Qwen3-VL-Thinking-8B to perform chain-of-thought reasoning over the source video and instruction and output a refined prompt, injects the hidden states into a Wan-2.1-5B DiT through a learnable-query connector, pairs this with Progressive Curriculum Training and Inference-Time Thinking Scaling, and builds the ThinkV2V-150K dataset and ThinkV2V-Bench; under two judge models, the 5B-scale model achieves the best overall scores on both complex and standard editing scenarios and surpasses several 10B-scale baselines.
  6. arXiv

    A hidden current date in system prompts shifts nine models' scores by up to 14% and reshuffles leaderboards

    Holding all other settings fixed and varying only the hidden current date injected into the system prompt across every day of 2024, this study of 9 recent LLMs and 6 datasets finds accuracy swings of up to 6% on multiple-choice QA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation, with model rankings reordering, an effect larger than batch size or numerical precision and one that chain-of-thought prompting amplifies rather than reduces.
  7. arXiv

    OASIS restricts self-distillation supervision to verified trajectories, widening its average gain over OPSD from 0.59 to 3.05 points across Qwen3-1.7B, 4B and 8B

    Using truncated-continuation probes, the work finds that in on-policy self-distillation (OPSD) the teacher's advantage over the student is concentrated on failed trajectories (77% versus 37% recovery at 1.7B) and shrinks with scale, and therefore proposes OASIS: keeping the OPSD objective unchanged, it applies supervision only along the shortest outcome-verified on-policy trajectory and replaces the written reference solution with a distinct same-problem self-generated rollout, typically the student's own unverified attempt, improving mean Avg@12 over OPSD by 0.59, 1.86 and 3.05 points on Qwen3-1.7B, 4B and 8B across AIME 2024, AIME 2025 and HMMT 2025 while requiring no written solutions.
  8. arXiv

    When one observation admits several valid actions, CVAE KL regularization and flow-model Lipschitz smoothness decide whether a policy keeps multimodality, while standard robot simulation benchmarks turn out to be nearly unimodal

    The work formalizes multimodality in behavioral cloning, proves that latent-variable policies preserve demonstrated modes only if the latent carries action-conditioned information and that excessive posterior-prior regularization suppresses it, shows that action-space generative policies are limited by the Lipschitz constant of the base-to-action map, validates these mechanisms on synthetic multimodal navigation and a real-robot bimodal tissue-grasping task, and uses a GMM modal-clustering diagnostic to find limited conditional multimodality in Push-T, UR3, LIBERO, and Meta-World, where deterministic regression remains competitive.
  9. arXiv

    WUSH-KV cuts 2-bit KV-cache perplexity from OSCAR's 13.74 to 10.51 and leads OSCAR on all four 8B downstream tasks

    The work adapts the WUSH transform, originally for weight-activation quantization, to the KV cache of grouped-query attention: calibration data builds separate key and value transforms per KV head (the value transform folded into the weights, the key transform applied after RoPE), the transform is shown to be near-optimal among sensitivity-balanced transforms for the QuEST quantizer, and end-to-end evaluation in SGLang with OSCAR-style percentile-clipped affine quantization gives 2-bit WikiText-2 perplexity of 10.51 on Qwen3-8B versus OSCAR's 13.74, with the 8B model scoring higher than OSCAR on AIME 2025, MATH-500, GPQA Diamond, and LiveCodeBench v6.
  10. arXiv

    NavHarness carries maps, search records and house notes into the next conversation, lifting GOAT-Bench s-SR by 18.6 to 34.8 points

    NavHarness is a training-free embodied harness that makes memory processing part of the navigation loop: each task opens a fresh multi-round agentic session that inherits prior experience through maps, task records and house notes, while outcome verification and run-end consolidation decide what later sessions inherit; on GOAT-Bench it improves s-SR over context-only independent sessions by 18.6 points with GPT-6 Astra, 22.6 with Opus 5, 34.8 with GPT-4o and 30.3 with Qwen3.8-27B, and with SLAM-estimated poses Astra reaches 83.7 s-SR with 36.9 e-SR on GOAT-Bench and 85.9 s-SR on IR2R-CE.

Page 7 · showing 10