Skip to main content

AI Core

906 items

  1. arXiv

    AREX-2 trains a 27B agent on long-horizon reflective trajectories, reaching 81.8 on MLE-bench Lite and 84.0 on BrowseComp

    AREX-2 defines agent self-improvement as turning more test-time rounds into a better solution, synthesizes long-horizon improvement trajectories from machine learning engineering and algorithmic programming tasks that retain failures and regressions, and after fine-tuning Qwen3.8-27B reaches 81.8 on MLE-bench Lite and 70.7 on Frontier-CS while lifting BrowseComp to 84.0, HLE to 52.6, GAIA to 92.2 and DeepSearchQA to 93.8 without adding new deep-research data, with continued gains as the round budget grows.
  2. arXiv

    Splitting masked-diffusion decoding into five axes, researchers find state adaptation pays off only in a few predictable states, and selective intervention lifts task-macro utility by 3.2, 4.0, and 5.6 points across three models

    The work factorizes masked diffusion language model (MDM) inference into five axes—score, cardinality, region, commitment, and planning—defines adaptation opportunity as the one-step utility advantage of the best candidate action over a validation-selected fixed action, finds across three models (LLaDA-8B-Instruct, LLaDA-1.5, Dream-7B) and ten tasks that adaptation opportunities are highly heterogeneous, and proposes selective adaptation in which lightweight detectors calibrated on validation prompts identify high-opportunity states, capturing for example 56.9% of the candidate-set oracle opportunity by adapting only the top 10% of states on LLaDA-8B constrained JSON filling, with end-to-end selective decoding raising task-macro utility from 0.392 to 0.424 (LLaDA-8B), 0.344 to 0.
  3. arXiv

    PatchHolmes lets an agent read all 100 candidate commits at once, lifting patch-retrieval Recall@1 from 32.63% to 59.95%

    PatchHolmes is a two-phase patch retrieval system whose first stage fuses BM25 with time decay and Qwen3-Embedding-8B dense retrieval via RRF into a top-100 candidate set, and whose second stage runs a frozen open-weight LLM agent that views all 100 candidates listwise through four tools (list_candidates, read_commit, read_file_diff, submit_answer) and submits a single best commit; on 809 CVEs from GitHubAD it reaches 59.95% Recall@1, 25.34% above the pointwise classifier Favia and 31.40% above IRCoT, adds 27.32% Recall@1 over the retriever's top pick on identical candidates, and transfers unchanged to PatchFinder_top10 to lift Recall@1 from 24.28% to 39.86%.
  4. arXiv

    LEAP keeps hour-scale audio-video QA out of one giant context: 4.5–16.8% over the Qwen3-Omni-30B baseline and transfer to MiniCPM-o 4.5

    LEAP splits an hour-scale recording into fixed-duration blocks, runs a lightweight localization pass per block to score short candidate windows, then re-encodes only the top-ranked windows in a single bounded answer pass, so the answer input and peak context stay independent of recording duration; across several AVQA benchmarks it improves over the Qwen3-Omni-30B-A3B baseline by 4.5–16.8% and transfers to MiniCPM-o 4.5, surpassing its published results by 3.1–13.0%.
  5. arXiv

    EvoDuet lets LLMs retrieve the web by knowledge gap during evolutionary search, raising OpenEvolve's normalized discovery gain from 74.1% to 78.0% (GPT-5.6-Luna) and from 61.3% to 82.3% (Gemini-3.8-Flash) across 21 optimization tasks and surpassing prior best scores on eight

    The work introduces EvoDuet, a bi-level method that co-evolves solutions and web search queries with fixed model parameters: at each iteration a knowledge-gap-based retrieval gate decides whether to retrieve new documents, reuse stored ones, or proceed without them, an inner loop refines queries and ranks documents by the solution scores they are predicted to yield, and an outer loop generates candidates in parallel and records evaluated outcomes for later searches; across 21 optimization tasks with one candidate per iteration it raises OpenEvolve's normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash, Qwen3.
  6. arXiv

    Mid-Harness verifies candidate actions between model and harness, lifting TMAX-9B Pass@1 on TerminalBench-Lite from 50.00% to 68.03%

    The work introduces Mid-Harness, which samples and verifies candidate actions at the model-harness boundary before forwarding one for execution while keeping the generator and harness unchanged, and studies action-level test-time compute scaling: on TerminalBench-Lite a GPT-5.6 Sol verifier raises TMAX-9B Pass@1 from 50.00% to 68.03% with 8 sampled actions, wider sampling yields little benefit under weak verification, pairwise verification performs best among the evaluated self-verification mechanisms, distilling the stronger verifier's responses into TMAX-9B further improves Pass@1, and combining action scaling with trajectory scaling (Best-of-N and Sequential Refine) reaches higher success at lower estimated token cost.
  7. arXiv

    Removing the overlapping-window timing shortcut lets non-invasive brain-to-text reach 36.6% word error rate with five observations per word

    The work shows that most of the reported gain from jointly decoding all words in a sentence (d'Ascoli et al., 2025) is reproducible on synthetic signals containing no brain information (22.0% versus 22.3% on real MEG), because neighbouring three-second word-aligned windows overlap and implicitly leak word duration; decoding words independently instead (SimpleB2T) makes aggregation across observations and an LLM prior effective, reaching 36.6% word error rate with five observations per word on a clinically motivated benchmark.
  8. arXiv

    RegLLM diagnostic harness exposes run-to-run variance in two nominally identical GPU pilots: escalation recall 1.0 versus 0.5, with the same adapter's effect flipping direction

    The authors propose RegLLM, a diagnostic harness for bounded autonomy in regulated agentic workflows that instruments six trustworthiness signals (citation validity, source grounding, schema compliance, escalation correctness, constitutional alignment, unsafe-action rate) alongside a deterministic runtime supervisor; an offline reference run (n=12) lifts escalation recall from 0 to 0.67 and cuts the unsafe-action rate from 0.33 to 0.08, while two nominally identical single-GPU Qwen2.5-3B LoRA/DPO pilots (n=8, same seed and eval split) show task success of 0.25 versus 0.12 and escalation recall of 1.0 versus 0.
  9. arXiv

    Six-domain evaluation of GPT-6 Astra as an embodied policy: navigation leads, hybrid control lifts manipulation success, but direct in-hand control and dense-reference locomotion remain unreliable

    This work systematically evaluates GPT-6 Astra as an embodied policy across six domains—gripper manipulation, dexterous manipulation, mobile manipulation, navigation, locomotion, and humanoid loco-manipulation—comparing direct control with hybrid control that cooperates with learned policies or whole-body controllers, finding leading navigation results (92% on RxR instruction following, 82% on HM3D object search), hybrid success of 48% on RoboDojo, 50% on DexJoCo, and 38.7% on RoboCasa365, while direct in-hand control and dense-reference locomotion remain unreliable, with substantial token use and inference latency recorded.
  10. arXiv

    AIM treats research ideas as explicit search objects: 67.0% and 55.8% average scores across 10 AutoLab tasks, matching the strongest baseline up to 3.1x faster in wall-clock time

    The work introduces the Agentic Idea Manager (AIM), a fully autonomous framework that makes research ideas explicit search objects: a Bayesian-optimization-inspired Agentic Surrogate organizes ideas into semantic clusters and produces ordinal promisingness estimates, an Agentic Acquisition mechanism balances exploration and exploitation at cluster and idea level, a Solution Auditor checks idea-solution integrity, and a Resource Planner adaptively allocates parallel branches under a fixed budget; on 10 AutoLab tasks AIM reaches 67.0% average on System Optimization and 55.8% on long-horizon Model Development & CUDA, exceeding the strongest baseline ScientistOne by 1.6 and 4.9 percentage points and reaching that baseline's best score up to 3.1x faster in wall-clock time.

Page 6 · showing 10