Skip to main content

AI Core

974 items

  1. arXiv

    Real2Gym turns human demonstration videos into executable simulation gyms and trains a failed Franka task into success

    Real2Gym is an agentic Real2Sim2Real framework that reconstructs human and robot demonstration videos into visually aligned, natively physics-validated Blender and MuJoCo interactive environments, where an agent generates executable code and distills successes and failures into reusable skills, reaching 87.5% task success with roughly 75% fewer policy-execution tokens than GPT-6 Astra across 24 reconstructed DROID and EgoDex environments and turning a zero-shot real-robot failure on narrow-clearance plate placement into success after simulation-based evolution on a Franka arm.
  2. arXiv

    Sparse crosscoders show on-policy distillation adds no new student features but reweights features the student already shares with the teacher

    Using sparse crosscoders and a proposed swap readout, the study reads how each student checkpoint uses every feature before and after on-policy distillation (OPD) across three OPD settings, finding that OPD neither creates the student's own features nor passes on the teacher's, but slightly reweights features the student already shares with the teacher (over 98% of frequently used features change firing rate by less than 20%), with the largest changes concentrated on decision tokens such as Wait, Hmm, and So; the SFT warm-up on the teacher's rollouts that commonly precedes OPD also adds no features, reweighting shared features partly along OPD's direction and partly beyond it, and imposing this reweighting on a directly distilled student's features without changing its weights brings its a
  3. arXiv

    FurE uses human-hair priors to cut per-strand fur training from 10.5 hours to 52 minutes and works on a real bison sequence

    FurE is a strand-based animal fur reconstruction method that optimizes a root-conditioned low-dimensional latent field over multi-view images, decodes it into strand geometry with a PCA-based decoder, and reconstructs a defurred body from surface-constrained Gaussian Frosting thickness cues plus part-based priors; without any animal-fur dataset it cuts strand training from NeuralFur's 10.5 hours to 52 minutes (about a 10x speedup), matches or improves rendering and strand metrics on Artemis synthetic scenes, and, to the authors' knowledge, is the first to reconstruct instance-specific strand-based fur from a noisy real-world bison multi-view sequence.
  4. arXiv

    SoL-Refiner refines low-resolution generated video to 4K in one denoising step, cutting 2K refinement latency from 57.461 to 6.447 seconds

    The work presents SoL-Refiner, a one-step video refiner that turns low-resolution generator outputs into 4K video through a three-stage recipe of high-resolution continual training, reinforcement-learning post-training with frame-based reward models, and final one-step distillation, and introduces Refiner-Bench, a video refinement benchmark using a shared-input protocol to compare refiners at roughly 2K output resolution; at 2K the one-step model outperforms all evaluated external refiners on VBench and UniPercept averages, at 3840×2176 it improves both metrics over the three-step LTX-2.3 Refiner, and with the complete acceleration stack it achieves an 8.91× speedup in refinement latency over the same baseline in the 2K latency setting.
  5. arXiv

    PlaylistEval tests video-language judges on ~100-hour playlists: best judge reaches only 75.4% pairwise accuracy against 93.0% human agreement

    The work introduces PlaylistEval, an agentic framework that builds video-language judge benchmarks without human annotation by automatically generating questions whose evidence is scattered across two distant segments of ~100-hour playlists and whose wrong answers differ only visually, yielding PlaylistBench with 630 pairs across seven domains; on a stratified subset of 152 pairs it agrees with human judgments 93.0% of the time (IAA 0.781), while evaluating 17 omnimodal and multimodal models from eight families shows frontier judges reach only 75.4% pairwise accuracy, open-source judge models perform far behind, retrieval improves accuracy by up to 10.5 points yet the best retriever finds both relevant segments in the top 10 only 37.
  6. arXiv

    Meituan's LongCat team splits deep research into planning, parallel section research, and global-to-local editing via ResearchSpec, scoring 55.25, 51.35, and 79.83 on three public benchmarks

    Meituan's LongCat team presents LongCat-DeepResearch, which shifts early research iteration from the full report to an executable ResearchSpec: multiple planning agents first search external sources and consolidate a research plan, researchers then investigate and draft citation-bearing sections in parallel independent contexts, and a Global Editor assigns cross-section ownership while Local Editors make targeted revisions, yielding 55.25 on DeepResearchBench, 51.35 on DeepResearchBench II, and 79.83 on ResearchRubrics, and 76.04 on an in-house benchmark, second among four compared systems.
  7. arXiv

    StructRL lifts long-horizon VLA success from 41.5% to 49.1% with verifiable subtask rewards

    StructRL is an online reinforcement learning framework that automatically decomposes each long-horizon instruction into subtasks verifiable by binary environment-state criteria and organizes them into ordered dependency groups, granting intermediate rewards only after all prerequisites of a subtask are complete and scaling each reward by completion pace; across RoboCasa365 and LIBERO-Long with the GR00T-N1.5 and π0.5 VLA backbones it consistently outperforms the evaluated online RL baselines (e.g., on GR00T-N1.5, RoboCasa365 rises from 41.5% to 49.1% and LIBERO-Long from 92.4% to 96.6%).
  8. arXiv

    Yandex team's AsyncLLM lets Qwen 3.x models watch, think and act concurrently without fine-tuning, speeding up streaming video, games and system monitoring

    The work introduces AsyncLLM, a general framework built on the Python asyncio async/await programming model that lets pretrained LLMs observe, reason and act as concurrent coroutines, sharing memory in real time through CacheBlocks (attention KV caches and Gated DeltaNet recurrent states) and cache views, and demonstrates asynchronous operation with the same family of Qwen 3.x vision-language models on streaming video understanding, ViZDoom games and DevOps-Gym system monitoring without task-specific training.
  9. arXiv

    EngiWorld tests 1,301 real engineering tasks and finds the strongest model reaches only 44.3 EngiScore, with 3.6% success on multi-software attempts

    The authors built EngiWorld, a benchmark structured around the complete engineering design loop, with 1,301 expert-curated tasks across 6 domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, evaluated through a unified domain-verifier suite that checks geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts; across seven frontier models the strongest reaches an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed.
  10. arXiv

    ANTMAN replaces static partitioning with a revisable Need Graph: a 16x larger search space raises active coordination only 1.23x versus over 15x for partition-driven baselines

    The work introduces ANTMAN, a multi-agent information-seeking framework that treats evolving unresolved information needs, rather than input partitions, as the unit of runtime coordination, maintaining a revisable Need Graph that tracks unresolved requirements, accumulated evidence, prior attempts, and search progress to control worker selection, routing, and task-local recovery; across multi-document QA, controlled long-context scaling, and structured navigation, ANTMAN increases active coordination by only 1.23x under a 16x increase in searchable context, compared with more than 15x for partition-driven baselines, while preserving strong answer quality and retaining most of its performance when execution is delegated to substantially smaller 8B worker models.

Page 23 · showing 10