Skip to main content

AI Core

999 items

  1. arXiv

    GEAR treats geometry as a memory address: sparse correspondence attention cuts ATE by 80.3% for minute-long camera-controlled video generation

    The work introduces GEAR, which avoids fusing history into a persistent global 3D representation and instead uses per-frame geometry to build token-level history-to-target correspondences, treating geometry as an explicit address for visual memory; Geometric Correspondence Attention then sparsely injects only geometrically matched historical features, while an Invisible Octree rejects projectable but occluded correspondences, reporting state-of-the-art visual quality, camera control, and revisit consistency on DL3DV-Evaluation and WorldScore-Static and enabling minute-long generation.
  2. arXiv

    DroneWAM cuts drone visual-navigation inference to 483 ms with JEPA latent prediction and adaptive rollout, lifting closed-loop progress from 0.510 to 0.774

    The work presents DroneWAM, an efficient world-action model for drone visual navigation that uses a JEPA-based architecture to predict future states directly in representation space, a pretrained Resampler to compress dense encoder features into fewer latent tokens, and a preference-trained Gate that adaptively allocates prediction depth per scene, alongside DroneNav-6D, a simulated dataset with synchronized RGB observations, 6-DoF flight trajectories, control commands, and randomized wind disturbances; on DroneNav-6D it achieves the best trajectory accuracy among the compared methods, and adaptive rollout reduces the average prediction depth from 8 to 4.58 while improving trajectory accuracy.
  3. arXiv

    Matched evaluation finds DPO leads overall on safety control and text monitoring, while representation probes stay competitive at far lower marginal compute

    Under one matched setup, this work compares behavioral safeguards (DPO alignment and text monitors) with representation engineering (three steering methods and four probes), finding that DPO gives the strongest overall control and improves with more training data but can lose safety after benign fine-tuning, that specialized text monitors achieve the strongest detection accuracy, and that representation probes remain competitive at substantially lower marginal cost when reusing native activations, with probe-guided interventions recovering much of the safety DPO loses after benign fine-tuning at little added over-refusal.
  4. arXiv

    Writing robot tasks as code: HexaAnything lifts RoboCasa365 Composite-Unseen success from 34.3% to 38.3% and turns Harness traces into a stronger HexaModel

    The work proposes Physical Coding, representing physical world state (Code as World) and execution procedure (Code as Policy) as executable programs, and builds the physical coding agent HexaAnything, which treats VLA/WAM policies as callable action tools while a Harness owns state bookkeeping, verification, and recovery; on RoboCasa365 with XR-1 as the action model it raises Composite-Unseen success from 34.3% to 38.3% and overall success from 56.6% to 61.1%, a HexaModel v0.1 fine-tuned on Harness-collected traces beats its base model on every split in the same Harness, PhyBench simulated physics experiments are completed with mean relative errors below 5%, and on a dual-arm AgileX robot five of seven tabletop tasks succeed in all three trials.
  5. arXiv

    BaRe-Mem uses Bayesian reliability memory to shift a central model toward autonomous reasoning as advice turns misleading, and finds capable workers earlier on MuSiQue

    The work introduces BaRe-Mem, an online Bayesian reliability memory that estimates advisor reliability conditioned on the central model's internal belief representations, uses those estimates to modulate attention and to decide whether to consult; across nine benchmarks and six central models it is more robust to misleading advisor information than debate and majority voting, stays above autonomous reasoning on the more challenging tasks across all tested misleading levels, and in MuSiQue agent-team routing identifies capable workers earlier than routing by historical success counts.
  6. arXiv

    GeoVerse injects video generative priors into geometric latent space, raising PSNR by 2.23 dB on DL3DV and cutting ATE by 32.4% on Mip-NeRF360

    GeoVerse performs diffusion generation inside the geometric latent space of a pretrained 3D foundation model (DA3), injects frozen Wan2.2 VACE video features into the denoiser through a ControlNet-style adapter, and uses a continuously updated colored point-cloud spatial memory to reproject history as target-aligned guidance, synthesizing cross-view consistent novel views from sparse inputs; across DL3DV, RealEstate10K, Mip-NeRF 360, and ScanNetV2 it reports better visual quality and geometric consistency, with PSNR 2.23 dB higher than GLD on DL3DV and ATE 32.4% lower than GLD on Mip-NeRF 360, while reflow distillation brings inference from over 100 seconds to under 10 seconds.
  7. arXiv

    CIS recasts training-inference mismatch as a log-odds displacement and tops the five-benchmark average on all three MoE models

    The work studies training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and introduces calibrated importance sampling (CIS): the mismatch is characterized as an additive displacement in log-odds, and large positive displacements are truncated at a single constant threshold, which maps back to an importance-ratio cap that tightens as token confidence increases; across three mixture-of-experts models and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average on all three models among the evaluated baselines, and diagnostic analyses show it places less truncation bias on low-confidence tokens than truncat
  8. arXiv

    PyroAdapt lifts daily California wildfire average precision from 21.62% to 24.35–24.57% and captures 344 extra positive cell-days under a fixed 34-cell budget

    The work proposes PyroAdapt, a pretrain–retrieve–rank framework that pretrains a wildfire risk model on historical data, retrieves similar historical locations using unlabeled target-year covariates, and fine-tunes through same-day fire–nonfire cell-pair ranking losses (direct ranking, residual pairwise DPO, and selective ranking); over 666 0.25° California grid cells, the three ranking objectives raise daily average precision from 21.62% under continued focal fine-tuning to 24.35–24.57% and Top5% recall from 18.70% to 22.11–22.79%, selective ranking captures 344 additional positive cell-days under a fixed daily budget of 34 cells (5% of area), and rolling evaluations over Yosemite show the ranking gains persist under temporal distribution shift.
  9. arXiv

    NTU team finds speaker-verification EER does not track human voice similarity, and an embedding-dimensionality bottleneck lifts alignment correlation from 0.08 to 0.74

    Using 16,905 different-speaker pairs from VoxSim, the study builds a human perceptual alignment metric and systematically compares seven training objectives across five model conditions, finding that verification accuracy (EER) does not track human similarity judgments while embedding effective dimensionality correlates with perceptual alignment at about -0.95 rank correlation, and that a dimensionality bottleneck on ECAPA-TDNN raises AM-Softmax alignment from 0.08 to 0.74.
  10. arXiv

    Tencent Hunyuan and collaborators compared 11 MoE rungs and found encoder-free multimodal loss falls faster with compute, crossing over near 10^22 FLOPs

    Using a matched ladder of 11 sparse MoE decoders (1.1B–44B total, 71M–2.4B active non-embedding parameters) with identical data mixture and optimization, the work fits separate text and multimodal scaling laws for encoder-free and encoder-based MLLMs and finds that removing the visual encoder shifts compute-optimal multimodal allocation toward larger models (model allocation exponent rising from about 0.464 to about 0.570) while leaving text nearly unchanged; encoder-free models underperform at small scales but their multimodal loss falls faster with compute, with extrapolation predicting an efficiency crossover on the order of 10^22 FLOPs (point estimate about 3.

Page 34 · showing 10