AI Core
999 items
DroneWAM cuts drone visual-navigation inference to 483 ms with JEPA latent prediction and adaptive rollout, lifting closed-loop progress from 0.510 to 0.774
The work presents DroneWAM, an efficient world-action model for drone visual navigation that uses a JEPA-based architecture to predict future states directly in representation space, a pretrained Resampler to compress dense encoder features into fewer latent tokens, and a preference-trained Gate that adaptively allocates prediction depth per scene, alongside DroneNav-6D, a simulated dataset with synchronized RGB observations, 6-DoF flight trajectories, control commands, and randomized wind disturbances; on DroneNav-6D it achieves the best trajectory accuracy among the compared methods, and adaptive rollout reduces the average prediction depth from 8 to 4.58 while improving trajectory accuracy.
Matched evaluation finds DPO leads overall on safety control and text monitoring, while representation probes stay competitive at far lower marginal compute
Under one matched setup, this work compares behavioral safeguards (DPO alignment and text monitors) with representation engineering (three steering methods and four probes), finding that DPO gives the strongest overall control and improves with more training data but can lose safety after benign fine-tuning, that specialized text monitors achieve the strongest detection accuracy, and that representation probes remain competitive at substantially lower marginal cost when reusing native activations, with probe-guided interventions recovering much of the safety DPO loses after benign fine-tuning at little added over-refusal.
Writing robot tasks as code: HexaAnything lifts RoboCasa365 Composite-Unseen success from 34.3% to 38.3% and turns Harness traces into a stronger HexaModel
The work proposes Physical Coding, representing physical world state (Code as World) and execution procedure (Code as Policy) as executable programs, and builds the physical coding agent HexaAnything, which treats VLA/WAM policies as callable action tools while a Harness owns state bookkeeping, verification, and recovery; on RoboCasa365 with XR-1 as the action model it raises Composite-Unseen success from 34.3% to 38.3% and overall success from 56.6% to 61.1%, a HexaModel v0.1 fine-tuned on Harness-collected traces beats its base model on every split in the same Harness, PhyBench simulated physics experiments are completed with mean relative errors below 5%, and on a dual-arm AgileX robot five of seven tabletop tasks succeed in all three trials.
BaRe-Mem uses Bayesian reliability memory to shift a central model toward autonomous reasoning as advice turns misleading, and finds capable workers earlier on MuSiQue
The work introduces BaRe-Mem, an online Bayesian reliability memory that estimates advisor reliability conditioned on the central model's internal belief representations, uses those estimates to modulate attention and to decide whether to consult; across nine benchmarks and six central models it is more robust to misleading advisor information than debate and majority voting, stays above autonomous reasoning on the more challenging tasks across all tested misleading levels, and in MuSiQue agent-team routing identifies capable workers earlier than routing by historical success counts.
GeoVerse injects video generative priors into geometric latent space, raising PSNR by 2.23 dB on DL3DV and cutting ATE by 32.4% on Mip-NeRF360
GeoVerse performs diffusion generation inside the geometric latent space of a pretrained 3D foundation model (DA3), injects frozen Wan2.2 VACE video features into the denoiser through a ControlNet-style adapter, and uses a continuously updated colored point-cloud spatial memory to reproject history as target-aligned guidance, synthesizing cross-view consistent novel views from sparse inputs; across DL3DV, RealEstate10K, Mip-NeRF 360, and ScanNetV2 it reports better visual quality and geometric consistency, with PSNR 2.23 dB higher than GLD on DL3DV and ATE 32.4% lower than GLD on Mip-NeRF 360, while reflow distillation brings inference from over 100 seconds to under 10 seconds.
CIS recasts training-inference mismatch as a log-odds displacement and tops the five-benchmark average on all three MoE models
The work studies training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and introduces calibrated importance sampling (CIS): the mismatch is characterized as an additive displacement in log-odds, and large positive displacements are truncated at a single constant threshold, which maps back to an importance-ratio cap that tightens as token confidence increases; across three mixture-of-experts models and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average on all three models among the evaluated baselines, and diagnostic analyses show it places less truncation bias on low-confidence tokens than truncat
PyroAdapt lifts daily California wildfire average precision from 21.62% to 24.35–24.57% and captures 344 extra positive cell-days under a fixed 34-cell budget
The work proposes PyroAdapt, a pretrain–retrieve–rank framework that pretrains a wildfire risk model on historical data, retrieves similar historical locations using unlabeled target-year covariates, and fine-tunes through same-day fire–nonfire cell-pair ranking losses (direct ranking, residual pairwise DPO, and selective ranking); over 666 0.25° California grid cells, the three ranking objectives raise daily average precision from 21.62% under continued focal fine-tuning to 24.35–24.57% and Top5% recall from 18.70% to 22.11–22.79%, selective ranking captures 344 additional positive cell-days under a fixed daily budget of 34 cells (5% of area), and rolling evaluations over Yosemite show the ranking gains persist under temporal distribution shift.
NTU team finds speaker-verification EER does not track human voice similarity, and an embedding-dimensionality bottleneck lifts alignment correlation from 0.08 to 0.74
Using 16,905 different-speaker pairs from VoxSim, the study builds a human perceptual alignment metric and systematically compares seven training objectives across five model conditions, finding that verification accuracy (EER) does not track human similarity judgments while embedding effective dimensionality correlates with perceptual alignment at about -0.95 rank correlation, and that a dimensionality bottleneck on ECAPA-TDNN raises AM-Softmax alignment from 0.08 to 0.74.
Tencent Hunyuan and collaborators compared 11 MoE rungs and found encoder-free multimodal loss falls faster with compute, crossing over near 10^22 FLOPs
Using a matched ladder of 11 sparse MoE decoders (1.1B–44B total, 71M–2.4B active non-embedding parameters) with identical data mixture and optimization, the work fits separate text and multimodal scaling laws for encoder-free and encoder-based MLLMs and finds that removing the visual encoder shifts compute-optimal multimodal allocation toward larger models (model allocation exponent rising from about 0.464 to about 0.570) while leaving text nearly unchanged; encoder-free models underperform at small scales but their multimodal loss falls faster with compute, with extrapolation predicting an efficiency crossover on the order of 10^22 FLOPs (point estimate about 3.
Page 34 · showing 10