Skip to main content

AI Core

970 items

  1. The latest research from Google

    Google Research introduces Diffusion Controller: a lightweight steering-damper network that beats LoRA on HPS-v2 win rates in gray-box settings, with a white-box version reaching a 90% win rate over baseline

    Google Research engineers Chih-wei Hsu and Moonkyung Ryu present the Diffusion Controller framework, which reframes the diffusion denoising process as a smooth continuous control problem and uses a lightweight steering-damper network to dynamically correct the generation trajectory while the base model stays frozen; evaluated on a Stable Diffusion v1.4 backbone across SFT, RWL, and PPO regimes with the standardized Human Preference Score (HPS-v2), the framework is reported to outperform corresponding baselines in both white-box and gray-box settings, with the gray-box version beating LoRA on HPS-v2 win rates in the SFT and RWL tracks while manipulating significantly fewer internal model layers, and the white-box version achieving a 90% win rate over the baseline, all with a single inferenc
  2. NVIDIA Technical Blog

    NVIDIA VSS Blueprint 3.3 builds a visual AI agent from one prompt in under 30 minutes and cuts VLM input tokens by 80%

    NVIDIA released VSS Blueprint 3.3 with a Build Vision Agent skill (vss-build-vision-ai) and Adaptive Efficient Video Sampling (Adaptive EVS): the former lets a coding agent turn a natural-language request into a deployment by starting from one of four validated profiles and computing the smallest delta, delivering an orange-juice bottling-line overflow agent as a live, previewable deployment in under 30 minutes on a two-GPU RTX PRO 6000 Blackwell host; the latter, running Cosmos 3 Super FP8 on an RTX PRO 6000 Blackwell, cut alert contextualization latency from 1,021 ms to 844 ms (17%), raised concurrent real-time VLM streams from 13 to 19 (46%), and summarized a 60-minute video in about half the time with 80% fewer VLM input tokens.
  3. arXiv

    MaLiang-Harness generates images and video from executable programs: GPT-6-Astra reaches 100% generation success on both benchmarks, with 96.0% of image and 76.9% of video tasks meeting all quality thresholds

    The work introduces MaLiang-Harness, a framework that organizes MLLM-driven image and video generation as a persistent process of construction, inspection, and revision, in which a Persistent Executable Generation state, a Traceable Generation Process, and Revision-aware Editing and Verification share a common revision reference; it evaluates 11 and 4 closed-source MLLMs on MaLiang-IBench (50 text-to-image prompts) and MaLiang-VBench (13 text-to-video prompts), measuring GPT-6-Astra at 100% generation success on both benchmarks, with 96.0% of image tasks and 76.9% of video tasks meeting all quality thresholds.
  4. arXiv

    VoxMem tests 15 audio LLMs on 3,196 questions: none tops 40% at 32K, and swapping audio for transcripts drops speaker accuracy from 69.8% to 10.3%

    The authors propose a two-axis taxonomy of spoken conversational memory — acoustic evidence type (speech semantics, speaker identity, paralinguistic cues, environmental sound) crossed with memory operation (information extraction, multi-session reasoning, temporal evolution tracking, answer refusal) — and build VoxMem on it: 3,196 evaluation instances over 34,743 spoken sessions (177 hours) at four context budgets from 8K to 64K tokens; evaluating 15 large audio language models, no model exceeds 40% overall accuracy at 32K (best 38.5%), models remember what was said far better than who said it, how it was said, or what was audible, and replacing audio with exact transcripts drops speaker accuracy from 69.8% to 10.3% while speech-semantics accuracy barely moves (75.9% to 71.0%).
  5. arXiv

    ActFirst-OPD lets multi-turn agents act before reasoning, speeding on-policy distillation training 2.3x, 1.8x and 4.9x on ALFWorld, WebShop and ScienceWorld

    The work proposes ActFirst-OPD, which decouples environment interaction from full-response generation: the student infers and executes actions through reference-conditioned inverse dynamics using its current interaction context and a reference next observation, switches to autonomous next-action prediction once the transition deviates from the reference trajectory, and asynchronously generates full think-then-act responses from the collected interaction contexts for token-level teacher supervision; across 0.6B, 1.7B and 4B Qwen3 students it achieves average wall-clock training speedups of 2.3x on ALFWorld, 1.8x on WebShop and 4.9x on ScienceWorld over Vanilla OPD, while matching or exceeding mean task success rate in eight of nine benchmark-model settings.
  6. arXiv

    SaveRouter Cuts LLM Routing's Break-Even Deployment Volume by Up to About 9.5x Using Roughly a Third of the Supervision

    The work proposes SaveRouter, a sparse-supervision routing framework that selectively acquires informative query-model feedback and shares capability information across related queries, using only about 33-41% of available training feedback across four routing benchmarks while maintaining competitive or better routing quality and reducing the break-even deployment volume by roughly 1.9-9.5x relative to the fastest conventional fully supervised router.
  7. arXiv

    PanoVLN lifts R2R-CE success rate to 77.3%, 11.9 points above the previous best, using panoramic vision

    PanoVLN performs vision-and-language navigation in continuous environments with a 4B vision-language model and RGB-only panoramic input, combining longer action-sequence supervision, confidence-guided execution, decision-centric training data, and fused panoramic geometric features to reach 77.3% and 78.0% success rate on R2R-CE and RxR-CE Val-Unseen, exceeding the previous best by 11.9 and 8.7 percentage points, and enabling real-world quadruped navigation with fewer pauses.
  8. arXiv

    HDL locates branch points by hindsight divergence, cutting generated tokens 35–61% and lifting agent tasks by up to 12.46 points

    The work introduces Hindsight-Divergence Localization (HDL), which picks branch points in a trajectory by how much token log-likelihoods change once verifier feedback and a reflection are added, then builds each training group from a few complete root trajectories plus continuations that reuse the root prefix; across math, code, and agent tasks with three models it cuts generated tokens by 35–61% and rollout wall-clock time by 18–45% relative to GRPO while improving task performance, with gains up to 12.46 percentage points on agent tasks.
  9. arXiv

    EasyPPO traces two critic failure modes in PPO and stays stable across three tasks, gaining 14.89%, 2.28% and 9.47% over PPO

    The work identifies two critic failure modes that destabilize PPO for LLM reinforcement learning — filtering truncated rollouts from both actor and critic shifts the policy objective to reward conditioned on completion, and heterogeneous return noise across prompts lets high-variance prompts dominate critic updates in finite batches — and introduces EasyPPO, which combines actor-only overlong filtering, noise-normalized critic regression weighting each prompt by the inverse standard deviation of its sampled returns, and moderately smaller critic mini-batches with gradient clipping, remaining stable across continuous-reward coding on FrontierCS, binary-reward mathematical reasoning on AIME24 and multi-turn search on Search-R1, with best validation scores improving over PPO by 14.89%, 2.
  10. arXiv

    WorldLine pretrains on 10,000 hours of action-free robot video and grounds 2,000 hours of action trajectories across ten embodiments, lifting robot-mask IoU by 0.1626 on failed trajectories and predicting trajectory success at 74% mean accuracy

    WorldLine introduces an action-driven visual simulator that decouples transferable robot-object dynamics learning from heterogeneous action grounding: it pretrains on more than 10,000 hours of action-free robot videos, grounds them with over 2,000 hours of action trajectories across more than ten embodiments through image-space action maps, adds multi-view, failure-enriched and relational-regularized training, and distills a robot-focused few-step causal rollout; across held-out and out-of-domain settings it maintains visual quality and robot-motion agreement, improves robot-mask IoU by 0.1626 over the strongest baseline on failed trajectories, predicts trajectory success with 74% mean accuracy on RoboTwin and AgiBot, and improves task success by up to 21.

Page 19 · showing 10