AI Core
993 items
VoxPolyMem pairs interaction-aware hierarchical memory with an EG-GRPO retrieval policy to score 85.0 on the multi-party spoken-memory benchmark VoxPolyBench, 23.6 points above the strongest baseline
The work proposes VoxPolyMem, an interaction-aware multimodal long-term memory framework for multi-party spoken conversations that combines incremental speaker identification with a memory hierarchy of interaction memory, fact memory, and participant profiles, formulates retrieval as sequential decision-making, and trains it with Evidence-Gain GRPO (EG-GRPO) to reward newly acquired supporting evidence round by round; it also builds VoxPolyBench (18 scenarios, 176 sessions, 18.9 hours of synthesized speech, 1,527 QA pairs), on which VoxPolyMem scores 85.0 overall, surpassing the strongest evaluated baseline by 23.6 points, and scores 89.6 and 74.4 on Mem-Gallery and H2HMem-Multi, exceeding the strongest public memory baselines by more than 8 points each.
Deferring particle expansion to late transformer layers and adding time-dependent scoring-rule schedules lets a DDM trained from scratch in one stage reach 4.48 FID at 4 steps and 2.38 at 50 steps on ImageNet-256²
To address two obstacles in scaling Distributional Diffusion Models (DDMs)—multi-particle training overhead that grows with the number of particles, and globally fixed scoring-rule hyperparameters that force a single trade-off across sampling budgets—the work defers particle expansion to late transformer layers and introduces time-dependent scoring-rule schedules informed by the dynamical regimes of Biroli2024; combined with a DiT-based latent setup, this makes DDM training practical on class-conditional ImageNet-256², reaching 4.48 FID at 4 steps and 2.38 at 50 steps with DiT-XL/2 from a single model trained from scratch in one stage, without a teacher, self-distillation or JVPs, and the same recipe transfers to text-to-image generation.
A 100-page, 412-reference survey splits robot in-context learning into four interfaces and argues for evaluating 'did it infer the teaching' separately from 'can it execute after objects change'
Organized around how contextual evidence connects to execution, this survey sorts the robot in-context learning (ICL) literature into four families — context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution — compares their transfer assumptions and the roles of training, correspondence, and memory across manipulation and navigation, and proposes evaluation practices that separate responsiveness to teaching, physical transfer, and benefits from retained experience, together with an agenda linking compositional task acquisition, faithful transfer, and physical recursive self-improvement.
Ranking positions by chat structure lets auditors explain 5% of tokens and keep nearly all threat-detection success
Across about 4.7 million natural language autoencoder (NLA) explanations on four models and four auditing datasets, the study compares thirteen position-scoring signals computable in a single forward pass against a ranker trained only on chat structure, finding that chat structure beats the best individual signal in twelve of fourteen model-dataset cells, that explaining just 5% of positions retains nearly all of the success rate from explaining every position on three of four datasets, and that pretrained verbalizers recover words a fine-tuned model learned to conceal without additional verbalizer training.
AnyStep-WAM distills frozen-teacher trajectories and schedules budgets by risk and benefit, cutting denoising steps by roughly half to 85% across three world-action models while holding success rates
The work introduces AnyStep World Action Model, a framework that performs budget-aligned flow-map distillation from frozen-teacher trajectory intervals and trains a lightweight risk-benefit scheduler to predict teacher-trajectory difficulty and budget-specific student fidelity from a single one-step preview, selecting the smallest denoising budget that meets a fidelity requirement; on RoboTwin 2.0 it reduces average denoising steps by 60.2%, 49.8%, and 85.28% on Motus, FastWAM, and LingBotVA while keeping average success within 0.24 percentage points of full-budget baselines, raises one-step success by 7.07, 12.08, and 8.94 percentage points respectively, and achieves 1.67-6.14x per-call speedups on six real-world manipulation tasks.
Same bytes, two kinds of authority: splitting forged chat-template markers into ordinary subwords cuts injection success by 39 to 66 percentage points on Llama-3.1, GLM-4.5 and Seed-OSS-36B
Holding bytes identical and changing only whether forged chat-template markers are encoded as reserved token ids or as ordinary subwords, this study measures the resulting gap in indirect prompt-injection success on LLM agents in InjecAgent, finding an identity gap of 39 to 66 percentage points on Llama-3.1, GLM-4.5 and Seed-OSS-36B while Qwen3-8B draws most of its attack from marker text, and further showing that the authority lives in the single learned input vector at the marker position and that the standard tokenizer-side mitigation misses the tool-protocol tokens that carry tool output.
AI 'speech clock' estimates ageing speed from four minutes of speech and rates cognitively impaired voices as older
Ibáñez and colleagues report in Science Advances a 'speech clock': they recorded 2,928 Spanish speakers from Argentina, Chile, Colombia, Mexico and Peru across various speech tasks, used machine learning to extract more than 700 speech features that change with ageing and dementia (such as pitch and vocabulary range), trained a model to predict age and compute a 'speech age gap', and found the clock could distinguish healthy individuals from those with some form of cognitive impairment, rating the speech of people with cognitive issues as older than expected for their chronological age while healthy people's speech generally matched their age.
WorldAttention pairs hierarchical KV caching with hybrid sparse attention to reach 22 FPS on a single H100 and 0.9472 subject consistency on VBench-Long
The work proposes WorldAttention, an attention architecture for interactive video world models that combines a Hierarchical KV Cache (HKV) storing and retrieving page-level KV across GPU, CPU, and NVMe tiers with a Hybrid Sparse Attention (HSA) that fuses a linear global branch and a head-adaptive block-sparse branch, reporting subject consistency of 0.9472 on VBench-Long and 0.9668 on InterVBench, a 14.02x sparse-kernel speedup over FlashAttention-3, a 2.21x end-to-end speedup, and 22.0 FPS on a single NVIDIA H100.
Org-Agent uses a task dependency graph and constraint-aware execution to extend single-user assistants into organizational agents serving multiple users
The work introduces Org-Agent, a constraint-centric reasoning framework that decomposes a task into atomic subtasks and builds a Task Dependency Graph (TDG), schedules subtasks by topological order, and executes each subtask under constraint rules with evidence-acquisition and memory-management tools; it reaches 44.03% average accuracy on GroupMemBench versus 37.72% for BM25 and 38.79% for the strongest memory system Hindsight, and improves MUSES-Bench averages by 8.27, 4.24, and 5.29 points across three backbones, with ablations supporting both dependency modeling and tool use.
Page 29 · showing 10