AI Core
903 items
MemLife builds entity-grounded first-person text memories read by a time-indexed agentic reader, beating the strongest training-free baseline by 4.6–12.0% on four long-horizon egocentric video benchmarks, with MemOpt adding 2.7–5.0% by training only the memory writer under the FIRM reward
The work introduces MemLife, a multimodal memory system that compacts egocentric video into time- and entity-anchored first-person text episodes retrieved by a time-indexed agentic reader, improving over the strongest training-free baseline by 4.6–12.0% across four long-horizon benchmarks without training or query-time video access, and further proposes MemOpt, which applies reinforcement learning only to the memory writer under the faithful, informative, and retrievable FIRM reward, adding 2.7–5.0% with gains that transfer across writer and reader backbones, memory systems, and out-of-domain video and question distributions.
A survey recasts attention evolution as contextual-memory organization, using 59 release records and 11 open-weight endpoints
Treating model-internal contextual memory as the shared unit of analysis, this survey introduces five dimensions—Memory Representation, Memory Update, Access, Readout, and Integration—and reviews five research lines (Softmax Attention, Sparse Attention, Linear Attention, State Space Models, and Hybrid Architectures), using 59 release-level records across 14 major model lineages and a frozen comparison of 11 high-performing open-weight endpoints through September 22, 2026 to argue that attention design is diversifying rather than converging, with hybridization and cross-layer reuse making network depth a dimension along which contextual memory is organized.
Across six pretrained language models, residual streams favor their own endpoint early while competitor sets shrink with changing membership
By comparing each context's intermediate residual states with its own final state and an empirical bank of final states from other contexts, the study finds across Gemma-2B, Gemma-7B, Qwen2.5-1.5B, Qwen2.5-7B, Mistral-7B, and Llama-3-8B that the own endpoint becomes preferable to the average alternative at the earliest measured layer while many individual endpoints remain closer, that competitor sets generally shrink with depth yet show both entries and exits, and that directional alignment can improve while Euclidean distance to the final state changes little; the authors add a high-dimensional model separating norm, alignment, and endpoint geometry and prove that a straight path toward the own endpoint cannot introduce new competitors.
MIST stress test: adding an image shifts about a fifth of VLM judge labels, regardless of what the image shows
The authors built MIST (the Misleading-Image Stress Test), 200 English sentences each containing a phrase readable either figuratively or literally and shown with an aligned image, a misleading image, or no image; across thirteen VLM judges an aligned image changed 20.5% of labels and a misleading one 19.4%, versus 11.6% when only the ignore-the-image instruction was deleted with the image left in place, and only 37% of the labels that differ between the two images moved toward the sense shown, while agreement with the human annotators was unchanged whether the image was absent, aligned, or misleading.
JEV-like direct-decision models use only 67–76% of the effective ordinal label space across 36 datasets, and BA-LoRA post-training lifts utilization from about 47% to 86%
Analyzing JEV 1.13 and three open KEV models (0.8B/4B/9B), the study finds that on ANLI JEV reaches 74.95% accuracy yet assigns 38.8% of predictions and 51.3% of errors to Neutral, and across 36 ordinal datasets final decisions cover only 67–76% of the effective gold support (versus 87–102% on four nominal tasks); through candidate-order randomization, same-item scale refinement from K=2 to K=14 (utilization falling to 26–75% at K=14), and targeted BA-LoRA post-training (gold-relative utilization rising from roughly 47% to 86% on eight supervised scales), it shows this ordinal scale-utilization bias is a learned and modifiable decision-stage compression of the candidate space.
Scaling synthetic pre-pretraining from 1B to 7B and 100B tokens keeps the token savings, but the authors trace them to long-range retrieval rather than a grammatical prior
The study systematically evaluates synthetic pre-pretraining (PPT) across four parameter scales from 500M to 7B, four pre-training data mixtures, five PPT tasks, and pre-training budgets up to 100B tokens, finding that downstream performance and token-efficiency gains persist at scale (saving at least 21B pre-training tokens at 3B) but that there is no consistent evidence the gains come from a grammatical prior; instead they arise from PPT tasks that improve long-range retrieval, are insensitive to code and math share, and diminish only when web text is absent.
Keeping only the top-attention tokens without retraining, nine checkpoints need effective attention sets of 7 to 283 tokens, and longer context raises the required set size
Without retraining, this work retains only the tokens with the highest attention weights at each head, layer, and query while keeping their original weights unchanged, and estimates the effective attention set size needed to stay within a chosen loss tolerance by measuring the increase in negative log-likelihood (NLL); it finds that relatively small selected sets keep NLL close to the full-attention baseline, that attention-based selection substantially outperforms random selection, that the required set size grows with context while its fraction of context decreases, that in BABILong experiments with a fixed annotated supporting fact additional background pushes support tokens down the attention ranking and reduces their attention mass, and that renormalizing the retained weights can subs
Galahad turns LLM document reading into a one-time cost with byte-exact KV memory: 100 of 100 on a 97,000-token corpus at 0.59–0.64 s per question
Sietse Schelpe of Corbenic AI presents Galahad, a memory layer in which Taliesin saves and reloads byte-exact KV state and Blaise passes the model only the section a question needs; on a recall test with 100 facts hidden in a 97,000-token corpus (Gemma 4 31B), Taliesin alone answered 98 of 100 on llama.cpp at 3.0 s and 572 J per question against 10 of 100, 9.3 s and 2,754 J without Galahad, and with Blaise added the model answered 100 of 100 on all three runtimes at 0.59–0.64 s and 200–213 J, while a tuned RAGFlow pipeline answered 77.
VoxParity tests 28 voice agents on 183 scenarios: when audio calls for protection, every system more often executes the routine request (41% against 12%)
The work introduces VoxParity, a benchmark of 183 scenarios from 14 sectors in which one transcript stays fixed while the audio changes (a coaching voice, a medical monitor beeping, a mayday under a radio check, noise over a drug name, a child's voice placing a bet, a frightened whisper) and the correct executable tool call changes with it; across 206 cue-bearing cells, all 28 systems (including nine production realtime agents) execute the routine request on protective calls at 41% against 12% over-triggering on clean calls, and only 11 of the 23 systems with a transcript path pass the words-only null test, with passes coming almost entirely from items that state the rule.
Page 4 · showing 10