Skip to main content

AI Core

995 items

  1. arXiv

    WorldAttention pairs hierarchical KV caching with hybrid sparse attention to reach 22 FPS on a single H100 and 0.9472 subject consistency on VBench-Long

    The work proposes WorldAttention, an attention architecture for interactive video world models that combines a Hierarchical KV Cache (HKV) storing and retrieving page-level KV across GPU, CPU, and NVMe tiers with a Hybrid Sparse Attention (HSA) that fuses a linear global branch and a head-adaptive block-sparse branch, reporting subject consistency of 0.9472 on VBench-Long and 0.9668 on InterVBench, a 14.02x sparse-kernel speedup over FlashAttention-3, a 2.21x end-to-end speedup, and 22.0 FPS on a single NVIDIA H100.
  2. arXiv

    Org-Agent uses a task dependency graph and constraint-aware execution to extend single-user assistants into organizational agents serving multiple users

    The work introduces Org-Agent, a constraint-centric reasoning framework that decomposes a task into atomic subtasks and builds a Task Dependency Graph (TDG), schedules subtasks by topological order, and executes each subtask under constraint rules with evidence-acquisition and memory-management tools; it reaches 44.03% average accuracy on GroupMemBench versus 37.72% for BM25 and 38.79% for the strongest memory system Hindsight, and improves MUSES-Bench averages by 8.27, 4.24, and 5.29 points across three backbones, with ablations supporting both dependency modeling and tool use.
  3. arXiv

    GEB links visually grounded observations into entity biographies, lifting EgoLifeQA accuracy to 72.0%, 4.4 points above the strongest published memory framework

    The work introduces Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies and retrieves a biography alongside episodic evidence at question time; across four benchmarks including day-long and week-long recordings it improves both multiple-choice and open-ended question answering over prior memory frameworks, reaching 72.0% accuracy on EgoLifeQA, 4.4 percentage points above the best published result, with ablations showing that grounded identity association and biography reading each contribute and that additional descriptions alone do not fully recover the gains.
  4. arXiv

    EpiCon builds a shared multimodal memory bank with two 2B models, letting different agent systems reuse each other's experience and lifting macro-average scores by 1.7 to 4.9 points

    EpiCon introduces a shared multimodal memory framework in which two independently trained 2B models, a memory controller and a tree self-organizer, co-evolve question-level textual guidance with visual evidence and link it to a persistent experience bank, so different multi-agent systems can reuse and contribute experience without updating host model parameters; across eleven benchmarks, four multimodal task domains, two harnesses and multiple backbones, a frozen bank improves other systems with a single solving attempt, a second harness raises the original system's macro-average by 2.6 points, the 2B variant improves macro-average scores by 1.7 to 4.9 points over No Memory across four host configurations, and memory-operation time drops 67% to 74% relative to backbone-sized memory models.
  5. arXiv

    HybridCUA lets a 9B agent reach 53.6% on OSWorld by mixing GUI and CLI, 14.8 points above its base model

    The work builds HybridCUA-8K, a data pipeline of 5,000 GUI-only, CLI-only, and interleaved GUI–CLI trajectories plus 3,000 verified RLVR tasks, and a two-stage framework of supervised fine-tuning followed by online reinforcement learning with CLI-aware rewards; the resulting HybridCUA-9B reaches 53.6% accuracy on OSWorld (14.8 points above its Qwen3.5-9B base, with average steps falling from 31.6 to 14.0) and 36.0% on WindowsAgentArena (a 4.0-point gain).
  6. arXiv

    TAPS adapts recurrent step size to the trajectory, lifting Sudoku accuracy to 91.39% and Maze to 79.90% while cutting the loops needed for matched quality

    The work introduces the Trajectory Adaptive Progress–Fluctuation Scheduler (TAPS), which uses exponential moving averages to estimate persistent progress and centered fluctuation in recurrent updates and adapts the per-step update scale online; it proves that under stated conditions TAPS reduces expected terminal loss and reaches a target quality in fewer loops, and empirically improves terminal accuracy with up to roughly 1.5x wall-clock speedup across Sudoku, Maze, language-model recurrence, and intermediate-layer recurrence.
  7. arXiv

    CorpusMap links a corpus into an entity graph, raising answer quality by 6.4–11.7 points and cutting input tokens by 34–57% across 7 models and 3 benchmarks

    The work introduces CorpusMap, an offline-built, entity-anchored navigation layer that renders each recurring cross-document entity as a source-attributed Entity Page linked to every document mentioning it, and shows across 7 models and 3 multi-document QA benchmarks (EnterpriseRAG-Bench, WixQA, HERB) that it improves both evidence discovery and answer quality over raw-corpus agentic search while using fewer tokens on average, and further outperforms 4 alternative navigation layers.
  8. arXiv

    Omni-IO Skills lifts GPT-5.6 Sol and Claude Sonnet 5 multimodal input support from about 40% to 100% with 27 skills

    The work presents Omni-IO Skills, a plug-and-play Agent Harness that makes existing agents omni-native through hierarchical Skills, a standardized multimodal execution interface, dependency-aware orchestration, and a persistent Asset Registry, representing multi-asset workflows as Declare Execution Graphs; on UniM-90 it raises the input-support rates of GPT-5.6 Sol and Claude Sonnet 5 from 40.00% and 38.89% to 100%, increases relative Semantic–Quality Coupled Score from 26.99 to 74.94 and from 27.82 to 77.78, and reaches Strict Structure Scores of 100.00 and 99.78.
  9. Microsoft Research

    Microsoft Research introduces Quine: a multimodal biology world model that prioritized compounds driving tumor-state shifts in pancreatic cancer, narrowing candidates in a single weekend

    Microsoft Research introduces Quine, a research system combining a world model of biology trained jointly across modalities including sequence, structure, function, cellular state, and imaging with a harness connecting scientific tools, literature, the wet lab, and researchers; with the Broad Institute it used the system in pancreatic ductal adenocarcinoma (PDAC) to predict and prioritize thousands of compounds by their potential to shift tumor cells between therapeutically relevant states, wet-lab assays showed Quine's highest-ranked compounds produced the largest intended shifts, the whole process from narrowing the search space to prioritizing a handful of candidates took just one weekend, and the experiments also bore out the model's prediction of a distinct third phenotype.
  10. MIT News - Artificial intelligence

    MIT researchers systematically assess objections to algorithmic monoculture: systematic exclusion fails, but information echo chambers hinder exploration, and ensembling can partly offset it

    In Philosophical Perspectives, MIT's Brian Hedden and Manish Raghavan systematically evaluate major objections to algorithmic monoculture—where one algorithm makes all decisions in a domain—arguing that objections such as systematic exclusion do not hold, proving mathematically that monoculture tends to create information echo chambers that hinder exploration, and showing through a series of hiring simulations that bundling algorithms into an "ensemble" can sometimes overcome this limitation so that monoculture performs as well as or better than polyculture.

Page 30 · showing 10