AI Core
1006 items
AdaTutoRank turns a set-level rubric score into token-level credit via adaptive tutoring, reaching the best overall score across ten benchmarks with fewer retrieval calls
The work proposes AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level, nine-dimension rubric hierarchy that assigns each rollout one of three hint forms—rubrics, a corrective reflection, or a sibling set—matched to its reward, and distills the hint's effect into a token-level advantage combined with the group-relative outcome advantage; across ten benchmarks spanning RAG, deep research, and setwise evaluation it attains the best overall performance while reducing the agent's retrieval calls.
ZJU team releases EMem-Bench: across 2,554 long-horizon embodied episodes the strongest model Gemini-3-Flash reaches only 64.2% average success, while its EMem memory system lifts GPT-5.4-mini by 23.6 points
The OmniAI Group of ZJU ACES Lab introduces EMem-Bench, a long-horizon embodied memory benchmark of 2,554 executable episodes spanning Passive Observation, Dynamic Tracking, Interaction Failure, and Experience Generalization, together with an external memory system EMem and an 8B policy EMem-8B; evaluating 16 open-source and proprietary MLLMs plus representative multimodal memory systems shows the strongest proprietary model Gemini-3-Flash reaches only 64.2% average success rate and most open-source models fall below 45%, while EMem achieves the best overall performance among matched backbones, improving Mistral-Small-3.1-24B by 16.4 points and GPT-5.4-mini by 23.6 points, with EMem-8B improving over Qwen3-VL-8B by 20.6 points.
EAPO couples entropy with the advantage sign for asymmetric credit assignment, topping average accuracy across four backbones
The work introduces Entropic Advantage Policy Optimization (EAPO), which couples normalized policy entropy with the sign of the response advantage to reinforce high-entropy decisions in successful responses, penalize low-entropy decisions in failed responses, and attenuate penalties at uncertain positions, redistributing the response advantage to token level without auxiliary models, extra sampling, or privileged information, and achieving the best overall performance on mathematical, logical, and algorithmic reasoning tasks.
MBZUAI team evaluates GPT-6 Astra and five other general systems across 34 capabilities and 55 benchmarks, finding semantic understanding near reference levels while precise geometry and faithful reconstruction remain gaps
Researchers at MBZUAI and Apertix placed GPT-6 Astra alongside Fable 5, Kimi K3, Gemini 3.1 Pro, Qwen 3.8-Max, and Muse Spark 1.3 under identical instructions, inputs, and scoring across 34 capabilities, 55 benchmarks, and nine areas of computer vision, comparing them with dedicated models and humans, and found that semantic interpretation, reasoning, and object-centric prediction approach or reach reference levels while metric geometric accuracy, faithful reconstruction, temporally consistent dense prediction, and specialized fine-grained visual knowledge retain larger gaps.
SentZero pairs abstract-level sentence mapping with patch-level false-negative alignment to beat prior multi-task zero-shot methods on chest X-ray classification and grounding after MIMIC-CXR pretraining
SentZero is a sentence-centric vision-language pretraining framework for chest X-ray: it uses an LLM to extract phrases from radiology reports and map them to concise topic-presence sentences that expand positive-pair diversity, adds an auxiliary loss that attracts the highest-similarity image patches toward clinical sentences recurring across studies to mitigate false negatives, and applies sentence-conditioned residual modulation to visual embeddings; after pretraining on MIMIC-CXR it outperforms prior multi-task zero-shot methods on most metrics for zero-shot multi-label classification on Open-I, ChestXray14, PadChest, ChestXDet10 and CheXpert and for zero-shot grounding on ChestXDet10 and MS-CXR.
SMAT decomposes merging into Scale, Mask and Perturb simulations, lifting the five-merger average by 1.07–2.16 points across four backbones with under 2% training overhead
The work introduces SMAT (Simple MAT), which from a single expert's perspective describes common model-merging methods through three operations—Scale (reweighting its own update), Mask (removing selected coordinates) and Perturb (adding other experts' updates)—and during training samples scaling coefficients, masks and additive noise to simulate merged parameters while jointly optimizing the expert loss and the expected loss at simulated merged parameters, made efficient by periodic scheduling, kernel fusion and parameter storage switching with one forward and one backward pass per step; across four backbones (Llama-3.2-1B-Instruct, Llama-3.1-8B-Instruct, CLIP ViT-B/32 and ViT-L/14), SMAT improves the mean score over five merging methods by 1.07–2.
PCD pins the primary gradient into the update direction and holds 70.9% accuracy at 90% sparsity on ResNet-34/CIFAR-100 while symmetric multi-objective baselines collapse
The work introduces Priority-Constrained Descent (PCD), a gradient-based framework that anchors on the primary gradient and applies only the minimal Euclidean distortion needed to guarantee each secondary objective a τ-fraction of first-order progress, reporting better per-objective performance and Pareto dominance over symmetric multi-objective methods in structured pruning, unstructured sparsity with low-rankness, and synthetic experiments, with closed-form solutions for two and three objectives and scale invariance.
δ-Vision rebuilds layer-wise visual memory with lightweight MLP adapters, cutting FLOPs to 17.90% while keeping every visual token
Using low-rank interventions, the work finds that visual-to-text information flow in multimodal LLMs is concentrated in a low-dimensional subspace and that layer-specific visual states are highly predictable by lightweight MLPs; it therefore proposes δ-Vision, which freezes the language-model backbone and uses low-rank adapters to construct each layer's visual memory while preserving all visual tokens for text retrieval, reaching an average score of 74.4 on Qwen3-VL-4B (92.9% of uncompressed performance), 9.1 points above the strongest 5%-retention pruning baseline DivPrune, and requiring only 17.90% of the uncompressed model's FLOPs on Video-MME.
RoboFoundry treats the agent system itself as the policy to evolve, lifting GPT-5.5 by 27.8% on EmbodiedBench and transferring zero-shot to real robots
RoboFoundry proposes a Self-Evolving System-as-Policy framework that treats the entire supporting system around a foundation model—its context system and hierarchical skill system—rather than model parameters as the policy to evolve, diagnosing capability gaps from execution traces, converting them into validated task-specific system updates, and promoting recurring improvements into general system capabilities; it achieves state-of-the-art results on EmbodiedBench (improving GPT-5.5 by 27.8% and bringing Qwen3.7-Plus to 70.3% versus GPT-5.5's 72.7%), outperforms all baselines on RoboMemArena long-horizon memory by at least 39.0%, beats Cap-Agent0 on LIBERO-PRO by 243.8%–679.7%, and demonstrates zero-shot transfer and online evolution on real robots.
Page 37 · showing 10