AI Core
1006 items
LT-OPD uses on-policy self-distillation on its own generated trajectories to lift average retained performance at 5% visual tokens from 68.6% to 82.3%
The work proposes LT-OPD, which supervises a student keeping only 5% of visual tokens on its own generated trajectories using a frozen full-token copy of the same model as teacher, together with a curriculum that progressively lowers the token budget, raising average retained performance on nine benchmarks with Qwen3.5-4B from 68.6% to 82.3% while cutting KV cache by 85.2% and prefill FLOPs by 85.4% with no added inference overhead.
AdaGuard uses an adaptive guard model to review agent trajectories against user-defined policies, with its 4B model reaching 89.30% binary accuracy on AdaptiveSafety
The work introduces AdaptiveSafety, a dataset of 10,939 training and 1,000 test examples covering policies with 1 to 100 rules, together with SafePO, a reinforcement learning algorithm, and uses them to train the AdaGuard family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time; the 4B model achieves binary accuracies of 89.30% on AdaptiveSafety and 71.82% on DynaBench.
Treating EEG masking geometry as the only variable: 58 pre-trained models point to a moderate spatial radius with short temporal blocks, and expose a JEPA-specific bias-inflation collapse
The study unifies EEG self-supervised masking strategies into a three-parameter framework of spatial radius, temporal length and mask ratio, trains 58 models under a fixed REVE-Small backbone and corpus across MAE and JEPA, evaluates them with a linear probe on the 12 OpenEEGBench datasets, and finds that both frameworks agree on a moderate spatial radius with short temporal blocks as the optimum while identifying a JEPA-specific bias-inflation collapse at full-channel masking.
LSPD recasts on-policy distillation as KL-regularized RL: +1.59 Avg@16 across six math reasoning benchmarks, with a replay variant matching vanilla OPD in about a quarter of rollout batches
The work reinterprets the reverse-KL objective of on-policy distillation (OPD) as KL-regularized policy optimization and introduces Least-Square Policy Distillation (LSPD), which replaces policy-gradient updates with robust quadratic log-probability matching plus entropy regularization, enabling multiple updates per rollout and reuse of historical trajectories; across six mathematical reasoning benchmarks and three teacher-student settings it improves Avg@16 by +1.59 and Pass@16 by +1.87 on average, and its fully off-policy replay variant LSPD-RB reaches performance comparable to vanilla OPD using roughly one quarter of the rollout batches.
DISCO splits long context into parallel Worker grounding plus Driver reasoning, holding 78.4% on 1M-token RULER-QA while cutting inference cost by over 80%
The work proposes DISCO, which separates long-context processing into local evidence grounding executed in parallel by lightweight Worker LLMs and global reasoning handled by a central Driver LLM whose dynamic DAG planning is optimized with GRPO reinforcement learning; with Qwen3-8B it holds 78.4% accuracy on RULER-QA at 1M tokens (where the RAG baseline collapses to 10.9%), improves up to 9.8 points over full-context baselines on LongBench v2, and matches Gemini-3-Pro-Preview while reducing inference cost by over 80%.
CoWA splits causal-history coverage across KV heads, cutting 128K training forward and backward latency to roughly one-seventh of FullAttn
The work introduces CoWindow Attention (CoWA), in which all KV heads share near-diagonal and prefix-sink windows while complementary long-range windows partition the remaining history across heads, making full causal coverage a collective property of the head ensemble; in a window-matched 8K ablation CoWA reaches 89.73% versus 89.97% for FullAttn while duplicated long-range windows perform substantially worse, a 128K tensor-parallel operator benchmark reduces training forward and backward latency by 7.4x and 8.6x and decoding latency by 3.0x, and scaling-law training from 0.6B to 14B tracks FullAttn perplexity while cutting total training FLOPs by 28.5% in the 32K long-context stage.
Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction
This monograph develops a unified account of training and inference in Power Law Decoder Representation language models (PLDR-LLMs): exact finite work identities decompose changes in the absolute energy of the row-centered learned map into parameter contributions, signed interactions, and numerical observation defects; positive affine blocking retains restarts at the row-constant face while the augmented AdamW state supplies the complete dynamical description; predictive renormalization acts on the complete conditional training law for a single pass over distinct corpus target blocks; experiments reveal observer and optimizer dependence, reject the tested autonomous row-state candidates, and support finite conditional prediction and state-specific operator reduction, while independent sing
WorldPlay2 pairs a factorized hybrid control interface with compressed memory and stable distillation, reaching 83.1 average on WBench and 0.105 MEt3R on RevisitBench
WorldPlay2 is a real-time interactive world model that combines frame-aligned action control with structured semantic control disentangling scene appearance, character identity, and dynamic semantic events into a factorized hybrid control interface, compresses historical context into memory tokens shared by the autoregressive student and the bidirectional teacher, and uses Stable Forcing distillation built from few-step initialization and full-rollout replay, achieving 83.1 average on WBench (above Alaya-Evoke-Turbo's 82.0), 19.71 PSNR and 0.105 MEt3R on RevisitBench, and 16 FPS real-time generation on 8 H20 GPUs.
Imprint Reader uses SMaRT to read frozen weight updates into natural language, and MetaEdit turns those descriptions into behavioral intervention
The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
Page 38 · showing 10