Skip to main content

AI Core

1006 items

  1. arXiv

    WaveFront Decoding fuses drafting and verification into the same recurrent call, speeding up decoding 2.42x on Ouro-2.6B and 3.54x on Huginn-3.5B

    The work introduces WaveFront Decoding (WFD), a training-free self-speculative decoding framework for looped language models that exploits two properties of these architectures, namely that intermediate recurrence outputs provide effective draft predictions and that weight sharing lets token states at different positions and recurrence depths be processed in one batched recurrent-block call, organizing these mixed-depth states into a diagonal wavefront so that drafting and verification proceed concurrently within the same recurrent calls while rejected drafts are corrected using full-depth predictions; across six Spec-Bench task categories it achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, rising to 4.81x on Huginn-3.
  2. arXiv

    LT-OPD uses on-policy self-distillation on its own generated trajectories to lift average retained performance at 5% visual tokens from 68.6% to 82.3%

    The work proposes LT-OPD, which supervises a student keeping only 5% of visual tokens on its own generated trajectories using a frozen full-token copy of the same model as teacher, together with a curriculum that progressively lowers the token budget, raising average retained performance on nine benchmarks with Qwen3.5-4B from 68.6% to 82.3% while cutting KV cache by 85.2% and prefill FLOPs by 85.4% with no added inference overhead.
  3. arXiv

    AdaGuard uses an adaptive guard model to review agent trajectories against user-defined policies, with its 4B model reaching 89.30% binary accuracy on AdaptiveSafety

    The work introduces AdaptiveSafety, a dataset of 10,939 training and 1,000 test examples covering policies with 1 to 100 rules, together with SafePO, a reinforcement learning algorithm, and uses them to train the AdaGuard family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time; the 4B model achieves binary accuracies of 89.30% on AdaptiveSafety and 71.82% on DynaBench.
  4. arXiv

    Treating EEG masking geometry as the only variable: 58 pre-trained models point to a moderate spatial radius with short temporal blocks, and expose a JEPA-specific bias-inflation collapse

    The study unifies EEG self-supervised masking strategies into a three-parameter framework of spatial radius, temporal length and mask ratio, trains 58 models under a fixed REVE-Small backbone and corpus across MAE and JEPA, evaluates them with a linear probe on the 12 OpenEEGBench datasets, and finds that both frameworks agree on a moderate spatial radius with short temporal blocks as the optimum while identifying a JEPA-specific bias-inflation collapse at full-channel masking.
  5. arXiv

    LSPD recasts on-policy distillation as KL-regularized RL: +1.59 Avg@16 across six math reasoning benchmarks, with a replay variant matching vanilla OPD in about a quarter of rollout batches

    The work reinterprets the reverse-KL objective of on-policy distillation (OPD) as KL-regularized policy optimization and introduces Least-Square Policy Distillation (LSPD), which replaces policy-gradient updates with robust quadratic log-probability matching plus entropy regularization, enabling multiple updates per rollout and reuse of historical trajectories; across six mathematical reasoning benchmarks and three teacher-student settings it improves Avg@16 by +1.59 and Pass@16 by +1.87 on average, and its fully off-policy replay variant LSPD-RB reaches performance comparable to vanilla OPD using roughly one quarter of the rollout batches.
  6. arXiv

    DISCO splits long context into parallel Worker grounding plus Driver reasoning, holding 78.4% on 1M-token RULER-QA while cutting inference cost by over 80%

    The work proposes DISCO, which separates long-context processing into local evidence grounding executed in parallel by lightweight Worker LLMs and global reasoning handled by a central Driver LLM whose dynamic DAG planning is optimized with GRPO reinforcement learning; with Qwen3-8B it holds 78.4% accuracy on RULER-QA at 1M tokens (where the RAG baseline collapses to 10.9%), improves up to 9.8 points over full-context baselines on LongBench v2, and matches Gemini-3-Pro-Preview while reducing inference cost by over 80%.
  7. arXiv

    CoWA splits causal-history coverage across KV heads, cutting 128K training forward and backward latency to roughly one-seventh of FullAttn

    The work introduces CoWindow Attention (CoWA), in which all KV heads share near-diagonal and prefix-sink windows while complementary long-range windows partition the remaining history across heads, making full causal coverage a collective property of the head ensemble; in a window-matched 8K ablation CoWA reaches 89.73% versus 89.97% for FullAttn while duplicated long-range windows perform substantially worse, a 128K tensor-parallel operator benchmark reduces training forward and backward latency by 7.4x and 8.6x and decoding latency by 3.0x, and scaling-law training from 0.6B to 14B tracks FullAttn perplexity while cutting total training FLOPs by 28.5% in the 32K long-context stage.
  8. arXiv

    Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction

    This monograph develops a unified account of training and inference in Power Law Decoder Representation language models (PLDR-LLMs): exact finite work identities decompose changes in the absolute energy of the row-centered learned map into parameter contributions, signed interactions, and numerical observation defects; positive affine blocking retains restarts at the row-constant face while the augmented AdamW state supplies the complete dynamical description; predictive renormalization acts on the complete conditional training law for a single pass over distinct corpus target blocks; experiments reveal observer and optimizer dependence, reject the tested autonomous row-state candidates, and support finite conditional prediction and state-specific operator reduction, while independent sing
  9. arXiv

    WorldPlay2 pairs a factorized hybrid control interface with compressed memory and stable distillation, reaching 83.1 average on WBench and 0.105 MEt3R on RevisitBench

    WorldPlay2 is a real-time interactive world model that combines frame-aligned action control with structured semantic control disentangling scene appearance, character identity, and dynamic semantic events into a factorized hybrid control interface, compresses historical context into memory tokens shared by the autoregressive student and the bidirectional teacher, and uses Stable Forcing distillation built from few-step initialization and full-rollout replay, achieving 83.1 average on WBench (above Alaya-Evoke-Turbo's 82.0), 19.71 PSNR and 0.105 MEt3R on RevisitBench, and 16 FPS real-time generation on 8 H20 GPUs.
  10. arXiv

    Imprint Reader uses SMaRT to read frozen weight updates into natural language, and MetaEdit turns those descriptions into behavioral intervention

    The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.

Page 38 · showing 10