Skip to main content

AI Core

1008 items

  1. arXiv

    WorldPlay2 pairs a factorized hybrid control interface with compressed memory and stable distillation, reaching 83.1 average on WBench and 0.105 MEt3R on RevisitBench

    WorldPlay2 is a real-time interactive world model that combines frame-aligned action control with structured semantic control disentangling scene appearance, character identity, and dynamic semantic events into a factorized hybrid control interface, compresses historical context into memory tokens shared by the autoregressive student and the bidirectional teacher, and uses Stable Forcing distillation built from few-step initialization and full-rollout replay, achieving 83.1 average on WBench (above Alaya-Evoke-Turbo's 82.0), 19.71 PSNR and 0.105 MEt3R on RevisitBench, and 16 FPS real-time generation on 8 H20 GPUs.
  2. arXiv

    Imprint Reader uses SMaRT to read frozen weight updates into natural language, and MetaEdit turns those descriptions into behavioral intervention

    The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
  3. arXiv

    Across nearly half a million names and 12 tokenizers, whether a name becomes a single token shifts concept accessibility in fellowship, hiring, clinical, and lending judgments

    The study introduces NameTrace, measures direct lexical support for names across nearly half a million first names and 12 LLM-associated tokenizers, and, on atomic versus short-fragmented names matched within the same race/ethnicity–gender strata, finds systematically higher task-aligned concept accessibility for atomic names across fellowship, hiring, clinical assessment, and lending; the differences persist in all eight strata, transfer to unseen names, and hidden-state interventions along the measured task directions shift later constrained choices.
  4. arXiv

    FlashForward reuses in-flight KV cache with sparse clean anchors to speed 20s+ video generation by 1.16–1.69x on four GPUs while reaching 0.838 VBench Total

    FlashForward publishes the in-flight KV cache already computed by each ordinary denoising forward for reuse by later chunks and complements it with sparse clean anchor KV generated ahead of time for long-range structural guidance, removing cache-update-only forwards; with up to four GPUs it runs 1.16–1.69x faster than HiAR and 1.42–2.92x faster than Self-Forcing for 16 FPS videos of 20 seconds or longer across 1.3B and 14B backbones at 480p and 720p, and the 1.3B model at 480p scores 0.838 on VBench-1.0 while remaining stable at 20s, 35s, and 65s.
  5. arXiv

    WhiteMatter lets every Transformer layer read past-token representations from any depth, beating matched baselines at half KV cache

    The work introduces WhiteMatter, which uses a learned router to mix past-token representations from any depth into shared key-value (KV) channels so every layer can draw on cross-layer information, together with cyclic Gauss–Seidel iteration to speed up training and prefill; under matched training-token budgets, half-cache WhiteMatter improves perplexity and average downstream scores over matched standard Transformers at two scales up to about 1.3B parameters, while the full-cache version performs comparably to a deeper standard Transformer.
  6. arXiv

    FlyBy lets a 4B small reasoning model query stronger models at knowledge bottlenecks, beating Qwen3-14B on 1,158 hard problems at 2.7x lower serving cost

    Through counterfactual interventions at intermediate reasoning states across two model families and multiple scales, the work finds that self-refinement largely consolidates probability mass onto solutions already reachable from the current state rather than making new ones reachable, distinguishes execution bottlenecks from knowledge bottlenecks, and introduces FlyBy, a selective querying framework that trains 4B and 8B models to reason first, diagnose what remains unresolved, and query stronger models at knowledge bottlenecks; on 1,158 hard problems across six benchmarks FlyBy-4B reaches 45.96% pass@8, surpassing Qwen3-14B's 41.64% at 2.7x lower serving cost, and FlyBy-8B raises pass@8 to 51.81%.
  7. arXiv

    Peking University-led team distills coding ability from one-word answers: a Qwen2.5-1.5B student beats an exact nuisance-matched control by 5.34 points on HumanEval+

    The work introduces Active Taskless Distillation (ATD), which uses a public ancestor model to select unrelated prompts on which it is nearly indifferent between two ordinary words, has a privately post-trained teacher return a single word per prompt, and trains a same-ancestor student only on those prompt-word pairs; in the primary Qwen2.5-1.5B coding experiment, 5,664 single-word responses yield a 5.34 percentage-point gain on HumanEval+ over an exact nuisance-matched control, with positive mean gains also observed in scientific knowledge, commonsense reasoning, and reading comprehension across additional model generations, sizes, and families.
  8. arXiv

    Adobe Research and KAIST introduce FlowTool: treating retouching tool parameters as conditional generation, cutting L1 by up to 28.8% and L2 by up to 47.4% on reference metrics with at least 50x lower latency

    Researchers at Adobe Research and KAIST introduce FlowTool, which uses conditional rectified flow to directly model the distribution of high-quality retouching tool parameters conditioned on the input image and user instruction, combining a vision-language model backbone with a Diffusion Transformer parameter generator and a tool-presence head instead of autoregressive MLLM reasoning and token-by-token numeric generation; across MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, and MIT-Adobe5K it achieves stronger reference-based performance than specialized MLLM editing agents and proprietary MLLMs, remains competitive with proprietary models under reference-free evaluation, and reduces inference latency by at least 50x while requiring nearly 2x less memory.
  9. arXiv

    InfiniHand estimates hands and camera trajectory together from egocentric video with one streaming feed-forward network, cutting ARCTIC PA-p to 7.72 mm at 11.19 FPS

    InfiniHand is an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video, trained in two progressive stages on roughly 5,000 hours of aggregated egocentric data; it reduces ARCTIC PA-p by 21.4% versus ViDiHand (9.82 to 7.72 mm), lowers EgoDex W-MPJPE from 78.41 to 28.77 mm versus Dyn-HaMR, and runs at 11.19 FPS, more than twice the throughput of HaWoR.
  10. arXiv

    Intervention experiments across DeepSeekMoE, OLMoE and Qwen3-MoE show post-merge expert routing drift is mostly input-representation-induced yet fails to predict task gains from source-route restoration

    The work introduces a routing analysis toolkit and, across DeepSeekMoE, OLMoE and Qwen3-MoE, crosses source and merged router inputs and parameters and runs paired token- and task-level evaluations, finding that most expert reassignments are input-shift-induced rather than parameter-induced, that source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and that different expert selections can produce directionally similar mixture outputs; it therefore redefines routing failure as task loss recoverable under a specified routing intervention with non-routing parameters fixed, and proposes Selective Router Repair as a case study that finds no reliable evidence that source-specialist token-likelihood advantages identify beneficial local corr

Page 39 · showing 10