AI Core
1008 items
Imprint Reader uses SMaRT to read frozen weight updates into natural language, and MetaEdit turns those descriptions into behavioral intervention
The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
Across nearly half a million names and 12 tokenizers, whether a name becomes a single token shifts concept accessibility in fellowship, hiring, clinical, and lending judgments
The study introduces NameTrace, measures direct lexical support for names across nearly half a million first names and 12 LLM-associated tokenizers, and, on atomic versus short-fragmented names matched within the same race/ethnicity–gender strata, finds systematically higher task-aligned concept accessibility for atomic names across fellowship, hiring, clinical assessment, and lending; the differences persist in all eight strata, transfer to unseen names, and hidden-state interventions along the measured task directions shift later constrained choices.
FlashForward reuses in-flight KV cache with sparse clean anchors to speed 20s+ video generation by 1.16–1.69x on four GPUs while reaching 0.838 VBench Total
FlashForward publishes the in-flight KV cache already computed by each ordinary denoising forward for reuse by later chunks and complements it with sparse clean anchor KV generated ahead of time for long-range structural guidance, removing cache-update-only forwards; with up to four GPUs it runs 1.16–1.69x faster than HiAR and 1.42–2.92x faster than Self-Forcing for 16 FPS videos of 20 seconds or longer across 1.3B and 14B backbones at 480p and 720p, and the 1.3B model at 480p scores 0.838 on VBench-1.0 while remaining stable at 20s, 35s, and 65s.
WhiteMatter lets every Transformer layer read past-token representations from any depth, beating matched baselines at half KV cache
The work introduces WhiteMatter, which uses a learned router to mix past-token representations from any depth into shared key-value (KV) channels so every layer can draw on cross-layer information, together with cyclic Gauss–Seidel iteration to speed up training and prefill; under matched training-token budgets, half-cache WhiteMatter improves perplexity and average downstream scores over matched standard Transformers at two scales up to about 1.3B parameters, while the full-cache version performs comparably to a deeper standard Transformer.
FlyBy lets a 4B small reasoning model query stronger models at knowledge bottlenecks, beating Qwen3-14B on 1,158 hard problems at 2.7x lower serving cost
Through counterfactual interventions at intermediate reasoning states across two model families and multiple scales, the work finds that self-refinement largely consolidates probability mass onto solutions already reachable from the current state rather than making new ones reachable, distinguishes execution bottlenecks from knowledge bottlenecks, and introduces FlyBy, a selective querying framework that trains 4B and 8B models to reason first, diagnose what remains unresolved, and query stronger models at knowledge bottlenecks; on 1,158 hard problems across six benchmarks FlyBy-4B reaches 45.96% pass@8, surpassing Qwen3-14B's 41.64% at 2.7x lower serving cost, and FlyBy-8B raises pass@8 to 51.81%.
Peking University-led team distills coding ability from one-word answers: a Qwen2.5-1.5B student beats an exact nuisance-matched control by 5.34 points on HumanEval+
The work introduces Active Taskless Distillation (ATD), which uses a public ancestor model to select unrelated prompts on which it is nearly indifferent between two ordinary words, has a privately post-trained teacher return a single word per prompt, and trains a same-ancestor student only on those prompt-word pairs; in the primary Qwen2.5-1.5B coding experiment, 5,664 single-word responses yield a 5.34 percentage-point gain on HumanEval+ over an exact nuisance-matched control, with positive mean gains also observed in scientific knowledge, commonsense reasoning, and reading comprehension across additional model generations, sizes, and families.
Adobe Research and KAIST introduce FlowTool: treating retouching tool parameters as conditional generation, cutting L1 by up to 28.8% and L2 by up to 47.4% on reference metrics with at least 50x lower latency
Researchers at Adobe Research and KAIST introduce FlowTool, which uses conditional rectified flow to directly model the distribution of high-quality retouching tool parameters conditioned on the input image and user instruction, combining a vision-language model backbone with a Diffusion Transformer parameter generator and a tool-presence head instead of autoregressive MLLM reasoning and token-by-token numeric generation; across MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, and MIT-Adobe5K it achieves stronger reference-based performance than specialized MLLM editing agents and proprietary MLLMs, remains competitive with proprietary models under reference-free evaluation, and reduces inference latency by at least 50x while requiring nearly 2x less memory.
InfiniHand estimates hands and camera trajectory together from egocentric video with one streaming feed-forward network, cutting ARCTIC PA-p to 7.72 mm at 11.19 FPS
InfiniHand is an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video, trained in two progressive stages on roughly 5,000 hours of aggregated egocentric data; it reduces ARCTIC PA-p by 21.4% versus ViDiHand (9.82 to 7.72 mm), lowers EgoDex W-MPJPE from 78.41 to 28.77 mm versus Dyn-HaMR, and runs at 11.19 FPS, more than twice the throughput of HaWoR.
Intervention experiments across DeepSeekMoE, OLMoE and Qwen3-MoE show post-merge expert routing drift is mostly input-representation-induced yet fails to predict task gains from source-route restoration
The work introduces a routing analysis toolkit and, across DeepSeekMoE, OLMoE and Qwen3-MoE, crosses source and merged router inputs and parameters and runs paired token- and task-level evaluations, finding that most expert reassignments are input-shift-induced rather than parameter-induced, that source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and that different expert selections can produce directionally similar mixture outputs; it therefore redefines routing failure as task loss recoverable under a specified routing intervention with non-routing parameters fixed, and proposes Selective Router Repair as a case study that finds no reliable evidence that source-specialist token-likelihood advantages identify beneficial local corr
Page 39 · showing 10