Skip to main content

Search

“All disciplines” · 1449 results

Page 23 · showing 20
arXiv

A strong agent writes its intervention experience into a playbook for a light agent, lifting real-world manipulation success from 37.3% to 64.0%

The work proposes Recursive Harness Distillation (RHD): a strong agent (GPT-6 Astra) first interacts with a frozen vision-language-action model and distills effective intervention strategies into a reusable playbook, then recursively revises that playbook using the light agent's (GPT-5.6 Luna) execution feedback, improving manipulation without updating any model parameters—real-world success rises from 37.3% to 64.0%, the light agent with the playbook reaches 66.7% on SimplerEnv Bridge versus 41.7% for GR00T alone, and the same playbook also lifts the strong agent to 79.2%.
arXiv

Tencent Hunyuan and collaborators compared 11 MoE rungs and found encoder-free multimodal loss falls faster with compute, crossing over near 10^22 FLOPs

Using a matched ladder of 11 sparse MoE decoders (1.1B–44B total, 71M–2.4B active non-embedding parameters) with identical data mixture and optimization, the work fits separate text and multimodal scaling laws for encoder-free and encoder-based MLLMs and finds that removing the visual encoder shifts compute-optimal multimodal allocation toward larger models (model allocation exponent rising from about 0.464 to about 0.570) while leaving text nearly unchanged; encoder-free models underperform at small scales but their multimodal loss falls faster with compute, with extrapolation predicting an efficiency crossover on the order of 10^22 FLOPs (point estimate about 3.
arXiv

G²PTQ refreshes gradients and Hessians before quantizing each Transformer block, halving KL divergence versus GPTAQ and adding 6.45% QA accuracy at 2-bit

The work introduces G²PTQ, a post-training quantization framework that combines first- and second-order information under a globally supervised, block-wise objective: it recomputes gradient and Hessian estimates before quantizing each Transformer block to avoid staleness and uses trust-region scaling to bound the exact first-order compensation step, achieving lower KL divergence and higher downstream QA accuracy than baselines such as GPTQ, GuidedQuant, and GPTAQ across 13 dense and 2 MoE models from 0.6B to 125B parameters at 2/3/4-bit weight and 4-bit weight-activation settings.
arXiv

CoWA splits causal-history coverage across KV heads, cutting 128K training forward and backward latency to roughly one-seventh of FullAttn

The work introduces CoWindow Attention (CoWA), in which all KV heads share near-diagonal and prefix-sink windows while complementary long-range windows partition the remaining history across heads, making full causal coverage a collective property of the head ensemble; in a window-matched 8K ablation CoWA reaches 89.73% versus 89.97% for FullAttn while duplicated long-range windows perform substantially worse, a 128K tensor-parallel operator benchmark reduces training forward and backward latency by 7.4x and 8.6x and decoding latency by 3.0x, and scaling-law training from 0.6B to 14B tracks FullAttn perplexity while cutting total training FLOPs by 28.5% in the 32K long-context stage.
arXiv

Auditing 1,813 AmE–BrE variant pairs across the pipeline: six pretraining corpora, 21 post-training datasets and ten checkpoints all lean American English

The study builds a curated resource of 1,813 matched American English–British English variant pairs and introduces DiAlign, a training-free method for estimating regional alignment, then uses them to audit six pretraining corpora, 21 post-training datasets, nine tokenizers, ten model checkpoints and generations under two prompting conditions, finding that American English is systematically favored across data exposure, representation and generation, and that British-English prompting only partially shifts this default.
arXiv

PReCache shares KV caches across multi-LoRA agents via low-rank precomputation and neutral reconstruction, cutting TTFT by up to several times with almost no accuracy loss

The work presents PReCache, a training-free KV cache sharing framework with two designs: PreLRShared precomputes a compact agent-specific low-rank cache for every agent when each segment of the shared context is first processed, removing later agents' repeated prefill over the accumulated trajectory, while ReBaseShared further reconstructs the shared base cache from adapter-free hidden states to reduce dependence on the previous agent, with lazy prefill for single-stream inference and double batching for concurrent serving; across LLaMA-3.
arXiv

UMM-Reflection lifts BAGEL's GenEval from 0.72 to 0.84 and its repair rate from 20.59% to 64.94% by training on whole reflection trajectories

The work introduces UMM-Reflection, which cold-starts a unified multimodal model (BAGEL) with supervised fine-tuning on reflection trajectories and then applies reinforcement learning with a single trajectory-level advantage that updates both the reflection text and the flow-based image revisions, raising GenEval by 12.05 points over SFT and the conditional repair rate from 20.59% to 64.94%, with gains transferring to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none used in training.
arXiv

FlyBy lets a 4B small reasoning model query stronger models at knowledge bottlenecks, beating Qwen3-14B on 1,158 hard problems at 2.7x lower serving cost

Through counterfactual interventions at intermediate reasoning states across two model families and multiple scales, the work finds that self-refinement largely consolidates probability mass onto solutions already reachable from the current state rather than making new ones reachable, distinguishes execution bottlenecks from knowledge bottlenecks, and introduces FlyBy, a selective querying framework that trains 4B and 8B models to reason first, diagnose what remains unresolved, and query stronger models at knowledge bottlenecks; on 1,158 hard problems across six benchmarks FlyBy-4B reaches 45.96% pass@8, surpassing Qwen3-14B's 41.64% at 2.7x lower serving cost, and FlyBy-8B raises pass@8 to 51.81%.
arXiv

InfiniHand estimates hands and camera trajectory together from egocentric video with one streaming feed-forward network, cutting ARCTIC PA-p to 7.72 mm at 11.19 FPS

InfiniHand is an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video, trained in two progressive stages on roughly 5,000 hours of aggregated egocentric data; it reduces ARCTIC PA-p by 21.4% versus ViDiHand (9.82 to 7.72 mm), lowers EgoDex W-MPJPE from 78.41 to 28.77 mm versus Dyn-HaMR, and runs at 11.19 FPS, more than twice the throughput of HaWoR.
arXiv

REALM adds a delay prior and stochastic expression refinement to turn listening from a frozen stare into blinks and smiles on a robot

REALM is a coarse-to-fine framework for audio-driven reactive listening that fuses listener motion history with speaker audio through a Reactive Gated Speaker–Listener Fusion module using a delay-centered attention prior and adaptive gating, then predicts a base motion trajectory with a coarse decoder and augments it with audio-conditioned stochastic residuals in the expression subspace, improving over the evaluated baselines on multiple motion-quality metrics on ViCo and L2L and deploying to an Ameca humanoid robot with a perceptual user study.
arXiv

WhiteMatter lets every Transformer layer read past-token representations from any depth, beating matched baselines at half KV cache

The work introduces WhiteMatter, which uses a learned router to mix past-token representations from any depth into shared key-value (KV) channels so every layer can draw on cross-layer information, together with cyclic Gauss–Seidel iteration to speed up training and prefill; under matched training-token budgets, half-cache WhiteMatter improves perplexity and average downstream scores over matched standard Transformers at two scales up to about 1.3B parameters, while the full-cache version performs comparably to a deeper standard Transformer.
arXiv

SentZero pairs abstract-level sentence mapping with patch-level false-negative alignment to beat prior multi-task zero-shot methods on chest X-ray classification and grounding after MIMIC-CXR pretraining

SentZero is a sentence-centric vision-language pretraining framework for chest X-ray: it uses an LLM to extract phrases from radiology reports and map them to concise topic-presence sentences that expand positive-pair diversity, adds an auxiliary loss that attracts the highest-similarity image patches toward clinical sentences recurring across studies to mitigate false negatives, and applies sentence-conditioned residual modulation to visual embeddings; after pretraining on MIMIC-CXR it outperforms prior multi-task zero-shot methods on most metrics for zero-shot multi-label classification on Open-I, ChestXray14, PadChest, ChestXDet10 and CheXpert and for zero-shot grounding on ChestXDet10 and MS-CXR.
arXiv

Rewriting public documents into about 10k synthetic samples lifts a 35B student from 13.7% to 24.6% on CL-bench, near a trillion-parameter frontier model

The work proposes a human-annotation-free synthetic supervision pipeline: it deeply rewrites 3,515 public documents via entity renaming and numeric perturbation, has LLMs generate questions and rubrics that require reasoning over the document, and admits only samples that genuinely depend on the document via a with-document versus without-document rubric gap check, yielding about 9,625 training samples; SFT raises Qwen3.6-35B-A3B on CL-bench from 13.7% to 22.8%, a subsequent rubric-reward RL stage reaches 24.6%, comparable to Qwen3.8-2.4T at 23.9%, with broad transfer to long-context understanding, instruction following, and reasoning, while code generation and knowledge stay mostly flat.
arXiv

VGGT-Diff routes VGGT-Ω geometry latents into Wan video diffusion, topping PSNR, LPIPS and DreamSim on 6,188 DL3DV targets

VGGT-Diff is a geometry-routed multi-view diffusion framework that uses a frozen VGGT-Ω to transform source-view features together with 3D points and confidence into query-aligned conditions via a confidence-aware Visual Geometry Router, injects them into a pretrained Wan2.1-I2V-14B video diffusion model, and regularizes predicted-clean residuals along reliable 3D tracks with Point-Track Residual Consistency; trained on only 980 DL3DV scenes, it ranks first in PSNR, LPIPS and DreamSim and second in SSIM on 6,188 DL3DV-Benchmark targets, and first in LPIPS and DreamSim with second-place PSNR on zero-shot Mip-NeRF 360.
arXiv

QwenGyre pairs elastic GPU scheduling with trajectory-tree processing to lift Qwen 3.8 2.4T from 52.5% to 58.5% on NL2RepoBench in 48 steps

QwenGyre is an end-to-end framework for extreme-long-horizon (xLong) online agent RL that elastically reallocates GPUs between rollout and training without interrupting live harness executions, while its trajectory processor reconstructs branching executions into shared-prefix trajectory trees, scores partial progress after timeouts, and deduplicates sampled trajectories by role priority; scaled to Qwen 3.8 2.4T with 700K tokens per rollout, it raises NL2RepoBench pass rate from 52.5% to 58.5% in 48 steps and delivers up to 1.85x and 1.78x end-to-end speedups over Colocate and Async.
Google AI 与 Gemini 产品博客

Four builders used Gemini 3.8 Flash to map rocket orbits, animate an ink painting, build a T. rex skeleton, and simulate an automatic transmission

Google describes its workhorse model Gemini 3.8 Flash as delivering significant improvements over 3.7 Flash in software engineering, agentic tasks, and multistep reasoning in specialized domains, and showcases four builder projects: live path mapping of satellites, space stations, and orbital rockets; turning Seigaiha waves into a moving ink painting; a T. rex skeleton built with a four-phase prompt and accuracy checks; and an interactive automatic transmission model with 10 camera views and four display modes.
arXiv

8B multimodal verifier SciGen-Verifier tops Qwen3.8-Max with 86.28% judgement accuracy on a scientific image verification benchmark and serves as an online critic for iterative correction

The work builds SciGen-Verify, a benchmark for explainable verification of scientific image generation (1,350 samples: 406 instruction following, 287 multidisciplinary reasoning, 657 world knowledge, with a three-tier cascading protocol over binary judgement, explanation, and editing instruction), and develops SciGen-Verifier on Qwen3-VL-8B-Instruct trained by cold-start supervised fine-tuning (48,311 instances) followed by a curriculum-based two-stage reinforcement learning pipeline (rubric-guided process reward, then outcome reward), reaching 86.28% overall judgement accuracy, 81.79% explanation accuracy, and 57.42% editing-instruction accuracy on the benchmark while also serving as an online critic for iterative image rectification.
arXiv

FactorEngram factorizes n-gram memory over a shared dictionary, lifting 8K long-context retrieval by up to 43.3 points at 340M and 1B backbones

The work proposes FactorEngram, which replaces monolithic per-pattern n-gram embeddings with sparsity-regularized coefficients over a dictionary shared across patterns and reuses that same dictionary for basis-level contextual gating, improving language modeling, downstream tasks, and long-context retrieval on 340M- and 1B-parameter Transformer backbones.
arXiv

Tsinghua LeapLab replaces KL divergence with a binary directional reward, matching or beating on-policy distillation on math and code, and proposes C-MOPD for multi-teacher training

This work revisits whether reverse KL divergence is necessary for on-policy distillation (OPD), proposes BinaryOPD, which keeps only the update direction as a binary reward and matches or slightly exceeds OPD across multiple student-teacher pairs on math and code, further finds that only the direction of a small subset of high-disagreement tokens is critical, and builds on this to propose C-MOPD, in which every sample is supervised by all teachers and which consistently outperforms MOPD on math and code benchmarks.
arXiv

WaveFront Decoding fuses drafting and verification into the same recurrent call, speeding up decoding 2.42x on Ouro-2.6B and 3.54x on Huginn-3.5B

The work introduces WaveFront Decoding (WFD), a training-free self-speculative decoding framework for looped language models that exploits two properties of these architectures, namely that intermediate recurrence outputs provide effective draft predictions and that weight sharing lets token states at different positions and recurrence depths be processed in one batched recurrent-block call, organizing these mixed-depth states into a diagonal wavefront so that drafting and verification proceed concurrently within the same recurrent calls while rejected drafts are corrected using full-depth predictions; across six Spec-Bench task categories it achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, rising to 4.81x on Huginn-3.