Skip to main content

Search

“All disciplines” · 1445 results

Page 22 · showing 20
arXiv

Relic turns recurring collaboration failures into executable protocols, lifting complete-contract delivery from 14.06% to 19.76% across 360 controlled runs

Relic lets multi-agent organizations turn recurring collaboration failures into organization-owned, executable, revisable protocols, raising complete-contract delivery on mainline from 14.06% to 19.76% (+5.71 percentage points) across 360 controlled runs over ten software workloads and three models, and raising behavioral correctness under fresh-member transfer from 25.4% with no inherited protocol and 34.6% with text-only rules to 41.2% with executable bindings.
arXiv

LSPD recasts on-policy distillation as KL-regularized RL: +1.59 Avg@16 across six math reasoning benchmarks, with a replay variant matching vanilla OPD in about a quarter of rollout batches

The work reinterprets the reverse-KL objective of on-policy distillation (OPD) as KL-regularized policy optimization and introduces Least-Square Policy Distillation (LSPD), which replaces policy-gradient updates with robust quadratic log-probability matching plus entropy regularization, enabling multiple updates per rollout and reuse of historical trajectories; across six mathematical reasoning benchmarks and three teacher-student settings it improves Avg@16 by +1.59 and Pass@16 by +1.87 on average, and its fully off-policy replay variant LSPD-RB reaches performance comparable to vanilla OPD using roughly one quarter of the rollout batches.
arXiv

DepthBench sweeps width–depth ratio across 10 architectures at fixed parameter budget: HC and Full AttnRes keep lowering validation loss at extreme deep-narrow shapes while Pre-LN degrades

The work introduces DepthBench, a controlled benchmark that varies the width–depth aspect ratio (d_model/n_layer) from shallow-wide to deep-narrow under a fixed parameter budget and pre-training recipe, comparing 10 representative residual and normalization architectures, and finds that the benefit of depth is strongly architecture-dependent: Pre-LN and most of its norm- and scaling-based variants worsen as models get deeper and narrower (Pre-LN rises from 2.759 to 2.782), whereas HC and Full AttnRes improve consistently (HC from 2.729 to 2.699, Full AttnRes from 2.751 to 2.718), accompanied by stronger cross-layer causal dependence and order sensitivity, though deeper models also incur compute and memory overhead.
arXiv

YuE2 writes a readable score before rendering, and experts prefer symbolic planning 49.3% to 34.6% on overall quality

YuE2 uses a single roughly 3.58B-parameter AR–NAR Mixture-of-Transformers to first write an ABC score specifying melody, chords, key, meter, tempo, and form, then expand it into semantic music tokens and render full-song audio, with companion models MERT2 and SheetSage2 constructing semantic and symbolic supervision from recordings; with the same checkpoint, experts prefer symbolic planning 49.3% to 34.6% for overall quality, YuE2 scores 6.7316 on SongBench Global Avg on WildSongBench (6.9632 with best-of-8), MERT2 surpasses previous best results on 14 of 15 MARBLE metrics, and SheetSage2-AR leads 12 of 15 benchmark–metric pairs.
arXiv

Asked the same prompt twice, eight models returned identical structure node sets in only 28% of cells, and the authors' self-audit withdrew six of seven claims

The authors ran an end-to-end self-audit of an evaluation of LLM-inferred prompt structure: across eight open model variants spanning five families and 8B to 675B parameters, with caching disabled and 293 raw intermediate representations persisted, repeated identical calls yielded mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells were never node-set-perfect; a joint cluster bootstrap showed only the bottom of the ranking is firm (99% and 86% rank retention), the middle four at 27% to 48% and the top two at 68% each; two equally defensible merge rules changed four of eight rows and moved the study-wide headline by 7 percentage points; reproducibility could not be read as accuracy; and four of eight endpoints were withdrawn within ten weeks, so the study as specified can
arXiv

FlashForward reuses in-flight KV cache with sparse clean anchors to speed 20s+ video generation by 1.16–1.69x on four GPUs while reaching 0.838 VBench Total

FlashForward publishes the in-flight KV cache already computed by each ordinary denoising forward for reuse by later chunks and complements it with sparse clean anchor KV generated ahead of time for long-range structural guidance, removing cache-update-only forwards; with up to four GPUs it runs 1.16–1.69x faster than HiAR and 1.42–2.92x faster than Self-Forcing for 16 FPS videos of 20 seconds or longer across 1.3B and 14B backbones at 480p and 720p, and the 1.3B model at 480p scores 0.838 on VBench-1.0 while remaining stable at 20s, 35s, and 65s.
arXiv

KernelZero's 7B Proposer–Coder co-evolution reaches 75.8 CUDA and 77.2 Triton pass@1 on KernelBench

KernelZero introduces a co-evolution framework in which a Proposer generates Torch modules at the Coder's current capability frontier and the Coder is trained with CA-GRPO so that performance is optimized only after correctness becomes reliable; both models start from Qwen2.5-Coder-7B and reach 75.8/69.6 pass@1 on KernelBench CUDA Level 1/2 and 77.2/72.5 on Triton, surpassing Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton.
arXiv

DRM replaces the scalar reward head with a diffusion head: matching or beating same-data, same-backbone baselines on five benchmarks while its reward distribution turns multimodal as human disagreement grows

The work introduces DRM (Diffusion Reward Model), which recasts reward modeling as conditional density estimation: on top of a frozen LLM encoder, a lightweight Diffusion Transformer denoises Gaussian noise into a reward vector with no parametric family assumed, a single architecture handles both multi-attribute regression and pairwise preference data, it matches or surpasses the same-data, same-backbone scalar head ArmoRM and quantile head QRM across five benchmarks, and it exhibits multimodal structure tied to human disagreement on repeated-annotation data, with uncertainty-aware rejection, LCB aggregation, and downstream RLHF experiments demonstrating the practical value of distributional information.
arXiv

LT-OPD uses on-policy self-distillation on its own generated trajectories to lift average retained performance at 5% visual tokens from 68.6% to 82.3%

The work proposes LT-OPD, which supervises a student keeping only 5% of visual tokens on its own generated trajectories using a frozen full-token copy of the same model as teacher, together with a curriculum that progressively lowers the token budget, raising average retained performance on nine benchmarks with Qwen3.5-4B from 68.6% to 82.3% while cutting KV cache by 85.2% and prefill FLOPs by 85.4% with no added inference overhead.
arXiv

GeoVerse injects video generative priors into geometric latent space, raising PSNR by 2.23 dB on DL3DV and cutting ATE by 32.4% on Mip-NeRF360

GeoVerse performs diffusion generation inside the geometric latent space of a pretrained 3D foundation model (DA3), injects frozen Wan2.2 VACE video features into the denoiser through a ControlNet-style adapter, and uses a continuously updated colored point-cloud spatial memory to reproject history as target-aligned guidance, synthesizing cross-view consistent novel views from sparse inputs; across DL3DV, RealEstate10K, Mip-NeRF 360, and ScanNetV2 it reports better visual quality and geometric consistency, with PSNR 2.23 dB higher than GLD on DL3DV and ATE 32.4% lower than GLD on Mip-NeRF 360, while reflow distillation brings inference from over 100 seconds to under 10 seconds.
arXiv

AlignOPSD aligns long-horizon agent supervision to functional decisions rather than timestamps, beating GRPO by 5.5–8.7% on ALFWorld, WebShop, and Search-QA

The work identifies Decision–Timestamp Mismatch in on-policy self-distillation for long-horizon agents—privileged guidance may correspond to a decision at a different timestep, and a student decision may span multiple timesteps—and introduces AlignOPSD, which first rectifies privileged supervision using functional correspondences across sibling rollouts and then applies semi-Markov hierarchical credit assignment over variable-duration decision spans; with Qwen2.5-3B/7B it outperforms GRPO on all eight backbone–aggregate-metric comparisons across ALFWorld, WebShop, and Search-QA by 5.5–8.7%, ranking first in six.
arXiv

CompoWorld composes 448 reusable services into cross-service tasks, lifting Qwen3.6-35B-A3B by 9.17 points on average across eight benchmarks

CompoWorld introduces compositional environment scaling: coding agents turn MCP tool specifications into executable services with typed states and shared interfaces, a random walk connects services through dependency graphs to generate and verify cross-service tasks, verified trajectories support SFT while a Completion-Focused Rubric Reward guides GRPO-based RL, and with 448 services exposing 10,130 tools plus 3K SFT trajectories and 1K RL tasks, Qwen3.6-35B-A3B improves by 9.17 points on average across eight benchmarks and raises AutomationBench task success from 10.33% to 32.33%.
arXiv

DN-MOPD normalizes per-domain teacher feedback and lifts the six-task average over MOPD from 58.4/50.3/26.6 to 59.6/52.5/29.0 across three Qwen3.5 sizes

The work finds that multi-teacher on-policy distillation (MOPD) with domain-label routing alone fails to beat the strongest single-teacher student and transfers little of the mathematics expert's gain, because domain feedback is unbalanced (in the first training batch instruction-following log-ratios are 2.3-4.4 times as dispersed as the pooled signal while mathematics is about half as dispersed, and for the initial 4B student the instruction-following loss supplies 94% of the combined gradient), and proposes DN-MOPD, which keeps label routing and rescales each domain's distillation advantages by the clipped ratio of pooled to domain log-ratio spread estimated every batch, improving the six-task average over MOPD at every size and both evaluation budgets (1.17-2.36 points at 16K, 2.47-3.
arXiv

MT-OPSD uses on-policy self-distillation on self-generated states to lift open-source image editors from 0.03–0.15 to 0.38–0.52 ten-turn success

The work finds that instruction-based image editing models degrade rapidly under recursive multi-turn editing and attributes this to a train–test mismatch in the conditioning distribution, since models are trained on clean source images but must repeatedly condition on their own imperfect outputs at inference; it proposes MT-OPSD, which rolls out the current student under an identity instruction to produce self-generated states containing its own errors, then uses the same model conditioned on the clean source as a teacher to transfer clean-condition editing behavior onto those states via sparse velocity matching, without multi-turn annotations; it also builds LME-Bench of 100 ten-turn editing sessions, raising SR@10 to 0.38–0.52 and cutting CR@10 to at most 0.
arXiv

MALA lets attention allocate its own post-score compute from normalized contribution, cutting 128K training forward/backward latency to 1/2.2 and 1/3.0 of FullAttn

The work introduces MassAlloc Attention (MALA), a fused attention primitive that keeps QK score discovery over every legal causal interaction and uses normalized online-softmax contribution to decide which tiles execute post-score computation; under exactly matched post-score work at 8K it reaches 0.0188% mean omitted mass versus 0.0182% for a per-instance reference-mass oracle, holds low output and gradient errors from 1K to 32K, reaches 89.67% associative-recall accuracy at 8K versus 89.97% for FullAttn, reduces training forward and backward latency by 2.2x and 3.0x and decoding latency by 1.6x at 128K with tensor parallelism on 8 GPUs, tracks FullAttn perplexity from 0.6B to 14B while cutting total training FLOPs by 23.
arXiv

AI Night-Scientist uses reinforcement learning to teach models when to depart from predictable reasoning, widening research-direction coverage by 27.8% and contribution types by 14.9%

The work introduces AI Night-Scientist, a cognitive-science-grounded agentic framework that represents creativity at the action, process, and outcome levels and uses GRPO to train Qwen3-8B/14B models for long-horizon research proposal generation, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model, improving predicted citation impact by up to 32.0 percentage points and originality by 66.2 points, and showing that these gains cannot be reproduced by raising decoding temperature alone, with semantic guidance about what kind of creativity to pursue being critical.
arXiv

A strong agent writes its intervention experience into a playbook for a light agent, lifting real-world manipulation success from 37.3% to 64.0%

The work proposes Recursive Harness Distillation (RHD): a strong agent (GPT-6 Astra) first interacts with a frozen vision-language-action model and distills effective intervention strategies into a reusable playbook, then recursively revises that playbook using the light agent's (GPT-5.6 Luna) execution feedback, improving manipulation without updating any model parameters—real-world success rises from 37.3% to 64.0%, the light agent with the playbook reaches 66.7% on SimplerEnv Bridge versus 41.7% for GR00T alone, and the same playbook also lifts the strong agent to 79.2%.
arXiv

Tencent Hunyuan and collaborators compared 11 MoE rungs and found encoder-free multimodal loss falls faster with compute, crossing over near 10^22 FLOPs

Using a matched ladder of 11 sparse MoE decoders (1.1B–44B total, 71M–2.4B active non-embedding parameters) with identical data mixture and optimization, the work fits separate text and multimodal scaling laws for encoder-free and encoder-based MLLMs and finds that removing the visual encoder shifts compute-optimal multimodal allocation toward larger models (model allocation exponent rising from about 0.464 to about 0.570) while leaving text nearly unchanged; encoder-free models underperform at small scales but their multimodal loss falls faster with compute, with extrapolation predicting an efficiency crossover on the order of 10^22 FLOPs (point estimate about 3.
arXiv

G²PTQ refreshes gradients and Hessians before quantizing each Transformer block, halving KL divergence versus GPTAQ and adding 6.45% QA accuracy at 2-bit

The work introduces G²PTQ, a post-training quantization framework that combines first- and second-order information under a globally supervised, block-wise objective: it recomputes gradient and Hessian estimates before quantizing each Transformer block to avoid staleness and uses trust-region scaling to bound the exact first-order compensation step, achieving lower KL divergence and higher downstream QA accuracy than baselines such as GPTQ, GuidedQuant, and GPTAQ across 13 dense and 2 MoE models from 0.6B to 125B parameters at 2/3/4-bit weight and 4-bit weight-activation settings.
arXiv

CoWA splits causal-history coverage across KV heads, cutting 128K training forward and backward latency to roughly one-seventh of FullAttn

The work introduces CoWindow Attention (CoWA), in which all KV heads share near-diagonal and prefix-sink windows while complementary long-range windows partition the remaining history across heads, making full causal coverage a collective property of the head ensemble; in a window-matched 8K ablation CoWA reaches 89.73% versus 89.97% for FullAttn while duplicated long-range windows perform substantially worse, a 128K tensor-parallel operator benchmark reduces training forward and backward latency by 7.4x and 8.6x and decoding latency by 3.0x, and scaling-law training from 0.6B to 14B tracks FullAttn perplexity while cutting total training FLOPs by 28.5% in the 32K long-context stage.