Skip to main content

Search

“All disciplines” · 1400 results

Page 12 · showing 20
arXiv

ROSS reuses discarded historical self-generated rollouts, lifting Qwen3.6-35B-A3B's six-benchmark MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%

ROSS introduces a selective-supervision relearning procedure that takes historical self-generated rollouts saved from domain-specific RL, multi-teacher on-policy distillation (MOPD), or agentic RL, uses an outcome verifier to keep successful trajectories and an LLM reviewer to mark which model-generated spans are worth imitating, then applies loss only to those selected tokens while keeping the full trajectory as context, improving upstream checkpoints through an offline SFT stage without new policy rollouts and raising Qwen3.6-35B-A3B's six-benchmark MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%.
arXiv

TabFM-Auto pairs an LLM agent with frozen TabFM to evolve data pipelines, lifting Elo from 1785 to 2013 across all 51 TabArena datasets

TabFM-Auto pairs the frozen tabular foundation model TabFM with a language-model coding agent that iteratively rewrites four pipeline stages—data cleaning, feature engineering, context selection, and post-processing—guided by dataset metadata and validation feedback; across all 51 TabArena datasets five configurations take the top five overall positions, the best (Codex with Opus 5) raises TabFM from 1785.3 to 2013.0 Elo, the discovered pipelines transfer without further search to TabPFN-3, TabICLv2, and EXAONE-Tabular (+69 to +143 Elo), and TabFM-Auto ranks first overall among MLE agents on the 8 tabular competitions of MLE-Bench.
arXiv

Tsinghua team turns the SIS proposal into a learning problem with GFlowNets: one network zero-shot matches or beats the post-hoc best of 31 analytic proposals on 1,190 unseen margins

The work shows that the zero-variance sequential importance sampling (SIS) proposal for binary matrices with fixed margins is exactly the policy of a unit-reward GFlowNet, and proposes MarginFlow, a set transformer that reads the remaining margins and, trained on 1,904 margins, runs zero-shot on 1,190 held-out margins, matching or beating the post-hoc best of 31 analytically designed configurations on 1,187 of them with a median effective sample fraction of 99.8%.
arXiv

HDL locates branch points by hindsight divergence, cutting generated tokens 35–61% and lifting agent tasks by up to 12.46 points

The work introduces Hindsight-Divergence Localization (HDL), which picks branch points in a trajectory by how much token log-likelihoods change once verifier feedback and a reflection are added, then builds each training group from a few complete root trajectories plus continuations that reuse the root prefix; across math, code, and agent tasks with three models it cuts generated tokens by 35–61% and rollout wall-clock time by 18–45% relative to GRPO while improving task performance, with gains up to 12.46 percentage points on agent tasks.
arXiv

HSTA tracks technology diffusion across 30,000 arXiv preprints and USPTO patents, finding the highest semantic drift in Large Language Models (0.332) while paper-volume velocity fails to Granger-cause frontier compute surges

The study introduces Hyperspherical Semantic Trajectory Analysis (HSTA), an unsupervised pipeline that encodes 20,000 arXiv preprints and 10,000 USPTO patent abstracts with Sentence-BERT, projects them onto a unit hypersphere, clusters them into eight sub-topics with Spherical K-Means alongside UMAP reduction, and defines two metrics, Semantic Centroid Vector Drift and Commercialization Offset, then links them to Epoch AI compute data through Vector Autoregressive Granger tests, finding that Large Language Models (drift 0.332) and Artificial Intelligence Systems (0.234) evolve fastest semantically while quarterly paper volume velocity alone does not Granger-cause frontier training compute surges at conventional significance levels.
arXiv

TRM first writes a case-adaptive rubric before scoring, beating open-source reward models and nearing proprietary ones on image generation and editing benchmarks

The work introduces the "Think Before You Score" paradigm and the Thinking Reward Model (TRM), which first generates a case-adaptive rubric for each task condition and candidate output, then inspects the candidate criterion by criterion and aggregates the evidence into a fine-grained pointwise reward; it also proposes PD-GRPO to use pairwise preference supervision for better discrimination while mitigating score polarization, and reports state-of-the-art results among open-source reward models on image generation and editing reward-modeling benchmarks, highly competitive with proprietary alternatives, with TRM-guided reinforcement learning consistently improving diverse visual generation models.
arXiv

PMOPD projects parameter updates away from protected task subspaces to ease the multi-teacher distillation capability seesaw, lifting the three-task average by 2.54 points on Qwen2.5-7B and 2.09 on Llama-3.1-8B

The work proposes PMOPD, which builds per-matrix low-dimensional subspace memories from the cumulative parameter displacements of completed task blocks and projects both gradients and Adafactor optimizer updates of later tasks to remove components conflicting with protected task directions, while a lightweight conflict probe sets task order and a subspace-consistency criterion sets the cycle count; across Code, Reason, and Math, PMOPD improves every evaluated capability over MOPD, raising the three-task average by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B.
Nature News

Clark fired the Z machine at water-infused glass and found melt may grow easier to compress under deep pressure, offering a clue to how early Earth kept its water

Mineral physicist Alisha Clark used the Z machine at Sandia National Laboratories to send shockwaves through water-infused glass samples, recreating pressures near Earth's core during its formation; earlier experiments showed the wet glass became easier to compress as pressure rose, leading her to propose that molten rock could have locked water inside Earth during its magma-ocean phase rather than losing it all or receiving it later from comets and asteroids.
arXiv

PrismQuant rotates dominant activation energy into the null space of grouped quantization, reaching 3.85 perplexity at W4A4KV4 on Llama-3.1-70B, only 0.22 points below full precision

The work introduces PrismQuant, a quantizer-aware rotation framework that aligns the leading activation eigenspace with the group-constant subspace of asymmetric grouped INT4, formulates rotation design as a Ky Fan trace maximization with a provably optimal closed-form solution, realizes it with compact Householder (compact-WY) transforms without gradient training, and evaluates W4A4KV4 post-training quantization on Llama, Qwen, and Mistral (including dense models up to 70B and a 30B mixture-of-experts model): on Llama-3.2-3B it sets the state of the art among compared methods in perplexity and accuracy, on Llama-3.1-70B it attains 3.85 perplexity and 72.46% average zero-shot accuracy (0.22 percentage points below full precision), and in a Llama-3.
arXiv

Braco reframes visual-token compression from picking tokens to re-parameterizing them, reaching 95.2% normalized accuracy at 23–64x compression while cutting prefill FLOPs by 84.2%–86.7%

The work recasts extreme visual-token compression as a token-parameterization problem, separating basis transformation and structured truncation (which fix the retained subspace and compressibility) from coordinate organization (which affects optimization and cross-modal alignment, i.e. learnability), and designs Braco, a lightweight four-step coder combining transform-basis truncation, input-independent basis-coordinate embeddings, budget-dependent orthogonal re-parameterization, and learned spatial residual tokens from lightweight pooling; experiments show Braco forms the favorable empirical accuracy–efficiency frontier under 23x–64x compression and remains competitive at 144x, reaching 95.2% accuracy while reducing prefill FLOPs by 84.2%–86.
arXiv

WISE-ATTA shifts active test-time adaptation from which sample to label to when to label, reaching lower error on ImageNet-C/R/K/A with fewer labels

The work introduces budgeted active test-time adaptation (ATTA), in which labels are available for only a fraction of batches in a long test stream, and proposes WISE-ATTA: lightweight online signals (the fraction of low-entropy predictions) plus a sliding-history quantile threshold and a budget-debt controller decide when to request supervision, while prediction drift relative to an EMA anchor selects a single sample to label; on ImageNet-C under CTTA/FTTA and on ImageNet-R/K/A it matches or improves on recent ATTA methods at one label per batch and reports up to 50% fewer labels.
arXiv

FRAC swaps exponential forgetting for power-law long memory in SSMs, beating Mamba and GDN on long-context 1.3B language modeling

The work introduces FRAC, a selective state space model architecture derived from fractional dynamics that approximates a heavy-tailed fractional kernel with a finite-state, log-spaced sum of exponential modes, replacing exponential forgetting with power-law long memory and improving long-context performance over SSM baselines such as Mamba2, GDN, and Mamba3 on synthetic long-tail and recall tasks, 1.3B-parameter language modeling, and DNA modeling, while staying competitive on short-context tasks.
arXiv

SaveRouter Cuts LLM Routing's Break-Even Deployment Volume by Up to About 9.5x Using Roughly a Third of the Supervision

The work proposes SaveRouter, a sparse-supervision routing framework that selectively acquires informative query-model feedback and shares capability information across related queries, using only about 33-41% of available training feedback across four routing benchmarks while maintaining competitive or better routing quality and reducing the break-even deployment volume by roughly 1.9-9.5x relative to the fastest conventional fully supervised router.
arXiv

Deferring particle expansion to late transformer layers and adding time-dependent scoring-rule schedules lets a DDM trained from scratch in one stage reach 4.48 FID at 4 steps and 2.38 at 50 steps on ImageNet-256²

To address two obstacles in scaling Distributional Diffusion Models (DDMs)—multi-particle training overhead that grows with the number of particles, and globally fixed scoring-rule hyperparameters that force a single trade-off across sampling budgets—the work defers particle expansion to late transformer layers and introduces time-dependent scoring-rule schedules informed by the dynamical regimes of Biroli2024; combined with a DiT-based latent setup, this makes DDM training practical on class-conditional ImageNet-256², reaching 4.48 FID at 4 steps and 2.38 at 50 steps with DiT-XL/2 from a single model trained from scratch in one stage, without a teacher, self-distillation or JVPs, and the same recipe transfers to text-to-image generation.
arXiv

Meituan's LongCat team splits deep research into planning, parallel section research, and global-to-local editing via ResearchSpec, scoring 55.25, 51.35, and 79.83 on three public benchmarks

Meituan's LongCat team presents LongCat-DeepResearch, which shifts early research iteration from the full report to an executable ResearchSpec: multiple planning agents first search external sources and consolidate a research plan, researchers then investigate and draft citation-bearing sections in parallel independent contexts, and a Global Editor assigns cross-section ownership while Local Editors make targeted revisions, yielding 55.25 on DeepResearchBench, 51.35 on DeepResearchBench II, and 79.83 on ResearchRubrics, and 76.04 on an in-house benchmark, second among four compared systems.
arXiv

APM-Bench tests cross-session persistent memory for streaming video assistants across 549 sessions and 104 life trajectories, finding existing methods struggle to combine recall, latency, and storage

The work introduces APM-Bench, which reformulates egocentric streaming interaction as multi-session life trajectories (549 sessions, 104 trajectories, 2,719 candidates, averaging 69 minutes of video per trajectory) and evaluates general video models and eight specialized memory systems on cross-session understanding, real-time perception, and adaptive response, plus an evidence-availability-aware test; results reveal a clear utility-latency-storage trade-off: raw video memory gives the strongest cross-session performance but needs GiB-scale storage and high latency, text summaries cut storage to the KiB scale at the cost of cross-session performance, event-structured memory shows the strongest utility among specialized systems, and adaptive response remains difficult even with rich history
arXiv

VoxMem tests 15 audio LLMs on 3,196 questions: none tops 40% at 32K, and swapping audio for transcripts drops speaker accuracy from 69.8% to 10.3%

The authors propose a two-axis taxonomy of spoken conversational memory — acoustic evidence type (speech semantics, speaker identity, paralinguistic cues, environmental sound) crossed with memory operation (information extraction, multi-session reasoning, temporal evolution tracking, answer refusal) — and build VoxMem on it: 3,196 evaluation instances over 34,743 spoken sessions (177 hours) at four context budgets from 8K to 64K tokens; evaluating 15 large audio language models, no model exceeds 40% overall accuracy at 32K (best 38.5%), models remember what was said far better than who said it, how it was said, or what was audible, and replacing audio with exact transcripts drops speaker accuracy from 69.8% to 10.3% while speech-semantics accuracy barely moves (75.9% to 71.0%).
arXiv

ZJU team runs 25 same-family on-policy distillation pairs on Qwen2.5 from 0.5B to 14B, finding peak capability is predictable from student scale and teacher score, and that weaker teachers can teach stronger students

The study systematically characterizes the scaling properties of same-family on-policy distillation (OPD) on Qwen2.5 (0.5B–14B) math reasoning, finding a regular early useful-transfer regime in which held-out accuracy rises approximately linearly with the square root of token-level reverse KL from the student initialization, and fitting power laws in student scale, teacher scale, and teacher gold score that predict peak accuracy and transfer rate, with every weak-to-strong student peaking above its own teacher.
arXiv

SoL-Refiner refines low-resolution generated video to 4K in one denoising step, cutting 2K refinement latency from 57.461 to 6.447 seconds

The work presents SoL-Refiner, a one-step video refiner that turns low-resolution generator outputs into 4K video through a three-stage recipe of high-resolution continual training, reinforcement-learning post-training with frame-based reward models, and final one-step distillation, and introduces Refiner-Bench, a video refinement benchmark using a shared-input protocol to compare refiners at roughly 2K output resolution; at 2K the one-step model outperforms all evaluated external refiners on VBench and UniPercept averages, at 3840×2176 it improves both metrics over the three-step LTX-2.3 Refiner, and with the complete acceleration stack it achieves an 8.91× speedup in refinement latency over the same baseline in the 2K latency setting.
arXiv

PanoVLN lifts R2R-CE success rate to 77.3%, 11.9 points above the previous best, using panoramic vision

PanoVLN performs vision-and-language navigation in continuous environments with a 4B vision-language model and RGB-only panoramic input, combining longer action-sequence supervision, confidence-guided execution, decision-centric training data, and fused panoramic geometric features to reach 77.3% and 78.0% success rate on R2R-CE and RxR-CE Val-Unseen, exceeding the previous best by 11.9 and 8.7 percentage points, and enabling real-world quadruped navigation with fewer pauses.