Skip to main content

Search

“All disciplines” · 1474 results

Page 31 · showing 20
arXiv

CARD pairs cluster-level LoRA with decoding-time preference vectors, taking 10 of 12 metric settings across six LaMP and LongLaMP tasks

CARD introduces a hierarchical framework for personalized text generation: it clusters users by shared stylistic patterns and trains group-specific LoRA adapters, derives lightweight user preference vectors through implicit preference learning that contrasts user-authored text with cluster-level generations, and injects personalization at inference only via low-rank logit corrections, ranking first in 10 of 12 settings across six LaMP and LongLaMP tasks and two metrics while remaining stable for low-resource users, across model scales, and in storage efficiency.
arXiv

MOPD-Router replaces domain-label hard routing with token-level routing, lifting multi-teacher distillation overall score by 5.88 points on unlabeled mixtures

The work introduces MOPD-Router, a multi-teacher on-policy distillation framework that needs neither domain labels nor a separately trained routing model, selecting and weighting the full teacher pool at every token, and proposes ExpertAlign, a metric that scores each teacher by the positive cosine alignment between the specialization it acquired relative to the shared pre-RL base and the teaching direction it would apply to the current student; across unlabeled and domain-labeled training mixtures under strong-to-weak and same-size distillation, ExpertAlign achieves the strongest overall performance in all four settings, improving the overall score by 5.88 points (+12.3%) over Mean aggregation on unlabeled data and by 3.95 points (+7.
arXiv

CodeGraph annotates 145 million source files into a knowledge graph with about 1 billion typed edges and grounds its concepts in Wikidata

The work presents a pipeline that uses a code-specialised LLM (based on Qwen3-Coder-30B-A3B-Instruct) to annotate source files under an open taxonomy, extracting algorithms, paradigms, design patterns, and application domains, then grounds these in Wikidata through a three-stage procedure (deterministic SPARQL, a Deep Research Agent for the long tail, and parent-of hierarchy rollup), with a calibrated quality-assurance protocol combining a human gold set and an LLM-as-a-judge filter; applied to the Stack-Edu corpus it yields CodeGraph with roughly 158 million nodes (about 145 million file nodes, about 63,000 extracted concept entities, and roughly 19,800 grounded Wikidata entities) and about 1 billion typed edges across 14 programming languages.
arXiv

D-JEPA supervises relations among candidate futures with executed outcomes, reaching 87.89% on PushT and a 17-point gain on physical robots

D-JEPA introduces a decision-aligned latent world model: it first quantifies a decision-local prediction gap in which, among the few futures competing for execution, a candidate predicted closer to the goal can fail while an available alternative succeeds; it then uses executed outcomes to supervise a bounded, permutation-equivariant set operator that learns decision-relevant relations among candidate futures, and realizes the learned decision structure in JEPA-compatible future representations so aligned actions can be read out through native latent-distance planning; it reaches 87.89% success on PushT, a 15.04-point average gain on RoboTwin, a 17-point gain on physical robot tasks, and raises mean driving PDMS from 57.36 to 95.34.
arXiv

FuseReg trains decoders and DiTs on random encoder-layer subsets, letting one decoder reconstruct from full, sparse, and single-layer fusions and lowering gFID with the generator unchanged

The work introduces FuseReg, which replaces heuristic layer selection in representation autoencoders (RAEs) with training over random subsets of encoder layers: on ImageNet-256 with a frozen DINOv3-L, a single FuseReg decoder reconstructs from full, sparse, and single-layer fusions without retraining and achieves higher PSNR than decoders specialized to fixed fusions, while decoder replacement alone (with an unchanged RAEv2 DiT-XL generator) reduces unguided gFID and joint regularization of decoder and generator further reduces unguided gFID on DiT-Base.
arXiv

RayOrch separates cross-parent batching from ordered gathering via lineage state, cutting MinerU end-to-end time 13.1% below Ray Data on 64 H20 GPUs

RayOrch presents a programming model and Ray-based distributed execution engine that uses compiler-validated F.expand/F.reduce pairs and runtime-maintained structural lineage state (child set, immediate parent, immutable ordinal, terminal state) for multi-grain dataflows, so GPUs can batch children across parents while still reconstructing parent results in order and containing failures per parent; on NVIDIA H20 GPUs, MinerU scales from 4 to 64 GPUs with a 15.14x processing-time speedup and finishes 64-GPU end-to-end in 4295.7 seconds at 40.6788 pages/s, 13.1% less time than Ray Data and 29.0% less than Daft, Docling is 16.0% faster than Ray Data, FIFO dispatch lowers ablation wall time from 634.1 to 579.3 seconds (8.
arXiv

PISA cuts block-sparse attention selection to O(N log N) with pyramid Top-K, running 9.95x faster than BSA at 256K

Researchers from Shanghai Jiao Tong University and ByteDance Seed propose PISA, a block-sparse attention mechanism that uses coarse-to-fine pyramid Top-K selection with LogSumExp scoring and hardware-aware Triton kernels to reduce block-selection complexity from O(N²/C) to O(N log N); across 418M, 1.47B, and 2.67B scales it matches BSA, NSA, and HiLS on language modeling and commonsense reasoning, achieves the highest average accuracy among sparse methods on six containment tasks, and speeds up block selection over BSA by 2.86x, 5.31x, and 9.95x at 64K, 128K, and 256K.
arXiv

Adding a paragraph coordinate to attention makes real paragraphs compress deeper than random labels, but depth varies by corpus and no corpus-only statistic fully reproduces it

Using a hierarchical rotary positional encoding (hRoPE) that splits paragraph, sentence, and token indices into separate channels, the authors intervene on the paragraph coordinate while holding the token sequence fixed and measure cross-paragraph attention with a token-distance-exact estimator, finding that attention is compressed in all three corpora but that a density-matched random-label channel is compressed too; what distinguishes real structure is the depth of compression, which is greater and corpus-dependent while the control's is not, and across eight corpus-only quantities spanning lexical persistence, paragraph length, and embedding-based coherence, none fully reproduces the cross-corpus ordering of depth, though embedding-based coherence comes closest.
arXiv

Berkeley team's Morphometric Imitation retargets human hand-object interaction to three-, four-, and five-fingered robot hands, reaching 89.3% zero-shot success in 300 real-world trials

The work presents Morphometric Imitation, a three-stage framework that first kinematically retargets human hand-object interaction to robot hands of different morphologies while preserving demonstrated contacts via morphometric optimization (MMO), then uses residual reinforcement learning with object pose and contact information to produce dynamically feasible demonstrations, and finally distills them into visuomotor policies; across three robot hands and ten GRAB trajectories, MMO improves contact F1 over the strongest of five baselines by at least 8 points, downstream dynamic retargeting success by as much as 35 points, and the distilled policies achieve 89.3% zero-shot success in 300 real-world trials on 30 objects.
arXiv

VLA-Precision reaches 98.3% mean success across nine precision chemistry tasks with 45.8 minutes of online training per task

The work presents VLA-Precision, a framework for real-world online reinforcement learning of large vision-language-action (VLA) models, combining the Asymmetric Co-Bootstrapping (ACoB) algorithm with the ACoB-Stream training architecture, and reports 98.3% mean success across nine high-precision chemistry tasks in four categories and four robot platforms, with 45.8 minutes of online training per task, 27.6-second episodes, and up to 10.9x improvements in throughput and computational efficiency.
arXiv

PsPLUG uses a "personalization residual" plug-in to curb style-instruction suppression of user preferences, outperforming baselines on LaMP

The work identifies that explicit style instructions can erode the user-specific traits that personalized LLMs aim to preserve (a failure mode it calls personalization collapse), proposes modeling personalization as a distributional residual between the user's true linguistic distribution and the base model's neutral distribution under the same input, and builds PsPLUG: a lightweight plug-in that prepends a 3-token prefix (system instruction vector, user vector, input vector) to a frozen Qwen3-8B backbone, learns the user-specific residual with a style-conditioned Bradley–Terry preference objective, and tunes personalization strength at inference via a scaling coefficient; on the LaMP benchmark, PsPLUG generally outperforms non-personalized, RAG, PAG, PPlug, and OPPU baselines without styl
Nature News

Third-party firms openly recruit paid reviewers: about 1,000 job ads and five interviewees reveal a market for outsourced journal peer review

An as-yet un-peer-reviewed MetaArXiv preprint manually scraped about 1,000 online advertisements for 'freelance' peer-reviewer jobs from LinkedIn, company websites and other recruitment platforms and interviewed five working and retired researchers in India, Spain and the United Kingdom, finding a market in which third-party firms offer paid reviewing services to scholarly publishers; interviewees often had 2–3 days or even 24 hours to assess a manuscript and most received €30–40 per review, while one company's website listed more than 10,000 registered freelance peer reviewers, including 5,669 marked 'active' who had each reviewed up to 50 manuscripts, though the study does not provide conclusive evidence about which journals use these services or how widespread the practice is.
arXiv

TRACE audits streaming video understanding across 1,240 records from 517 videos, finding that near-identical accuracy hides large gaps in completion, invalid output, and response timing

TRACE introduces a condition-aware benchmark and evaluation framework for streaming video understanding that makes temporal validity, execution conditions, and operational outcomes explicit, evaluating eight publicly available models or systems in eight configurations on 1,240 records from 517 videos, and finds that nearly identical QA accuracy can mask substantial differences in completion, answer validity, and generation workload, while proactive performance separates into response quality, response delay, false alarms, and missed target windows.
arXiv

ISTA-DASLab proposes disaggregated quantization: separate formats for prefill and decode, more than doubling 1-bit decode accuracy

The work proposes disaggregated quantization (DQ), which specializes computation formats, weights and storage placement to the prefill and decode phases of LLM inference, together with the QADD training method; on Qwen 3 and Gemma 3, removing activation quantization only on decode improves accuracy on decode-heavy tasks without increasing inference cost, training separate NVFP4 prefill weights accelerates prefill while raising accuracy at 2–3-bit decode, and on Qwen3.8-27B 1-bit GGUF decoders it more than doubles MMLU-Pro and MMMU-Pro accuracy, with offloaded disaggregated prefill streaming prefill weights from SSD delivering a time-to-first-token speedup over the weight-only baseline at 8K context.
arXiv

Samsung Research's Net Utility allocates LoRA merging rank budgets per singular direction, lifting vision tasks by 2.1% and language tasks by 2.2% on average

The work identifies the uniform assumption that every layer and every task receives the same rank budget as a major source of the gap between merged and per-task LoRAs, and introduces Net Utility, a data-free metric that decomposes each task LoRA by SVD, scores every singular direction by its benefit to its own task minus its interference with other tasks, and then globally selects the highest-scoring directions under a total rank budget; applied on top of five merging methods across three merging spaces on 7 vision and 6 language tasks, it improves performance by 2.1% on average for vision and 2.2% for language.
arXiv

IndicBankBench tests 799 Indian retail-banking cases and finds eleven models reach only 43.7%–58.2% strict three-run reliability while at-least-once success runs 60%–74%

The authors release IndicBankBench, a 799-case benchmark for Indian retail banking spanning five operational domains, a capability/refusal domain, and twenty primary axes, graded in four stages—safety, action and tool use, response adequacy, and advisory quality—with deterministic tool-use and most safety checks, an LLM judge only for semantic response adequacy, and a narrow resolver for ambiguous confirmation-before-write; running every case three times across eleven models yields strict pass3 of 43.7%–58.2% versus at-least-once success of 60%–74%, a gap of 10.8–21.4 percentage points.
arXiv

Refining photogrammetric DSMs with a pretrained diffusion model and multimodal conditioning cut Dense Urban RMSE from 6.00 m to 3.45 m in French cities

The study adapts pretrained Stable Diffusion 3 into an image-only generative backbone with a pruned text stream and patch-wise normalization, conditioning on both photogrammetric DSMs and Pléiades-HR imagery via two ControlNets to refine vertically co-registered DSMs, reducing Dense Urban RMSE from 6.00 m to 3.45 m across eight in-context French cities and from 4.16 m to 2.77 m in the geographically held-out city of Bordeaux.
Nature News

Transplanted hearts shift their biological age toward the host: mouse grafts and 11 human transplants show donor hearts ageing or rejuvenating

Jesse Poganik's team at Harvard grafted hearts from young, middle-aged and old mice into mice of all three age groups and measured DNA methylation across roughly 320,000 genomic regions, finding that grafted hearts shifted their biological age toward the recipient; the same effect appeared in biopsy samples from 11 historical heart transplants at Brigham and Women's Hospital with large donor-recipient age gaps, while some functional measures such as heart rate, posterior wall thickness and exercise capacity tracked recipient age rather than donor age.
arXiv

An equivalent reparameterization of the output head cuts Phi-4-mini W4 AW-MSE KL from 0.936 to 0.256 while keeping a 10.8% batch-one latency gain

The work proposes softmax reparameterization: before quantization, subtract a scalar multiple of the vocabulary-row mean from every output-head row and pick the coefficient by validation KL separately for RTN, AW-MSE and full-Hessian GPTQ, preserving the full-precision softmax distribution and the trained decoder while improving low-bit output-head fidelity, for example lowering Phi-4-mini W4 AW-MSE KL from 0.936 to 0.256 and reducing batch-one generation latency by 10.8% relative to a BF16-head baseline in packed W4 deployment.
arXiv

TrackEverything pushes dense 3D point tracking past 1000 frames by de-duplicating scene representations, beating open-source all-frame dense trackers by over 20% APD on short clips

TrackEverything represents videos as persistent 3D scene tracks in world coordinates and, through voxelization-based de-duplication at sliding-window boundaries, an endpoint-then-trajectory decomposition, and 3D WAFT feature sampling, becomes the first 3D tracker to follow all visible points across videos exceeding 1000 frames within 40 GB of GPU memory, outperforming all open-source all-frame dense 3D trackers by more than 20% APD on TAPVid-3D short clips while remaining competitive with state-of-the-art sparse trackers on long sequences.