Skip to main content

Search

“All disciplines” · 1452 results

Page 25 · showing 20
arXiv

TokenCast forecasts token consumption during agent runs with composable segment costs, cutting MAE by 14.5% on average across 96 comparisons

TokenCast learns a composable cost triple for each execution segment of an LLM agent (call count, net input-length change, and a cost residual), uses an exact composition identity to fold the cost of earlier context being re-read by every later call into a cumulative estimate, and refreshes the forecast from newly observed execution evidence without additional LLM calls; on SWE-bench Verified the mean cumulative prediction time is 32.8 ms per run, mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations across 4 task suites and 6 agent models, and offline budget-control replay uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion.
arXiv

Imprint Reader uses SMaRT to read frozen weight updates into natural language, and MetaEdit turns those descriptions into behavioral intervention

The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
arXiv

Moving bidding on device: one tick of sync staleness overspends budgets by 17.77%, and 50 ticks by 1,669.31%

In an auction-logic-faithful on-device simulation with 36 campaigns, 50 devices, and 30 paired demand paths, the study examines two economic misalignments created when privacy-preserving ML decisions move onto clients, finding that proportional Even pacing overspends 17.77% after one tick of staleness and 1,669.31% after 50 ticks, that 50-tick overspend remains 106.95% at two-times budget pressure, that a visible-budget no-sale guard makes zero-lag compliance exact yet leaves 11.88% overspend at one tick, and that 98.23% of rival auctions at one tick admit a profitable deviation once the ML/pacing score transformation is allowed to change payment units.
arXiv

Swapping matrix multiplication for an associative-algebra product: a 110M-parameter model gains 6.2–7.8% generation throughput while GSM8K, MBPP and IFEval all drop

The work replaces ordinary matrix multiplication in Transformer projections with an associative-algebra multiplication table that keeps the full learned weight bank and parameter count, lowers the bilinear rank for the q=2 case from 7 (Strassen's 2×2 algorithm) to 6, and in a controlled pretraining run of two approximately 110M-parameter models over 12.3B tokens observes a 6.2–7.8% end-to-end generation throughput gain across four prompt domains together with lower scores on GSM8K, IFEval and MBPP than the dense baseline.
Harvard University

Harvard commits $150 million to a research initiative, allocating $100 million to Schools in proportion to federal funding expenditures and adding climate, neuroscience, and immunology support

Harvard announced a $150 million strategic research initiative in which $100 million is distributed to Schools in proportion to their federal research funding expenditures and $50 million supports cross-disciplinary internal award programs, covering neuroscience, immunology, inflammation and infectious diseases, energy and climate (two new Salata Institute cross-faculty clusters at up to $600,000 per year for three years), the Frontiers innovation fund, and expanded computing infrastructure.
arXiv

ExpVoyager lets a skill curator navigate raw experience on demand, lifting ALFWorld success from 41.4% to 63.7%

The work reframes agent skill synthesis as a dynamic navigation problem over past experience and proposes ExpVoyager, in which a skill curator uses a Navigable Interface to inspect raw trajectories across views and resolutions while a Navigation State tracks reusable procedural knowledge and open questions, producing task-specific skills for a frozen executor; across ALFWorld, WebShop, and ScienceWorld it improves task success and reduces execution steps over no-skill and existing skill-based baselines, with gains that grow as the experience space scales.
arXiv

SMAT decomposes merging into Scale, Mask and Perturb simulations, lifting the five-merger average by 1.07–2.16 points across four backbones with under 2% training overhead

The work introduces SMAT (Simple MAT), which from a single expert's perspective describes common model-merging methods through three operations—Scale (reweighting its own update), Mask (removing selected coordinates) and Perturb (adding other experts' updates)—and during training samples scaling coefficients, masks and additive noise to simulate merged parameters while jointly optimizing the expert loss and the expected loss at simulated merged parameters, made efficient by periodic scheduling, kernel fusion and parameter storage switching with one forward and one backward pass per step; across four backbones (Llama-3.2-1B-Instruct, Llama-3.1-8B-Instruct, CLIP ViT-B/32 and ViT-L/14), SMAT improves the mean score over five merging methods by 1.07–2.
arXiv

SkillDRE evolves malicious agent skills through a dual-stage loop of pre-execution scanning and runtime feedback, reaching a 45.28% average attack success rate across four victim models on SkillsBench

SkillDRE is a fully automated framework that, given a benign task and its associated skills, autonomously constructs and validates a task-conditioned malicious objective and a verifiable judge rule, then holds both fixed while alternating scanner-guided evolution with runtime-defense-guided refinement, forming a closed loop between the pre-execution and runtime stages; on SkillsBench across four victim models it achieves an average attack success rate of 45.28%, exceeding the strongest baseline by 40.3%, while its final submitted skills receive no SkillScan findings and largely preserve benign-task performance.
arXiv

PyroAdapt lifts daily California wildfire average precision from 21.62% to 24.35–24.57% and captures 344 extra positive cell-days under a fixed 34-cell budget

The work proposes PyroAdapt, a pretrain–retrieve–rank framework that pretrains a wildfire risk model on historical data, retrieves similar historical locations using unlabeled target-year covariates, and fine-tunes through same-day fire–nonfire cell-pair ranking losses (direct ranking, residual pairwise DPO, and selective ranking); over 666 0.25° California grid cells, the three ranking objectives raise daily average precision from 21.62% under continued focal fine-tuning to 24.35–24.57% and Top5% recall from 18.70% to 22.11–22.79%, selective ranking captures 344 additional positive cell-days under a fixed daily budget of 34 cells (5% of area), and rolling evaluations over Yosemite show the ranking gains persist under temporal distribution shift.
arXiv

BaRe-Mem uses Bayesian reliability memory to shift a central model toward autonomous reasoning as advice turns misleading, and finds capable workers earlier on MuSiQue

The work introduces BaRe-Mem, an online Bayesian reliability memory that estimates advisor reliability conditioned on the central model's internal belief representations, uses those estimates to modulate attention and to decide whether to consult; across nine benchmarks and six central models it is more robust to misleading advisor information than debate and majority voting, stays above autonomous reasoning on the more challenging tasks across all tested misleading levels, and in MuSiQue agent-team routing identifies capable workers earlier than routing by historical success counts.
arXiv

Treating EEG masking geometry as the only variable: 58 pre-trained models point to a moderate spatial radius with short temporal blocks, and expose a JEPA-specific bias-inflation collapse

The study unifies EEG self-supervised masking strategies into a three-parameter framework of spatial radius, temporal length and mask ratio, trains 58 models under a fixed REVE-Small backbone and corpus across MAE and JEPA, evaluates them with a linear probe on the 12 OpenEEGBench datasets, and finds that both frameworks agree on a moderate spatial radius with short temporal blocks as the optimum while identifying a JEPA-specific bias-inflation collapse at full-channel masking.
Science News

The Deep End's new season Talk to Me devotes six episodes to people who lost their voices and regained speech through AI voice clones and brain implants

Science News reporter Laura Sanders hosts a new season of The Deep End called Talk to Me, six episodes asking whether a lost voice can be found again, told through people who use AI-cloned voices, a brain implant that pulls words from thoughts, and a synthetic voice for singing, and through what those experiences reveal about why voices matter.
arXiv

Across nearly half a million names and 12 tokenizers, whether a name becomes a single token shifts concept accessibility in fellowship, hiring, clinical, and lending judgments

The study introduces NameTrace, measures direct lexical support for names across nearly half a million first names and 12 LLM-associated tokenizers, and, on atomic versus short-fragmented names matched within the same race/ethnicity–gender strata, finds systematically higher task-aligned concept accessibility for atomic names across fellowship, hiring, clinical assessment, and lending; the differences persist in all eight strata, transfer to unseen names, and hidden-state interventions along the measured task directions shift later constrained choices.
arXiv

SpatialSpeak's QA-style reconstruction pretraining lifts the spatial chain-of-thought gain from 2.6 to 6.9 points, reaching 62.8 on ReVSI

The work introduces SpatialSpeak, a two-stage framework that first jointly trains local marked-point 3D coordinates and global object-center reconstruction as question answering (QA-RP), then supervises the model with spatial chain-of-thought plus visual compensation (CoT-VC) to express and use geometric estimates; on ReVSI it raises the gain from CoT-VC from 2.6 to 6.9 points and, with a 4B backbone, reaches the best results among compared methods on ReVSI (62.8), VSI-Bench (73.0), and SPAR-Bench (76.0).
arXiv

Allspark lets a weak teacher pass reasoning to frozen strong models via alternating chains of thought, lifting ARC-AGI-2 accuracy from 78.1% to 81.6%

The work introduces Allspark, a training and inference framework in which a weak teacher and a frozen copy of the same model alternate reasoning segments during training while only the teacher is updated from weak-model rollouts, and at inference a stronger frozen student replaces the frozen partner; across Qwen3-1.7B/4B math and reasoning tasks and Inkling-Small teacher evaluations on ARC-AGI-2 with Inkling, Kimi-K2.6, and Nemotron-3-Ultra students, accuracy gains are observed along with an explicit accuracy-token tradeoff.
arXiv

Skill2Env synthesizes 2,963 executable environments from skills, and 1.5K fine-tuning trajectories lift a 35B agent from 36.6 to 45.0 across seven benchmarks

Skill2Env is a capability-oriented framework that curates skills from public skill libraries, uses reusable difficulty patterns and task blueprints to jointly synthesize task instructions, execution substrates, workspaces, and rubric-based evaluators, and iteratively hardens tasks using solver execution evidence, yielding 2,963 executable tasks; supervised fine-tuning of Qwen3.6-35B-A3B on 1.5K high-scoring trajectories from these environments raises the unweighted average across seven agent benchmarks from 36.6 to 45.0.
arXiv

AdaTutoRank turns a set-level rubric score into token-level credit via adaptive tutoring, reaching the best overall score across ten benchmarks with fewer retrieval calls

The work proposes AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level, nine-dimension rubric hierarchy that assigns each rollout one of three hint forms—rubrics, a corrective reflection, or a sibling set—matched to its reward, and distills the hint's effect into a token-level advantage combined with the group-relative outcome advantage; across ten benchmarks spanning RAG, deep research, and setwise evaluation it attains the best overall performance while reducing the agent's retrieval calls.
arXiv

VQS computes answers with programs over structured image records, cutting self-evolving VLM pseudo-label errors from 24% to 6%

VQS has a model first parse an image into a structured record such as a scene graph, chart table, or diagram graph, then uses fixed program templates to write questions and compute their answers from that record while the model only confirms the atomic facts the program reads one at a time; human raters find 94.4% of VQS answers correct versus 76.4% for majority voting and 82.2% for a model judge, and across ten benchmarks it improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales, reaching 3.84 points at 2B after three training cycles.
arXiv

TT-VidT splits appearance and motion into two pathways: among 24 architecture-objective pairings, only TT3D with Diff Compression leads Jester, SSv2, ARID and Diving48 at once

The work proposes TT-VidT, which pairs a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer and trains it by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens; under a matched recipe at roughly 170M-190M encoder scale on 1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, a 24-cell architecture-objective sweep shows that only TT3D with Diff Compression enters the strongest motion-sensitive regime, and in the final comparison it leads Jester, Something-Something V2, ARID and Diving48 fine-tuning simultaneously while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA 2.
Mistral AI

Mistral opens a Munich hub with Physics AI and Industrial AI teams, partnering with BMW, Siemens Energy and TUM

Mistral announced a new German hub in Munich housing research teams dedicated to Physics AI and Industrial AI plus applied engineers serving enterprise partners, and disclosed that it acquired Emmi AI (bringing in more than 30 physicists, researchers and engineers), is working with BMW on crash simulations and engineering AI and with Siemens Energy on industrial AI applications, and has formed a research partnership with the Technical University Munich (TUM) to use TUM's wind tunnel facilities with Prof. Dr. Nikolaus A.