arXiv TokenCast learns a composable cost triple for each execution segment of an LLM agent (call count, net input-length change, and a cost residual), uses an exact composition identity to fold the cost of earlier context being re-read by every later call into a cumulative estimate, and refreshes the forecast from newly observed execution evidence without additional LLM calls; on SWE-bench Verified the mean cumulative prediction time is 32.8 ms per run, mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations across 4 task suites and 6 agent models, and offline budget-control replay uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion.
TokenCast learns a composable cost triple for each execution segment of an LLM agent (call count, net input-length change, and a cost residual), uses an exact composition identity to fold the cost of earlier context being re-read by every later call into a cumulative estimate, and refreshes the forecast from newly observed execution evidence without additional LLM calls; on SWE-bench Verified the mean cumulative prediction time is 32.8 ms per run, mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations across 4 task suites and 6 agent models, and offline budget-control replay uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion.
TokenCast learns a composable cost triple for each execution segment of an LLM agent (call count, net input-length change, and a cost residual), uses an exact composition identity to fold the cost of earlier context being re-read by every later call into a cumulative estimate, and refreshes the forecast from newly observed execution evidence without additional LLM calls; on SWE-bench Verified the mean cumulative prediction time is 32.8 ms per run, mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations across 4 task suites and 6 agent models, and offline budget-control replay uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion.
TokenCast learns a composable cost triple for each execution segment of an LLM agent (call count, net input-length change, and a cost residual), uses an exact composition identity to fold the cost of earlier context being re-read by every later call into a cumulative estimate, and refreshes the forecast from newly observed execution evidence without additional LLM calls; on SWE-bench Verified the mean cumulative prediction time is 32.8 ms per run, mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations across 4 task suites and 6 agent models, and offline budget-control replay uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion.
arXiv The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
arXiv In an auction-logic-faithful on-device simulation with 36 campaigns, 50 devices, and 30 paired demand paths, the study examines two economic misalignments created when privacy-preserving ML decisions move onto clients, finding that proportional Even pacing overspends 17.77% after one tick of staleness and 1,669.31% after 50 ticks, that 50-tick overspend remains 106.95% at two-times budget pressure, that a visible-budget no-sale guard makes zero-lag compliance exact yet leaves 11.88% overspend at one tick, and that 98.23% of rival auctions at one tick admit a profitable deviation once the ML/pacing score transformation is allowed to change payment units.
In an auction-logic-faithful on-device simulation with 36 campaigns, 50 devices, and 30 paired demand paths, the study examines two economic misalignments created when privacy-preserving ML decisions move onto clients, finding that proportional Even pacing overspends 17.77% after one tick of staleness and 1,669.31% after 50 ticks, that 50-tick overspend remains 106.95% at two-times budget pressure, that a visible-budget no-sale guard makes zero-lag compliance exact yet leaves 11.88% overspend at one tick, and that 98.23% of rival auctions at one tick admit a profitable deviation once the ML/pacing score transformation is allowed to change payment units.
In an auction-logic-faithful on-device simulation with 36 campaigns, 50 devices, and 30 paired demand paths, the study examines two economic misalignments created when privacy-preserving ML decisions move onto clients, finding that proportional Even pacing overspends 17.77% after one tick of staleness and 1,669.31% after 50 ticks, that 50-tick overspend remains 106.95% at two-times budget pressure, that a visible-budget no-sale guard makes zero-lag compliance exact yet leaves 11.88% overspend at one tick, and that 98.23% of rival auctions at one tick admit a profitable deviation once the ML/pacing score transformation is allowed to change payment units.
In an auction-logic-faithful on-device simulation with 36 campaigns, 50 devices, and 30 paired demand paths, the study examines two economic misalignments created when privacy-preserving ML decisions move onto clients, finding that proportional Even pacing overspends 17.77% after one tick of staleness and 1,669.31% after 50 ticks, that 50-tick overspend remains 106.95% at two-times budget pressure, that a visible-budget no-sale guard makes zero-lag compliance exact yet leaves 11.88% overspend at one tick, and that 98.23% of rival auctions at one tick admit a profitable deviation once the ML/pacing score transformation is allowed to change payment units.
arXiv The work replaces ordinary matrix multiplication in Transformer projections with an associative-algebra multiplication table that keeps the full learned weight bank and parameter count, lowers the bilinear rank for the q=2 case from 7 (Strassen's 2×2 algorithm) to 6, and in a controlled pretraining run of two approximately 110M-parameter models over 12.3B tokens observes a 6.2–7.8% end-to-end generation throughput gain across four prompt domains together with lower scores on GSM8K, IFEval and MBPP than the dense baseline.
The work replaces ordinary matrix multiplication in Transformer projections with an associative-algebra multiplication table that keeps the full learned weight bank and parameter count, lowers the bilinear rank for the q=2 case from 7 (Strassen's 2×2 algorithm) to 6, and in a controlled pretraining run of two approximately 110M-parameter models over 12.3B tokens observes a 6.2–7.8% end-to-end generation throughput gain across four prompt domains together with lower scores on GSM8K, IFEval and MBPP than the dense baseline.
The work replaces ordinary matrix multiplication in Transformer projections with an associative-algebra multiplication table that keeps the full learned weight bank and parameter count, lowers the bilinear rank for the q=2 case from 7 (Strassen's 2×2 algorithm) to 6, and in a controlled pretraining run of two approximately 110M-parameter models over 12.3B tokens observes a 6.2–7.8% end-to-end generation throughput gain across four prompt domains together with lower scores on GSM8K, IFEval and MBPP than the dense baseline.
The work replaces ordinary matrix multiplication in Transformer projections with an associative-algebra multiplication table that keeps the full learned weight bank and parameter count, lowers the bilinear rank for the q=2 case from 7 (Strassen's 2×2 algorithm) to 6, and in a controlled pretraining run of two approximately 110M-parameter models over 12.3B tokens observes a 6.2–7.8% end-to-end generation throughput gain across four prompt domains together with lower scores on GSM8K, IFEval and MBPP than the dense baseline.
Harvard University Harvard announced a $150 million strategic research initiative in which $100 million is distributed to Schools in proportion to their federal research funding expenditures and $50 million supports cross-disciplinary internal award programs, covering neuroscience, immunology, inflammation and infectious diseases, energy and climate (two new Salata Institute cross-faculty clusters at up to $600,000 per year for three years), the Frontiers innovation fund, and expanded computing infrastructure.
Harvard announced a $150 million strategic research initiative in which $100 million is distributed to Schools in proportion to their federal research funding expenditures and $50 million supports cross-disciplinary internal award programs, covering neuroscience, immunology, inflammation and infectious diseases, energy and climate (two new Salata Institute cross-faculty clusters at up to $600,000 per year for three years), the Frontiers innovation fund, and expanded computing infrastructure.
Harvard announced a $150 million strategic research initiative in which $100 million is distributed to Schools in proportion to their federal research funding expenditures and $50 million supports cross-disciplinary internal award programs, covering neuroscience, immunology, inflammation and infectious diseases, energy and climate (two new Salata Institute cross-faculty clusters at up to $600,000 per year for three years), the Frontiers innovation fund, and expanded computing infrastructure.
Harvard announced a $150 million strategic research initiative in which $100 million is distributed to Schools in proportion to their federal research funding expenditures and $50 million supports cross-disciplinary internal award programs, covering neuroscience, immunology, inflammation and infectious diseases, energy and climate (two new Salata Institute cross-faculty clusters at up to $600,000 per year for three years), the Frontiers innovation fund, and expanded computing infrastructure.
arXiv The work reframes agent skill synthesis as a dynamic navigation problem over past experience and proposes ExpVoyager, in which a skill curator uses a Navigable Interface to inspect raw trajectories across views and resolutions while a Navigation State tracks reusable procedural knowledge and open questions, producing task-specific skills for a frozen executor; across ALFWorld, WebShop, and ScienceWorld it improves task success and reduces execution steps over no-skill and existing skill-based baselines, with gains that grow as the experience space scales.
The work reframes agent skill synthesis as a dynamic navigation problem over past experience and proposes ExpVoyager, in which a skill curator uses a Navigable Interface to inspect raw trajectories across views and resolutions while a Navigation State tracks reusable procedural knowledge and open questions, producing task-specific skills for a frozen executor; across ALFWorld, WebShop, and ScienceWorld it improves task success and reduces execution steps over no-skill and existing skill-based baselines, with gains that grow as the experience space scales.
The work reframes agent skill synthesis as a dynamic navigation problem over past experience and proposes ExpVoyager, in which a skill curator uses a Navigable Interface to inspect raw trajectories across views and resolutions while a Navigation State tracks reusable procedural knowledge and open questions, producing task-specific skills for a frozen executor; across ALFWorld, WebShop, and ScienceWorld it improves task success and reduces execution steps over no-skill and existing skill-based baselines, with gains that grow as the experience space scales.
The work reframes agent skill synthesis as a dynamic navigation problem over past experience and proposes ExpVoyager, in which a skill curator uses a Navigable Interface to inspect raw trajectories across views and resolutions while a Navigation State tracks reusable procedural knowledge and open questions, producing task-specific skills for a frozen executor; across ALFWorld, WebShop, and ScienceWorld it improves task success and reduces execution steps over no-skill and existing skill-based baselines, with gains that grow as the experience space scales.
arXiv The work introduces SMAT (Simple MAT), which from a single expert's perspective describes common model-merging methods through three operations—Scale (reweighting its own update), Mask (removing selected coordinates) and Perturb (adding other experts' updates)—and during training samples scaling coefficients, masks and additive noise to simulate merged parameters while jointly optimizing the expert loss and the expected loss at simulated merged parameters, made efficient by periodic scheduling, kernel fusion and parameter storage switching with one forward and one backward pass per step; across four backbones (Llama-3.2-1B-Instruct, Llama-3.1-8B-Instruct, CLIP ViT-B/32 and ViT-L/14), SMAT improves the mean score over five merging methods by 1.07–2.
The work introduces SMAT (Simple MAT), which from a single expert's perspective describes common model-merging methods through three operations—Scale (reweighting its own update), Mask (removing selected coordinates) and Perturb (adding other experts' updates)—and during training samples scaling coefficients, masks and additive noise to simulate merged parameters while jointly optimizing the expert loss and the expected loss at simulated merged parameters, made efficient by periodic scheduling, kernel fusion and parameter storage switching with one forward and one backward pass per step; across four backbones (Llama-3.2-1B-Instruct, Llama-3.1-8B-Instruct, CLIP ViT-B/32 and ViT-L/14), SMAT improves the mean score over five merging methods by 1.07–2.
The work introduces SMAT (Simple MAT), which from a single expert's perspective describes common model-merging methods through three operations—Scale (reweighting its own update), Mask (removing selected coordinates) and Perturb (adding other experts' updates)—and during training samples scaling coefficients, masks and additive noise to simulate merged parameters while jointly optimizing the expert loss and the expected loss at simulated merged parameters, made efficient by periodic scheduling, kernel fusion and parameter storage switching with one forward and one backward pass per step; across four backbones (Llama-3.2-1B-Instruct, Llama-3.1-8B-Instruct, CLIP ViT-B/32 and ViT-L/14), SMAT improves the mean score over five merging methods by 1.07–2.
The work introduces SMAT (Simple MAT), which from a single expert's perspective describes common model-merging methods through three operations—Scale (reweighting its own update), Mask (removing selected coordinates) and Perturb (adding other experts' updates)—and during training samples scaling coefficients, masks and additive noise to simulate merged parameters while jointly optimizing the expert loss and the expected loss at simulated merged parameters, made efficient by periodic scheduling, kernel fusion and parameter storage switching with one forward and one backward pass per step; across four backbones (Llama-3.2-1B-Instruct, Llama-3.1-8B-Instruct, CLIP ViT-B/32 and ViT-L/14), SMAT improves the mean score over five merging methods by 1.07–2.
arXiv SkillDRE is a fully automated framework that, given a benign task and its associated skills, autonomously constructs and validates a task-conditioned malicious objective and a verifiable judge rule, then holds both fixed while alternating scanner-guided evolution with runtime-defense-guided refinement, forming a closed loop between the pre-execution and runtime stages; on SkillsBench across four victim models it achieves an average attack success rate of 45.28%, exceeding the strongest baseline by 40.3%, while its final submitted skills receive no SkillScan findings and largely preserve benign-task performance.
SkillDRE is a fully automated framework that, given a benign task and its associated skills, autonomously constructs and validates a task-conditioned malicious objective and a verifiable judge rule, then holds both fixed while alternating scanner-guided evolution with runtime-defense-guided refinement, forming a closed loop between the pre-execution and runtime stages; on SkillsBench across four victim models it achieves an average attack success rate of 45.28%, exceeding the strongest baseline by 40.3%, while its final submitted skills receive no SkillScan findings and largely preserve benign-task performance.
SkillDRE is a fully automated framework that, given a benign task and its associated skills, autonomously constructs and validates a task-conditioned malicious objective and a verifiable judge rule, then holds both fixed while alternating scanner-guided evolution with runtime-defense-guided refinement, forming a closed loop between the pre-execution and runtime stages; on SkillsBench across four victim models it achieves an average attack success rate of 45.28%, exceeding the strongest baseline by 40.3%, while its final submitted skills receive no SkillScan findings and largely preserve benign-task performance.
SkillDRE is a fully automated framework that, given a benign task and its associated skills, autonomously constructs and validates a task-conditioned malicious objective and a verifiable judge rule, then holds both fixed while alternating scanner-guided evolution with runtime-defense-guided refinement, forming a closed loop between the pre-execution and runtime stages; on SkillsBench across four victim models it achieves an average attack success rate of 45.28%, exceeding the strongest baseline by 40.3%, while its final submitted skills receive no SkillScan findings and largely preserve benign-task performance.
arXiv The work proposes PyroAdapt, a pretrain–retrieve–rank framework that pretrains a wildfire risk model on historical data, retrieves similar historical locations using unlabeled target-year covariates, and fine-tunes through same-day fire–nonfire cell-pair ranking losses (direct ranking, residual pairwise DPO, and selective ranking); over 666 0.25° California grid cells, the three ranking objectives raise daily average precision from 21.62% under continued focal fine-tuning to 24.35–24.57% and Top5% recall from 18.70% to 22.11–22.79%, selective ranking captures 344 additional positive cell-days under a fixed daily budget of 34 cells (5% of area), and rolling evaluations over Yosemite show the ranking gains persist under temporal distribution shift.
The work proposes PyroAdapt, a pretrain–retrieve–rank framework that pretrains a wildfire risk model on historical data, retrieves similar historical locations using unlabeled target-year covariates, and fine-tunes through same-day fire–nonfire cell-pair ranking losses (direct ranking, residual pairwise DPO, and selective ranking); over 666 0.25° California grid cells, the three ranking objectives raise daily average precision from 21.62% under continued focal fine-tuning to 24.35–24.57% and Top5% recall from 18.70% to 22.11–22.79%, selective ranking captures 344 additional positive cell-days under a fixed daily budget of 34 cells (5% of area), and rolling evaluations over Yosemite show the ranking gains persist under temporal distribution shift.
The work proposes PyroAdapt, a pretrain–retrieve–rank framework that pretrains a wildfire risk model on historical data, retrieves similar historical locations using unlabeled target-year covariates, and fine-tunes through same-day fire–nonfire cell-pair ranking losses (direct ranking, residual pairwise DPO, and selective ranking); over 666 0.25° California grid cells, the three ranking objectives raise daily average precision from 21.62% under continued focal fine-tuning to 24.35–24.57% and Top5% recall from 18.70% to 22.11–22.79%, selective ranking captures 344 additional positive cell-days under a fixed daily budget of 34 cells (5% of area), and rolling evaluations over Yosemite show the ranking gains persist under temporal distribution shift.
The work proposes PyroAdapt, a pretrain–retrieve–rank framework that pretrains a wildfire risk model on historical data, retrieves similar historical locations using unlabeled target-year covariates, and fine-tunes through same-day fire–nonfire cell-pair ranking losses (direct ranking, residual pairwise DPO, and selective ranking); over 666 0.25° California grid cells, the three ranking objectives raise daily average precision from 21.62% under continued focal fine-tuning to 24.35–24.57% and Top5% recall from 18.70% to 22.11–22.79%, selective ranking captures 344 additional positive cell-days under a fixed daily budget of 34 cells (5% of area), and rolling evaluations over Yosemite show the ranking gains persist under temporal distribution shift.
arXiv The work introduces BaRe-Mem, an online Bayesian reliability memory that estimates advisor reliability conditioned on the central model's internal belief representations, uses those estimates to modulate attention and to decide whether to consult; across nine benchmarks and six central models it is more robust to misleading advisor information than debate and majority voting, stays above autonomous reasoning on the more challenging tasks across all tested misleading levels, and in MuSiQue agent-team routing identifies capable workers earlier than routing by historical success counts.
The work introduces BaRe-Mem, an online Bayesian reliability memory that estimates advisor reliability conditioned on the central model's internal belief representations, uses those estimates to modulate attention and to decide whether to consult; across nine benchmarks and six central models it is more robust to misleading advisor information than debate and majority voting, stays above autonomous reasoning on the more challenging tasks across all tested misleading levels, and in MuSiQue agent-team routing identifies capable workers earlier than routing by historical success counts.
The work introduces BaRe-Mem, an online Bayesian reliability memory that estimates advisor reliability conditioned on the central model's internal belief representations, uses those estimates to modulate attention and to decide whether to consult; across nine benchmarks and six central models it is more robust to misleading advisor information than debate and majority voting, stays above autonomous reasoning on the more challenging tasks across all tested misleading levels, and in MuSiQue agent-team routing identifies capable workers earlier than routing by historical success counts.
The work introduces BaRe-Mem, an online Bayesian reliability memory that estimates advisor reliability conditioned on the central model's internal belief representations, uses those estimates to modulate attention and to decide whether to consult; across nine benchmarks and six central models it is more robust to misleading advisor information than debate and majority voting, stays above autonomous reasoning on the more challenging tasks across all tested misleading levels, and in MuSiQue agent-team routing identifies capable workers earlier than routing by historical success counts.
arXiv The study unifies EEG self-supervised masking strategies into a three-parameter framework of spatial radius, temporal length and mask ratio, trains 58 models under a fixed REVE-Small backbone and corpus across MAE and JEPA, evaluates them with a linear probe on the 12 OpenEEGBench datasets, and finds that both frameworks agree on a moderate spatial radius with short temporal blocks as the optimum while identifying a JEPA-specific bias-inflation collapse at full-channel masking.
The study unifies EEG self-supervised masking strategies into a three-parameter framework of spatial radius, temporal length and mask ratio, trains 58 models under a fixed REVE-Small backbone and corpus across MAE and JEPA, evaluates them with a linear probe on the 12 OpenEEGBench datasets, and finds that both frameworks agree on a moderate spatial radius with short temporal blocks as the optimum while identifying a JEPA-specific bias-inflation collapse at full-channel masking.
The study unifies EEG self-supervised masking strategies into a three-parameter framework of spatial radius, temporal length and mask ratio, trains 58 models under a fixed REVE-Small backbone and corpus across MAE and JEPA, evaluates them with a linear probe on the 12 OpenEEGBench datasets, and finds that both frameworks agree on a moderate spatial radius with short temporal blocks as the optimum while identifying a JEPA-specific bias-inflation collapse at full-channel masking.
The study unifies EEG self-supervised masking strategies into a three-parameter framework of spatial radius, temporal length and mask ratio, trains 58 models under a fixed REVE-Small backbone and corpus across MAE and JEPA, evaluates them with a linear probe on the 12 OpenEEGBench datasets, and finds that both frameworks agree on a moderate spatial radius with short temporal blocks as the optimum while identifying a JEPA-specific bias-inflation collapse at full-channel masking.
Science News Science News reporter Laura Sanders hosts a new season of The Deep End called Talk to Me, six episodes asking whether a lost voice can be found again, told through people who use AI-cloned voices, a brain implant that pulls words from thoughts, and a synthetic voice for singing, and through what those experiences reveal about why voices matter.
Science News reporter Laura Sanders hosts a new season of The Deep End called Talk to Me, six episodes asking whether a lost voice can be found again, told through people who use AI-cloned voices, a brain implant that pulls words from thoughts, and a synthetic voice for singing, and through what those experiences reveal about why voices matter.
Science News reporter Laura Sanders hosts a new season of The Deep End called Talk to Me, six episodes asking whether a lost voice can be found again, told through people who use AI-cloned voices, a brain implant that pulls words from thoughts, and a synthetic voice for singing, and through what those experiences reveal about why voices matter.
Science News reporter Laura Sanders hosts a new season of The Deep End called Talk to Me, six episodes asking whether a lost voice can be found again, told through people who use AI-cloned voices, a brain implant that pulls words from thoughts, and a synthetic voice for singing, and through what those experiences reveal about why voices matter.
arXiv The study introduces NameTrace, measures direct lexical support for names across nearly half a million first names and 12 LLM-associated tokenizers, and, on atomic versus short-fragmented names matched within the same race/ethnicity–gender strata, finds systematically higher task-aligned concept accessibility for atomic names across fellowship, hiring, clinical assessment, and lending; the differences persist in all eight strata, transfer to unseen names, and hidden-state interventions along the measured task directions shift later constrained choices.
The study introduces NameTrace, measures direct lexical support for names across nearly half a million first names and 12 LLM-associated tokenizers, and, on atomic versus short-fragmented names matched within the same race/ethnicity–gender strata, finds systematically higher task-aligned concept accessibility for atomic names across fellowship, hiring, clinical assessment, and lending; the differences persist in all eight strata, transfer to unseen names, and hidden-state interventions along the measured task directions shift later constrained choices.
The study introduces NameTrace, measures direct lexical support for names across nearly half a million first names and 12 LLM-associated tokenizers, and, on atomic versus short-fragmented names matched within the same race/ethnicity–gender strata, finds systematically higher task-aligned concept accessibility for atomic names across fellowship, hiring, clinical assessment, and lending; the differences persist in all eight strata, transfer to unseen names, and hidden-state interventions along the measured task directions shift later constrained choices.
The study introduces NameTrace, measures direct lexical support for names across nearly half a million first names and 12 LLM-associated tokenizers, and, on atomic versus short-fragmented names matched within the same race/ethnicity–gender strata, finds systematically higher task-aligned concept accessibility for atomic names across fellowship, hiring, clinical assessment, and lending; the differences persist in all eight strata, transfer to unseen names, and hidden-state interventions along the measured task directions shift later constrained choices.
arXiv The work introduces SpatialSpeak, a two-stage framework that first jointly trains local marked-point 3D coordinates and global object-center reconstruction as question answering (QA-RP), then supervises the model with spatial chain-of-thought plus visual compensation (CoT-VC) to express and use geometric estimates; on ReVSI it raises the gain from CoT-VC from 2.6 to 6.9 points and, with a 4B backbone, reaches the best results among compared methods on ReVSI (62.8), VSI-Bench (73.0), and SPAR-Bench (76.0).
The work introduces SpatialSpeak, a two-stage framework that first jointly trains local marked-point 3D coordinates and global object-center reconstruction as question answering (QA-RP), then supervises the model with spatial chain-of-thought plus visual compensation (CoT-VC) to express and use geometric estimates; on ReVSI it raises the gain from CoT-VC from 2.6 to 6.9 points and, with a 4B backbone, reaches the best results among compared methods on ReVSI (62.8), VSI-Bench (73.0), and SPAR-Bench (76.0).
The work introduces SpatialSpeak, a two-stage framework that first jointly trains local marked-point 3D coordinates and global object-center reconstruction as question answering (QA-RP), then supervises the model with spatial chain-of-thought plus visual compensation (CoT-VC) to express and use geometric estimates; on ReVSI it raises the gain from CoT-VC from 2.6 to 6.9 points and, with a 4B backbone, reaches the best results among compared methods on ReVSI (62.8), VSI-Bench (73.0), and SPAR-Bench (76.0).
The work introduces SpatialSpeak, a two-stage framework that first jointly trains local marked-point 3D coordinates and global object-center reconstruction as question answering (QA-RP), then supervises the model with spatial chain-of-thought plus visual compensation (CoT-VC) to express and use geometric estimates; on ReVSI it raises the gain from CoT-VC from 2.6 to 6.9 points and, with a 4B backbone, reaches the best results among compared methods on ReVSI (62.8), VSI-Bench (73.0), and SPAR-Bench (76.0).
arXiv The work introduces Allspark, a training and inference framework in which a weak teacher and a frozen copy of the same model alternate reasoning segments during training while only the teacher is updated from weak-model rollouts, and at inference a stronger frozen student replaces the frozen partner; across Qwen3-1.7B/4B math and reasoning tasks and Inkling-Small teacher evaluations on ARC-AGI-2 with Inkling, Kimi-K2.6, and Nemotron-3-Ultra students, accuracy gains are observed along with an explicit accuracy-token tradeoff.
The work introduces Allspark, a training and inference framework in which a weak teacher and a frozen copy of the same model alternate reasoning segments during training while only the teacher is updated from weak-model rollouts, and at inference a stronger frozen student replaces the frozen partner; across Qwen3-1.7B/4B math and reasoning tasks and Inkling-Small teacher evaluations on ARC-AGI-2 with Inkling, Kimi-K2.6, and Nemotron-3-Ultra students, accuracy gains are observed along with an explicit accuracy-token tradeoff.
The work introduces Allspark, a training and inference framework in which a weak teacher and a frozen copy of the same model alternate reasoning segments during training while only the teacher is updated from weak-model rollouts, and at inference a stronger frozen student replaces the frozen partner; across Qwen3-1.7B/4B math and reasoning tasks and Inkling-Small teacher evaluations on ARC-AGI-2 with Inkling, Kimi-K2.6, and Nemotron-3-Ultra students, accuracy gains are observed along with an explicit accuracy-token tradeoff.
The work introduces Allspark, a training and inference framework in which a weak teacher and a frozen copy of the same model alternate reasoning segments during training while only the teacher is updated from weak-model rollouts, and at inference a stronger frozen student replaces the frozen partner; across Qwen3-1.7B/4B math and reasoning tasks and Inkling-Small teacher evaluations on ARC-AGI-2 with Inkling, Kimi-K2.6, and Nemotron-3-Ultra students, accuracy gains are observed along with an explicit accuracy-token tradeoff.
arXiv Skill2Env is a capability-oriented framework that curates skills from public skill libraries, uses reusable difficulty patterns and task blueprints to jointly synthesize task instructions, execution substrates, workspaces, and rubric-based evaluators, and iteratively hardens tasks using solver execution evidence, yielding 2,963 executable tasks; supervised fine-tuning of Qwen3.6-35B-A3B on 1.5K high-scoring trajectories from these environments raises the unweighted average across seven agent benchmarks from 36.6 to 45.0.
Skill2Env is a capability-oriented framework that curates skills from public skill libraries, uses reusable difficulty patterns and task blueprints to jointly synthesize task instructions, execution substrates, workspaces, and rubric-based evaluators, and iteratively hardens tasks using solver execution evidence, yielding 2,963 executable tasks; supervised fine-tuning of Qwen3.6-35B-A3B on 1.5K high-scoring trajectories from these environments raises the unweighted average across seven agent benchmarks from 36.6 to 45.0.
Skill2Env is a capability-oriented framework that curates skills from public skill libraries, uses reusable difficulty patterns and task blueprints to jointly synthesize task instructions, execution substrates, workspaces, and rubric-based evaluators, and iteratively hardens tasks using solver execution evidence, yielding 2,963 executable tasks; supervised fine-tuning of Qwen3.6-35B-A3B on 1.5K high-scoring trajectories from these environments raises the unweighted average across seven agent benchmarks from 36.6 to 45.0.
Skill2Env is a capability-oriented framework that curates skills from public skill libraries, uses reusable difficulty patterns and task blueprints to jointly synthesize task instructions, execution substrates, workspaces, and rubric-based evaluators, and iteratively hardens tasks using solver execution evidence, yielding 2,963 executable tasks; supervised fine-tuning of Qwen3.6-35B-A3B on 1.5K high-scoring trajectories from these environments raises the unweighted average across seven agent benchmarks from 36.6 to 45.0.
arXiv The work proposes AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level, nine-dimension rubric hierarchy that assigns each rollout one of three hint forms—rubrics, a corrective reflection, or a sibling set—matched to its reward, and distills the hint's effect into a token-level advantage combined with the group-relative outcome advantage; across ten benchmarks spanning RAG, deep research, and setwise evaluation it attains the best overall performance while reducing the agent's retrieval calls.
The work proposes AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level, nine-dimension rubric hierarchy that assigns each rollout one of three hint forms—rubrics, a corrective reflection, or a sibling set—matched to its reward, and distills the hint's effect into a token-level advantage combined with the group-relative outcome advantage; across ten benchmarks spanning RAG, deep research, and setwise evaluation it attains the best overall performance while reducing the agent's retrieval calls.
The work proposes AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level, nine-dimension rubric hierarchy that assigns each rollout one of three hint forms—rubrics, a corrective reflection, or a sibling set—matched to its reward, and distills the hint's effect into a token-level advantage combined with the group-relative outcome advantage; across ten benchmarks spanning RAG, deep research, and setwise evaluation it attains the best overall performance while reducing the agent's retrieval calls.
The work proposes AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level, nine-dimension rubric hierarchy that assigns each rollout one of three hint forms—rubrics, a corrective reflection, or a sibling set—matched to its reward, and distills the hint's effect into a token-level advantage combined with the group-relative outcome advantage; across ten benchmarks spanning RAG, deep research, and setwise evaluation it attains the best overall performance while reducing the agent's retrieval calls.
arXiv VQS has a model first parse an image into a structured record such as a scene graph, chart table, or diagram graph, then uses fixed program templates to write questions and compute their answers from that record while the model only confirms the atomic facts the program reads one at a time; human raters find 94.4% of VQS answers correct versus 76.4% for majority voting and 82.2% for a model judge, and across ten benchmarks it improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales, reaching 3.84 points at 2B after three training cycles.
VQS has a model first parse an image into a structured record such as a scene graph, chart table, or diagram graph, then uses fixed program templates to write questions and compute their answers from that record while the model only confirms the atomic facts the program reads one at a time; human raters find 94.4% of VQS answers correct versus 76.4% for majority voting and 82.2% for a model judge, and across ten benchmarks it improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales, reaching 3.84 points at 2B after three training cycles.
VQS has a model first parse an image into a structured record such as a scene graph, chart table, or diagram graph, then uses fixed program templates to write questions and compute their answers from that record while the model only confirms the atomic facts the program reads one at a time; human raters find 94.4% of VQS answers correct versus 76.4% for majority voting and 82.2% for a model judge, and across ten benchmarks it improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales, reaching 3.84 points at 2B after three training cycles.
VQS has a model first parse an image into a structured record such as a scene graph, chart table, or diagram graph, then uses fixed program templates to write questions and compute their answers from that record while the model only confirms the atomic facts the program reads one at a time; human raters find 94.4% of VQS answers correct versus 76.4% for majority voting and 82.2% for a model judge, and across ten benchmarks it improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales, reaching 3.84 points at 2B after three training cycles.
arXiv The work proposes TT-VidT, which pairs a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer and trains it by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens; under a matched recipe at roughly 170M-190M encoder scale on 1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, a 24-cell architecture-objective sweep shows that only TT3D with Diff Compression enters the strongest motion-sensitive regime, and in the final comparison it leads Jester, Something-Something V2, ARID and Diving48 fine-tuning simultaneously while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA 2.
The work proposes TT-VidT, which pairs a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer and trains it by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens; under a matched recipe at roughly 170M-190M encoder scale on 1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, a 24-cell architecture-objective sweep shows that only TT3D with Diff Compression enters the strongest motion-sensitive regime, and in the final comparison it leads Jester, Something-Something V2, ARID and Diving48 fine-tuning simultaneously while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA 2.
The work proposes TT-VidT, which pairs a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer and trains it by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens; under a matched recipe at roughly 170M-190M encoder scale on 1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, a 24-cell architecture-objective sweep shows that only TT3D with Diff Compression enters the strongest motion-sensitive regime, and in the final comparison it leads Jester, Something-Something V2, ARID and Diving48 fine-tuning simultaneously while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA 2.
The work proposes TT-VidT, which pairs a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer and trains it by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens; under a matched recipe at roughly 170M-190M encoder scale on 1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, a 24-cell architecture-objective sweep shows that only TT3D with Diff Compression enters the strongest motion-sensitive regime, and in the final comparison it leads Jester, Something-Something V2, ARID and Diving48 fine-tuning simultaneously while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA 2.
Mistral AI Mistral announced a new German hub in Munich housing research teams dedicated to Physics AI and Industrial AI plus applied engineers serving enterprise partners, and disclosed that it acquired Emmi AI (bringing in more than 30 physicists, researchers and engineers), is working with BMW on crash simulations and engineering AI and with Siemens Energy on industrial AI applications, and has formed a research partnership with the Technical University Munich (TUM) to use TUM's wind tunnel facilities with Prof. Dr. Nikolaus A.
Mistral announced a new German hub in Munich housing research teams dedicated to Physics AI and Industrial AI plus applied engineers serving enterprise partners, and disclosed that it acquired Emmi AI (bringing in more than 30 physicists, researchers and engineers), is working with BMW on crash simulations and engineering AI and with Siemens Energy on industrial AI applications, and has formed a research partnership with the Technical University Munich (TUM) to use TUM's wind tunnel facilities with Prof. Dr. Nikolaus A.
Mistral announced a new German hub in Munich housing research teams dedicated to Physics AI and Industrial AI plus applied engineers serving enterprise partners, and disclosed that it acquired Emmi AI (bringing in more than 30 physicists, researchers and engineers), is working with BMW on crash simulations and engineering AI and with Siemens Energy on industrial AI applications, and has formed a research partnership with the Technical University Munich (TUM) to use TUM's wind tunnel facilities with Prof. Dr. Nikolaus A.
Mistral announced a new German hub in Munich housing research teams dedicated to Physics AI and Industrial AI plus applied engineers serving enterprise partners, and disclosed that it acquired Emmi AI (bringing in more than 30 physicists, researchers and engineers), is working with BMW on crash simulations and engineering AI and with Siemens Energy on industrial AI applications, and has formed a research partnership with the Technical University Munich (TUM) to use TUM's wind tunnel facilities with Prof. Dr. Nikolaus A.