arXiv Using low-rank interventions, the work finds that visual-to-text information flow in multimodal LLMs is concentrated in a low-dimensional subspace and that layer-specific visual states are highly predictable by lightweight MLPs; it therefore proposes δ-Vision, which freezes the language-model backbone and uses low-rank adapters to construct each layer's visual memory while preserving all visual tokens for text retrieval, reaching an average score of 74.4 on Qwen3-VL-4B (92.9% of uncompressed performance), 9.1 points above the strongest 5%-retention pruning baseline DivPrune, and requiring only 17.90% of the uncompressed model's FLOPs on Video-MME.
Using low-rank interventions, the work finds that visual-to-text information flow in multimodal LLMs is concentrated in a low-dimensional subspace and that layer-specific visual states are highly predictable by lightweight MLPs; it therefore proposes δ-Vision, which freezes the language-model backbone and uses low-rank adapters to construct each layer's visual memory while preserving all visual tokens for text retrieval, reaching an average score of 74.4 on Qwen3-VL-4B (92.9% of uncompressed performance), 9.1 points above the strongest 5%-retention pruning baseline DivPrune, and requiring only 17.90% of the uncompressed model's FLOPs on Video-MME.
Using low-rank interventions, the work finds that visual-to-text information flow in multimodal LLMs is concentrated in a low-dimensional subspace and that layer-specific visual states are highly predictable by lightweight MLPs; it therefore proposes δ-Vision, which freezes the language-model backbone and uses low-rank adapters to construct each layer's visual memory while preserving all visual tokens for text retrieval, reaching an average score of 74.4 on Qwen3-VL-4B (92.9% of uncompressed performance), 9.1 points above the strongest 5%-retention pruning baseline DivPrune, and requiring only 17.90% of the uncompressed model's FLOPs on Video-MME.
Using low-rank interventions, the work finds that visual-to-text information flow in multimodal LLMs is concentrated in a low-dimensional subspace and that layer-specific visual states are highly predictable by lightweight MLPs; it therefore proposes δ-Vision, which freezes the language-model backbone and uses low-rank adapters to construct each layer's visual memory while preserving all visual tokens for text retrieval, reaching an average score of 74.4 on Qwen3-VL-4B (92.9% of uncompressed performance), 9.1 points above the strongest 5%-retention pruning baseline DivPrune, and requiring only 17.90% of the uncompressed model's FLOPs on Video-MME.
arXiv The work proposes DISCO, which separates long-context processing into local evidence grounding executed in parallel by lightweight Worker LLMs and global reasoning handled by a central Driver LLM whose dynamic DAG planning is optimized with GRPO reinforcement learning; with Qwen3-8B it holds 78.4% accuracy on RULER-QA at 1M tokens (where the RAG baseline collapses to 10.9%), improves up to 9.8 points over full-context baselines on LongBench v2, and matches Gemini-3-Pro-Preview while reducing inference cost by over 80%.
The work proposes DISCO, which separates long-context processing into local evidence grounding executed in parallel by lightweight Worker LLMs and global reasoning handled by a central Driver LLM whose dynamic DAG planning is optimized with GRPO reinforcement learning; with Qwen3-8B it holds 78.4% accuracy on RULER-QA at 1M tokens (where the RAG baseline collapses to 10.9%), improves up to 9.8 points over full-context baselines on LongBench v2, and matches Gemini-3-Pro-Preview while reducing inference cost by over 80%.
The work proposes DISCO, which separates long-context processing into local evidence grounding executed in parallel by lightweight Worker LLMs and global reasoning handled by a central Driver LLM whose dynamic DAG planning is optimized with GRPO reinforcement learning; with Qwen3-8B it holds 78.4% accuracy on RULER-QA at 1M tokens (where the RAG baseline collapses to 10.9%), improves up to 9.8 points over full-context baselines on LongBench v2, and matches Gemini-3-Pro-Preview while reducing inference cost by over 80%.
The work proposes DISCO, which separates long-context processing into local evidence grounding executed in parallel by lightweight Worker LLMs and global reasoning handled by a central Driver LLM whose dynamic DAG planning is optimized with GRPO reinforcement learning; with Qwen3-8B it holds 78.4% accuracy on RULER-QA at 1M tokens (where the RAG baseline collapses to 10.9%), improves up to 9.8 points over full-context baselines on LongBench v2, and matches Gemini-3-Pro-Preview while reducing inference cost by over 80%.
arXiv The work introduces a routing analysis toolkit and, across DeepSeekMoE, OLMoE and Qwen3-MoE, crosses source and merged router inputs and parameters and runs paired token- and task-level evaluations, finding that most expert reassignments are input-shift-induced rather than parameter-induced, that source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and that different expert selections can produce directionally similar mixture outputs; it therefore redefines routing failure as task loss recoverable under a specified routing intervention with non-routing parameters fixed, and proposes Selective Router Repair as a case study that finds no reliable evidence that source-specialist token-likelihood advantages identify beneficial local corr
The work introduces a routing analysis toolkit and, across DeepSeekMoE, OLMoE and Qwen3-MoE, crosses source and merged router inputs and parameters and runs paired token- and task-level evaluations, finding that most expert reassignments are input-shift-induced rather than parameter-induced, that source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and that different expert selections can produce directionally similar mixture outputs; it therefore redefines routing failure as task loss recoverable under a specified routing intervention with non-routing parameters fixed, and proposes Selective Router Repair as a case study that finds no reliable evidence that source-specialist token-likelihood advantages identify beneficial local corr
The work introduces a routing analysis toolkit and, across DeepSeekMoE, OLMoE and Qwen3-MoE, crosses source and merged router inputs and parameters and runs paired token- and task-level evaluations, finding that most expert reassignments are input-shift-induced rather than parameter-induced, that source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and that different expert selections can produce directionally similar mixture outputs; it therefore redefines routing failure as task loss recoverable under a specified routing intervention with non-routing parameters fixed, and proposes Selective Router Repair as a case study that finds no reliable evidence that source-specialist token-likelihood advantages identify beneficial local corr
The work introduces a routing analysis toolkit and, across DeepSeekMoE, OLMoE and Qwen3-MoE, crosses source and merged router inputs and parameters and runs paired token- and task-level evaluations, finding that most expert reassignments are input-shift-induced rather than parameter-induced, that source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and that different expert selections can produce directionally similar mixture outputs; it therefore redefines routing failure as task loss recoverable under a specified routing intervention with non-routing parameters fixed, and proposes Selective Router Repair as a case study that finds no reliable evidence that source-specialist token-likelihood advantages identify beneficial local corr
arXiv Keeping the 6.5M-parameter v0.3 architecture unchanged, NanoForecast v0.5 retrains after fixing loss-scope handling, tensor shape alignment, and augmentation coverage, cutting overall MASE from 3.030 to 1.704 (a 43.8% reduction) under one fixed protocol on the same data and compute budget, and beating 200M-parameter TimesFM on ETTh1, ETTh2, ETTm1, and exchange rate while beating 15M+-parameter PatchTST on all three ETT sets.
Keeping the 6.5M-parameter v0.3 architecture unchanged, NanoForecast v0.5 retrains after fixing loss-scope handling, tensor shape alignment, and augmentation coverage, cutting overall MASE from 3.030 to 1.704 (a 43.8% reduction) under one fixed protocol on the same data and compute budget, and beating 200M-parameter TimesFM on ETTh1, ETTh2, ETTm1, and exchange rate while beating 15M+-parameter PatchTST on all three ETT sets.
Keeping the 6.5M-parameter v0.3 architecture unchanged, NanoForecast v0.5 retrains after fixing loss-scope handling, tensor shape alignment, and augmentation coverage, cutting overall MASE from 3.030 to 1.704 (a 43.8% reduction) under one fixed protocol on the same data and compute budget, and beating 200M-parameter TimesFM on ETTh1, ETTh2, ETTm1, and exchange rate while beating 15M+-parameter PatchTST on all three ETT sets.
Keeping the 6.5M-parameter v0.3 architecture unchanged, NanoForecast v0.5 retrains after fixing loss-scope handling, tensor shape alignment, and augmentation coverage, cutting overall MASE from 3.030 to 1.704 (a 43.8% reduction) under one fixed protocol on the same data and compute budget, and beating 200M-parameter TimesFM on ETTh1, ETTh2, ETTm1, and exchange rate while beating 15M+-parameter PatchTST on all three ETT sets.
arXiv The work introduces ColNanoVDR, which uses OTW, an entropic optimal-transport objective with learned token weights, to distill the query encoder of multi-vector visual document retrieval teachers into 149M text-only students without encoding or reading any page during training, retaining about 95% of the teachers' NDCG@5 on ViDoRe v1-v3, encoding queries up to 26x faster on a single CPU thread, and matching score distillation under identical training while reading 12.6x less cached teacher data.
The work introduces ColNanoVDR, which uses OTW, an entropic optimal-transport objective with learned token weights, to distill the query encoder of multi-vector visual document retrieval teachers into 149M text-only students without encoding or reading any page during training, retaining about 95% of the teachers' NDCG@5 on ViDoRe v1-v3, encoding queries up to 26x faster on a single CPU thread, and matching score distillation under identical training while reading 12.6x less cached teacher data.
The work introduces ColNanoVDR, which uses OTW, an entropic optimal-transport objective with learned token weights, to distill the query encoder of multi-vector visual document retrieval teachers into 149M text-only students without encoding or reading any page during training, retaining about 95% of the teachers' NDCG@5 on ViDoRe v1-v3, encoding queries up to 26x faster on a single CPU thread, and matching score distillation under identical training while reading 12.6x less cached teacher data.
The work introduces ColNanoVDR, which uses OTW, an entropic optimal-transport objective with learned token weights, to distill the query encoder of multi-vector visual document retrieval teachers into 149M text-only students without encoding or reading any page during training, retaining about 95% of the teachers' NDCG@5 on ViDoRe v1-v3, encoding queries up to 26x faster on a single CPU thread, and matching score distillation under identical training while reading 12.6x less cached teacher data.
arXiv The work introduces Pruned CTC, which exploits the fact that every valid CTC alignment uses only target tokens and blank to restrict alignment computation to that batch-level subset while retaining full-vocabulary normalization, and proves this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and first-order gradients; with a Zipformer-M encoder and a 180K vocabulary it cuts full-step memory by 5.1× at only 17% step-time overhead and matches standard CTC accuracy across three corpora; built on it, LLM-CTC adapts pretrained causal LLMs to non-autoregressive ASR over native vocabularies, staying within 7% relative WER of LLM-CE across six Qwen3 sizes from 0.
The work introduces Pruned CTC, which exploits the fact that every valid CTC alignment uses only target tokens and blank to restrict alignment computation to that batch-level subset while retaining full-vocabulary normalization, and proves this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and first-order gradients; with a Zipformer-M encoder and a 180K vocabulary it cuts full-step memory by 5.1× at only 17% step-time overhead and matches standard CTC accuracy across three corpora; built on it, LLM-CTC adapts pretrained causal LLMs to non-autoregressive ASR over native vocabularies, staying within 7% relative WER of LLM-CE across six Qwen3 sizes from 0.
The work introduces Pruned CTC, which exploits the fact that every valid CTC alignment uses only target tokens and blank to restrict alignment computation to that batch-level subset while retaining full-vocabulary normalization, and proves this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and first-order gradients; with a Zipformer-M encoder and a 180K vocabulary it cuts full-step memory by 5.1× at only 17% step-time overhead and matches standard CTC accuracy across three corpora; built on it, LLM-CTC adapts pretrained causal LLMs to non-autoregressive ASR over native vocabularies, staying within 7% relative WER of LLM-CE across six Qwen3 sizes from 0.
The work introduces Pruned CTC, which exploits the fact that every valid CTC alignment uses only target tokens and blank to restrict alignment computation to that batch-level subset while retaining full-vocabulary normalization, and proves this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and first-order gradients; with a Zipformer-M encoder and a 180K vocabulary it cuts full-step memory by 5.1× at only 17% step-time overhead and matches standard CTC accuracy across three corpora; built on it, LLM-CTC adapts pretrained causal LLMs to non-autoregressive ASR over native vocabularies, staying within 7% relative WER of LLM-CE across six Qwen3 sizes from 0.
arXiv The OmniAI Group of ZJU ACES Lab introduces EMem-Bench, a long-horizon embodied memory benchmark of 2,554 executable episodes spanning Passive Observation, Dynamic Tracking, Interaction Failure, and Experience Generalization, together with an external memory system EMem and an 8B policy EMem-8B; evaluating 16 open-source and proprietary MLLMs plus representative multimodal memory systems shows the strongest proprietary model Gemini-3-Flash reaches only 64.2% average success rate and most open-source models fall below 45%, while EMem achieves the best overall performance among matched backbones, improving Mistral-Small-3.1-24B by 16.4 points and GPT-5.4-mini by 23.6 points, with EMem-8B improving over Qwen3-VL-8B by 20.6 points.
The OmniAI Group of ZJU ACES Lab introduces EMem-Bench, a long-horizon embodied memory benchmark of 2,554 executable episodes spanning Passive Observation, Dynamic Tracking, Interaction Failure, and Experience Generalization, together with an external memory system EMem and an 8B policy EMem-8B; evaluating 16 open-source and proprietary MLLMs plus representative multimodal memory systems shows the strongest proprietary model Gemini-3-Flash reaches only 64.2% average success rate and most open-source models fall below 45%, while EMem achieves the best overall performance among matched backbones, improving Mistral-Small-3.1-24B by 16.4 points and GPT-5.4-mini by 23.6 points, with EMem-8B improving over Qwen3-VL-8B by 20.6 points.
The OmniAI Group of ZJU ACES Lab introduces EMem-Bench, a long-horizon embodied memory benchmark of 2,554 executable episodes spanning Passive Observation, Dynamic Tracking, Interaction Failure, and Experience Generalization, together with an external memory system EMem and an 8B policy EMem-8B; evaluating 16 open-source and proprietary MLLMs plus representative multimodal memory systems shows the strongest proprietary model Gemini-3-Flash reaches only 64.2% average success rate and most open-source models fall below 45%, while EMem achieves the best overall performance among matched backbones, improving Mistral-Small-3.1-24B by 16.4 points and GPT-5.4-mini by 23.6 points, with EMem-8B improving over Qwen3-VL-8B by 20.6 points.
The OmniAI Group of ZJU ACES Lab introduces EMem-Bench, a long-horizon embodied memory benchmark of 2,554 executable episodes spanning Passive Observation, Dynamic Tracking, Interaction Failure, and Experience Generalization, together with an external memory system EMem and an 8B policy EMem-8B; evaluating 16 open-source and proprietary MLLMs plus representative multimodal memory systems shows the strongest proprietary model Gemini-3-Flash reaches only 64.2% average success rate and most open-source models fall below 45%, while EMem achieves the best overall performance among matched backbones, improving Mistral-Small-3.1-24B by 16.4 points and GPT-5.4-mini by 23.6 points, with EMem-8B improving over Qwen3-VL-8B by 20.6 points.
arXiv The work studies training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and introduces calibrated importance sampling (CIS): the mismatch is characterized as an additive displacement in log-odds, and large positive displacements are truncated at a single constant threshold, which maps back to an importance-ratio cap that tightens as token confidence increases; across three mixture-of-experts models and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average on all three models among the evaluated baselines, and diagnostic analyses show it places less truncation bias on low-confidence tokens than truncat
The work studies training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and introduces calibrated importance sampling (CIS): the mismatch is characterized as an additive displacement in log-odds, and large positive displacements are truncated at a single constant threshold, which maps back to an importance-ratio cap that tightens as token confidence increases; across three mixture-of-experts models and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average on all three models among the evaluated baselines, and diagnostic analyses show it places less truncation bias on low-confidence tokens than truncat
The work studies training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and introduces calibrated importance sampling (CIS): the mismatch is characterized as an additive displacement in log-odds, and large positive displacements are truncated at a single constant threshold, which maps back to an importance-ratio cap that tightens as token confidence increases; across three mixture-of-experts models and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average on all three models among the evaluated baselines, and diagnostic analyses show it places less truncation bias on low-confidence tokens than truncat
The work studies training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and introduces calibrated importance sampling (CIS): the mismatch is characterized as an additive displacement in log-odds, and large positive displacements are truncated at a single constant threshold, which maps back to an importance-ratio cap that tightens as token confidence increases; across three mixture-of-experts models and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average on all three models among the evaluated baselines, and diagnostic analyses show it places less truncation bias on low-confidence tokens than truncat
arXiv The work introduces Decoupled Credit Self-Distillation (DCSD), which uses belief-margin probing to set each reasoning step's credit direction and marginal information gain to quantify its contribution magnitude, leaving the privileged teacher to allocate credit only within steps; across 11 mathematical and multimodal reasoning benchmarks DCSD achieves the best overall scores against GRPO, OPSD, RLSD and RLCSD, improving overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning over base models, while correcting credit direction for about 6% of tokens and reducing token credit magnitude by roughly 1.5 times.
The work introduces Decoupled Credit Self-Distillation (DCSD), which uses belief-margin probing to set each reasoning step's credit direction and marginal information gain to quantify its contribution magnitude, leaving the privileged teacher to allocate credit only within steps; across 11 mathematical and multimodal reasoning benchmarks DCSD achieves the best overall scores against GRPO, OPSD, RLSD and RLCSD, improving overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning over base models, while correcting credit direction for about 6% of tokens and reducing token credit magnitude by roughly 1.5 times.
The work introduces Decoupled Credit Self-Distillation (DCSD), which uses belief-margin probing to set each reasoning step's credit direction and marginal information gain to quantify its contribution magnitude, leaving the privileged teacher to allocate credit only within steps; across 11 mathematical and multimodal reasoning benchmarks DCSD achieves the best overall scores against GRPO, OPSD, RLSD and RLCSD, improving overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning over base models, while correcting credit direction for about 6% of tokens and reducing token credit magnitude by roughly 1.5 times.
The work introduces Decoupled Credit Self-Distillation (DCSD), which uses belief-margin probing to set each reasoning step's credit direction and marginal information gain to quantify its contribution magnitude, leaving the privileged teacher to allocate credit only within steps; across 11 mathematical and multimodal reasoning benchmarks DCSD achieves the best overall scores against GRPO, OPSD, RLSD and RLCSD, improving overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning over base models, while correcting credit direction for about 6% of tokens and reducing token credit magnitude by roughly 1.5 times.
arXiv The work systematically analyzes Diffusion Transformer internal representations, finding a latent preference for early-layer feature reuse and symmetric layer pairing, and on that basis proposes a structured connectivity design that turns residual connections from passive summation into active retrieval: the DiT is split evenly along depth into an encoder and a decoder phase, encoder-side representations serve as differentiable residual sources, and each decoder sublayer adaptively fuses its mirrored encoder source through lightweight softmax routing; on ImageNet class-conditional generation, with under 0.1% additional parameters, the method lowers FID from 69.81 to 62.48 on DiT-S/2, from 44.21 to 36.33 on DiT-B/2, and from 18.85 to 15.
The work systematically analyzes Diffusion Transformer internal representations, finding a latent preference for early-layer feature reuse and symmetric layer pairing, and on that basis proposes a structured connectivity design that turns residual connections from passive summation into active retrieval: the DiT is split evenly along depth into an encoder and a decoder phase, encoder-side representations serve as differentiable residual sources, and each decoder sublayer adaptively fuses its mirrored encoder source through lightweight softmax routing; on ImageNet class-conditional generation, with under 0.1% additional parameters, the method lowers FID from 69.81 to 62.48 on DiT-S/2, from 44.21 to 36.33 on DiT-B/2, and from 18.85 to 15.
The work systematically analyzes Diffusion Transformer internal representations, finding a latent preference for early-layer feature reuse and symmetric layer pairing, and on that basis proposes a structured connectivity design that turns residual connections from passive summation into active retrieval: the DiT is split evenly along depth into an encoder and a decoder phase, encoder-side representations serve as differentiable residual sources, and each decoder sublayer adaptively fuses its mirrored encoder source through lightweight softmax routing; on ImageNet class-conditional generation, with under 0.1% additional parameters, the method lowers FID from 69.81 to 62.48 on DiT-S/2, from 44.21 to 36.33 on DiT-B/2, and from 18.85 to 15.
The work systematically analyzes Diffusion Transformer internal representations, finding a latent preference for early-layer feature reuse and symmetric layer pairing, and on that basis proposes a structured connectivity design that turns residual connections from passive summation into active retrieval: the DiT is split evenly along depth into an encoder and a decoder phase, encoder-side representations serve as differentiable residual sources, and each decoder sublayer adaptively fuses its mirrored encoder source through lightweight softmax routing; on ImageNet class-conditional generation, with under 0.1% additional parameters, the method lowers FID from 69.81 to 62.48 on DiT-S/2, from 44.21 to 36.33 on DiT-B/2, and from 18.85 to 15.
arXiv Drawing on 1,729,171 pull requests across 103 software ecosystems, the authors built WideSWE, 120 real tasks (60 bug fixes, 60 features) that require coordinated changes across multiple repositories, and evaluated seven coding-agent configurations: full task success ranges from 10.83% to 42.50%, with Codex CLI paired with GPT-5.6-sol highest, while joint and per-repository independent execution each help in different ways.
Drawing on 1,729,171 pull requests across 103 software ecosystems, the authors built WideSWE, 120 real tasks (60 bug fixes, 60 features) that require coordinated changes across multiple repositories, and evaluated seven coding-agent configurations: full task success ranges from 10.83% to 42.50%, with Codex CLI paired with GPT-5.6-sol highest, while joint and per-repository independent execution each help in different ways.
Drawing on 1,729,171 pull requests across 103 software ecosystems, the authors built WideSWE, 120 real tasks (60 bug fixes, 60 features) that require coordinated changes across multiple repositories, and evaluated seven coding-agent configurations: full task success ranges from 10.83% to 42.50%, with Codex CLI paired with GPT-5.6-sol highest, while joint and per-repository independent execution each help in different ways.
Drawing on 1,729,171 pull requests across 103 software ecosystems, the authors built WideSWE, 120 real tasks (60 bug fixes, 60 features) that require coordinated changes across multiple repositories, and evaluated seven coding-agent configurations: full task success ranges from 10.83% to 42.50%, with Codex CLI paired with GPT-5.6-sol highest, while joint and per-repository independent execution each help in different ways.
arXiv The work introduces AdaptiveSafety, a dataset of 10,939 training and 1,000 test examples covering policies with 1 to 100 rules, together with SafePO, a reinforcement learning algorithm, and uses them to train the AdaGuard family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time; the 4B model achieves binary accuracies of 89.30% on AdaptiveSafety and 71.82% on DynaBench.
The work introduces AdaptiveSafety, a dataset of 10,939 training and 1,000 test examples covering policies with 1 to 100 rules, together with SafePO, a reinforcement learning algorithm, and uses them to train the AdaGuard family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time; the 4B model achieves binary accuracies of 89.30% on AdaptiveSafety and 71.82% on DynaBench.
The work introduces AdaptiveSafety, a dataset of 10,939 training and 1,000 test examples covering policies with 1 to 100 rules, together with SafePO, a reinforcement learning algorithm, and uses them to train the AdaGuard family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time; the 4B model achieves binary accuracies of 89.30% on AdaptiveSafety and 71.82% on DynaBench.
The work introduces AdaptiveSafety, a dataset of 10,939 training and 1,000 test examples covering policies with 1 to 100 rules, together with SafePO, a reinforcement learning algorithm, and uses them to train the AdaGuard family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time; the 4B model achieves binary accuracies of 89.30% on AdaptiveSafety and 71.82% on DynaBench.
arXiv RoboFoundry proposes a Self-Evolving System-as-Policy framework that treats the entire supporting system around a foundation model—its context system and hierarchical skill system—rather than model parameters as the policy to evolve, diagnosing capability gaps from execution traces, converting them into validated task-specific system updates, and promoting recurring improvements into general system capabilities; it achieves state-of-the-art results on EmbodiedBench (improving GPT-5.5 by 27.8% and bringing Qwen3.7-Plus to 70.3% versus GPT-5.5's 72.7%), outperforms all baselines on RoboMemArena long-horizon memory by at least 39.0%, beats Cap-Agent0 on LIBERO-PRO by 243.8%–679.7%, and demonstrates zero-shot transfer and online evolution on real robots.
RoboFoundry proposes a Self-Evolving System-as-Policy framework that treats the entire supporting system around a foundation model—its context system and hierarchical skill system—rather than model parameters as the policy to evolve, diagnosing capability gaps from execution traces, converting them into validated task-specific system updates, and promoting recurring improvements into general system capabilities; it achieves state-of-the-art results on EmbodiedBench (improving GPT-5.5 by 27.8% and bringing Qwen3.7-Plus to 70.3% versus GPT-5.5's 72.7%), outperforms all baselines on RoboMemArena long-horizon memory by at least 39.0%, beats Cap-Agent0 on LIBERO-PRO by 243.8%–679.7%, and demonstrates zero-shot transfer and online evolution on real robots.
RoboFoundry proposes a Self-Evolving System-as-Policy framework that treats the entire supporting system around a foundation model—its context system and hierarchical skill system—rather than model parameters as the policy to evolve, diagnosing capability gaps from execution traces, converting them into validated task-specific system updates, and promoting recurring improvements into general system capabilities; it achieves state-of-the-art results on EmbodiedBench (improving GPT-5.5 by 27.8% and bringing Qwen3.7-Plus to 70.3% versus GPT-5.5's 72.7%), outperforms all baselines on RoboMemArena long-horizon memory by at least 39.0%, beats Cap-Agent0 on LIBERO-PRO by 243.8%–679.7%, and demonstrates zero-shot transfer and online evolution on real robots.
RoboFoundry proposes a Self-Evolving System-as-Policy framework that treats the entire supporting system around a foundation model—its context system and hierarchical skill system—rather than model parameters as the policy to evolve, diagnosing capability gaps from execution traces, converting them into validated task-specific system updates, and promoting recurring improvements into general system capabilities; it achieves state-of-the-art results on EmbodiedBench (improving GPT-5.5 by 27.8% and bringing Qwen3.7-Plus to 70.3% versus GPT-5.5's 72.7%), outperforms all baselines on RoboMemArena long-horizon memory by at least 39.0%, beats Cap-Agent0 on LIBERO-PRO by 243.8%–679.7%, and demonstrates zero-shot transfer and online evolution on real robots.
arXiv The work proposes a more realistic threat model for distillation attacks in which attackers continue training with reinforcement learning after distillation, and shows experimentally that distillation followed by RL outperforms either alone, that a well-known defense (antidistillation sampling) can appear effective after distillation but be broken after RL, and that distilling on approximate reasoning traces reconstructed from closed-source API summaries and final answers, followed by RL, yields reasoning gains similar to using full hidden traces.
The work proposes a more realistic threat model for distillation attacks in which attackers continue training with reinforcement learning after distillation, and shows experimentally that distillation followed by RL outperforms either alone, that a well-known defense (antidistillation sampling) can appear effective after distillation but be broken after RL, and that distilling on approximate reasoning traces reconstructed from closed-source API summaries and final answers, followed by RL, yields reasoning gains similar to using full hidden traces.
The work proposes a more realistic threat model for distillation attacks in which attackers continue training with reinforcement learning after distillation, and shows experimentally that distillation followed by RL outperforms either alone, that a well-known defense (antidistillation sampling) can appear effective after distillation but be broken after RL, and that distilling on approximate reasoning traces reconstructed from closed-source API summaries and final answers, followed by RL, yields reasoning gains similar to using full hidden traces.
The work proposes a more realistic threat model for distillation attacks in which attackers continue training with reinforcement learning after distillation, and shows experimentally that distillation followed by RL outperforms either alone, that a well-known defense (antidistillation sampling) can appear effective after distillation but be broken after RL, and that distilling on approximate reasoning traces reconstructed from closed-source API summaries and final answers, followed by RL, yields reasoning gains similar to using full hidden traces.
arXiv RenderRank renders document text as images and encodes them into compressed visual tokens, aligning relevance scores to a text-based teacher and then refining positive-versus-negative scores within each query, reaching an average NDCG@10 of 55.96 on 11 BEIR datasets with 290.07 average input tokens, 16.5–35.5% fewer than the evaluated text-based rerankers, and an average NDCG@10 of 88.27 on four long-document datasets at roughly half their average input token count.
RenderRank renders document text as images and encodes them into compressed visual tokens, aligning relevance scores to a text-based teacher and then refining positive-versus-negative scores within each query, reaching an average NDCG@10 of 55.96 on 11 BEIR datasets with 290.07 average input tokens, 16.5–35.5% fewer than the evaluated text-based rerankers, and an average NDCG@10 of 88.27 on four long-document datasets at roughly half their average input token count.
RenderRank renders document text as images and encodes them into compressed visual tokens, aligning relevance scores to a text-based teacher and then refining positive-versus-negative scores within each query, reaching an average NDCG@10 of 55.96 on 11 BEIR datasets with 290.07 average input tokens, 16.5–35.5% fewer than the evaluated text-based rerankers, and an average NDCG@10 of 88.27 on four long-document datasets at roughly half their average input token count.
RenderRank renders document text as images and encodes them into compressed visual tokens, aligning relevance scores to a text-based teacher and then refining positive-versus-negative scores within each query, reaching an average NDCG@10 of 55.96 on 11 BEIR datasets with 290.07 average input tokens, 16.5–35.5% fewer than the evaluated text-based rerankers, and an average NDCG@10 of 88.27 on four long-document datasets at roughly half their average input token count.
arXiv Using a fixed-context Pass@K analysis, this work shows that degradation after visual-token pruning is not only lost information but also unreliable use of surviving evidence (a representation-utilization gap), and introduces SCOPD and SCOPD+ sparse-context on-policy self-distillation, raising the normalized average over 13 benchmarks at 10% visual-token retention from 86.37% (Vanilla) to 90.49% and 92.43% without architectural changes, extra inference-time computation, or ground-truth answers.
Using a fixed-context Pass@K analysis, this work shows that degradation after visual-token pruning is not only lost information but also unreliable use of surviving evidence (a representation-utilization gap), and introduces SCOPD and SCOPD+ sparse-context on-policy self-distillation, raising the normalized average over 13 benchmarks at 10% visual-token retention from 86.37% (Vanilla) to 90.49% and 92.43% without architectural changes, extra inference-time computation, or ground-truth answers.
Using a fixed-context Pass@K analysis, this work shows that degradation after visual-token pruning is not only lost information but also unreliable use of surviving evidence (a representation-utilization gap), and introduces SCOPD and SCOPD+ sparse-context on-policy self-distillation, raising the normalized average over 13 benchmarks at 10% visual-token retention from 86.37% (Vanilla) to 90.49% and 92.43% without architectural changes, extra inference-time computation, or ground-truth answers.
Using a fixed-context Pass@K analysis, this work shows that degradation after visual-token pruning is not only lost information but also unreliable use of surviving evidence (a representation-utilization gap), and introduces SCOPD and SCOPD+ sparse-context on-policy self-distillation, raising the normalized average over 13 benchmarks at 10% visual-token retention from 86.37% (Vanilla) to 90.49% and 92.43% without architectural changes, extra inference-time computation, or ground-truth answers.
arXiv This monograph develops a unified account of training and inference in Power Law Decoder Representation language models (PLDR-LLMs): exact finite work identities decompose changes in the absolute energy of the row-centered learned map into parameter contributions, signed interactions, and numerical observation defects; positive affine blocking retains restarts at the row-constant face while the augmented AdamW state supplies the complete dynamical description; predictive renormalization acts on the complete conditional training law for a single pass over distinct corpus target blocks; experiments reveal observer and optimizer dependence, reject the tested autonomous row-state candidates, and support finite conditional prediction and state-specific operator reduction, while independent sing
This monograph develops a unified account of training and inference in Power Law Decoder Representation language models (PLDR-LLMs): exact finite work identities decompose changes in the absolute energy of the row-centered learned map into parameter contributions, signed interactions, and numerical observation defects; positive affine blocking retains restarts at the row-constant face while the augmented AdamW state supplies the complete dynamical description; predictive renormalization acts on the complete conditional training law for a single pass over distinct corpus target blocks; experiments reveal observer and optimizer dependence, reject the tested autonomous row-state candidates, and support finite conditional prediction and state-specific operator reduction, while independent sing
This monograph develops a unified account of training and inference in Power Law Decoder Representation language models (PLDR-LLMs): exact finite work identities decompose changes in the absolute energy of the row-centered learned map into parameter contributions, signed interactions, and numerical observation defects; positive affine blocking retains restarts at the row-constant face while the augmented AdamW state supplies the complete dynamical description; predictive renormalization acts on the complete conditional training law for a single pass over distinct corpus target blocks; experiments reveal observer and optimizer dependence, reject the tested autonomous row-state candidates, and support finite conditional prediction and state-specific operator reduction, while independent sing
This monograph develops a unified account of training and inference in Power Law Decoder Representation language models (PLDR-LLMs): exact finite work identities decompose changes in the absolute energy of the row-centered learned map into parameter contributions, signed interactions, and numerical observation defects; positive affine blocking retains restarts at the row-constant face while the augmented AdamW state supplies the complete dynamical description; predictive renormalization acts on the complete conditional training law for a single pass over distinct corpus target blocks; experiments reveal observer and optimizer dependence, reject the tested autonomous row-state candidates, and support finite conditional prediction and state-specific operator reduction, while independent sing
arXiv TokenCast learns a composable cost triple for each execution segment of an LLM agent (call count, net input-length change, and a cost residual), uses an exact composition identity to fold the cost of earlier context being re-read by every later call into a cumulative estimate, and refreshes the forecast from newly observed execution evidence without additional LLM calls; on SWE-bench Verified the mean cumulative prediction time is 32.8 ms per run, mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations across 4 task suites and 6 agent models, and offline budget-control replay uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion.
TokenCast learns a composable cost triple for each execution segment of an LLM agent (call count, net input-length change, and a cost residual), uses an exact composition identity to fold the cost of earlier context being re-read by every later call into a cumulative estimate, and refreshes the forecast from newly observed execution evidence without additional LLM calls; on SWE-bench Verified the mean cumulative prediction time is 32.8 ms per run, mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations across 4 task suites and 6 agent models, and offline budget-control replay uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion.
TokenCast learns a composable cost triple for each execution segment of an LLM agent (call count, net input-length change, and a cost residual), uses an exact composition identity to fold the cost of earlier context being re-read by every later call into a cumulative estimate, and refreshes the forecast from newly observed execution evidence without additional LLM calls; on SWE-bench Verified the mean cumulative prediction time is 32.8 ms per run, mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations across 4 task suites and 6 agent models, and offline budget-control replay uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion.
TokenCast learns a composable cost triple for each execution segment of an LLM agent (call count, net input-length change, and a cost residual), uses an exact composition identity to fold the cost of earlier context being re-read by every later call into a cumulative estimate, and refreshes the forecast from newly observed execution evidence without additional LLM calls; on SWE-bench Verified the mean cumulative prediction time is 32.8 ms per run, mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations across 4 task suites and 6 agent models, and offline budget-control replay uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion.
arXiv The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
arXiv In an auction-logic-faithful on-device simulation with 36 campaigns, 50 devices, and 30 paired demand paths, the study examines two economic misalignments created when privacy-preserving ML decisions move onto clients, finding that proportional Even pacing overspends 17.77% after one tick of staleness and 1,669.31% after 50 ticks, that 50-tick overspend remains 106.95% at two-times budget pressure, that a visible-budget no-sale guard makes zero-lag compliance exact yet leaves 11.88% overspend at one tick, and that 98.23% of rival auctions at one tick admit a profitable deviation once the ML/pacing score transformation is allowed to change payment units.
In an auction-logic-faithful on-device simulation with 36 campaigns, 50 devices, and 30 paired demand paths, the study examines two economic misalignments created when privacy-preserving ML decisions move onto clients, finding that proportional Even pacing overspends 17.77% after one tick of staleness and 1,669.31% after 50 ticks, that 50-tick overspend remains 106.95% at two-times budget pressure, that a visible-budget no-sale guard makes zero-lag compliance exact yet leaves 11.88% overspend at one tick, and that 98.23% of rival auctions at one tick admit a profitable deviation once the ML/pacing score transformation is allowed to change payment units.
In an auction-logic-faithful on-device simulation with 36 campaigns, 50 devices, and 30 paired demand paths, the study examines two economic misalignments created when privacy-preserving ML decisions move onto clients, finding that proportional Even pacing overspends 17.77% after one tick of staleness and 1,669.31% after 50 ticks, that 50-tick overspend remains 106.95% at two-times budget pressure, that a visible-budget no-sale guard makes zero-lag compliance exact yet leaves 11.88% overspend at one tick, and that 98.23% of rival auctions at one tick admit a profitable deviation once the ML/pacing score transformation is allowed to change payment units.
In an auction-logic-faithful on-device simulation with 36 campaigns, 50 devices, and 30 paired demand paths, the study examines two economic misalignments created when privacy-preserving ML decisions move onto clients, finding that proportional Even pacing overspends 17.77% after one tick of staleness and 1,669.31% after 50 ticks, that 50-tick overspend remains 106.95% at two-times budget pressure, that a visible-budget no-sale guard makes zero-lag compliance exact yet leaves 11.88% overspend at one tick, and that 98.23% of rival auctions at one tick admit a profitable deviation once the ML/pacing score transformation is allowed to change payment units.