arXiv Relic lets multi-agent organizations turn recurring collaboration failures into organization-owned, executable, revisable protocols, raising complete-contract delivery on mainline from 14.06% to 19.76% (+5.71 percentage points) across 360 controlled runs over ten software workloads and three models, and raising behavioral correctness under fresh-member transfer from 25.4% with no inherited protocol and 34.6% with text-only rules to 41.2% with executable bindings.
Relic lets multi-agent organizations turn recurring collaboration failures into organization-owned, executable, revisable protocols, raising complete-contract delivery on mainline from 14.06% to 19.76% (+5.71 percentage points) across 360 controlled runs over ten software workloads and three models, and raising behavioral correctness under fresh-member transfer from 25.4% with no inherited protocol and 34.6% with text-only rules to 41.2% with executable bindings.
Relic lets multi-agent organizations turn recurring collaboration failures into organization-owned, executable, revisable protocols, raising complete-contract delivery on mainline from 14.06% to 19.76% (+5.71 percentage points) across 360 controlled runs over ten software workloads and three models, and raising behavioral correctness under fresh-member transfer from 25.4% with no inherited protocol and 34.6% with text-only rules to 41.2% with executable bindings.
Relic lets multi-agent organizations turn recurring collaboration failures into organization-owned, executable, revisable protocols, raising complete-contract delivery on mainline from 14.06% to 19.76% (+5.71 percentage points) across 360 controlled runs over ten software workloads and three models, and raising behavioral correctness under fresh-member transfer from 25.4% with no inherited protocol and 34.6% with text-only rules to 41.2% with executable bindings.
arXiv The work reinterprets the reverse-KL objective of on-policy distillation (OPD) as KL-regularized policy optimization and introduces Least-Square Policy Distillation (LSPD), which replaces policy-gradient updates with robust quadratic log-probability matching plus entropy regularization, enabling multiple updates per rollout and reuse of historical trajectories; across six mathematical reasoning benchmarks and three teacher-student settings it improves Avg@16 by +1.59 and Pass@16 by +1.87 on average, and its fully off-policy replay variant LSPD-RB reaches performance comparable to vanilla OPD using roughly one quarter of the rollout batches.
The work reinterprets the reverse-KL objective of on-policy distillation (OPD) as KL-regularized policy optimization and introduces Least-Square Policy Distillation (LSPD), which replaces policy-gradient updates with robust quadratic log-probability matching plus entropy regularization, enabling multiple updates per rollout and reuse of historical trajectories; across six mathematical reasoning benchmarks and three teacher-student settings it improves Avg@16 by +1.59 and Pass@16 by +1.87 on average, and its fully off-policy replay variant LSPD-RB reaches performance comparable to vanilla OPD using roughly one quarter of the rollout batches.
The work reinterprets the reverse-KL objective of on-policy distillation (OPD) as KL-regularized policy optimization and introduces Least-Square Policy Distillation (LSPD), which replaces policy-gradient updates with robust quadratic log-probability matching plus entropy regularization, enabling multiple updates per rollout and reuse of historical trajectories; across six mathematical reasoning benchmarks and three teacher-student settings it improves Avg@16 by +1.59 and Pass@16 by +1.87 on average, and its fully off-policy replay variant LSPD-RB reaches performance comparable to vanilla OPD using roughly one quarter of the rollout batches.
The work reinterprets the reverse-KL objective of on-policy distillation (OPD) as KL-regularized policy optimization and introduces Least-Square Policy Distillation (LSPD), which replaces policy-gradient updates with robust quadratic log-probability matching plus entropy regularization, enabling multiple updates per rollout and reuse of historical trajectories; across six mathematical reasoning benchmarks and three teacher-student settings it improves Avg@16 by +1.59 and Pass@16 by +1.87 on average, and its fully off-policy replay variant LSPD-RB reaches performance comparable to vanilla OPD using roughly one quarter of the rollout batches.
arXiv The work introduces DepthBench, a controlled benchmark that varies the width–depth aspect ratio (d_model/n_layer) from shallow-wide to deep-narrow under a fixed parameter budget and pre-training recipe, comparing 10 representative residual and normalization architectures, and finds that the benefit of depth is strongly architecture-dependent: Pre-LN and most of its norm- and scaling-based variants worsen as models get deeper and narrower (Pre-LN rises from 2.759 to 2.782), whereas HC and Full AttnRes improve consistently (HC from 2.729 to 2.699, Full AttnRes from 2.751 to 2.718), accompanied by stronger cross-layer causal dependence and order sensitivity, though deeper models also incur compute and memory overhead.
The work introduces DepthBench, a controlled benchmark that varies the width–depth aspect ratio (d_model/n_layer) from shallow-wide to deep-narrow under a fixed parameter budget and pre-training recipe, comparing 10 representative residual and normalization architectures, and finds that the benefit of depth is strongly architecture-dependent: Pre-LN and most of its norm- and scaling-based variants worsen as models get deeper and narrower (Pre-LN rises from 2.759 to 2.782), whereas HC and Full AttnRes improve consistently (HC from 2.729 to 2.699, Full AttnRes from 2.751 to 2.718), accompanied by stronger cross-layer causal dependence and order sensitivity, though deeper models also incur compute and memory overhead.
The work introduces DepthBench, a controlled benchmark that varies the width–depth aspect ratio (d_model/n_layer) from shallow-wide to deep-narrow under a fixed parameter budget and pre-training recipe, comparing 10 representative residual and normalization architectures, and finds that the benefit of depth is strongly architecture-dependent: Pre-LN and most of its norm- and scaling-based variants worsen as models get deeper and narrower (Pre-LN rises from 2.759 to 2.782), whereas HC and Full AttnRes improve consistently (HC from 2.729 to 2.699, Full AttnRes from 2.751 to 2.718), accompanied by stronger cross-layer causal dependence and order sensitivity, though deeper models also incur compute and memory overhead.
The work introduces DepthBench, a controlled benchmark that varies the width–depth aspect ratio (d_model/n_layer) from shallow-wide to deep-narrow under a fixed parameter budget and pre-training recipe, comparing 10 representative residual and normalization architectures, and finds that the benefit of depth is strongly architecture-dependent: Pre-LN and most of its norm- and scaling-based variants worsen as models get deeper and narrower (Pre-LN rises from 2.759 to 2.782), whereas HC and Full AttnRes improve consistently (HC from 2.729 to 2.699, Full AttnRes from 2.751 to 2.718), accompanied by stronger cross-layer causal dependence and order sensitivity, though deeper models also incur compute and memory overhead.
arXiv YuE2 uses a single roughly 3.58B-parameter AR–NAR Mixture-of-Transformers to first write an ABC score specifying melody, chords, key, meter, tempo, and form, then expand it into semantic music tokens and render full-song audio, with companion models MERT2 and SheetSage2 constructing semantic and symbolic supervision from recordings; with the same checkpoint, experts prefer symbolic planning 49.3% to 34.6% for overall quality, YuE2 scores 6.7316 on SongBench Global Avg on WildSongBench (6.9632 with best-of-8), MERT2 surpasses previous best results on 14 of 15 MARBLE metrics, and SheetSage2-AR leads 12 of 15 benchmark–metric pairs.
YuE2 uses a single roughly 3.58B-parameter AR–NAR Mixture-of-Transformers to first write an ABC score specifying melody, chords, key, meter, tempo, and form, then expand it into semantic music tokens and render full-song audio, with companion models MERT2 and SheetSage2 constructing semantic and symbolic supervision from recordings; with the same checkpoint, experts prefer symbolic planning 49.3% to 34.6% for overall quality, YuE2 scores 6.7316 on SongBench Global Avg on WildSongBench (6.9632 with best-of-8), MERT2 surpasses previous best results on 14 of 15 MARBLE metrics, and SheetSage2-AR leads 12 of 15 benchmark–metric pairs.
YuE2 uses a single roughly 3.58B-parameter AR–NAR Mixture-of-Transformers to first write an ABC score specifying melody, chords, key, meter, tempo, and form, then expand it into semantic music tokens and render full-song audio, with companion models MERT2 and SheetSage2 constructing semantic and symbolic supervision from recordings; with the same checkpoint, experts prefer symbolic planning 49.3% to 34.6% for overall quality, YuE2 scores 6.7316 on SongBench Global Avg on WildSongBench (6.9632 with best-of-8), MERT2 surpasses previous best results on 14 of 15 MARBLE metrics, and SheetSage2-AR leads 12 of 15 benchmark–metric pairs.
YuE2 uses a single roughly 3.58B-parameter AR–NAR Mixture-of-Transformers to first write an ABC score specifying melody, chords, key, meter, tempo, and form, then expand it into semantic music tokens and render full-song audio, with companion models MERT2 and SheetSage2 constructing semantic and symbolic supervision from recordings; with the same checkpoint, experts prefer symbolic planning 49.3% to 34.6% for overall quality, YuE2 scores 6.7316 on SongBench Global Avg on WildSongBench (6.9632 with best-of-8), MERT2 surpasses previous best results on 14 of 15 MARBLE metrics, and SheetSage2-AR leads 12 of 15 benchmark–metric pairs.
arXiv The authors ran an end-to-end self-audit of an evaluation of LLM-inferred prompt structure: across eight open model variants spanning five families and 8B to 675B parameters, with caching disabled and 293 raw intermediate representations persisted, repeated identical calls yielded mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells were never node-set-perfect; a joint cluster bootstrap showed only the bottom of the ranking is firm (99% and 86% rank retention), the middle four at 27% to 48% and the top two at 68% each; two equally defensible merge rules changed four of eight rows and moved the study-wide headline by 7 percentage points; reproducibility could not be read as accuracy; and four of eight endpoints were withdrawn within ten weeks, so the study as specified can
The authors ran an end-to-end self-audit of an evaluation of LLM-inferred prompt structure: across eight open model variants spanning five families and 8B to 675B parameters, with caching disabled and 293 raw intermediate representations persisted, repeated identical calls yielded mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells were never node-set-perfect; a joint cluster bootstrap showed only the bottom of the ranking is firm (99% and 86% rank retention), the middle four at 27% to 48% and the top two at 68% each; two equally defensible merge rules changed four of eight rows and moved the study-wide headline by 7 percentage points; reproducibility could not be read as accuracy; and four of eight endpoints were withdrawn within ten weeks, so the study as specified can
The authors ran an end-to-end self-audit of an evaluation of LLM-inferred prompt structure: across eight open model variants spanning five families and 8B to 675B parameters, with caching disabled and 293 raw intermediate representations persisted, repeated identical calls yielded mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells were never node-set-perfect; a joint cluster bootstrap showed only the bottom of the ranking is firm (99% and 86% rank retention), the middle four at 27% to 48% and the top two at 68% each; two equally defensible merge rules changed four of eight rows and moved the study-wide headline by 7 percentage points; reproducibility could not be read as accuracy; and four of eight endpoints were withdrawn within ten weeks, so the study as specified can
The authors ran an end-to-end self-audit of an evaluation of LLM-inferred prompt structure: across eight open model variants spanning five families and 8B to 675B parameters, with caching disabled and 293 raw intermediate representations persisted, repeated identical calls yielded mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells were never node-set-perfect; a joint cluster bootstrap showed only the bottom of the ranking is firm (99% and 86% rank retention), the middle four at 27% to 48% and the top two at 68% each; two equally defensible merge rules changed four of eight rows and moved the study-wide headline by 7 percentage points; reproducibility could not be read as accuracy; and four of eight endpoints were withdrawn within ten weeks, so the study as specified can
arXiv FlashForward publishes the in-flight KV cache already computed by each ordinary denoising forward for reuse by later chunks and complements it with sparse clean anchor KV generated ahead of time for long-range structural guidance, removing cache-update-only forwards; with up to four GPUs it runs 1.16–1.69x faster than HiAR and 1.42–2.92x faster than Self-Forcing for 16 FPS videos of 20 seconds or longer across 1.3B and 14B backbones at 480p and 720p, and the 1.3B model at 480p scores 0.838 on VBench-1.0 while remaining stable at 20s, 35s, and 65s.
FlashForward publishes the in-flight KV cache already computed by each ordinary denoising forward for reuse by later chunks and complements it with sparse clean anchor KV generated ahead of time for long-range structural guidance, removing cache-update-only forwards; with up to four GPUs it runs 1.16–1.69x faster than HiAR and 1.42–2.92x faster than Self-Forcing for 16 FPS videos of 20 seconds or longer across 1.3B and 14B backbones at 480p and 720p, and the 1.3B model at 480p scores 0.838 on VBench-1.0 while remaining stable at 20s, 35s, and 65s.
FlashForward publishes the in-flight KV cache already computed by each ordinary denoising forward for reuse by later chunks and complements it with sparse clean anchor KV generated ahead of time for long-range structural guidance, removing cache-update-only forwards; with up to four GPUs it runs 1.16–1.69x faster than HiAR and 1.42–2.92x faster than Self-Forcing for 16 FPS videos of 20 seconds or longer across 1.3B and 14B backbones at 480p and 720p, and the 1.3B model at 480p scores 0.838 on VBench-1.0 while remaining stable at 20s, 35s, and 65s.
FlashForward publishes the in-flight KV cache already computed by each ordinary denoising forward for reuse by later chunks and complements it with sparse clean anchor KV generated ahead of time for long-range structural guidance, removing cache-update-only forwards; with up to four GPUs it runs 1.16–1.69x faster than HiAR and 1.42–2.92x faster than Self-Forcing for 16 FPS videos of 20 seconds or longer across 1.3B and 14B backbones at 480p and 720p, and the 1.3B model at 480p scores 0.838 on VBench-1.0 while remaining stable at 20s, 35s, and 65s.
arXiv KernelZero introduces a co-evolution framework in which a Proposer generates Torch modules at the Coder's current capability frontier and the Coder is trained with CA-GRPO so that performance is optimized only after correctness becomes reliable; both models start from Qwen2.5-Coder-7B and reach 75.8/69.6 pass@1 on KernelBench CUDA Level 1/2 and 77.2/72.5 on Triton, surpassing Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton.
KernelZero introduces a co-evolution framework in which a Proposer generates Torch modules at the Coder's current capability frontier and the Coder is trained with CA-GRPO so that performance is optimized only after correctness becomes reliable; both models start from Qwen2.5-Coder-7B and reach 75.8/69.6 pass@1 on KernelBench CUDA Level 1/2 and 77.2/72.5 on Triton, surpassing Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton.
KernelZero introduces a co-evolution framework in which a Proposer generates Torch modules at the Coder's current capability frontier and the Coder is trained with CA-GRPO so that performance is optimized only after correctness becomes reliable; both models start from Qwen2.5-Coder-7B and reach 75.8/69.6 pass@1 on KernelBench CUDA Level 1/2 and 77.2/72.5 on Triton, surpassing Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton.
KernelZero introduces a co-evolution framework in which a Proposer generates Torch modules at the Coder's current capability frontier and the Coder is trained with CA-GRPO so that performance is optimized only after correctness becomes reliable; both models start from Qwen2.5-Coder-7B and reach 75.8/69.6 pass@1 on KernelBench CUDA Level 1/2 and 77.2/72.5 on Triton, surpassing Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton.
arXiv The work introduces DRM (Diffusion Reward Model), which recasts reward modeling as conditional density estimation: on top of a frozen LLM encoder, a lightweight Diffusion Transformer denoises Gaussian noise into a reward vector with no parametric family assumed, a single architecture handles both multi-attribute regression and pairwise preference data, it matches or surpasses the same-data, same-backbone scalar head ArmoRM and quantile head QRM across five benchmarks, and it exhibits multimodal structure tied to human disagreement on repeated-annotation data, with uncertainty-aware rejection, LCB aggregation, and downstream RLHF experiments demonstrating the practical value of distributional information.
The work introduces DRM (Diffusion Reward Model), which recasts reward modeling as conditional density estimation: on top of a frozen LLM encoder, a lightweight Diffusion Transformer denoises Gaussian noise into a reward vector with no parametric family assumed, a single architecture handles both multi-attribute regression and pairwise preference data, it matches or surpasses the same-data, same-backbone scalar head ArmoRM and quantile head QRM across five benchmarks, and it exhibits multimodal structure tied to human disagreement on repeated-annotation data, with uncertainty-aware rejection, LCB aggregation, and downstream RLHF experiments demonstrating the practical value of distributional information.
The work introduces DRM (Diffusion Reward Model), which recasts reward modeling as conditional density estimation: on top of a frozen LLM encoder, a lightweight Diffusion Transformer denoises Gaussian noise into a reward vector with no parametric family assumed, a single architecture handles both multi-attribute regression and pairwise preference data, it matches or surpasses the same-data, same-backbone scalar head ArmoRM and quantile head QRM across five benchmarks, and it exhibits multimodal structure tied to human disagreement on repeated-annotation data, with uncertainty-aware rejection, LCB aggregation, and downstream RLHF experiments demonstrating the practical value of distributional information.
The work introduces DRM (Diffusion Reward Model), which recasts reward modeling as conditional density estimation: on top of a frozen LLM encoder, a lightweight Diffusion Transformer denoises Gaussian noise into a reward vector with no parametric family assumed, a single architecture handles both multi-attribute regression and pairwise preference data, it matches or surpasses the same-data, same-backbone scalar head ArmoRM and quantile head QRM across five benchmarks, and it exhibits multimodal structure tied to human disagreement on repeated-annotation data, with uncertainty-aware rejection, LCB aggregation, and downstream RLHF experiments demonstrating the practical value of distributional information.
arXiv The work proposes LT-OPD, which supervises a student keeping only 5% of visual tokens on its own generated trajectories using a frozen full-token copy of the same model as teacher, together with a curriculum that progressively lowers the token budget, raising average retained performance on nine benchmarks with Qwen3.5-4B from 68.6% to 82.3% while cutting KV cache by 85.2% and prefill FLOPs by 85.4% with no added inference overhead.
The work proposes LT-OPD, which supervises a student keeping only 5% of visual tokens on its own generated trajectories using a frozen full-token copy of the same model as teacher, together with a curriculum that progressively lowers the token budget, raising average retained performance on nine benchmarks with Qwen3.5-4B from 68.6% to 82.3% while cutting KV cache by 85.2% and prefill FLOPs by 85.4% with no added inference overhead.
The work proposes LT-OPD, which supervises a student keeping only 5% of visual tokens on its own generated trajectories using a frozen full-token copy of the same model as teacher, together with a curriculum that progressively lowers the token budget, raising average retained performance on nine benchmarks with Qwen3.5-4B from 68.6% to 82.3% while cutting KV cache by 85.2% and prefill FLOPs by 85.4% with no added inference overhead.
The work proposes LT-OPD, which supervises a student keeping only 5% of visual tokens on its own generated trajectories using a frozen full-token copy of the same model as teacher, together with a curriculum that progressively lowers the token budget, raising average retained performance on nine benchmarks with Qwen3.5-4B from 68.6% to 82.3% while cutting KV cache by 85.2% and prefill FLOPs by 85.4% with no added inference overhead.
arXiv GeoVerse performs diffusion generation inside the geometric latent space of a pretrained 3D foundation model (DA3), injects frozen Wan2.2 VACE video features into the denoiser through a ControlNet-style adapter, and uses a continuously updated colored point-cloud spatial memory to reproject history as target-aligned guidance, synthesizing cross-view consistent novel views from sparse inputs; across DL3DV, RealEstate10K, Mip-NeRF 360, and ScanNetV2 it reports better visual quality and geometric consistency, with PSNR 2.23 dB higher than GLD on DL3DV and ATE 32.4% lower than GLD on Mip-NeRF 360, while reflow distillation brings inference from over 100 seconds to under 10 seconds.
GeoVerse performs diffusion generation inside the geometric latent space of a pretrained 3D foundation model (DA3), injects frozen Wan2.2 VACE video features into the denoiser through a ControlNet-style adapter, and uses a continuously updated colored point-cloud spatial memory to reproject history as target-aligned guidance, synthesizing cross-view consistent novel views from sparse inputs; across DL3DV, RealEstate10K, Mip-NeRF 360, and ScanNetV2 it reports better visual quality and geometric consistency, with PSNR 2.23 dB higher than GLD on DL3DV and ATE 32.4% lower than GLD on Mip-NeRF 360, while reflow distillation brings inference from over 100 seconds to under 10 seconds.
GeoVerse performs diffusion generation inside the geometric latent space of a pretrained 3D foundation model (DA3), injects frozen Wan2.2 VACE video features into the denoiser through a ControlNet-style adapter, and uses a continuously updated colored point-cloud spatial memory to reproject history as target-aligned guidance, synthesizing cross-view consistent novel views from sparse inputs; across DL3DV, RealEstate10K, Mip-NeRF 360, and ScanNetV2 it reports better visual quality and geometric consistency, with PSNR 2.23 dB higher than GLD on DL3DV and ATE 32.4% lower than GLD on Mip-NeRF 360, while reflow distillation brings inference from over 100 seconds to under 10 seconds.
GeoVerse performs diffusion generation inside the geometric latent space of a pretrained 3D foundation model (DA3), injects frozen Wan2.2 VACE video features into the denoiser through a ControlNet-style adapter, and uses a continuously updated colored point-cloud spatial memory to reproject history as target-aligned guidance, synthesizing cross-view consistent novel views from sparse inputs; across DL3DV, RealEstate10K, Mip-NeRF 360, and ScanNetV2 it reports better visual quality and geometric consistency, with PSNR 2.23 dB higher than GLD on DL3DV and ATE 32.4% lower than GLD on Mip-NeRF 360, while reflow distillation brings inference from over 100 seconds to under 10 seconds.
arXiv The work identifies Decision–Timestamp Mismatch in on-policy self-distillation for long-horizon agents—privileged guidance may correspond to a decision at a different timestep, and a student decision may span multiple timesteps—and introduces AlignOPSD, which first rectifies privileged supervision using functional correspondences across sibling rollouts and then applies semi-Markov hierarchical credit assignment over variable-duration decision spans; with Qwen2.5-3B/7B it outperforms GRPO on all eight backbone–aggregate-metric comparisons across ALFWorld, WebShop, and Search-QA by 5.5–8.7%, ranking first in six.
The work identifies Decision–Timestamp Mismatch in on-policy self-distillation for long-horizon agents—privileged guidance may correspond to a decision at a different timestep, and a student decision may span multiple timesteps—and introduces AlignOPSD, which first rectifies privileged supervision using functional correspondences across sibling rollouts and then applies semi-Markov hierarchical credit assignment over variable-duration decision spans; with Qwen2.5-3B/7B it outperforms GRPO on all eight backbone–aggregate-metric comparisons across ALFWorld, WebShop, and Search-QA by 5.5–8.7%, ranking first in six.
The work identifies Decision–Timestamp Mismatch in on-policy self-distillation for long-horizon agents—privileged guidance may correspond to a decision at a different timestep, and a student decision may span multiple timesteps—and introduces AlignOPSD, which first rectifies privileged supervision using functional correspondences across sibling rollouts and then applies semi-Markov hierarchical credit assignment over variable-duration decision spans; with Qwen2.5-3B/7B it outperforms GRPO on all eight backbone–aggregate-metric comparisons across ALFWorld, WebShop, and Search-QA by 5.5–8.7%, ranking first in six.
The work identifies Decision–Timestamp Mismatch in on-policy self-distillation for long-horizon agents—privileged guidance may correspond to a decision at a different timestep, and a student decision may span multiple timesteps—and introduces AlignOPSD, which first rectifies privileged supervision using functional correspondences across sibling rollouts and then applies semi-Markov hierarchical credit assignment over variable-duration decision spans; with Qwen2.5-3B/7B it outperforms GRPO on all eight backbone–aggregate-metric comparisons across ALFWorld, WebShop, and Search-QA by 5.5–8.7%, ranking first in six.
arXiv CompoWorld introduces compositional environment scaling: coding agents turn MCP tool specifications into executable services with typed states and shared interfaces, a random walk connects services through dependency graphs to generate and verify cross-service tasks, verified trajectories support SFT while a Completion-Focused Rubric Reward guides GRPO-based RL, and with 448 services exposing 10,130 tools plus 3K SFT trajectories and 1K RL tasks, Qwen3.6-35B-A3B improves by 9.17 points on average across eight benchmarks and raises AutomationBench task success from 10.33% to 32.33%.
CompoWorld introduces compositional environment scaling: coding agents turn MCP tool specifications into executable services with typed states and shared interfaces, a random walk connects services through dependency graphs to generate and verify cross-service tasks, verified trajectories support SFT while a Completion-Focused Rubric Reward guides GRPO-based RL, and with 448 services exposing 10,130 tools plus 3K SFT trajectories and 1K RL tasks, Qwen3.6-35B-A3B improves by 9.17 points on average across eight benchmarks and raises AutomationBench task success from 10.33% to 32.33%.
CompoWorld introduces compositional environment scaling: coding agents turn MCP tool specifications into executable services with typed states and shared interfaces, a random walk connects services through dependency graphs to generate and verify cross-service tasks, verified trajectories support SFT while a Completion-Focused Rubric Reward guides GRPO-based RL, and with 448 services exposing 10,130 tools plus 3K SFT trajectories and 1K RL tasks, Qwen3.6-35B-A3B improves by 9.17 points on average across eight benchmarks and raises AutomationBench task success from 10.33% to 32.33%.
CompoWorld introduces compositional environment scaling: coding agents turn MCP tool specifications into executable services with typed states and shared interfaces, a random walk connects services through dependency graphs to generate and verify cross-service tasks, verified trajectories support SFT while a Completion-Focused Rubric Reward guides GRPO-based RL, and with 448 services exposing 10,130 tools plus 3K SFT trajectories and 1K RL tasks, Qwen3.6-35B-A3B improves by 9.17 points on average across eight benchmarks and raises AutomationBench task success from 10.33% to 32.33%.
arXiv The work finds that multi-teacher on-policy distillation (MOPD) with domain-label routing alone fails to beat the strongest single-teacher student and transfers little of the mathematics expert's gain, because domain feedback is unbalanced (in the first training batch instruction-following log-ratios are 2.3-4.4 times as dispersed as the pooled signal while mathematics is about half as dispersed, and for the initial 4B student the instruction-following loss supplies 94% of the combined gradient), and proposes DN-MOPD, which keeps label routing and rescales each domain's distillation advantages by the clipped ratio of pooled to domain log-ratio spread estimated every batch, improving the six-task average over MOPD at every size and both evaluation budgets (1.17-2.36 points at 16K, 2.47-3.
The work finds that multi-teacher on-policy distillation (MOPD) with domain-label routing alone fails to beat the strongest single-teacher student and transfers little of the mathematics expert's gain, because domain feedback is unbalanced (in the first training batch instruction-following log-ratios are 2.3-4.4 times as dispersed as the pooled signal while mathematics is about half as dispersed, and for the initial 4B student the instruction-following loss supplies 94% of the combined gradient), and proposes DN-MOPD, which keeps label routing and rescales each domain's distillation advantages by the clipped ratio of pooled to domain log-ratio spread estimated every batch, improving the six-task average over MOPD at every size and both evaluation budgets (1.17-2.36 points at 16K, 2.47-3.
The work finds that multi-teacher on-policy distillation (MOPD) with domain-label routing alone fails to beat the strongest single-teacher student and transfers little of the mathematics expert's gain, because domain feedback is unbalanced (in the first training batch instruction-following log-ratios are 2.3-4.4 times as dispersed as the pooled signal while mathematics is about half as dispersed, and for the initial 4B student the instruction-following loss supplies 94% of the combined gradient), and proposes DN-MOPD, which keeps label routing and rescales each domain's distillation advantages by the clipped ratio of pooled to domain log-ratio spread estimated every batch, improving the six-task average over MOPD at every size and both evaluation budgets (1.17-2.36 points at 16K, 2.47-3.
The work finds that multi-teacher on-policy distillation (MOPD) with domain-label routing alone fails to beat the strongest single-teacher student and transfers little of the mathematics expert's gain, because domain feedback is unbalanced (in the first training batch instruction-following log-ratios are 2.3-4.4 times as dispersed as the pooled signal while mathematics is about half as dispersed, and for the initial 4B student the instruction-following loss supplies 94% of the combined gradient), and proposes DN-MOPD, which keeps label routing and rescales each domain's distillation advantages by the clipped ratio of pooled to domain log-ratio spread estimated every batch, improving the six-task average over MOPD at every size and both evaluation budgets (1.17-2.36 points at 16K, 2.47-3.
arXiv The work finds that instruction-based image editing models degrade rapidly under recursive multi-turn editing and attributes this to a train–test mismatch in the conditioning distribution, since models are trained on clean source images but must repeatedly condition on their own imperfect outputs at inference; it proposes MT-OPSD, which rolls out the current student under an identity instruction to produce self-generated states containing its own errors, then uses the same model conditioned on the clean source as a teacher to transfer clean-condition editing behavior onto those states via sparse velocity matching, without multi-turn annotations; it also builds LME-Bench of 100 ten-turn editing sessions, raising SR@10 to 0.38–0.52 and cutting CR@10 to at most 0.
The work finds that instruction-based image editing models degrade rapidly under recursive multi-turn editing and attributes this to a train–test mismatch in the conditioning distribution, since models are trained on clean source images but must repeatedly condition on their own imperfect outputs at inference; it proposes MT-OPSD, which rolls out the current student under an identity instruction to produce self-generated states containing its own errors, then uses the same model conditioned on the clean source as a teacher to transfer clean-condition editing behavior onto those states via sparse velocity matching, without multi-turn annotations; it also builds LME-Bench of 100 ten-turn editing sessions, raising SR@10 to 0.38–0.52 and cutting CR@10 to at most 0.
The work finds that instruction-based image editing models degrade rapidly under recursive multi-turn editing and attributes this to a train–test mismatch in the conditioning distribution, since models are trained on clean source images but must repeatedly condition on their own imperfect outputs at inference; it proposes MT-OPSD, which rolls out the current student under an identity instruction to produce self-generated states containing its own errors, then uses the same model conditioned on the clean source as a teacher to transfer clean-condition editing behavior onto those states via sparse velocity matching, without multi-turn annotations; it also builds LME-Bench of 100 ten-turn editing sessions, raising SR@10 to 0.38–0.52 and cutting CR@10 to at most 0.
The work finds that instruction-based image editing models degrade rapidly under recursive multi-turn editing and attributes this to a train–test mismatch in the conditioning distribution, since models are trained on clean source images but must repeatedly condition on their own imperfect outputs at inference; it proposes MT-OPSD, which rolls out the current student under an identity instruction to produce self-generated states containing its own errors, then uses the same model conditioned on the clean source as a teacher to transfer clean-condition editing behavior onto those states via sparse velocity matching, without multi-turn annotations; it also builds LME-Bench of 100 ten-turn editing sessions, raising SR@10 to 0.38–0.52 and cutting CR@10 to at most 0.
arXiv The work introduces MassAlloc Attention (MALA), a fused attention primitive that keeps QK score discovery over every legal causal interaction and uses normalized online-softmax contribution to decide which tiles execute post-score computation; under exactly matched post-score work at 8K it reaches 0.0188% mean omitted mass versus 0.0182% for a per-instance reference-mass oracle, holds low output and gradient errors from 1K to 32K, reaches 89.67% associative-recall accuracy at 8K versus 89.97% for FullAttn, reduces training forward and backward latency by 2.2x and 3.0x and decoding latency by 1.6x at 128K with tensor parallelism on 8 GPUs, tracks FullAttn perplexity from 0.6B to 14B while cutting total training FLOPs by 23.
The work introduces MassAlloc Attention (MALA), a fused attention primitive that keeps QK score discovery over every legal causal interaction and uses normalized online-softmax contribution to decide which tiles execute post-score computation; under exactly matched post-score work at 8K it reaches 0.0188% mean omitted mass versus 0.0182% for a per-instance reference-mass oracle, holds low output and gradient errors from 1K to 32K, reaches 89.67% associative-recall accuracy at 8K versus 89.97% for FullAttn, reduces training forward and backward latency by 2.2x and 3.0x and decoding latency by 1.6x at 128K with tensor parallelism on 8 GPUs, tracks FullAttn perplexity from 0.6B to 14B while cutting total training FLOPs by 23.
The work introduces MassAlloc Attention (MALA), a fused attention primitive that keeps QK score discovery over every legal causal interaction and uses normalized online-softmax contribution to decide which tiles execute post-score computation; under exactly matched post-score work at 8K it reaches 0.0188% mean omitted mass versus 0.0182% for a per-instance reference-mass oracle, holds low output and gradient errors from 1K to 32K, reaches 89.67% associative-recall accuracy at 8K versus 89.97% for FullAttn, reduces training forward and backward latency by 2.2x and 3.0x and decoding latency by 1.6x at 128K with tensor parallelism on 8 GPUs, tracks FullAttn perplexity from 0.6B to 14B while cutting total training FLOPs by 23.
The work introduces MassAlloc Attention (MALA), a fused attention primitive that keeps QK score discovery over every legal causal interaction and uses normalized online-softmax contribution to decide which tiles execute post-score computation; under exactly matched post-score work at 8K it reaches 0.0188% mean omitted mass versus 0.0182% for a per-instance reference-mass oracle, holds low output and gradient errors from 1K to 32K, reaches 89.67% associative-recall accuracy at 8K versus 89.97% for FullAttn, reduces training forward and backward latency by 2.2x and 3.0x and decoding latency by 1.6x at 128K with tensor parallelism on 8 GPUs, tracks FullAttn perplexity from 0.6B to 14B while cutting total training FLOPs by 23.
arXiv The work introduces AI Night-Scientist, a cognitive-science-grounded agentic framework that represents creativity at the action, process, and outcome levels and uses GRPO to train Qwen3-8B/14B models for long-horizon research proposal generation, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model, improving predicted citation impact by up to 32.0 percentage points and originality by 66.2 points, and showing that these gains cannot be reproduced by raising decoding temperature alone, with semantic guidance about what kind of creativity to pursue being critical.
The work introduces AI Night-Scientist, a cognitive-science-grounded agentic framework that represents creativity at the action, process, and outcome levels and uses GRPO to train Qwen3-8B/14B models for long-horizon research proposal generation, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model, improving predicted citation impact by up to 32.0 percentage points and originality by 66.2 points, and showing that these gains cannot be reproduced by raising decoding temperature alone, with semantic guidance about what kind of creativity to pursue being critical.
The work introduces AI Night-Scientist, a cognitive-science-grounded agentic framework that represents creativity at the action, process, and outcome levels and uses GRPO to train Qwen3-8B/14B models for long-horizon research proposal generation, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model, improving predicted citation impact by up to 32.0 percentage points and originality by 66.2 points, and showing that these gains cannot be reproduced by raising decoding temperature alone, with semantic guidance about what kind of creativity to pursue being critical.
The work introduces AI Night-Scientist, a cognitive-science-grounded agentic framework that represents creativity at the action, process, and outcome levels and uses GRPO to train Qwen3-8B/14B models for long-horizon research proposal generation, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model, improving predicted citation impact by up to 32.0 percentage points and originality by 66.2 points, and showing that these gains cannot be reproduced by raising decoding temperature alone, with semantic guidance about what kind of creativity to pursue being critical.
arXiv The work proposes Recursive Harness Distillation (RHD): a strong agent (GPT-6 Astra) first interacts with a frozen vision-language-action model and distills effective intervention strategies into a reusable playbook, then recursively revises that playbook using the light agent's (GPT-5.6 Luna) execution feedback, improving manipulation without updating any model parameters—real-world success rises from 37.3% to 64.0%, the light agent with the playbook reaches 66.7% on SimplerEnv Bridge versus 41.7% for GR00T alone, and the same playbook also lifts the strong agent to 79.2%.
The work proposes Recursive Harness Distillation (RHD): a strong agent (GPT-6 Astra) first interacts with a frozen vision-language-action model and distills effective intervention strategies into a reusable playbook, then recursively revises that playbook using the light agent's (GPT-5.6 Luna) execution feedback, improving manipulation without updating any model parameters—real-world success rises from 37.3% to 64.0%, the light agent with the playbook reaches 66.7% on SimplerEnv Bridge versus 41.7% for GR00T alone, and the same playbook also lifts the strong agent to 79.2%.
The work proposes Recursive Harness Distillation (RHD): a strong agent (GPT-6 Astra) first interacts with a frozen vision-language-action model and distills effective intervention strategies into a reusable playbook, then recursively revises that playbook using the light agent's (GPT-5.6 Luna) execution feedback, improving manipulation without updating any model parameters—real-world success rises from 37.3% to 64.0%, the light agent with the playbook reaches 66.7% on SimplerEnv Bridge versus 41.7% for GR00T alone, and the same playbook also lifts the strong agent to 79.2%.
The work proposes Recursive Harness Distillation (RHD): a strong agent (GPT-6 Astra) first interacts with a frozen vision-language-action model and distills effective intervention strategies into a reusable playbook, then recursively revises that playbook using the light agent's (GPT-5.6 Luna) execution feedback, improving manipulation without updating any model parameters—real-world success rises from 37.3% to 64.0%, the light agent with the playbook reaches 66.7% on SimplerEnv Bridge versus 41.7% for GR00T alone, and the same playbook also lifts the strong agent to 79.2%.
arXiv Using a matched ladder of 11 sparse MoE decoders (1.1B–44B total, 71M–2.4B active non-embedding parameters) with identical data mixture and optimization, the work fits separate text and multimodal scaling laws for encoder-free and encoder-based MLLMs and finds that removing the visual encoder shifts compute-optimal multimodal allocation toward larger models (model allocation exponent rising from about 0.464 to about 0.570) while leaving text nearly unchanged; encoder-free models underperform at small scales but their multimodal loss falls faster with compute, with extrapolation predicting an efficiency crossover on the order of 10^22 FLOPs (point estimate about 3.
Using a matched ladder of 11 sparse MoE decoders (1.1B–44B total, 71M–2.4B active non-embedding parameters) with identical data mixture and optimization, the work fits separate text and multimodal scaling laws for encoder-free and encoder-based MLLMs and finds that removing the visual encoder shifts compute-optimal multimodal allocation toward larger models (model allocation exponent rising from about 0.464 to about 0.570) while leaving text nearly unchanged; encoder-free models underperform at small scales but their multimodal loss falls faster with compute, with extrapolation predicting an efficiency crossover on the order of 10^22 FLOPs (point estimate about 3.
Using a matched ladder of 11 sparse MoE decoders (1.1B–44B total, 71M–2.4B active non-embedding parameters) with identical data mixture and optimization, the work fits separate text and multimodal scaling laws for encoder-free and encoder-based MLLMs and finds that removing the visual encoder shifts compute-optimal multimodal allocation toward larger models (model allocation exponent rising from about 0.464 to about 0.570) while leaving text nearly unchanged; encoder-free models underperform at small scales but their multimodal loss falls faster with compute, with extrapolation predicting an efficiency crossover on the order of 10^22 FLOPs (point estimate about 3.
Using a matched ladder of 11 sparse MoE decoders (1.1B–44B total, 71M–2.4B active non-embedding parameters) with identical data mixture and optimization, the work fits separate text and multimodal scaling laws for encoder-free and encoder-based MLLMs and finds that removing the visual encoder shifts compute-optimal multimodal allocation toward larger models (model allocation exponent rising from about 0.464 to about 0.570) while leaving text nearly unchanged; encoder-free models underperform at small scales but their multimodal loss falls faster with compute, with extrapolation predicting an efficiency crossover on the order of 10^22 FLOPs (point estimate about 3.
arXiv The work introduces G²PTQ, a post-training quantization framework that combines first- and second-order information under a globally supervised, block-wise objective: it recomputes gradient and Hessian estimates before quantizing each Transformer block to avoid staleness and uses trust-region scaling to bound the exact first-order compensation step, achieving lower KL divergence and higher downstream QA accuracy than baselines such as GPTQ, GuidedQuant, and GPTAQ across 13 dense and 2 MoE models from 0.6B to 125B parameters at 2/3/4-bit weight and 4-bit weight-activation settings.
The work introduces G²PTQ, a post-training quantization framework that combines first- and second-order information under a globally supervised, block-wise objective: it recomputes gradient and Hessian estimates before quantizing each Transformer block to avoid staleness and uses trust-region scaling to bound the exact first-order compensation step, achieving lower KL divergence and higher downstream QA accuracy than baselines such as GPTQ, GuidedQuant, and GPTAQ across 13 dense and 2 MoE models from 0.6B to 125B parameters at 2/3/4-bit weight and 4-bit weight-activation settings.
The work introduces G²PTQ, a post-training quantization framework that combines first- and second-order information under a globally supervised, block-wise objective: it recomputes gradient and Hessian estimates before quantizing each Transformer block to avoid staleness and uses trust-region scaling to bound the exact first-order compensation step, achieving lower KL divergence and higher downstream QA accuracy than baselines such as GPTQ, GuidedQuant, and GPTAQ across 13 dense and 2 MoE models from 0.6B to 125B parameters at 2/3/4-bit weight and 4-bit weight-activation settings.
The work introduces G²PTQ, a post-training quantization framework that combines first- and second-order information under a globally supervised, block-wise objective: it recomputes gradient and Hessian estimates before quantizing each Transformer block to avoid staleness and uses trust-region scaling to bound the exact first-order compensation step, achieving lower KL divergence and higher downstream QA accuracy than baselines such as GPTQ, GuidedQuant, and GPTAQ across 13 dense and 2 MoE models from 0.6B to 125B parameters at 2/3/4-bit weight and 4-bit weight-activation settings.
arXiv The work introduces CoWindow Attention (CoWA), in which all KV heads share near-diagonal and prefix-sink windows while complementary long-range windows partition the remaining history across heads, making full causal coverage a collective property of the head ensemble; in a window-matched 8K ablation CoWA reaches 89.73% versus 89.97% for FullAttn while duplicated long-range windows perform substantially worse, a 128K tensor-parallel operator benchmark reduces training forward and backward latency by 7.4x and 8.6x and decoding latency by 3.0x, and scaling-law training from 0.6B to 14B tracks FullAttn perplexity while cutting total training FLOPs by 28.5% in the 32K long-context stage.
The work introduces CoWindow Attention (CoWA), in which all KV heads share near-diagonal and prefix-sink windows while complementary long-range windows partition the remaining history across heads, making full causal coverage a collective property of the head ensemble; in a window-matched 8K ablation CoWA reaches 89.73% versus 89.97% for FullAttn while duplicated long-range windows perform substantially worse, a 128K tensor-parallel operator benchmark reduces training forward and backward latency by 7.4x and 8.6x and decoding latency by 3.0x, and scaling-law training from 0.6B to 14B tracks FullAttn perplexity while cutting total training FLOPs by 28.5% in the 32K long-context stage.
The work introduces CoWindow Attention (CoWA), in which all KV heads share near-diagonal and prefix-sink windows while complementary long-range windows partition the remaining history across heads, making full causal coverage a collective property of the head ensemble; in a window-matched 8K ablation CoWA reaches 89.73% versus 89.97% for FullAttn while duplicated long-range windows perform substantially worse, a 128K tensor-parallel operator benchmark reduces training forward and backward latency by 7.4x and 8.6x and decoding latency by 3.0x, and scaling-law training from 0.6B to 14B tracks FullAttn perplexity while cutting total training FLOPs by 28.5% in the 32K long-context stage.
The work introduces CoWindow Attention (CoWA), in which all KV heads share near-diagonal and prefix-sink windows while complementary long-range windows partition the remaining history across heads, making full causal coverage a collective property of the head ensemble; in a window-matched 8K ablation CoWA reaches 89.73% versus 89.97% for FullAttn while duplicated long-range windows perform substantially worse, a 128K tensor-parallel operator benchmark reduces training forward and backward latency by 7.4x and 8.6x and decoding latency by 3.0x, and scaling-law training from 0.6B to 14B tracks FullAttn perplexity while cutting total training FLOPs by 28.5% in the 32K long-context stage.